Author: Evolving AI - Bewertung: 1x - Views:2
A Chinese startup just built its own “Google-style” TPU—and it might quietly rewrite the AI hardware game. In this video, we break down Ghana, a custom AI chip from Zhonghao Xinying, founded by a former Google TPU engineer. On paper, it’s already 1.5× faster than Nvidia’s A100 and up to 75% more efficient, even though it’s built on an older 12 nm-class process. Instead of trying to do everything like a GPU, Ghana is a pure tensor engine: fixed-function systolic arrays, FP16/BF16 only, big on-chip SRAM, and a compiler-driven execution model that keeps almost every transistor busy doing real AI math. We dig into how this “General-Purpose TPU” platform sacrifices flexibility to win on utilization, efficiency, and control, why it matters that the entire stack—chip, compiler, runtime, and interconnect—is built without Western IP, and how their Taize system can stitch together up to 1,024 Ghana chips into one cluster. Rather than chasing Nvidia’s latest Blackwell GPUs, Ghana is targeting the vast A100-class market in China, where export controls, power limits, and cost make sovereignty and performance-per-watt more important than raw peak FLOPs. Over time, every compiler update makes the same silicon faster, just like Google’s TPUs evolved from v1 to v4. The real question isn’t whether Ghana beats Nvidia today—it’s what happens after five years of iteration inside China’s own data centers. And if you want the real story behind the world’s fastest-moving tech and AI breakthroughs, make sure to like and subscribe to Evolving AI for daily coverage.
SOCIAL SHARE CARD GENERATOR