Semi Analysis: n-gram, AgentX, TPU v7 and Vera Rubin in the race for fast tokens
In the Semi Analysis Weekly podcast, the InferenceX team covered hybrid n-gram architectures, the AgentX agentic-coding benchmark, Google's TPU v7 and Nvidia's Vera Rubin. TPU v7 is Google's first externally sold accelerator with a TCO of about $1.21 per million tokens, while Vera Rubin triples inference throughput via HBM4.
- Hybrid MoE and n-gram layers reach the lowest loss in DeepSeek tests
- AgentX shows cache hit rates above 98%, DRAM offload adds 20–30%
- TPU v7 is Google's first publicly sold accelerator, TCO around $1.21
- Vera Rubin is 3x faster than GB300 at 100 tokens/s thanks to HBM4
Read next
Hardware