openTPU: AI agents designed the chip that runs their inference
openTPU is an AI inference accelerator whose SystemVerilog, ISA, compiler, simulator and host stack were written by AI agents. On an Inspur card with a Xilinx Kintex-7 xc7k480t, the hardware produces bit-identical tokens to its reference simulator; the code is open under Apache 2.0.
- openTPU's RTL, ISA, compiler and simulator were written by AI agents
- On the Kintex-7 card tokens and logits matched the simulator bit for bit
- A single 133.33 MHz bitstream serves every model
- The 4-bit format speeds up decode by 40–45% at a perplexity cost
Read next
Hardware