Ito: speech synthesis in 4.89 MB per voice on ESP32-S3
Lokutor released the Ito inference engine, ESP32-S3 firmware and two English voice models: 4.05M parameters and 4.89 MB of weights per voice, producing 24 kHz speech with integer arithmetic. The firmware is verified in Espressif's QEMU emulator, but nothing has run on a physical board yet.
- 4.05M parameters and 4.89 MB of weights per English voice
- 24 kHz synthesis with integer math on a 240 MHz chip without an NPU
- QEMU output is bit-identical to the host engine's PCM
- Estimated RTF 0.53–1.27; first audio 137–318 ms
Read next
AI