Forlinx launches 20-TOPS M.2 AI accelerator for local LLM inference
Forlinx has listed an M.2 AI accelerator card built on Rockchip's RK1820 and RK1828 coprocessors, delivering 20 TOPS of INT8 performance and up to 5GB of stacked DRAM. The M.2 2280 module targets local LLM, VLM and computer vision inference on Linux and Android, and a four-card PCIe cascade on the OK3588-C board can run 27B-31B parameter models at roughly 40W.
- NPU rated at 20 TOPS INT8, supports INT4/INT8/INT16/FP8/FP16/BF16
- RK1820 packs 2.5GB DRAM, RK1828 packs 5GB in 3D-stacked package
- Single PCIe 2.1 lane, hosts include RK3568, RK3572, RK3576, RK3588
- Four-card cascade: Qwen2.5-3B at 102 tok/s, 27B up to 13 tok/s
Read next
Hardware