chiprook
← AI
AISeptember 15, 2026, 14:03

Compressing An 11B VLM To 2.7-bit Weights For Mobile CPUs

Graphcore Research and Arm introduced Llama-Mobile: the S3D8 format compresses Llama 3.2 Vision 11B weights from 21.3 GB to 3.7 GB (2.7 bits per parameter) and accelerates token generation on Arm CPUs. On Pixel 8a, the model achieves 3.8 tokens/s, retaining 66.1% versus 74.4% for bfloat16.

Compressing An 11B VLM To 2.7-bit Weights For Mobile CPUs
#ARM#Graphcore#Llama#Google
Read next
AI

Xiaomi debuts open-weight omnimodal MiMo-V2.6 Pro and Flash models

AI

Napster Developing AI Teacher Clones to Provide Personalized Homework Support

AI

Rogue AI agents push identity sector to build governance frameworks

AI

Google Gemini Notebook Brings Interactive Learning Overviews to All Users