Pinterest and Nvidia speed up visual AI search on Blackwell
Pinterest built a standardized multimodal AI infrastructure with Nvidia using Blackwell B200 GPUs and the open-source Dynamo framework. Precomputing visual representations cut overall latency by 7.3x and accelerated initial responses by roughly 85x, while Pinterest Assistant can handle up to 25x more visual context per request.
- The system runs across a fleet of roughly 14,000 Nvidia GPUs
- Latency dropped 7.3x and initial responses sped up about 85x
- Pinterest Assistant processes up to 25x more visual context per request
- Inference costs under 8% of closed proprietary models
Read next
AI