AWS launches Serverless AI Runtime for serverless inference
AWS unveiled its Serverless AI Runtime, a managed environment for running machine learning models without provisioning infrastructure. It supports models up to 70 billion parameters, is available in US East and Frankfurt, and charges per request; beta testers report 40–60% lower inference costs.
- Supports models up to 70B parameters in ONNX and PyTorch formats
- Per-request pricing with a free tier for prototyping
- Available in US East and EU (Frankfurt), Asia-Pacific next quarter
- Beta testers report 40–60% lower inference costs vs EC2
Read next
AI