Changelog

Speculative decoding on by default

Median latency reduced ~46% across chat models. No action required.

New models: added 8 open-weight models

Including new multilingual and code-specialized checkpoints.

Multi-region failover

Automatic routing to the nearest healthy region with transparent failover.

RAG collections GA

Managed retrieval is now generally available on all paid plans.