Speculative decoding on by default
Median latency reduced ~46% across chat models. No action required.
Median latency reduced ~46% across chat models. No action required.
Including new multilingual and code-specialized checkpoints.
Automatic routing to the nearest healthy region with transparent failover.
Managed retrieval is now generally available on all paid plans.