<unit>_long_context usage keys so they can be billed at the higher rate — see the OpenAI page. All models support tool/function calling unless noted.
Anthropic
Claude Fable 5.1, Opus 5, Sonnet 5, and Haiku
BytePlus
Seed 2.0 and Seed 1.8 models
Cerebras
Wafer-scale inference at world-record token speeds
Fireworks
Open-source models via Fireworks
Gemini 3.6, 3.5, 3.1 and 2.5 series
Groq
Ultra-low latency via Groq LPU
Moonshot (Kimi)
Kimi K3 / K2.x via Moonshot’s OpenAI-compatible API
OpenAI
GPT-6 Astra, GPT-5, GPT-4, and o-series reasoning models
SambaNova
High-throughput inference on custom RDU hardware
TogetherAI
Open-source models via TogetherAI
xAI
Grok 4.6, Grok 4.5, and Grok 4.3
Xiaomi MiMo
MiMo V2 Pro, Omni, and Flash
Scoring
Each model is rated on two axes using a 1-5 scale:- Reasoning — depth of analytical and chain-of-thought capability
- Speed — relative latency and throughput for its class