Source: Fireworks model docs. All model IDs use the prefix
fireworks/accounts/fireworks/models/. Only serverless models are listed here — other Fireworks models require a dedicated deployment.All Models
deepseek-v4-flash-0731
Reasoning · Speed
fireworks/accounts/fireworks/models/deepseek-v4-flash-0731Official DeepSeek V4 Flash on Fireworks serverless (supersedes the April preview). Fast, cost-efficient MoE with near-Pro reasoning at 1M context.- $0.22 / $0.66
- 1M context
- Text input
qwen3p8-max
Reasoning · Speed
fireworks/accounts/fireworks/models/qwen3p8-maxAlibaba Qwen 3.8 Max on Fireworks serverless: flagship MoE for long-horizon coding and agentic work. Replacement for retired qwen3p7-plus.- $2.00 / $6.00
- 262K context
- Text input
kimi-k2p6
Reasoning · Speed
fireworks/accounts/fireworks/models/kimi-k2p6Moonshot Kimi K2.6 on Fireworks: multimodal MoE for high-quality tool use and long-context workloads.- $0.95 / $4.00
- 262K context
- Text, Image input
kimi-k2p7-code
Reasoning · Speed
fireworks/accounts/fireworks/models/kimi-k2p7-codeMoonshot Kimi K2.7 Code on Fireworks serverless for long-context coding agents with thinking mode.- $0.95 / $4.00
- 262K context
- Text, Image input
- Thinking
glm-5p2
Reasoning · Speed
fireworks/accounts/fireworks/models/glm-5p2Z.ai GLM-5.2 on Fireworks serverless: 1M-token context, strongest open coding model with prompt caching (cached input $0.14 / 1M).- $1.40 / $4.40
- 1M context
- Text input
minimax-m3
Reasoning · Speed
fireworks/accounts/fireworks/models/minimax-m3MiniMax M3 multimodal MoE on Fireworks serverless for chat, agents, and long-context workloads.- $0.30 / $1.20
- 512K context
- Text, Image input
gpt-oss-120b
Reasoning · Speed
fireworks/accounts/fireworks/models/gpt-oss-120bOpenAI’s open-weight 120B MoE achieving near-parity with o4-mini on reasoning benchmarks. Apache 2.0.- $0.15 / $0.60
- 128K context
- Text input