Skip to main content
Source: Fireworks model docs. All model IDs use the prefix fireworks/accounts/fireworks/models/. Only serverless models are listed here — other Fireworks models require a dedicated deployment.

All Models

deepseek-v4-pro

Reasoning · Speedfireworks/accounts/fireworks/models/deepseek-v4-proDeepSeek V4 Pro on Fireworks serverless: frontier open MoE for coding, reasoning, and up to ~1M token context with function calling.
  • $1.74 / $3.48
  • 1M context
  • Text input

deepseek-v4-flash

Reasoning · Speedfireworks/accounts/fireworks/models/deepseek-v4-flashDeepSeek V4 Flash on Fireworks serverless: fast, cost-efficient MoE with near-Pro reasoning at 1M context.
  • $0.14 / $0.28
  • 1M context
  • Text input

qwen3p7-plus

Reasoning · Speedfireworks/accounts/fireworks/models/qwen3p7-plusAlibaba Qwen 3.7 Plus on Fireworks serverless: multimodal flagship with strong agentic and coding benchmarks.
  • $0.40 / $1.60
  • 262K context
  • Text, Image input

qwen3p6-plus

Reasoning · Speedfireworks/accounts/fireworks/models/qwen3p6-plusQwen 3.6 multimodal plus tier on Fireworks for vision-language and general agent tasks.
  • $0.50 / $3.00
  • 262K context
  • Text, Image input
  • Serverless deprecated — prefer qwen3p7-plus

kimi-k2p6

Reasoning · Speedfireworks/accounts/fireworks/models/kimi-k2p6Moonshot Kimi K2.6 on Fireworks: multimodal MoE for high-quality tool use and long-context workloads.
  • $0.95 / $4.00
  • 262K context
  • Text, Image input

kimi-k2p7-code

Reasoning · Speedfireworks/accounts/fireworks/models/kimi-k2p7-codeMoonshot Kimi K2.7 Code on Fireworks serverless for long-context coding agents with thinking mode.
  • $0.95 / $4.00
  • 262K context
  • Text, Image input
  • Thinking

kimi-k2p5

Reasoning · Speedfireworks/accounts/fireworks/models/kimi-k2p5Moonshot Kimi K2.5 on Fireworks serverless: multimodal 1T-parameter MoE with strong agentic tool use.
  • $0.60 / $3.00
  • 256K context
  • Text, Image input
  • Serverless deprecated — prefer kimi-k2p6

glm-5p1

Reasoning · Speedfireworks/accounts/fireworks/models/glm-5p1Z.ai GLM-5.1 on Fireworks serverless: post-training upgrade with stronger coding, reasoning, and agentic tool use.
  • $1.40 / $4.40
  • 203K context
  • Text input

glm-5p2

Reasoning · Speedfireworks/accounts/fireworks/models/glm-5p2Z.ai GLM-5.2 on Fireworks serverless: 1M-token context, strongest open coding model with prompt caching (cached input $0.14 / 1M).
  • $1.40 / $4.40
  • 1M context
  • Text input

minimax-m2p5

Reasoning · Speedfireworks/accounts/fireworks/models/minimax-m2p5MiniMax M2.5 MoE (230B total, 10B active) with SOTA coding and agentic tool use on Fireworks serverless.
  • $0.30 / $1.20
  • 200K context
  • Text input
  • Serverless deprecated — prefer minimax-m2p7 or minimax-m3

minimax-m2p7

Reasoning · Speedfireworks/accounts/fireworks/models/minimax-m2p7MiniMax M2.7 MoE on Fireworks serverless: improved agent harnesses, complex skills, and dynamic tool search.
  • $0.30 / $1.20
  • 196K context
  • Text input

minimax-m3

Reasoning · Speedfireworks/accounts/fireworks/models/minimax-m3MiniMax M3 multimodal MoE on Fireworks serverless for chat, agents, and long-context workloads.
  • $0.30 / $1.20
  • 512K context
  • Text, Image input

gpt-oss-120b

Reasoning · Speedfireworks/accounts/fireworks/models/gpt-oss-120bOpenAI’s open-weight 120B MoE achieving near-parity with o4-mini on reasoning benchmarks. Apache 2.0.
  • $0.15 / $0.60
  • 128K context
  • Text input

gpt-oss-20b

Reasoning · Speedfireworks/accounts/fireworks/models/gpt-oss-20bOpenAI’s compact 20B MoE similar to o3-mini, running on edge devices with 16GB memory. Apache 2.0.
  • $0.07 / $0.30
  • 128K context
  • Text input