Skip to main content
Source: Groq model docs. All model IDs use the prefix groq/. Ultra-low latency inference via custom LPU hardware.

All Models

qwen/qwen3.8-27b

Reasoning · Speedgroq/qwen/qwen3.8-27bQwen3.8 27B multimodal with hybrid thinking/instruct modes, delivered at Groq’s ultra-low latency. Preview tier; successor to Qwen3.6 27B in this catalog (Groq still lists 3.6 as Preview).
  • $0.80 / $4.00
  • 131K context
  • Text, Image input
  • Hybrid thinking

openai/gpt-oss-120b

Reasoning · Speedgroq/openai/gpt-oss-120bOpenAI’s open-weight MoE model with 120B total parameters (5.1B active per token), running at Groq speeds. Near-parity with o4-mini on reasoning benchmarks. Apache 2.0.
  • $0.15 / $0.60
  • 128K context
  • Text input
  • Thinking

openai/gpt-oss-20b

Reasoning · Speedgroq/openai/gpt-oss-20bOpenAI’s compact open-weight MoE model with 20B total parameters (3.6B active), delivering results similar to o3-mini at Groq’s ultra-low latency. Apache 2.0.
  • $0.075 / $0.30
  • 128K context
  • Text input
  • Thinking