Skip to main content
Source: Cerebras model docs. All model IDs use the prefix cerebras/. Powered by Cerebras wafer-scale chips — the world’s largest AI accelerator — delivering up to 3000+ tokens/second.

All Models

gpt-oss-120b

Reasoning · Speedcerebras/gpt-oss-120bOpenAI’s open-weight MoE model with 120B total parameters (5.1B active per token), running at up to 3000 tokens/s on Cerebras wafer-scale hardware. Near-parity with o4-mini on reasoning benchmarks. Supports extended thinking. Apache 2.0.
  • $0.35 / $0.75
  • 128K context
  • Text input
  • Thinking
  • Knowledge cutoff Jun 2024