Skip to main content

How cache and fast-mode pricing is tracked

Every 1M-context Claude model bills the full window at standard rates — there is no long-context tier. The pricing modifiers Anthropic does apply are all reported in the response usage, and Timbal’s collector emits one unit per disjoint bucket: output_tokens_details.thinking_tokens is a subset of output_tokens and is not emitted as a separate unit (it would be billed twice). On Opus 5 / Opus 4.8, usage.speed == "fast" (fast mode, 2x) suffixes every token unit with _fast, e.g. anthropic/claude-opus-5:output_tokens_fast; pass model_params={"extra_body": {"speed": "fast"}, "extra_headers": {"anthropic-beta": "fast-mode-2026-02-01"}} to opt in. If a stream is interrupted before the final message_delta, the prompt-side units are still recorded from message_start (Anthropic charges them regardless); output is left unbilled.

Latest

claude-fable-5-1

Reasoning · Speedanthropic/claude-fable-5-1Anthropic’s most capable generally available model. Successor to Fable 5 for long-running agentic coding, knowledge work, and research, with fewer false-positive safety refusals in multi-turn agent loops. Same price as Fable 5; cache reads are 75% cheaper.
  • $10 / $50 per 1M tokens (input / output); cache reads $0.25 (0.025x input); batch $5 / $25
  • 1M context
  • 128K max output
  • Text, Image input
  • Adaptive thinking (always on)
  • Web search
  • Knowledge cutoff Jun 2026
  • tool_choice any / tool not supported (400); requires 30-day data retention (no ZDR)

claude-opus-5

Reasoning · Speedanthropic/claude-opus-5Latest Opus and the recommended default for complex agentic coding and enterprise work — near-Fable-5 intelligence at half the price, with large gains in deep reasoning, long-horizon tool loops, and test-time compute scaling.
  • $5 / $25 per 1M tokens (input / output; Fast mode $10 / $50 at ~2.5x speed)
  • 1M context (default and maximum)
  • 128K max output
  • Text, Image input
  • Adaptive thinking (on by default)
  • Web search
  • Knowledge cutoff May 2026

claude-sonnet-5

Reasoning · Speedanthropic/claude-sonnet-5Current-generation Sonnet — near-Opus intelligence at Sonnet pricing for coding, agents, and everyday professional work. Drop-in replacement for Sonnet 4.6.
  • $2 / $10 per 1M tokens (input / output)
  • 1M context
  • 128K max output
  • Text, Image input
  • Adaptive thinking
  • Web search
  • Knowledge cutoff Jan 2026

claude-opus-4-8

Reasoning · Speedanthropic/claude-opus-4-8Previous Opus generation for agentic coding, long-running tasks, and complex reasoning. Same price as Opus 5 — migrate unless you depend on thinking being off by default.
  • $5 / $25 per 1M tokens (input / output)
  • 1M context
  • 128K max output
  • Text, Image input
  • Adaptive thinking
  • Web search
  • Knowledge cutoff Jan 2026

claude-opus-4-7

Reasoning · Speedanthropic/claude-opus-4-7Frontier Opus for agentic coding and complex reasoning: stronger software engineering, verification, and high-resolution vision than 4.6, with adaptive thinking and a 1M-token context window at standard per-token rates.
  • $5 / $25 per 1M tokens (input / output)
  • 1M context
  • 128K max output
  • Text, Image input
  • Adaptive thinking
  • Web search
  • Knowledge cutoff Jan 2026

claude-opus-4-6

Reasoning · Speedanthropic/claude-opus-4-6Anthropic’s most intelligent model for building agents and coding. Exceptional at planning, code review, debugging, and operating reliably within large codebases.
  • $5 / $25
  • 1M context
  • 128K max output
  • Text, Image input
  • Extended thinking
  • Web search
  • Knowledge cutoff May 2025

claude-haiku-4-5

Reasoning · Speedanthropic/claude-haiku-4-5The fastest Claude model with near-frontier intelligence for office files, strategy planning, and business analysis.
  • $1 / $5
  • 200K context
  • 64K max output
  • Text, Image input
  • Thinking
  • Web search
  • Knowledge cutoff Feb 2025

Legacy

claude-fable-5

Reasoning · Speedanthropic/claude-fable-5Previous Fable generation for long-horizon agentic work, complex reasoning, and ambitious multi-day coding tasks. Superseded by Fable 5.1 at the same price.
  • $10 / $50 per 1M tokens (input / output)
  • 1M context
  • 128K max output
  • Text, Image input
  • Adaptive thinking (always on)
  • Web search
  • Knowledge cutoff Jan 2026

claude-opus-4-5

Reasoning · Speedanthropic/claude-opus-4-5Previous flagship Opus model with top-tier reasoning and coding capabilities.
  • $5 / $25
  • 200K context
  • 64K max output
  • Text, Image input
  • Extended thinking
  • Knowledge cutoff May 2025

claude-opus-4-1

Reasoning · Speedanthropic/claude-opus-4-1Earlier Opus variant with strong reasoning, higher pricing tier.
  • $15 / $75
  • 200K context
  • 32K max output
  • Text, Image input
  • Extended thinking
  • Knowledge cutoff Jan 2025
  • Deprecated — retires August 5, 2026

claude-sonnet-4-6

Reasoning · Speedanthropic/claude-sonnet-4-6Previous Sonnet generation with strong speed/intelligence balance and computer use skills. Superseded by Sonnet 5.
  • $3 / $15
  • 1M context
  • 64K max output
  • Text, Image input
  • Thinking
  • Knowledge cutoff Aug 2025

claude-sonnet-4-5

Reasoning · Speedanthropic/claude-sonnet-4-5Previous Sonnet generation with strong all-around performance and 1M beta context support.
  • $3 / $15
  • 200K context (1M beta)
  • 64K max output
  • Text, Image input
  • Thinking
  • Knowledge cutoff Jan 2025