> ## Documentation Index
> Fetch the complete documentation index at: https://docs.timbal.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# TogetherAI

> Open-source models via TogetherAI inference with specs, pricing, and capabilities

<Note>Source: [TogetherAI model docs](https://docs.together.ai/docs/chat-models). All model IDs use the prefix `togetherai/`. Models with a dedicated-only warning are not serverless — you must create and start a dedicated endpoint first.</Note>

## Meta LLaMA

<CardGroup cols={2}>
  <Card title="Llama-3.3-70B-Instruct-Turbo">
    <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> Reasoning · <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> Speed

    `togetherai/meta-llama/Llama-3.3-70B-Instruct-Turbo`

    Multilingual instruction-tuned model with 70B parameters, delivering enhanced performance relative to Llama 3.1 70B and matching Llama 3.2 90B on text-only tasks.

    * \$0.88 / \$0.88
    * <Icon icon="window-maximize" size={14} /> 128K context
    * <Icon icon="keyboard" size={14} /> Text input
    * <Icon icon="calendar" size={14} /> Knowledge cutoff Dec 2023
  </Card>
</CardGroup>

***

## Qwen

<CardGroup cols={2}>
  <Card title="Qwen/Qwen3.5-397B-A17B">
    <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> Reasoning · <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> Speed

    `togetherai/Qwen/Qwen3.5-397B-A17B`

    Multimodal foundation model with 397B total parameters (17B active) featuring a Hybrid MoE architecture with early fusion vision-language training. State-of-the-art across chat, RAG, vision-language, and agentic workflows.

    * \$0.30 / \$1.20
    * <Icon icon="window-maximize" size={14} /> 262K context
    * <Icon icon="keyboard" size={14} /> Text, Image input
    * <Icon icon="brain" size={14} /> Hybrid thinking
    * <Icon icon="calendar" size={14} /> Knowledge cutoff \~2025
  </Card>

  <Card title="Qwen3-235B-A22B-Instruct-2507-tput">
    <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> Reasoning · <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> Speed

    `togetherai/Qwen/Qwen3-235B-A22B-Instruct-2507-tput`

    MoE model with 235B total parameters (22B active) in non-thinking mode, optimized for throughput. Supports multilingual dialogue across 100+ languages.

    * \$0.20 / \$0.60
    * <Icon icon="window-maximize" size={14} /> 262K context
    * <Icon icon="keyboard" size={14} /> Text input
    * <Icon icon="calendar" size={14} /> Knowledge cutoff \~early 2025
  </Card>

  <Card title="Qwen3-Coder-480B-A35B-Instruct-FP8">
    <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> Reasoning · <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> Speed

    `togetherai/Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8`

    Qwen's most agentic code model, a 480B-parameter MoE (35B active) achieving results comparable to Claude Sonnet on agentic coding, browser-use, and repository-scale tasks.

    <Warning>Dedicated only — create and start a [dedicated endpoint](https://api.together.ai/models/Qwen/Qwen3-Coder-480B-A35B-Instruct-FP8) before use.</Warning>

    * \$0.22 / \$1
    * <Icon icon="window-maximize" size={14} /> 262K context
    * <Icon icon="keyboard" size={14} /> Text input
  </Card>

  <Card title="Qwen3-Coder-Next-FP8">
    <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> Reasoning · <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> Speed

    `togetherai/Qwen/Qwen3-Coder-Next-FP8`

    Next-generation coding model with hybrid thinking mode for adaptive reasoning depth.

    <Warning>Dedicated only — create and start a [dedicated endpoint](https://api.together.ai/models/Qwen/Qwen3-Coder-Next-FP8) before use.</Warning>

    * \$0.50 / \$1.20
    * <Icon icon="window-maximize" size={14} /> 256K context
    * <Icon icon="keyboard" size={14} /> Text input
    * <Icon icon="brain" size={14} /> Hybrid thinking
  </Card>

  <Card title="Qwen3-Next-80B-A3B-Instruct">
    <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> Reasoning · <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> Speed

    `togetherai/Qwen/Qwen3-Next-80B-A3B-Instruct`

    First model in the Qwen3-Next series with 80B total parameters (3.9B active), featuring hybrid attention. Matches Qwen3-235B performance while using less than 10% training cost.

    <Warning>Dedicated only — create and start a [dedicated endpoint](https://api.together.ai/models/Qwen/Qwen3-Next-80B-A3B-Instruct) before use.</Warning>

    * \$0.15 / \$1.50
    * <Icon icon="window-maximize" size={14} /> 262K context
    * <Icon icon="keyboard" size={14} /> Text input
    * <Icon icon="brain" size={14} /> Hybrid thinking
  </Card>

  <Card title="Qwen2.5-7B-Instruct-Turbo">
    <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> Reasoning · <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> Speed

    `togetherai/Qwen/Qwen2.5-7B-Instruct-Turbo`

    Part of the Qwen2.5 family with 7B parameters, featuring improvements in coding, mathematics, instruction following, and structured data understanding.

    * \$0.30 / \$1.20
    * <Icon icon="window-maximize" size={14} /> 128K context
    * <Icon icon="keyboard" size={14} /> Text input
    * <Icon icon="calendar" size={14} /> Knowledge cutoff \~Oct 2023
  </Card>
</CardGroup>

***

## DeepSeek

<CardGroup cols={2}>
  <Card title="DeepSeek-V3.1">
    <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> Reasoning · <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> Speed

    `togetherai/deepseek-ai/DeepSeek-V3.1`

    Hybrid model supporting both thinking and non-thinking modes. Features significantly improved tool usage and agent task performance, with quality comparable to DeepSeek-R1-0528 in thinking mode.

    <Warning>Dedicated only — create and start a [dedicated endpoint](https://api.together.ai/models/deepseek-ai/DeepSeek-V3.1) before use.</Warning>

    * \$0.60 / \$1.70
    * <Icon icon="window-maximize" size={14} /> 128K context
    * <Icon icon="keyboard" size={14} /> Text input
    * <Icon icon="brain" size={14} /> Hybrid thinking
    * <Icon icon="calendar" size={14} /> Knowledge cutoff \~mid 2025
  </Card>

  <Card title="DeepSeek-V4-Pro">
    <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> Reasoning · <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> Speed

    `togetherai/deepseek-ai/DeepSeek-V4-Pro`

    DeepSeek V4 Pro on Together serverless: hybrid attention, up to 512K context, strong coding and agent benchmarks.

    * \$1.74 / \$3.48
    * <Icon icon="window-maximize" size={14} /> 512K context
    * <Icon icon="keyboard" size={14} /> Text input
    * <Icon icon="brain" size={14} /> Hybrid thinking
    * <Icon icon="calendar" size={14} /> Knowledge cutoff \~2025
  </Card>
</CardGroup>

***

## Kimi / MiniMax / GLM / Other

<CardGroup cols={2}>
  <Card title="Kimi-K2.6">
    <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> Reasoning · <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> Speed

    `togetherai/moonshotai/Kimi-K2.6`

    Moonshot Kimi K2.6 on Together serverless: 1T-scale MoE with tool calling and JSON mode for agentic and multimodal workloads.

    * \$1.20 / \$4.50
    * <Icon icon="window-maximize" size={14} /> 262K context
    * <Icon icon="keyboard" size={14} /> Text, Image input
    * <Icon icon="calendar" size={14} /> Knowledge cutoff \~2025
  </Card>

  <Card title="Kimi-K2.7-Code">
    <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> Reasoning · <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> Speed

    `togetherai/moonshotai/Kimi-K2.7-Code`

    Moonshot Kimi K2.7 Code on Together for long-context programming agents with thinking mode.

    * \$0.95 / \$4.00
    * <Icon icon="window-maximize" size={14} /> 262K context
    * <Icon icon="keyboard" size={14} /> Text, Image input
    * <Icon icon="brain" size={14} /> Thinking
  </Card>

  <Card title="MiniMax-M2.7">
    <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> Reasoning · <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> Speed

    `togetherai/MiniMaxAI/MiniMax-M2.7`

    MiniMax successor MoE (\~229B) with improved coding and agentic tool use, JSON mode, and prompt caching on Together serverless.

    * \$0.30 / \$1.20
    * <Icon icon="window-maximize" size={14} /> 203K context
    * <Icon icon="keyboard" size={14} /> Text input
  </Card>

  <Card title="MiniMax-M3">
    <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> Reasoning · <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> Speed

    `togetherai/MiniMaxAI/MiniMax-M3`

    MiniMax M3 multimodal model on Together for chat, agents, and long-context workloads.

    * \$0.30 / \$1.20
    * <Icon icon="window-maximize" size={14} /> 512K context
    * <Icon icon="keyboard" size={14} /> Text, Image input
  </Card>

  <Card title="GLM-5">
    <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> Reasoning · <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> Speed

    `togetherai/zai-org/GLM-5`

    Zhipu AI's fifth-generation model with \~745B parameters in a MoE architecture (44B active), designed for complex system engineering and long-range agent tasks. Trained entirely on Huawei Ascend chips.

    * \$1 / \$3.20
    * <Icon icon="window-maximize" size={14} /> 200K context
    * <Icon icon="keyboard" size={14} /> Text input
    * <Icon icon="brain" size={14} /> Thinking
    * <Icon icon="calendar" size={14} /> Knowledge cutoff late 2025
  </Card>

  <Card title="GLM-5.1">
    <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> Reasoning · <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> Speed

    `togetherai/zai-org/GLM-5.1`

    Z.ai post-training upgrade to GLM-5: 754B MoE (40B active), 200K context, thinking mode, tool calling, and stronger coding via RL.

    * \$1.40 / \$4.40
    * <Icon icon="window-maximize" size={14} /> 200K context
    * <Icon icon="keyboard" size={14} /> Text input
    * <Icon icon="brain" size={14} /> Thinking
    * <Icon icon="calendar" size={14} /> Knowledge cutoff late 2025
  </Card>

  <Card title="GLM-5.2">
    <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> Reasoning · <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> Speed

    `togetherai/zai-org/GLM-5.2`

    Z.ai GLM-5.2 flagship on Together for coding, reasoning, and agentic tool use.

    * \$1.40 / \$4.40
    * <Icon icon="window-maximize" size={14} /> 200K context
    * <Icon icon="keyboard" size={14} /> Text input
    * <Icon icon="brain" size={14} /> Thinking
  </Card>

  <Card title="GLM-4.7">
    <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> Reasoning · <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> Speed

    `togetherai/zai-org/GLM-4.7`

    Zhipu AI's foundation model with \~400B parameters and 200K context, designed for real-world development environments with strong coding, reasoning, and agentic capabilities.

    <Warning>Dedicated only — create and start a [dedicated endpoint](https://api.together.ai/models/zai-org/GLM-4.7) before use.</Warning>

    * \$0.45 / \$2
    * <Icon icon="window-maximize" size={14} /> 200K context
    * <Icon icon="keyboard" size={14} /> Text input
    * <Icon icon="brain" size={14} /> Thinking
    * <Icon icon="calendar" size={14} /> Knowledge cutoff \~mid 2024
  </Card>

  <Card title="gpt-oss-120b">
    <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> Reasoning · <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> Speed

    `togetherai/openai/gpt-oss-120b`

    OpenAI's open-weight MoE model with 120B total parameters (5.1B active per token). Achieves near-parity with o4-mini on core reasoning benchmarks while running on a single 80GB GPU. Apache 2.0.

    * \$0.15 / \$0.60
    * <Icon icon="window-maximize" size={14} /> 128K context
    * <Icon icon="keyboard" size={14} /> Text input
    * <Icon icon="brain" size={14} /> Thinking
    * <Icon icon="calendar" size={14} /> Knowledge cutoff Jun 2024
  </Card>

  <Card title="gpt-oss-20b">
    <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> Reasoning · <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> Speed

    `togetherai/openai/gpt-oss-20b`

    OpenAI's compact 20B MoE delivering o3-mini-level results on Together serverless. Apache 2.0.

    * \$0.05 / \$0.20
    * <Icon icon="window-maximize" size={14} /> 128K context
    * <Icon icon="keyboard" size={14} /> Text input
    * <Icon icon="brain" size={14} /> Thinking
    * <Icon icon="calendar" size={14} /> Knowledge cutoff Jun 2024
  </Card>

  <Card title="gemma-3n-E4B-it">
    <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> Reasoning · <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> Speed

    `togetherai/google/gemma-3n-E4B-it`

    Google's on-device multimodal model with 8B raw parameters but an effective 4B memory footprint. First sub-10B model to exceed 1300 on LMArena, running with as little as 3GB of memory.

    * \$0.02 / \$0.04
    * <Icon icon="window-maximize" size={14} /> 32K context
    * <Icon icon="keyboard" size={14} /> Text, Image, Audio, Video input
  </Card>

  <Card title="gemma-3-27b-it">
    <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> Reasoning · <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> Speed

    `togetherai/google/gemma-3-27b-it`

    Google's multimodal open model with 27B parameters, built from the same technology as Gemini 2.0. Supports 128K context, 140+ languages, and runs on a single GPU/TPU.

    <Warning>Dedicated only — create and start a [dedicated endpoint](https://api.together.ai/models/google/gemma-3-27b-it) before use.</Warning>

    * \~\$0.10 / \~\$0.10
    * <Icon icon="window-maximize" size={14} /> 128K context
    * <Icon icon="keyboard" size={14} /> Text, Image input
    * <Icon icon="calendar" size={14} /> Knowledge cutoff Aug 2024
  </Card>

  <Card title="cogito-v2-1-671b">
    <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> Reasoning · <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> Speed

    `togetherai/deepcogito/cogito-v2-1-671b`

    DeepCogito's MoE model with 671B total parameters (37B active), trained via a novel process supervision approach that guides reasoning chains. Competitive with frontier closed models while using fewer tokens.

    * \$1.25 / \$1.25
    * <Icon icon="window-maximize" size={14} /> 128K context
    * <Icon icon="keyboard" size={14} /> Text input
    * <Icon icon="brain" size={14} /> Thinking
  </Card>

  <Card title="Mistral-Small-24B-Instruct-2501">
    <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> <Icon icon="lightbulb" size={14} /> Reasoning · <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> <Icon icon="bolt" size={14} /> Speed

    `togetherai/mistralai/Mistral-Small-24B-Instruct-2501`

    A 24B-parameter dense model setting new benchmarks in the sub-70B category, with native function calling, JSON output, and support for dozens of languages. Fits on a single RTX 4090.

    <Warning>Dedicated only — create and start a [dedicated endpoint](https://api.together.ai/models/mistralai/Mistral-Small-24B-Instruct-2501) before use.</Warning>

    * \$0.10 / \$0.30
    * <Icon icon="window-maximize" size={14} /> 32K context
    * <Icon icon="keyboard" size={14} /> Text input
    * <Icon icon="calendar" size={14} /> Knowledge cutoff Oct 2023
  </Card>
</CardGroup>
