Skip to main content
Start with Your first voice agent for a working conversation. This page is the reference for changing providers and resolving configuration after the basic lessons. Set Agent(voice_config=...) to configure sessions created by the server. Use a sparse dict or VoiceConfig so the fields you omit inherit operator defaults. The server applies environment defaults, then the agent’s declared settings, then allowed per-call client overrides. The client cannot supply credentials, arbitrary provider hosts, recording paths, or Python detector objects.

Transcription providers

Set stt_provider and optionally stt_model. These are adapter IDs and speech model IDs, separate from the provider/model IDs used by Agent.model.
  • elevenlabs: defaults to scribe_v2_realtime; requires ELEVENLABS_API_KEY.
  • deepgram or deepgram-flux: defaults to flux-general-multi; requires DEEPGRAM_API_KEY. Flux reports native end-of-turn events.
  • deepgram-nova: defaults to nova-3; requires DEEPGRAM_API_KEY. Select Nova explicitly or supply a nova-* model.
  • munsit: defaults to munsit-en-ar; requires MUNSIT_API_KEY. Reports native utterance boundaries.
  • openai: defaults to gpt-transcribe; requires OPENAI_API_KEY. Uses a transcription-only Realtime session with server VAD.
Set language to a provider-supported language hint or omit it for automatic detection. Provider-specific transcription options belong in stt_extra. Each adapter supports its own options; switching providers does not make an option portable.

Speech providers

Set tts_provider, tts_model, and voice for the output provider:
  • elevenlabs: defaults to eleven_flash_v2_5; requires ELEVENLABS_API_KEY and an accessible ElevenLabs voice ID. Supports incremental text through one streaming context per reply.
  • deepgram: defaults to aura-2-thalia-en; requires DEEPGRAM_API_KEY. The full Aura model ID identifies the voice and language.
  • fishaudio: defaults to s2.1-pro; requires FISH_API_KEY. Use a Fish reference voice ID for voice when selecting a reference voice.
  • munsit: defaults to faseeh-v1-preview; requires MUNSIT_API_KEY. The fallback voice is MUNSIT_VOICE_ID or ar-uae-male-1.
  • openai: defaults to gpt-4o-mini-tts and coral; requires OPENAI_API_KEY. Synthesizes complete text segments as they become ready.
Use the chosen provider’s voice ID. language controls transcription; it does not select a Deepgram Aura voice. Output style options belong in tts_extra. See Speech examples for detailed OpenAI and Deepgram settings. The default session audio is mono PCM16 little-endian at 16,000 Hz (sample_rate=16000, encoding="pcm_s16le"). Transports convert their wire formats where needed. OpenAI uses 24 kHz on the provider connection and requires the voice extra to resample other session rates.

Mix providers

With OPENAI_API_KEY configured, this selects OpenAI for transcription and speech while keeping agent.model as the conversation model:
This assignment replaces the agent’s declared config. Keep any greetings or other behavior fields you want to retain. The earlier lessons use a dict and add individual keys to preserve previous behavior; a VoiceConfig instance or callable needs its own update pattern.

Turn detection

See Turn-taking and interruptions for mode names, native provider boundaries, VAD endpointing, and a listening test.

Greetings and silence

See Greetings for fixed and generated openers, timing, and outbound behavior. Silence and timeouts covers idle prompts, hangup, and the voice turn deadline.

Tool-call fillers

See Spoken fillers to enable a waiting phrase, tune generation, or repeat it during a long tool call. Thinking sounds are a separate browser feature.

Operator defaults

The server reads TIMBAL_STT_PROVIDER, TIMBAL_STT_MODEL, TIMBAL_TTS_MODEL, ELEVENLABS_VOICE_ID (before TIMBAL_VOICE_ID), TIMBAL_VOICE_LANGUAGE, and TIMBAL_VOICE_GREETING. Use voice_config.tts_provider to select the output provider; there is no TIMBAL_TTS_PROVIDER reader in the server. TIMBAL_VOICE_FILLER=1 enables default fillers. The TIMBAL_VOICE_FILLER_MODEL, TIMBAL_VOICE_FILLER_SYSTEM_PROMPT, TIMBAL_VOICE_FILLER_DELAY_SECS, and TIMBAL_VOICE_FILLER_REPEAT_SECS variables customize them. See Recording calls and Background audio for their operator settings and transport limits.