Agent(voice_config=...) to configure sessions created by the server. Use a sparse dict or VoiceConfig so the fields you omit inherit operator defaults. The server applies environment defaults, then the agent’s declared settings, then allowed per-call client overrides. The client cannot supply credentials, arbitrary provider hosts, recording paths, or Python detector objects.
Transcription providers
Setstt_provider and optionally stt_model. These are adapter IDs and speech model IDs, separate from the provider/model IDs used by Agent.model.
elevenlabs: defaults toscribe_v2_realtime; requiresELEVENLABS_API_KEY.deepgramordeepgram-flux: defaults toflux-general-multi; requiresDEEPGRAM_API_KEY. Flux reports native end-of-turn events.deepgram-nova: defaults tonova-3; requiresDEEPGRAM_API_KEY. Select Nova explicitly or supply anova-*model.munsit: defaults tomunsit-en-ar; requiresMUNSIT_API_KEY. Reports native utterance boundaries.openai: defaults togpt-transcribe; requiresOPENAI_API_KEY. Uses a transcription-only Realtime session with server VAD.
language to a provider-supported language hint or omit it for automatic detection. Provider-specific transcription options belong in stt_extra. Each adapter supports its own options; switching providers does not make an option portable.
Speech providers
Settts_provider, tts_model, and voice for the output provider:
elevenlabs: defaults toeleven_flash_v2_5; requiresELEVENLABS_API_KEYand an accessible ElevenLabs voice ID. Supports incremental text through one streaming context per reply.deepgram: defaults toaura-2-thalia-en; requiresDEEPGRAM_API_KEY. The full Aura model ID identifies the voice and language.fishaudio: defaults tos2.1-pro; requiresFISH_API_KEY. Use a Fish reference voice ID forvoicewhen selecting a reference voice.munsit: defaults tofaseeh-v1-preview; requiresMUNSIT_API_KEY. The fallback voice isMUNSIT_VOICE_IDorar-uae-male-1.openai: defaults togpt-4o-mini-ttsandcoral; requiresOPENAI_API_KEY. Synthesizes complete text segments as they become ready.
language controls transcription; it does not select a Deepgram Aura voice. Output style options belong in tts_extra. See Speech examples for detailed OpenAI and Deepgram settings.
The default session audio is mono PCM16 little-endian at 16,000 Hz (sample_rate=16000, encoding="pcm_s16le"). Transports convert their wire formats where needed. OpenAI uses 24 kHz on the provider connection and requires the voice extra to resample other session rates.
Mix providers
WithOPENAI_API_KEY configured, this selects OpenAI for transcription and speech while keeping agent.model as the conversation model:
VoiceConfig instance or callable needs its own update pattern.
Turn detection
See Turn-taking and interruptions for mode names, native provider boundaries, VAD endpointing, and a listening test.Greetings and silence
See Greetings for fixed and generated openers, timing, and outbound behavior. Silence and timeouts covers idle prompts, hangup, and the voice turn deadline.Tool-call fillers
See Spoken fillers to enable a waiting phrase, tune generation, or repeat it during a long tool call. Thinking sounds are a separate browser feature.Operator defaults
The server readsTIMBAL_STT_PROVIDER, TIMBAL_STT_MODEL, TIMBAL_TTS_MODEL, ELEVENLABS_VOICE_ID (before TIMBAL_VOICE_ID), TIMBAL_VOICE_LANGUAGE, and TIMBAL_VOICE_GREETING. Use voice_config.tts_provider to select the output provider; there is no TIMBAL_TTS_PROVIDER reader in the server.
TIMBAL_VOICE_FILLER=1 enables default fillers. The TIMBAL_VOICE_FILLER_MODEL, TIMBAL_VOICE_FILLER_SYSTEM_PROMPT, TIMBAL_VOICE_FILLER_DELAY_SECS, and TIMBAL_VOICE_FILLER_REPEAT_SECS variables customize them. See Recording calls and Background audio for their operator settings and transport limits.