Agent runnable and install timbal[server,voice]; LiveKit also needs timbal[voice-livekit].
First get a conversation working in the quickstart. Then choose one transport below. Browser-local thinking sounds and background audio are not automatically included in a carrier’s audio stream; spoken fillers are part of the session’s speech output.
Browser playground
Open/voice on the running server. The playground exposes provider and model selection, microphone controls, transcripts, and turn metrics. GET /voice/meta describes the loaded runnable and its declared configuration. GET /voice_config returns {"voice_config": ...} with portable declared settings; it is not the fully resolved environment configuration.
WebSocket clients
Connect to/voice/ws and send a JSON configuration hello without a type field:
{"type": "audio", "data": "..."} in the same PCM format. Decode and queue it for playback. Report cumulative milliseconds of output audio actually played:
interrupted, stop playback and discard queued audio. The session uses playback acknowledgements to align the caller’s heard reply with conversation memory; without them it estimates the position from the buffered playback schedule. Closing the socket closes the session.
Other JSON messages include session_started, transcript_partial, transcript_committed, agent_status, agent_text_delta, agent_text_done, filler, metrics, voice_usage, interrupted, error, session_transcript, and session_ended. agent_status reports tool activity. A committed transcript with replace=true replaces the previous user transcript bubble rather than appending another.
WebRTC
Create a peer connection with a microphone audio track and a data channel before creating the offer. Complete ICE gathering, then send:TIMBAL_STUN_URL and, when needed, TIMBAL_TURN_URL, TIMBAL_TURN_USERNAME, and TIMBAL_TURN_PASSWORD for your network. TIMBAL_VOICE_RTC_FORCE_RELAY=1 filters to relay candidates only when TURN is configured. The route returns 501 if WebRTC dependencies are unavailable.
LiveKit
Install the LiveKit extra:TIMBAL_VOICE_DIAL_SECRET to validate this header and TIMBAL_LIVEKIT_URL to restrict the destination. Clients join and publish microphone audio using LiveKit’s client SDK. Audio uses room tracks; session events use the timbal.events data topic.
hello_wait_secs controls the browser configuration-hello window (default two seconds), and sip_hello_wait_secs controls the SIP caller window (default zero). Phone callers normally have no data channel to send a hello.
The separate boot-environment path uses TIMBAL_VOICE_TRANSPORT=livekit, TIMBAL_LIVEKIT_URL, and TIMBAL_LIVEKIT_TOKEN to join one room at startup. Reserve that path for a process created for a single call; long-lived deployments should dial per request.
Twilio and Telnyx
Expose the server through HTTPS with WebSocket upgrades and configure the carrier’s voice webhook:- Twilio:
POST https://YOUR_HOST/voice/twilio/incoming. - Telnyx:
POST https://YOUR_HOST/voice/telnyx/incoming.
/voice/twilio/stream or /voice/telnyx/stream. The bridge decodes carrier 8 kHz G.711 μ-law audio into the session format and converts speech back to the carrier format. Interruption clears carrier playback; mark echoes track playback position.
Set TWILIO_AUTH_TOKEN or TELNYX_PUBLIC_KEY to enable incoming-webhook signature verification. Without these variables, the webhook accepts requests with a development-mode warning. Configure your reverse proxy so the externally visible URL used by the carrier is the URL validated by the webhook.
The routes answer and bridge calls. Placing an outbound call requires your carrier integration to originate it and connect the media stream. A host that knows the call direction can select outbound_greeting; setting the greeting alone does not originate a call.
Approval and input requests
Voice sessions forwardagent_approval and agent_interaction events when the agent pauses. Render the request and resume through POST /stream with the event’s run_id as parent_id and the decision in resume. Speaking into the microphone is not an approval resolution. See Client integration.
An agent_text_done event also carries the agent run_id, allowing an authorized application to continue that conversation through the text endpoint. Identity and prior-run continuation require host-controlled authorization; they are separate from speech tuning in the hello.