Skip to main content
A voice agent listens to you and speaks its reply. Start with a simple browser conversation, then add one behavior at a time. The lessons reuse the same agent.py, credentials, and server command.

From first conversation to a deployed application

  1. Your first voice agent: install, start, and talk. No tools, greeting, or added audio.
  2. Greetings: add a fixed hello, then control timing and generated wording.
  3. Using tools: look up a mock delivery and hear the result.
  4. Thinking sounds, spoken fillers, and background audio: try each optional sound behavior separately.
  5. Turn-taking and interruptions and silence and timeouts: test how the conversation handles pauses, interruptions, and waiting.
  6. Recording calls, transports and phone calls, and events and usage: connect a client or carrier and operate the session.
Use Providers and configuration as a reference when changing speech providers. Continue to Custom pipelines when you need to embed the session or implement an adapter.

How it works

The built-in pipeline runs a regular Timbal Agent between transcription and speech synthesis. It keeps the agent’s tools, system prompt, memory, guardrails, and tracing. Speech providers and transports are configured separately from the language model. The built-in server serves the same agent through text and voice endpoints. Open /voice for the browser playground, or connect your own client using WebSocket, WebRTC, LiveKit, or a Twilio/Telnyx media stream. Keep replies concise and ask one question at a time. Tables, URLs, and long lists are difficult to follow aloud. Start with a conversation you can understand and measure before adding optional behavior.

Scope of the built-in pipeline

The shipped provider adapters use a transcription → text agent → speech pipeline. OpenAIRealtimeSTT is a transcription-only adapter; it does not run the conversation model or synthesize responses. RealtimeModel and RealtimeSession define an extension interface for speech-to-speech models. No concrete speech-to-speech provider adapter ships yet, and the built-in server constructs VoiceSession. See Custom pipelines before planning an OpenAI Realtime or Gemini Live integration. For processing a recorded file, see Audio files. For generating an audio file, see Text-to-speech tools.