agent.py, credentials, and server command.
From first conversation to a deployed application
- Your first voice agent: install, start, and talk. No tools, greeting, or added audio.
- Greetings: add a fixed hello, then control timing and generated wording.
- Using tools: look up a mock delivery and hear the result.
- Thinking sounds, spoken fillers, and background audio: try each optional sound behavior separately.
- Turn-taking and interruptions and silence and timeouts: test how the conversation handles pauses, interruptions, and waiting.
- Recording calls, transports and phone calls, and events and usage: connect a client or carrier and operate the session.
How it works
The built-in pipeline runs a regular TimbalAgent between transcription and speech synthesis. It keeps the agent’s tools, system prompt, memory, guardrails, and tracing. Speech providers and transports are configured separately from the language model.
The built-in server serves the same agent through text and voice endpoints. Open /voice for the browser playground, or connect your own client using WebSocket, WebRTC, LiveKit, or a Twilio/Telnyx media stream.
Keep replies concise and ask one question at a time. Tables, URLs, and long lists are difficult to follow aloud. Start with a conversation you can understand and measure before adding optional behavior.
Scope of the built-in pipeline
The shipped provider adapters use a transcription → text agent → speech pipeline.OpenAIRealtimeSTT is a transcription-only adapter; it does not run the conversation model or synthesize responses.
RealtimeModel and RealtimeSession define an extension interface for speech-to-speech models. No concrete speech-to-speech provider adapter ships yet, and the built-in server constructs VoiceSession. See Custom pipelines before planning an OpenAI Realtime or Gemini Live integration.
For processing a recorded file, see Audio files. For generating an audio file, see Text-to-speech tools.