Skip to main content

Install

Use Python 3.11 or newer and install the server and voice extras:
Or add them to a uv project:
The server extra provides HTTP and WebSocket support. The voice extra adds WebRTC and local turn detection. Local detection models are downloaded on first use; the first session can take longer while they load.

Configure credentials

Create a .env beside agent.py:
.env
The example uses OpenAI for the text agent and ElevenLabs for transcription and speech. Choose a voice accessible to your ElevenLabs account. Credentials remain on the server.

Create the agent

agent.py
voice_config configures the server’s speech session. The defaults are ElevenLabs scribe_v2_realtime for transcription and eleven_flash_v2_5 for speech. It can be a dict, a VoiceConfig, or a callable returning configuration. Unknown keys fail validation at server startup.

Start and speak

With uv, use uv run python -m timbal.server with the same arguments. The server loads .env from the working directory. TIMBAL_RUNNABLE=agent.py::agent can supply the import spec instead of the flag. Open http://127.0.0.1:4444/voice, allow microphone access, and start a session. Say “Hello, what can you help me with?” The agent should transcribe your words and speak a short reply. Ask a follow-up question to try a conversation. A browser microphone requires localhost or HTTPS. The agent waits for you to speak first. This example has no greeting, tools, fillers, or background audio. Leave the playground’s thinking sound and ambience controls off for this first conversation. The agent is also available at POST /run and POST /stream. Voice routes require an Agent runnable; a Workflow or standalone Tool cannot be used as the voice conversation runnable.

Next: let the agent say hello

Keep this agent.py, its .env, and the same server command for the following lessons. After changing the file, restart the server and start a new browser session. Continue to Greetings to make the agent speak first, then Using tools to give it a useful action. You can change speech providers later in Providers and configuration.