Skip to main content

Events and transcripts

VoiceSession.run() yields typed events from timbal.voice. The server converts them to transport JSON; WebRTC and LiveKit carry speech on audio tracks instead of JSON audio frames.
session.transcript contains committed user and assistant entries with timestamps. Fillers are marked filler=True. At session end, the server sends session_transcript followed by session_ended. Transcript text reflects interruption handling and may contain only the heard prefix of a cancelled assistant reply.

Latency

TurnMetricsEvent reports millisecond durations and a run_id for joining a voice turn to the agent trace. Unobserved stages have None, not zero.
  • speech_end_to_transcript_ms: speech-end evidence to the committed transcript. speech_end_source identifies VAD or the coarser partial-transcript estimate.
  • eou_to_llm_first_token_ms: committed transcript to the first model text delta.
  • eou_to_first_audio_ms: committed transcript to the first emitted speech bytes. This excludes transcription delay and network/client playback delay.
  • turn_total_ms, llm_total_ms, and tts_total_ms: turn and stage durations.
  • interrupted, heard_bytes, and playback_acks_received: interruption and playback evidence.
  • filler_spoken and filler_count: whether a latency-masking phrase was spoken. When present, first-audio latency can measure the filler rather than the substantive answer.
Measure time to the useful answer as well as first audio. Compare providers on representative calls, including slow tools, short acknowledgements, background noise, and interruptions.

Provider-reported speech usage

OpenAI transcription and mini-TTS adapters emit VoiceUsageEvent independently of transcripts and audio. These events carry provider quantities, not dollar amounts. Other providers and custom adapters emit nothing unless they implement the usage capability. For billing, register a non-blocking listener on the adapters before the session starts. Delivery through the client connection alone can lose teardown events:
Listeners are synchronous and must not block. Listener failures are logged and isolated from speech. Delivery is in-process; your consumer must persist records durably. provider, operation (stt or tts), and model identify the resolved operation. Deduplicate by usage_id; prefer a complete report over an earlier incomplete report with the same ID. request_id and item_id, when present, help reconcile provider operations. status="complete" means terminal provider counters arrived. status="incomplete" means quantities are missing, invalid, or unavailable after cancellation, error, or closure. Incomplete is not zero cost. usage preserves the provider object, including any partial details. OpenAI transcription can report token usage with separate input audio/text details, or duration usage with seconds. Mini-TTS reports input text and output audio tokens through the Speech API’s SSE completion event. Legacy tts-1 and tts-1-hd use PCM streaming and do not supply these SSE usage reports. Do not apply one text rate to total_tokens or charge both an estimate and the reported quantities. Generated or cancelled audio can be billed even when the caller never hears it. Apply the resolved speech model’s published rates separately from the text agent’s OutputEvent.usage. The text model catalog and cost SQL generator do not provide a complete speech rate card. Use an explicit estimate policy for incomplete operations, and label those estimates.

Recordings and client audio

See Recording calls to save session audio locally or separate the speakers. Thinking sounds and background audio are played by the browser and do not appear in the server’s audio recording. Spoken fillers use speech synthesis and do appear in session audio and usage.