Scenario format¶
A scenario is a YAML file that describes one scripted conversation. Validate a file without running it:
voicebench validate my-scenario.yaml
Unknown keys are errors, so a typo never silently falls back to a default.
Top-level keys¶
| Key | Default | Description |
|---|---|---|
name |
required | Name used in reports and in the default output directory. |
description |
"" |
Free text, copied into the report. |
sample_rate |
16000 |
Sample rate in Hz for all audio in the session. Minimum 8000. |
frame_ms |
20 |
Size of each user audio frame sent to the agent. Minimum 5. |
tail_ms |
3000 |
How long to keep recording after the last turn ends. The session also waits for the agent to finish speaking. |
max_duration_ms |
120000 |
Hard limit on session length. |
vad |
Detector settings: user and agent blocks, each with threshold_db, hysteresis_db, frame_ms, min_speech_ms, and min_silence_ms. See metrics. |
|
adapter |
{type: mock} |
type is a built-in adapter name or module.path:ClassName. options are passed to the adapter as keyword arguments. |
defaults |
Default response_timeout_ms and max_stop_ms for every turn. |
|
turns |
required | Non-empty list of turns. |
assert |
Pass or fail limits. See assertions. |
Turns¶
| Key | Default | Description |
|---|---|---|
id |
turn-N |
Unique name for the turn. |
audio |
required | Exactly one of file or synth. See audio. |
start |
see below | When the turn starts. See start rules. |
expect |
respond |
respond: the agent should answer. stop: the turn interrupts the agent, which should stop talking. ignore: the audio is noise, and the agent should neither stop nor answer. |
text |
What the audio says. Enables WER when the agent reports transcripts. | |
response_timeout_ms |
5000 |
How long after the user's speech ends to wait for a reply. |
max_stop_ms |
1000 |
For stop, the longest acceptable stop time. For ignore, stopping within this time counts as a false barge-in. |
Audio¶
A file path is relative to the scenario file:
audio: {file: audio/question.wav}
voicebench reads 8-, 16-, and 32-bit PCM WAV files, downmixes them to mono, and resamples them to the session rate.
A synthetic signal needs no file:
audio: {synth: speech, duration_ms: 1500, level_db: -20, seed: 2}
synth |
Signal |
|---|---|
speech |
Harmonic complex on a wandering pitch, amplitude modulated at a syllable rate. Energy detectors treat it as speech. Speech-to-text engines and neural VADs do not. |
tone |
Sine tone at freq_hz (default 440). |
noise |
White noise. Useful for ignore turns. |
silence |
Digital silence. |
level_db sets the RMS level in dBFS (default -20). seed changes the random parts of speech and noise.
Use synthetic signals with the mock agent and for transport smoke tests. For a real agent, record your prompts, because a real agent runs speech recognition on them.
Start rules¶
start: {after: agent_speech_end, delay_ms: 400, timeout_ms: 10000}
A turn never starts before the previous turn's audio ends. After that, it starts delay_ms after its anchor:
after |
Anchor |
|---|---|
session_start |
The start of the session. Default for the first turn, with delay_ms: 500. |
previous_turn_end |
The end of the previous turn's audio. |
agent_speech_start |
The next agent speech onset after the previous turn ended. Use it for barge-in and noise turns. |
agent_speech_end |
The next agent speech offset after the previous turn ended. Default for later turns, with delay_ms: 0. |
The agent anchors come from a live VAD on the agent track, using the vad.agent settings. If the anchor never happens, the turn starts timeout_ms (default 10000) after the previous turn ended, and the report marks its trigger as timeout.
Assertions¶
| Key | Passes when |
|---|---|
max_response_latency_ms |
Maximum response latency is at most the limit. |
p50_response_latency_ms |
Median response latency is at most the limit. |
p90_response_latency_ms |
90th percentile response latency is at most the limit. |
p50_ttfa_ms |
Median time to first audio is at most the limit. |
max_barge_in_stop_ms |
Slowest barge-in stop time is at most the limit. |
min_barge_in_success_rate |
Barge-in success rate (0 to 1) is at least the limit. |
max_false_barge_ins |
Count of false barge-ins is at most the limit. |
max_missed_responses |
Count of missed responses is at most the limit. |
max_wer |
Mean WER is at most the limit. |
An assertion on a metric with no data fails. With --repeat, assertions apply to the pooled turns of all sessions.