Python API reference¶
The docs build generates this page from the docstrings in the source. Most users need only run_benchmark and load_scenario:
import asyncio
from voicebench.bench import run_benchmark
from voicebench.scenario import load_scenario
result = asyncio.run(run_benchmark(load_scenario("examples/mock-conversation.yaml")))
print(result.summary["response_latency_ms"]["p50"], result.passed)
Running a benchmark¶
High-level API: run a scenario one or more times and build the report.
run_benchmark(scenario, adapter=None, options=None, repeat=1, realtime=False)
async
¶
Run scenario repeat times, each with a fresh adapter instance.
adapter and options override the scenario's adapter block. When
adapter names a different type than the scenario, the scenario's
options are not used.
Scenarios¶
Scenario files: YAML descriptions of a scripted conversation.
ScenarioError
¶
Bases: ValueError
Raised when a scenario file is invalid. The message names the offending key.
AudioSpec
dataclass
¶
Where a turn's audio comes from: a WAV file or a synthesized signal.
StartRule
dataclass
¶
When a turn starts: delay_ms after after, or after timeout_ms regardless.
Turn
dataclass
¶
One scripted user utterance and what the agent should do about it.
Scenario
dataclass
¶
A parsed and validated scenario.
load_scenario(path)
¶
Load and validate a scenario from a YAML file.
parse_scenario(data, base_dir=None, source=None)
¶
Validate a scenario from already-parsed YAML data.
Adapters¶
The adapter interface between the runner and an agent under test.
AgentAudio
dataclass
¶
Agent audio at the session sample rate, received at time seconds.
AgentClear
dataclass
¶
The agent asked the client to drop any audio it has not played yet.
Transcript
dataclass
¶
Text reported by the agent, for example its speech-to-text result.
SessionInfo
dataclass
¶
What an adapter needs to know about the session it joins.
AdapterError
¶
Bases: RuntimeError
Raised when an adapter cannot connect to or talk to the agent.
Adapter
¶
Bases: ABC
Base class for transports.
Subclasses implement :meth:connect, :meth:send_audio, and :meth:close,
and report everything the agent sends through :meth:emit. The runner
pulls reported events with :meth:drain once per frame.
start(session)
async
¶
Bind the session and connect. The runner calls this once.
connect()
abstractmethod
async
¶
Open the connection to the agent.
send_audio(pcm)
abstractmethod
async
¶
Send one frame of user audio at the session sample rate.
close()
abstractmethod
async
¶
Release all resources. Must be safe to call more than once.
emit(event)
¶
Record an event from the agent.
drain()
¶
Return and clear the events reported since the last call.
describe()
¶
Adapter details for the report. Never include secrets.
Voice activity detection¶
Energy-based voice activity detection with hysteresis and hangover.
The detector works on fixed analysis frames. A segment opens when the frame
level stays at or above threshold_db for min_speech_ms, and closes when
the level stays below threshold_db - hysteresis_db for min_silence_ms.
Reported onsets and offsets are backdated to the first and last active frame,
so the confirmation delay does not bias latency numbers.
VadConfig
dataclass
¶
Tuning parameters for :class:StreamingVad.
VadEvent
dataclass
¶
A confirmed speech boundary. time is in seconds from the stream origin.
Segment
dataclass
¶
A span of detected speech, in seconds.
overlaps(other)
¶
Return the overlap with other in seconds (0 if disjoint).
StreamingVad
¶
Incremental detector. Feed audio with :meth:push, then call :meth:flush.
events_to_segments(events)
¶
Pair onset and offset events into segments.
detect_segments(audio, sample_rate, config=None)
¶
Run the detector over a whole buffer and return the speech segments.
Metrics¶
Turn-level and session-level metrics computed from a recorded session.
TurnResult
dataclass
¶
Measurements for one scripted turn. Times are ms unless the name says _s.
SessionMetrics
dataclass
¶
Everything measured in one session.
analyze(record)
¶
Compute all metrics for a recorded session.
percentile(values, q)
¶
Linear-interpolated percentile, or None for an empty list.
summarize(sessions)
¶
Aggregate turn results across one or more sessions.
check_assertions(summary, assertions)
¶
Compare the summary against the scenario's assert block.
Word error rate¶
Word error rate.
normalize(text)
¶
Lowercase, strip punctuation (keeping apostrophes), and split into words.
edit_distance(ref, hyp)
¶
Levenshtein distance between two word sequences.
wer(reference, hypothesis)
¶
Return (substitutions + deletions + insertions) / reference words.
An empty reference scores 0.0 against an empty hypothesis and 1.0 otherwise.