Command-line reference¶
This page is generated from the evalkit argument parser when the site builds, so it always
matches the released code.
Exit codes: 0 passed, 1 failed the pass-rate bar or regressed, 2 usage or config error.
NO_COLOR and FORCE_COLOR control terminal colors.
evalkit¶
usage: evalkit [-h] [--version] COMMAND ...
Evals as code for LLM apps and agents.
positional arguments:
COMMAND
run run a suite and report results
compare diff two results files and gate on regressions
validate check suite files without running them
init write a starter evals.yaml
schema print the suite JSON Schema (for editor support)
cache inspect or clear the response cache
options:
-h, --help show this help message and exit
--version show program's version number and exit
evalkit run¶
usage: evalkit run [-h] [-o PATH] [--junit PATH] [--markdown PATH] [--badge PATH]
[--badge-label BADGE_LABEL] [--format {table,json}] [-q] [--no-color] [-t TAG]
[-k TEXT] [--limit N] [-j N] [-n N] [--seed SEED] [--no-cache]
[--cache-path PATH] [--fail-under RATE] [--baseline PATH]
[--threshold THRESHOLD] [--strict]
suite
positional arguments:
suite path to the suite YAML file
options:
-h, --help show this help message and exit
outputs:
-o PATH, --output PATH
write the full results JSON
--junit PATH write JUnit XML
--markdown PATH write a Markdown summary
--badge PATH write a shields.io endpoint badge JSON
--badge-label BADGE_LABEL
badge label (default: evals)
--format {table,json}
stdout format
-q, --quiet print nothing on stdout
--no-color disable ANSI colors
selection:
-t TAG, --tag TAG only tasks with this tag (repeatable)
-k TEXT, --match TEXT
only tasks whose id contains TEXT
--limit N run at most N tasks
execution:
-j N, --concurrency N
parallel requests
-n N, --repeat N runs per task
--seed SEED base seed
--no-cache ignore and do not write the cache
--cache-path PATH SQLite cache location
gating:
--fail-under RATE minimum pass rate, 0 to 1
--baseline PATH results JSON to compare against
--threshold THRESHOLD
allowed drop vs baseline
--strict fail if any task regresses
evalkit compare¶
usage: evalkit compare [-h] [--threshold THRESHOLD] [--strict] [--format {text,markdown,json}]
base head
positional arguments:
base baseline results JSON
head new results JSON
options:
-h, --help show this help message and exit
--threshold THRESHOLD
allowed drop in pass rate or score, as a fraction (0.05 = 5 points)
--strict fail if any single task regresses
--format {text,markdown,json}
evalkit validate¶
usage: evalkit validate [-h] SUITE [SUITE ...]
positional arguments:
SUITE
options:
-h, --help show this help message and exit
evalkit init¶
usage: evalkit init [-h] [--force] [directory]
positional arguments:
directory
options:
-h, --help show this help message and exit
--force overwrite an existing file
evalkit schema¶
usage: evalkit schema [-h]
options:
-h, --help show this help message and exit
evalkit cache¶
usage: evalkit cache [-h] [--cache-path PATH] {stats,clear}
positional arguments:
{stats,clear}
options:
-h, --help show this help message and exit
--cache-path PATH