Skip to content

Command-line reference

This page is generated from the evalkit argument parser when the site builds, so it always matches the released code.

Exit codes: 0 passed, 1 failed the pass-rate bar or regressed, 2 usage or config error. NO_COLOR and FORCE_COLOR control terminal colors.

evalkit

usage: evalkit [-h] [--version] COMMAND ...

Evals as code for LLM apps and agents.

positional arguments:
  COMMAND
    run       run a suite and report results
    compare   diff two results files and gate on regressions
    validate  check suite files without running them
    init      write a starter evals.yaml
    schema    print the suite JSON Schema (for editor support)
    cache     inspect or clear the response cache

options:
  -h, --help  show this help message and exit
  --version   show program's version number and exit

evalkit run

usage: evalkit run [-h] [-o PATH] [--junit PATH] [--markdown PATH] [--badge PATH]
                   [--badge-label BADGE_LABEL] [--format {table,json}] [-q] [--no-color] [-t TAG]
                   [-k TEXT] [--limit N] [-j N] [-n N] [--seed SEED] [--no-cache]
                   [--cache-path PATH] [--fail-under RATE] [--baseline PATH]
                   [--threshold THRESHOLD] [--strict]
                   suite

positional arguments:
  suite                 path to the suite YAML file

options:
  -h, --help            show this help message and exit

outputs:
  -o PATH, --output PATH
                        write the full results JSON
  --junit PATH          write JUnit XML
  --markdown PATH       write a Markdown summary
  --badge PATH          write a shields.io endpoint badge JSON
  --badge-label BADGE_LABEL
                        badge label (default: evals)
  --format {table,json}
                        stdout format
  -q, --quiet           print nothing on stdout
  --no-color            disable ANSI colors

selection:
  -t TAG, --tag TAG     only tasks with this tag (repeatable)
  -k TEXT, --match TEXT
                        only tasks whose id contains TEXT
  --limit N             run at most N tasks

execution:
  -j N, --concurrency N
                        parallel requests
  -n N, --repeat N      runs per task
  --seed SEED           base seed
  --no-cache            ignore and do not write the cache
  --cache-path PATH     SQLite cache location

gating:
  --fail-under RATE     minimum pass rate, 0 to 1
  --baseline PATH       results JSON to compare against
  --threshold THRESHOLD
                        allowed drop vs baseline
  --strict              fail if any task regresses

evalkit compare

usage: evalkit compare [-h] [--threshold THRESHOLD] [--strict] [--format {text,markdown,json}]
                       base head

positional arguments:
  base                  baseline results JSON
  head                  new results JSON

options:
  -h, --help            show this help message and exit
  --threshold THRESHOLD
                        allowed drop in pass rate or score, as a fraction (0.05 = 5 points)
  --strict              fail if any single task regresses
  --format {text,markdown,json}

evalkit validate

usage: evalkit validate [-h] SUITE [SUITE ...]

positional arguments:
  SUITE

options:
  -h, --help  show this help message and exit

evalkit init

usage: evalkit init [-h] [--force] [directory]

positional arguments:
  directory

options:
  -h, --help  show this help message and exit
  --force     overwrite an existing file

evalkit schema

usage: evalkit schema [-h]

options:
  -h, --help  show this help message and exit

evalkit cache

usage: evalkit cache [-h] [--cache-path PATH] {stats,clear}

positional arguments:
  {stats,clear}

options:
  -h, --help         show this help message and exit
  --cache-path PATH