Concepts¶
Suites¶
A suite is a YAML (or JSON) file that lives in your repository next to the code it tests. It names
a target, lists tasks, and attaches graders. Unknown keys are errors, so a typo fails loudly
instead of silently disabling a check. Run evalkit validate evals.yaml to check a file without
running it.
Tasks¶
A task is one input plus the checks its output must pass. Tasks can live inline in the suite or in
a separate .yaml, .json, or .jsonl file, which is convenient for datasets. Each task has an
id, an input or vars for the prompt template, an optional expected value, tags for
selection, and its own graders.
Targets¶
The target is what you evaluate:
- A model:
openai(OpenAI and any compatible API, including gateways and local servers) oranthropic. - Your program:
commandruns a process and sends the prompt on stdin. - Your service:
httpposts a templated JSON body to your endpoint and reads the answer from the response. - Nothing at all:
mockreturns canned or echoed answers. It's deterministic and offline, which makes it the right target for demos and for testing a suite itself.
Testing your app instead of the bare model means you grade what your users actually see, after retrieval, tools, and post-processing.
Graders¶
A grader turns one output into a verdict: pass or fail, a score from 0 to 1, and a reason. evalkit
ships exact, contains, regex, json_schema, numeric, python (your own function), and
llm_judge (a model scores the output against a rubric). Any grader can be negated for "must not"
checks and weighted in the task score.
Suite-level graders apply to every task, before each task's own graders.
Scoring¶
- A run passes when every grader passes. Its score is the weighted mean of the grader scores.
- A task passes when at least
task_pass_rateof its runs pass. - The suite passes when at least
fail_underof its tasks pass.
Task statuses are pass, FAIL, FLAKY (some runs passed, but too few), and ERROR (the target
or a grader raised an error).
Repeats and flakiness¶
Model output isn't deterministic. Set repeat (or pass -n 10) to run each task several times.
evalkit reports each task's pass rate across its runs and counts a task as flaky when its runs
disagree. This separates a task that's broken from one that's unreliable.
The response cache¶
evalkit stores model responses in a SQLite file (.evalkit/cache.sqlite by default). The key is
a hash of the provider identity (provider, model, and settings, never secrets), the prompt, the
system prompt, the seed, and the repeat index. A rerun with unchanged prompts makes no API calls,
so iterating on graders is free. command and http targets aren't cached by default, because
your app changes between runs. Pass --no-cache to bypass the cache for one run.
Baselines and the regression gate¶
evalkit run -o results.json writes everything about a run. evalkit compare base.json head.json
diffs two such files and fails when the pass rate or the mean score drops by more than
--threshold, or, with --strict, when any single task goes from pass to fail. Save a baseline on
your main branch and compare every pull request against it.
Outputs¶
One run can write a terminal table, the full results JSON, JUnit XML for your CI's test report view, a GitHub-flavored Markdown summary, and a shields.io endpoint badge.