Skip to content

Suite file reference

A suite is a YAML (or JSON) file. Unknown keys are errors, so typos fail loudly. Run evalkit validate evals.yaml to check a file without running it, and evalkit schema for a JSON Schema your editor can use. The full schema is on the Suite JSON Schema page.

version: 1                      # optional, the only version is 1
name: support-bot               # required
description: Optional text shown in reports.

target: {provider: openai, model: gpt-4o-mini}   # required, see "Providers"
judge: {provider: openai, model: gpt-4o-mini}    # default judge for llm_judge graders

system: You are a helpful support agent.         # system prompt for the target
prompt: "Customer says: {{message}}"             # template; omit it to send each task's input as-is

graders:                        # applied to every task, before each task's own graders
  - type: contains
    value: ["Thanks"]

tasks:                          # a list, or a path to a .yaml, .json, or .jsonl file
  - id: refund                  # defaults to task-1, task-2, ...
    vars: {message: "I want a refund"}
    input: null                 # string prompt, or any value when you use `prompt`
    expected: Billing > Refunds # default value for exact, contains, and numeric
    tags: [billing]             # select with `evalkit run -t billing`
    metadata: {owner: payments} # free-form, passed to python graders
    graders:
      - type: regex
        pattern: 'Billing\s*>\s*Refunds'

settings:
  concurrency: 4                # parallel requests
  repeat: 1                     # runs per task, for flakiness stats
  retries: 2                    # retries on timeouts, 429, and 5xx
  retry_backoff: 0.5            # seconds, doubles per attempt
  seed: 0                       # passed to providers that accept a seed
  cache: true                   # cache model responses in SQLite
  cache_path: .evalkit/cache.sqlite   # relative to the directory you run evalkit from
  fail_under: 1.0               # minimum share of passing tasks for exit code 0
  task_pass_rate: 1.0           # share of a task's repeats that must pass

Templates

Prompts, grader values, regex patterns, rubrics, and http bodies accept {{ name }} placeholders. The variables are input, expected, id, metadata, vars, and each key of vars (and of input, when it's a mapping) at the top level. Dotted paths such as {{ vars.city }} and {{ items.0 }} walk into nested values. A value that is exactly one placeholder keeps its type, so value: "{{ expected }}" can resolve to a number.

Provider settings in target and judge also expand environment variables: ${VAR} or ${VAR:-default}. Task data is never expanded.

Providers

Every provider accepts timeout (seconds, default 60), cache (override the default), and pricing: {input_per_mtok, output_per_mtok} in USD per million tokens. evalkit ships no price table, so cost is reported only when you set prices.

Provider Key fields Cached by default
openai model, base_url (default https://api.openai.com/v1), api_key_env (default OPENAI_API_KEY, null for none), temperature, max_tokens, headers, params Yes
anthropic model, base_url, api_key_env (default ANTHROPIC_API_KEY), anthropic_version, temperature, max_tokens (default 1024), headers, params Yes
command command (string or list), cwd, env. The prompt goes to stdin unless an argument contains {{prompt}}. stdout is the output. EVALKIT_SEED, EVALKIT_REPEAT, and EVALKIT_SYSTEM are set. No
http url, method (POST, PUT, or GET), headers, body (JSON template, default {"input": "{{prompt}}"}), output_path (dotted path into the response, such as data.answer) No
mock responses (list of {match: regex, output, flaky}), default (echoes the prompt if unset), latency_ms, flaky, flaky_output, fail_times Yes

The openai provider works with any server that speaks the chat completions API, including gateways such as LiteLLM and OpenRouter and local servers such as vLLM, llama.cpp, and Ollama. Set base_url and, for servers without auth, api_key_env: null.

params merges extra fields into the request body, for example params: {top_p: 0.9}.

Scoring

Each grader returns pass or fail, a score from 0 to 1, and a reason. A run passes when every grader passes. Its score is the weighted mean of the grader scores. A task passes when at least task_pass_rate of its runs pass. The suite passes when at least fail_under of its tasks pass.

Task statuses in reports are pass, FAIL, FLAKY (some runs passed, too few to pass the task), and ERROR (the target or a grader raised an error).

Graders

Every grader accepts name (label in reports), negate: true (invert the verdict, for "must not" checks), and weight (default 1, used in the task score).

Type Passes when Options
exact The output equals value. value (defaults to expected), ignore_case, strip (default true)
contains The output contains value, or each item of a list. value, mode: all or any, ignore_case. Score is the share of items found.
regex pattern matches anywhere in the output. pattern, ignore_case, fullmatch (match the whole stripped output)
json_schema The output is JSON that validates against the schema. schema (inline) or schema_file (JSON or YAML, relative to the suite), extract (default true: find JSON in code fences or prose)
numeric A number in the output is within tolerance of value. value (defaults to expected), abs_tol, rel_tol, pick: last or first. Commas such as 1,250.50 are handled.
python Your function says so. function (checks.py:name relative to the suite, or package.module:name), args (keyword arguments)
llm_judge A judge model scores the output at or above min_score. rubric (templated), provider (defaults to the suite judge), scale (default 5), min_score (default scale - 1)

Python graders

A python grader receives the output and a context dict with input, expected, vars, id, metadata, prompt, repeat, and seed. It can be sync or async and returns one of:

def within_length(output, context, max_words=40):
    words = len(output.split())
    return {
        "pass": words <= max_words,
        "score": min(1.0, max_words / max(words, 1)),
        "reason": f"{words} words (max {max_words})",
    }


# Also valid: True / False, a score in [0, 1] (passes at 0.5), or (passed, "reason").

LLM-as-judge

The judge sees the rubric, the task input, the task's expected value as a reference when there is one, and the output. It replies with {"score": <1..scale>, "reason": "..."}. evalkit normalizes the score to 0 to 1 and passes the grade at min_score. Judge calls go through the same retries and cache as the target. A reply without a readable score fails the grade and shows the raw reply in the reason.

judge:
  provider: anthropic       # reads ANTHROPIC_API_KEY
  model: ${JUDGE_MODEL}
graders:
  - type: llm_judge
    name: tone
    rubric: The reply is polite, acknowledges the customer, and gives a concrete next step.
    min_score: 4