Skip to content

Quickstart

This page takes you from nothing to a passing suite, then to a regression gate. Every command here runs offline, because the starter suite uses the built-in mock provider.

Install evalkit

pip install sic-evalkit

Or install the standalone executable, which bundles Python:

curl -fsSL https://raw.githubusercontent.com/superintelligenceco/evalkit/main/install.sh | sh

Check the install:

evalkit --version

Run the starter suite

mkdir my-evals && cd my-evals
evalkit init
evalkit run evals.yaml

evalkit init writes evals.yaml. evalkit run prints one row per task and exits with status 0 when the pass rate meets the suite's fail_under bar.

Point it at a real model

Replace the target in evals.yaml:

target:
  provider: openai          # any OpenAI-compatible API
  model: gpt-4o-mini        # reads OPENAI_API_KEY

For Anthropic models, use provider: anthropic, which reads ANTHROPIC_API_KEY. To test your own program or service instead of a model, use a command or http target. See Concepts.

Save a baseline and catch a regression

evalkit run evals.yaml -o base.json          # on main
evalkit run evals.yaml -o head.json          # on your branch
evalkit compare base.json head.json --threshold 0.05

compare exits with status 1 when the pass rate or the mean score drops by more than the threshold. To gate in one step, run evalkit run evals.yaml --baseline base.json --threshold 0.05.

Next steps