Skip to content

evalkit

Evals as code for LLM apps and agents. You write tasks and graders in a YAML suite, run them against a model or your own app, and fail the build when quality regresses.

evalkit running the support-bot example suite

evalkit sends each task to a target, grades the answers, and gives you a results table, JSON and JUnit output, a regression diff against a saved baseline, and a badge. It runs the same on your laptop and in GitHub Actions.

Install

pip install sic-evalkit

The distribution is named sic-evalkit. The command and the Python package are both evalkit.

curl -fsSL https://raw.githubusercontent.com/superintelligenceco/evalkit/main/install.sh | sh

The installer downloads the executable for your OS and CPU from the latest GitHub Release, checks it against SHA256SUMS, and puts it in ~/.local/bin.

docker run --rm ghcr.io/superintelligenceco/evalkit run evals.yaml

The image's working directory holds a starter suite, so this command runs offline.

Where to go next