Skip to content

CI integration

GitHub Actions

name: Evals
on: [pull_request]
permissions:
  contents: read
jobs:
  evals:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v5
      - uses: superintelligenceco/evalkit@v0
        id: evals
        with:
          suite: evals.yaml
          python-version: "3.12"
        env:
          OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
      - run: echo "Pass rate ${{ steps.evals.outputs.pass-rate }}"

The release workflow moves a major version tag, such as v0, to each new release, so superintelligenceco/evalkit@v0 tracks the latest 0.x release. Pin the action to a full commit SHA in production workflows. The action writes results.json, junit.xml, summary.md, and badge.json to output-dir (default evalkit-results) and appends the Markdown summary to the job summary.

Input Default Description
suite required Path to the suite file.
baseline Results JSON to compare against.
threshold 0 Allowed drop against the baseline, as a fraction.
strict false Fail when any single task regresses.
fail-under Minimum pass rate. Overrides the suite setting.
output-dir evalkit-results Where the reports go.
args Extra evalkit run arguments, such as --tag smoke.
python-version Install this Python first. Leave empty to use the runner's python3.
install the action's own copy A pip requirement to install instead.
job-summary true Write the Markdown summary to the job summary.

Outputs: passed, total, pass-rate, ok, and results-file.

This repository runs its own examples through the action on every push. See .github/workflows/evals.yml.

Baselines in CI

A simple pattern: upload results.json as an artifact on main, download the latest one in pull request workflows, and pass it as baseline. The response cache makes the main-branch run cheap, and you can cache .evalkit/ between jobs with actions/cache to skip repeat API calls.

Other CI systems

Run the CLI and publish the JUnit file with your system's test report feature. Download a release executable or install with pip:

pip install sic-evalkit
evalkit run evals.yaml --junit evalkit-junit.xml -o results.json

Badge

--badge badge.json writes a shields.io endpoint badge such as {"schemaVersion": 1, "label": "evals", "message": "7/8 passed", "color": "yellow"}. Publish the file anywhere public, such as GitHub Pages or a gist, and point https://img.shields.io/endpoint?url=<url-of-badge.json> at it.

Results format

-o results.json writes everything evalkit knows about a run: the suite, target, seed, a summary (pass rate, score, flaky and errored counts, tokens, cost, p50 and p95 latency, cached runs), and every task with every run, output, grade, and reason. format: 1 versions the layout. compare reads only the summary and each task's id, status, passed, and score.