Skip to content

Datasets & scripted eval

care dataset manages a small eval dataset attached to one chain — the terminal twin of the TUI’s /dataset. Each entry pairs an input task with an expected substring; run replays every entry through the chain and scores it. The sub-action is required: list, add, run, or export.

  1. add cases with --expected — the substring the chain’s answer should contain.
  2. run the dataset — every case is replayed through the chain (the CARL executor) and scored by case-insensitive substring match. It prints a per-case ✓/✗ and a score: N/M passed line, and exits 1 if any case fails — so it gates a CI job out of the box.
  3. export the dataset as JSONL when you want to score it in an external eval framework instead.

Add one test case to a chain’s dataset.

Terminal window
care dataset add weather "weather in Paris tomorrow" --expected "Paris"
FlagDefaultPurpose
--expected EXPECTEDrequiredSubstring the answer must contain (the run gate matches this case-insensitively).
--rubric RUBRIC""LLM-judge rubric used only by the TUI run — care dataset run ignores it and scores by substring.

List a chain’s dataset entries — one row per case with its status and a truncated task.

Terminal window
care dataset list weather
care dataset list weather --json
FlagDefaultPurpose
--jsonoffEmit the entries as JSON instead of the row view.

Replay every entry through the chain and score it.

Terminal window
care dataset run weather
✓ weather in Paris tomorrow
✗ five-day forecast for SF
score: 1/2 passed (substring)
FlagDefaultPurpose
--jsonoffEmit structured results instead of the per-case lines.

Exit code is 0 only when every scored case passes; otherwise it is 1.

Export the dataset as JSONL — one entry per line — for an external eval framework.

Terminal window
care dataset export weather weather-eval.jsonl

Because run exits non-zero the moment a case fails, a CI step is just the command itself — no extra glue:

Terminal window
# seed the dataset once (or commit the cases via `add` in a setup step)
care dataset add weather "weather in Paris tomorrow" --expected "Paris"
care dataset add weather "is it raining in London" --expected "London"
# gate: non-zero exit fails the pipeline
care dataset run weather
# .github/workflows/eval.yml (excerpt)
- name: Eval the weather chain
run: care dataset run weather
env:
OPENROUTER_API_KEY: ${{ secrets.OPENROUTER_API_KEY }}

Run care dataset --help for the authoritative, up-to-date flag set.