Datasets & scripted eval
care dataset manages a small eval dataset attached to one chain — the terminal
twin of the TUI’s /dataset. Each entry pairs an input task with an
expected substring; run replays every entry through the chain and scores it.
The sub-action is required: list, add, run, or export.
The eval loop
Section titled “The eval loop”addcases with--expected— the substring the chain’s answer should contain.runthe dataset — every case is replayed through the chain (the CARL executor) and scored by case-insensitive substring match. It prints a per-case ✓/✗ and ascore: N/M passedline, and exits 1 if any case fails — so it gates a CI job out of the box.exportthe dataset as JSONL when you want to score it in an external eval framework instead.
care dataset add <chain_id> <task>
Section titled “care dataset add <chain_id> <task>”Add one test case to a chain’s dataset.
care dataset add weather "weather in Paris tomorrow" --expected "Paris"| Flag | Default | Purpose |
|---|---|---|
--expected EXPECTED | required | Substring the answer must contain (the run gate matches this case-insensitively). |
--rubric RUBRIC | "" | LLM-judge rubric used only by the TUI run — care dataset run ignores it and scores by substring. |
care dataset list <chain_id>
Section titled “care dataset list <chain_id>”List a chain’s dataset entries — one row per case with its status and a truncated task.
care dataset list weathercare dataset list weather --json| Flag | Default | Purpose |
|---|---|---|
--json | off | Emit the entries as JSON instead of the row view. |
care dataset run <chain_id>
Section titled “care dataset run <chain_id>”Replay every entry through the chain and score it.
care dataset run weather ✓ weather in Paris tomorrow ✗ five-day forecast for SF
score: 1/2 passed (substring)| Flag | Default | Purpose |
|---|---|---|
--json | off | Emit structured results instead of the per-case lines. |
Exit code is 0 only when every scored case passes; otherwise it is 1.
care dataset export <chain_id> <output>
Section titled “care dataset export <chain_id> <output>”Export the dataset as JSONL — one entry per line — for an external eval framework.
care dataset export weather weather-eval.jsonlBecause run exits non-zero the moment a case fails, a CI step is just the
command itself — no extra glue:
# seed the dataset once (or commit the cases via `add` in a setup step)care dataset add weather "weather in Paris tomorrow" --expected "Paris"care dataset add weather "is it raining in London" --expected "London"
# gate: non-zero exit fails the pipelinecare dataset run weather# .github/workflows/eval.yml (excerpt)- name: Eval the weather chain run: care dataset run weather env: OPENROUTER_API_KEY: ${{ secrets.OPENROUTER_API_KEY }}Run care dataset --help for the authoritative, up-to-date flag set.
See also
Section titled “See also”- Production Commands — the
/dataset, evolution, and publish flow in the TUI. - CLI Overview — every headless subcommand at a glance.