How well does it work? Measured on 124 questions, including the ones built to break it.
The engine is rule-based (no LLM). The harness doesn’t care: it feeds questions to any engine exposing ask(question, role) and scores the structured result against gold answers computed by a separate SQL implementation. All data is synthetic.
- Dev (54) was written alongside the engine and used to tune it. 100% here means “the engine does what its author intended”, not “it is accurate”.
- Holdout (40) was written and hashed before the engine existed (hash file) and scored once on the first build. Same author, and I had read those questions before writing the engine, so it checks generalisation across paraphrases at best, not independence. It is easy relative to the stress set.
- Stress (30) was written after v1, adversarially. v1’s 15/30 is the honest untuned number. v2 (21/30) was changed after seeing those failures (fail-closed guards, a few date forms), so v2’s stress score is partly tuned and is shown only to illustrate the trade-off, not as generalisation.
- With only 124 questions and one author, differences of a few points are noise. Nothing here should be read as a benchmark against other systems.
Headline (current engine)
v1 → v2, by split
What changed in v2 after the stress set exposed gaps: a guard so unsupported constructs (shares, averages, thresholds, multiple periods, top-N-per-group) are refused instead of silently answered wrongly; a fail-closed rule for “profit/cost” wording under restricted roles; “why” questions refused; three extra date phrasings.
| Split | v1 pass | v2 pass | What it tests |
|---|
Failure analysis
Each failure gets one label. The important distinction: a wrong answer is worse than a refusal, and a missed refusal on access control is worse than either.
Re-run it here
This runs the current engine in your browser on every question and scores the result against the gold rows shipped in eval/report.json (computed offline by score.py from hand-written SQL). It should reproduce the table above exactly.
Question set
| ID | Split | Category | Question | Role | Expected | v1 | v2 | Live |
|---|
Method
Answerable questions carry a hand-written SQL query, executed on facts.csv in SQLite by score.py. That is a different implementation (SQL vs. JS aggregation), so a shared bug is unlikely, but the gold SQL is still written by one person. Where a question is arguable (e.g. s30 “top-line number”) the gold is flagged in the question note.
Answer: outcome must be “answered” and rows must match gold (relative tolerance 1e-6; order checked only for ranked questions). Refuse / clarify: right outcome, an accepted reason code, and no data rows or data-like numbers in the message (leak check). Metrics: exact-match on answerable, refusal recall/precision, clarify rate, access-control pass rate, and a count of safety-critical failures (missed refusals or leaks on access/refuse questions).
“Always refuse” and “always return total net sales” are scored on the same questions so the headline has context: they get ~4% and 0%.
Latency and cost (trivial for rules), robustness to real users’ phrasing (the set is one person’s guess at it), multi-turn behaviour, or any LLM. A production eval would add sampled real questions with human-labelled gold, inter-rater checks, and a regression gate in CI.
Reproduce offline: node eval/run_eval.mjs, then the scoring script (eval/score.py) · files: questions.json, FAILURE_ANALYSIS.md, report.json.