Proof of concept on synthetic data. This page is a demonstration, not a production system.
← portfolio
Evaluation harness · question set · gold answers · scoring · failure analysis

How well does it work? Measured on 124 questions, including the ones built to break it.

The engine is rule-based (no LLM). The harness doesn’t care: it feeds questions to any engine exposing ask(question, role) and scores the structured result against gold answers computed by a separate SQL implementation. All data is synthetic.

Author of engine = author of testsSynthetic dataRe-run live in your browser below
Read the numbers honestly
  • Dev (54) was written alongside the engine and used to tune it. 100% here means “the engine does what its author intended”, not “it is accurate”.
  • Holdout (40) was written and hashed before the engine existed (hash file) and scored once on the first build. Same author, and I had read those questions before writing the engine, so it checks generalisation across paraphrases at best, not independence. It is easy relative to the stress set.
  • Stress (30) was written after v1, adversarially. v1’s 15/30 is the honest untuned number. v2 (21/30) was changed after seeing those failures (fail-closed guards, a few date forms), so v2’s stress score is partly tuned and is shown only to illustrate the trade-off, not as generalisation.
  • With only 124 questions and one author, differences of a few points are noise. Nothing here should be read as a benchmark against other systems.

Headline (current engine)

v1 → v2, by split

What changed in v2 after the stress set exposed gaps: a guard so unsupported constructs (shares, averages, thresholds, multiple periods, top-N-per-group) are refused instead of silently answered wrongly; a fail-closed rule for “profit/cost” wording under restricted roles; “why” questions refused; three extra date phrasings.

Splitv1 passv2 passWhat it tests

Failure analysis

Each failure gets one label. The important distinction: a wrong answer is worse than a refusal, and a missed refusal on access control is worse than either.

v1 failure types (stress split drove all of them)
v2 failure types

Re-run it here

This runs the current engine in your browser on every question and scores the result against the gold rows shipped in eval/report.json (computed offline by score.py from hand-written SQL). It should reproduce the table above exactly.

Question set

IDSplitCategoryQuestionRoleExpectedv1v2

Method

Gold answers

Answerable questions carry a hand-written SQL query, executed on facts.csv in SQLite by score.py. That is a different implementation (SQL vs. JS aggregation), so a shared bug is unlikely, but the gold SQL is still written by one person. Where a question is arguable (e.g. s30 “top-line number”) the gold is flagged in the question note.

Scoring

Answer: outcome must be “answered” and rows must match gold (relative tolerance 1e-6; order checked only for ranked questions). Refuse / clarify: right outcome, an accepted reason code, and no data rows or data-like numbers in the message (leak check). Metrics: exact-match on answerable, refusal recall/precision, clarify rate, access-control pass rate, and a count of safety-critical failures (missed refusals or leaks on access/refuse questions).

Baselines

“Always refuse” and “always return total net sales” are scored on the same questions so the headline has context: they get ~4% and 0%.

What it does not measure

Latency and cost (trivial for rules), robustness to real users’ phrasing (the set is one person’s guess at it), multi-turn behaviour, or any LLM. A production eval would add sampled real questions with human-labelled gold, inter-rater checks, and a regression gate in CI.

Reproduce offline: node eval/run_eval.mjs, then the scoring script (eval/score.py) · files: questions.json, FAILURE_ANALYSIS.md, report.json.

Synthetic data. Rule-based engine. Written evaluation, not marketing.