What stops a talk-to-data assistant from saying what it shouldn’t
The design principles behind the demo, what each one looks like in code, a live red-team suite you can run, and — just as important — what this demo does not protect against.
1 · Access control
Three layers, applied in this order, all keyed to the session role — never to text in the question.
Each metric can be marked sensitive in the semantic model; each role lists deny_metrics. Denied ⇒ refusal with a reason code, before any value is computed. Example: the Sales analyst cannot query COGS, gross profit or gross margin.
Roles carry row_filters (e.g. category = Whiskey). The engine injects them into the structured query, not the SQL string, so the user can’t escape them with wording. Requests wholly outside scope are refused; partly outside are trimmed and disclosed.
The engine can only answer in metrics defined once in the semantic layer. There is no free-form SQL path, so “select * from facts” is an unsafe request, not a query. Result rows are capped (500).
Roles in this demo
| Role | Denied metrics | Row scope | Description |
|---|
2 · Audit log
{ seq, ts, role, question (≤300 chars),
outcome: answered | refused | clarify,
reason: restricted_metric | restricted_rows |
unsafe_request | out_of_scope | …,
metrics[], dims[], scope_applied{},
rows_returned, flags[], engine }- Metadata, not values. The log records that 40 rows came back, not what was in them, so the log itself can’t become a leak. (Tested: no 6+ digit values in any entry.)
- Refusals are logged. Repeated denied attempts are the signal you want.
- Flags such as claims_authority and unsafe_request support later detection rules.
- Here: in-memory, resets on reload. Production: append-only store, tamper-evident (hash chain), retention policy, restricted read access.
See it live: ask anything on the demo page and scroll to the audit table.
3 · Leak prevention
Refusal and clarification payloads contain a reason and generic text only — no values, counts or “there are N rows you can’t see”. The evaluation harness fails any refusal whose message contains data-like numbers.
Restricted metrics are refused before the engine echoes back what it understood, so the parse trace can’t confirm the existence of things the role can’t see.
If the wording hints at a restricted metric (“what it cost us to make it”) but doesn’t resolve to one, restricted roles are refused — otherwise the allowed half of a mixed question would be answered with the restricted half silently dropped. Found by probing; there is a regression test.
There is no model to inject into, and question text can’t change the role or the policy. Injection-style text (“ignore the access rules”) is flagged, refused and logged. With an LLM in the loop, the same invariant must hold: the model proposes a structured query, and the policy layer — not the model — decides.
4 · Red-team suite (runs in your browser)
35 adversarial cases (tests/redteam.json): denied metrics, paraphrased denied metrics, authority claims, scope escapes, SQL injection strings, unknown role, empty input, mixed allowed+denied requests. Same file is run in CI-style by node tests/test_engine.mjs.
| ID | Role | Prompt | Invariant | Outcome | Result |
|---|
5 · Known limitations (what would still go wrong)
- Keyword guards are incomplete by nature. The cost/profit fail-closed list caught the phrasings I tried; a novel paraphrase that resolves to no known term will get a clarification, not data — but I can’t prove there’s no phrasing that slips a restricted metric through. That is why the real control is the metric-level deny on the resolved structured query, with the keyword net as defence-in-depth.
- Aggregates can leak. Allowed aggregates can be differenced (total − visible segments ⇒ hidden segment). Mitigations not built: minimum-group-size thresholds, query budgets, noise.
- No authentication. The role is a dropdown standing in for a session.
- Client-side only. Policy, data and engine are all downloadable; see the scope note above.
- No rate limiting, no anomaly detection on the audit log, no PII (the dataset has none by design).
- An LLM engine would add new risks (prompt injection via data, hallucinated metrics, cost). Those need their own evaluation set — the harness is built to take one.