# Data quality: checks, faults, and lessons

`src/dq.py` runs 19 checks across seven dimension labels and scores itself against `data/fault_manifest.json` (possible because the faults are planted).

| Dimension | Checks | Notes |
|---|---|---|
| uniqueness | U1 | duplicate `event_id` |
| completeness | C_user_id, C_ts, C_dst_resource, C_outcome, C_auth_method, C_source_system, C_reason | C_reason is a warning: kept, relabelled `unknown` |
| validity | V_outcome, V_latency, V_ip | outcome must be exactly `success`/`failure` |
| timeliness | T_window | future / pre-window timestamps (clock skew) |
| integrity | R_user, R_resource | foreign keys |
| consistency | L_mfa1, L_mfa2, L_block, L_device | domain invariants; L_device is a *finding*, not a data error |
| freshness / volume | F_gap | per-source hourly volume ≤ 15% of the median for the same hour-of-day and weekday/weekend, requiring median ≥ 8 events/hour |

## Planted faults and results (default seed)
Duplicates (900), NULL users, corrupted outcome enum in one connector, 40 future timestamps, ~265 failures with no reason, and a ~5-hour VPN ingestion gap. Recall is 100% on all six; there are no extra flags beyond the planted ones except the counted explanations in `dq_report.json` (e.g. duplicated rows inherit their original's fault).

## Limits of the checks (honest)
* **Volume check needs volume.** With a median below 8 events/hour a real gap is invisible; on a 300-user world it is missed (tested, documented). On other seeds it raises about one false-alarm hour in 1,440.
* **Planted faults are the ones I thought of.** 100% recall on faults I wrote says the checks work for those faults, not that the data is clean. Unknown-unknowns need profiling against source documentation and drift monitors over time.
* **Thresholds are hand-set** (0 for most checks). In production each would have an owner, an SLA and an alert route.

## Lesson: a "repair" that flipped security outcomes
The first version of the trusted view repaired corrupted `outcome` values (`ok`, `succes`, `SUCCESS` → success; `fail` → failure). The consistency checks then flagged successes carrying `blocked_by = MFA` and an unsatisfied MFA requirement. Cause: the corruption was applied to *failures* too (a connector bug that garbles the field regardless of truth), so the repair turned failed MFA timeouts into successes — exactly the error that overstates a control's failure rate or hides an attack. Fix: non-canonical outcomes are **quarantined, not repaired**; `tests/test_pipeline.py::test_corrupted_outcomes_are_quarantined_not_repaired` pins the behaviour. Rule of thumb adopted: never auto-repair a field whose value changes a security decision unless the repair can be verified against a second source.

## A control gap the consistency check surfaced (L_device)
In the default world, 5 successes reach a high-sensitivity resource (`code_repo`) from an unmanaged device by a non-service user. I first assumed these were lateral movement through a managed device; querying the truth labels showed I was wrong: **all 5 are `session_hijack` events using a stolen `sso_token`**. The simulation lets token-based sessions skip the device-trust control (a modelled assumption that I had not listed as a deliberate gap). So the consistency rule found a real-in-the-model bypass path: *token replay is not subject to device trust*. It is not a data error, and it is reported in the memo as residual risk, with the caveat that the gap exists because of how I wrote the generator.
