Data quality: checks, faults, and lessons
src/dq.py runs 19 checks across seven dimension labels and scores itself against data/fault_manifest.json (possible because the faults are planted).
| Dimension | Checks | Notes |
|---|---|---|
| uniqueness | U1 | duplicate event_id |
| completeness | C_user_id, C_ts, C_dst_resource, C_outcome, C_auth_method, C_source_system, C_reason | C_reason is a warning: kept, relabelled unknown |
| validity | V_outcome, V_latency, V_ip | outcome must be exactly success/failure |
| timeliness | T_window | future / pre-window timestamps (clock skew) |
| integrity | R_user, R_resource | foreign keys |
| consistency | L_mfa1, L_mfa2, L_block, L_device | domain invariants; L_device is a finding, not a data error |
| freshness / volume | F_gap | per-source hourly volume ≤ 15% of the median for the same hour-of-day and weekday/weekend, requiring median ≥ 8 events/hour |
Planted faults and results (default seed)
Duplicates (900), NULL users, corrupted outcome enum in one connector, 40 future timestamps, ~265 failures with no reason, and a ~5-hour VPN ingestion gap. Recall is 100% on all six; there are no extra flags beyond the planted ones except the counted explanations in dq_report.json (e.g. duplicated rows inherit their original's fault).
Limits of the checks (honest)
- Volume check needs volume. With a median below 8 events/hour a real gap is invisible; on a 300-user world it is missed (tested, documented). On other seeds it raises about one false-alarm hour in 1,440.
- Planted faults are the ones I thought of. 100% recall on faults I wrote says the checks work for those faults, not that the data is clean. Unknown-unknowns need profiling against source documentation and drift monitors over time.
- Thresholds are hand-set (0 for most checks). In production each would have an owner, an SLA and an alert route.
Lesson: a "repair" that flipped security outcomes
The first version of the trusted view repaired corrupted outcome values (ok, succes, SUCCESS → success; fail → failure). The consistency checks then flagged successes carrying blocked_by = MFA and an unsatisfied MFA requirement. Cause: the corruption was applied to failures too (a connector bug that garbles the field regardless of truth), so the repair turned failed MFA timeouts into successes — exactly the error that overstates a control's failure rate or hides an attack. Fix: non-canonical outcomes are quarantined, not repaired; tests/test_pipeline.py::test_corrupted_outcomes_are_quarantined_not_repaired pins the behaviour. Rule of thumb adopted: never auto-repair a field whose value changes a security decision unless the repair can be verified against a second source.
A control gap the consistency check surfaced (L_device)
In the default world, 5 successes reach a high-sensitivity resource (code_repo) from an unmanaged device by a non-service user. I first assumed these were lateral movement through a managed device; querying the truth labels showed I was wrong: all 5 are session_hijack events using a stolen sso_token. The simulation lets token-based sessions skip the device-trust control (a modelled assumption that I had not listed as a deliberate gap). So the consistency rule found a real-in-the-model bypass path: token replay is not subject to device trust. It is not a data error, and it is reported in the memo as residual risk, with the caveat that the gap exists because of how I wrote the generator.