Proof of concept on synthetic data. This page is a demonstration, not a production system.

← home · raw markdown

Data quality: checks, faults, and lessons

src/dq.py runs 19 checks across seven dimension labels and scores itself against data/fault_manifest.json (possible because the faults are planted).

Dimension Checks Notes
uniqueness U1 duplicate event_id
completeness C_user_id, C_ts, C_dst_resource, C_outcome, C_auth_method, C_source_system, C_reason C_reason is a warning: kept, relabelled unknown
validity V_outcome, V_latency, V_ip outcome must be exactly success/failure
timeliness T_window future / pre-window timestamps (clock skew)
integrity R_user, R_resource foreign keys
consistency L_mfa1, L_mfa2, L_block, L_device domain invariants; L_device is a finding, not a data error
freshness / volume F_gap per-source hourly volume ≤ 15% of the median for the same hour-of-day and weekday/weekend, requiring median ≥ 8 events/hour

Planted faults and results (default seed)

Duplicates (900), NULL users, corrupted outcome enum in one connector, 40 future timestamps, ~265 failures with no reason, and a ~5-hour VPN ingestion gap. Recall is 100% on all six; there are no extra flags beyond the planted ones except the counted explanations in dq_report.json (e.g. duplicated rows inherit their original's fault).

Limits of the checks (honest)

Lesson: a "repair" that flipped security outcomes

The first version of the trusted view repaired corrupted outcome values (ok, succes, SUCCESS → success; fail → failure). The consistency checks then flagged successes carrying blocked_by = MFA and an unsatisfied MFA requirement. Cause: the corruption was applied to failures too (a connector bug that garbles the field regardless of truth), so the repair turned failed MFA timeouts into successes — exactly the error that overstates a control's failure rate or hides an attack. Fix: non-canonical outcomes are quarantined, not repaired; tests/test_pipeline.py::test_corrupted_outcomes_are_quarantined_not_repaired pins the behaviour. Rule of thumb adopted: never auto-repair a field whose value changes a security decision unless the repair can be verified against a second source.

A control gap the consistency check surfaced (L_device)

In the default world, 5 successes reach a high-sensitivity resource (code_repo) from an unmanaged device by a non-service user. I first assumed these were lateral movement through a managed device; querying the truth labels showed I was wrong: all 5 are session_hijack events using a stolen sso_token. The simulation lets token-based sessions skip the device-trust control (a modelled assumption that I had not listed as a deliberate gap). So the consistency rule found a real-in-the-model bypass path: token replay is not subject to device trust. It is not a data error, and it is reported in the memo as residual risk, with the caveat that the gap exists because of how I wrote the generator.