Proof of concept on synthetic data. This page is a demonstration, not a production system.

← home · raw markdown

Memo: Are our authentication controls working, and what do they cost?

To: Security leadership (hypothetical) From: Max Mattox Status: self-directed project on synthetic data (371,226 events, 60 days, 1,200 users). Not employment experience; describes no employer's systems.

Recommendation

  1. Keep the staged MFA rollout and finish it. It reduced account compromises by about 12.4 per 1,000 active user-days (difference-in-differences; 95% interval 8.9 to 15.8), against a baseline of roughly 8.0 per 1,000. The generator's exact counterfactual for the same window is 15.1, inside the interval.
  2. Do not judge MFA by a before/after chart. Attack volume roughly doubled mid-rollout. Comparing the whole population before vs. during the rollout window shows a change of only +0.2 per 1,000 (it looks like nothing happened, partly because only half the users were treated in that window and the attack surge offsets the benefit); the wave-1-only before/after shows -6.3 per 1,000, which is biased toward zero for the same reason. Across 16 regenerated worlds the wave-1 before/after estimate was off by 5.0 per 1,000 on average (too small an effect), the DiD by 0.2.
  3. Price the friction and the noise. MFA cost 176 user-hours of added challenge latency (9.8 s per challenge) and 2,117 legitimate MFA failures (timeouts, un-enrolled users), i.e. 7.8 legitimate failures per compromise averted (271 averted in this world). DEVICE_TRUST blocked 5,084 legitimate events for 687 attacker events. It needs tuning or an enrolment fix, not more enforcement.
  4. Fix the token-replay gap. All 5 successes on a high-sensitivity resource from an unmanaged device were stolen-SSO-token sessions that skipped device trust (a gap in how this world is modelled, and a realistic one to test for).
  5. Deploy detection with an alert budget, not a threshold. A logistic regression over per-user-day features caught 95% of compromised user-days (95% Wilson interval 83% to 99%) at 10 alerts/day, with 18% precision. The union of six hand-written rules caught 69% at 64 alerts/day and 2.1% precision (about 10.6 analyst-hours/day at an assumed 10 min/alert). Base rate: 0.2% of user-days.

Evidence

Question Method Result
Is the data trustworthy? 19 checks on 7 dimension labels, scored vs planted faults 100% recall on 6 planted faults; 2,082 rows (0.6%) quarantined
Did MFA reduce compromises? Staged-rollout DiD, wave 1 vs not-yet-treated wave 2, days 25 to 44 -12.4 per 1,000 user-days [-15.8, -8.9]; placebo +1.7 [-2.8, 6.2]
Is the estimator sound? 16 regenerated worlds vs exact counterfactual mean error 0.2, RMSE 2.0; interval covered truth in 81% of worlds
Does label quality matter? Same DiD using only "confirmed" labels (44% of compromise user-days) -7.0 vs -12.4 with full labels: the effect is under-stated by 43%
Which detector? Time split: train days 0 to 39, test 40 to 59 table on the project page

What could make this wrong

Next steps (would do with real data)

Run the same pipeline on LANL's multi-source cyber dataset (its red-team file gives ground truth for the detection half; it has no rollout, so the DiD half needs internal data); add an event-study plot and a sensitivity analysis for unobserved shocks; replace hand-set DQ thresholds with drift monitors; measure alert-to-resolution time, not just alert counts.