Memo: Are our authentication controls working, and what do they cost?
To: Security leadership (hypothetical) From: Max Mattox Status: self-directed project on synthetic data (371,226 events, 60 days, 1,200 users). Not employment experience; describes no employer's systems.
Recommendation
- Keep the staged MFA rollout and finish it. It reduced account compromises by about 12.4 per 1,000 active user-days (difference-in-differences; 95% interval 8.9 to 15.8), against a baseline of roughly 8.0 per 1,000. The generator's exact counterfactual for the same window is 15.1, inside the interval.
- Do not judge MFA by a before/after chart. Attack volume roughly doubled mid-rollout. Comparing the whole population before vs. during the rollout window shows a change of only +0.2 per 1,000 (it looks like nothing happened, partly because only half the users were treated in that window and the attack surge offsets the benefit); the wave-1-only before/after shows -6.3 per 1,000, which is biased toward zero for the same reason. Across 16 regenerated worlds the wave-1 before/after estimate was off by 5.0 per 1,000 on average (too small an effect), the DiD by 0.2.
- Price the friction and the noise. MFA cost 176 user-hours of added challenge latency (9.8 s per challenge) and 2,117 legitimate MFA failures (timeouts, un-enrolled users), i.e. 7.8 legitimate failures per compromise averted (271 averted in this world). DEVICE_TRUST blocked 5,084 legitimate events for 687 attacker events. It needs tuning or an enrolment fix, not more enforcement.
- Fix the token-replay gap. All 5 successes on a high-sensitivity resource from an unmanaged device were stolen-SSO-token sessions that skipped device trust (a gap in how this world is modelled, and a realistic one to test for).
- Deploy detection with an alert budget, not a threshold. A logistic regression over per-user-day features caught 95% of compromised user-days (95% Wilson interval 83% to 99%) at 10 alerts/day, with 18% precision. The union of six hand-written rules caught 69% at 64 alerts/day and 2.1% precision (about 10.6 analyst-hours/day at an assumed 10 min/alert). Base rate: 0.2% of user-days.
Evidence
| Question | Method | Result |
|---|---|---|
| Is the data trustworthy? | 19 checks on 7 dimension labels, scored vs planted faults | 100% recall on 6 planted faults; 2,082 rows (0.6%) quarantined |
| Did MFA reduce compromises? | Staged-rollout DiD, wave 1 vs not-yet-treated wave 2, days 25 to 44 | -12.4 per 1,000 user-days [-15.8, -8.9]; placebo +1.7 [-2.8, 6.2] |
| Is the estimator sound? | 16 regenerated worlds vs exact counterfactual | mean error 0.2, RMSE 2.0; interval covered truth in 81% of worlds |
| Does label quality matter? | Same DiD using only "confirmed" labels (44% of compromise user-days) | -7.0 vs -12.4 with full labels: the effect is under-stated by 43% |
| Which detector? | Time split: train days 0 to 39, test 40 to 59 | table on the project page |
What could make this wrong
- It is a simulation. Simulated attackers are easier to separate than real ones; treat the detector's recall as an upper bound for this design. Every rate is an assumption in
generate.py. - Few positives. The test window has 39 compromised user-days; intervals are wide (Wilson intervals on the project page).
- Randomised waves. Real rollouts are seldom randomised, so parallel-trends would have to be argued, not assumed. The placebo check and event-study plots would be the first things to add.
- Bootstrap over users ignores day-level clustering of attacks; interval coverage of 81% (nominal 95%, n=16 worlds) suggests it is a little optimistic.
- Friction is a floor. Abandonment, help-desk load and workarounds are not modelled.
- Labels. "Observed" labels are a censored subset. In real data even the "true" denominator is unknown; capture-recapture or red-team injection would be needed.
Next steps (would do with real data)
Run the same pipeline on LANL's multi-source cyber dataset (its red-team file gives ground truth for the detection half; it has no rollout, so the DiD half needs internal data); add an event-study plot and a sensitivity analysis for unobserved shocks; replace hand-set DQ thresholds with drift monitors; measure alert-to-resolution time, not just alert counts.