Which security controls actually work — and what do they cost the people they protect?
A self-directed project: a documented telemetry schema, data-quality checks scored against planted faults, a difference-in-differences estimate of what a staged MFA rollout did to account compromises, the friction it cost, and a detection baseline evaluated against both true and censored labels. Built on synthetic data so the true answer is known and every method can be checked against it.
Headline numbers
1 · Schema
Raw table is append-only and keeps bad rows; analytics read a trusted view. Labels live in separate tables so a detector can’t touch them by accident (tested). Full column documentation: docs/SCHEMA.md.
2 · Data-quality checks
| Check | Dimension | What it tests | Rows flagged | Status |
|---|
Scored against the planted faults
| Planted fault | Planted | Detected | Recall |
|---|
3 · What each control blocks — and who
Effectiveness needs a denominator on both sides: how much attacker activity a control stopped, and how much legitimate activity it stopped with it. (Attacker vs. legitimate uses the synthetic truth labels; in real data this is the hard part.)
| Control | Events blocked | Attacker events | Legitimate events | Attacker share of blocks |
|---|
4 · Did the MFA rollout reduce compromises?
MFA was enforced for remote-facing apps in two randomised waves (wave 1 at day 25, wave 2 at day 45). Attack intensity roughly doubles on days 20–40 and hits both waves, which is the trap: a naive before/after comparison is biased. Comparing wave 1 to the not-yet-treated wave 2 in the window between waves (difference-in-differences) removes the shared shock.
| Estimate (change in compromises per 1,000 user-days) | Value | Reading |
|---|
5 · Detection baseline: rules vs. logistic regression, and what the label choice does to the score
Unit: user-day. Trained on days 0–39, tested on days 40–59 (time split, no shuffling). Features come from SQL over the trusted view only. Precision here is only meaningful next to the base rate — .
| Detector | Alerts / day | Precision (true compromise) | Recall | Precision (observed labels) | Benign-oddity FPs |
|---|
6 · Limits — read before quoting any number
- Synthetic world. The detector’s strong result is partly because simulated attackers are easier to separate than real ones (e.g. attackers use a few IP classes). Treat the logistic-regression numbers as an upper bound on this design, not a forecast.
- Small positive counts. The test window has few compromised user-days, so recall and precision carry wide intervals (Wilson 95% intervals shown in the memo).
- Single generator family. The estimator sweep shows DiD is sound under this design (random waves, parallel trends by construction). Real rollouts are rarely randomised.
- Friction is a floor. Only added challenge latency and counted failures are modelled; abandonment, help-desk load and shadow IT are not.
- Labels. “Observed” labels are a censored subset of true ones (~half of campaigns confirmed) to mimic real IR; the truth-based numbers are only possible because this is synthetic.
- Not employment experience. Nothing here describes any employer’s systems or data.
Files
README · one-page memo · schema · data-quality notes · results.json · dq_report.json · seed_sweep.json · 3,000-row event sample (CSV)