Proof of concept on synthetic data. This page is a demonstration, not a production system.
← portfolio
Security data science · authentication telemetry · measurement under limited ground truth

Which security controls actually work — and what do they cost the people they protect?

A self-directed project: a documented telemetry schema, data-quality checks scored against planted faults, a difference-in-differences estimate of what a staged MFA rollout did to account compromises, the friction it cost, and a detection baseline evaluated against both true and censored labels. Built on synthetic data so the true answer is known and every method can be checked against it.

Synthetic data — not real telemetrySelf-directed project, not employment experienceScripts + SQL · runs locally
Why synthetic, and what that does and doesn’t buy
Public sets such as LANL’s Comprehensive, Multi-Source Cyber-Security Events (58 days, ~1.6B events, red-team labels; download needs a request form) are the right thing to run this pipeline on next. I used a generator because it gives what real data never does: the exact counterfactual (what would have happened without MFA), the full labels, and a manifest of planted data faults — so I can measure whether each method is right, not just whether it produces a number. The cost: results describe how the methods behave in this simulated world, not how real attackers or enterprises behave. Every rate is an explicit assumption in generate.py.

Headline numbers

loading…

1 · Schema

Raw table is append-only and keeps bad rows; analytics read a trusted view. Labels live in separate tables so a detector can’t touch them by accident (tested). Full column documentation: docs/SCHEMA.md.

2 · Data-quality checks

CheckDimensionWhat it testsRows flaggedStatus

Scored against the planted faults

Planted faultPlantedDetectedRecall
A bug the checks caught in my own pipeline
My first trusted view repaired obvious typos in outcome (“ok”, “succes”, “SUCCESS”, “fail”). The consistency checks then flagged successes that carried a blocking control and an unsatisfied MFA requirement: the corrupted rows were really failures, and “repairing” them flipped a security outcome. The view now quarantines non-canonical outcomes instead, and a regression test pins it.

3 · What each control blocks — and who

Effectiveness needs a denominator on both sides: how much attacker activity a control stopped, and how much legitimate activity it stopped with it. (Attacker vs. legitimate uses the synthetic truth labels; in real data this is the hard part.)

ControlEvents blockedAttacker eventsLegitimate eventsAttacker share of blocks

4 · Did the MFA rollout reduce compromises?

MFA was enforced for remote-facing apps in two randomised waves (wave 1 at day 25, wave 2 at day 45). Attack intensity roughly doubles on days 20–40 and hits both waves, which is the trap: a naive before/after comparison is biased. Comparing wave 1 to the not-yet-treated wave 2 in the window between waves (difference-in-differences) removes the shared shock.

Wave 1 (MFA enforced day 25)Wave 2 (enforced day 45)Window used for DiD (days 25–44)
Weekly initial compromises per 1,000 active user-days, by rollout wave (synthetic). Small weekly counts: expect noise.
Estimate (change in compromises per 1,000 user-days)ValueReading
How the estimator did across 16 independently generated worlds

5 · Detection baseline: rules vs. logistic regression, and what the label choice does to the score

Unit: user-day. Trained on days 0–39, tested on days 40–59 (time split, no shuffling). Features come from SQL over the trusted view only. Precision here is only meaningful next to the base rate — .

DetectorAlerts / dayPrecision (true compromise)RecallPrecision (observed labels)Benign-oddity FPs

6 · Limits — read before quoting any number

  • Synthetic world. The detector’s strong result is partly because simulated attackers are easier to separate than real ones (e.g. attackers use a few IP classes). Treat the logistic-regression numbers as an upper bound on this design, not a forecast.
  • Small positive counts. The test window has few compromised user-days, so recall and precision carry wide intervals (Wilson 95% intervals shown in the memo).
  • Single generator family. The estimator sweep shows DiD is sound under this design (random waves, parallel trends by construction). Real rollouts are rarely randomised.
  • Friction is a floor. Only added challenge latency and counted failures are modelled; abandonment, help-desk load and shadow IT are not.
  • Labels. “Observed” labels are a censored subset of true ones (~half of campaigns confirmed) to mimic real IR; the truth-based numbers are only possible because this is synthetic.
  • Not employment experience. Nothing here describes any employer’s systems or data.

Files

README · one-page memo · schema · data-quality notes · results.json · dq_report.json · seed_sweep.json · 3,000-row event sample (CSV)

Synthetic data. All IP addresses are RFC1918-private. No real users, hosts or organisations.