Proof of concept on synthetic data. This page is a demonstration, not a production system.

← home · raw markdown

Measuring security-control effectiveness from authentication telemetry (synthetic)

A self-directed project. All data is synthetic (generated by src/generate.py, fixed seed); it is not employment experience and describes no employer's systems. Every IP address is RFC1918-private, so nothing can point at a real host.

Question: which authentication controls (MFA, device trust, lockout) actually reduce account compromise, what do they cost the legitimate users they sit in front of, and how well can we detect what gets through, when ground truth is limited?

Why synthetic (and the public alternative)

LANL's Comprehensive, Multi-Source Cyber-Security Events dataset (58 days, ~1.6B events, authentication + red-team labels; request form required, ~10 GB) is the right public dataset for the detection half. It has no controlled rollout and only red-team labels, so it can't answer "did MFA work?". A simulator can: it yields the exact counterfactual, full labels, and a manifest of planted data faults, so each method is checked against the truth. The price is that results describe this simulated world only. Every assumption is a named parameter in PARAMS at the top of src/generate.py.

What's here

Path What
src/generate.py Synthetic world: 1,200 users, 60 days, ~370k events; staged MFA rollout in two random waves; credential stuffing, password spray, brute force, MFA fatigue, session hijack; legitimate look-alikes (travel, project bursts, Monday NAT rush); injected data faults.
sql/01_schema.sql, 02_quality.sql, 03_metrics.sql Raw / trusted / quarantine layers and the metric queries.
src/dq.py 19 data-quality checks (7 dimension labels), scored against planted faults.
src/analysis.py Control blocking (attacker vs legitimate), staged-rollout difference-in-differences, friction costs, rule and logistic-regression detection baselines, time-split evaluation vs true and censored labels. Needs numpy only.
src/seed_sweep.py Regenerates the world under 16 seeds and checks the DiD against each world's exact truth.
src/render_memo.py Renders docs/MEMO.md from results, so no number in the memo is hand-typed.
docs/SCHEMA.md, DATA_QUALITY.md, MEMO.md Documented schema; DQ notes with a bug found in my own pipeline; one-page memo.
web/ Static, no-build project page (web/index.html), reads web/data/*.json.
tests/test_pipeline.py 17 tests: determinism, RFC1918-only IPs, counterfactual identity, DQ recall, label-leakage guards, estimator recovers a planted effect, end-to-end.

Run it

# needs one numeric library (numpy) and the standard library's sqlite3
./run_all.sh                           # ~2 min; SKIP_SWEEP=1 to skip the 16-world sweep; also runs the tests (~15 s)
# to view the page: serve web/ with any static file server and open it in a browser

data/telemetry.db (~60 MB) is generated, not committed; web/data/ holds the small result files the page needs.

Method summary

Headline results (default seed; regenerate for exact values, see data/results.json)

Honest limits

Synthetic world, one generator family; the detector's recall is an upper bound; few positives in the test window; friction is a floor; randomised waves are unrealistic; the DQ checks are only proven for faults I planted; the volume check needs enough volume and has a false-alarm rate of about one hour in 1,440. See docs/MEMO.md and the project page.

Next (with real data)

Run the trusted layer, DQ checks and detection baseline on LANL's dataset; add event-study plots and an unobserved-confounding sensitivity analysis; add an LLM-assisted triage evaluation (accuracy, duplication, actionability vs a hand-labelled sample).