Articles
Notes on backend architecture, performance work, and the WordPress and Laravel systems I build.
13 min readYour cron ran twice: why exactly-once scheduling does not exist
A job runs at midnight for four years, then runs twice with no deploy. The prime suspect is a daylight-saving fall-back, but there are eight others — clock steps, duplicate schedulers, catch-up controllers, lost acks. How to tell them apart from the logs in ten minutes, and why the real fix is a run key with a unique index rather than a better schedule.
Intermediate- Scheduling
- Reliability
- Distributed Systems
- Operations
12 min readYour authenticator works in airplane mode: how TOTP really works
The 6-digit code is never transmitted — it is derived. A full walk through RFC 6238: the shared secret exchanged once at enrolment, the time step, HMAC-SHA1 and dynamic truncation, the verification window and clock drift, replay prevention and rate limiting, plus the attacks TOTP does and does not stop.
Intermediate- Security
- Authentication
- Cryptography
- MFA
12 min readThe auth service is down and users are still logging in: how stateless auth works
A backend interview question with a precise answer — the auth service sits on the issuance path, not the verification path. How locally verified signatures, JWKS caching and asymmetric keys keep requests flowing, exactly which operations are already broken, the revocation you traded away, and how to design the degradation deliberately.
Intermediate- Authentication
- Security
- Reliability
- System Design
14 min readOne product page, 201 queries: fixing N+1 and making it visible in review
The N+1 query problem, answered end to end — why 201 fast queries add up to 9 seconds while the slow log stays empty, the fix ladder from eager loading to batch loaders and counter caches, the traps (cartesian joins, per-parent limits, huge IN lists), and how to turn query count from a runtime accident into something CI fails on.
Intermediate- Databases
- Performance
- Laravel
- ORM
14 min readThe cache scheduled the outage: stampedes, avalanches and how to survive midnight
A backend interview classic — 10,000 rps in front of a 1,000 rps database, one cache, one shared TTL, and a database that dies at 00:00:00. What the failure is actually called, why it does not recover on its own, and the full ladder of fixes: TTL jitter, single-flight coalescing, stale-while-revalidate, probabilistic early expiry and load shedding.
Intermediate- Caching
- Reliability
- System Design
- Redis
17 min read800 million requests, one corrupted row: debugging what you cannot reproduce
A senior backend interview question, answered as an investigation. Why "1 in 800 million" and "cannot reproduce" are themselves evidence, how to read the shape of the corruption, how to reconstruct the timeline from binlog and WAL, how to turn one bad row into a population you can query — and how to make the whole bug class impossible afterwards.
Advanced- Debugging
- Concurrency
- Databases
- Reliability
16 min readSELECT COUNT(*) is slow on 30M rows: how to make counting fast at scale
Why COUNT(*) degrades on a 30M-row table under MVCC, and the full ladder of fixes — keyset pagination, planner estimates, capped counts, covering indexes, sharded summary counters, rollup tables and OLAP replicas — with the MySQL and PostgreSQL code for each and a table for picking the right one.
Intermediate- Databases
- Performance
- MySQL
- PostgreSQL
15 min readMillions of rows nobody uses: how to delete them without breaking production
You found a table with millions of rows and no obvious owner. This is the playbook: prove nothing reads it, make every step reversible before any step is destructive, then delete in batches your replicas can survive — with the MySQL and PostgreSQL queries that produce the evidence.
Intermediate- Databases
- MySQL
- PostgreSQL
- Operations
10 min readOne movie, 190 countries, zero buffering: how Netflix answers the interview
A system-design case study of the classic Netflix question — 120 million simultaneous streams, 4–15 Mbps each, no buffering. We do the napkin math, show why a data centre cannot serve it, and walk through Open Connect, adaptive bitrate and the control-plane / data-plane split that actually make it work.
Intermediate- System Design
- Networking
- CDN
- Streaming