An operator's notes on scaling systems, building platforms and teams, and how engineering changes with AI agents. Written from the production floor at Probo, Gojek and now KoshaX.
Sudden, spiky, event-driven traffic. How Probo stayed up when all of India opened the app at the same moment.
Human approval, bounded briefs and context layers for supervising multi-agent software work.
Google plays by different rules on Android. Why I'm building an assistant for India anyway.
Heroku, the AWS CLI, K9s and Claude Code: from exposing complexity to structuring it.
I don't type most of my code anymore. The joy was never in the typing.
Tools come and go. What compounds, what doesn't, and how to last in a career built on change.
The biggest wins came from deleting APIs, trimming features and writing match-day playbooks.
Karpenter in 2023: nodes in 40 seconds instead of 2 minutes, 40% cheaper compute, and a lot of IAM pain.
Redis kept p99 under 100 ms at peak. It also hid our tech debt, until it didn't.
A shared Google Calendar and an Airflow DAG let marketing schedule our infrastructure scaling.
EC2 autoscaling was too slow for match-day spikes and Kubernetes wasn't ready. So we built a switcher.
In India, a cricket match is a stress test. What we learned keeping Probo up through it.
A fraud rule engine that paid off, a meme engine that didn't, and knowing the difference.
Teams that know the why, the ideal state and the trade-offs don't wait to be told what's next.
Why the same few people always answered the alerts, and how we made reliability everyone's job.
Internal platforms succeed or fail on developer experience. Treat it as the core product.
Long feedback loops, becoming the bottleneck, becoming the dumping ground, and what helped.
How about 30 engineers kept Probo on one monolith from zero to 250Mn+ trades a day, and what it cost us.