Sandeep Sidhu

About

Sandeep Sidhu

I'm Sandeep Sidhu, a site reliability engineer in the UK. I've spent 20+ years running production infrastructure - everything from crimping cat5 cables and racking servers, to helping run tens of thousands of hypervisors on a public cloud, to SRE work on large Kubernetes and AWS estates at companies like GoDaddy, Sky, and NET-A-PORTER.

These days I work at the intersection of three things: security, observability, and AI. I run a multi-region production fleet where AI agents do real work - deploys, diagnostics, change management - and I'm the founder of AlertKick, a security monitoring and compliance platform built for teams that don't have a dedicated security function. Almost everything I write comes straight from that work: eBPF detections, SSH lockdowns, monitoring that catches real problems, and what actually happens when an AI agent has hands on your servers.

If you run production without a security team, want monitoring you don't have to babysit, or are working out where AI agents fit in your workflows, you're exactly who I write for. I offer a free 30-minute infrastructure and security review if you want a second pair of eyes on where you stand.

The blog goes back to 2010, so the older posts are a time capsule of the early cloud era. I keep them up because that's the point: this stuff is learned by running systems for a long time.

Thanks for dropping in :)

X / Twitter · LinkedIn · GitHub