Senior Site Reliability Engineer · AI Infrastructure

I keep GPU inference boring.

I run the platforms that serve models in production: GPU fleets that scale from zero, inference that stays fast under load, and latency budgets that hold when traffic doesn’t. Crypto validators and bare-metal clusters before this. A few OSS tools alongside.