Waleed Malik

Platform engineer

I build infrastructure for teams running Kubernetes at scale: networking, GitOps, cluster lifecycle, and the unglamorous stuff that just has to work.

Lately I've been exploring the less magical side of AI: inference, model serving, training, and the actual economics of running models yourself instead of renting them forever.

Off the keyboard, it's football, video games, and too much coffee.

GitHub · LinkedIn · X · Email

Kubernetes LLM Inference Platform

Self-hosted LLM and inference platform layered with real developer/team AI workflows: GPU substrate, vLLM/KServe, GIE routing, LiteLLM budgets, observability, and developer AI workflows.

Best read: The Missing Control Plane Above vLLM