KubeFlow: Twelve Months of Pain
After a year of running KubeFlow across multiple client deployments, I can confidently say: the platform that promises end-to-end ML on Kubernetes delivers end-to-end frustration instead. Here's the post-mortem.
After a year of running KubeFlow across multiple client deployments, I can confidently say: the platform that promises end-to-end ML on Kubernetes delivers end-to-end frustration instead. Here's the post-mortem.
Vendor consultants optimise for showcasing their platform's capabilities, not for your team's ability to maintain it. They build reference architectures, not products. The demo passes. Production doesn't.
Every organisation wanted an MLOps platform circa 2020. Most had data scientists running Jupyter notebooks on their laptops and calling it a pipeline. Building an airport for one flight a week.
Nobody is running workloads portably across clouds. They're running different things on different clouds and calling it a strategy. Multi-cloud is what happens when nobody's in charge.
Everyone adopted Istio because everyone else was adopting Istio. Complex, resource-hungry, and solving problems most teams didn't have — it was the poster child for resume-driven development.
Service Fabric got a lot right — stateful services, the actor model, rolling upgrades with health checks. But it lost the ecosystem war to Kubernetes, and being technically superior wasn't enough to save it.