Spotlight
Mijndert Stuij
This tutorial shows how to put dev, QA and staging clusters to sleep overnight with KEDA's cron scaler, including the Argo CD gotchas around CRDs, RoleBindings and replica drift.
Nerav Doshi
This article explains how to design a production-grade MCP server for platform teams, with governance, backend clients, tool definitions and auth as four separate layers, plus the RBAC and deployment work needed before it touches a real cluster.
Roman Glushko
This article explains why Ingress is being replaced by the Gateway API, what the new Gateway, Route and Policy resources actually do, and how to choose a gateway controller before migrating.
Ishay Ezagouri
This article explains the two hard parts of backing up databases on Kubernetes: getting a consistent snapshot while writes continue, and patching what the app points at during a Velero restore.
Tools and utilities
Rune is a fast native macOS app for people who debug clusters every day, with unified pod logs, YAML editing with server dry-run, port-forwarding and a command palette.
Kogaro continuously validates Kubernetes config with 60+ checks across reference, resource, security, image, and network domains, catching silent failures before they impact production.
kubectl tree is a kubectl plugin that follows ownerReferences and prints the whole object tree under a Deployment or custom resource, so you can see what created what.
k8s-overcommit Operator is a Kubernetes operator that trims pod CPU and memory requests based on per-workload overcommit classes, so you can pack more onto a cluster without touching the limits.
Kubernetes production readiness checklist is a free, open checklist for taking applications to production on Kubernetes, covering app design, manifests, security, scaling and going live.
Events starting soon
October 1, 2026
Location: Philadelphia, PA, USA
This event requires an entrance fee
Use DOD26KUBEEVENTS to get $30 off
October 1, 2026
This is a virtual event
This is a free event.
October 1, 2026
This is a virtual event
This is a free event.
October 1, 2026
Location: Frankfurt am Main, DE
This is a free event.
October 1, 2026
This is a virtual event
This is a free event.
October 1, 2026
Location: St. Louis, MO, USA
This event requires an entrance fee
Learn from production
Aswin A
This case study shows how a team ran one Kafka cluster stretched across three separate Kubernetes clusters with Strimzi, Submariner and Cilium, and kept it alive when a whole cluster went down.
Ayush Agarwal
This case study shows how a payments company rebuilt deployments around GitOps with Argo CD on Kubernetes, so every production change is a reviewed Git commit and PCI audits stop being guesswork.
Chinedu Mba
This case study shows how to migrate from fragile EC2-based infrastructure to AWS EKS with GitOps using ArgoCD, Jenkins, and SonarQube for a URL shortening platform, achieving 80% faster deployments and self-healing infrastructure.
Harish
This case study shows how a nightly export job kept getting OOM-killed on Kubernetes, and how switching to streaming dropped its memory use from 700 MB to about 30 MB.
Matching jobs
DevOps Engineer with Valarian Technologies Limited
Salary: US$72K to US$324.5K a year
Location: based in the office (and remote from home) in London, GB
Tech stack: Kubernetes, On-premise, Helm, ArgoCD, Flux, Go, Python, Rust, Terraform, Tekton
Infrastructure Architect with North Point Technology
Salary: $95.67K to $253K a year
Location: based in the office in Chantilly, VA, USA
Tech stack: Kubernetes, OpenShift, Java, Python
Machine Learning Engineer with Valarian Technologies Limited
Salary: US$112.5K to US$330K a year
Location: based in the office (and remote from home) in London, GB
Tech stack: Kubernetes, Python
Platform Engineer with New Relic
Salary: $126K to $158K a year
Location: based in the office in Portland, OR, USA
Tech stack: Kubernetes, Helm, ArgoCD, Flux, Go
Platform Engineer with Valarian Technologies Limited
Salary: US$66.6K to US$303.6K a year
Location: based in the office in London, GB
Tech stack: Kubernetes, Bare-metal, GCP, On-premise, Helm, Kustomize, Flux, Docker, Go, Javascript
Build something
Npanchal
This tutorial shows how to chain KEDA cron scaling with Karpenter so non-prod EKS clusters drop to zero nodes at night and on weekends, cutting the compute bill by 65 to 85%.
Kuldeep Paul
This tutorial shows how to run Bifrost, an open-source AI gateway, on Kubernetes with Helm, cluster mode and autoscaling so thousands of LLM requests stay fast under load.
Mahir Berkan Oğuz
This tutorial walks you through a full GitOps setup on minikube, where Argo CD deploys a Go service straight from Git and an HPA scales the pods under load.
This tutorial shows how to link three GKE clusters with Linkerd multicluster, mixing federation and mirroring so a service keeps answering even when a whole cluster disappears.
More articles
Vtulasinadh
This article lays out a GitOps reference architecture for running one app across 25 regions, where every region reconciles itself with Argo CD and a rollback is a git revert scoped to that region alone.
Varun Varde
This article explains how to make Kubernetes spend traceable to each team, using a consistent label taxonomy, Prometheus recording rules on resource requests, policy enforcement and chargeback reporting.
Abhishek Joshi
This article uses a food delivery startup story to explain how a .NET API on one EC2 box moves onto EKS, with Docker, Kubernetes objects and Terraform each given a plain-language role.
Sagar Parmar
This article explains why the default Kubernetes scheduler struggles with AI training jobs, and how Volcano adds gang scheduling, queues and GPU topology awareness on top of it.