Spotlight
Netflix Technology Blog
This case study shows how Netflix moved millions of batch jobs from its homegrown queueing system to Kueue, without the people submitting those jobs noticing any change.
Deepakravi
This case study shows how to stabilize Harbor on VMware VKS by expanding storage, upgrading the Supervisor Service, and configuring Trivy scanning to receive vulnerability results.
EXANTE Technology
This case study shows how EXANTE replaced manual Saturday releases with a fully automated GitLab CI + Flux + Jira pipeline across 60+ Django modules, 7 GKE environments, and 30+ services to meet fintech regulatory audit requirements.
Saeedafzal2030
This case study shows how one team ran LiteLLM as a single gateway to many model providers on EKS, kept it highly available, and managed the whole thing with ArgoCD.
Tools and utilities
Warden is a secure gateway that brokers connections between AI agents and enterprise systems by authenticating agent identity and injecting short-lived credentials at request time.
Praesto is a Kubernetes-native operator and CSI driver that automates the caching and mounting of AI model artifacts from Hugging Face into workloads using local node storage or shared PVCs.
A native iOS and Android app for monitoring and managing Argo CD deployments from your phone.
l9gpu is a GPU telemetry agent that emits OpenTelemetry metrics with workload attribution built in, so you can see which pod, team or Slurm job is using each NVIDIA, AMD or Intel GPU.
CruiseKube is a Kubernetes-native resource optimization controller that automates CPU and memory right-sizing for workloads at runtime and admission time.
Events starting soon
August 27, 2026
Location: Barton, AU
This is a free event.
August 27, 2026
This is a virtual event
This event requires an entrance fee
August 27, 2026
Location: Austin, TX, USA
This is a free event.
August 31, 2026
Location: Seoul, KR
This is a free event.
September 1, 2026
Location: San Francisco, US
This event requires an entrance fee
September 2, 2026
Location: Melbourne, AU and virtual
This is a free event.
Kube Signals starts where the keynote ends: with the trends that platform teams will have to operationalize next.
In this special episode, Brian Teller speaks with Saiyam Pathak about his KubeCon India keynote and the shift from developer platforms to AI factories. They examine what GPU scarcity, shared accelerators, and AI workloads mean after the conference slides meet real infrastructure.
In this interview:
Learn from production
Jacob Amar
This case study traces why Kafka started reading from disk after moving from EC2 to EKS, and lands on cgroup v2 reclaim pressure.
Anastassios Nanos
This case study shows how etcd CrashLoopBack pods on a Karmada/k3s demo cluster turned out to be a ZFS I/O latency problem, and how four ZFS tuning settings fixed it — including the exact etcd Prometheus metrics to watch.
Krishnakanth E
This case study explains how an ECS workload was migrated to EKS across two AWS regions with zero production downtime.
It covers KEDA autoscaling, HashiCorp Vault, IRSA, disaster recovery, and a coordinated production cutover.
Matching jobs
DevOps Engineer with Mark43
Salary: $155K to $170K a year
Location: remote from
Tech stack: Kubernetes, Docker, Terraform
DevOps Engineer with RobCo
Salary: US$70.74K to US$440K a year
Location: based in the office in Munich, DE
Tech stack: Kubernetes, AWS, Go, Python, Terraform, Datadog, Grafana, Prometheus
Site Reliability Engineer with MyFitnessPal
Salary: $120K to $165K a year
Location: remote from
Tech stack: Kubernetes, AWS, Docker, Go, Python, Typescript, Terraform, GitHub Actions, Datadog
AI Enterprise Technical Program Manager with Redhorse Corporation
Salary: $37 to $485.65K a year
Location: based in the office in Arlington, VA, USA
Tech stack: Kubernetes, AWS, Azure, Docker, Spark
Commercial Account Executive with Vantage
Salary: $100K to $200K a year
Location: remote from
Tech stack: Kubernetes, AWS, CircleCI, Datadog
Build something
Osura Viduranga
This tutorial walks through deploying the WSO2 API Platform Gateway on OpenShift, from building the images and installing cert-manager to applying the gateway operator, a sample REST API, and the route that exposes it.
Dharmendra Yadav
This tutorial explains why Gateway API replaces Ingress on GKE, walking through GatewayClass, Gateway, and HTTPRoute, then sets up traffic splitting and TLS without vendor-specific annotations.
Matthew Wimpelberg
This tutorial builds a real k6 test suite against Google's Online Boutique running on a home lab Kubernetes cluster, with a shared client library, smoke, load, stress, and browser tests, and the two bugs the first run exposed.
Timur Nizamutdinov
This tutorial shows how to replace always-on EC2 runners with autoscaling GitLab runners using the AWS fleeting plugin and an Auto Scaling Group, so instances start per job and shut down when the queue is empty.
Call for Papers closing soon
1
days
This is a virtual event
Online conference organized by ScyllaDB.
The conference starts on the 11 March 2027.
1
days
Location: Las Vegas, NV, USA
In-person conference organized by Dynatrace.
The conference starts on the 16 February 2027.
4
days
Location: Warsaw, PL
In-person conference organized by Devopsdays.
The conference starts on the 23 November 2026.
4
days
Location: Recife, BR
In-person conference organized by Devopsdays.
The conference starts on the 12 December 2026.
4
days
Location: Florianópolis, BR
In-person conference organized by Devopsdays.
The conference starts on the 24 October 2026.
4
days
Location: Bucharest, RO
In-person conference organized by OmniOpenCon.
The conference starts on the 18 October 2026.
4
days
Location: Boston, MA, USA
In-person conference organized by Devopsdays.
The conference starts on the 19 October 2026.
More articles
mirusser
This article argues that MCP servers for Kubernetes ship the accelerator without the brake, and walks through a two-tier design in which an agent can only propose a change, and a human approves the exact plan before it is applied.
zolty
This case study covers migrating a live k3s cluster from a flat network to a VLAN architecture, including an etcd quorum loss caused by moving too many nodes at once and the recovery steps using k3s server --cluster-reset.
Rute C. Sofia
This article explains the centralise-or-not trade-off behind CODECO's federated Kubernetes architecture across cloud, edge and IoT, and how neighbourhood-scoped control, decentralised AI and federated scheduling answer it.
Sujith K. Surendran
This article explains why CPU is the wrong scaling signal for Pub/Sub consumers, since pods look healthy while the backlog grows, and argues for scaling on queue depth with KEDA instead.