Cover image for LiveWyer blog post: Kubernetes is complex on purpose
Engineering • 17min read

Kubernetes is complex on purpose

Most of what gets called Kubernetes bloat is the work of running distributed systems. Here is how to tell when a cluster will pay off, and when it never will.

Written by:

Avatar Louise Champ Louise Champ

Published on:

Last updated on:

Every few months, a post titled some version of “We ripped out Kubernetes and never looked back” tops the aggregators. By the end of the afternoon, it has been pasted into someone’s architecture channel with a simple one-line commentary: do we need all of this?

To be fair, it is a valid piece to question. After all, a lot of people do find Kubernetes hard to learn, hard to operate, and unforgiving when misconfigured.

But the next step is where the reasoning goes wrong. The complexity is in and of itself read as an unintended consequence, and that a simpler tool can do the same job with less of it. Much of it is actually the complexity of running distributed services, made visible in an API. We’ve made a separate argument before that simply running a workload in a cluster does not make that workload in itself cloud native.

So in this post, we are going to argue from how Kubernetes works under the hood. We have been building and operating Kubernetes platforms since its early days, and we will happily tell a team not to use it when the use case does not call for it.

In short, for anyone skimming:

  • Most of what gets called Kubernetes bloat is essential complexity. It is the work of running distributed systems, made explicit in one API.
  • The rest is usually chosen by the team. Kubernetes will express almost any structure you can imagine needing, and it will not stop you building more than you need.
  • The costs are mostly fixed, so what decides it is how many workloads you spread them across. By the twentieth workload, adding another costs close to nothing.
  • Managed Kubernetes stops at the cluster boundary, even in its most managed tiers. Everything you build on top, from certificates to observability, is still your team’s to run, and to keep compatible with each new minor version.
  • Every extra cluster pays those costs all over again, so draw your boundaries as low as the requirement allows. Before creating a cluster, try to name what would go wrong if it were a namespace instead. If nobody can, then it should be a namespace.
  • None of this means you have to run a cluster. If you are still running three stateless services after three years, you are never going to reach the point where the model pays off.

Better tooling only removes the accidental part

In his 1986 essay No Silver Bullet, Fred Brooks split complexity into two parts.

In it, he said that essential complexity is inherent in the problem. The complexity which is added by our tools is termed as accidental complexity. Better tooling can drive the accidental part towards zero, but nothing can be done about the essential part, as this is why the software has been built to address it.

Kubernetes itself needs a third category which Brooks had no reason to call out. Let’s call it chosen complexity: that is, added by a team’s own architectural decisions. Kubernetes will readily soak it up, because the model will express almost any structure a team imagines it needs.

Two forms of it keep coming up in our client work. One is a service mesh adopted for mTLS nobody asked for, which puts a proxy in the path of every request and another control plane on the upgrade path. The other is a cluster per team, where a namespace and an RBAC policy would have separated the teams just as well. This one is costly enough to get its own section below.

The essential part, on the other hand, is there whether or not Kubernetes is in the picture:

  • Something has to decide which machine a workload should run on.
  • Something has to notice when a process dies, and then start it again elsewhere.
  • Something has to route traffic to instances whose addresses keep changing.
  • Something has to roll out a new version, without dropping requests, and roll it back if it misbehaves.
  • Something has to hold the desired state, so that a reboot or a partition can converge back to what was asked for.

Kubernetes is complex largely because it tackles all of these in just one model. If a different tool looks simpler, that is usually because it has left part of the list for the Platform team to handle out of band.

One Deployment replaces five separate systems

A Deployment is pretty much the first thing anyone learns in Kubernetes.

When you declare one, you are handing a controller the job of keeping a set number of replicas running: it replaces a pod when one dies, reschedules when a node fails, brings a new image up bit by bit while retiring the old pods, and stops if the rollout stalls.

Put a Service in front of it and traffic gets balanced across whichever pods are healthy, so the rest of the system can talk to one stable name while the instances underneath come and go.

To achieve these, under the hood they require a supervisor, a health checker, a scheduler with failover, a rolling deployment system, and a dynamic load balancer, reconciling against one declared source of truth.

On a virtual machine, each of these is a separate decision: systemd, a health check on a cron entry, a drain-and-update playbook, a load balancer kept in sync by yet another mechanism. Each of these works fine on its own, but none of them know about each other, so the Platform is required to become the integration layer, holding in their head (and documentation) the invariants Kubernetes holds in the API server.

The pattern underneath all of this is reconciliation: look at the current state, compare it to the declared state, take one action to close the gap, and repeat. It is the same loop whether the object is a Deployment, a StatefulSet, or a custom resource for a database.

Fixed costs decide whether a cluster will pay off

The usual argument against Kubernetes goes like this: someone puts a forty-line Deployment manifest next to a git push to a managed platform, sees that both get one web service into production, and decides Kubernetes is overkill.

On day one, this argument does kind of make sense. The trouble is that nobody runs the comparison again 18 months later, when there are 20+ services instead of one.

Almost every cost in a Kubernetes estate is paid per cluster, as opposed to per workload. You learn the model once, and, more usefully, each new capability you add inherits the machinery that is already there.

Let’s say we add some services. We use cert-manager to issue certificates, we run databases using an operator which can handle its own failover, and we add a controller that syncs secrets out of a vault. All of these benefit from the existing machinery which is wrapped around them:

  • kubectl can describe them.
  • RBAC governs who can change them.
  • Admission control applies to them.
  • The audit log records what happens to them.
  • The pipeline that deploys everything else deploys them too.

Buying the same capabilities as separate managed services means each one comes with its own console, access model, audit trail, and deployment mechanism, none of which talk to each other.

The platform needs an ingress controller, a certificate flow, a delivery pipeline, and a policy layer regardless of whether it carries three services or three hundred. With only three services to spread it across, that is a lot of overhead for each one, but at three hundred, it barely registers.

That ratio is what decides whether a cluster is worth running, and we tend to check three things, the first two of which build on each other:

  1. Enough workloads to divide the costs across. This is a count of the distinct things the platform carries.
  2. Few enough clusters that those costs are not paid repeatedly.
  3. A requirement that rules the managed platforms out. A stateful system they will not run, hardware they do not offer, or a regulator or sovereignty requirement that removes them from the list. A requirement that might turn up one day does not count.

If the third one holds, you need a cluster regardless, and the only question left is how few you can get away with. If it does not, the first two decide it on cost: when both hold, the cluster works out cheaper, and when either one fails, you are better off on a managed platform.

Kubernetes vs ECS: the teams who remove it often rebuild it

Most “we removed Kubernetes” stories follow the same shape: the team was small, the cluster was more machinery than the workload needed, and, for where they were at the time, moving to something lighter was the right call.

Where they usually end up is rarely exotic: Amazon ECS or Fargate, AWS App Runner, Cloud Run, or plain containers on an autoscaling group. A team that moves everything onto ECS has adopted a managed platform, and AWS has done the integration work for it.

However, the problems Kubernetes was solving still exist on ECS, and services still have to find each other as tasks come and go. AWS offers three different ways to do that:

  • Service Connect runs an agent beside each task and replaces clients on deployment.
  • Cloud Map points DNS straight at task addresses. It is the simplest to set up and the one that bites: a record’s TTL can hand a client a task that has already gone, and the retry logic is yours to write.
  • VPC Lattice puts tasks behind a managed networking layer. This is the most capable, but also the heaviest.

Kubernetes has service discovery built in, in three pieces: a Service to give the thing a stable name, an EndpointSlice to track which pods are healthy right now, and a readiness probe to define what healthy means. The first two come free; the readiness probe is the one you have to write yourself, and without it a pod counts as ready the moment its containers start.

On ECS, choosing between the three is left to you, and Cloud Map, the easiest one to set up, is the one that tends to cost the most later on.

Where it goes wrong is a partial move, and that is a more common scenario than a clean one; Move some services go to ECS, keep the awkward ones on virtual machines, and suddenly discovery, deployment, and health checking now span across both. The team now owns the gap between two models, and ends up rebuilding by hand exactly what a cluster used to do for them.

Managed Kubernetes stops at the cluster boundary

When people say “Kubernetes is hard”, they are lumping two costs together.

The first is learning and using the model. The second is running the platform underneath it: etcd, control plane availability, certificate rotation, and the patching schedule.

Most of the complaints in those removal posts are about the second one, and that is the part you can buy your way out of without giving up the Kubernetes mental model. This is exactly what EKS, GKE, and AKS sell.

However, that does not make the cluster somebody else’s problem completely. If we can use EKS as an example, managed node groups are not automatically upgraded when the control plane is, nor are Fargate pods, and add-ons such as the VPC CNI, kube-proxy, and CoreDNS update separately again.

You also do not get much say over the Kubernetes version upgrade cadence, which is:

  • 14 months of standard support per version.
  • Then extended support, automatically, at an extra charge.
  • 12 months later, when extended support ends, AWS upgrades the control plane at a time it declines to predict, without notice, and leaves managed node groups on the old version.

EKS Auto Mode and GKE Autopilot draw that line even further up, and additionally manage the nodes and the core add-ons as well (pod networking, cluster DNS, load balancing, and block storage). At the end of the day, both of these still stop at the cluster boundary.

Most of the expense, however, lies above the cluster itself. A fresh, conformant Kubernetes cluster runs almost nothing which anyone would put in front of their end-users. What turns it into a production platform is everything you bolt on top: Ingress, certificates and DNS, secrets, observability, CI/CD, policy and admission control, and autoscaling.

Each of these services has their own version, cadence, CVE stream, and ways of falling over at 03:00, and sometimes an end-of-life date coming out of nowhere, as teams migrating off ingress-nginx found out this year. Choosing which services to use is what everyone talks about, but most of the effort goes into running them afterwards.

That said, Kubernetes itself did not invent any of this work. A PaaS such as Cloud Run or Heroku needs the same things and includes them in the price; they are just not yours to operate or to break. A cluster, even a managed one, unbundles them, so your team pays for them in engineering hours instead.

Deciding which layers to own and which to adopt is a whole question of its own, and one we discuss in our buying, building, or assembling an internal developer platform blog post.

Demarcate as low as the requirement allows

Of all the chosen complexity, the most expensive is drawing the environment boundaries higher than they need to be.

You can separate workloads at several levels, listed here from the most expensive to the cheapest:

  • A separate cluster, a complete second platform with every component reproduced and maintained.
  • A separate node pool, sharing the control plane and platform components but running workloads on their own hardware.
  • A namespace, with RBAC, resource quotas, and network policy doing the isolating.
  • Labels and selectors, where the requirement is really about identifying workloads.

A separate cluster does buy you some things: a blast radius that contains a control plane failure, the ability to run different versions during an upgrade, and hard isolation where workloads are untrusted or a regulator requires it. Keeping production apart from everything else clears that bar comfortably.

Upgrades can justify a separate cluster too. Plenty of teams decide replacing a cluster is safer than upgrading one in place, and they are often right, though blue/green at the cluster level means standing up a second complete platform and running both while workloads move. Moving the workloads is the easy part; it is everything holding state that does not move quite so obligingly.

What a separate cluster does not buy is most of what it gets used for:

  • Separating teams needs a namespace and an RBAC policy.
  • Attributing cost to each team needs a label, which tools such as OpenCost can report on.
  • Keeping a noisy workload away from a sensitive one needs a quota, a network policy, and possibly a node pool.

The test we would suggest before creating a cluster is to ask what would go wrong if the workload ran in a namespace instead. “A control plane failure would take production down with everything else” and “the regulator requires separate infrastructure” justify a cluster. “Each team wants its own bill”, however, does not, because labelling each team’s workloads already lets you split the bill. If nobody can give a proper answer, then it should be a namespace.

The reasons for creating a cluster are not always technical, either; the decision can also get caught up in corporate politics. When cluster count becomes a KPI, a bigger number starts to look like success, and new cluster builds get encouraged. Every one of those clusters carries the full supporting stack, and every one has to be upgraded and supported.

Getting this wrong hurts in more ways than the multiplied bill. Copies of infrastructure drift, because keeping them identical is a job nobody has been given. Before long, the staging environment is on a different ingress version and has a policy nobody backported, so a release that passes in staging can still fail in production.

Much of that pain goes away if standing up a cluster is not a project in itself. Where the platform layer is itself declared and reconciled, the second cluster is one apply, and so is the tenth, and because every cluster comes from the same source, there is far less room for drift to creep in.

Without that, what you have is a cluster somebody built by hand which nobody can reproduce, and the team only finds this out on the day it needs to build second cluster.

When Kubernetes is overkill

Accepting that the complexity is mostly essential does not mean you have to take all of it on yourself.

A managed platform such as Google Cloud Run, AWS Fargate, or Fly.io can carry the work described above on your behalf. The scheduling, the failover, the rolling deployments, and the load balancing are still happening, they are just behind an interface operated by somebody else.

If you have a handful of services, traffic a managed platform can handle comfortably, and no plans to grow into multi-team territory any time soon, that trade is frequently the right call.

Who carries the complexity?

The complexity of running distributed services does not itself go away, but what a team does get to decide is where it lives, whether that be:

  • inside a platform you run, where each part of it is an object you can declare and inspect,
  • inside a managed service that hides the complexity behind someone else’s interface,
  • or spread across a team’s collective memory as a set of scripts and conventions.

Either of the first two can be the right answer, depending on how many workloads you have and whether anything rules the managed platforms out.

The third is where teams end up when they treat that essential complexity as if it were accidental, and assume a simpler tool will make it go away.

Where to start

Admittedly, the decision is not always made on technical grounds. We have had a client move off Kubernetes onto Heroku, then decide to come back to managed EKS. As their services grew into larger dynos, the premium they paid over the same capacity on AWS grew with them, until running the same services on EKS was the cheaper option, with no change to the architecture either way.

If the operational load has outgrown your team, or you are still weighing up a cluster against a managed platform, a second opinion from people who build both will pay for itself. Our Cloud Platform Engineering work settles the question by building the thing, so your team comes away either running a platform it understands, or with a clear case for not running one at all.

Frequently asked questions

Is Kubernetes overkill for a small team?

For a small team shipping a few services with no plan to grow into multi-team infrastructure, very possibly. A managed PaaS carries the same essential complexity on the team’s behalf and removes the operational load of a cluster. The case strengthens when specific needs appear: stateful or hardware-specific workloads, cost at fleet scale, sovereignty constraints, or many teams sharing one platform.

Is ECS simpler than Kubernetes?

For a team that moves wholesale, largely yes, though it leaves the choice of service discovery to the team. ECS is a managed platform and AWS has done the integration work. The trouble is the partial move, where some services run on ECS and the rest stay on virtual machines, so discovery, deployment, and health checking span both models and the code bridging them is undocumented.

What else do you need to run Kubernetes in production?

More than the cluster. A production platform normally adds ingress, certificate management, DNS, a secrets store, metrics, logs and traces, a delivery mechanism such as Argo CD or Flux, policy and admission control, autoscaling, a CSI driver, and backup. Each carries its own upgrade cadence and failure modes. A managed platform includes equivalents in the price; a cluster hands them over to operate.

Do we need a separate cluster for each environment?

Not necessarily. Production apart from everything else usually earns its own cluster. Beyond that, name the specific thing that would go wrong if an environment were a namespace instead; if nobody can, it should be a namespace. Teams end up with several through some mix of environment separation, regional requirements, tenancy boundaries, and treating cluster replacement as a safer upgrade path. The supporting stack is fixed per cluster rather than per organisation, so cluster count multiplies it. It stays manageable only where the platform layer is declared and reconciled like the workloads.

Thank you for reading

If you enjoy our writing and would like to see more, add us as a Preferred Source on Google and you'll be more likely to see us in your results.

LiveWyer is a London-based consultancy specialising in deploying small, agile engineering teams to design and implement modern Cloud and Platform solutions.

Talk to a specialist