# One node became a cluster, and no application had to change

Company: Webanion
Service: DevOps Engineering
Period: 2026-07 - 2026-09
Tech: Kubernetes, Traefik, Linux, systemd, Go, Grafana, VictoriaMetrics, GitHub Actions
Canonical: https://webanion.com/portfolio/scaling-a-homelab-from-one-node-to-a-multi-node-cluster

A second machine that can be switched off at any moment joined a running single-node cluster without a change to any application. An always-on server keeps the control plane, every database and volume and the first copy of every service; the build machine takes the builds and second copies of stateless services behind a taint, and those copies wait rather than fall back when it sleeps. It powers itself off when the work runs out and comes back when a job is queued, a console shows every node, workload and availability span, and every CPU request on the always-on server was measured, taking its reservations from eighty-four to fifty-six percent of what it can allocate.

## The Full Story

How capacity was added to a running estate without a migration, and what it proves about running infrastructure rather than renting it.

### One machine did everything, and it ran out of reservations

<p>For its first months the cluster was one always-on server holding everything: the control plane, every database and volume, every public site, the registry and, on every deploy, the build itself. On any node a CPU request is a reservation the scheduler sets aside whether or not the work uses it, and a busy node runs out of reservations long before it runs out of work. The always-on server had reached eighty-four percent of the CPU it could allocate, so the next workload would not fit, however quiet the machine looked.</p><p>Adding capacity is simple when the new machine never goes away. The harder version is a second machine that can be switched off at any moment, and it is only safe if the cluster knows exactly what may run on that machine and what must never depend on it. The rule that made it safe fits in one line: nothing that holds data ever runs on a machine that can disappear.</p>

### Machines with opposite jobs

<p>One server is never switched off and holds the control plane, every database and volume, the registries and the first copy of every service. The build machine can be off at any moment and takes the builds and the second copies while it is on. A role label and a taint decide what may land where, and nothing lands on the build machine unless it asks to.</p>

### Every node's state on one screen, including those that sleep

<p>The console's estate view gives each node a card: its role, what its workloads have requested against what they use, how many pods it carries, and an availability strip over the last day or week with every span of Ready and Away listed. The cards read as opposites by design. On the capture day the always-on server had been ready for the whole of the previous seven days, and the build machine for thirty-seven percent of them, with thirty recorded changes, because it sleeps whenever there is no work.</p><p>Roles and a taint, not hope, keep workloads where they belong. The always-on server carries its role and no taint, so ordinary workloads land there. The build machine carries a taint that turns away every pod that has not declared it can live with a node that disappears, so nothing lands on it by accident, and the workloads that keep data on a disk are pinned to the server that disk is in.</p>

### Nothing lands on the wrong node, and nothing waits for one

<p>A stateless service that benefits from more capacity gets a second copy rather than a move. Its original stays on the always-on server and serves alone whenever the build machine is off. The second copy may run on the build machine and nowhere else: while that machine sleeps the copy waits, parked, and it never falls back onto the always-on server, whose reservations are the scarce resource. When the machine wakes, the copies start by themselves and rejoin the traffic.</p><p>The edge is tuned for a node that vanishes without saying so. A copy on a machine that lost power does not refuse connections, it simply stops answering, so the router gives up on it within a few seconds instead of half a minute and retries once, on page requests only, and a visitor gets the page a moment late rather than a spinner. Eight services had a second copy on 22 September, the website and the public MCP server among them; the databases, the CMS and every backend stay single on the always-on server. The captures show both states: the copies parked while the build machine slept, and the website and the MCP server served from every node while it was awake.</p>

### Capacity that sleeps until there is work

<p>The build machine draws sixty to seventy watts doing nothing, so it does not sit there doing nothing. When the work runs out it powers itself off, and when a job is queued it is back in the cluster within a couple of minutes, with no one having to remember to switch it on. The Hosts view puts the machines side by side: load with a day of history, memory, disk, temperature and the build cache against its cap.</p><p>The phone view shows the other half of the design. While the build machine is away its last readings stay on screen, marked with their age, so an empty slot reads as a machine resting rather than a fault. What that buys is capacity that is awake when work arrives and off when it does not.</p>

### Requests were the scarce resource, so every one was measured

<p>Every request on the always-on server was measured over forty samples thirty seconds apart and set to each workload's peak plus headroom. The server had been reserving nearly nine times what its workloads used; afterwards its reservations fell from eighty-four to fifty-six percent of what it can allocate, room for the next service without another machine.</p>

### The next machine is a role, not a project

<p>No application changed along the way. Placement lives in a role label, a taint, affinities and each service's own overlay, so the services never learned how many machines there are. That is also what makes the next step routine: the next machine joins by being given a role and a share of the build runners, a service opts into a second copy from its own repository, and another cluster on the same pattern repeats the same steps rather than needing a new design. It scales in either direction, because switching the build machine off loses nothing and needs no drain; only the extra speed goes with it.</p><p>The limit is stated rather than discovered: a power cut in the middle of a build fails that build, and it is run again. This suits a business with real hardware that needs more capacity without a migration, and it answers the question a CTO asks before handing over infrastructure, which is whether the person in front of them can run it as well as write code for it. The platform underneath is the story of <a href="/portfolio/self-hosted-homelab-that-runs-the-studio-infrastructure">the self-hosted homelab that runs the studio</a>, and what the build machine did for delivery is the story of <a href="/portfolio/replacing-hosted-ci-with-a-home-runner-pool-and-private-registries">replacing hosted CI with a home runner pool</a>.</p>
