Skip to main content
The cluster's nodes and workloads in the console

One node became a cluster, and no application had to change

Webanion brand logoWebanionDevOps EngineeringJul 13, 2026 - Sep 22, 20263 months

A second machine that can be switched off at any moment joined a running single-node cluster without a change to any application. An always-on server keeps the control plane, every database and volume and the first copy of every service; the build machine takes the builds and second copies of stateless services behind a taint, and those copies wait rather than fall back when it sleeps. It powers itself off when the work runs out and comes back when a job is queued, a console shows every node, workload and availability span, and every CPU request on the always-on server was measured, taking its reservations from eighty-four to fifty-six percent of what it can allocate.

KubernetesTraefikLinuxsystemdGoGrafanaVictoriaMetricsGitHub Actions
Md Moniruzzaman Image

Md Moniruzzaman

The Full Story

How capacity was added to a running estate without a migration, and what it proves about running infrastructure rather than renting it.

One machine did everything, and it ran out of reservations

For its first months the cluster was one always-on server holding everything: the control plane, every database and volume, every public site, the registry and, on every deploy, the build itself. On any node a CPU request is a reservation the scheduler sets aside whether or not the work uses it, and a busy node runs out of reservations long before it runs out of work. The always-on server had reached eighty-four percent of the CPU it could allocate, so the next workload would not fit, however quiet the machine looked.

Adding capacity is simple when the new machine never goes away. The harder version is a second machine that can be switched off at any moment, and it is only safe if the cluster knows exactly what may run on that machine and what must never depend on it. The rule that made it safe fits in one line: nothing that holds data ever runs on a machine that can disappear.

Machines with opposite jobs

One server is never switched off and holds the control plane, every database and volume, the registries and the first copy of every service. The build machine can be off at any moment and takes the builds and the second copies while it is on. A role label and a taint decide what may land where, and nothing lands on the build machine unless it asks to.

Brand colour backdrop
Drawn placement diagram: the always-on server holds the control plane, every database and volume, the registries and the first copy of each service; the tainted build machine takes the runners and second copies, which wait rather than fall back

Every node's state on one screen, including those that sleep

The console's estate view gives each node a card: its role, what its workloads have requested against what they use, how many pods it carries, and an availability strip over the last day or week with every span of Ready and Away listed. The cards read as opposites by design. On the capture day the always-on server had been ready for the whole of the previous seven days, and the build machine for thirty-seven percent of them, with thirty recorded changes, because it sleeps whenever there is no work.

Roles and a taint, not hope, keep workloads where they belong. The always-on server carries its role and no taint, so ordinary workloads land there. The build machine carries a taint that turns away every pod that has not declared it can live with a node that disappears, so nothing lands on it by accident, and the workloads that keep data on a disk are pinned to the server that disk is in.

Brand colour backdrop
The estate view's Nodes tab: the tabs for nodes, workloads, pods, volumes, edge, schedules, registry and namespaces, and one card per node with its role, CPU and memory requested against in use, pods, and a seven day availability strip
The build machine's availability over seven days: thirty-seven percent ready, thirty changes recorded, and its latest Ready and Away spans with their times

Nothing lands on the wrong node, and nothing waits for one

A stateless service that benefits from more capacity gets a second copy rather than a move. Its original stays on the always-on server and serves alone whenever the build machine is off. The second copy may run on the build machine and nowhere else: while that machine sleeps the copy waits, parked, and it never falls back onto the always-on server, whose reservations are the scarce resource. When the machine wakes, the copies start by themselves and rejoin the traffic.

The edge is tuned for a node that vanishes without saying so. A copy on a machine that lost power does not refuse connections, it simply stops answering, so the router gives up on it within a few seconds instead of half a minute and retries once, on page requests only, and a visitor gets the page a moment late rather than a spinner. Eight services had a second copy on 22 September, the website and the public MCP server among them; the databases, the CMS and every backend stay single on the always-on server. The captures show both states: the copies parked while the build machine slept, and the website and the MCP server served from every node while it was awake.

Brand colour backdrop
The Workloads tab while the build machine was away: the website and the public MCP server served by one copy on the always-on server, their second copies parked, the CMS, the console and the databases pinned, and each row with its run, version and place
Recent deploys with where each image serves: the public MCP server and the website each on the always-on server and, as a second copy, on the build machine

Capacity that sleeps until there is work

The build machine draws sixty to seventy watts doing nothing, so it does not sit there doing nothing. When the work runs out it powers itself off, and when a job is queued it is back in the cluster within a couple of minutes, with no one having to remember to switch it on. The Hosts view puts the machines side by side: load with a day of history, memory, disk, temperature and the build cache against its cap.

The phone view shows the other half of the design. While the build machine is away its last readings stay on screen, marked with their age, so an empty slot reads as a machine resting rather than a fault. What that buys is capacity that is awake when work arrives and off when it does not.

Brand colour backdrop
The Hosts view: the build machine's live readings, CPU load with a day of history, memory, disk, temperature and its build cache against the cap, with the always-on server's card below
The Hosts view on a phone while the build machine is away: its last readings stay on screen, marked two hours old, with the always-on server's card starting below

Requests were the scarce resource, so every one was measured

Every request on the always-on server was measured over forty samples thirty seconds apart and set to each workload's peak plus headroom. The server had been reserving nearly nine times what its workloads used; afterwards its reservations fell from eighty-four to fifty-six percent of what it can allocate, room for the next service without another machine.

Brand colour backdrop
Drawn chart of the always-on server's CPU as shares of what it can allocate: eighty-four reserved before the measuring pass, fifty-six after, and about ten used by its workloads, measured on 19 September 2026
The always-on server's card on the capture day: CPU and memory requested against what is in use, and its pods

The next machine is a role, not a project

No application changed along the way. Placement lives in a role label, a taint, affinities and each service's own overlay, so the services never learned how many machines there are. That is also what makes the next step routine: the next machine joins by being given a role and a share of the build runners, a service opts into a second copy from its own repository, and another cluster on the same pattern repeats the same steps rather than needing a new design. It scales in either direction, because switching the build machine off loses nothing and needs no drain; only the extra speed goes with it.

The limit is stated rather than discovered: a power cut in the middle of a build fails that build, and it is run again. This suits a business with real hardware that needs more capacity without a migration, and it answers the question a CTO asks before handing over infrastructure, which is whether the person in front of them can run it as well as write code for it. The platform underneath is the story of the self-hosted homelab that runs the studio, and what the build machine did for delivery is the story of replacing hosted CI with a home runner pool.

You May Also Like

Webanion brand logo

When hosted CI became the slowest step, the builds came home

Builds and deploys moved from GitHub's hosted runners, where every run started on a fresh machine and releases waited on queues and incidents, to a runner pool at home: a build machine that takes every job while it is up, and standby runners on an always-on server behind a watchdog that fails towards shipping. Images go to a private container registry and the studio's libraries to a private npm registry, both in the building, and every run is archived past GitHub's retention. The slowest deploy went from eight minutes twenty-four seconds to three minutes five, and CI came out a third faster.

Md Moniruzzaman Image

Md Moniruzzaman

Sep 24, 2026

27 Pearls App Logo

27 Pearls - Mobile Release and Store Submission

A finished 27 Pearls app became approved listings on the App Store and Google Play inside a one-week engagement, without a round of rejections. The week covered release builds and signing, listing copy written around the student's benefit, screenshots for every required device size, privacy declarations that matched what the app collects, and a reviewer account that worked. The app remains on Google Play today.

Md Moniruzzaman Image

Md Moniruzzaman

May 05, 2023

Dr. Kris Valenza
fully booked brand logo

Fully Booked - Zero-Downtime Migration to Self-Hosted Kubernetes

Fully Booked ran on managed cloud infrastructure with a bill that grew every month and a stack nobody fully owned. We moved the whole platform onto a self-hosted Kubernetes cluster without a minute of downtime for the people using it. The cutover covered DNS, ingress, databases and the release pipeline, staged so that traffic shifted only once each layer had been proven on the new cluster. Nobody using the app that week knew anything had happened, which was the point. The recurring bill became owned hardware, and the same cluster now carries the platform's weekly releases.

Md Moniruzzaman Image

Md Moniruzzaman

Jun 23, 2026

+1
Dominic Dormer
Read the Md Moniruzzaman blog

The blog is field notes from production: architecture decisions, AI systems that survived contact with real users, and infrastructure that pays for itself. Written from the work, not about it.