
One node became a cluster, and no application had to change
A second machine that can be switched off at any moment joined a running single-node cluster without a change to any application. An always-on server keeps the control plane, every database and volume and the first copy of every service; the build machine takes the builds and second copies of stateless services behind a taint, and those copies wait rather than fall back when it sleeps. It powers itself off when the work runs out and comes back when a job is queued, a console shows every node, workload and availability span, and every CPU request on the always-on server was measured, taking its reservations from eighty-four to fifty-six percent of what it can allocate.
The Full Story
How capacity was added to a running estate without a migration, and what it proves about running infrastructure rather than renting it.
One machine did everything, and it ran out of reservations
For its first months the cluster was one always-on server holding everything: the control plane, every database and volume, every public site, the registry and, on every deploy, the build itself. On any node a CPU request is a reservation the scheduler sets aside whether or not the work uses it, and a busy node runs out of reservations long before it runs out of work. The always-on server had reached eighty-four percent of the CPU it could allocate, so the next workload would not fit, however quiet the machine looked.
Adding capacity is simple when the new machine never goes away. The harder version is a second machine that can be switched off at any moment, and it is only safe if the cluster knows exactly what may run on that machine and what must never depend on it. The rule that made it safe fits in one line: nothing that holds data ever runs on a machine that can disappear.
Machines with opposite jobs
One server is never switched off and holds the control plane, every database and volume, the registries and the first copy of every service. The build machine can be off at any moment and takes the builds and the second copies while it is on. A role label and a taint decide what may land where, and nothing lands on the build machine unless it asks to.


Every node's state on one screen, including those that sleep
The console's estate view gives each node a card: its role, what its workloads have requested against what they use, how many pods it carries, and an availability strip over the last day or week with every span of Ready and Away listed. The cards read as opposites by design. On the capture day the always-on server had been ready for the whole of the previous seven days, and the build machine for thirty-seven percent of them, with thirty recorded changes, because it sleeps whenever there is no work.
Roles and a taint, not hope, keep workloads where they belong. The always-on server carries its role and no taint, so ordinary workloads land there. The build machine carries a taint that turns away every pod that has not declared it can live with a node that disappears, so nothing lands on it by accident, and the workloads that keep data on a disk are pinned to the server that disk is in.



Nothing lands on the wrong node, and nothing waits for one
A stateless service that benefits from more capacity gets a second copy rather than a move. Its original stays on the always-on server and serves alone whenever the build machine is off. The second copy may run on the build machine and nowhere else: while that machine sleeps the copy waits, parked, and it never falls back onto the always-on server, whose reservations are the scarce resource. When the machine wakes, the copies start by themselves and rejoin the traffic.
The edge is tuned for a node that vanishes without saying so. A copy on a machine that lost power does not refuse connections, it simply stops answering, so the router gives up on it within a few seconds instead of half a minute and retries once, on page requests only, and a visitor gets the page a moment late rather than a spinner. Eight services had a second copy on 22 September, the website and the public MCP server among them; the databases, the CMS and every backend stay single on the always-on server. The captures show both states: the copies parked while the build machine slept, and the website and the MCP server served from every node while it was awake.



Capacity that sleeps until there is work
The build machine draws sixty to seventy watts doing nothing, so it does not sit there doing nothing. When the work runs out it powers itself off, and when a job is queued it is back in the cluster within a couple of minutes, with no one having to remember to switch it on. The Hosts view puts the machines side by side: load with a day of history, memory, disk, temperature and the build cache against its cap.
The phone view shows the other half of the design. While the build machine is away its last readings stay on screen, marked with their age, so an empty slot reads as a machine resting rather than a fault. What that buys is capacity that is awake when work arrives and off when it does not.



Requests were the scarce resource, so every one was measured
Every request on the always-on server was measured over forty samples thirty seconds apart and set to each workload's peak plus headroom. The server had been reserving nearly nine times what its workloads used; afterwards its reservations fell from eighty-four to fifty-six percent of what it can allocate, room for the next service without another machine.



The next machine is a role, not a project
No application changed along the way. Placement lives in a role label, a taint, affinities and each service's own overlay, so the services never learned how many machines there are. That is also what makes the next step routine: the next machine joins by being given a role and a share of the build runners, a service opts into a second copy from its own repository, and another cluster on the same pattern repeats the same steps rather than needing a new design. It scales in either direction, because switching the build machine off loses nothing and needs no drain; only the extra speed goes with it.
The limit is stated rather than discovered: a power cut in the middle of a build fails that build, and it is run again. This suits a business with real hardware that needs more capacity without a migration, and it answers the question a CTO asks before handing over infrastructure, which is whether the person in front of them can run it as well as write code for it. The platform underneath is the story of the self-hosted homelab that runs the studio, and what the build machine did for delivery is the story of replacing hosted CI with a home runner pool.
You May Also Like
When hosted CI became the slowest step, the builds came home
Builds and deploys moved from GitHub's hosted runners, where every run started on a fresh machine and releases waited on queues and incidents, to a runner pool at home: a build machine that takes every job while it is up, and standby runners on an always-on server behind a watchdog that fails towards shipping. Images go to a private container registry and the studio's libraries to a private npm registry, both in the building, and every run is archived past GitHub's retention. The slowest deploy went from eight minutes twenty-four seconds to three minutes five, and CI came out a third faster.
27 Pearls - Mobile Release and Store Submission
A finished 27 Pearls app became approved listings on the App Store and Google Play inside a one-week engagement, without a round of rejections. The week covered release builds and signing, listing copy written around the student's benefit, screenshots for every required device size, privacy declarations that matched what the app collects, and a reviewer account that worked. The app remains on Google Play today.
Fully Booked - Zero-Downtime Migration to Self-Hosted Kubernetes
Fully Booked ran on managed cloud infrastructure with a bill that grew every month and a stack nobody fully owned. We moved the whole platform onto a self-hosted Kubernetes cluster without a minute of downtime for the people using it. The cutover covered DNS, ingress, databases and the release pipeline, staged so that traffic shifted only once each layer had been proven on the new cluster. Nobody using the app that week knew anything had happened, which was the point. The recurring bill became owned hardware, and the same cluster now carries the platform's weekly releases.

The blog is field notes from production: architecture decisions, AI systems that survived contact with real users, and infrastructure that pays for itself. Written from the work, not about it.


