
When hosted CI became the slowest step, the builds came home
Builds and deploys moved from GitHub's hosted runners, where every run started on a fresh machine and releases waited on queues and incidents, to a runner pool at home: a build machine that takes every job while it is up, and standby runners on an always-on server behind a watchdog that fails towards shipping. Images go to a private container registry and the studio's libraries to a private npm registry, both in the building, and every run is archived past GitHub's retention. The slowest deploy went from eight minutes twenty-four seconds to three minutes five, and CI came out a third faster.
The Full Story
What a hosted pipeline costs a business in waiting, and what changed when builds, images and packages moved under one roof.
The pipeline was somebody else's machine, and somebody else's bad day
Until the middle of September 2026 every build and deploy ran on GitHub's hosted runners, and a hosted runner is a fresh virtual machine every time: the base image, every dependency and every build layer downloaded again for a one-line change, then thrown away. The image was then built on the always-on production server, because that was the only place the result could run, so the machine serving the databases and the public sites was compiling while it did. The slowest deploy took eight minutes twenty-four seconds at the median, nearly all of it the build, merges queued behind each other, and a release could fail on a dependency download timing out, a failure that has nothing to do with the code. Even the record did not last: GitHub keeps workflow logs for ninety days by default.
The service underneath had a hard year of its own. On 2 February GitHub reported its hosted runners unavailable, and "Actions jobs queued and timed out while waiting to acquire a hosted runner" (incident). On 5 March ninety-five percent of workflow runs "failed to start within 5 minutes with an average delay of 30 minutes" (incident). On 9 July, for about ten hours, Actions "experienced delayed and failed job starts on GitHub-hosted runners" (incident). On 26 August "Actions jobs failed to start", and for two hours after that runs "were delayed starting by more than 5 minutes as the system caught up with delayed load" (incident). On 13 September GitHub reported "degraded availability across approximately 28 services, including Issues, Pull Requests, Actions" (incident). Counted from its status page, more than fifty incidents touched Actions between January and September. The repositories, the reviews and the workflow definitions stay on GitHub; what moved was the machine the work runs on, and everything that machine had been downloading and throwing away.
Deploys in three minutes instead of eight, measured
Measured on 17 September against the hosted medians: the CMS deploy went from 8 min 24 s to 3 min 05 s, the website from 6 min 12 s to 2 min 27 s and the MCP server from 1 min 45 s to 45 s. Two releases merged together landed in 3 min 22 s of wall clock, and CI running all three at once came out a third faster.


Every run on the machine that built the last one
A hosted runner is a fresh machine every time. The home pool is the same machines every job, so their image layers and package caches are already there, the image never crosses the internet on its way to the cluster, and the build runs on a machine meant for building instead of the one serving the sites. The same build measured 528 seconds from a cold cache and 192 from a warm one. The first home runs were three times slower until one step was removed: each build was uploading its dependency cache back to GitHub over a home connection, 411 seconds of it on the website, which makes sense when the runner is somebody else's machine and none when the cache already lives where the next build will run.
The Runs view lists every run of every watched repository and archives each one at completion, so it outlives GitHub's retention. On the capture day it showed deploys of 49 seconds, 3 min 15 s and 2 min 51 s and CI runs of two to three minutes, and it kept the failures in view: two failed attempts of one deploy, then the third that passed.


A deploy, step by step, on its own time axis
Opening a run shows every job and step on the run's own time axis, measured from GitHub's timestamps, beside the image the run pushed with its digest, size and platforms, and where that image runs now. Retries are kept rather than overwritten: the deploy in the captures failed on its first two attempts and passed on the third, each attempt opens on its own, and the second one shows the deploy step failing after forty-six seconds, so the reason for a failure is still there after the fix.
The passing attempt shows where the time goes on the home pool: seconds for the checkout and the version bump, and a little under two and a half minutes for the build and the deploy to the cluster. It also shows what a console can say that a CI log cannot: that image is no longer in the registry, because retention keeps only the recent builds, and nothing in the cluster runs that digest any more.



When the build machine goes away, the work still happens
The runner pool has a primary and a standby. The build machine's runners take every job while it is up, and the always-on server keeps standby runners that start only when a watchdog decides the build machine is gone. The watchdog needs more than one failed check before it calls the machine down and more than one passing check before it calls it back, so a network blip never flaps anything, and the handover happens within a couple of minutes in either direction. It never cuts a job in half: a shutdown waits for the running job, and the standby stands down only once it is idle.
If the watchdog cannot reach GitHub at all, it starts the standby anyway, because a monitoring failure that stops deploys is worse than the outage it was watching for. The runner pool panel on the console's overview says which side takes the jobs and which standby runners are ready.



Images and packages, served from under one roof
A container registry and an npm registry run side by side on the always-on server, because container images and npm packages speak protocols that share nothing. The container registry takes every image the pool builds over the local network, with one account that can push, read-only accounts for everything that pulls and no anonymous access, and one tag can carry both processor architectures so each node pulls the build that matches it. Retention keeps the two most recent builds of each service plus the one last pulled, which is the rollback window, stated rather than assumed.
The npm registry serves the studio's own library family. Each library's release is published into it byte for byte from its GitHub Release by a sync inside the cluster every five minutes, every read needs a login, the public name refuses writes, and a nightly job compares every published version with its release. The registries have no interface of their own, so the console is their screen: the Estate view lists the container registry's repositories with their sizes and platforms, and an install in any consumer resolves from the building.



From a merge to a running pod, and one line back out
A merge queues the build on the pool, the image goes to the registry over the local network, and the deploy runs as an identity scoped to one namespace, with no right to delete and one named secret. Each node pulls the build that matches it, and one variable per repository sends a service back to hosted runners on its next push.



What it did to the working week, and who should do it
Merges no longer wait for each other: with the build machine up, releases build and deploy side by side, and on the standby a lock takes the heavy steps one at a time so nothing collides. A fix goes out in about three minutes rather than eight, the always-on server compiles only when the build machine is away, and the record of every run stays in the building after GitHub lets it go. The repositories, the reviews and the workflow definitions all stayed on GitHub; only the machines, the caches, the images and the packages moved.
The limit is stated up front: a power cut in the middle of a build fails that build, and it is run again. This is the right move for a team whose releases have started to queue, for anyone whose code and images should not leave the building, and for a founder who has watched a fix wait in someone else's queue. It is the wrong move for a team with nobody to own the machines, and that is worth saying plainly. The cluster underneath is the story of one node becoming a cluster, the platform around it is the homelab that runs the studio, and the same approach applied to a client's platform is the zero-downtime migration to self-hosted Kubernetes.
You May Also Like
27 Pearls - Mobile Release and Store Submission
A finished 27 Pearls app became approved listings on the App Store and Google Play inside a one-week engagement, without a round of rejections. The week covered release builds and signing, listing copy written around the student's benefit, screenshots for every required device size, privacy declarations that matched what the app collects, and a reviewer account that worked. The app remains on Google Play today.
Fully Booked - Zero-Downtime Migration to Self-Hosted Kubernetes
Fully Booked ran on managed cloud infrastructure with a bill that grew every month and a stack nobody fully owned. We moved the whole platform onto a self-hosted Kubernetes cluster without a minute of downtime for the people using it. The cutover covered DNS, ingress, databases and the release pipeline, staged so that traffic shifted only once each layer had been proven on the new cluster. Nobody using the app that week knew anything had happened, which was the point. The recurring bill became owned hardware, and the same cluster now carries the platform's weekly releases.
One node became a cluster, and no application had to change
A second machine that can be switched off at any moment joined a running single-node cluster without a change to any application. An always-on server keeps the control plane, every database and volume and the first copy of every service; the build machine takes the builds and second copies of stateless services behind a taint, and those copies wait rather than fall back when it sleeps. It powers itself off when the work runs out and comes back when a job is queued, a console shows every node, workload and availability span, and every CPU request on the always-on server was measured, taking its reservations from eighty-four to fifty-six percent of what it can allocate.

The blog is field notes from production: architecture decisions, AI systems that survived contact with real users, and infrastructure that pays for itself. Written from the work, not about it.


