# When hosted CI became the slowest step, the builds came home

Company: Webanion
Service: DevOps Engineering
Period: 2026-08 - 2026-09
Tech: GitHub Actions, CI/CD, Docker, zot, Verdaccio, Kubernetes, Go, SQLite, Linux, systemd
Canonical: https://webanion.com/portfolio/replacing-hosted-ci-with-a-home-runner-pool-and-private-registries

Builds and deploys moved from GitHub's hosted runners, where every run started on a fresh machine and releases waited on queues and incidents, to a runner pool at home: a build machine that takes every job while it is up, and standby runners on an always-on server behind a watchdog that fails towards shipping. Images go to a private container registry and the studio's libraries to a private npm registry, both in the building, and every run is archived past GitHub's retention. The slowest deploy went from eight minutes twenty-four seconds to three minutes five, and CI came out a third faster.

## The Full Story

What a hosted pipeline costs a business in waiting, and what changed when builds, images and packages moved under one roof.

### The pipeline was somebody else's machine, and somebody else's bad day

<p>Until the middle of September 2026 every build and deploy ran on GitHub's hosted runners, and a hosted runner is a fresh virtual machine every time: the base image, every dependency and every build layer downloaded again for a one-line change, then thrown away. The image was then built on the always-on production server, because that was the only place the result could run, so the machine serving the databases and the public sites was compiling while it did. The slowest deploy took eight minutes twenty-four seconds at the median, nearly all of it the build, merges queued behind each other, and a release could fail on a dependency download timing out, a failure that has nothing to do with the code. Even the record did not last: GitHub keeps workflow logs for ninety days by default.</p><p>The service underneath had a hard year of its own. On 2 February GitHub reported its hosted runners unavailable, and "Actions jobs queued and timed out while waiting to acquire a hosted runner" (<a href="https://www.githubstatus.com/incidents/xwn6hjps36ty">incident</a>). On 5 March ninety-five percent of workflow runs "failed to start within 5 minutes with an average delay of 30 minutes" (<a href="https://www.githubstatus.com/incidents/g5gnt5l5hf56">incident</a>). On 9 July, for about ten hours, Actions "experienced delayed and failed job starts on GitHub-hosted runners" (<a href="https://www.githubstatus.com/incidents/cstx3v63mklm">incident</a>). On 26 August "Actions jobs failed to start", and for two hours after that runs "were delayed starting by more than 5 minutes as the system caught up with delayed load" (<a href="https://www.githubstatus.com/incidents/y1t7p9fzrlj2">incident</a>). On 13 September GitHub reported "degraded availability across approximately 28 services, including Issues, Pull Requests, Actions" (<a href="https://www.githubstatus.com/incidents/0rn90wk115q9">incident</a>). Counted from <a href="https://www.githubstatus.com/history">its status page</a>, more than fifty incidents touched Actions between January and September. The repositories, the reviews and the workflow definitions stay on GitHub; what moved was the machine the work runs on, and everything that machine had been downloading and throwing away.</p>

### Deploys in three minutes instead of eight, measured

<p>Measured on 17 September against the hosted medians: the CMS deploy went from 8 min 24 s to 3 min 05 s, the website from 6 min 12 s to 2 min 27 s and the MCP server from 1 min 45 s to 45 s. Two releases merged together landed in 3 min 22 s of wall clock, and CI running all three at once came out a third faster.</p>

### Every run on the machine that built the last one

<p>A hosted runner is a fresh machine every time. The home pool is the same machines every job, so their image layers and package caches are already there, the image never crosses the internet on its way to the cluster, and the build runs on a machine meant for building instead of the one serving the sites. The same build measured 528 seconds from a cold cache and 192 from a warm one. The first home runs were three times slower until one step was removed: each build was uploading its dependency cache back to GitHub over a home connection, 411 seconds of it on the website, which makes sense when the runner is somebody else's machine and none when the cache already lives where the next build will run.</p><p>The Runs view lists every run of every watched repository and archives each one at completion, so it outlives GitHub's retention. On the capture day it showed deploys of 49 seconds, 3 min 15 s and 2 min 51 s and CI runs of two to three minutes, and it kept the failures in view: two failed attempts of one deploy, then the third that passed.</p>

### A deploy, step by step, on its own time axis

<p>Opening a run shows every job and step on the run's own time axis, measured from GitHub's timestamps, beside the image the run pushed with its digest, size and platforms, and where that image runs now. Retries are kept rather than overwritten: the deploy in the captures failed on its first two attempts and passed on the third, each attempt opens on its own, and the second one shows the deploy step failing after forty-six seconds, so the reason for a failure is still there after the fix.</p><p>The passing attempt shows where the time goes on the home pool: seconds for the checkout and the version bump, and a little under two and a half minutes for the build and the deploy to the cluster. It also shows what a console can say that a CI log cannot: that image is no longer in the registry, because retention keeps only the recent builds, and nothing in the cluster runs that digest any more.</p>

### When the build machine goes away, the work still happens

<p>The runner pool has a primary and a standby. The build machine's runners take every job while it is up, and the always-on server keeps standby runners that start only when a watchdog decides the build machine is gone. The watchdog needs more than one failed check before it calls the machine down and more than one passing check before it calls it back, so a network blip never flaps anything, and the handover happens within a couple of minutes in either direction. It never cuts a job in half: a shutdown waits for the running job, and the standby stands down only once it is idle.</p><p>If the watchdog cannot reach GitHub at all, it starts the standby anyway, because a monitoring failure that stops deploys is worse than the outage it was watching for. The runner pool panel on the console's overview says which side takes the jobs and which standby runners are ready.</p>

### Images and packages, served from under one roof

<p>A container registry and an npm registry run side by side on the always-on server, because container images and npm packages speak protocols that share nothing. The container registry takes every image the pool builds over the local network, with one account that can push, read-only accounts for everything that pulls and no anonymous access, and one tag can carry both processor architectures so each node pulls the build that matches it. Retention keeps the two most recent builds of each service plus the one last pulled, which is the rollback window, stated rather than assumed.</p><p>The npm registry serves the studio's own library family. Each library's release is published into it byte for byte from its GitHub Release by a sync inside the cluster every five minutes, every read needs a login, the public name refuses writes, and a nightly job compares every published version with its release. The registries have no interface of their own, so the console is their screen: the Estate view lists the container registry's repositories with their sizes and platforms, and an install in any consumer resolves from the building.</p>

### From a merge to a running pod, and one line back out

<p>A merge queues the build on the pool, the image goes to the registry over the local network, and the deploy runs as an identity scoped to one namespace, with no right to delete and one named secret. Each node pulls the build that matches it, and one variable per repository sends a service back to hosted runners on its next push.</p>

### What it did to the working week, and who should do it

<p>Merges no longer wait for each other: with the build machine up, releases build and deploy side by side, and on the standby a lock takes the heavy steps one at a time so nothing collides. A fix goes out in about three minutes rather than eight, the always-on server compiles only when the build machine is away, and the record of every run stays in the building after GitHub lets it go. The repositories, the reviews and the workflow definitions all stayed on GitHub; only the machines, the caches, the images and the packages moved.</p><p>The limit is stated up front: a power cut in the middle of a build fails that build, and it is run again. This is the right move for a team whose releases have started to queue, for anyone whose code and images should not leave the building, and for a founder who has watched a fix wait in someone else's queue. It is the wrong move for a team with nobody to own the machines, and that is worth saying plainly. The cluster underneath is the story of <a href="/portfolio/scaling-a-homelab-from-one-node-to-a-multi-node-cluster">one node becoming a cluster</a>, the platform around it is <a href="/portfolio/self-hosted-homelab-that-runs-the-studio-infrastructure">the homelab that runs the studio</a>, and the same approach applied to a client's platform is the <a href="/portfolio/fully-booked-zero-downtime-migration-to-self-hosted-kubernetes">zero-downtime migration to self-hosted Kubernetes</a>.</p>
