# A self-hosted homelab that runs the studio, maintained every day

Company: Webanion
Service: DevOps Engineering
Period: 2026-02 - present
Tech: Kubernetes, Traefik, Cloudflare, Go, React.js, TypeScript, SQLite, VictoriaMetrics, Grafana, PostgreSQL, MongoDB, zot, Verdaccio, systemd, Linux, GitHub Actions
Canonical: https://webanion.com/portfolio/self-hosted-homelab-that-runs-the-studio-infrastructure

The studio's own infrastructure on a self-hosted Kubernetes cluster, planned in February 2026 and maintained every day since: its website, CMS and MCP servers and the client platforms it looks after, served through Cloudflare tunnels, with every database dumped and checked each night, metrics and alert rules that reach a phone, an outage monitor outside the house, a container registry and an npm registry, and one private console that joins merges, runs, images, pods and nodes on a single page.

## The Full Story

What it takes to own the whole stack instead of renting a dashboard per vendor, and what a studio gets back for looking after it every day.

### Everything the studio ships runs on servers it owns

<p>A studio that builds products for clients also runs a surprising amount of its own. Webanion serves its website in all of its languages, the CMS behind it, public and private MCP servers that let AI assistants read and manage that content, and the client platforms it builds and looks after, among them a finance platform, a booking platform, a ledger app and a research group's site. Each of those needs somewhere to run, a database that is backed up, a certificate that is renewed and someone who notices when it breaks. Renting each piece from its own vendor means a bill, a dashboard and a blind spot per vendor, and no single place that answers the three questions that decide a working day: what is running, is it healthy, and what needs attention today.</p><p>The answer here is a Kubernetes cluster on servers the studio owns, planned in February 2026 and maintained every day since. Every site and platform is served from it through Cloudflare tunnels, every database is dumped every night, metrics and alert rules watch the cluster from the inside and an outage monitor watches it from outside the house, a container registry and an npm registry hold the studio's own images and packages, and one private console sits over all of it. On the capture day, 30 September 2026, the console counted two nodes, fifty-four workloads and ninety-three pods.</p>

### One page that says what is running and what needs attention

<p>The console is the first page opened in the morning and the one left open during a merge. Its overview reads the cluster, GitHub, the registries and the hosts together: nodes ready, workloads and pods healthy, pull requests pending and merged, the share of recent runs that passed, what is left of the GitHub request budget, each host's load, memory, disk, temperature and build cache, the latest actions taken, and a short list of what needs a hand today, most serious first.</p><p>What it adds is the join: this merge, the run it started, the image that run pushed, the pod running that image and the node the pod is on, on one page instead of five tabs. It sits behind Cloudflare Access and works at phone width, because the moment something needs checking is rarely the moment a laptop is open. On the capture day it read two of two nodes ready, fifty-four of fifty-four workloads, ninety-three of ninety-three pods healthy and thirteen merges in the previous two days.</p>

### Every merge lands from one board

<p>Pull requests from the studio's GitHub organisations sit on one board, pending, running and done, and each card says what merging it will do: deploy to live, deploy to staging, or nothing. Below the board a table goes repository by repository through whether it may be merged from here, why not when it may not, its concurrency group and whether its workflow can be dispatched.</p><p>A merge from the board is guarded rather than trusted. The server reads the pull request again at the moment of the press, never merges past a red check or a draft, always merges with a merge commit, and wants the repository's full name typed before anything that deploys to live. For a studio shipping several products in a week, that is the difference between knowing what a merge will do and finding out afterwards. On the capture day the board held two pending pull requests and thirteen done, each done card carrying the image it deployed.</p>

### Backups that check in every night

<p>Every database on the cluster is dumped every night, a quarter of an hour apart so no two start together, and each dump writes the row count of every table beside it, which is what a restore proof checks against when it restores a run into a throwaway namespace. The console's own archive, the one store nothing else could rebuild, is copied with SQLite's online backup and checked for integrity and row counts before it is kept. After a verified run every job pings a dead man's switch outside the house, so a job that fails, hangs or stops being scheduled is noticed the next morning rather than on the day a restore is needed.</p><p>The Backups view charts each claim's growth from hourly measurements. On the capture day it showed 664.2 MB of backup data across ten claims, ten backup jobs, the newest nine hours old and none failing, and the Schedules tab listed all fourteen scheduled jobs with their last run succeeded.</p>

### The hosts, the metrics, and an alarm outside the house

<p>Each host reports load, memory, disk, temperature and build cache every five minutes, and the metrics store is read through Grafana inside the page. The cluster's alert rules reach a phone, an hourly check flags merged manifests nobody applied, and an outage monitor on Cloudflare's edge checks every public site each minute, because a monitor on the cluster cannot report the cluster down.</p>

### Every facility, on one map

<p>What the homelab provides, drawn by role rather than by machine: the tunnels and the router at the edge, the cluster, the databases and their volumes, the container and npm registries, the runner pool, the nightly backups and their switches, the metrics store and its alerts, the outage monitor outside the house, and the console reading all of it.</p>

### Every action written down, refusals included

<p>The console does more than read. Merges, re-runs, dispatches, restarts, scaling a second copy up or down and roll backs are pressed from the page, and each one is written through a second identity whose permissions name every object it may touch, with a token requested for that one action and dropped after it. Before anything happens the console reads the live state again and lists every precondition as passed, a warning or a refusal, and anything that deploys to live needs its name typed to confirm.</p><p>Every press becomes a row that is only ever appended: when, the action, the target, the outcome and what was said, refusals included, so the record of an action that was refused is as complete as the record of one that went through. The same archive keeps what GitHub deletes: every run is archived at completion, a year of log bodies and the rows for good, past GitHub's default ninety-day retention.</p>

### Maintained every day, and what that is worth

<p>The daily routine is short because the tooling carries most of it: the console's list of what needs a hand in the morning, a drift check every hour that catches a merged manifest nobody applied, the phone for anything that fails at night, backup switches that stay quiet while every job reports in, and certificate checks that warn weeks before one expires. On the capture day that list held one item, a build cache over its cap that the nightly prune brings back, the kind of small drift that is cheap to fix on the day and expensive to find a month later.</p><p>The honest cost is that someone has to own it. Owning the stack means certificates, disks and upgrades are the studio's problem rather than a vendor's, and a team with nobody to look after its infrastructure is better off renting. For a founder or a small team who want their infrastructure to be something they understand rather than a stack of invoices, and for a client who wants evidence that the platform they pay for is looked after every day, this is what that looks like in practice. The <a href="/portfolio/webanion-headless-cms-portfolio-platform-with-mcp-server">Webanion platform</a> runs on it, and the things it made possible have stories of their own, among them <a href="/portfolio/scaling-a-homelab-from-one-node-to-a-multi-node-cluster">growing one node into a cluster without changing an application</a> and <a href="/portfolio/replacing-hosted-ci-with-a-home-runner-pool-and-private-registries">moving builds from hosted CI to a home runner pool and private registries</a>.</p>

### The alarm reaches the phone, and the alarm is watched too

<p>Nobody has to be watching a dashboard for trouble to be noticed: an alarm reaches the phone on Telegram, Discord and ntfy, from the outage monitor and the cluster's own alert rules on Telegram and from the monitor's watcher on Discord and ntfy. The monitor is watched in turn from outside the house, so when it fell silent for four minutes on 30 September the silence itself arrived as an alarm, down and then back up, on Discord and ntfy at once.</p>
