A live console for a 4-node k3s cluster and a smart home, built and run alone. A Rust api gathers vitals from gRPC agents on every node and pushes one live snapshot to every screen over a WebSocket; a Svelte 5 web UI, four Android apps and a voice path all drive it; and its own ship pipeline proves every image, walks every page and can roll back in one command.
rolesole engineerdesign, code, infra, ops
stackRust · Svelte 5Kotlin · gRPC · Flux · Caddy
scale4-node k3sagents on every node
tests~4,300~2,700 Rust · ~1,100 web · ~440 Kotlin
1 · the problem
Many moving parts, no single place to see them or act on them.
A small cluster still has everything a big one has: node health, etcd quorum, disks, certificates, backups, workloads, GitOps. Add connected devices across three radio protocols and a family of apps, and the answer to "is everything fine?" was spread across a dozen tools.
Deploys were the other half. Every check in today's pipeline exists because a hand deploy once got it wrong: a stale cached layer shipped an old binary under a new label; two sessions imported over each other; a restart fought the GitOps controller.
2 · what I built
One api, one live snapshot, every screen.
Select any box to see what it does and why it is there.
architecture
k3s · 4 nodeshttps⇄ gRPC stream, once a second/ws live snapshot, through Caddycontrolreconciles from Gitships
selectedRust api[RUST] [AXUM] [TONIC] [KUBE]
The centre of everything. Polls the Kubernetes API, receives node vitals from the agents over gRPC, keeps metric history in SQLite, evaluates alert rules with quiet hours, and exposes a token-gated control surface where every write is audited. ~2,600 tests.
$ make api # :8080 http/ws · :50051 grpc
3 · the ship pipeline
One command from a commit to live, and one command back.
$ scripts/ship api web
05 · walk
Walk all 176 pages, then run the design check.
An automated browser walks every page, in a tab opened before the deploy and in a fresh one. The design check fails the ship on a page that scrolls sideways at 390px or on unreadable text that was not there before; everything else lands on a ship review.
why it exists
Green tests do not prove a page still renders for a real person.
4 · decisions & trade-offs
Every choice had a cost. Here is what I paid.
4.1 · Rust for the api
chose
Rust: axum, tonic, kube
because
A long-running process holds live state for every screen and takes a stream from every node each second. The compiler catches a whole class of mistakes before the walk does, and one static binary makes the hash proof simple.
cost
Slower builds, and fewer people who can pick the code up.
4.2 · Svelte 5
chose
Svelte 5 runes, Vite, TypeScript
because
A UI that changes every second needs fine-grained reactivity, not whole-tree re-renders. Small bundles help on a phone.
cost
Rune effects do not run under the unit-test runner, so logic lives in plain tested modules and components stay thin.
4.3 · Self-hosted GitOps
chose
Flux on my own k3s
because
Git is the source of truth: a rebuilt node is a re-apply, not an afternoon of remembering what was done by hand.
cost
I own every upgrade, and the pipeline had to learn to work with the controller instead of against it.
4.4 · Terminal-native UI
chose
One mono family, panes, ⌘K, block meters
because
Dense information read at a glance and driven from the keyboard, with the CLI equivalent shown so nothing is magic.
cost
Only size, weight and ink are available for hierarchy, so the walk’s design check guards readability on every ship.
4.5 · LLM features, fast paths first
chose
Deterministic first, model second
because
Voice: light commands are matched without a model and handled in milliseconds; everything else goes to a model. The same rule runs through the Telegram agent: predictable and cheap for the common case, flexible for the long tail.
cost
Two paths to keep consistent, so both are covered by tests.