← cd ~/work

case study · personal project · live

substation/

A live console for a 4-node k3s cluster and a smart home, built and run alone. A Rust api gathers vitals from gRPC agents on every node and pushes one live snapshot to every screen over a WebSocket; a Svelte 5 web UI, four Android apps and a voice path all drive it; and its own ship pipeline proves every image, walks every page and can roll back in one command.

role sole engineer design, code, infra, ops
stack Rust · Svelte 5 Kotlin · gRPC · Flux · Caddy
scale 4-node k3s agents on every node
tests ~4,300 ~2,700 Rust · ~1,100 web · ~440 Kotlin

1 · the problem

Many moving parts, no single place to see them or act on them.

A small cluster still has everything a big one has: node health, etcd quorum, disks, certificates, backups, workloads, GitOps. Add connected devices across three radio protocols and a family of apps, and the answer to "is everything fine?" was spread across a dozen tools.

Deploys were the other half. Every check in today's pipeline exists because a hand deploy once got it wrong: a stale cached layer shipped an old binary under a new label; two sessions imported over each other; a restart fought the GitOps controller.

2 · what I built

One api, one live snapshot, every screen.

Select any box to see what it does and why it is there.

architecture
selected Rust api [RUST] [AXUM] [TONIC] [KUBE]

The centre of everything. Polls the Kubernetes API, receives node vitals from the agents over gRPC, keeps metric history in SQLite, evaluates alert rules with quiet hours, and exposes a token-gated control surface where every write is audited. ~2,600 tests.

$ make api # :8080 http/ws · :50051 grpc

3 · the ship pipeline

One command from a commit to live, and one command back.

$ scripts/ship api web
05 · walk

Walk all 176 pages, then run the design check.

An automated browser walks every page, in a tab opened before the deploy and in a fresh one. The design check fails the ship on a page that scrolls sideways at 390px or on unreadable text that was not there before; everything else lands on a ship review.

why it exists

Green tests do not prove a page still renders for a real person.

4 · decisions & trade-offs

Every choice had a cost. Here is what I paid.

4.1 · Rust for the api

chose
Rust: axum, tonic, kube
because
A long-running process holds live state for every screen and takes a stream from every node each second. The compiler catches a whole class of mistakes before the walk does, and one static binary makes the hash proof simple.
cost
Slower builds, and fewer people who can pick the code up.

4.2 · Svelte 5

chose
Svelte 5 runes, Vite, TypeScript
because
A UI that changes every second needs fine-grained reactivity, not whole-tree re-renders. Small bundles help on a phone.
cost
Rune effects do not run under the unit-test runner, so logic lives in plain tested modules and components stay thin.

4.3 · Self-hosted GitOps

chose
Flux on my own k3s
because
Git is the source of truth: a rebuilt node is a re-apply, not an afternoon of remembering what was done by hand.
cost
I own every upgrade, and the pipeline had to learn to work with the controller instead of against it.

4.4 · Terminal-native UI

chose
One mono family, panes, ⌘K, block meters
because
Dense information read at a glance and driven from the keyboard, with the CLI equivalent shown so nothing is magic.
cost
Only size, weight and ink are available for hierarchy, so the walk’s design check guards readability on every ship.

4.5 · LLM features, fast paths first

chose
Deterministic first, model second
because
Voice: light commands are matched without a model and handled in milliseconds; everything else goes to a model. The same rule runs through the Telegram agent: predictable and cheap for the common case, flexible for the long tail.
cost
Two paths to keep consistent, so both are covered by tests.

5 · results

$ substation --facts
tests in total
~4,300
rust · web · kotlin tests
~2,700 · ~1,100 · ~440
commits
1,128
lines of Rust
~210k
lines of Svelte + TypeScript
~98k
lines of Kotlin
~59k
nodes under watch
4 · k3s
pages walked per console ship
176
rollback
1 command
android apps from one pipeline
4
console writes audited
every one
metrics exported
Prometheus /metrics

6 · what this means for your project

The same habits, pointed at your problem.

Want this kind of care on your codebase?

$ curl sala.dev/contact -d "re: substation"
NORMAL /work/substation available UK · remote tony@sala.dev updated
available Start a project →