Install — RKE2
The production install path on RKE2 — Helm, manifests, kube-vip IP allocation, ingress, secrets.
This is the path that actually runs in production. The fleet runs on
RKE2 today, so this page assumes a healthy RKE2 cluster with
persistent storage, ingress, kube-vip service IPs, and a working kubectl/helm
context.
The route remains /install/k8s because RKE2 is the Kubernetes distribution we run.
When this page says "the cluster", read that as "the RKE2 cluster".
192.168.3.200 is the control-plane VIP and must not be allocated to a service.
The kube-vip ConfigMap excludes it — never edit the range without preserving that
exclusion. A service that lands on .200 will collide with the control plane and
break the cluster.
Prerequisites
| Thing | Version we run | Notes |
|---|---|---|
| Kubernetes distribution | RKE2 stable channel | Use the official RKE2 quick start and requirements for host bootstrap |
| RKE2 access | /etc/rancher/rke2/rke2.yaml | Export KUBECONFIG or pass --kubeconfig; bundled tools live under /var/lib/rancher/rke2/bin |
| RKE2 HA endpoint | 192.168.3.200 | Fixed control-plane registration/API VIP; must be present in RKE2 tls-san |
| Persistent storage | Longhorn | Any RWO-capable provisioner works for r2d2-fleet itself |
| Ingress | RKE2 ingress-nginx today | Replaceable; RKE2 docs flag ingress-nginx EOL, so treat Traefik migration as future work |
| LB IPs | kube-vip (DaemonSet, ARP mode) | MetalLB also fine — see the IP-allocation note |
| Container registry | any | The dashboard image is pulled by the K8s nodes; private registries need an imagePullSecret |
kubectl + helm | recent | Helm needs no RKE2-specific flags once kubeconfig is correct |
RKE2 operator baseline
RKE2 installs and runs as systemd services. On a server node, the important paths are:
| Path | Why it matters |
|---|---|
/etc/rancher/rke2/config.yaml | Primary RKE2 config file; create it manually before first start |
/etc/rancher/rke2/config.yaml.d/*.yaml | Optional ordered drop-in config files |
/etc/rancher/rke2/rke2.yaml | Kubeconfig for kubectl and helm |
/var/lib/rancher/rke2/bin/ | Bundled kubectl, crictl, and ctr binaries |
/var/lib/rancher/rke2/server/node-token | Token for joining additional server or agent nodes |
/var/lib/rancher/rke2/server/manifests/ | RKE2 auto-applied AddOn manifests and HelmChart resources |
Use this shell profile when operating directly from an RKE2 server:
export PATH="$PATH:/var/lib/rancher/rke2/bin"
export KUBECONFIG=/etc/rancher/rke2/rke2.yaml
kubectl get nodes -o wide
kubectl -n kube-system get pods
helm ls --all-namespacesWhen accessing the cluster from an operator workstation, copy
/etc/rancher/rke2/rke2.yaml into your local kubeconfig and replace 127.0.0.1 with
the control-plane VIP or DNS name. Do not hand-edit a workstation kubeconfig back onto
the server.
RKE2 server and agent health is checked through systemd:
systemctl status rke2-server
journalctl -u rke2-server -f
systemctl status rke2-agent
journalctl -u rke2-agent -fFor HA, RKE2 needs an odd number of server nodes, with three recommended. Additional
server nodes and all agents register through the fixed endpoint on port 9345; the
Kubernetes API is served on port 6443.
# /etc/rancher/rke2/config.yaml on the first server
token: <shared-registration-token>
tls-san:
- 192.168.3.200
- rke2.office.ilab.zone# /etc/rancher/rke2/config.yaml on additional servers/agents
server: https://192.168.3.200:9345
token: <shared-registration-token>RKE2 host checks that matter before installing R2-D2:
| Check | Command |
|---|---|
| Unique node names | kubectl get nodes -o custom-columns=NAME:.metadata.name |
| RKE2 packaged AddOns | kubectl get addon -A |
| Ingress controller | kubectl -n kube-system get pods -l app.kubernetes.io/name=ingress-nginx |
| kube-vip | `kubectl -n kube-system get ds,cm |
| Longhorn storage class | kubectl get storageclass |
| Supervisor/API reachability | nc -vz 192.168.3.200 9345 && nc -vz 192.168.3.200 6443 |
Minimum RKE2 network ports to keep open inside the cluster:
| Port | Protocol | Scope | Purpose |
|---|---|---|---|
6443 | TCP | RKE2 nodes → server nodes | Kubernetes API |
9345 | TCP | RKE2 nodes → server nodes | RKE2 supervisor / node registration |
10250 | TCP | RKE2 nodes → all nodes | kubelet metrics |
2379-2381 | TCP | server nodes → server nodes | embedded etcd |
8472 | UDP | nodes → nodes | Canal/Flannel VXLAN when using default Canal |
30000-32767 | TCP | where needed | NodePort range |
Never expose VXLAN ports to the public internet. If NetworkManager is installed, make sure it ignores CNI-managed interfaces before blaming pods for networking failures.
The official RKE2 docs currently warn that ingress-nginx is going end-of-life and that RKE2 v1.36 will default new clusters to Traefik. R2-D2 still documents the current nginx-ingress + Caddy path below; schedule and test any Traefik migration deliberately instead of letting an upgrade surprise operators.
Architecture recap for the installer
RKE2 API/VIP kube-vip (ARP LB) nginx-ingress controller Caddy edge
│ │ │ │
▼ ▼ ▼ ▼
192.168.3.200 service IPs cluster.local routes *.office.ilab.zone
6443 / 9345 (192.168.3.x range) (Ingress objects) (TLS terminates here)Two paths reach the dashboard from outside the cluster:
- Internal users hit
https://r2d2.office.ilab.zone, which Caddy proxies to the ingress controller listening on192.168.3.211–214:8080. The ingress object routes to ther2d2-fleetservice. - In-cluster services (e.g. Node-RED instances calling back into the dashboard)
hit the cluster-internal
r2d2-fleet.<namespace>.svc.cluster.local.
There is no path "Caddy → kube-vip LB IP → service" for the dashboard. That convention
applies broadly: *.office.ilab.zone Caddy entries proxy to the nginx-ingress nodes,
not to kube-vip-assigned LB IPs.
1. Reserve an IP range for kube-vip
192.168.3.200 is the control-plane VIP and must not be allocated to a service.
The kube-vip ConfigMap carries the allocatable pool — exclude .200:
# kube-system/kubevip ConfigMap
data:
range-global: "192.168.3.120-199,201-220,10"The .10 at the end is a separate reserved slot (legacy). The point is: anything in
.200 will collide with the control plane.
RKE2 also needs the control-plane VIP in tls-san, or remote kubectl and joining
nodes will eventually hit certificate-name failures.
Recommended kube-vip env tuning for stable lease behavior (set on the kube-vip DaemonSet):
vip_leaseduration = 15
vip_renewdeadline = 10
vip_retryperiod = 2These keep failover crisp without thrashing the leader election. Rollback is one
kubectl rollout undo daemonset/kube-vip -n kube-system away.
2. Stand up the dependencies
The dashboard talks to four external systems. Each one is its own helm release; they're listed here in dependency order.
2a. TrailBase
helm upgrade --install trailbase ./trailbase \
--namespace trailbase --create-namespace \
-f trailbase/values.yamlInitial bootstrap: open the TrailBase admin (it'll come up on the LB IP from
values.yaml) and create the first operator account. That account becomes the seed for
the dashboard's AUTH_ENABLED=true path.
2b. YugabyteDB
Used for relational data outside TrailBase (Matrix, future telemetry). The dashboard itself does not hold a direct YB connection.
helm upgrade --install yugabytedb-matrix ./yugabytedb-matrix \
--namespace yugabytedb-matrix --create-namespace \
-f yugabytedb-matrix/values.override.yamlBe patient with first-time bootstrap — the master quorum takes a minute to elect.
If you've previously had a tserver disk wiped and rebuilt, you may need to clean up
stale Raft UUIDs with yb-ts-cli unsafe_config_change (the runbook lives in the team's
bd knowledge base — search "yb tserver rebuild").
2c. Restreamer (datarhei/core)
helm upgrade --install restreamer ./restreamer \
--namespace restreamer --create-namespace \
-f restreamer/values.yamlSet RESTREAMER_API_URL on the dashboard to the in-cluster service DNS, e.g.
http://restreamer.restreamer.svc.cluster.local:8080.
2d. TBMQ (MQTT broker)
helm upgrade --install tbmq ./tbmq \
--namespace tbmq --create-namespace \
-f tbmq/values.yamlThe redis-cluster behind TBMQ is the failure-mode that bites operators most often. If
the host node carrying a redis pod restarts hard, the pod's nodes.conf can be left
in a state where the cluster can't re-form. Recovery is FORGET + MEET + REPLICATE
against the broken peer — runbook in bd.
3. Deploy the dashboard itself
kubectl create namespace nodered
# imagePullSecret if you're pulling from a private registry
kubectl -n nodered create secret docker-registry r2d2-pull \
--docker-server=<registry> \
--docker-username=<user> --docker-password=<token>
# secrets the dashboard needs
kubectl -n nodered create secret generic r2d2-fleet \
--from-literal=AUTH_ENABLED=true \
--from-literal=NODE_RED_ADMIN_TOKEN='...' \
--from-literal=RESTREAMER_API_URL='http://restreamer.restreamer.svc.cluster.local:8080' \
--from-literal=TRAILBASE_URL='http://trailbase.trailbase.svc.cluster.local:4000' \
--from-literal=HOLONET_URL='https://holonet.office.ilab.zone' \
--from-literal=HOLOCHRON_URL='https://holochron.office.ilab.zone' \
--from-literal=COP_URL='https://cop.office.ilab.zone' \
--from-literal=PROBE_URL='https://probe.office.ilab.zone' \
--from-literal=RIFT_GATES_URL='https://rift-gates.office.ilab.zone'
# the deployment, service, ingress
kubectl -n nodered apply -f r2d2-fleet/k8s/deployment.yaml
kubectl -n nodered apply -f r2d2-fleet/k8s/service.yaml
kubectl -n nodered apply -f r2d2-fleet/k8s/ingress.yamlThe service.yaml is a ClusterIP (the dashboard exits via ingress, not a per-service
LB IP). The ingress.yaml carries nginx.ingress.kubernetes.io/proxy-body-size: "10m"
(for YAML imports on the Restreamer page) and is otherwise plain.
4. Wire Caddy at the edge
Public TLS terminates at Caddy. Caddy forwards https://r2d2.office.ilab.zone to the
nginx-ingress backends on port 8080 — not to a kube-vip LB IP. This is the same
pattern every *.office.ilab.zone host follows.
Example Caddyfile fragment:
r2d2.office.ilab.zone {
encode zstd gzip
reverse_proxy {
to 192.168.3.211:8080 192.168.3.212:8080 192.168.3.213:8080 192.168.3.214:8080
health_uri /healthz
lb_policy round_robin
}
tls {
# ...your TLS config
}
}Port 80 is Traefik on the same Caddy box, separately. Don't confuse the two.
5. Bootstrap content
The first time the dashboard starts, two things need to be present in the running pod:
instances/registry.yaml— the Node-RED registry. Today the deployment mounts this from a git clone bake-step. There's a known drift issue when the ConfigMap version of this file diverges from the in-pod copy; the Settings → Fleet page surfaces the drift when it sees it.republic-config.yaml,videos-config.yaml,holocron-config.yaml,archives-config.yaml— each plane's config. These can start empty; the Settings pages let you fill them in and commit back to git.
6. Verify
| Check | What good looks like |
|---|---|
kubectl -n nodered get pods | r2d2-fleet-... is Ready |
curl -sv https://r2d2.office.ilab.zone/healthz | 200 |
| Dashboard sidebar | All four sidebar status tiles (Node Red, Restreaming, The Force) show coloured dots; "Updated" timestamp is recent |
GET /api/openapi.json | returns the full spec — Scalar should render at /droidspeak/api |
GET /api/instances | returns the registry with instances: [] (empty fleet) or your registered instances |
Cleanup / rollback
The dashboard is stateless — kubectl delete it, redeploy the previous tag.
TrailBase, YugabyteDB, Restreamer, and TBMQ each have their own backup/restore stories
(check their helm chart values.yaml for the volume and snapshot config).
The dashboard image and the OpenAPI spec ship together — they cannot get out of sync mid-deploy. The fastest rollback for a regression in the API surface is a tag bump on the deployment image.