R2-D2
Dashboard
Node Red
Restreaming
The Force
UpdatedNever
v1.0.0

Overview

Loading…
R2-D2
Dashboard
Node Red
Restreaming
The Force
UpdatedNever
v1.0.0
DDroidspeak / Docs
Operator handbook
Droidspeak
What is R2-D2?
Runtime architectureAuth — Keycloak migration plan
FleetRestreamingGalaxy MapThe RepublicTemple ArchivesSTANAG 4817The Force
Tech StackNext.js 15 + React 19Tailwind v4ZustandTanStack Table + DataViewhls.jsLow-latency playerreagraphoglreact-grid-layoutMonacoScalar API Referencefumadocs
Cluster InfraTrailBaseReductStoreRestreamer (datarhei/core)TBMQKeycloakLonghornkube-vipIngress (Caddy + nginx-ingress + Traefik)Netbird
Operator QA Runbook
Install — RKE2Install — Dokploy
Contributor guideRelease Notes
Cluster Infra

Longhorn

Block storage — PVCs, csi-attacher leader recovery, and Multi-Attach error decoding.

Longhorn is the cluster's block-storage layer. R2-D2's stateful pieces — TrailBase, ReductStore, Restreamer, Keycloak, TBMQ's redis-cluster — all live on Longhorn PVCs.

Where it sits

Chartlonghorn/ in the parent repo
NodesStorage on every worker
Replicas3 by default (per-volume)

csi-attacher stuck-leader recovery

The most common Longhorn pain point: a workload shows Multi-Attach error for volume "<pv>" and a VolumeAttachment is stuck with a finalizer.

Decision tree:

flowchart TB
  A[Multi-Attach error + stuck VA finalizer] --> B[Check csi-attacher pods]
  B --> C{Lease holder healthy?}
  C -- yes --> D[Investigate the actual workload]
  C -- no, crashlooping --> E[Force-delete the csi-attacher pod]
  E --> F[Lease re-elects to a healthy pod]
  F --> G[VA finalizer resolves]

The exact pods + lease are namespace-prefixed; identify them via kubectl get leases -n longhorn-system and cross-reference with kubectl get pods -n longhorn-system.

Force-delete is only safe when the holder is genuinely crashlooping. If the holder is healthy and the VA is still stuck, the problem is downstream (node loss, taint mismatch, finalizer race) — investigate before swinging the hammer.

Operator quick-reference

SymptomFirst check
Pod stuck ContainerCreating with Multi-Attachkubectl get volumeattachment for the PV; identify lease holder
Volume "degraded" in Longhorn UIOne replica is down — usually a worker that's cordoned or storage drift
Snapshots failingLonghorn-side snapshot quota, not a workload issue
New PVCs not provisioningkubectl get sc — is longhorn the default?

R2-D2 workload notes

  • Restreamer — small PVC for the process DB; can survive a fresh PVC at the cost of re-creating processes from the dashboard.
  • TrailBase / ReductStore — both rely on PVCs surviving pod restarts; treat them as durable.
  • TBMQ redis-cluster — see the TBMQ page for the prevention block already applied and the nodes.conf recovery recipe.

Scale-downs during Longhorn maintenance

For R2-D2's workloads, scale down anything multi-replica before maintenance and rebuild after. The redis-cluster behind TBMQ is the most sensitive — see TBMQ for the prevention block + recovery procedure.

See also

  • TrailBase, ReductStore, Restreamer, TBMQ — R2-D2's PVC consumers
  • kube-vip — sibling cluster infra

Keycloak

OIDC identity — R2-D2 theme, hardened realm, and the dashboard's auth integration.

kube-vip

LoadBalancer VIP allocation — IP reservations, lease tuning, and the .200 invariant.

On this page

Where it sitscsi-attacher stuck-leader recoveryOperator quick-referenceR2-D2 workload notesScale-downs during Longhorn maintenanceSee also