Longhorn
Block storage — PVCs, csi-attacher leader recovery, and Multi-Attach error decoding.
Longhorn is the cluster's block-storage layer. R2-D2's stateful pieces — TrailBase, ReductStore, Restreamer, Keycloak, TBMQ's redis-cluster — all live on Longhorn PVCs.
Where it sits
| Chart | longhorn/ in the parent repo |
| Nodes | Storage on every worker |
| Replicas | 3 by default (per-volume) |
csi-attacher stuck-leader recovery
The most common Longhorn pain point: a workload shows Multi-Attach error for volume "<pv>" and a VolumeAttachment is stuck with a finalizer.
Decision tree:
flowchart TB
A[Multi-Attach error + stuck VA finalizer] --> B[Check csi-attacher pods]
B --> C{Lease holder healthy?}
C -- yes --> D[Investigate the actual workload]
C -- no, crashlooping --> E[Force-delete the csi-attacher pod]
E --> F[Lease re-elects to a healthy pod]
F --> G[VA finalizer resolves]The exact pods + lease are namespace-prefixed; identify them via kubectl get leases -n longhorn-system and cross-reference with kubectl get pods -n longhorn-system.
Force-delete is only safe when the holder is genuinely crashlooping. If the holder is healthy and the VA is still stuck, the problem is downstream (node loss, taint mismatch, finalizer race) — investigate before swinging the hammer.
Operator quick-reference
| Symptom | First check |
|---|---|
Pod stuck ContainerCreating with Multi-Attach | kubectl get volumeattachment for the PV; identify lease holder |
| Volume "degraded" in Longhorn UI | One replica is down — usually a worker that's cordoned or storage drift |
| Snapshots failing | Longhorn-side snapshot quota, not a workload issue |
| New PVCs not provisioning | kubectl get sc — is longhorn the default? |
R2-D2 workload notes
- Restreamer — small PVC for the process DB; can survive a fresh PVC at the cost of re-creating processes from the dashboard.
- TrailBase / ReductStore — both rely on PVCs surviving pod restarts; treat them as durable.
- TBMQ redis-cluster — see the TBMQ page for the
prevention block already applied and the
nodes.confrecovery recipe.
Scale-downs during Longhorn maintenance
For R2-D2's workloads, scale down anything multi-replica before maintenance and rebuild after. The redis-cluster behind TBMQ is the most sensitive — see TBMQ for the prevention block + recovery procedure.
See also
- TrailBase, ReductStore, Restreamer, TBMQ — R2-D2's PVC consumers
- kube-vip — sibling cluster infra