Stale CSI Mount — Workspace Stuck Pending
Summary
Workspace pods (b70553ae, ebafc5da, d048256b) were stuck in Creating phase for 7–17 days. Root cause: a deleted Longhorn PV (pvc-81b676e6) left a stale device mount on worker-04, which blocked the kubelet's serialized NestedPendingOperations queue. Every pod needing a CSI volume mount on that node was deadlocked.
Symptoms
- Workspace CRD stuck in
Creatingphase - Pod is
PendingwithPodScheduled=Truebut zero containers - PVC shows
Boundand volume attachment showsATTACHED: true - Kubelet logs:
NestedPendingOperations: GetDeviceMountRefs check failedwith exponential backoff - No
MountDevice/MapVolumeevents for the stuck pods - Controller logs
RestartGeneration bumped in Creating phaseevery 2s (hot loop bug)
Diagnosis Commands
# Check pod state — PodScheduled=True but no containerStatuses
kubectl get pod <POD> -n llmsafespaces -o jsonpath='{range .status.conditions[*]}{.type}={.status} {end}'
# Check kubelet logs for volume operation errors
talosctl --nodes <NODE_IP> logs kubelet | grep -iE 'NestedPendingOperations|GetDeviceMountRefs|context deadline exceeded'
# Check for stale mounts
talosctl --nodes <NODE_IP> mounts | grep longhorn
# Check CSI plugin health
kubectl get pods -n longhorn-system -o wide | grep csi-plugin
# Check controller hot-loop
kubectl logs -n llmsafespaces deploy/llmsafespaces-controller | grep 'RestartGeneration bumped'
Remediation
Option 1: Kubelet restart (first attempt)
talosctl --nodes <NODE_IP> service kubelet restart
Works if the stale mount is the only issue. CSI plugin and volume queue reset.
Option 2: Restart CSI plugin (if kubelet restart insufficient)
kubectl delete pod <csi-plugin-pod> -n longhorn-system
Option 3: Clear stale snapshot pods (if kubelet workers saturated)
# Delete stale non-running pods blocking kubelet sync workers
kubectl get pods -A --field-selector spec.nodeName=<NODE>,status.phase!=Running
kubectl delete pod <stale-pod> -n <ns> --force --grace-period=0
Option 4: Full node reboot (nuclear option)
If kubelet internal state is irrecoverably corrupted (broken prober manager, stuck volume cleanup):
kubectl cordon <NODE>
talosctl --nodes <NODE_IP> reboot
# Wait for node to return
kubectl uncordon <NODE>
If Talos is stuck in Rebooting stage (hung NFS, disk I/O errors),
a hard power cycle (IPMI/Redfish/physical) is required.
Post-recovery: verify workspaces
kubectl get pods -n llmsafespaces | grep <WORKSPACE-ID>
kubectl get workspace -n llmsafespaces | grep <WORKSPACE-ID>
# Should be Running 1/1 and Active
Root Cause Analysis
-
Stale CSI mount: Deleted PV left a device mount reference on disk that the kubelet couldn't clean up (the owning pod was already gone from the API). The
NestedPendingOperationsqueue is serialized — one stuck operation blocks ALL volume operations on the node. -
Controller hot-loop bug (
phase_creating.go):ObservedRestartGenerationwas set in-memory but never persisted to etcd on the Pending-scheduled path. The controller re-read the stale value every 2s, detected the bump again, and logged indefinitely. -
No stuck-scheduled detection: FN3 recovery only caught
Unschedulablepods. A pod that was scheduled but couldn't mount volumes (stale CSI, dead plugin, kubelet queue saturation) was never detected — it sat Pending with zero containers indefinitely.
Prevention (PR #601)
Deploy controller fixes from https://github.com/lenaxia/LLMSafeSpaces/pull/601:
-
Hot-loop fix: Persist
ObservedRestartGenerationviaStatus().Update()before the catch-all return inhandleCreating. -
Stuck-scheduled detection (FN3b): New check —
PodScheduled=True+ zero container/init statuses + >10min → Infrastructure recovery (delete pod, exponential backoff). This would have caught the issue in 10 minutes instead of 7–17 days. -
Node version alignment: worker-04 runs containerd 2.2.5 / Talos v1.13.6 while the rest of the fleet runs 2.1.6 / v1.12.x. Align versions to reduce CSI/volume handling risk.
Timeline
| Time | Event |
|---|---|
| -7d to -17d | Workspaces entered Creating, pods scheduled to worker-04 |
| T+0 | Diagnosis: stale CSI mount blocking kubelet volume queue |
| T+15m | Kubelet restart — cleared stale mount, volumes attached |
| T+30m | CSI plugin restart — cleared stale CSI registrations |
| T+45m | Pods stuck: prober manager corrupted, backlog of stale snapshot pods |
| T+60m | Deleted stale pods, kubelet still degraded |
| T+90m | Cordoned + attempted drain (pods stuck terminating) |
| T+95m | Talos reboot — stuck on hung NFS + disk I/O errors |
| T+120m | Hard power cycle — node came back clean |
| T+125m | All 3 workspaces Active, node Ready, 2 GPUs detected |