Like 0

Diagnosis and remediation for workspace pods stuck in Creating due to stale CSI mounts / kubelet volume queue deadlocks

stale-csi-mount-runbook.md Raw

Stale CSI Mount — Workspace Stuck Pending

Summary

Workspace pods (b70553ae, ebafc5da, d048256b) were stuck in Creating phase for 7–17 days. Root cause: a deleted Longhorn PV (pvc-81b676e6) left a stale device mount on worker-04, which blocked the kubelet's serialized NestedPendingOperations queue. Every pod needing a CSI volume mount on that node was deadlocked.

Symptoms

  • Workspace CRD stuck in Creating phase
  • Pod is Pending with PodScheduled=True but zero containers
  • PVC shows Bound and volume attachment shows ATTACHED: true
  • Kubelet logs: NestedPendingOperations: GetDeviceMountRefs check failed with exponential backoff
  • No MountDevice/MapVolume events for the stuck pods
  • Controller logs RestartGeneration bumped in Creating phase every 2s (hot loop bug)

Diagnosis Commands

# Check pod state — PodScheduled=True but no containerStatuses
kubectl get pod <POD> -n llmsafespaces -o jsonpath='{range .status.conditions[*]}{.type}={.status} {end}'

# Check kubelet logs for volume operation errors
talosctl --nodes <NODE_IP> logs kubelet | grep -iE 'NestedPendingOperations|GetDeviceMountRefs|context deadline exceeded'

# Check for stale mounts
talosctl --nodes <NODE_IP> mounts | grep longhorn

# Check CSI plugin health
kubectl get pods -n longhorn-system -o wide | grep csi-plugin

# Check controller hot-loop
kubectl logs -n llmsafespaces deploy/llmsafespaces-controller | grep 'RestartGeneration bumped'

Remediation

Option 1: Kubelet restart (first attempt)

talosctl --nodes <NODE_IP> service kubelet restart

Works if the stale mount is the only issue. CSI plugin and volume queue reset.

Option 2: Restart CSI plugin (if kubelet restart insufficient)

kubectl delete pod <csi-plugin-pod> -n longhorn-system

Option 3: Clear stale snapshot pods (if kubelet workers saturated)

# Delete stale non-running pods blocking kubelet sync workers
kubectl get pods -A --field-selector spec.nodeName=<NODE>,status.phase!=Running
kubectl delete pod <stale-pod> -n <ns> --force --grace-period=0

Option 4: Full node reboot (nuclear option)

If kubelet internal state is irrecoverably corrupted (broken prober manager, stuck volume cleanup):

kubectl cordon <NODE>
talosctl --nodes <NODE_IP> reboot
# Wait for node to return
kubectl uncordon <NODE>

If Talos is stuck in Rebooting stage (hung NFS, disk I/O errors), a hard power cycle (IPMI/Redfish/physical) is required.

Post-recovery: verify workspaces

kubectl get pods -n llmsafespaces | grep <WORKSPACE-ID>
kubectl get workspace -n llmsafespaces | grep <WORKSPACE-ID>
# Should be Running 1/1 and Active

Root Cause Analysis

  1. Stale CSI mount: Deleted PV left a device mount reference on disk that the kubelet couldn't clean up (the owning pod was already gone from the API). The NestedPendingOperations queue is serialized — one stuck operation blocks ALL volume operations on the node.

  2. Controller hot-loop bug (phase_creating.go): ObservedRestartGeneration was set in-memory but never persisted to etcd on the Pending-scheduled path. The controller re-read the stale value every 2s, detected the bump again, and logged indefinitely.

  3. No stuck-scheduled detection: FN3 recovery only caught Unschedulable pods. A pod that was scheduled but couldn't mount volumes (stale CSI, dead plugin, kubelet queue saturation) was never detected — it sat Pending with zero containers indefinitely.

Prevention (PR #601)

Deploy controller fixes from https://github.com/lenaxia/LLMSafeSpaces/pull/601:

  1. Hot-loop fix: Persist ObservedRestartGeneration via Status().Update() before the catch-all return in handleCreating.

  2. Stuck-scheduled detection (FN3b): New check — PodScheduled=True + zero container/init statuses + >10min → Infrastructure recovery (delete pod, exponential backoff). This would have caught the issue in 10 minutes instead of 7–17 days.

  3. Node version alignment: worker-04 runs containerd 2.2.5 / Talos v1.13.6 while the rest of the fleet runs 2.1.6 / v1.12.x. Align versions to reduce CSI/volume handling risk.

Timeline

Time Event
-7d to -17d Workspaces entered Creating, pods scheduled to worker-04
T+0 Diagnosis: stale CSI mount blocking kubelet volume queue
T+15m Kubelet restart — cleared stale mount, volumes attached
T+30m CSI plugin restart — cleared stale CSI registrations
T+45m Pods stuck: prober manager corrupted, backlog of stale snapshot pods
T+60m Deleted stale pods, kubelet still degraded
T+90m Cordoned + attempted drain (pods stuck terminating)
T+95m Talos reboot — stuck on hung NFS + disk I/O errors
T+120m Hard power cycle — node came back clean
T+125m All 3 workspaces Active, node Ready, 2 GPUs detected