Like 0

Diagnosis and remediation for workspace pods stuck in Creating due to stale CSI mounts / kubelet volume queue deadlocks

mike revised this gist 1 month ago · a0e1b00

1 file changed, 128 insertions

stale-csi-mount-runbook.md (file created)
@@ -0,0 +1,128 @@
1 + # Stale CSI Mount — Workspace Stuck Pending
2 +
3 + ## Summary
4 +
5 + Workspace pods (b70553ae, ebafc5da, d048256b) were stuck in `Creating` phase for 7–17 days. Root cause: a deleted Longhorn PV (`pvc-81b676e6`) left a stale device mount on `worker-04`, which blocked the kubelet's serialized `NestedPendingOperations` queue. Every pod needing a CSI volume mount on that node was deadlocked.
6 +
7 + ## Symptoms
8 +
9 + - Workspace CRD stuck in `Creating` phase
10 + - Pod is `Pending` with `PodScheduled=True` but zero containers
11 + - PVC shows `Bound` and volume attachment shows `ATTACHED: true`
12 + - Kubelet logs: `NestedPendingOperations: GetDeviceMountRefs check failed` with exponential backoff
13 + - No `MountDevice`/`MapVolume` events for the stuck pods
14 + - Controller logs `RestartGeneration bumped in Creating phase` every 2s (hot loop bug)
15 +
16 + ## Diagnosis Commands
17 +
18 + ```bash
19 + # Check pod state — PodScheduled=True but no containerStatuses
20 + kubectl get pod <POD> -n llmsafespaces -o jsonpath='{range .status.conditions[*]}{.type}={.status} {end}'
21 +
22 + # Check kubelet logs for volume operation errors
23 + talosctl --nodes <NODE_IP> logs kubelet | grep -iE 'NestedPendingOperations|GetDeviceMountRefs|context deadline exceeded'
24 +
25 + # Check for stale mounts
26 + talosctl --nodes <NODE_IP> mounts | grep longhorn
27 +
28 + # Check CSI plugin health
29 + kubectl get pods -n longhorn-system -o wide | grep csi-plugin
30 +
31 + # Check controller hot-loop
32 + kubectl logs -n llmsafespaces deploy/llmsafespaces-controller | grep 'RestartGeneration bumped'
33 + ```
34 +
35 + ## Remediation
36 +
37 + ### Option 1: Kubelet restart (first attempt)
38 +
39 + ```bash
40 + talosctl --nodes <NODE_IP> service kubelet restart
41 + ```
42 +
43 + Works if the stale mount is the only issue. CSI plugin and volume queue reset.
44 +
45 + ### Option 2: Restart CSI plugin (if kubelet restart insufficient)
46 +
47 + ```bash
48 + kubectl delete pod <csi-plugin-pod> -n longhorn-system
49 + ```
50 +
51 + ### Option 3: Clear stale snapshot pods (if kubelet workers saturated)
52 +
53 + ```bash
54 + # Delete stale non-running pods blocking kubelet sync workers
55 + kubectl get pods -A --field-selector spec.nodeName=<NODE>,status.phase!=Running
56 + kubectl delete pod <stale-pod> -n <ns> --force --grace-period=0
57 + ```
58 +
59 + ### Option 4: Full node reboot (nuclear option)
60 +
61 + If kubelet internal state is irrecoverably corrupted (broken prober manager,
62 + stuck volume cleanup):
63 +
64 + ```bash
65 + kubectl cordon <NODE>
66 + talosctl --nodes <NODE_IP> reboot
67 + # Wait for node to return
68 + kubectl uncordon <NODE>
69 + ```
70 +
71 + **If Talos is stuck in `Rebooting` stage** (hung NFS, disk I/O errors),
72 + a hard power cycle (IPMI/Redfish/physical) is required.
73 +
74 + ### Post-recovery: verify workspaces
75 +
76 + ```bash
77 + kubectl get pods -n llmsafespaces | grep <WORKSPACE-ID>
78 + kubectl get workspace -n llmsafespaces | grep <WORKSPACE-ID>
79 + # Should be Running 1/1 and Active
80 + ```
81 +
82 + ## Root Cause Analysis
83 +
84 + 1. **Stale CSI mount**: Deleted PV left a device mount reference on disk that
85 + the kubelet couldn't clean up (the owning pod was already gone from the API).
86 + The `NestedPendingOperations` queue is serialized — one stuck operation blocks
87 + ALL volume operations on the node.
88 +
89 + 2. **Controller hot-loop bug** (`phase_creating.go`): `ObservedRestartGeneration`
90 + was set in-memory but never persisted to etcd on the Pending-scheduled path.
91 + The controller re-read the stale value every 2s, detected the bump again,
92 + and logged indefinitely.
93 +
94 + 3. **No stuck-scheduled detection**: FN3 recovery only caught `Unschedulable`
95 + pods. A pod that was scheduled but couldn't mount volumes (stale CSI, dead
96 + plugin, kubelet queue saturation) was never detected — it sat Pending with
97 + zero containers indefinitely.
98 +
99 + ## Prevention (PR #601)
100 +
101 + Deploy controller fixes from https://github.com/lenaxia/LLMSafeSpaces/pull/601:
102 +
103 + 1. **Hot-loop fix**: Persist `ObservedRestartGeneration` via `Status().Update()`
104 + before the catch-all return in `handleCreating`.
105 +
106 + 2. **Stuck-scheduled detection (FN3b)**: New check — `PodScheduled=True` +
107 + zero container/init statuses + >10min → Infrastructure recovery (delete pod,
108 + exponential backoff). This would have caught the issue in 10 minutes instead
109 + of 7–17 days.
110 +
111 + 3. **Node version alignment**: worker-04 runs containerd 2.2.5 / Talos v1.13.6
112 + while the rest of the fleet runs 2.1.6 / v1.12.x. Align versions to reduce
113 + CSI/volume handling risk.
114 +
115 + ## Timeline
116 +
117 + | Time | Event |
118 + |------|-------|
119 + | -7d to -17d | Workspaces entered Creating, pods scheduled to worker-04 |
120 + | T+0 | Diagnosis: stale CSI mount blocking kubelet volume queue |
121 + | T+15m | Kubelet restart — cleared stale mount, volumes attached |
122 + | T+30m | CSI plugin restart — cleared stale CSI registrations |
123 + | T+45m | Pods stuck: prober manager corrupted, backlog of stale snapshot pods |
124 + | T+60m | Deleted stale pods, kubelet still degraded |
125 + | T+90m | Cordoned + attempted drain (pods stuck terminating) |
126 + | T+95m | Talos reboot — stuck on hung NFS + disk I/O errors |
127 + | T+120m | Hard power cycle — node came back clean |
128 + | T+125m | All 3 workspaces Active, node Ready, 2 GPUs detected |