| @@ -0,0 +1,128 @@ | |||
| 1 | + | # Stale CSI Mount — Workspace Stuck Pending | |
| 2 | + | ||
| 3 | + | ## Summary | |
| 4 | + | ||
| 5 | + | Workspace pods (b70553ae, ebafc5da, d048256b) were stuck in `Creating` phase for 7–17 days. Root cause: a deleted Longhorn PV (`pvc-81b676e6`) left a stale device mount on `worker-04`, which blocked the kubelet's serialized `NestedPendingOperations` queue. Every pod needing a CSI volume mount on that node was deadlocked. | |
| 6 | + | ||
| 7 | + | ## Symptoms | |
| 8 | + | ||
| 9 | + | - Workspace CRD stuck in `Creating` phase | |
| 10 | + | - Pod is `Pending` with `PodScheduled=True` but zero containers | |
| 11 | + | - PVC shows `Bound` and volume attachment shows `ATTACHED: true` | |
| 12 | + | - Kubelet logs: `NestedPendingOperations: GetDeviceMountRefs check failed` with exponential backoff | |
| 13 | + | - No `MountDevice`/`MapVolume` events for the stuck pods | |
| 14 | + | - Controller logs `RestartGeneration bumped in Creating phase` every 2s (hot loop bug) | |
| 15 | + | ||
| 16 | + | ## Diagnosis Commands | |
| 17 | + | ||
| 18 | + | ```bash | |
| 19 | + | # Check pod state — PodScheduled=True but no containerStatuses | |
| 20 | + | kubectl get pod <POD> -n llmsafespaces -o jsonpath='{range .status.conditions[*]}{.type}={.status} {end}' | |
| 21 | + | ||
| 22 | + | # Check kubelet logs for volume operation errors | |
| 23 | + | talosctl --nodes <NODE_IP> logs kubelet | grep -iE 'NestedPendingOperations|GetDeviceMountRefs|context deadline exceeded' | |
| 24 | + | ||
| 25 | + | # Check for stale mounts | |
| 26 | + | talosctl --nodes <NODE_IP> mounts | grep longhorn | |
| 27 | + | ||
| 28 | + | # Check CSI plugin health | |
| 29 | + | kubectl get pods -n longhorn-system -o wide | grep csi-plugin | |
| 30 | + | ||
| 31 | + | # Check controller hot-loop | |
| 32 | + | kubectl logs -n llmsafespaces deploy/llmsafespaces-controller | grep 'RestartGeneration bumped' | |
| 33 | + | ``` | |
| 34 | + | ||
| 35 | + | ## Remediation | |
| 36 | + | ||
| 37 | + | ### Option 1: Kubelet restart (first attempt) | |
| 38 | + | ||
| 39 | + | ```bash | |
| 40 | + | talosctl --nodes <NODE_IP> service kubelet restart | |
| 41 | + | ``` | |
| 42 | + | ||
| 43 | + | Works if the stale mount is the only issue. CSI plugin and volume queue reset. | |
| 44 | + | ||
| 45 | + | ### Option 2: Restart CSI plugin (if kubelet restart insufficient) | |
| 46 | + | ||
| 47 | + | ```bash | |
| 48 | + | kubectl delete pod <csi-plugin-pod> -n longhorn-system | |
| 49 | + | ``` | |
| 50 | + | ||
| 51 | + | ### Option 3: Clear stale snapshot pods (if kubelet workers saturated) | |
| 52 | + | ||
| 53 | + | ```bash | |
| 54 | + | # Delete stale non-running pods blocking kubelet sync workers | |
| 55 | + | kubectl get pods -A --field-selector spec.nodeName=<NODE>,status.phase!=Running | |
| 56 | + | kubectl delete pod <stale-pod> -n <ns> --force --grace-period=0 | |
| 57 | + | ``` | |
| 58 | + | ||
| 59 | + | ### Option 4: Full node reboot (nuclear option) | |
| 60 | + | ||
| 61 | + | If kubelet internal state is irrecoverably corrupted (broken prober manager, | |
| 62 | + | stuck volume cleanup): | |
| 63 | + | ||
| 64 | + | ```bash | |
| 65 | + | kubectl cordon <NODE> | |
| 66 | + | talosctl --nodes <NODE_IP> reboot | |
| 67 | + | # Wait for node to return | |
| 68 | + | kubectl uncordon <NODE> | |
| 69 | + | ``` | |
| 70 | + | ||
| 71 | + | **If Talos is stuck in `Rebooting` stage** (hung NFS, disk I/O errors), | |
| 72 | + | a hard power cycle (IPMI/Redfish/physical) is required. | |
| 73 | + | ||
| 74 | + | ### Post-recovery: verify workspaces | |
| 75 | + | ||
| 76 | + | ```bash | |
| 77 | + | kubectl get pods -n llmsafespaces | grep <WORKSPACE-ID> | |
| 78 | + | kubectl get workspace -n llmsafespaces | grep <WORKSPACE-ID> | |
| 79 | + | # Should be Running 1/1 and Active | |
| 80 | + | ``` | |
| 81 | + | ||
| 82 | + | ## Root Cause Analysis | |
| 83 | + | ||
| 84 | + | 1. **Stale CSI mount**: Deleted PV left a device mount reference on disk that | |
| 85 | + | the kubelet couldn't clean up (the owning pod was already gone from the API). | |
| 86 | + | The `NestedPendingOperations` queue is serialized — one stuck operation blocks | |
| 87 | + | ALL volume operations on the node. | |
| 88 | + | ||
| 89 | + | 2. **Controller hot-loop bug** (`phase_creating.go`): `ObservedRestartGeneration` | |
| 90 | + | was set in-memory but never persisted to etcd on the Pending-scheduled path. | |
| 91 | + | The controller re-read the stale value every 2s, detected the bump again, | |
| 92 | + | and logged indefinitely. | |
| 93 | + | ||
| 94 | + | 3. **No stuck-scheduled detection**: FN3 recovery only caught `Unschedulable` | |
| 95 | + | pods. A pod that was scheduled but couldn't mount volumes (stale CSI, dead | |
| 96 | + | plugin, kubelet queue saturation) was never detected — it sat Pending with | |
| 97 | + | zero containers indefinitely. | |
| 98 | + | ||
| 99 | + | ## Prevention (PR #601) | |
| 100 | + | ||
| 101 | + | Deploy controller fixes from https://github.com/lenaxia/LLMSafeSpaces/pull/601: | |
| 102 | + | ||
| 103 | + | 1. **Hot-loop fix**: Persist `ObservedRestartGeneration` via `Status().Update()` | |
| 104 | + | before the catch-all return in `handleCreating`. | |
| 105 | + | ||
| 106 | + | 2. **Stuck-scheduled detection (FN3b)**: New check — `PodScheduled=True` + | |
| 107 | + | zero container/init statuses + >10min → Infrastructure recovery (delete pod, | |
| 108 | + | exponential backoff). This would have caught the issue in 10 minutes instead | |
| 109 | + | of 7–17 days. | |
| 110 | + | ||
| 111 | + | 3. **Node version alignment**: worker-04 runs containerd 2.2.5 / Talos v1.13.6 | |
| 112 | + | while the rest of the fleet runs 2.1.6 / v1.12.x. Align versions to reduce | |
| 113 | + | CSI/volume handling risk. | |
| 114 | + | ||
| 115 | + | ## Timeline | |
| 116 | + | ||
| 117 | + | | Time | Event | | |
| 118 | + | |------|-------| | |
| 119 | + | | -7d to -17d | Workspaces entered Creating, pods scheduled to worker-04 | | |
| 120 | + | | T+0 | Diagnosis: stale CSI mount blocking kubelet volume queue | | |
| 121 | + | | T+15m | Kubelet restart — cleared stale mount, volumes attached | | |
| 122 | + | | T+30m | CSI plugin restart — cleared stale CSI registrations | | |
| 123 | + | | T+45m | Pods stuck: prober manager corrupted, backlog of stale snapshot pods | | |
| 124 | + | | T+60m | Deleted stale pods, kubelet still degraded | | |
| 125 | + | | T+90m | Cordoned + attempted drain (pods stuck terminating) | | |
| 126 | + | | T+95m | Talos reboot — stuck on hung NFS + disk I/O errors | | |
| 127 | + | | T+120m | Hard power cycle — node came back clean | | |
| 128 | + | | T+125m | All 3 workspaces Active, node Ready, 2 GPUs detected | | |
mike / Stale CSI Mount — Workspace Stuck Pending Runbook
Last active 1 month ago
Diagnosis and remediation for workspace pods stuck in Creating due to stale CSI mounts / kubelet volume queue deadlocks