# Stale CSI Mount — Workspace Stuck Pending ## Summary Workspace pods (b70553ae, ebafc5da, d048256b) were stuck in `Creating` phase for 7–17 days. Root cause: a deleted Longhorn PV (`pvc-81b676e6`) left a stale device mount on `worker-04`, which blocked the kubelet's serialized `NestedPendingOperations` queue. Every pod needing a CSI volume mount on that node was deadlocked. ## Symptoms - Workspace CRD stuck in `Creating` phase - Pod is `Pending` with `PodScheduled=True` but zero containers - PVC shows `Bound` and volume attachment shows `ATTACHED: true` - Kubelet logs: `NestedPendingOperations: GetDeviceMountRefs check failed` with exponential backoff - No `MountDevice`/`MapVolume` events for the stuck pods - Controller logs `RestartGeneration bumped in Creating phase` every 2s (hot loop bug) ## Diagnosis Commands ```bash # Check pod state — PodScheduled=True but no containerStatuses kubectl get pod -n llmsafespaces -o jsonpath='{range .status.conditions[*]}{.type}={.status} {end}' # Check kubelet logs for volume operation errors talosctl --nodes logs kubelet | grep -iE 'NestedPendingOperations|GetDeviceMountRefs|context deadline exceeded' # Check for stale mounts talosctl --nodes mounts | grep longhorn # Check CSI plugin health kubectl get pods -n longhorn-system -o wide | grep csi-plugin # Check controller hot-loop kubectl logs -n llmsafespaces deploy/llmsafespaces-controller | grep 'RestartGeneration bumped' ``` ## Remediation ### Option 1: Kubelet restart (first attempt) ```bash talosctl --nodes service kubelet restart ``` Works if the stale mount is the only issue. CSI plugin and volume queue reset. ### Option 2: Restart CSI plugin (if kubelet restart insufficient) ```bash kubectl delete pod -n longhorn-system ``` ### Option 3: Clear stale snapshot pods (if kubelet workers saturated) ```bash # Delete stale non-running pods blocking kubelet sync workers kubectl get pods -A --field-selector spec.nodeName=,status.phase!=Running kubectl delete pod -n --force --grace-period=0 ``` ### Option 4: Full node reboot (nuclear option) If kubelet internal state is irrecoverably corrupted (broken prober manager, stuck volume cleanup): ```bash kubectl cordon talosctl --nodes reboot # Wait for node to return kubectl uncordon ``` **If Talos is stuck in `Rebooting` stage** (hung NFS, disk I/O errors), a hard power cycle (IPMI/Redfish/physical) is required. ### Post-recovery: verify workspaces ```bash kubectl get pods -n llmsafespaces | grep kubectl get workspace -n llmsafespaces | grep # Should be Running 1/1 and Active ``` ## Root Cause Analysis 1. **Stale CSI mount**: Deleted PV left a device mount reference on disk that the kubelet couldn't clean up (the owning pod was already gone from the API). The `NestedPendingOperations` queue is serialized — one stuck operation blocks ALL volume operations on the node. 2. **Controller hot-loop bug** (`phase_creating.go`): `ObservedRestartGeneration` was set in-memory but never persisted to etcd on the Pending-scheduled path. The controller re-read the stale value every 2s, detected the bump again, and logged indefinitely. 3. **No stuck-scheduled detection**: FN3 recovery only caught `Unschedulable` pods. A pod that was scheduled but couldn't mount volumes (stale CSI, dead plugin, kubelet queue saturation) was never detected — it sat Pending with zero containers indefinitely. ## Prevention (PR #601) Deploy controller fixes from https://github.com/lenaxia/LLMSafeSpaces/pull/601: 1. **Hot-loop fix**: Persist `ObservedRestartGeneration` via `Status().Update()` before the catch-all return in `handleCreating`. 2. **Stuck-scheduled detection (FN3b)**: New check — `PodScheduled=True` + zero container/init statuses + >10min → Infrastructure recovery (delete pod, exponential backoff). This would have caught the issue in 10 minutes instead of 7–17 days. 3. **Node version alignment**: worker-04 runs containerd 2.2.5 / Talos v1.13.6 while the rest of the fleet runs 2.1.6 / v1.12.x. Align versions to reduce CSI/volume handling risk. ## Timeline | Time | Event | |------|-------| | -7d to -17d | Workspaces entered Creating, pods scheduled to worker-04 | | T+0 | Diagnosis: stale CSI mount blocking kubelet volume queue | | T+15m | Kubelet restart — cleared stale mount, volumes attached | | T+30m | CSI plugin restart — cleared stale CSI registrations | | T+45m | Pods stuck: prober manager corrupted, backlog of stale snapshot pods | | T+60m | Deleted stale pods, kubelet still degraded | | T+90m | Cordoned + attempted drain (pods stuck terminating) | | T+95m | Talos reboot — stuck on hung NFS + disk I/O errors | | T+120m | Hard power cycle — node came back clean | | T+125m | All 3 workspaces Active, node Ready, 2 GPUs detected |