feat: GPU process list with per-process VRAM/utilization and kill

Add a GPU process table to the Performance tab showing which processes
use the selected GPU, with per-process VRAM, GPU utilization, CPU, host
memory, and command. Includes a kill endpoint (SIGTERM/SIGKILL, with
optional parent kill) behind a confirm dialog.

- nvcurve/hal/processes.py: NVML + /proc process enumeration
- server.py: GET /api/processes, POST /api/processes/kill (refuses to
  signal init/kthreadd)
- frontend: ProcessList component, API client, types
- tests: test_processes.py
- docs: API reference for the new endpoints
This commit is contained in:
ARIA committed 2026-10-03 20:25:33 +02:00
1 parent b75b9d43e9
commit c11ea73ead
10 files changed
+1008 -23

No files matched your search

+2
View File
@@ -125,6 +125,8 @@ Notes:
| `GET /api/gpu?gpu_index=0` | Name, driver version, VRAM |
| `GET /api/dashboard?gpu_index=0` | Static dashboard info (VBIOS, CUDA cores, PCIe, BAR1, …). Live values come from the monitor WebSocket |
| `GET /api/monitor?gpu_index=0` | One-shot monitoring snapshot: voltage, clocks, temp, power, fans, p-state, VRAM, utilization, throttle reasons |
| `GET /api/processes?gpu_index=0` | Processes using this GPU, sorted by VRAM (descending): pid, user, dev, type (C/G/C+G), gpu_util_pct + mem_util_pct (last ~1 s only), vram_bytes, vram_pct, cpu_pct, mem_host_bytes, command |
| `POST /api/processes/kill` | Body `{pid, signal: "TERM" \| "KILL", parent: false}` (default TERM) → send signal. `parent: true` signals the process's parent instead (response reports the signalled pid). Errors: `400` bad pid/signal or no parent, `404` process not found, `403` permission denied. **Note:** the pid need not be a GPU process — this is an arbitrary-pid signal primitive. The server runs as root, so with no users configured (open API) any reachable client can signal any process; enable auth before exposing the port on a shared system |
### Curve