feat: experimental NVIDIA power control via RM ioctl interface

Adds an experimental power-cap mode using the undocumented RM ioctl
interface (based on panchovix's LACT PR #1205) to set power limits
below the VBIOS minimum (down to 30 W).

- hal/rm_power.py: RM ioctl power-cap read/write/reset + runtime probe
- limits.py: power_cap_mode (nvml/ioctl) with support detection
- config.py: persist power_cap_mode per GPU
- profiles: record/apply power_cap_mode
- server.py: POST /api/limits validates ioctl support (409 on failure)
- cli.py: profile save falls back to persisted mode
- client.py: power_cap_mode in Limits
- frontend: toggle + warning with panchovix attribution (LACT #1205)
- tests: test_rm_power.py (unit) + integration coverage
- Makefile: add test_rm_power.py to make test

Also includes automated linter reformatting (prettier, ruff, shellcheck,
isort, markdownlint) that the linter would apply anyway.
This commit is contained in:
ARIA committed 2026-09-17 22:44:27 +02:00
1 parent a8462e696c
commit e148c83622
21 files changed
+1641 -271

No files matched your search

+17
View File
@@ -84,6 +84,23 @@ The monitoring panel shows real-time GPU metrics:
Data is streamed via WebSocket from the backend at a configurable poll interval (default: 1 second).
### Performance Limits
The Performance panel controls the board power limit and the memory clock offset. The power limit slider is bounded by the GPU's VBIOS minimum and maximum (shown at the slider ends); changes are applied on **Apply** and reset to the hardware default on **Reset**.
#### Experimental NVIDIA power control
On compatible drivers, an **Experimental NVIDIA power control** checkbox appears in the Performance panel. Enabling it switches power-limit application from the standard NVML call to an undocumented driver (RM) interface, which **permits caps below the VBIOS minimum, down to 30 W**. The native maximum still applies.
> **Warning.** This uses an undocumented driver interface for *all* power limits, including resets. It may cause instability or stop working after a driver update. Enable it at your own risk. The option is clearly labelled with a red warning in the UI, and profiles saved while it is enabled are marked accordingly.
Notes:
- The checkbox only appears when the driver exposes a compatible RM power layout (detected with a read-only probe — no writes).
- In this mode there is **no automatic fallback** to NVML: if the RM route fails, the error is reported rather than silently switching backends.
- The mode is per-GPU and persisted across server restarts. Reset restores the default through the same route, so it can also clear a previously-set below-minimum cap.
- The CLI reports availability via `nvcurve read --diag` ("Experimental RM power: available").
### Multi-GPU
When multiple NVIDIA GPUs are detected, a GPU selector dropdown appears in the status bar. Switching GPUs resets pending edits, selection state, and monitoring for the new target.