Files
nvcurve/docs/Usage-Guide.md
T
ARIA e148c83622 feat: experimental NVIDIA power control via RM ioctl interface
Adds an experimental power-cap mode using the undocumented RM ioctl
interface (based on panchovix's LACT PR #1205) to set power limits
below the VBIOS minimum (down to 30 W).

- hal/rm_power.py: RM ioctl power-cap read/write/reset + runtime probe
- limits.py: power_cap_mode (nvml/ioctl) with support detection
- config.py: persist power_cap_mode per GPU
- profiles: record/apply power_cap_mode
- server.py: POST /api/limits validates ioctl support (409 on failure)
- cli.py: profile save falls back to persisted mode
- client.py: power_cap_mode in Limits
- frontend: toggle + warning with panchovix attribution (LACT #1205)
- tests: test_rm_power.py (unit) + integration coverage
- Makefile: add test_rm_power.py to make test

Also includes automated linter reformatting (prettier, ruff, shellcheck,
isort, markdownlint) that the linter would apply anyway.
2026-09-17 22:44:27 +02:00

304 lines
12 KiB
Markdown

# Usage Guide
This guide covers the day-to-day use of NVCurve, from launching the web UI to managing profiles and using the CLI.
## First-Time Setup
Before using NVCurve, always run the hardware compatibility check:
```bash
nvcurve setup
```
This verifies that your GPU and driver support the required NvAPI functions, performs a safe test write, and restores your GPU to its original state. Only proceed if the result is **Compatible**.
## Web UI
The web UI is the primary interface for curve editing and monitoring.
### Launching the Web UI
```bash
nvcurve
```
This starts the backend server and opens your browser at `http://127.0.0.1:8042`.
### Server Management
```bash
nvcurve serve start # Start the web server
nvcurve serve start --detach # Start in the background
nvcurve serve status # Check if the server is running
nvcurve serve stop # Stop the server
```
### Curve Editor
The curve editor displays your GPU's V/F curve as an interactive graph with draggable points.
| Action | How |
| --- | --- |
| Select a point | Click on it |
| Multi-select | Shift+click to add/remove points |
| Select all active points | Ctrl/Cmd+A |
| Clear selection | Escape |
| Edit a point | Drag it to stage a frequency offset |
| Move multiple points | Drag with multiple points selected |
| Exact value input | Select one point, press Enter |
| Nudge ±1 MHz | Arrow keys (↑ / ↓) |
| Nudge ±10 MHz | Ctrl/Cmd + arrow keys |
| Step through points | Tab / Shift+Tab |
| Box select | Shift+drag on the background |
| Pan the graph | Drag on the background |
| Zoom | Alt+scroll |
### Point Table
The point table provides a spreadsheet-like view of all curve points with their frequency, voltage, and offset values.
| Action | How |
| --- | --- |
| Select a point | Click a row |
| Toggle selection | Ctrl/Cmd+click |
| Range select | Shift+click or drag across rows |
| Select before/after | Click "Select Before" or "Select After" (appears for single selection) |
### Toolbar
The toolbar provides bulk operations:
- **Global Offset Slider** — Appears when all active points share a uniform delta. Adjusts every point simultaneously.
- **Flatten** — When two or more points are selected, flattens them to the anchor point's frequency. The anchor is the last explicitly clicked point (highlighted with an amber halo).
- **Apply** — Commits staged changes to the GPU.
- **Reset** — Clears all offsets back to zero.
### Live Monitoring
The monitoring panel shows real-time GPU metrics:
- Voltage
- Clock speed
- Temperature
- Power draw
Data is streamed via WebSocket from the backend at a configurable poll interval (default: 1 second).
### Performance Limits
The Performance panel controls the board power limit and the memory clock offset. The power limit slider is bounded by the GPU's VBIOS minimum and maximum (shown at the slider ends); changes are applied on **Apply** and reset to the hardware default on **Reset**.
#### Experimental NVIDIA power control
On compatible drivers, an **Experimental NVIDIA power control** checkbox appears in the Performance panel. Enabling it switches power-limit application from the standard NVML call to an undocumented driver (RM) interface, which **permits caps below the VBIOS minimum, down to 30 W**. The native maximum still applies.
> **Warning.** This uses an undocumented driver interface for *all* power limits, including resets. It may cause instability or stop working after a driver update. Enable it at your own risk. The option is clearly labelled with a red warning in the UI, and profiles saved while it is enabled are marked accordingly.
Notes:
- The checkbox only appears when the driver exposes a compatible RM power layout (detected with a read-only probe — no writes).
- In this mode there is **no automatic fallback** to NVML: if the RM route fails, the error is reported rather than silently switching backends.
- The mode is per-GPU and persisted across server restarts. Reset restores the default through the same route, so it can also clear a previously-set below-minimum cap.
- The CLI reports availability via `nvcurve read --diag` ("Experimental RM power: available").
### Multi-GPU
When multiple NVIDIA GPUs are detected, a GPU selector dropdown appears in the status bar. Switching GPUs resets pending edits, selection state, and monitoring for the new target.
## Authentication (Multi-User)
By default the server runs **without** authentication — anyone who can reach the port can use it. This is fine for a single-user workstation, but on a shared AI server you will want to lock it down. nvcurve uses a **dual mode**:
- **No users configured** → the API and web UI are open, exactly as before.
- **One or more users configured** → every API and WebSocket endpoint requires a login. The web UI shows a sign-in screen first.
There is no "single shared password" mode — once you add a user, each person gets their own account.
### Managing users
Users are stored as **bcrypt** hashes in `/etc/nvcurve/users.json` (mode `0600`, root-owned). Plaintext passwords are never written to disk. Manage them with the CLI (root required to add/remove/change):
```bash
# Add a user (prompts for the password twice). Never pass the password as an
# argument — it would be visible in the process list and recorded in sudo logs.
sudo nvcurve user add alice
sudo nvcurve user add bob
# List users
nvcurve user list
# Change a user's password
sudo nvcurve user set-password alice
# Remove a user
sudo nvcurve user remove bob
```
Adding the first user **switches the server into authenticated mode immediately** (no restart needed). Removing the last user switches it back to open mode.
### Signing in
- **Web UI:** open the app as usual; if authentication is enabled you will see a sign-in screen. Enter your username and password.
- **CLI / scripts:** the Python client can authenticate and reuse the session:
```python
from nvcurve.client import NvCurveClient
client = NvCurveClient(base="http://127.0.0.1:8042")
client.login("alice", "S3cret!") # stores the session token
print(client.gpu()) # subsequent calls are authorized
```
Or pass a token you already have: `NvCurveClient(base=..., token="...")`.
### Sessions
- A successful login creates a session that **lasts 24 hours**, after which a new login is required.
- Browsers receive the session as an `HttpOnly` cookie; CLI/scripts use the returned token as an `Authorization: Bearer <token>` header.
- Sessions are kept in server memory, so a server restart invalidates them (users must sign in again).
- A per-IP lockout (10 failed attempts within 5 minutes → 15-minute lockout) slows down brute-force guessing.
### Security notes
- The user store file should stay root-owned and `0600` (the CLI enforces this).
- The web UI and API are still only as safe as the network path to the server — bind to a trusted interface (`--host`) and/or firewall the port. Authentication protects against casual access, not a determined network attacker.
- The `nvcurve user` commands and the user store require root; day-to-day sign-in does not.
- The login lockout is keyed by client IP. Behind a reverse proxy all clients share the proxy's IP — set `trusted_proxies` in `/etc/nvcurve/config.json` (e.g. `["127.0.0.1"]`) so the lockout uses the real client IP from `X-Forwarded-For`. The header is only honoured for peers you list there (it is spoofable otherwise).
- On shared systems consider setting `allow_api_shutdown: false` so users cannot stop the server via the API (manage it with systemd instead).
- The frequency safety cap (`max_delta_khz`) is enforced by the server from its config; API clients cannot raise it per request. The CLI's `--max-delta` (root-only, direct hardware path) can still override it for a single write.
## TLS (HTTPS)
The server speaks **plain HTTP by default**. When you expose it beyond localhost, enable TLS so credentials and session cookies are not sent in cleartext:
```bash
# Persistent (stored in /etc/nvcurve/config.json, used by the daemon too)
sudo nvcurve service configure --ssl-certfile /path/to/cert.pem --ssl-keyfile /path/to/key.pem
# One-off
nvcurve serve start --ssl-certfile /path/to/cert.pem --ssl-keyfile /path/to/key.pem
```
- Both files must be set for TLS to activate; the UI then lives at `https://<host>:8042` and the WebSocket upgrades to `wss://` automatically.
- To disable TLS again: `sudo nvcurve service configure --no-ssl` (removes the certificate/key from the config).
- The session cookie gets the `Secure` flag, so it is only sent over HTTPS.
- A self-signed certificate is fine for a home LAN (the browser shows a warning); for multi-user setups use a certificate your browser trusts.
- The CLI detects TLS from the config/runtime info and switches to `https://` automatically.
## CLI Reference
The CLI is designed for scripting, headless use, and quick operations. All write commands support `--dry-run` to preview changes.
### Reading the Curve
```bash
nvcurve read # Condensed V/F curve
nvcurve read --full # All points including zeros
nvcurve read --json # JSON output for scripting
```
### Writing Offsets
```bash
# Preview changes without applying
nvcurve write --global --delta 50 --dry-run
nvcurve write --point 80 --delta 100 --dry-run
# Apply changes
nvcurve write --global --delta 50 # All active GPU points
nvcurve write --point 80 --delta 100 # Single point
nvcurve write --range 70-90 --delta 75 # Range of points
nvcurve write --reset # Reset all to zero
```
> If you use LACT or similar tools, disable them before writing. Concurrent writes will overwrite each other.
### Snapshots
Snapshots are saved automatically before every write. You can also manage them manually:
```bash
nvcurve snapshot save
nvcurve snapshot list
nvcurve snapshot restore
```
Snapshots are stored in `/var/cache/nvcurve/snapshots`.
### Profiles
Profiles save a named set of curve offsets (plus optional memory offset and power limit settings). Profile commands work whether or not the server is running.
```bash
nvcurve profile save my_profile # Save current curve state
nvcurve profile apply my_profile # Apply a saved profile
nvcurve profile list # List all profiles
nvcurve profile default my_profile # Set as auto-load on startup
nvcurve profile default --clear # Clear auto-load setting
```
Profiles are stored as JSON files in `/etc/nvcurve/profiles`.
### Diagnostics
```bash
nvcurve read --diag # Full NvAPI function probe + system info
nvcurve inspect --point 80 # Raw ClockBoostTable fields
nvcurve inspect --range 78-82 # Inspect a range of points
nvcurve gpus # List detected GPUs
```
### Global Arguments
The `--gpu N` flag can be placed before or after any subcommand to target a specific GPU in multi-GPU setups:
```bash
nvcurve --gpu 1 read
nvcurve read --gpu 1
nvcurve --gpu 1 profile default perf
```
## Systemd Service
Install NVCurve as a systemd service for automatic profile loading on boot:
### Install
```bash
nvcurve service install
```
With optional web server auto-start:
```bash
nvcurve service install --auto-serve --host 0.0.0.0 --port 8042
```
### Manage
```bash
nvcurve service start
nvcurve service stop
nvcurve service restart
nvcurve service status
nvcurve service uninstall
```
### Reconfigure
```bash
sudo nvcurve service configure --auto-serve
sudo nvcurve service configure --no-auto-serve
sudo nvcurve service configure --host 0.0.0.0 --port 8042
sudo nvcurve service configure --ssl-certfile /path/to/cert.pem --ssl-keyfile /path/to/key.pem
```
## Configuration Files
| File | Purpose |
| --- | --- |
| `/etc/nvcurve/config.json` | Persistent config (host, port, auto-serve, TLS, safety cap, default profiles) |
| `/etc/nvcurve/profiles/*.json` | Saved profiles |
| `/var/cache/nvcurve/snapshots/` | Auto-saved snapshots before writes |
| `/run/nvcurve.json` | Runtime server info (host, port, PID) |
| `/etc/systemd/system/nvcurve.service` | Systemd unit file |