Adds an experimental power-cap mode using the undocumented RM ioctl interface (based on panchovix's LACT PR #1205) to set power limits below the VBIOS minimum (down to 30 W). - hal/rm_power.py: RM ioctl power-cap read/write/reset + runtime probe - limits.py: power_cap_mode (nvml/ioctl) with support detection - config.py: persist power_cap_mode per GPU - profiles: record/apply power_cap_mode - server.py: POST /api/limits validates ioctl support (409 on failure) - cli.py: profile save falls back to persisted mode - client.py: power_cap_mode in Limits - frontend: toggle + warning with panchovix attribution (LACT #1205) - tests: test_rm_power.py (unit) + integration coverage - Makefile: add test_rm_power.py to make test Also includes automated linter reformatting (prettier, ruff, shellcheck, isort, markdownlint) that the linter would apply anyway.
304 lines
12 KiB
Markdown
304 lines
12 KiB
Markdown
# Usage Guide
|
|
|
|
This guide covers the day-to-day use of NVCurve, from launching the web UI to managing profiles and using the CLI.
|
|
|
|
## First-Time Setup
|
|
|
|
Before using NVCurve, always run the hardware compatibility check:
|
|
|
|
```bash
|
|
nvcurve setup
|
|
```
|
|
|
|
This verifies that your GPU and driver support the required NvAPI functions, performs a safe test write, and restores your GPU to its original state. Only proceed if the result is **Compatible**.
|
|
|
|
## Web UI
|
|
|
|
The web UI is the primary interface for curve editing and monitoring.
|
|
|
|
### Launching the Web UI
|
|
|
|
```bash
|
|
nvcurve
|
|
```
|
|
|
|
This starts the backend server and opens your browser at `http://127.0.0.1:8042`.
|
|
|
|
### Server Management
|
|
|
|
```bash
|
|
nvcurve serve start # Start the web server
|
|
nvcurve serve start --detach # Start in the background
|
|
nvcurve serve status # Check if the server is running
|
|
nvcurve serve stop # Stop the server
|
|
```
|
|
|
|
### Curve Editor
|
|
|
|
The curve editor displays your GPU's V/F curve as an interactive graph with draggable points.
|
|
|
|
| Action | How |
|
|
| --- | --- |
|
|
| Select a point | Click on it |
|
|
| Multi-select | Shift+click to add/remove points |
|
|
| Select all active points | Ctrl/Cmd+A |
|
|
| Clear selection | Escape |
|
|
| Edit a point | Drag it to stage a frequency offset |
|
|
| Move multiple points | Drag with multiple points selected |
|
|
| Exact value input | Select one point, press Enter |
|
|
| Nudge ±1 MHz | Arrow keys (↑ / ↓) |
|
|
| Nudge ±10 MHz | Ctrl/Cmd + arrow keys |
|
|
| Step through points | Tab / Shift+Tab |
|
|
| Box select | Shift+drag on the background |
|
|
| Pan the graph | Drag on the background |
|
|
| Zoom | Alt+scroll |
|
|
|
|
### Point Table
|
|
|
|
The point table provides a spreadsheet-like view of all curve points with their frequency, voltage, and offset values.
|
|
|
|
| Action | How |
|
|
| --- | --- |
|
|
| Select a point | Click a row |
|
|
| Toggle selection | Ctrl/Cmd+click |
|
|
| Range select | Shift+click or drag across rows |
|
|
| Select before/after | Click "Select Before" or "Select After" (appears for single selection) |
|
|
|
|
### Toolbar
|
|
|
|
The toolbar provides bulk operations:
|
|
|
|
- **Global Offset Slider** — Appears when all active points share a uniform delta. Adjusts every point simultaneously.
|
|
- **Flatten** — When two or more points are selected, flattens them to the anchor point's frequency. The anchor is the last explicitly clicked point (highlighted with an amber halo).
|
|
- **Apply** — Commits staged changes to the GPU.
|
|
- **Reset** — Clears all offsets back to zero.
|
|
|
|
### Live Monitoring
|
|
|
|
The monitoring panel shows real-time GPU metrics:
|
|
|
|
- Voltage
|
|
- Clock speed
|
|
- Temperature
|
|
- Power draw
|
|
|
|
Data is streamed via WebSocket from the backend at a configurable poll interval (default: 1 second).
|
|
|
|
### Performance Limits
|
|
|
|
The Performance panel controls the board power limit and the memory clock offset. The power limit slider is bounded by the GPU's VBIOS minimum and maximum (shown at the slider ends); changes are applied on **Apply** and reset to the hardware default on **Reset**.
|
|
|
|
#### Experimental NVIDIA power control
|
|
|
|
On compatible drivers, an **Experimental NVIDIA power control** checkbox appears in the Performance panel. Enabling it switches power-limit application from the standard NVML call to an undocumented driver (RM) interface, which **permits caps below the VBIOS minimum, down to 30 W**. The native maximum still applies.
|
|
|
|
> **Warning.** This uses an undocumented driver interface for *all* power limits, including resets. It may cause instability or stop working after a driver update. Enable it at your own risk. The option is clearly labelled with a red warning in the UI, and profiles saved while it is enabled are marked accordingly.
|
|
|
|
Notes:
|
|
|
|
- The checkbox only appears when the driver exposes a compatible RM power layout (detected with a read-only probe — no writes).
|
|
- In this mode there is **no automatic fallback** to NVML: if the RM route fails, the error is reported rather than silently switching backends.
|
|
- The mode is per-GPU and persisted across server restarts. Reset restores the default through the same route, so it can also clear a previously-set below-minimum cap.
|
|
- The CLI reports availability via `nvcurve read --diag` ("Experimental RM power: available").
|
|
|
|
### Multi-GPU
|
|
|
|
When multiple NVIDIA GPUs are detected, a GPU selector dropdown appears in the status bar. Switching GPUs resets pending edits, selection state, and monitoring for the new target.
|
|
|
|
## Authentication (Multi-User)
|
|
|
|
By default the server runs **without** authentication — anyone who can reach the port can use it. This is fine for a single-user workstation, but on a shared AI server you will want to lock it down. nvcurve uses a **dual mode**:
|
|
|
|
- **No users configured** → the API and web UI are open, exactly as before.
|
|
- **One or more users configured** → every API and WebSocket endpoint requires a login. The web UI shows a sign-in screen first.
|
|
|
|
There is no "single shared password" mode — once you add a user, each person gets their own account.
|
|
|
|
### Managing users
|
|
|
|
Users are stored as **bcrypt** hashes in `/etc/nvcurve/users.json` (mode `0600`, root-owned). Plaintext passwords are never written to disk. Manage them with the CLI (root required to add/remove/change):
|
|
|
|
```bash
|
|
# Add a user (prompts for the password twice). Never pass the password as an
|
|
# argument — it would be visible in the process list and recorded in sudo logs.
|
|
sudo nvcurve user add alice
|
|
sudo nvcurve user add bob
|
|
|
|
# List users
|
|
nvcurve user list
|
|
|
|
# Change a user's password
|
|
sudo nvcurve user set-password alice
|
|
|
|
# Remove a user
|
|
sudo nvcurve user remove bob
|
|
```
|
|
|
|
Adding the first user **switches the server into authenticated mode immediately** (no restart needed). Removing the last user switches it back to open mode.
|
|
|
|
### Signing in
|
|
|
|
- **Web UI:** open the app as usual; if authentication is enabled you will see a sign-in screen. Enter your username and password.
|
|
- **CLI / scripts:** the Python client can authenticate and reuse the session:
|
|
|
|
```python
|
|
from nvcurve.client import NvCurveClient
|
|
client = NvCurveClient(base="http://127.0.0.1:8042")
|
|
client.login("alice", "S3cret!") # stores the session token
|
|
print(client.gpu()) # subsequent calls are authorized
|
|
```
|
|
|
|
Or pass a token you already have: `NvCurveClient(base=..., token="...")`.
|
|
|
|
### Sessions
|
|
|
|
- A successful login creates a session that **lasts 24 hours**, after which a new login is required.
|
|
- Browsers receive the session as an `HttpOnly` cookie; CLI/scripts use the returned token as an `Authorization: Bearer <token>` header.
|
|
- Sessions are kept in server memory, so a server restart invalidates them (users must sign in again).
|
|
- A per-IP lockout (10 failed attempts within 5 minutes → 15-minute lockout) slows down brute-force guessing.
|
|
|
|
### Security notes
|
|
|
|
- The user store file should stay root-owned and `0600` (the CLI enforces this).
|
|
- The web UI and API are still only as safe as the network path to the server — bind to a trusted interface (`--host`) and/or firewall the port. Authentication protects against casual access, not a determined network attacker.
|
|
- The `nvcurve user` commands and the user store require root; day-to-day sign-in does not.
|
|
- The login lockout is keyed by client IP. Behind a reverse proxy all clients share the proxy's IP — set `trusted_proxies` in `/etc/nvcurve/config.json` (e.g. `["127.0.0.1"]`) so the lockout uses the real client IP from `X-Forwarded-For`. The header is only honoured for peers you list there (it is spoofable otherwise).
|
|
- On shared systems consider setting `allow_api_shutdown: false` so users cannot stop the server via the API (manage it with systemd instead).
|
|
- The frequency safety cap (`max_delta_khz`) is enforced by the server from its config; API clients cannot raise it per request. The CLI's `--max-delta` (root-only, direct hardware path) can still override it for a single write.
|
|
|
|
## TLS (HTTPS)
|
|
|
|
The server speaks **plain HTTP by default**. When you expose it beyond localhost, enable TLS so credentials and session cookies are not sent in cleartext:
|
|
|
|
```bash
|
|
# Persistent (stored in /etc/nvcurve/config.json, used by the daemon too)
|
|
sudo nvcurve service configure --ssl-certfile /path/to/cert.pem --ssl-keyfile /path/to/key.pem
|
|
|
|
# One-off
|
|
nvcurve serve start --ssl-certfile /path/to/cert.pem --ssl-keyfile /path/to/key.pem
|
|
```
|
|
|
|
- Both files must be set for TLS to activate; the UI then lives at `https://<host>:8042` and the WebSocket upgrades to `wss://` automatically.
|
|
- To disable TLS again: `sudo nvcurve service configure --no-ssl` (removes the certificate/key from the config).
|
|
- The session cookie gets the `Secure` flag, so it is only sent over HTTPS.
|
|
- A self-signed certificate is fine for a home LAN (the browser shows a warning); for multi-user setups use a certificate your browser trusts.
|
|
- The CLI detects TLS from the config/runtime info and switches to `https://` automatically.
|
|
|
|
## CLI Reference
|
|
|
|
The CLI is designed for scripting, headless use, and quick operations. All write commands support `--dry-run` to preview changes.
|
|
|
|
### Reading the Curve
|
|
|
|
```bash
|
|
nvcurve read # Condensed V/F curve
|
|
nvcurve read --full # All points including zeros
|
|
nvcurve read --json # JSON output for scripting
|
|
```
|
|
|
|
### Writing Offsets
|
|
|
|
```bash
|
|
# Preview changes without applying
|
|
nvcurve write --global --delta 50 --dry-run
|
|
nvcurve write --point 80 --delta 100 --dry-run
|
|
|
|
# Apply changes
|
|
nvcurve write --global --delta 50 # All active GPU points
|
|
nvcurve write --point 80 --delta 100 # Single point
|
|
nvcurve write --range 70-90 --delta 75 # Range of points
|
|
nvcurve write --reset # Reset all to zero
|
|
```
|
|
|
|
> If you use LACT or similar tools, disable them before writing. Concurrent writes will overwrite each other.
|
|
|
|
### Snapshots
|
|
|
|
Snapshots are saved automatically before every write. You can also manage them manually:
|
|
|
|
```bash
|
|
nvcurve snapshot save
|
|
nvcurve snapshot list
|
|
nvcurve snapshot restore
|
|
```
|
|
|
|
Snapshots are stored in `/var/cache/nvcurve/snapshots`.
|
|
|
|
### Profiles
|
|
|
|
Profiles save a named set of curve offsets (plus optional memory offset and power limit settings). Profile commands work whether or not the server is running.
|
|
|
|
```bash
|
|
nvcurve profile save my_profile # Save current curve state
|
|
nvcurve profile apply my_profile # Apply a saved profile
|
|
nvcurve profile list # List all profiles
|
|
nvcurve profile default my_profile # Set as auto-load on startup
|
|
nvcurve profile default --clear # Clear auto-load setting
|
|
```
|
|
|
|
Profiles are stored as JSON files in `/etc/nvcurve/profiles`.
|
|
|
|
### Diagnostics
|
|
|
|
```bash
|
|
nvcurve read --diag # Full NvAPI function probe + system info
|
|
nvcurve inspect --point 80 # Raw ClockBoostTable fields
|
|
nvcurve inspect --range 78-82 # Inspect a range of points
|
|
nvcurve gpus # List detected GPUs
|
|
```
|
|
|
|
### Global Arguments
|
|
|
|
The `--gpu N` flag can be placed before or after any subcommand to target a specific GPU in multi-GPU setups:
|
|
|
|
```bash
|
|
nvcurve --gpu 1 read
|
|
nvcurve read --gpu 1
|
|
nvcurve --gpu 1 profile default perf
|
|
```
|
|
|
|
## Systemd Service
|
|
|
|
Install NVCurve as a systemd service for automatic profile loading on boot:
|
|
|
|
### Install
|
|
|
|
```bash
|
|
nvcurve service install
|
|
```
|
|
|
|
With optional web server auto-start:
|
|
|
|
```bash
|
|
nvcurve service install --auto-serve --host 0.0.0.0 --port 8042
|
|
```
|
|
|
|
### Manage
|
|
|
|
```bash
|
|
nvcurve service start
|
|
nvcurve service stop
|
|
nvcurve service restart
|
|
nvcurve service status
|
|
nvcurve service uninstall
|
|
```
|
|
|
|
### Reconfigure
|
|
|
|
```bash
|
|
sudo nvcurve service configure --auto-serve
|
|
sudo nvcurve service configure --no-auto-serve
|
|
sudo nvcurve service configure --host 0.0.0.0 --port 8042
|
|
sudo nvcurve service configure --ssl-certfile /path/to/cert.pem --ssl-keyfile /path/to/key.pem
|
|
```
|
|
|
|
## Configuration Files
|
|
|
|
| File | Purpose |
|
|
| --- | --- |
|
|
| `/etc/nvcurve/config.json` | Persistent config (host, port, auto-serve, TLS, safety cap, default profiles) |
|
|
| `/etc/nvcurve/profiles/*.json` | Saved profiles |
|
|
| `/var/cache/nvcurve/snapshots/` | Auto-saved snapshots before writes |
|
|
| `/run/nvcurve.json` | Runtime server info (host, port, PID) |
|
|
| `/etc/systemd/system/nvcurve.service` | Systemd unit file |
|