Compare commits

...
25 Commits
Author SHA1 Message Date
Pakobbix bf00e1149a Merge pull request 'feat: wireview pin-imbalance warning, per-pin rating bars, pin_imbalance_a' (#18) from feat/wireview-imbalance-warning into main
Reviewed-on: #18
2026-10-03 20:15:52 +00:00
ARIA b69b7502b3 feat: wireview pin-imbalance warning, per-pin rating bars, pin_imbalance_a
- Per-pin bar now scales against the 9.2 A per-pin rating (10 A full
  scale, red at/over rating) instead of the 55 A total connector OCP
- Debounced app-wide warning banner (visible on all tabs) when the
  current between two pins diverges by > 2 A for ~7 s - the signature
  of a bad pin contact that can overheat and melt the 12VHPWR
  connector; 5 s grace period after clearing
- New pin_imbalance_a sample field (max - min pin current) in both the
  serial and hwmon sample builders, plus frontend type
- Docs: README + WireView.md cover the warning and threshold; fresh
  screenshot with the new per-pin labels
- Tests: vitest setup + 9 unit tests for the debounce state machine;
  3 new backend checks for pin_imbalance_a
2026-10-03 22:13:13 +02:00
Pakobbix 92356d1409 Merge pull request 'fix: set SO_REUSEADDR on port pre-check so restarts don't wait for TIME_WAIT' (#16) from fix/reuseaddr-port-check into main
Reviewed-on: #16
2026-10-03 19:18:59 +00:00
ARIA 5cfadb4d9b fix: set SO_REUSEADDR on port pre-check so restarts don't wait for TIME_WAIT
The fail-fast port check bound a plain socket, so TIME_WAIT sockets left
by a just-stopped instance (~60s) made it report 'port already in use'
even though uvicorn (which sets SO_REUSEADDR) could bind fine. Mirror
uvicorn's bind semantics: the check now only fails when the port is
genuinely held by a live listener.
2026-10-03 21:18:22 +02:00
Pakobbix 3c0fee3ec8 Merge pull request 'feat: Report Bug button that creates Gitea issues' (#15) from feat/report-bug into main
Reviewed-on: #15
2026-10-03 19:07:44 +00:00
ARIA 1bafbc80b5 feat: Report Bug button that creates Gitea issues
Adds a 'Report Bug' button to the web UI (top bar) that files an issue
on Gitea via the daemon. The issue body is enriched with auto-collected
context: nvcurve version, hostname, GPU name, applied offset summary,
active profile, and a collapsed tail of the recent server log.

Security model:
- The Gitea token (nvcurve_bot, scope write:issue) is embedded as a
  built-in default so the feature works out of the box for every
  install of this public repo — no configuration required.
- The token is only ever used server-side; it never reaches the
  browser, logs, or API responses.
- The endpoint is behind the existing session auth when enabled, and
  rate-limited to one report per 5 minutes per client IP.
- config.json can override gitea_url/gitea_repo/gitea_token (e.g.
  forks pointing at their own repo); gitea_token: '' disables it.
2026-10-03 21:06:50 +02:00
Pakobbix e89135ef94 Merge pull request 'feat: GPU process list with per-process VRAM/utilization and kill' (#12) from feat/gpu-process-list into main
Reviewed-on: #12
2026-10-03 18:26:19 +00:00
ARIA c11ea73ead feat: GPU process list with per-process VRAM/utilization and kill
Add a GPU process table to the Performance tab showing which processes
use the selected GPU, with per-process VRAM, GPU utilization, CPU, host
memory, and command. Includes a kill endpoint (SIGTERM/SIGKILL, with
optional parent kill) behind a confirm dialog.

- nvcurve/hal/processes.py: NVML + /proc process enumeration
- server.py: GET /api/processes, POST /api/processes/kill (refuses to
  signal init/kthreadd)
- frontend: ProcessList component, API client, types
- tests: test_processes.py
- docs: API reference for the new endpoints
2026-10-03 20:25:33 +02:00
ARIA b75b9d43e9 docs: add API reference and enrich OpenAPI metadata
- docs/API.md: full REST + WebSocket reference with auth usage
  (cookie/Bearer login flow, curl examples), conventions, endpoint
  tables, and WS protocol (WS routes are not in the OpenAPI schema)
- server.py: app-level description, openapi_tags grouping, and
  tags/responses metadata on all REST routes documenting the error
  codes each route raises; hide the SPA catch-all from the schema
- README: link the API reference in the docs index
2026-10-03 14:51:58 +02:00
ARIA db0e676501 docs: add WireView Pro II setup guide, note one-time host setup in README
The device is not usable out of the box on Linux: without the udev rule
the port is root:uucp 0660 with no stable name, and ModemManager probes
it for ~30 s after every plug. New docs/WireView.md covers the one-time
setup (udev rule, dialout group, port access, verification,
troubleshooting); README links to it from the feature note and the
documentation list.
2026-10-03 13:52:51 +02:00
Pakobbix 38c3a30017 Merge pull request 'feat: native WireView Pro II monitoring (detection, serial protocol, live tab)' (#11) from feat/wireview-monitoring into main
Reviewed-on: #11
2026-10-03 11:43:59 +00:00
ARIA 457ca5c3a6 docs: add WireView tab screenshot to README 2026-10-03 13:43:14 +02:00
ARIA 6f5c046a59 docs: mention WireView Pro II monitoring in README fork-features block 2026-10-03 13:37:39 +02:00
ARIA b8e91e3c91 fix: degrade gracefully when pyserial is missing
A stale venv after a code update (editable install + git pull) left
pyserial out of the environment, and the hard module-level import in
wireview.py took the entire web server down — even on machines without
a WireView device.

Guard the import: without pyserial the serial transport is disabled
(one-time warning, connect fails, reads return None) while the rest of
the server keeps running. The hwmon transport is unaffected.

Also simplify the except clauses to OSError (SerialException is an
OSError subclass) so they no longer reference the possibly-None module.
2026-10-03 13:32:43 +02:00
ARIA 526b71b45f feat: native WireView Pro II monitoring (detection, serial protocol, live tab)
Standalone support for the ThermalGrizzly WireView Pro II without
depending on the external wireview_reporter exporter:

- wireview.py: serial protocol (STX/ETX + 16-bit CRC16-CCITT) with
  vendor-data product identification, config version, UID, build
  string, screen layout, and temperature/power/current sensor reads;
  hwmon fallback with per-channel index resolution; udev-based
  detection (vendor 0x2560 / product 0x0101) with fallback port
  probing; serial read timeout and watchdog reconnect
- server.py: startup detection (root only), 1 Hz poller gated on
  subscribed clients, WS 'wireview' channel, GET /api/wireview,
  rejected-product memoization, stale-read race guard
- frontend: WireView tab (visible when a device is detected) with
  live temperature/power/current cards, sparkline history, and
  device info; nullable temp channels
- tests: 76 tests covering parser, CRC, fault classification,
  hwmon resolution, JSON safety, and the serial transport against
  a pty-based fake device

Verified live: tab appears with the device connected, disappears
when unplugged, reconnects on re-plug, no serial traffic when idle.
2026-10-03 13:20:56 +02:00
Pakobbix cc102f26c1 Merge pull request 'feat: experimental NVIDIA power control via RM ioctl interface' (#10) from feat/rm-power-control into main
Reviewed-on: #10
2026-09-17 20:48:46 +00:00
ARIA e148c83622 feat: experimental NVIDIA power control via RM ioctl interface
Adds an experimental power-cap mode using the undocumented RM ioctl
interface (based on panchovix's LACT PR #1205) to set power limits
below the VBIOS minimum (down to 30 W).

- hal/rm_power.py: RM ioctl power-cap read/write/reset + runtime probe
- limits.py: power_cap_mode (nvml/ioctl) with support detection
- config.py: persist power_cap_mode per GPU
- profiles: record/apply power_cap_mode
- server.py: POST /api/limits validates ioctl support (409 on failure)
- cli.py: profile save falls back to persisted mode
- client.py: power_cap_mode in Limits
- frontend: toggle + warning with panchovix attribution (LACT #1205)
- tests: test_rm_power.py (unit) + integration coverage
- Makefile: add test_rm_power.py to make test

Also includes automated linter reformatting (prettier, ruff, shellcheck,
isort, markdownlint) that the linter would apply anyway.
2026-09-17 22:44:27 +02:00
ARIA a8462e696c Add single-command installation: install.sh, Makefile, frontend build hook
- hatch_build.py: custom hatchling build hook that compiles the React
  frontend (npm ci + build) when frontend/dist is missing or stale, so
  'uv tool install git+https://gitea.zephyre.one/Pakobbix/nvcurve.git'
  works as a single command
- install.sh: curl|bash installer (checks prerequisites, auto-installs uv,
  clones and installs)
- Makefile: dev targets (frontend, install, dev, test, clean)
- README/docs: document the one-liner, clone, and direct-git installs
2026-09-16 10:43:23 +02:00
Pakobbix d810c44478 Merge pull request 'security: harden web server, daemon socket, and write paths' (#9) from security/hardening into main
Reviewed-on: #9
2026-09-10 14:20:57 +00:00
ARIA 39701c12ff security: harden web server, daemon socket, and write paths
Security review findings, fixed and verified:

Critical
- Fix unauthenticated arbitrary file read: the SPA catch-all route
  joined the raw URL path onto the dist dir without containment, so
  encoded '..' segments (/%2e%2e/etc/passwd) leaked any file readable
  by the root server. Resolve with realpath and reject paths outside
  the dist dir (fail-closed 404).

High
- Daemon socket: serve_start no longer accepts caller-chosen
  host/port. The socket is world-connectable (unprivileged CLI users),
  so callers could previously rebind the root web server to 0.0.0.0.
  The daemon now always binds the operator-configured address and
  reports it in the response; the CLI warns on mismatch.

Medium
- Remove the per-request max_delta_khz override from the API: the
  server-enforced safety cap is now authoritative. CLI direct paths
  (write, profile apply, verify) honor the configured cap; --max-delta
  still overrides for explicit root use.
- Snapshot restore: confine filepath to the snapshot directory
  (realpath containment; blocks symlink escapes).
- Login lockout: honor X-Forwarded-For only for peers listed in the
  new trusted_proxies config (rightmost untrusted hop), so the
  per-IP lockout works behind a reverse proxy. Spoofed headers from
  untrusted peers are ignored.
- /api/shutdown: new allow_api_shutdown config (default true);
  shared systems can disable the API shutdown path.

TLS (opt-in, like auth)
- New ssl_certfile/ssl_keyfile config + CLI flags (serve start,
  service install/configure, --no-ssl to disable). When active:
  HTTPS for UI/API, wss:// for WebSockets, Secure session cookie,
  CLI auto-switches to https://. Cert/key paths are validated up
  front with a clear error instead of a silent uvicorn crash.

Tests & docs
- tests/test_security.py: standalone regression tests (no new deps)
  covering SPA containment, snapshot containment, cap removal,
  client-IP derivation, proxy normalization, TLS scheme detection,
  and daemon host/port hardening.
- README + Usage-Guide: TLS section, new config keys, updated
  security notes.
2026-09-10 16:19:21 +02:00
Pakobbix 8956fc9d7b Merge pull request 'feat: full fan control — all fans or individual fans' (#8) from feat/multi-fan-control into main
Reviewed-on: #8
2026-09-10 13:50:55 +00:00
ARIA 6cb33187d3 feat: full fan control — all fans or individual fans
The fan curve previously only controlled fan index 0; secondary fans
stayed on driver control. The curve can now target all fans (new
default) or any individual fan(s).

Backend:
- hal/fans.py: get_num_fans() via nvmlDeviceGetNumFans; get_fan_info()
  returns per-fan speeds; set_fan_speed() accepts a fan index list
  (None = all fans; all-fans mode is lenient toward driver-locked
  fans, explicit lists are strict); reset_fan() restores all fans.
- server.py: per-GPU fan_targets state; the poller applies the curve to
  all target fans and logs write failures (once per distinct error);
  activation validates targets against the hardware (stale indices fall
  back to all fans); POST /api/fans accepts fans, POST /api/fans/speed
  accepts a fan index, GET /api/fans returns num_fans/fans/fan_targets.
- Persistence format is now {"curve": ..., "fans": ...}; legacy
  bare-curve entries migrate to "all fans" at startup.
- Profiles save/apply fan_targets alongside fan_curve.
- MonitoringSample carries per-fan speeds for live gauges.

Frontend:
- Fans tab: All / Fan 1 / Fan 2 / ... selector with live per-fan %;
  the selection is applied together with the curve.
- Live monitor: per-fan gauges with sparklines for multi-fan GPUs.
2026-09-10 15:48:28 +02:00
Pakobbix 34a9bc6d6e Merge pull request 'Clean up LSP diagnostics across backend and frontend' (#7) from cleanup/lint-modernization into main
Reviewed-on: #7
2026-09-08 21:58:42 +00:00
ARIA 930e56bd07 Clean up LSP diagnostics across backend and frontend
Backend (nvcurve/):
- hal/fans.py, hal/limits.py, hal/gpu.py: replace conditional pynvml
  imports with the established 'pynvml: Any = _pynvml_import' pattern
  (fixes ~50 'possibly unbound' errors); type the result dicts; guard
  query_interface() results; explicit uuid/pci-bus parsing (int, hex
  convention documented); modernize Optional[T] -> T | None
- cli.py: fix 'curve_state' possibly-unbound and snap_path None handling
  in cmd_setup; wrap unchecked int()/open()/makedirs() calls in
  try/except with clean CLI errors; add module logger for silent
  except-pass blocks; raise ... from exc; fix unused loop vars and
  set-comprehension
- hal/snapshot.py: filepath: str | None; wrap all file ops; sorted
  imports; remove unused CT_POINTS import
- daemon.py: extract 0o666 to _SOCKET_MODE constant (intentional for
  /run sockets) with nosemgrep
- server.py: nosemgrep for Python 3.7-compat false positive (project
  requires >= 3.12); log previously-swallowed exception
- profiles/native.py, profiles/apply.py: wrap file ops and int(k)
  profile-key parsing; sorted imports; modernize typing

Frontend (frontend/src):
- Add .js extensions to all relative imports (standard TS-ESM; Vite
  resolves .js -> .ts)
- React.FormEvent (deprecated in React 19 types) -> React.SubmitEvent
- catch (e: any) -> catch (e: unknown) + instanceof Error narrowing
- React-hooks: move ref writes from render into effects; convert
  viewport reset to render-phase state adjustment; split
  selectPoint(index, multi) into selectPoint + togglePoint (no flag
  argument); remove non-null assertion
- Static inline styles -> Tailwind classes (dynamic positioning/cursor
  styles kept)
- Remove non-standard 'container' option from scrollIntoView (browsers
  ignore unknown options) which had orphaned a @ts-expect-error
- Object.fromEntries for Map -> Record conversion

Tooling:
- .gitignore: ignore .codegraph/ local tool data

Verified: tsc --noEmit, vite production build, python imports, and
full LSP scan (0 errors/warnings in both projects).
2026-09-08 23:57:30 +02:00
Pakobbix fa944c9576 Update README.md 2026-09-02 15:44:49 +00:00
68 changed files with 9087 additions and 1283 deletions

No files matched your search

+4 -1
View File
@@ -6,6 +6,7 @@ __pycache__/
# Node # Node
frontend/node_modules/ frontend/node_modules/
node_modules/ node_modules/
.vitest/
# IDE/Misc # IDE/Misc
.vscode/ .vscode/
@@ -14,4 +15,6 @@ node_modules/
/dist/ /dist/
/build/ /build/
*.egg-info/ *.egg-info/
.claude .claude
# Local tool data
.codegraph/
+32
View File
@@ -0,0 +1,32 @@
# NVCurve — developer convenience targets.
#
# End users don't need make: run ./install.sh (see README "Installation").
UV ?= uv
NPM ?= npm
.PHONY: help frontend frontend-dev install dev test clean
help: ## Show available targets
@grep -E '^[a-zA-Z_-]+:.*?## ' $(MAKEFILE_LIST) | awk 'BEGIN {FS = ":.*?## "}; {printf " \033[36m%-15s\033[0m %s\n", $$1, $$2}'
frontend: ## Build the React frontend into frontend/dist
cd frontend && $(NPM) ci && $(NPM) run build
frontend-dev: ## Run the Vite dev server (hot reload)
cd frontend && $(NPM) run dev
install: ## Install nvcurve as a uv tool (builds frontend if missing/stale)
$(UV) tool install --force .
dev: ## Create/refresh the dev environment (uv sync)
$(UV) sync
test: ## Run the test suite
$(UV) run python tests/test_security.py
$(UV) run python tests/test_rm_power.py
$(UV) run python tests/test_wireview.py
$(UV) run python tests/test_processes.py
clean: ## Remove build artifacts
rm -rf frontend/dist frontend/node_modules
+50 -13
View File
@@ -13,8 +13,10 @@ NVCurve brings MSI Afterburner-style per-point voltage-frequency curve control t
> [!IMPORTANT] > [!IMPORTANT]
> **Blackwell GPU memory** — This is a specialized fork with extended memory offset support (up to +3000 MHz) for Blackwell GPUs (RTX 50-series). \ > **Blackwell GPU memory** — This is a specialized fork with extended memory offset support (up to +3000 MHz) for Blackwell GPUs (RTX 50-series). \
> **Fan Controls** — There is an additional "Fans" tab to setup a customized fan curve. \ > **Fan Controls** — There is an additional "Fans" tab to setup a customized fan curve, controlling all fans or individual fans. \
> **Dashboard** — The default tab is an Dashboard with additional information (PCIe link speed, VBIOS information, Max Core Clock, Throttle Reason and much much more.) \ > **Dashboard** — The default tab is an Dashboard with additional information (PCIe link speed, VBIOS information, Max Core Clock, Throttle Reason and much much more.) \
> **WireView Pro II** — Native monitoring of the Thermal Grizzly WireView Pro II 12VHPWR connector: a dedicated tab shows live temperature, power and current readings — no external exporter or kernel module required. It also raises a big warning in the app header (visible on all tabs) when the current between two power pins diverges by more than 2 A — the signature of a bad pin contact (faulty cable/adapter) that can overheat and melt the connector. The host needs a one-time setup (udev rule + port access, see [WireView Pro II Setup](docs/WireView.md)); after that the device is auto-detected over USB and the tab only appears while it is connected. \
> **Authentication** — For production deployment, I added authentication with bcrypt hashing to allow only one or multiple people to have access. \
> Installing the pre-built PyPI package will NOT include these features. You must build from source. > Installing the pre-built PyPI package will NOT include these features. You must build from source.
<table> <table>
@@ -26,6 +28,9 @@ NVCurve brings MSI Afterburner-style per-point voltage-frequency curve control t
<td align="center"><img src="docs/performance.png" width="480" alt="Performance"></td> <td align="center"><img src="docs/performance.png" width="480" alt="Performance"></td>
<td align="center"><img src="docs/fans.png" width="480" alt="Fans"></td> <td align="center"><img src="docs/fans.png" width="480" alt="Fans"></td>
</tr> </tr>
<tr>
<td align="center"><img src="docs/wireview.png" width="480" alt="WireView"></td>
</tr>
</table> </table>
## Prerequisites ## Prerequisites
@@ -36,20 +41,28 @@ NVCurve brings MSI Afterburner-style per-point voltage-frequency curve control t
- **[uv](https://docs.astral.sh/uv/)** — Python package manager - **[uv](https://docs.astral.sh/uv/)** — Python package manager
- **Root/sudo access** (required for GPU hardware interactions) - **Root/sudo access** (required for GPU hardware interactions)
## Installation from Source ## Installation
### One-liner
```bash ```bash
git clone <this-repo-url>.git curl -fsSL https://gitea.zephyre.one/Pakobbix/nvcurve/raw/branch/main/install.sh | bash
```
The script checks prerequisites (installs `uv` if missing), clones the repo, and installs NVCurve — the React frontend is compiled automatically during the build.
### From a clone
```bash
git clone https://gitea.zephyre.one/Pakobbix/nvcurve.git
cd nvcurve cd nvcurve
./install.sh
```
# Build the React frontend ### Direct from git (no clone, no script)
cd frontend
npm install
npm run build
cd ..
# Install the Python package (includes bundled frontend) ```bash
uv tool install . uv tool install "git+https://gitea.zephyre.one/Pakobbix/nvcurve.git"
``` ```
After installation, verify hardware compatibility: After installation, verify hardware compatibility:
@@ -77,6 +90,20 @@ sudo nvcurve user remove alice # remove a user
Adding the first user enables authentication immediately; removing the last user disables it. See the [Usage Guide](docs/Usage-Guide.md#authentication-multi-user) for details. Adding the first user enables authentication immediately; removing the last user disables it. See the [Usage Guide](docs/Usage-Guide.md#authentication-multi-user) for details.
## TLS (HTTPS)
The server speaks **plain HTTP by default**. For network access (e.g. behind a reverse proxy or on a LAN), you can enable TLS so the web UI, API, and WebSocket all run over HTTPS — the session cookie is then marked `Secure`.
```bash
# One-off (this server run only)
nvcurve serve start --ssl-certfile /path/to/cert.pem --ssl-keyfile /path/to/key.pem
# Persistent (stored in /etc/nvcurve/config.json; used by the daemon too)
sudo nvcurve service configure --ssl-certfile /path/to/cert.pem --ssl-keyfile /path/to/key.pem
```
With TLS enabled the UI is at `https://<host>:8042` and the CLI switches to `https://` automatically. A self-signed certificate works for local use (the browser will warn); for multi-user setups use a certificate your browser trusts (e.g. via your internal CA or a reverse proxy).
## Systemd Service ## Systemd Service
Install the daemon for automatic profile loading on boot and optional web server auto-start: Install the daemon for automatic profile loading on boot and optional web server auto-start:
@@ -148,6 +175,10 @@ The daemon reads settings from `/etc/nvcurve/config.json`:
"max_delta_khz": 3000000, "max_delta_khz": 3000000,
"auto_snapshot": true, "auto_snapshot": true,
"max_snapshots": 20, "max_snapshots": 20,
"ssl_certfile": null,
"ssl_keyfile": null,
"trusted_proxies": [],
"allow_api_shutdown": true,
"auto_load_profiles": { "auto_load_profiles": {
"idx:0": "my_profile" "idx:0": "my_profile"
} }
@@ -159,9 +190,12 @@ The daemon reads settings from `/etc/nvcurve/config.json`:
| `host` | Web server bind address (`0.0.0.0` for network access) | | `host` | Web server bind address (`0.0.0.0` for network access) |
| `port` | Web server port (default `8042`) | | `port` | Web server port (default `8042`) |
| `auto_serve` | Auto-start web server on boot | | `auto_serve` | Auto-start web server on boot |
| `max_delta_khz` | Safety cap for frequency offsets (default 3000 MHz) | | `max_delta_khz` | Safety cap for frequency offsets (default 3000 MHz). Enforced server-side; API clients cannot raise it per request |
| `auto_snapshot` | Save snapshot before every write | | `auto_snapshot` | Save snapshot before every write |
| `max_snapshots` | Max snapshots to keep (`0` = unlimited) | | `max_snapshots` | Max snapshots to keep (`0` = unlimited) |
| `ssl_certfile` / `ssl_keyfile` | TLS certificate/key — enables HTTPS when both are set (default: off) |
| `trusted_proxies` | Proxy IPs whose `X-Forwarded-For` is trusted for the login lockout (e.g. `["127.0.0.1"]` for a local reverse proxy) |
| `allow_api_shutdown` | Allow authenticated users to stop the server via `POST /api/shutdown` (set `false` on shared systems; use systemd instead) |
| `auto_load_profiles` | Per-GPU profile to apply on boot (`{gpu_key: profile_name}`) | | `auto_load_profiles` | Per-GPU profile to apply on boot (`{gpu_key: profile_name}`) |
The GPU key can be a UUID, `pci:XXXX`, or `idx:N` fallback. Find your GPU key with `nvcurve gpus`. The GPU key can be a UUID, `pci:XXXX`, or `idx:N` fallback. Find your GPU key with `nvcurve gpus`.
@@ -171,17 +205,20 @@ The GPU key can be a UUID, `pci:XXXX`, or `idx:N` fallback. Find your GPU key wi
- **[Overview](docs/Overview.md)** — What it does, capabilities, architecture - **[Overview](docs/Overview.md)** — What it does, capabilities, architecture
- **[Installation](docs/Installation.md)** — Prerequisites, source build, troubleshooting - **[Installation](docs/Installation.md)** — Prerequisites, source build, troubleshooting
- **[Usage Guide](docs/Usage-Guide.md)** — Web UI, CLI reference, systemd service - **[Usage Guide](docs/Usage-Guide.md)** — Web UI, CLI reference, systemd service
- **[API Reference](docs/API.md)** — REST + WebSocket API, authentication, endpoint summary (interactive docs at `/docs`)
- **[Tips and Tricks](docs/Tips-and-Tricks.md)** — Workflows, curve flattening, safety - **[Tips and Tricks](docs/Tips-and-Tricks.md)** — Workflows, curve flattening, safety
- **[WireView Pro II Setup](docs/WireView.md)** — One-time host setup (udev rule + port access) for the 12VHPWR connector monitor
## Upgrading ## Upgrading
```bash ```bash
cd nvcurve cd nvcurve
git pull git pull
cd frontend && npm run build && cd .. uv tool install --force .
uv tool install .
``` ```
The frontend is rebuilt automatically if it is missing or older than the frontend sources. If you modified frontend code locally, run `make frontend` first (or `rm -rf frontend/dist`).
If running as a systemd service: If running as a systemd service:
```bash ```bash
+245
View File
@@ -0,0 +1,245 @@
# NVCurve API Reference
The web server (`nvcurve serve`) exposes a **REST + WebSocket API** on
`http://127.0.0.1:8042` by default. The same server also serves the web UI
SPA, so the API and UI share one port.
**Interactive docs** (served by the running server, no login required):
| URL | Description |
| --- | --- |
| `/docs` | Swagger UI (interactive, per-endpoint request/response metadata) |
| `/redoc` | ReDoc (read-friendly reference) |
| `/openapi.json` | Machine-readable OpenAPI 3.1 schema |
> WebSocket endpoints are **not** part of the OpenAPI schema (OpenAPI has no
> WebSocket operation type) — they are documented in the
> [WebSockets](#websockets) section below.
---
## Authentication
The server has two modes, decided automatically:
- **No users configured** → the API is fully open (no auth).
- **One or more users configured** → every `/api/*` and `/ws/*` endpoint
requires a valid session, except the public endpoints listed below.
Sessions last **24 hours**. Passwords are stored as bcrypt hashes.
### Public endpoints (always reachable)
| Endpoint | Purpose |
| --- | --- |
| `GET /api/ping` | Liveness probe (CLI uses it to detect a running server) |
| `GET /api/auth/status` | Whether auth is required and whether *this* request is authenticated |
| `POST /api/auth/login` | Authenticate, obtain session |
| `POST /api/auth/logout` | End the current session |
### Obtaining a session
```bash
curl -s -X POST http://127.0.0.1:8042/api/auth/login \
-H 'Content-Type: application/json' \
-d '{"username": "alice", "password": "secret"}'
```
Response:
```json
{
"ok": true,
"username": "alice",
"expires_at": 1750000000.0,
"token": "9f2c1a…"
}
```
On success the server also sets an `HttpOnly` cookie `nvcurve_session`
(`Secure` when TLS is enabled). You can use either credential:
- **Cookie** — what browsers use automatically.
- **Bearer token** — what scripts/CLI use:
```bash
TOKEN=$(curl -s -X POST http://127.0.0.1:8042/api/auth/login \
-H 'Content-Type: application/json' \
-d '{"username": "alice", "password": "secret"}' | jq -r .token)
curl -s http://127.0.0.1:8042/api/curve -H "Authorization: Bearer $TOKEN"
```
Notes:
- For WebSockets, the token is read from the `Authorization` header or the
cookie on the handshake. It is **deliberately not** accepted via query
string (the access log records full paths).
- Failed logins are rate-limited per client IP; lockouts return `429`.
Behind a reverse proxy, set `trusted_proxies` in the config so the real
client IP is used.
- `GET /api/auth/status` reports your current state:
```json
{"auth_required": true, "authenticated": true, "username": "alice", "expires_at": 1750000000.0}
```
---
## Conventions
- **`gpu_index`** — GPU-selecting routes take a 0-based `gpu_index` query
parameter (default `0`). Unknown indexes → `404`; an uninitialized GPU →
`503`.
- **Errors** — JSON `{"detail": "..."}`. Validation errors on write endpoints
return `400` with `detail.errors[]` listing the offending points.
- **Safety cap** — curve writes are capped by the server-side
`max_delta_khz` (default ±3000 MHz, set in `/etc/nvcurve/config.json`).
Clients cannot raise it per request.
- **Units** — curve deltas are in **kHz**; clock offsets and memory offsets
are in **MHz**; voltages in µV/mV; power in watts.
- **Side effects** — write endpoints (`/api/curve/*`, `/api/limits`,
`/api/profiles/*/apply`, `/api/snapshot/restore`) change hardware state.
With `auto_snapshot: true` (default) a ClockBoostTable snapshot is saved
before every curve write.
---
## REST endpoints
### Auth
| Method & path | Description |
| --- | --- |
| `GET /api/ping` | Liveness probe → `{"ok": true}` |
| `GET /api/auth/status` | Auth mode + current session state |
| `POST /api/auth/login` | Body `{"username", "password"}` → session token + cookie. Errors: `401` bad credentials, `404` auth disabled, `429` locked out |
| `POST /api/auth/logout` | End session (idempotent) |
| `GET /api/auth/users` | List configured usernames (requires session) |
### GPU
| Method & path | Description |
| --- | --- |
| `GET /api/gpus` | List discovered GPUs → `[{index, name, uuid, pci_bus_id}]` |
| `GET /api/gpu?gpu_index=0` | Name, driver version, VRAM |
| `GET /api/dashboard?gpu_index=0` | Static dashboard info (VBIOS, CUDA cores, PCIe, BAR1, …). Live values come from the monitor WebSocket |
| `GET /api/monitor?gpu_index=0` | One-shot monitoring snapshot: voltage, clocks, temp, power, fans, p-state, VRAM, utilization, throttle reasons |
| `GET /api/processes?gpu_index=0` | Processes using this GPU, sorted by VRAM (descending): pid, user, dev, type (C/G/C+G), gpu_util_pct + mem_util_pct (last ~1 s only), vram_bytes, vram_pct, cpu_pct, mem_host_bytes, command |
| `POST /api/processes/kill` | Body `{pid, signal: "TERM" \| "KILL", parent: false}` (default TERM) → send signal. `parent: true` signals the process's parent instead (response reports the signalled pid). Errors: `400` bad pid/signal or no parent, `404` process not found, `403` permission denied. **Note:** the pid need not be a GPU process — this is an arbitrary-pid signal primitive. The server runs as root, so with no users configured (open API) any reachable client can signal any process; enable auth before exposing the port on a shared system |
### Curve
| Method & path | Description |
| --- | --- |
| `GET /api/curve?gpu_index=0` | Full curve: `{gpu_name, timestamp, points: [{index, freq_khz, volt_mv, delta_khz, effective_freq_khz, domain, …}]}` |
| `GET /api/curve/{point}?gpu_index=0` | Single V/F point detail (`400` if index out of range) |
| `GET /api/ranges?gpu_index=0` | Clock boost domain ranges (min/max offset per domain) |
| `GET /api/voltage?gpu_index=0` | Current core voltage → `{voltage_uv, voltage_mv}` |
| `POST /api/curve/write?gpu_index=0` | Body `{"deltas": {"<point>": <delta_khz>}}` — write per-point frequency offsets. Response includes `ok`, `return_code`, optional `warning` (external change detected) and `freq_warnings` (negative-frequency risk) |
| `POST /api/curve/write/global?gpu_index=0` | Body `{"delta_khz": N}` — uniform offset on all GPU-domain points |
| `POST /api/curve/reset?gpu_index=0` | Reset all frequency offsets to zero |
| `POST /api/curve/verify?gpu_index=0` | Body `{"deltas": {...}}` — write-verify-read cycle; returns per-point `match` results and `collateral_changes` on other points |
### Profiles
Profiles capture the current curve deltas, power limit, memory offset and
active fan curve as a named JSON file.
| Method & path | Description |
| --- | --- |
| `GET /api/profiles?gpu_index=0` | `{profiles: [...], active, auto_load}` |
| `POST /api/profiles?gpu_index=0` | Body `{"name": "..."}` — save current GPU state as a profile |
| `POST /api/profiles/{name}/apply?gpu_index=0` | Apply a saved profile to hardware. `404` if missing; `500` with per-part errors if any part fails |
| `DELETE /api/profiles/{name}` | Delete a profile (also clears auto-load references) |
| `POST /api/profiles/{name}/rename` | Body `{"new_name": "..."}` |
### Config
| Method & path | Description |
| --- | --- |
| `GET /api/config?gpu_index=0` | `{auto_load_profile}` for the GPU |
| `POST /api/config` | Body `{"auto_load_profile": "name" \| null, "gpu_index": 0}` — set/clear the profile auto-applied at server start. Persists to `/etc/nvcurve/config.json` if present |
### Limits
| Method & path | Description |
| --- | --- |
| `GET /api/limits?gpu_index=0` | Power limit (current/default/min/max), `gpc_offset_mhz`, `mem_offset_mhz`, memory offset range |
| `POST /api/limits?gpu_index=0` | Body `{power_limit_w?, mem_offset_mhz?, power_cap_mode?}` — any subset. `power_cap_mode` is `"nvml"` (default) or `"ioctl"` (experimental RM power control; `409` if unsupported). Changing the memory offset may reset the curve table — the server re-applies the last known curve offsets |
| `POST /api/limits/reset?gpu_index=0` | Power limit → hardware default, memory offset → 0 |
### Fans
| Method & path | Description |
| --- | --- |
| `GET /api/fans?gpu_index=0` | Per-fan state: `{fan_pct, fans: [{index, fan_pct}], num_fans, min_fan_pct, max_fan_pct, fan_mode: "auto" \| "curve", curve, curve_active, fan_targets}` |
| `POST /api/fans?gpu_index=0` | Body `{"curve": [{"temp_c": 60, "fan_pct": 40}, …], "fans": [0] \| null}` — set the fan curve and start the control poller. `fans: null` drives all fans. `400` for invalid curves (e.g. non-monotonic temps) |
| `POST /api/fans/reset?gpu_index=0` | Deactivate curve control, restore automatic fan mode |
| `POST /api/fans/speed?gpu_index=0` | Body `{"fan_pct": 50, "fan": 0 \| null}` — one-shot exact speed (bypasses curve) |
Active fan curves persist across server restarts.
### WireView Pro II
| Method & path | Description |
| --- | --- |
| `GET /api/wireview` | `{available, connected, info, sample}` — `available` is true when a Thermal Grizzly WireView Pro II is connected; `sample` is the most recent reading (`null` until the first) |
### Snapshots
| Method & path | Description |
| --- | --- |
| `GET /api/snapshots` | List saved ClockBoostTable snapshots → `[{filepath, timestamp, gpu, nonzero_offsets, size}]` |
| `POST /api/snapshot/save?gpu_index=0` | Save the current ClockBoostTable → `{ok, filepath}` |
| `POST /api/snapshot/restore?gpu_index=0` | Body `{"filepath": "…" \| null}` — restore a snapshot (most recent if omitted) |
### Server
| Method & path | Description |
| --- | --- |
| `POST /api/shutdown` | Gracefully stop the server process. `403` when `allow_api_shutdown: false` (recommended on shared systems — use systemd instead) |
---
## WebSockets
All three endpoints use the same handshake: connect, then send a subscribe
message. The server sends an initial state payload, then keeps streaming.
| Endpoint | Subscribe message | Stream |
| --- | --- | --- |
| `/ws/monitor` | `{"action": "subscribe", "gpu_index": 0}` | Monitoring sample every `poll_interval_s` (default 1 s): voltage, clocks, temp, power, fans, p-state, VRAM, utilization, throttle reasons |
| `/ws/curve` | `{"action": "subscribe", "gpu_index": 0}` | Full curve state, pushed whenever the curve changes (after any write) |
| `/ws/wireview` | `{"action": "subscribe"}` | `{"type": "unavailable"}` or `{"type": "sample", "info": {...}, "sample": {...}}` per poll tick |
Authentication (when enabled): the session token is taken from the
`Authorization: Bearer <token>` header or the `nvcurve_session` cookie on the
handshake. Unauthenticated connections are closed with code `1008`.
Example (Python `websockets`):
```python
import asyncio, json, websockets
async def main():
async with websockets.connect(
"ws://127.0.0.1:8042/ws/monitor",
extra_headers={"Authorization": f"Bearer {token}"},
) as ws:
await ws.send(json.dumps({"action": "subscribe", "gpu_index": 0}))
async for msg in ws:
sample = json.loads(msg)
print(sample["temp_c"], sample["clock_mhz"])
asyncio.run(main())
```
---
## CLI
The `nvcurve` CLI talks to this same API (default base
`http://127.0.0.1:8042`, override with `--base-url`), so any scripted use
covered by the CLI works against the API as well. See
[Usage-Guide.md](Usage-Guide.md) for CLI details.
+29 -19
View File
@@ -29,43 +29,52 @@ sudo pacman -S uv
pip install uv pip install uv
``` ```
## Installation from Source ## Installation
### Step 1: Clone the Repository ### Option 1: One-liner (recommended)
```bash ```bash
git clone <this-repo-url>.git curl -fsSL https://gitea.zephyre.one/Pakobbix/nvcurve/raw/branch/main/install.sh | bash
```
The script checks prerequisites (installs `uv` if missing), clones the repository, and installs NVCurve. The React frontend is compiled automatically during the build by a hatchling build hook (`hatch_build.py`).
### Option 2: From a clone
```bash
git clone https://gitea.zephyre.one/Pakobbix/nvcurve.git
cd nvcurve cd nvcurve
./install.sh
``` ```
### Step 2: Build the Frontend Equivalent manual steps (what the script does):
The frontend is a React + TypeScript + Vite application in the `frontend/` directory.
```bash ```bash
cd frontend git clone https://gitea.zephyre.one/Pakobbix/nvcurve.git
npm install cd nvcurve
npm run build uv tool install . # frontend is built automatically if missing/stale
cd ..
``` ```
This produces a `dist/` directory with the compiled static assets. The hatch build system bundles `frontend/dist` into the Python package. ### Option 3: Direct from git (no clone, no script)
### Step 3: Install the Python Package
```bash ```bash
uv tool install . uv tool install "git+https://gitea.zephyre.one/Pakobbix/nvcurve.git"
``` ```
This installs `nvcurve` as a system-wide tool with the bundled frontend. To install a specific branch:
### Step 4: Verify ```bash
uv tool install "git+https://gitea.zephyre.one/Pakobbix/nvcurve.git@<branch>"
```
### Verify
```bash ```bash
nvcurve setup nvcurve setup
``` ```
This performs four checks: This performs four checks:
1. **NvAPI function probe** — verifies all required functions resolve in your driver 1. **NvAPI function probe** — verifies all required functions resolve in your driver
2. **Curve read** — reads and displays your current V/F curve as a baseline 2. **Curve read** — reads and displays your current V/F curve as a baseline
3. **Write-verify** — writes `+5 MHz` to a safe point, reads it back, and confirms the change 3. **Write-verify** — writes `+5 MHz` to a safe point, reads it back, and confirms the change
@@ -94,10 +103,11 @@ nvcurve serve start
```bash ```bash
cd nvcurve cd nvcurve
git pull git pull
cd frontend && npm run build && cd .. uv tool install --force .
uv tool install .
``` ```
The frontend is rebuilt automatically if it is missing or older than the frontend sources. If you modified frontend code locally, run `make frontend` first (or `rm -rf frontend/dist`).
If running as a systemd service: If running as a systemd service:
```bash ```bash
@@ -116,7 +126,7 @@ source ~/.local/bin/env # or wherever uv installed
### Frontend not loading in the web UI ### Frontend not loading in the web UI
Verify that `frontend/dist` exists and contains built assets. If the directory is empty or missing, rebuild with `npm run build` and reinstall with `uv tool install .`. Verify that the installed package contains the frontend. If `frontend/dist` is empty or missing, rebuild with `make frontend` (or `cd frontend && npm ci && npm run build`) and reinstall with `uv tool install --force .`.
### NvAPI functions not found ### NvAPI functions not found
+2 -2
View File
@@ -17,7 +17,7 @@ NVCurve provides two ways to interact with your GPU:
- **Per-Point Curve Editing** — Adjust the frequency offset for any individual voltage point on the V/F curve. - **Per-Point Curve Editing** — Adjust the frequency offset for any individual voltage point on the V/F curve.
- **Extended Memory Offset** — Memory clock offset up to +3000 MHz (Blackwell GPUs). Standard NVCurve caps at +1000 MHz. - **Extended Memory Offset** — Memory clock offset up to +3000 MHz (Blackwell GPUs). Standard NVCurve caps at +1000 MHz.
- **Fan Curve Control** — Custom temperature-to-fan-speed curves via NVML, adjustable through the web UI and savable in profiles. - **Fan Curve Control** — Custom temperature-to-fan-speed curves via NVML (all fans or individual fans), adjustable through the web UI and savable in profiles.
- **Curve Flattening** — Select multiple points and flatten them to a common frequency using anchor-point targeting. - **Curve Flattening** — Select multiple points and flatten them to a common frequency using anchor-point targeting.
- **Live Monitoring** — Track GPU voltage, clock speed, temperature, and power draw in real time via NvAPI and NVML. - **Live Monitoring** — Track GPU voltage, clock speed, temperature, and power draw in real time via NvAPI and NVML.
- **Profile Management** — Save, load, and switch between named profiles. Set a default profile that auto-applies on startup. - **Profile Management** — Save, load, and switch between named profiles. Set a default profile that auto-applies on startup.
@@ -29,7 +29,7 @@ NVCurve provides two ways to interact with your GPU:
NVCurve consists of two components: NVCurve consists of two components:
| Component | Description | | Component | Description |
|---|---| | --- | --- |
| **Python Backend** | Talks directly to `libnvidia-api.so` (via ctypes) and `libnvidia-ml.so` to read and write GPU hardware state. Exposes functionality through a FastAPI REST + WebSocket server. | | **Python Backend** | Talks directly to `libnvidia-api.so` (via ctypes) and `libnvidia-ml.so` to read and write GPU hardware state. Exposes functionality through a FastAPI REST + WebSocket server. |
| **React Frontend** | Runs in the browser and communicates with the backend over HTTP and WebSockets. Handles curve visualization, point editing, live monitoring, and profile management. | | **React Frontend** | Runs in the browser and communicates with the backend over HTTP and WebSockets. Handles curve visualization, point editing, live monitoring, and profile management. |
+40 -1
View File
@@ -84,6 +84,23 @@ The monitoring panel shows real-time GPU metrics:
Data is streamed via WebSocket from the backend at a configurable poll interval (default: 1 second). Data is streamed via WebSocket from the backend at a configurable poll interval (default: 1 second).
### Performance Limits
The Performance panel controls the board power limit and the memory clock offset. The power limit slider is bounded by the GPU's VBIOS minimum and maximum (shown at the slider ends); changes are applied on **Apply** and reset to the hardware default on **Reset**.
#### Experimental NVIDIA power control
On compatible drivers, an **Experimental NVIDIA power control** checkbox appears in the Performance panel. Enabling it switches power-limit application from the standard NVML call to an undocumented driver (RM) interface, which **permits caps below the VBIOS minimum, down to 30 W**. The native maximum still applies.
> **Warning.** This uses an undocumented driver interface for *all* power limits, including resets. It may cause instability or stop working after a driver update. Enable it at your own risk. The option is clearly labelled with a red warning in the UI, and profiles saved while it is enabled are marked accordingly.
Notes:
- The checkbox only appears when the driver exposes a compatible RM power layout (detected with a read-only probe — no writes).
- In this mode there is **no automatic fallback** to NVML: if the RM route fails, the error is reported rather than silently switching backends.
- The mode is per-GPU and persisted across server restarts. Reset restores the default through the same route, so it can also clear a previously-set below-minimum cap.
- The CLI reports availability via `nvcurve read --diag` ("Experimental RM power: available").
### Multi-GPU ### Multi-GPU
When multiple NVIDIA GPUs are detected, a GPU selector dropdown appears in the status bar. Switching GPUs resets pending edits, selection state, and monitoring for the new target. When multiple NVIDIA GPUs are detected, a GPU selector dropdown appears in the status bar. Switching GPUs resets pending edits, selection state, and monitoring for the new target.
@@ -145,6 +162,27 @@ Adding the first user **switches the server into authenticated mode immediately*
- The user store file should stay root-owned and `0600` (the CLI enforces this). - The user store file should stay root-owned and `0600` (the CLI enforces this).
- The web UI and API are still only as safe as the network path to the server — bind to a trusted interface (`--host`) and/or firewall the port. Authentication protects against casual access, not a determined network attacker. - The web UI and API are still only as safe as the network path to the server — bind to a trusted interface (`--host`) and/or firewall the port. Authentication protects against casual access, not a determined network attacker.
- The `nvcurve user` commands and the user store require root; day-to-day sign-in does not. - The `nvcurve user` commands and the user store require root; day-to-day sign-in does not.
- The login lockout is keyed by client IP. Behind a reverse proxy all clients share the proxy's IP — set `trusted_proxies` in `/etc/nvcurve/config.json` (e.g. `["127.0.0.1"]`) so the lockout uses the real client IP from `X-Forwarded-For`. The header is only honoured for peers you list there (it is spoofable otherwise).
- On shared systems consider setting `allow_api_shutdown: false` so users cannot stop the server via the API (manage it with systemd instead).
- The frequency safety cap (`max_delta_khz`) is enforced by the server from its config; API clients cannot raise it per request. The CLI's `--max-delta` (root-only, direct hardware path) can still override it for a single write.
## TLS (HTTPS)
The server speaks **plain HTTP by default**. When you expose it beyond localhost, enable TLS so credentials and session cookies are not sent in cleartext:
```bash
# Persistent (stored in /etc/nvcurve/config.json, used by the daemon too)
sudo nvcurve service configure --ssl-certfile /path/to/cert.pem --ssl-keyfile /path/to/key.pem
# One-off
nvcurve serve start --ssl-certfile /path/to/cert.pem --ssl-keyfile /path/to/key.pem
```
- Both files must be set for TLS to activate; the UI then lives at `https://<host>:8042` and the WebSocket upgrades to `wss://` automatically.
- To disable TLS again: `sudo nvcurve service configure --no-ssl` (removes the certificate/key from the config).
- The session cookie gets the `Secure` flag, so it is only sent over HTTPS.
- A self-signed certificate is fine for a home LAN (the browser shows a warning); for multi-user setups use a certificate your browser trusts.
- The CLI detects TLS from the config/runtime info and switches to `https://` automatically.
## CLI Reference ## CLI Reference
@@ -251,13 +289,14 @@ nvcurve service uninstall
sudo nvcurve service configure --auto-serve sudo nvcurve service configure --auto-serve
sudo nvcurve service configure --no-auto-serve sudo nvcurve service configure --no-auto-serve
sudo nvcurve service configure --host 0.0.0.0 --port 8042 sudo nvcurve service configure --host 0.0.0.0 --port 8042
sudo nvcurve service configure --ssl-certfile /path/to/cert.pem --ssl-keyfile /path/to/key.pem
``` ```
## Configuration Files ## Configuration Files
| File | Purpose | | File | Purpose |
| --- | --- | | --- | --- |
| `/etc/nvcurve/config.json` | Persistent config (host, port, auto-serve, default profiles) | | `/etc/nvcurve/config.json` | Persistent config (host, port, auto-serve, TLS, safety cap, default profiles) |
| `/etc/nvcurve/profiles/*.json` | Saved profiles | | `/etc/nvcurve/profiles/*.json` | Saved profiles |
| `/var/cache/nvcurve/snapshots/` | Auto-saved snapshots before writes | | `/var/cache/nvcurve/snapshots/` | Auto-saved snapshots before writes |
| `/run/nvcurve.json` | Runtime server info (host, port, PID) | | `/run/nvcurve.json` | Runtime server info (host, port, PID) |
+118
View File
@@ -0,0 +1,118 @@
# WireView Pro II Setup
NVCurve reads the [Thermal Grizzly WireView Pro II](https://www.thermal-grizzly.com/en/wireview-pro-ii-gpu/s-tg-wv-p2) (12 VHPWR connector monitor) directly over its USB CDC/ACM serial port — **no exporter, no GUI, no kernel module required**.
Before the first use, the host needs a one-time setup: a udev rule so the serial port is accessible under a stable name and left alone by other tools. After that, NVCurve auto-detects the device over USB — the WireView tab appears while the device is connected and disappears when it is unplugged.
## Why a one-time setup is needed
The WireView is an STM32 CDC/ACM virtual serial port (VID `0483`, PID `5740`). Out of the box, Linux:
- creates the port as `root:uucp 0660` with no stable device name, and
- lets **ModemManager** probe every new CDC-ACM port with AT/QCDM commands for up to ~30 s after each plug — it holds the port open and sends bytes the device does not expect.
The udev rule below fixes both: the node becomes `root:dialout 0660`, a stable `/dev/wireview-pro2` symlink is created, and ModemManager is told to ignore the device.
## 1. Install the udev rule
```sh
sudo tee /etc/udev/rules.d/99-wireview.rules > /dev/null <<'EOF'
# WireView Pro II udev rules
#
# Same access policy as wireview-hwmon's 99-wireview-hwmon.rules: the node is
# 0660 root:dialout, and the user logged in at the local seat gets an ACL via
# uaccess. systemd applies that tag on hotplug from 73-seat-late.rules, which
# runs before this 99- file and so never sees it, so each rule also runs the
# uaccess builtin itself. The tag is still needed: logind uses it to move the
# ACL to the new active session on a user switch.
ACTION=="remove", GOTO="wireview_end"
# Normal operation (STM32 CDC/ACM virtual serial port)
SUBSYSTEM=="tty", ATTRS{idVendor}=="0483", ATTRS{idProduct}=="5740", GROUP="dialout", MODE="0660", TAG+="uaccess", RUN{builtin}+="uaccess", SYMLINK+="wireview-pro2"
# Keep ModemManager away: it probes every new CDC-ACM port with AT and QCDM
# commands for about half a minute, holding the port and sending the device
# bytes it does not expect.
SUBSYSTEM=="usb", ENV{DEVTYPE}=="usb_device", ATTR{idVendor}=="0483", ATTR{idProduct}=="5740", ENV{ID_MM_DEVICE_IGNORE}="1"
SUBSYSTEM=="tty", ATTRS{idVendor}=="0483", ATTRS{idProduct}=="5740", ENV{ID_MM_PORT_IGNORE}="1"
# DFU bootloader mode (firmware update)
SUBSYSTEM=="usb", ENV{DEVTYPE}=="usb_device", ATTR{idVendor}=="0483", ATTR{idProduct}=="df11", GROUP="dialout", MODE="0660", TAG+="uaccess", RUN{builtin}+="uaccess"
LABEL="wireview_end"
EOF
```
> [!WARNING]
> The rule sets `GROUP="dialout"`. If the `dialout` group does not exist on your distro, udev silently ignores the rule (the node stays `root:uucp`). Create it first: `sudo groupadd dialout`.
Then reload the rules and apply them to an already-plugged device (a reload alone only affects future hotplugs):
```sh
sudo udevadm control --reload-rules
sudo udevadm trigger
```
## 2. Port access
- **NVCurve web server (systemd):** runs as root, so it can open the port as soon as the rule is in place — nothing else to do.
- **CLI as a regular user:** add your user to the `dialout` group, then log out/in:
```sh
sudo usermod -aG dialout $USER
```
A running process only picks up new groups after a re-login; `sg dialout -c 'nvcurve ...'` works for a one-off.
## 3. Verify
```sh
lsusb | grep 0483 # 0483:5740 present
ls -l /dev/wireview-pro2 # -> /dev/ttyACM0
```
With the device plugged in, start (or restart) the NVCurve server. The **WireView** tab appears in the web UI and `GET /api/wireview` returns the device info and live samples. The tab disappears when the device is unplugged and reappears on re-plug — no restart needed.
## Pin current imbalance warning
The 12 VHPWR connector is the most fragile part of a PC's power delivery: if one power pin has a bad contact (worn cable, faulty adapter, loose plug), the remaining pins are forced to carry the missing current. Under heavy load this can overheat and melt the connector — in rare cases even start a fire.
NVCurve watches for this: if the current between any two power pins diverges by more than **2 A**, a big warning appears in the app header (visible on every tab, not just WireView) naming the deviating pins (e.g. `Pin 5 (5.6 A) vs Pin 6 (3.5 A)`). It is debounced — the imbalance must persist for several seconds (~7 s) before the warning shows, so single-sample glitches and short load ramps (e.g. AI workloads) do not flash it. When you see it, reduce the load and replace the cable or adapter — do not keep using the connector.
## Alternative: wireview-hwmon kernel module
If the official `wireview-hwmon` kernel module and its `wireviewd` daemon are installed, the daemon owns the serial port and NVCurve automatically reads the sysfs node instead. No udev rule is needed in that case.
## USB device IDs
| Mode | VID | PID | Description |
| --- | --- | --- | --- |
| Normal | `0483` | `5740` | STM32 CDC/ACM virtual serial port |
| DFU bootloader | `0483` | `df11` | STM32 bootloader (firmware updates only) |
## Troubleshooting
### WireView tab does not appear
- `lsusb | grep 0483` — is the device visible at all (cable, USB port)?
- `ls -l /dev/wireview-pro2` — if the node is `root:uucp`, the `dialout` group is missing or the rule did not load: `sudo groupadd dialout && sudo udevadm control --reload-rules && sudo udevadm trigger`.
- Restart the NVCurve server — detection runs at startup, and the watchdog re-checks periodically afterwards.
### `Permission denied` on the port (CLI)
The user must be in the `dialout` group; a running process only picks up new groups after a restart/re-login.
### Device visible but never connects
The firmware occasionally stops answering the RTS welcome handshake (observed after USB state changes, e.g. udev re-triggers or boot) while still answering every data command. NVCurve falls back to the vendor-data reply and retries, but if it still does not connect, reset the USB device — physically unplug/replug, or without reaching for it:
```sh
P=$(readlink -f /sys/class/tty/ttyACM0)
DEV=$(dirname "$(dirname "$(dirname "$(dirname "$P")")")")
sudo bash -c "echo 0 > $DEV/authorized; sleep 1; echo 1 > $DEV/authorized"
```
---
*The udev rule is taken from the wireview-reporter project, which uses the same access policy as the official `wireview-hwmon` module.*
BIN
View File
Binary file not shown.

After

Width:  |  Height:  |  Size: 151 KiB

+524 -10
View File
@@ -18,6 +18,7 @@
"devDependencies": { "devDependencies": {
"@eslint/js": "^9.39.1", "@eslint/js": "^9.39.1",
"@tailwindcss/vite": "^4.2.1", "@tailwindcss/vite": "^4.2.1",
"@testing-library/react": "^16.3.3",
"@types/d3": "^7.4.3", "@types/d3": "^7.4.3",
"@types/node": "^24.10.1", "@types/node": "^24.10.1",
"@types/react": "^19.2.7", "@types/react": "^19.2.7",
@@ -27,10 +28,12 @@
"eslint-plugin-react-hooks": "^7.0.1", "eslint-plugin-react-hooks": "^7.0.1",
"eslint-plugin-react-refresh": "^0.4.24", "eslint-plugin-react-refresh": "^0.4.24",
"globals": "^16.5.0", "globals": "^16.5.0",
"happy-dom": "^20.14.5",
"tailwindcss": "^4.2.1", "tailwindcss": "^4.2.1",
"typescript": "~5.9.3", "typescript": "~5.9.3",
"typescript-eslint": "^8.48.0", "typescript-eslint": "^8.48.0",
"vite": "^7.3.1" "vite": "^7.3.1",
"vitest": "^5.0.3"
} }
}, },
"node_modules/@babel/code-frame": { "node_modules/@babel/code-frame": {
@@ -225,6 +228,16 @@
"node": ">=6.0.0" "node": ">=6.0.0"
} }
}, },
"node_modules/@babel/runtime": {
"version": "7.29.7",
"resolved": "https://registry.npmjs.org/@babel/runtime/-/runtime-7.29.7.tgz",
"integrity": "sha512-Nq8OhGWiZIZGV6hLHoyAKLLcJihP/xFeBMGJoUrxTX2psI8dCifzLhZISFb+VWS3wFMRDmCGw5R+dOySCqPLhw==",
"dev": true,
"license": "MIT",
"engines": {
"node": ">=6.9.0"
}
},
"node_modules/@babel/template": { "node_modules/@babel/template": {
"version": "7.28.6", "version": "7.28.6",
"resolved": "https://registry.npmjs.org/@babel/template/-/template-7.28.6.tgz", "resolved": "https://registry.npmjs.org/@babel/template/-/template-7.28.6.tgz",
@@ -971,9 +984,9 @@
} }
}, },
"node_modules/@jridgewell/sourcemap-codec": { "node_modules/@jridgewell/sourcemap-codec": {
"version": "1.5.5", "version": "1.6.0",
"resolved": "https://registry.npmjs.org/@jridgewell/sourcemap-codec/-/sourcemap-codec-1.5.5.tgz", "resolved": "https://registry.npmjs.org/@jridgewell/sourcemap-codec/-/sourcemap-codec-1.6.0.tgz",
"integrity": "sha512-cYQ9310grqxueWbl+WuIUIaiUaDcj7WOq5fVhEljNVgRfOUhY9fy2zTvfoqWsnebh8Sl70VScFbICvJnLKB0Og==", "integrity": "sha512-T7jf+5zgsZHwNJ4lvQ7/aezbyk0nNX+zJVWpmHA7VYsEx7a7qr5Rg5IbtJFqkgze5Y2sruq1RUY8Q837Od7iFw==",
"dev": true, "dev": true,
"license": "MIT" "license": "MIT"
}, },
@@ -1879,6 +1892,74 @@
"vite": "^5.2.0 || ^6 || ^7 || ^8" "vite": "^5.2.0 || ^6 || ^7 || ^8"
} }
}, },
"node_modules/@testing-library/dom": {
"version": "10.4.2",
"resolved": "https://registry.npmjs.org/@testing-library/dom/-/dom-10.4.2.tgz",
"integrity": "sha512-yzr2S9HyAIdhz2/6qHgbs665Q7PKVcDF05vsOlHPxG1mo36gKVesdYVeDLnXgfjJ03CrKRk08knc6+E/9m8v2Q==",
"dev": true,
"license": "MIT",
"peer": true,
"dependencies": {
"@babel/code-frame": "^7.10.4",
"@babel/runtime": "^7.12.5",
"@types/aria-query": "^5.0.1",
"aria-query": "5.3.0",
"dom-accessibility-api": "^0.5.9",
"lz-string": "^1.5.0",
"picocolors": "1.1.1",
"pretty-format": "^27.0.2"
},
"engines": {
"node": ">=18"
}
},
"node_modules/@testing-library/react": {
"version": "16.3.3",
"resolved": "https://registry.npmjs.org/@testing-library/react/-/react-16.3.3.tgz",
"integrity": "sha512-Uo193NgQbPMz6lrrhtRQQFcMC6Re/ELLFbbuVL30WDlZxlpZf9/lMHTAVxPRLw1q1iu9OJmR1c2BLiENRstdBg==",
"dev": true,
"license": "MIT",
"dependencies": {
"@babel/runtime": "^7.12.5"
},
"engines": {
"node": ">=18"
},
"peerDependencies": {
"@testing-library/dom": "^10.0.0",
"@types/react": "^18.0.0 || ^19.0.0",
"@types/react-dom": "^18.0.0 || ^19.0.0",
"react": "^18.0.0 || ^19.0.0",
"react-dom": "^18.0.0 || ^19.0.0"
},
"peerDependenciesMeta": {
"@types/react": {
"optional": true
},
"@types/react-dom": {
"optional": true
}
}
},
"node_modules/@types/aria-query": {
"version": "5.0.4",
"resolved": "https://registry.npmjs.org/@types/aria-query/-/aria-query-5.0.4.tgz",
"integrity": "sha512-rfT93uj5s0PRL7EzccGMs3brplhcrghnDoV26NqKhCAS1hVo+WdNsPvE/yb6ilfr5hi2MEk6d5EWJTKdxg8jVw==",
"dev": true,
"license": "MIT",
"peer": true
},
"node_modules/@types/chai": {
"version": "5.2.3",
"resolved": "https://registry.npmjs.org/@types/chai/-/chai-5.2.3.tgz",
"integrity": "sha512-Mw558oeA9fFbv65/y4mHtXDs9bPnFMZAL/jxdPFUpOHHIXX91mcgEHbS5Lahr+pwZFR8A7GQleRWeI6cGFC2UA==",
"dev": true,
"license": "MIT",
"dependencies": {
"@types/deep-eql": "*",
"assertion-error": "^2.0.1"
}
},
"node_modules/@types/d3": { "node_modules/@types/d3": {
"version": "7.4.3", "version": "7.4.3",
"resolved": "https://registry.npmjs.org/@types/d3/-/d3-7.4.3.tgz", "resolved": "https://registry.npmjs.org/@types/d3/-/d3-7.4.3.tgz",
@@ -2163,6 +2244,13 @@
"@types/d3-selection": "*" "@types/d3-selection": "*"
} }
}, },
"node_modules/@types/deep-eql": {
"version": "4.0.2",
"resolved": "https://registry.npmjs.org/@types/deep-eql/-/deep-eql-4.0.2.tgz",
"integrity": "sha512-c9h9dVVMigMPc4bwTvC5dxqtqJZwQPePsWjPlpSOnojbor6pGqdk541lfA7AqFQr5pB1BRdq0juY9db81BwyFw==",
"dev": true,
"license": "MIT"
},
"node_modules/@types/estree": { "node_modules/@types/estree": {
"version": "1.0.9", "version": "1.0.9",
"resolved": "https://registry.npmjs.org/@types/estree/-/estree-1.0.9.tgz", "resolved": "https://registry.npmjs.org/@types/estree/-/estree-1.0.9.tgz",
@@ -2214,6 +2302,23 @@
"@types/react": "^19.2.0" "@types/react": "^19.2.0"
} }
}, },
"node_modules/@types/whatwg-mimetype": {
"version": "3.0.2",
"resolved": "https://registry.npmjs.org/@types/whatwg-mimetype/-/whatwg-mimetype-3.0.2.tgz",
"integrity": "sha512-c2AKvDT8ToxLIOUlN51gTiHXflsfIFisS4pO7pDPoKouJCESkhZnEy623gwP9laCy5lnLDAw1vAzu2vM2YLOrA==",
"dev": true,
"license": "MIT"
},
"node_modules/@types/ws": {
"version": "8.18.2",
"resolved": "https://registry.npmjs.org/@types/ws/-/ws-8.18.2.tgz",
"integrity": "sha512-67MQl+fpWKVTT1NYdnmo3U4sc/xPo/zQBncVnI74qmQa0z/b+1g6iYqNmGCPbxO+zz2aklb08a0oHfegiVd0/w==",
"dev": true,
"license": "MIT",
"dependencies": {
"@types/node": "*"
}
},
"node_modules/@typescript-eslint/eslint-plugin": { "node_modules/@typescript-eslint/eslint-plugin": {
"version": "8.59.2", "version": "8.59.2",
"resolved": "https://registry.npmjs.org/@typescript-eslint/eslint-plugin/-/eslint-plugin-8.59.2.tgz", "resolved": "https://registry.npmjs.org/@typescript-eslint/eslint-plugin/-/eslint-plugin-8.59.2.tgz",
@@ -2526,6 +2631,54 @@
"vite": "^4 || ^5 || ^6 || ^7 || ^8" "vite": "^4 || ^5 || ^6 || ^7 || ^8"
} }
}, },
"node_modules/@vitest/mocker": {
"version": "5.0.3",
"resolved": "https://registry.npmjs.org/@vitest/mocker/-/mocker-5.0.3.tgz",
"integrity": "sha512-T8sWAIbkSyAjkwTcaEc3Iu0o9A27X1/kdXrizhZkGuSKScRQtRzclfAMpOTcGdXCsqxeWlpGy3XjqaW8CpLORg==",
"dev": true,
"license": "MIT",
"dependencies": {
"@jridgewell/trace-mapping": "0.3.31",
"@vitest/spy": "5.0.3",
"estree-walker": "^3.0.3",
"magic-string": "^1.2.3"
},
"funding": {
"url": "https://opencollective.com/vitest"
},
"peerDependencies": {
"msw": "^2.4.9",
"vite": "^6.0.0 || ^7.0.0 || ^8.0.0"
},
"peerDependenciesMeta": {
"msw": {
"optional": true
},
"vite": {
"optional": true
}
}
},
"node_modules/@vitest/mocker/node_modules/magic-string": {
"version": "1.4.2",
"resolved": "https://registry.npmjs.org/magic-string/-/magic-string-1.4.2.tgz",
"integrity": "sha512-vG+rjFRj1PqdIBozIxAGMjPlOhaVe+GXpbttY/iSK7rGcJRMlwNJO7dcUwmUqkymsFLJiNGI06t4D7Fr7yRC9g==",
"dev": true,
"license": "MIT",
"dependencies": {
"@jridgewell/sourcemap-codec": "^1.6.0"
}
},
"node_modules/@vitest/spy": {
"version": "5.0.3",
"resolved": "https://registry.npmjs.org/@vitest/spy/-/spy-5.0.3.tgz",
"integrity": "sha512-XhFysQTB8AZ+P4gMi+Lpo99vg2AZi0qKpaB9yXQl37+CaMEAPO3iH/wGVnSyL5MPERiLezpqTVtrR6UZH5GCXg==",
"dev": true,
"license": "MIT",
"funding": {
"url": "https://opencollective.com/vitest"
}
},
"node_modules/acorn": { "node_modules/acorn": {
"version": "8.16.0", "version": "8.16.0",
"resolved": "https://registry.npmjs.org/acorn/-/acorn-8.16.0.tgz", "resolved": "https://registry.npmjs.org/acorn/-/acorn-8.16.0.tgz",
@@ -2566,6 +2719,17 @@
"url": "https://github.com/sponsors/epoberezkin" "url": "https://github.com/sponsors/epoberezkin"
} }
}, },
"node_modules/ansi-regex": {
"version": "5.0.1",
"resolved": "https://registry.npmjs.org/ansi-regex/-/ansi-regex-5.0.1.tgz",
"integrity": "sha512-quJQXlTSUGL2LH9SUXo8VwsY4soanhgo6LNSm84E1LBcE8s3O0wpdiRzyR9z/ZZJMlMWv37qOOb9pdJlMUEKFQ==",
"dev": true,
"license": "MIT",
"peer": true,
"engines": {
"node": ">=8"
}
},
"node_modules/ansi-styles": { "node_modules/ansi-styles": {
"version": "4.3.0", "version": "4.3.0",
"resolved": "https://registry.npmjs.org/ansi-styles/-/ansi-styles-4.3.0.tgz", "resolved": "https://registry.npmjs.org/ansi-styles/-/ansi-styles-4.3.0.tgz",
@@ -2589,6 +2753,27 @@
"dev": true, "dev": true,
"license": "Python-2.0" "license": "Python-2.0"
}, },
"node_modules/aria-query": {
"version": "5.3.0",
"resolved": "https://registry.npmjs.org/aria-query/-/aria-query-5.3.0.tgz",
"integrity": "sha512-b0P0sZPKtyu8HkeRAfCq0IfURZK+SuwMjY1UXGBU27wpAiTwQAIlq56IbIO+ytk/JjS1fMR14ee5WBBfKi5J6A==",
"dev": true,
"license": "Apache-2.0",
"peer": true,
"dependencies": {
"dequal": "^2.0.3"
}
},
"node_modules/assertion-error": {
"version": "2.0.1",
"resolved": "https://registry.npmjs.org/assertion-error/-/assertion-error-2.0.1.tgz",
"integrity": "sha512-Izi8RQcffqCeNVgFigKli1ssklIbpHnCYc6AknXGYoB6grJqyeby7jv12JUQgmTAnIDnbck1uxksT4dzN3PWBA==",
"dev": true,
"license": "MIT",
"engines": {
"node": ">=12"
}
},
"node_modules/balanced-match": { "node_modules/balanced-match": {
"version": "1.0.2", "version": "1.0.2",
"resolved": "https://registry.npmjs.org/balanced-match/-/balanced-match-1.0.2.tgz", "resolved": "https://registry.npmjs.org/balanced-match/-/balanced-match-1.0.2.tgz",
@@ -2654,6 +2839,19 @@
"node": "^6 || ^7 || ^8 || ^9 || ^10 || ^11 || ^12 || >=13.7" "node": "^6 || ^7 || ^8 || ^9 || ^10 || ^11 || ^12 || >=13.7"
} }
}, },
"node_modules/buffer-image-size": {
"version": "0.6.4",
"resolved": "https://registry.npmjs.org/buffer-image-size/-/buffer-image-size-0.6.4.tgz",
"integrity": "sha512-nEh+kZOPY1w+gcCMobZ6ETUp9WfibndnosbpwB1iJk/8Gt5ZF2bhS6+B6bPYz424KtwsR6Rflc3tCz1/ghX2dQ==",
"dev": true,
"license": "MIT",
"dependencies": {
"@types/node": "*"
},
"engines": {
"node": ">=4.0"
}
},
"node_modules/callsites": { "node_modules/callsites": {
"version": "3.1.0", "version": "3.1.0",
"resolved": "https://registry.npmjs.org/callsites/-/callsites-3.1.0.tgz", "resolved": "https://registry.npmjs.org/callsites/-/callsites-3.1.0.tgz",
@@ -2685,6 +2883,16 @@
], ],
"license": "CC-BY-4.0" "license": "CC-BY-4.0"
}, },
"node_modules/chai": {
"version": "6.3.0",
"resolved": "https://registry.npmjs.org/chai/-/chai-6.3.0.tgz",
"integrity": "sha512-XWAtwJ6OHO+tj0EKCs0Y2UamnyOxseZWltU4x2U2wh8g4AigdjwvtUjvLP2tqkA/avxHEtzxNaqGq/YGNwckKg==",
"dev": true,
"license": "MIT",
"engines": {
"node": ">=18"
}
},
"node_modules/chalk": { "node_modules/chalk": {
"version": "4.1.2", "version": "4.1.2",
"resolved": "https://registry.npmjs.org/chalk/-/chalk-4.1.2.tgz", "resolved": "https://registry.npmjs.org/chalk/-/chalk-4.1.2.tgz",
@@ -3202,6 +3410,17 @@
"robust-predicates": "^3.0.2" "robust-predicates": "^3.0.2"
} }
}, },
"node_modules/dequal": {
"version": "2.0.3",
"resolved": "https://registry.npmjs.org/dequal/-/dequal-2.0.3.tgz",
"integrity": "sha512-0je+qPKHEMohvfRTCEo3CrPG6cAzAYgmzKyxRiYSSDkS6eGJdyVJm7WaYA5ECaAD9wLB2T4EEeymA5aFVcYXCA==",
"dev": true,
"license": "MIT",
"peer": true,
"engines": {
"node": ">=6"
}
},
"node_modules/detect-libc": { "node_modules/detect-libc": {
"version": "2.1.2", "version": "2.1.2",
"resolved": "https://registry.npmjs.org/detect-libc/-/detect-libc-2.1.2.tgz", "resolved": "https://registry.npmjs.org/detect-libc/-/detect-libc-2.1.2.tgz",
@@ -3212,6 +3431,14 @@
"node": ">=8" "node": ">=8"
} }
}, },
"node_modules/dom-accessibility-api": {
"version": "0.5.16",
"resolved": "https://registry.npmjs.org/dom-accessibility-api/-/dom-accessibility-api-0.5.16.tgz",
"integrity": "sha512-X7BJ2yElsnOJ30pZF4uIIDfBEVgF4XEBxL9Bxhy6dnrm5hkzqmsWHGTiHqRiITNhMyFLyAiWndIJP7Z1NTteDg==",
"dev": true,
"license": "MIT",
"peer": true
},
"node_modules/electron-to-chromium": { "node_modules/electron-to-chromium": {
"version": "1.5.353", "version": "1.5.353",
"resolved": "https://registry.npmjs.org/electron-to-chromium/-/electron-to-chromium-1.5.353.tgz", "resolved": "https://registry.npmjs.org/electron-to-chromium/-/electron-to-chromium-1.5.353.tgz",
@@ -3233,6 +3460,26 @@
"node": ">=10.13.0" "node": ">=10.13.0"
} }
}, },
"node_modules/entities": {
"version": "7.0.1",
"resolved": "https://registry.npmjs.org/entities/-/entities-7.0.1.tgz",
"integrity": "sha512-TWrgLOFUQTH994YUyl1yT4uyavY5nNB5muff+RtWaqNVCAK408b5ZnnbNAUEWLTCpum9w6arT70i1XdQ4UeOPA==",
"dev": true,
"license": "BSD-2-Clause",
"engines": {
"node": ">=0.12"
},
"funding": {
"url": "https://github.com/fb55/entities?sponsor=1"
}
},
"node_modules/es-module-lexer": {
"version": "2.3.2",
"resolved": "https://registry.npmjs.org/es-module-lexer/-/es-module-lexer-2.3.2.tgz",
"integrity": "sha512-poHGpORABojJJucnV9KbOavETW8lBVnphkW77ER5/BQ5Fz7oXSoCNek7IH3vR5nRjdsEz926ibFYX8KtLQmdyw==",
"dev": true,
"license": "MIT"
},
"node_modules/esbuild": { "node_modules/esbuild": {
"version": "0.27.7", "version": "0.27.7",
"resolved": "https://registry.npmjs.org/esbuild/-/esbuild-0.27.7.tgz", "resolved": "https://registry.npmjs.org/esbuild/-/esbuild-0.27.7.tgz",
@@ -3472,6 +3719,16 @@
"node": ">=4.0" "node": ">=4.0"
} }
}, },
"node_modules/estree-walker": {
"version": "3.0.3",
"resolved": "https://registry.npmjs.org/estree-walker/-/estree-walker-3.0.3.tgz",
"integrity": "sha512-7RUKfXgSMMkzt6ZuXmqapOurLGPPfgj6l9uRZ7lRGolvk0y2yocc35LdcxKC5PQZdn2DMqioAQ2NoWcrTKmm6g==",
"dev": true,
"license": "MIT",
"dependencies": {
"@types/estree": "^1.0.0"
}
},
"node_modules/esutils": { "node_modules/esutils": {
"version": "2.0.3", "version": "2.0.3",
"resolved": "https://registry.npmjs.org/esutils/-/esutils-2.0.3.tgz", "resolved": "https://registry.npmjs.org/esutils/-/esutils-2.0.3.tgz",
@@ -3482,6 +3739,16 @@
"node": ">=0.10.0" "node": ">=0.10.0"
} }
}, },
"node_modules/expect-type": {
"version": "1.4.0",
"resolved": "https://registry.npmjs.org/expect-type/-/expect-type-1.4.0.tgz",
"integrity": "sha512-KfYbmpRm0VbLjEvVa9yGwCi9GI34xvi7A/HXYWQO65CSD2u3MczUJSuwXKFIxlGsgBQizV9q5J9NHj4VG0n+pA==",
"dev": true,
"license": "Apache-2.0",
"engines": {
"node": ">=12.0.0"
}
},
"node_modules/fast-deep-equal": { "node_modules/fast-deep-equal": {
"version": "3.1.3", "version": "3.1.3",
"resolved": "https://registry.npmjs.org/fast-deep-equal/-/fast-deep-equal-3.1.3.tgz", "resolved": "https://registry.npmjs.org/fast-deep-equal/-/fast-deep-equal-3.1.3.tgz",
@@ -3630,6 +3897,25 @@
"dev": true, "dev": true,
"license": "ISC" "license": "ISC"
}, },
"node_modules/happy-dom": {
"version": "20.14.5",
"resolved": "https://registry.npmjs.org/happy-dom/-/happy-dom-20.14.5.tgz",
"integrity": "sha512-x/RzkpWO40bTjIoT30iQtt64FLLmH/iRcUCN2X//bLx7H3ifkdfPXyqsro/OYtqzIAhiLMMA7mmiOR9C3NOKjQ==",
"dev": true,
"license": "MIT",
"dependencies": {
"@types/node": ">=20.0.0",
"@types/whatwg-mimetype": "^3.0.2",
"@types/ws": "^8.18.1",
"buffer-image-size": "^0.6.4",
"entities": "^7.0.1",
"whatwg-mimetype": "^3.0.0",
"ws": "^8.21.0"
},
"engines": {
"node": ">=20.0.0"
}
},
"node_modules/has-flag": { "node_modules/has-flag": {
"version": "4.0.0", "version": "4.0.0",
"resolved": "https://registry.npmjs.org/has-flag/-/has-flag-4.0.0.tgz", "resolved": "https://registry.npmjs.org/has-flag/-/has-flag-4.0.0.tgz",
@@ -4149,6 +4435,17 @@
"react": "^16.5.1 || ^17.0.0 || ^18.0.0 || ^19.0.0" "react": "^16.5.1 || ^17.0.0 || ^18.0.0 || ^19.0.0"
} }
}, },
"node_modules/lz-string": {
"version": "1.5.0",
"resolved": "https://registry.npmjs.org/lz-string/-/lz-string-1.5.0.tgz",
"integrity": "sha512-h5bgJWpxJNswbU7qCrV0tIKQCaS3blPDrqKWx+QxzuzL1zGUzij9XCWLrSLsJPu5t+eWA/ycetzYAO5IOMcWAQ==",
"dev": true,
"license": "MIT",
"peer": true,
"bin": {
"lz-string": "bin/bin.js"
}
},
"node_modules/magic-string": { "node_modules/magic-string": {
"version": "0.30.21", "version": "0.30.21",
"resolved": "https://registry.npmjs.org/magic-string/-/magic-string-0.30.21.tgz", "resolved": "https://registry.npmjs.org/magic-string/-/magic-string-0.30.21.tgz",
@@ -4212,6 +4509,20 @@
"dev": true, "dev": true,
"license": "MIT" "license": "MIT"
}, },
"node_modules/obug": {
"version": "2.2.1",
"resolved": "https://registry.npmjs.org/obug/-/obug-2.2.1.tgz",
"integrity": "sha512-XrsrhT5sybtKI6wakr2SPOlGZWWYbUXZ7a0jT8/QOeAPau+1X/bSegNe5YR75oJmEZQbKningirmGOEJCIk61Q==",
"dev": true,
"funding": [
"https://github.com/sponsors/sxzz",
"https://opencollective.com/debug"
],
"license": "MIT",
"engines": {
"node": ">=12.20.0"
}
},
"node_modules/optionator": { "node_modules/optionator": {
"version": "0.9.4", "version": "0.9.4",
"resolved": "https://registry.npmjs.org/optionator/-/optionator-0.9.4.tgz", "resolved": "https://registry.npmjs.org/optionator/-/optionator-0.9.4.tgz",
@@ -4303,9 +4614,9 @@
"license": "ISC" "license": "ISC"
}, },
"node_modules/picomatch": { "node_modules/picomatch": {
"version": "4.0.4", "version": "4.0.7",
"resolved": "https://registry.npmjs.org/picomatch/-/picomatch-4.0.4.tgz", "resolved": "https://registry.npmjs.org/picomatch/-/picomatch-4.0.7.tgz",
"integrity": "sha512-QP88BAKvMam/3NxH6vj2o21R6MjxZUAd6nlwAS/pnGvN9IVLocLHxGYIzFhg6fUQ+5th6P4dv4eW9jX3DSIj7A==", "integrity": "sha512-qcJu88Q2IWqJsDD529JKMdwGm/dvInW4HvQnRwiH9JtihJvzGOscDtHE3x1pBKeUOTysQ8kVmLnJ2kJu7yhcGA==",
"dev": true, "dev": true,
"license": "MIT", "license": "MIT",
"engines": { "engines": {
@@ -4354,6 +4665,36 @@
"node": ">= 0.8.0" "node": ">= 0.8.0"
} }
}, },
"node_modules/pretty-format": {
"version": "27.5.1",
"resolved": "https://registry.npmjs.org/pretty-format/-/pretty-format-27.5.1.tgz",
"integrity": "sha512-Qb1gy5OrP5+zDf2Bvnzdl3jsTf1qXVMazbvCoKhtKqVs4/YK4ozX4gKQJJVyNe+cajNPn0KoC0MC3FUmaHWEmQ==",
"dev": true,
"license": "MIT",
"peer": true,
"dependencies": {
"ansi-regex": "^5.0.1",
"ansi-styles": "^5.0.0",
"react-is": "^17.0.1"
},
"engines": {
"node": "^10.13.0 || ^12.13.0 || ^14.15.0 || >=15.0.0"
}
},
"node_modules/pretty-format/node_modules/ansi-styles": {
"version": "5.2.0",
"resolved": "https://registry.npmjs.org/ansi-styles/-/ansi-styles-5.2.0.tgz",
"integrity": "sha512-Cxwpt2SfTzTtXcfOlzGEee8O+c+MmUgGrNiBcXnuWxuFJHe6a5Hz7qwhwe5OgaSYI0IJvkLqWX1ASG+cJOkEiA==",
"dev": true,
"license": "MIT",
"peer": true,
"engines": {
"node": ">=10"
},
"funding": {
"url": "https://github.com/chalk/ansi-styles?sponsor=1"
}
},
"node_modules/punycode": { "node_modules/punycode": {
"version": "2.3.1", "version": "2.3.1",
"resolved": "https://registry.npmjs.org/punycode/-/punycode-2.3.1.tgz", "resolved": "https://registry.npmjs.org/punycode/-/punycode-2.3.1.tgz",
@@ -4385,6 +4726,14 @@
"react": "^19.2.6" "react": "^19.2.6"
} }
}, },
"node_modules/react-is": {
"version": "17.0.2",
"resolved": "https://registry.npmjs.org/react-is/-/react-is-17.0.2.tgz",
"integrity": "sha512-w2GsyukL62IJnlaff/nRegPQR94C/XXamvMWmSHRJ4y7Ts/4ocGRmTHvOs8PSE6pB3dWOrD/nueuU5sduBsQ4w==",
"dev": true,
"license": "MIT",
"peer": true
},
"node_modules/resolve-from": { "node_modules/resolve-from": {
"version": "4.0.0", "version": "4.0.0",
"resolved": "https://registry.npmjs.org/resolve-from/-/resolve-from-4.0.0.tgz", "resolved": "https://registry.npmjs.org/resolve-from/-/resolve-from-4.0.0.tgz",
@@ -4524,6 +4873,13 @@
"node": ">=0.10.0" "node": ">=0.10.0"
} }
}, },
"node_modules/std-env": {
"version": "4.3.0",
"resolved": "https://registry.npmjs.org/std-env/-/std-env-4.3.0.tgz",
"integrity": "sha512-OtU/EgQ1kIm5KwqQpBC6ZEMXrZRui11w8zgfTWp8cdO9B8OaPsbA8bTHO2P+HNo1VlUTGMVBwPhydu6poeXiag==",
"dev": true,
"license": "MIT"
},
"node_modules/strip-json-comments": { "node_modules/strip-json-comments": {
"version": "3.1.1", "version": "3.1.1",
"resolved": "https://registry.npmjs.org/strip-json-comments/-/strip-json-comments-3.1.1.tgz", "resolved": "https://registry.npmjs.org/strip-json-comments/-/strip-json-comments-3.1.1.tgz",
@@ -4571,10 +4927,30 @@
"url": "https://opencollective.com/webpack" "url": "https://opencollective.com/webpack"
} }
}, },
"node_modules/tinybench": {
"version": "6.2.0",
"resolved": "https://registry.npmjs.org/tinybench/-/tinybench-6.2.0.tgz",
"integrity": "sha512-78U2TlB2CnVenajOFzf3BKSm0J6oz5L0NV7g32LCPccvYc0lbWvys4d3uUUCS2B1N8PAf2+aekR8i1KbC3HO7Q==",
"dev": true,
"license": "MIT",
"engines": {
"node": ">=20.0.0"
}
},
"node_modules/tinyexec": {
"version": "1.3.1",
"resolved": "https://registry.npmjs.org/tinyexec/-/tinyexec-1.3.1.tgz",
"integrity": "sha512-GCvB3aoys96IuDFBMcTB46JOR6mdMtAToqwiW8JlWhsoh1mhHi/xn9ss/Dg7N555GiJyEt2qzoG/NHCwM6h1EA==",
"dev": true,
"license": "MIT",
"engines": {
"node": ">=18"
}
},
"node_modules/tinyglobby": { "node_modules/tinyglobby": {
"version": "0.2.16", "version": "0.2.17",
"resolved": "https://registry.npmjs.org/tinyglobby/-/tinyglobby-0.2.16.tgz", "resolved": "https://registry.npmjs.org/tinyglobby/-/tinyglobby-0.2.17.tgz",
"integrity": "sha512-pn99VhoACYR8nFHhxqix+uvsbXineAasWm5ojXoN8xEwK5Kd3/TrhNn1wByuD52UxWRLy8pu+kRMniEi6Eq9Zg==", "integrity": "sha512-wXR/dYpcqKmfWpEdZjiKJOwCNFndD0DMnrW/cYjVGttEkBfVgcLFHoNrlj47mjOVic9yyNu65alsgF4NQyTa2g==",
"dev": true, "dev": true,
"license": "MIT", "license": "MIT",
"dependencies": { "dependencies": {
@@ -4775,6 +5151,109 @@
} }
} }
}, },
"node_modules/vitest": {
"version": "5.0.3",
"resolved": "https://registry.npmjs.org/vitest/-/vitest-5.0.3.tgz",
"integrity": "sha512-xMw97S3rjdtj5dkVat7jCsqWBpvchs3RlpQctUqwJD0KkERk40vz2fJ77lDwW/Vzh/pk18eItYAzkodhSes3jQ==",
"dev": true,
"license": "MIT",
"dependencies": {
"@types/chai": "^5.2.2",
"@vitest/mocker": "5.0.3",
"chai": "^6.2.2",
"es-module-lexer": "^2.3.2",
"expect-type": "^1.4.0",
"magic-string": "^1.2.3",
"obug": "^2.1.4",
"picomatch": "^4.0.7",
"std-env": "^4.2.0",
"tinybench": "^6.1.4",
"tinyexec": "^1.3.0",
"tinyglobby": "^0.2.17",
"why-is-node-running": "3.2.1"
},
"bin": {
"vitest": "vitest.mjs"
},
"engines": {
"node": "^22.12.0 || ^24.0.0 || >=26.0.0"
},
"funding": {
"url": "https://opencollective.com/vitest"
},
"peerDependencies": {
"@edge-runtime/vm": "*",
"@opentelemetry/api": "^1.9.0",
"@types/node": "^22.0.0 || >=24.0.0",
"@vitest/browser-playwright": "5.0.3",
"@vitest/browser-preview": "5.0.3",
"@vitest/browser-webdriverio": "^5.0.0-beta.5 || >=5.0.0",
"@vitest/coverage-istanbul": "5.0.3",
"@vitest/coverage-v8": "5.0.3",
"@vitest/ui": "5.0.3",
"happy-dom": "*",
"jsdom": "*",
"vite": "^6.4.0 || ^7.0.0 || ^8.0.0"
},
"peerDependenciesMeta": {
"@edge-runtime/vm": {
"optional": true
},
"@opentelemetry/api": {
"optional": true
},
"@types/node": {
"optional": true
},
"@vitest/browser-playwright": {
"optional": true
},
"@vitest/browser-preview": {
"optional": true
},
"@vitest/browser-webdriverio": {
"optional": true
},
"@vitest/coverage-istanbul": {
"optional": true
},
"@vitest/coverage-v8": {
"optional": true
},
"@vitest/ui": {
"optional": true
},
"happy-dom": {
"optional": true
},
"jsdom": {
"optional": true
},
"vite": {
"optional": false
}
}
},
"node_modules/vitest/node_modules/magic-string": {
"version": "1.4.2",
"resolved": "https://registry.npmjs.org/magic-string/-/magic-string-1.4.2.tgz",
"integrity": "sha512-vG+rjFRj1PqdIBozIxAGMjPlOhaVe+GXpbttY/iSK7rGcJRMlwNJO7dcUwmUqkymsFLJiNGI06t4D7Fr7yRC9g==",
"dev": true,
"license": "MIT",
"dependencies": {
"@jridgewell/sourcemap-codec": "^1.6.0"
}
},
"node_modules/whatwg-mimetype": {
"version": "3.0.0",
"resolved": "https://registry.npmjs.org/whatwg-mimetype/-/whatwg-mimetype-3.0.0.tgz",
"integrity": "sha512-nt+N2dzIutVRxARx1nghPKGv1xHikU7HKdfafKkLNLindmPU/ch3U31NOCGGA/dmPcmb1VlofO0vnKAcsm0o/Q==",
"dev": true,
"license": "MIT",
"engines": {
"node": ">=12"
}
},
"node_modules/which": { "node_modules/which": {
"version": "2.0.2", "version": "2.0.2",
"resolved": "https://registry.npmjs.org/which/-/which-2.0.2.tgz", "resolved": "https://registry.npmjs.org/which/-/which-2.0.2.tgz",
@@ -4791,6 +5270,19 @@
"node": ">= 8" "node": ">= 8"
} }
}, },
"node_modules/why-is-node-running": {
"version": "3.2.1",
"resolved": "https://registry.npmjs.org/why-is-node-running/-/why-is-node-running-3.2.1.tgz",
"integrity": "sha512-Tb2FUhB4vUsGQlfSquQLYkApkuPAFQXGFzxWKHHumVz2dK+X1RUm/HnID4+TfIGYJ1kTcwOaCk/buYCEJr6YjQ==",
"dev": true,
"license": "MIT",
"bin": {
"why-is-node-running": "cli.js"
},
"engines": {
"node": ">=20.11"
}
},
"node_modules/word-wrap": { "node_modules/word-wrap": {
"version": "1.2.5", "version": "1.2.5",
"resolved": "https://registry.npmjs.org/word-wrap/-/word-wrap-1.2.5.tgz", "resolved": "https://registry.npmjs.org/word-wrap/-/word-wrap-1.2.5.tgz",
@@ -4801,6 +5293,28 @@
"node": ">=0.10.0" "node": ">=0.10.0"
} }
}, },
"node_modules/ws": {
"version": "8.22.0",
"resolved": "https://registry.npmjs.org/ws/-/ws-8.22.0.tgz",
"integrity": "sha512-Ydggc987+RO0AnWtZ/7Wq9FtNvcrL1b/RO0ud9mWjUPgDrsAAwQSF51sm2hm1XofbU/4jkpGEsLFsZZxU+1DOg==",
"dev": true,
"license": "MIT",
"engines": {
"node": ">=10.0.0"
},
"peerDependencies": {
"bufferutil": "^4.0.1",
"utf-8-validate": ">=5.0.2"
},
"peerDependenciesMeta": {
"bufferutil": {
"optional": true
},
"utf-8-validate": {
"optional": true
}
}
},
"node_modules/yallist": { "node_modules/yallist": {
"version": "3.1.1", "version": "3.1.1",
"resolved": "https://registry.npmjs.org/yallist/-/yallist-3.1.1.tgz", "resolved": "https://registry.npmjs.org/yallist/-/yallist-3.1.1.tgz",
+5 -1
View File
@@ -7,6 +7,7 @@
"dev": "vite", "dev": "vite",
"build": "tsc -b && vite build", "build": "tsc -b && vite build",
"lint": "eslint .", "lint": "eslint .",
"test": "vitest run",
"preview": "vite preview" "preview": "vite preview"
}, },
"dependencies": { "dependencies": {
@@ -20,6 +21,7 @@
"devDependencies": { "devDependencies": {
"@eslint/js": "^9.39.1", "@eslint/js": "^9.39.1",
"@tailwindcss/vite": "^4.2.1", "@tailwindcss/vite": "^4.2.1",
"@testing-library/react": "^16.3.3",
"@types/d3": "^7.4.3", "@types/d3": "^7.4.3",
"@types/node": "^24.10.1", "@types/node": "^24.10.1",
"@types/react": "^19.2.7", "@types/react": "^19.2.7",
@@ -29,9 +31,11 @@
"eslint-plugin-react-hooks": "^7.0.1", "eslint-plugin-react-hooks": "^7.0.1",
"eslint-plugin-react-refresh": "^0.4.24", "eslint-plugin-react-refresh": "^0.4.24",
"globals": "^16.5.0", "globals": "^16.5.0",
"happy-dom": "^20.14.5",
"tailwindcss": "^4.2.1", "tailwindcss": "^4.2.1",
"typescript": "~5.9.3", "typescript": "~5.9.3",
"typescript-eslint": "^8.48.0", "typescript-eslint": "^8.48.0",
"vite": "^7.3.1" "vite": "^7.3.1",
"vitest": "^5.0.3"
} }
} }
+72 -28
View File
@@ -1,24 +1,28 @@
import { useGpu } from "./hooks/useGpu"; import { useGpu } from "./hooks/useGpu.js";
import { useCurve } from "./hooks/useCurve"; import { useCurve } from "./hooks/useCurve.js";
import { useMonitor } from "./hooks/useMonitor"; import { useMonitor } from "./hooks/useMonitor.js";
import { useDashboard } from "./hooks/useDashboard"; import { useDashboard } from "./hooks/useDashboard.js";
import { StatusBar } from "./components/Monitor/StatusBar"; import { StatusBar } from "./components/Monitor/StatusBar.js";
import { LiveMonitor } from "./components/Monitor/LiveMonitor"; import { LiveMonitor } from "./components/Monitor/LiveMonitor.js";
import { Dashboard } from "./components/Dashboard/Dashboard"; import { Dashboard } from "./components/Dashboard/Dashboard.js";
import { CurveEditor } from "./components/CurveEditor/CurveEditor"; import { CurveEditor } from "./components/CurveEditor/CurveEditor.js";
import { PointTable } from "./components/PointTable/PointTable"; import { PointTable } from "./components/PointTable/PointTable.js";
import { PerformancePanel } from "./components/Limits/PerformancePanel"; import { PerformancePanel } from "./components/Limits/PerformancePanel.js";
import { PerformanceMonitor } from "./components/Monitor/PerformanceMonitor"; import { PerformanceMonitor } from "./components/Monitor/PerformanceMonitor.js";
import { FanMonitor } from "./components/Monitor/FanMonitor"; import { ProcessList } from "./components/Monitor/ProcessList.js";
import { FanCurveEditor } from "./components/Fans/FanCurveEditor"; import { FanMonitor } from "./components/Monitor/FanMonitor.js";
import { ProfilePanel } from "./components/Profiles/ProfilePanel"; import { FanCurveEditor } from "./components/Fans/FanCurveEditor.js";
import { api, onUnauthorized } from "./api/client"; import { WireViewPanel } from "./components/WireView/WireViewPanel.js";
import { LoginScreen } from "./components/Auth/LoginScreen"; import { ProfilePanel } from "./components/Profiles/ProfilePanel.js";
import { useCurveStore } from "./store/curveStore"; import { api, onUnauthorized } from "./api/client.js";
import { LoginScreen } from "./components/Auth/LoginScreen.js";
import { useCurveStore } from "./store/curveStore.js";
import { useWireview } from "./hooks/useWireview.js";
import { useWireviewImbalance } from "./hooks/useWireviewImbalance.js";
import { Toaster } from "sonner"; import { Toaster } from "sonner";
import { Loader, ChevronDown } from "lucide-react"; import { Loader, ChevronDown } from "lucide-react";
import { useState, useRef, useEffect } from "react"; import { useState, useRef, useEffect } from "react";
import type { FanState } from "./types"; import type { FanState } from "./types.js";
type AuthState = "checking" | "login" | "ok"; type AuthState = "checking" | "login" | "ok";
@@ -94,9 +98,11 @@ function MainApp({
const { dashboard, loading: dashboardLoading } = useDashboard(); const { dashboard, loading: dashboardLoading } = useDashboard();
const { setCurve, activeProfile, setActiveProfile, selectedGpuIndex } = const { setCurve, activeProfile, setActiveProfile, selectedGpuIndex } =
useCurveStore(); useCurveStore();
const wireview = useWireview();
const wireImbalance = useWireviewImbalance(wireview.sample);
const [activeTab, setActiveTab] = useState< const [activeTab, setActiveTab] = useState<
"dashboard" | "curve" | "performance" | "fans" "dashboard" | "curve" | "performance" | "fans" | "wireview"
>("dashboard"); >("dashboard");
const [fanState, setFanState] = useState<FanState | null>(null); const [fanState, setFanState] = useState<FanState | null>(null);
const [activeDomain, setActiveDomain] = useState<"gpu" | "memory">("gpu"); const [activeDomain, setActiveDomain] = useState<"gpu" | "memory">("gpu");
@@ -167,6 +173,26 @@ function MainApp({
</div> </div>
)} )}
{wireImbalance && (
<div className="bg-red-500/15 border-b-2 border-red-500 text-red-300 px-4 py-2.5 flex items-center justify-center gap-3 flex-wrap">
<span className="text-xl leading-none">⚠</span>
<span className="text-base font-bold">
WireView: current imbalance {wireImbalance.diff.toFixed(1)} A —
Pin {wireImbalance.hiPin} ({wireImbalance.maxI.toFixed(1)} A) vs
Pin {wireImbalance.loPin} ({wireImbalance.minI.toFixed(1)} A)
</span>
<span className="text-sm text-red-200/80">
Bad pin contact — reduce load, replace cable/adapter
</span>
<button
onClick={() => setActiveTab("wireview")}
className="text-sm font-semibold text-red-200 hover:text-white underline underline-offset-2"
>
View
</button>
</div>
)}
<main className="flex-1 p-4 flex flex-col gap-4 w-full max-w-[1600px] mx-auto"> <main className="flex-1 p-4 flex flex-col gap-4 w-full max-w-[1600px] mx-auto">
{/* Tab Header and Profile Selector */} {/* Tab Header and Profile Selector */}
<div className="flex justify-between items-end border-b border-zinc-800 pb-2"> <div className="flex justify-between items-end border-b border-zinc-800 pb-2">
@@ -195,6 +221,15 @@ function MainApp({
> >
Fans Fans
</button> </button>
{wireview.available && (
<button
onClick={() => setActiveTab("wireview")}
className={`text-lg font-medium pb-2 -mb-[9px] border-b-2 transition-colors flex items-center gap-2 ${activeTab === "wireview" ? "border-pink-500 text-zinc-100" : "border-transparent text-zinc-500 hover:text-zinc-300"}`}
>
WireView
<span className="w-1.5 h-1.5 rounded-full bg-emerald-400" />
</button>
)}
</div> </div>
<div className="relative" ref={profileRef}> <div className="relative" ref={profileRef}>
<button <button
@@ -278,17 +313,26 @@ function MainApp({
)} )}
</> </>
) : activeTab === "performance" ? ( ) : activeTab === "performance" ? (
<div className="flex gap-4 items-start w-full"> <div className="flex flex-col gap-4 w-full">
<div className="flex-1 min-w-0"> <div className="flex gap-4 items-start w-full">
<PerformancePanel /> <div className="flex-1 min-w-0">
</div> <PerformancePanel />
<div className="w-80 shrink-0 flex flex-col"> </div>
<PerformanceMonitor <div className="w-80 shrink-0 flex flex-col">
monitor={monitor} <PerformanceMonitor
history={monitorHistory} monitor={monitor}
/> history={monitorHistory}
/>
</div>
</div> </div>
<ProcessList />
</div> </div>
) : activeTab === "wireview" ? (
<WireViewPanel
info={wireview.info}
sample={wireview.sample}
history={wireview.history}
/>
) : ( ) : (
<div className="flex gap-4 items-start w-full"> <div className="flex gap-4 items-start w-full">
<div className="flex-1 min-w-0"> <div className="flex-1 min-w-0">
+25 -5
View File
@@ -8,7 +8,8 @@ import type {
FanState, FanState,
FanPoint, FanPoint,
DashboardInfo, DashboardInfo,
} from "../types"; ProcessListResponse,
} from "../types.js";
export class ApiError extends Error { export class ApiError extends Error {
status: number; status: number;
@@ -119,6 +120,17 @@ export const api = {
voltage: (gpuIndex: number) => voltage: (gpuIndex: number) =>
get<{ voltage_uv: number; voltage_mv: number }>("/voltage", gpuIndex), get<{ voltage_uv: number; voltage_mv: number }>("/voltage", gpuIndex),
monitor: (gpuIndex: number) => get<MonitoringSample>("/monitor", gpuIndex), monitor: (gpuIndex: number) => get<MonitoringSample>("/monitor", gpuIndex),
processes: (gpuIndex: number) =>
get<ProcessListResponse>("/processes", gpuIndex),
killProcess: (
pid: number,
signal: "TERM" | "KILL" = "TERM",
parent = false,
) =>
post<{ ok: boolean; pid: number; signal: string; parent: boolean }>(
"/processes/kill",
{ pid, signal, parent },
),
snapshots: (gpuIndex: number) => get<SnapshotInfo[]>("/snapshots", gpuIndex), snapshots: (gpuIndex: number) => get<SnapshotInfo[]>("/snapshots", gpuIndex),
/** Write per-point frequency deltas. deltas: { pointIndex: deltaKhz } */ /** Write per-point frequency deltas. deltas: { pointIndex: deltaKhz } */
@@ -168,9 +180,17 @@ export const api = {
/** Fan control */ /** Fan control */
fans: (gpuIndex: number) => get<FanState>("/fans", gpuIndex), fans: (gpuIndex: number) => get<FanState>("/fans", gpuIndex),
updateFans: (curve: FanPoint[], gpuIndex: number) => updateFans: (curve: FanPoint[], gpuIndex: number, fans?: number[] | null) =>
post("/fans", { curve }, gpuIndex), post("/fans", { curve, fans: fans ?? null }, gpuIndex),
resetFans: (gpuIndex: number) => post("/fans/reset", undefined, gpuIndex), resetFans: (gpuIndex: number) => post("/fans/reset", undefined, gpuIndex),
setFanSpeed: (fanPct: number, gpuIndex: number) => setFanSpeed: (fanPct: number, gpuIndex: number, fan?: number | null) =>
post("/fans/speed", { fan_pct: fanPct }, gpuIndex), post("/fans/speed", { fan_pct: fanPct, fan: fan ?? null }, gpuIndex),
/** Report Bug (Gitea) */
reportBugStatus: () => get<{ configured: boolean }>("/report-bug/status"),
reportBug: (title: string, body: string) =>
post<{ ok: boolean; issue_number: number | null; issue_url: string | null }>(
"/report-bug",
{ title, body },
),
}; };
+2 -2
View File
@@ -1,6 +1,6 @@
import { useState } from "react"; import { useState } from "react";
import { Loader, Lock, User } from "lucide-react"; import { Loader, Lock, User } from "lucide-react";
import { api, ApiError } from "../../api/client"; import { api, ApiError } from "../../api/client.js";
interface Props { interface Props {
onSuccess: (username: string) => void; onSuccess: (username: string) => void;
@@ -12,7 +12,7 @@ export function LoginScreen({ onSuccess }: Props) {
const [error, setError] = useState<string | null>(null); const [error, setError] = useState<string | null>(null);
const [busy, setBusy] = useState(false); const [busy, setBusy] = useState(false);
async function submit(e: React.FormEvent) { async function submit(e: React.SubmitEvent) {
e.preventDefault(); e.preventDefault();
if (busy) return; if (busy) return;
setBusy(true); setBusy(true);
File diff suppressed because it is too large. Load diff
@@ -1,7 +1,7 @@
import { useState, useMemo, useEffect } from 'react'; import { useState, useMemo } from "react";
import { ZoomIn, RotateCcw, Minus } from 'lucide-react'; import { ZoomIn, RotateCcw, Minus } from "lucide-react";
import { useCurveStore } from '../../store/curveStore'; import { useCurveStore } from "../../store/curveStore.js";
import type { VFPoint } from '../../types'; import type { VFPoint } from "../../types.js";
interface Props { interface Props {
/** All curve points — used by global offset slider */ /** All curve points — used by global offset slider */
@@ -18,23 +18,44 @@ interface Props {
onZoomChange: (factor: number) => void; onZoomChange: (factor: number) => void;
} }
export function CurveToolbar({ activePts, onResetZoom, isZoomed, readOnly, zoomFactor, onZoomChange }: Props) { export function CurveToolbar({
const { pendingDeltas, selectedPoints, anchorPoint, curve, stageRangeEdit, flattenToAnchor } = useCurveStore(); activePts,
onResetZoom,
isZoomed,
readOnly,
zoomFactor,
onZoomChange,
}: Props) {
const {
pendingDeltas,
selectedPoints,
anchorPoint,
curve,
stageRangeEdit,
flattenToAnchor,
} = useCurveStore();
const [offsetMhz, setOffsetMhz] = useState(0); const [offsetMhz, setOffsetMhz] = useState(0);
const uniformDeltaMhz = useMemo(() => { const uniformDeltaMhz = useMemo(() => {
if (activePts.length === 0) return 0; if (activePts.length === 0) return 0;
const firstD = pendingDeltas.get(activePts[0].index) ?? activePts[0].delta_khz; const firstD =
const uniform = activePts.every((p) => (pendingDeltas.get(p.index) ?? p.delta_khz) === firstD); pendingDeltas.get(activePts[0].index) ?? activePts[0].delta_khz;
const uniform = activePts.every(
(p) => (pendingDeltas.get(p.index) ?? p.delta_khz) === firstD,
);
return uniform ? firstD / 1000 : null; return uniform ? firstD / 1000 : null;
}, [activePts, pendingDeltas]); }, [activePts, pendingDeltas]);
useEffect(() => { // Sync the slider to the uniform delta when it changes (adjust state during
if (uniformDeltaMhz !== null) { // render instead of an effect; undefined sentinel so the first render syncs).
setOffsetMhz(uniformDeltaMhz); const [lastUniformDelta, setLastUniformDelta] = useState<
} number | null | undefined
}, [uniformDeltaMhz]); >();
if (uniformDeltaMhz !== null && lastUniformDelta !== uniformDeltaMhz) {
setLastUniformDelta(uniformDeltaMhz);
setOffsetMhz(uniformDeltaMhz);
}
function handleOffsetChange(mhz: number) { function handleOffsetChange(mhz: number) {
setOffsetMhz(mhz); setOffsetMhz(mhz);
@@ -44,7 +65,10 @@ export function CurveToolbar({ activePts, onResetZoom, isZoomed, readOnly, zoomF
return ( return (
<div className="flex flex-wrap items-center gap-2 px-1 pb-2"> <div className="flex flex-wrap items-center gap-2 px-1 pb-2">
{/* Zoom control */} {/* Zoom control */}
<div className="flex items-center gap-1.5 px-2 py-1 rounded bg-zinc-800/60 border border-zinc-700/40" title="Zoom x-axis (Alt+scroll also works)"> <div
className="flex items-center gap-1.5 px-2 py-1 rounded bg-zinc-800/60 border border-zinc-700/40"
title="Zoom x-axis (Alt+scroll also works)"
>
<ZoomIn size={11} className="text-zinc-500 shrink-0" /> <ZoomIn size={11} className="text-zinc-500 shrink-0" />
<input <input
type="range" type="range"
@@ -55,7 +79,9 @@ export function CurveToolbar({ activePts, onResetZoom, isZoomed, readOnly, zoomF
onChange={(e) => onZoomChange(Number(e.target.value))} onChange={(e) => onZoomChange(Number(e.target.value))}
className="w-20 h-1 cursor-pointer accent-cyan-400" className="w-20 h-1 cursor-pointer accent-cyan-400"
/> />
<span className={`text-xs font-mono w-8 tabular-nums ${isZoomed ? 'text-cyan-400' : 'text-zinc-600'}`}> <span
className={`text-xs font-mono w-8 tabular-nums ${isZoomed ? "text-cyan-400" : "text-zinc-600"}`}
>
{zoomFactor.toFixed(1)}× {zoomFactor.toFixed(1)}×
</span> </span>
{isZoomed && ( {isZoomed && (
@@ -75,7 +101,9 @@ export function CurveToolbar({ activePts, onResetZoom, isZoomed, readOnly, zoomF
{/* Global offset slider — GPU only */} {/* Global offset slider — GPU only */}
{!readOnly && uniformDeltaMhz !== null && ( {!readOnly && uniformDeltaMhz !== null && (
<div className="flex items-center gap-1.5 min-w-[260px]"> <div className="flex items-center gap-1.5 min-w-[260px]">
<span className="text-zinc-500 text-xs whitespace-nowrap">Global Offset</span> <span className="text-zinc-500 text-xs whitespace-nowrap">
Global Offset
</span>
<input <input
type="range" type="range"
min={-1000} min={-1000}
@@ -84,46 +112,63 @@ export function CurveToolbar({ activePts, onResetZoom, isZoomed, readOnly, zoomF
value={offsetMhz} value={offsetMhz}
onChange={(e) => handleOffsetChange(Number(e.target.value))} onChange={(e) => handleOffsetChange(Number(e.target.value))}
className="w-32 accent-cyan-400" className="w-32 accent-cyan-400"
title={`${offsetMhz > 0 ? '+' : ''}${offsetMhz} MHz`} title={`${offsetMhz > 0 ? "+" : ""}${offsetMhz} MHz`}
/> />
<span <span
className={[ className={[
'text-xs font-mono w-16', "text-xs font-mono w-16",
offsetMhz > 0 ? 'text-cyan-400' : offsetMhz < 0 ? 'text-orange-400' : 'text-zinc-500', offsetMhz > 0
].join(' ')} ? "text-cyan-400"
: offsetMhz < 0
? "text-orange-400"
: "text-zinc-500",
].join(" ")}
> >
{offsetMhz > 0 ? '+' : ''}{offsetMhz} MHz {offsetMhz > 0 ? "+" : ""}
{offsetMhz} MHz
</span> </span>
</div> </div>
)} )}
{/* Flatten — visible when 2+ points are selected */} {/* Flatten — visible when 2+ points are selected */}
{!readOnly && selectedPoints.size >= 2 && (() => { {!readOnly &&
const anchor = anchorPoint !== null && selectedPoints.has(anchorPoint) selectedPoints.size >= 2 &&
? anchorPoint (() => {
: Math.min(...selectedPoints); const anchor =
const anchorDelta = anchorPoint !== null && selectedPoints.has(anchorPoint)
pendingDeltas.get(anchor) ?? ? anchorPoint
curve?.points.find(p => p.index === anchor)?.delta_khz ?? : Math.min(...selectedPoints);
0; const anchorDelta =
const label = `·${anchor} ${anchorDelta >= 0 ? '+' : ''}${anchorDelta / 1000} MHz`; pendingDeltas.get(anchor) ??
return ( curve?.points.find((p) => p.index === anchor)?.delta_khz ??
<button 0;
onClick={flattenToAnchor} const label = `·${anchor} ${anchorDelta >= 0 ? "+" : ""}${anchorDelta / 1000} MHz`;
className="flex items-center gap-1.5 px-2 py-1 rounded text-xs font-medium text-amber-400 hover:text-amber-300 hover:bg-zinc-800 border border-zinc-700/40 transition" return (
title={`Flatten all selected points to anchor point ${anchor} (${anchorDelta >= 0 ? '+' : ''}${anchorDelta / 1000} MHz)`} <button
> onClick={flattenToAnchor}
<Minus size={11} /> className="flex items-center gap-1.5 px-2 py-1 rounded text-xs font-medium text-amber-400 hover:text-amber-300 hover:bg-zinc-800 border border-zinc-700/40 transition"
Flatten to {label} title={`Flatten all selected points to anchor point ${anchor} (${anchorDelta >= 0 ? "+" : ""}${anchorDelta / 1000} MHz)`}
</button> >
); <Minus size={11} />
})()} Flatten to {label}
</button>
);
})()}
{/* Legend — right-aligned */} {/* Legend — right-aligned */}
<div className="flex items-center gap-3 text-xs text-zinc-500 ml-auto"> <div className="flex items-center gap-3 text-xs text-zinc-500 ml-auto">
<span className="flex items-center gap-1"><span className="inline-block w-3 h-0.5 bg-emerald-400 rounded" /> effective</span> <span className="flex items-center gap-1">
<span className="flex items-center gap-1"><span className="inline-block w-3 h-px border-t-2 border-dashed border-cyan-400" /> pending</span> <span className="inline-block w-3 h-0.5 bg-emerald-400 rounded" />{" "}
<span className="flex items-center gap-1"><span className="inline-block w-2 h-2 rounded-full bg-yellow-400" /> current</span> effective
</span>
<span className="flex items-center gap-1">
<span className="inline-block w-3 h-px border-t-2 border-dashed border-cyan-400" />{" "}
pending
</span>
<span className="flex items-center gap-1">
<span className="inline-block w-2 h-2 rounded-full bg-yellow-400" />{" "}
current
</span>
</div> </div>
</div> </div>
); );
@@ -1,12 +1,12 @@
import { fmt } from '../../utils/units'; import { fmt } from "../../utils/units.js";
import type { VFPoint } from '../../types'; import type { VFPoint } from "../../types.js";
interface Props { interface Props {
point: VFPoint; point: VFPoint;
/** Pending delta in kHz (from store), if any */ /** Pending delta in kHz (from store), if any */
pendingDeltaKhz?: number; pendingDeltaKhz?: number;
/** True if this point is being held up by monotonicity enforcement */ /** True if this point is being held up by monotonicity enforcement */
isClamped?: boolean; isClamped?: boolean;
} }
/** /**
@@ -14,57 +14,70 @@ interface Props {
* SVG so it's never clipped by the SVG viewport. * SVG so it's never clipped by the SVG viewport.
*/ */
export function CurveTooltip({ point, pendingDeltaKhz, isClamped }: Props) { export function CurveTooltip({ point, pendingDeltaKhz, isClamped }: Props) {
const hasPending = pendingDeltaKhz !== undefined; const hasPending = pendingDeltaKhz !== undefined;
const pendingMhz = hasPending ? pendingDeltaKhz! / 1000 : 0; const pendingMhz = hasPending ? pendingDeltaKhz! / 1000 : 0;
const deltaChange = hasPending ? pendingDeltaKhz! - point.delta_khz : 0; const deltaChange = hasPending ? pendingDeltaKhz! - point.delta_khz : 0;
const pendingEffMhz = hasPending ? point.freq_mhz + deltaChange / 1000 : null; const pendingEffMhz = hasPending
? point.freq_mhz + deltaChange / 1000
: null;
return ( return (
<div <div className="absolute right-4 bottom-4 pointer-events-none z-50 w-[172px] bg-zinc-800 border border-zinc-700 rounded-md p-2 text-xs shadow-xl">
style={{ <div className="text-zinc-400 mb-1">Point {point.index}</div>
position: 'absolute',
right: 16,
bottom: 16,
pointerEvents: 'none',
zIndex: 50,
width: 172,
}}
className="bg-zinc-800 border border-zinc-700 rounded-md p-2 text-xs shadow-xl"
>
<div className="text-zinc-400 mb-1">Point {point.index}</div>
<div className="text-zinc-200">
<span className="text-zinc-400">Volt: </span>{fmt.mv(point.volt_mv, 1)}
</div>
<div className="text-zinc-200">
<span className="text-zinc-400">Offset: </span>
<span className={point.delta_khz > 0 ? 'text-emerald-400' : point.delta_khz < 0 ? 'text-red-400' : 'text-zinc-400'}>
{point.delta_khz > 0 ? '+' : ''}{fmt.mhz(point.delta_mhz, 1)}
</span>
</div>
<div className="text-emerald-300 font-semibold">
<span className="text-zinc-400">Eff.: </span>{fmt.mhz(point.freq_mhz, 0)}
{isClamped && <span className="text-amber-500 ml-1">⇡</span>}
</div>
{isClamped && (
<div className="text-amber-500/80 text-[10px] mt-0.5">
Clamped by lower-voltage point
</div>
)}
{hasPending && (
<>
<div className="border-t border-zinc-700 mt-1.5 pt-1.5">
<div className="text-zinc-200"> <div className="text-zinc-200">
<span className="text-zinc-400">Pending: </span> <span className="text-zinc-400">Volt: </span>
<span className={pendingMhz > 0 ? 'text-cyan-400' : pendingMhz < 0 ? 'text-orange-400' : 'text-zinc-400'}> {fmt.mv(point.volt_mv, 1)}
{pendingMhz > 0 ? '+' : ''}{pendingMhz.toFixed(1)} MHz
</span>
</div> </div>
<div className="text-cyan-300 font-semibold"> <div className="text-zinc-200">
<span className="text-zinc-400">→ Eff.: </span>{fmt.mhz(pendingEffMhz, 0)} <span className="text-zinc-400">Offset: </span>
<span
className={
point.delta_khz > 0
? "text-emerald-400"
: point.delta_khz < 0
? "text-red-400"
: "text-zinc-400"
}
>
{point.delta_khz > 0 ? "+" : ""}
{fmt.mhz(point.delta_mhz, 1)}
</span>
</div> </div>
</div> <div className="text-emerald-300 font-semibold">
</> <span className="text-zinc-400">Eff.: </span>
)} {fmt.mhz(point.freq_mhz, 0)}
</div> {isClamped && <span className="text-amber-500 ml-1">⇡</span>}
); </div>
{isClamped && (
<div className="text-amber-500/80 text-[10px] mt-0.5">
Clamped by lower-voltage point
</div>
)}
{hasPending && (
<>
<div className="border-t border-zinc-700 mt-1.5 pt-1.5">
<div className="text-zinc-200">
<span className="text-zinc-400">Pending: </span>
<span
className={
pendingMhz > 0
? "text-cyan-400"
: pendingMhz < 0
? "text-orange-400"
: "text-zinc-400"
}
>
{pendingMhz > 0 ? "+" : ""}
{pendingMhz.toFixed(1)} MHz
</span>
</div>
<div className="text-cyan-300 font-semibold">
<span className="text-zinc-400">→ Eff.: </span>
{fmt.mhz(pendingEffMhz, 0)}
</div>
</div>
</>
)}
</div>
);
} }
@@ -1,7 +1,7 @@
import { Loader } from "lucide-react"; import { Loader } from "lucide-react";
import { GaugeCard } from "../Monitor/GaugeCard"; import { GaugeCard } from "../Monitor/GaugeCard.js";
import { fmt } from "../../utils/units"; import { fmt } from "../../utils/units.js";
import type { MonitoringSample, DashboardInfo } from "../../types"; import type { MonitoringSample, DashboardInfo } from "../../types.js";
interface Props { interface Props {
monitor: MonitoringSample | null; monitor: MonitoringSample | null;
+117 -17
View File
@@ -1,10 +1,10 @@
import { useState, useEffect, useRef, useCallback } from "react"; import { useState, useEffect, useRef, useCallback } from "react";
import { Check, X, RotateCcw, Plus } from "lucide-react"; import { Check, X, RotateCcw, Plus } from "lucide-react";
import { api } from "../../api/client"; import { api } from "../../api/client.js";
import { useCurveStore } from "../../store/curveStore"; import { useCurveStore } from "../../store/curveStore.js";
import type { FanPoint, FanState } from "../../types"; import type { FanInfo, FanPoint, FanState } from "../../types.js";
import { toast } from "sonner"; import { toast } from "sonner";
import { ConfirmDialog } from "../common/ConfirmDialog"; import { ConfirmDialog } from "../common/ConfirmDialog.js";
function defaultCurve(): FanPoint[] { function defaultCurve(): FanPoint[] {
return [ return [
@@ -44,6 +44,29 @@ function yToFan(y: number) {
return Math.round(FAN_MAX - ((y - PAD.top) / PLOT_H) * (FAN_MAX - FAN_MIN)); return Math.round(FAN_MAX - ((y - PAD.top) / PLOT_H) * (FAN_MAX - FAN_MIN));
} }
function sameTargets(
a: number[] | null | undefined,
b: number[] | null | undefined,
) {
if (a === null || a === undefined) return b === null || b === undefined;
if (b === null || b === undefined) return false;
if (a.length !== b.length) return false;
return a.every((v, i) => v === b[i]);
}
function fanLabel(targets: number[] | null): string {
return targets === null
? "all fans"
: targets.map((i) => `Fan ${i + 1}`).join(", ");
}
const chipCls = (selected: boolean) =>
`px-2 py-0.5 rounded-full border text-xs font-medium transition-colors ${
selected
? "bg-cyan-500/15 border-cyan-500/40 text-cyan-300"
: "bg-zinc-800 border-zinc-700 text-zinc-500 hover:text-zinc-300"
}`;
export function FanCurveEditor({ onChanged }: { onChanged?: () => void }) { export function FanCurveEditor({ onChanged }: { onChanged?: () => void }) {
const { selectedGpuIndex } = useCurveStore(); const { selectedGpuIndex } = useCurveStore();
const [fanState, setFanState] = useState<FanState | null>(null); const [fanState, setFanState] = useState<FanState | null>(null);
@@ -54,6 +77,9 @@ export function FanCurveEditor({ onChanged }: { onChanged?: () => void }) {
const [confirmApply, setConfirmApply] = useState(false); const [confirmApply, setConfirmApply] = useState(false);
const [confirmReset, setConfirmReset] = useState(false); const [confirmReset, setConfirmReset] = useState(false);
const [dragIdx, setDragIdx] = useState<number | null>(null); const [dragIdx, setDragIdx] = useState<number | null>(null);
// Fan selection: undefined = no pending change (follow server state),
// null = all fans, list = specific fan indices.
const [fanSel, setFanSel] = useState<number[] | null | undefined>(undefined);
const svgRef = useRef<SVGSVGElement>(null); const svgRef = useRef<SVGSVGElement>(null);
async function fetchFans() { async function fetchFans() {
@@ -61,6 +87,7 @@ export function FanCurveEditor({ onChanged }: { onChanged?: () => void }) {
setLoading(true); setLoading(true);
const data = await api.fans(selectedGpuIndex); const data = await api.fans(selectedGpuIndex);
setFanState(data); setFanState(data);
setFanSel(undefined);
if (data.curve && data.curve.length > 0) { if (data.curve && data.curve.length > 0) {
setPending(null); setPending(null);
} }
@@ -71,7 +98,10 @@ export function FanCurveEditor({ onChanged }: { onChanged?: () => void }) {
} }
} }
// Data fetch on GPU change — setState calls happen after the await, not
// synchronously in the effect body (rule false-positive on async fetch).
useEffect(() => { useEffect(() => {
// eslint-disable-next-line react-hooks/set-state-in-effect
fetchFans(); fetchFans();
}, [selectedGpuIndex]); }, [selectedGpuIndex]);
@@ -80,20 +110,47 @@ export function FanCurveEditor({ onChanged }: { onChanged?: () => void }) {
const curveActive = fanState?.curve_active ?? false; const curveActive = fanState?.curve_active ?? false;
const isDefaults = !pending && !fanState?.curve; const isDefaults = !pending && !fanState?.curve;
const fanTargets =
fanSel === undefined ? (fanState?.fan_targets ?? null) : fanSel;
const fansChanged =
fanSel !== undefined && !sameTargets(fanSel, fanState?.fan_targets ?? null);
const fanList: FanInfo[] =
fanState?.fans && fanState.fans.length > 0
? fanState.fans
: [{ index: 0, fan_pct: null }];
function toggleFan(idx: number) {
const numFans = fanState?.num_fans ?? 1;
let list: number[];
if (fanTargets === null) {
// Start from all fans, then drop the toggled one
list = Array.from({ length: numFans }, (_, i) => i).filter(
(i) => i !== idx,
);
} else {
list = fanTargets.includes(idx)
? fanTargets.filter((i) => i !== idx)
: [...fanTargets, idx].sort((a, b) => a - b);
}
if (list.length === 0) return; // keep at least one fan selected
setFanSel(list.length === numFans ? null : list);
}
async function handleApply() { async function handleApply() {
const curveToApply = pending ?? fanState?.curve ?? defaultCurve(); const curveToApply = pending ?? fanState?.curve ?? defaultCurve();
if (curveToApply.length < 2) return; if (curveToApply.length < 2) return;
setBusy(true); setBusy(true);
setError(null); setError(null);
try { try {
await api.updateFans(curveToApply, selectedGpuIndex); await api.updateFans(curveToApply, selectedGpuIndex, fanTargets);
setPending(null); setPending(null);
setFanSel(undefined);
setConfirmApply(false); setConfirmApply(false);
await fetchFans(); await fetchFans();
onChanged?.(); onChanged?.();
toast.success("Fan curve applied"); toast.success("Fan curve applied");
} catch (e: any) { } catch (e: unknown) {
setError(e.message ?? String(e)); setError(e instanceof Error ? e.message : String(e));
setConfirmApply(false); setConfirmApply(false);
} finally { } finally {
setBusy(false); setBusy(false);
@@ -106,12 +163,13 @@ export function FanCurveEditor({ onChanged }: { onChanged?: () => void }) {
try { try {
await api.resetFans(selectedGpuIndex); await api.resetFans(selectedGpuIndex);
setPending(null); setPending(null);
setFanSel(undefined);
setConfirmReset(false); setConfirmReset(false);
await fetchFans(); await fetchFans();
onChanged?.(); onChanged?.();
toast.success("Fan control reset to automatic"); toast.success("Fan control reset to automatic");
} catch (e: any) { } catch (e: unknown) {
setError(e.message ?? String(e)); setError(e instanceof Error ? e.message : String(e));
setConfirmReset(false); setConfirmReset(false);
} finally { } finally {
setBusy(false); setBusy(false);
@@ -248,9 +306,10 @@ export function FanCurveEditor({ onChanged }: { onChanged?: () => void }) {
<button <button
onClick={() => { onClick={() => {
setPending(null); setPending(null);
setFanSel(undefined);
setError(null); setError(null);
}} }}
disabled={!hasPending || busy} disabled={(!hasPending && !fansChanged) || busy}
className="flex items-center gap-1.5 px-2 py-1 rounded bg-zinc-800 hover:bg-zinc-700 text-zinc-300 text-xs transition-colors disabled:opacity-40 disabled:cursor-not-allowed" className="flex items-center gap-1.5 px-2 py-1 rounded bg-zinc-800 hover:bg-zinc-700 text-zinc-300 text-xs transition-colors disabled:opacity-40 disabled:cursor-not-allowed"
> >
<X size={12} /> <X size={12} />
@@ -267,6 +326,48 @@ export function FanCurveEditor({ onChanged }: { onChanged?: () => void }) {
</div> </div>
</div> </div>
{/* Fan selection */}
<div className="px-4 py-2 border-b border-zinc-800/60 flex items-center gap-2 flex-wrap">
<span className="text-xs text-zinc-500 uppercase tracking-wider font-semibold">
Fans
</span>
{fanState?.num_fans === 0 ? (
<span className="text-xs text-zinc-600">
No fans available on this GPU
</span>
) : (
<>
<button
onClick={() => setFanSel(null)}
className={chipCls(fanTargets === null)}
title="Control all fans"
>
All
</button>
{fanList.map((f) => (
<button
key={f.index}
onClick={() => toggleFan(f.index)}
className={chipCls(
fanTargets === null || fanTargets.includes(f.index),
)}
title={`Control Fan ${f.index + 1} with the curve`}
>
Fan {f.index + 1}
<span className="ml-1.5 font-mono text-[10px] opacity-80">
{f.fan_pct !== null ? `${Math.round(f.fan_pct)}%` : "—"}
</span>
</button>
))}
</>
)}
{fansChanged && (
<span className="inline-flex items-center gap-1 px-2 py-0.5 rounded-full bg-cyan-500/15 border border-cyan-500/30 text-cyan-400 text-xs">
Fans: {fanLabel(fanTargets)}
</span>
)}
</div>
{/* Error banner */} {/* Error banner */}
{error && ( {error && (
<div className="px-3 py-1.5 bg-red-900/40 border-b border-red-700 text-red-300 text-xs flex items-center justify-between"> <div className="px-3 py-1.5 bg-red-900/40 border-b border-red-700 text-red-300 text-xs flex items-center justify-between">
@@ -294,8 +395,7 @@ export function FanCurveEditor({ onChanged }: { onChanged?: () => void }) {
ref={svgRef} ref={svgRef}
width="100%" width="100%"
viewBox={`0 0 ${CHART_W} ${CHART_H}`} viewBox={`0 0 ${CHART_W} ${CHART_H}`}
className="max-w-full cursor-crosshair select-none" className="max-w-full cursor-crosshair select-none touch-none"
style={{ touchAction: "none" }}
onClick={handleCanvasClick} onClick={handleCanvasClick}
onPointerMove={handlePointerMove} onPointerMove={handlePointerMove}
> >
@@ -410,8 +510,7 @@ export function FanCurveEditor({ onChanged }: { onChanged?: () => void }) {
fill={hasPending ? "#22d3ee" : "#fb923c"} fill={hasPending ? "#22d3ee" : "#fb923c"}
stroke="#09090b" stroke="#09090b"
strokeWidth="2" strokeWidth="2"
className="cursor-grab active:cursor-grabbing" className="cursor-grab active:cursor-grabbing touch-none"
style={{ touchAction: "none" }}
onPointerDown={(e) => { onPointerDown={(e) => {
e.stopPropagation(); e.stopPropagation();
handlePointerDown(i); handlePointerDown(i);
@@ -556,13 +655,14 @@ export function FanCurveEditor({ onChanged }: { onChanged?: () => void }) {
{/* Info */} {/* Info */}
<div className="px-4 pb-3 text-[10px] text-zinc-600"> <div className="px-4 pb-3 text-[10px] text-zinc-600">
{curveActive {curveActive
? "Fan curve is active. Server adjusts fan speed based on GPU temperature." ? `Fan curve is active — controlling ${fanLabel(fanTargets)}. Server adjusts fan speed based on GPU temperature.`
: isDefaults : isDefaults
? "These are default values. Click Apply to enable curve control, or edit points first." ? "These are default values. Click Apply to enable curve control, or edit points first."
: "Apply a curve to enable automatic fan control based on temperature."} : "Apply a curve to enable automatic fan control based on temperature."}
<span className="block mt-1 text-zinc-700"> <span className="block mt-1 text-zinc-700">
Drag points to adjust · click the chart to add a point · click the ✕ Drag points to adjust · click the chart to add a point · click the ✕
(chart or table) to remove one. (chart or table) to remove one · pick which fans the curve controls
above.
</span> </span>
</div> </div>
</div> </div>
@@ -570,7 +670,7 @@ export function FanCurveEditor({ onChanged }: { onChanged?: () => void }) {
{confirmApply && ( {confirmApply && (
<ConfirmDialog <ConfirmDialog
message="Apply fan curve?" message="Apply fan curve?"
detail="Fan control will switch to curve-based mode. The server will adjust fan speed based on GPU temperature." detail={`Fan control will switch to curve-based mode for ${fanLabel(fanTargets)}. The server will adjust fan speed based on GPU temperature.`}
confirmLabel="Apply" confirmLabel="Apply"
onConfirm={handleApply} onConfirm={handleApply}
onCancel={() => setConfirmApply(false)} onCancel={() => setConfirmApply(false)}
@@ -1,12 +1,12 @@
import { useState, useEffect } from 'react'; import { useState, useEffect } from "react";
import { Check, X, RotateCcw } from 'lucide-react'; import { Check, X, RotateCcw } from "lucide-react";
import { api } from '../../api/client'; import { api } from "../../api/client.js";
import { useCurveStore } from '../../store/curveStore'; import { useCurveStore } from "../../store/curveStore.js";
import type { LimitsState } from '../../types'; import type { LimitsState } from "../../types.js";
import { toast } from 'sonner'; import { toast } from "sonner";
import { ConfirmDialog } from '../common/ConfirmDialog'; import { ConfirmDialog } from "../common/ConfirmDialog.js";
type Pending = Pick<LimitsState, 'power_limit_w' | 'mem_offset_mhz'>; type Pending = Pick<LimitsState, "power_limit_w" | "mem_offset_mhz">;
export function PerformancePanel() { export function PerformancePanel() {
const { selectedGpuIndex } = useCurveStore(); const { selectedGpuIndex } = useCurveStore();
@@ -23,13 +23,18 @@ export function PerformancePanel() {
setLoading(true); setLoading(true);
setLimits(await api.limits(selectedGpuIndex)); setLimits(await api.limits(selectedGpuIndex));
} catch { } catch {
toast.error('Failed to load performance limits'); toast.error("Failed to load performance limits");
} finally { } finally {
setLoading(false); setLoading(false);
} }
} }
useEffect(() => { fetchLimits(); }, [selectedGpuIndex]); // Data fetch on GPU change — setState calls happen after the await, not
// synchronously in the effect body (rule false-positive on async fetch).
useEffect(() => {
// eslint-disable-next-line react-hooks/set-state-in-effect
fetchLimits();
}, [selectedGpuIndex]);
async function handleApply() { async function handleApply() {
setBusy(true); setBusy(true);
@@ -40,9 +45,9 @@ export function PerformancePanel() {
setPending({}); setPending({});
setConfirmApply(false); setConfirmApply(false);
await fetchLimits(); await fetchLimits();
toast.success('Performance limits applied'); toast.success("Performance limits applied");
} catch (e: any) { } catch (e: unknown) {
setError(e.message ?? String(e)); setError(e instanceof Error ? e.message : String(e));
setConfirmApply(false); setConfirmApply(false);
} finally { } finally {
setBusy(false); setBusy(false);
@@ -58,15 +63,30 @@ export function PerformancePanel() {
setPending({}); setPending({});
setConfirmReset(false); setConfirmReset(false);
await fetchLimits(); await fetchLimits();
toast.success('Performance limits reset to defaults'); toast.success("Performance limits reset to defaults");
} catch (e: any) { } catch (e: unknown) {
setError(e.message ?? String(e)); setError(e instanceof Error ? e.message : String(e));
setConfirmReset(false); setConfirmReset(false);
} finally { } finally {
setBusy(false); setBusy(false);
} }
} }
async function handleModeChange(enabled: boolean) {
setBusy(true);
try {
await api.updateLimits(
{ power_cap_mode: enabled ? "ioctl" : "nvml" },
selectedGpuIndex,
);
await fetchLimits();
} catch (e: unknown) {
toast.error(e instanceof Error ? e.message : String(e));
} finally {
setBusy(false);
}
}
if (loading && !limits) { if (loading && !limits) {
return ( return (
<div className="bg-zinc-900 rounded-lg overflow-hidden flex flex-col animate-pulse"> <div className="bg-zinc-900 rounded-lg overflow-hidden flex flex-col animate-pulse">
@@ -96,10 +116,11 @@ export function PerformancePanel() {
return ( return (
<> <>
<div className="bg-zinc-900 rounded-lg overflow-hidden flex flex-col"> <div className="bg-zinc-900 rounded-lg overflow-hidden flex flex-col">
{/* ── Header ─────────────────────────────────────────────────────── */} {/* ── Header ─────────────────────────────────────────────────────── */}
<div className="flex items-center gap-2 px-3 py-2 border-b border-zinc-800 shrink-0"> <div className="flex items-center gap-2 px-3 py-2 border-b border-zinc-800 shrink-0">
<span className="text-xs text-zinc-500 uppercase tracking-wider font-semibold">Performance</span> <span className="text-xs text-zinc-500 uppercase tracking-wider font-semibold">
Performance
</span>
{hasPending && ( {hasPending && (
<span className="inline-flex items-center gap-1 px-2 py-0.5 rounded-full bg-cyan-500/15 border border-cyan-500/30 text-cyan-400 text-xs"> <span className="inline-flex items-center gap-1 px-2 py-0.5 rounded-full bg-cyan-500/15 border border-cyan-500/30 text-cyan-400 text-xs">
@@ -117,7 +138,10 @@ export function PerformancePanel() {
Apply Apply
</button> </button>
<button <button
onClick={() => { setPending({}); setError(null); }} onClick={() => {
setPending({});
setError(null);
}}
disabled={!hasPending || busy} disabled={!hasPending || busy}
className="flex items-center gap-1.5 px-2 py-1 rounded bg-zinc-800 hover:bg-zinc-700 text-zinc-300 text-xs transition-colors disabled:opacity-40 disabled:cursor-not-allowed" className="flex items-center gap-1.5 px-2 py-1 rounded bg-zinc-800 hover:bg-zinc-700 text-zinc-300 text-xs transition-colors disabled:opacity-40 disabled:cursor-not-allowed"
> >
@@ -139,25 +163,100 @@ export function PerformancePanel() {
{error && ( {error && (
<div className="px-3 py-1.5 bg-red-900/40 border-b border-red-700 text-red-300 text-xs flex items-center justify-between"> <div className="px-3 py-1.5 bg-red-900/40 border-b border-red-700 text-red-300 text-xs flex items-center justify-between">
<span>⚠ {error}</span> <span>⚠ {error}</span>
<button onClick={() => setError(null)} className="ml-2 text-red-400 hover:text-red-200">✕</button> <button
onClick={() => setError(null)}
className="ml-2 text-red-400 hover:text-red-200"
>
✕
</button>
</div> </div>
)} )}
<div className="flex flex-col divide-y divide-zinc-800"> <div className="flex flex-col divide-y divide-zinc-800">
{/* ── Experimental NVIDIA power control ─────────────────────────── */}
{(limits.rm_power_supported || limits.power_cap_mode === "ioctl") && (
<div className="px-4 py-3 flex flex-col gap-2">
<label className="flex items-center gap-2 cursor-pointer select-none">
<input
type="checkbox"
checked={limits.power_cap_mode === "ioctl"}
disabled={busy}
onChange={(e) => handleModeChange(e.target.checked)}
className="accent-red-500"
/>
<span className="text-xs text-zinc-300">
Experimental NVIDIA power control
</span>
</label>
{limits.rm_power_supported ? (
<div
role="alert"
className={
"px-2.5 py-1.5 rounded border text-xs leading-relaxed " +
(limits.power_cap_mode === "ioctl"
? "bg-red-950/80 border-red-500 text-red-300"
: "bg-red-950/40 border-red-800 text-red-400")
}
>
<span className="font-bold">⚠ WARNING:</span> uses an
undocumented driver interface for ALL power limits, including
resets. Allows values below the VBIOS minimum (down to 30 W)
and may cause instability or stop working after driver
updates. Enable at your own risk. Based on the work of{" "}
<a
href="https://github.com/ilya-zlobintsev/LACT/pull/1205"
target="_blank"
rel="noopener noreferrer"
className="underline hover:text-red-200"
>
panchovix
</a>{" "}
(LACT PR #1205).
</div>
) : (
<div
role="alert"
className="px-2.5 py-1.5 rounded border border-red-500 bg-red-950/80 text-red-300 text-xs leading-relaxed"
>
<span className="font-bold">
⚠ Interface not currently available.
</span>
The driver no longer exposes the RM power interface (it may
have been updated). Experimental mode is still enabled, so
power-limit changes will fail. Uncheck to switch back to the
standard NVML mode.
</div>
)}
</div>
)}
{/* ── Board Power Limit ─────────────────────────────────────────── */} {/* ── Board Power Limit ─────────────────────────────────────────── */}
<div className="px-4 py-4 flex flex-col gap-3"> <div className="px-4 py-4 flex flex-col gap-3">
<div className="flex items-center justify-between"> <div className="flex items-center justify-between">
<span className="text-xs text-zinc-500 uppercase tracking-wider">Board Power Limit</span> <div className="flex items-center gap-2">
<span className="text-xs text-zinc-500 uppercase tracking-wider">
Board Power Limit
</span>
{limits.power_cap_mode === "ioctl" &&
limits.min_power_limit_w_native != null && (
<span
className="text-xs text-red-400/80 font-mono"
title="Native VBIOS minimum — experimental mode allows lower"
>
VBIOS min {limits.min_power_limit_w_native} W
</span>
)}
</div>
<div className="flex items-center gap-1.5"> <div className="flex items-center gap-1.5">
<input <input
type="number" type="number"
min={pwrMin} min={pwrMin}
max={pwrMax} max={pwrMax}
value={pwrVal} value={pwrVal}
onChange={e => { onChange={(e) => {
const v = parseInt(e.target.value); const v = parseInt(e.target.value);
if (!isNaN(v)) setPending(p => ({ ...p, power_limit_w: v })); if (!isNaN(v))
setPending((p) => ({ ...p, power_limit_w: v }));
}} }}
className="w-14 bg-zinc-950 border border-zinc-800 rounded text-xs px-2 py-1 text-right font-mono focus:outline-none focus:border-cyan-500" className="w-14 bg-zinc-950 border border-zinc-800 rounded text-xs px-2 py-1 text-right font-mono focus:outline-none focus:border-cyan-500"
/> />
@@ -165,32 +264,44 @@ export function PerformancePanel() {
</div> </div>
</div> </div>
<div className="flex items-center gap-2"> <div className="flex items-center gap-2">
<span className="text-xs text-zinc-600 font-mono w-8 text-right">{pwrMin}</span> <span className="text-xs text-zinc-600 font-mono w-8 text-right">
{pwrMin}
</span>
<input <input
type="range" type="range"
min={pwrMin} min={pwrMin}
max={pwrMax} max={pwrMax}
value={pwrVal} value={pwrVal}
onChange={e => setPending(p => ({ ...p, power_limit_w: parseInt(e.target.value) }))} onChange={(e) =>
setPending((p) => ({
...p,
power_limit_w: parseInt(e.target.value),
}))
}
className="flex-1 accent-cyan-400 h-1 cursor-pointer" className="flex-1 accent-cyan-400 h-1 cursor-pointer"
/> />
<span className="text-xs text-zinc-600 font-mono w-8">{pwrMax}</span> <span className="text-xs text-zinc-600 font-mono w-8">
{pwrMax}
</span>
</div> </div>
</div> </div>
{/* ── Memory Clock Offset ───────────────────────────────────────── */} {/* ── Memory Clock Offset ───────────────────────────────────────── */}
<div className="px-4 py-4 flex flex-col gap-3"> <div className="px-4 py-4 flex flex-col gap-3">
<div className="flex items-center justify-between"> <div className="flex items-center justify-between">
<span className="text-xs text-zinc-500 uppercase tracking-wider">Memory Clock Offset</span> <span className="text-xs text-zinc-500 uppercase tracking-wider">
Memory Clock Offset
</span>
<div className="flex items-center gap-1.5"> <div className="flex items-center gap-1.5">
<input <input
type="number" type="number"
min={memMin} min={memMin}
max={memMax} max={memMax}
value={memVal} value={memVal}
onChange={e => { onChange={(e) => {
const v = parseInt(e.target.value); const v = parseInt(e.target.value);
if (!isNaN(v)) setPending(p => ({ ...p, mem_offset_mhz: v })); if (!isNaN(v))
setPending((p) => ({ ...p, mem_offset_mhz: v }));
}} }}
className="w-16 bg-zinc-950 border border-zinc-800 rounded text-xs px-2 py-1 text-right font-mono focus:outline-none focus:border-cyan-500" className="w-16 bg-zinc-950 border border-zinc-800 rounded text-xs px-2 py-1 text-right font-mono focus:outline-none focus:border-cyan-500"
/> />
@@ -198,20 +309,28 @@ export function PerformancePanel() {
</div> </div>
</div> </div>
<div className="flex items-center gap-2"> <div className="flex items-center gap-2">
<span className="text-xs text-zinc-600 font-mono w-10 text-right">{memMin}</span> <span className="text-xs text-zinc-600 font-mono w-10 text-right">
{memMin}
</span>
<input <input
type="range" type="range"
min={memMin} min={memMin}
max={memMax} max={memMax}
step={1} step={1}
value={memVal} value={memVal}
onChange={e => setPending(p => ({ ...p, mem_offset_mhz: parseInt(e.target.value) }))} onChange={(e) =>
setPending((p) => ({
...p,
mem_offset_mhz: parseInt(e.target.value),
}))
}
className="flex-1 accent-cyan-400 h-1 cursor-pointer" className="flex-1 accent-cyan-400 h-1 cursor-pointer"
/> />
<span className="text-xs text-zinc-600 font-mono w-10">+{memMax}</span> <span className="text-xs text-zinc-600 font-mono w-10">
+{memMax}
</span>
</div> </div>
</div> </div>
</div> </div>
</div> </div>
+45 -18
View File
@@ -1,6 +1,6 @@
import { GaugeCard } from './GaugeCard'; import { GaugeCard } from "./GaugeCard.js";
import { fmt } from '../../utils/units'; import { fmt } from "../../utils/units.js";
import type { MonitoringSample, FanPoint } from '../../types'; import type { MonitoringSample, FanPoint } from "../../types.js";
interface Props { interface Props {
monitor: MonitoringSample | null; monitor: MonitoringSample | null;
@@ -16,7 +16,10 @@ function pluck<K extends keyof MonitoringSample>(
return history.map((s) => (s[key] as number | null) ?? 0); return history.map((s) => (s[key] as number | null) ?? 0);
} }
function computeTargetFan(curve: FanPoint[] | null, tempC: number | null): number | null { function computeTargetFan(
curve: FanPoint[] | null,
tempC: number | null,
): number | null {
if (!curve || !tempC || curve.length < 2) return null; if (!curve || !tempC || curve.length < 2) return null;
for (let i = 0; i < curve.length - 1; i++) { for (let i = 0; i < curve.length - 1; i++) {
@@ -36,19 +39,30 @@ function computeTargetFan(curve: FanPoint[] | null, tempC: number | null): numbe
return curve[curve.length - 1].fan_pct; return curve[curve.length - 1].fan_pct;
} }
export function FanMonitor({ monitor, history, fanCurve, fanCurveActive }: Props) { export function FanMonitor({
const fanHistory = pluck(history, 'fan_pct'); monitor,
const tempHistory = pluck(history, 'temp_c'); history,
fanCurve,
fanCurveActive,
}: Props) {
const fanHistory = pluck(history, "fan_pct");
const tempHistory = pluck(history, "temp_c");
const fans = monitor?.fans ?? null;
const multiFan = fans !== null && fans.length > 1;
const currentTemp = monitor?.temp_c ?? null; const currentTemp = monitor?.temp_c ?? null;
const targetFan = computeTargetFan(fanCurve, currentTemp); const targetFan = computeTargetFan(fanCurve, currentTemp);
const targetFanHistory = history.map((s) => computeTargetFan(fanCurve, s.temp_c) ?? 0); const targetFanHistory = history.map(
(s) => computeTargetFan(fanCurve, s.temp_c) ?? 0,
);
return ( return (
<div className="flex flex-col gap-2 w-full h-full"> <div className="flex flex-col gap-2 w-full h-full">
<div className="bg-zinc-900 rounded-lg p-3 flex flex-col gap-2 h-full"> <div className="bg-zinc-900 rounded-lg p-3 flex flex-col gap-2 h-full">
<div className="flex items-center justify-between pb-2 border-b border-zinc-800"> <div className="flex items-center justify-between pb-2 border-b border-zinc-800">
<span className="text-xs text-zinc-500 uppercase tracking-wider font-semibold">Live Monitor</span> <span className="text-xs text-zinc-500 uppercase tracking-wider font-semibold">
Live Monitor
</span>
{fanCurveActive && ( {fanCurveActive && (
<span className="inline-flex items-center gap-1 px-1.5 py-0.5 rounded-full bg-orange-500/15 border border-orange-500/30 text-orange-400 text-[10px] font-semibold"> <span className="inline-flex items-center gap-1 px-1.5 py-0.5 rounded-full bg-orange-500/15 border border-orange-500/30 text-orange-400 text-[10px] font-semibold">
Curve Active Curve Active
@@ -56,13 +70,26 @@ export function FanMonitor({ monitor, history, fanCurve, fanCurveActive }: Props
)} )}
</div> </div>
<div className="flex flex-col gap-2 mt-1"> <div className="flex flex-col gap-2 mt-1">
<GaugeCard {multiFan ? (
label="Fan Speed" fans!.map((_, i) => (
value={fmt.pct(monitor?.fan_pct)} <GaugeCard
history={fanHistory} key={i}
color="#fb923c" label={`Fan ${i + 1}`}
max={100} value={fmt.pct(fans![i])}
/> history={history.map((s) => s.fans?.[i] ?? s.fan_pct ?? 0)}
color="#fb923c"
max={100}
/>
))
) : (
<GaugeCard
label="Fan Speed"
value={fmt.pct(monitor?.fan_pct)}
history={fanHistory}
color="#fb923c"
max={100}
/>
)}
<GaugeCard <GaugeCard
label="GPU Temp" label="GPU Temp"
value={fmt.celsius(monitor?.temp_c)} value={fmt.celsius(monitor?.temp_c)}
@@ -73,7 +100,7 @@ export function FanMonitor({ monitor, history, fanCurve, fanCurveActive }: Props
{fanCurveActive && ( {fanCurveActive && (
<GaugeCard <GaugeCard
label="Target Fan" label="Target Fan"
value={targetFan !== null ? `${targetFan}%` : '—'} value={targetFan !== null ? `${targetFan}%` : "—"}
history={targetFanHistory} history={targetFanHistory}
color="#fbbf24" color="#fbbf24"
max={100} max={100}
@@ -82,7 +109,7 @@ export function FanMonitor({ monitor, history, fanCurve, fanCurveActive }: Props
<div className="mt-auto"> <div className="mt-auto">
<GaugeCard <GaugeCard
label="Fan Mode" label="Fan Mode"
value={fanCurveActive ? 'Curve' : 'Auto'} value={fanCurveActive ? "Curve" : "Auto"}
/> />
</div> </div>
</div> </div>
+41 -39
View File
@@ -1,6 +1,6 @@
import { GaugeCard } from './GaugeCard'; import { GaugeCard } from "./GaugeCard.js";
import { fmt } from '../../utils/units'; import { fmt } from "../../utils/units.js";
import type { MonitoringSample } from '../../types'; import type { MonitoringSample } from "../../types.js";
interface Props { interface Props {
monitor: MonitoringSample | null; monitor: MonitoringSample | null;
@@ -19,48 +19,50 @@ export function LiveMonitor({ monitor, history }: Props) {
<div className="flex flex-col gap-2 w-full h-full"> <div className="flex flex-col gap-2 w-full h-full">
<div className="bg-zinc-900 rounded-lg p-3 flex flex-col gap-2 h-full"> <div className="bg-zinc-900 rounded-lg p-3 flex flex-col gap-2 h-full">
<div className="flex items-center justify-between pb-2 border-b border-zinc-800"> <div className="flex items-center justify-between pb-2 border-b border-zinc-800">
<span className="text-xs text-zinc-500 uppercase tracking-wider font-semibold">Live Monitor</span> <span className="text-xs text-zinc-500 uppercase tracking-wider font-semibold">
Live Monitor
</span>
</div> </div>
<div className="flex flex-col gap-2 mt-1"> <div className="flex flex-col gap-2 mt-1">
<GaugeCard
label="Core Clock"
value={fmt.mhz(monitor?.clock_mhz)}
history={pluck(history, 'clock_mhz')}
color="#34d399"
max={3000}
/>
<GaugeCard
label="Voltage"
value={fmt.mv(monitor?.voltage_mv)}
history={pluck(history, 'voltage_mv')}
color="#a78bfa"
max={1100}
/>
<GaugeCard
label="Power Draw"
value={fmt.watts(monitor?.power_w)}
history={pluck(history, 'power_w')}
color="#f472b6"
max={600}
/>
<GaugeCard
label="GPU Util"
value={fmt.pct(monitor?.gpu_util_pct)}
history={pluck(history, 'gpu_util_pct')}
color="#facc15"
max={100}
/>
<div className="mt-auto">
<GaugeCard <GaugeCard
label="P-State" label="Core Clock"
value={monitor?.pstate_label ?? 'Unknown'} value={fmt.mhz(monitor?.clock_mhz)}
history={pluck(history, 'pstate')} history={pluck(history, "clock_mhz")}
color="#a8a29e" color="#34d399"
max={15} max={3000}
/> />
<GaugeCard
label="Voltage"
value={fmt.mv(monitor?.voltage_mv)}
history={pluck(history, "voltage_mv")}
color="#a78bfa"
max={1100}
/>
<GaugeCard
label="Power Draw"
value={fmt.watts(monitor?.power_w)}
history={pluck(history, "power_w")}
color="#f472b6"
max={600}
/>
<GaugeCard
label="GPU Util"
value={fmt.pct(monitor?.gpu_util_pct)}
history={pluck(history, "gpu_util_pct")}
color="#facc15"
max={100}
/>
<div className="mt-auto">
<GaugeCard
label="P-State"
value={monitor?.pstate_label ?? "Unknown"}
history={pluck(history, "pstate")}
color="#a8a29e"
max={15}
/>
</div>
</div> </div>
</div> </div>
</div>
</div> </div>
); );
} }
@@ -1,6 +1,6 @@
import { GaugeCard } from './GaugeCard'; import { GaugeCard } from "./GaugeCard.js";
import { fmt } from '../../utils/units'; import { fmt } from "../../utils/units.js";
import type { MonitoringSample } from '../../types'; import type { MonitoringSample } from "../../types.js";
interface Props { interface Props {
monitor: MonitoringSample | null; monitor: MonitoringSample | null;
@@ -17,35 +17,38 @@ function pluck<K extends keyof MonitoringSample>(
export function PerformanceMonitor({ monitor, history }: Props) { export function PerformanceMonitor({ monitor, history }: Props) {
const memUsed = monitor?.mem_used_mib ?? null; const memUsed = monitor?.mem_used_mib ?? null;
const memTotal = monitor?.mem_total_mib ?? null; const memTotal = monitor?.mem_total_mib ?? null;
const memLabel = memUsed != null && memTotal != null const memLabel =
? `${memUsed.toFixed(0)} / ${memTotal.toFixed(0)} MiB` memUsed != null && memTotal != null
: '—'; ? `${memUsed.toFixed(0)} / ${memTotal.toFixed(0)} MiB`
: "—";
return ( return (
<div className="flex flex-col gap-2 w-full h-full"> <div className="flex flex-col gap-2 w-full h-full">
<div className="bg-zinc-900 rounded-lg p-3 flex flex-col gap-2 h-full"> <div className="bg-zinc-900 rounded-lg p-3 flex flex-col gap-2 h-full">
<div className="flex items-center justify-between pb-2 border-b border-zinc-800"> <div className="flex items-center justify-between pb-2 border-b border-zinc-800">
<span className="text-xs text-zinc-500 uppercase tracking-wider font-semibold">Live Monitor</span> <span className="text-xs text-zinc-500 uppercase tracking-wider font-semibold">
Live Monitor
</span>
</div> </div>
<div className="flex flex-col gap-2 mt-1"> <div className="flex flex-col gap-2 mt-1">
<GaugeCard <GaugeCard
label="Mem Clock" label="Mem Clock"
value={fmt.mhz(monitor?.mem_clock_mhz)} value={fmt.mhz(monitor?.mem_clock_mhz)}
history={pluck(history, 'mem_clock_mhz')} history={pluck(history, "mem_clock_mhz")}
color="#67e8f9" color="#67e8f9"
max={20000} max={20000}
/> />
<GaugeCard <GaugeCard
label="Power Draw" label="Power Draw"
value={fmt.watts(monitor?.power_w)} value={fmt.watts(monitor?.power_w)}
history={pluck(history, 'power_w')} history={pluck(history, "power_w")}
color="#f472b6" color="#f472b6"
max={600} max={600}
/> />
<GaugeCard <GaugeCard
label="VRAM Used" label="VRAM Used"
value={memLabel} value={memLabel}
history={pluck(history, 'mem_used_mib')} history={pluck(history, "mem_used_mib")}
color="#a78bfa" color="#a78bfa"
max={memTotal ?? 32768} max={memTotal ?? 32768}
/> />
@@ -0,0 +1,254 @@
import { useEffect, useRef, useState } from "react";
import { Loader, X } from "lucide-react";
import { api } from "../../api/client.js";
import { useCurveStore } from "../../store/curveStore.js";
import type { GpuProcess } from "../../types.js";
import { fmt } from "../../utils/units.js";
import { toast } from "sonner";
import { ConfirmDialog } from "../common/ConfirmDialog.js";
const POLL_MS = 3000;
const TYPE_STYLES: Record<GpuProcess["type"], { label: string; cls: string }> =
{
C: { label: "C", cls: "bg-cyan-500/15 text-cyan-400 border-cyan-500/30" },
G: {
label: "G",
cls: "bg-emerald-500/15 text-emerald-400 border-emerald-500/30",
},
"C+G": {
label: "C+G",
cls: "bg-violet-500/15 text-violet-400 border-violet-500/30",
},
};
const TH =
"px-3 py-1.5 font-semibold text-left text-[10px] uppercase tracking-wider text-zinc-500 whitespace-nowrap";
const TD = "px-3 py-1.5 whitespace-nowrap";
export function ProcessList() {
const { selectedGpuIndex } = useCurveStore();
const [processes, setProcesses] = useState<GpuProcess[] | null>(null);
const [memTotal, setMemTotal] = useState<number | null>(null);
const [error, setError] = useState<string | null>(null);
const [killTarget, setKillTarget] = useState<GpuProcess | null>(null);
const [pendingAction, setPendingAction] = useState<
"TERM" | "KILL" | "KILL_PARENT" | null
>(null);
// Ref so handleKill can trigger an immediate refresh after a kill.
const refreshRef = useRef<() => void>(() => {});
useEffect(() => {
let cancelled = false;
async function fetchOnce() {
try {
const data = await api.processes(selectedGpuIndex);
if (cancelled) return;
setProcesses(data.processes);
setMemTotal(data.mem_total_bytes);
setError(null);
} catch (e) {
if (!cancelled)
setError(e instanceof Error ? e.message : String(e));
}
}
refreshRef.current = () => {
void fetchOnce();
};
fetchOnce();
const timer = setInterval(fetchOnce, POLL_MS);
return () => {
cancelled = true;
clearInterval(timer);
};
}, [selectedGpuIndex]);
async function handleKill(action: "TERM" | "KILL" | "KILL_PARENT") {
if (!killTarget || pendingAction) return;
const pid = killTarget.pid;
setPendingAction(action);
try {
const signal = action === "TERM" ? "TERM" : "KILL";
const parent = action === "KILL_PARENT";
const res = await api.killProcess(pid, signal, parent);
toast.success(
parent
? `Sent SIGKILL to parent process ${res.pid} of ${pid}`
: `Sent SIG${signal} to process ${pid}`,
);
setKillTarget(null);
refreshRef.current();
} catch (e) {
toast.error(e instanceof Error ? e.message : String(e));
} finally {
setPendingAction(null);
}
}
return (
<div className="bg-zinc-900 rounded-lg overflow-hidden flex flex-col">
{/* ── Header ─────────────────────────────────────────────────────── */}
<div className="flex items-center gap-2 px-3 py-2 border-b border-zinc-800">
<span className="text-xs text-zinc-500 uppercase tracking-wider font-semibold">
GPU Processes
</span>
{processes && processes.length > 0 && (
<span className="text-xs text-zinc-600">{processes.length}</span>
)}
<div className="ml-auto flex items-center gap-2">
{memTotal != null && (
<span className="text-xs text-zinc-600 font-mono">
VRAM total {fmt.bytes(memTotal)}
</span>
)}
</div>
</div>
{/* ── Error banner ────────────────────────────────────────────────── */}
{error && (
<div className="px-3 py-1.5 bg-red-900/40 border-b border-red-700 text-red-300 text-xs flex items-center justify-between">
<span>⚠ {error}</span>
<button
onClick={() => setError(null)}
className="ml-2 text-red-400 hover:text-red-200"
>
✕
</button>
</div>
)}
{/* ── Table ───────────────────────────────────────────────────────── */}
{processes === null && !error ? (
<div className="p-8 flex items-center justify-center gap-2 text-zinc-500 text-xs">
<Loader size={14} className="animate-spin" />
Loading processes…
</div>
) : processes && processes.length === 0 && !error ? (
<div className="p-8 text-center text-zinc-600 text-xs">
{memTotal === null
? "GPU process data unavailable (NVML not initialized)."
: "No processes using this GPU."}
</div>
) : (
<div className="overflow-x-auto">
<table className="w-full text-xs">
<thead>
<tr className="border-b border-zinc-800">
<th className={TH}>PID</th>
<th className={TH}>User</th>
<th className={TH}>Dev</th>
<th className={TH}>Type</th>
<th
className={`${TH} text-right`}
title="NVML per-process utilization — smoothed over the driver's sample window, so it lags instantaneous usage (accurate when steady)"
>
GPU
</th>
<th className={`${TH} text-right`}>VRAM</th>
<th className={`${TH} text-right`}>VRAM %</th>
<th className={`${TH} text-right`}>CPU</th>
<th className={`${TH} text-right`}>MEM</th>
<th className={TH}>Command</th>
<th className={`${TH} text-right`}>Kill</th>
</tr>
</thead>
<tbody className="divide-y divide-zinc-800/60 font-mono">
{processes?.map((p) => (
<tr key={p.pid} className="hover:bg-zinc-800/40">
<td className={TD}>{p.pid}</td>
<td className={TD}>{p.user}</td>
<td className={TD}>{p.dev}</td>
<td className={TD}>
<span
className={`inline-block px-1.5 py-0.5 rounded border text-[10px] font-semibold ${TYPE_STYLES[p.type].cls}`}
title={
p.type === "C"
? "Compute"
: p.type === "G"
? "Graphics"
: "Compute + Graphics"
}
>
{TYPE_STYLES[p.type].label}
</span>
</td>
<td className={`${TD} text-right`}>
{fmt.pct(p.gpu_util_pct ?? 0)}
</td>
<td className={`${TD} text-right`}>
{fmt.bytes(p.vram_bytes)}
</td>
<td className={`${TD} text-right`}>
{fmt.pct(p.vram_pct, 1)}
</td>
<td className={`${TD} text-right`}>
{fmt.pct(p.cpu_pct, 1)}
</td>
<td className={`${TD} text-right`}>
{fmt.bytes(p.mem_host_bytes)}
</td>
<td className="px-3 py-1.5 max-w-[320px] min-w-[160px]">
<span
className="block truncate text-zinc-300"
title={p.command}
>
{p.command}
</span>
</td>
<td className={`${TD} text-right`}>
<button
onClick={() => setKillTarget(p)}
className="inline-flex items-center gap-1 px-2 py-1 rounded bg-zinc-800 hover:bg-red-900 text-zinc-400 hover:text-red-300 text-xs transition-colors"
title={`Kill process ${p.pid}`}
>
<X size={11} />
Kill
</button>
</td>
</tr>
))}
</tbody>
</table>
</div>
)}
{/* ── Kill confirmation ───────────────────────────────────────────── */}
{killTarget && (
<ConfirmDialog
message={`Kill process ${killTarget.pid}?`}
detail={`${killTarget.command || "(unknown command)"} — owned by ${killTarget.user}, using ${fmt.bytes(
killTarget.vram_bytes,
)} VRAM. SIGTERM asks the process to exit; SIGKILL kills it immediately; SIGKILL Parent kills the parent process, which may take down the whole process family.`}
onCancel={() => setKillTarget(null)}
actions={[
{
label: pendingAction === "TERM" ? "Sending…" : "SIGTERM",
onClick: () => handleKill("TERM"),
disabled: pendingAction != null,
className:
"px-3 py-1.5 rounded text-xs font-semibold transition-colors bg-zinc-800 hover:bg-zinc-700 text-zinc-200",
},
{
label: pendingAction === "KILL" ? "Sending…" : "SIGKILL",
onClick: () => handleKill("KILL"),
disabled: pendingAction != null,
className:
"px-3 py-1.5 rounded text-xs font-semibold transition-colors bg-red-600 hover:bg-red-500 text-white",
},
{
label:
pendingAction === "KILL_PARENT" ? "Sending…" : "SIGKILL Parent",
onClick: () => handleKill("KILL_PARENT"),
disabled: pendingAction != null,
className:
"px-3 py-1.5 rounded text-xs font-semibold transition-colors bg-red-900 hover:bg-red-800 text-red-100",
},
]}
/>
)}
</div>
);
}
+29 -4
View File
@@ -7,10 +7,13 @@ import {
ChevronDown, ChevronDown,
LogOut, LogOut,
User, User,
Bug,
} from "lucide-react"; } from "lucide-react";
import type { GpuInfo, MonitoringSample } from "../../types"; import type { GpuInfo, MonitoringSample } from "../../types.js";
import { fmt } from "../../utils/units"; import { fmt } from "../../utils/units.js";
import { useCurveStore } from "../../store/curveStore"; import { useCurveStore } from "../../store/curveStore.js";
import { api } from "../../api/client.js";
import { ReportBugDialog } from "../common/ReportBugDialog.js";
import { useState, useRef, useEffect } from "react"; import { useState, useRef, useEffect } from "react";
interface Props { interface Props {
@@ -43,8 +46,17 @@ export function StatusBar({
const { availableGpus, selectedGpuIndex, setSelectedGpuIndex } = const { availableGpus, selectedGpuIndex, setSelectedGpuIndex } =
useCurveStore(); useCurveStore();
const [isOpen, setIsOpen] = useState(false); const [isOpen, setIsOpen] = useState(false);
const [bugReportConfigured, setBugReportConfigured] = useState(false);
const [reportBugOpen, setReportBugOpen] = useState(false);
const dropdownRef = useRef<HTMLDivElement>(null); const dropdownRef = useRef<HTMLDivElement>(null);
useEffect(() => {
api
.reportBugStatus()
.then((s) => setBugReportConfigured(s.configured))
.catch(() => {});
}, []);
useEffect(() => { useEffect(() => {
function handleClickOutside(event: MouseEvent) { function handleClickOutside(event: MouseEvent) {
if ( if (
@@ -169,8 +181,18 @@ export function StatusBar({
/> />
</div> </div>
{/* Right: user + connection status */} {/* Right: report bug + user + connection status */}
<div className="flex items-center gap-3 shrink-0"> <div className="flex items-center gap-3 shrink-0">
{bugReportConfigured && (
<button
onClick={() => setReportBugOpen(true)}
title="Report a bug (creates a Gitea issue)"
className="flex items-center gap-1.5 px-2 py-1 rounded hover:bg-zinc-800 text-zinc-400 hover:text-zinc-200 transition-colors"
>
<Bug size={14} />
<span className="hidden sm:inline text-xs">Report Bug</span>
</button>
)}
{user && ( {user && (
<div className="flex items-center gap-1.5 text-zinc-400"> <div className="flex items-center gap-1.5 text-zinc-400">
<User size={14} className="text-zinc-500" /> <User size={14} className="text-zinc-500" />
@@ -192,6 +214,9 @@ export function StatusBar({
</div> </div>
</div> </div>
</div> </div>
{reportBugOpen && (
<ReportBugDialog onClose={() => setReportBugOpen(false)} />
)}
</header> </header>
); );
} }
+58 -22
View File
@@ -1,7 +1,7 @@
import { useState, useRef, useEffect } from 'react'; import { useState, useRef, useEffect } from "react";
import { fmt } from '../../utils/units'; import { fmt } from "../../utils/units.js";
import type { VFPoint } from '../../types'; import type { VFPoint } from "../../types.js";
import { useCurveStore } from '../../store/curveStore'; import { useCurveStore } from "../../store/curveStore.js";
interface Props { interface Props {
point: VFPoint; point: VFPoint;
@@ -14,17 +14,28 @@ interface Props {
onMouseEnter?: () => void; onMouseEnter?: () => void;
} }
export function PointRow({ point, isCurrent, isSelected, isClamped, pendingDeltaKhz, shouldAutoScroll, onMouseDown, onMouseEnter }: Props) { export function PointRow({
point,
isCurrent,
isSelected,
isClamped,
pendingDeltaKhz,
shouldAutoScroll,
onMouseDown,
onMouseEnter,
}: Props) {
const { stageEdit } = useCurveStore(); const { stageEdit } = useCurveStore();
const [editing, setEditing] = useState(false); const [editing, setEditing] = useState(false);
const [inputValue, setInputValue] = useState(''); const [inputValue, setInputValue] = useState("");
const inputRef = useRef<HTMLInputElement>(null); const inputRef = useRef<HTMLInputElement>(null);
const trRef = useRef<HTMLTableRowElement>(null); const trRef = useRef<HTMLTableRowElement>(null);
useEffect(() => { useEffect(() => {
if (shouldAutoScroll && trRef.current) { if (shouldAutoScroll && trRef.current) {
// @ts-expect-error: the typing seems to not include the valid 'container' option trRef.current.scrollIntoView({
trRef.current.scrollIntoView({ behavior: 'smooth', block: 'nearest', container: 'nearest' }); behavior: "smooth",
block: "nearest",
});
} }
}, [shouldAutoScroll]); }, [shouldAutoScroll]);
@@ -36,11 +47,19 @@ export function PointRow({ point, isCurrent, isSelected, isClamped, pendingDelta
const displayEffMhz = point.freq_mhz + deltaChange / 1000; const displayEffMhz = point.freq_mhz + deltaChange / 1000;
const deltaColor = hasPending const deltaColor = hasPending
? displayDeltaKhz > 0 ? 'text-cyan-400' : displayDeltaKhz < 0 ? 'text-orange-400' : 'text-zinc-400' ? displayDeltaKhz > 0
: point.delta_khz > 0 ? 'text-emerald-400' : point.delta_khz < 0 ? 'text-red-400' : 'text-zinc-500'; ? "text-cyan-400"
: displayDeltaKhz < 0
? "text-orange-400"
: "text-zinc-400"
: point.delta_khz > 0
? "text-emerald-400"
: point.delta_khz < 0
? "text-red-400"
: "text-zinc-500";
function startEdit() { function startEdit() {
setInputValue((displayDeltaMhz).toFixed(1)); setInputValue(displayDeltaMhz.toFixed(1));
setEditing(true); setEditing(true);
setTimeout(() => { setTimeout(() => {
inputRef.current?.select(); inputRef.current?.select();
@@ -64,9 +83,13 @@ export function PointRow({ point, isCurrent, isSelected, isClamped, pendingDelta
<tr <tr
ref={trRef} ref={trRef}
className={[ className={[
'border-b border-zinc-800 text-xs font-mono cursor-pointer', "border-b border-zinc-800 text-xs font-mono cursor-pointer",
isCurrent ? 'bg-yellow-400/10' : isSelected ? 'bg-cyan-500/10' : 'hover:bg-zinc-800/50', isCurrent
].join(' ')} ? "bg-yellow-400/10"
: isSelected
? "bg-cyan-500/10"
: "hover:bg-zinc-800/50",
].join(" ")}
onMouseDown={(e) => { onMouseDown={(e) => {
if (editing) return; if (editing) return;
onMouseDown?.(e); onMouseDown?.(e);
@@ -80,7 +103,13 @@ export function PointRow({ point, isCurrent, isSelected, isClamped, pendingDelta
<td className="px-3 py-1 text-zinc-300">{fmt.mv(point.volt_mv, 0)}</td> <td className="px-3 py-1 text-zinc-300">{fmt.mv(point.volt_mv, 0)}</td>
{/* Offset — click to edit inline */} {/* Offset — click to edit inline */}
<td className={`px-3 py-1 ${deltaColor}`} onClick={(e) => { e.stopPropagation(); startEdit(); }}> <td
className={`px-3 py-1 ${deltaColor}`}
onClick={(e) => {
e.stopPropagation();
startEdit();
}}
>
{editing ? ( {editing ? (
<input <input
ref={inputRef} ref={inputRef}
@@ -90,28 +119,35 @@ export function PointRow({ point, isCurrent, isSelected, isClamped, pendingDelta
onChange={(e) => setInputValue(e.target.value)} onChange={(e) => setInputValue(e.target.value)}
onBlur={commitEdit} onBlur={commitEdit}
onKeyDown={(e) => { onKeyDown={(e) => {
if (e.key === 'Enter' || e.key === 'Tab') { e.preventDefault(); commitEdit(); } if (e.key === "Enter" || e.key === "Tab") {
if (e.key === 'Escape') cancelEdit(); e.preventDefault();
commitEdit();
}
if (e.key === "Escape") cancelEdit();
}} }}
className="w-20 bg-zinc-700 text-cyan-300 rounded px-1 py-0 border border-cyan-500 outline-none text-xs" className="w-20 bg-zinc-700 text-cyan-300 rounded px-1 py-0 border border-cyan-500 outline-none text-xs font-mono"
style={{ fontFamily: 'monospace' }}
/> />
) : ( ) : (
<span title="Click to edit"> <span title="Click to edit">
{hasPending && <span className="text-cyan-500 mr-0.5">✎</span>} {hasPending && <span className="text-cyan-500 mr-0.5">✎</span>}
{displayDeltaKhz > 0 ? '+' : ''}{displayDeltaMhz.toFixed(1)} MHz {displayDeltaKhz > 0 ? "+" : ""}
{displayDeltaMhz.toFixed(1)} MHz
</span> </span>
)} )}
</td> </td>
{/* Eff. Freq */} {/* Eff. Freq */}
<td className={`px-3 py-1 font-semibold ${hasPending ? 'text-cyan-200' : 'text-zinc-100'}`}> <td
className={`px-3 py-1 font-semibold ${hasPending ? "text-cyan-200" : "text-zinc-100"}`}
>
{fmt.mhz(displayEffMhz, 0)} {fmt.mhz(displayEffMhz, 0)}
{isClamped && !hasPending && ( {isClamped && !hasPending && (
<span <span
className="ml-1 text-amber-500 cursor-help" className="ml-1 text-amber-500 cursor-help"
title="Clamped by monotonicity — a lower-voltage point with a higher offset is holding this frequency up" title="Clamped by monotonicity — a lower-voltage point with a higher offset is holding this frequency up"
>⇡</span> >
⇡
</span>
)} )}
</td> </td>
@@ -1,8 +1,11 @@
import { useState, useMemo, useEffect } from 'react'; import { useState, useMemo, useEffect } from "react";
import { PointRow } from './PointRow'; import { PointRow } from "./PointRow.js";
import type { VFPoint } from '../../types'; import type { VFPoint } from "../../types.js";
import { findCurrentPoint, detectClampedPoints } from '../../utils/curveHelpers'; import {
import { useCurveStore } from '../../store/curveStore'; findCurrentPoint,
detectClampedPoints,
} from "../../utils/curveHelpers.js";
import { useCurveStore } from "../../store/curveStore.js";
interface Props { interface Props {
points: VFPoint[]; points: VFPoint[];
@@ -11,22 +14,32 @@ interface Props {
} }
export function PointTable({ points, currentVoltageMv, readOnly }: Props) { export function PointTable({ points, currentVoltageMv, readOnly }: Props) {
const { pendingDeltas, selectedPoints, selectPoint, selectRange } = useCurveStore(); const {
pendingDeltas,
selectedPoints,
selectPoint,
togglePoint,
selectRange,
} = useCurveStore();
const currentPoint = findCurrentPoint(points, currentVoltageMv); const currentPoint = findCurrentPoint(points, currentVoltageMv);
const clampedPoints = useMemo(() => detectClampedPoints(points), [points]); const clampedPoints = useMemo(() => detectClampedPoints(points), [points]);
const [dragStartIdx, setDragStartIdx] = useState<number | null>(null); const [dragStartIdx, setDragStartIdx] = useState<number | null>(null);
useEffect(() => { useEffect(() => {
function onUp() { setDragStartIdx(null); } function onUp() {
window.addEventListener('mouseup', onUp); setDragStartIdx(null);
return () => window.removeEventListener('mouseup', onUp); }
window.addEventListener("mouseup", onUp);
return () => window.removeEventListener("mouseup", onUp);
}, []); }, []);
return ( return (
<div className="bg-zinc-900 rounded-lg overflow-hidden flex flex-col"> <div className="bg-zinc-900 rounded-lg overflow-hidden flex flex-col">
<div className="flex items-center gap-2 px-3 py-2 border-b border-zinc-800 shrink-0"> <div className="flex items-center gap-2 px-3 py-2 border-b border-zinc-800 shrink-0">
<span className="text-xs text-zinc-500 uppercase tracking-wider mr-2">Points</span> <span className="text-xs text-zinc-500 uppercase tracking-wider mr-2">
Points
</span>
{!readOnly && pendingDeltas.size > 0 && ( {!readOnly && pendingDeltas.size > 0 && (
<span className="inline-flex items-center gap-1 px-2 py-0.5 rounded-full bg-cyan-500/15 border border-cyan-500/30 text-cyan-400 text-xs"> <span className="inline-flex items-center gap-1 px-2 py-0.5 rounded-full bg-cyan-500/15 border border-cyan-500/30 text-cyan-400 text-xs">
{pendingDeltas.size} staged {pendingDeltas.size} staged
@@ -35,13 +48,17 @@ export function PointTable({ points, currentVoltageMv, readOnly }: Props) {
{readOnly && ( {readOnly && (
<span className="text-xs text-zinc-600 italic">read-only</span> <span className="text-xs text-zinc-600 italic">read-only</span>
)} )}
<span className="ml-auto text-xs text-zinc-600">{points.length} points</span> <span className="ml-auto text-xs text-zinc-600">
{points.length} points
</span>
{!readOnly && selectedPoints.size === 1 && ( {!readOnly && selectedPoints.size === 1 && (
<div className="flex gap-1 ml-4 border-l border-zinc-800 pl-4"> <div className="flex gap-1 ml-4 border-l border-zinc-800 pl-4">
<button <button
onClick={() => { onClick={() => {
const idx = Array.from(selectedPoints)[0]; const idx = Array.from(selectedPoints)[0];
const beforeIdxs = points.filter(p => p.index <= idx).map(p => p.index); const beforeIdxs = points
.filter((p) => p.index <= idx)
.map((p) => p.index);
selectRange(beforeIdxs); selectRange(beforeIdxs);
}} }}
className="px-2 py-0.5 rounded text-xs bg-zinc-800 text-zinc-400 hover:bg-zinc-700 transition-colors whitespace-nowrap" className="px-2 py-0.5 rounded text-xs bg-zinc-800 text-zinc-400 hover:bg-zinc-700 transition-colors whitespace-nowrap"
@@ -51,7 +68,9 @@ export function PointTable({ points, currentVoltageMv, readOnly }: Props) {
<button <button
onClick={() => { onClick={() => {
const idx = Array.from(selectedPoints)[0]; const idx = Array.from(selectedPoints)[0];
const afterIdxs = points.filter(p => p.index >= idx).map(p => p.index); const afterIdxs = points
.filter((p) => p.index >= idx)
.map((p) => p.index);
selectRange(afterIdxs); selectRange(afterIdxs);
}} }}
className="px-2 py-0.5 rounded text-xs bg-zinc-800 text-zinc-400 hover:bg-zinc-700 transition-colors whitespace-nowrap" className="px-2 py-0.5 rounded text-xs bg-zinc-800 text-zinc-400 hover:bg-zinc-700 transition-colors whitespace-nowrap"
@@ -66,9 +85,15 @@ export function PointTable({ points, currentVoltageMv, readOnly }: Props) {
<thead className="sticky top-0 bg-zinc-900 z-10"> <thead className="sticky top-0 bg-zinc-900 z-10">
<tr className="text-xs text-zinc-500 uppercase tracking-wider border-b border-zinc-800"> <tr className="text-xs text-zinc-500 uppercase tracking-wider border-b border-zinc-800">
<th className="px-3 py-2 text-left font-normal bg-zinc-900">#</th> <th className="px-3 py-2 text-left font-normal bg-zinc-900">#</th>
<th className="px-3 py-2 text-left font-normal bg-zinc-900">Voltage</th> <th className="px-3 py-2 text-left font-normal bg-zinc-900">
<th className="px-3 py-2 text-left font-normal bg-zinc-900">Offset</th> Voltage
<th className="px-3 py-2 text-left font-normal bg-zinc-900">Eff. Freq</th> </th>
<th className="px-3 py-2 text-left font-normal bg-zinc-900">
Offset
</th>
<th className="px-3 py-2 text-left font-normal bg-zinc-900">
Eff. Freq
</th>
<th className="px-3 py-2 text-left font-normal bg-zinc-900" /> <th className="px-3 py-2 text-left font-normal bg-zinc-900" />
</tr> </tr>
</thead> </thead>
@@ -81,34 +106,48 @@ export function PointTable({ points, currentVoltageMv, readOnly }: Props) {
isSelected={selectedPoints.has(p.index)} isSelected={selectedPoints.has(p.index)}
isClamped={clampedPoints.has(p.index)} isClamped={clampedPoints.has(p.index)}
pendingDeltaKhz={pendingDeltas.get(p.index)} pendingDeltaKhz={pendingDeltas.get(p.index)}
shouldAutoScroll={selectedPoints.size === 1 && selectedPoints.has(p.index)} shouldAutoScroll={
onMouseDown={readOnly ? undefined : (e) => { selectedPoints.size === 1 && selectedPoints.has(p.index)
if (e.shiftKey) { }
const currentSelected = Array.from(selectedPoints); onMouseDown={
if (currentSelected.length > 0) { readOnly
const last = Math.max(...currentSelected); ? undefined
const min = Math.min(last, p.index); : (e) => {
const max = Math.max(last, p.index); if (e.shiftKey) {
const toSelect = points.filter(a => a.index >= min && a.index <= max).map(a => a.index); const currentSelected = Array.from(selectedPoints);
selectRange(toSelect); if (currentSelected.length > 0) {
} else { const last = Math.max(...currentSelected);
selectPoint(p.index, false); const min = Math.min(last, p.index);
} const max = Math.max(last, p.index);
} else if (e.ctrlKey || e.metaKey) { const toSelect = points
selectPoint(p.index, true); .filter((a) => a.index >= min && a.index <= max)
} else { .map((a) => a.index);
setDragStartIdx(p.index); selectRange(toSelect);
selectPoint(p.index, false); } else {
} selectPoint(p.index);
}} }
onMouseEnter={readOnly ? undefined : () => { } else if (e.ctrlKey || e.metaKey) {
if (dragStartIdx !== null) { togglePoint(p.index);
const min = Math.min(dragStartIdx, p.index); } else {
const max = Math.max(dragStartIdx, p.index); setDragStartIdx(p.index);
const toSelect = points.filter(a => a.index >= min && a.index <= max).map(a => a.index); selectPoint(p.index);
selectRange(toSelect); }
} }
}} }
onMouseEnter={
readOnly
? undefined
: () => {
if (dragStartIdx !== null) {
const min = Math.min(dragStartIdx, p.index);
const max = Math.max(dragStartIdx, p.index);
const toSelect = points
.filter((a) => a.index >= min && a.index <= max)
.map((a) => a.index);
selectRange(toSelect);
}
}
}
/> />
))} ))}
</tbody> </tbody>
@@ -8,10 +8,14 @@ import {
Star, Star,
Fan, Fan,
} from "lucide-react"; } from "lucide-react";
import { api } from "../../api/client"; import { api } from "../../api/client.js";
import type { ProfileData } from "../../types"; import type { ProfileData } from "../../types.js";
import { toast } from "sonner"; import { toast } from "sonner";
import { useCurveStore } from "../../store/curveStore"; import { useCurveStore } from "../../store/curveStore.js";
function errMsg(e: unknown): string {
return e instanceof Error ? e.message : String(e);
}
interface ProfilePanelProps { interface ProfilePanelProps {
activeProfile: string | null; activeProfile: string | null;
@@ -65,37 +69,42 @@ export function ProfilePanel({
setAutoLoadProfile(name); setAutoLoadProfile(name);
if (name) toast.success(`"${name}" will load on server start`); if (name) toast.success(`"${name}" will load on server start`);
else toast.success("Auto-load cleared"); else toast.success("Auto-load cleared");
} catch (e: any) { } catch (e: unknown) {
toast.error( toast.error("Failed to update default profile: " + errMsg(e));
"Failed to update default profile: " + (e.message || String(e)),
);
} }
} }
// Data fetch on GPU change — setState calls happen after the await, not
// synchronously in the effect body (rule false-positive on async fetch).
useEffect(() => { useEffect(() => {
// eslint-disable-next-line react-hooks/set-state-in-effect
fetchProfiles(); fetchProfiles();
}, [selectedGpuIndex]); }, [selectedGpuIndex]);
useEffect(() => { useEffect(() => {
if (isSaveOpen) saveInputRef.current?.focus(); if (isSaveOpen) saveInputRef.current?.focus();
else setNewName("");
}, [isSaveOpen]); }, [isSaveOpen]);
useEffect(() => { useEffect(() => {
if (renamingName) renameInputRef.current?.focus(); if (renamingName) renameInputRef.current?.focus();
}, [renamingName]); }, [renamingName]);
async function handleSave(e: React.FormEvent) { function closeSaveForm() {
setIsSaveOpen(false);
setNewName("");
}
async function handleSave(e: React.SubmitEvent) {
e.preventDefault(); e.preventDefault();
if (!newName.trim()) return; if (!newName.trim()) return;
try { try {
setIsSaving(true); setIsSaving(true);
await api.saveProfile(newName.trim(), selectedGpuIndex); await api.saveProfile(newName.trim(), selectedGpuIndex);
toast.success(`Profile "${newName.trim()}" saved`); toast.success(`Profile "${newName.trim()}" saved`);
setIsSaveOpen(false); closeSaveForm();
await fetchProfiles(); await fetchProfiles();
} catch (e: any) { } catch (e: unknown) {
toast.error("Failed to save: " + (e.message || String(e))); toast.error("Failed to save: " + errMsg(e));
} finally { } finally {
setIsSaving(false); setIsSaving(false);
} }
@@ -107,8 +116,8 @@ export function ProfilePanel({
await api.applyProfile(name, selectedGpuIndex); await api.applyProfile(name, selectedGpuIndex);
onProfileApplied(name); onProfileApplied(name);
toast.success(`"${name}" applied`); toast.success(`"${name}" applied`);
} catch (e: any) { } catch (e: unknown) {
toast.error(`Failed to apply "${name}": ` + (e.message || String(e))); toast.error(`Failed to apply "${name}": ` + errMsg(e));
} finally { } finally {
setApplyingName(null); setApplyingName(null);
} }
@@ -123,14 +132,14 @@ export function ProfilePanel({
if (autoLoadProfile === name) setAutoLoadProfile(null); if (autoLoadProfile === name) setAutoLoadProfile(null);
setDeletingName(null); setDeletingName(null);
await fetchProfiles(); await fetchProfiles();
} catch (e: any) { } catch (e: unknown) {
toast.error("Failed to delete: " + (e.message || String(e))); toast.error("Failed to delete: " + errMsg(e));
} finally { } finally {
setIsDeleting(false); setIsDeleting(false);
} }
} }
async function handleRename(e: React.FormEvent, oldName: string) { async function handleRename(e: React.SubmitEvent, oldName: string) {
e.preventDefault(); e.preventDefault();
if (!renameValue.trim() || renameValue.trim() === oldName) { if (!renameValue.trim() || renameValue.trim() === oldName) {
setRenamingName(null); setRenamingName(null);
@@ -144,8 +153,8 @@ export function ProfilePanel({
if (autoLoadProfile === oldName) setAutoLoadProfile(renameValue.trim()); if (autoLoadProfile === oldName) setAutoLoadProfile(renameValue.trim());
setRenamingName(null); setRenamingName(null);
await fetchProfiles(); await fetchProfiles();
} catch (e: any) { } catch (e: unknown) {
toast.error("Failed to rename: " + (e.message || String(e))); toast.error("Failed to rename: " + errMsg(e));
} finally { } finally {
setIsRenaming(false); setIsRenaming(false);
} }
@@ -179,7 +188,7 @@ export function ProfilePanel({
value={newName} value={newName}
onChange={(e) => setNewName(e.target.value)} onChange={(e) => setNewName(e.target.value)}
disabled={isSaving} disabled={isSaving}
onKeyDown={(e) => e.key === "Escape" && setIsSaveOpen(false)} onKeyDown={(e) => e.key === "Escape" && closeSaveForm()}
className="flex-1 min-w-0 bg-zinc-950 border border-zinc-700 rounded px-3 py-1.5 text-sm focus:outline-none focus:border-pink-500 focus:ring-1 focus:ring-pink-500 disabled:opacity-50" className="flex-1 min-w-0 bg-zinc-950 border border-zinc-700 rounded px-3 py-1.5 text-sm focus:outline-none focus:border-pink-500 focus:ring-1 focus:ring-pink-500 disabled:opacity-50"
/> />
<button <button
@@ -191,7 +200,7 @@ export function ProfilePanel({
</button> </button>
<button <button
type="button" type="button"
onClick={() => setIsSaveOpen(false)} onClick={closeSaveForm}
className="px-3 py-1.5 text-zinc-400 hover:text-zinc-200 rounded text-sm transition shrink-0" className="px-3 py-1.5 text-zinc-400 hover:text-zinc-200 rounded text-sm transition shrink-0"
> >
Cancel Cancel
@@ -336,6 +345,11 @@ export function ProfilePanel({
{badges && ( {badges && (
<p className="text-xs text-zinc-500">{badges}</p> <p className="text-xs text-zinc-500">{badges}</p>
)} )}
{p.power_cap_mode === "ioctl" && (
<p className="text-xs text-red-400 font-medium">
⚠ experimental power (below VBIOS min)
</p>
)}
</div> </div>
</div> </div>
@@ -0,0 +1,315 @@
import { GaugeCard } from "../Monitor/GaugeCard.js";
import { fmt } from "../../utils/units.js";
import type { WireViewInfo, WireViewSample } from "../../types.js";
interface Props {
info: WireViewInfo | null;
sample: WireViewSample | null;
history: WireViewSample[];
}
// Per-pin current rating of the 12 VHPWR connector (A). The bar's
// full scale is 10 A (rating + headroom); red at/above the 9.2 A rating.
const PIN_RATED_A = 9.2;
const PIN_BAR_MAX_A = 10;
const FAULT_NAMES: Record<number, string> = {
0: "Chip over-temperature",
1: "Sensor over-temperature",
2: "Over-current (OCP)",
3: "Wire over-current",
4: "Over-power (OPP)",
5: "Current imbalance",
};
function decodeFaults(mask: number): string[] {
return Object.entries(FAULT_NAMES)
.filter(([bit]) => mask & (1 << Number(bit)))
.map(([, name]) => name);
}
function pluck(history: WireViewSample[], key: keyof WireViewSample): number[] {
return history.map((s) => (s[key] as number | null) ?? 0);
}
function FaultBadge({
label,
mask,
}: {
label: string;
mask: number;
}) {
const faults = decodeFaults(mask);
return (
<div
className={`rounded-lg p-3 border flex flex-col gap-1.5 min-w-0 ${
faults.length > 0
? "bg-red-500/10 border-red-500/40"
: "bg-zinc-900 border-zinc-800"
}`}
>
<div className="flex items-center justify-between gap-2">
<span className="text-xs text-zinc-500 uppercase tracking-wider">
{label}
</span>
<span
className={`text-xs font-mono ${
faults.length > 0 ? "text-red-400" : "text-zinc-600"
}`}
>
0x{mask.toString(16).toUpperCase().padStart(4, "0")}
</span>
</div>
{faults.length > 0 ? (
<div className="flex flex-wrap gap-1">
{faults.map((f) => (
<span
key={f}
className="px-1.5 py-0.5 rounded bg-red-500/20 border border-red-500/40 text-red-300 text-[11px] font-medium"
>
{f}
</span>
))}
</div>
) : (
<div className="text-sm text-emerald-400 font-medium">No faults</div>
)}
</div>
);
}
function PinCard({
index,
voltage,
current,
power,
}: {
index: number;
voltage: number;
current: number;
power: number;
}) {
const loadPct = Math.min(100, (current / PIN_BAR_MAX_A) * 100);
const barColor =
current >= PIN_RATED_A
? "#f87171"
: loadPct >= 75
? "#fbbf24"
: "#34d399";
return (
<div className="bg-zinc-900 rounded-lg p-3 flex flex-col gap-1.5 min-w-0">
<div className="flex items-center justify-between">
<span className="text-xs text-zinc-500 uppercase tracking-wider">
Pin {index + 1}
</span>
<span className="text-[10px] font-mono text-zinc-600">
{current.toFixed(1)} / {PIN_RATED_A} A
</span>
</div>
<div className="grid grid-cols-3 gap-1 text-center">
<div>
<div className="text-[10px] text-zinc-600">V</div>
<div className="text-sm font-mono font-semibold text-violet-300">
{voltage.toFixed(2)}
</div>
</div>
<div>
<div className="text-[10px] text-zinc-600">A</div>
<div className="text-sm font-mono font-semibold text-cyan-300">
{current.toFixed(2)}
</div>
</div>
<div>
<div className="text-[10px] text-zinc-600">W</div>
<div className="text-sm font-mono font-semibold text-pink-300">
{power.toFixed(1)}
</div>
</div>
</div>
<div className="h-1.5 rounded-full bg-zinc-800 overflow-hidden">
<div
className="h-full rounded-full transition-all duration-500"
style={{ width: `${loadPct}%`, backgroundColor: barColor }}
/>
</div>
</div>
);
}
function InfoItem({
label,
value,
}: {
label: string;
value: string | null | undefined;
}) {
if (!value) return null;
return (
<div className="flex flex-col gap-0.5 min-w-0">
<span className="text-xs text-zinc-500">{label}</span>
<span
className="text-sm font-mono text-zinc-200 truncate"
title={value}
>
{value}
</span>
</div>
);
}
export function WireViewPanel({ info, sample, history }: Props) {
if (!sample) {
return (
<div className="bg-zinc-900 rounded-lg p-10 text-center text-zinc-500 flex flex-col items-center gap-2">
<div className="text-lg font-medium text-zinc-400">
WireView Pro II not connected
</div>
<div className="text-sm">
Plug in the device — the tab appears automatically once it is
detected.
</div>
</div>
);
}
const psuCap = sample.psu_capability_w > 0 ? sample.psu_capability_w : 600;
return (
<div className="flex flex-col gap-4 w-full">
{/* ── Header ─────────────────────────────────────────────────── */}
<div className="bg-zinc-900 rounded-lg p-4 border border-zinc-800 flex flex-wrap items-center gap-x-6 gap-y-2">
<div className="flex items-center gap-2">
<span className="w-2 h-2 rounded-full bg-emerald-400" />
<span className="text-lg font-semibold text-zinc-100">
{info?.device_name ?? "WireView Pro II"}
</span>
</div>
<span className="text-sm font-mono text-zinc-400">
{info?.hw_rev ? `HW ${info.hw_rev}` : ""}
{info?.firmware_version ? ` · FW ${info.firmware_version}` : ""}
</span>
<span className="text-sm font-mono text-zinc-400">
PSU {psuCap} W
</span>
<span className="text-xs font-mono text-zinc-600 ml-auto">
{info?.transport === "hwmon"
? `hwmon: ${info.port}`
: `serial: ${info?.port ?? ""}`}
</span>
</div>
{/* ── Main gauges ────────────────────────────────────────────── */}
<div className="bg-zinc-900 rounded-lg p-4">
<div className="text-xs text-zinc-500 uppercase tracking-wider font-semibold pb-3 border-b border-zinc-800">
12 VHPWR Connector
</div>
<div className="grid grid-cols-2 md:grid-cols-4 gap-3 mt-3">
<GaugeCard
label="Total Power"
value={fmt.watts(sample.power_total_w, 1)}
history={pluck(history, "power_total_w")}
color="#f472b6"
max={psuCap}
/>
<GaugeCard
label="Total Current"
value={`${sample.current_total_a.toFixed(2)} A`}
history={pluck(history, "current_total_a")}
color="#38bdf8"
max={(psuCap / 12) * 1.1}
/>
<GaugeCard
label="Avg Voltage"
value={`${sample.voltage_avg_v.toFixed(3)} V`}
history={pluck(history, "voltage_avg_v")}
color="#a78bfa"
max={13}
/>
<GaugeCard
label="Fan Duty"
value={fmt.pct(sample.fan_duty_pct)}
history={pluck(history, "fan_duty_pct")}
color="#fb923c"
max={100}
/>
</div>
</div>
{/* ── Per-pin ────────────────────────────────────────────────── */}
<div className="bg-zinc-900 rounded-lg p-4">
<div className="text-xs text-zinc-500 uppercase tracking-wider font-semibold pb-3 border-b border-zinc-800">
Per-Pin Readings
</div>
<div className="grid grid-cols-2 md:grid-cols-3 lg:grid-cols-6 gap-3 mt-3">
{sample.pins.map((pin, i) => (
<PinCard
key={i}
index={i}
voltage={pin.voltage_v}
current={pin.current_a}
power={pin.power_w}
/>
))}
</div>
</div>
{/* ── Temperatures ───────────────────────────────────────────── */}
<div className="bg-zinc-900 rounded-lg p-4">
<div className="text-xs text-zinc-500 uppercase tracking-wider font-semibold pb-3 border-b border-zinc-800">
Temperatures
</div>
<div className="grid grid-cols-2 md:grid-cols-4 gap-3 mt-3">
<GaugeCard
label="Temp In"
value={fmt.celsius(sample.temp_in_c, 1)}
history={pluck(history, "temp_in_c")}
color="#f87171"
max={90}
/>
<GaugeCard
label="Temp Out"
value={fmt.celsius(sample.temp_out_c, 1)}
history={pluck(history, "temp_out_c")}
color="#fb923c"
max={90}
/>
<GaugeCard
label="Ext Sensor 1"
value={fmt.celsius(sample.temp_ext1_c, 1)}
history={pluck(history, "temp_ext1_c")}
color="#fbbf24"
max={90}
/>
<GaugeCard
label="Ext Sensor 2"
value={fmt.celsius(sample.temp_ext2_c, 1)}
history={pluck(history, "temp_ext2_c")}
color="#a78bfa"
max={90}
/>
</div>
</div>
{/* ── Faults + device info ───────────────────────────────────── */}
<div className="grid grid-cols-1 md:grid-cols-2 gap-4">
<div className="flex flex-col gap-3">
<FaultBadge label="Fault Status" mask={sample.fault_status} />
<FaultBadge label="Fault Log (latched)" mask={sample.fault_log} />
</div>
<div className="bg-zinc-900 rounded-lg p-4">
<div className="text-xs text-zinc-500 uppercase tracking-wider font-semibold pb-3 border-b border-zinc-800">
Device
</div>
<div className="grid grid-cols-2 gap-x-6 gap-y-3 mt-3">
<InfoItem label="Hardware Revision" value={info?.hw_rev} />
<InfoItem label="Firmware" value={info?.firmware_version} />
<InfoItem label="Build" value={info?.build} />
<InfoItem label="UID" value={info?.uid} />
<InfoItem label="Transport" value={info?.transport} />
<InfoItem label="Port" value={info?.port} />
</div>
</div>
</div>
</div>
);
}
@@ -1,25 +1,37 @@
export interface DialogAction {
label: string;
onClick: () => void;
className?: string;
disabled?: boolean;
}
interface Props { interface Props {
message: string; message: string;
detail?: string; detail?: string;
confirmLabel?: string; confirmLabel?: string;
isDestructive?: boolean; isDestructive?: boolean;
onConfirm: () => void; /** Required unless `actions` is provided (which replaces the single confirm button). */
onConfirm?: () => void;
onCancel: () => void; onCancel: () => void;
/** When provided, rendered after Cancel instead of the single confirm button. */
actions?: DialogAction[];
} }
export function ConfirmDialog({ export function ConfirmDialog({
message, message,
detail, detail,
confirmLabel = 'Confirm', confirmLabel = "Confirm",
isDestructive = false, isDestructive = false,
onConfirm, onConfirm,
onCancel, onCancel,
actions,
}: Props) { }: Props) {
return ( return (
<div <div
className="fixed inset-0 z-50 flex items-center justify-center" className="fixed inset-0 z-50 flex items-center justify-center bg-black/70"
style={{ background: 'rgba(0,0,0,0.7)' }} onMouseDown={(e) => {
onMouseDown={(e) => { if (e.target === e.currentTarget) onCancel(); }} if (e.target === e.currentTarget) onCancel();
}}
> >
<div className="bg-zinc-900 border border-zinc-700 rounded-xl shadow-2xl p-6 w-96 max-w-[90vw]"> <div className="bg-zinc-900 border border-zinc-700 rounded-xl shadow-2xl p-6 w-96 max-w-[90vw]">
<h2 className="text-zinc-100 font-semibold text-sm mb-1">{message}</h2> <h2 className="text-zinc-100 font-semibold text-sm mb-1">{message}</h2>
@@ -32,17 +44,34 @@ export function ConfirmDialog({
> >
Cancel Cancel
</button> </button>
<button {actions ? (
onClick={onConfirm} actions.map((a) => (
className={[ <button
'px-3 py-1.5 rounded text-xs font-semibold transition-colors', key={a.label}
isDestructive onClick={a.onClick}
? 'bg-red-600 hover:bg-red-500 text-white' disabled={a.disabled}
: 'bg-emerald-600 hover:bg-emerald-500 text-white', className={[
].join(' ')} a.className ??
> "px-3 py-1.5 rounded bg-zinc-800 hover:bg-zinc-700 text-zinc-200 text-xs font-semibold transition-colors",
{confirmLabel} a.disabled ? "opacity-40 cursor-not-allowed" : "",
</button> ].join(" ")}
>
{a.label}
</button>
))
) : (
<button
onClick={() => onConfirm?.()}
className={[
"px-3 py-1.5 rounded text-xs font-semibold transition-colors",
isDestructive
? "bg-red-600 hover:bg-red-500 text-white"
: "bg-emerald-600 hover:bg-emerald-500 text-white",
].join(" ")}
>
{confirmLabel}
</button>
)}
</div> </div>
</div> </div>
</div> </div>
@@ -0,0 +1,120 @@
import { useState } from "react";
import { Bug, Loader } from "lucide-react";
import { toast } from "sonner";
import { api } from "../../api/client.js";
interface Props {
onClose: () => void;
}
export function ReportBugDialog({ onClose }: Props) {
const [title, setTitle] = useState("");
const [body, setBody] = useState("");
const [submitting, setSubmitting] = useState(false);
const [done, setDone] = useState<{
number: number | null;
url: string | null;
} | null>(null);
async function handleSubmit() {
if (!title.trim() || submitting) return;
setSubmitting(true);
try {
const res = await api.reportBug(title.trim(), body.trim());
setDone({ number: res.issue_number, url: res.issue_url });
toast.success(
res.issue_number != null ? `Issue #${res.issue_number} created` : "Issue created",
);
} catch (err) {
toast.error(err instanceof Error ? err.message : "Failed to create issue");
} finally {
setSubmitting(false);
}
}
return (
<div
className="fixed inset-0 z-50 flex items-center justify-center bg-black/70"
onMouseDown={(e) => {
if (e.target === e.currentTarget) onClose();
}}
>
<div className="bg-zinc-900 border border-zinc-700 rounded-xl shadow-2xl p-6 w-[480px] max-w-[90vw]">
{done ? (
<>
<h2 className="text-zinc-100 font-semibold text-sm mb-2 flex items-center gap-2">
<Bug size={16} className="text-emerald-400" /> Issue created
</h2>
<p className="text-zinc-400 text-xs mb-4">
{done.number != null ? `Issue #${done.number} ` : "The issue "}
was filed on Gitea with the auto-collected system context.{" "}
{done.url ? (
<a
href={done.url}
target="_blank"
rel="noreferrer"
className="text-pink-400 hover:underline"
>
Open it →
</a>
) : (
"."
)}
</p>
<div className="flex justify-end">
<button
onClick={onClose}
className="px-3 py-1.5 rounded bg-zinc-800 hover:bg-zinc-700 text-zinc-200 text-xs font-semibold transition-colors"
>
Close
</button>
</div>
</>
) : (
<>
<h2 className="text-zinc-100 font-semibold text-sm mb-1 flex items-center gap-2">
<Bug size={16} className="text-pink-400" /> Report a bug
</h2>
<p className="text-zinc-400 text-xs mb-4">
Creates a Gitea issue. System info (version, GPU, active profile,
recent log) is attached automatically.
</p>
<input
autoFocus
value={title}
onChange={(e) => setTitle(e.target.value)}
placeholder="Title / Subject (required)"
maxLength={200}
className="w-full bg-zinc-800 border border-zinc-700 rounded px-3 py-2 text-sm text-zinc-100 placeholder:text-zinc-500 focus:outline-none focus:border-pink-500 mb-3"
/>
<textarea
value={body}
onChange={(e) => setBody(e.target.value)}
placeholder="What happened? Steps to reproduce, expected vs. actual…"
rows={6}
maxLength={5000}
className="w-full bg-zinc-800 border border-zinc-700 rounded px-3 py-2 text-sm text-zinc-100 placeholder:text-zinc-500 focus:outline-none focus:border-pink-500 mb-4 resize-y"
/>
<div className="flex justify-end gap-2">
<button
onClick={onClose}
disabled={submitting}
className="px-3 py-1.5 rounded bg-zinc-800 hover:bg-zinc-700 text-zinc-300 text-xs transition-colors disabled:opacity-40"
>
Cancel
</button>
<button
onClick={handleSubmit}
disabled={!title.trim() || submitting}
className="flex items-center gap-1.5 px-3 py-1.5 rounded bg-pink-600 hover:bg-pink-500 text-white text-xs font-semibold transition-colors disabled:opacity-40 disabled:cursor-not-allowed"
>
{submitting && <Loader size={12} className="animate-spin" />}
Submit
</button>
</div>
</>
)}
</div>
</div>
);
}
+9 -7
View File
@@ -1,12 +1,14 @@
import { useEffect, useRef, useState } from 'react'; import { useEffect, useRef, useState } from "react";
import { api } from '../api/client'; import { api } from "../api/client.js";
import { createWsConnection } from '../api/websocket'; import { createWsConnection } from "../api/websocket.js";
import { useCurveStore } from '../store/curveStore'; import { useCurveStore } from "../store/curveStore.js";
import type { CurveState } from '../types'; import type { CurveState } from "../types.js";
export function useCurve() { export function useCurve() {
const { curve, setCurve, selectedGpuIndex } = useCurveStore(); const { curve, setCurve, selectedGpuIndex } = useCurveStore();
const [wsStatus, setWsStatus] = useState<'connecting' | 'connected' | 'disconnected'>('connecting'); const [wsStatus, setWsStatus] = useState<
"connecting" | "connected" | "disconnected"
>("connecting");
const wsRef = useRef<ReturnType<typeof createWsConnection> | null>(null); const wsRef = useRef<ReturnType<typeof createWsConnection> | null>(null);
useEffect(() => { useEffect(() => {
@@ -15,7 +17,7 @@ export function useCurve() {
// Subscribe to /ws/curve for push updates after writes // Subscribe to /ws/curve for push updates after writes
wsRef.current = createWsConnection<CurveState>( wsRef.current = createWsConnection<CurveState>(
'/ws/curve', "/ws/curve",
(data) => setCurve(data), (data) => setCurve(data),
setWsStatus, setWsStatus,
selectedGpuIndex, selectedGpuIndex,
+9 -9
View File
@@ -1,7 +1,7 @@
import { useEffect, useState } from "react"; import { useEffect, useState } from "react";
import { api } from "../api/client"; import { api } from "../api/client.js";
import { useCurveStore } from "../store/curveStore"; import { useCurveStore } from "../store/curveStore.js";
import type { DashboardInfo } from "../types"; import type { DashboardInfo } from "../types.js";
interface DashboardState { interface DashboardState {
gpuIndex: number; gpuIndex: number;
@@ -23,17 +23,17 @@ export function useDashboard() {
useEffect(() => { useEffect(() => {
let cancelled = false; let cancelled = false;
api (async () => {
.dashboard(selectedGpuIndex) try {
.then((data) => { const data = await api.dashboard(selectedGpuIndex);
if (!cancelled) if (!cancelled)
setState({ gpuIndex: selectedGpuIndex, data, done: true }); setState({ gpuIndex: selectedGpuIndex, data, done: true });
}) } catch (err) {
.catch((err) => {
console.error("Failed to load dashboard info:", err); console.error("Failed to load dashboard info:", err);
if (!cancelled) if (!cancelled)
setState({ gpuIndex: selectedGpuIndex, data: null, done: true }); setState({ gpuIndex: selectedGpuIndex, data: null, done: true });
}); }
})();
return () => { return () => {
cancelled = true; cancelled = true;
}; };
+5 -4
View File
@@ -1,9 +1,10 @@
import { useEffect } from 'react'; import { useEffect } from "react";
import { api } from '../api/client'; import { api } from "../api/client.js";
import { useCurveStore } from '../store/curveStore'; import { useCurveStore } from "../store/curveStore.js";
export function useGpu() { export function useGpu() {
const { gpuInfo, setGpuInfo, setAvailableGpus, selectedGpuIndex } = useCurveStore(); const { gpuInfo, setGpuInfo, setAvailableGpus, selectedGpuIndex } =
useCurveStore();
useEffect(() => { useEffect(() => {
api.gpus().then(setAvailableGpus).catch(console.error); api.gpus().then(setAvailableGpus).catch(console.error);
+10 -7
View File
@@ -1,16 +1,19 @@
import { useEffect, useRef, useState } from 'react'; import { useEffect, useRef, useState } from "react";
import { createWsConnection } from '../api/websocket'; import { createWsConnection } from "../api/websocket.js";
import { useCurveStore } from '../store/curveStore'; import { useCurveStore } from "../store/curveStore.js";
import type { MonitoringSample } from '../types'; import type { MonitoringSample } from "../types.js";
export function useMonitor() { export function useMonitor() {
const { monitor, monitorHistory, pushMonitor, selectedGpuIndex } = useCurveStore(); const { monitor, monitorHistory, pushMonitor, selectedGpuIndex } =
const [wsStatus, setWsStatus] = useState<'connecting' | 'connected' | 'disconnected'>('connecting'); useCurveStore();
const [wsStatus, setWsStatus] = useState<
"connecting" | "connected" | "disconnected"
>("connecting");
const wsRef = useRef<ReturnType<typeof createWsConnection> | null>(null); const wsRef = useRef<ReturnType<typeof createWsConnection> | null>(null);
useEffect(() => { useEffect(() => {
wsRef.current = createWsConnection<MonitoringSample>( wsRef.current = createWsConnection<MonitoringSample>(
'/ws/monitor', "/ws/monitor",
pushMonitor, pushMonitor,
setWsStatus, setWsStatus,
selectedGpuIndex, selectedGpuIndex,
+40
View File
@@ -0,0 +1,40 @@
import { useEffect, useRef, useState } from "react";
import { createWsConnection } from "../api/websocket.js";
import type {
WireViewInfo,
WireViewSample,
WireViewWsMessage,
} from "../types.js";
const MAX_HISTORY = 120; // ~2 min at 1 s polling
export function useWireview() {
const [available, setAvailable] = useState(false);
const [info, setInfo] = useState<WireViewInfo | null>(null);
const [sample, setSample] = useState<WireViewSample | null>(null);
const [history, setHistory] = useState<WireViewSample[]>([]);
const wsRef = useRef<ReturnType<typeof createWsConnection> | null>(null);
useEffect(() => {
wsRef.current = createWsConnection<WireViewWsMessage>(
"/ws/wireview",
(msg) => {
if (msg.type === "sample") {
setAvailable(true);
setInfo(msg.info);
setSample(msg.sample);
setHistory((h) => [...h.slice(-(MAX_HISTORY - 1)), msg.sample]);
} else {
setAvailable(false);
setInfo(null);
setSample(null);
setHistory([]);
}
},
() => {},
);
return () => wsRef.current?.close();
}, []);
return { available, info, sample, history };
}
@@ -0,0 +1,129 @@
import { useState } from "react";
import { renderHook, act } from "@testing-library/react";
import { describe, it, expect } from "vitest";
import { useWireviewImbalance } from "./useWireviewImbalance.js";
import type { WireViewSample } from "../types.js";
/** Build a minimal sample with the given per-pin currents (A). */
function sample(pinCurrents: number[]): WireViewSample {
return {
timestamp: Date.now(),
power_total_w: 0,
current_total_a: pinCurrents.reduce((a, b) => a + b, 0),
pin_imbalance_a: 0,
voltage_avg_v: 12,
pins: pinCurrents.map((a) => ({
voltage_v: 12,
current_a: a,
power_w: 12 * a,
})),
temp_in_c: null,
temp_out_c: null,
temp_ext1_c: null,
temp_ext2_c: null,
fan_duty_pct: 0,
fault_status: 0,
fault_log: 0,
psu_capability_w: 600,
};
}
const BALANCED = [4.5, 4.6, 4.4, 4.8, 5.0, 4.2]; // spread 0.8 A
const IMBALANCED = [4.5, 4.6, 4.4, 4.8, 5.6, 3.5]; // spread 2.1 A
function useTestHook() {
const [s, setS] = useState<WireViewSample | null>(null);
const alert = useWireviewImbalance(s);
return { alert, setSample: setS };
}
/**
* Feed `times` fresh samples (new object identity, like the websocket).
* Each sample must be a separate state update — the hook counts per sample,
* so the loop is intentional (not batchable).
*/
function feed(
setSample: (s: WireViewSample) => void,
currents: number[],
times: number,
) {
for (let i = 0; i < times; i++) {
act(() => setSample(sample(currents)));
}
}
describe("useWireviewImbalance", () => {
it("returns null for a null sample", () => {
const { result } = renderHook(() => useTestHook());
expect(result.current.alert).toBeNull();
});
it("stays hidden while pins are balanced", () => {
const { result } = renderHook(() => useTestHook());
feed(result.current.setSample, BALANCED, 10);
expect(result.current.alert).toBeNull();
});
it("shows only after 7 consecutive imbalanced samples", () => {
const { result } = renderHook(() => useTestHook());
feed(result.current.setSample, IMBALANCED, 6);
expect(result.current.alert).toBeNull();
feed(result.current.setSample, IMBALANCED, 1);
expect(result.current.alert).not.toBeNull();
});
it("reports the correct pins and spread", () => {
const { result } = renderHook(() => useTestHook());
feed(result.current.setSample, IMBALANCED, 7);
const alert = result.current.alert!;
expect(alert.diff).toBeCloseTo(2.1, 5);
expect(alert.hiPin).toBe(5); // 5.6 A
expect(alert.loPin).toBe(6); // 3.5 A
expect(alert.maxI).toBeCloseTo(5.6, 5);
expect(alert.minI).toBeCloseTo(3.5, 5);
});
it("a single balanced sample in the streak resets the counter", () => {
const { result } = renderHook(() => useTestHook());
feed(result.current.setSample, IMBALANCED, 6);
feed(result.current.setSample, BALANCED, 1);
feed(result.current.setSample, IMBALANCED, 6);
expect(result.current.alert).toBeNull();
feed(result.current.setSample, IMBALANCED, 1);
expect(result.current.alert).not.toBeNull();
});
it("keeps showing for 5 samples after the imbalance clears, then hides", () => {
const { result } = renderHook(() => useTestHook());
feed(result.current.setSample, IMBALANCED, 7);
expect(result.current.alert).not.toBeNull();
feed(result.current.setSample, BALANCED, 4);
expect(result.current.alert).not.toBeNull(); // grace period
feed(result.current.setSample, BALANCED, 1);
expect(result.current.alert).toBeNull();
});
it("stays shown (with updated values) while the imbalance persists", () => {
const { result } = renderHook(() => useTestHook());
feed(result.current.setSample, IMBALANCED, 7);
feed(result.current.setSample, [4.5, 4.6, 4.4, 4.8, 6.1, 3.2], 1);
const alert = result.current.alert!;
expect(alert.diff).toBeCloseTo(2.9, 5);
expect(alert.hiPin).toBe(5);
expect(alert.loPin).toBe(6);
});
it("clears immediately when the sample disappears (device unplugged)", () => {
const { result } = renderHook(() => useTestHook());
feed(result.current.setSample, IMBALANCED, 7);
expect(result.current.alert).not.toBeNull();
act(() => result.current.setSample(null));
expect(result.current.alert).toBeNull();
});
it("ignores samples with fewer than 2 pins", () => {
const { result } = renderHook(() => useTestHook());
feed(result.current.setSample, [10], 5);
expect(result.current.alert).toBeNull();
});
});
@@ -0,0 +1,78 @@
import { useEffect, useRef, useState } from "react";
import type { WireViewSample } from "../types.js";
// Current difference between any two pins that indicates a bad contact.
const PIN_IMBALANCE_A = 2;
// Show only after the imbalance holds for this many consecutive samples (~1 s each).
// AI workloads produce frequent short load ramps that skew one pin transiently,
// so the threshold is deliberately high: a bad contact stays imbalanced, a ramp does not.
const SHOW_AFTER = 7;
// Keep showing this many samples after the imbalance clears (grace period).
const HIDE_AFTER = 5;
export interface ImbalanceAlert {
diff: number;
maxI: number;
minI: number;
/** 1-based pin index carrying the most current */
hiPin: number;
/** 1-based pin index carrying the least current */
loPin: number;
}
/**
* Debounced pin-current imbalance detection.
*
* A single sample can transiently skew (load change, measurement glitch), so
* the alert only appears after the imbalance persists for a few samples, and
* stays visible briefly after it clears instead of flickering.
*/
export function useWireviewImbalance(
sample: WireViewSample | null,
): ImbalanceAlert | null {
const [alert, setAlert] = useState<ImbalanceAlert | null>(null);
const alertRef = useRef<ImbalanceAlert | null>(null);
const streak = useRef(0);
const grace = useRef(0);
useEffect(() => {
function update(next: ImbalanceAlert | null) {
alertRef.current = next;
setAlert(next);
}
if (!sample || sample.pins.length < 2) {
streak.current = 0;
grace.current = 0;
if (alertRef.current) update(null);
return;
}
const currents = sample.pins.map((p) => p.current_a);
const maxI = Math.max(...currents);
const minI = Math.min(...currents);
const diff = maxI - minI;
if (diff > PIN_IMBALANCE_A) {
streak.current += 1;
grace.current = 0;
if (streak.current >= SHOW_AFTER) {
update({
diff,
maxI,
minI,
hiPin: currents.indexOf(maxI) + 1,
loPin: currents.indexOf(minI) + 1,
});
}
} else {
streak.current = 0;
if (alertRef.current) {
grace.current += 1;
if (grace.current >= HIDE_AFTER) update(null);
}
}
}, [sample]);
return alert;
}
+95 -45
View File
@@ -1,7 +1,12 @@
import { create } from 'zustand'; import { create } from "zustand";
import type { CurveState, GpuInfo, MonitoringSample, VFPoint } from '../types'; import type {
import { api } from '../api/client'; CurveState,
import { toast } from 'sonner'; GpuInfo,
MonitoringSample,
VFPoint,
} from "../types.js";
import { api } from "../api/client.js";
import { toast } from "sonner";
const HISTORY_SIZE = 120; // ~60s at 2Hz const HISTORY_SIZE = 120; // ~60s at 2Hz
@@ -50,7 +55,9 @@ interface CurveStore {
resetAllDeltas: (onSuccess: () => void) => Promise<void>; resetAllDeltas: (onSuccess: () => void) => Promise<void>;
// Selection actions // Selection actions
selectPoint: (index: number, multi?: boolean) => void; selectPoint: (index: number) => void;
/** Toggle a point in the selection (Shift/Ctrl+click); updates the anchor. */
togglePoint: (index: number) => void;
selectRange: (indices: number[]) => void; selectRange: (indices: number[]) => void;
clearSelection: () => void; clearSelection: () => void;
/** /**
@@ -80,7 +87,16 @@ export const useCurveStore = create<CurveStore>()((set, get) => ({
setAvailableGpus: (availableGpus) => set({ availableGpus }), setAvailableGpus: (availableGpus) => set({ availableGpus }),
setSelectedGpuIndex: (selectedGpuIndex) => { setSelectedGpuIndex: (selectedGpuIndex) => {
set({ selectedGpuIndex, curve: null, gpuInfo: null, monitor: null, monitorHistory: [], pendingDeltas: new Map(), selectedPoints: new Set(), anchorPoint: null }); set({
selectedGpuIndex,
curve: null,
gpuInfo: null,
monitor: null,
monitorHistory: [],
pendingDeltas: new Map(),
selectedPoints: new Set(),
anchorPoint: null,
});
}, },
setCurve: (curve) => set({ curve }), setCurve: (curve) => set({ curve }),
setGpuInfo: (gpuInfo) => set({ gpuInfo }), setGpuInfo: (gpuInfo) => set({ gpuInfo }),
@@ -95,7 +111,7 @@ export const useCurveStore = create<CurveStore>()((set, get) => ({
stageEdit: (pointIndex, deltaKhz) => stageEdit: (pointIndex, deltaKhz) =>
set((s) => { set((s) => {
const next = new Map(s.pendingDeltas); const next = new Map(s.pendingDeltas);
const point = s.curve?.points.find(p => p.index === pointIndex); const point = s.curve?.points.find((p) => p.index === pointIndex);
if (point && point.delta_khz === deltaKhz) { if (point && point.delta_khz === deltaKhz) {
next.delete(pointIndex); next.delete(pointIndex);
} else { } else {
@@ -108,7 +124,7 @@ export const useCurveStore = create<CurveStore>()((set, get) => ({
set((s) => { set((s) => {
const next = new Map(s.pendingDeltas); const next = new Map(s.pendingDeltas);
edits.forEach((deltaKhz, index) => { edits.forEach((deltaKhz, index) => {
const point = s.curve?.points.find(p => p.index === index); const point = s.curve?.points.find((p) => p.index === index);
if (point && point.delta_khz === deltaKhz) { if (point && point.delta_khz === deltaKhz) {
next.delete(index); next.delete(index);
} else { } else {
@@ -128,27 +144,39 @@ export const useCurveStore = create<CurveStore>()((set, get) => ({
}), }),
discardEdits: () => discardEdits: () =>
set({ pendingDeltas: new Map(), selectedPoints: new Set(), anchorPoint: null }), set({
pendingDeltas: new Map(),
selectedPoints: new Set(),
anchorPoint: null,
}),
applyEdits: async (onSuccess) => { applyEdits: async (onSuccess) => {
const { pendingDeltas, selectedGpuIndex } = get(); const { pendingDeltas, selectedGpuIndex } = get();
if (pendingDeltas.size === 0) return; if (pendingDeltas.size === 0) return;
// Convert Map to plain record for the API // Convert Map to plain record for the API
const deltas: Record<number, number> = {}; const deltas: Record<number, number> = Object.fromEntries(pendingDeltas);
pendingDeltas.forEach((v, k) => { deltas[k] = v; });
try { try {
const result = await api.writeDeltas(deltas, selectedGpuIndex); const result = await api.writeDeltas(deltas, selectedGpuIndex);
set({ pendingDeltas: new Map(), selectedPoints: new Set(), activeProfile: null }); set({
pendingDeltas: new Map(),
selectedPoints: new Set(),
activeProfile: null,
});
if (result?.freq_warnings?.length) { if (result?.freq_warnings?.length) {
toast.warning('Curve applied — driver clamped some points to 0 MHz (negative freq delta)'); toast.warning(
"Curve applied — driver clamped some points to 0 MHz (negative freq delta)",
);
} else { } else {
toast.success('Curve applied successfully'); toast.success("Curve applied successfully");
} }
onSuccess(); onSuccess();
} catch (e: any) { } catch (e: unknown) {
toast.error('Failed to apply curve: ' + (e.message || String(e))); toast.error(
"Failed to apply curve: " +
(e instanceof Error ? e.message : String(e)),
);
} }
}, },
@@ -156,35 +184,48 @@ export const useCurveStore = create<CurveStore>()((set, get) => ({
const { selectedGpuIndex } = get(); const { selectedGpuIndex } = get();
try { try {
await api.resetCurve(selectedGpuIndex); await api.resetCurve(selectedGpuIndex);
set({ pendingDeltas: new Map(), selectedPoints: new Set(), activeProfile: null }); set({
toast.success('Curve reset to hardware defaults'); pendingDeltas: new Map(),
selectedPoints: new Set(),
activeProfile: null,
});
toast.success("Curve reset to hardware defaults");
onSuccess(); onSuccess();
} catch (e: any) { } catch (e: unknown) {
toast.error('Failed to reset curve: ' + (e.message || String(e))); toast.error(
"Failed to reset curve: " +
(e instanceof Error ? e.message : String(e)),
);
} }
}, },
selectPoint: (index, multi = false) => selectPoint: (index) =>
set((s) => {
const next = new Set(s.selectedPoints);
let anchor: number | null = index;
if (next.size === 1 && next.has(index)) {
next.clear();
anchor = null;
} else {
next.clear();
next.add(index);
}
return { selectedPoints: next, anchorPoint: anchor };
}),
togglePoint: (index) =>
set((s) => { set((s) => {
const next = new Set(s.selectedPoints); const next = new Set(s.selectedPoints);
let anchor = s.anchorPoint; let anchor = s.anchorPoint;
if (multi) { if (next.has(index)) {
if (next.has(index)) { next.delete(index);
next.delete(index); if (anchor === index) {
if (anchor === index) anchor = next.size > 0 ? [...next].at(-1)! : null; const last = [...next].at(-1);
} else { anchor = last ?? null;
next.add(index);
anchor = index; // last explicitly added point is the new anchor
} }
} else { } else {
if (next.size === 1 && next.has(index)) { next.add(index);
next.clear(); anchor = index; // last explicitly added point is the new anchor
anchor = null;
} else {
next.clear();
next.add(index);
anchor = index;
}
} }
return { selectedPoints: next, anchorPoint: anchor }; return { selectedPoints: next, anchorPoint: anchor };
}), }),
@@ -193,29 +234,38 @@ export const useCurveStore = create<CurveStore>()((set, get) => ({
// Bulk selects don't change the anchor — preserve it if still in the new selection. // Bulk selects don't change the anchor — preserve it if still in the new selection.
set((s) => { set((s) => {
const next = new Set(indices); const next = new Set(indices);
const anchor = s.anchorPoint !== null && next.has(s.anchorPoint) ? s.anchorPoint : null; const anchor =
s.anchorPoint !== null && next.has(s.anchorPoint)
? s.anchorPoint
: null;
return { selectedPoints: next, anchorPoint: anchor }; return { selectedPoints: next, anchorPoint: anchor };
}), }),
clearSelection: () => clearSelection: () => set({ selectedPoints: new Set(), anchorPoint: null }),
set({ selectedPoints: new Set(), anchorPoint: null }),
flattenToAnchor: () => { flattenToAnchor: () => {
const { selectedPoints, anchorPoint, pendingDeltas, curve, stageMultiEdit } = get(); const {
selectedPoints,
anchorPoint,
pendingDeltas,
curve,
stageMultiEdit,
} = get();
if (selectedPoints.size < 2 || !curve) return; if (selectedPoints.size < 2 || !curve) return;
const anchor = anchorPoint !== null && selectedPoints.has(anchorPoint) const anchor =
? anchorPoint anchorPoint !== null && selectedPoints.has(anchorPoint)
: Math.min(...selectedPoints); ? anchorPoint
: Math.min(...selectedPoints);
const anchorPt = curve.points.find(p => p.index === anchor); const anchorPt = curve.points.find((p) => p.index === anchor);
if (!anchorPt) return; if (!anchorPt) return;
const anchorPendingDelta = pendingDeltas.get(anchor) ?? anchorPt.delta_khz; const anchorPendingDelta = pendingDeltas.get(anchor) ?? anchorPt.delta_khz;
const anchorEffectiveKhz = anchorPt.freq_khz + anchorPendingDelta; const anchorEffectiveKhz = anchorPt.freq_khz + anchorPendingDelta;
const edits = new Map<number, number>(); const edits = new Map<number, number>();
for (const idx of selectedPoints) { for (const idx of selectedPoints) {
const pt = curve.points.find(p => p.index === idx); const pt = curve.points.find((p) => p.index === idx);
if (!pt) continue; if (!pt) continue;
edits.set(idx, anchorEffectiveKhz - pt.freq_khz); edits.set(idx, anchorEffectiveKhz - pt.freq_khz);
} }
+86
View File
@@ -25,6 +25,7 @@ export interface MonitoringSample {
temp_c: number | null; temp_c: number | null;
power_w: number | null; power_w: number | null;
fan_pct: number | null; fan_pct: number | null;
fans: (number | null)[] | null;
pstate: number | null; pstate: number | null;
pstate_label: string | null; pstate_label: string | null;
mem_used_bytes: number | null; mem_used_bytes: number | null;
@@ -84,6 +85,26 @@ export interface GpuInfo {
vram_gib: number | null; vram_gib: number | null;
} }
export interface GpuProcess {
pid: number;
user: string;
dev: number;
type: "C" | "G" | "C+G";
// GPU utilization, only reported for processes active in the last ~1 s.
gpu_util_pct: number | null;
mem_util_pct: number | null;
vram_bytes: number | null;
vram_pct: number | null;
cpu_pct: number | null;
mem_host_bytes: number | null;
command: string;
}
export interface ProcessListResponse {
processes: GpuProcess[];
mem_total_bytes: number | null;
}
export interface SnapshotInfo { export interface SnapshotInfo {
filepath: string; filepath: string;
timestamp: string; timestamp: string;
@@ -97,7 +118,11 @@ export interface LimitsState {
power_limit_w: number | null; power_limit_w: number | null;
default_power_limit_w: number | null; default_power_limit_w: number | null;
min_power_limit_w: number | null; min_power_limit_w: number | null;
min_power_limit_w_native: number | null;
max_power_limit_w: number | null; max_power_limit_w: number | null;
// "nvml" (default) or "ioctl" (experimental RM power control)
power_cap_mode: "nvml" | "ioctl";
rm_power_supported: boolean;
// Clock offsets — current values // Clock offsets — current values
gpc_offset_mhz: number | null; gpc_offset_mhz: number | null;
mem_offset_mhz: number | null; mem_offset_mhz: number | null;
@@ -111,13 +136,22 @@ export interface FanPoint {
fan_pct: number; fan_pct: number;
} }
export interface FanInfo {
index: number;
fan_pct: number | null;
}
export interface FanState { export interface FanState {
fan_pct: number | null; fan_pct: number | null;
fans: FanInfo[] | null;
num_fans: number | null;
fan_mode: "auto" | "curve" | null; fan_mode: "auto" | "curve" | null;
min_fan_pct: number | null; min_fan_pct: number | null;
max_fan_pct: number | null; max_fan_pct: number | null;
curve: FanPoint[] | null; curve: FanPoint[] | null;
curve_active: boolean; curve_active: boolean;
// Fan indices controlled by the active curve; null = all fans.
fan_targets: number[] | null;
} }
export interface ProfileData { export interface ProfileData {
@@ -126,5 +160,57 @@ export interface ProfileData {
curve_deltas: Record<string, number>; curve_deltas: Record<string, number>;
mem_offset_mhz: number | null; mem_offset_mhz: number | null;
power_limit_w: number | null; power_limit_w: number | null;
// "nvml" (default) or "ioctl" (experimental RM power control)
power_cap_mode: "nvml" | "ioctl" | null;
fan_curve: FanPoint[] | null; fan_curve: FanPoint[] | null;
fan_targets: number[] | null;
} }
// ── WireView Pro II (Thermal Grizzly 12 VHPWR connector monitor) ──
export interface WireViewPin {
voltage_v: number;
current_a: number;
power_w: number;
}
export interface WireViewSample {
timestamp: number;
power_total_w: number;
current_total_a: number;
/** Spread (A) between the highest- and lowest-loaded pin */
pin_imbalance_a: number;
voltage_avg_v: number;
pins: WireViewPin[];
// Temperatures are null when the source has no channel for them
// (e.g. an hwmon module without external-sensor channels).
temp_in_c: number | null;
temp_out_c: number | null;
temp_ext1_c: number | null;
temp_ext2_c: number | null;
fan_duty_pct: number;
fault_status: number;
fault_log: number;
psu_capability_w: number;
}
export interface WireViewInfo {
device_name: string;
hw_rev: string;
firmware_version: string;
uid: string;
build: string;
transport: "serial" | "hwmon";
port: string;
}
export interface WireViewState {
available: boolean;
connected: boolean;
info: WireViewInfo | null;
sample: WireViewSample | null;
}
export type WireViewWsMessage =
| { type: "unavailable" }
| { type: "sample"; info: WireViewInfo; sample: WireViewSample };
+31 -29
View File
@@ -1,4 +1,4 @@
import type { VFPoint } from '../types'; import type { VFPoint } from "../types.js";
/** /**
* Approximate reference frequency (MHz) for a point: effective − delta. * Approximate reference frequency (MHz) for a point: effective − delta.
@@ -14,35 +14,37 @@ import type { VFPoint } from '../types';
* useful as a faint reference line ("where this point would sit with no boost"). * useful as a faint reference line ("where this point would sit with no boost").
*/ */
export function refBaseMhz(p: VFPoint): number { export function refBaseMhz(p: VFPoint): number {
return p.freq_mhz - p.delta_mhz; return p.freq_mhz - p.delta_mhz;
} }
/** Find which VF point the GPU is currently near based on voltage reading */ /** Find which VF point the GPU is currently near based on voltage reading */
export function findCurrentPoint( export function findCurrentPoint(
points: VFPoint[], points: VFPoint[],
voltage_mv: number | null, voltage_mv: number | null,
): VFPoint | null { ): VFPoint | null {
if (voltage_mv == null || points.length === 0) return null; if (voltage_mv == null || points.length === 0) return null;
return points.reduce((best, p) => return points.reduce((best, p) =>
Math.abs(p.volt_mv - voltage_mv) < Math.abs(best.volt_mv - voltage_mv) ? p : best, Math.abs(p.volt_mv - voltage_mv) < Math.abs(best.volt_mv - voltage_mv)
); ? p
: best,
);
} }
/** Voltage domain extent, with padding */ /** Voltage domain extent, with padding */
export function voltExtent(points: VFPoint[], padMv = 20): [number, number] { export function voltExtent(points: VFPoint[], padMv = 20): [number, number] {
if (points.length === 0) return [600, 1100]; if (points.length === 0) return [600, 1100];
const min = Math.min(...points.map((p) => p.volt_mv)); const min = Math.min(...points.map((p) => p.volt_mv));
const max = Math.max(...points.map((p) => p.volt_mv)); const max = Math.max(...points.map((p) => p.volt_mv));
return [min - padMv, max + padMv]; return [min - padMv, max + padMv];
} }
/** Frequency domain extent for the effective (boosted) curve, with padding */ /** Frequency domain extent for the effective (boosted) curve, with padding */
export function freqExtent(points: VFPoint[], padMhz = 50): [number, number] { export function freqExtent(points: VFPoint[], padMhz = 50): [number, number] {
if (points.length === 0) return [1000, 3000]; if (points.length === 0) return [1000, 3000];
const allFreqs = points.flatMap((p) => [p.freq_mhz, refBaseMhz(p)]); const allFreqs = points.flatMap((p) => [p.freq_mhz, refBaseMhz(p)]);
const min = Math.min(...allFreqs); const min = Math.min(...allFreqs);
const max = Math.max(...points.map((p) => p.freq_mhz)); const max = Math.max(...points.map((p) => p.freq_mhz));
return [min - padMhz, max + padMhz]; return [min - padMhz, max + padMhz];
} }
/** /**
@@ -59,19 +61,19 @@ export function freqExtent(points: VFPoint[], padMhz = 50): [number, number] {
* point 104 has +315 MHz → effective also 3907 MHz (clamped). * point 104 has +315 MHz → effective also 3907 MHz (clamped).
*/ */
export function detectClampedPoints(points: VFPoint[]): Set<number> { export function detectClampedPoints(points: VFPoint[]): Set<number> {
const clamped = new Set<number>(); const clamped = new Set<number>();
let ceiling = -Infinity; let ceiling = -Infinity;
let ceilingOffset = -Infinity; let ceilingOffset = -Infinity;
for (const p of points) { for (const p of points) {
if (p.freq_mhz <= ceiling && p.delta_khz < ceilingOffset) { if (p.freq_mhz <= ceiling && p.delta_khz < ceilingOffset) {
clamped.add(p.index); clamped.add(p.index);
}
if (p.freq_mhz > ceiling) {
ceiling = p.freq_mhz;
ceilingOffset = p.delta_khz;
}
} }
if (p.freq_mhz > ceiling) {
ceiling = p.freq_mhz;
ceilingOffset = p.delta_khz;
}
}
return clamped; return clamped;
} }
+8
View File
@@ -0,0 +1,8 @@
import { defineConfig } from "vitest/config";
export default defineConfig({
test: {
environment: "happy-dom",
include: ["src/**/*.test.ts"],
},
});
+74
View File
@@ -0,0 +1,74 @@
"""Custom hatchling build hook: build the React frontend if it is missing or stale.
This makes NVCurve installable with a single command, e.g.::
uv tool install "git+https://gitea.zephyre.one/Pakobbix/nvcurve.git"
The hook runs inside the isolated build environment right before the wheel
(or sdist) is assembled. If ``frontend/dist`` does not exist yet — or is older
than the frontend sources — it compiles the frontend using the host's ``npm``
(PATH is inherited from the environment).
"""
from __future__ import annotations
import os
import shutil
import subprocess
import sys
from hatchling.builders.hooks.plugin.interface import BuildHookInterface
# Frontend inputs that must be newer than dist/index.html to trigger a rebuild.
_FRONTEND_INPUTS = (
"src",
"index.html",
"vite.config.ts",
"package.json",
"tsconfig.json",
)
class FrontendBuildHook(BuildHookInterface):
"""Build ``frontend/dist`` with npm when it is missing or stale."""
PLUGIN_NAME = "custom"
def initialize(self, version: str, build_data: dict) -> None:
frontend = os.path.join(self.root, "frontend")
dist_index = os.path.join(frontend, "dist", "index.html")
if not self._needs_build(frontend, dist_index):
return
npm = shutil.which("npm")
if npm is None:
raise RuntimeError(
"npm not found on PATH. Node.js 18+ and npm are required to build "
"the NVCurve frontend. Install them and retry, or use install.sh "
"which checks prerequisites for you."
)
print(
"[nvcurve] frontend/dist missing or stale — building frontend with npm ...",
file=sys.stderr,
)
subprocess.run([npm, "ci", "--no-audit", "--no-fund"], cwd=frontend, check=True)
subprocess.run([npm, "run", "build"], cwd=frontend, check=True)
@staticmethod
def _needs_build(frontend: str, dist_index: str) -> bool:
if not os.path.isfile(dist_index):
return True
dist_mtime = os.path.getmtime(dist_index)
for name in _FRONTEND_INPUTS:
path = os.path.join(frontend, name)
if os.path.isfile(path):
if os.path.getmtime(path) > dist_mtime:
return True
elif os.path.isdir(path):
for root, _dirs, files in os.walk(path):
for file in files:
if os.path.getmtime(os.path.join(root, file)) > dist_mtime:
return True
return False
Executable
+62
View File
@@ -0,0 +1,62 @@
#!/usr/bin/env bash
#
# NVCurve single-command installer.
#
# curl -fsSL https://gitea.zephyre.one/Pakobbix/nvcurve/raw/branch/main/install.sh | bash
#
# or from a local clone:
#
# git clone https://gitea.zephyre.one/Pakobbix/nvcurve.git && cd nvcurve && ./install.sh
#
# The React frontend is compiled automatically during the Python build
# (see hatch_build.py), so Node.js 18+ and npm must be available.
#
# Environment:
# NVCURVE_BRANCH branch to install (default: main)
set -euo pipefail
REPO_URL="https://gitea.zephyre.one/Pakobbix/nvcurve.git"
BRANCH="${NVCURVE_BRANCH:-main}"
fail() {
echo "error: $*" >&2
exit 1
}
# --- prerequisites -----------------------------------------------------------
command -v git >/dev/null 2>&1 ||
fail "git is required. Install it first."
command -v node >/dev/null 2>&1 ||
fail "Node.js 18+ is required (e.g. 'sudo pacman -S nodejs npm' or 'sudo apt install nodejs npm')."
command -v npm >/dev/null 2>&1 ||
fail "npm is required (usually installed together with Node.js)."
if ! command -v uv >/dev/null 2>&1; then
echo "uv not found — installing it from https://astral.sh/uv ..."
curl -LsSf https://astral.sh/uv/install.sh | sh
export PATH="$HOME/.local/bin:$PATH"
command -v uv >/dev/null 2>&1 ||
fail "uv installation failed. Install uv manually: https://docs.astral.sh/uv/"
fi
# --- locate the source tree ---------------------------------------------------
if [ -f pyproject.toml ] && [ -d frontend ]; then
src="$(pwd)"
echo "Installing from current directory: $src"
else
tmp="$(mktemp -d)"
trap 'rm -rf "$tmp"' EXIT
echo "Cloning $REPO_URL (branch: $BRANCH) ..."
git clone --quiet --depth 1 --branch "$BRANCH" "$REPO_URL" "$tmp/nvcurve"
src="$tmp/nvcurve"
fi
# --- install -------------------------------------------------------------------
# The frontend is built automatically by the build hook (hatch_build.py).
uv tool install --force "$src"
echo
echo "NVCurve installed."
echo " Verify your GPU: nvcurve setup"
echo " Start the web UI: nvcurve"
+312 -73
View File
@@ -31,13 +31,14 @@ First-time / diagnostic commands (bypass server, escalate to root):
import argparse import argparse
import json import json
import logging
import os import os
import struct import struct
import sys import sys
import time import time
from .client import ApiError, NvCurveClient, ServerNotRunning from .client import ApiError, NvCurveClient, ServerNotRunning
from .config import Config, default_config from .config import Config, default_config, normalize_trusted_proxies, tls_enabled
from .nvapi.constants import ( from .nvapi.constants import (
CT_BASE, CT_BASE,
CT_POINTS, CT_POINTS,
@@ -48,6 +49,8 @@ from .nvapi.constants import (
VFP_STRIDE, VFP_STRIDE,
) )
log = logging.getLogger("nvcurve.cli")
# ── Utilities ───────────────────────────────────────────────────────────────── # ── Utilities ─────────────────────────────────────────────────────────────────
@@ -69,8 +72,8 @@ def parse_range(s: str):
raise argparse.ArgumentTypeError(f"Expected A-B format, got '{s}'") raise argparse.ArgumentTypeError(f"Expected A-B format, got '{s}'")
try: try:
a, b = int(parts[0]), int(parts[1]) a, b = int(parts[0]), int(parts[1])
except ValueError: except ValueError as exc:
raise argparse.ArgumentTypeError(f"Non-integer in range: '{s}'") raise argparse.ArgumentTypeError(f"Non-integer in range: '{s}'") from exc
if a > b: if a > b:
raise argparse.ArgumentTypeError(f"Start > end in range: {a}-{b}") raise argparse.ArgumentTypeError(f"Start > end in range: {a}-{b}")
if a < 0 or b >= CT_POINTS: if a < 0 or b >= CT_POINTS:
@@ -98,7 +101,7 @@ def print_curve(points, offsets, voltage, domains=None, full=False):
current_idx = None current_idx = None
if voltage: if voltage:
for i, (f, v) in enumerate(points): for i, (_f, v) in enumerate(points):
if v > 0 and abs(v - voltage) < 10000: if v > 0 and abs(v - voltage) < 10000:
current_idx = i current_idx = i
break break
@@ -214,7 +217,7 @@ def print_curve(points, offsets, voltage, domains=None, full=False):
if offsets: if offsets:
nonzero = sum(1 for o in offsets if o != 0) nonzero = sum(1 for o in offsets if o != 0)
if nonzero > 0: if nonzero > 0:
vals = set(o for o in offsets if o != 0) vals = {o for o in offsets if o != 0}
if len(vals) == 1: if len(vals) == 1:
print( print(
f"Global offset: {next(iter(vals)) / 1000:+.0f} MHz " f"Global offset: {next(iter(vals)) / 1000:+.0f} MHz "
@@ -316,7 +319,7 @@ def run_diagnostics(gpu, gpu_name, gpu_index: int = 0):
("SetClockBoostTable", FUNC["SetClockBoostTable"], CT_SIZE, 1, True), ("SetClockBoostTable", FUNC["SetClockBoostTable"], CT_SIZE, 1, True),
] ]
for name, fid, size, ver, needs_mask in probes: for name, fid, size, ver, _needs_mask in probes:
ptr = query_interface(fid) ptr = query_interface(fid)
resolved = "resolved" if ptr else "NOT FOUND" resolved = "resolved" if ptr else "NOT FOUND"
print(f" {name:30s} 0x{fid:08X} size=0x{size:04X} ver={ver} {resolved}") print(f" {name:30s} 0x{fid:08X} size=0x{size:04X} ver={ver} {resolved}")
@@ -402,6 +405,11 @@ def run_diagnostics(gpu, gpu_name, gpu_index: int = 0):
print(f" Default: {fmt_w(def_w)}") print(f" Default: {fmt_w(def_w)}")
if min_w is not None and max_w is not None: if min_w is not None and max_w is not None:
print(f" Range: {min_w} – {max_w} W") print(f" Range: {min_w} – {max_w} W")
if pwr.get("rm_power_supported"):
print(
" Experimental RM power: available (opt-in via web UI or profile;"
" extends range to 30 W)"
)
# ── Privilege / browser helpers ─────────────────────────────────────────────── # ── Privilege / browser helpers ───────────────────────────────────────────────
@@ -420,8 +428,8 @@ def _open_browser_as_user(url: str) -> None:
stderr=subprocess.DEVNULL, stderr=subprocess.DEVNULL,
) )
return return
except Exception: except Exception as exc:
pass log.debug("runuser xdg-open failed, falling back to webbrowser: %s", exc)
import webbrowser import webbrowser
webbrowser.open(url) webbrowser.open(url)
@@ -445,7 +453,7 @@ def require_root():
] ]
try: try:
# PYTHONDONTWRITEBYTECODE prevents root-owned __pycache__ in site-packages. # PYTHONDONTWRITEBYTECODE prevents root-owned __pycache__ in site-packages.
os.execvp( os.execvp( # noqa: S606 — intentional re-exec via sudo
"sudo", "sudo",
[ [
"sudo", "sudo",
@@ -469,6 +477,21 @@ _PERSISTENT_CONFIG_FILE = (
) )
_DAEMON_SOCKET_PATH = "/run/nvcurve-daemon.sock" _DAEMON_SOCKET_PATH = "/run/nvcurve-daemon.sock"
def _configured_max_delta() -> int:
"""Frequency safety cap from the persistent config (built-in default fallback).
The operator-configured cap is authoritative for all direct-hardware
write paths (write, profile apply, verify) unless explicitly overridden
with --max-delta.
"""
try:
with open(_PERSISTENT_CONFIG_FILE) as f:
return json.load(f).get("max_delta_khz", default_config.max_delta_khz)
except (FileNotFoundError, json.JSONDecodeError, OSError):
return default_config.max_delta_khz
_ALLOWED_HOSTS = {"127.0.0.1", "::1", "localhost"} _ALLOWED_HOSTS = {"127.0.0.1", "::1", "localhost"}
@@ -501,7 +524,7 @@ def _safe_host(host: str, cfg: Config) -> str:
0.0.0.0 (bind-all) is silently remapped to 127.0.0.1 — it's a valid local 0.0.0.0 (bind-all) is silently remapped to 127.0.0.1 — it's a valid local
server address, just not usable as a client connection target. server address, just not usable as a client connection target.
""" """
if host in ("0.0.0.0", "::"): if host in ("0.0.0.0", "::"): # noqa: S104 — comparison only, no binding here
return "127.0.0.1" return "127.0.0.1"
if host not in _ALLOWED_HOSTS: if host not in _ALLOWED_HOSTS:
print( print(
@@ -514,7 +537,7 @@ def _safe_host(host: str, cfg: Config) -> str:
def _log_file() -> str: def _log_file() -> str:
return "/var/log/nvcurve.log" if os.geteuid() == 0 else "/tmp/nvcurve.log" return "/var/log/nvcurve.log" if os.geteuid() == 0 else "/tmp/nvcurve.log" # noqa: S108
def _read_server_info() -> dict | None: def _read_server_info() -> dict | None:
@@ -545,12 +568,18 @@ def _discover_server_url(cfg: Config) -> str:
1. /run/nvcurve.json — runtime info written by the running server process 1. /run/nvcurve.json — runtime info written by the running server process
2. /etc/nvcurve/config.json — persistent config written by `service install` 2. /etc/nvcurve/config.json — persistent config written by `service install`
3. Config defaults — 127.0.0.1:8042 3. Config defaults — 127.0.0.1:8042
The scheme is https when TLS is configured (or reported by the running
server), http otherwise.
""" """
scheme = "https" if tls_enabled(cfg) else "http"
# 1. Runtime info (most accurate — reflects the actual running port) # 1. Runtime info (most accurate — reflects the actual running port)
info = _read_server_info() info = _read_server_info()
if info: if info:
host = _safe_host(info["host"], cfg) host = _safe_host(info["host"], cfg)
return f"http://{host}:{info['port']}" s = "https" if info.get("tls") else scheme
return f"{s}://{host}:{info['port']}"
# 2. Persistent config (survives reboots; written by `service install`) # 2. Persistent config (survives reboots; written by `service install`)
try: try:
@@ -558,12 +587,12 @@ def _discover_server_url(cfg: Config) -> str:
data = json.load(f) data = json.load(f)
host = _safe_host(data.get("host", cfg.host), cfg) host = _safe_host(data.get("host", cfg.host), cfg)
port = data.get("port", cfg.port) port = data.get("port", cfg.port)
return f"http://{host}:{port}" return f"{scheme}://{host}:{port}"
except (FileNotFoundError, json.JSONDecodeError, KeyError): except (FileNotFoundError, json.JSONDecodeError, KeyError):
pass pass
# 3. Hardcoded defaults # 3. Hardcoded defaults
return f"http://{cfg.host}:{cfg.port}" return f"{scheme}://{cfg.host}:{cfg.port}"
# ── Subcommand handlers ─────────────────────────────────────────────────────── # ── Subcommand handlers ───────────────────────────────────────────────────────
@@ -737,8 +766,14 @@ def cmd_inspect(args):
def cmd_write(args): def cmd_write(args):
delta_khz = int(args.delta * 1000) try:
max_delta_khz = int(args.max_delta * 1000) if args.max_delta is not None else None delta_khz = int(args.delta * 1000)
max_delta_khz = (
int(args.max_delta * 1000) if args.max_delta is not None else None
)
except (TypeError, ValueError, OverflowError) as exc:
print(f"Error: invalid numeric argument: {exc}", file=sys.stderr)
sys.exit(1)
point_deltas = {} point_deltas = {}
if args.reset: if args.reset:
@@ -814,7 +849,7 @@ def cmd_write(args):
} }
effective_max = ( effective_max = (
max_delta_khz if max_delta_khz is not None else default_config.max_delta_khz max_delta_khz if max_delta_khz is not None else _configured_max_delta()
) )
errors = validate_write(point_deltas, effective_max) errors = validate_write(point_deltas, effective_max)
if errors: if errors:
@@ -837,14 +872,15 @@ def cmd_write(args):
print(f"Write OK — {len(point_deltas)} point(s) updated.") print(f"Write OK — {len(point_deltas)} point(s) updated.")
try: try:
curve_state = None
if not args.glob: if not args.glob:
curve_state, _ = read_curve(gpu, gpu_name) curve_state, _ = read_curve(gpu, gpu_name)
if curve_state: if curve_state:
vfp_freqs = [p.freq_khz for p in curve_state.points] vfp_freqs = [p.freq_khz for p in curve_state.points]
for w in check_negative_freq_warnings(point_deltas, vfp_freqs, []): for w in check_negative_freq_warnings(point_deltas, vfp_freqs, []):
print(f"WARNING: {w}") print(f"WARNING: {w}")
except Exception: except Exception as exc:
pass log.debug("Post-write curve check failed: %s", exc)
def cmd_verify(args): def cmd_verify(args):
@@ -854,8 +890,13 @@ def cmd_verify(args):
from .hal.gpu import get_gpu from .hal.gpu import get_gpu
from .hal.snapshot import save as snapshot_save from .hal.snapshot import save as snapshot_save
from .hal.vfcurve import read_clock_offsets, write_offsets from .hal.vfcurve import read_clock_offsets, write_offsets
from .safety import validate_write
delta_khz = int(args.delta * 1000) try:
delta_khz = int(args.delta * 1000)
except (TypeError, ValueError, OverflowError) as exc:
print(f"Error: invalid numeric argument: {exc}", file=sys.stderr)
sys.exit(1)
if args.point is not None: if args.point is not None:
points = [args.point] points = [args.point]
@@ -867,6 +908,13 @@ def cmd_verify(args):
point_deltas = dict.fromkeys(points, delta_khz) point_deltas = dict.fromkeys(points, delta_khz)
# Enforce the operator-configured safety cap (same as cmd_write).
errors = validate_write(point_deltas, _configured_max_delta())
if errors:
for e in errors:
print(f"Error: {e}", file=sys.stderr)
sys.exit(1)
gpu, gpu_name = get_gpu(index=getattr(args, "gpu_index", 0)) gpu, gpu_name = get_gpu(index=getattr(args, "gpu_index", 0))
print("=== Write-Verify Cycle ===") print("=== Write-Verify Cycle ===")
@@ -1016,8 +1064,11 @@ def _profile_config_write(key: str, value) -> None:
data.pop(key, None) data.pop(key, None)
else: else:
data[key] = value data[key] = value
with open(_PERSISTENT_CONFIG_FILE, "w") as f: try:
_json.dump(data, f, indent=2) with open(_PERSISTENT_CONFIG_FILE, "w") as f:
_json.dump(data, f, indent=2)
except OSError as exc:
raise RuntimeError(f"Cannot write {_PERSISTENT_CONFIG_FILE}: {exc}") from exc
def _gpu_stable_key_offline(gpu_index: int) -> str | None: def _gpu_stable_key_offline(gpu_index: int) -> str | None:
@@ -1066,8 +1117,11 @@ def _profile_config_set_default(gpu_index: int, name: str | None) -> None:
profiles[gpu_key] = name profiles[gpu_key] = name
if not profiles: if not profiles:
data.pop("auto_load_profiles", None) data.pop("auto_load_profiles", None)
with open(_PERSISTENT_CONFIG_FILE, "w") as f: try:
_json.dump(data, f, indent=2) with open(_PERSISTENT_CONFIG_FILE, "w") as f:
_json.dump(data, f, indent=2)
except OSError as exc:
raise RuntimeError(f"Cannot write {_PERSISTENT_CONFIG_FILE}: {exc}") from exc
def cmd_profile(args): def cmd_profile(args):
@@ -1094,8 +1148,8 @@ def cmd_profile(args):
profiles.append( profiles.append(
{"name": name, "curve_deltas": p.get("curve_deltas", {})} {"name": name, "curve_deltas": p.get("curve_deltas", {})}
) )
except Exception: except Exception as exc:
pass log.debug("Skipping unreadable profile %s: %s", path, exc)
if not profiles: if not profiles:
print("No profiles found.") print("No profiles found.")
return return
@@ -1117,7 +1171,7 @@ def cmd_profile(args):
require_root() require_root()
try: try:
_profile_config_set_default(gpu_index, None if clearing else args.name) _profile_config_set_default(gpu_index, None if clearing else args.name)
except ValueError as e: except (ValueError, RuntimeError) as e:
print(f"Error: {e}", file=sys.stderr) print(f"Error: {e}", file=sys.stderr)
return return
if clearing: if clearing:
@@ -1158,12 +1212,33 @@ def cmd_profile(args):
power_limit_w = None power_limit_w = None
mem_offset_mhz = None mem_offset_mhz = None
# Capture the GPU's power-cap mode: prefer the running server (most
# current), else fall back to the persisted per-GPU mode from config
# (so a profile saved while the server is down or auth is enabled
# still records the GPU's actual mode rather than assuming nvml).
power_cap_mode = "nvml"
try:
from .client import NvCurveClient
base = getattr(args, "server", None) or _discover_server_url(default_config)
limits = NvCurveClient(base=base, gpu_index=gpu_index).limits()
if limits.get("power_cap_mode") in ("nvml", "ioctl"):
power_cap_mode = limits["power_cap_mode"]
except Exception as exc:
log.debug("Could not read power-cap mode from server: %s", exc)
gpu_key = _gpu_stable_key_offline(gpu_index)
if gpu_key is not None:
persisted = default_config.power_cap_modes.get(gpu_key)
if persisted in ("nvml", "ioctl"):
power_cap_mode = persisted
data = ProfileData( data = ProfileData(
name=args.name, name=args.name,
gpu_name=gpu_name, gpu_name=gpu_name,
curve_deltas=curve_deltas, curve_deltas=curve_deltas,
mem_offset_mhz=mem_offset_mhz, mem_offset_mhz=mem_offset_mhz,
power_limit_w=power_limit_w, power_limit_w=power_limit_w,
power_cap_mode=power_cap_mode,
) )
filepath = save_profile(default_config.profile_dir, data) filepath = save_profile(default_config.profile_dir, data)
print(f"Saved profile '{args.name}' to {filepath}") print(f"Saved profile '{args.name}' to {filepath}")
@@ -1200,13 +1275,21 @@ def cmd_profile(args):
errs.append(f"Mem offset: {msg}") errs.append(f"Mem offset: {msg}")
if profile.power_limit_w is not None: if profile.power_limit_w is not None:
ok, msg = set_power_limit(profile.power_limit_w, gpu_index) mode = profile.power_cap_mode or "nvml"
ok, msg = set_power_limit(profile.power_limit_w, gpu_index, mode)
if not ok: if not ok:
errs.append(f"Power limit: {msg}") errs.append(f"Power limit: {msg}")
if profile.curve_deltas: if profile.curve_deltas:
deltas = {int(k): v for k, v in profile.curve_deltas.items()} try:
errors = validate_write(deltas, default_config.max_delta_khz) deltas = {int(k): v for k, v in profile.curve_deltas.items()}
except ValueError:
print(
f"Profile '{args.name}' has invalid curve point keys.",
file=sys.stderr,
)
sys.exit(1)
errors = validate_write(deltas, _configured_max_delta())
if errors: if errors:
errs.append("Curve: " + "; ".join(errors)) errs.append("Curve: " + "; ".join(errors))
else: else:
@@ -1328,7 +1411,11 @@ def cmd_setup(args):
"""One-shot hardware compatibility check: diag → read → write-verify → restore.""" """One-shot hardware compatibility check: diag → read → write-verify → restore."""
explicit_point = getattr(args, "point", None) explicit_point = getattr(args, "point", None)
verify_delta_mhz = getattr(args, "delta", 5.0) or 5.0 verify_delta_mhz = getattr(args, "delta", 5.0) or 5.0
verify_delta_khz = int(verify_delta_mhz * 1000) try:
verify_delta_khz = int(verify_delta_mhz * 1000)
except (TypeError, ValueError, OverflowError) as exc:
print(f"Error: invalid numeric argument: {exc}", file=sys.stderr)
sys.exit(1)
require_root() require_root()
@@ -1452,13 +1539,20 @@ def cmd_setup(args):
print() print()
print("Step 4/4 Restoring snapshot") print("Step 4/4 Restoring snapshot")
print() print()
ok = snapshot_restore(gpu, default_config.snapshot_dir, snap_path) if snap_path is None:
if ok:
print(" Hardware state restored to baseline.")
else:
print( print(
" WARNING: Restore failed. Run: nvcurve snapshot restore", file=sys.stderr " WARNING: Snapshot save failed — cannot restore baseline.",
file=sys.stderr,
) )
else:
ok = snapshot_restore(gpu, default_config.snapshot_dir, snap_path)
if ok:
print(" Hardware state restored to baseline.")
else:
print(
" WARNING: Restore failed. Run: nvcurve snapshot restore",
file=sys.stderr,
)
print() print()
print(sep) print(sep)
@@ -1509,41 +1603,59 @@ def cmd_service(args):
"WantedBy=multi-user.target\n" "WantedBy=multi-user.target\n"
) )
with open(unit_path, "w") as f: try:
f.write(unit) with open(unit_path, "w") as f:
f.write(unit)
except OSError as exc:
print(f"Failed to write {unit_path}: {exc}", file=sys.stderr)
return
print(f"Unit file written to {unit_path}") print(f"Unit file written to {unit_path}")
# Write persistent config. # Write persistent config.
os.makedirs("/etc/nvcurve", exist_ok=True) try:
os.makedirs("/etc/nvcurve", exist_ok=True)
except OSError as exc:
print(f"Failed to create /etc/nvcurve: {exc}", file=sys.stderr)
return
persistent_cfg: dict = {} persistent_cfg: dict = {}
try: try:
with open(_PERSISTENT_CONFIG_FILE) as f: with open(_PERSISTENT_CONFIG_FILE) as f:
persistent_cfg = json.load(f) persistent_cfg = json.load(f)
except Exception: except Exception as exc:
pass log.debug("Could not read persistent config: %s", exc)
host = getattr(args, "host", "127.0.0.1") host = getattr(args, "host", "127.0.0.1")
port = getattr(args, "port", 8042) port = getattr(args, "port", 8042)
auto_serve = getattr(args, "auto_serve", False) auto_serve = getattr(args, "auto_serve", False)
persistent_cfg.update({"host": host, "port": port, "auto_serve": auto_serve}) persistent_cfg.update({"host": host, "port": port, "auto_serve": auto_serve})
with open(_PERSISTENT_CONFIG_FILE, "w") as f: if getattr(args, "ssl_certfile", None):
json.dump(persistent_cfg, f, indent=2) persistent_cfg["ssl_certfile"] = args.ssl_certfile
if getattr(args, "ssl_keyfile", None):
persistent_cfg["ssl_keyfile"] = args.ssl_keyfile
try:
with open(_PERSISTENT_CONFIG_FILE, "w") as f:
json.dump(persistent_cfg, f, indent=2)
except OSError as exc:
print(f"Failed to write {_PERSISTENT_CONFIG_FILE}: {exc}", file=sys.stderr)
return
print(f"Persistent config written to {_PERSISTENT_CONFIG_FILE}") print(f"Persistent config written to {_PERSISTENT_CONFIG_FILE}")
scheme = (
"https"
if persistent_cfg.get("ssl_certfile") and persistent_cfg.get("ssl_keyfile")
else "http"
)
if auto_serve: if auto_serve:
print(f" Web server will auto-start on boot at {host}:{port}") print(f" Web server will auto-start on boot at {scheme}://{host}:{port}")
else: else:
print( print(
f" Web server default: {host}:{port} (start on demand: nvcurve serve start)" f" Web server default: {scheme}://{host}:{port} "
"(start on demand: nvcurve serve start)"
) )
try: try:
subprocess.run(["systemctl", "daemon-reload"], check=True) subprocess.run(["systemctl", "daemon-reload"], check=True)
was_active = ( probe = subprocess.run(["systemctl", "is-active", "--quiet", "nvcurve"])
subprocess.run( was_active = probe.returncode == 0
["systemctl", "is-active", "--quiet", "nvcurve"],
).returncode
== 0
)
subprocess.run(["systemctl", "enable", "--now", "nvcurve"], check=True) subprocess.run(["systemctl", "enable", "--now", "nvcurve"], check=True)
print("Service enabled and started.") print("Service enabled and started.")
@@ -1560,7 +1672,7 @@ def cmd_service(args):
print(" systemctl status nvcurve") print(" systemctl status nvcurve")
print(" journalctl -u nvcurve -f") print(" journalctl -u nvcurve -f")
print(" nvcurve service uninstall") print(" nvcurve service uninstall")
except subprocess.CalledProcessError as e: except (subprocess.CalledProcessError, FileNotFoundError) as e:
print(f"systemctl failed: {e}", file=sys.stderr) print(f"systemctl failed: {e}", file=sys.stderr)
elif action == "uninstall": elif action == "uninstall":
@@ -1660,25 +1772,31 @@ def cmd_service(args):
auto_serve = pcfg.get("auto_serve", False) auto_serve = pcfg.get("auto_serve", False)
host = pcfg.get("host", "127.0.0.1") host = pcfg.get("host", "127.0.0.1")
port = pcfg.get("port", 8042) port = pcfg.get("port", 8042)
tls = bool(pcfg.get("ssl_certfile") and pcfg.get("ssl_keyfile"))
print() print()
print(f"web server auto-start: {'on' if auto_serve else 'off'}") print(f"web server auto-start: {'on' if auto_serve else 'off'}")
print(f"web server address: {host}:{port}") print(f"web server address: {'https' if tls else 'http'}://{host}:{port}")
print(f"web server TLS: {'on' if tls else 'off'}")
print() print()
print( print(
"Change with: nvcurve service configure [--auto-serve|--no-auto-serve] [--host H] [--port P]" "Change with: nvcurve service configure [--auto-serve|--no-auto-serve] [--host H] [--port P] [--ssl-certfile C --ssl-keyfile K]"
) )
elif action == "configure": elif action == "configure":
require_root() require_root()
import subprocess import subprocess
os.makedirs("/etc/nvcurve", exist_ok=True) try:
os.makedirs("/etc/nvcurve", exist_ok=True)
except OSError as exc:
print(f"Failed to create /etc/nvcurve: {exc}", file=sys.stderr)
return
pcfg: dict = {} pcfg: dict = {}
try: try:
with open(_PERSISTENT_CONFIG_FILE) as f: with open(_PERSISTENT_CONFIG_FILE) as f:
pcfg = json.load(f) pcfg = json.load(f)
except Exception: except Exception as exc:
pass log.debug("Could not read persistent config: %s", exc)
if hasattr(args, "auto_serve") and args.auto_serve is not None: if hasattr(args, "auto_serve") and args.auto_serve is not None:
pcfg["auto_serve"] = args.auto_serve pcfg["auto_serve"] = args.auto_serve
@@ -1686,13 +1804,27 @@ def cmd_service(args):
pcfg["host"] = args.host pcfg["host"] = args.host
if hasattr(args, "port") and args.port is not None: if hasattr(args, "port") and args.port is not None:
pcfg["port"] = args.port pcfg["port"] = args.port
if getattr(args, "ssl_certfile", None):
pcfg["ssl_certfile"] = args.ssl_certfile
if getattr(args, "ssl_keyfile", None):
pcfg["ssl_keyfile"] = args.ssl_keyfile
if getattr(args, "no_ssl", False):
pcfg.pop("ssl_certfile", None)
pcfg.pop("ssl_keyfile", None)
with open(_PERSISTENT_CONFIG_FILE, "w") as f: try:
json.dump(pcfg, f, indent=2) with open(_PERSISTENT_CONFIG_FILE, "w") as f:
json.dump(pcfg, f, indent=2)
except OSError as exc:
print(f"Failed to write {_PERSISTENT_CONFIG_FILE}: {exc}", file=sys.stderr)
return
print(f"Config updated ({_PERSISTENT_CONFIG_FILE}):") print(f"Config updated ({_PERSISTENT_CONFIG_FILE}):")
print(f" auto-serve: {'on' if pcfg.get('auto_serve', False) else 'off'}") print(f" auto-serve: {'on' if pcfg.get('auto_serve', False) else 'off'}")
print(f" host: {pcfg.get('host', '127.0.0.1')}") print(f" host: {pcfg.get('host', '127.0.0.1')}")
print(f" port: {pcfg.get('port', 8042)}") print(f" port: {pcfg.get('port', 8042)}")
print(
f" TLS: {'on' if pcfg.get('ssl_certfile') and pcfg.get('ssl_keyfile') else 'off'}"
)
if os.path.exists(unit_path): if os.path.exists(unit_path):
try: try:
@@ -1712,11 +1844,31 @@ def _cmd_serve_start(args, cfg: Config, open_browser: bool = False) -> None:
host = getattr(args, "host", cfg.host) host = getattr(args, "host", cfg.host)
port = getattr(args, "port", cfg.port) port = getattr(args, "port", cfg.port)
# Optional TLS (CLI flags override the persistent config).
ssl_certfile = getattr(args, "ssl_certfile", None)
ssl_keyfile = getattr(args, "ssl_keyfile", None)
if ssl_certfile:
cfg.ssl_certfile = ssl_certfile
if ssl_keyfile:
cfg.ssl_keyfile = ssl_keyfile
# --direct: skip daemon round-trip (used when the daemon itself spawns us). # --direct: skip daemon round-trip (used when the daemon itself spawns us).
if getattr(args, "direct", False): if getattr(args, "direct", False):
require_root() require_root()
with open(_SERVER_INFO_FILE, "w") as f: try:
json.dump({"pid": os.getpid(), "host": host, "port": port}, f) with open(_SERVER_INFO_FILE, "w") as f:
json.dump(
{
"pid": os.getpid(),
"host": host,
"port": port,
"tls": tls_enabled(cfg),
},
f,
)
except OSError as exc:
print(f"Failed to write {_SERVER_INFO_FILE}: {exc}", file=sys.stderr)
return
try: try:
from .server import run as server_run from .server import run as server_run
@@ -1733,13 +1885,34 @@ def _cmd_serve_start(args, cfg: Config, open_browser: bool = False) -> None:
return return
# Prefer daemon socket: no root required, daemon manages the server process. # Prefer daemon socket: no root required, daemon manages the server process.
resp = _daemon_send({"cmd": "serve_start", "host": host, "port": port}) # The daemon always binds the configured host/port (callers cannot choose
# the interface), so report the address from the daemon's response.
resp = _daemon_send({"cmd": "serve_start"})
if resp is not None: if resp is not None:
if resp.get("ok"): if resp.get("ok"):
print(f"Web server starting (PID {resp['pid']}) at http://{host}:{port}") rhost = resp.get("host", host)
rport = resp.get("port", port)
if (rhost, rport) != (host, port):
print(
f"Note: daemon uses the configured bind address {rhost}:{rport} "
"(change with: nvcurve service configure --host/--port)",
file=sys.stderr,
)
if (ssl_certfile or ssl_keyfile) and not resp.get("tls"):
print(
"Note: --ssl-certfile/--ssl-keyfile are ignored while the daemon "
"manages the server — the daemon uses the TLS settings from "
"/etc/nvcurve/config.json (set with: nvcurve service configure "
"--ssl-certfile/--ssl-keyfile)",
file=sys.stderr,
)
scheme = "https" if resp.get("tls") else "http"
print(
f"Web server starting (PID {resp['pid']}) at {scheme}://{rhost}:{rport}"
)
if open_browser: if open_browser:
time.sleep(1.5) time.sleep(1.5)
_open_browser_as_user(f"http://{host}:{port}") _open_browser_as_user(f"{scheme}://{rhost}:{rport}")
else: else:
print(f"Daemon: {resp.get('error')}", file=sys.stderr) print(f"Daemon: {resp.get('error')}", file=sys.stderr)
return return
@@ -1749,7 +1922,8 @@ def _cmd_serve_start(args, cfg: Config, open_browser: bool = False) -> None:
info = _read_server_info() info = _read_server_info()
if info: if info:
url = f"http://{info['host']}:{info['port']}" scheme = "https" if info.get("tls") else "http"
url = f"{scheme}://{info['host']}:{info['port']}"
print(f"Server is already running (PID {info['pid']}) at {url}.") print(f"Server is already running (PID {info['pid']}) at {url}.")
if open_browser: if open_browser:
_open_browser_as_user(url) _open_browser_as_user(url)
@@ -1769,12 +1943,20 @@ def _cmd_serve_start(args, cfg: Config, open_browser: bool = False) -> None:
"--port", "--port",
str(port), str(port),
] ]
if ssl_certfile:
cmd += ["--ssl-certfile", ssl_certfile]
if ssl_keyfile:
cmd += ["--ssl-keyfile", ssl_keyfile]
if getattr(args, "gpu_index", 0): if getattr(args, "gpu_index", 0):
cmd += ["--gpu", str(args.gpu_index)] cmd += ["--gpu", str(args.gpu_index)]
log_path = _log_file() log_path = _log_file()
print("Starting nvcurve server in background...") print("Starting nvcurve server in background...")
with open(log_path, "a") as lf: try:
p = subprocess.Popen(cmd, stdout=lf, stderr=lf, start_new_session=True) with open(log_path, "a") as lf:
p = subprocess.Popen(cmd, stdout=lf, stderr=lf, start_new_session=True)
except OSError as exc:
print(f"Failed to open log file {log_path}: {exc}", file=sys.stderr)
return
print(f"Server starting (PID {p.pid}). Logs: {log_path}") print(f"Server starting (PID {p.pid}). Logs: {log_path}")
if open_browser: if open_browser:
time.sleep(1.5) time.sleep(1.5)
@@ -1782,8 +1964,20 @@ def _cmd_serve_start(args, cfg: Config, open_browser: bool = False) -> None:
return return
# Foreground mode — write info file so clients can discover host:port. # Foreground mode — write info file so clients can discover host:port.
with open(_SERVER_INFO_FILE, "w") as f: try:
json.dump({"pid": os.getpid(), "host": host, "port": port}, f) with open(_SERVER_INFO_FILE, "w") as f:
json.dump(
{
"pid": os.getpid(),
"host": host,
"port": port,
"tls": tls_enabled(cfg),
},
f,
)
except OSError as exc:
print(f"Failed to write {_SERVER_INFO_FILE}: {exc}", file=sys.stderr)
return
try: try:
from .server import run as server_run from .server import run as server_run
@@ -1990,6 +2184,12 @@ Examples:
"--host", default="127.0.0.1", help="Bind address (default 127.0.0.1)" "--host", default="127.0.0.1", help="Bind address (default 127.0.0.1)"
) )
p_start.add_argument("--port", type=int, default=8042, help="Port (default 8042)") p_start.add_argument("--port", type=int, default=8042, help="Port (default 8042)")
p_start.add_argument(
"--ssl-certfile", default=None, help="TLS certificate (enables HTTPS)"
)
p_start.add_argument(
"--ssl-keyfile", default=None, help="TLS private key (enables HTTPS)"
)
p_start.add_argument( p_start.add_argument(
"--detach", "-d", action="store_true", help="Run in background" "--detach", "-d", action="store_true", help="Run in background"
) )
@@ -2024,6 +2224,16 @@ Examples:
default=8042, default=8042,
help="Default web server port (stored in config)", help="Default web server port (stored in config)",
) )
p_install.add_argument(
"--ssl-certfile",
default=None,
help="TLS certificate (stored in config; enables HTTPS)",
)
p_install.add_argument(
"--ssl-keyfile",
default=None,
help="TLS private key (stored in config; enables HTTPS)",
)
p_configure = s_svc.add_parser( p_configure = s_svc.add_parser(
"configure", help="Update config and restart daemon (escalates to root)" "configure", help="Update config and restart daemon (escalates to root)"
@@ -2043,6 +2253,19 @@ Examples:
) )
p_configure.add_argument("--host", default=None, help="Web server bind address") p_configure.add_argument("--host", default=None, help="Web server bind address")
p_configure.add_argument("--port", type=int, default=None, help="Web server port") p_configure.add_argument("--port", type=int, default=None, help="Web server port")
p_configure.add_argument(
"--ssl-certfile",
default=None,
help="TLS certificate (stored in config; enables HTTPS)",
)
p_configure.add_argument(
"--ssl-keyfile", default=None, help="TLS private key (stored in config)"
)
p_configure.add_argument(
"--no-ssl",
action="store_true",
help="Disable TLS (remove certificate/key from config)",
)
s_svc.add_parser("uninstall", help="Remove systemd service (escalates to root)") s_svc.add_parser("uninstall", help="Remove systemd service (escalates to root)")
s_svc.add_parser("start", help="Start systemd service (escalates to root)") s_svc.add_parser("start", help="Start systemd service (escalates to root)")
@@ -2083,9 +2306,17 @@ def main():
"users_file", "users_file",
"host", "host",
"port", "port",
"ssl_certfile",
"ssl_keyfile",
"allow_api_shutdown",
"gitea_url",
"gitea_repo",
"gitea_token",
): ):
if key in data: if key in data:
setattr(cfg, key, data[key]) setattr(cfg, key, data[key])
if "trusted_proxies" in data:
cfg.trusted_proxies = normalize_trusted_proxies(data["trusted_proxies"])
if "auto_load_profiles" in data: if "auto_load_profiles" in data:
# Keys are stable GPU identifiers (UUID, "pci:XXXX", or "idx:N") # Keys are stable GPU identifiers (UUID, "pci:XXXX", or "idx:N")
cfg.auto_load_profiles = dict(data["auto_load_profiles"]) cfg.auto_load_profiles = dict(data["auto_load_profiles"])
@@ -2095,8 +2326,16 @@ def main():
if "fan_curves" in data: if "fan_curves" in data:
# Per-GPU active fan curves, restored on server startup. # Per-GPU active fan curves, restored on server startup.
cfg.fan_curves = dict(data["fan_curves"]) cfg.fan_curves = dict(data["fan_curves"])
except Exception: if "power_cap_modes" in data:
pass # Per-GPU experimental power-cap mode. "nvml" is the default
# (the server treats it as unset); keep only valid values.
cfg.power_cap_modes = {
str(k): str(v)
for k, v in dict(data["power_cap_modes"]).items()
if str(v) in ("nvml", "ioctl")
}
except Exception as exc:
log.debug("Could not load user config: %s", exc)
base_url = args.server or _discover_server_url(cfg) base_url = args.server or _discover_server_url(cfg)
client = NvCurveClient(base=base_url, gpu_index=getattr(args, "gpu_index", 0)) client = NvCurveClient(base=base_url, gpu_index=getattr(args, "gpu_index", 0))
@@ -2148,8 +2387,8 @@ def main():
if os.path.exists(_SERVER_INFO_FILE): if os.path.exists(_SERVER_INFO_FILE):
try: try:
os.remove(_SERVER_INFO_FILE) os.remove(_SERVER_INFO_FILE)
except OSError: except OSError as exc:
pass log.debug("Could not remove %s: %s", _SERVER_INFO_FILE, exc)
except ApiError as e: except ApiError as e:
if e.status_code == 401: if e.status_code == 401:
print( print(
+8 -10
View File
@@ -120,18 +120,13 @@ class NvCurveClient:
def write_curve( def write_curve(
self, self,
deltas: dict[int, int], deltas: dict[int, int],
max_delta_khz: int | None = None,
) -> dict: ) -> dict:
body: dict = {"deltas": deltas} # The server enforces its configured safety cap; clients cannot
if max_delta_khz is not None: # override it per request.
body["max_delta_khz"] = max_delta_khz return self._post("/api/curve/write", {"deltas": deltas})
return self._post("/api/curve/write", body)
def write_global(self, delta_khz: int, max_delta_khz: int | None = None) -> dict: def write_global(self, delta_khz: int) -> dict:
body: dict = {"delta_khz": delta_khz} return self._post("/api/curve/write/global", {"delta_khz": delta_khz})
if max_delta_khz is not None:
body["max_delta_khz"] = max_delta_khz
return self._post("/api/curve/write/global", body)
def reset_curve(self) -> dict: def reset_curve(self) -> dict:
return self._post("/api/curve/reset") return self._post("/api/curve/reset")
@@ -150,6 +145,9 @@ class NvCurveClient:
def snapshots(self) -> list: def snapshots(self) -> list:
return self._get("/api/snapshots") return self._get("/api/snapshots")
def limits(self) -> dict:
return self._get("/api/limits")
# ── Profiles ───────────────────────────────────────────────────────────── # ── Profiles ─────────────────────────────────────────────────────────────
def profiles(self) -> dict: def profiles(self) -> dict:
+60 -2
View File
@@ -17,6 +17,21 @@ class Config:
host: str = "127.0.0.1" host: str = "127.0.0.1"
port: int = 8042 port: int = 8042
# Optional TLS: when both are set, the server serves HTTPS and the
# session cookie is marked Secure. Off by default (plain HTTP).
ssl_certfile: str | None = None
ssl_keyfile: str | None = None
# Proxy IPs (e.g. a reverse proxy on 127.0.0.1) whose X-Forwarded-For
# header is trusted for the login brute-force lockout. JSON array in
# config.json (a comma-separated string is also accepted and normalized).
# Without this, all proxied clients share the proxy's IP.
trusted_proxies: list[str] = field(default_factory=list)
# Allow any authenticated user to stop the server via POST /api/shutdown.
# Set false on shared systems; manage the service via systemd instead.
allow_api_shutdown: bool = True
snapshot_dir: str = "/var/cache/nvcurve/snapshots" snapshot_dir: str = "/var/cache/nvcurve/snapshots"
profile_dir: str = "/etc/nvcurve/profiles" profile_dir: str = "/etc/nvcurve/profiles"
@@ -34,9 +49,52 @@ class Config:
# curve applied via the UI survives server restarts (fan control itself is # curve applied via the UI survives server restarts (fan control itself is
# volatile — the driver reverts to automatic mode on reboot). # volatile — the driver reverts to automatic mode on reboot).
# Key = stable GPU identifier (same as auto_load_profiles). # Key = stable GPU identifier (same as auto_load_profiles).
# Value = list of {"temp_c": int, "fan_pct": int} sorted by temp_c. # Value = {"curve": [{"temp_c": int, "fan_pct": int}, ...] sorted by temp_c,
fan_curves: dict[str, list] = field(default_factory=dict) # "fans": [fan indices] | None (None = all fans)}.
# Legacy entries (bare curve list) are migrated at load time.
fan_curves: dict[str, object] = field(default_factory=dict)
# Per-GPU power-cap mode: "nvml" (default, never stored) or "ioctl"
# (experimental RM power control — permits caps below the VBIOS minimum).
# Key = stable GPU identifier (same as auto_load_profiles).
power_cap_modes: dict[str, str] = field(default_factory=dict)
# Gitea "Report Bug" integration: the built-in defaults point at the
# upstream project's Gitea, so the feature works out of the box for every
# install — no configuration required. Forks can override these in
# config.json (own instance/repo/token); setting gitea_token to ""
# disables the feature. The token is scoped to write:issue on a public
# repo, so it can only create/read issues there.
gitea_url: str = "https://gitea.zephyre.one"
gitea_repo: str = "Pakobbix/nvcurve"
# Intentionally embedded: zero-config bug reporting for a public repo.
# The token can only create/read issues on Pakobbix/nvcurve (scope
# write:issue) and is revocable/rotatable by the repo owner.
# pi-lens-ignore: S105
gitea_token: str = "bfe416043b633f8f3c4df103272ad0ccf91c8f22"
# Module-level default config instance. # Module-level default config instance.
default_config = Config() default_config = Config()
def tls_enabled(cfg: Config) -> bool:
"""True when both TLS files are configured (server serves HTTPS)."""
return bool(cfg.ssl_certfile and cfg.ssl_keyfile)
def normalize_trusted_proxies(value) -> list[str]:
"""Normalize a trusted_proxies config value to a list of IP strings.
Accepts a JSON array (the documented format) or a comma-separated string
(tolerated for convenience). Normalizing matters because the server does
exact list membership tests — a raw string would degrade to substring
matching (e.g. "127.0.0.1" in "127.0.0.10").
"""
if value is None:
return []
if isinstance(value, str):
return [h.strip() for h in value.split(",") if h.strip()]
if isinstance(value, (list, tuple)):
return [str(h).strip() for h in value if str(h).strip()]
return []
+34 -10
View File
@@ -7,10 +7,16 @@ Protocol: newline-delimited JSON, one request → one response, connection close
Commands: Commands:
{"cmd": "ping"} {"cmd": "ping"}
{"cmd": "serve_start", "host": "127.0.0.1", "port": 8042} {"cmd": "serve_start"}
{"cmd": "serve_stop"} {"cmd": "serve_stop"}
{"cmd": "serve_status"} {"cmd": "serve_status"}
The socket is world-connectable (unprivileged users drive it via the CLI),
so the command surface is deliberately minimal: serve_start ALWAYS binds the
configured host/port from /etc/nvcurve/config.json — callers cannot choose
the bind address (no ad-hoc 0.0.0.0 exposure). Changing the bind address is
an operator action via `nvcurve service configure`.
Requires root. Requires root.
""" """
@@ -23,7 +29,7 @@ import signal
import subprocess import subprocess
import sys import sys
from .config import Config from .config import Config, normalize_trusted_proxies
log = logging.getLogger("nvcurve.daemon") log = logging.getLogger("nvcurve.daemon")
@@ -38,7 +44,13 @@ _cfg: Config | None = None # Config instance, set in run()
# ── Socket command handlers ──────────────────────────────────────────────────── # ── Socket command handlers ────────────────────────────────────────────────────
async def _handle_serve_start(host: str, port: int) -> dict: async def _handle_serve_start() -> dict:
"""Start the web server on the *configured* host/port.
The bind address is taken from /etc/nvcurve/config.json only — the
socket is reachable by unprivileged users, so callers must not be able
to choose the interface (e.g. binding 0.0.0.0 to expose the API).
"""
global _server_proc global _server_proc
if _server_proc is not None and _server_proc.poll() is None: if _server_proc is not None and _server_proc.poll() is None:
return { return {
@@ -47,6 +59,8 @@ async def _handle_serve_start(host: str, port: int) -> dict:
"pid": _server_proc.pid, "pid": _server_proc.pid,
} }
host = _cfg.host if _cfg is not None else "127.0.0.1"
port = _cfg.port if _cfg is not None else 8042
cmd = [ cmd = [
sys.executable, sys.executable,
"-m", "-m",
@@ -72,7 +86,13 @@ async def _handle_serve_start(host: str, port: int) -> dict:
except OSError as exc: except OSError as exc:
return {"ok": False, "error": f"cannot open log file {log_path}: {exc}"} return {"ok": False, "error": f"cannot open log file {log_path}: {exc}"}
log.info("Web server started (PID %d)", _server_proc.pid) log.info("Web server started (PID %d)", _server_proc.pid)
return {"ok": True, "pid": _server_proc.pid} return {
"ok": True,
"pid": _server_proc.pid,
"host": host,
"port": port,
"tls": bool(_cfg and _cfg.ssl_certfile and _cfg.ssl_keyfile),
}
async def _handle_serve_stop() -> dict: async def _handle_serve_stop() -> dict:
@@ -105,9 +125,7 @@ async def _dispatch(req: dict) -> dict:
elif cmd == "serve_start": elif cmd == "serve_start":
if _cfg is None: if _cfg is None:
return {"ok": False, "error": "config not initialized"} return {"ok": False, "error": "config not initialized"}
host = req.get("host", _cfg.host) return await _handle_serve_start()
port = req.get("port", _cfg.port)
return await _handle_serve_start(host, port)
elif cmd == "serve_stop": elif cmd == "serve_stop":
return await _handle_serve_stop() return await _handle_serve_stop()
elif cmd == "serve_status": elif cmd == "serve_status":
@@ -174,9 +192,13 @@ def run() -> None:
"profile_dir", "profile_dir",
"host", "host",
"port", "port",
"ssl_certfile",
"ssl_keyfile",
): ):
if key in cfg_data: if key in cfg_data:
setattr(_cfg, key, cfg_data[key]) setattr(_cfg, key, cfg_data[key])
if "trusted_proxies" in cfg_data:
_cfg.trusted_proxies = normalize_trusted_proxies(cfg_data["trusted_proxies"])
# Apply auto-load profiles in a subprocess so the daemon process itself # Apply auto-load profiles in a subprocess so the daemon process itself
# never loads NvAPI/NVML/HAL modules — keeps steady-state RSS low. # never loads NvAPI/NVML/HAL modules — keeps steady-state RSS low.
@@ -207,10 +229,12 @@ async def _serve_socket(auto_serve: bool = False) -> None:
server = await asyncio.start_unix_server(_handle_client, path=SOCKET_PATH) server = await asyncio.start_unix_server(_handle_client, path=SOCKET_PATH)
# The socket must be connectable by unprivileged users: the CLI runs as the # The socket must be connectable by unprivileged users: the CLI runs as the
# regular user and talks to this root daemon over the socket. 0o666 is # regular user and talks to this root daemon over the socket. 0o666 is
# intentional (standard for /run daemon sockets). # intentional — the command surface is restricted accordingly (serve_start
# always uses the configured host/port; see module docstring).
# pi-lens-ignore: S103 # pi-lens-ignore: S103
_SOCKET_MODE = 0o666
os.chmod( os.chmod(
SOCKET_PATH, 0o666 SOCKET_PATH, _SOCKET_MODE
) # nosemgrep: python.lang.security.audit.insecure-file-permissions.insecure-file-permissions ) # nosemgrep: python.lang.security.audit.insecure-file-permissions.insecure-file-permissions
log.info("Daemon listening on %s", SOCKET_PATH) log.info("Daemon listening on %s", SOCKET_PATH)
@@ -219,7 +243,7 @@ async def _serve_socket(auto_serve: bool = False) -> None:
log.warning("auto_serve requested but config not initialized") log.warning("auto_serve requested but config not initialized")
else: else:
log.info("auto_serve enabled — starting web server on boot") log.info("auto_serve enabled — starting web server on boot")
await _handle_serve_start(_cfg.host, _cfg.port) await _handle_serve_start()
stop_event = asyncio.Event() stop_event = asyncio.Event()
loop = asyncio.get_running_loop() loop = asyncio.get_running_loop()
+161 -49
View File
@@ -1,26 +1,33 @@
"""Hardware Abstraction Layer for Fan Control. """Hardware Abstraction Layer for Fan Control.
Uses NVML (via pynvml) for all operations: Uses NVML (via pynvml) for all operations:
- nvmlDeviceGetFanSpeed_v2 : read current fan speed % for a fan index - nvmlDeviceGetNumFans : number of fans on the device
- nvmlDeviceSetFanSpeed_v2 : set fan speed % for a fan index - nvmlDeviceGetFanSpeed_v2 : read current fan speed % for a fan index
- nvmlDeviceGetMinMaxFanSpeed: get min/max fan speed constraints - nvmlDeviceSetFanSpeed_v2 : set fan speed % for a fan index
- nvmlDeviceGetTemperature : read GPU temp for curve interpolation - nvmlDeviceSetDefaultFanSpeed_v2 : restore automatic control for a fan index
- nvmlDeviceGetMinMaxFanSpeed : get min/max fan speed constraints
- nvmlDeviceGetTemperature : read GPU temp for curve interpolation
Fans are addressed by 0-based index. Passing ``fans=None`` to the set/reset
helpers means "all fans on the device".
""" """
import ctypes import ctypes
import logging import logging
from typing import List, Optional from typing import Any
try: try:
import pynvml import pynvml as _pynvml_import
_NVML_AVAILABLE = True _NVML_AVAILABLE = True
except ImportError: except ImportError:
_pynvml_import = None
_NVML_AVAILABLE = False _NVML_AVAILABLE = False
log = logging.getLogger("nvcurve.hal.fans") # Aliased as Any so attribute access is not flagged when the import failed.
pynvml: Any = _pynvml_import
# We use fan index 0 (first/primary fan) for all operations. log = logging.getLogger("nvcurve.hal.fans")
_FAN_INDEX = 0
def _get_handle(gpu_index: int): def _get_handle(gpu_index: int):
@@ -30,13 +37,42 @@ def _get_handle(gpu_index: int):
return pynvml.nvmlDeviceGetHandleByIndex(gpu_index) return pynvml.nvmlDeviceGetHandleByIndex(gpu_index)
def _num_fans(handle) -> int:
"""Return the number of fans on the device (>= 1 on query failure)."""
try:
return max(0, int(pynvml.nvmlDeviceGetNumFans(handle)))
except pynvml.NVMLError:
# GetNumFans unsupported: assume at least the primary fan exists.
return 1
def get_num_fans(gpu_index: int = 0) -> int:
"""Return the number of fans on the GPU (0 if NVML is unavailable)."""
if not _NVML_AVAILABLE:
return 0
try:
return _num_fans(_get_handle(gpu_index))
except pynvml.NVMLError as exc:
log.warning("get_num_fans: %s", exc)
return 0
def get_fan_info(gpu_index: int = 0) -> dict: def get_fan_info(gpu_index: int = 0) -> dict:
"""Return current fan state: fan_pct, fan_mode, min_fan_pct, max_fan_pct. """Return current fan state for all fans.
Returns a dict with:
fan_pct : current speed % of fan 0 (legacy, None on failure)
fans : [{"index": i, "fan_pct": pct | None}, ...] per fan
num_fans : number of fans on the device
fan_mode : None (the server derives "auto"/"curve")
min_fan_pct / max_fan_pct : device-wide speed constraints
Returns None values on failure. Returns None values on failure.
""" """
out = { out: dict[str, Any] = {
"fan_pct": None, "fan_pct": None,
"fans": [],
"num_fans": 0,
"fan_mode": None, "fan_mode": None,
"min_fan_pct": None, "min_fan_pct": None,
"max_fan_pct": None, "max_fan_pct": None,
@@ -45,42 +81,103 @@ def get_fan_info(gpu_index: int = 0) -> dict:
return out return out
try: try:
handle = _get_handle(gpu_index) handle = _get_handle(gpu_index)
# Get current fan speed using v2 API (fan index 0)
try:
out["fan_pct"] = float(pynvml.nvmlDeviceGetFanSpeed_v2(handle, _FAN_INDEX))
except pynvml.NVMLError:
# Fallback to legacy v1 API
try:
out["fan_pct"] = float(pynvml.nvmlDeviceGetFanSpeed(handle))
except pynvml.NVMLError:
pass
# Get min/max fan speed constraints
try:
min_s = ctypes.c_uint(0)
max_s = ctypes.c_uint(0)
pynvml.nvmlDeviceGetMinMaxFanSpeed(handle, min_s, max_s)
out["min_fan_pct"] = int(min_s.value)
out["max_fan_pct"] = int(max_s.value)
except pynvml.NVMLError:
out["min_fan_pct"] = 0
out["max_fan_pct"] = 100
except pynvml.NVMLError as exc: except pynvml.NVMLError as exc:
log.warning("get_fan_info: %s", exc) log.warning("get_fan_info: %s", exc)
return out
out["num_fans"] = _num_fans(handle)
# Per-fan speeds via the v2 API; fan 0 falls back to the legacy v1 API.
for i in range(out["num_fans"]):
pct: float | None = None
try:
pct = float(pynvml.nvmlDeviceGetFanSpeed_v2(handle, i))
except pynvml.NVMLError:
if i == 0:
try:
pct = float(pynvml.nvmlDeviceGetFanSpeed(handle))
except pynvml.NVMLError:
pct = None
out["fans"].append({"index": i, "fan_pct": pct})
out["fan_pct"] = out["fans"][0]["fan_pct"] if out["fans"] else None
# Get min/max fan speed constraints
try:
min_s = ctypes.c_uint(0)
max_s = ctypes.c_uint(0)
pynvml.nvmlDeviceGetMinMaxFanSpeed(handle, min_s, max_s)
out["min_fan_pct"] = int(min_s.value)
out["max_fan_pct"] = int(max_s.value)
except pynvml.NVMLError:
out["min_fan_pct"] = 0
out["max_fan_pct"] = 100
return out return out
def set_fan_speed(gpu_index: int, pct: int) -> tuple[bool, str]: def set_fan_speed(
"""Set fan speed to a percentage (0-100) on the primary fan.""" gpu_index: int,
pct = max(0, min(100, int(pct))) pct: int,
fans: list[int] | None = None,
) -> tuple[bool, str]:
"""Set fan speed to a percentage (0-100).
fans=None targets every fan on the device; fans=[0, 1] targets the
listed fan indices. In all-fans mode, fans the driver does not allow
manual control of (e.g. driver-mirrored secondary fans) are skipped
with a warning instead of failing the whole operation; explicit fan
lists are strict and fail if any selected fan cannot be set.
"""
try:
pct = max(0, min(100, int(pct)))
except (TypeError, ValueError):
return False, "Invalid fan speed"
if not _NVML_AVAILABLE: if not _NVML_AVAILABLE:
return False, "NVML not available" return False, "NVML not available"
try: try:
handle = _get_handle(gpu_index) handle = _get_handle(gpu_index)
pynvml.nvmlDeviceSetFanSpeed_v2(handle, _FAN_INDEX, pct) num_fans = _num_fans(handle)
if fans is None:
targets = list(range(num_fans))
strict = False
else:
targets: list[int] = []
for f in fans:
try:
f = int(f)
except (TypeError, ValueError):
return False, f"Invalid fan index: {f!r}"
if f < 0 or f >= num_fans:
return False, f"Fan index {f} out of range (0-{num_fans - 1})"
targets.append(f)
strict = True
if not targets:
return False, "No fans selected"
if not targets:
return False, "No fans available on this GPU"
skipped: list[str] = []
for i in targets:
try:
pynvml.nvmlDeviceSetFanSpeed_v2(handle, i, pct)
except pynvml.NVMLError as exc:
if strict:
log.warning(
"set_fan_speed(%d, fan %d, %d): %s", gpu_index, i, pct, exc
)
return False, str(exc)
# All-fans mode: secondary fans may be driver-controlled and
# reject manual writes; skip them and report in the message.
log.debug("set_fan_speed: fan %d not settable: %s", i, exc)
skipped.append(f"fan {i + 1}")
if skipped:
return (
True,
f"OK ({len(skipped)} fan(s) not manually controllable: {', '.join(skipped)})",
)
return True, "OK" return True, "OK"
except pynvml.NVMLError as exc: except pynvml.NVMLError as exc:
log.warning("set_fan_speed(%d, %d): %s", gpu_index, pct, exc) log.warning("set_fan_speed(%d, %d): %s", gpu_index, pct, exc)
@@ -88,10 +185,11 @@ def set_fan_speed(gpu_index: int, pct: int) -> tuple[bool, str]:
def reset_fan(gpu_index: int = 0) -> tuple[bool, str]: def reset_fan(gpu_index: int = 0) -> tuple[bool, str]:
"""Restore automatic fan control. """Restore automatic fan control for all fans.
Tries nvidia-smi --fan=default first (most reliable), then falls back to Tries nvidia-smi --fan=default first (most reliable, resets all fans on
NVML nvmlDeviceSetDefaultFanSpeed_v2. the device), then falls back to NVML nvmlDeviceSetDefaultFanSpeed_v2
per fan index.
""" """
if not _NVML_AVAILABLE: if not _NVML_AVAILABLE:
return False, "NVML not available" return False, "NVML not available"
@@ -102,7 +200,9 @@ def reset_fan(gpu_index: int = 0) -> tuple[bool, str]:
try: try:
ret = subprocess.run( ret = subprocess.run(
["nvidia-smi", "-i", str(gpu_index), "-fan", "default"], ["nvidia-smi", "-i", str(gpu_index), "-fan", "default"],
capture_output=True, text=True, timeout=10, capture_output=True,
text=True,
timeout=10,
) )
if ret.returncode == 0: if ret.returncode == 0:
return True, "OK" return True, "OK"
@@ -114,28 +214,34 @@ def reset_fan(gpu_index: int = 0) -> tuple[bool, str]:
except Exception as exc: except Exception as exc:
log.debug("nvidia-smi -fan default error: %s", exc) log.debug("nvidia-smi -fan default error: %s", exc)
# Fallback: use NVML to reset to default fan speed # Fallback: use NVML to reset every fan to default speed
try: try:
handle = _get_handle(gpu_index) handle = _get_handle(gpu_index)
pynvml.nvmlDeviceSetDefaultFanSpeed_v2(handle, _FAN_INDEX) for i in range(_num_fans(handle)):
try:
pynvml.nvmlDeviceSetDefaultFanSpeed_v2(handle, i)
except pynvml.NVMLError as exc:
log.debug("reset_fan: fan %d: %s", i, exc)
return True, "OK" return True, "OK"
except pynvml.NVMLError as exc: except pynvml.NVMLError as exc:
return False, f"Failed to reset fan: {exc}" return False, f"Failed to reset fan: {exc}"
def get_temp(gpu_index: int = 0) -> Optional[float]: def get_temp(gpu_index: int = 0) -> float | None:
"""Read current GPU temperature in °C.""" """Read current GPU temperature in °C."""
if not _NVML_AVAILABLE: if not _NVML_AVAILABLE:
return None return None
try: try:
handle = _get_handle(gpu_index) handle = _get_handle(gpu_index)
return float(pynvml.nvmlDeviceGetTemperature(handle, pynvml.NVML_TEMPERATURE_GPU)) return float(
pynvml.nvmlDeviceGetTemperature(handle, pynvml.NVML_TEMPERATURE_GPU)
)
except pynvml.NVMLError as exc: except pynvml.NVMLError as exc:
log.debug("get_temp: %s", exc) log.debug("get_temp: %s", exc)
return None return None
def interpolate_fan_speed(curve: List[dict], temp_c: float) -> Optional[int]: def interpolate_fan_speed(curve: list[dict], temp_c: float) -> int | None:
"""Interpolate target fan speed from a curve at a given temperature. """Interpolate target fan speed from a curve at a given temperature.
curve: list of {temp_c: int, fan_pct: int} sorted by temp_c curve: list of {temp_c: int, fan_pct: int} sorted by temp_c
@@ -144,7 +250,10 @@ def interpolate_fan_speed(curve: List[dict], temp_c: float) -> Optional[int]:
if not curve or len(curve) < 2: if not curve or len(curve) < 2:
return None return None
temp = float(temp_c) try:
temp = float(temp_c)
except (TypeError, ValueError):
return None
# Find the two surrounding points # Find the two surrounding points
for i in range(len(curve) - 1): for i in range(len(curve) - 1):
@@ -157,7 +266,10 @@ def interpolate_fan_speed(curve: List[dict], temp_c: float) -> Optional[int]:
if t0 <= temp <= t1: if t0 <= temp <= t1:
fraction = (temp - t0) / (t1 - t0) fraction = (temp - t0) / (t1 - t0)
result = f0 + fraction * (f1 - f0) result = f0 + fraction * (f1 - f0)
return max(0, min(100, int(round(result)))) try:
return max(0, min(100, int(round(result))))
except (TypeError, ValueError):
return None
# Outside range: clamp to first or last point # Outside range: clamp to first or last point
if temp <= curve[0]["temp_c"]: if temp <= curve[0]["temp_c"]:
@@ -165,7 +277,7 @@ def interpolate_fan_speed(curve: List[dict], temp_c: float) -> Optional[int]:
return max(0, min(100, curve[-1]["fan_pct"])) return max(0, min(100, curve[-1]["fan_pct"]))
def validate_curve(curve: List[dict]) -> tuple[bool, str]: def validate_curve(curve: list[dict]) -> tuple[bool, str]:
"""Validate a fan curve. """Validate a fan curve.
Returns (True, "OK") or (False, error_message). Returns (True, "OK") or (False, error_message).
+38 -20
View File
@@ -1,12 +1,17 @@
"""GPU discovery and initialization.""" """GPU discovery and initialization."""
import contextlib
import ctypes import ctypes
import logging
import sys import sys
from typing import Any
from ..nvapi.bootstrap import query_interface from ..nvapi.bootstrap import query_interface
from ..nvapi.constants import FUNC from ..nvapi.constants import FUNC
from ..nvapi.types import GpuInfo from ..nvapi.types import GpuInfo
log = logging.getLogger("nvcurve.hal.gpu")
def init_nvapi() -> None: def init_nvapi() -> None:
"""Initialize NvAPI. Must be called before any GPU operations.""" """Initialize NvAPI. Must be called before any GPU operations."""
@@ -19,7 +24,10 @@ def enumerate_gpus() -> tuple[ctypes.Array, int]:
"""Return (gpu_handles_array, count). Exits if no GPUs found.""" """Return (gpu_handles_array, count). Exits if no GPUs found."""
gpus = (ctypes.c_void_p * 64)() gpus = (ctypes.c_void_p * 64)()
ngpu = ctypes.c_int32() ngpu = ctypes.c_int32()
query_interface(FUNC["EnumPhysicalGPUs"])(ctypes.byref(gpus), ctypes.byref(ngpu)) enum_fn = query_interface(FUNC["EnumPhysicalGPUs"])
if enum_fn is None:
raise RuntimeError("NvAPI function EnumPhysicalGPUs not available")
enum_fn(ctypes.byref(gpus), ctypes.byref(ngpu))
if ngpu.value == 0: if ngpu.value == 0:
print("No NVIDIA GPUs found") print("No NVIDIA GPUs found")
sys.exit(1) sys.exit(1)
@@ -29,7 +37,10 @@ def enumerate_gpus() -> tuple[ctypes.Array, int]:
def get_gpu_name(gpu) -> str: def get_gpu_name(gpu) -> str:
"""Return the full name string for a GPU handle.""" """Return the full name string for a GPU handle."""
name_buf = ctypes.create_string_buffer(256) name_buf = ctypes.create_string_buffer(256)
query_interface(FUNC["GetFullName"])(gpu, name_buf) fn = query_interface(FUNC["GetFullName"])
if fn is None:
raise RuntimeError("NvAPI function GetFullName not available")
fn(gpu, name_buf)
return name_buf.value.decode(errors="replace") return name_buf.value.decode(errors="replace")
@@ -40,38 +51,45 @@ def discover_gpus() -> list[GpuInfo]:
infos = [] infos = []
try: try:
import pynvml import pynvml as _pynvml
pynvml.nvmlInit()
has_nvml = True _pynvml.nvmlInit()
except Exception: except Exception:
has_nvml = False _pynvml = None
# Aliased as Any so attribute access is not flagged when the import failed.
pynvml: Any = _pynvml
for i in range(count): for i in range(count):
name = get_gpu_name(gpus[i]) name = get_gpu_name(gpus[i])
uuid = None uuid = None
pci_bus_id = None pci_bus_id = None
if has_nvml: if pynvml is not None:
try: try:
handle = pynvml.nvmlDeviceGetHandleByIndex(i) handle = pynvml.nvmlDeviceGetHandleByIndex(i)
uuid = pynvml.nvmlDeviceGetUUID(handle) raw_uuid = pynvml.nvmlDeviceGetUUID(handle)
# NVML might return bytes # NVML might return bytes
if isinstance(uuid, bytes): if isinstance(raw_uuid, bytes):
uuid = uuid.decode('utf-8', errors='ignore') uuid = raw_uuid.decode("utf-8", errors="ignore")
elif raw_uuid is not None:
uuid = str(raw_uuid)
pci_info = pynvml.nvmlDeviceGetPciInfo(handle) pci_info = pynvml.nvmlDeviceGetPciInfo(handle)
# Parse something like "00000000:01:00.0" -> bus is 1 # Parse something like "00000000:01:00.0" -> bus is 1.
if isinstance(pci_info.bus, bytes): # PCI bus numbers are hex by convention (pynvml's field is an
pci_bus_id = int(pci_info.bus.decode('utf-8', errors='ignore'), 16) # int; the str/bytes branches are defensive).
bus = pci_info.bus
if isinstance(bus, bytes):
pci_bus_id = int(bus.decode("utf-8", errors="ignore"), 16)
elif isinstance(bus, str):
pci_bus_id = int(bus, 16)
else: else:
pci_bus_id = pci_info.bus pci_bus_id = int(bus)
except Exception: except Exception as exc:
pass log.debug("NVML query for GPU %d failed: %s", i, exc)
infos.append(GpuInfo(name=name, index=i, uuid=uuid, pci_bus_id=pci_bus_id)) infos.append(GpuInfo(name=name, index=i, uuid=uuid, pci_bus_id=pci_bus_id))
if has_nvml: if pynvml is not None:
try: with contextlib.suppress(Exception):
pynvml.nvmlShutdown() pynvml.nvmlShutdown()
except Exception:
pass
return infos return infos
+125 -43
View File
@@ -11,21 +11,26 @@ that are explicitly specified, leaving others unchanged on hardware.
""" """
import ctypes import ctypes
import subprocess
import logging import logging
from typing import Optional import subprocess
from typing import Any
try: try:
import pynvml import pynvml as _pynvml_import
_NVML_AVAILABLE = True _NVML_AVAILABLE = True
except ImportError: except ImportError:
_pynvml_import = None
_NVML_AVAILABLE = False _NVML_AVAILABLE = False
# Aliased as Any so attribute access is not flagged when the import failed.
pynvml: Any = _pynvml_import
log = logging.getLogger("nvcurve.hal.limits") log = logging.getLogger("nvcurve.hal.limits")
# ── NVML library / handle helpers ───────────────────────────────────────────── # ── NVML library / handle helpers ─────────────────────────────────────────────
_nvml_lib: Optional[ctypes.CDLL] = None _nvml_lib: ctypes.CDLL | None = None
def _nvml_cdll() -> ctypes.CDLL: def _nvml_cdll() -> ctypes.CDLL:
@@ -34,12 +39,12 @@ def _nvml_cdll() -> ctypes.CDLL:
if _nvml_lib is not None: if _nvml_lib is not None:
return _nvml_lib return _nvml_lib
# Prefer to reuse the library already loaded by pynvml to avoid dlopen races. # Prefer to reuse the library already loaded by pynvml to avoid dlopen races.
for attr in ("nvml", "_nvml"): # attribute name varies by pynvml version for attr in ("nvml", "_nvml"): # attribute name varies by pynvml version
mod = getattr(pynvml, attr, None) mod = getattr(pynvml, attr, None)
lib = getattr(mod, "_lib", None) or getattr(mod, "_nvmlLib", None) lib = getattr(mod, "_lib", None) or getattr(mod, "_nvmlLib", None)
if lib is not None: if lib is not None:
_nvml_lib = lib _nvml_lib = lib
return _nvml_lib return lib
_nvml_lib = ctypes.CDLL("libnvidia-ml.so.1") _nvml_lib = ctypes.CDLL("libnvidia-ml.so.1")
return _nvml_lib return _nvml_lib
@@ -53,33 +58,79 @@ def _get_handle(gpu_index: int):
# ── Power limit ─────────────────────────────────────────────────────────────── # ── Power limit ───────────────────────────────────────────────────────────────
def get_power_limit(gpu_index: int = 0) -> dict:
"""Return dict with power_limit_w, default_power_limit_w, min_power_limit_w, max_power_limit_w.""" def get_power_limit(gpu_index: int = 0, mode: str = "nvml") -> dict:
out = { """Return dict with power limit info.
Keys: power_limit_w, default_power_limit_w, min_power_limit_w,
min_power_limit_w_native, max_power_limit_w, rm_power_supported,
power_cap_mode.
mode: "nvml" (default) or "ioctl" (experimental RM power control).
min_power_limit_w is the effective minimum: in ioctl mode it is
extended to the experimental floor (30 W) when the RM interface is
present and validated; min_power_limit_w_native is always the VBIOS
minimum. The RM probe is GET-only (no writes) and safe to run on
every call.
"""
out: dict[str, int | bool | str | None] = {
"power_limit_w": None, "power_limit_w": None,
"default_power_limit_w": None, "default_power_limit_w": None,
"min_power_limit_w": None, "min_power_limit_w": None,
"min_power_limit_w_native": None,
"max_power_limit_w": None, "max_power_limit_w": None,
"rm_power_supported": False,
"power_cap_mode": mode,
} }
try: try:
handle = _get_handle(gpu_index) handle = _get_handle(gpu_index)
limit = pynvml.nvmlDeviceGetPowerManagementLimit(handle) limit = pynvml.nvmlDeviceGetPowerManagementLimit(handle)
constrs = pynvml.nvmlDeviceGetPowerManagementLimitConstraints(handle) constrs = pynvml.nvmlDeviceGetPowerManagementLimitConstraints(handle)
out["power_limit_w"] = limit // 1000 out["power_limit_w"] = limit // 1000
out["min_power_limit_w"] = constrs[0] // 1000 native_min = constrs[0] // 1000
out["min_power_limit_w"] = native_min
out["min_power_limit_w_native"] = native_min
out["max_power_limit_w"] = constrs[1] // 1000 out["max_power_limit_w"] = constrs[1] // 1000
try: try:
default = pynvml.nvmlDeviceGetPowerManagementDefaultLimit(handle) default = pynvml.nvmlDeviceGetPowerManagementDefaultLimit(handle)
out["default_power_limit_w"] = default // 1000 out["default_power_limit_w"] = default // 1000
except Exception: except Exception as exc:
pass log.debug("nvmlDeviceGetPowerManagementDefaultLimit: %s", exc)
except Exception as exc: except Exception as exc:
log.warning("get_power_limit: %s", exc) log.warning("get_power_limit: %s", exc)
return out
# GET-only RM discovery — reported so the UI can offer the experimental
# mode; the effective minimum only changes in ioctl mode.
try:
from . import rm_power
bounds = rm_power.probe_gpu(gpu_index)
out["rm_power_supported"] = bounds is not None
if bounds is not None and mode == "ioctl":
out["min_power_limit_w"] = bounds.lower_min_mw() // 1000
except Exception as exc:
log.debug("RM power probe failed: %s", exc)
return out return out
def set_power_limit(limit_w: int, gpu_index: int = 0) -> tuple[bool, str]: def set_power_limit(
"""Set the board power limit (Watts).""" limit_w: int, gpu_index: int = 0, mode: str = "nvml"
) -> tuple[bool, str]:
"""Set the board power limit (Watts).
mode "ioctl" (experimental) applies the limit through the undocumented
RM interface, which permits values below the VBIOS minimum. It has no
fallback: failures are reported, never silently switched to NVML.
"""
if mode == "ioctl":
from . import rm_power
try:
rm_power.set_power_limit_w(gpu_index, limit_w)
return True, "OK"
except rm_power.RmPowerError as exc:
return False, str(exc)
try: try:
handle = _get_handle(gpu_index) handle = _get_handle(gpu_index)
pynvml.nvmlDeviceSetPowerManagementLimit(handle, limit_w * 1000) pynvml.nvmlDeviceSetPowerManagementLimit(handle, limit_w * 1000)
@@ -89,7 +140,8 @@ def set_power_limit(limit_w: int, gpu_index: int = 0) -> tuple[bool, str]:
ret = subprocess.run( ret = subprocess.run(
["nvidia-smi", "-i", str(gpu_index), "-pl", str(limit_w)], ["nvidia-smi", "-i", str(gpu_index), "-pl", str(limit_w)],
capture_output=True, text=True, capture_output=True,
text=True,
) )
if ret.returncode == 0: if ret.returncode == 0:
return True, "OK" return True, "OK"
@@ -111,34 +163,38 @@ def set_power_limit(limit_w: int, gpu_index: int = 0) -> tuple[bool, str]:
# pynvml (nvidia-ml-py ≥ 12) exposes c_nvmlClockOffset_t and nvmlClockOffset_v1 # pynvml (nvidia-ml-py ≥ 12) exposes c_nvmlClockOffset_t and nvmlClockOffset_v1
# as ctypes objects; we use them when available and fall back to our own definition. # as ctypes objects; we use them when available and fall back to our own definition.
class _ClockOffset(ctypes.Structure): class _ClockOffset(ctypes.Structure):
_fields_ = [ _fields_ = [
("version", ctypes.c_uint), ("version", ctypes.c_uint),
("type", ctypes.c_uint), # nvmlClockType_t ("type", ctypes.c_uint), # nvmlClockType_t
("pstate", ctypes.c_uint), # nvmlPstates_t ("pstate", ctypes.c_uint), # nvmlPstates_t
("clockOffsetMHz", ctypes.c_int), ("clockOffsetMHz", ctypes.c_int),
] ]
_CLOCK_OFFSET_VER = (1 << 24) | ctypes.sizeof(_ClockOffset) # = 0x01000010 (16 bytes) _CLOCK_OFFSET_VER = (1 << 24) | ctypes.sizeof(_ClockOffset) # = 0x01000010 (16 bytes)
# NVML clock-type constants (same values as pynvml). # NVML clock-type constants (same values as pynvml).
_NVML_CLOCK_GRAPHICS = 0 _NVML_CLOCK_GRAPHICS = 0
_NVML_CLOCK_MEM = 2 _NVML_CLOCK_MEM = 2
def _make_clock_offset(clock_type: int, pstate: int = 0, offset_mhz: int = 0) -> ctypes.Structure: def _make_clock_offset(
clock_type: int, pstate: int = 0, offset_mhz: int = 0
) -> ctypes.Structure:
"""Return a populated nvmlClockOffset_t struct, using pynvml's type when available.""" """Return a populated nvmlClockOffset_t struct, using pynvml's type when available."""
if hasattr(pynvml, "c_nvmlClockOffset_t") and hasattr(pynvml, "nvmlClockOffset_v1"): if hasattr(pynvml, "c_nvmlClockOffset_t") and hasattr(pynvml, "nvmlClockOffset_v1"):
info = pynvml.c_nvmlClockOffset_t() info = pynvml.c_nvmlClockOffset_t()
info.version = pynvml.nvmlClockOffset_v1 info.version = pynvml.nvmlClockOffset_v1
info.type = clock_type info.type = clock_type
info.pstate = pstate info.pstate = pstate
info.clockOffsetMHz = offset_mhz info.clockOffsetMHz = offset_mhz
return info return info
info = _ClockOffset() info = _ClockOffset()
info.version = _CLOCK_OFFSET_VER info.version = _CLOCK_OFFSET_VER
info.type = clock_type info.type = clock_type
info.pstate = pstate info.pstate = pstate
info.clockOffsetMHz = offset_mhz info.clockOffsetMHz = offset_mhz
return info return info
@@ -158,7 +214,7 @@ def get_clock_offsets(gpu_index: int = 0) -> dict:
Keys: gpc_offset_mhz, mem_offset_mhz (both int or None on failure). Keys: gpc_offset_mhz, mem_offset_mhz (both int or None on failure).
Calls nvmlDeviceGetClockOffsets once per clock domain (GRAPHICS, MEM). Calls nvmlDeviceGetClockOffsets once per clock domain (GRAPHICS, MEM).
""" """
out = {"gpc_offset_mhz": None, "mem_offset_mhz": None} out: dict[str, int | None] = {"gpc_offset_mhz": None, "mem_offset_mhz": None}
if not _NVML_AVAILABLE: if not _NVML_AVAILABLE:
return out return out
try: try:
@@ -167,11 +223,15 @@ def get_clock_offsets(gpu_index: int = 0) -> dict:
# Try pynvml wrapper first (nvidia-ml-py ≥ 12 exposes it correctly). # Try pynvml wrapper first (nvidia-ml-py ≥ 12 exposes it correctly).
# Fall back to ctypes-direct if pynvml doesn't have it. # Fall back to ctypes-direct if pynvml doesn't have it.
_pynvml_get = getattr(pynvml, "nvmlDeviceGetClockOffsets", None) _pynvml_get = getattr(pynvml, "nvmlDeviceGetClockOffsets", None)
fn_get = _try_nvml_fn("nvmlDeviceGetClockOffsets") if _pynvml_get is None else None fn_get = (
_try_nvml_fn("nvmlDeviceGetClockOffsets") if _pynvml_get is None else None
)
used_new_api = False used_new_api = False
for clock_type, key in ((_NVML_CLOCK_GRAPHICS, "gpc_offset_mhz"), for clock_type, key in (
(_NVML_CLOCK_MEM, "mem_offset_mhz")): (_NVML_CLOCK_GRAPHICS, "gpc_offset_mhz"),
(_NVML_CLOCK_MEM, "mem_offset_mhz"),
):
info = _make_clock_offset(clock_type, pstate=0) info = _make_clock_offset(clock_type, pstate=0)
try: try:
if _pynvml_get is not None: if _pynvml_get is not None:
@@ -184,7 +244,9 @@ def get_clock_offsets(gpu_index: int = 0) -> dict:
out[key] = int(info.clockOffsetMHz) out[key] = int(info.clockOffsetMHz)
used_new_api = True used_new_api = True
else: else:
log.debug("nvmlDeviceGetClockOffsets(type=%d) returned %d", clock_type, rc) log.debug(
"nvmlDeviceGetClockOffsets(type=%d) returned %d", clock_type, rc
)
except Exception as exc: except Exception as exc:
log.debug("nvmlDeviceGetClockOffsets(type=%d): %s", clock_type, exc) log.debug("nvmlDeviceGetClockOffsets(type=%d): %s", clock_type, exc)
@@ -200,7 +262,9 @@ def get_clock_offsets(gpu_index: int = 0) -> dict:
if hasattr(pynvml, "nvmlDeviceGetMemClkVfOffset"): if hasattr(pynvml, "nvmlDeviceGetMemClkVfOffset"):
try: try:
res = pynvml.nvmlDeviceGetMemClkVfOffset(handle) res = pynvml.nvmlDeviceGetMemClkVfOffset(handle)
out["mem_offset_mhz"] = int(res[0] if isinstance(res, (list, tuple)) else res) out["mem_offset_mhz"] = int(
res[0] if isinstance(res, (list, tuple)) else res
)
except Exception as exc: except Exception as exc:
log.debug("nvmlDeviceGetMemClkVfOffset: %s", exc) log.debug("nvmlDeviceGetMemClkVfOffset: %s", exc)
@@ -210,8 +274,8 @@ def get_clock_offsets(gpu_index: int = 0) -> dict:
def set_clock_offsets( def set_clock_offsets(
gpc_offset_mhz: Optional[int] = None, gpc_offset_mhz: int | None = None,
mem_offset_mhz: Optional[int] = None, mem_offset_mhz: int | None = None,
gpu_index: int = 0, gpu_index: int = 0,
) -> tuple[bool, str]: ) -> tuple[bool, str]:
"""Set clock offsets (MHz) for the specified domains only. """Set clock offsets (MHz) for the specified domains only.
@@ -235,20 +299,35 @@ def set_clock_offsets(
domains.append((_NVML_CLOCK_MEM, mem_offset_mhz)) domains.append((_NVML_CLOCK_MEM, mem_offset_mhz))
_pynvml_set = getattr(pynvml, "nvmlDeviceSetClockOffsets", None) _pynvml_set = getattr(pynvml, "nvmlDeviceSetClockOffsets", None)
fn_set = _try_nvml_fn("nvmlDeviceSetClockOffsets") if _pynvml_set is None else None fn_set = (
_try_nvml_fn("nvmlDeviceSetClockOffsets") if _pynvml_set is None else None
)
if _pynvml_set is not None or fn_set is not None: if _pynvml_set is not None or fn_set is not None:
all_ok = True all_ok = True
for clock_type, offset in domains: for clock_type, offset in domains:
info = _make_clock_offset(clock_type, pstate=0, offset_mhz=offset) info = _make_clock_offset(clock_type, pstate=0, offset_mhz=offset)
try: try:
rc = _pynvml_set(handle, ctypes.byref(info)) if _pynvml_set else fn_set(handle, ctypes.byref(info)) if _pynvml_set is not None:
rc = _pynvml_set(handle, ctypes.byref(info))
elif fn_set is not None:
rc = fn_set(handle, ctypes.byref(info))
else:
break
if rc != 0: if rc != 0:
log.debug("nvmlDeviceSetClockOffsets(type=%d) returned %d — trying fallback", clock_type, rc) log.debug(
"nvmlDeviceSetClockOffsets(type=%d) returned %d — trying fallback",
clock_type,
rc,
)
all_ok = False all_ok = False
break break
except Exception as exc: except Exception as exc:
log.debug("nvmlDeviceSetClockOffsets(type=%d): %s — trying fallback", clock_type, exc) log.debug(
"nvmlDeviceSetClockOffsets(type=%d): %s — trying fallback",
clock_type,
exc,
)
all_ok = False all_ok = False
break break
if all_ok: if all_ok:
@@ -257,12 +336,16 @@ def set_clock_offsets(
# Deprecated per-domain fallback (works on Blackwell/driver 590.x). # Deprecated per-domain fallback (works on Blackwell/driver 590.x).
errs = [] errs = []
if gpc_offset_mhz is not None and hasattr(pynvml, "nvmlDeviceSetGpcClkVfOffset"): if gpc_offset_mhz is not None and hasattr(
pynvml, "nvmlDeviceSetGpcClkVfOffset"
):
try: try:
pynvml.nvmlDeviceSetGpcClkVfOffset(handle, gpc_offset_mhz) pynvml.nvmlDeviceSetGpcClkVfOffset(handle, gpc_offset_mhz)
except Exception as exc: except Exception as exc:
errs.append(f"GPC: {exc}") errs.append(f"GPC: {exc}")
if mem_offset_mhz is not None and hasattr(pynvml, "nvmlDeviceSetMemClkVfOffset"): if mem_offset_mhz is not None and hasattr(
pynvml, "nvmlDeviceSetMemClkVfOffset"
):
try: try:
pynvml.nvmlDeviceSetMemClkVfOffset(handle, mem_offset_mhz) pynvml.nvmlDeviceSetMemClkVfOffset(handle, mem_offset_mhz)
except Exception as exc: except Exception as exc:
@@ -278,6 +361,7 @@ def set_clock_offsets(
# ── Range queries ───────────────────────────────────────────────────────────── # ── Range queries ─────────────────────────────────────────────────────────────
def get_mem_offset_range(gpu_index: int = 0) -> dict: def get_mem_offset_range(gpu_index: int = 0) -> dict:
"""Return the min/max allowed memory clock offset (MHz). """Return the min/max allowed memory clock offset (MHz).
@@ -285,7 +369,7 @@ def get_mem_offset_range(gpu_index: int = 0) -> dict:
Uses nvmlDeviceGetMemClkMinMaxVfOffset; falls back to observed RTX values. Uses nvmlDeviceGetMemClkMinMaxVfOffset; falls back to observed RTX values.
""" """
# Observed RTX 5090 defaults (NvAPI GetClockBoostRanges says -1000/+3000). # Observed RTX 5090 defaults (NvAPI GetClockBoostRanges says -1000/+3000).
out = {"min_mem_offset_mhz": -2000, "max_mem_offset_mhz": 3000} out: dict[str, int] = {"min_mem_offset_mhz": -2000, "max_mem_offset_mhz": 3000}
if not _NVML_AVAILABLE: if not _NVML_AVAILABLE:
return out return out
try: try:
@@ -317,5 +401,3 @@ def get_mem_offset_range(gpu_index: int = 0) -> dict:
except Exception as exc: except Exception as exc:
log.debug("get_mem_offset_range: %s", exc) log.debug("get_mem_offset_range: %s", exc)
return out return out
+14
View File
@@ -162,6 +162,20 @@ def _nvml_read(gpu_index: int) -> dict:
with contextlib.suppress(_pynvml.NVMLError): with contextlib.suppress(_pynvml.NVMLError):
out["fan_pct"] = float(_pynvml.nvmlDeviceGetFanSpeed(handle)) out["fan_pct"] = float(_pynvml.nvmlDeviceGetFanSpeed(handle))
# Per-fan speeds via the v2 API (fan_pct above stays fan 0 for legacy clients).
try:
num_fans = int(_pynvml.nvmlDeviceGetNumFans(handle))
fan_list: list[float | None] = []
for i in range(num_fans):
try:
fan_list.append(float(_pynvml.nvmlDeviceGetFanSpeed_v2(handle, i)))
except _pynvml.NVMLError:
fan_list.append(None)
if fan_list:
out["fans"] = fan_list
except _pynvml.NVMLError:
pass
with contextlib.suppress(_pynvml.NVMLError): with contextlib.suppress(_pynvml.NVMLError):
out["throttle_reasons"] = int( out["throttle_reasons"] = int(
_pynvml.nvmlDeviceGetCurrentClocksThrottleReasons(handle) _pynvml.nvmlDeviceGetCurrentClocksThrottleReasons(handle)
+237
View File
@@ -0,0 +1,237 @@
"""GPU process list — which processes are using VRAM on a GPU.
Per-PID VRAM and process type (compute/graphics) come from NVML; user,
CPU usage, host memory and the full command come from /proc. The server
runs as root, so it can read /proc entries of other users' processes.
"""
import os
import pwd
import time
from typing import Any
# NVML state lives in monitoring.py (nvmlInit is process-wide); import the
# module so attribute lookups see the current values, not import-time copies.
from . import monitoring as _mon
_BOOT_TIME: float | None = None
def _get_boot_time() -> float:
"""System boot time as a Unix timestamp (cached; read from /proc/stat)."""
global _BOOT_TIME
if _BOOT_TIME is None:
_BOOT_TIME = time.time() # fallback if /proc/stat is unreadable
try:
with open("/proc/stat") as f:
for line in f:
if line.startswith("btime "):
_BOOT_TIME = float(line.split()[1])
break
except OSError:
pass
return _BOOT_TIME
def _proc_info(pid: int) -> dict[str, Any] | None:
"""Read user, CPU usage, host memory and command for a pid from /proc.
CPU% is the average over the process lifetime (total jiffies / age),
which needs a single /proc read — no sampling state. Returns None when
the process no longer exists.
"""
base = f"/proc/{pid}"
try:
with open(f"{base}/stat", "rb") as f:
stat_raw = f.read().decode("ascii", "replace")
with open(f"{base}/statm") as f:
statm = f.read().split()
with open(f"{base}/cmdline", "rb") as f:
cmdline = f.read()
except OSError:
return None
# comm (field 2) may contain spaces and parentheses, so split at the
# last ')' — everything after it is fields 3..N.
lp = stat_raw.rfind(")")
if lp < 0:
return None
comm = stat_raw[stat_raw.index("(") + 1 : lp]
fields = stat_raw[lp + 2 :].split()
# field N -> fields[N-3]: utime=14, stime=15, starttime=22
if len(fields) < 20:
return None
try:
utime = int(fields[11])
stime = int(fields[12])
starttime = int(fields[19])
except ValueError:
return None
uid: int | None = None
try:
with open(f"{base}/status") as f:
for line in f:
if line.startswith("Uid:"):
uid = int(line.split()[1])
break
except OSError:
pass
try:
user = pwd.getpwuid(uid).pw_name if uid is not None else "unknown"
except KeyError:
user = str(uid)
hertz = os.sysconf("SC_CLK_TCK")
age_s = ((time.time() - _get_boot_time()) * hertz - starttime) / hertz
cpu_pct = (
round((utime + stime) / hertz / age_s * 100.0, 1) if age_s > 0 else None
)
page = os.sysconf("SC_PAGE_SIZE")
try:
mem_bytes = int(statm[1]) * page
except (IndexError, ValueError):
mem_bytes = None
parts = [p for p in cmdline.split(b"\0") if p]
command = " ".join(p.decode("utf-8", "replace") for p in parts) or comm
return {
"user": user,
"cpu_pct": cpu_pct,
"mem_bytes": mem_bytes,
"command": command,
}
def _nvml_process_lists(handle: Any) -> dict[int, dict[str, Any]]:
"""Map pid -> {"type": "C"|"G"|"C+G", "vram_bytes": int} for one GPU."""
out: dict[int, dict[str, Any]] = {}
pynvml = _mon._pynvml
try:
for p in pynvml.nvmlDeviceGetComputeRunningProcesses_v2(handle):
entry = out.setdefault(p.pid, {"type": "C", "vram_bytes": 0})
entry["vram_bytes"] = max(entry["vram_bytes"], p.usedGpuMemory or 0)
except pynvml.NVMLError:
pass
try:
for p in pynvml.nvmlDeviceGetGraphicsRunningProcesses_v2(handle):
entry = out.setdefault(p.pid, {"type": "G", "vram_bytes": 0})
entry["vram_bytes"] = max(entry["vram_bytes"], p.usedGpuMemory or 0)
if entry["type"] == "C":
entry["type"] = "C+G"
except pynvml.NVMLError:
pass
return out
def _nvml_process_utilization(
handle: Any,
) -> dict[int, tuple[float | None, float | None]]:
"""Map pid -> (gpu_util_pct, mem_util_pct).
Only processes that used the GPU within the last ~1 s are reported
(driver >= 525; older drivers raise NOT_SUPPORTED -> empty map).
"""
out: dict[int, tuple[float | None, float | None]] = {}
pynvml = _mon._pynvml
try:
for u in pynvml.nvmlDeviceGetProcessUtilization(handle, 0):
# Newer NVML renamed gpuUtil -> smUtil; accept both.
gpu_util = getattr(u, "smUtil", None)
if gpu_util is None:
gpu_util = getattr(u, "gpuUtil", None)
out[u.pid] = (
float(gpu_util) if gpu_util is not None else None,
float(u.memUtil) if u.memUtil is not None else None,
)
except pynvml.NVMLError:
pass
return out
def list_gpu_processes(gpu_index: int = 0) -> dict[str, Any]:
"""List processes using the given GPU, with per-PID details.
Returns {"processes": [...], "mem_total_bytes": int | None}. Processes
are sorted by VRAM usage (descending). Each process:
pid, user, dev (gpu_index), type ("C"/"G"/"C+G"), gpu_util_pct,
mem_util_pct, vram_bytes, vram_pct, cpu_pct, mem_host_bytes, command.
"""
from .monitoring import get_vram_total
mem_total = get_vram_total(gpu_index)
processes: list[dict[str, Any]] = []
if _mon._NVML_AVAILABLE and _mon._nvml_initialized:
pynvml = _mon._pynvml
try:
handle = pynvml.nvmlDeviceGetHandleByIndex(gpu_index)
nvml_procs = _nvml_process_lists(handle)
util = _nvml_process_utilization(handle)
except pynvml.NVMLError:
nvml_procs, util = {}, {}
for pid, info in sorted(
nvml_procs.items(), key=lambda kv: kv[1]["vram_bytes"], reverse=True
):
proc = _proc_info(pid)
if proc is None:
continue # process exited between the NVML read and /proc read
gpu_util, mem_util = util.get(pid, (None, None))
vram = info["vram_bytes"]
processes.append(
{
"pid": pid,
"user": proc["user"],
"dev": gpu_index,
"type": info["type"],
"gpu_util_pct": gpu_util,
"mem_util_pct": mem_util,
"vram_bytes": vram,
"vram_pct": (
round(vram / mem_total * 100.0, 1)
if mem_total and vram is not None
else None
),
"cpu_pct": proc["cpu_pct"],
"mem_host_bytes": proc["mem_bytes"],
"command": proc["command"],
}
)
return {"processes": processes, "mem_total_bytes": mem_total}
def kill_process(pid: int, sig: int) -> None:
"""Send a signal to a pid.
Raises ProcessLookupError if the pid does not exist, PermissionError if
the caller may not signal it.
"""
os.kill(pid, sig)
def get_parent_pid(pid: int) -> int | None:
"""Return the parent pid of a process.
Returns None when the process no longer exists, or 0 when it has no
parent (kernel threads, init).
"""
try:
with open(f"/proc/{pid}/stat") as f:
raw = f.read()
except OSError:
return None
lp = raw.rfind(")")
if lp < 0:
return None
fields = raw[lp + 2 :].split()
# field 4 (ppid) -> fields[1]
if len(fields) < 2:
return None
try:
return int(fields[1])
except ValueError:
return None
+627
View File
@@ -0,0 +1,627 @@
"""Undocumented NVIDIA RM power-limit interface (EXPERIMENTAL).
Port of the approach from LACT PR #1205 (ilya-zlobintsev/LACT): applies board
power limits through the private NV2080 power-limit "ordinary client"
interface on /dev/nvidiactl, which permits caps below the VBIOS minimum
(down to 30 W). The native maximum still applies.
EXPERIMENTAL — uses an undocumented driver interface. It may break after
driver updates. Discovery is GET-only and validates the RM payload against
NVML before any write is issued; a failed write restores the previous
request (even if it was below the VBIOS minimum).
"""
from __future__ import annotations
import contextlib
import ctypes
import fcntl
import logging
import os
import struct
import sys
from collections.abc import Callable
from dataclasses import dataclass
log = logging.getLogger("nvcurve.hal.rm_power")
# ── ioctl constants (nv-ioctl.h / nv-ioctl-numbers.h) ─────────────────────────
NV_IOCTL_MAGIC = ord("N") # 0x4E — user-space RM interface
NV_ESC_RM_ALLOC = 0x2B
NV_ESC_RM_CONTROL = 0x2A
# 'F' magic interface (kernel-open/common/inc/nv-ioctl-numbers.h) —
# NV_ESC_REGISTER_FD lives here, not in the 'N' RM interface.
NV_IOCTL_MAGIC_F = ord("F") # 0x46
NV_IOCTL_BASE_F = 200
NV_ESC_REGISTER_FD = NV_IOCTL_BASE_F + 1 # 201
# RM class IDs (nv0080.h / nv2080.h)
NV01_DEVICE_0 = 0x0080
NV20_SUBDEVICE_0 = 0x2080
# NV01_ROOT GPU queries (ctrl0000gpu.h) — resolve PCI identity to the RM
# device/subdevice instance numbers used by NV0080 and NV2080 allocations;
# neither number is a Linux device minor.
_CTRL_GPU_GET_ATTACHED_IDS = 0x201
_CTRL_GPU_GET_ID_INFO_V2 = 0x205
_CTRL_GPU_GET_PCI_INFO = 0x21B
_MAX_GPUS = 32
_INVALID_GPU_ID = 0xFFFFFFFF
# Private NV2080 power-limit client commands. Payloads compared against
# NvAPI and GSP from R595, R610 and R615 (native RM payloads, without
# NvAPI's 0x10-byte transport prefix).
_PWR_GET_INFO = 0x2080_A630
_PWR_GET_CONTROL = 0x2080_A632
_PWR_SET_CONTROL = 0x2080_E633
_ORDINARY_CLIENT = 0xFE
_LOWER_LIMIT_MW = 30_000 # experimental floor: 30 W
def _ioctl_rw(size: int, nr: int, magic: int = NV_IOCTL_MAGIC) -> int:
"""Linux ioctl request code: dir=RW, given size/type/nr."""
return (2 << 30) | (size << 16) | (magic << 8) | nr
def _ioctl_call(fd: int, code: int, arg) -> None:
"""Issue an ioctl, converting errno failures to RmPowerError.
The driver normally reports failures as an RM status in the parameter
struct, but an experimental interface can also fail at the kernel level
(ENOTTY/EBADF/EPERM across driver versions). Converting to RmPowerError
keeps the module's error contract uniform and lets callers clean up fds.
"""
try:
fcntl.ioctl(fd, code, arg)
except OSError as exc:
raise RmPowerError(f"ioctl 0x{code:x} failed: {exc}") from exc
# ── NVOS parameter structs (nvos.h) ──────────────────────────────────────────
class _NVOS21(ctypes.Structure):
_fields_ = [
("hRoot", ctypes.c_uint32),
("hObjectParent", ctypes.c_uint32),
("hObjectNew", ctypes.c_uint32),
("hClass", ctypes.c_uint32),
("pAllocParms", ctypes.c_uint64),
("paramsSize", ctypes.c_uint32),
("status", ctypes.c_uint32),
]
class _NVOS64(ctypes.Structure):
_fields_ = [
("hRoot", ctypes.c_uint32),
("hObjectParent", ctypes.c_uint32),
("hObjectNew", ctypes.c_uint32),
("hClass", ctypes.c_uint32),
("pAllocParms", ctypes.c_uint64),
("pRightsRequested", ctypes.c_uint64),
("paramsSize", ctypes.c_uint32),
("flags", ctypes.c_uint32),
("status", ctypes.c_uint32),
]
class _NVOS54(ctypes.Structure):
_fields_ = [
("hClient", ctypes.c_uint32),
("hObject", ctypes.c_uint32),
("cmd", ctypes.c_uint32),
("flags", ctypes.c_uint32),
("params", ctypes.c_uint64),
("paramsSize", ctypes.c_uint32),
("status", ctypes.c_uint32),
]
class _NV0080_ALLOC(ctypes.Structure):
_fields_ = [
("deviceId", ctypes.c_uint32),
("deviceFlags", ctypes.c_uint32),
("vgpuInstance", ctypes.c_uint32),
("pad", ctypes.c_uint32),
]
class _NV2080_ALLOC(ctypes.Structure):
_fields_ = [
("subDeviceId", ctypes.c_uint32),
("clientShare", ctypes.c_uint32),
("flags", ctypes.c_uint32),
("pad", ctypes.c_uint32),
]
# ── Errors ───────────────────────────────────────────────────────────────────
class RmPowerError(RuntimeError):
"""Raised when the RM power-limit interface is unavailable or fails."""
# ── Power-limit layouts and bounds ───────────────────────────────────────────
@dataclass(frozen=True)
class PowerLimitLayout:
"""Byte offsets of the private power-limit payloads for one wire format."""
name: str
info_size: int
control_size: int
info_min_at: int
request_at: int
client_at: int
mask_end: int
EXTENDED_LAYOUT = PowerLimitLayout(
name="extended",
info_size=0x924,
control_size=0x328,
info_min_at=0x28,
request_at=0x2C,
client_at=0x30,
mask_end=0x24,
)
LEGACY_LAYOUT = PowerLimitLayout(
name="legacy",
info_size=0x488,
control_size=0x188,
info_min_at=0xC,
request_at=0xC,
client_at=0x10,
mask_end=0x8,
)
@dataclass(frozen=True)
class PowerLimitBounds:
"""Power limit bounds in milliwatts (NVML/RM units)."""
min_mw: int
default_mw: int
max_mw: int
def lower_min_mw(self) -> int:
"""Effective minimum when the experimental route is active."""
return min(self.min_mw, _LOWER_LIMIT_MW)
@dataclass(frozen=True)
class LowerPowerLimit:
"""A validated RM power-limit layout that can be written."""
bounds: PowerLimitBounds
layout: PowerLimitLayout
def lower_min_mw(self) -> int:
return self.bounds.lower_min_mw()
# ── PCI identity → RM instance resolution ────────────────────────────────────
@dataclass(frozen=True)
class PciLocation:
domain: int
bus: int
dev: int
func: int = 0
def resolve_gpu_instance(
pci: PciLocation,
query: Callable[[int, bytearray], None],
) -> tuple[int, int]:
"""Resolve (device_instance, subdevice_instance) by PCI identity.
/dev/nvidiaN minors and RM device instances can have different orders;
the RM object must be matched by PCI domain/bus/slot, not by index.
The RM query exposes domain/bus/slot but no PCI function, so only
function-zero devices can be matched (never another function of a
multifunction device).
"""
if pci.func != 0:
raise RmPowerError("RM GPU lookup requires PCI function zero")
attached = bytearray(_MAX_GPUS * 4)
query(_CTRL_GPU_GET_ATTACHED_IDS, attached)
for i in range(_MAX_GPUS):
gpu_id = struct.unpack_from("<I", attached, i * 4)[0]
if gpu_id == _INVALID_GPU_ID:
continue
# NV0000_CTRL_GPU_GET_PCI_INFO_PARAMS: u32 gpuId, u32 domain,
# u16 bus, u16 slot.
location = bytearray(12)
location[0:4] = struct.pack("<I", gpu_id)
query(_CTRL_GPU_GET_PCI_INFO, location)
domain, bus, slot = struct.unpack_from("<IHH", location, 4)
if (domain, bus, slot) != (pci.domain, pci.bus, pci.dev):
continue
# NV0000_CTRL_GPU_GET_ID_INFO_V2_PARAMS: eight u32 fields, with
# deviceInstance/subDeviceInstance at +8/+12.
info = bytearray(32)
info[0:4] = struct.pack("<I", gpu_id)
query(_CTRL_GPU_GET_ID_INFO_V2, info)
device, subdevice = struct.unpack_from("<II", info, 8)
return device, subdevice
raise RmPowerError(f"no RM GPU matches PCI location {pci}")
# ── RM handle ────────────────────────────────────────────────────────────────
def _rm_control(fd: int, client: int, obj: int, cmd: int, buf: bytearray) -> None:
"""Issue an NVOS54 RM control whose parameter block is a byte buffer."""
arr = (ctypes.c_uint8 * len(buf)).from_buffer(buf)
req = _NVOS54(
hClient=client,
hObject=obj,
cmd=cmd,
flags=0,
params=ctypes.addressof(arr),
paramsSize=len(buf),
status=0,
)
_ioctl_call(fd, _ioctl_rw(ctypes.sizeof(_NVOS54), NV_ESC_RM_CONTROL), req)
if req.status != 0:
raise RmPowerError(
f"RM control 0x{cmd:08x} failed with status 0x{req.status:x}"
)
def _alloc_client(fd: int) -> int:
"""Allocate an RM client (NVOS21, all-zero parameters)."""
req = _NVOS21()
_ioctl_call(fd, _ioctl_rw(ctypes.sizeof(_NVOS21), NV_ESC_RM_ALLOC), req)
if req.status != 0:
raise RmPowerError(f"could not allocate RM client (status 0x{req.status:x})")
return req.hObjectNew
def _alloc_object(
fd: int, client: int, parent: int, class_id: int, alloc_params: ctypes.Structure
) -> int:
"""Allocate an RM object (NVOS64) and return its handle."""
req = _NVOS64(
hRoot=client,
hObjectParent=parent,
hObjectNew=0,
hClass=class_id,
pAllocParms=ctypes.addressof(alloc_params),
pRightsRequested=0,
paramsSize=ctypes.sizeof(alloc_params),
flags=0,
status=0,
)
_ioctl_call(fd, _ioctl_rw(ctypes.sizeof(_NVOS64), NV_ESC_RM_ALLOC), req)
if req.status != 0:
raise RmPowerError(
f"RM class 0x{class_id:x} allocation failed (status 0x{req.status:x})"
)
return req.hObjectNew
def _register_fd(device_fd: int, nvidiactl_fd: int) -> None:
"""Register the nvidiactl client with the device fd (NV_ESC_REGISTER_FD).
The ioctl is issued on the /dev/nvidiaN fd; the argument is the
nvidiactl fd to associate with it.
"""
_ioctl_call(
device_fd,
_ioctl_rw(4, NV_ESC_REGISTER_FD, NV_IOCTL_MAGIC_F),
struct.pack("i", nvidiactl_fd),
)
class RmHandle:
"""An NVIDIA RM client with device + subdevice objects for one GPU."""
def __init__(
self,
nvidiactl_fd: int,
device_fd: int,
client_handle: int,
device_handle: int,
subdevice_handle: int,
) -> None:
self._nvidiactl_fd = nvidiactl_fd
self._device_fd = device_fd
self.client_handle = client_handle
self.device_handle = device_handle
self.subdevice_handle = subdevice_handle
@classmethod
def open(cls, gpu_index: int) -> RmHandle:
"""Open an RM handle for the GPU at the given NVML index.
The RM device/subdevice instances are resolved by PCI identity
(minors and RM instances can have different orders).
"""
pynvml = _ensure_nvml()
try:
handle = pynvml.nvmlDeviceGetHandleByIndex(gpu_index)
minor = int(pynvml.nvmlDeviceGetMinorNumber(handle))
pci_info = pynvml.nvmlDeviceGetPciInfo(handle)
pci = PciLocation(
domain=int(pci_info.domain),
bus=int(pci_info.bus),
dev=int(pci_info.device),
)
except Exception as exc:
raise RmPowerError(f"NVML query for GPU {gpu_index} failed: {exc}") from exc
try:
nvidiactl_fd = os.open("/dev/nvidiactl", os.O_RDWR)
except OSError as exc:
raise RmPowerError(f"could not open /dev/nvidiactl: {exc}") from exc
try:
client_handle = _alloc_client(nvidiactl_fd)
device_instance, subdevice_instance = resolve_gpu_instance(
pci,
lambda cmd, buf: _rm_control(
nvidiactl_fd, client_handle, client_handle, cmd, buf
),
)
except RmPowerError:
os.close(nvidiactl_fd)
raise
try:
device_fd = os.open(f"/dev/nvidia{minor}", os.O_RDWR)
except OSError as exc:
os.close(nvidiactl_fd)
raise RmPowerError(f"could not open /dev/nvidia{minor}: {exc}") from exc
try:
_register_fd(device_fd, nvidiactl_fd)
device_handle = _alloc_object(
nvidiactl_fd,
client_handle,
client_handle,
NV01_DEVICE_0,
_NV0080_ALLOC(deviceId=device_instance),
)
subdevice_handle = _alloc_object(
nvidiactl_fd,
client_handle,
device_handle,
NV20_SUBDEVICE_0,
_NV2080_ALLOC(subDeviceId=subdevice_instance),
)
except RmPowerError:
os.close(device_fd)
os.close(nvidiactl_fd)
raise
return cls(
nvidiactl_fd, device_fd, client_handle, device_handle, subdevice_handle
)
def control(self, cmd: int, buf: bytearray) -> None:
"""Issue an NVOS54 RM control on the subdevice with a byte buffer."""
_rm_control(
self._nvidiactl_fd, self.client_handle, self.subdevice_handle, cmd, buf
)
def close(self) -> None:
"""Close the fds; the driver reclaims the RM client objects."""
with contextlib.suppress(OSError):
os.close(self._device_fd)
with contextlib.suppress(OSError):
os.close(self._nvidiactl_fd)
# ── Power-limit probe / set (pure logic, testable with a fake query) ─────────
def _u32(data: bytearray | bytes, offset: int) -> int:
return struct.unpack_from("<I", data, offset)[0]
def _validate_header(layout: PowerLimitLayout, data: bytearray) -> None:
if _u32(data, 0) != 0xFF or _u32(data, 4) != 1:
raise RmPowerError("unrecognized RM power client layout")
# The extended layout has additional mask words; accepting only its low
# word would allow an unexpected client to be included in a later SET.
if any(byte != 0 for byte in data[8 : layout.mask_end]):
raise RmPowerError("unrecognized RM power client layout")
def _read_bounds(
layout: PowerLimitLayout, query: Callable[[int, bytearray], None]
) -> PowerLimitBounds:
info = bytearray(layout.info_size)
query(_PWR_GET_INFO, info)
_validate_header(layout, info)
bounds = PowerLimitBounds(
min_mw=_u32(info, layout.info_min_at),
default_mw=_u32(info, layout.info_min_at + 4),
max_mw=_u32(info, layout.info_min_at + 8),
)
if not (
bounds.min_mw > 0
and bounds.min_mw <= bounds.default_mw
and bounds.default_mw <= bounds.max_mw
):
raise RmPowerError("invalid RM power limit bounds")
return bounds
def _read_control(
layout: PowerLimitLayout, query: Callable[[int, bytearray], None]
) -> bytearray:
control = bytearray(layout.control_size)
control[4:8] = struct.pack("<I", 1)
control[layout.client_at] = _ORDINARY_CLIENT
query(_PWR_GET_CONTROL, control)
_validate_header(layout, control)
if control[layout.client_at] != _ORDINARY_CLIENT:
raise RmPowerError("unexpected power client")
if _u32(control, layout.request_at) in (0, 0xFFFFFFFF):
raise RmPowerError("no ordinary power request available")
return control
def probe(
nvml_bounds: PowerLimitBounds,
nvml_current_mw: int,
query: Callable[[int, bytearray], None],
) -> LowerPowerLimit:
"""GET-only discovery of the RM power-limit layout.
Probes the two known wire formats using GETs only. A driver version
number is not evidence that the payload still has the same layout or
units, so the bounds and the current request are validated against
NVML. Discovery never issues a SET.
"""
if sys.byteorder != "little":
raise RmPowerError("little-endian host required")
errors: list[str] = []
for layout in (EXTENDED_LAYOUT, LEGACY_LAYOUT):
try:
bounds = _read_bounds(layout, query)
if bounds != nvml_bounds:
raise RmPowerError("RM power bounds differ from NVML")
control = _read_control(layout, query)
if _u32(control, layout.request_at) != nvml_current_mw:
raise RmPowerError("RM ordinary power request differs from NVML")
return LowerPowerLimit(bounds=bounds, layout=layout)
except RmPowerError as exc:
errors.append(f"{layout.name}: {exc}")
raise RmPowerError("no compatible RM power layout: " + "; ".join(errors))
def set_limit(
limit_mw: int,
support: LowerPowerLimit,
query: Callable[[int, bytearray], None],
) -> None:
"""Set the ordinary-client power request with readback verification.
Keeps the entire current payload, changing only entry 0's request.
Mask 1 and selector 0xFE prevent modifying any other entry or the
additional F8 client. A failed SET can have side effects, so the
previous request is restored even on transport failure — and the
restore uses 0xFE so a previous limit below the VBIOS minimum can
also be restored.
"""
layout = support.layout
bounds = _read_bounds(layout, query)
if bounds != support.bounds:
raise RmPowerError("RM power bounds changed since discovery")
lower = bounds.lower_min_mw()
if not (lower <= limit_mw <= bounds.max_mw):
raise RmPowerError(
f"power limit {limit_mw} mW outside supported range "
f"{lower}..{bounds.max_mw} mW"
)
before = _read_control(layout, query)
if _u32(before, layout.request_at) == limit_mw:
return
expected = bytearray(before)
expected[layout.request_at : layout.request_at + 4] = struct.pack("<I", limit_mw)
try:
request = bytearray(expected)
query(_PWR_SET_CONTROL, request)
if _read_control(layout, query) != expected:
raise RmPowerError("power request readback differs")
except RmPowerError as apply_error:
try:
restore = bytearray(before)
query(_PWR_SET_CONTROL, restore)
if _read_control(layout, query) != before:
raise RmPowerError("restored power request differs")
except RmPowerError as restore_error:
raise RmPowerError(
f"power request failed: {apply_error}; "
f"restoration also failed: {restore_error}"
) from None
raise RmPowerError(f"{apply_error} (previous power request restored)") from None
# ── High-level API (wires NVML state + RmHandle to the pure logic) ───────────
def _ensure_nvml():
"""Return pynvml with NVML initialized (nvmlInit is refcounted)."""
import pynvml
pynvml.nvmlInit()
return pynvml
def _nvml_power_state(gpu_index: int) -> tuple[PowerLimitBounds, int]:
"""Return (bounds, current_mw) from NVML for the given GPU."""
pynvml = _ensure_nvml()
try:
handle = pynvml.nvmlDeviceGetHandleByIndex(gpu_index)
min_mw, max_mw = pynvml.nvmlDeviceGetPowerManagementLimitConstraints(handle)
default_mw = pynvml.nvmlDeviceGetPowerManagementDefaultLimit(handle)
current_mw = pynvml.nvmlDeviceGetPowerManagementLimit(handle)
bounds = PowerLimitBounds(int(min_mw), int(default_mw), int(max_mw))
current = int(current_mw)
except Exception as exc:
raise RmPowerError(
f"NVML power state for GPU {gpu_index} unavailable: {exc}"
) from exc
return bounds, current
def probe_gpu(gpu_index: int = 0) -> PowerLimitBounds | None:
"""GET-only discovery of the RM power-limit interface for a GPU.
Returns the validated power bounds (milliwatts) when a compatible RM
layout is present, else None. Never issues a write.
"""
try:
bounds, current = _nvml_power_state(gpu_index)
except Exception as exc:
log.debug("RM probe: NVML state unavailable: %s", exc)
return None
try:
handle = RmHandle.open(gpu_index)
except RmPowerError as exc:
log.debug("RM probe: handle open failed: %s", exc)
return None
try:
probe(bounds, current, handle.control)
return bounds
except RmPowerError as exc:
log.debug("RM probe: %s", exc)
return None
finally:
handle.close()
def set_power_limit_w(gpu_index: int, limit_w: int) -> None:
"""Set the board power limit (watts) via the RM interface.
Raises RmPowerError on any failure (probe, range, write, readback).
A failed write restores the previous request.
"""
bounds, current = _nvml_power_state(gpu_index)
handle = RmHandle.open(gpu_index)
try:
support = probe(bounds, current, handle.control)
set_limit(int(limit_w) * 1000, support, handle.control)
finally:
handle.close()
+70 -30
View File
@@ -2,18 +2,20 @@
import ctypes import ctypes
import json import json
import logging
import os import os
import struct import struct
from datetime import datetime from datetime import datetime
from typing import Optional
from ..nvapi.bootstrap import nvcall_raw from ..nvapi.bootstrap import nvcall_raw
from ..nvapi.constants import FUNC, CT_SIZE, CT_BASE, CT_STRIDE, CT_DELTA_OFF, CT_POINTS from ..nvapi.constants import CT_BASE, CT_DELTA_OFF, CT_SIZE, CT_STRIDE, FUNC
from ..nvapi.types import SnapshotInfo from ..nvapi.types import SnapshotInfo
from .vfcurve import read_clock_table_raw, get_boost_mask from .vfcurve import get_boost_mask, read_clock_table_raw
log = logging.getLogger("nvcurve.hal.snapshot")
def save(gpu, gpu_name: str, snapshot_dir: str, max_snapshots: int = 0) -> Optional[str]: def save(gpu, gpu_name: str, snapshot_dir: str, max_snapshots: int = 0) -> str | None:
"""Save the current ClockBoostTable to disk. """Save the current ClockBoostTable to disk.
Writes both a binary .bin file and a human-readable .json metadata file. Writes both a binary .bin file and a human-readable .json metadata file.
@@ -25,13 +27,21 @@ def save(gpu, gpu_name: str, snapshot_dir: str, max_snapshots: int = 0) -> Optio
print(f"Failed to read ClockBoostTable: {err}") print(f"Failed to read ClockBoostTable: {err}")
return None return None
os.makedirs(snapshot_dir, exist_ok=True) try:
os.makedirs(snapshot_dir, exist_ok=True)
except OSError as exc:
print(f"Failed to create snapshot dir {snapshot_dir}: {exc}")
return None
ts = datetime.now().strftime("%Y%m%d_%H%M%S") ts = datetime.now().strftime("%Y%m%d_%H%M%S")
bin_path = os.path.join(snapshot_dir, f"clock_boost_table_{ts}.bin") bin_path = os.path.join(snapshot_dir, f"clock_boost_table_{ts}.bin")
meta_path = os.path.join(snapshot_dir, f"clock_boost_table_{ts}.json") meta_path = os.path.join(snapshot_dir, f"clock_boost_table_{ts}.json")
with open(bin_path, "wb") as f: try:
f.write(raw) with open(bin_path, "wb") as f:
f.write(raw)
except OSError as exc:
print(f"Failed to write snapshot {bin_path}: {exc}")
return None
offsets = [] offsets = []
max_entries = (len(raw) - CT_BASE) // CT_STRIDE max_entries = (len(raw) - CT_BASE) // CT_STRIDE
@@ -48,10 +58,14 @@ def save(gpu, gpu_name: str, snapshot_dir: str, max_snapshots: int = 0) -> Optio
"offsets_kHz": offsets, "offsets_kHz": offsets,
"nonzero_offsets": sum(1 for o in offsets if o != 0), "nonzero_offsets": sum(1 for o in offsets if o != 0),
} }
with open(meta_path, "w") as f: try:
json.dump(meta, f, indent=2) with open(meta_path, "w") as f:
json.dump(meta, f, indent=2)
except OSError as exc:
print(f"Failed to write snapshot metadata {meta_path}: {exc}")
return None
print(f"Snapshot saved:") print("Snapshot saved:")
print(f" Binary: {bin_path}") print(f" Binary: {bin_path}")
print(f" Metadata: {meta_path}") print(f" Metadata: {meta_path}")
print(f" Size: {len(raw)} bytes") print(f" Size: {len(raw)} bytes")
@@ -65,20 +79,22 @@ def save(gpu, gpu_name: str, snapshot_dir: str, max_snapshots: int = 0) -> Optio
def _prune_snapshots(snapshot_dir: str, max_snapshots: int) -> None: def _prune_snapshots(snapshot_dir: str, max_snapshots: int) -> None:
"""Delete oldest snapshots (both .bin and .json) to stay within max_snapshots.""" """Delete oldest snapshots (both .bin and .json) to stay within max_snapshots."""
bins = sorted( # Oldest first (lexicographic = chronological for our timestamp format).
f for f in os.listdir(snapshot_dir) if f.endswith(".bin") try:
) # oldest first (lexicographic = chronological for our timestamp format) bins = sorted(f for f in os.listdir(snapshot_dir) if f.endswith(".bin"))
except OSError:
return
excess = len(bins) - max_snapshots excess = len(bins) - max_snapshots
for fname in bins[:excess]: for fname in bins[:excess]:
stem = fname[:-4] # strip .bin stem = fname[:-4] # strip .bin
for ext in (".bin", ".json"): for ext in (".bin", ".json"):
try: try:
os.remove(os.path.join(snapshot_dir, stem + ext)) os.remove(os.path.join(snapshot_dir, stem + ext))
except OSError: except OSError as exc:
pass log.debug("Could not remove %s: %s", stem + ext, exc)
def restore(gpu, snapshot_dir: str, filepath: str = None) -> bool: def restore(gpu, snapshot_dir: str, filepath: str | None = None) -> bool:
"""Restore a ClockBoostTable snapshot from disk. """Restore a ClockBoostTable snapshot from disk.
If no filepath is given, uses the most recent snapshot in snapshot_dir. If no filepath is given, uses the most recent snapshot in snapshot_dir.
@@ -88,10 +104,14 @@ def restore(gpu, snapshot_dir: str, filepath: str = None) -> bool:
if not os.path.isdir(snapshot_dir): if not os.path.isdir(snapshot_dir):
print(f"No snapshots found in {snapshot_dir}") print(f"No snapshots found in {snapshot_dir}")
return False return False
bins = sorted( try:
[f for f in os.listdir(snapshot_dir) if f.endswith(".bin")], bins = sorted(
reverse=True, [f for f in os.listdir(snapshot_dir) if f.endswith(".bin")],
) reverse=True,
)
except OSError:
print(f"No snapshots found in {snapshot_dir}")
return False
if not bins: if not bins:
print(f"No snapshot .bin files in {snapshot_dir}") print(f"No snapshot .bin files in {snapshot_dir}")
return False return False
@@ -101,8 +121,21 @@ def restore(gpu, snapshot_dir: str, filepath: str = None) -> bool:
print(f"Snapshot file not found: {filepath}") print(f"Snapshot file not found: {filepath}")
return False return False
with open(filepath, "rb") as f: # Contain the path inside the snapshot directory — callers (in
raw = f.read() # particular the HTTP API) must not be able to point the restore at
# arbitrary files on the filesystem.
snap_dir = os.path.realpath(snapshot_dir)
resolved = os.path.realpath(filepath)
if not resolved.startswith(snap_dir + os.sep):
print(f"Snapshot path outside snapshot directory: {filepath}")
return False
try:
with open(filepath, "rb") as f:
raw = f.read()
except OSError as exc:
print(f"Failed to read snapshot {filepath}: {exc}")
return False
if len(raw) != CT_SIZE: if len(raw) != CT_SIZE:
print(f"Snapshot size mismatch: expected {CT_SIZE}, got {len(raw)}") print(f"Snapshot size mismatch: expected {CT_SIZE}, got {len(raw)}")
@@ -134,8 +167,13 @@ def list_snapshots(snapshot_dir: str) -> list[SnapshotInfo]:
if not os.path.isdir(snapshot_dir): if not os.path.isdir(snapshot_dir):
return [] return []
try:
fnames = sorted(os.listdir(snapshot_dir), reverse=True)
except OSError:
return []
results = [] results = []
for fname in sorted(os.listdir(snapshot_dir), reverse=True): for fname in fnames:
if not fname.endswith(".json"): if not fname.endswith(".json"):
continue continue
meta_path = os.path.join(snapshot_dir, fname) meta_path = os.path.join(snapshot_dir, fname)
@@ -143,13 +181,15 @@ def list_snapshots(snapshot_dir: str) -> list[SnapshotInfo]:
with open(meta_path) as f: with open(meta_path) as f:
meta = json.load(f) meta = json.load(f)
bin_path = meta.get("file", meta_path.replace(".json", ".bin")) bin_path = meta.get("file", meta_path.replace(".json", ".bin"))
results.append(SnapshotInfo( results.append(
filepath=bin_path, SnapshotInfo(
timestamp=meta.get("timestamp", ""), filepath=bin_path,
gpu=meta.get("gpu", ""), timestamp=meta.get("timestamp", ""),
nonzero_offsets=meta.get("nonzero_offsets", 0), gpu=meta.get("gpu", ""),
size=meta.get("size", 0), nonzero_offsets=meta.get("nonzero_offsets", 0),
)) size=meta.get("size", 0),
)
)
except (json.JSONDecodeError, KeyError): except (json.JSONDecodeError, KeyError):
continue continue
+1 -1
View File
@@ -9,7 +9,7 @@ from .errors import NVAPI_ERRORS
def load_nvapi() -> ctypes.CDLL: def load_nvapi() -> ctypes.CDLL:
"""Load libnvidia-api.so from the NVIDIA driver.""" """Load libnvidia-api.so from the NVIDIA driver."""
for name in ("libnvidia-api.so", "libnvidia-api.so.1"): for name in ("libnvidia-api.so", "libnvidia-api.so.1"): # gitleaks:allow
try: try:
return ctypes.CDLL(name) return ctypes.CDLL(name)
except OSError: except OSError:
+1
View File
@@ -61,6 +61,7 @@ class MonitoringSample:
pcie_link_width: int | None = None # Current PCIe link width (x1..x16) pcie_link_width: int | None = None # Current PCIe link width (x1..x16)
pcie_link_generation: int | None = None # Current PCIe link generation (1..5) pcie_link_generation: int | None = None # Current PCIe link generation (1..5)
mem_temp_c: float | None = None # VRAM temperature (if the GPU exposes it) mem_temp_c: float | None = None # VRAM temperature (if the GPU exposes it)
fans: list[float | None] | None = None # Per-fan speed % (index-aligned)
@dataclass @dataclass
+91 -39
View File
@@ -21,12 +21,12 @@ def _gpu_stable_key(info) -> str:
def apply_profile(gpu_index: int, name: str, cfg) -> list[str]: def apply_profile(gpu_index: int, name: str, cfg) -> list[str]:
"""Apply a named profile to the given GPU. Returns a list of error strings.""" """Apply a named profile to the given GPU. Returns a list of error strings."""
from .native import load_profile
from ..hal.gpu import get_gpu from ..hal.gpu import get_gpu
from ..hal.limits import set_clock_offsets, set_power_limit from ..hal.limits import set_clock_offsets, set_power_limit
from ..hal.vfcurve import write_offsets, reset_offsets
from ..hal.snapshot import save as snapshot_save from ..hal.snapshot import save as snapshot_save
from ..hal.vfcurve import reset_offsets, write_offsets
from ..safety import validate_write from ..safety import validate_write
from .native import load_profile
safe_name = "".join(c for c in name if c.isalnum() or c in " _-()").strip() safe_name = "".join(c for c in name if c.isalnum() or c in " _-()").strip()
filepath = os.path.join(cfg.profile_dir, f"{safe_name}.json") filepath = os.path.join(cfg.profile_dir, f"{safe_name}.json")
@@ -43,24 +43,31 @@ def apply_profile(gpu_index: int, name: str, cfg) -> list[str]:
errs.append(f"Mem offset: {msg}") errs.append(f"Mem offset: {msg}")
if profile.power_limit_w is not None: if profile.power_limit_w is not None:
ok, msg = set_power_limit(profile.power_limit_w, gpu_index) mode = profile.power_cap_mode or "nvml"
ok, msg = set_power_limit(profile.power_limit_w, gpu_index, mode)
if not ok: if not ok:
errs.append(f"Power limit: {msg}") errs.append(f"Power limit: {msg}")
if profile.curve_deltas: if profile.curve_deltas:
deltas = {int(k): v for k, v in profile.curve_deltas.items()} try:
errors = validate_write(deltas, cfg.max_delta_khz) deltas = {int(k): v for k, v in profile.curve_deltas.items()}
if errors: except ValueError:
errs.append("Curve: " + "; ".join(errors)) errs.append("Curve: invalid point keys in profile")
else: else:
if cfg.auto_snapshot: errors = validate_write(deltas, cfg.max_delta_khz)
try: if errors:
snapshot_save(gpu, gpu_name, cfg.snapshot_dir, cfg.max_snapshots) errs.append("Curve: " + "; ".join(errors))
except Exception as exc: else:
log.warning("Auto-snapshot failed: %s", exc) if cfg.auto_snapshot:
ret, desc = write_offsets(gpu, deltas) try:
if ret != 0: snapshot_save(
errs.append(f"Curve write failed ({ret}): {desc}") gpu, gpu_name, cfg.snapshot_dir, cfg.max_snapshots
)
except Exception as exc:
log.warning("Auto-snapshot failed: %s", exc)
ret, desc = write_offsets(gpu, deltas)
if ret != 0:
errs.append(f"Curve write failed ({ret}): {desc}")
else: else:
reset_offsets(gpu) reset_offsets(gpu)
@@ -69,9 +76,9 @@ def apply_profile(gpu_index: int, name: str, cfg) -> list[str]:
def apply_with_retry(gpu_index: int, name: str, cfg, max_retries: int = 3) -> bool: def apply_with_retry(gpu_index: int, name: str, cfg, max_retries: int = 3) -> bool:
"""Apply a named profile with read-back verification, retrying on mismatch.""" """Apply a named profile with read-back verification, retrying on mismatch."""
from .native import load_profile
from ..hal.gpu import get_gpu from ..hal.gpu import get_gpu
from ..hal.vfcurve import read_clock_offsets from ..hal.vfcurve import read_clock_offsets
from .native import load_profile
safe_name = "".join(c for c in name if c.isalnum() or c in " _-()").strip() safe_name = "".join(c for c in name if c.isalnum() or c in " _-()").strip()
filepath = os.path.join(cfg.profile_dir, f"{safe_name}.json") filepath = os.path.join(cfg.profile_dir, f"{safe_name}.json")
@@ -82,51 +89,88 @@ def apply_with_retry(gpu_index: int, name: str, cfg, max_retries: int = 3) -> bo
log.warning("Auto-load profile %r not found — skipping GPU %d", name, gpu_index) log.warning("Auto-load profile %r not found — skipping GPU %d", name, gpu_index)
return False return False
expected: dict[int, int] = ( try:
{int(k): v for k, v in profile.curve_deltas.items()} expected: dict[int, int] = (
if profile.curve_deltas else {} {int(k): v for k, v in profile.curve_deltas.items()}
) if profile.curve_deltas
else {}
)
except ValueError:
log.warning(
"Profile %r has invalid curve point keys — skipping GPU %d",
name,
gpu_index,
)
return False
for attempt in range(max_retries): for attempt in range(max_retries):
try: try:
errs = apply_profile(gpu_index, name, cfg) errs = apply_profile(gpu_index, name, cfg)
except Exception as exc: except Exception as exc:
log.warning("Auto-load attempt %d/%d exception: %s", attempt + 1, max_retries, exc) log.warning(
"Auto-load attempt %d/%d exception: %s", attempt + 1, max_retries, exc
)
errs = [str(exc)] errs = [str(exc)]
if errs: if errs:
log.warning("Auto-load attempt %d/%d errors: %s", log.warning(
attempt + 1, max_retries, "; ".join(errs)) "Auto-load attempt %d/%d errors: %s",
attempt + 1,
max_retries,
"; ".join(errs),
)
elif expected: elif expected:
gpu, _ = get_gpu(index=gpu_index) gpu, _ = get_gpu(index=gpu_index)
offsets, err = read_clock_offsets(gpu) offsets, err = read_clock_offsets(gpu)
if offsets is None: if offsets is None:
log.warning("Auto-load attempt %d/%d: read-back failed: %s", log.warning(
attempt + 1, max_retries, err) "Auto-load attempt %d/%d: read-back failed: %s",
attempt + 1,
max_retries,
err,
)
else: else:
mismatches = [ mismatches = [
f"pt{idx}: expected {val/1000:+.0f}MHz got {offsets[idx]/1000:+.0f}MHz" f"pt{idx}: expected {val / 1000:+.0f}MHz got {offsets[idx] / 1000:+.0f}MHz"
for idx, val in expected.items() for idx, val in expected.items()
if idx < len(offsets) and offsets[idx] != val if idx < len(offsets) and offsets[idx] != val
] ]
if not mismatches: if not mismatches:
log.info("Auto-load profile %r verified on GPU %d (attempt %d/%d)", log.info(
name, gpu_index, attempt + 1, max_retries) "Auto-load profile %r verified on GPU %d (attempt %d/%d)",
name,
gpu_index,
attempt + 1,
max_retries,
)
return True return True
log.warning("Auto-load attempt %d/%d: read-back mismatch — %s", log.warning(
attempt + 1, max_retries, "; ".join(mismatches)) "Auto-load attempt %d/%d: read-back mismatch — %s",
attempt + 1,
max_retries,
"; ".join(mismatches),
)
else: else:
log.info("Auto-load profile %r applied on GPU %d (attempt %d/%d)", log.info(
name, gpu_index, attempt + 1, max_retries) "Auto-load profile %r applied on GPU %d (attempt %d/%d)",
name,
gpu_index,
attempt + 1,
max_retries,
)
return True return True
if attempt < max_retries - 1: if attempt < max_retries - 1:
delay = 2 ** attempt # 1s, 2s, 4s delay = 2**attempt # 1s, 2s, 4s
log.info("Retrying auto-load in %ds…", delay) log.info("Retrying auto-load in %ds…", delay)
time.sleep(delay) time.sleep(delay)
log.warning("Auto-load profile %r failed after %d attempts on GPU %d", log.warning(
name, max_retries, gpu_index) "Auto-load profile %r failed after %d attempts on GPU %d",
name,
max_retries,
gpu_index,
)
return False return False
@@ -153,13 +197,19 @@ def run_autoload() -> None:
return return
from ..config import Config from ..config import Config
cfg = Config() cfg = Config()
for key in ("max_delta_khz", "auto_snapshot", "max_snapshots", for key in (
"snapshot_dir", "profile_dir"): "max_delta_khz",
"auto_snapshot",
"max_snapshots",
"snapshot_dir",
"profile_dir",
):
if key in cfg_data: if key in cfg_data:
setattr(cfg, key, cfg_data[key]) setattr(cfg, key, cfg_data[key])
from ..hal.gpu import init_nvapi, discover_gpus from ..hal.gpu import discover_gpus, init_nvapi
from ..hal.monitoring import init_nvml, shutdown_nvml from ..hal.monitoring import init_nvml, shutdown_nvml
# Retry NvAPI init — the driver may not be fully ready at early boot. # Retry NvAPI init — the driver may not be fully ready at early boot.
@@ -189,7 +239,9 @@ def run_autoload() -> None:
if gpu_idx is None: if gpu_idx is None:
log.warning("Auto-load: no GPU found with key %r — skipping", gpu_key) log.warning("Auto-load: no GPU found with key %r — skipping", gpu_key)
continue continue
log.info("Auto-loading profile %r on GPU %d (%s)", profile_name, gpu_idx, gpu_key) log.info(
"Auto-loading profile %r on GPU %d (%s)", profile_name, gpu_idx, gpu_key
)
apply_with_retry(gpu_idx, profile_name, cfg) apply_with_retry(gpu_idx, profile_name, cfg)
shutdown_nvml() shutdown_nvml()
+38 -18
View File
@@ -1,49 +1,70 @@
"""Native profile storage and schema.""" """Native profile storage and schema."""
import json
import os
import glob import glob
from dataclasses import dataclass, asdict import json
from typing import Dict, Optional, List import logging
import os
from dataclasses import asdict, dataclass
log = logging.getLogger("nvcurve.profiles.native")
@dataclass @dataclass
class ProfileData: class ProfileData:
name: str name: str
gpu_name: str gpu_name: str
curve_deltas: Dict[str, int] # { "index": delta_khz } curve_deltas: dict[str, int] # { "index": delta_khz }
mem_offset_mhz: Optional[int] = None mem_offset_mhz: int | None = None
power_limit_w: Optional[int] = None power_limit_w: int | None = None
fan_curve: Optional[List[Dict[str, int]]] = None # How power_limit_w is applied: "nvml" (default) or "ioctl" (experimental
# RM power control, permits values below the VBIOS minimum).
power_cap_mode: str | None = None
fan_curve: list[dict[str, int]] | None = None
# Fan indices controlled by fan_curve (0-based); None = all fans.
fan_targets: list[int] | None = None
def save_profile(profile_dir: str, data: ProfileData) -> str: def save_profile(profile_dir: str, data: ProfileData) -> str:
"""Save profile to JSON, sanitising the filename.""" """Save profile to JSON, sanitising the filename."""
os.makedirs(profile_dir, exist_ok=True) try:
os.makedirs(profile_dir, exist_ok=True)
except OSError as exc:
raise RuntimeError(f"Cannot create profile dir {profile_dir}: {exc}") from exc
safe_name = "".join(c for c in data.name if c.isalnum() or c in " _-()").strip() safe_name = "".join(c for c in data.name if c.isalnum() or c in " _-()").strip()
if not safe_name: if not safe_name:
safe_name = "Unnamed" safe_name = "Unnamed"
filepath = os.path.join(profile_dir, f"{safe_name}.json") filepath = os.path.join(profile_dir, f"{safe_name}.json")
with open(filepath, "w", encoding="utf-8") as f: try:
json.dump(asdict(data), f, indent=2) with open(filepath, "w", encoding="utf-8") as f:
json.dump(asdict(data), f, indent=2)
except OSError as exc:
raise RuntimeError(f"Cannot write profile {filepath}: {exc}") from exc
return filepath return filepath
def load_profile(filepath: str) -> ProfileData: def load_profile(filepath: str) -> ProfileData:
"""Load profile from JSON.""" """Load profile from JSON."""
with open(filepath, "r", encoding="utf-8") as f: try:
data = json.load(f) with open(filepath, encoding="utf-8") as f:
data = json.load(f)
except FileNotFoundError:
raise
except (OSError, json.JSONDecodeError) as exc:
raise RuntimeError(f"Cannot read profile {filepath}: {exc}") from exc
# Migrate old field names. # Migrate old field names.
if "vram_p0_offset_mhz" in data and "mem_offset_mhz" not in data: if "vram_p0_offset_mhz" in data and "mem_offset_mhz" not in data:
data["mem_offset_mhz"] = data.pop("vram_p0_offset_mhz") data["mem_offset_mhz"] = data.pop("vram_p0_offset_mhz")
# Drop removed fields so old profiles don't cause TypeError. # Drop removed fields so old profiles don't cause TypeError.
for obsolete in ("gpu_locked_min_mhz", "gpu_locked_max_mhz", "vram_p0_offset_mhz"): for obsolete in ("gpu_locked_min_mhz", "gpu_locked_max_mhz", "vram_p0_offset_mhz"):
data.pop(obsolete, None) data.pop(obsolete, None)
# Normalize the experimental power-cap mode; unknown values fall back to NVML.
if data.get("power_cap_mode") not in (None, "nvml", "ioctl"):
data["power_cap_mode"] = None
return ProfileData(**data) return ProfileData(**data)
def list_profiles(profile_dir: str) -> List[ProfileData]: def list_profiles(profile_dir: str) -> list[ProfileData]:
"""Return a list of all safely readable profiles.""" """Return a list of all safely readable profiles."""
if not os.path.exists(profile_dir): if not os.path.exists(profile_dir):
return [] return []
@@ -51,9 +72,8 @@ def list_profiles(profile_dir: str) -> List[ProfileData]:
for fp in glob.glob(os.path.join(profile_dir, "*.json")): for fp in glob.glob(os.path.join(profile_dir, "*.json")):
try: try:
profiles.append(load_profile(fp)) profiles.append(load_profile(fp))
except Exception as e: except Exception as exc:
# log warning ideally, but swallowing for robustness log.debug("Skipping unreadable profile %s: %s", fp, exc)
pass
# Sort alphabetically by name # Sort alphabetically by name
profiles.sort(key=lambda p: p.name.lower()) profiles.sort(key=lambda p: p.name.lower())
return profiles return profiles
+937 -82
View File
File diff suppressed because it is too large. Load diff
+691
View File
@@ -0,0 +1,691 @@
"""Native reader for the Thermal Grizzly WireView Pro II.
Talks to the 12 VHPWR connector monitor directly over its USB CDC/ACM
serial port (STM32, VID 0483 / PID 5740) — no exporter, no GUI, no kernel
module required. When the wireview-hwmon kernel module is loaded, the
sysfs node is used instead (the wireviewd daemon owns the port in that
case, so direct serial would corrupt frames).
Protocol notes (matches the firmware's DEVICE_STR_LEN=32 layout):
* No framing/CRC — a desynced read corrupts arbitrary fields for one
poll. Real frames always carry zero padding bytes and a fan duty
<= 100; anything else is discarded and the next poll realigns.
* The firmware occasionally stops answering the RTS welcome handshake
(observed after USB state changes) while still answering every data
command, so identification falls back to the vendor-data reply.
"""
import logging
import os
import struct
import time
from typing import Any
try:
import serial
except ImportError: # pyserial missing — the serial transport is disabled,
# but the rest of the server keeps running (a stale venv after a code
# update must not take the whole web server down).
serial = None
log = logging.getLogger("nvcurve.wireview")
_serial_missing_warned = False
# ── Device identification ─────────────────────────────────────────────────────
USB_VENDOR_ID = "0483" # STMicroelectronics (CDC/ACM)
USB_PRODUCT_ID = "5740" # WireView Pro II normal mode
WELCOME_MESSAGE = "Thermal Grizzly WireView Pro II"
MAX_WELCOME_LENGTH = 64
# Vendor/product ids reported by CMD_READ_VENDOR_DATA (not the USB ids).
VENDOR_ID_THERMAL_GRIZZLY = 0xEF
PRODUCT_ID_PRO2 = 0x05
PRODUCT_ID_PRO2_NOCTUA = 0x06
DEVICE_NAMES = {
PRODUCT_ID_PRO2: "WireView Pro II",
PRODUCT_ID_PRO2_NOCTUA: "WireView Pro II Noctua Edition",
}
BAUD_RATE = 115200
READ_TIMEOUT_S = 1.0
# ── Serial protocol commands ──────────────────────────────────────────────────
CMD_READ_VENDOR_DATA = 0x01
CMD_READ_UID = 0x02
CMD_READ_SENSOR_VALUES = 0x04
CMD_READ_CONFIG = 0x05
CMD_SCREEN_CHANGE = 0x0C
CMD_READ_BUILD_INFO = 0x0D
SCREEN_RESUME_UPDATES = 0xF1
# ── Wire layout ───────────────────────────────────────────────────────────────
# SensorStruct (100 bytes, little-endian, pack=4):
# 4x int16 temperatures (0.1 °C): in, out, ext1, ext2
# uint16 Vdd (mV)
# uint8 fan duty (%)
# pad
# 6x { int16 voltage (mV), pad, uint32 current (mA), uint32 power (mW) }
# uint32 total power (mW)
# uint32 total current (mA)
# uint16 avg voltage (mV)
# uint8 PSU capability (0=600W, 1=450W, 2=300W, 3=150W)
# pad
# uint16 fault status mask
# uint16 fault log mask
SENSOR_STRUCT = struct.Struct(
"<4hHBx" + "".join("hxxII" for _ in range(6)) + "IIHBxHH"
)
SENSOR_STRUCT_SIZE = SENSOR_STRUCT.size # 100
# BuildStruct: VendorData(3) + ProductName(32) + BuildInfo(32) + NameLength(1)
BUILD_STRUCT_SIZE = 3 + 32 + 32 + 1
BUILD_INFO_OFFSET = 3 + 32 # 35
PSU_CAPABILITY_W = {0: 600, 1: 450, 2: 300, 3: 150}
# Fault bitmask (both the active status and the latched log use these bits).
FAULT_BITS = {
0: "Chip over-temperature",
1: "Sensor over-temperature",
2: "Over-current (OCP)",
3: "Wire over-current",
4: "Over-power (OPP)",
5: "Current imbalance",
}
def decode_faults(mask: int) -> list[str]:
"""Human-readable names of the active fault bits in a status/log mask."""
return [
FAULT_BITS[bit]
for bit in sorted(FAULT_BITS)
if mask & (1 << bit)
]
def is_supported_product(vendor_id: int, product_id: int) -> bool:
"""True for the products the Pro II protocol serves (5 and 6)."""
return (
vendor_id == VENDOR_ID_THERMAL_GRIZZLY
and product_id in (PRODUCT_ID_PRO2, PRODUCT_ID_PRO2_NOCTUA)
)
def parse_sensor_struct(buf: bytes) -> dict:
"""Decode a 100-byte sensor frame into a JSON-serializable sample.
Totals are computed from the per-pin readings (voltage * current),
matching the exporter's output; the device's own total fields are
not used.
"""
fields = SENSOR_STRUCT.unpack(buf)
ts_in, ts_out, ts_ext1, ts_ext2, _vdd, fan_duty = fields[:6]
pin_fields = fields[6:24]
(
_total_power,
_total_current,
_avg_voltage,
psu_cap,
fault_status,
fault_log,
) = fields[24:]
pins = []
power_total = 0.0
current_total = 0.0
for i in range(6):
voltage_v = pin_fields[i * 3] / 1000.0
current_a = pin_fields[i * 3 + 1] / 1000.0
power_w = pin_fields[i * 3 + 2] / 1000.0
pins.append(
{
"voltage_v": round(voltage_v, 3),
"current_a": round(current_a, 3),
"power_w": round(power_w, 3),
}
)
power_total += voltage_v * current_a
current_total += current_a
# Spread between the highest- and lowest-loaded pin; a growing spread
# indicates a bad pin contact (see the frontend imbalance warning).
pin_currents = [p["current_a"] for p in pins]
pin_imbalance = round(max(pin_currents) - min(pin_currents), 3)
return {
"timestamp": time.time(),
"power_total_w": round(power_total, 3),
"current_total_a": round(current_total, 3),
"pin_imbalance_a": pin_imbalance,
"voltage_avg_v": (
round(power_total / current_total, 3) if current_total > 0 else 0.0
),
"pins": pins,
"temp_in_c": ts_in / 10.0,
"temp_out_c": ts_out / 10.0,
"temp_ext1_c": ts_ext1 / 10.0,
"temp_ext2_c": ts_ext2 / 10.0,
"fan_duty_pct": fan_duty,
"fault_status": fault_status,
"fault_log": fault_log,
"psu_capability_w": PSU_CAPABILITY_W.get(psu_cap, 0),
}
def sensor_frame_is_corrupt(buf: bytes) -> bool:
"""Corruption check for a sensor frame: real frames always carry zero
padding bytes and a fan duty <= 100. Wire layout: fan duty at offset
10, pad1 at 11, pad2 five bytes from the end (before the two 16-bit
fault masks)."""
if len(buf) < SENSOR_STRUCT_SIZE:
return True
return buf[10] > 100 or buf[11] != 0 or buf[SENSOR_STRUCT_SIZE - 5] != 0
# ── Device discovery ──────────────────────────────────────────────────────────
def _sysfs_matches_wireview(tty_sysfs_dir: str) -> bool:
"""Walk up from a resolved tty sysfs path to the USB device and check
its idVendor/idProduct."""
path = tty_sysfs_dir
while path and path != "/":
vid_file = os.path.join(path, "idVendor")
pid_file = os.path.join(path, "idProduct")
if os.path.isfile(vid_file) and os.path.isfile(pid_file):
try:
with open(vid_file) as f:
vid = f.read().strip().lower()
with open(pid_file) as f:
pid = f.read().strip().lower()
except OSError:
return False
return vid == USB_VENDOR_ID and pid == USB_PRODUCT_ID
path = os.path.dirname(path)
return False
def find_wireview_ports() -> list[str]:
"""Find /dev nodes of connected WireView Pro II devices.
Checks the stable /dev/wireview-pro2 symlink (created by the
99-wireview.rules udev rule) and falls back to a sysfs scan of all
ttyACM* ports matched by USB VID/PID.
"""
ports: set[str] = set()
link = "/dev/wireview-pro2"
if os.path.islink(link) or os.path.exists(link):
try:
target = os.path.realpath(link)
if os.path.exists(target):
ports.add(target)
except OSError:
pass
sys_class = "/sys/class/tty"
if os.path.isdir(sys_class):
try:
entries = os.listdir(sys_class)
except OSError:
entries = []
for entry in entries:
if not entry.startswith("ttyACM"):
continue
tty_dir = os.path.join(sys_class, entry)
try:
resolved = os.path.realpath(tty_dir)
except OSError:
continue
if _sysfs_matches_wireview(resolved):
ports.add(f"/dev/{entry}")
return sorted(ports)
def find_hwmon_path() -> str | None:
"""Find the wireview-hwmon sysfs node, if the kernel module is loaded."""
base = "/sys/class/hwmon"
if not os.path.isdir(base):
return None
try:
entries = os.listdir(base)
except OSError:
return None
for entry in entries:
name_path = os.path.join(base, entry, "name")
try:
with open(name_path) as f:
if f.read().strip().lower() == "wireview":
return os.path.join(base, entry)
except OSError:
continue
return None
# ── Serial transport ──────────────────────────────────────────────────────────
class WireViewSerialDevice:
"""Direct serial access to a WireView Pro II.
The port is opened and closed per transaction (open → flush → write →
read → close), matching the proven behavior of the exporter: it keeps
the port unheld between polls so other tools (official GUI, wireviewd)
can share the device, and a fresh open realigns a desynced stream.
"""
def __init__(self, port: str, baud: int = BAUD_RATE) -> None:
self._port = port
self._baud = baud
self._connected = False
self._rejected = False
self._vendor_id = 0
self._product_id = 0
self._firmware_version = ""
self._uid = ""
self._build = ""
self._config_version = -1
# ── Identity ──
@property
def connected(self) -> bool:
return self._connected
@property
def rejected(self) -> bool:
"""True when the device reported an unsupported product id. Callers
can memoize this so the port is not re-probed (and re-logged) on
every watchdog tick."""
return self._rejected
@property
def transport(self) -> str:
return "serial"
@property
def port(self) -> str:
return self._port
def info(self) -> dict:
return {
"device_name": DEVICE_NAMES.get(
self._product_id, "WireView Pro II"
),
"hw_rev": f"{self._vendor_id:02X}{self._product_id:02X}",
"firmware_version": self._firmware_version,
"uid": self._uid,
"build": self._build,
"transport": self.transport,
"port": self._port,
}
def node_exists(self) -> bool:
"""Whether the serial device node still exists (unplug check)."""
return os.path.exists(self._port)
# ── Connection ──
def connect(self) -> bool:
"""Identify the device and prepare it for sensor reads.
The welcome handshake (RTS edge) is the primary identification,
but the firmware occasionally stops answering it while still
answering every command — a supported vendor-data reply is
equally conclusive, so accept either.
"""
if self._connected:
return True
# The welcome handshake (RTS edge) is the primary identification, but
# the firmware occasionally stops answering it while still answering
# every command — the supported vendor-data reply below is equally
# conclusive, so the welcome is read for logging only.
welcome = self._read_welcome()
if welcome and welcome != WELCOME_MESSAGE:
log.debug(
"WireView: unexpected welcome string %r on %s",
welcome,
self._port,
)
vd = self._transaction(bytes([CMD_READ_VENDOR_DATA]), 3)
if vd is None or len(vd) < 3:
return False
vendor, product, fw = vd[0], vd[1], vd[2]
if not is_supported_product(vendor, product):
self._rejected = True
log.info(
"WireView: unsupported product %02X%02X on %s, skipped",
vendor,
product,
self._port,
)
return False
self._vendor_id = vendor
self._product_id = product
self._firmware_version = str(fw)
cfg = self._transaction(bytes([CMD_READ_CONFIG]), 4)
if cfg is None or len(cfg) < 3:
return False
self._config_version = cfg[2]
uid = self._transaction(bytes([CMD_READ_UID]), 12)
if uid is not None and len(uid) == 12:
self._uid = uid.hex().upper()
# Enable display updates just in case.
self._transaction(
bytes([CMD_SCREEN_CHANGE, SCREEN_RESUME_UPDATES]), 0
)
build = self._transaction(bytes([CMD_READ_BUILD_INFO]), BUILD_STRUCT_SIZE)
if build is not None and len(build) >= BUILD_INFO_OFFSET + 1:
self._build = (
build[BUILD_INFO_OFFSET : BUILD_INFO_OFFSET + 32]
.split(b"\x00")[0]
.decode("ascii", errors="replace")
)
self._connected = True
log.info(
"WireView connected on %s (%s, fw %s)",
self._port,
self.info()["hw_rev"],
self._firmware_version,
)
return True
def close(self) -> None:
self._connected = False
# ── Sensor reads ──
def read_sample(self) -> dict | None:
"""Read one sensor sample, or None when the device is unresponsive
or the frame is corrupt."""
if not self._connected:
return None
buf = self._transaction(bytes([CMD_READ_SENSOR_VALUES]), SENSOR_STRUCT_SIZE)
if buf is None or sensor_frame_is_corrupt(buf):
return None
return parse_sensor_struct(buf)
# ── Transport ──
def _open_port(self):
"""Open the serial port, or None when pyserial is missing or the
port is unavailable."""
global _serial_missing_warned
if serial is None:
if not _serial_missing_warned:
_serial_missing_warned = True
log.warning(
"pyserial is not installed — WireView serial transport "
"disabled (install pyserial, e.g. `uv sync`)"
)
return None
try:
return serial.Serial(self._port, self._baud, timeout=READ_TIMEOUT_S)
except OSError: # SerialException is an OSError
return None
def _read_welcome(self) -> str | None:
"""Assert RTS and read the NUL-terminated welcome string the device
answers with. Null when nothing (or no terminator) arrives in time."""
ser = self._open_port()
if ser is None:
return None
try:
ser.reset_input_buffer()
ser.rts = False
time.sleep(0.01)
ser.rts = True
time.sleep(0.01)
buf = bytearray()
deadline = time.monotonic() + READ_TIMEOUT_S
while len(buf) < MAX_WELCOME_LENGTH:
remaining = deadline - time.monotonic()
if remaining <= 0:
break
ser.timeout = min(READ_TIMEOUT_S, remaining)
chunk = ser.read(MAX_WELCOME_LENGTH - len(buf))
if not chunk:
break
buf.extend(chunk)
if b"\x00" in chunk:
break
time.sleep(0.01)
ser.rts = False
nul = buf.find(b"\x00")
if nul >= 0:
return bytes(buf[:nul]).decode("ascii", errors="replace")
return None
except OSError: # SerialException is an OSError
return None
finally:
ser.close()
def _transaction(self, cmd: bytes, response_size: int) -> bytes | None:
"""Open the port, send cmd, read exactly response_size bytes (one
second budget), close the port. None when the port is unavailable
or the reply is incomplete."""
ser = self._open_port()
if ser is None:
return None
try:
ser.reset_input_buffer()
if cmd:
ser.write(cmd)
if response_size == 0:
return b""
return self._read_exact(ser, response_size)
except OSError: # SerialException is an OSError
return None
finally:
ser.close()
@staticmethod
def _read_exact(ser: Any, size: int) -> bytes | None:
"""Read exactly size bytes within one second, or None."""
buf = bytearray()
deadline = time.monotonic() + READ_TIMEOUT_S
while len(buf) < size:
remaining = deadline - time.monotonic()
if remaining <= 0:
return None
ser.timeout = min(READ_TIMEOUT_S, remaining)
chunk = ser.read(size - len(buf))
if not chunk:
return None
buf.extend(chunk)
return bytes(buf)
# ── hwmon (sysfs) transport ───────────────────────────────────────────────────
class WireViewHwmonDevice:
"""Reads the wireview-hwmon sysfs node (kernel module + wireviewd).
Used when the module is loaded: the daemon owns the serial port in
that case, so direct serial would corrupt frames.
"""
def __init__(self, hwmon_path: str) -> None:
self._path = hwmon_path
self._connected = False
@property
def connected(self) -> bool:
return self._connected
@property
def rejected(self) -> bool:
return False
@property
def transport(self) -> str:
return "hwmon"
@property
def port(self) -> str:
return self._path
def info(self) -> dict:
return {
"device_name": "WireView Pro II",
"hw_rev": "",
"firmware_version": "",
"uid": "",
"build": "",
"transport": self.transport,
"port": self._path,
}
def node_exists(self) -> bool:
return os.path.isdir(self._path)
def connect(self) -> bool:
if self._connected:
return True
name_path = os.path.join(self._path, "name")
try:
with open(name_path) as f:
if f.read().strip().lower() != "wireview":
return False
# Probe that the node actually serves data.
with open(os.path.join(self._path, "in0_input")) as f:
f.read().strip()
except OSError:
return False
self._connected = True
log.info("WireView connected via hwmon (%s)", self._path)
return True
def close(self) -> None:
self._connected = False
def read_sample(self) -> dict | None:
if not self._connected:
return None
try:
pin_voltage = [
self._read_int(f"in{i}_input") / 1000.0 for i in range(6)
]
pin_current = [
self._read_int(f"curr{i + 1}_input") / 1000.0 for i in range(6)
]
temp_in = self._read_temp("temp1_input")
temp_out = self._read_temp("temp2_input")
temp_ext1 = self._read_temp("temp3_input")
temp_ext2 = self._read_temp("temp4_input")
fault_status = self._read_int_or("fault_status_raw")
if fault_status is None:
fault_status = 0xFFFF if self._read_int("intrusion0_alarm") else 0
fault_log = self._read_int_or("fault_log_raw")
if fault_log is None:
fault_log = 0xFFFF if self._read_int("intrusion1_alarm") else 0
psu_cap_uw = self._read_int_or("power1_cap")
if psu_cap_uw is not None:
psu_capability = int(round(psu_cap_uw / 1_000_000.0))
else:
psu_cap = self._read_int_or("psu_cap")
psu_capability = PSU_CAPABILITY_W.get(psu_cap or 0, 0)
pwm = self._read_int_or("pwm1")
if pwm is not None:
fan_duty = int(round(min(255, max(0, pwm)) * 100 / 255.0))
else:
fan_duty = self._read_int("fan1_input")
power_total = sum(
v * i
for v, i in zip(pin_voltage, pin_current, strict=True)
)
current_total = sum(pin_current)
pin_currents = [round(i, 3) for i in pin_current]
pin_imbalance = round(max(pin_currents) - min(pin_currents), 3)
return {
"timestamp": time.time(),
"power_total_w": round(power_total, 3),
"current_total_a": round(current_total, 3),
"pin_imbalance_a": pin_imbalance,
"voltage_avg_v": (
round(power_total / current_total, 3)
if current_total > 0
else 0.0
),
"pins": [
{
"voltage_v": round(v, 3),
"current_a": round(i, 3),
"power_w": round(v * i, 3),
}
for v, i in zip(pin_voltage, pin_current, strict=True)
],
"temp_in_c": temp_in,
"temp_out_c": temp_out,
"temp_ext1_c": temp_ext1,
"temp_ext2_c": temp_ext2,
"fan_duty_pct": fan_duty,
"fault_status": fault_status,
"fault_log": fault_log,
"psu_capability_w": psu_capability,
}
except OSError:
return None
def _read_int(self, filename: str) -> int:
"""Read an integer sysfs attribute; 0 when missing or unreadable
(matches the exporter's ReadIntFile)."""
try:
with open(os.path.join(self._path, filename)) as f:
return int(f.read().strip())
except (OSError, ValueError):
return 0
def _read_int_or(self, filename: str) -> int | None:
try:
with open(os.path.join(self._path, filename)) as f:
return int(f.read().strip())
except (OSError, ValueError):
return None
def _read_temp(self, filename: str) -> float | None:
"""Read a temperature sysfs attribute (m°C); None when missing or
unreadable. NaN would poison the whole sample: the WebSocket
serializer emits a bare NaN token (invalid JSON) and the REST
JSONResponse rejects it with a 500."""
try:
with open(os.path.join(self._path, filename)) as f:
return int(f.read().strip()) / 1000.0
except (OSError, ValueError):
return None
def create_device() -> WireViewSerialDevice | WireViewHwmonDevice | None:
"""Create a device for the first available WireView, or None.
Preference: the wireview-hwmon sysfs node (the wireviewd daemon owns
the serial port in that case), otherwise direct serial on the first
matching /dev/ttyACM*.
"""
hwmon_path = find_hwmon_path()
if hwmon_path:
return WireViewHwmonDevice(hwmon_path)
ports = find_wireview_ports()
if ports:
return WireViewSerialDevice(ports[0])
return None
+12
View File
@@ -14,14 +14,26 @@ dependencies = [
"pydantic>=2.0", "pydantic>=2.0",
"httpx>=0.27", "httpx>=0.27",
"bcrypt>=4.0", "bcrypt>=4.0",
"pyserial>=3.5",
] ]
[project.scripts] [project.scripts]
nvcurve = "nvcurve.cli:main" nvcurve = "nvcurve.cli:main"
[dependency-groups]
dev = [
"hatchling", # enables local `hatch build` and resolves hatch_build.py imports
]
[tool.hatch.build.hooks.custom]
[tool.hatch.build.targets.wheel] [tool.hatch.build.targets.wheel]
packages = ["nvcurve"] packages = ["nvcurve"]
# The custom build hook (hatch_build.py) compiles the React frontend when
# frontend/dist is missing or stale, so `uv tool install git+<repo-url>` works
# as a single command. It runs for both wheel and sdist builds.
[tool.hatch.build.targets.wheel.force-include] [tool.hatch.build.targets.wheel.force-include]
"frontend/dist" = "nvcurve/frontend/dist" "frontend/dist" = "nvcurve/frontend/dist"
+267 -166
View File
@@ -55,20 +55,20 @@ Key findings:
See NvAPI_VF_Curve_Documentation.md for full technical details. See NvAPI_VF_Curve_Documentation.md for full technical details.
""" """
import argparse
import ctypes import ctypes
import struct
import sys
import json import json
import os import os
import struct
import sys
import time import time
import argparse
from datetime import datetime from datetime import datetime
from typing import Optional, List, Tuple, Set, Dict
# ═══════════════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════════════
# NvAPI bootstrap # NvAPI bootstrap
# ═══════════════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════════════
def load_nvapi(): def load_nvapi():
"""Load libnvidia-api.so from the NVIDIA driver.""" """Load libnvidia-api.so from the NVIDIA driver."""
for name in ("libnvidia-api.so", "libnvidia-api.so.1"): for name in ("libnvidia-api.so", "libnvidia-api.so.1"):
@@ -145,23 +145,20 @@ def nvcall_raw(fid: int, gpu, buf: ctypes.Array):
FUNC = { FUNC = {
# Bootstrap # Bootstrap
"Initialize": 0x0150E828, "Initialize": 0x0150E828,
"EnumPhysicalGPUs": 0xE5AC921F, "EnumPhysicalGPUs": 0xE5AC921F,
"GetFullName": 0xCEEE8E9F, "GetFullName": 0xCEEE8E9F,
# V/F curve (read) # V/F curve (read)
"GetVFPCurve": 0x21537AD4, # ClkVfPointsGetStatus "GetVFPCurve": 0x21537AD4, # ClkVfPointsGetStatus
"GetClockBoostMask": 0x507B4B59, # ClkVfPointsGetInfo "GetClockBoostMask": 0x507B4B59, # ClkVfPointsGetInfo
"GetClockBoostTable": 0x23F1B133, # ClkVfPointsGetControl "GetClockBoostTable": 0x23F1B133, # ClkVfPointsGetControl
"GetCurrentVoltage": 0x465F9BCF, # ClientVoltRailsGetStatus "GetCurrentVoltage": 0x465F9BCF, # ClientVoltRailsGetStatus
"GetClockBoostRanges": 0x64B43A6A, # ClkDomainsGetInfo "GetClockBoostRanges": 0x64B43A6A, # ClkDomainsGetInfo
# Additional read # Additional read
"GetPerfLimits": 0xE440B867, # PerfClientLimitsGetStatus "GetPerfLimits": 0xE440B867, # PerfClientLimitsGetStatus
"GetVoltBoostPercent": 0x9DF23CA1, # ClientVoltRailsGetControl "GetVoltBoostPercent": 0x9DF23CA1, # ClientVoltRailsGetControl
# Write # Write
"SetClockBoostTable": 0x0733E009, # ClkVfPointsSetControl "SetClockBoostTable": 0x0733E009, # ClkVfPointsSetControl
} }
# ═══════════════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════════════
@@ -170,33 +167,33 @@ FUNC = {
# ═══════════════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════════════
# GetVFPCurve (0x21537AD4) # GetVFPCurve (0x21537AD4)
VFP_SIZE = 0x1C28 VFP_SIZE = 0x1C28
VFP_BASE = 0x48 VFP_BASE = 0x48
VFP_STRIDE = 0x1C # 28 bytes VFP_STRIDE = 0x1C # 28 bytes
VFP_MAX_ENTRIES = (VFP_SIZE - VFP_BASE) // VFP_STRIDE # 255 VFP_MAX_ENTRIES = (VFP_SIZE - VFP_BASE) // VFP_STRIDE # 255
# Get/SetClockBoostTable (0x23F1B133 / 0x0733E009) # Get/SetClockBoostTable (0x23F1B133 / 0x0733E009)
CT_SIZE = 0x2420 CT_SIZE = 0x2420
CT_BASE = 0x44 CT_BASE = 0x44
CT_STRIDE = 0x24 # 36 bytes CT_STRIDE = 0x24 # 36 bytes
CT_DELTA_OFF = 0x14 # freqDelta offset within entry CT_DELTA_OFF = 0x14 # freqDelta offset within entry
CT_MAX_ENTRIES = (CT_SIZE - CT_BASE) // CT_STRIDE # 255 CT_MAX_ENTRIES = (CT_SIZE - CT_BASE) // CT_STRIDE # 255
# GetClockBoostMask (0x507B4B59) # GetClockBoostMask (0x507B4B59)
MASK_SIZE = 0x182C MASK_SIZE = 0x182C
# Other structs # Other structs
VOLT_SIZE = 0x004C VOLT_SIZE = 0x004C
RANGES_SIZE = 0x0928 RANGES_SIZE = 0x0928
PERF_SIZE = 0x030C PERF_SIZE = 0x030C
VBOOST_SIZE = 0x0028 VBOOST_SIZE = 0x0028
# Mask location within VFP/CT structs # Mask location within VFP/CT structs
MASK_OFFSET = 0x04 MASK_OFFSET = 0x04
MASK_BYTES = 32 # 256 bits — covers up to 256 points MASK_BYTES = 32 # 256 bits — covers up to 256 points
# Safety constants # Safety constants
MAX_DELTA_KHZ = 300_000 # ±300 MHz hard cap for safety MAX_DELTA_KHZ = 300_000 # ±300 MHz hard cap for safety
SNAPSHOT_DIR = os.path.expanduser("~/.cache/nv_vfcurve") SNAPSHOT_DIR = os.path.expanduser("~/.cache/nv_vfcurve")
@@ -205,6 +202,7 @@ SNAPSHOT_DIR = os.path.expanduser("~/.cache/nv_vfcurve")
# GPU initialization # GPU initialization
# ═══════════════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════════════
def init_gpu() -> tuple: def init_gpu() -> tuple:
"""Initialize NvAPI, enumerate GPUs, return (handle, name).""" """Initialize NvAPI, enumerate GPUs, return (handle, name)."""
init_fn = nvfunc(FUNC["Initialize"], 0) init_fn = nvfunc(FUNC["Initialize"], 0)
@@ -236,18 +234,20 @@ def init_gpu() -> tuple:
# also distinguishes GPU core vs memory clock domains. # also distinguishes GPU core vs memory clock domains.
# ═══════════════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════════════
class BoostMask: class BoostMask:
"""Parsed GetClockBoostMask data. """Parsed GetClockBoostMask data.
Provides the raw mask bytes for copying into other calls, plus Provides the raw mask bytes for copying into other calls, plus
parsed per-entry enabled info for filtering. parsed per-entry enabled info for filtering.
""" """
def __init__(self, raw: bytes): def __init__(self, raw: bytes):
self.raw = raw self.raw = raw
self.size = len(raw) self.size = len(raw)
# The mask field at offset 0x04, 16 bytes — same position as in VFP/CT structs # The mask field at offset 0x04, 16 bytes — same position as in VFP/CT structs
self.mask_bytes = raw[MASK_OFFSET:MASK_OFFSET + MASK_BYTES] self.mask_bytes = raw[MASK_OFFSET : MASK_OFFSET + MASK_BYTES]
self.entries = [] self.entries = []
self._parse_entries() self._parse_entries()
@@ -260,7 +260,7 @@ class BoostMask:
enabled = bool(self.mask_bytes[byte_idx] & (1 << bit_idx)) enabled = bool(self.mask_bytes[byte_idx] & (1 << bit_idx))
self.entries.append({"index": i, "enabled": enabled}) self.entries.append({"index": i, "enabled": enabled})
def get_enabled_indices(self) -> List[int]: def get_enabled_indices(self) -> list[int]:
"""Return list of point indices that are enabled in the mask.""" """Return list of point indices that are enabled in the mask."""
return [e["index"] for e in self.entries if e["enabled"]] return [e["index"] for e in self.entries if e["enabled"]]
@@ -273,12 +273,13 @@ class BoostMask:
buf[offset + i] = self.mask_bytes[i] buf[offset + i] = self.mask_bytes[i]
def read_boost_mask(gpu) -> Tuple[Optional[BoostMask], str]: def read_boost_mask(gpu) -> tuple[BoostMask | None, str]:
"""Read the clock boost mask — the canonical source of active point info. """Read the clock boost mask — the canonical source of active point info.
Per nvapioc, this mask must be copied into VFP and ClockBoostTable calls. Per nvapioc, this mask must be copied into VFP and ClockBoostTable calls.
Using all-0xFF works on some GPUs (Blackwell) but fails on others (Pascal). Using all-0xFF works on some GPUs (Blackwell) but fails on others (Pascal).
""" """
def fill(buf): def fill(buf):
for i in range(MASK_OFFSET, MASK_OFFSET + MASK_BYTES): for i in range(MASK_OFFSET, MASK_OFFSET + MASK_BYTES):
buf[i] = 0xFF buf[i] = 0xFF
@@ -294,20 +295,22 @@ def read_boost_mask(gpu) -> Tuple[Optional[BoostMask], str]:
# Point classification — GPU core vs memory # Point classification — GPU core vs memory
# ═══════════════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════════════
class CurveInfo: class CurveInfo:
"""Holds classified point information for the GPU's V/F curve. """Holds classified point information for the GPU's V/F curve.
Combines data from GetClockBoostMask, GetVFPCurve, and GetClockBoostTable Combines data from GetClockBoostMask, GetVFPCurve, and GetClockBoostTable
to determine which points are GPU core and which are memory. to determine which points are GPU core and which are memory.
""" """
def __init__(self): def __init__(self):
self.gpu_points: List[int] = [] # GPU core V/F point indices self.gpu_points: list[int] = [] # GPU core V/F point indices
self.mem_points: List[int] = [] # Memory V/F point indices self.mem_points: list[int] = [] # Memory V/F point indices
self.total_points: int = 0 # Total populated entries self.total_points: int = 0 # Total populated entries
self.mask: Optional[BoostMask] = None self.mask: BoostMask | None = None
@staticmethod @staticmethod
def build(gpu, mask: Optional[BoostMask] = None) -> 'CurveInfo': def build(gpu, mask: BoostMask | None = None) -> "CurveInfo":
"""Classify all points by reading CT field_00 and VFP data. """Classify all points by reading CT field_00 and VFP data.
field_00 == 0: GPU core data point field_00 == 0: GPU core data point
@@ -340,7 +343,7 @@ class CurveInfo:
has_vfp_data = False has_vfp_data = False
if vfp_points and i < len(vfp_points): if vfp_points and i < len(vfp_points):
f, v = vfp_points[i] f, v = vfp_points[i]
has_vfp_data = (f > 0 or v > 0) has_vfp_data = f > 0 or v > 0
has_ct_data = False has_ct_data = False
for j in range(9): for j in range(9):
@@ -378,13 +381,15 @@ class CurveInfo:
# Data readers (mask-aware) # Data readers (mask-aware)
# ═══════════════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════════════
def _fill_mask_from_boost(buf, mask: BoostMask): def _fill_mask_from_boost(buf, mask: BoostMask):
"""Copy boost mask into buffer.""" """Copy boost mask into buffer."""
mask.copy_mask_into(buf) mask.copy_mask_into(buf)
def _read_vfp_with_mask(gpu, mask: Optional[BoostMask]) -> Optional[List[Tuple[int, int]]]: def _read_vfp_with_mask(gpu, mask: BoostMask | None) -> list[tuple[int, int]] | None:
"""Read VFP curve using the canonical boost mask.""" """Read VFP curve using the canonical boost mask."""
def fill(buf): def fill(buf):
_fill_mask_from_boost(buf, mask) _fill_mask_from_boost(buf, mask)
@@ -403,8 +408,9 @@ def _read_vfp_with_mask(gpu, mask: Optional[BoostMask]) -> Optional[List[Tuple[i
return points return points
def _read_clock_table_raw_with_mask(gpu, mask: Optional[BoostMask]) -> Optional[bytes]: def _read_clock_table_raw_with_mask(gpu, mask: BoostMask | None) -> bytes | None:
"""Read raw ClockBoostTable using the canonical boost mask.""" """Read raw ClockBoostTable using the canonical boost mask."""
def fill(buf): def fill(buf):
_fill_mask_from_boost(buf, mask) _fill_mask_from_boost(buf, mask)
@@ -412,13 +418,14 @@ def _read_clock_table_raw_with_mask(gpu, mask: Optional[BoostMask]) -> Optional[
return d if d else None return d if d else None
def read_vfp_curve(gpu, mask: Optional[BoostMask] = None, def read_vfp_curve(
curve_info: Optional[CurveInfo] = None gpu, mask: BoostMask | None = None, curve_info: CurveInfo | None = None
) -> Tuple[Optional[List[Tuple[int, int]]], str]: ) -> tuple[list[tuple[int, int]] | None, str]:
"""Read V/F curve (frequency + voltage pairs). """Read V/F curve (frequency + voltage pairs).
Returns up to 255 entries. Use curve_info to determine which are GPU/mem. Returns up to 255 entries. Use curve_info to determine which are GPU/mem.
""" """
def fill(buf): def fill(buf):
_fill_mask_from_boost(buf, mask) _fill_mask_from_boost(buf, mask)
@@ -442,18 +449,20 @@ def read_vfp_curve(gpu, mask: Optional[BoostMask] = None,
return points, "OK" return points, "OK"
def read_clock_table_raw(gpu, mask: Optional[BoostMask] = None def read_clock_table_raw(
) -> Tuple[Optional[bytes], str]: gpu, mask: BoostMask | None = None
) -> tuple[bytes | None, str]:
"""Read the raw ClockBoostTable buffer.""" """Read the raw ClockBoostTable buffer."""
def fill(buf): def fill(buf):
_fill_mask_from_boost(buf, mask) _fill_mask_from_boost(buf, mask)
return nvcall(FUNC["GetClockBoostTable"], gpu, CT_SIZE, ver=1, pre_fill=fill) return nvcall(FUNC["GetClockBoostTable"], gpu, CT_SIZE, ver=1, pre_fill=fill)
def read_clock_offsets(gpu, mask: Optional[BoostMask] = None, def read_clock_offsets(
curve_info: Optional[CurveInfo] = None gpu, mask: BoostMask | None = None, curve_info: CurveInfo | None = None
) -> Tuple[Optional[List[int]], str]: ) -> tuple[list[int] | None, str]:
"""Read per-point frequency offsets from the ClockBoostTable.""" """Read per-point frequency offsets from the ClockBoostTable."""
d, err = read_clock_table_raw(gpu, mask) d, err = read_clock_table_raw(gpu, mask)
if not d: if not d:
@@ -482,14 +491,18 @@ def read_clock_entry_full(data: bytes, point: int) -> dict:
for j in range(9): for j in range(9):
off = base + j * 4 off = base + j * 4
if j == 5: if j == 5:
fields[f"field_{j:02d}_0x{j*4:02X}"] = struct.unpack_from("<i", data, off)[0] fields[f"field_{j:02d}_0x{j * 4:02X}"] = struct.unpack_from(
"<i", data, off
)[0]
else: else:
fields[f"field_{j:02d}_0x{j*4:02X}"] = struct.unpack_from("<I", data, off)[0] fields[f"field_{j:02d}_0x{j * 4:02X}"] = struct.unpack_from(
"<I", data, off
)[0]
fields["freqDelta_kHz"] = fields["field_05_0x14"] fields["freqDelta_kHz"] = fields["field_05_0x14"]
return fields return fields
def read_voltage(gpu) -> Tuple[Optional[int], str]: def read_voltage(gpu) -> tuple[int | None, str]:
"""Read current GPU core voltage in µV.""" """Read current GPU core voltage in µV."""
d, err = nvcall(FUNC["GetCurrentVoltage"], gpu, VOLT_SIZE, ver=1) d, err = nvcall(FUNC["GetCurrentVoltage"], gpu, VOLT_SIZE, ver=1)
if not d: if not d:
@@ -497,7 +510,7 @@ def read_voltage(gpu) -> Tuple[Optional[int], str]:
return struct.unpack_from("<I", d, 0x28)[0], "OK" return struct.unpack_from("<I", d, 0x28)[0], "OK"
def read_clock_ranges(gpu) -> Tuple[Optional[dict], str]: def read_clock_ranges(gpu) -> tuple[dict | None, str]:
"""Read clock domain min/max offset ranges.""" """Read clock domain min/max offset ranges."""
d, err = nvcall(FUNC["GetClockBoostRanges"], gpu, RANGES_SIZE, ver=1) d, err = nvcall(FUNC["GetClockBoostRanges"], gpu, RANGES_SIZE, ver=1)
if not d: if not d:
@@ -508,8 +521,7 @@ def read_clock_ranges(gpu) -> Tuple[Optional[dict], str]:
base = 0x08 + i * 0x48 base = 0x08 + i * 0x48
if base + 0x48 > len(d): if base + 0x48 > len(d):
break break
words = [struct.unpack_from("<i", d, base + j)[0] words = [struct.unpack_from("<i", d, base + j)[0] for j in range(0, 0x48, 4)]
for j in range(0, 0x48, 4)]
domains.append(words) domains.append(words)
return {"num_domains": num, "domains": domains}, "OK" return {"num_domains": num, "domains": domains}, "OK"
@@ -518,14 +530,17 @@ def read_clock_ranges(gpu) -> Tuple[Optional[dict], str]:
# Mask bit helpers # Mask bit helpers
# ═══════════════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════════════
def set_mask_bit(buf, point: int, offset=MASK_OFFSET): def set_mask_bit(buf, point: int, offset=MASK_OFFSET):
"""Set a single bit in the mask field.""" """Set a single bit in the mask field."""
byte_idx = offset + (point // 8) byte_idx = offset + (point // 8)
bit_idx = point % 8 bit_idx = point % 8
buf[byte_idx] = int.from_bytes(buf[byte_idx:byte_idx+1], 'little') | (1 << bit_idx) buf[byte_idx] = int.from_bytes(buf[byte_idx : byte_idx + 1], "little") | (
1 << bit_idx
)
def set_mask_bits(buf, points: Set[int], offset=MASK_OFFSET): def set_mask_bits(buf, points: set[int], offset=MASK_OFFSET):
"""Set mask bits for a set of points.""" """Set mask bits for a set of points."""
for p in points: for p in points:
set_mask_bit(buf, p, offset) set_mask_bit(buf, p, offset)
@@ -535,11 +550,12 @@ def set_mask_bits(buf, points: Set[int], offset=MASK_OFFSET):
# Write operations # Write operations
# ═══════════════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════════════
def build_write_buffer( def build_write_buffer(
gpu, gpu,
point_deltas: dict, point_deltas: dict,
mask: Optional[BoostMask] = None, mask: BoostMask | None = None,
) -> Tuple[Optional[ctypes.Array], str]: ) -> tuple[ctypes.Array | None, str]:
"""Build a SetClockBoostTable buffer with specified per-point deltas. """Build a SetClockBoostTable buffer with specified per-point deltas.
Strategy: read the current ClockBoostTable (using canonical mask), Strategy: read the current ClockBoostTable (using canonical mask),
@@ -576,9 +592,9 @@ def build_write_buffer(
def write_clock_offsets( def write_clock_offsets(
gpu, gpu,
point_deltas: dict, point_deltas: dict,
mask: Optional[BoostMask] = None, mask: BoostMask | None = None,
dry_run: bool = False, dry_run: bool = False,
) -> Tuple[int, str]: ) -> tuple[int, str]:
"""Write per-point frequency offsets via SetClockBoostTable.""" """Write per-point frequency offsets via SetClockBoostTable."""
buf, err = build_write_buffer(gpu, point_deltas, mask) buf, err = build_write_buffer(gpu, point_deltas, mask)
if buf is None: if buf is None:
@@ -595,9 +611,10 @@ def write_clock_offsets(
# Safety checks # Safety checks
# ═══════════════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════════════
def validate_write_request(point_deltas: dict,
curve_info: Optional[CurveInfo] = None def validate_write_request(
) -> Optional[str]: point_deltas: dict, curve_info: CurveInfo | None = None
) -> str | None:
"""Return an error message if the write request is unsafe, else None.""" """Return an error message if the write request is unsafe, else None."""
mem_points = set() mem_points = set()
if curve_info: if curve_info:
@@ -608,14 +625,18 @@ def validate_write_request(point_deltas: dict,
return f"Point {point} out of range (0–{CT_MAX_ENTRIES - 1})" return f"Point {point} out of range (0–{CT_MAX_ENTRIES - 1})"
if point in mem_points: if point in mem_points:
return (f"Point {point} is a memory clock entry. " return (
"Memory offsets use a different mechanism (NVML). " f"Point {point} is a memory clock entry. "
"Use --force if you really mean it.") "Memory offsets use a different mechanism (NVML). "
"Use --force if you really mean it."
)
if abs(delta_khz) > MAX_DELTA_KHZ: if abs(delta_khz) > MAX_DELTA_KHZ:
return (f"Delta {delta_khz/1000:+.0f} MHz for point {point} exceeds " return (
f"safety limit of ±{MAX_DELTA_KHZ/1000:.0f} MHz. " f"Delta {delta_khz / 1000:+.0f} MHz for point {point} exceeds "
"Use --max-delta to raise the limit if needed.") f"safety limit of ±{MAX_DELTA_KHZ / 1000:.0f} MHz. "
"Use --max-delta to raise the limit if needed."
)
return None return None
@@ -624,11 +645,12 @@ def validate_write_request(point_deltas: dict,
# Hex dump utility # Hex dump utility
# ═══════════════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════════════
def hexdump(data: bytes, start: int, length: int, cols: int = 16) -> str: def hexdump(data: bytes, start: int, length: int, cols: int = 16) -> str:
lines = [] lines = []
end = min(start + length, len(data)) end = min(start + length, len(data))
for off in range(start, end, cols): for off in range(start, end, cols):
chunk = data[off:off + cols] chunk = data[off : off + cols]
hx = " ".join(f"{b:02x}" for b in chunk) hx = " ".join(f"{b:02x}" for b in chunk)
asc = "".join(chr(b) if 32 <= b < 127 else "." for b in chunk) asc = "".join(chr(b) if 32 <= b < 127 else "." for b in chunk)
lines.append(f" {off:04x}: {hx:<{cols * 3}} {asc}") lines.append(f" {off:04x}: {hx:<{cols * 3}} {asc}")
@@ -639,7 +661,8 @@ def hexdump(data: bytes, start: int, length: int, cols: int = 16) -> str:
# Snapshot save/restore # Snapshot save/restore
# ═══════════════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════════════
def snapshot_save(gpu, gpu_name: str, mask: Optional[BoostMask] = None):
def snapshot_save(gpu, gpu_name: str, mask: BoostMask | None = None):
"""Save the current ClockBoostTable to disk.""" """Save the current ClockBoostTable to disk."""
raw, err = read_clock_table_raw(gpu, mask) raw, err = read_clock_table_raw(gpu, mask)
if not raw: if not raw:
@@ -674,7 +697,7 @@ def snapshot_save(gpu, gpu_name: str, mask: Optional[BoostMask] = None):
with open(meta_fname, "w") as f: with open(meta_fname, "w") as f:
json.dump(meta, f, indent=2) json.dump(meta, f, indent=2)
print(f"Snapshot saved:") print("Snapshot saved:")
print(f" Binary: {fname}") print(f" Binary: {fname}")
print(f" Metadata: {meta_fname}") print(f" Metadata: {meta_fname}")
print(f" Size: {len(raw)} bytes") print(f" Size: {len(raw)} bytes")
@@ -682,7 +705,7 @@ def snapshot_save(gpu, gpu_name: str, mask: Optional[BoostMask] = None):
return True return True
def snapshot_restore(gpu, mask: Optional[BoostMask] = None, filepath: str = None): def snapshot_restore(gpu, mask: BoostMask | None = None, filepath: str = None):
"""Restore a ClockBoostTable snapshot from disk.""" """Restore a ClockBoostTable snapshot from disk."""
if filepath is None: if filepath is None:
if not os.path.isdir(SNAPSHOT_DIR): if not os.path.isdir(SNAPSHOT_DIR):
@@ -730,7 +753,8 @@ def snapshot_restore(gpu, mask: Optional[BoostMask] = None, filepath: str = None
# Diagnostics # Diagnostics
# ═══════════════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════════════
def run_diagnostics(gpu, gpu_name, mask: Optional[BoostMask] = None):
def run_diagnostics(gpu, gpu_name, mask: BoostMask | None = None):
"""Probe all known functions and report results.""" """Probe all known functions and report results."""
print(f"GPU: {gpu_name}") print(f"GPU: {gpu_name}")
print() print()
@@ -739,14 +763,14 @@ def run_diagnostics(gpu, gpu_name, mask: Optional[BoostMask] = None):
print("=== Function probe ===") print("=== Function probe ===")
print() print()
probes = [ probes = [
("GetVFPCurve", FUNC["GetVFPCurve"], VFP_SIZE, 1), ("GetVFPCurve", FUNC["GetVFPCurve"], VFP_SIZE, 1),
("GetClockBoostMask", FUNC["GetClockBoostMask"], MASK_SIZE, 1), ("GetClockBoostMask", FUNC["GetClockBoostMask"], MASK_SIZE, 1),
("GetClockBoostTable", FUNC["GetClockBoostTable"], CT_SIZE, 1), ("GetClockBoostTable", FUNC["GetClockBoostTable"], CT_SIZE, 1),
("GetCurrentVoltage", FUNC["GetCurrentVoltage"], VOLT_SIZE, 1), ("GetCurrentVoltage", FUNC["GetCurrentVoltage"], VOLT_SIZE, 1),
("GetClockBoostRanges", FUNC["GetClockBoostRanges"], RANGES_SIZE, 1), ("GetClockBoostRanges", FUNC["GetClockBoostRanges"], RANGES_SIZE, 1),
("GetPerfLimits", FUNC["GetPerfLimits"], PERF_SIZE, 2), ("GetPerfLimits", FUNC["GetPerfLimits"], PERF_SIZE, 2),
("GetVoltBoostPercent", FUNC["GetVoltBoostPercent"], VBOOST_SIZE, 1), ("GetVoltBoostPercent", FUNC["GetVoltBoostPercent"], VBOOST_SIZE, 1),
("SetClockBoostTable", FUNC["SetClockBoostTable"], CT_SIZE, 1), ("SetClockBoostTable", FUNC["SetClockBoostTable"], CT_SIZE, 1),
] ]
for name, fid, size, ver in probes: for name, fid, size, ver in probes:
ptr = QI(fid) ptr = QI(fid)
@@ -768,7 +792,9 @@ def run_diagnostics(gpu, gpu_name, mask: Optional[BoostMask] = None):
# Step 3: test reads with the proper mask # Step 3: test reads with the proper mask
needs_mask_fns = { needs_mask_fns = {
FUNC["GetVFPCurve"], FUNC["GetClockBoostMask"], FUNC["GetClockBoostTable"] FUNC["GetVFPCurve"],
FUNC["GetClockBoostMask"],
FUNC["GetClockBoostTable"],
} }
print() print()
@@ -817,8 +843,10 @@ def run_diagnostics(gpu, gpu_name, mask: Optional[BoostMask] = None):
# Output formatting # Output formatting
# ═══════════════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════════════
def print_curve(points, offsets, voltage, curve_info: Optional[CurveInfo] = None,
full=False): def print_curve(
points, offsets, voltage, curve_info: CurveInfo | None = None, full=False
):
"""Print formatted V/F curve table.""" """Print formatted V/F curve table."""
if voltage: if voltage:
print(f"Current voltage: {voltage / 1000:.1f} mV") print(f"Current voltage: {voltage / 1000:.1f} mV")
@@ -845,9 +873,7 @@ def print_curve(points, offsets, voltage, curve_info: Optional[CurveInfo] = None
for i, (f, v) in enumerate(points): for i, (f, v) in enumerate(points):
if f == 0 and v == 0: if f == 0 and v == 0:
continue continue
if i in mem_set: if i in mem_set or f != prev_freq or i == len(points) - 1:
show.append(i)
elif f != prev_freq or i == len(points) - 1:
show.append(i) show.append(i)
prev_freq = f prev_freq = f
@@ -882,52 +908,78 @@ def print_curve(points, offsets, voltage, curve_info: Optional[CurveInfo] = None
# Summary # Summary
if curve_info and curve_info.gpu_points: if curve_info and curve_info.gpu_points:
gpu_data = [(points[i][0], points[i][1]) for i in curve_info.gpu_points gpu_data = [
if i < len(points) and points[i][0] > 0] (points[i][0], points[i][1])
for i in curve_info.gpu_points
if i < len(points) and points[i][0] > 0
]
if gpu_data: if gpu_data:
freqs = [f for f, v in gpu_data] freqs = [f for f, v in gpu_data]
volts = [v for f, v in gpu_data] volts = [v for f, v in gpu_data]
print() print()
print(f"GPU core: {min(freqs)/1000:.0f} – {max(freqs)/1000:.0f} MHz, " print(
f"{min(volts)/1000:.0f} – {max(volts)/1000:.0f} mV " f"GPU core: {min(freqs) / 1000:.0f} – {max(freqs) / 1000:.0f} MHz, "
f"({len(gpu_data)} points)") f"{min(volts) / 1000:.0f} – {max(volts) / 1000:.0f} mV "
f"({len(gpu_data)} points)"
)
if curve_info and curve_info.mem_points: if curve_info and curve_info.mem_points:
mem_data = [(points[i][0], points[i][1]) for i in curve_info.mem_points mem_data = [
if i < len(points) and points[i][0] > 0] (points[i][0], points[i][1])
for i in curve_info.mem_points
if i < len(points) and points[i][0] > 0
]
if mem_data: if mem_data:
freqs = [f for f, v in mem_data] freqs = [f for f, v in mem_data]
volts = [v for f, v in mem_data] volts = [v for f, v in mem_data]
print(f"Memory: {min(freqs)/1000:.0f} – {max(freqs)/1000:.0f} MHz, " print(
f"{min(volts)/1000:.0f} – {max(volts)/1000:.0f} mV " f"Memory: {min(freqs) / 1000:.0f} – {max(freqs) / 1000:.0f} MHz, "
f"({len(mem_data)} points)") f"{min(volts) / 1000:.0f} – {max(volts) / 1000:.0f} mV "
f"({len(mem_data)} points)"
)
if offsets: if offsets:
gpu_indices = set(curve_info.gpu_points) if curve_info else set(range(len(offsets))) gpu_indices = (
gpu_offsets = [offsets[i] for i in gpu_indices set(curve_info.gpu_points) if curve_info else set(range(len(offsets)))
if i < len(offsets) and offsets[i] != 0] )
gpu_offsets = [
offsets[i] for i in gpu_indices if i < len(offsets) and offsets[i] != 0
]
if gpu_offsets: if gpu_offsets:
vals = set(gpu_offsets) vals = set(gpu_offsets)
if len(vals) == 1: if len(vals) == 1:
print(f"GPU offset: {next(iter(vals))/1000:+.0f} MHz " print(
f"(uniform across {len(gpu_offsets)} points)") f"GPU offset: {next(iter(vals)) / 1000:+.0f} MHz "
f"(uniform across {len(gpu_offsets)} points)"
)
else: else:
print(f"GPU offsets: {len(gpu_offsets)} points active " print(
f"(range: {min(vals)/1000:+.0f} to {max(vals)/1000:+.0f} MHz)") f"GPU offsets: {len(gpu_offsets)} points active "
f"(range: {min(vals) / 1000:+.0f} to {max(vals) / 1000:+.0f} MHz)"
)
def output_json(gpu_name, points, offsets, voltage, def output_json(
curve_info: Optional[CurveInfo] = None): gpu_name, points, offsets, voltage, curve_info: CurveInfo | None = None
):
"""Output JSON format.""" """Output JSON format."""
data = { data = {
"gpu": gpu_name, "gpu": gpu_name,
"current_voltage_uV": voltage, "current_voltage_uV": voltage,
"layout": { "layout": {
"vfp_curve": {"size": VFP_SIZE, "base": VFP_BASE, "vfp_curve": {
"stride": VFP_STRIDE, "max_entries": VFP_MAX_ENTRIES}, "size": VFP_SIZE,
"clock_table": {"size": CT_SIZE, "base": CT_BASE, "base": VFP_BASE,
"stride": CT_STRIDE, "delta_offset": CT_DELTA_OFF, "stride": VFP_STRIDE,
"max_entries": CT_MAX_ENTRIES}, "max_entries": VFP_MAX_ENTRIES,
},
"clock_table": {
"size": CT_SIZE,
"base": CT_BASE,
"stride": CT_STRIDE,
"delta_offset": CT_DELTA_OFF,
"max_entries": CT_MAX_ENTRIES,
},
}, },
"curve_info": { "curve_info": {
"gpu_points": curve_info.gpu_points if curve_info else [], "gpu_points": curve_info.gpu_points if curve_info else [],
@@ -958,6 +1010,7 @@ def output_json(gpu_name, points, offsets, voltage,
# Write command handler # Write command handler
# ═══════════════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════════════
def cmd_write(gpu, gpu_name, args, mask, curve_info): def cmd_write(gpu, gpu_name, args, mask, curve_info):
"""Handle write subcommand.""" """Handle write subcommand."""
delta_khz = int(args.delta * 1000) delta_khz = int(args.delta * 1000)
@@ -977,21 +1030,27 @@ def cmd_write(gpu, gpu_name, args, mask, curve_info):
elif args.point is not None: elif args.point is not None:
point_deltas[args.point] = delta_khz point_deltas[args.point] = delta_khz
print(f"Target: point {args.point}, delta {args.delta:+.0f} MHz " print(
f"({delta_khz:+d} kHz)") f"Target: point {args.point}, delta {args.delta:+.0f} MHz "
f"({delta_khz:+d} kHz)"
)
elif args.range: elif args.range:
start, end = args.range start, end = args.range
for i in range(start, end + 1): for i in range(start, end + 1):
point_deltas[i] = delta_khz point_deltas[i] = delta_khz
print(f"Target: points {start}–{end} ({len(point_deltas)} points), " print(
f"delta {args.delta:+.0f} MHz") f"Target: points {start}–{end} ({len(point_deltas)} points), "
f"delta {args.delta:+.0f} MHz"
)
elif args.glob: elif args.glob:
for i in gpu_points: for i in gpu_points:
point_deltas[i] = delta_khz point_deltas[i] = delta_khz
print(f"Target: all {len(point_deltas)} GPU core points, " print(
f"delta {args.delta:+.0f} MHz") f"Target: all {len(point_deltas)} GPU core points, "
f"delta {args.delta:+.0f} MHz"
)
else: else:
print("Error: specify --point N, --range A-B, --global, or --reset") print("Error: specify --point N, --range A-B, --global, or --reset")
@@ -1013,11 +1072,17 @@ def cmd_write(gpu, gpu_name, args, mask, curve_info):
changed = 0 changed = 0
for point in sorted(point_deltas.keys()): for point in sorted(point_deltas.keys()):
new = point_deltas[point] new = point_deltas[point]
old = current_offsets[point] if current_offsets and point < len(current_offsets) else 0 old = (
current_offsets[point]
if current_offsets and point < len(current_offsets)
else 0
)
if old != new: if old != new:
changed += 1 changed += 1
if changed <= 20: if changed <= 20:
print(f" Point {point:3d}: {old/1000:+8.0f} MHz → {new/1000:+8.0f} MHz") print(
f" Point {point:3d}: {old / 1000:+8.0f} MHz → {new / 1000:+8.0f} MHz"
)
if changed > 20: if changed > 20:
print(f" ... and {changed - 20} more points") print(f" ... and {changed - 20} more points")
if changed == 0: if changed == 0:
@@ -1036,8 +1101,10 @@ def cmd_write(gpu, gpu_name, args, mask, curve_info):
first_pt = min(point_deltas.keys()) first_pt = min(point_deltas.keys())
entry_off = CT_BASE + first_pt * CT_STRIDE entry_off = CT_BASE + first_pt * CT_STRIDE
print(f"\nEntry for point {first_pt} (offset 0x{entry_off:04X}, " print(
f"stride 0x{CT_STRIDE:02X}):") f"\nEntry for point {first_pt} (offset 0x{entry_off:04X}, "
f"stride 0x{CT_STRIDE:02X}):"
)
print(hexdump(bytes(buf), entry_off, CT_STRIDE)) print(hexdump(bytes(buf), entry_off, CT_STRIDE))
return return
@@ -1070,8 +1137,10 @@ def cmd_write(gpu, gpu_name, args, mask, curve_info):
actual = new_offsets[point] if point < len(new_offsets) else 0 actual = new_offsets[point] if point < len(new_offsets) else 0
if actual != expected: if actual != expected:
mismatches += 1 mismatches += 1
print(f" MISMATCH point {point}: expected {expected/1000:+.0f} MHz, " print(
f"got {actual/1000:+.0f} MHz") f" MISMATCH point {point}: expected {expected / 1000:+.0f} MHz, "
f"got {actual / 1000:+.0f} MHz"
)
if mismatches == 0: if mismatches == 0:
print(f"Verified: all {len(point_deltas)} points match expected values.") print(f"Verified: all {len(point_deltas)} points match expected values.")
@@ -1083,6 +1152,7 @@ def cmd_write(gpu, gpu_name, args, mask, curve_info):
# Verify command handler # Verify command handler
# ═══════════════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════════════
def cmd_verify(gpu, gpu_name, args, mask, curve_info): def cmd_verify(gpu, gpu_name, args, mask, curve_info):
"""Write-verify-read cycle for a single point or range.""" """Write-verify-read cycle for a single point or range."""
delta_khz = int(args.delta * 1000) delta_khz = int(args.delta * 1000)
@@ -1095,14 +1165,14 @@ def cmd_verify(gpu, gpu_name, args, mask, curve_info):
print("Error: --point or --range required for verify mode") print("Error: --point or --range required for verify mode")
return return
point_deltas = {p: delta_khz for p in points} point_deltas = dict.fromkeys(points, delta_khz)
err = validate_write_request(point_deltas, curve_info) err = validate_write_request(point_deltas, curve_info)
if err: if err:
print(f"Safety check FAILED: {err}") print(f"Safety check FAILED: {err}")
return return
print(f"=== Write-Verify Cycle ===") print("=== Write-Verify Cycle ===")
print(f"GPU: {gpu_name}") print(f"GPU: {gpu_name}")
if curve_info: if curve_info:
print(f"Curve: {curve_info.describe()}") print(f"Curve: {curve_info.describe()}")
@@ -1121,7 +1191,7 @@ def cmd_verify(gpu, gpu_name, args, mask, curve_info):
for p in points[:5]: for p in points[:5]:
entry = read_clock_entry_full(before_raw, p) if before_raw else {} entry = read_clock_entry_full(before_raw, p) if before_raw else {}
off_val = before_offsets[p] if p < len(before_offsets) else 0 off_val = before_offsets[p] if p < len(before_offsets) else 0
print(f" Point {p:3d}: freqDelta = {off_val/1000:+8.0f} MHz") print(f" Point {p:3d}: freqDelta = {off_val / 1000:+8.0f} MHz")
if entry: if entry:
print(f" All fields: {entry}") print(f" All fields: {entry}")
@@ -1156,8 +1226,10 @@ def cmd_verify(gpu, gpu_name, args, mask, curve_info):
match = "OK" if actual == expected else "MISMATCH" match = "OK" if actual == expected else "MISMATCH"
if actual != expected: if actual != expected:
all_ok = False all_ok = False
print(f" Point {p:3d}: expected {expected/1000:+8.0f} MHz, " print(
f"got {actual/1000:+8.0f} MHz [{match}]") f" Point {p:3d}: expected {expected / 1000:+8.0f} MHz, "
f"got {actual / 1000:+8.0f} MHz [{match}]"
)
# Step 5: Check for collateral damage # Step 5: Check for collateral damage
print() print()
@@ -1169,8 +1241,10 @@ def cmd_verify(gpu, gpu_name, args, mask, curve_info):
continue continue
if before_offsets[i] != after_offsets[i]: if before_offsets[i] != after_offsets[i]:
collateral += 1 collateral += 1
print(f" WARNING: Point {i} changed unexpectedly: " print(
f"{before_offsets[i]/1000:+.0f} → {after_offsets[i]/1000:+.0f} MHz") f" WARNING: Point {i} changed unexpectedly: "
f"{before_offsets[i] / 1000:+.0f} → {after_offsets[i] / 1000:+.0f} MHz"
)
if collateral == 0: if collateral == 0:
print(" No unintended changes detected.") print(" No unintended changes detected.")
@@ -1187,14 +1261,16 @@ def cmd_verify(gpu, gpu_name, args, mask, curve_info):
continue continue
if before_entry[key] != after_entry[key]: if before_entry[key] != after_entry[key]:
field_changes += 1 field_changes += 1
print(f" Point {p}, {key}: {before_entry[key]} → {after_entry[key]}") print(
f" Point {p}, {key}: {before_entry[key]} → {after_entry[key]}"
)
if field_changes == 0: if field_changes == 0:
print(" No unknown fields changed.") print(" No unknown fields changed.")
# Step 7: Read voltage # Step 7: Read voltage
voltage, _ = read_voltage(gpu) voltage, _ = read_voltage(gpu)
if voltage: if voltage:
print(f"\nCurrent voltage after write: {voltage/1000:.1f} mV") print(f"\nCurrent voltage after write: {voltage / 1000:.1f} mV")
# Summary # Summary
print() print()
@@ -1215,6 +1291,7 @@ def cmd_verify(gpu, gpu_name, args, mask, curve_info):
# Inspect command # Inspect command
# ═══════════════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════════════
def cmd_inspect(gpu, gpu_name, args, mask, curve_info): def cmd_inspect(gpu, gpu_name, args, mask, curve_info):
"""Show detailed field-level data for specific points.""" """Show detailed field-level data for specific points."""
raw, err = read_clock_table_raw(gpu, mask) raw, err = read_clock_table_raw(gpu, mask)
@@ -1246,8 +1323,9 @@ def cmd_inspect(gpu, gpu_name, args, mask, curve_info):
print(f"GPU: {gpu_name}") print(f"GPU: {gpu_name}")
if curve_info: if curve_info:
print(f"Curve: {curve_info.describe()}") print(f"Curve: {curve_info.describe()}")
print(f"ClockBoostTable entry detail (stride=0x{CT_STRIDE:02X}, " print(
f"9 fields × 4 bytes)") f"ClockBoostTable entry detail (stride=0x{CT_STRIDE:02X}, 9 fields × 4 bytes)"
)
print() print()
for p in indices: for p in indices:
@@ -1267,7 +1345,7 @@ def cmd_inspect(gpu, gpu_name, args, mask, curve_info):
freq_str = "" freq_str = ""
if vfp_points and p < len(vfp_points): if vfp_points and p < len(vfp_points):
f, v = vfp_points[p] f, v = vfp_points[p]
freq_str = f" (VFP: {f/1000:.0f} MHz @ {v/1000:.0f} mV)" freq_str = f" (VFP: {f / 1000:.0f} MHz @ {v / 1000:.0f} mV)"
print(f"Point {p:3d} — buffer offset 0x{off:04X}{domain}{freq_str}") print(f"Point {p:3d} — buffer offset 0x{off:04X}{domain}{freq_str}")
for key, val in entry.items(): for key, val in entry.items():
@@ -1275,8 +1353,10 @@ def cmd_inspect(gpu, gpu_name, args, mask, curve_info):
continue continue
marker = " ← freqDelta" if "0x14" in key else "" marker = " ← freqDelta" if "0x14" in key else ""
if "0x14" in key: if "0x14" in key:
print(f" {key}: {val:12d} (0x{val & 0xFFFFFFFF:08X})" print(
f" = {val/1000:+.0f} MHz{marker}") f" {key}: {val:12d} (0x{val & 0xFFFFFFFF:08X})"
f" = {val / 1000:+.0f} MHz{marker}"
)
else: else:
print(f" {key}: {val:12d} (0x{val:08X})") print(f" {key}: {val:12d} (0x{val:08X})")
print() print()
@@ -1286,6 +1366,7 @@ def cmd_inspect(gpu, gpu_name, args, mask, curve_info):
# Read command handler # Read command handler
# ═══════════════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════════════
def cmd_read(gpu, gpu_name, args, mask, curve_info): def cmd_read(gpu, gpu_name, args, mask, curve_info):
"""Handle read subcommand.""" """Handle read subcommand."""
if args.diag: if args.diag:
@@ -1309,10 +1390,13 @@ def cmd_read(gpu, gpu_name, args, mask, curve_info):
print(f"GPU: {gpu_name}") print(f"GPU: {gpu_name}")
if args.raw: if args.raw:
def fill_vfp(buf): def fill_vfp(buf):
_fill_mask_from_boost(buf, mask) _fill_mask_from_boost(buf, mask)
vfp_raw, _ = nvcall(FUNC["GetVFPCurve"], gpu, VFP_SIZE,
ver=1, pre_fill=fill_vfp) vfp_raw, _ = nvcall(
FUNC["GetVFPCurve"], gpu, VFP_SIZE, ver=1, pre_fill=fill_vfp
)
ct_raw, _ = read_clock_table_raw(gpu, mask) ct_raw, _ = read_clock_table_raw(gpu, mask)
if vfp_raw: if vfp_raw:
@@ -1348,7 +1432,8 @@ def cmd_read(gpu, gpu_name, args, mask, curve_info):
# Argument parsing # Argument parsing
# ═══════════════════════════════════════════════════════════════════════════ # ═══════════════════════════════════════════════════════════════════════════
def parse_range(s: str) -> Tuple[int, int]:
def parse_range(s: str) -> tuple[int, int]:
"""Parse 'A-B' into (A, B) tuple.""" """Parse 'A-B' into (A, B) tuple."""
parts = s.split("-") parts = s.split("-")
if len(parts) != 2: if len(parts) != 2:
@@ -1360,7 +1445,9 @@ def parse_range(s: str) -> Tuple[int, int]:
if a > b: if a > b:
raise argparse.ArgumentTypeError(f"Start > end in range: {a}-{b}") raise argparse.ArgumentTypeError(f"Start > end in range: {a}-{b}")
if a < 0 or b >= CT_MAX_ENTRIES: if a < 0 or b >= CT_MAX_ENTRIES:
raise argparse.ArgumentTypeError(f"Range {a}-{b} outside 0–{CT_MAX_ENTRIES - 1}") raise argparse.ArgumentTypeError(
f"Range {a}-{b} outside 0–{CT_MAX_ENTRIES - 1}"
)
return (a, b) return (a, b)
@@ -1390,14 +1477,16 @@ Examples:
# --- read --- # --- read ---
p_read = sub.add_parser("read", help="Read V/F curve (default)") p_read = sub.add_parser("read", help="Read V/F curve (default)")
p_read.add_argument("--full", action="store_true", p_read.add_argument(
help="Show all points including empty slots") "--full", action="store_true", help="Show all points including empty slots"
p_read.add_argument("--json", action="store_true", )
help="JSON output with domain classification") p_read.add_argument(
p_read.add_argument("--raw", action="store_true", "--json", action="store_true", help="JSON output with domain classification"
help="Include hex dumps") )
p_read.add_argument("--diag", action="store_true", p_read.add_argument("--raw", action="store_true", help="Include hex dumps")
help="Probe all functions with mask comparison") p_read.add_argument(
"--diag", action="store_true", help="Probe all functions with mask comparison"
)
# --- inspect --- # --- inspect ---
p_insp = sub.add_parser("inspect", help="Show detailed entry fields") p_insp = sub.add_parser("inspect", help="Show detailed entry fields")
@@ -1409,30 +1498,42 @@ Examples:
tgt = p_write.add_mutually_exclusive_group() tgt = p_write.add_mutually_exclusive_group()
tgt.add_argument("--point", type=int, help="Single point index") tgt.add_argument("--point", type=int, help="Single point index")
tgt.add_argument("--range", type=parse_range, help="Point range A-B") tgt.add_argument("--range", type=parse_range, help="Point range A-B")
tgt.add_argument("--global", dest="glob", action="store_true", tgt.add_argument(
help="All GPU core points") "--global", dest="glob", action="store_true", help="All GPU core points"
tgt.add_argument("--reset", action="store_true", )
help="Reset all GPU core offsets to 0") tgt.add_argument(
p_write.add_argument("--delta", type=float, default=0.0, "--reset", action="store_true", help="Reset all GPU core offsets to 0"
help="Frequency offset in MHz (e.g. 15, -30)") )
p_write.add_argument("--dry-run", action="store_true", p_write.add_argument(
help="Preview changes without applying") "--delta",
p_write.add_argument("--force", action="store_true", type=float,
help="Allow modifying memory points") default=0.0,
p_write.add_argument("--max-delta", type=float, default=300.0, help="Frequency offset in MHz (e.g. 15, -30)",
help="Override safety limit (MHz, default 300)") )
p_write.add_argument(
"--dry-run", action="store_true", help="Preview changes without applying"
)
p_write.add_argument(
"--force", action="store_true", help="Allow modifying memory points"
)
p_write.add_argument(
"--max-delta",
type=float,
default=300.0,
help="Override safety limit (MHz, default 300)",
)
# --- verify --- # --- verify ---
p_ver = sub.add_parser("verify", help="Write-verify-read cycle") p_ver = sub.add_parser("verify", help="Write-verify-read cycle")
p_ver.add_argument("--point", type=int, help="Single point index") p_ver.add_argument("--point", type=int, help="Single point index")
p_ver.add_argument("--range", type=parse_range, help="Point range A-B") p_ver.add_argument("--range", type=parse_range, help="Point range A-B")
p_ver.add_argument("--delta", type=float, required=True, p_ver.add_argument(
help="Frequency offset in MHz") "--delta", type=float, required=True, help="Frequency offset in MHz"
)
# --- snapshot --- # --- snapshot ---
p_snap = sub.add_parser("snapshot", help="Save/restore ClockBoostTable") p_snap = sub.add_parser("snapshot", help="Save/restore ClockBoostTable")
p_snap.add_argument("action", choices=["save", "restore"], p_snap.add_argument("action", choices=["save", "restore"], help="save or restore")
help="save or restore")
p_snap.add_argument("--file", help="Snapshot file path (for restore)") p_snap.add_argument("--file", help="Snapshot file path (for restore)")
args = parser.parse_args() args = parser.parse_args()
+346
View File
@@ -0,0 +1,346 @@
"""Tests for the GPU process list + kill endpoints.
Standalone (no pytest required):
python tests/test_processes.py
Also works under pytest if available. Covers the /proc reader (including
comm values with spaces and parentheses), the kill endpoint's input
validation, a real SIGTERM round-trip, and the GET /api/processes payload.
"""
import contextlib
import os
import subprocess
import sys
import time
sys.path.insert(0, os.path.join(os.path.dirname(__file__), ".."))
from nvcurve import server # noqa: E402
from nvcurve.hal import processes # noqa: E402
PASS = 0
FAIL = 0
def check(name: str, cond: bool) -> None:
global PASS, FAIL
if cond:
PASS += 1
print(f" PASS {name}")
else:
FAIL += 1
print(f" FAIL {name}")
def test_proc_info_self() -> None:
"""_proc_info must return sane data for our own process."""
info = processes._proc_info(os.getpid())
check("self info not None", info is not None)
if info is None:
return
check(
"user is non-empty string",
bool(isinstance(info["user"], str) and info["user"]),
)
check(
"cpu_pct is None or non-negative number",
info["cpu_pct"] is None or (isinstance(info["cpu_pct"], float) and info["cpu_pct"] >= 0),
)
check("mem_bytes > 0", info["mem_bytes"] is not None and info["mem_bytes"] > 0)
check("command non-empty", isinstance(info["command"], str) and len(info["command"]) > 0)
def test_proc_info_missing() -> None:
"""A pid that does not exist must yield None (not an exception)."""
check("missing pid -> None", processes._proc_info(99999999) is None)
def test_proc_info_comm_with_parens() -> None:
"""comm containing spaces and parentheses must not break /proc/stat parsing.
The kernel comm is set via prctl(PR_SET_NAME); `exec -a` only changes
argv[0] and would leave comm as the executable name, so the parens path
would never be exercised.
"""
proc = subprocess.Popen(
[
sys.executable,
"-c",
"import ctypes, time\n"
"ctypes.CDLL(None).prctl(15, b'proc (test) x', 0, 0, 0)\n"
"time.sleep(30)",
]
)
try:
# Wait until /proc/<pid>/stat exists and comm actually carries the
# parenthesised name (prctl may take a moment after exec).
comm = ""
for _ in range(100):
try:
with open(f"/proc/{proc.pid}/stat") as f:
raw = f.read()
comm = raw[raw.index("(") + 1 : raw.rfind(")")]
except OSError:
pass
if "(" in comm:
break
time.sleep(0.05)
check("fixture comm contains parens", "(" in comm)
info = processes._proc_info(proc.pid)
check("weird-comm info not None", info is not None)
if info is not None:
check("weird-comm command parsed", "prctl" in info["command"])
finally:
proc.kill()
proc.wait()
def test_kill_endpoint_validation() -> None:
"""The kill endpoint must reject bad input and unknown pids."""
from fastapi.testclient import TestClient
client = TestClient(server.app)
r = client.post(
"/api/processes/kill", json={"pid": os.getpid(), "signal": "BOGUS"}
)
check("invalid signal -> 400", r.status_code == 400)
r = client.post("/api/processes/kill", json={"pid": -5, "signal": "TERM"})
check("negative pid -> 400", r.status_code == 400)
r = client.post("/api/processes/kill", json={"pid": 0, "signal": "TERM"})
check("zero pid -> 400", r.status_code == 400)
r = client.post("/api/processes/kill", json={"pid": 99999999, "signal": "TERM"})
check("unknown pid -> 404", r.status_code == 404)
def test_kill_endpoint_real() -> None:
"""A real SIGTERM round-trip: the process must actually exit."""
from fastapi.testclient import TestClient
client = TestClient(server.app)
proc = subprocess.Popen(["sleep", "30"])
try:
r = client.post(
"/api/processes/kill", json={"pid": proc.pid, "signal": "TERM"}
)
check("kill -> 200", r.status_code == 200)
check("response ok", r.json().get("ok") is True)
try:
proc.wait(timeout=5)
check("process exited", True)
except subprocess.TimeoutExpired:
check("process exited", False)
finally:
proc.kill()
proc.wait()
def test_kill_endpoint_refuses_init_kthreadd() -> None:
"""The endpoint must refuse to signal init (pid 1) or kthreadd (pid 2).
SIGTERM to init is a system-shutdown vector; kernel threads don't handle
signals anyway. This fires before any NVML/state access, so no GPU needed.
"""
from fastapi.testclient import TestClient
client = TestClient(server.app)
for pid in (1, 2):
r = client.post(
"/api/processes/kill", json={"pid": pid, "signal": "KILL"}
)
check(f"pid {pid} refused -> 400", r.status_code == 400)
def test_get_parent_pid() -> None:
"""get_parent_pid must return the kernel ppid and degrade gracefully."""
check("self ppid matches", processes.get_parent_pid(os.getpid()) == os.getppid())
check("missing pid -> None", processes.get_parent_pid(99999999) is None)
check("pid 1 has no parent", processes.get_parent_pid(1) == 0)
def test_kill_parent_endpoint() -> None:
"""parent=true must signal the parent, leaving the child alive."""
from fastapi.testclient import TestClient
client = TestClient(server.app)
# Validation cases first (no fixture needed).
r = client.post(
"/api/processes/kill", json={"pid": 99999999, "signal": "KILL", "parent": True}
)
check("parent of missing pid -> 404", r.status_code == 404)
r = client.post(
"/api/processes/kill", json={"pid": 1, "signal": "KILL", "parent": True}
)
check("pid 1 has no parent -> 400", r.status_code == 400)
# bash forks a background sleep (child), then sleeps itself (parent).
parent = subprocess.Popen(["bash", "-c", "sleep 30 & sleep 30"])
child: int | None = None
for _ in range(100):
try:
with open(f"/proc/{parent.pid}/task/{parent.pid}/children") as f:
kids = f.read().split()
if kids:
child = int(kids[0])
break
except OSError:
pass
time.sleep(0.05)
try:
check("fixture has a child", child is not None)
if child is None:
return
r = client.post(
"/api/processes/kill",
json={"pid": child, "signal": "KILL", "parent": True},
)
check("kill parent -> 200", r.status_code == 200)
check("response targets the parent", r.json().get("pid") == parent.pid)
check("response reports parent flag", r.json().get("parent") is True)
try:
parent.wait(timeout=5)
check("parent died", True)
except subprocess.TimeoutExpired:
check("parent died", False)
time.sleep(0.2)
try:
os.kill(child, 0)
check("child survived (orphaned)", True)
except ProcessLookupError:
check("child survived (orphaned)", False)
finally:
for victim in (parent,):
with contextlib.suppress(ProcessLookupError, subprocess.TimeoutExpired):
victim.kill()
victim.wait(timeout=5)
if child is not None:
with contextlib.suppress(ProcessLookupError):
os.kill(child, 9)
def test_processes_endpoint() -> None:
"""GET /api/processes must return the HAL payload for a known GPU."""
from fastapi.testclient import TestClient
fake = {
"processes": [
{
"pid": 1,
"user": "root",
"dev": 3,
"type": "C",
"gpu_util_pct": 42.0,
"mem_util_pct": 1.0,
"vram_bytes": 1000,
"vram_pct": 0.1,
"cpu_pct": 1.5,
"mem_host_bytes": 2048,
"command": "/usr/bin/fake",
}
],
"mem_total_bytes": 1000000,
}
orig = server.list_gpu_processes
server._state["gpus"][3] = {"gpu": object(), "gpu_name": "fake"}
server.list_gpu_processes = lambda idx: fake
try:
client = TestClient(server.app)
r = client.get("/api/processes", params={"gpu_index": 3})
check("processes -> 200", r.status_code == 200)
check("payload matches HAL", r.json() == fake)
r = client.get("/api/processes", params={"gpu_index": 99})
check("unknown gpu -> 404", r.status_code == 404)
finally:
server.list_gpu_processes = orig
server._state["gpus"].pop(3, None)
def test_list_gpu_processes_shape() -> None:
"""With NVML available, list_gpu_processes must return a well-shaped payload."""
from nvcurve.hal import monitoring
if not monitoring._NVML_AVAILABLE:
print(" SKIP NVML not available")
return
if not monitoring.init_nvml():
print(" SKIP NVML init failed (no driver?)")
return
try:
data = processes.list_gpu_processes(0)
check("has processes list", isinstance(data.get("processes"), list))
check("has mem_total_bytes", "mem_total_bytes" in data)
keys = {
"pid",
"user",
"dev",
"type",
"gpu_util_pct",
"mem_util_pct",
"vram_bytes",
"vram_pct",
"cpu_pct",
"mem_host_bytes",
"command",
}
for p in data["processes"]:
if not keys.issubset(p.keys()):
check(f"process {p.get('pid')} has all keys", False)
break
else:
check("all processes have all keys", True)
# Sorted by VRAM descending.
vram = [p["vram_bytes"] or 0 for p in data["processes"]]
check("sorted by VRAM desc", vram == sorted(vram, reverse=True))
finally:
monitoring.shutdown_nvml()
def test_list_gpu_processes_no_nvml() -> None:
"""With NVML unavailable, list_gpu_processes returns an empty payload.
This is the branch the shape test skips (no driver): the UI uses
mem_total_bytes is None to render 'data unavailable' instead of
'no processes'.
"""
from nvcurve.hal import monitoring
orig_avail = monitoring._NVML_AVAILABLE
orig_init = monitoring._nvml_initialized
monitoring._NVML_AVAILABLE = False
monitoring._nvml_initialized = False
try:
data = processes.list_gpu_processes(0)
check("empty processes list", data["processes"] == [])
check("mem_total_bytes is None", data["mem_total_bytes"] is None)
finally:
monitoring._NVML_AVAILABLE = orig_avail
monitoring._nvml_initialized = orig_init
def main() -> None:
tests = [
test_proc_info_self,
test_proc_info_missing,
test_proc_info_comm_with_parens,
test_kill_endpoint_validation,
test_kill_endpoint_real,
test_kill_endpoint_refuses_init_kthreadd,
test_get_parent_pid,
test_kill_parent_endpoint,
test_processes_endpoint,
test_list_gpu_processes_shape,
test_list_gpu_processes_no_nvml,
]
for t in tests:
print(f"== {t.__name__} ==")
t()
print()
print(f"{PASS} passed, {FAIL} failed")
if FAIL:
sys.exit(1)
if __name__ == "__main__":
main()
+355
View File
@@ -0,0 +1,355 @@
"""Unit tests for the RM power-limit interface (fake RM, no hardware).
Standalone (no pytest required):
python tests/test_rm_power.py
Also works under pytest if available. Ports the test battery from LACT PR
#1205 (ilya-zlobintsev/LACT): layout discovery, NVML cross-validation,
write minimality, readback verification, and failure restoration.
"""
import os
import sys
sys.path.insert(0, os.path.join(os.path.dirname(__file__), ".."))
from nvcurve.hal.rm_power import ( # noqa: E402
_CTRL_GPU_GET_ATTACHED_IDS,
_CTRL_GPU_GET_ID_INFO_V2,
_CTRL_GPU_GET_PCI_INFO,
_PWR_GET_CONTROL,
_PWR_GET_INFO,
_PWR_SET_CONTROL,
EXTENDED_LAYOUT,
LEGACY_LAYOUT,
PciLocation,
PowerLimitBounds,
RmPowerError,
_u32,
probe,
resolve_gpu_instance,
set_limit,
)
PASS = 0
FAIL = 0
def check(name: str, cond: bool) -> None:
global PASS, FAIL
if cond:
PASS += 1
print(f" PASS {name}")
else:
FAIL += 1
print(f" FAIL {name}")
BOUNDS = PowerLimitBounds(min_mw=250_000, default_mw=300_000, max_mw=325_000)
class FakeRm:
"""In-memory fake of the RM power-limit client (both wire layouts)."""
def __init__(self, layout, current: int) -> None:
self.layout = layout
self.control = bytearray(layout.control_size)
self.control[0:8] = bytes([0xFF, 0, 0, 0, 1, 0, 0, 0])
self.control[layout.request_at - 4 : layout.request_at] = bytes(
[0x67, 0x67, 0, 0]
)
self.control[layout.request_at : layout.request_at + 4] = current.to_bytes(
4, "little"
)
self.control[layout.client_at] = 0xFE
self.reads: list[tuple[int, int]] = []
self.writes: list[bytes] = []
self.fail_first_write = False
self.fail_readback = False
self.fail_restore = False
def query(self, cmd: int, data: bytearray) -> None:
if cmd == _PWR_GET_INFO:
self.reads.append((cmd, len(data)))
if len(data) != self.layout.info_size:
raise RmPowerError("Unsupported INFO size")
data[0:8] = bytes([0xFF, 0, 0, 0, 1, 0, 0, 0])
for index, value in enumerate([250_000, 300_000, 325_000]):
offset = self.layout.info_min_at + 4 * index
data[offset : offset + 4] = value.to_bytes(4, "little")
elif cmd == _PWR_GET_CONTROL:
self.reads.append((cmd, len(data)))
if len(data) != self.layout.control_size:
raise RmPowerError("Unsupported CONTROL size")
if data[self.layout.client_at] != 0xFE:
raise AssertionError("unexpected client selector on GET")
if self.fail_readback and len(self.writes) == 1:
raise RmPowerError("readback unavailable")
data[:] = self.control
elif cmd == _PWR_SET_CONTROL:
if len(data) != self.layout.control_size:
raise AssertionError("bad SET size")
if data[4:8] != (1).to_bytes(4, "little"):
raise AssertionError("bad SET mask")
if data[self.layout.client_at] != 0xFE:
raise AssertionError("bad SET client selector")
# Only the request field may differ from the current state.
for i, (a, b) in enumerate(zip(data, self.control, strict=True)):
if self.layout.request_at <= i < self.layout.request_at + 4:
continue
if a != b:
raise AssertionError(f"SET modified byte {i:#x}")
self.writes.append(bytes(data))
if self.fail_restore and len(self.writes) > 1:
raise RmPowerError("restore unavailable")
self.control[:] = data
if self.fail_first_write and len(self.writes) == 1:
raise RmPowerError("SET failed after modifying hardware")
else:
raise AssertionError(f"unexpected command {cmd:#x}")
def test_detects_both_layouts_with_gets_without_a_driver_version() -> None:
for layout in (EXTENDED_LAYOUT, LEGACY_LAYOUT):
rm = FakeRm(layout, 250_000)
support = probe(BOUNDS, 250_000, rm.query)
check(
f"{layout.name}: detected",
support.bounds == BOUNDS and support.layout == layout,
)
expected = (
[(_PWR_GET_INFO, 0x924), (_PWR_GET_CONTROL, 0x328)]
if layout == EXTENDED_LAYOUT
else [
(_PWR_GET_INFO, 0x924),
(_PWR_GET_INFO, 0x488),
(_PWR_GET_CONTROL, 0x188),
]
)
check(f"{layout.name}: GETs only, expected sequence", rm.reads == expected)
check(f"{layout.name}: no writes during discovery", rm.writes == [])
def test_unknown_layout_and_nvml_mismatches_never_write() -> None:
calls: list[tuple[int, int]] = []
def failing(cmd: int, data: bytearray) -> None:
calls.append((cmd, len(data)))
raise RmPowerError("Unsupported payload")
try:
probe(BOUNDS, 250_000, failing)
check("unknown layout rejected", False)
except RmPowerError:
check("unknown layout rejected", True)
check(
"unknown layout: only GETs attempted",
calls == [(_PWR_GET_INFO, 0x924), (_PWR_GET_INFO, 0x488)],
)
for layout in (EXTENDED_LAYOUT, LEGACY_LAYOUT):
rm = FakeRm(layout, 250_000)
try:
probe(BOUNDS, 300_000, rm.query)
check(f"{layout.name}: current mismatch rejected", False)
except RmPowerError:
check(f"{layout.name}: current mismatch rejected", True)
other_bounds = PowerLimitBounds(
min_mw=BOUNDS.min_mw, default_mw=BOUNDS.default_mw, max_mw=350_000
)
try:
probe(other_bounds, 250_000, rm.query)
check(f"{layout.name}: bounds mismatch rejected", False)
except RmPowerError:
check(f"{layout.name}: bounds mismatch rejected", True)
check(f"{layout.name}: no writes on mismatch", rm.writes == [])
def test_rejects_unrecognized_headers_masks_and_client_values() -> None:
for layout in (EXTENDED_LAYOUT, LEGACY_LAYOUT):
for at, value in [(0, 0), (4, 3), (layout.client_at, 0xF8)]:
rm = FakeRm(layout, 250_000)
rm.control[at] = value
try:
probe(BOUNDS, 250_000, rm.query)
check(f"{layout.name}: bad header/client rejected", False)
except RmPowerError:
check(f"{layout.name}: bad header/client rejected", True)
check(f"{layout.name}: no writes on bad header", rm.writes == [])
for current in (0, 0xFFFFFFFF):
rm = FakeRm(layout, current)
try:
probe(BOUNDS, current, rm.query)
check(f"{layout.name}: empty request rejected", False)
except RmPowerError:
check(f"{layout.name}: empty request rejected", True)
# The extended layout has additional mask words. Accepting only its low
# word would allow an unexpected client to be included in a later SET.
rm = FakeRm(EXTENDED_LAYOUT, 250_000)
rm.control[8] = 1
try:
probe(BOUNDS, 250_000, rm.query)
check("extended: nonzero mask word rejected", False)
except RmPowerError:
check("extended: nonzero mask word rejected", True)
check("extended: no writes on mask violation", rm.writes == [])
def test_changes_only_fe_request_and_keeps_vbios_maximum() -> None:
for layout in (EXTENDED_LAYOUT, LEGACY_LAYOUT):
rm = FakeRm(layout, 250_000)
support = probe(BOUNDS, 250_000, rm.query)
for cap in (150_000, 30_000, 250_000):
set_limit(cap, support, rm.query)
check(
f"{layout.name}: set {cap} mW",
_u32(rm.control, layout.request_at) == cap,
)
writes = len(rm.writes)
for cap in (0, 29_999, 325_001, 350_000, 0xFFFFFFFF):
try:
set_limit(cap, support, rm.query)
check(f"{layout.name}: out-of-range {cap} rejected", False)
except RmPowerError:
check(f"{layout.name}: out-of-range {cap} rejected", True)
check(
f"{layout.name}: no writes for out-of-range caps",
len(rm.writes) == writes,
)
def test_restores_previous_below_minimum_request_after_set_or_readback_failure() -> (
None
):
for layout in (EXTENDED_LAYOUT, LEGACY_LAYOUT):
for fail_set in (False, True):
rm = FakeRm(layout, 100_000)
original = bytes(rm.control)
rm.fail_first_write = fail_set
rm.fail_readback = not fail_set
support = probe(BOUNDS, 100_000, rm.query)
try:
set_limit(150_000, support, rm.query)
check(f"{layout.name}: failure reported", False)
except RmPowerError:
check(f"{layout.name}: failure reported", True)
check(f"{layout.name}: restore issued", len(rm.writes) == 2)
check(
f"{layout.name}: previous request restored",
bytes(rm.control) == original,
)
def test_reports_restore_failure_and_rejects_wrong_client_before_writing() -> None:
for layout in (EXTENDED_LAYOUT, LEGACY_LAYOUT):
rm = FakeRm(layout, 100_000)
rm.fail_first_write = True
rm.fail_restore = True
support = probe(BOUNDS, 100_000, rm.query)
try:
set_limit(150_000, support, rm.query)
check(f"{layout.name}: restore failure reported", False)
except RmPowerError as exc:
check(
f"{layout.name}: restore failure reported",
"restoration also failed" in str(exc),
)
rm = FakeRm(layout, 250_000)
rm.control[layout.client_at] = 0xF8
try:
set_limit(150_000, support, rm.query)
check(f"{layout.name}: wrong client rejected", False)
except RmPowerError:
check(f"{layout.name}: wrong client rejected", True)
check(f"{layout.name}: no writes for wrong client", rm.writes == [])
# ── PCI identity → RM instance resolution ────────────────────────────────────
def test_resolves_pci_identity_when_minor_and_rm_orders_differ() -> None:
# This host has Ada at minor 5/RM 4 and the 5090 at minor 4/RM 5.
# IDs are opaque and enumeration order must not select the device.
pci = PciLocation(domain=0, bus=0x0D, dev=0, func=0)
instances = resolve_gpu_instance(pci, lambda cmd, data: _fake_root(cmd, data))
check("resolves by PCI identity", instances == (5, 2))
def _fake_root(cmd: int, data: bytearray) -> None:
if cmd == _CTRL_GPU_GET_ATTACHED_IDS:
data[0:4] = (0x2E00).to_bytes(4, "little")
data[4:8] = (0x0D00).to_bytes(4, "little")
elif cmd == _CTRL_GPU_GET_PCI_INFO:
gpu_id = _u32(data, 0)
bus = 0x2E if gpu_id == 0x2E00 else 0x0D
data[8:10] = bus.to_bytes(2, "little")
elif cmd == _CTRL_GPU_GET_ID_INFO_V2:
if _u32(data, 0) != 0x0D00:
raise AssertionError("unexpected gpu id in ID_INFO_V2")
data[8:12] = (5).to_bytes(4, "little")
data[12:16] = (2).to_bytes(4, "little")
else:
raise AssertionError(f"unexpected command {cmd:#x}")
def test_does_not_fall_back_to_another_gpu_when_pci_is_missing() -> None:
pci = PciLocation(domain=1, bus=0x0D, dev=0, func=0)
def query(cmd: int, data: bytearray) -> None:
if cmd == _CTRL_GPU_GET_ATTACHED_IDS:
data[0:4] = (0x0D00).to_bytes(4, "little")
elif cmd == _CTRL_GPU_GET_PCI_INFO:
data[8:10] = (0x0D).to_bytes(2, "little")
else:
raise AssertionError("must not allocate a GPU from another PCI domain")
try:
resolve_gpu_instance(pci, query)
check("foreign PCI domain rejected", False)
except RmPowerError:
check("foreign PCI domain rejected", True)
def test_rejects_nonzero_pci_function() -> None:
pci = PciLocation(domain=0, bus=0x0D, dev=0, func=1)
try:
resolve_gpu_instance(pci, lambda cmd, data: None)
check("nonzero function rejected", False)
except RmPowerError:
check("nonzero function rejected", True)
def test_propagates_rm_query_failure() -> None:
pci = PciLocation(domain=0, bus=0x0D, dev=0, func=0)
try:
resolve_gpu_instance(
pci, lambda cmd, data: (_ for _ in ()).throw(RmPowerError("RM unavailable"))
)
check("RM query failure propagated", False)
except RmPowerError as exc:
check("RM query failure propagated", "RM unavailable" in str(exc))
def main() -> int:
tests = [
test_detects_both_layouts_with_gets_without_a_driver_version,
test_unknown_layout_and_nvml_mismatches_never_write,
test_rejects_unrecognized_headers_masks_and_client_values,
test_changes_only_fe_request_and_keeps_vbios_maximum,
test_restores_previous_below_minimum_request_after_set_or_readback_failure,
test_reports_restore_failure_and_rejects_wrong_client_before_writing,
test_resolves_pci_identity_when_minor_and_rm_orders_differ,
test_does_not_fall_back_to_another_gpu_when_pci_is_missing,
test_rejects_nonzero_pci_function,
test_propagates_rm_query_failure,
]
for t in tests:
print(f"== {t.__name__} ==")
t()
print(f"\n{PASS} passed, {FAIL} failed")
return 1 if FAIL else 0
if __name__ == "__main__":
sys.exit(main())
+263
View File
@@ -0,0 +1,263 @@
"""Security regression tests for nvcurve.
Standalone (no pytest required):
python tests/test_security.py
Also works under pytest if available. Covers the security-critical logic:
SPA path containment, snapshot restore containment, login lockout client-IP
derivation, TLS scheme detection, daemon socket hardening, and the
server-enforced safety cap.
"""
import asyncio
import builtins
import io
import os
import sys
import tempfile
sys.path.insert(0, os.path.join(os.path.dirname(__file__), ".."))
from nvcurve import (
daemon, # noqa: E402
server, # noqa: E402
)
from nvcurve.cli import _discover_server_url # noqa: E402
from nvcurve.config import ( # noqa: E402
Config,
normalize_trusted_proxies,
tls_enabled,
)
from nvcurve.hal.snapshot import restore as snapshot_restore # noqa: E402
PASS = 0
FAIL = 0
def check(name: str, cond: bool) -> None:
global PASS, FAIL
if cond:
PASS += 1
print(f" PASS {name}")
else:
FAIL += 1
print(f" FAIL {name}")
def _encoded_traversal(target: str = "/etc/hostname") -> str:
"""Build an encoded '..'-based traversal path deep enough to escape any dist dir."""
dist = os.path.abspath(server._dist_dir)
depth = len(dist.rstrip("/").split("/"))
enc = lambda s: s.replace("/", "%2f") # noqa: E731
return "/" + enc("../*" + str(depth + 2) + target)
def test_spa_path_containment() -> None:
"""The SPA catch-all must not serve files outside the dist directory."""
from fastapi.testclient import TestClient
client = TestClient(server.app)
r = client.get(_encoded_traversal())
check("SPA traversal -> 404", r.status_code == 404)
r = client.get(_encoded_traversal("/etc/passwd"))
check("SPA traversal /etc/passwd -> 404", r.status_code == 404)
# Normal static files must still be served.
r = client.get("/index.html")
check("SPA /index.html -> 200", r.status_code == 200)
r = client.get("/api/nonexistent")
check("unknown /api/ path -> 404", r.status_code == 404)
def test_snapshot_restore_containment() -> None:
"""Snapshot restore must only read files inside the snapshot dir."""
tmp = tempfile.mkdtemp()
snap_dir = os.path.join(tmp, "snaps")
os.makedirs(snap_dir)
outside = os.path.join(tmp, "evil.bin")
with open(outside, "wb") as f:
f.write(b"\x00" * 9248)
check(
"restore(outside file) rejected",
snapshot_restore(None, snap_dir, outside) is False,
)
check(
"restore(nonexistent) rejected",
snapshot_restore(None, snap_dir, "/etc/hostname") is False,
)
link = os.path.join(snap_dir, "link.bin")
os.symlink(outside, link)
check(
"restore(symlink escape) rejected",
snapshot_restore(None, snap_dir, link) is False,
)
def test_safety_cap_not_client_overridable() -> None:
"""The API must not accept a per-request safety cap override."""
check(
"WriteRequest has no max_delta_khz field",
"max_delta_khz" not in server.WriteRequest.model_fields,
)
check(
"GlobalOffsetRequest has no max_delta_khz field",
"max_delta_khz" not in server.GlobalOffsetRequest.model_fields,
)
def test_client_ip_derivation() -> None:
"""X-Forwarded-For is only honoured for configured trusted proxies."""
class FakeReq:
def __init__(self, peer: str, headers: dict):
self.client = type("C", (), {"host": peer})()
self.headers = headers
check(
"trusted proxy -> XFF used",
server._client_ip(
FakeReq("127.0.0.1", {"x-forwarded-for": "9.9.9.9"}), ["127.0.0.1"]
)
== "9.9.9.9",
)
check(
"untrusted peer -> XFF ignored",
server._client_ip(
FakeReq("8.8.8.8", {"x-forwarded-for": "9.9.9.9"}), ["127.0.0.1"]
)
== "8.8.8.8",
)
check(
"rightmost untrusted hop",
server._client_ip(
FakeReq("127.0.0.1", {"x-forwarded-for": "127.0.0.1, 9.9.9.9"}),
["127.0.0.1"],
)
== "9.9.9.9",
)
check(
"no XFF -> peer",
server._client_ip(FakeReq("127.0.0.1", {}), ["127.0.0.1"]) == "127.0.0.1",
)
check(
"all-trusted chain -> peer",
server._client_ip(
FakeReq("127.0.0.1", {"x-forwarded-for": "10.0.0.1, 10.0.0.2"}),
["127.0.0.1", "10.0.0.1", "10.0.0.2"],
)
== "127.0.0.1",
)
def test_trusted_proxies_normalization() -> None:
"""String values must be normalized to lists (no substring matching)."""
check("list passthrough", normalize_trusted_proxies(["1.2.3.4"]) == ["1.2.3.4"])
check(
"comma string split",
normalize_trusted_proxies("1.2.3.4, 5.6.7.8") == ["1.2.3.4", "5.6.7.8"],
)
check("None -> []", normalize_trusted_proxies(None) == [])
check("junk -> []", normalize_trusted_proxies(42) == [])
# The original bug: substring membership. After normalization, "127.0.0.1"
# must NOT be trusted when only "127.0.0.10" is listed.
trusted = normalize_trusted_proxies(["127.0.0.10"])
check("no substring trust", "127.0.0.1" not in trusted)
def test_tls_scheme_detection() -> None:
"""_discover_server_url must pick https when TLS is configured."""
from nvcurve import cli as cli_mod
# Hermetic: hide any live server's runtime info file and any existing
# persistent config so the defaults level of the priority chain is
# exercised.
real_info_file = cli_mod._SERVER_INFO_FILE
real_persistent_cfg = cli_mod._PERSISTENT_CONFIG_FILE
hidden = tempfile.mkdtemp()
cli_mod._SERVER_INFO_FILE = os.path.join(hidden, "nvcurve.json")
cli_mod._PERSISTENT_CONFIG_FILE = os.path.join(hidden, "config.json")
try:
cfg = Config(
host="10.0.0.5", port=9000, ssl_certfile="/x/c.pem", ssl_keyfile="/x/k.pem"
)
check("tls_enabled true", tls_enabled(cfg) is True)
check("https url", _discover_server_url(cfg) == "https://10.0.0.5:9000")
cfg2 = Config(host="10.0.0.5", port=9000)
check("http url", _discover_server_url(cfg2) == "http://10.0.0.5:9000")
finally:
cli_mod._SERVER_INFO_FILE = real_info_file
cli_mod._PERSISTENT_CONFIG_FILE = real_persistent_cfg
def test_daemon_ignores_caller_host_port() -> None:
"""serve_start must bind the configured host/port, never the caller's."""
daemon._cfg = Config(host="10.1.1.1", port=9999)
captured: dict = {}
class FakePopen:
def __init__(self, cmd, **kw):
captured["cmd"] = cmd
self.pid = 4242
def poll(self):
return 0
def terminate(self):
pass
def wait(self, *a):
return 0
real_open = builtins.open
def fake_open(path, *a, **kw):
if str(path).endswith("nvcurve-server.log"):
return io.StringIO()
return real_open(path, *a, **kw)
orig_popen = daemon.subprocess.Popen
daemon.subprocess.Popen = FakePopen
builtins.open = fake_open
try:
loop = asyncio.new_event_loop()
resp = loop.run_until_complete(
# "0.0.0.0" is a test payload proving the daemon ignores caller
# host/port — no socket is bound here.
daemon._dispatch({"cmd": "serve_start", "host": "0.0.0.0", "port": 12345}) # noqa: S104
)
finally:
builtins.open = real_open
daemon.subprocess.Popen = orig_popen
cmd = " ".join(captured.get("cmd", []))
check("config host/port used", "10.1.1.1" in cmd and "9999" in cmd)
check("caller host/port ignored", "0.0.0.0" not in cmd and "12345" not in cmd) # noqa: S104
check(
"response reports configured values",
resp.get("ok") is True
and resp.get("host") == "10.1.1.1"
and resp.get("port") == 9999,
)
check("response includes tls flag", "tls" in resp)
def main() -> int:
tests = [
test_spa_path_containment,
test_snapshot_restore_containment,
test_safety_cap_not_client_overridable,
test_client_ip_derivation,
test_trusted_proxies_normalization,
test_tls_scheme_detection,
test_daemon_ignores_caller_host_port,
]
for t in tests:
print(f"== {t.__name__} ==")
t()
print(f"\n{PASS} passed, {FAIL} failed")
return 1 if FAIL else 0
if __name__ == "__main__":
sys.exit(main())
+541
View File
@@ -0,0 +1,541 @@
"""Tests for the native WireView Pro II reader (nvcurve/wireview.py).
Standalone (no pytest required):
python tests/test_wireview.py
Also works under pytest. Covers the sensor-frame parser, corruption
detection, fault decoding, USB port discovery, and the hwmon (sysfs)
transport.
"""
import builtins
import io
import os
import struct
import sys
import tempfile
from unittest import mock
sys.path.insert(0, os.path.join(os.path.dirname(__file__), ".."))
from nvcurve import wireview as wv # noqa: E402
PASS = 0
FAIL = 0
def check(name: str, cond: bool) -> None:
global PASS, FAIL
if cond:
PASS += 1
print(f" PASS {name}")
else:
FAIL += 1
print(f" FAIL {name}")
def make_frame(
ts=(392, 352, 273, 395),
vdd=11900,
fan=0,
pins=((11904, 6221, 74055),) * 6,
total_power=450568,
total_current=37863,
avg_voltage=11900,
psu_cap=0,
fault_status=0,
fault_log=0,
pad1=0,
pad2=0,
) -> bytes:
"""Build a 100-byte sensor frame with controllable fields."""
frame = struct.pack("<4hHB", *ts, vdd, fan)
frame += bytes([pad1])
for voltage, current, power in pins:
frame += struct.pack("<hxxII", voltage, current, power)
frame += struct.pack("<IIHB", total_power, total_current, avg_voltage, psu_cap)
frame += bytes([pad2])
frame += struct.pack("<HH", fault_status, fault_log)
assert len(frame) == wv.SENSOR_STRUCT_SIZE
return frame
# ── Sensor frame parser ───────────────────────────────────────────────────────
def test_parse_sensor_struct():
s = wv.parse_sensor_struct(make_frame())
check("temp in", abs(s["temp_in_c"] - 39.2) < 1e-9)
check("temp out", abs(s["temp_out_c"] - 35.2) < 1e-9)
check("temp ext1", abs(s["temp_ext1_c"] - 27.3) < 1e-9)
check("temp ext2", abs(s["temp_ext2_c"] - 39.5) < 1e-9)
check("fan duty", s["fan_duty_pct"] == 0)
check("psu capability", s["psu_capability_w"] == 600)
check("fault status", s["fault_status"] == 0)
check("fault log", s["fault_log"] == 0)
check("6 pins", len(s["pins"]) == 6)
pin = s["pins"][0]
check("pin voltage", abs(pin["voltage_v"] - 11.904) < 1e-9)
check("pin current", abs(pin["current_a"] - 6.221) < 1e-9)
check("pin power", abs(pin["power_w"] - 74.055) < 1e-9)
# Totals are computed from the pins (exporter behavior), not the
# device's total fields.
check(
"total current",
abs(s["current_total_a"] - round(6 * 6.221, 3)) < 1e-9,
)
check(
"total power",
abs(s["power_total_w"] - round(6 * 11.904 * 6.221, 3)) < 1e-9,
)
check(
"pin imbalance (uniform)",
s["pin_imbalance_a"] == 0.0,
)
# Imbalance is the spread between the highest- and lowest-loaded pin.
uneven = wv.parse_sensor_struct(
make_frame(pins=((11904, 6221, 74055),) * 5 + ((11900, 1100, 13090),))
)
check(
"pin imbalance (uneven)",
abs(uneven["pin_imbalance_a"] - 5.121) < 1e-9,
)
check(
"avg voltage",
abs(s["voltage_avg_v"] - 11.904) < 1e-3,
)
def test_parse_psu_capabilities():
for cap, watts in ((0, 600), (1, 450), (2, 300), (3, 150)):
s = wv.parse_sensor_struct(make_frame(psu_cap=cap))
check(f"psu cap {cap} -> {watts}W", s["psu_capability_w"] == watts)
def test_parse_negative_temps():
s = wv.parse_sensor_struct(make_frame(ts=(-10, 0, 555, 999)))
check("negative temp", abs(s["temp_in_c"] - (-1.0)) < 1e-9)
check("zero temp", s["temp_out_c"] == 0.0)
check("high temp", abs(s["temp_ext2_c"] - 99.9) < 1e-9)
# ── Corruption detection ──────────────────────────────────────────────────────
def test_corruption_check():
check("clean frame accepted", not wv.sensor_frame_is_corrupt(make_frame()))
check(
"fan duty > 100 rejected",
wv.sensor_frame_is_corrupt(make_frame(fan=101)),
)
check(
"fan duty 100 accepted",
not wv.sensor_frame_is_corrupt(make_frame(fan=100)),
)
check("pad1 dirty rejected", wv.sensor_frame_is_corrupt(make_frame(pad1=1)))
check("pad2 dirty rejected", wv.sensor_frame_is_corrupt(make_frame(pad2=1)))
check(
"short frame rejected",
wv.sensor_frame_is_corrupt(make_frame()[:50]),
)
# ── Fault decoding ────────────────────────────────────────────────────────────
def test_decode_faults():
check("no faults", wv.decode_faults(0) == [])
check(
"chip over-temp",
wv.decode_faults(1) == ["Chip over-temperature"],
)
check(
"over-power",
wv.decode_faults(16) == ["Over-power (OPP)"],
)
check(
"multiple faults",
wv.decode_faults(1 | 32)
== ["Chip over-temperature", "Current imbalance"],
)
check(
"unknown bits ignored",
wv.decode_faults(1 | 0x80) == ["Chip over-temperature"],
)
def test_product_support():
check("Pro II supported", wv.is_supported_product(0xEF, 0x05))
check("Noctua Edition supported", wv.is_supported_product(0xEF, 0x06))
check("WireView II unsupported", not wv.is_supported_product(0xEF, 0x07))
check("other vendor unsupported", not wv.is_supported_product(0x12, 0x05))
# ── USB port discovery ────────────────────────────────────────────────────────
def test_find_wireview_ports():
"""Fake a sysfs tree with one WireView (ttyACM0) and one other CDC
device (ttyACM1), plus the udev symlink."""
contents = {
"/sys/devices/fake/usb0/idVendor": "0483\n",
"/sys/devices/fake/usb0/idProduct": "5740\n",
"/sys/devices/fake/usb1/idVendor": "1234\n",
"/sys/devices/fake/usb1/idProduct": "5678\n",
}
def fake_islink(p):
return p == "/dev/wireview-pro2"
def fake_realpath(p):
return {
"/dev/wireview-pro2": "/dev/ttyACM0",
"/sys/class/tty/ttyACM0": "/sys/devices/fake/usb0",
"/sys/class/tty/ttyACM1": "/sys/devices/fake/usb1",
}.get(p, p)
def fake_isdir(p):
return p == "/sys/class/tty"
def fake_listdir(p):
return ["ttyACM0", "ttyACM1"] if p == "/sys/class/tty" else []
def fake_isfile(p):
return p in contents
def fake_exists(p):
return p == "/dev/ttyACM0"
def fake_open(p, *a, **k):
if p in contents:
return io.StringIO(contents[p])
return real_open(p, *a, **k)
real_open = open
with mock.patch.object(wv.os.path, "islink", fake_islink), mock.patch.object(
wv.os.path, "realpath", fake_realpath
), mock.patch.object(wv.os.path, "isdir", fake_isdir), mock.patch.object(
wv.os, "listdir", fake_listdir
), mock.patch.object(wv.os.path, "isfile", fake_isfile), mock.patch.object(
wv.os.path, "exists", fake_exists
), mock.patch.object(
wv.os.path, "dirname", os.path.dirname
), mock.patch.object(
builtins, "open", fake_open
):
ports = wv.find_wireview_ports()
check("exactly one port found", ports == ["/dev/ttyACM0"])
def test_find_wireview_ports_none():
with mock.patch.object(wv.os.path, "islink", lambda p: False), mock.patch.object(
wv.os.path, "isdir", lambda p: False
), mock.patch.object(wv.os.path, "exists", lambda p: False):
check("no ports", wv.find_wireview_ports() == [])
def test_find_hwmon_path():
with tempfile.TemporaryDirectory() as tmp:
# A non-wireview hwmon and a wireview one.
for name, dev in (("hwmon0", "coretemp"), ("hwmon1", "wireview")):
d = os.path.join(tmp, name)
os.makedirs(d)
with open(os.path.join(d, "name"), "w") as f:
f.write(dev + "\n")
real_join = os.path.join
real_listdir = os.listdir
def fake_join(*parts):
if parts and parts[0] == "/sys/class/hwmon":
parts = (tmp,) + parts[1:]
return real_join(*parts)
def fake_listdir(p):
if p == "/sys/class/hwmon":
return real_listdir(tmp)
return real_listdir(p)
with mock.patch.object(wv.os.path, "join", fake_join), mock.patch.object(
wv.os, "listdir", fake_listdir
):
found = wv.find_hwmon_path()
check("hwmon path found", found == os.path.join(tmp, "hwmon1"))
with mock.patch.object(wv.os.path, "isdir", lambda p: False):
check("hwmon absent", wv.find_hwmon_path() is None)
# ── hwmon transport ───────────────────────────────────────────────────────────
def test_hwmon_device():
with tempfile.TemporaryDirectory() as tmp:
p = os.path.join(tmp, "hwmon2")
os.makedirs(p)
with open(os.path.join(p, "name"), "w") as f:
f.write("wireview\n")
for i in range(6):
with open(os.path.join(p, f"in{i}_input"), "w") as f:
f.write("11904\n")
with open(os.path.join(p, f"curr{i + 1}_input"), "w") as f:
f.write("6221\n")
for name, val in (
("temp1_input", "39200"),
("temp2_input", "35200"),
("temp3_input", "27300"),
("temp4_input", "39500"),
("fault_status_raw", "0"),
("fault_log_raw", "0"),
("power1_cap", "600000000"),
("pwm1", "128"),
):
with open(os.path.join(p, name), "w") as f:
f.write(val + "\n")
dev = wv.WireViewHwmonDevice(p)
check("hwmon connect", dev.connect())
check("hwmon node exists", dev.node_exists())
s = dev.read_sample()
check("hwmon sample", s is not None)
if s:
check("hwmon temp in", abs(s["temp_in_c"] - 39.2) < 1e-9)
check("hwmon fan duty ~50%", 49 <= s["fan_duty_pct"] <= 51)
check("hwmon psu cap", s["psu_capability_w"] == 600)
check("hwmon 6 pins", len(s["pins"]) == 6)
check(
"hwmon total current",
abs(s["current_total_a"] - round(6 * 6.221, 3)) < 1e-9,
)
check("hwmon pin imbalance (uniform)", s["pin_imbalance_a"] == 0.0)
dev.close()
check("hwmon closed", not dev.connected)
# A missing temp channel must yield null (not NaN): NaN would break
# the WebSocket JSON (bare NaN token) and 500 the REST endpoint.
os.remove(os.path.join(p, "temp3_input"))
dev3 = wv.WireViewHwmonDevice(p)
dev3.connect()
s3 = dev3.read_sample()
check("hwmon missing temp sample", s3 is not None)
if s3:
check("hwmon missing temp is None", s3["temp_ext1_c"] is None)
check("hwmon missing temp present", s3["temp_in_c"] == 39.2)
import json
check(
"hwmon sample JSON-safe",
json.dumps(s3, allow_nan=False) is not None,
)
dev3.close()
# Wrong device name is rejected.
with open(os.path.join(p, "name"), "w") as f:
f.write("coretemp\n")
dev2 = wv.WireViewHwmonDevice(p)
check("wrong name rejected", not dev2.connect())
def test_hwmon_legacy_attrs():
"""Older modules expose psu_cap/fan1_input instead of power1_cap/pwm1."""
with tempfile.TemporaryDirectory() as tmp:
p = os.path.join(tmp, "hwmon3")
os.makedirs(p)
with open(os.path.join(p, "name"), "w") as f:
f.write("wireview\n")
for i in range(6):
with open(os.path.join(p, f"in{i}_input"), "w") as f:
f.write("12000\n")
with open(os.path.join(p, f"curr{i + 1}_input"), "w") as f:
f.write("1000\n")
for name, val in (
("temp1_input", "30000"),
("temp2_input", "30000"),
("temp3_input", "30000"),
("temp4_input", "30000"),
("psu_cap", "2"),
("fan1_input", "42"),
("intrusion0_alarm", "0"),
("intrusion1_alarm", "0"),
):
with open(os.path.join(p, name), "w") as f:
f.write(val + "\n")
dev = wv.WireViewHwmonDevice(p)
check("legacy connect", dev.connect())
s = dev.read_sample()
check("legacy sample", s is not None)
if s:
check("legacy psu cap 300W", s["psu_capability_w"] == 300)
check("legacy fan pct", s["fan_duty_pct"] == 42)
# ── Serial transport (pty-based fake device) ──────────────────────────────────
def _build_info_struct(product_name: str, build_info: str) -> bytes:
"""BuildStruct: VendorData(3) + ProductName(32) + BuildInfo(32) + NameLength(1)."""
return (
bytes([0xEF, 0x05, 1])
+ product_name.encode().ljust(32, b"\x00")
+ build_info.encode().ljust(32, b"\x00")
+ bytes([len(product_name)])
)
def _pty_fake_device(responses: dict):
"""Create a pty pair whose master side answers protocol commands in a
background thread. Returns (slave_path, master_fd, thread).
The fake never sends the welcome string, so the device exercises its
documented fallback: identification via the vendor-data reply.
"""
import pty
import threading
master, slave = pty.openpty()
slave_path = os.ttyname(slave)
def run():
while True:
try:
cmd = os.read(master, 1)
except OSError:
return
if not cmd:
return
resp = responses.get(cmd[0])
if resp:
try:
os.write(master, resp)
except OSError:
return
t = threading.Thread(target=run, daemon=True)
t.start()
return slave_path, master, t
def test_serial_device_protocol():
"""Full connect handshake + sensor read against a fake device."""
uid = bytes.fromhex("A7003100015045324B383120")
responses = {
wv.CMD_READ_VENDOR_DATA: bytes([0xEF, 0x05, 1]),
wv.CMD_READ_CONFIG: bytes([0, 0, 1, 0]), # version at offset 2
wv.CMD_READ_UID: uid,
wv.CMD_READ_BUILD_INFO: _build_info_struct(
"WireView Pro II", "TG-WV-PRO2-FW_20251211_1547"
),
wv.CMD_READ_SENSOR_VALUES: make_frame(),
# CMD_SCREEN_CHANGE expects no response.
}
slave_path, master, t = _pty_fake_device(responses)
try:
dev = wv.WireViewSerialDevice(slave_path)
check("serial connect", dev.connect())
check("serial not rejected", not dev.rejected)
info = dev.info()
check("serial device name", info["device_name"] == "WireView Pro II")
check("serial hw_rev", info["hw_rev"] == "EF05")
check("serial firmware", info["firmware_version"] == "1")
check("serial uid", info["uid"] == "A7003100015045324B383120")
check("serial build", info["build"] == "TG-WV-PRO2-FW_20251211_1547")
check("serial transport", info["transport"] == "serial")
s = dev.read_sample()
check("serial sample", s is not None)
if s:
check("serial sample temp", abs(s["temp_in_c"] - 39.2) < 1e-9)
check("serial sample pins", len(s["pins"]) == 6)
dev.close()
check("serial closed", not dev.connected)
finally:
os.close(master)
t.join(timeout=2)
def test_serial_device_rejected_product():
"""An unsupported product id is rejected and memoized via .rejected."""
responses = {wv.CMD_READ_VENDOR_DATA: bytes([0xEF, 0x07, 1])}
slave_path, master, t = _pty_fake_device(responses)
try:
dev = wv.WireViewSerialDevice(slave_path)
check("serial reject connect", not dev.connect())
check("serial rejected flag", dev.rejected)
finally:
os.close(master)
t.join(timeout=2)
def test_serial_device_no_response():
"""A silent port fails the connect (vendor-data read times out)."""
slave_path, master, t = _pty_fake_device({})
try:
dev = wv.WireViewSerialDevice(slave_path)
check("serial no-response connect", not dev.connect())
check("serial no-response not rejected", not dev.rejected)
check("serial no-response sample", dev.read_sample() is None)
finally:
os.close(master)
t.join(timeout=2)
def test_serial_transport_missing_pyserial():
"""A missing pyserial degrades gracefully: connect fails, reads are
None, and the warning is logged once — not on every attempt."""
import logging
records: list[str] = []
class Capture(logging.Handler):
def emit(self, record: logging.LogRecord) -> None:
records.append(record.getMessage())
logger = logging.getLogger("nvcurve.wireview")
handler = Capture()
old_level = logger.level
logger.addHandler(handler)
logger.setLevel(logging.WARNING)
try:
with mock.patch.object(wv, "serial", None):
wv._serial_missing_warned = False
dev = wv.WireViewSerialDevice("/dev/ttyACM99")
check("no-pyserial connect", not dev.connect())
check("no-pyserial not rejected", not dev.rejected)
check("no-pyserial sample", dev.read_sample() is None)
# A second attempt must not re-warn.
dev2 = wv.WireViewSerialDevice("/dev/ttyACM99")
check("no-pyserial second attempt", not dev2.connect())
warnings = [r for r in records if "pyserial" in r]
check("no-pyserial warns once", len(warnings) == 1)
finally:
logger.removeHandler(handler)
logger.setLevel(old_level)
wv._serial_missing_warned = False
def main() -> int:
print("wireview tests:")
test_parse_sensor_struct()
test_parse_psu_capabilities()
test_parse_negative_temps()
test_corruption_check()
test_decode_faults()
test_product_support()
test_find_wireview_ports()
test_find_wireview_ports_none()
test_find_hwmon_path()
test_hwmon_device()
test_hwmon_legacy_attrs()
test_serial_device_protocol()
test_serial_device_rejected_product()
test_serial_device_no_response()
test_serial_transport_missing_pyserial()
print(f"\n{PASS} passed, {FAIL} failed")
return 1 if FAIL else 0
if __name__ == "__main__":
sys.exit(main())
Generated
+80
View File
@@ -154,6 +154,22 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/04/4b/29cac41a4d98d144bf5f6d33995617b185d14b22401f75ca86f384e87ff1/h11-0.16.0-py3-none-any.whl", hash = "sha256:63cf8bbe7522de3bf65932fda1d9c2772064ffb3dae62d55932da54b31cb6c86", size = 37515, upload-time = "2025-04-24T03:35:24.344Z" }, { url = "https://files.pythonhosted.org/packages/04/4b/29cac41a4d98d144bf5f6d33995617b185d14b22401f75ca86f384e87ff1/h11-0.16.0-py3-none-any.whl", hash = "sha256:63cf8bbe7522de3bf65932fda1d9c2772064ffb3dae62d55932da54b31cb6c86", size = 37515, upload-time = "2025-04-24T03:35:24.344Z" },
] ]
[[package]]
name = "hatchling"
version = "1.32.0"
source = { registry = "https://pypi.org/simple" }
dependencies = [
{ name = "packaging" },
{ name = "pathspec" },
{ name = "pluggy" },
{ name = "tomlkit" },
{ name = "trove-classifiers" },
]
sdist = { url = "https://files.pythonhosted.org/packages/69/08/33331757185504aae48b8d9bd78cec03a76e3aecfb52e549d05a2347c0dd/hatchling-1.32.0.tar.gz", hash = "sha256:0bdbde4a52b06c37e3eca395f85a762bf0ef06fe374fd8ae429dc6be10230f5f", size = 57783, upload-time = "2026-08-11T05:03:44.114Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/a9/84/1798b6d85ecde0e31546004efd25c5de1b1f49250644a60cce460e12593a/hatchling-1.32.0-py3-none-any.whl", hash = "sha256:0e17c9c3b9aa7c625acc8d0f5b622f107d5049af9ecf5ada4de1aada5be7cdbc", size = 78435, upload-time = "2026-08-11T05:03:42.644Z" },
]
[[package]] [[package]]
name = "httpcore" name = "httpcore"
version = "1.0.9" version = "1.0.9"
@@ -230,9 +246,15 @@ dependencies = [
{ name = "httpx" }, { name = "httpx" },
{ name = "nvidia-ml-py" }, { name = "nvidia-ml-py" },
{ name = "pydantic" }, { name = "pydantic" },
{ name = "pyserial" },
{ name = "uvicorn", extra = ["standard"] }, { name = "uvicorn", extra = ["standard"] },
] ]
[package.dev-dependencies]
dev = [
{ name = "hatchling" },
]
[package.metadata] [package.metadata]
requires-dist = [ requires-dist = [
{ name = "bcrypt", specifier = ">=4.0" }, { name = "bcrypt", specifier = ">=4.0" },
@@ -240,9 +262,13 @@ requires-dist = [
{ name = "httpx", specifier = ">=0.27" }, { name = "httpx", specifier = ">=0.27" },
{ name = "nvidia-ml-py", specifier = ">=12.0" }, { name = "nvidia-ml-py", specifier = ">=12.0" },
{ name = "pydantic", specifier = ">=2.0" }, { name = "pydantic", specifier = ">=2.0" },
{ name = "pyserial", specifier = ">=3.5" },
{ name = "uvicorn", extras = ["standard"], specifier = ">=0.30" }, { name = "uvicorn", extras = ["standard"], specifier = ">=0.30" },
] ]
[package.metadata.requires-dev]
dev = [{ name = "hatchling" }]
[[package]] [[package]]
name = "nvidia-ml-py" name = "nvidia-ml-py"
version = "13.590.48" version = "13.590.48"
@@ -252,6 +278,33 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/fd/72/fb2af0d259a651affdce65fd6a495f0e07a685a0136baf585c5065204ee7/nvidia_ml_py-13.590.48-py3-none-any.whl", hash = "sha256:fd43d30ee9cd0b7940f5f9f9220b68d42722975e3992b6c21d14144c48760e43", size = 50680, upload-time = "2026-01-22T01:14:55.281Z" }, { url = "https://files.pythonhosted.org/packages/fd/72/fb2af0d259a651affdce65fd6a495f0e07a685a0136baf585c5065204ee7/nvidia_ml_py-13.590.48-py3-none-any.whl", hash = "sha256:fd43d30ee9cd0b7940f5f9f9220b68d42722975e3992b6c21d14144c48760e43", size = 50680, upload-time = "2026-01-22T01:14:55.281Z" },
] ]
[[package]]
name = "packaging"
version = "26.3"
source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/7d/fa/3944b40b07da9ce895c0e6303a5ab7d53da063554f534556b134a54d6093/packaging-26.3.tar.gz", hash = "sha256:94edc256424af38762eb31306eed28beb9f0efc50a8837492c9d6fd6004aed79", size = 313412, upload-time = "2026-08-04T18:15:28.737Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/63/34/ba1c580383c9eada3711951fef0795c80b829a078d72188184bcab9dd527/packaging-26.3-py3-none-any.whl", hash = "sha256:d7193f7c8e4e93f444fde0262bf90af30e16fa0ad0ad44cb553c87339b23cd1c", size = 129956, upload-time = "2026-08-04T18:15:27.159Z" },
]
[[package]]
name = "pathspec"
version = "1.1.1"
source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/5a/82/42f767fc1c1143d6fd36efb827202a2d997a375e160a71eb2888a925aac1/pathspec-1.1.1.tar.gz", hash = "sha256:17db5ecd524104a120e173814c90367a96a98d07c45b2e10c2f3919fff91bf5a", size = 135180, upload-time = "2026-04-27T01:46:08.907Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/f1/d9/7fb5aa316bc299258e68c73ba3bddbc499654a07f151cba08f6153988714/pathspec-1.1.1-py3-none-any.whl", hash = "sha256:a00ce642f577bf7f473932318056212bc4f8bfdf53128c78bbd5af0b9b20b189", size = 57328, upload-time = "2026-04-27T01:46:07.06Z" },
]
[[package]]
name = "pluggy"
version = "1.6.0"
source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/f9/e2/3e91f31a7d2b083fe6ef3fa267035b518369d9511ffab804f839851d2779/pluggy-1.6.0.tar.gz", hash = "sha256:7dcc130b76258d33b90f61b658791dede3486c3e6bfb003ee5c9bfb396dd22f3", size = 69412, upload-time = "2025-05-15T12:30:07.975Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/54/20/4d324d65cc6d9205fabedc306948156824eb9f0ee1633355a8f7ec5c66bf/pluggy-1.6.0-py3-none-any.whl", hash = "sha256:e920276dd6813095e9377c0bc5566d94c932c33b27a3e3945d8389c374dd4746", size = 20538, upload-time = "2025-05-15T12:30:06.134Z" },
]
[[package]] [[package]]
name = "pydantic" name = "pydantic"
version = "2.12.5" version = "2.12.5"
@@ -338,6 +391,15 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/f7/07/34573da085946b6a313d7c42f82f16e8920bfd730665de2d11c0c37a74b5/pydantic_core-2.41.5-graalpy312-graalpy250_312_native-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:76d0819de158cd855d1cbb8fcafdf6f5cf1eb8e470abe056d5d161106e38062b", size = 2139017, upload-time = "2025-11-04T13:42:59.471Z" }, { url = "https://files.pythonhosted.org/packages/f7/07/34573da085946b6a313d7c42f82f16e8920bfd730665de2d11c0c37a74b5/pydantic_core-2.41.5-graalpy312-graalpy250_312_native-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:76d0819de158cd855d1cbb8fcafdf6f5cf1eb8e470abe056d5d161106e38062b", size = 2139017, upload-time = "2025-11-04T13:42:59.471Z" },
] ]
[[package]]
name = "pyserial"
version = "3.5"
source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/1e/7d/ae3f0a63f41e4d2f6cb66a5b57197850f919f59e558159a4dd3a818f5082/pyserial-3.5.tar.gz", hash = "sha256:3c77e014170dfffbd816e6ffc205e9842efb10be9f58ec16d3e8675b4925cddb", size = 159125, upload-time = "2020-11-23T03:59:15.045Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/07/bc/587a445451b253b285629263eb51c2d8e9bcea4fc97826266d186f96f558/pyserial-3.5-py2.py3-none-any.whl", hash = "sha256:c4451db6ba391ca6ca299fb3ec7bae67a5c55dde170964c7a14ceefec02f2cf0", size = 90585, upload-time = "2020-11-23T03:59:13.41Z" },
]
[[package]] [[package]]
name = "python-dotenv" name = "python-dotenv"
version = "1.2.2" version = "1.2.2"
@@ -406,6 +468,24 @@ wheels = [
{ url = "https://files.pythonhosted.org/packages/81/0d/13d1d239a25cbfb19e740db83143e95c772a1fe10202dda4b76792b114dd/starlette-0.52.1-py3-none-any.whl", hash = "sha256:0029d43eb3d273bc4f83a08720b4912ea4b071087a3b48db01b7c839f7954d74", size = 74272, upload-time = "2026-01-18T13:34:09.188Z" }, { url = "https://files.pythonhosted.org/packages/81/0d/13d1d239a25cbfb19e740db83143e95c772a1fe10202dda4b76792b114dd/starlette-0.52.1-py3-none-any.whl", hash = "sha256:0029d43eb3d273bc4f83a08720b4912ea4b071087a3b48db01b7c839f7954d74", size = 74272, upload-time = "2026-01-18T13:34:09.188Z" },
] ]
[[package]]
name = "tomlkit"
version = "0.15.1"
source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/94/96/e07752635b98536177fa1f37671c8f3cdde2e724c6bcf6034b2cfb571565/tomlkit-0.15.1.tar.gz", hash = "sha256:e25bbf38843005246210a12982776f27f99cb9be67160e14434d0c0d21ee1e97", size = 180129, upload-time = "2026-07-17T01:48:04.562Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/13/bc/8c13eb66537dce1d2bd3a57132902f38d0e7f5bb46fa9f4daed9fe9d76ee/tomlkit-0.15.1-py3-none-any.whl", hash = "sha256:177a05aece5a8ca5266fd3c448abb47b8d352f09d477d3ca8332db4d89b24304", size = 49449, upload-time = "2026-07-17T01:48:05.728Z" },
]
[[package]]
name = "trove-classifiers"
version = "2026.6.1.19"
source = { registry = "https://pypi.org/simple" }
sdist = { url = "https://files.pythonhosted.org/packages/c2/e3/7ca82ee24c82d344584abd5b8637b3bd056f2900226e8d82fc22f1184b92/trove_classifiers-2026.6.1.19.tar.gz", hash = "sha256:c5132b4b61a829d11cfbd2d72e97f20a45ed6edb95e45c5efdeb5e00836b2745", size = 17059, upload-time = "2026-06-01T19:41:34.649Z" }
wheels = [
{ url = "https://files.pythonhosted.org/packages/7c/a4/81502f486f01db95bc8320646a8a12511f5e556cb63d5e224d91816605c4/trove_classifiers-2026.6.1.19-py3-none-any.whl", hash = "sha256:ab4c4ec93cc4a4e7815fa759906e05e6bb3f2fbd92ea0f897288c6a43efd15b3", size = 14211, upload-time = "2026-06-01T19:41:33.434Z" },
]
[[package]] [[package]]
name = "typing-extensions" name = "typing-extensions"
version = "4.15.0" version = "4.15.0"