Compare commits
9
Commits
d810c44478
..
main
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
db0e676501 | ||
|
|
38c3a30017 | ||
|
|
457ca5c3a6 | ||
|
|
6f5c046a59 | ||
|
|
b8e91e3c91 | ||
|
|
526b71b45f | ||
|
|
cc102f26c1 | ||
|
|
e148c83622 | ||
|
|
a8462e696c |
No files matched your search
@@ -0,0 +1,31 @@
|
||||
# NVCurve — developer convenience targets.
|
||||
#
|
||||
# End users don't need make: run ./install.sh (see README "Installation").
|
||||
|
||||
UV ?= uv
|
||||
NPM ?= npm
|
||||
|
||||
.PHONY: help frontend frontend-dev install dev test clean
|
||||
|
||||
help: ## Show available targets
|
||||
@grep -E '^[a-zA-Z_-]+:.*?## ' $(MAKEFILE_LIST) | awk 'BEGIN {FS = ":.*?## "}; {printf " \033[36m%-15s\033[0m %s\n", $$1, $$2}'
|
||||
|
||||
frontend: ## Build the React frontend into frontend/dist
|
||||
cd frontend && $(NPM) ci && $(NPM) run build
|
||||
|
||||
frontend-dev: ## Run the Vite dev server (hot reload)
|
||||
cd frontend && $(NPM) run dev
|
||||
|
||||
install: ## Install nvcurve as a uv tool (builds frontend if missing/stale)
|
||||
$(UV) tool install --force .
|
||||
|
||||
dev: ## Create/refresh the dev environment (uv sync)
|
||||
$(UV) sync
|
||||
|
||||
test: ## Run the test suite
|
||||
$(UV) run python tests/test_security.py
|
||||
$(UV) run python tests/test_rm_power.py
|
||||
$(UV) run python tests/test_wireview.py
|
||||
|
||||
clean: ## Remove build artifacts
|
||||
rm -rf frontend/dist frontend/node_modules
|
||||
@@ -15,6 +15,7 @@ NVCurve brings MSI Afterburner-style per-point voltage-frequency curve control t
|
||||
> **Blackwell GPU memory** — This is a specialized fork with extended memory offset support (up to +3000 MHz) for Blackwell GPUs (RTX 50-series). \
|
||||
> **Fan Controls** — There is an additional "Fans" tab to setup a customized fan curve, controlling all fans or individual fans. \
|
||||
> **Dashboard** — The default tab is an Dashboard with additional information (PCIe link speed, VBIOS information, Max Core Clock, Throttle Reason and much much more.) \
|
||||
> **WireView Pro II** — Native monitoring of the Thermal Grizzly WireView Pro II 12VHPWR connector: a dedicated tab shows live temperature, power and current readings — no external exporter or kernel module required. The host needs a one-time setup (udev rule + port access, see [WireView Pro II Setup](docs/WireView.md)); after that the device is auto-detected over USB and the tab only appears while it is connected. \
|
||||
> **Authentification** — For production deplyoment, I added authentification with bcrypt hashing to allow only one or multiple people to have access. \
|
||||
> Installing the pre-built PyPI package will NOT include these features. You must build from source.
|
||||
|
||||
@@ -27,6 +28,9 @@ NVCurve brings MSI Afterburner-style per-point voltage-frequency curve control t
|
||||
<td align="center"><img src="docs/performance.png" width="480" alt="Performance"></td>
|
||||
<td align="center"><img src="docs/fans.png" width="480" alt="Fans"></td>
|
||||
</tr>
|
||||
<tr>
|
||||
<td align="center"><img src="docs/wireview.png" width="480" alt="WireView"></td>
|
||||
</tr>
|
||||
</table>
|
||||
|
||||
## Prerequisites
|
||||
@@ -37,20 +41,28 @@ NVCurve brings MSI Afterburner-style per-point voltage-frequency curve control t
|
||||
- **[uv](https://docs.astral.sh/uv/)** — Python package manager
|
||||
- **Root/sudo access** (required for GPU hardware interactions)
|
||||
|
||||
## Installation from Source
|
||||
## Installation
|
||||
|
||||
### One-liner
|
||||
|
||||
```bash
|
||||
git clone <this-repo-url>.git
|
||||
curl -fsSL https://gitea.zephyre.one/Pakobbix/nvcurve/raw/branch/main/install.sh | bash
|
||||
```
|
||||
|
||||
The script checks prerequisites (installs `uv` if missing), clones the repo, and installs NVCurve — the React frontend is compiled automatically during the build.
|
||||
|
||||
### From a clone
|
||||
|
||||
```bash
|
||||
git clone https://gitea.zephyre.one/Pakobbix/nvcurve.git
|
||||
cd nvcurve
|
||||
./install.sh
|
||||
```
|
||||
|
||||
# Build the React frontend
|
||||
cd frontend
|
||||
npm install
|
||||
npm run build
|
||||
cd ..
|
||||
### Direct from git (no clone, no script)
|
||||
|
||||
# Install the Python package (includes bundled frontend)
|
||||
uv tool install .
|
||||
```bash
|
||||
uv tool install "git+https://gitea.zephyre.one/Pakobbix/nvcurve.git"
|
||||
```
|
||||
|
||||
After installation, verify hardware compatibility:
|
||||
@@ -194,16 +206,18 @@ The GPU key can be a UUID, `pci:XXXX`, or `idx:N` fallback. Find your GPU key wi
|
||||
- **[Installation](docs/Installation.md)** — Prerequisites, source build, troubleshooting
|
||||
- **[Usage Guide](docs/Usage-Guide.md)** — Web UI, CLI reference, systemd service
|
||||
- **[Tips and Tricks](docs/Tips-and-Tricks.md)** — Workflows, curve flattening, safety
|
||||
- **[WireView Pro II Setup](docs/WireView.md)** — One-time host setup (udev rule + port access) for the 12VHPWR connector monitor
|
||||
|
||||
## Upgrading
|
||||
|
||||
```bash
|
||||
cd nvcurve
|
||||
git pull
|
||||
cd frontend && npm run build && cd ..
|
||||
uv tool install .
|
||||
uv tool install --force .
|
||||
```
|
||||
|
||||
The frontend is rebuilt automatically if it is missing or older than the frontend sources. If you modified frontend code locally, run `make frontend` first (or `rm -rf frontend/dist`).
|
||||
|
||||
If running as a systemd service:
|
||||
|
||||
```bash
|
||||
|
||||
+29
-19
@@ -29,43 +29,52 @@ sudo pacman -S uv
|
||||
pip install uv
|
||||
```
|
||||
|
||||
## Installation from Source
|
||||
## Installation
|
||||
|
||||
### Step 1: Clone the Repository
|
||||
### Option 1: One-liner (recommended)
|
||||
|
||||
```bash
|
||||
git clone <this-repo-url>.git
|
||||
curl -fsSL https://gitea.zephyre.one/Pakobbix/nvcurve/raw/branch/main/install.sh | bash
|
||||
```
|
||||
|
||||
The script checks prerequisites (installs `uv` if missing), clones the repository, and installs NVCurve. The React frontend is compiled automatically during the build by a hatchling build hook (`hatch_build.py`).
|
||||
|
||||
### Option 2: From a clone
|
||||
|
||||
```bash
|
||||
git clone https://gitea.zephyre.one/Pakobbix/nvcurve.git
|
||||
cd nvcurve
|
||||
./install.sh
|
||||
```
|
||||
|
||||
### Step 2: Build the Frontend
|
||||
|
||||
The frontend is a React + TypeScript + Vite application in the `frontend/` directory.
|
||||
Equivalent manual steps (what the script does):
|
||||
|
||||
```bash
|
||||
cd frontend
|
||||
npm install
|
||||
npm run build
|
||||
cd ..
|
||||
git clone https://gitea.zephyre.one/Pakobbix/nvcurve.git
|
||||
cd nvcurve
|
||||
uv tool install . # frontend is built automatically if missing/stale
|
||||
```
|
||||
|
||||
This produces a `dist/` directory with the compiled static assets. The hatch build system bundles `frontend/dist` into the Python package.
|
||||
|
||||
### Step 3: Install the Python Package
|
||||
### Option 3: Direct from git (no clone, no script)
|
||||
|
||||
```bash
|
||||
uv tool install .
|
||||
uv tool install "git+https://gitea.zephyre.one/Pakobbix/nvcurve.git"
|
||||
```
|
||||
|
||||
This installs `nvcurve` as a system-wide tool with the bundled frontend.
|
||||
To install a specific branch:
|
||||
|
||||
### Step 4: Verify
|
||||
```bash
|
||||
uv tool install "git+https://gitea.zephyre.one/Pakobbix/nvcurve.git@<branch>"
|
||||
```
|
||||
|
||||
### Verify
|
||||
|
||||
```bash
|
||||
nvcurve setup
|
||||
```
|
||||
|
||||
This performs four checks:
|
||||
|
||||
1. **NvAPI function probe** — verifies all required functions resolve in your driver
|
||||
2. **Curve read** — reads and displays your current V/F curve as a baseline
|
||||
3. **Write-verify** — writes `+5 MHz` to a safe point, reads it back, and confirms the change
|
||||
@@ -94,10 +103,11 @@ nvcurve serve start
|
||||
```bash
|
||||
cd nvcurve
|
||||
git pull
|
||||
cd frontend && npm run build && cd ..
|
||||
uv tool install .
|
||||
uv tool install --force .
|
||||
```
|
||||
|
||||
The frontend is rebuilt automatically if it is missing or older than the frontend sources. If you modified frontend code locally, run `make frontend` first (or `rm -rf frontend/dist`).
|
||||
|
||||
If running as a systemd service:
|
||||
|
||||
```bash
|
||||
@@ -116,7 +126,7 @@ source ~/.local/bin/env # or wherever uv installed
|
||||
|
||||
### Frontend not loading in the web UI
|
||||
|
||||
Verify that `frontend/dist` exists and contains built assets. If the directory is empty or missing, rebuild with `npm run build` and reinstall with `uv tool install .`.
|
||||
Verify that the installed package contains the frontend. If `frontend/dist` is empty or missing, rebuild with `make frontend` (or `cd frontend && npm ci && npm run build`) and reinstall with `uv tool install --force .`.
|
||||
|
||||
### NvAPI functions not found
|
||||
|
||||
|
||||
@@ -84,6 +84,23 @@ The monitoring panel shows real-time GPU metrics:
|
||||
|
||||
Data is streamed via WebSocket from the backend at a configurable poll interval (default: 1 second).
|
||||
|
||||
### Performance Limits
|
||||
|
||||
The Performance panel controls the board power limit and the memory clock offset. The power limit slider is bounded by the GPU's VBIOS minimum and maximum (shown at the slider ends); changes are applied on **Apply** and reset to the hardware default on **Reset**.
|
||||
|
||||
#### Experimental NVIDIA power control
|
||||
|
||||
On compatible drivers, an **Experimental NVIDIA power control** checkbox appears in the Performance panel. Enabling it switches power-limit application from the standard NVML call to an undocumented driver (RM) interface, which **permits caps below the VBIOS minimum, down to 30 W**. The native maximum still applies.
|
||||
|
||||
> **Warning.** This uses an undocumented driver interface for *all* power limits, including resets. It may cause instability or stop working after a driver update. Enable it at your own risk. The option is clearly labelled with a red warning in the UI, and profiles saved while it is enabled are marked accordingly.
|
||||
|
||||
Notes:
|
||||
|
||||
- The checkbox only appears when the driver exposes a compatible RM power layout (detected with a read-only probe — no writes).
|
||||
- In this mode there is **no automatic fallback** to NVML: if the RM route fails, the error is reported rather than silently switching backends.
|
||||
- The mode is per-GPU and persisted across server restarts. Reset restores the default through the same route, so it can also clear a previously-set below-minimum cap.
|
||||
- The CLI reports availability via `nvcurve read --diag` ("Experimental RM power: available").
|
||||
|
||||
### Multi-GPU
|
||||
|
||||
When multiple NVIDIA GPUs are detected, a GPU selector dropdown appears in the status bar. Switching GPUs resets pending edits, selection state, and monitoring for the new target.
|
||||
|
||||
@@ -0,0 +1,112 @@
|
||||
# WireView Pro II Setup
|
||||
|
||||
NVCurve reads the [Thermal Grizzly WireView Pro II](https://www.thermal-grizzly.com/en/wireview-pro-ii-gpu/s-tg-wv-p2) (12 VHPWR connector monitor) directly over its USB CDC/ACM serial port — **no exporter, no GUI, no kernel module required**.
|
||||
|
||||
Before the first use, the host needs a one-time setup: a udev rule so the serial port is accessible under a stable name and left alone by other tools. After that, NVCurve auto-detects the device over USB — the WireView tab appears while the device is connected and disappears when it is unplugged.
|
||||
|
||||
## Why a one-time setup is needed
|
||||
|
||||
The WireView is an STM32 CDC/ACM virtual serial port (VID `0483`, PID `5740`). Out of the box, Linux:
|
||||
|
||||
- creates the port as `root:uucp 0660` with no stable device name, and
|
||||
- lets **ModemManager** probe every new CDC-ACM port with AT/QCDM commands for up to ~30 s after each plug — it holds the port open and sends bytes the device does not expect.
|
||||
|
||||
The udev rule below fixes both: the node becomes `root:dialout 0660`, a stable `/dev/wireview-pro2` symlink is created, and ModemManager is told to ignore the device.
|
||||
|
||||
## 1. Install the udev rule
|
||||
|
||||
```sh
|
||||
sudo tee /etc/udev/rules.d/99-wireview.rules > /dev/null <<'EOF'
|
||||
# WireView Pro II udev rules
|
||||
#
|
||||
# Same access policy as wireview-hwmon's 99-wireview-hwmon.rules: the node is
|
||||
# 0660 root:dialout, and the user logged in at the local seat gets an ACL via
|
||||
# uaccess. systemd applies that tag on hotplug from 73-seat-late.rules, which
|
||||
# runs before this 99- file and so never sees it, so each rule also runs the
|
||||
# uaccess builtin itself. The tag is still needed: logind uses it to move the
|
||||
# ACL to the new active session on a user switch.
|
||||
|
||||
ACTION=="remove", GOTO="wireview_end"
|
||||
|
||||
# Normal operation (STM32 CDC/ACM virtual serial port)
|
||||
SUBSYSTEM=="tty", ATTRS{idVendor}=="0483", ATTRS{idProduct}=="5740", GROUP="dialout", MODE="0660", TAG+="uaccess", RUN{builtin}+="uaccess", SYMLINK+="wireview-pro2"
|
||||
|
||||
# Keep ModemManager away: it probes every new CDC-ACM port with AT and QCDM
|
||||
# commands for about half a minute, holding the port and sending the device
|
||||
# bytes it does not expect.
|
||||
SUBSYSTEM=="usb", ENV{DEVTYPE}=="usb_device", ATTR{idVendor}=="0483", ATTR{idProduct}=="5740", ENV{ID_MM_DEVICE_IGNORE}="1"
|
||||
SUBSYSTEM=="tty", ATTRS{idVendor}=="0483", ATTRS{idProduct}=="5740", ENV{ID_MM_PORT_IGNORE}="1"
|
||||
|
||||
# DFU bootloader mode (firmware update)
|
||||
SUBSYSTEM=="usb", ENV{DEVTYPE}=="usb_device", ATTR{idVendor}=="0483", ATTR{idProduct}=="df11", GROUP="dialout", MODE="0660", TAG+="uaccess", RUN{builtin}+="uaccess"
|
||||
|
||||
LABEL="wireview_end"
|
||||
EOF
|
||||
```
|
||||
|
||||
> [!WARNING]
|
||||
> The rule sets `GROUP="dialout"`. If the `dialout` group does not exist on your distro, udev silently ignores the rule (the node stays `root:uucp`). Create it first: `sudo groupadd dialout`.
|
||||
|
||||
Then reload the rules and apply them to an already-plugged device (a reload alone only affects future hotplugs):
|
||||
|
||||
```sh
|
||||
sudo udevadm control --reload-rules
|
||||
sudo udevadm trigger
|
||||
```
|
||||
|
||||
## 2. Port access
|
||||
|
||||
- **NVCurve web server (systemd):** runs as root, so it can open the port as soon as the rule is in place — nothing else to do.
|
||||
- **CLI as a regular user:** add your user to the `dialout` group, then log out/in:
|
||||
|
||||
```sh
|
||||
sudo usermod -aG dialout $USER
|
||||
```
|
||||
|
||||
A running process only picks up new groups after a re-login; `sg dialout -c 'nvcurve ...'` works for a one-off.
|
||||
|
||||
## 3. Verify
|
||||
|
||||
```sh
|
||||
lsusb | grep 0483 # 0483:5740 present
|
||||
ls -l /dev/wireview-pro2 # -> /dev/ttyACM0
|
||||
```
|
||||
|
||||
With the device plugged in, start (or restart) the NVCurve server. The **WireView** tab appears in the web UI and `GET /api/wireview` returns the device info and live samples. The tab disappears when the device is unplugged and reappears on re-plug — no restart needed.
|
||||
|
||||
## Alternative: wireview-hwmon kernel module
|
||||
|
||||
If the official `wireview-hwmon` kernel module and its `wireviewd` daemon are installed, the daemon owns the serial port and NVCurve automatically reads the sysfs node instead. No udev rule is needed in that case.
|
||||
|
||||
## USB device IDs
|
||||
|
||||
| Mode | VID | PID | Description |
|
||||
| --- | --- | --- | --- |
|
||||
| Normal | `0483` | `5740` | STM32 CDC/ACM virtual serial port |
|
||||
| DFU bootloader | `0483` | `df11` | STM32 bootloader (firmware updates only) |
|
||||
|
||||
## Troubleshooting
|
||||
|
||||
### WireView tab does not appear
|
||||
|
||||
- `lsusb | grep 0483` — is the device visible at all (cable, USB port)?
|
||||
- `ls -l /dev/wireview-pro2` — if the node is `root:uucp`, the `dialout` group is missing or the rule did not load: `sudo groupadd dialout && sudo udevadm control --reload-rules && sudo udevadm trigger`.
|
||||
- Restart the NVCurve server — detection runs at startup, and the watchdog re-checks periodically afterwards.
|
||||
|
||||
### `Permission denied` on the port (CLI)
|
||||
|
||||
The user must be in the `dialout` group; a running process only picks up new groups after a restart/re-login.
|
||||
|
||||
### Device visible but never connects
|
||||
|
||||
The firmware occasionally stops answering the RTS welcome handshake (observed after USB state changes, e.g. udev re-triggers or boot) while still answering every data command. NVCurve falls back to the vendor-data reply and retries, but if it still does not connect, reset the USB device — physically unplug/replug, or without reaching for it:
|
||||
|
||||
```sh
|
||||
P=$(readlink -f /sys/class/tty/ttyACM0)
|
||||
DEV=$(dirname "$(dirname "$(dirname "$(dirname "$P")")")")
|
||||
sudo bash -c "echo 0 > $DEV/authorized; sleep 1; echo 1 > $DEV/authorized"
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
*The udev rule is taken from the wireview-reporter project, which uses the same access policy as the official `wireview-hwmon` module.*
|
||||
Binary file not shown.
|
After Width: | Height: | Size: 131 KiB |
+19
-1
@@ -11,10 +11,12 @@ import { PerformancePanel } from "./components/Limits/PerformancePanel.js";
|
||||
import { PerformanceMonitor } from "./components/Monitor/PerformanceMonitor.js";
|
||||
import { FanMonitor } from "./components/Monitor/FanMonitor.js";
|
||||
import { FanCurveEditor } from "./components/Fans/FanCurveEditor.js";
|
||||
import { WireViewPanel } from "./components/WireView/WireViewPanel.js";
|
||||
import { ProfilePanel } from "./components/Profiles/ProfilePanel.js";
|
||||
import { api, onUnauthorized } from "./api/client.js";
|
||||
import { LoginScreen } from "./components/Auth/LoginScreen.js";
|
||||
import { useCurveStore } from "./store/curveStore.js";
|
||||
import { useWireview } from "./hooks/useWireview.js";
|
||||
import { Toaster } from "sonner";
|
||||
import { Loader, ChevronDown } from "lucide-react";
|
||||
import { useState, useRef, useEffect } from "react";
|
||||
@@ -94,9 +96,10 @@ function MainApp({
|
||||
const { dashboard, loading: dashboardLoading } = useDashboard();
|
||||
const { setCurve, activeProfile, setActiveProfile, selectedGpuIndex } =
|
||||
useCurveStore();
|
||||
const wireview = useWireview();
|
||||
|
||||
const [activeTab, setActiveTab] = useState<
|
||||
"dashboard" | "curve" | "performance" | "fans"
|
||||
"dashboard" | "curve" | "performance" | "fans" | "wireview"
|
||||
>("dashboard");
|
||||
const [fanState, setFanState] = useState<FanState | null>(null);
|
||||
const [activeDomain, setActiveDomain] = useState<"gpu" | "memory">("gpu");
|
||||
@@ -195,6 +198,15 @@ function MainApp({
|
||||
>
|
||||
Fans
|
||||
</button>
|
||||
{wireview.available && (
|
||||
<button
|
||||
onClick={() => setActiveTab("wireview")}
|
||||
className={`text-lg font-medium pb-2 -mb-[9px] border-b-2 transition-colors flex items-center gap-2 ${activeTab === "wireview" ? "border-pink-500 text-zinc-100" : "border-transparent text-zinc-500 hover:text-zinc-300"}`}
|
||||
>
|
||||
WireView
|
||||
<span className="w-1.5 h-1.5 rounded-full bg-emerald-400" />
|
||||
</button>
|
||||
)}
|
||||
</div>
|
||||
<div className="relative" ref={profileRef}>
|
||||
<button
|
||||
@@ -289,6 +301,12 @@ function MainApp({
|
||||
/>
|
||||
</div>
|
||||
</div>
|
||||
) : activeTab === "wireview" ? (
|
||||
<WireViewPanel
|
||||
info={wireview.info}
|
||||
sample={wireview.sample}
|
||||
history={wireview.history}
|
||||
/>
|
||||
) : (
|
||||
<div className="flex gap-4 items-start w-full">
|
||||
<div className="flex-1 min-w-0">
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
import { fmt } from '../../utils/units.js';
|
||||
import type { VFPoint } from '../../types.js';
|
||||
import { fmt } from "../../utils/units.js";
|
||||
import type { VFPoint } from "../../types.js";
|
||||
|
||||
interface Props {
|
||||
point: VFPoint;
|
||||
@@ -17,24 +17,35 @@ export function CurveTooltip({ point, pendingDeltaKhz, isClamped }: Props) {
|
||||
const hasPending = pendingDeltaKhz !== undefined;
|
||||
const pendingMhz = hasPending ? pendingDeltaKhz! / 1000 : 0;
|
||||
const deltaChange = hasPending ? pendingDeltaKhz! - point.delta_khz : 0;
|
||||
const pendingEffMhz = hasPending ? point.freq_mhz + deltaChange / 1000 : null;
|
||||
const pendingEffMhz = hasPending
|
||||
? point.freq_mhz + deltaChange / 1000
|
||||
: null;
|
||||
|
||||
return (
|
||||
<div
|
||||
className="absolute right-4 bottom-4 pointer-events-none z-50 w-[172px] bg-zinc-800 border border-zinc-700 rounded-md p-2 text-xs shadow-xl"
|
||||
>
|
||||
<div className="absolute right-4 bottom-4 pointer-events-none z-50 w-[172px] bg-zinc-800 border border-zinc-700 rounded-md p-2 text-xs shadow-xl">
|
||||
<div className="text-zinc-400 mb-1">Point {point.index}</div>
|
||||
<div className="text-zinc-200">
|
||||
<span className="text-zinc-400">Volt: </span>{fmt.mv(point.volt_mv, 1)}
|
||||
<span className="text-zinc-400">Volt: </span>
|
||||
{fmt.mv(point.volt_mv, 1)}
|
||||
</div>
|
||||
<div className="text-zinc-200">
|
||||
<span className="text-zinc-400">Offset: </span>
|
||||
<span className={point.delta_khz > 0 ? 'text-emerald-400' : point.delta_khz < 0 ? 'text-red-400' : 'text-zinc-400'}>
|
||||
{point.delta_khz > 0 ? '+' : ''}{fmt.mhz(point.delta_mhz, 1)}
|
||||
<span
|
||||
className={
|
||||
point.delta_khz > 0
|
||||
? "text-emerald-400"
|
||||
: point.delta_khz < 0
|
||||
? "text-red-400"
|
||||
: "text-zinc-400"
|
||||
}
|
||||
>
|
||||
{point.delta_khz > 0 ? "+" : ""}
|
||||
{fmt.mhz(point.delta_mhz, 1)}
|
||||
</span>
|
||||
</div>
|
||||
<div className="text-emerald-300 font-semibold">
|
||||
<span className="text-zinc-400">Eff.: </span>{fmt.mhz(point.freq_mhz, 0)}
|
||||
<span className="text-zinc-400">Eff.: </span>
|
||||
{fmt.mhz(point.freq_mhz, 0)}
|
||||
{isClamped && <span className="text-amber-500 ml-1">⇡</span>}
|
||||
</div>
|
||||
{isClamped && (
|
||||
@@ -47,12 +58,22 @@ export function CurveTooltip({ point, pendingDeltaKhz, isClamped }: Props) {
|
||||
<div className="border-t border-zinc-700 mt-1.5 pt-1.5">
|
||||
<div className="text-zinc-200">
|
||||
<span className="text-zinc-400">Pending: </span>
|
||||
<span className={pendingMhz > 0 ? 'text-cyan-400' : pendingMhz < 0 ? 'text-orange-400' : 'text-zinc-400'}>
|
||||
{pendingMhz > 0 ? '+' : ''}{pendingMhz.toFixed(1)} MHz
|
||||
<span
|
||||
className={
|
||||
pendingMhz > 0
|
||||
? "text-cyan-400"
|
||||
: pendingMhz < 0
|
||||
? "text-orange-400"
|
||||
: "text-zinc-400"
|
||||
}
|
||||
>
|
||||
{pendingMhz > 0 ? "+" : ""}
|
||||
{pendingMhz.toFixed(1)} MHz
|
||||
</span>
|
||||
</div>
|
||||
<div className="text-cyan-300 font-semibold">
|
||||
<span className="text-zinc-400">→ Eff.: </span>{fmt.mhz(pendingEffMhz, 0)}
|
||||
<span className="text-zinc-400">→ Eff.: </span>
|
||||
{fmt.mhz(pendingEffMhz, 0)}
|
||||
</div>
|
||||
</div>
|
||||
</>
|
||||
|
||||
@@ -72,6 +72,21 @@ export function PerformancePanel() {
|
||||
}
|
||||
}
|
||||
|
||||
async function handleModeChange(enabled: boolean) {
|
||||
setBusy(true);
|
||||
try {
|
||||
await api.updateLimits(
|
||||
{ power_cap_mode: enabled ? "ioctl" : "nvml" },
|
||||
selectedGpuIndex,
|
||||
);
|
||||
await fetchLimits();
|
||||
} catch (e: unknown) {
|
||||
toast.error(e instanceof Error ? e.message : String(e));
|
||||
} finally {
|
||||
setBusy(false);
|
||||
}
|
||||
}
|
||||
|
||||
if (loading && !limits) {
|
||||
return (
|
||||
<div className="bg-zinc-900 rounded-lg overflow-hidden flex flex-col animate-pulse">
|
||||
@@ -158,12 +173,80 @@ export function PerformancePanel() {
|
||||
)}
|
||||
|
||||
<div className="flex flex-col divide-y divide-zinc-800">
|
||||
{/* ── Experimental NVIDIA power control ─────────────────────────── */}
|
||||
{(limits.rm_power_supported || limits.power_cap_mode === "ioctl") && (
|
||||
<div className="px-4 py-3 flex flex-col gap-2">
|
||||
<label className="flex items-center gap-2 cursor-pointer select-none">
|
||||
<input
|
||||
type="checkbox"
|
||||
checked={limits.power_cap_mode === "ioctl"}
|
||||
disabled={busy}
|
||||
onChange={(e) => handleModeChange(e.target.checked)}
|
||||
className="accent-red-500"
|
||||
/>
|
||||
<span className="text-xs text-zinc-300">
|
||||
Experimental NVIDIA power control
|
||||
</span>
|
||||
</label>
|
||||
{limits.rm_power_supported ? (
|
||||
<div
|
||||
role="alert"
|
||||
className={
|
||||
"px-2.5 py-1.5 rounded border text-xs leading-relaxed " +
|
||||
(limits.power_cap_mode === "ioctl"
|
||||
? "bg-red-950/80 border-red-500 text-red-300"
|
||||
: "bg-red-950/40 border-red-800 text-red-400")
|
||||
}
|
||||
>
|
||||
<span className="font-bold">⚠ WARNING:</span> uses an
|
||||
undocumented driver interface for ALL power limits, including
|
||||
resets. Allows values below the VBIOS minimum (down to 30 W)
|
||||
and may cause instability or stop working after driver
|
||||
updates. Enable at your own risk. Based on the work of{" "}
|
||||
<a
|
||||
href="https://github.com/ilya-zlobintsev/LACT/pull/1205"
|
||||
target="_blank"
|
||||
rel="noopener noreferrer"
|
||||
className="underline hover:text-red-200"
|
||||
>
|
||||
panchovix
|
||||
</a>{" "}
|
||||
(LACT PR #1205).
|
||||
</div>
|
||||
) : (
|
||||
<div
|
||||
role="alert"
|
||||
className="px-2.5 py-1.5 rounded border border-red-500 bg-red-950/80 text-red-300 text-xs leading-relaxed"
|
||||
>
|
||||
<span className="font-bold">
|
||||
⚠ Interface not currently available.
|
||||
</span>
|
||||
The driver no longer exposes the RM power interface (it may
|
||||
have been updated). Experimental mode is still enabled, so
|
||||
power-limit changes will fail. Uncheck to switch back to the
|
||||
standard NVML mode.
|
||||
</div>
|
||||
)}
|
||||
</div>
|
||||
)}
|
||||
|
||||
{/* ── Board Power Limit ─────────────────────────────────────────── */}
|
||||
<div className="px-4 py-4 flex flex-col gap-3">
|
||||
<div className="flex items-center justify-between">
|
||||
<div className="flex items-center gap-2">
|
||||
<span className="text-xs text-zinc-500 uppercase tracking-wider">
|
||||
Board Power Limit
|
||||
</span>
|
||||
{limits.power_cap_mode === "ioctl" &&
|
||||
limits.min_power_limit_w_native != null && (
|
||||
<span
|
||||
className="text-xs text-red-400/80 font-mono"
|
||||
title="Native VBIOS minimum — experimental mode allows lower"
|
||||
>
|
||||
VBIOS min {limits.min_power_limit_w_native} W
|
||||
</span>
|
||||
)}
|
||||
</div>
|
||||
<div className="flex items-center gap-1.5">
|
||||
<input
|
||||
type="number"
|
||||
|
||||
@@ -345,6 +345,11 @@ export function ProfilePanel({
|
||||
{badges && (
|
||||
<p className="text-xs text-zinc-500">{badges}</p>
|
||||
)}
|
||||
{p.power_cap_mode === "ioctl" && (
|
||||
<p className="text-xs text-red-400 font-medium">
|
||||
⚠ experimental power (below VBIOS min)
|
||||
</p>
|
||||
)}
|
||||
</div>
|
||||
</div>
|
||||
|
||||
|
||||
@@ -0,0 +1,309 @@
|
||||
import { GaugeCard } from "../Monitor/GaugeCard.js";
|
||||
import { fmt } from "../../utils/units.js";
|
||||
import type { WireViewInfo, WireViewSample } from "../../types.js";
|
||||
|
||||
interface Props {
|
||||
info: WireViewInfo | null;
|
||||
sample: WireViewSample | null;
|
||||
history: WireViewSample[];
|
||||
}
|
||||
|
||||
// Per-pin over-current protection of the 12 VHPWR connector (A).
|
||||
const PIN_OCP_A = 55;
|
||||
|
||||
const FAULT_NAMES: Record<number, string> = {
|
||||
0: "Chip over-temperature",
|
||||
1: "Sensor over-temperature",
|
||||
2: "Over-current (OCP)",
|
||||
3: "Wire over-current",
|
||||
4: "Over-power (OPP)",
|
||||
5: "Current imbalance",
|
||||
};
|
||||
|
||||
function decodeFaults(mask: number): string[] {
|
||||
return Object.entries(FAULT_NAMES)
|
||||
.filter(([bit]) => mask & (1 << Number(bit)))
|
||||
.map(([, name]) => name);
|
||||
}
|
||||
|
||||
function pluck(history: WireViewSample[], key: keyof WireViewSample): number[] {
|
||||
return history.map((s) => (s[key] as number | null) ?? 0);
|
||||
}
|
||||
|
||||
function FaultBadge({
|
||||
label,
|
||||
mask,
|
||||
}: {
|
||||
label: string;
|
||||
mask: number;
|
||||
}) {
|
||||
const faults = decodeFaults(mask);
|
||||
return (
|
||||
<div
|
||||
className={`rounded-lg p-3 border flex flex-col gap-1.5 min-w-0 ${
|
||||
faults.length > 0
|
||||
? "bg-red-500/10 border-red-500/40"
|
||||
: "bg-zinc-900 border-zinc-800"
|
||||
}`}
|
||||
>
|
||||
<div className="flex items-center justify-between gap-2">
|
||||
<span className="text-xs text-zinc-500 uppercase tracking-wider">
|
||||
{label}
|
||||
</span>
|
||||
<span
|
||||
className={`text-xs font-mono ${
|
||||
faults.length > 0 ? "text-red-400" : "text-zinc-600"
|
||||
}`}
|
||||
>
|
||||
0x{mask.toString(16).toUpperCase().padStart(4, "0")}
|
||||
</span>
|
||||
</div>
|
||||
{faults.length > 0 ? (
|
||||
<div className="flex flex-wrap gap-1">
|
||||
{faults.map((f) => (
|
||||
<span
|
||||
key={f}
|
||||
className="px-1.5 py-0.5 rounded bg-red-500/20 border border-red-500/40 text-red-300 text-[11px] font-medium"
|
||||
>
|
||||
{f}
|
||||
</span>
|
||||
))}
|
||||
</div>
|
||||
) : (
|
||||
<div className="text-sm text-emerald-400 font-medium">No faults</div>
|
||||
)}
|
||||
</div>
|
||||
);
|
||||
}
|
||||
|
||||
function PinCard({
|
||||
index,
|
||||
voltage,
|
||||
current,
|
||||
power,
|
||||
}: {
|
||||
index: number;
|
||||
voltage: number;
|
||||
current: number;
|
||||
power: number;
|
||||
}) {
|
||||
const loadPct = Math.min(100, (current / PIN_OCP_A) * 100);
|
||||
const barColor =
|
||||
loadPct >= 90 ? "#f87171" : loadPct >= 75 ? "#fbbf24" : "#34d399";
|
||||
return (
|
||||
<div className="bg-zinc-900 rounded-lg p-3 flex flex-col gap-1.5 min-w-0">
|
||||
<div className="flex items-center justify-between">
|
||||
<span className="text-xs text-zinc-500 uppercase tracking-wider">
|
||||
Pin {index + 1}
|
||||
</span>
|
||||
<span className="text-[10px] font-mono text-zinc-600">
|
||||
{current.toFixed(1)} / {PIN_OCP_A} A
|
||||
</span>
|
||||
</div>
|
||||
<div className="grid grid-cols-3 gap-1 text-center">
|
||||
<div>
|
||||
<div className="text-[10px] text-zinc-600">V</div>
|
||||
<div className="text-sm font-mono font-semibold text-violet-300">
|
||||
{voltage.toFixed(2)}
|
||||
</div>
|
||||
</div>
|
||||
<div>
|
||||
<div className="text-[10px] text-zinc-600">A</div>
|
||||
<div className="text-sm font-mono font-semibold text-cyan-300">
|
||||
{current.toFixed(2)}
|
||||
</div>
|
||||
</div>
|
||||
<div>
|
||||
<div className="text-[10px] text-zinc-600">W</div>
|
||||
<div className="text-sm font-mono font-semibold text-pink-300">
|
||||
{power.toFixed(1)}
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
<div className="h-1.5 rounded-full bg-zinc-800 overflow-hidden">
|
||||
<div
|
||||
className="h-full rounded-full transition-all duration-500"
|
||||
style={{ width: `${loadPct}%`, backgroundColor: barColor }}
|
||||
/>
|
||||
</div>
|
||||
</div>
|
||||
);
|
||||
}
|
||||
|
||||
function InfoItem({
|
||||
label,
|
||||
value,
|
||||
}: {
|
||||
label: string;
|
||||
value: string | null | undefined;
|
||||
}) {
|
||||
if (!value) return null;
|
||||
return (
|
||||
<div className="flex flex-col gap-0.5 min-w-0">
|
||||
<span className="text-xs text-zinc-500">{label}</span>
|
||||
<span
|
||||
className="text-sm font-mono text-zinc-200 truncate"
|
||||
title={value}
|
||||
>
|
||||
{value}
|
||||
</span>
|
||||
</div>
|
||||
);
|
||||
}
|
||||
|
||||
export function WireViewPanel({ info, sample, history }: Props) {
|
||||
if (!sample) {
|
||||
return (
|
||||
<div className="bg-zinc-900 rounded-lg p-10 text-center text-zinc-500 flex flex-col items-center gap-2">
|
||||
<div className="text-lg font-medium text-zinc-400">
|
||||
WireView Pro II not connected
|
||||
</div>
|
||||
<div className="text-sm">
|
||||
Plug in the device — the tab appears automatically once it is
|
||||
detected.
|
||||
</div>
|
||||
</div>
|
||||
);
|
||||
}
|
||||
|
||||
const psuCap = sample.psu_capability_w > 0 ? sample.psu_capability_w : 600;
|
||||
|
||||
return (
|
||||
<div className="flex flex-col gap-4 w-full">
|
||||
{/* ── Header ─────────────────────────────────────────────────── */}
|
||||
<div className="bg-zinc-900 rounded-lg p-4 border border-zinc-800 flex flex-wrap items-center gap-x-6 gap-y-2">
|
||||
<div className="flex items-center gap-2">
|
||||
<span className="w-2 h-2 rounded-full bg-emerald-400" />
|
||||
<span className="text-lg font-semibold text-zinc-100">
|
||||
{info?.device_name ?? "WireView Pro II"}
|
||||
</span>
|
||||
</div>
|
||||
<span className="text-sm font-mono text-zinc-400">
|
||||
{info?.hw_rev ? `HW ${info.hw_rev}` : ""}
|
||||
{info?.firmware_version ? ` · FW ${info.firmware_version}` : ""}
|
||||
</span>
|
||||
<span className="text-sm font-mono text-zinc-400">
|
||||
PSU {psuCap} W
|
||||
</span>
|
||||
<span className="text-xs font-mono text-zinc-600 ml-auto">
|
||||
{info?.transport === "hwmon"
|
||||
? `hwmon: ${info.port}`
|
||||
: `serial: ${info?.port ?? ""}`}
|
||||
</span>
|
||||
</div>
|
||||
|
||||
{/* ── Main gauges ────────────────────────────────────────────── */}
|
||||
<div className="bg-zinc-900 rounded-lg p-4">
|
||||
<div className="text-xs text-zinc-500 uppercase tracking-wider font-semibold pb-3 border-b border-zinc-800">
|
||||
12 VHPWR Connector
|
||||
</div>
|
||||
<div className="grid grid-cols-2 md:grid-cols-4 gap-3 mt-3">
|
||||
<GaugeCard
|
||||
label="Total Power"
|
||||
value={fmt.watts(sample.power_total_w, 1)}
|
||||
history={pluck(history, "power_total_w")}
|
||||
color="#f472b6"
|
||||
max={psuCap}
|
||||
/>
|
||||
<GaugeCard
|
||||
label="Total Current"
|
||||
value={`${sample.current_total_a.toFixed(2)} A`}
|
||||
history={pluck(history, "current_total_a")}
|
||||
color="#38bdf8"
|
||||
max={(psuCap / 12) * 1.1}
|
||||
/>
|
||||
<GaugeCard
|
||||
label="Avg Voltage"
|
||||
value={`${sample.voltage_avg_v.toFixed(3)} V`}
|
||||
history={pluck(history, "voltage_avg_v")}
|
||||
color="#a78bfa"
|
||||
max={13}
|
||||
/>
|
||||
<GaugeCard
|
||||
label="Fan Duty"
|
||||
value={fmt.pct(sample.fan_duty_pct)}
|
||||
history={pluck(history, "fan_duty_pct")}
|
||||
color="#fb923c"
|
||||
max={100}
|
||||
/>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
{/* ── Per-pin ────────────────────────────────────────────────── */}
|
||||
<div className="bg-zinc-900 rounded-lg p-4">
|
||||
<div className="text-xs text-zinc-500 uppercase tracking-wider font-semibold pb-3 border-b border-zinc-800">
|
||||
Per-Pin Readings
|
||||
</div>
|
||||
<div className="grid grid-cols-2 md:grid-cols-3 lg:grid-cols-6 gap-3 mt-3">
|
||||
{sample.pins.map((pin, i) => (
|
||||
<PinCard
|
||||
key={i}
|
||||
index={i}
|
||||
voltage={pin.voltage_v}
|
||||
current={pin.current_a}
|
||||
power={pin.power_w}
|
||||
/>
|
||||
))}
|
||||
</div>
|
||||
</div>
|
||||
|
||||
{/* ── Temperatures ───────────────────────────────────────────── */}
|
||||
<div className="bg-zinc-900 rounded-lg p-4">
|
||||
<div className="text-xs text-zinc-500 uppercase tracking-wider font-semibold pb-3 border-b border-zinc-800">
|
||||
Temperatures
|
||||
</div>
|
||||
<div className="grid grid-cols-2 md:grid-cols-4 gap-3 mt-3">
|
||||
<GaugeCard
|
||||
label="Temp In"
|
||||
value={fmt.celsius(sample.temp_in_c, 1)}
|
||||
history={pluck(history, "temp_in_c")}
|
||||
color="#f87171"
|
||||
max={90}
|
||||
/>
|
||||
<GaugeCard
|
||||
label="Temp Out"
|
||||
value={fmt.celsius(sample.temp_out_c, 1)}
|
||||
history={pluck(history, "temp_out_c")}
|
||||
color="#fb923c"
|
||||
max={90}
|
||||
/>
|
||||
<GaugeCard
|
||||
label="Ext Sensor 1"
|
||||
value={fmt.celsius(sample.temp_ext1_c, 1)}
|
||||
history={pluck(history, "temp_ext1_c")}
|
||||
color="#fbbf24"
|
||||
max={90}
|
||||
/>
|
||||
<GaugeCard
|
||||
label="Ext Sensor 2"
|
||||
value={fmt.celsius(sample.temp_ext2_c, 1)}
|
||||
history={pluck(history, "temp_ext2_c")}
|
||||
color="#a78bfa"
|
||||
max={90}
|
||||
/>
|
||||
</div>
|
||||
</div>
|
||||
|
||||
{/* ── Faults + device info ───────────────────────────────────── */}
|
||||
<div className="grid grid-cols-1 md:grid-cols-2 gap-4">
|
||||
<div className="flex flex-col gap-3">
|
||||
<FaultBadge label="Fault Status" mask={sample.fault_status} />
|
||||
<FaultBadge label="Fault Log (latched)" mask={sample.fault_log} />
|
||||
</div>
|
||||
<div className="bg-zinc-900 rounded-lg p-4">
|
||||
<div className="text-xs text-zinc-500 uppercase tracking-wider font-semibold pb-3 border-b border-zinc-800">
|
||||
Device
|
||||
</div>
|
||||
<div className="grid grid-cols-2 gap-x-6 gap-y-3 mt-3">
|
||||
<InfoItem label="Hardware Revision" value={info?.hw_rev} />
|
||||
<InfoItem label="Firmware" value={info?.firmware_version} />
|
||||
<InfoItem label="Build" value={info?.build} />
|
||||
<InfoItem label="UID" value={info?.uid} />
|
||||
<InfoItem label="Transport" value={info?.transport} />
|
||||
<InfoItem label="Port" value={info?.port} />
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
</div>
|
||||
);
|
||||
}
|
||||
@@ -0,0 +1,40 @@
|
||||
import { useEffect, useRef, useState } from "react";
|
||||
import { createWsConnection } from "../api/websocket.js";
|
||||
import type {
|
||||
WireViewInfo,
|
||||
WireViewSample,
|
||||
WireViewWsMessage,
|
||||
} from "../types.js";
|
||||
|
||||
const MAX_HISTORY = 120; // ~2 min at 1 s polling
|
||||
|
||||
export function useWireview() {
|
||||
const [available, setAvailable] = useState(false);
|
||||
const [info, setInfo] = useState<WireViewInfo | null>(null);
|
||||
const [sample, setSample] = useState<WireViewSample | null>(null);
|
||||
const [history, setHistory] = useState<WireViewSample[]>([]);
|
||||
const wsRef = useRef<ReturnType<typeof createWsConnection> | null>(null);
|
||||
|
||||
useEffect(() => {
|
||||
wsRef.current = createWsConnection<WireViewWsMessage>(
|
||||
"/ws/wireview",
|
||||
(msg) => {
|
||||
if (msg.type === "sample") {
|
||||
setAvailable(true);
|
||||
setInfo(msg.info);
|
||||
setSample(msg.sample);
|
||||
setHistory((h) => [...h.slice(-(MAX_HISTORY - 1)), msg.sample]);
|
||||
} else {
|
||||
setAvailable(false);
|
||||
setInfo(null);
|
||||
setSample(null);
|
||||
setHistory([]);
|
||||
}
|
||||
},
|
||||
() => {},
|
||||
);
|
||||
return () => wsRef.current?.close();
|
||||
}, []);
|
||||
|
||||
return { available, info, sample, history };
|
||||
}
|
||||
@@ -98,7 +98,11 @@ export interface LimitsState {
|
||||
power_limit_w: number | null;
|
||||
default_power_limit_w: number | null;
|
||||
min_power_limit_w: number | null;
|
||||
min_power_limit_w_native: number | null;
|
||||
max_power_limit_w: number | null;
|
||||
// "nvml" (default) or "ioctl" (experimental RM power control)
|
||||
power_cap_mode: "nvml" | "ioctl";
|
||||
rm_power_supported: boolean;
|
||||
// Clock offsets — current values
|
||||
gpc_offset_mhz: number | null;
|
||||
mem_offset_mhz: number | null;
|
||||
@@ -136,6 +140,55 @@ export interface ProfileData {
|
||||
curve_deltas: Record<string, number>;
|
||||
mem_offset_mhz: number | null;
|
||||
power_limit_w: number | null;
|
||||
// "nvml" (default) or "ioctl" (experimental RM power control)
|
||||
power_cap_mode: "nvml" | "ioctl" | null;
|
||||
fan_curve: FanPoint[] | null;
|
||||
fan_targets: number[] | null;
|
||||
}
|
||||
|
||||
// ── WireView Pro II (Thermal Grizzly 12 VHPWR connector monitor) ──
|
||||
|
||||
export interface WireViewPin {
|
||||
voltage_v: number;
|
||||
current_a: number;
|
||||
power_w: number;
|
||||
}
|
||||
|
||||
export interface WireViewSample {
|
||||
timestamp: number;
|
||||
power_total_w: number;
|
||||
current_total_a: number;
|
||||
voltage_avg_v: number;
|
||||
pins: WireViewPin[];
|
||||
// Temperatures are null when the source has no channel for them
|
||||
// (e.g. an hwmon module without external-sensor channels).
|
||||
temp_in_c: number | null;
|
||||
temp_out_c: number | null;
|
||||
temp_ext1_c: number | null;
|
||||
temp_ext2_c: number | null;
|
||||
fan_duty_pct: number;
|
||||
fault_status: number;
|
||||
fault_log: number;
|
||||
psu_capability_w: number;
|
||||
}
|
||||
|
||||
export interface WireViewInfo {
|
||||
device_name: string;
|
||||
hw_rev: string;
|
||||
firmware_version: string;
|
||||
uid: string;
|
||||
build: string;
|
||||
transport: "serial" | "hwmon";
|
||||
port: string;
|
||||
}
|
||||
|
||||
export interface WireViewState {
|
||||
available: boolean;
|
||||
connected: boolean;
|
||||
info: WireViewInfo | null;
|
||||
sample: WireViewSample | null;
|
||||
}
|
||||
|
||||
export type WireViewWsMessage =
|
||||
| { type: "unavailable" }
|
||||
| { type: "sample"; info: WireViewInfo; sample: WireViewSample };
|
||||
@@ -1,4 +1,4 @@
|
||||
import type { VFPoint } from '../types.js';
|
||||
import type { VFPoint } from "../types.js";
|
||||
|
||||
/**
|
||||
* Approximate reference frequency (MHz) for a point: effective − delta.
|
||||
@@ -24,7 +24,9 @@ export function findCurrentPoint(
|
||||
): VFPoint | null {
|
||||
if (voltage_mv == null || points.length === 0) return null;
|
||||
return points.reduce((best, p) =>
|
||||
Math.abs(p.volt_mv - voltage_mv) < Math.abs(best.volt_mv - voltage_mv) ? p : best,
|
||||
Math.abs(p.volt_mv - voltage_mv) < Math.abs(best.volt_mv - voltage_mv)
|
||||
? p
|
||||
: best,
|
||||
);
|
||||
}
|
||||
|
||||
|
||||
@@ -0,0 +1,74 @@
|
||||
"""Custom hatchling build hook: build the React frontend if it is missing or stale.
|
||||
|
||||
This makes NVCurve installable with a single command, e.g.::
|
||||
|
||||
uv tool install "git+https://gitea.zephyre.one/Pakobbix/nvcurve.git"
|
||||
|
||||
The hook runs inside the isolated build environment right before the wheel
|
||||
(or sdist) is assembled. If ``frontend/dist`` does not exist yet — or is older
|
||||
than the frontend sources — it compiles the frontend using the host's ``npm``
|
||||
(PATH is inherited from the environment).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import os
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
|
||||
from hatchling.builders.hooks.plugin.interface import BuildHookInterface
|
||||
|
||||
# Frontend inputs that must be newer than dist/index.html to trigger a rebuild.
|
||||
_FRONTEND_INPUTS = (
|
||||
"src",
|
||||
"index.html",
|
||||
"vite.config.ts",
|
||||
"package.json",
|
||||
"tsconfig.json",
|
||||
)
|
||||
|
||||
|
||||
class FrontendBuildHook(BuildHookInterface):
|
||||
"""Build ``frontend/dist`` with npm when it is missing or stale."""
|
||||
|
||||
PLUGIN_NAME = "custom"
|
||||
|
||||
def initialize(self, version: str, build_data: dict) -> None:
|
||||
frontend = os.path.join(self.root, "frontend")
|
||||
dist_index = os.path.join(frontend, "dist", "index.html")
|
||||
|
||||
if not self._needs_build(frontend, dist_index):
|
||||
return
|
||||
|
||||
npm = shutil.which("npm")
|
||||
if npm is None:
|
||||
raise RuntimeError(
|
||||
"npm not found on PATH. Node.js 18+ and npm are required to build "
|
||||
"the NVCurve frontend. Install them and retry, or use install.sh "
|
||||
"which checks prerequisites for you."
|
||||
)
|
||||
|
||||
print(
|
||||
"[nvcurve] frontend/dist missing or stale — building frontend with npm ...",
|
||||
file=sys.stderr,
|
||||
)
|
||||
subprocess.run([npm, "ci", "--no-audit", "--no-fund"], cwd=frontend, check=True)
|
||||
subprocess.run([npm, "run", "build"], cwd=frontend, check=True)
|
||||
|
||||
@staticmethod
|
||||
def _needs_build(frontend: str, dist_index: str) -> bool:
|
||||
if not os.path.isfile(dist_index):
|
||||
return True
|
||||
dist_mtime = os.path.getmtime(dist_index)
|
||||
for name in _FRONTEND_INPUTS:
|
||||
path = os.path.join(frontend, name)
|
||||
if os.path.isfile(path):
|
||||
if os.path.getmtime(path) > dist_mtime:
|
||||
return True
|
||||
elif os.path.isdir(path):
|
||||
for root, _dirs, files in os.walk(path):
|
||||
for file in files:
|
||||
if os.path.getmtime(os.path.join(root, file)) > dist_mtime:
|
||||
return True
|
||||
return False
|
||||
Executable
+62
@@ -0,0 +1,62 @@
|
||||
#!/usr/bin/env bash
|
||||
#
|
||||
# NVCurve single-command installer.
|
||||
#
|
||||
# curl -fsSL https://gitea.zephyre.one/Pakobbix/nvcurve/raw/branch/main/install.sh | bash
|
||||
#
|
||||
# or from a local clone:
|
||||
#
|
||||
# git clone https://gitea.zephyre.one/Pakobbix/nvcurve.git && cd nvcurve && ./install.sh
|
||||
#
|
||||
# The React frontend is compiled automatically during the Python build
|
||||
# (see hatch_build.py), so Node.js 18+ and npm must be available.
|
||||
#
|
||||
# Environment:
|
||||
# NVCURVE_BRANCH branch to install (default: main)
|
||||
|
||||
set -euo pipefail
|
||||
|
||||
REPO_URL="https://gitea.zephyre.one/Pakobbix/nvcurve.git"
|
||||
BRANCH="${NVCURVE_BRANCH:-main}"
|
||||
|
||||
fail() {
|
||||
echo "error: $*" >&2
|
||||
exit 1
|
||||
}
|
||||
|
||||
# --- prerequisites -----------------------------------------------------------
|
||||
command -v git >/dev/null 2>&1 ||
|
||||
fail "git is required. Install it first."
|
||||
command -v node >/dev/null 2>&1 ||
|
||||
fail "Node.js 18+ is required (e.g. 'sudo pacman -S nodejs npm' or 'sudo apt install nodejs npm')."
|
||||
command -v npm >/dev/null 2>&1 ||
|
||||
fail "npm is required (usually installed together with Node.js)."
|
||||
|
||||
if ! command -v uv >/dev/null 2>&1; then
|
||||
echo "uv not found — installing it from https://astral.sh/uv ..."
|
||||
curl -LsSf https://astral.sh/uv/install.sh | sh
|
||||
export PATH="$HOME/.local/bin:$PATH"
|
||||
command -v uv >/dev/null 2>&1 ||
|
||||
fail "uv installation failed. Install uv manually: https://docs.astral.sh/uv/"
|
||||
fi
|
||||
|
||||
# --- locate the source tree ---------------------------------------------------
|
||||
if [ -f pyproject.toml ] && [ -d frontend ]; then
|
||||
src="$(pwd)"
|
||||
echo "Installing from current directory: $src"
|
||||
else
|
||||
tmp="$(mktemp -d)"
|
||||
trap 'rm -rf "$tmp"' EXIT
|
||||
echo "Cloning $REPO_URL (branch: $BRANCH) ..."
|
||||
git clone --quiet --depth 1 --branch "$BRANCH" "$REPO_URL" "$tmp/nvcurve"
|
||||
src="$tmp/nvcurve"
|
||||
fi
|
||||
|
||||
# --- install -------------------------------------------------------------------
|
||||
# The frontend is built automatically by the build hook (hatch_build.py).
|
||||
uv tool install --force "$src"
|
||||
|
||||
echo
|
||||
echo "NVCurve installed."
|
||||
echo " Verify your GPU: nvcurve setup"
|
||||
echo " Start the web UI: nvcurve"
|
||||
+36
-1
@@ -405,6 +405,11 @@ def run_diagnostics(gpu, gpu_name, gpu_index: int = 0):
|
||||
print(f" Default: {fmt_w(def_w)}")
|
||||
if min_w is not None and max_w is not None:
|
||||
print(f" Range: {min_w} – {max_w} W")
|
||||
if pwr.get("rm_power_supported"):
|
||||
print(
|
||||
" Experimental RM power: available (opt-in via web UI or profile;"
|
||||
" extends range to 30 W)"
|
||||
)
|
||||
|
||||
|
||||
# ── Privilege / browser helpers ───────────────────────────────────────────────
|
||||
@@ -1207,12 +1212,33 @@ def cmd_profile(args):
|
||||
power_limit_w = None
|
||||
mem_offset_mhz = None
|
||||
|
||||
# Capture the GPU's power-cap mode: prefer the running server (most
|
||||
# current), else fall back to the persisted per-GPU mode from config
|
||||
# (so a profile saved while the server is down or auth is enabled
|
||||
# still records the GPU's actual mode rather than assuming nvml).
|
||||
power_cap_mode = "nvml"
|
||||
try:
|
||||
from .client import NvCurveClient
|
||||
|
||||
base = getattr(args, "server", None) or _discover_server_url(default_config)
|
||||
limits = NvCurveClient(base=base, gpu_index=gpu_index).limits()
|
||||
if limits.get("power_cap_mode") in ("nvml", "ioctl"):
|
||||
power_cap_mode = limits["power_cap_mode"]
|
||||
except Exception as exc:
|
||||
log.debug("Could not read power-cap mode from server: %s", exc)
|
||||
gpu_key = _gpu_stable_key_offline(gpu_index)
|
||||
if gpu_key is not None:
|
||||
persisted = default_config.power_cap_modes.get(gpu_key)
|
||||
if persisted in ("nvml", "ioctl"):
|
||||
power_cap_mode = persisted
|
||||
|
||||
data = ProfileData(
|
||||
name=args.name,
|
||||
gpu_name=gpu_name,
|
||||
curve_deltas=curve_deltas,
|
||||
mem_offset_mhz=mem_offset_mhz,
|
||||
power_limit_w=power_limit_w,
|
||||
power_cap_mode=power_cap_mode,
|
||||
)
|
||||
filepath = save_profile(default_config.profile_dir, data)
|
||||
print(f"Saved profile '{args.name}' to {filepath}")
|
||||
@@ -1249,7 +1275,8 @@ def cmd_profile(args):
|
||||
errs.append(f"Mem offset: {msg}")
|
||||
|
||||
if profile.power_limit_w is not None:
|
||||
ok, msg = set_power_limit(profile.power_limit_w, gpu_index)
|
||||
mode = profile.power_cap_mode or "nvml"
|
||||
ok, msg = set_power_limit(profile.power_limit_w, gpu_index, mode)
|
||||
if not ok:
|
||||
errs.append(f"Power limit: {msg}")
|
||||
|
||||
@@ -2296,6 +2323,14 @@ def main():
|
||||
if "fan_curves" in data:
|
||||
# Per-GPU active fan curves, restored on server startup.
|
||||
cfg.fan_curves = dict(data["fan_curves"])
|
||||
if "power_cap_modes" in data:
|
||||
# Per-GPU experimental power-cap mode. "nvml" is the default
|
||||
# (the server treats it as unset); keep only valid values.
|
||||
cfg.power_cap_modes = {
|
||||
str(k): str(v)
|
||||
for k, v in dict(data["power_cap_modes"]).items()
|
||||
if str(v) in ("nvml", "ioctl")
|
||||
}
|
||||
except Exception as exc:
|
||||
log.debug("Could not load user config: %s", exc)
|
||||
|
||||
|
||||
@@ -145,6 +145,9 @@ class NvCurveClient:
|
||||
def snapshots(self) -> list:
|
||||
return self._get("/api/snapshots")
|
||||
|
||||
def limits(self) -> dict:
|
||||
return self._get("/api/limits")
|
||||
|
||||
# ── Profiles ─────────────────────────────────────────────────────────────
|
||||
|
||||
def profiles(self) -> dict:
|
||||
|
||||
@@ -54,6 +54,11 @@ class Config:
|
||||
# Legacy entries (bare curve list) are migrated at load time.
|
||||
fan_curves: dict[str, object] = field(default_factory=dict)
|
||||
|
||||
# Per-GPU power-cap mode: "nvml" (default, never stored) or "ioctl"
|
||||
# (experimental RM power control — permits caps below the VBIOS minimum).
|
||||
# Key = stable GPU identifier (same as auto_load_profiles).
|
||||
power_cap_modes: dict[str, str] = field(default_factory=dict)
|
||||
|
||||
|
||||
# Module-level default config instance.
|
||||
default_config = Config()
|
||||
|
||||
+51
-6
@@ -59,20 +59,37 @@ def _get_handle(gpu_index: int):
|
||||
# ── Power limit ───────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def get_power_limit(gpu_index: int = 0) -> dict:
|
||||
"""Return dict with power_limit_w, default_power_limit_w, min_power_limit_w, max_power_limit_w."""
|
||||
out: dict[str, int | None] = {
|
||||
def get_power_limit(gpu_index: int = 0, mode: str = "nvml") -> dict:
|
||||
"""Return dict with power limit info.
|
||||
|
||||
Keys: power_limit_w, default_power_limit_w, min_power_limit_w,
|
||||
min_power_limit_w_native, max_power_limit_w, rm_power_supported,
|
||||
power_cap_mode.
|
||||
|
||||
mode: "nvml" (default) or "ioctl" (experimental RM power control).
|
||||
min_power_limit_w is the effective minimum: in ioctl mode it is
|
||||
extended to the experimental floor (30 W) when the RM interface is
|
||||
present and validated; min_power_limit_w_native is always the VBIOS
|
||||
minimum. The RM probe is GET-only (no writes) and safe to run on
|
||||
every call.
|
||||
"""
|
||||
out: dict[str, int | bool | str | None] = {
|
||||
"power_limit_w": None,
|
||||
"default_power_limit_w": None,
|
||||
"min_power_limit_w": None,
|
||||
"min_power_limit_w_native": None,
|
||||
"max_power_limit_w": None,
|
||||
"rm_power_supported": False,
|
||||
"power_cap_mode": mode,
|
||||
}
|
||||
try:
|
||||
handle = _get_handle(gpu_index)
|
||||
limit = pynvml.nvmlDeviceGetPowerManagementLimit(handle)
|
||||
constrs = pynvml.nvmlDeviceGetPowerManagementLimitConstraints(handle)
|
||||
out["power_limit_w"] = limit // 1000
|
||||
out["min_power_limit_w"] = constrs[0] // 1000
|
||||
native_min = constrs[0] // 1000
|
||||
out["min_power_limit_w"] = native_min
|
||||
out["min_power_limit_w_native"] = native_min
|
||||
out["max_power_limit_w"] = constrs[1] // 1000
|
||||
try:
|
||||
default = pynvml.nvmlDeviceGetPowerManagementDefaultLimit(handle)
|
||||
@@ -83,9 +100,37 @@ def get_power_limit(gpu_index: int = 0) -> dict:
|
||||
log.warning("get_power_limit: %s", exc)
|
||||
return out
|
||||
|
||||
# GET-only RM discovery — reported so the UI can offer the experimental
|
||||
# mode; the effective minimum only changes in ioctl mode.
|
||||
try:
|
||||
from . import rm_power
|
||||
|
||||
def set_power_limit(limit_w: int, gpu_index: int = 0) -> tuple[bool, str]:
|
||||
"""Set the board power limit (Watts)."""
|
||||
bounds = rm_power.probe_gpu(gpu_index)
|
||||
out["rm_power_supported"] = bounds is not None
|
||||
if bounds is not None and mode == "ioctl":
|
||||
out["min_power_limit_w"] = bounds.lower_min_mw() // 1000
|
||||
except Exception as exc:
|
||||
log.debug("RM power probe failed: %s", exc)
|
||||
return out
|
||||
|
||||
|
||||
def set_power_limit(
|
||||
limit_w: int, gpu_index: int = 0, mode: str = "nvml"
|
||||
) -> tuple[bool, str]:
|
||||
"""Set the board power limit (Watts).
|
||||
|
||||
mode "ioctl" (experimental) applies the limit through the undocumented
|
||||
RM interface, which permits values below the VBIOS minimum. It has no
|
||||
fallback: failures are reported, never silently switched to NVML.
|
||||
"""
|
||||
if mode == "ioctl":
|
||||
from . import rm_power
|
||||
|
||||
try:
|
||||
rm_power.set_power_limit_w(gpu_index, limit_w)
|
||||
return True, "OK"
|
||||
except rm_power.RmPowerError as exc:
|
||||
return False, str(exc)
|
||||
try:
|
||||
handle = _get_handle(gpu_index)
|
||||
pynvml.nvmlDeviceSetPowerManagementLimit(handle, limit_w * 1000)
|
||||
|
||||
@@ -0,0 +1,627 @@
|
||||
"""Undocumented NVIDIA RM power-limit interface (EXPERIMENTAL).
|
||||
|
||||
Port of the approach from LACT PR #1205 (ilya-zlobintsev/LACT): applies board
|
||||
power limits through the private NV2080 power-limit "ordinary client"
|
||||
interface on /dev/nvidiactl, which permits caps below the VBIOS minimum
|
||||
(down to 30 W). The native maximum still applies.
|
||||
|
||||
EXPERIMENTAL — uses an undocumented driver interface. It may break after
|
||||
driver updates. Discovery is GET-only and validates the RM payload against
|
||||
NVML before any write is issued; a failed write restores the previous
|
||||
request (even if it was below the VBIOS minimum).
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import contextlib
|
||||
import ctypes
|
||||
import fcntl
|
||||
import logging
|
||||
import os
|
||||
import struct
|
||||
import sys
|
||||
from collections.abc import Callable
|
||||
from dataclasses import dataclass
|
||||
|
||||
log = logging.getLogger("nvcurve.hal.rm_power")
|
||||
|
||||
# ── ioctl constants (nv-ioctl.h / nv-ioctl-numbers.h) ─────────────────────────
|
||||
|
||||
NV_IOCTL_MAGIC = ord("N") # 0x4E — user-space RM interface
|
||||
NV_ESC_RM_ALLOC = 0x2B
|
||||
NV_ESC_RM_CONTROL = 0x2A
|
||||
|
||||
# 'F' magic interface (kernel-open/common/inc/nv-ioctl-numbers.h) —
|
||||
# NV_ESC_REGISTER_FD lives here, not in the 'N' RM interface.
|
||||
NV_IOCTL_MAGIC_F = ord("F") # 0x46
|
||||
NV_IOCTL_BASE_F = 200
|
||||
NV_ESC_REGISTER_FD = NV_IOCTL_BASE_F + 1 # 201
|
||||
|
||||
# RM class IDs (nv0080.h / nv2080.h)
|
||||
NV01_DEVICE_0 = 0x0080
|
||||
NV20_SUBDEVICE_0 = 0x2080
|
||||
|
||||
# NV01_ROOT GPU queries (ctrl0000gpu.h) — resolve PCI identity to the RM
|
||||
# device/subdevice instance numbers used by NV0080 and NV2080 allocations;
|
||||
# neither number is a Linux device minor.
|
||||
_CTRL_GPU_GET_ATTACHED_IDS = 0x201
|
||||
_CTRL_GPU_GET_ID_INFO_V2 = 0x205
|
||||
_CTRL_GPU_GET_PCI_INFO = 0x21B
|
||||
_MAX_GPUS = 32
|
||||
_INVALID_GPU_ID = 0xFFFFFFFF
|
||||
|
||||
# Private NV2080 power-limit client commands. Payloads compared against
|
||||
# NvAPI and GSP from R595, R610 and R615 (native RM payloads, without
|
||||
# NvAPI's 0x10-byte transport prefix).
|
||||
_PWR_GET_INFO = 0x2080_A630
|
||||
_PWR_GET_CONTROL = 0x2080_A632
|
||||
_PWR_SET_CONTROL = 0x2080_E633
|
||||
_ORDINARY_CLIENT = 0xFE
|
||||
_LOWER_LIMIT_MW = 30_000 # experimental floor: 30 W
|
||||
|
||||
|
||||
def _ioctl_rw(size: int, nr: int, magic: int = NV_IOCTL_MAGIC) -> int:
|
||||
"""Linux ioctl request code: dir=RW, given size/type/nr."""
|
||||
return (2 << 30) | (size << 16) | (magic << 8) | nr
|
||||
|
||||
|
||||
def _ioctl_call(fd: int, code: int, arg) -> None:
|
||||
"""Issue an ioctl, converting errno failures to RmPowerError.
|
||||
|
||||
The driver normally reports failures as an RM status in the parameter
|
||||
struct, but an experimental interface can also fail at the kernel level
|
||||
(ENOTTY/EBADF/EPERM across driver versions). Converting to RmPowerError
|
||||
keeps the module's error contract uniform and lets callers clean up fds.
|
||||
"""
|
||||
try:
|
||||
fcntl.ioctl(fd, code, arg)
|
||||
except OSError as exc:
|
||||
raise RmPowerError(f"ioctl 0x{code:x} failed: {exc}") from exc
|
||||
|
||||
|
||||
# ── NVOS parameter structs (nvos.h) ──────────────────────────────────────────
|
||||
|
||||
|
||||
class _NVOS21(ctypes.Structure):
|
||||
_fields_ = [
|
||||
("hRoot", ctypes.c_uint32),
|
||||
("hObjectParent", ctypes.c_uint32),
|
||||
("hObjectNew", ctypes.c_uint32),
|
||||
("hClass", ctypes.c_uint32),
|
||||
("pAllocParms", ctypes.c_uint64),
|
||||
("paramsSize", ctypes.c_uint32),
|
||||
("status", ctypes.c_uint32),
|
||||
]
|
||||
|
||||
|
||||
class _NVOS64(ctypes.Structure):
|
||||
_fields_ = [
|
||||
("hRoot", ctypes.c_uint32),
|
||||
("hObjectParent", ctypes.c_uint32),
|
||||
("hObjectNew", ctypes.c_uint32),
|
||||
("hClass", ctypes.c_uint32),
|
||||
("pAllocParms", ctypes.c_uint64),
|
||||
("pRightsRequested", ctypes.c_uint64),
|
||||
("paramsSize", ctypes.c_uint32),
|
||||
("flags", ctypes.c_uint32),
|
||||
("status", ctypes.c_uint32),
|
||||
]
|
||||
|
||||
|
||||
class _NVOS54(ctypes.Structure):
|
||||
_fields_ = [
|
||||
("hClient", ctypes.c_uint32),
|
||||
("hObject", ctypes.c_uint32),
|
||||
("cmd", ctypes.c_uint32),
|
||||
("flags", ctypes.c_uint32),
|
||||
("params", ctypes.c_uint64),
|
||||
("paramsSize", ctypes.c_uint32),
|
||||
("status", ctypes.c_uint32),
|
||||
]
|
||||
|
||||
|
||||
class _NV0080_ALLOC(ctypes.Structure):
|
||||
_fields_ = [
|
||||
("deviceId", ctypes.c_uint32),
|
||||
("deviceFlags", ctypes.c_uint32),
|
||||
("vgpuInstance", ctypes.c_uint32),
|
||||
("pad", ctypes.c_uint32),
|
||||
]
|
||||
|
||||
|
||||
class _NV2080_ALLOC(ctypes.Structure):
|
||||
_fields_ = [
|
||||
("subDeviceId", ctypes.c_uint32),
|
||||
("clientShare", ctypes.c_uint32),
|
||||
("flags", ctypes.c_uint32),
|
||||
("pad", ctypes.c_uint32),
|
||||
]
|
||||
|
||||
|
||||
# ── Errors ───────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
class RmPowerError(RuntimeError):
|
||||
"""Raised when the RM power-limit interface is unavailable or fails."""
|
||||
|
||||
|
||||
# ── Power-limit layouts and bounds ───────────────────────────────────────────
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class PowerLimitLayout:
|
||||
"""Byte offsets of the private power-limit payloads for one wire format."""
|
||||
|
||||
name: str
|
||||
info_size: int
|
||||
control_size: int
|
||||
info_min_at: int
|
||||
request_at: int
|
||||
client_at: int
|
||||
mask_end: int
|
||||
|
||||
|
||||
EXTENDED_LAYOUT = PowerLimitLayout(
|
||||
name="extended",
|
||||
info_size=0x924,
|
||||
control_size=0x328,
|
||||
info_min_at=0x28,
|
||||
request_at=0x2C,
|
||||
client_at=0x30,
|
||||
mask_end=0x24,
|
||||
)
|
||||
LEGACY_LAYOUT = PowerLimitLayout(
|
||||
name="legacy",
|
||||
info_size=0x488,
|
||||
control_size=0x188,
|
||||
info_min_at=0xC,
|
||||
request_at=0xC,
|
||||
client_at=0x10,
|
||||
mask_end=0x8,
|
||||
)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class PowerLimitBounds:
|
||||
"""Power limit bounds in milliwatts (NVML/RM units)."""
|
||||
|
||||
min_mw: int
|
||||
default_mw: int
|
||||
max_mw: int
|
||||
|
||||
def lower_min_mw(self) -> int:
|
||||
"""Effective minimum when the experimental route is active."""
|
||||
return min(self.min_mw, _LOWER_LIMIT_MW)
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class LowerPowerLimit:
|
||||
"""A validated RM power-limit layout that can be written."""
|
||||
|
||||
bounds: PowerLimitBounds
|
||||
layout: PowerLimitLayout
|
||||
|
||||
def lower_min_mw(self) -> int:
|
||||
return self.bounds.lower_min_mw()
|
||||
|
||||
|
||||
# ── PCI identity → RM instance resolution ────────────────────────────────────
|
||||
|
||||
|
||||
@dataclass(frozen=True)
|
||||
class PciLocation:
|
||||
domain: int
|
||||
bus: int
|
||||
dev: int
|
||||
func: int = 0
|
||||
|
||||
|
||||
def resolve_gpu_instance(
|
||||
pci: PciLocation,
|
||||
query: Callable[[int, bytearray], None],
|
||||
) -> tuple[int, int]:
|
||||
"""Resolve (device_instance, subdevice_instance) by PCI identity.
|
||||
|
||||
/dev/nvidiaN minors and RM device instances can have different orders;
|
||||
the RM object must be matched by PCI domain/bus/slot, not by index.
|
||||
The RM query exposes domain/bus/slot but no PCI function, so only
|
||||
function-zero devices can be matched (never another function of a
|
||||
multifunction device).
|
||||
"""
|
||||
if pci.func != 0:
|
||||
raise RmPowerError("RM GPU lookup requires PCI function zero")
|
||||
|
||||
attached = bytearray(_MAX_GPUS * 4)
|
||||
query(_CTRL_GPU_GET_ATTACHED_IDS, attached)
|
||||
|
||||
for i in range(_MAX_GPUS):
|
||||
gpu_id = struct.unpack_from("<I", attached, i * 4)[0]
|
||||
if gpu_id == _INVALID_GPU_ID:
|
||||
continue
|
||||
|
||||
# NV0000_CTRL_GPU_GET_PCI_INFO_PARAMS: u32 gpuId, u32 domain,
|
||||
# u16 bus, u16 slot.
|
||||
location = bytearray(12)
|
||||
location[0:4] = struct.pack("<I", gpu_id)
|
||||
query(_CTRL_GPU_GET_PCI_INFO, location)
|
||||
domain, bus, slot = struct.unpack_from("<IHH", location, 4)
|
||||
if (domain, bus, slot) != (pci.domain, pci.bus, pci.dev):
|
||||
continue
|
||||
|
||||
# NV0000_CTRL_GPU_GET_ID_INFO_V2_PARAMS: eight u32 fields, with
|
||||
# deviceInstance/subDeviceInstance at +8/+12.
|
||||
info = bytearray(32)
|
||||
info[0:4] = struct.pack("<I", gpu_id)
|
||||
query(_CTRL_GPU_GET_ID_INFO_V2, info)
|
||||
device, subdevice = struct.unpack_from("<II", info, 8)
|
||||
return device, subdevice
|
||||
|
||||
raise RmPowerError(f"no RM GPU matches PCI location {pci}")
|
||||
|
||||
|
||||
# ── RM handle ────────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def _rm_control(fd: int, client: int, obj: int, cmd: int, buf: bytearray) -> None:
|
||||
"""Issue an NVOS54 RM control whose parameter block is a byte buffer."""
|
||||
arr = (ctypes.c_uint8 * len(buf)).from_buffer(buf)
|
||||
req = _NVOS54(
|
||||
hClient=client,
|
||||
hObject=obj,
|
||||
cmd=cmd,
|
||||
flags=0,
|
||||
params=ctypes.addressof(arr),
|
||||
paramsSize=len(buf),
|
||||
status=0,
|
||||
)
|
||||
_ioctl_call(fd, _ioctl_rw(ctypes.sizeof(_NVOS54), NV_ESC_RM_CONTROL), req)
|
||||
if req.status != 0:
|
||||
raise RmPowerError(
|
||||
f"RM control 0x{cmd:08x} failed with status 0x{req.status:x}"
|
||||
)
|
||||
|
||||
|
||||
def _alloc_client(fd: int) -> int:
|
||||
"""Allocate an RM client (NVOS21, all-zero parameters)."""
|
||||
req = _NVOS21()
|
||||
_ioctl_call(fd, _ioctl_rw(ctypes.sizeof(_NVOS21), NV_ESC_RM_ALLOC), req)
|
||||
if req.status != 0:
|
||||
raise RmPowerError(f"could not allocate RM client (status 0x{req.status:x})")
|
||||
return req.hObjectNew
|
||||
|
||||
|
||||
def _alloc_object(
|
||||
fd: int, client: int, parent: int, class_id: int, alloc_params: ctypes.Structure
|
||||
) -> int:
|
||||
"""Allocate an RM object (NVOS64) and return its handle."""
|
||||
req = _NVOS64(
|
||||
hRoot=client,
|
||||
hObjectParent=parent,
|
||||
hObjectNew=0,
|
||||
hClass=class_id,
|
||||
pAllocParms=ctypes.addressof(alloc_params),
|
||||
pRightsRequested=0,
|
||||
paramsSize=ctypes.sizeof(alloc_params),
|
||||
flags=0,
|
||||
status=0,
|
||||
)
|
||||
_ioctl_call(fd, _ioctl_rw(ctypes.sizeof(_NVOS64), NV_ESC_RM_ALLOC), req)
|
||||
if req.status != 0:
|
||||
raise RmPowerError(
|
||||
f"RM class 0x{class_id:x} allocation failed (status 0x{req.status:x})"
|
||||
)
|
||||
return req.hObjectNew
|
||||
|
||||
|
||||
def _register_fd(device_fd: int, nvidiactl_fd: int) -> None:
|
||||
"""Register the nvidiactl client with the device fd (NV_ESC_REGISTER_FD).
|
||||
|
||||
The ioctl is issued on the /dev/nvidiaN fd; the argument is the
|
||||
nvidiactl fd to associate with it.
|
||||
"""
|
||||
_ioctl_call(
|
||||
device_fd,
|
||||
_ioctl_rw(4, NV_ESC_REGISTER_FD, NV_IOCTL_MAGIC_F),
|
||||
struct.pack("i", nvidiactl_fd),
|
||||
)
|
||||
|
||||
|
||||
class RmHandle:
|
||||
"""An NVIDIA RM client with device + subdevice objects for one GPU."""
|
||||
|
||||
def __init__(
|
||||
self,
|
||||
nvidiactl_fd: int,
|
||||
device_fd: int,
|
||||
client_handle: int,
|
||||
device_handle: int,
|
||||
subdevice_handle: int,
|
||||
) -> None:
|
||||
self._nvidiactl_fd = nvidiactl_fd
|
||||
self._device_fd = device_fd
|
||||
self.client_handle = client_handle
|
||||
self.device_handle = device_handle
|
||||
self.subdevice_handle = subdevice_handle
|
||||
|
||||
@classmethod
|
||||
def open(cls, gpu_index: int) -> RmHandle:
|
||||
"""Open an RM handle for the GPU at the given NVML index.
|
||||
|
||||
The RM device/subdevice instances are resolved by PCI identity
|
||||
(minors and RM instances can have different orders).
|
||||
"""
|
||||
pynvml = _ensure_nvml()
|
||||
|
||||
try:
|
||||
handle = pynvml.nvmlDeviceGetHandleByIndex(gpu_index)
|
||||
minor = int(pynvml.nvmlDeviceGetMinorNumber(handle))
|
||||
pci_info = pynvml.nvmlDeviceGetPciInfo(handle)
|
||||
pci = PciLocation(
|
||||
domain=int(pci_info.domain),
|
||||
bus=int(pci_info.bus),
|
||||
dev=int(pci_info.device),
|
||||
)
|
||||
except Exception as exc:
|
||||
raise RmPowerError(f"NVML query for GPU {gpu_index} failed: {exc}") from exc
|
||||
|
||||
try:
|
||||
nvidiactl_fd = os.open("/dev/nvidiactl", os.O_RDWR)
|
||||
except OSError as exc:
|
||||
raise RmPowerError(f"could not open /dev/nvidiactl: {exc}") from exc
|
||||
|
||||
try:
|
||||
client_handle = _alloc_client(nvidiactl_fd)
|
||||
device_instance, subdevice_instance = resolve_gpu_instance(
|
||||
pci,
|
||||
lambda cmd, buf: _rm_control(
|
||||
nvidiactl_fd, client_handle, client_handle, cmd, buf
|
||||
),
|
||||
)
|
||||
except RmPowerError:
|
||||
os.close(nvidiactl_fd)
|
||||
raise
|
||||
|
||||
try:
|
||||
device_fd = os.open(f"/dev/nvidia{minor}", os.O_RDWR)
|
||||
except OSError as exc:
|
||||
os.close(nvidiactl_fd)
|
||||
raise RmPowerError(f"could not open /dev/nvidia{minor}: {exc}") from exc
|
||||
|
||||
try:
|
||||
_register_fd(device_fd, nvidiactl_fd)
|
||||
device_handle = _alloc_object(
|
||||
nvidiactl_fd,
|
||||
client_handle,
|
||||
client_handle,
|
||||
NV01_DEVICE_0,
|
||||
_NV0080_ALLOC(deviceId=device_instance),
|
||||
)
|
||||
subdevice_handle = _alloc_object(
|
||||
nvidiactl_fd,
|
||||
client_handle,
|
||||
device_handle,
|
||||
NV20_SUBDEVICE_0,
|
||||
_NV2080_ALLOC(subDeviceId=subdevice_instance),
|
||||
)
|
||||
except RmPowerError:
|
||||
os.close(device_fd)
|
||||
os.close(nvidiactl_fd)
|
||||
raise
|
||||
|
||||
return cls(
|
||||
nvidiactl_fd, device_fd, client_handle, device_handle, subdevice_handle
|
||||
)
|
||||
|
||||
def control(self, cmd: int, buf: bytearray) -> None:
|
||||
"""Issue an NVOS54 RM control on the subdevice with a byte buffer."""
|
||||
_rm_control(
|
||||
self._nvidiactl_fd, self.client_handle, self.subdevice_handle, cmd, buf
|
||||
)
|
||||
|
||||
def close(self) -> None:
|
||||
"""Close the fds; the driver reclaims the RM client objects."""
|
||||
with contextlib.suppress(OSError):
|
||||
os.close(self._device_fd)
|
||||
with contextlib.suppress(OSError):
|
||||
os.close(self._nvidiactl_fd)
|
||||
|
||||
|
||||
# ── Power-limit probe / set (pure logic, testable with a fake query) ─────────
|
||||
|
||||
|
||||
def _u32(data: bytearray | bytes, offset: int) -> int:
|
||||
return struct.unpack_from("<I", data, offset)[0]
|
||||
|
||||
|
||||
def _validate_header(layout: PowerLimitLayout, data: bytearray) -> None:
|
||||
if _u32(data, 0) != 0xFF or _u32(data, 4) != 1:
|
||||
raise RmPowerError("unrecognized RM power client layout")
|
||||
# The extended layout has additional mask words; accepting only its low
|
||||
# word would allow an unexpected client to be included in a later SET.
|
||||
if any(byte != 0 for byte in data[8 : layout.mask_end]):
|
||||
raise RmPowerError("unrecognized RM power client layout")
|
||||
|
||||
|
||||
def _read_bounds(
|
||||
layout: PowerLimitLayout, query: Callable[[int, bytearray], None]
|
||||
) -> PowerLimitBounds:
|
||||
info = bytearray(layout.info_size)
|
||||
query(_PWR_GET_INFO, info)
|
||||
_validate_header(layout, info)
|
||||
bounds = PowerLimitBounds(
|
||||
min_mw=_u32(info, layout.info_min_at),
|
||||
default_mw=_u32(info, layout.info_min_at + 4),
|
||||
max_mw=_u32(info, layout.info_min_at + 8),
|
||||
)
|
||||
if not (
|
||||
bounds.min_mw > 0
|
||||
and bounds.min_mw <= bounds.default_mw
|
||||
and bounds.default_mw <= bounds.max_mw
|
||||
):
|
||||
raise RmPowerError("invalid RM power limit bounds")
|
||||
return bounds
|
||||
|
||||
|
||||
def _read_control(
|
||||
layout: PowerLimitLayout, query: Callable[[int, bytearray], None]
|
||||
) -> bytearray:
|
||||
control = bytearray(layout.control_size)
|
||||
control[4:8] = struct.pack("<I", 1)
|
||||
control[layout.client_at] = _ORDINARY_CLIENT
|
||||
query(_PWR_GET_CONTROL, control)
|
||||
_validate_header(layout, control)
|
||||
if control[layout.client_at] != _ORDINARY_CLIENT:
|
||||
raise RmPowerError("unexpected power client")
|
||||
if _u32(control, layout.request_at) in (0, 0xFFFFFFFF):
|
||||
raise RmPowerError("no ordinary power request available")
|
||||
return control
|
||||
|
||||
|
||||
def probe(
|
||||
nvml_bounds: PowerLimitBounds,
|
||||
nvml_current_mw: int,
|
||||
query: Callable[[int, bytearray], None],
|
||||
) -> LowerPowerLimit:
|
||||
"""GET-only discovery of the RM power-limit layout.
|
||||
|
||||
Probes the two known wire formats using GETs only. A driver version
|
||||
number is not evidence that the payload still has the same layout or
|
||||
units, so the bounds and the current request are validated against
|
||||
NVML. Discovery never issues a SET.
|
||||
"""
|
||||
if sys.byteorder != "little":
|
||||
raise RmPowerError("little-endian host required")
|
||||
errors: list[str] = []
|
||||
for layout in (EXTENDED_LAYOUT, LEGACY_LAYOUT):
|
||||
try:
|
||||
bounds = _read_bounds(layout, query)
|
||||
if bounds != nvml_bounds:
|
||||
raise RmPowerError("RM power bounds differ from NVML")
|
||||
control = _read_control(layout, query)
|
||||
if _u32(control, layout.request_at) != nvml_current_mw:
|
||||
raise RmPowerError("RM ordinary power request differs from NVML")
|
||||
return LowerPowerLimit(bounds=bounds, layout=layout)
|
||||
except RmPowerError as exc:
|
||||
errors.append(f"{layout.name}: {exc}")
|
||||
raise RmPowerError("no compatible RM power layout: " + "; ".join(errors))
|
||||
|
||||
|
||||
def set_limit(
|
||||
limit_mw: int,
|
||||
support: LowerPowerLimit,
|
||||
query: Callable[[int, bytearray], None],
|
||||
) -> None:
|
||||
"""Set the ordinary-client power request with readback verification.
|
||||
|
||||
Keeps the entire current payload, changing only entry 0's request.
|
||||
Mask 1 and selector 0xFE prevent modifying any other entry or the
|
||||
additional F8 client. A failed SET can have side effects, so the
|
||||
previous request is restored even on transport failure — and the
|
||||
restore uses 0xFE so a previous limit below the VBIOS minimum can
|
||||
also be restored.
|
||||
"""
|
||||
layout = support.layout
|
||||
bounds = _read_bounds(layout, query)
|
||||
if bounds != support.bounds:
|
||||
raise RmPowerError("RM power bounds changed since discovery")
|
||||
lower = bounds.lower_min_mw()
|
||||
if not (lower <= limit_mw <= bounds.max_mw):
|
||||
raise RmPowerError(
|
||||
f"power limit {limit_mw} mW outside supported range "
|
||||
f"{lower}..{bounds.max_mw} mW"
|
||||
)
|
||||
|
||||
before = _read_control(layout, query)
|
||||
if _u32(before, layout.request_at) == limit_mw:
|
||||
return
|
||||
|
||||
expected = bytearray(before)
|
||||
expected[layout.request_at : layout.request_at + 4] = struct.pack("<I", limit_mw)
|
||||
|
||||
try:
|
||||
request = bytearray(expected)
|
||||
query(_PWR_SET_CONTROL, request)
|
||||
if _read_control(layout, query) != expected:
|
||||
raise RmPowerError("power request readback differs")
|
||||
except RmPowerError as apply_error:
|
||||
try:
|
||||
restore = bytearray(before)
|
||||
query(_PWR_SET_CONTROL, restore)
|
||||
if _read_control(layout, query) != before:
|
||||
raise RmPowerError("restored power request differs")
|
||||
except RmPowerError as restore_error:
|
||||
raise RmPowerError(
|
||||
f"power request failed: {apply_error}; "
|
||||
f"restoration also failed: {restore_error}"
|
||||
) from None
|
||||
raise RmPowerError(f"{apply_error} (previous power request restored)") from None
|
||||
|
||||
|
||||
# ── High-level API (wires NVML state + RmHandle to the pure logic) ───────────
|
||||
|
||||
|
||||
def _ensure_nvml():
|
||||
"""Return pynvml with NVML initialized (nvmlInit is refcounted)."""
|
||||
import pynvml
|
||||
|
||||
pynvml.nvmlInit()
|
||||
return pynvml
|
||||
|
||||
|
||||
def _nvml_power_state(gpu_index: int) -> tuple[PowerLimitBounds, int]:
|
||||
"""Return (bounds, current_mw) from NVML for the given GPU."""
|
||||
pynvml = _ensure_nvml()
|
||||
|
||||
try:
|
||||
handle = pynvml.nvmlDeviceGetHandleByIndex(gpu_index)
|
||||
min_mw, max_mw = pynvml.nvmlDeviceGetPowerManagementLimitConstraints(handle)
|
||||
default_mw = pynvml.nvmlDeviceGetPowerManagementDefaultLimit(handle)
|
||||
current_mw = pynvml.nvmlDeviceGetPowerManagementLimit(handle)
|
||||
bounds = PowerLimitBounds(int(min_mw), int(default_mw), int(max_mw))
|
||||
current = int(current_mw)
|
||||
except Exception as exc:
|
||||
raise RmPowerError(
|
||||
f"NVML power state for GPU {gpu_index} unavailable: {exc}"
|
||||
) from exc
|
||||
return bounds, current
|
||||
|
||||
|
||||
def probe_gpu(gpu_index: int = 0) -> PowerLimitBounds | None:
|
||||
"""GET-only discovery of the RM power-limit interface for a GPU.
|
||||
|
||||
Returns the validated power bounds (milliwatts) when a compatible RM
|
||||
layout is present, else None. Never issues a write.
|
||||
"""
|
||||
try:
|
||||
bounds, current = _nvml_power_state(gpu_index)
|
||||
except Exception as exc:
|
||||
log.debug("RM probe: NVML state unavailable: %s", exc)
|
||||
return None
|
||||
try:
|
||||
handle = RmHandle.open(gpu_index)
|
||||
except RmPowerError as exc:
|
||||
log.debug("RM probe: handle open failed: %s", exc)
|
||||
return None
|
||||
try:
|
||||
probe(bounds, current, handle.control)
|
||||
return bounds
|
||||
except RmPowerError as exc:
|
||||
log.debug("RM probe: %s", exc)
|
||||
return None
|
||||
finally:
|
||||
handle.close()
|
||||
|
||||
|
||||
def set_power_limit_w(gpu_index: int, limit_w: int) -> None:
|
||||
"""Set the board power limit (watts) via the RM interface.
|
||||
|
||||
Raises RmPowerError on any failure (probe, range, write, readback).
|
||||
A failed write restores the previous request.
|
||||
"""
|
||||
bounds, current = _nvml_power_state(gpu_index)
|
||||
handle = RmHandle.open(gpu_index)
|
||||
try:
|
||||
support = probe(bounds, current, handle.control)
|
||||
set_limit(int(limit_w) * 1000, support, handle.control)
|
||||
finally:
|
||||
handle.close()
|
||||
@@ -9,7 +9,7 @@ from .errors import NVAPI_ERRORS
|
||||
|
||||
def load_nvapi() -> ctypes.CDLL:
|
||||
"""Load libnvidia-api.so from the NVIDIA driver."""
|
||||
for name in ("libnvidia-api.so", "libnvidia-api.so.1"):
|
||||
for name in ("libnvidia-api.so", "libnvidia-api.so.1"): # gitleaks:allow
|
||||
try:
|
||||
return ctypes.CDLL(name)
|
||||
except OSError:
|
||||
|
||||
@@ -43,7 +43,8 @@ def apply_profile(gpu_index: int, name: str, cfg) -> list[str]:
|
||||
errs.append(f"Mem offset: {msg}")
|
||||
|
||||
if profile.power_limit_w is not None:
|
||||
ok, msg = set_power_limit(profile.power_limit_w, gpu_index)
|
||||
mode = profile.power_cap_mode or "nvml"
|
||||
ok, msg = set_power_limit(profile.power_limit_w, gpu_index, mode)
|
||||
if not ok:
|
||||
errs.append(f"Power limit: {msg}")
|
||||
|
||||
|
||||
@@ -16,6 +16,9 @@ class ProfileData:
|
||||
curve_deltas: dict[str, int] # { "index": delta_khz }
|
||||
mem_offset_mhz: int | None = None
|
||||
power_limit_w: int | None = None
|
||||
# How power_limit_w is applied: "nvml" (default) or "ioctl" (experimental
|
||||
# RM power control, permits values below the VBIOS minimum).
|
||||
power_cap_mode: str | None = None
|
||||
fan_curve: list[dict[str, int]] | None = None
|
||||
# Fan indices controlled by fan_curve (0-based); None = all fans.
|
||||
fan_targets: list[int] | None = None
|
||||
@@ -55,6 +58,9 @@ def load_profile(filepath: str) -> ProfileData:
|
||||
# Drop removed fields so old profiles don't cause TypeError.
|
||||
for obsolete in ("gpu_locked_min_mhz", "gpu_locked_max_mhz", "vram_p0_offset_mhz"):
|
||||
data.pop(obsolete, None)
|
||||
# Normalize the experimental power-cap mode; unknown values fall back to NVML.
|
||||
if data.get("power_cap_mode") not in (None, "nvml", "ioctl"):
|
||||
data["power_cap_mode"] = None
|
||||
return ProfileData(**data)
|
||||
|
||||
|
||||
|
||||
+268
-6
@@ -72,6 +72,7 @@ from .profiles.native import (
|
||||
rename_profile,
|
||||
save_profile,
|
||||
)
|
||||
from .wireview import create_device, find_wireview_ports
|
||||
from .safety import check_negative_freq_warnings, validate_write
|
||||
|
||||
log = logging.getLogger("nvcurve.server")
|
||||
@@ -103,6 +104,15 @@ def _open_browser_as_user(url: str) -> None:
|
||||
_state: dict[str, Any] = {
|
||||
"gpus": {}, # dict[int, dict] mapping gpu_index -> gpu state
|
||||
"config": default_config,
|
||||
"wireview": {
|
||||
"clients": set(), # connected /ws/wireview clients
|
||||
"device": None, # WireViewSerialDevice | WireViewHwmonDevice | None
|
||||
"info": None, # device identity (static per connection)
|
||||
"last_sample": None, # most recent sensor sample
|
||||
"connected": False, # last read succeeded
|
||||
"failures": 0, # consecutive failed reads
|
||||
"rejected_ports": set(), # ports reporting an unsupported product
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
@@ -258,6 +268,126 @@ async def _fan_poller(gpu_index: int) -> None:
|
||||
await asyncio.sleep(2.0)
|
||||
|
||||
|
||||
# ── WireView Pro II (Thermal Grizzly) ─────────────────────────────────────────
|
||||
# Consecutive failed reads after which a connected device is given up. Same
|
||||
# rule as the exporter: one corrupt frame or a slow read never trips it, and
|
||||
# an unplug is caught at once by the port-node check.
|
||||
WIREVIEW_MAX_FAILED_READS = 6
|
||||
WIREVIEW_WATCHDOG_INTERVAL_S = 5.0
|
||||
|
||||
|
||||
async def _wireview_disconnect() -> None:
|
||||
"""Drop the connected WireView device and tell clients it is gone."""
|
||||
wv = _state["wireview"]
|
||||
device = wv["device"]
|
||||
if device is None:
|
||||
return
|
||||
wv["device"] = None
|
||||
wv["connected"] = False
|
||||
wv["info"] = None
|
||||
wv["last_sample"] = None
|
||||
wv["failures"] = 0
|
||||
await _run(device.close)
|
||||
log.info("WireView disconnected")
|
||||
await _broadcast(wv["clients"], {"type": "unavailable"})
|
||||
|
||||
|
||||
async def _wireview_connect() -> None:
|
||||
"""Detect and connect a WireView Pro II. No-op when none is attached or
|
||||
one is already connected."""
|
||||
wv = _state["wireview"]
|
||||
if wv["device"] is not None:
|
||||
return
|
||||
# Forget rejected ports that disappeared, so a re-plug gets a fresh probe.
|
||||
if wv["rejected_ports"]:
|
||||
present = set(await _run(find_wireview_ports))
|
||||
wv["rejected_ports"] &= present
|
||||
device = await _run(create_device)
|
||||
if device is None:
|
||||
return
|
||||
# A port that reported an unsupported product is not re-probed (and
|
||||
# re-logged) on every watchdog tick.
|
||||
if device.port in wv["rejected_ports"]:
|
||||
return
|
||||
ok = await _run(device.connect)
|
||||
if not ok:
|
||||
await _run(device.close)
|
||||
if device.rejected:
|
||||
wv["rejected_ports"].add(device.port)
|
||||
return
|
||||
wv["device"] = device
|
||||
wv["failures"] = 0
|
||||
wv["connected"] = True
|
||||
wv["info"] = await _run(device.info)
|
||||
# Push the first sample right away so clients don't wait for the next tick.
|
||||
sample = await _run(device.read_sample)
|
||||
if sample is not None:
|
||||
wv["last_sample"] = sample
|
||||
await _broadcast(
|
||||
wv["clients"], {"type": "sample", "info": wv["info"], "sample": sample}
|
||||
)
|
||||
else:
|
||||
wv["connected"] = False
|
||||
|
||||
|
||||
async def _wireview_poller() -> None:
|
||||
"""Read the WireView at poll_interval_s and push to connected WS clients.
|
||||
|
||||
Skips reads while no client is subscribed (like the monitor poller), so
|
||||
idle polling never contends for the port with other tools (official GUI,
|
||||
wireviewd)."""
|
||||
cfg: Config = _state["config"]
|
||||
while True:
|
||||
try:
|
||||
wv = _state["wireview"]
|
||||
device = wv["device"]
|
||||
if device is not None and wv["clients"]:
|
||||
sample = await _run(device.read_sample)
|
||||
if wv["device"] is not device:
|
||||
# The device was disconnected (unplug) while the read was
|
||||
# in flight — drop the stale result instead of
|
||||
# resurrecting state.
|
||||
pass
|
||||
elif sample is not None:
|
||||
wv["failures"] = 0
|
||||
wv["connected"] = True
|
||||
wv["last_sample"] = sample
|
||||
await _broadcast(
|
||||
wv["clients"],
|
||||
{"type": "sample", "info": wv["info"], "sample": sample},
|
||||
)
|
||||
else:
|
||||
wv["failures"] += 1
|
||||
if (
|
||||
not await _run(device.node_exists)
|
||||
or wv["failures"] >= WIREVIEW_MAX_FAILED_READS
|
||||
):
|
||||
await _wireview_disconnect()
|
||||
except asyncio.CancelledError:
|
||||
return
|
||||
except Exception as exc:
|
||||
log.warning("WireView poller error: %s", exc)
|
||||
await asyncio.sleep(cfg.poll_interval_s)
|
||||
|
||||
|
||||
async def _wireview_watchdog() -> None:
|
||||
"""Hot-plug detection: connect when a WireView appears, disconnect when
|
||||
its device node disappears."""
|
||||
while True:
|
||||
await asyncio.sleep(WIREVIEW_WATCHDOG_INTERVAL_S)
|
||||
try:
|
||||
wv = _state["wireview"]
|
||||
device = wv["device"]
|
||||
if device is None:
|
||||
await _wireview_connect()
|
||||
elif not await _run(device.node_exists):
|
||||
await _wireview_disconnect()
|
||||
except asyncio.CancelledError:
|
||||
return
|
||||
except Exception as exc:
|
||||
log.warning("WireView watchdog error: %s", exc)
|
||||
|
||||
|
||||
async def _activate_fan_curve(
|
||||
gpu_index: int, curve: list, fans: list[int] | None = None
|
||||
) -> None:
|
||||
@@ -380,6 +510,16 @@ async def lifespan(app: FastAPI):
|
||||
except Exception as exc:
|
||||
log.error("Failed to initialize GPU %d: %s", idx, exc)
|
||||
|
||||
# ── WireView Pro II detection ─────────────────────────────────────────────
|
||||
# Independent of GPU discovery: the tab appears whenever the connector
|
||||
# monitor is attached, and the watchdog handles hot-plug afterwards.
|
||||
try:
|
||||
await _wireview_connect()
|
||||
except Exception as exc:
|
||||
log.warning("WireView initial detection failed: %s", exc)
|
||||
poller_tasks.append(asyncio.create_task(_wireview_poller()))
|
||||
poller_tasks.append(asyncio.create_task(_wireview_watchdog()))
|
||||
|
||||
# ── Backward Compatibility Bridge ──────────────────────────────────────────
|
||||
# NOTE: This auto-load path is for users running the server directly (e.g.
|
||||
# via an old systemd unit file that lacks the new daemon mode).
|
||||
@@ -465,6 +605,14 @@ async def lifespan(app: FastAPI):
|
||||
with suppress(asyncio.CancelledError):
|
||||
await task
|
||||
|
||||
# Release the WireView port (if connected) so other tools can use it.
|
||||
wv = _state["wireview"]
|
||||
if wv["device"] is not None:
|
||||
with suppress(Exception):
|
||||
await _run(wv["device"].close)
|
||||
wv["device"] = None
|
||||
wv["connected"] = False
|
||||
|
||||
for gpu_index, g_state in _state["gpus"].items():
|
||||
if g_state.get("fan_poller_task"):
|
||||
g_state["fan_poller_task"].cancel()
|
||||
@@ -569,6 +717,8 @@ class SnapshotRestoreRequest(BaseModel):
|
||||
class LimitsRequest(BaseModel):
|
||||
power_limit_w: int | None = None
|
||||
mem_offset_mhz: int | None = None
|
||||
# "nvml" (default) or "ioctl" (experimental RM power control).
|
||||
power_cap_mode: str | None = None
|
||||
|
||||
|
||||
class ProfileSaveRequest(BaseModel):
|
||||
@@ -914,6 +1064,17 @@ def _persist_fan_curves(fan_curves: dict) -> None:
|
||||
_persist_config_field("fan_curves", fan_curves if fan_curves else None)
|
||||
|
||||
|
||||
def _persist_power_cap_modes(modes: dict[str, str]) -> None:
|
||||
"""Persist the per-GPU experimental power-cap mode dict to config.json."""
|
||||
_persist_config_field("power_cap_modes", modes if modes else None)
|
||||
|
||||
|
||||
def _power_cap_mode(cfg: Config, gpu_index: int) -> str:
|
||||
"""Return the effective power-cap mode for a GPU ("nvml" or "ioctl")."""
|
||||
mode = cfg.power_cap_modes.get(_gpu_stable_key(gpu_index), "nvml")
|
||||
return mode if mode in ("nvml", "ioctl") else "nvml"
|
||||
|
||||
|
||||
@app.get("/api/profiles")
|
||||
async def api_profiles(gpu_index: int = 0):
|
||||
"""List saved native profiles, the active profile name, and the auto-load profile name."""
|
||||
@@ -941,13 +1102,15 @@ async def api_profile_save(req: ProfileSaveRequest, gpu_index: int = 0):
|
||||
curve_deltas = {str(p.index): p.delta_khz for p in state.points if p.delta_khz != 0}
|
||||
|
||||
try:
|
||||
power_info = await _run(get_power_limit, gpu_index)
|
||||
mode = _power_cap_mode(cfg, gpu_index)
|
||||
power_info = await _run(get_power_limit, gpu_index, mode)
|
||||
offsets = await _run(get_clock_offsets, gpu_index)
|
||||
power_limit_w = power_info.get("power_limit_w")
|
||||
mem_offset_mhz = offsets.get("mem_offset_mhz")
|
||||
except Exception:
|
||||
power_limit_w = None
|
||||
mem_offset_mhz = None
|
||||
mode = "nvml"
|
||||
|
||||
data = ProfileData(
|
||||
name=req.name,
|
||||
@@ -955,6 +1118,7 @@ async def api_profile_save(req: ProfileSaveRequest, gpu_index: int = 0):
|
||||
curve_deltas=curve_deltas,
|
||||
mem_offset_mhz=mem_offset_mhz,
|
||||
power_limit_w=power_limit_w,
|
||||
power_cap_mode=mode,
|
||||
fan_curve=g_state.get("fan_curve") if g_state.get("fan_curve_active") else None,
|
||||
fan_targets=g_state.get("fan_targets")
|
||||
if g_state.get("fan_curve_active")
|
||||
@@ -1075,7 +1239,8 @@ async def _apply_profile(name: str, gpu_index: int = 0) -> list[str]:
|
||||
errs.append(f"Mem offset: {msg}")
|
||||
|
||||
if profile.power_limit_w is not None:
|
||||
ok, msg = await _run(set_power_limit, profile.power_limit_w, gpu_index)
|
||||
mode = profile.power_cap_mode or "nvml"
|
||||
ok, msg = await _run(set_power_limit, profile.power_limit_w, gpu_index, mode)
|
||||
if not ok:
|
||||
errs.append(f"Power limit: {msg}")
|
||||
|
||||
@@ -1204,7 +1369,9 @@ async def api_config_update(req: ConfigUpdateRequest):
|
||||
@app.get("/api/limits")
|
||||
async def api_limits(gpu_index: int = 0):
|
||||
"""Current performance limits: power and clock offsets."""
|
||||
power = await _run(get_power_limit, gpu_index)
|
||||
cfg: Config = _state["config"]
|
||||
mode = _power_cap_mode(cfg, gpu_index)
|
||||
power = await _run(get_power_limit, gpu_index, mode)
|
||||
offsets = await _run(get_clock_offsets, gpu_index)
|
||||
mem_off_range = await _run(get_mem_offset_range, gpu_index)
|
||||
return {
|
||||
@@ -1218,10 +1385,36 @@ async def api_limits(gpu_index: int = 0):
|
||||
async def api_limits_update(req: LimitsRequest, gpu_index: int = 0):
|
||||
"""Update performance limits."""
|
||||
g_state = _get_gpu_state(gpu_index)
|
||||
cfg: Config = _state["config"]
|
||||
errs = []
|
||||
|
||||
if req.power_cap_mode is not None:
|
||||
if req.power_cap_mode not in ("nvml", "ioctl"):
|
||||
raise HTTPException(
|
||||
status_code=400, detail="power_cap_mode must be 'nvml' or 'ioctl'"
|
||||
)
|
||||
if req.power_cap_mode == "ioctl":
|
||||
# Verify the GPU actually exposes the RM interface before enabling,
|
||||
# so a client can't lock a GPU into a mode where every power
|
||||
# operation fails (ioctl mode has no NVML fallback by design).
|
||||
info = await _run(get_power_limit, gpu_index, "ioctl")
|
||||
if not info.get("rm_power_supported"):
|
||||
raise HTTPException(
|
||||
status_code=409,
|
||||
detail="Experimental RM power control is not supported "
|
||||
"on this GPU/driver",
|
||||
)
|
||||
key = _gpu_stable_key(gpu_index)
|
||||
if req.power_cap_mode == "nvml":
|
||||
cfg.power_cap_modes.pop(key, None)
|
||||
else:
|
||||
cfg.power_cap_modes[key] = "ioctl"
|
||||
_persist_power_cap_modes(cfg.power_cap_modes)
|
||||
|
||||
mode = _power_cap_mode(cfg, gpu_index)
|
||||
|
||||
if req.power_limit_w is not None:
|
||||
ok, msg = await _run(set_power_limit, req.power_limit_w, gpu_index)
|
||||
ok, msg = await _run(set_power_limit, req.power_limit_w, gpu_index, mode)
|
||||
if not ok:
|
||||
errs.append(f"Power Limit: {msg}")
|
||||
|
||||
@@ -1284,12 +1477,17 @@ async def _update_offsets_and_broadcast(gpu_index: int) -> None:
|
||||
async def api_limits_reset(gpu_index: int = 0):
|
||||
"""Reset power limit to hardware default and memory clock offset to 0."""
|
||||
g_state = _get_gpu_state(gpu_index)
|
||||
cfg: Config = _state["config"]
|
||||
errs = []
|
||||
|
||||
power = await _run(get_power_limit, gpu_index)
|
||||
# Reset uses the GPU's current mode: in ioctl mode the default is
|
||||
# restored through the RM route (which can also restore a previous
|
||||
# below-VBIOS-minimum cap).
|
||||
mode = _power_cap_mode(cfg, gpu_index)
|
||||
power = await _run(get_power_limit, gpu_index, mode)
|
||||
default_w = power.get("default_power_limit_w")
|
||||
if default_w is not None:
|
||||
ok, msg = await _run(set_power_limit, default_w, gpu_index)
|
||||
ok, msg = await _run(set_power_limit, default_w, gpu_index, mode)
|
||||
if not ok:
|
||||
errs.append(f"Power Limit: {msg}")
|
||||
|
||||
@@ -1392,6 +1590,25 @@ async def api_fans_speed(req: FanSpeedRequest, gpu_index: int = 0):
|
||||
return {"ok": True}
|
||||
|
||||
|
||||
# ── WireView Pro II (Thermal Grizzly) ─────────────────────────────────────────
|
||||
|
||||
|
||||
@app.get("/api/wireview")
|
||||
async def api_wireview():
|
||||
"""WireView Pro II availability and the most recent sensor sample.
|
||||
|
||||
available: a device is connected (the UI shows the WireView tab).
|
||||
sample: None until the first successful read.
|
||||
"""
|
||||
wv = _state["wireview"]
|
||||
return {
|
||||
"available": wv["device"] is not None,
|
||||
"connected": wv["connected"],
|
||||
"info": wv["info"],
|
||||
"sample": wv["last_sample"],
|
||||
}
|
||||
|
||||
|
||||
# ── Write endpoints ────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
@@ -1782,6 +1999,51 @@ async def ws_curve(ws: WebSocket):
|
||||
g_state["curve_clients"].discard(ws)
|
||||
|
||||
|
||||
@app.websocket("/ws/wireview")
|
||||
async def ws_wireview(ws: WebSocket):
|
||||
"""Stream WireView Pro II sensor samples at poll_interval_s.
|
||||
|
||||
Messages:
|
||||
{"type": "unavailable"} — no device connected
|
||||
{"type": "sample", "info": {...}, "sample": {...}} — new reading
|
||||
"""
|
||||
if not _ws_authenticated(ws):
|
||||
await ws.close(code=1008)
|
||||
return
|
||||
await ws.accept()
|
||||
try:
|
||||
data = await ws.receive_json()
|
||||
if data.get("action") != "subscribe":
|
||||
await ws.close()
|
||||
return
|
||||
except WebSocketDisconnect:
|
||||
return
|
||||
except Exception:
|
||||
await ws.close()
|
||||
return
|
||||
|
||||
wv = _state["wireview"]
|
||||
wv["clients"].add(ws)
|
||||
try:
|
||||
# Send the current state immediately so the client does not have to
|
||||
# wait for the next poll tick.
|
||||
if wv["device"] is not None and wv["last_sample"] is not None:
|
||||
await ws.send_json(
|
||||
{"type": "sample", "info": wv["info"], "sample": wv["last_sample"]}
|
||||
)
|
||||
else:
|
||||
await ws.send_json({"type": "unavailable"})
|
||||
|
||||
while True:
|
||||
await ws.receive_text()
|
||||
except WebSocketDisconnect:
|
||||
pass
|
||||
except Exception:
|
||||
log.debug("wireview ws client error", exc_info=True)
|
||||
finally:
|
||||
wv["clients"].discard(ws)
|
||||
|
||||
|
||||
# ── Frontend SPA ──────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
|
||||
@@ -0,0 +1,682 @@
|
||||
"""Native reader for the Thermal Grizzly WireView Pro II.
|
||||
|
||||
Talks to the 12 VHPWR connector monitor directly over its USB CDC/ACM
|
||||
serial port (STM32, VID 0483 / PID 5740) — no exporter, no GUI, no kernel
|
||||
module required. When the wireview-hwmon kernel module is loaded, the
|
||||
sysfs node is used instead (the wireviewd daemon owns the port in that
|
||||
case, so direct serial would corrupt frames).
|
||||
|
||||
Protocol notes (matches the firmware's DEVICE_STR_LEN=32 layout):
|
||||
* No framing/CRC — a desynced read corrupts arbitrary fields for one
|
||||
poll. Real frames always carry zero padding bytes and a fan duty
|
||||
<= 100; anything else is discarded and the next poll realigns.
|
||||
* The firmware occasionally stops answering the RTS welcome handshake
|
||||
(observed after USB state changes) while still answering every data
|
||||
command, so identification falls back to the vendor-data reply.
|
||||
"""
|
||||
|
||||
import logging
|
||||
import os
|
||||
import struct
|
||||
import time
|
||||
from typing import Any
|
||||
|
||||
try:
|
||||
import serial
|
||||
except ImportError: # pyserial missing — the serial transport is disabled,
|
||||
# but the rest of the server keeps running (a stale venv after a code
|
||||
# update must not take the whole web server down).
|
||||
serial = None
|
||||
|
||||
log = logging.getLogger("nvcurve.wireview")
|
||||
|
||||
_serial_missing_warned = False
|
||||
|
||||
# ── Device identification ─────────────────────────────────────────────────────
|
||||
|
||||
USB_VENDOR_ID = "0483" # STMicroelectronics (CDC/ACM)
|
||||
USB_PRODUCT_ID = "5740" # WireView Pro II normal mode
|
||||
WELCOME_MESSAGE = "Thermal Grizzly WireView Pro II"
|
||||
MAX_WELCOME_LENGTH = 64
|
||||
|
||||
# Vendor/product ids reported by CMD_READ_VENDOR_DATA (not the USB ids).
|
||||
VENDOR_ID_THERMAL_GRIZZLY = 0xEF
|
||||
PRODUCT_ID_PRO2 = 0x05
|
||||
PRODUCT_ID_PRO2_NOCTUA = 0x06
|
||||
|
||||
DEVICE_NAMES = {
|
||||
PRODUCT_ID_PRO2: "WireView Pro II",
|
||||
PRODUCT_ID_PRO2_NOCTUA: "WireView Pro II Noctua Edition",
|
||||
}
|
||||
|
||||
BAUD_RATE = 115200
|
||||
READ_TIMEOUT_S = 1.0
|
||||
|
||||
# ── Serial protocol commands ──────────────────────────────────────────────────
|
||||
|
||||
CMD_READ_VENDOR_DATA = 0x01
|
||||
CMD_READ_UID = 0x02
|
||||
CMD_READ_SENSOR_VALUES = 0x04
|
||||
CMD_READ_CONFIG = 0x05
|
||||
CMD_SCREEN_CHANGE = 0x0C
|
||||
CMD_READ_BUILD_INFO = 0x0D
|
||||
|
||||
SCREEN_RESUME_UPDATES = 0xF1
|
||||
|
||||
# ── Wire layout ───────────────────────────────────────────────────────────────
|
||||
# SensorStruct (100 bytes, little-endian, pack=4):
|
||||
# 4x int16 temperatures (0.1 °C): in, out, ext1, ext2
|
||||
# uint16 Vdd (mV)
|
||||
# uint8 fan duty (%)
|
||||
# pad
|
||||
# 6x { int16 voltage (mV), pad, uint32 current (mA), uint32 power (mW) }
|
||||
# uint32 total power (mW)
|
||||
# uint32 total current (mA)
|
||||
# uint16 avg voltage (mV)
|
||||
# uint8 PSU capability (0=600W, 1=450W, 2=300W, 3=150W)
|
||||
# pad
|
||||
# uint16 fault status mask
|
||||
# uint16 fault log mask
|
||||
SENSOR_STRUCT = struct.Struct(
|
||||
"<4hHBx" + "".join("hxxII" for _ in range(6)) + "IIHBxHH"
|
||||
)
|
||||
SENSOR_STRUCT_SIZE = SENSOR_STRUCT.size # 100
|
||||
|
||||
# BuildStruct: VendorData(3) + ProductName(32) + BuildInfo(32) + NameLength(1)
|
||||
BUILD_STRUCT_SIZE = 3 + 32 + 32 + 1
|
||||
BUILD_INFO_OFFSET = 3 + 32 # 35
|
||||
|
||||
PSU_CAPABILITY_W = {0: 600, 1: 450, 2: 300, 3: 150}
|
||||
|
||||
# Fault bitmask (both the active status and the latched log use these bits).
|
||||
FAULT_BITS = {
|
||||
0: "Chip over-temperature",
|
||||
1: "Sensor over-temperature",
|
||||
2: "Over-current (OCP)",
|
||||
3: "Wire over-current",
|
||||
4: "Over-power (OPP)",
|
||||
5: "Current imbalance",
|
||||
}
|
||||
|
||||
|
||||
def decode_faults(mask: int) -> list[str]:
|
||||
"""Human-readable names of the active fault bits in a status/log mask."""
|
||||
return [
|
||||
FAULT_BITS[bit]
|
||||
for bit in sorted(FAULT_BITS)
|
||||
if mask & (1 << bit)
|
||||
]
|
||||
|
||||
|
||||
def is_supported_product(vendor_id: int, product_id: int) -> bool:
|
||||
"""True for the products the Pro II protocol serves (5 and 6)."""
|
||||
return (
|
||||
vendor_id == VENDOR_ID_THERMAL_GRIZZLY
|
||||
and product_id in (PRODUCT_ID_PRO2, PRODUCT_ID_PRO2_NOCTUA)
|
||||
)
|
||||
|
||||
|
||||
def parse_sensor_struct(buf: bytes) -> dict:
|
||||
"""Decode a 100-byte sensor frame into a JSON-serializable sample.
|
||||
|
||||
Totals are computed from the per-pin readings (voltage * current),
|
||||
matching the exporter's output; the device's own total fields are
|
||||
not used.
|
||||
"""
|
||||
fields = SENSOR_STRUCT.unpack(buf)
|
||||
ts_in, ts_out, ts_ext1, ts_ext2, _vdd, fan_duty = fields[:6]
|
||||
pin_fields = fields[6:24]
|
||||
(
|
||||
_total_power,
|
||||
_total_current,
|
||||
_avg_voltage,
|
||||
psu_cap,
|
||||
fault_status,
|
||||
fault_log,
|
||||
) = fields[24:]
|
||||
|
||||
pins = []
|
||||
power_total = 0.0
|
||||
current_total = 0.0
|
||||
for i in range(6):
|
||||
voltage_v = pin_fields[i * 3] / 1000.0
|
||||
current_a = pin_fields[i * 3 + 1] / 1000.0
|
||||
power_w = pin_fields[i * 3 + 2] / 1000.0
|
||||
pins.append(
|
||||
{
|
||||
"voltage_v": round(voltage_v, 3),
|
||||
"current_a": round(current_a, 3),
|
||||
"power_w": round(power_w, 3),
|
||||
}
|
||||
)
|
||||
power_total += voltage_v * current_a
|
||||
current_total += current_a
|
||||
|
||||
return {
|
||||
"timestamp": time.time(),
|
||||
"power_total_w": round(power_total, 3),
|
||||
"current_total_a": round(current_total, 3),
|
||||
"voltage_avg_v": (
|
||||
round(power_total / current_total, 3) if current_total > 0 else 0.0
|
||||
),
|
||||
"pins": pins,
|
||||
"temp_in_c": ts_in / 10.0,
|
||||
"temp_out_c": ts_out / 10.0,
|
||||
"temp_ext1_c": ts_ext1 / 10.0,
|
||||
"temp_ext2_c": ts_ext2 / 10.0,
|
||||
"fan_duty_pct": fan_duty,
|
||||
"fault_status": fault_status,
|
||||
"fault_log": fault_log,
|
||||
"psu_capability_w": PSU_CAPABILITY_W.get(psu_cap, 0),
|
||||
}
|
||||
|
||||
|
||||
def sensor_frame_is_corrupt(buf: bytes) -> bool:
|
||||
"""Corruption check for a sensor frame: real frames always carry zero
|
||||
padding bytes and a fan duty <= 100. Wire layout: fan duty at offset
|
||||
10, pad1 at 11, pad2 five bytes from the end (before the two 16-bit
|
||||
fault masks)."""
|
||||
if len(buf) < SENSOR_STRUCT_SIZE:
|
||||
return True
|
||||
return buf[10] > 100 or buf[11] != 0 or buf[SENSOR_STRUCT_SIZE - 5] != 0
|
||||
|
||||
|
||||
# ── Device discovery ──────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def _sysfs_matches_wireview(tty_sysfs_dir: str) -> bool:
|
||||
"""Walk up from a resolved tty sysfs path to the USB device and check
|
||||
its idVendor/idProduct."""
|
||||
path = tty_sysfs_dir
|
||||
while path and path != "/":
|
||||
vid_file = os.path.join(path, "idVendor")
|
||||
pid_file = os.path.join(path, "idProduct")
|
||||
if os.path.isfile(vid_file) and os.path.isfile(pid_file):
|
||||
try:
|
||||
with open(vid_file) as f:
|
||||
vid = f.read().strip().lower()
|
||||
with open(pid_file) as f:
|
||||
pid = f.read().strip().lower()
|
||||
except OSError:
|
||||
return False
|
||||
return vid == USB_VENDOR_ID and pid == USB_PRODUCT_ID
|
||||
path = os.path.dirname(path)
|
||||
return False
|
||||
|
||||
|
||||
def find_wireview_ports() -> list[str]:
|
||||
"""Find /dev nodes of connected WireView Pro II devices.
|
||||
|
||||
Checks the stable /dev/wireview-pro2 symlink (created by the
|
||||
99-wireview.rules udev rule) and falls back to a sysfs scan of all
|
||||
ttyACM* ports matched by USB VID/PID.
|
||||
"""
|
||||
ports: set[str] = set()
|
||||
|
||||
link = "/dev/wireview-pro2"
|
||||
if os.path.islink(link) or os.path.exists(link):
|
||||
try:
|
||||
target = os.path.realpath(link)
|
||||
if os.path.exists(target):
|
||||
ports.add(target)
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
sys_class = "/sys/class/tty"
|
||||
if os.path.isdir(sys_class):
|
||||
try:
|
||||
entries = os.listdir(sys_class)
|
||||
except OSError:
|
||||
entries = []
|
||||
for entry in entries:
|
||||
if not entry.startswith("ttyACM"):
|
||||
continue
|
||||
tty_dir = os.path.join(sys_class, entry)
|
||||
try:
|
||||
resolved = os.path.realpath(tty_dir)
|
||||
except OSError:
|
||||
continue
|
||||
if _sysfs_matches_wireview(resolved):
|
||||
ports.add(f"/dev/{entry}")
|
||||
|
||||
return sorted(ports)
|
||||
|
||||
|
||||
def find_hwmon_path() -> str | None:
|
||||
"""Find the wireview-hwmon sysfs node, if the kernel module is loaded."""
|
||||
base = "/sys/class/hwmon"
|
||||
if not os.path.isdir(base):
|
||||
return None
|
||||
try:
|
||||
entries = os.listdir(base)
|
||||
except OSError:
|
||||
return None
|
||||
for entry in entries:
|
||||
name_path = os.path.join(base, entry, "name")
|
||||
try:
|
||||
with open(name_path) as f:
|
||||
if f.read().strip().lower() == "wireview":
|
||||
return os.path.join(base, entry)
|
||||
except OSError:
|
||||
continue
|
||||
return None
|
||||
|
||||
|
||||
# ── Serial transport ──────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
class WireViewSerialDevice:
|
||||
"""Direct serial access to a WireView Pro II.
|
||||
|
||||
The port is opened and closed per transaction (open → flush → write →
|
||||
read → close), matching the proven behavior of the exporter: it keeps
|
||||
the port unheld between polls so other tools (official GUI, wireviewd)
|
||||
can share the device, and a fresh open realigns a desynced stream.
|
||||
"""
|
||||
|
||||
def __init__(self, port: str, baud: int = BAUD_RATE) -> None:
|
||||
self._port = port
|
||||
self._baud = baud
|
||||
self._connected = False
|
||||
self._rejected = False
|
||||
self._vendor_id = 0
|
||||
self._product_id = 0
|
||||
self._firmware_version = ""
|
||||
self._uid = ""
|
||||
self._build = ""
|
||||
self._config_version = -1
|
||||
|
||||
# ── Identity ──
|
||||
|
||||
@property
|
||||
def connected(self) -> bool:
|
||||
return self._connected
|
||||
|
||||
@property
|
||||
def rejected(self) -> bool:
|
||||
"""True when the device reported an unsupported product id. Callers
|
||||
can memoize this so the port is not re-probed (and re-logged) on
|
||||
every watchdog tick."""
|
||||
return self._rejected
|
||||
|
||||
@property
|
||||
def transport(self) -> str:
|
||||
return "serial"
|
||||
|
||||
@property
|
||||
def port(self) -> str:
|
||||
return self._port
|
||||
|
||||
def info(self) -> dict:
|
||||
return {
|
||||
"device_name": DEVICE_NAMES.get(
|
||||
self._product_id, "WireView Pro II"
|
||||
),
|
||||
"hw_rev": f"{self._vendor_id:02X}{self._product_id:02X}",
|
||||
"firmware_version": self._firmware_version,
|
||||
"uid": self._uid,
|
||||
"build": self._build,
|
||||
"transport": self.transport,
|
||||
"port": self._port,
|
||||
}
|
||||
|
||||
def node_exists(self) -> bool:
|
||||
"""Whether the serial device node still exists (unplug check)."""
|
||||
return os.path.exists(self._port)
|
||||
|
||||
# ── Connection ──
|
||||
|
||||
def connect(self) -> bool:
|
||||
"""Identify the device and prepare it for sensor reads.
|
||||
|
||||
The welcome handshake (RTS edge) is the primary identification,
|
||||
but the firmware occasionally stops answering it while still
|
||||
answering every command — a supported vendor-data reply is
|
||||
equally conclusive, so accept either.
|
||||
"""
|
||||
if self._connected:
|
||||
return True
|
||||
|
||||
# The welcome handshake (RTS edge) is the primary identification, but
|
||||
# the firmware occasionally stops answering it while still answering
|
||||
# every command — the supported vendor-data reply below is equally
|
||||
# conclusive, so the welcome is read for logging only.
|
||||
welcome = self._read_welcome()
|
||||
if welcome and welcome != WELCOME_MESSAGE:
|
||||
log.debug(
|
||||
"WireView: unexpected welcome string %r on %s",
|
||||
welcome,
|
||||
self._port,
|
||||
)
|
||||
|
||||
vd = self._transaction(bytes([CMD_READ_VENDOR_DATA]), 3)
|
||||
if vd is None or len(vd) < 3:
|
||||
return False
|
||||
vendor, product, fw = vd[0], vd[1], vd[2]
|
||||
if not is_supported_product(vendor, product):
|
||||
self._rejected = True
|
||||
log.info(
|
||||
"WireView: unsupported product %02X%02X on %s, skipped",
|
||||
vendor,
|
||||
product,
|
||||
self._port,
|
||||
)
|
||||
return False
|
||||
|
||||
self._vendor_id = vendor
|
||||
self._product_id = product
|
||||
self._firmware_version = str(fw)
|
||||
|
||||
cfg = self._transaction(bytes([CMD_READ_CONFIG]), 4)
|
||||
if cfg is None or len(cfg) < 3:
|
||||
return False
|
||||
self._config_version = cfg[2]
|
||||
|
||||
uid = self._transaction(bytes([CMD_READ_UID]), 12)
|
||||
if uid is not None and len(uid) == 12:
|
||||
self._uid = uid.hex().upper()
|
||||
|
||||
# Enable display updates just in case.
|
||||
self._transaction(
|
||||
bytes([CMD_SCREEN_CHANGE, SCREEN_RESUME_UPDATES]), 0
|
||||
)
|
||||
|
||||
build = self._transaction(bytes([CMD_READ_BUILD_INFO]), BUILD_STRUCT_SIZE)
|
||||
if build is not None and len(build) >= BUILD_INFO_OFFSET + 1:
|
||||
self._build = (
|
||||
build[BUILD_INFO_OFFSET : BUILD_INFO_OFFSET + 32]
|
||||
.split(b"\x00")[0]
|
||||
.decode("ascii", errors="replace")
|
||||
)
|
||||
|
||||
self._connected = True
|
||||
log.info(
|
||||
"WireView connected on %s (%s, fw %s)",
|
||||
self._port,
|
||||
self.info()["hw_rev"],
|
||||
self._firmware_version,
|
||||
)
|
||||
return True
|
||||
|
||||
def close(self) -> None:
|
||||
self._connected = False
|
||||
|
||||
# ── Sensor reads ──
|
||||
|
||||
def read_sample(self) -> dict | None:
|
||||
"""Read one sensor sample, or None when the device is unresponsive
|
||||
or the frame is corrupt."""
|
||||
if not self._connected:
|
||||
return None
|
||||
buf = self._transaction(bytes([CMD_READ_SENSOR_VALUES]), SENSOR_STRUCT_SIZE)
|
||||
if buf is None or sensor_frame_is_corrupt(buf):
|
||||
return None
|
||||
return parse_sensor_struct(buf)
|
||||
|
||||
# ── Transport ──
|
||||
|
||||
def _open_port(self):
|
||||
"""Open the serial port, or None when pyserial is missing or the
|
||||
port is unavailable."""
|
||||
global _serial_missing_warned
|
||||
if serial is None:
|
||||
if not _serial_missing_warned:
|
||||
_serial_missing_warned = True
|
||||
log.warning(
|
||||
"pyserial is not installed — WireView serial transport "
|
||||
"disabled (install pyserial, e.g. `uv sync`)"
|
||||
)
|
||||
return None
|
||||
try:
|
||||
return serial.Serial(self._port, self._baud, timeout=READ_TIMEOUT_S)
|
||||
except OSError: # SerialException is an OSError
|
||||
return None
|
||||
|
||||
def _read_welcome(self) -> str | None:
|
||||
"""Assert RTS and read the NUL-terminated welcome string the device
|
||||
answers with. Null when nothing (or no terminator) arrives in time."""
|
||||
ser = self._open_port()
|
||||
if ser is None:
|
||||
return None
|
||||
try:
|
||||
ser.reset_input_buffer()
|
||||
ser.rts = False
|
||||
time.sleep(0.01)
|
||||
ser.rts = True
|
||||
time.sleep(0.01)
|
||||
buf = bytearray()
|
||||
deadline = time.monotonic() + READ_TIMEOUT_S
|
||||
while len(buf) < MAX_WELCOME_LENGTH:
|
||||
remaining = deadline - time.monotonic()
|
||||
if remaining <= 0:
|
||||
break
|
||||
ser.timeout = min(READ_TIMEOUT_S, remaining)
|
||||
chunk = ser.read(MAX_WELCOME_LENGTH - len(buf))
|
||||
if not chunk:
|
||||
break
|
||||
buf.extend(chunk)
|
||||
if b"\x00" in chunk:
|
||||
break
|
||||
time.sleep(0.01)
|
||||
ser.rts = False
|
||||
nul = buf.find(b"\x00")
|
||||
if nul >= 0:
|
||||
return bytes(buf[:nul]).decode("ascii", errors="replace")
|
||||
return None
|
||||
except OSError: # SerialException is an OSError
|
||||
return None
|
||||
finally:
|
||||
ser.close()
|
||||
|
||||
def _transaction(self, cmd: bytes, response_size: int) -> bytes | None:
|
||||
"""Open the port, send cmd, read exactly response_size bytes (one
|
||||
second budget), close the port. None when the port is unavailable
|
||||
or the reply is incomplete."""
|
||||
ser = self._open_port()
|
||||
if ser is None:
|
||||
return None
|
||||
try:
|
||||
ser.reset_input_buffer()
|
||||
if cmd:
|
||||
ser.write(cmd)
|
||||
if response_size == 0:
|
||||
return b""
|
||||
return self._read_exact(ser, response_size)
|
||||
except OSError: # SerialException is an OSError
|
||||
return None
|
||||
finally:
|
||||
ser.close()
|
||||
|
||||
@staticmethod
|
||||
def _read_exact(ser: Any, size: int) -> bytes | None:
|
||||
"""Read exactly size bytes within one second, or None."""
|
||||
buf = bytearray()
|
||||
deadline = time.monotonic() + READ_TIMEOUT_S
|
||||
while len(buf) < size:
|
||||
remaining = deadline - time.monotonic()
|
||||
if remaining <= 0:
|
||||
return None
|
||||
ser.timeout = min(READ_TIMEOUT_S, remaining)
|
||||
chunk = ser.read(size - len(buf))
|
||||
if not chunk:
|
||||
return None
|
||||
buf.extend(chunk)
|
||||
return bytes(buf)
|
||||
|
||||
|
||||
# ── hwmon (sysfs) transport ───────────────────────────────────────────────────
|
||||
|
||||
|
||||
class WireViewHwmonDevice:
|
||||
"""Reads the wireview-hwmon sysfs node (kernel module + wireviewd).
|
||||
|
||||
Used when the module is loaded: the daemon owns the serial port in
|
||||
that case, so direct serial would corrupt frames.
|
||||
"""
|
||||
|
||||
def __init__(self, hwmon_path: str) -> None:
|
||||
self._path = hwmon_path
|
||||
self._connected = False
|
||||
|
||||
@property
|
||||
def connected(self) -> bool:
|
||||
return self._connected
|
||||
|
||||
@property
|
||||
def rejected(self) -> bool:
|
||||
return False
|
||||
|
||||
@property
|
||||
def transport(self) -> str:
|
||||
return "hwmon"
|
||||
|
||||
@property
|
||||
def port(self) -> str:
|
||||
return self._path
|
||||
|
||||
def info(self) -> dict:
|
||||
return {
|
||||
"device_name": "WireView Pro II",
|
||||
"hw_rev": "",
|
||||
"firmware_version": "",
|
||||
"uid": "",
|
||||
"build": "",
|
||||
"transport": self.transport,
|
||||
"port": self._path,
|
||||
}
|
||||
|
||||
def node_exists(self) -> bool:
|
||||
return os.path.isdir(self._path)
|
||||
|
||||
def connect(self) -> bool:
|
||||
if self._connected:
|
||||
return True
|
||||
name_path = os.path.join(self._path, "name")
|
||||
try:
|
||||
with open(name_path) as f:
|
||||
if f.read().strip().lower() != "wireview":
|
||||
return False
|
||||
# Probe that the node actually serves data.
|
||||
with open(os.path.join(self._path, "in0_input")) as f:
|
||||
f.read().strip()
|
||||
except OSError:
|
||||
return False
|
||||
self._connected = True
|
||||
log.info("WireView connected via hwmon (%s)", self._path)
|
||||
return True
|
||||
|
||||
def close(self) -> None:
|
||||
self._connected = False
|
||||
|
||||
def read_sample(self) -> dict | None:
|
||||
if not self._connected:
|
||||
return None
|
||||
try:
|
||||
pin_voltage = [
|
||||
self._read_int(f"in{i}_input") / 1000.0 for i in range(6)
|
||||
]
|
||||
pin_current = [
|
||||
self._read_int(f"curr{i + 1}_input") / 1000.0 for i in range(6)
|
||||
]
|
||||
temp_in = self._read_temp("temp1_input")
|
||||
temp_out = self._read_temp("temp2_input")
|
||||
temp_ext1 = self._read_temp("temp3_input")
|
||||
temp_ext2 = self._read_temp("temp4_input")
|
||||
|
||||
fault_status = self._read_int_or("fault_status_raw")
|
||||
if fault_status is None:
|
||||
fault_status = 0xFFFF if self._read_int("intrusion0_alarm") else 0
|
||||
fault_log = self._read_int_or("fault_log_raw")
|
||||
if fault_log is None:
|
||||
fault_log = 0xFFFF if self._read_int("intrusion1_alarm") else 0
|
||||
|
||||
psu_cap_uw = self._read_int_or("power1_cap")
|
||||
if psu_cap_uw is not None:
|
||||
psu_capability = int(round(psu_cap_uw / 1_000_000.0))
|
||||
else:
|
||||
psu_cap = self._read_int_or("psu_cap")
|
||||
psu_capability = PSU_CAPABILITY_W.get(psu_cap or 0, 0)
|
||||
|
||||
pwm = self._read_int_or("pwm1")
|
||||
if pwm is not None:
|
||||
fan_duty = int(round(min(255, max(0, pwm)) * 100 / 255.0))
|
||||
else:
|
||||
fan_duty = self._read_int("fan1_input")
|
||||
|
||||
power_total = sum(
|
||||
v * i
|
||||
for v, i in zip(pin_voltage, pin_current, strict=True)
|
||||
)
|
||||
current_total = sum(pin_current)
|
||||
|
||||
return {
|
||||
"timestamp": time.time(),
|
||||
"power_total_w": round(power_total, 3),
|
||||
"current_total_a": round(current_total, 3),
|
||||
"voltage_avg_v": (
|
||||
round(power_total / current_total, 3)
|
||||
if current_total > 0
|
||||
else 0.0
|
||||
),
|
||||
"pins": [
|
||||
{
|
||||
"voltage_v": round(v, 3),
|
||||
"current_a": round(i, 3),
|
||||
"power_w": round(v * i, 3),
|
||||
}
|
||||
for v, i in zip(pin_voltage, pin_current, strict=True)
|
||||
],
|
||||
"temp_in_c": temp_in,
|
||||
"temp_out_c": temp_out,
|
||||
"temp_ext1_c": temp_ext1,
|
||||
"temp_ext2_c": temp_ext2,
|
||||
"fan_duty_pct": fan_duty,
|
||||
"fault_status": fault_status,
|
||||
"fault_log": fault_log,
|
||||
"psu_capability_w": psu_capability,
|
||||
}
|
||||
except OSError:
|
||||
return None
|
||||
|
||||
def _read_int(self, filename: str) -> int:
|
||||
"""Read an integer sysfs attribute; 0 when missing or unreadable
|
||||
(matches the exporter's ReadIntFile)."""
|
||||
try:
|
||||
with open(os.path.join(self._path, filename)) as f:
|
||||
return int(f.read().strip())
|
||||
except (OSError, ValueError):
|
||||
return 0
|
||||
|
||||
def _read_int_or(self, filename: str) -> int | None:
|
||||
try:
|
||||
with open(os.path.join(self._path, filename)) as f:
|
||||
return int(f.read().strip())
|
||||
except (OSError, ValueError):
|
||||
return None
|
||||
|
||||
def _read_temp(self, filename: str) -> float | None:
|
||||
"""Read a temperature sysfs attribute (m°C); None when missing or
|
||||
unreadable. NaN would poison the whole sample: the WebSocket
|
||||
serializer emits a bare NaN token (invalid JSON) and the REST
|
||||
JSONResponse rejects it with a 500."""
|
||||
try:
|
||||
with open(os.path.join(self._path, filename)) as f:
|
||||
return int(f.read().strip()) / 1000.0
|
||||
except (OSError, ValueError):
|
||||
return None
|
||||
|
||||
|
||||
def create_device() -> WireViewSerialDevice | WireViewHwmonDevice | None:
|
||||
"""Create a device for the first available WireView, or None.
|
||||
|
||||
Preference: the wireview-hwmon sysfs node (the wireviewd daemon owns
|
||||
the serial port in that case), otherwise direct serial on the first
|
||||
matching /dev/ttyACM*.
|
||||
"""
|
||||
hwmon_path = find_hwmon_path()
|
||||
if hwmon_path:
|
||||
return WireViewHwmonDevice(hwmon_path)
|
||||
ports = find_wireview_ports()
|
||||
if ports:
|
||||
return WireViewSerialDevice(ports[0])
|
||||
return None
|
||||
@@ -14,14 +14,26 @@ dependencies = [
|
||||
"pydantic>=2.0",
|
||||
"httpx>=0.27",
|
||||
"bcrypt>=4.0",
|
||||
"pyserial>=3.5",
|
||||
]
|
||||
|
||||
[project.scripts]
|
||||
nvcurve = "nvcurve.cli:main"
|
||||
|
||||
[dependency-groups]
|
||||
dev = [
|
||||
"hatchling", # enables local `hatch build` and resolves hatch_build.py imports
|
||||
]
|
||||
|
||||
[tool.hatch.build.hooks.custom]
|
||||
|
||||
[tool.hatch.build.targets.wheel]
|
||||
packages = ["nvcurve"]
|
||||
|
||||
# The custom build hook (hatch_build.py) compiles the React frontend when
|
||||
# frontend/dist is missing or stale, so `uv tool install git+<repo-url>` works
|
||||
# as a single command. It runs for both wheel and sdist builds.
|
||||
|
||||
[tool.hatch.build.targets.wheel.force-include]
|
||||
"frontend/dist" = "nvcurve/frontend/dist"
|
||||
|
||||
|
||||
+227
-126
@@ -55,20 +55,20 @@ Key findings:
|
||||
See NvAPI_VF_Curve_Documentation.md for full technical details.
|
||||
"""
|
||||
|
||||
import argparse
|
||||
import ctypes
|
||||
import struct
|
||||
import sys
|
||||
import json
|
||||
import os
|
||||
import struct
|
||||
import sys
|
||||
import time
|
||||
import argparse
|
||||
from datetime import datetime
|
||||
from typing import Optional, List, Tuple, Set, Dict
|
||||
|
||||
# ═══════════════════════════════════════════════════════════════════════════
|
||||
# NvAPI bootstrap
|
||||
# ═══════════════════════════════════════════════════════════════════════════
|
||||
|
||||
|
||||
def load_nvapi():
|
||||
"""Load libnvidia-api.so from the NVIDIA driver."""
|
||||
for name in ("libnvidia-api.so", "libnvidia-api.so.1"):
|
||||
@@ -148,18 +148,15 @@ FUNC = {
|
||||
"Initialize": 0x0150E828,
|
||||
"EnumPhysicalGPUs": 0xE5AC921F,
|
||||
"GetFullName": 0xCEEE8E9F,
|
||||
|
||||
# V/F curve (read)
|
||||
"GetVFPCurve": 0x21537AD4, # ClkVfPointsGetStatus
|
||||
"GetClockBoostMask": 0x507B4B59, # ClkVfPointsGetInfo
|
||||
"GetClockBoostTable": 0x23F1B133, # ClkVfPointsGetControl
|
||||
"GetCurrentVoltage": 0x465F9BCF, # ClientVoltRailsGetStatus
|
||||
"GetClockBoostRanges": 0x64B43A6A, # ClkDomainsGetInfo
|
||||
|
||||
# Additional read
|
||||
"GetPerfLimits": 0xE440B867, # PerfClientLimitsGetStatus
|
||||
"GetVoltBoostPercent": 0x9DF23CA1, # ClientVoltRailsGetControl
|
||||
|
||||
# Write
|
||||
"SetClockBoostTable": 0x0733E009, # ClkVfPointsSetControl
|
||||
}
|
||||
@@ -205,6 +202,7 @@ SNAPSHOT_DIR = os.path.expanduser("~/.cache/nv_vfcurve")
|
||||
# GPU initialization
|
||||
# ═══════════════════════════════════════════════════════════════════════════
|
||||
|
||||
|
||||
def init_gpu() -> tuple:
|
||||
"""Initialize NvAPI, enumerate GPUs, return (handle, name)."""
|
||||
init_fn = nvfunc(FUNC["Initialize"], 0)
|
||||
@@ -236,12 +234,14 @@ def init_gpu() -> tuple:
|
||||
# also distinguishes GPU core vs memory clock domains.
|
||||
# ═══════════════════════════════════════════════════════════════════════════
|
||||
|
||||
|
||||
class BoostMask:
|
||||
"""Parsed GetClockBoostMask data.
|
||||
|
||||
Provides the raw mask bytes for copying into other calls, plus
|
||||
parsed per-entry enabled info for filtering.
|
||||
"""
|
||||
|
||||
def __init__(self, raw: bytes):
|
||||
self.raw = raw
|
||||
self.size = len(raw)
|
||||
@@ -260,7 +260,7 @@ class BoostMask:
|
||||
enabled = bool(self.mask_bytes[byte_idx] & (1 << bit_idx))
|
||||
self.entries.append({"index": i, "enabled": enabled})
|
||||
|
||||
def get_enabled_indices(self) -> List[int]:
|
||||
def get_enabled_indices(self) -> list[int]:
|
||||
"""Return list of point indices that are enabled in the mask."""
|
||||
return [e["index"] for e in self.entries if e["enabled"]]
|
||||
|
||||
@@ -273,12 +273,13 @@ class BoostMask:
|
||||
buf[offset + i] = self.mask_bytes[i]
|
||||
|
||||
|
||||
def read_boost_mask(gpu) -> Tuple[Optional[BoostMask], str]:
|
||||
def read_boost_mask(gpu) -> tuple[BoostMask | None, str]:
|
||||
"""Read the clock boost mask — the canonical source of active point info.
|
||||
|
||||
Per nvapioc, this mask must be copied into VFP and ClockBoostTable calls.
|
||||
Using all-0xFF works on some GPUs (Blackwell) but fails on others (Pascal).
|
||||
"""
|
||||
|
||||
def fill(buf):
|
||||
for i in range(MASK_OFFSET, MASK_OFFSET + MASK_BYTES):
|
||||
buf[i] = 0xFF
|
||||
@@ -294,20 +295,22 @@ def read_boost_mask(gpu) -> Tuple[Optional[BoostMask], str]:
|
||||
# Point classification — GPU core vs memory
|
||||
# ═══════════════════════════════════════════════════════════════════════════
|
||||
|
||||
|
||||
class CurveInfo:
|
||||
"""Holds classified point information for the GPU's V/F curve.
|
||||
|
||||
Combines data from GetClockBoostMask, GetVFPCurve, and GetClockBoostTable
|
||||
to determine which points are GPU core and which are memory.
|
||||
"""
|
||||
|
||||
def __init__(self):
|
||||
self.gpu_points: List[int] = [] # GPU core V/F point indices
|
||||
self.mem_points: List[int] = [] # Memory V/F point indices
|
||||
self.gpu_points: list[int] = [] # GPU core V/F point indices
|
||||
self.mem_points: list[int] = [] # Memory V/F point indices
|
||||
self.total_points: int = 0 # Total populated entries
|
||||
self.mask: Optional[BoostMask] = None
|
||||
self.mask: BoostMask | None = None
|
||||
|
||||
@staticmethod
|
||||
def build(gpu, mask: Optional[BoostMask] = None) -> 'CurveInfo':
|
||||
def build(gpu, mask: BoostMask | None = None) -> "CurveInfo":
|
||||
"""Classify all points by reading CT field_00 and VFP data.
|
||||
|
||||
field_00 == 0: GPU core data point
|
||||
@@ -340,7 +343,7 @@ class CurveInfo:
|
||||
has_vfp_data = False
|
||||
if vfp_points and i < len(vfp_points):
|
||||
f, v = vfp_points[i]
|
||||
has_vfp_data = (f > 0 or v > 0)
|
||||
has_vfp_data = f > 0 or v > 0
|
||||
|
||||
has_ct_data = False
|
||||
for j in range(9):
|
||||
@@ -378,13 +381,15 @@ class CurveInfo:
|
||||
# Data readers (mask-aware)
|
||||
# ═══════════════════════════════════════════════════════════════════════════
|
||||
|
||||
|
||||
def _fill_mask_from_boost(buf, mask: BoostMask):
|
||||
"""Copy boost mask into buffer."""
|
||||
mask.copy_mask_into(buf)
|
||||
|
||||
|
||||
def _read_vfp_with_mask(gpu, mask: Optional[BoostMask]) -> Optional[List[Tuple[int, int]]]:
|
||||
def _read_vfp_with_mask(gpu, mask: BoostMask | None) -> list[tuple[int, int]] | None:
|
||||
"""Read VFP curve using the canonical boost mask."""
|
||||
|
||||
def fill(buf):
|
||||
_fill_mask_from_boost(buf, mask)
|
||||
|
||||
@@ -403,8 +408,9 @@ def _read_vfp_with_mask(gpu, mask: Optional[BoostMask]) -> Optional[List[Tuple[i
|
||||
return points
|
||||
|
||||
|
||||
def _read_clock_table_raw_with_mask(gpu, mask: Optional[BoostMask]) -> Optional[bytes]:
|
||||
def _read_clock_table_raw_with_mask(gpu, mask: BoostMask | None) -> bytes | None:
|
||||
"""Read raw ClockBoostTable using the canonical boost mask."""
|
||||
|
||||
def fill(buf):
|
||||
_fill_mask_from_boost(buf, mask)
|
||||
|
||||
@@ -412,13 +418,14 @@ def _read_clock_table_raw_with_mask(gpu, mask: Optional[BoostMask]) -> Optional[
|
||||
return d if d else None
|
||||
|
||||
|
||||
def read_vfp_curve(gpu, mask: Optional[BoostMask] = None,
|
||||
curve_info: Optional[CurveInfo] = None
|
||||
) -> Tuple[Optional[List[Tuple[int, int]]], str]:
|
||||
def read_vfp_curve(
|
||||
gpu, mask: BoostMask | None = None, curve_info: CurveInfo | None = None
|
||||
) -> tuple[list[tuple[int, int]] | None, str]:
|
||||
"""Read V/F curve (frequency + voltage pairs).
|
||||
|
||||
Returns up to 255 entries. Use curve_info to determine which are GPU/mem.
|
||||
"""
|
||||
|
||||
def fill(buf):
|
||||
_fill_mask_from_boost(buf, mask)
|
||||
|
||||
@@ -442,18 +449,20 @@ def read_vfp_curve(gpu, mask: Optional[BoostMask] = None,
|
||||
return points, "OK"
|
||||
|
||||
|
||||
def read_clock_table_raw(gpu, mask: Optional[BoostMask] = None
|
||||
) -> Tuple[Optional[bytes], str]:
|
||||
def read_clock_table_raw(
|
||||
gpu, mask: BoostMask | None = None
|
||||
) -> tuple[bytes | None, str]:
|
||||
"""Read the raw ClockBoostTable buffer."""
|
||||
|
||||
def fill(buf):
|
||||
_fill_mask_from_boost(buf, mask)
|
||||
|
||||
return nvcall(FUNC["GetClockBoostTable"], gpu, CT_SIZE, ver=1, pre_fill=fill)
|
||||
|
||||
|
||||
def read_clock_offsets(gpu, mask: Optional[BoostMask] = None,
|
||||
curve_info: Optional[CurveInfo] = None
|
||||
) -> Tuple[Optional[List[int]], str]:
|
||||
def read_clock_offsets(
|
||||
gpu, mask: BoostMask | None = None, curve_info: CurveInfo | None = None
|
||||
) -> tuple[list[int] | None, str]:
|
||||
"""Read per-point frequency offsets from the ClockBoostTable."""
|
||||
d, err = read_clock_table_raw(gpu, mask)
|
||||
if not d:
|
||||
@@ -482,14 +491,18 @@ def read_clock_entry_full(data: bytes, point: int) -> dict:
|
||||
for j in range(9):
|
||||
off = base + j * 4
|
||||
if j == 5:
|
||||
fields[f"field_{j:02d}_0x{j*4:02X}"] = struct.unpack_from("<i", data, off)[0]
|
||||
fields[f"field_{j:02d}_0x{j * 4:02X}"] = struct.unpack_from(
|
||||
"<i", data, off
|
||||
)[0]
|
||||
else:
|
||||
fields[f"field_{j:02d}_0x{j*4:02X}"] = struct.unpack_from("<I", data, off)[0]
|
||||
fields[f"field_{j:02d}_0x{j * 4:02X}"] = struct.unpack_from(
|
||||
"<I", data, off
|
||||
)[0]
|
||||
fields["freqDelta_kHz"] = fields["field_05_0x14"]
|
||||
return fields
|
||||
|
||||
|
||||
def read_voltage(gpu) -> Tuple[Optional[int], str]:
|
||||
def read_voltage(gpu) -> tuple[int | None, str]:
|
||||
"""Read current GPU core voltage in µV."""
|
||||
d, err = nvcall(FUNC["GetCurrentVoltage"], gpu, VOLT_SIZE, ver=1)
|
||||
if not d:
|
||||
@@ -497,7 +510,7 @@ def read_voltage(gpu) -> Tuple[Optional[int], str]:
|
||||
return struct.unpack_from("<I", d, 0x28)[0], "OK"
|
||||
|
||||
|
||||
def read_clock_ranges(gpu) -> Tuple[Optional[dict], str]:
|
||||
def read_clock_ranges(gpu) -> tuple[dict | None, str]:
|
||||
"""Read clock domain min/max offset ranges."""
|
||||
d, err = nvcall(FUNC["GetClockBoostRanges"], gpu, RANGES_SIZE, ver=1)
|
||||
if not d:
|
||||
@@ -508,8 +521,7 @@ def read_clock_ranges(gpu) -> Tuple[Optional[dict], str]:
|
||||
base = 0x08 + i * 0x48
|
||||
if base + 0x48 > len(d):
|
||||
break
|
||||
words = [struct.unpack_from("<i", d, base + j)[0]
|
||||
for j in range(0, 0x48, 4)]
|
||||
words = [struct.unpack_from("<i", d, base + j)[0] for j in range(0, 0x48, 4)]
|
||||
domains.append(words)
|
||||
return {"num_domains": num, "domains": domains}, "OK"
|
||||
|
||||
@@ -518,14 +530,17 @@ def read_clock_ranges(gpu) -> Tuple[Optional[dict], str]:
|
||||
# Mask bit helpers
|
||||
# ═══════════════════════════════════════════════════════════════════════════
|
||||
|
||||
|
||||
def set_mask_bit(buf, point: int, offset=MASK_OFFSET):
|
||||
"""Set a single bit in the mask field."""
|
||||
byte_idx = offset + (point // 8)
|
||||
bit_idx = point % 8
|
||||
buf[byte_idx] = int.from_bytes(buf[byte_idx:byte_idx+1], 'little') | (1 << bit_idx)
|
||||
buf[byte_idx] = int.from_bytes(buf[byte_idx : byte_idx + 1], "little") | (
|
||||
1 << bit_idx
|
||||
)
|
||||
|
||||
|
||||
def set_mask_bits(buf, points: Set[int], offset=MASK_OFFSET):
|
||||
def set_mask_bits(buf, points: set[int], offset=MASK_OFFSET):
|
||||
"""Set mask bits for a set of points."""
|
||||
for p in points:
|
||||
set_mask_bit(buf, p, offset)
|
||||
@@ -535,11 +550,12 @@ def set_mask_bits(buf, points: Set[int], offset=MASK_OFFSET):
|
||||
# Write operations
|
||||
# ═══════════════════════════════════════════════════════════════════════════
|
||||
|
||||
|
||||
def build_write_buffer(
|
||||
gpu,
|
||||
point_deltas: dict,
|
||||
mask: Optional[BoostMask] = None,
|
||||
) -> Tuple[Optional[ctypes.Array], str]:
|
||||
mask: BoostMask | None = None,
|
||||
) -> tuple[ctypes.Array | None, str]:
|
||||
"""Build a SetClockBoostTable buffer with specified per-point deltas.
|
||||
|
||||
Strategy: read the current ClockBoostTable (using canonical mask),
|
||||
@@ -576,9 +592,9 @@ def build_write_buffer(
|
||||
def write_clock_offsets(
|
||||
gpu,
|
||||
point_deltas: dict,
|
||||
mask: Optional[BoostMask] = None,
|
||||
mask: BoostMask | None = None,
|
||||
dry_run: bool = False,
|
||||
) -> Tuple[int, str]:
|
||||
) -> tuple[int, str]:
|
||||
"""Write per-point frequency offsets via SetClockBoostTable."""
|
||||
buf, err = build_write_buffer(gpu, point_deltas, mask)
|
||||
if buf is None:
|
||||
@@ -595,9 +611,10 @@ def write_clock_offsets(
|
||||
# Safety checks
|
||||
# ═══════════════════════════════════════════════════════════════════════════
|
||||
|
||||
def validate_write_request(point_deltas: dict,
|
||||
curve_info: Optional[CurveInfo] = None
|
||||
) -> Optional[str]:
|
||||
|
||||
def validate_write_request(
|
||||
point_deltas: dict, curve_info: CurveInfo | None = None
|
||||
) -> str | None:
|
||||
"""Return an error message if the write request is unsafe, else None."""
|
||||
mem_points = set()
|
||||
if curve_info:
|
||||
@@ -608,14 +625,18 @@ def validate_write_request(point_deltas: dict,
|
||||
return f"Point {point} out of range (0–{CT_MAX_ENTRIES - 1})"
|
||||
|
||||
if point in mem_points:
|
||||
return (f"Point {point} is a memory clock entry. "
|
||||
return (
|
||||
f"Point {point} is a memory clock entry. "
|
||||
"Memory offsets use a different mechanism (NVML). "
|
||||
"Use --force if you really mean it.")
|
||||
"Use --force if you really mean it."
|
||||
)
|
||||
|
||||
if abs(delta_khz) > MAX_DELTA_KHZ:
|
||||
return (f"Delta {delta_khz/1000:+.0f} MHz for point {point} exceeds "
|
||||
return (
|
||||
f"Delta {delta_khz / 1000:+.0f} MHz for point {point} exceeds "
|
||||
f"safety limit of ±{MAX_DELTA_KHZ / 1000:.0f} MHz. "
|
||||
"Use --max-delta to raise the limit if needed.")
|
||||
"Use --max-delta to raise the limit if needed."
|
||||
)
|
||||
|
||||
return None
|
||||
|
||||
@@ -624,6 +645,7 @@ def validate_write_request(point_deltas: dict,
|
||||
# Hex dump utility
|
||||
# ═══════════════════════════════════════════════════════════════════════════
|
||||
|
||||
|
||||
def hexdump(data: bytes, start: int, length: int, cols: int = 16) -> str:
|
||||
lines = []
|
||||
end = min(start + length, len(data))
|
||||
@@ -639,7 +661,8 @@ def hexdump(data: bytes, start: int, length: int, cols: int = 16) -> str:
|
||||
# Snapshot save/restore
|
||||
# ═══════════════════════════════════════════════════════════════════════════
|
||||
|
||||
def snapshot_save(gpu, gpu_name: str, mask: Optional[BoostMask] = None):
|
||||
|
||||
def snapshot_save(gpu, gpu_name: str, mask: BoostMask | None = None):
|
||||
"""Save the current ClockBoostTable to disk."""
|
||||
raw, err = read_clock_table_raw(gpu, mask)
|
||||
if not raw:
|
||||
@@ -674,7 +697,7 @@ def snapshot_save(gpu, gpu_name: str, mask: Optional[BoostMask] = None):
|
||||
with open(meta_fname, "w") as f:
|
||||
json.dump(meta, f, indent=2)
|
||||
|
||||
print(f"Snapshot saved:")
|
||||
print("Snapshot saved:")
|
||||
print(f" Binary: {fname}")
|
||||
print(f" Metadata: {meta_fname}")
|
||||
print(f" Size: {len(raw)} bytes")
|
||||
@@ -682,7 +705,7 @@ def snapshot_save(gpu, gpu_name: str, mask: Optional[BoostMask] = None):
|
||||
return True
|
||||
|
||||
|
||||
def snapshot_restore(gpu, mask: Optional[BoostMask] = None, filepath: str = None):
|
||||
def snapshot_restore(gpu, mask: BoostMask | None = None, filepath: str = None):
|
||||
"""Restore a ClockBoostTable snapshot from disk."""
|
||||
if filepath is None:
|
||||
if not os.path.isdir(SNAPSHOT_DIR):
|
||||
@@ -730,7 +753,8 @@ def snapshot_restore(gpu, mask: Optional[BoostMask] = None, filepath: str = None
|
||||
# Diagnostics
|
||||
# ═══════════════════════════════════════════════════════════════════════════
|
||||
|
||||
def run_diagnostics(gpu, gpu_name, mask: Optional[BoostMask] = None):
|
||||
|
||||
def run_diagnostics(gpu, gpu_name, mask: BoostMask | None = None):
|
||||
"""Probe all known functions and report results."""
|
||||
print(f"GPU: {gpu_name}")
|
||||
print()
|
||||
@@ -768,7 +792,9 @@ def run_diagnostics(gpu, gpu_name, mask: Optional[BoostMask] = None):
|
||||
|
||||
# Step 3: test reads with the proper mask
|
||||
needs_mask_fns = {
|
||||
FUNC["GetVFPCurve"], FUNC["GetClockBoostMask"], FUNC["GetClockBoostTable"]
|
||||
FUNC["GetVFPCurve"],
|
||||
FUNC["GetClockBoostMask"],
|
||||
FUNC["GetClockBoostTable"],
|
||||
}
|
||||
|
||||
print()
|
||||
@@ -817,8 +843,10 @@ def run_diagnostics(gpu, gpu_name, mask: Optional[BoostMask] = None):
|
||||
# Output formatting
|
||||
# ═══════════════════════════════════════════════════════════════════════════
|
||||
|
||||
def print_curve(points, offsets, voltage, curve_info: Optional[CurveInfo] = None,
|
||||
full=False):
|
||||
|
||||
def print_curve(
|
||||
points, offsets, voltage, curve_info: CurveInfo | None = None, full=False
|
||||
):
|
||||
"""Print formatted V/F curve table."""
|
||||
if voltage:
|
||||
print(f"Current voltage: {voltage / 1000:.1f} mV")
|
||||
@@ -845,9 +873,7 @@ def print_curve(points, offsets, voltage, curve_info: Optional[CurveInfo] = None
|
||||
for i, (f, v) in enumerate(points):
|
||||
if f == 0 and v == 0:
|
||||
continue
|
||||
if i in mem_set:
|
||||
show.append(i)
|
||||
elif f != prev_freq or i == len(points) - 1:
|
||||
if i in mem_set or f != prev_freq or i == len(points) - 1:
|
||||
show.append(i)
|
||||
prev_freq = f
|
||||
|
||||
@@ -882,52 +908,78 @@ def print_curve(points, offsets, voltage, curve_info: Optional[CurveInfo] = None
|
||||
|
||||
# Summary
|
||||
if curve_info and curve_info.gpu_points:
|
||||
gpu_data = [(points[i][0], points[i][1]) for i in curve_info.gpu_points
|
||||
if i < len(points) and points[i][0] > 0]
|
||||
gpu_data = [
|
||||
(points[i][0], points[i][1])
|
||||
for i in curve_info.gpu_points
|
||||
if i < len(points) and points[i][0] > 0
|
||||
]
|
||||
if gpu_data:
|
||||
freqs = [f for f, v in gpu_data]
|
||||
volts = [v for f, v in gpu_data]
|
||||
print()
|
||||
print(f"GPU core: {min(freqs)/1000:.0f} – {max(freqs)/1000:.0f} MHz, "
|
||||
print(
|
||||
f"GPU core: {min(freqs) / 1000:.0f} – {max(freqs) / 1000:.0f} MHz, "
|
||||
f"{min(volts) / 1000:.0f} – {max(volts) / 1000:.0f} mV "
|
||||
f"({len(gpu_data)} points)")
|
||||
f"({len(gpu_data)} points)"
|
||||
)
|
||||
|
||||
if curve_info and curve_info.mem_points:
|
||||
mem_data = [(points[i][0], points[i][1]) for i in curve_info.mem_points
|
||||
if i < len(points) and points[i][0] > 0]
|
||||
mem_data = [
|
||||
(points[i][0], points[i][1])
|
||||
for i in curve_info.mem_points
|
||||
if i < len(points) and points[i][0] > 0
|
||||
]
|
||||
if mem_data:
|
||||
freqs = [f for f, v in mem_data]
|
||||
volts = [v for f, v in mem_data]
|
||||
print(f"Memory: {min(freqs)/1000:.0f} – {max(freqs)/1000:.0f} MHz, "
|
||||
print(
|
||||
f"Memory: {min(freqs) / 1000:.0f} – {max(freqs) / 1000:.0f} MHz, "
|
||||
f"{min(volts) / 1000:.0f} – {max(volts) / 1000:.0f} mV "
|
||||
f"({len(mem_data)} points)")
|
||||
f"({len(mem_data)} points)"
|
||||
)
|
||||
|
||||
if offsets:
|
||||
gpu_indices = set(curve_info.gpu_points) if curve_info else set(range(len(offsets)))
|
||||
gpu_offsets = [offsets[i] for i in gpu_indices
|
||||
if i < len(offsets) and offsets[i] != 0]
|
||||
gpu_indices = (
|
||||
set(curve_info.gpu_points) if curve_info else set(range(len(offsets)))
|
||||
)
|
||||
gpu_offsets = [
|
||||
offsets[i] for i in gpu_indices if i < len(offsets) and offsets[i] != 0
|
||||
]
|
||||
if gpu_offsets:
|
||||
vals = set(gpu_offsets)
|
||||
if len(vals) == 1:
|
||||
print(f"GPU offset: {next(iter(vals))/1000:+.0f} MHz "
|
||||
f"(uniform across {len(gpu_offsets)} points)")
|
||||
print(
|
||||
f"GPU offset: {next(iter(vals)) / 1000:+.0f} MHz "
|
||||
f"(uniform across {len(gpu_offsets)} points)"
|
||||
)
|
||||
else:
|
||||
print(f"GPU offsets: {len(gpu_offsets)} points active "
|
||||
f"(range: {min(vals)/1000:+.0f} to {max(vals)/1000:+.0f} MHz)")
|
||||
print(
|
||||
f"GPU offsets: {len(gpu_offsets)} points active "
|
||||
f"(range: {min(vals) / 1000:+.0f} to {max(vals) / 1000:+.0f} MHz)"
|
||||
)
|
||||
|
||||
|
||||
def output_json(gpu_name, points, offsets, voltage,
|
||||
curve_info: Optional[CurveInfo] = None):
|
||||
def output_json(
|
||||
gpu_name, points, offsets, voltage, curve_info: CurveInfo | None = None
|
||||
):
|
||||
"""Output JSON format."""
|
||||
data = {
|
||||
"gpu": gpu_name,
|
||||
"current_voltage_uV": voltage,
|
||||
"layout": {
|
||||
"vfp_curve": {"size": VFP_SIZE, "base": VFP_BASE,
|
||||
"stride": VFP_STRIDE, "max_entries": VFP_MAX_ENTRIES},
|
||||
"clock_table": {"size": CT_SIZE, "base": CT_BASE,
|
||||
"stride": CT_STRIDE, "delta_offset": CT_DELTA_OFF,
|
||||
"max_entries": CT_MAX_ENTRIES},
|
||||
"vfp_curve": {
|
||||
"size": VFP_SIZE,
|
||||
"base": VFP_BASE,
|
||||
"stride": VFP_STRIDE,
|
||||
"max_entries": VFP_MAX_ENTRIES,
|
||||
},
|
||||
"clock_table": {
|
||||
"size": CT_SIZE,
|
||||
"base": CT_BASE,
|
||||
"stride": CT_STRIDE,
|
||||
"delta_offset": CT_DELTA_OFF,
|
||||
"max_entries": CT_MAX_ENTRIES,
|
||||
},
|
||||
},
|
||||
"curve_info": {
|
||||
"gpu_points": curve_info.gpu_points if curve_info else [],
|
||||
@@ -958,6 +1010,7 @@ def output_json(gpu_name, points, offsets, voltage,
|
||||
# Write command handler
|
||||
# ═══════════════════════════════════════════════════════════════════════════
|
||||
|
||||
|
||||
def cmd_write(gpu, gpu_name, args, mask, curve_info):
|
||||
"""Handle write subcommand."""
|
||||
delta_khz = int(args.delta * 1000)
|
||||
@@ -977,21 +1030,27 @@ def cmd_write(gpu, gpu_name, args, mask, curve_info):
|
||||
|
||||
elif args.point is not None:
|
||||
point_deltas[args.point] = delta_khz
|
||||
print(f"Target: point {args.point}, delta {args.delta:+.0f} MHz "
|
||||
f"({delta_khz:+d} kHz)")
|
||||
print(
|
||||
f"Target: point {args.point}, delta {args.delta:+.0f} MHz "
|
||||
f"({delta_khz:+d} kHz)"
|
||||
)
|
||||
|
||||
elif args.range:
|
||||
start, end = args.range
|
||||
for i in range(start, end + 1):
|
||||
point_deltas[i] = delta_khz
|
||||
print(f"Target: points {start}–{end} ({len(point_deltas)} points), "
|
||||
f"delta {args.delta:+.0f} MHz")
|
||||
print(
|
||||
f"Target: points {start}–{end} ({len(point_deltas)} points), "
|
||||
f"delta {args.delta:+.0f} MHz"
|
||||
)
|
||||
|
||||
elif args.glob:
|
||||
for i in gpu_points:
|
||||
point_deltas[i] = delta_khz
|
||||
print(f"Target: all {len(point_deltas)} GPU core points, "
|
||||
f"delta {args.delta:+.0f} MHz")
|
||||
print(
|
||||
f"Target: all {len(point_deltas)} GPU core points, "
|
||||
f"delta {args.delta:+.0f} MHz"
|
||||
)
|
||||
|
||||
else:
|
||||
print("Error: specify --point N, --range A-B, --global, or --reset")
|
||||
@@ -1013,11 +1072,17 @@ def cmd_write(gpu, gpu_name, args, mask, curve_info):
|
||||
changed = 0
|
||||
for point in sorted(point_deltas.keys()):
|
||||
new = point_deltas[point]
|
||||
old = current_offsets[point] if current_offsets and point < len(current_offsets) else 0
|
||||
old = (
|
||||
current_offsets[point]
|
||||
if current_offsets and point < len(current_offsets)
|
||||
else 0
|
||||
)
|
||||
if old != new:
|
||||
changed += 1
|
||||
if changed <= 20:
|
||||
print(f" Point {point:3d}: {old/1000:+8.0f} MHz → {new/1000:+8.0f} MHz")
|
||||
print(
|
||||
f" Point {point:3d}: {old / 1000:+8.0f} MHz → {new / 1000:+8.0f} MHz"
|
||||
)
|
||||
if changed > 20:
|
||||
print(f" ... and {changed - 20} more points")
|
||||
if changed == 0:
|
||||
@@ -1036,8 +1101,10 @@ def cmd_write(gpu, gpu_name, args, mask, curve_info):
|
||||
|
||||
first_pt = min(point_deltas.keys())
|
||||
entry_off = CT_BASE + first_pt * CT_STRIDE
|
||||
print(f"\nEntry for point {first_pt} (offset 0x{entry_off:04X}, "
|
||||
f"stride 0x{CT_STRIDE:02X}):")
|
||||
print(
|
||||
f"\nEntry for point {first_pt} (offset 0x{entry_off:04X}, "
|
||||
f"stride 0x{CT_STRIDE:02X}):"
|
||||
)
|
||||
print(hexdump(bytes(buf), entry_off, CT_STRIDE))
|
||||
return
|
||||
|
||||
@@ -1070,8 +1137,10 @@ def cmd_write(gpu, gpu_name, args, mask, curve_info):
|
||||
actual = new_offsets[point] if point < len(new_offsets) else 0
|
||||
if actual != expected:
|
||||
mismatches += 1
|
||||
print(f" MISMATCH point {point}: expected {expected/1000:+.0f} MHz, "
|
||||
f"got {actual/1000:+.0f} MHz")
|
||||
print(
|
||||
f" MISMATCH point {point}: expected {expected / 1000:+.0f} MHz, "
|
||||
f"got {actual / 1000:+.0f} MHz"
|
||||
)
|
||||
|
||||
if mismatches == 0:
|
||||
print(f"Verified: all {len(point_deltas)} points match expected values.")
|
||||
@@ -1083,6 +1152,7 @@ def cmd_write(gpu, gpu_name, args, mask, curve_info):
|
||||
# Verify command handler
|
||||
# ═══════════════════════════════════════════════════════════════════════════
|
||||
|
||||
|
||||
def cmd_verify(gpu, gpu_name, args, mask, curve_info):
|
||||
"""Write-verify-read cycle for a single point or range."""
|
||||
delta_khz = int(args.delta * 1000)
|
||||
@@ -1095,14 +1165,14 @@ def cmd_verify(gpu, gpu_name, args, mask, curve_info):
|
||||
print("Error: --point or --range required for verify mode")
|
||||
return
|
||||
|
||||
point_deltas = {p: delta_khz for p in points}
|
||||
point_deltas = dict.fromkeys(points, delta_khz)
|
||||
|
||||
err = validate_write_request(point_deltas, curve_info)
|
||||
if err:
|
||||
print(f"Safety check FAILED: {err}")
|
||||
return
|
||||
|
||||
print(f"=== Write-Verify Cycle ===")
|
||||
print("=== Write-Verify Cycle ===")
|
||||
print(f"GPU: {gpu_name}")
|
||||
if curve_info:
|
||||
print(f"Curve: {curve_info.describe()}")
|
||||
@@ -1156,8 +1226,10 @@ def cmd_verify(gpu, gpu_name, args, mask, curve_info):
|
||||
match = "OK" if actual == expected else "MISMATCH"
|
||||
if actual != expected:
|
||||
all_ok = False
|
||||
print(f" Point {p:3d}: expected {expected/1000:+8.0f} MHz, "
|
||||
f"got {actual/1000:+8.0f} MHz [{match}]")
|
||||
print(
|
||||
f" Point {p:3d}: expected {expected / 1000:+8.0f} MHz, "
|
||||
f"got {actual / 1000:+8.0f} MHz [{match}]"
|
||||
)
|
||||
|
||||
# Step 5: Check for collateral damage
|
||||
print()
|
||||
@@ -1169,8 +1241,10 @@ def cmd_verify(gpu, gpu_name, args, mask, curve_info):
|
||||
continue
|
||||
if before_offsets[i] != after_offsets[i]:
|
||||
collateral += 1
|
||||
print(f" WARNING: Point {i} changed unexpectedly: "
|
||||
f"{before_offsets[i]/1000:+.0f} → {after_offsets[i]/1000:+.0f} MHz")
|
||||
print(
|
||||
f" WARNING: Point {i} changed unexpectedly: "
|
||||
f"{before_offsets[i] / 1000:+.0f} → {after_offsets[i] / 1000:+.0f} MHz"
|
||||
)
|
||||
if collateral == 0:
|
||||
print(" No unintended changes detected.")
|
||||
|
||||
@@ -1187,7 +1261,9 @@ def cmd_verify(gpu, gpu_name, args, mask, curve_info):
|
||||
continue
|
||||
if before_entry[key] != after_entry[key]:
|
||||
field_changes += 1
|
||||
print(f" Point {p}, {key}: {before_entry[key]} → {after_entry[key]}")
|
||||
print(
|
||||
f" Point {p}, {key}: {before_entry[key]} → {after_entry[key]}"
|
||||
)
|
||||
if field_changes == 0:
|
||||
print(" No unknown fields changed.")
|
||||
|
||||
@@ -1215,6 +1291,7 @@ def cmd_verify(gpu, gpu_name, args, mask, curve_info):
|
||||
# Inspect command
|
||||
# ═══════════════════════════════════════════════════════════════════════════
|
||||
|
||||
|
||||
def cmd_inspect(gpu, gpu_name, args, mask, curve_info):
|
||||
"""Show detailed field-level data for specific points."""
|
||||
raw, err = read_clock_table_raw(gpu, mask)
|
||||
@@ -1246,8 +1323,9 @@ def cmd_inspect(gpu, gpu_name, args, mask, curve_info):
|
||||
print(f"GPU: {gpu_name}")
|
||||
if curve_info:
|
||||
print(f"Curve: {curve_info.describe()}")
|
||||
print(f"ClockBoostTable entry detail (stride=0x{CT_STRIDE:02X}, "
|
||||
f"9 fields × 4 bytes)")
|
||||
print(
|
||||
f"ClockBoostTable entry detail (stride=0x{CT_STRIDE:02X}, 9 fields × 4 bytes)"
|
||||
)
|
||||
print()
|
||||
|
||||
for p in indices:
|
||||
@@ -1275,8 +1353,10 @@ def cmd_inspect(gpu, gpu_name, args, mask, curve_info):
|
||||
continue
|
||||
marker = " ← freqDelta" if "0x14" in key else ""
|
||||
if "0x14" in key:
|
||||
print(f" {key}: {val:12d} (0x{val & 0xFFFFFFFF:08X})"
|
||||
f" = {val/1000:+.0f} MHz{marker}")
|
||||
print(
|
||||
f" {key}: {val:12d} (0x{val & 0xFFFFFFFF:08X})"
|
||||
f" = {val / 1000:+.0f} MHz{marker}"
|
||||
)
|
||||
else:
|
||||
print(f" {key}: {val:12d} (0x{val:08X})")
|
||||
print()
|
||||
@@ -1286,6 +1366,7 @@ def cmd_inspect(gpu, gpu_name, args, mask, curve_info):
|
||||
# Read command handler
|
||||
# ═══════════════════════════════════════════════════════════════════════════
|
||||
|
||||
|
||||
def cmd_read(gpu, gpu_name, args, mask, curve_info):
|
||||
"""Handle read subcommand."""
|
||||
if args.diag:
|
||||
@@ -1309,10 +1390,13 @@ def cmd_read(gpu, gpu_name, args, mask, curve_info):
|
||||
print(f"GPU: {gpu_name}")
|
||||
|
||||
if args.raw:
|
||||
|
||||
def fill_vfp(buf):
|
||||
_fill_mask_from_boost(buf, mask)
|
||||
vfp_raw, _ = nvcall(FUNC["GetVFPCurve"], gpu, VFP_SIZE,
|
||||
ver=1, pre_fill=fill_vfp)
|
||||
|
||||
vfp_raw, _ = nvcall(
|
||||
FUNC["GetVFPCurve"], gpu, VFP_SIZE, ver=1, pre_fill=fill_vfp
|
||||
)
|
||||
ct_raw, _ = read_clock_table_raw(gpu, mask)
|
||||
|
||||
if vfp_raw:
|
||||
@@ -1348,7 +1432,8 @@ def cmd_read(gpu, gpu_name, args, mask, curve_info):
|
||||
# Argument parsing
|
||||
# ═══════════════════════════════════════════════════════════════════════════
|
||||
|
||||
def parse_range(s: str) -> Tuple[int, int]:
|
||||
|
||||
def parse_range(s: str) -> tuple[int, int]:
|
||||
"""Parse 'A-B' into (A, B) tuple."""
|
||||
parts = s.split("-")
|
||||
if len(parts) != 2:
|
||||
@@ -1360,7 +1445,9 @@ def parse_range(s: str) -> Tuple[int, int]:
|
||||
if a > b:
|
||||
raise argparse.ArgumentTypeError(f"Start > end in range: {a}-{b}")
|
||||
if a < 0 or b >= CT_MAX_ENTRIES:
|
||||
raise argparse.ArgumentTypeError(f"Range {a}-{b} outside 0–{CT_MAX_ENTRIES - 1}")
|
||||
raise argparse.ArgumentTypeError(
|
||||
f"Range {a}-{b} outside 0–{CT_MAX_ENTRIES - 1}"
|
||||
)
|
||||
return (a, b)
|
||||
|
||||
|
||||
@@ -1390,14 +1477,16 @@ Examples:
|
||||
|
||||
# --- read ---
|
||||
p_read = sub.add_parser("read", help="Read V/F curve (default)")
|
||||
p_read.add_argument("--full", action="store_true",
|
||||
help="Show all points including empty slots")
|
||||
p_read.add_argument("--json", action="store_true",
|
||||
help="JSON output with domain classification")
|
||||
p_read.add_argument("--raw", action="store_true",
|
||||
help="Include hex dumps")
|
||||
p_read.add_argument("--diag", action="store_true",
|
||||
help="Probe all functions with mask comparison")
|
||||
p_read.add_argument(
|
||||
"--full", action="store_true", help="Show all points including empty slots"
|
||||
)
|
||||
p_read.add_argument(
|
||||
"--json", action="store_true", help="JSON output with domain classification"
|
||||
)
|
||||
p_read.add_argument("--raw", action="store_true", help="Include hex dumps")
|
||||
p_read.add_argument(
|
||||
"--diag", action="store_true", help="Probe all functions with mask comparison"
|
||||
)
|
||||
|
||||
# --- inspect ---
|
||||
p_insp = sub.add_parser("inspect", help="Show detailed entry fields")
|
||||
@@ -1409,30 +1498,42 @@ Examples:
|
||||
tgt = p_write.add_mutually_exclusive_group()
|
||||
tgt.add_argument("--point", type=int, help="Single point index")
|
||||
tgt.add_argument("--range", type=parse_range, help="Point range A-B")
|
||||
tgt.add_argument("--global", dest="glob", action="store_true",
|
||||
help="All GPU core points")
|
||||
tgt.add_argument("--reset", action="store_true",
|
||||
help="Reset all GPU core offsets to 0")
|
||||
p_write.add_argument("--delta", type=float, default=0.0,
|
||||
help="Frequency offset in MHz (e.g. 15, -30)")
|
||||
p_write.add_argument("--dry-run", action="store_true",
|
||||
help="Preview changes without applying")
|
||||
p_write.add_argument("--force", action="store_true",
|
||||
help="Allow modifying memory points")
|
||||
p_write.add_argument("--max-delta", type=float, default=300.0,
|
||||
help="Override safety limit (MHz, default 300)")
|
||||
tgt.add_argument(
|
||||
"--global", dest="glob", action="store_true", help="All GPU core points"
|
||||
)
|
||||
tgt.add_argument(
|
||||
"--reset", action="store_true", help="Reset all GPU core offsets to 0"
|
||||
)
|
||||
p_write.add_argument(
|
||||
"--delta",
|
||||
type=float,
|
||||
default=0.0,
|
||||
help="Frequency offset in MHz (e.g. 15, -30)",
|
||||
)
|
||||
p_write.add_argument(
|
||||
"--dry-run", action="store_true", help="Preview changes without applying"
|
||||
)
|
||||
p_write.add_argument(
|
||||
"--force", action="store_true", help="Allow modifying memory points"
|
||||
)
|
||||
p_write.add_argument(
|
||||
"--max-delta",
|
||||
type=float,
|
||||
default=300.0,
|
||||
help="Override safety limit (MHz, default 300)",
|
||||
)
|
||||
|
||||
# --- verify ---
|
||||
p_ver = sub.add_parser("verify", help="Write-verify-read cycle")
|
||||
p_ver.add_argument("--point", type=int, help="Single point index")
|
||||
p_ver.add_argument("--range", type=parse_range, help="Point range A-B")
|
||||
p_ver.add_argument("--delta", type=float, required=True,
|
||||
help="Frequency offset in MHz")
|
||||
p_ver.add_argument(
|
||||
"--delta", type=float, required=True, help="Frequency offset in MHz"
|
||||
)
|
||||
|
||||
# --- snapshot ---
|
||||
p_snap = sub.add_parser("snapshot", help="Save/restore ClockBoostTable")
|
||||
p_snap.add_argument("action", choices=["save", "restore"],
|
||||
help="save or restore")
|
||||
p_snap.add_argument("action", choices=["save", "restore"], help="save or restore")
|
||||
p_snap.add_argument("--file", help="Snapshot file path (for restore)")
|
||||
|
||||
args = parser.parse_args()
|
||||
|
||||
@@ -0,0 +1,355 @@
|
||||
"""Unit tests for the RM power-limit interface (fake RM, no hardware).
|
||||
|
||||
Standalone (no pytest required):
|
||||
|
||||
python tests/test_rm_power.py
|
||||
|
||||
Also works under pytest if available. Ports the test battery from LACT PR
|
||||
#1205 (ilya-zlobintsev/LACT): layout discovery, NVML cross-validation,
|
||||
write minimality, readback verification, and failure restoration.
|
||||
"""
|
||||
|
||||
import os
|
||||
import sys
|
||||
|
||||
sys.path.insert(0, os.path.join(os.path.dirname(__file__), ".."))
|
||||
|
||||
from nvcurve.hal.rm_power import ( # noqa: E402
|
||||
_CTRL_GPU_GET_ATTACHED_IDS,
|
||||
_CTRL_GPU_GET_ID_INFO_V2,
|
||||
_CTRL_GPU_GET_PCI_INFO,
|
||||
_PWR_GET_CONTROL,
|
||||
_PWR_GET_INFO,
|
||||
_PWR_SET_CONTROL,
|
||||
EXTENDED_LAYOUT,
|
||||
LEGACY_LAYOUT,
|
||||
PciLocation,
|
||||
PowerLimitBounds,
|
||||
RmPowerError,
|
||||
_u32,
|
||||
probe,
|
||||
resolve_gpu_instance,
|
||||
set_limit,
|
||||
)
|
||||
|
||||
PASS = 0
|
||||
FAIL = 0
|
||||
|
||||
|
||||
def check(name: str, cond: bool) -> None:
|
||||
global PASS, FAIL
|
||||
if cond:
|
||||
PASS += 1
|
||||
print(f" PASS {name}")
|
||||
else:
|
||||
FAIL += 1
|
||||
print(f" FAIL {name}")
|
||||
|
||||
|
||||
BOUNDS = PowerLimitBounds(min_mw=250_000, default_mw=300_000, max_mw=325_000)
|
||||
|
||||
|
||||
class FakeRm:
|
||||
"""In-memory fake of the RM power-limit client (both wire layouts)."""
|
||||
|
||||
def __init__(self, layout, current: int) -> None:
|
||||
self.layout = layout
|
||||
self.control = bytearray(layout.control_size)
|
||||
self.control[0:8] = bytes([0xFF, 0, 0, 0, 1, 0, 0, 0])
|
||||
self.control[layout.request_at - 4 : layout.request_at] = bytes(
|
||||
[0x67, 0x67, 0, 0]
|
||||
)
|
||||
self.control[layout.request_at : layout.request_at + 4] = current.to_bytes(
|
||||
4, "little"
|
||||
)
|
||||
self.control[layout.client_at] = 0xFE
|
||||
self.reads: list[tuple[int, int]] = []
|
||||
self.writes: list[bytes] = []
|
||||
self.fail_first_write = False
|
||||
self.fail_readback = False
|
||||
self.fail_restore = False
|
||||
|
||||
def query(self, cmd: int, data: bytearray) -> None:
|
||||
if cmd == _PWR_GET_INFO:
|
||||
self.reads.append((cmd, len(data)))
|
||||
if len(data) != self.layout.info_size:
|
||||
raise RmPowerError("Unsupported INFO size")
|
||||
data[0:8] = bytes([0xFF, 0, 0, 0, 1, 0, 0, 0])
|
||||
for index, value in enumerate([250_000, 300_000, 325_000]):
|
||||
offset = self.layout.info_min_at + 4 * index
|
||||
data[offset : offset + 4] = value.to_bytes(4, "little")
|
||||
elif cmd == _PWR_GET_CONTROL:
|
||||
self.reads.append((cmd, len(data)))
|
||||
if len(data) != self.layout.control_size:
|
||||
raise RmPowerError("Unsupported CONTROL size")
|
||||
if data[self.layout.client_at] != 0xFE:
|
||||
raise AssertionError("unexpected client selector on GET")
|
||||
if self.fail_readback and len(self.writes) == 1:
|
||||
raise RmPowerError("readback unavailable")
|
||||
data[:] = self.control
|
||||
elif cmd == _PWR_SET_CONTROL:
|
||||
if len(data) != self.layout.control_size:
|
||||
raise AssertionError("bad SET size")
|
||||
if data[4:8] != (1).to_bytes(4, "little"):
|
||||
raise AssertionError("bad SET mask")
|
||||
if data[self.layout.client_at] != 0xFE:
|
||||
raise AssertionError("bad SET client selector")
|
||||
# Only the request field may differ from the current state.
|
||||
for i, (a, b) in enumerate(zip(data, self.control, strict=True)):
|
||||
if self.layout.request_at <= i < self.layout.request_at + 4:
|
||||
continue
|
||||
if a != b:
|
||||
raise AssertionError(f"SET modified byte {i:#x}")
|
||||
self.writes.append(bytes(data))
|
||||
if self.fail_restore and len(self.writes) > 1:
|
||||
raise RmPowerError("restore unavailable")
|
||||
self.control[:] = data
|
||||
if self.fail_first_write and len(self.writes) == 1:
|
||||
raise RmPowerError("SET failed after modifying hardware")
|
||||
else:
|
||||
raise AssertionError(f"unexpected command {cmd:#x}")
|
||||
|
||||
|
||||
def test_detects_both_layouts_with_gets_without_a_driver_version() -> None:
|
||||
for layout in (EXTENDED_LAYOUT, LEGACY_LAYOUT):
|
||||
rm = FakeRm(layout, 250_000)
|
||||
support = probe(BOUNDS, 250_000, rm.query)
|
||||
check(
|
||||
f"{layout.name}: detected",
|
||||
support.bounds == BOUNDS and support.layout == layout,
|
||||
)
|
||||
expected = (
|
||||
[(_PWR_GET_INFO, 0x924), (_PWR_GET_CONTROL, 0x328)]
|
||||
if layout == EXTENDED_LAYOUT
|
||||
else [
|
||||
(_PWR_GET_INFO, 0x924),
|
||||
(_PWR_GET_INFO, 0x488),
|
||||
(_PWR_GET_CONTROL, 0x188),
|
||||
]
|
||||
)
|
||||
check(f"{layout.name}: GETs only, expected sequence", rm.reads == expected)
|
||||
check(f"{layout.name}: no writes during discovery", rm.writes == [])
|
||||
|
||||
|
||||
def test_unknown_layout_and_nvml_mismatches_never_write() -> None:
|
||||
calls: list[tuple[int, int]] = []
|
||||
|
||||
def failing(cmd: int, data: bytearray) -> None:
|
||||
calls.append((cmd, len(data)))
|
||||
raise RmPowerError("Unsupported payload")
|
||||
|
||||
try:
|
||||
probe(BOUNDS, 250_000, failing)
|
||||
check("unknown layout rejected", False)
|
||||
except RmPowerError:
|
||||
check("unknown layout rejected", True)
|
||||
check(
|
||||
"unknown layout: only GETs attempted",
|
||||
calls == [(_PWR_GET_INFO, 0x924), (_PWR_GET_INFO, 0x488)],
|
||||
)
|
||||
|
||||
for layout in (EXTENDED_LAYOUT, LEGACY_LAYOUT):
|
||||
rm = FakeRm(layout, 250_000)
|
||||
try:
|
||||
probe(BOUNDS, 300_000, rm.query)
|
||||
check(f"{layout.name}: current mismatch rejected", False)
|
||||
except RmPowerError:
|
||||
check(f"{layout.name}: current mismatch rejected", True)
|
||||
other_bounds = PowerLimitBounds(
|
||||
min_mw=BOUNDS.min_mw, default_mw=BOUNDS.default_mw, max_mw=350_000
|
||||
)
|
||||
try:
|
||||
probe(other_bounds, 250_000, rm.query)
|
||||
check(f"{layout.name}: bounds mismatch rejected", False)
|
||||
except RmPowerError:
|
||||
check(f"{layout.name}: bounds mismatch rejected", True)
|
||||
check(f"{layout.name}: no writes on mismatch", rm.writes == [])
|
||||
|
||||
|
||||
def test_rejects_unrecognized_headers_masks_and_client_values() -> None:
|
||||
for layout in (EXTENDED_LAYOUT, LEGACY_LAYOUT):
|
||||
for at, value in [(0, 0), (4, 3), (layout.client_at, 0xF8)]:
|
||||
rm = FakeRm(layout, 250_000)
|
||||
rm.control[at] = value
|
||||
try:
|
||||
probe(BOUNDS, 250_000, rm.query)
|
||||
check(f"{layout.name}: bad header/client rejected", False)
|
||||
except RmPowerError:
|
||||
check(f"{layout.name}: bad header/client rejected", True)
|
||||
check(f"{layout.name}: no writes on bad header", rm.writes == [])
|
||||
for current in (0, 0xFFFFFFFF):
|
||||
rm = FakeRm(layout, current)
|
||||
try:
|
||||
probe(BOUNDS, current, rm.query)
|
||||
check(f"{layout.name}: empty request rejected", False)
|
||||
except RmPowerError:
|
||||
check(f"{layout.name}: empty request rejected", True)
|
||||
# The extended layout has additional mask words. Accepting only its low
|
||||
# word would allow an unexpected client to be included in a later SET.
|
||||
rm = FakeRm(EXTENDED_LAYOUT, 250_000)
|
||||
rm.control[8] = 1
|
||||
try:
|
||||
probe(BOUNDS, 250_000, rm.query)
|
||||
check("extended: nonzero mask word rejected", False)
|
||||
except RmPowerError:
|
||||
check("extended: nonzero mask word rejected", True)
|
||||
check("extended: no writes on mask violation", rm.writes == [])
|
||||
|
||||
|
||||
def test_changes_only_fe_request_and_keeps_vbios_maximum() -> None:
|
||||
for layout in (EXTENDED_LAYOUT, LEGACY_LAYOUT):
|
||||
rm = FakeRm(layout, 250_000)
|
||||
support = probe(BOUNDS, 250_000, rm.query)
|
||||
for cap in (150_000, 30_000, 250_000):
|
||||
set_limit(cap, support, rm.query)
|
||||
check(
|
||||
f"{layout.name}: set {cap} mW",
|
||||
_u32(rm.control, layout.request_at) == cap,
|
||||
)
|
||||
writes = len(rm.writes)
|
||||
for cap in (0, 29_999, 325_001, 350_000, 0xFFFFFFFF):
|
||||
try:
|
||||
set_limit(cap, support, rm.query)
|
||||
check(f"{layout.name}: out-of-range {cap} rejected", False)
|
||||
except RmPowerError:
|
||||
check(f"{layout.name}: out-of-range {cap} rejected", True)
|
||||
check(
|
||||
f"{layout.name}: no writes for out-of-range caps",
|
||||
len(rm.writes) == writes,
|
||||
)
|
||||
|
||||
|
||||
def test_restores_previous_below_minimum_request_after_set_or_readback_failure() -> (
|
||||
None
|
||||
):
|
||||
for layout in (EXTENDED_LAYOUT, LEGACY_LAYOUT):
|
||||
for fail_set in (False, True):
|
||||
rm = FakeRm(layout, 100_000)
|
||||
original = bytes(rm.control)
|
||||
rm.fail_first_write = fail_set
|
||||
rm.fail_readback = not fail_set
|
||||
support = probe(BOUNDS, 100_000, rm.query)
|
||||
try:
|
||||
set_limit(150_000, support, rm.query)
|
||||
check(f"{layout.name}: failure reported", False)
|
||||
except RmPowerError:
|
||||
check(f"{layout.name}: failure reported", True)
|
||||
check(f"{layout.name}: restore issued", len(rm.writes) == 2)
|
||||
check(
|
||||
f"{layout.name}: previous request restored",
|
||||
bytes(rm.control) == original,
|
||||
)
|
||||
|
||||
|
||||
def test_reports_restore_failure_and_rejects_wrong_client_before_writing() -> None:
|
||||
for layout in (EXTENDED_LAYOUT, LEGACY_LAYOUT):
|
||||
rm = FakeRm(layout, 100_000)
|
||||
rm.fail_first_write = True
|
||||
rm.fail_restore = True
|
||||
support = probe(BOUNDS, 100_000, rm.query)
|
||||
try:
|
||||
set_limit(150_000, support, rm.query)
|
||||
check(f"{layout.name}: restore failure reported", False)
|
||||
except RmPowerError as exc:
|
||||
check(
|
||||
f"{layout.name}: restore failure reported",
|
||||
"restoration also failed" in str(exc),
|
||||
)
|
||||
rm = FakeRm(layout, 250_000)
|
||||
rm.control[layout.client_at] = 0xF8
|
||||
try:
|
||||
set_limit(150_000, support, rm.query)
|
||||
check(f"{layout.name}: wrong client rejected", False)
|
||||
except RmPowerError:
|
||||
check(f"{layout.name}: wrong client rejected", True)
|
||||
check(f"{layout.name}: no writes for wrong client", rm.writes == [])
|
||||
|
||||
|
||||
# ── PCI identity → RM instance resolution ────────────────────────────────────
|
||||
|
||||
|
||||
def test_resolves_pci_identity_when_minor_and_rm_orders_differ() -> None:
|
||||
# This host has Ada at minor 5/RM 4 and the 5090 at minor 4/RM 5.
|
||||
# IDs are opaque and enumeration order must not select the device.
|
||||
pci = PciLocation(domain=0, bus=0x0D, dev=0, func=0)
|
||||
instances = resolve_gpu_instance(pci, lambda cmd, data: _fake_root(cmd, data))
|
||||
check("resolves by PCI identity", instances == (5, 2))
|
||||
|
||||
|
||||
def _fake_root(cmd: int, data: bytearray) -> None:
|
||||
if cmd == _CTRL_GPU_GET_ATTACHED_IDS:
|
||||
data[0:4] = (0x2E00).to_bytes(4, "little")
|
||||
data[4:8] = (0x0D00).to_bytes(4, "little")
|
||||
elif cmd == _CTRL_GPU_GET_PCI_INFO:
|
||||
gpu_id = _u32(data, 0)
|
||||
bus = 0x2E if gpu_id == 0x2E00 else 0x0D
|
||||
data[8:10] = bus.to_bytes(2, "little")
|
||||
elif cmd == _CTRL_GPU_GET_ID_INFO_V2:
|
||||
if _u32(data, 0) != 0x0D00:
|
||||
raise AssertionError("unexpected gpu id in ID_INFO_V2")
|
||||
data[8:12] = (5).to_bytes(4, "little")
|
||||
data[12:16] = (2).to_bytes(4, "little")
|
||||
else:
|
||||
raise AssertionError(f"unexpected command {cmd:#x}")
|
||||
|
||||
|
||||
def test_does_not_fall_back_to_another_gpu_when_pci_is_missing() -> None:
|
||||
pci = PciLocation(domain=1, bus=0x0D, dev=0, func=0)
|
||||
|
||||
def query(cmd: int, data: bytearray) -> None:
|
||||
if cmd == _CTRL_GPU_GET_ATTACHED_IDS:
|
||||
data[0:4] = (0x0D00).to_bytes(4, "little")
|
||||
elif cmd == _CTRL_GPU_GET_PCI_INFO:
|
||||
data[8:10] = (0x0D).to_bytes(2, "little")
|
||||
else:
|
||||
raise AssertionError("must not allocate a GPU from another PCI domain")
|
||||
|
||||
try:
|
||||
resolve_gpu_instance(pci, query)
|
||||
check("foreign PCI domain rejected", False)
|
||||
except RmPowerError:
|
||||
check("foreign PCI domain rejected", True)
|
||||
|
||||
|
||||
def test_rejects_nonzero_pci_function() -> None:
|
||||
pci = PciLocation(domain=0, bus=0x0D, dev=0, func=1)
|
||||
try:
|
||||
resolve_gpu_instance(pci, lambda cmd, data: None)
|
||||
check("nonzero function rejected", False)
|
||||
except RmPowerError:
|
||||
check("nonzero function rejected", True)
|
||||
|
||||
|
||||
def test_propagates_rm_query_failure() -> None:
|
||||
pci = PciLocation(domain=0, bus=0x0D, dev=0, func=0)
|
||||
try:
|
||||
resolve_gpu_instance(
|
||||
pci, lambda cmd, data: (_ for _ in ()).throw(RmPowerError("RM unavailable"))
|
||||
)
|
||||
check("RM query failure propagated", False)
|
||||
except RmPowerError as exc:
|
||||
check("RM query failure propagated", "RM unavailable" in str(exc))
|
||||
|
||||
|
||||
def main() -> int:
|
||||
tests = [
|
||||
test_detects_both_layouts_with_gets_without_a_driver_version,
|
||||
test_unknown_layout_and_nvml_mismatches_never_write,
|
||||
test_rejects_unrecognized_headers_masks_and_client_values,
|
||||
test_changes_only_fe_request_and_keeps_vbios_maximum,
|
||||
test_restores_previous_below_minimum_request_after_set_or_readback_failure,
|
||||
test_reports_restore_failure_and_rejects_wrong_client_before_writing,
|
||||
test_resolves_pci_identity_when_minor_and_rm_orders_differ,
|
||||
test_does_not_fall_back_to_another_gpu_when_pci_is_missing,
|
||||
test_rejects_nonzero_pci_function,
|
||||
test_propagates_rm_query_failure,
|
||||
]
|
||||
for t in tests:
|
||||
print(f"== {t.__name__} ==")
|
||||
t()
|
||||
print(f"\n{PASS} passed, {FAIL} failed")
|
||||
return 1 if FAIL else 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -0,0 +1,528 @@
|
||||
"""Tests for the native WireView Pro II reader (nvcurve/wireview.py).
|
||||
|
||||
Standalone (no pytest required):
|
||||
|
||||
python tests/test_wireview.py
|
||||
|
||||
Also works under pytest. Covers the sensor-frame parser, corruption
|
||||
detection, fault decoding, USB port discovery, and the hwmon (sysfs)
|
||||
transport.
|
||||
"""
|
||||
|
||||
import builtins
|
||||
import io
|
||||
import os
|
||||
import struct
|
||||
import sys
|
||||
import tempfile
|
||||
from unittest import mock
|
||||
|
||||
sys.path.insert(0, os.path.join(os.path.dirname(__file__), ".."))
|
||||
|
||||
from nvcurve import wireview as wv # noqa: E402
|
||||
|
||||
PASS = 0
|
||||
FAIL = 0
|
||||
|
||||
|
||||
def check(name: str, cond: bool) -> None:
|
||||
global PASS, FAIL
|
||||
if cond:
|
||||
PASS += 1
|
||||
print(f" PASS {name}")
|
||||
else:
|
||||
FAIL += 1
|
||||
print(f" FAIL {name}")
|
||||
|
||||
|
||||
def make_frame(
|
||||
ts=(392, 352, 273, 395),
|
||||
vdd=11900,
|
||||
fan=0,
|
||||
pins=((11904, 6221, 74055),) * 6,
|
||||
total_power=450568,
|
||||
total_current=37863,
|
||||
avg_voltage=11900,
|
||||
psu_cap=0,
|
||||
fault_status=0,
|
||||
fault_log=0,
|
||||
pad1=0,
|
||||
pad2=0,
|
||||
) -> bytes:
|
||||
"""Build a 100-byte sensor frame with controllable fields."""
|
||||
frame = struct.pack("<4hHB", *ts, vdd, fan)
|
||||
frame += bytes([pad1])
|
||||
for voltage, current, power in pins:
|
||||
frame += struct.pack("<hxxII", voltage, current, power)
|
||||
frame += struct.pack("<IIHB", total_power, total_current, avg_voltage, psu_cap)
|
||||
frame += bytes([pad2])
|
||||
frame += struct.pack("<HH", fault_status, fault_log)
|
||||
assert len(frame) == wv.SENSOR_STRUCT_SIZE
|
||||
return frame
|
||||
|
||||
|
||||
# ── Sensor frame parser ───────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def test_parse_sensor_struct():
|
||||
s = wv.parse_sensor_struct(make_frame())
|
||||
check("temp in", abs(s["temp_in_c"] - 39.2) < 1e-9)
|
||||
check("temp out", abs(s["temp_out_c"] - 35.2) < 1e-9)
|
||||
check("temp ext1", abs(s["temp_ext1_c"] - 27.3) < 1e-9)
|
||||
check("temp ext2", abs(s["temp_ext2_c"] - 39.5) < 1e-9)
|
||||
check("fan duty", s["fan_duty_pct"] == 0)
|
||||
check("psu capability", s["psu_capability_w"] == 600)
|
||||
check("fault status", s["fault_status"] == 0)
|
||||
check("fault log", s["fault_log"] == 0)
|
||||
check("6 pins", len(s["pins"]) == 6)
|
||||
pin = s["pins"][0]
|
||||
check("pin voltage", abs(pin["voltage_v"] - 11.904) < 1e-9)
|
||||
check("pin current", abs(pin["current_a"] - 6.221) < 1e-9)
|
||||
check("pin power", abs(pin["power_w"] - 74.055) < 1e-9)
|
||||
# Totals are computed from the pins (exporter behavior), not the
|
||||
# device's total fields.
|
||||
check(
|
||||
"total current",
|
||||
abs(s["current_total_a"] - round(6 * 6.221, 3)) < 1e-9,
|
||||
)
|
||||
check(
|
||||
"total power",
|
||||
abs(s["power_total_w"] - round(6 * 11.904 * 6.221, 3)) < 1e-9,
|
||||
)
|
||||
check(
|
||||
"avg voltage",
|
||||
abs(s["voltage_avg_v"] - 11.904) < 1e-3,
|
||||
)
|
||||
|
||||
|
||||
def test_parse_psu_capabilities():
|
||||
for cap, watts in ((0, 600), (1, 450), (2, 300), (3, 150)):
|
||||
s = wv.parse_sensor_struct(make_frame(psu_cap=cap))
|
||||
check(f"psu cap {cap} -> {watts}W", s["psu_capability_w"] == watts)
|
||||
|
||||
|
||||
def test_parse_negative_temps():
|
||||
s = wv.parse_sensor_struct(make_frame(ts=(-10, 0, 555, 999)))
|
||||
check("negative temp", abs(s["temp_in_c"] - (-1.0)) < 1e-9)
|
||||
check("zero temp", s["temp_out_c"] == 0.0)
|
||||
check("high temp", abs(s["temp_ext2_c"] - 99.9) < 1e-9)
|
||||
|
||||
|
||||
# ── Corruption detection ──────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def test_corruption_check():
|
||||
check("clean frame accepted", not wv.sensor_frame_is_corrupt(make_frame()))
|
||||
check(
|
||||
"fan duty > 100 rejected",
|
||||
wv.sensor_frame_is_corrupt(make_frame(fan=101)),
|
||||
)
|
||||
check(
|
||||
"fan duty 100 accepted",
|
||||
not wv.sensor_frame_is_corrupt(make_frame(fan=100)),
|
||||
)
|
||||
check("pad1 dirty rejected", wv.sensor_frame_is_corrupt(make_frame(pad1=1)))
|
||||
check("pad2 dirty rejected", wv.sensor_frame_is_corrupt(make_frame(pad2=1)))
|
||||
check(
|
||||
"short frame rejected",
|
||||
wv.sensor_frame_is_corrupt(make_frame()[:50]),
|
||||
)
|
||||
|
||||
|
||||
# ── Fault decoding ────────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def test_decode_faults():
|
||||
check("no faults", wv.decode_faults(0) == [])
|
||||
check(
|
||||
"chip over-temp",
|
||||
wv.decode_faults(1) == ["Chip over-temperature"],
|
||||
)
|
||||
check(
|
||||
"over-power",
|
||||
wv.decode_faults(16) == ["Over-power (OPP)"],
|
||||
)
|
||||
check(
|
||||
"multiple faults",
|
||||
wv.decode_faults(1 | 32)
|
||||
== ["Chip over-temperature", "Current imbalance"],
|
||||
)
|
||||
check(
|
||||
"unknown bits ignored",
|
||||
wv.decode_faults(1 | 0x80) == ["Chip over-temperature"],
|
||||
)
|
||||
|
||||
|
||||
def test_product_support():
|
||||
check("Pro II supported", wv.is_supported_product(0xEF, 0x05))
|
||||
check("Noctua Edition supported", wv.is_supported_product(0xEF, 0x06))
|
||||
check("WireView II unsupported", not wv.is_supported_product(0xEF, 0x07))
|
||||
check("other vendor unsupported", not wv.is_supported_product(0x12, 0x05))
|
||||
|
||||
|
||||
# ── USB port discovery ────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def test_find_wireview_ports():
|
||||
"""Fake a sysfs tree with one WireView (ttyACM0) and one other CDC
|
||||
device (ttyACM1), plus the udev symlink."""
|
||||
contents = {
|
||||
"/sys/devices/fake/usb0/idVendor": "0483\n",
|
||||
"/sys/devices/fake/usb0/idProduct": "5740\n",
|
||||
"/sys/devices/fake/usb1/idVendor": "1234\n",
|
||||
"/sys/devices/fake/usb1/idProduct": "5678\n",
|
||||
}
|
||||
|
||||
def fake_islink(p):
|
||||
return p == "/dev/wireview-pro2"
|
||||
|
||||
def fake_realpath(p):
|
||||
return {
|
||||
"/dev/wireview-pro2": "/dev/ttyACM0",
|
||||
"/sys/class/tty/ttyACM0": "/sys/devices/fake/usb0",
|
||||
"/sys/class/tty/ttyACM1": "/sys/devices/fake/usb1",
|
||||
}.get(p, p)
|
||||
|
||||
def fake_isdir(p):
|
||||
return p == "/sys/class/tty"
|
||||
|
||||
def fake_listdir(p):
|
||||
return ["ttyACM0", "ttyACM1"] if p == "/sys/class/tty" else []
|
||||
|
||||
def fake_isfile(p):
|
||||
return p in contents
|
||||
|
||||
def fake_exists(p):
|
||||
return p == "/dev/ttyACM0"
|
||||
|
||||
def fake_open(p, *a, **k):
|
||||
if p in contents:
|
||||
return io.StringIO(contents[p])
|
||||
return real_open(p, *a, **k)
|
||||
|
||||
real_open = open
|
||||
with mock.patch.object(wv.os.path, "islink", fake_islink), mock.patch.object(
|
||||
wv.os.path, "realpath", fake_realpath
|
||||
), mock.patch.object(wv.os.path, "isdir", fake_isdir), mock.patch.object(
|
||||
wv.os, "listdir", fake_listdir
|
||||
), mock.patch.object(wv.os.path, "isfile", fake_isfile), mock.patch.object(
|
||||
wv.os.path, "exists", fake_exists
|
||||
), mock.patch.object(
|
||||
wv.os.path, "dirname", os.path.dirname
|
||||
), mock.patch.object(
|
||||
builtins, "open", fake_open
|
||||
):
|
||||
ports = wv.find_wireview_ports()
|
||||
check("exactly one port found", ports == ["/dev/ttyACM0"])
|
||||
|
||||
|
||||
def test_find_wireview_ports_none():
|
||||
with mock.patch.object(wv.os.path, "islink", lambda p: False), mock.patch.object(
|
||||
wv.os.path, "isdir", lambda p: False
|
||||
), mock.patch.object(wv.os.path, "exists", lambda p: False):
|
||||
check("no ports", wv.find_wireview_ports() == [])
|
||||
|
||||
|
||||
def test_find_hwmon_path():
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
# A non-wireview hwmon and a wireview one.
|
||||
for name, dev in (("hwmon0", "coretemp"), ("hwmon1", "wireview")):
|
||||
d = os.path.join(tmp, name)
|
||||
os.makedirs(d)
|
||||
with open(os.path.join(d, "name"), "w") as f:
|
||||
f.write(dev + "\n")
|
||||
|
||||
real_join = os.path.join
|
||||
real_listdir = os.listdir
|
||||
|
||||
def fake_join(*parts):
|
||||
if parts and parts[0] == "/sys/class/hwmon":
|
||||
parts = (tmp,) + parts[1:]
|
||||
return real_join(*parts)
|
||||
|
||||
def fake_listdir(p):
|
||||
if p == "/sys/class/hwmon":
|
||||
return real_listdir(tmp)
|
||||
return real_listdir(p)
|
||||
|
||||
with mock.patch.object(wv.os.path, "join", fake_join), mock.patch.object(
|
||||
wv.os, "listdir", fake_listdir
|
||||
):
|
||||
found = wv.find_hwmon_path()
|
||||
check("hwmon path found", found == os.path.join(tmp, "hwmon1"))
|
||||
|
||||
with mock.patch.object(wv.os.path, "isdir", lambda p: False):
|
||||
check("hwmon absent", wv.find_hwmon_path() is None)
|
||||
|
||||
|
||||
# ── hwmon transport ───────────────────────────────────────────────────────────
|
||||
|
||||
|
||||
def test_hwmon_device():
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
p = os.path.join(tmp, "hwmon2")
|
||||
os.makedirs(p)
|
||||
with open(os.path.join(p, "name"), "w") as f:
|
||||
f.write("wireview\n")
|
||||
for i in range(6):
|
||||
with open(os.path.join(p, f"in{i}_input"), "w") as f:
|
||||
f.write("11904\n")
|
||||
with open(os.path.join(p, f"curr{i + 1}_input"), "w") as f:
|
||||
f.write("6221\n")
|
||||
for name, val in (
|
||||
("temp1_input", "39200"),
|
||||
("temp2_input", "35200"),
|
||||
("temp3_input", "27300"),
|
||||
("temp4_input", "39500"),
|
||||
("fault_status_raw", "0"),
|
||||
("fault_log_raw", "0"),
|
||||
("power1_cap", "600000000"),
|
||||
("pwm1", "128"),
|
||||
):
|
||||
with open(os.path.join(p, name), "w") as f:
|
||||
f.write(val + "\n")
|
||||
|
||||
dev = wv.WireViewHwmonDevice(p)
|
||||
check("hwmon connect", dev.connect())
|
||||
check("hwmon node exists", dev.node_exists())
|
||||
s = dev.read_sample()
|
||||
check("hwmon sample", s is not None)
|
||||
if s:
|
||||
check("hwmon temp in", abs(s["temp_in_c"] - 39.2) < 1e-9)
|
||||
check("hwmon fan duty ~50%", 49 <= s["fan_duty_pct"] <= 51)
|
||||
check("hwmon psu cap", s["psu_capability_w"] == 600)
|
||||
check("hwmon 6 pins", len(s["pins"]) == 6)
|
||||
check(
|
||||
"hwmon total current",
|
||||
abs(s["current_total_a"] - round(6 * 6.221, 3)) < 1e-9,
|
||||
)
|
||||
dev.close()
|
||||
check("hwmon closed", not dev.connected)
|
||||
|
||||
# A missing temp channel must yield null (not NaN): NaN would break
|
||||
# the WebSocket JSON (bare NaN token) and 500 the REST endpoint.
|
||||
os.remove(os.path.join(p, "temp3_input"))
|
||||
dev3 = wv.WireViewHwmonDevice(p)
|
||||
dev3.connect()
|
||||
s3 = dev3.read_sample()
|
||||
check("hwmon missing temp sample", s3 is not None)
|
||||
if s3:
|
||||
check("hwmon missing temp is None", s3["temp_ext1_c"] is None)
|
||||
check("hwmon missing temp present", s3["temp_in_c"] == 39.2)
|
||||
import json
|
||||
|
||||
check(
|
||||
"hwmon sample JSON-safe",
|
||||
json.dumps(s3, allow_nan=False) is not None,
|
||||
)
|
||||
dev3.close()
|
||||
|
||||
# Wrong device name is rejected.
|
||||
with open(os.path.join(p, "name"), "w") as f:
|
||||
f.write("coretemp\n")
|
||||
dev2 = wv.WireViewHwmonDevice(p)
|
||||
check("wrong name rejected", not dev2.connect())
|
||||
|
||||
|
||||
def test_hwmon_legacy_attrs():
|
||||
"""Older modules expose psu_cap/fan1_input instead of power1_cap/pwm1."""
|
||||
with tempfile.TemporaryDirectory() as tmp:
|
||||
p = os.path.join(tmp, "hwmon3")
|
||||
os.makedirs(p)
|
||||
with open(os.path.join(p, "name"), "w") as f:
|
||||
f.write("wireview\n")
|
||||
for i in range(6):
|
||||
with open(os.path.join(p, f"in{i}_input"), "w") as f:
|
||||
f.write("12000\n")
|
||||
with open(os.path.join(p, f"curr{i + 1}_input"), "w") as f:
|
||||
f.write("1000\n")
|
||||
for name, val in (
|
||||
("temp1_input", "30000"),
|
||||
("temp2_input", "30000"),
|
||||
("temp3_input", "30000"),
|
||||
("temp4_input", "30000"),
|
||||
("psu_cap", "2"),
|
||||
("fan1_input", "42"),
|
||||
("intrusion0_alarm", "0"),
|
||||
("intrusion1_alarm", "0"),
|
||||
):
|
||||
with open(os.path.join(p, name), "w") as f:
|
||||
f.write(val + "\n")
|
||||
|
||||
dev = wv.WireViewHwmonDevice(p)
|
||||
check("legacy connect", dev.connect())
|
||||
s = dev.read_sample()
|
||||
check("legacy sample", s is not None)
|
||||
if s:
|
||||
check("legacy psu cap 300W", s["psu_capability_w"] == 300)
|
||||
check("legacy fan pct", s["fan_duty_pct"] == 42)
|
||||
|
||||
|
||||
# ── Serial transport (pty-based fake device) ──────────────────────────────────
|
||||
|
||||
|
||||
def _build_info_struct(product_name: str, build_info: str) -> bytes:
|
||||
"""BuildStruct: VendorData(3) + ProductName(32) + BuildInfo(32) + NameLength(1)."""
|
||||
return (
|
||||
bytes([0xEF, 0x05, 1])
|
||||
+ product_name.encode().ljust(32, b"\x00")
|
||||
+ build_info.encode().ljust(32, b"\x00")
|
||||
+ bytes([len(product_name)])
|
||||
)
|
||||
|
||||
|
||||
def _pty_fake_device(responses: dict):
|
||||
"""Create a pty pair whose master side answers protocol commands in a
|
||||
background thread. Returns (slave_path, master_fd, thread).
|
||||
|
||||
The fake never sends the welcome string, so the device exercises its
|
||||
documented fallback: identification via the vendor-data reply.
|
||||
"""
|
||||
import pty
|
||||
import threading
|
||||
|
||||
master, slave = pty.openpty()
|
||||
slave_path = os.ttyname(slave)
|
||||
|
||||
def run():
|
||||
while True:
|
||||
try:
|
||||
cmd = os.read(master, 1)
|
||||
except OSError:
|
||||
return
|
||||
if not cmd:
|
||||
return
|
||||
resp = responses.get(cmd[0])
|
||||
if resp:
|
||||
try:
|
||||
os.write(master, resp)
|
||||
except OSError:
|
||||
return
|
||||
|
||||
t = threading.Thread(target=run, daemon=True)
|
||||
t.start()
|
||||
return slave_path, master, t
|
||||
|
||||
|
||||
def test_serial_device_protocol():
|
||||
"""Full connect handshake + sensor read against a fake device."""
|
||||
uid = bytes.fromhex("A7003100015045324B383120")
|
||||
responses = {
|
||||
wv.CMD_READ_VENDOR_DATA: bytes([0xEF, 0x05, 1]),
|
||||
wv.CMD_READ_CONFIG: bytes([0, 0, 1, 0]), # version at offset 2
|
||||
wv.CMD_READ_UID: uid,
|
||||
wv.CMD_READ_BUILD_INFO: _build_info_struct(
|
||||
"WireView Pro II", "TG-WV-PRO2-FW_20251211_1547"
|
||||
),
|
||||
wv.CMD_READ_SENSOR_VALUES: make_frame(),
|
||||
# CMD_SCREEN_CHANGE expects no response.
|
||||
}
|
||||
slave_path, master, t = _pty_fake_device(responses)
|
||||
try:
|
||||
dev = wv.WireViewSerialDevice(slave_path)
|
||||
check("serial connect", dev.connect())
|
||||
check("serial not rejected", not dev.rejected)
|
||||
info = dev.info()
|
||||
check("serial device name", info["device_name"] == "WireView Pro II")
|
||||
check("serial hw_rev", info["hw_rev"] == "EF05")
|
||||
check("serial firmware", info["firmware_version"] == "1")
|
||||
check("serial uid", info["uid"] == "A7003100015045324B383120")
|
||||
check("serial build", info["build"] == "TG-WV-PRO2-FW_20251211_1547")
|
||||
check("serial transport", info["transport"] == "serial")
|
||||
|
||||
s = dev.read_sample()
|
||||
check("serial sample", s is not None)
|
||||
if s:
|
||||
check("serial sample temp", abs(s["temp_in_c"] - 39.2) < 1e-9)
|
||||
check("serial sample pins", len(s["pins"]) == 6)
|
||||
|
||||
dev.close()
|
||||
check("serial closed", not dev.connected)
|
||||
finally:
|
||||
os.close(master)
|
||||
t.join(timeout=2)
|
||||
|
||||
|
||||
def test_serial_device_rejected_product():
|
||||
"""An unsupported product id is rejected and memoized via .rejected."""
|
||||
responses = {wv.CMD_READ_VENDOR_DATA: bytes([0xEF, 0x07, 1])}
|
||||
slave_path, master, t = _pty_fake_device(responses)
|
||||
try:
|
||||
dev = wv.WireViewSerialDevice(slave_path)
|
||||
check("serial reject connect", not dev.connect())
|
||||
check("serial rejected flag", dev.rejected)
|
||||
finally:
|
||||
os.close(master)
|
||||
t.join(timeout=2)
|
||||
|
||||
|
||||
def test_serial_device_no_response():
|
||||
"""A silent port fails the connect (vendor-data read times out)."""
|
||||
slave_path, master, t = _pty_fake_device({})
|
||||
try:
|
||||
dev = wv.WireViewSerialDevice(slave_path)
|
||||
check("serial no-response connect", not dev.connect())
|
||||
check("serial no-response not rejected", not dev.rejected)
|
||||
check("serial no-response sample", dev.read_sample() is None)
|
||||
finally:
|
||||
os.close(master)
|
||||
t.join(timeout=2)
|
||||
|
||||
|
||||
def test_serial_transport_missing_pyserial():
|
||||
"""A missing pyserial degrades gracefully: connect fails, reads are
|
||||
None, and the warning is logged once — not on every attempt."""
|
||||
import logging
|
||||
|
||||
records: list[str] = []
|
||||
|
||||
class Capture(logging.Handler):
|
||||
def emit(self, record: logging.LogRecord) -> None:
|
||||
records.append(record.getMessage())
|
||||
|
||||
logger = logging.getLogger("nvcurve.wireview")
|
||||
handler = Capture()
|
||||
old_level = logger.level
|
||||
logger.addHandler(handler)
|
||||
logger.setLevel(logging.WARNING)
|
||||
try:
|
||||
with mock.patch.object(wv, "serial", None):
|
||||
wv._serial_missing_warned = False
|
||||
dev = wv.WireViewSerialDevice("/dev/ttyACM99")
|
||||
check("no-pyserial connect", not dev.connect())
|
||||
check("no-pyserial not rejected", not dev.rejected)
|
||||
check("no-pyserial sample", dev.read_sample() is None)
|
||||
# A second attempt must not re-warn.
|
||||
dev2 = wv.WireViewSerialDevice("/dev/ttyACM99")
|
||||
check("no-pyserial second attempt", not dev2.connect())
|
||||
warnings = [r for r in records if "pyserial" in r]
|
||||
check("no-pyserial warns once", len(warnings) == 1)
|
||||
finally:
|
||||
logger.removeHandler(handler)
|
||||
logger.setLevel(old_level)
|
||||
wv._serial_missing_warned = False
|
||||
|
||||
|
||||
def main() -> int:
|
||||
print("wireview tests:")
|
||||
test_parse_sensor_struct()
|
||||
test_parse_psu_capabilities()
|
||||
test_parse_negative_temps()
|
||||
test_corruption_check()
|
||||
test_decode_faults()
|
||||
test_product_support()
|
||||
test_find_wireview_ports()
|
||||
test_find_wireview_ports_none()
|
||||
test_find_hwmon_path()
|
||||
test_hwmon_device()
|
||||
test_hwmon_legacy_attrs()
|
||||
test_serial_device_protocol()
|
||||
test_serial_device_rejected_product()
|
||||
test_serial_device_no_response()
|
||||
test_serial_transport_missing_pyserial()
|
||||
print(f"\n{PASS} passed, {FAIL} failed")
|
||||
return 1 if FAIL else 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
sys.exit(main())
|
||||
@@ -154,6 +154,22 @@ wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/04/4b/29cac41a4d98d144bf5f6d33995617b185d14b22401f75ca86f384e87ff1/h11-0.16.0-py3-none-any.whl", hash = "sha256:63cf8bbe7522de3bf65932fda1d9c2772064ffb3dae62d55932da54b31cb6c86", size = 37515, upload-time = "2025-04-24T03:35:24.344Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "hatchling"
|
||||
version = "1.32.0"
|
||||
source = { registry = "https://pypi.org/simple" }
|
||||
dependencies = [
|
||||
{ name = "packaging" },
|
||||
{ name = "pathspec" },
|
||||
{ name = "pluggy" },
|
||||
{ name = "tomlkit" },
|
||||
{ name = "trove-classifiers" },
|
||||
]
|
||||
sdist = { url = "https://files.pythonhosted.org/packages/69/08/33331757185504aae48b8d9bd78cec03a76e3aecfb52e549d05a2347c0dd/hatchling-1.32.0.tar.gz", hash = "sha256:0bdbde4a52b06c37e3eca395f85a762bf0ef06fe374fd8ae429dc6be10230f5f", size = 57783, upload-time = "2026-08-11T05:03:44.114Z" }
|
||||
wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/a9/84/1798b6d85ecde0e31546004efd25c5de1b1f49250644a60cce460e12593a/hatchling-1.32.0-py3-none-any.whl", hash = "sha256:0e17c9c3b9aa7c625acc8d0f5b622f107d5049af9ecf5ada4de1aada5be7cdbc", size = 78435, upload-time = "2026-08-11T05:03:42.644Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "httpcore"
|
||||
version = "1.0.9"
|
||||
@@ -230,9 +246,15 @@ dependencies = [
|
||||
{ name = "httpx" },
|
||||
{ name = "nvidia-ml-py" },
|
||||
{ name = "pydantic" },
|
||||
{ name = "pyserial" },
|
||||
{ name = "uvicorn", extra = ["standard"] },
|
||||
]
|
||||
|
||||
[package.dev-dependencies]
|
||||
dev = [
|
||||
{ name = "hatchling" },
|
||||
]
|
||||
|
||||
[package.metadata]
|
||||
requires-dist = [
|
||||
{ name = "bcrypt", specifier = ">=4.0" },
|
||||
@@ -240,9 +262,13 @@ requires-dist = [
|
||||
{ name = "httpx", specifier = ">=0.27" },
|
||||
{ name = "nvidia-ml-py", specifier = ">=12.0" },
|
||||
{ name = "pydantic", specifier = ">=2.0" },
|
||||
{ name = "pyserial", specifier = ">=3.5" },
|
||||
{ name = "uvicorn", extras = ["standard"], specifier = ">=0.30" },
|
||||
]
|
||||
|
||||
[package.metadata.requires-dev]
|
||||
dev = [{ name = "hatchling" }]
|
||||
|
||||
[[package]]
|
||||
name = "nvidia-ml-py"
|
||||
version = "13.590.48"
|
||||
@@ -252,6 +278,33 @@ wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/fd/72/fb2af0d259a651affdce65fd6a495f0e07a685a0136baf585c5065204ee7/nvidia_ml_py-13.590.48-py3-none-any.whl", hash = "sha256:fd43d30ee9cd0b7940f5f9f9220b68d42722975e3992b6c21d14144c48760e43", size = 50680, upload-time = "2026-01-22T01:14:55.281Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "packaging"
|
||||
version = "26.3"
|
||||
source = { registry = "https://pypi.org/simple" }
|
||||
sdist = { url = "https://files.pythonhosted.org/packages/7d/fa/3944b40b07da9ce895c0e6303a5ab7d53da063554f534556b134a54d6093/packaging-26.3.tar.gz", hash = "sha256:94edc256424af38762eb31306eed28beb9f0efc50a8837492c9d6fd6004aed79", size = 313412, upload-time = "2026-08-04T18:15:28.737Z" }
|
||||
wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/63/34/ba1c580383c9eada3711951fef0795c80b829a078d72188184bcab9dd527/packaging-26.3-py3-none-any.whl", hash = "sha256:d7193f7c8e4e93f444fde0262bf90af30e16fa0ad0ad44cb553c87339b23cd1c", size = 129956, upload-time = "2026-08-04T18:15:27.159Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "pathspec"
|
||||
version = "1.1.1"
|
||||
source = { registry = "https://pypi.org/simple" }
|
||||
sdist = { url = "https://files.pythonhosted.org/packages/5a/82/42f767fc1c1143d6fd36efb827202a2d997a375e160a71eb2888a925aac1/pathspec-1.1.1.tar.gz", hash = "sha256:17db5ecd524104a120e173814c90367a96a98d07c45b2e10c2f3919fff91bf5a", size = 135180, upload-time = "2026-04-27T01:46:08.907Z" }
|
||||
wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/f1/d9/7fb5aa316bc299258e68c73ba3bddbc499654a07f151cba08f6153988714/pathspec-1.1.1-py3-none-any.whl", hash = "sha256:a00ce642f577bf7f473932318056212bc4f8bfdf53128c78bbd5af0b9b20b189", size = 57328, upload-time = "2026-04-27T01:46:07.06Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "pluggy"
|
||||
version = "1.6.0"
|
||||
source = { registry = "https://pypi.org/simple" }
|
||||
sdist = { url = "https://files.pythonhosted.org/packages/f9/e2/3e91f31a7d2b083fe6ef3fa267035b518369d9511ffab804f839851d2779/pluggy-1.6.0.tar.gz", hash = "sha256:7dcc130b76258d33b90f61b658791dede3486c3e6bfb003ee5c9bfb396dd22f3", size = 69412, upload-time = "2025-05-15T12:30:07.975Z" }
|
||||
wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/54/20/4d324d65cc6d9205fabedc306948156824eb9f0ee1633355a8f7ec5c66bf/pluggy-1.6.0-py3-none-any.whl", hash = "sha256:e920276dd6813095e9377c0bc5566d94c932c33b27a3e3945d8389c374dd4746", size = 20538, upload-time = "2025-05-15T12:30:06.134Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "pydantic"
|
||||
version = "2.12.5"
|
||||
@@ -338,6 +391,15 @@ wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/f7/07/34573da085946b6a313d7c42f82f16e8920bfd730665de2d11c0c37a74b5/pydantic_core-2.41.5-graalpy312-graalpy250_312_native-manylinux_2_17_x86_64.manylinux2014_x86_64.whl", hash = "sha256:76d0819de158cd855d1cbb8fcafdf6f5cf1eb8e470abe056d5d161106e38062b", size = 2139017, upload-time = "2025-11-04T13:42:59.471Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "pyserial"
|
||||
version = "3.5"
|
||||
source = { registry = "https://pypi.org/simple" }
|
||||
sdist = { url = "https://files.pythonhosted.org/packages/1e/7d/ae3f0a63f41e4d2f6cb66a5b57197850f919f59e558159a4dd3a818f5082/pyserial-3.5.tar.gz", hash = "sha256:3c77e014170dfffbd816e6ffc205e9842efb10be9f58ec16d3e8675b4925cddb", size = 159125, upload-time = "2020-11-23T03:59:15.045Z" }
|
||||
wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/07/bc/587a445451b253b285629263eb51c2d8e9bcea4fc97826266d186f96f558/pyserial-3.5-py2.py3-none-any.whl", hash = "sha256:c4451db6ba391ca6ca299fb3ec7bae67a5c55dde170964c7a14ceefec02f2cf0", size = 90585, upload-time = "2020-11-23T03:59:13.41Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "python-dotenv"
|
||||
version = "1.2.2"
|
||||
@@ -406,6 +468,24 @@ wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/81/0d/13d1d239a25cbfb19e740db83143e95c772a1fe10202dda4b76792b114dd/starlette-0.52.1-py3-none-any.whl", hash = "sha256:0029d43eb3d273bc4f83a08720b4912ea4b071087a3b48db01b7c839f7954d74", size = 74272, upload-time = "2026-01-18T13:34:09.188Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "tomlkit"
|
||||
version = "0.15.1"
|
||||
source = { registry = "https://pypi.org/simple" }
|
||||
sdist = { url = "https://files.pythonhosted.org/packages/94/96/e07752635b98536177fa1f37671c8f3cdde2e724c6bcf6034b2cfb571565/tomlkit-0.15.1.tar.gz", hash = "sha256:e25bbf38843005246210a12982776f27f99cb9be67160e14434d0c0d21ee1e97", size = 180129, upload-time = "2026-07-17T01:48:04.562Z" }
|
||||
wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/13/bc/8c13eb66537dce1d2bd3a57132902f38d0e7f5bb46fa9f4daed9fe9d76ee/tomlkit-0.15.1-py3-none-any.whl", hash = "sha256:177a05aece5a8ca5266fd3c448abb47b8d352f09d477d3ca8332db4d89b24304", size = 49449, upload-time = "2026-07-17T01:48:05.728Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "trove-classifiers"
|
||||
version = "2026.6.1.19"
|
||||
source = { registry = "https://pypi.org/simple" }
|
||||
sdist = { url = "https://files.pythonhosted.org/packages/c2/e3/7ca82ee24c82d344584abd5b8637b3bd056f2900226e8d82fc22f1184b92/trove_classifiers-2026.6.1.19.tar.gz", hash = "sha256:c5132b4b61a829d11cfbd2d72e97f20a45ed6edb95e45c5efdeb5e00836b2745", size = 17059, upload-time = "2026-06-01T19:41:34.649Z" }
|
||||
wheels = [
|
||||
{ url = "https://files.pythonhosted.org/packages/7c/a4/81502f486f01db95bc8320646a8a12511f5e556cb63d5e224d91816605c4/trove_classifiers-2026.6.1.19-py3-none-any.whl", hash = "sha256:ab4c4ec93cc4a4e7815fa759906e05e6bb3f2fbd92ea0f897288c6a43efd15b3", size = 14211, upload-time = "2026-06-01T19:41:33.434Z" },
|
||||
]
|
||||
|
||||
[[package]]
|
||||
name = "typing-extensions"
|
||||
version = "4.15.0"
|
||||
|
||||
Reference in new issue
Block a user