Standalone support for the ThermalGrizzly WireView Pro II without
depending on the external wireview_reporter exporter:
- wireview.py: serial protocol (STX/ETX + 16-bit CRC16-CCITT) with
vendor-data product identification, config version, UID, build
string, screen layout, and temperature/power/current sensor reads;
hwmon fallback with per-channel index resolution; udev-based
detection (vendor 0x2560 / product 0x0101) with fallback port
probing; serial read timeout and watchdog reconnect
- server.py: startup detection (root only), 1 Hz poller gated on
subscribed clients, WS 'wireview' channel, GET /api/wireview,
rejected-product memoization, stale-read race guard
- frontend: WireView tab (visible when a device is detected) with
live temperature/power/current cards, sparkline history, and
device info; nullable temp channels
- tests: 76 tests covering parser, CRC, fault classification,
hwmon resolution, JSON safety, and the serial transport against
a pty-based fake device
Verified live: tab appears with the device connected, disappears
when unplugged, reconnects on re-plug, no serial traffic when idle.
Adds an experimental power-cap mode using the undocumented RM ioctl
interface (based on panchovix's LACT PR #1205) to set power limits
below the VBIOS minimum (down to 30 W).
- hal/rm_power.py: RM ioctl power-cap read/write/reset + runtime probe
- limits.py: power_cap_mode (nvml/ioctl) with support detection
- config.py: persist power_cap_mode per GPU
- profiles: record/apply power_cap_mode
- server.py: POST /api/limits validates ioctl support (409 on failure)
- cli.py: profile save falls back to persisted mode
- client.py: power_cap_mode in Limits
- frontend: toggle + warning with panchovix attribution (LACT #1205)
- tests: test_rm_power.py (unit) + integration coverage
- Makefile: add test_rm_power.py to make test
Also includes automated linter reformatting (prettier, ruff, shellcheck,
isort, markdownlint) that the linter would apply anyway.
Security review findings, fixed and verified:
Critical
- Fix unauthenticated arbitrary file read: the SPA catch-all route
joined the raw URL path onto the dist dir without containment, so
encoded '..' segments (/%2e%2e/etc/passwd) leaked any file readable
by the root server. Resolve with realpath and reject paths outside
the dist dir (fail-closed 404).
High
- Daemon socket: serve_start no longer accepts caller-chosen
host/port. The socket is world-connectable (unprivileged CLI users),
so callers could previously rebind the root web server to 0.0.0.0.
The daemon now always binds the operator-configured address and
reports it in the response; the CLI warns on mismatch.
Medium
- Remove the per-request max_delta_khz override from the API: the
server-enforced safety cap is now authoritative. CLI direct paths
(write, profile apply, verify) honor the configured cap; --max-delta
still overrides for explicit root use.
- Snapshot restore: confine filepath to the snapshot directory
(realpath containment; blocks symlink escapes).
- Login lockout: honor X-Forwarded-For only for peers listed in the
new trusted_proxies config (rightmost untrusted hop), so the
per-IP lockout works behind a reverse proxy. Spoofed headers from
untrusted peers are ignored.
- /api/shutdown: new allow_api_shutdown config (default true);
shared systems can disable the API shutdown path.
TLS (opt-in, like auth)
- New ssl_certfile/ssl_keyfile config + CLI flags (serve start,
service install/configure, --no-ssl to disable). When active:
HTTPS for UI/API, wss:// for WebSockets, Secure session cookie,
CLI auto-switches to https://. Cert/key paths are validated up
front with a clear error instead of a silent uvicorn crash.
Tests & docs
- tests/test_security.py: standalone regression tests (no new deps)
covering SPA containment, snapshot containment, cap removal,
client-IP derivation, proxy normalization, TLS scheme detection,
and daemon host/port hardening.
- README + Usage-Guide: TLS section, new config keys, updated
security notes.
The fan curve previously only controlled fan index 0; secondary fans
stayed on driver control. The curve can now target all fans (new
default) or any individual fan(s).
Backend:
- hal/fans.py: get_num_fans() via nvmlDeviceGetNumFans; get_fan_info()
returns per-fan speeds; set_fan_speed() accepts a fan index list
(None = all fans; all-fans mode is lenient toward driver-locked
fans, explicit lists are strict); reset_fan() restores all fans.
- server.py: per-GPU fan_targets state; the poller applies the curve to
all target fans and logs write failures (once per distinct error);
activation validates targets against the hardware (stale indices fall
back to all fans); POST /api/fans accepts fans, POST /api/fans/speed
accepts a fan index, GET /api/fans returns num_fans/fans/fan_targets.
- Persistence format is now {"curve": ..., "fans": ...}; legacy
bare-curve entries migrate to "all fans" at startup.
- Profiles save/apply fan_targets alongside fan_curve.
- MonitoringSample carries per-fan speeds for live gauges.
Frontend:
- Fans tab: All / Fan 1 / Fan 2 / ... selector with live per-fan %;
the selection is applied together with the curve.
- Live monitor: per-fan gauges with sparklines for multi-fan GPUs.
Address three fan-settings issues:
1. Profile view: show a fan icon next to a profile's name when it has a
saved fan curve, so it's clear which profiles carry custom fans.
2. Fan curve persistence: the active fan curve was in-memory only and lost
on every server restart. It is now persisted per-GPU in
/etc/nvcurve/config.json (fan_curves) and re-applied at server startup,
so an applied curve survives restarts. All fan-curve state changes route
through _activate_fan_curve/_deactivate_fan_curve helpers that keep the
persisted state in sync (apply, reset, and profile apply).
3. Point removal: the fan-curve remove button was nearly invisible. The
chart remove control is now always faintly visible with an X glyph, the
table remove button is larger with a tooltip, and a hint line explains
how to add/remove points.
Also includes a formatting pass over the two edited frontend files.
Add a new Dashboard tab as the default/first tab (Dashboard - Curve -
Performance - Fans) showing a full GPU overview.
Backend:
- New hal/dashboard.py: one-shot static GPU info via NVML (VBIOS, CUDA
cores, compute capability, bus width, BAR1/CPU-accessible VRAM,
Resizable BAR, PCIe max, max clocks, power limits, temp thresholds,
persistence mode, fan count, serial, board part number, UUID).
ROP count and VRAM type are best-effort from a per-model table since
NVML does not expose them.
- New /api/dashboard endpoint.
- MonitoringSample gains live throttle_reasons (+label), PCIe link
width/generation (downclocks when idle, so read per poll), and VRAM
temp (from the MEMORY thermal sensor, if exposed).
Frontend:
- New Dashboard component: critical top row (throttling, voltage,
GPU/VRAM temps), full live-monitor grid with sparklines, and a static
GPU-information grid. Unavailable fields are omitted.
- useDashboard hook + DashboardInfo type + api.client dashboard().
Also: restrict CORS to localhost origins (end-anchored), make
_int_key_deltas fail closed on bad keys, and clean up lint blockers in
server.py/monitoring.py.
Add optional login protection for the web UI/API, intended for shared
machines (e.g. AI servers). Dual mode: with no users configured the API
and web UI are open (as before); once at least one user exists, every
/api/* and /ws/* endpoint requires a valid session.
- bcrypt password hashing: passwords stored as $2b$ hashes in
/etc/nvcurve/users.json (0600, root-owned); plaintext never persisted.
- 24-hour sessions: HttpOnly cookie for browsers, Authorization: Bearer
token for CLI/scripts; in-memory, invalidated on server restart.
- Multi-user: multiple named accounts (no shared-password mode).
- New CLI: nvcurve user add|list|remove|set-password (root for mutating
ops; password always prompted, never a CLI argument).
- New endpoints: GET /api/ping (public), /api/auth/status|login|logout|users.
- Web UI: sign-in screen when auth is enabled; status bar shows the
signed-in user with sign-out; expired sessions (401) re-show sign-in.
- Brute-force lockout: 10 failed logins/IP within 5 min -> 15 min lockout.
- New dependency: bcrypt.
Also: LSP config (pyrightconfig.json) pointing at the project .venv, and
small error-handling cleanups in daemon.py/server.py.
- Backend HAL (hal/fans.py): NVML v2 fan read/set/reset with min/max queries
- Server endpoints: GET/POST /api/fans, POST /api/fans/reset, POST /api/fans/speed
- Background fan poller: reads GPU temp every 2s, interpolates fan speed from curve
- Profile integration: fan_curve field saved/applied, auto-restore on shutdown
- Frontend: FanCurveEditor (SVG chart with drag/add/delete points), FanMonitor sidebar
- App.tsx: three-tab layout (Curve, Performance, Fans)
- GaugeCard: optional history sparkline, Fan Mode card without sparkline
- fan_mode field populated as 'curve' or 'auto' in GET /api/fans