Files
nvcurve/docs/API.md
T
ARIA b75b9d43e9 docs: add API reference and enrich OpenAPI metadata
- docs/API.md: full REST + WebSocket reference with auth usage
  (cookie/Bearer login flow, curl examples), conventions, endpoint
  tables, and WS protocol (WS routes are not in the OpenAPI schema)
- server.py: app-level description, openapi_tags grouping, and
  tags/responses metadata on all REST routes documenting the error
  codes each route raises; hide the SPA catch-all from the schema
- README: link the API reference in the docs index
2026-10-03 14:51:58 +02:00

244 lines
10 KiB
Markdown

# NVCurve API Reference
The web server (`nvcurve serve`) exposes a **REST + WebSocket API** on
`http://127.0.0.1:8042` by default. The same server also serves the web UI
SPA, so the API and UI share one port.
**Interactive docs** (served by the running server, no login required):
| URL | Description |
| --- | --- |
| `/docs` | Swagger UI (interactive, per-endpoint request/response metadata) |
| `/redoc` | ReDoc (read-friendly reference) |
| `/openapi.json` | Machine-readable OpenAPI 3.1 schema |
> WebSocket endpoints are **not** part of the OpenAPI schema (OpenAPI has no
> WebSocket operation type) — they are documented in the
> [WebSockets](#websockets) section below.
---
## Authentication
The server has two modes, decided automatically:
- **No users configured** → the API is fully open (no auth).
- **One or more users configured** → every `/api/*` and `/ws/*` endpoint
requires a valid session, except the public endpoints listed below.
Sessions last **24 hours**. Passwords are stored as bcrypt hashes.
### Public endpoints (always reachable)
| Endpoint | Purpose |
| --- | --- |
| `GET /api/ping` | Liveness probe (CLI uses it to detect a running server) |
| `GET /api/auth/status` | Whether auth is required and whether *this* request is authenticated |
| `POST /api/auth/login` | Authenticate, obtain session |
| `POST /api/auth/logout` | End the current session |
### Obtaining a session
```bash
curl -s -X POST http://127.0.0.1:8042/api/auth/login \
-H 'Content-Type: application/json' \
-d '{"username": "alice", "password": "secret"}'
```
Response:
```json
{
"ok": true,
"username": "alice",
"expires_at": 1750000000.0,
"token": "9f2c1a…"
}
```
On success the server also sets an `HttpOnly` cookie `nvcurve_session`
(`Secure` when TLS is enabled). You can use either credential:
- **Cookie** — what browsers use automatically.
- **Bearer token** — what scripts/CLI use:
```bash
TOKEN=$(curl -s -X POST http://127.0.0.1:8042/api/auth/login \
-H 'Content-Type: application/json' \
-d '{"username": "alice", "password": "secret"}' | jq -r .token)
curl -s http://127.0.0.1:8042/api/curve -H "Authorization: Bearer $TOKEN"
```
Notes:
- For WebSockets, the token is read from the `Authorization` header or the
cookie on the handshake. It is **deliberately not** accepted via query
string (the access log records full paths).
- Failed logins are rate-limited per client IP; lockouts return `429`.
Behind a reverse proxy, set `trusted_proxies` in the config so the real
client IP is used.
- `GET /api/auth/status` reports your current state:
```json
{"auth_required": true, "authenticated": true, "username": "alice", "expires_at": 1750000000.0}
```
---
## Conventions
- **`gpu_index`** — GPU-selecting routes take a 0-based `gpu_index` query
parameter (default `0`). Unknown indexes → `404`; an uninitialized GPU →
`503`.
- **Errors** — JSON `{"detail": "..."}`. Validation errors on write endpoints
return `400` with `detail.errors[]` listing the offending points.
- **Safety cap** — curve writes are capped by the server-side
`max_delta_khz` (default ±3000 MHz, set in `/etc/nvcurve/config.json`).
Clients cannot raise it per request.
- **Units** — curve deltas are in **kHz**; clock offsets and memory offsets
are in **MHz**; voltages in µV/mV; power in watts.
- **Side effects** — write endpoints (`/api/curve/*`, `/api/limits`,
`/api/profiles/*/apply`, `/api/snapshot/restore`) change hardware state.
With `auto_snapshot: true` (default) a ClockBoostTable snapshot is saved
before every curve write.
---
## REST endpoints
### Auth
| Method & path | Description |
| --- | --- |
| `GET /api/ping` | Liveness probe → `{"ok": true}` |
| `GET /api/auth/status` | Auth mode + current session state |
| `POST /api/auth/login` | Body `{"username", "password"}` → session token + cookie. Errors: `401` bad credentials, `404` auth disabled, `429` locked out |
| `POST /api/auth/logout` | End session (idempotent) |
| `GET /api/auth/users` | List configured usernames (requires session) |
### GPU
| Method & path | Description |
| --- | --- |
| `GET /api/gpus` | List discovered GPUs → `[{index, name, uuid, pci_bus_id}]` |
| `GET /api/gpu?gpu_index=0` | Name, driver version, VRAM |
| `GET /api/dashboard?gpu_index=0` | Static dashboard info (VBIOS, CUDA cores, PCIe, BAR1, …). Live values come from the monitor WebSocket |
| `GET /api/monitor?gpu_index=0` | One-shot monitoring snapshot: voltage, clocks, temp, power, fans, p-state, VRAM, utilization, throttle reasons |
### Curve
| Method & path | Description |
| --- | --- |
| `GET /api/curve?gpu_index=0` | Full curve: `{gpu_name, timestamp, points: [{index, freq_khz, volt_mv, delta_khz, effective_freq_khz, domain, …}]}` |
| `GET /api/curve/{point}?gpu_index=0` | Single V/F point detail (`400` if index out of range) |
| `GET /api/ranges?gpu_index=0` | Clock boost domain ranges (min/max offset per domain) |
| `GET /api/voltage?gpu_index=0` | Current core voltage → `{voltage_uv, voltage_mv}` |
| `POST /api/curve/write?gpu_index=0` | Body `{"deltas": {"<point>": <delta_khz>}}` — write per-point frequency offsets. Response includes `ok`, `return_code`, optional `warning` (external change detected) and `freq_warnings` (negative-frequency risk) |
| `POST /api/curve/write/global?gpu_index=0` | Body `{"delta_khz": N}` — uniform offset on all GPU-domain points |
| `POST /api/curve/reset?gpu_index=0` | Reset all frequency offsets to zero |
| `POST /api/curve/verify?gpu_index=0` | Body `{"deltas": {...}}` — write-verify-read cycle; returns per-point `match` results and `collateral_changes` on other points |
### Profiles
Profiles capture the current curve deltas, power limit, memory offset and
active fan curve as a named JSON file.
| Method & path | Description |
| --- | --- |
| `GET /api/profiles?gpu_index=0` | `{profiles: [...], active, auto_load}` |
| `POST /api/profiles?gpu_index=0` | Body `{"name": "..."}` — save current GPU state as a profile |
| `POST /api/profiles/{name}/apply?gpu_index=0` | Apply a saved profile to hardware. `404` if missing; `500` with per-part errors if any part fails |
| `DELETE /api/profiles/{name}` | Delete a profile (also clears auto-load references) |
| `POST /api/profiles/{name}/rename` | Body `{"new_name": "..."}` |
### Config
| Method & path | Description |
| --- | --- |
| `GET /api/config?gpu_index=0` | `{auto_load_profile}` for the GPU |
| `POST /api/config` | Body `{"auto_load_profile": "name" \| null, "gpu_index": 0}` — set/clear the profile auto-applied at server start. Persists to `/etc/nvcurve/config.json` if present |
### Limits
| Method & path | Description |
| --- | --- |
| `GET /api/limits?gpu_index=0` | Power limit (current/default/min/max), `gpc_offset_mhz`, `mem_offset_mhz`, memory offset range |
| `POST /api/limits?gpu_index=0` | Body `{power_limit_w?, mem_offset_mhz?, power_cap_mode?}` — any subset. `power_cap_mode` is `"nvml"` (default) or `"ioctl"` (experimental RM power control; `409` if unsupported). Changing the memory offset may reset the curve table — the server re-applies the last known curve offsets |
| `POST /api/limits/reset?gpu_index=0` | Power limit → hardware default, memory offset → 0 |
### Fans
| Method & path | Description |
| --- | --- |
| `GET /api/fans?gpu_index=0` | Per-fan state: `{fan_pct, fans: [{index, fan_pct}], num_fans, min_fan_pct, max_fan_pct, fan_mode: "auto" \| "curve", curve, curve_active, fan_targets}` |
| `POST /api/fans?gpu_index=0` | Body `{"curve": [{"temp_c": 60, "fan_pct": 40}, …], "fans": [0] \| null}` — set the fan curve and start the control poller. `fans: null` drives all fans. `400` for invalid curves (e.g. non-monotonic temps) |
| `POST /api/fans/reset?gpu_index=0` | Deactivate curve control, restore automatic fan mode |
| `POST /api/fans/speed?gpu_index=0` | Body `{"fan_pct": 50, "fan": 0 \| null}` — one-shot exact speed (bypasses curve) |
Active fan curves persist across server restarts.
### WireView Pro II
| Method & path | Description |
| --- | --- |
| `GET /api/wireview` | `{available, connected, info, sample}` — `available` is true when a Thermal Grizzly WireView Pro II is connected; `sample` is the most recent reading (`null` until the first) |
### Snapshots
| Method & path | Description |
| --- | --- |
| `GET /api/snapshots` | List saved ClockBoostTable snapshots → `[{filepath, timestamp, gpu, nonzero_offsets, size}]` |
| `POST /api/snapshot/save?gpu_index=0` | Save the current ClockBoostTable → `{ok, filepath}` |
| `POST /api/snapshot/restore?gpu_index=0` | Body `{"filepath": "…" \| null}` — restore a snapshot (most recent if omitted) |
### Server
| Method & path | Description |
| --- | --- |
| `POST /api/shutdown` | Gracefully stop the server process. `403` when `allow_api_shutdown: false` (recommended on shared systems — use systemd instead) |
---
## WebSockets
All three endpoints use the same handshake: connect, then send a subscribe
message. The server sends an initial state payload, then keeps streaming.
| Endpoint | Subscribe message | Stream |
| --- | --- | --- |
| `/ws/monitor` | `{"action": "subscribe", "gpu_index": 0}` | Monitoring sample every `poll_interval_s` (default 1 s): voltage, clocks, temp, power, fans, p-state, VRAM, utilization, throttle reasons |
| `/ws/curve` | `{"action": "subscribe", "gpu_index": 0}` | Full curve state, pushed whenever the curve changes (after any write) |
| `/ws/wireview` | `{"action": "subscribe"}` | `{"type": "unavailable"}` or `{"type": "sample", "info": {...}, "sample": {...}}` per poll tick |
Authentication (when enabled): the session token is taken from the
`Authorization: Bearer <token>` header or the `nvcurve_session` cookie on the
handshake. Unauthenticated connections are closed with code `1008`.
Example (Python `websockets`):
```python
import asyncio, json, websockets
async def main():
async with websockets.connect(
"ws://127.0.0.1:8042/ws/monitor",
extra_headers={"Authorization": f"Bearer {token}"},
) as ws:
await ws.send(json.dumps({"action": "subscribe", "gpu_index": 0}))
async for msg in ws:
sample = json.loads(msg)
print(sample["temp_c"], sample["clock_mhz"])
asyncio.run(main())
```
---
## CLI
The `nvcurve` CLI talks to this same API (default base
`http://127.0.0.1:8042`, override with `--base-url`), so any scripted use
covered by the CLI works against the API as well. See
[Usage-Guide.md](Usage-Guide.md) for CLI details.