diff --git a/README.md b/README.md index a386458..e6c0ba3 100644 --- a/README.md +++ b/README.md @@ -205,6 +205,7 @@ The GPU key can be a UUID, `pci:XXXX`, or `idx:N` fallback. Find your GPU key wi - **[Overview](docs/Overview.md)** — What it does, capabilities, architecture - **[Installation](docs/Installation.md)** — Prerequisites, source build, troubleshooting - **[Usage Guide](docs/Usage-Guide.md)** — Web UI, CLI reference, systemd service +- **[API Reference](docs/API.md)** — REST + WebSocket API, authentication, endpoint summary (interactive docs at `/docs`) - **[Tips and Tricks](docs/Tips-and-Tricks.md)** — Workflows, curve flattening, safety - **[WireView Pro II Setup](docs/WireView.md)** — One-time host setup (udev rule + port access) for the 12VHPWR connector monitor diff --git a/docs/API.md b/docs/API.md new file mode 100644 index 0000000..16b187d --- /dev/null +++ b/docs/API.md @@ -0,0 +1,243 @@ +# NVCurve API Reference + +The web server (`nvcurve serve`) exposes a **REST + WebSocket API** on +`http://127.0.0.1:8042` by default. The same server also serves the web UI +SPA, so the API and UI share one port. + +**Interactive docs** (served by the running server, no login required): + +| URL | Description | +| --- | --- | +| `/docs` | Swagger UI (interactive, per-endpoint request/response metadata) | +| `/redoc` | ReDoc (read-friendly reference) | +| `/openapi.json` | Machine-readable OpenAPI 3.1 schema | + +> WebSocket endpoints are **not** part of the OpenAPI schema (OpenAPI has no +> WebSocket operation type) — they are documented in the +> [WebSockets](#websockets) section below. + +--- + +## Authentication + +The server has two modes, decided automatically: + +- **No users configured** → the API is fully open (no auth). +- **One or more users configured** → every `/api/*` and `/ws/*` endpoint + requires a valid session, except the public endpoints listed below. + +Sessions last **24 hours**. Passwords are stored as bcrypt hashes. + +### Public endpoints (always reachable) + +| Endpoint | Purpose | +| --- | --- | +| `GET /api/ping` | Liveness probe (CLI uses it to detect a running server) | +| `GET /api/auth/status` | Whether auth is required and whether *this* request is authenticated | +| `POST /api/auth/login` | Authenticate, obtain session | +| `POST /api/auth/logout` | End the current session | + +### Obtaining a session + +```bash +curl -s -X POST http://127.0.0.1:8042/api/auth/login \ + -H 'Content-Type: application/json' \ + -d '{"username": "alice", "password": "secret"}' +``` + +Response: + +```json +{ + "ok": true, + "username": "alice", + "expires_at": 1750000000.0, + "token": "9f2c1a…" +} +``` + +On success the server also sets an `HttpOnly` cookie `nvcurve_session` +(`Secure` when TLS is enabled). You can use either credential: + +- **Cookie** — what browsers use automatically. +- **Bearer token** — what scripts/CLI use: + +```bash +TOKEN=$(curl -s -X POST http://127.0.0.1:8042/api/auth/login \ + -H 'Content-Type: application/json' \ + -d '{"username": "alice", "password": "secret"}' | jq -r .token) + +curl -s http://127.0.0.1:8042/api/curve -H "Authorization: Bearer $TOKEN" +``` + +Notes: + +- For WebSockets, the token is read from the `Authorization` header or the + cookie on the handshake. It is **deliberately not** accepted via query + string (the access log records full paths). +- Failed logins are rate-limited per client IP; lockouts return `429`. + Behind a reverse proxy, set `trusted_proxies` in the config so the real + client IP is used. +- `GET /api/auth/status` reports your current state: + + ```json + {"auth_required": true, "authenticated": true, "username": "alice", "expires_at": 1750000000.0} + ``` + +--- + +## Conventions + +- **`gpu_index`** — GPU-selecting routes take a 0-based `gpu_index` query + parameter (default `0`). Unknown indexes → `404`; an uninitialized GPU → + `503`. +- **Errors** — JSON `{"detail": "..."}`. Validation errors on write endpoints + return `400` with `detail.errors[]` listing the offending points. +- **Safety cap** — curve writes are capped by the server-side + `max_delta_khz` (default ±3000 MHz, set in `/etc/nvcurve/config.json`). + Clients cannot raise it per request. +- **Units** — curve deltas are in **kHz**; clock offsets and memory offsets + are in **MHz**; voltages in µV/mV; power in watts. +- **Side effects** — write endpoints (`/api/curve/*`, `/api/limits`, + `/api/profiles/*/apply`, `/api/snapshot/restore`) change hardware state. + With `auto_snapshot: true` (default) a ClockBoostTable snapshot is saved + before every curve write. + +--- + +## REST endpoints + +### Auth + +| Method & path | Description | +| --- | --- | +| `GET /api/ping` | Liveness probe → `{"ok": true}` | +| `GET /api/auth/status` | Auth mode + current session state | +| `POST /api/auth/login` | Body `{"username", "password"}` → session token + cookie. Errors: `401` bad credentials, `404` auth disabled, `429` locked out | +| `POST /api/auth/logout` | End session (idempotent) | +| `GET /api/auth/users` | List configured usernames (requires session) | + +### GPU + +| Method & path | Description | +| --- | --- | +| `GET /api/gpus` | List discovered GPUs → `[{index, name, uuid, pci_bus_id}]` | +| `GET /api/gpu?gpu_index=0` | Name, driver version, VRAM | +| `GET /api/dashboard?gpu_index=0` | Static dashboard info (VBIOS, CUDA cores, PCIe, BAR1, …). Live values come from the monitor WebSocket | +| `GET /api/monitor?gpu_index=0` | One-shot monitoring snapshot: voltage, clocks, temp, power, fans, p-state, VRAM, utilization, throttle reasons | + +### Curve + +| Method & path | Description | +| --- | --- | +| `GET /api/curve?gpu_index=0` | Full curve: `{gpu_name, timestamp, points: [{index, freq_khz, volt_mv, delta_khz, effective_freq_khz, domain, …}]}` | +| `GET /api/curve/{point}?gpu_index=0` | Single V/F point detail (`400` if index out of range) | +| `GET /api/ranges?gpu_index=0` | Clock boost domain ranges (min/max offset per domain) | +| `GET /api/voltage?gpu_index=0` | Current core voltage → `{voltage_uv, voltage_mv}` | +| `POST /api/curve/write?gpu_index=0` | Body `{"deltas": {"": }}` — write per-point frequency offsets. Response includes `ok`, `return_code`, optional `warning` (external change detected) and `freq_warnings` (negative-frequency risk) | +| `POST /api/curve/write/global?gpu_index=0` | Body `{"delta_khz": N}` — uniform offset on all GPU-domain points | +| `POST /api/curve/reset?gpu_index=0` | Reset all frequency offsets to zero | +| `POST /api/curve/verify?gpu_index=0` | Body `{"deltas": {...}}` — write-verify-read cycle; returns per-point `match` results and `collateral_changes` on other points | + +### Profiles + +Profiles capture the current curve deltas, power limit, memory offset and +active fan curve as a named JSON file. + +| Method & path | Description | +| --- | --- | +| `GET /api/profiles?gpu_index=0` | `{profiles: [...], active, auto_load}` | +| `POST /api/profiles?gpu_index=0` | Body `{"name": "..."}` — save current GPU state as a profile | +| `POST /api/profiles/{name}/apply?gpu_index=0` | Apply a saved profile to hardware. `404` if missing; `500` with per-part errors if any part fails | +| `DELETE /api/profiles/{name}` | Delete a profile (also clears auto-load references) | +| `POST /api/profiles/{name}/rename` | Body `{"new_name": "..."}` | + +### Config + +| Method & path | Description | +| --- | --- | +| `GET /api/config?gpu_index=0` | `{auto_load_profile}` for the GPU | +| `POST /api/config` | Body `{"auto_load_profile": "name" \| null, "gpu_index": 0}` — set/clear the profile auto-applied at server start. Persists to `/etc/nvcurve/config.json` if present | + +### Limits + +| Method & path | Description | +| --- | --- | +| `GET /api/limits?gpu_index=0` | Power limit (current/default/min/max), `gpc_offset_mhz`, `mem_offset_mhz`, memory offset range | +| `POST /api/limits?gpu_index=0` | Body `{power_limit_w?, mem_offset_mhz?, power_cap_mode?}` — any subset. `power_cap_mode` is `"nvml"` (default) or `"ioctl"` (experimental RM power control; `409` if unsupported). Changing the memory offset may reset the curve table — the server re-applies the last known curve offsets | +| `POST /api/limits/reset?gpu_index=0` | Power limit → hardware default, memory offset → 0 | + +### Fans + +| Method & path | Description | +| --- | --- | +| `GET /api/fans?gpu_index=0` | Per-fan state: `{fan_pct, fans: [{index, fan_pct}], num_fans, min_fan_pct, max_fan_pct, fan_mode: "auto" \| "curve", curve, curve_active, fan_targets}` | +| `POST /api/fans?gpu_index=0` | Body `{"curve": [{"temp_c": 60, "fan_pct": 40}, …], "fans": [0] \| null}` — set the fan curve and start the control poller. `fans: null` drives all fans. `400` for invalid curves (e.g. non-monotonic temps) | +| `POST /api/fans/reset?gpu_index=0` | Deactivate curve control, restore automatic fan mode | +| `POST /api/fans/speed?gpu_index=0` | Body `{"fan_pct": 50, "fan": 0 \| null}` — one-shot exact speed (bypasses curve) | + +Active fan curves persist across server restarts. + +### WireView Pro II + +| Method & path | Description | +| --- | --- | +| `GET /api/wireview` | `{available, connected, info, sample}` — `available` is true when a Thermal Grizzly WireView Pro II is connected; `sample` is the most recent reading (`null` until the first) | + +### Snapshots + +| Method & path | Description | +| --- | --- | +| `GET /api/snapshots` | List saved ClockBoostTable snapshots → `[{filepath, timestamp, gpu, nonzero_offsets, size}]` | +| `POST /api/snapshot/save?gpu_index=0` | Save the current ClockBoostTable → `{ok, filepath}` | +| `POST /api/snapshot/restore?gpu_index=0` | Body `{"filepath": "…" \| null}` — restore a snapshot (most recent if omitted) | + +### Server + +| Method & path | Description | +| --- | --- | +| `POST /api/shutdown` | Gracefully stop the server process. `403` when `allow_api_shutdown: false` (recommended on shared systems — use systemd instead) | + +--- + +## WebSockets + +All three endpoints use the same handshake: connect, then send a subscribe +message. The server sends an initial state payload, then keeps streaming. + +| Endpoint | Subscribe message | Stream | +| --- | --- | --- | +| `/ws/monitor` | `{"action": "subscribe", "gpu_index": 0}` | Monitoring sample every `poll_interval_s` (default 1 s): voltage, clocks, temp, power, fans, p-state, VRAM, utilization, throttle reasons | +| `/ws/curve` | `{"action": "subscribe", "gpu_index": 0}` | Full curve state, pushed whenever the curve changes (after any write) | +| `/ws/wireview` | `{"action": "subscribe"}` | `{"type": "unavailable"}` or `{"type": "sample", "info": {...}, "sample": {...}}` per poll tick | + +Authentication (when enabled): the session token is taken from the +`Authorization: Bearer ` header or the `nvcurve_session` cookie on the +handshake. Unauthenticated connections are closed with code `1008`. + +Example (Python `websockets`): + +```python +import asyncio, json, websockets + +async def main(): + async with websockets.connect( + "ws://127.0.0.1:8042/ws/monitor", + extra_headers={"Authorization": f"Bearer {token}"}, + ) as ws: + await ws.send(json.dumps({"action": "subscribe", "gpu_index": 0})) + async for msg in ws: + sample = json.loads(msg) + print(sample["temp_c"], sample["clock_mhz"]) + +asyncio.run(main()) +``` + +--- + +## CLI + +The `nvcurve` CLI talks to this same API (default base +`http://127.0.0.1:8042`, override with `--base-url`), so any scripted use +covered by the CLI works against the API as well. See +[Usage-Guide.md](Usage-Guide.md) for CLI details. diff --git a/nvcurve/server.py b/nvcurve/server.py index d433abe..8fd4b40 100644 --- a/nvcurve/server.py +++ b/nvcurve/server.py @@ -638,7 +638,88 @@ async def lifespan(app: FastAPI): # ── App ─────────────────────────────────────────────────────────────────────── -app = FastAPI(title="nvcurve", version="0.5.0", lifespan=lifespan) +# OpenAPI tag groups, rendered as sections in /docs. +OPENAPI_TAGS = [ + { + "name": "Auth", + "description": ( + "Session management. ping, status, login and logout are public; " + "all other /api/* and /ws/* endpoints require a session once at " + "least one user is configured." + ), + }, + { + "name": "GPU", + "description": ( + "GPU discovery, static info and live monitoring. Routes accept a " + "0-based gpu_index query parameter (default 0)." + ), + }, + { + "name": "Curve", + "description": ( + "V/F curve read and write (per-point deltas, global offset, reset, " + "write-verify). Writes are capped by the server-side safety limit " + "(max_delta_khz in /etc/nvcurve/config.json)." + ), + }, + { + "name": "Profiles", + "description": ( + "Named profiles: save, apply, rename, delete. A profile captures " + "curve deltas, limits and the active fan curve." + ), + }, + { + "name": "Config", + "description": "Mutable per-GPU server configuration (auto-load profile).", + }, + { + "name": "Limits", + "description": "Power limit and memory clock offset (read/set/reset).", + }, + { + "name": "Fans", + "description": ( + "Fan curve control: set a curve, one-shot speed, reset to automatic " + "fan mode." + ), + }, + { + "name": "WireView", + "description": "Thermal Grizzly WireView Pro II sensor access.", + }, + { + "name": "Snapshots", + "description": "ClockBoostTable snapshots: list, save, restore.", + }, + { + "name": "Server", + "description": "Server lifecycle.", + }, +] + +app = FastAPI( + title="nvcurve", + version="0.5.0", + lifespan=lifespan, + description=( + "REST + WebSocket API for the nvcurve GPU V/F curve editor.\n\n" + "Base URL defaults to http://127.0.0.1:8042 (see `nvcurve serve --help`).\n\n" + "**Authentication** — with no users configured the API is open. Once at " + "least one user exists, every /api/* and /ws/* endpoint requires a " + "session: either the `nvcurve_session` cookie set by " + "`POST /api/auth/login` (browsers) or the returned token as " + "`Authorization: Bearer ` (scripts/CLI). Sessions last 24 hours.\n\n" + "**gpu_index** — GPU-selecting routes take a 0-based `gpu_index` query " + "parameter (default `0`). Unknown indexes return 404; an uninitialized " + "GPU returns 503.\n\n" + "**Errors** — JSON `{'detail': '...'}`. Write endpoints return 400 with " + "`detail.errors[]` when a delta exceeds the server-side safety cap " + "(`max_delta_khz` in /etc/nvcurve/config.json)." + ), + openapi_tags=OPENAPI_TAGS, +) app.add_middleware( CORSMiddleware, @@ -764,16 +845,29 @@ async def _run(fn, *args): return await loop.run_in_executor(None, fn, *args) +# ── Shared OpenAPI response metadata ────────────────────────────────────────── +# Reused `responses=` entries so every route documents the same error shapes. +R_UNAUTHORIZED = {"description": "No valid session (auth is enabled)."} +R_GPU_NOT_FOUND = {"description": "Unknown gpu_index."} +R_GPU_NOT_INIT = {"description": "GPU not initialized."} +R_HW_FAILED = {"description": "Hardware operation failed (NvAPI/NVML error)."} + +# Convenience bundles for routes that take a gpu_index. Annotated to match +# FastAPI's `responses` parameter type (keys may be int or str). +R_AUTH: dict[int | str, dict[str, Any]] = {401: R_UNAUTHORIZED} +R_GPU: dict[int | str, dict[str, Any]] = {404: R_GPU_NOT_FOUND, 503: R_GPU_NOT_INIT} + + # ── Auth endpoints ──────────────────────────────────────────────────────────── -@app.get("/api/ping") +@app.get("/api/ping", tags=["Auth"]) async def api_ping(): """Public liveness probe (no auth). Used by the CLI to detect a running server.""" return {"ok": True} -@app.get("/api/auth/status") +@app.get("/api/auth/status", tags=["Auth"]) async def api_auth_status(request: Request): """Report whether auth is required and whether this request is authenticated.""" cfg: Config = _state["config"] @@ -789,7 +883,15 @@ async def api_auth_status(request: Request): } -@app.post("/api/auth/login") +@app.post( + "/api/auth/login", + tags=["Auth"], + responses={ + 401: {"description": "Invalid username or password."}, + 404: {"description": "Authentication is not enabled (no users configured)."}, + 429: {"description": "Too many failed login attempts — client is temporarily locked out."}, + }, +) async def api_auth_login(req: LoginRequest, request: Request): """Authenticate with username+password. Sets a 24-hour session cookie. @@ -834,7 +936,7 @@ async def api_auth_login(req: LoginRequest, request: Request): raise HTTPException(status_code=401, detail="Invalid username or password") -@app.post("/api/auth/logout") +@app.post("/api/auth/logout", tags=["Auth"]) async def api_auth_logout(request: Request): """End the current session (idempotent).""" token = auth.extract_token(request) @@ -844,7 +946,7 @@ async def api_auth_logout(request: Request): return response -@app.get("/api/auth/users") +@app.get("/api/auth/users", tags=["Auth"], responses=R_AUTH) async def api_auth_users(): """List configured usernames (requires auth).""" cfg: Config = _state["config"] @@ -892,7 +994,7 @@ def _client_ip(request: _ClientIpSource, trusted_proxies: list[str]) -> str: # ── REST endpoints ──────────────────────────────────────────────────────────── -@app.get("/api/gpus") +@app.get("/api/gpus", tags=["GPU"]) async def api_gpus(): """List all discovered GPUs.""" from .hal.gpu import discover_gpus @@ -909,7 +1011,7 @@ async def api_gpus(): ] -@app.get("/api/gpu") +@app.get("/api/gpu", tags=["GPU"], responses={**R_AUTH, **R_GPU}) async def api_gpu(gpu_index: int = 0): """GPU info: name, driver version, VRAM.""" gpu, g_state = _require_gpu(gpu_index) @@ -924,7 +1026,7 @@ async def api_gpu(gpu_index: int = 0): } -@app.get("/api/dashboard") +@app.get("/api/dashboard", tags=["GPU"], responses={**R_AUTH, **R_GPU}) async def api_dashboard(gpu_index: int = 0): """Static GPU info for the Dashboard tab (VBIOS, CUDA cores, PCIe, BAR1, etc.). @@ -935,7 +1037,11 @@ async def api_dashboard(gpu_index: int = 0): return info -@app.get("/api/curve") +@app.get( + "/api/curve", + tags=["Curve"], + responses={**R_AUTH, **R_GPU, 500: R_HW_FAILED}, +) async def api_curve(gpu_index: int = 0): """Full CurveState: all V/F points with base freq, voltage, delta, effective freq.""" gpu, g_state = _require_gpu(gpu_index) @@ -948,7 +1054,16 @@ async def api_curve(gpu_index: int = 0): return _curve_state_dict(state) -@app.get("/api/curve/{point}") +@app.get( + "/api/curve/{point}", + tags=["Curve"], + responses={ + **R_AUTH, + **R_GPU, + 400: {"description": "Point index out of range."}, + 500: R_HW_FAILED, + }, +) async def api_curve_point(point: int, gpu_index: int = 0): """Single V/F point detail.""" gpu, g_state = _require_gpu(gpu_index) @@ -962,7 +1077,7 @@ async def api_curve_point(point: int, gpu_index: int = 0): return _vfpoint_dict(state.points[point]) -@app.get("/api/ranges") +@app.get("/api/ranges", tags=["Curve"], responses={**R_AUTH, **R_GPU, 500: R_HW_FAILED}) async def api_ranges(gpu_index: int = 0): """Clock boost domain ranges (min/max offset per domain).""" gpu, g_state = _require_gpu(gpu_index) @@ -972,7 +1087,7 @@ async def api_ranges(gpu_index: int = 0): return ranges -@app.get("/api/voltage") +@app.get("/api/voltage", tags=["Curve"], responses={**R_AUTH, **R_GPU, 500: R_HW_FAILED}) async def api_voltage(gpu_index: int = 0): """Current GPU core voltage.""" from .hal.monitoring import read_voltage @@ -984,7 +1099,7 @@ async def api_voltage(gpu_index: int = 0): return {"voltage_uv": voltage_uv, "voltage_mv": voltage_uv / 1000.0} -@app.get("/api/monitor") +@app.get("/api/monitor", tags=["GPU"], responses={**R_AUTH, **R_GPU}) async def api_monitor(gpu_index: int = 0): """One-shot monitoring snapshot: voltage, clock, temp, power, fan, p-state, VRAM, utilization.""" gpu, g_state = _require_gpu(gpu_index) @@ -992,7 +1107,7 @@ async def api_monitor(gpu_index: int = 0): return _sample_dict(sample) -@app.get("/api/snapshots") +@app.get("/api/snapshots", tags=["Snapshots"], responses=R_AUTH) async def api_snapshots(): """List saved ClockBoostTable snapshots.""" cfg: Config = _state["config"] @@ -1075,7 +1190,7 @@ def _power_cap_mode(cfg: Config, gpu_index: int) -> str: return mode if mode in ("nvml", "ioctl") else "nvml" -@app.get("/api/profiles") +@app.get("/api/profiles", tags=["Profiles"], responses=R_AUTH) async def api_profiles(gpu_index: int = 0): """List saved native profiles, the active profile name, and the auto-load profile name.""" cfg: Config = _state["config"] @@ -1089,7 +1204,7 @@ async def api_profiles(gpu_index: int = 0): } -@app.post("/api/profiles") +@app.post("/api/profiles", tags=["Profiles"], responses={**R_AUTH, **R_GPU, 500: R_HW_FAILED}) async def api_profile_save(req: ProfileSaveRequest, gpu_index: int = 0): """Save current GPU state (curve deltas + limits) as a named profile.""" gpu, g_state = _require_gpu(gpu_index) @@ -1284,7 +1399,16 @@ async def _apply_profile(name: str, gpu_index: int = 0) -> list[str]: return errs -@app.post("/api/profiles/{name}/apply") +@app.post( + "/api/profiles/{name}/apply", + tags=["Profiles"], + responses={ + **R_AUTH, + **R_GPU, + 404: {"description": "Profile not found."}, + 500: {"description": "Profile failed to apply (per-part errors in detail)."}, + }, +) async def api_profile_apply(name: str, gpu_index: int = 0): """Apply a saved profile to hardware (curve deltas + limits).""" _require_gpu(gpu_index) @@ -1303,7 +1427,11 @@ async def api_profile_apply(name: str, gpu_index: int = 0): return {"ok": True} -@app.delete("/api/profiles/{name}") +@app.delete( + "/api/profiles/{name}", + tags=["Profiles"], + responses={**R_AUTH, 404: {"description": "Profile not found."}}, +) async def api_profile_delete(name: str): """Delete a saved profile by name.""" cfg: Config = _state["config"] @@ -1322,7 +1450,15 @@ async def api_profile_delete(name: str): return {"ok": True} -@app.post("/api/profiles/{name}/rename") +@app.post( + "/api/profiles/{name}/rename", + tags=["Profiles"], + responses={ + **R_AUTH, + 400: {"description": "New name is empty."}, + 404: {"description": "Profile not found."}, + }, +) async def api_profile_rename(name: str, req: ProfileRenameRequest): """Rename a profile.""" cfg: Config = _state["config"] @@ -1344,14 +1480,14 @@ async def api_profile_rename(name: str, req: ProfileRenameRequest): return {"ok": True} -@app.get("/api/config") +@app.get("/api/config", tags=["Config"], responses=R_AUTH) async def api_config_get(gpu_index: int = 0): """Get mutable server configuration for a specific GPU.""" cfg: Config = _state["config"] return {"auto_load_profile": cfg.auto_load_profiles.get(_gpu_stable_key(gpu_index))} -@app.post("/api/config") +@app.post("/api/config", tags=["Config"], responses={**R_AUTH, 404: R_GPU_NOT_FOUND}) async def api_config_update(req: ConfigUpdateRequest): """Update mutable server configuration. Changes persist to /etc/nvcurve/config.json if present.""" if req.gpu_index not in _state["gpus"]: @@ -1366,7 +1502,7 @@ async def api_config_update(req: ConfigUpdateRequest): return {"ok": True, "auto_load_profile": cfg.auto_load_profiles.get(key)} -@app.get("/api/limits") +@app.get("/api/limits", tags=["Limits"], responses={**R_AUTH, 404: R_GPU_NOT_FOUND}) async def api_limits(gpu_index: int = 0): """Current performance limits: power and clock offsets.""" cfg: Config = _state["config"] @@ -1381,7 +1517,19 @@ async def api_limits(gpu_index: int = 0): } -@app.post("/api/limits") +@app.post( + "/api/limits", + tags=["Limits"], + responses={ + **R_AUTH, + 404: R_GPU_NOT_FOUND, + 400: {"description": "Invalid power_cap_mode (must be 'nvml' or 'ioctl')."}, + 409: { + "description": "ioctl power mode requested but not supported on this GPU/driver." + }, + 500: R_HW_FAILED, + }, +) async def api_limits_update(req: LimitsRequest, gpu_index: int = 0): """Update performance limits.""" g_state = _get_gpu_state(gpu_index) @@ -1473,7 +1621,7 @@ async def _update_offsets_and_broadcast(gpu_index: int) -> None: g_state["last_offsets"] = offsets -@app.post("/api/limits/reset") +@app.post("/api/limits/reset", tags=["Limits"], responses={**R_AUTH, 404: R_GPU_NOT_FOUND, 500: R_HW_FAILED}) async def api_limits_reset(gpu_index: int = 0): """Reset power limit to hardware default and memory clock offset to 0.""" g_state = _get_gpu_state(gpu_index) @@ -1508,7 +1656,7 @@ async def api_limits_reset(gpu_index: int = 0): # ── Fan endpoints ────────────────────────────────────────────────────────────── -@app.get("/api/fans") +@app.get("/api/fans", tags=["Fans"], responses={**R_AUTH, 404: R_GPU_NOT_FOUND}) async def api_fans(gpu_index: int = 0): """Current fan state: per-fan %, curve, and whether curve control is active.""" _get_gpu_state(gpu_index) @@ -1524,7 +1672,16 @@ async def api_fans(gpu_index: int = 0): } -@app.post("/api/fans") +@app.post( + "/api/fans", + tags=["Fans"], + responses={ + **R_AUTH, + 404: R_GPU_NOT_FOUND, + 400: {"description": "Invalid fan curve (e.g. non-monotonic temperatures)."}, + 500: {"description": "Fan control not available on this GPU."}, + }, +) async def api_fans_update(req: FanCurveRequest, gpu_index: int = 0): """Set or update the fan curve. Starts the fan control poller. @@ -1563,7 +1720,7 @@ async def api_fans_update(req: FanCurveRequest, gpu_index: int = 0): return {"ok": True} -@app.post("/api/fans/reset") +@app.post("/api/fans/reset", tags=["Fans"], responses={**R_AUTH, 404: R_GPU_NOT_FOUND}) async def api_fans_reset(gpu_index: int = 0): """Deactivate fan curve control and restore automatic fan mode.""" _get_gpu_state(gpu_index) @@ -1575,7 +1732,7 @@ async def api_fans_reset(gpu_index: int = 0): return {"ok": True} -@app.post("/api/fans/speed") +@app.post("/api/fans/speed", tags=["Fans"], responses={**R_AUTH, 404: R_GPU_NOT_FOUND, 500: R_HW_FAILED}) async def api_fans_speed(req: FanSpeedRequest, gpu_index: int = 0): """One-shot set fan(s) to an exact percentage (bypasses curve). @@ -1593,7 +1750,7 @@ async def api_fans_speed(req: FanSpeedRequest, gpu_index: int = 0): # ── WireView Pro II (Thermal Grizzly) ───────────────────────────────────────── -@app.get("/api/wireview") +@app.get("/api/wireview", tags=["WireView"], responses=R_AUTH) async def api_wireview(): """WireView Pro II availability and the most recent sensor sample. @@ -1645,7 +1802,16 @@ async def _reconcile_check(gpu_index: int) -> dict | None: } -@app.post("/api/curve/write") +@app.post( + "/api/curve/write", + tags=["Curve"], + responses={ + **R_AUTH, + **R_GPU, + 400: {"description": "Delta exceeds the server-side safety cap (max_delta_khz); detail.errors[] lists the offending points."}, + 500: R_HW_FAILED, + }, +) async def api_curve_write(req: WriteRequest, gpu_index: int = 0): """Write per-point frequency offsets. {deltas: {point_index: delta_kHz}}""" gpu, g_state = _require_gpu(gpu_index) @@ -1696,7 +1862,16 @@ async def api_curve_write(req: WriteRequest, gpu_index: int = 0): return result -@app.post("/api/curve/write/global") +@app.post( + "/api/curve/write/global", + tags=["Curve"], + responses={ + **R_AUTH, + **R_GPU, + 400: {"description": "Offset exceeds the server-side safety cap (max_delta_khz); detail.errors[] lists the offending points."}, + 500: R_HW_FAILED, + }, +) async def api_curve_write_global(req: GlobalOffsetRequest, gpu_index: int = 0): """Apply a uniform frequency offset to all curve points.""" gpu, g_state = _require_gpu(gpu_index) @@ -1746,7 +1921,7 @@ async def api_curve_write_global(req: GlobalOffsetRequest, gpu_index: int = 0): return result -@app.post("/api/curve/reset") +@app.post("/api/curve/reset", tags=["Curve"], responses={**R_AUTH, **R_GPU, 500: R_HW_FAILED}) async def api_curve_reset(gpu_index: int = 0): """Reset all frequency offsets to zero.""" gpu, g_state = _require_gpu(gpu_index) @@ -1777,7 +1952,16 @@ async def api_curve_reset(gpu_index: int = 0): return result -@app.post("/api/curve/verify") +@app.post( + "/api/curve/verify", + tags=["Curve"], + responses={ + **R_AUTH, + **R_GPU, + 400: {"description": "Delta exceeds the server-side safety cap (max_delta_khz); detail.errors[] lists the offending points."}, + 500: R_HW_FAILED, + }, +) async def api_curve_verify(req: VerifyRequest, gpu_index: int = 0): """Write-verify-read cycle. Returns per-point match results and collateral changes.""" gpu, g_state = _require_gpu(gpu_index) @@ -1847,7 +2031,14 @@ async def api_curve_verify(req: VerifyRequest, gpu_index: int = 0): } -@app.post("/api/shutdown") +@app.post( + "/api/shutdown", + tags=["Server"], + responses={ + **R_AUTH, + 403: {"description": "API shutdown disabled (allow_api_shutdown: false in config)."}, + }, +) async def api_shutdown(): """Gracefully shut down the server process. @@ -1869,7 +2060,7 @@ async def api_shutdown(): return {"ok": True} -@app.post("/api/snapshot/save") +@app.post("/api/snapshot/save", tags=["Snapshots"], responses={**R_AUTH, **R_GPU, 500: R_HW_FAILED}) async def api_snapshot_save(gpu_index: int = 0): """Save a ClockBoostTable snapshot.""" gpu, g_state = _require_gpu(gpu_index) @@ -1882,7 +2073,7 @@ async def api_snapshot_save(gpu_index: int = 0): return {"ok": True, "filepath": path} -@app.post("/api/snapshot/restore") +@app.post("/api/snapshot/restore", tags=["Snapshots"], responses={**R_AUTH, **R_GPU, 500: R_HW_FAILED}) async def api_snapshot_restore(req: SnapshotRestoreRequest, gpu_index: int = 0): """Restore a ClockBoostTable snapshot. Uses most recent if filepath not specified.""" gpu, g_state = _require_gpu(gpu_index) @@ -2087,7 +2278,7 @@ if os.path.isdir(os.path.join(_dist_dir, "assets")): ) -@app.get("/{catchall:path}") +@app.get("/{catchall:path}", include_in_schema=False) async def serve_spa(catchall: str): if catchall.startswith(("api/", "ws/")): raise HTTPException(status_code=404, detail="Not Found")