Files

1.4 KiB

MCQ Benchmark

LLM benchmark for evaluating models on multiple-choice questions from a parquet dataset (financial reasoning).

Features

  • Checkpoint/Resume — progress is saved after each question. Stop and resume anytime.
  • Live Status — real-time progress with running accuracy percentage.
  • Configurable — YAML config for model endpoint, parameters, and benchmark settings.

Setup

  1. Copy the example config:

    cp config.example.yaml config.yaml
    
  2. Edit config.yaml with your model endpoint and parameters.

Usage

# Run benchmark (prompts to resume if checkpoint exists)
python3 benchmark.py

# Auto-resume from latest checkpoint
python3 benchmark.py --resume

# Start fresh, ignoring existing checkpoint
python3 benchmark.py --fresh

# Use custom config
python3 benchmark.py --config my_config.yaml

Output

Results are saved to runs/<model_name>_YYYYMMDD_HHMMSS.json with full details:

  • Overall accuracy and timing
  • Per-question results (correct, incorrect, timeouts)
  • Model configuration used

Checkpoints (*.checkpoint.json) are stored in runs/ and auto-deleted on completion.

Live Output Example

  [3048/5074] what is AAPL expecting in return for Q4 2016 based on this sheet:...  CORRECT  [67.4%]
  [3049/5074] calculate the total revenue for the fiscal year 2015...  WRONG (expected: B, got: A)  [67.4%]

License

Apache 2.0