LLM Benchmarker

Dashboard New Test Run

Test Run #5 Analysis

Comparing model performance for the GPQA 2026 benchmark.

Global Filters

Languages

Models

Tags

Overall Avg. Score

0.515

Best Model

GPT O3

Highest Model Score

0.554

Model Scores per Language

© 2026 LLM Benchmarker. All rights reserved.