Test Run #8 Analysis
Comparing model performance for the GPQA 2026 benchmark.
Global Filters
Languages
Models
Tags
Overall Avg. Score
0.470
Best Model
GPT O3
Highest Model Score
0.493
Comparing model performance for the GPQA 2026 benchmark.
0.470
GPT O3
0.493