Test Run #10 Analysis
Comparing model performance for the GPQA 2026 benchmark.
Global Filters
Languages
Models
Tags
Overall Avg. Score
0.428
Best Model
Claude 4 Sonnet
Highest Model Score
0.505
Comparing model performance for the GPQA 2026 benchmark.
0.428
Claude 4 Sonnet
0.505