15 October 2025
Model additions
Added Claude Haiku 4.5 (standard and Thinking variants) to the benchmark.
Added Claude Haiku 4.5 (standard and Thinking variants) to the benchmark.
Added Claude Sonnet 4.5 (standard and Thinking variants) and Grok 4 Fast to the benchmark.
Added DeepSeek V3.1‑Terminus and GPT‑5 Codex (High).
First public release of CompileBench: 21 models evaluated across 15 tasks. Read the announcement blog post: Introducing CompileBench.