Changelog
Notable changes to the benchmark, dataset, and site.
15 October 2025

Model additions

Added Claude Haiku 4.5 (standard and Thinking variants) to the benchmark.

29 September 2025

Model additions

Added Claude Sonnet 4.5 (standard and Thinking variants) and Grok 4 Fast to the benchmark.

23 September 2025

Model additions

Added DeepSeek V3.1‑Terminus and GPT‑5 Codex (High).

17 September 2025

Initial public release

First public release of CompileBench: 21 models evaluated across 15 tasks. Read the announcement blog post: Introducing CompileBench.