BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
✓ DL4C @ NeurIPS 2025 Workshop
TL;DR A realistic benchmark for LLM agents that compile real-world open-source software.
03 / Research output
Peer-reviewed papers, workshop contributions and preprints. 10 total — expand Abstract or BibTeX on any entry.
2025
✓ DL4C @ NeurIPS 2025 Workshop
TL;DR A realistic benchmark for LLM agents that compile real-world open-source software.
✓ ICLR 2025 Conference
TL;DR A diagnostic benchmark for reasoning about actions, change and ramifications.
✓ ACL 2025 Conference
TL;DR A contamination-free benchmark that forces genuine temporal reasoning, not memorization.