← All publications

BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software

Zehua Zhang, A. P. Bajaj, Divij Handa, S. Liu, A. S. Raj, H. Chen, H. Wang, Y. Liu, et al.

DL4C @ NeurIPS 2025 2025 Evaluation & Safety

Abstract

A challenging, realistic benchmark of diverse real-world open-source software for evaluating LLM agents on automatically compiling projects, plus OSS-Build-Agent, a strong baseline with an enhanced build-instruction retrieval module.

AgentsEvaluation