BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software
✓ DL4C @ NeurIPS 2025 Workshop
TL;DR A realistic benchmark for LLM agents that compile real-world open-source software.
03 / Research output
Peer-reviewed papers, workshop contributions and preprints. 10 total — expand Abstract or BibTeX on any entry.
2025
✓ DL4C @ NeurIPS 2025 Workshop
TL;DR A realistic benchmark for LLM agents that compile real-world open-source software.
Preprint arXiv
TL;DR Uses multi-agent shopper simulation as a reward to optimize e-commerce query rewriting.
✓ Findings of NAACL 2025 Conference
TL;DR Goal-driven, constraint-guided LLM agents that generate materials-science hypotheses.