03 / Research output

Publications

Peer-reviewed papers, workshop contributions and preprints. 10 total — expand Abstract or BibTeX on any entry.

2025

[001]

BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software

Zehua Zhang, A. P. Bajaj, Divij Handa, S. Liu, A. S. Raj, H. Chen, H. Wang, Y. Liu, et al.

DL4C @ NeurIPS 2025 Workshop

TL;DR A realistic benchmark for LLM agents that compile real-world open-source software.

A challenging, realistic benchmark of diverse real-world open-source software for evaluating LLM agents on automatically compiling projects, plus OSS-Build-Agent, a strong baseline with an enhanced build-instruction retrieval module.
[002]

OptAgent: Optimizing Query Rewriting for E-Commerce via Multi-Agent Simulation

Divij Handa, David Blincoe, Orson Adams, Yinlin Fu

Preprint arXiv

TL;DR Uses multi-agent shopper simulation as a reward to optimize e-commerce query rewriting.

Combines multi-agent simulation with an evolutionary algorithm for query rewriting: multiple LLM agents act as simulated shoppers, and their averaged scores form a dynamic reward that iteratively refines the query — improving fitness by ~22%.