03 / Research output

Publications

Peer-reviewed papers, workshop contributions and preprints. 10 total — expand Abstract or BibTeX on any entry.

2025

[001]

BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software

Zehua Zhang, A. P. Bajaj, Divij Handa, S. Liu, A. S. Raj, H. Chen, H. Wang, Y. Liu, et al.

DL4C @ NeurIPS 2025 Workshop

TL;DR A realistic benchmark for LLM agents that compile real-world open-source software.

A challenging, realistic benchmark of diverse real-world open-source software for evaluating LLM agents on automatically compiling projects, plus OSS-Build-Agent, a strong baseline with an enhanced build-instruction retrieval module.
[002]

ActionReasoningBench: Reasoning about Actions with and without Ramification Constraints

Divij Handa, Pavel Dolin, Shrinidhi Kumbhar, Tran Cao Son, Chitta Baral

ICLR 2025 Conference

TL;DR A diagnostic benchmark for reasoning about actions, change and ramifications.

A diagnostic benchmark spanning eight domains that evaluates LLMs on six dimensions of reasoning about actions and change, including new ramification constraints for indirect effects. State-of-the-art models struggle across all dimensions, especially ramifications.
[003]

UnSeenTimeQA: Time-Sensitive Question-Answering Beyond LLMs' Memorization

Md Nayem Uddin, Amir Saeidi, Divij Handa, Agastya Seth, Tran Cao Son, Eduardo Blanco, Steven R. Corman, Chitta Baral

ACL 2025 Conference

TL;DR A contamination-free benchmark that forces genuine temporal reasoning, not memorization.

A contamination-free, time-sensitive QA benchmark built on synthetic facts, forcing models to do genuine temporal reasoning rather than recalling pre-training knowledge.