03 / Research output

Publications

Peer-reviewed papers, workshop contributions and preprints. 10 total — expand Abstract or BibTeX on any entry.

2026

[002]

GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time

Divij Handa, Mihir Parmar, Aswin RRV, Md Nayem Uddin, Hamid Palangi, Chitta Baral

ICLR 2026 Conference

TL;DR Makes inference-time sampling produce genuinely diverse solutions — +21.6% pass@50.

Decouples exploration from generation at inference time so repeated sampling produces genuinely diverse candidate solutions, improving pass@50 by ~21.6% over standard repeated sampling.

2025

[001]

BuildBench: Benchmarking LLM Agents on Compiling Real-World Open-Source Software

Zehua Zhang, A. P. Bajaj, Divij Handa, S. Liu, A. S. Raj, H. Chen, H. Wang, Y. Liu, et al.

DL4C @ NeurIPS 2025 Workshop

TL;DR A realistic benchmark for LLM agents that compile real-world open-source software.

A challenging, realistic benchmark of diverse real-world open-source software for evaluating LLM agents on automatically compiling projects, plus OSS-Build-Agent, a strong baseline with an enhanced build-instruction retrieval module.
[002]

When “Competency” in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphers

Divij Handa, Zehua Zhang, Amir Saeidi, Shrinidhi Kumbhar, Md Nayem Uddin, Aswin RRV, Chitta Baral

NeurIPS 2025 Workshop (Reliable ML from Unreliable Data) Workshop

TL;DR Stronger reasoning makes LLMs easier to jailbreak — via custom ciphers (ACE/LACE).

Shows that as LLMs get better at reasoning they become more susceptible to novel jailbreaks. Introduces ACE and LACE — attacks that encode malicious queries with custom and layered ciphers — and CipherBench to measure cipher-decoding ability.
[003]

ThinkTuning: Instilling Cognitive Reflections without Distillation

Aswin RRV, Jacob Dineen, Divij Handa, Md Nayem Uddin, Mihir Parmar, Chitta Baral, Ben Zhou

EMNLP 2025 Conference

TL;DR Teaches models to self-reflect through teacher feedback, without distillation.

A GRPO-based interactive training method where a teacher model gives corrective feedback on a student model’s rollouts, instilling self-reflective reasoning without distillation.
[004]

OptAgent: Optimizing Query Rewriting for E-Commerce via Multi-Agent Simulation

Divij Handa, David Blincoe, Orson Adams, Yinlin Fu

Preprint arXiv

TL;DR Uses multi-agent shopper simulation as a reward to optimize e-commerce query rewriting.

Combines multi-agent simulation with an evolutionary algorithm for query rewriting: multiple LLM agents act as simulated shoppers, and their averaged scores form a dynamic reward that iteratively refines the query — improving fitness by ~22%.
[005]

ActionReasoningBench: Reasoning about Actions with and without Ramification Constraints

Divij Handa, Pavel Dolin, Shrinidhi Kumbhar, Tran Cao Son, Chitta Baral

ICLR 2025 Conference

TL;DR A diagnostic benchmark for reasoning about actions, change and ramifications.

A diagnostic benchmark spanning eight domains that evaluates LLMs on six dimensions of reasoning about actions and change, including new ramification constraints for indirect effects. State-of-the-art models struggle across all dimensions, especially ramifications.
[006]

UnSeenTimeQA: Time-Sensitive Question-Answering Beyond LLMs' Memorization

Md Nayem Uddin, Amir Saeidi, Divij Handa, Agastya Seth, Tran Cao Son, Eduardo Blanco, Steven R. Corman, Chitta Baral

ACL 2025 Conference

TL;DR A contamination-free benchmark that forces genuine temporal reasoning, not memorization.

A contamination-free, time-sensitive QA benchmark built on synthetic facts, forcing models to do genuine temporal reasoning rather than recalling pre-training knowledge.

2023

[001]

Can NLP Models Correctly Reason Over Contexts That Break the Common Assumptions?

Neeraj Varshney, Mihir Parmar, Nisarg Patel, Divij Handa, Sayantan Sarkar, Man Luo, Chitta Baral

Preprint arXiv

TL;DR Models reason well — until you break the common assumptions behind a context.

Systematically constructs contexts that break common assumptions and shows that, while models reason well over assumption-following contexts, performance drops by up to 20% when those assumptions are broken.