03 / Research output

Publications

Peer-reviewed papers, workshop contributions and preprints. 10 total — expand Abstract or BibTeX on any entry.

2026

[001]

GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time

Divij Handa, Mihir Parmar, Aswin RRV, Md Nayem Uddin, Hamid Palangi, Chitta Baral

ICLR 2026 Conference

TL;DR Makes inference-time sampling produce genuinely diverse solutions — +21.6% pass@50.

Decouples exploration from generation at inference time so repeated sampling produces genuinely diverse candidate solutions, improving pass@50 by ~21.6% over standard repeated sampling.

2025

[001]

ThinkTuning: Instilling Cognitive Reflections without Distillation

Aswin RRV, Jacob Dineen, Divij Handa, Md Nayem Uddin, Mihir Parmar, Chitta Baral, Ben Zhou

EMNLP 2025 Conference

TL;DR Teaches models to self-reflect through teacher feedback, without distillation.

A GRPO-based interactive training method where a teacher model gives corrective feedback on a student model’s rollouts, instilling self-reflective reasoning without distillation.
[002]

ActionReasoningBench: Reasoning about Actions with and without Ramification Constraints

Divij Handa, Pavel Dolin, Shrinidhi Kumbhar, Tran Cao Son, Chitta Baral

ICLR 2025 Conference

TL;DR A diagnostic benchmark for reasoning about actions, change and ramifications.

A diagnostic benchmark spanning eight domains that evaluates LLMs on six dimensions of reasoning about actions and change, including new ramification constraints for indirect effects. State-of-the-art models struggle across all dimensions, especially ramifications.
[003]

UnSeenTimeQA: Time-Sensitive Question-Answering Beyond LLMs' Memorization

Md Nayem Uddin, Amir Saeidi, Divij Handa, Agastya Seth, Tran Cao Son, Eduardo Blanco, Steven R. Corman, Chitta Baral

ACL 2025 Conference

TL;DR A contamination-free benchmark that forces genuine temporal reasoning, not memorization.

A contamination-free, time-sensitive QA benchmark built on synthetic facts, forcing models to do genuine temporal reasoning rather than recalling pre-training knowledge.

2023

[001]

Can NLP Models Correctly Reason Over Contexts That Break the Common Assumptions?

Neeraj Varshney, Mihir Parmar, Nisarg Patel, Divij Handa, Sayantan Sarkar, Man Luo, Chitta Baral

Preprint arXiv

TL;DR Models reason well — until you break the common assumptions behind a context.

Systematically constructs contexts that break common assumptions and shows that, while models reason well over assumption-following contexts, performance drops by up to 20% when those assumptions are broken.