PhD researcher · LLM agents · open to collaborations & internships

Divij Handa

research focus ▸ .

I'm a PhD researcher in Computer Science at Arizona State University. My core research is on LLM agents — how they're trained, how they reason and act at test time, and how we evaluate and keep them safe. My work spans these stages, with papers at ICLR, ACL, NAACL, EMNLP and NeurIPS workshops.

scroll

Core research · LLM agents

The agent lifecycle

Selected work

Featured publications

All publications →
[001]

GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time

Divij Handa, Mihir Parmar, Aswin RRV, Md Nayem Uddin, Hamid Palangi, Chitta Baral

ICLR 2026 Conference

TL;DR Makes inference-time sampling produce genuinely diverse solutions — +21.6% pass@50.

Decouples exploration from generation at inference time so repeated sampling produces genuinely diverse candidate solutions, improving pass@50 by ~21.6% over standard repeated sampling.
[002]

When “Competency” in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphers

Divij Handa, Zehua Zhang, Amir Saeidi, Shrinidhi Kumbhar, Md Nayem Uddin, Aswin RRV, Chitta Baral

NeurIPS 2025 Workshop (Reliable ML from Unreliable Data) Workshop

TL;DR Stronger reasoning makes LLMs easier to jailbreak — via custom ciphers (ACE/LACE).

Shows that as LLMs get better at reasoning they become more susceptible to novel jailbreaks. Introduces ACE and LACE — attacks that encode malicious queries with custom and layered ciphers — and CipherBench to measure cipher-decoding ability.
[003]

OptAgent: Optimizing Query Rewriting for E-Commerce via Multi-Agent Simulation

Divij Handa, David Blincoe, Orson Adams, Yinlin Fu

Preprint arXiv

TL;DR Uses multi-agent shopper simulation as a reward to optimize e-commerce query rewriting.

Combines multi-agent simulation with an evolutionary algorithm for query rewriting: multiple LLM agents act as simulated shoppers, and their averaged scores form a dynamic reward that iteratively refines the query — improving fitness by ~22%.
[004]

ActionReasoningBench: Reasoning about Actions with and without Ramification Constraints

Divij Handa, Pavel Dolin, Shrinidhi Kumbhar, Tran Cao Son, Chitta Baral

ICLR 2025 Conference

TL;DR A diagnostic benchmark for reasoning about actions, change and ramifications.

A diagnostic benchmark spanning eight domains that evaluates LLMs on six dimensions of reasoning about actions and change, including new ramification constraints for indirect effects. State-of-the-art models struggle across all dimensions, especially ramifications.

Path

Experience

Adobe May 2026 — Aug 2026
Research Scientist Intern · College Park, MD

Building a multi-agent system that turns any uploaded document into an interactive webpage, and an optimization pipeline that autonomously improves multi-agent systems — tuning their context, topology, harness, cost, and latency.

Etsy May 2025 — Aug 2025
AI Research Intern · Brooklyn, NY

Designed an agentic framework for e-commerce query rewriting via a genetic algorithm (OptAgent), using a multi-agent shopper simulation for offline semantic evaluation — a 3.28% gain over test-time baselines. Scaled data collection and analysis over large query logs with BigQuery.

Arizona State University Present Jan 2024 — now
Research Assistant · Tempe, AZ

Built an LLM system translating decompiler output into high-level source code for reverse engineering; fine-tuned with SFT + RL (GRPO) and a graph-edit-distance reward for structural fidelity, achieving a 2.75× improvement on held-out repositories.

Arizona State University Aug 2022 — May 2023
Graduate Teaching Assistant · Tempe, AZ

Teaching assistant for graduate machine-learning coursework.

Nagarro Oct 2020 — Apr 2021
Associate Engineer · Gurugram, India

Built and maintained enterprise web applications.

Nagarro Jun 2019 — Aug 2019
Software Engineering Intern · Gurugram, India

Full-stack web development internship.

Updates

News