Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models
Preprint arXiv
TL;DR Mid-training on self-generated data boosts downstream RL and reasoning.
03 / Research output
Peer-reviewed papers, workshop contributions and preprints. 10 total — expand Abstract or BibTeX on any entry.
2026
Preprint arXiv
TL;DR Mid-training on self-generated data boosts downstream RL and reasoning.
✓ ICLR 2026 Conference
TL;DR Makes inference-time sampling produce genuinely diverse solutions — +21.6% pass@50.
2025
✓ DL4C @ NeurIPS 2025 Workshop
TL;DR A realistic benchmark for LLM agents that compile real-world open-source software.
✓ NeurIPS 2025 Workshop (Reliable ML from Unreliable Data) Workshop
TL;DR Stronger reasoning makes LLMs easier to jailbreak — via custom ciphers (ACE/LACE).
✓ EMNLP 2025 Conference
TL;DR Teaches models to self-reflect through teacher feedback, without distillation.
Preprint arXiv
TL;DR Uses multi-agent shopper simulation as a reward to optimize e-commerce query rewriting.
✓ ICLR 2025 Conference
TL;DR A diagnostic benchmark for reasoning about actions, change and ramifications.
✓ ACL 2025 Conference
TL;DR A contamination-free benchmark that forces genuine temporal reasoning, not memorization.
✓ Findings of NAACL 2025 Conference
TL;DR Goal-driven, constraint-guided LLM agents that generate materials-science hypotheses.
2023
Preprint arXiv
TL;DR Models reason well — until you break the common assumptions behind a context.