[001]
When “Competency” in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphers
✓ NeurIPS 2025 Workshop (Reliable ML from Unreliable Data) Workshop
TL;DR Stronger reasoning makes LLMs easier to jailbreak — via custom ciphers (ACE/LACE).
Shows that as LLMs get better at reasoning they become more susceptible to novel jailbreaks. Introduces ACE and LACE — attacks that encode malicious queries with custom and layered ciphers — and CipherBench to measure cipher-decoding ability.