Can NLP Models Correctly Reason Over Contexts That Break the Common Assumptions?
Abstract
Systematically constructs contexts that break common assumptions and shows that, while models reason well over assumption-following contexts, performance drops by up to 20% when those assumptions are broken.