- Entangled skills overestimate ability. Popular benchmarks (ReClor, LogiQA) mix commonsense and other reasoning with logic, so they can overestimate pure logical-reasoning ability. Prompting a model to avoid logical reasoning can paradoxically improve its score on them, but not on DivLogicEval.
- Counterintuitive, diverse statements. DivLogicEval composes natural sentences from SNLI/MNLI into propositional logic problems connected in a deliberately counterintuitive way, isolating logic from pretraining shortcuts.
- A debiasing metric. We propose PartialCircular alongside Accuracy and Circular to mitigate the bias and randomness inherent in LLMs, giving a cleaner ranking of models.

