SAFE: Enhancing Mathematical Reasoning in Large Language Models via Retrospective Step-aware Formal Verification

Published in ACL 2025, 2025

SAFE uses Lean 4 proofs to identify hallucinations in natural-language mathematical reasoning at the step level.