What Makes Mathematicians Believe Unproven Mathematical Statements?
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
RealMath: A Continuous Benchmark for Evaluating Language Models on Research-Level Mathematics
Measuring Mathematical Problem Solving With the MATH Dataset
Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap
Mathematical Reasoning via Self-supervised Skip-tree Training
What is the point of computers? A question for pure mathematicians