Share this page:

TaoAnalysisBench: Do Automated Theorem Prover LLMs Generalize Beyond MathLib?

Alexander K. Taylor, Junyi Zhang, Ethan Ji, Vigyan Sahai, Haikang Deng, Yuanzhou Chen, Yifan Yuan, Di Wu, Jia-Chen Gu, Kai-Wei Chang, Nanyun Peng, Amit Sahai, and Wei Wang, in EMNLP-Findings, 2026.

Download the full text


Abstract


Bib Entry

@inproceedings{taylor2026taobench,
  title = {TaoAnalysisBench: Do Automated Theorem Prover LLMs Generalize Beyond MathLib?},
  author = {Taylor, Alexander K and Zhang, Junyi and Ji, Ethan and Sahai, Vigyan and Deng, Haikang and Chen, Yuanzhou and Yuan, Yifan and Wu, Di and Gu, Jia-Chen and Chang, Kai-Wei and Peng, Nanyun and Sahai, Amit and Wang, Wei},
  booktitle = {EMNLP-Findings},
  year = {2026}
}

Related Publications

  1. OpenThoughts: Data Recipes for Reasoning Models, ICLR, 2026
  2. Gold-Medal-Level Olympiad Geometry Solving with Efficient Heuristic Auxiliary Constructions, ACL, 2026
  3. OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles, NeurIPS, 2025
  4. When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoning, COLM 2025, 2025
  5. MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?, ECCV, 2024
  6. MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts, ICLR, 2024