TaoAnalysisBench: Do Automated Theorem Prover LLMs Generalize Beyond MathLib?
Alexander K. Taylor, Junyi Zhang, Ethan Ji, Vigyan Sahai, Haikang Deng, Yuanzhou Chen, Yifan Yuan, Di Wu, Jia-Chen Gu, Kai-Wei Chang, Nanyun Peng, Amit Sahai, and Wei Wang, in EMNLP-Findings, 2026.
Abstract
Bib Entry
@inproceedings{taylor2026taobench,
title = {TaoAnalysisBench: Do Automated Theorem Prover LLMs Generalize Beyond MathLib?},
author = {Taylor, Alexander K and Zhang, Junyi and Ji, Ethan and Sahai, Vigyan and Deng, Haikang and Chen, Yuanzhou and Yuan, Yifan and Wu, Di and Gu, Jia-Chen and Chang, Kai-Wei and Peng, Nanyun and Sahai, Amit and Wang, Wei},
booktitle = {EMNLP-Findings},
year = {2026}
}
Related Publications
-
OpenThoughts: Data Recipes for Reasoning Models, ICLR, 2026
-
Gold-Medal-Level Olympiad Geometry Solving with Efficient Heuristic Auxiliary Constructions, ACL, 2026
-
OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles, NeurIPS, 2025
-
When To Solve, When To Verify: Compute-Optimal Problem Solving and Generative Verification for LLM Reasoning, COLM 2025, 2025
-
MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?, ECCV, 2024
-
MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts, ICLR, 2024