- Sept. 9, 2026, 3:30 pm US/Central
- Moritz Munchmeyer, University of Wisconsin - Madison
Large language models are increasingly being used by physicists to assist with mathematical reasoning at the research frontier. In this talk, I will briefly introduce TPBench, a benchmark dataset designed to evaluate and improve AI models on theoretical-physics reasoning tasks. I will then discuss how test-time scaling and symbolic verification can improve both performance and reliability. Despite recent progress, current LLMs have limitations in robustness and depth of understanding, which limits their ability to generate novel research. A promising path toward overcoming these limitations is to use models on problems that are verifiable, or that have a clearly defined performance metric. Many algorithmic tasks in theoretical and computational physics naturally admit objective evaluation. I will describe our MadEvolve project, which combines an outer loop of LLM-based conceptual search with an inner loop of conventional numerical optimization. We show that this framework can improve several cosmological algorithms beyond human-designed baselines. Finally, I will show our recent work on fine-tuning LLMs for theoretical physics using academia-scale computing resources.
