Physical Sciences › Computer Science › Computational Theory and Mathematics
Mathematics, Computing, and Information Processing
158 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- United States54% · 44 papers
- China37% · 30 papers
- Canada13% · 11 papers
- Switzerland9.8% · 8 papers
- South Korea8.5% · 7 papers
- France7.3% · 6 papers
- United Kingdom6.1% · 5 papers
- Bulgaria4.9% · 4 papers
Across 82 papers on this subject with at least one lab located. 34 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- Objects Without Morphisms: What LLMs for Mathematics Do Not Represent
Yanli Wang, Suijin Wang, Xiaopeng Yuan, Haohan Wang · 5 October 2026
Large language models (LLMs) have reached expert-level performance on competition mathematics largely through the volume of search placed around them: candidate solutions are sampled in quantity and retained only when an external criterion accepts them. Such a procedure improves the outcome that sur…
- Counterexample Generation via Per-Theorem Symbolic Verifiers: When Imitation Hurts and Reinforcement Repairs
Omar Farouk Zouak, Houssam Eddine Boukhalfa, Soumaya Lakehal, Shiv Katiyar, Samia Nefti-Meziani · 5 October 2026
Large language models often solve a theorem forward yet fail to disprove a closely related false one: a falsification gap that supervised fine-tuning does not close and can actively worsen. We frame counterexample generation as constrained witness emission against a deterministic per-theorem Python …
- Continual Graph Memory for Mathematical Research Agents
Junyi Zhang, Jinxi Yu, Eric Hanchen Jiang, Jiachen Lu, Zhi Zhang, Xinjie He, Hyunsik Chae, Ethan Ji, Alexander K Taylor, Vigyan Sahai, Yiwen Kou, Kai-Wei Chang, Raghu Meka, Nanyun Peng, Amit Sahai, Terence Tao, Wei Wang · 5 October 2026
Using frontier agent harnesses to tackle mathematical research problems has emerged as an effective means of advancing mathematics. However, solving frontier problems in mathematics may require a massive number of agents working in parallel for extended periods to construct proofs, thereby generatin…
- Mathematical Transfer in LLMs Follows Reasoning Approach More Than Topic
Sajad Goudarzi, Samaneh Zamanifard, Seyed Amin Seyed Haeri, Moloud Nasiri, Hamed Rahimian · 2 October 2026
When selecting mathematical training data for LLMs, a natural organizing principle is topic: probability examples for probability targets. An alternative is reasoning approach: worked solutions that share a solution method with the target, even when the mathematical domain differs. We ask which rela…
- ANI: Adaptive Numerical Injection for Unifying Semantic and Arithmetic Representations in Numerical Reasoning
Jinsung Jeon, Seung-won Hwang · 1 October 2026
Precise numerical reasoning with Large Language Models (LLMs) is essential for expanding their applicability to complex real-world tasks. However, text-based tokenization often fragments numbers, significantly hindering precise arithmetic reasoning. Meanwhile, numerical embeddings, despite arithmeti…
- Fyan: A Human--AI Harness with Semantic Auditing for Document-Level Formalization
Wei Zhao, Yangshuo Zou, Chengxiang Ding, Yifan Wu, Xuchuan Wang, Zimu Mao, Lei Zhang, Tao Luo · 1 October 2026
We present FYAN, a human--AI harness for document-level mathematical formalization. Rather than treating theorems in isolation, FYAN coordinates an end-to-end workflow spanning specification, proof planning, logical review, Lean proof construction, knowledge curation, and validation, with support fo…
- Recipe-Matching, Not Equivalence
Ali Habibullah, Mohammad Alshiekh, Yazan Alshoibi, Salman Khan, Naeemullah Khan · 30 September 2026
MathNet-Retrieve asks a retriever to find, for a math problem, a document stating the same problem. An LLM under one fixed prompt writes each gold document and its near-miss distractors; LLM judges filter them. We call this procedure the "recipe", training on pairs built the same way "recipe-matchin…
- LANTERN: Illuminating Hidden Mathematical Knowledge in Language Models
Pavel Tikhonov, Elena Tutubalina, Ivan Oseledets, Dmitry I. Ignatov, Mikhail Seleznyov · 30 September 2026
Language models can now prove theorems, but people still decide which problems to pursue. We ask whether a model's internal representations can help identify promising mathematical connections. We develop LANTERN, a fast, cost-efficient pipeline that uses a classifier over pretrained-model activatio…
- Solving Every Step Is Not Enough: Milestone Oracles Reveal a Composition Gap in LLM Math Reasoning
Zhuohan Wang, Haoran Ma, Tianyu Wu, Yuanlin Duan, Zichun Liao, Jieming Yu · 30 September 2026
Large language models (LLMs) can solve every intermediate step of a multi-step math problem on its own and still fail the full problem, even when given a roadmap of the steps and all of their answers. We introduce OracleLadder, a diagnostic evaluation that locates where LLM math reasoning fails by g…
- When Does Structured Knowledge Help Neural Theorem Proving?
Sareh Nabi, Roland Vogl, Marzieh Nabi · 29 September 2026
Does structured mathematical knowledge help LLMs prove theorems in Lean 4? If so, for which models, and does the answer vary by problem? Formal libraries such as Mathlib encode 285,000+ verified theorems with syntactic dependencies, but the semantic layer mathematicians rely on for discovery (analog…
- Math Reasoning in LLMs is Organized by Approach, Not Topic
Sajad Goudarzi, Samaneh Zamanifard, Moloud Nasiri, Hamed Rahimian · 25 September 2026
Mathematical reasoning benchmarks are typically organized by topic, but language models may organize their internal computation by reusable reasoning approach instead. In this paper, we investigate whether open math-capable LLMs organize internally by topical sub-skill or by reasoning approach, and …
- Learning to Discover Interesting Mathematics
Niket Patel, Ahmad Rammal, Amaury Hayat, Remi Munos, Julia Kempe · 25 September 2026
Recently, Large Language Models (LLMs) have been increasingly able to solve advanced mathematical problems, including many that have been open for decades. This opens the door to expansion of mathematical knowledge at unprecedented scale. Yet, while LLMs may be able to conjecture and prove more and …
- PermuFormer: Multi-Task Pretraining for Permutation Representation in Algebraic Combinatorics
Henry Kvinge · 23 September 2026
Diverse pretraining has been shown to be an effective method for learning reusable, domain-aware representations that provide a starting point for fine-tuning on downstream tasks. While much of the excitement in AI for math has been concentrated in the use of frontier reasoning models to solve well-…
- Same Quantity, Different Answer: Numerical Representation Invariance in Language Models
Ephraim Atta-Duncan · 23 September 2026
Numerically equivalent word problems should yield the same canonical answer whether a quantity is written as a decimal, fraction, percentage, number word, scientific notation, or an exactly converted unit. We generate 3,600 exact-rational problems and 8,600 prompts spanning five identity-preserving …
- FrontierMath Erd\H{o}s
Tom Adamczewski (Epoch AI), Thomas F. Bloom (University of Manchester) · 23 September 2026
We introduce FrontierMath Erd\H{o}s (FME), a benchmark of 68 Erd\H{o}s problems that are open as of August 2026. To solve a task in FME, AI systems must resolve (prove or disprove) one of the 68 conjectures in the proof assistant Lean. Our 68 problems were selected by the second author among 652 ope…
- Math2Visual-X: A Modular Framework for Pedagogically Aligned Lower-Primary Math Visuals Generation
H. D. E. Maduranga, S. K. Munasinghe, K. P. T. I. Weerasekara, Surangika Ranathunga, Nisansa de Silva · 22 September 2026
Visual representations can help lower-primary learners understand Math Word Problems, but generating classroom-usable visuals remains difficult. Existing symbolic systems are controllable but limited in coverage, while end-to-end text-to-image systems often fail to satisfy exact mathematical constra…
- LLM-Based FORM Code Generation with Verification-Driven Fine-Tuning
Bakar Chargeishvili · 22 September 2026
FORM is a domain-specific symbolic manipulation language widely used in particle physics for processing the very large algebraic expressions arising from multi-loop Feynman diagram calculations. Despite its central role in precision theoretical physics, no artificial-intelligence tooling exists, to …
- Euston: Training Away Mathematical Sycophancy Without Losing the Mathematics
Zehua Cheng, Wei Dai, Jiahao Sun · 22 September 2026
Reasoning language models are trained to produce solutions, not to refuse them, and this bias persists when the problem they are handed is false. Asked to prove a corrupted theorem, a strong model will typically comply and produce a confident derivation of something untrue. We present Euston, an 8B …
- A Four-Stage Decomposition of Word-Problem Solving and Mechanistic Fragility in LLM Math Reasoning
Zhongdi Qu, Carla P. Gomes · 17 September 2026
Large language models solve grade-school math word problems with high accuracy, yet a single irrelevant clause inserted into the problem can collapse it. We reconcile these observations with a mechanistic account. We show that the model's internal computation decomposes into a four-stage sequential …
- AI and Human Approaches to Mathematical Problem Solving
Yang Ding · 17 September 2026
AI systems have begun to report solutions, disproofs, and substantive advances on long-standing mathematical problems, raising questions about whether they approach research in the same way as mathematicians. This study compares public AI research accounts with the human literature on 11 such proble…
- StochBench: A Domain-Specific Benchmark for Stochastic Processes in Lean
Idan Davidovich, Debargha Ganguly, Vikash Singh, Vipin Chaudhary · 10 September 2026
Leading benchmarks for formal theorem proving with large language models are small collections drawn from competition math, such as the IMO and Putnam, that poorly represent field-specific applications. We introduce StochBench, a Lean 4 benchmark of 450 graduate stochastic-processes problems at vary…
- From Symbolic Perception to Logical Deduction: A Framework for Guiding Language Models in Geometric Reasoning
Weichen Dai, Rafael Medeiros Cabral, Ziyi Shou, Yan Cao, Xin Shen, Dongcai Lu, Yi Zhou · 10 September 2026
Plane geometry remains a significant challenge in AI, requiring the integration of visual perception and mathematical reasoning. While Large Multimodal Models (LMMs) naturally handle visuo-linguistic inputs, they are often computationally intensive and opaque. We demonstrate that a pure Large Langua…
- AxQM: A Textbook-Scale Benchmark for Formal Proof Synthesis in a Library of Finite-Dimensional Quantum Mechanics
Weichen Winston Yin, Jacob M. Taylor, Dirk R. Englund, Frank H. L. Koppens · 7 September 2026
Formalizing mathematics in a proof assistant, where a machine checks every definition, statement and proof, has set a new standard of rigor. Large language models are now capable of formalizing autonomously, even at the scale of whole textbooks. We bring this standard of rigor to physics, where theo…
- Shared circuits predict whether LLMs generalize across formats in arithmetic reasoning
Andrea Gregor de Varda, Sana Pandey, Pengrui Han, Jacob Andreas, Evelina Fedorenko · 7 September 2026
In many forms of reasoning, including arithmetic reasoning, generalizing across superficial changes in input format is effortless for humans: anyone who can solve 2+5 can also solve 'two plus five'. In contrast, LLMs are more brittle to surface variations of the prompts: for example, they solve nume…
- Where the Verifier Fails: A Category-Level Audit of Reward Signals in RLVR
Esther Xin · 2 September 2026
Reinforcement learning with verifiable rewards (RLVR) and standard benchmark evaluation both rely on an automatic verifier that turns a free text answer into a binary reward. Prior work reports that one evaluation harness accepts only about 94% of its own ground truth answers, blaming LaTeX parsing.…
