TL;DR [中文]

In Terence Tao’s ICM 2026 public lecture Mathematics in the Age of AI, his central question was not simply, “Can AI do mathematics?” but rather: If AI can soon complete a significant portion of research-level mathematical tasks, what should the mathematical community optimise, and what must it protect?

As proofs move from scarcity to abundance, the scarce resources will no longer be answers alone, but validation, explanation, understanding, responsibility, community recognition and the ability to consolidate scattered results into stable theories.


On 24 July 2026, Professor Terence Tao delivered a public lecture titled “Mathematics in the Age of AI” at the International Congress of Mathematicians (ICM 2026). He did not focus on model leaderboards, nor did he attempt to predict ‘when AI will replace mathematicians’.

For readers unfamiliar with Terence Tao, he is one of the world’s leading and most influential mathematicians. Born in Adelaide, Australia, he earned his PhD from Princeton University at 20, and was promoted to full professor at UCLA at 24, becoming the youngest full professor in UCLA’s history. Today, he is a Distinguished Professor in the UCLA Department of Mathematics, with research spanning areas such as harmonic analysis, partial differential equations, combinatorics, and number theory. In 2006, at the age of 31, he received the Fields Medal for his contributions to partial differential equations, combinatorics, harmonic analysis, and additive number theory.

People often describe Tao as a ‘genius’, but emphasising only his precocity understates his decades of sustained research breadth, academic contributions, and commitment to public mathematical discourse. He is not only renowned for his problem-solving abilities but has also long discussed how mathematics is discovered, expressed, and validated through papers, books, and his research blog. Therefore, his perspective on mathematics in the age of AI is not simply a prominent mathematician commenting on a technological trend. It is a serious reflection from a scholar working at the front line of knowledge production on the goals, values, and responsibilities of his profession.

To summarise the core question I gleaned from the lecture in one sentence:

When machines can produce more and more correct answers, do we still clearly understand what mathematical research truly aims to produce?

This was ostensibly a lecture about mathematics, but at a deeper level it was about the system of knowledge production. It concerned not just the boundaries of AI’s capabilities, but how a community, after a fundamental change in production methods, reconsiders its goals, values, responsibilities and division of labour. What I particularly admired was that even with his extraordinary reputation, which could easily sway the discussion, Tao remained notably restrained: he set conditions and clarified objectives in a way familiar to mathematicians, then invited the entire community to consider the answers collectively, rather than presenting his personal judgement as a definitive conclusion.

Setting a Working Hypothesis: Focusing on the Real Question

Discussions about AI and mathematics often quickly turn into a capability race: can models solve IMO problems, can they complete research-level proofs, can they surpass a certain benchmark, when can they conduct independent research?

While these questions are important, Tao did not let the entire lecture dwell on the capability race. Instead, he proposed a ‘working hypothesis’ deliberately broad, roughly as follows:

In the near future, certain AI tools will be able to complete a significant proportion of research-level tasks in some mathematical fields, at a reasonable cost, with a reasonable success rate and a reasonable level of human oversight.

Listeners are not required to believe or welcome this hypothesis. It is more like a conditional step in a mathematical proof: assume it holds, then analyse the consequences.

Tao cited preliminary results from the First Proof challenge as background material: in a controlled test, four AI harnesses attempted to solve ten new research-level mathematical problems, of which seven had at least one answer reaching publishable quality, with a computational cost of approximately US$10–US$1,000 per problem. This does not show that artificial general intelligence has arrived. It does, however, suggest that AI participation in research-level mathematics is no longer merely a distant philosophical possibility.

Once this premise is temporarily accepted, the question shifts from ‘what machines can do’ to:

If machines can indeed produce more proofs, can the original goals, values, and institutions of mathematical research remain unchanged?

Math in the Age of AI

A New “Foundational Crisis”

Tao likened the current situation to the foundational crisis of mathematics in the early twentieth century.

Results like Russell’s paradox and Gödel’s incompleteness theorems forced mathematicians to re-examine previously taken-for-granted foundations such as sets, infinity, formal systems and axioms. That period was full of debate, but ultimately led to a more explicit, rigorous, and standardised mathematical framework.

Today, what is being challenged is no longer primarily ‘what is a number’ or ‘what is a set’, but another set of long-standing assumptions:

  • Is the sole purpose of mathematical research to solve more problems?
  • Does a correct proof automatically amount to valuable knowledge?
  • Who should be held responsible for results generated with AI involvement?
  • What process must a result undergo before it genuinely belongs to the mathematical community?

Therefore, this current crisis is more like a foundational crisis of mathematical values and practice.

Previously, such interrogations were often handled by philosophy, history, and sociology, while mathematicians focused more on technical work. The advent of AI, however, forces the entire mathematical community to confront them anew: once production capabilities fundamentally change, objectives that could previously be tacitly understood must now be explicitly articulated.

Mathematical Research Is More Than Problem-Solving

Mathematical research certainly includes solving unsolved problems, but it extends far beyond that. Proof generation has long been considered one of the scarcest and most highly rewarded parts of the research process, but this does not mean it is the entirety of mathematical research. Tao listed other goals, including: developing new theories and techniques, understanding the real world, building academic communities, training the next generation, expanding shared knowledge networks, and creating works of lasting aesthetic value.

In the past, these goals largely reinforced each other.

An important proof often brings new methods, new language, and new research directions; solving a problem can also train students, connect different fields, and drive updates in textbooks and theoretical frameworks. Precisely because these goals were highly correlated, people could use “how many important problems have been solved” as a proxy metric for mathematical progress.

The arrival of AI may change this correlation.

If machines can generate proofs at high speed, the “number of problems solved” might rise rapidly, but understanding, education, theory building, and community absorption capacity may not grow proportionally. We might have more conclusions without gaining an equivalent level of understanding; more papers without enough people to check, explain, and integrate them.

This is the risk captured by Goodhart’s Law:

When a measure becomes a target, it ceases to be a good measure.

If the entire system only rewards ‘who solved the problem first’ or ‘how many problems were solved’, a previously effective proxy metric might rapidly lose its original meaning of measurement under the impetus of AI.

How a Proof Becomes Common Knowledge

The most illuminating part of the lecture was Tao’s step-by-step expansion of what it means to ‘solve a problem’.

The initial goal seems simple:

Open Problem → Proof Generation

But if the answer cannot withstand scrutiny, quantity alone does not constitute true progress, so verification is also needed:

Open Problem → Proof Generation → Correctness Verification

Formal methods and proof assistants such as Lean, Rocq, and HOL can make verification more automated and reviewable, significantly speeding up the verification process in many scenarios. However, a formally correct proof can still be a piece of machine output that no one truly understands.

Thus, after correctness, explanation is also needed:

Open Problem → Proof Generation → Correctness Verification → Clear Exposition

Good mathematical writing not only demonstrates the validity of each step but also answers higher-level questions: Where are the key difficulties? What is the genuinely new idea? How does it relate to existing results? Which steps can be reused?

Even if a proof is correct and fluently written, it will not automatically become common knowledge. Other researchers still need to read, discuss, verify, cite, generalise, and reorganise it. Ultimately, truly important results must enter textbooks, standard references, and the theoretical framework of the field.

I summarize the several goal revisions in the lecture as follows:

Problem Identification → Proof Generation → Correctness Verification → Clear Exposition → Publication & Peer Review → Community Understanding & Acceptance → Integration into the Field’s Standard Theory

This is not Tao’s verbatim statement, but a generalisation of his sequential additions of verification, exposition, publication, digestion, and canonicalization across multiple slides. It reveals an often-overlooked fact:

Proof completion is not the end of research.

A more complete endpoint of research is the transformation of results produced by individuals or machines into public knowledge that the community can understand, verify, inherit, and continue to develop.

From Proof Scarcity to Proof Indigestion

Tao used a vivid term: proof indigestion.

If AI generates proofs faster than humans can verify them, unverified results will continuously accumulate; if verification outpaces explanation, correct but difficult-to-read proofs will continue to pile up; if publication outpaces community absorption, papers will overwhelm the peer review system; even if all results are published, truly important content might not be organised into stable, teachable, reusable theories in time.

Mathematics could thus transition from “proof scarcity” to “proof abundance,” with the bottleneck shifting from upstream discovery to downstream digestion:

Past bottleneck: Can a proof be found?

Future bottleneck: Can a proof be verified, explained, selected, absorbed and transmitted?

What is most noteworthy here is not just the familiar concern that “AI will produce junk papers,” but the deeper institutional problem behind it.

The existing academic system often values original discovery but tends to underestimate the importance of reviewing, editing, explaining, reproducing, textbook writing, and theoretical consolidation. These tasks may not be at the cutting edge, but they bear the heavy responsibility of transforming individual achievements into collective progress.

When generative capabilities improve significantly, these previously considered auxiliary tasks may instead become the most crucial infrastructure for the entire knowledge system.

Over-Smoothness: A Potential Barrier to Understanding

Another intriguing observation from the lecture was that AI-written mathematical texts are often near-perfect in spelling, grammar, and typography, yet they might gloss over genuinely difficult steps while extensively explaining obvious parts.

More subtly, proofs might also be “over-polished.”

In human writing, places where the author deliberated extensively usually retain some natural friction, prompting the reader to slow down. Over-smooth AI text might present difficult and routine steps with equal ease, causing readers to lose clues to gauge importance.

This is not an attempt to romanticise errors or suggest that ambiguity is preferable. It reminds us: good explanation does not mean superficial frictionlessness.

Truly excellent exposition should help readers establish the correct cognitive rhythm: where to skim quickly, where to pause and re-derive, and where the core ideas of the entire work reside.

As AI-assisted writing becomes more common, ‘linguistic fluency’ will become less scarce; the ability to accurately present the difficulty structure, origin of ideas, and logical flow of reasoning will become a more important writing skill.

Lecture and Leiden Declaration Practical Recommendations

Tao does not attempt to cover every institutional issue in a single lecture. The following practical recommendations draw partly on Tao’s lecture and partly on the Leiden Declaration, which he explicitly recommended.

1. Transparent Disclosure of AI Use

Researchers should state clearly in which stages AI was used, what tools were employed, and how it influenced the research process. Responsible disclosure should become an academic norm, and researchers should not conceal their AI use due to concerns about peer evaluation.

2. Human Authors Remain Responsible

The Leiden Declaration recommended by the lecture emphasises that even with the use of automated tools, the correctness of the results, the adequacy of the arguments and the completeness of the citations remain entirely the responsibility of the human authors. AI can participate in proof generation, formalisation, writing, and literature review, but it cannot become a void into which responsibility disappears.

3. Actively Support Review and Verification

Authors should not only pursue generation speed but also strive to reduce the cost of verification for their peers, including providing complete citations, clearly explaining AI involvement, and, when appropriate, submitting formal proofs or reproducible materials.

4. Reduce the Overemphasis on ‘Being the First to Solve a Problem’

When proof generation becomes increasingly inexpensive, simply rewarding ‘the first answer’ will amplify rushing, concealment, and low-quality output. Academic evaluation should elevate the status of explanation, validation, editing, reproduction, and theoretical consolidation.

5. Treat the Ability to Explain Clearly as a Publication Threshold

Tao proposed a simple yet weighty judgement criterion: if an author cannot explain their results with expert-level clarity, correctness, and appropriate attribution, then the work is not yet suitable for formal publication.

6. Retain Necessary Human Training in Education

In education and research training, AI use needs to be strictly limited in some scenarios. Learners still need to retain the ability to independently derive, identify errors, and perceive difficulty; otherwise, even with powerful tools, they will be unable to judge when the tools deviate from the goal.

7. Build New Workflows and Infrastructure

Traditional journals, peer review, and textbook systems struggle to directly absorb large-scale machine-generated results. The mathematical community needs to proactively design new mechanisms for verification, discussion, formalisation, knowledge consolidation and public communication, rather than completely entrusting the rules to commercial model providers.

The common principle behind these recommendations is:

AI can expand the capabilities of mathematicians, but it cannot replace transparency, responsibility, judgement and community collaboration.

Why I Admire This Lecture

In my view, there are three points that make this lecture most admirable.

First, it showed unusual restraint regarding short-term capability predictions, elevating the problem to a more enduring institutional level. Model performance will change rapidly, but understanding, responsibility, education, and community will not automatically be resolved by the next model release.

Second, it accurately identified the bottleneck shift. AI lowers the cost of generation, but it might transfer costs to verification, review, maintenance and integration. In software engineering, we already see similar cost shifts: after code became easier to generate, the expensive work shifted towards requirements judgement, architectural constraints, testing, deployment, and long-term maintenance.

Third, it reassessed the ‘invisible labour’ within the academic system. Reviewers, editors, teachers, textbook authors, and knowledge maintainers were often seen as supporting roles after innovation; in an era of proof abundance, they may become central to sustaining the quality of knowledge.

Precisely because I take the lecture seriously, I also think several issues deserve further exploration.

First, the lecture intentionally bundled differences across fields, tasks, and supervision methods into a broad working hypothesis to make the conditional analysis manageable. However, when the discussion turns to policy and resource allocation, these differences still need to be unpacked: problems with a high degree of formalisability and clear verification feedback will not be affected at the same rate as problems that rely on conceptual creation, long-term theoretical development, and research judgement.

Second, while ‘community recognition’ is irreplaceable, it is also not inherently fair. Academic communities can also be conservative, form hierarchies, and be influenced by resources, prestige and personal networks. The AI era requires both valuing human expert judgement and continuously improving human institutions, rather than idealising existing peer review.

Tao also explicitly stated at the end that similar analyses should be conducted for teaching, mentoring, hiring, grant applications, and public communication. Building on this point, I believe academic authorship, the concentration of computational power, closed commercial models and participation opportunities among different countries and research institutions also need to be incorporated into the same framework.

Raising these extended questions is not to demand that one lecture answer everything; on the contrary, that these questions naturally arise from the lecture demonstrates Tao’s success in opening up a space for long-term discussion within the mathematical community. For me, this is the most valuable aspect of the lecture: it is not a self-proclaimed complete, definitive answer, but a research agenda solemnly proposed by a leading scholar, inviting the entire community to continue advancing it.

Insights for AI Engineering and Academic Work

As someone involved in both AI engineering and academic review, I prefer to interpret this lecture as a thinking framework that can be transferred to many forms of knowledge work.

Greater generative capability does not transfer responsibility. Models can draft code, proofs, reports and review comments, but ultimate responsibility must still rest with individuals who can explain, verify, and bear the consequences.

Investing in generation capabilities must be accompanied by investments in validation and digestion capabilities. Stronger models, without corresponding testing, review, knowledge management, provenance tracking and quality thresholds, will only push the bottleneck downstream. Reviewing, editing, teaching, and maintenance should also receive evaluation commensurate with their importance.

Educational stages require retaining independent judgement. The goal of training is not just to get answers faster, but also to develop a sense of problems, errors, and judgement. Some difficulties cannot be entirely outsourced, as they are part of the process of capability formation itself.

Conclusion: Answers Become Cheaper, Judgement Becomes More Expensive

AI will not necessarily end mathematics, nor will it render mathematicians meaningless. It is more likely to prompt the mathematical community to re-answer a question long obscured by technological progress: Is mathematics merely about producing more theorems, or is it about helping humanity understand and think more clearly and effectively?

When proofs were still scarce, these two goals seemed almost identical; when proofs began to become abundant, they truly diverged.

The Chinese term xuéwèn (学问), often translated as ‘learning’ or ‘knowledge’, suggests more than answers alone: it combines ‘learning’ and ‘inquiring’ (学 and 问). It is about seeking truth, but also about articulating principles clearly and passing them on. From this perspective, the understanding, responsibility and community that Tao cherishes are not mere add-ons to mathematical discovery, but the very foundation upon which learning can grow and be passed down.

The most valuable capabilities in the future may no longer be just producing answers faster, but judging which questions are worth asking, which results are worth believing, which ideas are worth transmitting, and who is willing to take responsibility for these judgements.

In this sense, AI brings not merely an upgrade of mathematical tools, but a repricing of knowledge, value and community.

When answers become cheaper, understanding, judgement and responsibility become more expensive.


References


Leave a comment

Trending