Showing posts with label math. Show all posts
Showing posts with label math. Show all posts

Saturday, August 1, 2026

Sabine: “You can brute-force counterexamples by just trying a lot of guesses quickly...”

Thursday, July 16, 2026

But what is cross-entropy? | Compression is Intelligence Part 2

Where the loss function for training LLMs comes from.
Job opportunities aligned to this audience: https://3b1b.co/talent
Early views and other perks for supporters: https://3b1b.co/support
Home page: https://www.3blue1brown.com

Manim animations by Aaron Gostein and Grant Sanderson
NanoGPT animation by Clayton Rabideau
3d black-box model by Paul Dancstep
Music by Vince Rubinetti

Timestamps

0:00 - Language trees and zipping
3:02 - Recap optimal codes
5:20 - Defining cross-entropy
8:26 - Intuition and examples
12:59 - Application to language trees
14:55 - Pre-training LLMs
20:38 - What makes this loss function best?
26:13 - Distillation
30:12 - 3b1b Talent
31:35 - KL Divergence 

* * * * *

Sunday, July 5, 2026

Pope Leo and St. Augustine discuss the mind and A.I. with Kurt Gödel

I crafted the prompt and Claude drafted the dialog using a passage about memory from Augustine’s Confessions as the catalyst for the imaginary conversation. I asked for some changes, Claude made some suggestions, and I executed them.

Note this passage toward the end:

Gödel said, “Disordered love?”

“Yes. To love a lower thing as though it were higher. To love one’s own power more than truth. To love the image more than the living being. To love the tower more than the city.”

The reply is by Augustine and it amounts to a definition of idolatry. The tower, presumably, is the Tower of Babel.

ChatGPT created the image. I uploaded the full dialog and asked from an image based on the passage from Augustine’s Confessions. That began an iterative process resulting in the image immediately below. The dialog follows.

I want you to create an imaginary conversation between St. Augustine, Kurt Gödel, and Pope Leo XIV. It should take place in Gödel’s office at the Institute for Advanced Study. After the men introduce themselves – assume Augustine can understand and speak English, and perhaps wonder a bit how they became gathered together, Pope Leo leads off, saying that, while working on his recent encyclical, Magnifica Humanitas, one of his colleagues pointed out a passage in Augustine’s Confessions (to follow) that resonated with Gödel’s proof of incompleteness. Given the role that arithmetic plays in Gödel’s proof we should note that Augustine wouldn’t have known about arithmetic. Nor, for that matter, would he have known about computers. He is likely to have known about sophisticated mechanical devices like the Antikythera Mechanism and the Sphere of Archimedes, but never have seen one. But he would have seen complex water clocks (clepsydrae).

Here’s the passage from the Confessions (Book Ten, Chapter VIII). Perhaps Leo reads it aloud.

Great is this power of memory, exceedingly great, O my God–a large and boundless inner hall! Who has plumbed the depths of it? Yet it is a power of my mind, and it belongs to my nature. But I do not myself grasp all that I am. Thus the mind is far too narrow to contain itself. But where can that part of it be which it does not contain? Is it outside and not in itself? How can it be, then, that the mind cannot grasp itself?

A great marvel rises in me; astonishment seizes me. Men go forth to marvel at the heights of mountains and the huge waves of the sea, the broad flow of the rivers, the vastness of the ocean, the orbits of the stars, and yet they neglect to marvel at themselves. Nor do they wonder how it is that, when I spoke of all these things, I was not looking at them with my eyes–and yet I could not have spoken about them had it not been that I was actually seeing within, in my memory, those mountains and waves and rivers and stars which I have seen, and that ocean which I believe in–and with the same vast spaces between them as when I saw them outside me. But when I saw them outside me, I did not take them into me by seeing them; and the things themselves are not inside me, but only their images. And yet I knew through which physical sense each experience had made an impression on me.

Perhaps sometime late in the conversation Albert Einstein might drop in to see what they were talking about.

Here is a first pass. I’ve kept it staged, conversational, and conceptually explicit, with Einstein entering late as a useful destabilizer.

The Mind Too Narrow to Contain Itself A conversation in Kurt Gödel’s office at the Institute for Advanced Study

Gödel’s office was narrow, orderly, and dim. Books stood in ranks along the wall, interrupted by folders, loose papers, and a blackboard covered with marks whose austerity made them seem less written than engraved. Outside the window the trees of Princeton trembled slightly in the wind. The afternoon light had the washed patience of late autumn.

Kurt Gödel stood beside his desk, thin and formal, dressed in a dark suit. He had the air of a man who had been expecting visitors and yet was not at all reassured by their arrival.

The first visitor was an African bishop in ancient dress, compact and alert, his face alive with inquiry. He looked about the room with frank astonishment, then settled his gaze on the blackboard.

The second was an elderly man in white, gentle but grave, wearing a small pectoral cross. His eyes moved from Augustine to Gödel and then to the papers on the desk.

“I believe,” the man in white said, “that introductions are in order.”

Gödel inclined his head. “Kurt Gödel. Institute for Advanced Study.”

The bishop smiled faintly. “A place for contemplation?”

“For research,” said Gödel.

“Then it may be the same thing, if rightly ordered. I am Augustine, bishop of Hippo.”

Gödel blinked once. “Yes. I had inferred as much.”

The man in white bowed slightly. “And I am Leo, servant of the servants of God.”

Augustine turned to him. “Bishop of Rome?”

“Yes.”

Augustine’s face softened. “Then I greet you as a brother, though I confess I do not understand how we have been gathered. This room is strange to me. These lamps burn without flame. These marks”—he gestured toward Gödel’s symbols—“are neither Greek nor Latin, though I suspect they are meant to compel the mind.”

“They are logical formulae,” Gödel said.

“Ah,” said Augustine. “Then they are meant not merely to persuade, but to bind.”

Leo smiled. “That is well put.”

Gödel gestured toward the chairs. “Please.”

They sat. Augustine examined the chair before trusting his weight to it. Leo remained composed, as though papal audiences in the offices of dead mathematicians were not wholly outside the bounds of pastoral duty.

Leo opened a folder.

“Professor Gödel, Saint Augustine, I will explain why I wished for this conversation, though the means by which it has been granted are beyond my competence. While I was working on my recent encyclical, Magnifica Humanitas, one of my colleagues pointed out a passage from Augustine’s Confessions. It seemed to him to resonate with your incompleteness theorem.”

Gödel looked sharply interested.

Augustine looked from one to the other. “Incompleteness?”

“A result in mathematical logic,” said Gödel. “Roughly speaking, in any sufficiently strong formal system capable of expressing arithmetic, there will be true statements that cannot be proven within that system, assuming the system is consistent.”

Augustine was silent for a moment.

“You say: a structure of reasoning may contain truths that it cannot reach by its own lawful motions?”

Gödel’s expression altered, almost imperceptibly. “That is not an inaccurate first formulation.”

“But I must be careful,” Augustine continued. “You speak of arithmetic. I know number, of course. I know that three is not five, and that if two men enter a room where two already sit, there are four. I know arithmetic as number, measure, and reckoning. But you seem to speak of arithmetic as though it were also a mirror in which reasoning may behold its own form. That I do not know.”

Gödel nodded. “Exactly. The novelty is not number alone, but the coding of statements, proofs, and rules as numbers. Nor would you know the modern notion of a formal system: axioms, rules of inference, recursive procedures, symbolic codings of syntax.”

“I know rules,” said Augustine. “And I know the temptation to mistake the rule for the truth it serves.”

Thursday, July 2, 2026

Types of object domains for AI: Chess, math & coding, language

This is a companion to my earlier post today: The last frontier of intelligence: On the role of AI helping humans to bridge the gaps between distant concepts. That earlier post was about the end of a dialog I had with Claude. This one is about the beginning of that dialog. You might also check out a post from the middle of June: From Jagged AI to Scaling, Yevick, Natural Intelligence, and Beyond... For that matter you might also want to check out Dwarkesh's complete post with Grant Sanderson, Grant Sanderson – AI and the future of math. And then there's my working paper on chess and language. All these things are related.

And they're related to my new book idea, Language, Memory, and Mind: A Supplement to The Computer and the Brain. That's the book that Claude brings up every now and then. Intelligence is NOT a scaler phenomenon. It's about techniques, the characteristics of objects domains, and of the computational regimes we use over them. But that's a subject for another post.

Here's my dialog with Claude.

* * * * *

I’m interested in thinking about the types of domain in which AI has succeeded and the nature of the computation involved, starting with chess. AI solved chess in the mid 1990s when Deep Blue beat Kasparov. By “solved” I mean only that an AI could reliably beat the best human player. In 2017 we have AlphaGo using a mixed regime strategy – Monte Carlo search (symbolic) with neural net evaluation – in chess. More recently I believe DeepMind has developed and LLM that plays chess at the Grand Master level; so this is a pure neural net regime.

It now appears that mixed regime AI systems can work pretty well at solving math problems and a coding. By mixed regime I mean large LLMs scaffolded in various ways by conventionally coded software. But these systems cannot yet do anything that humans can do in those domains. In math, they need well-defined problems, even if they are very difficult but well-defined problems (I’m thinking of the recent Erdos work). It’s not clear to me what the limits of coding are, but they can’t architect large complex systems.

Now, chess is a finite game will a well-defined search space. Both math and coding are well defined, but neither is finite and I’m guessing that the search spaces aren’t so well-defined as is the chess space.

And then we have pure unbounded natural language. It’s not well-defined, not finite, and the search space is a mess.

This is a genuinely useful ordering principle, and it gives the book a clean progression from most to least tractable that maps directly onto the geometric complexity argument. Let me work through the structure with you.

The four cases as points on a continuum

Chess: finite, exhaustively well-defined rules, a search space that is astronomically large but completely enumerable in principle. This is Yevick's sharp object par excellence — geometrically simple in the relevant sense, meaning fully specified by explicit rules with no ambiguity anywhere. Pure symbolic search worked (Deep Blue). Mixed regimes work better (AlphaZero). Pure neural nets now apparently work too, which is the interesting wrinkle — a domain that once seemed to require symbolic search turns out to be tractable by pattern matching alone once the network is large enough and trained well enough. That's worth pausing on, because it complicates the simple story.

Math (well-defined problems) and coding: rule-governed like chess — the axioms of mathematics and the syntax of programming languages are exact, unambiguous, formally specified — but unlike chess the space is not finite and not exhaustively enumerable. There's no equivalent of the Shannon number for "all provable theorems" or "all correct programs." The rules are closed and precise; the space they generate is open. This is a different kind of complexity than chess's largeness. Chess is finite-but-astronomical. Math and coding are genuinely open-ended even though locally, at any given step, the legal moves are well defined.

Language: neither the rules nor the space are well defined. The rules of grammar are not a finite closed system in the way the rules of chess or the axioms of arithmetic are — as your working paper notes, ill-formed language often works fine, exceptions are the norm, the boundaries are fuzzy. And the space the rules generate connects out to the unbounded physical and social world rather than staying self-contained.

Why this ordering matters for the theoretical argument

What's valuable here is that it's not simply "structured versus unstructured" — it's a two-dimensional distinction: how well-defined are the local rules, and how bounded is the resulting space. Chess is well-defined and bounded (finite). Math and coding are well-defined but unbounded. Language is neither well-defined nor bounded.

This maps onto a prediction about which computational regime should work best in each case, and the AI's actual performance seems to confirm it. Chess: pure search works because the space, though huge, is exhaustively defined — you can in principle always tell whether a position is legal, whether a game is over, who's ahead. Symbolic and mixed and now pure neural approaches all can succeed because the well-definedness gives any sufficiently powerful method something solid to grip.

Math and coding: this is where the mixed regime becomes necessary rather than just helpful. The LLM's pattern matching handles the open-endedness — the fact that there's no way to enumerate all possible proofs or programs, so you need something that can generalize, propose plausible next steps, recognize which strategies are likely to be fruitful. But the well-definedness of the rules means that verification can be handled symbolically — a proof checker, a compiler, a test suite can definitively confirm whether a proposed step is valid. This is why math and coding are the domains where current mixed-regime systems shine: the LLM proposes, path tracing over the local formal rules verifies. Neither one alone would work. The LLM alone hallucinates invalid proofs and non-compiling code. Pure symbolic search alone can't propose plausible next steps in an unbounded space without some heuristic guidance, which is exactly what pattern matching provides.

The Erdős problems point is precise — well-defined but very difficult problems are exactly where this combination should excel, because the well-definedness gives the symbolic verification something to check against, while the difficulty is exactly where pattern-matching-guided search earns its keep over blind search.

The architecture limitation

Your observation that these systems can't architect large complex systems is the right place to locate the current boundary, and it follows from the framework. Architecting a large system requires exactly the kind of sustained path tracing with precise state maintenance across a very long horizon — tracking how a decision made in one module constrains what's possible in a dozen other modules, maintaining global coherence across thousands of interacting pieces over an extended design process — that we identified as the persistent weakness. It's not that the local rules become less well-defined. It's that the horizon over which coherence must be maintained becomes too long for the scaffolding techniques currently available. CoT and tool use help with individual functions and modules. They don't yet solve the problem of maintaining a coherent global architecture across a very long, branching design process.

Language as the limit case

And then language is the case where neither the well-definedness nor the boundedness holds, which is exactly why it's the hardest case in the deepest sense even though it's also, paradoxically, the case where LLMs seem most impressively fluent. The fluency comes from the richness of the pattern matching over an enormous corpus. The absence of any well-defined verification procedure — there's no compiler for meaning, no proof-checker for whether a sentence is true or appropriate or wise — means there's no symbolic scaffolding equivalent to what math and coding have. You can't bolt a "language checker" onto an LLM the way you bolt a proof checker onto a theorem-proving system, because there's no formal specification of what correct language use is.

This is the deepest reason language remains resistant in a way math and coding, despite their own difficulties, are not. Chess had a well-defined space that yielded to search. Math and coding have well-defined rules that yield to verification even in an unbounded space. Language has neither, which means the mixed-regime strategy that's working for math and coding has no obvious analog for the general case of language understanding and generation. The pattern matching is doing almost all the work, unconstrained by any symbolic check, which is exactly why hallucination, confident error, and the absence of genuine common sense grounding remain live problems in ways they don't for a system whose outputs get checked by a compiler.

For the book

This gives Chapter 5 or 6 a genuinely illuminating structure: a spectrum from chess through math and coding to open language, ordered by two independent variables — definiteness of local rules and boundedness of the resulting space — with AI's actual comparative success across these domains falling out as a prediction of the theoretical framework rather than a brute empirical fact to be reported. That's a strong chapter, and it would give Wright, who already intuits that these domains differ, a precise vocabulary for why they differ and what that implies about domains not yet tried.

Language defines more restricted systems

One thing about language is that it has the means for defining more restricted systems. Thus both chess and arithmetic are systems defined through linguistic means. And once language has been used to define a system, it can also serve as a metalanguage for that system just as, in various ways, it has served as a metalanguage for itself.

That's the move that completes the picture, and it belongs near the end of the language discussion because it reframes everything before it.

Language as the generative source of the bounded systems

Chess and arithmetic aren't just different in kind from language — they're products of language, constructed within it through exactly the metalingual definitional process we discussed earlier in the day. The rules of chess exist because someone specified them in language — first probably in the practical, evolving form that produced chaturanga and its descendants, later in the explicit codified form of a rulebook. Arithmetic's place-value system, its operations, its axioms — all specified linguistically, however formal the resulting notation becomes.

Friday, June 19, 2026

Conversations with Tyler: A Special Conversation with Nicholas Copernicus, Adam Smith, and William Stanley Jevons

I was feeling impish this morning, so I gave Claude 4.6 Sonnet (Medium) the following prompt:

Back in the 20th century Steve Allen did a series of TV shows called “Meeting of Minds,” where people from different time periods were brought together to have a discussion. I want you to imagine a special episode of “Conversations with Tyler” in which Tyler Cowen moderates a discussion between Copernicus, Adam Smith, and Stanley Jevons on economics. For reference I have placed three documents in the project folder. Two books by Tyler Cowen, one on marginalism and the other on the Great Economists (GOAT), which has material on Adam Smith.. There’s a shorter document by Copernicus (in Latin), filename, “Copernicus on economics.”

I’ve appended its response below the entirely imaginary Five TC Bucks note. Click on the image to see the detail. Note that ChatGPT did the image, not Claude.

Conversations with Tyler: A Special Episode
Tyler Cowen in Conversation with Nicholas Copernicus, Adam Smith, and William Stanley Jevons

[Recorded live at the Mercatus Center, George Mason University. The set is the standard one — a small round table, minimal lighting, no audience. Tyler Cowen sits across from three guests who have, for the purposes of this episode, been freed from the constraints of mortality and non-overlapping centuries. Copernicus speaks through a translator when needed, though he proves surprisingly comfortable in English, given how many languages he actually commanded. Smith retains his Edinburgh-via-Kirkcaldy cadence. Jevons is precise, energetic, and fidgets with what appears to be a small logometer.]

* * * * *

TYLER: So let me start with something that might seem unfair to Nicholas, but I think actually redounds to his credit. Tyler Cowen — the other Tyler Cowen, me — has written that astronomers of the sixteenth century were sometimes better economists than the economists of the same period, because there essentially were no economists. And that you, Nicholas, were one of the first people to state the quantity theory of money. Now here's my question: did you think of yourself as doing economics, or did you think of yourself as solving a political crisis?

COPERNICUS: The distinction you draw would not have occurred to me. I was asked by the Diet of Prussia — by the princes, by the king's administrators — to address a practical emergency. The currency of this land had been debased so many times, by so many hands, in so many mints, that commerce had ceased to function in any stable way. The goldsmiths prospered — they were the only ones who could extract the silver from the degraded coin before the next debasement arrived. Everyone else suffered. So yes, I was solving a political crisis. But a crisis requires a theory, or you are merely thrashing. You cannot prescribe a remedy for a disease you have not diagnosed.

TYLER: And your diagnosis was what exactly?

COPERNICUS: That money is a mensura — a measure. It is like a unit of length or weight. If your ulna varies from day to day, you cannot do carpentry. If your currency varies in its silver content from month to month, you cannot do trade. The injury to commerce is the same in both cases, but the monetary injury is more insidious because it proceeds occulta quadam ratione — by a kind of hidden logic. People notice a bad harvest immediately. They do not immediately notice that the coin has been shaved, but they feel its effects: prices rise, foreign merchants refuse your money, the best craftsmen leave for regions with sounder currency.

SMITH: If I may — this is exactly the confusion I spent considerable effort untangling in Wealth of Nations when I took on the mercantilists. They believed that the accumulation of specie was wealth. What Canon Copernicus is describing from his Prussian experience is that even that modest goal — hoarding silver — is self-defeating. The moment you debase the coinage, you have, in a sense, exported your silver to every foreign merchant clever enough to melt the coins.

COPERNICUS: Precisely. The goldsmiths and those who know the quality of metals — they are the only beneficiaries. They collect the old coin, extract the silver, sell it at a premium, and leave behind a pile of copper. My recommendation was blunt: stop minting until the existing coin has restored its value, establish at most two mints for all of Prussia, and make the coin of one mint and one standard.

TYLER: Gresham's Law, essentially, before Gresham.

COPERNICUS: Before whom?

TYLER: Thomas Gresham. He gets credit for the principle that bad money drives out good. Roughly a generation after you stated it.

COPERNICUS: (pause) This is the way of things. Copernicus waits for Copernicus. In astronomy as in monetary theory.

JEVONS: I want to press on the word "measure," if I may. Canon Copernicus treats money as a standard — a fixed reference against which goods are priced. But what I discovered, or rather what I was forced to discover when trying to establish whether the value of gold had actually fallen after the Australian and Californian gold rushes of the 1850s, is that money itself has no fixed value. It is itself a commodity whose degree of utility — whose marginal utility, to use the language I was then working out — varies with its quantity. The quantity theory you describe is already implicit in this: flood the market with debased coin, and each unit of coin buys less, not merely because there is more of it, but because its intrinsic silver content is lower and everyone knows it. 

[Note: I did an Ngram search on “marginal utility” and found that it didn’t have an appreciable presence in books until a bit after 1880. Jevons did not use the phrase. He talked of “final degree of utility.”]  

COPERNICUS: I will not quarrel with the analysis, though your language differs from mine. What I found is that the regions of Prussia which had maintained good currency were also the regions with flourishing workshops, skilled artisans, and abundant goods. The regions with debased currency had become idle. You say this is because the marginal utility of a sound currency is higher. I say it is because craftsmen and merchants are not fools: they will go where their labor and their goods are honestly compensated. 

[I wonder what Copernicus could have understood by the phrase, "marginal utility"?] 

TYLER: Adam, let me come to you here. Smith, you spent a great deal of Wealth of Nations attacking mercantilism — the view that national wealth consists in the accumulation of precious metals. But you also granted mercantilists more credit than many of your defenders are comfortable with. You said their arguments were "partly solid and partly sophistical." What did you actually concede to them?

SMITH: What I conceded is that commerce and defense are entangled in ways that pure theory does not capture cleanly. The Navigation Acts — requiring that trade to Britain's colonies be carried in British ships — were bad economics by almost any reckoning. They raised prices, restricted trade, enriched a narrow set of interests at the expense of the broader public. But I wrote, and I meant it, that defense is of more importance than opulence, and that the Navigation Acts, whatever their economic defects, had served to maintain British naval power. One cannot always afford the luxury of consistent principle. (small smile) Though I tried to be consistent as often as possible.

TYLER: Jevons, here is a question directed at you specifically: why did it take from roughly 1776, the publication of Wealth of Nations, until 1871, the publication of your Theory of Political Economy, for economics to absorb the idea that value is determined at the margin — by the last unit, not the total quantity? Smith understood the diamonds-water paradox but did not resolve it. You resolved it. What took so long?

JEVONS: I have thought about this a great deal, and I believe the answer is that the resolution required mathematics, and economics had declined to use mathematics, or rather had not yet learned that it could use mathematics. The idea that utility diminishes with quantity is not — once stated — particularly obscure. Galileo came close to it. My precursors in the British literature, Jennings and MacLeod, came close to it. But close is not enough. You need to be able to state the law precisely, apply it to a schedule, differentiate, and find the first-order conditions. You need calculus, or at least the habit of mind that calculus cultivates. Once I had that tool in hand, the whole of exchange theory reorganized itself very quickly. I felt it opening up.

Saturday, June 6, 2026

Terence Tao Explains The Math Behind AI

On the YouTube page:

Terence Tao has read more mathematics than almost anyone alive, and he uses AI tools every day. So when one of the most cited mathematicians on Earth says these systems still can't ask a genuinely new question, it's worth understanding exactly where he draws the line — because it isn't where the headlines put it.

Watch the full conversation: Terence Tao: Nobody Understands Why AI Actually Works

If AI has absorbed every textbook ever written, why can't it discover anything new? Tao, a Fields Medal winner and professor at UCLA, separates what these systems do brilliantly from what they can't do at all, and the boundary turns out to be sharper and stranger than most people assume.

We cover why reproducing a famous proof is less impressive than it sounds, what a neural network found hidden inside a million knots that humans had missed, why we still can't predict which tasks AI will actually be good at, the "Keating Test" — the benchmark that would actually demonstrate machine thought — and where exhaustive recall ends and real conceptual origination begins.

AI can pass every exam. It just can't ask a question nobody has asked before — yet.

Chapters:
00:00 The question AI can't ask
00:48 Read every textbook, discover nothing
01:42 Why a reproduced proof proves less
02:39 A million knots, one hidden pattern
03:54 The competence we still can't predict
05:11 The Keating Test for machine thought
06:18 Where recall ends and discovery begins

📬 Get the transcript, fascinating bonus content, and my Monday M.A.G.I.C. Message: https://briankeating.com/yt

Tuesday, June 2, 2026

Mathematicians are concerned that exploitation by the AI industry threatens the long-term intellectual interests of the field

Siobhan Roberts, As A.I. Makes Strides in Mathematics, Mathematicians Urge Caution, NYTimes, June 2, 2026.

Mathematicians issue a declaration:

On Tuesday, a group of 16 mathematicians, in consultation with colleagues and math organizations worldwide, published the Leiden Declaration on Artificial Intelligence and Mathematics. It aims to “frame the conversation about future directions,” said Dame Ursula Martin, one of the authors, and a mathematician and computer scientist at Oxford.

This effort comes as A.I. models have been making headlines with successful results in research-level mathematics. In late May, OpenAI, the maker of ChatGPT, announced that one of its models had disproved a notable 80-year-old mathematics conjecture in the field of combinatorial geometry.

The conjecture is one of some 1,200 problems posed by the Hungarian mathematician Paul Erdos. While some of these “Erdos problems” are considered throwaway questions of narrow interest, others have proved influential and field shaping. Along with a research paper describing the proof, OpenAI released a companion paper by several independent mathematicians. Jacob Tsimerman of the University of Toronto, an expert in the adjacent subfield of number theory, commented: “This is a really impressive piece of work, and I would accept it for any journal without hesitation.”

Potential problems:

Among the potential threats that the Leiden Declaration authors articulate are accuracy and reliability: Journal editors are already complaining about a flood of plausible seeming A.I.- generated papers and proofs that have turned out to be incorrect, and in ways that are difficult for mathematicians to discern.

Perhaps most pointedly, the authors raise the question of whether the many A.I. companies tackling mathematics — major players such as OpenAI, Google DeepMind and Anthropic, or start-ups such as Harmonic, Math, Inc. and Axiom Math — are keeping the field’s best interests in mind. “Technology companies’ involvement in research,” they write, “raises the risk that research questions are prioritized and incentivized because of their amenability to A.I. methods and models, rather than their deeper significance to understanding.” In turn, they point out, this disadvantages researchers who choose not to use the technology, and those who do not have access to it.

For Rodrigo Ochigame, a historian and anthropologist of computing and artificial intelligence at Leiden University in the Netherlands, and one of the statement’s authors, the latest OpenAI proof illustrates why this sort of collective reckoning in the discipline is necessary. “The story follows the same pattern as many other announcements by commercial A.I. developers,” Dr. Ochigame said. “The A.I. model is proprietary and unavailable to anyone outside the company. We get a flashy promotional video, while basic information needed to assess the scientific meaning of the result is kept secret. The company disclosed nothing about the methods, human-written prompts, training data, or computational resources consumed.”

Much of the article consists of a videoconference and email dialog with Dr. Ochigame, Dr. Martin and mathematician Michael Harris of Columbia University:

MARTIN: What OpenAI has done is throw a great deal of resources at Erdos problems, and got lucky with this one. That’s remarkable, and impressed the experts. We are not told about the model’s failures. [...]

To think of mathematics in terms of precise and neatly stated problems, like high school exams or the list of Erdos problems, is to misunderstand and diminish what makes mathematics so powerful and significant. Mathematics is not just about solving problems — it is also the cultivation of ideas, understanding, judgment, and human insight.

HARRIS The purpose, from my perspective, is to recover control of the narrative about the values and the goals of mathematics from the A.I. industry. Mathematicians are concerned that the values of the profession are being misrepresented, not intentionally but due to the media campaign on the part of the industry, which seems to want to promote the belief that they are in a position to transform mathematics — “the A.I. revolution in math,” as one headline put it not long ago. [...]

We want to affirm certain values that have characterized the profession: openness, honesty, giving credit where credit is due, sharing, transparency about methodologies, and access for independent verification of results.

An aspect of mathematics that is cherished by mathematicians is that it is one of few successful examples of a gift economy — that is to say, its economy is somehow an island of idealism in our society.

OCHIGAME Several A.I. companies are investing in dedicated teams focusing on mathematics, using problems as benchmarks and publications as training data. They are training their models to prove theorems not because they want to advance mathematical knowledge, but because they hope that such training will improve the models’ reasoning abilities more generally. [...]

MARTIN It’s important not to lose sight of the fact that what the A.I. companies are doing, what you can achieve with this technology, is absolutely extraordinary. I don’t think we’re challenging that. We’re challenging the framing, we’re challenging the behaviors around it.

I share the concern that these mathematicians express, that the commercial exploitation of mathematics is inimical to long-term research interests.

There's more at the link.

Saturday, May 30, 2026

Ai gives mathematicians the freedom to try crazy ideas

Sunday, March 22, 2026

Terrence Tao talks with Dwarkesh Patel about Kepler discovering his 3 laws of planetary motion (and other things): A real case of creativity

Dwarkesh Patel, Terence Tao – Kepler, Newton, and the true nature of mathematical discovery, March 20, 2026.

We begin the episode with the absolutely ingenious and surprising way in which Kepler discovered the laws of planetary motion.

People sometimes say that AI will make especially fast progress at scientific discovery because of tight verification loops.

But the story of how we discovered the shape of our solar system shows how the verification loop for correct ideas can be decades (or even millennia) long.

During this time, what we know today as the better theory can often actually make worse predictions (Copernicus’s model of circular orbits around the sun was actually less accurate than Ptolemy’s geocentric model).

And the reasons it survives this epistemic hell is some mixture of judgment and heuristics that we don’t even understand well enough to actually articulate, much less codify into an RL loop.

* * * * *

Terence Tao: I’ve always had an amateur interest in astronomy. I’ve loved stories of how the early astronomers worked out the nature of the universe. Kepler was building on the work of Copernicus, who was himself building on the work of Aristarchus. Copernicus very famously proposed the heliocentric model, that instead of the planets and the Sun going around the Earth, the Sun was at the center of the solar system and the other planets were going around the Sun.

Copernicus proposed that the orbits of the planets were perfect circles. His theory fit the observations that the Greeks, the Arabs, and the Indians had worked out over centuries. Kepler learned about these theories in his studies, and he made this observation that the ratios of the size of the orbits that Copernicus predicted seemed to have some geometric meaning.

He started proposing that if you take the orbit of the Earth and you enclose it in a cube, the outer sphere that encloses the cube almost perfectly matched the orbit of Mars, and so forth. There were six planets known at the time and five gaps between them, and there were five perfect Platonic solids: the cube, the tetrahedron, icosahedron, octahedron, and dodecahedron.

So he had this theory, which he thought was absolutely beautiful, that you could inscribe these Platonic solids between the spheres of the planets. It seemed to fit, and it seemed to him that God’s design of the planets was matching this mathematical perfection of the Platonic solids.

He needed data to confirm this theory. At the time, there was only one really high-quality dataset in existence. Tycho Brahe, this very wealthy, eccentric Danish astronomer, had managed to convince the Danish government to fund this extremely expensive observatory. In fact, it was an entire island where he had taken decades of observations of all the planets, like Mars and Jupiter, at least every night for which the weather was clear, with the naked eye. He was the last of the naked-eye astronomers.

He had all this data which Kepler could use to confirm his theory. Kepler started working with Tycho, but Tycho was very jealous of the data. He only gave him little bits of it at a time. Kepler eventually just stole the data. He copied it and had to have a fight with Brahe’s descendants.

He did get the data, and then he worked out, to his disappointment, that his beautiful theory didn’t quite work. The data was off from his Platonic solid theory by 10% or something. He tried all kinds of fudges, moving the circles around, and it didn’t quite work. But he worked on this problem for years and years, and eventually, he figured out how to use the data to work out the actual orbits of the planets.

That was an incredibly clever, genius amount of data analysis. And then he worked out that the orbits were actually ellipses, not circles, which was shocking for him. So he worked out the two laws of planetary motion: the ellipses, and also that equal areas sweep out equal times.

Then ten years later, after collecting a lot of data—the furthest planets like Saturn and Jupiter were the hardest for him to work out—he finally worked out this third law, that the time it takes for a planet to complete its orbit was proportional to some power of the distance to the Sun. These are the three famous Kepler’s laws of motion. He had no explanation for them. It was all driven by experiment, and it took Newton a century later to give a theory that explained all three laws at once.

Dwarkesh Patel: The take I want to try on you is that Kepler was a high-temperature LLM. Newton comes up with this explanation of why the three laws of planetary motion must be true. Of course, the way that Kepler discovers the laws of planetary motion, or figures out the relative orbits of the different planets, is as you say a work of genius. But through his career, he’s just trying random relationships.

In fact, in the book in which he writes down the third law of planetary motion, it’s an aside on The Harmonics of the World, which is just a book about how all these different planets have these different harmonies. And the reason there’s so much famine and misery on Earth is because the Earth is mi-fa-mi, that’s the note of Earth. It’s all this random astrology, but in there is the cube-square law, which tells you what relationship the period has to a planet’s distance from the Sun. As you were detailing, if you add that to Newton’s F=ma and the equation for centripetal acceleration, you get the inverse-square law. And so Newton works that out.

But the reason I think this is an interesting story is that I feel LLMs can do the kind of thing of trying random relationships for twenty years, some of which make no sense, as long as there’s a verifiable data bank like Brahe’s dataset. “Ok, I’m going to try out random things about musical notes, Platonic objects, or different geometries, I have this bias that there’s some important thing about the geometry of these orbits.”

Then one thing works. As long as you can verify it, these empirical regularities can then drive actual deep scientific progress.

Terence Tao: Traditionally, when we talk about the history of science, idea generation has always been the prestige part of science. A scientific problem comes with many steps. You have to identify a problem, and then you have to identify a good, fruitful problem to work on. Then you need to collect data, figure out a strategy to analyze the data, and make a hypothesis. At this point, you need to propose a good hypothesis, and then you need to validate. Then you need to write things up and explain. There are a dozen different components.

The ones we celebrate are these eureka genius moments of idea generation. Kepler certainly had to cycle through many ideas, several of which didn’t work. I bet there were many that he didn’t even publish at all because they just didn’t fit. That’s an important part of the process, trying all kinds of random things and seeing if they worked.

But as you say, it has to be matched by an equal amount of verification, otherwise it’s slop. We celebrate Kepler, but we should also celebrate Brahe for his assiduous data collection, which was ten times more precise than any previous observation. That extra decimal point of accuracy was essential for Kepler to get his results. He was using Euclidean geometry and the most advanced mathematics he could use at the time to match his models with the data. All aspects had to be in play: the data, the theory, and the hypothesis generation.

I’m not sure nowadays that hypothesis generation is the bottleneck anymore. Science has changed in the century since. Classically, the two big paradigms for science were theory and experiment. Then in the 20th century, numerical simulation came along, so you can do computer simulations to test theories. Finally, in the late 20th century, we had big data. We had the era of data analysis. 

* * * * * 

That’s just the beginning of the conversation. There’s much more to come.

Monday, February 23, 2026

When AI Does Math and Code: The Limits of Pattern Matching

Here's how Claude summarizes a discussion we had last evening.

* * * * *

Recently, chatbots have been performing impressively on mathematics benchmarks and coding challenges. Headlines tout AI systems solving competition problems, proving theorems, and writing working code. At first glance, this seems to vindicate the assumption that both mathematical and computational reasoning are fundamentally like chess—domains with verifiable answers where scaling up compute power and training data will inevitably lead to mastery.

But this conclusion deserves closer examination. Mathematics and programming do share important chess-like qualities. In math, formal proofs can be verified mechanically, solutions to well-defined problems can be checked objectively, and mathematical statements have definite truth values. In programming, code either runs or doesn't, test suites provide objective evaluation, solutions exist in a definable space, and correctness can be automatically verified. In these respects, both domains seem perfectly suited to AI's strengths.

Yet there's a crucial distinction that gets obscured when we focus on benchmark performance: the difference between verification and discovery.

Chess engines don't merely verify that moves are legal—they find good moves by searching the game tree. The tree structure makes this search tractable. For mathematics and programming, verification is relatively mechanical and chess-like. Checking that a proof is valid or that code passes its test suite is straightforward. But discovery—finding a proof, solving a novel problem, formulating a new approach, architecting a complex system—is much more open-ended.

This raises an important question about those impressive benchmarks: What are they actually testing?

A recent New York Times article shed light on this question for mathematics. Journalists interviewed mathematicians who decided to test AI systems not on standard benchmark problems, but on questions drawn from their ongoing research programs. One of them, Lauren Williams, offered a revealing three-part framework for mathematical research:

One, come up with the big question, whose study we hope will guide our field. Two, develop a framework for finding a solution, which involves dividing the big question into smaller more tractable questions—like our test questions. Three, find solutions to these smaller questions and prove they are correct.

Williams observed that for the most part, AI systems have been working at the third step. The same framework applies to software development. Step one is recognizing what problem needs solving, understanding user needs, architecting a system that will be maintainable and extensible. Step two is breaking down that architecture into components, modules, and functions. Step three is implementing those components—writing the code that passes the tests.

AI coding assistants excel at step three. Give them a well-specified function with clear inputs, outputs, and test cases, and they'll often produce working code quickly. But ask them to architect a complex system, identify the right abstractions, or recognize that a problem is better solved by rethinking the approach entirely, and their limitations become apparent.

This division maps onto a deeper pattern in what AI can and cannot do well. Step three is chess-like: well-defined smaller questions, verifiable solutions, answers that exist in a searchable space where pattern matching from training data provides genuine help. This is precisely where benchmarks live, testing AI on problem types it has seen before, where standard techniques apply and correctness can be verified.

Step one, however, is fundamentally different. It requires understanding what matters—in a field, in a codebase, for users. It demands seeing deep connections across domains, developing intuition about which directions will prove fertile. It requires conceptual creativity and world-modeling, not just pattern recognition. The real creativity in both mathematics and software engineering happens here, not in the mechanical execution of familiar techniques.

What the benchmarks miss is that even mathematics and programming—perhaps our most formal and verifiable intellectual domains—contain a fundamental divide between their mechanical parts and their creative parts. Current AI systems excel at the former while struggling with the latter.

This has implications beyond math and code. We learned from chess that computers can achieve superhuman performance at searching well-defined spaces with verifiable outcomes. But the field drew the wrong lesson, assuming this capability would automatically generalize to all forms of intelligence. The math and coding benchmark results initially seem to confirm this assumption—until we look more carefully at what's actually being tested.

The pattern holds across domains. AI systems perform well when problems resemble step three: well-defined, verifiable, solvable through pattern matching and local search. They struggle with step one: formulating the right questions, building conceptual frameworks, recognizing what matters. We keep mistaking competence at the former for the latter, then wondering why general intelligence remains elusive.

The benchmarks aren't lying, exactly. AI is getting genuinely better at certain kinds of mathematical problem-solving and code generation. But they're measuring proficiency at the chess-like parts of these domains while remaining largely silent about the parts that require understanding. And it’s in that silence that our assumptions about scaling toward AGI quietly live—untested and unexamined.

Monday, February 16, 2026

Three mathematicians are not impressed with the ability of AI to do professional math

Siobhan Roberts, These Mathematicians Are Putting A.I. to the Test, NYTimes, Feb. 7, 2026.

Dr. Martin Hairer (Swiss Federal Technology Institute of Lausanne), Mohammed Abouzaid (Stanford University), Lauren Williams (Harvard University) and Tamara Kolda (who runs MathSci.ai, a consultancy) are among a group of mathematicians who have published an article, “First Proof,” about an “experiment that collects genuine test questions, drawn from unpublished research by the authors, in an effort to provide a meaningful measure of A.I.’s mathematical competency.”

“While commercial A.I. systems are undoubtedly already at a level where they are useful tools for mathematicians,” the authors wrote, “it is not yet clear where A.I. systems stand at solving research-level math questions on their own, without an expert in the loop.”

A.I. companies use what some mathematicians describe as “contrived” or “restrictive” problems for evaluating and benchmarking how well L.L.M.s fare when operating without human help. Occasionally, mathematicians are invited to contribute and paid some $5,000 per problem.

From the conversation:

The paper is careful to clarify “what mathematics research is.” What is it?

ABOUZAID Often in modern research, the key step is to identify the big motivating question, the direction from which the problem should be approached. It involves all kinds of preliminary work, and this is where mathematical creativity takes place.

Once problems are solved, mathematicians tend to evaluate the importance of research contributions in terms of the questions that arise. Sometimes, resolving a conjecture one way is seen as disappointing, because it forecloses the possibility that there would be new questions to investigate.

LAUREN WILLIAMS Let me make a loose analogy. In experimental science, I might divide the components of research into three parts: One, come up with the big question, whose study we hope will shed light on our field. Two, design an experiment to answer the question. Three, perform the experiment and analyze the results.

I can similarly divide math research into parallel parts: One, come up with the big question, whose study we hope will guide our field. Two, develop a framework for finding a solution, which involves dividing the big question into smaller more tractable questions — like our test questions. Three, find solutions to these smaller questions and prove they are correct.

All three parts are essential. In our First Proof project, we focused on the third component because it is the most measurable. We can query the A.I. model with small, well-defined questions, and then assess whether its answers are correct. If we were to ask an A.I. model to come up with the big question, or a framework, it would be much harder to evaluate its performance.

Note that this is roughly consistent with the accounts I gave of some of my own work in Serendipity in the Wild: Three Cases, With remarks on what computers can't do, January 8, 2026. That they focused on the third component is consistent with my impression that the problems LLMs solve successfully are in well-specified more or less closed domains. But, as Abouzaid noted, the creativity takes place before such problems have been identified.

MARTIN HAIRER One thing I noticed, in general, was that the model tended to give a lot of details on the things that were easy, where you would be like: “Yeah, sure, go a bit faster. I’m bored with what you’re saying.” And then it would give very little detail with the crux of the argument. Sometimes it would be like reading a paper by a bad undergraduate student, where they sort of know where they’re starting from, they know where they want to go, but they don’t really know how get there. So they wander around here and there, and then at some point they just stick in “and therefore” and pray.

Sounds like the classic hand-waving — lacking rigor, skipping over complexities.

HAIRER Yeah, it’s pretty good at giving hand-wavy answers.

So, you weren’t impressed?

HAIRER No, I wouldn’t say that. At times I was actually quite impressed — for example, with the way it could string together a bunch of known arguments, with a few calculations in between. It was really good at doing that correctly.

In your dream world, what would the A.I. be doing for you?

HAIRER Currently the output of the L.L.M.’s is hard to trust. They display absolute confidence, but it requires a lot of effort to convince yourself whether their answers are correct or not; I find it intellectually painful. Again, it’s like a graduate student where you don’t quite know whether they are strong or whether they’re just good at B.S. The ideal thing would be a model that you can trust.

KOLDA A.I. is touted as being like a colleague or a collaborator, but I don’t find it to be true. My human colleagues have particular outlooks, and I especially enjoy when we debate different points of views. An A.I. has whatever viewpoint I tell it to have, which is not interesting at all!

Wednesday, January 14, 2026

Can AI generate new mathematical ideas?

Cade Metz, Can A.I. Generate New Ideas? NYTimes, Jan 14, 2026 

Here's how Claude 4.5 summarized the article:

AI Tackles Mathematical Problems - But Questions Remain

An AI startup called Harmonic, using OpenAI's GPT-5 technology, recently solved a mathematical "Erdos problem" - a collection of challenging puzzles designed to test mathematical ingenuity. This sparked debate about whether AI has reached the point of conducting legitimate academic research.

Key Points:

The achievement generated mixed reactions. While some experts see it as evidence that AI can do genuine research, others, including renowned mathematician Terence Tao, argue the solution resembles "a clever student who memorized everything" rather than demonstrating true understanding. The solution was largely based on well-known methods. Initially, OpenAI claimed GPT-5 had solved 10 Erdos problems, but researchers discovered it had merely identified existing solutions buried in decades-old papers. Despite this, the technology proved valuable - it could find obscure sources humans might never locate.

Current Capabilities:

Modern AI systems use "reinforcement learning" to reason through problems, sometimes working for hours. While they can't yet generate truly novel ideas, they've become powerful research tools that can:

  • Analyze and store far more information than humans
  • Suggest hypotheses researchers hadn't considered
  • Help scientists narrow down experiments from 50 to 5

The Verdict:

Experts agree AI is a rapidly improving research assistant, but it still requires experienced human collaborators to guide it, interpret results, and separate useful insights from noise. Whether AI can independently generate breakthrough ideas remains an open question.

Thursday, August 21, 2025

GPT-5 invents a new proof

Monday, August 4, 2025

BygteDance (China) drops automated theorem prover

Sunday, July 13, 2025

Do LLMs require statistical foundations?

Weijie Su, Do Large Language Models (Really) Need Statistical Foundations?, arXiv:2505.19145v2 [stat.ME], June 2, 2025.

Abstract: Large language models (LLMs) represent a new paradigm for processing unstructured data, with applications across an unprecedented range of domains. In this paper, we address, through two arguments, whether the development and application of LLMs would genuinely benefit from foundational contributions from the statistics discipline. First, we argue affirmatively, beginning with the observation that LLMs are inherently statistical models due to their profound data dependency and stochastic generation processes, where statistical insights are naturally essential for handling variability and uncertainty. Second, we argue that the persistent black-box nature of LLMs -- stemming from their immense scale, architectural complexity, and development practices often prioritizing empirical performance over theoretical interpretability -- renders closed-form or purely mechanistic analyses generally intractable, thereby necessitating statistical approaches due to their flexibility and often demonstrated effectiveness. To substantiate these arguments, the paper outlines several research areas -- including alignment, watermarking, uncertainty quantification, evaluation, and data mixture optimization -- where statistical methodologies are critically needed and are already beginning to make valuable contributions. We conclude with a discussion suggesting that statistical research concerning LLMs will likely form a diverse “mosaic” of specialized topics rather than deriving from a single unifying theory, and highlighting the importance of timely engagement by our statistics community in LLM research.

H/t Jessica Hullman:

Something a bit cringey that becomes clearer when you see the various statistical challenges laid out like this is that sometimes they arise not just because LLMs are too complex for us to understand, but also because they are proprietary objects. E.g., once a large LLM has been trained, it’s been found that it can be more efficient to distill its knowledge into a smaller model than to train a smaller model from scratch. This motivates developers of big models to figure out ways to make their outputs resistant to distillation by competitors. It’s all just statistics I suppose, but I’d much prefer to work on problems like uncertainty quantification or watermarking outputs than how to resist sharing knowledge! Similarly, secrecy around training data curation can make it harder to theorize about dependencies between data mixtures and model capabilities.

In reading this alongside other recent takes on the state of stats in ML, it’s interesting to me that despite a growing consensus that we need to develop interpretable models to make sense of LLMs, there still seems to be a contingent of ML researchers who dismiss further integration of classical stats. For example, Su cites evaluation of LLMs as a place where we need statistically grounded methods to avoid an evaluation crisis with similarities to the replication crisis in social science, where researchers game the evaluations they present (there are various reasons to worry about this, some of which we summarized here a few years ago). But others refer to attempts to incentivize more thorough reporting of uncertainty in ML evaluation as “a weird obsession with statistics.” What’s up with that, I wonder?

Monday, June 9, 2025

And the math came tumbling down. o4-mini stuns mathematicians with its prowess [but I’m not worried]

Lyndie Chiou, edited by Clara Moskowitz, At Secret Math Meeting, Researchers Struggle to Outsmart AI, Scientific American, June 6, 2025.

Here's the opening paragraph:

On a weekend in mid-May, a clandestine mathematical conclave convened. Thirty of the world’s most renowned mathematicians traveled to Berkeley, Calif., with some coming from as far away as the U.K. The group’s members faced off in a showdown with a “reasoning” chatbot that was tasked with solving problems they had devised to test its mathematical mettle. After throwing professor-level questions at the bot for two days, the researchers were stunned to discover it was capable of answering some of the world’s hardest solvable problems. “I have colleagues who literally said these models are approaching mathematical genius,” says Ken Ono, a mathematician at the University of Virginia and a leader and judge at the meeting.

And then we go on for several paragraphs about the whole thing was set up. And then we're making progress:

By the end of that Saturday night, Ono was frustrated with the bot, whose unexpected mathematical prowess was foiling the group’s progress. “I came up with a problem which experts in my field would recognize as an open question in number theory—a good Ph.D.-level problem,” he says. He asked o4-mini to solve the question. Over the next 10 minutes, Ono watched in stunned silence as the bot unfurled a solution in real time, showing its reasoning process along the way. The bot spent the first two minutes finding and mastering the related literature in the field. Then it wrote on the screen that it wanted to try solving a simpler “toy” version of the question first in order to learn. A few minutes later, it wrote that it was finally prepared to solve the more difficult problem. Five minutes after that, o4-mini presented a correct but sassy solution. “It was starting to get really cheeky,” says Ono, who is also a freelance mathematical consultant for Epoch AI. “And at the end, it says, ‘No citation necessary because the mystery number was computed by me!’”

Defeated, Ono jumped onto Signal early that Sunday morning and alerted the rest of the participants. “I was not prepared to be contending with an LLM like this,” he says, “I’ve never seen that kind of reasoning before in models. That’s what a scientist does. That’s frightening.”

Although the group did eventually succeed in finding 10 questions that stymied the bot, the researchers were astonished by how far AI had progressed in the span of one year. Ono likened it to working with a “strong collaborator.” Yang Hui He, a mathematician at the London Institute for Mathematical Sciences and an early pioneer of using AI in math, says, “This is what a very, very good graduate student would be doing—in fact, more.”

The bot was also much faster than a professional mathematician, taking mere minutes to do what it would take such a human expert weeks or months to complete.

The future:

By the end of the meeting, the group started to consider what the future might look like for mathematicians. Discussions turned to the inevitable “tier five”—questions that even the best mathematicians couldn't solve. If AI reaches that level, the role of mathematicians would undergo a sharp change. For instance, mathematicians may shift to simply posing questions and interacting with reasoning-bots to help them discover new mathematical truths, much the same as a professor does with graduate students. As such, Ono predicts that nurturing creativity in higher education will be a key in keeping mathematics going for future generations.

I remember back at Johns Hopkins, either 68-69 or perhaps fall of 69. I was talking with Prof. Earl Wasserman, a Romanticist. In 68-69 I took his course on the Romantics; that's where I discovered and became hooked on "Kubla Khan." But I was still working with him in the fall, when I'd begin my master's thesis on "Kubla Khan." Anyhow we were conversing in his office, I might even have been leaning back with me feet up on his desk (really, but I forget how long I kept them up). Anyhow, I made some remark about asking questions and he replied that that's perhaps where the real work is, asking the right questions. I've kept that with me to this day.

Claude and ChatGPT can come up with reasonable answers to a wide range of questions. But they lack agency and so lack the capacity to form questions of their own. And once we manage to endow an AI with agency, how then will it develop its ability to ask questions? Where does that come from? How is it grounded in one's experience of the world?

Consider my most recent working paper: Claude on the Eightfold Way: Sonnet 129, Fantasia, Kubla Khan. I asked for both an esoteric (Hindu/Buddhist) reading of Shakespeare's Sonnet 129. Under what circumstances would Claude come up with that juxtaposition on its own? What prompt would lead it to "think" that such a juxtaposition was a reasonable thing to produce in response to the prompt? For example, imagine that I ask it: Pick a famous poem, any famous poem, and do both a Hindu and a Christian interpretation? What poem would it pick and how would the interpretations go? What would be the point of such an odd open-ended prompt? When I asked for those two interpretations (this was the second time around, first time was with ChatGPT), I had a specific reason for asking.

And so on through the rest of that working paper. In what world is it logical to go from a reading of Sonnet 129 to a reading of Disney's Fantasia? I my world, with my experiences, that's what world. I had a very good reason for making that particular connection, and it proved fruitful. And so with the final connection to "Kubla Khan."

I don't care what tests these chatbots can ace. The fact that they know way more that I do? Not a problem. It's up to me to figure out how I can but that breadth of knowledge to use solving problems which interest me, problems grounded in a half-century of serious intellectual work across a half-dozen disciplines.

But, now things get really interesting. You see, way back in 1976 I'd published an article in which I presented a semantic network model for that Shakespeare sonnet: Cognitive Networks and Literary Semantics. So, I wondered, given that ChatGPT now has visual capacity (I was using 4o to produce a mandala for the title page of that my working paper) what would it do if I asked it to create a semantic network model for Sonnet 129? So:

Let’s look at Shakespeare sonnet in the uploaded document. How would you use a semantic networks model to analyze its semantic structure? That is to say, could you create a semantic networks model that would provide a plausible foundation for the semantics of the sonnet?

Yes, we can construct a semantic network model for Shakespeare’s Sonnet 129—as presented in your document—by identifying the core concepts, relations, and dynamic transformations within the poem. Semantic networks, which consist of nodes (concepts) and edges (relations), allow us to map how meanings emerge through interaction rather than merely cataloging lexical items.

OK, so it knows what a semantic network is. It then proceeded to reason its way toward a model. When it finally coughed up a diagram, that diagram was pathetic. It may have been able to give me a “dictionary definition” of a semantic network, but it didn’t have the foggiest idea of how to produce such a model for a particular stretch of language. That’s something worth thinking about while marveling at how wonderful o4-mini is at solving expert-level math problems. 

When I was a graduate student back in the mid-1970s it took me the better part of a semester to learn how to create semantic network diagrams that modeled reasonable chunks of English. It's not a cookbook type of process. I did that learning while being tutored by David Hays for one session a week. I'd produce a diagram or three, show them to Hays, and he'd explain why they didn't really work. Then we set out doing more adequate diagrams. After three months I began to get the hang of it. I don't recall what was involved in working on Sonnet 129, but I certainly did that well after that initial period of learning. I probably covered 30, 40, or more sheets of paper with diagrams before I settled on the ones I used for that 1976 paper.

What kind of cognitive model was I building up during the period when I learned how to create semantic network structures? What are the characteristics of a model that can move fluently between the verbal and visual domains? Whatever that model is, it isn’t the stuff of rocket science. In subsequent years I had no trouble teaching the basics to undergraduates at the Rensselaer Polytechnic Institute. They were smart, but not likely world-class mathematician smart. It’s one thing to more around fluidly within one domain, but moving back and forth between domains, that’s a whole different ballgame.

Finally, note that the process that led me to be able to create semantic network models was a learning process. I would make proposals, get critiques, revise my understanding and then make new proposals. That process took place over months. Chatbots can't do that, not at the level of the underlying LLM. And that, I suspect, is where the deep learning must occur. Not in during the RLHF (reinforcement learning from human feedback) fine-tuning phase, and certainly not during in-context learning in a single session.

H/t to Tyler Cowen for the link to the article.

Wednesday, January 17, 2024

AlphaGeometry uses a neuro-symbolic architecture to solve Olympiad geometry math problems.

Monday, October 9, 2023

The life and ideas of Miriam Lipschutz Yevick @3QD

My most recent piece for 3 Quarks Daily:

Next year in Jerusalem: The brilliant ideas and radiant legacy of Miriam Lipschutz Yevick [in relation to current AI debates]

It’s about Miriam Yevick, a remarkable thinker. No one has conceptualized the phenomenon of human thought in technical terms more deeply than Yevick did in 1975 (see bibliography below). Her conception was certainly incomplete – she knew that. It may also be wrong, but that's something that has yet to be determined, through analysis and discussion.

Moreover, as I read her memoir, A Testament to Ariela, I realized that she was a remarkable woman as well. This essay is about both her ideas and her life and quotes extensively from one of her technical papers and from her memoir.

On women, Boston Commons, uninterrupted work

A Testament for Ariela consists of letters that Yevick wrote to her young grand-daughter, Ariela. In the following passage Yevick is reminiscing about visiting relatives in Boston and, in particular, about Harvard Yard.

I reminisced. I had entered this yard for the first time, Ariela, in l944, hand-in-hand with your grandfather [that is, Yevick’s husband, George]. We were two physics graduate students at MIT. I had fled a Europe under Nazi occupation a little more than three years before and he had left home in a steel town of Pennsylvania, immersed in the raw struggles caused by the depression. We were dreaming of a better world to come. We believed in science to help bring it about. We slipped into this yard deliberately hard of access, through one of the few entrances and left the monstrous world outside. Here all that ailed the world could be transmuted into abstractions, amenable to fascinating discussions. We listened awe-struck to debates and lectures by the giants of our time: the mathematicians, Norbert Wiener and John Von Neumann, who foretold the power of the computer in society; Niels Bohr, the physicist, who intimated what atomic energy was to create; Philip Frank, the philosopher of science, who stimulated us to argue about the meaning of science; Peter Debye, the biophysicist, who revealed the place of the quantum in biology. They projected to us a radically transformed post-war world, our world to be. We stood for hours with our fellow students on the steps of the bulky library building and commented to each other about what we had heard. Ideas were what made the world move, and if nothing is as powerful as an idea whose time has come, this yard with its old trees, its somewhat decrepit buildings, and its deep-furrowed scholars was to us the quintessential incubator of this new world. We were determined to study hard so we could give it our best!

More recent memories came to me as I sat in the yard: of helping your daddy in the fall of 1971 cart his numerous possessions into the dilapidated Harvard dormitory which was to be his home for the school year. He, too, was to master physics, he dreamed confidently.

We returned some six weeks later in pouring rain to a desperate young scholar. “Now you get into the driver’s seat,” your daddy would say to his father as they took turns in heroically confronting the murderous problem set, which yielded but slowly to their joint onslaught. We left behind a chastened but not defeated aspiring scientist when we said good-bye to our son, who was still bent on making his contribution to a world that one day would be yours, Ariela’s, and Catherine’s and Helen’s.

A bit later:

I lovingly thought of these, my female successors, and recognized my younger self in them. Once again I reminisced about the war years I had spent studying here in Cambridge. Hitler was raging in Europe, murdering my Jewish classmates who had been less lucky than I. America was united in a war effort to defeat this scourge of humanity. There was no time for niceties: Boston was disorderly, dilapidated and rambunctious. The city swarmed with soldiers on leave from the fronts, with sailors taking a respite on shore from the submarine infested oceans. Spirited, newly liberated women of all ages, working in factories and as technicians, replacing the men in combat, energetically mingled in the crowds. Love was made fearlessly in the Boston Commons, passions enhanced by the uncertainty of the service men’s return. Pregnant women stuck to their tasks; there was a job to be done!

Everyone’s life had a purpose. I, Ariela, was going to beat the men at their game and be a physicist. I was to be an independent woman: no man was to spread his protective cover over me again!

The spring term ended; the summer term followed close upon its heels. We had to learn hard and fast. The break between terms was of one week’s duration only. I felt that I needed a holiday from the intense pressures of my studies. My Master’s thesis adviser forbade me to leave: “There is no better way to produce results than by an uninterrupted week of work in the laboratory,” he said.

[Kindle Edition. Locations 663-699]

Remarks Yevick makes here and there make it clear that she often had to scramble for uninterrupted time in which to conduct her research.

Back when I was on the faculty at the Rensselaer Polytechnic Institute it would be two or three weeks after then Spring semester had ended before I’d be in a good frame of mind for research. My brain had to readjust the wiring on millions if not billions of synapses for that to happen.

Some of her papers (the ones I could find)

The Lebesgue Density Theorem in Abstract Measure Spaces, Ph.D. Dissertation, MIT, 1947. https://dspace.mit.edu/handle/1721.1/85715

Probability and Determinism, American Journal of Physics 25, 570 (1957); doi: 10.1119/1.1934551

Holographic or Fourier Logic, Pattern Recognition, Vol. 7, 1975, 197-213. https://doi.org/10.1016/0031-3203(75)90005-9

Two modes of identifying objects: descriptive and holistic for concrete objects; recursive and ostensive for abstract objects, The Behavioral and Brain Sciences, Vol.1, No. 2, 1978, 253-254, https://doi.org/10.1017/S0140525X00074458

Lipschutz-Yevick, Miriam (1991) "Mathematics for Life and Society," Humanistic Mathematics Network Journal: Iss. 6, Article 20. Available at: http://scholarship.claremont.edu/hmnj/vol1/iss6/20

Powell, Arthur B. and Yevick, Miriam L., Letter to The New York Times, Humanistic Mathematics Network Journal: Iss. 19, Article 4, 1999, Available at: http://scholarship.claremont.edu/hmnj/vol1/iss19/4

A Mathematics Manifesto: Think Differently! Think Quantitatively!Quantitative Awareness as a Fresh Thinking Cap, Humanistic Mathematics Network Journal: Iss. 20, Article 20, 1999, Available at: http://scholarship.claremont.edu/hmnj/vol1/iss20/20

Use Your Head: Mathematics As Therapy, Humanistic Mathematics Network Journal: Iss. 22, Article 12, 2000, Available at: http://scholarship.claremont.edu/hmnj/vol1/iss22/12

Semiotic impediments to formalizations, Semiotica, 120(1-2), 1998, doi:10.1515/semi.1998.120.1-2.109

Social Influences on Quantum Mechanics? – II, The Mathematical Intelligencer, 23(4), 2001, 18-22. doi:10.1007/bf03024597

Memoir: To Agnes Berger (1916-2002) and Our Friendship, Newsletter, Association for Women in Mathematics, Volume 33, Number 1, January-February 2003, 7-10.

Poetic metaphor and mathematical demonstration: A shallow analogy. The Mathematical Intelligencer, 29(3), 2007, 11–13. doi:10.1007/bf02985683

More later.