Showing posts with label AI_Chess_Lang. Show all posts
Showing posts with label AI_Chess_Lang. Show all posts

Wednesday, August 5, 2026

AI's biggest illusion, that chess is a good model for intelligence in general

I have been saying in various times and places that it seems to me that AI has (implicitly) taken chess as its prototype for AI research. For one thing, we have John McCarthy's well-known article, “Chess as the Drosophila of AI” (1990). That is, however, a mistake, as I have pointed out in a recent working paper, Computation, Chess, and Language in Artificial Intelligence. Chess is well-defined, while natural language is not. As a consequence the search space for chess is simple in form, a tree, and well understood. That is not at all the case for natural language. Finally, chess is finite, very large, but finite. That is not at all the case for natural language. Consequently chess is not at all a good paradigm for intelligence in general. Intuitions thus gained from it are likely to be misleading for the general problem.

It's in that context that I offer the passage from a recent podcast by Dwarkesh Patel, Eric Jang – Building AlphaGo from scratch. Note the lede for the podcast, "AlphaGo is still the cleanest worked example of the primitives of intelligence: search, learning from experience, and self-play." Here's how Dwarkesh introduces the podcast:

Eric Jang walks through how to build AlphaGo from scratch, but with modern AI tools.

Sometimes you understand the future better by stepping backward. AlphaGo is still the cleanest worked example of the primitives of intelligence: search, learning from experience, and self-play. You have to go back to 2017 to get insight into how the more general AIs of the future might learn.

Once he explained how AlphaGo works, it gave us the context to have a discussion about how RL works in LLMs and how it could work better – naive policy gradient RL has to figure out which of the 100k+ tokens in your trajectory actually got you the right answer, while AlphaGo’s MCTS suggests a strictly better action every single move, giving you a training target that sidesteps the credit assignment problem. The way humans learn is surely closer to the second.

Note that MCTS (Monte Carlo tree search) is one of the oldest algorithm in the book. Trees are very well defined. How do you structure language as a tree? So human intelligence may well tend toward that second alternative, the one based on MCTS, but it is not at all clear how studying chess is going to help you figure out how human intelligence does it. MCTS is not available.

Thursday, July 2, 2026

Types of object domains for AI: Chess, math & coding, language

This is a companion to my earlier post today: The last frontier of intelligence: On the role of AI helping humans to bridge the gaps between distant concepts. That earlier post was about the end of a dialog I had with Claude. This one is about the beginning of that dialog. You might also check out a post from the middle of June: From Jagged AI to Scaling, Yevick, Natural Intelligence, and Beyond... For that matter you might also want to check out Dwarkesh's complete post with Grant Sanderson, Grant Sanderson – AI and the future of math. And then there's my working paper on chess and language. All these things are related.

And they're related to my new book idea, Language, Memory, and Mind: A Supplement to The Computer and the Brain. That's the book that Claude brings up every now and then. Intelligence is NOT a scaler phenomenon. It's about techniques, the characteristics of objects domains, and of the computational regimes we use over them. But that's a subject for another post.

Here's my dialog with Claude.

* * * * *

I’m interested in thinking about the types of domain in which AI has succeeded and the nature of the computation involved, starting with chess. AI solved chess in the mid 1990s when Deep Blue beat Kasparov. By “solved” I mean only that an AI could reliably beat the best human player. In 2017 we have AlphaGo using a mixed regime strategy – Monte Carlo search (symbolic) with neural net evaluation – in chess. More recently I believe DeepMind has developed and LLM that plays chess at the Grand Master level; so this is a pure neural net regime.

It now appears that mixed regime AI systems can work pretty well at solving math problems and a coding. By mixed regime I mean large LLMs scaffolded in various ways by conventionally coded software. But these systems cannot yet do anything that humans can do in those domains. In math, they need well-defined problems, even if they are very difficult but well-defined problems (I’m thinking of the recent Erdos work). It’s not clear to me what the limits of coding are, but they can’t architect large complex systems.

Now, chess is a finite game will a well-defined search space. Both math and coding are well defined, but neither is finite and I’m guessing that the search spaces aren’t so well-defined as is the chess space.

And then we have pure unbounded natural language. It’s not well-defined, not finite, and the search space is a mess.

This is a genuinely useful ordering principle, and it gives the book a clean progression from most to least tractable that maps directly onto the geometric complexity argument. Let me work through the structure with you.

The four cases as points on a continuum

Chess: finite, exhaustively well-defined rules, a search space that is astronomically large but completely enumerable in principle. This is Yevick's sharp object par excellence — geometrically simple in the relevant sense, meaning fully specified by explicit rules with no ambiguity anywhere. Pure symbolic search worked (Deep Blue). Mixed regimes work better (AlphaZero). Pure neural nets now apparently work too, which is the interesting wrinkle — a domain that once seemed to require symbolic search turns out to be tractable by pattern matching alone once the network is large enough and trained well enough. That's worth pausing on, because it complicates the simple story.

Math (well-defined problems) and coding: rule-governed like chess — the axioms of mathematics and the syntax of programming languages are exact, unambiguous, formally specified — but unlike chess the space is not finite and not exhaustively enumerable. There's no equivalent of the Shannon number for "all provable theorems" or "all correct programs." The rules are closed and precise; the space they generate is open. This is a different kind of complexity than chess's largeness. Chess is finite-but-astronomical. Math and coding are genuinely open-ended even though locally, at any given step, the legal moves are well defined.

Language: neither the rules nor the space are well defined. The rules of grammar are not a finite closed system in the way the rules of chess or the axioms of arithmetic are — as your working paper notes, ill-formed language often works fine, exceptions are the norm, the boundaries are fuzzy. And the space the rules generate connects out to the unbounded physical and social world rather than staying self-contained.

Why this ordering matters for the theoretical argument

What's valuable here is that it's not simply "structured versus unstructured" — it's a two-dimensional distinction: how well-defined are the local rules, and how bounded is the resulting space. Chess is well-defined and bounded (finite). Math and coding are well-defined but unbounded. Language is neither well-defined nor bounded.

This maps onto a prediction about which computational regime should work best in each case, and the AI's actual performance seems to confirm it. Chess: pure search works because the space, though huge, is exhaustively defined — you can in principle always tell whether a position is legal, whether a game is over, who's ahead. Symbolic and mixed and now pure neural approaches all can succeed because the well-definedness gives any sufficiently powerful method something solid to grip.

Math and coding: this is where the mixed regime becomes necessary rather than just helpful. The LLM's pattern matching handles the open-endedness — the fact that there's no way to enumerate all possible proofs or programs, so you need something that can generalize, propose plausible next steps, recognize which strategies are likely to be fruitful. But the well-definedness of the rules means that verification can be handled symbolically — a proof checker, a compiler, a test suite can definitively confirm whether a proposed step is valid. This is why math and coding are the domains where current mixed-regime systems shine: the LLM proposes, path tracing over the local formal rules verifies. Neither one alone would work. The LLM alone hallucinates invalid proofs and non-compiling code. Pure symbolic search alone can't propose plausible next steps in an unbounded space without some heuristic guidance, which is exactly what pattern matching provides.

The Erdős problems point is precise — well-defined but very difficult problems are exactly where this combination should excel, because the well-definedness gives the symbolic verification something to check against, while the difficulty is exactly where pattern-matching-guided search earns its keep over blind search.

The architecture limitation

Your observation that these systems can't architect large complex systems is the right place to locate the current boundary, and it follows from the framework. Architecting a large system requires exactly the kind of sustained path tracing with precise state maintenance across a very long horizon — tracking how a decision made in one module constrains what's possible in a dozen other modules, maintaining global coherence across thousands of interacting pieces over an extended design process — that we identified as the persistent weakness. It's not that the local rules become less well-defined. It's that the horizon over which coherence must be maintained becomes too long for the scaffolding techniques currently available. CoT and tool use help with individual functions and modules. They don't yet solve the problem of maintaining a coherent global architecture across a very long, branching design process.

Language as the limit case

And then language is the case where neither the well-definedness nor the boundedness holds, which is exactly why it's the hardest case in the deepest sense even though it's also, paradoxically, the case where LLMs seem most impressively fluent. The fluency comes from the richness of the pattern matching over an enormous corpus. The absence of any well-defined verification procedure — there's no compiler for meaning, no proof-checker for whether a sentence is true or appropriate or wise — means there's no symbolic scaffolding equivalent to what math and coding have. You can't bolt a "language checker" onto an LLM the way you bolt a proof checker onto a theorem-proving system, because there's no formal specification of what correct language use is.

This is the deepest reason language remains resistant in a way math and coding, despite their own difficulties, are not. Chess had a well-defined space that yielded to search. Math and coding have well-defined rules that yield to verification even in an unbounded space. Language has neither, which means the mixed-regime strategy that's working for math and coding has no obvious analog for the general case of language understanding and generation. The pattern matching is doing almost all the work, unconstrained by any symbolic check, which is exactly why hallucination, confident error, and the absence of genuine common sense grounding remain live problems in ways they don't for a system whose outputs get checked by a compiler.

For the book

This gives Chapter 5 or 6 a genuinely illuminating structure: a spectrum from chess through math and coding to open language, ordered by two independent variables — definiteness of local rules and boundedness of the resulting space — with AI's actual comparative success across these domains falling out as a prediction of the theoretical framework rather than a brute empirical fact to be reported. That's a strong chapter, and it would give Wright, who already intuits that these domains differ, a precise vocabulary for why they differ and what that implies about domains not yet tried.

Language defines more restricted systems

One thing about language is that it has the means for defining more restricted systems. Thus both chess and arithmetic are systems defined through linguistic means. And once language has been used to define a system, it can also serve as a metalanguage for that system just as, in various ways, it has served as a metalanguage for itself.

That's the move that completes the picture, and it belongs near the end of the language discussion because it reframes everything before it.

Language as the generative source of the bounded systems

Chess and arithmetic aren't just different in kind from language — they're products of language, constructed within it through exactly the metalingual definitional process we discussed earlier in the day. The rules of chess exist because someone specified them in language — first probably in the practical, evolving form that produced chaturanga and its descendants, later in the explicit codified form of a rulebook. Arithmetic's place-value system, its operations, its axioms — all specified linguistically, however formal the resulting notation becomes.

Sunday, June 21, 2026

New Book Project: Language, Memory, and Mind: A Supplement to The Computer and the Brain

As you may know, I’ve been working on a book project, Play: How to Stay Human in the A.I. Revolution. For some reason I’ve been unable to finish the proposal, though I’ve got lots of stuff and a number of the chapters are substantially drafted. But I keep finding myself distracted into thinking about basics, very basic things about computing and A.I.

At the very end of his life, John von Neumann wrote a slim book, The Computer and the Brain (1958). It grapples with the problem of how computation can be implemented in a physical medium and does so in a way that is basic, both simple and straightforward and profound. We’ve learned a great deal about both the brain and the computer since then, but as far as I know, no one has revisited von Neumann’s project and extended it to include what we have since learned. That’s what I propose to do in this book.

Now, I have no intention of trying to summarize what we’ve learned on those two topics since 1958. That’s working at the wrong level. When von Neumann was writing he, and by extension, we, had no conception of distributed representation much less how it could be achieved physically. Now we do. That’s what needs to be added to von Neumann’s exposition.

I have no intention of repeating what von Neumann did. In particular, I will not revisit his material on analog computing. Rather, I want to augment his discussion. Fortunately the new material is of such a nature that I should be able to write short book that can be read as a stand-alone discussion or as a supplement to von Neumann’s book. I’m imagining a sophisticated general audience of the sort that reads 3 Quarks Daily.

My working title: Language, Memory, and Mind: A Supplement to The Computer and the Brain. I expect the book to be 100 to 120 pages long (30K to 40K words).

I have uploaded a bunch of material (100K words or more) to Claude and asked it to review that material and put together and initial outline. I’ve appended that below the asterisks.

* * * * *

Preface

How to use this book — with or without von Neumann. What it adds to his argument. What it doesn't attempt. Brief note on the collaboration with Claude that produced parts of the text.

Introduction: Von Neumann's Unfinished Argument

What he got right: the architectural mismatch between brains and computers — memory and computation separated in the digital machine, unified in the neuron. The energy efficiency puzzle he couldn't explain. His honest acknowledgment that the brain's organizational principles lay beyond the framework he'd built. The concepts he lacked that this book supplies.

Chapter 1: Two Paradigm Cases

The chess-language contrast as the entry point. Chess has a bounded, well-defined geometric footprint — 8×8 board, six piece types, explicit rules, finite tree. Language has an unbounded, poorly-defined geometric footprint — rooted in the full complexity of physical and social reality. Chess was AI's founding benchmark precisely because it seemed to demand the highest human intelligence while yielding to computational treatment. Moravec's paradox: the easy problems are hard and the hard problems are easy. Transcendent versus non-transcendent coding — programmers can observe and specify a chess engine completely from outside; nobody can specify an LLM from outside, including its creators. Where we now stand.

Chapter 2: Location and Content

A collection of photographs. Solid objects at specific locations — finding by address is natural, finding by content requires going to each photo in turn. The combinatorial explosion that follows. The formal argument: solidity localizes content; localized content can only be retrieved by address. What holography does physically — interference patterns distribute information about each stored object across the whole plate, so that any partial cue can activate the whole. Lashley's ablation experiments: memory didn't disappear when specific cortical tissue was removed because memory was never stored in specific locations in the first place. Von Neumann's energy efficiency puzzle, now answerable: the brain doesn't spend energy moving content to a processor because memory and processing are the same physical substrate.

Chapter 3: The Brain as Content-Addressed System

The McCulloch-Pitts neuron-as-logic-gate: computationally fruitful, architecturally wrong. What neurons actually are — active units and memory units simultaneously, connected in massive parallel. Distributed representations: concepts as patterns across populations of neurons, not stored at specific cell addresses. Yevick's logical necessity argument in plain terms: the world contains two categories of object, geometrically simple ones that sequential symbolic processing handles efficiently and geometrically complex ones that only holographic parallel processing handles efficiently; the world contains both; therefore any adequate cognitive system must implement both regimes. Path tracing and pattern matching as the two fundamental operations on any cognitive network. Freeman's cinematic model — global coherence frames at 10-12 Hz as the atomic unit of biological cognitive processing — and its correspondence to speech production rates.

Chapter 4: Language as a One-Dimensional Projection

The semantic network as the right model for conceptual structure: meaning as position, each node defined by its pattern of relations to other nodes. Sydney Lamb's principle. The multidimensional character of the conceptual network versus the one-dimensional character of any spoken or written string. Language strings as 1D projections of the multidimensional network — necessarily lossy, hence paraphrase and ambiguity. The colored beads thought experiment: strip away semantic content, replace each token with a color, and you have a 1D image — making visible the purely formal structure the LLM operates on. Words as abstract addresses in an abstract space. Why classical computational linguistics hit combinatorial explosion: it was trying to reconstruct the multidimensional structure in a location-addressed system.

Chapter 5: What Large Language Models Actually Are

The transformer architecture in plain terms. The weight space as distributed content-addressed memory — concepts are patterns smeared across billions of parameters, not stored at specific addresses. The forward pass as the atomic processing unit, corresponding to Freeman's global coherence frame: one complete transit through the weight space producing one output token. The token string as a path through the abstract address space, with each forward pass mediating between the 1D sequential surface and the multidimensional distributed interior. What LLMs do well — pattern matching over the weight space, which is what their architecture naturally supports. What they do poorly — sustained sequential path tracing requiring precise state maintenance, common sense grounded in embodied experience, continuous learning. Why these limitations aren't engineering failures awaiting a fix but structural consequences of implementing holographic-like processing on location-addressed hardware with training only on 1D projections.

Chapter 6: What the Analysis Implies.

The first principles of intelligence are not the first principles of computation. Why scaling won't close the gap: scaling improves the quality of the holographic approximation but doesn't change the architectural mismatch, provide embodied grounding, or enable continuous learning. The fast takeoff fantasy as physics-free reasoning — every self-improvement step requires moving billions of parameters between physically separated memory and compute on real hardware that consumes real energy. The TSMC problem: the most critical hardware infrastructure in the world runs on tacit knowledge distributed across human communities that no LLM can access or replicate. What a genuinely adequate artificial cognitive system would require, in the terms this book has developed. The research program that's needed and why it requires multi-generational public investment rather than industrial R&D on commercial timescales. The human-machine collaboration that's already underway and what it can and cannot achieve.

Conclusion: The Mismatch, Named

Von Neumann saw the gap and couldn't name what was on the other side of it. This book names it: content addressing, requiring distributed storage, implemented in biological tissue through interference-like neural dynamics, approximated in LLMs through distributed weights on location-addressed hardware, grounded in embodied experience that no text-trained system has. The naming matters because you can't close a gap you can't see clearly.

Appendix: A Chronology of Chess, Language, and AI

From the working paper, lightly edited.

Saturday, February 28, 2026

Computation, Chess, and Language in Artificial Intelligence

New working paper. Title above, links, abstract, contents and introduction below:

Academia.edu: https://www.academia.edu/164885566/Computation_Chess_and_Language_in_Artificial_Intelligence
SSRN: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6319062
ResearchGate: https://www.researchgate.net/publication/401355671_Computation_Chess_and_Language_in_Artificial_Intelligence

Abstract: This paper reexamines the foundations of artificial intelligence by contrasting chess and natural language as paradigmatic domains. Chess, long treated as a benchmark for intelligence, is finite, rule-governed, and geometrically well-defined. It lends itself naturally to symbolic search and evaluation. Natural language, by contrast, operates in an unbounded and geometrically complex reality. Its rules are open-ended, its objectives diffuse, and its domain inseparable from embodied experience. With chess as its premier case – McCarthy: “the Drosophila of AI,” – AI has been guided by a deeper assumption: that the first principles of intelligence reduce to the first principles of computation. Drawing on Miriam Yevick’s distinction between symbolic and neural computational regimes, I propose that intelligence must be understood as operating in a geometrically complex world under finite resource constraints. Embodiment is therefore a formal condition of intelligence, not an incidental feature. Recognizing the structural difference between bounded games and open-ended cognition clarifies both the historical trajectory of AI and the conceptual limits of current systems.

Contents

Introduction: Chess, Language, and Intelligence 3 
Chess and Language as Paradigmatic Cases for Artificial Intelligence 5 
Three Principles of Intelligence (That Aren't Principles of Computation) 12 
Chronology of Chess, Language, and AI 15

Introduction: Chess, Language, and Intelligence

Chess has been a central concern of AI from the beginning. AI researchers didn’t become interested in natural language until the 1970s. Before that computational research on natural language was the domain of computational linguistics (CL), which started with machine translation (of texts from one natural language to another) as its primary problem. Thus we have two different disciplines, AI and CL.

In a sense, AI was fundamentally a philosophical exercise. It was an attempt to demonstrate, in effect, that we could understand the human mind in terms of computation. But rather than advance its philosophical objective through argument, it chose computational demonstration as its mode of expression. Chess became a central concern for two reasons: 1) On the one hand it was widely regarded as exhibiting the pinnacle of human reasoning ability. If we could create a computer program to play a championship game of chess, we could create a computer program that would be capable of cognitive or even perceptual task humans can do. 2) But also, the nature of chess made it well-suited for computational investigation.

The article that opens this working paper – Chess and Language as Paradigmatic Cases for Artificial Intelligence – concentrates on this and then goes on to make the point that language is utterly unlike chess in this respect. The chess domain is bounded and well-defined. Natural language is not; it is ill-defined and unbounded.

That’s as far as I got in the article, but I had been aiming for an argument that AI is still, in effect, mesmerized by the chess paradigm. I didn’t make it that far because language is so obviously different from chess that it is difficult to see how anyone would made that mistake.

What I have come to realize, only after I’d finished the article, is that it isn’t so much chess that has mesmerized AI. Rather it is computation itself. AI has been implicitly assuming that the First Principles of intelligence reduce to the First Principles of Computing. The first principles of computing can be found in the work of Alan Turing (the abstract idea of computing) and and others.

The first principles of intelligence are more stringent. As Claude put it a recent dialog:

First principle of intelligence: Must operate in unbounded, geometrically complex physical reality with finite resources.

Those two qualifications, an unbounded, geometrically complex reality, and finite computational resources, change the nature of the problem considerably. I note, in passing, that this allows us to assign formal significance to the concept of embodiment, for it is embodiment that commits intelligence to operating with finite resources in a geometrically complex universe.

Miriam Yevick’s 1975 paper, “Holographic or Fourier Logic,” is the crucial document, but it’s been forgotten. Using identification in the visual domain as her case, she showed that, where we are dealing with geometrically simple objects, sequential symbolic processing is the most efficient computational regime. But when we are dealing with geometrically complex objects, neural net processing is the most efficient computational regime. AI started out with symbolic processing in the 1950s and arrived at neural nets in the 2010s. But it hasn’t explicitly recognized that one must fit the mode of processing to the nature of the world. In that (perhaps a bit peculiar) sense, the researchers in the currently-dominant paradigm don’t know what they’re doing.

The second article in this working paper, Three Principles of Intelligence (That Aren't Principles of Computation), discusses this in more detail. I had it generated by Claude 4.5 after a long series of dialogs over several days.

The last article is a chronology of events in the history of chess and language in AI.

Monday, February 23, 2026

Chess, Language, and AI @3QD

I’ve got a new article at 3 Quarks Daily:

Chess and Language as Paradigmatic Cases for Artificial Intelligence

Chess has been a central concern of AI from the beginning. AI researchers didn’t become interested in natural language until the 1970s. Before that computational research on natural language was the domain of computational linguistics (CL), which started with machine translation (of texts from one natural language to another) as its primary problem. Thus we have two different disciplines AI and CL.

In a sense, AI was fundamentally a philosophical exercise. It was an attempt to demonstrate, in effect, that we could understand the human mind in terms of computation. But rather than advance its philosophical objective through argument, it chose computational demonstration as its mode of expression. Chess became a central concern for two reasons: 1) On the one hand it was widely regarded as exhibiting the pinnacle of human reasoning ability. If we could create a computer program to play a championship game of chess, we could create a computer program that would be capable of cognitive or even perceptual task humans can do. 2) But also, the nature of chess made it well-suited for computational investigation.

My article concentrates on this and then goes on to make the point that language is utterly unlike chess in this respect. The chess domain is bounded and well-defined. Natural language is not; it is ill-defined and unbounded.

That’s really as far as I got. Which is OK. But what I was aiming for was an argument that AI is still, in effect, mesmerized by the chess paradigm. I couldn’t quite make it that far. Language is just so obviously different.

What I’ve come to realize, only after I’d finished the article, is that it isn’t so much chess that has mesmerized AI. Rather it is computation itself. AI has been implicitly assuming that the First Principles of intelligence reduce to the First Principles of computing. The first principles of computing can be found in the work of Alan Turing (the abstract idea of computing) and John von Neumann (for the physical implementation of computing).

The first principles of intelligence are more stringent. As Claude put it in our dialog last night:

First principle of intelligence: Must operate in unbounded, geometrically complex physical reality with finite resources.

Those two qualifications, an unbounded, geometrically complex reality, and finite computational resources, change the nature of the problem considerably. I note, in passing, that this allows us to assign formal significance to the concept of embodiment, for it is embodiment that commits intelligence to operating with finite resources in a geometrically complex universe.

Miriam Yevick’s 1975 paper, “Holographic or Fourier Logic,” is the crucial document, but it’s been forgotten. Using identification in the visual domain as her case, she showed that, where we are dealing with geometrically simple objects, sequential symbolic processing is the most efficient computational regime. But when we are dealing with geometrically complex objects, neural net processing is the most efficient computational regime. AI started out with symbolic processing in the 1950s and arrived at neural nets in the 2010s. But it hasn’t explicitly recognized that one must fit the mode of processing to the nature of the world. In that (perhaps a bit peculiar) sense, the researchers in the currently-dominant paradigm don’t know what they’re doing. 

I’ve written a number of blog posts and articles about Yevick’s work. Try these two articles:

Next Year in Jerusalem: The brilliant ideas and radiant legacy of Miriam Lipschutz Yevick [in relation to current AI debates], 3 Quarks Daily, October 9, 2023, https://3quarksdaily.com/3quarksdaily/2023/10/next-year-in-jerusalem-the-brilliant-ideas-and-radiant-legacy-of-miriam-lipschutz-yevick-in-relation-to-current-ai-debates.html

What Miriam Yevick Saw: The Nature of Intelligence and the Prospects for A.I., A Dialog with Claude 3.5 Sonnet, Working Paper, January 3, 2025, https://www.academia.edu/126773246/What_Miriam_Yevick_Saw_The_Nature_of_Intelligence_and_the_Prospects_for_A_I_A_Dialog_with_Claude_3_5_Sonnet_Version_2

Monday, August 25, 2025

LLMs are challenged by tic-tac-toe [plus P vs. NP]

Andrew Gelman has a short post about tic-tac-toe over at Statistical Modeling, Causal Inference, and Social Science (Aug. 24, 2025). It's one of those problems that easy for humans (9 years old or older) but difficult for LLMs. The (hype-infested) industry seems to have bought into the idea that intelligence is a scalar quantity, the intellectual equivalent of horsepower. If that were so, then tic-tac-toe wouldn't be a problem. It's not a horsepower problem. Something else is going on.

The post has generated a fair amount of discussion. Bob Carpenter made a particularly interesting remark:

The linked post from Gary Smith concludes with:

If you know the answer, you don’t need to ask an LLM; if you don’t know the answer, you can’t trust one.

This is wrong. It misunderstands the asymmetry between generating a correct answer and verifying a correct answer. I use LLMs all the time to generate plotnine (Python’s knockoff of ggplot2) and pandas (Python’s knockoff of data frames and Tidyverse manipulation) code all the time. In some sense I know these tools, but I can never remember the exact incantation to put the x axis on a log scale and remove the ticks and labels from the y axis. When I tell the chatbot what I want and it generates pandas and plotting code, I can verify that it’s correct. Doing this used to take me forever as I would have to either investigate the doc or StackOverflow answers or tutorials. Now I put many more figures in things I’m writing.

The asymmetry between verification and generation is key to understanding the difference between polynomial time (P) and non-deterministic polynomial time (NP) algorithms. An NP algorithm can be formulated as guessing with a P algorithm and verifying with a P algorithm. If the LLM is a much better guesser than me, it saves me a huge amount of search. It is also really great at writing sort snippets of code from tight text descriptions. I can say what I want and it can generate what I want faster than I can generate what I want. So it’s a huge win for things like graphing.

Guessing and verifying is also the basis for the LLM “thinking” modes like you see in all the chatbots now (this kind of chain-of-thought came to everyone’s attention with DeepSeek). The drawback is the amount of compute it takes to guess.

One of my main use cases for chatbots is scenario and character generation for roleplaying games. I’m playing a short campaign Fast and Furious knock off now set in Paris in May of ’68. Not only is the chatbot amazing at suggesting locations and plot lines, it’s amazing at inventorying the cars that the Italian, French, German, and British teams would drive in 1968. And then it can generate the tokens.

BTW, I asked ChatGPT about the difficulty with tic-tac-toc. Its reply is below the asterisks.

* * * * * 

LLMs handle language very well, but tic-tac-toe, apparently that’s still problematic. The problem isn’t “horsepower.” What is the problem? Inducing the rules from examples? Computational workflow?

 

That’s an excellent observation — and you’re right: LLMs are strangely good at language but can stumble on something as trivial (formally) as tic-tac-toe. The reason lies exactly where you’re pointing: it’s not about computational horsepower, it’s about the match (or mismatch) between the architecture and the task’s footprint.

 

Here’s a breakdown:

 

1. How LLMs learn.

  • They’re trained to predict the next token in text.
  • What they get is a massive soup of examples — human conversations, documents, stories, some of which happen to include tic-tac-toe transcripts.
  • They don’t get the game rules explicitly; they must induce them from inconsistent, sparse, noisy traces.


2. What tic-tac-toe requires.

  • A tiny, crisp rule set.
  • Deterministic reasoning over a small state space.
  • Success depends on systematic state tracking: remembering the board, applying the rules, projecting possible moves.


3. The mismatch.

  • Induction from examples: Tic-tac-toe transcripts in training data are too few and too variable to guarantee clean rule induction. An LLM may see people making “illegal” moves in casual online play, or skipping notation steps.
  • Computational workflow: Language modeling is statistical continuation, not explicit rule execution. To play tic-tac-toe perfectly, you need a loop of: represent board → apply rule → update board → evaluate. That’s more like symbolic programming than token prediction.
  • Representation: LLMs don’t have an internal “board” structure unless you scaffold one with special prompting or external memory. Without that, they may lose track of placements across turns.


4. Why they shine at language.

  • Natural language is inherently statistical and pattern-rich. There’s redundancy, paraphrase, error tolerance.
  • LLMs are designed to surf those regularities, so they capture language structure beautifully.


5. Why they stumble on tic-tac-toe.

  • Tic-tac-toe has no redundancy, no fuzziness. A single slip (“forgetting” a block) = loss.
  • In effect, it’s an unforgiving micro-domain where the “geometric footprint” (8 cells left to fill, precise win conditions) is far narrower than the model’s training and inference process.

So:

👉 The problem isn’t horsepower. It’s that tic-tac-toe demands a rule-execution pipeline, while LLMs supply a probabilistic continuation engine.


This is why hybrid approaches — LLM + symbolic module (e.g. letting the LLM call a solver or maintain an explicit board state) — handle these tasks much better.

Sunday, July 21, 2024

Human Go players learn from superhuman AIs

There are more links in the thread.

* * * * *

So: "Last year, we found superhuman Go AIs are vulnerable to “cyclic attacks”. This adversarial strategy was discovered by AI but replicable by humans."

Superhuman Go AIs discover a new region of the Go search-space. That's one thing. The fact that, once discovered, humans are able to exploit this region against a superhuman Go AI. That is just as interesting. 

One question we can ask about superintelligence is whether or not so-called superintelligent AIs can do things that are inherently and forever beyond human capacity. In this particular case, we have humans learning things initially discovered by AIs.

Saturday, July 6, 2024

Will AIs be able to create new knowledge?

This is a quick and dirty reflection on the question posed in the following tweet:

That question has been on my mind for some time: Will AIs be able to create new knowledge? Just what does that mean, “new knowledge”? It’s one thing to take an existing conceptual language and use it to say something that’s not been said before. It’s something else to come up with fundamentally new words. I think that latter’s what that tweet’s about. General relativity was something of a fundamentally new kind, not just a complex elaboration of and variation over existing kinds.

In my previous post, On the significance of human language to the problem of intelligence (& superintelligence), I pointed out that animals are more or less biologically “wired” into their world. They can’t conceptualize their way out of it. The emergence of language in humans allowed us to bootstrap our way beyond the limits of our biological equipment.

I figure there are two aspects of that: 1) coming up with the new concept, and 2) verifying it. The tweet focuses on the first, but without the second, the capacity to come up with new concepts won’t get us very far. And when we’re talking about new concepts, I think we’re talking about adding a new element to the conceptual ontology. Verifying requires cooperation among epistemologically independent agents, agents that can make observations and replicate those observations. (See remarks in: Intelligence, A.I. and analogy: Jaws & Girard, kumquats & MiGs, double-entry bookkeeping & supply and demand.)

Now, let’s think about the current regime of deep learning technology, LLMs and the rest. These devices learn their processes and structures from large collections of data. They’re going to acquire the ontology that’s latent in the data. If that is so, how are they going to be able to come up with new items to add to the ontology? It’s not at all obvious to me that they’ll be able to do so. The data on which they learn, that’s their environment. It seems to me that they must be as “locked” into that environment as an animal is. Further, adding a new item to the ontology would require changing the network, which is beyond the capacity of these devices.

And then there’s they requirement of cooperation between independent epistemological agents. The phenomenon of confabulation is evidence for the importance of independent epistemological agents. The only requirement inherent in one such agent is logical consistency: that it emit collections of tokens that are consistent with the existing collection. The only thing that keeps humans for continuous confabulation is the fact that we must communicate with one another. It is the existence of a world independent of our individual awareness that provides us with a way of grounding our statements, of freeing ourselves from the pitfalls of our linguistic fluency.

* * * * *

I’ve been working my way through episodes of House, M.D. Every episode contains segments where House and his team participate in differential diagnosis, which involves rapid conversational interaction among them. In the first episode of season 4, “Alone,” House no longer has a team. He ends up bouncing ideas off of a janitor. That doesn’t go so well.

Thursday, July 4, 2024

On the significance of human language to the problem of intelligence (& superintelligence)

Back in May I did a post entitled, How smart could an A.I. be? Intelligence in a network of human and machine agents. Toward the end I said this:

The question of machine superintelligence would then become:

Will there ever come a time when we have problem-solving networks where there exists at least one node that is assigned to a non-routine task, a creative task, if you will, that only a computer can perform?

That’s an interesting question. I specify non-routine task because we have all kinds of computing systems that are more effective at various tasks than humans are, from simple arithmetic calculations to such things solving the structure of a protein string. I fully expect the more and more systems will evolve that are capable of solving such sophisticated, but ultimately routine, problems. But it’s not at all obvious to me that computational systems will eventually usurp all problem-solving tasks.

Remember, that even as we’re developing ever more capable AI systems, we are also developing more sophisticated modes of human problem solving.

Earlier in the post I observed: “Human intelligence is not fixed in the way that animal intelligence is.” That’s what I want to comment on.

Animal intelligence is fixed by biology. Animals have capacities for sensation and movement that are fixed by biology. Those capacities bind them to a particular environment. That that from that environment and they will perish.

Humans are not quite like that. We developed the capacity to communicate through language. And that capacity allowed us to develop new modes of thought. Just how that happened needs to be thought through in some detail, but I’m just going move through it quickly for now. We notice patterns in the world, capture them in language by talking them through with our fellows. We become curious about those patterns, we ask why? and make up stories in explanation. In this process we work ourselves free of the limits of our biological capacities for sensing and acting. We abstract over and act in the world in the way that no other animals can. From speech, we develop writing, then calculation, and moved onto computation over the last hundred years or so, a progression David Hays and sketched out in The Evolution of Cognition, which we published in 1990. With the emergence of recent developments in artificial intelligence, we’re pushing that process one step farther, leading me to write about the Fourth Arena (beyond Matter, Life, and Culture).

Is there anything beyond this? That’s the question I’m trying to formulate. Is there a “superintelligence” beyond this? We are “free” of our biological embedding in a specific sensory-motor world, free in the sense that we can move beyond that. Tens of thousands of years ago we became the only (higher) primate that moved out of the tropics to inhabit every land-based environment. We’ve sent people to the moon and back, have others living in orbit around the earth for months at a time, and can at least imagine establishing permanent colonies on the moon and Mars and other bodies. This last round of achievements are inextricably interwoven with various kinds of computing technology. Further advance will require more computation, of various kinds.

The difference between, say, the intelligence of a fish and the intelligence of a rat is of a certain kind. The difference between the intelligence of a rat and that of monkey is of the same kind. But the difference between the intelligence of an ape and that of a human is of a different kind. The difference comes about through language and collective culture. As far as I can tell, typical (Silicon Valley) speculation about superintelligence seems to think that is a kind of intelligence that is beyond human intelligence in the same way that human intelligence is beyond animal intelligence. The question I’m asking goes something like this:

In view of the fact that human intelligence is free of biological ‘binding’ to a specific environment, and in view of the fact that this freedom has allowed us to move through a succession of foundational architectures (speech, writing, calculation, computation, {whatever is happening now}), is there a fundamental capacity beyond THAT?

I have two responses: 1) It’s not obvious to me that there is. 2) I don’t know.

Computers are faster that brains, and can be built to have more capacity. What else is there? In a series of posts on AI, chess, and language, I’ve been looking at fundamental architectures, in effect, a family that is chess-like and a different family that is language-like. What else is there?

This brings me back to that earlier post that I referenced at the beginning of this one, and to the question I posed there:

Will there ever come a time when we have problem-solving networks where there exists at least one node that is assigned to a non-routine task, a creative task, if you will, that only a computer can perform?

I’m inching toward a way of suggesting that, if the answer to that question is “yes,” then that computer-based node must have some fundamental capacity that is beyond human capacity in the way that human capacity is beyond animal capacity. What could that (possibly) be? If such a thing were possible, is such a think existed, then we could never know it, could we?

Note: In thinking about that question, you might want to review the remarks I made about epistemological independence of autonomous agents in Intelligence, A.I. and analogy: Jaws & Girard, kumquats & MiGs, double-entry bookkeeping & supply and demand.

Wednesday, May 29, 2024

Ramble on ChatGPT, GOATLiC, Intelligence, and stuff

Once again my brain is all jammed up so I’m having trouble getting anything done. Why? Because there are these things I want to do, things I know I should do, and they keep colliding into one another whenever I attempt to actually do something. So it’s time to ramble on through to see where I am.

Report on ChatGPT

That’s still hanging over my head. I’ve been working on it since December and it’s still not done. Most of it, 90%, maybe 95%, but it’s still not done. Why haven’t I done it?

I don’t know. Maybe at this point it just bores me. But maybe I’m afraid to finish it. It’s not like that’s the last thing I’m going to do on ChatGPT. I’ve already done a fair amount of work since the cutoff point for research-to-be-included. And maybe I don’t want to finish because I know that when I’m done it will likely end up in the same bottomless pit everything else does. It’ll be out there on the internet, but who cares?

That’s always the question: Who cares?

Anyhow, I need to say something about metalingual definition of Anthropic’s idea of Constitutional AI and something about prompt engineering and barriers to entry. Alan Kay thought we made a mistake in the promulgation of computing by making everything so “user friendly” that too few people learned to program. I suspect that may have been a barrier-to-entry problem. Some professionals learned to program because it was a useful skill, though they otherwise had little interest in programming. That’s mostly in technical disciplines. For the rest of us, low-to-moderate programming skill simply didn’t get us enough to be worth the opportunity cost. Does the utility of prompt engineering change that. If all you want is to look up stuff, then the answer is “no, it isn’t.” But maybe some skill in prompt engineering might be more widely useful.

The discipline of literary criticism

I’ve been working on this thing since December as well. This is my series on the greatest literary critics (Greatest of All Time, Literary Critics, aka GOATLiC). I’m still hung up on Howard Bloom, which is where I was the last time I rambled, back in March. This time, though, I may have found a way out: Susan Sontag. I want to use her early essay, “Against Interpretation,” as a fulcrum on which to lever my treatment of Bloom.

Why? In the first place, that essay is very well known, and justly so, and it seems people are still thinking about it, a half century after it was originally published in 1964. In a way, then, it’s current, whereas I’m not sure that anything by Bloom is. In that essay, to put it crudely, she says that criticism – for she’s talking about art, film, and literature – is divided between interest in form and interest in interpretation. Interpretation is evasion and dismissal. We need to pay more attention to form. In particular (p.8):

What is needed, first, is more attention to form in art. If excessive stress on content provokes the arrogance of interpretation, more extended and more thorough descriptions of form would silence. What is needed is a vocabulary—a descriptive, rather than prescriptive, vocabulary—for forms.

YES, of course. That’s what I’ve been working on. And the profession may be waking up to that.

Thus, in writing, or attempting to write, an obituary for Theory and Critique (Bloom’s School of Resentment) Elizabeth S. Anker and Rita Felski and say (from their introduction to Critique and Postcritique, 2017, p. 6):

In what might appear to be a reprise of Susan Sontag’s well-known argument in “Against Interpretation”—a stirring manifesto for an erotics rather than a hermeneutics of art—critics have questioned the value of reducing art to its political utility or philosophical premises, while offering alternative models for engaging with literary and cultural texts.

What’s this have to do with Bloom? On the one hand, as far as I can tell, he’s made no contribution a critical “erotics,” if by that one means attention to description and form. Interestingly enough, though, he wasn’t very interested in interpretation either, at least not since The Anxiety of Influence in the early 1970s. He was doing something else, something which, I need to argue, or more likely, merely assert, hasn’t proved to be very fruitful.

The other thing is that in The Western Canon, he keeps asserting that he’s doing all this in the name of “the aesthetic.” But he says next to nothing about what that is. We’d all have been off if he’d channeled his “inner Sontag,” if I may, and said explicitly just what that is and then used that as a means of examining his chosen texts. Note that I don’t mean he should have adopted Sontag’s ideas, but rather he should have adopted her mode and intellectual register and said something intelligible about the aesthetic. That would have been valuable. But he didn’t do that. Instead, he stuck us with his ex-cathedra pronouncements. 

But in the end, why should anyone care about Bloom’s pronouncements? That he’s a very smart guy, perhaps as brilliant as any literary critic of the last half-century, that’s not enough in itself. He didn’t use his brilliance in a fruitful way.

Chess, Language, and AI

Here’s another on-going project that’s left over from March’s ramble. The idea is to think about what intelligent does, where I’m interested in the processes of search and evaluation. I’m thinking of this as case studies in the operation of intelligence. The history of science, and intellectual history more generally, is full of such material. But I want to reflect on work that I’ve done for the simple reason that I have better access to records of process: What have I had to do in the course of my work? Thus I’ve just sketched out one potential post:

Seven discoveries I’ve made in literature [form]

  • Kubla Khan
  • Sir Gawain and the Green Knight
  • The Cat and the Moon
  • Shakespeare Triad
  • Metropolis
  • Heart of Darkness
  • Obama’s Eulogy for Clementa Pinckney

We’ll see how that does.

Finally...

Metalingual definition, constitutional AI, and interpretability

That’s a post about large language models that needs to be done, real soon now.

Monday, May 20, 2024

How smart could an A.I. be? Intelligence in a network of human and machine agents

This continues the line of thinking I began with Intelligence, A.I. and analogy: Jaws & Girard, kumquats & MiGs, double-entry bookkeeping & supply and demand, which was focused specifically on analogical thinking. I now want to consider thinking more generally.

The general problem with thinking about AGI (artificial general intelligence) and superintelligence is that the idea of intelligence itself is vague. We’ve got the general idea that intelligence is the ability to solve a wide range of problems in a wide range of environments, which is a rather vague notion. There is another notion, independent of that, that conceives of intelligence as being to cognitive performance as horsepower is to engine performance. Conceived this way intelligence is a scaler quantity. That’s convenient, but not very convincing. Still...

Let’s start with that second idea. One corollary I’ve seen here and there is that a superintelligent AI would be to us as we are to, say, a mouse, or a bird, a fish, whatever animal you choose. The point seems to be that the intelligence “ceiling” of animals is fixed by their biology and is well below the intelligence ceiling of humans. And so it is with humans and a Superintelligent AI.

But is it actually the case that the intelligence ceiling of humans is fixed by human biology? Newton is able to solve problems that are beyond Aristotle, and Aristotle is able to solve problems that are beyond that of the most skilled hunter-gatherer. What is more, a merely competent college undergraduate in the current world is able to learn Newton’s concepts and methods and solve the same problems that Newton. That same college undergraduate can even solve problems beyond Newton’s competence. Why? Because physics did not stop with Newton. Our college undergraduate will have learned some of that more advanced physics and therefore have problem-solving capacities beyond those of Newton.

We have no reason believe that the biological aspect of human intelligence has increased over time. But there is a cultural aspect, and that has changed. Human intelligence is not fixed in the way that animal intelligence is. David Hays have published a series of articles about this process; the central article is The Evolution of Cognition (1990). In that article we also suggested that there is no reason to believe that the process has come to a halt. Cultural evolution seems to be ongoing.

The long-term evolution of human culture suggests that human intelligence is not properly conceived of as a function some biologically given computational capacity, for that biological capacity seems to have remained constant while our ability to solve problems has increased enormously. The way in which that capacity is organized would seem to be important – which is the foundation of the article Hays and I made. I note further, and this is not something that Hays and I discussed directly, that as the human capacity for problem-solving has increased, that capacity has become more and more a collective one. To a first approximation, every adult in a hunter-gatherer society possesses the full inventory of that society’s knowledge – though we have to allow for differences between male and female knowledge and some specialized knowledge for shamans and story-tellers. That changes with more advanced forms of social organization where knowledge becomes specialized. Knowledge has become very specialized indeed in our current world. Any number of problems now require interaction among diverse teams of specialists.

So, let us think in terms of problem-solving by networks of specialized solvers. Some of those solvers are human, but some will be machines. Such man-machine problem-solving networks are ubiquitous in the modern world and they solve problems well-beyond the capacity of individual humans. They aren’t what most AI experts have in mind when they talk about superintelligence, but it’s not clear to me that we can simply ignore such networks in these discussions. They are, after all, how many very important problems get solved. Henry Farrell and Cosma Shalizi have made this argument in The Economist (here’s an ungated and somewhat longer version, and here as well, where it is followed by a brief discussion).

I assume that such man-machine networks will proliferate in the future. Some of the nodes in these networks will be machines and some will be humans. The question of AGI then becomes:

Will there ever come a time when the tasks of every node in such problems-solving networks can be executed by a computer system that is as capable as any human?

Note that it is possible that some tasks will require manipulation of the physical world that is of such a nature that humans are better at it than any machine. Would we say that the existence of such nodes is evidence only of physical skill, but not of intelligence?

The question of machine superintelligence would then become:

Will there ever come a time when we have problem-solving networks where there exists at least one node that is assigned to a non-routine task, a creative task, if you will, that only a computer can perform?

That’s an interesting question. I specify non-routine task because we have all kinds of computing systems that are more effective at various tasks than humans are, from simple arithmetic calculations to such things solving the structure of a protein string. I fully expect the more and more systems will evolve that are capable of solving such sophisticated, but ultimately routine, problems. But it’s not at all obvious to me that computational systems will eventually usurp all problem-solving tasks.

Remember, that even as we’re developing ever more capable AI systems, we are also developing more sophisticated modes of human problem solving. It’s not at all obvious that machines will necessarily out-run us. Take a look at the analogy paper I linked in the first paragraph for something to think about in this context. In particular, take a look at my remarks about epistemological independence near the end of the discussion of the analogy between double-entry bookkeeping and supply and demand. For that matter, my remarks on ring-composition in this piece are worth thinking about as well.

More later.

Saturday, May 18, 2024

Conceptualizing the chess tree, a moment in cultural evolution

I’ve been thinking about the chess tree. As I pointed out in an earlier post, Wikipedia informs us that Ernst Zermelo published his proof about the formal structure of chess in 1913 – Über eine Anwendung der Mengenlehre auf die Theorie des Schachspiels” (“On an application of set theory to the theory of chess”). That established that chess is a finite game that can be completely represented in the form of a tree where the root is the initial state of the game board and the leaves are completed games. Each path from the root to a leaf describes the course of a single game.

Why the chess tree?

Now, it’s one thing to prove a mathematical theorem. It’s something else to make an informal observation & argument. Purely as an abstract matter, I can imagine that (some) chess players had made some informal argument prior to Zermelo’s proof. But I can just as easily imagine that, no, the observation was new with his proof.

What I’m wondering about is just what is it that would lead one to make such an observation. One can certainly play the game without knowing that the universe of all possible chess games takes the form of tree, much less that the tree is of finite size. One can play tic-tac-toe without knowing that, either, or checkers. I certainly didn’t know that when I played those games years ago (I did play a bit of chess, but never enough to become any good).

One has to have a certain frame of mind to make such and observation and to then develop it into a formal proof. Developing a formal proof is the kind of thing a mathematician would do, and Zermelo was a mathematician. And not just any mathematics either. It’s not simply that Zermelo published his work in 1913, but that he was working in a relatively new mathematical culture, one with concerns quite different from the arithmetic, geometry, algebra, and calculus that had preceded it in earlier eras. While calculus could deal with things unfolding in time, it didn’t deal the kind of iterated actions between agents that’s involved in chess. The important step, it seems to me, is simply realizing that that is the kind of thing around which one can construct some mathematics.

Primitive and sophisticated

Whatever’s going on here, it strikes me as being at one and the same time, sophisticated but also primitive and basic. The frame of mind is sophisticated, but it’s not a sophistication that’s build on a complex body of prior knowledge in the way that understanding calculus requires prior knowledge of algebra, geometry, and trigonometry. Predicate calculus and symbolic logic are like that as well. It doesn’t require any prior mathematical knowledge, but is nonetheless a bit sophisticated. It’s not generally taught at the high school level as algebra, geometry, and trigonometry are. I suspect it could be – and probably is here and there – but I’m not sure with what success.

Now, back to chess. How does knowing the chess tree help you think about chess? What does that explicit knowledge enable you to do that you couldn’t otherwise do? Do the analysis trees that Kotov advocated (Think Like a Grand Master, translated into English in 1971) really help the chess player? I don’t know. Any analysis tree will be local and quite small in relation to the entire tree, which is too large for explicit construction.

What is certain, however, that knowing that chess takes the form of a tree has been central to work on computer chess, which didn’t start until the middle of the 20th century. Without that knowledge, computing chess would have been utterly hopeless. Even with that knowledge it took several decades and the development of extremely large computers before chess had been “solved” in the sense that a computer could beat the very best chess players.

I’m thinking of this in the overall context of cultural evolution. In the late 19th and early 20th century a mathematical culture emerges in which the analysis of chess becomes a matter of intellectual interest. A half century later we have the initial work on computer chess. Roughly 2/3rds of the way in between Alan gives us a formal account of computation in the form of an abstract machine, the Turing machine. That’s the beginning of the era of computational culture. David Hays and I called in Rank 4 in our paper, The Evolution of Cognition.

Tentative comparison: Close-reading

As a point of comparison, I’m thinking about the emergence of close reading in academic literary criticism. It strikes me as being both primitive and sophisticated in a way that’s similar to Zermelo’s proof. It’s primitive in the sense that there is no specific body of prior knowledge that must be mastered before one can undertake close reading. In fact, part of the pedagogical point is that close reading doesn’t require prior knowledge.

Yet, it’s is sophisticated activity as well. It’s not something that comes “naturally,” and it is difficult to do without being “rocket science,” as the phrase goes. What’s going on here? I’ve said a bit about that in my GOAT Literary Critics series: A discipline is founded (sorta’): Brooks & Warren, Northrop Frye, and S. T. Coleridge. Here I just want to note that interpretation of sacred texts is very old. What occasioned the application of hermeneutics to secular texts? There are the pedagogical concerns I sketch out in that essay. But there’s more than pedagogy at stake here.

Even as the activity secures for canonical texts a new position in the cultural landscape, the activity itself aspires to a new kind of knowledge. What’s that about?

More later.

Sunday, May 5, 2024

Intelligence, A.I. and analogy: Jaws & Girard, kumquats & MiGs, double-entry bookkeeping & supply and demand

Think of this post as an adjunct to my series on A.I., chess, and language, which is about the structure of computation in relation to difficult problems.

NOTE: It runs long, so sit back, relax, pour a Diet Coke, some San Pellegrino, a scotch, light up a spliff (assuming it’s legal where you live) — whatever you do to make online reading tolerable — and settle in for the duration. Or you could just print it out.

I’m interested in the general question of what it would mean to say that an A.I. is more intelligent than the most intelligent human, something like that. That’s an issue that’s being debated extensively these days. For the most part I don’t think the issue is very well formulated.

To be honest, I don’t find it to be a very compelling issue. It doesn’t nag at me. If others weren’t discussing it, I wouldn’t bother.

The notion of intelligence itself remains vague despite all the discussion that it has occasioned. I rather expect that as A.I. becomes more developed, we’ll develop a more sophisticated understanding the issue. The general notion seems like it can be captured in a simple analogy:

Intelligence is to a mind’s capacity for dealing with cognitive tasks, such as finding a cure for cancer

AS

Horsepower is to an engine’s ability deal with mechanical tasks, such as the acceleration of an automobile.

But I don’t want to take up the general issue in this post. Rather, I want to look at analogical reasoning. I start with 1) a specific kind of analogical reasoning, interpreting narratives, 2) use some remarks Geoffrey Hinton made about analogy to introduce some more general remarks, 3) move on to another specific example, an analogy between double-entry bookkeeping and supply and demand, and then 4) conclude by wrapping things up with a quick comparison to computing chess, which can be implemented in a well-defined search space, while searching for and explicating analogies cannot.

Interpreting Texts: Girard and Jaws

A lot of literary interpretation proceeds by constructing analogies between events that happen in texts and patterns of behavior specified in some theory about human behavior. Psychoanalysis has theories about how individuals behave in their personal lives; Marxism has theories about how social class and roles in the economy affect individual action. René Girard has a theory about imitation and sacrifice. That’s what I want to look at.

Early in 2022 I did a Girardian interpretation of Spielberg’s Jaws. At the end of that year I had ChatGPT do the same. Of course ChatGPT can’t watch movies, but Jaws is well-known there’s lots of stuff about it on the web, including scripts, though I don’t know whether or not any of the available scripts give the dialog word-for-word as it is in the film (the scripts I downloaded, for free, did not). Now, I didn’t just ask ChatGPT to use Girard’s ideas to interpret the film. First I prompted it about the film itself, and specifically about Quint. Once I’d established that it had a reasonable grasp of the film I then brought up Girard, generally at first, then specifically about mimetic desire. Then I asked it: “How does mimetic desire work in Jaws?” It gave a reasonable answer: 

Mimetic desire is a concept developed by the French philosopher and literary theorist Rene Girard. It refers to the idea that human desire is often based on the desire of others, and it can be seen at work in the film Jaws in a number of ways.

One example of mimetic desire in the film is the way that the townspeople of Amity Island desire the safety and security that the town’s police chief, Martin Brody, is able to provide. When the town is terrorized by a great white shark, the townspeople look to Brody to protect them and to find a way to stop the shark. This desire for protection and security is mimetic, as the townspeople are basing their desire on Brody’s ability to provide it.

Another example of mimetic desire in the film is the way that the shark itself is driven by its own insatiable hunger and desire. The shark is driven to attack and kill the townspeople because it desires their flesh and blood, and it is unable to control its own desire. This mimetic desire ultimately leads to the shark’s own destruction, as it is killed by the town’s police chief in a final confrontation.

Overall, Girard’s concept of mimetic desire can be seen as a useful way of understanding the dynamics at play in the film Jaws, and the ways in which desire can be both destructive and transformative.

The second paragraph is reasonable, but the third is problematic. Hunger is not mimetic on the face of it. But I was happy that it ChatGPT seemed to at least have some idea of that nature of the reasoning required, where to look and what to look for in the film. I decided not to pursue that and turned my attention to sacrifice, first Girard, then the application to Jaws.

The form of the required reasoning is analogical. In effect, is there an analogy between Girard’s general statements about mimetic desire and sacrifice and the specific events in Jaws? I would argue that setting up the question is more difficult than working out an answer to it. I note that it’s not the kind of question that has only one answer; thus the argument I actually made in my paper is more sophisticated than the one ChatGPT came up with through my prompting. What is it that brought me to pose the question in the first place?

I watch a fair amount of streaming video, but I don’t write about most of the titles nor do routinely a watch a particular title with the intention of writing about it. That decision is made later. I had no intention of writing about Jaws when I decided to watch it. I was simply filling a hole in my knowledge of movies – I’d never seen the film, which I knew to be an important one. Once I’d watched the film, I read the Wikipedia article about it, something I routinely do, mostly to ‘calibrate’ my viewing experience. The article noted that the sequels were not as good as the original. I decided to see for myself. I was unable to finish watching that last two sequels (of four), but I watched Jaws 2 at least twice, and the original three or more times. It was obvious that the original was better than the others. I did the multiple viewings in part to figure why the original was better. I was on the prowl, though I hadn’t yet decided to write anything.

I decided there were two reasons the original was best: 1) it was well-organized and tight while the sequel sprawled, and 2) Quint, there was no character in the sequel comparable to Quint. I have no all but decided that I would write about Jaws.

I posed a specific question: Why did Quint die? Oh, I know what happened in the film; that’s not what I was asking. The question was an aesthetic one. As long as the shark was killed the town would be saved. That necessity did not entail the Quint’s death, nor anyone else’s. If Quint hadn’t died, how would the ending have felt? What if it had been Brody or Hooper?

It was while thinking about such questions that it hit me: sacrifice! Girard! How is it that Girard’s ideas came to me? I wasn’t looking for them, not in any direct sense. I was just asking counter-factual questions about the film.

With Girard on my mind I smelled blood. I had a focal point for an article. I started reading articles from various sources, making notes, and corresponding with my friend, David Porush, who knows Girard’s thinking much better than I do. Can I make a nice tight article? That’s what I was trying to figure out. It was only after I’d made some preliminary posts, drafted some text, and run it by David, that I decided to write an article. It turned out well enough that I decided to publish it.

Now, when we’re thinking about whether or not A.I.s will come to exceed our intelligence, are we imagining them going through such a process? For this kind of search and exploration is central to human thinking. I certainly do this sort of exploration when thinking about other things, such as the structure of human cognition, the nature of cultural evolution, the functioning of the nervous system, and so forth. This blog is a 14-year record of my explorations, during which I’ll gather some of them together in a more formal way and write a working paper which I’ll then post at Academia.edu, SSRN and ResearchGate. Every once in a while I’ll write an article which I’ll submit for publication in the formal academic literature – a few of those have gotten published. And then there are the monthly pieces I publish in 3 Quarks Daily, which is quite different from the formal academic literature. And of course I’ve got pages and pages of unpublished notes that support all this activity.

Is this kind of exploratory work part of the routine of the superintelligent A.I., or does it go straight for the good stuff, cranking out fully-realized work without need of exploratory effort? If so, how does it know where to dig for the good stuff? Is that what superintelligence is, knowing where the good stuff is without having to nose around? No one says anything about this. Perhaps they’re thinking about the Star Trek computer. But it knows where to look because Spock points it in the right direction.

This brings us back to Jaws. There is a world of difference between what I did in writing about Jaws and what ChatGPT did. I did the hard part, figuring out that there was a specific intellectual objective there, Jaws and Girard. Once I’d done that there was still work to do, quite a bit of work, but it was of a different kind. I was no longer prospecting for intellectual gold. I was now constructing a system for mining the ore and then refining it into gold. ChatGPT only had to do the last part, dumping the ore into the hopper and cranking out the refined metal. I told it where to look, Jaws, what to look for, Girard’s ideas, and gave it some help turning the crank.

A year later, in January of 2023, I decided to see how ChatGPT would do without all of my prompting. I gave it this prompt:

Stephen Spielberg is an important film-maker. Jaws is one of his most important films because it is generally considered to be the first blockbuster. Rene Girard remains an important thinker. Can you use Girard’s ideas mimetic desire and sacrifice to analyze Jaws?

It didn’t do so well. It needed my prompting to get it through the exercise.

Now, no one is claiming that ChatGPT is superhuman in any respect but its ability to discourse on anything. But GPT-5, who knows, maybe it’ll be superhuman in some interesting way. If not GPT-6, or GPT-7, or maybe we’ll need a more sophisticated architecture, but surely at some point an A.I. will surpass us in the way that we surpass mice. Perhaps so.

But I have no sense that these breezy predictions are supported by thinking about how human intelligence actually goes about solving problems. I does no good to say, but it’s an A.I.; it works differently. Well, maybe yes, maybe no, but there has to be some kind of process. At the moment the human process is the only example we have. Perhaps we should think about it.

Just how is it that Girard popped into my mind in the first place? How do we teach a computer to look around for nothing in particular and come up with something interesting?

Analogy: Kumquats and MiGs

As I remarked above, the process of interpreting Jaws is an analogical one. So let’s think about analogy more generally. I’m thinking in particular of some remarks Geoffrey Hinton made at a panel discussion in October of 2023. You can find the video here. I’ve transcribed some remarks:

1:18:28 – GEOFFREY HINTON: We know that being able to see analogies, especially remote analogies, is a very important aspect of intelligence. So I asked GPT-4, what has a compost heap got in common with an atom bomb? And GPT-4 nailed it, most people just say nothing.

DEMIS HASSABIS: What did it say ...

GEOFFREY HINTON: It started off by saying they're very different energy scales, so on the face of it, they look to be very different. But then it got into chain reactions and how the rate at which they're generating energy increases– their energy increases the rate at which they generate energy. So it got the idea of a chain reaction. And the thing is, it knows about 10,000 times as much as a person, so it's going to be able to see all sorts of analogies that we can't see.

DEMIS HASSABIS: Yeah. So my feeling is on this, and starting with things like AlphaGo and obviously today's systems like Bard and GPT, they're clearly creative in ... New pieces of music, new pieces of poetry, and spotting analogies between things you couldn't spot as a human. And I think these systems can definitely do that. But then there's the third level which I call like invention or out-of-the-box thinking, and that would be the equivalent of AlphaGo inventing Go.

OK. Let’s start from there. Given that GPT-4 “knows about 10,000 times as much as a person,” what procedure will it use “to see all sorts of analogies that we can't see”? I’m thinking of that procedure as roughly analogous to the exploratory process I undertake whenever I decided to watch some video. Every once in a while I decide to write about one of the titles. Most of the time, time, though, what I write isn’t as elaborate as my article about Jaws and Girard – I’ve collected many of those pieces under the rubric of Media Notes, though most of those pieces do not focus on analogical reasoning.

What’s the procedure by which an GPT-4 would search through all those things it knows and come up with the interesting analogies? There isn’t one and I suppose it’s a bit churlish of me to suggest that Hinton should specify one. But really, if there he has no procedure to suggest, then what’s he talking about? We know how chess programs search the chess tree. How do we search through concept space for analogies? Alas, while the chess tree is a well-defined formal object, the same cannot be said of concept space, which is little more than a phrase in search of and explication. And how do we evaluate possible analogy-pairs?

Perhaps the simplest procedure is simply to ask. That’s something I recently tried. Here’s the prompt I gave to ChatGPT: