Showing posts with label neurosymbolic. Show all posts
Showing posts with label neurosymbolic. Show all posts

Thursday, August 6, 2026

What we’ve got in frontier models is now neurosymbolic

Monday, April 13, 2026

Language is a lower-dimensional projection of high-dimensional neural dynamics.

But it also allows for content addressed memory. That’s very important, for it gives fine-grain control over the memory and planning systems. That’s the job of sentence-level syntax together with discourse structure.

“Classical” semantic or cognitive networks had a problem with coming up with just the right set of node types and arc types. David Hays dissolved the problem in his 1981 book, Cognitive Structures (scan down the page), by grounding cognition in an analog system modeled on William Powers perceptual control stack (in Behavior: The Control of Perception, 1973). The identity of a cognitive node is a function of its parameter values, where the parameters are derived from the control stack. The identity of the arcs is a function of the difference in parameter values between the nodes it connects.

Concerning Chomsky’s approach to syntax: It depends on a sharp distinction between grammatical and ungrammatical sentences. A generative grammar, in Chomsky’s theory, must account for all and only the grammatical sentences.

However, there are no explicit criteria for separating sentences into the two categories, grammatical and ungrammatical. Rather, the separation depends on the intuitions of the linguist. Naturally enough, different syntacticians have different intuitions. The problem is insoluble.

Moreover, anyone who pays close attention to real speech soon realizes that people do not (always) speak in complete grammatically correct sentences. Real language is sloppy, but nonetheless effective. A neural net of very high dimensionality can deal with this readily enough. A purely symbolic system cannot. Augmenting the system through fuzzy logic and the like doesn’t fix the problem.

LLMs provide a very useful simulacrum of the natural language system. Since LLMs are trained on written texts, the resulting model necessarily conflates the functions of semantics and syntax/discourse. Thus they cannot achieve the flexibility and precision of the full system, where semantics and syntax/discourse are separated.

Neuro-symbolic computing sneaks in through the "back door" [a transitional stage in the evolution of AI]

Wednesday, March 11, 2026

Maybe the AI industry is waking up, at long last

From the tweet:

This tells you everything about how the smart money is actually modeling AI’s future. They’re not pricing AMI on a revenue multiple. They’re pricing it on the probability that LLMs hit a ceiling. And if you look at the investor list, Nvidia, Samsung, Toyota Ventures, Dassault, Sea, these are companies that need AI to understand physics, geometry, and force dynamics. 

Saturday, February 21, 2026

The Molecular Structure of Thought

Qiguang Chen et al., The Molecular Structure of Thought: Mapping the Topology of Long Chain-of-Thought Reasoning, arXiv:2601.06002v2 [cs.CL], https://doi.org/10.48550/arXiv.2601.06002

Abstract: Large language models (LLMs) often fail to learn effective long chain-of-thought (Long CoT) reasoning from human or non-Long-CoT LLMs imitation. To understand this, we propose that effective and learnable Long CoT trajectories feature stable molecular-like structures in unified view, which are formed by three interaction types: Deep-Reasoning (covalent-like), Self-Reflection (hydrogen-bond-like), and Self-Exploration (van der Waals-like). Analysis of distilled trajectories reveals these structures emerge from Long CoT fine-tuning, not keyword imitation. We introduce Effective Semantic Isomers and show that only bonds promoting fast entropy convergence support stable Long CoT learning, while structural competition impairs training. Drawing on these findings, we present Mole-Syn, a distribution-transfer-graph method that guides synthesis of effective Long CoT structures, boosting performance and RL stability across benchmarks.

Tuesday, February 3, 2026

Agentic coding is neurosymbolic AI

Tuesday, March 18, 2025

Into the digital wilderness and beyond LLMs

Alexander Doria has a recent post, The Model is the Product, that’s gotten me to think about future developments in a different way:

There were a lot of speculation over the past years about what the next cycle of AI development could be. Agents? Reasoners? Actual multimodality?

I think it’s time to call it: the model is the product.

All current factors in research and market development push in this direction.

  • Generalist scaling is stalling. This was the whole message behind the release of GPT-4.5: capacities are growing linearly while compute costs are on a geometric curve. Even with all the efficiency gains in training and infrastructure of the past two years, OpenAI can’t deploy this giant model with a remotely affordable pricing.
  • Opinionated training is working much better than expected. The combination of reinforcement learning and reasoning means that models are suddenly learning tasks. It’s not machine learning, it’s not base model either, it’s a secret third thing. It’s even tiny models getting suddenly scary good at math. It’s coding model no longer just generating code but managing an entire code base by themselves. It’s Claude playing Pokemon with very poor contextual information and no dedicated training.
  • Inference cost are in free fall. The recent optimizations from DeepSeek means that all the available GPUs could cover a demand of 10k tokens per day from a frontier model for… the entire earth population. There is nowhere this level of demand. The economics of selling tokens does not work anymore for model providers: they have to move higher up in the value chain.

This is also an uncomfortable direction. All investors have been betting on the application layer. In the next stage of AI evolution, the application layer is likely to be the first to be automated and disrupted.

He then goes on to explain why he’s saying that. He then says:

...the current training ecosystem is very tiny. You can count all theses companies on your hands: Prime Intellect, Moondream, Arcee, Nous, Pleias, Jina, the HuggingFace pretraining team (actually tiny)… Along with a few more academic actors (Allen AI, Eleuther…) they build and support most of the current open infrastructure for training. In Europe, I know that at least 7-8 LLM projects will integrate the Common Corpus and some of the pretraining tools we developed at Pleias — and the rest will be fineweb, and likely post-training instruction sets from Nous or Arcee.

There is something deeply wrong in the current funding environment. Even OpenAI senses it now. Lately, there was some felt irritation at the lack of “vertical RL” in the current Silicon Valley startup landscape. I believe the message comes straight from Sam Altman and will likely result in some adjustment in the next YC batch but pinpoint to a larger shift: soon the big labs select partners won’t be API customers but associated contractors involved in the earlier training stage.

If the model is the product, you cannot necessarily build it alone. Search and code are easy low hanging fruits: major use cases for two years, the market is nearly mature and you can ship a new cursor in a few months. Now many of the most lucrative AI uses cases in the future are not at this advanced stage of development — typically, think about all these rule based system that still rule most of the world economy… Small dedicated teams with a cross-expertise and a high level of focus may be best positioned to tackle this—eventually becoming potential acquihire once the initial ground work is done. We could see the same pipeline in the UI side. Some preferred partner, getting exclusive API access to close specialized models, provided they get on the road for business acquisition.

Note I have been arguing that LLMs are a digital wilderness. In post from December 29, 2022, not long after ChatGPT was released, I quoted from an article by Ted Underwood:

In his penultimate paragraph Underwood notes:

I have suggested that approaching neural models as models of culture rather than intelligence or individual language use gives us even more reason to worry. But it also gives us more reason to hope. It is not entirely clear what we plan to gain by modeling intelligence, since we already have more than seven billion intelligences on the planet. By contrast, it’s easy to see how exploring spaces of possibility implied by the human past could support a more reflective and more adventurous approach to our future. I can imagine a world where generative models of culture are used grotesquely or locked down as IP for Netflix. But I can also imagine a world where fan communities use them to remix plot tropes and gender norms, making “mass culture” a more self-conscious, various, and participatory phenomenon than the twentieth century usually allowed it to become.

These digital wildness regions thus represent opportunities for discovery and elaboration. Alignment is simply one aspect of that process.

And by alignment I mean more than aligning the AI’s values with human values; I mean aligning its conceptual structure as well. That’s where “old school” symbolic computing enters the picture, especially language. Language – not the mere word forms available in digital corpora, but word forms plus semantics and syntactic affordances – is one of the chief ‘tools’ through which young humans are acculturated and through which human communities maintain their beliefs and practices. The full powers of language, as treated by classical symbolic systems, will be essential for “domesticating” the digital wilderness and developing it for human use.

What do you do with a wildness? You explore it, map it, enclose parts of it, and then develop them. That’s a job for “small dedicated teams with a cross-expertise and a high level of focus.”

I returned to the wilderness them in the report I recently posted on the year and-a-half I spent investigating ChatGPT: ChatGPT: Exploring the Digital Wilderness, Findings and Prospects. There I said:

To return to the metaphor with which I began this report, these LLMs, these so-called Foundation Models, they are a digital wildness. Wild and untamed, but rich in resources. Ongoing esearch in mechantistic interpretatility is one way to explore that wildnerness, to map it. I have been presenting a different, a complementary mode of exploration and mapping. We need to map the territory, settle it, and domesticate it. By that I mean develop symbolic models to operate on and in the territory. While we may start developing those models through hand-coding, we will have to develop programmatic techniques if we want to cover the territory, which, of course, will be ever expanding.

As far as I can tell, Doria is imagining that his small focused teams will be working within existing architectural and programmatic frameworks or obvious extensions of them. I’m imagining something a bit different, something that requires us to understand the internal operations of LLMs. Yet, I have little trouble imagining that here and there one or more of his small focused teams will begin making sense of those internal operations beyond that displayed in current interpretability research. Perhaps they will lay the foundations for the kind of programmatic techniques I am imagining.

H/t Tyler Cowen.

Friday, January 3, 2025

What Miriam Yevick Saw: The Nature of Intelligence and the Prospects for A.I., A Dialog with Claude 3.5 Sonnet

New working paper posted. Title above, links, summary (abstract), TOC, and introduction below.

Claude’s Summary

1. Miriam Yevick's 1975 work proposed a fundamental distinction between two types of computing: A) Holographic/parallel processing suited for pattern recognition, and B) Sequential/symbolic processing for logical operations. Crucially, she explicitly connected these computational approaches to different types of objects in the world.

2. This connects to the current debate about neural vs. symbolic AI approaches: A) The dominant view suggests brain-inspired (neural) approaches are sufficient. B) Others argue for neuro-symbolic approaches combining both paradigms, and C) Systems like AlphaZero demonstrate the value of hybrid approaches (Monte Carlo tree search for game space exploration, neural networks for position evaluation).

3. The 1988 Benzon/Hays paper formalized “Yevick's Law” showing: A) Simple objects are best represented symbolically, B) Complex objects are best represented holographically, and C) Many real-world problems require both approaches.

4. This framework helps explain: A) Why Chain of Thought prompting works well for math/programming (neural system navigating symbolic space), B) Why AlphaZero works well for chess (symbolic system navigating game tree, neural evaluation of positions), and C) The complementary relationship between these approaches.

5. These insights led to a new definition of intelligence: “The capacity to assign computational capacity to propositional (symbolic) and/or holographic (neural) processes as the nature of the problem requires.”

6. This definition has implications for super-intelligence: A) Current LLMs' breadth of knowledge doesn't constitute true super-intelligence, B) Real super-intelligence would require superior ability to switch between computational modes, and C) This suggests the need for fundamental architectural innovations, not just scaling up existing approaches.

The conversation highlighted how Yevick's prescient insights about the relationship between object types and computational approaches remain highly relevant to current AI development and our understanding of intelligence itself.

CONTENTS

Introduction: Speculative Engineering 2

Yevick’s Idea 2
Why Hasn’t Yevick’s Work Been Cited? 4
From Claude to Speculative Engineering 5
What’s in the Dialog 6
Attribution 7

From Holographic Logic 1975 to Natural Intelligence 1988 8

Yevick’s Holographic Logic 8
Natural Intelligence 13
Chain of Thought 16
Chain of Thought Is the Inverse of the Alphazero Architecture 17
Define Intelligence 18
Implications of My Proposed Definition of Intelligence 20
Summary 22

Claude Reviews Yevick’s Full 1975 Text 25
Generalizing Yevick’s Results 26
Explaining Yevick to Non-Experts 28
Uzkeki Peasants Have Trouble With Simple Geometric Objects 31
Why Yevick’s Work Has Been Neglected 33

Introduction: Speculative Engineering

In the course of the dialog with Anthropic’s Claude 3.5 about artificial intelligence I propose a novel definition:

Intelligence is the capacity to assign computational capacity to propositional (symbolic) and/or holographic (neural) processes as the nature of the problem requires.

Compare this with a version of the conventional definition:

Intelligence is the ability to learn and perform a range of strategies and techniques to solve problems and achieve goals.

Those definitions are very different. My proposal suggests specific mechanisms while the conventional definition does not. While my proposal must be considered speculative, it can be used to guide research (that’s what speculation is for). In fact it can guide two research programs: 1) a scientific program about the mechanisms of human intelligence, and 2) an engineering program about the construction of A.I. systems. The conventional definition is probably correct, in some sense. But it provides very little guidance for any research. Its use is primarily rhetorical, for use in general or philosophical discussions of artificial intelligence.

Which is preferable, a definition that is speculative and useful in guiding research, or one that is (weakly) correct but of little value in guiding research?

This is the issue that hovers over the dialog I have constructed with Claude.

I will get around to explaining why I’ve consulted Claude soon enough. But first I want to talk about where that definition of intelligence came from. The words are mine, but the basic idea is not.

Yevick’s idea

The idea belongs to a mathematician, the late Miriam Yevick. Consider this remark that Yevick published in 1978:

If we consider that both of these modes of identification enter into our mental processes, we might speculate that there is a constant movement (a shifting across boundaries) from one mode to the other: the compacting into one unit of the description of a scene, event, and so forth that has become familiar to us, and the analysis of such into its parts by description. Mastery, skill and holistic grasp of some aspect of the world are attained when this object becomes identifiable as one whole complex unit; new rational knowledge is derived when the arbitrary complex object apprehended is analytically described.

She thought of one of those computational modes as holographic and the other as logical or sequential. She regarded both as essential to human mentation. Yevick analyzed those two modes in mathematical detail in a 1975 article I used to I begin my discourse with Claude (p. 8 below). David Hays and I employed those modes in the 1988 article that I quote from later in the dialog (p. 13) and in a more informal article (about metaphor) that we published at roughly the same time. Though they use different terms, Bengio, LeCun, and Hinton recognize the same distinction in their Turing Award paper.

But Yevick’s distinction entails something that Bengio, LeCun, and Hinton do not talk about. She regards the holographic mode as best suited for dealing with one kind of object, one that is geometrically complex, and the logical mode as best suited for a different kind of object, one that is geometrically simple. Now, in the passage I’ve quoted above Yevick has generalized from the visual mode, which she argued in her 1975 paper, to mental processes in general. Hays and I do so in our 1988 article as well. As far as I know, that mathematics has yet to be done – which is one reason I am trying to bring attention to Yevick’s work.

If Yevick is correct, then the current debate we are having about how to proceed is ill-posed. As far as I can tell, the majority view is that scaling up machine learning procedures will prove sufficient to achieving human-scale intelligence (not to mention super-intelligence). Gary Marcus, Subbarao Kambhampati, and various others have been arguing that, no, we also need symbolic (logical) mechanisms. Yevick’s work comes down on this side of the debate as well, and contributes to it by specifying that the two modes are each suited to a different aspect of that world. That would naturally lead to a discussion about the nature of the world, but that is beyond the scope of this document.

In view of the importance of this debate, it does not seem unreasonable. It is not only that billions of dollars are currently being wagered–and that is what it is, gambling, but that they are being wagered on technology that may well bring about fundamental changes in the way we live. Can we not take a step back and think about what we are doing?

Thursday, March 7, 2024

AI, Chess, and Language 1: Two VERY Different Beasts

Chess has been with AI since the beginning. In fact, computer chess all but beats the origins of the term “artificial intelligence,” which was coined for the well-known 1956 Dartmouth summer research program. According to this timeline in Wikipedia the possibility of a mechanical chess device dates back to the 18th century with a chess-playing automaton which, however, was operated by a human concealed inside it. A number of other mechanical hacks appeared before Norbert Wiener, Claude Shannon, and Alan Turing theorized about computational chess engines. John McCarthy invented the alpha-beta search algorithm in 1956, the same year as the AI conference, and the first programs to play a full game emerged a year later, in 1957. That was also the year that Chomsky published Syntactic Structures and the Russians launched Sputnik (which I was able to observe from my backyard).

Meanwhile, in 1949 Roberto Busa got IBM to sponsor a project to create a computer-generated concordance to the works of Thomas Acquinas, the Index Thomisticus. Thus the so-called digital humanities were born. That same year Warren Weaver, who was head of the Rockefeller Foundation at the time, wrote a memorandum in which he proposed a statistical rationale for machine translation. In 1952 Yehoshua Bar-Hillel, an Israeli logician, organized the first conference in machine translation at MIT’s Research Laboratory for Electronics and two years later IBM demonstrated the automatic translation of bits of Russian text into English (Nilsson 2010, p. 148).

The chess gang theorized that, as chess exemplified the highest form of human intelligence, when AI had succeeded in beating the best humans at chess, full artificial intelligence would have been achieved. In 1997 IBM’s Deep Blue beat Garry Kasparov decisively. Ever since then computers have been the best chess players in the world. But computer performance on language tasks has lagged far behind chess performance. The recent development of transformer-based large language models (LLMs) has resulted in a quantum leap in linguistic performance for computers, but the writing, though fluent, is also pedestrian (with various exceptions we need not go into). It would seem that there is a profound difference between the computational requirements of chess and those of language.The difference is not simply a matter of raw compute, with language requiring more, much more. There is also a fundamental, and perhaps even irreducible, difference in the way that compute is orchestrated computationally.

This post offers some quick observations about that difference. I discuss chess first, then language. While I’m a fluent speaker and writer of English, and know a bit about computational linguistics as well, I don’t play chess (though I do know the rules) and I know little about computer chess. I had a brief session with ChatGPT about chess. I’ve included that as an appendix.

Chess as a computational problem

While chess is a fairly cerebral game, it is a game played in the physical world using a board and game pieces. Let us call that chess’s geometric footprint. When we get to language we’ll talk about language’s geometric footprint, which is quite different from chess’s. The rules of chess can be defined with respect to its geometric footprint:

  • There are two players, who alternate moves.
  • It is played on an 8 by 8 board.
  • Each player has 16 pieces distributed over 6 types.
  • The moves of each type are rigidly and unambiguously specified.
  • There are other rules regarding games play among those pieces on the board.

The only constraint the players are subject to is the constraint that they obey the rules of the game. Any move that is consistent with the rules is permitted.

Given that the number of squares on the board is finite, the number of pieces is finite, that each play involves finite movement on the board, and that a convention is adopted to terminate play in the case no pieces are being exchanged, the total number of possible chess games is finite and takes the form of a tree.

To play chess at even a quite modest level one must have a repertoire of tactics and strategies that are not specified by the rules. Chess is a game where there is a game where there is a strict distinction between the basic rules and what we might call the elaboration, that is, the tactics and strategies governing games play. As a practical matter, large ill-defined areas of the chess tree are unexplored because, once a player moves into any of those areas, they will lose to a superior opponent. In particular, the opening of a game is quite restricted, not by the rules, but by well-known tactical considerations.

Given that the game is defined in terms of its geometric footprint, and that all possible chess games can be enumerated in the form of a tree, it follows that a person’s ability to play chess depends on how well they know the chess tree. However it is that they represent this tree in their minds is, at this point, a secondary issue. Given two players, if one of them knows a larger and more interesting (however one specifies interesting, a difficult problem) region of the chess board than the other, they will consistently win over the other.

It follows therefore, that since computers now consistently beat even the best of human players, they have explored regions of the chess tree that no human player has.

Language as a computational problem

Now let us consider human language. It has a geometric footprint as well, which is given in language’s relationship to the natural world. The meaning of a good many words is given directly by physical phenomena; much of so-called common-sense knowledge is like this. Unlike the geometric footprint of chess, which is small, simple, and well-defined, the geometric footprint of language is large, complex, and poorly defined. I would note further that words that are not given their primary meaning in physical terms as given in the human sensorium can be given meaning by various means, including patterns of words and patterns which include formal symbols as well, symbols from mathematics, chemistry, physics, and so forth. On this last point, see various posts tagged metalingual definition, and two papers that I wrote with David Hays:

William Benzon and David Hays, Metaphor, Recognition, and Neural Process, The American Journal of Semiotics, Vol. 5, No. 1 (1987), 59-80, https://www.academia.edu/238608/Metaphor_Recognition_and_Neural_Process

William Benzon and David Hays, The Evolution of Cognition, Journal of Social and Biological Structures. 13(4): 297-320, 1990, https://www.academia.edu/243486/The_Evolution_of_Cognition

It is thus difficult to make a firm distinction between the basic rules of language and the elaboration. While the number of possible chess games is finite, though it is so large that we cannot list it. It makes little sense to talk of listing all possible language texts. The number is unbounded and the set is not enumerable.

The basic rules of chess are so simple that computers need not play chess by moving physical pieces around on a board; a purely symbolic notation is entirely adequate. The training of LLMs does not involve access to the physical world. But the geometric footprint of language so constrains semantic relationships that LLMs can induce a suitable approximation of those relationships given a sufficiently large training corpus. But we have no way of determining whether or not an LLM can generate any possible text. Nor do we have any reason to believe that LLMs can generate any text that can be generated by human having full access to the physical world. In fact, given the problems that LLMs have with common-sense knowledge, we have reason to believe that LLMs are more limited in their capacities than the collectivity of humans.

Computationally, chess and natural language are very different beasts. The former is tractable and considered solved while the latter is still somewhat mysterious. Chess was solved by the methods of classical symbolic computing, methods which we understand. We know how classical chess engines work. But the newer neural network chess engines are black boxes, as are LLMs. We don’t (yet) know how they work. Moreover, while it is clear the LLMs are more generally successful with language than the old symbolic systems were, it does seem that the older symbolic systems are better at some kinds of reasoning and planning tasks and it is not obvious that LLMs can acquire those capacities.

Two kinds of computing and an exercise for the reader

I have written various posts about the work of Miriam Yevick. See, for example: Toward a Theory of Intelligence: Did Miriam Yevick know something in 1975 that Bengio, LeCun, and Hinton did not know in 2018? In particular, Yevick proved that in the task of visual identification, symbolic systems (like old-style chess engines) are better than neural nets. Does her result generalize across the whole range of tasks we would require of our intelligence engines? If so, then it means that we cannot reach AGI (whatever that is) simply by making larger and larger LLMs.

Appendix: ChatGPT on computer chess

Is there any evidence that good computer chess programs sometimes play differently than the best humans?

Yes, there is evidence that computer chess programs, especially the strongest ones, can play differently than the best human players. Computer chess engines rely on brute-force calculations and sophisticated algorithms to evaluate positions and make decisions, while human players often rely on intuition, pattern recognition, and strategic understanding.

Tuesday, February 27, 2024

The Fastest Way to AGI: LLMs + Tree Search [Neuro-symbolic]

Demis Hassabis (Google DeepMind CEO)

About 0:23:

....we've got to carry on
um basically making the more and more
accurate predictors of the world so in
effect making them more more reliable
World models that's clearly a necessary
but I would say probably not sufficient
component of an AGI system um and then
on top of that I would you know we're
working on things like Alpha zero like
planning mechanisms on top that make use
of that model in order to make concrete
plans to achieve certain goals in the
world um and and perhaps sort of chain
you know chain thought together or lines
of reasoning together and maybe use
search to kind of explore massive spaces
of possibility I think that's kind of
missing from our current large models...

About 2:15

my betting would be is that um you know
the final AGI system will have these
large multimodels um models as part of
the the overall solution but probably U
won't be enough on their own you will
need this additional planning search on
top

Tuesday, November 14, 2023

Enhancing LM common-sense reasoning through neurosymbolic techniques

Abstract of paper linked in the GitHub site:

Yuling Gu and Bhavana Dalvi Mishra and Peter Clark, Do language models have coherent mental models of everyday things?, arXiv:2212.10029v3 [cs.CL].

When people think of everyday things like an egg, they typically have a mental image associated with it. This allows them to correctly judge, for example, that “the yolk surrounds the shell” is a false statement. Do language models similarly have a coherent picture of such everyday things? To investigate this, we propose a benchmark dataset consisting of 100 every- day things, their parts, and the relationships between these parts, expressed as 11,720 “X relation Y?” true/false questions. Using these questions as probes, we observe that state-of- the-art pretrained language models (LMs) like GPT-3 and Macaw have fragments of knowledge about these everyday things, but do not have fully coherent “parts mental models” (54- 59% accurate, 19-43% conditional constraint violation). We propose an extension where we add a constraint satisfaction layer on top of the LM’s raw predictions to apply common-sense constraints. As well as removing inconsistencies, we find that this also significantly improves accuracy (by 16-20%), suggesting how the incoherence of the LM’s pictures of everyday things can be significantly reduced.

Friday, November 25, 2022

Meta announces CICERO, an AI that plays Diplomacy

Gary Marcus and Ernest Davis have posted an interesting evaluation of Cicero: What does Meta AI’s Diplomacy-winning Cicero Mean for AI? [Hint: It’s not all about scaling]

First:

The first thing to realize is that Cicero is a very complex system. Its high-level structure is considerably more complex than systems like AlphaZero, which mastered Go and chess, or GPT-3 which focuses purely on sequences of words. Some of that complexity is immediately apparent in the flowchart; whereas a lot of recent models are something like data-in, action out, with some kind of unified system (say a Transformer) in between, Cicero is heavily prestructured, in advance of any learning or training, with a carefully-designed bespoke architecture that is divided into multiple modules and streams, each with their own specialization.

A marvel, but...

Cicero is in many ways a marvel; it has achieved by far the deepest and most extensive integration of language and action in a dynamic world of any AI system built to date. It has also succeeded in carrying out complex interactions with humans of a form not previously seen.

But it is also striking in how it does that. Strikingly, and in opposition to much of the Zeitgeist, Cicero relies quite heavily on hand-crafting, both in the data sets, and in the architecture; in this sense it is in many ways more reminiscent of classical “Good Old Fashioned AI” than deep learning systems that tend to be less structured, and less customized to particular problems. There is far more innateness here than we have typically seen in recent AI systems

Also, it is worth noting that some aspects of Cicero use a neurosymbolic approach to AI, such as the association of messages in language with symbolic representation of actions, the built-in (innate) understanding of dialogue structure, the nature of lying as a phenomenon that modifies the significance of utterances, and so forth.

That said, it’s less clear to us how generalizable the particulars of Cicero are.

In sum:

Cicero makes extensive use of machine learning, but is hardly a poster child for simply making ever bigger models (so-called “scaling maximalism”), nor for the currently popular view of “end-to-end” machine learning of in which some single general learning algorithm applies across the board, with little internal structure and zero innate knowledge. At execution time, Cicero consists of a complex array of separate hand-crafted modules with complex interactions. At training time, it draws on a wide range of training materials, some built by experts specifically for Cicero, some synthesized in programs hand-crafted by experts. [...]

Our final takeaway? We have known for some time that machine learning is valuable; but too often nowadays ML is a taken as universal solvent—as if the rest of AI was irrelevant—and left to do everything on its own. Cicero may change that calculus. If Cicero is any guide, machine learning may ultimately prove to be even more valuable if it is embedded in highly structured systems, with a fair amount of innate, sometimes neurosymbolic machinery.

There's much more in their article.

Sunday, August 28, 2022

Elemental Cognition is ready to deploy hybrid AI technology in practical systems

Steve Lohr, One Man's Dream of Fusing A.I. With Common Sense, NYTimes, Aug. 28, 2022

David Ferrucci is best-known as the researcher who led the team that developed IBM's Watson, which beat the best human players of Jeopardy in 2011. He left IBM a year later and formed his own company, Elemental Cognition, in 2015. Elemental cognition is taking a hybrid approach, combining aspects of machine learning and symbolic computation.

Elemental Cognition has recently developed a system that helps people plan and book round-the-world airline tickets:

The round-the-world ticket is a project for oneworld, an alliance of 13 airlines including American Airlines, British Airways, Qantas, Cathay Pacific and Japan Airlines. Its round-the-world tickets can have up to 16 different flights with stops of varying lengths over the course of a year.

Elemental Cognition supplies the technology behind a trip-planning intelligent agent on oneworld’s website. It was developed over the past year and introduced in April.

The user sees a global route map on the left and a chatbot dialogue begins on the right. A traveler starting from New York types in the desired locations — say, London, Rome and Tokyo. “OK,” replies the chatbot, “I have added London, Rome and Tokyo to the itinerary.”

Then, the customer wants to make changes — “add Paris before London,” and “replace Rome with Berlin.” That goes smoothly, too, before the system moves on to travel times and lengths of stays in each city.

Rob Gurney, chief executive of oneworld, is a former Qantas and British Airways executive familiar with the challenges of online travel planning and booking. Most chatbots are rigid systems that often repeat canned answers or make irrelevant suggestions, a frustrating “spiral of misery.”

Instead, Mr. Gurney said, the Elemental Cognition technology delivers a problem-solving dialogue on the fly. The rates of completing an itinerary online are three to four times higher than without the company’s software.

Elemental Cognition has developed an approach that all-but eliminates the hand-coding typcial of symbolic A.I.:

For example, the rules and options for a global airline ticket are spelled out in many pages of documents, which are scanned.

Dr. Ferrucci and his team use machine learning algorithms to convert them into suggested statements in a form a computer can interpret. Those statements can be facts, concepts, rules or relationships: Qantas is an airline, for example. When a person says “go to” a city, that means add a flight to that city. If a traveler adds four more destinations, that adds a certain amount to the cost of the ticket.

In training the round-the-world ticket assistant, an airline expert reviews the computer-generated statements, as a final check. The process eliminates most of the need for hand coding knowledge into a computer, a crippling handicap of the old expert systems.

There's more at the link.

* * * * *

Lex Fridman interviews David Ferrucci (2019).

0:00 - Introduction
1:06 - Biological vs computer systems
8:03 - What is intelligence?
31:49 - Knowledge frameworks
52:02 - IBM Watson winning Jeopardy
1:24:21 - Watson vs human difference in approach
1:27:52 - Q&A vs dialogue
1:35:22 - Humor
1:41:33 - Good test of intelligence
1:46:36 - AlphaZero, AlphaStar accomplishments
1:51:29 - Explainability, induction, deduction in medical diagnosis
1:59:34 - Grand challenges
2:04:03 - Consciousness
2:08:26 - Timeline for AGI
2:13:55 - Embodied AI
2:17:07 - Love and companionship
2:18:06 - Concerns about AI
2:21:56 - Discussion with AGI

Sunday, August 7, 2022

AGI as shibboleth, symbols [reacting to Jack Clark]

Jack Clark has a LONG tweet stream on AI policy. Though I don’t agree with every tweet – would anyone? – it’s worth at least a quick look. I want to comment on two of the tweets.

AGI as shibboleth, and beyond

Has AGI ever been anything other than a shibboleth? I believe the term was coined in the 1990s because some researchers felt that AI had become stale and focused on specialized domains, so-called “narrow” AI. The phrase “artificial general intelligence” (AGI) was as a banner under which to revive the founding goal of AI, to construct the artificial equivalent of human intelligence.

What researchers actually construct are mechanisms. But no one knows how to specify a mechanism or set of mechanisms for AGI. Oh, sure, there’s the Universal Turing machine which can, in point of abstract theory, compute any computable function. It may be a mechanism, but the idea so abstract that it provides little to no guidance in the construction of computer systems.

AGI, like AI before it, is an abstract goal, a beacon, without a procedure that will lead to it. No matter how vigorously you chase over the surface of the earth for the North Star, you’re never going to get there. And so AGI simply functions as a shibboleth. If you want into the club, you have to pledge allegiance to AGI.

But you don’t need to pledge allegiance in order to construct interesting and even useful systems. So why invent this unreachable goal? Is it just to define a club?

Meanwhile I’ve written a paper in which I define the idea of an artificial mind. I begin by defining mind:

A MIND is a relational network of logic gates over the attractor landscape of a partitioned neural network. A partitioned network is one loosely divided into regions where the interaction within a region is (much) stronger than the interactions between regions. Each of these regions will have many basins of attraction. The relational network specifies relations between basins in different regions.

Note that the definition takes the form of specifying a mechanism involving logic gates and a neural network. Given that:

A NATURAL MIND is one where the substrate is the nervous system of a living animal.

And:

An ARTIFICIAL MIND is one where the substrate is inanimate matter engineered by humans to be a mind.

There are other definitions as well as some caveats and qualifications.

However, those definitions come after 50 pages of text and diagrams in which I lay out the mechanisms that support those definitions. The paper is primarily about the human brain, but one can imagine constructing artificial devices that meet those specifications. Now, whether those specifications are the right specifications, that’s open for discussion. However that discussion turns out, it is a discussion about mechanisms, not myths and magic.

The paper:

Relational Nets Over Attractors, A Primer: Part 1, Design for a Mind, Version 2, Working Paper, July 13, 2022, pp. 76, https://www.academia.edu/81911617/Relational_Nets_Over_Attractors_A_Primer_Part_1_Design_for_a_Mind

Ah, symbols

Here’s a twofer:

It's the first tweet that interests me, but let’s look Richard Sutton’s bitter lesson. Here’s his opening paragraph:

The biggest lesson that can be read from 70 years of AI research is that general methods that leverage computation are ultimately the most effective, and by a large margin. The ultimate reason for this is Moore's law, or rather its generalization of continued exponentially falling cost per unit of computation. Most AI research has been conducted as if the computation available to the agent were constant (in which case leveraging human knowledge would be one of the only ways to improve performance) but, over a slightly longer time than a typical research project, massively more computation inevitably becomes available. Seeking an improvement that makes a difference in the shorter term, researchers seek to leverage their human knowledge of the domain, but the only thing that matters in the long run is the leveraging of computation. These two need not run counter to each other, but in practice they tend to. Time spent on one is time not spent on the other. There are psychological commitments to investment in one approach or the other. And the human-knowledge approach tends to complicate methods in ways that make them less suited to taking advantage of general methods leveraging computation. There were many examples of AI researchers' belated learning of this bitter lesson, and it is instructive to review some of the most prominent.

Sutton then goes on to list domains where there has proven so: chess, Go, speech recognition, and computer vision. He then draws some conclusions, which I want to bracket.

Note, however, that Sutton talks of researchers seeking “to leverage their human knowledge of the domain.” Is that what’s going on symbolic AI? Perhaps in expert systems, which may have been the most pervasive practical result of GOFAI. But I don’t think that’s an accurate general characterization. That’s not what was going on in computational linguistics, for example, or in much of the work on knowledge representation. That research was based on the belief that much of human knowledge is inherently symbolic in character and therefore that we must create models that capture that symbolic character.

Why did those models collapse? I think there are several factors involved:

1. Combinatorial explosion: Symbolic systems tend to generate large numbers of alternative with little or no way of choosing among them.

2. Hand coding: Symbolic systems have to be painstakingly hand-coded, which takes time.

3. Too many models, difficult to choose among them: This exacerbates the hand-coding problem.

4. Common sense has proven elusive: But then it has proven elusive for deep learning as well.

Perhaps the first problem can be solved through more computing power, though exponential search can easily outstrip the addition of CPU cycles and memory. The third problem is one for science, and is, I believe, entangled with the fourth one. The second problem is inconvenient, but, alas, if hand-coding is necessary, then it’s necessary. But perhaps if we’re clever....

On the fourth one, here’s what I said in my GPT-3 paper:

A lot of common-sense reasoning takes place “close” to the physical world. I have come to believe, but will not here argue, that much of our basic (‘common sense’) knowledge of the physical world is grounded in analogue and quasi-analogue representations. This gives us the power to generate language about such matters on the fly. Old school symbolic machines did not have this capacity nor do current statistical models, such as GPT-3.

Thus the problem is not specific to symbolic systems. It is quite general. It’s not at all clear that we can deal with this problem without having robots out and about in the world. I note that the working paper I mentioned in the previous section, Relational Nets Over Attractors, is about constructing symbolic structures over quasi-analog representations, which, following the terminology of Saty Chary, I characterize as structured physical systems.

Let’s return to Sutton’s paper. Here’s his final paragraph:

The second general point to be learned from the bitter lesson is that the actual contents of minds are tremendously, irredeemably complex; we should stop trying to find simple ways to think about the contents of minds, such as simple ways to think about space, objects, multiple agents, or symmetries. All these are part of the arbitrary, intrinsically-complex, outside world. They are not what should be built in, as their complexity is endless; instead we should build in only the meta-methods that can find and capture this arbitrary complexity. Essential to these methods is that they can find good approximations, but the search for them should be by our methods, not by us. We want AI agents that can discover like we can, not which contain what we have discovered. Building in our discoveries only makes it harder to see how the discovering process can be done.

I’m hesitant to think of symbol systems as being “simple ways to think about the contents of minds.” That strikes me as rhetorical overkill. But Sutton is right about “the arbitrary, intrinsically-complex, outside world.” He says that “we should build in only the meta-methods that can find and capture this arbitrary complexity.” Well, sure, why not?

But are we doing that now? That’s not at all obvious to me. it seems likely to me that the DL community is hoping that they’ve discovered the metamethods, or will do so in the near future, and so we don’t have to think about what’s going on inside either human minds or the machines we’re building. Well, if human minds use symbols, and it seems all but self-evident that we do – if language isn’t a symbol system, what is? – then the current repertoire of DL methods is not up to the task.

What meta-methods are needed to detect patterns of symbolic meaning and construct those quasi-analog representations?

My GPT-3 paper:

GPT-3: Waterloo or Rubicon? Here be Dragons, Version 4.1, Working Paper, May 7, 2022, 38 pp., https://www.academia.edu/43787279/GPT_3_Waterloo_or_Rubicon_Here_be_Dragons_Version_4_1

Monday, July 11, 2022

My main response to LeCun's position paper: Why are symbols important? Because they index cognitive space.

Yann LeCun recently posted a major position paper that has been receiving quite a bit of discussion:

Yann LeCun, A Path Towards Autonomous Machine Intelligence, Version 0.9.2, 2022-06-27, https://openreview.net/forum?id=BZ5a1r-kVsf

I posted the following remarks to the LeCun discussion on 3 July 2022.

* * * * *

I want to address the issue that your raise at the very end of your paper: Do We Need Symbols for Reasoning? I think we do. Why? 1) Symbols form an index over cognitive space that, 2) facilitates flexible (aka ‘random’) access to that space during complex reasoning.

Let me quote a passage from the paper you recently published with Jacob Browning:

For the empiricist tradition, symbols and symbolic reasoning is a useful invention for communication purposes, which arose from general learning abilities and our complex social world. This treats the internal calculations and inner monologue — the symbolic stuff happening in our heads — as derived from the external practices of mathematics and language use.[1]

I agree with the second sentence. Symbols are not primitive to the nervous system, they are derived. Initially, from linguistic communication, but then, as culture evolves, from mathematics as well.

The first sentence is true, but not entirely adequate for understanding language (where I consider arithmetic, for example, to be a very specialized form of language). Back in the 1930s the Russian psychologist, Lev Vygotsky, gave an account of language acquisition that moves through three phases: 1) adults (and others) use language to direct the very young child’s attention and actions, 2) gradually the child learns to use speech to direct their own attention and action, and finally 3) the process becomes completely internalized, e.g. inner monologues. I spell this out in more detail in a wide-ranging working paper I’ve recently posted to the web [2].

Now, what is the nature of cognitive space? That’s a complicated question, but much of it is defined directly over physical objects, events, and processes and that is, I believe, differentiable in the way you desire. Here I believe the geometric semantics developed by Peter Gärdenfors [3] may prove useful in seeing how cognition is linked to symbols and I utilize it in my working paper.

Still, let me mention one complication. Here’s an example that was much discussed in the Old Symbolic Days: What’s a chair? Chairs are obviously physical objects, but when you consider the range of objects that are recruited to serve as chairs, it becomes difficult to imagine a single physical description that characterizes all of them, even a fairly abstract description. Perhaps chairs are best characterized by their function, that is, by the role they play in a simple action. The concept of “poison” presents a similar problem. There’s no doubt that poisons are physical substances, but they don’t have a common physical appearance. Nor, for that matter, do fruits and vegetables. Fruits and vegetables play certain roles in cuisines and poisons are most generally characterized known by their effects. And so forth. It’s a complicated problem, but a secondary one at the moment.

Will your proposed H-JEPA architecture support such symbols? I find the following passage suggestive [p. 7]:

The world model may predict natural evolutions of the world, or may predict future world states resulting from a sequence of actions proposed by the actor module. The world model may predict multiple plausible world states, parameterized by latent variables that represent the uncertainty about the world state. The world model is a kind of “simulator” of the relevant aspects of world. What aspects of the world state is relevant depends on the task at hand.

That sounds a bit like natural language parsers, where partial parses will be developed and maintained until enough information is obtained to decide on one of them. I tentatively conclude that, yes, your architecture can accommodate symbols, though you will have to deal with the discrete nature of the symbols themselves.

I really should say something about how symbols facilitate flexible access to cognition, but well, that’s tricky. Let me offer up a fake example that points in the direction I’m thinking. Imagine that you’ve arrived at a local maximum in your progression toward some goal but you’ve not yet reached the goal. How do you get unstuck? The problem is, of course, well known and extensively studied. Imagine that your local maximum has a name1, and that name1 is close to name2 of some other location in the space you are searching. That other location may or may not get you closer to the goal; you won’t know until you try. But it is easy to get to name2 and then see where where that puts you in the search space. If you’re not better off, well, go back to name1 and try name3. And so forth. Symbol space indexes cognitive space and provides you with an ordering over cognitive space that is different from and somewhat independent of the gradients within cognitive space. It’s another way to move around. More than that, however, it provides you with ways of constructing abstract concepts, and that’s a vast, but poorly studied subject [4].

[1] Yann LeCun and Jacob Browning, What AI Can Tell Us About Intelligence, Noema, June 16, 2022, https://www.noemamag.com/what-ai-can-tell-us-about-intelligence/

[2] William Benzon, Relational Nets Over Attractors, A Primer: Part 1, Design for a Mind, June 20, 2022, https://ssrn.com/abstract=4141479

[3] Peter Gärdenfors, Conceptual Spaces, MIT 2000, The Geometry of Meaning, MIT 2014. For a quick introduction see Peter Gärdenfors, An Epigenetic Approach to Semantic Categories, IEEE Transactions on Cognitive and Developmental Systems (Volume: 12, Issue: 2, June 2020) 139 – 147. DOI: 10.1109/TCDS.2018.2833387

[4] For some thoughts on various mechanisms for constructing abstract concepts, see William Benzon and David Hays, The Evolution of Cognition, Journal of Social and Biological Structures, 13(4): 297-320, 1990, https://doi.org/10.1016/0140-1750(90)90490-W

Monday, May 30, 2022

Eureka! Have I Found It? How to Model the Mind, that Is. [Symbols and Nets]

Since roughly the last week in April, when I applied for an Emergent Ventures grant (which was quickly, but politely, turned down), I have been working hard on revising and updating work on a system of notation which I sketched out in 2003 and posted to the web in 2010, 2011. I am referring to what I then called called an Attractor Network, but now call a Relational Network over Attractors (RNA) because I found out that neuroscientists already talk about attractor networks, which are not the same as what I’ve got in mind. The neuroscientists are referring to a network of neurons whose dynamics tend toward an attractor. I am referring to a network that specifies relationships between a very large number of attractors (hence, it is constructed over them).

Anyhow, by the time Emergent Ventures had turned me down, I was committed to the project, which has gone well so far. I had no particular expectations, just a general direction. I’ve been looking, and I’ve found some interesting things, encouraging things. Or, if you will, I’ve been puttering around, assembling bits and pieces here and there, and an interesting structure has begun to emerge.

Lamb Notation

The idea has been to develop a new notation for representing semantic structures in network form. Actually, the notation is not new; it had already been developed by Sydney Lamb in the 1960s. He developed it to model the structures of a stratificational grammer. I’ve been adapting it to model semantics.

I am doing that by assuming that the cerebral cortex is loosely divided into functionally distinct regions which I call neurofunctional areas (NFAs). The activity of these NFAs is to be modeled by complex dynamics (Walter Freeman) and a low-dimensional projection of each NFA phase space can be modeled by a conceptual space (Peter Gärdenfors). Each NFA is thus characterized by an attractor landscape.

The RNA (relational net over attractors) is a network where the nodes are logical operators (AND, OR) and the edges are basins of attraction in the NFA attractor landscapes. This is not the place to explain what that actually means, but I can give you a taste by showing you three pictures.

This is a simple semantic structure expressed in a “classical” notation from the 1970s:

It depicts the fact that both beagles and collies are varieties (VAR) of dog. The light gray nodes at the bottom are perceptual schemas, while the dark gray nodes at the right are lexemes. The white nodes are cognitive.

Here’s a fragment of one of Lamb’s networks:

The triangular nodes are AND while the brackets (both pointing up and down) are OR. The content is carried on the edges.

This RNA network takes the information expressed in the semantic network and expresses it using AND and OR nodes.

I am not even going to attempt to explain just how that works. Suffice it to say that it seems a bit more visually complicated than the old notation and thus harder to read. It also expresses more informatation. Those AND and OR nodes specify processing while no processing is specified in the classical diagram.

I am finding it more demanding to work with. In part that is because I haven’t drawn nearly so many RNA diagrams, perhaps 100 or so as compared to 1000s. But also, in drawing RNAs I have to imagine these structures being somehow laid out on a sheet of cortex, which is tricky. It would be even trickier if I were working with data about the regional functional anatomy of the cortex at my elbow, trying to figure just where each NFA is on the cortical sheet. Eventually, that will have to be done, but right now I’m satisfied just to draw some diagrams.

Crazy and Not So Crazy

The fact that I intend these diagrams as a very abstract sketch of functional cortical anatomy means that they have fairly direct empirical implications that the old diagrams never had. Of course, we were always committed to the view that we were figuring out how the human mind worked and so  eventually someone would have to figure out where and how those structures were implemented in the brain. Well, now is eventually and these new diagrams are a tool for figuring out the where and how.

And that, I suppose, is a crazy assertion. Everyone who knows anything knows that the brain is fiercely complicated and we’re never going to figure it out in a million years but anyhow we have to a waste a billion euros building a damned brain model that tells us a bit more than diddly squat, but not a whole hell of a lot more. But then what I’m doing costs nothing more than my time. Excuse the rant.

As I said, it’s crazy of me to propose a way of thinking about how high-level cognitive processes are organized in the brain. But I’m only proposing, and I’m doing it by offering a conceptual tool, a notation, that helps us think about the problem in a new way. I don’t expect that the constructions I propose are correct. I ask only that they are coherent enough to lead us to better ones.

There’s one further thing and this is not so crazy: This notation, in conjunction with 1) my assertation that it is about complex cortical dynamics, and 2) and Lev Vygotsky’s account of language development, gives us a new way of thinking about a debate that is currently blazing away in a small region of the internet: How do we model the mind, neural vectors, symbols, or both? If both, how? I am opting for both and making a fairly specific proposal about how the human brain does it. The question then becomes: What will it take to craft an artificial device that does it? If my proposal ends up taking 14K or 15K words and maybe 30 diagrams, well it deals with a very a complicated problem.

Here is the draft introduction, Symbols, holograms, and diagrams, to the working paper. With that, I’ll leave you with a brief sketch of my proposal.

The Model in 14 Propositions

1. I assume that the cortex is organized into NeuroFunctional Areas (NFAs), each of which has its own characteristic pattern of inputs and outputs. As far as I can tell, these NFAs are not sharply distinct from one another. The boundaries can be revised – think of cerebral plasticity.

2. I assume that the operations of each NFA are those of complex dynamics. I have been influenced by Walter Freeman in this.)

3. A low dimensional projection of each NFA phase space can be modeled by a conceptual space as outlined by Peter Gärdenfors.

4. Each NFA has its own attractor landscape. A primary NFA is one driven primarily by subcortical inputs. Then we have secondary and tertiary NFAs, which involve a mixture of cortical and subcortical inputs. (I’m thinking of the standard notions of primary, secondary, and tertiary cortex.)

5. Interaction between NFAs is defined by a Relational Network over Attractors (RNA), which is a relational network defined over basins in multiple linked attractor landscapes.

6. The RNA network employs a notation developed by Sydney Lamb in which the nodes are logical operators, AND & OR, while ‘content’ of the network is carried on the arcs. [REF/LINK to his paper.]

7. Each arc corresponds to a basin of attraction in some attractor landscape.

8. The output of a source NFA is ‘governed’ by an OR relationship (actually exclusive OR, XOR) over its basins. Only one basin can be active at a time. [Provision needs to be made for the situation in which no basin is entered.]

9. Inputs to a basin in a target NFA are regulated by an AND relationship over outputs from source NFAs.

10. Symbolic computation arises with the advent of language. It adds new primary attractor landscapes (phonetics & phonology, morphology?) and extends the existing RNA. Thus overall RNA is roughly divided into a general network and a lingistic network.

11. Word forms (signifiers) exist as basins in the linguistic network. A word form whose meaning is given by physical phenomena are coupled with an attractor basin (signifier) in the general network. This linkage yields a symbol (or sign). Word forms are said to index the general RNA.

12. Not all word forms are defined in that way. Some are defined by cognitive metaphor (Lakoff and Johnson). Others are defined by metalingual definition (David Hays). I assume there are other forms of definition as well (see e.g. Benzon and Hays 1990). It is not clear to me how we are to handle these forms.

13. Words can be said to index the general RNA (Benzon & Hays 1988).

14. The common-sense concept of thinking refers to the process by which one uses indices to move through the general RNA to 1) add new attractors to some landscape, and 2) construct new patterns over attractors, new or existing.

Monday, May 16, 2022

Is GPT-3 a structuralist? A brave new world in which language exists beyond the human?

Tobias Rees, Non-Human Words: On GPT-3 as a Philosophical Laboratory, Dædalus, Spring 2020.

In [Saussure's] words, “language is a system of signs that expresses ideas.”

Put differently, language is a freestanding arbitrary system organized by an inner combinatorial logic. If one wishes to understand this system, one must discover the structure of its logic. De Saussure, effectively, separated language from the human. There is much to be said about the history of structuralism post de Saussure.

However, for my purposes here, it is perhaps sufficient to highlight that every thinker that came after the Swiss linguist, from Jakobson (who developed Saussure’s original ideas into a consistent research program) to Claude Lévi-Strauss (who moved Jakobson’s method outside of linguistics and into cultural anthropology) to Michel Foucault (who developed a quasi-structuralist understanding of history that does not ground in an intentional subject), ultimately has built on the two key insights already provided by de Saussure: 1) the possibility to understand language, culture, or history as a structure organized by a combinatorial logics that 2) can be–must be–understood independent of the human subject.

GPT-3, wittingly or not, is an heir to structuralism. Both in terms of the concept of language that structuralism produced and in terms of the antisubject philosophy that it gave rise to. GPT-3 is a machine learning (ML) system that assigns arbitrary numerical values to words and then, after analyzing large amounts of texts, calculates the likelihood that one particular word will follow another. This analysis is done by a neural network, each layer of which analyzes a different aspect of the samples it was provided with: meanings of words, relations of words, sentence structures, and so on. It can be used for translation from one language to another, for predicting what words are likely to come next in a series, and for writing coherent text all by itself.

GPT-3, then, is arguably a structural analysis of and a structuralist production of language. It stands in direct continuity with the work of de Saussure: language comes into view here as a logical system to which the speaker is merely incidental.

That view has some similarity with the one I advanced on Pages 15-19 of my 2020 working paper about GPT-3.

Moreover,

All prior structuralists were at home in the human sciences and analyzed what they themselves considered human-specific phenomena: language, culture, history, thought. They may have embraced cybernetics, they may have conducted a formal, computer-based analysis of speech or art or kinship systems. And yet their focus was on things human, not on machines. GPT-3, in short, extends structuralism beyond the human.

The second, in some ways even more far-reaching, difference is that the structuralism that informs LLMs like GPT-3 is not a theoretical analysis of something. Quite to the contrary, it is a practical way of building things. If up until the early 2010s the term structuralism referred to a way of analyzing, of decoding, of relating to language, then now it refers to the actual practice of building machines “that have words.”

A new ontology?

Machine learning engineers in companies like OpenAI, Google, Facebook, or Microsoft have experimentally established a concept of language at the center of which does not need to be the human, either as a knowing thing or as an existential subject. According to this new concept, language is a system organized by an internal combinatorial logic that is independent from whomever speaks (human or machine). Indeed, they have shown, in however rudimentary a way, that if a machine discovers this combinatorial logic, it can produce and participate in language (have words). By doing so, they have effectively undermined and rendered untenable the idea that only humans have language–or words.

What is more, they have undermined the key logical assumptions that organized the modern Western experience and understanding of reality: the idea that humans have what animals and machines do not have, language and logos. [...]

In fact, the new concept of language–the structuralist concept of language–that they make practically available makes possible a whole new ontology.

What is this new ontology? Here is a rough, tentative sketch, based on my current understanding.

By undoing the formerly exclusive link between language and humans, GPT-3 created the condition of the possibility of elaborating a much more general concept of language: as long as language needed human subjects, only humans could have language. But once language is understood as a communication system, then there is in principle nothing that separates human language from the language of animals or microbes or machines.

A brave new world?

Language, almost certainly, is just a first field of application, a first radical transformation of the human provoked by experimental structuralism. That is, we are likely to see the transformation of aspects previously thought of as exclusive human qualities–intelligence, thought, language, creativity–into general themes: into series of which humans are but one entry.

What will it mean to be surrounded by a multitude of non-human forms of intelligence? What is the alternative to building large-scale collaborations between philosophers and technologists that ground in engineering as well as an acute awareness of the philosophical stakes of building LLMs and other foundational models?

It is naive to think we can simply navigate–or regulate–the new world that surrounds us with the help of the old concepts. And it is equally naive to assume engineers can do a good job at building the new epoch without making these philosophical questions part of the building itself: for articulating new concepts is not a theoretical but a practical challenge; it is at stake in the experiments happening in (the West at) places like OpenAI, Google, Microsoft, Facebook, and Amazon.

Addendum, 5.17-19.22:  We're not quite there yet, but Rees' final point still stands, we do need new concepts.  It's not at all clear in what sense GPT-3 is "arguably a structural analysis of" language, or any kind of analysis at all, and it certainly is not a stand-alone language automaton. It does not in fact constitute a/the language system divorced from a human agent. There is only a partial separation, a distancing. We're heading in that direction, but I doubt that we'll get there on extensions of current tech alone. We're going to need something new. Neurosymbolic? Maybe. 

For a different kind of analysis of GPT, but also deeper because it gets closer to the mechanism, see my working paper, GPT-3: Waterloo or Rubicon? Here be Dragons, Version 4.1.