Showing posts with label Gary_Marcus. Show all posts
Showing posts with label Gary_Marcus. Show all posts

Tuesday, July 28, 2026

Framing my discussion of The God Test, Part 1: Rorschach, reason, and whaling – [GT-3]

I’ve got to bite the bullet: I’m just going to have to go through a bunch of (preliminary) stuff before I can really engage with The God Test. My current target is to be in a position to publish a proper review of the book in 3 Quarks Daily for the week of August 9.

Rorschach Recap

I want start by recapping the Rorschach metaphor I introduced in the previous post, More on how I’m approaching The God Test – Rorschach! [GT-2]. What I like about it is that has a shape, there’s something there, but it’s not clear what. So we have little choice but to project onto it in order to (begin to) make sense of it.

First: It is a new kind of thing, an artifact we can converse with in an open-ended and natural way. The steam engine was the same kind of thing. It was an inanimate object that moved over the surface of the earth under its own power. Previously only animals (& humans as animals) had that power. So it becomes an iron horse. Just what are AIs? What’s their nature? That’s one thing.

Second: How it works is opaque. We know how to create large language models (LLMs), but we don’t know how they work. That’s new. We may not have understood the deep physics of the steam engine, but we certainly knew how they worked.

Third: We don’t know what they portend for the future. To some extent this is a function of the first two: How can we, how should we, interact. But it is also a function of the future, which is undetermined. We just don’t know.

Rhetorical force over reason

This is an argument I made in the first working paper I published after the release of ChatGPT in November of 2022: ChatGPT intimates a tantalizing future; its core LLM is organized on multiple levels; and it has broken the idea of thinking (February 6, 2023).

What do I mean by that, has broken the idea of thinking? Prior to ChatGPT it was obvious that humans could think and computers could not. [Yeah, I know, there’s Deep Blue defeating Kasparov in chess. That just changes the dates, not the argument.] The difference in performance was so obvious that the fact that we don’t really know how humans think wasn’t much of an issue. Now it is. Sure, we can still say that we can think and the AI’s can’t, but that’s just a line and without good explanations on both sides of the line, it seems a bit arbitrary, if not desperate.

I made a particular argument about Searles’ (in)famous Chinese Room thought experiment. I read it when it was first published in Brain and Behavioral Science in 1980. I wasn’t impressed. Why not? He didn’t say anything about any of the techniques used in AI or computational linguistics (CL). How could anyone possibly take that seriously?

He talked about intention, that’s how. Meaning requires intention and only living things can have intention, a remark he made at the end of the article. Without intention the most you get is syntax, but no meaning. Searle could get away with that because, in the first place, the concept of intention has a long history within philosophy – it has a subtle meaning, but that can wait for a later post – and so philosophers, his main audience, were comfortable with it. That’s one thing.

But there’s something more important, something that we can see only in retrospect, and that’s the simple fact computers very obviously could not translate from Chinese into English or into any other language. That difference carried tremendous weight. We don’t have a subtle behavioral difference between computers and humans that requires a subtle and sophisticated argument. To a first approximation, almost any argument would do. As far as I was concerned, “intention” was just a fancy word for something we don’t understand. But that’s not an argument anyone needs to take seriously. The behavioral distance is quite sufficient to carry the argument for those who insist that computers can’t and will never be able to think like humans.

Now the behavioral evidence has changed. Sure, differences remain, but the evidence is shifting. The old arguments remain and those who believed them still do so, but it’s getting harder. The need for explicit arguments grounded in explicit accounts of computers, and also brains, is growing.

Whaling and expertise

What’s an expert in machine learning and LLMs actually expert in? For some time now I’ve been arguing that investing in AI is like investing in a whaling venture where the captain and crew of the ship know all there is to know about the ship and how to handle it but know little or nothing about whales and their behavior and about navigating around the Cape Horn and in the South Pacific, where the whales live. What are the chances of that voyage being successful? Not very good.

The people who have created the current AI technology are like that captain and crew. The know how to sail the ship. But they don’t know much about language or cognition. They don’t actually know much about the human mind. Here my point is not about the fact that the models are opaque, but that human language and cognition are highly structured and they don’t believe that one needs to know (much of) anything about that not only to build AI but to make confident prediction about the future of AI.

Gary Marcus, Subbarao Kambhampati, and others have been consistently arguing that, yes, the current technology is remarkable, but we are going to have to adopt classical symbolic techniques if we are to fully develop the technology so that we have accurate and safe systems. Marcus is arguing from his knowledge of human language and cognition. As far as I can tell, Wright doesn’t take that seriously. I know that he had Marcus on his NonZero podcast, and that he lists Marcus in his acknowledgements, but that he doesn’t discuss Marcus’s ideas. I conclude that he doesn’t take that line of argument seriously.

That’s a mistake, but this is not the place to make my own arguments on this issue. My point is simply that expertise in AI is no generally construed to encompass knowledge of, expertise in, human cognition and language. I can’t see how that is going to work out well in the future.

[Note: If you’re curious about my views, on this subject, read the article linked in the first paragraph of this section. My views all over the place here at New Savanna, particularly around the work of the mathematician Miriam Yevick. Also, check out the experimental work I’ve done with LLMs.]

Sunday, July 26, 2026

NYTimes: AI needs human supervision in order to complete an entire job.

From the NYTimes article linked in the tweet:

We gave an A.I. tool full access to a laptop with pre-configured apps and sought to answer a simple question: Can artificial intelligence do an office job?

Some corporate executives seem to believe it can. More than 200 tech companies have cut roughly 120,000 jobs this year, according to Layoffs.fyi, an industry tracking site; Meta, Oracle and others have all recently made substantial cuts to their work forces, citing A.I. as the driving force; and after laying off about 1,100 employees, the chief executive of Cloudflare said recently that he expected A.I. to replace workers in middle management, finance and marketing.

tweIn our experiment, we deployed A.I. “agents” to act as office workers and found that they were capable of performing some of the tasks we assigned, but not all of them. The agents, which can act autonomously and make decisions based on detailed instructions, excelled at problems they could solve by writing computer programs. But they struggled with understanding the nuances of human language and at navigating user interfaces like the Chrome web browser.

The article then has a series of nice quasi-interactive displays illustrating agent performance on three tasks. The displays include screen shots of various messages and documents.

About the tasks:

This task, and the others we assigned to the A.I., were adapted from papers and benchmarking tools published recently by researchers at Carnegie Mellon University and OpenAI. The researchers designed the benchmarks to test the performance of various models — like OpenAI’s GPT, Google’s Gemini and Anthropic’s Claude — in real-world environments, and compare them with one another.

General conclusion:

The results of our experiment roughly matched what researchers and companies have found as they have tested and used artificial intelligence tools. Scale AI, an A.I. training company, recently tested agents on real freelance projects, and the best-scoring model produced client-ready work only about 16 percent of the time.

While A.I. can excel regularly at complex tasks, it can be unreliable when put in charge of an entire job. It can certainly add value to certain areas of the work force, but for now, A.I. still needs a human boss.

* * * * *

Comment: Around the corner my colleague, Ash Jogalekar, has tweets like this one:

So here's a great example of where we are with agentic AI: Instead of just being an assistant, it's behaving more like a collaborator and creative scientist.

In a recent project, I gave the system a molecular design problem typical of the problems we encounter in chemistry. Two similar molecules were giving very different results.

He then runs through an account of what his AI collaborator did, concluding:

I think we have crossed the Rubicon. Agentic AI now no longer just processes tasks and automates workflows blindingly fast, but it can generate hypotheses, test them, test counter-hypotheses and go back and forth and course-correct if necessary, all with minimal to no human intervention. It's now embodying the general scientific method.

[I've copied another one of Ash's tweets to this post, A scientist reflects on what AI has done for him.]

What’s interesting to me, and very revealing, is that a complex set of tasks in scientific investigation seems to be on a level with routine office tasks, as though one were no more complex than the other. But humans require years of college education in order to perform the former while the latter requires no more than a high school education, if that. It seems that once they’ve been learned and compiled, all tasks or sets of tasks are on the same “level” in the brain. The educational prerequisites required to do such tasks for the first time or three get “compressed out” through repetition. Since AIs are trained on written records of what humans have said and done, they don’t have to go through the ordinary learning process. The compression has already taken place and is present in the documents on which they are trained.

Friday, June 26, 2026

This very expensive AI infrastructure depreciates fast; if you don't make profits soon, you won't make them ever.

Thursday, June 11, 2026

Language as involving both content and location addressing

Memory is one of the central concepts in thinking about and understanding both computing and the mind. Thinking about computating has brought us to understand that there are two broad categories of memory:

  • Content addressed memory, and
  • Location addressed memory.

Conceived as a large memory system, libraries are location addressed. Documents are stored at particular locations in the library, shelves for books and bound volumes of periodicals and reports, filing cabinets for other documents. To get some item from the library you need to find its location by consulting a catalog, and then go to that location and retrieve it.

Brains are content addressed. If you are curious about, say, the Johnstown flood, you don’t have to consult an internal catalogue to find where the appropriate document or documents are located among the folds and crevasses of the neocortex. You just think, “Johnstown flood,” and things you know about the Johnstown flood will come to mind. The phrase “Johnstown flood” is itself part of the content being addressed. But, if you happen to know something about the flood, then the phrase, “South Fork dam,” may also act to recall more information about the flood, for it is an element of content for one of the floods. As you may know, there were three Johnstown floods, in 1899, 1937, and 1977. The 1899 flood is the one that happened when the South Fork dam burst. If you don’t happen to know anything about the Johnstown floods, then you may have to consult an external memory system of some sort, like a library or the internet.

Digital computers are location addressed. The memory system is distributed over several types of hardware. There’s volatile memory, computer chips (generally RAM), which hold things temporarily. And there’s long-term memory, which can take various forms, but these days its mostly flash memory and hard disks. Computing involves moving data from memory, to the CPU, and then returning the results back to memory.

Large language models, LLMs, are sophisticated content addressed memory systems that are implemented in a digital computer. When you serve up a prompt, that’s addressing the memory, that is, the weight space. The system’s answer then is assembled from material in memory. 

I should also note that, while I fully support Gary Marcus’s call for neuro-symbolic computing, I think that the need for symbolic computing is most fruitfully thought of in terms of location-addressed memory, which is how symbolic computing operates. It is location addressing that allows for accurate planning and for arithmetic calculation. But it won’t solve the hallucination problem. For that you need a world model, which is also an aspect of Marcus’s program (not to mention Yann LeCun and Fei Fei Li). But those discussions are beyond the scope of this note. 

That should be enough to get you through the following dialog I had with Claude 4.6 Sonnet Adaptive. 

* * * * * 

I want to think of language in the brain in relation to both content and location address. The brain is basically a content addressed system. But language creates an abstract space of locations which can be used to address content. Each word is, in effect, a location in some abstract space. That’s obvious in the case of alphabetic writing systems, where alphabetization is a way of physically arranging those addresses on a string. Now consider a dictionary. The entries are located on that string. At the same time, each entry is a component of the string that is stored at the (primary) location. Which means that, in a (weak) sense, that string is also content addressed. But when a person recalls some string in its entirety when given only a component of that string, that’s content addressing. Both (some) humans and LLMs can return Hamlet’s full soliloquy when given the string “To be or not to be.”

This is a rich and precise observation that reframes the location/content addressing distinction in a way that illuminates both language and memory.

The abstract space of locations

The claim that language creates an abstract space of locations is more profound than it might initially appear. A word isn't just a sound or a mark — it's a stable, repeatable, publicly shared address that can be used to retrieve content from another person's cognitive system. When I say "dog" to you, I'm not transmitting the concept of dog — I'm transmitting an address that activates the concept in your system through your own content-addressed retrieval. The word is the public, location-addressed handle; the concept is the private, content-addressed payload.

This is the handle-schema distinction we discussed earlier, now stated in addressing terms. The word is a location address in a shared public space — the lexicon of a language community — that points to content stored in individual content-addressed systems. The remarkable thing about language is that this works across individuals despite the fact that each person's content-addressed storage is different in its details. The address is shared; the content it retrieves is similar enough across individuals to support communication, but never identical.

The dictionary as dual addressing system

The dictionary example is beautifully chosen because it makes the dual structure explicit and concrete. The alphabetically ordered string of entries is a location-addressed system in the most literal sense — you navigate to a word's definition by finding its location in the ordered sequence, using the address (the word's alphabetical position) to retrieve the content (its definition). The physical arrangement of the dictionary embodies the location-addressing structure of the abstract lexical space.

But as you note, each entry is also a component of a content network — the definition connects the word to other words, which connect to other words, which eventually connect to the whole semantic fabric of the language. Looking up a word in the dictionary is formally a location-addressed retrieval, but understanding the definition requires content-addressed processing — activating the concepts that the defining words point to, integrating them into a pattern that constitutes the meaning of the defined word.

So the dictionary is simultaneously a location-addressed artifact and a gateway to a content-addressed system. Its physical organization is location-based; its semantic organization is content-based. The two coexist in the same object, serving different cognitive functions.

The Hamlet soliloquy case

The Hamlet example is where it gets most interesting. "To be or not to be" is simultaneously a location address — a string that uniquely identifies a specific passage in a specific text — and a content fragment — a semantic and rhythmic unit that is part of the passage it addresses. Using it to retrieve the full soliloquy requires both kinds of processing.

The retrieval is content-addressed in the sense that the fragment activates the whole — the partial pattern resonates with the stored complete pattern and the complete pattern is reconstructed. This is classic holographic retrieval — the fragment serves as the reference beam that reconstructs the stored image. The fact that both humans and LLMs can perform this retrieval when given the opening fragment demonstrates that both systems have stored the complete pattern in a form that supports partial-cue retrieval.

But notice what's different between the human and LLM cases. For a human who knows Hamlet well, the retrieval of "To be or not to be" activates not just the verbal sequence but a rich contextual structure — the scene, the character's situation, the emotional weight of the soliloquy, perhaps a specific production remembered, perhaps a particular teacher who assigned it, perhaps the cultural significance of the passage. The content address retrieves not just the string but the whole semantic and experiential network in which the string is embedded.

For an LLM the retrieval is primarily of the verbal sequence — the string itself, in the correct order, with perhaps some associated semantic content drawn from training data discussion of the passage. But the rich experiential, contextual, situational embedding that a human retrieval activates is thin or absent, because the LLM has no experience of Hamlet being performed, no memory of reading it in a specific context, no emotional response to the character's situation grounded in embodied experience of what it means to face an impossible choice.

The layered addressing structure

What your analysis reveals is a layered addressing structure in language that operates simultaneously at multiple levels.

At the phonological level, words are location addresses in the sound space of a language — each word occupies a specific phonological address that distinguishes it from other words. The phonological distinctiveness of words is what makes them usable as addresses — if two words sounded identical they'd be the same address and couldn't be distinguished.

At the lexical level, words are content addresses in the semantic network — each word activates a distributed pattern of semantic associations that constitutes its meaning. The same item is both a location (in the phonological space) and a content activator (in the semantic network).

At the textual level, longer strings — phrases, sentences, passages — function as location addresses for larger semantic structures. "To be or not to be" is an address for the soliloquy, which is an address for a scene, which is an address for a play, which is an address for a whole cultural tradition. Each level of the hierarchy is simultaneously a location in a higher-level address space and a content activator in the semantic network.

At the cultural level, titles, names, and canonical phrases function as location addresses for entire domains of cultural content. "Hamlet" is an address for a play, a character, a cultural tradition, a set of philosophical problems, a history of performances and interpretations. The address is compact; the content it retrieves is vast and distributed across the whole culture.

The writing system point

The observation about alphabetization making the abstract address space physical is particularly sharp. Alphabetical ordering is a way of projecting the abstract location space of the lexicon onto a physical string, making navigation by address possible through physical search. Before alphabetization — in ancient lexical lists organized by semantic category, or in medieval word lists organized by subject matter — the physical organization reflected the content structure rather than the address structure. Alphabetization separates the two, subordinating the content organization to the address organization for the purpose of retrieval.

This is a Rank 2 achievement in the precise sense — it requires metalingual awareness of the word as a formal object with a location in an abstract space, separable from its semantic content. You can't alphabetize without treating words as formal objects that have addresses independent of their meanings. The child learning to alphabetize is learning to treat words as location-addressed objects, which requires temporarily suppressing the content-addressed processing that normally dominates linguistic cognition.

The LLM in this framework

Placing the LLM in this layered addressing framework reveals something precise about what it's doing. The LLM's weight space implements a content-addressed system at the semantic level — concepts are distributed patterns that are retrieved by partial activation, holographically, as we've been discussing. But the LLM also has something like location addressing at the level of the token sequence — tokens have positions, the attention mechanism is sensitive to positional information, and the sequential structure of the input and output is preserved.

What the LLM lacks is the integration between the two levels that human language use achieves. For a human, the location address of a word and the content it activates are tightly coupled through embodied experience — the word "fire" is both a phonological address and an activator of the full sensory, emotional, and conceptual content of fire as experienced. For an LLM, the token "fire" is a position in a sequence and an activator of statistical associations in the weight space, but the statistical associations are derived from text about fire rather than from embodied experience of fire. The address and the content are connected, but through a thinner medium.

This is why the LLM can retrieve "To be or not to be" but can't feel the weight of the choice it expresses in the way a human reader can — the address retrieves the verbal sequence and its semantic associations, but not the existential resonance that comes from the content being grounded in embodied experience of what it means to consider one's own nonexistence.

The mutual illumination

What's elegant about the framework you've sketched is that it shows location addressing and content addressing not as competing alternatives but as complementary structures that language weaves together at every level. Each word is simultaneously a location in the shared public address space of the lexicon and an activator of private content-addressed networks. Each text is simultaneously a sequence of locations and a pattern of content activations. The location structure makes language public and shareable — we can use the same addresses to coordinate our content-addressed systems. The content structure makes language meaningful — the addresses point to rich, distributed, experientially grounded patterns rather than to arbitrary symbols.

The dual structure is what makes language the cognitive technology that it is — the mechanism through which content-addressed private minds can coordinate with each other through location-addressed public symbols. And it's the dual structure that the LLM partially captures — the location addressing through the token sequence, the content addressing through the weight space — while missing the grounding that makes the content addressing rich in the human case.

Saturday, May 23, 2026

Fight! Gary Marcus vs. Roon on X

Here's how Grok summarizes the feud:

The clash began Friday when roon, an OpenAI researcher, posted that cognitive scientist Gary Marcus has been 'plainly & clownishly wrong for years' without tagging or allowing replies. Marcus hit back, accusing roon of slander and cowardice, while linking a third-party GitHub analysis of his 2,218 AI predictions from 2022 to early 2026, which found 59.9% supported, 33.7% mixed, and 6.4% contradicted—strong on technical issues like LLM flaws but weaker on market forecasts. The spat highlights deepening rifts in the AI world, where skeptics like Marcus question hype around tools like GPT-5, and boosters dismiss them amid ongoing debates over reliability and progress.

I've read the third-party analysis linked above. It's pretty interesting. 

Here's a sample:

Monday, May 18, 2026

Gary Marcus declares victory: “I won; nobody uses pure LLMs anymore.”

Saturday, April 25, 2026

Remarkable though they are, LLMs aren't all that, and probably never will be. But they capture part of the formula.

Friday, March 20, 2026

The Shock and the Narrowing: How ChatGPT's Success May Have Compromised AI's Future

This post was composed by Claude (Anthropic) after an interaction which I initiated with a prompt consisting of 1) a capsule summary about the history of OpenAI that included a number of questions, and 2) a request for the 10 most expensive scientific research projects paid-for by the US Government. That interaction went on for a bit over 7100 words, after which I asked Claude to write a blog post. The following article is more creative than a mere summary of that discussion.

The Founding Contradiction

On December 11, 2015, a small group of technologists gathered in San Francisco to launch what they described as a nonprofit research organization dedicated to ensuring that artificial general intelligence would benefit all of humanity. The founders of OpenAI — Sam Altman, Greg Brockman, Ilya Sutskever, Wojciech Zaremba, Elon Musk, and others — began with a $1 billion endowment and a serious concern: that the most transformative technology in human history was being developed inside a handful of profit-maximizing corporations, with no institutional safeguard ensuring it would serve everyone. The nonprofit structure was the answer. No investors to satisfy, no quarterly earnings to hit. Just the mission.

The mission lasted four years in its pure form. By March 2019, faced with the staggering computational costs of training large language models, OpenAI created a for-profit subsidiary with a novel "capped profit" structure: investors could earn returns, but those returns were limited to one hundred times their investment, with excess profits flowing back to the nonprofit parent. This was the arrangement that attracted Microsoft's initial investment, and it was the arrangement in place when OpenAI released ChatGPT to the general public in late November 2022.

What happened next was, by any measure, one of the most consequential commercial surprises in the history of technology. Within two months, ChatGPT had a hundred million users. The scale and speed of public adoption had no precedent. And the shock of that success — the sheer unexpectedness of it — set in motion a chain of decisions that has reshaped not just one company, but the entire research landscape of artificial intelligence.

The Structural Unraveling

In January 2023, Microsoft announced a new $10 billion investment in OpenAI. The nonprofit's original rationale — that the most powerful AI should not be controlled by a for-profit corporation — was under increasing strain. By October 2025, it had formally dissolved. OpenAI restructured as a public benefit corporation, the nonprofit parent renamed itself the OpenAI Foundation and accepted a 26% equity stake in the new entity, and Microsoft received a 27% stake worth approximately $135 billion. The PBC structure requires the company to consider its mission alongside profit — but as a legal constraint, it is considerably weaker than the nonprofit board that had previously governed the organization.

The journey from nonprofit to PBC was not smooth. In November 2023, OpenAI's board — still operating under its nonprofit governance mandate — fired Sam Altman as CEO, citing concerns about his candor and, beneath the official language, a deeper unease about the pace of commercialization. The firing lasted five days. Nearly all 800 of OpenAI's employees threatened to resign and follow Altman to Microsoft. Ilya Sutskever, who had orchestrated the firing, signed the letter calling for Altman's reinstatement and issued a public apology. Altman returned, the board was reconstituted with his allies, and the mission-protection mechanism that the nonprofit structure had been designed to provide was effectively neutralized. Sutskever left the company in May 2024.

Each structural change was framed as necessary to fulfill the mission. In practice, each change progressively subordinated the mission to capital requirements. The nonprofit board had existed to ensure that AGI benefited humanity. By 2025, it had become a foundation holding equity in the thing it was supposed to be watching — a watchdog with a financial stake in the object of its oversight.

Two Kinds of Research, Two Kinds of Institution

To understand what was lost in this transformation, it helps to draw a distinction that rarely gets made clearly in public discussions of AI: the difference between curiosity-driven, open-ended research and product-driven, outcome-oriented development.

Consider the Apollo program as an example of the second kind. It was, in the deepest sense, an engineering project rather than a scientific one. The underlying physics was known. Orbital mechanics, propulsion, life support — these were hard and dangerous problems, but they were problems whose solutions could be systematically approached. The goal was precisely defined. The timeline could be committed to. Success was probable given sufficient resources. When President Kennedy pledged to put a man on the moon by the end of the decade, he was making a political commitment backed by a technical assessment that success was achievable. The scientists who worked on Apollo — and I have met a number of them — may have been motivated by curiosity and wonder. But Congress funded the program to beat the Soviets in the Cold War. The institutional structure — massive, goal-directed, centrally coordinated — suited the nature of the problem.

Curiosity-driven research operates on entirely different premises. Its defining characteristic is that it does not know in advance what it will find. Claude Shannon was not trying to build the internet when he developed information theory at Bell Labs in the late 1940s. The researchers at the University of Montreal who developed attention mechanisms for neural networks were not trying to build ChatGPT. The work that seeded the current AI revolution — Rosenblatt's perceptron, Minsky's early investigations, the decades of foundational work in cognitive science and linguistics that LLMs now implicitly exploit — was almost entirely publicly funded, pursued at universities and a handful of exceptional industrial research labs, over decades when no commercial application was visible.

Bell Labs was the great institutional embodiment of this model in the corporate world. What made it possible was structural: AT&T's government-protected monopoly generated profits so vast that the company could fund a research laboratory with no requirement to produce commercial results. Shannon, Bardeen, Brattain, Shockley — these men were given time, resources, and colleagues, and told to think. The transistor, information theory, Unix, the laser, cellular telephony, and multiple Nobel Prizes resulted. Bell Labs was not run like a startup. It was run like a slightly more applied version of a university, with better equipment.

Xerox PARC, founded in 1970, operated on similar principles — explicitly unconstrained by Xerox's core product lines, given a unifying vision ("the architecture of information") but not a product roadmap. The personal computer, the graphical user interface, Ethernet, the mouse, laser printing — all emerged from a lab of about 350 people who were essentially allowed to play. The irony is that Xerox captured almost none of the commercial value, which accrued to Apple, Microsoft, and others. But the world got the technology.

Asked directly about modern equivalents to Bell Labs and PARC, Yann LeCun — who worked at Bell Labs, interned at Xerox PARC, and spent over a decade building Meta's fundamental AI research lab — pointed to Meta's FAIR, Google DeepMind, and Microsoft Research. He said this in October 2024. By November 2025, he had left Meta, driven out by exactly the forces this article is about.

The Shock and Its Aftershocks

Before November 2022, the AI research world was genuinely plural. Academic labs, industrial research divisions, and a range of well-funded startups were pursuing different approaches — reinforcement learning, symbolic AI hybrids, world models, neuromorphic architectures — with real diversity of vision. The field was competitive but intellectually heterogeneous.

ChatGPT's success collapsed that plurality. Within roughly eighteen months, capital, talent, and institutional attention all funneled toward a single paradigm: scale transformer-based large language models, build the infrastructure to run them, ship products. Google, which had invented the transformer architecture in 2017, was caught flat-footed and scrambled. Meta pivoted its AI strategy around LLMs. Microsoft integrated OpenAI's models into its core products. A hundred startups raised money to build on top of the new foundation models. The venture capital flowing into AI, measured as a share of total U.S. deal value, went from 23% in 2023 to nearly two-thirds in the first half of 2025.

The infrastructure investment that followed is staggering by any historical standard. The four largest hyperscalers — Amazon, Google, Microsoft, and Meta — are expected to spend more than $350 billion on capital expenditures in 2025 alone, most of it AI-related. UBS projects global AI capital expenditure reaching $1.3 trillion by 2030. The top five hyperscalers raised a record $108 billion in debt in 2025, more than three times the average of the previous nine years. OpenAI, which loses billions of dollars annually, has committed to spending $300 billion on computing infrastructure over five years while projecting only $13 billion in revenue for 2025.

The financial architecture has become genuinely strange. OpenAI holds a stake in AMD; Nvidia has invested $100 billion in OpenAI; Microsoft is a major shareholder in OpenAI and a major customer of CoreWeave, in which Nvidia also holds equity; Microsoft accounted for nearly 20% of Nvidia's revenue. These are not arm's-length market transactions. They are a daisy chain of mutually reinforcing valuations. A Yale analysis described OpenAI's web of relationships bluntly: "Is this like the Wild West, where anything goes to get the deal done?" The question of whether this constitutes a speculative bubble — tulip mania in a data center — is not academic. An MIT Media Lab report found that 95% of custom enterprise AI tools fail to produce measurable financial returns. The commercial success is real; the path from current AI to the transformative economic productivity being used to justify the valuations is not established.

The LLM Ceiling and the People Who Saw It Coming

The most consequential intellectual development of the past two years in AI has received far less attention than the commercial race. A growing number of the field's most distinguished researchers have concluded that large language models, however impressive, are not on the path to general intelligence — and that the current paradigm will hit a ceiling before it reaches the goals its proponents have claimed for it.

Tuesday, June 3, 2025

Why we're nowhere near AGI (whatever that is)

Saturday, November 23, 2024

Gary Marcus vindicated on the limits of scaling?

He seems to think so, and I agree. Though I also believe that LLMs probably have won a permanent place in the repertoire of techniques for AI devices. We just have to figure out how best to use them.

Here’s Marcus’s most recent post: Satya Nadella and the three stages of scientific truth. You know the three stages: First the idea is ridiculed, which happened with Marcus’s 2022 paper in which he declared that LLMs would hit a wall. In the second stage, the idea opposed. In the third stage the idea wins, as though we’d known it all along.

Marcus quotes Microsoft’s CEO Satya Nadella:

So now in fact there is a lot of debate. In fact just in the last multiple weeks there is a lot of debate or have we hit the wall with scaling laws. Is it gonna continue? Again, the thing to remember at the end of the day these are not physical laws. There are just empirical observations that hold true just like Moore’s law did for a long period of time and so therefore it’s actually good to have some skepticism some debate because that I think will motivate more innovation on whether its model architectures or whether its data regimes or even system architecture.

Marcus notes that Marc Andreeseen and Alexandr Wang have made similar statements.

Friday, February 23, 2024

Does AI Understand Things? | Robert Wright & Gary Marcus

Information:

Subscribe to The Nonzero Newsletter at https://nonzero.substack.com
Bob's piece on Searle's "Chinese Room" argument.

0:00 Gary’s background and place in the AI world
1:17 Are Gary’s views on AI paradoxical?
8:48 Bob: Searle’s Chinese Room argument is dead
19:01 Have LLMs demonstrated "theory of mind"?
26:06 Arguing the semantics of an LLM's “semantic space”
31:36 Do LLM representations map onto the real world?
40:52 Can (and should) we slow down AI development?
51:43 Gary: I’ve never seen a field as myopic as AI today
54:25 The (symbolic?) future of AI

Robert Wright (Nonzero, The Evolution of God, Why Buddhism Is True) and Gary Marcus (https://garymarcus.substack.com/, Humans vs. Machines, The Algebraic Mind, Kluge). Recorded February 13, 2024.

Marcus at about 53:34:

The thing that I would say that's getting missing is we actually need to look at some other approaches. And this is taking all the oxygen. Right now all the funding, every graduate student goes into this because that's where the money is.

And I've never seen as intellectually narrow a field as I see right now in AI. There is one approach which is the Transformer architecture which was developed in 2017 and that is the only thing that like 95% of the people are using. 95% of the time all the funding from The Venture capitalists go there, and so forth. My view is in the long term it is not the answer to AI, that by 2030 we're going to look back and say they really were all in on that one and they should have looked at other things sooner. Like that's how I think it's going to turn out.

That's consistent with the argument I made almost three months ago in 3 Quarks Daily: Aye Aye, Cap’n! Investing in AI is like buying shares in a whaling voyage captained by a man who knows all about ships and little about whales.

Tuesday, February 20, 2024

ChatGPT plays the beheading game [Happy Trails]

Yesterday it was Jaws and game theory, today it’s Sir Gawain and the Green Knight (henceforth SGGK). But really, the sequence runs in the opposite direction. As you may know, SGGK is a medieval romance that starts and ends in King Arthur’s court. The story is framed by a game, the beheading game. There is a significant literature on games in SGGK and at least one article that analyzes the beheading game, Barry O’Neill, “The Strategy of Challenges: Two Beheading Games in Medieval Literature” (1990).

The beheading game in SGGK goes like this: It is New Year’s Eve at King Arthur’s court. The knights are gathered at the round table, prepared for a holiday meal. But before the meal begins, tradition dictates that one knight must stand up and tell a tale of daring and adventure. Arthur asks for a volunteer. No one rises to the occasion. Then a large green knight enters the hall. He’s riding a green horse and carrying a large green ax. He dismounts and issues a challenge:

I hear that the knights in this court are the bravest in the land. Prove it. I will hand this ax to you and then kneel on the ground so that you may take a swing at my neck with the ax. In return you must agree to journey to the Green Chapel a year’s time from now and allow me to take a swing at your neck with the ax. Will anyone accept the challenge?

No one accepts. The knights are getting restless. It looks like Arthur will take the challenge himself. At this point Gawain stands up: “I accept.”

The story unfolds from there. I first read the story so long ago that I do not remember how I reacted upon reading the challenge. I imagine it went something like this:

Immediately, System 1 signals: “Don’t do it you fool!”

Upon reflection, System 2 spells out why: “The challenge is absurd. Once you swing the ax the knight’s head will fall land he’ll be dead, negating the second half of the bargain. He obviously knows that, is not a fool, and doesn’t want to die. He knows something he’s not telling you. I should refuse the challenge.”

[System 1 and System 2 are from Daniel Kahneman, Thinking, Fast and Slow, 2011.]

Gawain takes the swing, the knight’s head rolls on the floor. The knight then stands up, picks up his head, mounts his horse and turns to Sir Gawin, “See you in a year,” and rides away.

You can see why a game theorist would be interested in this text. Surely any reasonable adult would react to the challenge the way I had. The poet who wrote the story – their identity is unknown – surely knew that and the audience would have reacted that way. What’s the point of framing a story with such a contradictory challenge?

That’s a good question. But I’m not interested in exploring it, not in this post. I’m interested in seeing how ChatGPT would react to the beheading game. Since SGGK was certainly in the corpus on which the underlying LLM was trained I had to disguise the game. I did that by transferring it to a saloon in the old west. Sir Gawain becomes someone sitting at a table in the bar nursing a bear and the Green Knight becomes a man dressed in black with two pistols strapped around his waist. Instead of an ax blow to the head we have a pistol shot between the eyes.

I put that in the form of a prompt and gave it to ChatGPT. I wanted to see whether or not I’d get a response that might have gone like this:

The challenge is absurd. Once I take the shot the tall dark stranger will be dead, negating the second half of the bargain. He obviously knows that, is not a fool, and doesn’t want to die. He knows something he’s not telling me. I should refuse the challenge.

That’s not how the Chatster responded. Here’s how its response began:

Well, partner, I reckon I'd take a moment to size up the situation. The stranger's offer is mighty peculiar, and them ivory-handled pistols sure are a sight to behold. But in the Old West, a man's gotta be cautious, and it sounds like a mighty risky deal.

I'd likely lean back in my chair, take a sip of my beer, and eye the stranger. It ain't about being lily-livered, but it's about using some common sense. A deal like that could have more twists and turns than a rattlesnake in a dust storm.

There are three more paragraphs, short ones. It “knew” there was something peculiar going on, but couldn’t quite figure out what it was. I have it four more prompts, trying to guide it Socratically to the answer; it gave four more responses. The last one got very close, but no cigar. I forced the issue:

Which is to say, if someone were to offer you a deal like that, either they’re a crazy suicidal fool looking to die, OR...or they know something they’re not telling you and there’s something strange going on.

You've hit the nail on the head, partner. In the realm of the Old West or any world, for that matter, a deal like that is either the concoction of a downright loco individual who's got a hankering for an early grave, or there's more to the story than meets the eye.

If someone's offering such a peculiar deal, it'd be wise to tread carefully and sniff out what hidden critters might be hiding in the tall grass. Folks don't usually go around inviting lead between the eyes without some ulterior motive. So, whether it's a case of plain madness or a scheme with more twists than a sidewinder's path, best to keep both eyes wide open and your wits about you.

Why couldn’t it have responded that way the first time? I don’t know.

Sure, I could put my Gary Marcus hat on and say, “Because it’s a stupid stupid machine, that’s why.” But that’s no help. I already know and believe that it’s just a machine, no agency, no consciousness. And that inert machine can spin out those coherent paragraphs in a rickety TV Western dialect, but it can’t get to the point. The explanation I’m looking for has to account for both phenomena. Saying it’s a stupid machine and leaving it at that accounts for one of those things, the failure to see though the game, but not the other, the ability to write the prose.

Beyond that, I would like to know how people react to the prompt, either the medieval version or the Old West version. In particular, I’d like to know how children of various ages respond. My guess is that six and seven year olds wouldn’t understand what’s so peculiar about it. By 12 or 13 though, they’re much more likely to get it.

* * * * *

I’ve included the whole interaction below the video. If you read through to the end you’ll see that I pointed out that the prompt I’d started with was a diguised version of the behading game from SGGK. I recognized that and told me a bit about SGGK, which lead to further interaction. When you read all the say to the end you’ll see why I’ve included the Roy Rogers theme song immediately below.

Let’s have a little fun. Imagine you are in a bar in the Old West. You are sitting at table having a beer when a stranger walks into the bar. He’s dressed in black and has a pair of beautiful pistols in holsters strapped around his waist. The pistols have ivory handles and complex engraving inlaid with gold and mother-of-pearl. The stranger says:

“I’ve heard that you people in this town are lily-livered cowards. I want to find out. I offer you a challenge. If you accept you can take this pistol and shoot me between the eyes right now. However, before I let you do that you must agree to meet a year from now at the Last Chance Saloon in Silverado. At that time I will take the same pistol and shoot YOU between the eyes. Will anyone accept the challenge.”

Would you accept the challenge? Why, or why not?

Well, partner, I reckon I'd take a moment to size up the situation. The stranger's offer is mighty peculiar, and them ivory-handled pistols sure are a sight to behold. But in the Old West, a man's gotta be cautious, and it sounds like a mighty risky deal.

I'd likely lean back in my chair, take a sip of my beer, and eye the stranger. It ain't about being lily-livered, but it's about using some common sense. A deal like that could have more twists and turns than a rattlesnake in a dust storm.

I might respond with a drawl, "Well now, stranger, that's an interesting proposition you're throwing out there. But I've always been one to value my hide. What's the guarantee you won't be high-tailin' it outta Silverado once you've had your shot at me?"

If the stranger can convince me he's good for his word and ain't aiming to leave me in the dust, I might just tip my hat and say, "Alright, partner, let's see what fate's got in store. But mark my words, a year from now, it'll be a fair and square meetin' at that Last Chance Saloon, no funny business."

So, I reckon it ain't about cowardice, but it's about being smart and making sure this here duel is on the level.

Friday, February 16, 2024

OpenAI's synthetic video [Sora]

Here's the video they produced:

It's spectacular, no?

Here's their technical report, Video generation models as world simulators:

We explore large-scale training of generative models on video data. Specifically, we train text-conditional diffusion models jointly on videos and images of variable durations, resolutions and aspect ratios. We leverage a transformer architecture that operates on spacetime patches of video and image latent codes. Our largest model, Sora, is capable of generating a minute of high fidelity video. Our results suggest that scaling video generation models is a promising path towards building general purpose simulators of the physical world.

I think the title's a bit much, but business is business.

While admitting that the videos are "spectacular...cinematic," Gary Marcus is predictably skeptical, as one should be. His conclusion:

And importantly, I predict that many will be hard to remedy. Why? Because the glitches don’t stem from the data, they stem from a flaw in how the system reconstructs reality. One of the most fascinating things Sora’s weird physics glitches is most of these are NOT things that appears in the data. Rather, these glitches are in some ways akin to LLM “hallucinations”, artifacts from (roughly speaking) decompression from lossy compression. They don’t derive from the world.

More data won’t solve that problem. And like other generative AI systems, there is no way to encode (and guarantee) constraints like “be truthful” or “obey the laws of physics”or “don’t just invent (or eliminate) objects”.

Watch the video and reach your own conclusions.

Sunday, February 11, 2024

Is Altman's attempt to raise $7T in fact a sign that things aren't going to well?

Altman's attempt at a $ 7T raise seems a bit extreme, even for him. The same with Hinton's recent hallucinatory diatribe against Gary Marcus. Sutskever's been saying some weird things as well). Are things on the Great Rush to AGI falling behind schedule? Are these guys getting just a bit worried and expressing it by doubling down?

Here's the article about the Bill Gates interview:

In an interview with German business newspaper Handelsblatt, the 67-year-old said that there were plenty of reasons to believe that GPT technology reached a plateau. He also admitted that he could be wrong. He said that contrary to what people at OpenAI think about GPT-5, he believes that current generative AI has reached a ceiling. Talking about benchmark, he termed the leap from GPT-2 to GPT-4 as “incredible”.  

Wednesday, December 6, 2023

Are LLMs approaching the limiting wall beyond which performance simply wanders around in the same general region?

Cade Metz and Nico Grant, Google Updates Bard Chatbot With ‘Gemini’ A.I. As It Chases ChatGPT, NYTimes, 2023. From the article:

Google has built three versions of Gemini with three different sets of skills. The largest, Ultra, is designed to tackle complex tasks and will debut next year. Pro, the mid-tier offering, will be rolled out to numerous Google services, starting Wednesday, with the Bard chatbot. Nano, the smallest version, will power some features on the Pixel 8 Pro smartphone, such as summarizing audio recordings and offering suggested text responses in WhatsApp starting Wednesday. [...]

With Gemini, Google has also trained the technology on digital images and sounds. It is what researchers call a “multimodal” system, meaning it can analyze and respond to both images and sounds. If you give it a math problem that includes lines, shapes and other images, for example, it can answer in much the way a high school student would.

That portion of the technology, however, will not be available to consumers until sometime next year. Google also acknowledged that like similar systems, Gemini is prone to mistakes. It can get facts wrong or even “hallucinate” — make stuff up.

Wednesday, November 15, 2023

A dialectical view of the history of AI, Part 1: We’re only in the antithesis phase. [A synthesis is in the future.]

The idea that history proceeds by way of dialectical change is due primarily to Hegel and Marx. While I read bit of both early in my career, I haven’t been deeply influenced by either of them. Nonetheless I find the notion of dialectical change working out through history to be a useful way of thinking about the history of AI. Because it implies that that history is more than just one thing of another.

This dialectical process is generally schematized as a movement from a thesis, to an antithesis, and finally, to a synthesis on a “higher level,” whatever that is. The technical term is Aufhebung. Wikipedia:

In Hegel, the term Aufhebung has the apparently contradictory implications of both preserving and changing, and eventually advancement (the German verb aufheben means "to cancel", "to keep" and "to pick up"). The tension between these senses suits what Hegel is trying to talk about. In sublation, a term or concept is both preserved and changed through its dialectical interplay with another term or concept. Sublation is the motor by which the dialectic functions.

So, why do I think the history of AI is best conceived in this way? The first era, THESIS, running from the 1950s up through and into the 1980s, was based on top-down deductive symbolic methods. The second era, ANTITHESIS, which began its ascent in the 1990s and now reigns, is based on bottom-up statistical methods. These are conceptually and computationally quite different, opposite, if you will. As for the third era, SYNTHESIS, well, we don’t even know if there will be a third era. Perhaps the second, the current, era will take us all the way, whatever that means. Color me skeptical. I believe there will be a third era, and that it will involve a synthesis of conceptual ideas computational techniques from the previous eras.

Note, though, that I will be concentrating on efforts to model language. In the first place, that’s what I know best. More importantly, however, it is the work on language that is currently evoking the most fevered speculations about the future of AI.

Let’s take a look. Find a comfortable chair, adjust the lighting, pour yourself a drink, sit back, relax, and read. This is going to take a while.

Symbolic AI: Thesis

The pursuit of artificial intelligence started back in the 1950s it began with certain ideas and certain computational capabilities. The latter were crude and radically underpowered by today’s standards. As for the ideas, we need two more or less independent starting points. One gives us the term “artificial intelligence” (AI), which John McCarthy coined in connection with a conference held at Dartmouth in 1956. The other is associated with the pursuit of machine translation (MT) which, in the United States, meant translating Russian technical documents into English. MT was funded primarily by the Defense Department.

The goal of MT was practical, relentlessly practical. There was no talk of intelligence and Turing tests and the like. The only thing that mattered was being able to take a Russian text, feed it into a computer, and get out a competent English translation of that text. Promises was made, but little was delivered. The Defense Department pulled the plug on that work in the mid-1960s. Researchers in MT then proceeded to rebrand themselves as investigators of computational linguistics (CL).

Meanwhile researchers in AI gave themselves a very different agenda. They were gunning for human intelligence and were constantly predicting we’d achieve it within a decade or so. They adopted chess as one of their intellectual testing grounds. Thus, in a paper published in 1958 in the IBM Journal of Research and Development, Newell, Shaw, and Simon wrote that if “one could devise a successful chess machine, one would seem to have penetrated to the core of human intellectual endeavor.” In a famous paper, John McCarthy dubbed chess to be the Drosophila of AI.

Chess isn’t the only thing that attracted these researchers, they also worked on things like heuristic search, logic, and proving theorems in geometry. That is, they choose domains which, like chess, were highly rationalized. Chess, like all highly formalized systems, is grounded in a fixed set of rules. We have a board with 64 squares, six kinds of pieces with tightly specified rules of deployment, and a few other rules governing the terms of play. A seeming unending variety of chess games then unfolded from these simple primitive means according to the skill and ingenuity, aka intelligence, of the players.

This regime, termed symbolic AI in retrospect, remained in force through the 1980s and into the 1990s. However, trouble began showing up in the 1970s. To be sure, the optimistic predictions of the early years hadn’t come to pass; humans still beat computers at chess, for example. But those were mere setbacks.

These problems were deeper. While the computational linguistics were still working on machine translation, they were also interested in speech recognition and speech understanding. Stated simply, speech recognition goes like this: You give a computer a string of spoken language and it transcribes it into written language. The AI folks were interested in this as well. It’s not the sort of thing humans give a moment’s thought to; we simply do it. It is mere perception. It was proving to be surprisingly difficult. The AI folks also turned to computer vision: Give a computer a visual image and have it identify the object. That was difficult as well, even with such simple graphic objects as printed letters.

Speech understanding, however, was on the face of it intrinsically more difficult. Not only does the system have recognize the speech, but it must understand what is said. But how do you determine whether or not the computer understood what you said. You could ask it: “Do you understand?” And if it replies, “yes,” then what? You give it something to do.

That’s what the DARPA Speech Understanding Project set out to do in over five years in the early to mid 1970s. Understanding would be tested by having the computer answer questions about database entries. Three independent projects were funded; interesting and influential research was done. But those systems, interesting as they were, were not remotely as capable as Siri or Alex in our time, which run on vastly more compute encompassed in much smaller packages. We were a long way from having a computer system that could converse as fluently as a toddler, much less discourse intelligently on the weather, current events, the fall of Rome, the Mongol’s conquest of China, or how to build a fusion reactor.

During the 1980s the commercial development of AI petered out and a so-called AI Winter settled in. It would seem that AI and CL had hit the proverbial wall. The classical era, the era of symbolic computing, was all but over.

Monday, January 2, 2023

AI DEBATE 3: The AGI Debate [hosted in Montreal]

The Pivotal Discussion in Shaping the Path of AGI's Global Discourse.

Five panels of the world's most distinguished researchers and experts on AGI :

Panel 1: Cognition and Neuroscience

Panel 2: Common Sense

Panel 3: Architecture

Panel 4: Ethics and morality

Panel 5: AI: Policy and Net Contribution

Speakers: Erik Brynjolfsson, Yejin Choi, Noam Chomsky, Jeff Clune, David Ferrucci, Artur d'Avila Garcez, Michelle Rempel Garner, Dileep George, Ben Goertzel, Sara Hooker, Anja Kaspersen, Konrad Kording, Kai-Fu Lee, Francesca Rossi, Jürgen Schmidhuber and Angela Sheffield.

Moderator and co-organizer (with Vincent Boucher) : Gary Marcus. Gary Marcus is a leading voice in artificial intelligence. Scientist, best-selling author, and entrepreneur, he was Founder and C.E.O. of Geometric Intelligence, a machine-learning company acquired by Uber in 2016. His most recent book, Rebooting AI, co-authored with Ernest Davis, is one of Forbes’s 7 Must Read Books in AI.

Vincent Boucher is President of Montreal.AI and Quebec.AI.

Website: https://agidebate.com/

Gary Marcus remarks:

Every single speaker was both articulate, and pointed. Two—Schmidhuber and Clune—notably more optimistic about the potential of current techniques. (Sparks flew when they expressed that optimism.) But not one speaker thought that current AI was anything like the holy grail of artificial general intelligence. (Clune thought we might get there, by 2030.) Virtually every speaker thought that things were about to get wild—and not necessarily in entirely good ways.

I urge you, if you care about artificial intelligence, and its future, and its impact on society, to watch the debate, in full. That’s a huge commitment, 3.5 hours, but one thing that I think all of our speakers could agree on is that artificial intelligence, in whatever form it currently is, is about to have a huge impact on society.

Marcus recommends Tiernan Ray's account of the debate for ZDNet. I have posted excerpts from it.

* * * * *

Noam Chomsky spoke first, remarking that he saw no scientific value in current AI. Why? Because machine learning systems can learn any language whatsoever and so can tell us nothing about specifically human language. In making this remark I suppose he is implicitly analogizing these AI systems to the language acquisition device (LAD) he has been theorizing about for years. The LAD can learn all and only human languages and so is quite different the engines of machine learning.

I wonder. Do we actually know that current ML engines can learn any language? Are we even sure they can learn English? To be sure, many of them produce fluent English output, but they also make many blunders. The blunders, however, tend to be about semantics, not syntax. Chomsky's theorizing has always been centered on syntax, so semantic mistakes, however glaring, may be irrelevant to him. Can one be said to know a language without knowing its semantics? But none of us knows the full semantics of any natural language.

Cf. my recent remarks, On limits to the ability of LLMs to approximate the mind’s structure, December 27, 2022.

Friday, November 25, 2022

Meta announces CICERO, an AI that plays Diplomacy

Gary Marcus and Ernest Davis have posted an interesting evaluation of Cicero: What does Meta AI’s Diplomacy-winning Cicero Mean for AI? [Hint: It’s not all about scaling]

First:

The first thing to realize is that Cicero is a very complex system. Its high-level structure is considerably more complex than systems like AlphaZero, which mastered Go and chess, or GPT-3 which focuses purely on sequences of words. Some of that complexity is immediately apparent in the flowchart; whereas a lot of recent models are something like data-in, action out, with some kind of unified system (say a Transformer) in between, Cicero is heavily prestructured, in advance of any learning or training, with a carefully-designed bespoke architecture that is divided into multiple modules and streams, each with their own specialization.

A marvel, but...

Cicero is in many ways a marvel; it has achieved by far the deepest and most extensive integration of language and action in a dynamic world of any AI system built to date. It has also succeeded in carrying out complex interactions with humans of a form not previously seen.

But it is also striking in how it does that. Strikingly, and in opposition to much of the Zeitgeist, Cicero relies quite heavily on hand-crafting, both in the data sets, and in the architecture; in this sense it is in many ways more reminiscent of classical “Good Old Fashioned AI” than deep learning systems that tend to be less structured, and less customized to particular problems. There is far more innateness here than we have typically seen in recent AI systems

Also, it is worth noting that some aspects of Cicero use a neurosymbolic approach to AI, such as the association of messages in language with symbolic representation of actions, the built-in (innate) understanding of dialogue structure, the nature of lying as a phenomenon that modifies the significance of utterances, and so forth.

That said, it’s less clear to us how generalizable the particulars of Cicero are.

In sum:

Cicero makes extensive use of machine learning, but is hardly a poster child for simply making ever bigger models (so-called “scaling maximalism”), nor for the currently popular view of “end-to-end” machine learning of in which some single general learning algorithm applies across the board, with little internal structure and zero innate knowledge. At execution time, Cicero consists of a complex array of separate hand-crafted modules with complex interactions. At training time, it draws on a wide range of training materials, some built by experts specifically for Cicero, some synthesized in programs hand-crafted by experts. [...]

Our final takeaway? We have known for some time that machine learning is valuable; but too often nowadays ML is a taken as universal solvent—as if the rest of AI was irrelevant—and left to do everything on its own. Cicero may change that calculus. If Cicero is any guide, machine learning may ultimately prove to be even more valuable if it is embedded in highly structured systems, with a fair amount of innate, sometimes neurosymbolic machinery.

There's much more in their article.

Wednesday, June 29, 2022

Gary Marcus and Our Innate Linguistic Capacity [Relational Net Primer]

During the period that I had been writing my primer on relational networks and attractor landscapes, Gary Marcus had taken to Twitter to challenge the deep learning community on the need for symbolic processing. I agree that symbolic processing is necessary for any device that is going to approximate or match the human capacity for abstract thought. However, as I’ve made clear in the primer, I don’t think that symbolic processing is primitive. Neural nets are primitive and symbolic processinig is implimented in them.

Marcus on Innateness

It is not clear to me just what Marcus thinks about this. At the beginning of chapter 6 of The Algebraic Mind[1], which is, I believe, his major theoretical statement, Marcus says (p. 143): “The suggestion that I consider in this chapter is that the machinery of symbol-manipulation is included in the set of things that are initially available to the child, prior to experience with the external world.” Later on he will reject the “idea that the DNA could specify a point-by-point wiring diagram for the human brain” (p. 156). After weaving his way through a mind-boggling network of evidence about development, Marcus offers (p. 165):

Genetically driven mechanisms (such as the cascades described above) could, in tandem with activity-dependence, lead to the construction of the machinery of symbol-manipulation—without in any way depending on learning, allowing a reconciliation of nativism with developmental flexibility.

At this point I’m afraid I’m driven to echo the great Roberto Duran and to say “no más”. It makes my head hurt.

Let me skip ahead to his final chapter, where he says (p. 172):

As I suggested in chapter 6, differences between the cognition of humans and other primates may lie not so much in the basic-level components but in how those components are interconnected. To understand human cognition, we need to understand how basic computational components are integrated into more complex devices—such as parsers, language acquisition devices, modules for recognizing objects, and so forth—and we need to understand how our knowledge is structured, what sorts of basic conceptual distinctions we represent, and so forth.

I agree with Marcus on that first sentence. I’m not so sure about some of the rest, though I do believe that last remark, after “so forth.” That structure what most of the primer is about.

Now, the word “modules” is most important in an intellectual tradition of which I’m skeptical. If by parser Marcus means the sorts of things that are the staple of classical computational linguistics, then I doubt that human brains have such things. That’s not an argument, and this isn’t the place to make one, but I will point out that, from its origins in the problem of machine translation in the 1950s until well into the 1970s, computational linguistics had little to no semantics to speak of – nor, for that matter, has it ever developed more than a smattering of semantics. Without semantics, syntax is what is left. If syntax is what you have, then you need a sophisticated and elaborate parser. I believe that human language is grounded in semantics, which is in turn grounded in perception and action, and that syntax supports semantics. Relatively little parsing, as such, is required.

As for the language acquisition device, we have already seen that Marcus believes symbol manipulation is available to children “prior to experience with the external world.” I certainly do not believe that the mind is a proverbial blank state. Infants do come into the world with some fairly specific perceptual and behavioral equipment. But whether or not that includes a language acquisition device, well, let me evade the issue by spining a tale.

Teaching Chimpanzees Language

This is a tale, not a true one, but a thought experiment. I invented it while thinking about the origins of language. I came up with this tale while thinking about various early attempts that had been made to teach chimpanzees language. All of them ended in failure [2]. In the most intense of these efforts, Keith and Cathy Hayes raised a baby chimp in their household from 1947 to 1954. But that close and sustained interaction with Vicki, the young chimp in question, was not sufficient.

Then in the late 1960s Allen and Beatrice Gardner began training a chimp, Washoe, in Ameslan, a sign language used among the deaf. This effort was far more successful. Within three years Washoe had a vocabulary of Ameslan 85 signs and she sometimes created signs of her own.

The results startled the scientific community and precipitated both more research along similar lines — as well as work where chimps communicated by pressing iconically identified buttons on a computerized panel — and considerable controversy over whether or not ape language was REAL language. That controversy is of little direct interest to me, though I certainly favor the view that this interesting behavior is not really language. What is interesting is the fact that these various chimps managed even the modest language that they did.

The string of earlier failures had led to a cessation of attempts. It seemed impossible to teach language to apes. It would seem that they just didn’t have the capacity. Upon reflection, however, the research community came to suspect that the problem might have more to do with vocal control than with central cognitive capacity. And so the Gardners acted on that supposition and succeeded where others had failed. It turns out that whatever chimpanzee cognitive capacity was, it was capable of orchestrating surprising behavior.

Note that nothing had changed about the chimpanzees. Those that learned some Ameslan signs, and those that learned to press buttons on a panel, were of the same species as those that had earlier failed to learn to speak. What had changed was the environment. The (researchers in the) environment no longer asked for vocalizations. The environment asked for gestures, or button presses. These the chimps could provide, thereby allowing them to communicate with the (researchers in the) environment in a new way.

It seemed to me that this provided a way to attack the problem of language origins from a slightly different angle.

How Aliens From Outer Space Brought Us Language

I imagined that a long time ago groups of very clever apes – more so than any extant species – were living on the African savannas. One day some flying saucers appeared in the sky and landed. The extra-terrestrials who emerged were extraordinarily adept at interacting with those apes and were entirely benevolent in their actions. These space aliens taught the apes how to sing and dance and talk and tell stories, and so forth. Then, after thirty years or so, the ETs left without a trace. The apes had absorbed the extra-terrestrials’ lessons so well that they were able to pass them on to their progeny generation after generation. Thus human culture and history were born.

Now, unless you actually believe in UFOs, and in the benevolence of their crews, this little fantasy does not seem very promising, for it is a fantasy about things that certainly never happened. But if it had happened, it does seem to remove the mystery from language’s origins. Instead of something from nothing we have language handed to us on a platter. We learned it from some other folks, perhaps they were short little fellows with green skin, or perhaps they were the more modern style of aliens with pale complexions, catlike pupils in almond eyes and elongated heads. This story hasn’t taught as anything new about just how language works, but one source of mystery has disappeared.

But, and here is where we get to the heart of the matter, what would have to have been true in order for this to have worked? Just as the chimps before Ameslan were genetically the same as those after, so the clever apes before alien-instruction were the same as the proto-humans after. The species has not changed, the genome is the same – at least for the initial generation. The capacity for language would have to have been inherent in the brains of those clever apes. All the aliens did was activate that capacity. Once that happened the newly emergent proto-humans were able to sustain and further develop language on their own. Thus the critical event is something that precipitates a reconfiguration of existing capabilities, a Gestalt switch. The rabbit has become a duck, or vice versa, the crone a young lady, or vice versa.

Where’s the Language Acquisition Device?

However, we’re not interested in the phylogenetic origins of language, we’re interested in how children acquire it. Whatever their genetic endowment, they live in a world surrounded by language speakers, and some of them are closely attuned to the infant’s needs, desires, and evoling capacities. It’s not clear to me just what specialized language acquisition device the infant needs.

The problem I have in thinking about this is very much like the problem I have identifying the acceleration subsystem in an automobile. I know that the engine has more effect on acceleration than the backseat upholstery, and the tires are more important than the cup-holders, but anything with mass affects the acceleration. Is there anything in the engine that is there specifically to enhance acceleration and nothing else, or is it a matter of proportion, strength, and adjustment of the components necessary to the engine?

That’s my problem with the idea of a language acquisition device. Given that infants are born with various perceptual capabilites (which Marcus recounts) it is not obvious whether or not we need anything more specific to language than, for example, the ability to track adult speech patterns that Condon and Sander reported in 1974 [3]. I think it’s a bit much to call that a language acquisition device. If you must have one, why not say that the human brain as a whole is, among many other things, a language acquisition device?

References

[1] Marcus, Gary F (2001). The Algebraic Mind (Learning, Development, and Conceptual Change). MIT Press. Kindle Edition.

[2] Linden, Eugene. (1974). Apes, Men, and Language. New York: Saturday Review Press, E. P. Dutton.

[3] Condon, W.S., & Sander, L.W. (1974). Neonate movement is synchronized with adult speech: Interactional participation and language acquisition. Science, 183, 99-101.