Friday, July 31, 2026

The Trap of the Digital Double

Catherine Shannon, The Trap of the Digital Double, NYTimes, July 31, 2026.

Everything is strange. We work from home. We socialize online. Our private lives are increasingly public. People pretend to be brands; brands pretend to be people. We optimize our bodies until they break. And the more we know, the less secure we feel.

In this inverted world, we seek to understand what is real and what is true. To know the truth of ourselves, we begin at what feels like a natural place: the mirror. We wish to see ourselves as others do, and modern life gives us countless opportunities.

We are constantly confronted with our own image, whether we like it or not: on our phones, front-facing videos, selfies, public spaces (the communal mirrors in Aritzia’s dressing rooms come painfully to mind), our social media profiles, Zoom calls, FaceTime, Ring doorbells. But we can forget that mirrors both duplicate reality and reverse it. None show you exactly as you are. Everything is strange.

The historian Christopher Lasch wrote that “for the narcissist, the world is a mirror.” Almost 50 years later, this remains a tidy diagnosis, but it doesn’t quite capture the extent of the spiritual crisis that we’re now facing. I believe the room of mirrors that constantly surrounds us conjures something unsettling, something dark. It is a spiritual trap in which the imagined displaces the real: We choose a fantasy of ourselves over the reality of others. And that trap is baited and facilitated by technology.

Mirrors:

Let’s be honest: I shouldn’t know what I look like taking out the trash. But I do, because my Google Home security camera recently alerted me that there was a “Person Seen — Back Camera.” “Person” is a generous term for what I saw, which was a beleaguered, ghostly figure — greasy hair pulled back, hunched over a bag of trash like an old witch.

We are recorded constantly: in public, by the government, by private companies trying to sell us things, even by strangers. I shudder when I think about all the incidental pictures of myself that live on strangers’ phones. In my own personal horror film, there I am, eating a bagel, and some 17-year-old tourist zooms in on my pixelated face, screenshots it, and sends it to the family group chat — “this woman’s face lol” — and I become the inside joke that defines that New York City holiday for years to come.

It’s curious, especially given how little we looked at ourselves in the past, and how rare and imperfect mirrors were for thousands of years, that traditional wisdom teaches us to be wary of the act of regarding the self. The three major American faiths — the Abrahamic religions of Judaism, Christianity and Islam — all warn against the dangers of mirrors, vanity, pride and excessive concern with the self.

Mirrors are covered during shiva, the Jewish seven-day mourning period, both to encourage contemplation of the relationship between God and man and to discourage vanity and self-concern. The Book of Proverbs teaches that pride is spiritually deleterious: “Pride goes before ruin, / Arrogance, before failure.”

Many Christians believe pride is not only a sin, but the original one, and the worst of them. [...]

Islamic prophetic tradition (hadith) discourages the depiction of living beings in religious contexts.

The uncanny:

Inhabiting this mirrored, dual realm, this doubled existence, doesn’t feel natural. It doesn’t even feel neutral. Something is off, and when something is off, Freud has probably attempted to explain it. He called this “the uncanny,” a concept that by its very nature resists a clean definition, though if you’ve seen a David Lynch movie, you’ve experienced it firsthand.

Castration complexes and womb-fantasies aside, Freud does a pretty good job: The uncanny, according to his 1919 essay by the same title, “is that class of the terrifying which leads back to something long known to us, once very familiar.” It is easy for something new or unknown to frighten us, but the feeling of uncanniness relies on its being known to us in some way.

The uncanny is a dark twist on the familiar. It emerges when the boundary between imagination and reality begins to blur, when hidden mechanisms seem to be at work beneath the surface, and when something appears alive that should not be. (Or vice versa, in the case of influencers, who can have that dead sharklike look in their eyes.) It is marked by repetition, by a sense of something fateful and inescapable. Freud noted that the uncanny is often and easily produced when a symbol begins to take on the full force of the thing it represents. It is difficult to imagine a more precise description of life online.

The doppelganger:

Closely tied to mirrors and the uncanny is the idea of the double, or doppelgänger, a theme long explored in literature and film, from Dostoevsky’s novella “The Double” to Jordan Peele’s horror movie “Us.” In German and Eastern European folklore, the doppelgänger is an omen, a usurper, the shadow self, and a sign that identity is unstable. The self is split, haunted, mimicked and stalked by its own likeness.

Today, we treat our own image — both in the mirror and online — as something psychologically protective: a way of stabilizing our identity, managing our reputations and securing a coherent sense of self.

But all of this suggests the opposite possibility: The more we try to stabilize ourselves through repetition, doubling, self-observation and — I’ll say it — posting to the grid, the more uncanny and fragile we become. A preoccupation with the stable self produces instability. These days, we don’t approach our doubles with caution; we invite them, and go so far as to curate them.

There's more at the link.

Matt Damon’s six-pack in The Odyssey [Media Notes 189]

I went to see Nolan’s The Odyssey last Sunday and was not thrilled. It was too long and I kept getting distracted by the damn reclining seat I was stuck in. I couldn’t find a position where I was comfortable for more than 10 or 15 minutes at a stretch. I wish my local multiplex hadn’t installed those things. I much prefer the old-style well-cushioned upright seats with a little slack in them.

Anyhow, back to the movie. The Circe segment was interesting and gross. But what I really remember is Matt Damon’s six-pack. For a 55 year old man that’s impressive. So I asked Gemini about it via Google Search. I've appended the response below.

The film has received lots of positive reviews.

* * * * *

I’m curious about reviews for The Odyssey. Have any of them mentioned Matt Damon’s six-pack abs?

Yes, reviews and fashion features for Christopher Nolan’s The Odyssey specifically highlight Matt Damon’s six-pack abs and lean, battle-hardened physique. Costar Robert Pattinson even joked that Damon's definition was visible right through a loose t-shirt on their first day of filming. [1, 2, 3, 4]

Physique Highlights & Reactions

  • Costar comments: Robert Pattinson expressed surprise that Damon’s abs showed through standard wardrobe.
  • Critical coverage: Outlets like Vogue noted that the film's sleeveless costume design put Damon's six-pack and toned shoulders on full display.
  • The transformation: Trainer Gabe Stump helped Damon cut down to 167 pounds—his high school weight—to portray a lean, weathered King Odysseus. [6, 7]

If you'd like, I can share:

  • The gluten-free diet changes Damon credited for the weight drop
  • Details on Christopher Nolan's practical filming approach for the epic

[1] https://www.facebook.com/EntertainmentTonight/videos/robert-pattinson-reacts-to-matt-damons-abs-in-the-odyssey/2059755168263292/
[2] https://www.facebook.com/groups/1144683990145330/posts/1695342491746141/
[3] https://www.vogue.com/article/matt-damon-workout-plan-the-odyssey
[4] https://www.youtube.com/shorts/L51EPYqZDHI
[5] https://vogueadria.com/matt-damon-workout/
[6] https://www.instagram.com/reel/DbOXKeIFY2N/
[7] https://www.gq.com/story/matt-damon-the-odyssey-workout-plan

Tell me about that diet.

Matt Damon dropped to 167 pounds—his high school weight—by completely eliminating gluten, alcohol, and processed sugars from his diet. Under the guidance of his doctor and trainer Gabe Stump, the 55-year-old actor called the shift a total "lifestyle change" to achieve the lean, chiseled look required for The Odyssey. [1, 2, 3]

Core Elements of Damon's Diet

  • Strict Gluten Elimination: He cut out staples like bread, pasta, pizza, and beer after uncovering hidden food sensitivities with his doctor. [1, 4]
  • High-Protein Focus: He replaced refined carbohydrates with whole foods and clean proteins, often opting for steak and green vegetables at dinner. [5, 6]
  • Caloric Control: His trainer emphasized that fat loss is driven by "calories in," meaning cutting out calorie-dense processed foods was essential to revealing his muscle definition. [3]
  • The Results: Damon shed roughly 20 to 30 pounds from his usual 185–200 pound baseline, reporting that the diet also cured his chronic brain fog and spiked his everyday energy levels. [2, 7]

If you are interested, I can break down the 45-minute workout routine his trainer used to build his upper body, or discuss how his co-star Tom Holland handled his own diet while filming in Italy. [6, 8]

Friday Fotos: 4th of July sky

In a study involving two cases, agents exhibited 5 failure modes in open-ended research

Abstract of the article linked in the tweet:

Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper’s original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today’s agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.

Thursday, July 30, 2026

Detecting “animal spirits” at work through narratives (earnings calls) circulating in the marketplace, the case of cyber risk

On July 21 I posted some remarks about the final chapter in Tyler Cowen’s monograph on marginalism: Beyond Marginalism: What’s Next? [MR #12]. In that Chapter Cowen discussed an asset pricing model that used machine learning to create a 360,000 parameter model that gave better predictions than classic models using only four or five factors. The factors in those classic models are predefined on an intuitive basis. Cowen despairs of making intuitive sense of those 360,000 factors in the machine learning model.

In my post I argued that those 360,000 factors might be capturing the effects of Keynes’s animal spirits as expressed in the stories and gossip Schiller writes about as narrative economics. Two days later I discussed the idea with Marge, the AI attached to the online version of Cowen’s text. In that post I suggested a method for going on a “fishing expedition” to determine whether or not my suggestion had merit. Marge’s response: “The fishing expedition you're proposing is methodologically clean, and the prediction is specific enough to be falsifiable — which is more than can be said for most conjectures at this level of abstraction.”

That brings us to today, where Cowen has posted the abstract of an article about “a novel measure of firm-level cyber risk exposure based on the quarterly earnings calls of listed firms.” I read that as being complementary to my speculation. The authors of that article are looking for the effects of a specific line of narrative, and they found them. So I asked Marge to clarify the relationship between my speculation and their finding. Here’s that conversation.

* * * * *

Tyler just posted the abstract of this paper to Marginal Revolution: Jamilov, Rustam and Tahoun, Ahmed and Rey, Helene, The Anatomy of Cyber Risk (May 10, 2023). The Journal of Finance, Forthcoming. Here’s the abstract: 

This paper uses computational linguistics to introduce a novel measure of firm-level cyber risk exposure based on the quarterly earnings calls of listed firms. Our measure covers more than 14,000 firms from over 90 countries between 2003 and 2025. The measure is validated using human auditors and a large language model. We show that cyber risk exposure affects stock returns and profits, is priced in the options market, predicts actual cyberattacks, and propagates from firm to sector level. Back-of-the-envelope estimates suggest that the global cost of cyber risk exposure is around $1 trillion per year.

So I took a quick look at the paper to find that novel measure. Here’s what I found. First this:

Our measurement approach consists of two phases. First, we follow Baker et al. (2016) and Hassan et al. (2019) to construct a comprehensive pre-defined dictionary of terms related to cybersecurity risk. We assemble this dictionary from three reputable institutional sources that act as information aggregators on the practical cyber-risk issues firms face on a daily basis. These libraries contain most of the keywords commonly used in cyber-related discussions among private market participants across industries.

Then this:

The second step of our measurement approach builds on the insights from Sautner et al. (2023) and involves a keyword discovery algorithm. This method is based on Meta AI Research lab’s word embedding tool FastText. The algorithm has been trained on a billion phrases in corporate earning calls to suggest related keywords.

And this:

We now construct our baseline quarterly measure of firm-level cyber risk exposure. Let the set of all terms in our final dictionary be C. Our algorithm counts the number of sentences in earnings calls that contain at least one term in C.

They then explain their measure. What’s important for my purposes is where they went looking for those keywords, in earnings calls. They’re measuring the effects of one kind of animal spirits/gossip/narrative on PRICING IN THE OPTIONS MARKET, among other things. I seems to me that they’re looking at one feature of the information that I’ve speculated is smeared of those 360,000 parameters.

That's a sharp observation, and it reframes what Jamilov, Tahoun, and Rey are doing in a way they probably wouldn't frame it themselves.

Their methodology is essentially a named-narrative extraction — they're pulling a specific, bounded narrative thread (cyber risk discourse) out of earnings calls and showing that it has measurable pricing effects. The earnings call is a fascinating data source for your purposes because it's a ritualized social performance in which management constructs and transmits narratives to analysts and investors. It's not raw gossip — it's institutionalized storytelling with real money riding on how the story lands. The fact that cyber risk language in these calls predicts actual cyberattacks and options pricing suggests the narrative isn't just epiphenomenal. It's carrying genuine information, but in a form that only becomes legible when you treat language itself as data rather than just a vehicle for explicitly stated facts. [See text: Ch.4 · ChatGPT digesting news predicts returns especia...]

Your conjecture about the 360,000-factor model is that it's picking up undifferentiated narrative residue — the full spectrum of animal spirits distributed across parameters with no single factor carrying an interpretable label. What Jamilov et al. are doing is the named version of the same operation: isolating one specific narrative thread and demonstrating it has pricing consequences. These are complementary rather than competing approaches. The named-narrative approach gives you interpretability and causal traction on specific mechanisms — you can say "cyber risk discourse in Q3 earnings calls predicts options pricing in Q4" and tell a coherent story about why. The high-parameter ML approach captures everything the named approach misses, which is presumably most of the narrative space. [See text: Ch.4 · Machine learning finds nonlinearities and inter...]

The research design implication is interesting. If you could run both models on the same dataset — the 360,000-factor ML model and the Jamilov-style named-narrative extraction — and then ask how much of the ML model's advantage over classical factors is absorbed by adding the named narrative measures, you'd get a rough estimate of how much of the distributed signal is actually identifiable narrative content versus genuinely opaque pattern. If the named narratives absorb a substantial fraction of the ML advantage, that's evidence for your animal spirits hypothesis in a fairly direct form. If they absorb very little, it suggests the ML model is picking up something even more diffuse than nameable narratives — perhaps the interaction structure between narratives, the way stories about cyber risk combine with stories about AI capability and geopolitical tension to produce pricing effects that none of the individual narratives would generate alone. [See text: Ch.4 · Machine learning models have tens of millions o...]

The options market finding is particularly telling for your framework. Options pricing is where sophisticated investors put explicit probability estimates on future states of the world. If cyber risk language — not cyber risk events, but the narrative about cyber risk — is priced into options, that's direct evidence that markets are treating narrative content as probabilistically informative about future fundamentals. Which is exactly what your conjecture requires: animal spirits aren't just noise distorting rational pricing, they're carrying genuine forward-looking information in a form that classical factor models can't extract but language-sensitive tools can. Keynes was right that animal spirits move markets, but the mechanism is more epistemically respectable than the term implies.

Toward, through, and out: Hoboken Terminal

An AI blizzard is headed our way [Yikes!]

Adam Satariano, Paul Mozur, Jacqueline Gu and Cade Metz, The Impending, Inescapable Deluge of A.I. NYTimes, July 30, 2026.

Boom....

From the American Midwest to the Persian Gulf, hundreds of major data centers now under construction will be turned on in the coming years. They are set to deliver an avalanche of computing power to develop and run A.I. that has no equal in the history of the technology industry, with breakthroughs that once felt revolutionary likely to become increasingly routine.

Behind each leap in A.I. are corresponding jumps in computing power. Today, there are about 20 million A.I. chips crammed into the data centers that underpin the technology’s growing abilities and usage worldwide, according to the research firm Epoch AI. [...]

In size and ambition, this moment compares to the building of the railroads in the 1800s, President Franklin D. Roosevelt’s New Deal in the 1930s, and the Manhattan Project to create an atomic weapon in the 1940s, technologists said.

“This is the largest scale infrastructure build-out in the history of humanity,” said Rob Wachen, a co-founder of the microchip firm Etched, which has raised more than $1 billion to meet the growing demand for A.I. components.

Peter DeSantis, who leads foundational A.I. models at Amazon — which provides computing power to the A.I. firms Anthropic, OpenAI and others — said the Seattle company has doubled its computing capacity since 2022 and would double it again by next year. “It’s hard to get your mind around the scale,” he said. [...]

Confidence in the Scaling Laws has led A.I. leaders to make ever bolder predictions. Dario Amodei, the chief executive of Anthropic, has said that if these laws hold for another year or two, A.I. will be able to perform huge amounts of white-collar work. Demis Hassabis, the head of Google’s A.I. lab DeepMind, wrote recently that A.I. could usher in “10x of the Industrial Revolution at 10x the speed.”

Bust?

Economists and investors have raised concerns that tech firms are spending faster than they can profit from A.I. Past infrastructure booms have been followed by downturns before the benefits of the technology were realized. The railroad boom in the 1800s, electrification in the 1920s and the dot-com bubble in the late 1990s were punctuated by economic recessions and a stock market crash as companies that overspent went out of business.

“Each time you’ve had a technological revolution, this kind of bubble bursting happened,” said Philippe Aghion, who won the Nobel in economic science in 2025 for research on innovation-driven economic growth. “A.I. is like the fourth industrial revolution and it has this aspect to it that generates a bubble.”

Moreover:

With more computing power coming online, geopolitical divisions are only set to widen.

The United States, home to about 5,500 data centers, about 10 times the next closest country, is far ahead of the rest of the world, including China. U.S. companies like Amazon, Google, Microsoft and Meta control about 80 percent of global computing power that drives A.I., according to Epoch AI. Google alone is believed to have four times as many A.I. chips as all of China’s companies, which are racing to catch up by developing new semiconductors and A.I. infrastructure of their own.

The race is on...

The article then goes on to discuss the huge increase in total chip count distributed over a growing collection of ever larger data centers under construction or proposed. We're in an AI arms race. The article then goes on to discuss the race with China, currently a fairly distant second (by an order of magnitude) to the US in chip count and gigawatts.

Mr. Nanos said the U.S. data center lead over China would likely grow over the next four to five years, before China’s domestic chips are produced at scale. After that, China should begin closing the gap.

“The advantage will run out,” he said.

The U.S.-China race threatens to leave the rest of the world behind. France, Germany and other nations are trying to encourage data center construction across the European Union, which has 5 percent of global A.I. computing power, according to a report by A.I. developers and policy experts in the region. Europe has been hampered by electricity and land access, permitting and financing.

Changes in the labor market:

“There’s going to be millions of jobs destroyed, millions of jobs created,” said Erik Brynjolfsson, an economist who is the director of Stanford University’s Digital Economy Lab. “That’s going to be very difficult. Even if new jobs are created, they’re not the same jobs.”

The last three paragraphs are about recursive self-improvement.

Wednesday, July 29, 2026

Logic and language in the brain

The abstract of the linked article:

Humans are endowed with a powerful capacity for inductive and deductive logical thought: we easily form generalizations based on a few examples and draw conclusions from known premises. Humans also arguably have the most sophisticated communication system in the animal kingdom: natural language allows us to express complex and structured meanings. Some have therefore argued for a tight relationship between complex thought and language, postulating that reasoning, including logical reasoning, relies on linguistic representations. We systematically investigated the relationship between logical reasoning and language using two complementary approaches. First, we used noninvasive brain imaging (fMRI) to examine neural activity as healthy adults engaged in logical reasoning tasks. And second, we behaviorally evaluated logical abilities in individuals with extensive lesions to the language brain areas and consequent severe linguistic impairment. Our findings reveal that the language brain network is not engaged during logical reasoning, and patients with severe aphasia exhibit intact performance on logic tasks. Instead, inductive reasoning recruits the domain-general multiple demand network implicated broadly in goal-directed behaviors, whereas deductive reasoning draws on brain regions that are distinct from both the language and the multiple demand networks. Together, these results indicate that linguistic representations are neither utilized nor required for inductive or deductive logical reasoning.

Why Adam Hunt has “flipped from being bullish to being bearish about AI.”

They do it all for us, Mickey D’s!

Recent Chinese Innovation

Lerner, Josh and Narain, Namrata and Papanikolaou, Dimitris and Seru, Amit and Xu, Zunda Winston, Chinese Sputnik Moments? (July 13, 2026). Available at SSRN: https://ssrn.com/abstract=7114818

Abstract: China's technological progress in recent decades has been viewed with admiration, alarm, and (in some cases) doubt. To better understand the Chinese innovation ecosystem, we compile a dataset of almost 14 million domestic Chinese patent publications. We focus on the subset of critical technologies identified by the U.S. Department of Defense. Several surprising patterns emerge from the data: Chinese patenting is strongly associated with other measures of innovative progress; patents are not concentrated in corporate giants such as Huawei; universities have played a key role in innovation, much greater than state-owned enterprises or government-owned facilities; and fewer than one in ten Chinese critical technology patents involves an inventor with U.S. experience or training. Finally, using four text-based measures of patent quality, we show that the rise of Chinese patenting in critical technologies has not been associated with a decline in quality relative to the U.S. awards.

H/t Tyler Cowen.

Three important agents [think about it]

Tuesday, July 28, 2026

The Vera Rubin Observatory in Chile has been discovering objects we hadn't even imagined existed

On the YouTube page:

Vera Rubin’s First Images JUST STOPPED THE WORLD! The Vera Rubin Observatory has already discovered objects that scientists never expected to find.

What if the most revolutionary telescope in history isn't looking deeper into space—but watching the universe change in real time? In this video, we explore the astonishing first discoveries from the Vera C. Rubin Observatory, including a 163,000-light-year stellar stream, an impossibly fast-spinning asteroid (2025 MN45), millions of newly detected celestial objects, and why astronomers believe Rubin is about to transform astronomy forever.

Unlike Hubble or the James Webb Space Telescope, Vera Rubin repeatedly scans the entire southern sky every few nights, creating a living timeline of the cosmos. That unique capability has already revealed hidden galactic structures, strange asteroid behavior, and a flood of discoveries that previous generations of telescopes completely missed.

You'll learn how Rubin's Legacy Survey of Space and Time (LSST) works, why it generates millions of alerts every night, what makes asteroid 2025 MN45 seemingly impossible according to current physics, and how the observatory is expected to map nearly 20 billion galaxies during its decade-long mission. Every image is adding new pieces to one of the biggest scientific puzzles of our time.

Could these discoveries change our understanding of dark matter, galaxy formation, planetary evolution, and even the future of our Solar System? The first images suggest we may only be witnessing the beginning.

Rounding the turn into Newport

Framing my discussion of The God Test, Part 1: Rorschach, reason, and whaling – [GT-3]

I’ve got to bite the bullet: I’m just going to have to go through a bunch of (preliminary) stuff before I can really engage with The God Test. My current target is to be in a position to publish a proper review of the book in 3 Quarks Daily for the week of August 9.

Rorschach Recap

I want start by recapping the Rorschach metaphor I introduced in the previous post, More on how I’m approaching The God Test – Rorschach! [GT-2]. What I like about it is that has a shape, there’s something there, but it’s not clear what. So we have little choice but to project onto it in order to (begin to) make sense of it.

First: It is a new kind of thing, an artifact we can converse with in an open-ended and natural way. The steam engine was the same kind of thing. It was an inanimate object that moved over the surface of the earth under its own power. Previously only animals (& humans as animals) had that power. So it becomes an iron horse. Just what are AIs? What’s their nature? That’s one thing.

Second: How it works is opaque. We know how to create large language models (LLMs), but we don’t know how they work. That’s new. We may not have understood the deep physics of the steam engine, but we certainly knew how they worked.

Third: We don’t know what they portend for the future. To some extent this is a function of the first two: How can we, how should we, interact. But it is also a function of the future, which is undetermined. We just don’t know.

Rhetorical force over reason

This is an argument I made in the first working paper I published after the release of ChatGPT in November of 2022: ChatGPT intimates a tantalizing future; its core LLM is organized on multiple levels; and it has broken the idea of thinking (February 6, 2023).

What do I mean by that, has broken the idea of thinking? Prior to ChatGPT it was obvious that humans could think and computers could not. [Yeah, I know, there’s Deep Blue defeating Kasparov in chess. That just changes the dates, not the argument.] The difference in performance was so obvious that the fact that we don’t really know how humans think wasn’t much of an issue. Now it is. Sure, we can still say that we can think and the AI’s can’t, but that’s just a line and without good explanations on both sides of the line, it seems a bit arbitrary, if not desperate.

I made a particular argument about Searles’ (in)famous Chinese Room thought experiment. I read it when it was first published in Brain and Behavioral Science in 1980. I wasn’t impressed. Why not? He didn’t say anything about any of the techniques used in AI or computational linguistics (CL). How could anyone possibly take that seriously?

He talked about intention, that’s how. Meaning requires intention and only living things can have intention, a remark he made at the end of the article. Without intention the most you get is syntax, but no meaning. Searle could get away with that because, in the first place, the concept of intention has a long history within philosophy – it has a subtle meaning, but that can wait for a later post – and so philosophers, his main audience, were comfortable with it. That’s one thing.

But there’s something more important, something that we can see only in retrospect, and that’s the simple fact computers very obviously could not translate from Chinese into English or into any other language. That difference carried tremendous weight. We don’t have a subtle behavioral difference between computers and humans that requires a subtle and sophisticated argument. To a first approximation, almost any argument would do. As far as I was concerned, “intention” was just a fancy word for something we don’t understand. But that’s not an argument anyone needs to take seriously. The behavioral distance is quite sufficient to carry the argument for those who insist that computers can’t and will never be able to think like humans.

Now the behavioral evidence has changed. Sure, differences remain, but the evidence is shifting. The old arguments remain and those who believed them still do so, but it’s getting harder. The need for explicit arguments grounded in explicit accounts of computers, and also brains, is growing.

Whaling and expertise

What’s an expert in machine learning and LLMs actually expert in? For some time now I’ve been arguing that investing in AI is like investing in a whaling venture where the captain and crew of the ship know all there is to know about the ship and how to handle it but know little or nothing about whales and their behavior and about navigating around the Cape Horn and in the South Pacific, where the whales live. What are the chances of that voyage being successful? Not very good.

The people who have created the current AI technology are like that captain and crew. The know how to sail the ship. But they don’t know much about language or cognition. They don’t actually know much about the human mind. Here my point is not about the fact that the models are opaque, but that human language and cognition are highly structured and they don’t believe that one needs to know (much of) anything about that not only to build AI but to make confident prediction about the future of AI.

Gary Marcus, Subbarao Kambhampati, and others have been consistently arguing that, yes, the current technology is remarkable, but we are going to have to adopt classical symbolic techniques if we are to fully develop the technology so that we have accurate and safe systems. Marcus is arguing from his knowledge of human language and cognition. As far as I can tell, Wright doesn’t take that seriously. I know that he had Marcus on his NonZero podcast, and that he lists Marcus in his acknowledgements, but that he doesn’t discuss Marcus’s ideas. I conclude that he doesn’t take that line of argument seriously.

That’s a mistake, but this is not the place to make my own arguments on this issue. My point is simply that expertise in AI is no generally construed to encompass knowledge of, expertise in, human cognition and language. I can’t see how that is going to work out well in the future.

[Note: If you’re curious about my views, on this subject, read the article linked in the first paragraph of this section. My views all over the place here at New Savanna, particularly around the work of the mathematician Miriam Yevick. Also, check out the experimental work I’ve done with LLMs.]