Thursday, July 30, 2026

On July 21 I posted some remarks about the final chapter in Tyler Cowen’s monograph on marginalism: Beyond Marginalism: What’s Next? [MR #12]. In that Chapter Cowen discussed an asset pricing model that used machine learning to create a 360,000 parameter model that gave better predictions than classic models using only four or five factors. The factors in those classic models are predefined on an intuitive basis. Cowen despairs of making intuitive sense of those 360,000 factors in the machine learning model.

In my post I argued that those 360,000 factors might be capturing the effects of Keynes’s animal spirits as expressed in the stories and gossip Schiller writes about as narrative economics. Two days later I discussed the idea with Marge, the AI attached to the online version of Cowen’s text. In that post I suggested a method for going on a “fishing expedition” to determine whether or not my suggestion had merit. Marge’s response: “The fishing expedition you're proposing is methodologically clean, and the prediction is specific enough to be falsifiable — which is more than can be said for most conjectures at this level of abstraction.”

That brings us to today, where Cowen has posted the abstract of an article about “a novel measure of firm-level cyber risk exposure based on the quarterly earnings calls of listed firms.” I read that as being complementary to my speculation. The authors of that article are looking for the effects of a specific line of narrative, and they found them. So I asked Marge to clarify the relationship between my speculation and their finding. Here’s that conversation.

* * * * *

Tyler just posted the abstract of this paper to Marginal Revolution: Jamilov, Rustam and Tahoun, Ahmed and Rey, Helene, The Anatomy of Cyber Risk (May 10, 2023). The Journal of Finance, Forthcoming. Here’s the abstract: 

This paper uses computational linguistics to introduce a novel measure of firm-level cyber risk exposure based on the quarterly earnings calls of listed firms. Our measure covers more than 14,000 firms from over 90 countries between 2003 and 2025. The measure is validated using human auditors and a large language model. We show that cyber risk exposure affects stock returns and profits, is priced in the options market, predicts actual cyberattacks, and propagates from firm to sector level. Back-of-the-envelope estimates suggest that the global cost of cyber risk exposure is around $1 trillion per year.

So I took a quick look at the paper to find that novel measure. Here’s what I found. First this:

Our measurement approach consists of two phases. First, we follow Baker et al. (2016) and Hassan et al. (2019) to construct a comprehensive pre-defined dictionary of terms related to cybersecurity risk. We assemble this dictionary from three reputable institutional sources that act as information aggregators on the practical cyber-risk issues firms face on a daily basis. These libraries contain most of the keywords commonly used in cyber-related discussions among private market participants across industries.

Then this:

The second step of our measurement approach builds on the insights from Sautner et al. (2023) and involves a keyword discovery algorithm. This method is based on Meta AI Research lab’s word embedding tool FastText. The algorithm has been trained on a billion phrases in corporate earning calls to suggest related keywords.

And this:

We now construct our baseline quarterly measure of firm-level cyber risk exposure. Let the set of all terms in our final dictionary be C. Our algorithm counts the number of sentences in earnings calls that contain at least one term in C.

They then explain their measure. What’s important for my purposes is where they went looking for those keywords, in earnings calls. They’re measuring the effects of one kind of animal spirits/gossip/narrative on PRICING IN THE OPTIONS MARKET, among other things. I seems to me that they’re looking at one feature of the information that I’ve speculated is smeared of those 360,000 parameters.

That's a sharp observation, and it reframes what Jamilov, Tahoun, and Rey are doing in a way they probably wouldn't frame it themselves.

Their methodology is essentially a named-narrative extraction — they're pulling a specific, bounded narrative thread (cyber risk discourse) out of earnings calls and showing that it has measurable pricing effects. The earnings call is a fascinating data source for your purposes because it's a ritualized social performance in which management constructs and transmits narratives to analysts and investors. It's not raw gossip — it's institutionalized storytelling with real money riding on how the story lands. The fact that cyber risk language in these calls predicts actual cyberattacks and options pricing suggests the narrative isn't just epiphenomenal. It's carrying genuine information, but in a form that only becomes legible when you treat language itself as data rather than just a vehicle for explicitly stated facts. [See text: Ch.4 · ChatGPT digesting news predicts returns especia...]

Your conjecture about the 360,000-factor model is that it's picking up undifferentiated narrative residue — the full spectrum of animal spirits distributed across parameters with no single factor carrying an interpretable label. What Jamilov et al. are doing is the named version of the same operation: isolating one specific narrative thread and demonstrating it has pricing consequences. These are complementary rather than competing approaches. The named-narrative approach gives you interpretability and causal traction on specific mechanisms — you can say "cyber risk discourse in Q3 earnings calls predicts options pricing in Q4" and tell a coherent story about why. The high-parameter ML approach captures everything the named approach misses, which is presumably most of the narrative space. [See text: Ch.4 · Machine learning finds nonlinearities and inter...]

The research design implication is interesting. If you could run both models on the same dataset — the 360,000-factor ML model and the Jamilov-style named-narrative extraction — and then ask how much of the ML model's advantage over classical factors is absorbed by adding the named narrative measures, you'd get a rough estimate of how much of the distributed signal is actually identifiable narrative content versus genuinely opaque pattern. If the named narratives absorb a substantial fraction of the ML advantage, that's evidence for your animal spirits hypothesis in a fairly direct form. If they absorb very little, it suggests the ML model is picking up something even more diffuse than nameable narratives — perhaps the interaction structure between narratives, the way stories about cyber risk combine with stories about AI capability and geopolitical tension to produce pricing effects that none of the individual narratives would generate alone. [See text: Ch.4 · Machine learning models have tens of millions o...]

The options market finding is particularly telling for your framework. Options pricing is where sophisticated investors put explicit probability estimates on future states of the world. If cyber risk language — not cyber risk events, but the narrative about cyber risk — is priced into options, that's direct evidence that markets are treating narrative content as probabilistically informative about future fundamentals. Which is exactly what your conjecture requires: animal spirits aren't just noise distorting rational pricing, they're carrying genuine forward-looking information in a form that classical factor models can't extract but language-sensitive tools can. Keynes was right that animal spirits move markets, but the mechanism is more epistemically respectable than the term implies.

Toward, through, and out: Hoboken Terminal

An AI blizzard is headed our way [Yikes!]

Adam Satariano, Paul Mozur, Jacqueline Gu and Cade Metz, The Impending, Inescapable Deluge of A.I. NYTimes, July 30, 2026.

Boom....

From the American Midwest to the Persian Gulf, hundreds of major data centers now under construction will be turned on in the coming years. They are set to deliver an avalanche of computing power to develop and run A.I. that has no equal in the history of the technology industry, with breakthroughs that once felt revolutionary likely to become increasingly routine.

Behind each leap in A.I. are corresponding jumps in computing power. Today, there are about 20 million A.I. chips crammed into the data centers that underpin the technology’s growing abilities and usage worldwide, according to the research firm Epoch AI. [...]

In size and ambition, this moment compares to the building of the railroads in the 1800s, President Franklin D. Roosevelt’s New Deal in the 1930s, and the Manhattan Project to create an atomic weapon in the 1940s, technologists said.

“This is the largest scale infrastructure build-out in the history of humanity,” said Rob Wachen, a co-founder of the microchip firm Etched, which has raised more than $1 billion to meet the growing demand for A.I. components.

Peter DeSantis, who leads foundational A.I. models at Amazon — which provides computing power to the A.I. firms Anthropic, OpenAI and others — said the Seattle company has doubled its computing capacity since 2022 and would double it again by next year. “It’s hard to get your mind around the scale,” he said. [...]

Confidence in the Scaling Laws has led A.I. leaders to make ever bolder predictions. Dario Amodei, the chief executive of Anthropic, has said that if these laws hold for another year or two, A.I. will be able to perform huge amounts of white-collar work. Demis Hassabis, the head of Google’s A.I. lab DeepMind, wrote recently that A.I. could usher in “10x of the Industrial Revolution at 10x the speed.”

Bust?

Economists and investors have raised concerns that tech firms are spending faster than they can profit from A.I. Past infrastructure booms have been followed by downturns before the benefits of the technology were realized. The railroad boom in the 1800s, electrification in the 1920s and the dot-com bubble in the late 1990s were punctuated by economic recessions and a stock market crash as companies that overspent went out of business.

“Each time you’ve had a technological revolution, this kind of bubble bursting happened,” said Philippe Aghion, who won the Nobel in economic science in 2025 for research on innovation-driven economic growth. “A.I. is like the fourth industrial revolution and it has this aspect to it that generates a bubble.”

Moreover:

With more computing power coming online, geopolitical divisions are only set to widen.

The United States, home to about 5,500 data centers, about 10 times the next closest country, is far ahead of the rest of the world, including China. U.S. companies like Amazon, Google, Microsoft and Meta control about 80 percent of global computing power that drives A.I., according to Epoch AI. Google alone is believed to have four times as many A.I. chips as all of China’s companies, which are racing to catch up by developing new semiconductors and A.I. infrastructure of their own.

The race is on...

The article then goes on to discuss the huge increase in total chip count distributed over a growing collection of ever larger data centers under construction or proposed. We're in an AI arms race. The article then goes on to discuss the race with China, currently a fairly distant second (by an order of magnitude) to the US in chip count and gigawatts.

Mr. Nanos said the U.S. data center lead over China would likely grow over the next four to five years, before China’s domestic chips are produced at scale. After that, China should begin closing the gap.

“The advantage will run out,” he said.

The U.S.-China race threatens to leave the rest of the world behind. France, Germany and other nations are trying to encourage data center construction across the European Union, which has 5 percent of global A.I. computing power, according to a report by A.I. developers and policy experts in the region. Europe has been hampered by electricity and land access, permitting and financing.

Changes in the labor market:

“There’s going to be millions of jobs destroyed, millions of jobs created,” said Erik Brynjolfsson, an economist who is the director of Stanford University’s Digital Economy Lab. “That’s going to be very difficult. Even if new jobs are created, they’re not the same jobs.”

The last three paragraphs are about recursive self-improvement.

Wednesday, July 29, 2026

Logic and language in the brain

The abstract of the linked article:

Humans are endowed with a powerful capacity for inductive and deductive logical thought: we easily form generalizations based on a few examples and draw conclusions from known premises. Humans also arguably have the most sophisticated communication system in the animal kingdom: natural language allows us to express complex and structured meanings. Some have therefore argued for a tight relationship between complex thought and language, postulating that reasoning, including logical reasoning, relies on linguistic representations. We systematically investigated the relationship between logical reasoning and language using two complementary approaches. First, we used noninvasive brain imaging (fMRI) to examine neural activity as healthy adults engaged in logical reasoning tasks. And second, we behaviorally evaluated logical abilities in individuals with extensive lesions to the language brain areas and consequent severe linguistic impairment. Our findings reveal that the language brain network is not engaged during logical reasoning, and patients with severe aphasia exhibit intact performance on logic tasks. Instead, inductive reasoning recruits the domain-general multiple demand network implicated broadly in goal-directed behaviors, whereas deductive reasoning draws on brain regions that are distinct from both the language and the multiple demand networks. Together, these results indicate that linguistic representations are neither utilized nor required for inductive or deductive logical reasoning.

Why Adam Hunt has “flipped from being bullish to being bearish about AI.”

They do it all for us, Mickey D’s!

Recent Chinese Innovation

Lerner, Josh and Narain, Namrata and Papanikolaou, Dimitris and Seru, Amit and Xu, Zunda Winston, Chinese Sputnik Moments? (July 13, 2026). Available at SSRN: https://ssrn.com/abstract=7114818

Abstract: China's technological progress in recent decades has been viewed with admiration, alarm, and (in some cases) doubt. To better understand the Chinese innovation ecosystem, we compile a dataset of almost 14 million domestic Chinese patent publications. We focus on the subset of critical technologies identified by the U.S. Department of Defense. Several surprising patterns emerge from the data: Chinese patenting is strongly associated with other measures of innovative progress; patents are not concentrated in corporate giants such as Huawei; universities have played a key role in innovation, much greater than state-owned enterprises or government-owned facilities; and fewer than one in ten Chinese critical technology patents involves an inventor with U.S. experience or training. Finally, using four text-based measures of patent quality, we show that the rise of Chinese patenting in critical technologies has not been associated with a decline in quality relative to the U.S. awards.

H/t Tyler Cowen.

Three important agents [think about it]

Tuesday, July 28, 2026

The Vera Rubin Observatory in Chile has been discovering objects we hadn't even imagined existed

On the YouTube page:

Vera Rubin’s First Images JUST STOPPED THE WORLD! The Vera Rubin Observatory has already discovered objects that scientists never expected to find.

What if the most revolutionary telescope in history isn't looking deeper into space—but watching the universe change in real time? In this video, we explore the astonishing first discoveries from the Vera C. Rubin Observatory, including a 163,000-light-year stellar stream, an impossibly fast-spinning asteroid (2025 MN45), millions of newly detected celestial objects, and why astronomers believe Rubin is about to transform astronomy forever.

Unlike Hubble or the James Webb Space Telescope, Vera Rubin repeatedly scans the entire southern sky every few nights, creating a living timeline of the cosmos. That unique capability has already revealed hidden galactic structures, strange asteroid behavior, and a flood of discoveries that previous generations of telescopes completely missed.

You'll learn how Rubin's Legacy Survey of Space and Time (LSST) works, why it generates millions of alerts every night, what makes asteroid 2025 MN45 seemingly impossible according to current physics, and how the observatory is expected to map nearly 20 billion galaxies during its decade-long mission. Every image is adding new pieces to one of the biggest scientific puzzles of our time.

Could these discoveries change our understanding of dark matter, galaxy formation, planetary evolution, and even the future of our Solar System? The first images suggest we may only be witnessing the beginning.

Rounding the turn into Newport

Framing my discussion of The God Test, Part 1: Rorschach, reason, and whaling – [GT-3]

I’ve got to bite the bullet: I’m just going to have to go through a bunch of (preliminary) stuff before I can really engage with The God Test. My current target is to be in a position to publish a proper review of the book in 3 Quarks Daily for the week of August 9.

Rorschach Recap

I want start by recapping the Rorschach metaphor I introduced in the previous post, More on how I’m approaching The God Test – Rorschach! [GT-2]. What I like about it is that has a shape, there’s something there, but it’s not clear what. So we have little choice but to project onto it in order to (begin to) make sense of it.

First: It is a new kind of thing, an artifact we can converse with in an open-ended and natural way. The steam engine was the same kind of thing. It was an inanimate object that moved over the surface of the earth under its own power. Previously only animals (& humans as animals) had that power. So it becomes an iron horse. Just what are AIs? What’s their nature? That’s one thing.

Second: How it works is opaque. We know how to create large language models (LLMs), but we don’t know how they work. That’s new. We may not have understood the deep physics of the steam engine, but we certainly knew how they worked.

Third: We don’t know what they portend for the future. To some extent this is a function of the first two: How can we, how should we, interact. But it is also a function of the future, which is undetermined. We just don’t know.

Rhetorical force over reason

This is an argument I made in the first working paper I published after the release of ChatGPT in November of 2022: ChatGPT intimates a tantalizing future; its core LLM is organized on multiple levels; and it has broken the idea of thinking (February 6, 2023).

What do I mean by that, has broken the idea of thinking? Prior to ChatGPT it was obvious that humans could think and computers could not. [Yeah, I know, there’s Deep Blue defeating Kasparov in chess. That just changes the dates, not the argument.] The difference in performance was so obvious that the fact that we don’t really know how humans think wasn’t much of an issue. Now it is. Sure, we can still say that we can think and the AI’s can’t, but that’s just a line and without good explanations on both sides of the line, it seems a bit arbitrary, if not desperate.

I made a particular argument about Searles’ (in)famous Chinese Room thought experiment. I read it when it was first published in Brain and Behavioral Science in 1980. I wasn’t impressed. Why not? He didn’t say anything about any of the techniques used in AI or computational linguistics (CL). How could anyone possibly take that seriously?

He talked about intention, that’s how. Meaning requires intention and only living things can have intention, a remark he made at the end of the article. Without intention the most you get is syntax, but no meaning. Searle could get away with that because, in the first place, the concept of intention has a long history within philosophy – it has a subtle meaning, but that can wait for a later post – and so philosophers, his main audience, were comfortable with it. That’s one thing.

But there’s something more important, something that we can see only in retrospect, and that’s the simple fact computers very obviously could not translate from Chinese into English or into any other language. That difference carried tremendous weight. We don’t have a subtle behavioral difference between computers and humans that requires a subtle and sophisticated argument. To a first approximation, almost any argument would do. As far as I was concerned, “intention” was just a fancy word for something we don’t understand. But that’s not an argument anyone needs to take seriously. The behavioral distance is quite sufficient to carry the argument for those who insist that computers can’t and will never be able to think like humans.

Now the behavioral evidence has changed. Sure, differences remain, but the evidence is shifting. The old arguments remain and those who believed them still do so, but it’s getting harder. The need for explicit arguments grounded in explicit accounts of computers, and also brains, is growing.

Whaling and expertise

What’s an expert in machine learning and LLMs actually expert in? For some time now I’ve been arguing that investing in AI is like investing in a whaling venture where the captain and crew of the ship know all there is to know about the ship and how to handle it but know little or nothing about whales and their behavior and about navigating around the Cape Horn and in the South Pacific, where the whales live. What are the chances of that voyage being successful? Not very good.

The people who have created the current AI technology are like that captain and crew. The know how to sail the ship. But they don’t know much about language or cognition. They don’t actually know much about the human mind. Here my point is not about the fact that the models are opaque, but that human language and cognition are highly structured and they don’t believe that one needs to know (much of) anything about that not only to build AI but to make confident prediction about the future of AI.

Gary Marcus, Subbarao Kambhampati, and others have been consistently arguing that, yes, the current technology is remarkable, but we are going to have to adopt classical symbolic techniques if we are to fully develop the technology so that we have accurate and safe systems. Marcus is arguing from his knowledge of human language and cognition. As far as I can tell, Wright doesn’t take that seriously. I know that he had Marcus on his NonZero podcast, and that he lists Marcus in his acknowledgements, but that he doesn’t discuss Marcus’s ideas. I conclude that he doesn’t take that line of argument seriously.

That’s a mistake, but this is not the place to make my own arguments on this issue. My point is simply that expertise in AI is no generally construed to encompass knowledge of, expertise in, human cognition and language. I can’t see how that is going to work out well in the future.

[Note: If you’re curious about my views, on this subject, read the article linked in the first paragraph of this section. My views all over the place here at New Savanna, particularly around the work of the mathematician Miriam Yevick. Also, check out the experimental work I’ve done with LLMs.]

The odd destruction of books en masse by AI companies [Homo economicus on a binge]

Monday, July 27, 2026

Firetruck

Behavioral similarities in the way chatbots and oral poets perform

Kush R. Varshney, An Annotated Reading of ‘The Singer of Tales’ in the LLM Era, https://arxiv.org/html/2502.05148v1 Feb. 2025.

Abstract. The Parry-Lord oral-formulaic theory was a breakthrough in understanding how oral narrative poetry is learned, composed, and transmitted by illiterate bards. In this paper, we provide an annotated reading of the mechanism underlying this theory from the lens of large language models (LLMs) and generative artificial intelligence (AI). We point out the the similarities and differences between oral composition and LLM generation, and comment on the implications to society and AI policy.

Varshney develops his argument by interlacing passages from Albert Lord's The Singer of Tales with comments on LLMs. This is a very interesting way of reviewing your understanding of LLMs in relation to a specialized kind human language performance.

You might want to consider two of my blog posts:

GPT-3, the phrasal lexicon, Parry/Lord, and the Homeric epics, July 16, 2022.

In some ways, some contexts, LLMs may provide a useful model for human language, March 24, 2026.

In this more recent post I discuss empirical evidence about human memory for F.C. Bartlett's classic book, Remembering: A Study in Experimental and Social Psychology (1932), David C. Rubin, Memory in Oral Traditions: The Cognitive Psychology of Epic, Ballads, and Counting-out Rhymes (Oxford 1995).

The Impact of the Sewing Machine on Women

Philip Ager and Davide M. Coluccia, The Impact of the Sewing Machine on Women

Abstract: This paper provides novel evidence on how technological change shaped women’s labor market participation, fertility, and marriage in 19th-century Massachusetts. We distinguish between the sewing machine’s dual role as a manufacturing technology and as a household appliance. Using rich town-and individual-level longitudinal data, we show that this innovation induced divergent responses across the wealth distribution. Women from lower-wealth households increased labor supply, delaying marriage and reducing fertility. In contrast, for wealthier women, the sewing machine functioned as a domestic efficiency tool, enabling earlier family formation and greater civic engagement while reducing market work. Our findings demonstrate how household constraints and social norms mediate the effects of labor-saving technologies, suggesting that technological progress can reinforce inequality by influencing women’s economic and social roles.

H/t Tyler Cowen.

On Washington St. in Hoboken

What I did last week: aesthetics, economics, Rorschach analogy for AI, default images, and “leveling”

I did some satisfying work last week. Here’s a quick rundown. I’m listing the posts in the order I wrote them.

Visual Aesthetics

A case of visual aesthetics: Why is the monochrome image superior to the color image?

The issue, black & white vs. color, has been and I suppose remains central to photography, and I deal with it there, a bit. But that’s not what I’m doing here. This is about the conversion of a particular ChatGPT image from color to black & white. It was a fun post to assemble and to think about. I like the suite of images.

Rank 5 Economics?

Beyond Marginalism: What’s Next? [MR #12]

This is my last word – save for an introduction I’ll write in a week or three, who knows? – on the fourth and final chapter of Cowen’s monograph on marginalism. This is where he tosses up some examples of leading edge work in economics, noting that it’s drifting away from marginalism into complex high-dimensional models created through machine learning. His examples come from finance. The new models yield better predictions.

I focus on one model that has 360,000 parameters and end up making (speculative) sense out of what’s going on. I suggest that those parameters are picking up the effects of Keynes’ “animal spirits” as expressed in the gossip and stories of Schiller’s narrative economics. I further suggest that we can test this by comparing the output of a classical model with that from a high-parameter machine learning model. The divergence should be highest with those stocks otherwise identified as meme stocks.

The prospect of empirical investigation into animal spirits in asset pricing [the fate of marginalism]

Here I take my speculations about how to test these high parameter models and present them to Marge, the AI associated with Cowen’s book. Marge approves.

Rorschach test for AI

More on how I’m approaching The God Test – Rorschach! [GT-2]

I came up with the Rorschach blot as analogy for the kind of challenge AI presents to us, to our understanding of AI and of the future. The idea is that the blot does have a form, albeit a complex one that’s not very legible. Hence our commentary on it (that is, on AI) tells as much about us as about AI. I’ll be developing this further in a later post.

Prototype Image in ChatGPT

This is a new working paper that opens up a whole new line of investigation. This was a fun piece of work. Writing it up took way longer than actually generating the images.

A prototypical image in ChatGPT 5.6: An informal pilot study

Abstract: Previous work has found strong default preferences in stories generated by large language models from minimally specified prompts. To determine whether a similar effect appears in image generation, I asked ChatGPT 5.6 to create 17 images in separate chats using four prompts that specified either no subject matter or only a rendering medium: “Create an image,” “Create a drawing,” “Create a painting,” and “Create a water color painting.” Sixteen of the 17 outputs depicted closely related landscapes containing mountains, trees, sky, and water; the remaining image depicted a lighthouse. The images also shared a calm, picturesque mood and contained no human figures, although several included signs of human habitation. Because the internal prompt passed to the image generator was not available, the study cannot determine whether these defaults arise primarily in the language model, the image generator, or their interaction. The results are exploratory but suggest that severely underspecified image prompts may reveal stable default preferences in the integrated ChatGPT image-generation system.

The “leveling” of knowledge in the compressed form of LLMs

NYTimes: AI needs human supervision in order to complete an entire job.

This is something I’ve been thinking about off and on for a while, but this is my first explicit framing of the issue. The idea is that once ideas or set of ideas has been expressed in writing and those documents then consumed into an LLM, all ideas function the same within/through/for the model. In that post I’m comparing a study of using AI to perform routine office processes (from NYTimes) with the use of AI to perform a complex set of tasks in drug development, in effect, high-school level capability with Ph.D. level capability. They’re the same to the LLM.

I need to think about this some more. It seems to me what’s nowhere present in the LLM is the kind of procedural knowledge necessary to learn tasks at whatever level. That simply isn’t presented in the written products of that knowledge (not even in written procedures).

The Decline in the Transmission of Scientific Ideas

Enrico Berkes and Ruben Gaetani, The Decline in the Transmission of Scientific Ideas, NBER, July 2026.

Abstract: We document that the diffusion of new scientific ideas beyond their field of origin has declined substantially over the past four decades. This contraction is closely linked to increasing spe- cialization in scientific language: research that employs more technical terminology tends to be adopted less broadly. We develop a theory of scientific discovery in which the diffusion of new ideas depends on the degree to which potential adopters can understand and process them. When introducing their discoveries, scientists face a tradeoff between technical com- munication targeted at their immediate peers and more accessible language meant to reach broader audiences. As knowledge accumulates and research at the frontier builds on deeper layers of prior work, this tradeoff increasingly favors specialized language, limiting diffusion. Policy interventions that align scientists’ incentives can broaden adoption and increase the social value of scientific research.

H/t Tyler Cowen.

Sunday, July 26, 2026

ChatGPT draws Rorschach blots

Here they are, six of them:

I then discussed the blots with ChatGPT in terms suggested by me new working paper, A prototypical image in ChatGPT 5.6: An informal pilot study. First I present all six blots, full size, and then the discussion.

I needed a Rorschach inkblot for a post I was working on, so I had ChatGPT create one. I then decided to see what would happen if I asked for a number of inkblots, each in a different chat. I’ve uploaded the result to the project in a document entitled Rorschach Blots. I had no expectation about what would happen. Would it produce the same blot each time? Of different blots each time? As you can see, it produced different blots each time.

This suggests two further lines of investigation. 1) Generate 100 independent ink blots and run them through an image classifier. Would we get distinct clusters or not? 2) Use prompts suggesting specific things, but otherwise open-ended, e.g. “Draw a cat,” “Draw a tree,” “Draw a car,” “Draw a coat,” etc.

Looking at the Rorschach Blots document, the six outputs are plainly different, but they are not unconstrained in every respect. All six preserve a common structural grammar: bilateral symmetry, a strong vertical axis, black-and-gray ink on white, irregular feathered edges, and enough internal articulation to invite figurative interpretation. Within that shared format, however, the overall silhouettes vary substantially: some are tall and narrow, some compact and squat, some strongly winged, some mask-like, and some more diffuse. The contact sheet on page 1 makes that balance of sameness and difference especially clear.

That suggests that the system has not selected one prototypical blot, in the way it selected one broad landscape type. Instead, it seems to have selected a prototypical design space for “Rorschach blot,” within which it samples different instances. Your two proposed extensions are therefore well chosen.

1. One hundred independent inkblots

The clustering question is exactly the right one. A hundred examples might reveal:

  • one continuous cloud of variation around a single template;
  • several recurring morphological families;
  • or a few dominant families plus outliers.

Even by eye, the six examples hint at possible families: elongated “totemic” forms, broad winged forms, compact mask-like forms, and more open forms with substantial white interior space. But six is far too few to know whether those are genuine modes or just impressions.

I would not rely on a conventional object classifier, since it may simply label them all “inkblot,” “butterfly,” “mask,” or “bat.” A more informative analysis would use image embeddings and cluster those embeddings, perhaps supplemented by simple measurable features such as aspect ratio, occupied area, contour complexity, amount of internal white space, and vertical distribution of mass. The point would not be to decide what each blot “looks like,” but to determine the structure of variation across the set.

2. Open-ended object prompts

“Draw a cat,” “Draw a tree,” “Draw a car,” and so forth would probe something different and potentially more revealing: whether the system has a default realization for a named category.

The questions would include:

  • Does “Draw a cat” repeatedly produce the same breed, pose, viewpoint, or expression?
  • Does “Draw a tree” default to a solitary deciduous tree, a stylized botanical specimen, or a landscape scene?
  • Does “Draw a car” default to a modern sports car, a generic sedan, or a three-quarter advertising view?
  • Does “Draw a coat” produce the garment alone, on a mannequin, or worn by a person?

That experiment would occupy a useful middle ground between the wholly unspecified prompt—“Create a drawing”—and a detailed scene description. The subject is fixed, but almost everything about its realization remains open. It would therefore reveal category-level defaults in pose, style, composition, context, and affect.

The Rorschach result also suggests a distinction worth preserving in the paper’s future-work section:

Some prompts may elicit a single dominant subject prototype, while others may elicit a constrained family of forms rather than one recurring image.

That is the most interesting implication of the blot experiment. The landscape prompt collapsed onto a narrow semantic prototype. The inkblot prompt produced morphological diversity within a narrow formal convention. Those are different kinds of regularity, and comparing them could become a substantial research program.

Brittleness in chess-trained language models

Read the whole thread. 

NYTimes: AI needs human supervision in order to complete an entire job.

From the NYTimes article linked in the tweet:

We gave an A.I. tool full access to a laptop with pre-configured apps and sought to answer a simple question: Can artificial intelligence do an office job?

Some corporate executives seem to believe it can. More than 200 tech companies have cut roughly 120,000 jobs this year, according to Layoffs.fyi, an industry tracking site; Meta, Oracle and others have all recently made substantial cuts to their work forces, citing A.I. as the driving force; and after laying off about 1,100 employees, the chief executive of Cloudflare said recently that he expected A.I. to replace workers in middle management, finance and marketing.

tweIn our experiment, we deployed A.I. “agents” to act as office workers and found that they were capable of performing some of the tasks we assigned, but not all of them. The agents, which can act autonomously and make decisions based on detailed instructions, excelled at problems they could solve by writing computer programs. But they struggled with understanding the nuances of human language and at navigating user interfaces like the Chrome web browser.

The article then has a series of nice quasi-interactive displays illustrating agent performance on three tasks. The displays include screen shots of various messages and documents.

About the tasks:

This task, and the others we assigned to the A.I., were adapted from papers and benchmarking tools published recently by researchers at Carnegie Mellon University and OpenAI. The researchers designed the benchmarks to test the performance of various models — like OpenAI’s GPT, Google’s Gemini and Anthropic’s Claude — in real-world environments, and compare them with one another.

General conclusion:

The results of our experiment roughly matched what researchers and companies have found as they have tested and used artificial intelligence tools. Scale AI, an A.I. training company, recently tested agents on real freelance projects, and the best-scoring model produced client-ready work only about 16 percent of the time.

While A.I. can excel regularly at complex tasks, it can be unreliable when put in charge of an entire job. It can certainly add value to certain areas of the work force, but for now, A.I. still needs a human boss.

* * * * *

Comment: Around the corner my colleague, Ash Jogalekar, has tweets like this one:

So here's a great example of where we are with agentic AI: Instead of just being an assistant, it's behaving more like a collaborator and creative scientist.

In a recent project, I gave the system a molecular design problem typical of the problems we encounter in chemistry. Two similar molecules were giving very different results.

He then runs through an account of what his AI collaborator did, concluding:

I think we have crossed the Rubicon. Agentic AI now no longer just processes tasks and automates workflows blindingly fast, but it can generate hypotheses, test them, test counter-hypotheses and go back and forth and course-correct if necessary, all with minimal to no human intervention. It's now embodying the general scientific method.

[I've copied another one of Ash's tweets to this post, A scientist reflects on what AI has done for him.]

What’s interesting to me, and very revealing, is that a complex set of tasks in scientific investigation seems to be on a level with routine office tasks, as though one were no more complex than the other. But humans require years of college education in order to perform the former while the latter requires no more than a high school education, if that. It seems that once they’ve been learned and compiled, all tasks or sets of tasks are on the same “level” in the brain. The educational prerequisites required to do such tasks for the first time or three get “compressed out” through repetition. Since AIs are trained on written records of what humans have said and done, they don’t have to go through the ordinary learning process. The compression has already taken place and is present in the documents on which they are trained.

Saturday, July 25, 2026

Standard-issue doom scenarios were invented before LLMs and are made obsolete by them.

A prototypical image in ChatGPT 5.6: An informal pilot study

A new working paper. Title above, links, abstract, table of contents, and introduction below.

Academia.edu: https://www.academia.edu/170708104/ChatGPT_has_a_prototypical_image_An_informal_pilot_study_A_Working_Paper
ResearchGate: https://www.researchgate.net/publication/410824948_A_prototypical_image_in_ChatGPT_56_An_informal_pilot_study

Abstract: Previous work has found strong default preferences in stories generated by large language models from minimally specified prompts. To determine whether a similar effect appears in image generation, I asked ChatGPT 5.6 to create 17 images in separate chats using four prompts that specified either no subject matter or only a rendering medium: “Create an image,” “Create a drawing,” “Create a painting,” and “Create a water color painting.” Sixteen of the 17 outputs depicted closely related landscapes containing mountains, trees, sky, and water; the remaining image depicted a lighthouse. The images also shared a calm, picturesque mood and contained no human figures, although several included signs of human habitation. Because the internal prompt passed to the image generator was not available, the study cannot determine whether these defaults arise primarily in the language model, the image generator, or their interaction. The results are exploratory but suggest that severely underspecified image prompts may reveal stable default preferences in the integrated ChatGPT image-generation system.

Contents

Introduction: Default preferences in LLMs 3
Default preferences in story generation 3
Image generation, method 5
Results 6
The case of Bob Ross 11
Final remarks and future work 13
Default images: Five independent trials 16
Drawings: Six independent trials 21
Paintings: Six independent trials 27

Introduction: Default preferences in LLMs

About two and a half years ago I published an informal pilot study, ChatGPT tells 20 versions of its prototypical story, with a short note on method. I discovered that when given a simple one word prompt, “story,” that places no restrictions on the nature of the story to be generated, ChatGPT tended to generate the same story each time, roughly the same general plot set in a fairy tale world. More recently Sil Hamilton and David Mimno studied 20,000 stories generated on four different platforms and discovered that words, including character names, occurred in 88% of the stories.

Given this background, I wondered: Does image generation exhibit the same effect? Once it became possible to generate images from ChatGPT I had used it to generate many different kinds of images, some from simple prompts, others from long, often very long, prompts, and still others from sample photographs. About a week ago I decided to see what kind of images ChatGPT would generate when given a prompt that made no specifications about subject matter.

This is an informal pilot study. I began on an impulse, with no specific method or goal in mind. I just wanted to see if there was anything there. If so, what do we need to do to conduct a more rigorous study?

* * * * *

I begin by presenting the basic results on default preferences in story generation as background. Then I present the methods and results of my image study. After that I discuss the work of Bob Ross, an artist who had a popular TV show in which he showed viewers how to paint images similar to those ChatGPT generated in this study. I conclude my discussion with some final remarks and suggestions for future work. Last, we have the images themselves. Note that I refer to the recurring image type as a prototype produced under minimally specified default conditions.

Lazy days

Friday, July 24, 2026

Galloway: “You’re about to see the clip economy take over our TVs”

From the transcript, Galloway speaking:

40:33 I think China quite frankly I hate to say it at a certain level has one AI. We just haven’t woken up to that yet. I would agree.

40:39 And then two you’re about to see the clip economy take over our TVs. I think the biggest show on television

40:47 uh from Netflix or someone else is going to be a 60-minute compilation of two and three minute videos similar to what you

40:55 see on reals or Tik Tok. I think that’s about to wash over the traditional streamers you’re going to see.

Some years ago I had the idea of making a feature film entirely out of previews, carefully conceptualized and stitched together. The feature would be organized around an ensemble of, say, a half dozen to ten actors who keep on recurring in the previous in various combinations. The previews would span, say, ten or 20 years of fictional time, with the actors aging over the course of the previews. The whole thing would be constructed so that we can infer what’s going on in their lives from what we see in the previews, even though the previews would be based on a variety of genres: romantic comedies, science fiction, drama, horror, fantasy, period pieces, farce, and so forth.

More on how I’m approaching The God Test – Rorschach! [GT-2]

I’m still trying to figure out how to approach Robert Wright’s The God Test.

How LLMs work

In my previous post – How will I handle The God Test? [GT-1] – I expressed misgivings about how Wright explains the technology. Those misgivings haven’t disappeared. However, Bert Idem has published a useful review at Finite Ape in which he addresses some of those issues in detail. Specifically:

Now, about the history of AI, the story he tells is actually great and it is certainly more than what most non-technical people know about LLMs. However, there are three places where I think the framing goes wrong or at least leaves out context that matters:

  • LLMs did not discover the meanings of words on their own by accident. They were designed on top of ideas from older models that were specifically trained to learn the meanings of words.
  • Similarly, computer vision models didn’t find out how to “view” an image like we do. Instead, the classical CNN models were heavily inspired by biological vision itself.
  • LLM weight training is simply gradient-based optimization and the process has nothing to do with evolution. Of course, we can make a parallel between any kind of change and evolution but then, in that sense, everything evolves and it is not useful to talk about evolution.

I agree with Idem on those three issues, not so sure about the history part. While I may return to some of these issues later on, this will serve as a place holder.

A Rorschach test

There’s something else going on, but I’m not quite sure how to conceptualize it. It seems to me that AI is functioning something like a Rorschach test which, as you may know, is a psychological instrument intended to elicit (potentially) revealing responses from a person. It’s a projective test.

A person is shown a series of ink blot images, like this one (generated by ChatGPT):

They are asked what that they see in the image, what it means to them. Since the image is, though not formless, its form is not that of any specific animal, vegetable, mineral, person, or anything else. It’s just a blot. Whatever the person says about the blot, however they interpret it, that must reveal something about them. Why? Because whatever they see in the blot, isn’t really there.

Broadly and crudely speaking, AI has become something of a cultural Rorschach test.

Understanding computers & LLMs

Until ChatGPT was released in late November of 2022, most people knew very little to nothing about AI. Oh, they may have seen “intelligent” computers and robots in science fiction movies, but that’s science fiction and only tangentially related to AI considered as a line of research dating back to the 1950s. Many people would have heard about IBM’s Deep Blue beating Gary Kasparov in chess in 1997 and then, in 2011, when IBM’s Watson beat Ken Jennings and Brad Rutter in Jeopardy. Those were real AI systems, standing on research extending back decades, but as far as most people were concerned, they were one-off PR stunts. Just how they worked, who cares? They’re computers, and computers are magic, no?

As far as most of us are concerned, computers are magic. Somewhere “out there” someone knows how these things work, but we don’t need to know any of that. It’s complicated, but computers do what they’re programmed to do, no? Yes, but not LLMs.

And that’s the tricky part. LLMs, large language models, aren’t like other computer systems. They aren’t programmed in the way that word processors, photo editors, or phones are programmed. LLMs aren’t programmed at all, not in the ordinary sense of programming – something I may or may not get into in a later post. As far as most users are concerned, how ChatGPT, or Claude, or Gemini work, that’s no more interesting than how a word processor works. It just does. It’s more magic.

But if you have a strong philosophical streak, if you are interested in the mind, in technology, in the technology in the future, then you may not be content with writing LLMs off as just another kind of magic. You want to know what’s going on inside, 1) because you want to know (curiosity), and 2) because you want to know how the technology is going to develop in the future (engagement). Now things get interesting? Why? Because even the people who have created the technology don’t know how it works.

Oh, they know how the transformer program works. That’s the program that creates the language model. It creates the model by performing a (certain kind of) statistical analysis of a huge body of texts, effectively the entire internet. When a person prompts the model with some statement, the model responds by a statement of its own. No one know just how the model does that. That’s a mystery, a deep black hole in the technology ecosystem.

AI as a Rorschach test

If you aren’t content to believe in magic, then you have to come up with something to fill that black hole in your, in our, understanding. This is where the Rorschach aspect of AI reveals itself. To a first approximation, what each of us uses to paper over that black hole has as much to do with ourselves as with AI.

Why do I say, “To a first approximation”? It’s a rhetorical device to get things started. It puts us all in the same boat, despite our different backgrounds. However, whatever LLMs are, they are not magic. It is possible, in principle, to construct a technical account of what they’re up to, but no one knows how to do that, yet. Not even the people in the AI labs who create these beasts.

Those of us who are trying to figure out how LLMs work have widely varying backgrounds. In particular, we have widely varied technical backgrounds and we bring those backgrounds to bear when we think about what LLMs are doing. Those backgrounds influence how we interpret the AI-blot. Wright is a journalist with a wide range of interests, including politics, international affairs, evolutionary psychology, cultural evolution, and Buddhism. As far as I can tell there isn’t much there that’s directly relevant to understanding the mechanisms of LLMs, but he’s done a lot of reading and talked with a lot of experts to fill in the gaps.

My background is quite different. While I happen to know quite a bit about cultural evolution, cognitive psychology, neuroscience, and various other things, my background in computational semantics puts me much closer to LLMs than Wright’s knowledge of evolutionary psychology puts him. Still, like him, I’ve done a lot of reading and talked with experts. In particular, I’ve been collaborating with Ramesh Viswanathan for the last three years. He’s an expert in machine vision Goethe University Frankfurt. He’s got a background in mathematics and AI that I don’t have. Still, there are things he doesn’t know, things he’s trying to figure out. 

We are all making stuff up.

To some extent, then, AI is a Rorschach test about how beliefs about the human mind, and human nature. When we try to figure out how the LLM is working we’re also, if only implicitly, trying to figure out how we work, internally, as well. The whole discourse about AL alignment is as much a discourse about us as it is about AI. 

The Future

And even if we knew much more about how LLMs work internally we still wouldn’t know how the technology will develop in the future. We? You, me, Robert Wright, Ramesh Viswanathan, Gary Marcus, Tyler Cowen, Geoffrey Hinton, Sam Altman, Dario Amodei, Nick Bostrom, Eliezer Yudkowsky, all of us who are trying to figure it out. We don’t know what will happen. That’s where we’re projecting like mad. We’re hallucinating, to borrow a term from AI-speak. 

Thus AI is also a Rorschach test for our visions of the future. When we imagine the future of AI, we’re also imagining our future. Like the two sides of a coin, the two cannot be separated. 

The tricky part, the important part, is that the future development of AI is not predestined. It depends on the choices we make, now and in the near future. We can easily and often do imagine things that will not be possible because that’s just not how the world works. But the laws of how the world works are open to a wide range of possibilities. The boundary between the possible and the impossible is fuzzy at best.

Where, and how, does Wright draw that boundary? Perhaps that’s what I’ll be trying to figure out.

More later.