NEW SAVANNA
“You won't get a wild heroic ride to heaven on pretty little sounds.”– George Ives
Tuesday, July 28, 2026
Framing my discussion of The God Test, Part 1: Rorschach, reason, and whaling – [GT-3]
I’ve got to bite the bullet: I’m just going to have to go through a bunch of (preliminary) stuff before I can really engage with The God Test. My current target is to be in a position to publish a proper review of the book in 3 Quarks Daily for the week of August 9.
Rorschach Recap
I want start by recapping the Rorschach metaphor I introduced in the previous post, More on how I’m approaching The God Test – Rorschach! [GT-2]. What I like about it is that has a shape, there’s something there, but it’s not clear what. So we have little choice but to project onto it in order to (begin to) make sense of it.
First: It is a new kind of thing, an artifact we can converse with in an open-ended and natural way. The steam engine was the same kind of thing. It was an inanimate object that moved over the surface of the earth under its own power. Previously only animals (& humans as animals) had that power. So it becomes an iron horse. Just what are AIs? What’s their nature? That’s one thing.
Second: How it works is opaque. We know how to create large language models (LLMs), but we don’t know how they work. That’s new. We may not have understood the deep physics of the steam engine, but we certainly knew how they worked.
Third: We don’t know what they portend for the future. To some extent this is a function of the first two: How can we, how should we, interact. But it is also a function of the future, which is undetermined. We just don’t know.
Rhetorical force over reason
This is an argument I made in the first working paper I published after the release of ChatGPT in November of 2022: ChatGPT intimates a tantalizing future; its core LLM is organized on multiple levels; and it has broken the idea of thinking (February 6, 2023).
What do I mean by that, has broken the idea of thinking? Prior to ChatGPT it was obvious that humans could think and computers could not. [Yeah, I know, there’s Deep Blue defeating Kasparov in chess. That just changes the dates, not the argument.] The difference in performance was so obvious that the fact that we don’t really know how humans think wasn’t much of an issue. Now it is. Sure, we can still say that we can think and the AI’s can’t, but that’s just a line and without good explanations on both sides of the line, it seems a bit arbitrary, if not desperate.
I made a particular argument about Searles’ (in)famous Chinese Room thought experiment. I read it when it was first published in Brain and Behavioral Science in 1980. I wasn’t impressed. Why not? He didn’t say anything about any of the techniques used in AI or computational linguistics (CL). How could anyone possibly take that seriously?
He talked about intention, that’s how. Meaning requires intention and only living things can have intention, a remark he made at the end of the article. Without intention the most you get is syntax, but no meaning. Searle could get away with that because, in the first place, the concept of intention has a long history within philosophy – it has a subtle meaning, but that can wait for a later post – and so philosophers, his main audience, were comfortable with it. That’s one thing.
But there’s something more important, something that we can see only in retrospect, and that’s the simple fact computers very obviously could not translate from Chinese into English or into any other language. That difference carried tremendous weight. We don’t have a subtle behavioral difference between computers and humans that requires a subtle and sophisticated argument. To a first approximation, almost any argument would do. As far as I was concerned, “intention” was just a fancy word for something we don’t understand. But that’s not an argument anyone needs to take seriously. The behavioral distance is quite sufficient to carry the argument for those who insist that computers can’t and will never be able to think like humans.
Now the behavioral evidence has changed. Sure, differences remain, but the evidence is shifting. The old arguments remain and those who believed them still do so, but it’s getting harder. The need for explicit arguments grounded in explicit accounts of computers, and also brains, is growing.
Whaling and expertise
What’s an expert in machine learning and LLMs actually expert in? For some time now I’ve been arguing that investing in AI is like investing in a whaling venture where the captain and crew of the ship know all there is to know about the ship and how to handle it but know little or nothing about whales and their behavior and about navigating around the Cape Horn and in the South Pacific, where the whales live. What are the chances of that voyage being successful? Not very good.
The people who have created the current AI technology are like that captain and crew. The know how to sail the ship. But they don’t know much about language or cognition. They don’t actually know much about the human mind. Here my point is not about the fact that the models are opaque, but that human language and cognition are highly structured and they don’t believe that one needs to know (much of) anything about that not only to build AI but to make confident prediction about the future of AI.
Gary Marcus, Subbarao Kambhampati, and others have been consistently arguing that, yes, the current technology is remarkable, but we are going to have to adopt classical symbolic techniques if we are to fully develop the technology so that we have accurate and safe systems. Marcus is arguing from his knowledge of human language and cognition. As far as I can tell, Wright doesn’t take that seriously. I know that he had Marcus on his NonZero podcast, and that he lists Marcus in his acknowledgements, but that he doesn’t discuss Marcus’s ideas. I conclude that he doesn’t take that line of argument seriously.
That’s a mistake, but this is not the place to make my own arguments on this issue. My point is simply that expertise in AI is no generally construed to encompass knowledge of, expertise in, human cognition and language. I can’t see how that is going to work out well in the future.
[Note: If you’re curious about my views, on this subject, read the article linked in the first paragraph of this section. My views all over the place here at New Savanna, particularly around the work of the mathematician Miriam Yevick. Also, check out the experimental work I’ve done with LLMs.]
The odd destruction of books en masse by AI companies [Homo economicus on a binge]
I'm a second-generation bookseller. My family runs Houston's largest used & rare bookstore and I'm building an AI tool for used bookstores. We got hit by exactly these orders, including a single order for 70 obscure books that made us pause online sales entirely. So I dug in. I… https://t.co/IK7OE8f1DG
— Charlie D. Becker (@charliedbecker) July 28, 2026
Monday, July 27, 2026
Behavioral similarities in the way chatbots and oral poets perform
Kush R. Varshney, An Annotated Reading of ‘The Singer of Tales’ in the LLM Era, https://arxiv.org/html/2502.05148v1 Feb. 2025.
Abstract. The Parry-Lord oral-formulaic theory was a breakthrough in understanding how oral narrative poetry is learned, composed, and transmitted by illiterate bards. In this paper, we provide an annotated reading of the mechanism underlying this theory from the lens of large language models (LLMs) and generative artificial intelligence (AI). We point out the the similarities and differences between oral composition and LLM generation, and comment on the implications to society and AI policy.
Varshney develops his argument by interlacing passages from Albert Lord's The Singer of Tales with comments on LLMs. This is a very interesting way of reviewing your understanding of LLMs in relation to a specialized kind human language performance.
You might want to consider two of my blog posts:
GPT-3, the phrasal lexicon, Parry/Lord, and the Homeric epics, July 16, 2022.
In some ways, some contexts, LLMs may provide a useful model for human language, March 24, 2026.
In this more recent post I discuss empirical evidence about human memory for F.C. Bartlett's classic book, Remembering: A Study in Experimental and Social Psychology (1932), David C. Rubin, Memory in Oral Traditions: The Cognitive Psychology of Epic, Ballads, and Counting-out Rhymes (Oxford 1995).
The Impact of the Sewing Machine on Women
Philip Ager and Davide M. Coluccia, The Impact of the Sewing Machine on Women
Abstract: This paper provides novel evidence on how technological change shaped women’s labor market participation, fertility, and marriage in 19th-century Massachusetts. We distinguish between the sewing machine’s dual role as a manufacturing technology and as a household appliance. Using rich town-and individual-level longitudinal data, we show that this innovation induced divergent responses across the wealth distribution. Women from lower-wealth households increased labor supply, delaying marriage and reducing fertility. In contrast, for wealthier women, the sewing machine functioned as a domestic efficiency tool, enabling earlier family formation and greater civic engagement while reducing market work. Our findings demonstrate how household constraints and social norms mediate the effects of labor-saving technologies, suggesting that technological progress can reinforce inequality by influencing women’s economic and social roles.
H/t Tyler Cowen.
What I did last week: aesthetics, economics, Rorschach analogy for AI, default images, and “leveling”
I did some satisfying work last week. Here’s a quick rundown. I’m listing the posts in the order I wrote them.
Visual Aesthetics
A case of visual aesthetics: Why is the monochrome image superior to the color image?
The issue, black & white vs. color, has been and I suppose remains central to photography, and I deal with it there, a bit. But that’s not what I’m doing here. This is about the conversion of a particular ChatGPT image from color to black & white. It was a fun post to assemble and to think about. I like the suite of images.
Rank 5 Economics?
Beyond Marginalism: What’s Next? [MR #12]
This is my last word – save for an introduction I’ll write in a week or three, who knows? – on the fourth and final chapter of Cowen’s monograph on marginalism. This is where he tosses up some examples of leading edge work in economics, noting that it’s drifting away from marginalism into complex high-dimensional models created through machine learning. His examples come from finance. The new models yield better predictions.
I focus on one model that has 360,000 parameters and end up making (speculative) sense out of what’s going on. I suggest that those parameters are picking up the effects of Keynes’ “animal spirits” as expressed in the gossip and stories of Schiller’s narrative economics. I further suggest that we can test this by comparing the output of a classical model with that from a high-parameter machine learning model. The divergence should be highest with those stocks otherwise identified as meme stocks.
Here I take my speculations about how to test these high parameter models and present them to Marge, the AI associated with Cowen’s book. Marge approves.
Rorschach test for AI
More on how I’m approaching The God Test – Rorschach! [GT-2]
I came up with the Rorschach blot as analogy for the kind of challenge AI presents to us, to our understanding of AI and of the future. The idea is that the blot does have a form, albeit a complex one that’s not very legible. Hence our commentary on it (that is, on AI) tells as much about us as about AI. I’ll be developing this further in a later post.
Prototype Image in ChatGPT
This is a new working paper that opens up a whole new line of investigation. This was a fun piece of work. Writing it up took way longer than actually generating the images.
A prototypical image in ChatGPT 5.6: An informal pilot study
Abstract: Previous work has found strong default preferences in stories generated by large language models from minimally specified prompts. To determine whether a similar effect appears in image generation, I asked ChatGPT 5.6 to create 17 images in separate chats using four prompts that specified either no subject matter or only a rendering medium: “Create an image,” “Create a drawing,” “Create a painting,” and “Create a water color painting.” Sixteen of the 17 outputs depicted closely related landscapes containing mountains, trees, sky, and water; the remaining image depicted a lighthouse. The images also shared a calm, picturesque mood and contained no human figures, although several included signs of human habitation. Because the internal prompt passed to the image generator was not available, the study cannot determine whether these defaults arise primarily in the language model, the image generator, or their interaction. The results are exploratory but suggest that severely underspecified image prompts may reveal stable default preferences in the integrated ChatGPT image-generation system.
The “leveling” of knowledge in the compressed form of LLMs
NYTimes: AI needs human supervision in order to complete an entire job.
This is something I’ve been thinking about off and on for a while, but this is my first explicit framing of the issue. The idea is that once ideas or set of ideas has been expressed in writing and those documents then consumed into an LLM, all ideas function the same within/through/for the model. In that post I’m comparing a study of using AI to perform routine office processes (from NYTimes) with the use of AI to perform a complex set of tasks in drug development, in effect, high-school level capability with Ph.D. level capability. They’re the same to the LLM.
I need to think about this some more. It seems to me what’s nowhere present in the LLM is the kind of procedural knowledge necessary to learn tasks at whatever level. That simply isn’t presented in the written products of that knowledge (not even in written procedures).
The Decline in the Transmission of Scientific Ideas
Enrico Berkes and Ruben Gaetani, The Decline in the Transmission of Scientific Ideas, NBER, July 2026.
Abstract: We document that the diffusion of new scientific ideas beyond their field of origin has declined substantially over the past four decades. This contraction is closely linked to increasing spe- cialization in scientific language: research that employs more technical terminology tends to be adopted less broadly. We develop a theory of scientific discovery in which the diffusion of new ideas depends on the degree to which potential adopters can understand and process them. When introducing their discoveries, scientists face a tradeoff between technical com- munication targeted at their immediate peers and more accessible language meant to reach broader audiences. As knowledge accumulates and research at the frontier builds on deeper layers of prior work, this tradeoff increasingly favors specialized language, limiting diffusion. Policy interventions that align scientists’ incentives can broaden adoption and increase the social value of scientific research.
H/t Tyler Cowen.
Sunday, July 26, 2026
ChatGPT draws Rorschach blots
Here they are, six of them:
I then discussed the blots with ChatGPT in terms suggested by me new working paper, A prototypical image in ChatGPT 5.6: An informal pilot study. First I present all six blots, full size, and then the discussion.
I needed a Rorschach inkblot for a post I was working on, so I had ChatGPT create one. I then decided to see what would happen if I asked for a number of inkblots, each in a different chat. I’ve uploaded the result to the project in a document entitled Rorschach Blots. I had no expectation about what would happen. Would it produce the same blot each time? Of different blots each time? As you can see, it produced different blots each time.
This suggests two further lines of investigation. 1) Generate 100 independent ink blots and run them through an image classifier. Would we get distinct clusters or not? 2) Use prompts suggesting specific things, but otherwise open-ended, e.g. “Draw a cat,” “Draw a tree,” “Draw a car,” “Draw a coat,” etc.
Looking at the Rorschach Blots document, the six outputs are plainly different, but they are not unconstrained in every respect. All six preserve a common structural grammar: bilateral symmetry, a strong vertical axis, black-and-gray ink on white, irregular feathered edges, and enough internal articulation to invite figurative interpretation. Within that shared format, however, the overall silhouettes vary substantially: some are tall and narrow, some compact and squat, some strongly winged, some mask-like, and some more diffuse. The contact sheet on page 1 makes that balance of sameness and difference especially clear.
That suggests that the system has not selected one prototypical blot, in the way it selected one broad landscape type. Instead, it seems to have selected a prototypical design space for “Rorschach blot,” within which it samples different instances. Your two proposed extensions are therefore well chosen.
1. One hundred independent inkblots
The clustering question is exactly the right one. A hundred examples might reveal:
- one continuous cloud of variation around a single template;
- several recurring morphological families;
- or a few dominant families plus outliers.
Even by eye, the six examples hint at possible families: elongated “totemic” forms, broad winged forms, compact mask-like forms, and more open forms with substantial white interior space. But six is far too few to know whether those are genuine modes or just impressions.
I would not rely on a conventional object classifier, since it may simply label them all “inkblot,” “butterfly,” “mask,” or “bat.” A more informative analysis would use image embeddings and cluster those embeddings, perhaps supplemented by simple measurable features such as aspect ratio, occupied area, contour complexity, amount of internal white space, and vertical distribution of mass. The point would not be to decide what each blot “looks like,” but to determine the structure of variation across the set.
2. Open-ended object prompts
“Draw a cat,” “Draw a tree,” “Draw a car,” and so forth would probe something different and potentially more revealing: whether the system has a default realization for a named category.
The questions would include:
- Does “Draw a cat” repeatedly produce the same breed, pose, viewpoint, or expression?
- Does “Draw a tree” default to a solitary deciduous tree, a stylized botanical specimen, or a landscape scene?
- Does “Draw a car” default to a modern sports car, a generic sedan, or a three-quarter advertising view?
- Does “Draw a coat” produce the garment alone, on a mannequin, or worn by a person?
That experiment would occupy a useful middle ground between the wholly unspecified prompt—“Create a drawing”—and a detailed scene description. The subject is fixed, but almost everything about its realization remains open. It would therefore reveal category-level defaults in pose, style, composition, context, and affect.
The Rorschach result also suggests a distinction worth preserving in the paper’s future-work section:
Some prompts may elicit a single dominant subject prototype, while others may elicit a constrained family of forms rather than one recurring image.
That is the most interesting implication of the blot experiment. The landscape prompt collapsed onto a narrow semantic prototype. The inkblot prompt produced morphological diversity within a narrow formal convention. Those are different kinds of regularity, and comparing them could become a substantial research program.
Brittleness in chess-trained language models
"We examine several claims made in existing literature regarding chess-trained language models and assert that their
— Chomba Bupe (@ChombaBupe) July 25, 2026
impressive benchmark performance is largely explained by pattern-matching"
Vision language models (VLM) are even worse at chess. pic.twitter.com/egzDDsnofC
Read the whole thread.
NYTimes: AI needs human supervision in order to complete an entire job.
“While A.I. can excel regularly at complex tasks, it can be unreliable when put in charge of an entire job. It can certainly add value to certain areas of the work force, but for now, A.I. still needs a human boss.”
— Gary Marcus (@GaryMarcus) July 25, 2026
When AGI comes, we will stop reading stuff like that. But that…
From the NYTimes article linked in the tweet:
We gave an A.I. tool full access to a laptop with pre-configured apps and sought to answer a simple question: Can artificial intelligence do an office job?
Some corporate executives seem to believe it can. More than 200 tech companies have cut roughly 120,000 jobs this year, according to Layoffs.fyi, an industry tracking site; Meta, Oracle and others have all recently made substantial cuts to their work forces, citing A.I. as the driving force; and after laying off about 1,100 employees, the chief executive of Cloudflare said recently that he expected A.I. to replace workers in middle management, finance and marketing.
tweIn our experiment, we deployed A.I. “agents” to act as office workers and found that they were capable of performing some of the tasks we assigned, but not all of them. The agents, which can act autonomously and make decisions based on detailed instructions, excelled at problems they could solve by writing computer programs. But they struggled with understanding the nuances of human language and at navigating user interfaces like the Chrome web browser.
The article then has a series of nice quasi-interactive displays illustrating agent performance on three tasks. The displays include screen shots of various messages and documents.
About the tasks:
This task, and the others we assigned to the A.I., were adapted from papers and benchmarking tools published recently by researchers at Carnegie Mellon University and OpenAI. The researchers designed the benchmarks to test the performance of various models — like OpenAI’s GPT, Google’s Gemini and Anthropic’s Claude — in real-world environments, and compare them with one another.
General conclusion:
The results of our experiment roughly matched what researchers and companies have found as they have tested and used artificial intelligence tools. Scale AI, an A.I. training company, recently tested agents on real freelance projects, and the best-scoring model produced client-ready work only about 16 percent of the time.
While A.I. can excel regularly at complex tasks, it can be unreliable when put in charge of an entire job. It can certainly add value to certain areas of the work force, but for now, A.I. still needs a human boss.
* * * * *
Comment: Around the corner my colleague, Ash Jogalekar, has tweets like this one:
So here's a great example of where we are with agentic AI: Instead of just being an assistant, it's behaving more like a collaborator and creative scientist.
In a recent project, I gave the system a molecular design problem typical of the problems we encounter in chemistry. Two similar molecules were giving very different results.
He then runs through an account of what his AI collaborator did, concluding:
I think we have crossed the Rubicon. Agentic AI now no longer just processes tasks and automates workflows blindingly fast, but it can generate hypotheses, test them, test counter-hypotheses and go back and forth and course-correct if necessary, all with minimal to no human intervention. It's now embodying the general scientific method.
[I've copied another one of Ash's tweets to this post, A scientist reflects on what AI has done for him.]
What’s interesting to me, and very revealing, is that a complex set of tasks in scientific investigation seems to be on a level with routine office tasks, as though one were no more complex than the other. But humans require years of college education in order to perform the former while the latter requires no more than a high school education, if that. It seems that once they’ve been learned and compiled, all tasks or sets of tasks are on the same “level” in the brain. The educational prerequisites required to do such tasks for the first time or three get “compressed out” through repetition. Since AIs are trained on written records of what humans have said and done, they don’t have to go through the ordinary learning process. The compression has already taken place and is present in the documents on which they are trained.
Saturday, July 25, 2026
Standard-issue doom scenarios were invented before LLMs and are made obsolete by them.
I’m going to take a crack at explaining this just a little, because it’s worth putting out there.
— Jon Stokes (@jon_stokes) July 24, 2026
The paperclip maximizer + related AI doom scenarios were mainly developed in a time when “AI” did not reduce to Large Language Models. The term was a lot wider and inherited a lot… https://t.co/c11tJibynH
A prototypical image in ChatGPT 5.6: An informal pilot study
A new working paper. Title above, links, abstract, table of contents, and introduction below.
Academia.edu: https://www.academia.edu/170708104/ChatGPT_has_a_prototypical_image_An_informal_pilot_study_A_Working_Paper
ResearchGate: https://www.researchgate.net/publication/410824948_A_prototypical_image_in_ChatGPT_56_An_informal_pilot_study
Abstract: Previous work has found strong default preferences in stories generated by large language models from minimally specified prompts. To determine whether a similar effect appears in image generation, I asked ChatGPT 5.6 to create 17 images in separate chats using four prompts that specified either no subject matter or only a rendering medium: “Create an image,” “Create a drawing,” “Create a painting,” and “Create a water color painting.” Sixteen of the 17 outputs depicted closely related landscapes containing mountains, trees, sky, and water; the remaining image depicted a lighthouse. The images also shared a calm, picturesque mood and contained no human figures, although several included signs of human habitation. Because the internal prompt passed to the image generator was not available, the study cannot determine whether these defaults arise primarily in the language model, the image generator, or their interaction. The results are exploratory but suggest that severely underspecified image prompts may reveal stable default preferences in the integrated ChatGPT image-generation system.
Contents
Introduction: Default preferences in LLMs 3
Default preferences in story generation 3
Image generation, method 5
Results 6
The case of Bob Ross 11
Final remarks and future work 13
Default images: Five independent trials 16
Drawings: Six independent trials 21
Paintings: Six independent trials 27
Introduction: Default preferences in LLMs
About two and a half years ago I published an informal pilot study, ChatGPT tells 20 versions of its prototypical story, with a short note on method. I discovered that when given a simple one word prompt, “story,” that places no restrictions on the nature of the story to be generated, ChatGPT tended to generate the same story each time, roughly the same general plot set in a fairy tale world. More recently Sil Hamilton and David Mimno studied 20,000 stories generated on four different platforms and discovered that words, including character names, occurred in 88% of the stories.
Given this background, I wondered: Does image generation exhibit the same effect? Once it became possible to generate images from ChatGPT I had used it to generate many different kinds of images, some from simple prompts, others from long, often very long, prompts, and still others from sample photographs. About a week ago I decided to see what kind of images ChatGPT would generate when given a prompt that made no specifications about subject matter.
This is an informal pilot study. I began on an impulse, with no specific method or goal in mind. I just wanted to see if there was anything there. If so, what do we need to do to conduct a more rigorous study?
* * * * *
I begin by presenting the basic results on default preferences in story generation as background. Then I present the methods and results of my image study. After that I discuss the work of Bob Ross, an artist who had a popular TV show in which he showed viewers how to paint images similar to those ChatGPT generated in this study. I conclude my discussion with some final remarks and suggestions for future work. Last, we have the images themselves. Note that I refer to the recurring image type as a prototype produced under minimally specified default conditions.























