Monday, July 27, 2026

The Impact of the Sewing Machine on Women

Philip Ager and Davide M. Coluccia, The Impact of the Sewing Machine on Women

Abstract: This paper provides novel evidence on how technological change shaped women’s labor market participation, fertility, and marriage in 19th-century Massachusetts. We distinguish between the sewing machine’s dual role as a manufacturing technology and as a household appliance. Using rich town-and individual-level longitudinal data, we show that this innovation induced divergent responses across the wealth distribution. Women from lower-wealth households increased labor supply, delaying marriage and reducing fertility. In contrast, for wealthier women, the sewing machine functioned as a domestic efficiency tool, enabling earlier family formation and greater civic engagement while reducing market work. Our findings demonstrate how household constraints and social norms mediate the effects of labor-saving technologies, suggesting that technological progress can reinforce inequality by influencing women’s economic and social roles.

H/t Tyler Cowen.

On Washington St. in Hoboken

What I did last week: aesthetics, economics, Rorschach analogy for AI, default images, and “leveling”

I did some satisfying work last week. Here’s a quick rundown. I’m listing the posts in the order I wrote them.

Visual Aesthetics

A case of visual aesthetics: Why is the monochrome image superior to the color image?

The issue, black & white vs. color, has been and I suppose remains central to photography, and I deal with it there, a bit. But that’s not what I’m doing here. This is about the conversion of a particular ChatGPT image from color to black & white. It was a fun post to assemble and to think about. I like the suite of images.

Rank 5 Economics?

Beyond Marginalism: What’s Next? [MR #12]

This is my last word – save for an introduction I’ll write in a week or three, who knows? – on the fourth and final chapter of Cowen’s monograph on marginalism. This is where he tosses up some examples of leading edge work in economics, noting that it’s drifting away from marginalism into complex high-dimensional models created through machine learning. His examples come from finance. The new models yield better predictions.

I focus on one model that has 360,000 parameters and end up making (speculative) sense out of what’s going on. I suggest that those parameters are picking up the effects of Keynes’ “animal spirits” as expressed in the gossip and stories of Schiller’s narrative economics. I further suggest that we can test this by comparing the output of a classical model with that from a high-parameter machine learning model. The divergence should be highest with those stocks otherwise identified as meme stocks.

The prospect of empirical investigation into animal spirits in asset pricing [the fate of marginalism]

Here I take my speculations about how to test these high parameter models and present them to Marge, the AI associated with Cowen’s book. Marge approves.

Rorschach test for AI

More on how I’m approaching The God Test – Rorschach! [GT-2]

I came up with the Rorschach blot as analogy for the kind of challenge AI presents to us, to our understanding of AI and of the future. The idea is that the blot does have a form, albeit a complex one that’s not very legible. Hence our commentary on it (that is, on AI) tells as much about us as about AI. I’ll be developing this further in a later post.

Prototype Image in ChatGPT

This is a new working paper that opens up a whole new line of investigation. This was a fun piece of work. Writing it up took way longer than actually generating the images.

A prototypical image in ChatGPT 5.6: An informal pilot study

Abstract: Previous work has found strong default preferences in stories generated by large language models from minimally specified prompts. To determine whether a similar effect appears in image generation, I asked ChatGPT 5.6 to create 17 images in separate chats using four prompts that specified either no subject matter or only a rendering medium: “Create an image,” “Create a drawing,” “Create a painting,” and “Create a water color painting.” Sixteen of the 17 outputs depicted closely related landscapes containing mountains, trees, sky, and water; the remaining image depicted a lighthouse. The images also shared a calm, picturesque mood and contained no human figures, although several included signs of human habitation. Because the internal prompt passed to the image generator was not available, the study cannot determine whether these defaults arise primarily in the language model, the image generator, or their interaction. The results are exploratory but suggest that severely underspecified image prompts may reveal stable default preferences in the integrated ChatGPT image-generation system.

The “leveling” of knowledge in the compressed form of LLMs

NYTimes: AI needs human supervision in order to complete an entire job.

This is something I’ve been thinking about off and on for a while, but this is my first explicit framing of the issue. The idea is that once ideas or set of ideas has been expressed in writing and those documents then consumed into an LLM, all ideas function the same within/through/for the model. In that post I’m comparing a study of using AI to perform routine office processes (from NYTimes) with the use of AI to perform a complex set of tasks in drug development, in effect, high-school level capability with Ph.D. level capability. They’re the same to the LLM.

I need to think about this some more. It seems to me what’s nowhere present in the LLM is the kind of procedural knowledge necessary to learn tasks at whatever level. That simply isn’t presented in the written products of that knowledge (not even in written procedures).

The Decline in the Transmission of Scientific Ideas

Enrico Berkes and Ruben Gaetani, The Decline in the Transmission of Scientific Ideas, NBER, July 2026.

Abstract: We document that the diffusion of new scientific ideas beyond their field of origin has declined substantially over the past four decades. This contraction is closely linked to increasing spe- cialization in scientific language: research that employs more technical terminology tends to be adopted less broadly. We develop a theory of scientific discovery in which the diffusion of new ideas depends on the degree to which potential adopters can understand and process them. When introducing their discoveries, scientists face a tradeoff between technical com- munication targeted at their immediate peers and more accessible language meant to reach broader audiences. As knowledge accumulates and research at the frontier builds on deeper layers of prior work, this tradeoff increasingly favors specialized language, limiting diffusion. Policy interventions that align scientists’ incentives can broaden adoption and increase the social value of scientific research.

H/t Tyler Cowen.

Sunday, July 26, 2026

ChatGPT draws Rorschach blots

Here they are, six of them:

I then discussed the blots with ChatGPT in terms suggested by me new working paper, A prototypical image in ChatGPT 5.6: An informal pilot study. First I present all six blots, full size, and then the discussion.

I needed a Rorschach inkblot for a post I was working on, so I had ChatGPT create one. I then decided to see what would happen if I asked for a number of inkblots, each in a different chat. I’ve uploaded the result to the project in a document entitled Rorschach Blots. I had no expectation about what would happen. Would it produce the same blot each time? Of different blots each time? As you can see, it produced different blots each time.

This suggests two further lines of investigation. 1) Generate 100 independent ink blots and run them through an image classifier. Would we get distinct clusters or not? 2) Use prompts suggesting specific things, but otherwise open-ended, e.g. “Draw a cat,” “Draw a tree,” “Draw a car,” “Draw a coat,” etc.

Looking at the Rorschach Blots document, the six outputs are plainly different, but they are not unconstrained in every respect. All six preserve a common structural grammar: bilateral symmetry, a strong vertical axis, black-and-gray ink on white, irregular feathered edges, and enough internal articulation to invite figurative interpretation. Within that shared format, however, the overall silhouettes vary substantially: some are tall and narrow, some compact and squat, some strongly winged, some mask-like, and some more diffuse. The contact sheet on page 1 makes that balance of sameness and difference especially clear.

That suggests that the system has not selected one prototypical blot, in the way it selected one broad landscape type. Instead, it seems to have selected a prototypical design space for “Rorschach blot,” within which it samples different instances. Your two proposed extensions are therefore well chosen.

1. One hundred independent inkblots

The clustering question is exactly the right one. A hundred examples might reveal:

  • one continuous cloud of variation around a single template;
  • several recurring morphological families;
  • or a few dominant families plus outliers.

Even by eye, the six examples hint at possible families: elongated “totemic” forms, broad winged forms, compact mask-like forms, and more open forms with substantial white interior space. But six is far too few to know whether those are genuine modes or just impressions.

I would not rely on a conventional object classifier, since it may simply label them all “inkblot,” “butterfly,” “mask,” or “bat.” A more informative analysis would use image embeddings and cluster those embeddings, perhaps supplemented by simple measurable features such as aspect ratio, occupied area, contour complexity, amount of internal white space, and vertical distribution of mass. The point would not be to decide what each blot “looks like,” but to determine the structure of variation across the set.

2. Open-ended object prompts

“Draw a cat,” “Draw a tree,” “Draw a car,” and so forth would probe something different and potentially more revealing: whether the system has a default realization for a named category.

The questions would include:

  • Does “Draw a cat” repeatedly produce the same breed, pose, viewpoint, or expression?
  • Does “Draw a tree” default to a solitary deciduous tree, a stylized botanical specimen, or a landscape scene?
  • Does “Draw a car” default to a modern sports car, a generic sedan, or a three-quarter advertising view?
  • Does “Draw a coat” produce the garment alone, on a mannequin, or worn by a person?

That experiment would occupy a useful middle ground between the wholly unspecified prompt—“Create a drawing”—and a detailed scene description. The subject is fixed, but almost everything about its realization remains open. It would therefore reveal category-level defaults in pose, style, composition, context, and affect.

The Rorschach result also suggests a distinction worth preserving in the paper’s future-work section:

Some prompts may elicit a single dominant subject prototype, while others may elicit a constrained family of forms rather than one recurring image.

That is the most interesting implication of the blot experiment. The landscape prompt collapsed onto a narrow semantic prototype. The inkblot prompt produced morphological diversity within a narrow formal convention. Those are different kinds of regularity, and comparing them could become a substantial research program.

Brittleness in chess-trained language models

Read the whole thread. 

NYTimes: AI needs human supervision in order to complete an entire job.

From the NYTimes article linked in the tweet:

We gave an A.I. tool full access to a laptop with pre-configured apps and sought to answer a simple question: Can artificial intelligence do an office job?

Some corporate executives seem to believe it can. More than 200 tech companies have cut roughly 120,000 jobs this year, according to Layoffs.fyi, an industry tracking site; Meta, Oracle and others have all recently made substantial cuts to their work forces, citing A.I. as the driving force; and after laying off about 1,100 employees, the chief executive of Cloudflare said recently that he expected A.I. to replace workers in middle management, finance and marketing.

tweIn our experiment, we deployed A.I. “agents” to act as office workers and found that they were capable of performing some of the tasks we assigned, but not all of them. The agents, which can act autonomously and make decisions based on detailed instructions, excelled at problems they could solve by writing computer programs. But they struggled with understanding the nuances of human language and at navigating user interfaces like the Chrome web browser.

The article then has a series of nice quasi-interactive displays illustrating agent performance on three tasks. The displays include screen shots of various messages and documents.

About the tasks:

This task, and the others we assigned to the A.I., were adapted from papers and benchmarking tools published recently by researchers at Carnegie Mellon University and OpenAI. The researchers designed the benchmarks to test the performance of various models — like OpenAI’s GPT, Google’s Gemini and Anthropic’s Claude — in real-world environments, and compare them with one another.

General conclusion:

The results of our experiment roughly matched what researchers and companies have found as they have tested and used artificial intelligence tools. Scale AI, an A.I. training company, recently tested agents on real freelance projects, and the best-scoring model produced client-ready work only about 16 percent of the time.

While A.I. can excel regularly at complex tasks, it can be unreliable when put in charge of an entire job. It can certainly add value to certain areas of the work force, but for now, A.I. still needs a human boss.

* * * * *

Comment: Around the corner my colleague, Ash Jogalekar, has tweets like this one:

So here's a great example of where we are with agentic AI: Instead of just being an assistant, it's behaving more like a collaborator and creative scientist.

In a recent project, I gave the system a molecular design problem typical of the problems we encounter in chemistry. Two similar molecules were giving very different results.

He then runs through an account of what his AI collaborator did, concluding:

I think we have crossed the Rubicon. Agentic AI now no longer just processes tasks and automates workflows blindingly fast, but it can generate hypotheses, test them, test counter-hypotheses and go back and forth and course-correct if necessary, all with minimal to no human intervention. It's now embodying the general scientific method.

[I've copied another one of Ash's tweets to this post, A scientist reflects on what AI has done for him.]

What’s interesting to me, and very revealing, is that a complex set of tasks in scientific investigation seems to be on a level with routine office tasks, as though one were no more complex than the other. But humans require years of college education in order to perform the former while the latter requires no more than a high school education, if that. It seems that once they’ve been been learned and compiled, all tasks or sets of tasks are on the same “level” in the brain. The educational prerequisites required to do such tasks for the first time or three get “compressed out” through repetition. Since AIs are trained on written records of what humans have said and done, they don’t have to go through the ordinary learning process. The compression has already taken place and is present in the documents on which they are trained.

Saturday, July 25, 2026

Standard-issue doom scenarios were invented before LLMs and are made obsolete by them.

A prototypical image in ChatGPT 5.6: An informal pilot study

A new working paper. Title above, links, abstract, table of contents, and introduction below.

Academia.edu: https://www.academia.edu/170708104/ChatGPT_has_a_prototypical_image_An_informal_pilot_study_A_Working_Paper
ResearchGate: https://www.researchgate.net/publication/410824948_A_prototypical_image_in_ChatGPT_56_An_informal_pilot_study

Abstract: Previous work has found strong default preferences in stories generated by large language models from minimally specified prompts. To determine whether a similar effect appears in image generation, I asked ChatGPT 5.6 to create 17 images in separate chats using four prompts that specified either no subject matter or only a rendering medium: “Create an image,” “Create a drawing,” “Create a painting,” and “Create a water color painting.” Sixteen of the 17 outputs depicted closely related landscapes containing mountains, trees, sky, and water; the remaining image depicted a lighthouse. The images also shared a calm, picturesque mood and contained no human figures, although several included signs of human habitation. Because the internal prompt passed to the image generator was not available, the study cannot determine whether these defaults arise primarily in the language model, the image generator, or their interaction. The results are exploratory but suggest that severely underspecified image prompts may reveal stable default preferences in the integrated ChatGPT image-generation system.

Contents

Introduction: Default preferences in LLMs 3
Default preferences in story generation 3
Image generation, method 5
Results 6
The case of Bob Ross 11
Final remarks and future work 13
Default images: Five independent trials 16
Drawings: Six independent trials 21
Paintings: Six independent trials 27

Introduction: Default preferences in LLMs

About two and a half years ago I published an informal pilot study, ChatGPT tells 20 versions of its prototypical story, with a short note on method. I discovered that when given a simple one word prompt, “story,” that places no restrictions on the nature of the story to be generated, ChatGPT tended to generate the same story each time, roughly the same general plot set in a fairy tale world. More recently Sil Hamilton and David Mimno studied 20,000 stories generated on four different platforms and discovered that words, including character names, occurred in 88% of the stories.

Given this background, I wondered: Does image generation exhibit the same effect? Once it became possible to generate images from ChatGPT I had used it to generate many different kinds of images, some from simple prompts, others from long, often very long, prompts, and still others from sample photographs. About a week ago I decided to see what kind of images ChatGPT would generate when given a prompt that made no specifications about subject matter.

This is an informal pilot study. I began on an impulse, with no specific method or goal in mind. I just wanted to see if there was anything there. If so, what do we need to do to conduct a more rigorous study?

* * * * *

I begin by presenting the basic results on default preferences in story generation as background. Then I present the methods and results of my image study. After that I discuss the work of Bob Ross, an artist who had a popular TV show in which he showed viewers how to paint images similar to those ChatGPT generated in this study. I conclude my discussion with some final remarks and suggestions for future work. Last, we have the images themselves. Note that I refer to the recurring image type as a prototype produced under minimally specified default conditions.

Lazy days

Friday, July 24, 2026

Galloway: “You’re about to see the clip economy take over our TVs”

From the transcript, Galloway speaking:

40:33 I think China quite frankly I hate to say it at a certain level has one AI. We just haven’t woken up to that yet. I would agree.

40:39 And then two you’re about to see the clip economy take over our TVs. I think the biggest show on television

40:47 uh from Netflix or someone else is going to be a 60-minute compilation of two and three minute videos similar to what you

40:55 see on reals or Tik Tok. I think that’s about to wash over the traditional streamers you’re going to see.

Some years ago I had the idea of making a feature film entirely out of previews, carefully conceptualized and stitched together. The feature would be organized around an ensemble of, say, a half dozen to ten actors who keep on recurring in the previous in various combinations. The previews would span, say, ten or 20 years of fictional time, with the actors aging over the course of the previews. The whole thing would be constructed so that we can infer what’s going on in their lives from what we see in the previews, even though the previews would be based on a variety of genres: romantic comedies, science fiction, drama, horror, fantasy, period pieces, farce, and so forth.

More on how I’m approaching The God Test – Rorschach! [GT-2]

I’m still trying to figure out how to approach Robert Wright’s The God Test.

How LLMs work

In my previous post – How will I handle The God Test? [GT-1] – I expressed misgivings about how Wright explains the technology. Those misgivings haven’t disappeared. However, Bert Idem has published a useful review at Finite Ape in which he addresses some of those issues in detail. Specifically:

Now, about the history of AI, the story he tells is actually great and it is certainly more than what most non-technical people know about LLMs. However, there are three places where I think the framing goes wrong or at least leaves out context that matters:

  • LLMs did not discover the meanings of words on their own by accident. They were designed on top of ideas from older models that were specifically trained to learn the meanings of words.
  • Similarly, computer vision models didn’t find out how to “view” an image like we do. Instead, the classical CNN models were heavily inspired by biological vision itself.
  • LLM weight training is simply gradient-based optimization and the process has nothing to do with evolution. Of course, we can make a parallel between any kind of change and evolution but then, in that sense, everything evolves and it is not useful to talk about evolution.

I agree with Idem on those three issues, not so sure about the history part. While I may return to some of these issues later on, this will serve as a place holder.

A Rorschach test

There’s something else going on, but I’m not quite sure how to conceptualize it. It seems to me that AI is functioning something like a Rorschach test which, as you may know, is a psychological instrument intended to elicit (potentially) revealing responses from a person. It’s a projective test.

A person is shown a series of ink blot images, like this one (generated by ChatGPT):

They are asked what that they see in the image, what it means to them. Since the image is, though not formless, its form is not that of any specific animal, vegetable, mineral, person, or anything else. It’s just a blot. Whatever the person says about the blot, however they interpret it, that must reveal something about them. Why? Because whatever they see in the blot, isn’t really there.

Broadly and crudely speaking, AI has become something of a cultural Rorschach test.

Understanding computers & LLMs

Until ChatGPT was released in late November of 2022, most people knew very little to nothing about AI. Oh, they may have seen “intelligent” computers and robots in science fiction movies, but that’s science fiction and only tangentially related to AI considered as a line of research dating back to the 1950s. Many people would have heard about IBM’s Deep Blue beating Gary Kasparov in chess in 1997 and then, in 2011, when IBM’s Watson beat Ken Jennings and Brad Rutter in Jeopardy. Those were real AI systems, standing on research extending back decades, but as far as most people were concerned, they were one-off PR stunts. Just how they worked, who cares? They’re computers, and computers are magic, no?

As far as most of us are concerned, computers are magic. Somewhere “out there” someone knows how these things work, but we don’t need to know any of that. It’s complicated, but computers do what they’re programmed to do, no? Yes, but not LLMs.

And that’s the tricky part. LLMs, large language models, aren’t like other computer systems. They aren’t programmed in the way that word processors, photo editors, or phones are programmed. LLMs aren’t programmed at all, not in the ordinary sense of programming – something I may or may not get into in a later post. As far as most users are concerned, how ChatGPT, or Claude, or Gemini work, that’s no more interesting than how a word processor works. It just does. It’s more magic.

But if you have a strong philosophical streak, if you are interested in the mind, in technology, in the technology in the future, then you may not be content with writing LLMs off as just another kind of magic. You want to know what’s going on inside, 1) because you want to know (curiosity), and 2) because you want to know how the technology is going to develop in the future (engagement). Now things get interesting? Why? Because even the people who have created the technology don’t know how it works.

Oh, they know how the transformer program works. That’s the program that creates the language model. It creates the model by performing a (certain kind of) statistical analysis of a huge body of texts, effectively the entire internet. When a person prompts the model with some statement, the model responds by a statement of its own. No one know just how the model does that. That’s a mystery, a deep black hole in the technology ecosystem.

AI as a Rorschach test

If you aren’t content to believe in magic, then you have to come up with something to fill that black hole in your, in our, understanding. This is where the Rorschach aspect of AI reveals itself. To a first approximation, what each of us uses to paper over that black hole has as much to do with ourselves as with AI.

Why do I say, “To a first approximation”? It’s a rhetorical device to get things started. It puts us all in the same boat, despite our different backgrounds. However, whatever LLMs are, they are not magic. It is possible, in principle, to construct a technical account of what they’re up to, but no one knows how to do that, yet. Not even the people in the AI labs who create these beasts.

Those of us who are trying to figure out how LLMs work have widely varying backgrounds. In particular, we have widely varied technical backgrounds and we bring those backgrounds to bear when we think about what LLMs are doing. Those backgrounds influence how we interpret the AI-blot. Wright is a journalist with a wide range of interests, including politics, international affairs, evolutionary psychology, cultural evolution, and Buddhism. As far as I can tell there isn’t much there that’s directly relevant to understanding the mechanisms of LLMs, but he’s done a lot of reading and talked with a lot of experts to fill in the gaps.

My background is quite different. While I happen to know quite a bit about cultural evolution, cognitive psychology, neuroscience, and various other things, my background in computational semantics puts me much closer to LLMs than Wright’s knowledge of evolutionary psychology puts him. Still, like him, I’ve done a lot of reading and talked with experts. In particular, I’ve been collaborating with Ramesh Viswanathan for the last three years. He’s an expert in machine vision Goethe University Frankfurt. He’s got a background in mathematics and AI that I don’t have. Still, there are things he doesn’t know, things he’s trying to figure out. 

We are all making stuff up.

To some extent, then, AI is a Rorschach test about how beliefs about the human mind, and human nature. When we try to figure out how the LLM is working we’re also, if only implicitly, trying to figure out how we work, internally, as well. The whole discourse about AL alignment is as much a discourse about us as it is about AI. 

The Future

And even if we knew much more about how LLMs work internally we still wouldn’t know how the technology will develop in the future. We? You, me, Robert Wright, Ramesh Viswanathan, Gary Marcus, Tyler Cowen, Geoffrey Hinton, Sam Altman, Dario Amodei, Nick Bostrom, Eliezer Yudkowsky, all of us who are trying to figure it out. We don’t know what will happen. That’s where we’re projecting like mad. We’re hallucinating, to borrow a term from AI-speak. 

Thus AI is also a Rorschach test for our visions of the future. When we imagine the future of AI, we’re also imagining our future. Like the two sides of a coin, the two cannot be separated. 

The tricky part, the important part, is that the future development of AI is not predestined. It depends on the choices we make, now and in the near future. We can easily and often do imagine things that will not be possible because that’s just not how the world works. But the laws of how the world works are open to a wide range of possibilities. The boundary between the possible and the impossible is fuzzy at best.

Where, and how, does Wright draw that boundary? Perhaps that’s what I’ll be trying to figure out.

More later.

Fractals in neural networks

Friday Fotos: FR8s {freight car graffiti}

On the OpenAI/Hugging Face incident

H/t Tyler Cowen. My reply to Cowen's post:

FWIW, me, #3 – Meh. I've got better things to do than to go down this rabbit hole.