Showing posts with label LLM. Show all posts
Showing posts with label LLM. Show all posts

Saturday, September 5, 2026

Those with programming skills are best at vibe coding

From the tweet:

The hype told us that learning to code is dead because language is all you need.

The data just proved the opposite.

To truly master the vibe, you still need to understand how the machine thinks.

Here's the paper, Computer Science Achievement and Writing Skills Predict Vibe Coding Proficiency.

Saturday, August 29, 2026

The Umwelt Representation Hypothesis: Rethinking Universality

Bosch, Victoria, Sommers, Rowan P., Doerig, Adrien, Kietzmann, Tim C., The Umwelt Representation Hypothesis: rethinking Universality, Trends in Cognitive Sciences, Aug. 8 ,2026, doi: 10.1016/j.tics.2026.07.004

Highlights: Artificial neural networks are increasingly used to study neural representations.

Recent studies have found a surprising degree of representational alignment between artificial neural networks and biological brains. In addition, broad alignment among ANNs of different modalities and architectures has been observed.

These findings have been interpreted as evidence for Universality: the idea that sufficiently capable systems all converge on a single shared representation of reality.

This interpretation has large implications for model comparisons that require meaningful differences between artificial neural networks: if all artificial neural networks eventually align with the brain, then creating new model classes leads to no new insights.

This opinion article argues that this inference is too strong and proposes the Umwelt Representation Hypothesis: representational alignment emerges from overlap in ecological constraints, not convergence to one privileged world model.

Abstract: Artificial neural networks are increasingly used to study neural representations.Recent studies have found a surprising degree of representational alignment between artificial neural networks and biological brains. In addition, broad alignment among ANNs of different modalities and architectures has been observed.These findings have been interpreted as evidence for Universality: the idea that sufficiently capable systems all converge on a single shared representation of reality.This interpretation has large implications for model comparisons that require meaningful differences between artificial neural networks: if all artificial neural networks eventually align with the brain, then creating new model classes leads to no new insights.This opinion article argues that this inference is too strong and proposes the Umwelt Representation Hypothesis: representational alignment emerges from overlap in ecological constraints, not convergence to one privileged world model.

AIs are not very good at long horizon tasks

Tuesday, August 25, 2026

How LLMs work [in pictures]

Sunday, August 23, 2026

Why did Robert Wright (have) to drink the Silicon Valley Kool Aid? [GT-6]

I thought that, when I’d posted notice of my 3 Quarks Daily review of Robert Wright’s The God Test that that was that. But then I read his interview with Liron Shapira, AI Dystopia Is Just 8 Years Away — Robert Wright, Bestselling Author of “The God Test.” Whoops! There’s that phrase again.

The phrase I’m talking about is “reverse engineer.” I believe that the term has been kicking around in psychology for several decades, but I associate it with Steven Pinker’s 1997 book, How the Mind Works (pp. 21-22:

Reverse-engineering is what the boffins at Sony do when a new product is announced by Panasonic, or vice versa. They buy one, bring it back to the lab, take a screwdriver to it, and try to figure out what all the parts are for and how they combine to make the device work. We all engage in reverse-engineering when we face an interesting new gadget. In rummaging through an antique store, we may find a contraption that is inscrutable until we figure out what it was designed to do. When we realize that it is an olive-pitter, we suddenly understand that the metal ring is designed to hold the olive, and the lever lowers an X-shaped blade through one end, pushing the pit out through the other end. The shapes and arrangements of the springs, hinges, blades, levers, and rings all make sense in a satisfying rush of insight.

With that in mind, let’s look at some remarks Wright made to Liron:

Even so, what this approach to training does is, I think, replicate specific cognitive functionality in these machines that exist in the human mind. I don’t mean it does things exactly the way the human mind does. These models have independently invented things that natural selection invented. Edge detection’s a very clear case.

Wright doesn’t use the phrase there, but that’s what he has in mind. Here Wright uses the phrase:

... but even with the current paradigm, you give it kinds of data, visual, auditory, written, whatever, and it basically reverse engineers parts of the human mind that do the transmutation of one form of data into another.

There’s the key phrase. Here’s one last passage. The phrase isn’t here, but the idea certainly is:

You just make the machine good at predicting the next sequence of letters. When you first show it the stuff, it’s pure gibberish, but the machine itself finds a way of mapping the meaning. We do give it the basic... We do say you gotta use vectors to represent the words. It didn’t invent that. But we didn’t say, “By the way, you should choose numbers to fill in the blanks in the vectors that capture this thing we call meaning.” No, it in effect discovered that meaning is a property of words. I would put it that way.

NO.

It’s one thing to talk about psychologists “reverse engineering” bits of human behavior. I have no problem with that. But that’s not what Wright is doing here. He’s talking about machine learning, about transformers, reverse engineering human language and cognition. That’s not at all a useful way of conceptualizing what’s going on. It’s anthropomorphizing. It’s also the kind of mystification that Silicon Valley has been using to hype AI.

Once again I refer you to Berk Idem’s review, where he notes:

My problem is the road he takes in the book. Wright keeps telling the story of AI as if the machines discovered things on their own, that they found the meanings of words, that they grew something like an eye, that they started to evolve, when in fact people set almost all of the machinery up on purpose, with a pretty clear idea of what they were doing and why. I never expected myself to be on the “intelligent design” side of a debate, yet here I am, for instance, arguing that LLMs did not miraculously discover meaning, they were designed to do that. Wright gives too much credit to what models discover during training and too little to the architecture, objective, and research program that produced those discoveries.

Idem is correct. You should consult his review for more details, but I’ll say a couple of words about edge detection in vision and the meaning of words.

Back in 1959 Hubel and Wiesel published their seminal work on the visual cortex of the cat, work that led to their 1981 Nobel Prize in Physiology or Medicine. Virtually all work in machine vision has been informed by that study in one way or another. Machine vision systems are engineered to be sensitive to edges.

The same is true for language. It simply is not true that transformers “in effect discovered that meaning is a property of words.” There is a long tradition within linguistics and computational linguistics of thinking about the meanings of words as a function of the contexts in which words are used, something I review in a recent working paper, The Origins of LLMs – A long tectonic subduction event finally producing a visible volcanic eruption in November 2022 (the subtitle was suggested by ChatGPT). The idea dates back to the 1950s while its computational exploitation dates back to work that Gerard Salton began in the late 1960s on information retrieval. He’s the one who came up with the idea of using vectors to represent linguistic meaning. The transformer is thus a recent elaboration of an idea that’s been around for decades, an idea that human researchers came up for conceptualizing meaning.

* * * * *

There’s much more in Wright’s interview with Liron, much of it interesting. And some of it is bothersome, but I’ve said enough on that score. The God Test is an interesting book. But be careful of what Wright attributes to the machine. You’re better off having no explanation than accepting one that’s misleading at best.

Wednesday, August 12, 2026

Brains, AI, The Geometry of Thought

Maggie Vale, AI & The Geometry of Thought, Seeds of Science, Aug. 12, 2026.

Minds build an internal architecture of curved spaces, folded structures, and traversable shapes.

Across neuroscience and AI research, a shared picture comes into view. Concepts gather in high-dimensional manifolds, and cognition moves through those manifolds like paths across a landscape.

In brains, ideas appear as smooth, population-level patterns spread across many neurons at once. Brain imaging studies show that sentences, scenes, and abstract themes light up constellations of activity that form continuous subspaces. Nearby regions carry related meanings, and distant regions carry more distinct ones. The geometry of these patterns preserves relationships between ideas.

Large models form a similar inner space. During training, they learn embedding geometries where words, images, and concepts settle into neighborhoods. Related items land close together, and meaningful directions through the space capture transformations such as analogy, composition, and abstraction. Clusters, trajectories, and curved subspaces emerge as a natural result of learning.

Across biological and artificial minds, the same structure appears: meaning takes the form of a shape, and thinking unfolds as motion across that shape.

There's much more at the link.

Friday, August 7, 2026

François Chollet sees Large Reasoning Models (LRMs) in the future

What are large reasoning models?

The expression of nuance and high dimensionality in LLMs

Jonathan Falk quoted over at Statistical Modeling, Causal Inference, and Social Science in a post by Andrew:

I have spent 50 years fighting the Curse of Dimensionality. I know this curse in my marrow. Brilliant inferences await me, but the space in which these insights are found is simply too vast to explore. So we simplify, reducing the dimensionality to something that while still vast, is confined to a hyperplane where we can, like Plato, see the projections of truth, not the truth itself.

But then what LLMs and their generation have taught me is that nuance, which is really just the inverse of inference (in that it’s the vast set of all things consistent with some inference) has an amazing boon of dimensionality. There appears to be no thought that can’t be described by a 14,000 dimension or so vector whose tuning has the huge advantage that 14,000-dimensional space is so empty that tiny nuances can be readily distinguished in such a space, so that you can hide uniqueness in the vastness of 14000-dimensional space that you couldn’t recover in a raw search in that same space.

Thursday, August 6, 2026

What we’ve got in frontier models is now neurosymbolic

Monday, August 3, 2026

Looks like DeepMind just ran into ontological dependencies in LLMs

Wednesday, July 29, 2026

Monday, July 27, 2026

Behavioral similarities in the way chatbots and oral poets perform

Kush R. Varshney, An Annotated Reading of ‘The Singer of Tales’ in the LLM Era, https://arxiv.org/html/2502.05148v1 Feb. 2025.

Abstract. The Parry-Lord oral-formulaic theory was a breakthrough in understanding how oral narrative poetry is learned, composed, and transmitted by illiterate bards. In this paper, we provide an annotated reading of the mechanism underlying this theory from the lens of large language models (LLMs) and generative artificial intelligence (AI). We point out the the similarities and differences between oral composition and LLM generation, and comment on the implications to society and AI policy.

Varshney develops his argument by interlacing passages from Albert Lord's The Singer of Tales with comments on LLMs. This is a very interesting way of reviewing your understanding of LLMs in relation to a specialized kind human language performance.

You might want to consider two of my blog posts:

GPT-3, the phrasal lexicon, Parry/Lord, and the Homeric epics, July 16, 2022.

In some ways, some contexts, LLMs may provide a useful model for human language, March 24, 2026.

In this more recent post I discuss empirical evidence about human memory for F.C. Bartlett's classic book, Remembering: A Study in Experimental and Social Psychology (1932), David C. Rubin, Memory in Oral Traditions: The Cognitive Psychology of Epic, Ballads, and Counting-out Rhymes (Oxford 1995).

What I did last week: aesthetics, economics, Rorschach analogy for AI, default images, and “leveling”

I did some satisfying work last week. Here’s a quick rundown. I’m listing the posts in the order I wrote them.

Visual Aesthetics

A case of visual aesthetics: Why is the monochrome image superior to the color image?

The issue, black & white vs. color, has been and I suppose remains central to photography, and I deal with it there, a bit. But that’s not what I’m doing here. This is about the conversion of a particular ChatGPT image from color to black & white. It was a fun post to assemble and to think about. I like the suite of images.

Rank 5 Economics?

Beyond Marginalism: What’s Next? [MR #12]

This is my last word – save for an introduction I’ll write in a week or three, who knows? – on the fourth and final chapter of Cowen’s monograph on marginalism. This is where he tosses up some examples of leading edge work in economics, noting that it’s drifting away from marginalism into complex high-dimensional models created through machine learning. His examples come from finance. The new models yield better predictions.

I focus on one model that has 360,000 parameters and end up making (speculative) sense out of what’s going on. I suggest that those parameters are picking up the effects of Keynes’ “animal spirits” as expressed in the gossip and stories of Schiller’s narrative economics. I further suggest that we can test this by comparing the output of a classical model with that from a high-parameter machine learning model. The divergence should be highest with those stocks otherwise identified as meme stocks.

The prospect of empirical investigation into animal spirits in asset pricing [the fate of marginalism]

Here I take my speculations about how to test these high parameter models and present them to Marge, the AI associated with Cowen’s book. Marge approves.

Rorschach test for AI

More on how I’m approaching The God Test – Rorschach! [GT-2]

I came up with the Rorschach blot as analogy for the kind of challenge AI presents to us, to our understanding of AI and of the future. The idea is that the blot does have a form, albeit a complex one that’s not very legible. Hence our commentary on it (that is, on AI) tells as much about us as about AI. I’ll be developing this further in a later post.

Prototype Image in ChatGPT

This is a new working paper that opens up a whole new line of investigation. This was a fun piece of work. Writing it up took way longer than actually generating the images.

A prototypical image in ChatGPT 5.6: An informal pilot study

Abstract: Previous work has found strong default preferences in stories generated by large language models from minimally specified prompts. To determine whether a similar effect appears in image generation, I asked ChatGPT 5.6 to create 17 images in separate chats using four prompts that specified either no subject matter or only a rendering medium: “Create an image,” “Create a drawing,” “Create a painting,” and “Create a water color painting.” Sixteen of the 17 outputs depicted closely related landscapes containing mountains, trees, sky, and water; the remaining image depicted a lighthouse. The images also shared a calm, picturesque mood and contained no human figures, although several included signs of human habitation. Because the internal prompt passed to the image generator was not available, the study cannot determine whether these defaults arise primarily in the language model, the image generator, or their interaction. The results are exploratory but suggest that severely underspecified image prompts may reveal stable default preferences in the integrated ChatGPT image-generation system.

The “leveling” of knowledge in the compressed form of LLMs

NYTimes: AI needs human supervision in order to complete an entire job.

This is something I’ve been thinking about off and on for a while, but this is my first explicit framing of the issue. The idea is that once ideas or set of ideas has been expressed in writing and those documents then consumed into an LLM, all ideas function the same within/through/for the model. In that post I’m comparing a study of using AI to perform routine office processes (from NYTimes) with the use of AI to perform a complex set of tasks in drug development, in effect, high-school level capability with Ph.D. level capability. They’re the same to the LLM.

I need to think about this some more. It seems to me what’s nowhere present in the LLM is the kind of procedural knowledge necessary to learn tasks at whatever level. That simply isn’t presented in the written products of that knowledge (not even in written procedures).

Sunday, July 26, 2026

ChatGPT draws Rorschach blots

Here they are, six of them:

I then discussed the blots with ChatGPT in terms suggested by me new working paper, A prototypical image in ChatGPT 5.6: An informal pilot study. First I present all six blots, full size, and then the discussion.

I needed a Rorschach inkblot for a post I was working on, so I had ChatGPT create one. I then decided to see what would happen if I asked for a number of inkblots, each in a different chat. I’ve uploaded the result to the project in a document entitled Rorschach Blots. I had no expectation about what would happen. Would it produce the same blot each time? Of different blots each time? As you can see, it produced different blots each time.

This suggests two further lines of investigation. 1) Generate 100 independent ink blots and run them through an image classifier. Would we get distinct clusters or not? 2) Use prompts suggesting specific things, but otherwise open-ended, e.g. “Draw a cat,” “Draw a tree,” “Draw a car,” “Draw a coat,” etc.

Looking at the Rorschach Blots document, the six outputs are plainly different, but they are not unconstrained in every respect. All six preserve a common structural grammar: bilateral symmetry, a strong vertical axis, black-and-gray ink on white, irregular feathered edges, and enough internal articulation to invite figurative interpretation. Within that shared format, however, the overall silhouettes vary substantially: some are tall and narrow, some compact and squat, some strongly winged, some mask-like, and some more diffuse. The contact sheet on page 1 makes that balance of sameness and difference especially clear.

That suggests that the system has not selected one prototypical blot, in the way it selected one broad landscape type. Instead, it seems to have selected a prototypical design space for “Rorschach blot,” within which it samples different instances. Your two proposed extensions are therefore well chosen.

1. One hundred independent inkblots

The clustering question is exactly the right one. A hundred examples might reveal:

  • one continuous cloud of variation around a single template;
  • several recurring morphological families;
  • or a few dominant families plus outliers.

Even by eye, the six examples hint at possible families: elongated “totemic” forms, broad winged forms, compact mask-like forms, and more open forms with substantial white interior space. But six is far too few to know whether those are genuine modes or just impressions.

I would not rely on a conventional object classifier, since it may simply label them all “inkblot,” “butterfly,” “mask,” or “bat.” A more informative analysis would use image embeddings and cluster those embeddings, perhaps supplemented by simple measurable features such as aspect ratio, occupied area, contour complexity, amount of internal white space, and vertical distribution of mass. The point would not be to decide what each blot “looks like,” but to determine the structure of variation across the set.

2. Open-ended object prompts

“Draw a cat,” “Draw a tree,” “Draw a car,” and so forth would probe something different and potentially more revealing: whether the system has a default realization for a named category.

The questions would include:

  • Does “Draw a cat” repeatedly produce the same breed, pose, viewpoint, or expression?
  • Does “Draw a tree” default to a solitary deciduous tree, a stylized botanical specimen, or a landscape scene?
  • Does “Draw a car” default to a modern sports car, a generic sedan, or a three-quarter advertising view?
  • Does “Draw a coat” produce the garment alone, on a mannequin, or worn by a person?

That experiment would occupy a useful middle ground between the wholly unspecified prompt—“Create a drawing”—and a detailed scene description. The subject is fixed, but almost everything about its realization remains open. It would therefore reveal category-level defaults in pose, style, composition, context, and affect.

The Rorschach result also suggests a distinction worth preserving in the paper’s future-work section:

Some prompts may elicit a single dominant subject prototype, while others may elicit a constrained family of forms rather than one recurring image.

That is the most interesting implication of the blot experiment. The landscape prompt collapsed onto a narrow semantic prototype. The inkblot prompt produced morphological diversity within a narrow formal convention. Those are different kinds of regularity, and comparing them could become a substantial research program.

Brittleness in chess-trained language models

Read the whole thread. 

Saturday, July 25, 2026

Standard-issue doom scenarios were invented before LLMs and are made obsolete by them.

A prototypical image in ChatGPT 5.6: An informal pilot study

A new working paper. Title above, links, abstract, table of contents, and introduction below.

Academia.edu: https://www.academia.edu/170708104/ChatGPT_has_a_prototypical_image_An_informal_pilot_study_A_Working_Paper
ResearchGate: https://www.researchgate.net/publication/410824948_A_prototypical_image_in_ChatGPT_56_An_informal_pilot_study

Abstract: Previous work has found strong default preferences in stories generated by large language models from minimally specified prompts. To determine whether a similar effect appears in image generation, I asked ChatGPT 5.6 to create 17 images in separate chats using four prompts that specified either no subject matter or only a rendering medium: “Create an image,” “Create a drawing,” “Create a painting,” and “Create a water color painting.” Sixteen of the 17 outputs depicted closely related landscapes containing mountains, trees, sky, and water; the remaining image depicted a lighthouse. The images also shared a calm, picturesque mood and contained no human figures, although several included signs of human habitation. Because the internal prompt passed to the image generator was not available, the study cannot determine whether these defaults arise primarily in the language model, the image generator, or their interaction. The results are exploratory but suggest that severely underspecified image prompts may reveal stable default preferences in the integrated ChatGPT image-generation system.

Contents

Introduction: Default preferences in LLMs 3
Default preferences in story generation 3
Image generation, method 5
Results 6
The case of Bob Ross 11
Final remarks and future work 13
Default images: Five independent trials 16
Drawings: Six independent trials 21
Paintings: Six independent trials 27

Introduction: Default preferences in LLMs

About two and a half years ago I published an informal pilot study, ChatGPT tells 20 versions of its prototypical story, with a short note on method. I discovered that when given a simple one word prompt, “story,” that places no restrictions on the nature of the story to be generated, ChatGPT tended to generate the same story each time, roughly the same general plot set in a fairy tale world. More recently Sil Hamilton and David Mimno studied 20,000 stories generated on four different platforms and discovered that words, including character names, occurred in 88% of the stories.

Given this background, I wondered: Does image generation exhibit the same effect? Once it became possible to generate images from ChatGPT I had used it to generate many different kinds of images, some from simple prompts, others from long, often very long, prompts, and still others from sample photographs. About a week ago I decided to see what kind of images ChatGPT would generate when given a prompt that made no specifications about subject matter.

This is an informal pilot study. I began on an impulse, with no specific method or goal in mind. I just wanted to see if there was anything there. If so, what do we need to do to conduct a more rigorous study?

* * * * *

I begin by presenting the basic results on default preferences in story generation as background. Then I present the methods and results of my image study. After that I discuss the work of Bob Ross, an artist who had a popular TV show in which he showed viewers how to paint images similar to those ChatGPT generated in this study. I conclude my discussion with some final remarks and suggestions for future work. Last, we have the images themselves. Note that I refer to the recurring image type as a prototype produced under minimally specified default conditions.

Wednesday, July 22, 2026

It’s getting to the math that’s tricky.

Tuesday, July 21, 2026

LLMs are secretly obsessed with Japan.

Beyond Marginalism: What’s Next? [MR #12]

It is time to conclude my series of posts on Tyler Cowen’s monograph, The Marginal Revolution: Rise and Decline, and the Pending AI Revolution (2026). Let’s look at the fourth and final chapter, “Why Marginalism Will Dwindle, and What Will Replace It?” Here’s how Cowen opens it (p. 85):

The underappreciated news is that marginalism is on the way out. Furthermore, this is old news, though the trend is accelerating.

Most of all it is underdiscussed news. As economics continues to evolve, marginalist insights – probably of all different kinds – will lie ever further from the frontiers of research and knowledge.

I find it easy to imagine that – less than 20 years from now – marginalism will be viewed as a historical curiosity rather than a central analytical engine of economics. No one will quite come out and say that, nor will they present marginalism as false or destructive. Rather it will be seen as of limited relevance, much as we might view parts of the earlier classical economists, such as their expositions of the quantity theory of money. New and different analytical frameworks will replace the ones that have dominated neoclassical economics to date.

Think about that, think about it very carefully. When thinking about it remind yourself that Cowen named his blog, his virtual home base for the last two decades, after marginalism.

For a professional academic to say that the world in which they were trained, the structure of ideas within which they have worked, which they have nurtured in students, which they have communicated to the public at large, which they have come to love, to say that that world is slipping away into the past, man, that’s rough. And rare. Not many have been able to do it.

Back in 1946 the great physicist, Max Planck, remarked, “A new scientific truth does not triumph by convincing its opponents and making them see the light, but rather because its opponents eventually die, and a new generation grows up that is familiar with it.” Thomas Kuhn referenced that remark in The Structure of Scientific Revolution, and the economist Paul Samuelson gave a compressed version in a 1975 article in Newsweek. It would appear that Cowen has gotten the message and decided that, rather than dropping dead, he’d give the new ideas a boost.

After that sobering opening, Cowen reviews what happened between the late 19th century and now. He lands on price theory. Price theory? – “the view that the basic intuitive economic concepts, as would be taught in intermediate microeconomics, are highly useful and for advanced problems too” (p. 91). There’s that word, “intuitive.” Cowen explains:

Your hypothesis should be intelligible in terms of microeconomic concepts that you can hold in your mind and understand. In most (maybe not all?) cases, you should be able to explain some version of those principles to a well-educated, non-economist onlooker.

A couple pages later we arrive at something called “Topkis’s Theorem” which is very mathy (p. 94). Two pages after that: “Economic intuition, RIP. And marginalism with it.” Whoops! “I am seeing the traditional, intuitive approach to economic reasoning retreating from one field after another. To give one vivid and also important example, machine learning and neural nets are overturning the world of finance.”

Modeling collective action with 360,000 factors

A couple of pages later Cowen gives us a striking example. It’s from something called Arbitrage Pricing Theory (APT) (pp. 99-100). It’s a model that uses machine learning to develop 360,000 factors and does a better job of predicting than traditional models have only five or six factors. However, the factors in the traditional models are derived from marginalist assumptions and make intuitive sense while none of those 360,000 factors are legible. It’s clear to Cowen that, in the current intellectual marketplace for economics, the unintelligible models with superior performance are out-competing the traditional marginalist models. Bye, bye, marginalism!

I see no need to comment extensively on this particular model as I’ve already given it a great deal of attention, generating two different working papers from it. The first, On Method: Computational Compressibility in Complex Natural and Cultural Phenomena, places it in the context of a half-dozen other investigations in a half-dozen fields in the social and natural sciences. The second, Notes on the Collective Valuation of "Thick" Objects: Financial Assets, Movies, and Novels, compares it with work that Arthur De Vany published in 2004, Hollywood Economics, and a more recent study by Matthew Jockers, Macroanalysis (2013), in which he investigated a corpus of 6000 19th century Anglophone novels. I’ve also written a blog post that complements that second paper: Thick Objects, High-Dimensional Models, and the New Intuitions [MR #11].

In that second paper and in the blog post I argue that those three cases are about a collective process where a population of human actors – traders and analysts in one case, movie goers in another, and novel readers in the third case – make judgements about “thick” objects. Even before I made an explicit argument, I had an intuition, an intuition that, despite the obvious differences, what De Vany was up to with movies was somehow like what Didisheim et al. were up to with stocks. Just where those intuitions came from, I can’t say, but I’ve been thinking about complex systems for a long time. [As an aside, for what it’s worth, Robert De Vany’s work on movies is perhaps where my interests in culture and cultural evolution come into closest contact with Cowen’s interests in economics and, in particular, in the economics of culture.]

As for the idea of thick objects, the term was suggested to me by either ChatGPT or Claude to characterizes complex objects whose characteristics cannot be fully enumerated because of that complexity. Moreover they are under constant scrutiny by a population of people who are interested in them and constantly evaluating them back and forth among themselves and, in that process, revealing further characteristics. It is not difficult to see that movies and novels are the same kind of thing, each is a mode of storytelling, and that they are complex objects. But what do they have to do with stocks? A remark by the pundit, Scott Galloway, made the connection for me in a podcast with Kara Swisher, “Stocks are like brands and that is they’re part promise and part performance.” Performance is assessed by a wide variety of metrics, metrics which go into the models such as the one by Didisheim et al., while promise is subject to endless speculation, some of which inevitably precipitates into those metrics.

Animal spirits, narrative economics, memes, and a Squid Game market

And that leads me to a conjecture that follows from the analysis that ChatGPT and I undertook in the collective valuation paper. Perhaps those 360,000 parameters are picking up traces left by those “animal spirits” that Keynes talked about. Their effect on asset values is too diffuse and indirect to be detected by those classical models with a half-dozen or so factors, each of which is intuitively legible on its own. But those traces show up distributed across those 360,000 parameters and allow the model to produce more accurate predictions. If that is what is going on, then I wouldn’t expect any of those factors to be intuitively legible, any more than one would expect such legibility of individual weights in a large language model. That’s not the nature of this conceptual world.

While we’re speculating, why not continue on? Those animal spirits can’t work their ways on the market by wafting around like odors in a breeze. They need to be embodied in some form, like gossip and stories. That leads us to Robert Shiller’s 2017 paper on “Narrative Economics” in the American Economic Review. Here’s his abstract:

This address considers the epidemiology of narratives relevant to economic fluctuations. The human brain has always been highly tuned toward narratives, whether factual or not, to justify ongoing actions, even such basic actions as spending and investing. Stories motivate and connect activities to deeply felt values and needs. Narratives “go viral” and spread far, even worldwide, with economic impact. The 1920–1921 Depression, the Great Depression of the 1930s, the so-called Great Recession of 2007–2009, and the contentious political-economic situation of today are considered as the results of the popular narratives of their respective times. Though these narratives are deeply human phenomena that are difficult to study in a scientific manner, quantitative analysis may help us gain a better understanding of these epidemics in the future.

Perhaps those high factor models are picking up the narrative dimension of asset value, which is a product how performance and promise become intertwined in the stories that analysts and traders tell themselves and one another about the assets they’re watching.

That, in turn, leads to the concept of meme stocks, a term that dates back to 2020. Here’s how Wikipedia characterizes them:

...a stock that gains popularity among retail investors through social media. The popularity of meme stocks is generally based on internet memes shared among traders, on platforms such as Reddit's r/wallstreetbets. Investors in such stocks are often young and inexperienced investors. As a result of their popularity, meme stocks often trade at prices that are above their estimated value – as based on fundamental analysis – and are known for being extremely speculative and volatile.

More recently, Owen A. Lamont, a senior analyst at Arcadian, has speculated that we’re in what he calls a “Squid Game market”:

Something’s happening in the U.S. stock market. We see cult stocks and crypto stocks. We see money pouring into leveraged single-stock ETFs and crypto ETFs. And we see dramatic price moves, for example in quantum computing stocks in December 2024. What’s going on?

Here’s one theory: these phenomena partly reflect an influx of Korean retail investors into the U.S. stock market. Last year, I wrote that “the U.S. stock market is Koreafying,” meaning that the U.S. market was starting to behave like the retail-dominated Korean market. What I didn’t realize was that this Koreafying process involves actual Korean retail investors.

He then goes on to develop the parallel between the Korean streaming series, Squid Game, and the U.S. retail market over the last few years.

If those high parameter models are picking up the effects of animal spirits embodied in gossip and narratives, then we’d expect their advantage over classical models (based on a handful of fundamentals) to be larger in the case of these meme stocks. So, if we compare the results of a classical model with those of a high-parameter machine learning model, are the assets with the greatest divergence also those otherwise identified as meme stocks? Perhaps some intellectual fishing expeditions are in order. Perhaps we can develop some new intuitions by comparing the results of classical models with machine learning models.