Showing posts with label macroanalysis. Show all posts
Showing posts with label macroanalysis. Show all posts

Monday, June 29, 2026

Notes on the Collective Valuation of “Thick” Objects: Financial Assets, Movies, and Novels

New working paper. Title above, links, abstract, TOC, and introduction below.

Links:

Academia.edu: https://www.academia.edu/169390494/Notes_on_the_Collective_Valuation_of_Thick_Objects_Financial_Assets_Movies_and_Novels
ResearchGate: https://www.researchgate.net/publication/408219138_Notes_on_the_Collective_Valuation_of_Thick_Objects_Financial_Assets_Movies_and_Novels

Abstract: Machine learning is creating a methodological bridge between disciplines that previously seemed far apart, especially economics and literary criticism. The bridge is the analysis of how populations deal with “thick objects.” A thick object is not exhausted by a few visible traits. It gathers interpretation, expectation, memory, value, narrative, and social response. A toaster is usually a thin object. A firm that manufactures toasters is thick: it has assets, debt, brands, patents, management, supply chains, analyst coverage, market expectations, and future promises. Scott Galloway’s remark that stocks are like brands — part promise, part performance — links stock, movies and novels. Each is a thick object moving through a field of collective judgment. Its value reflects both measurable performance and imagined future promise. They are thus as neighboring cases in a general problem: how populations perceive, classify, value, and transform thick objects. Machine learning constructs object-spaces from the traces minds leave behind. The task now is to learn how to interpret those spaces without mistaking the model for the world.

High-dimensional asset-pricing models start with many stock characteristics — price, returns, volume, profitability, leverage, liquidity, analyst revisions, momentum, volatility, investment, and so on. These characteristics are traces of firm activity, accounting conventions, analyst judgment, and trader behavior. New models then generate hundreds of thousands of nonlinear transformations from those characteristics in order to approximate the market’s pricing kernel, the structure through which future payoffs are priced under uncertainty. The individual factors are analytic objects approximating the valuation geometry produced by collective market activity.

That sounds strange in economics, but it is familiar from Matthew Jockers’ work on nineteenth-century Anglophone novels. Jockers created a high-dimensional design space from thousands of novels, using stylistic features and topic models. His topics are not literal thoughts in anyone’s mind. They are model-derived approximations to recurrent regions of culturally circulating thought. Yet the model revealed historical direction: novels arranged by similarity formed a temporal diagonal, a computationally disciplined proxy for population-level cultural cognition.

Arthur De Vany’s model of Hollywood adds the dynamic bridge. Movies are thick expressive-market objects. Their success cannot be predicted simply from stars, director, budget, genre, or advertising. Once released, they enter an audience field where word of mouth, imitation, and nonlinear cascades determine their fate. Most fail, some profit, a few become blockbusters. The dynamics are heavy-tailed, interactive, and collective.

Contents

Introduction: Using ChatGPT for focused intellectual exploration across disciplines 3
Thick Objects: Ground Shared by Economics and Cultural Analysis [Summary] 9
AIPT, Large Factor Models [First Session] 17
Hollywood Economics 23
Macroanalysis 27
The emerging triad 30
Direction over time 31
Doing a Jockers style analysis for financial assets 38
Thinking about thick objects 40
Stocks are like brands [Session Two] 42
Algorithmic and Causal models [Session Three] 52
Those empirical APT models [Session Four] 56
Decision space 63
A bridge between disparate disciplines 67

Introduction: Using ChatGPT for focused intellectual exploration across disciplines

This document serves two purposes. It presents a specific argument leading to the following provisional formulation:

High-dimensional models of novels, movies, and assets disclose the population-level geometry of collective interpretation around thick objects, turning literary criticism and economics into neighboring sciences of modeled valuation.

How I arrived at the speculation, however, is as important as the idea itself, perhaps more so. I did not arrive at that idea unaided. ChatGPT helped me. Those aren’t my words; they’re ChatGPT’s. I know a great deal about literary criticism and about movies, but not much about economics. I need ChatGPT to bridge the conceptual distance between the humanities, literary criticism, and the social sciences, economics.

Methodological curiosity

Fortunately the peculiar circumstances of my career have forced me to be interested in method and epistemology: How is it that we can come to know about the world and what methods can we use to arrive at that knowledge? When I entered Johns Hopkins as a freshman in 1965 the discipline of literary criticism was in a state of crisis, though I didn’t know that. How could I? I’d only just graduated high school and I still pretty much knowledge as it was handed to me.

That soon changed. The details of just how, when, and why don’t matter much at the moment. That it happened is sufficient for my present purposes. The upshot is that I became interested in Coleridge’s “Kubla Khan” in my senior year. I investigated the poem with standard interpretive methods augmented by avant garde structuralism and found patterns I could not explain. But they “smelled” of the nested loops I learned about in a course in computer programming.

That sent me to the English Department at SUNY Buffalo, which had the best experimental program in the nation. I found a fellow graduate student, Ralph Henry Reese, who pointed me around a corner and down the hall to David Hays in Linguistics. Hays had been a first generation researcher in machine translation at the RAND Corp. and, as such, was one of the founders of computational linguistics. While I wasn’t able to resolve my issues with “Kubla Khan” – they’re still hanging fire – I became hooked on cognitive science. Consequently my dissertation in the English Department was also a quasi-technical exercise in knowledge representation, the discipline within cognitive science and artificial intelligence about the representation of human knowledge in computable form.

Given that that is where I had arrived in the late 1970s it is perhaps not so strange that now, decades later, I find myself staring down some pretty formidable economics despite never having studied the subject. For the last 15 years, however, I have been reading the Marginal Revolution blog hosted by Tyler Cowen and Alex Tabarrok and I have been reading my way through Cowen’s recent monograph, The Marginal Revolution: Rise and Decline, and the Pending AI Revolution (2026). Cowen’s theme in the fourth (and last) chapter is that the economics he was trained in, the economics which followed from the Marginal Revolution, is rapidly being eclipsed by a more determinedly empirical discipline based on machine learning.

Bombed by 360,000 factors

Here is Cowen’s premier example. It’s from something called Arbitrage Pricing Theory (APT) (pp. 99-100):

There is a recent working paper which is perhaps more striking yet, by Antoine Didisheim, Shikun (Barry) Ke, Bryan T. Kelly, and Semyon Malamud. They pick up from Arbitrage Pricing Theory (APT), a well-established idea from financial economics. APT typically looks for “factors” in the data which predict excess returns, and a traditional APT model might have found five or six such factors. Are “inflation” or perhaps “the term structure of interest rates” useful factors? Well, that can be debated, but if so, those results sound pretty intuitive. But those intuitions seem to be disappearing. In a paper by these authors, they apply machine learning methods to look for more factors. As we know, machine learning is very good at finding non-obvious relationships in the data. The largest model they built has 360,000 (!) factors, and it reduces pricing errors by 54.8 percent relative to the classic six-factor model from Fama and French. Bravo to the authors, but what kinds of intuitions do you think possibly can be supported by those 360,000 factors?

When I read that, it “looked like Greek to me,” as the cliché has it. But I took a deep breath and thought carefully, step by step and concluded that the assets in question are stocks. What you need to pay attention to is 1) the contrast between six factors and 360,000 factors, 2) the fact that one set of factors is intuitive while the other certainly is not, 3) but the unintelligible, unintuitive, collection of factors does a better job of pricing. That’s the new world toward which economics is moving. While the old intuitions are gasping for breath the new-fangled numbers are fit as a fiddle and ready for duty.

I thought some more and realized that what’s really going on is that people are evaluating those stocks, communicating with one another directly about them, and making decisions about buying and selling, thereby communicating indirectly with one another. That’s what those 360,000 factors are capturing, the actions of a dispersed community of analysts and traders. “Could this be roughly similar to the decisions movie-goers make about the movies they see based, not only on their preferences, but on information they get from reviews, and perhaps more importantly, from their friends?” “If so,” I conjectured, “then perhaps Cowen’s old colleague from Irvine, Arthur De Vany, can shed some light on the situation.” That is to say, can give me some intuitions that I can apply to the situation.

For De Vany had written a very interesting book, Hollywood Economics (2004), about the fate of movies once they have been released. Just as those intuitive “classical” models in economics aren’t as accurate as the new high-factor models, so you can’t predict the box-office performance of movies on such simple factors as the identities of the producer, screen writers, or stars in the movies. Now, De Vany didn’t produce a high-factor model that improved matters, he did something quite different (which is discussed below, pp. 23 ff.), but that’s secondary at the moment. The point is that we seem to have a gross similarity, the behavior of some object that interests a lot of people, a stock or a movie, cannot be reliably predicted using a simple model.

Meme stocks and novels

The similarity was reinforced when I heard a remark by Scott Galloway on the Pivot podcast: “Stocks are like brands and that is they’re part promise and part performance.” Consider the recent phenomenon of meme stocks, which Wikipedia glosses this way:

A meme stock is a stock that gains popularity among retail investors through social media. The popularity of meme stocks is generally based on internet memes shared among traders, on platforms such as Reddit's r/wallstreetbets. Investors in such stocks are often young and inexperienced investors. As a result of their popularity, meme stocks often trade at prices that are above their estimated value – as based on fundamental analysis – and are known for being extremely speculative and volatile.

Meme stocks are assets where promise overwhelms performance, more story than substance.

That’s what movies are. You are purchasing the story and the experience, not the seat in the theater, or the DVD, or the stream, those are the vehicles that carry the story. Claude calls these things “thick” objects (perhaps borrowing from the anthropological concept of “thick” description? ), as opposed to “thin” objects like toasters and drills. Novels are thick objects as well, which led me to Matthew Jockers’ 2013 book, Macroanalysis, where he uses machine learning to develop a high dimensional model (a mere 600 dimensions rather than 360,000) of a corpus of 3000 19th century Anglophone novels. Just as read De Vany’s book quite closely, so I’ve written a series of posts about Jockers’ book. I bring his model into the mix as well (pp. 27 ff.).

Thus I am now in a position to take two models in subjects I know well, movies and novels, and bring them to bear on contemporary machine learning in financial economics, a subject I do not know at all. And, for that matter, still don’t. But I’ve got some intuitions. And one of those intuitions led me to focus on the fact that, while Jockers’ model did not contain any dates, upon inspection it turned out to have a diagonal (p. 27) that is correlated with direction in time. Not only did 19th century novels change in theme and motif over time, there is a direction to that change. The system seems to exhibit directional evolution. And so I directed Claude to explore the possibility of that this might be a general characteristic of thick-objects being used by a large population of interested parties (pp. 31 ff). Here is the conjecture Claude arrived at (p. 35):

In thick-object domains, low-dimensional intuitive factors often fail to explain individual outcomes. But high-dimensional representation can reveal population-level structure: outcome basins in movies, pricing kernels in finance, and temporal direction in novels. The next step is to ask whether all such artifact systems exhibit historical vectors in feature space, generated by a generational ratchet in which each cohort of producers is shaped by the artifact ecology inherited from its predecessors.

Notice the territory we have traversed in conceptual space. We started with an undergraduate at Johns Hopkins (me) using interpretive methods to study a poem, “Kubla Khan.” That investigation led to problems that forced me to study computational semantics in graduate school, a distinctly different mode of intellectual work, one based on formulating an elaborate system of structural rules. We then zipped through time and over intellectual space to a social scientist, Tyler Cowen, who was trained in the used of causal models to generate statistically controlled observations about economic behavior. He is now confronted with multifactor machine learning models with no intuitively discernible causal structure that nonetheless have superior predictive power. Cowen got me interested in one of those models and I, in turn, summoned Anthropic’s Claude to explain it to me.

The way I see, and I’ve seen it this way for a long time, the human sciences – more a European notion than American, les sciences humaines – can be arranged into three camps according to methodological focus: interpretive or hermeneutic (roughly, the humanities), causal modeling (roughly, the social sciences), and structural rules (roughly, the “classical” cognitive sciences). We’ve spanned them all in the course of this introduction. What will the future bring?

Bonus: I leave it as an exercise for the reader to consider the relevance of Keynes’s talk of “animal spirits” and to incorporate Robert Shiller’s narrative economics into this picture.

What’s in this document

The rest of this document is devoted to the dialogs where I used ChatGPT to work through the connections between these three models, two I knew quite well (De Vany on movies and Jockers on novels), and one I did not (Didisheim et al. on asset pricing). Claude knows them all, for some non-trivial meaning of “know,” and many others as well. The purpose of the dialog, then, is to link something I do not know to something that I do. The dialog took place in four sessions over the course of a week from the end of May into June.

Rather than comment on each of the sections listed in the outline, with one exception, I am commenting only on the sections that mark the beginning of a new session with ChatGPT. For what it’s worth, they mark how the subject evolved in my mind. The one exception? The summary was the last thing ChatGPT did, obviously, but I moved it to first place.

Thick Objects and the New Common Ground of Economics and Cultural Analysis [Summary] – I had ChatGPT prepare this summary and the very end of the process, on June 22. I put if first in case some might want to get the gist of the exercise without slogging through the details.

AIPT, Large Factor Models [First Session] – There is where I began on May 26. I started by asking ChatGPT to explain asset pricing to me. Once I had some sense of that, I then went on to the models I was familiar with, first De Vany on movies and the Jockers on 19th century Anglophone novels.

Stocks are like brands [Session Two] – I initiated this session on May 30 when I heard Galloway’s remark about stocks being like brands. That crystalized things for me so I needed to work back through the analysis. In the course of that discussion I focused on the concept of a brand as a distinct conceptual objects and ChatGPT’s response clarified the role of marginalism in clearing the way for asset models with a very large number of factors.

Algorithmic and Causal models [Session Three] – I don’t recall whether anything in particular prompted me to initiate this dialog. Perhaps mere methodological curiosity. This took place on June 2.

Those empirical APT models [Session Four] – It’s not entirely clear to me just whether anything in particular prompted this session. But what I was thinking was that, while I’m familiar with novels and movies and the academic discourse about them, asset pricing is unfamiliar territory. So I wanted to nail down as well as I could just what “ground truth” is in this area. Movies start with eyeballs in theaters and novels start with eyeballs scanning pages, where does asset pricing start? Once ChatGPT had gone through this I realized that I’d seen it earlier in the whole process. Still, I was happy to go through it again, this time coming at it after having thought about it. It’s as the end of this session that I asked ChatGPT to summarize the discussion.

Tuesday, June 16, 2026

Thick Objects, High-Dimensional Models, and the New Intuitions [MR #11]

Early in Chapter 4 of his marginalism monograph, “Why Marginalism Will Dwindle, and What Will Replace It?,” Cowen reviews what happened between the late 19th century and now. He then makes his way to a striking example from something called Arbitrage Pricing Theory (APT) (pp. 99-100):

There is a recent working paper which is perhaps more striking yet, by Antoine Didisheim, Shikun (Barry) Ke, Bryan T. Kelly, and Semyon Malamud. They pick up from Arbitrage Pricing Theory (APT), a well-established idea from financial economics. APT typically looks for “factors” in the data which predict excess returns, and a traditional APT model might have found five or six such factors. Are “inflation” or perhaps “the term structure of interest rates” useful factors? Well, that can be debated, but if so, those results sound pretty intuitive. But those intuitions seem to be disappearing. In a paper by these authors, they apply machine learning methods to look for more factors. As we know, machine learning is very good at finding non-obvious relationships in the data. The largest model they built has 360,000 (!) factors, and it reduces pricing errors by 54.8 percent relative to the classic six-factor model from Fama and French. Bravo to the authors, but what kinds of intuitions do you think possibly can be supported by those 360,000 factors?

That sent me through a loop. I thought about it, consulted Claude, and ended up writing a rather long working paper, On Method: Computational Compressibility in Complex Natural and Cultural Phenomena. I thought about it some more and then went to ChatGPT this time for another long dialog (14K words). I then asked ChatGPT to summarize the dialog. Here’s the prompt I gave it:

Would you summarize the preceding discussion in terms suitable for a blog post? I want it to start with some kind of quick introductory overview and then go on through asset pricing, De Vany on movies, and Jockers on novels in that order. Two things to bridge between asset pricing and movies: 1) Despite the different objects, movies vs. stocks, the objects are both ‘thick’ and the underlying structures are similar to a first approximation though the analytic methods are quite different. 2) Galloway’s remark about stocks as brands connects movies with stocks as thick objects.

I’ve appended that summary below. The title of this post is the one ChatGPT gave to its summary.

* * * * *

We have been circling around a question raised by Tyler Cowen’s recent reflections on economics and artificial intelligence: what happens when the old economic intuitions no longer seem adequate to the objects economists are trying to understand? Cowen’s worry is that marginalist reasoning, once the pride of economics, may have helped create a world too complex for marginalist intuition to master.

That is a real worry, but perhaps not a hopeless one. If we look across asset pricing, movie dynamics, and literary history, a different picture emerges. Machine learning and high-dimensional modeling are not simply replacing human understanding. They are revealing new conceptual objects. Just as telescopes and microscopes disclosed new physical objects in the early modern world, computational models now disclose new relational objects: pricing kernels, heavy-tailed outcome regimes, temporal diagonals in literary space.

The old intuitions are not enough. But the new models may help us build new intuitions.

Asset Pricing: From Marginalism to High-Dimensional Valuation

Classical asset-pricing theory begins with a powerful marginalist intuition: investors require compensation for bearing risk. An asset’s expected return should depend on its exposure to systematic risks. CAPM gave us beta; later factor models added size, value, momentum, profitability, investment, and other variables.

These models are attractive because they are low-dimensional and intuitively graspable. A few named factors are supposed to explain many asset returns. That is the dream: a compact causal vocabulary.

But recent high-dimensional models suggest that the true pricing structure may not be compressible into five or six factors. The AIPT work we discussed begins with roughly 130 stock characteristics and then generates hundreds of thousands of nonlinear factors. These are used to approximate the market’s pricing kernel, or stochastic discount factor: the hidden valuation structure through which future payoffs are priced under uncertainty.

This is no longer a simple causal story in which one named factor explains one outcome. It is more algorithmic. The model works by probing a vast feature space and finding structure there. The individual factors may not be interpretable as ordinary causes. But the model still has an economic frame: markets price future payoffs, and that pricing process appears to be high-dimensional.

This is where Cowen’s story becomes interesting. Marginalism did not merely explain markets. It helped construct modern finance. Pricing theory made derivatives, securitization, and the decomposition of risk into tradable claims possible. But once those instruments proliferated, they helped create a financial world too complex for the original marginalist intuitions to command.

In short: marginalism may have helped build the world that now requires machine learning to map.

Thick Objects: Stocks, Brands, and Movies

At first glance, financial assets seem very different from movies or novels. A stock is an economic claim; a movie is a cultural artifact. But the distinction begins to blur once we treat both as “thick objects.”

A thick object is not exhausted by one or two measurable properties. It has many dimensions. It is embedded in social interpretation. Its value depends on history, expectation, reputation, performance, and future promise.

Scott Galloway’s remark captures this beautifully: “Stocks are like brands and that is they’re part promise and part performance.” [YouTube, Pivot, May 29, 2026.]

That is a sophisticated observation. A stock is not merely a claim on current earnings. It is a socially circulating judgment about a company’s future. The performance side includes revenue, earnings, margins, growth, debt, liquidity, volatility, and so forth. The promise side includes brand power, technological imagination, managerial credibility, founder charisma, regulatory risk, and collective belief about what the company may become.

That makes stocks resemble brands. A brand is not just a name, logo, or handle. It is a socially stabilized promise. It condenses past performance and future expectation into a recognizable entity.

And that is the bridge to movies. A movie before release is also part promise and part performance. The performance side includes director, stars, genre, budget, studio, distribution, trailers, and reviews. The promise side is what audiences imagine the movie will deliver: spectacle, prestige, emotional satisfaction, social participation, novelty, nostalgia.

So although movies and stocks are different objects, both are thick. Both circulate through populations. Both are interpreted under uncertainty. Both depend on the conversion of signals into expectation and expectation into value.

The analytic methods differ. Asset-pricing models estimate valuation structure. Movie models track social diffusion and outcome distributions. But to a first approximation, the underlying situation is structurally similar: complex objects move through fields of collective judgment.

Tuesday, April 28, 2026

On Method: Computational Compressibility in Complex Natural and Cultural Phenomena

New working paper. Title above, abstract, contents, and introduction below:

Academia.edu: https://www.academia.edu/166054951/On_Method_Computational_Compressibility_in_Complex_Natural_and_Cultural_Phenomena
ResearchGate: https://www.researchgate.net/publication/404263330_On_Method_Computational_Compressibility_in_Complex_Natural_and_Cultural_Phenomena

Abstract: Various machine learning techniques have been used to develop models of complex systems from empirical data. Through discussions with Claude, this paper examines several examples, including: weather, protein folding, chess, language, asset pricing, ticket sales for movies, the 19th century English-language novel. These models differ from one another in various ways, but all are fundamentally descriptive in character. Explanations must necessarily reside with their respective disciplines. In some cases we already have fundamental accounts of the phenomena, while in other cases we do not. With respect to economics in particular, it is clear that such models reveal phenomena for which no explanations are currently available, presenting a challenge to economic theory.

Contents 

Part I: Computational Compressibility, Implications for Economics, Description and Explanation 5
Part II: Weather, Protein Folding, Chess, and Language 16
Part III: Interim Summary: Compressibility Without Reducibility 26
Part IV: Pricing Theory, Movies, 19th Century Novels, and Cultural Evolution 28
Part V: To Infinity and Beyond! – Hollywood Redux, Blockbusters, the Spreadsheet, Economics Going Forward 38

Introduction: Describing Computationally Compressible Systems 

This a transcription of a dialog I had with Claude 4.6 and 4.7 on April 21 - 23, 2026. While I started it with a specific case from Chapter 4 of Tyler Cowen’s recent monograph on marginalism ( The Marginal Revolution: Rise and Decline, and the Pending AI Revolution), now that the dialog has concluded with Chomsky’s distinction between descriptive and explanatory adequacy (Aspects of the Theory of Syntax), I realize that I’ve been thinking about the underlying issues for some time. While I read Aspects in about 1970, give or take a year, I didn’t think much about description as such until the 2000s, and then I was thinking about describing individual texts; but that’s not directly relevant to these cases in this paper. Then in the second decade of this century I began thinking about computational criticism, aka digital humanities, which typically involve some kind of statistical or machine learning investigation of a corpus of texts. In particular, I gave a great deal of attention to Macroanalysis (2013), where Matthew Jockers studied a corpus of roughly 3000 English-language novels published in the 19th century. That investigation culminated in a directed graph showing depicting relationships of close-similarity among the novels in the corpus. I decided that that graph, in effect, was fundamentally descriptive in character, depicting, in effect, the 19th century Anglophone Geist, or Spirit.

But Jockers’s graph wasn’t on my mind when I started my dialog with Claude. Rather, I was thinking about the distinction between computationally reducible and irreducible phenomena that Stephen Wolfram had introduced in his New Kind of Science (2001). As Claude notes in its summary of the dialog, “a reducible system admits shortcuts through its dynamics; an irreducible one must be simulated step by step.” My target was a paper about asset pricing that Cowen discussed in his monograph, which produced a model having 360,000 parameters but which defied intuitive understanding.

The weather is a canonical example of phenomenon that is computationally irreducible. Thus forecasting the weather generally involves running a simulation of the weather and stepping through it interval after interval. This requires enormous computing resources and takes time. But DeepMind has created a machine learning system, GraphCast, that abstracts over historical data in a way that allows more accurate forecasts with less compute. Thus the weather system is computationally irreducible, but it is also compressible.

I take that as my paradigm case of computational compressibility (pp. 16 ff.) and then move on to other examples: protein folding (another physical phenomenon, pp. 18 ff.), chess (human activity, pp. 22 ff.), and natural language (a different human activity, pp. 23 ff.), each of which is compressible using machine learning techniques. Each example sharpens and extends the idea of computational compressibility. At that point I asked Claude to summarize the discussion (pp. 26 ff)..

Then, and only then do I ask Claude to consider Cowen’s problematic example, AI Pricing Theory (pp. 28 ff.). In its analysis of the paper, Claude notice that it introduces something fundamentally new to the discussion, reflexivity. Asset pricing is done by a large group of actors over time who thus influence one another’s decisions. And that, in turn, brings up Arthur De Vany’s work on Hollywood Economics (pp. 30 ff.). De Vany discovered that box-office success cannot be predicted by such analytic variables as producer, screen writer, director, movie stars, or opening weekend box office. Rather the success of a film depends on a word-of-mouth cascade which cannot be predicted. That leads me, in turn, to suggest a thought experiment involve a hypothetical system capable to abstracting over entire films and developing a high-dimensional model which could be used to predict the success of individual films.

And that, in turn, led me to the work that Matthew Jockers had done on 19th century English-language novels (pp. 33 ff.), something that had not been on my mind when I began this dialog on April 21. Jockers used machine learning, albeit nothing so elaborate and computationally expensive as using a transformer to create an LLM – it only had roughly 600 parameters. What his model revealed, and what made it so fascinating to me, is that there is an inherent directionality to the production of novels over the course of a century. It’s not simply that later novels are systematically different from earlier ones, but that that difference has a direction in the 600-dimensional measurement space. What we’d really like to know, now, is a say to characterize that diction. The model shows us that there is a direction, but it doesn’t tell us what that direction is. Though the model is much simpler than that asset pricing model – it has three orders of magnitude fewer parameters – its significance is no more legible.

After that I have two discussions that are not based on existing models, but that do have implications for economists who want to study them. First, I consider the phenomenon of the blockbuster, arguing that it reveals audience preferences that had previously been unrecognized (pp. 41 ff.). Then I consider the spreadsheet (e.g. VisiCalc), which transformed the personal computer market from a small niche market into a large mainstream market (pp. 43 ff.). How do you create a model that allows you to predict markets that don’t even exist at the time you make your model? What kind of a problem is that? After that I took a brief look at Cowen’s argument in The Great Stagnation (pp. 45 ff.), where Claude remarked:

If the VisiCalc model is right, then what matters about ChatGPT and its successors is not primarily that they do existing things faster or cheaper—though they do—but whether they are constitutive technologies in the VisiCalc sense. Do they reorganize the space of possible wants, making new activities imaginable and practical that previously had no well-formed representation in anyone's preference space? With that I brought the exploration to a halt.

I then asked Claude to summarize the entire dialog, which I’ve placed immediately following these remarks (pp. XX ff), with a special emphasis on implications for economics (pp. 7 ff.). Then I introduce Chomsky’s distinction from the 1960s, description vs. explanation (pp. 9 ff.). Each of these cases involves a complex phenomenon that is irreducible, but can be compressed into a model that is descriptive in character. They have that in common. As for explanations, those must necessarily be specific to each phenomenon. Note that in some cases we have explanatory theories grounded in a fundamental understanding of the underlying system (weather, protein folding) while in others we do not (chess, asset pricing, cultural evolution).

Finally, I’ve added a coda from a different conversation with Claude (pp. 13 ff.), one I had with the AI that accompanied Cowen’s book. That conversation is about Hollywood Economics and Rational Ritual and argues that the factoring of intellectual space that we’ve inherited from the 19th century German university has outgrown its usefulness.

Monday, March 3, 2025

Confabulation, Dylan’s epistemic stance, and progress in the arts: “I’ll let you be in my dreams of I can be in yours.”

I continue to think about the tendency of LLMs to confabulate, that is, to make stuff up that is simply not true of the world. As I have remarked here and there, I tend to think that 1) confabulation is inherent in the architecture, and 2) that this “confabulation” is the default mode of human language. We just make things up.

However, we must live with one another and that requires cooperation. Effective communication requires agreement. It turns out that the external world is a convenient locus for that agreement. We agree that THAT tree over is a pine, that THAT apple is ripe, that THAT bird is a cardinal, that the stew is too salty, the earth is round and that the moon travels around the earth every 28 days. Some of these agreements may come easily, others are more difficult in the making.evolu

However true that may be, it does seem a bit odd to think of external reality as a vehicle for grounding agreement on language use. And, if I thought about it a bit, I could probably come up with some account of why that doesn’t seem quite right. But I’m stalking a different beast at the moment.

Consider this observation that Weston La Barre made in 1972 in The Ghost Dance: The Origins of Religion (p. 60):

... the Australian Bushman themselves equate dream-time with the myth-time that is mysteriously brought back in ritual; myth is as timeless as the unconscious mind. It is the delectability of dreams that makes them desirable, and it is their desirability that (along with lowered critical threshold) that gives them their intense “reality” and conviction. The fact that he dreams first force on man the need to epistemologize.

Let’s pretend we’ve just gotten up from a dream. What do we make of it? Or perhaps, how do we make anything of it?

We tell it to others, and they tell us their dreams. We talk and arrive at some agreement about what transpires in those dreams, about the nature of those events and their relationship to the waking world. I am thus reminded of Bob Dylan’s line in “Talkin’ World War III Blues”: “I’ll let you be in my dreams of I can be in yours.”

Is that what expressive culture is about, religious myth, ritual, folktales, drama, epics, novels, not to mention the visual and plastic arts as well? Are they vehicles for allowing us to reach agreement about certain of our, shall we say, confabulations?

This brings up a post from February, What is it about depicting sex in writing? From shared knowledge to common knowledge. Using three novels as examples (Pride and Prejudice, Tess of the D’Urbervilles, and Lady Chatterly’s Lover) I suggested that as we move from one to the other in historical order, we’re moving from shared knowledge of sexual matters to common knowledge. Those are terms from game theory. We talk of shared knowledge when everyone knows something, but aren’t aware that the others know that same thing. Once everyone comes to know that everyone knows that thing, whatever it is, that knowledge has become common knowledge.

I have thus begun to wonder: Is that how the arts function over time? I recently discussed the matter with Claude 3.7. Here’s its response:

It's not just that artistic styles change, but that art gradually makes explicit what was once implicit across many domains of human experience.

This pattern extends beyond sexuality. Consider how literature has treated:

  • Mental illness (from metaphorical to explicit)
  • Social class tensions (from coded to direct)
  • Religious doubt (from allegorical to overt)
  • Political criticism (from veiled to unmistakable)

In each case, what begins as shared knowledge eventually becomes common knowledge through artistic expression. This does suggest a kind of progress - not necessarily aesthetic progress, but progress in cultural honesty and transparency.

It is by no means obvious that this is the case. But I’m certainly sympathetic.

What is undeniable is that in the last 15 years or so digital humanists are found empirical evidence of unidirectional trends measured attributed of literary texts over long periods of time. Perhaps the most interesting example is in Matthew Jocker’s Macroanalysis, where he shows a unidirectional trend in a corpus of 3000 Anglophone novels from the 19th century. I discuss this in a number of posts. This working paper might be the place to start: On the direction of literary history: How should we interpret that 3300 node graph in Macroanalysis? There’s another working paper: On the Direction of 19th Century Poetic Style, Underwood and Sellers 2015. You might also look at this blog post from 2016, From Telling to Showing, by the Numbers, which is also about 19th century novels.

More later.

Thursday, May 5, 2022

Cohort succession and literary change [ --> Macroanalysis]

Note:  You should read the whole tweet stream, not just the first tweet. Underwood has another tweet stream about this article, noting "sometimes loose ends or new puzzles are more interesting to other researchers. So: short thread about those."

Abstract of the linked article: Many aspects of behavior are guided by dispositions that are relatively durable once formed. Political opinions and phonology, for instance, change largely through cohort succession. But evidence for cohort effects has been scarce in artistic and intellectual history; researchers in those fields more commonly explain change as an immediate response to recent innovations and events. We test these conflicting theories of change in a corpus of 10,830 works of fiction from 1880 to 1999 and find that slightly more than half (54.7 percent) of the variance explained by time is explained better by an author’s year of birth than by a book’s year of publication. Writing practices do change across an author’s career. But the pace of change declines steeply with age. This finding suggests that existing histories of literary culture have a large blind spot: the early experiences that form cohorts are pivotal but leave few traces in the historical record. 

* * * * *

The first diagram in this post is about cohort succesion: The Direction of Cultural Evolution, Macroanalysis at 3 Quarks Daily. That post is about Matthew Jockers's Macroanalysis. For further context, see mt working paper, On the Direction of Cultural Evolution: Lessons from the 19th Century Anglophone Novel, April 2015.

Sunday, December 15, 2019

Some informal remarks on Jockers’ 3300 node graph: Part 3, Signs and mechanisms

I want to start at the point where I ended my previous post in this series, Some informal remarks on Jockers’ 3300 node graph: Part 2, structure and computational process. I was arguing that the issue of scale was misconceived. The fundamental issue is NOT a corpus of texts versus one or a handful of texts. That’s trivial. The issue has to do with the terms of analysis. So-called close reading exists within a conceptual and methodological discourse which is quite different from that of so-called distant reading.

The difference is one of conceptual ontology, as that term has come to be understood in computer science and the cognitive sciences. In that context the issue is not the ultimately real, a philosophical question, but the kind of concepts we are we using. I start with a simple example, salt versus sodium chloride and then build out from there to words versus signifiers and then on to texts and meaning.

Conceptual ontology

Consider for a moment, salt on the one hand and NaCl (sodium Chloride) on the other. Physically they are (almost) the same substance, but conceptually they are quite different. Salt is defined and understood in terms of its physical appearance and, above all else, its taste. We can taste salt even when we cannot see it, and so can animals. NaCl, however, is defined in terms of an atomic theory of matter that didn’t exist until early in the 19th century. Moreover, NaCl is a pure substance, consisting of nothing by sodium and chlorine atoms; salt on the other hand will always have some impurities. Thus, strictly speaking, salt and NaCl are not physically the same, close, but not exactly the same.

Salt and sodium chloride, then, are used in different intellectual contexts, each of which has its own vocabulary. Salt belongs with sugar, pepper, cinnamon, flour, and so forth all substances having to do with food, food preparation, and eating all related concepts. Sodium chloride is related to potassium chloride, sodium hydroxide, and so forth, electrolysis, ion-exchange, and so forth, for a long list. Thus we have two different conceptual worlds, each coupled to characteristic actions and processes, but both ultimately grounded in the same physical reality. Someone whose occupation has them working in the sodium chloride world has no trouble with table salt at mealtime. The transition from one world to the other is seamless, or nearly so (there may be a change of clothes involved).

The worlds of “distant reading” and “close reading” differ from one another in the same way. The physical texts and the symbols imprinted on them are the same in both worlds, but the concepts and methods of description and analysis are different. It’s that (kind of) different that Geoffrey Hartman had in mind almost a half century ago when he observed: “modern ‘rithmatics’—semiotics, linguistics, and technical structuralism—are not the solution. They widen, if anything, the rift between reading and writing” (The Fate of Reading, Chicago 1973, p. 272). He expresses himself using the standard trope of distance, but he’s certainly not talking about distance in any physical sense – no one is.

These days we might want to gloss distance in terms of explicit mediating steps. In close-reading you read a primary text and then does some thinking, perhaps some secondary reading and research, and then you write about that primary text. Just how you break that down is somewhat arbitrary, but it’s nothing like what happens in distant-reading, which starts when a collection of primary texts is digitized. The digitized texts are then cleaned up, tagged with metadata, organized in a database or databases, and then subject to analysis, which might be a straightforward matter of word counts, or the somewhat more strenuous process of topic modeling, or perhaps we’re going to make use of vector semantics, or any of a number of things. This more complicated process is likely to involve two, three, or more people at various stages along with way. At some point the analytic process will produce some set of visualizations and they, in turn, will be interpreted in terms appropriate to the primary texts (and their contexts).

As I said, that’s how we might gloss the close vs. distant distinction these days. The distance isn’t one of physical steps, but operational steps. However, it’s doubtful that Hartman had anything like this in mind. To be sure, stylometrics existed in 1973, and Stanley Fish was doing his best to skewer it, but I doubt that Hartman was thinking about stylometrics when he made that remark. However, he might well have been thinking about the kinds of tables and diagrams that Lévi-Strauss used, or that showed in linguistics article, those diagrammatic signs that interrupted the linguistic flow. They are indices of a different mode of thought, a different conceptual ontology.

Words and signifiers

We can begin to appreciate that difference by noting the difference between words, which are signs in Saussure’s sense, and signifiers, which are components of signs. Literary critics generally talk of words, and occasionally of signifiers. The concept of word is transparent enough in ordinary casual discourse. As ordinarily understood, the concept of words encompasses pronunciation, spelling, grammatical usage, and meaning, often several meanings. Words in this sense are the things listed in dictionaries. And for the most part, literary critics deal in words, even those who’ve read a bit of Saussure.

Saussure distinguished between the signifier and the signified. The signifier is a physical entity, either sonic or visual. The signfied is a mental entity; it is the bearer of meaning. Signifiers are public; they can be transmitted between people. Meanings are, well, they’re not exactly private – Wittgenstein established that in a famous section in his Investigations – but as mental objects they exist in people’s heads and we do not have direct access to one another’s heads. Thus, as William Croft has argued in chapter 4 of Explaining Language Change [1], word meanings are negotiated in conversational interaction with one another.

Computational critics, in contrast, may talk of words, but what they are actually working with are signifiers, signifiers rendered in digital form. That is to say, that’s what’s in the databases at the heart of computational work, mere signifiers. The computer has no access to word meanings, to signifieds; it only knows the written signifier. Consider the following passage from the well-known 1949 memo on machine translation written by Warren Weaver, who headed the Natural Sciences Division of the Rockefeller Foundation:
First, let us think of a way in which the problem of multiple meaning can, in principle at least, be solved. If one examines the words in a book, one at a time as through an opaque mask with a hole in it one word wide, then it is obviously impossible to determine, one at a time, the meaning of the words. "Fast" may mean "rapid"; or it may mean "motionless"; and there is no way of telling which.

But if one lengthens the slit in the opaque mask, until one can see not only the central word in question, but also say N words on either side, then if N is large enough one can unambiguously decide the meaning of the central word. The formal truth of this statement becomes clear when one mentions that the middle word of a whole article or a whole book is unambiguous if one has read the whole article or book, providing of course that the article or book is sufficiently well written to communicate at all.

The practical question is, what minimum value of N will, at least in a tolerable fraction of cases, lead to the correct choice of meaning for the central word?
While that memo catalyzed early work on machine translation, the approach suggested in those paragraphs played no role in that work. The necessary computing power wasn’t available. That changed in the 1980s and especially 1990s.

The topic analysis technique that Jockers used follows from the insight expressed in those paragraphs, as does the work on vector semantics. The idea is simple. Words that appear together frequently must somehow have related meanings. Race, horse, saddle, jockey, and track, all mean different things, but they belong to the same discourse and so will appear in close proximity when that discourse is spoken, or written. And so it is with human, race, IQ, identity, American, black, and white, these too belong to a discourse and so will co-occur in texts using that discourse. Notice that race is common to those two very different discourses. Considered alone, without context, the meaning of race is ambiguous, we don’t know what it means. But if we see race in close proximity to horse, we’ll locate it in one discourse while if we see it in proximity to IQ we’ll locate it in that other discourse.

Signifiers in themselves have no meaning. But in actual usage signifiers are always bound to a meaning and that meaning connects them to related signifiers. Given a large enough body of texts, computers can determine which signifiers occur together. It is up to the investigator to make a judgment about why those signifiers are occurring together. The investigator will, of course, make that judgment on the basis of their knowledge of the language.

The conventional literary critic finds such judgments utterly trivial and has no way, no intellectual context, for understanding how remarkable it is that mere computation can discover such relationships. And since that critic is only interested in a handful of texts they have no reason to reconsider that judgment and consider the possibility that there is something to be learned in examining such patterns in collections of texts, such as the collection of 3346 Anglophone novels Jockers has been working with.

Text and meaning

But what is a text? For the linguist, or for the computational critic, and answer is simple: a text is a string of characters.

Things are not so simple for the literary critic. Yes, sure, a text is a physical object, a scroll or a codex, and it is a bunch of signifiers inscribed on such. That’s trivial and, for the most part, uninteresting except in very well defined contexts, such as physical preservation or, more interestingly, the preparation of an edition. And the preparation of critical editions has been very important in the development of digital humanities broadly conceived, but still, we can bracket that.

What I am interested in here is the text as an object of interpretation. THAT text is an enigma [2]. The following passage is from the introduction Rita Copeland and Frances Ferguson prepared for five essays from the 2012 English Institute devoted to the text:
Yet with the conceptual breadth that has come to characterize notions of text and textuality, literary criticism has found itself at a confluence of disciplines, including linguistics, anthropology, history, politics, and law. Thus, for example, notions of cultural text and social text have placed literary study in productive dialogue with fields in the social sciences. Moreover, text has come to stand for different and often contradictory things: linguistic data for philology; the unfolding “real time” of interaction for sociolinguistics; the problems of copy-text and markup in editorial theory; the objectified written work (“verbal icon”) for New Criticism; in some versions of poststructuralism the horizons of language that overcome the closure of the work; in theater studies the other of performance, ambiguously artifact and event. “Text” has been the subject of venerable traditions of scholarship centered on the establishment and critique of scriptural authority as well as the classical heritage. In the modern world it figures anew in the regulation of intellectual property. Has text become, or was it always, an ideal, immaterial object, a conceptual site for the investigation of knowledge, ownership and propriety, or authority? If so, what then is, or ever was, a “material” text? What institutions, linguistic procedures, commentary forms, and interpretive protocols stabilize text as an object of study?
What? “Linguistic data” sounds like it might be mere signifiers, and perhaps “copy-text” as well. But the rest of them, those various sites “for the investigation of knowledge, ownership and propriety, or authority”, those texts clearly consist of words, in the ordinary sense, complete with their multiple and often ambiguous and contradictory meanings.

That is the text that is conceptualized in vague spatial metaphors. Do we talk of meaning as being IN the text or as being somehow outside the text, in the CONTEXT? I think it would be a mistake to interpret that as a vague assertion about mechanism, though I am sorely tempted to do so. Rather, it’s a statement about critical methodology. If you think of meaning as IN the text, then you do not invoke anything but the text itself in the process of interpreting it. That’s what was originally meant by close reading. But if you mean that meaning is located in the context, well then you must bring other evidence to bear in your close reading. Depending on your particular methodology you may call on historical materials of various kinds – newspapers and other contemporary periodicals, correspondence, legal documents, and so forth – or you may want to mount a psychological argument in terms of this or that theory, diaries, and what have you. And of course the text itself.

And then we have the idea of “hidden” meaning. That’s clearly a spatial metaphor. But a metaphor for what? Not the text itself; it’s not as though anyone is imagining a secret text bound into the spine or sandwiched into the front or back boards. The idea seems to be something like a large and elaborate old house with hidden rooms and passages. Expressed in that way it seems just a bit foolish, but then no one ever expresses it in that way, do they? Again, we’re dealing with a covert methodological injunction. Words have many meanings and are linked to other meanings through various figures, such as the master figures of metaphor and metonymy – need I invoke, for example, Roman Jakobson’s work on the subject [4]? And so, as a methodological precept, one must explore those possibilities.

The fact of the matter is that, by contemporary standards, those of the so-called cognitive revolution and after, literary criticism lacks an account of language mechanism, whether in the reader, the author, or the critic. It is non-mechanistic or pre-mechanistic and certainly non-computational – an issue I’ll take up in the next and last post in this series. For now I want to end on the assertion that meaning, as the object of critical investigation, is inherently and irreducibly subjective. That doesn’t mean that critics cannot agree on the meaning of texts, for subjectivity necessarily implies intersubjectivity. It means only that meaning resides in subjects, in readers, writers, and critics and is not subject to objective determination. It is thus like color, which is also subjective in that sense. Sensations of color may be closely related to wavelengths of light, but cannot simply be reduced to them. Color arises within subjects, though psychologists have made a great deal of progress in figuring out how color perception functions in humans and animals.

The objective study of meaning, that is, of semantic mechanisms, has not yet been so successful. It was the search for semantic mechanisms that lead me to David Hays and computational linguistics, which I discussed in the previous post in this series, Some informal remarks on Jockers’ 3300 node graph: Part 2, structure and computational process.

References

[1] Croft, William (2000). Explaining Language Change: An Evolutionary Approach. Longman.

[2] I’ve written a number of posts at New Savanna about the concept of the text. They’re at this link: http://new-savanna.blogspot.com/search/label/text.

[3] Rita Copeland and Frances Ferguson, “Introduction”, ELH, Volume 81, Number 2, Summer 2014, p. 417.

[4] For example, Roman Jakobson and Morris Halle, Two Aspects of Language and Two Types of Aphasic Disturbances in Fundamentals of Language, The Hague & Paris: Mouton, 1956.

[5] I’ve got a good many posts on color, https://new-savanna.blogspot.com/search/label/color. They are of various kinds. Some are simply sets photographs where I was interested in colors and color contrasts. But others are about color perception in the mind and in photography. If you are interested in this topic you might what to search the web on “color perception”.

Thursday, December 5, 2019

Some informal remarks on Jockers’ 3300 node graph: Part 1, “a generic time trend” [#DH]

For some time I’ve thought my intellectual mission was to examine “stuff” that has been developed through largely discursive and informal methods and bring it to a form where real mathematics, of an appropriate kind, can be employed to gain further insight. I note that I have little mathematical training beyond high school, but I have developed sophisticated intuitions through reading and through interacting with thinkers having more mathematical knowledge than I have. My interests take me toward literature, art, music and the like, while my intellectual predilections move toward the mathematical. And so I find myself enacting the role of a bridge.

That’s what I’ve been doing with the 3300 node graph from chapter 9 of Matt Jockers’ Macroanalysis (2013). Here it is [1]:


In this post and the next I’d like to make some informal observations about how I think about Jockers’ 3300 node graph. This post is about time and evolution while the next one will be about structure and computation.

Generic time trends

In dismissing that graph, Nan Da said that Jocker had simply given us “a generic time trend” and, as supporting evidence, she informs us [2]:
If you take a similar dataset (texts over one hundred years) and regress absolute Euclidean distances (based on similarly determined features) on absolute distances in time, you will see super significant positive correlation.
But she doesn’t tell us what that “similar dataset” is. You have to go to the online material to find out. It turns out that it’s the set of journal articles that Goldstone and Underwood used for their study of academic literary criticism. So, it’s a different kind of text, expository prose vs. fiction, and a different century, 20th vs. 19th. But still, what you see is that texts that are close together in time are also highly similar according to the same kind of metric – assuming Da did her work well in this case, and I’m willing to grant her that.

I think she’s right about this being a “generic” time trend. Where I differ is that I don’t for a minute think we should take it for granted. Students of the nineteenth century novel – can’t say that I’m one, but I’ve read a bunch of them – will have read many texts. I rather suspect they believe (if only tacitly, without deliberate conscious reflection) that there is an element of necessity, if you will, in the order in which they were written and published, at least by the decade if not by the exact year. They’d be very surprised to see something like Pride and Prejudice coming out in 1893 and something like The Adventures of Huckleberry Finn coming out in 1829. That’s just not how it works. Similarly, no one was writing like Kenneth Burke in 1993 and critics would be deeply puzzled to find Jameson texts in 1910. That’s just not how it works.

In the case of critical texts, some critics might be willing to admit to, you know, intellectual “progress”, and not worry too much about being charged with promoting a Whig view of the profession’s history. But progress would be much more problematic in the case of those novels. Does Tess of the D’Urbervilles represent some kind of progress over The Heart of Midlothian? We can drop the word, “progress”, but the directionality remains to be accounted for. Before we can begin to account for it, though. we need to tease these texts away from the familiar comfortable “face” they present to us, the one we embed in Da’s notion of a generic time trend. And Jockers’ graph does that rather nicely, especially if we keep firmly in mind that each node ‘encapsulates’ a fairly rich – if of an odd an unfamiliar kind – representation of a text.

So that’s one thing.

Evolution

Back when I first read Macroanalysis, and after I’d had a chance to think about that graph a bit, I thought, “the rough temporal ordering of those representations (of texts) is evidence of an underlying evolutionary process: evolution, descent with modification.” And I set about rationalizing that original intuitive judgment. That process of rationalization, of course, is ongoing.

Now, I have a long standing interest in evolution, first biological evolution, and then, a bit later, cultural evolution. I’m sure this goes back to my undergraduate years at Johns Hopkins, where I also became interested in “the arrow of time”. It seems that many/most physical laws are time symmetrical; they work the same regardless of the direction of time. Thus, while you can use Newtonian mechanisms to predict future positions of the planets, you can also use them to recover past positions. But nonetheless the universe does seem to have a temporal direction. Why? That takes us to thermodynamics and entropy, which is sometimes thought to be about increasing disorder but, my scientist friend, Tim Perper, tells me that’s not quite it. Entropy is about the irreversibility of a process [3].
When I say that I see Jockers’ graph as a trace of the activity of a complex dynamical system, all I’ve got is a metaphor, but it’s a carefully considered metaphor. And it’s that metaphor that tells me that Da’s “generic time trend” is something that we should not take for granted.
And one thing that makes biological evolution so interesting is that it seems to be working against entropy, which it can do because the earth gets free energy from the sun which organisms capture and use in various ways. So the direction of time is a topic of direct relevance to biological evolution. Biologists have been very reluctant to countenance the idea that later life forms are somehow more complex than earlier ones mostly, I suspect, because they fear such ideas are just too close to the good old Chain of Being in a version that heads up beyond humans to the angels and finally, at the very top, the supreme deity. But it need not go there, not if you’re careful. So some years ago David Hays and I published a short article, A Note on Why Natural Selection Leads to Complexity [4]. We argue that evolution is able to produce ever more complex organisms because increased capacity for information processing is energetically cheap in relation to the increased energy capture it enables.

That seems rather remote from a unilinear direction in the trajectory of the 19th century novel. And it is. But this is informal, I'm just thinking out loud.

Is culture a complex dynamical system?

Still, I’ve got reasons. When I was working on my book on music, Beethoven’s Anvil (2001), I was in touch with and strongly influenced by the late Walter Freeman. He’d pioneered the use of complexity theory in modeling the dynamics of neural nets, real neural nets in animal brains, not the artificial neural nets of contemporary AI and machine learning. That’s pretty much the same mathematics the physicists use in statistical thermodynamics. Where they’re creating high-dimensional models with individual molecules (air, water, whatever) as the elements (a dimension [actually 6] for each individual molecule], Freeman created high-dimensional models where individual neurons are the elements (a dimension for each neuron). In the second and third chapters of that music book I extended this thinking to groups of people engaged in making music together. I didn’t actually do any math – which is beyond my technical capabilities – but I made a careful step-by-step argument. Though construction is a better word. I constructed a way of thinking about coupled brains as a single dynamical system [5].

It’s still some distance from a small band singing and dancing together to millions of people reading English language novels in the 19th century. But at least we’re now in the world of humans and human culture. So, when I say that I see Jockers’ graph as a trace of the activity of a complex dynamical system, all I’ve got is a metaphor, but it’s a carefully considered metaphor. And it’s that metaphor that tells me that Da’s “generic time trend” is something that we should not take for granted. It is something that must be explained.

Just what form that explanation will take, that’s not all obvious to me. I think a lot of construction is going to have to take place to do the job. On the one hand, we need to “open up” that graph and get a handle on what features seem to be driving the temporal trend. That’s one kind of intellectual work. Figuring out just WHY that happens, that’s something else. But surely we can make progress on figuring out just what has happened without having to work out the mechanisms driving it, no?


References

[1] For an explanation of the graph you can of course read Jockers’ book. In fact you should read the book, because understanding the earlier chapters will give you a richer understanding of the graph. For a refresher that doesn’t depend directly on the book see, for example, my post, Notes toward a theory of the corpus, Part 1: History [#DH], New Savanna, May 9, 2019, https://new-savanna.blogspot.com/2018/09/notes-toward-theory-of-corpus-part-1.html.

[2] Nan Z. Da, The Computational Case against Computational Literary Studies, Critical Inquiry 45, Spring 2019, 601-639.

[3] I’ve written an informal little paper on this, complete with photos, A Primer on Self-Organization: With some tabletop physics you can do at home, 2014, https://www.academia.edu/6238739/A_Primer_on_Self-Organization.

[4] William Benzon and David G. Hays, A Note on Why Natural Selection Leads to Complexity, Journal of Social and Biological Structures 1990, https://www.academia.edu/8488872/A_Note_on_Why_Natural_Selection_Leads_to_Complexity.

[5] William Benzon, Beethoven’s Anvil: Music in Mind and Culture, Basic Books 2001. You can download the final drafts of the second and third chapters here, https://www.academia.edu/232642/Beethovens_Anvil_Music_in_Mind_and_Culture. That’s where I do most, but not all, of the constructing.

Wednesday, November 27, 2019

Divergence and Reticulation in Cultural Evolution: Some draft text for an article in progress [#DH]

That's the title of my latest working paper. You can download it here: https://www.academia.edu/41095277/Divergence_and_Reticulation_in_Cultural_Evolution_Some_draft_text_for_an_article_in_progress.

And you can participate in a discussion of it here: https://www.academia.edu/s/9b97738023.

Abstract, Contents, and introductory material below.

* * * * *

Abstract: In a recent review of articles in computational criticism Franco Moretti and Oleg Sobchuk bring up the issue of tree-like (dendriform) vs. reticular phylogenies in biology and pose the question for the form taken by the evolution of cultural objects: How is cultural information transmitted, vertically (leading to trees) or horizontally (yielding webs)? Dendriform phylogenies are particularly interesting because one can infer the phylogenetic history of an ensemble of species by examining the current state. The horizontal transmission of information in webs obscures any historical signal. I examine a few cultural examples in some detail, including jazz styles and natural language, and then take up the 3300 node graph Matthew Jockers (Macroanalysis 2013) used to depict similarity relationships between 3300 19th century Anglophone novels. The graph depicts a web-like mesh of texts but, uncharacteristically of such patterns, also exhibits a strong historical signal. (Just how that is possible is the subject of another draft.)

Contents

What’s Up? 1
The need for theory: Cultural evolution 2
Trees, Nets, and Inheritance in Biology 4
Divergence and reticulation in culture 7
What kind of objects are we dealing with? 12
Jockers’ Graph, a reticulate network 18
Appendix: A quick guide to cultural evolution 22
 
What’s Up?

In the past year we have had two reviews of recent work in computational criticism:
Nan Z. Da, The Computational Case against Computational Literary Studies, Critical Inquiry 45, Spring 2019, 601-639.

Franco Moretti and Oleg Sobchuk, Hidden in Plain Sight: Data Visualization in the Humanities, New Left Review 118, July August 2019, 86-119.
Though both are critical of that work, they are quite different in tone and intent. Da is broadly dismissive and sees little value in it. Moretti and Sobchuk see considerable value in the work, but are disappointed that it is largely empirical in character, failing to articulate a theoretical superstructure that deepens our understanding of literary history.

I’ve been working on a critique of those papers which seems to have expanded into a primer on thinking about literary culture as an evolutionary phenomenon. I’m currently imaging that the final article will have five parts:
  1. Genealogy in literary history
  2. Unidirectional trends in cultural evolution
  3. Jockers’ Graph: Direction in the 19th century Anglophone novel
  4. Expressive culture as a force in history
  5. A quick guide to cultural evolution for humanists
I have already posted draft material for the second part of the article, which centers on a graph from Matthew Jockers’ Macroanalysis (2013) [1].

That graph was my central concern from the beginning. It is the most interesting conceptual object I’ve seen in computational criticism, but it is easily misunderestimated and glossed over – as far as I know Da’s understandable but unfortunate dismissal is the only treatment of it in the referred literature. The problem, it seems to me, is that a proper appreciation of it requires a conceptual framework that doesn’t exist in the literature. My objective, then, is to begin assembling such a framework.

Moretti and Sobchuk didn’t mention it at all as their review was confined to journal articles. But it merits consideration in a framework that did establish in their review, if only barely. The invoke a distinction from evolutionary biology, that between tree-like (dendriform) phylogenies and free-form or web-like phylogenies, and suggest that it is important for understanding the relationship between literary for and history (pp. 108 ff.). Jockers graph is web-like network of texts but it exhibits an important feature of dendriform phylogenies, it displays a strong temporal signal. Thus a discussion of issues raised by Moretti and Sobchuk is a good way to begin constructing the missing conceptual framework.

This document consists of draft material for the discussion, the first part of the planned article, and the fifth part. The fifth part, the appendix is straight forward, and I have included it the end of this document. Once I have discussed the issue of dendriform vs. web-like relationships I introduce Jockers’s graph.

In the second part of the article, unidirectional trends in cultural evolution, I plan to say a few words about time and directionality. I will then take up a number of the examples Moretti and Sobchuk review in their article. While they don’t frame them as evidence for unidirectional trends, that is what they are. From my point of view that’s the most interesting and important aspect of their review, they gather those articles into one place. I will be placing those articles in the context of other work showing unidirectional trends.

I don’t yet know whether I’ll post draft materials on the second and fourth sections before drafting the whole article.

The need for theory: Cultural evolution

Now let us turn to Moretti and Sobchuk. Here is their penultimate paragraph (112-113):
Tree-like, linear, reticulate . . . why should we even care about the shape of cultural history? We should, because that shape is implicitly a hypothesis about the forces that operate within history; the tentative, intuitive beginning of a theoretical framework. ‘Theories are, even more than laboratory instruments, the essential tools of the scientist’s trade’, wrote Thomas Kuhn over a half century ago; too bad we didn’t heed his advice. Although the crass anti-intellectualism of Wired—‘correlation is enough’, ‘the scientific method is obsolete’—has fortunately remained an exception, what seems to have happened is that, as the amount of quantitative evidence at our disposal was increasing, our attempts at in-depth explanations were losing their strength. Disclaimers, postponements, ad hoc reactions, false modesty, leaving inferences ‘for another day’ . . . such have been, far too often, our inconclusive conclusions.
Ah, “the forces that operate within history”, that’s what we’re after, no? And we’re not going to get there without theory, yes?

I believe that that theory will be about culture as an evolutionary phenomenon. It is clear that both Moretti and Sobchuk believe that as well, but they do not introduce or frame their essay that way. They introduce it as a methodological inquiry into the use of visualization. It is only as the essay unfolds that evolution emerges as an ideational engine parallel to if not quite driving their interest in visualization.

Accordingly it is necessary to make some preliminary remarks about cultural evolution. Work in cultural evolution has blossomed in the last quarter century but:
While humanities and social science scholars are interested in complex phenomena—often involving the interaction between behaviour rich in semantic information, networks of social interactions, material artefacts and persisting institutions—many prominent cultural evolutionary models focus on the evolution of a few select cultural traits, or traits that vary along a single dimension [...]. Moreover, when such models do build in more traits, these typically are taken to evolve independently of one another [...]. Within cultural evolutionary theory, this strategy holds that the dynamics and structure of cultural evolutionary phenomena can be extrapolated from models that represent a small number of cultural traits interacting in independent (or non-epistatic) processes. This kind of strategy licences the modelling of simple trait systems, either with an eye to describing the kinematics of those simple systems, or to illuminate the evolution and operation of mechanisms underpinning their transmission [...]. [2]
Hence, if students of literature want to think about culture as a phenomenon of evolutionary processes, we will not find suitable models and methods in existing work on cultural evolution. Though we certainly need to be aware of and conversant with that work, we are going to have to construct models and methods suitable to our material. That is the primary objective of this essay. To that end, then, I will be introducing a several of examples of work on cultural evolution in other domains.

Biologists, of course, has been developing evolutionary theory over the last half century. While they agree on basic issues, many details are still under contention. When we, then, as students of literary culture set out to adapt evolutionary theory to the analysis of literary phenomena, just what do we take from biological thinking and how do we do it? Various approaches exist in the general cultural literature, but this is hardly the place to sort through them – though I have prepared a brief appendix with pointers into those discussions. What Moretti and Sobchuk seem to have taken over is the distinction between tree-like lineages and more chaotic, network-like lineages. So that’s where I will start.

Where I am going, though, is toward an argument which says that that distinction is a reflection of the mechanisms that underlie the evolutionary process and it is to those mechanisms that we must look in adapting evolutionary theory to the study of human culture. Cultural evolution unfolds though collectivities of human minds, and they give cultural evolution a different texture, if you will, and different large scale patterns.

References

[1] On the direction of literary history: How should we interpret that 3300 node graph in Macroanalysis, Version 2, https://www.academia.edu/40550795/On_the_direction_of_literary_history_How_should_we_interpret_that_3300_node_graph_in_Macroanalysis_Version_2.

[2] Buskell, A., Enquist, M. & Jansson, F. A systems approach to cultural evolution. Palgrave Commun 5, 131 (2019) pp. 4-5, doi:10.1057/s41599-019-0343-5 https://rdcu.be/bVNtP.

* * * * *

Addendum 12.10.19: Cultural cross pollination is very old:
According to the Seshat team, the data also clearly undermine another of Jaspers’ key claims: that innovation arose independently in the five core societies, which he referred to as “islands of light”. These societies were engaged in a “ton of cross-cultural exchange,” says historian and Seshat project manager Daniel Hoyer at George Brown College in Toronto, Canada. “The Rabbinical tradition and even Plato’s writings aren’t really conceivable without the Zoroastrianism and Egyptian moral ideals and Hittite legalism that went before.”