Showing posts with label Jockers. Show all posts
Showing posts with label Jockers. Show all posts

Monday, June 29, 2026

Notes on the Collective Valuation of “Thick” Objects: Financial Assets, Movies, and Novels

New working paper. Title above, links, abstract, TOC, and introduction below.

Links:

Academia.edu: https://www.academia.edu/169390494/Notes_on_the_Collective_Valuation_of_Thick_Objects_Financial_Assets_Movies_and_Novels
ResearchGate: https://www.researchgate.net/publication/408219138_Notes_on_the_Collective_Valuation_of_Thick_Objects_Financial_Assets_Movies_and_Novels

Abstract: Machine learning is creating a methodological bridge between disciplines that previously seemed far apart, especially economics and literary criticism. The bridge is the analysis of how populations deal with “thick objects.” A thick object is not exhausted by a few visible traits. It gathers interpretation, expectation, memory, value, narrative, and social response. A toaster is usually a thin object. A firm that manufactures toasters is thick: it has assets, debt, brands, patents, management, supply chains, analyst coverage, market expectations, and future promises. Scott Galloway’s remark that stocks are like brands — part promise, part performance — links stock, movies and novels. Each is a thick object moving through a field of collective judgment. Its value reflects both measurable performance and imagined future promise. They are thus as neighboring cases in a general problem: how populations perceive, classify, value, and transform thick objects. Machine learning constructs object-spaces from the traces minds leave behind. The task now is to learn how to interpret those spaces without mistaking the model for the world.

High-dimensional asset-pricing models start with many stock characteristics — price, returns, volume, profitability, leverage, liquidity, analyst revisions, momentum, volatility, investment, and so on. These characteristics are traces of firm activity, accounting conventions, analyst judgment, and trader behavior. New models then generate hundreds of thousands of nonlinear transformations from those characteristics in order to approximate the market’s pricing kernel, the structure through which future payoffs are priced under uncertainty. The individual factors are analytic objects approximating the valuation geometry produced by collective market activity.

That sounds strange in economics, but it is familiar from Matthew Jockers’ work on nineteenth-century Anglophone novels. Jockers created a high-dimensional design space from thousands of novels, using stylistic features and topic models. His topics are not literal thoughts in anyone’s mind. They are model-derived approximations to recurrent regions of culturally circulating thought. Yet the model revealed historical direction: novels arranged by similarity formed a temporal diagonal, a computationally disciplined proxy for population-level cultural cognition.

Arthur De Vany’s model of Hollywood adds the dynamic bridge. Movies are thick expressive-market objects. Their success cannot be predicted simply from stars, director, budget, genre, or advertising. Once released, they enter an audience field where word of mouth, imitation, and nonlinear cascades determine their fate. Most fail, some profit, a few become blockbusters. The dynamics are heavy-tailed, interactive, and collective.

Contents

Introduction: Using ChatGPT for focused intellectual exploration across disciplines 3
Thick Objects: Ground Shared by Economics and Cultural Analysis [Summary] 9
AIPT, Large Factor Models [First Session] 17
Hollywood Economics 23
Macroanalysis 27
The emerging triad 30
Direction over time 31
Doing a Jockers style analysis for financial assets 38
Thinking about thick objects 40
Stocks are like brands [Session Two] 42
Algorithmic and Causal models [Session Three] 52
Those empirical APT models [Session Four] 56
Decision space 63
A bridge between disparate disciplines 67

Introduction: Using ChatGPT for focused intellectual exploration across disciplines

This document serves two purposes. It presents a specific argument leading to the following provisional formulation:

High-dimensional models of novels, movies, and assets disclose the population-level geometry of collective interpretation around thick objects, turning literary criticism and economics into neighboring sciences of modeled valuation.

How I arrived at the speculation, however, is as important as the idea itself, perhaps more so. I did not arrive at that idea unaided. ChatGPT helped me. Those aren’t my words; they’re ChatGPT’s. I know a great deal about literary criticism and about movies, but not much about economics. I need ChatGPT to bridge the conceptual distance between the humanities, literary criticism, and the social sciences, economics.

Methodological curiosity

Fortunately the peculiar circumstances of my career have forced me to be interested in method and epistemology: How is it that we can come to know about the world and what methods can we use to arrive at that knowledge? When I entered Johns Hopkins as a freshman in 1965 the discipline of literary criticism was in a state of crisis, though I didn’t know that. How could I? I’d only just graduated high school and I still pretty much knowledge as it was handed to me.

That soon changed. The details of just how, when, and why don’t matter much at the moment. That it happened is sufficient for my present purposes. The upshot is that I became interested in Coleridge’s “Kubla Khan” in my senior year. I investigated the poem with standard interpretive methods augmented by avant garde structuralism and found patterns I could not explain. But they “smelled” of the nested loops I learned about in a course in computer programming.

That sent me to the English Department at SUNY Buffalo, which had the best experimental program in the nation. I found a fellow graduate student, Ralph Henry Reese, who pointed me around a corner and down the hall to David Hays in Linguistics. Hays had been a first generation researcher in machine translation at the RAND Corp. and, as such, was one of the founders of computational linguistics. While I wasn’t able to resolve my issues with “Kubla Khan” – they’re still hanging fire – I became hooked on cognitive science. Consequently my dissertation in the English Department was also a quasi-technical exercise in knowledge representation, the discipline within cognitive science and artificial intelligence about the representation of human knowledge in computable form.

Given that that is where I had arrived in the late 1970s it is perhaps not so strange that now, decades later, I find myself staring down some pretty formidable economics despite never having studied the subject. For the last 15 years, however, I have been reading the Marginal Revolution blog hosted by Tyler Cowen and Alex Tabarrok and I have been reading my way through Cowen’s recent monograph, The Marginal Revolution: Rise and Decline, and the Pending AI Revolution (2026). Cowen’s theme in the fourth (and last) chapter is that the economics he was trained in, the economics which followed from the Marginal Revolution, is rapidly being eclipsed by a more determinedly empirical discipline based on machine learning.

Bombed by 360,000 factors

Here is Cowen’s premier example. It’s from something called Arbitrage Pricing Theory (APT) (pp. 99-100):

There is a recent working paper which is perhaps more striking yet, by Antoine Didisheim, Shikun (Barry) Ke, Bryan T. Kelly, and Semyon Malamud. They pick up from Arbitrage Pricing Theory (APT), a well-established idea from financial economics. APT typically looks for “factors” in the data which predict excess returns, and a traditional APT model might have found five or six such factors. Are “inflation” or perhaps “the term structure of interest rates” useful factors? Well, that can be debated, but if so, those results sound pretty intuitive. But those intuitions seem to be disappearing. In a paper by these authors, they apply machine learning methods to look for more factors. As we know, machine learning is very good at finding non-obvious relationships in the data. The largest model they built has 360,000 (!) factors, and it reduces pricing errors by 54.8 percent relative to the classic six-factor model from Fama and French. Bravo to the authors, but what kinds of intuitions do you think possibly can be supported by those 360,000 factors?

When I read that, it “looked like Greek to me,” as the cliché has it. But I took a deep breath and thought carefully, step by step and concluded that the assets in question are stocks. What you need to pay attention to is 1) the contrast between six factors and 360,000 factors, 2) the fact that one set of factors is intuitive while the other certainly is not, 3) but the unintelligible, unintuitive, collection of factors does a better job of pricing. That’s the new world toward which economics is moving. While the old intuitions are gasping for breath the new-fangled numbers are fit as a fiddle and ready for duty.

I thought some more and realized that what’s really going on is that people are evaluating those stocks, communicating with one another directly about them, and making decisions about buying and selling, thereby communicating indirectly with one another. That’s what those 360,000 factors are capturing, the actions of a dispersed community of analysts and traders. “Could this be roughly similar to the decisions movie-goers make about the movies they see based, not only on their preferences, but on information they get from reviews, and perhaps more importantly, from their friends?” “If so,” I conjectured, “then perhaps Cowen’s old colleague from Irvine, Arthur De Vany, can shed some light on the situation.” That is to say, can give me some intuitions that I can apply to the situation.

For De Vany had written a very interesting book, Hollywood Economics (2004), about the fate of movies once they have been released. Just as those intuitive “classical” models in economics aren’t as accurate as the new high-factor models, so you can’t predict the box-office performance of movies on such simple factors as the identities of the producer, screen writers, or stars in the movies. Now, De Vany didn’t produce a high-factor model that improved matters, he did something quite different (which is discussed below, pp. 23 ff.), but that’s secondary at the moment. The point is that we seem to have a gross similarity, the behavior of some object that interests a lot of people, a stock or a movie, cannot be reliably predicted using a simple model.

Meme stocks and novels

The similarity was reinforced when I heard a remark by Scott Galloway on the Pivot podcast: “Stocks are like brands and that is they’re part promise and part performance.” Consider the recent phenomenon of meme stocks, which Wikipedia glosses this way:

A meme stock is a stock that gains popularity among retail investors through social media. The popularity of meme stocks is generally based on internet memes shared among traders, on platforms such as Reddit's r/wallstreetbets. Investors in such stocks are often young and inexperienced investors. As a result of their popularity, meme stocks often trade at prices that are above their estimated value – as based on fundamental analysis – and are known for being extremely speculative and volatile.

Meme stocks are assets where promise overwhelms performance, more story than substance.

That’s what movies are. You are purchasing the story and the experience, not the seat in the theater, or the DVD, or the stream, those are the vehicles that carry the story. Claude calls these things “thick” objects (perhaps borrowing from the anthropological concept of “thick” description? ), as opposed to “thin” objects like toasters and drills. Novels are thick objects as well, which led me to Matthew Jockers’ 2013 book, Macroanalysis, where he uses machine learning to develop a high dimensional model (a mere 600 dimensions rather than 360,000) of a corpus of 3000 19th century Anglophone novels. Just as read De Vany’s book quite closely, so I’ve written a series of posts about Jockers’ book. I bring his model into the mix as well (pp. 27 ff.).

Thus I am now in a position to take two models in subjects I know well, movies and novels, and bring them to bear on contemporary machine learning in financial economics, a subject I do not know at all. And, for that matter, still don’t. But I’ve got some intuitions. And one of those intuitions led me to focus on the fact that, while Jockers’ model did not contain any dates, upon inspection it turned out to have a diagonal (p. 27) that is correlated with direction in time. Not only did 19th century novels change in theme and motif over time, there is a direction to that change. The system seems to exhibit directional evolution. And so I directed Claude to explore the possibility of that this might be a general characteristic of thick-objects being used by a large population of interested parties (pp. 31 ff). Here is the conjecture Claude arrived at (p. 35):

In thick-object domains, low-dimensional intuitive factors often fail to explain individual outcomes. But high-dimensional representation can reveal population-level structure: outcome basins in movies, pricing kernels in finance, and temporal direction in novels. The next step is to ask whether all such artifact systems exhibit historical vectors in feature space, generated by a generational ratchet in which each cohort of producers is shaped by the artifact ecology inherited from its predecessors.

Notice the territory we have traversed in conceptual space. We started with an undergraduate at Johns Hopkins (me) using interpretive methods to study a poem, “Kubla Khan.” That investigation led to problems that forced me to study computational semantics in graduate school, a distinctly different mode of intellectual work, one based on formulating an elaborate system of structural rules. We then zipped through time and over intellectual space to a social scientist, Tyler Cowen, who was trained in the used of causal models to generate statistically controlled observations about economic behavior. He is now confronted with multifactor machine learning models with no intuitively discernible causal structure that nonetheless have superior predictive power. Cowen got me interested in one of those models and I, in turn, summoned Anthropic’s Claude to explain it to me.

The way I see, and I’ve seen it this way for a long time, the human sciences – more a European notion than American, les sciences humaines – can be arranged into three camps according to methodological focus: interpretive or hermeneutic (roughly, the humanities), causal modeling (roughly, the social sciences), and structural rules (roughly, the “classical” cognitive sciences). We’ve spanned them all in the course of this introduction. What will the future bring?

Bonus: I leave it as an exercise for the reader to consider the relevance of Keynes’s talk of “animal spirits” and to incorporate Robert Shiller’s narrative economics into this picture.

What’s in this document

The rest of this document is devoted to the dialogs where I used ChatGPT to work through the connections between these three models, two I knew quite well (De Vany on movies and Jockers on novels), and one I did not (Didisheim et al. on asset pricing). Claude knows them all, for some non-trivial meaning of “know,” and many others as well. The purpose of the dialog, then, is to link something I do not know to something that I do. The dialog took place in four sessions over the course of a week from the end of May into June.

Rather than comment on each of the sections listed in the outline, with one exception, I am commenting only on the sections that mark the beginning of a new session with ChatGPT. For what it’s worth, they mark how the subject evolved in my mind. The one exception? The summary was the last thing ChatGPT did, obviously, but I moved it to first place.

Thick Objects and the New Common Ground of Economics and Cultural Analysis [Summary] – I had ChatGPT prepare this summary and the very end of the process, on June 22. I put if first in case some might want to get the gist of the exercise without slogging through the details.

AIPT, Large Factor Models [First Session] – There is where I began on May 26. I started by asking ChatGPT to explain asset pricing to me. Once I had some sense of that, I then went on to the models I was familiar with, first De Vany on movies and the Jockers on 19th century Anglophone novels.

Stocks are like brands [Session Two] – I initiated this session on May 30 when I heard Galloway’s remark about stocks being like brands. That crystalized things for me so I needed to work back through the analysis. In the course of that discussion I focused on the concept of a brand as a distinct conceptual objects and ChatGPT’s response clarified the role of marginalism in clearing the way for asset models with a very large number of factors.

Algorithmic and Causal models [Session Three] – I don’t recall whether anything in particular prompted me to initiate this dialog. Perhaps mere methodological curiosity. This took place on June 2.

Those empirical APT models [Session Four] – It’s not entirely clear to me just whether anything in particular prompted this session. But what I was thinking was that, while I’m familiar with novels and movies and the academic discourse about them, asset pricing is unfamiliar territory. So I wanted to nail down as well as I could just what “ground truth” is in this area. Movies start with eyeballs in theaters and novels start with eyeballs scanning pages, where does asset pricing start? Once ChatGPT had gone through this I realized that I’d seen it earlier in the whole process. Still, I was happy to go through it again, this time coming at it after having thought about it. It’s as the end of this session that I asked ChatGPT to summarize the discussion.

Sunday, December 15, 2019

Some informal remarks on Jockers’ 3300 node graph: Part 3, Signs and mechanisms

I want to start at the point where I ended my previous post in this series, Some informal remarks on Jockers’ 3300 node graph: Part 2, structure and computational process. I was arguing that the issue of scale was misconceived. The fundamental issue is NOT a corpus of texts versus one or a handful of texts. That’s trivial. The issue has to do with the terms of analysis. So-called close reading exists within a conceptual and methodological discourse which is quite different from that of so-called distant reading.

The difference is one of conceptual ontology, as that term has come to be understood in computer science and the cognitive sciences. In that context the issue is not the ultimately real, a philosophical question, but the kind of concepts we are we using. I start with a simple example, salt versus sodium chloride and then build out from there to words versus signifiers and then on to texts and meaning.

Conceptual ontology

Consider for a moment, salt on the one hand and NaCl (sodium Chloride) on the other. Physically they are (almost) the same substance, but conceptually they are quite different. Salt is defined and understood in terms of its physical appearance and, above all else, its taste. We can taste salt even when we cannot see it, and so can animals. NaCl, however, is defined in terms of an atomic theory of matter that didn’t exist until early in the 19th century. Moreover, NaCl is a pure substance, consisting of nothing by sodium and chlorine atoms; salt on the other hand will always have some impurities. Thus, strictly speaking, salt and NaCl are not physically the same, close, but not exactly the same.

Salt and sodium chloride, then, are used in different intellectual contexts, each of which has its own vocabulary. Salt belongs with sugar, pepper, cinnamon, flour, and so forth all substances having to do with food, food preparation, and eating all related concepts. Sodium chloride is related to potassium chloride, sodium hydroxide, and so forth, electrolysis, ion-exchange, and so forth, for a long list. Thus we have two different conceptual worlds, each coupled to characteristic actions and processes, but both ultimately grounded in the same physical reality. Someone whose occupation has them working in the sodium chloride world has no trouble with table salt at mealtime. The transition from one world to the other is seamless, or nearly so (there may be a change of clothes involved).

The worlds of “distant reading” and “close reading” differ from one another in the same way. The physical texts and the symbols imprinted on them are the same in both worlds, but the concepts and methods of description and analysis are different. It’s that (kind of) different that Geoffrey Hartman had in mind almost a half century ago when he observed: “modern ‘rithmatics’—semiotics, linguistics, and technical structuralism—are not the solution. They widen, if anything, the rift between reading and writing” (The Fate of Reading, Chicago 1973, p. 272). He expresses himself using the standard trope of distance, but he’s certainly not talking about distance in any physical sense – no one is.

These days we might want to gloss distance in terms of explicit mediating steps. In close-reading you read a primary text and then does some thinking, perhaps some secondary reading and research, and then you write about that primary text. Just how you break that down is somewhat arbitrary, but it’s nothing like what happens in distant-reading, which starts when a collection of primary texts is digitized. The digitized texts are then cleaned up, tagged with metadata, organized in a database or databases, and then subject to analysis, which might be a straightforward matter of word counts, or the somewhat more strenuous process of topic modeling, or perhaps we’re going to make use of vector semantics, or any of a number of things. This more complicated process is likely to involve two, three, or more people at various stages along with way. At some point the analytic process will produce some set of visualizations and they, in turn, will be interpreted in terms appropriate to the primary texts (and their contexts).

As I said, that’s how we might gloss the close vs. distant distinction these days. The distance isn’t one of physical steps, but operational steps. However, it’s doubtful that Hartman had anything like this in mind. To be sure, stylometrics existed in 1973, and Stanley Fish was doing his best to skewer it, but I doubt that Hartman was thinking about stylometrics when he made that remark. However, he might well have been thinking about the kinds of tables and diagrams that Lévi-Strauss used, or that showed in linguistics article, those diagrammatic signs that interrupted the linguistic flow. They are indices of a different mode of thought, a different conceptual ontology.

Words and signifiers

We can begin to appreciate that difference by noting the difference between words, which are signs in Saussure’s sense, and signifiers, which are components of signs. Literary critics generally talk of words, and occasionally of signifiers. The concept of word is transparent enough in ordinary casual discourse. As ordinarily understood, the concept of words encompasses pronunciation, spelling, grammatical usage, and meaning, often several meanings. Words in this sense are the things listed in dictionaries. And for the most part, literary critics deal in words, even those who’ve read a bit of Saussure.

Saussure distinguished between the signifier and the signified. The signifier is a physical entity, either sonic or visual. The signfied is a mental entity; it is the bearer of meaning. Signifiers are public; they can be transmitted between people. Meanings are, well, they’re not exactly private – Wittgenstein established that in a famous section in his Investigations – but as mental objects they exist in people’s heads and we do not have direct access to one another’s heads. Thus, as William Croft has argued in chapter 4 of Explaining Language Change [1], word meanings are negotiated in conversational interaction with one another.

Computational critics, in contrast, may talk of words, but what they are actually working with are signifiers, signifiers rendered in digital form. That is to say, that’s what’s in the databases at the heart of computational work, mere signifiers. The computer has no access to word meanings, to signifieds; it only knows the written signifier. Consider the following passage from the well-known 1949 memo on machine translation written by Warren Weaver, who headed the Natural Sciences Division of the Rockefeller Foundation:
First, let us think of a way in which the problem of multiple meaning can, in principle at least, be solved. If one examines the words in a book, one at a time as through an opaque mask with a hole in it one word wide, then it is obviously impossible to determine, one at a time, the meaning of the words. "Fast" may mean "rapid"; or it may mean "motionless"; and there is no way of telling which.

But if one lengthens the slit in the opaque mask, until one can see not only the central word in question, but also say N words on either side, then if N is large enough one can unambiguously decide the meaning of the central word. The formal truth of this statement becomes clear when one mentions that the middle word of a whole article or a whole book is unambiguous if one has read the whole article or book, providing of course that the article or book is sufficiently well written to communicate at all.

The practical question is, what minimum value of N will, at least in a tolerable fraction of cases, lead to the correct choice of meaning for the central word?
While that memo catalyzed early work on machine translation, the approach suggested in those paragraphs played no role in that work. The necessary computing power wasn’t available. That changed in the 1980s and especially 1990s.

The topic analysis technique that Jockers used follows from the insight expressed in those paragraphs, as does the work on vector semantics. The idea is simple. Words that appear together frequently must somehow have related meanings. Race, horse, saddle, jockey, and track, all mean different things, but they belong to the same discourse and so will appear in close proximity when that discourse is spoken, or written. And so it is with human, race, IQ, identity, American, black, and white, these too belong to a discourse and so will co-occur in texts using that discourse. Notice that race is common to those two very different discourses. Considered alone, without context, the meaning of race is ambiguous, we don’t know what it means. But if we see race in close proximity to horse, we’ll locate it in one discourse while if we see it in proximity to IQ we’ll locate it in that other discourse.

Signifiers in themselves have no meaning. But in actual usage signifiers are always bound to a meaning and that meaning connects them to related signifiers. Given a large enough body of texts, computers can determine which signifiers occur together. It is up to the investigator to make a judgment about why those signifiers are occurring together. The investigator will, of course, make that judgment on the basis of their knowledge of the language.

The conventional literary critic finds such judgments utterly trivial and has no way, no intellectual context, for understanding how remarkable it is that mere computation can discover such relationships. And since that critic is only interested in a handful of texts they have no reason to reconsider that judgment and consider the possibility that there is something to be learned in examining such patterns in collections of texts, such as the collection of 3346 Anglophone novels Jockers has been working with.

Text and meaning

But what is a text? For the linguist, or for the computational critic, and answer is simple: a text is a string of characters.

Things are not so simple for the literary critic. Yes, sure, a text is a physical object, a scroll or a codex, and it is a bunch of signifiers inscribed on such. That’s trivial and, for the most part, uninteresting except in very well defined contexts, such as physical preservation or, more interestingly, the preparation of an edition. And the preparation of critical editions has been very important in the development of digital humanities broadly conceived, but still, we can bracket that.

What I am interested in here is the text as an object of interpretation. THAT text is an enigma [2]. The following passage is from the introduction Rita Copeland and Frances Ferguson prepared for five essays from the 2012 English Institute devoted to the text:
Yet with the conceptual breadth that has come to characterize notions of text and textuality, literary criticism has found itself at a confluence of disciplines, including linguistics, anthropology, history, politics, and law. Thus, for example, notions of cultural text and social text have placed literary study in productive dialogue with fields in the social sciences. Moreover, text has come to stand for different and often contradictory things: linguistic data for philology; the unfolding “real time” of interaction for sociolinguistics; the problems of copy-text and markup in editorial theory; the objectified written work (“verbal icon”) for New Criticism; in some versions of poststructuralism the horizons of language that overcome the closure of the work; in theater studies the other of performance, ambiguously artifact and event. “Text” has been the subject of venerable traditions of scholarship centered on the establishment and critique of scriptural authority as well as the classical heritage. In the modern world it figures anew in the regulation of intellectual property. Has text become, or was it always, an ideal, immaterial object, a conceptual site for the investigation of knowledge, ownership and propriety, or authority? If so, what then is, or ever was, a “material” text? What institutions, linguistic procedures, commentary forms, and interpretive protocols stabilize text as an object of study?
What? “Linguistic data” sounds like it might be mere signifiers, and perhaps “copy-text” as well. But the rest of them, those various sites “for the investigation of knowledge, ownership and propriety, or authority”, those texts clearly consist of words, in the ordinary sense, complete with their multiple and often ambiguous and contradictory meanings.

That is the text that is conceptualized in vague spatial metaphors. Do we talk of meaning as being IN the text or as being somehow outside the text, in the CONTEXT? I think it would be a mistake to interpret that as a vague assertion about mechanism, though I am sorely tempted to do so. Rather, it’s a statement about critical methodology. If you think of meaning as IN the text, then you do not invoke anything but the text itself in the process of interpreting it. That’s what was originally meant by close reading. But if you mean that meaning is located in the context, well then you must bring other evidence to bear in your close reading. Depending on your particular methodology you may call on historical materials of various kinds – newspapers and other contemporary periodicals, correspondence, legal documents, and so forth – or you may want to mount a psychological argument in terms of this or that theory, diaries, and what have you. And of course the text itself.

And then we have the idea of “hidden” meaning. That’s clearly a spatial metaphor. But a metaphor for what? Not the text itself; it’s not as though anyone is imagining a secret text bound into the spine or sandwiched into the front or back boards. The idea seems to be something like a large and elaborate old house with hidden rooms and passages. Expressed in that way it seems just a bit foolish, but then no one ever expresses it in that way, do they? Again, we’re dealing with a covert methodological injunction. Words have many meanings and are linked to other meanings through various figures, such as the master figures of metaphor and metonymy – need I invoke, for example, Roman Jakobson’s work on the subject [4]? And so, as a methodological precept, one must explore those possibilities.

The fact of the matter is that, by contemporary standards, those of the so-called cognitive revolution and after, literary criticism lacks an account of language mechanism, whether in the reader, the author, or the critic. It is non-mechanistic or pre-mechanistic and certainly non-computational – an issue I’ll take up in the next and last post in this series. For now I want to end on the assertion that meaning, as the object of critical investigation, is inherently and irreducibly subjective. That doesn’t mean that critics cannot agree on the meaning of texts, for subjectivity necessarily implies intersubjectivity. It means only that meaning resides in subjects, in readers, writers, and critics and is not subject to objective determination. It is thus like color, which is also subjective in that sense. Sensations of color may be closely related to wavelengths of light, but cannot simply be reduced to them. Color arises within subjects, though psychologists have made a great deal of progress in figuring out how color perception functions in humans and animals.

The objective study of meaning, that is, of semantic mechanisms, has not yet been so successful. It was the search for semantic mechanisms that lead me to David Hays and computational linguistics, which I discussed in the previous post in this series, Some informal remarks on Jockers’ 3300 node graph: Part 2, structure and computational process.

References

[1] Croft, William (2000). Explaining Language Change: An Evolutionary Approach. Longman.

[2] I’ve written a number of posts at New Savanna about the concept of the text. They’re at this link: http://new-savanna.blogspot.com/search/label/text.

[3] Rita Copeland and Frances Ferguson, “Introduction”, ELH, Volume 81, Number 2, Summer 2014, p. 417.

[4] For example, Roman Jakobson and Morris Halle, Two Aspects of Language and Two Types of Aphasic Disturbances in Fundamentals of Language, The Hague & Paris: Mouton, 1956.

[5] I’ve got a good many posts on color, https://new-savanna.blogspot.com/search/label/color. They are of various kinds. Some are simply sets photographs where I was interested in colors and color contrasts. But others are about color perception in the mind and in photography. If you are interested in this topic you might what to search the web on “color perception”.

Tuesday, September 10, 2019

Reading Macroanalysis 7.3: Style, Genre, Time, and Influence

This post is from Aug. 31, 2014, but I'm bumping it to the top of the queue as I am thinking about these matters in connection with Moretti and Sobchuk, Hidden in Plain Sight: Data Visualization in the Humanities (New Left Review 118, 2019, 86-119). They don't discuss this visualization, but they should have.
In this post I suggest some studies I’d like to be done. I begin by recalling Moretti’s account of genre succession from Maps, Graphs, Trees in the context of Jockers’ massive graph of literary influence. Then I revisit the “Style” chapter and look at some of the work I passed over when I first posted on that chapter, the work related to Moretti’s generational observation. I then make some suggestions about how we could infer quasi-genres in the data assembled to build the influence graph and thereby extend Jockers’ work on style from his limited corpus of 106 texts to the larger corpus of 3346 texts. I conclude with some vague and tentative remarks about the pattern of reader interest betrayed in the record we’ve been examining, that of book publication.

Influence and Genre Succession

I’ve been thinking a lot about two things: 1) Moretti’s argument in Graphs, Maps, Trees that genres tend to cluster into 30 year cycles, and 2) Jockers’ massive graph in which all 3346 texts in his corpus are linked by relations of similarity, producing a graph that looks like this (which is Figure 9.3, p. 165; color version from the web):

9dot3

As Jockers points out, what’s remarkable about this graph is that the nodes are ordered in time from left (oldest) to right, but there is no temporal information in the data from which it was derived: “Books are being pulled together (and pushed apart) based on the similarity of their computed stylistic and thematic distances from each other” (p. 164).

That temporal ordering is a side effect of ordering by thematic and stylistic similarity. But, in the abstract, it could have been otherwise, no? Why should positioning texts near similar texts result in temporal ordering? (Would the same thing be true of 20th Century texts?) This ordering implies that the evolution of literary culture IS directional, but Jockers himself hasn’t posited any telos, nor do I see any need to do so. That directionality stems from the internal dynamics of the system. Authors, and I assume audiences as well, want to stick with what they know, and what they know was published in the previous years.

It seemed to me that Moretti’s cycles must somehow be in that graph, for all the texts in a given cycle are close together in time, by definition, as well as similarity. Alas, the whole corpus has not been coded for genre (p. 158). Is there some way we can back into genre since we’ve got this massive graph based on similarity relations among texts along 578 dimensions? Aren’t texts within the same genre more likely to resemble one another than texts in different genres?

The other thing on my mind is the fact that what really interests me is what’s on people’s minds and how that evolves over time. Some books will attract few readers, some books many readers; but the mere fact that a book has been published doesn’t speak to that. Moreover, books can be read long after they’ve been published. In the case of Moby Dick, it would seem that, for the most part it was read only long after it was published. Publication history is, at best, an indirect proxy measure of that.

And yet that history IS a history. Assuming that publishers are for the most part rational economic actors who want to turn a profit, their decisions on what to publish must take into account their sense of what people are reading and therefore what they’re buying. And the kinds of books that got published changed from one decade to the next. That record of  changes must reflect changes of reading taste.

Thursday, May 9, 2019

Notes toward a theory of the corpus, Part 1: History [#DH]

The recent discussion of Nan. Z. Da,  The Computational Case against Computational Literary Studies, has me thinking about Matt Jockers' Macroanalysis, in particular, about his high-dimensional graph of his 19th century corpus. Da dissmisses with with two paragraphs (pp. 610-611). Interestingly enough, though, when Critical Inquiry hosted a discussion forum on the article, it used tan image of that graph to head the forum. That, I assume, is because it is visually compelling.

Visually compelling, but nontheless trivial? I think not. I'm bumping this post to the top of the cue. It represents my thinking about the implications of that diagram as of late September of 2018. I've thought a bit more about that diagram in the past month, going over and over and over. I've got a few more thoughts.
That graph, of course, is constructed in a space of roughly 600 dimensions. We can think of that space as, shall we say, a design space, where each point represents a possible novel. Like most such spaces, most of it is empty. What's interesting about his space is that it emerged over time and we can more or less track that emergence. That follows from the fact that the texts in that space are ordered in time, more or less, from left to right. As novels were written in the course of the century, they enlarged the space, rather than moving about in the already existing space. That's what's interesting, new texts enlarged the space. What does that tell us about the history of the novel, and about cultural evolution?
A corpus

By corpus I mean a collection of texts. The texts can be of any kind, but I am interested in literature, so I’m interested in literary texts. What can we infer from a corpus of literary texts? In particular, what can we infer about history?

Well, to some extent, it depends on the corpus, no? I’m interested in an answer which is fairly general in some ways, in other ways not. The best thing to do is to pick an example and go from there.

The example I have in mind is the 3300 or so 19th century Anglophone novels that Matthew Jockers examined in Macroanalysis (2013 – so long ago, but it almost seems like yesterday). Of course, Jockers has already made plenty of inferences from that corpus. Let’s just accept them all more or less at face value. I’m after something different.

I’m thinking about the nature of historical process. Jockers' final study, the one about influence, tells us something about that process, more than Jockers seems to realize. I think it tells us that cultural evolution is a force in human history, but I don’t intend to make that argument here. Rather, my purpose is to argue that Jockers has created evidence that can be brought to bear on that kind of assertion. The purpose of this post is to indicate why I believe that.

A direction in a 600 dimension space

In his final study Jockers produced the following figure (I’ve superimposed the arrow):

direction of 19C lit history


Each node in that graph represents a single novel. The image is a 2D projection of a roughly 600 dimensional space, one dimension for each of the 600 features Jockers has identified for each novel. The length of each edge is proportional to the distance between the two nodes. Jockers has eliminated all edges above a certain relatively small value (as I recall he doesn’t tell us the cut off point). Thus two nodes are connected only if they are relatively close to one another, where Jockers takes closeness to indicate that the author of the more recent novel was influenced by the author of more distant one.

You may or may not find that to be a reasonable assumption, but let’s set it aside. What interests me is the fact that the novels in this graph are in rough temporal order, from 1800 at the left (gray) to 1900 at the right (purple). Where did that order come from? There were no dates in 600D description of each novel, so the software was not reading dates and ordering the nodes according to those dates. By process of elimination, that ordering must be a product of whatever historical process that produced the texts represented in the graph. What else is there? That process must therefore have a temporal direction.

I’ve spent a fair amount of effort explicitly arguing that point [1], but don’t want to reprise that argument here. I note, however, that the argument is a geometrical one. For the purposes of this piece, assume that that argument is at least a reasonable one to make.

What is that direction? I don’t have a name for it, but that’s what the arrow in the image indicates. One might call it Progress, especially with Hegel looking over your shoulder. And I admit to a bias in favor of progress, though I have no use for the notion of some ultimate telos toward which history tends. But saying that direction is progress is a gesture without substantial intellectual content because it doesn’t engage with the terms in which that 600D space is constructed. What are those terms? Some of them are topics of the sort identified in topic analysis, e.g. American slavery, beauty and affection, dreams and thoughts, Greek and Egyptian gods, knaves rogues and asses, life history, machines and industry, misery and despair, scenes of natural beauty, and so on [3]. Others are stylistic features, such as the frequency of specific words, e.g. the, heart, would, me, lady, which are the first five words in a list Jockers has in the “Style” chapter of Macroanalysis (p. 94).

The arrow I’ve imposed on Jockers’ graph is a diagonal in the 600D space whose dimensions are defined by those features and so its direction must specified in terms that are commensurate with such features. Would I like to have an intelligible interpretation of that direction? Sure. But let’s leave that aside. We’ve got an abstract space in which we can represent the characteristics of novels (Daniel Dennett might call this a design space) and we’ve got a vector in that space, a direction.

What’s that direction about? What is it about texts that is changing as we move along that vector? I don’t know. Can I speculate? Sure. But not here and now. What’s important now is that that vector exists. We can think about it without having to know exactly what it is.

A snapshot of Spirit of the 19th century

In a post back in 2014 I suggested that Jockers’ image depicts the Geist of 19th century Anglo-American literary culture [2]. That’s what interests me, the possibility that we’re looking at a 21st century operationalization of an idea from 19th century German idealism. Here’s what the Stanford Encyclopedia of Philosophy has to say about Hegel’s conception of history [4]:
In a sense Hegel’s phenomenology is a study of phenomena (although this is not a realm he would contrast with that of noumena) and Hegel’s Phenomenology of Spirit is likewise to be regarded as a type of propaedeutic to philosophy rather than an exercise in or work of philosophy. It is meant to function as an induction or education of the reader to the standpoint of purely conceptual thought from which philosophy can be done. As such, its structure has been compared to that of a Bildungsroman (educational novel), having an abstractly conceived protagonist—the bearer of an evolving series of so-called shapes of consciousness or the inhabitant of a series of successive phenomenal worlds—whose progress and set-backs the reader follows and learns from. Or at least this is how the work sets out: in the later sections the earlier series of shapes of consciousness becomes replaced with what seem more like configurations of human social life, and the work comes to look more like an account of interlinked forms of social existence and thought within which participants in such forms of social life conceive of themselves and the world. Hegel constructs a series of such shapes that maps onto the history of western European civilization from the Greeks to his own time.
Now, I am not proposing that Jockers’ has operationalized that conception, those “so-called shapes of consciousness”, in any way that could be used to buttress or refute Hegel’s philosophy of history – which, after all, posited a final end to history. But I am suggesting that can we reasonably interpret that image as depicting a (single) historical phenomenon, perhaps even something like an animating ‘force’, albeit one requiring a thoroughly material account. Whatever it is, it is as abstract as the Hegelian Geist.

How could that be?

Thursday, June 25, 2015

On the Direction of 19th Century Poetic Style, Underwood and Sellers 2015

Another working paper (title above). Download at:

Abstract, contents, and introduction below:

Abstract: Underwood and Sellers have discovered that over the course of roughly a century (1820-1919) Anglo-American poetry has undergone a consistent change in style in a direction favored by editors and reviewers of elite journals. This directional shift aligns with the one Matthew Jockers found in Angophone novels during roughly the same period (from the beginning of the 19th century to its end). I argue that this change is characteristic of a cultural evolutionary process and sketch a way to simulate such a process as an interaction between a population of texts and a population of writers where texts and writers. I suggest that such directionality is a sign of autonomy in the aesthetic system, that it is not completely coupled to and subsumed by surrounding historical events.

C O N T E N T S

0. Introduction: Looking at Cultural Evolution whether You Like It or Not 2
1. Cosmic Background Radiation, an Aesthetic Realm, and the Direction of 19thC Poetic Diction 8
2. Beyond Whig History to Evolutionary Thinking 14
3. Could Heart of Darkness have been published in 1813? – a digression 19
4. Beyond narrative we have simulation 22

0. Introduction: Looking at Cultural Evolution whether You Like It or Not

I was of course thrilled to read How Quickly Do Literary Standards Change? (Underwood and Sellers 2015). Why? Because they provide preliminary evidence that 19th century Anglophone poetic culture has a direction. Just what that direction, and how to characterize it, that’s something else. But there does appear to be a direction. And just why is that exciting? Because Matthew Jockers made the same discovery about the 19th century Anglophone novel. To be sure, that’s not what he claimed – I’ve had to reinterpret his work (see my working paper, On the Direction of Cultural Evolution: Lessons from the 19th Century Anglophone Novel) – but that’s what he has in fact done.

So we’ve got two investigations making the same observation: there is a long-term direction 19th century literary culture. But not the same, as Jockers looked at novels and Underwood and Sellers looked at poetry. Moreover their observational methods are quite different. Jockers uncovered direction by looking for similarity between texts where similarity judgments are based on a variety of stylistic measures and on topic analysis. Underwood and Smalls bumped into directionality by looking for differences between the general run of literary texts and texts selected for review by elite publications. Jockers’ work, almost by design, uncovered continuity between successive cohorts of texts, but simply ignored elite culture. Underwood and Smalls had no explicit interest in local continuity but, by looking at elite choice, uncovered a possible factor in directional cultural change: the “pressure” of elite preference on the system as a whole.

Wednesday, June 3, 2015

Where I’m at on cultural evolution, some quick remarks

I don’t know.

Some notes to myself.

1. Cultural Analogs to Genes and Phenotypes

I’ve spent a fair amount of time off and on over the last two decades hacking away at identifying cultural analogues to biological genes and phenotypes. In the past few years that effort has taken the form of an examination of Dan Dennett. I more or less like the current conceptual configuration, where I’ve got Cultural Beings as an analog to phenotypes and coordinators as analogs to genes. As far as I can tell – and I AM biased, of course, it’s the best such scheme going.

And it just lays there. So what? I don’t see that it allows me to explain anything that can’t otherwise be explained. Nor does it have obvious empirical consequences that one could test in obvious ways. It seems to me mostly a formal exercise at this point. In that it is not different from any version of memetics nor from Sperber’s cultural attractor theory. These are all formal exercises with little explanatory value that I can see.

That’s got to change. But how? I note that dealing with words as evolutionary objects seems somewhat different from treating literary works (or musical works and performances, works of visual art, etc.) as evolutionary objects.

Issues: Design, Human Communication

2. Cultural Direction

Perhaps the most interesting work I’ve done in the past year as been my work on Matt Jockers’ Macroanalysis and, just recently, on Underwood and Sellers’ paper on 19th century poetry. In the case of Jockers’ work on the novel, he’d done a study of influence which I’ve reconceptualized as a demonstration that the literary system as a direction. In the case of Underwood and Sellers, they’ve found themselves looking at directionality, but they hadn’t been looking for it. Their problem was to ward of the conceptual ‘threat’ of Whig historicism; they want to see if they can accept the directionality but not commit themselves to Whiggishness, and I’ve spent some time arguing that they need not worry.

What excites me is that two independent studies have come up with what looks like demonstrations of historical direction. I take this as an indication of the causal structure of the underlying historical process, which encompasses thousands upon thousands of people interaction with and through thousands of texts over the course of a century. What shows up in the texts can be thought of as a manifestation of Geist and so these studies are about the apparent direction of Geist.

Wednesday, May 27, 2015

Could Heart of Darkness have been published in 1813? – a digression from Underwood and Sellers 2015

Here I’m just thinking out loud. I want to play around a bit.
Conrad’s Heart of Darkness is well within the 1820-1919 time span covered by Underwood and Sellers in How Quickly Do Literary Standards Change?, while Austen’s Pride and Prejudice, published in 1813, is a bit before. And both are novels, while Underwood and Sellers wrote about poetry. But these are incidental matters. My purpose is to think about literary history and the direction of cultural change, which is front and center in their inquiry. But I want to think about that topic in a hypothetical mode that is quite different from their mode of inquiry.

So, how likely is it that a book like Heart of Darkness would have been published in the second decade of the 19th century, when Pride and Prejudice was published? A lot, obviously, hangs on that word “like”. For the purposes of this post likeness means similar in the sense that Matt Jockers defined in Chapter 9 of Macroanalysis. For all I know, such a book may well have been published; if so, I’d like to see it. But I’m going to proceed on the assumption that such a book doesn’t exist.

The question I’m asking is about whether or not the literary system operates in such a way that such a book is very unlikely to have been written. If that is so, then what happened that the literary system was able to produce such a book almost a century later?

What characteristics of Heart of Darkness would have made it unlikely/impossible to publish such a book in 1813? For one thing, it involved a steamship, and steamships didn’t exist at that time. This strikes me as a superficial matter given the existence of ships of all kinds and their extensive use for transport on rivers, canals, lakes, and oceans.

Another superficial impediment is the fact that Heart is set in the Belgian Congo, but the Congo hadn’t been colonized until the last quarter of the century. European colonialism was quite extensive by that time, and much of it was quite brutal. So far as I know, the British novel in the early 19th century did not concern itself with the brutality of colonialism. Why not? Correlatively, the British novel of the time was very much interested in courtship and marriage, topics not central to Heart, but not entirely absent either.

The world is a rich and complicated affair, bursting with stories of all kinds. But some kinds of stories are more salient in a given tradition than others. What determines the salience of a given story and what drives changes in salience over time? What had happened that colonial brutality had become highly salient at the turn of the 20th century?

Friday, May 22, 2015

Underwood and Sellers 2015: Cosmic Background Radiation, an Aesthetic Realm, and the Direction of 19thC Poetic Diction

I’ve read and been thinking about Underwood and Sellers 2015, How Quickly Do Literary Standards Change?, both the blog post and the working paper. I’ve got a good many thoughts about their work and its relation to the superficially quite different work that Matt Jockers did on influence in chapter nine of Macroanalysis. I am, however, somewhat reluctant to embark on what might become another series of long-form posts, which I’m likely to need in order to sort out the intuitions and half-thoughts that are buzzing about in my mind.

What to do?

I figure that at the least I can just get it out there, quick and crude, without a lot of explanation. Think of it as a mark in the sand. More detailed explanations and explorations can come later.

19th Century Literary Culture has a Direction

My central thought is this: Both Jockers on influence and Underwood and Sellers on literary standards are looking at the same thing: long-term change in 19th Century literary culture has a direction – where that culture is understood to include readers, writers, reviewers, publishers and the interactions among them. Underwood and Sellers weren’t looking for such a direction, but have (perhaps somewhat reluctantly) come to realize that that’s what they’ve stumbled upon. Jockers seems a bit puzzled by the model of influence he built (pp. 167-168); but in any event, he doesn’t recognize it as a model of directional change. That interpretation of his model is my own.

When I say “direction” what do I mean?

That’s a very tricky question. In their full paper Underwood and Sellers devote two long paragraphs (pp. 20-21) to warding off the spectre of Whig history – the horror! the horror! In the Whiggish view, history has a direction, and that direction is a progression from primitive barbarism to the wonders of (current Western) civilization. When they talk of direction, THAT’s not what Underwood and Sellers mean.

But just what DO they mean? Here’s a figure from their work:

19C Direction

Notice that we’re depicting time along the X-axis (horizontal), from roughly 1820 at the left to 1920 on the right. Each dot in the graph, regardless of color (red, gray) or shape (triangle, circle), represents a volume of poetry and its position on the X-axis is volume’s publication date.

But what about the Y-axis (vertical)? That’s tricky, so let us set that aside for a moment. The thing to pay attention to is the overall relation of these volumes of poetry to that axis. Notice that as we move from left to right, the volumes seem to drift upward along the Y-axis, a drift that’s easily seen in the trend line. That upward drift is the direction that Underwood and Sellers are talking about. That upward drift was not at all what they were expecting.

Drifting in Space

But what does the upward drift represent? What’s it about? It represents movement in some space, and that space represents poetic diction or language. What we see along the Y-axis is a one-dimensional reduction or projection of a space that in fact has 3200 dimensions. Now, that’s not how Underwood and Sellers characterize the Y-axis. That’s my reinterpretation of that axis. I may or may not get around to writing a post in which I explain why that’s a reasonable interpretation.

Monday, April 27, 2015

On the Direction of Cultural Evolution: Lessons from the 19th Century Anglophone Novel

I've got another working paper available (title above):

Most of the material in this document was in an earlier working paper, Cultural Evolution: Literary History, Popular Music, Cultural Beings, Temporality, and the Mesh, which also has a great deal of material that isn’t in this paper. I’ve created this version so that I can focus on the issue of directionality and so I’ve dropped all the material that didn’t related to that issue. The last section, The Universe and Time, is new, as is this introduction.

* * * * *

Abstract: Matthew Jockers has analyzed a corpus of 19th century American and British novels (Macroanalysis 2013). Using standard techniques from natural language processing (NLP) Jockers created a 600-dimensional design space for a corpus of 3300 novels. There is no temporal information in that space, but when the novels are grouped according to close similarity that grouping generates a diagonal through the space that, upon inspection, is aligned with the direction of time. That implies that the process that created those novels is a directional one. Certain (kinds of) novels are necessarily earlier than others because that is how the causal mechanism (whatever they are) work. This result has implications for our understanding of cultural evolution in general and of the relationship between cultural evolution and biological evolution.

1. Introduction: Direction in Design Space, Telos? 2
2. The Direction of Cultural Evolution: The Child is Father or the Man 6
3. Nineteenth Century English-Language Novels 9
4. Macroanalysis: Styles 10
5. Macroanalysis: Themes 13
6. Influence and Large Scale Direction 15
7. The 19th Century Anglophone Novel 18
8. Why Did Jockers Get That Result? 20
9. What Remains to be Done? 21
10. Literary History, Temporal Orders, and Many Worlds 22
11. The Universe and Time 30

Introduction: Evolving Along a Direction in Design Space

In 2013 Matthew Jockers published Macroanalysis: Digital Methods & Literary History (2013). I devoted considerable blogging effort to it 2014, including most, but not all, of the material in this working paper. In Jockers’ final study he operationalized the idea of influence by calculating the similarity between each pair of texts in his corpus of roughly 3300 19th century English-language novels. The rationale is obvious enough: If novelist K was influenced by novelist F, then you would expect her novels to resemble those of F more than those of C, who K had never even read.

Jockers examined this data by creating a directed graph in which each text was represented by a node and each text (node) was connected only to those texts to which it had a high degree of resemblance. This is the resulting graph:

9dot3

It is, alas, almost impossible to read this graph as represented here. But Jockers, of course, had interactive access to it and to all the data and calculations behind it. What is particularly interesting, though, is that the graph lays out the novels more or less in chronological order, from left to right (notice the coloring of the graph), though there was no temporal information in the underlying data. Much of the material in the rest of this working paper deals with that most interesting result (in particular, sections 2, 6, 7, 8, and 10).

What I want to do here is, first of all, reframe my treatment of Jockers’ analysis in terms of something we might call a design space (a phrase I take from Dan Dennett, though I believe it is a common one in certain intellectual circles). Then I emphasize the broader metaphysical implications of Jockers’ analysis.

Monday, September 22, 2014

The Direction of Cultural Evolution, Macroanalysis at 3 Quarks Daily

As soon as I finished up my series of posts about Matt Jockers, Macroanalysis: Digital Methods & Literary History, I set up a file on my Mac for further thoughts, knowing full well I’d keep thinking about the book. I’ve now posted the first of those continuing thoughts at 3 Quarks Daily: Macroanalysis and the Directional Evolution of Nineteenth Century English-Language Novels.

The issue is cultural evolution, a notion that Jockers flirts with, but rejects. Of course I’ve been committed to the idea for a long time and I’ve decided that his data, that is, the patterns he’s found in his data, constitute a very strong argument of conceptualizing literary history as an evolutionary phenomenon. That’s what my 3QD post is about, a fairly detailed (a handful of new visualizations) reanalysis of Jockers’ account of literary influence.

From Influence to Evolution

It is one thing to track influence among a handful of texts; that is the ordinary business of traditional literary history. You read the texts, look for similar passages and motifs, read correspondence and diaries by the authors, and so forth, and arrive at judgements about how the author of some later text was influenced by authors of earlier texts. It’s not practical to do that for over 3000 texts, most of which you’ve never read, nor has anyone read many or even most them in over 100 years.

Here, in brief, is what Jockers did: He assumed that, if Author X was influenced by Author Q, then X’s texts would be very similar to Q’s. Given the work he’d already done on stylistic and thematic features, it was easy for Jockers to combine those features into a single list comprising almost 600 features. With each text scored on all of those features it was then relatively easy for Jockers to calculate the similarity between texts and represent it in a directed graph where texts are represented by nodes and similarity by the edges between nodes. The length of the edge between two texts is proportional to their similarity.

Note, however, that when Jockers created the graph, he did not include all possible edges. With 3346 nodes in the graph, the full graph where each node is connected to all of the others would have contained millions of edges and been all but impossible to deal with. Jockers reasoned that only where a pair of books was highly similar could one reasonably conjecture and influence from the older to the newer. So he culled all edges below a certain threshold, leaving the final graph with only 165,770 edges (p. 163).

When Jockers visualized the graph (using Force Atlas 2 in the Gephi) he found, much to his delight, that the graph was laid out roughly in temporal order from left to right. And yet, as he points out, there is no date information in the data itself, only information about some 600 stylistic and thematic features of the novels. What I argue in my 3QD post is that that in itself is evidence that 19th century literary culture constitutes an evolutionary system. That’s what you would expect if literary change were an evolutionary process.

Cultural Evolution Has a Direction

What’s particularly striking, though, is that this change is clearly directional, a matter I examine closely in my post. Another way to characterize Jockers’ graph is this:
The literary system is evolving in a 600 dimensional feature matrix. As time unfolds, the links between highly similar books trace a diagonal through the matrix.
But why?

Wednesday, September 3, 2014

Reading Macroanalysis: Notes on the Evolution of Nineteenth Century Anglo-American Literary Culture

Matthew L. Jockers. Macroanalysis: Digital Methods & Literary History. University of Illinois Press, 2013. x + 192 pp. ISBN 978-0252-07907-8

I've compiled all the posts into a working paper. HERE's the SSRN link. Abstract and introduction below.

* * * * *

Abstract: Macroanalysis is a statistical study of a corpus of 3346 19th Century American, British, Irish, and Scottish novels. Jockers investigates metatdata; the stylometrics of authorship, gender, genre, and national origin; themes, using a 500 item topic model; and influence, developing a graph model of the entire corpus in a 578 dimensional feature space. I recast his model in terms of cultural evolution where the dynamics are those of blind variation and selective retention. Texts become phenotypical objects, words become genetic objects, and genres become species-like objects. The genetic elements combine and recombine in authors' minds but they are substantially blind to audience preferences. Audiences determine whether or not a text remains alive in society.

* * * * *

Introduction: Get in the Driver’s Seat

I knew it was going to be good. But not THIS good. A better formulation: I didn’t know it would good in THIS way, that it would put me in driver’s seat, if only in a limited way.

The driver’s seat, you ask, what do you mean? In this case it means that I could actively work with the data. When, for example, I read Moretti’s Graphs, Maps, Trees, I read it as I do pretty much any book, though this one had a bunch of charts and diagrams, which is unusual for literary criticism. There wasn’t anything for me to do other than just read.

If I didn’t have ready access to the web, reading Macroanalysis would have been the same. But I do have web access and I use it all the time. So, when I got to Chapter 8, “Theme,” I also accessed the topic browser that Jockers had put on the web. Through this browser I could explore the topic model Jockers used in the book and, in particular, I could use it to investigate matters that Jockers hadn’t considered.

So I moved from thinking about Jockers’ work to using his work for my own intellectual ends. I ended up writing four posts (6.1 – 6.4) on that material totaling almost 12,000 words and I don’t know how many charts and graphs, all of which I got from Jockers’ web site. Once I’d worked through an initial curiosity about a spike that looked like Call of the Wild (but wasn’t, because that text isn’t in the database) I settled into some explorations framed by Leslie Fiedler’s Love and Death in the American Novel, Melville’s Moby Dick, and Edward Said’s anxiety on behalf of the autonomous existence of the aesthetic realm.

Data is Independent of Interpretations

You can do that as well, or whatever you wish. While the web browser gives you only limited access to Jockers’ corpus, that access is real and useful. A lot of work in digital criticism, and digital humanities in general, is like that. It produces ‘knowledge utilities’ that are generally useful, not just the private preserves of the original investigator.

There is an important epistemological point here as well. Jockers was led to this work by a certain set of intellectual concerns. Some of those concerns are quite general–about literature and the novel–while others are more specific–he has a particular interest in Irish and Irish-American literature. But I had no trouble putting his results to use in service of my own somewhat different interests.

Reading Macroanalysis 7: Influence, or the evolving dynamic integrity of the aesthetic sphere [REVISED]

Note: I decided that we needed a more explicit account of how Jockers visualized his 3346-node influence graph. I've inserted that account into the middle of the text and added subheadings.
I opened this investigation of Macroanalysis with the following paragraph:
The book arrived midway last week, when I hadn’t even finished reading Tim Morton’s Hyperobjects, much less finished blogging about it. But that didn’t stop me from giving Macroanalysis a look-thru: contents, some of the figures, read a bit here and there. I ended up reading Chapter 9, “Influence”, first; I’d read Matt Wilkins’ review in the LA Review of Books:
It’s a nifty approach that produces a fascinatingly opaque result: Tristram Shandy, Laurence Sterne’s famously odd 18th-century bildungsroman, is judged to be the most influential member of the collection, followed by George Gissing’s unremarkable The Whirlpool (1897) and Benjamin Disraeli’s decidedly minor romance Venetia (1837). If you can make sense of this result, you’re ahead of Jockers himself, who more or less throws up his hands and ends both the chapter and the analytical portion of the book a paragraph later.
Would I be able to make sense of those results? thought I to myself as I read. Nope, I couldn’t. Better luck next time.
I am now prepared to offer a re-interpretation of those results. But before I do that I need to explain more or less what Jockers is doing in this the final analytical chapter of the book. How does he operationalize the concept of influence?

What is influence?

When we say that, for example, that J. K. Rowling was influenced by the Narnia novels of C. S. Lewis, what do we mean? We mean that she read them and has incorporated features of those books into her own work. There is a direct relationship between Rowling’s activities and those influential books.

Influence thus understood is something that ‘travels’ along certain paths in the enormous meshwork of reading and writing transactions that constitute literary culture. As there are only a relatively few writers in that network, and only a relatively few of their transactions are writing ones (let’s say that the writing of a book is a single transaction) most of the transactions in the network are readings. Only a few of the transactions in the meshwork carry influence.

But Jockers doesn’t have access to that meshwork. None of us does. To be sure, we can see bits and pieces of it here in there in diaries, letters, and published reviews, but most of the transactions are lost to history. We can only look for the effects of those transactions.

And that’s what Jockers does. He assumes, reasonably enough, that if one author is influenced by another, then we should see indicators of that influence in the work. There should be a noticeable resemblance between those works.

And that is something Jockers can look for. For each of his 3,346 texts he’s got a bunch of features, stylistic and thematic. Once he’s tossed out the uninterpretable thematic features he’s left with 578 features for each text. He then represents this information as a geometric space with 578 dimensions, one for each feature, in which we have 3346 points, one for each text. He can now calculate the distance between any two texts, that is, points, in this high dimensional feature space. That distance is a measure of the similarity between the texts.

That’s what he does, and he gives us examples of the results. For each of Pride and Prejudice, Tale of Two Cities, and Moby Dick he gives us a table listing the ten novels the shortest distance from them and thus most like them (in terms of these 578 features). Not surprisingly, other books by Austen, Dickens, and Melville respectively occupy the top slots on these similarity lists. For the Austen list, the other authors are female, but one (Thomas Lister). Similarly, the authors most like Dickens are male, though the author of Life’s Masquerade (10th) is unknown, hence gender unknown. All the authors on these two lists are British. In Melville’s case, the list is also all-male, but not all-American. Two Scots, Robert Ballantyne and Robert Louis Stevenson, also made the list.

As interesting as this is, Jockers points out that it’s a bit small scale. We need something else if we want to gauge influence throughout the century. What to do?

Monday, September 1, 2014

[8] From Macroanalysis to Cultural Evolution

The purpose of this post is to recast the work reported in Macroanalysis: Digital Methods & Literary History in terms appropriate to cultural evolution. The idea is to propose a model of cultural evolution and assign objects from Jockerss analysis to play roles in that model. I will leave Jockers’ work untouched. All I’m doing is reframing it.

Before doing that, however, I should note that in the last quarter of a century or so there has been quite a lot of work on cultural evolution in a variety of discipline including linguistics, anthropology, archaeology, and biology. Though it must be done at some time, I have no intention of even attempting to review that work here and so to place the scheme I propose in relation to it. That’s a job for another time and another venue. I note, however, that I have done quite a bit of work on cultural evolution myself and that some of that discussion can be found in documents I list at the end of this post.

Why Evolution?

First of all, why bother to recast the processes of literary history in evolutionary terms at all? Jockers wrote an excellent book without creating an evolutionary model, though he mentioned evolution here and there. What’s to be gained by this recasting?

As far as I can tell, much of the work that has been done on cultural evolution has been undertaken simply to exercise and extend the range of evolutionary discourse. It has not, as yet, resulted in an understanding of cultural process that is deeper than more conventional forms of historical discourse. Much of my own work has been undertaken in this spirit. I believe that, yes, at some point, evolutionary explanation will prove more robust that other forms of explanation, but we’re not there yet.

This work in effect is looking to evolutionary accounts as exhibiting something like formal cause in Aristotle’s sense. Evolutionary accounts are about distribution of traits across populations. In biology such accounts have a characteristic formal appearance so that, e.g. phylogenetic analysis of a population of entities tends to “look” a certain way. So, in the cultural sphere, let’s conduct a similar analysis and see how things look even if we don’t have our entities embedded in the kind of causal framework that genetics and population biology, molecular biology, and developmental biology provide the biologist.

That’s fine, as long as we remind ourselves periodically that that’s what we’re doing. But we must keep looking for the terms in which to construct a causal model.

What I specifically want from an evolutionary approach to culture is
  • a way to think about Said’s autonomous aesthetic realm,
  • a way to prove out Shelley’s assertion that “poets are the unacknowledged legislators of the world,”
  • a way of restoring agency to writers and readers rather than casting them as puppets of various vast and impersonal forces, and
  • a way of thinking about the canon in relation to the whole of literary culture.
That’s what I want. Those requirements imply having a causal model. Whether or not I’ll get it, that’s another matter.

Current critical approaches, however, in which individual humans are but nodal points in the machinations of vast and impersonal hegemonic forces, have trouble on all these points. Individual human beings are deprived of agency thus turning readers into zombies watching the ghosts of dead authors flicker on the remaining walls of Plato’s cave. The canon is captive to those same hegemonic forces, which have promulgated Shelley’s defense as an opiate for the masses, which R’ us.

The critical machine is broken. It’s time to start over. Before we do that, however, I need to dispense with one objection to seeking an evolutionary account of cultural phenomena.