Showing posts with label topic models. Show all posts
Showing posts with label topic models. Show all posts

Tuesday, January 23, 2024

GOAT Literary Critics: Part 3.1, René Girard prepares the way for the French invasion

I had originally intended this essay to cover three thinkers, René Girard, Jacques Derrida, and Claude Lévi-Strauss. As I began working on it the thinking grew knotty and the prose just grew and grew. So I’ve decided to break the essay into two parts. In this part I begin by negotiating a transition from the previous article in this series, which was about how the contemporary discipline of literary criticism emerged after World War Two. I then say a little about my years at Johns Hopkins by way of introducing our three thinkers. After that comes some empirical evidence about the mid-century transition in literary criticism. I then conclude with a look at René Girard. I’ll discuss Jacques Derrida, and Claude Lévi-Strauss in the next part and make some remarks about the aftermath of this mid-century Sturm und Drang.

Do we really have a discipline?

As the previous article ended, Michael Bérubé was informing us that that the disciplinary regime set in place by Northrup Frye would begin unraveling a decade later. That’s what this article is about. But I want to begin by reviewing where we’ve been.

Let us start with Cowen and his interest in the greatest economists. Cowen took the disciplinary existence of economics as a given: In the beginning there was Adam Smith, and the rest followed after. He had to do a bit of tap dancing to fit John Stuart Mill into his scheme, for we generally think of Mill as a philosopher, not an economist, but that is easily done. One might, I suppose, do the same for literary criticism. The deparments exist in colleges and universities; just go back as far as you can.

However, administrative continuity is one thing; intellectual continuity is another. The conceptual focus of academic literary criticism changed in the middle of the previous century. That’s what the previous article is about. Sure, the primary texts are still there, but what academics do with them has changed. There are continuities, to be sure, there always are. Think of astronomy – Lord! I hate what I’m about to do – and its Copernican revolution. The earth, moon, sun, and the other planets are the same things they were before and after the revolution, but our understanding of their relationships has changed. Something like that is what has happened to literary criticism, and the discipline almost knows it and is still thrashing about with the consequences.

Brooks & Warren, taken as a duo, are important because they focused the discipline’s attention on the texts themselves in the most concrete way possible, by gathering a bunch of them together for a reader intended for undergraduates. That in turn brought the teachers of those undergraduates to think about those texts in a new way, to search for the meaning held within each text. In his Anatomy of Criticism Northrup Frye both conducted an inductive survey of the literary field and, in his polemical introduction, explained how an interpretive focus on that field constituted a proper academic discipline. Then, a century and a half before them, Coleridge introduced concepts that brought the literary mind into conceptual focus. That’s a discipline.

It's the focus on meaning that became problematic. For one thing it turns out that critics kept coming up with different meanings for the same texts. How can we call ourselves an academic discipline if we can’t agree on the central objects of our discipline? The problem occasioned a lot of thinking about theory and method. What’s even more problematic, in time it became conceptually difficult to separate the roles of critics and writers in the literary system, if you will. It wasn’t at all obvious that that would happen, and it took a while for awareness of problem to come into view.

It’s in that context that we should consider the well-known symposium that took place at Johns Hopkins in the Fall of 1966: The Languages of Criticism and the Sciences of Man. Notice that phrase, “the Sciences of Man,” from the French “les Sciences de l’Homme.” The concept of the human sciences is European, not American, and covers a range of disciplines that would be divided between the humanities and society sciences in America.

In retrospect that event is recognized as a ‘tipping point’ in the course of American literary criticism. Such things do not tip in the course of four days (October 18-21). They are the culmination, in this case, of a decade of uncertainty about the conceptual nature of literary criticism. Before taking a look at that 1966 symposium, however, I want to step back and insert myself at the edges of the narrative.

A change in perspective

I wrote the previous posts in this series from the “view from nowhere” point-of-view that is the default stance for much intellectual writing. That stance really isn’t available to me for this post, which are about developments that happened early in my career, during my undergraduate years at Johns Hopkins and my graduate study in the English Department at the State University of New York at Buffalo. While I certainly did not play an active role in the events I will be describing, I was an interested, concerned, and involved bystander (see the appendix, "Skin in the game," in the next installment).

Though I never studied with René Girard, I heard him lecture in classes I took with Richard Macksey when I was an undergraduate at Johns Hopkins and I later met with him in connection with book collection that never came to fruition (it was to be a collection of structuralist essays on Shakespeare). The ideas of Jacques Derrida influenced me during my undergraduate years, especially “Structure, Sign, and Play in the Discourse of the Human Sciences,” the paper he delivered at the famous 1966 structuralism symposium. However, the influence of Claude Lévi-Strauss’s work on myth was to prove more intriguing, influential and enduring.

Those are the three thinkers at the center of this part of our story: René Girard, Jacques Derrida, and Claude Lévi-Strauss. All of them were French, through Girard spent most of his life in America. Jacques Derrida was frequently in residence at Yale and other schools and ended his career at the University of California at Irvine. The Nazi occupation of America caused Lévi-Strauss to leave France for in America between 1941, where he stayed until 1947 and then returned to France in 1948. None of them were primarily literary critics.

To be sure, Girard sojourned in literary studies from the late 1950s on into the 1970s. But he had trained as a historian and, by the time of that structuralism conference in the mid-1060s, he was moving through anthropology to become a grand social theorist. Derrida was a philosopher, though he often commented on literary texts. And Lévi-Strauss was an anthropologist, though he famously collaborated with Roman Jakobson, the great linguist, on an analysis of Baudelaire’s “Les Chats.

What, then, are three non-literary critics doing in a series of posts ostensibly about the GOAT literary critics? They are influencing the course of literary criticism, that’s what. And that’s ultimately what this series of posts is about (for me). Why is literary study like it is? That the answer to that question depends critically on major thinkers outside of literary studies, that tells us something about the peculiar nature of a discipline focused on teasing out the meaning of literary texts.

Turning toward our three thinkers, Derrida arguably had more influence on literary criticism as a whole after 1970 than anyone trained as and writing primarily as a literary critic. It’s not clear to me that that is true of Lévi-Strauss, though he certainly had an influence and Derrida arguably made his bones with a famous essay about Lévi-Strauss. As for Girard, Tyler Cowen thinks he’ll go down in intellectual history as one of the major French thinkers of the last half century. Perhaps so. But that would be more in his person as a social theorist than as a literary critic. There his influence has been real, but limited. Thus he plays a somewhat different role in this story.

Note: In Appendix 1 I present some empirical evidence about the relative importance of these three thinkers.

Transition, the 1970s in literary criticism

Before discussing these three thinkers, however, I want to talk about what happened to academic literary criticism during this period in a general and empirical way. This is possible because of an important article that Andrew Goldstone and Ted Underwood published in 2014 in New Literary History: “The Quiet Transformations of Literary Studies: What Thirteen Thousand Scholars Could Tell Us”. Using topic modeling, a machine learning technique, Goldstone and Underwood analyzed complete runs of seven mainstream journals going back to the late nineteenth century.

This is not the place to explain how topic analysis is done – you’ll find an explanation in the article, or you may consult Appendix 2, where you’ll find what ChatGPT said about topic modeling. In this context “topic” is a term of art and refers to a group of words that tend to occur together in a collection of documents. It is up to the analyst to interpret the overall meaning of the topic.

This chart shows how Topic 16 evolved over time:

Here are the words most prevalent for that topic; criticism work critical theory art critics critic nature method view. The topic seems to be oriented toward method and theory and is most prevalent between roughly 1935 and 1985. The first half of that period likely reflects the rise of the so-called New Criticism while the second half reflects the developments I’ll be examining in this post, which came to a head at the time of that 1966 structuralism conference at Johns Hopkins. I should note as well that the study of poetry was prominent during the era of the New Criticism while disciplinary interested shifted toward the novel in the post-structuralist era. This is not covered by Goldstone and Underwood. I take it as a casual observation from personal correspondence with Franco Moretti.

Now look at Topic 20. Judging from its most prominent words, it seems weighted toward current critical usage: reading text reader read readers texts textual woolf essay Virginia.

That usage, reading as interpretation, becomes more obvious when we compare it with Topic 117, which seems to reflect a more prosaic usage, where reading is mostly just reading, not hermeneutics: text ms line reading mss other two lines first scribe. Note in particular the terms – ms line mss lines – which clearly reference a physical text.

This topic is most prominent prior to 1960 while Topic 20 rises to prominence after 1970. For what it’s worth – not much, but it’s what I can offer – that accords with my personal sense of things going back to my undergraduate years at Johns Hopkins in the later 1960s. I have vague memories of remarking (to myself) on how odd it seemed to refer to interpretative criticism as mere reading when, really, it wasn’t that at all.

Reading isn’t the only word whose meaning shifted during that period. Theory changed as well. Up into the 1960s and even the early 1970s “literary theory” was thinking about the nature of literature, as exemplified by the venerable Theory of Literature (1949), by René Wellek and Austin Warren. By the mid-1970s or so literary theory had come to refer to the use of some kind of theory about mind and/or society in the interpretation of literary texts. That is to say, it wasn’t theoretical discourse about literature. Rather it was a method for creating interpretations of texts, which are readings in the newer sense of the term. That’s where Derrida and Lévi-Strauss come in, as sources of interpretive tools, along with Marx, Freud, Adorno, Barthes, Deleuze, Foucault, and a cast of thousands. Well, I exaggerate the number, but you get the idea.

The rest of this essay, especially the second part coming up in a later post, is about those shifts.

Finally, and in view of recent events at Harvard and elsewhere, I must admit that I’ve ‘plagiarized’ from myself in what I’ve said about Topics 20 and 177. Those words come from a working paper that has a more complete discussion of this transitional period based in part on the work of Goldstone and Underwood. Here’s the paper:

Transition! The 1970s in Literary Criticism, Version 2, January 2017, https://www.academia.edu/31012802/Transition_The_1970s_in_Literary_Criticism

Rene Girard and Mimetic Desire

René Girard was the first professor I heard lecture when I entered Johns Hopkins in the fall of 1965. At one point during freshman orientation we were given a choice of attending one of five lectures. I attended Girard’s lecture on cultural relativism – I forget what the other four were. I remember only two things from that lecture: Girard spoke with an accent and it was the craziest damn thing I ever heard. Yes, that’s how I thought of it at the time, but this guy was a Hopkins professor, so there must be something to it.

In the spring of 1966 I took Richard Macksey’s course on the autobiographical novel. He invited Girard to lecture on mimetic desire, the central idea of his first book, published as Mensonge romantique et verité in 1961 and translated into English as Deceit, Desire, and the Novel in 1966. Here’s a brief account of the argument by John Pistelli:

He posits, therefore, a fundamental geometric relation governing the plots of great novels: the protagonist seems to desire something (wealth, status, a lover, etc.), but in fact really desires the object through a mediator whom the protagonist wishes to emulate or usurp. Desire, then, is mimetic—in the sense of mimicry—and triangular—because it doesn’t go from subject to object but from the subject through the mediator to the object. Girard’s first and clearest example is Don Quixote, who learns to want what a chivalric knight wants through reading about the hero Amadis of Gaul in medieval romances; his quest is less to possess the knight’s rewards than to become Amadis through this possession. Given Don Quixote’s reputation as the inaugural European novel, it’s no surprise to find that the pattern continues in later fiction: Girard’s main examples are the works of Stendhal, Flaubert, Dostoevsky, and Proust.

That book put Girard on the intellectual map, making him a force to be reckoned with.

And he used that force to organize the 1966 symposium at Johns Hopkins: The Languages of Criticism and the Sciences of Man. Girard provided the visibility and the international connections while Richard Macksey and Eugenio Donato did much of the actual organizing. Here’s a few observations Cynthia Haven offered about that symposium in the chapter, “The French Invasion,” in her biography of Girard, Evolution of Desire: A Life of René Girard (2018):

At that historical moment, “structuralism” was the height of intellectual chic in France, and widely considered to be existentialism’s successor. Structuralism had been born in New York City nearly three decades earlier, when French anthropologist Claude Lévi-Strauss, one of many European scholars fleeing Nazi persecution to the United States, met another refugee scholar, the linguist Roman Jakobson, at the New School for Social Research. The interplay of the two disciplines, anthropology and linguistics, sparked a new intellectual movement. Linguistics became fashionable, and many of the symposium papers were cloaked in its vocabulary.

Girard never saw himself as a structuralist. “He saw himself as his own person, not one of the under-lieutenants of structuralism,” said Macksey. Yet structuralism would have had a natural pull for Girard, who was already moving away from literary concerns and toward more anthropological ones by the time of the symposium. Indeed, in this as in other matters, he was indebted to the structuralists. His own metanarratives strove toward universal truths, akin to the movement that endeavored to discover the basic structural patterns in all human phenomena, from myths to monuments, from economics to fashion.

It's that drive toward “more anthropological” issues that interests me.

For it is through his readings in anthropology that Girard was able to enlarge mimetic theory to embrace society-wide dynamics of conflict and the resolution of conflict through the mechanism of sacrifice. Just how he was able to do that is not my concern here, any more than I was concerned about the validity of Northrup Frye’s archetypal criticism in the previous essay in this series. That led Girard to write his second book, Violence and the Sacred, which came out in English translation in 1977. And from there, apparently – for my own attention was elsewhere – he came up with what is somewhat derisively called a “theory of everything,” albeit one that was convincing enough that he was elected to the Académie Française in 2005. I prefer to think of him has a grand social theorist in the 19th century manner of, say, Karl Marx, Herbert Spencer, or somewhat later, Oswald Spengler. Whether history will remember him as such in 50 years, that’s another matter.

“But,” one might ask, “why did he do it?”
“Well,” might come the response, “there is curiosity.”
“But that’s not enough given the strictures of the modern academy.”
“But it wasn’t always thus, was it? Earlier thinkers had more latitude, no?”
“Yes, but...”

Monday, July 20, 2020

Advances in topic modeling [digital humanities]

Thursday, August 23, 2018

Meaning, Theory, and the Disciplines of Criticism

Relevant to a discussion about the rise of "stars" in Academic lit crit in the 1980s and 90s. See especially discussions of Reading and Theory as Critique below, which speak to the erosion of boundaries that allows critics to become stars.
In the fifth post, It’s Time to Leave the Sandbox, in my series on the poverty of cognitive criticism I managed to rough out a sketch of academic literary criticism that I rather like. So I’ve decided to present that sketch in a post of its own. I’ve dropped most of the comments on cognitive criticism and added a few things.

The most important additions come from the study Andrew Goldstone and Ted Underwood did of a century-long run of articles in seven journals, The Quiet Transformation of Literary Studies [1]. These are the journals they used: Critical Inquiry (1974–2013), ELH (1934–2013), Modern Language Review (1905–2013), Modern Philology (1903–2013), New Literary History (1969–2012), PMLA (1889–2007), and the Review of English Studies (1925–2012).

I begin by presenting one aspect of the argument that Goldstone and Underwood make, that our current terms of critical art only because established after World War II. Then I present the material I’ve reworked from that older post, this time with some charts from Goldstone and Underwood, and so other things as well.

Note: If topic models are something of a mystery. I (attempt to) explain them in Topic Models: Strange Objects, New Worlds.

Post War Shift

Goldstone and Underwood demonstrate that perhaps the most significant change in the discipline happened in the quarter century or so after World War II. Here’s a summary statement (p. 372):
The model indicates that the conceptual building blocks of contemporary literary study become prominent as scholarly key terms only in the decades after the war—and some not until the 1980s. We suggest, speculatively, that this pattern testifies not to the rejection but to the naturalization of literary criticism in scholarship. It becomes part of the shared atmosphere of literary study, a taken-for-granted part of the doxa of literary scholarship. Whereas in the prewar decades, other, more descriptive modes of scholarship were important, the post-1970 discourses of the literary, interpretation, and reading all suggest a shared agreement that these are the true objects and aims of literary study—as the critics believed. If criticism itself was no longer the most prominent idea under discussion, this was likely due to the tacit acceptance of its premises, not their supersession.
Consider topic 16, where these are the ten most prominent words: criticism work critical theory art critics critic nature method view. This graph shows how that topic evolved in prominence over time:

criticism 16

Now consider topic 29: image time first images theme structure imagery final present pattern. Notice “structure” and “pattern” in that list. The (in)famous structualism conference was held at Johns Hopkins in 1966, just before the peak in the chart, while the conference proceedings (The Languages of Criticism and the Sciences of Man) were published in 1970, one year after the peak, 1969:

image pattern 29
At the time structuralism was regarded as The Next Big Thing, hence the conference. But Jacques Derrida was brought in as a last minute replacement and his critique of Lévi-Strauss turned out to be the beginning of the end of structuralism and the beginning of the various developments that came to be called capital “T” Theory. The structuralist moment, with its focus on language, signs, and system, was but a turning point.

Thursday, December 15, 2016

Darwin was an exploratory forager: Topic modeling our way through his notebooks and the books he mentions there

 2016 Dec 8;159:117-126. doi: 10.1016/j.cognition.2016.11.012. [Epub ahead of print]

Exploration and exploitation of Victorian science in Darwin's reading notebooks.

Abstract

Search in an environment with an uncertain distribution of resources involves a trade-off between exploitation of past discoveries and further exploration. This extends to information foraging, where a knowledge-seeker shifts between reading in depth and studying new domains. To study this decision-making process, we examine the reading choices made by one of the most celebrated scientists of the modern era: Charles Darwin. From the full-text of books listed in his chronologically-organized reading journals, we generate topic models to quantify his local (text-to-text) and global (text-to-past) reading decisions using Kullback-Liebler Divergence, a cognitively-validated, information-theoretic measure of relative surprise. Rather than a pattern of surprise-minimization, corresponding to a pure exploitation strategy, Darwin's behavior shifts from early exploitation to later exploration, seeking unusually high levels of cognitive surprise relative to previous eras. These shifts, detected by an unsupervised Bayesian model, correlate with major intellectual epochs of his career as identified both by qualitative scholarship and Darwin's own self-commentary. Our methods allow us to compare his consumption of texts with their publication order. We find Darwin's consumption more exploratory than the culture's production, suggesting that underneath gradual societal changes are the explorations of individual synthesis and discovery. Our quantitative methods advance the study of cognitive search through a framework for testing interactions between individual and collective behavior and between short- and long-term consumption choices. This novel application of topic modeling to characterize individual reading complements widespread studies of collective scientific behavior.

KEYWORDS: 

Cognitive search; Exploration-exploitation; History of science; Information foraging; Scientific discovery; Topic modeling

H/t Tyler Cowen.

Friday, September 25, 2015

Goldstone and Underwood, #149: Foreign Language Education

Another post based on: Andrew Goldstone and Ted Underwood , “The Quiet Transformations of Literary Studies: What Thirteen Thousand Scholars Could Tell Us”, New Literary History 45, no. 3, Summer 2014. Website: http://andrewgoldstone.com/blog/2014/05/29/quiet/
Topic 149 is interesting: language university students english education teaching study modern foreign school. That looks like foreign language education.

The temporal distribution is interesting as well – and it’s the only topic with this three-peak distribution (you can see thumbails of all topic distributions here):

149 language universtiy

That middle rise spans the 1950s and 1960s, the early decades of the Cold War. In 1957 the Russian’s put the first artificial satellite into earth orbit, Sputnik, and that spurred the government to put money into higher education.

Here we’ve got the top ten articles for the topic. Notice that they’re all from that middle period and that they’re all about foreign language education.
  1. Paquette, F. André. "Developing Guidelines for Teacher Education Programs in Modern Foreign Languages." PMLA 81, no. 2 (May 1966): 3–6.
  2. [Anon]. "English Teacher Preparation Study Guidelines for the Preparation of Teachers of English." PMLA 82, no. 4 (September 1967): 19–25.
  3. Hook, J. N. "Project English: The First Year." PMLA 78, no. 4 (September 1963): 33–35.
  4. [Anon]. "Qualifications for Secondary School Teachers of Modern Foreign Languages." PMLA 70, no. 4 (September 1955): 46–49.
  5. [Anon]. "Modern Foreign Languages in the Comprehensive Secondary School." PMLA 74, no. 4 (September 1959): 27–33.
  6. Walsh, Donald D. "The MLA FL Program in 1962." PMLA 78, no. 2 (May 1963): 20–24.
  7. Mildenberger, Kenneth W. "The MLA College Language Manual Project: History and Present Status." PMLA 72, no. 4 (September 1957): 11–18.
  8. Walsh, Donald D. "The Foreign Language Program in 1964." PMLA 80, no. 2 (May 1965): 29–32.
  9. Shugrue, Michael F., and Thomas F. Crawley. "The Conclusion of the Initial Phase: The English Program of the Usoe." PMLA 82, no. 6 (November 1967): 15–32.
  10. [Anon]. "The Preparation of College FL Teachers." PMLA 70, no. 4 (September 1955): 57–68.

Thursday, September 24, 2015

Topic Analysis: Goldstone & Underwood and the "economy" of word distribution across topics

Andrew Goldstone and Ted Underwood , “The Quiet Transformations of Literary Studies: What Thirteen Thousand Scholars Could Tell Us”, New Literary History 45, no. 3, Summer 2014. Website: http://andrewgoldstone.com/blog/2014/05/29/quiet/
A week or so ago when I was examining the topic model Goldstone and Underwood used in their study of the shifting interests of academic literary critics I became curious about the topics that accounted for the largest number of word tokens in the entire corpus. I’d noticed one, for example, that accounted for 8.7% of the tokens. What’s that about? What about the next highest three topics?

So I looked them up and here they are, the top four arranged in that order, from top to bottom, but without any information to identify what’s in the topics:

top4

These charts are quite different. The middle two are skewed to the left and the top and bottom one are skewed to the right, the though bottom one, while generally to the right, is strongest toward the center.

One would expect them to be quite different. But they’re not. This table lists the words in the topic along with the totals for the corpus:

top 4 total

Even a casual look at those words, only the top 10 for each topic, reveals quite a bit of overlap. This table gives more detail:

top 4 analysis

We’ve got only 22 different words. Two of them, “even” and “made”, are present in all four topics while two others, “other” and ”same”, are present in three topics. If these topics are so apparently similar, why are they distributed so differently in time?

Curious about this, I sent an email to Ted Underwood that included an earlier version of this document, suggesting that the difference has to be in the words at the middles and the bottom ends of the topics, which cannot be easily examined in the online presentation. I noted:
Those top ten words seem to be semantically neutral, sort of "connector" words. They don't identify any substantive intellectual interests. These topics must somehow be complements to those topics which show strong thematic interests concentrated either early in the discipline's history, or late.
Here’s Underwood’s reply:
I think you're really right about the problem here. These large topics are representing something of historical significance, but the standard way of labeling them doesn't reveal what.

One thing I tried (though it didn't make it into the final article) was labeling topics by selecting words that are prominent in the topic and that also correlate with its trajectory over time. That gets you down below the neutral "connector" words at the top of the list, and reveals what it is that's changing over time. I believe that labeling strategy makes these big "glob topics" a lot more interpretable. It works well with smaller topics too. But this is also a kind of an ad hoc trick, and we didn't want to complicate the story in the article.
So, let’s think about this for a minute. If we look at the topics that skew to the right (that is, the present) in these charts, they tend to be Theory-laden and are heavily about society and politics, as Goldstone and Underwood pointed out – I’ve presented a few of them below. Those that skew right (the past), are quite different in character – I’ve presented some of them as well.

Thursday, August 21, 2014

Reading Macroanalysis 6.2: Theme, Moby Dick in the Context of Literary Culture

In this post, which is a long one, I use two books to investigate Jockers’ themes, and vice versa. One of the books is a classic of fairly traditional, at least by now, literary criticism, Leslie Fiedler’s Love and Death in the American Novel. The other book is one that Fiedler examined, Herman Melville’s Moby Dick.

First I open with some of Jockers' charts, then I consider a particular passage from Moby Dick, and Fiedler’s gloss on it. I then return to Jockers, looking for themes and charts that resonate with or respond to the passage. What I’m up to is, in effect, using Fiedler to guide me in a close reading of Jockers’ distant reading, a term, by the way, that he doesn’t use. I’m using it, of course, for rhetorical purposes.

Note: I downloaded all the charts from Matt Jockers’ Macroanalysis website.

An Outlier in the 19th Century

Look at this graph:

PACIFC ISLANDS AKA SEAS AND WHALING nation year

That spike in the middle is Moby Dick. Not literally of course. But it is reasonable to believe that that spike in the data is caused mostly by Moby Dick–where cause is to be read as Aristotelian formal cause. The timing is right; Moby Dick was published in 1851, which is roughly where that spike peaks. The light grey line shows the use of the topic by Irish authors in the 19th century; the darker grey line shows use by British authors. And the black line, of course shows American authors.

And the topic is typical of Moby Dick. In his book Jockers calls it Seas and Whaling; on his website it’s PACIFC ISLANDS AKA SEAS AND WHALING. Here’s the corresponding word cloud:

PACIFC ISLANDS AKA SEAS AND WHALING cloud

Above island and islands in the middle you can see whale while near the top in the middle you can see Queequeg.
Note: for reasons that Jockers explains in the book, proper names are generally eliminated from topic models because they can recur in many contexts without there being any thematic similarity between those contexts; this one slipped through.
This particular topic constitutes almost 20% of the text of Moby Dick, though it constitutes a very small faction of the corpus as a whole (p. 130 and Figure 8.4 p. 132). Other topics having to do with ships and the sea are also prominent in the book while being relatively unimportant in the whole corpus of 3.346 British, American, and Irish novels.

Moby Dick is, as Jockers says, an outlier. It is also one of the world’s great novels. Back in 1976 Edward Mendelson included it as one in a very small and exclusive genre he called the encyclopedic narrative (“Encyclopedic Narrative: From Dante to Pynchon,” MLN 91, 1267-1275). The other examples are: Dante’s Divine Comedy, Rablais’ Gargantua and Pantagruel, Cervantes’ Don Quixote, Goethe’s Faust, Joyce’s Ulysses, and Pynchon’s Gravity’s Rainbow. In Mendelson’s account each of these works is unique in its national canon and each is encyclopedic in scope, encompassing the world as it was known at the time. (Mendelson does have a ‘fix’ for the fact the America gets two such books; he links Gravity’s Rainbow to a newly forming international culture.)

Tuesday, August 19, 2014

Reading Macroanalysis 6.1: Theme–Dogs, Gold, Slavery, and Awakening

Chapter 8 of Macroanalysis is about “Theme.” Jockers uses topic analysis to investigate the occurrence of 500 ‘themes’ in a corpus of 3,346 19th-century British, American, and Irish books. He opens with a bit of intellectual history, from the Russin Formalists to Google’s Ngrams; then he launches into topic analysis, which emerged at the turn of the millennium he gives some simple examples, and then he gets serious.

But I’m going to skip over all of that for now. For one thing, I’ve been through the topic analysis drill several times in the past year or so and don’t want to go through it again. If you need an introduction or a review, check out Topic Models: Strange Objects, New Worlds, or, in this series, Reading Macroanalysis 5: An Interlude on Scale: Micro, Meso, and Macro. For another, Jockers has put a topic tool online, 500 Themes from a corpus of 19th-Century Fiction. Those are the topics he discusses in this chapter.

Once I was done reading the chapter I started playing with the tool. I’d pick a topic and then look at the graphics:
  1. a word cloud to display the most frequent words in the topic,
  2. a bar chart indicating usage of the topic by author gender (male, female, and undetermined),
  3. a line graph showing gender usage over time,
  4. a bar chart indicating usage of topic by author nationality (American, British, Irish).,and
  5. a line graph showing national usage over time.
At first I was just browsing, moving from one theme to the next. But then I hit one that grabbed my attention. So I spent the next couple of hours looking at themes and thinking about them.

I’m going to devote the rest of this post and the next one showing what I found. Then I’ll do a third post where I review what Jockers found and recast the enterprise in terms of cultural evolution. Note that in all of this I’m just playing around, but in a serious way. It is all preliminary and provisional. I haven’t reached any firm conclusions on the particular themes I look at. The only thing I’m sure about is that this, and similar techniques, are going to revolutionize the way we do literary history.

Before proceeding on, however, two caveats are necessary. While the Jockers’ is substantial it isn’t every British, American, and Irish novel written in the 19th Century. Perhaps more important, it is natural to read these theme charts as reflecting the interests of the 19th Century reading public. And in some sense that is so.

But we have to be careful. For some of these books were more widely read than others and a few of them, the canonical ones, are still being read. But the extent of a books’ readership is not reflected in the data. The fact that a book was published at all implies, of course, that someone thought there was an audience for it. But a publisher’s interest isn’t quite the same as a reader’s interest. We simply don’t know how accurately publisher interest tracks reader interest.

With those reservations in mind, let’s take a look.

Of Dogs and Gold

In the course of browsing through Jockers’ themes menu I saw “DOGS.” Let’s look at that, I thought. Why dogs? you may ask. No deep reason, but some years ago, way back in graduate school in fact, I’d noticed that dogs figured as a significant motif in Wuthering Heights. Major transitions among humans were marked by violence between dogs and humans (e.g. Lockwood arrives and is greeted by a barking dog, Catherine gets bitten by Skulker; see this post). More recently, I’d read a handful of articles about the domestication of dogs during human evolutionary history. I was just curious.

Here’s the word cloud for the DOGS topic:

dog cloud

Friday, August 15, 2014

Reading Macroanalysis 5: An Interlude on Scale: Micro, Meso, and Macro

Before moving on to the last two major chapters, “Theme” (8), and “Influence” (9), I want to pause a bit and think about scale as I discussed it Toward a Computational Historicism. Part 1: Discourse and Conceptual Topology and a consequent Working Paper on Digital Historicism. While Jockers focuses on the macroscale, large populations of texts, he is also working at the mesoscale or the individual text, and his analytical work implies microscale phenomena.

From Micro to Meso: Paths in a Network

It has become common for cognitive scientists to think of the mind as a cognitive, or associative, network:

network

The nodes represent concepts, or even words, while the arcs or edges represent relations. We can now think of an utterance or a written statement as taking a path through the network:

path

An entire novel is simply a very long path, one that will pass through a large area of the network and that will go through various subnetworks many times, often along different paths and orientations.

Let us posit that the way an author moves through the net is the author’s style. The function words that are so very useful as features in identifying style can be thought of as regulating movement from one node (content) to another. That’s why they are so very useful in stylistic analysis.

No matter what you write about, you have to use those function words. Your choice of content words varies as your topic varies, but the set of function words is quite limited and you have to draw from that set regardless of topic. That is to say, while an author has to navigate many different subnetworks, that is, many different topics, the way they navigate each subnetwork is pretty much the same.

But why should that signal also be an individual one? Because, as the saying goes, there are many ways to skin a cat. That is, there are many ways to construct sentences and paragraphs about any given topic. Different writers will choose different ways of so doing.

Sunday, June 1, 2014

Hyperobjects and "Distant Reading"

Two successive tweets from Alan Liu:
So, topic models. Is the model the hyperobject or is it simply a description of the hyperobject? Consider one of Morton's core examples, global warming. Surely global warming is the hyperobject while a given climate model simply provides a description of that hyperobject. It's global warming that is "massively distributed in time and space," not a given climate model. The climate model is just a bunch of code running on a computer, which is fairly compact in time and space.

And so it must be with a topic model and whatever cultural hyperobject it models. Thus, when Goldstone and Underwood developed topic models for over a century's worth of articles in literary studies (The Quiet Transformations of Literary Studies: What Thirteen Thousand Scholars Could Tell Us  [PDF preprint]) literary studies is the hyperobject and, as such, is distinct from their topic models (which you can explore interactively HERE, introduction HERE). 

Sunday, May 4, 2014

Working Paper on Computational Historicism

I've now taken my recent series of four posts on computational historicism and edited them into a single working paper. This required some reorganizing. I took that Stephen Greenblatt material that had opened the second post and moved it to the beginning of the working paper. The paper is thus framed by Stephen Greenblatt on the directionality of literary history (not his phrase) at the beginning and Edward Said on the autonomous aesthetic realm at the end. I've also expanded and tweaked some of the other discussions. I've also added two appendices.

You can download the paper HERE (SSRN) and HERE (Academia.edu).  Abstract and contents below.


Toward a Computational Historicism: From Literary Networks to the Autonomous Aesthetic

Abstract: Stephen Greenblatt has identified pairs of moments in literary history such that the former moment must necessarily have preceded the later: literary history has a direction. This can be explained by asserting that the later texts required computational procedures capable of operating on the objects created by the earlier procedures, in the manner of Piaget’s reflective abstraction. Beyond Greenblatt’s examples two such pairs are examined, Amleth and Hamlet, The Winter’s Tale and Wuthering Heights. Also considered: Heuser and Le-Khac on 19th Century British novels. Three network-based network models are considered, at macro (topic modeling), meso (Moretti’s character networks), and micro scales (cognitive networks) of time.

Contents
Part 1: Greenblatt’s Difference: Time and History
Part 2: Conceptual Topology and Discourse
Part 3: Spatializing Time: Sonnet 129
Part 4: Abstraction at the Time Scale of History
Part 5: Into the Autonomous Aesthetic
Appendix 1: Hamlet as Ring, The Winter’s Tale also
Appendix 2: Nine Propositions in Computational Criticism

Monday, April 21, 2014

Toward a Computational Historicism. Part 1: Discourse and Conceptual Topology

Poets are the unacknowledged legislators of the world.
– Percy Bysshe Shelley

... it is precisely because we are talking about ordinary language that we need to adopt a notation as different from ordinary language as possible, to keep us from getting lost in confusion between the object of description and the means of description.
–Sydney Lamb


Worlds within worlds – that’s how Tim Perper, my friend and colleague, described biology. At the smallest scale we have individual molecules, with DNA being of prime importance. At the largest scale we have the earth as a whole, with all living beings interacting in a single ecosystem over billions of years. In between we have cells, tissues, and organs of various sizes, autonomous organisms, populations of organisms on various scales from the invisible to continent-spanning, and interactions among populations of organisms on various scales.

Literature too is like that, from single figures and tropes, even single words (think of Joyce’s portmanteaus) through complete works of various sizes, from haiku to oral epics, from short stories through multi-volume novels, onto whole bodies of literature circulating locally, regionally, across continents and between them, from weeks and years to centuries and millennia. Somehow we as humanists and literary critics must comprehend it all. Breathtaking, no?

In this essay I sketch a potential computational historicism operating at multiple scales, both in time and textual extent. In the first part I consider network models on three scale: 1) topic models at the macroscale, 2) Moretti’s plot networks at the mesoscale, and 3) cognitive networks, taken from computational linguistics, at the microscale. I give examples of each and conclude by sketching relationships among them. I open the second part by presenting an account of abstraction given by David Hays in the early 1970s; in this model abstract concepts are defined over stories. I then move on to Hauser and Le-Khac on 19th Century novels, Stephen Greenblatt on self and person, and consider several texts, Amleth, Hamlet, The Winter’s Tale, Wuthering Heights, and Heart of Darkness.

Graphs and Networks

To the mathematician the image below depicts a topological object called a graph. Civilians tend to call such objects networks. The nodes or vertices, as they are called, are connected by arcs or edges.

net

Such graphs can be used to represent many different kinds of phenomena, a road map is an obvious example, a kinship tree is another, sentence structure is a third example. The point is that such graphs are signs of phenomena, notations. They are not the phenomena itself.

Thursday, July 18, 2013

Corpus Linguistics for the Humanist: Notes of an Old Hand on Encountering New Tech

I've published another working paper, this one on "digital humanities." Specifically, corpus linguistics. Here's the abstract and the introduction:
Abstract: Corpus linguistics offers literary scholars a way of investigating large bodies of texts, but these tools require new modes of thinking. Literary scholars will have to recover a kind of interest in linguistics that was lost when the discipline abandoned philology. Scholars will need to think statistically and will have to start thinking about cultural evolution in all but Darwinian terms. This working paper develops these ideas in the context of a topic analysis of PMLA undertaken by Ted Underwood and Andrew Goldstone.
Introduction: Theory in a Digital Age

In reading about so-called “digital humanities” over the last year or two I kept coming up against the question: What about Theory? In the context of academic humanities over the last three or four decades the term “theory” does not mean quite what it means elsewhere. In particular, in literary studies it does not mean what “literary theory” once meant: a body of theory about literature, its texts, writers, readers, history, and social influence. Rather, Theory (often capitalized) is a body of techniques for interpreting texts, for explicating their meanings in terms of some body of thinking about the mind, society, history, or general comport of the cosmos.

What about Theory, then, is a plea to link these new techniques to older concerns, to the concerns of ethical criticism, as broadly construed by Wayne Booth in The Company We Keep: An Ethics of Fiction. Ethical criticism is a worthy, indeed a necessary enterprise, but it is not the only worthy enterprise one can imagine. It is not the only form of knowing.

And it is not one the follows naturally from these new computer-enabled modes of inquiry. Hence the question: What about Theory? In the short-term my answer is: Forget about it! There’s nothing there at the moment. Perhaps later on, but not now.

Thursday, January 10, 2013

Topic Models: Strange Objects, New Worlds

In the weeks since writing about the preliminary results Goldstone and Underwood have reported of their work on topic analysis of PMLA (Publications of the Modern Language Association) I’ve continued to think about Natalia Cecire’s reservations. You may recall that she’s comfortable using topic analysis to point toward interesting texts that she will then examine for herself. But she had doubts about using it as evidence itself. For that, she’d need to have “a convincing theory of what the math has to do with the structure of (English) language” and that, presumably, requires some detailed knowledge both of language and of math.

I’m sympathetic to her reservations. What I’ve been thinking about is this: Just WHAT would she want to know?

Ted Underwood offered a response to her in which he referenced a Wikipedia article on distributional semantics, which I’ve read. He summarized that article, which is a short one, thus: “It’s just … things that occur in the same contexts gotta have something in common.” I agree with that summary.

And that’s the problem. That is what has had me thinking about this matter. Underwood’s statement is brief and easy to understand. What’s the problem?

I believe, in fact that “a convincing theory of what the math has to do with the structure of (English) language” would not be terribly useful to Cecire, not nearly so useful as simply playing around with topic analysis over a period of time by going back and forth between the computer-generated topics and associated texts. Only doing this time after time Cecile will be able to verify, for herself, that the computer-identified topics are meaningful entities. Though I don’t know this, I suspect that, whatever they may know about the computational processing behind topic analysis, such play has been important to both Goldstone and Underwood and to anyone else who uses the technique.

Thus I know that this, yet another run at explaining topic modeling, is bound to fail, as it cannot possibly substitute for the requisite experience. All I’m after is a different way to thinking about the technique and the strange conceptual objects it creates.

Bags of Words

I’ve read several accounts of topic modeling, this one by Matt Jockers, this one by Scott Weingart, this one by Ted Underwood and, finally, this technical review by David Blei: Probabilistic topic models [PDF] (Communications of the ACM, 55(4): 77–84, 2012). The first three were written for humanists while the last was written for computer scientists. It contained a most useful phrase, “bag of words.” That’s how the basic topic modeling technique, something called Latent Dirichlet Allocation (LDA) treats individual texts, as bags of words.

Monday, December 17, 2012

Literary History, the Future: Kemp Malone, Corpus Linguistics, Digital Archaeology, and Cultural Evolution

In scientific prognostication we have a condition analogous to a fact of archery—the farther back you draw your longbow, the farther ahead you can shoot.
– Buckminster Fuller

The following remarks are rather speculative in nature, as many of my remarks tend to be. I’m sketching large conclusions on the basis of only a few anecdotes. But those conclusions aren’t really conclusions at all, not in the sense that they are based on arguments presented prior to them. I’ve been thinking about cultural evolution for years, and about the need to apply sophisticated statistical techniques to large bodies of text—really, all the texts we can get, in all languages—by way of investigating cultural evolution.

So it is no surprise that this post arrives at cultural evolution and concludes with remarks on how the human sciences will have to change their institutional ways to support that kind of research. Conceptually, I was there years ago. But now we have a younger generation of scholars who are going down this path, and it is by no means obvious that the profession is ready to support them. Sure, funding is there for “digital humanities,” so deans and department chairs can get funding and score points for successful hires. But you can’t build a new and profound intellectual enterprise on financially-driven institutional gamesmanship alone.

You need a vision, and though I’d like to be proved wrong, I don’t see that vision, certainly not on the web. That’s why I’m writing this post. Consider it a sequel to an article I published back in 1976 with my teacher and mentor, David Hays: Computational Linguistics and the Humanist. This post presupposes the conceptual framework of that article, but does not restate nor endorse its specific visionary recommendations (given in the form of a hypothetical computer program, called Prospero, for simulating the “reading” of texts).

The world has changed since then and in ways neither Hays nor I anticipated. This post reflects those changes and takes as its starting point a recent web discussion about recovering the history of literary studies by using the largely statistical techniques of corpus linguistics in a kind of digital archaeology. But like Tristram Shandy, I approach that starting point indirectly, by way of a digression.

Who’s Kemp Malone?

Back in the ancient days when I was still an undergraduate, and we tied an onion in our belts as was the style at the time, I was at an English Department function at Johns Hopkins when someone pointed to an old man and said, in hushed tones, “that’s Kemp Malone.” Who is Kemp Malone, I thought? From his Wikipedia bio:
Born in an academic family, Kemp Malone graduated from Emory College as it then was in 1907, with the ambition of mastering all the languages that impinged upon the development of Middle English. He spent several years in Germany, Denmark and Iceland. When World War I broke out he served two years in the United States Army and was discharged with the rank of Captain.

Malone served as President of the Modern Language Association, and other philological associations ... and was etymology editor of the American College Dictionary, 1947.
Who’d have thought the Modern Language Association was a philological association?