Showing posts with label Moretti. Show all posts
Showing posts with label Moretti. Show all posts

Monday, November 18, 2019

2 Comments on Moretti’s LitLab 15: Patterns and Interpretation [#DH] [now 3 comments]

And now let me add a third comment based on a 2016 lecture which covers much of the conceptual ground in this pamphlet. I've added the third comment after the first two. Accordingly, I've bumped this post to the top of the queue.
Franco Moretti has produced another pamphlet:
Franco Moretti, Patterns and Interpretation, Stanford Literary Lab, Pamphlet 15, September 2017, 10 pp. [In Canon/Archive (2017) as the conclusion, pp. 295-308.]
I have two comments to offer. In the first I contrast meaning with explicit semantic models while I the second I suggest that Lévi-Strauss was in pursuit of the form of myths.

Meaning and Semantics

Early in the pamphlet Moretti remarks (p. 2):
Now, meaning is not one of the things literary critics study; it is the thing. Here lies the great challenge of computational criticism: thinking about literature, removing meaning to the periphery of the picture. But of course this is also the great challenge for computational criticism: you discard meaning and replace it with – what?
Yes.

Meaning is a peculiar beast. It is inherently subjective, by which I mean primarily that it arises within subjects; that subjectivity is its mode of existence. As a practical matter assertions about the meaning of texts varies widely and seemingly idiosyncratically from one critic to another, though agree is possible. The former is subjectivity in the ontological sense while the later is subjectivity in the epistemic sense – a distinction I have from John Searle [1]. That the meaning of texts is ontologically subjective implies that, as investigators, we cannot grasp as an object of examination and investigation. All we can do is translate, paraphrase, wheedle and cajole, in an attempt to indicate a text’s meaning(s). It’s a slippery Heraclitean business.

Moretti, if I understand him correctly, would have us, as computational critics, set meaning aside in favor of pattern and, I believe, form. I’m OK with that in the context of contemporary digital criticism. However, there is a different, though by no means contradictory, “replacement” for meaning.

Semantics –

By which I mean an explicit account of how words have, well, you know, meaning. Back in the 1960s cognitive scientists began to create such explicit accounts, often in the form of computer simulations. I learned one such model when I studied with the late David Hays in the Linguistics Department at SUNY Buffalo and employed it in an account of Shakespeare’s Sonnet 129, “The Expense of Spirit” [2]. The model took the form of a semantic or cognitive network where the nodes represented ‘words’ (approximately, more or less) and the edges connecting them represented the relations between them.

This diagram depicts a small fragment of such a model:

Shame s129 mln76

One could then ‘account for’ the text as a path through the semantic model. This depicts a fragment of such a path, which arrows pointing from the text to the network fragments they realize:

heaven&hell2

I say account for because such a path doesn’t actually explain the text. It doesn’t tell you why that path was taken in the first place, nor why it takes this form rather than that form. That’s a different matter.

My point is that such a model doesn’t “give” you the meaning of the text. When you examine such a model, when you follow a text through such a model, you are not reading the text in any reasonable sense of the word. You are simply examining the operations of a model, and a very complex one at that.

Tuesday, September 10, 2019

Reading Macroanalysis 7.3: Style, Genre, Time, and Influence

This post is from Aug. 31, 2014, but I'm bumping it to the top of the queue as I am thinking about these matters in connection with Moretti and Sobchuk, Hidden in Plain Sight: Data Visualization in the Humanities (New Left Review 118, 2019, 86-119). They don't discuss this visualization, but they should have.
In this post I suggest some studies I’d like to be done. I begin by recalling Moretti’s account of genre succession from Maps, Graphs, Trees in the context of Jockers’ massive graph of literary influence. Then I revisit the “Style” chapter and look at some of the work I passed over when I first posted on that chapter, the work related to Moretti’s generational observation. I then make some suggestions about how we could infer quasi-genres in the data assembled to build the influence graph and thereby extend Jockers’ work on style from his limited corpus of 106 texts to the larger corpus of 3346 texts. I conclude with some vague and tentative remarks about the pattern of reader interest betrayed in the record we’ve been examining, that of book publication.

Influence and Genre Succession

I’ve been thinking a lot about two things: 1) Moretti’s argument in Graphs, Maps, Trees that genres tend to cluster into 30 year cycles, and 2) Jockers’ massive graph in which all 3346 texts in his corpus are linked by relations of similarity, producing a graph that looks like this (which is Figure 9.3, p. 165; color version from the web):

9dot3

As Jockers points out, what’s remarkable about this graph is that the nodes are ordered in time from left (oldest) to right, but there is no temporal information in the data from which it was derived: “Books are being pulled together (and pushed apart) based on the similarity of their computed stylistic and thematic distances from each other” (p. 164).

That temporal ordering is a side effect of ordering by thematic and stylistic similarity. But, in the abstract, it could have been otherwise, no? Why should positioning texts near similar texts result in temporal ordering? (Would the same thing be true of 20th Century texts?) This ordering implies that the evolution of literary culture IS directional, but Jockers himself hasn’t posited any telos, nor do I see any need to do so. That directionality stems from the internal dynamics of the system. Authors, and I assume audiences as well, want to stick with what they know, and what they know was published in the previous years.

It seemed to me that Moretti’s cycles must somehow be in that graph, for all the texts in a given cycle are close together in time, by definition, as well as similarity. Alas, the whole corpus has not been coded for genre (p. 158). Is there some way we can back into genre since we’ve got this massive graph based on similarity relations among texts along 578 dimensions? Aren’t texts within the same genre more likely to resemble one another than texts in different genres?

The other thing on my mind is the fact that what really interests me is what’s on people’s minds and how that evolves over time. Some books will attract few readers, some books many readers; but the mere fact that a book has been published doesn’t speak to that. Moreover, books can be read long after they’ve been published. In the case of Moby Dick, it would seem that, for the most part it was read only long after it was published. Publication history is, at best, an indirect proxy measure of that.

And yet that history IS a history. Assuming that publishers are for the most part rational economic actors who want to turn a profit, their decisions on what to publish must take into account their sense of what people are reading and therefore what they’re buying. And the kinds of books that got published changed from one decade to the next. That record of  changes must reflect changes of reading taste.

Tuesday, December 12, 2017

From distant reading to computational criticism: Canon/Archive @ 3QD [#DH]

n_1_moretti-shopify_1024x1024

As far as I can tell literary studies will remain committed to pre-computational intellectual formations for the foreseeable future and will do so from a position of quasi-aristocratic superiority over crass calculation, of which it will remain fitfully ignorant.

One might say that my participation in the academic blogosphere is framed, for the moment, by engagement with the work of Franco Moretti. It began with a so-called book event at The Valve, a now dormant group blog I joined late in 2005. Jonathan Goodwin had organized an online symposium about Moretti’s Graphs, Maps, Trees. It went live on the web on January 2, 2006 and had contributions by 15 people. Eight of us were Valve contributors: John Holbo, Ray Davis, Matt Greenfield, Amardeep Singh, Adam Roberts, Bill Benzon, Jonathan Goodwin, and Sean McCann. There were guest appearances by seven: Franco Moretti, Matthew Kirschenbaum, Timothy Burke, Eric Hayot, Steven Berlin Johnson, Jenny Davidson, and Cosma Shalizi.

I have no idea how many people contributed comments to those discussions, which unfolded over a period of three weeks. Most, but not all of these authors held academic posts (I did not). Two, I believe, did not have doctorates. The contributors came from all over the place and I wouldn’t hazard a guess about their backgrounds. Some were academics, I’m sure, and some where not. The symposium eventually became (self-published) book: Jonathan Goodwin and John Holbo, eds. <i>Reading Graphs, Maps, and Trees: Responses to Franco Moretti (Glassbead Books 2009).

That was over a decade ago. It was a collective event that took place outside the ordinary confines of the academic world – my use of confines is, of course, quite deliberate.

At the time Moretti was talking about distant reading. He had not yet become involved with computing and, necessarily, with people who know how to program then. He formed the Stanford Literary Lab in 2010 with Matthew Jockers, who had computer skills that he did not. The Lab issued its first pamphlet in, I believe, January of 2011: Quantitative Formalism: an Experiment. It was signed by Moretti and four others: Sarah Allison, Ryan Heuser, Matthew Jockers, and Michael Witmore. The most recent pamphlet, no. 16, was issued in November of this year: Totentanz. Operationalizing Aby Warburg’s Pathosformeln, by Leonardo Impett and Franco Moretti.

That first pamphlet had been submitted to a prestigious journal but was turned down in terms that suggested that the intellectual method itself was being rejected, not just that particular argument. And so the group decided to sidestep the world of formal academic publication and publish its work in the form of stand-alone pamphlets. This is common enough in technical disciplines, where work will be published in the form of technical reports – though that work will sometimes/often then by published in journal form as well; but it was unheard of in the humanities.

My point, then, is that the literary lab has worked at the edge of the academic world. It is located at a prestigious university, and its pamphlets bear the imprimatur of that university. But they do not exist in the world of formal academic publication.

Moretti, who is now retired from Stanford, and the lab have now gathered eleven of those pamphlets into a book:
Canon/Archive: Studies in Quantitative Formalism by Franco Moretti (Author, Editor), Mark Algee-Hewitt, Sarah Allison, Marissa Gemma, Ryan Heuser, Matthew Jockers, Holst Katsma, Long Le-Khac, Dominique Pestre, Erik Steiner, Amir Tevel, Hannah Walser, Michael Witmore, and Irena Yamboliev, published by n+1.
Note the term, quantitative formalism. Moretti now refers to this approach as computational criticism rather than distant reading, though some in the field use the latter term.

I have written an essay/review about the book and published it in 3 Quarks Daily, December 11, 2017, which is not an academic journal: Moretti and the Stanford Literary Lab: Computational criticism in two senses and the prospect of a new approach to literary studies.

The epigraph to this post is the final sentence of that essay/review.

Will the future of literary studies unfold outside departments of literature and their associated societies and journals?

More later.

Wednesday, November 29, 2017

Rejected @NLH! Part 4: Déjà vu all over again at New Literary History + Welcome to the club, Franco! [#DH #Canon/Archive]

I've been reading the new Canon/Archive by Franco Moretti and 13 others and decided to bump this post to the top of the queue for reasons that will become obvious when you read addition I've made to the head of the post.

EXTRA! EXTRA! READ ALL ABOUT IT! LATE ADDITION!

In the first paragraph of his preface to Canon/Archive Moretti tells us how the Literary Lab decided to publish its own pamphlets:
A well-known scholarly journal had been asking for an article on new critical approaches, and that’s where we sent the piece once it was finished. But it came back with so many requests for corrections that it felt like a straightforward rejection. It was dismaying; a few years ago, computational criticism was still shunned by the academic world, and we couldn’t help thinking that what was being turned down was not just an article, but a whole critical perspective. And since we also thought that the essay was perfectly fine as it was, we decided that—instead of trying our fortune with another journal (or, god forbid, making the required alterations)—we would publish it on our own, as a document of the Literary Lab. I cannot remember how the term “pamphlet” came up; and, frankly, it wasn’t even the right one: pamphlets have a public vocation that our work, with its heavily technical aspects, couldn’t possibly have. But the word captured the euphoria of being on our own; the freedom to publish what we wanted, when and how we wanted: short, long, even very long, our pamphlets never come out a minute earlier than they’re ready, nor a minute later, either; and without going through the grinder of editing “styles.” And all this, because “Quantitative Formalism” was rejected by—Never mind. They did us a favor.
YES! To all of it, been there, done that. The following post tells how I was rejected at MLN early in my career in 1980 and at NLH just last year. It's pretty clear in both cases that "what was being turned down was not just an article, but a whole critical perspective." Moreover, it is computation that was being rejected in those cases as well as Moretti/LitLab's. I had framed the MLN submission as being structuralist, which it was in a way, but it was computational at its heart. The NLH submission was explicitly computational. Silly me, I thought the critical world was changing.

As for the term "pamphlet", it's fine. But hundreds if not thousands of academic, government, and industrial labs issue things called "technical reports", some of which become articles in the formal literature, and many do not. I've read 100s if not thousands of these tech reports, and have even written two (back at the Center for Manufacturing Productivity and the Rensselaer Polytechnic Institute).

And I certainly understand "the freedom to publish what we wanted ... long, even very long". Some of my best work takes the form of pieces that are too long for journal publication but too short for monograph publication. The economics of hardcopy publication places restrictions on what can be published and therefore, if only indirectly, on what is thought. 

More later on Canon/Archive.

* * * * *

I’ve been discussing a manuscript of my that was rejected at New Literary History:
Sharing Experience: Computation, Form, and Meaning in the Work of Literature, https://www.academia.edu/28764246/Sharing_Experience_Computation_Form_and_Meaning_in_the_Work_of_Literature
In previous posts I’ve laid some groundwork, first discussing why I decided to submit to NLH, then positioning the article within my larger intellectual project, and, most recently, recounting the history of literary criticism in the 1970s as it moved from openness to closure.

That brings us to 1980, when I decided to submit an essay about “Kubla Khan” to MLN. It was turned down on the basis of a deeply conflicted set of reviewer’s comments. Some of those comments are resonant with comments made by the reviewer who rejected the current essay for NLH. It’s that resemblance that, in part, prompted me to once more re-examine the 1970s and to write this series of posts.

In this post I begin by telling the story of being rejected at MLN. Then I discuss my rejection at NLH, in two parts. In the first part I discuss the similarities between the two rejections. In the second part I suggest that the NLH reviewer is skeptical about computing for reasons that seem more ideological than the result of well-informed study.

“Kubla Khan” – Rejected at MLN

In 1972 I filed a master’s thesis with the Humanities Center at Johns Hopkins. I forget the exact title, but it was a more or less structuralist analysis of “Kubla Khan,” the work that prompted me to go all-in on the emerging cognitive sciences (though the term, “cognitive science”, wasn’t coined until 1973). It wasn’t until 1980 that I decided to publish that work. I deleted a lot of the philosophical discussion, added some new diagrams of a style owing more to cognitive science than structuralism, and sent it out under the title “Articulate Vision: A Structuralist Reading of ‘Kubla Khan’.” By that time I’d ceased thinking of myself as a structuralist, after all I written a 1978 dissertation entitled “Cognitive Science and Literary Theory”, but I presented the paper that way because I figured that a literary audience would at least recognize structuralism.

But where should I submit it? No one was publishing essays like that.

I decided to submit to the comparative literature issue of MLN. The basic reason was simple; Richard Macksey edited that issue and he’s the one who directed that master’s thesis. Moreover the comparative literature issue publishes theoretical pieces, which this more or less was. And, of course, MLN had published my first cognitive networks piece in the special Centennial Issue, “Cognitive Networks and Literary Semantics” (MLN 91: 952-982, 1976).

So I submitted the piece to MLN. Macksey had to turn it down because the reviewer’s report was unfavorable. The reader noted that “I found myself teetering on the edge of Kubla’s girdling wall, uncertain whether to tip one way and fall into Benzon’s enchanted ground, or the other way and run from his tables and charts”. Note the reviewer’s alarm at the diagrams [1], which were somewhat more complicated than the one’s that Mark Rose had apologized for in Shakespearean Design back in 1972 [2]. The reviewer goes on to register “surprise at encountering a straightforward, unembarrassed structuralist analysis” in the deconstructive era.

Yet the reviewer acknowledges that those same charts “have a real value in coming to terms with the text, and I will no doubt refer to them when I teach the poem.” That strikes me as a very strong positive remark, a clear statement of his approval. After all, you don’t – at least I didn’t – ordinarily base your teaching on far-out crazy ideas; you are conservative in what you present to students. The reviewer went on, however, to complain that the essay “ought to argue with itself, to put into question some of the patterns it establishes – or better, perhaps to let the poem talk back.” And after this that and the other, they [yes, I know, but I prefer that usage to the more awkward “his or her”] flatly recommend against publication, no chance for revision.

It was a strange and conflicted review. The analysis seems to have made sense to the reviewer but did so in terms so at odds with their sense of the proper (deconstructive) way to approach a poem that they were in the grip of cognitive dissonance. It shouldn’t have made sense at all. But it did, gosh darn it! What to do? The easiest way to resolve that dissonance was simply to wish my article out of existence, that is, to reject it. Whatever Macksey himself may have thought about the article, he had little or no choice but to follow the reviewer’s advice and reject it.

And you know, come to think of it, since I knew Macksey personally, I called him up and we discussed the rejection. I don’t recall the discussion in any detail, though I remember that his wife, Catherine picked up the phone, but it was amiable. Macksey acknowledged the review was strange, but I didn’t push him on it. And that was that. But I’m not in a position to call Rita Felski, the editor of NLH.

Rejection at NLH

The reviewer at NLH didn’t express any such conflict or ambivalence. The rejection was firm and unequivocal. That’s quite clear. Beyond that, however, I’m a bit up in the air since I don’t know who the reviewer was and so have no sense of what they know. In particular, what do they know of computing, which is how I framed by article?

Wednesday, November 15, 2017

Ted Underwood on Canon/Archive [#DH]

Ted Underwood reviews Canon/Archive, by Franco Moretti (Author and Editor) and 13 others (Mark Algee-Hewitt, Sarah Allison, Marissa Gemma, Ryan Heuser, Matthew Jockers, Holst Katsma, Long Le-Khac, Dominique Pestre, Erik Steiner, Amir Tevel, Hannah Walser, Michael Witmore, and Irena Yamboliev, all authors). How, you might ask, could that be? Simple, says Underwood, more or less. Moretti worked at Stanford's Literary Lab (he's now retired from Stanford) and laboratories are (typically) sites of collaborative and collective work. People at different stages in the intellectual life, with different sets of conceptual, craft, and rhetorical skills, work together in a common intellectual enterprise. I participated in such a group when I was doing by PhD work at Buffalo (English), where I was a member of David Hays's linguistics/cognitive science group (Linguistics). It can be a good way to work, but it is not how work in the humanities has ordinarily–classically, if you will–been done.

And, despite the sense one gets in the general media that quantitative work is new to the humanities, that's not the case, says Underwood. It's been around for several decades.
But over the last 30 years, the enclaves have joined to produce a practice of quantitative interpretation that is no longer purely sociological, or purely linguistic, but able to range freely across the spectrum from single words to social trends. Computers are certainly useful in this mode of interpretation, but they aren’t the new element.
What's new is a certain "human connection between scholars." Initially Moretti and Matthew Jockers, but others joined the party, and this volume presents their collective work.

And so:
Experiment is presented here not just as a test of reliable knowledge but as a style of intellectual growth: “By frustrating our expectations, failed experiments ‘estrange’ our natural habits of thought, offering a chance to transcend them.” At moments, the point of experiment seems to become entirely aesthetic. In the book’s introduction, Moretti admits that he set out to write “a scientific essay, composed like a Mahler symphony: discordant registers that barely manage to coexist; a forward movement endlessly diverted; the easiest of melodies, followed by leaps into the unknown.”
I wonder, did Moretti get the musical trope from Lévi-Strauss, who uses it to structure The Raw and the Cooked, with section, chapter, and subchapter titles all derived from music (e.g. "The Fugue of the Five Senses", "The Oppossum's Cantata").

Underwood continues:
The essays within are unified by a deliberately wandering structure, which keeps its distance both from scientists’ predictable sequences (methods → results → conclusions), and from the thesis-driven template that prevails in the humanities (counter-intuitive claim → evidence → I was right after all). Instead, these essays become stories of progressive disorientation, written in the first-person plural, and arriving at theses that were only dimly foreshadowed.

This narrative form has given the Literary Lab a coherent authorial persona, which may lead readers to assume that the experiments gathered here are also unified by shared methods and theories, more or less identified with Moretti. That would be a mistake. The Literary Lab is genuinely a collective project, and these essays have been shaped by many different approaches to the literary past.
Of numbers and meaning?
The most important fracture in the book involves the nature of the connection between numbers and interpretation. Several of the essays make this connection with statistical models. In chapter 1, for instance, Michael Witmore reveals that a model of textual similarity based purely on word frequencies can group Shakespeare’s plays into comedies, histories, and tragedies. But when Moretti describes the book’s methods in the conclusion, he downplays the interpretive connection provided by models, in order to tell a story that leaps from observation of opaque “patterns” to the “discovery of a causal mechanism.” This account again reveals Moretti’s commitment to framing the work of the Lab as a humanistic narrative. Scientists don’t usually understand their methods as a process of pure induction that produces meaning at the last moment; instead, they tend to begin with a hypothesis, and find a way to test it (often using a model).
Except that "humanistic narrative" tends to be short on causal mechanism.

But who cares? Will it spread? Underwood notes that New Historicist criticism spread rapidly because it didn't require new skills, just a somewhat different deployment of standard lit crit skills, "surprising connections between works of literature and historical events" plus anecdotes. Computational criticism of the sort done at the Literary Lab, however, is a different kind of beast. The work requires people with programming skills, statistical skills, and a nose for experimental design in addition to a feel for literary phenomena. No one individual has to have all these skills, but all skills must be available in the group, and all members of the group need to be able to talk with one another.

Thus, Underwood assures his (humanistic) reader, "there is no danger that a quantitative approach to literature will spread like wildfire". Whew!  However, "there will certainly be more books like this one." Yes, there will.

And even stranger ones. I've been exploring this territory for some time now, since the mid-1970s. Oh, not the particular regions Moretti and Underwood (and others) have been exploring so fruitfully for the last decade or two. That's new to me, as it is to the discipline. But that's not all there is here, wherever THIS is, not by a long shot. I've got to tell you, Guys, this isn't the East Indies, it's not even Kansas. It's something else, maybe not even terrestrial.

Stay tuned.

Sunday, November 5, 2017

Tracking “Xanadu” around the web

IMGP6920rd

I'm gearing up to write a review of Franco Moretti's latest; well, actually, not just Moretti, but Moretti plus thirteen others: Mark Algee-Hewitt, Sarah Allison, Marissa Gemma, Ryan Heuser, Matthew Jockers, Holst Katsma, Long Le-Khac, Dominique Pestre, Erik Steiner, Amir Tevel, Hannah Walser, Michael Witmore, Irena Yamboliev (notice the alphabetic ordering). N+1 is putting out a collection of papers from the Stanford Literary Lab, Canon/Archive: Studies in Quantitative Formalism, and I'll be reviewing it for 3 Quarks Daily.

I first learned about Moretti just over a decade ago when I was writing for The Valve. We had a "book event" organized around his Graphs, Maps, Trees (2007). A number of people contributed posts and Moretti responded. I contributed two posts, a short one, and a somewhat longer one, One Candle, a Thousand Points of Light: Moretti and the Individual Text. The format's gone bizarro and there's a lot of link rot, but the discussion was long and interested. What I did was google the term "xanadu" and investigated the fascinating results. Who would have thought that such an exotic word would get two to three million hits? That was 2006; now it gets almost 10 million hits.

Anyhow, sometime later I took that old post, refined it a bit, and published it as a working paper: One Candle, a Thousand Points of Light: The Xanadu Meme. You can download it here:
It’s rather different in spirit and technique from the text mining/corpus linguistics techniques that are currently raising the bar in computational humanities. But it is computationally intensive. It’s just that the intensity is not on my desktop, it’s in Google’s servers. I used Google to troll the web for “Xanadu” and identified two contexts, other than the Coleridge’s poem, where I got a lot of hits.

Here’s the abstract:
I treat a single word 'xanadu', as a 'meme' and follow it from a 17th century book, to a 19th century poem (Coleridge's "Kubla Khan"), into the 20th century where it was picked up by a classic movie ("Citizen Kane"), an ongoing software development project (Ted Nelson's Project Xanadu), and another movie and hit song, Olivia Newton-John's Xanadu. The aggregate result can be seen when you google the word, you get 6 million hits. What is interesting about those hits is that, while some of them are directly related to Coleridge's poem, more seem to be related to Nelson's software project, Olivia Newton-John's film and song, and (indirectly) to Welles' movie. Thus one cluster of Xanadu sites is high tech while another is about luxury and excess (and then there's the Manchester Swingers Club Xanadu).
We can call the first cluster of sites the cybernetic context while the second is the sybaritic context. This simple diagram shows the evolution:

Xanadu Cladogram

The Citizen Kane line is, of course, the sybaritic one, while the Ted Nelson line is the cybernetic one.

My point is simple and, I hope, obvious. That this one term, “Xanadu,” has taken on two different valences, which are active in two different contexts. The cybernetic doesn’t seem to exist in Coleridge’s text at all. Ted Nelson created that when he provided a new context for the poem by making it the inspiration for his hypertext system. Though it wasn’t the poem that was the context as much as it was Coleridge’s alleged inability to remember the whole thing. Nelson imagined his hypertext system as one where nothing would ever get lost.

The sybaritic context, however, is there in the poem—“pleasure dome,” “demon lover,” “damsel with a dulcimer” and so on. Welles’ film, in effect, stripped the rest of the poem away and then magnified the cultural presence sybaritic element by putting it on the motion picture screen. If we examine the historical record more closely, though, we have hints that that job had started before Welles made his film. Prior to 1940 The New York Times, for example, makes mentions a yacht, or perhaps two, named “Xanadu.”

Now, imagine contexts for thousands of words in thousands of documents. That’s what is examined in topic analysis. Documents provide contexts for topics, and topics provide contexts for works. Algorithms for topic analysis examine words in large collections of documents and infer the topics in which they belong.

* * * * *

Here's the table of contents:
Introduction: “Xanadu”— 2
Googling for Memes — 3
A Thousand Points of Light, a Metaphor — 4
Xanadu: A View from the Wikipedia — 5
Xanadu: A Google View. — 7
Examining the Xanadu System — 12
Beyond the Meme — 18
Beyond Interpretation — 19
Appendix 1: Googling Oedipus — 19
Appendix 2: Xanadu in Google Books — 21
References — 23

Monday, May 9, 2016

An interesting response to Moretti on digital humanities

On Saturday I ran up a post on Franco Moretti’s assertion that the results of computational criticism have “so far been below expectations”, as he remarked in his LARB interview, conducted by Melissa Dinsman. In response Ted Underwood assured me
I'm more sanguine than I have ever been about the intellectual and historical payoffs for computational criticism. What I've seen in recent papers, and in forthcoming ones, makes me very confident that we are going to engage, challenge, and in some cases frankly refute important existing theses about literary history. Big Ideas and broad theories of the nature of literary change won't be scarce.
I’ve found another response in an interview with Richard Jean So, also by Dinsman in LARB. He’s an “assistant professor in modern and contemporary American culture at the University of Chicago” and has done “computational work on race, language, and power dynamics.” In her last question Disman references Moretti’s reservations and asks So to “look backward and speak to what you think the digital in the humanities has accomplished so far.” Here’s how he responds:
As a younger person in the field, I don’t like the gesture of looking back. I find it problematic for someone to justify a field by saying to people, “Look how great my field is because of all these accomplishments.” If you are in a field and the accomplishments don’t speak for themselves, then you have more work to do. It is a position of potential weakness or insecurity to constantly say, “Look what we’ve done.” There is certainly a place for that, but in trying to build out the field, this looking back can be problematic when done in a defensive way. If our work and accomplishments are not instantly recognizable outside the field, then we have to do more work. So I would definitely say that I am more future minded. It isn’t obvious yet that DH is here to stay, so rather than meet critiques with a “look what we’ve done” mentality, we need to go back to our books and computers and do better work until we don’t have to answer this question any more.

Saturday, May 7, 2016

What’s Interesting? Is Moretti Getting Bored?

“The interesting” or “interestingness” came up in the Twittersphere in a response Ted Underwood made to Ryan Heuser:


This is something I think about from time to time – I blogged about it back in 2014: The Thinkable and the Interesting: Katherine Hayles Interviews Alan Liu – generally in connection with my own work on description and, in particular, the description of formal features of texts and films, such as ring-composition. I find this activity to be quite interesting, intrinsically interesting if you will. But, judging by what literary critics actually do, most critics aren’t particularly interested in such things.

Why not?

I don’t know. I assume, though, that I’ve got some unseen conceptual context to which I assimilate such description and in terms of which it is interesting. Whatever that context is, most literary critics don’t have it.

But enough about me in my possibly peculiar interests.

I want to think about Franco Moretti. In both a recent interview in the Los Angeles Review of Books and his most recent pamphlet, Literature, Measured (which is now the preface to Canon/Archive, 2107, pp. ix-xvii), Moretti has expressed misgivings about the current state of affairs in computational criticism. He ends the pamphlet thus (p. 7):
“Bourdieu” stands for a literary study that is empirical and sociological at once. Which, of course, is obvious. But he also stands for something less obvious, and rather perplexing: the near-absence from digital humanities, and from our own work as well, of that other sociological approach that is Marxist criticism (Raymond Williams, in “A Quantitative Literary History”, being the lone exception). This disjunction […] is puzzling, considering the vast social horizon which digital archives could open to historical materialism, and the critical depth which the latter could inject into the “programming imagination”. It’s a strange state of affairs; and it’s not clear what, if anything, may eventually change it. For now, let’s just acknowledge that this is how things stand; and that – for the present writer – something needs to be done. It would be nice if, one day, big data could lead us back to big questions. It would be nice if, one day, big data could lead us back to big questions.
What I’m wondering is if “big questions” is a marker for missing intellectual context. Is Moretti unable to see that these investigations will lead, even must eventually lead, to home truths? It’s not that I can see where things are going – I can’t – but, for whatever reason, I don’t share his feeling that things might be going amiss.

The curious thing is that, even as he expresses these misgivings in both these texts, he also expresses a certain fondness, even nostalgia, for traditional critical discourse. Thus, in Literature, Measured he says (p. 5):
Forget the hype about computation making everything faster. Yes, data are gathered and analyzed with amazing speed; but the explanation of those results – unless you’re happy with the first commonplace that crosses your mind – is a different story; here, only patience will do. For rapidity, nothing beats traditional interpretation: Verne’s “Nautilus” means – childhood; Count Dracula – monopoly capital. One second, and everything changes. In the lab, it takes months of work.
And maybe the explanation, if and when it comes at all, seems a bit weak. This, from the LARB interview:
But again, think of this: to make it better — it's a perfect expression because it's a comparative, it was good and now it's more good — this is not how the humanities think in general. It's usually much more of a polemical, an all-or-nothing affair. It's a conflict of interpretation. It's: you thought Hamlet was the protagonist of Hamlet, how foolish of you; the protagonist is Osric. Digital humanities doesn't work in this mode and I think there is something very adult and very sober in not working in this mode. There is also something, maybe especially for older people like me, which is always a little disappointing: the digital humanities lacks that free song — the bubbliness of the best example of the old humanities.
The thing is, that traditional critical discourse, the one in which the critic can construct a bubbly free song, is embedded in a conceptual matrix that is linked to some version of Home Truth, whether Christian humanism, Hegelian phenomenology, Marxist theory, or critique in its many flavors, whatever. That matrix gives the critic intellectual purchase on the Whole of Life.

That’s the context of the traditional humanities. That’s where the traditional humanities draw their interest: Life, the Universe, and All.

By contrast, computational criticism seems to have cut the cord. Can’t get there from here. The big questions seem impossibly distant, if not actually gone.

I wonder what sustained all those pre-Darwinian naturalists, the ones who went out in the field and were both content and even eager to describe the flora and fauna they found? For that matter, would Darwin have spent so much time describing barnacles if he didn’t find the activity itself interesting apart from whatever it might contribute to his larger intellectual enterprise? Could he even have found that larger enterprise if he hadn’t been fascinated by the details of the morphology, physiology, and lifeways of flora and fauna?

Tuesday, April 26, 2016

From Telling to Showing, by the Numbers

I've been thinking about some remarks Moretti made about the digital humanities in  a recent interview. Among other things he suggested that the results of computational criticism have so far been disappointing. But he also held up Lit Lab Pamphlet #4 as an example of "an intelligence that takes the form of writing a script, but in the writing of the script there is also the beginning of a concept, very often not expressed as a concept, but that you can see that it was there from the results that the coding produces." Here's what I wrote about that pamphlet back in October of 2012.
I’ve just looked at a pamphlet from Stanford’s Literary Lab: Ryan Heuser and Long Le-Khac, A Quantitative Literary History Of 2,958 Nineteenth-Century British Novels: The Semantic Cohort Method (68 page PDF), May 2012. I’ve not read it in detail, but only blitzed my way through, looking for the good parts. Well, not even all of those. I was just looking to get a sense of what’s going on.

Which I did. And I like it. THIS is the sort of work I want to see from ‘digital humanities.’ Not the only sort, but it’s one of the things we can do with ‘big data’ and pretty much only do with big data. If traditional humanists can’t see value in this kind of work, well, then forget about them.

First I’ll give you the abstract, then I’ll quote a bunch and make some comments.

Authors’ abstract
The nineteenth century in Britain saw tumultuous changes that reshaped the fabric of society and altered the course of modernization. It also saw the rise of the novel to the height of its cultural power as the most important literary form of the period. This paper reports on a long-term experiment in tracing such macroscopic changes in the novel during this crucial period. Specifically, we present findings on two interrelated transformations in novelistic language that reveal a systemic concretization in language and fundamental change in the social spaces of the novel. We show how these shifts have consequences for setting, characterization, and narration as well as implications for the responsiveness of the novel to the dramatic changes in British society.

This paper has a second strand as well. This project was simultaneously an experiment in developing quantitative and computational methods for tracing changes in literary language. We wanted to see how far quantifiable features such as word usage could be pushed toward the investigation of literary history. Could we leverage quantitative methods in ways that respect the nuance and complexity we value in the humanities? To this end, we present a second set of results, the techniques and methodological lessons gained in the course of designing and running this project.

Tuesday, March 8, 2016

Franco Moretti on digital humanities, again...

Melissa Dinsman interviews Franco Moretti in the LA Review of Books. On the humanities in the 20th century:
In the 20th century the natural sciences have produced some amazingly stunning and beautiful theories in physics, and genetics, and in biology. The humanities have produced nothing of this sort. Literature, art, in a sense even political history (mostly in a horrendous way), have produced enormously interesting objects, but the study of these objects, that is to say the disciplines of the humanities — the study of literature, the study of history — have lagged behind. The humanities have lagged behind in conceptual imagination and in boldness.
Whoops! But digital humanities isn't coming to the rescue:
No, to make the humanities relevant you need something much bigger than the digital humanities. What the humanities need are large theories and bold concepts.
Yep. In the value of coding:
It's an intelligence that takes the form of writing a script, but in the writing of the script there is also the beginning of a concept, very often not expressed as a concept, but that you can see that it was there from the results that the coding produces. Perhaps the best example in the case of the Literary Lab was Pamphlet #4,* which was written by two grad students who invented their own script. I envy them that form of intelligence, knowing that I will never have it. And I like it. I think that actually many of the most promising results in the future will come from scripts that are half scripts / half cultural, literary, historical concept.
Yes, of course. I would only add, and now I'm mounting one of my favorite hobby horses, is that literary critics must learn to think about mental and cultural processes as involving computation in some form.

On judgment vs. explanation:
You read reviews that tell you if a book or a film is good or bad. And the same for art shows, and so on. Digital humanities is as non-normative as one can get in the field of literature. It is much more towards the explanatory. So to make it interesting for the general public, a major revolution in the way in which literature is approached by the media would be necessary. Will this revolution happen? No. Should this revolution happen? I'm not even sure. I have devoted my life to explanation rather than value judgment. On the other hand, I am not sure that for society-at-large, for the world-at-large, explanation is more important than value judgment. I think it is more important for people who devote their lives to try to understand how things work.
On this I'll say, unequivocally, that explanation of cultural phenomena (works of art, music, literature, dance, and so forth) is certainly important for society-at-large, for the world-at-large.

Moretti poses his own question for DH: "Leave aside what it can do in the future; has it done anything?" What do you think his answer is:
...the results so far have been below expectations. Now, it's true that the field is at the beginning still. It's true that much of scientific research is so called normal science, and it's certainly true that traditional literary criticism is not sending off sparks every day. All of this is true, but it is also irrelevant because digital humanities are claiming to be the big novelty and so far I think I have produced little evidence about that. I don't want to push it too far. I don't want to say there is not evidence, because it is complicated. Evidence comes in many forms.
I'll go along with those last three sentences.

Tuesday, August 12, 2014

Reading Macroanalysis 3.1: Style, or Measuring the Autonomous Aesthetic Realm

Yesterday I posted on Chapter 6, “Style,” in which Jockers argued, in effect, that insofar as we can measure (or estimate) the factors that affect a text’s style, authorial identity is the strongest of those factors. The key is that phrase, “insofar as we can measure,” because that’s the intellectual world in which we are now functioning.

I now want to take Jockers’ arguments on that score and refit them for use as evidence that an autonomous aesthetic realm does indeed exist, as the late Edward Said believed but couldn’t quite explain.

First I want to take up the topic that was the subject of Moretti’s most recent pamphlet, operationalization. Then I’ll introduce Said’s conundrum about the existence of an autonomous aesthetic realm and discuss how we could operationalize it. I’ll conclude by arguing that Jockers has already, in effect, all but given us an operationalization of it.

Operationalization

If we can’t measure it, or operationalize it, to use a term Moretti adopted from physics (“Operationalizing”: or, the Function of Measurement in Modern Literary Theory, Literary Lab, Pamphlet, December 2013) then we can’t reason about it in this universe of discourse. In some other universe of discourse, sure, but not in this one.

Moretti glosses “operationalize” by a passage from P.W. Bridgeman (p. 2):
We may illustrate [the meaning of the term] by considering the concept of length: what do we mean by the length of an object? [...] To find the length of an object we have to perform certain physical operations. The concept of length is therefore fixed when the operations by which length is fixed are fixed: that is, the concept of length involves as much and nothing more than the set of operations by which length is determined. In general, we mean by any concept nothing more than a set of operations; the concept is synonymous with the corresponding set of operations [...] the proper definition of a concept is not in terms of its properties but in terms of actual operations.
If I might push the term a bit, since the middle of the last century, and a bit before, literary criticism has ‘operationalized’ the concept of meaning by the procedure of so-called close reading. That is to say, the meaning of a literary text, whether a sonnet by Shakespeare or a narrative by Murasaki Shikibu, is what the procedure of close reading determines it to be.

Alas!, or if you are so inclined, mirable dictu!, the result of this process tends to vary from one critic to another. In science, that would be a problem. In literary criticism it is merely a provocation to theory.

Many critics simply ignore it, perhaps covering it over with the anodyne topos that the multiplicity of meanings simply shows the richness of the text. Other critics assume that the concept has been inadequately operationalized and go in search of more adequate methods – the literary Darwinists, led by Joseph Carroll, are the most recent such school. And still other critics accept this as evidence that meaning is indeterminate.

But I digress. It’s not meaning we’re after. It’s style. Traditionally the concept of style has been operationalized – and here I’m again pushing things a bit – by describing texts in rhetorical, philological and, more recently, linguistic terms. Since the middle of the previous century computational humanists have operationalized the concept by counting textual features and undertaking a statistical analysis of the counts.

From the standpoint of traditional humanism that seems odd and terribly impoverished, and no doubt it is. But it is also reliable from one researcher to another and has allowed stylisticians to accomplish at least one task beyond the reach of traditional humanists, with their richer methodology. Namely, identifying the authors of otherwise anonymous texts. This is the tradition in which Jockers is working.

Wednesday, July 23, 2014

Beyond Quantification: Digital Criticism and the Search for Patterns

I've collected my recent posts on patterns into a working paper. It's online at SSRN. Here's the abstract and the introduction.
Abstract: Literary critics seek patterns, whether patterns in individual texts or patterns in large collections of texts. Valid patterns are taken as indices of causal mechanisms of one sort or another. Most abstractly, a pattern emerges or is enacted as some machine makes its way in an environment. An ecological niche is a pattern “traced” by an organism in its environment. Literary texts are themselves patterns traced by writers (and readers) through their life worlds. Patterns are frequently described through visualizations. The concept of pattern thus dissolves the apparent conflict between quantification and meaning, for quantification is but a means to describing a pattern. It is up to the critic to determine whether or not a pattern is meaningful by identifying the mechanism that produced the pattern. Examples from Shakespeare and Joseph Conrad.
Introduction: Patterns and Descriptions

There is a sense, of course, in which I’ve been aware of and have been perceiving and thinking about patterns all my life. They are ubiquitous after all. But it wasn’t until I began studying cognitive science with the late David Hays that “pattern” became a term of art. Hays and his students were developing a network model of cognitive structure – such works became common in the 1970s. Such networks admit of two general kinds of computational process, path tracing and pattern recognition. Path tracing is computationally easy, while the pattern recognition is not. Human beings, however, are very good at perceiving and recognizing patterns.

What put the idea before me, though, as something demanding specific thought, are remarks Franco Moretti made in coming to grips with his work on the network analysis of plot structure. In Network Theory, Plot Analysis (Literary Lab Pamphlet 2, 2011, p. 11) he noted that he “did not need network theory; but I probably needed networks.... What I took from network theory were less concepts than visualization.” We then examine the visualizations to determine whether or not they indicate patterns that are worth further exploration.

That, it seems to me, should put to rest fears about the incommensurability of numbers and meaning or, even worse, anxiety about infecting humanistic inquiry with quantitative evil. It’s not about numbers and counting. It’s about patterns. Numerical work is subordinate to and in service of looking for patterns, whether patterns in individual texts, as Moretti was doing in his work on plot structures, or patterns in collections of hundreds and thousands of texts spanning decades or more of historical time.

But, just what IS a pattern anyhow? How do we tell the difference between patterns and, well, non-patterns? Those are tricky questions, questions I pursue in the posts that make up this working paper. If what we’re looking for is some a priori way of specifying what patterns are so that we can then theorize about patterns in a general way, then I think we’re in trouble. In the sections, “Pattern” as a Term of Art and Patterns as Epistemological Objects, I suggest that there is no such thing. What emerges from those discussions is something like this: A pattern is something that emerges or is enacted as some machine makes its way in an environment in which it either survives or fails – where the italicized terms are understood in a very general and abstract sense. Thus understood, patterns are relations between machines and environments.

Saturday, June 7, 2014

Rens Bod on Patterns

WHO'S AFRAID OF PATTERNS?: THE PARTICULAR VERSUS THE UNIVERSAL AND THE MEANING OF HUMANITIES 3.0. BMGN - Low Countries Historical Review. Vol. 128, No. 4, 2013, 171-180.
Rens Bod

Abstract

The advent of Digital Humanities has enabled scholars to identify previously unknown patterns in the arts and letters; but the notion of pattern has also been subject to debate. In my response to the authors of this Forum, I argue that ‘pattern’ should not be confused with universal pattern. The term pattern itself is neutral with respect to being either particular or universal. Yet the testing and discovery of patterns – be they local or global – is greatly aided by digital tools. While such tools have been beneficial for the humanities, numerous scholars lack a sufficient grasp of the underlying assumptions and methods of these tools. I argue that in order to criticise and interpret the results of digital humanities properly, scholars must acquire a good working knowledge of the underlying tools and methods. Only then can digital humanities be fully integrated (humanities 3.0) with time-honoured (humanities 1.0) tools of hermeneutics and criticism.
What I'm wondering is whether or not pattern is emerging as a fundamental epistemological/ontological entity. I've broached this idea in an earlier post focussed on Moretti, From Quantification to Patterns in Digital Criticism, where I observed, of his network diagrams:
What are those diagrams about? Let me suggest that they are about patterns. Yes, I know, the word is absurdly general, but hear me out.

That bar chart depicts a pattern of quantitative relationships, and does so better and more usefully than a bunch of verbal statements. You look at it and see the pattern.

The pattern in the network diagram is harder to characterize. It’s a pattern of relationship among characters. What kind of relationships? Dramatic relationships? That, I admit, is weak. But if you read Moretti’s pamphlet, you’ll see what’s going on.

The important point is what happens when you get such diagrams based on a bunch of different texts. You can see, at a glance, that there are different patterns in different texts. While each such diagram represents the reduction of a text to a model, the patterns in themselves are irreducible. They are a primary object of description and analysis.

And that is my point: patterns.
As I said, the idea of patterns is very general. But that doesn’t make it useless. On the contrary, that generality makes the idea useful and powerful.

It is a commonplace in the cognitive science that the human mind (and brain) is very good at pattern recognition. But digital computers are not so good at it. I note also that the notion of design patterns has been popular in computer programming, which got it from the notion of pattern language articulated by the architect, Christopher Alexander.

Tuesday, May 27, 2014

Patterns: Ramsay on Shakespeare, and Beyond

I was looking though the syllabus for one of Alan Liu’s courses, Literature + (New Media & Literary Interpretation: Close, Distant, and Other Reading) and came across an older, and fascinating, paper by Stephen Ramsay, In Praise of Patterns, TEXT Technology, Number 2, 2005, pp. 177-190 (with accompanying figures). It’s an interesting piece of work, both for what it says about Shakespeare and for what it says about methodology.

Methodologically, Ramsay tells us he began playing around with graphs of Shakespeare plays because, well, he was interested in graphs and wanted to see what would show up. After a fair amount of work, including a presentation to some mathematicians and collaboration with some data miners, something very interesting showed up. Alas, Ramsay doesn’t know quite what to make of it, but it’s still interesting.

And that’s just fine. That’s the kind of world we’re in. We have the capacity to find interesting things, but once found, explaining them is sometimes/often a problem. If we can’t figure out how to operate in that world (this is me speaking now) we’re not going to get very far with digital humanities. Whatever digital humanities is, it is not a positivistic haven of certain knowledge (I’ve now returned to Ramsay).

This post is going to be a long one, over 2K words. First I present Ramsay’s work, then I present some work I did on Shakespeare some time ago, work that speaks to genre by looking at a comedy (Much Ado About Nothing), a tragedy (Othello), and a romance (The Winter’s Tale). I conclude by suggesting that we’ve left the world defined by existing methods of hermeneutic analysis and exegesis.

Patterns of Loci and Scenes

Ramsay looked at how Shakespeare’s plays moved from place to place as they moved from one scene to the next. Here’s the graph he produced for The Comedy of Errors:

plate_2

By standard mathematical convention such a diagram is called a graph; the ovals are called nodes and the arrows are called edges or arcs. Each node in one of Ramsay’s diagrams indicates a locus (my term) where a scene takes place and is labeled with a name or phrase designating that place. The edges indicate the transition from one scene to the following scene; the label on the edge indicates the following scene. It follows from the labeling convention that there will not be any edge labeled “1.1” as that is the first scene in any play and, as such, does not follow any other scene.

Tuesday, May 20, 2014

Negotiating the Manifold Nature of Literary Criticism

Once again, what is the (academic) discipline of literary criticism? It has, of course, changed a great deal over the last 100 years, though I don’t want to look quite that far back.

Let’s start with the present, with a comment by Franco Moretti in a recent interview with Laura Miller in Salon interview:
Literary history has two sides, I think. One is the normative side: deciding what is good and what is less good. The other is the explanatory side. It’s two very different modalities of thought, and I’ve always been inclined toward the explanatory. That’s what fires my mind. And in the study of the humanities, the normative modality has disappeared. It’s all explanatory now.
It’s that two-sidedness that’s so interesting to me. Moretti characterizes them as normative versus explanatory. There are other such characterizations. Whatever the sides, there, between them, that’s an interesting no-man’s land. I want to look at these matters through a hand-full of passages I’ve been pondering for some time.

Norms vs. Explanation

In the introduction to his 1957 Anatomy of Criticism Northrup Frye relegates the making of value judgments to the history of taste (p. 21):
Shakespeare, we say, was one of a group of English dramatists working around 1600, and also one of the great poets of the world. The first part of this is a statement of fact, the second a value-judgement so generally accepted as to pass for a statement of fact. But it is not a statement of fact. It remains a value judgement, and not a shred of systematic criticism can ever be attached to it.
The profession was happy to see in Frye’s statement a statement of its aims and desires, though of course, the journalistic reviewers of books and movies did not hesitate to make aesthetic and ethical judgements about the works they reviewed. And there are persistent calls to return value judgments to the fold.

Modes of Thought

I believe Moretti is right about norms and explanations requiring “two very different modalities of thought.” I even suspect brain-imaging studies would show different patterns of brain activity during these two modes.

The basic distinction, however, is between the reading of literary texts and the writing of criticism. Northrup Frye, again, makes this distinction in the introduction to his Anatomy of Criticism (pp. 27-28):
The reading of literature should, like prayer in the Gospels, step out of the talking world of criticism into the private and secret presence of literature. Otherwise the reading will not be a genuine literary experience, but a mere reflection of critical conventions, memories, and prejudices. The presence of incommunicable experienced in center of criticism will always keep criticism as art, as long as the critic recognized that criticism comes out of it but cannot be built on it.
But Frye would not have penned those words if the distinction between reading and criticism was not already so problematic that it had to be asserted at some length.

Monday, April 21, 2014

Toward a Computational Historicism. Part 1: Discourse and Conceptual Topology

Poets are the unacknowledged legislators of the world.
– Percy Bysshe Shelley

... it is precisely because we are talking about ordinary language that we need to adopt a notation as different from ordinary language as possible, to keep us from getting lost in confusion between the object of description and the means of description.
–Sydney Lamb


Worlds within worlds – that’s how Tim Perper, my friend and colleague, described biology. At the smallest scale we have individual molecules, with DNA being of prime importance. At the largest scale we have the earth as a whole, with all living beings interacting in a single ecosystem over billions of years. In between we have cells, tissues, and organs of various sizes, autonomous organisms, populations of organisms on various scales from the invisible to continent-spanning, and interactions among populations of organisms on various scales.

Literature too is like that, from single figures and tropes, even single words (think of Joyce’s portmanteaus) through complete works of various sizes, from haiku to oral epics, from short stories through multi-volume novels, onto whole bodies of literature circulating locally, regionally, across continents and between them, from weeks and years to centuries and millennia. Somehow we as humanists and literary critics must comprehend it all. Breathtaking, no?

In this essay I sketch a potential computational historicism operating at multiple scales, both in time and textual extent. In the first part I consider network models on three scale: 1) topic models at the macroscale, 2) Moretti’s plot networks at the mesoscale, and 3) cognitive networks, taken from computational linguistics, at the microscale. I give examples of each and conclude by sketching relationships among them. I open the second part by presenting an account of abstraction given by David Hays in the early 1970s; in this model abstract concepts are defined over stories. I then move on to Hauser and Le-Khac on 19th Century novels, Stephen Greenblatt on self and person, and consider several texts, Amleth, Hamlet, The Winter’s Tale, Wuthering Heights, and Heart of Darkness.

Graphs and Networks

To the mathematician the image below depicts a topological object called a graph. Civilians tend to call such objects networks. The nodes or vertices, as they are called, are connected by arcs or edges.

net

Such graphs can be used to represent many different kinds of phenomena, a road map is an obvious example, a kinship tree is another, sentence structure is a third example. The point is that such graphs are signs of phenomena, notations. They are not the phenomena itself.

Thursday, April 17, 2014

A digital humanist was walking to Damascus...

I’m wondering how many digital humanists set out to do one thing and ended up realizing they were doing something else, something they don’t quite understand. Some texts...

* * *

We who have been working in the field know that the digital humanities can provide better resources for scholarship and better access to them. We know that in the process of designing and constructing these resources our collaborators often undergo significant growth in understanding of digital tools and methods, and that this sometimes, perhaps even in a significant majority of cases, fosters insight into the originating scholarly questions. Sometimes secular metanoia is not too strong a term to describe the experience.

Willard McCarty, A Telescope for the Mind?

* * *

To Capture Infinity in a Bottle: The Digital Humanities and Cultural Criticism

The Gist: The only way the digital humanities are going to develop a cultural analytics that is sui generis is by thinking about the nature of computation itself in relation to human minds, as embodied in human brains, and as developing though interaction with other minds through various media and in groups of varying size and social structure. Otherwise the digital humanities will have no choice by to borrow its cultural concepts from other discourses, as it is now doing.

* * * * *

Let has start from some passages. First up, Alan Liu, from “Why I’m In It” x 2 – Antiphonal Response to Stephan Ramsay on Digital Humanities and Cultural Criticism (September 13, 2013):
The digital humanities can only take on their full importance when they are seen to serve the larger humanities (and arts, with affiliated social sciences) in helping them maintain their ability to contribute to the making of the full wealth of society, where “wealth” here has its older, classic sense of “well-being” or the good life woven together with the life of good.
Compare that with Willard McCarty, A telescope for the mind? (Debates in the Digital Humanities, ed. Matthew K. Gold. Minneapolis MN: University of Minnesota Press, 2012): “What can the digital humanities can do for the humanities as a whole that helps these disciplines improve the well-being of us all?”

Back to Liu:
It seems to me that digital humanists can and should evolve a mode of cultural criticism that is uniquely their own and not a mere echo of fading humanist cultural criticism by treating their immediate objects of inquiry (academically-oriented technologies and methods) as always also “mediate objects of inquiry” bearing on the way the human beings they wish their students could become (and they themselves could be on their best days) can really engage meaningfully with larger social agents and forces....The goal is to do research, to teach, and to live as if humanities technology is constantly intertwined with, reacts to, and acts on the way the links are now being forged between individuals (starting with those in the academy where we teach and conduct research) and the social-economic-political-technological constitution of contemporary society.

What it comes down to is that the digital humanities need both to work on tools and methods in their own institutional place (the academy) and to develop a capable imagination of the relation of that unique institutional place (or family of variant institutional spaces) to the other major institutions that play a part in enabling or thwarting the passageway from private human subjectivity to public social sensibility.
It seems that Liu is imagining a cultural criticism centered on institutions and society, one that treats computers and minds as black boxes whose inner workings remain unexamined. In this practice it seems to me that the ideas about culture and society would likely come from already existing bodies of work.

Monday, April 14, 2014

From Quantification to Patterns in Digital Criticism

I would like to continue the examination of fundamental presuppositions, conceptual matrices, which I began in The Fate of Reading and Theory. That post was concerned with how, in the context of academic literary criticism, 1) “reading” elides the distinction between (merely) reading some text – for enjoyment, edification, whatever – and writing up an interpretation of that text and 2) how “literary theory” became the use of theory in interpreting literary texts. This post is about the common sense association between computers and computing on the one hand and numbers and mathematics on the other.

* * * * *

Let’s start with a couple of sentences from one of the pamphlets published by Stanford’s Literary Lab, Ryan Heuser and Long Le-Khac, A Quantitative Literary History of 2,958 Nineteenth-Century British Novels: The Semantic Cohort Method (May 2012, 68 page PDF):
The general methodological problem of the digital humanities can be bluntly stated: How do we get from numbers to meaning? The objects being tracked, the evidence collected, the ways they’re analyzed—all of these are quantitative. How to move from this kind of evidence and object to qualitative arguments and insights about humanistic subjects—culture, literature, art, etc.—is not clear.
There we have it, numbers on the one hand and meaning on the other. It’s presented is a gulf which the digital humanities must somehow cross.

When first read that pamphlet most likely I thought nothing of that statement. It states, after all, a commonplace notion. But when I read those words in the context of writing a post about Alan Liu’s essay, “The Meaning of the Digital Humanities” (PMLA 128, 2013, 409-423) I came up short. “That’s not quite right,” I said to myself, it’s wrong to so casually identify computers and computing with numbers.”

* * * * *

Now let’s take a look at an essay by Kari Krauss, Conjectural Criticism: Computing Past and Future Texts (DHQ: Digital Humanities Quarterly, 2009, Volume 3 Number 4). Here’s her opening paragraph:
In an essay published in the Blackwell Companion to Digital Literary Studies, Stephen Ramsay argues that efforts to legitimate humanities computing within the larger discipline of literature have met with resistance because well-meaning advocates have tried too hard to brand their work as "scientific," a word whose positivistic associations conflict with traditional humanistic values of ambiguity, open-endedness, and indeterminacy [Ramsay 2007]. If, as Ramsay notes, the computer is perceived primarily as an instrument for quantizing, verifying, counting, and measuring, then what purpose does it serve in those disciplines committed to a view of knowledge that admits of no incorrigible truth somehow insulated from subjective interpretation and imaginative intervention [Ramsay 2007, 479–482]?