Showing posts with label StanfordLitLab. Show all posts
Showing posts with label StanfordLitLab. Show all posts

Monday, July 16, 2018

Continuity and discontinuity in literary history, H vs. DH? (Literary Lab 4) – Tales from the Twitterverse [#DH]

If, a few years ago, you’d asked me whether or not a half dozen scholars could have an interesting and fruitful conversation 280 characters at a time, I’d have said “Are you freakin’ kidding me?”–or words to that effect. But it happens, not all the time, but often enough, and it even happened when we were restricted to 140 characters at a time. Such is life in the academic Twittersphere.

Ted Underwood kicked one off on Friday (the 13th, FWIW) and it continued on into Saturday. A half dozen or so joined in and who knows how many followed along–a dozen, 30, 50, more, who knows? I don’t know what Ted expected when he threw that first tweet into the maelstrom, nor does that much matter. What matters is what happened, and that was unplanned, spontaneous, even fun.

Of course, it requires interlocutors who understand the issues at hand well enough that they can speak in code. And they need to read one another charitably and in good faith. But, given those conditions, it is possible to do some good work.

But enough of this. Let’s get to it.

In the next section of this post I reproduce three tweets from that conversation and add some commentary. These set the stage for the next two sections, where I use the well-known Stanford Literary Lab Pamphlet 4 to interrogate the issues raised in the first section; this fleshes out an argument I tossed into the conversation in a five-tweet string.

Dropping science in the Twitterverse

Here’s the tweet that started things off:

So, yeah, it’s a bit aggressive: maybe the received wisdom ain’t necessarily so (as the song goes). And periodization is just dropped in there at the end, as a for-example. Of course, it’s an important case because literary studies is more or less organized according to periods: fields of study, journals, conferences, professional organizations, coursework, all organized by period. To be, period isn’t the only parameter of organization, we’ve also got author, genre, and a blither or theoretical proclivities, but it’s an important one.

And it’s one that Underwood has investigated. Here’s the abstract of an article he and Jordan Sellers recently published, The Longue Durée of Literary Prestige [1]:
A history of literary prestige needs to study both works that achieved distinction and the mass of volumes from which they were distinguished. To understand how those patterns of preference changed across a century, we gathered two samples of English-language poetry from the period 1820–1919: one drawn from volumes reviewed in prominent periodicals and one selected at random from a large digital library (in which the majority of authors are relatively obscure). The stylistic differences associated with literary prominence turn out to be quite stable: a statistical model trained to distinguish reviewed from random volumes in any quarter of this century can make predictions almost as accurate about the rest of the period. The “poetic revolutions” described by many histories are not visible in this model; instead, there is a steady tendency for new volumes of poetry to change by slightly exaggerating certain features that defined prestige in the recent past.
Those “poetic revolution” imply period boundaries, but those boundaries didn’t show up. To be sure, the model did show change, but it was gradual and seemed to have a historical direction–which is a whole different kettle of conceptual and perhaps even ideological fish.

I rather liked that work, and took a close look at it [2]. I particularly liked the apparent directionality it revealed, but let’s set that aside. As for the lack of period boundaries...well, I don’t think periodization is simply a matter of disciplinary hallucination. So I’m inclined to think that the method Underwood and Sellers used simply isn’t sensitive to such matters. But that’s not an argument; it’s merely a casual assertion. Underwood is right, there IS work to be done.

And that’s what caught people’s attention. Daniel Shore entered and exchanged tweets with Ted. Then I suggest color perception as an analogy:

In making that suggestion I had more in mind than the mere presence of categories over continuity. Color perception is tricky, and much studied. We know, and have know so for quite awhile, that there isn’t a simple and direct relationship between perceived color and wavelength of light. This is not the place for a primer in color vision (the Wikipedia article is a reasonable place to begin). But I can offer a few remarks.

Blue, for example, does not map directly to a fixed region in the electromagnetic spectrum, with violet and green in adjacent fixed regions. The color of a given region in the visual field is calculated over input from three kinds of retinal receptors (cones) having different sensitivities. Moreover the color of one region is “normalized” over adjacent regions. The upshot is that a red apple will appear to be red under a wide range of viewing conditions. Different illumination means different wavelengths incident on the apple. Hence light that the apple reflects to the eye varies in different situation. Because the brain normalizes, the apple’s color remains constant.

We’ll return to part of that story a bit later.

Let’s return to the twitter conversation with a pair of remarks by Ryan Heuser:

The conversation continued on. Others joined. It forked here and there. And somewhere in there I introduced a string of tweets about Lit Lab 4.

Tuesday, April 26, 2016

From Telling to Showing, by the Numbers

I've been thinking about some remarks Moretti made about the digital humanities in  a recent interview. Among other things he suggested that the results of computational criticism have so far been disappointing. But he also held up Lit Lab Pamphlet #4 as an example of "an intelligence that takes the form of writing a script, but in the writing of the script there is also the beginning of a concept, very often not expressed as a concept, but that you can see that it was there from the results that the coding produces." Here's what I wrote about that pamphlet back in October of 2012.
I’ve just looked at a pamphlet from Stanford’s Literary Lab: Ryan Heuser and Long Le-Khac, A Quantitative Literary History Of 2,958 Nineteenth-Century British Novels: The Semantic Cohort Method (68 page PDF), May 2012. I’ve not read it in detail, but only blitzed my way through, looking for the good parts. Well, not even all of those. I was just looking to get a sense of what’s going on.

Which I did. And I like it. THIS is the sort of work I want to see from ‘digital humanities.’ Not the only sort, but it’s one of the things we can do with ‘big data’ and pretty much only do with big data. If traditional humanists can’t see value in this kind of work, well, then forget about them.

First I’ll give you the abstract, then I’ll quote a bunch and make some comments.

Authors’ abstract
The nineteenth century in Britain saw tumultuous changes that reshaped the fabric of society and altered the course of modernization. It also saw the rise of the novel to the height of its cultural power as the most important literary form of the period. This paper reports on a long-term experiment in tracing such macroscopic changes in the novel during this crucial period. Specifically, we present findings on two interrelated transformations in novelistic language that reveal a systemic concretization in language and fundamental change in the social spaces of the novel. We show how these shifts have consequences for setting, characterization, and narration as well as implications for the responsiveness of the novel to the dramatic changes in British society.

This paper has a second strand as well. This project was simultaneously an experiment in developing quantitative and computational methods for tracing changes in literary language. We wanted to see how far quantifiable features such as word usage could be pushed toward the investigation of literary history. Could we leverage quantitative methods in ways that respect the nuance and complexity we value in the humanities? To this end, we present a second set of results, the techniques and methodological lessons gained in the course of designing and running this project.

Friday, May 2, 2014

Method in DH: Signal and Concept, Operationalization

Signal and concept: distinction employed by Ryan Hauser and Long Le-Khac: A Quantitative Literary History of 2,958 Nineteenth-Century British Novels: The Semantic Cohort Method. Stanford Literary Lab, Pamphlet 4, May 2012.

Operationalization: concept discussed in Franco Moretti, “Operationalizing”: or, the Function of Measurement in Modern Literary Theory, Stanford Literary Lab, Pamphlet 6, December 2013.

Think like a social scientist: Arthur Stinchcombe, Constructing Social Theories, New York: Harcourt, Brace, and World, 1968.

• • • • • •

Toward the end of their pamphlet Hauser and Le-Khac observe (pp. 46-47):
How do we get from numbers to meaning? The objects being tracked, the evidence collected, the ways they’re analyzed—all of these are quantitative. How to move from this kind of evidence and object to qualitative arguments and insights about humanistic subjects— culture, literature, art, etc.—is not clear. In our research we’ve found it useful to think about this problem through two central terms: signal and concept. We define a signal as the behavior of the feature actually being tracked and analyzed. The signal could be any number of things that are readily tracked computationally: the term frequencies of the 50 most frequent words in a corpus, the average lengths of words in a text, the density of a network of dialogue exchanges, etc. A concept, on the other hand, is the phenomenon that we take a signal to stand for, or the phenomenon we take the signal to reveal. It’s always the concept that really matters to us. When we make arguments, we make arguments about concepts not signals.
I’ve already noted that I believe the equation of computing with quantitative is misleading (From Quantification to Patterns in Digital Criticism). That’s one thing.

The solution Heuser and Le-Khac propose is to distinguish between signals, which they can measure, and concepts, which are the phenomena being indicated by the signals. The distinction is useful and it’s also pretty much standard in social science, though not necessarily in those terms. While the mechanisms responsible for the operation of, for example, an automobile, can be observed directly, those operating in human society or the human mind are generally inaccessible to direct observation.

In order to investigate those mechanisms one must propose a theory or a model and then operationalize that theory. What does that mean? Moretti (p. 1):
Forget programs and visions; the operational approach refers specifically to concepts, and in a very specific way: it describes the process whereby concepts are transformed into a series of operations—which, in their turn, allow to measure all sorts of objects. Operationalizing means building a bridge from concepts to measurement, and then to the world. In our case: from the concepts of literary theory, through some form of quantification, to literary texts.
Just so.