Showing posts with label Gavin. Show all posts
Showing posts with label Gavin. Show all posts

Monday, December 31, 2018

Toward a Theory of the Corpus [#DH]

I've just posted a new working paper. Title above, abstract, table of contents, introduction and section introductions are below. Download from:

Academia.edu: https://www.academia.edu/38066424/Toward_a_Theory_of_the_Corpus_Toward_a_Theory_of_the_Corpus_-by.
SSRN: https://ssrn.com/abstract=3308601.

Abstract: Recent corpus techniques ask literary analysts to bracket the interpretation of meaning so that we may trace the motions of mind. These techniques allow us to think of the mind as being, in some aspect, a high-dimensional space of verbal meanings. Texts then become paths through such a space. The overarching argument is that by thinking of texts as just ordered collections of physical symbols that are meaningless in themselves we can examine those collections in ways that allow us to recover the motions of mind as it constructs meanings for itself. When we examine a corpus over historical time we can see the evolution of mind. The corpus thus becomes an arena in which we investigate the movements of mind at various scales.

Contents

Meaning, text, and mind: Notes toward a theory of the corpus 2
PART 1: MAPPING A NEW ONTOLOGY OF THE TEXT 4
1. Can you learn anything worthwhile about a text if you treat it, not as a TEXT, but as a string of marks on pages? 6
2. Computational linguistics & NLP: What’s in a corpus? – MT vs. topic analysis 13
3. Why computational critics need to know about constitutive computational semantics 18
PART 2: VIRTUAL READING: PATHS THROUGH THE MIND AND THE MIND OVER HISTORICAL TIME 21
4. Augustine’s Path, A note on virtual reading 23
5. Mapping the pathways of the mind 29
6. Inferring the direction of the historical process underlying a corpus 39

Meaning, text, and mind: Notes toward a theory of the corpus

The set of observations I’ve collected in this working paper has two sources; it was spun out to scratch two conceptual itches. One is my long-standing interest in literary form. The other is the opposition or tension between meaning and, well, computation that has been dogging computational criticism for, I don’t know, a decade. Even computational critics who otherwise refuse to take that opposition as a criticism nonetheless tend to treat their mathematical models as scaffolding to support, or as gadgets for detecting, what really interests them. In the end I spent more time scratching that second itch than the first.

Computational critics have an opportunity to map the human mind that is qualitatively different from what interpretive critics accomplish by uncovering meanings ‘hidden’ in literary texts. But to avail themselves of this opportunity computational critics must understand the broad disciplinary framework in which “meaning” is opposed to “distant reading”. It is not simply that these are two different phenomena, or that “distant reading” is not intended to replace or supplant the explication of “meaning”, but that yoking them together in that opposition makes no more sense than opposing “salt” to “NaCl”.

The first three sections – Part 1: Mapping a new ontology of the text – deal with that kind conceptual difference. The last three sections – Part 2: Virtual reading: Paths through the mind and the mind over historical time – are about those new conceptual possibilities. I’ve provided some introductory material for both parts that is intended to help stitch these various arguments together. The overarching argument is that by thinking of texts as just ordered collections of physical symbols that are meaningless in themselves we can examine those collections in ways that allow us to recover the motions of mind as it constructs meanings for itself. We bracket the interpretation of meaning so that we may trace the motions of mind.

* * * * *

Part 1: Mapping a new ontology of the text – My overall objective here is to outline a way of thinking about language and texts that is centered on form and mechanism (linguistics) rather than meaning (literary criticism).

1. Can you learn anything worthwhile about a text if you treat it, not as a TEXT, but as a string of marks on pages? – Conventional literary criticism talks a lot about the text, but has no coherent conception of it. That is because it is focused on meaning and meaning doesn’t exist in the marks on pages, the physical text. Corpus techniques, topic modeling for example, have nothing but those marks and yet manage to reconstitute something that looks like meaning (but really isn’t, not quite). How is that possible? Moreover, by focusing on certain kinds of patterns in those marks, we can uncover formal structure in texts, structure that is otherwise invisible to conventional criticism, which also talks a lot about form without offering a coherent account of it.

2. Computational linguistics & NLP: What’s in a corpus? – MT vs. topic analysis – Corpora play very different roles in topic modeling and in machine translation. In topic modeling a corpus is the object of investigation while in machine translation a corpus is used to build a tool which then, in turn, does the translation. In MT the corpus allows us to create that Martin Kay calls an “ignorance model”. We would really like to be able to us a robust account of natural language semantics in MT; alas, we don’t have such a model (ignorance), so we use corpus techniques to construct a very crude approximation of semantics.

3. Why computational critics need to know about constitutive computational semantics – Simple, you need to know the lay of the land. That can be expressed in four contrasts: 1) close reading vs. distant reading, 2) meaning vs. semantics, 3) statistical semantics vs. computational semantics, and 4) corpus as tool vs. corpus as object. More often than not, corpus as tool is a substitute for constitutive computational semantics.

Part 2: Virtual reading: Paths through the mind and the mind over historical time – Assuming that we can think of the mind as, in some aspect, a high-dimensional network of verbal meanings, we can use statistical techniques to reveal the paths different texts trace through the mind and, beyond that, follow the mind as it evolves over historical time.

4. Augustine’s Path, A note on virtual reading – If we think of the mind as a high-dimensional space that can be approximated by statistical techniques, including those in analyzing texts, then we can see Andrew Piper’s statistical analysis of conversion texts, chiefly Augustine’s Confessions, as an analysis of mental structure. The statistical structure uncovered in the location of the 13 books of the Confessions can thus be reinterpreted as a pathway in the mind, of Augustine, but also of his readers. What are these different mental regions that are traversed in just this way?

5. Mapping the pathways of the mind – Michael Gavin uses vector semantics to examine a passage from Paradise Lost. After arguing that a word-space model is, after all, a model of the mind, I suggest that vector semantics could be used to map paths through the mind. I illustrate this conjecture by drawing a path for the Milton passage by picking words that had been brought to my attention by Gavin’s analysis. There’s no reason why such a path couldn’t be traced computationally.

6. Inferring the direction of the historical process underlying a corpus – Mathew Jockers’ final study in Macroanalysis (2013) attempted to investigate influence in a corpus of 3300 19th century novels. I argue that what he in fact discovered is that the socio-cultural process that created those novels is inherently directional. Without intending to do so, Jockers had in effect operationalized the 19th century idealist notion of Spirit and provided a way of thinking about “an autonomous aesthetic realm” (in a phrase from Edward Said).

Monday, December 3, 2018

Notes toward a theory of the corpus, Part 2: Mind [#DH]

Way back at the end of September I posted the first of a two, possibly three, post series: Notes toward a theory of the corpus, Part 1: History. I figured the second post would follow within a couple days and perhaps a third a few days after that.

I got delayed, diverted by other matters. But here we are with the second post.

As I said at the beginning of that post, by corpus I mean a collection of texts. The texts can be of any kind, but I am interested in literature, so I’m interested in literary texts. What can we infer from a corpus of literary texts?

Then I was interested in history. This time around I’m interested in the mind. History and the mind, two different, but not unrelated, phenomena. The fact that a given corpus consists of text by many different authors is essential to making historical inferences. The different authors published at different times; we can use a corpus of those different texts to arrive at inferences about historical process. In that earlier post I used Mathew Jockers’ Macroanalysis as my example. He had a corpus of 3300 Anglophone novels from the 19th century.

In this post I’ll be using Matthew Gavin’s work on a passage from Paradise Lost [1]. He uses a corpus of 18,351 documents drawn from Early English Books Online and dating from 1649 to 1699 (which covers the years when Milton wrote and published his epic). I don’t know how many authors are included in the corpus, but it doesn’t matter. Gavin is interested in only one of those authors, John Milton. And while he says nothing about Milton’s mind, he writes only about his text, I will argue that he is, in fact, investigating Milton’s mind. But also the mind of any sympathetic reader of Paradise Lost.

How is that possible? On the contrary, how is it not possible?

The corpus in vector semantics
For reasons I’ll explain shortly, if only superficially, Gavin needed a large body of texts to create a vector semantics model he could use in investigating Milton. I don’t know how many words were in those 18,000+ texts, but if each text were, on average, 10,000 words long, that would be a total of 180 million words; 50K words each would yield a corpus of 900 million words. I figure we’re dealing with 100s of millions if not over a billion words or continuous text. If Milton had written a lot Gavin could have used a corpus consisting entirely of Milton’s own texts. Milton didn’t, so Gavin couldn’t.

However many authors are included in that corpus, each has their own mind. And, at the margin, their own idiolect as well. Still, they hold English as a common tongue; if it wasn’t pretty much the same for each of them it wouldn’t function as a medium for communication. As long as we remind ourselves of what we’re doing, we can treat the lot of them as one somewhat idealized corporate author of the texts in the corpus. It is the semantics of that author that Gavin is applying to Paradise Lost.

Given our corpus, how do we create a semantic model? At this point I am going to more or less assume that you’ve already been through a good explanation of how vector semantics works. But I’d just like to remind you of some aspects of the process.

The process is founded on the assumption that words that occur close together in texts share some aspect of meaning. How can we turn that insight into a usable model? We create a co-occurrence matrix.

Here I’m using an example from Magnus Sahlgren [2]. Consider this text from Wittgenstein, “Whereof one cannot speak thereof one must be silent.” It contains nine word tokens from eight word types, to use terms from logic (one appears twice). Now we need to define what we mean by “close together”. Given two words, are the considered to be close if they are, for example, 1) within they same 1000 word string, 2) immediately contiguous, or 3) something else? For this example let’s choose the second criterion.

Given that, we can produce the following table:


whereof
one
cannot
speak
thereof
must
be
silent
whereof
0
1
0
0
0
0
0
0
one
1
0
1
0
1
1
0
0
cannot
0
1
0
1
0
0
0
0
speak
0
0
1
0
1
0
0
0
thereof
0
1
0
1
0
1
0
0
must
0
1
0
0
0
0
1
0
be
0
0
0
0
0
1
0
1
silent
0
0
0
0
0
0
1
0

It makes no difference whether we read by rows or columns as they are the same, but let’s read it by rows. If the word in a column is next to the target word we place a “1” in the column, otherwise “0”. So, whereof does not occur next to itself; a 0 goes in the first column. It does occur next to ˆ; a 1 goes in the next column. An so it goes for the rest of that row and for the rest of the rows. 

Note that as whereof is the first word in our text it can have only one word next to it. The same is true for silent, the last word. Since one occurs twice it is next to four other words. The remaining five tokens – cannot, speak, thereof, must, be – each have two neighbors, one before and one after. 

Each word is not associated with a string of eight numbers, which we can call a vector. We can now use those numbers to associate each word with a point in a space of eight dimensions, one for each word. In practice we wouldn’t bother with a corpus consisting of only one short text. But the principle remains the same given a corpus of 100s of millions of words constructed of tokens drawn from a population of, say, 100,000 types. Define what you mean by context and construct a co-occurrence matrix for your 100,000 types. Now you can associate each type with a point in a space of 100,000 dimensions – a rather breath-taking notion. 

That’s the general principle. Various methods can be used to reduce the number of dimensions we have to deal with, but regardless of how it is done we still end up with many more than three dimensions, which is the limit of what we can conveniently visualize. None of that matters to us. What matters to us is simply that we can associate the meanings of words with points in space and perform various operations in that space. 

Thus we have what has become the paradigmatic example for vector semantics [3]:
1) king - man + woman = queen
Other kinds of operation are possible:
2) paris - france + poland = warsaw
3) cars - car + apple = apples
All of these are based on analogies where a word is missing:
A : B :: C : ?
Remember, where you and I see a word the computer model sees a point in space, a point defined by a vector, which is a string of numbers. In each of these three cases it starts with a vector (that is, a string of numbers), subtracts a second vector from the first, then adds a third vector to that result, yielding a fourth vector, another point in space. The point is our answer. 

You and I may know the meanings of these words (examples 1 and 2) and the rules for pluralization (example 3), but the computer model knows none of that. It knows only the relations between word types that it can infer from the word-space constructed from the co-occurrence matrix based on the contexts in which word tokens occur. 

Gavin, however, isn’t interested in such analogies. Well, yeah, I rather suspect that he’s very interested in the fact that one can do such things with vector semantics, but that’s not what he does with the model he’s constructed. He uses it to examine a passage from Milton. 

As I mentioned at the outset, his co-occurrence matrix is based on 18,351 documents drawn from Early English Books Online. The structure in each of those documents must necessarily come from the mind of the document’s author. As those authors speak a common language we may, as I’ve argued above, think of the structure in that co-occurrence matrix as coming from the mind of a somewhat idealized corporate author of the corpus. Gavin is, in effect, asking that author to read a passage from Paradise Lost while we look on over their shoulder, as it were. 

That may not be how Gavin thinks of what he’s doing, but it’s how I’ve come to think of it. Call it a virtual reading.

Thursday, September 6, 2018

Why computational critics need to know about constitutive computational semantics [#DH]

That, I suppose, is the rubric under which I wrote a recent series of posts about Mathew Gavin’s article on Empson, ambiguity, vector semantics, and Milton [1]. That’s a topic that’s been on my mind in a somewhat different form for decades: Why literary critics need to know about computational semantics. Computational semantics was very important to me early in my career; it changed the way I thought about language, literature, and the mind.

For awhile I even thought it would change the conduct of literary criticism. I’d used it to analyze the underlying semantics of a Shakespeare sonnet. As things turned out, however, that was the last literary work I analyzed using that kind of model. It was too complicated and strenuous and subsequent developments did little to change that.

And yet, as I said, it had changed my intellectual life? But how could I convince literary critics of the importance of an analytic tool that I only used on one text? That seems a rather hard sell.

But practical criticism isn’t the only game in town, and computational critics aren’t traditional literary critics. There’s another reason for learning something about constitutive computational semantics (notice that I’ve added ‘constitutive’ back to the phrase): You need to know the lay of the land.

For the last several years it’s looked like close reading vs. distant reading, meaning vs., versus what exactly? But that’s NOT how it is; that’s not the territory. I wrote all those pieces around and about Gavin’s article in part to better map out the territory.

This post clarifies that map. I reprise four contrasts:
close reading <> distant reading
meaning <> semantics
statistical semantics <> computational semantics
corpus as tool <> corpus as object
Call them the lay of the land.

Difference

Meaning and understanding arise in part through difference and contrast. That’s why we’ve been having this conversation about close reading vs. distant reading, to better understand both terms of the opposition through comparison and contrast. This opposition is often discussed in terms of scale, which I find uninteresting, and meaning, as though distant reading has nothing corresponding to meaning. But of course that’s not true. As Moretti, among others, has pointed out, it has models [2], giving us meaning vs. models.

In the course of my Gavin posts I introduced another contrast, two of them in fact [3]:
meaning vs. semantics
statistical semantics vs. computational semantics
Meaning is inherently subjective [4]. We analyze meaning through a process of interpretation. Semantics, as I am using the term, is different. Semantic analysis requires an explicit model of the language system. The model is in the domain of objects, it is ontologically objective in Searle’s usage [5]. The whole world of so-called distant reading is ontologically objective in that sense (which doesn’t imply that it is objectively true, for the determination of truth is a matter of epistemology).

THAT is the distinction behind the opposition of close and distant reading. On the one hand we have a world where one interprets the meaning of texts. On the other hand we have a world where one builds semantic models of various kinds. The models used by computational critics are statistical in character. But we also have models that are constitutive in character, hence the distinction between statistical semantics and constitutive computational semantics.

Here the use of ‘computation’ is tricky. The various statistical techniques used in distant reading require a great deal of computing power, but the computer isn’t used to model or simulate a linguistic processes. That’s quite different from the computational semantics I learned early in my career. In that kind of semantics language processes were conceived as computational processes. Computation is thus constitutive of semantics, hence the rather awkward phrase, “constitutive computational semantics”. Moreover this kind of computational semantics shares one of its originating thought streams with the vector semantics Gavin used, machine translation (MT) of natural language. That is, it goes with the territory.

Corpus as tool vs. corpus as object

This brings us to Monday’s post [6] where I contrasted the role of the corpus in topic analysis and in MT. The contrast is simple and obvious in retrospect, but it took a bit of work to get there. In topic analysis the corpus itself is the object of investigation whereas in machine translation it is not. Rather, the corpus is used to build a translation system. The same, by the way, is true of vector semantics as Gavin has used. He wasn’t analyzing a corpus, rather he used a corpus to build a Word-Space he could use to analyze a passage from Paradise Lost.

This gives us another distinction: corpus as tool builder vs. corpus as object of investigation.

What does this have to do with computational semantics of the constitutive kind? In that post I quoted extensively from Martin Kay, a first generation computational linguist. He argued that, in the context of MT, statistical methods constitute what he called an “ignorance model”. If we had a rich and robust constitutive semantics we would need the statistical methods. Hence “statistics are standing in for a vast number of things for which we have no computer model” [7].

That’s simply not the case for topic modeling. Nor is it the case for Gavin’s vector semantics. That, I feel, requires a bit of commentary. But not here and now.

Later.

Reprise: The lay of the land in four contrasts
close reading <> distant reading
meaning <> semantics
statistical semantics <> computational semantics
corpus as tool <> corpus as object
References

[1] Michael Gavin, Vector Semantics, William Empson, and the Study of Ambiguity, Critical Inquiry 44 (Summer 2018) 641-673: https://www.journals.uchicago.edu/doi/abs/10.1086/698174

[2] Franco Moretti, Network theory, Plot Analysis, Stanford Literary Lab, Pamphlet 2, May 1, 2011. https://litlab.stanford.edu/LiteraryLabPamphlet2.pdf

[3] William Benzon, Gavin 5: Three modes of literary investigation and two binary distinctions, blog post, August 16, 2018, http://new-savanna.blogspot.com/2018/08/gavin-5-three-modes-of-literary.html

[4] William Benzon, The subjective nature of meaning, blog post, July 31, 2018, https://new-savanna.blogspot.com/2018/07/the-subjective-nature-of-meaning.html

[5] John Searle, The Construction of Social Reality, Penguin Books, 1995.

[6] William Benzon, Computational linguistics & NLP: What’s in a corpus? – MT vs. topic analysis, blog post, July 31, 2018, https://new-savanna.blogspot.com/2018/09/computational-linguistics-nlp-whats-in.html

[7] [1] Kay, M.: A Life of Language. Computational Linguistics 31(4), 425-438 (2005). http://web.stanford.edu/~mjkay/LifeOfLanguage.pdf

Sunday, August 19, 2018

Computation, Semantics, and Meaning: Adding to an Argument by Michael Gavin [#DH]

I've completed a new working paper. Title above. You can download it here:
Abstract, table of contents, and introduction below.

network distribution 2

Abstract: Michael Gavin has published an article in which he uses ambiguity as a theme for juxtaposing close reading, a standard procedure in literary criticism, with vector semantics, a newer technique in statistical semantics with a lineage that includes machine translation. I take that essay as a framework of adding computational semantics to the comparison, which also derives from machine translation. After recounting and adding to Gavin’s account of 18 lines from Paradise Lost I use computational semantics to examine Shakespeare “The expense of spirit”. Close reading deals in meaning, which is ontologically subjective (in Searle’s usage) while both vector semantics and computational semantics are ontologically objective (though not necessarily objectively true, a matter of epistemology).

Introduction: I heard it in the Twitterverse 3
Warren Weaver, “Translation”, 1949 4
Multiple meaning 5
The distributional hypothesis 6
MT and Computational semantics 8
A connection between Weaver 1949 and semantic nets 9
Abstraction and topic analysis 12
Two kinds of computational semantics 13
The meanings of words are intimately interlinked 14
Concepts are separable from words 14
Models and graphics (a new ontology of the text?) 16
Comments on a passage from “Paradise Lost” 17
Some computational semantics for a Shakespeare sonnet 21
Three modes of literary investigation and two binary distinctions 25
Semantics and meaning 25
Two kinds of semantic model 26
Where are we? All roads lead to Rome 27
Appendix 1: Virtual Reading as a path through a high-dimensional semantic space 28
Appendix 2: The subjective nature of meaning 30
Walter Freeman’s neuroscience of meaning 30
From Word Space to World Spirit? 32

Introduction: I heard it in the Twitterverse

Michael Gavin recently published a fascinating article in Critical Inquiry, Vector Semantics, William Empson, and the Study of Ambiguity , one that has developed some buzz in the Twitterverse. I liked the article a lot. But I was thrown off balance by two things, his use of the term “computational semantics” and a hole in his account of machine translation (MT). The first problem is easily remedied by using a different term, “statistical semantics”. The second could probably be dealt with by the addition of a paragraph or two in which he points out that, while early work on MT failed and so was defunded, it did lead to work in computational semantics of a kind that’s quite different from statistical semantics, work that’s been quite influential in a variety of ways.

In terms of Gavin’s immediate purpose in his article, however, those are minor issues. But in a larger scope, things are different. And that is why I’d composed the posts I’ve gathered into the working paper. Digital humanists need to be aware of and in dialog with that other kind of computational semantics. Gavin’s article provides a useful framework for doing that.
Caveat: In this working paper I’m not going to attempt to explain Gavin’s statistical semantics from the ground up. He’s already done that. I assume a reader who is comfortable with such material.
Warren Weaver, “Translation”, 1949

Let us start with a famous memo Warren Weaver wrote in 1949. Weaver was director of the Natural Sciences division of the Rockefeller Foundation from 1932 to 1955. He collaborated Claude Shannon in the publication of a book which popularized Shannon’s seminal work in information theory, The Mathematical Theory of Communication. Weaver’s 1949 memorandum, simply entitled “Translation” , is regarded as the catalytic document in the origin of machine translation (MT).

He opens the memo with two paragraphs entitled “Preliminary Remarks” (p. 1).
There is no need to do more than mention the obvious fact that a multiplicity of language impedes cultural interchange between the peoples of the earth, and is a serious deterrent to international understanding. The present memorandum, assuming the validity and importance of this fact, contains some comments and suggestions bearing on the possibility of contributing at least something to the solution of the world-wide translation problem through the use of electronic computers of great capacity, flexibility, and speed.

The suggestions of this memorandum will surely be incomplete and naïve, and may well be patently silly to an expert in the field - for the author is certainly not such.
But then there were no experts in the field, were there? Weaver was attempting to conjure a field of investigation out of nothing.

I think it important to note, moreover, that language is one of the seminal fields of inquiry for computer science. Yes, the Defense Department was interested in artillery tables and atomic explosions, and, somewhat earlier, the Census Bureau funded Herman Hollerith in the development of machines for data tabulation, but language study was important too.

A bit later in his memo Weaver quotes from a 1947 letter he wrote to Norbert Weiner, a mathematician perhaps best known for his work in cybernetics (p. 4):
When I look at an article in Russian, I say "This is really written in English, but it has been coded in some strange symbols.
 I will now proceed to decode."
Weiner didn’t think much of the idea. Yet, as Gavins explains, crude though it was, that idea was the beginning of MT.

The code-breaker assumes the message is in a language they understand but that it has been disguised by a procedure that scrambles and transforms the expression of that message. Once you’ve broken the code you can read and understand the message. Human translators, however, don’t work that way. They read and understand the message in the source language–Russian, for example–and then re-express the message in the target language–perhaps English.

Toward the end of his memo Weaver remarks (p. 11):
Think, by analogy, of individuals living in a series of tall 
closed towers, all erected over a common foundation. When they try to communicate with one another they shout back and forth, each from his own closed tower. It is difficult to make the sound penetrate even the nearest towers, and communication proceeds very poorly indeed. But when an individual goes down his tower, he finds himself in a great open basement, common to all the towers. Here he establishes easy and useful communication with the persons who have also descended from their towers.

Thus may it be true that the way to translate from Chinese to Arabic,
or from Russian to Portuguese, is not to attempt the direct route, shouting 
from tower to tower. Perhaps the way is to descend, from each language, down
 to the common base of human communication - the real but as yet undiscovered universal language - and then re-emerge by whatever particular route is convenient.
The story of MT is, in effect, one in which researchers find themselves forced to reverse engineer the entire tower in computational terms, all the way down to the basement where we find, not a universal language, but a semantics constructed from and over perception and cognition.

Thursday, August 16, 2018

Gavin 5: Three modes of literary investigation and two binary distinctions [#DH]

network distribution 2

Now that we’ve got two modes of semantic investigation on the table, statistical semantics and computational semantics, it’s time to add a third to the group. It’s been there all along, of course: close-reading. But it’s not about semantics at all, at least not as I’m using the term in this discussion. It’s about meaning. Semantics and meaning, that’s our first binary distinction.

Semantics and meaning

I hold that meaning – the meaning of literary works or of any work of art – is inherently subjective. Semantic analysis, in contrast, takes place in the object realm, if you will. That doesn’t mean that semantic analysis is necessarily or inherently objective in the sense of objective truth. That’s a different issue.

Here we need a distinction made by John Searle [1]. Both subjective and objective must be considered in ontological and epistemological senses. Color is ontologically subjective in this sense though, allowing for color blindness (which is fairly well understood), it is epistemologically objective. That is, color exists in minds of perceivers, not in phenomena, but the perceptual apparatus of perceivers is such that they agree on colors, where agreement is ascertained by matching color swatches (are they the same or different?), rather than naming the colors, which can vary from one individual to another, not to mention variation in cultural conventions about color names.

The meaning of texts is subjective in the ontological sense. That is the kind of phenomenon meaning is; it exists within the minds of subjects. Attempts to assay meaning through critical essays, moreover, seem to be subjective in the epistemological sense; they vary from one literary critic to another. I note, however, subjects may and often do converse with one another; intersubjectivity is possible.

Now consider Gavin’s vector semantics, or any type of statistical semantics. Whatever the specific statistical technique, it is grounded in a corpus. In Gavin’s case the corpus consists of 18,351 documents from 1640 to 1699 [2]. Gavin then constructed a Word Space using an explicit set of procedures. It is that Word Space that Gavin queried with words and sets of words from Paradise Lost, again using explicit procedures.

First, there is a clear and explicit distinction between the words of Milton’s text and the constructions of Gavin’s Word Space model and his queries against that model. The visualizations Gavin uses are one thing, Milton’s texts is another. They are different objects, though related by an explicit procedure [3].

That is not the case with ordinary interpretation. Whatever the critic believes may being going on in an author’s or a reader’s mind, the critic has no way of talking about that except through the words in the text under discussion – plus any interpretive apparatus they bring to bear on those words. The relationship between the terms and concepts of that interpretive apparatus and the words in the text exists in the critic’s mind and is not open to public inspection and, for that matter, resists introspection as well.

Moreover what Gavin did is, at least in theory, reproducible. In practice, not quite. He has noted in supporting material [4]:
Here's where I must apologize to any readers interested in recreating the exact images that appear in the article. I failed to preserve the scripts used when generating the charts and so I'm not sure exactly which parameters I used. As a result, following the commands below will result in slightly different layouts and slightly different most-similar word lists. These differences do not, I trust, affect the main points I'm hoping to make in the article, but they are worth noting. Following the instructions below will closely but not perfectly replicate what appears in Critical Inquiry. Also, it's worth noting that when creating the images in R, I exported them to PDF and tweaked them in Inkscape, adjusting the font and spacing for readability.
If he’d preserved those scripts, then others could run them and get the same results.

The point then is that we’re working in a world where there’s a great deal of explicit inference and construction, far more than is the case with traditional criticism – a point David Ramsey has made somewhere, though I’ve forgotten the citation.

The some is true for computational semantics. However a computational model is constructed, it is clearly a different conceptual object from and text one applies it to. In the example I gave in my previous post [5] the computational model I outlined is clearly distinct from the words of Shakespeare’s sonnet. Just how and why a specific construction is included in the model – there may be a reason grounded in empirical evidence from psychology, or there may be a more or less arbitrary computational reason – that’s a secondary issue at this point. My only point here is that these models are explicit, open to public inspection, and distinctly different from any text. Whether or not they represent are valid accounts of the human mind, that’s a different issue. It’s an important issue, but not my immediate concern.

Two kinds of semantic model

As I’ve already discussed the differences between statistical semantics and computational semantics in previous posts [6] I won’t go into any detail here. Statistical semantics is a way of analyzing semantic relationships between words. Computational semantics was devised to enact a semantic process. Some investigators conceive of that process as a simulation of a human mind while others are content to think of it as an artificially conceived process designed to achieve a practical result. In the large that’s an important difference, but for the purposes of this post, that difference is not so significant. What’s important is that we conceive of language processes as being computational in kind.

I note that there is a discipline called computational narratology [7] that draws on this work. The people who program computer games draw on techniques pioneered in this work. And why not, they’re in the business of telling stories, albeit interactively with the help of game players.

Wednesday, August 8, 2018

Gavin 4: Some computational semantics for a Shakespeare sonnet

Now we’re ready to look at a computational model I developed for a Shakespeare sonnet, 129, “The Expense of Spirit”. When did this work I had no intention of modeling the whole thing, much less implementing a computer simulation of such a model. I just wanted to do something that felt useful – a vague criterion if ever there was one, but nonetheless real – that afforded some kind of insight. In such a model the nodes in a graph are generally taken to represent concepts while the links between them represent relations between concepts, see my earlier post, Gavin 1.2: From Warren Weaver 1949 to computational semantics, for further remarks.

With that in mind, here’s the sonnet with modernized spelling:
1  The expense of spirit in a waste of shame
2  Is lust in action, and till action, lust
3  Is perjured, murderous, bloody, full of blame,
4  Savage, extreme, rude, cruel, not to trust;
5  Enjoyed no sooner but despised straight,
6  Past reason hunted, and no sooner had,
7  Past reason hated as a swallowed bait
8  On purpose laid to make the taker mad:
9  Mad in pursuit and in possession so,
10 Had, having, and in quest to have, extreme;
11 A bliss in proof, and proved, a very woe,
12 Before, a joy proposed, behind, a dream.
13   All this the world well knows; yet none knows well
14   To shun the heaven that leads men to this hell.
Let’s begin at the beginning. The first line and a half is generally taken as a play words. In one sense, expense means ejaculation and spirit means semen, making lust in action an act of sexual intercourse. You can’t get more sensorimotor than that.

But the lines can also be mapped into Elizabethan faculty psychology, in which spirit was introduced as a tertium quid between the material body and the immaterial soul(s). The rational soul exerted control over the body through the intellectual spirit or spirits; the sensitive soul worked through the animal spirit; and the vegetative soul worked through the vital spirit. Madness could be rationalized as the loss of intellectual spirit causing a situation in which the rational soul can no longer control the body. The body is consequently under control by man’s lower nature, the sensitive and vegetative souls. That too is lust in action. But the conception is abstract. It is this abstract lust in action that allows uncontrolled pursuit of the physical pleasures (and disappointments) of sex.

The uncontrolled and unfulfilling pursuit of sex can be expressed as a narrative that is at the core of the sonnet. The first 12 lines direct our attention back and forth over the following sequence of actions and mental states:
Desire: Protagonist becomes consumed with sexual desire and purses the object of that desire using whatever means are necessary: “perjur'd, murderous, bloody . . . not to trust” (ll. 3-4).

Have Sex: Protagonist gets his way, having “a bliss in proof” (l. 11).

Shame: Desire satisfied, the protagonist is consumed with guilt: “despisèd straight” (l. 5), “no sooner had/ Past reason hated” (ll. 6-7).
We might diagram that narrative sequence like this, where that loop at the top indicates that the whole sequence repeats (the diagrams are simplified from "Cognitive Networks and Literary Semantics"):

129 lust sequence
Figure 1: Lust Sequence
Line 4 looks at Desire (“not to trust”), then line 5 evokes Have Sex followed by Shame. Line 6 begins in Desire then moves to Have Sex, followed by Shame at the beginning of line 7, whose second half begins a simile derived from hunting. Line 10 begins by pointing to Shame, then to Have Sex, then to Desire, thus moving through the sequence in reverse order. It concludes by characterizing the whole sordid business as “extreme.” I leave it as an exercise for the reader to trace the sequence in lines 11 and 12.

Now consider this Figure 2, in which I have taken lines from the poem and superimposed them on the lust sequence with pointers from words in the text to nodes in the network:

129 lexeme & concept
Figure 2: Words and concepts
Those words ARE the text of Shakespeare’s sonnet. In the diagram they are physically distinct from the mental machinery postulated to be supporting their meaningfulness. Notice that many words are without pointers. The diagram is incomplete, which is obvious at a glance. That’s a virtue, not the incompleteness, but that it is readily apparent. When pursued properly, such models are brutally unforgiving in such matters.

Tuesday, August 7, 2018

Gavin 3: Comments on a passage from “Paradise Lost”

What Michael Gavin [1] did was to place traditional close-reading and computational criticism in the same intellectual context by using the concept of ambiguity as a tertium quid. My goal in this post is to recount his analysis and, by raising the issue of temporality – poems, and our experiences of them, unfold in time – edge near to some work on I did on a Shakespeare sonnet using “old school” symbolic computation, which I will offer in my next post.

Gavin choose to examine eighteen lines from book 9 of Milton’s Paradise Lost (455-472):
1)  Such Pleasure took the Serpent to behold 
2)  This Flourie Plat, the sweet recess of Eve 
3)  Thus earlie, thus alone; her Heav’nly forme 
4)  Angelic, but more soft, and Feminine,
5)  Her graceful Innocence, her every Aire
6)  Of gesture or lest action overawd
7)  His Malice, and with rapine sweet bereav’d
8)  His fierceness of the fierce intent it brought: 
9)  That space the Evil one abstracted stood
10) From his own evil, and for the time remaind 
11) Stupidly good, of enmity disarm’d,
12) Of guile, of hate, of envie, of revenge;
13) But the hot Hell that alwayes in him burnes, 
14) Though in mid Heav’n, soon ended his delight, 
15) And tortures him now more, the more he sees 
16) Of pleasure not for him ordain’d: then soon 
17) Fierce hate he recollects, and all his thoughts 
18) Of mischief, gratulating, thus excites.
I’ve highlighted words that Gavin identified for specific attention, but before commenting on them I want to offer a paraphrase of the passage, a minimal reading as Attridge and Staton [2] call it:
Satan, in the form of the Serpent, comes upon Eve in the garden and is overcome by her beauty. For a moment thoughts of evil are gone and he harbors no ill intent toward her. But as he continues to view her, the knowledge that he cannot have her terminates his pleasure and he hates her (again).
That is to say, the passage follows a movement in Satan’s mind, movement I’ll consider in a bit.

But first, to Gavin. After citing some traditional criticism of this passage, Gavin, without really saying why, zeros in on line nine–though I note that the last bit of criticism he cited was about that line. Note that, as the whole passage is 18 lines long, that line is at the passage’s center. When you drop out the stop-words–that, the, and one–we are left with four terms for investigation, space, Evil, abstracted, and stood. Gavin then produces visualized conceptual spaces for each of those four terms, for the four terms together, and for the entire passage.

Let’s first consider the space for abstracted:

Abstracted
Gavin notes: “abstracted connects to highly specialized vocabularies for ontology (entity, essence, existence), epistemology (imagination, sensation), and physics (selfmotion)...” (p. 668).

Here’s the space for the four words taken together as a composite vector:

13 space evil stood
Notice that space, evil, and stood are set in bold type and collocates of each appear in this compound space. But neither abstracted nor any of its collocates appears. As Gavin noted to me in a tweet, “Abstracted is abstracted away. It’s in there in the data, but it’s just too idiosyncratic for its collocates to float to the top.” I note that one edition of Paradise Lost glosses abstracted as withdrawn. That’s apt enough, I suppose, but misses the metaphysical and, well, abstract meanings of the word. Withdrawn could be a mere matter of physical location, but there’s much more than that at stake. As Gavin notes (671-672):
The space of evil is a space of action, of doings and deeds, of eschewing and overcoming, while Satan at this moment occupies a very different kind of space, one characterized by abstract spatiotemporal dimensionality more appropriate to angels, an emptiness visible only as an absence of action and volition.
Notice angel at the center of the diagram.

Wednesday, August 1, 2018

Gavin 2: The meanings of words are intimately interlinked [#DH]

This post is part of an ongoing series I am doing about Michael Gavin’s recent article on Empson, statistical semantics, and Milton [1], but it can also be read independently of the others in the series.

Concepts are separable from words

Let’s step away from Gavin’s vector semantics article and take a look at a couple of paragraphs from Gavin’s essay-review of Peter de Bolla’s The Architecture of Concepts: The Historical Formation of Human Rights [2]. Here he separates words from their meanings (p. 250):
For de Bolla, concepts aren’t equivalent to the words people use to denote meanings, nor even the meanings that people might have in mind. Instead, they form the abstract substrate of language, connecting words in ideational structures that enable thought without necessarily rising to the level of consciousness or explicit expression. Concepts are affiliations among words, sometimes teased out logically but just as often left unspoken, so they demand a different kind of analysis. Whereas intellectual histories usually proceed through chronologically arranged close readings, de Bolla’s method attempts to get underneath the (misleading) history of how words were used to direct attention instead to concepts as they really are in an almost Platonic sense: “my aim is to parse the grammar of the concept of rights and to describe its distinctive architecture within the culture of the English language eighteenth century” (63). The “architecture” of a concept is its relationships to other concepts.
Such ideas were common among the computational semanticists of the 1970s and 1980s, thinkers in psychology, computational linguistics, and artificial intelligence (aka cognitive science). Words, more precisely, lexemes, were one thing, concepts another, with much attention being given to the relations of concepts among themselves. These concepts belonged to what came to be called “the cognitive unconscious” as opposed to, for example, the tangle of affect and desire that populations the unconscious of psychoanalysis.

Gavin goes on to assert (p. 252):
Without polemicizing, Architecture of Concepts lays out a theory of conceptuality that has the potential to upend large-scale quantitative research. If concepts exist culturally as lexical networks rather than as expressions contained in individual texts, the whole debate between distant and close reading needs reframing. Conceptual history should be traceable using techniques like lexical mapping and supervised probabilistic topic modeling.
Here I call attention to the word “networks”. Precisely, concepts form networks among themselves, networks that are independent of individual texts. Cognitive scientists turned to the investigation of semantic networks because they were interested in how we understand language or, more pragmatically, in creating computer systems that understand language. That seemed to require a semantics that was separate from the stuff of language–lexemes, syntax, and the rest–and linked to cognition and perception.

As for reframing the debate between distant and close reading, isn’t that what Gavin has done in his article on vector semantics [1]? In this passage he’s talking about a sense of the relationship between words and meaning that emerges from Empson’s critical practice (p. 647):
The playful exuberance of Empson’s method tries to capture something of words’ malleability and extensibility. Meanings aren’t discrete things, even though they inhere to words; instead they unfold over many dimensions of continuously scaled variation. Words have bodies and agency, Empson argues. Even a sort of personhood. They occupy an invisible lexical “thoughtspace” where they break apart and recombine to form superstructures, molding opinions and otherwise forging human experience.
A bit later (p. 659):
Empson always pushed against the dictionary he made use of. Meanings weren’t discrete objects for him but makeshift focal points across which the ambiguities of a poem could be viewed.
I like the assertion that “meanings aren’t discrete things”, something my teacher, David Hays, emphasized to me, referencing the work of his friend and colleague, Sydney Lamb [3] – both are of Chomsky’s generation, but with a very different predilections. Meanings are distributed in a network where the nodes are mutually defining [4]. They aren’t discrete objects.

Gavin finds (more or less) this conception in the vector semantics of Word Spaces. And I rather suspect that his work with vector semantics has influenced his reading of Empson. When he talks of meanings unfolding “many dimensions of continuously scaled variation”, that sounds like vector semantics, not Empson, which is fine.