Showing posts with label word illusion. Show all posts
Showing posts with label word illusion. Show all posts

Friday, November 22, 2024

Time for another ramble: Melancholy, Claude, Bloom, Claude, Ring composition, and Other Stuff

It’s been a while since I’ve done one of these; May 29th was the last one. If you look over there to right at the Blog Archive you’ll see I’ve been on a posting slump, with 3-figure monthly totals from January through June, then a dip to 61 for July, August: 18, September: 15, October: 30, and now 33 for November as I write this, and the month isn’t over. Maybe I’m pulling out of the slump.

Anyhow, I’m feeling a little backed up with things to post about, so it’s time to ramble on and see what’s up.

Melancholy, Mind (Mine), and Growth

That’s the tentative title for my next 3 Quarks Daily article. Starting back in November 2017 I’ve been making occasional posts about my monthly posting habits, which tend to drop during the winter. I’m thinking of using that as the point of departure for my next 3QD piece, which will go up on December 2nd.

During those down times I’m depressed to one degree or another (melancholy). But why? Since those down times have been in the winter, perhaps its seasonal affective disorder (SAD). But that doesn’t square with all of the evidence. There was no down-time in the winder of 2022-2023 and 2023-2024, but there was a slump in the summer of 2023. Something else is going on, and I think it has to do with creativity. To that end I want to discuss the ridiculous blither of tags here, 665 by November or 2023.

Claude

I’ve starting working with Claude, Anthropic’s chatbot. I want to do some posts where I verify some of the work I’ve done with ChatGPT. I’m thinking of posts on stories, ontological structure, abstract definition, and the Girardian analysis of Jaws. I can then gather those into a working paper.

I also want to look at other things. At the moment I’m thinking of seeing how Claude summarizes longish documents. I’m thinking of the Hamlet chapter from Bloom’s Shakespeare book and Heart of Darkness.

Harold Bloom and GOAT literary critics

A year ago I began a series of posts on the theme of the greatest literary critics. I got bogged down in discussing Harold Bloom. It’s time to finish it off.

Bloom may well be as brilliant a literary critic as we’ve had in the last 50 or 60 years. But brilliance is one thing, greatness is another. Brilliance is a function of the individual, while greatness is a function of the relationship between an individual’s work and the arena in which they’re working.

I’m not sure about Bloom’s fit. While he’s got a wide readership, it’s not clear to me that scholars have taken up his work in any significant way. They may cite him – perhaps especially is concept of influence – but that they don’t much use of his ideas in his work. But we’ve also got to consider his work in the larger public arena, where he is hands-down the most prominent literary critic. I’m not sure of how to handle that.

However, if History wants to declare that Harold Bloom is one of the great all-time literary critics, maybe even the GOAT, what do I care? What would really bother me is if future critics should decide to take his work as a model and (attempt to) do more like it. Like most literary critics he’s neglected the study of form and he’s been deaf to the cognitive sciences. There’s little in his work that’s worth amplifying. It’s a dead end.

ChatGPT report

About a year or so ago I started writing a report summarizing my work on ChatGPT. I need to finish that report. I’d estimate the three-fourths or more are done. I’d like to be able to include some work with Claude. I don’t intend a lot on this, just enough to say that I’ve verified some things.

I’d like to finish this by the end of this year.

Why’s ring-form composition important?

That’s tricky. It has to do with the fact that literary works are extended in time, unlike the visual arts, which are static in time but extended in space. You can’t take the whole thing in at a glance like you can a painting.

There are constraints on how a literary work can unfold in time. There is a sense in which (the nature of) the end is inherent in the beginning. Ring-compositions are even more tightly constrained. Dylan Thomas consciously and deliberately plotted the ring-composition of the rhyme scheme in his “Author’s Prologue.” Rhyme is not about meaning; its patterns are arbitrary with respect to meaning. But Coleridge did not consciously work out the ring-compositions in “Kubla Khan,” not Conrad in Heart of Darkness. These patterns ARE NOT arbitrary with respect to meaning. On the contrary, they are central to how meaning is constituted.

Other Stuff

More Cobra Kai: Follow-up on my post where I explore the Freudian angle, saying a bit more about Girard, and extending that into history.

* *

More on meaning in LLMs: I’ve suggested that Ilya Sutskyver conflates mistakenly (and unknowingly?) conflates cognitive and semantic structure with the structure of the world and so suggests that robust next-token prediction requires knowledge of the world. In this post I argue that making that distinction is, in fact, difficult, and involves what I’ve been calling the word illusion. I first confronted the problem as an undergraduate when I was trying to understanding the difference between the signified, and mental structure, and the reference, something in the world, of a sign. I may not have gotten deep intuitions about that until I began studying cognitive networks in graduate school.

* *

Gila-monster venom and computational irreducibility: The idea is to start with a NYTimes article on drug discovery that starts with Gila-monster venom and ends up with Wolfram’s concept of computational irreducibility. This is about search, computation, and the complex and irregular structure of the (natural world).

* *

My work with Ramesh: Notes on conceptual ontology and hypergraphs in conceptual space.

* *

LLMs and literary study: Can we use LLMs to analyze the thematic structure of literary texts?

In particular, can we use them to examine Bloom’s these about Shakespeare as “inventing” the human. Are there themes that appear first in Shakespeare? We need more than Bloom’s vigorous assertion on this. We need to examine the thematic structure of prior texts and of Shakespeare’s texts and show that there are things new in Shakespeare. Do those new things then continue if texts after Shakespeare? To do this properly we need to examine a lot of texts.

I’d also like to know if LLMs could be used to find ring-composition in literary texts. It is by no means obvious to me that they can.

Monday, February 5, 2024

OpenAI Co-Founder Ilya Sutskever on the mystical powers of artificial neural nets

Transcription (which I found here):

Ilya Sutskever: I challenge the claim that next-token prediction cannot surpass human performance. On the surface, it looks like it cannot. It looks like if you just learn to imitate, to predict what people do, it means that you can only copy people. But here is a counter argument for why it might not be quite so. If your base neural net is smart enough, you just ask it — What would a person with great insight, wisdom, and capability do? Maybe such a person doesn’t exist, but there’s a pretty good chance that the neural net will be able to extrapolate how such a person would behave. Do you see what I mean?

Dwarkesh Patel: Yes, although where would it get that sort of insight about what that person would do? If not from…

Ilya Sutskever: From the data of regular people. Because if you think about it, what does it mean to predict the next token well enough? It’s actually a much deeper question than it seems. Predicting the next token well means that you understand the underlying reality that led to the creation of that token. It’s not statistics. Like it is statistics but what is statistics? In order to understand those statistics to compress them, you need to understand what is it about the world that creates this set of statistics? And so then you say — Well, I have all those people. What is it about people that creates their behaviors? Well they have thoughts and their feelings, and they have ideas, and they do things in certain ways. All of those could be deduced from next-token prediction. And I’d argue that this should make it possible, not indefinitely but to a pretty decent degree to say — Well, can you guess what you’d do if you took a person with this characteristic and that characteristic? Like such a person doesn’t exist but because you’re so good at predicting the next token, you should still be able to guess what that person who would do. This hypothetical, imaginary person with far greater mental ability than the rest of us

Yikes! If a stream of tokens is the only thing the machine has access to, then just how is it to divine the underlying reality? It's basing its predictions on its experience of the token stream, nothing else, N O T H I N G. These folks seem deeply enmeshed in what I've been calling the word illusion in a number of posts. 

This is the A.I. equivalent of believing the earth is flat.

Wednesday, October 4, 2023

Entanglement and intuition about words and meaning

Two things have just occurred to me about my recent post, Word meaning and entanglement in LLMs:

1.) That the issue is one of intuition as well, and
2.) that we’re dealing with system 1 thinking, in the System 1/system 2 dichotomy popularized by Daniel Kahneman.

I’ve not yet read Kahneman’s book – Thinking Fast, Thinking Slow, though it’s on my “to be read someday” shelf – but I gather that System 1 is fast, intuitive and largely tacit (to use a word from Michael Polanyi) while System two is slow, deliberate, and logical.

My argument in that earlier post is that, in effect, our default notion of word meaning is that it is atomic and discrete. When words are linked in phrases, sentences, and paragraphs, it is liking beads on a thread, or freight cars in a train. The linkage is external and contingent. Without reflection, that’s just how we think about words (and meaning). That’s fine in informal discussions, but not so good in at least some technical contexts, such as large language models (LLMs).

Now, let’s take the idea that LLMs are “trained” by being asked to predict the next word. That is at least consistent with, if not actually reinforcing of, this default conceptualization of atoms-of-meaning. One can easily make predictions about the behavior of atoms. One simply observes them and notes down what they do from one moment to the next. There is no sense of “interiority.”

Whereas the idea that words are entangled with one another through their meanings, that’s all about “interiority.” Those vectors are “interior” to the token, and relate one token to another and, more generally, tokens among themselves. The idea of entanglement leads naturally to the idea of weaving, weaving a fabric of meaning. The so-called prediction procedure, then, is one of placing a word, with its 12K item vector, into the unfolding fabric of meaning. Backpropagation, in this view, is the act of fine-tuning the placement. Prediction is merely a means to an end, a device, not the point of the procedure.

To think in terms of atomic meaning is simply to gloss over all this. All of that may be implicit in the mathematics, but the atomic view of meaning stands in the way of allowing ones thought to be perspicuously guided by the mathematics. The mathematics becomes (and functions as) a secondary construction.

I note finally that traditional training in propositional and symbolic logic reinforces this atomic view of word meaning. Word meaning is reduced to variable names, Ps and Qs, having no intrinsic content whatsoever. That’s find for System 2 deliberative thinking, which is what logic was invented for. But it gets in the way of understanding how meaning works in collections of entangled, entangled what? What do we call them?

This leads to a final irony: The world of standard computer programming is close kin to that of symbolic and propositional logic, with their variables, bindings, and values. Thus the mode of thinking necessary for programming the computational engines that create LLMs, that mode of thought stands in the way of understanding how LLMs work. The AI/ML experts who create the models are thus crippled in understanding how they work. The intuitions that guide them in writing code render the operations of LLMs opaque and invisible when deployed in understanding them.

This opacity thus has two aspects:

1.) the sheer complexity of the models, and
2.) conceptual intractability.

I am suggesting, then, that thinking of meaning as entailing entanglement is a way to deal with the second issue (and this may also lead to a holographic account as well, but this is a secondary issue). On the first issue, complexity, that is there regardless of your conceptual instruments. Thinking in terms of entanglement will NOT eliminate the complexity, but it may well make it tractable

If your goal is mechanistic interpretability, then you need conceptual tools appropriate to the mechanisms you are trying to understand, no? You need to discard, or at least bracket, intuitions based on the idea of atomic-self-contained word meaning and develop intuitions that are consistent with the mathematics underlying the LLMs.

Monday, October 2, 2023

Word meaning and entanglement in LLMs

It is my impression that, unless someone has had experience with distributed accounts of word meaning, they’re likely to think of word meaning as an enclosed “atom” of meaning, distinct from other such atoms, but like word forms themselves. The meaning of a proposition or a sentence is just composed of a string of such atoms of meaning, as a freight train is composed of a string of cars. I like to oppose this with a different metaphor, dropping pebbles into a pond, one after the other. Each pebble sends ripples across the surface of the pond. The succession of ripples from each pebble interferes with the others. That growing interference pattern is the meaning of the string.

And that’s how we need to think about meaning in LLMs, sorta’. Each word consists of a token and the vector encoding its meaning as an embedding in a high-dimensional space – roughly 12K, I believe, for GPT-3. Given two words, we can compare their vectors, dimension by dimension. Where the words are closely related, they should have similar, perhaps even identical, values along some dimensions. Where the words are highly dissimilar, they will share few or no values.

I find the idea of entanglement useful here. Some words have meanings that are closely entangled, while others do not. We can think of an embedding model as an entanglement matrix. This matrix shows how the meaning of any one word is a function of it position in the matrix. When you present a prompt to, say, ChatGPT, it generates an output by calculating the entanglement of the prompt with the language model.

Contrast this way of thinking with the standard, “Generate the net token, and the one after that, and so on.” The standard way of thinking has you thinking in terms of atomic units, tokens, and obscures the nature of the process, making it seem deeply obscure, even magical. Just what’s going on when the underlying model is “calculating the entanglement” of the prompt with the model is not at all obvious – I can’t tell you what it is – but it has a different feel. Similarly, training by “predict the next word” is really a way of calculating the entanglement of the text with the whole model, for the whole text (in the context window) is involved in the calculation, not just the leading word.

More later.

Sunday, September 17, 2023

Cultural Evolution 8: Language Games 1, Speech

Once again I'm bumping this to the top, this time to emphasize the material I've highlighted in yellow, that the meaning of words is constantly being negotiated through interaction with others. I wish to posit, polemically and provisionally, that what happens within a single head, a single brain, a single mind, that that be thought of as purely mechanical, purely a matter of relationality and adhesion, to use terms I've recently adopted. What happens when people negotiate meaning through interaction is perhaps not so mechanical. This is where freedom and novelty enter the system, where epistemic difference forces us to renegotiate the world.

* * * * *

I'm bumping this post, from 2010, to the top of the queue for two reasons: 1) the section "Language Games and Game Theory" is germane to my recent post, Why do we need a genotype-phenotype distinction for cultural evolution? Because minds are built from the inside.The post proposed as an answer: minds are built from the inside. From that it follows that we can't read one another's minds, which is my point of departure in this post. 2) The following section, "What is a language and what are the memes?," is where I first worked out my current approach to the genetic component of culture, which I have since come to call coordinators. The rest of the posts in this particular series are gathered under the tag CE workshop. Note: You might want to read the comments for this post.
* * * * *

The key to the treasure is the treasure.
– John Barth

But I’m not talking of language games in Wittgenstein’s sense, though the Wittgenstein of the Tractatus had a considerable influence on me as an undergraduate. No, I’m thinking of game theory, not something I’ve studied, though I did have an undergraduate course on decision theory taught by R. B. Braithwaite. But I’m getting ahead of the game.

As the title says, this post is about language. There’s been a fair amount of work done on language from an evolutionary point of view, which is not surprising, as historical linguistics has well-developed treatments of language lineages and taxonomy, the “stuff” of large-scale evolutionary investigation. While this work is directly relevant to a consideration of cultural evolution, however, I will not be reviewing or discussing it. For it doesn’t deal with the theoretical issues which most concern me in these posts, namely, a conceptualization of the genetic and phenotypic entities of culture. This literature is empirically oriented in a way that doesn’t depend on such matters.

The Arbitrariness of the Sign
 
In particular, I want to deal with the arbitrariness of the sign. Given my approach to memes, that arbitrariness would appear to eliminate the possibility that word meanings could have memetic status. For, as you may recall, I’ve defined memes to be perceptual properties – albeit sometimes very complex and abstract ones – of physical things and events. Memes can be defined over speech sounds, language gestures, or printed words, but not over the meanings of words. Note that by “meaning” I mean the mental or neural event that is the meaning of the word, what Saussure called the signified. I don’t mean the referent of the word, which, in many cases, but by no means all, would have perceptible physical properties. I mean the meaning, the mental event. In this conception, it would seem that that cannot be memetic.

That seems right to me. Language is different from music and drawing and painting and sculpture and dance, it plays a different role in human society and culture. On that basis one would expect it to come out fundamentally different on a memetic analysis.

This, of course, leaves us with a problem. If word meaning is not memetic, then how is it that we can use language to communicate, and very effectively over a wide range of cases? Not only language, of course, but everything that depends on language. Literature obviously – which I’ll take up in the next post – but much else as well.

Speech as a Means of Communication
 
Willard van Orman Quine has given us a classic thought experiment that points up the problem of word meaning. He broaches the issue by considering the problem of radical translation, “translation of the language of a hitherto untouched people” (Quine 1960, 28). He asks to consider a “linguist who, unaided by and interpreter, is out to penetrate and translate a language hitherto unknown. All the objective data he has to go on are the forces that he sees impinging on the native’s surfaces and the observable behavior, focal and otherwise, of the native.” That is to say, he has no direct access to what is going on inside the native’s head, but utterances are available to him. Quine then asks us to imagine that “a rabbit scurries by, the native says ‘Gavagai’, and the linguist notes down the sentence ‘Rabbit’ (of ‘Lo, a rabbit’) as tentative translation, subject to testing in further cases” (p. 29). And thus begins one of the best known intellectual romps in the philosophy of language.

Quine goes on to argue that, in thus proposing that initial translation, the linguist is making illegitimate assumptions. Perhaps he begins his argument by noting that the native might, in fact, mean “white” or “animal” and later on offers more exotic possibilities, the sort of things only a philosopher would think of. Quine also notes that whatever gestures and utterances the native offers as the linguist attempts to clarify and verify will be subject to the same problem. Quine’s argument is thorough and convincing.

When he did that work, however, he did not, of course, have access to a range of more recent work in cognitive anthropology and evolutionary psychology that indicated that our adapted minds have a preferred way of parsing the world, as do baboons. To be sure, this is “overwritten” and augmented in culture-specific ways, but those underlying perceptual and cognitive systems do not disappear. To consider a specific example, the work on folk taxonomy (Berlin 1992) suggests that there is a so-called basic level of designation, and that is at the level of “rabbit” and not “animal” (in fact, many languages don’t even have a word at that level of generality). So the linguist is reasonable in assuming “rabbit” is a more likely translation than “animal.” Other considerations are likely to rule out “white” or Quine’s other suggestions. I have no reason to believe that this cognitive architecture so constrains matters that there is only one possible referent for “Gavagai.” But I do think that it is likely to turn out that, all other things being equal, “rabbit” is in fact the best guess.

This situation, of course, is rather different from that of ordinary speech between people who share a common language. In the common situation both parties would know the meaning of “Gavagai.” Yet, however effective it is, ordinary speech sometimes fails to secure understanding between people and, where such understanding is achieved, that achievement has required back-and-forth speech. The mutual understanding is achieved through a process of negotiation. As William Croft reiterates in chapter 4 of Explaining Language Change, we cannot get inside one another’s heads and so must negotiate meanings in conversation.

That is to say, communication through language is not a matter of sending information through a pipeline. It does not happen according to what Michael Reddy (1993) has called the conduit metaphor. Reddy’s article is based on 53 example sentences. Here are the first three (p. 166):
1. Try to get your thoughts across better
2. None of Mary’s feelings came through to me with any clarity
3. You still haven’t given me any idea of what you mean
Reddy’s argument is that many of our statements about communication seemed to be based on the notion of sending something (the thought, idea, feeling) through a conduit, hence he calls it the conduit metaphor. He knows that communication doesn’t work that way, but that’s not is central issue. His central concern is to detail the way we use the conduit metaphor to structure our thinking about communication.

Reddy’s argument is reminiscent of a somewhat earlier argument by Paul de Man, “Form and Intent in the American New Criticism” (1983, first published in 1971). Consider this passage (p. 25):
“Intent” is seen, by analogy with a physical model, as a transfer of a psychic or mental content that exists in the mind of the poet to the mind of a reader, somewhat as one would pour wine from a jar into a glass. A certain content has to be transferred elsewhere, and the energy necessary to effect the transfer has to come from an outside source called intention.
De Man’s point was that, when we read a text, the intention (de Man uses the term in its somewhat rarified philosophical sense) that gives life to those signs on the page is our intention, not the author’s. And he is right.

De Man’s insight, and similar ones by Derrida, Barthes, Foucault and others, had an electrifying effect on literary critics in the United States, leading to a tremendously fertile period in academic literary criticism that, however, became increasingly sclerotic in the 1990s. But that story’s neither here nor there. My point is simply that these thinkers were attempting to deal with a real problem and, ultimately, they failed.

What, for example, could Derrida (1976, p. 158) have possibly meant by proclaiming “There is nothing outside of the text”? What he did not mean is that the world is nothing but a text and a text created by more or less arbitrary social conventions. Read sympathetically, and in context, the phrase seems to mean something to the effect that there is no way we can “step outside” language so as to examine, in full omniscient and transcendental objectivity, the relationship between language and the world. And that, it seems to me, is true. We’re always going to be immersed in “language,” whether natural or the various languages of science and mathematics.

How, then, do we fly free of the bottle? We play games.

Language Games and Game Theory
 
Where de Man argues that intent cannot be transmitted from one speaker to another like pouring wine from a jar, William Croft points out that linguistic communication is tricky “precisely because our thoughts cannot leave our heads” (2000, p. 111). Croft is a linguist who has undertaken to explain language change using an evolutionary approach. He defines a language to be “the population of utterances in a speech community” (p. 26), thus focusing our attention, not on some abstract language system, but on the concrete production of speech.

How does Croft deal with the fact that we cannot transmit thoughts directly to another’s mind? He argues that meaning is negotiated in the back-and-forth of conversation and draws on game theory to make his argument (p. 95):
There is a problem here: the hearer cannot read the speaker’s mind, and she can’t read his. This is what is called a COORDINATION PROBLEM. In speaking and understanding, speaker and hearer are trying to coordinate on the same meaning.
Croft then introduces the notion of a third-party Schelling game in which two players “are presented by a third party with a set of stimuli” which helps them converge on the same meaning. Sometimes it works, sometimes not. One possibility, he argues, is to use “natural perceptual or cognitive distinctiveness [as] a COORDINATION DEVICE” (p. 96). That gives us the adapted mind that I invoked in discussing Quine’s problem. Croft goes on to discuss a variety of linguistic devices as non-conventional coordination devices.

While the details are interesting and important – I recommend his discussion to you – we need not worry about them now.

Save one. Croft notes that, in order for speaker and hearing to reach agreement in conversation their mental states “need not be identical, though it is assumed that they are systematically related” (p. 99). Later on he notes that (114):
successful communication involves not the recovery of and original, ‘correct’ interpretation of the speaker’s original intention, but instead an interpretation that evolves over the course of the conversation, and is assessed by the success or failure of the higher social-interactional goals that the interlocutors are striving to achieve.
One reason why this effort is not doomed to failure from the beginning is the fact that although we cannot read each other’s minds, we do inhabit a shared world.

Croft’s general point, then, is simple, speech communication is a two-way interaction, not the one-way transmission of meaning, information, whatever, though a channel. De Man’s problem is thus solved for the case of face-to-face interaction, a common case, and surely the most basic one. Note that this solution does not involve recourse to a transcendental signified nor to stepping outside the text, nothing like that. It involves the ordinary and obvious means of interactive speech. In this sense, the key to the treasure, is the treasure. Nothing else is required.

But what, you may ask, of written communication, where direct interaction is not possible? After all, de Man was a literary critic, writing about the reading of written texts. What about that?

Good question. I’m going to punt on it. But I observe that some written communication – correspondence – does involve interaction, but at a slower pace than conversation, often much slower. In the case of literary texts, yes, readers cannot ordinary interact with authors, but they can interact with one another. I’ll say a little about that in the next post. Beyond that, yes, there are issues, serious issues. But this is not the place to address them. My concern here is just to get things started.
Note: Mathematician and psychologist Mark Changizi (1999) has an interesting argument about why vagueness of word meaning is essential to the proper functioning of language. His argument is grounded in considerations of computability and I recommend it to you. It makes an interesting complement to the game-theoretic conception of speaking.
Addition: See subsequent post reporting an experiment that David Hays did at RAND in the mid-1950s. It’s relevant to the game theoretic treatment of conversation.
What is a language and what are the memes? 
 
Now I want to shift gears a bit and work my way back to the physical “side” of the linguistic sign, because that’s where we’re going to go looking for memetic entities.

Throughout this post I’ve been assuming that we know what a language is. Now I want to get picky. Here’s what Sidney Lamb has to say in Pathways of the Brain. He’s talking about Roman Jakobson, the great linguist (p. 41):
Using the term language in a way it is commonly used . . . we could say that he spoke six languages quite fluently: Russian, Czech, German, English, Swediksh, and French, and he had varying amounts of skill in a number of others. But each of them except Russian was spoken with a thick accent. It was said of him that “He speaks six languages, all of them in Russian.” . . . the evidence indicates that from a neurocognitive point of view there is no such unit as a language. What exists from a neurocognitive point of view is not so much one linguistic system as a group of interconnected systems, relatively independent from one another.
Lamb goes on to assert that (p. 42):
Professor Jakobson’s internal linguistic information included a single phonological system, that of his native Russian, together with separate systems of grammar and lexicon for Russian, Czech, English, German, French, and Swedish – with some overlap in these grammars and lexicons . . . along with his more limited abilities in various additional languages; plus a conceptual system connected to them all.
So far we’ve been concerned with how meaning is negotiated, where meaning is a matter of the conceptual system. That’s on one “side” of the arbitrary sign, the side inside the brain. Now we’re going to look at the other “side” of the sign, the side that’s in public view, the physical sign. It’s that physical side that most differs among languages.

The question before us is: How do we conceptualize the memetic elements of language? In glossing the emic/etic distinction in a comment to John Wilkins I remarked that (now I’m simply repeating that comment) the distinction originates in linguistics, in the distinction between phonetics and phonemics. The former is about the psychophsics of speech sound while the latter is about phoneme systems. These are obviously very closely related matters, but they aren’t the same. We tend to perceive the speech stream as consisting of discrete sound entities, syllables and phonemes; this is the domain of phonemics. But the speech signal is, in fact, continuous. If you look at a sonogram of some chunk of speech, you don’t draw a series of vertical lines through it separating one phoneme from another; nor can you snip a tape recording into phoneme-long or syllable-long segments and reassemble it into something that sounds like natural speech. The aspects of the speech stream which are phonemically active differ from one language to another, which is why foreign languages all sound like “Greek.” Independently of the fact that you don’t know what the words mean or how the syntax works, you can’t even hear the phonemes in the speech stream.

Now, that’s the distinction I’m after, between phonemes and the raw speech stream. That’s the distinction I drew in my discussion of music (third post). Phonemes are those properties of the speech stream that are linguistically active. We need, however, to distinguish between segmental phonemes and suprasegmental phonemes. The segmental phonemes are roughly parallel to the letters of an alphabetic writing system. Suprasegmentals include tone, stress, and prosodic patterns. And then we need to consider ordering as well, as the order in which elements occur is certainly a property of the speech stream, and a most important one.

Before thinking about order, thought, we need to think a bit more about what’s going on. Roughly speaking, two things need to be extracted from the speech signal: 1) word identities (to be somehow linked to word meanings), and 2) the relations between the words (syntax). My quick take on matters – I’m not a linguist and I’ve not thought this through – is that both segmental and suprasegmental phonemes are involved in both of those processes. Relations between words are often indicated by word affixes, which are realized through segmental phonemes. Word identities are certainly realized by segmental phonemes, but tone and accent are involved as well.

Beyond this, relations between words are signaled by word order. In linguistic typology, typical word order is the primary trait on which classification based. Thus one has SVO languages (subject-verb-object), VSO languages (verb-subject-object), and so forth. As those designations suggest, word order indicates grammatical function, that is, relations between words.

Thus between word order and phonemes we’ve got a rich set of memetic elements. And we could also consider morphology in here as well. Taken together these aspects of the speech signal seem to be as memetically rich and abstract as the musical properties we looked at in discussing Rhythm Changes (first post).

Tuesday, June 15, 2021

How symbolic AI helped us transcend the word illusion

By “symbolic AI” I mean the research done from roughly the beginnings of AI in the mid-1950s though the 1970s and into the 1980s. In particular, I mean the work done on language. In this I include computational linguistics (CL), which was a different research tradition, and cognitive science more generally.

By “word illusion” I mean our ordinary understanding of words, an understanding fostered by dictionaries. In this understanding words are complex things having various properties, including pronunciation, spelling, grammatical usage, and meaning, sometimes two or more meanings. In this ordinary sense meaning is both reified (it’s there in the dictionary entry) and obscured in that it is often difficult, when one thinks about it, to distinguish between word meaning and the word reference. Thus in my post, The Word Illusion in Literary Criticism, I talk about how very difficult it was for me to, as an undergraduate learning about semiotics, to distinguish between signifier and signified. Oh, I understood, in principle, that a sign consisted of both signifier and signfied, and that the signified was not the “thing” to which a word referred. But that understanding was brittle. The difference between word (sign) and thing was easy enough, but the signified was obscure.

The problem, I maintain, was not specific to me. No, it was a general one. And – here’s my point – I don’t believe we had a flexible and usable way of distinguishing the signifier as a distinct object until the cognitive sciences, AI and CL in particular, explored specific formal and quasi-formal proposals about the structure and mechanisms in the domain of signifieds. I would argue that, in a sense, it was only then that the world of signifieds emerged from the shadows and became fully REAL.

What, you might ask, about symbolic logic? Symbolic logic and the predicate calculus were never intended as an account of the meaning of natural language, though learning to produce logical formulae corresponding to specific sentences belongs to the discipline of logic. Logic, if I’m not mistaken, was introduced as something superior to, because more precise, than natural language. And, yes, I know that formal logic was central to symbolic AI, but, I note: 1) such logics had to be modified and extended in order to do the work, and 2) it wasn’t until AI and CL that we actually had to produce micro-worlds of interlinked propositions representing some aspect of the conceptual universe. Wittgenstein may have imagined such a body of interlinked propositions in his Tractatus Logico-Philosophicus, but he didn’t actually attempt to produce it. And, as you know, he soon abandoned the Tractatus as a way of thinking about language.

No, it was work in AI and CL that made the domain of signifieds one which we could investigate. The fact that, for various reasons, those systems failed to produce robust practical tools should not be allowed to obscure that significant accomplishment. When current investigators in machine learning (ML) and artificial neural networks (ANN) denigrate those systems they risk throwing out the baby with the bath water. After all, MA and ANN achieve their results, which often have great practical value, at the expense of once more rendering the domain of signifieds obscure and intractable.

I don’t for a minute believe that we should attempt to implement symbolic mechanisms directly IN the models created by ML or ANN engines. No, our task is more subtle and difficult: How do we create ML or ANN engines that will, through their own mechanisms, arrive at symbolic computation where it is necessary, and only there? That’s a problem way beyond the scope of a single blog post. For some pointers about my current thinking, see my post, Geoffrey Hinton says deep learning will do everything. I’m not sure what he means, but I offer some pointers. Version 2.

Monday, May 31, 2021

Geoffrey Hinton says deep learning will do everything. I’m not sure what he means, but I offer some pointers. Version 2.

This is updated from a previous version to include a passage by Sydney Lamb.

* * * * *

Late last year Geoffrey Hinton had an interview with Karen Hao [1] in which he said “I do believe deep learning is going to be able to do everything,” with the qualification that “there’s going to have to be quite a few conceptual breakthroughs.” I’m trying to figure out whether or not, to what extent, in what way I (might) agree with him.

Neural Vectors, Symbols, Reasoning, and Understanding

Hinton believes that “What’s inside the brain is these big vectors of neural activity” and that one of the breakthroughs we need is “how you get big vectors of neural activity to implement things like reason.” That will certainly require a massive increase in scale. Thus while GPT-3 has 175 billion parameters, the brain has trillion, where Hinton treats each synapse as a parameter. 

Correspondingly Hinton rejects the idea that symbolic reasoning is primitive to the nervous system (my formulation), rather “we do internal operations on big vectors.” What about language? He doesn’t address the issue directly but he does say that “symbols just exist out there in the external world.” I do think that covers language, speech sounds, written words, gestural signs, those things are out there in the external world. But the brain uses “big vectors of neural activity” to process those. 

Hinton’s remark about symbols bears comparison with a remark by Sydney Lamb: “the linguistic system is a relational network and as such does not contain lexemes or any objects at all. Rather it is a system that can produce and receive such objects. Those objects are external to the system, not within it”[2]. Lamb has come to think of his approach as neurocognitive linguistics and, while his sense of the nervous system is somewhat different from Hinton’s, they agree on this issue and, in the current intellectual climate, that agreement is of some significance. For Lamb is a first generation researcher in machine translation  and so was working when most AI research was committed to symbolic systems. We’ll return to Lamb later as I think the notation he developed is a way to bring symbolic reasoning within range of Hinton’s “big vectors of neural activity”.

But now let’s return to Hinton with a passage from an article he co-authored with Yann LeCun and Yoshua Bengio [3]:

In the logic-inspired paradigm, an instance of a symbol is something for which the only property is that it is either identical or non-identical to other symbol instances. It has no internal structure that is relevant to its use; and to reason with symbols, they must be bound to the variables in judiciously chosen rules of inference. By contrast, neural networks just use big activity vectors, big weight matrices and scalar non-linearities to perform the type of fast ‘intuitive’ inference that underpins effortless commonsense reasoning.

I note, moreover, commonsense reasoning seems to be problematic for everyone.[4]

Let’s look at one more passage from the interview:

For things like GPT-3, which generates this wonderful text, it’s clear it must understand a lot to generate that text, but it’s not quite clear how much it understands.

I’m not sure that it is at all useful to say that GPT-3 understands anything. I think that, in using that term, Hinton is displaying what I’ve come to think of as the word illusion.[5] Briefly, GPT-3’s language model is constructed over a corpus consisting entirely of word forms, of signifiers without signifieds, to use an old terminology. Hinton knows that, of course, but, after all, he understands texts from seeing or hearing word forms alone, as do we all, and so, in effect, credits GPT-3 with somehow having induced meaning from a statistical distribution. The text it generates looks pretty good, no? Yes. And that is something we do need to understand, just what is GPT-3 doing and how does it do it? But this is not the place to enter into that.[6]

I think that GPT-3’s remarkable performance based on such ‘shallow’ material should prompt us into reconsidering just what humans are doing when we produce everyday ‘boilerplate’ text. Consider this passage from LeCun, Bengio, and Hinton, where they are referring to the use of an RNN:

This rather naive way of performing machine translation has quickly become competitive with the state-of-the-art, and this raises serious doubts about whether understanding a sentence requires anything like the internal symbolic expressions that are manipulated by using inference rules. It is more compatible with the view that everyday reasoning involves many simultaneous analogies that each contribute plausibility to a conclusion
.

In dealing with these utterly remarkable devices, we would be rein in our narcissistic investment in the routine use of our ‘higher’ cognitive and linguistic capacities as opposed to our mere sensory-motor competence. It’s all neural vectors. 

Note, however, that it is one thing to say that “we do internal operations on big vectors.” I agree with that. That’s not quite the same as saying we can do everything with deep learning. Deep learning is a collection of architectures, but I’m not sure such architectures are adequate for internalizing the vectors needed to effectively mimic human perceptual and cognitive behavior. The necessary conceptual breakthroughs will likely take us considerably beyond deep learning engines. With that qualification, let’s continue.

How the brain might be doing it

I find that, with the caveats I’ve mentioned, this is rather congenial. Which is to say that I can make sense of it in terms of issues I’ve thought through in my own work.

Some years ago David Hays and I wanted to come to terms with neuroscience and ended up reviewing a wide range of work and writing a paper entitled, “Principles and Development of Natural Intelligence.”[7] The principles are ordered such that principle N assumed N-1. We called the fifth and last principle indexing:

The indexing principle is about computational geometry, by which we mean the geometry, that is, the architecture (Pylyshyn, 1980) of computation rather than computing geometrical structures. While the other four principles can be construed as being principles of computation, only the indexing principle deals with computing in the sense it has had since the advent of the stored program digital computer. Indexed computation requires (1) an alphabet of symbols and (2) relations over places, where tokens of the alphabet exist at the various places in the system. The alphabet of symbols encodes the contents of the calculation while the relations over places, i.e. addresses, provide the means of manipulating alphabet tokens in carrying out the computation. [...] Within the context of natural intelligence, indexing is embodied in language. Linguists talk of duality of patterning (Hockett, 1960), the fact that language patterns both sounds and sense. The system which patterns sound is used to index the system which patterns sense.

In short, “indexing gives computational geometry, and language enables the system to operate on its own geometry.” This is where we get symbols and complex reasoning.

I should note that, while we talked of “an alphabet of symbols” and “relations over places” we were not asserting that that’s what was going on in the brain. That’s what’s actually going on in computers, but it applies only figuratively to the brain. The system that is using sound patterns to index patterns of sense is using one set of neural vectors (though we didn’t use that term) to index a different set of neural vectors.

How do we get deep learning to figure that out? I note that automatic image annotation is a step in that direction [8], but have nothing to say about that here.

Instead I want to mention some informal work I did some years ago on something I call attractor nets.[9] The general idea was to use Sydney Lamb’s relational networks, in which nodes are logical operators, as a tertium quid between the symbol-based semantic networks Hays and I had worked on in the 1970s and the attractor landscapes of Walter Freeman’s neurodynamics. I showed – informally, using diagrams – how using logical operators (AND, OR) over attractor basins in different neurofunctional areas could reconstruct symbolic systems represented as directed graphs. Each node in a symbolic graph corresponds to a basin of attraction, that is, an attractor. In the present context we can think of each neurofunctional area as corresponding to a collection of neural vectors and the attractors as objects represented by those vectors. An attractor net would then become a way of thinking about how complex reasoning could be accomplished with neural vectors.

In the attractor net notation word forms, or signifiers, are distinct from word meanings, of signifieds. Is that distinction important for complex reasoning? I believe it is, though I’m not interested in constructing an argument at this point. That, I believe, puts a limit on what one can expect of engines like GPT-3. That too requires an argument.

So, what about natural vs. artificial intelligence?

The notion of intelligence is somewhat problematic. As a practical matter I believe that a formulation by Robin Hanson is adequate: “’Intelligence’ just means an ability to do mental/calculation tasks, averaged over many tasks.”[10] As for the difference between artificial and natural, that comes down to four things:

1) a living system vs. an inanimate system,
2) a carbon-based organic electro-chemical substrate vs. a silicon-based electronic substrate,
3) real neurons (having on average 10K connections with others) vs. considerably simpler artificial neurons realized in program code, and
4) the neurofunctional architecture and innate capacities of a real brain vs. the system architecture of a digital computing system.

Make no mistake, those differences are considerable. But I think we now have in hand a body of concepts and models that is rich enough to support ever more sophisticated interaction between students of neuroscience and students of artificial intelligence. To the extent that our research and teaching institutions can support that interaction I expect to see progress accelerate in the future. I offer no predictions about what will come of this interaction.

Some related posts

William Benzon, Showdown at the AI Corral, or: What kinds of mental structures are constructible by current ML/neural-net methods? [& Miriam Yevick 1975], New Savanna, June 3, 2020, https://new-savanna.blogspot.com/2020/06/showdown-at-ai-corral-or-what-kinds-of.html.

William Benzon, What’s AI? – Part 2, on the contrasting natures of symbolic and statistical semantics [can GPT-3 do this?], New Savanna, July 17, 2020, https://new-savanna.blogspot.com/2019/11/whats-ai-part-2-on-contrasting-natures.html.

William Benzon, A quick note on the ‘neural code’ [AI meets neuroscience], New Savanna, April 20, 2021, https://new-savanna.blogspot.com/2021/04/a-quick-note-on-neural-code-ai-meets.html.

References

[1] Interview with Karen Hao, AI pioneer Geoff Hinton: “Deep learning is going to be able to do everything”, MIT Technology Review, Nov. 3, 2020. https://www.technologyreview.com/2020/11/03/1011616/ai-godfather-geoffrey-hinton-deep-learning-will-do-everything/

[2] Sydney Lamb, Linguistic Structure: A Plausible Theory, Language Under Discussion, 4(1) 2016, 1–37, https://doi.org/10.31885/lud.4.1.229.

[3] From Yann LeCun, Yoshua Bengio, and Geoffrey Hinton, Deep Learning, Nature, 521 28 May 2015, 436-444, https://doi.org/10.1038/nature14539.

[4] As an example I offer a recent post in which I quiz GPT-3 about a Jerry Seinfeld bit: Analyze This! Screaming on the flat part of the roller coaster ride [Does GPT-3 get the joke?], May 7, 2021, https://new-savanna.blogspot.com/2021/05/analyze-this-screaming-on-flat-part-of.html.

[5] See my post, The Word Illusion, May 12, 2021, https://new-savanna.blogspot.com/2021/05/the-word-illusion.html.

[6] For some extended remarks, see my working paper, GPT-3: Waterloo or Rubicon? Here be Dragons, Working Paper, Version 3, August 20, 2020, 34 pp., https://www.academia.edu/43787279/GPT_3_Waterloo_or_Rubicon_Here_be_Dragons_Version_3.

[7] William Benzon and David Hays, Principles and Development of Natural Intelligence, Journal of Social and Biological Structures, Vol. 11, No. 8, July 1988, 293-322, https://www.academia.edu/235116/Principles_and_Development_of_Natural_Intelligence.

[8] Wikipedia, Automatic image annotation, https://en.wikipedia.org/wiki/Automatic_image_annotation.

[9] William Benzon, Attractor Nets, Series I: Notes Toward a New Theory of Mind, Logic, and Dynamics in Relational Networks, Working Paper, 52 pp., https://www.academia.edu/9012847/Attractor_Nets_Series_I_Notes_Toward_a_New_Theory_of_Mind_Logic_and_Dynamics_in_Relational_Networks.

William Benzon, Attractor Nets 2011: Diagrams for a New Theory of Mind, Working Paper, 55 pp., https://www.academia.edu/9012810/Attractor_Nets_2011_Diagrams_for_a_New_Theory_of_Mind.

William Benzon, From Associative Nets to the Fluid Mind, Working Paper. October 2013, 16 pp. https://www.academia.edu/9508938/From_Associative_Nets_to_the_Fluid_Mind.

[10] Robin Hanson, I Still Don’t Get Foom, Overcoming Bias, July 24, 2014, https://www.overcomingbias.com/2014/07/30855.html.

Sunday, May 23, 2021

Dictionaries, encyclopedias, and the word illusion

Something has been bugging me about my use of the term “word illusion.” Though I know how I mean to use it and why, it always seems a bit awkward and roundabout. I’ve figured out why.

It is our ordinary use and understanding that is (somewhat) illusory. When we see words on the page, we see and experience them as words-in-full, not as mere signifiers/word-forms. We don’t even experience words in a foreign language as mere signifiers/word-forms. Rather, we experience them as words we don’t understand. They aren’t word-forms, they’re merely words.

And the same is even more true for spoken language. We hear and understand the meanings. The word-forms come before consciousness only if something goes wrong. And then we experience them as words-gone-wrong.

When I talk of the word illusion, then, I’m talking of taking our ordinary understanding of and experience of words into a domain where that understanding can get us into trouble. In the case of literary criticism it leads us to interpret all over the damned place. In NLP it leads us to misinterpret or over interpret what the machine is doing.

And I suppose why I was a bit taken aback when I read Michael Gavin’s fascinating article, “Vector Semantics, William Empson, and the Study of Ambiguity.”[1] For he also referred to vector semantics as computational semantics, to which my immediate response was, “No no no, that’s not computational semantics. Computational semantics is different.” By that I meant that computational semantics (proper) is what I’d studied in graduate school, Old School symbolic constitutive semantics. Vector semantics is fine, but it is a very different beast from, well, constitutive semantics – a term I coined to contrast with vector semantics. People need to know the difference.

But enough of that, which is in a way a diversion. What set me off this morning is thinking back to the days when I was fascinated by dictionaries and encyclopedias. I’d guess it was middle school.

I’d look up a word in the dictionary, read its definition, and still be a bit puzzled. I found the usage examples particularly useful – that’s why they’re there, no? Often as not the definition itself would contain a word or two I didn’t understand, so I’d go look them up. And the same thing would happen. And so I’d happily loop my way through the dictionary reading word after word after word. And that’s certainly how I understood and experienced those words, as words-in-full, even though each dictionary and information about pronunciation and possibly alternative spellings. They were simply other attributes of the word.

I’d do the same thing with the encyclopedia. Of course, I didn’t loop through those entries quite so rapidly. And no doubt I went back and forth between the encyclopedia and the dictionary.

This was several years before I went to college and tried to puzzle out signifier, signified, and thing or referent. How would I have reacted to that three-way distinction back then when I was looping through the dictionary and the encyclopedia?

Reference

[1] Michael Gavin, Vector Semantics, William Empson, and the Study of Ambiguity, Critical Inquiry 44 (Summer 2018) 641-673: https://www.journals.uchicago.edu/doi/abs/10.1086/698174 Ungated PDF: http://modelingliteraryhistory.org/wp-content/uploads/2018/06/Gavin-2018-Vector-Semantics.pdf.

Tuesday, May 18, 2021

The Word Illusion in Literary Criticism

We all know that each word has a spelling, a pronunciation, and one or more meanings. Different words may have the same meaning. Spellings and pronunciations are physical things, but meanings are not. Meanings are elusive.

When I refer to the word illusion I mean to indicate difficulties that some specialized disciplines encounter when dealing with word meanings.The illusory quality results from the mistaken idea/intuition that, because you know what words mean, what this that or the other word means, you are in a position, in effect, to think about the semantic underpinnings of meaning. This post is about problems that literary criticism has as a consequence of the word illusion.

On learning the distinction between form and content

It is one thing to be told, and to believe, that words have form, spelling and/or pronunciation, and meaning. That is simple. But how you think about language and texts depends on the conceptual equipment you have for dealing with meaning.

I don’t recall whether or not the distinction between form and content was even mentioned in literature courses I took in my freshman year at college. But in my sophomore year I read Roland Barthes’ Elements of Semiology. There I learned that the sign consisted of a signifier (form) and signified (content). I found the difference between (external) thing, or referent, and (internal) signified difficult to grasp.

Here is how Barthes talked of the signified:

II.2.1. Nature of the signified: In linguistics, the nature of the signified has given rise to discussions which have centred chiefly on its degree of ‘reality’; all agree, however, on emphasising the fact that the signified is not ‘a thing’ but a mental representation of the ‘thing’. We have seen that in the definition of the sign by Wallon, this representative character was a relevant feature of the sign and the symbol (as opposed to the index and the signal). Saussure himself has clearly marked the mental nature of the signified by calling it a concept: the signified of the word ox is not the animal ox, but its mental image (this will prove important in the subsequent discussion on the nature of the sign). These discussions, however, still bear the stamp of psychologism, so the analysis of the Stoics will perhaps be thought preferable. They carefully distinguished the phantasia logiki (the mental representation), the tinganon (the real thing) and the lekton (the utterable). The signified is neither the phantasia nor the tinganon but rather the lekton; being neither an act of consciousness, nor a real thing, it can be defined only within the signifying process, in a quasi-tautological way: it is this 'something' which is meant by the person who uses the sign. In this way we are back again to a purely functional definition: the signified is one of the two relata of the sign; the only difference which opposes it to the signified is that the latter is a mediator.

Notice that Barthes begins by noting that the nature of this signified, whatever it is, is problematic. It is at least clear that the signified is to be distinguished from the thing, the referent. Beyond that, alas, I find it opaque. I have no idea what I would have made of it back then.

I am quite sure, however, that I had a rather clear sense of the distinction between signifier and thing, or reference, six years later in my second year of graduate school, when I’d undertaken the study of computational semantics with David Hays. Hays had worked out a complete linguistic system, which he’d sketched out in an unpublished book entitled Mechanisms of Language, from phonetics and phonology, though morphology and syntax, to semantics and discourse. He used diagrams, network diagrams, to indicate this system. In those diagrams semantic objects where clearly distinct from syntactic and morphological ones. There could be no doubt that signifiers (morphological objects) are distinctly different from signifieds (semantic objects), and that all those linguistic objects are distinct from the various things, referents, language can be used to talk about. Those diagrams made these matters very clear.

I have no recollection of what I made of the distinction between signifier and thing between my sophomore year and graduate school. In the second half of that period, however, I did read about semantics in the nascent cognitive sciences and that must have given me some way of dealing with that distinction. But the discipline of literary criticism never made it beyond that obscure paragraph by Barthes, that and others like it. Yes, literary critics certainly understand that there is a distinction between, say, the concept of an apple, which is in one’s head, and a physical apple out there in the world. But that understanding doesn’t give you much to work with when dealing with more complicated cases, and every literary text is a more complicated case.

The critics and the text

In the late 1930s and on into the 1950s a group of critics known as the New Critics argued that literary texts were autonomous. By that they meant that you need not, indeed you should not, refer either to socio-historical context or to the author’s life in interpreting a literary work. What made texts autonomous? Their form. How did form do that? I do not at this point know what the New Critics thought about that, but I’m sure they had reasons and those reasons no doubt became more sophisticated as literary criticism grew in philosophical sophistication after World War II.

Beyond the general idea of interpretation the New Critics had little intellectual equipment for dealing with meaning, the world of signifiers, of verbal content. That meant that in practice critics could wander far and wide as long as they avoided attempting to ground their interpretations in the author’s life and/or historical context. They were interpreting ‘universal’ meanings.

One was free to track down influences in earlier texts that the author might have read. Even if an author hadn’t read them, there were there floating around in the tradition. These influences ranged all the way back to the ancient Greeks and Hebrews. Biblical and classical references and symbols gave one leeway to traipse far and wide in discovering meanings somehow IN the text, in the text in the sense that they aren’t in the biographical and historical details of the author’s life and times. Beyond that, Freud and Jung afforded the critic other means of interpretation. Such ‘depth’ criticism certainly wasn’t authorized by the New Critics, but it did afford ways of interpreting texts that avoided context and biography, and so of finding meanings IN the texts.

It was fascinating stuff. I do remember, if only vaguely, that, as I read such criticism, I was puzzled about just how and in what sense these meanings were actually there in the text. I certainly didn’t think of them when I read the text. In what sense does one have to know such things in order properly to have read the text?

What is reading and where’s the text?

One consequence of this methodology is was to elide the distinction between reading, as the word is ordinarily used, and reading, as critics offer and explicit written interpretation of the text. Reading is reading. They are the same thing, reading. I recall being puzzled about that elision as it happened. A study of topics in seven mainstream literary journals suggests that the elision happened in the 1960s and 1970s (see my post Meaning, Theory, and the Disciplines of Criticism).

Another consequence is that the concept of the text became deeply problematic. That a text is a string of signifiers, of word forms, is obvious, and trivial. This is an idea/observation of little use to a literary critic. The critic wants to know, what is this thing I am interpreting? It’s not those dumb symbols. It’s something they represent, evoke, constitute, but what?

In a well-known article, for example,  Roland Barthes’ tells us (“From Work to Text, in The Rustle of Language, 1986):

1. The text must not be understood as a computable object. It would be futile to attempt a material separation of works from texts. In particular, we must not permit ourselves to say: the work is classical, the text is avant-garde; there is no question of establishing a trophy in modernity's name and declaring certain literary productions in and out by reason of their chronological situation: there can be “Text” in a very old work, and many products of contemporary literature are not texts at all. The difference is as follows: the work is a fragment of substance, it occupies a portion of the spaces of books (for example, in a library). The Text is a methodological field. The opposition may recall (though not reproduce term for term) a distinction proposed by Lacan: “reality” is shown [se montre], the “real” is proved [se démontre]; in the same way, the work is seen (in bookstores, in card catalogues, on examination syllabuses), the text is demonstrated, is spoken according to certain rules (or against certain rules); the work is held in the hand, the text is held in language: it exists only when caught up in a discourse (or rather it is Text for the very reason that it knows itself to be so); the Text is not the decomposition of the work, it is the work which is the Text's imaginary tail. Or again: the Text is experienced only in an activity, in a production. It follows that the Text cannot stop (for example, at a library shelf); its constitutive moment is traversal (notably, it can traverse the work, several works).

Just what does that mean?

Note, however, that in the world I explored with David Hays the text really is a computable object, a concept which does not, however, diminish it. In that world, the text considered as a computable object, is a very rich a complex object. It is rich and complex because that world contains tools for constructing a rich and complex account of semantics.

Correlatively, the notion of form is as problematic as that of the text. Why? Because the New Critics and their heirs were never very interested in literary form as itself an object of analysis and description – though they did comment on such matters in poetry, where they are unavoidable. The idea of form provided philosophical cover for their declaration of textual autonomy. They were focused on meaning and so looked right through, as it were, the text’s formal features.

Thus the word illusion led literary critics to mystify the nature of the text, of form, and to conflate the acts of reading and interpretation. But – and this is important – without an explicit and tractable account of the world of signifieds, of meaning, it is not clear that critics could have done anything else, not if they wanted to study the meanings of literary texts. The current intellectual situation is quite different, but that’s well beyond the scope of a blog post.

Moving forward

Indeed, much of what I’ve written in New Savanna in the last decade about those current possibilities. Two somewhat different arenas have opened up.

At the level of the individual text, one can analyze and describe formal elements [1]. I suppose this possibility has always been there, but the existence of sophisticated approaches to computational semantics lends them a plausibility that they hadn’t had before. Not, mind you, that one can apply these semantic models to whole texts. I tried that and got as far as a Shakespeare sonnet, 129 [2]. But the existence of these models allows for the clear separation of signifier, linguistic content, from referent, ‘thing’ in the world and that separation “liberates” the text itself from the its interpretive subordination to the search for meaning.

At the level of a corpus of texts, we have a great and growing variety of work in computational criticism. The nature of these methods necessarily means that they are focused on the signifieds, the word forms, as that is all that is accessible to the computational techniques employed. It turns out that a sophisticated analysis of patterns of signifiers can be interpreted into the domain of meaning [3].

There is every hope that a revivified study of literary criticism can emerge from the shadow cast by the word illusion.

References

[1] For an explicit methodological justification of a computational approach to literary form, see my Literary Morphology: Nine Propositions in a Naturalist Theory of Form, PsyArt: An Online Journal for the Psychological Study of the Arts, August 2006, Article 060608, https://www.academia.edu/235110/Literary_Morphology_Nine_Propositions_in_a_Naturalist_Theory_of_Form

In recent years I have chosen to focus on ring-forms as a specific kind of formal design. In particular, see my working paper, Ring Composition: Some Notes on a Particular Literary Morphology, September 11, 2017, 71 pp. https://www.academia.edu/8529105/Ring_Composition_Some_Notes_on_a_Particular_Literary_Morphology

[2] William Benzon, Cognitive Networks and Literary Semantics, MLN 91: 1976, 952-982, https://www.academia.edu/235111/Cognitive_Networks_and_Literary_Semantics

William Benzon, Lust in Action: An Abstraction, Language and Style 14, 1981, 251-270, https://www.academia.edu/7931834/Lust_in_Action_An_Abstraction

 [3] Perhaps my best systematic statement on this is From Canon/Archive to a REAL REVOLUTION in literary studies, Working Paper, December 21, 2017, 26 pp., https://www.academia.edu/35486902/From_Canon_Archive_to_a_REAL_REVOLUTION_in_literary_studies.

Saturday, May 15, 2021

Ramble: Machine learning & the brain, word illusion, FOOM, Models of the Mind

Machine learning, the brain and the future of software

Friday’s post on Geoffrey Hinton’s assertion (deep learning is all we need) didn’t quite go the way I’d planned, but that’s OK. I figured relatively straightforward to point out the obvious limitations of that statement – obvious from my POV; it would take no more than four or five paragraphs. Then I got into it and realized that he was likely OK on the core assertion, but that it had implications he may not have appreciated. And once I’d worked through that I realized that I could interpret my old work on attractor nets more or less in Hinton’s terms and that, in turn, provides a way to think about complex thought in Hinton’s terms. That in turn led to the paper Hays and I did on natural intelligence. It all fits together. Making the connections took more time and effort than I had planned. But the end result is much more interesting.

That leaves me with more work to do. Well, maybe I’ll do it, maybe I won’t, at least not immediately. What sort of work? I’ve sketched out a framework, but what could we do within that framework in now and in the near future? That’s something I’d like to think about. Maybe it has implications for Andreessen’s AI-as-a-platform. I note that my post on that subject uses my old PowerPoint Assistant idea, and that is, in turn, based on some of the ideas behind attractor nets.

So much to do!

Word illusion

One problem Hinton as, that many have, is that he’s subject to the word illusion. In think about that I recalled how difficult it was for me to understand the difference between a word’s meaning and its referent. I think it was my sophomore year in college, when I read Roland Barthes’ Elements of Semiology, not much more than a pamphlet. I was easy enough to grasp that the sign consisted of a signifier ¬– something spoken or written – and well, what? I was used to thinking of words as pointing to things; that’s easy enough. But that the meaning, the signifier, is NOT the thing, the referent, I really had trouble absorbing that idea.

By the time I was in graduate school at SUNY Buffalo and working with David Hays on computational semantics, by that time the distinction between signifier and signified was abundantly clear. For the semantic network formalism was a very ‘concrete’ way of thinking about the realm of signifieds. You could draw pictures, or write (quasi)formal statements. But I didn't have anything like that at my disposal when I encountered Barthes. I don’t know just when the signifier/signified distinction became clear to me. Perhaps it really wasn’t clear until I began studying with Hays. But there must have been something before that, after all I’d been reading in cognitive science and knew about cognitive networks. I’d seen those diagrams (Ross Qullian, Don Norman, Sydney Lamb).

That is, it is one thing to assert the distinction, as Saussure did early in the 20th century. But how you understand that distinction depends on the conceptual tools available to you. And those tools weren’t readily available until, well, the so-called cognitive revolution. I suppose symbolic logic is a precursor, but the formalism is so detached from ordinary language that it doesn’t really do the job.

The thing about machine learning is that it doesn’t really force you to deal with the signifier/signified distinction. Sure, they are know that the texts in the corpora used by the machine are just word forms, signifiers. The meanings/signifieds aren’t there. But they don’t really have to work with the distinction. And when the resulting engine does interesting things with language, like vector semantics, machine translation, and so forth, well, it must somehow in some degree ‘understand’ the language it’s spitting out, no? The fact that it is not at all clear just what the machine is doing only reinforces the illusion. Since we don’t actually know, it’s easy enough to, in effect, believe that the machine has some how induced signifiers for the signifieds. That’s what those vectors are, no? Not, really, but...

You see the problem.

FOOM vs. the real singularity

I’m thinking of using that as the starting point for a 3 Quarks Daily piece. By FOOM, which is a term of art in some circles, I mean the idea that at some point in the (perhaps not too distant) future that the machines will suddenly become super-intelligent, bootstrap themselves to even greater levels of intelligence and then, FOOM! take over the world. That is, FOOM is a term for a very popular version of the so-called Technological Singularity.

I’m using “singularity” in the sense Jon von Neurmann used it, as a point “in the history of the race beyond which human affairs, as we know them, could not continue.” I think we’re already there. It is real. It is not a point event, but an ongoing and ill-defined process. Think of the cyberwar we’ve seen so far, including election inteference, ransomewhere (think of the current attack on the oil pipeline to the NYC region), and this that and the other. It’s not going away. Think about the fact that Facebook, Google, and Twitter have (potentially) more control over public discourse than any national governments. And so forth.

I was thinking of doing this for May 24, but I think I’ll push it off a month. I want to simmer it.

Models of the Mind

Instead I’m going to review Grace Lindsay, Models of the Mind: How Physics, Engineering, and Mathematics Have Shaped Our Understanding of the Brain (2021). It’s a very good general introduction to some of the various mathematical ideas that have been and are being used in understanding the brain. I know some of this stuff, and some I didn’t.

* * * * *

I’ll continue blogging about Seinfeld bits and about kids and music. And then we have the irises.

More later.

Friday, May 14, 2021

Geoffrey Hinton says deep learning will do everything. I’m not sure what he means, but I offer some pointers.

Superseded by Version 2 with an additional paragraph about Sydney Lamb.

* * * * *

Late last year Geoffrey Hinton had an interview with Karen Hao [1] in which he said “I do believe deep learning is going to be able to do everything,” with the qualification that “there’s going to have to be quite a few conceptual breakthroughs.” I’m trying to figure out whether or not, to what extent, in what way I (might) agree with him.

Neural Vectors, Symbols, Reasoning, and Understanding

Hinton believes that “What’s inside the brain is these big vectors of neural activity” and that one of the breakthroughs we need is “how you get big vectors of neural activity to implement things like reason.” That will certainly require a massive increase in scale. Thus while GPT-3 has 175 billion parameters, the brain has trillion, where Hinton treats each synapse as a parameter.

Correspondingly Hinton rejects the idea that symbolic reasoning is primitive to the nervous system (my formulation), rather “we do internal operations on big vectors.” What about language? He doesn’t address the issue directly but he does say that “symbols just exist out there in the external world.” I do think that covers language, speech sounds, written words, gestural signs, those things are out there in the external world. But the brain uses “big vectors of neural activity” to process those.

Let’s look at a passage from an article Hinton co-authored with Yann LeCun and Yoshua Bengio [2]:

In the logic-inspired paradigm, an instance of a symbol is something for which the only property is that it is either identical or non-identical to other symbol instances. It has no internal structure that is relevant to its use; and to reason with symbols, they must be bound to the variables in judiciously chosen rules of inference. By contrast, neural networks just use big activity vectors, big weight matrices and scalar non-linearities to perform the type of fast ‘intuitive’ inference that underpins effortless commonsense reasoning.

I note, however, commonsense reasoning seems to be problematic for everyone.[3]

Let’s look at one more passage from the interview:

For things like GPT-3, which generates this wonderful text, it’s clear it must understand a lot to generate that text, but it’s not quite clear how much it understands.

I’m not sure that it is at all useful to say that GPT-3 understands anything. I think that, in using that term, Hinton is displaying what I’ve come to think of as the word illusion.[4] Briefly, GPT-3’s language model is constructed over a corpus consisting entirely of word forms, of signifiers without signifieds, to use an old terminology. But Hinton knows that, of course, but, after all, he understands texts on the basis of word forms alone, as do we all, and so, in effect, credits GPT-3 with somehow having induced meaning from a statistical distribution. The text it generates looks pretty good, no? Yes. And that is something we do need to understand, just what is GPT-3 doing and how does it do it? But this is not the place to enter into that.[5]

I think that GPT-3’s remarkable performance based on such ‘shallow’ material should prompt us into reconsidering just what humans are doing when we produce everyday ‘boilerplate’ text. Consider this passage from LeCun, Bengio, and Hinton, where they are referring to the use of an RNN:

This rather naive way of performing machine translation has quickly become competitive with the state-of-the-art, and this raises serious doubts about whether understanding a sentence requires anything like the internal symbolic expressions that are manipulated by using inference rules. It is more compatible with the view that everyday reasoning involves many simultaneous analogies that each contribute plausibility to a conclusion
.

In dealing with these utterly remarkable devices, we would be rein in our narcissistic investment in the routine use of our ‘higher’ cognitive and linguistic capacities as opposed to our mere sensory-motor competence. It’s all neural vectors. 

Note, however, that it is one thing to say that “we do internal operations on big vectors.” I agree with that. That’s not quite the same as saying we can do everything with deep learning. Deep learning is a collection of architectures, but I’m not sure such architectures are adequate for internalizing the vectors needed to effectively mimic human perceptual and cognitive behavior. The necessary conceptual breakthroughs will likely take us considerably beyond deep learning engines. With that qualification, let’s continue.

How the brain might be doing it

I find that, with the caveats I’ve mentioned, this is rather congenial. Which is to say that I can make sense of it in terms of issues I’ve thought through in my own work.

Some years ago David Hays and I wanted to come to terms with neuroscience and ended up reviewing a wide range of work and writing a paper entitled, “Principles and Development of Natural Intelligence.”[6] The principles are ordered such that principle N assumed N-1. We called the fifth and last principle indexing:

The indexing principle is about computational geometry, by which we mean the geometry, that is, the architecture (Pylyshyn, 1980) of computation rather than computing geometrical structures. While the other four principles can be construed as being principles of computation, only the indexing principle deals with computing in the sense it has had since the advent of the stored program digital computer. Indexed computation requires (1) an alphabet of symbols and (2) relations over places, where tokens of the alphabet exist at the various places in the system. The alphabet of symbols encodes the contents of the calculation while the relations over places, i.e. addresses, provide the means of manipulating alphabet tokens in carrying out the computation. [...] Within the context of natural intelligence, indexing is embodied in language. Linguists talk of duality of patterning (Hockett, 1960), the fact that language patterns both sounds and sense. The system which patterns sound is used to index the system which patterns sense.

In short, “indexing gives computational geometry, and language enables the system to operate on its own geometry.” This is where we get symbols and complex reasoning.

I should note that, while we talked of “an alphabet of symbols” and “relations over places” we were not asserting that that’s what was going on in the brain. That’s what’s actually going on in computers, but it applies only figuratively to the brain. The system that is using sound patterns to index patterns of sense is using one set of neural vectors (though we didn’t use that term) to index a different set of neural vectors.

How do we get deep learning to figure that out? I note that automatic image annotation is a step in that direction [7], but have nothing to say about that here.

Instead I want to mention some informal work I did some years ago on something I call attractor nets.[8] The general idea was to use Sydney Lamb’s relational networks, in which nodes are logical operators, as a tertium quid between the symbol-based semantic networks Hays and I had worked on in the 1970s and the attractor landscapes of Walter Freeman’s neurodynamics. I showed – informally, using diagrams – how using logical operators (AND, OR) over attractor basins in different neurofunctional areas could reconstruct symbolic systems represented as directed graphs. Each node in a symbolic graph corresponds to a basin of attraction, that is, an attractor. In the present context we can think of each neurofunctional area as corresponding to a collection of neural vectors and the attractors as objects represented by those vectors. An attractor net would then become a way of thinking about how complex reasoning could be accomplished with neural vectors.

In the attractor net notation word forms, or signifiers, are distinct from word meanings, of signifieds. Is that distinction important for complex reasoning? I believe it is, though I’m not interested in constructing an argument at this point. That, I believe, puts a limit on what one can expect of engines like GPT-3. That too requires an argument.

So, what about natural vs. artificial intelligence?

The notion of intelligence is somewhat problematic. As a practical matter I believe that a formulation by Robin Hanson is adequate: “’Intelligence’ just means an ability to do mental/calculation tasks, averaged over many tasks.”[9] As for the difference between artificial and natural, that comes down to four things:

1) a living system vs. an inanimate system,
2) a carbon-based organic electro-chemical substrate vs. a silicon-based electronic substrate,
3) real neurons (having on average 10K connections with others) vs. considerably simpler artificial neurons realized in program code, and
4) the neurofunctional architecture and innate capacities of a real brain vs. the system architecture of a digital computing system.

Make no mistake, those differences are considerable. But I think we now have in hand a body of concepts and models that is rich enough to support ever more sophisticated interaction between students of neuroscience and students of artificial intelligence. To the extent that our research and teaching institutions can support that interaction I expect to see progress accelerate in the future. I offer no predictions about what will come of this interaction.

Some related posts

William Benzon, Showdown at the AI Corral, or: What kinds of mental structures are constructible by current ML/neural-net methods? [& Miriam Yevick 1975], New Savanna, June 3, 2020, https://new-savanna.blogspot.com/2020/06/showdown-at-ai-corral-or-what-kinds-of.html.

William Benzon, What’s AI? – Part 2, on the contrasting natures of symbolic and statistical semantics [can GPT-3 do this?], New Savanna, July 17, 2020, https://new-savanna.blogspot.com/2019/11/whats-ai-part-2-on-contrasting-natures.html.

William Benzon, A quick note on the ‘neural code’ [AI meets neuroscience], New Savanna, April 20, 2021, https://new-savanna.blogspot.com/2021/04/a-quick-note-on-neural-code-ai-meets.html.

References

[1] Interview with Karen Hao, AI pioneer Geoff Hinton: “Deep learning is going to be able to do everything”, MIT Technology Review, Nov. 3, 2020. https://www.technologyreview.com/2020/11/03/1011616/ai-godfather-geoffrey-hinton-deep-learning-will-do-everything/.

[2] From Yann LeCun, Yoshua Bengio, and Geoffrey Hinton, Deep Learning, Nature, 521 28 May 2015, 436-444, https://doi.org/10.1038/nature14539.

[3] As an example I offer a recent post in which I quiz GPT-3 about a Jerry Seinfeld bit: Analyze This! Screaming on the flat part of the roller coaster ride [Does GPT-3 get the joke?], May 7, 2021, https://new-savanna.blogspot.com/2021/05/analyze-this-screaming-on-flat-part-of.html.

[4] See my post, The Word Illusion, May 12, 2021, https://new-savanna.blogspot.com/2021/05/the-word-illusion.html.

[5] For some extended remarks, see my working paper, GPT-3: Waterloo or Rubicon? Here be Dragons, Working Paper, Version 2, August 20, 2020, 34 pp., https://www.academia.edu/43787279/GPT_3_Waterloo_or_Rubicon_Here_be_Dragons_Version_2.

[6] William Benzon and David Hays, Principles and Development of Natural Intelligence, Journal of Social and Biological Structures, Vol. 11, No. 8, July 1988, 293-322, https://www.academia.edu/235116/Principles_and_Development_of_Natural_Intelligence.

[7] Wikipedia, Automatic image annotation, https://en.wikipedia.org/wiki/Automatic_image_annotation.

[8] William Benzon, Attractor Nets, Series I: Notes Toward a New Theory of Mind, Logic, and Dynamics in Relational Networks, Working Paper, 52 pp., https://www.academia.edu/9012847/Attractor_Nets_Series_I_Notes_Toward_a_New_Theory_of_Mind_Logic_and_Dynamics_in_Relational_Networks.

William Benzon, Attractor Nets 2011: Diagrams for a New Theory of Mind, Working Paper, 55 pp., https://www.academia.edu/9012810/Attractor_Nets_2011_Diagrams_for_a_New_Theory_of_Mind.

William Benzon, From Associative Nets to the Fluid Mind, Working Paper. October 2013, 16 pp. https://www.academia.edu/9508938/From_Associative_Nets_to_the_Fluid_Mind.

[9] Robin Hanson, I Still Don’t Get Foom, Overcoming Bias, July 24, 2014, https://www.overcomingbias.com/2014/07/30855.html.