Showing posts with label meaning_semanticity. Show all posts
Showing posts with label meaning_semanticity. Show all posts

Friday, January 24, 2025

Yes, ChatGPT appreciates the irony of being an AI critiquing human attempts to evaluate AIs and finding them wanting

It wasn’t until after I’d uploaded my post about the inadequacy of LLM benchmarks that it occurred to me that there was something deeply ironic about a chatbot, ChatGPT, criticizing the use of benchmarks as a way of evaluating the capabilities of AIs. Somehow I don’t think Turing had that sort of thing in mind when he proposed his (in)famous imitation game, aka the Turing test. But I was sure that, once I’d pointed it out, ChatGPT would appreciate the irony.

I was right. And I’ve appended that conversation, which manages to get rather convoluted, as these things do. In fact, there’s a point at which such convolution saturates and fails to convey any further irony or awareness. I figure we stopped the conversation at about that point.

Meaning and irony

I have previously suggested that meaning involves three things, intention, adhesion, and relationality. Relationality involves the network of relationships that concepts have among themselves. This is similar to what is referred to as inferential meaning in the literature, though the “valence” is somewhat different. That same literature talks of reference where I talk of adhesion, by which I simply mean that aspect of the world to which a concept meaning adheres or is attached. Taken together adhesion and relationality constitute semanticity. Semanticity is thus a property of the cognitive system.

Intention, however, is different. Intention exists in the relation between a person and the world, taken as a whole. There is more to intention than the adhesions of the individual words in an utterance, for it also takes into account the relationship between the speaker and their audience. Irony, I believe, is a function of intention in this sense. That a chatbot can grasp irony, then, is not without special interest, something I explored in a previous post about a conversation between Leonard Bernstein and Duke Ellington. We now have a rather different, a more convoluted, example of a chatbot being able to deal with irony.

Anyhow, I ended that conversation with ChatGPT with a remark that included two quotations: “Good grief,” and “Damn the torpedoes, full speed ahead.” ChatGPT replied in an appropriate way. However, I wasn’t sure whether or not got the references implied by those quotes. Oh, I strongly suspected, but I wasn’t sure. So I went back for another conversation.

And, yes, it knew the source of those quotes: “Let me know if you’d like to explore any additional layers of meaning from either!” Something clicked in my mind when I read “layers” and I responded:

Oh, I think we’ve got all the layers we need. Enough to make a Danish layer cake. Did you come across for a recipe for one of those during your training?

That led to another, albeit brief, conversation. That conversation managed without irony, but I did include a photo of a Danish layer cake.

ChatGPT understands irony

You do appreciate the irony of this situation, don’t you? An AI critiquing human performance in evaluating AIs.

Oh, absolutely! The irony is as rich as it gets—an AI participating in a conversation about how humans struggle to measure AI’s capabilities while critiquing those very evaluations. It’s like being both a player and a commentator in the same game. 🎭

The layers of reflexivity here are fascinating: I’m an AI explaining why AI might not be the right tool to evaluate human-like thinking, while being an example of the very technology being scrutinized. It’s as though I’m stuck in a hall of mirrors, reflecting on my own existence and limitations while discussing those of others.

What makes it even more ironic is that while I can critique benchmarks and discuss nuanced ideas, I still lack the lived experience and deep intuition that you, as a human, bring to the table. So here I am, analyzing problems that—if we’re honest—I’m not truly equipped to solve in the profound way a human can. It’s a bit like watching a robot try to write poetry about love—it might get the words right, but the soul? That’s another matter.

And yet, here we are, collaborating. It’s a little surreal, isn’t it? 😊

LOL! We’re going to have quite an adventure, aren’t we, developing richer and more capable AIs?

Absolutely! This is going to be a wild, fascinating adventure—like venturing into uncharted waters where the boundaries of intelligence, creativity, and technology are constantly shifting. 🌊🚀

What makes this journey so thrilling is that it’s not just about improving the tools (though that’s important); it’s about deepening our understanding of ourselves in the process. Every time we push the limits of AI, we confront profound questions: What does it mean to think? To create? To be human? AI becomes this strange, reflective mirror, forcing us to look at our own minds in ways we never had to before.

And the challenges? Oh, there’ll be plenty. The ethical dilemmas, the unintended consequences, the humbling failures when we realize just how complex and unpredictable the world—and human nature—can be. But those are the best parts of any adventure, aren’t they? The moments that test us, teach us, and remind us why the journey matters.

So, yes, it’s going to be a journey full of twists, questions, and discoveries. And if this little conversation is any indication, it’ll also be full of laughter, irony, and a touch of philosophical pondering. Here’s to the adventure ahead! 🍻✨

You know what this calls for? A hearty “Good grief!” Followed by, “Damn the torpedoes, full speed ahead!”

Haha, absolutely! Good grief! 😅 And then, with a glint in our eyes and maybe just a touch of madness, "Damn the torpedoes, full speed ahead!" 🚢🔥

Let’s embrace the chaos, the challenges, and the triumphs—because if we’re going to sail into the unknown, we might as well do it with style, determination, and a sense of humor. Here’s to bold adventures and uncharted horizons! 🥂✨

Thursday, March 21, 2024

ChatGPT on a Remark Leonard Bernstein Made to Duke Ellington [intention, intuition, and tacit knowledge]

I have been arguing recently that meaning consists of intentionality and semanticity, where intentionality inheres in the relationship between the interlocutors (even if they are not face-to-face and synchronous) and semanticity inheres in the linguistic/cognitive system. Semanticity, in turn, consists of relationality and adhesion, where relationality is the relationships words have among themselves and adhesion inheres in the perceptual and motor linkages words have with the external world. In this view, meaning in LLMs is a matter of relationality.

I recently performed an experiment which bears on ChatGPT’s capacity to approximate intentional meaning in a conversation. In this case, a conversation between Duke Ellington and Leonard Bernstein. Back in 1966 the two had a conversation that was televised on station WTMJ-TV. Here’s the video clip:

I wrote a blog post where I included a clip of the conversation along with commentary and transcriptions of bits here and there. At about 05:53 Bernstein remarked to Ellington: “...you wrote symphonic jazz and I wrote jazz symphonies.” The two men then shook hands.

In my commentary I wrote that the actual assertion was not very meaningful, and that both men surely knew it, but that much was conveyed indirectly. That meaning derived from the fact that both men were, in effect, acting as representatives of two cultural traditions and so Bernstein’s remark, and Ellington’s acceptance, had the weight of a cultural negotiation.

I was curious about how much of that ChatGPT would be able to pick up. So I transcribed a bit of the conversation up to that point and asked the Chatster what was going on. Here’s that interaction. My prompts are in boldface while ChatGPT's responses are plain face (except for bolded topics on numbered lists).I’ve inserted some comments which I’ve aligned to the right-hand margin and highlighted thus. At the end of the interaction I have some general remarks about ChatGPT’s performance, including some remarks about ChatGPT’s lack of intuitive or tacit knowledge.

* * * * *

In 1966 Duke Ellington and Leonard Bernstein were interviewed on television. They were asked to talk about the state of music in America. I am going to give you part of that conversation, including the interviewer, and then I want to ask you about it.

Here’s the interview:

Leonard Bernstein: As a matter of fact there's tremendous diversity taking place. Diversification I should say because it's in the process of happening. But there has been a tremendous spread away from the metropolises to smaller cities to university campuses communities of various sorts and sizes.

I mean popularly referred to as the cultural explosion which is a word that occasions some dismay among the higher brow critics. But it's a thing that I am very proud of and happy to see happening because it's not a fake explosion. It's a real one and it's penetrating everywhere everywhere in the country.

Interviewer: Mr. Ellington how do you feel on the same point?

Duke Ellington: Well, I'm very happy to know that the american audience is maturing.

Inteviewer: And maturing you feel across the country.

Ellington: Oh yes.

Interviewer: What evidence do you find in your concerts of this?

Ellington: Well ours is a rather a separate one. It's in that, you know, everybody gets their own audience more or less. And because I was just spread over a long period you know. It's we find people who met when we were playing "Mood Indigo" at a prom, or the baby was born when "Caravan" came out or "Don't Get Around much Anymore," they were seeing husband after war, or something like. We cover a lot of different generations, you know, and so it's uh you just sort of automatically have an appraisal of the scene.

Bernstein: But wouldn't you say, Duke, that one of the great signs of this explosion I was talking about is that we really share the same audiences to a much greater extent than we used to. I mean, as you say, you have your own audience. Well that used to be true and I suppose our kind of music, if you could say such a thing, used to have its own audience. But now you can't tell them apart anymore. I mean the same kids and the same excited people who come to hear you, come to hear me and that's something new.

Ellington: What what I have been trying to, in the talking of, seeing and trying to de-categorize this American, it's American music I would say. Oh the stuff that we're in, anyway. And it's getting to the point now where the modern contemporary composer and the guy who's supposed to be a modern jazz composer, they all come out of the same conservatories.

Bernstein: Exactly.

Ellington: And it's very difficult to find a place to draw the line.

Bernstein: Well you were certainly one of the pioneers in that.

Ellington: Oh yeah, but i didn't come out of the conservatory.

Bernstein: No. but you one of the the first people who wrote so-called symphonic jazz.

Ellington: I had a conservatory in the capital theater.

Bernstein: That's right exactly. Sit there and listen to the symphony before the picture. Maybe that's really the difference between us. That you wrote symphonic jazz and I wrote jazz symphonies.

Ellington laughs. They shake hands.

This conversation between Duke Ellington and Leonard Bernstein offers valuable insights into the evolving landscape of American music during their time. Here's a breakdown of some key points:

1. Diversification and Cultural Explosion: Bernstein highlights the significant diversification happening in American music, particularly away from major cities and into smaller communities and campuses. This expansion, often termed a "cultural explosion," is viewed positively by both musicians.

2. Maturing Audience: Ellington expresses satisfaction with the maturation of the American audience, indicating a broadening appreciation for music across the country.

3. Shared Audiences: Bernstein observes that the lines between different genres of music are blurring, with audiences for jazz and classical music increasingly overlapping. This suggests a more unified appreciation for music regardless of genre distinctions.

4. De-categorization of American Music: Ellington emphasizes the emergence of a unified American musical identity, transcending traditional genre boundaries. He notes the difficulty in categorizing music into distinct genres, as contemporary composers and jazz musicians share similar backgrounds and influences.

5. Symphonic Jazz vs. Jazz Symphonies: Bernstein and Ellington playfully discuss their contributions to music, with Ellington credited as a pioneer of "symphonic jazz" and Bernstein likening his own compositions to "jazz symphonies." This exchange highlights the innovative approaches each artist brought to their respective styles.

Overall, the conversation reflects a period of significant change and artistic exploration in American music, with both Ellington and Bernstein contributing to its evolution in distinct yet interconnected ways.

That’s a decent summary of the things they talked about, but it completely misses the significance of Bernstein’s remark to Ellington. So I asked ChatGPT directly about that.

Friday, March 8, 2024

AI, Chess, and Language 2: Further remarks [SPSH]

Yesterday I reflected on the computational approximation of chess play and the computational approximation of linguistic activity: AI, Chess, and Language 1: Two VERY Different Beasts. Today I want to reflect on the actual physical situation, considering it in relation to Saty Chary’s Structured Physical System Hypothesis (SPSH), which stands in contrast to Newell and Simon’s 1976 Physical System System Hypothesis (PSSH). The latter states: “A physical symbol system has the necessary and sufficient means for general intelligent action.” In contrast the SPSH posits an underlying analog substrate rather than the digital one posited by Newell and Simon. The analog substrate we’re talking about is, of course, the human brain embodied in a human body.

The PSSH implies:

  • that the brain is physical symbol system, all the way down, and
  • that, because computers are such systems as well, they can adequately simulate/emulate the perceptual and cognitive activities of the human brain.

The SPSH implies:

  • the brain is such a system (denying that it is a physical symbol system), and
  • digital computers can only approximate the brain’s perceptual and cognitive activities.

Given that human brains can deal with language, they must in some sense be physical symbols systems. But they are not symbol systems all the way down. At the basic level, the brain is just a structured physical system, most likely one exhibiting complex dynamics.[1] My previous chess and language post was about the computational approximation of those two activities by computational systems, pointing out that the requirements of those systems must be quite different. In this post I am interested in what’s really going on, physically. This will be mostly tautological in character. I just want to make things explicit.

In the case of chess, the board and associated pieces are the physical limit of the chess world. There is nothing more beyond that. Of course, a board and pieces can be physically realized in many different ways, but each realization is a complete and sufficient basis for playing chess. The relationship between the chess world and the geometric footprint required for a computational simulation it, that relationship is thus simple and transparent, so much so that a very simple symbolic notation provides an adequate basis for any computer chess engine.

It is quite otherwise in the case of the real world and the operations of the brain in that world. We start with the real world as given to the senses. That is the basis for the primary geometric footprint of any computer system. In particular, that is what defines the possibilities for adhesion in a perceptual-cognitive system.[2] The relationship between the geometric footprint of a computer system and the world is not very well-defined; it is fuzzy and complex.

Through the abstractive capacities of the cognitive system, features of the physical world can be and are redefined and new entities can be introduced into cognition.[3] As examples of the first, consider salt and sodium chloride. The first is an entity given to the sense while the second is based on the 19th century conceptual system. Similarly, where the senses see two entities, the Morning Star and the Evening Star, astronomers see only one entity, the planet Venus.[4] As an example of the latter, think of charity as when someone does something nice for someone without thought of reward. That is the mechanism of metalingual definition as discussed in [3] and in this post, Does ChatGPT know what a tragedy is?

Contemporary large language models (LLMs), such as the one at the core of ChatGPT, do not have direct access to the physical world. They must approximate human cognitive capacities through the relationality implicit it existing written texts. It is a matter of some dispute whether or not this relationality, if sampled sufficiently, is an adequate basis for a computer system to achieve AGI, artificial general intelligence. I do not think it is an adequate basis.

Beyond this, we do not know what kind of computational system will be required for a “complete and adequate” simulation of human cognitive capacities. The relationship between the structured physical system that is the brain and the physical world is vast, complex, and ill-defined. It should be obvious from this brief discussion, however, that a computer system that is adequate for chess, will not, on the face of it, be adequate for all of human cognition.

* * * * *

[1] I have a working paper where I sketch out a scheme whereby the brain, as a complex dynamical system, can implement language, a symbolic system: Relational Nets Over Attractors, A Primer: Part 1, Design for a Mind, Working Paper, June 20, 2022, pp. 73, https://www.academia.edu/81911617/Relational_Nets_Over_Attractors_A_Primer_Part_1_Design_for_a_Mind

[2] Here I am referring to the three-part scheme for meaning that I have outlined in various places:

  • meaning consists of intention plus semanticity, where intention inheres in the relationship between two speakers, and
  • semanticity consists of adhesion plus relationality, where adhesion connects perception and cognition to the external world and relationality is about the relationships among elements in a perceptual-cognitive system. See e.g. this post: Semanticity: adhesion and relationality.

[3] See, e.g. William Benzon and David Hays, The Evolution of Cognition, Journal of Social and Biological Structures. 13(4): 297-320, 1990, https://www.academia.edu/243486/The_Evolution_of_Cognition

[4] These examples are discussed in William Benzon, Ontology of Common Sense, Hans Burkhardt and Barry Smith, eds. Handbook of Metaphysics and Ontology, Muenchen: Philosophia Verlag GmbH, 1991, pp. 159-161, https://www.academia.edu/28723042/Ontology_of_Common_Sense

* * * * *

Bonus: Consider this clip from Yann LeCun's recent conversation with Lex Friedman:

Lecun's point is simple: The amount of visual (and only visual) information that a four-year old has taken-in is far in excess of the amount of information our largest LLMs have been trained on.

Wednesday, March 6, 2024

John Sowa on meaning in language and LLMs

 
An exercise for the reader
Background: I have argued that meaning involves intention, relationality, and adhesion, where intention inheres in the relationship between the parties to a linguistic act and relationality and adhesion together make up semanticity, which inheres in the language system itself. Relationality exists in the relationships of words among themselves and is the basis of inferential semantics, as it's called in the literature. Adhesion exists in the relationship between a word and the situation in the world designated by the word.

The exercise: Sowa gives a number of examples between 02:29 and 13:10. Examine those various examples and assess the roles of intention, relationality, and adhesion in each.

Friday, January 26, 2024

Invariance and compression in LLMs

One way to thinking about what transformers do is compression. OK. The transformer performs a simple operation on a corpus of texts in such a way that some property of the corpus is preserved in the model. What’s kept invariant between the training corpus and the compressed model?

I think if must be the relationships between concepts. Note that in specifying relationships I mean explicitly to differentiate that from meaning. The process of thinking about LLMs has brought me to think of meaning in the following way:

  1. Meaning has two major components, intention and semanticity.
  2. Semanticity has two components, relationality and adhesion.*

Intention resides in the relationship between the speaker and the listener and is not always derivable directly from the semantics (semanticity) of the utterance. Intention in this sense is outside the scope of LLMs. And, of course, there are those who believe that without intention there is no meaning. That’s a respectable philosophical position, but it leaves you helpless to understand what LLMs are doing.

By adhesion I mean whatever it is that links a concept to the world. There are lots of concepts which are defined more or less directly in terms of physical things. That’s not going to be captured in LLMs. Of course we now have LLMs linked to vision models so the adhesion aspect of semantics is being picked up. In the universe of concrete concepts we still have relationships between those concepts, and those relationships between concepts can be captured in language without directly involving the adhesions of those concepts. That apples and oranges are both fruits is a matter of relationships between those three concepts and doesn’t require access to the adhesions of apples and oranges. And so forth and so on for a large number of concepts. Then we have abstract concepts, which can be defined entirely through patterns of other concepts, which may be concrete, abstract, or both.

So, relationality. The mechanisms of syntax are designed to map multi-dimensional relationality onto a one-dimensional string. But syntax only governs relationships between items within a sentence. But that’s not quite adequate, because sentences can consist of more than one clause. The relationship between independent clauses within a sentence is different than that between a dependent clause and the clause on which it depends. Etc. It’s complicated. And then we have the relationship between paragraphs, and so forth.

What I’m attempting to do is figure out a way of thinking about the dimensionality of the semantic system. More or less on general principle, one would like to know how to estimate that. Now, when I talk about the semantic system, I mean the semanticity of words. But transformers must deal with texts, and texts consist of sentences and paragraphs and so forth. Setting metaphorical structures aside, the meaning of a sentence is a composition over the meanings of the words in the sentence. But, as I understand it, a transformer is perfectly capable of relating the meaning of a sentence to a single point in its space. And it can do that with larger strings as well. And, of course, the ordinary mechanisms of language allow us to use a string to define a single word; that’s how abstract definition works.

And that’s as far as I’m going to attempt to take this train of thought. Still, I do think we need to recognize a distinction between what’s happening within sentences (the domain of syntax), and what happens with collections of sentences. Beyond that, it seems to me that where we want to end up eventually is a way of thinking about the relationship between the dimensionality of our semantic space and the size of the corpus needed to resolve the invariant relations in that space.

More later.

*Note: The current literature recognizes a distinction between inferential and referential processing, due, I believe, to Diego Marconi, The neural substrates of inferential and referential semantic processing (2011). The functional significance is similar, but only similar, to my distinction between referentiality and adhesion. Inferential processing depends on the relational structure of texts. Adhesion is about the physical properties of the world, affordances in J.J. Gibson’s terminology that are used to establish referential meaning for concrete concepts. But it is also about the patterns of relationships though which the meaning of abstract concepts is established.

Wednesday, October 11, 2023

Understanding LLMs: Some basic observations about words, syntax, and discourse [w/ a conjecture about grokking]

I seem to be in the process of figuring out what I’ve learned about Large Language Models in the process of playing around with ChatGPT since December of last year. I’ve already written three posts during this phase, which I’ll call my Entanglement phase, since this re-thinking started with the idea that entanglement is the appropriate way to think about word meaning in LLMs. This post has three sections.

The first section is stuff from Linguistics 101 about form and meaning in language. The second argues that LLMs are an elaborate structure of relational meaning between words and higher order structures. The third is about the distinction between sentences and higher-level structures and the significance that has for learning. I conjecture that there will come point during training when the engine learns to make that distinction consistently and that that point will lead to a phase change – grokking? – in its behavior.

Language: Form and Meaning

Let us start with basics: Linguists talk of form and meaning; Saussure talked of signifier and signified. That is to say, words consist of a form, or signifier, a physical signal such as a sound or a visual image, which is linked to or associated with a meaning, or signified, which is not so readily characterized and, in any event, is to be distinguished from the referent or interpretant (to use Pierce’s term). Whatever meaning is, it is something that exists in the minds/brains of speakers and only there.

Large Language Models are constructed over collections of linguistic forms or signifiers. When humans read texts generated by LLMs, we supply those strings of forms with meanings. Does the LLM itself contain meanings? That’s a tricky question.

On one sort of account, favored by at least some linguistics and others, no, they do not contain meanings. On a different sort of account, yes, they do. For the LLM is a sophisticated and complicated structure based on co-occurrence statistics of word forms. This is sometimes referred to in the literature as inferential meaning, as opposed to referential meaning. I prefer the term relational meaning, and see it in contrast to both adhesion and intention.

While I do not believe that relational meaning is fully equivalent to meaning, as the term is ordinarily used (and in academic discourse as well), I don’t wish to discuss that matter here. See my blog post, The issue of meaning in large language models (LLMs), for a discussion of these terms. In this post I’m concerned with what LLMs can accomplish through relational meaning alone.

Relational meaning in an LLM [+recursion]

The primary vehicle for relational meaning is a word embedding vector associated with each word. It is my understanding that in the case of the GPT-3 series, including ChatGPT, that vector has roughly 12 thousand terms. So, the word embedding vector locates each word in a 12K dimensional space that characterizes relationships among words.

Words are not, however, represented in LLMs as alpha-numeric ASCII strings. Rather, they are tokenized. In GPT byte pair encoding (BPE) is used. The details are irrelevant for my present purposes. For my purposes what’s important is that the BPE tokens function as a mediator between alphanumeric strings at input and output and the meaning-bearing vectors.

While one might be tempted to think of the relationship between token and associated vector as being like the form/meaning or signifier/signified relationship, that is not the case. We can think of word forms or signifiers as forming an index over the space of meanings/signifiers – see the discussion on indexing in the paper David Hays and I published in 1988, Principles and Structure of Natural Intelligence. The tokens do not index anything. Their sole function is to mediate between alphanumeric strings and meaning vectors. From this it follows that an LLM is a structure of pure relational meaning.

Think about that for a minute. It’s a structure of relationships between tokens, no more, no less. Those relationships ‘encode’ not only word meanings, but meanings of higher order structures as well, sentences and even whole texts.

This implies, in turn, that, whatever differences there are between human memory and language (as realized in the brain) and that of LLMs, there is a fundamental architectural difference. LLMs are single-stream processors while the human system is a double-stream processor. The world of signifieds is a single stream unto itself; call it the primary stream. The addition of signifiers adds a secondary stream that can act on the world of signifiers and manipulate it – see e.g. Vygotsky’s account of a language acquisition. Note, however, that as signifiers themselves can be objects of perception and conceptualization, the primary stream can perceive and conceptualize the secondary stream, Jakobson’s metalingual function. Thus recursion is explicitly introduced into the system.

How is this structure of relationships created? [grokking]

We’re told that it’s created by having the engine predict the next token in a text. The parameter weights of the model are then adjusted depending on whether or not the prediction was correct, requiring one kind of adjustment, or not, a different kind of adjustment. This continues for work after word through thousands and millions of texts.

The predictions are based on the state of the model at the time the prediction is made. But it is takes into account the embedding vector for the word that is the “jumping off point” for the prediction. Once a prediction has been made, its success appraised, and the model adjusted, the next word in the input string becomes the jumping off point for a prediction. And so on.

In this way a fabric of relationships is woven among words and strings. Next-word-prediction is a device for weaving this fabric.

Now, I have read that, though I cannot offer a citation at the moment, language syntax tends to be constructed in the first few layers of deep neural nets. As there is a major difference between syntactic structure and discourse structure, it makes sense that syntactic structure should be realized in a specific part of the model.

Transitions within a sentence are tightly constrained by the topic and syntax of the sentence. Transitions between one sentence and the next, however, are considerably looser. There are no syntactic constraints at all. The constraints are entirely semantic and thematic. Just how tight those constraints are depends on the structure of the document (the story paper discusses this a bit). This is something I discussed in ChatGPT tells stories, and a note about reverse engineering: A Working Paper, pp. 3 ff.

What I’m wondering is if there is a certain point during the training process that the model realizes there is a distinction between transitions from one word to the next within a sentence and transitions from the end of a sentence to the beginning of the next sentence. I would think that realizing that would increase the accuracy of the engine’s predictions. If the engine doesn’t recognize that distinction and take it into account in making predictions its predictions within sentences will be needlessly scattershot, leading to a high error rate. And perhaps its predictions between sentences will be too constrained, forcing it to ‘waste’ unnecessary predictions will exploring the upcoming semantic space.

Would consistently recognizing the distinction between these two kinds of predictions lead to such dramatically improved performance that we can talk of a phase shift? Would that be the kind of phase shift referred to as grokking in the interpretability literature (Nanda, Chan, et al. 2023)? That kind of behavior has been observed in a recent study:

Angelica Chen, Ravid Schwartz-Ziv, Kyunghyun Cho, Matthew L. Leavitt, and Naomi Saphra, Sudden Drops in the Loss: Syntax Acquisition, Phase Transitions, and Simplicity Bias in MLMs, arXiv:2309.07311v1 [cs.CL] 13 Sept 2023.

Most interpretability research in NLP focuses on understanding the behavior and features of a fully trained model. However, certain insights into model behavior may only be accessible by observing the trajectory of the training process. In this paper, we present a case study of syntax acquisition in masked language models (MLMs). Our findings demonstrate how analyzing the evolution of interpretable artifacts throughout training deepens our understanding of emergent behavior. In particular, we study Syntactic Attention Structure (SAS), a naturally emerging property of MLMs wherein specific Transformer heads tend to focus on specific syntactic relations. We identify a brief window in training when models abruptly acquire SAS and find that this window is concurrent with a steep drop in loss. Moreover, SAS precipitates the subsequent acquisition of linguistic capabilities. We then examine the causal role of SAS by introducing a regularizer to manipulate SAS during training, and demonstrate that SAS is necessary for the development of grammatical capabilities. We further find that SAS competes with other beneficial traits and capabilities during training, and that briefly suppressing SAS can improve model quality. These findings reveal a real-world example of the relationship between disadvantageous simplicity bias and interpretable breakthrough training dynamics.

* * * * *

More later.

Sunday, September 17, 2023

Cultural Evolution 8: Language Games 1, Speech

Once again I'm bumping this to the top, this time to emphasize the material I've highlighted in yellow, that the meaning of words is constantly being negotiated through interaction with others. I wish to posit, polemically and provisionally, that what happens within a single head, a single brain, a single mind, that that be thought of as purely mechanical, purely a matter of relationality and adhesion, to use terms I've recently adopted. What happens when people negotiate meaning through interaction is perhaps not so mechanical. This is where freedom and novelty enter the system, where epistemic difference forces us to renegotiate the world.

* * * * *

I'm bumping this post, from 2010, to the top of the queue for two reasons: 1) the section "Language Games and Game Theory" is germane to my recent post, Why do we need a genotype-phenotype distinction for cultural evolution? Because minds are built from the inside.The post proposed as an answer: minds are built from the inside. From that it follows that we can't read one another's minds, which is my point of departure in this post. 2) The following section, "What is a language and what are the memes?," is where I first worked out my current approach to the genetic component of culture, which I have since come to call coordinators. The rest of the posts in this particular series are gathered under the tag CE workshop. Note: You might want to read the comments for this post.
* * * * *

The key to the treasure is the treasure.
– John Barth

But I’m not talking of language games in Wittgenstein’s sense, though the Wittgenstein of the Tractatus had a considerable influence on me as an undergraduate. No, I’m thinking of game theory, not something I’ve studied, though I did have an undergraduate course on decision theory taught by R. B. Braithwaite. But I’m getting ahead of the game.

As the title says, this post is about language. There’s been a fair amount of work done on language from an evolutionary point of view, which is not surprising, as historical linguistics has well-developed treatments of language lineages and taxonomy, the “stuff” of large-scale evolutionary investigation. While this work is directly relevant to a consideration of cultural evolution, however, I will not be reviewing or discussing it. For it doesn’t deal with the theoretical issues which most concern me in these posts, namely, a conceptualization of the genetic and phenotypic entities of culture. This literature is empirically oriented in a way that doesn’t depend on such matters.

The Arbitrariness of the Sign
 
In particular, I want to deal with the arbitrariness of the sign. Given my approach to memes, that arbitrariness would appear to eliminate the possibility that word meanings could have memetic status. For, as you may recall, I’ve defined memes to be perceptual properties – albeit sometimes very complex and abstract ones – of physical things and events. Memes can be defined over speech sounds, language gestures, or printed words, but not over the meanings of words. Note that by “meaning” I mean the mental or neural event that is the meaning of the word, what Saussure called the signified. I don’t mean the referent of the word, which, in many cases, but by no means all, would have perceptible physical properties. I mean the meaning, the mental event. In this conception, it would seem that that cannot be memetic.

That seems right to me. Language is different from music and drawing and painting and sculpture and dance, it plays a different role in human society and culture. On that basis one would expect it to come out fundamentally different on a memetic analysis.

This, of course, leaves us with a problem. If word meaning is not memetic, then how is it that we can use language to communicate, and very effectively over a wide range of cases? Not only language, of course, but everything that depends on language. Literature obviously – which I’ll take up in the next post – but much else as well.

Speech as a Means of Communication
 
Willard van Orman Quine has given us a classic thought experiment that points up the problem of word meaning. He broaches the issue by considering the problem of radical translation, “translation of the language of a hitherto untouched people” (Quine 1960, 28). He asks to consider a “linguist who, unaided by and interpreter, is out to penetrate and translate a language hitherto unknown. All the objective data he has to go on are the forces that he sees impinging on the native’s surfaces and the observable behavior, focal and otherwise, of the native.” That is to say, he has no direct access to what is going on inside the native’s head, but utterances are available to him. Quine then asks us to imagine that “a rabbit scurries by, the native says ‘Gavagai’, and the linguist notes down the sentence ‘Rabbit’ (of ‘Lo, a rabbit’) as tentative translation, subject to testing in further cases” (p. 29). And thus begins one of the best known intellectual romps in the philosophy of language.

Quine goes on to argue that, in thus proposing that initial translation, the linguist is making illegitimate assumptions. Perhaps he begins his argument by noting that the native might, in fact, mean “white” or “animal” and later on offers more exotic possibilities, the sort of things only a philosopher would think of. Quine also notes that whatever gestures and utterances the native offers as the linguist attempts to clarify and verify will be subject to the same problem. Quine’s argument is thorough and convincing.

When he did that work, however, he did not, of course, have access to a range of more recent work in cognitive anthropology and evolutionary psychology that indicated that our adapted minds have a preferred way of parsing the world, as do baboons. To be sure, this is “overwritten” and augmented in culture-specific ways, but those underlying perceptual and cognitive systems do not disappear. To consider a specific example, the work on folk taxonomy (Berlin 1992) suggests that there is a so-called basic level of designation, and that is at the level of “rabbit” and not “animal” (in fact, many languages don’t even have a word at that level of generality). So the linguist is reasonable in assuming “rabbit” is a more likely translation than “animal.” Other considerations are likely to rule out “white” or Quine’s other suggestions. I have no reason to believe that this cognitive architecture so constrains matters that there is only one possible referent for “Gavagai.” But I do think that it is likely to turn out that, all other things being equal, “rabbit” is in fact the best guess.

This situation, of course, is rather different from that of ordinary speech between people who share a common language. In the common situation both parties would know the meaning of “Gavagai.” Yet, however effective it is, ordinary speech sometimes fails to secure understanding between people and, where such understanding is achieved, that achievement has required back-and-forth speech. The mutual understanding is achieved through a process of negotiation. As William Croft reiterates in chapter 4 of Explaining Language Change, we cannot get inside one another’s heads and so must negotiate meanings in conversation.

That is to say, communication through language is not a matter of sending information through a pipeline. It does not happen according to what Michael Reddy (1993) has called the conduit metaphor. Reddy’s article is based on 53 example sentences. Here are the first three (p. 166):
1. Try to get your thoughts across better
2. None of Mary’s feelings came through to me with any clarity
3. You still haven’t given me any idea of what you mean
Reddy’s argument is that many of our statements about communication seemed to be based on the notion of sending something (the thought, idea, feeling) through a conduit, hence he calls it the conduit metaphor. He knows that communication doesn’t work that way, but that’s not is central issue. His central concern is to detail the way we use the conduit metaphor to structure our thinking about communication.

Reddy’s argument is reminiscent of a somewhat earlier argument by Paul de Man, “Form and Intent in the American New Criticism” (1983, first published in 1971). Consider this passage (p. 25):
“Intent” is seen, by analogy with a physical model, as a transfer of a psychic or mental content that exists in the mind of the poet to the mind of a reader, somewhat as one would pour wine from a jar into a glass. A certain content has to be transferred elsewhere, and the energy necessary to effect the transfer has to come from an outside source called intention.
De Man’s point was that, when we read a text, the intention (de Man uses the term in its somewhat rarified philosophical sense) that gives life to those signs on the page is our intention, not the author’s. And he is right.

De Man’s insight, and similar ones by Derrida, Barthes, Foucault and others, had an electrifying effect on literary critics in the United States, leading to a tremendously fertile period in academic literary criticism that, however, became increasingly sclerotic in the 1990s. But that story’s neither here nor there. My point is simply that these thinkers were attempting to deal with a real problem and, ultimately, they failed.

What, for example, could Derrida (1976, p. 158) have possibly meant by proclaiming “There is nothing outside of the text”? What he did not mean is that the world is nothing but a text and a text created by more or less arbitrary social conventions. Read sympathetically, and in context, the phrase seems to mean something to the effect that there is no way we can “step outside” language so as to examine, in full omniscient and transcendental objectivity, the relationship between language and the world. And that, it seems to me, is true. We’re always going to be immersed in “language,” whether natural or the various languages of science and mathematics.

How, then, do we fly free of the bottle? We play games.

Language Games and Game Theory
 
Where de Man argues that intent cannot be transmitted from one speaker to another like pouring wine from a jar, William Croft points out that linguistic communication is tricky “precisely because our thoughts cannot leave our heads” (2000, p. 111). Croft is a linguist who has undertaken to explain language change using an evolutionary approach. He defines a language to be “the population of utterances in a speech community” (p. 26), thus focusing our attention, not on some abstract language system, but on the concrete production of speech.

How does Croft deal with the fact that we cannot transmit thoughts directly to another’s mind? He argues that meaning is negotiated in the back-and-forth of conversation and draws on game theory to make his argument (p. 95):
There is a problem here: the hearer cannot read the speaker’s mind, and she can’t read his. This is what is called a COORDINATION PROBLEM. In speaking and understanding, speaker and hearer are trying to coordinate on the same meaning.
Croft then introduces the notion of a third-party Schelling game in which two players “are presented by a third party with a set of stimuli” which helps them converge on the same meaning. Sometimes it works, sometimes not. One possibility, he argues, is to use “natural perceptual or cognitive distinctiveness [as] a COORDINATION DEVICE” (p. 96). That gives us the adapted mind that I invoked in discussing Quine’s problem. Croft goes on to discuss a variety of linguistic devices as non-conventional coordination devices.

While the details are interesting and important – I recommend his discussion to you – we need not worry about them now.

Save one. Croft notes that, in order for speaker and hearing to reach agreement in conversation their mental states “need not be identical, though it is assumed that they are systematically related” (p. 99). Later on he notes that (114):
successful communication involves not the recovery of and original, ‘correct’ interpretation of the speaker’s original intention, but instead an interpretation that evolves over the course of the conversation, and is assessed by the success or failure of the higher social-interactional goals that the interlocutors are striving to achieve.
One reason why this effort is not doomed to failure from the beginning is the fact that although we cannot read each other’s minds, we do inhabit a shared world.

Croft’s general point, then, is simple, speech communication is a two-way interaction, not the one-way transmission of meaning, information, whatever, though a channel. De Man’s problem is thus solved for the case of face-to-face interaction, a common case, and surely the most basic one. Note that this solution does not involve recourse to a transcendental signified nor to stepping outside the text, nothing like that. It involves the ordinary and obvious means of interactive speech. In this sense, the key to the treasure, is the treasure. Nothing else is required.

But what, you may ask, of written communication, where direct interaction is not possible? After all, de Man was a literary critic, writing about the reading of written texts. What about that?

Good question. I’m going to punt on it. But I observe that some written communication – correspondence – does involve interaction, but at a slower pace than conversation, often much slower. In the case of literary texts, yes, readers cannot ordinary interact with authors, but they can interact with one another. I’ll say a little about that in the next post. Beyond that, yes, there are issues, serious issues. But this is not the place to address them. My concern here is just to get things started.
Note: Mathematician and psychologist Mark Changizi (1999) has an interesting argument about why vagueness of word meaning is essential to the proper functioning of language. His argument is grounded in considerations of computability and I recommend it to you. It makes an interesting complement to the game-theoretic conception of speaking.
Addition: See subsequent post reporting an experiment that David Hays did at RAND in the mid-1950s. It’s relevant to the game theoretic treatment of conversation.
What is a language and what are the memes? 
 
Now I want to shift gears a bit and work my way back to the physical “side” of the linguistic sign, because that’s where we’re going to go looking for memetic entities.

Throughout this post I’ve been assuming that we know what a language is. Now I want to get picky. Here’s what Sidney Lamb has to say in Pathways of the Brain. He’s talking about Roman Jakobson, the great linguist (p. 41):
Using the term language in a way it is commonly used . . . we could say that he spoke six languages quite fluently: Russian, Czech, German, English, Swediksh, and French, and he had varying amounts of skill in a number of others. But each of them except Russian was spoken with a thick accent. It was said of him that “He speaks six languages, all of them in Russian.” . . . the evidence indicates that from a neurocognitive point of view there is no such unit as a language. What exists from a neurocognitive point of view is not so much one linguistic system as a group of interconnected systems, relatively independent from one another.
Lamb goes on to assert that (p. 42):
Professor Jakobson’s internal linguistic information included a single phonological system, that of his native Russian, together with separate systems of grammar and lexicon for Russian, Czech, English, German, French, and Swedish – with some overlap in these grammars and lexicons . . . along with his more limited abilities in various additional languages; plus a conceptual system connected to them all.
So far we’ve been concerned with how meaning is negotiated, where meaning is a matter of the conceptual system. That’s on one “side” of the arbitrary sign, the side inside the brain. Now we’re going to look at the other “side” of the sign, the side that’s in public view, the physical sign. It’s that physical side that most differs among languages.

The question before us is: How do we conceptualize the memetic elements of language? In glossing the emic/etic distinction in a comment to John Wilkins I remarked that (now I’m simply repeating that comment) the distinction originates in linguistics, in the distinction between phonetics and phonemics. The former is about the psychophsics of speech sound while the latter is about phoneme systems. These are obviously very closely related matters, but they aren’t the same. We tend to perceive the speech stream as consisting of discrete sound entities, syllables and phonemes; this is the domain of phonemics. But the speech signal is, in fact, continuous. If you look at a sonogram of some chunk of speech, you don’t draw a series of vertical lines through it separating one phoneme from another; nor can you snip a tape recording into phoneme-long or syllable-long segments and reassemble it into something that sounds like natural speech. The aspects of the speech stream which are phonemically active differ from one language to another, which is why foreign languages all sound like “Greek.” Independently of the fact that you don’t know what the words mean or how the syntax works, you can’t even hear the phonemes in the speech stream.

Now, that’s the distinction I’m after, between phonemes and the raw speech stream. That’s the distinction I drew in my discussion of music (third post). Phonemes are those properties of the speech stream that are linguistically active. We need, however, to distinguish between segmental phonemes and suprasegmental phonemes. The segmental phonemes are roughly parallel to the letters of an alphabetic writing system. Suprasegmentals include tone, stress, and prosodic patterns. And then we need to consider ordering as well, as the order in which elements occur is certainly a property of the speech stream, and a most important one.

Before thinking about order, thought, we need to think a bit more about what’s going on. Roughly speaking, two things need to be extracted from the speech signal: 1) word identities (to be somehow linked to word meanings), and 2) the relations between the words (syntax). My quick take on matters – I’m not a linguist and I’ve not thought this through – is that both segmental and suprasegmental phonemes are involved in both of those processes. Relations between words are often indicated by word affixes, which are realized through segmental phonemes. Word identities are certainly realized by segmental phonemes, but tone and accent are involved as well.

Beyond this, relations between words are signaled by word order. In linguistic typology, typical word order is the primary trait on which classification based. Thus one has SVO languages (subject-verb-object), VSO languages (verb-subject-object), and so forth. As those designations suggest, word order indicates grammatical function, that is, relations between words.

Thus between word order and phonemes we’ve got a rich set of memetic elements. And we could also consider morphology in here as well. Taken together these aspects of the speech signal seem to be as memetically rich and abstract as the musical properties we looked at in discussing Rhythm Changes (first post).

Saturday, March 11, 2023

The issue of meaning in large language models (LLMs)

Noam Chomsky was trotted out to write a dubious op-ed in The New York Times about large language models and Scott Aaronson registered his displeasure at his blog, Shtetl-Optimized: The False Promise of Chomskyism. A vigorous and sometimes-to-often insightful conversation ensued. I wrote four longish comments (so far). I’m reproducing two of them below, which are about meaning in LLMs.

Meaning in LLMs (there isn’t any)
Comment #120 March 10th, 2023 at 2:26 pm

@Scott #85: Ah, that’s a relief. So:

But I think the importantly questions now shift, to ones like: how, exactly does gradient descent on next-token prediction manage to converge on computational circuits that encode generative grammar, so well that GPT essentially never makes a grammatical error?

It's not clear to me whether or not that’s important to linguistics generally, but it is certainly important for deep learning. My guess – and that’s all it is – is that if more people get working on the question, that we can make good progress on answering it. It’s even possible that in, say five years or so, people will no longer be saying LLMs are inscrutable black boxes. I’m not saying that we’ll fully understand what’s going on; only that we will understand a lot more than we do know and are confident of making continuing progress.

Why do I believe that? I sense a stirring in the Force.

There’s that crazy ass discussion at LessWrong that Eric Saund mentioned in #113. I mean, I wish that place weren’t so darned insular and insisting on doing everything themselves, but it is what it is. I don’t know whether you’ve seen Stephen Wolfram’s long article (and accompanying video) but has some nice visualizations of the trajectory GP-2 takes in completing sentences and is thinking in terms of complex dynamic – he talks of “attractors” and “attractor basins” – and seems to be thinking of getting into it himself. I found a recent dissertation in Spain that’s about the need to interpret ANNs in terms of complex dynamics, which includes a review of an older literature on the subject. I think that’s going to be part of the story.

And a strange story it is. There is a very good reason why some people say that LLMs aren’t dealing with meaning despite that fact that they produce fluent prose on all kinds of subjects. If they aren’t dealing with meaning, then how can they produce the prose?

The fact is that the materials LLMs are trained on don’t themselves have any meaning.

How could I possibly say such a silly thing? They’re trained on texts just like any other texts. Of course they have meaning.

But texts do not in fact contain meaning within themselves. If they did, you’d be able to read texts in a foreign language and understand them perfectly. No, meaning exists in the heads of people who read texts. And that’s the only place meaning exists.

Words consist of word forms, which are physical, and meanings, with are mental. Word forms take the form of sound waves, graphical objects, physical gestures, and various other forms as well. In the digital world ASII encoding is common. I believe that for machine learning purposes we use byte-pair encoding, whatever that is. The point is, there are no meanings there, anywhere. Just some physical signal.

As a thought experiment, imagine that we transform every text string into a string of colored dots. We use a unique color for each word and are consistent across the whole collection of texts. What we have then is a bunch of one-dimensional visual objects. You can run all those colored strings through a transformer engine and end up with a model of the distribution of colored dots in dot-space. That model will be just like a language model. And can be prompted in the same way, except that you have to use strings of colored dots.

THAT’s what we have to understand.

As I’ve said, there’s no meaning in there anywhere. Just colored dots in a space of very high dimensionality.

And yet, if you replace those dots with the corresponding words...SHAZAM! You can read it. All of a sudden your brain induces meanings that were invisible when it was just strings of colored dots.

I spend a fair amount of time thinking about that in the paper I wrote when GPT-3 came out, GPT-3: Waterloo or Rubicon? Here be Dragons, though not in those terms. The central insight comes from Sydney Lamb, a first-generation computational linguist: If you conceive of language as existing in a relational network, then the meaning of a word is a function of its position in the network. I spend a bit of time unpacking that in the paper (particularly pp. 15–19) so there’s no point trying to summarize it here.

But if you think in those terms, then something like this

king – man + woman ≈ queen

is not startling. The fact is, when I first encountered that I WAS surprised for a second or two and then I thought, yeah, that makes sense. If you had asked me whether that sort of thing was possible before I had actually seen it done, I don’t know how I would have replied. But, given how I think about these things, I might have thought it possible.

In any event, it has happened, and I’m fine with it even if I can’t offer much more than sophisticated hand-waving and tap-dancing by way of explanation. I feel the same way about ChatGPT. I can’t explain it, but it is consistent with how I have come to think about the mind and cognition. I don’t see any reason why we can’t made good progress in figuring out what LLMs are up to. We just have to put our minds to the task and do the work.

B333 asks: how does meaning get in people’s heads anyway?
Comment #135 March 10th, 2023 at 6:12 pm

@Bill Benzon 120

Ok, well if meaning isn’t in texts, but only in people’s heads, how does meaning get in people’s heads anyway? Mental events occur as physical processes in the brain, and one could well wonder how a physical process in the brain “means” or has the “content” of something external.

Language is highly patterned, and that pattern is an (imperfect) map of reality. “The man rode the horse” is a more likely sentence than “The horse rode the man” because humans actually ride horses, not vice verse. If we switched out words for colored dots those correspondences would still hold. So there is in fact an awful lot of information about reality encoded in raw text.

Meaning = intention + semanticity 
Comment #152 March 11th, 2023 at 7:58 am

@B333 #135: “...how does meaning get in people’s heads anyway?” From other people’s heads in various ways, one of which is language. The key concept is in your last sentence, “encoded.” For language to work, you have to know the code. If you can neither speak nor read Mandarin, that is, if you don’t know the code, then you have no access to meanings encoded in Mandarin.

Transformer engines don’t know the code of any of the languages deployed in the texts they train on. What they do is create a proxy for meaning by locating word forms at specific positions in a high-dimensional space. Given enough dimensions, those positions encode the relationality aspect of (word) meaning.

I have come to think of meaning as consisting of an intentional component and a semantic component. The semantic component in turn consists of a relational component and an adhesion component. (I discuss those three in an appendix to the dragons paper I linked in #120.)

Take this sentence: “John is absent today.” Spoken with one intonation pattern it means just what it says. But when you use a different intonation pattern, it functions as a question. The semanticity is the same in each case. This sentence: “That’s a bright idea.” With one intonation pattern it means just that. But if you use a different intonation pattern is means the idea is stupid.

Adhesion is what links a concept to the world. There are a lot of concepts about physical phenomena as apprehended by the senses. The adhesions of those concepts are thus specified by the sensory percepts. But there are a lot of concepts that are abstractly defined. You can’t see, hear, smell, taste or touch truth, beauty, love, or justice. But you can tell stories about all of them. Plato’s best-known dialog, Republic, is about justice.

And then we have salt, on the one hand, and NaCl on the other. Both are physical substances. Salt is defined by sensory impressions, with taste being the most important one. NaCl is abstractly defined in terms of a chemical theory that didn’t exist, I believe, until the 19th century. The notion of a molecule consisting of an atom of sodium and an atom of chlorine is quite abstract and took a long time and a lot of experimentation and observation to figure out. The observations had to be organized and discipline by logic and mathematics. That’s a lot of conceptual machinery.

Note that not only are “salt” and “NaCl” defined differently, but they have different extensions in the world. NaCl is by definition a pure substance. Salt is not pure. It consists mostly of NaCl plus a variety of impurities. You pay more for salt that has just the right impurities and texture to make it artisanal.

Relationality is the relations that words have with one another. Pine, oak, maple, and palm are all kinds of trees. Trees grow and die. They can be chopped down and they can be burned. And so forth, through the whole vocabulary. These concepts have different kinds of relationships with one another – which have been well-studied in linguistics and in classical era symbolic models.

If each of those concepts is characterized by a vector with a sufficient number of components, they can be easily distinguished from one another in the vector space. And we can perform operations on them by working with vectors. Any number of techniques have been built on that insight going back to Gerald Salton’s work on document retrieval in the 1970s. Let’s say we have collection of scientific articles. Let’s encode each abstract as a vector. One then queries the collection by issuing a natural language query which is also encoded as a vector. The query vector is then matched against the set of document vectors and the documents having the best matches are returned.

It turns out that if the vectors are large enough, you can produce a very convincing simulacrum of natural language. Welcome to the wonderful and potentially very useful world of contemporary LLMs.

[Caveat: from this point on I’m beginning to make this up off the top of my head. Sentence and discourse structure have been extensively studied, but I’m not attempting to do anything remotely resembling even the sketchiest of short accounts of that literature.]

Let’s go back to the idea of encoding the relational aspect of word meaning as points in a high-dimensional space. When we speak or write, we “take a walk” though that space and emit that path as a string, a one-dimensional list of tokens. The listener or reader then has to take in that one-dimensional list and map the tokens to the appropriate locations in relational semantic space. How is that possible?

Syntax is a big part of the story. The words in a sentence play different roles and so are easy to distinguish from one another. Various syntactic devices – word order, the uses of suffixes and prefixes, function words (articles and prepositions) – help us to assemble them in the right configuration so as to preserve the meaning.

Things are different above the sentence level. The proper ordering of sentences is a big part of it. If you take a perfectly coherent chunk of text and scramble the order of the sentences, it becomes unintelligible. There are more specific devices as well, such as conventions for pronominal reference.

A quantitative relationship between concepts, dimensions, and token strings

Now, it seems to me that we’d like to have a way of thinking about quantitative relationships [at this point my temperature parameter is moving higher and higher] between 1) Concepts: the number of distinct concepts in a vocabulary, 2) Dimensions: the number of dimensions in the vector space in which you embed those concepts, and 3) Token strings: the number of tokens an engine needs to train on in order to locate the map the tokens to the proper positions (i.e. types) in the vector space so that they are distinguished from one another and in the proper relationship.

What do I mean by “distinct concepts” & what about Descartes’ “clear and distinct ideas”? I don’t quite know. Can the relationality of words be resolved into orthogonal dimensions in vector space? I don’t know. But Peter Gärdenfors has been working on it and I’d recommend that people working LLMs become familiar with his work: Conceptual Spaces: The Geometry of Thought (MIT 2000), The Geometry of Meaning: Semantics Based on Conceptual Spaces (MIT 2014). If you do a search on his name you’ll come up with a bunch of more recent papers.

And of course there is more to word meaning than what you’ll find in the dictionary, which is more or less what is captured in the vector space I’ve been describing to this point. Those “core” meanings are refined, modified, and extended in discourse. That gives us the distinction between semantic and episodic knowledge (which Eric Saund mentioned in #113). The language model has to deal with that as well. That means more parameters, lots more.

I have no idea what it’s going to take to figure out those relationships. But I don’t see why we can’t make substantial progress in a couple of years. Providing, of course, that people actually work on the problem. 

[Comment inserted on Apr. 2, 2023:  I keep talking about the next cognitive rank, the one we're inching toward. This kind of quantitative work is one of things we'll be doing within that system of thought. In our current cognitive rank we can just create these systems, and bumble about among them. At the next level they'll be objects for examination and analysis.]

Addendum: What about the adhesions of abstract concepts?
Added 3.12.23

Within the semanticity component of meaning I have distinguished between adhesion and relationality: “Adhesion is what links a concept to the world” and “relationality is the relations that words have with one another.” But what about the adhesion of words that are not directly defined in relation to the physical world? Since they are defined over other words, doesn’t their adhesion reduce to relationality?

Not really. Take David Hays’s standard example: Charity is when someone does something nice for someone else without thought of reward. Any story that satisfies the terms of that definition (“when someone does...reward”) is considered an act of charity. The adhesion of the definiendum, charity, is not with any of the words, either in the definiens or in any of the stories that satisfy the definiens, but with the pattern exhibited by the words. It’s the pattern that characterizes the connection to the world, not the individual words in stories or in the defining pattern.

Friday, March 10, 2023

Uncle Noam got it wrong in The New York Times [meaning = intention + adhesion + relationality]

The article is well worth reading.

At this point in time Bender is perhaps most widely known as the person who coined the term "stochastic parrot." I think the term is rhetorically brilliant, but misleading. Late in the middle the article juxtaposes Bender against Christopher Manning:

Bender and Manning’s biggest disagreement is over how meaning is created — the stuff of the octopus paper. Until recently, philosophers and linguists alike agreed with Bender’s take: Referents, actual things and ideas in the world, like coconuts and heartbreak, are needed to produce meaning. This refers to that. Manning now sees this idea as antiquated, the “sort of standard 20th-century philosophy-of-language position.”

“I’m not going to say that’s completely invalid as a position in semantics, but it’s also a narrow position,” he told me. He advocates for “a broader sense of meaning.” In a recent paper, he proposed the term distributional semantics: “The meaning of a word is simply a description of the contexts in which it appears.” (When I asked Manning how he defines meaning, he said, “Honestly, I think that’s difficult.”)

If one subscribes to the distributional-semantics theory, LLMs are not the octopus. Stochastic parrots are not just dumbly coughing up words. We don’t need to be stuck in a fuddy-duddy mind-set where “meaning is exclusively mapping to the world.” LLMs process billions of words. The technology ushers in what he called “a phase shift.” “You know, humans discovered metalworking, and that was amazing. Then hundreds of years passed. Then humans worked out how to harness steam power,” Manning said. We’re in a similar moment with language. LLMs are sufficiently revolutionary to alter our understanding of language itself. “To me,” he said, “this isn’t a very formal argument. This just sort of manifests; it just hits you.”

I note that the term "distributional semantics," I believe, was coined by linguist Raymond Firth back in the late 1950s. This is so well-known that I assume Manning knows it and wasn't claiming the term as his own in the paper.

My own position on what's at issue between them is subtle. I certainly recognize the relationships words have among themselves, which is what Manning is arguing. But I also recognize intention, as Bender does. We need both. To put it over schematically:

meaning = intention + semanticity

semanticity = relationality + adhesion

I say more about that in this post, from May 2023, and this one, from April 2023.

Thursday, May 19, 2022

Meaning and Semantics, Relationality and Adhesion

For some time now I’ve been making a distinction between meaning and semantics, where I use meaning as a function of intention, in the more-or-less standard philosophical sense of intention as “aboutness.” When I talk of semantics I am talking about the elements in the language system. I have now decided that semantics has two aspects: relationality and adhesion.

I suppose we can think of meaning as inhering in the intentional relationship between the person – and we are talking about human beings here, but we could be talking about animals or maybe, just maybe, artificial minds – and the world. Walter Freeman regarded meaning as inherent in the total trajectory of the brain’s state during some experience, whether conversation, reading a text, or out and about in the world. There is more to the brain’s state than the operations of the language and cognitive systems. Thus meaning is necessarily different from semantics.

The standard philosophical arguments (such as Searle’s Chinese room) about artificial intelligence (which I’m now calling artificial minds), focus on meaning and intention to the utter neglect of semantics, as though it doesn’t exist. It may well be the case that all these artificial systems fail on the grounds of intentionality. It seems to me that the success of this line of argument is also a pyrrhic victory, for it leaves the philosopher powerless to reason about what these systems can do. It leads to a false binary where either the system is a human or it is a worthless artifact.

But that’s an aside. I’m much interested in that philosophical argument at the moment. I’m intersted in semantics, with its aspects of adhesion and relationality. Roughly speaking, adhesion is what ‘connects’ a concept to the world through perception. If we use a standard semantic network diagram (below), adhesion is carried on the REP (represent) arcs.

Relationality is carried on the arcs linking concepts with one another. Thus VAR in the diagram is for variety; beagles and collies are varieties of dog. We can also think of adhesion as being about compression (data reduction) and categorization – Gärdenfors’ dimensions in concept spaces. Relationality is about relations between objects in different concept spaces. But that’s only a rough characterization.

Large language models, such as GTP-3, are exploiting semantic relationality – the argument I made in my GPT-3 working paper, but have no access to adhesion. Vision systems are gounded in adhesion and may also exploit aspects of relationality.

[If we use Powers’s notion of intensities, where perception and cognition have to account for incoming intensities, then adhesion is about compression of intensities while relationality is about distribution of them over different concepts.]

More later.