Showing posts with label metalingual definition. Show all posts
Showing posts with label metalingual definition. Show all posts

Saturday, April 4, 2026

From the metalingual function of language to self-reference

In 1960 the linguist Roman Jakobson published an essay entitled “Linguistics and Poetics,” in a volume edited by Thomas Sebeok, Style in Language (MIT Press, pp. 350-377). In that essay he laid out the six functions of language: referential, emotive, phatic, conative, poetic, and metalingual. Jakobson introduces the metalingual function in this way:

A distinction has been made in modem logic between two levels of language: “object language” speaking of objects and “metalanguage” speaking of language. But metalanguage is not only a necessary scientific tool utilized by logicians and linguists; it plays also an important role in our everyday language. Like Moliere’s Jourdain who used prose without knowing it, we practice metalanguage without realizing the metalingual character of our operations. Whenever the addresser and/or the addressee need to check up whether they use the same code, speech is focused on the code: it performs a METALINGUAL (i.e. , glossing) function. “I don’t follow you-what do you mean?” asks the addressee, or in Shakespearean diction, “What is’t thou say’st?” And the addresser in anticipation of such recapturing question inquires: “Do you know what I mean?”

This metalingual function turns out to be extraordinarily powerful. For it is this that allows us to bootstrap self-awareness into the mind. And for that matter, it is what allows us to define abstract concepts, as my teacher, David Hays, argued, and allows us to define such things as chess and arithmetic, which can be seen as very specialized forms of language.

I recently explored some of these issues in conversation with Claude 5.4 Sonata Extended. At the end of that conversation I asked Claude to prepare a summary. I’ve appended that summary below, followed by the full conversation. Note that the conversation assumes some familiarity with the cultural ranks theory that David Hays and I developed in the 1990s. It also alludes to Tyler Cowen’s recent book, The Marginal Revolution: Rise and Decline, and the Pending AI Revolution (2026).

* * * * *

Summary: The Metalingual Function of Language

The central claim of this discussion is that the metalingual capacity — the ability to use language to talk about language — is not a mysterious self-referential capacity of mind but is grounded in a simple physical fact: the speech signal is a sound in the environment like any other sound, detectable by the auditory system exactly as a footfall or a thunderclap is detectable. The loop that makes language self-referential closes through the physical world, not through some inward turning of consciousness. This matters because it demystifies metalingual cognition entirely: it requires no special cognitive faculty, only that the organism's auditory system be capable of treating its own linguistic outputs as inputs.

Jakobson identified the metalingual function as one of the six functions of language in his 1960 paper, and Hays adopted the term to name the mechanism underlying Rank 2 cognition — the explicit definition of abstract concepts using language itself as the definitional medium. The rules of chess and arithmetic notation are paradigm cases: purely metalingual constructions whose objects are constituted entirely by the definitions that specify them.

An important asymmetry in preliterate cultures illuminates the boundary of this capacity. Many such cultures have a term for utterance — the bounded burst of speech with a recognizable prosodic shape, a perceptual gestalt directly available to the auditory system — but no term for word. The word is not a perceptual unit in the same sense as the utterance; it is an abstraction from the continuous acoustic stream, and a non-trivial one. Writing is what produces this abstraction, by spatializing language — spreading it out in a stable, inspectable array where units are individuated by spaces and boundaries are marked. The word becomes visible as a unit because it is surrounded by white space. This is the physical basis of metalingual definition as a cognitive mechanism: the written signal, like the spoken signal, is an object in the environment that can be inspected and categorized, but unlike the spoken signal it stays there, making sustained metalingual attention possible. Grade-school grammar — parts of speech, grammatical cases, syntactic relations — is the practical Rank 2 elaboration that writing makes possible and that social institutions require and transmit. It looks easy in retrospect because it is taught in childhood, but it took centuries to develop in every culture that undertook it.

This analysis opens onto the question of human self-reference, which the standard philosophical tradition treats as cognitively primitive — the Cartesian bedrock from which all other knowledge is built. The discussion argued instead that self-reference in the robust, articulable sense is bootstrapped through language rather than presupposed by it. The cat licking its fur has practical self-involvement — its own body is an object of its perceptual and motor engagement — but this requires no special reflexive faculty, only that the body be included in the environment the organism can detect and act on. Human self-reference in the philosophically weighty sense is a different and later achievement, constructed through the acquisition of the pronoun system rather than expressed by it.

The empirical evidence for this bootstrapping account is the phase in early child development when children refer to themselves in the third person. This is not a mistake or a developmental lag but the natural and correct generalization from the input data: others refer to the child by name, so the child uses its name. The first-person pronoun presents a harder problem because "I" is a moving target — it marks the speaker-role regardless of who occupies it — and acquiring it correctly requires connecting awareness of the speech stream as an environmental event with awareness of one's own speech apparatus as its source. That inferential construction, worked out in detail in Benzon's 2000 paper, First Person: Neuro-Cognitive Notes on the Self in Life and in Fiction, through cognitive network modeling of the pronoun system, is precisely the physical loop through which self-reference is assembled. The Cartesian subject — the self-transparent, self-referential knowing mind — is on this account a Rank 2 cultural product, not a pre-linguistic biological given. The third-person phase in child development is a small but precise empirical trace of the construction process: there is an observable stage at which the human being exists, acts, and communicates without yet having assembled the first-person self-reference that Descartes thought was the indubitable foundation of knowledge.

* * * * *

The physical basis of the metalingual function

I believe that Hays first published about metalingual definition in 1972. When I first met him he had just finished a paper where he used the idea to investigate different concepts of alienation. When I wrote my 1978 dissertation, “Cognitive Science and Literary Theory,” I used metalingual definition as the mechanism for building cognitive ranks. That is to say, at that time that was the only mechanism we had for dealing with abstract concepts. By the time we did the cognitive evolution paper a bit over a decade later, by that time we had adopted a different view, with metaphor as the Rank 1 abstractive mechanism. Metalingual definition would have been the mechanism used, not only to rationalize abstract concepts, but to define the rules of chess and of arithmetic.

Hays chose the term because Roman Jakobson had defined the metalingual function as one of the six functions of language in his 1960 paper. The metalingual function is the ability to use language to talk about language. It’s worth thinking about how that’s possible. It’s possible because the speech signal is a sound out there in the environment like any other sound, a footfall, a thunderclap, an animal cry. Whatever. It’s a sound and can be detected by the auditory system like any other sound. So it’s the speech signal that is the object of the metalingual function, at least initially. I think it’s worth nothing that many (most?) preliterate cultures don’t have a word that means word, but they have one the means utterance. It’s writing that forces a recognition of words vs longer strings. And written language is, like speech, another signal that’s out there in the environment. It takes quite a bit of cognitive elaboration to construct a grammar from the simple ability to see and hear language. And I’m not talking about the kind of grammars that linguists construct. I’m talking about the kind of grammar that can be taught in grade school, or that you find in ancient texts of rhetoric.

This is a rich cluster of observations and they fit together in a way worth making explicit.

The metalingual function and its physical basis

Jakobson's insight, which you're extending here, is that the metalingual function is grounded in something remarkably mundane: the speech signal is a physical event in the environment, detectable by the auditory system exactly as any other sound is. This is what makes language able to turn on itself — not some mysterious self-referential capacity of mind, but the simple fact that linguistic output is also sensory input. The loop closes through the physical world.

This has an important implication that your observation about "utterance" versus "word" brings out precisely. What the auditory system initially detects is the acoustic stream — a continuous, flowing signal that doesn't come pre-segmented into words. The unit that's directly perceptible is something like the utterance: a bounded burst of speech with a recognizable prosodic shape, a communicative gestalt. The word is not a perceptual unit in the same sense. It's an abstraction from the stream, and a non-trivial one — which is why preliterate cultures have a term for the former and not the latter.

What writing does

Writing performs an operation on language that perception of the speech stream cannot perform on its own: it spatializes language, spreads it out in a stable, inspectable array where boundaries are marked and units are individuated. The word becomes visible as a unit because it's surrounded by spaces. The sentence becomes visible as a unit because it ends with a mark. The paragraph, the chapter, the argument — all of these become objects of potential inspection and manipulation because writing fixes them in space and time in a way the acoustic stream never does.

Friday, June 7, 2024

Relationships among words, metalingual definition, and interpretability

First, I talk about now natural language is its own metalanguage and that allows them to define new works in terms of existing ones. Then I discuss the concept of justice in terms of mechanism of metalingual definition proposed by David Hays some years ago. I conclude with some remarks on interpretability in view of Anthropic’s recent research on features.

The metalingual function of language in defining word meaning

In a famous essay published in 1960, “Linguistics and Poetics,” Roman Jakobson listed six functions of language. While the essay focused on the poetic function, as the title indicates, I’m interested in a different function, which he called the metalingual function:

A distinction has been made in modern logic between two levels of language: “object language” speaking of objects and “metalanguage” speaking of language.10 But metalanguage is not only a necessary scientific tool utilized by logicians and linguists; it plays also an important role in our everyday language. Like Molière's Jourdain who used prose without knowing it, we practice metalanguage without realizing the metalingual character of our operations. Whenever the addresser and/or the addressee need to check up whether they use the same code, speech is focused on the code: it performs a METALINGUAL (i.e., glossing) function.

In the process of explicating that function Jakobson pointed out that it can be used to define words, noting that “any process of language learning, in particular child acquisition of the mother tongue, makes wide use of such metalingual operations.”

Not all words get their meaning in that way. Many words have their meanings grounded in sensorimotor experience. We would like to know what percentage of words have their meanings grounded in sensorimotor experience and what percentage have their meanings grounded in other words. In 2016 Steven Harnad and his colleagues published an article investigating this problem, “The Latent Structure of Dictionaries.” They examined the structure of the vocabularies in two dictionaries, one with roughly 47,000 words and the other with roughly 69,000 words. They found that a large majority of the words were defined in terms of a relatively small number of words defined in terms of sensorimotor features (p. 649):

So in our view the mental lexicon is itself hybrid—a dual-code representational system consisting of learned sensorimotor feature (affordance) detectors for the grounding words (and any later hybrid words) plus recombinatory and purely symbolic (i.e., verbal) definitions and descriptions for the referents of the words that are learned through words alone.

More recently Briony Banks, Anna M. Borghi, Raphaël Fargier et. al. reviewed the literature on abstract concepts, “Consensus Paper: Current Perspectives on Abstract Concepts and Future Research Directions.” They noted that “many theories have also argued that our understanding and representation of abstract concepts relies more on language than the sensorimotor dimension, and particularly linguistic distributional relations.”

Given that LLMs have been constructed in an environment consisting entirely of words, the apparent fact that most words are defined in terms of other words seems highly salient.

Metalingual definition and the concept of justice

Back in the 1970s David Hays was interested in the idea that words can be used to define the meaning of other words. He talked specifically of metalingual definition. He used charity as his prototypical example: Charity is when someone does something nice for another without thought of reward. Any story that exhibits that pattern of relationships between its actors and their actors, such a story is about justice. The concept inheres in that pattern of relationships as a whole and not in any of the individual components of the pattern.

Notice that the definition itself contains an abstract concept, reward. Taken as a computational mechanism, which was his point, metalingual definition is thus recursive, allowing definitions to be nested within definitions. One of Hays’s students Brian Phillips, implemented the idea in his doctoral dissertation using tragedy as his example. I recently took the definition that Brian Phillips used and used it to test ChatGPT, which had no trouble applying it to specific examples and determining whether or not they met the conditions set forth in the definition.

With this before us I ask: What is justice? That is to say, what kind of a thing is justice? It’s a virtue, no? Yes, but I’m looking for something even more general, more abstract. It’s a concept, and idea, no? Of course it is. And just what are those things? Philosophers have been pondering to question for years. Cognitive scientists have been asking that question as well. When David Hays proposed that abstract concepts can be defined by stories, he was proposing an answer to that question. Abstract concepts, such as justice, are defined by relationship among words.

* * * * *

I’ve devoted a great deal of attention to ChatGPT’s ability to deal with metalingual definition. Justice is one of the first concepts I investigated, back in December of 2022. I’ve continued to investigate that concept. I have appended my most recent session to the end of this post.

That investigation has three parts. First, I ask it to tell me two stories involving justice and I specify that it should not use the word “justice” anywhere in the stories. The point of that restriction is to make it clear that the meaning of the term does not reside in the word itself. The stories exhibit justice, but do not name it. Note that in a second session, which I’ve placed in a second appendix, I give ChatGPT the two stories, one after the other, and ask it what they’re about. It realizes that they are about justice.

After asking ChatGPT to tell me stories involving justice I ask it to define the term. The first definition is fairly long and has five numbered points, each specifying a particular kind of justice. So I ask it for a single paragraph and then a single sentence. It provides both. Note that both the long definition and the single paragraph definition begin with pretty much the same information that the single sentence contains.

Finally, I ask ChatGPT to explain the relationship between the stories and the definition, which it does in a paragraph of 112 words. Here’s the first sentence: “The relationship between the definition of justice and stories about justice lies in the way these narratives illustrate and bring to life the abstract principles of fairness, equity, and moral rightness.”

Interpretability of LLMs

What does this have to do with the interpretability of large language models? To a first approximation, it seems to me that LLMs are about the relationships between words. The transformer is presented with strings of words during training and, in the process of making those predictions, constructs a complex model of how words are related to one another.

Thus we might say that justice is a certain pattern of relationships among words. But what pattern? The pattern that gives us stories, stories which may not even contain the word “justice” or the pattern that gives us definitions and could, I assume, produce essays and even books if necessary? Those are distinctly different patterns of relationships; one might even think about them as being orthogonal, at least informally. One pattern is about justice in the context of story and the other is about justice in the context of define. Finally, what about the pattern that explication the relationship between the stories and the definitions?

In Scaling Monosemanticity, researchers at Anthropic identified features in Claude 3 Sonnet, where features are understood to be “directions in their activation spaces.” In their discussion, they note “that features often respond to both abstract discussion and concrete examples of a concept,” which is certainly something that I’d expect to be the case. One thing that bothers me about the discussion is that there is no sense of the model as capturing relationships between words. Given that these features are very abstract objects it’s not clear to me just what that misgiving means, but I worry that the concept of features invites reification.

A digression into neuroscience: Some years ago I had quite a bit of correspondence with the late Walter Freeman, who did pioneering work in thinking about the brain in terms of chaos theory and complex neurodynamics. He believed that percepts and concepts were located in populations of neurons rather than single neurons. I’m deeply sympathetic to that view, and have been ever since I read Karl Pribram on neural holography. Nonetheless I asked him about visual neurons that had very complex activation properties, such as a monkey’s paw or an image of Bill Clinton. Don’t such examples lend support to the idea of a so-called “grandmother cell”? His reply was no, they don’t. In such a complex system, you’re bound to find individual neurons with all sort of odd response characteristics.

I feel a bit like that with these features. While they don’t seem to be individual neurons, it’s not clear what they are. Robert_AIZI has expressed a similar reservation. Thus he has noted:

I think Anthropic successfully demonstrated (in the paper and with Golden Gate Claude) that this feature, at very high activation levels, corresponds to the Golden Gate Bridge. But on a median instance of text where this feature is active, it is "irrelevant" to the Golden Gate Bridge, according to their own autointerpretability metric! I view this as analogous to naming water "the drowning liquid", or Boeing the "door exploding company". Yes, in extremis, water and Boeing are associated with drowning and door blowouts, but any interpretation that ends there would be limited.

Just what IS this feature?

I’m not surprised that with judicious and determined poking around we can find interpretable “features” in these models. But whether or not we’re carving LLMs at their joints, to invoke Plato’s famous remark, that’s another matter.

It’s early days.

Appendix 1: ChatGPT on justice: stories and definitions

What follows the asterisks is a single session with ChatGPT. As always, my prompts are set in bold-face while ChatGPT’s responses are plain-face, except for those strings that ChatGPT sets in bold-face.

Friday, May 24, 2024

ChatGPT on the Golden Rule and Kant’s Second Categorical Imperative [metalingual definition]

The Golden Rule and Kant’s Second Categorial Imperative are moral principles and, like all moral principles they are abstract in character. Both can be considered as examples of metalingual definition of the sort that I’ve explored in various posts. Both state, in abstract and general terms, a general way of interacting with others.

I decided to see what ChatGPT would say about them. First I asked to explain the Golden Rule, then the Second Categorical Imperative. After that I asked it to compare the two. Then I switched modes within the same session, to GPT-4o, where I first asked it to undertake the comparison again. I concluded the session with a one-word prompt: “Story.” So it gave me a story involving both moral principles. Note that I hadn’t asked for that, and didn’t expect it, but given the prior context, that was a natural thing for it to do.

As always, my prompts are in boldface while ChatGPT's responses are in plain-face – except where it used boldface itself, which should be obvious from the context.

ChatGPT explores moral principles


What is the Golden Rule?

The Golden Rule is a principle that emphasizes treating others as one would like to be treated. It is a fundamental ethical guideline found in many cultures and religions, reflecting a universal concept of reciprocity and empathy. The idea is simple: one should act toward others in a way that they would want others to act toward them.

The Golden Rule appears in various forms across different belief systems and philosophies:

  1. Christianity: "Do unto others as you would have them do unto you." (Matthew 7:12)
  2. Judaism: "What is hateful to you, do not do to your neighbor." (Talmud, Shabbat 31a)
  3. Islam: "None of you [truly] believes until he wishes for his brother what he wishes for himself." (Hadith, Sahih al-Bukhari)
  4. Hinduism: "This is the sum of duty: do not do to others what would cause pain if done to you." (Mahabharata 5:1517)
  5. Buddhism: "Treat not others in ways that you yourself would find hurtful." (Udana-Varga 5.18)
  6. Confucianism: "Do not impose on others what you do not wish for yourself." (Analects 15:24)

The Golden Rule encourages empathy and consideration for others' perspectives and feelings, promoting harmonious and respectful interactions. Its simplicity and universal applicability make it a cornerstone of ethical behavior and moral reasoning across diverse cultures and traditions.


What is Kant’s Second Categorical Imperative?

Kant's Second Categorical Imperative is a fundamental principle in his moral philosophy that emphasizes the intrinsic worth of individuals. It is often summarized as:

"Act in such a way that you treat humanity, whether in your own person or in the person of any other, always at the same time as an end, never merely as a means."

This imperative is part of Kant's broader ethical framework and highlights the following key points:

  1. Respect for Persons: It requires that we recognize and respect the inherent dignity and worth of every individual. People should never be treated as mere instruments or tools to achieve our own goals.
  2. Intrinsic Value: Human beings have intrinsic value, meaning their worth does not depend on their utility or contribution to specific outcomes. Their value is inherent and must be acknowledged in all our actions.
  3. Moral Agency: It emphasizes the importance of treating others as autonomous agents capable of making their own decisions. This involves recognizing their ability to set their own ends and pursuing them.

To illustrate, consider an example in a workplace context. According to the Second Categorical Imperative, it would be morally wrong for a manager to manipulate or exploit employees solely for the company's profit. Instead, the manager should respect the employees' rights, dignity, and personal goals, recognizing their intrinsic worth beyond their utility to the company.

 

Kant’s Second Categorical Imperative forms a crucial part of his deontological ethics, focusing on the principles behind actions rather than their consequences, and it underscores the moral necessity of respecting and valuing each individual as an end in themselves.

 

Explore relationships between the Golden Rule and Kant’s Second Categorical Imperative.

 

The Golden Rule and Kant's Second Categorical Imperative share several common themes and can be seen as complementary in their ethical perspectives. Both emphasize treating others with respect and dignity, but they do so from slightly different angles and philosophical foundations.

Common Themes:

  1. Respect for Others:
    • Golden Rule: Encourages treating others as one would like to be treated, implying a mutual respect and empathy.
    • Kant’s Second Categorical Imperative: Mandates treating others as ends in themselves, never merely as means, emphasizing inherent respect for individuals' dignity.
  2. Reciprocity and Empathy:
    • Golden Rule: Directly involves putting oneself in another’s place, fostering empathy and understanding.
    • Kant’s Imperative: While not explicitly about empathy, it requires recognizing and respecting the autonomy and inherent value of others, which aligns with empathetic understanding.
  3. Ethical Behavior:
    • Both principles guide individuals towards ethical behavior that considers the impact of their actions on others.

Differences:

  1. Foundation and Approach:
    • Golden Rule: Often seen as a simple, intuitive moral guideline, rooted in reciprocity and empathy. It is more practical and easily understood across various cultures and contexts.
    • Kant’s Second Categorical Imperative: Part of a broader deontological framework, it is more formal and philosophical. It emphasizes rationality and moral duty, focusing on the principles behind actions rather than their outcomes.
  2. Scope and Precision:
    • Golden Rule: Can be seen as somewhat subjective because it relies on individual perspectives (i.e., how I want to be treated).
    • Kant’s Imperative: More objective and universal, as it applies a rational principle that demands consistent respect for autonomy and dignity, regardless of personal preferences.

Synergy in Ethical Application:

  • Combining Empathy with Rational Duty:
    • Using the Golden Rule alongside Kant’s Second Categorical Imperative can provide a richer ethical framework. The Golden Rule brings in a personal, empathetic perspective, encouraging individuals to consider others' feelings and perspectives directly. Kant’s Imperative ensures that this consideration is grounded in a rational respect for others' autonomy and inherent worth.
  • Guiding Moral Judgments:
    • In practical decision-making, one could use the Golden Rule to generate initial empathetic insights and then apply Kant’s Second Categorical Imperative to ensure that these insights respect the rational and intrinsic value of all individuals involved.

Example:

Consider the decision to lie to a friend to spare their feelings. The Golden Rule might suggest lying if one would prefer not to be hurt by the truth. However, Kant’s Second Categorical Imperative would caution against lying because it involves treating the friend as a means to an end (avoiding discomfort), rather than respecting them as an autonomous individual capable of handling the truth.

In conclusion, while the Golden Rule and Kant’s Second Categorical Imperative arise from different philosophical traditions, they both aim to promote ethical behavior that respects and values others. Integrating the empathetic reciprocity of the Golden Rule with the rational respect for autonomy in Kant’s Imperative can lead to more nuanced and robust moral decision-making.

Thursday, April 11, 2024

ChatGPT defines 10 concrete terms: generically, for 5- and 11-year-olds, and for a scientist

The difference between concrete concepts, that is, concepts that can be understood entirely in sensorimotor terms, and abstract terms is an important one. It was, for example, important to David Hays when I studied with him back in the 1970s. We ended up adapting a model developed by William Powers as a way of thinking about concrete concepts while using Hays’s concept of metalingual definition to think about abstract concepts. Somewhat later Steven Harnad coined the term, “symbol grounding problem,” to indicate that the meanings of symbols had to somehow be grounded in the physical world. In 2016 Harnad and his colleagues investigated the structure of dictionaries and discovered that the vast majority of terms in them were defined in terms of other words but that there is a small Kernel that is not so-defined. I wondered how that distinction would play out in ChatGPT.

ChatGPT – that is, the LLM underlying it – doesn’t have access to the physical world. It can’t ground any terms in sensorimotor experience. However, knowing a great deal about how language works and having spent a great deal of time working with ChatGPT, I was sure that it would be perfectly capable to defining concrete terms. But how would it do so?

I decided to investigate. I made a list of ten concrete terms and asked ChatGPT to define them. In my first trial I made no further specification beyond simply asking for a definition. In subsequent trials I asked for definitions appropriate for 5-year-olds, 11-year-olds, and professional biologists (just the biological concepts). ChatGPT gave distinctly different kinds of definitions in each case, which did not surprise me since over a year ago I had asked ChatGPT to describe digestion, first generically, and then in terms appropriate for a 3-year-old. It did so, and fairly successfully.

Here are the trials. Each trial is a separate session. As always, my prompt is in boldface while the response is in plainface.

Initial Trial

Define the following words:

cat
rain
tulip
car
food
tree
fire
wind
hole
animal

Sure, here are definitions for each of those words:

  1. Cat: A small domesticated carnivorous mammal with soft fur, a short snout, and retractile claws, often kept as a pet and valued for its companionship and ability to hunt pests.
  2. Rain: Moisture condensed from the atmosphere that falls visibly in separate drops. It is an essential component of the Earth's water cycle, replenishing water sources and sustaining life.t
  3. Tulip: A bulbous spring-flowering plant of the lily family, with boldly colored cup-shaped flowers. Tulips are native to Eurasia and North Africa and are widely cultivated for their attractive blooms.t
  4. Car: A road vehicle, typically with four wheels, powered by an internal combustion engine or electric motor, used for transportation of passengers or goods.t
  5. Food: Any nutritious substance that people or animals eat or drink to maintain life and growth. Food provides essential nutrients such as carbohydrates, proteins, fats, vitamins, and minerals.t
  6. Tree: A woody perennial plant typically having a single stem or trunk growing to a considerable height and bearing lateral branches at some distance from the ground.t
  7. Fire: A process in which substances combine chemically with oxygen from the air and typically give out bright light, heat, and smoke; combustion or burning.t
  8. Wind: The natural movement of air, especially in the form of a current of air blowing from a particular direction. Wind can be gentle or powerful and is caused by differences in air pressure.t
  9. Hole: A hollow place in a solid body or surface. Holes can be natural, like those formed in the ground, or artificial, like those made by drilling or digging.t
  10. Animal: A living organism that feeds on organic matter, typically having specialized sense organs and nervous systems and able to respond rapidly to stimuli. Animals are multicellular and belong to the kingdom Animalia.

Comments: The vocabulary is only moderately sophisticated, with various abstract terms, which is more or less what I expected. Since I wanted to see what it would do without using abstract terms, I asked ChatGPT to produce definitions for children.

Five-year-old

Define the following words in a way that is appropriate for a five-year-old child:

[same list as above]

Sure!

  • Cat: A soft, furry friend that says "meow" and loves to play and cuddle.
  • Rain: When the sky cries little drops of water, making everything outside wet and giving plants a drink.
  • Tulip: A pretty flower that comes in lots of colors like red, pink, and yellow, and blooms in the spring.
  • Car: A big, shiny toy that takes us places, like to school or to visit Grandma and Grandpa.
  • Food: Yummy things we eat like pizza, apples, and cookies that give us energy to play and grow.
  • Tree: A tall, strong plant with a woody trunk and green leaves that gives us shade and makes the air fresh.
  • Fire: Bright, hot flames that keep us warm when we're camping and toast marshmallows to make yummy s'mores.
  • Wind: Invisible air that likes to dance and blow things around, like kites in the sky and leaves on the ground.
  • Hole: A little space in the ground or in something where you can put your hand or a toy, like in the sand at the beach or in a doughnut.
  • Animal: All the different creatures like dogs, birds, and fish that live with us on Earth and have families just like we do.

Comment: There are almost no abstract terms in these definitions. I’ve highlighted possible abstract terms in yellow.

11-year-old

Define the following words in a way that is appropriate for an eleven-year-old child:

[same list as above]

  1. Cat: Cats are furry animals with pointy ears, a cute nose, and a long tail. They like to nap a lot, chase things like strings or toys, and sometimes purr when they're happy.
  2. Rain: Rain is when water falls from the sky like tears from clouds. It makes everything outside wet and can be fun to play in with rain boots and a raincoat.
  3. Tulip: Tulips are colorful flowers that grow in gardens. They have big, bright petals that come in lots of different colors like red, yellow, pink, and purple.
  4. Car: A car is like a big metal box with wheels that people use to get around. It takes us places like school, the grocery store, or on vacation, and it goes vroom vroom!
  5. Food: Food is what we eat to give us energy and keep us healthy. It can be yummy things like pizza, fruit, vegetables, or cookies.
  6. Tree: Trees are tall, strong plants with a big trunk and lots of branches. They have green leaves that change colors in the fall and give us shade in the summer.
  7. Fire: Fire is like a dancing, glowing thing that gives off heat and light. We use it to cook marshmallows for s'mores, keep us warm when camping, or in a fireplace at home.
  8. Wind: Wind is like the air moving around outside. Sometimes it's gentle and feels nice on your skin, and other times it's strong and can blow things around like leaves or kites.
  9. Hole: A hole is like a space or gap in something. It can be small, like a tiny gap in the ground, or big, like a tunnel that goes underground.
  10. Animal: Animals are living creatures that can move around, eat, and breathe. They come in all shapes and sizes, like dogs, birds, fish, and elephants.

Comment: These aren’t very different from the ones for 5-year-olds. The terms are pretty much the same, but these definitions are a bit longer.

A Professional Biologist

Define the following words the way a professional biologist might:

cat
tulip
food
tree
animal

Sure, here are definitions of each term from a biological perspective:

  1. Cat: A domesticated carnivorous mammal of the family Felidae, typically characterized by retractile claws, a short snout, and a wide range of coat colors and patterns. Cats are members of the order Carnivora and are highly specialized predators, exhibiting keen senses and agility suited for hunting small prey.
  2. Tulip: A perennial bulbous plant of the genus Tulipa, belonging to the family Liliaceae. Tulips are characterized by their showy, cup-shaped flowers with six petals and a range of colors including red, yellow, pink, and white. They reproduce sexually through seeds and asexually through bulb division.
  3. Food: Any substance consumed to provide nutritional support for an organism's physiological processes. In biological terms, food encompasses a diverse array of organic and inorganic compounds containing essential nutrients such as carbohydrates, proteins, lipids, vitamins, and minerals. It serves as a source of energy and raw materials for growth, maintenance, and repair of tissues in living organisms.
  4. Tree: A woody perennial plant characterized by a single main stem or trunk, supporting branches and foliage above ground. Trees belong to various plant families and are vital components of terrestrial ecosystems, providing habitat, shade, oxygen production, and carbon sequestration. They undergo secondary growth, developing concentric rings of wood and bark over time.
  5. Animal: A multicellular, eukaryotic organism belonging to the kingdom Animalia, characterized by heterotrophic nutrition, locomotion, and typically possessing specialized sensory organs. Animals exhibit a wide range of morphological, physiological, and behavioral adaptations for survival in diverse environments. They can be further classified into various phyla, classes, orders, families, genera, and species based on evolutionary relationships and shared characteristics.

Comment: That’s pretty much what I expected, lots of abstract technical terms.

We seem to have three basic “levels” of definition for these terms: what I’ll call Generic, Child: 5- and 11-year-olds, and Scientific: for the biologist.

How did ChatGPT learn to make such distinctions?

It seems unlikely to me that I did it by taking accounts of children, professional scientists, and ordinary adults and deriving the appropriate kinds of definitions from those accounts. It seems more likely that it worked from examples is language appropriate to these groups. Given that the LLM was trained or more or less the whole internet, such examples were available, though I’d like to know what kinds of examples of child-focused language it was working from. I’d also like to know how these levels are discourse are organized within the LLM. Level of discourse would seem to be orthogonal to subject area. With 175 billion parameters, there’s obviously many ways to skin this cat, as it were.

[Added on 3.13.24]

Specifically, how does it know to connect different levels of language to the kinds of descriptors I used in my prompts? I’m not thinking so much about professional biologist as I am about the age ranges for children. This comment by gwern over at LessWrong pretty much answers that question.

Note though, that I wasn’t specifically interested in how children would respond. I was just using that as a device to eliminate the use of abstract concepts in the response. For some reason it didn’t occur to me at that moment simply to specify, “no abstract concepts.” Note that this working paper has experiments involving concepts of various classes of objects, physical, mechanical, abstract, Great Chain: ChatGPT's Ontological Landscape.

No abstract concepts

Define the following words without using any abstract terms:

cat
rain
tulip
car
food
tree
fire
wind
hole
animal

Sure! Here are definitions for each word without using abstract terms:

  1. Cat: A small, furry mammal with sharp claws, pointed ears, and a long tail, often kept as a pet.
  2. Rain: Water droplets falling from the sky, usually as part of a weather phenomenon, often bringing moisture to the ground.
  3. Tulip: A colorful flower with cup-shaped petals that bloom in spring, growing from bulbs planted in the ground.
  4. Car: A wheeled vehicle with an engine used for transportation on roads, typically with seats for passengers and space for cargo.
  5. Food: Edible substances that provide nourishment and energy for living organisms, typically consumed to sustain life and promote growth.
  6. Tree: A tall, woody plant with a single main stem or trunk, branches, and leaves, often providing shade and producing oxygen.
  7. Fire: A chemical reaction that produces heat and light, typically accompanied by flames and the burning of combustible materials.
  8. Wind: Moving air in the atmosphere, varying in speed and direction, often felt as a breeze or gust.
  9. Hole: An opening or hollow space in a surface, often with a defined boundary, such as a cavity in the ground or a gap in an object.
  10. Animal: Living organisms that are capable of movement, typically having sensory organs and the ability to respond to their environment, including creatures such as mammals, birds, reptiles, amphibians, and fish.

Comment: There are some abstract concepts in those definitions. I’ve highlighted some of the most obvious in yellow. Note that in my conception of abstraction, there are many abstract concepts whose referents are concrete physical things. Organisms are all physical things, but the category, organism, is itself abstract.

Let’s give it one more try.

Physical terms

Define the following words by using only terms about physical objects, processes, and properties:

cat
rain
tulip
car
food
tree
fire
wind
hole
animal

Cat: A furry, four-legged mammal with retractable claws, pointed ears, and a long tail, often kept as a pet, known for its agility and hunting skills.

Rain: Water droplets falling from the sky due to condensation of atmospheric vapor, forming precipitation essential for the water cycle and nourishing plants and ecosystems.

Tulip: A flowering plant with a bulbous root, long slender stem, and colorful cup-shaped petals, cultivated for ornamental purposes in gardens and landscapes.

Car: A wheeled motor vehicle powered by an internal combustion engine or electric motor, used for transportation of people and goods, typically equipped with seats, a steering wheel, and controls.

Food: Edible substances consumed by living organisms to provide energy and nutrients for growth, metabolism, and sustenance, sourced from plants, animals, and other natural resources.

Tree: A woody perennial plant with a single trunk or multiple branches, characterized by a crown of leaves or needles, roots for anchorage and nutrient absorption, and annual growth rings.

Fire: A rapid oxidation process involving combustion of combustible materials, producing heat, light, and various byproducts such as smoke and ash, essential for cooking, heating, and industrial processes.

Wind: The movement of air molecules across the Earth's surface caused by differences in atmospheric pressure, resulting in the displacement of objects and the generation of kinetic energy, often harnessed for power generation and transportation.

Hole: An opening or hollow space in a solid object, created by excavation, erosion, or perforation, varying in size and shape, often used for containment, passage, or drainage.

Animal: A multicellular organism belonging to the kingdom Animalia, characterized by mobility, heterotrophic nutrition, and typically possessing sensory organs, nervous systems, and reproductive capabilities.

Comment: I haven’t read each definition in detail, but I note there are a lot of concepts there that are abstract by the criterion I have in the “No abstract concepts” trial. What we’ve got are a lot of physical phenomena characterized abstractly.

Tuesday, April 2, 2024

ChatGPT explores the metalingual definition of some concepts from science [MTL]

“Metalingual definition” is the term David Hays gave to the linguistic structure and process whereby some term, called the definiendum, is defined by a string of words, the definiens, where the string of words typically takes the form of a story. Charity was his standard teaching example: Charity (the definiendum) is when someone does something nice for someone without thought of reward (the definiens). Any particular story that matches that pattern is an example of charity. In this conception, charity inheres in the entire story; more precisely, in the pattern exhibited by that story. Note however that the definiens need not take the form of a story, though that may be the most perspicuous form. The point is simply that a pattern of words can be used to define a term.

The power of the construction becomes apparent when you realize that the definiens of any term can itself contain metalingually defined terms, each of which has its own definiens, and so on recursively through a large swath of the word stock. Thus in Hays’s standard example, charity, the term reward appears in the definiens. That term is abstract, and so must be linked to a definiens of its own.

From time to time I have investigated ChatGPT’s capacity explicitly to deal with metalingual definition. Here, for example, is a post in which I explore several such constructs with ChatGPT, Abstract concepts and metalingual definition: Does ChatGPT understand justice and charity?

This post collects some examples from science. My typical convention is to set my prompt in boldface and ChatGPT’s response in plain face.

* * * * *

Define gravity.

Gravity is a natural force that attracts any two objects with mass toward each other. The force of gravity is proportional to the product of the masses of the objects and inversely proportional to the square of the distance between them. This is known as the Law of Universal Gravitation, which was first described by Sir Isaac Newton in the late 17th century.

The force of gravity is what keeps planets in orbit around the sun, moons in orbit around their planets, and galaxies held together. It also causes objects to fall toward the ground when dropped.

Gravity is one of the four fundamental forces of nature, along with the strong nuclear force, the weak nuclear force, and the electromagnetic force. It is a non-contact force, meaning that it acts on objects at a distance without coming into direct contact with them.

Einstein's theory of general relativity, developed in the early 20th century, provides a more accurate and comprehensive description of gravity as a curvature of spacetime caused by the presence of matter and energy. This theory explains phenomena such as gravitational lensing and the behavior of black holes that are difficult to understand under the Newtonian theory of gravity.

Define entropy.

Entropy is a measure of disorder or randomness in a system. It is typically represented by the symbol S and has units of energy per temperature (typically joules per kelvin in SI units).

In thermodynamics, which is the branch of physics that deals with the relationships between heat, energy, and work, entropy is a measure of the amount of thermal energy in a system that is unavailable to do work.

The Second Law of Thermodynamics states that in any energy transfer or transformation, the total entropy of a closed system will always increase over time, this is also known as entropy increase principle.

In more general terms, entropy can also refer to the disorder, randomness, or uncertainty in a system. In information theory, entropy is a measure of the amount of uncertainty or randomness in a random variable. In statistical mechanics, entropy is a measure of the number of microscopic configurations of a system that correspond to a given macroscopic state.

Entropy, in the context of Thermodynamics is often referred to as a state function, which means that it depends only on the current state of the system, not on how that state was achieved.

Define atomic bond.

An atomic bond is the force that holds two or more atoms together to form a chemical compound. There are several types of atomic bonds, each with their own characteristics and properties. The main types of atomic bonds are:

  • Covalent bonds: Covalent bonds form when atoms share one or more pairs of electrons. This type of bond is typically found in compounds made up of non-metals.
  • Ionic bonds: Ionic bonds form when an electron is transferred from one atom to another. This type of bond is typically found in compounds made up of a metal and a non-metal.
  • Metallic bonds: Metallic bonds form when atoms share a sea of electrons. This type of bond is typically found in pure metals and alloys.
  • Hydrogen bond: Hydrogen bond is a type of dipole-dipole interaction, a weaker type of force that occurs between a hydrogen atom covalently bonded to a highly electronegative atom such as nitrogen, oxygen, or fluorine and another highly electronegative atom nearby.

All of these bonds involve the sharing or transfer of electrons between atoms, which results in the formation of a more stable configuration known as a chemical compound. The strength of these bonds depends on the nature of the atoms involved, their electronic configurations, and the distance between them.

The properties of the chemical compound formed will be determined by the type of bond and the arrangement of atoms that holds it together.

Define momentum.

Momentum is a measure of the motion of an object. It is defined as the product of an object's mass and velocity. Mathematically, the momentum of an object (p) is given by the formula:

p = m * v

where m is the mass of the object and v is its velocity. The unit of momentum is typically kilogram meters per second (kg·m/s) in the International System of Units (SI).

The momentum of an object is a vector quantity, meaning it has both a magnitude and a direction. The direction of momentum is the same as the direction of velocity. When an object is at rest, its momentum is zero.

Momentum is a measure of the "inertia" of an object in motion, meaning how difficult it is to change an object's speed or direction of motion. The larger the mass of an object or the greater its velocity, the greater its momentum.

In physics, the principle of conservation of momentum states that in a closed system, the total momentum of all objects before and after a collision or other interaction will remain the same, unless acted upon by an external force.

momentum is an important concept in physics, specially in mechanics, that plays a role in understanding how objects move, and it is widely used in fields like physics, engineering, and aeronautics.

Define gene.

A gene is a functional unit of DNA that contains the instructions for the synthesis of a specific protein or RNA molecule. It consists of a specific sequence of nucleotides, the building blocks of DNA, which are located on chromosomes.

Genes are inherited from our parents and play a central role in determining our physical and behavioral characteristics, such as eye color, hair color, height, blood type and many others. Each gene carries specific information for the synthesis of a particular protein, and the collection of all genes in an organism, known as its genome, contains the complete set of instructions for the development and function of that organism.

The proteins made by a gene play a vital role in the body, acting as enzymes, structural molecules, hormones, and other key components. They also help in controlling when and where certain genes are expressed.

Genes can also be studied in the context of evolution, where variations in genes can lead to different traits that can be beneficial or harmful to an organism's survival. Some variations can be passed on to next generations, leading to natural selection and genetic drift.

Mutations, changes in the DNA sequence of a gene, can also occur and have effects on the proteins it codes for and the way the gene is expressed, these mutations can have positive or negative effect, some being benign and others causing diseases.

What is a molecule.

A molecule is a chemical structure composed of two or more atoms that are chemically bonded together. The atoms in a molecule are held together by chemical bonds, which are forces that result from the attractive or repulsive interactions between the electrons of different atoms. Molecules can be composed of any type of atom, and can vary greatly in size and complexity.

The basic building block of a molecule is an atom and the smallest molecule is diatomic (made up of two atoms, such as H2,O2). Larger molecules can be made up of many atoms and can be quite complex, such as DNA and proteins.

Molecules can also be characterized by their physical properties such as boiling and melting point, solubility, density, and others. The properties of a molecule are often determined by the atoms that make it up, the arrangement of those atoms, and the types of bonds that hold the atoms together.

Molecules can be broadly classified into two groups: organic and inorganic. Organic molecules are those that contain carbon, while inorganic molecules are those that do not contain carbon. Organic molecules are found in living organisms and include compounds such as carbohydrates, lipids, nucleic acids, and proteins, while inorganic molecules include compounds such as water, gases, and minerals.

Molecules play an important role in many chemical processes, including those that occur within living organisms, and are essential for the study of chemistry and biochemistry.

Sunday, March 10, 2024

Does ChatGPT Understand the Concept of Tragedy? Symbolic AI & Neural Nets

New working paper. Title above, links, abstract and contents below:

Abstract: Tragedy is an abstract concept. David Hays proposed that abstract concepts can be metalingually defined. A term, such as tragedy, can be given its meaning by a pattern over a string of words. By using a cognitive network to define tragedy Brian Phillips investigated that computationally. In this paper I have taken Brian Phillips’ definition and presented it to ChatGPT in six different trials, where each trial also included a short story. Three stories satisfied the definition and three did not. ChatGPT was asked to indicate whether or not the story satisfied the definition it was given. It was correct in each case.

CONTENTS

Symbolic Computation and Abstract Definition 2
Tragedy Defined 2
ChatGPT Confronts “Tragedy” in Six Trials 4

Saturday, September 2, 2023

Steven Harnad: Symbol grounding and the structure of dictionaries

Stevan Harnad: AI's Symbol Grounding Problem, The Gradient podcast, August 31, 2023

Stevan Harnad is professor of psychology and cognitive science at Université du Québec à Montréal, adjunct professor of cognitive science at McGill University, and professor emeritus of cognitive science at the University of Southampton. His research is on category learning, categorical perception, symbol grounding, the evolution of language, and animal and human sentience (otherwise known as “consciousness”). He is also an advocate for open access and an activist for animal rights.

Outline:

  • (00:00) Intro
  • (05:20) Professor Harnad’s background: interests in cognitive psychobiology, editing Behavioral and Brain Sciences
    • (07:40) John Searle submits the Chinese Room article
    • (09:20) Early reactions to Searle and Prof. Harnad’s role
  • (13:38) The core of Searle’s argument and the generator of the Symbol Grounding Problem, “strong AI”
  • (19:00) Ways to ground symbols
  • (20:26) The acquisition of categories
  • (25:00) Pantomiming, non-linguistic category formation
  • (27:45) Mathematics, abstraction, and grounding
  • (36:20) Symbol manipulation and interpretation language
  • (40:40) On the Whorf Hypothesis
  • (48:39) Defining “grounding” and introducing the “T3” Turing Test
  • (53:22) Turing’s concerns, AI and reverse-engineering cognition
  • (59:25) Other Minds, T4 and zombies
  • (1:05:48) Degrees of freedom in solutions to the Turing Test, the easy and hard problems of cognition
  • (1:14:33) Over-interepretation of AI systems’ behavior, sentience concerns, T3 and evidence sentience
  • (1:24:35) Prof. Harnad’s commentary on claims in The Vector Grounding Problem
  • (1:28:05) RLHF and grounding, LLMs’ (ungrounded) capabilities, syntactic structure and propositions
  • (1:35:30) Multimodal AI systems (image-text and robotic) and grounding, compositionality
  • (1:42:50) Chomsky’s Universal Grammar, LLMs and T2
  • (1:50:55) T3 and cognitive simulation
  • (1:57:34) Outro

The podcast site also has links to Harnad’s webpages and to five selected articles. One of them in particular, about the structure of dictionaries, interested me. Here’s the citatioin, abstract, and a link:

Philippe Vincent-Lamarre, Alexandre Blondin Massé, Marcos Lopes,Mélanie Lord, Odile Marcotte, Stevan Harnad. The Latent Structure of Dictionaries. Topics in Cognitive Science 8 (2016) 625–659. DOI: 10.1111/tops.12211. (Open Access)

Abstract: How many words—and which ones—are sufficient to define all other words? When dictionaries are analyzed as directed graphs with links from defining words to defined words, they reveal a latent structure. Recursively removing all words that are reachable by definition but that do not define any further words reduces the dictionary to a Kernel of about 10% of its size. This is still not the small- est number of words that can define all the rest. About 75% of the Kernel turns out to be its Core, a “Strongly Connected Subset” of words with a definitional path to and from any pair of its words and no word’s definition depending on a word outside the set. But the Core cannot define all the rest of the dictionary. The 25% of the Kernel surrounding the Core consists of small strongly connected subsets of words: the Satellites. The size of the smallest set of words that can define all the rest— the graph’s “minimum feedback vertex set” or MinSet—is about 1% of the dictionary, about 15% of the Kernel, and part-Core/part-Satellite. But every dictionary has a huge number of MinSets. The Core words are learned earlier, more frequent, and less concrete than the Satellites, which are in turn learned earlier, more frequent, but more concrete than the rest of the Dictionary. In principle, only one MinSet’s words would need to be grounded through the sensorimotor capacity to recognize and categorize their referents. In a dual-code sensorimotor/symbolic model of the mental lexicon, the symbolic code could do all the rest through recombinatory definition.

Finally, somewhere latish in the conversation Harnad made an incisive remark about the vexed issue of whether or not LLMs really understand language. The issue, he remarked, is not whether or not they understand language as we do, but how they can do so much without such understanding. YES, a thousand times yes.

He also noted that he enjoys working with, what was it? ChatGPT. So do I, so do I. And I haven’t the slightest suspicion, worry, or hope that it might be sentient. It is what it is.

Thursday, January 5, 2023

Discursive Competence in ChatGPT, Part 1: Talking with Dragons

Version 1, January 5, 2022

Title above, URLs, abstract, contents, and introduction below:

Academia.edu: https://www.academia.edu/94409729/Discursive_Competence_in_ChatGPT_Part_1_Talking_with_Dragons
SSRN: https://ssrn.com/abstract=4318832
Research Gate: https://www.researchgate.net/publication/
366897197_Discursive_Competence_in_ChatGPT_Part_1_Talking_with_Dragons_Discursive_Competence_in_ChatGPT_Part_1_Talking_with_Dragons

Abstract: Noam Chomsky’s idea of linguistic competence suggests a new approach to understanding how LLMs work. This approach requires careful analysis of text. Such analysis indicates that ChatGPT has explicit control over sophisticated discourse skills: 1) It possesses the capacity to specify high-level structures that regulate the organization of language strings into specific patterns: e.g. conversational turn-taking, story frames, film interpretation, and metalingual definition of abstract concepts. 2) It is capable of analogical reasoning in the interpretation of films and stories, such as Spielberg’s Jaws and A.I., and Tezuka’s Astro Boy stories. It must establish an analogy between some abstract interpretive theory (e.g. the ideas of Rene Girard) and people and events in a story. 3) It has some understanding of abstract concepts such as justice and charity. Such concepts can be defined over concepts that exhibit them (metalingual definition). ChatGPT recognizes suitable stories and can revise them. 4) ChatGPT can adjust its level of discourse to accommodate children of various ages. Finally, much of ChatGPT’s discourse seems formulaic in a way similar to what Parry/Lord found in oral epic.

Contents

Introduction: Walking among dragons 2
What is in the rest of this document? 6
Calibration: Understanding a Seinfeld Bit 9
Conversing with ChatGPT about Jaws, Mimetic Desire, and Sacrifice 12
ChatGPT on Spielberg’s A.I. Artificial Intelligence and AI Alignment 23
Extra! Extra! In a discussion about Astro Boy, ChatGPT defends the rights of robots and advanced AI 27
Pumpkins, the Falcon Heavy, and Groucho Marx: High level discourse structure in ChatGPT 30
High level discourse structure in ChatGPT: Part 2 [Quasi-symbolic?] 37
Abstract concepts and metalingual definition: Does ChatGPT understand justice and charity? 42
Does ChatGPT’s performance warrant working on a tutor for children? 53
To the future and beyond 59
Coda: What’s going to be in Part 2: A Framework for Description and Analysis? 65 Appendix: ChatGPT gets confused about Sonnet 129 68

Introduction: Walking among dragons

“Your brain, your job, and your most fundamental beliefs will be challenged by AI like nothing ever before. Make sure you understand it, how it works, and where and how it is being used.” – David Ferrucci

When I first heard about ChatGPT on November 30, 2022, I figured I’d pass on it. After all, it was ultimately based on GPT-3 and I’d already had a little bit of fun with that, albeit through an intermediary. What more could there be? The next day, however, I thought, Why not, it’s free, no? I signed up for an account. I had no particular intentions. I just wanted to test the water.

Total Immersion

I’ve been swimming in it since then. I’ve copied every “conversation” into a text document that is 178 pages long. I can’t tell you how many times I’ve laughed out loud and danced in my seat in reaction to ChatGPT’s response to a prompt.

ChatGPT is more fun than a barrel of monkeys. But it is also work. When I started playing with it I had no specific intentions; I certainly did not intend to write about ChatGPT extensively. I just wanted to poke around. I became, if not hooked, perhaps entranced. I began systematically exploring it, not to find its faults, its weakness, as many are doing, but to test its strengths.

In the process I have changed, though it is difficult to characterize that change. I wouldn’t say that it has added given me any new indeas. Nor has it changed my views on the strengths and weakness of so-called deep learning (DL) technology. It is not an INTELLECTUAL change.

Like many others I have believed DL needs to be augmented by “classical” symbolic processing. I still believe that. I also believed that DL systems need direct interaction with the world if they are to exhibit real “intelligence.” That belief remains rock-solid. Some enthusiats have been saying that DL will take us all the way to full Artificial General Intelligence (AGI) simply by scaling up: more parameters, more data, and more compute. I disagree. Deeper changes are required.

What has changed is my ORIENTATION, my outlook. I have had a glimpse, however limited and provisional, into a new world, a world where we will be working with these “miracles of rare device” (to borrow a phrase from Coleridge) in ways we had not previously imagined. If you want to know what it’s like to drive a car, you can only do it from the driver’s seat. I have taken the driver’s seat and have been systematically exploring ChatGPT.

It is one thing to be amazed by this or that output from ChatGPT. There’s Lots of that going around. I am going beyond that to analyze some of the mechanims of discursive competence, to borrow a term from Noam Chomsky, that enable ChatGPT to function so well. This working paper is a preliminary report on these explorations. I believe that through the careful analysis of ChatGPT’s discursive output we can gain insight into its inner operations, allowing us to improve future technology and to develop benchmarks more tailored to the capacities of emerging LLM technology.

I responded to GPT-3 with a report entitled, GPT-3: Waterloo or Rubicon? Here be Dragons. I have crossed the Rubicon and have been walking among dragons. It is time we get to know them better, to talk with them.

To learn about dragons, describe and analyze them

When I started playing with ChatGPT on December 1, 2022, I had no specific intentions. I wanted to poke around, see what I could see, and then...As I said, I had no specific intentions. I certainly did not intend to spend hours interacting with it to produce a Microsoft Word document currently (1.5.22) containing 61,580 words of transcription – the vast majority from ChatGPT – on 178 pages.

One of the earliest things I did with ChatGPT – not THE first, it was my third session, on December 1, 2022 ¬– was to dialog about Steven Spielberg’s Jaws and the ideas of Rene Girard. I took that and wrote it up for 3 Quarks Daily. Then I had some fun with “Kubla Khan,” quizzed it about trumpets, had a long session about Gojira/Godzilla, and then returned to Spielberg, this time to A.I. Artificial Intelligence. By this time I was developing a feel for how ChatGPT responded. Both the Jaws and the A.I. posts are included in this paper.

I became more systematic, looking for specific things, testing them out. That led to a post with a rather baroque title, “Of pumpkins, the Falcon Heavy, and Groucho Marx: High level discourse structure in ChatGPT,” which I’ve also included in this paper. In that post I advanced the argument that there are parameters in the language model that govern ligher level discourse structures independently of the specific words and strings that realize them.

The alternation pattern is something like this:

A, B, A, B....

That can be repeated as often as one will. The text in the A sections is always drawn from one body of material while the text in the B sections is drawn from a different body of material. That’s the pattern ChatGPT has learned. Where is it in the net? How’s it encoded.

The frame structure is a bit more complicated:

A (B, C, B, C....) A’

The embedded alternation draws on two bodies of material, any two bodies. The second part of the frame, A’, must complement the first, A.

Again, it’s not a complex structure. But it’s not defined directly over particular words. It’s defined over groups of words, placing the groups, not the individual words, into specified relationships in the discourse string.

I then suggested that the patterns I had identified in Jaws and A.I. where similar, but, if anything, more complex.

I had become all but convinced that ChatGPT had explicit control over high-level discourse properties. When humans make statements like those, we take it as obvious that they have some “grammar” of high-level discourse structures. Narratologists, linguists, and psycholinguists study them. But ChatGPT is not a human. It is, shall we say, a machine, a machine that was trained to guess the next word, word after word after word....and so forth, for jillions of texts. All that’s in the resulting model is statistics about those texts. It’s seems to be a “stochastic parrot”, as one well-know paper argued.

Perhaps, in a sense, that is a true. But that is a terribly reductive characterization, and, I have come to believe, all but beside the point. Large language models issue one word at a time for the same reason that humans do: That’s the nature of the communication channel, and tells us relatively little about the device that is pushing words through the channel. LLMs develop rich and complicated structures of parameter weights during the training process. Yes, those structures are statistical in nature, but they are also structures. Perhaps there are aspects of those structures that we can investigate without having to “open the hood” and examine parameter weights.

I made that suggestion in a post, “Abstract concepts and metalingual definition: Does ChatGPT understand justice and charity?”, also included in this paper. Chomsky famously distinguished between competence and performance, where the study of linguistic performance is about the mechanism that produces and understands texts while the study of linguistic competence is about the structure of the texts independent of underlying mechanisms. When I analyze ChatGPT’s output I am investigating its competence. When researchers pop the hood and examine parameter weights, they are investigating performance mechanisms. I further suggest that a better understanding of an LLM’s competence will aid in studying those performance mechanisms by giving us clues about what they are doing.

Nor am I the only one who believes in the value of studying the output of these engines. Others have come to that conclusion as well, though perhaps not quite in those terms. Here is the abstract of a recent preprint from Marcel Binz and Eric Schulz from the Max Planck Institute:

We study GPT-3, a recent large language model, using tools from cognitive psychology. More specifically, we assess GPT-3’s decision-making, information search, deliberation, and causal reasoning abilities on a battery of canonical experiments from the literature. We find that much of GPT-3’s behavior is impressive: it solves vignette-based tasks similarly or better than human subjects, is able to make decent decisions from descriptions, outperforms humans in a multi-armed bandit task, and shows signatures of model-based reinforcement learning. Yet we also find that small perturbations to vignette-based tasks can lead GPT-3 vastly astray, that it shows no signatures of directed exploration, and that it fails miserably in a causal reasoning task. These results enrich our understanding of current large language models and pave the way for future investigations using tools from cognitive psychology to study increasingly capable and opaque artificial agents.[1]

My methods are different, but my purpose is the same, “to study increasingly capable and opaque artificial agents” and thus to render them less opaque. The insights we gain thereby will aid us to improve the capabilities of the next generation of artificial agents. 

* * * * *

[1] Marcel Binz and Eric Schulz, 2022. “Using Cognitive Psychology to Understand GPT-3,” PsyArXiv, June 21, 2022. doi:10.31234/osf.io/6dfgk.

Wednesday, December 28, 2022

Felix Culpa [the Fortunate Fall] – To justify the ways of God to man [ChatGPT, theologian]

While I was quizzing ChatGPT on its knowledge of basic legal concepts – ChatGPT the legal beagle: Concepts, Citizens United, Constitutional Interpretation – I decided, Why not? Let’s head toward Shakespeare – which turned out to be a disaster, did it not? So I decided to ease into it by way of an interesting bit of medieval doctrine, The Fortunate Fall, felix culpa (NOT Felix catus). Why felix culpa? Because, I argued way back in the ancient days, that’s the pattern that underlies that Shakespeare sonnet we’re stalking.

Let’s get to it!

* * * * *

Moving away from the law, what about Felix Culpa?

Felix culpa is a Latin phrase that means "happy fault" or "fortunate fall." In Christian theology, it refers to the belief that the fall of man in the Garden of Eden, as described in the Bible, was actually a good thing because it allowed for the possibility of salvation and redemption.

According to this belief, God allowed Adam and Eve to eat the fruit of the tree of knowledge of good and evil, despite having warned them not to do so, in order to create the opportunity for humanity to be saved through the death and resurrection of Jesus Christ. This belief is based on the idea that the fall of man was necessary for the fulfillment of God's plan for the salvation of humanity.

The phrase "felix culpa" is often used to refer to the idea that something that seems like a misfortune or an obstacle can actually be a blessing in disguise, as it can lead to unexpected opportunities or positive outcomes. This concept is similar to the idea of "making lemonade out of lemons," which suggests that it is possible to find value or benefit in difficult or adverse situations.

Is the idea of Felix Culpa relevant to Milton’s Paradise Lost?

Yes, the idea of felix culpa plays a significant role in John Milton's epic poem Paradise Lost. In the poem, Milton explores the biblical story of the fall of man, as described in the book of Genesis, and presents his own interpretation of the events and their significance.

Milton's portrayal of the fall of man in Paradise Lost is rooted in the idea of felix culpa. According to Milton, the fall was not a tragic event but rather a necessary step in the fulfillment of God's plan for humanity. In Milton's view, the fall allowed for the possibility of redemption and salvation through the death and resurrection of Jesus Christ, and it ultimately led to the triumph of good over evil.

Throughout the poem, Milton presents the fall of man as a complex and nuanced event that has both positive and negative consequences. On the one hand, the fall brings suffering and death into the world, and it causes humanity to be separated from God. On the other hand, the fall also gives humanity the freedom to choose between good and evil, and it allows for the possibility of redemption and salvation.

In this way, Milton's portrayal of the fall in Paradise Lost reflects the idea of felix culpa, as he presents the fall as a "happy fault" that ultimately leads to a greater good.

What does it mean "to justify the ways of God to men?"

"To justify the ways of God to men" is a phrase that appears in John Milton's epic poem Paradise Lost, in which the narrator (the archangel Michael) tells the story of the fall of man and the biblical account of the creation of the world. The phrase appears in the following lines:

"To justify the ways of God to men
Doth God Himself in human form display,
Revealing His own Son, and that His Son
By His own choice subjected to the law
Which he Himself had made."

In this passage, the phrase "to justify the ways of God to men" refers to the idea of explaining or making sense of God's actions or decisions to human beings. The phrase suggests that God's ways are sometimes difficult for humans to understand or comprehend, and that it is necessary for God to reveal Himself and His intentions in order to help humans understand and accept them.

The phrase has become famous and is often quoted or referred to as a way of expressing the idea that it can be challenging for humans to understand or reconcile the actions or decisions of a higher power, especially when they seem difficult or mysterious. It is often used in discussions about religion, faith, and the relationship between humans and a higher power.