First, I talk about now natural language is its own metalanguage and that allows them to define new works in terms of existing ones. Then I discuss the concept of justice in terms of mechanism of metalingual definition proposed by David Hays some years ago. I conclude with some remarks on interpretability in view of Anthropic’s recent research on features.
The metalingual function of language in defining word meaning
In a famous essay published in 1960, “Linguistics and Poetics,” Roman Jakobson listed six functions of language. While the essay focused on the poetic function, as the title indicates, I’m interested in a different function, which he called the metalingual function:
A distinction has been made in modern logic between two levels of language: “object language” speaking of objects and “metalanguage” speaking of language.10 But metalanguage is not only a necessary scientific tool utilized by logicians and linguists; it plays also an important role in our everyday language. Like Molière's Jourdain who used prose without knowing it, we practice metalanguage without realizing the metalingual character of our operations. Whenever the addresser and/or the addressee need to check up whether they use the same code, speech is focused on the code: it performs a METALINGUAL (i.e., glossing) function.
In the process of explicating that function Jakobson pointed out that it can be used to define words, noting that “any process of language learning, in particular child acquisition of the mother tongue, makes wide use of such metalingual operations.”
Not all words get their meaning in that way. Many words have their meanings grounded in sensorimotor experience. We would like to know what percentage of words have their meanings grounded in sensorimotor experience and what percentage have their meanings grounded in other words. In 2016 Steven Harnad and his colleagues published an article investigating this problem, “The Latent Structure of Dictionaries.” They examined the structure of the vocabularies in two dictionaries, one with roughly 47,000 words and the other with roughly 69,000 words. They found that a large majority of the words were defined in terms of a relatively small number of words defined in terms of sensorimotor features (p. 649):
So in our view the mental lexicon is itself hybrid—a dual-code representational system consisting of learned sensorimotor feature (affordance) detectors for the grounding words (and any later hybrid words) plus recombinatory and purely symbolic (i.e., verbal) definitions and descriptions for the referents of the words that are learned through words alone.
More recently Briony Banks, Anna M. Borghi, Raphaël Fargier et. al. reviewed the literature on abstract concepts, “Consensus Paper: Current Perspectives on Abstract Concepts and Future Research Directions.” They noted that “many theories have also argued that our understanding and representation of abstract concepts relies more on language than the sensorimotor dimension, and particularly linguistic distributional relations.”
Given that LLMs have been constructed in an environment consisting entirely of words, the apparent fact that most words are defined in terms of other words seems highly salient.
Metalingual definition and the concept of justice
Back in the 1970s David Hays was interested in the idea that words can be used to define the meaning of other words. He talked specifically of metalingual definition. He used charity as his prototypical example: Charity is when someone does something nice for another without thought of reward. Any story that exhibits that pattern of relationships between its actors and their actors, such a story is about justice. The concept inheres in that pattern of relationships as a whole and not in any of the individual components of the pattern.
Notice that the definition itself contains an abstract concept, reward. Taken as a computational mechanism, which was his point, metalingual definition is thus recursive, allowing definitions to be nested within definitions. One of Hays’s students Brian Phillips, implemented the idea in his doctoral dissertation using tragedy as his example. I recently took the definition that Brian Phillips used and used it to test ChatGPT, which had no trouble applying it to specific examples and determining whether or not they met the conditions set forth in the definition.
With this before us I ask: What is justice? That is to say, what kind of a thing is justice? It’s a virtue, no? Yes, but I’m looking for something even more general, more abstract. It’s a concept, and idea, no? Of course it is. And just what are those things? Philosophers have been pondering to question for years. Cognitive scientists have been asking that question as well. When David Hays proposed that abstract concepts can be defined by stories, he was proposing an answer to that question. Abstract concepts, such as justice, are defined by relationship among words.
* * * * *
I’ve devoted a great deal of attention to ChatGPT’s ability to deal with metalingual definition. Justice is one of the first concepts I investigated, back in December of 2022. I’ve continued to investigate that concept. I have appended my most recent session to the end of this post.
That investigation has three parts. First, I ask it to tell me two stories involving justice and I specify that it should not use the word “justice” anywhere in the stories. The point of that restriction is to make it clear that the meaning of the term does not reside in the word itself. The stories exhibit justice, but do not name it. Note that in a second session, which I’ve placed in a second appendix, I give ChatGPT the two stories, one after the other, and ask it what they’re about. It realizes that they are about justice.
After asking ChatGPT to tell me stories involving justice I ask it to define the term. The first definition is fairly long and has five numbered points, each specifying a particular kind of justice. So I ask it for a single paragraph and then a single sentence. It provides both. Note that both the long definition and the single paragraph definition begin with pretty much the same information that the single sentence contains.
Finally, I ask ChatGPT to explain the relationship between the stories and the definition, which it does in a paragraph of 112 words. Here’s the first sentence: “The relationship between the definition of justice and stories about justice lies in the way these narratives illustrate and bring to life the abstract principles of fairness, equity, and moral rightness.”
Interpretability of LLMs
What does this have to do with the interpretability of large language models? To a first approximation, it seems to me that LLMs are about the relationships between words. The transformer is presented with strings of words during training and, in the process of making those predictions, constructs a complex model of how words are related to one another.
Thus we might say that justice is a certain pattern of relationships among words. But what pattern? The pattern that gives us stories, stories which may not even contain the word “justice” or the pattern that gives us definitions and could, I assume, produce essays and even books if necessary? Those are distinctly different patterns of relationships; one might even think about them as being orthogonal, at least informally. One pattern is about justice in the context of story and the other is about justice in the context of define. Finally, what about the pattern that explication the relationship between the stories and the definitions?
In Scaling Monosemanticity, researchers at Anthropic identified features in Claude 3 Sonnet, where features are understood to be “directions in their activation spaces.” In their discussion, they note “that features often respond to both abstract discussion and concrete examples of a concept,” which is certainly something that I’d expect to be the case. One thing that bothers me about the discussion is that there is no sense of the model as capturing relationships between words. Given that these features are very abstract objects it’s not clear to me just what that misgiving means, but I worry that the concept of features invites reification.
A digression into neuroscience: Some years ago I had quite a bit of correspondence with the late Walter Freeman, who did pioneering work in thinking about the brain in terms of chaos theory and complex neurodynamics. He believed that percepts and concepts were located in populations of neurons rather than single neurons. I’m deeply sympathetic to that view, and have been ever since I read Karl Pribram on neural holography. Nonetheless I asked him about visual neurons that had very complex activation properties, such as a monkey’s paw or an image of Bill Clinton. Don’t such examples lend support to the idea of a so-called “grandmother cell”? His reply was no, they don’t. In such a complex system, you’re bound to find individual neurons with all sort of odd response characteristics.
I feel a bit like that with these features. While they don’t seem to be individual neurons, it’s not clear what they are. Robert_AIZI has expressed a similar reservation. Thus he has noted:
I think Anthropic successfully demonstrated (in the paper and with Golden Gate Claude) that this feature, at very high activation levels, corresponds to the Golden Gate Bridge. But on a median instance of text where this feature is active, it is "irrelevant" to the Golden Gate Bridge, according to their own autointerpretability metric! I view this as analogous to naming water "the drowning liquid", or Boeing the "door exploding company". Yes, in extremis, water and Boeing are associated with drowning and door blowouts, but any interpretation that ends there would be limited.
Just what IS this feature?
I’m not surprised that with judicious and determined poking around we can find interpretable “features” in these models. But whether or not we’re carving LLMs at their joints, to invoke Plato’s famous remark, that’s another matter.
It’s early days.
Appendix 1: ChatGPT on justice: stories and definitions
What follows the asterisks is a single session with ChatGPT. As always, my prompts are set in bold-face while ChatGPT’s responses are plain-face, except for those strings that ChatGPT sets in bold-face.


