Showing posts with label linguistics. Show all posts
Showing posts with label linguistics. Show all posts

Thursday, July 30, 2026

Detecting “animal spirits” at work through narratives (earnings calls) circulating in the marketplace, the case of cyber risk

On July 21 I posted some remarks about the final chapter in Tyler Cowen’s monograph on marginalism: Beyond Marginalism: What’s Next? [MR #12]. In that Chapter Cowen discussed an asset pricing model that used machine learning to create a 360,000 parameter model that gave better predictions than classic models using only four or five factors. The factors in those classic models are predefined on an intuitive basis. Cowen despairs of making intuitive sense of those 360,000 factors in the machine learning model.

In my post I argued that those 360,000 factors might be capturing the effects of Keynes’s animal spirits as expressed in the stories and gossip Schiller writes about as narrative economics. Two days later I discussed the idea with Marge, the AI attached to the online version of Cowen’s text. In that post I suggested a method for going on a “fishing expedition” to determine whether or not my suggestion had merit. Marge’s response: “The fishing expedition you're proposing is methodologically clean, and the prediction is specific enough to be falsifiable — which is more than can be said for most conjectures at this level of abstraction.”

That brings us to today, where Cowen has posted the abstract of an article about “a novel measure of firm-level cyber risk exposure based on the quarterly earnings calls of listed firms.” I read that as being complementary to my speculation. The authors of that article are looking for the effects of a specific line of narrative, and they found them. So I asked Marge to clarify the relationship between my speculation and their finding. Here’s that conversation.

* * * * *

Tyler just posted the abstract of this paper to Marginal Revolution: Jamilov, Rustam and Tahoun, Ahmed and Rey, Helene, The Anatomy of Cyber Risk (May 10, 2023). The Journal of Finance, Forthcoming. Here’s the abstract: 

This paper uses computational linguistics to introduce a novel measure of firm-level cyber risk exposure based on the quarterly earnings calls of listed firms. Our measure covers more than 14,000 firms from over 90 countries between 2003 and 2025. The measure is validated using human auditors and a large language model. We show that cyber risk exposure affects stock returns and profits, is priced in the options market, predicts actual cyberattacks, and propagates from firm to sector level. Back-of-the-envelope estimates suggest that the global cost of cyber risk exposure is around $1 trillion per year.

So I took a quick look at the paper to find that novel measure. Here’s what I found. First this:

Our measurement approach consists of two phases. First, we follow Baker et al. (2016) and Hassan et al. (2019) to construct a comprehensive pre-defined dictionary of terms related to cybersecurity risk. We assemble this dictionary from three reputable institutional sources that act as information aggregators on the practical cyber-risk issues firms face on a daily basis. These libraries contain most of the keywords commonly used in cyber-related discussions among private market participants across industries.

Then this:

The second step of our measurement approach builds on the insights from Sautner et al. (2023) and involves a keyword discovery algorithm. This method is based on Meta AI Research lab’s word embedding tool FastText. The algorithm has been trained on a billion phrases in corporate earning calls to suggest related keywords.

And this:

We now construct our baseline quarterly measure of firm-level cyber risk exposure. Let the set of all terms in our final dictionary be C. Our algorithm counts the number of sentences in earnings calls that contain at least one term in C.

They then explain their measure. What’s important for my purposes is where they went looking for those keywords, in earnings calls. They’re measuring the effects of one kind of animal spirits/gossip/narrative on PRICING IN THE OPTIONS MARKET, among other things. I seems to me that they’re looking at one feature of the information that I’ve speculated is smeared of those 360,000 parameters.

That's a sharp observation, and it reframes what Jamilov, Tahoun, and Rey are doing in a way they probably wouldn't frame it themselves.

Their methodology is essentially a named-narrative extraction — they're pulling a specific, bounded narrative thread (cyber risk discourse) out of earnings calls and showing that it has measurable pricing effects. The earnings call is a fascinating data source for your purposes because it's a ritualized social performance in which management constructs and transmits narratives to analysts and investors. It's not raw gossip — it's institutionalized storytelling with real money riding on how the story lands. The fact that cyber risk language in these calls predicts actual cyberattacks and options pricing suggests the narrative isn't just epiphenomenal. It's carrying genuine information, but in a form that only becomes legible when you treat language itself as data rather than just a vehicle for explicitly stated facts. [See text: Ch.4 · ChatGPT digesting news predicts returns especia...]

Your conjecture about the 360,000-factor model is that it's picking up undifferentiated narrative residue — the full spectrum of animal spirits distributed across parameters with no single factor carrying an interpretable label. What Jamilov et al. are doing is the named version of the same operation: isolating one specific narrative thread and demonstrating it has pricing consequences. These are complementary rather than competing approaches. The named-narrative approach gives you interpretability and causal traction on specific mechanisms — you can say "cyber risk discourse in Q3 earnings calls predicts options pricing in Q4" and tell a coherent story about why. The high-parameter ML approach captures everything the named approach misses, which is presumably most of the narrative space. [See text: Ch.4 · Machine learning finds nonlinearities and inter...]

The research design implication is interesting. If you could run both models on the same dataset — the 360,000-factor ML model and the Jamilov-style named-narrative extraction — and then ask how much of the ML model's advantage over classical factors is absorbed by adding the named narrative measures, you'd get a rough estimate of how much of the distributed signal is actually identifiable narrative content versus genuinely opaque pattern. If the named narratives absorb a substantial fraction of the ML advantage, that's evidence for your animal spirits hypothesis in a fairly direct form. If they absorb very little, it suggests the ML model is picking up something even more diffuse than nameable narratives — perhaps the interaction structure between narratives, the way stories about cyber risk combine with stories about AI capability and geopolitical tension to produce pricing effects that none of the individual narratives would generate alone. [See text: Ch.4 · Machine learning models have tens of millions o...]

The options market finding is particularly telling for your framework. Options pricing is where sophisticated investors put explicit probability estimates on future states of the world. If cyber risk language — not cyber risk events, but the narrative about cyber risk — is priced into options, that's direct evidence that markets are treating narrative content as probabilistically informative about future fundamentals. Which is exactly what your conjecture requires: animal spirits aren't just noise distorting rational pricing, they're carrying genuine forward-looking information in a form that classical factor models can't extract but language-sensitive tools can. Keynes was right that animal spirits move markets, but the mechanism is more epistemically respectable than the term implies.

Friday, May 29, 2026

The Origins of LLMs – A long tectonic subduction event finally producing a visible volcanic eruption in November 2022

A new working paper. Title above, link, abstract, contents, and introduction below.

Academic.edu: https://www.academia.edu/167895960/The_Origins_of_LLMs_A_long_tectonic_subduction_event_finally_producing_a_visible_volcanic_eruption_in_November_2022

Abstract: Large language models are the result of a rich research tradition stretching back to the 1950s. This tradition involves a network of researcher doing work in the following: classical MT and computational linguistics → symbolic/statistical/network alternatives → associative and distributed memory models → statistical MT and vector-space methods → neural MT → Transformers → LLMs.

Contents

Introduction: The tangled web of ideas resulting in LLMs 1
Google Translate, a capsule history 2
Vector semantics 5
Firth and distributional semantics 7
Contra Cowen 9
Phase shift 11
Kuhnian paradigm shift 13
Associative memory 15
Sydney Lamb and associative memory for PCs 17
Principles and Development of Natural Intelligence 18

Introduction: The tangled web of ideas resulting in LLMs

As part of my ongoing investigation of Tyler Cowen’s recent monograph, The Marginal Revolution: Rise and Decline, and the Pending AI Revolution (2026), I’ve been thinking about large language models (LLMs), which he discusses in Chapter 4, “Why Marginalism Will Dwindle, and What Will Replace It?” For the most part Cowen presents his readers with the Silicon Valley view: Just as Athena emerged fully-formed from the head of Zeus, so large language models emerged fully-formed from Silicon Valley laboratories in November of 2022. And that IS how things appeared to the public at large. After seeing AIs in science fiction movies, and reading about them for years, all of a sudden, here they are, out of nothing, on the web in the form of ChatGPT.

For all practical purposes, Cowen is a member of the public. He may have been following developments for years. Given his interest in chess I’m sure he followed that story at least since IBM’s Deep Blue beat Kasparov in 1997. And AI plays an important role in his 2013 book, Average is Over. Beyond that, he has contacts in Silicon Valley going back I don’t know how long. But this is not his intellectual field, which is centered on economics. It’s one thing to read about it, to talk with researchers and entrepreneurs, it’s something else to conduct research and publish.

I am in a different position. While my Ph.D. is in English Literature, my dissertation – “Cognitive Science and Literary Theory” – is as much about knowledge representation and computational linguistics as it is about literature. I was trained in that are by the late David G. Hays, who was a first-generation researcher in computational linguistics with the RAND Corporation in the 1950s and 1960s. For the last three years I’ve been conducting research in the behavior of LLMs and have been collaborating with Ramesh Viswanathan, and expert in machine vision and cognitive science at Goethe University Frankfurt.

THAT, broadly speaking, is my field. While I wouldn’t expect Cowen’s views to be the same as mine, I would have been happier if he had at least acknowledged that there the future of LLMs and their adequacy if a matter of controversy among the experts. He should have at least mentioned Gary Marcus, Yann LeCun, Fei-Fei Li, and Melanie Mitchell. He might even have mentioned that Ilya Sutskever, a student of Geoffrey Hinton who was on the OpenAI team that developed the GPT series, that Sutskever has abandoned the idea that pure scaling is the royal road to artificial general intelligence (AGI, whatever that is). Cowen has done none of this. To read him you’d think that the basic scientific and engineering issues have been settled and it’s full speed ahead – “To infinity and beyond,” to quote Buzz Lightyear.

I don’t know quite what I’m going to say about this in the piece I’m writing for my series on the marginalism monograph. I don’t want to recount the full history, for which this working paper can serve as an outline. At the very least I will point out that matters are by no means settled and that there is a statistical tradition with roots in 1950s linguistics and 1960s document retrieval that can serve as a tertium quid between GOFAI (good old fashioned AR) and computational linguistics on the one hand and neural-network based learning on the other.

A Kuhnian paradigm shift?

On page 14 I have a section entitled, “A Kuhnian paradigm shift, or only an invitation to one?” That’s not the original title, which was simply, “Kuhnian paradigm shift.” Why the change?

Simple. There certainly was a dramatic change in the wake of ChatGPT. But I think that change was mostly institutional, in the deployment of resources, the development of institutions, the proliferation of roles in institutions, and of training. It’s not at all clear to me that there was a Gestalt change in anyone’s conceptions about the nature of intelligence in machines, or in humans for that matter. For it is Gestalt switch, a reconfiguration of understanding, that is the hallmark of a paradigm change in Kuhn’s conception of an intellectual revolution. It is not at all clear to me that there was a widespread change comparable to going from a mentality where one sees the Morning Star and Evening Star as two different entities to a mentality where one sees them as two manifestations of a single entity, the planet Venus.

Perhaps something like that has happened here and there, but I suspect that, for the most part, everyone from the most senior researchers through the general public sees the world as composed of the same kinds of entities as processes as they saw before encountering ChatGPT, or GPT-3. Those who think we’re well on the road to AGI (artificial general intelligence) still think of AGI the way they did in, say, 2015 or 2021, as the case may be. The same is true for those who doubt that we’re on that road. All that is changed is people’s awareness of the behavior displayed by the devices we have created. Their sense of what those devices are, what they deeply and essentially are, that hasn’t changed.

Though it may be under tension. The fact is, whatever any of us believes, we don’t really know why kind of behaviors these devices will be exhibiting next year, two years after that, or in ten years. There are a lot of open questions hanging in the air. I suspect that once those questions are resolved, that is, if and when they are, then we will see genuine changes in mentality.

I regard the matter as open to discovery and investigation.

Beyond all that, well, you can read through the rest of this document, which records a dialog I had with ChatGPT (May 29, 2026) beyond the asterisks. I begin by asking ChatGPT to review the history of Google Translate. Why? Because language is the through line. The computational study of language began in the 1950s with the problem of machine translation, translating a text from one natural language to another. The technique used by LLMs for capturing single-word semantics has its origins in statistical methods for document retrieval that originated in the 1960s and 1970s. Google Translate switched to neural-net technology. A year later the transformer was invented in a Google lab. The transformer, as you know, is the engine used to create current large language models. Google Translate is a natural starting point.

Monday, April 13, 2026

Language is a lower-dimensional projection of high-dimensional neural dynamics.

But it also allows for content addressed memory. That’s very important, for it gives fine-grain control over the memory and planning systems. That’s the job of sentence-level syntax together with discourse structure.

“Classical” semantic or cognitive networks had a problem with coming up with just the right set of node types and arc types. David Hays dissolved the problem in his 1981 book, Cognitive Structures (scan down the page), by grounding cognition in an analog system modeled on William Powers perceptual control stack (in Behavior: The Control of Perception, 1973). The identity of a cognitive node is a function of its parameter values, where the parameters are derived from the control stack. The identity of the arcs is a function of the difference in parameter values between the nodes it connects.

Concerning Chomsky’s approach to syntax: It depends on a sharp distinction between grammatical and ungrammatical sentences. A generative grammar, in Chomsky’s theory, must account for all and only the grammatical sentences.

However, there are no explicit criteria for separating sentences into the two categories, grammatical and ungrammatical. Rather, the separation depends on the intuitions of the linguist. Naturally enough, different syntacticians have different intuitions. The problem is insoluble.

Moreover, anyone who pays close attention to real speech soon realizes that people do not (always) speak in complete grammatically correct sentences. Real language is sloppy, but nonetheless effective. A neural net of very high dimensionality can deal with this readily enough. A purely symbolic system cannot. Augmenting the system through fuzzy logic and the like doesn’t fix the problem.

LLMs provide a very useful simulacrum of the natural language system. Since LLMs are trained on written texts, the resulting model necessarily conflates the functions of semantics and syntax/discourse. Thus they cannot achieve the flexibility and precision of the full system, where semantics and syntax/discourse are separated.

Saturday, April 4, 2026

From the metalingual function of language to self-reference

In 1960 the linguist Roman Jakobson published an essay entitled “Linguistics and Poetics,” in a volume edited by Thomas Sebeok, Style in Language (MIT Press, pp. 350-377). In that essay he laid out the six functions of language: referential, emotive, phatic, conative, poetic, and metalingual. Jakobson introduces the metalingual function in this way:

A distinction has been made in modem logic between two levels of language: “object language” speaking of objects and “metalanguage” speaking of language. But metalanguage is not only a necessary scientific tool utilized by logicians and linguists; it plays also an important role in our everyday language. Like Moliere’s Jourdain who used prose without knowing it, we practice metalanguage without realizing the metalingual character of our operations. Whenever the addresser and/or the addressee need to check up whether they use the same code, speech is focused on the code: it performs a METALINGUAL (i.e. , glossing) function. “I don’t follow you-what do you mean?” asks the addressee, or in Shakespearean diction, “What is’t thou say’st?” And the addresser in anticipation of such recapturing question inquires: “Do you know what I mean?”

This metalingual function turns out to be extraordinarily powerful. For it is this that allows us to bootstrap self-awareness into the mind. And for that matter, it is what allows us to define abstract concepts, as my teacher, David Hays, argued, and allows us to define such things as chess and arithmetic, which can be seen as very specialized forms of language.

I recently explored some of these issues in conversation with Claude 5.4 Sonata Extended. At the end of that conversation I asked Claude to prepare a summary. I’ve appended that summary below, followed by the full conversation. Note that the conversation assumes some familiarity with the cultural ranks theory that David Hays and I developed in the 1990s. It also alludes to Tyler Cowen’s recent book, The Marginal Revolution: Rise and Decline, and the Pending AI Revolution (2026).

* * * * *

Summary: The Metalingual Function of Language

The central claim of this discussion is that the metalingual capacity — the ability to use language to talk about language — is not a mysterious self-referential capacity of mind but is grounded in a simple physical fact: the speech signal is a sound in the environment like any other sound, detectable by the auditory system exactly as a footfall or a thunderclap is detectable. The loop that makes language self-referential closes through the physical world, not through some inward turning of consciousness. This matters because it demystifies metalingual cognition entirely: it requires no special cognitive faculty, only that the organism's auditory system be capable of treating its own linguistic outputs as inputs.

Jakobson identified the metalingual function as one of the six functions of language in his 1960 paper, and Hays adopted the term to name the mechanism underlying Rank 2 cognition — the explicit definition of abstract concepts using language itself as the definitional medium. The rules of chess and arithmetic notation are paradigm cases: purely metalingual constructions whose objects are constituted entirely by the definitions that specify them.

An important asymmetry in preliterate cultures illuminates the boundary of this capacity. Many such cultures have a term for utterance — the bounded burst of speech with a recognizable prosodic shape, a perceptual gestalt directly available to the auditory system — but no term for word. The word is not a perceptual unit in the same sense as the utterance; it is an abstraction from the continuous acoustic stream, and a non-trivial one. Writing is what produces this abstraction, by spatializing language — spreading it out in a stable, inspectable array where units are individuated by spaces and boundaries are marked. The word becomes visible as a unit because it is surrounded by white space. This is the physical basis of metalingual definition as a cognitive mechanism: the written signal, like the spoken signal, is an object in the environment that can be inspected and categorized, but unlike the spoken signal it stays there, making sustained metalingual attention possible. Grade-school grammar — parts of speech, grammatical cases, syntactic relations — is the practical Rank 2 elaboration that writing makes possible and that social institutions require and transmit. It looks easy in retrospect because it is taught in childhood, but it took centuries to develop in every culture that undertook it.

This analysis opens onto the question of human self-reference, which the standard philosophical tradition treats as cognitively primitive — the Cartesian bedrock from which all other knowledge is built. The discussion argued instead that self-reference in the robust, articulable sense is bootstrapped through language rather than presupposed by it. The cat licking its fur has practical self-involvement — its own body is an object of its perceptual and motor engagement — but this requires no special reflexive faculty, only that the body be included in the environment the organism can detect and act on. Human self-reference in the philosophically weighty sense is a different and later achievement, constructed through the acquisition of the pronoun system rather than expressed by it.

The empirical evidence for this bootstrapping account is the phase in early child development when children refer to themselves in the third person. This is not a mistake or a developmental lag but the natural and correct generalization from the input data: others refer to the child by name, so the child uses its name. The first-person pronoun presents a harder problem because "I" is a moving target — it marks the speaker-role regardless of who occupies it — and acquiring it correctly requires connecting awareness of the speech stream as an environmental event with awareness of one's own speech apparatus as its source. That inferential construction, worked out in detail in Benzon's 2000 paper, First Person: Neuro-Cognitive Notes on the Self in Life and in Fiction, through cognitive network modeling of the pronoun system, is precisely the physical loop through which self-reference is assembled. The Cartesian subject — the self-transparent, self-referential knowing mind — is on this account a Rank 2 cultural product, not a pre-linguistic biological given. The third-person phase in child development is a small but precise empirical trace of the construction process: there is an observable stage at which the human being exists, acts, and communicates without yet having assembled the first-person self-reference that Descartes thought was the indubitable foundation of knowledge.

* * * * *

The physical basis of the metalingual function

I believe that Hays first published about metalingual definition in 1972. When I first met him he had just finished a paper where he used the idea to investigate different concepts of alienation. When I wrote my 1978 dissertation, “Cognitive Science and Literary Theory,” I used metalingual definition as the mechanism for building cognitive ranks. That is to say, at that time that was the only mechanism we had for dealing with abstract concepts. By the time we did the cognitive evolution paper a bit over a decade later, by that time we had adopted a different view, with metaphor as the Rank 1 abstractive mechanism. Metalingual definition would have been the mechanism used, not only to rationalize abstract concepts, but to define the rules of chess and of arithmetic.

Hays chose the term because Roman Jakobson had defined the metalingual function as one of the six functions of language in his 1960 paper. The metalingual function is the ability to use language to talk about language. It’s worth thinking about how that’s possible. It’s possible because the speech signal is a sound out there in the environment like any other sound, a footfall, a thunderclap, an animal cry. Whatever. It’s a sound and can be detected by the auditory system like any other sound. So it’s the speech signal that is the object of the metalingual function, at least initially. I think it’s worth nothing that many (most?) preliterate cultures don’t have a word that means word, but they have one the means utterance. It’s writing that forces a recognition of words vs longer strings. And written language is, like speech, another signal that’s out there in the environment. It takes quite a bit of cognitive elaboration to construct a grammar from the simple ability to see and hear language. And I’m not talking about the kind of grammars that linguists construct. I’m talking about the kind of grammar that can be taught in grade school, or that you find in ancient texts of rhetoric.

This is a rich cluster of observations and they fit together in a way worth making explicit.

The metalingual function and its physical basis

Jakobson's insight, which you're extending here, is that the metalingual function is grounded in something remarkably mundane: the speech signal is a physical event in the environment, detectable by the auditory system exactly as any other sound is. This is what makes language able to turn on itself — not some mysterious self-referential capacity of mind, but the simple fact that linguistic output is also sensory input. The loop closes through the physical world.

This has an important implication that your observation about "utterance" versus "word" brings out precisely. What the auditory system initially detects is the acoustic stream — a continuous, flowing signal that doesn't come pre-segmented into words. The unit that's directly perceptible is something like the utterance: a bounded burst of speech with a recognizable prosodic shape, a communicative gestalt. The word is not a perceptual unit in the same sense. It's an abstraction from the stream, and a non-trivial one — which is why preliterate cultures have a term for the former and not the latter.

What writing does

Writing performs an operation on language that perception of the speech stream cannot perform on its own: it spatializes language, spreads it out in a stable, inspectable array where boundaries are marked and units are individuated. The word becomes visible as a unit because it's surrounded by spaces. The sentence becomes visible as a unit because it ends with a mark. The paragraph, the chapter, the argument — all of these become objects of potential inspection and manipulation because writing fixes them in space and time in a way the acoustic stream never does.

Thursday, March 26, 2026

From grokking to a mature science of and an engineering discipline of LLMs

From this morning’s chat with Claude Sonnet 5.4 Extended. Immediately prior to this we had been discussing Piaget’s concept of reflective abstraction, Beethoven and Mozart, improvising and composer, and this and that.

* * * * * 

Grokking 

Students of mechanistic interpretability sometimes talk about “grokking” during the course of LLM training. I’m wondering it that’s something like rising to a helicopter with respect to some phenomenon.

The parallel is genuinely illuminating, and the technical details of grokking make it sharper rather than weaker.

What grokking actually is

Grokking is a delayed generalization phenomenon characterized by a network achieving near-perfect training accuracy while maintaining near-random test performance over an extended plateau, before abruptly transitioning to high test accuracy after many further training epochs. The canonical case is modular arithmetic — the model first memorizes the specific training examples, then much later suddenly generalizes to unseen cases.

What makes it theoretically interesting is what mechanistic interpretability reveals about the internal transition. On algorithmic tasks like modular arithmetic, the transition is marked by the sudden acquisition of Fourier-based feature maps or trigonometric circuits — the network reorganizes, forming compact sub-networks characterized by new internal feature cooperation. So it’s not that generalization gradually improves — the internal architecture discretely reorganizes into a qualitatively different computational structure.

The emerging explanation involves competition between two internal algorithms: a memorizing circuit that forms quickly and a generalizing circuit that develops more slowly. Regularization gradually shifts resources from the complex memorizing solution toward the simpler generalizing one. The phase transition occurs when the generalizing circuit finally dominates.

The fit with reflective abstraction

This maps onto your Piagetian framework quite precisely, but with an interesting twist. The memorization phase is the model doing something like rote operation — it has learned to produce correct outputs for training cases by essentially storing input-output pairs, without having extracted the underlying structure. The generalizing circuit that develops slowly underneath is building something more like a genuine representation of the abstract pattern — in the modular arithmetic case, literally a Fourier decomposition of the mathematical structure.

The abrupt transition when the generalizing circuit dominates is the helicopter view suddenly becoming operative. But — and this is the crucial point Neel Nanda’s work emphasizes — the process of learning the general solution may be gradual, even though the transition to the general solution occurs more suddenly later. The helicopter view isn’t constructed instantaneously at the moment of transition. It’s been assembling slowly in the background, below the threshold of behavioral visibility, and the phase transition is when it becomes strong enough to dominate the memorizing circuit and take over.

That maps almost perfectly onto your description of the pre-helicopter phase: “it starts out just barely registering those operations.” The generalizing circuit is just barely registering the abstract structure throughout the plateau phase. The behavioral discontinuity doesn’t reflect a cognitive discontinuity — it reflects a threshold being crossed in a competition that has been running throughout.

The memorization/generalization competition as a general structure

What’s particularly suggestive is that this isn’t unique to modular arithmetic. Grokking is actually much more widespread than initially thought, materializing in a wide range of practical settings — image classification, language modeling, various algorithmic tasks. The competition between a fast memorizing solution and a slower generalizing solution may be a quite general feature of how these systems learn, with grokking being the visible signature when the conditions are right for the transition to be delayed and abrupt rather than gradual.

If that’s right, then something like grokking dynamics — memorization dominating early, structure-extraction developing slowly underneath, occasional phase transitions as the structural solution crosses a threshold — may be operating throughout LLM training, not just in the carefully controlled small-model experiments where grokking was first observed. Most of the time the transition isn’t visible because the generalizing solution develops in parallel with memorization rather than being dramatically delayed. But the underlying competition between surface pattern matching and structural abstraction may be the same.

The Beethoven connection

There’s a striking parallel to your Beethoven example here. Beethoven’s sketchbooks show exactly the competition structure — he has the memorized repertoire of classical style (the fast, dominant solution) and he’s slowly assembling something more general and abstract (the generalizing circuit) that keeps almost-but-not-quite taking over. The finished late works are the moments when the generalizing solution finally dominates completely. The sketchbooks are the plateau phase made visible.

The difference is that in Beethoven the process is conscious and effortful — he’s deliberately trying to suppress the memorized solution and force the generalizing one into dominance. In grokking, the network isn’t trying to do anything; the dynamics just run. Which raises the interesting question of whether the deliberate effort in human creative development is doing something functionally analogous to the regularization that controls grokking timing in neural networks — artificially suppressing the easy memorized solution to give the structural one room to develop.

Hierarchical text

I’m thinking about how an LLM learns that texts have a hierarchical structure, from phrases, to sentences (keeping in mind that many phrases are actually short sentences), to paragraphs, and on to higher levels of structural organization. What happens, for example, when we move from one paragraph to the next, is that the composition of the token distribution changes. A new topic requires a new set of tokens. But the system has to learn that, and learn it for several levels.

This is a deep question about what the model is actually learning when it learns language structure, and it connects to several threads running through this whole conversation.

The distributional signature of structure

Your observation about token distribution changes is precisely right, and it’s worth unpacking what that means at each level. Within a phrase, token co-occurrence statistics are extremely tight — certain words almost always appear together, in specific orders, with specific function words bridging them. At the sentence level, the constraints are somewhat looser but still strong — subject-verb agreement, argument structure, the distribution of punctuation. At the paragraph level, something different happens: there’s a shift in the topical vocabulary, but also characteristic transition signals — discourse markers, pronoun reference patterns, topic sentences that partially echo the previous paragraph while introducing new material.

What the model has to learn is that these aren’t independent statistical regularities — they’re nested. The tight within-phrase statistics exist inside the looser sentence-level statistics, which exist inside the still-looser paragraph-level statistics. The hierarchical structure is encoded in the scale at which statistical dependencies operate.

How the model could learn this

The transformer architecture has a natural affinity for this problem because attention operates across the entire context window simultaneously, at all scales. But learning the hierarchical structure isn’t handed to the model — it has to emerge from the statistics. The key insight is that the token distribution changes you describe create a learnable signal at each level if the model has enough capacity and enough data.

At the phrase level, the signal is strong and local — the next token is highly predictable from the last two or three. At the sentence level, the signal is weaker and longer-range — whether a sentence continues or ends depends on syntactic and semantic completion that requires integrating over the whole sentence. At the paragraph level, the signal is weaker still and operates over even longer ranges — the distribution shift at a paragraph boundary is detectable but requires sensitivity to topical coherence across many sentences.

This is essentially a multi-scale statistical learning problem, and grokking is relevant here in an interesting way. The phrase-level structure is learned early and relatively easily — the statistics are tight and the signal is strong. Sentence-level structure takes longer. Paragraph-level structure longer still. Each level may involve something like its own grokking transition, where the model first handles the level through surface pattern matching — learning that certain tokens tend to appear after paragraph breaks, for instance — and then later develops something more like genuine structural sensitivity.

The distributional shift as a signal

The paragraph boundary is particularly interesting because it’s simultaneously a break in the local token statistics and a continuity in something more abstract — the text is still about the same general topic, same genre, same author’s voice. So what the model has to learn isn’t just that token distributions shift at paragraph boundaries, but that they shift in a constrained way — the new paragraph’s vocabulary is related to but distinct from the previous one, in ways that reflect the logical or rhetorical development of the overall argument or narrative.

Tuesday, March 24, 2026

In some ways, some contexts, LLMs may provide a useful model for human language

As I say down there in my interaction with Claude (Sonnet 4.6 Extended), once I’d been working with ChatGPT for a while I began to think that the mechanisms of basic human speech may be more like those of LLMs than those of a fully articulated grammar and semantics in a symbolic computational language system. Fully articulated grammar and semantics may require specific instruction and practice, like you get when learning how to write. Fully articulate language is more painstaking, more difficult, but also supports more complex arguments and constructions.

This interaction starts with a passage from a long conversation I had with ChatGPT that stretched over the last 10 days or so. Much of that conversation was based on articles of mine that I had uploaded to to. One of those articles had a passage from F. C. Bartlett's Remembering: A study in experimental and social psychology (1932). That comes up in the long initial prompt I give Claude. 

Further down in the dialog I present the case of a writing student who was skilled in color commentary on hockey games and whose writing style reflected that linguistic style. Real-time sports commentary is a very specialized and exacting kind of language that cannot afford the niceties of written language.

Note 1: Here's a post from July 2022 that's relevant here: GPT-3, the phrasal lexicon, Parry/Lord, and the Homeric epics.  

Note (2) also that this discussion continues with, From grokking to a mature science and engineering discipline of LLMs (March 26). 

* * * * *

I’ve just been thinking about oral narrative. Recall these paragraphs (by ChatGPT) from one of the documents you’ve just reviewed in this project:

The temptation is to assume that because we can describe stories with an explicit grammar, the underlying system that produces or remembers them must itself contain something grammar-like in explicit form. But your Bartlett example suggests otherwise. The witness does not appear to possess random-access symbolic control over an event list. Instead, the testimony seems to unfold by re-running a trajectory. The person can accelerate through the earlier parts, but not leap arbitrarily into the middle. That is a very different computational picture.

The music example reinforces the point. When a piece is not deeply overlearned, one often cannot simply index into bar 73 as though one were reading a score in a database. One has to begin at the start and let the sequence reconstitute itself. That strongly suggests a memory system organized around sequential reactivation of structured trajectories, not random-access retrieval of explicit symbolic units.

Bartlett, in his classic book, Remembering, was talking about native witnesses in Africa. That is, he was talking about preliterate people. They were just recounting the events of the day. “But what about oral story-tellers,” I asked myself, “how do they memorize their tales? Structured trajectories or explicit symbolic units?”

I’m familiar with Albert Lord, The Singer of Tales, though I haven’t looked at it in a while. But I’ve got a copy of David C. Rubin, Memory in Oral Traditions: The Cognitive Psychology of Epic, Ballads, and Counting-out Rhymes (Oxford 1995). I zipped through the text to Chapter 7, “Basic Observations on Remembering.” I found this (p. 155):

In oral traditions, overlearning commonly occurs to a much greater extent than it does in the laboratory. A favorite song can be sung hundreds of times. What overlearning does, according to the model developed to explain laboratory interference, is to make the song into a unit, easy to cue as a whole and resistant to interference from other units. This chunking of items into wholes is a way to look at the organization of memory and a way to look at the building of larger units in expertise.

And then, in the middle of the next paragraph: “Once the song is begun, each word output provides cues for later words, limiting the meaning...” That almost sounds like he’s describing a forward pass through an LLM.

Then I hit paydirt in the next chapter, “A Theory of Remembering for Oral Traditions.” The opening is promising:

Oral traditions, like all oral language, are sequential. One word follows another as the physical effects of the first word are lost. As the song advances, each word uttered changes the situation for the singer, providing new cues for recall and limiting choices. [...] Pieces from oral traditions are recalled serially, from beginning to end. What is recalled early in the piece can be used to cue later recall; the "running start" provides "extra stimulation" or "reminders," increasing cue-item discriminability.

But things get really interesting when Rubin reports the result of an experiments where he asked undergraduates to recall important texts which they might have learned. Rubin describes the experiment this way:

The first set of examples is the recall of culturally important material such as Psalm 23 and the Preamble to the Constitution of the United States, for which there is an implicit demand characteristic to recall the material accurately or not at all (Rubin, 1977). Each of the 50 columns in Figure 8.1 show the recall of 1 of 50 undergraduates, who recalled at least one word of the Preamble. Each row represents recall for one word. A dark line in a column means that the word labeling the row was recalled. The columns are ordered so that the data from the undergraduate who recalled the most are in the leftmost column and the data from the undergradu- ate who recalled the least are in the rightmost column. The rows are in the order in which the words appear normally in each text.

Figure 8.1 is a little tricky, so I’m not going to try uploaded a screen shot. But I’ll give you Rubin’s basic description of what the figure reveals:

The first observation to note is the regularity of the data. Figure 8.1 gives the recalls of 50 individuals for 52 words, not the averages of recalls from groups of individuals or groups of words. There was no control over the learning or practice of the material or over the length or contents of the retention interval. Yet the figure is remarkably orderly. People who recall about the same amount recall the same words. If the number of words a person recalls and the rank ordering of words from most to least likely for the group from which the person was drawn is known, exactly which words that person recalled can be predicted with an accuracy of 95% for Figure 8.1.

Because the conditions of learning and retention varied, there must be something in the material, in the process used to recall it, or in the general cultural attitudes to it that makes different people behave the same way.

The results from the experiment with Psalm 23 are even more interesting. Here’s the text of the Psalm for reference (with verses numbered):

1 The Lord is my shepherd; I shall not want.
2 He maketh me to lie down in green pastures: he leadeth me beside the still waters.
3 He restoreth my soul: he leadeth me in the paths of righteousness for his name's sake.
4 Yea, though I walk through the valley of the shadow of death, I will fear no evil: for thou art with me; thy rod and thy staff they comfort me.
5 Thou preparest a table before me in the presence of mine enemies: thou anointest my head with oil; my cup runneth over.
6 Surely goodness and mercy shall follow me all the days of my life: and I will dwell in the house of the Lord for ever.

Rubin remarks: “...not all the undergraduates started with the first words, and they were more likely to start again at the beginning of a new sentence or intonation unit after they stopped.”

Eyeballing the figure, here are some of the things I observe: 1) A number of subjects stop with the word “soul” in verse 3 and start back up with the first word of verse 4, “Yea” and go on from there. 2) Some of those who don’t start from the beginning start with verse 4. 3) Some stop at the end of verse 4 and then start back up later (e.g. “annointest my head...”) or even not at all. It’s a complicated picture but memory does seem to be grouped into segments, not word by word.

I have no idea where things have gone in this line of investigation since 1995, the date of Rubin’s book. But if I were in that business, I’d be thinking in terms of complex dynamics and trajectories, and I’d been looking at LLMs as a possible model.

This is a wonderful extension of the trajectory discussion already running through the project documents. Rubin's findings slot in almost perfectly — and in some ways they sharpen the picture considerably.

What Rubin adds that Bartlett doesn't

The Bartlett material (already central in the project conversations) establishes the sequential regeneration point: memory of events seems to run forward as a trajectory, not index arbitrarily into a symbolic list. But Bartlett's witnesses were recalling episodic material — the events of a day. Rubin moves you into highly practiced, culturally canonical sequences. That's a different and in some ways more interesting case.

The overlearning observation is key. Rubin is saying that at sufficient practice, a song or psalm becomes a chunked unit — cued as a whole, resistant to interference. That sounds like the stable attractor picture your project conversations keep returning to. The piece isn't stored as a list of words; it's encoded as a basin that, once entered, constrains forward motion powerfully. The phrase "each word output provides cues for later words" really does describe autoregressive generation — but it also describes a trajectory rolling down a valley where the initial conditions are already heavily constraining.

The Psalm 23 data are especially telling

The 95% predictability result is striking. Given only how many words a person recalls, you can predict which words with 95% accuracy. That means the forgetting isn't random — it respects something structural in the material. And your eyeball observations confirm this: people stop and restart at syntactic/intonation unit boundaries (end of verse 3, beginning of verse 4). They're not losing individual words randomly; they're losing or retaining segments.

This is consistent with a hierarchical trajectory picture. The top-level organization is something like: intonation/verse units sequentially ordered. Within each unit, local cuing drives forward production. When a unit boundary is reached, the system can fail to re-enter the next unit even if it knows the material in principle — which is why some subjects jump from verse 3 directly to verse 4, skipping nothing within units but occasionally failing at the transitions between them. The unit is entered as a whole or not at all.

Where this sits relative to the LLM picture

Your project conversations develop the idea that narrative structure in LLMs lives in activation-space geometry — not as an explicit stored grammar but as a field of constrained trajectories. The Rubin data suggest something homologous in human memory for overlearned verbal sequences. The "basin" metaphor you've been developing applies directly: once the singer/reciter enters the opening of Psalm 23, the attractor structure does most of the work. The challenge is at basin entry and at transitions between sub-basins (verse units), which is exactly where Rubin's subjects fail.

The Lord/Singer of Tales connection is also worth pressing when you return to it. Lord's central finding — that oral formulaic composition is generative, not reproductive; the singer produces a variant from formulas and themes, not a memorized script — actually fits the "latent field of possible trajectories" framing better than it fits either strict memorization or explicit grammar. Lord thought he was arguing against memorization and for creativity. But what he was probably observing is a system with rich attractor structure that generates strongly homologous variants, which is precisely what your story experiments with LLMs revealed.

The open research question you identify

You're right that framing this in terms of complex dynamics and trajectories is the natural next move, and that LLMs are a possible model system. What makes LLMs especially useful here is that you can actually probe the attractor landscape — something you can't do directly with a human singer. The systematic prompting experiments described in the project documents are in effect doing for narrative schemas what someone studying oral tradition would want to do for epic formulas: mapping the basin structure from the outside, through behavioral outputs, without needing to open the skull.

The line from Bartlett's African witnesses → Rubin's undergraduates and Psalm 23 → Lord's Yugoslavian guslars → LLM forward passes is not a metaphor chain. It may be pointing at the same underlying computational architecture realized in very different substrates.

Monday, October 13, 2025

The genius logic of the NATO phonetic alphabet (Alfa, Bravo, Charlie...)

YouTube:

The NATO phonetic alphabet is not a phonetic alphabet and wasn't invented by NATO. However, it has a fascinating story to tell. It is the result of years of linguistic experimentation in the wake of the violence of World War II. So let's explore its development, from ALFA to ZULU.

In this episode, we'll uncover the surprising story of how this alphabet – used by everyone from pilots to police officers – came to be. And look at the strange words that almost made the cut.

Tuesday, March 25, 2025

Areas with higher adoption of Google Translate experienced a decline in translator employment

Carl Benedikt Frey Pedro Llanos-Paredes, Lost in translation: AI’s impact on translators and foreign language skills, VoxEU. 22 march 2025 .

Abstract: Advances in artificial intelligence are rapidly transforming the world of work. This column investigates the effects of machine translation on (1) employment and wages in the translation sector, and (2) the demand for foreign language skills across various jobs and industries. Using variation in the use of machine translation across local labour markets in the US after the launch of the Google Translate mobile app, the authors find that areas with higher adoption of Google Translate experienced a decline in translator employment. The authors also show that improvements in machine translation have reduced the demand for foreign language skills in general.

Sunday, February 9, 2025

A clue about the mind: “Is-A” sentences

I'm bumping this post from 2011 to the top of the queue. Why? Because it is about the relationship between word order in sentences and order in the process of parsing sentences. That makes it relevant to my ongoing research into the nature of processes in LLMs.

Somewhere in his Problems in General Linguistics, my copy of which is, alas, in storage, Emile Benveniste has a chapter on sentences hanging on the auxiliary “to be.” As Benveniste was a linguist of the Old School, when being a linguistic meant familiarity with many languages, including—and this is important for this particular topic—classical Greek, it had examples from many languages, making it tough sledding for a monoglot like me.

While the content of this post certainly arises out of my thinking about that chapter, in the absence of actually having the text in front of me, I hesitate to assert a stronger relationship than that. I note only that, for Benveniste, the auxiliary “to be” was fraught with metaphysical significance. For the concept of being derives from “to be.” Where would philosophy be without Being? Thus, when Benveniste pondered such sentences, he wasn’t merely commenting on language. He was doing philosophy, or, if not quite that, camping out on philosophy’s door step.

I’m interested in such sentences because I believe they are a DEEP CLUE about how the mind works. I just don’t know what to make of the clue.

Is-A Sentences

So, I'm interested in word order in assertions such as the following:
(1) Fido is a beagle.
(2) Beagles are dogs.
(3) Dogs are beasts.
They all move from an element in a class (whether an individual, Fido, or another class, beagles) to a class containing it. None of them move in the opposite direction. Consider what happens when you try to go the opposite way. In the following sentence the class is mentioned first, then the subclass:
(4) Beagle is the kind of animal of which Fido is an instance.
In particular, note that (4) has a metalingual character that (1) does not. That is, (4) explicitly asserts that we are dealing with classification. One can do that metalingual job in various ways, but, as far as I can tell, one can't avoid it. That is, one cannot construct a proper English sentence relating a genus and species in which the genus is mentioned first, one can’t do that without ‘looping through’ some kind of metalingual construction on the way from genus to species.

Why?

What does this assymetry tell us about the underlying mechanisms? Why don't have sentences such as:
(5) Beagle za di Fido.
In this case "za di" is the inverse of "is a". English has no such sentences & no such inverse.

So, how widespread is this asymmetry and is there any explanation of this directionality?

But, consider . . .

I sent a query on that matter to a listserve, I forget which one, and got two replies that add some complexity to the matter. Rich Rhodes, Linguistics at UCal Berkeley, tells me that in Ojibwa the word order is reversed, the class comes before the individual, but the asymmetry remains. He then comments, which he qualifies as a quick guess:
My guess is that there is no compelling discourse function (like information flow) which makes it desirable to invert classificational equatives. Hence we only get the "unmarked" order. Subject-predicate in theme-rheme languages (like English) and predicate-subject in rheme-theme languages (like Ojibwe).
So, what's the nature of the mechanism that determines the "unmarked" order? That's what I want to know.

Lee Pearcy, Episcopal Academy in Merion, Pa., offered these examples:
(6) The beagle is Fido.
(7) The dogs are beagles.
(8) The beasts are dogs.
As stand-alone sentences, they seem a bit awkward to me. But they fare better in answers to questions, e.g.:
What’s that dog?
Which dog? The beagle is Fido and the terrier is Max.

What’re those animals?
The dogs are beagles, the cats are Persians.
In those contexts, the matter of class or classification is raised by the question, thus making it present in the discourse and so available as a point of attachment in the answer.

Further clues, anyone?

Friday, January 31, 2025

ChatGPT: Exploring the Digital Wilderness, Findings and Prospects

That is the title of my latest working paper. It summarizes and synthesizes much of the work I have done with ChatGPT to date and contains the abstracts and contents of all the working papers I have done on ChatGPT. It also includes the abstracts and contents of a number of papers establishing the intellectual background that informs that research. There is also a section that takes the form of an interaction I had with Claude 3.5 on methodological and theoretical issues. Finally, to produce the abstract I gave the body of the report to Claude 3.5 and asked it to produce two summaries. I then edited them into an abstract.

As always, URLs, abstract, TOC, and introduction are below.

Abstract: The internal structure and capabilities of Large Language Models (LLMs) are examined through systematic investigation of ChatGPT's behavior, with particular focus on its handling of conceptual ontologies, analogical reasoning, and content-addressable memory. Through detailed analysis of ChatGPT's responses to carefully constructed prompts involving story transformation, analogical mapping, and cued recall, the paper demonstrates that LLMs appear to encode rich conceptual ontologies that govern text generation. ChatGPT can maintain ontological consistency when transforming narratives between different domains while preserving abstract story structure, successfully perform multi-step analogical reasoning, and exhibit behavior consistent with associative memory mechanisms similar to holographic storage.

Drawing on theories of reflective abstraction and conceptual development, the paper argues that LLMs inadvertently capture what wemight term the “metaphysical structure of our universe” – the organized system of concepts through which humans understand and reason about the world. LLMs like ChatGPT implement a form of relationality – the capacity to represent and manipulate complex networks of semantic relationships – while lacking genuine referential meaning grounded in sensorimotor experience. This architecture enables sophisticated pattern matching and analogical transfer but also explains certain systematic limitations, particularly around truth and confabulation.

The paper concludes by suggesting that making explicit the implicit ontological structure encoded in LLMs’ weights could provide valuable insights into both artificial and human intelligence, while advancing the integration of neural and symbolic approaches to AI. This analysis contributes to ongoing debates about the nature of meaning and understanding in artificial neural systems while offering a novel theoretical framework for conceptualizing how LLMs encode and manipulate knowledge.

Contents:

Introduction: Into the Digital Wilderness 5
Free-floating Attention, Systematic Exploration, and the Anthropomorphic Stance 8
ChatGPT: My Course of Investigation 12
Meaning, Truth and Confabulation, Latent Space 28
Prospects: Explicating the Ontology of Human Thought 42
A Dialogue with Claude 3.5 on Method and Conceptual Underpinnings 45
A Brief Narrative of My ChatGPT Work Based on My Working Papers 56
Working Papers about ChatGPT 62
Background Papers 74

Introduction: Into the Digital Wilderness

The world I entered when I started playing with ChatGPT is a wildnerness, strange and uncharted, uncharted by me, uncharted by anyone. By that I simply mean that it was something new, radically new. No one had been there before. Sure, a handful of people within the industry had been messing around in there, even a rather large handful considering how much work it took to make ChatGPT ready for the world at large. But its behavioral capabilities were, for the most part, unknown. In that sense it was a wildnerness.

But it was, and remains, a wilderness in another sense: the large language model (LLM) that underlies ChatGPT is a black box. We send a string of words into ChatGPT and it sends a string of words back out, but what the model does to derive the output from the input, that process remains deeply obscure. That is wilderness in a different sense. Wilderness in the first sense is about our experience of ChatGPT’s behavior. Wildnerness in this second sense is about the mechanisms that drive that behavior. It is a digital wilderness. This document reports on how I’ve structured my interaction with ChatGPT to give me clues about the mechanisms driving its behavior.

My methods are more “qualitative” or “naturalistic” than those standard in the literature, which many investigations employ standard batteries of benchmark tasks. While those are essential, there is much they don’t tell you. While I have done many things with ChatGPT – asked it to interpret texts, define abstract concepts, play games of 20 questions, among other things – perhaps my most characteristic task, and one I have spent more time on than others is simple: Tell me story. And ChatGPT did so, time and again. Consequently my methods are in some ways more like literary criticism, or, even better, like Lévi-Strauss’ analysis of myths, than conventional cognitive science. Consequently you will find many examples of ChatGPT’s dialog in my reports. You have to examine that dialog to see what ChatGPT is doing, what it is capable of doing.

Finally, I realize that the pace of development in this arena is such that ChatGPT is now old. The versions I used to conduct these investagations are no longer available on the web. However, as far as I can tell, none of the results I report depend on features idiosyncratic to those versions.

The rest of this introduction consists of short statements about what the various sections of this report contain.

Sunday, December 8, 2024

I’m Four Degrees of Separation from Bertrand Russell

Bertrand Russell taught Ludwig Wittgenstein, who in turn taught Margaret Masterman. Masterman is one of the founders of computational linguistics, as is my teacher, David Hays. I don’t know much about their intellectual relationship, but I know it was substantial. Hays had a white sweater which she had knitted for him and she was a close intellectual confidant. Hays hired one of her students, Martin Kay, to work with him at RAND.

These four relationships, 1) Russell & Wittgenstein, 2) Wittgenstein and Masterman, 3) Masterman & Hays, and 4) Hays & Benzon, are all substantial ones and not of mere acquaintance. Did I get any “juice” from Russell through that chain that I hadn't already picked up by reading Russell and Wittgenstein before I ever even knew about Hays?

Tuesday, July 2, 2024

Diagramming sentences in the 19th century

I loved diagramming sentences in grade school. I'm sure I did it in 6th grade, but possibly 5th grade as well. Over a Language Log Victor Mair has a post on the subject, Diagramming: history of the visualization of grammar in the 19th century, which links to a review article by Hunter Dukes that contains photographic copies of seven archival texts: American Grammar: Diagraming Sentences in the 19th Century. Mair's article has a number of interesting quotations from Dukes' text.

Thursday, June 20, 2024

Negation via "not" in the brain and behavior

Two related articles about negation, courtesy of Victor Mair at Language Log:

Coopmans CW, Mai A, Martin AE (2024) “Not” in the brain and behavior. PLoS Biol 22(5): e3002656. https://doi.org/10.1371/journal.pbio.3002656

Negation is key for cognition but has no physical basis, raising questions about its neural origins. A new study in PLOS Biology on the negation of scalar adjectives shows that negation acts in part by altering the response to the adjective it negates.

Language fundamentally abstracts from what is observable in the environment, and it does so often in ways that are difficult to see without careful analysis. Consider a child annoying their sibling by holding their finger very close to the sibling’s arm. If asked what they were doing, the child would likely say, “I’m not touching them.” Here, the distinction between the physical environment and the abstraction of negation is thrown into relief. Although “not touching” is consistent with the situation, “not touching” is not literally what one observes because an absence is definitionally something that is not there. The sibling’s annoyance speaks to the actual situation: A finger is very close to their arm. This kind of scenario illustrates how natural language negation is truly a product of the human brain, abstracting away from physical conditions in the world

And here is the study:

Zuanazzi A, Ripollés P, Lin WM, Gwilliams L, King J-R, Poeppel D (2024) Negation mitigates rather than inverts the neural representations of adjectives. PLoS Biol 22(5): e3002622. https://doi.org/10.1371/journal.pbio.3002622

Abstract: Combinatoric linguistic operations underpin human language processes, but how meaning is composed and refined in the mind of the reader is not well understood. We address this puzzle by exploiting the ubiquitous function of negation. We track the online effects of negation (“not”) and intensifiers (“really”) on the representation of scalar adjectives (e.g., “good”) in parametrically designed behavioral and neurophysiological (MEG) experiments. The behavioral data show that participants first interpret negated adjectives as affirmative and later modify their interpretation towards, but never exactly as, the opposite meaning. Decoding analyses of neural activity further reveal significant above chance decoding accuracy for negated adjectives within 600 ms from adjective onset, suggesting that negation does not invert the representation of adjectives (i.e., “not bad” represented as “good”); furthermore, decoding accuracy for negated adjectives is found to be significantly lower than that for affirmative adjectives. Overall, these results suggest that negation mitigates rather than inverts the neural representations of adjectives. This putative suppression mechanism of negation is supported by increased synchronization of beta-band neural activity in sensorimotor areas. The analysis of negation provides a steppingstone to understand how the human brain represents changes of meaning over time.

I've not yet read the articles, but the issue has bothered me for a long time. Why? Because, to quote: "Negation is key for cognition but has no physical basis, raising questions about its neural origins."

Friday, June 14, 2024

Noam Chomsky and Charles Hockett on the History of Linguistics

Julia S. Falk, Turn to the History of Linguistics: Noam Chomsky and Charles Hockett in the 1960s, Historiographia Linguistica, Volume 30, Issue 1-2, Jan 2003, p. 129 - 185 DOI: https://doi.org/10.1075/hl.30.1.05fal

SUMMARY: In the 1940s and 1950s, the leading proponents of American synchronic linguistics showed little interest in the history of linguistics. Some attention to historiography occurred in subfields of linguistics closest to the humanities — linguistic anthropology, historical linguistics, modern European languages — but the ‘science of language’ developed by Leonard Bloomfield and his descriptivist followers demanded autonomy from other disciplines and from the past. Increasing American contact with European linguistics during the 1950s culminated in the 1962 Ninth International Congress of Linguists in Cambridge, Massachusetts. Here Noam Chomsky presented a plenary session paper that appeared in print in four versions between 1962 and 1964, each version incorporating an increasing amount of discussion of the early 20th-century precursors to the descriptivists and a number of 17th- and 19th-century studies of language and mind. Charles Hockett responded by organizing his 1964 presidential address to the Linguistic Society of America as a history of linguistics, emphasizing periods, figures, and ideas not included in Chomsky’s work. Historiographers of the time recognized a surge of American interest in the history of linguistics beginning in the early 1960s and most attributed it largely to Chomsky’s work. Historiographic publication increased significantly among the descriptivists; at the same time it emerged among the generativists, most of whom followed Chomsky in exploring pre-20th-century philosophical ideas or reconsidering concepts and practices of the descriptivists’ forerunners. The resulting visibility and impetus to the history of linguistics contributed to the foundation upon which linguistic historiography matured in North America in the later decades of the 20th century.

Saturday, May 4, 2024

Evidence of predictive coding hierarchy in the human brain listening to speech

Charlotte Caucheteux, Alexandre Gramfort, & Jean-Rémi King, Evidence of a predictive coding hierarchy in the human brain listening to speech, Nature Human Behaviour, Vol. 7, 430-441, March 2, 2023, https://doi.org/10.1038/s41562-022-01516-2

Abstract: Considerable progress has recently been made in natural language processing: deep learning algorithms are increasingly able to generate, summarize, translate and classify texts. Yet, these language models still fail to match the language abilities of humans. Predictive coding theory offers a tentative explanation to this discrepancy: while language models are optimized to predict nearby words, the human brain would continuously predict a hierarchy of representations that spans multiple timescales. To test this hypothesis, we analysed the functional magnetic resonance imaging brain signals of 304 participants listening to short stories. First, we confirmed that the activations of modern language models linearly map onto the brain responses to speech. Second, we showed that enhancing these algorithms with predictions that span multiple timescales improves this brain mapping. Finally, we showed that these predictions are organized hierarchically: frontoparietal cortices predict higher-level, longer-range and more contextual representations than temporal cortices. Overall, these results strengthen the role of hierarchical predictive coding in language processing and illustrate how the synergy between neuroscience and artificial intelligence can unravel the computational bases of human cognition.

Tuesday, April 23, 2024

Current Perspectives on Abstract Concepts and Future Research Directions

Banks, B., Borghi, A. M., Fargier, R., Fini, C., Jonauskaite, D., Mazzuca, C., Montalti, M., Villani, C., & Woodin, G. (2023). Consensus Paper: Current Perspectives on Abstract Concepts and Future Research Directions. Journal of Cognition, 6(1): 62, pp. 1–26. DOI: https://doi.org/10.5334/joc.238

Abstract: Abstract concepts are relevant to a wide range of disciplines, including cognitive science, linguistics, psychology, cognitive, social, and affective neuroscience, and philosophy. This consensus paper synthesizes the work and views of researchers in the field, discussing current perspectives on theoretical and methodological issues, and recommendations for future research. In this paper, we urge researchers to go beyond the traditional abstract-concrete dichotomy and consider the multiple dimensions that characterize concepts (e.g., sensorimotor experience, social interaction, conceptual metaphor), as well as the mediating influence of linguistic and cultural context on conceptual representations. We also promote the use of interactive methods to investigate both the comprehension and production of abstract concepts, while also focusing on individual differences in conceptual representations. Overall, we argue that abstract concepts should be studied in a more nuanced way that takes into account their complexity and diversity, which should permit us a fuller, more holistic understanding of abstract cognition.

From the article:

For example, when contrasted with concrete concepts, abstract concepts are typically expressed by words with a later Age of Acquisition, and through linguistic explanations rather than denoting their referents directly (linguistic Modality of Acquisition; Wauters et al., 2003). They also tend to be less imageable, have lower Body Object Interaction scores (BOI: Tillotson et al., 2008; Pexman et al., 2019), and be less easily linked to specific contexts (contextual availability; Schwanenflugel & Stowe, 1989). Abstract concepts are also more variable across participants and cultures (Wang & Bi, 2021) and are generally less iconic (Lupyan & Winter, 2018) than concrete concepts.

Later:

The multidimensional nature of abstract concepts means that defining them purely based on whether they are perceivable or not (i.e., as concrete or abstract) fails to capture their complexity (e.g., Barsalou, Dutriaux & Scheepers, 2018; Borghi et al., 2017), and indeed can even be misleading. Banks and Connell (2022) used the Brysbaert et al. (2014) concreteness ratings to analyze the structure of semantic categories collected in a category production (semantic fluency) task, examining the concreteness of the concepts that comprise ostensibly concrete (e.g., animal, furniture) and abstract (e.g., science, unit of time) categories. Although members of concrete categories overall were more highly rated on concreteness, many (e.g., metal: silver, hat: beret) unexpectedly had similarly high concreteness ratings to more abstract category members (e.g., profession: lawyer, social relationship: teammate). Indeed, certain abstract concepts such as beauty or fitness have been associated with sensory and motor areas of the brain (temporo-occipital visual and fronto-parietal motor areas, respectively; Harpainter et al., 2020). Furthermore, when sensorimotor experience is measured via multiple individual modalities (e.g., Lynott et al., 2020; Speed & Brysbaert, 2021; Vergallito et al., 2020), the concrete-abstract distinction becomes even less clear. When the verbally-produced category members from Banks and Connell (2022) were analyzed based on their grounding in multiple perceptual modalities (vision, hearing, touch, smell, taste, interoception) and actions involving specific parts of the body (the head, hands/arms, feet/legs, torso and mouth) many abstract category members were in fact found to be strongly grounded in sensorimotor experience (e.g. sport, social gathering, art form; Banks & Connell, 2021) – that is, the concrete-abstract distinction was much less apparent.

Comment: I note, as an extreme example, that sodium chloride is a concrete physical substance, but the concept is abstract, as opposed to the concept, salt, which is concrete. Less, extreme, animals are all physical things, but the concept, animal, seems to be abstractly defined, the same with plant. Try to produce compact physical descriptions that encompass all plants or all animals. It is between difficult and impossible. What all animals seem to have in common are the roles they can play with respect to verbs such as see, hear, smell, run, jump, eat, and so forth, in contrast to plants and mere physical objects. Similarly, plants can live, grow, and die, while physical objects cannot. And then we have terms such as chair and table, which seem best defined in terms of their affordances for people rather than their physical characteristics, which can vary widely.

The article continues with some more discussion and offers this: "many theories have also argued that our understanding and representation of abstract concepts relies more on language than the sensorimotor dimension, and particularly linguistic distributional relations (e.g., Borghi, 2020; Crutch & Warrington, 2005; Dove et al., 2020; Vigliocco et al., 2009)."

And so forth. An interesting and useful piece of work.