Showing posts with label emergence. Show all posts
Showing posts with label emergence. Show all posts

Friday, August 29, 2025

Talking Chimps and UFOs: A thought experiment

I'm bumping this to the top of the queue because I intend to put some version of it in my book in progress: PLAY: How to Stay Human in the AI Revolution.  
* * * * * 
 
This is an out-take from Beethoven’s Anvilmy book on music. It’s about a thought experiment that first occurred to me while in graduate school in the mid-1970s. Consider the often astounding and sometimes absurd things that trainers can get animals to do, things the don’t do naturally. Those acts are, in some sense, inherent in their neuro-muscular endowment, but not evoked by their natural habitat. But place them in an environment ruled by humans who take pleasure in watching dancing horses, and . . . Except that I’m not talking about horses.


It seems to me that what is so very remarkable about the evolution of our own species is that the behavioral differences between us and our nearest biological relatives are disproportionate to the physical and physiological differences. The physical and physiological differences are relatively small, but the behavioral differences are large.

In thinking about this problem I have found it useful to think about how at least some chimpanzees came to acquire a modicum of language. All of them ended in failure. In the most intense of these efforts, Keith and Cathy Hayes raised a baby chimp in their household from 1947 to 1954. But that close and sustained interaction with Vicki, the young chimp in question, was not sufficient. Then in the late 1960s Allen and Beatrice Gardner began training a chimp, Washoe, in Ameslan, a sign language used among the deaf. This effort was far more successful. Within three years Washoe had a vocabulary of Ameslan 85 signs and she sometimes created signs of her own.

The results startled the scientific community and precipitated both more research along similar lines—as well as work where chimps communicated by pressing ironically identified buttons on a computerized panel—and considerable controversy over whether or not ape language was REAL language. That controversy is of little direct interest to me, though I certainly favor the view that this interesting behavior is not really language. What is interesting is the fact that these various chimps managed even the modest language that they did.


The string of earlier failures had led to a cessation of attempts. It seemed impossible to teach language to apes. It would seem that they just didn’t have the capacity. Upon reflection, however, the research community came to suspect that the problem might have more to do with vocal control than with central cognitive capacity. And so the Gardners acted on that supposition and succeeded where others had failed. It turns out that whatever chimpanzee cognitive capacity was, it was capable of surprising things.

Note that nothing had changed about the chimpanzees. Those that learned some Ameslan signs, and those that learned to press buttons on a panel, were of the same species as those that had earlier failed to learn to speak. What had changed was the environment. The (researchers in the) environment no longer asked for vocalizations; the environment asked for gestures, or button presses. These the chimps could provide, thereby allowing them to communicate with the (researchers in the) environment in a new way.

It seemed to me that this provided a way to attack the problem of language origins from a slightly different angle. So I imagined that a long time ago groups of very clever apes – more so than any extant species – were living on the African savannas. One day some flying saucers appeared in the sky and landed. The extra-terrestrials who emerged were extraordinarily adept at interacting with those apes and were entirely benevolent in their actions. These creatures taught the apes how to sing and dance and talk and tell stories, and so forth. Then, after thirty years or so, the ETs left without a trace. The apes had absorbed the ETs’ lessons so well that they were able to pass them on to their progeny generation after generation. Thus human culture and history were born.


Now, unless you actually believe in UFOs, and in the benevolence of their crews, this little fantasy does not seem very promising, for it is a fantasy about things that certainly never happened. Further, even if this had happened, it does seem to remove the mystery from language’s origins. Instead of something from nothing we have language handed to us on a platter. We learned it from some other folks, perhaps they were little short fellows with green skin, or perhaps they were the modern style aliens with pale complexions, catlike pupils in almond eyes and elongated heads.

But, and here is where we get to the heart of the matter, what would have to have been true in order for this to have worked? Just as the chimps before Ameslan were genetically the same as those after, so the clever before alien-instruction were the same as the proto-humans after. The species has not changed, the genome is the same – at least for the initial generation. The capacity for language would have to have been inherent in the brains of those clever apes. All the aliens did was activate that capacity. Once that happened the newly emergent proto-humans were able to sustain and further develop language on their own. Thus the critical event is something the precipitates a reconfiguration of existing capabilities.

However, language origins is not our problem. We are searching for the origins of music. So, instead of alien instruction in Hebrew or Sanskrit we can imagine alien instruction in samba or polka. The basic configuration and dynamics of the story remains the same. However, to make it real we have to get rid of those aliens and their instruction. Instead of the aliens we have only our group of clever apes. They are going to have to instruct one another. What we are looking for is a a way to get a gestalt switch in group dynamics that supports new modes of neural dynamics in the brains of individuals who are interacting with one another in a group.


Let us call this the Gestalt Origins Hypothesis:
Gestalt Origins: The precursor to music arose when groups of hominids interacted in a way that triggered a new configuration of operation in their existing nervous system.
Notice that I talk of a precursor to music. I don’t think we can get from ape to music in a single bound. We need at least one precursor, something that is rhythmic, like music, but not yet fully formed. In order to get even that far, so my argument goes, our proto-humans need better control over their vocal cords than apes have, they need more rhythmic sophistication, and greater mimetic capacity.

Notice that this story says nothing about the adaptive value of music. I do intend to get around to that toward the end of the chapter, but that’s not my primary concern. My primary concern is getting our ancestors to the point where a gestalt switch can happen that will bring about a precursor to music, something we can call musicking. In order for that to happen we need to solve an adaptive problem or two. But those adaptations are not about music; they are about its precursors. Once music-making is going along smoothly we need another gestalt switch to differentiate it into language and music proper.

The photos show graffiti that’s at the western end of the Erie Cut (aka the Bergen Arches) in Jersey City. There are at least two layers of graffiti. The back layer contains what appear to be UFOs, though perhaps the two at the right are mushrooms. The top layer has a name, Hemlock, which is obscured by the grass to one degree or another.

Wednesday, June 18, 2025

Large Language Models and Emergence

David C. Krakauer, John W. Krakauer, and Melanie Mitchell, Large Language Models and Emergence: A Complex Systems Perspective, June 16, 2025, https://arxiv.org/pdf/2506.11135

Abstract: Emergence is a concept in complexity science that describes how many- body systems manifest novel higher-level properties, properties that can be described by replacing high-dimensional mechanisms with lower-dimensional effective variables and theories. This is captured by the idea “more is different”. Intelligence is a consummate emergent property manifesting increasingly efficient—cheaper and faster—uses of emergent capabilities to solve problems. This is captured by the idea “less is more”. In this paper, we first examine claims that Large Language Models exhibit emergent capabilities, reviewing several approaches to quantifying emergence, and secondly ask whether LLMs possess emergent intelligence.

From the conclusion:

We argued that in LLMs, the term emergence should be used not merely to signify surprising or unpredictable task performance, or abrupt changes in performance, but requires at minimum the identification of relevant coarse- grained variables that form effective mechanisms— reduced “internal degrees of freedom”—for this behavior, mechanisms that can explain or predict the be- havior of the system at this higher level, screening off details of lower level mechanisms such as weights and activations. More quantitative evidence for emergence includes the kinds of principles related to emergence in physical sys- tems, such as breaking of scaling through reorganization, evidence for the use of novel bases and manifolds formed through compression of regularities, and new forms of abstraction that lead to demonstrable efficiencies in prediction, prob- lem solving, generalization, and analogy-making. Identifying such principles would be an important step in understanding the seemingly novel capabilities that arise in LLMs.

Three types of emergence claims have been made for LLM capabilities: (1) sharp improvements in specific capabilities that occur as the system or training data is scaled; (2) capabilities are identified that the LLMs were not specifically trained for; and (3) internal “world models” emerging from autoregressive token prediction. Each of these cases, and particularly the last, present provocative evidence for emergence, but in all cases that evidence is incomplete. Cases (1) and (2) relies on several assumptions: that the capabilities tested are genuinely new, general, and don’t rely on memorized training data or other shortcuts; that these capabilities are not present in simpler models; and that the capabilities are unexpected or unpredictable given the training data and the models’ size. None of these assumptions has been conclusively verified. As for case (3) the complex- ity framework of [60] provides a principled approach to thinking about “world models” as these relate to discrete-time stochastic processes. To the extent that an LLM is effective at next-token prediction, and to the degree to which the model can be shown to exploit a minimum of information, they might be de- scribed as world models. However, the recent work by [61] demonstrates that recovering an accurate world model is very difficult, since next token prediction is a fragile metric.

That last is particularly important to me because it is a property of old-style semantic and cognitive networks. The network provides the world model (and should be linked to sensory and motor systems, as it was in the model David Hays developed in the mid-1970s) from which text can be generated through linguistic processes. LLMs conflate the two, text and cognition, into a single distributed representation.

Later:

There are three possible roles of language as it relates to training an LLM: (1) language itself provides a more or less complete and compressed representation of the world (including non-linguistic modalities); (2) spoken or written language mirrors an internal “language of thought”; and (3) language is a non-supervised “programming language”. If language does provide a complete representation of the world, then training on more language data would indeed enable an increasingly expansive and detailed representation of natural and cultural patterns and processes. If natural language is the language of thought (“mentalese”) then training on more language data would fill out the numerous ways that human- ity has historically reasoned about regularities in the world. And if language is a programming language, by combining detailed instruction tuning with next word prediction it can exploit principles of computational universality to imple- ment any computable function.

We do not have definitive evidence for any of these three claims, but they play a crucial role in any statement relating to how surprising the behavior of an LLM will be deemed.

The final paragraph:

Human intelligence is a low-bandwidth phenomenon, and is as much if not more about the scaling down of effort as the scaling up of capability [72]. As Einstein wrote, “The grand aim of all science is to cover the greatest number of empirical facts by logical deduction from the smallest number of hypotheses or axioms.” [73] We know that for any elegant algorithm there is an alternative brute force solution that does the job. It might even be the case that there are uncountable problems that require brute force and that this is a domain where LLMs and their cognitively alien relatives, including SAT solvers, will provide extraordinary utility [74]. What Donald Knuth said of programs might also be applied to intelligence: “Programs are meant to be read by humans and only incidentally for computers to execute.” [75]. Similarly, intelligence is a property of understanding and only incidentally a matter of capability.

I really like the first sentence of that last paragraph. It "resonates" with the definition of intelligence I gave in What Miriam Yevick Saw: The Nature of Intelligence and the Prospects for A.I.:

Intelligence is the capacity to assign computational capacity to propositional (symbolic) and/or holographic (neural) processes as the nature of the problem requires.

As Yevick herself observed:

If we consider that both of these modes of identification enter into our mental processes, we might speculate that there is a constant movement (a shifting across boundaries) from one mode to the other: the compacting into one unit of the description of a scene, event, and so forth that has become familiar to us, and the analysis of such into its parts by description. Mastery, skill and holistic grasp of some aspect of the world are attained when this object becomes identifiable as one whole complex unit; new rational knowledge is derived when the arbitrary complex object apprehended is analytically described.