Showing posts with label neural-holography. Show all posts
Showing posts with label neural-holography. Show all posts

Sunday, June 9, 2024

The Hologram Metaphor in Mind and Culture

I'm bumping this post from June of 2017 to the top of the queue as it is germane to my current thinking about large language models (aka LLMs) and associative memory. Here I extend the idea of neural holography to a whole population.

* * * * *
 
From some old notes:

Finally, I'd like to suggest cultures encode their master patterns like holograms encode images. If you rip a hologram of, for example, a coin, in half, you can still use it to view the entire coin. If you rip one of those halves in half, you can use either of the resulting quarters to view the entire coin. This is quite different from an ordinary photograph where, if you rip it in half, you get one half of the coin on one piece of photograph and the other half on the other piece. If you rip one of those pieces in half you will be down to photographic fragments showing a quarter of the coin. A piece of a photograph contains a piece of the image.

A piece of a hologram, however, still contains the entire image. The resolution of the image, that is, its sharpness, will be somewhat reduced—the smaller the piece, the lower the resolution—but the entire image is there. A hologram is thus a way of distributing the entire image throughout the representing medium (the piece of photographic film). Similarly, the pattern of a culture is distributed throughout all the artifacts and practices of the people who live that culture. Each piece and aspect reflects the pattern of the whole.

My use of the hologram metaphor is not accidental. There is a considerable body of research and theory which indicates that the brain stores information holographically (Karl Pribram is perhaps the most vigorous proponent of this; see Languages of the Brain, 1971). Thus it may be no accident that a man's essence shows in everything he is or does or touches. That is so because that's how the brain works. The brain also encodes culture.

For culture ultimately resides in the brains of those who carry the culture. If those brains store information so that each physical part reflects the pattern of the whole, then that is how they will organize culture. It is thus no accident that students of culture from Ruth Benedict through Claude Levi-Strauss and Clifford Geertz and Erik Erikson find correspondences, for example, between the pattern of a wedding ceremony and the layout of a dwelling, between patterns of infant weaning and adult aggression, and so forth. Cultures form coherent patterns because the brains of culture-bearers seek to impose coherence everywhere.

Similarly, each individual carries a low-res version of the entire culture. Oh, each person is an expert (high res) in this or that aspect of culture; but has only a nodding acquaintance with most of it. And there is some body of knowledge and practice all hold more or less in common in some reasonable detail. But the full culture in all its richness exists only in the interactions among all the individuals.

Monday, January 15, 2024

Toward a Theory of Intelligence: Did Miriam Yevick know something in 1975 that Bengio, LeCun, and Hinton did not know in 2018?

One theme that comes up in various discussions of artificial intelligence is that the discipline is primarily an empirical one that lacks theoretical grounding. The default view, and perhaps the dominant one as well, is that what we’re doing is producing results so damn the torpedoes – full speed ahead! But the call to theory keeps nagging, perhaps most recently in a panel discussion entitled Research on Intelligence in the Age of AI, and hosted by MIT’s Center for Minds, Brains, and Machines on its 10th Anniversary.

One theme that has been kicking around for several decades is that there are two styles of computational regime underlying perception, action, and cognition. My purpose is to compare the views that Miriam Lipschutz Yevick articulated about this dichotomy in 1975 and 1978 with those articulated by Yoshua Bengio, Yann LeCun, and Geoffrey Hinton in their 2018 Turing Award lecture, which was published in 2021.

Bengio, LeCun, and Hinton, 2018

Let’s start with Bengio, LeCun, and Hinton, who won the Turing Award in 2018. They published their paper, Deep learning for AI, in 2021. In that paper they asserted:

There are two quite different paradigms for AI. Put simply, the logic-inspired paradigm views sequential reasoning as the essence of intelligence and aims to implement reasoning in computers using hand-designed rules of inference that operate on hand-designed symbolic expressions that formalize knowledge. The brain-inspired paradigm views learning representations from data as the essence of intelligence and aims to implement learning by hand-designing or evolving rules for modifying the connection strengths in simulated networks of artificial neurons.

In the logic-inspired paradigm, a symbol has no meaningful internal structure: Its meaning resides in its relationships to other symbols which can be represented by a set of symbolic expressions or by a relational graph. By contrast, in the brain-in- spired paradigm the external symbols that are used for communication are converted into internal vectors of neural activity and these vectors have a rich similarity structure. Activity vectors can be used to model the structure inherent in a set of symbol strings by learning appropriate activity vectors for each symbol and learning non-linear transformations that allow the activity vectors that correspond to missing elements of a symbol string to be filled in. This was first demonstrated in Rumelhart et al. on toy data and then by Bengio et al. on real sentences. A very impressive recent demonstration is BERT, which also exploits self-attention to dynamically connect groups of units, as described later.

As I said at the beginning, some such characterization of two modes of thinking has been around for some time, though it is expressed in various ways. I have no problem recognizing such a distinction.

Yevick, 1975 and 1978

Miriam Yevick recognized that distinction in her 1975 paper, Holographic or Fourier Logic (Pattern Recognition 7, 187-213). That was at the peak of interest in logic-inspired AI. That was the year Newell and Simon won the Turing Award; their paper, Computer Science as Empirical Inquiry: Symbols and Search, was published the following year.

Yevick was a mathematician, not a cognitive scientist, and had become interested in optical holography though her extensive correspondence with David Bohm, the physicist, during the 1950s. During the 1960s a number of thinkers, including the neuroscientist, Karl Pribram, and the cognitive scientist, Chrisopher Longuet-Higgins, had become interested in holography as a model for neural processing. It’s that interest the Yevick had in mind when she wrote her article. Here is one statement from that article:

It has recently been conjectured that neural holograms enter as units in the thought process. If holographic processes do occur in the brain and are instrumental in thought, then the logical operations implicit in these processes could be considered as intuitive and enter as units in our mental and mathematical computations.

It has also been said that: “if we want the computer to have eyes, we shall first have to give him instruction in the facts of life”.

We maintain in this paper that a language of thought in which holographic operations enter as primitives is essentially different from one in which the same operations are carried out sequentially and hence over a finite time span [...] Our assumption is that “holographic thought” utilizes the associative properties of holograms in “one shot”. Similarly we maintain that apprehension proceeds from the very beginning via two modes, the aural and the optical; whereas the verbal string is natural to the first, the pattern as such is natural to the second: the essentially instantaneous nature of the optical process captures the apprehension as a global unit whose meaning is expressed in the first place in terms of “associations” with other such units.

There we have our distinction, between aural, verbal, and sequential on the one hand and optical, intuitive, and pattern on the other.

Having chosen visual objects as her domain, she argues thus (and here I am quoting from a 1978 restatement):

We can explicate this proposition on a theoretical level in the domain of optical patterns. [...] Such patterns or objects are thin, white regions on a black background. These can be simple (regular), like the outlines of rectangles; or complex, like the outlines of Chinese characters or random-like motions. The following holds true: a complex object requires a long (sequential, quasi-linguistic) description but yields a sharp recognition (auto-correlation) spot under holographic filtering; hence it is identified most readily by holographic recognition, or holistically. A simple object requires a short (quasi-linguistic) description but yields a diffuse recognition spot; hence it is identified most readily by quasi-linguistic representation or description.

Description and holographic recognition thus appear as two (complementary) modes of identifying an object: the more complex the object, the longer its description and the sharper its auto-correlation spot, and vice versa. The more complex they physiognomy of a person, the more unique, and hence sharper, its identity and ease of recall; the more simple, the more common and hence “unidentifiable.” Perfect holographic recognition obtains for a totally “random object”, that is, one with an infinitely long description; for a perfectly sharp point the opposite is true.

Suppose that one is given a store of objects with which one is familiar, a holographic recognition device, and a quasi-linguistic mode of representation; one is then presented with an arbitrary object to be “identified.” An approximate match is obtained either by producing a description of acceptable length or by holographic recognition of a subset of similar (associated) objects from the store. The mode of identification that will be more appropriate then depends on the complexity of the unknown object. If it is simple, we ”know” it by a short linguistic description; if it is complex, by the “associations” it evokes.

What Yevick is explicit about, and what is missing from Bengio, LeCun, and Hinton, is the relationship between some object of perception and cognition and the computational regime operating on that object. She recognizes the utility of both regimes, but associates them with different kinds of objects. As Bengio, LeCun, and Hinton simply do not conceptualize that relationship it is not clear how they would respond to Yevick’s work.

As I recall, that relationship was beginning to be recognized as an issue. Early AI had achieved its successes from dealing with sequential symbolic processing (e.g. theorem proving, expert systems), but faltered when dealing with visual perception and speech recognition. Though I can’t offer a citation, I recall David Marr, who died in 1980, mentioning the problem. The best-known statement of the problem is by Hans Moravec in his 1988 book, Mind Children, where he says “it is comparatively easy to make computers exhibit adult level performance on intelligence tests or playing checkers, and difficult or impossible to give them the skills of a one-year-old when it comes to perception and mobility”(as quoted in Wikipedia). While it is generally recognized within computer science at large, that different kinds of computational system are suited to different kinds of problems, so far as I know the issue has not be systematically investigated in the context of artificial intelligence and machine learning. More specifically, Miriam Yevick’s work from the 1970s has not been taken into account.

What needs to be done

Lots.

Let me repeat: Lots.

For one thing, Yevick’s work has been forgotten. It needs to be revived and vetted in view of more recent work.

Moreover, while her mathematics concentrated on one problem, object identification, in the visual domain, she informally generalized that result to thinking in general. For example (I’m quoting from her 1978 paper):

If we consider that both of these modes of identification enter into our mental processes, we might speculate that there is a constant movement (a shifting across boundaries) from one mode to the other: the compacting into one unit of the description of a scene, event, and so forth that has become familiar to us, and the analysis of such into its parts by description. Mastery, skill and holistic grasp of some aspect of the world are attained when this object becomes identifiable as one whole complex unit; new rational knowledge is derived when the arbitrary complex object apprehended is analytically described.

I’m certainly sympathetic to that generalization. It’s what David Hays and I had in mind when we called on Yevick’s ideas in a 1987 paper on metaphor [1] and a 1988 paper on the brain and human intelligence [2].

For all I know, Bengio, LeCun, and Hinton might be sympathetic as well. Here’s the final paragraph of their Turing Award paper:

How are the directions suggested by these open questions related to the symbolic AI research program from the 20th century? Clearly, this symbolic AI program aimed at achieving system 2 abilities, such as reasoning, being able to factorize knowledge into pieces which can easily recombined in a sequence of computational steps, and being able to manipulate abstract variables, types, and instances. We would like to design neural networks which can do all these things while working with real-valued vectors so as to preserve the strengths of deep learning which include efficient large-scale learning using differentiable computation and gradient-based adaptation, grounding of high-level concepts in low-level perception and action, handling uncertain data, and using distributed representations.

That sounds like a call to reconstruct symbolic capabilities in the context of more realistic models of real neural networks. That also sounds like intellectual work for several generations of researchers. As Charlie Parker was fond of saying, “Now’s the Time.”

References

[1] William Benzon and David Hays, Metaphor, Recognition, and Neural Process, The American Journal of Semiotics, Vol. 5, No. 1 (1987), 59-80. https://www.academia.edu/238608/Metaphor_Recognition_and_Neural_Process.

[2] William Benzon and David Hays, Principles and Development of Natural Intelligence, Journal of Social and Biological Structures, Vol. 11, No. 8, July 1988, 293-322. https://www.academia.edu/235116/Principles_and_Development_of_Natural_Intelligence.

Tuesday, November 28, 2023

Distinctive features in phonology and "polysemanticity" in neural networks

Scott Alexander has started a discussion of recent paper mechanical interpretability paper over at Astral Codex Ten: Towards Monosemanticity: Decomposing Language Models With Dictionary Learning. In a response to a comment by Hollis Robbins I offered these remarks:

Though it is true, Hollis, that the more sophisticated neuroscientists have long ago given up any idea of a one-to-one relationship between neurons and percepts and concepts (the so-called "grandmother cell") I think that Scott is right that "polysemanticity at the level of words and polysemanticity at the level of neurons are two totally different concepts/ideas." I think the idea of distinctive features in phonology is a much better idea.

Thus, for example, English has 24 consonant phonemes and between 14 and 25 vowel phonemes depending on the variety of English (American, Received Pronunciation, and Australian), for a total between 38 and 49 phonemes. But there are only 14 distinctive features in the account given by Roman Jakobson and Morris Halle in 1971. So, how is it the we can account for 38-49 phonemes with only 14 features?

Each phoneme is characterized by more than one feature. As you know, each phoneme is characterized by the presence (+) of absence (-) of a feature. The relationship between phonemes and features can thus be represented by matrix having 38-49 columns, one for each phoneme, and 14 rows, one for each row. Each cell is then marked +/- depending on whether or not the feature is present for that phoneme. Lévi-Strauss adopted a similar system in his treatment of myths in his 1955 paper, "The Structural Study of Myth." I used such a system in one of my first publications, "Sir Gawain and the Green Knight and the Semiotics of Ontology," where I was analyzing the exchanges in the third section of the poem.

Now, in the paper under consideration, we're dealing with many more features, but I suspect the principle is the same. Thus, from the paper: "Just 512 neurons can represent tens of thousands of features." The set of neurons representing a feature will be unique, but it will also be the case that features share neurons. Features are represented by populations, not individual neurons, and individual neurons can participate in many different populations. In the case of animal brains, Karl Pribram argued that over 50 years ago and he wasn't the first.

Pribram argued that perception and memory were holographic in nature. The idea was given considerable discussion back in the 1970s and into the 1980s. In 1982 John Hopfield published a very influential paper on a similar theme, "Neural networks and physical systems with emergent collective computational abilities." I'm all but convinced that LLMs are organized along these lines and have been saying so in recent posts and papers.

* * * * * 

Addendum: Another possible example: I use 655+ tags here at New Savanna for over 9600 posts. Some posts have only one or two tags, and some have a half dozen or more. Note, however, that I don't intend that an post's tag set function as an identifier for the post, or a proxy for an identifier. But the tags do characterize the posts in some way.

Thursday, November 9, 2023

November Ramble: AI, Yevick, State of AI, Much Ado, photos

There’s a mixed bag of tricks on my plate these days. Let’s run through them.

LLMs and holography

This is a major chunk of work. I’ve written a bunch of material on neural-holography. I need to pull that together and elaborate on some core ideas. Perhaps the single biggest chunk of work is a more careful and explicit account of the work that Miriam Yevick did in her 1975 paper on holographic logic, which was at the center of my 3QD piece about her. She was concerned about the relationship between the properties of some object or phenomenon that was to be treated computationally and the formal properties of a computational system, with one style of computation (holographic, ‘one shot’) being suitable for a certain class of phenomena ((geometrically) complex) and a different style (logical, sequential) being more suited for a contrasting class of phenomena ((geometrically) simple).

As far as I know she’s the first one to think explicitly in those terms, though the issue certainly lurks very close to the surface of what’s come to be called Moravec’s paradox. As Steven Pinker put it (I’m quoting him from the Wikipedia article), “the main lesson of thirty-five years of AI research is that the hard problems are easy and the easy problems are hard.” If you look closely at the various examples, the (almost) all have to do with the relationship between computational regime and application domain, but that doesn’t quite seem to have been stated, apart from Yevick’s work. So, give a fuller account of her work, but also talk about why her work has not been recognized and taken-up.

That’s one aspect of the project. A second aspect is to look over the work I’ve done with ChatGPT and explain how neural holography is a reasonable approach to accounting for 1) the story work, and 2) the memory work. And the third, is this that and the other, such as my recent post about IS-A sentences.

I have no idea how much work this will entail. For more details, see my last ramble, Ramble on ChatGPT: Coming up on a one-year anniversary, time to reflect on ChatGPT & LLMs. I hope to have this done by early December.

Yevick

I want to produce a working paper on Miriam Lipschutz Yevick, one oriented toward her life, as she talks of it in A Testament for Ariel. This will be based on my 3DD piece about her, Next Year in Jerusalem: The Brilliant Ideas and Radiant Legacy of Miriam Lipschutz Yevick. I’ve already written some blog posts which I can add to that. I plan one more, about the overall construction and literary merit of her Testament.

O Captain, My Captain! On the state of AI [3QD]

That’s the theme of my 3 Quarks Daily article for December 4, 2023. I’ll be opening with the following conceit: Would you invest in a whaling voyage captained by someone who knows all there is to know about his ship, and is able to helm it in a day sail to and from home port, but whose knowledge of sailing on the open ocean, of the weather, of navigation and, above all else, of whales and whaling, is no greater than that of the typical landlubber? I doubt that you’d consider that a prudent investment. But that, so the conceit goes, is the current state of artificial intelligence.

What’s the point of the conceit? The whaling captain is primarily a figure for an expert in machine learning, but also for executives of companies in the AI business. As far as I can tell, experts in machine learning know a great deal about how to construct, use, and maintain machine learning systems, but that don’t know much about the phenomena in the domains where those systems operate. Thus, experts in large language models (LLMs) don’t know much about language, cognition, or philosophy, no more, say, than a bright sophomore at a good school. And yet they make confident assertions about machine consciousness, when they’ll get to AGI, how scaling up is all we need to do, etc. So, I argue that.

What’s that mean for business? Well, if you want to invest in AI, that’s what you’ve got to invest in, because there isn’t much else. There aren’t many AI researchers with deep knowledge of both the technology and the application domain (language is my main interest), nor of research teams rich in both kinds of expertise.

What’s that mean? I don’t know. It could mean that most current investments will fail, though whether they do so at a higher rate than is typical of early investment is not something I’m willing to take a guess on. There are niches that will support successful products using current technology, built by those asymmetrically knowledgeable sea captains. Long-term, though, it’s a different story.

As a comparison, consider the scientific revolution in astronomy: Copernicus, to Kepler, to Newton (deriving orbits from his laws of motion), to the 19th century discovery of Uranus. Think of the current moment in AI as the Copernican moment. We’re going to need new ideas to reach the Keplerian and Newtonian moments, and those new ideas will have to include detailed knowledge of phenomena in the application domain (e.g. language and cognition). Perhaps we can think in Hegelian terms, with old school symbolic AI being the thesis, machine learning the antithesis, and something else the synthesis.

Much Ado About Nothing

This is mostly a reminder to myself. Years ago, when I was in graduate school, I wrote a nice paper on ritual structure in Shakespeare’s Much Ado About Nothing. At various times I’ve started work on developing that into a more sophisticated piece of work. I need to get back to it, though just when, I don’t know.

And beyond there a book based on my old article about Much Ado About Nothing, Othello, The Winter’s Tale and The Tempest. I have no idea when, if ever, I’ll get around to that.

Photo Exhibit

Another reminder. I’ve been planning to put together a small pop-up exhibit of some of my Hoboken photos. I need to make a final selection, from these photos (plus one or two others – click on right or left edge to scroll the photos):

Hoboken Final

I’ve got to get prints made, and frame the pictures. At the same time I’ve got to line up some venues to exhibit the photos. Lots of work to do.

‘Till later.

Wednesday, November 8, 2023

What’s going on? LLMs and IS-A sentences

For the moment I have decided that Waddington’s classic diagram of the epigenetic landscape is a useful way of thinking about when happens when an LLM responds to a prompt. Here’s the diagram:

The language model corresponds to the landscape. The prompt serves to position that ball at a certain place in the landscape – perhaps we can think of that ball as the prompt. The ball then rolls down the valley, going left and right as appropriate. It never reverses direction and goes up the hill. That path, or trajectory if you will, is the LLM’s response to the prompt.

Moreover, I have decided to think of the generation of each word (yes, I know, technically it spits out tokens, not words) as a single primitive operation. That is to say, it has no internal logical structure, no ANDs or ORs. It’s simply one (gigantic) calculation over roughly 175 billion values (in the case of ChatGPT). The generation of each word presents the system with a choice among alternatives, but that’s the only kind of choice involved in calculating the response to a prompt – though for qualification and elaboration, see ChatGPT tells stories, and a note about reverse engineering: A Working Paper, Version 3, pp. 3-6.

That brings me to something I’ve been puzzled about for years. We find it natural to say things like, Garfield is a cat. Now, express the same thought, but reverse the order of cat and Garfield in your sentence. It’s difficult to do. Oh, you can do it, but the resulting sentence is awkward and unnatural, something like, Cats are the kind of thing of which Garfield is a particular instance. No one would ever speak like that, nor write it either.

What’s the source of that asymmetry? As far as I can tell, we don’t know, but I take it as a clue about the mechanisms of language. The purpose of this note is to suggest that my crude model of LLM calculation would provide an answer: The linguistic landscape is structured so that the ball easily rolls from Garfield to cat, or cat to mammal, Snoopy to beagle, Tesla to EV, C. elegans to worm, etc. One might, of course, as why the landscape is arranged in that way, but that’s a different question, no?

Here’s some notes I made about IS-A sentences.

Notes on IS-A Sentences

Somewhere in his Problems in General Linguistics, my copy of which is, alas, in storage, Emile Benveniste has a chapter, “The Nominal Sentence,” on sentences hanging on the auxiliary “to be.” As Benveniste was a linguist of the Old School, when being a linguistic meant familiarity with many languages, including—and this is important for this particular topic—classical Greek, it had examples from many languages, making it tough sledding for a monoglot like me.

While the content of this post certainly arises out of my thinking about that chapter, in the absence of actually having the text in front of me, I hesitate to assert a stronger relationship than that. I note only that, for Benveniste, the auxiliary “to be” was fraught with metaphysical significance. For the concept of being derives from “to be.” Where would philosophy be without Being? Thus, when Benveniste pondered such sentences, he wasn’t merely commenting on language. He was doing philosophy, or, if not quite that, camping out on philosophy’s door step.

I’m interested in such sentences because I believe they are a DEEP CLUE about how the mind works. I just don’t know what to make of the clue.

So, I'm interested in word order in assertions such as the following:

(1) Fido is a beagle.
(2) Beagles are dogs.
(3) Dogs are beasts.

They all move from an element in a class (whether an individual, Fido, or another class, beagles) to a class containing it. None of them move in the opposite direction. Consider what happens when you try to go the opposite way. In the following sentence the class is mentioned first, then the subclass:

(4) Beagle is the kind of animal of which Fido is an instance.

In particular, note that (4) has a metalingual character that (1) does not. That is, (4) explicitly asserts that we are dealing with classification. One can do that metalingual job in various ways, but, as far as I can tell, one can't avoid it. That is, one cannot construct a proper English sentence relating a genus and species in which the genus is mentioned first, one can’t do that without ‘looping through’ some kind of metalingual construction on the way from genus to species.

Why?

What does this assymetry tell us about the underlying mechanisms? Why don't have sentences such as:

(5) Beagle za di Fido.

In this case "za di" is the inverse of "is a". English has no such sentences & no such inverse.

So, how widespread is this asymmetry and is there any explanation of this directionality?

I sent a query on that matter to a listserve, I forget which one, and got two replies that add some complexity to the matter. Rich Rhodes, Linguistics at UCal Berkeley, tells me that in Ojibwe the word order is reversed, the class comes before the individual, but the asymmetry remains. He then comments, which he qualifies as a quick guess:

My guess is that there is no compelling discourse function (like information flow) which makes it desirable to invert classificational equatives. Hence we only get the "unmarked" order. Subject-predicate in theme-rheme languages (like English) and predicate-subject in rheme-theme languages (like Ojibwe).

So, what's the nature of the mechanism that determines the "unmarked" order? That's what I want to know.

Lee Pearcy, Episcopal Academy in Merion, Pa. offered these examples:

(6) The beagle is Fido.
(7) The dogs are beagles.
(8) The beasts are dogs.

As stand-alone sentences, they seem a bit awkward to me. But they fare better as answers to questions, e.g.:

What’s that dog?
Which dog? The beagle is Fido and the terrier is Max.

What’re those animals?
The dogs are beagles, the cats are Persians.

In those contexts, the matter of class or classification is raised by the question, thus making it present in the discourse and so available as a point of attachment in the answer.

Further clues, anyone?

Do I believe this?

I don’t believe it, or disbelieve it. It’s a working hypothesis. One I think is worth investigating. It places relatively simple and severe constraints on our conception of what LLMs are doing. That, it seems to me, is a good thing. Should it turn out that those constraints are valid, well then, we’ve learned something, no? If they’re not valid, we’ve also learned something.

More later.

Friday, October 20, 2023

Are (at least some) Large Language Models Holographic Memory Stores?

That’s been on my mind for the last week or two, ever since my recent work on ChatGPT’s memory for texts [1]. On the other than, there’s a sense in which it’s been on my mind for my entire career, or, more accurately, it’s been growing in my mind ever since I read Karl Pribram on neural holography back in 1969 in Scientific American [2]. For the moment let’s think of it as a metaphor, just a metaphor, nothing we have to commit to. Just yet. But ultimately, yes, I think it’s more than a metaphor. To that end I note that cognitive psychologists have recently been developing the idea of verbal memory as holographic in nature [3]. 

Note: These are quick and dirty notes, a place-holder for more considered thought.

Holography in the mind

Let’s start with an article David Hays and I published on neural holography as the neural underpinning of metaphor [4]. Here’s where we explain the holographic process:

Holography is a photographic technique for making images. A beam of laser light is split into two beams. One beam strikes the object and is reflected to a photographic plate. The other beam, called a reference beam, goes from laser to plate directly. When they meet, the two beams create an interference pattern—imagine dropping two stones into a pond at different places; the waves propagating from each of these points will meet and the resulting pattern is an interference pattern. The photographic plate records the pattern of interference between the reference beam and the reflected beam.

The image recorded on the film doesn't look at all like an ordinary photographic image—it’s just a dense mass of fine dots. But when a beam of laser light having the same properties as the original reference beam is directed through the film an image appears in front of the film. The interaction of the laser beam and the hologram has recreated the wave form of the laser beam which bounced off the object when the hologram was made. The new beam has extracted the image from the plate.

Holography is, as its name suggests, holistic. Every part of the scene is represented in every part of the plate. (This situation is most unlike ordinary photography, which uses a good lens to focus infinitesimal parts of the scene onto equally infinitesimal parts of the plate.) With such a determinedly nondigital recording, certain mathematical possibilities can be realized more easily—we are tempted to say, infinitely more easily. For example, convolution. Take the holographic image of a printed page, and the image of a single word. Convolute them. The result is an image of the page with each occurrence of the word highlighted. We can think of visual recognition as a kind of convolution. The present scene, containing several horses, is convoluted with the memory of a horse and the present horses are immediately recognized. We can think of recognition this way, but we must admit that this process has not been achieved in any machine as yet.

Further, it is possible to record many different images on the same piece of film, using different reference beams. The reference beams may differ in color, in angle of incidence, or otherwise. We can think— although again we cannot cite a demonstration—of convoluting such a composite plate with a second plate. If the image in the second plate matches any one of the images in the composite, then it is recognized. For metaphor we want to convolute Achilles and the lion and to recognize, to elicit another image containing not Achilles, not the lion, but just that wherein they resemble one another. Such is the metaphor mechanism—but that must wait until the next section, on focal and residual schemas.

The 175 billion weights that constitute the LLM at the core of ChatGPT, that’s the holographic memory. It is the superposition of all the texts in the training corpus. The training procedure – predict the next word – is a device for calculating a correlation (entanglement [5]) between each word in context, and every other word in every other text, in context. It’s a tedious process, no? But it works, yes?

When one prompts a trained memory, the prompt serves as a reference beam. And the whole memory must be ‘swept’ to generate each character. Given the nature of digital computers, this is a somewhat sequential process, even given a warehouse full of GPUs, but conceptually it’s a single pass. When one accesses an optical hologram with a reference beam, the beam illuminates the whole holograph. This is what Miriam Yevick called “one-shot” access in her 1975 paper, Holographic or Fourier Logic [6]. The whole memory is searched in a single sweep.

Style transfer

So, that’s the general idea. Much detail remains to be supplied, most of it by people with more technical knowledge than I’ve got. But I want to get in one last idea from the metaphor paper. We’ve been explaining the concepts of focal and residual schemas:

Now consider a face. Everything we said about the chair applies here as well. But the expression on the face can vary widely and the identity of the face remains constant. This variability of expression can also be handled by the mechanism of focal and residual. There is a focal schema for face-in-neutral-expression and then we have various residuals which can operate on the focal schema to produce various expressions. (You might want to recall D'Arcy Thompson's coordinate transformations in On Growth and Form 1932.) We tend to discard presentation residuals such as lighting and angle of sight, but we respond to expression residuals

Our basic point about metaphor is that the ground which links tenor and vehicle is derived from residuals on them. Consider the following example, from Book Twenty of Homer's Iliad (Lattimore translation, 1951, ll. 163-175)—it has the verbal form of a simile, but the basic conceptual process is, of course, metaphorical:

                          From the other
side the son of Peleus rose like a lion against him,
the baleful beast, when men have been straining to kill him, the country
all in the hunt, and he at first pays them no attention
but goes his way, only when some one of the impetuous young men
has hit him with the spear he whirls, jaws open, over his teeth foam
breaks out, and in the depth of his chest the powerful heart groans;
he lashes his own ribs with his tail and the flanks on both sides
as he rouses himself to fury for the fight, eyes glaring,
and hurls himself straight onward on the chance of killing some one
of the men, or else being killed himself in the first onrush.
So the proud heart and fighting fury stirred on Achilleus 
to go forward in the face of great-hearted Aineias.

In short, Achilles was a lion in battle. Achilles is the tenor, lion the vehicle, and the ground is some martial virtue “proud heart and fighting fury”. But what of that detailed vignette about the lion's fighting style? Whatever its use in pacing the narrative, its real value, in our view, is that it contains the residuals on which the comparison rests, the residuals which give it life. The phrase “proud heart and fighting fury” is propositional while the fighting style is physiognomic. “Proud heart and fighting fury” may convey something of what is behind the fighting style, but only metaphoric interaction can foreground the complex schema by which we recognize and feel that style.

The cognitive problem is to isolate the physiognomy of style, to tease it apart from the entities which exhibit that style. [...] In the case of Achilles and the lion we have two complex physiognomies, each extended in space and time. Metaphoric comparison serves to isolate the style, to allow us to focus our attention on that style as distinct from the entities which exhibit it.

This comparison involves two foci, Achilles and the lion. The physical resemblance between them is not great—their body proportions are quite different and the lion is covered with fur while Achilles is, depending on the occasion, either naked or clothed in some one of many possible ways. The likeness shows up in the way they move in battle. A body in motion doesn't appear the same as a body at rest. The appearance presented by the focal body is modified by the many residuals which characterize that body's movement— twists and turns, foreshortenings and elongations (for an account of motion residuals, see Hay 1966). The movements of Achilles and the lion must differ at the grossest level, since the lion stands on four legs and fights with claws and teeth, while Achilles stands on two legs and fights with a spear or sword. But their movements are alike at a subtler level, at the level of what we call, in a dancer or a fighter, their style. Residuals can be stacked to many levels. “Proud heart and fighting fury” may be a good phrase to designate that style, but it doesn't allow us to attend to that style. Homer's extended simile does.

That’s a mouthful, I know. Notice our emphasis on style. That’s what’s got my attention.

One of the more interesting things LLMs can do is stylistic transfer. Take a piece of garden variety prose and present it in the style of Hemingway or Sontag, whomever you choose. Hays and I argued that that’s how metaphor is created, deep metaphor, that is, not metaphor so desiccated we no longer register its metaphorical nature, e.g. the mouth of the river. We made our argument about visual scenes: Achilles in batter, a lion in battle. LLMs apply the same process to texts, where style is considered to be a pattern of residuals over the conceptual content of the text.

More later.

References

[1] Discursive Competence in ChatGPT, Part 2: Memory for Texts, Version 3, https://www.academia.edu/107318793/Discursive_Competence_in_ChatGPT_Part_2_Memory_for_Texts_Version_3

[2] I recount that history here: Xanadu, GPT, and Beyond: An adventure of the mind, https://www.academia.edu/106001453/Xanadu_GPT_and_Beyond_An_adventure_of_the_mind

[3] Michael N. Jones and Douglas J. K. Mewhort, Representing Word Meaning and Order Information in a Composite Holographic Lexicon, Psychological Review, 2007, Vol. 114, No. 1, 1-37. DOI: https://doi.org/10.1037/0033-295X.114.1.1

Donald R. J. Frankin and D. J. K. Mewhort, Memory as a Holograpm: An Analysis of Learning and Recall, Canadian Journal of Experimental Psychology / Revue canadienne de psychologie expérimentale, Association 2015, Vol. 69, No. 1, 115–135, https://doi.org/10.1037/cep0000035

[4] Metaphor, Recognition, and Neural Process, https://www.academia.edu/238608/Metaphor_Recognition_and_Neural_Process

[5] See posts tagged with “entangle”, https://new-savanna.blogspot.com/search/label/entangle

[6] Miriam Lipschutz Yevick, Holographic or Fourier Logic, Pattern Recognition 7, 197-213, https://sci-hub.tw/10.1016/0031-3203(75)90005-9

Thursday, October 19, 2023

Ramble on ChatGPT: Coming up on a one year anniversary, time to reflect on ChatGPT & LLMs

ChatGPT was released on November 31, 2022, and I started playing around with it on December 1, 2022. Perhaps early December 2023 would be a good time for me to reflect on the work I’ve been doing on ChatGPT and related matters, no?

Yes.

The purpose of this post is to think that through in an informal way. What do I need to do between now and then? What will that document look like?

I figure that document will itself be modest, no more than, say, 30 or so pages of exposition. But it will link to all my working papers on ChatGPT and related matters. The rest of this post consists of thoughts about the work I need to do and concludes with a list of those working papers.

Ontology and meaning

I’ve recently done some posts on conceptual ontology in ChatGPT, 20 questions and ontological landscape. I’m in the process of combining those posts with some older material and producing a working paper, ChatGPT’s Ontological Landscape. That’s an important piece of work because ontology (“natural kinds”) is one of the major structuring principles underlying language and cognition.

This will include the particular idea that conceptual systems each has its own underlying ontology and that specialized ontologies supersede common sense ontology is specialized domains. My prototypical example is salt, a common-sense concept, and NaCl, a specialized concept. The common-sense is defined in terms of sensory perception while the specialized concept is defined in the language of chemistry, atoms and chemical bonds.

The concept of “meaning” is like this as well. It is a common-sense concept. But it is also used in various specialized domains in philosophy, linguistics, cognitive science, and computer science and AI. Much of the confusion about whether or not AI systems can deal with linguistic meaning results from the fact that these various concepts all fly under the same linguistic flag, meaning, but they are no more than same concept than salt and NaCl are.

Miriam Yevick and neural holography

I recently published a long article in 3 Quarks Daily about the life and ideas of a mathematician named Miriam Yevick. I need to write a (longish) more detailed post about her ideas that will also: 1) reflect on why they weren’t taken up, 2) get a bit more explicit about how they relate to current issues in machine learning and large language models. One thing I need to do is argue that her characterization of “one-shot” processing in holography applies to transformers in the following way: each sweep through the collection of weights in the process of selecting the next token, that is equivalent to “one-shot” in an optical holography apparatus. I also need to discuss recent work in cognitive psychology which uses a holographic model for language and verbal memory.

Discursive supplement

I need to finish this working paper, Discursive Competence in ChatGPT, Supplemental Examples to Part 1. This consists of a variety of transcripts of ChatGPT interactions. Some topics: grammatical knowledge and self-reference, the abstract concept of charity, haiku and Margaret Masterman, legal concepts, word associations and word clusters, and the Chinese Room thought experiment.

The Limitations of Large Language Models

I see that LLMs have three limitations that are inherent in the architecture. You can’t eliminate them by scaling up or by fine tuning, prompt engineering, and RLHF. This is, of course, a matter controversy. I don’t intend to argue the issue (at least not much), but rather I just want to state it and think about the consequences:

The limitations:

  1. Once it has been trained, the model is fixed and cannot (readily) be altered.
  2. It confabulates.
  3. It is confined to single-stream processing, which is the source of its weakness in arithmetic, ‘tight’ logical reasoning, and planning (among others).

As it is, LLMs can function as a processor in various configurations, where they are linked to various external applications. A great deal of working is being done in this area. Some of it may result in useful products, but this will never lead to the mythical AGI.

Providing LLMs with symbolic capabilities should deal with the third problem, and those capabilities could also be a means of linking it to world model, which will deal with the second problem. The world model itself must be maintained and ultimately be subject to human oversight. The first problem requires a different architecture, a ‘looser’ one, but one that doesn’t squander the power of LLMs to embrace a wide range of material – see, in particular, GPT-3: Waterloo or Rubicon? pp. 23-26 (link below).

I tend to think of LLMs as digital wilderness, large webs and tangles of conceptual structure that needs to be explored and ‘domesticated.’ Just what that entails....I note as well that this involves matters of social, political, and economic organization that are well beyond the scope of this review. I do see opportunities here for citizen science.

Beyond mechanistic understanding

Mechanistic understanding is necessary, but not sufficient for understanding what’s going on inside LLMs. I see them as being organized on three levels, which I’m currently calling, phenomenon, matrix, and engine. The phenomenal level is language and texts. That’s what we see, and how we interact with the LLM. The engine is the computer code that runs the device, both in training and inference mode. The matrix is the model itself, which is generally said to be opaque.

Mechanistic understanding, as I understand it, is focused on the interface between the engine and the model. Much of my work is directed at providing clues about the interface between the phenomenon and the model. I don’t believe that you can get at that interface through mechanistic understanding alone. It’s not the right conceptual tool, the right language.

Ultimately, I believe some kind of specialized conceptual tools will be needed to understand the interface between language and the model. I think that the geometric semantics of Peter Gärdenfors will be useful here (Conceptual Spaces: The Geometry of Thought, 2000; The Geometry of Meaning: Semantics Based on Conceptual Spaces, 2014). We’ll also need to develop some kind of graphic notation. I’m thinking of starting with Syd Lamb’s notation, which I’ve used in Relational Nets Over Attractors, A Primer (see list below). Discovering and developing these conceptual tools will be one objective of a research program.

Research program

And THAT’s the major objective of this exercise, to come up with a program for further research. I’ve tried a lot of things with ChatGPT this past year. At the moment the three most productive seem to be:

  • Systematic story variations, which I’ve written up in ChatGPT tells stories, and a note about reverse engineering: A Working Paper (link below)
  • Memory structures, which I’ve written up in Discursive Competence in ChatGPT, Part 2: Memory for Texts (link below)
  • Twenty questions, where I’ve got a long blog post, and which I’ll be discussing in the ontology working paper.

I’ll probably say a word or three about some other possibilities as well.

One important thing about this research is that it doesn’t require API access to the ChatGPT, or any other LLM, and it doesn’t require programming. Anyone who has internet access to ChatGPT can do it.

Working Papers on ChatGPT

GPT-3: Waterloo or Rubicon? Here be Dragons, Version 4.1, https://www.academia.edu/43787279/GPT_3_Waterloo_or_Rubicon_Here_be_Dragons_Version_4_1

Discursive Competence in ChatGPT, Part 1: Talking with Dragons, Version 2, https://www.academia.edu/94409729/Discursive_Competence_in_ChatGPT_Part_1_Talking_with_Dragons_Version_2

ChatGPT vs. the Towers of Warsaw, https://www.academia.edu/94517239/ChatGPT_vs_the_Towers_of_Warsaw

ChatGPT intimates a tantalizing future; its core LLM is organized on multiple levels; and it has broken the idea of thinking. Version 3, https://www.academia.edu/95608526/ChatGPT_intimates_a_tantalizing_future_its_core_LLM_is_organized_on_multiple_levels_and_it_has_broken_the_idea_of_thinking_Version_3

ChatGPT tells stories, and a note about reverse engineering: A Working Paper, Version 3, https://www.academia.edu/97862447/ChatGPT_tells_stories_and_a_note_about_reverse_engineering_A_Working_Paper_Version_3

Stories by ChatGPT: Fairy Tale, Realistic, and True, https://www.academia.edu/99985817/Stories_by_ChatGPT_Fairy_Tale_Realistic_and_True

ChatGPT tells 20 versions of its prototypical story, with a short note on method, https://www.academia.edu/108129357/ChatGPT_tells_20_versions_of_its_prototypical_story_with_a_short_note_on_method

Discursive Competence in ChatGPT, Part 2: Memory for Texts, Version 3, https://www.academia.edu/107318793/Discursive_Competence_in_ChatGPT_Part_2_Memory_for_Texts_Version_3

Working Papers on related issues

To Model the Mind: Speculative Engineering as Philosophy, https://www.academia.edu/75749826/To_Model_the_Mind_Speculative_Engineering_as_Philosophy

Xanadu, GPT, and Beyond: An adventure of the mind, https://www.academia.edu/106001453/Xanadu_GPT_and_Beyond_An_adventure_of_the_mind

Symbols and Nets: Calculating Meaning in "Kubla Khan", https://www.academia.edu/78967114/Symbols_and_Nets_Calculating_Meaning_in_Kubla_Khan_

Relational Nets Over Attractors, A Primer: Part 1, Design for a Mind, Version 3, https://www.academia.edu/81911617/Relational_Nets_Over_Attractors_A_Primer_Part_1_Design_for_a_Mind_Version_3

Direct Brain-to-Brain Thought Transfer A High Tech Fantasy that Won't Work, https://www.academia.edu/44109360/Direct_Brain_to_Brain_Thought_Transfer_A_High_Tech_Fantasy_that_Wont_Work

Principles and Development of Natural Intelligence, https://www.academia.edu/235116/Principles_and_Development_of_Natural_Intelligence

Metaphor, Recognition, and Neural Process, https://www.academia.edu/238608/Metaphor_Recognition_and_Neural_Process

Tuesday, October 3, 2023

What would it mean to understand how a large language model (LLM) works? Some quick notes.

I don’t mean “understand” in any deep philosophical sense. I mean only a rough and ready sense of the word. We understand how toasters work, automobiles, moon rockets, digital computers, and so forth. We know how to design and construct these things, how to diagnose problems, how to maintain and repair them. Not perfectly to be sure, but well enough to use these devices to get things done.

LLMs, however, are said to be opaque. We don’t know how they work. We feed them prompts, they produce output, but how the model works from the prompt to produce the output, that’s mysterious. There are people working on mechanical interpretability, trying to understand the LLM as though it were a machine, or at least, a computer program of the ordinary kind, where we know, more or less, how it works on data – if it is the kind of program that works from data – to produce output. But what would it mean to understand the operational characteristics of 175 billion parameters, as in the case of GPT-3.5?

It means, I suppose, how those parameters mediate between the input, a prompt, and the output, whatever “follows from” a given prompt. At the lowest level we are told that LLMs are prediction machines. So, the output string is simply a continuation of the input string. And I suppose that, technically, that’s true. But it’s not very helpful, as I’ve argued at some length.

Let’s set that aside.

What could we possibly want by way of understanding?

We’ve got three things: There is the underlying engine, let’s call it, which is a computer program like any other. It’s created by programmers working with some language or languages and is designed to achieve a certain purpose. In this case, it’s designed to create a language model over a corpus of texts and then to use that model in generating new chunks of language given an input prompt.

It's that model that’s problematic, that’s said to be opaque. We, us humans, didn’t create that model. The engine did. And, in the case of GPT-3, that model’s got 175 billion parameters. More recent models have even more. And there are also models with only millions of parameters. But even those smaller models are huge.

But, here’s the thing, how can we understand how that opaque model operates unless we understanding what it’s trying to do? Sure, we can pop the hood and take a look. We see a bunch of gizmos, widgets, framblasts, and other things, but so what? They’re just whirling around, engaging with one another, in intricate patterns? But what are they trying to do? We know what car engines are supposed to do; they supply power to the wheels (and the wheels move the car).

Well, LLMs are supposed to produce language – and computer code and math as well, but let’s stick with ordinary language for the purposes of these notes. But, alas, the mechanisms of language are themselves opaque. The relationship between a car's wheels and that car's motion is transparent. The relationship between nouns and verbs and adjectives and prepositions and sentences and, you know, knowledge, understanding, entertainment, the things language is for, those relationships are not so obvious.

Of course, linguists have been working on language mechanisms for years. But it’s not at all clear what the field has come up with. There are major disagreements on how one is to understand syntax. And when we move beyond sentences to discourse of various kinds, we know even less about mechanisms.

I figure that there’s almost zero chance that we’re going to find those mechanisms by mucking around in LLMs. Yes, I know that LLMs are quite different from the human brain and mind. But, the fact is, LLMs do a very convincing imitation of human language. Given the complexity of language, they wouldn’t be able to do that if they hadn’t absorbed some (perhaps) useful approximation to human mechanisms. I’m willing to proceed on the default understanding that, whatever the model is doing, it has some resemblance to what humans do. If I make that assumption, that gives me some tools to think with. Without it, I got nothing.

Still, a grammar is a large and complex thing. The Cambridge Grammar of the English Language is 1860 pages long, and it is merely a descriptive grammar and not meant to account for the underlying mechanisms, however they might best be characterized. Is that what we want from a mechanistic understanding of an LLM? And that only gets us sentences. What about paragraphs, stories, histories, repair manuals, accounts of exotic astronomical objects, and who knows what else? Do we expect students of mechanistic interpretability to eventually give us detailed accounts of such wonders?

Understanding stories

What would it mean to understand how ChatGPT tells stories?

This morning I logged onto ChatGPT, not GPT Plus, just plain old ChatGPT, and prompted it with one word: “Story.” What do you think it did? Right, it told me a story. The story began with this sentence: “Once upon a time, in a quaint little village nestled at the foot of a towering mountain range, there lived a young girl named Lily.” I don’t think it’s very useful to think of that sentence as the natural continuation of a string beginning with the word, “story.” Yes, I know, I’m not prompting the “naked” underling LLM. ChatGPT has been prompt-engineered and RLHFed (RLHF: reinforcement learning with human feedback) to death to be a congenial conversational partner. But that doesn’t change the basic situation.

In this case, the situation is that, in some sense, ChatGPT “knows” what a story is and knows how to tell one. By this time I’ve prompted it to produce 100s, though probably not yet 1000s of stories. In a few cases the prompt was just that one word. More often it was something like one of these:

Tell me a story.
Tell me a story about a hero.
Tell me a realistic story.
Tell me a true story about a hero.

ChatGPT also told me a well-formed story. The stories were relatively short and simple, and the first two prompts produced stories with a fairytale feel, supernatural creatures and events were typical. Those were absent in realistic stories. As for true stories, sometimes they read more like short newspaper articles than like stories.

But where did ChatGPT learn to tell stories? Well, it consumed I don’t know how many stories during training. Whatever it knows about story-telling was distilled from those stories. I note that, to a first approximation, that’s how humans learn to tell stories as well. We are told stories as toddlers and children and, in time, begin telling our own stories, based on the models we’ve been exposed to. New stories are based on old stories, on remembered and half-remembered stories.

Now, as you may know, at some point I began to have ChatGPT tell stories based on rather elaborate prompts of a simple form consisting of 1) a request to tell a new story based on an existing one, but with one change (which I specified) and 2) the existing story. For example:

Wednesday, September 27, 2023

Discursive Competence in ChatGPT, Part 2: Memory for Texts

I've finished a new working paper. Title above, links, abstract, table of contents, and introduction below.

Academia.edu: https://www.academia.edu/107318793/Discursive_Competence_in_ChatGPT_Part_2_Memory_for_Texts
SSRN: https://ssrn.com/abstract=4585825
ResearchGate: https://www.researchgate.net/publication/374229644_Discursive_Competence_in_ChatGPT_Part_2_Memory_for_Texts_2_Memory_for_Texts

Abstract: In a few cases ChatGPT responds to a prompt (e.g. “To be or not to be”) by returning a specific text word-for-word. More often (e.g. “Johnstown flood, 1889”) it returns with information, but the specific wording will vary from one occasion to the next. In some cases (e.g. “Miriam Yevick”) it doesn’t return anything, though the topic was (most likely) in the training corpus. When the prompt is the beginning of a line or a sentence in a famous text, ChatGPT always identifies the text. When the prompt is a phrase that is syntactically coherent, ChatGPT generally identifies the text, but may not properly locate the phrase within the text. When the prompt cuts across syntactic boundaries, ChatGPT almost never identifies the text. But when told it is from a “well-known speech” it is able to do so. ChatGPT’s response to these prompts is similar to associative memory in humans, possibly on a holographic model.

Contents

Introduction: What is memory? 2
What must be the case that ChatGPT would have memorized “To be or not to be”? – Three kinds of conceptual objects for LLMs 4
To be or not: Snippets from a soliloquy 16
Entry points into the memory stream: Lincoln’s Gettysburg Address 26
Notes on ChatGPT’s “memory” for strings and for events 36
Appendix: Table of prompts for soliloquy and Gettysburg Address 43   

Introduction: What is memory?

In various discussions about large language models (LLMs), such as the one powering ChatGPT, I have seen assertions that such as, “oh, it’s just memorized that.” What does that mean, “to memorize?”

I am a fairly talented and skilled musician. I can and have memorized a piece of music by practicing it over and over. There are the notes on the page. I start playing them until I am comfortable. Then I look away and see how far I can go. When I get lost, I look at the music, continue playing the notes on the page, and finish the piece – something like that. Then I start over from the beginning, again without the music. When I can play the whole piece without having to consult the written music, I have it memorized. At least for the moment.

But I don’t do that very often. More likely, I’ll hear a tune I like two, three, or five times and then I pick up my trumpet and playing, sometimes perfectly, sometimes with a glitch or two. I didn’t memorize it, and yet I’m playing it. From memory? No, by ear?

I’m speaking metaphorically of course. What does it mean to play by ear? I don’t really know, but I imagine it goes something like this: Music has an inner logic, a grammar, a set of rules through which it is structured. When I hear a tune I’m listening to it in term of that inner logic, as, for that matter, anyone is – at least if they’re familiar with the musical idiom. It’s that logic that I’m registering as I listen to the tune. Once I’ve heard the tune a couple of times, I’ve “absorbed” that logic, without even thinking about it or working on it. It just happens as a side-effect of (ordinary) listening. When the absorption is complete, I am able to play the tune “by ear.”

Those are two very different processes, absorbing a tune through listening vs. repeating it over and over until you have it “memorized.” Which, if either, or those two is an LLM doing when it is chewing its way through a corpus of texts? When I prompt ChatGPT with “To be or not to be,” it responds with Hamlet’s complete soliloquy, word-for-word. When it does that is the process more like what I do when playing music by ear or like I do when memorizing music? Or is it something else?

That’s the kind of issue I had in mind when I undertook the investigations I report in this working paper. In the first piece – What must be the case that ChatGPT would have memorized “To be or not to be”? – I start out with Hamlet’s famous soliloquy, initially prompting ChatGPT with first line, but then prompting it with other fragments from the soliloquy. Then I prompt it with the phrase, “Johnstown flood, 1889,” and it responds with information about that flood, by not a specific text word-for-word. Many prompts are like that, many more than elicit a specific text word-for-word. What leads to that difference? I conclude with two topics I have reason to believe were included in the training corpus, but which ChatGPT seems to know nothing about. Why not?

In the next section (To be or not: Snippets from a soliloquy) I create various prompts for the soliloquy. I do the same in the third section (Entry points into the memory stream: Lincoln’s Gettysburg Address), but more systematically. Finally, I do a bit of speculating about what’s going on: Notes on ChatGPT’s “memory” for strings and for events. I begin by quoting a passage from F. C, Bartlett’s classic 1932 study, Remembering, and conclude that ChatGPT may have an associative memory along the lines suggested by holography, which engendered a great deal of speculation in the 1970s and, in this millennum, specifically for word meaning and order.

Wednesday, September 20, 2023

Notes on ChatGPT’s “memory” for strings and for events

[Updated on Sept. 21 and 26, 2023]

Here I take a look at the results reported in three previous posts and begin the job of making sense of them analytically. Here are the posts:

What must be the case that ChatGPT would have memorized “To be or not to be”? – Three kinds of conceptual objects for LLMs, New Savanna, September 3, 2023.

To be or not: Snippets from a soliloquy, New Savanna, September 12, 2023.

Entry points into the memory stream: Lincoln’s Gettysburg Address, New Savanna, September 13, 2023.

I set the stage with a passage from F. C. Bartlett’s 1932 classic, Remembering. Then I consider the three cases I laid out in that first post and then go on to look at the results reported in the next two. I conclude by suggesting that we look to the psychological literature on memory and recall to begin making analytic sense of these results. Of course, we also need more observations.

F.C. Bartlett, memory, and schemas

Back in the ancient days of 1932 F. C. Bartlett published a classic study of human recall, Remembering: A Study in Experimental and Social Psychology (1932). He performed a variety of experiments, a number involving the familiar game of having people tell a story from person to person to a chain and then comparing the initial story with the final one. He made the general conclusion that memory is not passive, like a tape-recorder or a camera, but rather is active, involving schemas (I believe he may have been the one to introduce that term to psychology), which shape our recall. A story that corresponds to an existing schema will be more faithfully transmitted than one that does not.

However, I’m not interested in those experiments. I’m interested in something he reports in a later chapter, “Social Psychology and the Manner of Recall,” pp. 264-266:

As everybody knows, the examination by Europeans of a native witness in a court of law, among a relatively primitive people, is often a matter of much difficulty. The commonest alleged reason is that the essential differences between the sophisticated and the unsophisticated modes of recall set a great strain on the patience of any European official. It is interesting to consider an actual record, very much abbreviated, of a Swazi trial at law. A native was being examined for the attempted murder of a woman, and the woman herself was called as a necessary witness. The case proceeded in this way:

The Magistrate: Now tell me how you got that knock on the head.

The Woman: Well, I got up that morning at daybreak and I did... (here followed a long list of things done, and of people met, and things said). Then we went to so and so’s kraal and we... (further lists here) and had some beer, and so and so said....

The Magistrate: Never mind about that. I don’t want to know anything except how you got the knock on the head.

The Woman: All right, all right. I am coming to that. I have not got there yet. And so I said to so and so... (there followed again a great deal of conversational and other detail). And then after that we went on to so and so’s kraal.

The Magistrate: You look here; if we go on like this we shall take all day. What about that knock on the head?

The Woman: Yes; all right, all right. But I have not got there yet. So we... (on and on for a very long time relating all the initial details of the day). And then we went on to so and so’s kraal.. .and there was a dispute ... and he knocked me on the head, and I died, and that is all I know.

Practically all white administrators in undeveloped regions agree that this sort of procedure is typical of the native witness in regard to many questions of daily behaviour. Forcibly to interrupt a chain of apparently irrelevant detail is fatal. Either it pushes the witness into a state of sulky silence, or disconcerts him to the extent that he can hardly tell his story at all. Indeed, not the African native alone, but a member of any slightly educated community is likely to tell in this way a story which he has to try to recall.

What’s going on here? Keep in mind that the issue is not word-for-word recall. Rather, it is the incidents being recalled, in whatever verbal form is convenient. Why can’t the witness simply begin talking about the incident in question? And when, when asked to get on with it, must the witness return to be beginning of the day?

It's as though the memory stream of a day’s events can only be entered at the beginning of the day, and not at arbitrary points within the day. I note that we are dealing with people who do not have clocks and watches they can use to mark events during the day. Of course it’s not enough to have a watch, you must also take note of it at various times during the day. That will give you various points of entry into the memory stream.

This sort of thing is also quite familiar to me as a musician. While I have learned to read music, and have done so often, I am an improvising (jazz) musician and am quite used to playing things “by ear.” If I am practicing a melody by ear, and get lost at some point, I may not be able to restart at the point where I broke off. Rather, like the witness testifying in court, I have to go back to the beginning – in this case, the beginning of the melody rather than the beginning of the day.

To be or not, and beyond

Early in September I asked the question: “Given that [the LLM underlying ChaGPT] has been trained to predict [only] the next word, what MUST have been the case in order to ChatGPT to return the whole soliloquy when given the opening six words?” It must have encountered that soliloquy many different times in its training corpus. That’s the only way that predicting that exact sequence, word after word, not result in training loss.

However, given Bartlett’s observations about human memory, ChatGPT’s ability to rattle off a whole sequence word for word does raise a question. Is it just passively stringing one word after another, or does it recognize internal structure? How can we figure out which is the case?

I want to set those questions aside for a moment, but I will return to the question. Though interesting, such specific sequences are relatively rare. It is much more common for the training corpus to have many texts about the same event or set of events, but not expressed in the exact same words. Thus I have ChatGPT the prompt, “Johnstown flood, 1889.” Note that I specified the year because Johnstown (PA) was subsequently flooded in 1937 and 1977. But it’s the 1889 flood that made the national news, prompting national concern.

ChatGPT responded in a way I thought reasonable. Since I had grown up in Johnstown and was familiar with the flood, I didn’t bother to check the Chatster’s reply against reliable sources. But, for all I know, the Chatster was giving me some specific text word-for-word, implying that there was some specific text about the flood that had appeared many times in the training corpus. While that didn’t seem likely, I had to check. Later that same day I opened a new session and gave ChatGPT the same prompt. Again, it gave me a reasonable reply, but one that was expressed differently from the earlier one. This reply gave the sequence of events in seven numbered paragraphs. The earlier reply did not have a sequence of numbered paragraphs.

So we’ve got two cases so far: 1) a specific sequence of words that is repeated when prompted for, and 2) and flexible recall of an event using different word sequences in different sessions. There is a third case to consider: 3) and event that is in the training corpus, but in so very few times, perhaps only once, that it doesn’t register in ChatGPT’s model as a specific event. The text serves as evidence about word usage, but otherwise has no effect on the model.

Without access to the training corpus, how do you identify such things? You can’t. But you can make a plausible. I’d attended a Dizzy Gillespie concert back in the mid-1980s which I’d written about in two places which could have been in the training corpus. I prompted ChatGPT with that concert, naming the venue and city where the concert took place in addition to the artist (Diz). Apparently, it had no record of it.

Let me offer you a somewhat different example of this last case. I’m currently interested in a mathematician named Miriam Lipschutz Yevick. She published a paper back in 1975 (Holographic or fourier logic), which I think is interesting and important, but which has been forgotten. The paper is available on the web, and I have blogged about it. A few other papers are also available, as well as an obituary, all before ChatGPT’s cut-off point. I’ve asked Chatster about Yevick in several different sessions but it knows nothing about her. 

Let’s think about this a bit. GPT-3.5, the large language model underlying ChatGPT, may have been trained on (almost, a big chunk of) the entire internet, but its model does not incorporate everything that it has been trained on. It is abstracting over those texts, not memorizing them in any ordinary sense of the word. When a particular text occurs word-for-word many times and in various contexts, GPT-3.5 will learn it word-for-word; think of that as, in effect, an abstraction over those many contexts. When a particular topic, that is, a particular congeries of terms, occurs many times and in various contexts, GPT-3.5 abstracts over than congeries and meshes them together so they are mutually available. If neither of these things occurs to something that appears in a text, then that something just dissolves into the net.

Thus we’ve got three cases: 1) word-for-word recall of a text, 2) flexible recall of a specific topic, and 3) no recall of a topic that was in its training corpus. I want to return to the first case, where ChatGPT is generating a fixed text, and see what, if anything, we can learn about how it does it.

Sunday, August 27, 2023

Xanadu, GPT, and Beyond: An adventure of the mind

I've posted a new working paper. Title above; links, abstract, table of contents, and introductory material below.

Download at: 

Academia: https://www.academia.edu/106001453/Xanadu_GPT_and_Beyond_An_adventure_of_the_mind
SSRN: https://ssrn.com/abstract=4553351
ResearchGate: https://www.researchgate.net/publication/373433939_Xanadu_GPT_and_Beyond_An_adventure_of_the_mind

Abstract: This article recounts an intellectual journey that began in curiosity about the structure of Coleridge’s “Kubla Khan” in the late 1960s and has led to an interest in large language models at the present time. A close analysis of the poem revealed its two parts each to have a nested structure (think of a matryoshka doll) that suggested the operation of an underlying computational process (nested loops). That led to the study of computational linguistics (semantic networks), followed by neuroscience (Karl Pribram’s neural holography), and cultural evolution. In the 2010s I began following work digital humans had been doing with machine learning. When GPT-3 was released in 2020 I was ready, though it took me awhile to establish a link, however tentative, between that conceptual universe and that of “Kubla Khan.”

Encountering Coleridge’s “Kubla Khan” 3
Romantic states of consciousness 3
Matryoshka dolls and the escape from Xanadu 5
Semantic networks and a Shakespeare sonnet 8
Karl Pribram, neural holography, and the brain 10
The wandering years 11
Through GPT to the future 13
Mind and world in text 15
The Text of “Kubla Khan,” including Coleridge’s prefatory note 17
A note about the cover image 19

Encountering Coleridge’s “Kubla Khan”

I became hooked on Coleridge’s “Kubla Khan” in the Spring of 1969, my last semester as an undergraduate at Johns Hopkins. Three years later “Kubla Khan” had become the standard against which I measured my understanding of the human mind. That is why I am telling a story about how my interest in the mind has evolved through “Kubla Khan” to include, most recently, ChatGPT. Strange as it may seem, that poem is the vehicle through which I am coming to terms with this new technology and arriving at a sense of its potential.

There is a sense in which the story of that great poem can be traced back to the 11th century invasion of Britain by the Norman French, for that culture-crossing is what gave rise to the English language. A century or so later that story encountered a tale born of an encounter between an Italian merchant, Marco Polo, and a Mongolian warlord, Kubla Khan, which, when enlivened by the East India Company’s trade in opium, set fire to the mind of Samuel Taylor Coleridge in the late 18th and early 19th centuries. We need not trace that trajectory in any detail. I mention it only to give a sense of the scope of this 36-line poem, which is one of the best-known poems in the English language, and is perhaps unique in the annals of Western literature. It has made its mark on popular culture, from Orson Welles’s Citizen Kane, where it names Kane’s estate, Xanadu, thereby establishing the matrix for the whole film, to a hit song and film by Olivia Newton-John, Xanadu, subsequently made into a Broadway musical. It even provided that most vulgar of real-estate barons, Donald Trump, with the name for the nightclub, Xanadu, in his now defunct Atlantic City casino.

Romantic states of consciousness

I may well have read the poem prior to taking Romantic Literature with Professor Earl Wasserman in 1968-1969. But I have no memory of that. Though we didn’t study Coleridge until the second semester, it is probably best if I start my story with the first semester.

The course started with Keats. I decided to write my paper about a minor poem, “To–[Fanny Brawne],” and had delayed writing until the night before it was due. I was tired and my mind snapped. All of a sudden, I was typing a passage from one of Keats’ letters to Fanny, but I experienced the act of typing as though the words were my own. When I finished that passage, my mind was astir and found its way to the second stanza of “Ode on a Grecian Urn” – you know “Heard melodies are sweet, but those unheard Are sweeter...” I read those words as though they were my own.

I finished the paper, turned it in, got a grade, and...

I had a problem: What was that!? I didn’t know. But this was the 1960s and altered states of consciousness were all the rage, drug induced, but also meditation, and now it seems, the influence of late-night poetry on a tired mind.

Next up: Percy Bysshe Shelley, he who had declared poets to be “the unacknowledged Legislators of the world.” Again, I delayed writing my paper until the last minute. I was tired. The damned paper wrote itself, through me. But I didn’t experience anything of Shelley’s as though I had written it. It was different from the experience I had writing about Keats. The words just lined themselves up, one after the other and flowed down my arms, through my fingers, from the typewriter and onto the page. It was easy. No sweat. It was a good paper too.

Wordsworth was up in the Spring semester. That was the best paper I’d written as an undergraduate. Wasserman remarked that it “was a mature contemplation of the poem” – though I forget just what poem it was. There were no mental hijinks. I wrote it with the standard task-assemblage of a sentence or three here, a paragraph there, pace the room a bit, make a note or three, look up something, back to the typewriter, rinse, repeat, and so forth....it’s done.

And so it went with my “Kubla Khan” paper. The poem itself presents a number of problems. The first is: What is it about? There is no narrative. It has often been dismissed as word music. Word music it is, but that is no ground for dismissal.

Then we have Coleridge’s preface. He said the poem was incomplete. He had become lost in an opium reverie when two or three hundred lines came to him – “all the images rose up ... as things, with a parallel production of the correspondent expressions, without any sensation or consciousness of effort” – which was dashed when he was interrupted by a man from Porlock. When the Porlockian had gone, so had those two or three hundred lines. All that was left were the 54 lines of this, one of the most extraordinary poems in the world. In fact, nothing is obviously missing. If it weren’t for that preface, no one would even suspect that the poem was incomplete.

Critics have had various ways of dealing with the disparity between the poem itself and Coleridge’s claim. I invented another solution to the problem. It is easy and natural to interpret the second part of the poem as asserting that the poem is incomplete (I’ve appended a complete text to the end of this essay). The speaker says “Could I revive within me” (l. 42), clearly implying that he can’t, but if he could he would “build that dome in air” (p. 46). The dome is assumed to be Kubla’s pleasure-dome from the first part and is here being used as figure for the poem itself. That’s a perfectly respectable reading of those lines.

I pushed it a step further. I asserted that the poem paradoxically completes itself by asserting that it is incomplete. That kind of reading has it all. The wealthy English of Coleridge’s time were fond of placing newly built but incomplete or dilapidated structures in their gardens – “follies” they were called. An exquisitely dilapidated poem fit right in with that aesthetic. Moreover, such paradoxical readings fit right in with the rising tide of structuralist, post-structuralist and deconstructionist readings in American literary criticism. Despite all that that, Wasserman, who was more traditional in his conceptual leanings, Wasserman loved it.

A note about the image

The portrait of “Kubla Khan” was made by Araniko, a Nepalese artist, shortly after Kubla’s death in 1294. The image is from Wikimedia Commons and is in the public domain.

Araniko: https://en.wikipedia.org/wiki/Araniko.
Image: https://commons.wikimedia.org/w/index.php?curid=4126240.

I have overlaid it with an image I made in MacPaint on a Classic Macintosh in 1985.