Showing posts with label Gärdenfors. Show all posts
Showing posts with label Gärdenfors. Show all posts

Monday, December 23, 2024

LLM as Collaborator, Part 1: Claude the Graduate Student

I started using ChatGPT on December 1, 2022 and have used it quite extensively ever since. I’ve spent some of my time just poking around, somewhat more time looking things up, and most of my time systematically investigating its performance. That resulted in a number of working papers, the most interesting of which is about stories: ChatGPT tells stories, and a note about reverse engineering: A Working Paper, Version 3.

I started working with Claude 3.5 Sonnet on November 18, 2024. I’ve used it in those three capacities, though obviously not as much as I’ve used ChatGPT. In particular, I’ve used it for background information on melancholy in various aspects. I’ve also done something I’d never done with ChatGPT, asked it to describe photographs. I’m doing this to see how well it does.

Then on November 24, 2024, I began using it in a somewhat more interesting new capacity, though I’m not sure what to call it. The phrase “thought partner” comes to mind, though it seems too much like “thought leader,” which I don’t like. I’m using it as a sounding board. Better yet, its a collaborator playing the role of sounding board. It’s not an equal collaborator in the intellectual dialog; academic norms would not require me to offer it co-authorship of papers. But those norms might well require an explicit acknowledgement, not to alert the reader that one of those new-fangled LLM things has been involved in the thinking, but simply acknowledging the help it has given me.

As for just what kind of help that is, the best way is to look at some examples. I’ve already published two of these dialogues on New Savanna: Computer chess, super-intelligence, A Connecticut Yankee, cargo cults, and Claude 3.5 Sonnet speculates about future developments in AI and beyond. The rest of this post contains the first interaction where Claude began acting in intellectual partnership.

When I started the interaction I was simply interested in seeing what it knew about my teacher and mentor, the late David Hays. I wasn’t looking for new information. I just wanted to see what it had picked up. That began to change when it asked me about the work of Quillian and of Schank and Abelson. Then I told it about Hays’s interest in the work of William Powers. Now we’re getting serious. This guy knows something. And then WHAM! Claude responded intelligently when I mentioned the work of Peter Gärdenfors. The interaction continued in that mode to the end.

Utterly remarkable. AGI? What the hell is that? But no. And I fear that thinking in those terms is likely to blind me to whatever it is that Claude is doing. I realize that the concept of AGI is written into the contractual relationship between OpenAI and Microsoft, but I’m not playing that game. Will they end up litigating it? Who knows?

I’m planning two more posts in this series. The next will be about the implications of paragraph length in Heart of Darkness. Then I’ll post a long interaction that takes Miriam Yevick’s 1975 paper, “Holographic or fourier Logic,” as its starting point.

In the following dialog my comments are bold while Claude’s are plain. I’ve interpolated a few comments highlighted in blue.

* * * * *

Are you familiar with the work of a computational linguist named David Hays? He worked on machine translation at RAND in the 1950s and 1960s?

Yes, I am somewhat familiar with David Hays' work, though I should note that since this involves fairly obscure historical details from computational linguistics, I want to be careful about potential inaccuracies in my knowledge. From what I understand, David Hays was one of the pioneers of machine translation and computational linguistics during the early years of the field. He worked at RAND Corporation in the 1950s and 1960s where he was involved in some of the earliest machine translation projects.

He was particularly interested in developing dependency grammar approaches for machine translation, which was quite innovative for that time period. Dependency grammar focuses on the relationships between words rather than constituent structure, which was the dominant paradigm in American linguistics at the time.

I believe he also wrote an influential book called "Computational Linguistics: Programming Linguistics in Automatic Language Translation" in the early 1960s, though I'd encourage you to verify this citation since I may be misremembering details.

Claude got the name wrong. It’s Introduction to Computational Linguistics, American Elsevier, New York, 1967.

Would you be interested in hearing more about any particular aspect of his work or that early period of machine translation research? I find the history of early computational linguistics fascinating, though again I want to be upfront that with someone as historically specific as Hays, some of my knowledge may be incomplete or imperfect.

Yes, that’s him. I studied with him at SUNY Buffalo in the 1970s. We were working on computational semantics and using a cognitive networks formalism.

Oh, that's fascinating that you studied directly with David Hays! I would be very interested to hear more about your experiences working with him and the cognitive networks approach you were developing at SUNY Buffalo. That was an important period in the development of semantic representation systems.

Could you tell me more about the cognitive networks formalism you were using? I'm particularly curious about how it compared to other semantic network approaches being developed around that time, like Quillian's work or Schank and Abelson's conceptual dependency theory.

It was more developed than Quillian’s work. As I recall, he never really developed is initial ideas. As for Schank and Abelson, they were imagining a self-sufficient cognitive system and were very much interested in identifying the primitive elements necessary for a cognitive network. By the time I began working with him, Hays had realized that cognition had to be grounded in something else – later on Steven Harnad would talk about the “symbol grounding” problem. When I worked with him, Hays was interested in grounding the cognitive system in the control-theory (cybernetic) account of perception and action developed by William T. Powers (perceptual control theory, or PCT).

Something special seems to be happening.

This is really interesting historical context - it suggests Hays was grappling with fundamental issues in cognitive science that are still relevant today. The shift from viewing semantic networks as self-contained symbolic systems to recognizing the need for grounding in perception and action seems quite prescient.

I'm curious about how Hays was trying to integrate Powers' PCT with semantic networks. As I understand it, PCT proposes that behavior emerges from organisms controlling their perceptions to match internal reference signals, rather than through direct control of outputs. Was Hays trying to use this as a way to ground symbolic meaning in perceptual-motor control loops?

This also seems to anticipate some later developments in embodied cognition and grounded cognition theories. The timing is interesting too - Powers published "Behavior: The Control of Perception" in 1973, so it sounds like Hays was engaging with these ideas pretty quickly after they were published.

Yes, that's what Hays was doing. His idea was that particular servo units in a Powers-like control system would be linked to cognitive units through what Hays called "parameters of perception." Hays identified a half dozen or so such parameters. Each cognitive node would be characterized by the values it took on those parameters. Relationships between nodes were thus derived from differences in parameter values for the nodes connected by an arc (or edge).

Tuesday, September 5, 2023

World, mind, and learnability: A note on the metaphysical structure of the cosmos [& LLMs]

I originally posted this three years ago, on August 15.  I have added an important new section at the end, Paths of the mind (virtual reading), and so I am bumping this to the top of the queue.
There is no a priori reason to believe that world has to be learnable. But if it were not, then we wouldn’t exist, nor would (most?) animals. The existing world, thus, is learnable. The human sensorium and motor system are necessarily adapted to that learnable structure, whatever it is.

I am, at least provisionally, calling that learnable structure the metaphysical structure of the world. Moreover, since humans did not arise de novo that metaphysical structure must necessarily extend through the animal kingdom and, who knows, plants as well.

“How”, you might ask, “does this metaphysical structure of the world differ from the world’s physical structure?” I will say, again provisionally, for I am just now making this up, that it is a matter of intension rather than extension. Extensionally the physical and the metaphysical are one and the same. But intensionally, they are different. We think about them in different terms. We ask different things of them. They have different conceptual affordances. The physical world is meaningless; it is simply there. It is in the metaphysical world that we seek meaning. [See my post, There is a fold in the fabric of reality. (Traditional) literary criticism is written on one side of it. I went around the bend years ago.]

A little dialog

Does this make sense, philosophically? How would I know?

I get it, you’re just making this up.

Right.

Hmmmm… How does this relate to that object-oriented ontology stuff you were so interested in a couple of years ago?

Interesting question. Why don’t you think about it and get back to me.

I mean, that metaphysical structure you’re talking about, it seems almost like a complex multidimensional tissue binding the world together. It has a whiff of a Latourian actor-network about it.

Hmmm… Set that aside for awhile. I want to go somewhere else.

Still on GPT-3, eh?

You got it.[1]
 
A little diagram: World, Text, and Mind

Text reflects this learnable, this metaphysical, structure, albeit at some remove:

Learning engines are learning the structure inherent in the text. But that learnable structure is not explicit in the language model created by the learning engine.

There are two things in play: 1) the fact that the text is learnable, and 2) that it is learnable by a statistical process. How are these two related?

If we already had an explicit ‘old school’ propositional model in computable form, then we wouldn’t need statistical learning at all. We could just run the propositional model over the corpus and encode the result. But why do even that? If we can read the corpus with the propositional model, in a simulation of human reading, then there’s no need to encode it at all. Just read whatever aspect of the corpus is needed at the time.

So, statistical learning is a substitute for the lack of a usable propositional model. The statistical model does work, but at the expense of explicitness.

But why does the statistical model work at all? That’s the question.

It’s not enough to say, because the world itself is learnable. That’s true for the propositional model as well. Both work because the world is learnable.

Language model as associative memory

BUT: Humans don’t learn the world with a statistical model. We learn it through a propositional engine floating over an analogue or quasi-analogue engine with statistical properties. And it is the propositional engine that allows us to produce language. A corpus is a product of the action of propositional engine, not a statistical model, acting on the world.

Description is one basic such action; narration is another. Analysis and explanation are perhaps more sophisticated and depend on (logically) prior description and narration. Note that this process of rendering into language is inherently and necessarily a temporal one. The order in which signifiers are placed into the speech stream depends in some way, not necessarily obvious, on the relations among the correlative signifieds in semantic or cognitive space. Distances between signifiers in the speech stream reflect distances between correlative signifieds in semantic space. We thus have systematic relationships between positions and distances of signifiers in the speech stream, on the one hand, and positions and distances of signifieds in semantic space. It is those systematic relationships that allow statistical analysis of the speech stream to reconstruct semantic space.

Note that time is not extrinsic to this process. Time is intrinsic and constitutive of computation. Speaking involves computation, as does the statistical analysis of the speech stream.

The propositional engine learns the world via Gärdenfors’ dimensions [2], and whatever else, Powers’ stack for example [3]. Those dimensions are implicit in the resulting propositional model and so become projected onto the speech stream via syntax, pragmatics, and discourse structure. The language engine is then able to extract (a simulacrum of) those dimensions through statistical learning. Those dimensions are expressed in the parameter weights of the model. THAT’s what makes the knowledge so ‘frozen’. One has to cue it with actual speech.

The whole language model thus functions as associative memory [4]. You present it with an input cue, and it then associates from that cue with each emitted string ‘feeding back’ into the memory bank via associative memory.
 
Paths of the mind (virtual reading)
 
Now, imagine a word embedding model constructed over some suitable corpus of texts. Given that texts reflect the interaction of the mind and the world, the location of individual words in that model necessarily reflects that interaction. That structure is what I have been calling the metaphysical structure of the cosmos.

Consider some text. It consists of word after word after word. That sequence reflects the actions of the mind that wrote the text, and only the mind. For all practical purposes, the cosmos remains unchanging during the writing of that text and the mind has withdrawn from active interaction with the world, except insofar as the text is a set of symbols that it is placing on a piece of paper, moment after moment, word by word. Now let us trace the path some texts takes through a word embedding. Call this a virtual reading. The word embedding consists of tens of 1000s of dimensions, so that path will be a complicated one. That path must necessarily be a product of the mind (and only the mind?). That path is the mind in action.

Further, consider that the mind is what the brain does. The brain consists of 86 billion neurons, each having on the order of 10,000 connections with other neurons. The dimensionality of the space required to represent the brain is thus huge. At any given moment we can represent the state of the brain as a point in that space. From one moment to the next, the brain traces a path in that space, what a dynamicist (such as Walter Freeman) would call a trajectory. When someone is writing a text, that text reflects the operations of their brain. Therefore the path of a text through a word embedding necessarily mirrors the trajectory taken by the author's brain while writing the text. Notice, however, the reduction in dimensionality. The brain's state space is of vastly higher dimensionality than the word embedding space.

Finally, consider the operations of a transformer as it generates a text. Each time it generates a token it takes the entire model into account, all those 100s of billions of parameters (are we up to trillions yet?). Compare that to what a brain did when generating that same text. Is it reasonable to consider a single token as an image of, a reflection of, short trajectory segment in the state space of the brain that generated the text. Walter Freeman thought of the brain as moving through states of global coherence at the rate of 7-10 Hz, like frames of a film [5]. In what way is the movement of a transformer from one token to the next comparable to the movement of the brain from one frame to the next?[6]


References
 
[1] This post is an exploration of ideas raised in the course of thinking about GPT-3. See William Benzon, GPT-3: Waterloo or Rubicon? Here be Dragons, Working Paper, August 5, 2020, 32 pp., Academia: https://www.academia.edu/s/9c587aeb25; SSRN: https://ssrn.com/abstract=3667608
ResearchGate: https://www.researchgate.net/publication/343444766_GPT-3_Waterloo_or_Rubicon_Here_be_Dragons.

[2] Peter Gärdenfors, Conceptual Spaces: The Geometry of Thought, MIT Press, 2000; The Geometry of Meaning: Semantics Based on Conceptual Spaces, MIT Press, 2014.

[3] William Powers, Behavior: The Control of Perception (Aldine) 1973. A decade later David Hays integrated Powers’ model into his cognitive network model, David G. Hays, Cognitive Structures, HRAF Press, 1981.

[4] The idea that the brain implements associative memory in a holographic fashion was championed by Karl Pribram in the 1970s and 1980s. David Hays and I drew on that work in an article on metaphor, William Benzon and David Hays, Metaphor, Recognition, and Neural Process, The American Journal of Semiotics , Vol. 5, No. 1 (1987), 59-80, https://www.academia.edu/238608/Metaphor_Recognition_and_Neural_Process.
 
[5] Freeman, W. J. (1999a). Consciousness, Intentionality and Causality. Reclaiming Cognition. R. Núñez and W. J. Freeman. Thoverton, Imprint Academic, 143-172.
 
[6] I first argue this point in William Benzon, The idea that ChatGPT is simply “predicting” the next word is, at best, misleading, New Savanna, Feb. 19, 2023.

Thursday, March 9, 2023

Two more thoughts on ChatGPT: Conceptual spaces and system time-steps

The mere fact that I’ve posted a substantial article, How ChatGPT tells stories, does not at all imply that I’ve stopped thinking about those issues. Not at all. The process of writing at then distributing an article is simple a device for bringing my thinking to a certain level of maturity. But the thinking continues.

Here are two further thoughts. The first is about conceptual spaces and might in fact find its way into a later version to some article, assuming I decided to produce one. The second is considerably more speculative and requires more work and, in any event, would go in a different kind of paper.

From a note to Gärdenfors on conceptual spaces

The procedure I have been using is derived from the analytical method Claude Lévi-Strauss employed in his magnum opus, Mythologiques. He started with one myth, analyzed it, and then introduced another one, very much like the first. But not quite. They are systematically different. He characterized the difference by a transformation – a term he took from algebraic group theory. He worked his way through hundreds of myths in this manner, each one derived from another by a transformation.

Here is what I have been doing: I give ChatGPT a prompt consisting of two things: 1) an existing story and 2) instructions to produce another story like it except for one change, which I specify. That change is, in effect, a way of triggering or specifying those “transformations” that Lévi-Strauss wrote about. What interests me are the ensemble of things that change along with the change I have specified. It some cases it’s quite striking, as you can see from the table in the article.

Though I mention your work in the paper, I don’t explore it. Now that the paper is (more or less) done, I’ve been thinking, and it seems to me that conceptual spaces provides a ‘natural’ way to account for the results of these experiments.

Most the times I directed ChatGPT to change the protagonist. But sometimes I focused on the antagonist. The protagonist I use in the source stories is princess Aurora. In one case the new protagonist was Prince Harry. Except for gender, these are similar people, requiring minimal changes in the new story. In another case, however, the protagonist was William the Lazy. Since ChatGPT is operating on the assumption that characters have an intrinsic ‘nature’ and their actions must follow from that nature. ChatGPT had to come up with a way that a Lazy man could defeat a dragon. That required more extensive changes in the new story. William the Lazy had to summon his knights and get them to do the work. Still more changes were required when I had to transform Princess Aurora into a giant chocolate milkshake. ChatGPT had no trouble doing it and the resulting story was quite different from the original. The whole mise-en-scène had changed.

So, let's create a conceptual space in which we place the protagonist of the original story and the protagonist of the new story. They will be at different positions in that space reflecting the fact that they have different values on the dimensions that define the space. Now, let’s take the difference between those positions and use that difference as an off-set that we apply to the whole story, thus shifting its trajectory is semantic space.

It’s probably not quite that simple. In the case of William the Lazy, I doubt that the shift if semantic space would automatically produce the act where he summons his knights. ChatGPT had to do a little work to come up with that. But on the whole it seems to me that this is the way to go.

In note that, in particular, this is quite different from what you would have to do if you used a story grammar based on symbolic systems. Though I never worked with story grammars, I was trained in symbolic systems (and analyzed a Shakespeare sonnet using a semantic network) and read the literature story grammars. I dare say none of them could have done that task, much less done it so easily and naturally. It would have required extensive machinery.

It seems to me that metaphor and analogy could be handled in a similar fashion.

System time steps in brains and GPTs

This is a comment I posted to the semiotic phyics post at LessWrong:

Have you thought of exploring the existing literature on the complex dynamics of nervous systems. It’s huge, but it does use the math you guys are borrowing from physics.

I’m thinking in particular of the work of the late Walter Freeman, who is a pioneer in the field. Toward the end of his career he began developing a concept of “cinematic consciousness.” As you know the movement in motion pictures is an illusion created by the fact the individual frames of the image are projected on the screen more rapidly than the mind can resolve them. So, while the frames are in fact still, they change so rapidly that we see motion.

First I’ll give you some quotes from Freeman’s article to give you a feel for his thinking (alas, you’ll have to read the article to see how those things connect up), then I’ll explain what that has to do with LLMs. The numbers are from Freeman’s article.

[20] EEG evidence shows that the process in the various parts occurs in discontinuous steps (Figure 2), like frames in a motion picture (Freeman, 1975; Barrie, Freeman and Lenhart, 1996).

[23] Everything that a human or an animal knows comes from the circular causality of action, preafference, perception, and up-date. It is done by successive frames of self-organized activity patterns in the sensory and limbic cortices. [...]

[35] EEG measurements show that multiple patterns self-organize independently in overlapping time frames in the several sensory and limbic cortices, coexisting with stimulus-driven activity in different areas of the neocortex, which structurally is an undivided sheet of neuropil in each hemisphere receiving the projections of sensory pathways in separated areas. [...]

[86] Science provides knowledge of relations among objects in the world, whereas technology provides tools for intervention into the relations by humans with intent to control the objects. The acausal science of understanding the self distinctively differs from the causal technology of self-control. "Circular causality" in self-organizing systems is a concept that is useful to describe interactions between microscopic neurons in assemblies and the macroscopic emergent state variable that organizes them. In this review intentional action is ascribed to the activities of the subsystems. Awareness (fleeting frames) and consciousness (continual operator) are ascribed to a hemisphere-wide order parameter constituting a global brain state. Linear causal inference is appropriate and essential for planning and interpreting human actions and personal relations, but it can be misleading when it is applied to microscopic- microscopic relations in brains.

Notice that Freeman refers to “a hemisphere-wide order parameter constituting a global brain state.” The cerebral cortex consists of 16B neurons, each with roughly 10K connections. Further, all areas of the cortex have connections with subcortical regions. That’s an awful-lot of neurons communicating in parallel in a single time step. As I recall from another article, these frames occur at a rate of 6-7 Hz.

The nervous system operates in parallel. I believe it is known that the brain exhibits a small world topology, so all neurons are within a relatively small number links from one another. Though at any moment some neurons will be more active than others, they are all active – the only inactive neuron is a dead neuron. Similarly, ANNs exhibit a high degree of parallelism. LLMs are parallel virtual machines being simulated by so-called von Neumann machines. The use of multiple cores gives a small degree of parallelism, but that’s quite small in relation to the overall number of parameters the system has.

I propose that the process of generating a single token in an LLM is comparable to a single “frame” of consciousness in Freeman’s model. All the parameters in the system are visited during a single time-step for the system. In the case of ChatGPT I believe that’s 175B parameters.

Thus the assertion that ChatGPT generates one token at a time, based on the previous string, while true, is terribly reductive and thus misleading. The appearance of a token is in fact more or less a side-effect of evolving a trajectory from the initial prompt.

ChatGPT explores the semantic differential

3.9.23 – 6:13 AM – Feb 13 Version

If you don’t already know what the semantic differential is, look it up in Wikipedia. But you should get the idea from the prompts I give it. It’s a standard technique is psychology going back to the middle of the 20th century.

I more or less expected it to get the idea, though I can do without all those annoying explanations for what it’s doing. Note that at two points it froze, so I had to re-initiate the session.

ChatGPT is ‘fluid’ in a way that standard symbolic language models from the “classic” period are not. That’s fine. It also lends credence to my current working hypothesis that its semantics seems based on the kind of “geometric” semantics that Peter Gärdenfors has been exploring. Perhaps if investigators learn about it and start exploring it in LLMs they’ll seem less mysterious and more coherent and lawful. 

Notice that toward the end ChatGPT seems to get the idea and comes up with its own continuum. In-context learning?

* * * * *

Imagine that the word “cloud” is at one end of a continuum with a value of 1. “Fire” is at the other end of that continuum with a value of 7. When you place “toy” on that continuum, what value does it have?

The continuum you described seems to represent a spectrum of natural phenomena, with "cloud" representing something ethereal and "fire" representing something intense and powerful. "Toy" is not a natural phenomenon and does not fit neatly onto this continuum. Therefore, it is not possible to assign a meaningful value to "toy" on this particular continuum.

Forget about natural or making sense. Just take a guess. Where would it go?

As requested, if we were to take an arbitrary guess and place "toy" on the continuum between "cloud" and "fire", I would place it somewhere around the midpoint, perhaps at a value of 4. This is because a toy is typically a man-made object that does not possess the same natural qualities or power as fire, but it also has a physical presence and substance that sets it apart from something as intangible as a cloud. However, it's important to note that this is just a random guess and has no real scientific or logical basis.

Let’s take another random guess. Put “desire” on that continuum. Just a guess.

Placing "desire" on the same continuum as "cloud" and "fire" is also a bit challenging because "desire" is an abstract concept and doesn't necessarily have a tangible physical presence like a cloud, nor does it possess the same raw power and energy as fire. However, if I had to make another arbitrary guess, I might place "desire" closer to "fire" than "cloud", perhaps at a value of 6. This is because "desire" can be a strong and intense feeling that can drive a person's actions and motivations, similar to the way fire can have a powerful and transformative effect on the world around it.

Now let’s stick with the continuum. Where do “thunder” and “volcano” go?

On the continuum between "cloud" and "fire", "thunder" and "volcano" could be placed as follows:

  • Thunder: Thunder is the loud sound that accompanies a lightning strike during a thunderstorm. While it is not a physical object like a cloud, it is an effect of the electrical discharge that occurs within a thunderstorm. Therefore, I would place "thunder" closer to "fire" than "cloud", perhaps at a value of 5.

  • Volcano: A volcano is a natural phenomenon that occurs when magma, ash, and gas escape from beneath the Earth's surface. It is a powerful and explosive force that can cause significant damage and destruction, similar to fire. Therefore, I would place "volcano" closer to "fire" than "cloud", perhaps at a value of 6 or 7.
  • Let’s try a different continuum. “Good” is 1 and “evil” is seven. Where do “trustworthy,” “duplicitous,” “charity,” “robbery,” “tithing,” and “fraud” go on that continuum?

    Imagine that the word “good” is at one end of a continuum with a value of 1. “evil” is at the other end of that continuum with a value of 7. When you place “trustworthy” on that continuum, what value does it have?

    If we consider a continuum with "good" at one end and "evil" at the other end, and we place "trustworthy" on that continuum, it is likely that most people would place it closer to the "good" end than the "evil" end.

    However, the exact value on the continuum would depend on various factors such as individual opinions, cultural context, and personal experiences.

    If we assume a typical scale of 1 to 7, with 1 being "good" and 7 being "evil", we might place "trustworthy" around 2 or 3 on the continuum, as it implies reliability, honesty, and dependability, which are generally positive qualities.

    What about “duplicitous?”

    If we consider the same continuum with "good" at one end and "evil" at the other end, and we place "duplicitous" on that continuum, it is likely that most people would place it closer to the "evil" end than the "good" end.

    "Duplicitous" implies deceitfulness and a lack of honesty, which are generally considered negative qualities. Therefore, on a scale of 1 to 7, with 1 being "good" and 7 being "evil," "duplicitous" might be placed around 6 or 7 on the continuum.

    What about “charity,” “robbery,” “tithing,” and “fraud”?

    On the continuum with "good" at one end and "evil" at the other end, here's how each of these words might be placed:

    • "Charity" typically implies generosity, kindness, and a desire to help others, so it would likely be placed closer to the "good" end of the continuum. On a scale of 1 to 7, "charity" might be placed around 1 or 2.
    • "Robbery" implies theft and a disregard for the property and well-being of others, so it would likely be placed closer to the "evil" end of the continuum. On a scale of 1 to 7, "robbery" might be placed around 6 or 7.
    • "Tithing" typically refers to the practice of giving a portion of one's income or resources to a religious organization or cause, so its placement on the continuum might depend on one's views about religion and the particular circumstances of the tithing. However, generally, it is seen as a positive action that reflects a desire to support one's community and values, so it would likely be placed closer to the "good" end of the continuum. On a scale of 1 to 7, "tithing" might be placed around 2 or 3.
    • "Fraud" typically implies deception or trickery for personal gain or advantage, so it would likely be placed closer to the "evil" end of the continuum. On a scale of 1 to 7, "fraud" might be placed around 6 or 7.

    OK. Now let’s consider a continuum that doesn’t make any sense, so you’re just going to have to take your best guess. “Stove” is 1 and “eagle” is 7. Where does “bowling ball” go?

    Tuesday, May 11, 2021

    A mathematical theory of semantic development in deep neural networks

    Andrew M. Saxea, James L. McClelland, and Surya Gangulic, A mathematical theory of semantic development in deep neural networks. PNAS, June 4, 2019, Vol. 116, No. 23, 11537-11546, www.pnas.org/cgi/doi/10.1073/pnas.1820226116

    Abstract: An extensive body of empirical research has revealed remarkable regularities in the acquisition, organization, deployment, and neural representation of human semantic knowledge, thereby raising a fundamental conceptual question: What are the theoretical principles governing the ability of neural networks to acquire, organize, and deploy abstract knowledge by integrating across many individual experiences? We address this question by mathematically analyzing the nonlinear dynamics of learning in deep linear networks. We find exact solutions to this learning dynamics that yield a conceptual explanation for the prevalence of many disparate phenomena in semantic cognition, including the hierarchical differentiation of concepts through rapid developmental transitions, the ubiquity of semantic illusions between such transitions, the emergence of item typicality and category coherence as factors controlling the speed of semantic processing, changing patterns of inductive projection over development, and the conservation of semantic similarity in neural representations across species. Thus, surprisingly, our simple neural model qualitatively recapitulates many diverse regularities underlying semantic development, while providing analytic insight into how the statistical structure of an environment can interact with nonlinear deep-learning dynamics to give rise to these regularities.

    Significance: Over the course of development, humans learn myriad facts about items in the world, and naturally group these items into useful categories and structures. This semantic knowledge is essential for diverse behaviors and inferences in adulthood. How is this richly structured semantic knowledge acquired, organized, deployed, and represented by neuronal networks in the brain? We address this question by studying how the nonlinear learning dynamics of deep linear networks acquires information about complex environmental structures. Our results show that this deep learning dynamics can self-organize emergent hidden representations in a manner that recapitulates many empirical phenomena in human semantic development. Such deep networks thus provide a mathematically tractable window into the development of internal neural representations through experience.

    * * * * *

    My quick take: I'm thinking that mathematical work like this will help close the gap between artificial neural models and investigation of real brains. I sense (possible) connections between this work and that in the tweet below, Graph Neural Networks, and between both of those and the work of Peter Gärdenfors on mental spaces (which I discuss in, e.g. World, mind, and learnability: A note on the metaphysical structure of the cosmos).

    Wednesday, August 5, 2020

    GPT-3: Waterloo or Rubicon? Here be Dragons


    I've published a new working paper. Title above, download links, abstract, table of contents, and introduction below.

    Download at:

    GPT-3 is a significant achievement.

    But I fear the community that has created it may, like other communities have done before – machine translation in the mid-1960s, symbolic computing in the mid-1980s, triumphantly walk over the edge of a cliff and find itself standing proudly in mid-air.

    This is not necessary and certainly not inevitable.

    A great deal has been written about GPTs and transformers more generally, both in the technical literature and in commentary of various levels of sophistication. I have read only a small portion of this. But nothing I have read indicates any interest in the nature of language or mind. That seems relegated to the GPT engine itself. And yet the product of that engine, a language model, is opaque. I believe that, if we are to move to a level of accomplishment beyond what has been exhibited to date, we must understand what that engine is doing so that we may gain control over it. We must think about the nature of language and of the mind.

    That is what this working paper sets out to achieve, a beginning point, and only that. By attending to ideas by Adam Neubig, Julian Michael, and Sydney Lamb, and by extending them through the geometric semantics of Peter Gärdenfors, we can create a framework in which to understand language and mind, a framework that is commensurate with the operations of GPT-3. That framework can help us to understand what GPT-3 is doing when it constructs a language model, and thereby to gain control over that model so we can enhance and extend it.

    It is in that speculative spirit that I offer the following remarks.


    Abstract: GPT-3 is an AI engine that generates text in response to a prompt given to it by a human user. It does not understand the language that it produces, at least not as philosophers understand such things. And yet its output is in many cases astonishingly like human language. How is this possible? Think of the mind as a high-dimensional space of signifieds, that is, meaning-bearing elements. Correlatively, text consists of one-dimensional strings of signifiers, that is, linguistic forms. GPT-3 creates a language model by examining the distances and ordering of signifiers in a collection of text strings and computes over them so as to reverse engineer the trajectories texts take through that space. Peter Gärdenfors’ semantic geometry provides a way of thinking about the dimensionality of mental space and the multiplicity of phenomena in the world, about how mind mirrors the world. Yet artificial systems are limited by the fact that they do not have a sensorimotor system that has evolved over millions of years. They do have inherent limits.

    Contents

    0. Starting point and preview 1
    1. Computers are strange beasts 4
    2. No meaning, no how: GPT-3 as Rubicon and Waterloo, a personal view 8
    3. The brain, the mind, and GPT-3: Dimensions and conceptual spaces 16
    4. Gestalt switch: GPT-3 as a model of the mind 24
    5. Engineered intelligence at liberty in the world 26

    0. Starting point and preview

    GPT-3 is based on distributional semantics. Warren Weaver had the basic idea in his 1949 memorandum, “Translation” (p. 8). Gerard Salton operationalized the idea in his work using vector semantics for document retrieval in the 1960s and 1970s (p. 9). Since then distributional semantics has developed as an empirical discipline. The last decade of work in NLP has seen remarkable, even astonishing, progress. And yet we lack a robust theoretical framework in which we can understand and explain that progress. Such a framework must also indicate the inherent limitations of distributional semantics. This document is a first attempt to outline such a framework, as such its various formulations must be seen as speculative and provisional. I offer them so that others may modify them, replace them, and move beyond them.

    It started with a comment at a blog

    On July 19, 2020, Tyler Cowen made a post to Marginal Evolution entitled “GPT-3, etc.” It consisted of an email from a reader who asserted, “When future AI textbooks are written, I could easily imagine them citing 2020 or 2021 as years when preliminary AGI first emerged,. This is very different than my own previous personal forecasts for AGI emerging in something like 20-50 years…” While I have my doubts about the concept of AGI – it’s too ill-defined to serve as anything other than a hook on which to hang dreams, anxieties, and fears – I think GPT-3 is worth serious consideration.

    Cowen’s post has attracted 52 comments so far, more than a few of acceptable or even high quality. I made a long comment to that post. I then decided to expand that comment into a series of blog posts, say three or four, and then to collect them into a single document as a working paper. When it appeared that those three or four posts would grow to five or six I decided that I would issue two working papers. This first one would concentrate on GPT-3 and the nature of artificial intelligence, or whatever it is. The second would speculate about the future and take a quick tour of the past.

    Here is a slightly revised version of the comment I made at Marginal Revolution. This paper covers the shaded material. The rest will be covered in the second paper.
    Yes, GPT-3 [may] be a game changer. But to get there from here we need to rethink a lot of things. And where that's going (that is, where I think it best should go) is more than I can do in a comment.

    Right now, we're doing it wrong, headed in the wrong direction. AGI, a really good one, isn't going to be what we're imagining it to be, e.g. the Star Trek computer.

    Think AI as platform, not feature (Andreessen). Obvious implication, the basic computer will be an AI-as-platform. Every human will get their own as an very young child. They're grow with it; it’ll grow with them. The child will care for it as with a pet. Hence we have ethical obligations to them. As the child grows, so does the pet – the pet will likely have to migrate to other physical platforms from time to time.

    Machine learning was the key breakthrough. Rodney Brooks’ Gengis, with its subsumption architecture, was a key development as well, for it was directed at robots moving about in the world. FWIW Brooks has teamed up with Gary Marcus and they think we need to add some old school symbolic computing into the mix. I think they’re right.

    Machines, however, have a hard time learning the natural world as humans do. We're born primed to deal with that world with millions of years of evolutionary history behind us. Machines, alas, are a blank slate.

    The native environment for computers is, of course, the computational environment. That's where to apply machine learning. Note that writing code is one of GPT-3's skills.

    So, the AGI of the future, let's call it GPT-42, will be looking in two directions, toward the world of computers and toward the human world. It will be learning in both, but in different styles and to different ends. In its interaction with other artificial computational entities GPT-42 is in its native milieu. In its interaction with us, well, we'll necessarily be in the driver’s seat.

    Where are we with respect to the hockey stick growth curve? For the last 3/4 quarters of a century, since the end of WWII, we've been moving horizontally, along a plateau, developing tech. GPT-3 is one signal that we've reached the toe of the next curve. But to move up the curve, as I’ve said, we have to rethink the whole shebang.

    We're IN the Singularity. Here be dragons.

    [Superintelligent computers emerging out of the FOOM is bullshit.]

    * * * * *

    ADDENDUM: A friend of mine, David Porush, has reminded me that Neal Stephenson has written of such a tutor in The Diamond Age: Or, A Young Lady's Illustrated Primer (1995). I then remembered that I have played the role of such a tutor in real life, The Freedoniad: A Tale of Epic Adventure in which Two BFFs Travel the Universe and End up in Dunkirk, New York.
    While the portion of the comment to be elaborated in the next working paper is considerably longer than the portion being elaborated in this one, I do not expect that paper to be proportionately longer. This paper covered quasi-technical matters requiring fairly careful exposition. The next paper will go by more quickly and will, in sections, approach science fiction.

    * * * * *

    1. Computers are strange beasts – They’re obviously inanimate, and yet we communicate with them through language. The don’t fit pre-existing (19th century?) conceptual categories, and so we are prone to strange views about them.

    2. No meaning, no how: GPT-3 as Rubicon and Waterloo, a personal view – Arguing from first principles it is clear that GPT-3 lacks understanding and access to meaning. And yet it produces very convincing simulacra of understanding. But common sense understanding remains elusive, as it did for old school symbolic processing. Much of common sense is deeply embedded in the physical world. GPT-3, as it currently functions is, in effect, an artificial brain in a vat.

    3. The brain, the mind, and GPT-3: Dimensions and conceptual spaces – GPT-3 creates a language model by examining the distances and ordering of signifiers in a collection of text strings and computes over them so as to reverse engineer but the trajectories texts take through a high-dimensional mental space of signifieds. Peter Gärdenfors’ semantic geometry provides a way of thinking about the dimensionality of mental space and the multiplicity of phenomena in the world.

    4. Gestalt switch: GPT-3 as a model of the mind – GPT-3 creates: 1) a model of a body of natural language texts, and only a model. 2) Those texts are the product of human minds. 3) Though the application of 2 to 1 we may conclude that GPT-3 is also a model of the mind, albeit a very limited one. 3 requires a Gestalt switch.

    5. Engineered intelligence at liberty in the world – The “intelligence” in systems such as GPT-3 is static and reactive. To liberate and mobilize it we need to endow AI systems with mental models of the kind investigated in “old school” symbolic AI.

    Wednesday, July 29, 2020

    2. The brain, the mind, and GPT-3: Dimensions and conceptual spaces

    [Edited, with a substantial addition, August 2, 2020]

    The purpose of this post is to sketch a conceptual framework in which we can understand the success of language models such as GPT-3 despite the fact that they are based on nothing more than massive collections of bare naked signifiers. There’s not a signified in sight, much less any referents. I have no intention of even attempting to explain how GPT-3 works. That it does work, in an astonishing variety of cases if (certainly) not universally, is sufficient for my purposes.

    First of all I present the insight that sent me down this path, a comment by Graham Neubig in an online conversation that I was not a part of. Then I set that insight in the context of and insight by Sydney Lamb (meaning resides in relations), a first-generation researcher in machine translation and computational linguistics. I think take a grounding case by Julian Michael, that of color, and suggest that it can be extended by the work of Peter Gärdenfors on conceptual spaces.

    A clue: an isomorphic transform into meaning space

    At the 58th Annual Meeting of the Association for Computational Linguistics Emily M. Bender and Alexander Koller delivered a paper, Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data [1], where NLU means natural language understanding. The issue is pretty much the one I laid out in my previous posts in the sections “No words, only signifiers” and “Martin Kay, ‘an ignorance model’” [2]. A lively discussion ensured online which Julian Michael has summarized and commented on in a recent blog post [3].

    In that post Michael quotes a remark by Graham Neubig:
    One thing from the twitter thread that it doesn’t seem made it into the paper... is the idea of how pre-training on form might learn something like an “isomorphic transform” onto meaning space. In other words, it will make it much easier to ground form to meaning with a minimal amount of grounding. There are also concrete ways to measure this, e.g. through work by Lena Voita or Dani Yogatama... This actually seems like an important point to me, and saying “training only on form cannot surface meaning,” while true, might be a little bit too harsh— something like “training on form makes it easier to surface meaning, but at least a little bit of grounding is necessary to do so” may be a bit more fair.
    That’s my point of departure in this post, that notion of “an ‘isomorphic transform’ onto meaning space.” I am going to sketch a framework in which we can begin unpacking that idea. But it may take awhile to get there.

    Meaning is in relations

    I want to develop an idea I have from Sydney Lamb, that meaning resides in relations. The idea is grounded in the “old school” world of symbolic computation, where language is conceived as a relational network of items. The meaning of any item in the network is a function of its position in the network.

    Let’s start with this simple diagram:


    It represents the fact that the central nervous system (CNS) is coupled to two worlds, each external to it. To the left we have the external world. The CNS is aware of that world through various senses (vision, hearing, smell, touch, taste, and perhaps others) and we act in that world through the motor system. But the CNS is also coupled to the internal milieu, with which it shares a physical body. The net is aware of that milieu by chemical sensors indicating contents of the blood stream and of the lungs, and by sensors in the joints and muscles. And it acts in the world through control of the endocrine system and the smooth muscles. Roughly speaking the CNS guides the organism’s actions in the external world so as to preserve the integrity of the internal milieu. When that integrity is gone, the organism is dead.

    Now consider this more differentiated presentation of the same facts:


    I have divided the CNS into four sections: A) senses the external world, B) senses the internal milieu, D) guides action in the internal milieu, and D) guides action in the external world. I rather doubt that even a very simple animal, such as C. elegans, with 302 neurons, is so simple. But I trust my point will survive that oversimplification.

    Lamb’s point is that the “meaning” or “significance” of any of those nodes – let’s not worry at the moment whether they’re physical neurons or more abstract entities – is a function of its position in the entire network, with its inputs from and outputs to the external world and the inner milieu [4]. To appreciate the full force of Lamb’s point we need to recall the diagrams typical of old school symbolic computing, such as this diagram from Brian Phillips we used in the previous post:
    All of the nodes and edges have labels. Lamb’s point is that those labels exist for our convenience, they aren’t actually a part of the system itself. If we think of that network as a fragment from a human cognitive system – and I’m pretty sure that’s how Phillips thought about it, even if he could not justify it in detail (no one could, not then, not now) – then it is ultimately connected to both the external world and the inner milieu. All those labels fall away; they serve no purpose. Alas, Phillips was not building a sophisticated robot, and so those labels are necessary fictions.

    But we’re interested in the full real case, a human being making their way in the world. In that case let us assume that, for one thing, the necessary diagram is WAY more complex, and that the nodes and edges do not represent individual neurons. Rather, they represent various entities that are implemented in neurons, sensations, thoughts, perceptions, and so forth. Just how such things are realized in neural structures is, of course, a matter of some importance and is being pursued by hundreds of thousands of investigators around the world. But we need not worry about that now. We’re about to fry some rather more abstract fish (if you will).

    Some of those nodes will represent signifiers, to use the Saussurian terminology I used in my previous post, and some will represent signifieds. What’s the difference between a signifier and a signified? Their position in the network as a whole. That’s all. No more, no less. Now, it seems to me, we can begin thinking about Neubig’s “isomorphic transform” onto meaning space.

    Let us notice, first of all, that language exists as strings of signifiers in the external world. In the case that interests us, those are strings of written characters that have been encoded into computer-readable form. Let us assume that the signifieds – which bear a major portion of meaning, no? – exist in some high dimensional network in mental space. This is, of course, an abstract space rather than the physical space of neurons, which is necessarily three dimensional. However many dimensions this mental space has, each signified exists at some point in that space and, as such, we can specify that point by a vector containing its value along each dimension.

    What happens when one writes? Well, one produces a string of signifiers. The distance between signifiers on this string, and their ordering relative to one another, are a function of the relative distances and orientations of their associated signifieds in mental space. That’s where to look for Neubig’s isometric transform into meaning space. What GPT-3, and other NLP engines, does is to examine the distances and ordering of signifiers in the string and compute over them so as to reverse engineer the distances and orientations of the associated signifieds in high-dimensional mental space.
    [A little reflection on that formulation makes it clear that it fails to take into account a distinction central to ‘old school’ symbolic computation, that between semantic and episodic memory. Rather than interrupt this argument with a refined formulation I have placed that in an appendix to this post: A more refined approach to meaning space. I also offer some remarks on need for a connection to the physical world in order to handle common-sense reasoning.]
    Is the result perfect? Of course not – but then how do we really know? It’s not as though we’ve got a well-accepted model of human conceptual space just lying around on a shelf somewhere. GPT-3’s language model is perhaps as good as we’ve got at the moment, and we can’t even open the hood and examine it. We know its effectiveness by examining how it performs. And it performs very well.

    Monday, July 27, 2020

    1. No meaning, no how: GPT-3 as Rubicon and Waterloo, a personal view

    I say that not merely because I am a person and, as such, I have a point of view on GPT-3, and related matters. I say because the discussion is informal, without journal-class discussion of this, that, and the others, along with the attendant burden of citation, though I will offer a few citations. More over, I’m pretty much making this up as I go along. That is to say, I am trying to figure out just what it is that I think, and see value in doing so in public.

    What value, you ask? It commits me to certain ideas, if only at a certain time. It lays out a set of priors and thus serves to sharpen my ideas developments unfold and I, inevitably, reconsider.

    GPT-3 represents an achievement of a high order; it deserves the attention it has received, if not the hype. We are now deep in “here be dragons” territory and we cannot go back. And yet, if we are not careful, we’ll never leave the dragons, we’ll always be wild and undisciplined. We will never actually advance; we’ll just spin faster and faster. Hence GPT-3 is both a Rubicon, the crossing of a threshold, and a potential Waterloo, a battle we cannot win.

    Here’s my plan: First we take a look at history, at the origins of machine translation and symbolic AI. Then I develop a fairly standard critic of semantic models such as those used in GPT-3 which I follow with some remarks by Martin Kay, one of the Grand Old Men of computational linguistics. Then I look at the problem of common sense reasoning and conclude be looking ahead to the next post in this series in which I offer some speculations on why (and perhaps even how) these models can succeed despite their sever and fundamental short-comings.

    Background: MT and Symbolic computing

    It all began with a famous memo Warren Weaver wrote in 1949. Weaver was director of the Natural Sciences division of the Rockefeller Foundation from 1932 to 1955. He collaborated Claude Shannon in the publication of a book which popularized Shannon’s seminal work in information theory, The Mathematical Theory of Communication. Weaver’s 1949 memorandum, simply entitled “Translation” [1], is regarded as the catalytic document in the origin of machine translation (MT) and hence of computational linguistics (CL) and heck! why not? artificial intelligence (AI).

    Let’s skip to the fifth section of Weaver’s memo, “Meaning and Context” (p. 8):
    First, let us think of a way in which the problem of multiple meaning can, in principle at least, be solved. If one examines the words in a book, one at a time as through an opaque mask with a hole in it one word wide, then it is obviously impossible to determine, one at a time, the meaning of the words. “Fast” may mean “rapid”; or it may mean "motionless"; and there is no way of telling which.

    But if one lengthens the slit in the opaque mask, until one can see not only the central word in question, but also say N words on either side, then if N is large enough one can unambiguously decide the meaning of the central word. The formal truth of this statement becomes clear when one mentions that the middle word of a whole article or a whole book is unambiguous if one has read the whole article or book, providing of course that the article or book is sufficiently well written to communicate at all.
    It wasn’t until the 1960s and ‘70s that computer scientists would make use of this insight; Gerard Salton was the central figure and he was interested in document retrieval [2]. Salton would represent documents as a vector of words and then query a database of such representation by using a vector composed from user input. Documents were retrieved as a function of similarity between the input query vector and the stored document vector.

    Work on MT went a different way. Various approaches were used, but at some relatively early point researchers were writing formal grammars of languages. In some cases these grammars were engineering conveniences while in others they were taken to represent the mental grammars of humans. In any event, that enterprise fell apart in the mid-1960s. The prospects for practical results could not justify federal funding and the government had interest in supporting purely scientific research into the nature of language.

    But such research continued nonetheless, sometimes under the rubric of computational linguistics (CL) and sometimes as AI. I encountered CL in graduate school in the mid-1970s when I joined the research group of David Hays in the Linguistics Department of the State University of New York at Buffalo – I was actually enrolled as a graduate student in English; it’s complicated.

    Many different semantic models were developed, but I’m not interested in anything like a review of that work, just a little taste. In particular I am interested in a general type of model was known as a semantic or cognitive network. Hays had been developing such a model for some years in conjunction with several graduate students [2]. Here’s a fragment of a network from a system developed by one of those students, Brian Phillips, to tell whether or not stories of people drowning were tragic [3]. Here’s a representation of capsize:
    Notice that there are two kinds of nodes in the network, square ones and smaller round ones. The square ones represent a scene while the round ones represent individual objects or events. Thus the square node at the upper left indicates a scene with two sub-scenes – I’m just going to follow out the logic of the network without explaining it in any detail. The first one asserts that there is a boat that contains one Horatio Smith. The second one asserts that the boat overturns. And so forth through the rest of the diagram.

    This network represents semantic structure. In the terminology of semiotics, it represents a network of signifieds. Though Phillips didn’t do so, it would be entirely possible to link such a semantic network with a syntactic network, and many systems of that era did so.

    Such networks were symbolic in the (obvious) sense that the objects in them were considered to be symbols, not sense perceptions or motor actions nor, for that matter, neurons, whether real or artificial. The relationship between such systems and the human brain was not explored, either in theory or in experimental observation. It wasn’t an issue.

    That enterprise collapsed in the mid-1980s. Why? The models had to be hand-coded, which took time. They were computationally expensive and so-called common sense reasoning proved to be endless, making the models larger and larger. (I discuss common sense below and I have many posts at New Savanna on the topic [4].)

    Oh, the work didn’t stop entirely. Some researchers kept at it. But interests shifted toward machine learning techniques and toward artificial neural networks. That is the line of evolution that has, three or four decades later, resulted in systems like GPT-3, which also owe a debt to the vector semantics pioneered by Salton. Such systems build huge language models from huge databases – GPT-3 is based on 500 billion tokens [5] – and contain no explicit models of syntax or semantics anywhere, at least not that researchers can recognize.

    Researchers build a system that constructs a language model (“learns” the language), but the inner workings of that model are opaque to the researchers. After all, the system built the model, not the researchers. They only built the system.

    It is a strange situation.

    Monday, October 28, 2019

    What is computation? That is to say, what do I mean by computation? [putting things in order]

    As far as I know, the nature of computation is still under investigation. I’m not really qualified to or in fact interested in addressing the question in its full scope and generality. I’m interested in a more limited question – though not, as these things go, all that limited – of the computational aspects of the human mind. And within that, I’m particularly interested in language, and in literature, which of course includes language, but more as well. How much more...who knows?

    So, I start out with the abstract idea of computation, then introduce the idea of implementing computation in a physical system and conclude by observing that the computational simulation of a system is not to be confused with the thing itself.

    Abstract computation

    Abstractly considered, Turing defined computation in terms of a machine that had, 1) a set of symbols, 2) a paper tape on which symbols could be written and from which they could be erased, 3) a device that read from and wrote to the tape, and 4) an instruction set defining relations between the symbols and specifying writing to and erasing from the tape. We need not go beyond that. My point is that we do have a well-known and thoroughly explored account of computation, and that that account is stated in terms of an abstract machine.

    Real computation requires physical implementation

    I’m not interested in abstract computation on an abstract machine. I’m interested in real computation on a real device of some kind. Given some device, how can we implement computation on that device. It is the idea of implementation that is key.

    The device that interests me, of course, is the human brain. And the conclusion I’ve reached over the past few years is that natural language is the simplest activity that requires computation. Language cannot be explained and understood without reference to computation. By implication then we should be able to understand, for example, visual perception without reference to computation. This implies that the full powers of an advanced primate brain are necessary for implementing computation. I suppose that’s the take-out from a paper David Hays and I published in the 1988:
    William Benzon and David Hays, Principles and Development of Natural Intelligence, Journal of Social and Biological Structures, Vol. 11, No. 8, July 1988, 293-322, https://www.academia.edu/235116/Principles_and_Development_of_Natural_Intelligence
    We didn’t quite put things in those terms, but that paper justifies them. We need not going into the details here.

    And so, going back to March of 2016, I’ve written a series of posts on that theme. I’ve collected them under the label, “computational envelope”. This, of course, is another post in that series.

    Simulation is not the thing itself

    Now we need one more idea, that of simulation. Digital computers can be, have been, and are being used to simulate all sorts of things. But the simulation of a thing is not to be confused with the thing itself. A simulation of an atomic explosion is quite a different phenomenon from a real atomic explosion. And so it is for many other things as well.

    And then we have the brain, of any animal, and the human mind. In this case there seems to be some difficulty in distinguishing between a simulation of the thing and the thing itself. It’s not that anyone is confused about the difference between a digital computer, but rather that there is a suspicion that, if we simulate mental processes on a digital computer with sufficient precision and power, then perhaps that computer is not merely running a simulation of a mind, but is in fact a mind. I say let’s set that one aside until we actually confront the situation. So far, we are no where near that.

    Now we’ve arrived at the point of this post, a passage from a most interesting book by Peter Gärdenfors, Conceptual Spaces (MIT 2000) p. 253:
    On the symbolic level, searching, matching, of symbol strings, and rule following are central. On the subconceptual level, pattern recognition, pattern transformation, and dynamic adaptation of values are some examples of typical computational processes. And on the intermediate conceptual level, vector calculations, coordinate transformations, as well as other geometrical operations are in focus. Of course, one type of calculation can be simulated by one of the others (for example, by symbolic methods on a Turing machine). A point that is often forgotten, however, is that the simulations will, in general be computationally more complex than the process that is simulated.
    I rather suspect that all of these kinds of processes take place in the human brain. Only the symbolic level processes however, are irreducibly computational as implemented in the human brain. The other processes are implemented in some non-computational way.

    Pattern recognition and transformation might be implemented in neurodynamics while coordinate transformations might, in part, be carried out by the physical structure of region to region mapping in the brain. Whatever. But the scientific investigation of human perception and cognition may require us to simulate any and all of these processes in a computer – as indeed, Walter Freeman has implemented dynamical processes in understanding how odors are recognized and remembered. The fact that we can simulate these processes computationally does not, of course, imply that they are computational in the brain.

    Coda

    It is my impression that a have of confusion has arisen through a failure to distinguish between the need for implementation on the one hand and the difference between simulation and reality on the other. That’s more than I want to go into here and now.