Showing posts with label Rubicon-Waterloo. Show all posts
Showing posts with label Rubicon-Waterloo. Show all posts

Tuesday, September 5, 2023

World, mind, and learnability: A note on the metaphysical structure of the cosmos [& LLMs]

I originally posted this three years ago, on August 15.  I have added an important new section at the end, Paths of the mind (virtual reading), and so I am bumping this to the top of the queue.
There is no a priori reason to believe that world has to be learnable. But if it were not, then we wouldn’t exist, nor would (most?) animals. The existing world, thus, is learnable. The human sensorium and motor system are necessarily adapted to that learnable structure, whatever it is.

I am, at least provisionally, calling that learnable structure the metaphysical structure of the world. Moreover, since humans did not arise de novo that metaphysical structure must necessarily extend through the animal kingdom and, who knows, plants as well.

“How”, you might ask, “does this metaphysical structure of the world differ from the world’s physical structure?” I will say, again provisionally, for I am just now making this up, that it is a matter of intension rather than extension. Extensionally the physical and the metaphysical are one and the same. But intensionally, they are different. We think about them in different terms. We ask different things of them. They have different conceptual affordances. The physical world is meaningless; it is simply there. It is in the metaphysical world that we seek meaning. [See my post, There is a fold in the fabric of reality. (Traditional) literary criticism is written on one side of it. I went around the bend years ago.]

A little dialog

Does this make sense, philosophically? How would I know?

I get it, you’re just making this up.

Right.

Hmmmm… How does this relate to that object-oriented ontology stuff you were so interested in a couple of years ago?

Interesting question. Why don’t you think about it and get back to me.

I mean, that metaphysical structure you’re talking about, it seems almost like a complex multidimensional tissue binding the world together. It has a whiff of a Latourian actor-network about it.

Hmmm… Set that aside for awhile. I want to go somewhere else.

Still on GPT-3, eh?

You got it.[1]
 
A little diagram: World, Text, and Mind

Text reflects this learnable, this metaphysical, structure, albeit at some remove:

Learning engines are learning the structure inherent in the text. But that learnable structure is not explicit in the language model created by the learning engine.

There are two things in play: 1) the fact that the text is learnable, and 2) that it is learnable by a statistical process. How are these two related?

If we already had an explicit ‘old school’ propositional model in computable form, then we wouldn’t need statistical learning at all. We could just run the propositional model over the corpus and encode the result. But why do even that? If we can read the corpus with the propositional model, in a simulation of human reading, then there’s no need to encode it at all. Just read whatever aspect of the corpus is needed at the time.

So, statistical learning is a substitute for the lack of a usable propositional model. The statistical model does work, but at the expense of explicitness.

But why does the statistical model work at all? That’s the question.

It’s not enough to say, because the world itself is learnable. That’s true for the propositional model as well. Both work because the world is learnable.

Language model as associative memory

BUT: Humans don’t learn the world with a statistical model. We learn it through a propositional engine floating over an analogue or quasi-analogue engine with statistical properties. And it is the propositional engine that allows us to produce language. A corpus is a product of the action of propositional engine, not a statistical model, acting on the world.

Description is one basic such action; narration is another. Analysis and explanation are perhaps more sophisticated and depend on (logically) prior description and narration. Note that this process of rendering into language is inherently and necessarily a temporal one. The order in which signifiers are placed into the speech stream depends in some way, not necessarily obvious, on the relations among the correlative signifieds in semantic or cognitive space. Distances between signifiers in the speech stream reflect distances between correlative signifieds in semantic space. We thus have systematic relationships between positions and distances of signifiers in the speech stream, on the one hand, and positions and distances of signifieds in semantic space. It is those systematic relationships that allow statistical analysis of the speech stream to reconstruct semantic space.

Note that time is not extrinsic to this process. Time is intrinsic and constitutive of computation. Speaking involves computation, as does the statistical analysis of the speech stream.

The propositional engine learns the world via Gärdenfors’ dimensions [2], and whatever else, Powers’ stack for example [3]. Those dimensions are implicit in the resulting propositional model and so become projected onto the speech stream via syntax, pragmatics, and discourse structure. The language engine is then able to extract (a simulacrum of) those dimensions through statistical learning. Those dimensions are expressed in the parameter weights of the model. THAT’s what makes the knowledge so ‘frozen’. One has to cue it with actual speech.

The whole language model thus functions as associative memory [4]. You present it with an input cue, and it then associates from that cue with each emitted string ‘feeding back’ into the memory bank via associative memory.
 
Paths of the mind (virtual reading)
 
Now, imagine a word embedding model constructed over some suitable corpus of texts. Given that texts reflect the interaction of the mind and the world, the location of individual words in that model necessarily reflects that interaction. That structure is what I have been calling the metaphysical structure of the cosmos.

Consider some text. It consists of word after word after word. That sequence reflects the actions of the mind that wrote the text, and only the mind. For all practical purposes, the cosmos remains unchanging during the writing of that text and the mind has withdrawn from active interaction with the world, except insofar as the text is a set of symbols that it is placing on a piece of paper, moment after moment, word by word. Now let us trace the path some texts takes through a word embedding. Call this a virtual reading. The word embedding consists of tens of 1000s of dimensions, so that path will be a complicated one. That path must necessarily be a product of the mind (and only the mind?). That path is the mind in action.

Further, consider that the mind is what the brain does. The brain consists of 86 billion neurons, each having on the order of 10,000 connections with other neurons. The dimensionality of the space required to represent the brain is thus huge. At any given moment we can represent the state of the brain as a point in that space. From one moment to the next, the brain traces a path in that space, what a dynamicist (such as Walter Freeman) would call a trajectory. When someone is writing a text, that text reflects the operations of their brain. Therefore the path of a text through a word embedding necessarily mirrors the trajectory taken by the author's brain while writing the text. Notice, however, the reduction in dimensionality. The brain's state space is of vastly higher dimensionality than the word embedding space.

Finally, consider the operations of a transformer as it generates a text. Each time it generates a token it takes the entire model into account, all those 100s of billions of parameters (are we up to trillions yet?). Compare that to what a brain did when generating that same text. Is it reasonable to consider a single token as an image of, a reflection of, short trajectory segment in the state space of the brain that generated the text. Walter Freeman thought of the brain as moving through states of global coherence at the rate of 7-10 Hz, like frames of a film [5]. In what way is the movement of a transformer from one token to the next comparable to the movement of the brain from one frame to the next?[6]


References
 
[1] This post is an exploration of ideas raised in the course of thinking about GPT-3. See William Benzon, GPT-3: Waterloo or Rubicon? Here be Dragons, Working Paper, August 5, 2020, 32 pp., Academia: https://www.academia.edu/s/9c587aeb25; SSRN: https://ssrn.com/abstract=3667608
ResearchGate: https://www.researchgate.net/publication/343444766_GPT-3_Waterloo_or_Rubicon_Here_be_Dragons.

[2] Peter Gärdenfors, Conceptual Spaces: The Geometry of Thought, MIT Press, 2000; The Geometry of Meaning: Semantics Based on Conceptual Spaces, MIT Press, 2014.

[3] William Powers, Behavior: The Control of Perception (Aldine) 1973. A decade later David Hays integrated Powers’ model into his cognitive network model, David G. Hays, Cognitive Structures, HRAF Press, 1981.

[4] The idea that the brain implements associative memory in a holographic fashion was championed by Karl Pribram in the 1970s and 1980s. David Hays and I drew on that work in an article on metaphor, William Benzon and David Hays, Metaphor, Recognition, and Neural Process, The American Journal of Semiotics , Vol. 5, No. 1 (1987), 59-80, https://www.academia.edu/238608/Metaphor_Recognition_and_Neural_Process.
 
[5] Freeman, W. J. (1999a). Consciousness, Intentionality and Causality. Reclaiming Cognition. R. Núñez and W. J. Freeman. Thoverton, Imprint Academic, 143-172.
 
[6] I first argue this point in William Benzon, The idea that ChatGPT is simply “predicting” the next word is, at best, misleading, New Savanna, Feb. 19, 2023.

Monday, April 25, 2022

Semanticity: adhesion and relationality

For some time now I have been puzzled by the (astonishing) success of statistical models, especially artificial neural networks, in language processing, machine translation in particular – see, e.g. this post, Borges redux: Computing Babel – Is that what’s going on with these abstract spaces of high dimensionality? [#DH], which dates back to October 2017. Sure, it is statistics, but what are those statistics “grabbing on to?” There is no meaning there, just (naked) word forms. What is it about those word-forms-in-context that yields an approximation to, a simulacrum of, meaning? Better, what’s there that we readily read meaning, intention, into the results of these statistical techniques.

Semanticity and intention

My puzzlement reached a climax with the unveiling of GPT-3 in 2020 and I decided to take another run at the problem. I produced a working paper, GPT-3: Waterloo or Rubicon? Here be Dragons, which I liked very much. I made real progress. I now think I can nudge things forward another step. Look at this passage, where I discuss the Chinese Room thought-experiment (p. 28):

Yet if you would believe John Searle, no matter how rich and detailed those old school mental models, understanding would necessarily elude them. I am referring, of course, to his (in)famous Chinese Room argument. When I first encountered it years ago my reaction was something like: interesting, but irrelevant. Why irrelevant? Because it said absolutely nothing about the techniques AI or cognitive science investigators used and so would provide no guidance toward improving that work. He did, however, have a point: If the machine has no contact with the world, how can it possibly be said to understand anything at all? All it does is grind away on syntax.

What Searle misses, though, is the way in which meaning is a function of relations among concepts, as I pointed out earlier (pp. 18 ff.). It seems to me, however – and here I’m just making this up – we can think of meaning as having both an intentional aspect, the connection of signs to the world, and a relational aspect, the relations of signs among themselves. Searle’s argument concentrated on the former and said nothing about the latter.

What of the intentional aspect when a person is writing or talking about things not immediately present, which is, after all quite common? In this case the intentional aspect of meaning is not supported by the immediate world. Language use thus must necessarily be driven entirely by the relations signifiers have among themselves, Sydney Lamb’s point which we have already investigated (p. 18).

Those statistics are grabbing onto the relational aspect of meaning. The question is: How much of that can these methods recover from texts? Let’s set that aside for the moment.

That passage mentions intention and relation. Intention resides in the relationship between a person and the world. Relation resides in the relationships that signifiers have among themselves. It is a property of the cognitive system. I am now thinking that it must be paired with adhesion. Taken together they constitute semanticity. Thus we have semanticity and intention where semanticity is a general capacity inherent in the cognitive system, in a person’s mind, and intention inheres in the relation between a person and the world in a particular perceptual and/or cognitive activity.

What do I mean by adhesion? Adhesion is how words ‘cling’ to the world while relationality is the differential interaction of words among themselves within the linguistic system. Words whose meaning is defined directly over the physical world, but also, to some extent, the interpersonal world of signals and feeling, they adhere to the world through sensorimotor schemas. Words whose meaning is abstract are more problematic. Their adhesion operates though patterns of words and other signs and symbols (e.g. mathematics, data visualizations, illustrative diagrams of various kinds, and so forth). Teasing out these systems of adhesion has just barely begun.

The psychologist J.J. Gibson talked of the affordances an environment presents to the organism. Affordances as the features of the world which an organism can readily pick up during its life in the world. Adhesions are the organism’s complement to environmental affordances; they are the perceptual devices through which the organism relates to the affordances.

What this means for language models

Large language models built through deep neural networks, such as GPT-3, conflate the interaction of three phenomena: 1) the world-level relational aspect of semanticity as captured in the locations of word forms (signifiers) in a string, 2) the conventions of discourse structure, and 3) the world itself. The world is present in the model because the texts over which the model was constructed were created by people interacting in the world. They were in an intentional relationship with the world when they wrote those texts. The conventions of discourse are present simply because they organize the placement of word forms in a text, with special emphasis on the long-distance relationships of word. As for relationality, that’s all that can possibly be present in a text. Adhesions belong to the realm of signifieds, of concepts and ideas, and they aren’t in the text itself.

Would it somehow be possible to factor a language model into these three aspects? I have no idea. The point of doing so would be to reduce the overall size of the model.

Putting that aside, let us ask: Given a sufficiently large database of texts and tokens and a high enough number of parameters for our model, is it possible for a language model to extract all the relationality from the texts? How much of that multidimensional relational semanticity can be recovered from strings of word forms? Given a deep enough understanding of how relational semantics is reflected in the structure of texts, can we calculate what is possible with various text bases and model parameterization?

To answer those questions we need to have some account of semantic relationality which we can examine. The models of Old School symbolic AI and computational linguistics provide such accounts. Many such models have been created. Which ones would we choose as the basis for our analysis? The sort of question that interests me is how many word forms have their meanings given in adhesions to the physical world (that is, physical objects and events), to the interpersonal world (facial expressions, gestures, etc.) and how many word forms are defined abstractly?

So many questions. 

* * * * *

I have appended this to my GPT-3 working paper, which is now Version 3.

Friday, September 25, 2020

Gwern on the implications of GPT-3 ["no coherent model of why GPT-3 was possible"]

I'm not a regular follower of Gwern, though I did check out what he has to say about GPT-3 and poetry, so I only just now noticed this statement:
...GPT-3’s scaling curves, unpredicted meta-learning, and success on various anti-AI challenges suggests that in terms of futurology, AI researchers’ forecasts are an emperor sans garments: they have no coherent model of how AI progress happens or why GPT-3 was possible or what specific achievements should cause alarm, where intelligence comes from, and do not learn from any falsified predictions. Their primary concerns appear to be supporting the status quo, placating public concern, and remaining respectable. As such, their comments on AI risk are meaningless: they would make the same public statements if the scaling hypothesis were true or not.
While Gwern appears to believe in AI in a way that I do not, I agree with this. And that is what prompted my recent thinking on GPT-3 in the first place, in particular, my working papers, GPT-3: Waterloo or Rubicon? Here be Dragons, from August 5, and the more recent, What economic growth and statistical semantics tell us about the structure of the world, from August 24.

Gwern concludes that assessment with this question:  "Depending on what investments are made into scaling DL, and how fast compute grows, the 2020s should be quite interesting—sigmoid or singularity?" I do expect the 2020s to be interesting, but I don't expect sigmoidal from GPT-X and similar engines, and not singularity from anything. Though, as I've been arguing for awhile, we're already swimming in a singularity.

Tuesday, September 1, 2020

GPT-3 meets “Kubla Khan” and the results are interesting, but not encouraging for AI poetry

I experimented a bit with GPT-3 and poetry in conjunction with my interview with Hollis Robbins. As you recall, she had written a book about the African-American sonnet tradition. I suggested, then, that we re-enact the contest between John Henry and the steam-drill as a contest between a real poet, she chose Marcus Christian, and GPT-3. GPT-3 produced one sonnet, sorta’, which we covered in our interview on July 20 [1], and a second, which I commented on in a post, GPT-3 writes two sonnets, sorta’. The first is better than the second. [digital humanities], on July 21, 2020.

Quite independently of this Gwern Branwen had been experimenting with poetry using both GPT-2 and GPT-3. He offered this general remark about prompting GPT-3’s performance:
For the prompt, just using the title/author is slightly unsatisfactory, as GPT-3 has memorized many of these famous poems and will, if you do not specify otherwise, happily complete them (which is an entirely valid thing to do, and does not represent ‘overfitting’—even when a completion does start with the next real lines, it will typically diverge in an entirely acceptable way after a few real lines). After playing around for a while, I thought I’d embrace the “Transformer Poetry” theme and make that the prompt:

Transformer AI poetry: Poetry classics as reimagined and rewritten by an artificial intelligence.
Using Coleridge’s “Kubla Khan” as a prompt for GPT-3

One of the poems he choose to work with is “Kubla Khan” (scroll down the page):
Transformer AI poetry: Poetry classics as reimagined and rewritten by an artificial intelligence.
Kubla Khan
Samuel Taylor Coleridge
In Xanadu did Kubla Khan
A stately pleasure-dome decree:
Where Alph, the sacred river, ran
Through caverns measureless to man
Down to a sunless sea.
I
saw the sacred river Alph
Rolling darkly, angrily,
Its roaring waters hurled
Down to the sunless sea.
Then I saw the dwelling-place
Of Coleridge in Xanadu:
A hundred caverns lofty, vast,
Down to a sunless sea.
I saw the caverns where he wrote,
And where he played the immortal strains
And poured out all his soul,
In lavish abundance:
This cavern was his Dawl, his Malta,
And where his enchanted spirit
Broke from its bounds in triumph,
Filled with the Glory he foresaw,
Piping and singing of immortal things
And ever-rising to immortal heights,
A hundred fathom deep.
That vision had the Prophecy
Which filled the world with song:
The glory of Kubla Khan
And Coleridge’s Song of Xanadu.
Notice the single capital letter “I” at the beginning of the line after “sunless sea”. I asked Gwern about that and he responded that it was necessary to keep GPT-3 from simply completing the poem as Coleridge had written it as GPT-3 had certainly “memorized” the original, which would be the obvious completion for the prompt.

General Comments on GPT-3’s performance

GPT-3’s continuation falls roughly into two components, the first continues through “In lavish abundance” and the second picks up with the next line and continues on to the end. The first component more or less works within the ‘territory’ indicated in the prompt, which includes the first five lines of the original poem. It emphasizes that territory. The “sunless sea” line is repeated twice, Alph shows up again, and we have the waters, the cavern, Xanadu, and Coleridge himself. The poem then begins shifting toward the second component with “I saw the caverns...immortal strains...all his soul...”

The second component shifts toward the poet himself, Coleridge, and his poetizing. We have his “enchanted spirit” breaking free; the poem is foresees, pipes, sings, rises to immortal heights, albeit “A hundred fathom deep” and so forth, onto song, glory, and Xanadu. For what it’s worth, “Dawl” gave me a pause and I had to do a bit of digging around before Google Translate told me that it is is Maltese for light.

The whole thing is rather rough, but that loose two-part structure is interesting. It gives the whole thing a crude coherence. On other hand, the repetition of lines from the prompt and the inclusion of Coleridge’s name are annoying. They follow, of course, from the nature of this exercise, and GPT-3 composes by, in effect, ‘predicting’ the next word, and the next, and so on. Finally, I note that while the versification of “Kubla Khan” is intricate, with varying line lengths, alliteration at various points, and a complex rhyme scheme, there’s not much to be said for GPT-3’s versification.

More specific comments

“Kubla Khan” presents a peculiar challenge to an AI engine that is trained, as GPT-3 is, to guess what comes next. When faced with this kind of task, to create a poem by continuing on from the initial lines of a human-produced poem, one wants, I presume, something more or less like the original, but different. The world contains zillions of sonnets, for example, but only one “Kubla Khan”. Coleridge never wrote another poem like it – I’ve read them all, though some years a go – nor, as far as I know, has anyone else. So GPT-3 has no other models to go by.

I note further that the poem does in fact have a very elaborate formal structure, which I have described in some detail [2], so one can imagine another poet writing a poem in the manner of Coleridge’s original, call it a Kubla, though they’d have to decided just which aspects of that structure are important and which are not. For example, “Kubla Khan” is 54 lines long and in two parts. The first is 36 lines long and the second is 18 lines. Do we require the same of a new Kubla? Or is it sufficient that a Kubla have a first section that is twice as long as the second, say 30 and 15, or 20 and 10, or for that matter, 44 and 22, and so forth? Once we change the length, however, we’re going to have to alter the rhyme scheme. The poet makes such decisions, writes and poem, and we can judge the results.

GPT-3 obviously did nothing of the kind.

Wednesday, August 12, 2020

Reflections on and current status of my GPT-3 project

As I noted at the beginning of the month the project began with a long comment posted to Marginal Revolution on July 19 [see below for a copy of that comment]. My original idea was to elaborate on that comment in a series of posts. It soon became clear, however, that things were going to more complicated.

On August 5 I issued the first working paper in the project, with the expectation that there would be a second one. That first paper is entitled, GPT-3: Waterloo or Rubicon? Here be Dragons. At that time I expected to issue a second working paper to cover the rest of the material from that original comment.

And then things became even more complicated. What happened is that I started thinking over the material in the first working paper and reading more about GPT-3. It was like when I first came to terms with topic models. The hardcore technical literature is a bit beyond me, but the surrounding explanatory material wasn’t doing it for me. For a couple of days I didn’t know whether I was coming or going.

A Plan, three more papers

Now things have cleared up. I think. At any rate I now have a plan, which is to issue not one, but three more working papers, two shorter ones and a longer one. The longer one will cover the rest of the material from the original comment below while the other two will go into greater depth on specific issues. This is the plan:
GPT-3: The Star Trek computer, and beyond
GPT-3: Bounding the Space, toward a theory of minds
Why GPT-X will fail in creating literature
I’ve been working on all three, but my current plan is to issue the future-oriented one – Star Trek computer – next, thereby covering the full scope of that original comment. I will issue the other two papers as they are ready.

But who knows, things may change. There’s no telling where a mind will go once it’s got the scent. Here’s brief notes on the other two working papers.

GPT-3: Bounding the Space, toward a theory of minds

This is really re-working and expanding on two sections from the first paper: 3. The brain, the mind, and GPT-3: Dimensions and conceptual spaces, and 5. Engineered intelligence at liberty in the world. I’ll be making sense out of this:
1. Symbolic AI: Construct a model of what the mind’s doing and run that model.

2. Machine learning: Construct a learning architecture (e.g. GPT-3), feed it piles of examples, and let it figure out what’s going on inside.

3. The question I’ve been getting to: What’s the world have to be like in order for 2 to work at all.

4. And 3 reflects back on 1: If THAT’s how the world is, what kind of (symbolic) model will produce output such that 2 will work.

And so forth
The third proposition is particularly important. That’s where the semantics of Peter Gärdenfors comes into play.

Why GPT-X will fail in creating literature

There’s a pro forma discussion of that issue: GPT-3 is not human, doesn’t have emotion, and is not creative. I suppose we could think of that as Commander Data’s problem, since he was forever fretting about it.

I suppose it’s true enough. But it doesn’t interest me. I have a much narrower and more specific issue in mind: GPT-X can’t do rhyme and neither will GPT-X. It’s a limitation that is inherent in the technology. Rhyme is a feature of how a text sounds, and the text base on which GPT-3 is built doesn’t have sound in it, nor is it at all obvious how that deficiency can be remedied.

If it can’t do rhyme, then it can’t do meter either. Nor can it do prose rhythm, which also depends, if not directly on sound, certainly on timing. Without these, GPT-X cannot do literature. At least it can’t do good literature, much less great literature. Oh, it can crank out wacky language by the bucket full, but that’s not what poetry is, and it can tell stories too. But stories are only a beginning point, not the end.

Think about it: Computers play the best chess in the world, Go too. But it’s not at all clear whether or not they’ll ever do anything more than mediocre literature. And that, I supposed, brings us back to the fact that computers aren’t human.

And they aren’t. They’re computers.

Wednesday, August 5, 2020

GPT-3: Waterloo or Rubicon? Here be Dragons


I've published a new working paper. Title above, download links, abstract, table of contents, and introduction below.

Download at:

GPT-3 is a significant achievement.

But I fear the community that has created it may, like other communities have done before – machine translation in the mid-1960s, symbolic computing in the mid-1980s, triumphantly walk over the edge of a cliff and find itself standing proudly in mid-air.

This is not necessary and certainly not inevitable.

A great deal has been written about GPTs and transformers more generally, both in the technical literature and in commentary of various levels of sophistication. I have read only a small portion of this. But nothing I have read indicates any interest in the nature of language or mind. That seems relegated to the GPT engine itself. And yet the product of that engine, a language model, is opaque. I believe that, if we are to move to a level of accomplishment beyond what has been exhibited to date, we must understand what that engine is doing so that we may gain control over it. We must think about the nature of language and of the mind.

That is what this working paper sets out to achieve, a beginning point, and only that. By attending to ideas by Adam Neubig, Julian Michael, and Sydney Lamb, and by extending them through the geometric semantics of Peter Gärdenfors, we can create a framework in which to understand language and mind, a framework that is commensurate with the operations of GPT-3. That framework can help us to understand what GPT-3 is doing when it constructs a language model, and thereby to gain control over that model so we can enhance and extend it.

It is in that speculative spirit that I offer the following remarks.


Abstract: GPT-3 is an AI engine that generates text in response to a prompt given to it by a human user. It does not understand the language that it produces, at least not as philosophers understand such things. And yet its output is in many cases astonishingly like human language. How is this possible? Think of the mind as a high-dimensional space of signifieds, that is, meaning-bearing elements. Correlatively, text consists of one-dimensional strings of signifiers, that is, linguistic forms. GPT-3 creates a language model by examining the distances and ordering of signifiers in a collection of text strings and computes over them so as to reverse engineer the trajectories texts take through that space. Peter Gärdenfors’ semantic geometry provides a way of thinking about the dimensionality of mental space and the multiplicity of phenomena in the world, about how mind mirrors the world. Yet artificial systems are limited by the fact that they do not have a sensorimotor system that has evolved over millions of years. They do have inherent limits.

Contents

0. Starting point and preview 1
1. Computers are strange beasts 4
2. No meaning, no how: GPT-3 as Rubicon and Waterloo, a personal view 8
3. The brain, the mind, and GPT-3: Dimensions and conceptual spaces 16
4. Gestalt switch: GPT-3 as a model of the mind 24
5. Engineered intelligence at liberty in the world 26

0. Starting point and preview

GPT-3 is based on distributional semantics. Warren Weaver had the basic idea in his 1949 memorandum, “Translation” (p. 8). Gerard Salton operationalized the idea in his work using vector semantics for document retrieval in the 1960s and 1970s (p. 9). Since then distributional semantics has developed as an empirical discipline. The last decade of work in NLP has seen remarkable, even astonishing, progress. And yet we lack a robust theoretical framework in which we can understand and explain that progress. Such a framework must also indicate the inherent limitations of distributional semantics. This document is a first attempt to outline such a framework, as such its various formulations must be seen as speculative and provisional. I offer them so that others may modify them, replace them, and move beyond them.

It started with a comment at a blog

On July 19, 2020, Tyler Cowen made a post to Marginal Evolution entitled “GPT-3, etc.” It consisted of an email from a reader who asserted, “When future AI textbooks are written, I could easily imagine them citing 2020 or 2021 as years when preliminary AGI first emerged,. This is very different than my own previous personal forecasts for AGI emerging in something like 20-50 years…” While I have my doubts about the concept of AGI – it’s too ill-defined to serve as anything other than a hook on which to hang dreams, anxieties, and fears – I think GPT-3 is worth serious consideration.

Cowen’s post has attracted 52 comments so far, more than a few of acceptable or even high quality. I made a long comment to that post. I then decided to expand that comment into a series of blog posts, say three or four, and then to collect them into a single document as a working paper. When it appeared that those three or four posts would grow to five or six I decided that I would issue two working papers. This first one would concentrate on GPT-3 and the nature of artificial intelligence, or whatever it is. The second would speculate about the future and take a quick tour of the past.

Here is a slightly revised version of the comment I made at Marginal Revolution. This paper covers the shaded material. The rest will be covered in the second paper.
Yes, GPT-3 [may] be a game changer. But to get there from here we need to rethink a lot of things. And where that's going (that is, where I think it best should go) is more than I can do in a comment.

Right now, we're doing it wrong, headed in the wrong direction. AGI, a really good one, isn't going to be what we're imagining it to be, e.g. the Star Trek computer.

Think AI as platform, not feature (Andreessen). Obvious implication, the basic computer will be an AI-as-platform. Every human will get their own as an very young child. They're grow with it; it’ll grow with them. The child will care for it as with a pet. Hence we have ethical obligations to them. As the child grows, so does the pet – the pet will likely have to migrate to other physical platforms from time to time.

Machine learning was the key breakthrough. Rodney Brooks’ Gengis, with its subsumption architecture, was a key development as well, for it was directed at robots moving about in the world. FWIW Brooks has teamed up with Gary Marcus and they think we need to add some old school symbolic computing into the mix. I think they’re right.

Machines, however, have a hard time learning the natural world as humans do. We're born primed to deal with that world with millions of years of evolutionary history behind us. Machines, alas, are a blank slate.

The native environment for computers is, of course, the computational environment. That's where to apply machine learning. Note that writing code is one of GPT-3's skills.

So, the AGI of the future, let's call it GPT-42, will be looking in two directions, toward the world of computers and toward the human world. It will be learning in both, but in different styles and to different ends. In its interaction with other artificial computational entities GPT-42 is in its native milieu. In its interaction with us, well, we'll necessarily be in the driver’s seat.

Where are we with respect to the hockey stick growth curve? For the last 3/4 quarters of a century, since the end of WWII, we've been moving horizontally, along a plateau, developing tech. GPT-3 is one signal that we've reached the toe of the next curve. But to move up the curve, as I’ve said, we have to rethink the whole shebang.

We're IN the Singularity. Here be dragons.

[Superintelligent computers emerging out of the FOOM is bullshit.]

* * * * *

ADDENDUM: A friend of mine, David Porush, has reminded me that Neal Stephenson has written of such a tutor in The Diamond Age: Or, A Young Lady's Illustrated Primer (1995). I then remembered that I have played the role of such a tutor in real life, The Freedoniad: A Tale of Epic Adventure in which Two BFFs Travel the Universe and End up in Dunkirk, New York.
While the portion of the comment to be elaborated in the next working paper is considerably longer than the portion being elaborated in this one, I do not expect that paper to be proportionately longer. This paper covered quasi-technical matters requiring fairly careful exposition. The next paper will go by more quickly and will, in sections, approach science fiction.

* * * * *

1. Computers are strange beasts – They’re obviously inanimate, and yet we communicate with them through language. The don’t fit pre-existing (19th century?) conceptual categories, and so we are prone to strange views about them.

2. No meaning, no how: GPT-3 as Rubicon and Waterloo, a personal view – Arguing from first principles it is clear that GPT-3 lacks understanding and access to meaning. And yet it produces very convincing simulacra of understanding. But common sense understanding remains elusive, as it did for old school symbolic processing. Much of common sense is deeply embedded in the physical world. GPT-3, as it currently functions is, in effect, an artificial brain in a vat.

3. The brain, the mind, and GPT-3: Dimensions and conceptual spaces – GPT-3 creates a language model by examining the distances and ordering of signifiers in a collection of text strings and computes over them so as to reverse engineer but the trajectories texts take through a high-dimensional mental space of signifieds. Peter Gärdenfors’ semantic geometry provides a way of thinking about the dimensionality of mental space and the multiplicity of phenomena in the world.

4. Gestalt switch: GPT-3 as a model of the mind – GPT-3 creates: 1) a model of a body of natural language texts, and only a model. 2) Those texts are the product of human minds. 3) Though the application of 2 to 1 we may conclude that GPT-3 is also a model of the mind, albeit a very limited one. 3 requires a Gestalt switch.

5. Engineered intelligence at liberty in the world – The “intelligence” in systems such as GPT-3 is static and reactive. To liberate and mobilize it we need to endow AI systems with mental models of the kind investigated in “old school” symbolic AI.

Friday, July 31, 2020

3. Interlude: GPT-3 as a model of the mind

Here are the key paragraphs from the previous post:
Let us notice, first of all, that language exists as strings of signifiers in the external world. In the case that interests us, those are strings of written characters that have been encoded into computer-readable form. Let us assume that the signifieds – which bear a major portion of meaning, no? – exist in some high dimensional network in mental space. This is, of course, an abstract space rather than the physical space of neurons, which is necessarily three dimensional. However many dimensions this mental space has, each signified exists at some point in that space and, as such, we can specify that point by a vector containing its value along each dimension.

What happens when one writes? Well, one produces a string of signifiers. The distance between signifiers on this string, and their ordering relative to one another, are a function of the relative distances and orientations of their associated signifieds in mental space. That’s where to look for Neubig’s isometric transform into meaning space. What GPT-3, and other NLP engines, does is to examine the distances and ordering of signifiers in the string and compute over them so as to reverse engineer the distances and orientations of the associated signifieds in high-dimensional mental space.
The purpose of this post is simply to underline the seriousness of my assertion to treat the mind as a high-dimensional space and that, therefore, we should treat the high-dimensional parameter space of GPT-3 as a model of the mind. If you aren't comfortable with the idea, well, it takes a bit of time for it to settle down (I've been here before). This post is a way of occupying some of that time.

If it’s not a model of the mind, then what IS it a model of? “The language”, you say? Where does the language come from, where does it reside? That’s right, the mind.

It is certainly not a complete model of the mind. The mind, for example, is quite fluid and iscapable of autonomous action. GPT-3 seems static and is only reactive. It cannot initiate action. Nonetheless, it is still a rich model.

I built plastic models as a kid, models of rockets, of people, and of sailing ships. None of those models completely captured the things they modeled. I was quite clear on that. I have a cousin who builds museum-class ship models from wood of various kinds, metal, cloth, paper, thread and twine (and perhaps some plastic here and there). They are much more accurate and aesthetically pleasing than the models I assembled from plastic kits as a kid. But they are still only models.

So it is with GPT-3. It is a model of the mind. We need to get used to thinking of it in those terms, dangerous as they may be. But, really, can the field get more narcissistic and hubristic than it already is?

* * * * *

This is not the first time I’ve been through this drill. I’ve been thinking about this that and the other in so-called digital humanities since 2014; call it computational criticism. These particular investigations have been using various kinds of distributional semantics – topic modeling, vector space semantics – to examine literary texts and populations of texts. They don’t think about their language models as models of the mind; they’re just, well, you know, language models, models of texts. There’s some kind membrane, some kind of barrier, that keeps us – them, me, you – from moving from these statistical models of the mind. They’re not the real thing, they’re stop gaps, approximation. Yes. And they are also models, as much models of the mind as a plastic battleship is a model of the real thing.

Why am I saying this? Like I said, to underline the seriousness of my assertion to treat the mind as a high-dimensional space. In a common formulation, the mind is what the brain does. The brain is a three-dimensional physical object.

It consists of roughly 86 billion neurons, each of which has roughly 10,000 connections with other neurons. The action at each of those synaptic junctures is mediated by upward of a 100 neurochemicals. The number of states a system can take depends on 1) the number of elements it has, 2) the number of states each element can take, and 3) the dependencies among those elements. How many states can that system assume? We don't really know. Jillions, maybe zillions, maybe jillions of zillions. A lot.

That is a state space of very high dimensionality. That state space is the mind. GPT-3 is a model of that.

* * * * *

I’ve written quite a bit about computational criticism, though nothing for formal academic publication. Here’s one paper to look at:

William Benzon, Virtual Reading: The Prosper Project Redux, Working Paper, Version 2, October 2018, 37 pp., https://www.academia.edu/34551243/Virtual_Reading_The_Prospero_Project_Redux.
Abstract: Virtual reading is proposed as a computational strategy for investigating the structure of literary texts. A computer ‘reads’ a text by moving a window N-words wide through the text from beginning to end and follows the trajectory that window traces through a high-dimensional semantic space computed for the language used in the text. That space is created by using contemporary corpus-based machine learning techniques. Virtual reading is compared and contrasted with a 40 year old proposal grounded in the symbolic computation systems of the mid-1970s. High-dimensional mathematical spaces are contrasted with the standard spatial imagery employed in literary criticism (inside and outside the text, etc.). The “manual” descriptive skills of experienced literary critics, however, are essential to virtual reading, both for purposes of calibration and adjustment of the model, and for motivating low-dimensional projection of results. Examples considered: Augustine’s Confessions, Heart of Darkness, Much Ado About Nothing, Othello, The Winter’s Tale.
* * * * *

Posts in this series are gathered under this link: Rubicon-Waterloo.

Wednesday, July 29, 2020

2. The brain, the mind, and GPT-3: Dimensions and conceptual spaces

[Edited, with a substantial addition, August 2, 2020]

The purpose of this post is to sketch a conceptual framework in which we can understand the success of language models such as GPT-3 despite the fact that they are based on nothing more than massive collections of bare naked signifiers. There’s not a signified in sight, much less any referents. I have no intention of even attempting to explain how GPT-3 works. That it does work, in an astonishing variety of cases if (certainly) not universally, is sufficient for my purposes.

First of all I present the insight that sent me down this path, a comment by Graham Neubig in an online conversation that I was not a part of. Then I set that insight in the context of and insight by Sydney Lamb (meaning resides in relations), a first-generation researcher in machine translation and computational linguistics. I think take a grounding case by Julian Michael, that of color, and suggest that it can be extended by the work of Peter Gärdenfors on conceptual spaces.

A clue: an isomorphic transform into meaning space

At the 58th Annual Meeting of the Association for Computational Linguistics Emily M. Bender and Alexander Koller delivered a paper, Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data [1], where NLU means natural language understanding. The issue is pretty much the one I laid out in my previous posts in the sections “No words, only signifiers” and “Martin Kay, ‘an ignorance model’” [2]. A lively discussion ensured online which Julian Michael has summarized and commented on in a recent blog post [3].

In that post Michael quotes a remark by Graham Neubig:
One thing from the twitter thread that it doesn’t seem made it into the paper... is the idea of how pre-training on form might learn something like an “isomorphic transform” onto meaning space. In other words, it will make it much easier to ground form to meaning with a minimal amount of grounding. There are also concrete ways to measure this, e.g. through work by Lena Voita or Dani Yogatama... This actually seems like an important point to me, and saying “training only on form cannot surface meaning,” while true, might be a little bit too harsh— something like “training on form makes it easier to surface meaning, but at least a little bit of grounding is necessary to do so” may be a bit more fair.
That’s my point of departure in this post, that notion of “an ‘isomorphic transform’ onto meaning space.” I am going to sketch a framework in which we can begin unpacking that idea. But it may take awhile to get there.

Meaning is in relations

I want to develop an idea I have from Sydney Lamb, that meaning resides in relations. The idea is grounded in the “old school” world of symbolic computation, where language is conceived as a relational network of items. The meaning of any item in the network is a function of its position in the network.

Let’s start with this simple diagram:


It represents the fact that the central nervous system (CNS) is coupled to two worlds, each external to it. To the left we have the external world. The CNS is aware of that world through various senses (vision, hearing, smell, touch, taste, and perhaps others) and we act in that world through the motor system. But the CNS is also coupled to the internal milieu, with which it shares a physical body. The net is aware of that milieu by chemical sensors indicating contents of the blood stream and of the lungs, and by sensors in the joints and muscles. And it acts in the world through control of the endocrine system and the smooth muscles. Roughly speaking the CNS guides the organism’s actions in the external world so as to preserve the integrity of the internal milieu. When that integrity is gone, the organism is dead.

Now consider this more differentiated presentation of the same facts:


I have divided the CNS into four sections: A) senses the external world, B) senses the internal milieu, D) guides action in the internal milieu, and D) guides action in the external world. I rather doubt that even a very simple animal, such as C. elegans, with 302 neurons, is so simple. But I trust my point will survive that oversimplification.

Lamb’s point is that the “meaning” or “significance” of any of those nodes – let’s not worry at the moment whether they’re physical neurons or more abstract entities – is a function of its position in the entire network, with its inputs from and outputs to the external world and the inner milieu [4]. To appreciate the full force of Lamb’s point we need to recall the diagrams typical of old school symbolic computing, such as this diagram from Brian Phillips we used in the previous post:
All of the nodes and edges have labels. Lamb’s point is that those labels exist for our convenience, they aren’t actually a part of the system itself. If we think of that network as a fragment from a human cognitive system – and I’m pretty sure that’s how Phillips thought about it, even if he could not justify it in detail (no one could, not then, not now) – then it is ultimately connected to both the external world and the inner milieu. All those labels fall away; they serve no purpose. Alas, Phillips was not building a sophisticated robot, and so those labels are necessary fictions.

But we’re interested in the full real case, a human being making their way in the world. In that case let us assume that, for one thing, the necessary diagram is WAY more complex, and that the nodes and edges do not represent individual neurons. Rather, they represent various entities that are implemented in neurons, sensations, thoughts, perceptions, and so forth. Just how such things are realized in neural structures is, of course, a matter of some importance and is being pursued by hundreds of thousands of investigators around the world. But we need not worry about that now. We’re about to fry some rather more abstract fish (if you will).

Some of those nodes will represent signifiers, to use the Saussurian terminology I used in my previous post, and some will represent signifieds. What’s the difference between a signifier and a signified? Their position in the network as a whole. That’s all. No more, no less. Now, it seems to me, we can begin thinking about Neubig’s “isomorphic transform” onto meaning space.

Let us notice, first of all, that language exists as strings of signifiers in the external world. In the case that interests us, those are strings of written characters that have been encoded into computer-readable form. Let us assume that the signifieds – which bear a major portion of meaning, no? – exist in some high dimensional network in mental space. This is, of course, an abstract space rather than the physical space of neurons, which is necessarily three dimensional. However many dimensions this mental space has, each signified exists at some point in that space and, as such, we can specify that point by a vector containing its value along each dimension.

What happens when one writes? Well, one produces a string of signifiers. The distance between signifiers on this string, and their ordering relative to one another, are a function of the relative distances and orientations of their associated signifieds in mental space. That’s where to look for Neubig’s isometric transform into meaning space. What GPT-3, and other NLP engines, does is to examine the distances and ordering of signifiers in the string and compute over them so as to reverse engineer the distances and orientations of the associated signifieds in high-dimensional mental space.
[A little reflection on that formulation makes it clear that it fails to take into account a distinction central to ‘old school’ symbolic computation, that between semantic and episodic memory. Rather than interrupt this argument with a refined formulation I have placed that in an appendix to this post: A more refined approach to meaning space. I also offer some remarks on need for a connection to the physical world in order to handle common-sense reasoning.]
Is the result perfect? Of course not – but then how do we really know? It’s not as though we’ve got a well-accepted model of human conceptual space just lying around on a shelf somewhere. GPT-3’s language model is perhaps as good as we’ve got at the moment, and we can’t even open the hood and examine it. We know its effectiveness by examining how it performs. And it performs very well.

Monday, July 27, 2020

1. No meaning, no how: GPT-3 as Rubicon and Waterloo, a personal view

I say that not merely because I am a person and, as such, I have a point of view on GPT-3, and related matters. I say because the discussion is informal, without journal-class discussion of this, that, and the others, along with the attendant burden of citation, though I will offer a few citations. More over, I’m pretty much making this up as I go along. That is to say, I am trying to figure out just what it is that I think, and see value in doing so in public.

What value, you ask? It commits me to certain ideas, if only at a certain time. It lays out a set of priors and thus serves to sharpen my ideas developments unfold and I, inevitably, reconsider.

GPT-3 represents an achievement of a high order; it deserves the attention it has received, if not the hype. We are now deep in “here be dragons” territory and we cannot go back. And yet, if we are not careful, we’ll never leave the dragons, we’ll always be wild and undisciplined. We will never actually advance; we’ll just spin faster and faster. Hence GPT-3 is both a Rubicon, the crossing of a threshold, and a potential Waterloo, a battle we cannot win.

Here’s my plan: First we take a look at history, at the origins of machine translation and symbolic AI. Then I develop a fairly standard critic of semantic models such as those used in GPT-3 which I follow with some remarks by Martin Kay, one of the Grand Old Men of computational linguistics. Then I look at the problem of common sense reasoning and conclude be looking ahead to the next post in this series in which I offer some speculations on why (and perhaps even how) these models can succeed despite their sever and fundamental short-comings.

Background: MT and Symbolic computing

It all began with a famous memo Warren Weaver wrote in 1949. Weaver was director of the Natural Sciences division of the Rockefeller Foundation from 1932 to 1955. He collaborated Claude Shannon in the publication of a book which popularized Shannon’s seminal work in information theory, The Mathematical Theory of Communication. Weaver’s 1949 memorandum, simply entitled “Translation” [1], is regarded as the catalytic document in the origin of machine translation (MT) and hence of computational linguistics (CL) and heck! why not? artificial intelligence (AI).

Let’s skip to the fifth section of Weaver’s memo, “Meaning and Context” (p. 8):
First, let us think of a way in which the problem of multiple meaning can, in principle at least, be solved. If one examines the words in a book, one at a time as through an opaque mask with a hole in it one word wide, then it is obviously impossible to determine, one at a time, the meaning of the words. “Fast” may mean “rapid”; or it may mean "motionless"; and there is no way of telling which.

But if one lengthens the slit in the opaque mask, until one can see not only the central word in question, but also say N words on either side, then if N is large enough one can unambiguously decide the meaning of the central word. The formal truth of this statement becomes clear when one mentions that the middle word of a whole article or a whole book is unambiguous if one has read the whole article or book, providing of course that the article or book is sufficiently well written to communicate at all.
It wasn’t until the 1960s and ‘70s that computer scientists would make use of this insight; Gerard Salton was the central figure and he was interested in document retrieval [2]. Salton would represent documents as a vector of words and then query a database of such representation by using a vector composed from user input. Documents were retrieved as a function of similarity between the input query vector and the stored document vector.

Work on MT went a different way. Various approaches were used, but at some relatively early point researchers were writing formal grammars of languages. In some cases these grammars were engineering conveniences while in others they were taken to represent the mental grammars of humans. In any event, that enterprise fell apart in the mid-1960s. The prospects for practical results could not justify federal funding and the government had interest in supporting purely scientific research into the nature of language.

But such research continued nonetheless, sometimes under the rubric of computational linguistics (CL) and sometimes as AI. I encountered CL in graduate school in the mid-1970s when I joined the research group of David Hays in the Linguistics Department of the State University of New York at Buffalo – I was actually enrolled as a graduate student in English; it’s complicated.

Many different semantic models were developed, but I’m not interested in anything like a review of that work, just a little taste. In particular I am interested in a general type of model was known as a semantic or cognitive network. Hays had been developing such a model for some years in conjunction with several graduate students [2]. Here’s a fragment of a network from a system developed by one of those students, Brian Phillips, to tell whether or not stories of people drowning were tragic [3]. Here’s a representation of capsize:
Notice that there are two kinds of nodes in the network, square ones and smaller round ones. The square ones represent a scene while the round ones represent individual objects or events. Thus the square node at the upper left indicates a scene with two sub-scenes – I’m just going to follow out the logic of the network without explaining it in any detail. The first one asserts that there is a boat that contains one Horatio Smith. The second one asserts that the boat overturns. And so forth through the rest of the diagram.

This network represents semantic structure. In the terminology of semiotics, it represents a network of signifieds. Though Phillips didn’t do so, it would be entirely possible to link such a semantic network with a syntactic network, and many systems of that era did so.

Such networks were symbolic in the (obvious) sense that the objects in them were considered to be symbols, not sense perceptions or motor actions nor, for that matter, neurons, whether real or artificial. The relationship between such systems and the human brain was not explored, either in theory or in experimental observation. It wasn’t an issue.

That enterprise collapsed in the mid-1980s. Why? The models had to be hand-coded, which took time. They were computationally expensive and so-called common sense reasoning proved to be endless, making the models larger and larger. (I discuss common sense below and I have many posts at New Savanna on the topic [4].)

Oh, the work didn’t stop entirely. Some researchers kept at it. But interests shifted toward machine learning techniques and toward artificial neural networks. That is the line of evolution that has, three or four decades later, resulted in systems like GPT-3, which also owe a debt to the vector semantics pioneered by Salton. Such systems build huge language models from huge databases – GPT-3 is based on 500 billion tokens [5] – and contain no explicit models of syntax or semantics anywhere, at least not that researchers can recognize.

Researchers build a system that constructs a language model (“learns” the language), but the inner workings of that model are opaque to the researchers. After all, the system built the model, not the researchers. They only built the system.

It is a strange situation.

Friday, July 24, 2020

0. The road ahead, into "the Singularity" and beyond [GPT-3]

This is the first in a series of posts where I set out a vision for the evolution of artificial intelligence beyond GPT-3 (GPT = Generative Pre-trained Transformer). As I explain in the next post in the series, “No meaning, no how”, it is both a remarkable achievement – we are now at sea in the Singularity and there is no turning back – and a remarkable temptation, hence my name for the overall series, GPT-3: Rubicon and Waterloo. No doubt some will yield to that temptation, but others are already resisting it, and have been for awhile. What will happen?

Of course I don’t know how things will unfold, but I have preferences.

The purpose of this series is to lay out those preferences. In the next section of this post I quote extensively from an article David Hays and I published in 1990, The Evolution of Cognition [1], the first in a series of essays in which we outline a view of human cultural evolution over the longue durée. Then I reprise the sketch of my (current) vision that I tucked into a comment at Marginal Revolution. I conclude with some observations of the value of being old (priors!).

Beyond AGI

In “The Evolution of Cognition” David Hays and I argued that the long-term evolution of human culture follows from the architectural foundations of thought and communication: first speech, then writing, followed by systematized calculation, and most recently, computation. In discussing the importance of the computer we remark:
One of the problems we have with the computer is deciding what kind of thing it is, and therefore what sorts of tasks are suitable to it. The computer is ontologically ambiguous. Can it think, or only calculate? Is it a brain or only a machine?

The steam locomotive, the so-called iron horse, posed a similar problem for people at Rank 3. It is obviously a mechanism and it is inherently inanimate. Yet it is capable of autonomous motion, something heretofore only within the capacity of animals and humans. So, is it animate or not? Perhaps the key to acceptance of the iron horse was the adoption of a system of thought that permits separation of autonomous motion from autonomous decision. The iron horse is fearsome only if it may, at any time, choose to leave the tracks and come after you like a charging rhinoceros. Once the system of thought had shaken down in such a way that autonomous motion did not imply the capacity for decision, people made peace with the locomotive.

The computer is similarly ambiguous. It is clearly an inanimate machine. Yet we interact with it through language; a medium heretofore restricted to communication with other people. To be sure, computer languages are very restricted, but they are languages. They have words, punctuation marks, and syntactic rules. To learn to program computers we must extend our mechanisms for natural language.

As a consequence it is easy for many people to think of computers as people. Thus Joseph Weizenbaum, with considerable dis-ease and guilt, tells of discovering that his secretary “consults” Eliza—a simple program which mimics the responses of a psychotherapist—as though she were interacting with a real person (Weizenbaum 1976). Beyond this, there are researchers who think it inevitable that computers will surpass human intelligence and some who think that, at some time, it will be possible for people to achieve a peculiar kind of immortality by “downloading” their minds to a computer. As far as we can tell such speculation has no ground in either current practice or theory. It is projective fantasy, projection made easy, perhaps inevitable, by the ontological ambiguity of the computer. We still do, and forever will, put souls into things we cannot understand, and project onto them our own hostility and sexuality, and so forth.

A game of chess between a computer program and a human master is just as profoundly silly as a race between a horse-drawn stagecoach and a train. But the silliness is hard to see at the time. At the time it seems necessary to establish a purpose for humankind by asserting that we have capacities that it does not. It is truly difficult to give up the notion that one has to add “because . . . “ to the assertion “I’m important.” But the evolution of technology will eventually invalidate any claim that follows “because.” Sooner or later we will create a technology capable of doing what, heretofore, only we could.
That is where we are now. The notion of an AGI (artificial general intelligence) that will bootstrap itself into superintelligence is fantasy; it arises because, even after three-quarters of a century, computers are still strange to us. We design, build, and operate them; but they challenge us; they’ve got bugs, they crash, they don’t come to heel when we command. We don’t know what they are. That is certainly the case with GPT-3. We’ve built it; it’s performance amazes (but puzzles and disappoints as well). And we do not understand how it works. It is almost as puzzling to us as we are to ourselves. Surely we can change that, no?

We conclude the essay with this paragraph:
We know that children can learn to program, that they enjoy doing so, and that a suitable programming environment helps them to learn (Kay 1977, Pappert 1980). Seymour Pappert argues that programming allows children to master abstract concepts at an earlier age. In general it seems obvious to us that a generation of 20-year-olds who have been programming computers since they were 4 or 5 years old are going to think differently than we do. Most of what they have learned they will have learned from us. But they will have learned it in a different way. Their ontology will be different from ours. Concepts which tax our abilities may be routine for them, just as the calculus, which taxed the abilities of Leibniz and Newton, is routine for us. These children will have learned to learn Rank 4 concepts.
Frankly, I think we are a behind the curve on this one. Had Hays and I hazarded to predict the advance of computing into the lives of children –“The child is Father to the man”, as Wordsworth observed – I fear we would be disappointed by the current situation.

Yes, relatively young programmers have done remarkable things and Silicon Valley teems with young virtuosi. It is not the virtuosi I’m concerned about. It is the average, which is too low, by far.

Oddly enough, the current pandemic may help raise that average, though only marginally. With at-home schooling looming in the future, school districts are beginning to buy laptop machines for children whose families cannot afford them. For without those machines, those children will not be able to participate in the only education available to them. No doubt most of the instruction they receive through those machines will train them to be only passive consumers of computation, as most of us are and have been conditioned to be.

But some of them surely will be curious. They’ll take a look under the virtual hood – though some of them will undoubtedly open up the physical machine itself (not that there’s much to see, with so much action integrated on a single chip) – and begin tinkering around. And before you know, they’ll do interesting things and Peter Thiel is going to be handing out more of those $100,000 fellowships [2] to teens living in institutionally impoverished neighborhoods plagued by substandard infrastructure.

We’ll see.

The road ahead

On July 19, 2020, Tyler Cowen made a post to Marginal Evolution entitled “GPT-3, etc.” It consisted of an email from a reader who asserted, “When future AI textbooks are written, I could easily imagine them citing 2020 or 2021 as years when preliminary AGI first emerged,. This is very different than my own previous personal forecasts for AGI emerging in something like 20-50 years…” As I’ve already indicated, I have my doubts about the concept of AGI.