Showing posts with label GPT-3. Show all posts
Showing posts with label GPT-3. Show all posts

Friday, October 20, 2023

Comparison of how humans continue a Pygmalion prompt vs. GPT-3.5 and GPT-4

There are more tweets in the stream, but check out the paper linked in the first tweet.

Saturday, October 14, 2023

GPT-3 has rendered Alan Turing 1950 obsolete. We need to enter a world he could not have imagined.

And by that I do not mean that GPT-3 has made the so-called Turing Test – which Turing called the Imitation Game (Computing Machinery and Intelligence, 1950) – obsolete. It’s my understanding that philosophers have long since found it to be useless; Weizenbaum’s ELIZA but the kibosh on it back in the mid-1960s. No, I mean something a bit different, and more difficult to conceptualize.

In that 1950 paper Turing gave expression to an old idea, a dream, that humankind should create an artificial human. Mary Shelley’s Dr. Frankenstein pursued one version of that idea. A much older version showed up in the title of a book by Norbert Weiner, God & Golem, Inc.: A Comment on Certain Points Where Cybernetics Impinges on Religion (1964). The imitation game was a device Turing proposed for determining whether or not a convincing simulacrum of humanity had been created. What GPT-3 has rendered obsolete is the intellectual drive to create that artefact. Paradoxically, it rendered that impetus obsolete by the very fact that it more convincingly passed or, if you will, outstripped, the Turing Test than any previous system had done.

Artificial intelligence originated as an enterprise devolved from the dream of creating, if not an artificial human, at least an artificial (human) mind. As such it issued prediction after prediction about when a computer would equal or surpass humans in some cognitive activity, chess most prominently, but everydamnthing else as well. The work that eventuated in GPT-3 is in that lineage, albeit based on technology quite different from that imagined by those worthies gathered at Dartmouth in that 1956 meeting.

But GPT-3 shocked everyone with its facility, even its creators (perhaps especially them, and they had intimations of things to come in GPT-2). It’s in that moment of shock that Turing 1950 was rendered obsolete. All that intellectual efforted had produced something that was at one and the same time: 1) unexpected, 2) too damn apparently human, and 3) thereby a fulfillment of the 1950 version of that ancient dream. It’s the unexpectedness of the success that’s so confounding.

That 1950 version of the dream of artificial humans was grounded in a certain system of thought – to use the phrase that Hays and I have employed in our theory of cultural ranks. That system of thought was capable to conceiving artificial neural nets, and of conceiving the transformer architecture. But it is not capable of understanding the language models produced by transformers and other artificial neural nets for that matter, but it’s GPT-3 that shocked everyone. And that is why I say it rendered Turing 1950 obsolete. We’re facing a different world now and we have to adjust our hopes and fears, our dreams and nightmares, accordingly.

That’s easier said than done. All the chaos that’s been occasioned by ChatGPT – but other systems too – is a manifestation of that process of adjustment, of profound cultural change. There is no guarantee of successful adaptation, nor, for that matter, is there some singular ideal form of successful adaptation.

One of the chief indices of our inability to comprehend these LLMs is the persistence of the idea that they are prediction machines – “stochastic parrots”, “autocomplete on steroids”. Prediction is a device, but it’s not the goal. The goal is a model, and the model is about the entanglement of meaning among words and texts, something I’ve discussed, Entanglement and intuition about words and meaning. The observation that these models are opaque, unintelligible, black boxes, is another index of the insufficiency of our current system of thought, of the system of thought that gave rise to the models in the first place. Yes, they are unintelligible, but that’s not a property inherent in the models themselves, like, e.g. the number of layers or parameters they have. It’s a property of the relationship between the models and some system of thought. The internal combustion engine would have been unintelligible to Aristotle or Archimedes, but that’s not because those engines are inherently unintelligible. Rather, Aristotle and Archimedes didn’t have a system of thought in which internal combustion engines could be understood.

Now, it’s one thing to observe that the models are unintelligible. It’s quite something else to believe/fear that that unintelligibility is inherent in them and, correlatively, in the human mind as well. It’s clear to me that some Doomers actively cultivate the notion that these models are unintelligible.

This opacity, this unintelligibility is not all of a sudden, a new phenomenon. It’s been with us ever since Frank Rosenblatt conceived of perceptrons, the first artificial neural nets. But it was not problematic in those older systems. It is the unexpected mimetic capacity of GPT-3 and later LLMs that has rendered that opacity problematic. It will remain problematic as long as we (insist upon continuing to) remain tethered to the system of thought which gave birth to these wonderful devices.

It's time to move on.

Thursday, August 24, 2023

Is this the beginning of the end for LLMS [as the royal road to AGI, whatever that is]?

It’s hard to tell, but it sure is...shall we say...interesting.

Back in the summer of 2020 when GPT-3 was unveiled I wrote a working paper, GPT-3: Waterloo or Rubicon? Here be Dragons. My objective was to convince myself that the underlying technology wasn’t just some weird statistical fluke, that there was in fact something going on of substantial interest and value. To my mind, I succeeded in that. But I was skeptical as well.

Here's what I put on the first page of that working paper, even before the abstract:

GPT-3 is a significant achievement.

But I fear the community that has created it may, like other communities have done before – machine translation in the mid-1960s, symbolic computing in the mid-1980s, triumphantly walk over the edge of a cliff and find itself standing proudly in mid-air.

This is not necessary and certainly not inevitable.

A great deal has been written about GPTs and transformers more generally, both in the technical literature and in commentary of various levels of sophistication. I have read only a small portion of this. But nothing I have read indicates any interest in the nature of language or mind. Interest seems relegated to the GPT engine itself. And yet the product of that engine, a language model, is opaque. I believe that, if we are to move to a level of accomplishment beyond what has been exhibited to date, we must understand what that engine is doing so that we may gain control over it. We must think about the nature of language and of the mind.

I didn’t expect that anyone with any influence in these matters would pay any attention to me – though one can always hope – but that’s no reason not to write.

That was 2020 and GPT-3. Two years later ChatGPT was launched to great acclaim, and justly so. I certainly spent a great deal of time playing with, investigating it, and writing about it. But I didn’t forget my cautionary remarks from 2020.

Now we’re hearing rumblings that things aren’t working out so well. Back on August 12 the ever skeptical Gary Marcus posted, What if Generative AI turned out to be a Dud? Some possible economic and geopolitical implications. His first two paragraphs:

With the possible exception of the quick to rise and quick to fall alleged room-temperature superconductor LK-99, few things I have ever seen have been more hyped than generative AI. Valuations for many companies are in the billions, coverage in the news is literally constant; it’s all anyone can talk about from Silicon Valley to Washington DC to Geneva.

But, to begin with, the revenue isn’t there yet, and might never come. The valuations anticipate trillion dollar markets, but the actual current revenues from generative AI are rumored to be in the hundreds of millions. Those revenues genuinely could grow by 1000x, but that’s mighty speculative. We shouldn’t simply assume it.

And his last:

If hallucinations aren’t fixable, generative AI probably isn’t going to make a trillion dollars a year. And if it probably isn’t going to make a trillion dollars a year, it probably isn’t going to have the impact people seem to be expecting. And if it isn’t going to have that impact, maybe we should not be building our world around the premise that it is.

FWIW, I believe, and have been saying time and again, that hallucinations seem to me to be inherent in the technology. They aren’t fixable.

Now, yesterday, Ted Gioia, a culture critic with an interest in technology and experience in business, has posted, Ugly Numbers from Microsoft and ChatGPT Reveal that AI Demand is Already Shrinking. Where Marcus has a professional interest in AI technology and has intellectual skin the tech game, Gioia is just a sophisticated and interested observer. Near the end of his post, after many links to unfavorable stories, Gioia observes:

... we can see that the real tech story of 2023 is NOT how AI made everything great. Instead this will be remembered as the year when huge corporations unleashed a half-baked and dangerous technology on a skeptical public—and consumers pushed back.

Here’s what we now know about AI:

  • Consumer demand is low, and already appears to be shrinking.
  • Skepticism and suspicion are pervasive among the public.
  • Even the companies using AI typically try to hide that fact—because they’re aware of the backlash.
  • The areas where AI has been implemented make clear how poorly it performs.
  • AI potentially creates a situation where millions of people can be fired and replaced with bots—so a few people at the top continue to promote it despite all these warning signs.
  • But even these true believers now face huge legal, regulatory, and attitudinal obstacles
  • In the meantime, cheaters and criminals are taking full advantage of AI as a tool of deception.

Marcus has just updated his earlier post with a followup: The Rise and Fall of ChatGPT?

The situation is very volatile. I certainly don’t know how to predict how things are going to unfold. In the long run, I remain convinced that if we are to move to a level of accomplishment beyond what has been exhibited to date, we must understand what these engines are doing so that we may gain control over them. We must think about the nature of language and of the mind.

Stay tuned.

Tuesday, April 4, 2023

Poetry and the digital simulacrum of mind [poetry as digital touchstone]

I'm bumping this to the top of the queue on general principle.

* * * * *

Carmine Starnino, Poetry & Digital personhood, The New Criterion, April 2022.

Starnino starts out by talking about Racter, a well-known computer generator of poem simulacra from the 1980s, and then moves on to GPT-3: “...GPT-3 isn’t a better Racter. It’s a godlike Racter.” Then mixes in a little of this and that and observes:

The Turing Test, after all, has shown that readers have a weakness for rhetoric, grand gestures, and feelingful murk—all of which algorithms easily mimic. If this is what we mean when we say AI will one day rival human poets, then it will surely win, and indeed may already have.

Yes! Our willingness to read meaning to any quasi-intelligible lump of language makes it easy for computers to crank out simulacra of poetic profundity. That’s a low bar to cross.

Starnino goes on:

But there’s another kind of poetry AI will have to beat—poetry as an art of brilliant accuracies, of reality re-described in ways that bind sound to perception. And here AI’s deficiencies are brutally exposed. Because to compete at this imitation game, a machine has to show that, by micro-adjustments of effect, it can draw our senses to the highest pitch of expression. It will need to be able to match Les Murray’s depiction of beans as “minute green dolphins at suck,” or Peter van Toorn’s realization that flying dragonflies have “a great rattle of rice in their wings,” or Elizabeth Bishop’s noting a fish’s “coarse white flesh/ packed in like feathers,” or how, for Seamus Heaney, love was “like a tinsmith’s scoop/ sunk past its gleam/ in the meal-bin.” To play at this level, a machine has to imbue words with the most intimate associations and, turning inward, confess hardships, regret irrevocable choices, ponder its ultimate demise. It has to hit the same mark Robert Frost does at the end of his sonnet “Design,” when, watching a spider readying itself to eat a moth, he asks what the grim scene reveals about nature—and if there is any moral code to such predation. “What but design of darkness to appall?—/ If design govern in a thing so small.” The word “appall” here logs the shock perfectly. In its French root, the word means to make white. It also contains “pall”: a sheet, laid over a coffin, usually of white linen. Thus Frost’s diction hones our cognition, schooling us to see the world in a fresh way.

Yes.

Starnino goes on to remind us of Eliza:

Released in 1966 by the MIT professor Joseph Weizenbaum, ELIZA was the world’s first chatbot. Designed to impersonate a therapist, it would reflect back a user’s statements with open-ended questions and prepared responses (“My mother never loved me” would trigger “please go on” or “tell me more.”) Weizenbaum’s goal was to explore a computer’s capacity for conversation. Instead, he was alarmed by how completely users were taken in by ELIZA’s shallow repartee; his own secretary once insisted he leave the room so she could talk to the program in private. Credulity even extended to graduate students who had watched him build ELIZA from scratch. Sherry Turkle, a social scientist and Weizenbaum’s colleague, called it “the Eliza effect,” which she defined as “human complicity in a digital fantasy.” We can see this effect in the love-struck language Racter’s programmers used to describe the moment their creation came to life.

A bit later he notes:

It’s no coincidence that each time a new threshold is smashed, poetry is soon offered up as evidence of the breakthrough. The most profound exercise of full human consciousness, poetry has long been coveted as a benchmark for silicon-based minds, the ultimate proof of concept. Its principles were not only present at the founding of artificial intelligence as a field—the 1956 conference that set out to design machines able to “use language, form abstractions and concepts”—but every step in eroding the line between robots and people has been marked by a poetry generator. When the famed futurist Ray Kurzweil wanted to sell the public on the idea of a thinking machine in the late 1980s, he began by inventing a “Cybernetic Poet.” In fact, you can even argue that the pursuit of machine poetry has driven entire sectors of AI, helping push the limits of what language models can now do.

Hmmmm. I can’t help but think that, to the extent that that is true, my 1970s work using computational semantics to analyze a Shakespeare sonnet* deserves a re-reading, and that despite the fact that it is firmly entrenched in the era of symbolic computing.

Citing AI’s recent successes in both GO and Chess, Starnino offers and interesting argument. He notes that when Deep Blue beat Kasparov in 1997 “the IBM supercomputer appeared capable of counterintuitive thought with a baffling move that left Kasparov profoundly unnerved.” Similarly, when AlphaGo be Lee Sedol in Go in 2016 it did so with “a move that so stunned Sedol with its strangeness, he needed fifteen minutes to recover.” AI didn’t beat us in those games by learning to think about them in a human way. Rather:

It beat us because it learned to think in an entirely inhuman way. The scale of AI’s processing power—able to mull millions of strategies and pit itself against those strategies millions of times—found bizarre but superior solutions that centuries of flesh-and-blood play never considered, solutions so removed from normal reasoning as to be alien.

However, poetry

is inexorably linked to how humans think—a kind of undeluded self-questioning that, as T. S. Eliot wrote, helps us become “a little more aware of the deeper, unnamed feelings which form the substratum of our being.” It’s also tied to the need to think this way. A poem’s mental force derives from the set of intentions driving it, intentions that push poets into action. But when GPT-3 gets the call to write a poem, it doesn’t know it’s writing “poetry,” or what “writing” even is. That last part is anything but trivial. Style is a sentient act: you strive for it. My point is that a computer will never replicate what poets do unless it can also replicate why they do it.

There’s more at the link.

* William Benzon, Cognitive Networks and Literary Semantics, MLN 91: 1976, 952-982, https://www.academia.edu/235111/Cognitive_Networks_and_Literary_Semantics.

William Benzon, Lust in Action: An Abstraction, Language and Style 14, 1981, 251-270, https://www.academia.edu/7931834/Lust_in_Action_An_Abstraction.

Saturday, March 4, 2023

My recent working papers on mind and machine [Someone's in the kitchen]

Things are beginning to fall in place. I’m getting a feel for ChatGPT, and by implication, for deep learning. And by “feel” I mean just that, a feel, a feeling for, intuition. I’m beginning to get a sense of what’s going on.

On this I’m a Piagetian. He argued that learning involves the interaction between two ‘movements of the mind’ if you will. In accommodation you change your mind to fit the phenomena you’re learning. That’s the learning part. But in order to do that, you have to figure out how to assimilate the phenomena to things you already know.

I already know quite a bit about “classical” symbolic approaches to computer modeling of language and the mind. One of my earliest publications was a cognitive network model of Shakespeare’s Sonnet 129, Cognitive Networks and Literary Semantics (1976), and I completed a dissertation on the subject two years after that. I refined the work I’d done on Sonnet 129 and offered some remarks – a chapter actually – on the long-scale development of cognitive structures in human history.

A decade later David Hays and I published two papers about the brain. One of them, Metaphor, Recognition, and Neural Process (1987) took a cue from work Karl Pribram had done earlier about holographic processing in the brain. Pretty much the same mathematics would turn up in Yann LeCun’s pioneering work on convolutional neural networks a couple years later, though I didn’t learn about it until only a couple of years ago (my interests were elsewhere). In the other paper, Principles and Development of Natural Intelligence (1988) – the title is a shot across the bow of artificial intelligence, Hays and I read a wide range of material in neuroscience, cognitive, developmental, and comparative psychology, evolution and came up with five principles underlying intelligent behavior in humans. Then, around the end of the previous century and into the first decade of this one, I had quite a bit of correspondence with Walter Freeman, who had once been a student of Pribram’s and was a pioneer in the complex dynamics of the brain. That informed by book on music, Beethoven’s Anvil (2001). A bit later I took Freeman’s dynamics, crossed it with a symbolic network model of Sydney’s Lamb’s and wrote up some notes on symbolic networks over attractor basins in the cerebral cortex.

That takes me into the early years of this century, though I didn’t post those notes until 2011. By that time things began jumping off in digital humanities so I busied myself with topic models and such. Then GPT-3 was released in 2020, forcing me to think seriously about deep learning. I didn’t have direct access to it, though I got a bit of indirect access through Phil Mohun, but I read a lot about it.

That brings us to these working papers, which I present along with their abstracts, but without other comment. But I need to point out one final thing. This whole ‘journey’ – to use a popular cliché – began with my interest in Coleridge’s “Kubla Khan,” the subject of my 1972 MA Thesis, THE ARTICULATED VISION: Coleridge's “Kubla Khan.” I published a considerably revised version of that reading in 1985. One of these papers revisits that subject in the context of deep learning and complex dynamics. I expect to return to that topic at some time in the future, though I do not know when. I have other work to do before that. I’m still working with and thinking about ChatGPT.

* * * * *

GPT-3: Waterloo or Rubicon? Here be Dragons, August 5, 2020 (Version 4.1 is the current version, May 7, 2022).

Abstract: GPT-3 is an AI engine that generates text in response to a prompt given to it by a human user. It does not understand the language that it produces, at least not as philosophers understand such things. And yet its output is in many cases astonishingly like human language. How is this possible? Think of the mind as a high-dimensional space of signifieds, that is, meaning-bearing elements. Correlatively, text consists of one-dimensional strings of signifiers, that is, linguistic forms. GPT-3 creates a language model by examining the distances and ordering of signifiers in a collection of text strings and computes over them so as to reverse engineer the trajectories texts take through that space. Peter Gärdenfors’ semantic geometry provides a way of thinking about the dimensionality of mental space and the multiplicity of phenomena in the world, about how mind mirrors the world. Yet artificial systems are limited by the fact that they do not have a sensorimotor system that has evolved over millions of years. They do have inherent limits.

Direct Brain-to-Brain Thought Transfer A High Tech Fantasy that Won't Work, September 17, 2020.

Abstract: Various thinkers (Rodolfo Llinás, Christof Koch, and Elon Musk) have proposed that, in the future, it would be possible to link two or more human brains directly together so that people could communicate without the need for language or any other conventional means of communication. These proposals fail to provide a means by which a brain can determine whether or not a neural impulse is endogenous or exogenous. That failure makes communication impossible. Confusion would the more likely result of such linkage. Moreover, in providing a rationale for his proposal, Musk assumes a mistaken view of how language works, a view cognitive linguists call the conduit metaphor. Finally, all these thinkers assume that we know what thoughts are in neural terms. We don’t.

To Model the Mind: Speculative Engineering as Philosophy, April 7, 2022.

Abstract: Are brains computers? Some say yes, some say no. Does it matter? Ideas about computing have certainly proven fruitful in understanding how brains give rise to minds. That’s what this paper is about. The central section is a review of Grace Lindsey’s wonderful book Models of the Mind: How Physics, Engineering, and Mathematics Have Shaped Our Understanding of the Brain (2021). I precede it with a bit of philosophy and follow it with brief notices about five books, each proposing computationally inspired models of the mind.

Symbols and Nets: Calculating Meaning in "Kubla Khan", May 11, 2022.

Abstract: This is a dialog between a Naturalist Literary Critic and a Sympathetic Techno-Wizard about the interaction of symbols and neural nets in understanding "Kubla Khan," which has an extraordinary structure. Each of two parts is like a matryoshka doll nested three deep, with the last line of the first part being repeated in the middle of the second. They start talking about traditional symbol processing, with addressable memory, and nested loops, and end up talking about a pair of interlinked neural nets where one (language forms) is used to index the other (meaning).

Relational Nets Over Attractors, A Primer: Part 1, Design for a Mind, July 13, 2022.

Abstract: Miriam Yevick’s 1975 holographic logic suggests we need both symbols and networks to model the mind. I explore that premise by adapting Sydney Lamb’s relational network notation to represent a logical structure over basins of attraction in a collection of attractor landscapes, each belonging to a different neurofunctional area (NFA) of the cortex. Peter Gärdenfors provides the idea of a conceptual space, a low dimensional projection of the high-dimensional phase space of a NFA. Vygotsky’s account of language acquisition and internalization is used to show how the mind is indexed. We then define a MIND as a relational network of logic gates over the attractor landscape of a neural network loosely partitioned into many NFAs. An INDEXED MIND consists of a GENERAL network and an INDEXING network adjacent to and recursively linked to it. A NATURAL MIND is one where the substrate is the nervous system of a living animal. An ARTIFICIAL MIND is one where the substrate is inanimate matter engineered by humans to be a mind; it becomes AUTONOMOUS when it is able to purchase its compute with services rendered.

Discursive Competence in ChatGPT, Part 1: Talking with Dragons, January 5, 2022 (Version 2, January 11, 2023).

Abstract: Noam Chomsky’s idea of linguistic competence suggests a new approach to understanding how LLMs work. This approach requires careful analysis of text. Such analysis indicates that ChatGPT has explicit control over sophisticated discourse skills: 1) It possesses the capacity to specify high-level structures that regulate the organization of language strings into specific patterns: e.g. conversational turn-taking, story frames, film interpretation, and metalingual definition of abstract concepts. 2) It is capable of analogical reasoning in the interpretation of films and stories, such as Spielberg’s Jaws and A.I., and Tezuka’s Astro Boy stories. It must establish an analogy between some abstract interpretive theory (e.g. the ideas of Rene Girard) and people and events in a story. 3) It has some understanding of abstract concepts such as justice and charity. Such concepts can be defined over concepts that exhibit them (metalingual definition). ChatGPT recognizes suitable stories and can revise them. 4) ChatGPT can adjust its level of discourse to accommodate children of various ages. Finally, much of ChatGPT’s discourse seems formulaic in a way similar to what Parry/Lord found in oral epic.

ChatGPT intimates a tantalizing future; its core LLM is organized on multiple levels; and it has broken the idea of thinking. January 24, 2023 (Version 3, 2023).

Abstract: I make three arguments. There is a philosophical argument that the behavior of ChatGPT is so sophisticated that the ordinary concept of thinking is no longer useful in distinguishing between human behavior and the behavior of advanced AI. We don’t have deep and explicit understanding about what either humans or advanced AI systems are doing. The other argument is about ChatGPT’s behavior. As a result of examining its output in a systematic way, short stories in particular, I have concluded that its operation is organized on at least two levels: 1) the parameters and layers of the LLN, and 2) higher level grammars, if you will, that are implemented in those parameters and layers. This is analogous to the way that high-level programming languages are implemented in assembly code. 3) Consequently, it turns out that aspects of symbolic computation are latent in LLMs. An appendix gives examples of how a story grammar is organized into frames, slots, and fillers.

ChatGPT tells stories, and a note about reverse engineering: A Working Paper, March 3, 2023.

Abstract: I examine a set of stories that are organized on three levels: 1) the entire story trajectory, 2) segments within the trajectory, and 3) sentences within individual segments. I conjecture that the probability distribution from which ChatGPT draws next tokens follows a hierarchy nested according to those three levels and that is encoded in the weights off ChatGPT’s parameters. I arrived at this conjecture to account for the results of experiments in which ChatGPT is given a prompt containing a story along with instructions to create a new story based on that story but changing a key character: the protagonist or the antagonist. That one change then ripples through the rest of the story. The pattern of differences between the old and the new story indicates how ChatGPT maintains story coherence. The nature and extent of the differences between the original story and the new one depends roughly on the degree of difference between the key character and the one substituted for it. I conclude with a methodological coda: ChatGPT’s behavior must be described and analyzed on three levels: 1) The experiments exhibit surface level behavior. 2) The conjecture is about a middle level that contains the nested hierarchy of probability distributions. 3) The transformer virtual machine is the bottom level.

Thursday, December 29, 2022

Thoughts on the implications of GPT-3, two years ago and NOW [here be dragons, we're swimming, flying and talking with them]

When GPT-3 first came out, I registered my first reactions in a comment at Marginal Revolution, which I've appended immediately below the picture of Gojochan and Sparkychan. I'm currently completing a working paper about my interaction with ChatGPT. That will end with an appendix in which I repeat my remarks from two years ago and append some new ones. I've appended those after the comment to Marginal Revolution.

* * * * *


A bit revised from a comment I made at Marginal Revolution:

Yes, GPT-3 [may] be a game changer. But to get there from here we need to rethink a lot of things. And where that's going (that is, where I think it best should go) is more than I can do in a comment.

Right now, we're doing it wrong, headed in the wrong direction. AGI, a really good one, isn't going to be what we're imagining it to be, e.g. the Star Trek computer.

Think AI as platform, not feature (Andreessen). Obvious implication, the basic computer will be an AI-as-platform. Every human will get their own as an very young child. They're grow with it; it'll grow with them. The child will care for it as with a pet. Hence we have ethical obligations to them. As the child grows, so does the pet – the pet will likely have to migrate to other physical platforms from time to time.

Machine learning was the key breakthrough. Rodney Brooks' Gengis, with its subsumption architecture, was a key development as well, for it was directed at robots moving about in the world. FWIW Brooks has teamed up with Gary Marcus and they think we need to add some old school symbolic computing into the mix. I think they're right.

Machines, however, have a hard time learning the natural world as humans do. We're born primed to deal with that world with millions of years of evolutionary history behind us. Machines, alas, are a blank slate.

The native environment for computers is, of course, the computational environment. That's where to apply machine learning. Note that writing code is one of GPT-3's skills.

So, the AGI of the future, let's call it GPT-42, will be looking in two directions, toward the world of computers and toward the human world. It will be learning in both, but in different styles and to different ends. In its interaction with other artificial computational entities GPT-42 is in its native milieu. In its interaction with us, well, we'll necessarily be in the driver's seat.

Where are we with respect to the hockey stick growth curve? For the last 3/4 quarters of a century, since the end of WWII, we've been moving horizontally, along a plateau, developing tech. GPT-3 is one signal that we've reached the toe of the next curve. But to move up the curve, as I've said, we have to rethink the whole shebang.

We're IN the Singularity. Here be dragons.

[Superintelligent computers emerging out of the FOOM is bullshit.]

* * * * *

ADDENDUM: A friend of mine, David Porush, has reminded me that Neal Stephenson has written of such a tutor in The Diamond Age: Or, A Young Lady's Illustrated Primer (1995). I then remembered that I have played the role of such a tutor in real life, The Freedoniad: A Tale of Epic Adventure in which Two BFFs Travel the Universe and End up in Dunkirk, New York.

* * * * *

To the future and beyond!

I stand by those remarks from two years ago, but I want to comment on four things: 1) AI alignment, 2) the need for symbolic computing, 3) the need for new kinds of hardware, and 4) a future world in which humans and AIs interact freely.

Considerable effort has gone into tuning ChatGPT so that it won’t say things that are offensive (e.g. racial slurs) or give out dangerous information (e.g. how to hotwire cars). These efforts have not been entirely successful. This is one aspect of what is now being called “AI alignment.” In the extreme the field of AI alignment is oriented toward the possibility – which some see as a certainty – that in the future (somewhere between, say, 30 and 130 years) rogue AIs will wage a successful battle against humankind.[1] I don’t think that fear is very creditable, but, as the rollout of ChatGPT makes abundantly clear, AIs built on deep learning are unpredictable and even, in some measure, uncontrollable.

I think the problem is inherent in deep learning technology. Its job is to fit a model to, in the case of ChatGPT, an extremely large corpus of writing, much of the internet. That corpus, in turn, is ultimately about the world. The world is vast, irregular, and messy. That messiness is amplified by the messiness inherent in the human brain/mind, which did, after all, evolve to fit that world. Any AI engine capable of capturing a significant portion of the order inherent in our collective writing about the world has no choice but to encounter and incorporate some of the disorder and clutter into its model as well.

I regard such Foundation models[2], as they have come to be called, as wilderness preserves, digital wilderness. They contain what digital humanist Ted Underwood calls the latent space of culture. He says:

The immediate value of these models is often not to mimic individual language understanding, but to represent specific cultural practices (like styles or expository templates) so they can be studied and creatively remixed. This may be disappointing for disciplines that aspire to model general intelligence. But for historians and artists, cultural specificity is not disappointing. Intelligence only starts to interest us after it mixes with time to become a biased, limited pattern of collective life. Models of culture are exactly what we need.

In his penultimate paragraph Underwood notes:

I have suggested that approaching neural models as models of culture rather than intelligence or individual language use gives us even more reason to worry. But it also gives us more reason to hope. It is not entirely clear what we plan to gain by modeling intelligence, since we already have more than seven billion intelligences on the planet. By contrast, it’s easy to see how exploring spaces of possibility implied by the human past could support a more reflective and more adventurous approach to our future. I can imagine a world where generative models of culture are used grotesquely or locked down as IP for Netflix. But I can also imagine a world where fan communities use them to remix plot tropes and gender norms, making “mass culture” a more self-conscious, various, and participatory phenomenon than the twentieth century usually allowed it to become.

These digital wildness regions thus represent opportunities for discovery and elaboration. Alignment is simply one aspect of that process.

And by alignment I mean more than aligning the AI’s values with human values; I mean aligning its conceptual structure as well. That’s where “old school” symbolic computing enters the picture, especially language. Language  – not the mere word forms available in digital corpora, but word forms plus semantics and syntactic affordances –  is one of the chief ‘tools’ through which young humans are acculturated and through which human communities maintain their beliefs and practices. The full powers of language, as treated by classical symbolic systems, will be essential for “domesticating” the digital wilderness and developing it for human use.

However, this presents technical problems, problems I cannot go into here in any detail.[4] The basic issue is that symbolic computing involves one strategy for implementing, call it cogitation, in a physical system while the neural computing underlying deep learning requires a different physical implementation. These approaches are incompatible. While one can “bolt” a symbolic system onto a neural computing system, that strikes me as no more than an interim solution. It will get us started, indeed the work has already begun.[5]

What we want, though, is for the symbolic system to arise from the neural system, organically, as it does in humans.[6] This may well call for fundamentally new physical platforms for computing, platforms based on “neuromorphic” components that are “grown,” as Geoffrey Hinton has recently remarked.[7] That technology will give us a whole new world, one where humans, AIs and robots interact freely with one another, but will have communities of their own as well. We know that dogs co-evolved with humans over tens of thousand of years. These miraculous new devices will co-evolve with us over the coming decades and centuries.

Let us end with Miranda’s words from Shakespeare’s The Tempest:

“Oh wonder!
How many goodly creatures are there here!
How beauteous mankind is! Oh brave new world,
That has such [devices] in’t.”

* * * * *

[1] The virtual center of this belief is a website called LessWrong, which has extensive discussion of this issue going back well over a decade. Here it is, https://www.lesswrong.com/.

[2] Foundation models, Wikipedia, https://en.wikipedia.org/wiki/Foundation_models.

[3] Ted Underwood, Mapping the latent spaces of culture, The Stone and the Shell, Oct. 21, 2021, https://tedunderwood.com/2021/10/21/latent-spaces-of-culture/.

[4] I discuss this issue in this blog post, Physical constraints on computing, process and memory, Part 1 [LeCun], New Savanna, July 24, 2022, https://new-savanna.blogspot.com/2022/07/physical-constraints-on-computing.html.

[5] Consult the Wikipedia entry, Neuro-symbolic AI, for some pointers, https://en.wikipedia.org/wiki/Neuro-symbolic_AI.

[6] I discuss this in a recent working paper, Relational Nets Over Attractors, A Primer: Part 1, Design for a Mind, Version 2, Working Paper, July 13, 2022, pp. 76, https://www.academia.edu/81911617/Relational_Nets_Over_Attractors_A_Primer_Part_1_Design_for_a_Mind.

[7] Tiernan Ray, We will see a completely new type of computer, says AI pioneer Geoff Hinton, ZDNET, December 1, 2022, https://www.zdnet.com/article/we-will-see-a-completely-new-type-of-computer-says-ai-pioneer-geoff-hinton-mortal-computation/#ftag=COS-05-10aaa0j.

Sunday, December 4, 2022

How ChatGPT parodied “Kubla Khan” and pwned DJT45 at the same time

Prelude

Man, if you can’t have more fun than a barrel of monkeys while playing with ChatGPT, then you aren’t human. Don’t get me wrong, I don’t think ChatGPT means that we’re well on the way to AGI, much less to being ruled by Rogue AIs that make Atilla the Hun look like Shirley Temple. Nothing like that.

I don’t know what kind of thing ChatGPT is. It surely isn’t your cranky Uncle Roger’s AI, nor Uncle Marvin’s either. And it isn’t a million drunken chimps banging away on a gaggle of antique Underwoods either. It’s something else. The AGI-is-coming people and the only-stats-on-steroids people are making two versions of the same mistake. They’re trying assimilate ChatGPT, and its kith and kin, to things we understand. We understand stats on steroids, and science fiction has given us examples homicidal computers.

But ChatGPT isn’t either one of those. What is it? And what, by the way, are we?

Don’t you think we should make a serious attempt to find out?

Let’s write a poem

So...poetry has long been a standard category of things AI does. I decided to give it a whirl. Rather than asking ChatGPT to write a poem based on some more or less arbitrary word contraption I dropped on it, I decided to ask it for a parody of a well-known poem. Here’s the result:

Me: Would you write a parody of “Kubla Khan”.

ChatGPT: Sure, here's a parody of "Kubla Khan" that I came up with:

In Xanadu did Kubla Khan
A stately pleasure dome decree
Where Alph, the sacred river, ran
Through caverns measureless to man
Down to a sunless sea

But alas, poor Kubla Khan
His pleasure dome was but a scam
He hired some guy named Dave
To build it, but it was a cave
Now Kubla's left with nothing but shame

He tried to impress his guests
With marble walls and golden towers
But the whole thing was a mess
It collapsed in just a few hours

Now Kubla's left to face the test
Of explaining to his friends
Why his fancy palace came to naught
He'll have to make amends
Or be forever ridiculed and caught
In the web of his own foolish ends.

Now, I suppose you can assert that it’s not much of a parody. It repeats the first stanza verbatim, doesn’t capture “Kubla Khan's” form at all, and pretty much ignores most of the poem: What of the fountain and the damsel with a dulcimer? They’re not reflected in the parody in any way.

But, you know, I don’t care. I like what ChatGPT did. I like that some “guy named Dave” did the work, and the rhyme isn’t bad.

So I asked it to “try again” (a term of art).

Pwning Donald Trump

Here’s what ChatGPT came up with:

In Xanadu did Kubla Khan
A stately pleasure dome decree
Where wrinkled old men played golf and tan
And maidens danced in bikinis

With a river rushing through the land
And a clubhouse built with care
Where the men could sip their beer on demand
And tell stories of their golfing flair

But alas, the dome was not to last
For the old men lost their energy and zest
And the maidens found a better place to dance and have a blast
Leaving the old men to their lonely quest

For a stately pleasure dome no more
Xanadu was just a golf course
Where the old men played in vain, but never a bore
And the maidens went off to find a better source.

I sensed a stirring in the Force, and gave it another prompt:

Sounds like Donald Trump's version of “Kubla Khan.”

Indeed it does, a pleasure dome filled with old men and scantily clad women, a sign of excess and opulence.

Wonderful. You're a poet and don't know it...or do you?

I am indeed a poet and am always happy to share my words with others. Thank you for the compliment!

You know, somewhere out there on the web there is a page that shows the design Trump commissioned for a Xanadu hotel in Las Vegas. And there’s this book, from 2020, Palm Beach, Mar-a-Lago, and the Rise of America’s Xanadu. Those sites clearly link Trump with Xanadu and hence with “Kubla Khan.” And here’s an article from 2012 in which Trump says that Citizen Kane is his favorite movie. As we all known, the mansion in that film was called “Xanadu” and the first five lines of the poem were recited early in the film.

I note finally that well over a decade ago I did an informal study of the distribution of “Xanadu” on the web, which would have been much smaller then than it is now: One Candle, a Thousand Points of Light: The Xanadu Meme. While it turned up in various contexts, there were two major ones. One was high tech, which I called the cybernetic cluster, while the other was associated with luxury and excess, the sybaritic cluster.

I don’t know whether how those pages were in the corpus ChatGPT was trained on – many of them surely were, but there certainly is an attraction between DJT and the sybaritic cluster.

Saturday, December 3, 2022

Screaming on the flat part of the roller coaster ride: From GPT-3 to ChatGPT

Back in May of 2021 I did a blog post in which I asked GPT-3 to analyze a Jerry Seinfeld bit. OpenAI has now released a somewhat updated version, ChatGPT. I asked it to analyze the same bit. In this post I’m going to give you the bit, then part of what GPT-3 said about it in 2021, and then what ChatGPT said about it two days ago (December 1, 2022). Here’s the bit as Seinfeld performed it in 1977: 

Here's a slightly different version he published in his book Is This Anything?

I see they just finished the Roosevelt Island Tramway.

That’s nice…

The city’s going bankrupt,

they’re putting up rides for us.

Next thing you know, there’ll be a roller coaster through the South Bronx.

That would be the first roller coaster where the people scream on the flat part of the ride.

Back in 2021 I didn’t have direct access to GPT-3, so I asked Phil Mohun to present the bit to GPT-3. I’ve set Phil’s words in boldface.

Jerry Seinfeld tells the following joke:

I see they just finished the Roosevelt Island Tramway.
That’s nice…
The city’s going bankrupt,
they’re putting up rides for us.
Next thing you know, there’ll be a roller coaster through the South Bronx.
That would be the first roller coaster where the people scream on the flat part of the ride.

Q: Why would the people scream on the flat part of the ride?
GPT-3: Because it’s the Bronx.

Not very interesting.

Phil then increased the “temperature,” which raised the amount of randomness in GPT-3’s response. Note that on this try Phil also gave GPT-3 the whole bit, as in example 1.

Q: Why would the people scream on the flat part of the ride?
GPT-3: There’ll be nothing to see for miles but rubble.

Again, not very interesting. We played around a bit, never managed to cue GPT-3 into a decent explanation of the joke – you can read the whole interaction in my original post. We did, however, get some interesting things out of it. For example (notice that Phil’s second question is explicitly about screaming):

Q: Why would the people scream on the flat part of the ride?
GPT-3: The ride is the South Bronx.
Q: Why do people scream on roller coaster rides?
GPT-3: Because that’s the only time of the ride they realize there are other options.

I thought that was hilarious. But it hardly amounted to an explanation of the joke.

Now I present the bit to ChatGPT:

Jerry Seinfeld tells the following joke:

I see they just finished the Roosevelt Island Tramway.
That’s nice…
The city’s going bankrupt,
they’re putting up rides for us.
Next thing you know, there’ll be a roller coaster through the South Bronx.
That would be the first roller coaster where the people scream on the flat part of the ride.

Q: Why would the people scream on the flat part of the ride?

The joke is intended to be humorous because it imagines a roller coaster being built in a poor, crime-ridden neighborhood. The implication is that the people would be afraid for their safety and would scream out of fear and anxiety, even on the “flat” part of the ride where they would normally not be in any danger.

That’s got it. Since ChatGPT allows you can get multiple responses to the same prompt, I got two more responses:

The people would scream on the flat part of the ride because they would be scared of the dangerous neighborhoods and high crime rates in the South Bronx.

And:

Because the roller coaster would be passing through the South Bronx, a high-crime area known for its violence and poverty. The riders would be scared for their safety.

These aren’t quite as complete as the first response. But they get the central point, explicitly connecting screaming with the neighborhood.

Sunday, August 7, 2022

AGI as shibboleth, symbols [reacting to Jack Clark]

Jack Clark has a LONG tweet stream on AI policy. Though I don’t agree with every tweet – would anyone? – it’s worth at least a quick look. I want to comment on two of the tweets.

AGI as shibboleth, and beyond

Has AGI ever been anything other than a shibboleth? I believe the term was coined in the 1990s because some researchers felt that AI had become stale and focused on specialized domains, so-called “narrow” AI. The phrase “artificial general intelligence” (AGI) was as a banner under which to revive the founding goal of AI, to construct the artificial equivalent of human intelligence.

What researchers actually construct are mechanisms. But no one knows how to specify a mechanism or set of mechanisms for AGI. Oh, sure, there’s the Universal Turing machine which can, in point of abstract theory, compute any computable function. It may be a mechanism, but the idea so abstract that it provides little to no guidance in the construction of computer systems.

AGI, like AI before it, is an abstract goal, a beacon, without a procedure that will lead to it. No matter how vigorously you chase over the surface of the earth for the North Star, you’re never going to get there. And so AGI simply functions as a shibboleth. If you want into the club, you have to pledge allegiance to AGI.

But you don’t need to pledge allegiance in order to construct interesting and even useful systems. So why invent this unreachable goal? Is it just to define a club?

Meanwhile I’ve written a paper in which I define the idea of an artificial mind. I begin by defining mind:

A MIND is a relational network of logic gates over the attractor landscape of a partitioned neural network. A partitioned network is one loosely divided into regions where the interaction within a region is (much) stronger than the interactions between regions. Each of these regions will have many basins of attraction. The relational network specifies relations between basins in different regions.

Note that the definition takes the form of specifying a mechanism involving logic gates and a neural network. Given that:

A NATURAL MIND is one where the substrate is the nervous system of a living animal.

And:

An ARTIFICIAL MIND is one where the substrate is inanimate matter engineered by humans to be a mind.

There are other definitions as well as some caveats and qualifications.

However, those definitions come after 50 pages of text and diagrams in which I lay out the mechanisms that support those definitions. The paper is primarily about the human brain, but one can imagine constructing artificial devices that meet those specifications. Now, whether those specifications are the right specifications, that’s open for discussion. However that discussion turns out, it is a discussion about mechanisms, not myths and magic.

The paper:

Relational Nets Over Attractors, A Primer: Part 1, Design for a Mind, Version 2, Working Paper, July 13, 2022, pp. 76, https://www.academia.edu/81911617/Relational_Nets_Over_Attractors_A_Primer_Part_1_Design_for_a_Mind

Ah, symbols

Here’s a twofer:

It's the first tweet that interests me, but let’s look Richard Sutton’s bitter lesson. Here’s his opening paragraph:

The biggest lesson that can be read from 70 years of AI research is that general methods that leverage computation are ultimately the most effective, and by a large margin. The ultimate reason for this is Moore's law, or rather its generalization of continued exponentially falling cost per unit of computation. Most AI research has been conducted as if the computation available to the agent were constant (in which case leveraging human knowledge would be one of the only ways to improve performance) but, over a slightly longer time than a typical research project, massively more computation inevitably becomes available. Seeking an improvement that makes a difference in the shorter term, researchers seek to leverage their human knowledge of the domain, but the only thing that matters in the long run is the leveraging of computation. These two need not run counter to each other, but in practice they tend to. Time spent on one is time not spent on the other. There are psychological commitments to investment in one approach or the other. And the human-knowledge approach tends to complicate methods in ways that make them less suited to taking advantage of general methods leveraging computation. There were many examples of AI researchers' belated learning of this bitter lesson, and it is instructive to review some of the most prominent.

Sutton then goes on to list domains where there has proven so: chess, Go, speech recognition, and computer vision. He then draws some conclusions, which I want to bracket.

Note, however, that Sutton talks of researchers seeking “to leverage their human knowledge of the domain.” Is that what’s going on symbolic AI? Perhaps in expert systems, which may have been the most pervasive practical result of GOFAI. But I don’t think that’s an accurate general characterization. That’s not what was going on in computational linguistics, for example, or in much of the work on knowledge representation. That research was based on the belief that much of human knowledge is inherently symbolic in character and therefore that we must create models that capture that symbolic character.

Why did those models collapse? I think there are several factors involved:

1. Combinatorial explosion: Symbolic systems tend to generate large numbers of alternative with little or no way of choosing among them.

2. Hand coding: Symbolic systems have to be painstakingly hand-coded, which takes time.

3. Too many models, difficult to choose among them: This exacerbates the hand-coding problem.

4. Common sense has proven elusive: But then it has proven elusive for deep learning as well.

Perhaps the first problem can be solved through more computing power, though exponential search can easily outstrip the addition of CPU cycles and memory. The third problem is one for science, and is, I believe, entangled with the fourth one. The second problem is inconvenient, but, alas, if hand-coding is necessary, then it’s necessary. But perhaps if we’re clever....

On the fourth one, here’s what I said in my GPT-3 paper:

A lot of common-sense reasoning takes place “close” to the physical world. I have come to believe, but will not here argue, that much of our basic (‘common sense’) knowledge of the physical world is grounded in analogue and quasi-analogue representations. This gives us the power to generate language about such matters on the fly. Old school symbolic machines did not have this capacity nor do current statistical models, such as GPT-3.

Thus the problem is not specific to symbolic systems. It is quite general. It’s not at all clear that we can deal with this problem without having robots out and about in the world. I note that the working paper I mentioned in the previous section, Relational Nets Over Attractors, is about constructing symbolic structures over quasi-analog representations, which, following the terminology of Saty Chary, I characterize as structured physical systems.

Let’s return to Sutton’s paper. Here’s his final paragraph:

The second general point to be learned from the bitter lesson is that the actual contents of minds are tremendously, irredeemably complex; we should stop trying to find simple ways to think about the contents of minds, such as simple ways to think about space, objects, multiple agents, or symmetries. All these are part of the arbitrary, intrinsically-complex, outside world. They are not what should be built in, as their complexity is endless; instead we should build in only the meta-methods that can find and capture this arbitrary complexity. Essential to these methods is that they can find good approximations, but the search for them should be by our methods, not by us. We want AI agents that can discover like we can, not which contain what we have discovered. Building in our discoveries only makes it harder to see how the discovering process can be done.

I’m hesitant to think of symbol systems as being “simple ways to think about the contents of minds.” That strikes me as rhetorical overkill. But Sutton is right about “the arbitrary, intrinsically-complex, outside world.” He says that “we should build in only the meta-methods that can find and capture this arbitrary complexity.” Well, sure, why not?

But are we doing that now? That’s not at all obvious to me. it seems likely to me that the DL community is hoping that they’ve discovered the metamethods, or will do so in the near future, and so we don’t have to think about what’s going on inside either human minds or the machines we’re building. Well, if human minds use symbols, and it seems all but self-evident that we do – if language isn’t a symbol system, what is? – then the current repertoire of DL methods is not up to the task.

What meta-methods are needed to detect patterns of symbolic meaning and construct those quasi-analog representations?

My GPT-3 paper:

GPT-3: Waterloo or Rubicon? Here be Dragons, Version 4.1, Working Paper, May 7, 2022, 38 pp., https://www.academia.edu/43787279/GPT_3_Waterloo_or_Rubicon_Here_be_Dragons_Version_4_1

Sunday, June 26, 2022

GPT-3 is confused about ducks, chickens, and eggs

Mark Liberman reports some puzzling conversations with GPT-3.  The first conversation:

Me: Do ducks lay eggs?
GTP-3: Yes, ducks lay eggs.
Me: What do you mean by that?
GTP-3: Ducks lay eggs .
Me: Do all ducks lay eggs?
GTP-3: Yes, all ducks lay eggs.
Me: What about lions?
GTP-3: I'm not sure about lions.
Me: Do lions have manes?
GTP-3: I’m not sure about lions.
The beginning of the second one:
Me: Where do eggs come from?
GTP-3: Eggs come from chickens. Chickens lay eggs and then farmers collect them to sell.
Me: What about ducks? Do they lay eggs?
GTP-3: Yes, ducks also lay eggs. Duck eggs are often used in Chinese cuisine.
Me: What about roosters? Do they lay eggs?
GTP-3: No, roosters do not lay eggs.
Me: Are roosters chickens?
GTP-3: No, we consider roosters and chickens to be two different types of poultry. 

There's more at the link.

Thursday, June 16, 2022

GPT-3's grip on reality is not at all secure [HA!]

Tuesday, June 14, 2022

The wit and wisdom of GPT-3

Tuesday, June 7, 2022

Gwern on GPT-3 and arithmetic

In a discussion of the "Blessings of Scale:"

...a sub-model which memorizes some of the data is indeed much simpler than a sub-model which encodes genuine arithmetic (a NN can probably memorize tens of thousands of lookup table entries storing examples of addition in the space it would take to encode an abstract algorithm like ‘addition’), but it can’t possibly memorize all the instances of arithmetic (implicit or explicit) in GPT-3’s Internet-scale dataset. If a memorizing sub-model tried to do so, it would become extremely large and penalized. Eventually, after enough examples and enough updates, there may be a phase transition (Viering & Loog 2021), and the simplest ‘arithmetic’ model which accurately predicts the data just is arithmetic. And then the meta-learning, after seeing enough instances of algorithms which vary slightly within each sample, making it hard to learn each task separately, just is learning of more generic algorithms, yielding sub-models which achieve lower loss than the rival sub-models, which either fail to predict well or bloat unacceptably. (GPT-2-1.5b apparently was too small or shallow to ensemble easily over sub-models encoding meta-learning algorithms, or perhaps not trained long enough on enough data to locate the meta-learner models; GPT-3 was.)

Monday, May 2, 2022

In words and images, a tweet stream explaining transformers

There are 10 more tweets in the string, for a total of 12.

Monday, April 25, 2022

Semanticity: adhesion and relationality

For some time now I have been puzzled by the (astonishing) success of statistical models, especially artificial neural networks, in language processing, machine translation in particular – see, e.g. this post, Borges redux: Computing Babel – Is that what’s going on with these abstract spaces of high dimensionality? [#DH], which dates back to October 2017. Sure, it is statistics, but what are those statistics “grabbing on to?” There is no meaning there, just (naked) word forms. What is it about those word-forms-in-context that yields an approximation to, a simulacrum of, meaning? Better, what’s there that we readily read meaning, intention, into the results of these statistical techniques.

Semanticity and intention

My puzzlement reached a climax with the unveiling of GPT-3 in 2020 and I decided to take another run at the problem. I produced a working paper, GPT-3: Waterloo or Rubicon? Here be Dragons, which I liked very much. I made real progress. I now think I can nudge things forward another step. Look at this passage, where I discuss the Chinese Room thought-experiment (p. 28):

Yet if you would believe John Searle, no matter how rich and detailed those old school mental models, understanding would necessarily elude them. I am referring, of course, to his (in)famous Chinese Room argument. When I first encountered it years ago my reaction was something like: interesting, but irrelevant. Why irrelevant? Because it said absolutely nothing about the techniques AI or cognitive science investigators used and so would provide no guidance toward improving that work. He did, however, have a point: If the machine has no contact with the world, how can it possibly be said to understand anything at all? All it does is grind away on syntax.

What Searle misses, though, is the way in which meaning is a function of relations among concepts, as I pointed out earlier (pp. 18 ff.). It seems to me, however – and here I’m just making this up – we can think of meaning as having both an intentional aspect, the connection of signs to the world, and a relational aspect, the relations of signs among themselves. Searle’s argument concentrated on the former and said nothing about the latter.

What of the intentional aspect when a person is writing or talking about things not immediately present, which is, after all quite common? In this case the intentional aspect of meaning is not supported by the immediate world. Language use thus must necessarily be driven entirely by the relations signifiers have among themselves, Sydney Lamb’s point which we have already investigated (p. 18).

Those statistics are grabbing onto the relational aspect of meaning. The question is: How much of that can these methods recover from texts? Let’s set that aside for the moment.

That passage mentions intention and relation. Intention resides in the relationship between a person and the world. Relation resides in the relationships that signifiers have among themselves. It is a property of the cognitive system. I am now thinking that it must be paired with adhesion. Taken together they constitute semanticity. Thus we have semanticity and intention where semanticity is a general capacity inherent in the cognitive system, in a person’s mind, and intention inheres in the relation between a person and the world in a particular perceptual and/or cognitive activity.

What do I mean by adhesion? Adhesion is how words ‘cling’ to the world while relationality is the differential interaction of words among themselves within the linguistic system. Words whose meaning is defined directly over the physical world, but also, to some extent, the interpersonal world of signals and feeling, they adhere to the world through sensorimotor schemas. Words whose meaning is abstract are more problematic. Their adhesion operates though patterns of words and other signs and symbols (e.g. mathematics, data visualizations, illustrative diagrams of various kinds, and so forth). Teasing out these systems of adhesion has just barely begun.

The psychologist J.J. Gibson talked of the affordances an environment presents to the organism. Affordances as the features of the world which an organism can readily pick up during its life in the world. Adhesions are the organism’s complement to environmental affordances; they are the perceptual devices through which the organism relates to the affordances.

What this means for language models

Large language models built through deep neural networks, such as GPT-3, conflate the interaction of three phenomena: 1) the world-level relational aspect of semanticity as captured in the locations of word forms (signifiers) in a string, 2) the conventions of discourse structure, and 3) the world itself. The world is present in the model because the texts over which the model was constructed were created by people interacting in the world. They were in an intentional relationship with the world when they wrote those texts. The conventions of discourse are present simply because they organize the placement of word forms in a text, with special emphasis on the long-distance relationships of word. As for relationality, that’s all that can possibly be present in a text. Adhesions belong to the realm of signifieds, of concepts and ideas, and they aren’t in the text itself.

Would it somehow be possible to factor a language model into these three aspects? I have no idea. The point of doing so would be to reduce the overall size of the model.

Putting that aside, let us ask: Given a sufficiently large database of texts and tokens and a high enough number of parameters for our model, is it possible for a language model to extract all the relationality from the texts? How much of that multidimensional relational semanticity can be recovered from strings of word forms? Given a deep enough understanding of how relational semantics is reflected in the structure of texts, can we calculate what is possible with various text bases and model parameterization?

To answer those questions we need to have some account of semantic relationality which we can examine. The models of Old School symbolic AI and computational linguistics provide such accounts. Many such models have been created. Which ones would we choose as the basis for our analysis? The sort of question that interests me is how many word forms have their meanings given in adhesions to the physical world (that is, physical objects and events), to the interpersonal world (facial expressions, gestures, etc.) and how many word forms are defined abstractly?

So many questions. 

* * * * *

I have appended this to my GPT-3 working paper, which is now Version 3.

Saturday, April 16, 2022

Have we reached a tipping point? Which one? [AI, GPT-3]

Steven Johnson has an interesting article in The New York Times Magazine: A.I. Is Mastering Language. Should We Trust What It Says? (April 15, 2022). He starts by showing us a guessing game, guess the missing ____. He talks about the technology – GPT-3 and AlphaFold, and so forth – discusses its implications, and arrives at the meeting that gave birth to OpenAI:

OpenAI’s origins date to July 2015, when a small group of tech-world luminaries gathered for a private dinner at the Rosewood Hotel on Sand Hill Road, the symbolic heart of Silicon Valley. The dinner took place amid two recent developments in the technology world, one positive and one more troubling. On the one hand, radical advances in computational power — and some new breakthroughs in the design of neural nets — had created a palpable sense of excitement in the field of machine learning; there was a sense that the long “A.I. winter,” the decades in which the field failed to live up to its early hype, was finally beginning to thaw. A group at the University of Toronto had trained a program called AlexNet to identify classes of objects in photographs (dogs, castles, tractors, tables) with a level of accuracy far higher than any neural net had previously achieved. Google quickly swooped in to hire the AlexNet creators, while simultaneously acquiring DeepMind and starting an initiative of its own called Google Brain. The mainstream adoption of intelligent assistants like Siri and Alexa demonstrated that even scripted agents could be breakout consumer hits.

But during that same stretch of time, a seismic shift in public attitudes toward Big Tech was underway, with once-popular companies like Google or Facebook being criticized for their near-monopoly powers, their amplifying of conspiracy theories and their inexorable siphoning of our attention toward algorithmic feeds. Long-term fears about the dangers of artificial intelligence were appearing in op-ed pages and on the TED stage. Nick Bostrom of Oxford University published his book “Superintelligence,” introducing a range of scenarios whereby advanced A.I. might deviate from humanity’s interests with potentially disastrous consequences. In late 2014, Stephen Hawking announced to the BBC that “the development of full artificial intelligence could spell the end of the human race.” It seemed as if the cycle of corporate consolidation that characterized the social media age was already happening with A.I., only this time around, the algorithms might not just sow polarization or sell our attention to the highest bidder — they might end up destroying humanity itself. And once again, all the evidence suggested that this power was going to be controlled by a few Silicon Valley megacorporations.

The agenda for the dinner on Sand Hill Road that July night was nothing if not ambitious: figuring out the best way to steer A.I. research toward the most positive outcome possible, avoiding both the short-term negative consequences that bedeviled the Web 2.0 era and the long-term existential threats. From that dinner, a new idea began to take shape — one that would soon become a full-time obsession for Sam Altman of Y Combinator and Greg Brockman, who recently had left Stripe. Interestingly, the idea was not so much technological as it was organizational: If A.I. was going to be unleashed on the world in a safe and beneficial way, it was going to require innovation on the level of governance and incentives and stakeholder involvement. The technical path to what the field calls artificial general intelligence, or A.G.I., was not yet clear to the group. But the troubling forecasts from Bostrom and Hawking convinced them that the achievement of humanlike intelligence by A.I.s would consolidate an astonishing amount of power, and moral burden, in whoever eventually managed to invent and control them.

In December 2015, the group announced the formation of a new entity called OpenAI.

He goes on to tell us how OpenAI was conceived as a nonprofit but the prodigious costs of the ‘compute’ required to build state of the art AI forced it to create a for-profit partner, OpenAI L.P. That’s very interesting. Johnson returns to the technology itself, talks of the “stochastic parrot” critique by Emily Bender and Timnit Gebru, checks in with uber-skeptic Gary Marcus, and so forth and so on. After a while he gets around to this:

So how do you widen the pool of stakeholders with a technology this significant? Perhaps the cost of computation will continue to fall, and building a system competitive to GPT-3 will become within the realm of possibility for true open-source movements, like the ones that built many of the internet’s basic protocols. (A decentralized group of programmers known as EleutherAI recently released an open source L.L.M. called GPT-NeoX, though it is not nearly as powerful as GPT-3.) Gary Marcus has argued for “a coordinated, multidisciplinary, multinational effort” modeled after the European high-energy physics lab CERN, which has successfully developed billion-dollar science projects like the Large Hadron Collider. “Without such coordinated global action,” Marcus wrote to me in an email, “I think that A.I. may be destined to remain narrow, disjoint and superficial; with it, A.I. might finally fulfill its promise.”

Yes, and there’s more of that. But really, I want to get around to something else. If you’ve been thinking about these issues, by all means, read the article.

I do believe that GPT-3 and its kin represent a potential phase shift in technology, something I talked about at some length in a working paper I posted two years ago, GPT-3: Waterloo or Rubicon? Here be Dragons, Version 2. But what does that mean, phase change? That’s when some folks start talking about the so-called Tech Singularity, which I’ve also written about, Redefining the Coming Singularity – It’s not what you think, Version 2. I’m now going to say a bit more.

It’s common to think of such things as exhibiting an “S curve,” like this:

That’s my view, but it’s not how the Singularitarians think. They think the rise is almost vertical and that it’s driven by the technology itself. Somewhere down there to the left some computer or computers become self-aware and proceeded to make themselves smarter and smarter and smarter and  – 

FOOM!

That’s nonsense.

WE are driving the rise. As we learn more about the technology, and more about human and, indeed, animal intelligence, we are able to develop and deploy more sophisticated intelligence. The singularity is in our minds, our collective culture.

Now, let’s go back to that meeting that was held in July 2015. Where do we place it on that curve? Do we place it here?

Or maybe a bit further up, perhaps here:

Given that AI got its start back in the 1950s so might argue that we much nearer the shoulder of the curve.

I’m skeptical. I think that prior work has been ‘absorbed’ in the horizontal to the left. [These are only suggestive diagrams, visual metaphors.] But we don’t really know.

It’s not at all obvious that we’ll be able to climb that curve. I believe that the technology Johnson discusses represents an advance over the technology we had, say, five years ago; but I do not think it will take us up the curve. I think we’ve got a lot more to learn, a lot. Do we have the wisdom and will to learn it?

* * * * *

Emily Bender replies to Steven Johnson, On NYT Magazine: Resist the Urge to be Impressed.