Showing posts with label Scott_Aaronson. Show all posts
Showing posts with label Scott_Aaronson. Show all posts

Tuesday, February 13, 2024

What's so special about us? Why do we insist on specialness?

Scott Aaronson has just posted "The Problem of Human Specialness in the Age of AI."

For the past year and a half, I’ve been moonlighting at OpenAI, thinking about what theoretical computer science can do for AI safety. [...] In addition to “how do we stop AGI from going disastrously wrong?,” I find myself asking “what if it goes right? What if it just continues helping us with various mental tasks, but improves to where it can do just about any task as well as we can do it, or better? Is there anything special about humans in the resulting world? What are we still for?”

Here's a comment I posted over there:

@Matteo Villa, #31:

Should we not just let go of human specialness, just like humanity had to do when discovering that we are not at the centre of the universe and not even at the centre of our solar system?

I've been wondering the same thing. It's not as though the universe was made for us or is somehow ours to do with as we see fit. It just is.

From Benzon and Hays, The Evolution of Cognition, 1990:

A game of chess between a computer program and a human master is just as profoundly silly as a race between a horse-drawn stagecoach and a train. But the silliness is hard to see at the time. At the time it seems necessary to establish a purpose for humankind by asserting that we have capacities that it does not. It is truly difficult to give up the notion that one has to add “because . . . “ to the assertion “I’m important.” But the evolution of technology will eventually invalidate any claim that follows “because.” Sooner or later we will create a technology capable of doing what, heretofore, only we could.

Something I've just begun to think about: What role can these emerging AIs play in helping us to synthesize what we know? Ever since I entered college in the Jurassic era I've been hearing laments about how intellectual work is becoming more and more specialized. I've seen and see the specialization myself. How do we put it all together? That's a real and pressing problem. We need help.

I suppose one could say: "Well, when a superintelligent AI emerges it'll put it all together." That doesn't help me all that much, in part because I don't know how to think about superintelligent AI in any way I find interesting. No way to get any purchase on it. That discussion – and I suppose the OP (alas) fits right in – just seems to me rather like a rat chasing its own tail. A lot of sound and fury signifying, you know...

But trying to synthesize knowledge, trying to get a broader view. That's something I can think about – in part because I've spent a lot of time doing it – and we need help. Will GPT-5 be able to help with the job? GPT-6?

BTW, before the Copernican Revolution we weren't special (by "we" I mean Europeans and their descendants; I don't know off hand how the rest of the world thought about these matters). Earth was as the "bottom" of the cosmos. Of course that was a very different cosmos from the one we're imagining today. That was a cosmos ordered by God.

Maybe the evolutionary psychologists have something to day about this need to think of ourselves as the masters of the universe. 

ADDENDUM: Perhaps some of my thoughts about the (coming) Fourth Arena are relevant:

Monday, May 15, 2023

Two more AI videos: Geoffrey Hinton, Aaronson & Hanson

Perhaps I'll offer some commentary later in separate posts. I'm posting these here and now as a reminder to myself.

Tuesday, April 11, 2023

Some comments on “intelligence” for jinzo ninge (Japanese for “artificial beings”)

Back in mid-March Scott Aaronson made a post entitled: On overexcitable children. The post begins:

Wilbur and Orville are circumnavigating the Ohio cornfield in their Flyer. Children from the nearby farms have run over to watch, point, and gawk. But their parents know better.

An amusing toy, nothing more. Any talk of these small, brittle, crash-prone devices ferrying passengers across continents is obvious moonshine. One doesn’t know whether to laugh or cry that anyone could be so gullible.

Or if they were useful, then mostly for espionage and dropping bombs. They’re a negative contribution to the world, made by autistic nerds heedless of the dangers.

And so forth. He’s obviously using Wilbur and Orville as figures for the developers of AI and their toy as a figure for current devices.

The ensuing discussion was variously interesting, even exiting, mediocre and boring, and pointless, as such things are on the web tubes. But the good stuff makes it worthwhile for me to follow along and occasionally throw in my 2¢. Here’s a nickel’s worth.

General intelligence

@Scott #101: You observe:

On the other hand, with GPT, the world has just witnessed a striking empirical confirmation that, as you train up a language model, countless different abilities (coding, logic puzzles, math problems, language translation…) emerge around the same time, without having to be explicitly programmed. Doesn’t this support the idea that there is such a thing as “general intelligence” that continues to make sense for non-human entities?

I’m going to go all-out weasel and say, I don’t know what it does.

One problem is that “intelligence” isn’t just a word that means “can do a lot of cognitive stuff.” At some point in the last 100 years or so it has become surrounded by a quasi-mystical aura that gets in the way. Saying that “as we scale up GPTs they become more capable” doesn’t have quite the same oomph that “as we scale up GPTs they become more intelligent” does.

We know right now that GPT-3 to 4 can do a lot more stuff at a pretty high level than any human being can do. We’ve got other more specialized systems that can do a few things – play chess, Go, predict protein folding, design a control regime for plasma containment – better than any human can. For that matter, we’ve long had computers that can do arithmetic way better than any humans. We think nothing of that because it’s merely routine and long has been. But things were trickier before the adoption of the Arabic notation.

Getting back to the GPTs, are there any specialized cognitive tasks they can do better than the best human? I don’t know off hand, I’m just asking. But I suppose that’s what the discussion is about, better than the best human. What if GPT-X turns out to prove one of those theorems you’re interested in? What then? But what if that’s the only thing it does better than the best human, but in many other areas it’s better than GPT-4, but not up to merely superior (as opposed to the best) human performance? What then? I don’t know.

And I’m having trouble keeping track of my line of thought. Oh, OK, so GPT-4 knows lots more stuff than any one human. But it also messes up in simple ways. Given that it has some visual capabilities, I wonder if it would go “off the reservation” when confronted with The Towers of Warsaw* in the hilarious way that ChatGPT did? That’s a digression. Even within its range of capabilities, though, GPT-4 hallucinates. What are we to make of that?

What I make of it is that it’s a very difficult problem. I’m guessing that Gary Marcus would say that to solve the problem you need a world model. OK. But how do you keep the world model accurate and up-to-date? That, it seems to me, is a difficult problem for humans, very difficult. As far as I can tell, we deal with the problem by constantly communicating with one another on all sorts of things at all sorts of levels of sophistication.

Let me take another stab at it.

Given the interests of the people who comment here, examples of Ultimate Problems tend to be drawn from science and math. While I’ve got an educated person’s interest in those things, I’m driven by curiosity about other things. While I’ve got a general interest in language and the mind, I’m particularly interest in literature and, above all else, one poem in particular, Coleridge’s “Kubla Khan.” Why that poem?

Because I discovered that, by treating line-end punctuation like nested parentheses in a Lisp expression, the poem is structured like a pair of matryoshka dolls. The poem has two parts, the first twice as long as the second. Each of them is divided into three, the middle is in turn divided into three, and once more, divided into three. All other divisions are binary. And the last line of the first part turns up in the structural center of the second part. So: line 36, “A stately pleasure-dome with caves of ice!”, line 47: “That sunny dome! those caves of ice!”

The whole thing smelled like computation. But computation of what, and how? That’s what drove me to computational linguistics, which I found very interesting. But it didn’t solve my problem. So I’ve been working on that off and on ever since. Oh, I’ve spent a lot of time on other things, a lot of time, but I still check in with “Kubla Khan” every now and then.

I took another look last week (you'll find diagrams at the link that make this is lot clearer). Between vector semantics and in-context learning I’ve made a bit more progress. Who knows, maybe GPT-X will be able to tell me what’s going on. And if it can tell me that, it’ll be able to tell us all a lot more about the human mind and about language.

Short of that, it would be nice to have a GPT, or some other LLM, that’s able to examine a literary text and tell me whether or not it exhibits ring-composition, which is generally depicted like this:

A, B, C...X...C’, B’, A’

It’s an obscure and all-but forgotten topic in literary studies, more prominent among classicists and Biblical scholars. I learned about it from the late Mary Douglas, an important British anthropologist who got knighted, or whatever it is called for women, in recognition of her general work in anthropology. The two parts of “Kubla Khan” exhibit that form. But so do many other texts, like Conrad’s Heart of Darkness, Obama’s Eulogy for Clementa Pinckney, or, of all things, Pulp Fiction, perhaps Shakespeare’s Hamlet as well. Figuring that out is not rocket science. But it’s tricky and tedious.

I’m afraid I’ve strayed rather far afield, Scott. But that’s more or less how I think about “intelligence” or whatever the heck it is. My interest in these matters seems to be dominated by a search for mechanisms, like those in literary texts. Thus, while I’m willing to take ChatGPT’s performance at face value – and believe there’s more coming down the pike, I really want to know how it works. That’s a far more compelling issue that whatever the heck intelligence is. [BTW, the Chatster can tell stories exhibiting ring-composition. That’s one kind of skill, but entirely different from being able to analyze and identify ring-composition in texts (or movies).]

*A variant of The Towers of Hanoi. The classic version is posed with three pegs and five graduated rings. Back in the early 70s some wise guys at Carnegie-Mellon posed a variant with five pegs and three rings.

Scale & intelligence

Scott, let me take another crack at the question you posed in #101, if I may paraphrase: What do we make of the fact that all sorts of capabilities just keep showing up in GPTs without any explicit programming? First, let’s put on the table the work the Anthropic people have done on In-context Learning and Induction Heads. They are circuits that emerge at a certainly (relatively early) point during training and seem to be capable of copying a sequence of tokens, completing a sequence, even pattern matching. What can you do with general pattern matching? Lots of things.

What follows is hardly a rigorous argument, but it’s a place to start. Consider the idea of a tree, by which I mean a kind of plant, not a mathematical object. Though trees are physical objects, they are conceptually abstract. No one ever saw a generic tree. What you see are individual examples of maple trees, or palm trees, or pine trees. Maples, palms, and pines appear quite different. Why would anyone ever think of them as the same kind of thing? Well, consider them in the context of bushes and shrubs, grasses, and flowers. In that context, their size would bring them together in similarity space, just as bushes and shrubs would have their region, grasses would have theirs, flowers would have theirs, and we’ll have to have a region for vines as well. So now we have trees, and bushes, etc. and what are they? They’re all plants. As such, they are distinguished from animals.

So now we have all these abstract or general categories for plants and animals. Let’s take the whole abstraction process up a level and talk about species and genera and biological taxonomy in general. Perhaps we go up another level of abstraction from there and arrive at graph theory.

But to keep climbing to these higher levels of abstraction, we need more and more examples to deal with, and more compute to make all the comparisons and sort things out. And GPTs, of course, aren’t dealing with patterns over physical objects. They’re dealing with patterns over tokens. But those tokens encode all kinds of statements about physical and other kinds of objects. So it is not deeply surprising to me that when you through enough compute at enough texts, all sorts of interesting capabilities show up. I’m not saying or implying that I understand how this works. I don’t. But it doesn’t violate my sense of how the world works.

Now, Jean Piaget, the great developmental psychologist, had this idea of reflective abstraction. He’s the one who elaborated on the idea that cognitive development happens in stages during a child’s life. Children aren’t just learning more and more facts. They’re developing more sophisticated ways of thinking. Thinking processes at a higher level take as their objects, processes at a lower level. He was unclear on just how this works & I’m not sure anyone has tried to figure it out – but then I haven’t looked into the stuff in a while. The general idea is simply that more sophisticated levels of thinking are built on lower-level processes. And, while he was mostly interested in child development, he also applied the idea to the cultural development of ideas.

So maybe one thing we’re seeing as GPTs scale up is phase changes in capability as we use more compute and more examples. The basic architecture remains the same, but significant jumps in capacity allow for new capabilities. To return to maple tree and cows, it’s one thing to encompass more plants and animals in the system. But maybe he takes a major leap in compute to be able to abstract over that whole system and come up with the ideas of family, genus, species and so forth.

Setting that aside, Richard Hanania has just done a podcast with Robin Hanson. In the section on Intelligence and “Betterness” Hanson observes:

So the issue is the kind of meaning behind various abstractions we use. So abstractions are powerful. We use abstractions to organize the world, and abstractions embody similarities between the things out there. And we care about our abstractions, and which pool we use.

But for some abstractions, they well summarize our ambitions and our hopes, but they don’t necessarily correspond to a thing out there, where there’s a knob on them you can turn and change things. So it’s important to distinguish which of our abstractions correspond to things that we can have more direct influence over, and which abstractions are just abstractions about our view of the world and our desires about the world. So that’s the key distinction here. We could talk about a good world and a happy world and a nice world, but there isn’t a knob in the world to turn out and make the world nicer in some sense.

In the next two sections (Knowledge Hierarchy and Innovation, The History of Economic Growth) Hanson tosses out some ideas about abstraction that are useful in thinking about these matters. Later:

And now we have this parameter intelligence, and the question is, what’s that? How does that fit in with all these other parameters? We don’t usually use intelligence, say, as a measure of a country or a measure of a firm. We use wealth or other parameters. If it’s equivalent, then fine. If it’s something separate then we want to go, “Well, what is that exactly?”

For an individual, we have this measure of intelligence for an individual in the sense that there’s a correlation across mental tasks and which ones they can do better. And then the question is, what’s the cause of that correlation? One theory is that some people’s brains just trigger faster, and if I got a brain that triggers faster, it can just think faster and then overall it can do more.

There are other theories, but there are ways to cash out, what is it that makes one person smarter than another? Maybe they just have a bigger brain. That’s one of the stories, a brain that triggers faster. Maybe a brain with certain modules that are more emphasized than others. Then that’s a story of the particular features of that brain, that makes it be able to do many more tasks.

And so on.

Tuesday, August 23, 2022

Which comes first, AGI or new systems of thought? [Further thoughts on the Pinker/Aaronson debates]

Long-time readers of New Savanna know that David Hays and I have a model of cultural evolution built on the idea of cognitive rank, systems of thought embedded in large-scale cognitive architecture. Within that context I have argued that we are currently undergoing a large-scale transformation comparable to those that gave us the Industrial Revolution (Rank 3 in our model) and, more recently, the conceptual revolutions of the first half of the 20th century (Rank 4). Thus I have suggested that the REAL singularity is not the fabled tech singularity, but the consolidation of new conceptual architectures:

Redefining the Coming Singularity – It’s not what you think, Version 2, Working Paper, November 2015, https://www.academia.edu/8847096/Redefining_the_Coming_Singularity_It_s_not_what_you_think

I had occasion to introduce this idea into the recent AI debate between Steven Pinker and Scott Aaronson. Pinker was asking for specific mechanisms underpinning superintelligence while Aaronson was offering what Steve called “superpowers.” Thus Pinker remarked:

If you’ll forgive me one more analogy, I think “superintelligence” is like “superpower.” Anyone can define “superpower” as “flight, superhuman strength, X-ray vision, heat vision, cold breath, super-speed, enhance hearing, and nigh-invulnerability.” Anyone could imagine it, and recognize it when he or she sees it. But that does not mean that there exists a highly advanced physiology called “superpower” that is possessed by refugees from Krypton! It does not mean that anabolic steroids, because they increase speed and strength, can be “scaled” to yield superpowers. And a skeptic who makes these points is not quibbling over the meaning of the word superpower, nor would he or she balk at applying the word upon meeting a real-life Superman. Their point is that we almost certainly will never, in fact, meet a real-life Superman. That’s because he’s defined by human imagination, not by an understanding of how things work. We will, of course, encounter machines that are faster than humans, and that see X-rays, that fly, and so on, each exploiting the relevant technology, but “superpower” would be an utterly useless way of understanding them.

I’ve added my comment below the asterisks.

* * * * *

I’m sympathetic with Pinker and I think I know where he’s coming from. Thus he’s done a lot of work on verb forms, regular and irregular, that involves the details of (computational) mechanisms. I like mechanisms as well, though I’ve worried about different ones than he has. For example, I’m interested in (mostly) literary texts and movies that have the form: A, B, C...X...C’, B’, A’. Some examples: Gojira (1954), the original 1933 King Kong, Pulp Fiction, Obama’s eulogy for Clementa Pinkney, Joseph Conrad’s Heart of Darkness, Shakespeare’s Hamlet, and Osamu Tezuka’s Metropolis.

What kind of computational process produces such texts and what kind of computational process is involved in comprehending them? Whatever that process is, it’s running in the human brain, whose mechanisms are obscure. There was a time when I tried writing something like pseudo-code to generate one or two such texts, but that never got very far. So these days I’m satisfied identifying and describing such texts. It’s not rocket science, but it’s not trivial either. It involves a bit of luck and a lot of detail work.

So, like Steve, I have trouble with mechanism-free definitions of AGI and superintelligence. When he contrasts defining intelligence as mechanism vs. magic, as he did earlier, I like that, as I like his current contrast between “intelligence as an undefined superpower rather than a[s] mechanisms with a makeup that determines what it can and can’t do.”

In contrast Gary Marcus has been arguing for the importance of symbolic systems in AI in addition to neural networks, often with Yann LeCun as his target. I’ve followed this debate fairly carefully, and even weighed in here and there. This debate is about mechanisms, mechanisms for computers, in the mind, for the near-term and far-term.

Whatever your current debate with Steve is about, it’s not about this kind of mechanism vs. that kind. It has a different flavor. It’s more about definitions, even, if you will, metaphysics. But, for the sake of argument I’ll grant that, sure, the concept of intellectual superpowers is coherent (even if we have little idea about how’d they’d work beyond MORE COMPUTE!).

With that in mind, you say:

Not only does the concept of “superpowers” seem coherent to me, but from the perspective of someone a few centuries ago, we arguably have superpowers—the ability to summon any of several billion people onto a handheld video screen at a moment’s notice, etc. etc. You’d probably reply that AI should be thought of the same way: just more tools that will enhance our capabilities, like airplanes or smartphones, not some terrifying science-fiction fantasy.

I like the way you’ve introduced cultural evolution into the conversation, as that’s something I’ve thought about a great deal. Mark Twain wrote a very amusing book, A Connecticut Yankee in King Arthur’s Court. From the Wikipedia description:

In the book, a Yankee engineer from Connecticut named Hank Morgan receives a severe blow to the head and is somehow transported in time and space to England during the reign of King Arthur. After some initial confusion and his capture by one of Arthur's knights, Hank realizes that he is actually in the past, and he uses his knowledge to make people believe that he is a powerful magician.

Is it possible that in the future there will be human beings as far beyond us as that Yankee engineer was beyond King Arthur and Merlin? It seems to me that, providing we avoid disasters like nuking ourselves back to the Stone Age, catastrophic climate change exacerbated by pandemics, and getting paperclipped by an absentminded Superintelligence, it seems to me almost inevitable that that will happen. Of course science fiction is filled with such people but, alas, has not a hint of the theories that give them such powers. But I’m not talking about science fiction futures. I’m talking about the real future. Over the long haul we have produced ever more powerful accounts of how the world works and ever more sophisticated technologies through which we have transformed the world. I see no reason why that should come to a stop.

So, at the moment various researchers are investigating the parameters of scale in LLMs. What are the effects of differing numbers of tokens in the training corpus and number of parameters in the model? Others are poking around inside the models to see what’s going on in various layers. Still others are comparing the response characteristics of individual units in artificial neural nets with the response characteristics of neurons in biological visual systems. And so and on and so forth. We’re developing a lot of empirical knowledge about how these systems work, and models here and there.

I have no trouble at all imagining a future in which we will know a lot more about how these artificial models work internally and how natural brains work as well. Perhaps we’ll even be able to create new AI systems in the way we create new automobiles. We specify the desired performance characteristics and then use our accumulated engineering knowledge and scientific theory to craft a system that meets those specifications.

It seems to me that’s at least as likely as an AI system spontaneously tipping into the FOOM regime and then paperclipping us. Can I predict when this will happen? No. But then I regard various attempts to predict the arrival of AGI (whether through simple Moore’s Law type extrapolation or the more heroic efforts of Open Philanthropy’s biological anchors) as mostly epistemic theater.

Saturday, July 23, 2022

Steven Pinker and Scott Aaronson debate the nature of intelligence and superintelligence

They're at it again: More AI debate between me and Steven Pinker! Here's my contribution (so far):

I’m sympathetic with Pinker and I think I know where he’s coming from. Thus he’s done a lot of work on verb forms, regular and irregular, that involves the details of (computational) mechanisms. I like mechanisms as well, though I’ve worried about different ones than he has. For example, I’m interested in (mostly) literary texts and movies that have the form: A, B, C...X...C’, B’, A’. Some examples: Gojira (1954), the original 1933 King Kong, Pulp Fiction, Obama’s eulogy for Clementa Pinkney, Joseph Conrad’s Heart of Darkness, Shakespeare’s Hamlet, and Osamu Tezuka’s Metropolis.

What kind of computational process produces such texts and what kind of computational process is involved in comprehending them? Whatever that process is, it’s running in the human brain, whose mechanisms are obscure. There was a time when I tried writing something like pseudo-code to generate one or two such texts, but that never got very far. So these days I’m satisfied identifying and describing such texts. It’s not rocket science, but it’s not trivial either. It involves a bit of luck and a lot of detail work.

So, like Steve, I have trouble with mechanism-free definitions of AGI and superintelligence. When he contrasts defining intelligence as mechanism vs. magic, as he did earlier, I like that, as I like his current contrast between “intelligence as an undefined superpower rather than a[s] mechanisms with a makeup that determines what it can and can’t do.”

In contrast Gary Marcus has been arguing for the importance of symbolic systems in AI in addition to neural networks, often with Yann LeCun as his target. I’ve followed this debate fairly carefully, and even weighed in here and there. This debate is about mechanisms, mechanisms for computers, in the mind, for the near-term and far-term.

Whatever your current debate with Steve is about, it’s not about this kind of mechanism vs. that kind. It has a different flavor. It’s more about definitions, even, if you will, metaphysics. But, for the sake of argument I’ll grant that, sure, the concept of intellectual superpowers is coherent (even if we have little idea about how’d they’d work beyond MORE COMPUTE!).

With that in mind, you say:

Not only does the concept of “superpowers” seem coherent to me, but from the perspective of someone a few centuries ago, we arguably have superpowers—the ability to summon any of several billion people onto a handheld video screen at a moment’s notice, etc. etc. You’d probably reply that AI should be thought of the same way: just more tools that will enhance our capabilities, like airplanes or smartphones, not some terrifying science-fiction fantasy.

I like the way you’ve introduced cultural evolution into the conversation, as that’s something I’ve thought about a great deal. Mark Twain wrote a very amusing book, A Connecticut Yankee in King Arthur’s Court. From the Wikipedia description:

In the book, a Yankee engineer from Connecticut named Hank Morgan receives a severe blow to the head and is somehow transported in time and space to England during the reign of King Arthur. After some initial confusion and his capture by one of Arthur's knights, Hank realizes that he is actually in the past, and he uses his knowledge to make people believe that he is a powerful magician.

Is it possible that in the future there will be human beings as far beyond us as that Yankee engineer was beyond King Arthur and Merlin? It seems to me that, providing we avoid disasters like nuking ourselves back to the Stone Age, catastrophic climate change exacerbated by pandemics, and getting paperclipped by an absentminded Superintelligence, it seems to me almost inevitable that that will happen. Of course science fiction is filled with such people but, alas, has not a hint of the theories that give them such powers. But I’m not talking about science fiction futures. I’m talking about the real future. Over the long haul we have produced ever more power accounts of how the world works and ever more sophisticated technologies through which we have transformed the world. I see no reason why that should come to a stop.

So, at the moment various researchers are investigating the parameters of scale in LLMs. What are the effects of differing numbers of tokens in the training corpus and number of parameters in the model? Others are poking around inside the models to see what’s going on in various layers. Still others are comparing the response characteristics of individual units in artificial neural nets with the response characteristics of neurons in biological visual systems. And so and on and so forth. We’re developing a lot of empirical knowledge about how these systems work, and models here and there.

I have no trouble at all imagining a future in which we will know a lot more about how these artificial models work internally and how natural brains work as well. Perhaps we’ll even be able to create new AI systems in the way we create new automobiles. We specify the desired performance characteristics and then use our accumulated engineering knowledge and scientific theory to craft a system that meets those specifications. It seems to me that’s at least as likely as an AI system spontaneously tipping into the FOOM regime and then paperclipping us.

Can I predict when this will happen? No. But then I regard various attempts to predict the arrival of AGI as mostly epistemic theater. As far as I can tell, these attempts either involve asking experts to produce their best estimates (on whatever basis) or involve some method of extrapolating available compute, whether through simple Moore’s Law type extrapolation or Open Philanthropy’s heroic work on biological anchors – which, incidentally, I find interesting on its own independently of its use in predicting the arrival of "tranformative AI." But it's not like predicting when the next solar eclipse will happen (a stunt that that Yankee engineer used to fool those medieval rubes) or even predicting who'll win the next election. It's fancy guesswork.

Thursday, July 7, 2022

No one had predicted GPT-3. How do you update your priors? [Why learning from history is hard]

More from the Pinker/Aaronson debate on AI scale.

Aaronson at comment #240:

It’s true that I utterly failed to predict the deep learning revolution. I was certainly aware of the thesis, which I associated with Ray Kurzweil, that before long Moore’s Law would cause machines to have as many computing cycles as the human brain, and at that point we should expect human-level AI to “just miraculously emerge.” That struck me as one of the stupidest theses I’d ever heard! Computing cycles aren’t just magical pixie dust, I’d explain: you’d also need a whole research effort, which could take who knows how many centuries or millennia, to figure out what to do with the cycles!

Now it turns out that the thesis was … well, we still don’t know if it’s right all the way to AGI, and certainly great new ideas (GANs, transformer models, etc.) have also played a role, but it’s now clear that the “computing cycles as magic pixie-dust” thesis contained more rightness than almost anyone imagined back in 2000.

So, this is my excuse: I’m not contradicting myself (which is bad), I’m updating based on new evidence (which is good).

But my real excuse is that hardly any of the experts predicted this either. And I just had dinner with Eliezer a couple weeks ago, and he told me that he didn’t predict it. He was worried about AGI in general, of course, but not about the pure scaling of ML. The spectacular success of the latter has now, famously, caused him to say that we’re doomed; the timelines are even shorter than he’d thought.

While it caused Eliezer to update from “we should all worry about this” to “screw it, we’re doomed,” it caused me and quite a few others to update from “we shouldn’t all worry about this” to “we should all worry about this.”

Me at comment #280 after quoting from Scott’s #240:

It caused me to update from “the space of possible minds is huge” to “the space of possible minds is even larger than I thought it was.” My update is different from yours, but doesn’t necessarily contradict it. More like orthogonal to it.

This language of “updating” comes from Bayesian statistics, which I do not know on a technical level. But then, in this kind of context, it is not used technically. This usage is ubiquitous in the so-called rationalist community.

Roughly speaking, you have some idea of what’s going on in some domain, in this case, artificial intelligence. That idea is your prior and implicitly entails predictions about how that domain will unfold over time. If things unfold in a way that is consistent with your prior views, then things won’t surprise you. When something surprising happens, though, that’s a signal that your priors are wrong. You must now adjust your priors. That’s what Aaronson is talking about in the last three paragraphs I quoted from him and what I’m talking about in my paragraph. Now, while Baysianism tells you to update your priors, it doesn’t tell you just how to update your priors.

I continue my comment with my now standard analogy for dealing with large complex problems, Christmas tree lights:

I’m a bit more interested in understanding the brain than I am in scaling Mount AGI. Here’s how I’ve been thinking about understanding the brain. Imagine that understanding means takes the form of a string of serial-wired Christmas tree lights, 10,000 of them. To consider the problem solved all the lights have to be good so that the string lights up.

Instead of understanding the brain, apply the analogy to understanding how to create AGI (whatever that is). Let’s start at 1956, the year of the Dartmouth conference. It’s at that point we were handed the string and were told, “get this to light up and you’ve solved AI.” Since digital computing had been a going concern for over a decade at that point and work had already been done on chess and on natural language, some of the bulbs in that 10,000 bulb string were good. But we did’t know how many or where they were. Between 1956 and whenever OpenAI started working on GPT-3 we’d replaced, say, 2037 bad bulbs with good ones. Let’s say that in creating GPT-3 OpenAI replaced 100 bulbs, which we know about. So 2137 bulbs have been replaced. How many more bulbs to go before all of them are good?

Obviously we don’t know. Some people seem to think it’s only a couple of hundred or so, most of them having to do with scaling up even further. What, beyond wishful thinking, justifies both the belief that the unknowns cluster in one area and that their number is so low? Maybe we still have over 5000 or 6000 to go, maybe more. Who knows?

That is to say, it’s one thing to adjust your priors by thinking we’re now at long last on the right track and quite different to think, as I did, the world just got much larger. Those different updates reflect, in effect, two different sets of priors.

I go on to say something about where my priors come from:

By way of calibration, I should note that back in the era of symbolic computing I had once felt – though never published – that we were within 20 years of being able to build a system that read Shakespeare plays in an “interesting” way. By “interesting” I meant that we could have the system read, say, The Winter’s Tale, and then we’d open it up, trace what it did, and thereby learn what happens we humans read that play. That is to say I believed we could construct a system that could reasonably be construed as a simulation of the human mind. Alas, the AI Winter of the mid-1980s killed that dream. While these new post GPT-3 systems are wonderful, I see little prospect that any of them can be considered a simulation of the human mind nor that any of them will be able to shed insight into Shakespeare in the near or mid-term future. Beyond, say, 2140 (the year of Kim Stanley Robinson’s New York 2140) I’m not prepared to say.

My sense of such matters is that reading about such collapses of intellectual projects in a history book is not the same as living through one. The valence is much weaker. So I’m sticking with my new prior, “the space of possible minds is even larger than I thought it was.”

And THAT difference, I believe, is crucial. Living through a set of events, AI Winter, affects your priors in a way that is quite different from only knowing about those events from a historical account. In both cases we’re dealing with an ongoing stream of events, the evolution of AI research from the 1950s into the present and on to the future. I suspect the difference can ultimately be traced to the brain. What someone reads about events simply does not affect “deep circuitry” the same way as experiencing those same events, even if one really really believes what was read. 

More generally, this is one reason that learning from history is so very difficult. What you’ve lived through is much more potent, has a greater effect on your updating mechanism, that what you’ve only read about or heard from third parties. Everyone’s experience is necessarily limited. If there is any wisdom that does indeed accrue to age, this would be one source of it. But there is obviousy a limit to how long one person can live.

I’ve discussed this before, in particular, in a post from 2021, Things change, but sometimes they don’t: On the difference between learning about and living through [revising your priors and the way of the world].

Friday, July 1, 2022

Just what is intelligence, anyhow? [Is it simply a matter of scale?]

The Aaronson/Pinker debate on AI scaling generated a lot of commentary, including some from me and some from NYU’s Ernie Davis, who works closely with Gary Marcus. I’ve gathered some of those together in this post. But first....

What’s interesting is that that definition defines intelligence as a relation between some device (natural or artificial) and the environment in which it operates. That relationship has been dogging AI for some time.

Moravec’s paradox

Here is my first contribution to the debate (comment #81):

There is a song lyric, "Fools rush in, where angels fear to tread." Call me a fool.

Scott #33:

...stepping back: my exchanges with you, Steve, and others have been useful for me, in clarifying how “the power or powerlessness of pure intellectual ability to shape the world” is really at the heart of the entire AGI debate.

Well, yes, though the first time I read that I gave it a very reductive reading where "pure intellectual ability" was something like "computational horsepower". However, the relationship between computational horsepower and pure intellectual ability (whatever that might be) is at best unspecified. However, computational horsepower is certainly at the center of current debats about scaling. And it's quite clear that the abundance of relatively cheap compute has been extraordinarily important.

Take chess, which has been at the center of AI since before the 1956 Dartmouth conference. Chess is a rather special kind of problem. From an abstract point of view it is no more difficult that tic-tac-toe. Both are finite games played on a very simple physical platform. However, the chess-tree is so very much larger than the tic-tac-toe tree that playing the game is challenging for even the most practiced adults, while tic-tac-toe challenges no one over the age of, what? seven?

However, the fact that the chess tree is generated from a relatively simple basic structure (on 64 squares, 32 pieces, highly restrictive rules) means that compute can be thrown at the problem in a relatively straight-forward way. And the availability of compute has been important in the conquest of chess. It's certainly not the only thing, but without it, we'd be stuck where we were well before Big Blue beat Kasparov.

In contrast, things like image recognition, machine translation, or common sense knowledge, those are quite different in character from chess. The number of possible images is unbounded and they're in all forms. Language, the number of word types may be finite, but it's not well-defined, and the number of different texts is unbounded. Common sense, the same. Throwing more and more compute at the problem helps, but computational approaches to those problems, and others like them, has not produced computers that perform at the Kasparov level, and better, in those respective domains.

This has been known for a long time, it has a name, Moravec’s paradox. I think we should keep it in mind during these discussions.

Note that Moravec’s paradox is about the nature of the environment in which computation is tasked with achieving goals. Some environments are more amenable to computational regimes we understand than others.

Ernie Davis on computational speed and intelligence

Here is his comment #18, in full:

Let me suggest the following thought experiment. Suppose we take some mediocre, stick-in-the-mud scientist from 1910 who rejected not just special relativity but also atomic theory, the kinetic theory of heat, and Darwinian evolution — there were, of course, quite a few such. Now speed him up by a factor of 1000. One’s intuition is that result would be thousands of mediocre papers, and no great breakthroughs. On the other hand, it doesn’t seem right to say that Einstein, Planck and so on were 1000 times more intelligent than him; in terms of measures like IQ, they may not have been at all less intelligent than him. So I am really doubtful that this speeding up process has much to do with genius in the sense of Einstein et al. And therefore I think your intuition about speeding up Einstein by a factor of 1000 is also wrong. Had we speeded up Einstein by a factor of 1000 during his lifetime starting in 1905, we might have gotten the great papers of 1905 within a day (as fast as he could physically write them) and general relativity within a week, (ignoring the fact that that involved interactions with non-speeded up people) but I don’t think you can be confident about how much more we would have gotten.

And some passages from his comment #24:

On the last point: I think that the terminology does matter, because the view that “intelligence” is a well-defined, scalar, characteristic of minds, shown in its highest degree by people of exceptional intellectual accomplishment, is an error, and not an innocuous one. There is really very little reason to think that the qualities of mind that made Jane Austen exceptional had anything at all in common with the quality of mind that made Ramanujan exceptional; or the qualities of mind that made Chopin, Emily Dickinson, William James, or Rachel Carson exceptional. [...]

Of course, if you take all of human history and, so to speak, videotape it and then run the video tape at 1000 x speed, then things happen 1000 times as fast. So what?

Indeed, so what?

What if searching for ideas is like searching for diamonds?

This is an idea I explored more extensively in a post from 2020, Stagnation, Redux: It’s the way of the world [good ideas are not evenly distributed, no more so than diamonds]. I subsequently incorporated that post into a working paper, What economic growth and statistical semantics tell us about the structure of the world.

Comment #108:

I would like to elaborate on the comment Ernie Davis made at #18, because I suspect he’s correct. I suspect that 1000X Einstein would have given us his great work rather quickly but that [he] would [then] have proceeded out into the same intellectual desert the real Einstein explored, but managed to explore it much more thoroughly, with, alas, the same success.

Just how are ideas distributed in idea space? (Is that even a coherent question?)

Let me suggest an analogy, diamonds. We know that they are not evenly distributed on or near the earth’s surface. Most of them seem to be in kimberlite (a type of rock) and that’s where diamond mines are located. Even there, they are few, far between, and irregularly located. So it takes a great deal of labor to find each diamond.

Now, imagine we have a robot that can find diamonds at 1000 times the rate human miners can, but only costs, say, 10 times or even 100 more times per hour. Such robots would be very valuable. Now, let’s place a bunch of these 1000X robots on some arbitrary chunk of land and let them dig and sort away. What are they going to find? Probably nothing. Why, because there are no diamonds there. They may be very good at excavating, moving, crushing, and sorting through earth, but if there are no diamonds there, the effort is wasted.

Perhaps ideas and idea space are like that. The ideas are unevenly distributed. We have no maps to guide us to them. But we have theories, and hunches, an intellectual style. Think of them collectively as a mapping procedure. So, Einstein had his intellectual style, his mapping procedure. That led to roughly a decade of important discoveries in his 20s and 30s, like diamond miners working in kimberlite. And then, nothing, like diamond miners working, say, in the middle of Vermont. Nice country, but no diamonds.

As for idea space, we can imagine it by analogy with chess space. But we know how to construct chess space, though it is too large for anything approaching a complete construction. And that knowledge allows us to construct useful procedures for searching it. We haven’t a clue about how to construct idea space, much less how to search it effectively. If speed is all we’ve got, it’s not clear how much that gets us in the general case.

It’s not at all obvious that we need the notion of idea space in the case of Einstein, and similar cases. Einstein’s just searching the world for a fit between his best thinking and natural phenomena. Chess space, of course, is different. It is entirely artificial; we created it when we created the game. The world Einstein explored pre-existed him (and us).

Five factors of genius/intelligence

Further response to Davis, # 18:

Suppose we take some mediocre, stick-in-the-mud scientist from 1910 who rejected not just special relativity but also atomic theory, the kinetic theory of heat, and Darwinian evolution — there were, of course, quite a few such. Now speed him up by a factor of 1000. One’s intuition is that result would be thousands of mediocre papers, and no great breakthroughs. On the other hand, it doesn’t seem right to say that Einstein, Planck and so on were 1000 times more intelligent than him; in terms of measures like IQ, they may not have been at all less intelligent than him.

Speed is one thing. And IQ is another. Einstein had something else. I suppose we could call it genius, in fact we do, don’t we? But that doesn’t tell us much.

For the sake of argument – I’m just making this up as I type – let’s say one aspect of that something else is intellectual technique. Einstein had more effective intellectual tactics and strategies than those standard investigators. Intellectual technique may, in turn, have a genetic aspect that’s not covered by IQ, but almost certainly has a learned aspect as well.

So now we have four things: 1) speed/compute, 2) IQ, 3) an inherited component of technique, and 4) a learned component of technique.

I’m going to posit one more thing, again, thinking off the top of my head. We might call it luck. Or, if we’re thinking in terms of something like idea space, we could call it initial position. By virtue of 1, 2, 3 and perhaps 4 as well, the so-called genius is at a position in idea space that allows them to make major discoveries by deploying their cumulative capabilities. The point of this initial-position factor is to allow for the possibility of a cohort of thinkers more or less equally endowed with 1,2,3+4, but having very different initial position. As a consequence, some are able to achieve major discoveries quickly, while others take more time, and still others never get there. Their capabilities are comparable, but their outcomes are not.

To invoke the diamond mining metaphor I introduced in comment #108, we have two equally skilled geologists/prospectors. One just happens to be located within 100 miles of a major kimberlite deposit while the other is over 3000 miles away from such a deposit. If they start walking from where they are, who’s going to find diamonds first?

In the case of AI, we know a great deal about compute/speed; we have that under control. I’m not sure just how the distinction between innate vs. learned techniques applies to machines, perhaps hardware and software. In any case, we do have a large repertoire of techniques of various kinds. In some areas we can produce a combination of compute and technique that allows the machine to outperform the best human. In other areas we have machines that do things that are amazing in comparison with what machines did, say, a decade ago, but which are no more than standard human performances, with various failings here and there. And so on. As for starting position, I think it’s up to us to position the AI properly, at least at the start.

[But once and if it FOOMs, it’s on its own. I’m not holding my breath on this one.]

Figure 5 in the 2020 New Savanna post gives a visual illustration of the initial position idea.

We now have a total of five factors:

1) compute,
2) IQ,
3) an inherited component of technique,
4) a learned component of technique, and
5) initial position or luck.

The seductiveness of scale

The hope of the scaling side of the current debate is that we can get all the way to AGI – whatever that is – by throwing more compute at the problem. Well, it’s not that simple, the compute has to be channeled through an appropriate machine-learning architecture which then chews its way to a huge pile of (appropriately curated) data. That is, architecture+data will cover the ground I’ve indicated in factors 2-5 above, thereby relieving us of the need and responsibility to think about those things.

It’s a seductive prospect. Why? In part because it is easy to understand, even by people who have little or no technical knowledge of computing, cognitive science, and AI. Everyone knows and understands, “bigger is better.”

GOFAI (good old fashioned artificial intelligence) was mostly about technique, factors 3 and 4. That technique was generally taken to be mediated by symbolic systems and, as a practical matter, it required that ‘knowledge’ be painstakingly hand-crafted into systems. While I can understand the desire to avoid hand-crafted knowledge – there’s so very much of it and the crafting is tedious and error-prone — I don’t think symbolic computation can be avoided. Can it be architected, as it were, into a learning regime? We know one case where it has been, the human case, but that case tells us that learning requires a lot of close interaction between teachers and students, in both formal and informal settings. It’s not at all clear to me that such interaction can be architected.

More later. 

Addendum, 7.1.22, on superintelligence: Alex, comment #171:

I think Pinker’s definition of intelligence, “the ability to use information to attain a goal in an environment”, is reasonable, but it doesn’t give us any meaningful way to compare or rank intelligences (so how can we meaningfully discuss “superintelligence”?). Of course, you chose compute time as the metric, but I think that dodges the more meaningful aspects of intelligence. I think a metric like computational complexity – or even Kolmogorov complexity – is more appealing to me, but whatever the metric, I think it has to capture the mechanism of thought in some way, not just the output. [...]

As a final note, I think “intelligence” is a crude word that tries to capture too many aspects of behavior (many of them human-relatable, but not of great importance to discussion). My comment here has been an attempt to break up “intelligence” into constituent parts to focus discussion: clock speed, memory, algorithmic/time complexity, size/space complexity. There are surely more parts of “intelligence”, some parts that are combinations of simpler parts.

Scott, comment #172:

Fundamentally, I care, not about the definitions of words like “superintelligence,” but about what will actually happen in the real world once AIs become much more powerful. [...] So OK then, what happens when we can launch a billion processes in datacenters, each one with the individual insight of a Terry Tao or Edward Witten (or the literary talent of Philip Roth, or the musical talent of the Beatles…), and they can all communicate with one another, and they can work at superhuman speed? Is it not obvious that all important intellectual and artistic production shifts entirely to AIs, with humans continuing to engage in it (if they do) only as a hobby? That’s the main question I care about when I discuss “superintelligence,” and I’m still waiting for anyone to explain why I’m wrong about it.

Thursday, June 30, 2022

Steven Pinker and Scott Aaronson debate scaling

Scott Aaronson has hosted Steven Pinker to a discussion at Shtetl-Optimized.

Pinker on AGI:

Regarding the second, engineering question of whether scaling up deep-learning models will “get us to Artificial General Intelligence”: I think the question is probably ill-conceived, because I think the concept of “general intelligence” is meaningless. (I’m not referring to the psychometric variable g, also called “general intelligence,” namely the principal component of correlated variation across IQ subtests. This is a variable that aggregates many contributors to the brain’s efficiency such as cortical thickness and neural transmission speed, but it is not a mechanism (just as “horsepower” is a meaningful variable, but it doesn’t explain how cars move.) I find most characterizations of AGI to be either circular (such as “smarter than humans in every way,” begging the question of what “smarter” means) or mystical—a kind of omniscient, omnipotent, and clairvoyant power to solve any problem. No logician has ever outlined a normative model of what general intelligence would consist of, and even Turing swapped it out for the problem of fooling an observer, which spawned 70 years of unhelpful reminders of how easy it is to fool an observer.

If we do try to define “intelligence” in terms of mechanism rather than magic, it seems to me it would be something like “the ability to use information to attain a goal in an environment.” (“Use information” is shorthand for performing computations that embody laws that govern the world, namely logic, cause and effect, and statistical regularities. “Attain a goal” is shorthand for optimizing the attainment of multiple goals, since different goals trade off.) Specifying the goal is critical to any definition of intelligence: a given strategy in basketball will be intelligent if you’re trying to win a game and stupid if you’re trying to throw it. So is the environment: a given strategy can be smart under NBA rules and stupid under college rules.

Since a goal itself is neither intelligent or unintelligent (Hume and all that), but must be exogenously built into a system, and since no physical system has clairvoyance for all the laws of the world it inhabits down to the last butterfly wing-flap, this implies that there are as many intelligences as there are goals and environments. There will be no omnipotent superintelligence or wonder algorithm (or singularity or AGI or existential threat or foom), just better and better gadgets.

Aaronson responds:

Basically, one side says that, while GPT-3 is of course mind-bogglingly impressive, and while it refuted confident predictions that no such thing would work, in the end it’s just a text-prediction engine that will run with any absurd premise it’s given, and it fails to model the world the way humans do. The other side says that, while GPT-3 is of course just a text-prediction engine that will run with any absurd premise it’s given, and while it fails to model the world the way humans do, in the end it’s mind-bogglingly impressive, and it refuted confident predictions that no such thing would work.

Though I’m with Pinker on the definition of AGI, I also take the second of the two positions Aaronson set forth, which is, I take it, Aaronson’s position while the first is Pinker’s position. That’s why I wrote GPT-3: Waterloo or Rubicon? Here be Dragons (Version 4.1).

Aaronson continues:

I freely admit that I have no principled definition of “general intelligence,” let alone of “superintelligence.” To my mind, though, there’s a simple proof-of-principle that there’s something an AI could do that pretty much any of us would call “superintelligent.” Namely, it could say whatever Albert Einstein would say in a given situation, while thinking a thousand times faster. Feed the AI all the information about physics that the historical Einstein had in 1904, for example, and it would discover special relativity in a few hours, followed by general relativity a few days later. Give the AI a year, and it would think … well, whatever thoughts Einstein would’ve thought, if he’d had a millennium in peak mental condition to think them.

If nothing else, this AI could work by simulating Einstein’s brain neuron-by-neuron—provided we believe in the computational theory of mind, as I’m assuming we do. It’s true that we don’t know the detailed structure of Einstein’s brain in order to simulate it [...]. But that’s irrelevant to the argument. It’s also true that the AI won’t experience the same environment that Einstein would have—so, alright, imagine putting it in a very comfortable simulated study, and letting it interact with the world’s flesh-based physicists. A-Einstein can even propose experiments for the human physicists to do—he’ll just have to wait an excruciatingly long subjective time for their answers. But that’s OK: as an AI, he never gets old.

Next let’s throw into the mix AI Von Neumann, AI Ramanujan, AI Jane Austen, even AI Steven Pinker—all, of course, sped up 1,000x compared to their meat versions, even able to interact with thousands of sped-up copies of themselves and other scientists and artists. Do we agree that these entities quickly become the predominant intellectual force on earth—to the point where there’s little for the original humans left to do but understand and implement the AIs’ outputs (and, of course, eat, drink, and enjoy their lives, assuming the AIs can’t or don’t want to prevent that)?

Eh. Now that I have an explicit definition of artificial minds, I have no need for a definition of artificial intelligence. While my primer (Relational Nets Over Attractors, A Primer: Part 1, Design for a Mind) is mostly about the human mind and the human brain, the fact that I was able to propose a substrate-neutral definition of “mind” has the side-effect that I can talk about artificial minds as mechanisms, not magic, to use Pinker’s formulation.

Aaronson also notes:

I should clarify that, in practice, I don’t expect AGI to work by slavishly emulating humans—and not only because of the practical difficulties of scanning brains, especially deceased ones. Like with airplanes, like with existing deep learning, I expect future AIs to take some inspiration from the natural world but also to depart from it whenever convenient. The point is that, since there’s something that would plainly count as “superintelligence,” the question of whether it can be achieved is therefore “merely” an engineering question, not a philosophical one.

That is consistent with the view I have articulated in the primer.

Aaronson has more to say, as does Pinker. As of this moment, the dialog has attracted 100 comments (including two from me). It’s worth exploring.