Wednesday, July 6, 2022

Green Onions, Korean style

For comparison, here's the original, from 1962:

One of these days I'll have something to say about the fact that a Korean woman, Luna Lee, in the 21st century has devoted so much time an energy to performing classic rock and pop from the 1960s through the 1980s and later on a traditional Korean instrument, the gayageum, but I really don't know what to say. Something about the power of that music and East and West and, well, just what? She's from Seoul, South Koren, and is now living in Houston.

Some flowers in orange and yellow

Tuesday, July 5, 2022

Virtual Reading: The Prospero Project Redux [#DH]

I'm bumping this 2017 post to the top of the queue because, 1) I think the concept of virtual reading proposed here may be of some use in thinking about and evaluating the written output of large language models, such as GPT-3, and 2) the concept of literary form implicit the section, "In search of a small-world net," is relevant to my arguments about the value of symbols as being, in part, a vehicle for moving about in mental space in a way that "outside" the "standard" landscape of mental space (see my post earlier today, Why Are Symbols So Useful to Us?).
 
* * * * *
 
I've uploaded another working paper. Title above, abstract, table of contents, and introduction below. Note that it's a long way through the introduction, but there's some good stuff there.

Download at:

* * * * *
Abstract: Virtual reading is proposed as a computational strategy for investigating the structure of literary texts. A computer ‘reads’ a text by moving a window N-words wide through the text from beginning to end and follows the trajectory that window traces through a high-dimensional semantic space computed for the language used in the text. That space is created by using contemporary corpus-based machine learning techniques. Virtual reading is compared and contrasted with a 40 year old proposal grounded in the symbolic computation systems of the mid-1970s. High-dimensional mathematical spaces are contrasted with the standard spatial imagery employed in literary criticism (inside and outside the text, etc.). The “manual” descriptive skills of experienced literary critics, however, are essential to virtual reading, both for purposes of calibration and adjustment of the model, and for motivating low-dimensional projection of results. Examples considered: Heart of Darkness, Much Ado About Nothing, Othello, The Winter’s Tale.
Contents

Introduction: Prospero Redux and Virtual Reading 2
In search of a small-world net: Computing an emblem in Heart of Darkness 8
Virtual reading as a path through a multidimensional semantic space 11
Reply to a traditional critic about computational criticism: Or, It’s time to escape the prison-house of critical language [#DH] 17
After the thrill is gone...A cognitive/computational understanding of the text, and how it motivates the description of literary form [Description!] 23
Appendix: Prospero Elaborated 30

Introduction: Prospero Redux and Virtual Reading

In a way, this working paper is a reflection on four decades of work in the study of language, mind, and literature. Not specifically my work, though, yes, certainly including my work. I say in a way, for it certainly doesn’t attempt to survey the relevant literature, which is huge, well beyond the scope of a single scholar. Rather I compare a project I had imagined back then (Prospero), mostly as a thought experiment, but also with some hope that it would in time be realized, with what has turned out to be a somewhat revised version of that project (Prospero Redux), a version which I believe to be doable, though I don’t alone posess the skills, much less the resources, to do it.

The rest of this working paper is devoted to Prospero Redux, the revised version. This introduction compares it with the 40 year-old Prospero. This comparison is a way of thinking about an issue that’s been on my mind for some time: Just what have we learned in the human sciences over the last half-century or so? As far as I can tell, there is no single theoretical model on which a large majority of thinkers agree in the way that all biologists agree on evolution. The details are much in dispute, but there is no dispute that world of living things is characterized by evolutionary dynamics. The human sciences have nothing comparable (though there is a move afoot to adopt evolution as a unifying principle for the social and behavioral sciences). If we don’t have even ONE such theoretical model, just what DO we know? And yet there HAS been a lot of interesting and important work over the last half-century. We must have learned something, no?

Let’s take a look.

Prospero, 1976

Work in machine translation started in the early 1950s [1]; George Miller published his classic article, “The Magical Number Seven, Plus or Minus Two” in 1956; Chomsky published Syntactic Structures in 1957; and we can date artificial intelligence (AI) to a 1956 workshop at Dartmouth [2]. That’s enough to characterize the beginnings of the so-called “Cognitive Revolution” in the human sciences. I encountered that revolution, if you will, during my undergraduate years at Johns Hopkins in the 1960s, where I also encountered semiotics and structuralism. By the early 1970s I was in graduate school in the English Department at The State University of New York at Buffalo, where I joined the research group of David Hays in the Linguistics Department. Hays was a Harvard-educated cognitive scientist who’d headed the mamachine translation program at the RAND Corporation in the 1950s.

At that time a number of reasearch groups were working on cognitive or semantic network models for natural language semantics. It was bleeding edge research at the time. I learned the model Hays and his students had developed and applied it to Shakespeare’s Sonnet 129 (which I touch on a bit later, pp. 21 ff.). At the same time I was preparing abstracts of the current literature in computational linguistics for The American Journal of Computational Linguistics. Hays edited the journal and had a generous sense of the relevant literature.

Thus when Hays was invited to review the field of computational linguistics for Computers and the Humanities it was natural for him to ask me to draft the article. I wrote up the standard kind of review material, including reports and articles coming out on the Defense Department’s speech understanding project, which was perhaps the single largest research effort in the field (I discuss this as well, pp. 20 ff.). But we aspired to more than just a literature review. We wanted a forward-looking vision, something that might induce humanists to look deeper into the cognitive sciences.

We ended the article with a thought experiment (p. 271):
Let us create a fantasy, a system with a semantics so rich that it can read all of Shakespeare and help in investigating the processes and structures that comprise poetic knowledge. We desire, in short, to reconstruct Shakespeare the poet in a computer. Call the system Prospero.

How would we go about building it? Prospero is certainly well beyond the state of the art. The computers we have are not large enough to do the job and their architecture makes them awkward for our purpose. But we are thinking about Prospero now, and inviting any who will to do the same, because the blueprints have to be made before the machine can be built. [...]

The general idea is to represent the requisite world knowledge – what the poet had in his head – and then investigate the structure of the paths which are taken through that world view as we move through the object text, resolving the meaning of the text into the structure of conceptual interrelationships which is the semantic network. Thus the Prospero project includes the making of a semantic network to represent Shakespeare’s version of the Elizabethan world view.
But a model of the Elizabethan world view was “only the background”. We would also have to model Shakespeare’s mind (p. 272):
A program, our model of Shakespeare’s poetic competence, must move through the cognitive model and produce fourteen lines of text. [...] The advantage of Prospero is that it takes the cognitive model as given – clearly and precisely – and the poetic act as a motion through the model. Instead of asking how the words are related to one another, we ask how the words are related to an organized collection of ideas, and the organization of the poem is determined, then, by the world view and poetics in unison. [3]
We declined to predict when such a marvel might have been possible, though I expected to see something within my lifetime. Not something that would rival the Star Trek computer, mind you, not something that could actually think in some robust sense of the word. But something.
 
What we got some 35 years later was an IBM computer system called Watson that defeated humans in playing Jeopardy [4]. Watson was a marvel, but was and is nowhere near to doing what Hays and I had imagined for Prospero. Nor do I see that old vision coming to life in the forseeable future.

Moreover, Watson is based on newer kind of technology that is quite different from that which Hays and I had reviewed in our article and which we were imagining for Prospero. Prospero came out of a research program, symbolic computing, that all but collapsed a decade later. It was replaced by technology that had a more stochastic character, which involved machine learning, and which, in some increasingly popular versions, was (somewhat distanctly) inspired by real nervous systems. It is this newer technology that runs Google’s online machine translation system, that runs Apple’s Siri, and that is behind much of the work in computational literary criticism.

Before turning to that, however, I want to say just a bit more about what we most likely had in mind – I say “most likely” because that was a LONG time ago and I don’t remember all that was whizzing through my head at the time. We were out to simulate the human mind, to produce a system that was, in at least some of its parts and processes, like the parts and processes of the mind. One could have Prospero read and even write texts while keeping records of what it does. One could then examine those records and thus learn how the mind works. Ambitious? Yes. But the computer simulation of cognitive tasks is quite common in the cognitive sciences, though not on THAT scale. In contrast, Watson, for example, was not intended as a simulation of the mind. It was a straight-up engineering activity. What matters for such systems, and for AI generally, is whether or not the system produces useful results. Whether or not it does so in a human way is, at best, a secondary consideration.

Why didn’t Prospero, or anything like it, happen? For one thing, such systems tended to be brittle. If you get something even a little bit wrong, the whole thing collapses. Then there’s combinatorial explosion; so many alternatives have to be considered on the way to a good one that the system just runs out of time – that is, it just keeps computing and computing and computing [...] without reaching a result. That’s closely related to what is called the “common sense” problem. No text is ever complete. Something must always be inferred in order to make smooth connections between the words in the text. Humans have vast reserves of such common sense knowledge; computing systems do not. How do they get it? Hand coding – which takes time and time and time. And when the system calls on the common sense knowledge that’s been hand-coded into it, what happens? Combinatorial explosion.

The enterprise of simulating a mind through symbolic computing simply collapsed. In the case of something like Prospero I would specially add that it now seems to me that, to tell us something really useful about the mind, such a system would have to simulate the human brain. Hays and I didn’t realize it at the time – we’d just barely begun to think about the brain – but that became obvious some years later in retrospect.

The limitations of scale in deep learning

My trip into the suburbs for barbeque on the 4th

Why Are Symbols So Useful to Us? [Relational Nets]

I’ve been participating in the discussion of Yann LeCun’s recent position paper, A Path Towards Autonomous Machine Intelligence. My first comment was a long one, Why are symbols important? Because they index cognitive space.

My opening paragraph:

I want to address the issue that your raise at the very end of your paper: Do We Need Symbols for Reasoning? I think we do. Why? 1) Symbols form an index over cognitive space that, 2) facilitates flexible (aka ‘random’) access to that space during complex reasoning.

My final paragraph is addressed to that second issue:

I really should say something about how symbols facilitate flexible access to cognition, but well, that’s tricky. Let me offer up a fake example that points in the direction I’m thinking. Imagine that you’ve arrived at a local maximum in your progression toward some goal but you’ve not yet reached the goal. How do you get unstuck? The problem is, of course, well known and extensively studied. Imagine that your local maximum has a name1, and that name1 is close to name2 of some other location in the space you are searching. That other location may or may not get you closer to the goal; you won’t know until you try. But it is easy to get to name2 and then see where that puts you in the search space. If you’re not better off, well, go back to name1 and try name3. And so forth. Symbol space indexes cognitive space and provides you with an ordering over cognitive space that is different from and somewhat independent of the gradients within cognitive space. It’s another way to move around. More than that, however, it provides you with ways of constructing abstract concepts, and that’s a vast, but poorly studied subject [1].

I really need to elaborate on that. Two discussions are needed: 1) one elaborates on the hill-climbing problem I mention, and 2) the other talks about syntax.

On the first, in an unindexed neural net all inference must proceed locally. In a space with billions and billions of dimensions, locality is obviously a very tricky matter. A local move on one dimension can easily put you in touch with locations on other dimensions which had been quite distant from your starting point. Still, an index constructed within the space gives you a set of vantage points which are outside the gradient structure of the network.

Syntax is one mechanism you have for moving around in index space. That’s what the syntax discussion needs to be about, how syntactic motion in index space can make it easier to move outside the local gradients in semantic space. But not here and now.

Nor is syntax the only mechanism available to you. You could move through a simple alphabetized list of word forms. Such a path would be arbitrary with respect to the gradients in semantic space, which is to say, such a path takes you outside semantic space.

What other mechanisms are there? How does rhyme in poetry figure into this?

More later.

[1] For some thoughts on various mechanisms for constructing abstract concepts, see William Benzon and David Hays, The Evolution of Cognition, Journal of Social and Biological Structures. 13(4): 297-320, 1990, https://www.academia.edu/243486/The_Evolution_of_Cognition

Monday, July 4, 2022

Two different versions of Hoboken Peek-a-Boo

A recent victory for right-to-repair

Cory Doctorow, A Win For Harley Riders, Medium, June 26, 2022.

He opens:

Right-to-Repair is a no-brainer. You bought a thing, you want to fix it — or nominate someone else to fix it for you — and the manufacturer doesn’t. How ever can we resolve this intractable difference of opinion?

And then goes on to detail various cases where manufacturers are trying to block owners from getting their stuff repaired by anyone other than the manufacturer. He then tells us about:

In 1975, Congress passed the Magnuson–Moss Warranty Act; a complex law whose subtext can be summed up in seven words: fuck you, I bought it, it’s mine.

Specifically, Magnuson-Moss “prohibits a company from conditioning a consumer product warranty on the consumer’s using any article or service which is identified by brand name unless it is provided for free.” In other words, unless a company wants to provide free service to its customers, it can’t tell them to use its own repair facility or its affiliates’.

And concludes with a recent ruling by the FTC:

FTC chair Lina Khan is part of a trio of Biden appointees (along with Jonathan Kanter at the DoJ and Tim Wu in the White House) who are laser-focused on promoting the interest of people over corporations.

Khan’s FTC just voted 5–0 to open enforcement action against Harley-Davidson and Westinghouse, and they immediately caved, removing the illegal warranty language that threatened to take away your right to service if you dared to fix something on your own, or at a garage or depot of your choice.

The FTC has been doing amazing stuff on Right to Repair, starting with the landmark Nixing the Fix report, which documented the many ways in which companies were ripping off their customers and destroying the planet by blocking repairs. Then there was Biden’s amazing Executive Order on repair, and endorsing federal Right to Repair legislation.

But the vanquishing of Harley-Davidson and Westinghouse shows that we don’t need an EO or a bill to do a lot about repair: all we need is an FTC that’s willing to do it’s job.

Sunday, July 3, 2022

Let's cool it

What do rock guitar gods want most? Sex with women or to impress other guys?

DeLecce, T., Pazhoohi, F., Szala, A., & Shackelford, T. K. (2022). Extreme metal guitar skill: A case of male–male status seeking, mate attraction, or byproduct? Evolutionary Behavioral Sciences. Advance online publication. https://doi.org/10.1037/ebs0000304

Abstract: There has been much debate around the ultimate explanation of cultural displays such as music and art. There are two main competing hypotheses for the function of music: sexual selection or byproduct of the complexity of the human brain. Although there is evidence that playing music increases male attractiveness, the sexual selection explanation may not be mutually exclusive to all types of music. Extreme metal is a genre that is heavily male-biased, not only among the individuals that play this style of music, but also among the fans of the genre. Therefore, it is unlikely that extreme metal musicians are primarily trying to increase their mating success through their music. However, musicians in this genre heavily invest their time in building technical skills (e.g., dexterity, coordination, timing), which raises the question of the purpose behind this costly investment. It could be that men engage in this genre mainly for status-seeking purposes: to intimidate other males with their technical skills and speed and thus gain social status. To explore the reasoning behind investment in technical guitar skills, a sample of 44 heterosexual male metal guitarists was recruited and surveyed about their practicing habits (newly created survey for this study), sexual behavior (using the Sociosexual Orientation Inventory–Revised [SOI-R]; Penke & Asendorpf, 2008), and feelings of competitiveness toward the same sex (via the Intrasexual Competition Scale [ICS]; Buunk & Fisher, 2009). The survey results indicated that time spent playing chords predicted desire for casual sex with women whereas perceptions of playing speed positively predicted intrasexual competitiveness (a desire to impress other men). The discussion addresses how these results, and the extreme metal genre, might relate to the three competing hypotheses for the function of cultural displays. (PsycInfo Database Record (c) 2022 APA, all rights reserved)

H/t Tyler Cowen.

3-D printing comes of industrial age

Steve Lohr, 3-D Printing Grows Beyond Its Novelty Roots, NYTimes, July 3, 2022.

DEVENS, Mass. — The machines stand 20 feet high, weigh 60,000 pounds and represent the technological frontier of 3-D printing.

Each machine deploys 150 laser beams, projected from a gantry and moving quickly back and forth, making high-tech parts for corporate customers in fields including aerospace, semiconductors, defense and medical implants.

The parts of titanium and other materials are created layer by layer, each about as thin as a human hair, up to 20,000 layers, depending on a part’s design. The machines are hermetically sealed. Inside, the atmosphere is mainly argon, the least reactive of gases, reducing the chance of impurities that cause defects in a part.

The 3-D-printing foundry in Devens, Mass., about 40 miles northwest of Boston, is owned by VulcanForms, a start-up that came out of the Massachusetts Institute of Technology. It has raised $355 million in venture funding. And its work force has jumped sixfold in the past year to 360, with recruits from major manufacturers like General Electric and Pratt & Whitney and tech companies including Google and Autodesk.

“We have proven the technology works,” said John Hart, a co-founder of VulcanForms and a professor of mechanical engineering at M.I.T. “What we have to show now is strong financials as a company and that we can manage growth.”

For 3-D printing, whose origins stretch back to the 1980s, the technology, economic and investment trends may finally be falling into place for the industry’s commercial breakout, according to manufacturing experts, business executives and investors.

Additive manufacturing:

3-D printing refers to making something from the ground up, one layer at a time. Computer-guided laser beams melt powders of metal, plastic or composite material to create the layers. In traditional “subtractive” manufacturing, a block of metal, for example, is cast and then a part is carved down into shape with machine tools.

In recent years, some companies have used additive technology to make specialized parts. General Electric relies on 3-D printing to make fuel nozzles for jet engines, Stryker makes spinal implants and Adidas prints latticed soles for high-end running shoes. Dental implants and teeth-straightening devices are 3-D printed. During the Covid-19 pandemic, 3-D printers produced emergency supplies of face shields and ventilator parts.

Today, experts say, the potential is far broader than a relative handful of niche products. The 3-D printing market is expected to triple to nearly $45 billion worldwide by 2026, according to a report by Hubs, a marketplace for manufacturing services.

The Biden administration is looking to 3-D printing to help lead a resurgence of American manufacturing. Additive technology will be one of “the foundations of modern manufacturing in the 21st century,” along with robotics and artificial intelligence, said Elisabeth Reynolds, special assistant to the president for manufacturing and economic development.

There's more at the link.

And when each 3-D printer is equipped with a powerful artificial mind, and they communicate with one another....?

Friday, July 1, 2022

Carefully curating your data makes for more efficient machine learning

Later in the stream:

Just what is intelligence, anyhow? [Is it simply a matter of scale?]

The Aaronson/Pinker debate on AI scaling generated a lot of commentary, including some from me and some from NYU’s Ernie Davis, who works closely with Gary Marcus. I’ve gathered some of those together in this post. But first....

What’s interesting is that that definition defines intelligence as a relation between some device (natural or artificial) and the environment in which it operates. That relationship has been dogging AI for some time.

Moravec’s paradox

Here is my first contribution to the debate (comment #81):

There is a song lyric, "Fools rush in, where angels fear to tread." Call me a fool.

Scott #33:

...stepping back: my exchanges with you, Steve, and others have been useful for me, in clarifying how “the power or powerlessness of pure intellectual ability to shape the world” is really at the heart of the entire AGI debate.

Well, yes, though the first time I read that I gave it a very reductive reading where "pure intellectual ability" was something like "computational horsepower". However, the relationship between computational horsepower and pure intellectual ability (whatever that might be) is at best unspecified. However, computational horsepower is certainly at the center of current debats about scaling. And it's quite clear that the abundance of relatively cheap compute has been extraordinarily important.

Take chess, which has been at the center of AI since before the 1956 Dartmouth conference. Chess is a rather special kind of problem. From an abstract point of view it is no more difficult that tic-tac-toe. Both are finite games played on a very simple physical platform. However, the chess-tree is so very much larger than the tic-tac-toe tree that playing the game is challenging for even the most practiced adults, while tic-tac-toe challenges no one over the age of, what? seven?

However, the fact that the chess tree is generated from a relatively simple basic structure (on 64 squares, 32 pieces, highly restrictive rules) means that compute can be thrown at the problem in a relatively straight-forward way. And the availability of compute has been important in the conquest of chess. It's certainly not the only thing, but without it, we'd be stuck where we were well before Big Blue beat Kasparov.

In contrast, things like image recognition, machine translation, or common sense knowledge, those are quite different in character from chess. The number of possible images is unbounded and they're in all forms. Language, the number of word types may be finite, but it's not well-defined, and the number of different texts is unbounded. Common sense, the same. Throwing more and more compute at the problem helps, but computational approaches to those problems, and others like them, has not produced computers that perform at the Kasparov level, and better, in those respective domains.

This has been known for a long time, it has a name, Moravec’s paradox. I think we should keep it in mind during these discussions.

Note that Moravec’s paradox is about the nature of the environment in which computation is tasked with achieving goals. Some environments are more amenable to computational regimes we understand than others.

Ernie Davis on computational speed and intelligence

Here is his comment #18, in full:

Let me suggest the following thought experiment. Suppose we take some mediocre, stick-in-the-mud scientist from 1910 who rejected not just special relativity but also atomic theory, the kinetic theory of heat, and Darwinian evolution — there were, of course, quite a few such. Now speed him up by a factor of 1000. One’s intuition is that result would be thousands of mediocre papers, and no great breakthroughs. On the other hand, it doesn’t seem right to say that Einstein, Planck and so on were 1000 times more intelligent than him; in terms of measures like IQ, they may not have been at all less intelligent than him. So I am really doubtful that this speeding up process has much to do with genius in the sense of Einstein et al. And therefore I think your intuition about speeding up Einstein by a factor of 1000 is also wrong. Had we speeded up Einstein by a factor of 1000 during his lifetime starting in 1905, we might have gotten the great papers of 1905 within a day (as fast as he could physically write them) and general relativity within a week, (ignoring the fact that that involved interactions with non-speeded up people) but I don’t think you can be confident about how much more we would have gotten.

And some passages from his comment #24:

On the last point: I think that the terminology does matter, because the view that “intelligence” is a well-defined, scalar, characteristic of minds, shown in its highest degree by people of exceptional intellectual accomplishment, is an error, and not an innocuous one. There is really very little reason to think that the qualities of mind that made Jane Austen exceptional had anything at all in common with the quality of mind that made Ramanujan exceptional; or the qualities of mind that made Chopin, Emily Dickinson, William James, or Rachel Carson exceptional. [...]

Of course, if you take all of human history and, so to speak, videotape it and then run the video tape at 1000 x speed, then things happen 1000 times as fast. So what?

Indeed, so what?

What if searching for ideas is like searching for diamonds?

This is an idea I explored more extensively in a post from 2020, Stagnation, Redux: It’s the way of the world [good ideas are not evenly distributed, no more so than diamonds]. I subsequently incorporated that post into a working paper, What economic growth and statistical semantics tell us about the structure of the world.

Comment #108:

I would like to elaborate on the comment Ernie Davis made at #18, because I suspect he’s correct. I suspect that 1000X Einstein would have given us his great work rather quickly but that [he] would [then] have proceeded out into the same intellectual desert the real Einstein explored, but managed to explore it much more thoroughly, with, alas, the same success.

Just how are ideas distributed in idea space? (Is that even a coherent question?)

Let me suggest an analogy, diamonds. We know that they are not evenly distributed on or near the earth’s surface. Most of them seem to be in kimberlite (a type of rock) and that’s where diamond mines are located. Even there, they are few, far between, and irregularly located. So it takes a great deal of labor to find each diamond.

Now, imagine we have a robot that can find diamonds at 1000 times the rate human miners can, but only costs, say, 10 times or even 100 more times per hour. Such robots would be very valuable. Now, let’s place a bunch of these 1000X robots on some arbitrary chunk of land and let them dig and sort away. What are they going to find? Probably nothing. Why, because there are no diamonds there. They may be very good at excavating, moving, crushing, and sorting through earth, but if there are no diamonds there, the effort is wasted.

Perhaps ideas and idea space are like that. The ideas are unevenly distributed. We have no maps to guide us to them. But we have theories, and hunches, an intellectual style. Think of them collectively as a mapping procedure. So, Einstein had his intellectual style, his mapping procedure. That led to roughly a decade of important discoveries in his 20s and 30s, like diamond miners working in kimberlite. And then, nothing, like diamond miners working, say, in the middle of Vermont. Nice country, but no diamonds.

As for idea space, we can imagine it by analogy with chess space. But we know how to construct chess space, though it is too large for anything approaching a complete construction. And that knowledge allows us to construct useful procedures for searching it. We haven’t a clue about how to construct idea space, much less how to search it effectively. If speed is all we’ve got, it’s not clear how much that gets us in the general case.

It’s not at all obvious that we need the notion of idea space in the case of Einstein, and similar cases. Einstein’s just searching the world for a fit between his best thinking and natural phenomena. Chess space, of course, is different. It is entirely artificial; we created it when we created the game. The world Einstein explored pre-existed him (and us).

Five factors of genius/intelligence

Further response to Davis, # 18:

Suppose we take some mediocre, stick-in-the-mud scientist from 1910 who rejected not just special relativity but also atomic theory, the kinetic theory of heat, and Darwinian evolution — there were, of course, quite a few such. Now speed him up by a factor of 1000. One’s intuition is that result would be thousands of mediocre papers, and no great breakthroughs. On the other hand, it doesn’t seem right to say that Einstein, Planck and so on were 1000 times more intelligent than him; in terms of measures like IQ, they may not have been at all less intelligent than him.

Speed is one thing. And IQ is another. Einstein had something else. I suppose we could call it genius, in fact we do, don’t we? But that doesn’t tell us much.

For the sake of argument – I’m just making this up as I type – let’s say one aspect of that something else is intellectual technique. Einstein had more effective intellectual tactics and strategies than those standard investigators. Intellectual technique may, in turn, have a genetic aspect that’s not covered by IQ, but almost certainly has a learned aspect as well.

So now we have four things: 1) speed/compute, 2) IQ, 3) an inherited component of technique, and 4) a learned component of technique.

I’m going to posit one more thing, again, thinking off the top of my head. We might call it luck. Or, if we’re thinking in terms of something like idea space, we could call it initial position. By virtue of 1, 2, 3 and perhaps 4 as well, the so-called genius is at a position in idea space that allows them to make major discoveries by deploying their cumulative capabilities. The point of this initial-position factor is to allow for the possibility of a cohort of thinkers more or less equally endowed with 1,2,3+4, but having very different initial position. As a consequence, some are able to achieve major discoveries quickly, while others take more time, and still others never get there. Their capabilities are comparable, but their outcomes are not.

To invoke the diamond mining metaphor I introduced in comment #108, we have two equally skilled geologists/prospectors. One just happens to be located within 100 miles of a major kimberlite deposit while the other is over 3000 miles away from such a deposit. If they start walking from where they are, who’s going to find diamonds first?

In the case of AI, we know a great deal about compute/speed; we have that under control. I’m not sure just how the distinction between innate vs. learned techniques applies to machines, perhaps hardware and software. In any case, we do have a large repertoire of techniques of various kinds. In some areas we can produce a combination of compute and technique that allows the machine to outperform the best human. In other areas we have machines that do things that are amazing in comparison with what machines did, say, a decade ago, but which are no more than standard human performances, with various failings here and there. And so on. As for starting position, I think it’s up to us to position the AI properly, at least at the start.

[But once and if it FOOMs, it’s on its own. I’m not holding my breath on this one.]

Figure 5 in the 2020 New Savanna post gives a visual illustration of the initial position idea.

We now have a total of five factors:

1) compute,
2) IQ,
3) an inherited component of technique,
4) a learned component of technique, and
5) initial position or luck.

The seductiveness of scale

The hope of the scaling side of the current debate is that we can get all the way to AGI – whatever that is – by throwing more compute at the problem. Well, it’s not that simple, the compute has to be channeled through an appropriate machine-learning architecture which then chews its way to a huge pile of (appropriately curated) data. That is, architecture+data will cover the ground I’ve indicated in factors 2-5 above, thereby relieving us of the need and responsibility to think about those things.

It’s a seductive prospect. Why? In part because it is easy to understand, even by people who have little or no technical knowledge of computing, cognitive science, and AI. Everyone knows and understands, “bigger is better.”

GOFAI (good old fashioned artificial intelligence) was mostly about technique, factors 3 and 4. That technique was generally taken to be mediated by symbolic systems and, as a practical matter, it required that ‘knowledge’ be painstakingly hand-crafted into systems. While I can understand the desire to avoid hand-crafted knowledge – there’s so very much of it and the crafting is tedious and error-prone — I don’t think symbolic computation can be avoided. Can it be architected, as it were, into a learning regime? We know one case where it has been, the human case, but that case tells us that learning requires a lot of close interaction between teachers and students, in both formal and informal settings. It’s not at all clear to me that such interaction can be architected.

More later. 

Addendum, 7.1.22, on superintelligence: Alex, comment #171:

I think Pinker’s definition of intelligence, “the ability to use information to attain a goal in an environment”, is reasonable, but it doesn’t give us any meaningful way to compare or rank intelligences (so how can we meaningfully discuss “superintelligence”?). Of course, you chose compute time as the metric, but I think that dodges the more meaningful aspects of intelligence. I think a metric like computational complexity – or even Kolmogorov complexity – is more appealing to me, but whatever the metric, I think it has to capture the mechanism of thought in some way, not just the output. [...]

As a final note, I think “intelligence” is a crude word that tries to capture too many aspects of behavior (many of them human-relatable, but not of great importance to discussion). My comment here has been an attempt to break up “intelligence” into constituent parts to focus discussion: clock speed, memory, algorithmic/time complexity, size/space complexity. There are surely more parts of “intelligence”, some parts that are combinations of simpler parts.

Scott, comment #172:

Fundamentally, I care, not about the definitions of words like “superintelligence,” but about what will actually happen in the real world once AIs become much more powerful. [...] So OK then, what happens when we can launch a billion processes in datacenters, each one with the individual insight of a Terry Tao or Edward Witten (or the literary talent of Philip Roth, or the musical talent of the Beatles…), and they can all communicate with one another, and they can work at superhuman speed? Is it not obvious that all important intellectual and artistic production shifts entirely to AIs, with humans continuing to engage in it (if they do) only as a hobby? That’s the main question I care about when I discuss “superintelligence,” and I’m still waiting for anyone to explain why I’m wrong about it.

Friday Fotos: Which one doesn't belong? [flowers]