Showing posts with label G_Hinton. Show all posts
Showing posts with label G_Hinton. Show all posts

Friday, June 26, 2026

A Meeting of Minds on Mars: Charles Babbage • John von Neumann • Geoffrey Hinton

This discussion is about the implementation of computing in matter. That was the topic of John von Neumann’s last book, The Computer and the Brain. He died before he finished it, so it was published posthumously in 1958. I don’t know when I first learned about it, perhaps sometime in the mid-1970s. Though I knew about von Neumann and his role in early computing, it dismissed the book itself, figuring that we’d learned so much about the brain since then, and the nature of computers had changed so much, that it must be obsolete.

I was wrong. When I finally read the book, probably in the early to mid-1980s I was stunned. This was a profound book and free of mathematics beyond some simple back-of-the-envelope calculations. For one thing von Neumann talked of both analog and digital computation; that contrast was central. When I first started reading about computers in the mid-to-late 1960s that contrast was at the beginning of every article or book. But once personal computers appeared and proliferated in the late 1970s and 1980s, analog computing was all but forgotten.

It was good to see it back. For one thing, I had been strongly influenced by the work of William Powers, whose 1973 book, Behavior: The Control of Perception, offered an elegant analog model of the brain. That was central to the work of my teacher and mentor, David Hays, who had been one of the founders of computational linguistics. In 1974 and 1975 he worked out a scheme in which cognitive networks were grounded in Powers’s analog model. This was years before anyone else was working on the problem, before Steven Harnad coined the term, “symbol grounding” in his 1990 paper on the problem in Physica D.

Now, with the success of artificial neural networks, there’s another aspect of physical implementation we’ve got to deal with, distributed representation. You’re looking at an apple. Where is that apple represented in your brain? There might be some one neuron or a tightly clustered group of neurons that represent that apple. That’s a local representation. But the representation might be distributed across a loosely linked population of neurons. That’s a distributed representation. We now know that that’s how the brain works, though I assume there’s some laggards still stuck in the previous century where local representation was widely favored. And that’s how artificial neural nets work. The concept of apple isn’t localized in one or a small group of weights, it’s smeared over a population of weights. The same for the concepts of truth, beauty, and justice, or, for that matter, neuron.

Consequently von Neumann’s 1958 discussion must now be extended to distributed representation. It is for that purpose that I asked Claude to convene a conversation between Charles Babbage, who conceived of an elaborate mechanical computer in the 19th century, von Neumann himself, and Geoffrey Hinton, one of the so-called godfathers of artificial intelligence – a bit of journalistic and pop culture slight-of-hand that obscures a long and complex history. Moreover, since these three gentleman did not co-exist in time, though there’s a bit of overlap between von Neumann and Hinton, I decided that I might also have Claude transport them to Mars for the occasion. Anything is possible in virtual reality. 

Or is it? 

Image by ChatGPT. 

A sun-facing room on Mars. Three chairs arranged loosely, angled toward a broad window. Through it, the Martian plain stretches red and absolutely still, and the sun — recognizable but wrong, too small, too distant — hangs in a pale pink sky. Three men stir and look about them, bewildered but unhurt.

I. Awakening

Charles Babbage

I am not dead — I can confirm that much. My pulse is steady and my faculties appear intact. But this room is not any room I have entered by ordinary means. That light through the window — is that the sun? The colour is right but the disc is too small by half. And that plain outside... brick red, entirely without vegetation or track or any sign of habitation. Gentlemen, I should very much like an explanation.

John von Neumann

I was in Princeton not five minutes ago. I notice the gravity is reduced — perhaps sixty percent of normal. The sky has the ferric oxide colouration one would predict. The solar disc subtends roughly two-thirds the angle it does from Earth. I believe I know where we are, and I find the hypothesis almost too extraordinary to voice.

Geoffrey Hinton

Mars. We're on Mars. I was at a conference in Toronto. But that sky, the dust, the sun at that distance — I'd stake my reputation on it. He exhales slowly. I've recently left a position at Google partly so I could speak freely about alarming possibilities. I did not anticipate that the alarming possibilities would include this. May I ask what year it is, from each of your perspectives?

Charles Babbage

The year of Our Lord 1862. I am Charles Babbage, formerly Lucasian Professor of Mathematics at Cambridge, Fellow of the Royal Society. I have spent the better part of forty years attempting to build mechanical calculating engines, with results that have been — a pause — mixed, as regards the support of Her Majesty's government.

John von Neumann

John von Neumann. The year is 1955. I work at the Institute for Advanced Study in Princeton, and consult for various government and military bodies. Mr. Babbage — I know your name extremely well. Better, I suspect, than you know mine.

Geoffrey Hinton

Geoffrey Hinton. For me it is 2024. And I know both of you — Mr. Babbage, you designed the Difference Engine and the Analytical Engine. Dr. von Neumann, you are among the most formidable mathematicians of the twentieth century and you gave your name to the architecture that every conventional computer on Earth is built upon. You are, in a real sense, my ancestors. The field I work in — machine learning, artificial intelligence — descends directly from the problems both of you were grappling with.

Charles Babbage

A long pause, during which he stares at Hinton with an expression mixing hunger, pride, and something close to grief. The Analytical Engine. Did anyone build it?

Geoffrey Hinton

Not in your lifetime. The government never provided the funds. But the ideas were entirely right — and they were eventually built, first in relay and vacuum tube and then in silicon, by people who in some cases had read your work and in other cases had arrived at the same conclusions independently. You were approximately a century early.

Charles Babbage

Very quietly. A century. I had hoped twenty years would suffice. I petitioned the Chancellor. Three times.

II. The mill and the store

John von Neumann

Mr. Babbage, allow me to tell you what your Analytical Engine set in motion — because it bears directly on the work that has occupied all three of us. You made a distinction, in your design, between what you called the Mill and the Store. The Mill performs the operations — addition, subtraction, multiplication. The Store holds the numbers awaiting operation and the results of operations completed. That separation of active calculation from passive memory was the foundational insight.

Charles Babbage

It seemed to me the only sensible arrangement. The columns of number-wheels in the Store are passive — they merely hold values. The Mill acts upon them. To mix the two functions in the same mechanism would create hopeless confusion.

John von Neumann

And yet it is precisely that separation which I have spent the last years of my life questioning — not as an engineering choice, which was entirely sound, but as a principle of intelligence itself. My colleagues and I formalized your Mill-and-Store distinction into what is now called the stored-program architecture. The processor executes instructions sequentially; the memory holds both data and those instructions passively until called upon. It is a serial machine, one operation following another, orchestrated by a central clock. The computers being built in my era all follow this pattern.

Geoffrey Hinton

And in my era they still do, at bottom. But you've just described the tension at the heart of everything, Dr. von Neumann. Serial, precise, with a strict wall between computation and memory — that is the von Neumann architecture. And it is, in a sense, the architecture that human intelligence refuses to use.

Charles Babbage

You are saying the brain does not separate Mill from Store?

Geoffrey Hinton

Exactly. In the brain, every neuron is simultaneously a memory element and a processing element. It holds information in the strength of its connections to other neurons, and it also fires — it computes — based on what it receives. There is no central Mill. There is no passive Store. Computation and memory are fused at every node in the network, and the whole thing operates in parallel, millions of neurons active at once.

John von Neumann

This is precisely what troubles me — and I am glad to find the trouble is still alive in your era, Mr. Hinton, because it means I was not merely chasing a phantom. I have been writing a manuscript, unfinished I'm afraid, called "The Computer and the Brain." My central puzzle is this: where, in a neuron, is the Mill? A neuron receives signals, sums them, and if the sum exceeds a threshold, it fires. That threshold operation is a computation. But the synaptic weights — the strengths of the incoming connections — those are also the memory. The neuron is its own Mill and its own Store simultaneously. I have been calling such things active elements, to distinguish them from the passive memory elements of our conventional machines. But I confess I have not yet worked out the full implications.

Geoffrey Hinton

With genuine feeling. Dr. von Neumann, the full implications are what I have spent my career working out. And you had named the essential thing: active elements. That is exactly what the nodes of a neural network are. Each one holds a weight — that is its memory — and each one applies a non-linear function to its inputs — that is its computation. Fused. Inseparable. Replicated millions or billions of times, connected in layers, and trained by adjusting all those weights simultaneously until the network's outputs match the desired answers.

Wednesday, May 14, 2025

Are radiologists here to stay? Yes.

Steve Lohr, Your A.I. Radiologist Will Not Be With You Soon, NYTimes, May 14, 2025.

Nine years ago, one of the world’s leading artificial intelligence scientists singled out an endangered occupational species.

“People should stop training radiologists now,” Geoffrey Hinton said, adding that it was “just completely obvious” that within five years A.I. would outperform humans in that field.

Today, radiologists — the physician specialists in medical imaging who look inside the body to diagnose and treat disease — are still in high demand. A recent study from the American College of Radiology projected a steadily growing work force through 2055.

Dr. Hinton, who was awarded a Nobel Prize in Physics last year for pioneering research in A.I., was broadly correct that the technology would have a significant impact — just not as a job killer.

That’s true for radiologists at the Mayo Clinic, one of the nation’s premier medical systems, whose main campus is in Rochester, Minn. There, in recent years, they have begun using A.I. to sharpen images, automate routine tasks, identify medical abnormalities and predict disease. A.I. can also serve as “a second set of eyes.”

“But would it replace radiologists? We didn’t think so,” said Dr. Matthew Callstrom, the Mayo Clinic’s chair of radiology, recalling the 2016 prediction. “We knew how hard it is and all that is involved.”

Computer scientists, labor experts and policymakers have long debated how A.I. will ultimately play out in the work force. Will it be a clever helper, enhancing human performance, or a robotic surrogate, displacing millions of workers?

The debate has intensified as the leading-edge technology behind chatbots appears to be improving faster than anticipated. Leaders at OpenAI, Anthropic and other companies in Silicon Valley now predict that A.I. will eclipse humans in most cognitive tasks within a few years. But many researchers foresee a more gradual transformation in line with seismic inventions of the past, like electricity or the internet.

The predicted extinction of radiologists provides a telling case study. So far, A.I. is proving to be a powerful medical tool to increase efficiency and magnify human abilities, rather than take anyone’s job.

When it comes to developing and deploying A.I. in medicine, radiology has been a prime target. Of the more than 1,000 A.I. applications approved by the Food and Drug Administration for use in medicine, about three-fourths are in radiology. A.I. typically excels at identifying and measuring a specific abnormality, like a lung lesion or a breast lump. [...]

Predictions that A.I. will steal jobs often “underestimate the complexity of the work that people actually do — just as radiologists do a lot more than reading scans,” said David Autor, a labor economist at the Massachusetts Institute of Technology.

Note well:

Dr. Halamka, an A.I. optimist, believes the technology will transform medicine.

“Five years from now, it will be malpractice not to use A.I.,” he said. “But it will be humans and A.I. working together.”

Dr. Hinton agrees. In retrospect, he believes he spoke too broadly in 2016, he said in an email. He didn’t make clear that he was speaking purely about image analysis, and was wrong on timing but not the direction, he added.

There's more at the link.

Tuesday, March 12, 2024

On the similarities between a kumquat and a MiG: ChatCPT chases analogies

We all know that not so long-ago Geoffrey “AI Godfather” Hinton went up on a mountain where he received a heavy granite tablet. As he was holding it a bolt of lightning struck the tablet, leaving the words “You Are Doomed” carved deeply into its surface. Hinton came down from the mountain, tablet in tow, quit his post at Google, and proceeded to warn the world about the dangers of Superintelligent AI.

In early October of last year he took part in a panel discussion about AI and creativity:

One of his reasons for believing in the impending superintelligence of AI has to do with reasoning by analogy. Starting at about 1:18:10:

GEOFFREY HINTON: Let me give you-- let me give you an example of something creative that GPT-4 can already do that most people can't do.

So we're still trapped in the idea of thinking that logical reasoning is the essence of intelligence when we know that being able to--

TOMASO POGGIO: --but some people.

GEOFFREY HINTON: Well--

TOMASO POGGIO: Yeah.

GEOFFREY HINTON: We know that being able to see analogies, especially remote analogies, is a very important aspect of intelligence. So I asked GPT-4, what has a compost heap got in common with an atom bomb? And GPT-4 nailed it, most people just say nothing.

DEMIS HASSABIS: What did it say? [LAUGHTER]

GEOFFREY HINTON: It started off by saying they're very different energy scales, so on the face of it, they look to be very different. But then it got into chain reactions and how the rate at which they're generating energy increases-- their energy increases the rate at which they generate energy. So it got the idea of a chain reaction. And the thing is, it knows about 10,000 times as much as a person, so it's going to be able to see all sorts of analogies that we can't see.

My initial reaction was: Very clever, very clever indeed! But then I thought, I’ll bet GPT-4 could explicate any analogy you ask it to. The cleverness is in picking the right analogy, not in justifying it. Justifying is easy; coming up with an analogy worth the effort of justification, that’s difficult. If GPT-4 knows 10,000 times as much as a person and it has all that compute, then maybe it can waste time far more extravagantly than any person can afford to. Whatever that is, it’s not intelligence.

So I decided to run a little experiment, not with GPT-4, but with plain old ChatGPT, aka the Chatster. I presented it with ten analogies I thought of rather quickly and asked it to explicate them. It did so, and rather impressively I might add. I then asked it to pick the best ones; it evaded the issue. Rather than trying to prompt it into submission I asked it to tell me which of those things was most like a kumquat. As the title of this post indicates, it choose the MiG, and offered an amusing bit of argumentation on the matter. Is that an example of intelligence creativity? How would GPT-4 respond to those prompts? With even more elaborate justifications? Would it have chosen the MiG as well.

And what about GPT-5? Even more extraneous extravagance? Or will it be clever enough to object, saying:

Those analogies are silly. Don’t waste my valuable time and your electricity. Come back when you have something more intelligent to offer.

That’s how I feel about Hinton’s ideas about creativity and human accomplishment. He may be a genius at coming up with neural network architectures, but his thinking about human creativity and intelligence seems to be the work of an overgrown college sophomore. And he’s not the only one.

You might not want to read the whole transcript. But you should at least look at the analogies I proposed, dip into some of Chatster’s explications, read his typically annoying evasion of my penultimate prompt and then, by all means, read his account of the similarities between a kumquat and a MiG. It’s more fun than a barrel of monkeys.

ChatGPT explicates ten analogies

I am going to present you with a bunch of analogies and I want you to explain them. Some of them may same more plausible than others, but don’t let that bother. Just find analogous features.

Here’s the first one: A kitten and a bicycle.

Certainly! Let's explore the analogy between a kitten and a bicycle:

  1. Learning Curve:
    • Kitten: Like learning to ride a bicycle, a kitten goes through a learning curve as it develops coordination and balance skills. It starts by stumbling and gradually becomes more adept at moving around.
    • Bicycle: When someone first learns to ride a bicycle, there is a learning curve in terms of balance and coordination. Over time, the rider becomes more skilled and comfortable.
  2. Independence:
    • Kitten: As a kitten grows, it becomes more independent, learning to explore its surroundings and perform activities on its own.
    • Bicycle: Riding a bicycle grants a sense of independence, allowing the rider to travel to different places under their control.
  3. Care and Maintenance:
    • Kitten: Similar to a bicycle requiring maintenance, a kitten needs care, including feeding, grooming, and veterinary attention.
    • Bicycle: Regular maintenance is essential for a bicycle to function smoothly. This involves checking and adjusting various components, such as the tires, brakes, and gears.
  4. Playfulness:
    • Kitten: Kittens are known for their playful nature, exploring their environment and engaging in playful activities.
    • Bicycle: Riding a bicycle can be a form of recreation and play, especially for leisure rides or sports.
  5. Adaptability:
    • Kitten: Kittens can adapt to various environments and situations, showcasing their flexibility.
    • Bicycle: Bicycles are versatile and can be adapted for different terrains or purposes, such as road cycling, mountain biking, or commuting.

While a kitten and a bicycle may seem quite different at first glance, there are interesting similarities when considering aspects like learning, independence, care, playfulness, and adaptability.

Saturday, March 2, 2024

Of Lit Crit “Stars” and AI “Godfathers” – In what way is Geoffrey Hinton like Jacques Derrida?

Back in 1997 David Shumway published “The Star System in Literary Studies” in PMLA. He begins with a paragraph about George Lyman Kittredge, of Harvard’s English Department at the end of the 19th and beginning of the 20th century, noting that Kittredge was unknown to the public. Here’s the first sentence of his second paragraph:

Kittredge, who virtually founded Chaucer studies in the United States, stood at the head of a professional genealogy that controlled the field for many years after his death, but he was not a star. Nor were any of his illustrious contemporaries or near contemporaries, such as John Manly, John Livingston Lowes, and so on. Why they were not stars and Judith Butler, Jacques Derrida, Stanley Fish, Henry Louis Gates, Jr., Fredric Jameson, Gayatri Chakravorty Spivak, and other figures in the academy are is the subject of this essay.

What I’m wondering is whether or not the so-called AI “Godfathers” don’t represent a similar phenomenon in contemporary AI. Strictly speaking I believe the Godfather term applies to the three winners of the 2018 Turing Award, Yoshua Bengio, Yann Lecun, and Geoffrey Hinton, but I believe there are others in AI with a similar status, such as Ilya Sutskever, Hinton’s student and co-founder of OpenAI, Andrej Karpathy, the former director of AI at Tesla who just made waves, albeit little ones, by resigning from OpenAI, Demis Hassibis, cofounder of DeepMind, and perhaps even such figures as Nick Bostrom and Eliezer Yudkowsky, who aren’t AI researchers but are highly influential figures through their commentary. Perhaps Sam Altman, the heroic CEO who fought off a recalcitrant board, is a star as well.

But first let’s get back to literary criticism. Shumway notes that there have been literary scholars in the past (relative to 1997) who were powerful and influential and who “probably received disproportionate recognition for their contributions compared with that accorded less well known scholars for comparable work.” The lit crit stars, whom he analogizes to movie stars (hence the term), are a product of the last quarter of the century. He dates the public emergence of these stars to a 1987 New York Times Magazine profile of the “Yale Critics,” the so-called “Yale Mafia,” of Harold Bloom, Geoffrey Hartman, J. Hillis Miller, and Jacques Derrida. He goes on to note that “The star system in literary studies, like that of the studio era, involves identification with a person who represents an ideal.”

Most of these critical stars are identified with capital-T Theory, a catchall term for the variety of schools of thought that emerged in the last three decades of the century. Harold Bloom, himself a star, nonetheless came to separate himself from the rest, categorizing them as the School of Resentment. [This separation, by the way, seems anti-mimetic, noting that Girard himself was such a star.] In a crucial passage, Shumway notes:

Theory not only gave its most influential practitioners a broad professional audience but also cast them as a new sort of author. Theorists asserted an authority more personal than that of literary historians or even critics. As we have seen, the rhetoric of literary history denied personal authority; in principle, even Kittredge was just another contributor to the edifice of knowledge. Criticism was able to enter the academy only by claiming objectivity for itself, so academic critics could not revel in personal idiosyncrasy. They developed their own critical perspectives, to be sure, but all the while they continued to appeal to the text as the highest authority. In the past twenty years theory has undermined the authority of the text and of the author and replaced it with the authority of systems...

Note the word “author” at the end of that sentence. Remember, Shumway is a literary critic writing about literary criticism. In that field, the primary and most important authors are the creators of the literary works that the field tends to (by editorial work and creating critical editions) and studies. All of a sudden these lit crit stars are up there in the firmament with Dickins, Sappho, Dante, Austen, Faulkner, and – gasp! – the Blessed Bard his-own Bad Self. That’s the kind of authority they have. Perhaps not of the same magnitude, but of the same (apparently) self-generating kind.

Shumway then notes: “Because authority in the natural sciences is rooted in a consensus about such norms, the hierarchies in these fields have not developed into star systems of the sort I have described here.” That brings us to AI, which is not a science, though it includes some forms of scientific knowledge within its scope. It is an engineering discipline that lives or dies on what it builds.

The field is attempting to create computer systems that are as “intelligent” as human beings across a wide range of tasks. But the concept of “intelligence” is difficult to define, as is the idea of AGI (artificial general intelligence). It has created remarkable and dazzling technology for language and images. But the technology has a “black box” aspect that has so far resisted analysis. We don’t know how it works. Nor do we know how to assess its performance or to project performance into the future.

Concerning the rise of literary stars, Shumway noted: “As theory has called into question the traditional means by which knowledge has been authorized, it may be that the construction of the individual personality has become an epistemological necessity.” That seems like the state of AI today. We’ve got a very complicated technology involving a blend of engineering, science, and alchemy, lacking objective knowledge. Note only that, the technology is enormously important and will change the way we live. In the absence of objective knowledge, what choice do we have but to steer by the freakin' stars?

* * * * *

These remarks were occasioned by the deference Nathan Gardels gave to Geoffrey Hinton in the current issue of Noema:

Beyond the avid venture capitalists and digital giants promoting the rapid commercialization of generative AI in all its promise, more sober and critical voices, not least the pioneers of the very technology among them, worry that it can become an “existential threat to humanity.” But few of those in the know ever explain, in lay terms you and I might understand if we try, what that actually means and how it may come about.

Considered the “godfather of AI,” Geoffrey Hinton is more in the know than most — and thus more concerned than most over the dangers of fostering superintelligence smarter than we can ever be. When OpenAI’s ChatGPT4 was released last year, he experienced an “epiphany” that led him to defect from his research post at Google, expressing regret over much of his life’s work.

In the Romanes Lecture delivered at Oxford University last week, Hinton explained the logic of his fears with the same step-by-step rigor by which he helped devise the early artificial neural networks that are the foundation of the superintelligence that so concerns him.

On the first highlighted passage: And we are so very lucky that the Great Man has taken time out of his busy schedule to tell us what’s on his mind.

On the second highlighted passage: I’ve not watched this particular video, but I’ve seen other recent performances by Hinton and I do not hold out high hopes for the rigor of this effort. For my opinion of Hinton on such matters, see the section, “The Experts Speak for Themselves,” in my recent 3QD piece, Aye Aye, Cap’n! Investing in AI is like buying shares in a whaling voyage captained by a man who knows all about ships and little about whales. 

Beyond that, “step-by-step” is not how’d I’d characterize research in AI – or any other field for that matter. Yes, there is rigor, but there’s also chance and dumb luck. None of this has come about through a carefully executed plan in pursuit of a well-defined goal. “Step-by-step” is wishful thinking; it is star-struck. 

* * * * *

Time Magazine lists (anoints?) the AI stars: TIME Reveals Inaugural TIME100 AI List of the World's Most Influential People in Artificial Intelligence.

Monday, January 15, 2024

Toward a Theory of Intelligence: Did Miriam Yevick know something in 1975 that Bengio, LeCun, and Hinton did not know in 2018?

One theme that comes up in various discussions of artificial intelligence is that the discipline is primarily an empirical one that lacks theoretical grounding. The default view, and perhaps the dominant one as well, is that what we’re doing is producing results so damn the torpedoes – full speed ahead! But the call to theory keeps nagging, perhaps most recently in a panel discussion entitled Research on Intelligence in the Age of AI, and hosted by MIT’s Center for Minds, Brains, and Machines on its 10th Anniversary.

One theme that has been kicking around for several decades is that there are two styles of computational regime underlying perception, action, and cognition. My purpose is to compare the views that Miriam Lipschutz Yevick articulated about this dichotomy in 1975 and 1978 with those articulated by Yoshua Bengio, Yann LeCun, and Geoffrey Hinton in their 2018 Turing Award lecture, which was published in 2021.

Bengio, LeCun, and Hinton, 2018

Let’s start with Bengio, LeCun, and Hinton, who won the Turing Award in 2018. They published their paper, Deep learning for AI, in 2021. In that paper they asserted:

There are two quite different paradigms for AI. Put simply, the logic-inspired paradigm views sequential reasoning as the essence of intelligence and aims to implement reasoning in computers using hand-designed rules of inference that operate on hand-designed symbolic expressions that formalize knowledge. The brain-inspired paradigm views learning representations from data as the essence of intelligence and aims to implement learning by hand-designing or evolving rules for modifying the connection strengths in simulated networks of artificial neurons.

In the logic-inspired paradigm, a symbol has no meaningful internal structure: Its meaning resides in its relationships to other symbols which can be represented by a set of symbolic expressions or by a relational graph. By contrast, in the brain-in- spired paradigm the external symbols that are used for communication are converted into internal vectors of neural activity and these vectors have a rich similarity structure. Activity vectors can be used to model the structure inherent in a set of symbol strings by learning appropriate activity vectors for each symbol and learning non-linear transformations that allow the activity vectors that correspond to missing elements of a symbol string to be filled in. This was first demonstrated in Rumelhart et al. on toy data and then by Bengio et al. on real sentences. A very impressive recent demonstration is BERT, which also exploits self-attention to dynamically connect groups of units, as described later.

As I said at the beginning, some such characterization of two modes of thinking has been around for some time, though it is expressed in various ways. I have no problem recognizing such a distinction.

Yevick, 1975 and 1978

Miriam Yevick recognized that distinction in her 1975 paper, Holographic or Fourier Logic (Pattern Recognition 7, 187-213). That was at the peak of interest in logic-inspired AI. That was the year Newell and Simon won the Turing Award; their paper, Computer Science as Empirical Inquiry: Symbols and Search, was published the following year.

Yevick was a mathematician, not a cognitive scientist, and had become interested in optical holography though her extensive correspondence with David Bohm, the physicist, during the 1950s. During the 1960s a number of thinkers, including the neuroscientist, Karl Pribram, and the cognitive scientist, Chrisopher Longuet-Higgins, had become interested in holography as a model for neural processing. It’s that interest the Yevick had in mind when she wrote her article. Here is one statement from that article:

It has recently been conjectured that neural holograms enter as units in the thought process. If holographic processes do occur in the brain and are instrumental in thought, then the logical operations implicit in these processes could be considered as intuitive and enter as units in our mental and mathematical computations.

It has also been said that: “if we want the computer to have eyes, we shall first have to give him instruction in the facts of life”.

We maintain in this paper that a language of thought in which holographic operations enter as primitives is essentially different from one in which the same operations are carried out sequentially and hence over a finite time span [...] Our assumption is that “holographic thought” utilizes the associative properties of holograms in “one shot”. Similarly we maintain that apprehension proceeds from the very beginning via two modes, the aural and the optical; whereas the verbal string is natural to the first, the pattern as such is natural to the second: the essentially instantaneous nature of the optical process captures the apprehension as a global unit whose meaning is expressed in the first place in terms of “associations” with other such units.

There we have our distinction, between aural, verbal, and sequential on the one hand and optical, intuitive, and pattern on the other.

Having chosen visual objects as her domain, she argues thus (and here I am quoting from a 1978 restatement):

We can explicate this proposition on a theoretical level in the domain of optical patterns. [...] Such patterns or objects are thin, white regions on a black background. These can be simple (regular), like the outlines of rectangles; or complex, like the outlines of Chinese characters or random-like motions. The following holds true: a complex object requires a long (sequential, quasi-linguistic) description but yields a sharp recognition (auto-correlation) spot under holographic filtering; hence it is identified most readily by holographic recognition, or holistically. A simple object requires a short (quasi-linguistic) description but yields a diffuse recognition spot; hence it is identified most readily by quasi-linguistic representation or description.

Description and holographic recognition thus appear as two (complementary) modes of identifying an object: the more complex the object, the longer its description and the sharper its auto-correlation spot, and vice versa. The more complex they physiognomy of a person, the more unique, and hence sharper, its identity and ease of recall; the more simple, the more common and hence “unidentifiable.” Perfect holographic recognition obtains for a totally “random object”, that is, one with an infinitely long description; for a perfectly sharp point the opposite is true.

Suppose that one is given a store of objects with which one is familiar, a holographic recognition device, and a quasi-linguistic mode of representation; one is then presented with an arbitrary object to be “identified.” An approximate match is obtained either by producing a description of acceptable length or by holographic recognition of a subset of similar (associated) objects from the store. The mode of identification that will be more appropriate then depends on the complexity of the unknown object. If it is simple, we ”know” it by a short linguistic description; if it is complex, by the “associations” it evokes.

What Yevick is explicit about, and what is missing from Bengio, LeCun, and Hinton, is the relationship between some object of perception and cognition and the computational regime operating on that object. She recognizes the utility of both regimes, but associates them with different kinds of objects. As Bengio, LeCun, and Hinton simply do not conceptualize that relationship it is not clear how they would respond to Yevick’s work.

As I recall, that relationship was beginning to be recognized as an issue. Early AI had achieved its successes from dealing with sequential symbolic processing (e.g. theorem proving, expert systems), but faltered when dealing with visual perception and speech recognition. Though I can’t offer a citation, I recall David Marr, who died in 1980, mentioning the problem. The best-known statement of the problem is by Hans Moravec in his 1988 book, Mind Children, where he says “it is comparatively easy to make computers exhibit adult level performance on intelligence tests or playing checkers, and difficult or impossible to give them the skills of a one-year-old when it comes to perception and mobility”(as quoted in Wikipedia). While it is generally recognized within computer science at large, that different kinds of computational system are suited to different kinds of problems, so far as I know the issue has not be systematically investigated in the context of artificial intelligence and machine learning. More specifically, Miriam Yevick’s work from the 1970s has not been taken into account.

What needs to be done

Lots.

Let me repeat: Lots.

For one thing, Yevick’s work has been forgotten. It needs to be revived and vetted in view of more recent work.

Moreover, while her mathematics concentrated on one problem, object identification, in the visual domain, she informally generalized that result to thinking in general. For example (I’m quoting from her 1978 paper):

If we consider that both of these modes of identification enter into our mental processes, we might speculate that there is a constant movement (a shifting across boundaries) from one mode to the other: the compacting into one unit of the description of a scene, event, and so forth that has become familiar to us, and the analysis of such into its parts by description. Mastery, skill and holistic grasp of some aspect of the world are attained when this object becomes identifiable as one whole complex unit; new rational knowledge is derived when the arbitrary complex object apprehended is analytically described.

I’m certainly sympathetic to that generalization. It’s what David Hays and I had in mind when we called on Yevick’s ideas in a 1987 paper on metaphor [1] and a 1988 paper on the brain and human intelligence [2].

For all I know, Bengio, LeCun, and Hinton might be sympathetic as well. Here’s the final paragraph of their Turing Award paper:

How are the directions suggested by these open questions related to the symbolic AI research program from the 20th century? Clearly, this symbolic AI program aimed at achieving system 2 abilities, such as reasoning, being able to factorize knowledge into pieces which can easily recombined in a sequence of computational steps, and being able to manipulate abstract variables, types, and instances. We would like to design neural networks which can do all these things while working with real-valued vectors so as to preserve the strengths of deep learning which include efficient large-scale learning using differentiable computation and gradient-based adaptation, grounding of high-level concepts in low-level perception and action, handling uncertain data, and using distributed representations.

That sounds like a call to reconstruct symbolic capabilities in the context of more realistic models of real neural networks. That also sounds like intellectual work for several generations of researchers. As Charlie Parker was fond of saying, “Now’s the Time.”

References

[1] William Benzon and David Hays, Metaphor, Recognition, and Neural Process, The American Journal of Semiotics, Vol. 5, No. 1 (1987), 59-80. https://www.academia.edu/238608/Metaphor_Recognition_and_Neural_Process.

[2] William Benzon and David Hays, Principles and Development of Natural Intelligence, Journal of Social and Biological Structures, Vol. 11, No. 8, July 1988, 293-322. https://www.academia.edu/235116/Principles_and_Development_of_Natural_Intelligence.

Tuesday, December 5, 2023

Cage Match in the Valley 2: Google suits up, others join the fray

Karen Weise, Cade Metz, Nico Grant and Mike Isaac, Big Tech Muscles In: The 12 Months That Changed Silicon Valley Forever, NYTimes, Dec. 5, 2023:

What played out at Google was repeated at other tech giants after OpenAI released ChatGPT in late 2022. They all had technology in various stages of development that relied on neural networks — A.I. systems that recognized sounds, generated images and chatted like a human. That technology had been pioneered by Geoffrey Hinton, an academic who had worked briefly with Microsoft and was now at Google. But the tech companies had been slowed by fears of rogue chatbots, and economic and legal mayhem.

Once ChatGPT was unleashed, none of that mattered as much, according to interviews with more than 80 executives and researchers, as well as corporate documents and audio recordings. The instinct to be first or biggest or richest — or all three — took over. The leaders of Silicon Valley’s biggest companies set a new course and pulled their employees along with them.

Over 12 months, Silicon Valley was transformed. Turning artificial intelligence into actual products that individuals and companies could use became the priority. Worries about safety and whether machines would turn on their creators were not ignored, but they were shunted aside — at least for the moment.

The night before ChatMAS:

On Nov. 29, the night before the launch, Mr. Brockman hosted drinks for the team. He didn’t think ChatGPT would attract a lot of attention, he said. His prediction: “no more than one tweet thread with 5k likes.”

Mr. Brockman was wrong. On the morning of Nov. 30, Mr. Altman tweeted about OpenAI’s new product, and the company posted a jargon-heavy blog item. And then, ChatGPT took off. Almost immediately, sign-ups overwhelmed the company’s servers. Engineers rushed in and out of a messy space near the office kitchen, huddling over laptops to pull computing power from other projects. In five days, more than a million people had used ChatGPT. Within a few weeks, that number would top 100 million. Though nobody was quite sure why, it was a hit. Network news programs tried to explain how it worked. A late-night comedy show even used it to write (sort of funny) jokes.

There's much more at the link.

Sunday, December 3, 2023

Cage Match in the Valley 1: Many men enter, one man leaves.

Cade Metz, Karen Weise, Nico Grant and Mike Isaac, Ego, Fear and Money: How the A.I. Fuse Was Lit, December 3, 2023:

The question of whether artificial intelligence will elevate the world or destroy it — or at least inflict grave damage — has framed an ongoing debate among Silicon Valley founders, chatbot users, academics, legislators and regulators about whether the technology should be controlled or set free.

That debate has pitted some of the world’s richest men against one another: Mr. Musk, Mr. Page, Mark Zuckerberg of Meta, the tech investor Peter Thiel, Satya Nadella of Microsoft and Sam Altman of OpenAI. All have fought for a piece of the business — which one day could be worth trillions of dollars — and the power to shape it.

At the heart of this competition is a brain-stretching paradox. The people who say they are most worried about A.I. are among the most determined to create it and enjoy its riches. They have justified their ambition with their strong belief that they alone can keep A.I. from endangering Earth.

Mr. Musk and Mr. Page stopped speaking soon after the party that summer. A few weeks later, Mr. Musk dined with Mr. Altman, who was then running a tech incubator, and several researchers in a private room at the Rosewood hotel in Menlo Park, Calif., a favored deal-making spot close to the venture capital offices of Sand Hill Road.

That dinner led to the creation of a start-up called OpenAI later in the year. Backed by hundreds of millions of dollars from Mr. Musk and other funders, the lab promised to protect the world from Mr. Page’s vision.

Thanks to its ChatGPT chatbot, OpenAI has fundamentally changed the technology industry and has introduced the world to the risks and potential of artificial intelligence. OpenAI is valued at more than $80 billion, according to two people familiar with the company’s latest funding round, though Mr. Musk and Mr. Altman’s partnership didn’t make it. The two have since stopped speaking.

Altman goes on to make a Girardian point:

“There is disagreement, mistrust, egos,” Mr. Altman said. “The closer people are to being pointed in the same direction, the more contentious the disagreements are. You see this in sects and religious orders. There are bitter fights between the closest people.”

There's much more at the link: Eliezer Yudkowsky (matchmaker to the stars), Demis Hassibis, DeepMind, Geoffrey Hinton, Microsoft, Bill Gates, Mark Zukerberg, Yan LeCun, Sergi Brin, Eric Schmidt, Darlo Amodei, Anthropic, the names keep coming moths to light.

Thursday, November 30, 2023

What AI experts think about the possibility of AI Doom

Saturday, November 25, 2023

On possible cross-fertilization between AI and neuroscience [Creativity]

MIT Center for Minds, Brains, and Machines (CBMM), a panel discussion: CBMM10 - A Symposium on Intelligence: Brains, Minds, and Machines.

On which critical problems should Neuroscience, Cognitive Science, and Computer Science focus now? Do we need to understand fundamental principles of learning -- in the sense of theoretical understanding like in physics -- and apply this understanding to real natural and artificial systems? Similar questions concern neuroscience and human intelligence from the society, industry and science point of view.

Panel Chair: T. Poggio
Panelists: D. Hassabis, G. Hinton, P. Perona, D. Siegel, I. Sutskever

Quick Comments

1.) I’m a bit annoyed that Hassabis is giving neuroscience credit for the idea of episodic memory. As far as I know, the term was coined by a cognitive psychologist named Endel Tulving in the early 1970s, who stood it in opposition to semantic memory. That distinction was all over the place in the cognitive sciences in the 1970s and its second nature to me. When ChatGPT places a number of events in order to make a story, that’s episodic memory.

2.) Rather than theory, I like to think of what I call speculative engineering. I coined the phrase in the preface to my book about music (Beethoven’s Anvil), where I said:

Engineering is about design and construction: How does the nervous system design and construct music? It is speculative because it must be. The purpose of speculation is to clarify thought. If the speculation itself is clear and well-founded, it will achieve its end even when it is wrong, and many of my speculations must surely be wrong. If I then ask you to consider them, not knowing how to separate the prescient speculations from the mistaken ones, it is because I am confident that we have the means to sort these matters out empirically. My aim is to produce ideas interesting, significant, and clear enough to justify the hard work of investigation, both through empirical studies and through computer simulation.

3.) On Chomsky (Hinton & Hassabis): Yes, Chomsky is fundamentally wrong about language. Language is primarily a tool for conveying meaning from one person to another and only derivatively a tool for thinking. And he’s wrong that LLMs can learn any language and therefore they are useless for the scientific study of language. Another problem with Chomsky’s thinking is that he has no interest in process, which is in the realm of performance, not competence.

Let us assume for the sake of argument that the introduction of a single token into the output stream requires one primitive operation of the virtual system being emulated by an LLM. By that I mean that there is no logical operation within the process, no AND or OR, no shift of control; all that’s happening is one gigantic calculation involving all the parameters in the system. That means that the number of primitive operations required to produce a given output is equal to the number of tokens in that output. I suggest that that places severe constraints on the organization of the LLM’s associative memory.

Contrast that with what happens in a classical symbolic system. Let us posit that each time a word (not quite the same as a token in an LLM, but the difference is of no consequence) is emitted, that itself requires a single primitive operation in the classical system. Beyond that, however, a classical system has to execute numerous symbolic operations in order to arrive at each word. Regardless of just how those operations resolve into primitive symbolic operations, the number has to be larger, perhaps considerably larger, than the number of primitive operations an LLM requires. I suggest that this process places fewer constraints on the organization of a symbolic memory system.

At this point I’ve reached 45:11 in the video, but I have to stop and think. Perhaps I’ll offer some more comments later.

LATER: Creativity

4.) Near the end (01:20:00 or so) the question of creativity comes up. Hassibis says AIs aren't there yet. Hinton brings up analogy, pointing out that, with all the vast knowledge LLMs have ingested, they're got opportunities for coming up with analogy after analogy after analogy. I've got experience with ChatGPT that's directly relevant to those issues, analogy and creativity.

One of the first things I did once I started playing with ChatGPT was have it undertake a Girardian interpretation of Steven Spielberg's Jaws. To do that it has to determine whether or not there is an analogy between events in the film and the phenomena that Girard theorizes about. It did that fairly well. So I wrote that up and published it in 3 Quarks Daily, Conversing with ChatGPT about Jaws, Mimetic Desire, and Sacrifice. Near the end I remarked:

I was impressed with ChatGPT’s capabilities. Interacting with it was fun, so much fun that at times I was giggling and laughing out loud. But whether or not this is a harbinger of the much-touted Artificial General Intelligence (AGI), much less a warning of impending doom at the hands of an All-Knowing, All-Powerful Superintelligence – are you kidding? Nothing like that, nothing at all. A useful assistant for a variety of tasks, I can see that, and relatively soon. Maybe even a bit more than an assistant. But that’s as far as I can see.

We can compare what ChatGPT did in response to my prompting with what I did unprompted, freely and of my own volition. There’s nothing its replies that approaches my article, Shark City Sacrifice, nor the various blog posts I wrote about the film. That’s important. I was neither expecting, much less hoping, that ChatGPT would act like a full-on AGI. No, I have something else in mind.

What’s got my attention is what I had to do to write the article. In the first place I had to watch the film and make sense of it. As I’ve already indicated, have no artificial system with the required capabilities, visual, auditory, and cognitive. I watched the film several times in order to be sure of the details. I also consulted scripts I found on the internet. I also watched Jaws 2 more than once. Why did I do that? There’s curiosity and general principle. But there’s also the fact that the Wikipedia article for Jaws asserted that none of the three sequels were as good as the original. I had to watch the others to see for myself – though I was unable to finish watching either of that last two.

At this point I was on the prowl, though I hadn’t yet decided to write anything.

I now asked myself why the original was so much better than the first sequel, which was at least watchable. I came up with two things: 1) the original film was well-organized and tight while the sequel sprawled, and 2) Quint, there was no character in the sequel comparable to Quint.

Why did Quint die? Oh, I know what happened in the film; that’s not what I was asking. The question was an aesthetic one. As long as the shark was killed the town would be saved. That necessity did not entail the Quint’s death, nor anyone else’s. If Quint hadn’t died, how would the ending have felt? What if it had been Brody or Hooper?

It was while thinking about such questions that it hit me: sacrifice! Girard! How is it that Girard’s ideas came to me. I wasn’t looking for them, not in any direct sense. I was just asking counter-factual questions about the film.

Whatever.

Once Girard was on my mind I smelled blood, that is, the possibility of writing an interesting article. I started reading, making notes, and corresponding with my friend, David Porush, who knows Girard’s thinking much better than I do. Can I make a nice tight article? That’s what I was trying to figure out. I was only after I’d made some preliminary posts, drafted some text, and run it by David that I decided to go for it. The article turned out well enough that I decided to publish it. And so I did.

It’s one thing to figure out whether or not such and such a text/film exhibits such and such pattern when you are given the text and the pattern. That’s what ChatGPT did. Since I had already made the connection between Girard and Jaws it didn’t have to do that. I was just prompting ChatGPT to verify the connection, which it did (albeit in a weak way). That’s the kind of task we set for high school students and lower division college students. […]

I don’t really think that ChatGPT is operating at a high school level in this context. Nor do I NOT think that. I don’t know quite what to think. And I’m happy with that.

The deeper point is that there is a world of difference between what ChatGPT was doing when I piloted it into Jaws and Girard and what I eventually did when I watched Jaws and decided to look around to see what I could see. How is it that, in that process, Girard came to me? I wasn’t looking for Girard. I wasn’t looking for anything in particular. How do we teach a computer to look around for nothing in particular and come up with something interesting?

These observations are informal and are only about a single example. Given those limitations it's difficult to imagine a generalization. But I didn't hear anything from those experts that was comparably rich.

Hinton gave an example of an analogy that he posed to GPT-4 (01:18:30): “What has a compost heap got in common with an atom bomb?” It got the answer he was looking for, chain reaction, albeit at different energy levels and different rates. That's interesting. Why wasn't the panel ready with 20 such examples among them? Perhaps more to the point, doesn't Hinton see that it is one thing for GPT-4 to explain an analogy he presents to it, but that coming up with the analogy in the first place is a different kind of mental process?

Do they not have more such examples from their own work? Don't they think about their own work process, all the starts and stops, the wandering around, the dead ends and false starts, the open-ended exploration, that came before final success. And even then, no success is final, but only provisional pending further investigation. Can they not see the difference between what they do and what their machines do? Do they think all the need for exploration will just vanish in the face of machine superintelligence. Do they really believe that the universe is that small?

STILL LATER: Hinton and Hassabis on analogies

Hinton continues with analogies and Hassabis weights in:

1:18:28 – GEOFFREY HINTON: We know that being able to see analogies, especially remote analogies, is a very important aspect of intelligence. So I asked GPT-4, what has a compost heap got in common with an atom bomb? And GPT-4 nailed it, most people just say nothing.

DEMIS HASSABIS: What did it say ...

GEOFFREY HINTON: It started off by saying they're very different energy scales, so on the face of it, they look to be very different. But then it got into chain reactions and how the rate at which they're generating energy increases-- their energy increases the rate at which they generate energy. So it got the idea of a chain reaction. And the thing is, it knows about 10,000 times as much as a person, so it's going to be able to see all sorts of analogies that we can't see.

DEMIS HASSABIS: Yeah. So my feeling is on this, and starting with things like AlphaGo and obviously today's systems like Bard and GPT, they're clearly creative in ...

1:20:18 – New pieces of music, new pieces of poetry, and spotting analogies between things you couldn't spot as a human. And I think these systems can definitely do that. But then there's the third level which I call like invention or out-of-the-box thinking, and that would be the equivalent of AlphaGo inventing Go.

Well, yeah, sure, GPT-4 has all this stuff in its model, way more topics than any one human. But where’s GPT-4 going to “stand” so it can “look over” all that stuff and spot the analogies? That requires some kind of procedure. What is it?

For example, it might partition all that knowledge into discrete bits and then set up a 2D matrix with a column and a row for each discrete chunk of knowledge. Then it can move systematically through the matrix, checking each cell to see whether or not the pair in that cell is a useful analogy. What kind of tests does it apply to make that determination? I can imagine there might be a test or tests that allows a quick and dirty rejection for many candidates. But those that remain, what can you do but see if any useful knowledge follows from trying out the analogy. How long will that determination take? And so forth.

That’s absurd on the face of it. What else is there? I just explained what I went through to come up with an analogy between Jaws and Girard. But that’s just my behavior, not the mental process that’s behind the behavior. I have no trouble imagining that, in principle, having these machines will help speed up the process, but in the end I think we’re going to end up with a community of human investigators communicating with one another while they make sense of the world. The idea, which, judging from remarks he’s made elsewhere, Hinton seems to hold, that one of these days we’ll have a machine that takes humans out of the process all together, that’s an idle fantasy.

Wednesday, November 15, 2023

A dialectical view of the history of AI, Part 1: We’re only in the antithesis phase. [A synthesis is in the future.]

The idea that history proceeds by way of dialectical change is due primarily to Hegel and Marx. While I read bit of both early in my career, I haven’t been deeply influenced by either of them. Nonetheless I find the notion of dialectical change working out through history to be a useful way of thinking about the history of AI. Because it implies that that history is more than just one thing of another.

This dialectical process is generally schematized as a movement from a thesis, to an antithesis, and finally, to a synthesis on a “higher level,” whatever that is. The technical term is Aufhebung. Wikipedia:

In Hegel, the term Aufhebung has the apparently contradictory implications of both preserving and changing, and eventually advancement (the German verb aufheben means "to cancel", "to keep" and "to pick up"). The tension between these senses suits what Hegel is trying to talk about. In sublation, a term or concept is both preserved and changed through its dialectical interplay with another term or concept. Sublation is the motor by which the dialectic functions.

So, why do I think the history of AI is best conceived in this way? The first era, THESIS, running from the 1950s up through and into the 1980s, was based on top-down deductive symbolic methods. The second era, ANTITHESIS, which began its ascent in the 1990s and now reigns, is based on bottom-up statistical methods. These are conceptually and computationally quite different, opposite, if you will. As for the third era, SYNTHESIS, well, we don’t even know if there will be a third era. Perhaps the second, the current, era will take us all the way, whatever that means. Color me skeptical. I believe there will be a third era, and that it will involve a synthesis of conceptual ideas computational techniques from the previous eras.

Note, though, that I will be concentrating on efforts to model language. In the first place, that’s what I know best. More importantly, however, it is the work on language that is currently evoking the most fevered speculations about the future of AI.

Let’s take a look. Find a comfortable chair, adjust the lighting, pour yourself a drink, sit back, relax, and read. This is going to take a while.

Symbolic AI: Thesis

The pursuit of artificial intelligence started back in the 1950s it began with certain ideas and certain computational capabilities. The latter were crude and radically underpowered by today’s standards. As for the ideas, we need two more or less independent starting points. One gives us the term “artificial intelligence” (AI), which John McCarthy coined in connection with a conference held at Dartmouth in 1956. The other is associated with the pursuit of machine translation (MT) which, in the United States, meant translating Russian technical documents into English. MT was funded primarily by the Defense Department.

The goal of MT was practical, relentlessly practical. There was no talk of intelligence and Turing tests and the like. The only thing that mattered was being able to take a Russian text, feed it into a computer, and get out a competent English translation of that text. Promises was made, but little was delivered. The Defense Department pulled the plug on that work in the mid-1960s. Researchers in MT then proceeded to rebrand themselves as investigators of computational linguistics (CL).

Meanwhile researchers in AI gave themselves a very different agenda. They were gunning for human intelligence and were constantly predicting we’d achieve it within a decade or so. They adopted chess as one of their intellectual testing grounds. Thus, in a paper published in 1958 in the IBM Journal of Research and Development, Newell, Shaw, and Simon wrote that if “one could devise a successful chess machine, one would seem to have penetrated to the core of human intellectual endeavor.” In a famous paper, John McCarthy dubbed chess to be the Drosophila of AI.

Chess isn’t the only thing that attracted these researchers, they also worked on things like heuristic search, logic, and proving theorems in geometry. That is, they choose domains which, like chess, were highly rationalized. Chess, like all highly formalized systems, is grounded in a fixed set of rules. We have a board with 64 squares, six kinds of pieces with tightly specified rules of deployment, and a few other rules governing the terms of play. A seeming unending variety of chess games then unfolded from these simple primitive means according to the skill and ingenuity, aka intelligence, of the players.

This regime, termed symbolic AI in retrospect, remained in force through the 1980s and into the 1990s. However, trouble began showing up in the 1970s. To be sure, the optimistic predictions of the early years hadn’t come to pass; humans still beat computers at chess, for example. But those were mere setbacks.

These problems were deeper. While the computational linguistics were still working on machine translation, they were also interested in speech recognition and speech understanding. Stated simply, speech recognition goes like this: You give a computer a string of spoken language and it transcribes it into written language. The AI folks were interested in this as well. It’s not the sort of thing humans give a moment’s thought to; we simply do it. It is mere perception. It was proving to be surprisingly difficult. The AI folks also turned to computer vision: Give a computer a visual image and have it identify the object. That was difficult as well, even with such simple graphic objects as printed letters.

Speech understanding, however, was on the face of it intrinsically more difficult. Not only does the system have recognize the speech, but it must understand what is said. But how do you determine whether or not the computer understood what you said. You could ask it: “Do you understand?” And if it replies, “yes,” then what? You give it something to do.

That’s what the DARPA Speech Understanding Project set out to do in over five years in the early to mid 1970s. Understanding would be tested by having the computer answer questions about database entries. Three independent projects were funded; interesting and influential research was done. But those systems, interesting as they were, were not remotely as capable as Siri or Alex in our time, which run on vastly more compute encompassed in much smaller packages. We were a long way from having a computer system that could converse as fluently as a toddler, much less discourse intelligently on the weather, current events, the fall of Rome, the Mongol’s conquest of China, or how to build a fusion reactor.

During the 1980s the commercial development of AI petered out and a so-called AI Winter settled in. It would seem that AI and CL had hit the proverbial wall. The classical era, the era of symbolic computing, was all but over.

Monday, May 15, 2023

Two more AI videos: Geoffrey Hinton, Aaronson & Hanson

Perhaps I'll offer some commentary later in separate posts. I'm posting these here and now as a reminder to myself.

Wednesday, December 28, 2022

Geoffrey Hinton predicts the evolution of "neuromorphic" computers that are "mortal"

From the article:

Future computer systems, said Hinton, will be take a different approach: they will be "neuromorphic," and they will be "mortal," meaning that every computer will be a close bond of the software that represents neural nets with hardware that is messy, in the sense of having analog rather than digital elements, which can incorporate elements of uncertainty and can develop over time. 

"Now, the alternative to that, which computer scientists really don't like because it's attacking one of their foundational principles, is to say we're going to give up on the separation of hardware and software," explained Hinton. 

"We're going to do what I call mortal computation, where the knowledge that the system has learned and the hardware, are inseparable."

These mortal computers could be "grown," he said, getting rid of expensive chip fabrication plants.

"If we do that, we can use very low power analog computation, you can have trillion way parallelism using things like memristors for the weights," he said, referring to a decades-old kind of experimental chip that is based on non-linear circuit elements. 

"And also you could grow hardware without knowing the precise quality of the exact behavior of different bits of the hardware."

The new mortal computers won't replace traditional digital computers, Hilton told the NeurIPS crowd. "It won't be the computer that is in charge of your bank account and knows exactly how much money you've got," said Hinton.

"It'll be used for putting something else: It'll be used for putting something like GPT-3 in your toaster for one dollar, so running on a few watts, you can have a conversation with your toaster."*

I have a number of posts that speak to this, for example, The structured physical system hypothesis (SPSH), Polyviscous connectivity [The brain as a physical system], one of the various posts tagged with the label "polyviscous". 

*Umm, err...You're kidding, right? Who'd want to chat with their toaster? Now, your wristwatch, that's something else.

Wednesday, June 1, 2022

Interesting interview with Geoffrey Hinton [+follow-up Q&A]

The Robot Brains Podcast

Season 2 Ep 22 Geoff Hinton on revolutionizing artificial intelligence... again

240 views Jun 1, 2022 Over the past ten years, AI has experienced breakthrough after breakthrough in everything from computer vision to speech recognition, protein folding prediction, and so much more.

Many of these advancements hinge on the deep learning work conducted by our guest, Geoff Hinton, who has fundamentally changed the focus and direction of the field. A recipient of the Turing Award, the equivalent of the Nobel prize for computer science, he has over half a million citations of his work.

Hinton has spent about half a century on deep learning, most of the time researching in relative obscurity. But that all changed in 2012 when Hinton and his students showed deep learning is better at image recognition than any other approaches to computer vision, and by a very large margin. That result, that moment, known as the ImageNet moment, changed the whole AI field. Pretty much everyone dropped what they had been doing and switched to deep learning.

Geoff joins Pieter in our two-part season finale for a wide-ranging discussion inspired by insights gleaned from Hinton’s journey from academia to Google Brain. The episode covers how existing neural networks and backpropagation models operate differently than how the brain actually works; the purpose of sleep; and why it’s better to grow our computers than manufacture them.

What's in this episode:

00:00:00 - Introduction
00:02:48 - Understanding how the brain works
00:06:59 - Why we need unsupervised local objective functions
00:09:39 - Mass auto-encoders
00:10:55 - Current methods in end to end learning
00:18:36 - Spiking neural networks
00:23:00 - Leveraging spike times
00:29:55 - The story behind AlexNet
00:36:15 - Transition from pure academia to Google
00:40:23 - The secret auction of Hinton’s company at NIPS
00:44:18 - Hinton’s start in psychology and carpentry
00:54:34 - Why computers should be grown rather than manufactured
01:06:57 - The function of sleep and Boltzmann Machines
01:11:49 - Need for negative data
01:19:35 - Visualizing data using t-SNE

Links:
Geoff's Bio: https://en.wikipedia.org/wiki/Geoffrey_Hinton
Geoff's Twitter: https://twitter.com/geoffreyhinton?la...
Research and Publications: https://bit.ly/3z3M54e
Google Scholar Citations: https://bit.ly/3N892HJ
Story Behind the 2012 NIPS Auction: https://bit.ly/3t9xsIN
GLOM: https://bit.ly/3lYgWr6
Vector Institute: https://vectorinstitute.ai/

Follow-up Q/A to the previous video:

 

403 views Jun 8, 2022 Last week, we were honored to have Professor Geoff Hinton join the show for a wide-ranging discussion inspired by insights gleaned from Geoff's journey in academia, as well as past 10 years with Google Brain. The episode covers how existing neural networks and backpropagation models operate differently than how the brain actually works; the ImageNet/AlexNet breakthrough moment; the purpose of sleep; and why it’s better to grow our computers than manufacture them.

As you might recall, we also gave our audience an opportunity to contribute questions for Geoff via Twitter. We received so many amazing questions from our audience that we had to break down our time with Geoff into two parts! In this episode, we’ll discuss some of these questions with Geoff.

Tune in to get Geoff’s answers to the following questions AND MORE:

Are you concerned with AI becoming too successful?
What is the connection between mania and genius?
What childhood experiences shaped him the most?
What is next in AI?
What should PhD students focus on?
How conscious do you think today's neural nets are?
How important is embodiment for intelligence?
How does the brain work?

Monday, May 30, 2022

Eureka! Have I Found It? How to Model the Mind, that Is. [Symbols and Nets]

Since roughly the last week in April, when I applied for an Emergent Ventures grant (which was quickly, but politely, turned down), I have been working hard on revising and updating work on a system of notation which I sketched out in 2003 and posted to the web in 2010, 2011. I am referring to what I then called called an Attractor Network, but now call a Relational Network over Attractors (RNA) because I found out that neuroscientists already talk about attractor networks, which are not the same as what I’ve got in mind. The neuroscientists are referring to a network of neurons whose dynamics tend toward an attractor. I am referring to a network that specifies relationships between a very large number of attractors (hence, it is constructed over them).

Anyhow, by the time Emergent Ventures had turned me down, I was committed to the project, which has gone well so far. I had no particular expectations, just a general direction. I’ve been looking, and I’ve found some interesting things, encouraging things. Or, if you will, I’ve been puttering around, assembling bits and pieces here and there, and an interesting structure has begun to emerge.

Lamb Notation

The idea has been to develop a new notation for representing semantic structures in network form. Actually, the notation is not new; it had already been developed by Sydney Lamb in the 1960s. He developed it to model the structures of a stratificational grammer. I’ve been adapting it to model semantics.

I am doing that by assuming that the cerebral cortex is loosely divided into functionally distinct regions which I call neurofunctional areas (NFAs). The activity of these NFAs is to be modeled by complex dynamics (Walter Freeman) and a low-dimensional projection of each NFA phase space can be modeled by a conceptual space (Peter Gärdenfors). Each NFA is thus characterized by an attractor landscape.

The RNA (relational net over attractors) is a network where the nodes are logical operators (AND, OR) and the edges are basins of attraction in the NFA attractor landscapes. This is not the place to explain what that actually means, but I can give you a taste by showing you three pictures.

This is a simple semantic structure expressed in a “classical” notation from the 1970s:

It depicts the fact that both beagles and collies are varieties (VAR) of dog. The light gray nodes at the bottom are perceptual schemas, while the dark gray nodes at the right are lexemes. The white nodes are cognitive.

Here’s a fragment of one of Lamb’s networks:

The triangular nodes are AND while the brackets (both pointing up and down) are OR. The content is carried on the edges.

This RNA network takes the information expressed in the semantic network and expresses it using AND and OR nodes.

I am not even going to attempt to explain just how that works. Suffice it to say that it seems a bit more visually complicated than the old notation and thus harder to read. It also expresses more informatation. Those AND and OR nodes specify processing while no processing is specified in the classical diagram.

I am finding it more demanding to work with. In part that is because I haven’t drawn nearly so many RNA diagrams, perhaps 100 or so as compared to 1000s. But also, in drawing RNAs I have to imagine these structures being somehow laid out on a sheet of cortex, which is tricky. It would be even trickier if I were working with data about the regional functional anatomy of the cortex at my elbow, trying to figure just where each NFA is on the cortical sheet. Eventually, that will have to be done, but right now I’m satisfied just to draw some diagrams.

Crazy and Not So Crazy

The fact that I intend these diagrams as a very abstract sketch of functional cortical anatomy means that they have fairly direct empirical implications that the old diagrams never had. Of course, we were always committed to the view that we were figuring out how the human mind worked and so  eventually someone would have to figure out where and how those structures were implemented in the brain. Well, now is eventually and these new diagrams are a tool for figuring out the where and how.

And that, I suppose, is a crazy assertion. Everyone who knows anything knows that the brain is fiercely complicated and we’re never going to figure it out in a million years but anyhow we have to a waste a billion euros building a damned brain model that tells us a bit more than diddly squat, but not a whole hell of a lot more. But then what I’m doing costs nothing more than my time. Excuse the rant.

As I said, it’s crazy of me to propose a way of thinking about how high-level cognitive processes are organized in the brain. But I’m only proposing, and I’m doing it by offering a conceptual tool, a notation, that helps us think about the problem in a new way. I don’t expect that the constructions I propose are correct. I ask only that they are coherent enough to lead us to better ones.

There’s one further thing and this is not so crazy: This notation, in conjunction with 1) my assertation that it is about complex cortical dynamics, and 2) and Lev Vygotsky’s account of language development, gives us a new way of thinking about a debate that is currently blazing away in a small region of the internet: How do we model the mind, neural vectors, symbols, or both? If both, how? I am opting for both and making a fairly specific proposal about how the human brain does it. The question then becomes: What will it take to craft an artificial device that does it? If my proposal ends up taking 14K or 15K words and maybe 30 diagrams, well it deals with a very a complicated problem.

Here is the draft introduction, Symbols, holograms, and diagrams, to the working paper. With that, I’ll leave you with a brief sketch of my proposal.

The Model in 14 Propositions

1. I assume that the cortex is organized into NeuroFunctional Areas (NFAs), each of which has its own characteristic pattern of inputs and outputs. As far as I can tell, these NFAs are not sharply distinct from one another. The boundaries can be revised – think of cerebral plasticity.

2. I assume that the operations of each NFA are those of complex dynamics. I have been influenced by Walter Freeman in this.)

3. A low dimensional projection of each NFA phase space can be modeled by a conceptual space as outlined by Peter Gärdenfors.

4. Each NFA has its own attractor landscape. A primary NFA is one driven primarily by subcortical inputs. Then we have secondary and tertiary NFAs, which involve a mixture of cortical and subcortical inputs. (I’m thinking of the standard notions of primary, secondary, and tertiary cortex.)

5. Interaction between NFAs is defined by a Relational Network over Attractors (RNA), which is a relational network defined over basins in multiple linked attractor landscapes.

6. The RNA network employs a notation developed by Sydney Lamb in which the nodes are logical operators, AND & OR, while ‘content’ of the network is carried on the arcs. [REF/LINK to his paper.]

7. Each arc corresponds to a basin of attraction in some attractor landscape.

8. The output of a source NFA is ‘governed’ by an OR relationship (actually exclusive OR, XOR) over its basins. Only one basin can be active at a time. [Provision needs to be made for the situation in which no basin is entered.]

9. Inputs to a basin in a target NFA are regulated by an AND relationship over outputs from source NFAs.

10. Symbolic computation arises with the advent of language. It adds new primary attractor landscapes (phonetics & phonology, morphology?) and extends the existing RNA. Thus overall RNA is roughly divided into a general network and a lingistic network.

11. Word forms (signifiers) exist as basins in the linguistic network. A word form whose meaning is given by physical phenomena are coupled with an attractor basin (signifier) in the general network. This linkage yields a symbol (or sign). Word forms are said to index the general RNA.

12. Not all word forms are defined in that way. Some are defined by cognitive metaphor (Lakoff and Johnson). Others are defined by metalingual definition (David Hays). I assume there are other forms of definition as well (see e.g. Benzon and Hays 1990). It is not clear to me how we are to handle these forms.

13. Words can be said to index the general RNA (Benzon & Hays 1988).

14. The common-sense concept of thinking refers to the process by which one uses indices to move through the general RNA to 1) add new attractors to some landscape, and 2) construct new patterns over attractors, new or existing.