Showing posts with label common sense knowledge. Show all posts
Showing posts with label common sense knowledge. Show all posts

Tuesday, November 14, 2023

Enhancing LM common-sense reasoning through neurosymbolic techniques

Abstract of paper linked in the GitHub site:

Yuling Gu and Bhavana Dalvi Mishra and Peter Clark, Do language models have coherent mental models of everyday things?, arXiv:2212.10029v3 [cs.CL].

When people think of everyday things like an egg, they typically have a mental image associated with it. This allows them to correctly judge, for example, that “the yolk surrounds the shell” is a false statement. Do language models similarly have a coherent picture of such everyday things? To investigate this, we propose a benchmark dataset consisting of 100 every- day things, their parts, and the relationships between these parts, expressed as 11,720 “X relation Y?” true/false questions. Using these questions as probes, we observe that state-of- the-art pretrained language models (LMs) like GPT-3 and Macaw have fragments of knowledge about these everyday things, but do not have fully coherent “parts mental models” (54- 59% accurate, 19-43% conditional constraint violation). We propose an extension where we add a constraint satisfaction layer on top of the LM’s raw predictions to apply common-sense constraints. As well as removing inconsistencies, we find that this also significantly improves accuracy (by 16-20%), suggesting how the incoherence of the LM’s pictures of everyday things can be significantly reduced.

Thursday, December 29, 2022

Yejin Choi on common sense and value pluralism in AI

David Marchese, An A.I. Pioneer on What We Should Really Fear, NYTimes, December 21, 2022.

Common sense

Can you explain what “common sense” means in the context of teaching it to A.I.? A way of describing it is that common sense is the dark matter of intelligence. Normal matter is what we see, what we can interact with. We thought for a long time that that’s what was there in the physical world — and just that. It turns out that’s only 5 percent of the universe. Ninety-five percent is dark matter and dark energy, but it’s invisible and not directly measurable. We know it exists, because if it doesn’t, then the normal matter doesn’t make sense. So we know it’s there, and we know there’s a lot of it. We’re coming to that realization with common sense. It’s the unspoken, implicit knowledge that you and I have. It’s so obvious that we often don’t talk about it. For example, how many eyes does a horse have? Two. We don’t talk about it, but everyone knows it. We don’t know the exact fraction of knowledge that you and I have that we didn’t talk about — but still know — but my speculation is that there’s a lot. Let me give you another example: You and I know birds can fly, and we know penguins generally cannot. So A.I. researchers thought, we can code this up: Birds usually fly, except for penguins. But in fact, exceptions are the challenge for common-sense rules. Newborn baby birds cannot fly, birds covered in oil cannot fly, birds who are injured cannot fly, birds in a cage cannot fly. The point being, exceptions are not exceptional, and you and I can think of them even though nobody told us. It’s a fascinating capability, and it’s not so easy for A.I.

Value Pluralism

So what’s most exciting to you right now about your work in A.I.? I’m excited about value pluralism, the fact that value is not singular. Another way to put it is that there’s no universal truth. A lot of people feel uncomfortable about this. As scientists, we’re trained to be very precise and strive for one truth. Now I’m thinking, well, there’s no universal truth — can birds fly or not? Or social and cultural norms: Is it OK to leave a closet door open? Some tidy person might think, always close it. I’m not tidy, so I might keep it open. But if the closet is temperature-controlled for some reason, then I will keep it closed; if the closet is in someone else’s house, I’ll probably behave. These rules basically cannot be written down as universal truths, because when applied in your context versus in my context, that truth will have to be bent. Moral rules: There must be some moral truth, you know? Don’t kill people, for example. But what if it’s a mercy killing? Then what? [...]

Is the ultimate hope that A.I. could someday make ethical decisions that might be sort of neutral or even contrary to its designers’ potentially unethical goals — like an A.I. designed for use by social media companies that could decide not to exploit children’s privacy? Or is there just always going to be some person or private interest on the back end tipping the ethical-value scale? The former is what we wish to aspire to achieve. The latter is what actually inevitably happens. In fact, Delphi is left-leaning in this regard because many of the crowd workers who do annotation for us are a little bit left-leaning. Both the left and right can be unhappy about this, because for people on the left Delphi is not left enough, and for people on the right it’s potentially not inclusive enough. But Delphi was just a first shot. There’s a lot of work to be done, and I believe that if we can somehow solve value pluralism for A.I., that would be really exciting. To have A.I. values not be one systematic thing but rather something that has multidimensions just like a group of humans. [...]

Could it be that if humans are in situations where we’re relying on A.I. to make moral decisions then we’ve already screwed up? Isn’t morality something we probably shouldn’t be outsourcing in the first place? You’re touching on a common — sorry to be blunt — misunderstanding that people seem to have about the Delphi model we made. It’s a Q. and A. model. We made it clear, we thought, that this is not for people to take moral advice from. This is more of a first step to test what A.I. can or cannot do. My primary motivation was that A.I. does need to learn moral decision-making in order to be able to interact with humans in a safer and more respectful way.

Take that, Nick Bostrom!

Like the Nick Bostrom paper clip example, which I know is maybe alarmist. But is an example like that concerning? No, but that’s why I am working on research like Delphi and social norms, because it is a concern if you deploy stupid A.I. to optimize for one thing. That’s more of a human error than an A.I. error. But that’s why human norms and values become important as background knowledge for A.I. Some people naïvely think if we teach A.I. “Don’t kill people while maximizing paper-clip production,” that will take care of it. But the machine might then kill all the plants. That’s why it also needs common sense. It’s common sense not to kill all the plants in order to preserve human lives; it’s common sense not to go with extreme, degenerative solutions.

There’s more in the interview.

Sunday, August 28, 2022

Elemental Cognition is ready to deploy hybrid AI technology in practical systems

Steve Lohr, One Man's Dream of Fusing A.I. With Common Sense, NYTimes, Aug. 28, 2022

David Ferrucci is best-known as the researcher who led the team that developed IBM's Watson, which beat the best human players of Jeopardy in 2011. He left IBM a year later and formed his own company, Elemental Cognition, in 2015. Elemental cognition is taking a hybrid approach, combining aspects of machine learning and symbolic computation.

Elemental Cognition has recently developed a system that helps people plan and book round-the-world airline tickets:

The round-the-world ticket is a project for oneworld, an alliance of 13 airlines including American Airlines, British Airways, Qantas, Cathay Pacific and Japan Airlines. Its round-the-world tickets can have up to 16 different flights with stops of varying lengths over the course of a year.

Elemental Cognition supplies the technology behind a trip-planning intelligent agent on oneworld’s website. It was developed over the past year and introduced in April.

The user sees a global route map on the left and a chatbot dialogue begins on the right. A traveler starting from New York types in the desired locations — say, London, Rome and Tokyo. “OK,” replies the chatbot, “I have added London, Rome and Tokyo to the itinerary.”

Then, the customer wants to make changes — “add Paris before London,” and “replace Rome with Berlin.” That goes smoothly, too, before the system moves on to travel times and lengths of stays in each city.

Rob Gurney, chief executive of oneworld, is a former Qantas and British Airways executive familiar with the challenges of online travel planning and booking. Most chatbots are rigid systems that often repeat canned answers or make irrelevant suggestions, a frustrating “spiral of misery.”

Instead, Mr. Gurney said, the Elemental Cognition technology delivers a problem-solving dialogue on the fly. The rates of completing an itinerary online are three to four times higher than without the company’s software.

Elemental Cognition has developed an approach that all-but eliminates the hand-coding typcial of symbolic A.I.:

For example, the rules and options for a global airline ticket are spelled out in many pages of documents, which are scanned.

Dr. Ferrucci and his team use machine learning algorithms to convert them into suggested statements in a form a computer can interpret. Those statements can be facts, concepts, rules or relationships: Qantas is an airline, for example. When a person says “go to” a city, that means add a flight to that city. If a traveler adds four more destinations, that adds a certain amount to the cost of the ticket.

In training the round-the-world ticket assistant, an airline expert reviews the computer-generated statements, as a final check. The process eliminates most of the need for hand coding knowledge into a computer, a crippling handicap of the old expert systems.

There's more at the link.

* * * * *

Lex Fridman interviews David Ferrucci (2019).

0:00 - Introduction
1:06 - Biological vs computer systems
8:03 - What is intelligence?
31:49 - Knowledge frameworks
52:02 - IBM Watson winning Jeopardy
1:24:21 - Watson vs human difference in approach
1:27:52 - Q&A vs dialogue
1:35:22 - Humor
1:41:33 - Good test of intelligence
1:46:36 - AlphaZero, AlphaStar accomplishments
1:51:29 - Explainability, induction, deduction in medical diagnosis
1:59:34 - Grand challenges
2:04:03 - Consciousness
2:08:26 - Timeline for AGI
2:13:55 - Embodied AI
2:17:07 - Love and companionship
2:18:06 - Concerns about AI
2:21:56 - Discussion with AGI

Tuesday, April 5, 2022

PaLM: Scaling Language Modeling with Pathways [common sense, explains jokes]

Since it has jokes within its repertoire, I wonder if it could explain Seinfeld's tramway bit?

Abstract from the linked paper (second link in the tweet):

Large language models have been shown to achieve remarkable performance across a variety of natural language tasks using few-shot learning, which drastically reduces the number of task-specific training examples needed to adapt the model to a particular application. To further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model (PaLM).

We trained PaLM on 6144 TPU v4 chips using Pathways, a new ML system which enables highly efficient training across multiple TPU Pods. We demonstrate continued benefits of scaling by achieving state-of- the-art few-shot learning results on hundreds of language understanding and generation benchmarks. On a number of these tasks, PaLM 540B achieves breakthrough performance, outperforming the finetuned state- of-the-art on a suite of multi-step reasoning tasks, and outperforming average human performance on the recently released BIG-bench benchmark. A significant number of BIG-bench tasks showed discontinuous improvements from model scale, meaning that performance steeply increased as we scaled to our largest model. PaLM also has strong capabilities in multilingual tasks and source code generation, which we demonstrate on a wide array of benchmarks. We additionally provide a comprehensive analysis on bias and toxicity, and study the extent of training data memorization with respect to model scale. Finally, we discuss the ethical considerations related to large language models and discuss potential mitigation strategies.

Thursday, February 17, 2022

Sean Carroll interviews Gary Marcus about AI and common sense

As you may know, Marcus is skeptical about the ability of (deep) learning approaches to go all the way. Here's one bit of the conversation:

0:11:58.7 Sean Carroll: And maybe it’s good to… We’re able to get into details a little bit. The audience likes the details. So let’s try to understand why there has been this progress. And as far as I can tell, the overwhelming majority of recent progress in AI has been driven by neural networks and deep learning algorithms. Is that fair? And what does that mean?

0:12:18.9 Gary Marcus: It’s true, but with some caveats. So, first of all, there are older techniques that everybody takes for granted but are real and are already out there. Second of all, there are things like AlphaGo, they’re actually hybrid models that use classical tree search techniques enhanced with Monte Carlo techniques in order to do what they’re doing. So they’re not just straight multi-layer perception is a kind of stereotype that people have with neural networks. We have some inputs, they feed into a hidden layer that does some summation and activation function goes to an output. They’re not just that. They actually borrow some important ideas about search, for example, and symbols from classical AI. And so they’re actually hybrid systems and people don’t acknowledge that. So this is the second caveat I would give you. The third caveat I would give you… We can come back to the second, but the third caveat I’ll give you is, yeah, most of the progress has been with deep learning lately, but most of the money has been there too, and it was really interesting to see… And I don’t just mean like 60% versus 40%. I mean, like 99.9% of the investment right now literally is in deep learning and classic symbol manipulation AI is really out of favor, and people like Geoff Hinton don’t spend any money on it at all. And so it was really interesting.

0:13:42.7 GM: There was this competition presented at the NeurIPS Conference which is the biggest conference these days in the AI field just a month or so ago, on a game called NetHack, it has various complications in it, and a symbolic system actually won in an upset victory over all this deep learning stuff. And so if you look back at the history of AI, in the history of science more generally, sometimes things get counted out too soon. It is true the deep learning has made a bunch of progress, but the question is, what follows from there?

0:14:13.9 SC: No, I’m not actually trying to make any value judgements. I would like to explain for our audience what the options are. What do you mean by deep learning? What is that and what is that in comparison to symbolic manipulation?

0:14:25.2 GM: So deep learning is fundamentally a way of doing statistical analysis on large quantities of data, at least that’s… It’s Forte. You can actually use it in a bunch of different ways, but most of the progress has come from that. And what’s impressive about the recent work is it allows us to learn from very large quantities of data. The classical AI system really didn’t do a lot of learning at all. They’re mostly hand-coded and sometimes that’s the right thing to do. So we don’t need to learn how to do navigation. We need to learn some details, but we don’t need to learn how to do navigation for the purpose of one of the most useful AI things out there, which is route planning, telling you how to get home from whatever crazy place you wound up in. Right? That’s not a deep learning-driven system. But there are other systems where if you can glom on to all the data that’s out there, you can solve certain problems very effectively, and that’s what deep learning has been good for.

The way forward:

0:20:35.4 GM: Yeah, well let me, before I do that, let me say that I think that we need elements of the symbolic approach, I think we need elements of the deep learning approach or something like it, but the… Neither by itself is sufficient. And so, I’m a big fan of what I call hybrid systems that bring together in ways that we haven’t really even figured out yet, the best of both worlds, but with that preface, ’cause people often in the field like to misrepresent me as that symbolic guy, and I’m more like the guy who said, Don’t forget about the symbolic stuff, we need it to be part of the answer. Okay, so the symbolic stuff is basically the essence of computer programming your algebra or something like that, what’s really about is having functions where you have variables that you bind to particular instances and calculate the values out, so simplest example would be an equation in Algebra, Y equals X plus 2, I tell you what X is, you can figure out what y is… And there it doesn’t matter, which Xs you have seen before, you have this thing that is defined universally is the way a logician might put it, universally for everything in some domain, any physicist would grasp that immediately or any program or any logician.

Note: Marcus remarks here and there that some systems that are presented as deep learning systems, such as AlphaGo (chess playing), Rubik's cube, or protein-folding, are actually hybrid systems, employing aspects of symbolic technology.

There's more at the link.

H/t 3QD.

Tuesday, June 1, 2021

Some thoughts on why systems like GPT-3 will always have trouble with common sense knowledge

I develop an analogical argument about why natural language systems trained only on text will never be able to deal with common-sense reasoning. I begin by presenting Herbert Simon’s famous parable of the ant and follow it with some information about sensory deprivation. From those I conclude that our mental apparatus depends on access to the world to achieve stability. I then tap-dance my way to the assertion that common sense reasoning depends on those sensory motor systems which are, in turn, dependent on the world.

AI language engines are enmeshed in language, with no access to the physical world. Consequently common sense reasoning will forever be elusive. Common sense grounds us in the physical world.

Simon’s ant

In Chapter 3, “The Psychology of Thinking: Embedding Artifice in Nature,” of The Sciences of the Artificial (2nd Ed., 1981), Herbert Simon gives us a parable, a story to think with. Simon asks us to imagine an ant moving about on a beach:

We watch an ant make his laborious way across a wind- and wave-molded beach. he moves ahead, angles to the right to ease his climb up a steep dunelet, detours around a pebble, stops for a moment to exchange information with a compatriot. Thus he makes his weaving, halting way back to his home. So as not to anthropomorphize about his purposes, I sketch the path on a piece of paper. It is a sequence of irregular, angular segments – not quite a random walk, for it has an underlying sense of direction, of aiming toward a goal.

After introducing a friend, to whom he shows the sketch and to whom he addresses a series of unanswered questions about the sketched path, Simon goes on to observe:

Viewed as a geometric figure, the ant’s path is irregular, complex, hard to describe. But its complexity is really a complexity in the surface of the beach, not a complexity in the ant. On that same beach another small creature with a home at the same place as the ant might well follow a very similar path.

That is, because the beach has a complex surface, the ant is able to walk a complex path on that surface using rather simple mechanisms. In posing this parable Simon is, of course, asking us to think of the beach as the world in full and that ant is us. Relative to the world’s complexity, our conceptual apparatus is relatively simple.

I would like to propose that the nervous system requires environmental support if it is to maintain its physical stability and coherence. Note that Simon was not at all interested in the physical requirements of the nervous system. Rather, he was interested in suggesting that we can get complex behavior from relatively simple devices, and simplicity translates into design requirements for a nervous system. That’s fine, but I’m suggesting that the nervous system actively seeks out the world and so is dependent upon finding it, in a more or less orderly fashion.

Our sensory systems don’t ‘represent’ (if that’s the right word, many reject it) the world in great detail. Their apprehension of the world is rough and ready. One doesn't need to represent apples and oranges in full detail in order to distinguish them, nor cats and dogs, cars and bicycles, and so forth. Our systems need only ‘grab on’ to the things and events in the world. The world itself will ‘fill out’ our perceptions in real time.

Now, consider this variation on Simon’s story. What would happen if we put the ant on an absolutely featureless surface and let it walk about? What kind of paths would it trace then? As that surface lacks any of the normal cues in the ant’s environment I would imagine the ant would either not move at all or move in a genuinely random or perhaps a rigidly stereotypic way (e.g. around and around in a circle). Or perhaps the ant would hallucinate.

Sensory deprivation

That is what seems to happen to humans when we are deprived of sensory input. Early on in The Ghost Dance, a classic anthropological study of the origins of religion, Weston La Barre considers what happens under various conditions of deprivation. Consider this passage about Captain Joshua Slocum, who sailed around the world alone at the turn of the 20th Century:

Once in a South Atlantic gale, he double-reefed his mainsail and left a whole jib instead of laying-to, then set the vessel on course and went below, because of a severe illness. Looking out, he suddenly saw a tall bearded man, he thought at first a pirate, take over the wheel. this man gently refused Slocum’s request to take down the sails and instead reassured the sick man he would pilot the boat safely through the storm. Next day Slocum found his boat ninety-three miles further along on a true course. That night the same red-capped and bearded man, who said he was the pilot of Columbus’ Pinta, came again in a dream and told Slocum he would reappear whenever needed.

La Barre goes on to cite similar experiences happening to other explorers and to people living in isolation, whether by choice, as in the case of religious meditation, or force, as in the case of prisoners being brainwashed.

In the early 1950s Woodburn Heron, a psychologist in the laboratory of Donald Hebb, conducted some of the earliest research on the effects of sensorimotor deprivation [2]. The subjects were placed on a bed in a small cubicle. They wore translucent goggles that transmitted light, but no visual patterns. Sound was masked by the pillow on which they rested their heads and by the continuous hum of air-conditioning equipment. Their arms and hands were covered with cardboard cuffs and long cotton gloves to blunt tactile perception. They stayed in the cubicle as long as they could, 24 hours a day, with brief breaks for eating and going to the bathroom.

The results were simple and dramatic. Mental functioning as measured by simple tests administered after 12, 24, and 48 hours or isolation deteriorated. Subjects lost their ability to concentrate and to think coherently. Most dramatically, subjects began hallucinating. They would begin with simple forms and designs and evolve into whole scenes. One subject saw dogs, another saw eyeglasses, and they had little control over what they saw; no matter how hard they tried, they couldn’t change what they were seeing. A few subjects had auditory and tactile hallucinations. Upon emerging from isolation the visual world appeared distorted with some subjects reporting that the room appeared to be moving. Woodburn concluded, as have other investigators, that the waking brain requires a constant flux of sensory input in order to function properly.

Of course, one might object to this conclusion by pointing out that, in particular, these people were deprived interaction with other people and that is what causes the instability, not mere sensory deprivation. But, from our point of view, that is no objection at all. For other people are a major part of the environment in which human beings live. The rhythms of our intentional structures are stable only if they are supported by the rhythms of the external world. Similarly, one might object that, while these people were cut off from the external physical world, their brains, of course, were still operating in the interior milieu. Consequently the instabilities they experienced reflect “pressure” from the interior milieu that is not balanced by activity in the external world. This may well be true, I suspect that it is, but it is no objection to the idea that the waking brain requires constant input from the external world in order to remain stable. Rather, this is simply another aspect of that requirement.

Thus I suggest that detaching one’s attention from the immediate world to “think” may cause problems. And yet it is the capacity for such thought that is one aspect of the mental agility that distinguishes us from our more primitive ancestors. How do we keep the nervous system stable enough to think coherently? The answer to that question depends, of course, on just what is causing the instability. Part of the answer may well be that we periodically “tune” our cortical circuits through music and dance. As long as the brain gets such tuning on a regular basis it can maintain its stability well enough during episodes of extending thinking--whether the merest day dreaming, or concentrated intellectual activity of one sort or another. But, without regular tuning, the brain begins to lose its stability.

What does that have to do with GPT-3?

GPT-3, and many other contemporary systems, is trained on large bodies of text. That text, of course, consists of words. Well, they are words to us, who can say them, spell them, offer definitions, and use them correctly in utterances and writing. To the computer those are merely word forms, symbols that are not attached to meanings in the way that word forms are attached to meanings in the human mind. The object of these systems is to somehow approximate word meanings by calculating over the distribution of words in texts. The underlying assumption, which Warren Weaver articulated in his famous memo on machine translation back in 1949 [1], is that words that appear together share some aspect of meaning. Thus if we can ‘examine’ words in a sufficiently large number of contexts each, we should be able to approximate their meanings.

And indeed, GPT-3’s ability to generate coherent text of some length suggests it has manages a remarkable approximation. But it exhibits failings in commonsense reasoning [3], a failing that the symbolic systems of four decades ago exhibited as well. I believe that much commonsense reasoning takes place ‘close to the ground’ as it were [4]. Because GPT-3 only has access to word forms, not to the sensory-motor schemas that directly support, give meaning to, many of them it lacks the basis on which common sense reasoning functions. It is, in effect, lost.

It is in the situation Slocum found himself when isolated at sea. Lacking people to talk to, his mind began unraveling. And so it happens with people in sensory deprivation. Without sensory input to stabilize their perceptual system they begin hallucinating.

Let’s return to Simon’s ant. There is the world, and there is the ant with its mental apparatus. The human situation is more complex. There is the world, and there is our sensorimotor apparatus. And for the first two years of life, that’s pretty much it. But then language begins to develop, and language depends on both those sensorimotor systems and on interaction with others. GPT-3 lacks both sensorimotor apparatus and conversation with others. That is, during training GPT-3 has no contact with human interlocutors. Once trained GPT-3 is given prompts to which it responds. But, it is my understanding that it does not learn from those interactions.

How can GPT-3, and similar engines, possibly make sense of language that functions ‘close to the world’? I suppose one can hope that by considering a very very large number of texts an AI Engine can somehow ‘fill in’ the information it is missing because it lacks direct access to the world. That has not worked so far. What reason do we have think that some day it will?

Consider Simon’s ant once again. By examining the paths it traces we can approximate the beach’s micro-geography. How many paths must we examine and superimpose in order fix the location of every pebble and dunelet – forget about grains of sand?

Now, until AI language systems and rich and flexible access to the world and the capacity to develop a rich analog or quasi-analog account of the world, until that happens, common sense reasoning will remain elusive.

References

[1] Warren Weaver, “Translation”, Carlsbad, NM, July 15, 1949, 12. pp. Online: http://www.mt-archive.info/Weaver-1949.pdf.

[2] Heron, W. (1957). The pathology of boredom. Scientific American, 196, 52–56. https://doi.org/10.1038/scientificamerican0157-52.

[3] This has been much discussed in the literature. I have offered a modest example in a post where I had GPT-3 explain the punch line to a Jerry Seinfeld joke: Analyze This! Screaming on the flat part of the roller coaster ride [Does GPT-3 get the joke?], May 7, 2021, https://new-savanna.blogspot.com/2021/05/analyze-this-screaming-on-flat-part-of.html.

[4] I have argued this in a working paper, GPT-3: Waterloo or Rubicon? Here be Dragons, Version 2, August 20, 2020, 34 pp., https://www.academia.edu/43787279/GPT_3_Waterloo_or_Rubicon_Here_be_Dragons_Version_2. See pp. 21-25. See also my post, Computation, Mind, and the World [bounding AI], September 28, 2019, https://new-savanna.blogspot.com/2019/12/computation-mind-and-world-bounding-ai.html.

Friday, May 7, 2021

Analyze This! Screaming on the flat part of the roller coaster ride [Does GPT-3 get the joke?]

Here’s Seinfeld’s first television appearance. It is from 1977 on Celebrity Cabaret, a nationally syndicated show. He’s doing a bit that starts with the Roosevelt Island tramway. You know what that is?

Background knowledge and common sense

Or don’t you? Just to be sure Seinfeld helpfully explains what a tramway is. What it is, really, is something he uses to set up the joke, but that’s not what I’m interested in. I’m interested in background knowledge, often known as common sense knowledge in the rarified world of artificial intelligence (AI).

If you look at the version of this bit that Seinfeld published in his book, Is This Anything?, you’ll see that he doesn’t explain the tramway at all (I’ve placed it immediately below). He just assumes – Bam! – you know it is. He also assumes you know that the South Bronx is a rather sketchy neighborhood – back in the day it was said that “the Bronx is burning.” Because sometimes it was. But if you didn’t already know that you wouldn’t get the joke.

But Seinfeld doesn’t say “the South Bronx” in the version in the clip. He simply refers to “the ghetto.” That tells you want you need to know to get the joke. Though I’ve not consulted him on this, I assume he did that because he figured that most people in a national audience would not know about the sad state of the South Bronx. New Yorkers would know that; it’s background knowledge for them – unless of course they’re actually in the South Bronx, in which case it’s in their face. But others are not likely to know that.

So that’s what interests me, the background knowledge, the common sense knowledge, that holds the bit together. You also have to know that roller coasters go up and down (did you notice the gesture he made during the bit?), that they’re a little scary on the downslope, that bankruptcy isn’t consistent with amusement park rides, that cities have governments and that it’s those governments that do things, etc. We know all this stuff without thinking about it.

But computers do not. So we’re going to quiz a computer about the punch line.

The Bit: Roosevelt Island Tramway

I see they just finished the Roosevelt Island Tramway.

That’s nice…

The city’s going bankrupt,

they’re putting up rides for us.

Next thing you know, there’ll be a roller coaster through the South Bronx.

That would be the first roller coaster where the people scream on the flat part of the ride.

Analysis and Commentary from GPT-3

GPT-3 is a massive AI engine that OpenAI revealed in June 2020. Massive? It’s got 175 billion parameters. Parameters? Don’t worry about it. Think of a parameter as a control knob on your TV – do TV’s have knobs anymore? Imagine a TV with 175 billion, with a “B”, knobs. That’s a lot of twiddling. It was trained on 499 billion tokens (think, roughly, words) of text sucked in from the internet. So it has processed text about Roosevelt Island, tramways, roller coasters, the South Bronx and a lot of other things, such as Antarctica, green cheese, Mata Hari, redwoods, giant Komodo dragons, fried rice, the Saturn V rocket, Jerry Seinfeld, pencils, The Tale of Genji, whales, catapults, looms, fruit flies, and anything else laying around on the Internet.

What GPT-3 does is to produce text. You feed it a bit of text and it takes it and runs with it by producing a continuation of the text. The text it produces is often astonishingly natural. It seems like a human wrote it. But no human did, just a dumb, but huge, just ginormous, machine.

If you feed GPT-3 some text and then ask a question, it will answer the question. That’s what I did with GPT-3. Well, not me, my friend Phil Mohun is the one who quizzed the machine. He gave it Seinfeld’s bit, and then asked a question. GPT-3 answered. Since I was particularly interested in the last line, the punch line, that’s what Phil asked about.

1] Here’s the first interaction. I’ve set Phil’s words in boldface.

Jerry Seinfeld tells the following joke:

I see they just finished the Roosevelt Island Tramway.
That’s nice…
The city’s going bankrupt,
they’re putting up rides for us.
Next thing you know, there’ll be a roller coaster through the South Bronx.
That would be the first roller coaster where the people scream on the flat part of the ride.

Q: Why would the people scream on the flat part of the ride?
GPT-3: Because it’s the Bronx.

Well, yeah. So what? I wanted to see if GPT-3 understood the relationship between being scared on a roller coaster and being scared on the ground in the South Bronx. So Phil tried again.

2] He increases the “temperature.” Don’t worry about just what that is, but increasing the temperature raises the amount of randomness in GPT-3’s response. Note that on this try Phil also gave GPT-3 the whole bit, as in example 1. That’s true of the other examples as well, but I’m omitting that part because it’s redundant. Just assume GPT-3 got the whole deal.

Q: Why would the people scream on the flat part of the ride?
GPT-3: There’ll be nothing to see for miles but rubble.

We’re getting somewhere. It seems like GPT-3 “knows” something about the South Bronx. But what does rubble have to do with screaming?

3] Here’s another response:

Q: Why would the people scream on the flat part of the ride?
GPT-3: So the people in Queens can see them.

That’s very interesting, and indicates some knowledge of geography as Queens and the South Bronx adjoin one another. But the roller coaster would have to be pretty close to Queens in order for the screams to reach people there. How far does GPT-3 think screams can be heard, under what wind conditions, and so forth. It’s not clear what’s going on inside GPT-3 – why would people on the roller coaster want to be noticed by people in Queens? – but the reply is not flat-out nonsense.

Monday, July 27, 2020

1. No meaning, no how: GPT-3 as Rubicon and Waterloo, a personal view

I say that not merely because I am a person and, as such, I have a point of view on GPT-3, and related matters. I say because the discussion is informal, without journal-class discussion of this, that, and the others, along with the attendant burden of citation, though I will offer a few citations. More over, I’m pretty much making this up as I go along. That is to say, I am trying to figure out just what it is that I think, and see value in doing so in public.

What value, you ask? It commits me to certain ideas, if only at a certain time. It lays out a set of priors and thus serves to sharpen my ideas developments unfold and I, inevitably, reconsider.

GPT-3 represents an achievement of a high order; it deserves the attention it has received, if not the hype. We are now deep in “here be dragons” territory and we cannot go back. And yet, if we are not careful, we’ll never leave the dragons, we’ll always be wild and undisciplined. We will never actually advance; we’ll just spin faster and faster. Hence GPT-3 is both a Rubicon, the crossing of a threshold, and a potential Waterloo, a battle we cannot win.

Here’s my plan: First we take a look at history, at the origins of machine translation and symbolic AI. Then I develop a fairly standard critic of semantic models such as those used in GPT-3 which I follow with some remarks by Martin Kay, one of the Grand Old Men of computational linguistics. Then I look at the problem of common sense reasoning and conclude be looking ahead to the next post in this series in which I offer some speculations on why (and perhaps even how) these models can succeed despite their sever and fundamental short-comings.

Background: MT and Symbolic computing

It all began with a famous memo Warren Weaver wrote in 1949. Weaver was director of the Natural Sciences division of the Rockefeller Foundation from 1932 to 1955. He collaborated Claude Shannon in the publication of a book which popularized Shannon’s seminal work in information theory, The Mathematical Theory of Communication. Weaver’s 1949 memorandum, simply entitled “Translation” [1], is regarded as the catalytic document in the origin of machine translation (MT) and hence of computational linguistics (CL) and heck! why not? artificial intelligence (AI).

Let’s skip to the fifth section of Weaver’s memo, “Meaning and Context” (p. 8):
First, let us think of a way in which the problem of multiple meaning can, in principle at least, be solved. If one examines the words in a book, one at a time as through an opaque mask with a hole in it one word wide, then it is obviously impossible to determine, one at a time, the meaning of the words. “Fast” may mean “rapid”; or it may mean "motionless"; and there is no way of telling which.

But if one lengthens the slit in the opaque mask, until one can see not only the central word in question, but also say N words on either side, then if N is large enough one can unambiguously decide the meaning of the central word. The formal truth of this statement becomes clear when one mentions that the middle word of a whole article or a whole book is unambiguous if one has read the whole article or book, providing of course that the article or book is sufficiently well written to communicate at all.
It wasn’t until the 1960s and ‘70s that computer scientists would make use of this insight; Gerard Salton was the central figure and he was interested in document retrieval [2]. Salton would represent documents as a vector of words and then query a database of such representation by using a vector composed from user input. Documents were retrieved as a function of similarity between the input query vector and the stored document vector.

Work on MT went a different way. Various approaches were used, but at some relatively early point researchers were writing formal grammars of languages. In some cases these grammars were engineering conveniences while in others they were taken to represent the mental grammars of humans. In any event, that enterprise fell apart in the mid-1960s. The prospects for practical results could not justify federal funding and the government had interest in supporting purely scientific research into the nature of language.

But such research continued nonetheless, sometimes under the rubric of computational linguistics (CL) and sometimes as AI. I encountered CL in graduate school in the mid-1970s when I joined the research group of David Hays in the Linguistics Department of the State University of New York at Buffalo – I was actually enrolled as a graduate student in English; it’s complicated.

Many different semantic models were developed, but I’m not interested in anything like a review of that work, just a little taste. In particular I am interested in a general type of model was known as a semantic or cognitive network. Hays had been developing such a model for some years in conjunction with several graduate students [2]. Here’s a fragment of a network from a system developed by one of those students, Brian Phillips, to tell whether or not stories of people drowning were tragic [3]. Here’s a representation of capsize:
Notice that there are two kinds of nodes in the network, square ones and smaller round ones. The square ones represent a scene while the round ones represent individual objects or events. Thus the square node at the upper left indicates a scene with two sub-scenes – I’m just going to follow out the logic of the network without explaining it in any detail. The first one asserts that there is a boat that contains one Horatio Smith. The second one asserts that the boat overturns. And so forth through the rest of the diagram.

This network represents semantic structure. In the terminology of semiotics, it represents a network of signifieds. Though Phillips didn’t do so, it would be entirely possible to link such a semantic network with a syntactic network, and many systems of that era did so.

Such networks were symbolic in the (obvious) sense that the objects in them were considered to be symbols, not sense perceptions or motor actions nor, for that matter, neurons, whether real or artificial. The relationship between such systems and the human brain was not explored, either in theory or in experimental observation. It wasn’t an issue.

That enterprise collapsed in the mid-1980s. Why? The models had to be hand-coded, which took time. They were computationally expensive and so-called common sense reasoning proved to be endless, making the models larger and larger. (I discuss common sense below and I have many posts at New Savanna on the topic [4].)

Oh, the work didn’t stop entirely. Some researchers kept at it. But interests shifted toward machine learning techniques and toward artificial neural networks. That is the line of evolution that has, three or four decades later, resulted in systems like GPT-3, which also owe a debt to the vector semantics pioneered by Salton. Such systems build huge language models from huge databases – GPT-3 is based on 500 billion tokens [5] – and contain no explicit models of syntax or semantics anywhere, at least not that researchers can recognize.

Researchers build a system that constructs a language model (“learns” the language), but the inner workings of that model are opaque to the researchers. After all, the system built the model, not the researchers. They only built the system.

It is a strange situation.

Thursday, December 19, 2019

AI at its best, pratfalls and all [the common sense problem is the resistance that the world presents to us]

Once again Mark Liberman presents us with an example of modern AI at its best, pratfalls and all, Mrs. Transformer-XL Tittlemouse, December 18, 2019. It is about this article:
Zihang Da et al., "Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context", arXiv 6/2/2019.

Abstract: Transformers have a potential of learning longer-term dependency, but are limited by a fixed-length context in the setting of language modeling. We propose a novel neural architecture Transformer-XL that enables learning dependency beyond a fixed length without disrupting temporal coherence. It consists of a segment-level recurrence mechanism and a novel positional encoding scheme. Our method not only enables capturing longer-term dependency, but also resolves the context fragmentation problem. As a result, TransformerXL learns dependency that is 80% longer than RNNs and 450% longer than vanilla Transformers, achieves better performance on both short and long sequences, and is up to 1,800+ times faster than vanilla Transformers during evaluation. Notably, we improve the state-ofthe-art results of bpc/perplexity to 0.99 on enwiki8, 1.08 on text8, 18.3 on WikiText-103, 21.8 on One Billion Word, and 54.5 on Penn Treebank (without finetuning). When trained only on WikiText-103, Transformer-XL manages to generate reasonably coherent, novel text articles with thousands of tokens. Our code, pretrained models, and hyperparameters are available in both Tensorflow and PyTorch.
Liberman presents an example of Generated Text from Appendix E:
= The Tale of Mrs. Tittlemouse =
The Tale of Mrs. Tittlemouse is a children's book written and illustrated by H. M. Potter. It was first published in February 1910 by Frederick Warne & Co. and included illustrations by Frederick Warne. Mrs. Tittlemouse was intended to be a children's book, rather than a musical, and the tale describes her attempts to find her own home. Mrs. Tittlemouse is an average kitten and the tale follows a mischievous rabbit named Mr. Tittlemouse who tries to kill her in an attempt to get a ride on a horse. Potter later wrote of the rabbit and her attempts to kill her, " If [ she ] were a fox, I should like to have been able to show her the way. " Potter's work was inspired by The Tale of Little Black Sambo and her two nieces, Petunia and Petunia. It was the first book in the Potter family to be published in a hardback format, and was also the first book in the Peter Rabbit series. The tale was released on 13 November 1910 and was accompanied by a greeting card by Potter's brother-in-law, Frederick Warne. The tale was well received by critics and children, and sold over 600,000 copies in its first few years. In the first year of its publication, Potter's The Tale of Mrs. Tittlemouse won children's book awards in England, America and Australia. It was selected as a " recommended reading " by children in the US and Canada.
= = Plot summary = =
The tale opens with the pet rabbit Mrs. Tittlemouse wandering through a forest in a small litter of four. He is shot and badly injured by a deer after she attempts to escape. Mrs. Tittlemouse decides that she wants to find her own home, because she is eager to go on her own. She goes alone to the farm where she makes a little money by selling a few seeds and building a small cabin in the woods. She is approached by a wealthy hunter named Mr. Tittlemouse, who tries to kill her but Mrs. Tittlemouse kills him by stuffing a rope into his nose and killing him. She is rescued by Mr. Tittlemouse's wife Ruth, but Mrs. Tittlemouse then leaves the woodland with the baby. When she is spotted by
Liberman introduces Transformer-XL with these words:
This is another note on the amazing ability of modern AI learning techniques to imitate some aspects of natural-language patterning almost perfectly, while managing to miss common sense almost entirely. This probably tells us something about modern AI and also about language, though we probably won't understand what it's telling us until many years in the future.
Yes.

Readers of New Savanna will know that I've been thinking something like that for a couple of years now, and have even attempted to conceptualize it – a number of posts I've labeled with "AI Limit" are of this kind, particularly Computational linguistics & NLP and Borges Redux: Computing Babel. In his speech upon accepting an award from the ACL Martin Kay (PDF) notes that contemporary AI is using statistics over word distributions as a substitute/proxy for a model of the world. I think that's right. I note as well that the problem of common sense knowledge is one of the problems that put the brakes on old-style symbolic AI. That, of course, is a problem about modeling the (surface of) the world in all it's trivial but inescapable variety. There were just so many bits of it to hand-code and, once coded, all those trivial bits exacerbated the problem of combinatorial explosion.

It is thus interesting that these new techniques, run on machines that dwarf those machines from the 1970s and 1980s, are now running into the common sense problem. It is not at all obvious to me that the problem can be solved by using ever more text as fodder and more computing power to digest that fodder. The world is just too big and too irreducibly complex to be mastered in that way.

The other side of the issue is that these statistical techniques work very well in closed domains, like chess and Go. In those domains there is hardly any world to speak of and hence there is no common sense problem. Moreover, abstractly considered, those games are finite. Given enough time and memory it would be possible to calculate every possible game and then list them all. What's interesting is that the best chess programs seem to have broken into regions of the chess space that human players had not explored, so they exhibit new styles of play.

It's as though Go and chess embody the abstract mental powers we bring to bear on the world (Chomskyian generativity? Cartesian rationality?) while the common sense problem, in effect, represents the resistance that the world presents to us. It is the world exerting its existence by daring us: "parse this, and this, and this, and...!"

Tuesday, November 5, 2019

A short post in my ongoing series about why I find most anti-AI arguments to be conceptually empty

My basic complaint about such arguments (by, e.g., Dreyfus, Searle, see my most recent post on this topic) is that they don’t engage with the actual techniques used by computer systems nor, for that matter, do they engage with what the various psychologies have to say about the mind/brain. So, if I am an AI researcher, these arguments don’t tell me anything I can use to improve the systems I design. And if I research the human mind/brain, they don’t tell me anything about that either.

The arguments are mostly rhetorical in aim and force, not conceptual. And they get their rhetorical force from the fact that current systems are far from being intelligent, whatever that is, in the way that humans are.

Thus an argument couched in theological terms could potentially have the same rhetorical force though its conceptual equipment would be quite different. One might argue, for example, that human actions are manifestations of divine design while the actions of computers are those of mere uninspired machines. For all I know, someone somewhere is offering such arguments. But that kind of argument would have to be directed to an audience for whom such arguments have conceptual currency. That’s not the case with that audience Dreyfus and Searle are addressing, with Dennett and others on the other side of the argument. In either case it is true that no existing computer system provides an inescapable counter example

Both arguments depend on the relative incapacity of current systems and neither engages with actual work on computer systems or on the mind. They are BOTH conceptually empty though rhetorically strong. Thus, for example, neither of these accounts have anything useful to say about the problem of commonsense knowledge, which certainly exists for AI of whatever kind and which hasn’t really been addressed in psychology at all, though a great deal of psychology is relevant to the problem.

More later.

Annals of machine intelligence: Play-by-play commentary about cat-on-field [the problem of common sense knowledge]


I'd like to hear a computer deliver commentary like that. What's required? It's got to make sense of the visual scene in real-time. And it's got to be able to focus on the cat while noticing the actions of the people as well. It must also know of course that the people are responding to and acting in relationship to what the cat is doing. While this perceptual and cognitive activity is going on the system must improvise appropriate commentary and then utter it, with appropriate intonation.

As someone with no more than a casual interest in football, I can tell you that I simply do not see what's happening on the field with anything like the detail and comprehension of play-by-play announcers. To be sure, they have the advantage of being there in the stands watching while I'm only watching on the TV. But still, they've watched many more football games than I have, and have analyzed those games, and so know how to follow the action.

Let us, however, for the sake of argument, imagine that we've got a computer system specialized for delivering football play-by-play commentary. In this case Harlan, the announcer, isn't calling a football play. Would a game calling computer program even know about the existence of cats? If not then it wouldn't be able to describe what's going on. Would such a system know about state troopers, who and what they are, why they'd be at a game, and what they're doing on the field? Routine play-by-play doesn't call for such knowledge. Nor does it call for knowledge of the tunnel leading through the stands to and from the field itself. No, even if we had a very good play-by-play system, it is not likely equipped to handle cat-on-the-field, not unless it also has knowledge of things that are, for the most part, peripheral to the game itself.

Umm, err, maybe it would be able to improvise something about an unidentified prowling animate object?

You think so, eh? What would THAT require? How'd it come up with that weird generalization,  "unidentified prowling animate object"?

Imagine that you're the developer of a football play-by-play system. You realize that it may well be called on to provide chatter about non-football things, like cat-on-the-field, of maybe just about the weather. What non-football information and knowledge are you going to equip your system with? Non-football knowledge, that's a vast and unbounded category.

Setting aside the sensory-motor aspect of the problem, this is what we mean by the problem of commonsense knowledge, all those little bits of routine knowledge we have about the world. It's one of the problems that contributed to the implosion of symbolic AI back in the mid-1980s. Marvelous though current state-of-the-art AI systems can be, they've got problems with commonsense knowledge as well, and at least some commentators recognize it.

Wednesday, September 4, 2019

AI scores an 'A' on the N.Y. Regents Science Exams

Peter Clark, et al., From 'F' to 'A' on the N.Y. Regents Science Exams: An Overview of the Aristo Project, arXiv:1909.01958 [cs.CL]
Abstract: AI has achieved remarkable mastery over games such as Chess, Go, and Poker, and even Jeopardy, but the rich variety of standardized exams has remained a landmark challenge. Even in 2016, the best AI system achieved merely 59.3% on an 8th Grade science exam challenge.

This paper reports unprecedented success on the Grade 8 New York Regents Science Exam, where for the first time a system scores more than 90% on the exam's non-diagram, multiple choice (NDMC) questions. In addition, our Aristo system, building upon the success of recent language models, exceeded 83% on the corresponding Grade 12 Science Exam NDMC questions. The results, on unseen test questions, are robust across different test years and different variations of this kind of test. They demonstrate that modern NLP methods can result in mastery on this task. While not a full solution to general question-answering (the questions are multiple choice, and the domain is restricted to 8th Grade science), it represents a significant milestone for the field.
From the introduction:
Instead of a binary pass/fail, machine intelligence is more appropriately viewed as a diverse collection of capabilities associated with intelligent behavior. Finding appropriate benchmarks to test such capabilities is challenging; ideally,a benchmark should test a variety of capabilities in a natural and unconstrained way, while additionally being clearly measurable, understandable, accessible, and motivating.

Standardized tests, in particular science exams, are a rare example of a challenge that meets these requirements.While not a full test of machine intelligence, they do explore several capabilities strongly associated with intelligence, including language understanding, reasoning, and use of common-sense knowledge. One of the most interesting and appealing aspects of science exams is their graduated and multifaceted nature; different questions explore different types of knowledge, varying substantially in difficulty. For this reason, they have been used as a compelling—and challenging—task for the field for many years (Brachmanet al., 2005; Clark and Etzioni, 2016).
Here's a story about this in the NYTimes.

Friday, February 22, 2019

Structural inequality in the Wikimedia universe

Jinhyuk Yun, Sang Hoon Lee & Hawoong Jeong, Early onset of structural inequality in the formation of collaborative knowledge in all Wikimedia projects, Nature Human Behaviour, volume 3, pages 155–163 (2019)
Abstract: The Wikimedia project, including Wikipedia, is one of the largest communal data sets and has served as a representative medium to convey collective knowledge in the twenty-first century. Researchers have believed that the analysis of these collaborative digital data sets provides a unique window into the processes of collaborative knowledge formation; yet, in reality, most previous studies have usually focused on its narrow subsets. Here, by analysing all 863 Wikimedia projects (various types and in different languages), we find evidence for a universal growth pattern in communal data formation. We observe that inequality arises early in the development of Wikimedia projects and stabilizes at high levels. To understand the mechanism behind the observed structural inequality, we develop an agent-based model that considers the characteristics of the editors and successfully reproduces the empirical results. Our findings from the Wikimedia projects data, along with other types of collaboration data, such as patents and academic papers, show that a small number of editors have a disproportionately large influence on the formation of collective knowledge. This analysis offers insights into how various collaboration environments can be sustained in the future.

Tuesday, November 6, 2018

State of the art AI systems are vulnerable to adversarial examples

Yes, AI has come a long way, and it still has a long way to go. In particular, it fails to comprehend the meaning of the phenomena it deals with, says Melanie Mitchell in a current NYTimes op-ed, Artificial Intelligence Hits the Barrier of Meaning.
Even more worrisome are recent demonstrations of the vulnerability of A.I. systems to so-called adversarial examples. In these, a malevolent hacker can make specific changes to images, sound waves or text documents that while imperceptible or irrelevant to humans will cause a program to make potentially catastrophic errors.

The possibility of such attacks has been demonstrated in nearly every application domain of A.I., including computer vision, medical image processing, speech recognition and language processing. Numerous studies have demonstrated the ease with which hackers could, in principle, fool face- and object-recognition systems with specific minuscule changes to images, put inconspicuous stickers on a stop sign to make a self-driving car’s vision system mistake it for a yield sign or modify an audio signal so that it sounds like background music to a human but instructs a Siri or Alexa system to perform a silent command.

These potential vulnerabilities illustrate the ways in which current progress in A.I. is stymied by the barrier of meaning. Anyone who works with A.I. systems knows that behind the facade of humanlike visual abilities, linguistic fluency and game-playing prowess, these programs do not — in any humanlike way — understand the inputs they process or the outputs they produce. The lack of such understanding renders these programs susceptible to unexpected errors and undetectable attacks.

What would be required to surmount this barrier, to give machines the ability to more deeply understand the situations they face, rather than have them rely on shallow features? To find the answer, we need to look to the study of human cognition.

Our own understanding of the situations we encounter is grounded in broad, intuitive “common-sense knowledge” about how the world works, and about the goals, motivations and likely behavior of other living creatures, particularly other humans. Additionally, our understanding of the world relies on our core abilities to generalize what we know, to form abstract concepts, and to make analogies — in short, to flexibly adapt our concepts to new situations. Researchers have been experimenting for decades with methods for imbuing A.I. systems with intuitive common sense and robust humanlike generalization abilities, but there has been little progress in this very difficult endeavor.

A.I. programs that lack common sense and other key aspects of human understanding are increasingly being deployed for real-world applications. While some people are worried about “superintelligent” A.I., the most dangerous aspect of A.I. systems is that we will trust them too much and give them too much autonomy while not being fully aware of their limitations.

Saturday, September 8, 2018

How stupid is machine translation? No common sense at all, none


So I tweeted a clip of an interview and performance by the Japanese pianist, Hiromi Uehara. The interview is in Japanese, a language I don't know, and the clip is labeled in Japanese. But I'd listened to the clip and figured out what was going on. She was asked to play a song, "My Way", a grandiloquent turkey which Frank Sinatra had made into a hit. She gigled, complied, and then, as any jazz musician would, proceeded to improvise.

So, when I posted the clip I added a short English preface to the Japanese, "Hiromi does it her way". Out of curiosity, when I saw my tweet, I asked for the translation Twitter offered. Not knowing any better, the computer also translated my bit of prefatory English, despite the fact that it was already in English, into "Hiromi do it her way". Why, I ask, why? Because it doesn't know any better.

The DOD is going to invest in "explainable AI"

And that requires common sense, among other things.

Right now, for example, if a soldier asks an AI system like a target identification platform to explain its selection, it can only provide the confidence estimate for its decision, DARPA’s director Steven Walker told reporters after a speech announcing the new investment – an estimate often given in percentage terms, as in the fractional likelihood that an object the system has singled out is actually what the operator was looking for.

“What we’re trying to do with explainable AI is have the machine tell the human ‘here’s the answer, and here’s why I think this is the right answer’ and explain to the human being how it got to that answer,” Walker said.

DARPA officials have been opaque about exactly how its newly-financed research will result in computers being able to explain key decisions to humans on the battlefield, amidst all the clamor and urgency of a conflict, but the officials said that being able to do so is critical to AI’s future in the military.

Vaulting over that hurdle, by explaining AI reasoning to operators in real time, could be a major challenge. Human decision-making and rationality depend on a lot more than just following rules, which machines are good at. It takes years for humans to build a moral compass and commonsense thinking abilities, characteristics that technologists are still struggling to design into digital machines.

“We probably need some gigantic Manhattan Project to create an AI system that has the competence of a three year old,” Ron Brachman, who spent three years managing DARPA’s AI programs ending in 2005, said earlier during the DARPA conference. “We’ve had expert systems in the past, we’ve had very robust robotic systems to a degree, we know how to recognize images in giant databases of photographs, but the aggregate, including what people have called commonsense from time to time, it’s still quite elusive in the field.”

Thursday, April 12, 2018

Common Sense in Artificial Intelligence

Sometime back in the 1970s, I believe it was, David Marr observed something of a paradox (I believed he used that word) in the development of artificial intelligence (AI). Much of the early work, which did meet with some success, involved modeling fairly sophisticated forms of knowledge, mathematics and science, but when researchers started working in simple domain, like ordinary narrative, things got more difficult. That is, it seemed easier to model the specialized knowledge of a highly trained scientist than the general knowledge of a six year old. That problem has come to be known in AI as the problem of common sense, and its intractability has was one reason that old school research programs grounded in symbolic reasoning fell apart in the mid-1980s. During the 1990s and continuing on to the present various machine learning techniques have become quite successful in domains that had eluded symbolic AI. But common sense reasoning has continued to elude researchers.

Earlier this year Microsoft co-founder Paul Allen announced that he was giving $125 million to his nonprofit Allen Institute for Artificial Intelligence (AI2) to study common sense reasoning. Here's a short paper from that lab that gives and overview of the problem.
Niket Tandon, Aparna S. Varde, Gerard de Melo, Commonsense Knowledge in Machine Intelligence, SIGMOD Records 2018.

Abstract: There is growing conviction that the future of computing depends on our ability to exploit big data on the Web to enhance intelligent systems. This includes encyclopedic knowledge for factual details, common sense for human-like reasoning and natural language generation for smarter communication. With recent chatbots conceivably at the verge of passing the Turing Test, there are calls for more common sense oriented alternatives, e.g., the Winograd Schema Challenge. The Aristo QA system demonstrates the lack of common sense in cur- rent systems in answering fourth-grade science exam questions. On the language generation front, despite the progress in deep learning, current models are easily confused by subtle distinctions that may require linguistic common sense, e.g. quick food vs. fast food. These issues bear on tasks such as machine translation and should be addressed using common sense acquired from text. Mining common sense from massive amounts of data and applying it in intelligent systems, in several respects, appears to be the next frontier in computing. Our brief overview of the state of Commonsense Knowledge (CSK) in Machine Intelligence provides insights into CSK acquisition, CSK in natural language, applications of CSK and discussion of open issues. This paper provides a report of a tutorial at a recent conference with a brief survey of topics.

Friday, October 6, 2017

Computation, games, humor, and common sense

Mark Seidenberg has an interesting post today, Cartoonist walks into a language lab… He poses a task for machine learning:
New Yorker cartoons are usually captioned these days, with fewer in the lovely mute style of a William Steig.  A general theory of language use should be able to explain how cartoon captions, a genre of text, are understood. The cartoons illustrate (sic) the dependence of language comprehension on context (the one created by the drawing) and background knowledge (about, for example, rats running mazes, guys marooned on islands, St. Peter’s gate, corporate culture, New Yorkers). The popular Caption Contest is an image-labeling task, generating humorous labels for an incongruous scene.
A bit later:
The weekly caption contest has yielded a massive amount of data that is being analyzed using NLP and machine learning techniques.  (The contest: Entrants submit captions for a cartoon; from the 5000 or so entries the editors pick three finalists; readers pick the winner by voting on-line.)  Just think of the studies that can be done with this goldmine of a data set!  Identify the linguistic properties that distinguish the winning captions from the two losers. Build a classifier that can estimate relative funniness from properties such as word choices, grammatical complexity, affective valence (“sentiment”), readability, structure of the joke, etc.  Use the classifier to predict the winners on other weeks. Or the rated humorosity of other cartoons.

Heavy hitters from places like Microsoft, Google, Michigan, Columbia, Yale, and Yahoo have taken swings at this ... The results (from the few published studies I found) have been uninspiring.
Liberman then asks a crucial question: “Is labeling New Yorker cartoons harder than playing Go?” – you may recall that not too long ago Google’s DeepMind made a media splash when one of its programs defeated the best human expert on Go. Liberman’s answer to his question is, of course, yes, labeling cartoons is harder.

He explains:
Go has a conventionalized set-up and explicit rules. A captioning model has to figure out what game is being played.  Captioning is a type of scene labeling but that requires recognizing what’s in the scene which in this case is, literally, ridiculous: exaggerated, crude, eccentric, stylized renderings of the world. Quite different from the naturalistic scenes that have been the focus of so much attention in AI. 
He goes on to observe:
... humor turns on a vast amount of background knowledge. That feeling when you just don’t get it happens when we either don’t know the relevant stuff or can’t figure out what’s relevant to that cartoon.  A deep learning network might well acquire the requisite knowledge of the world but not from 80,000 drawings: insufficient data.  Same for analyzing the captions: it’s necessary to know the language.  People have acquired most of the knowledge that's required by other means.
That, common sense background knowledge, is one of the problems that squelched classic symbolic processing several decades ago. Researchers realized that in dealing with even very simple language, we draw on a vast back of common sense knowledge. How are we doing to hand-code all that into a computer? And once we've done it, what of combinatorial explosion?

Saturday, July 29, 2017

Gary Marcus is skeptical of AI, however...we need a new funding paradigm

Artificial Intelligence is colossally hyped these days, but the dirty little secret is that it still has a long, long way to go. Sure, A.I. systems have mastered an array of games, from chess and Go to “Jeopardy” and poker, but the technology continues to struggle in the real world. Robots fall over while opening doors, prototype driverless cars frequently need human intervention, and nobody has yet designed a machine that can read reliably at the level of a sixth grader, let alone a college student. Computers that can educate themselves — a mark of true intelligence — remain a dream.

Even the trendy technique of “deep learning,” which uses artificial neural networks to discern complex statistical correlations in huge amounts of data, often comes up short. Some of the best image-recognition systems, for example, can successfully distinguish dog breeds, yet remain capable of major blunders, like mistaking a simple pattern of yellow and black stripes for a school bus. Such systems can neither comprehend what is going on in complex visual scenes (“Who is chasing whom and why?”) nor follow simple instructions (“Read this story and summarize what it means”).
What we need:
To get computers to think like humans, we need a new A.I. paradigm, one that places “top down” and “bottom up” knowledge on equal footing. Bottom-up knowledge is the kind of raw information we get directly from our senses, like patterns of light falling on our retina. Top-down knowledge comprises cognitive models of the world and how it works.
Marcus goes on to argue that the current methods of funding AI–small academic labs and somewhat larger industrial labs–are inadequate. Academic labs are too small to gather the wide variety of talent needed to make significant progress. Industrial labs are too focused on short-term results.
I look with envy at my peers in high-energy physics, and in particular at CERN, the European Organization for Nuclear Research, a huge, international collaboration, with thousands of scientists and billions of dollars of funding. [...]

An international A.I. mission focused on teaching machines to read could genuinely change the world for the better — the more so if it made A.I. a public good, rather than the property of a privileged few.