Showing posts with label predict. Show all posts
Showing posts with label predict. Show all posts

Thursday, August 14, 2025

Nate Silver on AI and AI as superforcasters

Nate Silver on Life’s Mixed Strategies, Conversations with Tyler, August 13, 2025.

In his third appearance on Conversations with Tyler, Nate Silver looks back at past predictions, weighs how academic ideas such as expected utility theory fare in practice, and examines the world of sports through the lens of risk and prediction.

Tyler and Nate dive into expected utility theory and random Nash equilibria in poker, whether Silver’s tell-reading abilities transfer to real-world situations like NBA games, why academic writing has disappointed him, his move from atheism to agnosticism, the meta-rationality of risk-taking, electoral systems and their flaws, 2028 presidential candidates, why he thinks superforecasters will continue to outperform AI for the next decade, why more athletes haven’t come out as gay, redesigning the NBA, what mentors he needs now, the cultural and psychological peculiarities of Bay area intellectual communities, why Canada can’t win a Stanley Cup, the politics of immigration in Europe and America, what he’ll work on next, and more.

Here's most of the discussion of AI:

COWEN: Now, speaking of predictions, a year ago, we talked about how long will it take AIs to be as good as human superforecasters? You made a prediction, where you said at least 10 to 15 years. Now, a year later, do you want to revisit that and revise?

SILVER: I think it’s probably about right. I would say, relative to a year ago, AI is about at the 40th pe rcentile of progress I would’ve expected. I’d be curious what you would think.

COWEN: A year ago, I said two to three years. Right now, I’m going to say one to two years, which is the same prediction. I think you’re way too pessimistic in your timetable.

SILVER: It depends on how competitive the exercise is. If it’s like a —

COWEN: Like a Math Olympiad tournament. They just did gold medal performance. I said this last year. I said, “In a year, they’re going to do gold medal.” A year ago, they weren’t sure how many r’s were in the word strawberry. You don’t think on superforecasting, they can —

SILVER: I think it’s very different when you’re dealing with a static problem as compared to a dynamic system where the inputs are changing all the time. Currently, the large language models are very, very bad at poker. They’re not trained on poker data. I’m sure if you did train them . . .

There are these things called solvers that are trained on poker data that do very well, but they cannot quite impute the general patterns just from mediocre text data or an amateurish kind of hand analysis. If you probe them on why they’re bad, they’re like, “Yes, maybe it’s tough for us when you have a complicated, evolving game theory dynamic, and you have to develop exploitative strategies very quickly.”

If you have a computer solve a poker hand to get to one where there’s enough loss minimization, a very strong computer can take minutes, whereas a poker player makes those calculations implicitly in a handful of seconds, for example. I don’t know. I worry, with the Math Olympiad stuff — there’s a little bit of teaching to the test where, because you set this as the goal that a large language model should have, that therefore there’s a lot of prestige when you meet that goal, potentially —

COWEN: Isn’t teaching to the test what we should do, even with humans in a sense. The test is what you think is important, and that’s what you ought to teach.

SILVER: Well, but the poker example, or chess — I think AI models are very poor at chess, from everything that I’ve heard, for example.

COWEN: Other AI models play chess great.

SILVER: Correct. This gets me to the question of, what’s it mean to be generally intelligent? We’ll probably have scaffolding of model on top of model, and you’ll now patch different things. Right now, you can’t really very effectively make a plane reservation using ChatGPT, but I’m sure if you dedicate a resource to that, then you have these agentic models now, or agent models are just creeping into the system, a little bit.

COWEN: Those will work in less than a year, I think. There’s an agentic model now from OpenAI.

SILVER: This is where I get to rough AGI versus superintelligence. I am less convinced that we’re going to have some intelligence explosion than I would have been maybe . . . I don’t think I was ever convinced of it, but this emergent superintelligence, where you train it on relatively simple data and it extrapolates beyond the data set. I think they do reason. Sometimes I think they’re quite smart, and I no longer am bashful about saying, “Oh, ChatGPT thinks this.” I used to avoid that term, think.

But there’s a big gap between approximate general intelligence for desk jobs, and then superintelligence on the one hand, or AGI for physical labor on the other hand. I think people are much too quick to make that leap. I think the Math Olympiad, in part because maybe the answers are somewhere latent in the training data, but even if they’re not, if you try to solve, what’s a Lucas critique? Whatever else, right? There’s a version of that, I think, for AI models.

COWEN: My intuition is that if you took five superforecasters and just had them write a five-page prompt for GPT-5, which will be out this summer, that we’d be there already. I don’t think it would be superintelligence. You could say it’s not AGI. But the human superforecasters, they’re not that impressive. They’re not Einsteins. They just have good methods, and they’re disciplined.

There's much more in the discussion.

Thursday, April 17, 2025

AGI, really? Does it really matter? [Tyler Cowan again + 5 predictions from Rodney Brooks]

Cowen has just run up a short post over at Marginal Revolution, A note on o3 and AGI. Here’s what he says:

Basically it wipes the floor with the humans, pretty much across the board. [...] I don’t mind if you don’t want to call it AGI. And no it doesn’t get everything right, and there are some ways to trick it, typically with quite simple (for humans) questions. But let’s not fool ourselves about what is going on here. On a vast array of topics and methods, it wipes the floor with the humans. It is time to just fess up and admit that.

I felt I had no choice but to make a longish reply, which follows immediately. I then add some further thoughts.

My reply to Tyler Cowen on o3 & AGI

Hmmmm... I have at various times and places, including in the comment section here [at Marginal Revolution], expressed the view that we’ll understand how LLMs work before we reach AGI. If I take Tyler’s assertions at face value, then I’d have to admit I’m wrong on that. Because we certainly do not understand how LLMs work. We don’t know any more about that today than we did yesterday or a week ago. Nor are “we” even trying very hard to figure it out. Oh, sure, the folks at Anthropic are spending a great deal of time on that problem. I’m sure others are working on it as well. But if OpenAI is, they’re not tell us or giving us any results of their work. Why not?

Anyhow, I don’t feel as though the “spirit” of my view has been falsified by o3. I’m willing to believe Tyler when he says its performance is spectacular, even when I apply the fanboy discount to his assertion. What these various LLM-based chatbots and reasoning-bots can do really IS spectacular.

I note that Tyler has said he “mind if you don’t want to call it AGI.” I certainly don’t care about that either.

My basic intellectual commitment in all of this, however, is to the question: How does it work? For me that question is primarily one about the human mind-brain. That’s what I want to understand. If I’ve spent a great deal of time (over the course of five decades) dealing with computational models of intelligence, it’s because I’m interested in how the mind works. And if, in the course of trying to figure that out, we manage to produce computer systems that have practical benefits, mazel tov! What’s not to like?

Now, it so happens that Rodney Brooks just coughed up five dated predictions for the next decade. The last two seem most relevant here:

4. Neural computation. There will be small and impactful academic forays into neuralish systems that are well beyond the linear threshold systems, developed by 1960, that are the foundation of recent successes. Clear winners will not yet emerge by 2036 but there will be multiple candidates.

5. LLMs that can explain which data led to what outputs will be key to non annoying/dangerous/stupid deployments. They will be surrounded by lots of mechanism to keep them boxed in, and those mechanisms, not yet invented for most applications, will be where the arms races occur.

If we’re going to achieve #5, it seems to me that we’re going to have to know how LLMs work. As for #4, I assume that Brooks is talking about systems with new non-LLM architectures. That’s fine. We need such systems. I figure that unraveling the inner workings of LLMs will contribute to work on such systems, and vice versa.

Question: There’s a 2023 agreement between OpenAI and Microsoft that sets a $100 billion profit threshold on AGI. When OpenAI produces a system that crosses that threshold, that system will be declared to be AGI. How long before that threshold is reached?

Further thoughts: What about interaction with the physical world?

One David Khoo replied: “Also, don’t forget Moravec’s Paradox. Reasoning is easy, sensorimotor is hard.” Yes. And here’s how RAD replied to Khoo: “Embodied AGI is a separate problem from the type of sapience required to perform knowledge work using digital tools.” Yes. They ARE different kinds of problems. But why?

That strikes me as being a deep observation. But what’s the explanation? As far as I know Miriam Yevick is the only one who’s thought about that, and she didn’t think about the issue in those terms. Here’s a post where I address Yevick’s insight: What Miriam Yevick Saw: The Nature of Intelligence and the Prospects for A.I., A Dialog with Claude 3.5 Sonnet. That post links to a PDF containing the full debate, which you can download in three places: Academic.edu, at Social Science Research Network (SSRN), or at ResearchGate.

I mentioned Rodney Brooks’ latest prediction. Here’s the third:

3. Humanoid Robots. Deployable dexterity will remain pathetic compared to human hands beyond 2036. Without new types of mechanical systems walking humanoids will remain too unsafe to be in close proximity to real humans.

That’s manipulation of the physical world. We know it’s a difficult problem, but why? Yevick was thinking about perception. Can her insight be transformed into one that applies to physical action?

Here’s Brook’s second prediction:

2. Self driving cars. In the US the players that will determine whether self driving cars are successful or abandoned are #1 Waymo (Google) and #2 Zoox (Amazon). No one else matters. The key metric will be human intervention rate as that will determine profitability.

That too is about interacting with the physical world. But at a different scale from manual dexterity. Are the same fundamental abilities operative at both these scales, manual dexterity and medium- and large-scale movement through the world? Certainly in the case of humans we have different effectors and different senses involved. Manual dexterity involves both hapsis and kinesis as well as vision while large-scale movement is primarily guided by vision, though hearing does come into play as well. We have similar perceptual systems in the machine world. But are the underlying computational principles different?

Here’s Brooks’ first prediction:

1. Quantum computers. The successful ones will emulate physical systems directly for specialized classes of problems rather than translating conventional general computation into quantum hardware. Think of them as 21st century analog computers. Impact will be on materials and physics computations.

Again, we’re dealing with the physical world.

I’m tempted to offer a final “prediction” of my own:

We won’t have a machine that thinks profound thoughts, that’s capable of profound discoveries, until that same machine is comfortable dealing with the physical world.

As for why I think that, I can’t quite tell you. But I do know that the deepest scientific and mathematical thinkers often rely on physical intuition, on visual thinking. I discuss this in my 1990 encyclopedia article, Visual Thinking.

Tuesday, January 7, 2025

Beyond prediction: Yes, we need to work on societal resilience, but we need to work on intellectual resilience as well.

Just around the corner from here Malcolm Murray has a nice article in 3 Quarks Daily, O3 and the Death of Prediction. Here’s some excerpts from his article:

A lot of the focus over the past years has been to try to pinpoint when AI will be able to do certain tasks. Various surveys have been run estimating when AI will be able to write a best-selling novel or win math competitions. [...] The real game is to prepare for advanced AI capabilities, whether they are called AGI or not. o3 shows two things clearly – that AI evolution seems set to continue apace, and that we can not predict it and should not attempt to.

Note that the death of prediction in the AI space does not mean the death of forecasting. Forecasting will still have its place. Whereas prediction is the non-scientific, crystal ball, finger-in-the-air activity beloved by media pundits, forecasting is a more scientific endeavor, which will still be valuable. Forecasting, especially in the form invented by Philip Tetlock – Superforecasting (disclosure: I am a Superforecaster) – means careful thinking regarding the applicability of historical base rates and adjusting them based on clear current trends. This can yield still very accurate forecasts of future events, at least a few years out.

But the fields of AI safety and AI risk management should turn their focus to resilience. Normally in risk management, risks are analyzed by their potential impact as well as their probability and their likely time to materialize. Focusing on resilience, however, means putting aside the probability and time frame and focusing on the impact. This is a different mindset. It is saying that we don’t know if or when this risk will arise, but if the impact of the risk is large enough, we should make adequate preparations regardless. Preparation takes time, so it is high time to start.

I wrote a rather long reply:

I agree with you, we can’t predict, and we should certainly give more attention to societal resilience, much more attention. But I’d like to say a word in defense of Marcus and co. Because we need more than societal resilience. We also need intellectual and technological resilience. Marcus is certainly calling for that.

Marcus has studied human cognition and language. Many (of us) skeptics have. They know something about how the mind works. These deep learning folks, not so much as far as I can tell. They don’t even know how their very clever devices work – and, if you’ve been reading 3QD, you know I’m a fan of those clever devices. I use them all the time myself.

It’s like a whaling voyage captained by someone who knows everything there is to know about their ship and seamanship, but little to nothing about whales. Just because they can make their ship do fancy and unexpected things doesn’t mean that sooner or later they’re going to find a whale. Nor does it mean you should discount those who actually know something about whales but may not be so enamored of fancy ships.

Predicting technology developments is very difficult, as you point out. No one can do it. Predicting in the AI space, broadly considered, has been going on for a long time. And it’s failed before. Let's take a quick and crude look at that history.

Back in the early days of computing the federal government spent a lot of money for research on computer systems that could translate natural language. They were specifically interested in translating technical documents from Russian to English. This was the 1950s and 60s and the Cold War was in high gear. So researchers made promises, and promises, and promises, and by the early 1960s those promises were looking pretty thin. So a commission was appointed to study the problem. They arrived at two conclusions: 1) There is no immediate prospect of high-quality machine translation. 2) We now have theories and models we didn’t have when this work started. More theoretical research looks promising. As you can imagine, the government paid attention to the first conclusion and ignored the second. The field of machine translation was dead for lack of funds. But not completely dead. It rebranded itself as computational linguistics and continued on.

My teacher, David Hays, was one of those first generation researchers. He headed the program at the RAND corporation. He was on the commission that made those recommendations. And he’s the one who coined the term “computational linguistics,” which is how the field rebranded itself.

In the mid-1980s pretty much the same thing happened with AI. We had the so-called AI Winter. The work was going well, commercial ventures were started. But the desired/expected results weren’t forth-coming.

I figure that the so-called “design space” for this technology is huge. The early researchers in machine translation explored one region of the space, and ultimately failed. Call it Region Alpha. A bit later other researchers explored another region of the space, and they too failed. Call it Region Beta. The unexpected success of AlexNet in 2012 opened up a whole new region of the design space, one made available by the use of GPUs. Call this Region Gamma. The unexpected success of GPT-3 showed us new areas within this new Gamma region. The same with o3.

I think what Marcus and others are saying is that AGI, whatever it is, isn’t going to be found in Region Gamma. It’s in some other region. Why are they saying that? Because they know, we know, something about language and cognition and don’t believe it is to be found in Region Gamma. And while you're at it, read a paper Miriam Yevick published back in 1975; she knew something back then that I don't think even Marcus knows about. Will the grail of AGI be found in the next region, Delta, or will it be Epsilon, Zeta...? Who knows.

Thursday, January 2, 2025

Rodney Brooks has his tech predictions up

As you may know, roboticist Rodney Brooks has been publishing and systematically updating tech predictions every January 1st since 2018. He makes predictions in the following categories: (1) self driving cars, (2) robotics, AI , and machine learning, and (3) human space travel. Here’s the link for 2025. I’ve got some Rodney Brooks stuff here on the Savanna, including, but not limited to, excerpts from earlier prediction.

In general

In these excerpts he talks about where we are and where we aren’t in a general way (coloring in the original):

I want to be clear, as there has been for almost seventy years now, there has been significant progress in Artificial Intelligence over the last decade. There are new tools and they are being applied widely in science and technology, and are changing the way we think about ourselves, and how to make further progress.

That being said, we are not on the verge of replacing and eliminating humans in either white collar jobs or blue collar jobs. Their tasks may shift in both styles of jobs, but the jobs are not going away. We are not on the verge of a revolution in medicine and the role of human doctors. We are not on the verge of the elimination of coding as a job. We are not on the verge of replacing humans with humanoid robots to do jobs that involve physical interactions in the world. We are not on the verge of replacing human automobile and truck drivers world wide. We are not on the verge of replacing scientists with AI programs.

Breathless predictions such as these have happened for seven decades in a row, and each time people have thought the end is in sight and that it is all over for humans, that we have figured out the secrets intelligence and it will all just scale. The only difference this time is that these expectations have leaked out into the world at large. [...]

Today I get asked about humanoid robots taking away people’s jobs. In March 2023 I was at a cocktail party and there was a humanoid robot behind the bar making jokes with people and shakily (in a bad way) mixing drinks. A waiter was standing about 20 feet away silently staring at the robot with mouth hanging open. I went over and told her it was tele-operated. “Thank God” she said. (And I didn’t need to explain what “tele-operated” meant). Humanoids are not going to be taking away jobs anytime soon (and by that I mean not for decades).

You, you people!, are all making fundamental errors in understanding the technologies and where their boundaries lie. Many of them will be useful technologies but their imagined capabilities are just not going to come about in the time frames the majority of the technology and prognosticator class, deeply driven by FOBAWTPALSL, think.

But this time it is different you say. This time it is really going to happen. You just don’t understand how powerful AI is now, you say. All the early predictions were clearly wrong and premature as the AI programs were clearly not as good as now and we had much less computation back then. This time it is all different and it is for sure now.

Humanoid robots

I think we are a long way off from being able to for-real deploy humanoid robots which have even minimal performance to be useable and even further off from ones that have enough ROI for people want to use them for anything beyond marketing the forward thinking outlook of the buyer.

Despite this, many people have predicted that the cost of humanoid robots will drop exponentially as their numbers grow, and so they will get dirt cheap. I have seen people refer to the cost of integrated circuits having dropped so much over the last few decades as proof. Not so.

They are committing the sin of exponentialism in an obviously dumb way. As I explained above the first integrated circuits were far from working at the limits of physics of representing information. But today’s robots use mechanical components and motors that are not too far at all from physics based limits, about mass, force, and energy. You can’t just halve the size of a motor and have a robot lift the same sized payload. Perhaps you can halve it once to get rid of inefficiencies in current designs. Perhaps. But you certainly can’t do it twice. Physical robots are not ripe for exponential cost reduction by burning wastes in current designs. And it won’t happen just because we start (perhaps) mass producing humanoid robots (oh, but the way, I already did this a decade ago–see my parting shot below). We know that from a century of mass producing automobiles. They did not get exponentially cheaper, except in the computing systems. Engines still have mass and still need the same amount of energy to accelerate good old fashioned mass.

Human spaceflight

Brooks has a couple of paragraphs on the SpaceX Starship. This is the next to the last of those (my highlighting):

This is the vehicle that the CEO of SpaceX recently said would be launched to Mars and attempt a soft landing there. He also said that if successful the humans would fly to Mars on it in 2030. These are enormously ambitious goals just from a maturity of technology standpoint. The real show stopper however may be human physiology as evidence accumulates that humans would not survive three years (the minimum duration of a Mars mission, due to orbital mechanics) in space with current shielding practices and current lack of gravity on board designs. Those two challenges may take decades, or even centuries to overcome (recall that Leonardo Da Vinci had designs for flying machines that took centuries to be developed…).

If and when we produce AGI-level AIs and robots, perhaps they’ll take over the space travel mission, as I’ve recently suggested in a conversation with Claude 3.5. I’ve also published an excerpt from that conversation over at 3 Quarks Daily. At the moment I can imagine a future in which humans regularly spend time in near earth orbit, and perhaps we’ll have a few on the Moon, but their presence there will be more ritual than practical. Maybe robots and AIs will have a permanent presence on Mars, and perhaps other planets as well. And perhaps we’ll have Jeff Bezos’s rotating cities as well. If one of them is near Mars, people can travel back and forth between it and Mars. But that’s a long way off.

Thursday, February 8, 2024

Rodney Brooks' most recent views on LLMs [ho hum, steady as she goes]

Rodney Brooks has published his most recent set of tech predictions: Predictions Scorecard, 2024 January 01.

He's got predictions and commentary for Self-Driving Cars, (humanoid) Robots, Artificial Intelligence and Machine learning, Human Spaceflight, and comments on electric cars, flying cars, and hyperloop.

On predicting developments in AI:

I had predicted that the “next big thing” in AI, beyond deep learning, would show up no earlier than 2023, but certainly by 2027. I also said in the table of predictions in my January 1st, 2018, that for sure someone was already working on that next big thing, and that papers were most likely already published about it. I just didn’t know what it would be; but I was quite sure that of the hundreds or thousands of AI projects that groups of people were already successfully working hard on, one would turn out to be that next big thing that everyone hopes is just around the corner. I was right about both 2023 being when it might show up, and that there were already papers about it before 2018.

Why was I successful in those predictions? Because it always happens that way and I just found the common thread in all “next big things” in AI, and their time constants.

The next big thing, Generative AI and Large Language Models started to enter the general AI consciousness last December, and indeed I talked about it a little in last year’s prediction update. I said that it was neither the savior nor the destroyer of mankind, as different camps had started to proclaim right at the end of 2022, and that both sides should calm down. I also said that perhaps the next big thing would be neuro-symbolic Artificial Intelligence.

By March of 2023, it was clear that the next big thing had arrived in AI, and that it was Large Language Models. The key innovation had been published before 2018, in 2017, in fact.

Vaswani, Ashish; Shazeer, Noam; Parmar, Niki; Uszkoreit, Jakob; Jones, Llion; Gomez, Aidan N; Kaiser, Łukasz; Polosukhin, Illia (2017). “Attention is All you Need”. Advances in Neural Information Processing Systems. Curran Associates, Inc. 30.

So I am going to claim victory on that particular prediction, with the bracketed years (OK, so I was a little lucky…) and that a major paper for the next big thing had already been published by the beginning of 2018 (OK, so I was even luckier…).

On generative AI and LLMS he points to a video of a talk he gave at MIT, and a blog post based on that talk, telling us

the talk is about what the existence of these “valuable cultural tools” (due to Alison Gopnik at UC Berkeley) tells us about deeper philosophical questions about how human intelligence works, and how they are following a well worn hype cycle that we have seen again, and again, during the 60+ year history of AI.

I concluded my talk encouraging people to do good things with LLMs but to not believe the conceit that their existence means we are on the verge of Artificial General Intelligence.

By the way, there are the initial signs that perhaps LLMs have already passed peak hype. And the ever interesting Cory Doctorow has written a piece on what will be the remnants after the LLM bubble has burst. He says there was lots of useful stuff left after the dot com bubble burst in 2000, but not much beyond the fraud in the case of the burst crypto bubble.

He tends to be pessimistic about how much will be left to harvest after the LLM bubble is gone. Meanwhile right at year’s end the lawsuits around LLM training are starting to get serious.

The concluding paragraphs of the Doctorow piece:

All the big, exciting uses for AI are either low-dollar (helping kids cheat on their homework, generating stock art for bottom-feeding publications) or high-stakes and fault-intolerant (self-driving cars, radiology, hiring, etc.).

Every bubble pops eventually. When this one goes, what will be left behind?

Well, there will be little models – Hugging Face, Llama, etc – that run on commodity hardware. The people who are learning to “prompt engineer” these “toy models” have gotten far more out of them than even their makers imagined possible. They will continue to eke out new marginal gains from these little models, possibly enough to satisfy most of those low-stakes, low-dollar ap­plications. But these little models were spun out of big models, and without stupid bubble money and/or a viable business case, those big models won’t survive the bubble and be available to make more capable little models.

There are some promising avenues, like “feder­ated learning,” that hypothetically combine a lot of commodity consumer hardware to replicate some of the features of those big, capital-intensive models from the bubble’s beneficiaries. It may be that – as with the interregnum after the dotcom bust – AI practitioners will use their all-expenses-paid education in PyTorch and TensorFlow (AI’s answer to Perl and Python) to push the limits on federated learning and small-scale AI models to new places, driven by playfulness, scientific curiosity, and a desire to solve real problems.

There will also be a lot more people who un­derstand statistical analysis at scale and how to wrangle large amounts of data. There will be a lot of people who know PyTorch and TensorFlow, too – both of these are “open source” projects, but are effectively controlled by Meta and Google, respectively. Perhaps they’ll be wrestled away from their corporate owners, forked and made more broadly applicable, after those corporate behemoths move on from their money-losing Big AI bets.

Our policymakers are putting a lot of energy into thinking about what they’ll do if the AI bubble doesn’t pop – wrangling about “AI ethics” and “AI safety.” But – as with all the previous tech bubbles – very few people are talking about what we’ll be able to salvage when the bubble is over.

Tuesday, November 14, 2023

New AI-based weather forecasting is superior to traditional methods

Dan Stillman, Why your weather forecasts may soon become more accurate, Washington Post, Nov. 14, 2023.

Google DeepMind’s AI model, named “GraphCast,” was trained on nearly 40 years of historical data and can make a 10-day forecast at six-hour intervals for locations spread around the globe in less than a minute on a computer the size of a small box. It takes a traditional model an hour or more on a supercomputer the size of a school bus to accomplish the same feat. GraphCast was about 10 percent more accurate than the European model on more than 90 percent of the weather variables evaluated.

The study’s results are similar to those in an academic article published in August to the online database arXiv.

“To be competitive with arguably the best global prediction system, if not outperforming it, is astonishing,” Aaron Hill, lead developer of Colorado State University’s machine learning prediction system, said in an email. “You can safely add GraphCast to a growing list of AI-based weather prediction models that should see continued evaluation for their application in industry, research and operational forecasting.”

AI weather models have drawn increasing attention from government weather agencies because of their speed, efficiency and potential cost savings.

Traditional weather models, such as “the European,” operated by the European Center for Medium-Range Weather Forecasts (ECMWF) in Reading, Britain, and “the American,” by the National Oceanic and Atmospheric Administration, make forecasts based on complex mathematical equations. Such models underpin forecasts and lifesaving warnings worldwide but are expensive to run because they require tremendous amounts of computing power.

AI models use a different approach. They are first trained to recognize patterns in vast amounts of historical weather data, then generate forecasts by ingesting current conditions and applying what they learned from the historical patterns. The process is much less computationally intensive and can be completed in minutes or even seconds on much smaller computers.

However:

Researchers have expressed concerns about the ability of AI to accurately forecast extreme weather, in part because there are relatively few such events to learn from in the past. Yet GraphCast reduced cyclone forecast track errors by around 10 to 15 miles at a lead time of two to four days, improved forecasts of water vapor associated with atmospheric rivers by 10 to 25 percent, and provided more precise forecasts of extreme heat and cold five to 10 days ahead of time.

And:

Most experts, including the study’s authors, agree that traditional models aren’t about to be replaced by AI models, which still depend on the older models to supply training data and to generate the current conditions they use as a starting point to make a forecast.

Here's a link to the article in Science reporting the research underlying this article.

Sunday, November 5, 2023

ChatGPT on predicting the weather [what are the limits of prediction for various phenomena?]

I’m interested in things we can predict and how we go about it. Newtonian mechanics gave us the means to predict the motion of the planets, so that’s now a solved problem – though I understand there are some nuances (chaos at the margins). Enormous effort goes into predicting the stock market. Much of that work is proprietary. In any event, I assume it’s an open problem.

What I’m really interested in is culture, such as Arthur De Vany’s work on predicting the success of a film once it’s released. He claims it’s impossible. I assume the same holds for books, records, and the like. Here we’ve got complex phenomena. Each title (movie, book, etc.) is an object of potentially unbounded complexity. The market is a population of people who have varying preferences for those complex objects. Perhaps the most interesting thing about this market place is that it has latent preferences (my coinage, though I wouldn’t be surprised if the concept, and even the term, is already in use; come to think of it, I’d be surprised if it isn’t). These preferences become visible only when an object is produced that satisfies them.

Now, one might observe of De Vany’s work that his method is crude. He trying to make predictions based on knowledge of producers, actors, directors, writers, and opening box-office. Maybe a more sophisticated set of variables, variables that tell us more about the films, would make prediction possible. Maybe, maybe some improvement is possible. But how much? What I’m wondering is whether or not these kinds of systems are computationally irreducible in the sense that Wolfram talks about. If so, then really accurate prediction would require a model as complex as the phenomenon itself.

All that’s by way of introduction. It provides context for this morning’s session with ChatGPT, where I asked it about weather prediction. My last question is about theoretical limits to accuracy. ChatGPT’s answers suggests that the weather system may ultimately be computationally irreducible. And, if you’re interested in long-term predictions, no matter how accurate your current model, if a large volcano blows the hurls enough dust into the atmosphere to affect the earth’s temperature, your predictions are now shot to hell.

* * * * *

I’m interested in how we predict the weather. I know we use computer models. I know that weather is chaotic, in the mathematical sense, that is, it is sensitive to small variations in initial conditions. I’m interested in two things in particular: 1) How accurate are the forecasts over time? For example what’s the accuracy of a 3-day forecast vs. a 10-day forecast? 2) What’s the size of the data grid on which the models are based? My impression is that the resolution of the grid is stated in kilometers.

Predicting the weather is a complex task, and it relies on a combination of observational data, computer models, and our understanding of atmospheric physics. Let me address your two questions:

1. Accuracy of Weather Forecasts Over Time:
  • Weather forecasts are generally more accurate for shorter timeframes and become less accurate as you extend the forecast period. The accuracy of a forecast depends on various factors, including the region, weather conditions, and the lead time.
  • Short-term forecasts (e.g., 1-3 days) tend to be more accurate, with high precision in predicting temperature, precipitation, and wind patterns. These forecasts are often reliable for planning daily activities.
  • Medium-range forecasts (around 4-7 days) are less precise and are subject to greater uncertainty. While they can provide general trends, they may not be accurate for specific details.
  • Long-range forecasts (beyond 7-10 days) are even less accurate and are more like climatological trends rather than detailed weather predictions.

2. Grid Resolution of Weather Models:

  • Numerical weather models divide the atmosphere into a three-dimensional grid to simulate the behavior of the atmosphere. The resolution of these grids is indeed stated in kilometers and is known as the "grid spacing" or "grid resolution."
  • High-resolution models have smaller grid spacing (e.g., 1-10 kilometers) and can provide more detailed forecasts, especially for local or short-term weather events. These models are typically used for regional weather predictions.
  • Coarser-resolution models have larger grid spacing (e.g., 10-50 kilometers or more) and are used for global weather forecasts, as they cover larger geographic areas. They are less detailed but can provide information on large-scale weather patterns.
  • The choice of grid resolution depends on the specific goals of the forecast and the available computational resources. Weather agencies use a hierarchy of models with different resolutions to cater to various forecasting needs.

It's important to note that advances in technology and improved understanding of atmospheric processes have led to better weather predictions over the years. Weather forecast accuracy has significantly improved for short-term forecasts, and even medium-range forecasts have seen improvements. However, long-range forecasting remains challenging due to the chaotic nature of the atmosphere and the limitations of current modeling capabilities.

What’s the largest computers used in weather forecasting? How long does it take to run a simulation? Are there any weather models so complicated that it takes almost as much time to run the model as it takes for the weather to evolve?

Weather forecasting requires powerful supercomputers to run the complex numerical models that simulate the atmosphere. Some of the largest computers used in weather forecasting include:

  1. IBM Blue Gene: The IBM Blue Gene supercomputers, like Blue Gene/P and Blue Gene/Q, have been used for weather modeling. These supercomputers are known for their high processing power and are used by various meteorological agencies.
  2. Cray Supercomputers: Cray has also provided supercomputing solutions for weather forecasting, such as the Cray XT series and the Cray XC40 series. These supercomputers are designed to handle large-scale simulations.
  3. Fujitsu Supercomputers: The Japanese Meteorological Agency (JMA) has used Fujitsu's supercomputers, such as the K computer, for high-resolution weather modeling.
  4. Various National Supercomputing Centers: Many countries operate their own supercomputing centers for weather forecasting, and these centers use some of the world's most powerful supercomputers. For example, the European Centre for Medium-Range Weather Forecasts (ECMWF) uses its own supercomputer for global weather prediction.

The time it takes to run a weather simulation depends on several factors, including the model's resolution, the complexity of the simulation, and the available computational resources. High-resolution regional models, which cover smaller geographic areas, generally require less time compared to global models. On average, it may take anywhere from several minutes to a few hours to run a single simulation.

Thursday, July 7, 2022

No one had predicted GPT-3. How do you update your priors? [Why learning from history is hard]

More from the Pinker/Aaronson debate on AI scale.

Aaronson at comment #240:

It’s true that I utterly failed to predict the deep learning revolution. I was certainly aware of the thesis, which I associated with Ray Kurzweil, that before long Moore’s Law would cause machines to have as many computing cycles as the human brain, and at that point we should expect human-level AI to “just miraculously emerge.” That struck me as one of the stupidest theses I’d ever heard! Computing cycles aren’t just magical pixie dust, I’d explain: you’d also need a whole research effort, which could take who knows how many centuries or millennia, to figure out what to do with the cycles!

Now it turns out that the thesis was … well, we still don’t know if it’s right all the way to AGI, and certainly great new ideas (GANs, transformer models, etc.) have also played a role, but it’s now clear that the “computing cycles as magic pixie-dust” thesis contained more rightness than almost anyone imagined back in 2000.

So, this is my excuse: I’m not contradicting myself (which is bad), I’m updating based on new evidence (which is good).

But my real excuse is that hardly any of the experts predicted this either. And I just had dinner with Eliezer a couple weeks ago, and he told me that he didn’t predict it. He was worried about AGI in general, of course, but not about the pure scaling of ML. The spectacular success of the latter has now, famously, caused him to say that we’re doomed; the timelines are even shorter than he’d thought.

While it caused Eliezer to update from “we should all worry about this” to “screw it, we’re doomed,” it caused me and quite a few others to update from “we shouldn’t all worry about this” to “we should all worry about this.”

Me at comment #280 after quoting from Scott’s #240:

It caused me to update from “the space of possible minds is huge” to “the space of possible minds is even larger than I thought it was.” My update is different from yours, but doesn’t necessarily contradict it. More like orthogonal to it.

This language of “updating” comes from Bayesian statistics, which I do not know on a technical level. But then, in this kind of context, it is not used technically. This usage is ubiquitous in the so-called rationalist community.

Roughly speaking, you have some idea of what’s going on in some domain, in this case, artificial intelligence. That idea is your prior and implicitly entails predictions about how that domain will unfold over time. If things unfold in a way that is consistent with your prior views, then things won’t surprise you. When something surprising happens, though, that’s a signal that your priors are wrong. You must now adjust your priors. That’s what Aaronson is talking about in the last three paragraphs I quoted from him and what I’m talking about in my paragraph. Now, while Baysianism tells you to update your priors, it doesn’t tell you just how to update your priors.

I continue my comment with my now standard analogy for dealing with large complex problems, Christmas tree lights:

I’m a bit more interested in understanding the brain than I am in scaling Mount AGI. Here’s how I’ve been thinking about understanding the brain. Imagine that understanding means takes the form of a string of serial-wired Christmas tree lights, 10,000 of them. To consider the problem solved all the lights have to be good so that the string lights up.

Instead of understanding the brain, apply the analogy to understanding how to create AGI (whatever that is). Let’s start at 1956, the year of the Dartmouth conference. It’s at that point we were handed the string and were told, “get this to light up and you’ve solved AI.” Since digital computing had been a going concern for over a decade at that point and work had already been done on chess and on natural language, some of the bulbs in that 10,000 bulb string were good. But we did’t know how many or where they were. Between 1956 and whenever OpenAI started working on GPT-3 we’d replaced, say, 2037 bad bulbs with good ones. Let’s say that in creating GPT-3 OpenAI replaced 100 bulbs, which we know about. So 2137 bulbs have been replaced. How many more bulbs to go before all of them are good?

Obviously we don’t know. Some people seem to think it’s only a couple of hundred or so, most of them having to do with scaling up even further. What, beyond wishful thinking, justifies both the belief that the unknowns cluster in one area and that their number is so low? Maybe we still have over 5000 or 6000 to go, maybe more. Who knows?

That is to say, it’s one thing to adjust your priors by thinking we’re now at long last on the right track and quite different to think, as I did, the world just got much larger. Those different updates reflect, in effect, two different sets of priors.

I go on to say something about where my priors come from:

By way of calibration, I should note that back in the era of symbolic computing I had once felt – though never published – that we were within 20 years of being able to build a system that read Shakespeare plays in an “interesting” way. By “interesting” I meant that we could have the system read, say, The Winter’s Tale, and then we’d open it up, trace what it did, and thereby learn what happens we humans read that play. That is to say I believed we could construct a system that could reasonably be construed as a simulation of the human mind. Alas, the AI Winter of the mid-1980s killed that dream. While these new post GPT-3 systems are wonderful, I see little prospect that any of them can be considered a simulation of the human mind nor that any of them will be able to shed insight into Shakespeare in the near or mid-term future. Beyond, say, 2140 (the year of Kim Stanley Robinson’s New York 2140) I’m not prepared to say.

My sense of such matters is that reading about such collapses of intellectual projects in a history book is not the same as living through one. The valence is much weaker. So I’m sticking with my new prior, “the space of possible minds is even larger than I thought it was.”

And THAT difference, I believe, is crucial. Living through a set of events, AI Winter, affects your priors in a way that is quite different from only knowing about those events from a historical account. In both cases we’re dealing with an ongoing stream of events, the evolution of AI research from the 1950s into the present and on to the future. I suspect the difference can ultimately be traced to the brain. What someone reads about events simply does not affect “deep circuitry” the same way as experiencing those same events, even if one really really believes what was read. 

More generally, this is one reason that learning from history is so very difficult. What you’ve lived through is much more potent, has a greater effect on your updating mechanism, that what you’ve only read about or heard from third parties. Everyone’s experience is necessarily limited. If there is any wisdom that does indeed accrue to age, this would be one source of it. But there is obviousy a limit to how long one person can live.

I’ve discussed this before, in particular, in a post from 2021, Things change, but sometimes they don’t: On the difference between learning about and living through [revising your priors and the way of the world].

Thursday, June 23, 2022

The Two Voices of Scott Alexander on Rogue AI

Back at the end of February, Tyler Cowen had a post entitled, “Are nuclear weapons or Rogue AI the more dangerous risk?” He linked to a long Scott Alexander post, “Biological Anchors: A Trick That Might or Might Not Work,” which was about some recent web discourse around and about predicting the emergence of human-level AI. Then, at the very end on his long post, seemingly out of nowhere, Alexander was fretting about danger of rogue AI.

I seem to have gotten trapped in Alexander’s post. I have read it several times, even taking notes. It’s a very interesting document and merits some discussion of how it is constructed. I’m not so much concerned about whether or not Alexander’s assessment of the prospects of human level AI is valid as I am about the convoluted nature of his post.

Two voices

Alexander writes the post in two voices. While I’ve not read a lot of his material, I’ve read enough to know that he’s a careful and skilled writer. If he spoke through two voices it must be because whatever he wants to convey arises from the interaction between them and cannot be stated within a single voice.

Let’s call one of the voices the Impersonal voice. Most of the post is written in that voice, which is the voice in which he’s written most of the posts I’m familiar with. Let’s call the other voice the Personal voice. By word-count it’s by far the lesser voice, but it packs a strong rhetorical punch.

Let’s look at the two strongest statements from the Personal voice. The first, and I believe longest, section of the report is Alexander’s account of Ajeya Cotra’s long report, Forecasting TAI with biological anchors, which I’ve not read (though I’ve read some of what Holden Karnofsky says about it). Very near the end of this section the Personal voice makes a strong statement (though this is not the first appearance of this voice):

One more question: what if this is all bullshit? What if it’s an utterly useless total garbage steaming pile of grade A crap?

Our second example comes near the end of the post, when Alexander begins his own assessment of things:

Oh God, I have to write some kind of conclusion to this post, in some way that suggests I have an opinion, or that I’m at all qualified to assess this kind of research. Oh God oh God.

Phrases like “total garbage,” “grade A crap” and “Oh God” are not appropriate to the work Alexander is doing through his (standard and) Impersonal voice. They signal us that we are listening to a different voice. This voice expresses a merely personal attitude and is quite different from the objectivity sought in the Impersonal voice.

Taken at face value the second quoted statement says Alexander doesn’t feel (technically) qualified to judge this material. As such, it also tells us why he’d made that first statement and in that voice. That first statement places the assertion, Cotra’s report is nonsense, into the record. By couching that assertion in the words and manner of the Personal voice Alexander separates it from his Impersonal summary of the report. In effect, The guy who summarized the report is not the guy who thinks it’s nonsense. The guy who summarized the report doesn’t feel competent to assess it, but happens to be closely coupled to the guy who has deep doubts.

From parody to grudging affirmation to confusion

So, Alexander has gotten a statement of deep doubt into the record. What happens next? He goes back into the Impersonal voice and invites us to

Imagine a scientist in Victorian Britain, speculating on when humankind might invent ships that travel through space. He finds a natural anchor: the moon travels through space! He can observe things about the moon: for example, it is 220 miles in diameter (give or take an order of magnitude). So when humankind invents ships that are 220 miles in diameter, they can travel through space!

He then spins out that tale and includes a helpful chart and a picture. It’s absurd – he does slip in a wink or two. It’s a parody of the methodology in Cotra’s report. A parody is not an argument, but it clarifies Alexander’s fears about the report.

Then he goes into his second major section, an account of Eliezer Yudkowsky’s critique, which takes the form of a long post in dialogue form. I’ve taken a look at the post, but haven’t read the whole thing. Alexander’s opens this section by asserting, “Eliezer Yudkowsky presents a more subtle version of these kinds of objection in an essay…” I won’t bother to say anything about that beyond noting the Yudkowsky thinks Cotra’s method is useless for estimating the arrival of Transformative AI. Alexander may not feel qualified to critique Cotra’s work, but Yudkowsky certainly does. And why not? As Alexander notes: “...he did found the field [AI alignment], so I guess everyone has to listen to him.”[1]

When he’s finished with Yudkowsky, Alexander discusses comments from other places (LessWrong, AI Impacts, and OpenPhil) and finally offers his own evaluation, which he opens with the “Oh God” statement I’ve already quoted. He says a thing or two and arrives at this:

Given these two assumptions - that natural artifacts usually have efficiencies within a few OOM [orders of magnitude] of artificial ones, and that compute drives progress pretty reliably - I am proud to be able to give Ajeya’s report the coveted honor of “I do not make an update of literally zero upon reading it”.

That still leaves the question of “how much of an update do I make?” Also “what are we even doing here?”

I take it that the passages he puts in quotes are being spoken though the Personal voice.

Let’s look at the first one. It’s stated in informal Baysian terms. It also feels arch and indirect. “I do not make an update of literally zero”? What’s that? He’s refrained from writing “0” in a ledger somewhere? Whereas if the report had been less convincing, he’d have

  • opened up that ledger,
  • added a line,
  • placed “Ajeya’s report” in the Argument column, and
  • written “0” in the Effect on Me column.

On the contrary, the report has had some non-zero effect on him. But then he attempts to run away: “what are we even doing here?”

A couple paragraphs later: “This report was insufficiently different from what I already believed for me to need to worry about updating from one to the other.” I’m not sure what to make of this. If his prior belief had been quite different from the report’s conclusion, then a decision to stick with that prior belief would represent lack of faith in the report (which he doesn’t feel competent to judge). In contrast, a decision to revise his belief in the direction indicated by the report would represent faith in Cotra (and her colleagues) despite his lack of (technical) qualifications for judging the report. So, he can’t judge the report, doesn’t think it’s wrong, but doesn’t think it’s right enough to lead him to change his mind.

It's as though he’s playing the role of Penelope in Odyssey. She tells her suitors she’ll pick one when she’s done weaving a burial shroud for Laertes. They see her diligently weaving during the day. But then at night, what does she do? She undoes the weaving she’d done during the day so that she can put off the day she has to pick one.

“I’m already scared”

And now, at long last, Alexander gets around to what was obviously on his mind from the very beginning, fear of a rogue AI. He‘s very little that up to this point, which is strange in itself, but now he comes out with it.

Alexander circles back to Yudkowsky: “The more interesting question, then, is whether I should update towards Eliezer’s slightly different distribution, which places more probability mass on earlier decades.” Yudkowsky, however, refuses to give dates: “I consider naming particular years to be a cognitively harmful sort of activity...” Incidentally, sounds like self-regarding grandstanding from Yudkowsky. Perhaps, as people surround him asking for the date, he passes out gilded fortune cookies as souvenirs.

Alexander:

So, should I update from my current distribution towards a black box with “EARLY” scrawled on it?

What would change if I did? I’d get scared? I’m already scared. I’d get even more scared? Seems bad.

That’s not the end. We’ve got two or three more paragraphs. But those paragraphs don’t change the fact that Alexander is scared. They just mix a bit of wit into the contemplation of doom.

Alexander and his community

What are we to make of all this?

I don’t quite know. But then neither does Alexander.

I note, however, that he wrote that post for a community he’s been cultivating for almost a decade. He’s written on a wide variety of topics. He’s conducted surveys, had book review contests, and interacted with that community in various says. That long ambivalent post, spoken through two voices, that’s the post he felt he owed that community. To what extent is Alexander’s ambivalence a reflection of attitudes in that community?

On the one hand there’s the apparent assumption that Transformative AI is on the way come hell and high water. That is coupled with interest in the arcane technical minutiae and leaps of epistemic faith required to issue a long report predicting that arrival by comparison with biological information processing in 1) the human brain, 2) a human life, 3) the evolution of life on earth, and 4) the genome, all measured in FLOPS (floating-point operations per second).[2] That’s one thing.

And there there’s the correlative assumption, perhaps not shared by all, but nonetheless widespread, that the arrival of Transformative AI brings with it the danger that that AI will turn against humanity and transform the earth into a paperclip factory, metaphorically speaking. To the extent that the rhetorical structure of his post is responding to what? a facture, ambivalence? in his audience, why does Alexander hold it in reserve until the end, like it is a shameful secret?

Notes

[1] That depends on what one thinks of the field of AI alignment. Color me skeptical. The idea that we are under not-so-distant threat from an AI hell-bent on our destruction strikes me as conspiracy theorizing directed at technology no one knows how to build.

[2] In his history of technology, David Hays tells us that

... the wheel was used for ritual over many years before it was put to use in war and, still later, work. The motivation for improvement of astronomical instruments in the late Middle Ages was to obtain measurements accurate enough for astrology. Critics wrote that even if the dubious doctrines of astrology were valid, the measurements were not close enough for their predictions to be meaningful. So they set out to make their instruments better, and all kinds of instrumentation followed from this beginning.

I feel a bit like that about the topics of Ajeya Cotra’s report. They are interesting and important in themselves for what they tell us about the world. The need not be yoked to the task of predicting future technology.

Wednesday, June 1, 2022

Gary Marcus has raised a $350K bet against Musk's prediction of AGI by 2029. Will it go higher?

Super-AGI is pumped!

The prediction:

The $100 K bet is in the linked article, which also sets forth the 5 criteria Marcus suggests:

Another $100 K:

And another $100 K:

And another $50 K:

That's where the bet stands as of 6:02 AM EDT on Wednesday June 1, 2022. How high can it go? Will it get to a million?

Wednesday, May 4, 2022

Predicting high-level textual phenomena from low-level feartures

Friday, March 4, 2022

Rodney Brooks has been making predictions: Concerning AI, “We’re still back in phlogiston land…”

Back on January 1, 2018 Rodney Brooks issued fairly specific predictions in three areas: 1) self-driving cars, 2) Artificial Intelligence, machine learning, and robotics, and 3) progress in the space industry. There are over a dozen predictions in each of those three areas. Brooks has updated those predictions each year since and plans to do so until 2050. You can find the most recent update, for 1.1.22, here: https://rodneybrooks.com/predictions-scorecard-2022-january-01/.

I’m not going to reprise any of those specific updates here, but I’d like to copy over some of his commentary for that second area, Artificial Intelligence, machine learning, and robotics.

Where’s the next big thing?

Back in 2018 I predicted that “the next big thing”, to replace Deep Learning, as the go to hot topic in AI would arrive somewhere between 2023 and 2027. I was convinced of this as there has always been a next big thing in AI. Neural networks have been the next big thing three times already. But others have had their shot at that title too, including (in no particular order) Bayesian inference, reinforcement learning, the primal sketch, shape from shading, frames, constraint programming, heuristic search, etc.

We are starting to get close to my window for the next big thing. Are there any candidates? I must admit that so far they all seem to be derivatives of deep learning in one way or another. If that is all we get I will be terribly disappointed, and probably have to give myself a bad grade on this prediction.

So far the things that I see bubbling around and getting people excited are transformers, foundation models, and unsupervised learning.

Concerning transformers:

These language models are over interpreted by people as understanding what they are spitting out, especially when the press writes stories where they have cherry picked responses. But they come with incredible problems, including copyright violations, intellectual theft of code, and even outright life threatening danger when they find their way into consumer products. Tech companies have a real problem in rushing some of these systems to market.

Continuing on:

Foundation models are large trained models that start out as a basis for tuning particular applications. There has been some self important announcements with a sort of me too feel (“Hey, I produced a foundation model too!!”), which don’t amount to much of an intellectual contribution. If this turns out to be the next big thing I am going to have to rip off my mask of equanimity and revert to my natural state of being a grumpy old man.

Unsupervised learning is an idea that has been around for a long time. Not a big intellectual jump to want to get it into deep learning–may be a hard technical problem, but not an intellectual breakthrough this time around.

The problem with AI

I have often stated that I think the field of AI, despite the great practical successes recently of Deep Learning, is probably a few hundred years away from where most people think it is. We’re still back in phlogiston land, not having yet figured out the elements, including oxygen.

Read that again and think about it. Does he really mean that? Why would he say such a thing? Is he nuts?

Let us assume that he’s correct. Given how impressive some current AI demonstrations are, can we not take Brooks’s view as implying that we have learned, or at least have the potential to learn, about ourselves and our own capacities? [Yeah, I know, that needs some unpacking. Maybe later.]

After he goes through his 14 specific predictions, Brooks reminds us of his bona fides:

AI, Robotics, and Machine Learning are areas that I have a real personal investment in. I wrote a terrible Masters thesis on ML back in 1977. I joined the Stanford AI Lab later that year, then the MIT AI Lab four years later, and became director of that lab in 1997, merging it with LCS (Lab for Computer Science) to form MIT CSAIL in 2003, the largest lab at MIT, still today. I have founded six AI and robotics companies. After 45 years in the academic and industry trenches can I be unbiased? Probably not.

I know that many who disagree with me will dismiss me for all that experience that I have. Perhaps those who agree with me should also dismiss me for the same reason!!

That last paragraph is interesting. Why would someone dismiss him for all his experience? He really knows this stuff, no? How can anyone look at this area without being biased in some way? Doesn’t naivete impose its own biases?

As you know, I’m of the belief that we’re in transition from one intellectual era to another. To which era does AI, robotics, and machine learning belong, the old one or the new. Maybe it straddles both. Maybe AI and robotics are old, machine learning new. Or maybe the perceptron is old, transformers new? Are we talking phlogiston or oxygen? How do you tell?

He goes on to state:

My current belief is that it all gets back to the symbol grounding problem, and even more deeply to adopting a computational approach to AI, Robotics, and ML (and I expect almost no one will agree with that latter claim).

Color me sympathetic to that last claim, that the computational approach is problematic. I’ve written a post on Brooks’s views: Has the computer metaphor for the mind run out of steam? New Savanna, June 19, 2019, https://new-savanna.blogspot.com/2019/06/has-computer-metaphor-for-mind-run-out.html.

He concludes by mentioning Brian Cantwell Smith, The Promise of Artificial Intelligence.

In this book Smith introduces the idea of registration, as a maintained relationship between an object outside of us and what goes on inside our head (and he would have it also in a classical computer) despite changes in perception and even context.

I’ve not read the book, but I’ve read reviews. I believe Smith introduces a distinction between reckoning and judgement. Reckoning is what computers do, but only humans are capable of judgement, at least so far. Intelligence requires judgement. I think we do need a fairly specific term for what it is that AI systems do. I kind of like “reckoning”. Note: Smith talks about registration in the video I've embedded here.

Friday, January 28, 2022

Neurons learn by predicting future activity

Wednesday, July 28, 2021

Things change, but sometimes they don’t: On the difference between learning about and living through [revising your priors and the way of the world]

This is a year old (first published in July 26 of last year). I'm bumping it to the top of the queue because I'm thinking about these things once again.
It has long been obvious to me that there is an epistemic difference between knowing history and living through history. I now have half a notion of how to think about this. That’s what this post is about.

I begin by talking about the effect the fall of the Soviet Union had on me. That’s my paradigm example of the phenomenon I’m talking about. Call it a RUPTURE in my sense of the world. Then I consider my intellectual career as a series of ruptures where I had to reconsider my priors, if you will, and so rethink my intellectual foundations. Finally I move to a graphics revolution that could have happened, SHOULD have happened, as a result to digital technology. I saw it coming, but it never really got started. What was I missing? I conclude with some reflections on the current situation.
 
I list five such ruptures in all. Three INSIDE my primary intellectual trajectory, my intellectual career, while two are OUTSIDE that trajectory, but nonetheless affected me.

World history and the fall of the Soviet Union

It’s about having to revise your priors – I’m talking Bayes, as least informally, your prior commitments. Living through events forces that on you, or at any rate, gives you the opportunity to revise those priors. Merely learning about the past doesn’t do that. Changing how you think is deeper than learning about change.

My paradigmatic example is the fall of the Soviet Union. I grew up in the 1950s and ‘60s, when the Cold War was going strong. I remember reading about how to construct bomb shelters, and thinking about where the family shelter should be. I remember talk of the missile gap; I remember the Cuban missile crisis. I fully expected to be living the Cold War when I died.

And then the Berlin Wall came down in 1989 and the Cold War was all but over. Though some had anticipated this – I’m thinking particularly of U.S. Senator Daniel Patrick Moynihan – I certainly did not. It took me by surprise.

What did I learn from that? Simple, that The World can change. Sure, I knew world history, I knew that things changed deeply and fundamentally, time and time again. But I hadn’t seen it for myself. I suppose I might have taken that simple lesson from the Civil Rights movement, but perhaps the persistence of racism blunted that achievement. The end of the Vietnam War? No, and I’d marched against that one and been a conscientious objector.

It was the fall of the Soviet Union that reached me, that forced me to abandon a set of simple, but fundamental – I was about to type “epistemic”, “epistemic commitments”, but no, it’s deeper – ontological commitments. For me the Cold War had ontological force. It was simply the way of the world. I learned about it in childhood and lived it through a quarter century of adulthood.

And then the world changed.

Let’s call that an OUTSIDE RUPTURE in my sense of the world, outside because it was outside my primary sphere of action and commitment.

My intellectual life

My intellectual life has had a number of “ruptures”, if you will. In the late 1960s I learned to interpret literary texts as an undergraduate at Johns Hopkins. I was trained in so-called “close reading”, but was interested in structuralist analysis as well. In the spring of my senior year I became interested in “Kubla Khan.” I decided to undertake a structuralist analysis of the poem for a master’s thesis.

Worksheet, first 36 lines of "Kubla Khan"

As things worked out that project fell apart in 1970 or ‘71. “Kubla Khan” refused to submit to the methods I applied to it. My commitment to the poem won out and I was forced to abandon those methods. I tell the story in considerable autobiographical detail in “Touchstones” [1] and supply the intellectual detail in a working paper on Lévi-Strauss [2].

FIRST INSIDE RUPTURE, inside because it is dead center in my primary sphere of action and commitment. I could no longer believe that existing methods of literary criticism were adequate to the task of understanding how literary texts work. New methods, incidentally, that required visual thinking. Standard literary criticism has been and remains committed to expository prose as its fundamental intellectual medium.

In 1973 I went off to get a Ph. D. in the English Department at SUNY at Buffalo and ended up spending considerable time studying computational semantics with David Hays in the Linguistics Department. My dissertation ended up as a quasi-technical exercise in cognitive semantics as applied to literature, one where, for large sections, I drew the diagrams first and then wrote prose to explain them. In 1976 Hays and I published a humanities-oriented review of computational linguistics in which we proposed the development of symbolic systems capable of “reading” a Shakespeare play in an intellectually interesting way [3]. I fully expected to be working with such a system later in my career. That hasn’t happened, nor do I expect it to.

Part of a semantic network for Shakespeare's Sonnet 129

SECOND INSIDE RUPTURE. Symbolic computational systems are not adequate for understanding the mind.

But this was not so drastic as that first inside rupture. For one thing, Hays and I didn’t really believe that symbolic systems were adequate to the task. We’d been exploring how to ground them in analog systems and in distributed neural computation [4]. The failure, rather, was one of degree rather than kind. I had overestimated the power of symbolic systems. I would later pick up my interest in neural systems in conversations I had with the late Walter Freeman and I incorporated into the early chapters of my book on music, Beethoven’s Anvil (2001) [5].

I suppose there was a THIRD INSIDE RUPTURE, or quasi-rupture, as well. In 1995 I discovered that a bunch of literary scholars was interested in “the cognitive revolution” – you know, the ideas I’d begun investigating in graduate school. That was in the Stanford Humanities Review. My work was forgotten but, and more importantly, computation was nowhere to be seen in this newer. Their version of cognitive science was thus very different from mine. It was in working through that difference that I figured out that I had been chasing form all along.