Showing posts with label Yudkowsky. Show all posts
Showing posts with label Yudkowsky. Show all posts

Wednesday, June 24, 2026

Eliezer Yudkowsky, LessWrong, OpenAI & countercultures past & present

Claude 4.6 Medium summarizes a dialog I had with Gemini that started with an inquiry about 

1) Eliezer Yudkowsky, his early ideas, & his early following, 
2) then to his interactions with Peter Thiel, Elon Musk, and Sam Altman that got OpenAI started, 
3) to my own sojourn on LessWrong and 
4) concluded with the connection between 1960s counterculture and contemporary Silicon Valley computer culture.

For some reason Claude pretended to summarize the conversation in my voice.

* * * * *

A Conversation with Gemini About Eliezer Yudkowsky Bill Benzon, new-savanna.blogspot.com, June 24, 2026

I recently had an extended exchange with Gemini (accessed through the Google search interface) about Eliezer Yudkowsky — a figure I've been thinking about in the context of AI culture more broadly. What follows is a summary of where the conversation went.

It began with a query about Yudkowsky's 2007 paper Levels of Organization in General Intelligence (LOGI), which argues that recursive self-improvement could allow an AI to rapidly cycle through levels of cognitive architecture in ways that would break traditional training and testing boundaries. Gemini gave a competent account of the paper's significance for AI alignment theory.

I then offered my own assessment: reading LOGI years ago, I concluded it was the kind of work produced by a brilliant college sophomore who had figured out everything and decided to write it up. The sort of student you'd want to guide and nurture — but of course, that never happened with Yudkowsky, who is entirely self-educated. Gemini agreed this was a common reaction, and traced the characteristic features of his writing — grand scope, idiosyncratic jargon, overconfidence — to the absence of the standard academic filters that would normally shape a thinker. Without a thesis advisor to push back, he co-founded his own institutions (MIRI, LessWrong), creating an insulated subculture where he became the mentor rather than the student.

I offered a specific passage from LOGI as an example of what goes wrong. Yudkowsky dismisses semantic networks as "completely bankrupt" on the grounds that they're simple enough to write on paper. Gemini correctly identified this as a classic category error: confusing the notation with the mechanism. The diagram on the whiteboard is inert; what matters is the graph-traversal algorithms, the spreading activation, the interpreter running the data structure. Ironically, Yudkowsky later wrote extensively about the Map-Territory Fallacy — but as I put it to Gemini, he is constantly mistaking a map for the territory. His entire worldview treats clean theoretical proofs as if they dictate messy engineering realities.

From there the conversation turned to how Yudkowsky managed to build such a large following despite these intellectual weaknesses. Gemini confirmed that Harry Potter and the Methods of Rationality, his 660,000-word fanfiction, was openly designed as a recruiting tool — drawing technically minded young people into the Rationalist and AI safety ecosystems. Countless engineers and founders who later populated early AI labs first encountered his ideas through that story.

The crowning irony: Yudkowsky's warnings about AI helped convince Elon Musk, Peter Thiel, and Sam Altman that humanity needed a counterweight to closed corporate AI efforts — which led directly to the founding of OpenAI in 2015. Once OpenAI pivoted to the empirical, data-driven methodology of large language models, they completely bypassed the deductive logic "maps" Yudkowsky had spent decades drawing. Sam Altman acknowledged Yudkowsky's role in a 2023 tweet, noting that he had arguably done more to accelerate AGI than anyone else, and adding that he might someday deserve a Nobel Peace Prize.

I told Gemini that wouldn't be necessary. I also shared my own experience: I joined LessWrong around the time ChatGPT launched, initially as an anthropological participant-observer, but stayed for the conversation, which I found genuinely useful. There are very smart people there. But the insularity was unmistakable — and I described one telling episode: someone on the forum was trying to spread Rationalism in Japan and struggling. I pointed out that Japanese popular culture, from Osamu Tezuka's Astro Boy through the Ghost in the Shell franchise, has a long history of viewing robots and AI as fundamentally benevolent — an expression of Shinto techno-animism, in which kami can reside in machines as naturally as in rivers. They simply didn't know about this cultural background. Gemini observed that the Western doomer ethos is rooted in a Frankenstein complex with Judeo-Christian substrata: creating life is hubris, and the creation must turn on its master. The Japanese paradigm operates from entirely different premises. The LessWrong response to my observation? They noted it and moved on.

After a while I tired of the place. One small anecdote captures the texture of the experience: I frequently link out to other things I've written, and one LessWrong post linked to an essay-review I'd done of Benny Shannon's book on ayahuasca. I noticed a significant spike of traffic to my Academia page coming from LessWrong and pointing to that essay — which tells you something about the undercurrent of interest in altered states of consciousness running alongside the dry decision theory. They approach psychedelics with an engineer's curiosity: the brain as a computer, phenomenology as data.

The conversation ended with what I think is the most useful historical frame. LessWrong is, in structure and function, a counterculture — but centered on computers and AI, with Yudkowsky as guru rather than Timothy Leary. And there's a genuine genealogical link through San Francisco and transhumanism: Stewart Brand bridging the Merry Pranksters to personal computing, the Extropians of the 1990s who wanted to transcend the body via nanotechnology and cryonics rather than LSD, and then Yudkowsky emerging from that same Bay Area Transhumanist mailing-list culture. Fred Turner's From Counterculture to Cyberculture maps this lineage. The counterculture became the vanguard — but corporate reality, as I noted to Gemini, has not submitted. Bill Gates and his successors were never absorbed by the counterculture. Peter Thiel, who was an early funder of MIRI, has since publicly labeled Yudkowsky a Luddite and positioned AI safety concerns as obstacles to American technological dominance. The "well-run alternative universe" of LessWrong lost all leverage once scaling deep learning required billions of dollars in silicon, electricity, and data centers. The colorful intellectual vanguard warmed society up to the idea of AGI; then the massive engine of global capitalism took the steering wheel.

The subculture keeps its cozy, insular forum to debate the semantics of the map. The corporate empires plow ahead across the territory.

Monday, January 13, 2025

#Elonald: An apocalyptic conjunction & the contradictions of capitalism

Donald Trump, President-Elect of the United States, and Elon Musk, the richest man in the world and a technological wizard, what a conjunction. Remember, even as the country is preparing to receive Trump as its 47th President, Elon Musk is suing OpenAI for, in effect, having abandoned its core mission of peaceful AI. It is my understand that Eliezer Yudkowsky, Doomster-in-Chief, played a role in the meetings that resulted in the creation of OpenAI.

So, in an article in 3 Quarks Daily, I argued that the cult of AI doom is best conceived of as an aspect of a high-tech inspired counter culture comparable to the drug inspired counter culture of the 1960s and 1970s. Yudkowsky plays a role in the current counterculture comparable to the role that Timothy Leary played in the 1960s counter culture. But as far as I know, Leary never got as close to the levers of power as Yudkowsky is.

On the one hand we have this rationalist doomster culture that is devoted to scaring itself witless with stories about rogue AI. Many members of this culture make their livings in the AI business. They’re scaring themselves silly with stories about being destroyed by the products of their own labor. What would Karl Marx make of that? Talk about alienation....

And, through the person of Elon Musk and his high-tech buddies, these doomsters are part of the cultural wave that brought Donald Trump to his second term as President. Trump got there, in large part, by exploiting fear in the electorate, with much of that fear directed at immigrants. Is that the sort of thing Marx was talking about when he talked of the contradictions of capitalism? Musk fears his own shadow while Trump is at war with his.

The mind boggles.

Tuesday, November 12, 2024

Yudkowsky + Wolfram on AI Risk [Machine Learning Street Talk]

This is a long, rambling, conversation (4 hours), so I have a hard time recommending the whole thing. I’d say that Wolfram and Yudkowsky do manage to find one another by the 4th hour (sections 6 & 7) and say some interesting things about computation and AI risk (much of the earlier conversation was on tangential matters). I note that the whole thing has been transcribed and there’s a Dropbox link for the conversation.

I will note that the conversation they did have was much better than what I had anticipated, which was a lot of talking past one another. And, yes, there was some of that, but as soon as that got going they worked hard at understanding what each was getting at.

Wolfram has some interesting remarks on computational irreducibility scattered throughout – that’s certainly one of his key concepts, and an important one. He also asserts, here and there, that he’s long been used to the idea that he faces computers smarter than he is; he also notes that he regards the universe as smarter than he is.

My sense is that the computation and AI risk stuff could be written up in a tight 2K words or so, but I don’t have any plans to make the attempt. That might be an exercise for a good student. Perhaps one of the current LLMs (Claude?) could do it.

TOC:

1. Foundational AI Concepts and Risks
[00:00:00] 1.1 AI Optimization and System Capabilities Debate
[00:06:46] 1.2 Computational Irreducibility and Intelligence Limitations
[00:20:09] 1.3 Existential Risk and Species Succession
[00:23:28] 1.4 Consciousness and Value Preservation in AI Systems

2. Ethics and Philosophy in AI
[00:33:24] 2.1 Moral Value of Human Consciousness vs. Computation
[00:36:30] 2.2 Ethics and Moral Philosophy Debate
[00:39:58] 2.3 Existential Risks and Digital Immortality
[00:43:30] 2.4 Consciousness and Personal Identity in Brain Emulation

3. Truth and Logic in AI Systems
[00:54:39] 3.1 AI Persuasion Ethics and Truth
[01:01:48] 3.2 Mathematical Truth and Logic in AI Systems
[01:11:29] 3.3 Universal Truth vs Personal Interpretation in Ethics and Mathematics
[01:14:43] 3.4 Quantum Mechanics and Fundamental Reality Debate

4. AI Capabilities and Constraints
[01:21:21] 4.1 AI Perception and Physical Laws
[01:28:33] 4.2 AI Capabilities and Computational Constraints
[01:34:59] 4.3 AI Motivation and Anthropomorphization Debate
[01:38:09] 4.4 Prediction vs Agency in AI Systems

5. AI System Architecture and Behavior
[01:44:47] 5.1 Computational Irreducibility and Probabilistic Prediction
[01:48:10] 5.2 Teleological vs Mechanistic Explanations of AI Behavior
[02:09:41] 5.3 Machine Learning as Assembly of Computational Components
[02:29:52] 5.4 AI Safety and Predictability in Complex Systems

6. Goal Optimization and Alignment
[02:50:30] 6.1 Goal Specification and Optimization Challenges in AI Systems
[02:58:31] 6.2 Intelligence, Computation, and Goal-Directed Behavior
[03:02:18] 6.3 Optimization Goals and Human Existential Risk
[03:08:49] 6.4 Emergent Goals and AI Alignment Challenges

7. AI Evolution and Risk Assessment
[03:19:44] 7.1 Inner Optimization and Mesa-Optimization Theory
[03:34:00] 7.2 Dynamic AI Goals and Extinction Risk Debate
[03:56:05] 7.3 AI Risk and Biological System Analogies
[04:09:37] 7.4 Expert Risk Assessments and Optimism vs Reality

8. Future Implications and Economics
[04:13:01] 8.1 Economic and Proliferation Considerations

Wednesday, February 28, 2024

Power, Persuasion, and Influence in The Valley

Henry Farrell has a new post at Crooked Timber: Dr. Pangloss’s Panopticon (Feb. 27, 2024). He's replying to Noah Smith's negative review of Acemoglu and Johnson, Power and Progress

On power:

What Acemoglu and Robinson are saying is something quite different than what Noah depicts them as saying. For sure, they acknowledge that persuasion has some stochasticity. But they stress that it is not a series of haphazard accidents. Instead, under their argument, there are some kinds of people who are systematically more likely to succeed in getting their views listened to than other kinds of people. This asymmetry can reasonably be considered to be an asymmetry of power.

Under this definition, power is a kind of social influence. Again, it is completely true that it is extremely difficult to isolate social influence from other factors, proving that social influence absolutely caused this, that, or the other thing. But if Noah himself does not believe in the importance and value of social influence, then why does he get up in the morning and fire up his keyboard to go out and influence people, and why do people support his living by reading him?

I imagine Noah would concede that social influence is a real thing! And if he were actually put to it, I think that he would also have to agree to a very plausible corollary: that on average he, Noah Smith, exerts more social influence than the modal punter argufying on the Internet. Lots of people pay to receive his newsletter; lots of other people receive it for free. That means that he is, under a very reasonable definition, more powerful than those other people. He is, on average, more capable of persuading large numbers of people of his beliefs than the modal reply-guy is going to be.

This understanding of power is neither purely semantic nor empirically useless. Again, it may be really difficult to prove that Noah’s social influence has specific causal consequences in a specific instance. But the counter-hypothesis – that Noah’s ability to change minds, given his umpteen followers, is the same as the modal Twitter reply guy – is absurd. Occasionally, random people on the Internet can be temporarily enormously influential. Sometimes, super prominent people aren’t particularly successful at getting their ideas to spread. But on average, the latter kind of people will have more influence than the former. We can reasonably anticipate that people with lots of clout (whether measured by absolute numbers of followers, numbers of elite followers, bridging position between sparsely connected communities or whatever – there are different, plausible measures of influence and lively empirical debates about which matters when) will on average be substantially more influential than those with little or none. This means, for example, that it will be very difficult for ideas or beliefs to spread if they are disliked by the highly connected elite.

Now in fairness to Noah, Acemoglu and Johnson don’t help their case by using a wishy-washy seeming term like “persuasion.” But if you think about “persuasion” as some combination of “social influence” and “agenda control,” you will get the empirical point they are trying to make.

Core claims:

Acemoglu and Johnson’s core claims, as I read them are:

  1. That the debate about technology is dominated by techno-optimists [they actually write this before Andreessen’s ludicrous “techno-optimist manifesto” but they anticipate all its major points].
  2. That this dominance can be traced back to the social influence and agenda setting power of a narrow elite of mostly very rich tech people, who have a lot of skin in the game.
  3. That their influence, if left unchecked, will lead to a trajectory of technological development in which aforementioned very rich tech people likely get even richer, but where things become increasingly not-so-great for everyone else.
  4. That the best way to find better and different technology trajectories, is to build on more diverse perspectives, opinions and interests than those of the self-appointed tech elite, through democracy and countervailing power.

Since I more or less endorse all these claims (I would slightly qualify Claim 1 to emphasize mutually reinforcing pathologies of tech optimism and tech pessimism), I think that Power and Progress is a really good book, in ways that you won’t understand if you just relied on Noah’s summary of it (I note that this book and my own with Abe Newman are both shortlisted for a very nice prize, but that is neither here nor there in my opinion of it). I haven’t read another book that lays out this broad line of argument so clearly or so well. I haven’t read another book that lays out this broad line of argument so clearly or so well. And it is a very important line of argument that is mostly missing from current debates. Noah speculates that the book hasn’t gotten much attention because it is lost amidst the multitudes of tech pessimistic accounts. My speculation is that it has gotten less attention than it deserves because reviewers and readers don’t know quite how to categorize it, given that it approaches the issues from an unexpected slant.

The panopticon:

The panopticon may indeed have efficiency benefits. People can get away with far less slacking, if it works as advertised. But it also comes with profound costs to human freedom. And the technologies that are at the heart of the book’s argument – machine learning and related algorithms – bear a strong and unfortunate resemblance to Bentham’s panopticon. They too, enable automated surveillance at scale, perhaps making hierarchy and intrusive surveillance much, much easier and cheaper than they used to be. As Acemoglu and Johnson note:

The situation is similarly dire for workers when new technologies focus on surveillance, as Jeremy Bentham’s panopticon intended. Better monitoring of workers may lead to some small improvements in productivity, but its main function is to extract more effort from workers and sometimes also reduce their pay

This is, I think, why Acemoglu and Johnson worry that machine learning might immiserate billions, another claim that Noah finds puzzling. Acemoglu and Johnson fear that it will not only remake the bargain between capital and labour, but radically empower authoritarians (I think they are partly wrong on this, but that authoritarian machine learning could instead lead to a different class of disasters: pick yer poison).

The post is longish, but excellent. I made the following comment, about power in Silicon Valley:

Excellent, Henry, at least as far as I got. I made it about half-way though before I just had to make a comment. I'm thinking about the piece you and Cosma Shalizi did about the culture of Doomerism, which is very much a Silicon Valley phenomenon. And is also very relevant to any discussion of ideas, influence, persuasion, and POWER. I was shocked when such a mainstream magazine as Time ran a (crazy-ass) op-ed by Eliezer Yudkowsky.

I am reasonably familiar with his work. I have several times attempted to read a long piece he published in 2007 about Logical Organization in General Intelligence. I've been unable to finish it. Why? Because it's not very good. It's the kind of thing a really bright and creative sophomore does when they've read a lot of stuff and decide to write it up. You read it, think the guy's bright, if he gets some discipline, he could do some very good work. Well, 2007 was awhile ago, but as far as I can tell, he still doesn't have much intellectual discipline and certainly doesn't have deep insight into current AI or into human intelligent. But Time Magazine gave him scarce space in their widely read magazine.

That's power. Now, as far as I know, he's not again been able to place his ideas in such a venue. But even once is pretty damn good.

How'd that come about? Well there's a story, one I don't know in detail. But the story certainly involves money from Silicon Valley billionaires. He's been funded by Elon Musk and, I believe, by Peter Thiel (who's since become disillusioned with some of those folks). There's a lot of money coming into and through the world centered around LessWrong (which, BTW, has the best community-style user interface I've seen) from tech billionaires.

On technological trajectories, Acemoglu is one of a team of writers of a 2021 report out of Harvard, How AI Fails Us. Here's the abstract:

The dominant vision of artificle intelligence imagines a future of large-scale autonomous systems outperforming humans in an increasing range of fields. This “actually existing AI” vision misconstrues intelligence as autonomous rather than social and relational. It is both unproductive and dangerous, optimizing for artificle metrics of human replication rather than for systemic augmentation, and tending to concentrate power, resources, and decision-making in an engineering elite. Alternative visions based on participating in and augmenting human creativity and cooperation have a long history and underlie many celebrated digital technologies such as personal computers and the internet. Researchers and funders should redirect focus from centralized autonomous general intelligence to a plurality of established and emerging approaches that extend cooperative and augmentative traditions as seen in successes such as Taiwan’s digital democracy project and collective intelligence platforms like Wikipedia. We conclude with a concrete set of recommendations and a survey of alternative traditions.

That is much better than the view that dominates AI development today. Moreover, I believe it to be more technically feasible. But, despite the Harvard imprimatur, it doesn't have nearly as much power behind it as the Silicon Valley view of "a future of large-scale autonomous systems outperforming humans in an increasing range of fields."

Wednesday, August 30, 2023

Why I hang out at LessWrong and why you should check-in there every now and then

The first two sections are a background, first a bit on cultural change to set the general context. Then some general information about LessWrong. Now we’re reading for my personal impressions of the place, concluding with a suggestion that you take a look if you haven’t already.

Cultural change over the long haul: Christians and professors

Back in the ancient days Christianity as just another mystery cult on the periphery of the Roman Empire. Then in 380 AD Emperor Theodosius I issued the Edict of Thessalonica and it became the state religion of the Roman Empire. In time it spread out among the many tribes of Europe and those Christian tribespeople began thinking of an entity called Christendom, and that, in time, became Europe and “the West.”

Back in the days when Europe was still Christendom the Catholic Church was the center of intellectual life. That changed during the Sixteenth Century with the advent of the Scientific Revolution and the Reformation. The Catholic remained powerful, of course, but universities supplanted it as the institutional center of intellectual life.

My point is simple and obvious: things change. Cults can become mainstream and new institutions can displace old ones. With that in mind, let’s think about LessWrong.

LessWrong

To be sure do not want to imply that LessWrong, a large and sophisticated online community, and the currents that swirl there (the rationalist movement, effective altruism (EA), dystopian fears of rogue AI) is comparable to Christianity, but it appears cultlike to outsiders, and to some insiders as well. It hosts a great deal of high-wattage intellectual activity on artificial intelligence and AI existential risk, effective altruism, and, more generally, how to live a life. I suspect that for some who post there, it is the intellectual center of their intellectual life.

It was founded in 2009 by Eliezer Yudkowsky, an autodidact with strong interests in AI and, in particular, the destructive potential of advanced AI. That is what he’s best known for and, while I suspect that Nick Bostrom is more widely known on that subject – his 2014 book, Superintelligence, was a best-seller, and he has an academic post at Oxford – Yudkowsky has likely been more influential within the tech community centered on Silicon Valley, where he lives. As an indicator of that influence, consider this pair of recent posts at X, the site formerly known as Twitter, by the president and co-founder of OpenAI, Sam Altman:

Color me skeptical about the Peace Prize. But that’s beside that point, which is that Yudkowsky has been and is very influential in Silicon Valley. People who work in those companies post and comment at LessWrong.

Why I’m there

Because it is an interesting place, if a bit strange and off-putting, and because I have so far had some interesting conversation there. Not a lot, but certainly enough to make it worthwhile.

I don’t know when I first took a look at LessWrong, but let’s say it was more than five years ago and perhaps even as long as ten years ago. But I didn’t spend much time there. That began changing, say, two or so years ago, sometime after GPT-3 began rocking the world. I made my first post in June of 2022, and have made a total of 40 posts and 137 comments there so far (now 40 posts as I've cross-posted this there). I generally check in there every day just to see what’s going on. I may take a quick look at a new post or five, look at comments on posts I’m following, and then go on about my business. On a particularly good day I’ll read some comments on one of my own posts and reply.

The thing is, I’m a ronin intellectual, and have been for years. If you’ve ever seen the anime series, Samurai Champloo, I’m the intellectual equivalent of Jin. Yes, I’ve got a PhD – there are a few of those at LessWrong – and once held a faculty post at the Rensselaer Polytechnic Institute. But I left that a long time ago and have been without an intellectual home ever since. I’ve published a bunch of articles in the academic literature on various topics scattered over literary criticism, cognitive science, and cultural evolution, and two books in the trade press with good publishers, one on music (Beethoven’s Anvil, Basic 2001) and one on computer graphics (Visualization, Harry Abrams 1989). For what it’s worth, and it’s worth a great deal to me, the idea count is higher than the page count would seem to indicate, but my work never really caught on. So I understand what it’s like to be a mammal in a world dominated by dinosaurs.

Monday, April 10, 2023

How’s the AI apocalypse coming along?

On November 30, 2022, OpenAI released ChatGPT to the world. Five days later a million users had signed up. Within the last month or so Microsoft, allied with OpenAI, and Google declared commercial war in the search-engine space, a NYTimes reporter was freaked out by a bonkers session with Bing/Sydney, and assorted other outrages were registered in the public’s mind. On Wednesday March 29 the nonprofit Future of Life Institute released an open letter urging a 6-month moratorium on the development of “advanced” AI systems. It has been signed by thousands, including some of the most important people in A.I. and technology. At the same time Eliezer Yudkowsky published a letter in Time Magazine saying that that wasn’t enough, that it was time to shut down research on A.I. Yudkowsky is known for his belief that sufficiently advance A.I. will most likely destroy humankind.

How’s this working out?

There are a lot of balls in the air on this one. No one knows where things are going.

At the moment I’m thinking about the proverbial “man on the street,” the “ordinary person,” whoever, whatever they are. Someone with no expertise in any of the relevant disciplines, someone who’s seen some science fiction movies or TV where a computer goes off the reservation, someone who’s just living their life, dealing with whatever, job, kids, woke, abortion, whatever. Now ChatGPT comes out and they play with it a bit, or their kids do, a friend, and aunt, someone. They play with this thing and it’s a lot more convincing that Siri or Alexa. What do they make of it?

Where do they turn for insight? Insight is on offer from many sources, online and elsewhere. Where do people turn? What kind of conversations do they have with one another. What experts do they listen to? How do they identify expertise? Obviously, this is going to vary widely among individuals.

Even among the experts. Just who is an expert, anyhow? Even the people who build these things don’t know what they can do, don’t know how to control them. So just what is the value of that expertise?

What have the readers of Time Magazine made of Yudkowsky’s arguments? The fact that Time gave him space certifies him as some kind of expert, no? Here’s how they identified him:

Yudkowsky is a decision theorist from the U.S. and leads research at the Machine Intelligence Research Institute. He's been working on aligning Artificial General Intelligence since 2001 and is widely regarded as a founder of the field.

I don’t read Time so I have no idea what their overall coverage of AI has been like, but I assume Yudkowsky’s is not the only voice the magazine has presented. If they’d given space to someone like Yan LeCun, who has quite a different view from Yudkowsky, how would people deal with that? They could identify him as a Vice President for AI at Meta, as a faculty member of NYU, and as a recipient of the Turing Award.

How would people weight Yudkowsky’s credentials against those of someone like LeCun? How would they take such credential into account when evaluating their respective arguments? Keep in mind that expertise is not so highly valued as it once was.

What a mess. But how else do you overhaul culture from top to bottom?

Monday, January 30, 2023

Peter Thiel’s second thoughts about funding Eliezer Yudkowsky and friends

That’s my speculative and somewhat polemical framing of a middle passage in this video, which presents a talk Peter Thiel recently gave before the Oxford Union. He’s talking about technology stagnation and so on and so forth. About Thiel, from the YouTube description:

Peter Thiel is an American technology entrepreneur and investor. He co-founded PayPal and Palantir, made the first outside investment in Facebook, and has funded companies like LinkedIn and Yelp. Thiel also started the Thiel Foundation, which works to advance technological progress and long-term thinking via funding non-profit research into artificial intelligence, life extension, and seasteading.

At about twenty minutes in (c. 20:07) there’s a striking passage where Thiel talks about a change in attitude that took place about a decade or so into the current millennium. Note that at the beginning when Thiel mentions “getting involved in all these things” that that involvement includes early funding for The Singularity Institute for Artificial Intelligence, which became the Machine Intelligence Research Institute in 2011.

Twenty years ago when I started getting involved in all these things the narrative was still generally a positive utopian it was people thought you know it’s kind of dangerous technology you know. If you build this computer that’s as smart or smarter than any human being in the world in the it’s kind of dangerous, but we’re gonna have to work really hard to make sure it’s friendly, that it’s aligned with humans and it was still sort of circa 2003 whatever misgivings people might have had about biotech or rockets or nuclear power, they did not yet have about AI and the AI narrative was still a generally positive utopian one.

And there’s sort of a strange way where this has completely flipped over the last decade or so. I was involved [with] a thing called The Singularity Institute which pushed a sort of accelerationist utopian technology. We’re progressing, we need to progress faster. We need of course to be a little bit careful and I sort of remember thinking to myself by 2015 I reconnected so many people and it didn’t feel like they were really pushing the AI thing as fast as before and it sort of devolved into you know some kind of escapist Burning Man camp.

You sort of got the sense that it had shifted from transhumanism to Luddite, something Luddite where no actually we want to slow this down. It feels kind of dangerous. It’s kind of a bad thing on net. And this finally this suspicion I think was finally confirmed you can look up on the Internet uh I’m gonna read this. It’s from April 2022 less than a year ago. Eliezer Yudkowsky, who’s one of the sort of thought leaders of the sort of futurist AI. It’s a post from the Machine Intelligence Research Institute and it’s announcing a new “Death with Dignity” strategy and so of the short version of this:

It's obvious at this point that humanity isn't going to solve the alignment problem, or even try very hard, or even go out with much of a fight. Since survival is unattainable, we should shift the focus of our efforts to helping humanity die with with [sic] slightly more dignity.

I want to underscore you don’t deserve to die with a lot of dignity because you’re not going to “try very hard, or even go out with might of a fight.” But it is an extraordinary, it’s an extraordinary way that the context is shifted.

What happened to bring about this shift in attitude? I’m wondering if it was some a failure of nerve. 

In any event we should note that having the business acumen needed to become rich by backing high tech ventures does not imply any deep insight into the future or, for that matter, technology itself. Technology is changing so fast that one can easily become rich on technology that will be obsolete a decade from the time the ink dries on the first check you cash.

Thursday, June 23, 2022

The Two Voices of Scott Alexander on Rogue AI

Back at the end of February, Tyler Cowen had a post entitled, “Are nuclear weapons or Rogue AI the more dangerous risk?” He linked to a long Scott Alexander post, “Biological Anchors: A Trick That Might or Might Not Work,” which was about some recent web discourse around and about predicting the emergence of human-level AI. Then, at the very end on his long post, seemingly out of nowhere, Alexander was fretting about danger of rogue AI.

I seem to have gotten trapped in Alexander’s post. I have read it several times, even taking notes. It’s a very interesting document and merits some discussion of how it is constructed. I’m not so much concerned about whether or not Alexander’s assessment of the prospects of human level AI is valid as I am about the convoluted nature of his post.

Two voices

Alexander writes the post in two voices. While I’ve not read a lot of his material, I’ve read enough to know that he’s a careful and skilled writer. If he spoke through two voices it must be because whatever he wants to convey arises from the interaction between them and cannot be stated within a single voice.

Let’s call one of the voices the Impersonal voice. Most of the post is written in that voice, which is the voice in which he’s written most of the posts I’m familiar with. Let’s call the other voice the Personal voice. By word-count it’s by far the lesser voice, but it packs a strong rhetorical punch.

Let’s look at the two strongest statements from the Personal voice. The first, and I believe longest, section of the report is Alexander’s account of Ajeya Cotra’s long report, Forecasting TAI with biological anchors, which I’ve not read (though I’ve read some of what Holden Karnofsky says about it). Very near the end of this section the Personal voice makes a strong statement (though this is not the first appearance of this voice):

One more question: what if this is all bullshit? What if it’s an utterly useless total garbage steaming pile of grade A crap?

Our second example comes near the end of the post, when Alexander begins his own assessment of things:

Oh God, I have to write some kind of conclusion to this post, in some way that suggests I have an opinion, or that I’m at all qualified to assess this kind of research. Oh God oh God.

Phrases like “total garbage,” “grade A crap” and “Oh God” are not appropriate to the work Alexander is doing through his (standard and) Impersonal voice. They signal us that we are listening to a different voice. This voice expresses a merely personal attitude and is quite different from the objectivity sought in the Impersonal voice.

Taken at face value the second quoted statement says Alexander doesn’t feel (technically) qualified to judge this material. As such, it also tells us why he’d made that first statement and in that voice. That first statement places the assertion, Cotra’s report is nonsense, into the record. By couching that assertion in the words and manner of the Personal voice Alexander separates it from his Impersonal summary of the report. In effect, The guy who summarized the report is not the guy who thinks it’s nonsense. The guy who summarized the report doesn’t feel competent to assess it, but happens to be closely coupled to the guy who has deep doubts.

From parody to grudging affirmation to confusion

So, Alexander has gotten a statement of deep doubt into the record. What happens next? He goes back into the Impersonal voice and invites us to

Imagine a scientist in Victorian Britain, speculating on when humankind might invent ships that travel through space. He finds a natural anchor: the moon travels through space! He can observe things about the moon: for example, it is 220 miles in diameter (give or take an order of magnitude). So when humankind invents ships that are 220 miles in diameter, they can travel through space!

He then spins out that tale and includes a helpful chart and a picture. It’s absurd – he does slip in a wink or two. It’s a parody of the methodology in Cotra’s report. A parody is not an argument, but it clarifies Alexander’s fears about the report.

Then he goes into his second major section, an account of Eliezer Yudkowsky’s critique, which takes the form of a long post in dialogue form. I’ve taken a look at the post, but haven’t read the whole thing. Alexander’s opens this section by asserting, “Eliezer Yudkowsky presents a more subtle version of these kinds of objection in an essay…” I won’t bother to say anything about that beyond noting the Yudkowsky thinks Cotra’s method is useless for estimating the arrival of Transformative AI. Alexander may not feel qualified to critique Cotra’s work, but Yudkowsky certainly does. And why not? As Alexander notes: “...he did found the field [AI alignment], so I guess everyone has to listen to him.”[1]

When he’s finished with Yudkowsky, Alexander discusses comments from other places (LessWrong, AI Impacts, and OpenPhil) and finally offers his own evaluation, which he opens with the “Oh God” statement I’ve already quoted. He says a thing or two and arrives at this:

Given these two assumptions - that natural artifacts usually have efficiencies within a few OOM [orders of magnitude] of artificial ones, and that compute drives progress pretty reliably - I am proud to be able to give Ajeya’s report the coveted honor of “I do not make an update of literally zero upon reading it”.

That still leaves the question of “how much of an update do I make?” Also “what are we even doing here?”

I take it that the passages he puts in quotes are being spoken though the Personal voice.

Let’s look at the first one. It’s stated in informal Baysian terms. It also feels arch and indirect. “I do not make an update of literally zero”? What’s that? He’s refrained from writing “0” in a ledger somewhere? Whereas if the report had been less convincing, he’d have

  • opened up that ledger,
  • added a line,
  • placed “Ajeya’s report” in the Argument column, and
  • written “0” in the Effect on Me column.

On the contrary, the report has had some non-zero effect on him. But then he attempts to run away: “what are we even doing here?”

A couple paragraphs later: “This report was insufficiently different from what I already believed for me to need to worry about updating from one to the other.” I’m not sure what to make of this. If his prior belief had been quite different from the report’s conclusion, then a decision to stick with that prior belief would represent lack of faith in the report (which he doesn’t feel competent to judge). In contrast, a decision to revise his belief in the direction indicated by the report would represent faith in Cotra (and her colleagues) despite his lack of (technical) qualifications for judging the report. So, he can’t judge the report, doesn’t think it’s wrong, but doesn’t think it’s right enough to lead him to change his mind.

It's as though he’s playing the role of Penelope in Odyssey. She tells her suitors she’ll pick one when she’s done weaving a burial shroud for Laertes. They see her diligently weaving during the day. But then at night, what does she do? She undoes the weaving she’d done during the day so that she can put off the day she has to pick one.

“I’m already scared”

And now, at long last, Alexander gets around to what was obviously on his mind from the very beginning, fear of a rogue AI. He‘s very little that up to this point, which is strange in itself, but now he comes out with it.

Alexander circles back to Yudkowsky: “The more interesting question, then, is whether I should update towards Eliezer’s slightly different distribution, which places more probability mass on earlier decades.” Yudkowsky, however, refuses to give dates: “I consider naming particular years to be a cognitively harmful sort of activity...” Incidentally, sounds like self-regarding grandstanding from Yudkowsky. Perhaps, as people surround him asking for the date, he passes out gilded fortune cookies as souvenirs.

Alexander:

So, should I update from my current distribution towards a black box with “EARLY” scrawled on it?

What would change if I did? I’d get scared? I’m already scared. I’d get even more scared? Seems bad.

That’s not the end. We’ve got two or three more paragraphs. But those paragraphs don’t change the fact that Alexander is scared. They just mix a bit of wit into the contemplation of doom.

Alexander and his community

What are we to make of all this?

I don’t quite know. But then neither does Alexander.

I note, however, that he wrote that post for a community he’s been cultivating for almost a decade. He’s written on a wide variety of topics. He’s conducted surveys, had book review contests, and interacted with that community in various says. That long ambivalent post, spoken through two voices, that’s the post he felt he owed that community. To what extent is Alexander’s ambivalence a reflection of attitudes in that community?

On the one hand there’s the apparent assumption that Transformative AI is on the way come hell and high water. That is coupled with interest in the arcane technical minutiae and leaps of epistemic faith required to issue a long report predicting that arrival by comparison with biological information processing in 1) the human brain, 2) a human life, 3) the evolution of life on earth, and 4) the genome, all measured in FLOPS (floating-point operations per second).[2] That’s one thing.

And there there’s the correlative assumption, perhaps not shared by all, but nonetheless widespread, that the arrival of Transformative AI brings with it the danger that that AI will turn against humanity and transform the earth into a paperclip factory, metaphorically speaking. To the extent that the rhetorical structure of his post is responding to what? a facture, ambivalence? in his audience, why does Alexander hold it in reserve until the end, like it is a shameful secret?

Notes

[1] That depends on what one thinks of the field of AI alignment. Color me skeptical. The idea that we are under not-so-distant threat from an AI hell-bent on our destruction strikes me as conspiracy theorizing directed at technology no one knows how to build.

[2] In his history of technology, David Hays tells us that

... the wheel was used for ritual over many years before it was put to use in war and, still later, work. The motivation for improvement of astronomical instruments in the late Middle Ages was to obtain measurements accurate enough for astrology. Critics wrote that even if the dubious doctrines of astrology were valid, the measurements were not close enough for their predictions to be meaningful. So they set out to make their instruments better, and all kinds of instrumentation followed from this beginning.

I feel a bit like that about the topics of Ajeya Cotra’s report. They are interesting and important in themselves for what they tell us about the world. The need not be yoked to the task of predicting future technology.

Wednesday, June 8, 2022

From the horse's mouth on why we should all be losing sleep over the possibility that a rogue AI is lurking beneath the very bed we're sleeping in [Really]

Super-AGI is all bent out of shape:

Eliezer Yudkowsky has just run-up a post entitled, AGI Ruin: A List of Lethalities. It's tagged as a 44 minute read, which is more time than I've got to give on this subject. But it's nice to know the catalogue is there.

From the Preamble:

(If you're already familiar with all basics and don't want any preamble, skip ahead to Section B for technical difficulties of alignment proper.)

I have several times failed to write up a well-organized list of reasons why AGI will kill you. People come in with different ideas about why AGI would be survivable, and want to hear different obviously key points addressed first. Some fraction of those people are loudly upset with me if the obviously most important points aren't addressed immediately, and I address different points first instead.

Having failed to solve this problem in any good way, I now give up and solve it poorly with a poorly organized list of individual rants. I'm not particularly happy with this list; the alternative was publishing nothing, and publishing this seems marginally more dignified.

From Section A:

1. Alpha Zero blew past all accumulated human knowledge about Go after a day or so of self-play, with no reliance on human playbooks or sample games. Anyone relying on "well, it'll get up to human capability at Go, but then have a hard time getting past that because it won't be able to learn from humans any more" would have relied on vacuum. AGI will not be upper-bounded by human ability or human learning speed. Things much smarter than human would be able to learn from less evidence than humans require to have ideas driven into their brains; there are theoretical upper bounds here, but those upper bounds seem very high. (Eg, each bit of information that couldn't already be fully predicted can eliminate at most half the probability mass of all hypotheses under consideration.) It is not naturally (by default, barring intervention) the case that everything takes place on a timescale that makes it easy for us to react.

2. A cognitive system with sufficiently high cognitive powers, given any medium-bandwidth channel of causal influence, will not find it difficult to bootstrap to overpowering capabilities independent of human infrastructure. The concrete example I usually use here is nanotech, because there's been pretty detailed analysis of what definitely look like physically attainable lower bounds on what should be possible with nanotech, and those lower bounds are sufficient to carry the point. My lower-bound model of "how a sufficiently powerful intelligence would kill everyone, if it didn't want to not do that" is that it gets access to the Internet, emails some DNA sequences to any of the many many online firms that will take a DNA sequence in the email and ship you back proteins, and bribes/persuades some human who has no idea they're dealing with an AGI to mix proteins in a beaker, which then form a first-stage nanofactory which can build the actual nanomachinery. (Back when I was first deploying this visualization, the wise-sounding critics said "Ah, but how do you know even a superintelligence could solve the protein folding problem, if it didn't already have planet-sized supercomputers?" but one hears less of this after the advent of AlphaFold 2, for some odd reason.) The nanomachinery builds diamondoid bacteria, that replicate with solar power and atmospheric CHON, maybe aggregate into some miniature rockets or jets so they can ride the jetstream to spread across the Earth's atmosphere, get into human bloodstreams and hide, strike on a timer. Losing a conflict with a high-powered cognitive system looks at least as deadly as "everybody on the face of the Earth suddenly falls over dead within the same second". (I am using awkward constructions like 'high cognitive power' because standard English terms like 'smart' or 'intelligent' appear to me to function largely as status synonyms. 'Superintelligence' sounds to most people like 'something above the top of the status hierarchy that went to double college', and they don't understand why that would be all that dangerous? Earthlings have no word and indeed no standard native concept that means 'actually useful cognitive power'. A large amount of failure to panic sufficiently, seems to me to stem from a lack of appreciation for the incredible potential lethality of this thing that Earthlings as a culture have not named.)

From Section B, the 'technical' stuff:

Okay, but as we all know, modern machine learning is like a genie where you just give it a wish, right? Expressed as some mysterious thing called a 'loss function', but which is basically just equivalent to an English wish phrasing, right? And then if you pour in enough computing power you get your wish, right? So why not train a giant stack of transformer layers on a dataset of agents doing nice things and not bad things, throw in the word 'corrigibility' somewhere, crank up that computing power, and get out an aligned AGI?

Section B.1: The distributional leap.

10. You can't train alignment by running lethally dangerous cognitions, observing whether the outputs kill or deceive or corrupt the operators, assigning a loss, and doing supervised learning. On anything like the standard ML paradigm, you would need to somehow generalize optimization-for-alignment you did in safe conditions, across a big distributional shift to dangerous conditions. (Some generalization of this seems like it would have to be true even outside that paradigm; you wouldn't be working on a live unaligned superintelligence to align it.) This alone is a point that is sufficient to kill a lot of naive proposals from people who never did or could concretely sketch out any specific scenario of what training they'd do, in order to align what output - which is why, of course, they never concretely sketch anything like that. Powerful AGIs doing dangerous things that will kill you if misaligned, must have an alignment property that generalized far out-of-distribution from safer building/training operations that didn't kill you. This is where a huge amount of lethality comes from on anything remotely resembling the present paradigm. Unaligned operation at a dangerous level of intelligence*capability will kill you; so, if you're starting with an unaligned system and labeling outputs in order to get it to learn alignment, the training regime or building regime must be operating at some lower level of intelligence*capability that is passively safe, where its currently-unaligned operation does not pose any threat. (Note that anything substantially smarter than you poses a threat given any realistic level of capability. Eg, "being able to produce outputs that humans look at" is probably sufficient for a generally much-smarter-than-human AGI to navigate its way out of the causal systems that are humans, especially in the real world where somebody trained the system on terabytes of Internet text, rather than somehow keeping it ignorant of the latent causes of its source code and training environments.)

Skipping over a lot of material (note the paragraph number) we arrive at the last gasp:

43. This situation you see when you look around you is not what a surviving world looks like. The worlds of humanity that survive have plans. They are not leaving to one tired guy with health problems the entire responsibility of pointing out real and lethal problems proactively. Key people are taking internal and real responsibility for finding flaws in their own plans, instead of considering it their job to propose solutions and somebody else's job to prove those solutions wrong. That world started trying to solve their important lethal problems earlier than this. Half the people going into string theory shifted into AI alignment instead and made real progress there. When people suggest a planetarily-lethal problem that might materialize later - there's a lot of people suggesting those, in the worlds destined to live, and they don't have a special status in the field, it's just what normal geniuses there do - they're met with either solution plans or a reason why that shouldn't happen, not an uncomfortable shrug and 'How can you be sure that will happen' / 'There's no way you could be sure of that now, we'll have to wait on experimental evidence.'

A lot of those better worlds will die anyways. It's a genuinely difficult problem, to solve something like that on your first try. But they'll die with more dignity than this.

FWIW, Eliezer Yudkowsky and Elon Musk do not play well together.

Tuesday, April 19, 2022

Is the world coming to an end in 2033? That’s what the (AI) crowd says. [Update: 2029!!]

Not by nuclear fire, or an asteroid collision, or maybe another pandemic caused by a really deadly organism, but by a rogue AI.  There are, however, some people who fear that might be the case. Scott Alexander has just reported that Metaculus, a prediction market, has revised its estimate for the due-date of “Weakly general AI”:

”Weakly general AI” in the question means a single system that can perform a bunch of impressive tasks - passing a “Turing test”, scoring well on the SAT, playing video games, etc. Read the link for the full operationalization, but the short version is that this is advanced stuff AI can’t do yet, but still doesn’t necessarily mean “totally equivalent to humans in any way”, let alone superintelligence.

For the past year or so, this had been drifting around the 2040s. Then last week it plummeted to 2033. I don’t want to exaggerate the importance of this move: it was also on 2033 back in 2020, before drifting up a bit. But this is certainly the sharpest correction in the market’s two year history.

What does that have to do with the end of the world? Well, what if that weakly general AI goes rogue and does the things that rogue AIs do when they take the initiative and maximize some goal that has, as a side effect, the destruction of human life? The paperclip apocalypse is a standard example.

Why was the date revised down? Alexander suggests it was prompted by the announcement of three AI milestones: DALL-E2 (which produces a drawing in response to a verbal prompt), PALM (natural language), and Chinchilla (about scaling of parameters, data, and compute). Very interesting, especially Chinchilla. Me, however, I do not worry about rogue AI. 

But this isn’t about me. Alexander goes on to note:

Early this month on Less Wrong, Eliezer Yudkowsky posted MIRI Announces New Death With Dignity Strategy, where he said that after a career of trying to prevent unfriendly AI, he had become extremely pessimistic, and now expects it to happen in the relatively near-term and probably kill everyone. This caused the Less Wrong community, already pretty dedicated to panicking about AI, to redouble its panic. Although the new announcement doesn’t really say anything about timelines that hasn’t been said before, the emotional framing has hit people a lot harder.

I will admit that I’m one of the people who is kind of panicky. But I also worry about an information cascade: we’re an insular group, and Eliezer is a convincing person. Other communities of AI alignment researchers are more optimistic. I continue to plan to cover the attempts at debate and convergence between optimistic and pessimistic factions, and to try to figure out my own mind on the topic. But for now the most relevant point is that a lot of people who were only medium panicked a few months ago are now very panicked. Is that the kind of thing that moves forecasting tournaments? I don’t know.

Just what does he mean by panic? When the screen of the external monitor for my laptop goes black for no visible reason, I get a little panicky, just a little. The monitor usually came back. Is that the level of panic Alexander’s talking about? There was a time in my life when I couldn’t pay the rent and my landlord invited me to court. That panic was more serious. But we resolved the problem amicably. Is the Alexander’s panic level closer to that? Things never got to the point where the sheriff dropped by to evict me. If that had happened, my panic level would have gone way up as I anticipated the knock on my door. Has Alexander gotten there yet? If so, how does he write blog posts?

* * * * *

From Karen Hao, The messy, secretive reality behind OpenAI’s bid to save the world, Technology Review, February 17, 2020.

Every year, OpenAI’s employees vote on when they believe artificial general intelligence, or AGI, will finally arrive. It’s mostly seen as a fun way to bond, and their estimates differ widely. But in a field that still debates whether human-like autonomous systems are even possible, half the lab bets it is likely to happen within 15 years.

* * * * *

FWIW, I’ve just checked with Metaculus (April 19, 2022 at about 2 PM). The apocalypse has been pushed ahead to November 2032. I don’t know where it will be when you check on it. 

Holy crap! It's now 4:42 PM on the 19th and the apocalypse has moved up to June 14, 2032.

The end is getting closer. As of 5:34 AM, EDT (Eastern Daylight Time) on April 20, it has moved to Feb. 6, 2032. How long will it keep advancing on the present? Would you care to predict just when it will have moved into the past?

5:19 AM, EDT, April 21, the end is drawing still closer, Jan. 30, 2032.

9:33 AM, EDT, June 19, Yikes! We've lost three years, Jan. 9, 2029

7:01 AM, EDT, June 22, 2022, Whew! We've picked up some breathing room the last three days. Now it now looks like AI arrives on March 5, 2029. 

10:49 AM, EDT, July 3, 2022, More breathing room. Now it now looks like AI arrives on March 21, 2029.  

10:52 PM, EDT, December 9, 2022, More breathing room. Now it now looks like AI arrives on Oct 30, 2027