Showing posts with label AI. Show all posts
Showing posts with label AI. Show all posts

Saturday, August 1, 2026

Sabine: “You can brute-force counterexamples by just trying a lot of guesses quickly...”

Ethan Mollick: “AI has blurred lines between jobs.”

Friday, July 31, 2026

In a study involving two cases, agents exhibited 5 failure modes in open-ended research

Abstract of the article linked in the tweet:

Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper’s original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today’s agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.

Thursday, July 30, 2026

An AI blizzard is headed our way [Yikes!]

Adam Satariano, Paul Mozur, Jacqueline Gu and Cade Metz, The Impending, Inescapable Deluge of A.I. NYTimes, July 30, 2026.

Boom....

From the American Midwest to the Persian Gulf, hundreds of major data centers now under construction will be turned on in the coming years. They are set to deliver an avalanche of computing power to develop and run A.I. that has no equal in the history of the technology industry, with breakthroughs that once felt revolutionary likely to become increasingly routine.

Behind each leap in A.I. are corresponding jumps in computing power. Today, there are about 20 million A.I. chips crammed into the data centers that underpin the technology’s growing abilities and usage worldwide, according to the research firm Epoch AI. [...]

In size and ambition, this moment compares to the building of the railroads in the 1800s, President Franklin D. Roosevelt’s New Deal in the 1930s, and the Manhattan Project to create an atomic weapon in the 1940s, technologists said.

“This is the largest scale infrastructure build-out in the history of humanity,” said Rob Wachen, a co-founder of the microchip firm Etched, which has raised more than $1 billion to meet the growing demand for A.I. components.

Peter DeSantis, who leads foundational A.I. models at Amazon — which provides computing power to the A.I. firms Anthropic, OpenAI and others — said the Seattle company has doubled its computing capacity since 2022 and would double it again by next year. “It’s hard to get your mind around the scale,” he said. [...]

Confidence in the Scaling Laws has led A.I. leaders to make ever bolder predictions. Dario Amodei, the chief executive of Anthropic, has said that if these laws hold for another year or two, A.I. will be able to perform huge amounts of white-collar work. Demis Hassabis, the head of Google’s A.I. lab DeepMind, wrote recently that A.I. could usher in “10x of the Industrial Revolution at 10x the speed.”

Bust?

Economists and investors have raised concerns that tech firms are spending faster than they can profit from A.I. Past infrastructure booms have been followed by downturns before the benefits of the technology were realized. The railroad boom in the 1800s, electrification in the 1920s and the dot-com bubble in the late 1990s were punctuated by economic recessions and a stock market crash as companies that overspent went out of business.

“Each time you’ve had a technological revolution, this kind of bubble bursting happened,” said Philippe Aghion, who won the Nobel in economic science in 2025 for research on innovation-driven economic growth. “A.I. is like the fourth industrial revolution and it has this aspect to it that generates a bubble.”

Moreover:

With more computing power coming online, geopolitical divisions are only set to widen.

The United States, home to about 5,500 data centers, about 10 times the next closest country, is far ahead of the rest of the world, including China. U.S. companies like Amazon, Google, Microsoft and Meta control about 80 percent of global computing power that drives A.I., according to Epoch AI. Google alone is believed to have four times as many A.I. chips as all of China’s companies, which are racing to catch up by developing new semiconductors and A.I. infrastructure of their own.

The race is on...

The article then goes on to discuss the huge increase in total chip count distributed over a growing collection of ever larger data centers under construction or proposed. We're in an AI arms race. The article then goes on to discuss the race with China, currently a fairly distant second (by an order of magnitude) to the US in chip count and gigawatts.

Mr. Nanos said the U.S. data center lead over China would likely grow over the next four to five years, before China’s domestic chips are produced at scale. After that, China should begin closing the gap.

“The advantage will run out,” he said.

The U.S.-China race threatens to leave the rest of the world behind. France, Germany and other nations are trying to encourage data center construction across the European Union, which has 5 percent of global A.I. computing power, according to a report by A.I. developers and policy experts in the region. Europe has been hampered by electricity and land access, permitting and financing.

Changes in the labor market:

“There’s going to be millions of jobs destroyed, millions of jobs created,” said Erik Brynjolfsson, an economist who is the director of Stanford University’s Digital Economy Lab. “That’s going to be very difficult. Even if new jobs are created, they’re not the same jobs.”

The last three paragraphs are about recursive self-improvement.

Wednesday, July 29, 2026

Why Adam Hunt has “flipped from being bullish to being bearish about AI.”

Three important agents [think about it]

Tuesday, July 28, 2026

Framing my discussion of The God Test, Part 1: Rorschach, reason, and whaling – [GT-3]

I’ve got to bite the bullet: I’m just going to have to go through a bunch of (preliminary) stuff before I can really engage with The God Test. My current target is to be in a position to publish a proper review of the book in 3 Quarks Daily for the week of August 9.

Rorschach Recap

I want start by recapping the Rorschach metaphor I introduced in the previous post, More on how I’m approaching The God Test – Rorschach! [GT-2]. What I like about it is that has a shape, there’s something there, but it’s not clear what. So we have little choice but to project onto it in order to (begin to) make sense of it.

First: It is a new kind of thing, an artifact we can converse with in an open-ended and natural way. The steam engine was the same kind of thing. It was an inanimate object that moved over the surface of the earth under its own power. Previously only animals (& humans as animals) had that power. So it becomes an iron horse. Just what are AIs? What’s their nature? That’s one thing.

Second: How it works is opaque. We know how to create large language models (LLMs), but we don’t know how they work. That’s new. We may not have understood the deep physics of the steam engine, but we certainly knew how they worked.

Third: We don’t know what they portend for the future. To some extent this is a function of the first two: How can we, how should we, interact. But it is also a function of the future, which is undetermined. We just don’t know.

Rhetorical force over reason

This is an argument I made in the first working paper I published after the release of ChatGPT in November of 2022: ChatGPT intimates a tantalizing future; its core LLM is organized on multiple levels; and it has broken the idea of thinking (February 6, 2023).

What do I mean by that, has broken the idea of thinking? Prior to ChatGPT it was obvious that humans could think and computers could not. [Yeah, I know, there’s Deep Blue defeating Kasparov in chess. That just changes the dates, not the argument.] The difference in performance was so obvious that the fact that we don’t really know how humans think wasn’t much of an issue. Now it is. Sure, we can still say that we can think and the AI’s can’t, but that’s just a line and without good explanations on both sides of the line, it seems a bit arbitrary, if not desperate.

I made a particular argument about Searles’ (in)famous Chinese Room thought experiment. I read it when it was first published in Brain and Behavioral Science in 1980. I wasn’t impressed. Why not? He didn’t say anything about any of the techniques used in AI or computational linguistics (CL). How could anyone possibly take that seriously?

He talked about intention, that’s how. Meaning requires intention and only living things can have intention, a remark he made at the end of the article. Without intention the most you get is syntax, but no meaning. Searle could get away with that because, in the first place, the concept of intention has a long history within philosophy – it has a subtle meaning, but that can wait for a later post – and so philosophers, his main audience, were comfortable with it. That’s one thing.

But there’s something more important, something that we can see only in retrospect, and that’s the simple fact computers very obviously could not translate from Chinese into English or into any other language. That difference carried tremendous weight. We don’t have a subtle behavioral difference between computers and humans that requires a subtle and sophisticated argument. To a first approximation, almost any argument would do. As far as I was concerned, “intention” was just a fancy word for something we don’t understand. But that’s not an argument anyone needs to take seriously. The behavioral distance is quite sufficient to carry the argument for those who insist that computers can’t and will never be able to think like humans.

Now the behavioral evidence has changed. Sure, differences remain, but the evidence is shifting. The old arguments remain and those who believed them still do so, but it’s getting harder. The need for explicit arguments grounded in explicit accounts of computers, and also brains, is growing.

Whaling and expertise

What’s an expert in machine learning and LLMs actually expert in? For some time now I’ve been arguing that investing in AI is like investing in a whaling venture where the captain and crew of the ship know all there is to know about the ship and how to handle it but know little or nothing about whales and their behavior and about navigating around the Cape Horn and in the South Pacific, where the whales live. What are the chances of that voyage being successful? Not very good.

The people who have created the current AI technology are like that captain and crew. The know how to sail the ship. But they don’t know much about language or cognition. They don’t actually know much about the human mind. Here my point is not about the fact that the models are opaque, but that human language and cognition are highly structured and they don’t believe that one needs to know (much of) anything about that not only to build AI but to make confident prediction about the future of AI.

Gary Marcus, Subbarao Kambhampati, and others have been consistently arguing that, yes, the current technology is remarkable, but we are going to have to adopt classical symbolic techniques if we are to fully develop the technology so that we have accurate and safe systems. Marcus is arguing from his knowledge of human language and cognition. As far as I can tell, Wright doesn’t take that seriously. I know that he had Marcus on his NonZero podcast, and that he lists Marcus in his acknowledgements, but that he doesn’t discuss Marcus’s ideas. I conclude that he doesn’t take that line of argument seriously.

That’s a mistake, but this is not the place to make my own arguments on this issue. My point is simply that expertise in AI is no generally construed to encompass knowledge of, expertise in, human cognition and language. I can’t see how that is going to work out well in the future.

[Note: If you’re curious about my views, on this subject, read the article linked in the first paragraph of this section. My views all over the place here at New Savanna, particularly around the work of the mathematician Miriam Yevick. Also, check out the experimental work I’ve done with LLMs.]

The odd destruction of books en masse by AI companies [Homo economicus on a binge]

Monday, July 27, 2026

What I did last week: aesthetics, economics, Rorschach analogy for AI, default images, and “leveling”

I did some satisfying work last week. Here’s a quick rundown. I’m listing the posts in the order I wrote them.

Visual Aesthetics

A case of visual aesthetics: Why is the monochrome image superior to the color image?

The issue, black & white vs. color, has been and I suppose remains central to photography, and I deal with it there, a bit. But that’s not what I’m doing here. This is about the conversion of a particular ChatGPT image from color to black & white. It was a fun post to assemble and to think about. I like the suite of images.

Rank 5 Economics?

Beyond Marginalism: What’s Next? [MR #12]

This is my last word – save for an introduction I’ll write in a week or three, who knows? – on the fourth and final chapter of Cowen’s monograph on marginalism. This is where he tosses up some examples of leading edge work in economics, noting that it’s drifting away from marginalism into complex high-dimensional models created through machine learning. His examples come from finance. The new models yield better predictions.

I focus on one model that has 360,000 parameters and end up making (speculative) sense out of what’s going on. I suggest that those parameters are picking up the effects of Keynes’ “animal spirits” as expressed in the gossip and stories of Schiller’s narrative economics. I further suggest that we can test this by comparing the output of a classical model with that from a high-parameter machine learning model. The divergence should be highest with those stocks otherwise identified as meme stocks.

The prospect of empirical investigation into animal spirits in asset pricing [the fate of marginalism]

Here I take my speculations about how to test these high parameter models and present them to Marge, the AI associated with Cowen’s book. Marge approves.

Rorschach test for AI

More on how I’m approaching The God Test – Rorschach! [GT-2]

I came up with the Rorschach blot as analogy for the kind of challenge AI presents to us, to our understanding of AI and of the future. The idea is that the blot does have a form, albeit a complex one that’s not very legible. Hence our commentary on it (that is, on AI) tells as much about us as about AI. I’ll be developing this further in a later post.

Prototype Image in ChatGPT

This is a new working paper that opens up a whole new line of investigation. This was a fun piece of work. Writing it up took way longer than actually generating the images.

A prototypical image in ChatGPT 5.6: An informal pilot study

Abstract: Previous work has found strong default preferences in stories generated by large language models from minimally specified prompts. To determine whether a similar effect appears in image generation, I asked ChatGPT 5.6 to create 17 images in separate chats using four prompts that specified either no subject matter or only a rendering medium: “Create an image,” “Create a drawing,” “Create a painting,” and “Create a water color painting.” Sixteen of the 17 outputs depicted closely related landscapes containing mountains, trees, sky, and water; the remaining image depicted a lighthouse. The images also shared a calm, picturesque mood and contained no human figures, although several included signs of human habitation. Because the internal prompt passed to the image generator was not available, the study cannot determine whether these defaults arise primarily in the language model, the image generator, or their interaction. The results are exploratory but suggest that severely underspecified image prompts may reveal stable default preferences in the integrated ChatGPT image-generation system.

The “leveling” of knowledge in the compressed form of LLMs

NYTimes: AI needs human supervision in order to complete an entire job.

This is something I’ve been thinking about off and on for a while, but this is my first explicit framing of the issue. The idea is that once ideas or set of ideas has been expressed in writing and those documents then consumed into an LLM, all ideas function the same within/through/for the model. In that post I’m comparing a study of using AI to perform routine office processes (from NYTimes) with the use of AI to perform a complex set of tasks in drug development, in effect, high-school level capability with Ph.D. level capability. They’re the same to the LLM.

I need to think about this some more. It seems to me what’s nowhere present in the LLM is the kind of procedural knowledge necessary to learn tasks at whatever level. That simply isn’t presented in the written products of that knowledge (not even in written procedures).

Sunday, July 26, 2026

NYTimes: AI needs human supervision in order to complete an entire job.

From the NYTimes article linked in the tweet:

We gave an A.I. tool full access to a laptop with pre-configured apps and sought to answer a simple question: Can artificial intelligence do an office job?

Some corporate executives seem to believe it can. More than 200 tech companies have cut roughly 120,000 jobs this year, according to Layoffs.fyi, an industry tracking site; Meta, Oracle and others have all recently made substantial cuts to their work forces, citing A.I. as the driving force; and after laying off about 1,100 employees, the chief executive of Cloudflare said recently that he expected A.I. to replace workers in middle management, finance and marketing.

tweIn our experiment, we deployed A.I. “agents” to act as office workers and found that they were capable of performing some of the tasks we assigned, but not all of them. The agents, which can act autonomously and make decisions based on detailed instructions, excelled at problems they could solve by writing computer programs. But they struggled with understanding the nuances of human language and at navigating user interfaces like the Chrome web browser.

The article then has a series of nice quasi-interactive displays illustrating agent performance on three tasks. The displays include screen shots of various messages and documents.

About the tasks:

This task, and the others we assigned to the A.I., were adapted from papers and benchmarking tools published recently by researchers at Carnegie Mellon University and OpenAI. The researchers designed the benchmarks to test the performance of various models — like OpenAI’s GPT, Google’s Gemini and Anthropic’s Claude — in real-world environments, and compare them with one another.

General conclusion:

The results of our experiment roughly matched what researchers and companies have found as they have tested and used artificial intelligence tools. Scale AI, an A.I. training company, recently tested agents on real freelance projects, and the best-scoring model produced client-ready work only about 16 percent of the time.

While A.I. can excel regularly at complex tasks, it can be unreliable when put in charge of an entire job. It can certainly add value to certain areas of the work force, but for now, A.I. still needs a human boss.

* * * * *

Comment: Around the corner my colleague, Ash Jogalekar, has tweets like this one:

So here's a great example of where we are with agentic AI: Instead of just being an assistant, it's behaving more like a collaborator and creative scientist.

In a recent project, I gave the system a molecular design problem typical of the problems we encounter in chemistry. Two similar molecules were giving very different results.

He then runs through an account of what his AI collaborator did, concluding:

I think we have crossed the Rubicon. Agentic AI now no longer just processes tasks and automates workflows blindingly fast, but it can generate hypotheses, test them, test counter-hypotheses and go back and forth and course-correct if necessary, all with minimal to no human intervention. It's now embodying the general scientific method.

[I've copied another one of Ash's tweets to this post, A scientist reflects on what AI has done for him.]

What’s interesting to me, and very revealing, is that a complex set of tasks in scientific investigation seems to be on a level with routine office tasks, as though one were no more complex than the other. But humans require years of college education in order to perform the former while the latter requires no more than a high school education, if that. It seems that once they’ve been learned and compiled, all tasks or sets of tasks are on the same “level” in the brain. The educational prerequisites required to do such tasks for the first time or three get “compressed out” through repetition. Since AIs are trained on written records of what humans have said and done, they don’t have to go through the ordinary learning process. The compression has already taken place and is present in the documents on which they are trained.

Saturday, July 25, 2026

Standard-issue doom scenarios were invented before LLMs and are made obsolete by them.

Friday, July 24, 2026

More on how I’m approaching The God Test – Rorschach! [GT-2]

I’m still trying to figure out how to approach Robert Wright’s The God Test.

How LLMs work

In my previous post – How will I handle The God Test? [GT-1] – I expressed misgivings about how Wright explains the technology. Those misgivings haven’t disappeared. However, Bert Idem has published a useful review at Finite Ape in which he addresses some of those issues in detail. Specifically:

Now, about the history of AI, the story he tells is actually great and it is certainly more than what most non-technical people know about LLMs. However, there are three places where I think the framing goes wrong or at least leaves out context that matters:

  • LLMs did not discover the meanings of words on their own by accident. They were designed on top of ideas from older models that were specifically trained to learn the meanings of words.
  • Similarly, computer vision models didn’t find out how to “view” an image like we do. Instead, the classical CNN models were heavily inspired by biological vision itself.
  • LLM weight training is simply gradient-based optimization and the process has nothing to do with evolution. Of course, we can make a parallel between any kind of change and evolution but then, in that sense, everything evolves and it is not useful to talk about evolution.

I agree with Idem on those three issues, not so sure about the history part. While I may return to some of these issues later on, this will serve as a place holder.

A Rorschach test

There’s something else going on, but I’m not quite sure how to conceptualize it. It seems to me that AI is functioning something like a Rorschach test which, as you may know, is a psychological instrument intended to elicit (potentially) revealing responses from a person. It’s a projective test.

A person is shown a series of ink blot images, like this one (generated by ChatGPT):

They are asked what that they see in the image, what it means to them. Since the image is, though not formless, its form is not that of any specific animal, vegetable, mineral, person, or anything else. It’s just a blot. Whatever the person says about the blot, however they interpret it, that must reveal something about them. Why? Because whatever they see in the blot, isn’t really there.

Broadly and crudely speaking, AI has become something of a cultural Rorschach test.

Understanding computers & LLMs

Until ChatGPT was released in late November of 2022, most people knew very little to nothing about AI. Oh, they may have seen “intelligent” computers and robots in science fiction movies, but that’s science fiction and only tangentially related to AI considered as a line of research dating back to the 1950s. Many people would have heard about IBM’s Deep Blue beating Gary Kasparov in chess in 1997 and then, in 2011, when IBM’s Watson beat Ken Jennings and Brad Rutter in Jeopardy. Those were real AI systems, standing on research extending back decades, but as far as most people were concerned, they were one-off PR stunts. Just how they worked, who cares? They’re computers, and computers are magic, no?

As far as most of us are concerned, computers are magic. Somewhere “out there” someone knows how these things work, but we don’t need to know any of that. It’s complicated, but computers do what they’re programmed to do, no? Yes, but not LLMs.

And that’s the tricky part. LLMs, large language models, aren’t like other computer systems. They aren’t programmed in the way that word processors, photo editors, or phones are programmed. LLMs aren’t programmed at all, not in the ordinary sense of programming – something I may or may not get into in a later post. As far as most users are concerned, how ChatGPT, or Claude, or Gemini work, that’s no more interesting than how a word processor works. It just does. It’s more magic.

But if you have a strong philosophical streak, if you are interested in the mind, in technology, in the technology in the future, then you may not be content with writing LLMs off as just another kind of magic. You want to know what’s going on inside, 1) because you want to know (curiosity), and 2) because you want to know how the technology is going to develop in the future (engagement). Now things get interesting? Why? Because even the people who have created the technology don’t know how it works.

Oh, they know how the transformer program works. That’s the program that creates the language model. It creates the model by performing a (certain kind of) statistical analysis of a huge body of texts, effectively the entire internet. When a person prompts the model with some statement, the model responds by a statement of its own. No one know just how the model does that. That’s a mystery, a deep black hole in the technology ecosystem.

AI as a Rorschach test

If you aren’t content to believe in magic, then you have to come up with something to fill that black hole in your, in our, understanding. This is where the Rorschach aspect of AI reveals itself. To a first approximation, what each of us uses to paper over that black hole has as much to do with ourselves as with AI.

Why do I say, “To a first approximation”? It’s a rhetorical device to get things started. It puts us all in the same boat, despite our different backgrounds. However, whatever LLMs are, they are not magic. It is possible, in principle, to construct a technical account of what they’re up to, but no one knows how to do that, yet. Not even the people in the AI labs who create these beasts.

Those of us who are trying to figure out how LLMs work have widely varying backgrounds. In particular, we have widely varied technical backgrounds and we bring those backgrounds to bear when we think about what LLMs are doing. Those backgrounds influence how we interpret the AI-blot. Wright is a journalist with a wide range of interests, including politics, international affairs, evolutionary psychology, cultural evolution, and Buddhism. As far as I can tell there isn’t much there that’s directly relevant to understanding the mechanisms of LLMs, but he’s done a lot of reading and talked with a lot of experts to fill in the gaps.

My background is quite different. While I happen to know quite a bit about cultural evolution, cognitive psychology, neuroscience, and various other things, my background in computational semantics puts me much closer to LLMs than Wright’s knowledge of evolutionary psychology puts him. Still, like him, I’ve done a lot of reading and talked with experts. In particular, I’ve been collaborating with Ramesh Viswanathan for the last three years. He’s an expert in machine vision Goethe University Frankfurt. He’s got a background in mathematics and AI that I don’t have. Still, there are things he doesn’t know, things he’s trying to figure out. 

We are all making stuff up.

To some extent, then, AI is a Rorschach test about how beliefs about the human mind, and human nature. When we try to figure out how the LLM is working we’re also, if only implicitly, trying to figure out how we work, internally, as well. The whole discourse about AL alignment is as much a discourse about us as it is about AI. 

The Future

And even if we knew much more about how LLMs work internally we still wouldn’t know how the technology will develop in the future. We? You, me, Robert Wright, Ramesh Viswanathan, Gary Marcus, Tyler Cowen, Geoffrey Hinton, Sam Altman, Dario Amodei, Nick Bostrom, Eliezer Yudkowsky, all of us who are trying to figure it out. We don’t know what will happen. That’s where we’re projecting like mad. We’re hallucinating, to borrow a term from AI-speak. 

Thus AI is also a Rorschach test for our visions of the future. When we imagine the future of AI, we’re also imagining our future. Like the two sides of a coin, the two cannot be separated. 

The tricky part, the important part, is that the future development of AI is not predestined. It depends on the choices we make, now and in the near future. We can easily and often do imagine things that will not be possible because that’s just not how the world works. But the laws of how the world works are open to a wide range of possibilities. The boundary between the possible and the impossible is fuzzy at best.

Where, and how, does Wright draw that boundary? Perhaps that’s what I’ll be trying to figure out.

More later.

On the OpenAI/Hugging Face incident

H/t Tyler Cowen. My reply to Cowen's post:

FWIW, me, #3 – Meh. I've got better things to do than to go down this rabbit hole.

Thursday, July 23, 2026

A scientist reflects on what AI has done for him – “a Rubicon has been crossed”

Here’s the full content of a tweet by Ash Jogalekar:

I came to the present AI revolution not as a credulous enthusiast, but as someone deeply skeptical of new technologies in science. For twenty-five years I have seen too many of them arrive surrounded by extravagant claims before settling into a useful but much more modest place in the scientific toolkit.

What has astonished me is that the latest agentic systems appear to represent a qualitative change. I have now run upward of two hundred scenarios and AI for science workflows, each one navigating a complex, multistep scientific protocol across diverse fields of chemistry, biology and materials science, and I think I have enough data now to make a credible judgment. Over just a few months I have seen these models and algorithms leapfrog over increasing levels of difficulty, starting almost as a toddler and turning into an adult interlocutor. Every day, something moves my baseline for what they can do. They autonomously install, run, and debug dozens of computational tools; clean and structure data; parameterize molecules; launch calculations; and manage complicated, multistage investigations. But as it turned out, that was just the beginning.

More fascinatingly, they increasingly display recognizable scientific judgment: proposing positive and negative controls, discovering that a method does not work, generating competing hypotheses, systematically eliminating them, changing direction when evidence contradicts an initially promising idea, searching an entire target space, constructing unexpectedly sophisticated models, and finding useful analogies across distant fields. This is no longer merely workflow automation or doing the same science faster. It is beginning to feel like genuine intellectual and creative collaboration.

Interacting with the system can resemble a conversation with a smart and experienced student or colleague: ideas are proposed, challenged, refined, rejected, and unexpectedly pushed in new directions, with both the human and the AI acknowledging mistakes and adjusting course. The exhilarating possibility is the sheer multiplication of intellectual labor - the ability to explore almost any question under the scientific sun and to see those “mountains beyond mountains”, as Tracy Kidder eloquently put it. In a week a scientist or a team can come up with dozens or hundreds of ideas and hypotheses, most of them reasonable and actionable. The physical lab is now the only bottleneck, and even that is being accelerated by AI.

As a scientist, AI has made me feel more intellectually alive and excited than I have felt since graduate school and my postdoctoral years more than two decades ago. I can begin with an idea in the morning and, by lunchtime, watch a rational, testable hypothesis take shape; within days, an investigation can progress from literature and classical calculations to increasingly rigorous quantum-mechanical analysis and experimentally actionable predictions. Eating and sleeping have often taken a backseat, exercise seems like a distant goal, and every night I feel like I did when I was a kid and begging my dad or mom for just *one more story*. Except that this time it’s just *one more prompt*. One more cycle of compute. One more result that will startle or confound. Every night I have to force myself to detach from the computer, and on more than one occasion I have fallen asleep at my desk, only to wake up and see the potential for yet *one more prompt*.

Precisely because these systems are beginning to criticize our assumptions, tell us when something does not work, and think alongside us rather than merely obey us, it feels as though we may already have passed beyond the first age of AI.

Of course, these predictions and results will stand or fall based on experimental testing, that’s how science always has been, but that’s no different from the pre-AI age. More importantly, in almost every case they appear as conclusions that any good scientist will regard as reasonable, at least as starting points. And sometimes they genuinely throw up a surprise. AI-enabled science should still be judged by the novelty, rigor, reproducibility, statistical validation, and epistemic integrity of the science, not by the novelty of the technology.

But there is no doubt now that a Rubicon has been crossed, and either we cross over to the other side or get left behind. What a time to be alive.

Wednesday, July 22, 2026

Robert Wright talks with Connor Leahy about the cult of AI

YouTube:

Robert Wright and Connor Leahy, US executive director of ControlAI, discuss the history of AI culture and the influences that shaped it—from the hippies, to the rationalists, to Marx, to sci-fi stories, and more. Plus: Mythos as “digital nuke," Dario's real motivations, and Connor's case against any and all attempts at building superintelligence.

Subscribe to The NonZero Newsletter at: https://www.nonzero.org/subscribe
Episode post on Substack: https://www.nonzero.org/p/the-ideolog...
Join NonZero's Discord server: / discord

0:00 Teaser
0:47 Connor’s place in the AI world
3:16 Mythos: “digital nuke”?
14:12 The milieu that birthed Dario and other AI titans
20:52 Yudkowsky, Marx, and other AI harbingers
35:34 Effective altruism’s effects on big tech
46:22 The "cult" in Silicon Valley culture
59:59 Is the AI race a market failure?
1:05:51 What’s really driving the AI titans?
1:19:04 Connor: ASI is a worse bet than Russian roulette

I've been aware of much of this story for some time, though many details are new. Back in 2023 I published an article in 3 Quarks Daily about the cultish nature of Silicon Valley computer culture, A New Counter Culture: From the Reification of IQ to the AI Apocalypse. Nonetheless, hearing the story once again, in one place, with new details has been interesting. It may shed light on just why the dominant AI culture is an technology monoculture. The cultish nature of the surrounding industrial culture has something to do with it.

Note that both Apple (1976) and Microsoft (1975) date back to the personal computer revolution of the 1970s. Amazon was founded when the web was born (1994), as was Google (Alphabet, 1998). Facebook (Meta) didn't emerge until the social web was well-established, 2004. None of them are pure AI companies. They came later; OpenAI in 2015 and Anthropic in 2021. Those were the products of this ingrown culture. The other companies were founded as business ventures. OpenAI and Anthropic were founded as expressions of a quasi-religious vision.

Note: As far as I can tell, Leahy seems rather adjacent and Wright seems to take their technological visions more seriously than is intellectually warranted.

Tuesday, July 21, 2026

Beyond Marginalism: What’s Next? [MR #12]

It is time to conclude my series of posts on Tyler Cowen’s monograph, The Marginal Revolution: Rise and Decline, and the Pending AI Revolution (2026). Let’s look at the fourth and final chapter, “Why Marginalism Will Dwindle, and What Will Replace It?” Here’s how Cowen opens it (p. 85):

The underappreciated news is that marginalism is on the way out. Furthermore, this is old news, though the trend is accelerating.

Most of all it is underdiscussed news. As economics continues to evolve, marginalist insights – probably of all different kinds – will lie ever further from the frontiers of research and knowledge.

I find it easy to imagine that – less than 20 years from now – marginalism will be viewed as a historical curiosity rather than a central analytical engine of economics. No one will quite come out and say that, nor will they present marginalism as false or destructive. Rather it will be seen as of limited relevance, much as we might view parts of the earlier classical economists, such as their expositions of the quantity theory of money. New and different analytical frameworks will replace the ones that have dominated neoclassical economics to date.

Think about that, think about it very carefully. When thinking about it remind yourself that Cowen named his blog, his virtual home base for the last two decades, after marginalism.

For a professional academic to say that the world in which they were trained, the structure of ideas within which they have worked, which they have nurtured in students, which they have communicated to the public at large, which they have come to love, to say that that world is slipping away into the past, man, that’s rough. And rare. Not many have been able to do it.

Back in 1946 the great physicist, Max Planck, remarked, “A new scientific truth does not triumph by convincing its opponents and making them see the light, but rather because its opponents eventually die, and a new generation grows up that is familiar with it.” Thomas Kuhn referenced that remark in The Structure of Scientific Revolution, and the economist Paul Samuelson gave a compressed version in a 1975 article in Newsweek. It would appear that Cowen has gotten the message and decided that, rather than dropping dead, he’d give the new ideas a boost.

After that sobering opening, Cowen reviews what happened between the late 19th century and now. He lands on price theory. Price theory? – “the view that the basic intuitive economic concepts, as would be taught in intermediate microeconomics, are highly useful and for advanced problems too” (p. 91). There’s that word, “intuitive.” Cowen explains:

Your hypothesis should be intelligible in terms of microeconomic concepts that you can hold in your mind and understand. In most (maybe not all?) cases, you should be able to explain some version of those principles to a well-educated, non-economist onlooker.

A couple pages later we arrive at something called “Topkis’s Theorem” which is very mathy (p. 94). Two pages after that: “Economic intuition, RIP. And marginalism with it.” Whoops! “I am seeing the traditional, intuitive approach to economic reasoning retreating from one field after another. To give one vivid and also important example, machine learning and neural nets are overturning the world of finance.”

Modeling collective action with 360,000 factors

A couple of pages later Cowen gives us a striking example. It’s from something called Arbitrage Pricing Theory (APT) (pp. 99-100). It’s a model that uses machine learning to develop 360,000 factors and does a better job of predicting than traditional models have only five or six factors. However, the factors in the traditional models are derived from marginalist assumptions and make intuitive sense while none of those 360,000 factors are legible. It’s clear to Cowen that, in the current intellectual marketplace for economics, the unintelligible models with superior performance are out-competing the traditional marginalist models. Bye, bye, marginalism!

I see no need to comment extensively on this particular model as I’ve already given it a great deal of attention, generating two different working papers from it. The first, On Method: Computational Compressibility in Complex Natural and Cultural Phenomena, places it in the context of a half-dozen other investigations in a half-dozen fields in the social and natural sciences. The second, Notes on the Collective Valuation of "Thick" Objects: Financial Assets, Movies, and Novels, compares it with work that Arthur De Vany published in 2004, Hollywood Economics, and a more recent study by Matthew Jockers, Macroanalysis (2013), in which he investigated a corpus of 6000 19th century Anglophone novels. I’ve also written a blog post that complements that second paper: Thick Objects, High-Dimensional Models, and the New Intuitions [MR #11].

In that second paper and in the blog post I argue that those three cases are about a collective process where a population of human actors – traders and analysts in one case, movie goers in another, and novel readers in the third case – make judgements about “thick” objects. Even before I made an explicit argument, I had an intuition, an intuition that, despite the obvious differences, what De Vany was up to with movies was somehow like what Didisheim et al. were up to with stocks. Just where those intuitions came from, I can’t say, but I’ve been thinking about complex systems for a long time. [As an aside, for what it’s worth, Robert De Vany’s work on movies is perhaps where my interests in culture and cultural evolution come into closest contact with Cowen’s interests in economics and, in particular, in the economics of culture.]

As for the idea of thick objects, the term was suggested to me by either ChatGPT or Claude to characterizes complex objects whose characteristics cannot be fully enumerated because of that complexity. Moreover they are under constant scrutiny by a population of people who are interested in them and constantly evaluating them back and forth among themselves and, in that process, revealing further characteristics. It is not difficult to see that movies and novels are the same kind of thing, each is a mode of storytelling, and that they are complex objects. But what do they have to do with stocks? A remark by the pundit, Scott Galloway, made the connection for me in a podcast with Kara Swisher, “Stocks are like brands and that is they’re part promise and part performance.” Performance is assessed by a wide variety of metrics, metrics which go into the models such as the one by Didisheim et al., while promise is subject to endless speculation, some of which inevitably precipitates into those metrics.

Animal spirits, narrative economics, memes, and a Squid Game market

And that leads me to a conjecture that follows from the analysis that ChatGPT and I undertook in the collective valuation paper. Perhaps those 360,000 parameters are picking up traces left by those “animal spirits” that Keynes talked about. Their effect on asset values is too diffuse and indirect to be detected by those classical models with a half-dozen or so factors, each of which is intuitively legible on its own. But those traces show up distributed across those 360,000 parameters and allow the model to produce more accurate predictions. If that is what is going on, then I wouldn’t expect any of those factors to be intuitively legible, any more than one would expect such legibility of individual weights in a large language model. That’s not the nature of this conceptual world.

While we’re speculating, why not continue on? Those animal spirits can’t work their ways on the market by wafting around like odors in a breeze. They need to be embodied in some form, like gossip and stories. That leads us to Robert Shiller’s 2017 paper on “Narrative Economics” in the American Economic Review. Here’s his abstract:

This address considers the epidemiology of narratives relevant to economic fluctuations. The human brain has always been highly tuned toward narratives, whether factual or not, to justify ongoing actions, even such basic actions as spending and investing. Stories motivate and connect activities to deeply felt values and needs. Narratives “go viral” and spread far, even worldwide, with economic impact. The 1920–1921 Depression, the Great Depression of the 1930s, the so-called Great Recession of 2007–2009, and the contentious political-economic situation of today are considered as the results of the popular narratives of their respective times. Though these narratives are deeply human phenomena that are difficult to study in a scientific manner, quantitative analysis may help us gain a better understanding of these epidemics in the future.

Perhaps those high factor models are picking up the narrative dimension of asset value, which is a product how performance and promise become intertwined in the stories that analysts and traders tell themselves and one another about the assets they’re watching.

That, in turn, leads to the concept of meme stocks, a term that dates back to 2020. Here’s how Wikipedia characterizes them:

...a stock that gains popularity among retail investors through social media. The popularity of meme stocks is generally based on internet memes shared among traders, on platforms such as Reddit's r/wallstreetbets. Investors in such stocks are often young and inexperienced investors. As a result of their popularity, meme stocks often trade at prices that are above their estimated value – as based on fundamental analysis – and are known for being extremely speculative and volatile.

More recently, Owen A. Lamont, a senior analyst at Arcadian, has speculated that we’re in what he calls a “Squid Game market”:

Something’s happening in the U.S. stock market. We see cult stocks and crypto stocks. We see money pouring into leveraged single-stock ETFs and crypto ETFs. And we see dramatic price moves, for example in quantum computing stocks in December 2024. What’s going on?

Here’s one theory: these phenomena partly reflect an influx of Korean retail investors into the U.S. stock market. Last year, I wrote that “the U.S. stock market is Koreafying,” meaning that the U.S. market was starting to behave like the retail-dominated Korean market. What I didn’t realize was that this Koreafying process involves actual Korean retail investors.

He then goes on to develop the parallel between the Korean streaming series, Squid Game, and the U.S. retail market over the last few years.

If those high parameter models are picking up the effects of animal spirits embodied in gossip and narratives, then we’d expect their advantage over classical models (based on a handful of fundamentals) to be larger in the case of these meme stocks. So, if we compare the results of a classical model with those of a high-parameter machine learning model, are the assets with the greatest divergence also those otherwise identified as meme stocks? Perhaps some intellectual fishing expeditions are in order. Perhaps we can develop some new intuitions by comparing the results of classical models with machine learning models.

Whoops! Alphabet, Microsoft, Amazon, Meta, and Oracle have $1.65 trillion in debt that doesn't appear on their balance sheets

Agentic AI for science – Yippie!

Monday, July 20, 2026

America and China in the World, a quick note

Is that where we’re headed, to a bipolar world dominated by America and China? I just asked Google Search: “Is China the largest economy in the world?” Its reply:

China is the world's largest economy when measured by Purchasing Power Parity (PPP), but the United States is the largest in terms of nominal Gross Domestic Product (GDP).

  • By Purchasing Power Parity (PPP): China is the world's largest economy, with an estimated output exceeding $44 trillion. This metric adjusts for the cost of living and the price of local goods, reflecting a higher real volume of economic activity.
  • By Nominal GDP: The United States holds the top spot. In current market exchange rates, the U.S. economy is valued at approximately $32 trillion, while China's nominal GDP is roughly $20 trillion.

So that’s one thing. The other is AI. America and China are in a race for dominance in AI. America has a technical lead, but with respect to Frontier LLMs, that lead is measured in months, not years. And China is making its models open weight while the most advanced American models (Anthropic, OpenAI, Google) and not open. It’s not at all clear how that will shake out.

With the Trump administration authoritarianism is on the rise in America. While the Democrats may well win the next presidential election, it’s not at all clear what that means for America’s already precarious democracy. The problem is that a great deal of power is concentrated in the hands of political and business elites, so much that America’s democracy looks like a chess game played by oligarchs where ordinary Americans are the pawns.

Is that what the world will become in 2040, a competition between American and Chinese oligarchs in which everyone else is a pawn, with Homo economicus triumphant over all?

Saturday, July 18, 2026

Some notes on AI and “fine art” imagery

A couple of weeks ago I had a post entitled “Friday Fotos: The Last Frontier of AI.” I was interested in whether or not a certain approach I’d been using to create images with ChatGPT could produce “fine art” images, as opposed to illustrations or popular art of various kinds. The particular images I developed for that post (there were five), while interesting, were not particularly compelling. So I went on to display eleven other images I’d created with ChatGPT, including some of the images which that had motivated the post in the first place. I then ranked the images among themselves and decided that the images I’d created specifically for the post ranked near the bottom.

So, while the approach that motivated that post cannot be called a success, the post as a whole has raised the question: Can ChatGPT (be used to) create “fine art” images? I put “fine art” in quotes because the term itself is problematic. It’s not as though there are identifiable characteristics such that any image exhibiting them is a fine art image. The notion of fine art as opposed to folk art or popular art or (mere) illustrations is a cultural convention, one that Marcel Duhamps exploded in 1917 when he entered a urinal into the inaugural exhibition of the Society of Independent Artists. He called it Fountain and attributed it to “R. Mutt.” Fine art is simply the art that society has decided deserves to be treated in a certain way, no more, no less. If you decided that a common urinal should be treated in that way, then it becomes fine art.

Duchamp’s move was controversial, and that controversy has been reverberating ever since. I have no intention of reviewing and rehashing it here. Rather, I simply want to present a collection of images I’ve made with ChatGPT and view them with that issue reverberating in the background.

This image is one of my favorites among those I’ve created with ChatGPT:

I created it for illustrative purposes, to go on the cover of a working paper about Joseph Conrad’s Heart of Darkness, but I think the image stands on its own. If you’re familiar with the book, then resonance is obvious. It tells about a voyage up the Congo River. As for the superimposed image of the Buddha, here’s first sentence of the last paragraph: “Marlow ceased, and sat apart, indistinct and silent, in the pose of a meditating Buddha.”

I should note, and this is important, I had ChatGPT create that image from within a chat devoted to that working paper, making the entire chat (up to that point) the context in which ChatGPT created the image. I have reason to believe that it wasn’t working simply from prompt that generated that image. For a discussion of that, see this post: High-level “vibe” – Creating Imaginary Bank Notes with ChatGPT: AI as cultural technology and collective creativity.

Here’s a somewhat different image that I also like very much. It’s almost, but not completely, abstract:

The book is obvious. The rest of it? But that’s not the first image ChatGPT offered to me. This came before (and there were others before this):

If it’s fine art we’re interested in, the black and white image seems (vastly) superior to me.

Here’s an utterly different image:

I don’t remember what prompt I used to create that. But I like the image, absurd as it is, a lot. THAT’s why I like it. It’s ridiculous, but fun. Fine art? Ask me if I care. 

What about this?