I just read this article: Kevin Roose, When A.I. Passes This Test, Look Out, NYTimes, Jan. 23, 2025.
“Fiddlesticks! Fiddlesticks and moonbeams!” said I to myself. “Not a benchmark-to-end-all-benchmarks. Will these people never learn? Why do they keep looking for the house keys under the streetlight when they lost them at the front door? Are they drunk or something?” Oh, they don’t know what they’re doing, do they?
So I put the article aside, intending to write up a post about it and to register my complaint about the inadequacy of benchmarks. Then I had an idea: I’ve not used ChatGPT in a while, so why don’t I have it summarize the article for me.” Once the Chatster had done that, I had an idea: Why not ask it what it thinks about bench-marking?
I didn’t have any particular expectations about what kind of answer it would give. To my surprise and delight it came up with a substantial critique. I decided to press on and told the story about how I came to write about Spielberg’s Jaws. Again, the Chatster rose to the occasion and wrote a mini-essay, “The Role of Discovery in Intellectual Work.” It was on point, so I gave it my whaling analogy. Chatster broke it down very nicely.
“Damn!” thought I to myself. “Chatster’s on FIRE!”
And away we go.
NOTE: See my post, Competence and Performance in Bench-marking LLMs.
* * * * *
Summarize this article, please. Assume I am sophisticated and well-educated.
The article, "A Test So Hard No AI System Can Pass It — Yet," examines the development and significance of a new benchmark called "Humanity’s Last Exam," designed to evaluate the limits of artificial intelligence (AI) capabilities. Created by Dan Hendrycks and his team at the Center for AI Safety in collaboration with Scale AI, the test comprises 3,000 highly challenging questions spanning disciplines like philosophy, physics, and engineering, contributed by experts. Its goal is to assess AI's ability to perform complex intellectual tasks, approximating a general intelligence score.
The impetus for this exam arises from the rapid advancements in AI, which have rendered traditional benchmarks obsolete as models like those from OpenAI and Google surpass graduate-level challenges. Despite their advanced abilities, current AI models failed Humanity’s Last Exam, with the best-performing model achieving only 8.3%. Hendrycks anticipates significant improvement, potentially marking a point where AI systems surpass human experts in answering complex questions.
The article highlights concerns over AI's uneven progress—astounding in specialized tasks but deficient in others—and discusses the broader implications for fields requiring nuanced problem-solving beyond exams, such as research. It concludes with a call for innovative methods to measure AI’s impact beyond standardized testing, particularly as these systems continue to evolve.
What do you think of this whole “benchmark” business? How useful are they?
Benchmarks like Humanity's Last Exam have value, but their utility is nuanced. Here’s my take:
Strengths of Benchmarks:
- Tracking Progress: Benchmarks serve as yardsticks, offering a quantitative way to measure the trajectory of AI systems over time. They allow researchers to compare different models and identify strengths and weaknesses systematically.
- Challenge Motivation: By creating hard problems, benchmarks inspire innovation. They push AI researchers to design models that can tackle tasks requiring deeper reasoning, creativity, and generalization.
- Defining Boundaries: Benchmarks like this can help delineate areas where AI falls short, highlighting tasks that still require human expertise, such as novel discovery or conceptual synthesis.
Limitations and Challenges:
- Artificial Metrics: Benchmarks may overfit AI development to specific tasks, leading to “training to the test” rather than true generalization. Models can excel at benchmarks without demonstrating real-world problem-solving abilities.
- Narrow Framing: Many benchmarks, including even sophisticated ones, focus on clearly-defined problems with deterministic answers. But much of human intellectual labor—creative writing, ethical decision-making, or forming new scientific hypotheses—doesn’t fit neatly into these molds.
- Temporary Relevance: As seen with earlier AI challenges, benchmarks can become obsolete quickly. Once AI models surpass them, they lose their diagnostic power, necessitating constant updates or more difficult tests.
- Underestimating Complexity: Benchmarks often fail to account for the broader, messier contexts in which humans operate. AI might pass a theoretical test on physics but struggle to apply those principles in practical, experimental settings.
Broader Perspective:
Benchmarks are useful for building a snapshot of what AI can do at any moment, but they must evolve to remain meaningful. To go beyond mere scoring, we might need:
- Dynamic Benchmarks: Tasks that adapt to the model’s performance, such as questions whose answers are not yet known or which require multi-modal reasoning (e.g., combining visual and textual information).
- Real-World Applications: Shifting focus from exams to practical contributions, like assessing AI’s impact on medicine, climate science, or engineering breakthroughs.
- Human-Centric Evaluation: Considering AI’s utility in collaborative settings, where it aids but does not replace human expertise.
Benchmarks like Humanity’s Last Exam have an important place in measuring AI progress. However, we must also look beyond them to gauge AI’s real-world significance and its potential to address humanity’s most pressing challenges. They’re a piece of the puzzle, not the whole picture.
* * * * *
I agree. The problem I have with benchmarks is that it asks for something specific. But that’s not always how life is lived. Often we have to figure out whether there’s a problem or something of interest. Once we’ve done that, we have to figure out an approach.
I’m trained as a literary and cultural critic. I read a lot of books, see a lot of films and TV shows. I write about many of these texts, broadly considered. It’s up to me to notice that there’s something interesting about a text. For example, a couple of years ago I was watching Steven Spielberg’s movie “Jaws.” After I had watched the movie I read the Wikipedia entry, which mentioned various interpretations. It also mentioned that the sequels were not as good. So I looked at them (couldn’t finish 3 and 4). All the while I was thinking.
I started comparing Jaws 2 with the original. The original was much tighter. I started thinking about the role of Quint. “Why did he have to die?” I asked myself. Sure, he wasn’t a nice man, but he hadn’t hurt anyone. All that had to happen to give the movie a “happy” ending was that the shark gets killed. It wasn’t necessary to have anyone else die. So, what’s the point of having Quint die? How would the film be different if he hadn’t been killed by the shark?
It was at that point that it hit me: Girard, sacrifice, mimetic desire. NOW I had a specific question: How do Girard’s ideas apply to Jaws? It took me awhile to come up with an answer that I liked. But that’s how it goes. The really important thing was coming up with the idea to apply Girard’s ideas to the movie. That’s what took the hard work.
These benchmarks don’t test that kind of work at all. But that’s where the really hard and fruitful problems lie.
You’ve pinpointed a profound challenge in evaluating AI: life—and much of human intellectual labor—is far messier and more open-ended than what benchmarks can capture. Your example of watching Jaws illustrates exactly the kind of complex, emergent thinking that is central to creative, critical, and interpretive work. Let’s unpack your critique a bit further.