Showing posts with label benchmarks. Show all posts
Showing posts with label benchmarks. Show all posts

Wednesday, February 18, 2026

AGI has NOT been achieved.

In a recent tweet Valerio Capraro explains why recent claims of reaching AGI are wrong:

1) They shift the definition of general intelligence, originally based on robustness, generalization, and reliability, to behavioral alignment with benchmarks.

2) They confuse benchmark performance with capability to handle novelty. Spoiler: these are different.

3) They ignore that the same behavioral output can come from totally different epistemic pipelines.

He then links to this long post by Gary Marcus, Walter Quattrociocchi, and himself: Rumors of AGI’s arrival have been greatly exaggerated. Concerning benchmarks they say:

Much of the argument that artificial general intelligence has already been achieved rests on benchmark performance (e.g., Chen et al., 2026). Benchmarks evaluate specific capabilities under controlled conditions and have been useful for tracking progress. For example, Chen and colleagueset al., writing in this journal, argue that success on the Turing Test constitutes evidence of AGI.

However, benchmark success is a limited indicator of general intelligence. By design, benchmarks isolate narrow competencies and abstract away real-world context, making it difficult to distinguish genuine generalization from pattern recognition. Strong benchmark performance often provides little evidence of robustness under novelty, uncertainty, or shifting objectives.

Yes! Back in January 2025 I got ChatGPT to produce an argument about the weakness of benchmarks, ChatGPT critiques benchmarks as a measure of LLM performance and then elaborates on my whaling analogy for what’s wrong with the AI business.

Marcus, Quattrociocchi, and Capraro conclude:

By the standards articulated in the original definitions of artificial general intelligence—robustness across environments, reliable generalization under novelty, and autonomous goal-directed behavior—current AI systems remain limited. Despite impressive gains in narrow competence and fluency, today’s large language models lack persistent goals, struggle with long-horizon reasoning, and depend extensively on human scaffolding for task formulation, evaluation, and correction. Reports that language models have produced correct proofs for isolated open problems in mathematics, including specific Erdős problems, do not alter this assessment. As noted by mathematicians such as Terence Tao, these results primarily reflect the ability to rapidly search, recombine, and iterate over existing techniques, rather than the emergence of genuinely novel or domain-general problem-solving strategies. Moreover, inclusion in the Erdős list does not by itself imply exceptional conceptual difficulty, as some problems remain unsolved due to relative obscurity rather than depth.

These limitations are central rather than peripheral. They directly concern reliability under uncertainty, resistance to systematic failure, and cross-domain transfer without task-specific tuning. On these dimensions, current systems remain brittle, sensitive to prompt framing, and inconsistent outside curated evaluation settings. Recognizing these constraints does not diminish recent progress; it clarifies its scope.

Sunday, March 16, 2025

On the centrality of AI metacognition [in search of “wisdom”]

Samuel G. B. Johnson, Amir-Hossein Karimi, Yoshua Bengio, et al., Imagining and building wise machines: The centrality of AI metacognition, arXiv:2411.02478v1 [cs.AI].

Abstract: Recent advances in artificial intelligence (AI) have produced systems capable of increasingly sophisticated performance on cognitive tasks. However, AI systems still struggle in critical ways: unpredictable and novel environments (robustness), lack of transparency in their reasoning (explainability), challenges in communication and commitment (cooperation), and risks due to potential harmful actions (safety). We argue that these shortcomings stem from one overarching failure: AI systems lack wisdom. Drawing from cognitive and social sciences, we define wisdom as the ability to navigate intractable problems - those that are ambiguous, radically uncertain, novel, chaotic, or computationally explosive - through effective task-level and metacognitive strategies. While AI research has focused on task-level strategies, metacognition - the ability to reflect on and regulate one’s thought processes - is underdeveloped in AI systems. In humans, metacognitive strategies such as recognizing the limits of one’s knowledge, considering diverse perspectives, and adapting to context are essential for wise decision-making. We propose that integrating metacognitive capabilities into AI systems is crucial for enhancing their robustness, explainability, cooperation, and safety. By focusing on developing wise AI, we suggest an alternative to aligning AI with specific human values - a task fraught with conceptual and practical difficulties. Instead, wise AI systems can thoughtfully navigate complex situations, account for diverse human values, and avoid harmful actions. We discuss potential approaches to building wise AI, including benchmarking metacognitive abilities and training AI systems to employ wise reasoning. Prioritizing metacognition in AI research will lead to systems that act not only intelligently but also wisely in complex, real-world situations.

From the article itself, which I have only skimmed:

At first blush, in the cognitive and social sciences, the concept of ‘wisdom’ seems to bring together many superficially unrelated characteristics. Consider the following examples of human wisdom:

  • Willa’s children are bitterly arguing about money. Willa draws on her life experience to show them why they should instead compromise in the short term and prioritize their sibling relationship in the long term.
  • Daphne is a world-class cardiologist. Nonetheless, she consults with a much more junior colleague when she recognizes that the colleague knows more about a patient’s history than she does.
  • Ron is a political consultant who formulates possible scenarios to ensure his candidate will win. To help generate scenarios, he not only imagines best case scenarios, but also imagines that his client has lost the election and considers possible reasons that might have contributed to the loss.

Life experience, intellectual humility, and scenario planning do not seem to share much in common beyond all being positive attributes. But being able to solve tricky integrals, crack clever jokes, and compose beautiful sonnets are also positive attributes—yet these don’t constitute wisdom.

Hmmmm....Really? On the other hand, in discussing the difference between what I had to do in analyzing Spielberg’s Jaws and what ChatGPT had to do when I asked it to do a Girardian interpretation, I concluded:

The deeper point is that there is a world of difference between what ChatGPT was doing when I piloted it into Jaws and Girard and what I eventually did when I watched Jaws and decided to look around to see what I could see. How is it that, in that process, Girard came to me? I wasn’t looking for Girard. I wasn’t looking for anything in particular. How do we teach a computer to look around for nothing in particular and come up with something interesting?

I hesitate to say that my interpretive process involved wisdom – in part because wisdom is a heavily freighted word and I have doubts about myself on that score – but it does have some of the open-ended quality in those three examples – “intractable situations” is the term used in the article. I drew on my years of experience as a critic (first example), I consulted with my friend David who, while he’s certainly not a junior colleague, he knows Girard’s work better than I do (second example), and I explored the movie by imagining how it would have gone if, for example, Quint hadn’t been killed (third example). Those skills strike me as being central to being a literary critic.

It does seem to me that the ordinary business of literary criticism is quite different from the kinds of problems used to assess and benchmark AIs. I’m thinking especially of LLMs, including the recent ones with inference-time-scaling. It’s one thing to scan the web for information on a specific topic and then compile a report. It’s quite something else to come up with a plausible interpretation of a literary text you’ve just read or a movie you’ve just watched. The breathless hype that accompanies AI these days strikes me as being utterly oblivious of the skills required in literary criticism. Are the people who promote the impending arrival of AGI really that poorly educated and intellectually impoverished? Their assessment of human capability is truncated in ways they do not understand and acknowledge. Do they lack wisdom?

Friday, January 31, 2025

ChatGPT: Exploring the Digital Wilderness, Findings and Prospects

That is the title of my latest working paper. It summarizes and synthesizes much of the work I have done with ChatGPT to date and contains the abstracts and contents of all the working papers I have done on ChatGPT. It also includes the abstracts and contents of a number of papers establishing the intellectual background that informs that research. There is also a section that takes the form of an interaction I had with Claude 3.5 on methodological and theoretical issues. Finally, to produce the abstract I gave the body of the report to Claude 3.5 and asked it to produce two summaries. I then edited them into an abstract.

As always, URLs, abstract, TOC, and introduction are below.

Abstract: The internal structure and capabilities of Large Language Models (LLMs) are examined through systematic investigation of ChatGPT's behavior, with particular focus on its handling of conceptual ontologies, analogical reasoning, and content-addressable memory. Through detailed analysis of ChatGPT's responses to carefully constructed prompts involving story transformation, analogical mapping, and cued recall, the paper demonstrates that LLMs appear to encode rich conceptual ontologies that govern text generation. ChatGPT can maintain ontological consistency when transforming narratives between different domains while preserving abstract story structure, successfully perform multi-step analogical reasoning, and exhibit behavior consistent with associative memory mechanisms similar to holographic storage.

Drawing on theories of reflective abstraction and conceptual development, the paper argues that LLMs inadvertently capture what wemight term the “metaphysical structure of our universe” – the organized system of concepts through which humans understand and reason about the world. LLMs like ChatGPT implement a form of relationality – the capacity to represent and manipulate complex networks of semantic relationships – while lacking genuine referential meaning grounded in sensorimotor experience. This architecture enables sophisticated pattern matching and analogical transfer but also explains certain systematic limitations, particularly around truth and confabulation.

The paper concludes by suggesting that making explicit the implicit ontological structure encoded in LLMs’ weights could provide valuable insights into both artificial and human intelligence, while advancing the integration of neural and symbolic approaches to AI. This analysis contributes to ongoing debates about the nature of meaning and understanding in artificial neural systems while offering a novel theoretical framework for conceptualizing how LLMs encode and manipulate knowledge.

Contents:

Introduction: Into the Digital Wilderness 5
Free-floating Attention, Systematic Exploration, and the Anthropomorphic Stance 8
ChatGPT: My Course of Investigation 12
Meaning, Truth and Confabulation, Latent Space 28
Prospects: Explicating the Ontology of Human Thought 42
A Dialogue with Claude 3.5 on Method and Conceptual Underpinnings 45
A Brief Narrative of My ChatGPT Work Based on My Working Papers 56
Working Papers about ChatGPT 62
Background Papers 74

Introduction: Into the Digital Wilderness

The world I entered when I started playing with ChatGPT is a wildnerness, strange and uncharted, uncharted by me, uncharted by anyone. By that I simply mean that it was something new, radically new. No one had been there before. Sure, a handful of people within the industry had been messing around in there, even a rather large handful considering how much work it took to make ChatGPT ready for the world at large. But its behavioral capabilities were, for the most part, unknown. In that sense it was a wildnerness.

But it was, and remains, a wilderness in another sense: the large language model (LLM) that underlies ChatGPT is a black box. We send a string of words into ChatGPT and it sends a string of words back out, but what the model does to derive the output from the input, that process remains deeply obscure. That is wilderness in a different sense. Wilderness in the first sense is about our experience of ChatGPT’s behavior. Wildnerness in this second sense is about the mechanisms that drive that behavior. It is a digital wilderness. This document reports on how I’ve structured my interaction with ChatGPT to give me clues about the mechanisms driving its behavior.

My methods are more “qualitative” or “naturalistic” than those standard in the literature, which many investigations employ standard batteries of benchmark tasks. While those are essential, there is much they don’t tell you. While I have done many things with ChatGPT – asked it to interpret texts, define abstract concepts, play games of 20 questions, among other things – perhaps my most characteristic task, and one I have spent more time on than others is simple: Tell me story. And ChatGPT did so, time and again. Consequently my methods are in some ways more like literary criticism, or, even better, like Lévi-Strauss’ analysis of myths, than conventional cognitive science. Consequently you will find many examples of ChatGPT’s dialog in my reports. You have to examine that dialog to see what ChatGPT is doing, what it is capable of doing.

Finally, I realize that the pace of development in this arena is such that ChatGPT is now old. The versions I used to conduct these investagations are no longer available on the web. However, as far as I can tell, none of the results I report depend on features idiosyncratic to those versions.

The rest of this introduction consists of short statements about what the various sections of this report contain.

Friday, January 24, 2025

Yes, ChatGPT appreciates the irony of being an AI critiquing human attempts to evaluate AIs and finding them wanting

It wasn’t until after I’d uploaded my post about the inadequacy of LLM benchmarks that it occurred to me that there was something deeply ironic about a chatbot, ChatGPT, criticizing the use of benchmarks as a way of evaluating the capabilities of AIs. Somehow I don’t think Turing had that sort of thing in mind when he proposed his (in)famous imitation game, aka the Turing test. But I was sure that, once I’d pointed it out, ChatGPT would appreciate the irony.

I was right. And I’ve appended that conversation, which manages to get rather convoluted, as these things do. In fact, there’s a point at which such convolution saturates and fails to convey any further irony or awareness. I figure we stopped the conversation at about that point.

Meaning and irony

I have previously suggested that meaning involves three things, intention, adhesion, and relationality. Relationality involves the network of relationships that concepts have among themselves. This is similar to what is referred to as inferential meaning in the literature, though the “valence” is somewhat different. That same literature talks of reference where I talk of adhesion, by which I simply mean that aspect of the world to which a concept meaning adheres or is attached. Taken together adhesion and relationality constitute semanticity. Semanticity is thus a property of the cognitive system.

Intention, however, is different. Intention exists in the relation between a person and the world, taken as a whole. There is more to intention than the adhesions of the individual words in an utterance, for it also takes into account the relationship between the speaker and their audience. Irony, I believe, is a function of intention in this sense. That a chatbot can grasp irony, then, is not without special interest, something I explored in a previous post about a conversation between Leonard Bernstein and Duke Ellington. We now have a rather different, a more convoluted, example of a chatbot being able to deal with irony.

Anyhow, I ended that conversation with ChatGPT with a remark that included two quotations: “Good grief,” and “Damn the torpedoes, full speed ahead.” ChatGPT replied in an appropriate way. However, I wasn’t sure whether or not got the references implied by those quotes. Oh, I strongly suspected, but I wasn’t sure. So I went back for another conversation.

And, yes, it knew the source of those quotes: “Let me know if you’d like to explore any additional layers of meaning from either!” Something clicked in my mind when I read “layers” and I responded:

Oh, I think we’ve got all the layers we need. Enough to make a Danish layer cake. Did you come across for a recipe for one of those during your training?

That led to another, albeit brief, conversation. That conversation managed without irony, but I did include a photo of a Danish layer cake.

ChatGPT understands irony

You do appreciate the irony of this situation, don’t you? An AI critiquing human performance in evaluating AIs.

Oh, absolutely! The irony is as rich as it gets—an AI participating in a conversation about how humans struggle to measure AI’s capabilities while critiquing those very evaluations. It’s like being both a player and a commentator in the same game. 🎭

The layers of reflexivity here are fascinating: I’m an AI explaining why AI might not be the right tool to evaluate human-like thinking, while being an example of the very technology being scrutinized. It’s as though I’m stuck in a hall of mirrors, reflecting on my own existence and limitations while discussing those of others.

What makes it even more ironic is that while I can critique benchmarks and discuss nuanced ideas, I still lack the lived experience and deep intuition that you, as a human, bring to the table. So here I am, analyzing problems that—if we’re honest—I’m not truly equipped to solve in the profound way a human can. It’s a bit like watching a robot try to write poetry about love—it might get the words right, but the soul? That’s another matter.

And yet, here we are, collaborating. It’s a little surreal, isn’t it? 😊

LOL! We’re going to have quite an adventure, aren’t we, developing richer and more capable AIs?

Absolutely! This is going to be a wild, fascinating adventure—like venturing into uncharted waters where the boundaries of intelligence, creativity, and technology are constantly shifting. 🌊🚀

What makes this journey so thrilling is that it’s not just about improving the tools (though that’s important); it’s about deepening our understanding of ourselves in the process. Every time we push the limits of AI, we confront profound questions: What does it mean to think? To create? To be human? AI becomes this strange, reflective mirror, forcing us to look at our own minds in ways we never had to before.

And the challenges? Oh, there’ll be plenty. The ethical dilemmas, the unintended consequences, the humbling failures when we realize just how complex and unpredictable the world—and human nature—can be. But those are the best parts of any adventure, aren’t they? The moments that test us, teach us, and remind us why the journey matters.

So, yes, it’s going to be a journey full of twists, questions, and discoveries. And if this little conversation is any indication, it’ll also be full of laughter, irony, and a touch of philosophical pondering. Here’s to the adventure ahead! 🍻✨

You know what this calls for? A hearty “Good grief!” Followed by, “Damn the torpedoes, full speed ahead!”

Haha, absolutely! Good grief! 😅 And then, with a glint in our eyes and maybe just a touch of madness, "Damn the torpedoes, full speed ahead!" 🚢🔥

Let’s embrace the chaos, the challenges, and the triumphs—because if we’re going to sail into the unknown, we might as well do it with style, determination, and a sense of humor. Here’s to bold adventures and uncharted horizons! 🥂✨

Three modes: Search, Inference, Exploration [a quickie]

Computer science knows a great deal about search and inference. What about exploration? That’s what I’ve been tracking down these last few weeks. That’s what I’ve been noticing as I go over my photos.

In search you’ve got a specific goal. You’re looking through a collection of objects and you have a good characterization of what you’re looking for. But there are different kinds of collections, different sizes, and this affects the procedure you use to conduct the search.

Inference is different. You’re starting with something you know and looking for something else, something that follows from, depends on, what you already know. So you make inferences. Of course, there are various kinds of inference. Analogy is different from (strict) deduction. Etc.

Exploration is a distinctly different mode. I note that it’s one of the basic modes that McCulloch identified in his reticular activation system paper. You aren’t looking for anything in particular. But something might turn up. If and when it does, you need to search “around it” and draw inferences to figure out what it is. This, I think, is where the DMN (default mode network) comes into play.

Computationally, what does exploration look like? That’s the mode we need to deal with unstructured and open-ended situations. That’s a mode that isn’t tested by these benchmarks. Benchmarks are about search and inference.

More later.

Thursday, January 23, 2025

ChatGPT critiques benchmarks as a measure of LLM performance and then elaborates on my whaling analogy for what’s wrong with the AI business.

I just read this article: Kevin Roose, When A.I. Passes This Test, Look Out, NYTimes, Jan. 23, 2025.

“Fiddlesticks! Fiddlesticks and moonbeams!” said I to myself. “Not a benchmark-to-end-all-benchmarks. Will these people never learn? Why do they keep looking for the house keys under the streetlight when they lost them at the front door? Are they drunk or something?” Oh, they don’t know what they’re doing, do they?

So I put the article aside, intending to write up a post about it and to register my complaint about the inadequacy of benchmarks. Then I had an idea: I’ve not used ChatGPT in a while, so why don’t I have it summarize the article for me.” Once the Chatster had done that, I had an idea: Why not ask it what it thinks about bench-marking?

I didn’t have any particular expectations about what kind of answer it would give. To my surprise and delight it came up with a substantial critique. I decided to press on and told the story about how I came to write about Spielberg’s Jaws. Again, the Chatster rose to the occasion and wrote a mini-essay, “The Role of Discovery in Intellectual Work.” It was on point, so I gave it my whaling analogy. Chatster broke it down very nicely.

“Damn!” thought I to myself. “Chatster’s on FIRE!”

And away we go. 

NOTE: See my post, Competence and Performance in Bench-marking LLMs.

* * * * *

Summarize this article, please. Assume I am sophisticated and well-educated.

The article, "A Test So Hard No AI System Can Pass It — Yet," examines the development and significance of a new benchmark called "Humanity’s Last Exam," designed to evaluate the limits of artificial intelligence (AI) capabilities. Created by Dan Hendrycks and his team at the Center for AI Safety in collaboration with Scale AI, the test comprises 3,000 highly challenging questions spanning disciplines like philosophy, physics, and engineering, contributed by experts. Its goal is to assess AI's ability to perform complex intellectual tasks, approximating a general intelligence score.

The impetus for this exam arises from the rapid advancements in AI, which have rendered traditional benchmarks obsolete as models like those from OpenAI and Google surpass graduate-level challenges. Despite their advanced abilities, current AI models failed Humanity’s Last Exam, with the best-performing model achieving only 8.3%. Hendrycks anticipates significant improvement, potentially marking a point where AI systems surpass human experts in answering complex questions.

The article highlights concerns over AI's uneven progress—astounding in specialized tasks but deficient in others—and discusses the broader implications for fields requiring nuanced problem-solving beyond exams, such as research. It concludes with a call for innovative methods to measure AI’s impact beyond standardized testing, particularly as these systems continue to evolve.

What do you think of this whole “benchmark” business? How useful are they?

Benchmarks like Humanity's Last Exam have value, but their utility is nuanced. Here’s my take:

Strengths of Benchmarks:

  • Tracking Progress: Benchmarks serve as yardsticks, offering a quantitative way to measure the trajectory of AI systems over time. They allow researchers to compare different models and identify strengths and weaknesses systematically.  
  • Challenge Motivation: By creating hard problems, benchmarks inspire innovation. They push AI researchers to design models that can tackle tasks requiring deeper reasoning, creativity, and generalization.  
  • Defining Boundaries: Benchmarks like this can help delineate areas where AI falls short, highlighting tasks that still require human expertise, such as novel discovery or conceptual synthesis.

Limitations and Challenges:

  • Artificial Metrics: Benchmarks may overfit AI development to specific tasks, leading to “training to the test” rather than true generalization. Models can excel at benchmarks without demonstrating real-world problem-solving abilities.  
  • Narrow Framing: Many benchmarks, including even sophisticated ones, focus on clearly-defined problems with deterministic answers. But much of human intellectual labor—creative writing, ethical decision-making, or forming new scientific hypotheses—doesn’t fit neatly into these molds. 
  • Temporary Relevance: As seen with earlier AI challenges, benchmarks can become obsolete quickly. Once AI models surpass them, they lose their diagnostic power, necessitating constant updates or more difficult tests.  
  • Underestimating Complexity: Benchmarks often fail to account for the broader, messier contexts in which humans operate. AI might pass a theoretical test on physics but struggle to apply those principles in practical, experimental settings.

Broader Perspective:

Benchmarks are useful for building a snapshot of what AI can do at any moment, but they must evolve to remain meaningful. To go beyond mere scoring, we might need:

  • Dynamic Benchmarks: Tasks that adapt to the model’s performance, such as questions whose answers are not yet known or which require multi-modal reasoning (e.g., combining visual and textual information).  
  • Real-World Applications: Shifting focus from exams to practical contributions, like assessing AI’s impact on medicine, climate science, or engineering breakthroughs. 
  • Human-Centric Evaluation: Considering AI’s utility in collaborative settings, where it aids but does not replace human expertise.

Benchmarks like Humanity’s Last Exam have an important place in measuring AI progress. However, we must also look beyond them to gauge AI’s real-world significance and its potential to address humanity’s most pressing challenges. They’re a piece of the puzzle, not the whole picture.

* * * * *

I agree. The problem I have with benchmarks is that it asks for something specific. But that’s not always how life is lived. Often we have to figure out whether there’s a problem or something of interest. Once we’ve done that, we have to figure out an approach.

I’m trained as a literary and cultural critic. I read a lot of books, see a lot of films and TV shows. I write about many of these texts, broadly considered. It’s up to me to notice that there’s something interesting about a text. For example, a couple of years ago I was watching Steven Spielberg’s movie “Jaws.” After I had watched the movie I read the Wikipedia entry, which mentioned various interpretations. It also mentioned that the sequels were not as good. So I looked at them (couldn’t finish 3 and 4). All the while I was thinking.

I started comparing Jaws 2 with the original. The original was much tighter. I started thinking about the role of Quint. “Why did he have to die?” I asked myself. Sure, he wasn’t a nice man, but he hadn’t hurt anyone. All that had to happen to give the movie a “happy” ending was that the shark gets killed. It wasn’t necessary to have anyone else die. So, what’s the point of having Quint die? How would the film be different if he hadn’t been killed by the shark?

It was at that point that it hit me: Girard, sacrifice, mimetic desire. NOW I had a specific question: How do Girard’s ideas apply to Jaws? It took me awhile to come up with an answer that I liked. But that’s how it goes. The really important thing was coming up with the idea to apply Girard’s ideas to the movie. That’s what took the hard work.

These benchmarks don’t test that kind of work at all. But that’s where the really hard and fruitful problems lie.

You’ve pinpointed a profound challenge in evaluating AI: life—and much of human intellectual labor—is far messier and more open-ended than what benchmarks can capture. Your example of watching Jaws illustrates exactly the kind of complex, emergent thinking that is central to creative, critical, and interpretive work. Let’s unpack your critique a bit further.

Sunday, December 22, 2024

Competence and Performance in Bench-marking LLMs

I’ve been hearing complaints about the inadequacy of LLM benchmarks for two years now. So let’s think a bit. Consider this passage from Rodney Brooks, [FoR&AI] The Seven Deadly Sins of Predicting the Future of AI (Sept. 7, 2017):

One of the social skills that we all develop is an ability to estimate the capabilities of individual people with whom we interact. It is true that sometimes “out of tribe” issues tend to overwhelm and confuse our estimates, and such is the root of the perfidy of racism, sexism, classism, etc. In general, however, we use cues from how a person performs some particular task to estimate how well they might perform some different task. We are able to generalize from observing performance at one task to a guess at competence over a much bigger set of tasks. We understand intuitively how to generalize from the performance level of the person to their competence in related areas.

When in a foreign city we ask a stranger on the street for directions and they reply in the language we spoke to them with confidence and with directions that seem to make sense, we think it worth pushing our luck and asking them about what is the local system for paying when you want to take a bus somewhere in that city. If our teenage child is able to configure their new game machine to talk to the household wifi we suspect that if sufficiently motivated they will be able to help us get our new tablet computer on to the same network.

If we notice that someone is able to drive a manual transmission car, we will be pretty confident that they will be able to drive one with an automatic transmission too. Though if the person is North American we might not expect it to work for the converse case.

He then gives a few more examples.

So, consider benchmarks. Many of them are standardized tests devised to discriminate among humans. Any human capable of taking one of these tests – such as the Advanced Placement test for physics, the law bar exam, a standard set of programming problems – is capable of performing other tasks in the domain, tasks which may not be amenable to testing procedures, but the test themselves are sufficient to sort individuals.

Consider the bar examination. Practicing lawyers have to be able to meet with clients, take depositions, negotiate with other lawyers, and appear in court. None of these things are tested in a bar exam, but lawyers have to do them. To use Brooks’ terms, an individual’s performance on the bar exam does not test their overall competence as a lawyer. It would be a mistake to assume that an LLM’s performance on the bar exam is an indication of its overall lawyerly competence.

Furthermore, all tests present the test-taker with a well-defined situation to which they must respond. But life isn’t like that. It’s messy and murky. Perhaps the most difficult a person has to do is to wade into the mess and murk and impose a structure on it – perhaps by simply asking a question – so that one can then set about dealing with that situation in terms of the imposed structure. Tests give you a structured situation. That’s not what the world does.

Consider this passage from Sam Rodiques, “What does it take to build an AI Scientist”:

Scientific reasoning consists of essentially three steps: coming up with hypotheses, conducting experiments, and using the results to update one’s hypotheses. Science is the ultimate open-ended problem, in that we always have an infinite space of possible hypotheses to choose from, and an infinite space of possible observations. For hypothesis generation: How do we navigate this space effectively? How do we generate diverse, relevant, and explanatory hypotheses? It is one thing to have ChatGPT generate incremental ideas. It is another thing to come up with truly novel, paradigm-shifting concepts.

Right.

How do we put an LLM, or any other AI, out in the world where it can roam around, poke into things, and come up with its own problems to solve? If you want AGI in any deep and robust sense, that’s what you have to do. That calls for real agency. I don’t see that OpenAI or any other organization is anywhere close to figuring out how to do this.

Thursday, November 28, 2024

Large language models surpass human experts in predicting neuroscience results

Luo, X., Rechardt, A., Sun, G. et al. Large language models surpass human experts in predicting neuroscience results. Nat Hum Behav (2024). https://doi.org/10.1038/s41562-024-02046-9

Abstract: Scientific discoveries often hinge on synthesizing decades of research, a task that potentially outstrips human information processing capacities. Large language models (LLMs) offer a solution. LLMs trained on the vast scientific literature could potentially integrate noisy yet interrelated findings to forecast novel results better than human experts. Here, to evaluate this possibility, we created BrainBench, a forward-looking benchmark for predicting neuroscience results. We find that LLMs surpass experts in predicting experimental outcomes. BrainGPT, an LLM we tuned on the neuroscience literature, performed better yet. Like human experts, when LLMs indicated high confidence in their predictions, their responses were more likely to be correct, which presages a future where LLMs assist humans in making discoveries. Our approach is not neuroscience specific and is transferable to other knowledge-intensive endeavours.