Friday, August 7, 2026

François Chollet sees Large Reasoning Models (LRMs) in the future

What are large reasoning models?

2 comments:

  1. Follett may have accidentally created Rorschach blots for AI... "Chollet created 1,000 ARC tasks, and released 800 of them." From Melanie Mitchell's "Why the Abstraction and Reasoning Corpus is interesting and important for AI' including links and images.

    Follet's prior...
    "... Algorithmic Information Theory, ... Using this definition, we propose a set of guidelines for what a general AI benchmark should look like. Finally, we present a benchmark closely following these guidelines, the Abstraction and Reasoning Corpus (ARC), built upon an explicit set of priors designed to be as close as possible to innate human priors. We argue that ARC can be used to measure a human-like form of general fluid intelligence and that it enables fair general intelligence comparisons between AI systems and humans."
    last revised 25 Nov 2019 (this version, v2)]
    "On the Measure of Intelligence"
    François Chollet
    https://arxiv.org/abs/1911.01547

    "Is AI Reasoning Right for the Wrong Reasons?
    ByJOHN PAVLUS
    July 31, 2026
    "The idea that artificial intelligence can “reason” is more intuitive than ever. But intuitions can be wrong, and the science is far from settled.
    ...
    "Melanie Mitchell’s career in AI stretches back to the 1980s, ...
    "“Number one: It works. It improves things,” she said, referring to LRMs’ superior accuracy on reasoning tasks compared to LLMs. “Number two: The actual text that’s generated” — i.e., the chain of thought that every LRM is trained to produce to improve its performance — “isn’t necessarily faithful to what’s going on [inside the model]. And number three: A lot of that text isn’t even useful. You can actually take it out.”
    ...
    " Subbarao Kambhampati calls “mumblings” — bits of language, yes, but ones whose meaning may be entirely incidental to any reasoning that might have occurred. Kambhampati’s lab showed in 2025 that fully replacing a model’s correct “traces” with incorrect or irrelevant ones didn’t degrade its performance on a formal reasoning task. Meanwhile, training the model only on correct trace data still led it to occasionally generate invalid records of its reasoning — even when it produced a correct solution to the original problem it was given. A 2024 paper from researchers at New York University showed that “meaningless filler tokens” — literally, strings of dots — could function effectively in place of a human-readable “chain of thought.”
    William Merrill, one of the authors on that paper ..., put the matter plainly: “There’s no guarantee the chain of thought has to be meaningful in any sense.” 
    ...
    https://www.quantamagazine.org/is-ai-reasoning-right-for-the-wrong-reasons-20260731/

    "Why the Abstraction and Reasoning Corpus is interesting and important for AI
    Melanie Mitchell
    Mar 01, 2023
    ...
    "There has been considerable research on getting AI systems to make abstractions and analogies, a lot of it using idealized domains such as Raven’s Progressive Matrices or Bongard Problems. My 2021 paper Abstraction and Analogy in Artificial Intelligence surveys some of this work.
    The Abstraction and Reasoning Corpus
    In this post I want to talk about one specific idealized domain, the Abstraction and Reasoning Corpus, created by François Chollet. I’ll talk about what the domain is, why I think it’s particularly interesting as a research program for AI, and why it’s still quite challenging, and perhaps impossible, for current AI systems, including everyone’s favorite large language models. 
    ...
    Chollet created 1,000 ARC tasks, and released 800 of them.   Two hundred tasks were reserved as a “hidden” test set.

    "In his paper describing the goals of ARC, Chollet cites the work of developmental psychologist Elizabeth Spelke on the “core knowledge systems” that are either innate or developed early on in humans (and likely some non-human animals).  These core systems include:
    ...
    https://aiguide.substack.com/p/why-the-abstraction-and-reasoning

    SD

    ReplyDelete