Showing posts with label Ash_Jogalekar. Show all posts
Showing posts with label Ash_Jogalekar. Show all posts

Sunday, July 26, 2026

NYTimes: AI needs human supervision in order to complete an entire job.

From the NYTimes article linked in the tweet:

We gave an A.I. tool full access to a laptop with pre-configured apps and sought to answer a simple question: Can artificial intelligence do an office job?

Some corporate executives seem to believe it can. More than 200 tech companies have cut roughly 120,000 jobs this year, according to Layoffs.fyi, an industry tracking site; Meta, Oracle and others have all recently made substantial cuts to their work forces, citing A.I. as the driving force; and after laying off about 1,100 employees, the chief executive of Cloudflare said recently that he expected A.I. to replace workers in middle management, finance and marketing.

tweIn our experiment, we deployed A.I. “agents” to act as office workers and found that they were capable of performing some of the tasks we assigned, but not all of them. The agents, which can act autonomously and make decisions based on detailed instructions, excelled at problems they could solve by writing computer programs. But they struggled with understanding the nuances of human language and at navigating user interfaces like the Chrome web browser.

The article then has a series of nice quasi-interactive displays illustrating agent performance on three tasks. The displays include screen shots of various messages and documents.

About the tasks:

This task, and the others we assigned to the A.I., were adapted from papers and benchmarking tools published recently by researchers at Carnegie Mellon University and OpenAI. The researchers designed the benchmarks to test the performance of various models — like OpenAI’s GPT, Google’s Gemini and Anthropic’s Claude — in real-world environments, and compare them with one another.

General conclusion:

The results of our experiment roughly matched what researchers and companies have found as they have tested and used artificial intelligence tools. Scale AI, an A.I. training company, recently tested agents on real freelance projects, and the best-scoring model produced client-ready work only about 16 percent of the time.

While A.I. can excel regularly at complex tasks, it can be unreliable when put in charge of an entire job. It can certainly add value to certain areas of the work force, but for now, A.I. still needs a human boss.

* * * * *

Comment: Around the corner my colleague, Ash Jogalekar, has tweets like this one:

So here's a great example of where we are with agentic AI: Instead of just being an assistant, it's behaving more like a collaborator and creative scientist.

In a recent project, I gave the system a molecular design problem typical of the problems we encounter in chemistry. Two similar molecules were giving very different results.

He then runs through an account of what his AI collaborator did, concluding:

I think we have crossed the Rubicon. Agentic AI now no longer just processes tasks and automates workflows blindingly fast, but it can generate hypotheses, test them, test counter-hypotheses and go back and forth and course-correct if necessary, all with minimal to no human intervention. It's now embodying the general scientific method.

[I've copied another one of Ash's tweets to this post, A scientist reflects on what AI has done for him.]

What’s interesting to me, and very revealing, is that a complex set of tasks in scientific investigation seems to be on a level with routine office tasks, as though one were no more complex than the other. But humans require years of college education in order to perform the former while the latter requires no more than a high school education, if that. It seems that once they’ve been been learned and compiled, all tasks or sets of tasks are on the same “level” in the brain. The educational prerequisites required to do such tasks for the first time or three get “compressed out” through repetition. Since AIs are trained on written records of what humans have said and done, they don’t have to go through the ordinary learning process. The compression has already taken place and is present in the documents on which they are trained.

Thursday, July 23, 2026

A scientist reflects on what AI has done for him – “a Rubicon has been crossed”

Here’s the full content of a tweet by Ash Jogalekar:

I came to the present AI revolution not as a credulous enthusiast, but as someone deeply skeptical of new technologies in science. For twenty-five years I have seen too many of them arrive surrounded by extravagant claims before settling into a useful but much more modest place in the scientific toolkit.

What has astonished me is that the latest agentic systems appear to represent a qualitative change. I have now run upward of two hundred scenarios and AI for science workflows, each one navigating a complex, multistep scientific protocol across diverse fields of chemistry, biology and materials science, and I think I have enough data now to make a credible judgment. Over just a few months I have seen these models and algorithms leapfrog over increasing levels of difficulty, starting almost as a toddler and turning into an adult interlocutor. Every day, something moves my baseline for what they can do. They autonomously install, run, and debug dozens of computational tools; clean and structure data; parameterize molecules; launch calculations; and manage complicated, multistage investigations. But as it turned out, that was just the beginning.

More fascinatingly, they increasingly display recognizable scientific judgment: proposing positive and negative controls, discovering that a method does not work, generating competing hypotheses, systematically eliminating them, changing direction when evidence contradicts an initially promising idea, searching an entire target space, constructing unexpectedly sophisticated models, and finding useful analogies across distant fields. This is no longer merely workflow automation or doing the same science faster. It is beginning to feel like genuine intellectual and creative collaboration.

Interacting with the system can resemble a conversation with a smart and experienced student or colleague: ideas are proposed, challenged, refined, rejected, and unexpectedly pushed in new directions, with both the human and the AI acknowledging mistakes and adjusting course. The exhilarating possibility is the sheer multiplication of intellectual labor - the ability to explore almost any question under the scientific sun and to see those “mountains beyond mountains”, as Tracy Kidder eloquently put it. In a week a scientist or a team can come up with dozens or hundreds of ideas and hypotheses, most of them reasonable and actionable. The physical lab is now the only bottleneck, and even that is being accelerated by AI.

As a scientist, AI has made me feel more intellectually alive and excited than I have felt since graduate school and my postdoctoral years more than two decades ago. I can begin with an idea in the morning and, by lunchtime, watch a rational, testable hypothesis take shape; within days, an investigation can progress from literature and classical calculations to increasingly rigorous quantum-mechanical analysis and experimentally actionable predictions. Eating and sleeping have often taken a backseat, exercise seems like a distant goal, and every night I feel like I did when I was a kid and begging my dad or mom for just *one more story*. Except that this time it’s just *one more prompt*. One more cycle of compute. One more result that will startle or confound. Every night I have to force myself to detach from the computer, and on more than one occasion I have fallen asleep at my desk, only to wake up and see the potential for yet *one more prompt*.

Precisely because these systems are beginning to criticize our assumptions, tell us when something does not work, and think alongside us rather than merely obey us, it feels as though we may already have passed beyond the first age of AI.

Of course, these predictions and results will stand or fall based on experimental testing, that’s how science always has been, but that’s no different from the pre-AI age. More importantly, in almost every case they appear as conclusions that any good scientist will regard as reasonable, at least as starting points. And sometimes they genuinely throw up a surprise. AI-enabled science should still be judged by the novelty, rigor, reproducibility, statistical validation, and epistemic integrity of the science, not by the novelty of the technology.

But there is no doubt now that a Rubicon has been crossed, and either we cross over to the other side or get left behind. What a time to be alive.