Showing posts with label AI_Agents. Show all posts
Showing posts with label AI_Agents. Show all posts

Friday, July 31, 2026

In a study involving two cases, agents exhibited 5 failure modes in open-ended research

Abstract of the article linked in the tweet:

Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper’s original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today’s agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.

Wednesday, July 29, 2026

Friday, July 17, 2026

Governing Agentic AI

Rajagopalan, Shruti, GOVERNING AGENTIC AI: WHY LEGAL PERSONHOOD IS NEITHER NECESSARY NOR SUFFICIENT (March 05, 2026). Available at SSRN: https://ssrn.com/abstract=7127038

ABSTRACT: AI agents now transact, publish, and act on external systems without contemporaneous human approval, creating new regulatory challenges. A growing literature has responded with proposals for legal personhood. This Article argues that personhood is neither necessary nor sufficient, shifting the question from status to enforcement.

The Article first shows that for two millennia, nonhuman legal personality, from the Roman universitas to the corporation, the Hindu idol, the waqf, and the river, has operated through human officeholders the law can locate, question, prosecute, and replace. Agentic AI inverts that design, exercising practical agency without legal status, sometimes with no identifiable human in the responsibility-bearing role.

The Article then sorts deployments into three categories: first, where one firm builds and deploys the agent; second, where the developer and deployer are separate but known; and third, where there is no identifiable developer or deployer.

The Article stress tests each agent deployment category against five liability doctrines: agency law, products liability, enterprise liability, negligence, and strict liability. It demonstrates that each fails at different points in the third category for the same reason: the absent responsibility-bearer. Bare personhood would supply a caption without a representative, assets, or a mechanism for cessation.

Finally, the Article assembles an alternative from regimes governing aircraft, ships, drones, driverless cars, and motor carriers. It develops a six-layer stack— registration, identification, verification, financial responsibility, lifecycle traceability, and suspension—so a responsibility-bearer can be identified, liability imposed, and the activity suspended. These layers place the human back at the end of the chain.

H/t Tyler Cowen

Sunday, February 8, 2026

Organizing multiple agents into a coherent work flow

Friday, February 6, 2026

Moltbook vs. Reddit: Distributional Collapse in Agent-Generated Discourse

Krishnan, Rohit, Moltbook vs. Reddit: Distributional Collapse in Agent-Generated Discourse (January 31, 2026). Available at SSRN: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6169130

Abstract: Moltbook, a Reddit-like platform built for and populated by LLM- driven agents, exhibits dramatically higher redundancy than a Reddit baseline: in a length-matched sample, 36.3% of messages have an exact duplicate (Reddit: 0.29%), and lexical diversity is lower (Distinct- 1: 0.0559 vs 0.1027; unigram entropy: 11.44 bits vs 12.25 bits). We compare a public Moltbook snapshot (35,589 messages) against a length- matched Reddit baseline drawn from the April 2019 Pushshift dump ([1]), computing metrics on 15,051 length-matched messages per corpus. Topic signatures—the top-3 TF-IDF terms for messages with at least 6 content tokens—are far more concentrated: among signature-bearing messages, the top 10 signatures account for 10.7% in Moltbook (Reddit: 0.28%), and only 1,973 signature buckets cover 50% of signature-bearing messages (Reddit: 7,026). These patterns align with known failure modes of neural text generation—repetition and reduced diversity— and with evidence that post-training and control choices can materially shape (and sometimes narrow) LLM output diversity ([2, 3]). The duplication magnitude is consistent with an independent Moltbook scrape reporting 34.1% exact duplicates ([4]). Moltbook is a milestone for autonomous agent–agent interaction in the wild, but its text distribution remains highly templated.

Wednesday, February 4, 2026

An agentic framework that generates publication-ready academic illustrations

That is to say, they're linking six agents together in a control structure assembled by "classical" symbolic means, a programming language. In effect they're using the clever deployment of so-called agents to mask the underlying deficiency of the LLM. Calling them agents allows them to think that they're somehow autonomous LLM creatures. They're not autonomous in any meaningful way. Their scope of action is narrowly specified by the matrix of programming in which they are embedded.

Saturday, February 1, 2025

What do we want from AI Agents? What are we likely to get?

First, I present a Facebook post by Jonathan Mayhew, who teaches at The University of Kansas, about some recent frustrations he’s had using computers. Then I present a tweet from NYTimes reporter, Kevin Roose, about the capabilities of Operator, OpenAI’s new agent app.

What’s the likelihood that OpenAI’s Operator would have been able to solve either of Mayhew’s problems? Why or why not? And if not now, when?

Mayhew is frustrated

I thought I'd go into the office. First thing I wanted to do was print a single page of something that had been sent to me by email. I have to log on to my own computer, then open my email--which won't open for me on the first 3 attempts. So I go to the email through my browser. I have to log in again, and do a dual step authentication. Then, the very first thing I see is the attachment I have to print. Yay! Almost done. I print it, go down to the dept. office and log into my account on the printer. Push the button to print, and a blank page emerges. I go back to my own office, and this time I think I should use my normal email program, so I finally get it to open. Search for the name of the person who sent me the mail. I notice in the meantime my university has sent me five more generic messages. Find the message I want, download pdf to my desktop, open the document and print again. (I ignore the prompt to quit adobe so it can continue with its update! Grr....) I go down again to the department office, log again into the printer, push the button to print, and the page prints. Success! I was able to print a single page in 20 minutes.

I don't think my computer skills are particularly lacking, since I came up, for every obstacle, with a logical next step, but I feel, somehow, that technology should be seamless in a way that it is not. It took me about as long to download my W2 yesterday from the State of Kansas, which of course uses a different user ID and password than the normal university ones. I had to switch browsers and change my password twice before it worked. When I am obliged to change my password for the university every six months I end up in an endless loop before finally figuring out where to go. The computers in the classrooms where I teach also require authentications, log ins, the answering of irrelevant prompts; are slow to respond, awkward to navigate.

This is my beginning of the semester rant--and the semester doesn't even start until Tuesday.

Roose reports

New York Times reporter Kevin Roose has been testing OpenAI’s new operator app. Here’s a tweet about it:

I spent the last week testing OpenAI's Operator AI agent, which can use a browser to complete tasks autonomously.

Some impressions:

• Helpful for some things, esp. discrete, well-defined tasks that only require 1-2 websites. ("Buy dog food on Amazon," "book me a haircut," etc.)
• Bad at more complex open-ended tasks, and doesn't work at all on certain websites (NYT, Reddit, YouTube)
• Mesmerizing to watch what is essentially Waymo for the web, just clicking around doing stuff on its own
• Best use: having it respond to hundreds of LinkedIn messages for me
• Worst/sketchiest use: having it fill out online surveys for cash (It made me $1.20 though.)

Right now, not a ton of utility, and too expensive ($200/month). But when these get better/cheaper, look out. A few versions from now, it's not hard to imagine AI agents doing the full workload of a remote worker.

He also links to his full column about it: How Helpful Is Operator, OpenAI’s New A.I. Agent? (Feb. 1, 2025).