This paper is a brutal reality check for long-horizon AI. Give an agent a year of interconnected decisions, delayed feedback, and consequences from its own past actions, and its performance collapses relative to humans.
— Rohan Paul (@rohanpaul_ai) August 28, 2026
The researchers tested eight leading models, including… https://t.co/rGBfuPZEtv pic.twitter.com/6CPvVIQCQX
Saturday, August 29, 2026
AIs are not very good at long horizon tasks
Subscribe to:
Post Comments (Atom)
No comments:
Post a Comment