The Long-Horizon Problem: Why Capable Agents Still Cannot Finish Long Tasks
Systems that handle a three-step task competently fall apart on a thirty-step version of the same task, and the reason is closer to arithmetic than to intelligence. Research through 2026 has moved past the naive compounding-error model toward something more diagnostic — and more difficult.

