Giao diện
TeguNews
Công nghệ

149 Pages Mapping the Long-Horizon Agent Frontier: Multi-University Survey Proposes Harness Engineering and Model Optimization as Two Main Evolution Lines for Next-Generation AI Agents

Renmin University GAIR leads multi-institution 149-page survey on long-horizon agents, proposing H1-H3 task difficulty hierarchy and C1-C3 capability tiers, with task span doubling every 4-7 months.

Pandaily (China Startup/AI)1 phút đọc

149 Pages Mapping the Long-Horizon Agent Frontier: Multi-University Survey Proposes Harness Engineering and Model Optimization as Two Main Evolution Lines for Next-Generation AI Agents

The 149-page survey, Towards Long-Horizon Agents: A Survey, released by the Renmin University GAIR Research Center in collaboration with Peking University, Tsinghua University, Sun Yat-sen University, Hong Kong University of Science and Technology, and National University of Singapore, provides the first systematic framework for understanding the emerging field of long-horizon AI agents. The survey organizes the field around two parallel evolution lines: externalized Harness Engineering for runtime capability enhancement and internalized Model Optimization that progressively writes execution capability back into the model through training.The survey identifies the defining characteristic as the ability to organize interdependent decisions into coherent trajectories.

The core contribution is a three-tier task difficulty and corresponding capability hierarchy. H1 Window-Level tasks completed within a single context window require sustained plan-act-feedback-correct loops, corresponding to C1 Interactive Reasoning capability. H2 Cross-Window tasks spanning hours to days require history compression, progress persistence, and cross-session handover beyond what expanded context windows alone can provide, with Context Rot identified as a fundamental bottleneck, corresponding to C2 State and Memory.

H3 Cross-Task-Flow tasks where tasks arrive continuously with changing environments, tools, and objectives require experience accumulation into reusable skills and continuous improvement, corresponding to C3 Experience Accumulation extending to continuous learning and self-evolution.The survey provides quantitative evidence of capability advancement using METR public data. The 50% task completion time span, measuring the human expert time equivalent for tasks the agent completes with approximately 50% success rate, shows a doubling every 196.

5 days (approximately 7 months) across the full dataset, accelerating to approximately 130.8 days (4 months) when focusing on the post-2023

Đọc thêm từ Công nghệ