149 Pages Mapping the Long-Horizon Agent Frontier: Multi-University Survey Proposes Harness Engineering and Model Optimization as Two Main Evolution Lines for Next-Generation AI Agents
Renmin University GAIR leads multi-institution 149-page survey on long-horizon agents, proposing H1-H3 task difficulty hierarchy and C1-C3 capability tiers, with task span doubling every 4-7 months.

The 149-page survey, Towards Long-Horizon Agents: A Survey, released by the Renmin University GAIR Research Center in collaboration with Peking University, Tsinghua University, Sun Yat-sen University, Hong Kong University of Science and Technology, and National University of Singapore, provides the first systematic framework for understanding the emerging field of long-horizon AI agents. The survey organizes the field around two parallel evolution lines: externalized Harness Engineering for runtime capability enhancement and internalized Model Optimization that progressively writes execution capability back into the model through training.The survey identifies the defining characteristic as the ability to organize interdependent decisions into coherent trajectories.
The core contribution is a three-tier task difficulty and corresponding capability hierarchy. H1 Window-Level tasks completed within a single context window require sustained plan-act-feedback-correct loops, corresponding to C1 Interactive Reasoning capability. H2 Cross-Window tasks spanning hours to days require history compression, progress persistence, and cross-session handover beyond what expanded context windows alone can provide, with Context Rot identified as a fundamental bottleneck, corresponding to C2 State and Memory.
H3 Cross-Task-Flow tasks where tasks arrive continuously with changing environments, tools, and objectives require experience accumulation into reusable skills and continuous improvement, corresponding to C3 Experience Accumulation extending to continuous learning and self-evolution.The survey provides quantitative evidence of capability advancement using METR public data. The 50% task completion time span, measuring the human expert time equivalent for tasks the agent completes with approximately 50% success rate, shows a doubling every 196.
5 days (approximately 7 months) across the full dataset, accelerating to approximately 130.8 days (4 months) when focusing on the post-2023
Đọc thêm từ Công nghệ

Samsung Electronics, Broadcom sign AI chip partnership
Samsung shipped commercial HBM4 and is expanding capacity for AI computing.

SK hynix eyes record Q2 profit on AI memory boom
SK hynix held a 57% share of the high-bandwidth memory market for AI servers.

Tên lửa Trung Quốc bị sét đánh ngay sau khi cất cánh
Tên lửa Trường Chinh 3B bị sét đánh trúng trong lúc bay lên quỹ đạo, song vẫn triển khai thành công một vệ tinh viễn thông.
Humans Haven't Stopped Evolving
Article URL: https://www.harvardmagazine.com/research/harvard-human-evolution-genes-selective-pressure Comments URL: https://news.ycombinator.com/item?id=49054307 Points: 6 # Comments: 0