149 Pages Mapping the Long-Horizon Agent Frontier: Multi-University Survey Proposes Harness Engineering and Model Optimization as Two Main Evolution Lines for Next-Generation AI Agents
Renmin University GAIR leads multi-institution 149-page survey on long-horizon agents, proposing H1-H3 task difficulty hierarchy and C1-C3 capability tiers, with task span doubling every 4-7 months.

The 149-page survey, Towards Long-Horizon Agents: A Survey, released by the Renmin University GAIR Research Center in collaboration with Peking University, Tsinghua University, Sun Yat-sen University, Hong Kong University of Science and Technology, and National University of Singapore, provides the first systematic framework for understanding the emerging field of long-horizon AI agents. The survey organizes the field around two parallel evolution lines: externalized Harness Engineering for runtime capability enhancement and internalized Model Optimization that progressively writes execution capability back into the model through training.The survey identifies the defining characteristic as the ability to organize interdependent decisions into coherent trajectories.
The core contribution is a three-tier task difficulty and corresponding capability hierarchy. H1 Window-Level tasks completed within a single context window require sustained plan-act-feedback-correct loops, corresponding to C1 Interactive Reasoning capability. H2 Cross-Window tasks spanning hours to days require history compression, progress persistence, and cross-session handover beyond what expanded context windows alone can provide, with Context Rot identified as a fundamental bottleneck, corresponding to C2 State and Memory.
H3 Cross-Task-Flow tasks where tasks arrive continuously with changing environments, tools, and objectives require experience accumulation into reusable skills and continuous improvement, corresponding to C3 Experience Accumulation extending to continuous learning and self-evolution.The survey provides quantitative evidence of capability advancement using METR public data. The 50% task completion time span, measuring the human expert time equivalent for tasks the agent completes with approximately 50% success rate, shows a doubling every 196.
5 days (approximately 7 months) across the full dataset, accelerating to approximately 130.8 days (4 months) when focusing on the post-2023
Đọc thêm từ Công nghệ

Google Messages : cette fonctionnalité cachée évite les impairs dans vos conversations
Avec des dizaines de conversations en cours, il vous est sans doute déjà arrivé d'envoyer la mauvaise réponse à la mauvaise personne. J'ai trouvé une solution pour éviter ce problème.

Where To Get Safe Eclipse Glasses For Aug. 12’s Solar Eclipse
Here’s how to avoid counterfeit and unsafe eclipse glasses and ensure you buy safe products for Aug’ 12’s partial and total solar eclipse in Europe and North America.

Amazon EKS Adds Kubernetes Version Rollback Within 7 Days of an Upgrade
Amazon EKS has recently introduced support for Kubernetes version rollbacks, letting practitioners revert a cluster's control plane to its previous Kubernetes version within 7 days of an upgrade if issues arise. The feature reduces the risk of in-place cluster upgrades by giving
What links Horrible Histories and Gordon Brown’s neighbour? The bumper summer quiz
From award-winning Julies to golfers who were nearly the best, test your knowledge with this bumper summer quiz1 What is now the world’s most populous city, according to the UN?2 Which two Julies won best actress Oscars in consecutive years?3 What was thrown by the goddess Eris a