New benchmark reveals top AI agents fail catastrophically at real-world work
New research shows leading AI models fail catastrophically, completing less than three percent of realistic remote-work projects.
June 19, 2026

The promise of artificial intelligence completely automating white-collar work has hit a formidable roadblock. While Silicon Valley tech giants continue to pitch a future run by autonomous digital agents capable of managing complex business operations, a newly developed industry benchmark has delivered a sobering reality check. According to a comprehensive evaluation, even the most advanced generative AI models and agentic frameworks fail catastrophically when tasked with executing real-world knowledge work. The research reveals that the very best artificial intelligence systems on the market can fully resolve less than three percent of realistic remote-work tasks, highlighting a massive gap between controlled laboratory success and actual economic value[1].
This rigorous new evaluation is known as the Remote Labor Index, a collaborative benchmarking project designed by Scale AI and the Center for AI Safety[2]. Unlike traditional academic datasets or standardized coding tests that rely on curated, textbook-style problems, the Remote Labor Index evaluates artificial intelligence against the exact demands of the modern freelance economy[3]. The creators of the index sourced hundreds of real-world projects directly from online freelance platforms, capturing the messy, unscripted reality of paid remote labor[3]. The final dataset consists of hundreds of active projects spanning diverse domains, including software engineering, data analysis, graphic design, game development, video editing, and architectural drafting[3][4]. In total, these projects represent thousands of hours of actual human labor and substantial monetary value, establishing a highly authentic baseline for testing AI capability[5].
The results of the evaluation are striking in their uniformity: current artificial intelligence systems perform near the absolute floor[2]. Graded manually by human experts who asked whether a paying client would accept the delivered work, the best-performing agent framework achieved an automation rate of just two and a half percent[6][5]. This top-tier performance was achieved by Manus, a prominent agentic startup that was recently acquired by a major technology parent company[7][6]. Other highly touted frontier models fared even worse. Advanced systems, including the latest iterations of Claude Sonnet, Grok, and GPT, registered success rates hovering between one and two percent, while other leading enterprise models successfully automated less than one percent of the assigned projects[8][9]. Out of thousands of dollars in potential earnings represented by the benchmarked projects, the collective efforts of the world's most advanced artificial intelligence models captured only a tiny fraction of the total economic value[6].
The primary driver of this widespread failure is the immense complexity of executing end-to-end multi-step projects[5]. In a typical freelance task, a human worker is expected to parse a loose project brief, manage files, troubleshoot unexpected technical errors, and synthesize different media formats[10][5]. Artificial intelligence models, by contrast, struggle immensely with long-horizon planning and lack persistent memory[10][5]. Because these models do not truly learn on the job, they cannot maintain consistency over tasks that require a dozen hours of active processing[10]. A minor mistake made by an AI agent in the early phases of a project will compound over time, leading to a domino effect of errors that ultimately results in broken files, missing features, and unusable deliverables[8][5]. Graders frequently reported that AI agents delivered projects that looked complete at first glance but were plagued by underlying structural issues that would never pass professional scrutiny[8][5].
Furthermore, the benchmark highlights a glaring discrepancy between academic AI capabilities and real-world utility. Tech companies frequently boast of models scoring above eighty percent on standardized software engineering benchmarks or acing graduate-level exams[5]. However, these tests are highly structured, rely on predictable patterns, and are often contaminated, meaning the models may have already encountered the questions during their initial training phases[11]. The Remote Labor Index bypasses this benchmark illusion by utilizing novel, multi-sector tasks that require genuine reasoning and multimodal output[5]. To successfully complete an architectural drawing or a customized data dashboard, an agent must manipulate CAD tools, coordinate spatial layouts, and correctly interpret subjective client preferences[10][5]. Current large language models are fundamentally mimicry engines; they excel at predicting the next word or mimicking human coding patterns, but they lack the cognitive architecture required to reason through novel, unstructured physical or digital tasks[12].
Industry experts and researchers are divided on what these findings mean for the trajectory of artificial intelligence. Critics of the current generative AI boom point to the dismal performance as evidence that the technology has hit a critical capability wall, suggesting that massive labor displacement is far further in the future than corporate marketing suggests[13]. Other analysts argue that evaluating general-purpose agents on such complex tasks is a flawed approach[8]. They contend that the path to real economic automation lies not in generalist models, but in highly specialized, fine-tuned agents trained specifically for narrow, well-defined corporate roles[8]. Technology developers are already shifting their strategies, moving away from broad assistant tools and toward heavily structured workflows where human supervisors remain firmly in the loop to catch and correct the compounding errors that derail autonomous models.
Ultimately, the findings of this benchmark provide a crucial, data-driven foundation for the ongoing debate over the future of labor[2]. The dramatic gap between expectations and reality suggests that human knowledge workers possess a subtle but powerful suite of skills—adaptability, systemic reasoning, and quality control—that remains safely out of reach for digital agents. Rather than predicting a sudden, disruptive wave of immediate unemployment, the transition to an automated economy is shaping up to be a slow, highly iterative process. As artificial intelligence continues to evolve, benchmarks like the Remote Labor Index will serve as vital yardsticks, tracking whether these systems can transition from impressive conversational novelties into truly reliable, economically productive digital colleagues[2][14].
Sources
[2]
[3]
[4]
[5]
[6]
[7]
[10]
[11]
[13]
[14]