We value your privacy

    We use cookies to analyse traffic, improve our website and show relevant content. You choose what we may use. Read our privacy policy.

    Live AI news
    Back to articles// AI Trends

    Digital labor automation quadruples in under eight months

    The newly published Remote Labor Index shows AI systems successfully completing up to 16.1 percent of freelance knowledge tasks across a deeply jagged frontier.

    ai.nl editorial Published 6 juli 2026 6 min read
    Abstract 3D render of rising blue bars visualising accelerating AI automation rates

    On July 1, 2026 the Center for AI Safety and Scale Labs published new results for the Remote Labor Index (RLI). The benchmark does not measure abstract language skills. It measures how often autonomous AI agents deliver real freelance projects at a quality a paying client would accept, across 3D and CAD, architecture, graphic design, video and animation, audio, data analysis and web applications. The latest round shows the landscape shifting faster than expected: the frontier has more than quadrupled in under eight months.

    The setup is deliberately strict. Human evaluators compare every AI deliverable against work produced by a paid professional. The automation rate is the share of projects where the AI output is judged at least as good as the human baseline. At launch, the strongest model reached 2.5 percent. The previous leader, Opus 4.6 running the Claude Cowork scaffold, sat at 4.17 percent.

    New scores on the board

    The current numbers redraw the picture. OpenAI''s GPT-5.5 climbs to 6.3 percent. Anthropic''s Opus 4.8 reaches 8.3 percent. The undisputed leader is Fable 5 at 16.1 percent, measured across 218 of the 240 benchmark projects. The gap is due to US government restrictions on model access. Even under the worst-case assumption that every missing project failed, Fable 5 still lands at 14.6 percent, higher than any other model. Compared to 2.5 percent at launch, the frontier has more than quadrupled in eight months.

    Serious system access

    The evaluation goes well beyond simple API calls. Each model runs inside the scaffold developers actually use: Anthropic in Claude Code, OpenAI in Codex CLI, both extended with a native computer-use tool (screenshot, click or type, screenshot). Everything executes inside a Linux VM with more than thirty professional applications, including Blender, FreeCAD, GIMP, Inkscape, Kdenlive, Audacity, LibreOffice and the LaTeX suite. Projects get up to 24 hours of wall-clock time, an A100 GPU when the task warrants it, and a worker-critic loop where a second agent audits the deliverable. Budgets are $50 per project, raised to $150 for Fable 5 given its higher per-token pricing.

    The automated judge falls short

    The team also built an automated LLM judge to see whether human evaluation could be replaced. According to the RLI paper, that judge overshoots badly on the newest models: roughly 3× for GPT-5.5 and ~2.5× for Opus 4.8. The reason is structural. Judging an RLI deliverable is itself an agentic task, requiring the evaluator to open files in the right professional application and inspect them the way a client would. Those are precisely the computer-use skills that even frontier models still lack. Encouragingly, the judge does rank models correctly (Spearman ρ = 0.90), which makes it useful for tracking relative progress, but not for absolute capability claims.

    The results also refute time-horizon analysis on this benchmark. The intuition that longer human tasks must be harder for AI does not survive contact with the RLI distribution. This is the jagged frontier at work: some short tasks — transcribing music, playtesting a real-time game — remain out of reach, while long-form work such as digital art or coding is finished in minutes.

    What this means for knowledge work

    Despite the sharp climb, sober framing matters. None of the Fable 5 example deliverables in the study would be accepted as finished professional work without revision. Ring designs look qualitatively better but carry unprofessional details; GPT-5.5''s architectural renders turn out to be image-generated stand-ins rather than usable 3D models. The autonomous stack is improving fast in bulk, but still lacks the last-mile finish a paid professional delivers.

    For knowledge-heavy organisations — consultancies, design studios, engineering firms — the question shifts from "what can AI replace" to "how do we organise the collaboration". A 16 percent fully-automated delivery rate sounds modest, yet it cuts across every discipline RLI covers. The sensible response is not to swap humans for agents (or the reverse), but to design workflows where an agent produces a first pass, a second agent critiques it, and a professional finishes it — under quality criteria that go well beyond what today''s automated graders can see.

    // GET STARTED// How we can help

    Beyond reading — let AI work for you.

    // CONTINUE READINGAll articles

    More from AI Trends.

    Newsletter

    Always up to date on AI.

    Once a month: cases, frameworks and concrete examples of what works in practice. No noise.

    No spam. Unsubscribe any time.