GMAsia
    AI News·9 Sept 2026·via The New Stack

    Claude did best on a new benchmark for agents that build agents. It still passed fewer than a quarter of the tests.

    Anthropic's Claude AI model achieved the top performance on a new benchmark designed to test the capabilities of 'agents that build agents'. Despite its leading score, Claude successfully passed fewer than 25% of the tests in this specialized evaluation. The benchmark assesses the ability of AI agents to autonomously create and refine other AI agents, a critical step toward more sophisticated and self-improving AI systems. This development highlights both the progress in AI agent development and the significant challenges that remain in achieving robust autonomous agent creation. The New Stack reported on these findings, emphasizing the current limitations even among the best-performing models.

    Nexa's Summary

    The benchmark results for 'agents that build agents' show that even the top-performing AI, Anthropic's Claude, only passed a small fraction of tests. This indicates that while progress is being made in AI's ability to autonomously develop other AI systems, the technology is still in its early stages. For Asian tech companies, particularly those investing heavily in AI research and development, this suggests that fully autonomous AI agent creation is still a distant goal, requiring continued significant investment in foundational AI research rather than immediate application. The observed limitations mean that human oversight and intervention will remain crucial in AI development pipelines for the foreseeable future, impacting resource allocation and skill requirements across the region. The focus should remain on incremental advancements and robust testing, as the current state of the art still has substantial room for improvement.

    #large language models#ai#ai agents#ai engineering
    Go deeper
    Original reporting by The New StackWe don't republish, read the full story →

    Related reading

    6 stories