GMAsia
    AI News·14 Sept 2026·via The New Stack

    AI’s best coding agent fails 60% of the time — and the data backs it up

    Claude Fable 5.1 achieved a new coding benchmark, marking a notable performance in AI agent capabilities. Despite its advanced design, the AI agent demonstrated a 38.8% success rate on the test. This means it failed more than six out of ten times when attempting coding tasks. The data from the benchmark confirms the agent's performance, providing a clear picture of its current effectiveness in real-world scenarios. This outcome offers critical insights into the present state of AI in software development.

    Nexa's Summary

    The core finding is that even the best AI coding agents fail 60% of the time. Claude Fable 5.1’s 38.8% success rate on a new benchmark shows current limitations. This is not a story about AI’s triumph, but about its persistent challenges in practical coding applications. Developers still face significant debugging overhead.

    For Asian tech companies, this means reliance on AI coding tools must be tempered with realistic expectations. Firms investing in developer productivity tools, especially in markets like India and Vietnam, should plan for substantial human oversight. The promise of fully autonomous coding remains distant, impacting development timelines and costs.

    The thing to watch is how quickly these failure rates improve over the next 12 to 18 months. If success rates do not climb significantly past 50%, adoption in critical enterprise environments will slow. Companies in Seoul and Singapore will continue to prioritize human-in-the-loop AI integration.

    #ai engineering#software testing#ai models
    Go deeper
    Original reporting by The New StackWe don't republish, read the full story →

    Related reading

    6 stories