GMAsia
    AI News·2 Oct 2026·via Techtimes

    Gemini 4 Argon Leads on Knowledge Work: Where It Still Trails Before You Commit

    Google DeepMind has launched Gemini 4 Argon, a new AI model that reportedly outperforms GPT-6 Astra and Claude Opus 5.5 on 13 of 19 internal benchmarks. The model, unveiled on September 30, 2026, features a 1 million output token ceiling and an introductory API price of $2 per million input tokens.

    Nexa's Summary

    Gemini 4 Argon's reported strengths are concentrated in enterprise knowledge work and long-context retrieval. It shows a lead on benchmarks like the Vals Index, which weights performance across finance, coding, legal, and tax work by each sector's contribution to US GDP. Its ability to process and retrieve information across prompts up to 1 million tokens, reflected in a 12.4-point lead over Astra on GraphWalks BFS F1, is significant for applications requiring extensive document or codebase analysis.

    Despite these leads, Argon trails competitors in specific areas. It falls behind GPT-6 Astra by 10.5 percentage points on the harder of two published software engineering evaluations, and Claude Opus 5.5 by nine points on terminal-driven agent work. These are critical deficits for teams building shell-based automation pipelines, where AI agents must drive a terminal or shell environment over multiple steps.

    A key financial uncertainty for potential adopters is the actual per-task cost at scale. Google has not confirmed whether the model's internal reasoning tokens bill at the output rate. This means that while the input token price is known, the full cost for complex operations requiring extensive internal processing remains undefined, impacting budget predictability for large-scale deployments.

    Share this article

    Go deeper
    Original reporting by TechtimesWe don't republish, read the full story →

    Related reading

    6 stories