GMAsia
    AI News·26 Sept 2026·via Testing Catalog

    OpenAI prepares to expand Ultrafast API to more users

    OpenAI is preparing a broader rollout of its Ultrafast API mode, likely after DevDay on September 29. The new mode offers speeds up to 750 output tokens per second, 14 times faster than Standard. It is currently powered by Cerebras technology and accessible to selected customers. The expansion could introduce tiered pricing, allowing developers to choose latency based on workload value. OpenAI’s 750 MW Cerebras partnership through 2028 supports this infrastructure strategy.

    Nexa's Summary

    OpenAI’s Ultrafast API mode is less about raw speed and more about infrastructure control. Speeds of 750 output tokens per second with GPT-5.6 Sol are impressive. OpenAI is integrating deeper with Cerebras and shifting towards comprehensive compute tiers. This suggests a future where API access is granular, priced by latency, and tied to specific hardware partnerships.

    For Asia’s AI developers, this means new cost-benefit calculations for deploying models. Companies in markets like Singapore or South Korea, focused on high-throughput applications, could see significant benefits. However, the economic trade-off will be critical. If Ultrafast mode carries a substantial premium, it will limit adoption for many regional startups with tighter budgets. The test for Asian developers is optimizing for value, not just speed.

    The thing to watch is the pricing structure announced at DevDay. If Ultrafast is priced aggressively, it will accelerate adoption across Asia. If it is a high-cost tier, it will remain niche. OpenAI’s capacity deployment through 2028 with Cerebras also bears watching. This will determine how quickly Ultrafast scales beyond initial limited access.

    Share this article

    Go deeper
    Original reporting by Testing CatalogWe don't republish, read the full story →

    Related reading

    6 stories