DeepSeek details DSec sandbox infrastructure for agent training
DeepSeek has detailed DSec, a new sandbox platform designed for large-scale agent training. The platform unifies function-call, container, microVM, and full-VM sandboxes, coordinating their lifecycle with reinforcement learning workloads. DeepSeek founder Liang Wenfeng is listed among the paper's 130-plus authors. A production-scale unit of DSec spans around 160 nodes, supporting approximately 3 million sandboxes per day and more than 380,000 concurrent sandboxes. It can also create over 5,000 sandboxes per second.
DeepSeek's DSec sandbox infrastructure points to the growing complexity of agent training. Unifying four sandbox types and handling 3 million daily sandboxes shows the scale needed for advanced AI development. This effort involves orchestrating diverse environments for reinforcement learning workloads, a critical bottleneck for many AI labs.
For Chinese AI companies, DSec offers a blueprint for internal infrastructure development. Building such a robust, scalable platform in-house reduces reliance on external cloud providers for sensitive agent training. This could give DeepSeek an edge in iterating faster on AI agents, a key area for competitive differentiation against global players.
The thing to watch is whether this internal infrastructure can be productized or licensed. If DeepSeek offers DSec as a service, it could become a significant player in the AI agent development toolchain across Asia. The test will be whether its 5,000 sandboxes per second creation rate can meet diverse enterprise needs.
Related reading
6 stories
Alibaba unveils ‘pragmatic’ AI road map to drive monetisation, infrastructure efficiency

China’s humanoid robot IPO slowdown no threat to firms with ‘genuine strength’: Deloitte

Mind Lab builds Mint Recursive to help companies train their own AI

Investors pivot to selective China bets in technology as property growth fades: DBS Bank

Tencent rolls out payment app for foreign travellers ahead of Apec summit

