GMAsia
    🇨🇳China·AI News·9 Oct 2026·via Tech Node

    ByteDance researchers identify cause of inconsistent long-context retrieval in DeepSeek models

    Researchers from ByteDance’s Seed team identified a mechanism causing inconsistent long-context retrieval in DeepSeek models. Their study found that chunked KV-cache compression makes retrieval accuracy sensitive to where information falls within a compression window, leading to differences of up to 40 percentage points across positions.

    Nexa's Summary

    The ByteDance Seed team's findings reveal a specific limitation introduced by chunked KV-cache compression, a technique designed to reduce memory and attention costs in large language models. While efficient, this compression method makes a model's ability to recall information highly dependent on the information's exact placement within a compressed segment. This variability means the same data can be easily retrieved in one position but difficult in another, a pattern the researchers call "phase sensitivity."

    The study reproduced this behavior in models trained from scratch, suggesting it is an inherent characteristic of the compression technique rather than a specific issue with DeepSeek models alone. The researchers found that different attention components within these models specialize in retrieving information from distinct positions, contributing to the observed inconsistencies. This specialization, when combined with chunked compression, explains how strong average benchmark results can obscure recurring retrieval weaknesses.

    This work offers a detailed explanation for a specific challenge in long-context understanding for models employing KV-cache compression. It implies that assessing model performance solely on average retrieval accuracy may overlook significant practical limitations. For applications requiring consistent and precise information recall from long contexts, addressing this "phase sensitivity" could be crucial for more reliable model behavior.

    Share this article

    Go deeper
    Original reporting by Tech NodeWe don't republish, read the full story →

    Related reading

    6 stories