GMAsia
    🇰🇷South Korea·AI News·18 Sept 2026·via 서울경제

    OpenAI Knowingly Used 10 Million News Articles to Train ChatGPT, Court Papers Show

    OpenAI knowingly trained its AI models on over 10 million news articles without permission. Court documents show about a third of these articles came from The New York Times. OpenAI co-founder Greg Brockman acknowledged bypassing paywalls for content collection. This occurred despite internal concerns that the practice constituted theft and damaged the news industry.

    Nexa's Summary

    OpenAI's internal communications reveal a clear awareness of copyright infringement risks. Microsoft's Brent Hecht called the content strategy "the largest exploitation of labor in human history." This undercuts OpenAI's public stance of fair use and its subsequent licensing agreements with publishers. The company prioritized model training over ethical sourcing, even as it recognized the "vicious cycle" harming news organizations.

    For Asia, this case underscores the precarious position of local news publishers. Many Asian outlets lack the legal resources of The New York Times to pursue copyright claims. This could accelerate the decline of traditional media revenue in markets like South Korea and Japan, where AI adoption is high. Regulators in these countries may face pressure to clarify AI training data legality, potentially impacting local AI model development.

    The U.S. Justice Department's support for OpenAI's fair use argument is a critical factor. If this position holds in court, it will set a precedent for AI developers globally. The thing to watch is how Asian courts and policymakers respond to similar copyright challenges in 2027 and beyond. A ruling against OpenAI would force a significant shift in data acquisition strategies for all AI companies.

    #international
    Original reporting by 서울경제We don't republish, read the full story →

    Related reading

    6 stories