Samsung's ChatGPT leak and the birth of the enterprise AI policy
Engineers pasted proprietary code into ChatGPT, Samsung banned the tools, and vendors are now selling the fix. A deployment failure case study on why the first enterprise AI policy at most companies was written by a leak.
What happened
Samsung's device solutions division, the part of the company that makes semiconductors, allowed staff to use generative AI tools from March 11. Within three weeks there were three incidents. One engineer pasted source code from a faulty semiconductor database into ChatGPT to ask for help debugging it. A second pasted code for defective equipment and asked for a fix. A third fed a recorded meeting into the tool and asked for minutes. Each of those is a completely ordinary use of the product, which is the point.
The company's first response, reported on April 6, was to cap each employee's prompt at 1,024 bytes. That is a length limit standing in for a data classification policy, and it tells you how little machinery existed for the problem. On May 1 the restriction became a ban on ChatGPT, Bing and Bard on company owned computers, tablets and phones, and on any device connected to the internal network. Samsung called it temporary, lasting until it builds security measures for a safe environment. An internal survey in April found about 65 percent of participants agreed the tools carry a security risk.
Why the leak was a leak
The specific worry Samsung named is that data sent to an external server is hard to retrieve and delete, and could be disclosed to other users. At the time of the incidents that was an accurate reading of OpenAI's consumer terms. The Gizmodo report quotes OpenAI's own statement that it may use data submitted to ChatGPT and other consumer services to improve its models, and that it is not able to delete specific prompts from a user's history. A chip company's database schema in a training corpus is a real exposure even if nobody ever sees it surface.
It is worth separating the two failure modes because the fixes differ. The first is retention, where the vendor keeps the text and could be compelled or breached. The second is training, where the text might shape a future model's outputs for other customers. Samsung's policy language covered both. Most of the products that followed address them separately, and a company evaluating a vendor should ask about each on its own.
The vendor response arrived before the ban
On April 25, six days before Samsung's ban took effect, OpenAI shipped a toggle that lets a user turn off chat history. Conversations started with history off are not used for training and are kept for thirty days for abuse review, then deleted. In the same announcement OpenAI previewed a ChatGPT Business subscription that follows the API's data usage policy, under which customer data is not used for training by default, and said it would arrive in the coming months.
So the product that would have satisfied Samsung's stated concern was announced within a month of the leak being reported. That timing is the case study. The consumer product was built for individuals and shipped with individual defaults, and the enterprise controls were built after companies discovered the defaults by accident. A company that waited for the controls before allowing use would have been fine. A company that allowed use because the tool was useful found out what the defaults were.
Not only Samsung
TechCrunch lists the other companies that had already restricted the tools by the time Samsung acted. Bank of America, Citigroup, Deutsche Bank, Goldman Sachs, Wells Fargo and JPMorgan on the finance side, where the retention question overlaps with regulatory record keeping. LG and SK Hynix in Korea, both of which handle the same kind of process data Samsung does. The bans clustered in industries where the legal cost of a disclosed document is already known, which suggests the companies without bans mostly had not yet worked out what theirs would be.
What we take from the list is that the first enterprise AI policy at most large firms was written in reaction to a headline rather than in advance of a rollout. That is normal for a new tool. It is also expensive, because a blanket ban removes the productivity that motivated the use in the first place, and the survey figure suggests the employees themselves saw the risk clearly enough to have followed a policy if one had existed.
What a policy looks like now
The version we would write for a research group is short. Classify what may leave the building, and treat a chat window like email to an outside party. Require a vendor tier whose terms say the data is not used for training and state a retention period, and check the terms rather than the marketing. Where the data cannot leave at all, run a model locally, which is now practical for the kind of debugging help the Samsung engineers wanted. And log usage, because the alternative to a log is a survey after the fact.
The thing to watch this year is whether the enterprise tiers actually differ in retention and training terms or only in price. The consumer history toggle already gives an individual most of what Samsung asked for. If the business product is the same promise with an admin console, then the ban era will end quickly. If it comes with contractual retention commitments and audit rights, then the leak will have done what a year of policy papers did not, which is to make data handling a line item in the purchase.
Sources
From the foundation