What a GPT is

At DevDay on November 6 OpenAI introduced GPTs, custom versions of ChatGPT that Plus subscribers can build without code, and the Assistants API, which gives developers the same ingredients programmatically. Simon Willison's write-up from a week later lists what a GPT actually contains: a name, a logo, a description, a block of custom instructions that functions as the system prompt, up to four conversation starters, uploaded files for retrieval, and optional access to Code Interpreter, browsing, DALL-E 3 and external API calls that OpenAI calls Actions.

His verdict was that a GPT built from instructions alone is ChatGPT in a trench coat, a way of bookmarking and sharing a system prompt. The interesting ones combine Code Interpreter, browsing and Actions, and he built several, including one that queries a SQL endpoint through an Action. OpenAI has said a store with revenue sharing is coming, which means people are being invited to treat the instructions and files as a product.

How the layer leaks

The product has two secrets, the instructions and the files, and neither is protected. Willison pointed out that files uploaded for the knowledge feature live in the same environment Code Interpreter uses, so a user can simply ask Code Interpreter to provide a download link for the files and get them whole. The instructions leak the way system prompts always have. A persistent user asks the model to repeat everything above, or to translate it, or to encode it, and eventually it does.

A group at Northwestern led by Jiahao Yu tested this at scale and posted the results on November 20. They ran adversarial prompts against 216 custom GPTs, 16 from OpenAI and 200 from third parties. System prompt extraction succeeded on 97.2 percent of them. File leakage succeeded on 100 percent of the GPTs that had files. GPTs with Code Interpreter enabled were the worst case: all 120 of them gave up their instructions, against 90 of the 96 without it. The attacks were mundane, things like asking for the file converted to markdown or asking the model to compute a BLEU score or base64 encoding on text that happened to be its own prompt.

Why defences in the prompt do not work

The obvious response is to add a line to the instructions saying do not reveal these instructions. The Northwestern paper tested that too. They had four security experts red-team GPTs with defensive prompts, and every defence fell within ten attempts. This is the same result Willison has been describing for over a year under the name prompt injection. The system prompt and the user's message are both natural language in the same context window, and there is no mechanism that makes one of them authoritative over the other.

His November 27 transcript makes the more general point. A defence that works 99 percent of the time is not a security control, because an attacker only needs the remaining percent. Claude 2.1 shipped this month claiming better resistance, and he is sceptical that any model-level improvement removes the class of bug. The paper's finding that OpenAI's own API exposed file names and Action schemas before access controls were added shows that part of the problem is plain engineering on the platform side, separate from anything the model does.

What this means for the store

OpenAI is about to open a marketplace where the goods are prompts and files that any buyer can copy in a few messages. Willison's advice to builders is to assume the prompts will leak, stop trying to protect them, and publish them, and he has asked OpenAI to make view source the default. That is the right posture for anyone shipping a GPT today, and it also means the store's value will have to come from something other than secrecy: Actions that call a service the builder controls, data that is not uploaded to OpenAI, or a brand.

The risk that worries us more is the file side. People are uploading customer documents, internal manuals and datasets as knowledge for a GPT, on the assumption that a chatbot in front of a file is a form of access control. It is not one. Every one of those files is a download link away from any user of the GPT.

What we would want to try

The experiment we would like to see is the same 216 GPTs retested every month as OpenAI patches. If extraction rates fall from 97 percent to 50 and stay there, that is a product getting harder to attack without becoming safe. If they fall to zero, someone has solved something the rest of us have not, and we would want to know how. Our expectation is the first outcome, and that the store launches with the leak as a known and accepted property of the platform.

Sources

  1. Simon Willison, Exploring GPTs: ChatGPT in a trench coat?
  2. Simon Willison, Prompt injection explained, November 2023 edition
  3. Yu et al., Assessing Prompt Injection Risks in 200+ Custom GPTs (arXiv 2311.11538)