The influence-operations report that predicted the year before it happened
Reading notes on the January 2023 report from OpenAI, Georgetown CSET and the Stanford Internet Observatory on language models and propaganda. What its four intervention stages actually propose, which unknowns it flagged, and where we think the framing will hold.
A workshop from 2021 published into a different world
The report came out on January 10 and it was written mostly before ChatGPT existed. Six authors from Georgetown's Center for Security and Emerging Technology, OpenAI and the Stanford Internet Observatory built it on a two-day workshop they convened in October 2021 with 30 disinformation and machine learning researchers, a whitepaper circulated to those participants, and what the authors describe as subsequent months of research. The authors are careful to say the report does not represent the consensus of that room.
That timing matters for how we read it. The paper's claims about what a language model could do for a propagandist were speculative when they were drafted. By the time the PDF went up, anyone with a browser had spent six weeks watching a model write fluent paragraphs on demand. So the interesting question for us is which of the report's structures survive contact with that reality, and which of its worries were about a capability that has now simply arrived.
Actors, behaviours, content
The threat side of the report borrows Camille Francois's ABC framework from the disinformation literature and asks how a text generator changes each letter. On actors, the claim is that lower cost brings a larger and more diverse group of propagandists into the game, and that propagandists for hire who automate text production gain a competitive edge. On behaviour, the report expects automation to increase the scale of campaigns, to make expensive tactics like cross-platform testing cheaper, and to enable new ones, with dynamic and personalised one-on-one chatbot content given as the example. On content, it expects messaging to become more credible to targets whose language and culture the propagandist does not share, and campaigns to become harder to spot because the copy-and-paste text that gets them caught today would be replaced by linguistically distinct outputs.
That last point is the one we find most convincing and most under-discussed. The report notes that existing operations are frequently discovered because they reuse text, and that identifying inauthentic accounts often relies on subtle cues such as a misused idiom, a repeated grammatical error, or a backtick where a native speaker would type an apostrophe. Every one of those cues is a thing a language model removes for free. The authors summarise their position as a high-confidence judgment that language models will be useful for influence operations, while saying the exact nature of that use is unclear.
The kill chain and its four stages
The mitigation side is organised as a kill chain. To run a campaign with a language model, a propagandist needs a model to exist, needs access to it, needs infrastructure to distribute what it generates, and needs the output to actually move a target audience. The report names these four stages model design and construction, model access, content dissemination, and belief formation, and sorts illustrative mitigations under each.
Under construction it lists building models that are more fact-sensitive, spreading so-called radioactive data so that generated text becomes detectable, government restrictions on data collection, and access controls on AI hardware. Under access it lists stricter usage restrictions from providers and new norms around model release. Under dissemination it lists platform and provider coordination to identify AI content, proof-of-personhood requirements to post, steps by entities that rely on public input to reduce exposure to misleading submissions, and widely adopted digital provenance standards. Under belief formation it lists media literacy campaigns and consumer-focused AI tools. The authors state plainly that they are not endorsing any of these, only showing where in the pipeline each one bites.
Reading the table, we notice how unevenly the stages are populated with things anyone can do today. The construction and access mitigations sit with a handful of labs and governments. The dissemination mitigations require platforms and labs to cooperate, and the report says outright that a platform will struggle to know whether a campaign used a language model unless it can work with the developer to attribute text. The belief-formation mitigations are the only ones that do not depend on the technology at all, and they are also the ones with the weakest evidence of working.
The unknowns the authors chose to name
The section we would hand to a colleague is the list of critical unknowns. The authors ask which capabilities relevant to influence will emerge as a side effect of ordinary research, giving long-form persuasive argument as an example of something that could arrive without anyone aiming for it. They ask whether it will be more effective for a propagandist to fine-tune a smaller model for persuasion than to use a generic large one, and admit they do not know how much cheaper that would be. They ask whether many actors will invest in building or stealing models, whether governments or industries will set norms against propaganda use, and when easy-to-use text tools will become publicly available.
That last unknown has already been answered, and the answer arrived between the workshop and the publication date. The report describes language models as still requiring operational know-how and infrastructure to use skilfully, and worries about the moment when tweet-length and paragraph-length generation becomes a consumer product. That moment is now behind us. What we take from this is that the report's other unknowns should be read with the assumption that the cheapest one resolved first, and that the harder ones about fine-tuning and investment by state actors are the ones worth tracking through the year.
What we would want measured
The report is honest that impact from influence operations is hard to measure at all, and it cites the mismatch between the engagement metrics researchers can see and the opinion change they care about. It also cites one 2020 study that found no substantial effect on attitudes from Twitter users interacting with Internet Research Agency accounts in late 2017. If language models change anything, it will show up first in the cost and detectability of operations that platforms take down, and only later, if ever, in measurable belief change.
So the experiment we would like to see someone run is a boring one. Take the takedown reports that Meta and Twitter have published since 2016, which the report says cover well over a hundred operations from dozens of countries, and track over 2023 whether the text-reuse and grammatical-error signals that currently catch these campaigns start to disappear from new takedowns. If the report's content predictions are right, that signal decays before any of the four mitigation stages has been built out. If it does not decay, the propagandists are slower to adopt than the authors feared, and that would be worth knowing too.
Sources
From the foundation