Gwern on a $12K a year salary: the outsider who saw scaling coming
Dwarkesh Patel's interview with the anonymous writer behind gwern.net covers how the 2020 scaling hypothesis essay came about, what it costs to live on about $12,000 a year, and why he thinks independent writing still matters. Notes on what a researcher without a lab can actually contribute.
An interview with a voice actor
The interview Dwarkesh Patel published on November 13 has an unusual production note. It was recorded in person, and then Gwern's words were read by Chris Painter so that neither his voice nor his face appears. The person being interviewed has written under a pseudonym for more than a decade and has decided that the anonymity is worth protecting even at the cost of a strange listening experience. His own explanation is that when people cannot see you they project onto you less, and, in his words, you get a hearing at all.
We found that framing useful for thinking about the rest of the conversation. Almost everything Gwern has contributed to AI discourse, including the piece most people know him for, was assessed on the page rather than on credentials. He has no lab, no grant, no affiliation to point at. That is either a limitation or an experiment in what a reader can produce alone, and the interview is mostly evidence for the second reading.
How the scaling essay was actually written
The 2020 essay on the scaling hypothesis is often described as a prediction, and Gwern is careful in the interview to describe it as something slower. He says he first encountered compute-centric arguments about AI from Moravec and Kurzweil in the mid 2000s and thought they were implausible. What changed his mind was a sequence of results rather than an argument, AlexNet and DanNet, then GPT-1 and GPT-2. The moment he names is looking at the few-shot learning charts in the GPT-3 paper and thinking, as he puts it, holy shit, we are living in the scaling world.
The essay itself was a synthesis. He describes reading the deep learning literature and noticing that datasets and models kept getting larger in every subfield, and that most commentators at the time had not read the papers that made the trend legible. He singles out the 2017 Baidu scaling laws work as one that had been widely overlooked. None of this required running an experiment. It required reading a large number of papers carefully, remembering them, and being willing to state a conclusion that most of the field was not stating.
The economics of doing this
The number in the title comes from the interview. Gwern says he lives on roughly $12,000 a year, made up of $900 to $1,000 a month from Patreon plus savings from Bitcoin he held early. He lives in a rural area, cooks his own food, uses a free gym and has no health insurance. His own description is that it is like being a grad student but with better ramen. Asked what it would take to move to San Francisco, he estimates $50,000 to $100,000 a year, which is a plain statement of how far the current setup is from what the field considers a normal research salary.
He also admits he almost never promotes the Patreon because he is reluctant to shill it harder. By the end of the interview Dwarkesh has persuaded him to set up a Stripe donation page. We mention this because it is the least glamorous part of the story and the part most worth noticing. The most cited independent writer on scaling has been operating on a budget that would not cover one month of a frontier training run's electricity, and that budget was set by his reluctance to ask rather than by anyone's assessment of his value.
What he thinks an outsider can still do
The obvious objection to independent research in 2024 is that the interesting experiments cost millions of dollars. Gwern's answer is that the essay form was never the only option and that its reproducibility is limited anyway. He says he would be perfectly happy if someone simply wrote more Reddit comments, never took a dollar for it, and wrote better Reddit comments. He also mentions Twitter threads and hosting PDFs so that links do not rot as contributions he takes seriously.
He also argues that this is an unusually good time to write, because what gets written now shapes what future models are trained on. His phrase is that there has never been a more vital, hinge-y time to write. Whether or not you accept that framing, it is a concrete claim about influence that a person with no compute has and that a lab does not have more of. A well argued document that many people read and cite influences the corpus more than a private experiment does.
The part of the interview that stuck with us is about motivation rather than method. Asked what drives him, he says he maximises rabbit holes, that he loves falling into a new one more than anything else. The scaling essay was a rabbit hole that happened to be about the most important trend of the decade. Most of his other rabbit holes were not, and he seems fine with that ratio.
What we take from it
We are a small foundation and we cannot train frontier models either. The interview is a reminder that the two things Gwern did in 2020, reading everything and being willing to state a conclusion out loud, are both cheap and both rare. The bottleneck on that kind of contribution is time and nerve, and the field rewards it slowly and unreliably, which is why the salary figure matters. If the next person doing this work needs $12,000 a year and a Patreon they are embarrassed to mention, that is a funding problem the rest of us could fix for less than the price of one GPU.
The thing we would want to see is someone attempt the same exercise on the literature of 2024, reading across subfields for the trend that most practitioners have not yet noticed, and publishing the synthesis under their own name or not. The 2020 essay was checkable in hindsight. A 2024 version would be checkable by 2028, and we would rather have three wrong ones to argue with than none.
Sources
From the foundation