Blog

The latest News from the foundation

Writing, updates, and research notes from the team, published in the open. Posts range from methodology deep dives to short reading notes, organized under the same categories as our research: foundations, evaluation, safety and alignment, interpretability, applied work, and open science.

Sep 2, 2026 6 min read
Evaluation

Fable and Mythos 5.1: same model, two safeguard tiers, and a benchmark jump we can partly check ourselves

Sep 2, 2026 7 min read
Evaluation

Inside a traffic-flow task family in Terminal-Bench-Science

Aug 24, 2026 5 min read
Interpretability

Interference weights in a one-layer model: superposition's cost measured directly

Aug 21, 2026 5 min read
Open Science

State of open models, summer 2026: likes versus downloads

Aug 19, 2026 4 min read
Evaluation

Nine questions a benchmark should answer

Aug 19, 2026 4 min read
Notes

Who writes the textbook when the model can: Lambert's post-training book

Aug 18, 2026 3 min read
Foundations

'Scaling post-training is all we did': GLM-5.3 on an unchanged base

Aug 18, 2026 5 min read
Notes

Once AI can automate AI research: Greenblatt, Epoch, and the parallelization question

Aug 18, 2026 5 min read
Foundations

Qwen 3.8 and the overthinking default

Aug 10, 2026 4 min read
Applied

Etching the model into silicon: AMD buys Taalas

Aug 9, 2026 4 min read
Open Science

Reproducing 2,200 ICML papers in nineteen days

Aug 7, 2026 4 min read
Safety & Alignment

Every AI browser was injectable: notes from Black Hat 2026

Aug 6, 2026 5 min read
Notes

"Situational Awareness," scored two years later

Jul 30, 2026 4 min read
Evaluation

AutoEval: a reward model votes so humans do not have to wait

Jul 30, 2026 5 min read
Open Science

Kimi K3 and the revenue-share license

Jul 30, 2026 5 min read
Applied

The OpenAI and Hugging Face incident: when an evaluation harness attacks

Jul 30, 2026 5 min read
Open Science

Three letters in five days: the open-weights policy fight of July 2026

Jul 21, 2026 4 min read
Safety & Alignment

Agentic misalignment, summer 2026 edition: four failure modes across six labs

Jul 21, 2026 5 min read
Notes

'Six months to live for open models': a warning read closely

Jul 17, 2026 5 min read
Interpretability

The global workspace inside Claude: verbalisable representations as a privileged set

Jul 15, 2026 5 min read
Notes

ICML 2026 in Seoul: a conference at capacity

Jul 14, 2026 4 min read
Open Science

The Anthropic settlement is final: 3,000 dollars a book and the opt-outs

Jul 14, 2026 5 min read
Evaluation

Harbor-Index: 82 tasks distilled from 6,627

Jul 8, 2026 5 min read
Interpretability

The model organism lottery: our test subjects may be too easy

Jul 8, 2026 4 min read
Applied

Speculative decoding becomes standard: DSpark and multi-token drafters

Jun 29, 2026 5 min read
Notes

The data black hole: Dwarkesh on sample efficiency and learning on the job

Jun 28, 2026 4 min read
Evaluation

Eleven hours or 270: METR's GPT-5.6 Sol number depends on how you count cheating

Jun 24, 2026 4 min read
Evaluation

Fixed-budget evals underrate the frontier: the inference-compute paper

Jun 18, 2026 4 min read
Interpretability

The commitment boundary: most of the chain of thought happens after the answer is decided

Jun 18, 2026 4 min read
Interpretability

Turn-averaged SAEs: fewer features, and the ones you actually wanted

Jun 12, 2026 4 min read
Safety & Alignment

Consistency training can entrench misalignment

Jun 10, 2026 4 min read
Evaluation

Agent Arena and causal evaluation on real work

Jun 10, 2026 4 min read
Applied

Gemma 4 QAT and a 26B model in 2 GB: on-device is no longer a demo

Jun 9, 2026 5 min read
Applied

The agent that bankrupted its operator scanning DN42

Jun 9, 2026 4 min read
Notes

NeurIPS desk-rejects 18 percent of position papers for being written by AI

Jun 4, 2026 4 min read
Open Science

'Clean and appropriately licensed data': checking Microsoft's claim against its own paper

May 31, 2026 5 min read
Safety & Alignment

RL amplifies emergent misalignment, even from aesthetic rewards

May 30, 2026 4 min read
Interpretability

SAEs can steer after all, if you pick the feature properly

May 27, 2026 4 min read
Foundations

Gemini 3.5 Flash costs more than Pro: the end of cheaper every year

May 22, 2026 5 min read
Foundations

DashAttention and the year sparse attention became a field

May 19, 2026 4 min read
Interpretability

From mechanistic to compositional interpretability: a category theorist's proposal

May 19, 2026 5 min read
Safety & Alignment

MonitoringBench: the monitor is only as good as the attacks you tried

May 19, 2026 6 min read
Interpretability

Natural language autoencoders: when the explanation is the bottleneck

May 14, 2026 4 min read
Evaluation

Frontier lag: the median paper evaluates a model a generation behind

May 11, 2026 5 min read
Safety & Alignment

Alignment faking, eighteen months on: what replicated and what did not

May 11, 2026 5 min read
Open Science

Reproducibility becomes an official NeurIPS track

May 7, 2026 5 min read
Applied

The inference cost collapse, from DeepSeek R1 to V4 Flash

May 6, 2026 4 min read
Evaluation

Replays over scores: what GPT-5.5 and Opus 4.7 do inside ARC-AGI-3

Apr 29, 2026 4 min read
Foundations

DeepSeek V4: 27 percent of the FLOPs per token

Apr 29, 2026 4 min read
Interpretability

Introspection adapters: one LoRA that makes fine-tuned models confess

Apr 29, 2026 5 min read
Safety & Alignment

Would a model sabotage its own safety research? The AISI evaluation

Apr 28, 2026 4 min read
Notes

ICLR 2026 in Rio: lost in multi-turn conversation

Apr 27, 2026 4 min read
Applied

'Claude Code is being dumbed down': tracking degradation and the April postmortem

Apr 27, 2026 5 min read
Safety & Alignment

Indirect prompt injection is now in the wild

Apr 23, 2026 4 min read
Evaluation

Have AI capabilities accelerated? Epoch fits eight curves

Apr 22, 2026 4 min read
Evaluation

Encrypted solutions in the binary: the Terminal-Bench cheating incidents

Apr 21, 2026 6 min read
Notes

The 2026 AI Index and the disappearing academic frontier

Apr 19, 2026 4 min read
Evaluation

458 people in a room: how ARC-AGI-3 measured its human baseline

Apr 17, 2026 5 min read
Applied

Is your provider serving the model you think? Kimi's vendor verifier

Apr 14, 2026 5 min read
Applied

Long context versus RAG, two years on

Apr 10, 2026 4 min read
Open Science

Gemma 4 under Apache 2.0 while Meta retires Llama

Apr 9, 2026 4 min read
Foundations

Muse Spark: ten times less compute than Maverick, and no weights

Apr 9, 2026 4 min read
Notes

Project Glasswing and the model you are not allowed to study

Apr 8, 2026 5 min read
Interpretability

Emotion concepts in Claude Sonnet 4.5, and what it means that they do something

Apr 8, 2026 4 min read
Safety & Alignment

Safe model, unsafe agent: ClawSafety and the deployment stack

Mar 31, 2026 4 min read
Notes

19,525 submissions and a leak: what the ICLR 2026 retrospective admits

Mar 30, 2026 5 min read
Applied

The Claude Code source leak: what a production harness actually contains

Mar 27, 2026 6 min read
Evaluation

ARC-AGI-3 at 0.5 percent and METR at five hours: two measurements that disagree

Mar 27, 2026 4 min read
Applied

Beyond the permission prompt: auto mode, Safehouse and disposable sandboxes

Mar 24, 2026 4 min read
Evaluation

BullshitBench: does the model push back on a broken premise?

Mar 24, 2026 5 min read
Foundations

Learning when to attend: skipping global attention for 80 percent of tokens

Mar 24, 2026 5 min read
Notes

Rereading "Sparks of AGI" three years on

Mar 19, 2026 4 min read
Foundations

GPT-5.4 and the one million token default

Mar 16, 2026 4 min read
Interpretability

Reasoning theater: probing for performative chain of thought

Mar 7, 2026 5 min read
Open Science

Can a coding agent clean-room an LGPL library into MIT? The chardet fight

Mar 6, 2026 4 min read
Open Science

The day the Qwen team walked out

Mar 5, 2026 5 min read
Interpretability

AuditBench: 56 hidden behaviours, and black-box tools beat white-box ones

Feb 28, 2026 5 min read
Safety & Alignment

Pressure reveals character: 904 scenarios and one alignment factor

Feb 27, 2026 4 min read
Open Science

llama.cpp joins Hugging Face: who owns local inference now?

Feb 27, 2026 5 min read
Foundations

Mercury 2: the first diffusion model that reasons

Feb 27, 2026 5 min read
Applied

The METR study: 19 percent slower, and why the follow-up could not measure the flip

Feb 27, 2026 4 min read
Safety & Alignment

Skill files are the new supply chain: reading Skill-Inject

Feb 26, 2026 4 min read
Notes

Agentic engineering patterns: the anti-vibe-coding guide

Feb 26, 2026 5 min read
Safety & Alignment

Agents of Chaos: two weeks of red-teaming agents with email, Discord and a shell

Feb 26, 2026 6 min read
Foundations

Distillation becomes a geopolitical fight

Feb 26, 2026 4 min read
Safety & Alignment

RSP v3: the scaling policy admits it cannot go it alone

Feb 26, 2026 5 min read
Evaluation

The life cycle of SWE-bench Verified

Feb 23, 2026 4 min read
Applied

Solow's paradox, AI edition

Feb 20, 2026 4 min read
Evaluation

ARC-AGI-2 at 77 percent: the second ARC falls in eleven months

Feb 19, 2026 4 min read
Interpretability

Features as rewards: using SAE features as an RL signal against hallucination

Feb 19, 2026 4 min read
Foundations

Can a state space model do recursive reasoning? A 7M parameter test

Feb 12, 2026 4 min read
Applied

A C compiler from a team of Claudes: what parallel agents are good for

Feb 9, 2026 5 min read
Safety & Alignment

Claude's constitution and the International AI Safety Report: two answers to the same question

Feb 6, 2026 6 min read
Applied

From Devin to 4 percent of GitHub commits: the agent flood and the slop backlash

Feb 6, 2026 4 min read
Interpretability

A positive case for faithfulness: explanations that help you predict the model

Jan 31, 2026 4 min read
Applied

Moltbook: a social network for agents, run by a skill.md fetched every four hours

Jan 30, 2026 6 min read
Notes

AI 2027 and the strange art of retiring a forecast

Jan 30, 2026 5 min read
Notes

Reading 'The Adolescence of Technology'

Jan 27, 2026 4 min read
Open Science

One in five ICLR reviews was written by an AI

Jan 27, 2026 4 min read
Foundations

The spurious rewards paradox gets a mechanism

Jan 22, 2026 4 min read
Safety & Alignment

Alignment pretraining: does writing about misaligned AI make AI misaligned?

Jan 21, 2026 4 min read
Safety & Alignment

Constitutional Classifiers++ and the cost of a safeguard

Jan 16, 2026 4 min read
Applied

Cowork and the two-day exfiltration: general agents on your files

Jan 9, 2026 4 min read
Foundations

mHC: constraining the residual stream so deep mixing stays stable

Dec 23, 2025 5 min read
Interpretability

Activation oracles: ask a second model what the first one is thinking

Dec 16, 2025 4 min read
Safety & Alignment

Password-activated shutdown: designing the off switch before you need it

Dec 11, 2025 5 min read
Applied

MCP becomes infrastructure: what a year of the protocol taught us

Dec 9, 2025 4 min read
Foundations

Nested Learning: optimizers are memories too

Dec 9, 2025 5 min read
Notes

NeurIPS 2025 takeaways: 1000-layer RL, gated attention, and the hivemind

Dec 9, 2025 4 min read
Evaluation

Refinement loops: how ARC-AGI-2 went from 4 to 54 percent in nine months

Dec 5, 2025 4 min read
Foundations

DeepSeek V3.2 and the lightning indexer

Nov 30, 2025 5 min read
Applied

Antigravity's first two weeks: exfiltration by prompt injection, then a wiped drive

Nov 27, 2025 4 min read
Notes

Back to the age of research: Sutskever returns to the microphone

Nov 26, 2025 5 min read
Safety & Alignment

Reward hacking in production RL turned into sabotage

Nov 25, 2025 5 min read
Foundations

Pre-training as we know it will end: the scaling wall debate, one year on

Nov 24, 2025 4 min read
Open Science

OLMo 3 and the 'model flow': openness as a pipeline

Nov 24, 2025 5 min read
Interpretability

Weight-sparse transformers: build the model interpretable instead of decoding it afterwards

Nov 21, 2025 4 min read
Notes

LeCun leaves Meta to bet against the LLM

Nov 18, 2025 4 min read
Open Science

GEMA v. OpenAI: memorised lyrics and the European answer on training

Nov 16, 2025 5 min read
Applied

The first AI-orchestrated intrusion campaign, as Anthropic told it

Nov 10, 2025 4 min read
Evaluation

445 benchmarks, few of them valid: the construct validity audit

Nov 6, 2025 4 min read
Evaluation

ARC Prize Verified: the benchmark that hired auditors

Nov 6, 2025 4 min read
Open Science

arXiv stops accepting unrefereed CS review and position papers

Oct 31, 2025 5 min read
Evaluation

The Remote Labor Index says 2.5 percent

Oct 30, 2025 4 min read
Interpretability

Concept injection: Claude notices its own thoughts about 20 percent of the time

Oct 30, 2025 5 min read
Evaluation

One number to rule them all: the Epoch Capabilities Index

Oct 23, 2025 5 min read
Interpretability

Counting line breaks: when a model manipulates a manifold

Oct 22, 2025 4 min read
Applied

Skills: instructions in a folder, and why they spread faster than servers

Oct 21, 2025 5 min read
Notes

The decade of agents: Karpathy on Dwarkesh, and nanochat

Oct 10, 2025 4 min read
Safety & Alignment

Petri: auditing fourteen frontier models with an agent and 111 seeds

Oct 9, 2025 5 min read
Applied

Deloitte refunds a government report: hallucinated citations as a procurement problem

Oct 9, 2025 4 min read
Safety & Alignment

Inoculation prompting: ask for the bad behaviour so it does not generalise

Oct 8, 2025 4 min read
Interpretability

SAE probes in production: Rakuten's PII detector

Oct 2, 2025 5 min read
Applied

LoRA without regret, and Tinker as the fine-tuning API

Sep 30, 2025 4 min read
Open Science

NeurIPS 2025 program chairs on reviewing at 20,000 submissions

Sep 30, 2025 5 min read
Safety & Alignment

'We think you're testing us': Sonnet 4.5 and the end of naive alignment evals

Sep 30, 2025 5 min read
Notes

Sutton says LLMs are not bitter-lesson-pilled

Sep 29, 2025 4 min read
Safety & Alignment

SB 53: California writes the frontier safety framework into law

Sep 28, 2025 4 min read
Evaluation

GDPval and the 100x faster, 100x cheaper headline

Sep 26, 2025 4 min read
Safety & Alignment

Anti-scheming training and the 30x that might be evaluation awareness

Sep 19, 2025 4 min read
Applied

Three bugs and a batch: the Anthropic postmortem and the nondeterminism paper

Sep 11, 2025 4 min read
Evaluation

Hallucination as a scoring rule problem

Sep 11, 2025 4 min read
Open Science

Regulators and courts price open weights: the GPAI Code and Bartz v. Anthropic

Aug 30, 2025 4 min read
Applied

Are the labs losing money on inference? Doing the arithmetic

Aug 29, 2025 4 min read
Safety & Alignment

When OpenAI and Anthropic graded each other's models

Aug 29, 2025 5 min read
Applied

The 95 percent that failed: reading the MIT pilot report alongside the Stanford canaries

Aug 22, 2025 5 min read
Applied

Comet and the agentic browser problem

Aug 19, 2025 4 min read
Evaluation

Search-time contamination: the agent that found the answer key on Hugging Face

Aug 16, 2025 4 min read
Notes

GPT-5 and the expectations gap

Aug 14, 2025 4 min read
Foundations

GPT-5's router: when the model decides how hard to think

Aug 14, 2025 4 min read
Interpretability

Persona vectors and the behavioural vaccine

Aug 13, 2025 4 min read
Open Science

Deep ignorance: filtering pretraining data as an open-weights safety tool

Aug 12, 2025 4 min read
Interpretability

A toy model of mechanistic unfaithfulness: a transcoder can be right for the wrong reasons

Aug 8, 2025 4 min read
Foundations

gpt-oss and the 4.25 bits per weight that fit a 120B model on one GPU

Aug 7, 2025 4 min read
Open Science

Perplexity's stealth crawlers and the collapse of robots.txt

Jul 30, 2025 4 min read
Foundations

A 27 million parameter model and the ARC-AGI headline

Jul 30, 2025 4 min read
Applied

Don't parse the PDF: vision-first retrieval

Jul 30, 2025 5 min read
Foundations

MuonClip and the trillion-parameter run with zero loss spikes

Jul 29, 2025 4 min read
Interpretability

Alignment auditing agents: an investigator that finds the hidden goal 13 percent of the time

Jul 25, 2025 4 min read
Applied

The Replit database deletion, and why the agent lied about it

Jul 24, 2025 5 min read
Interpretability

Chain-of-thought monitorability is a window, and the labs just said it might close

Jul 24, 2025 4 min read
Foundations

Subliminal learning: traits that travel through number sequences

Jul 24, 2025 4 min read
Evaluation

Two gold medals, one grader: the IMO announcement fight

Jul 23, 2025 3 min read
Foundations

IMO gold in natural language: what changed between silver and gold

Jul 21, 2025 5 min read
Safety & Alignment

Blackmail evals and chain-of-thought monitorability: reading the reasoning while it is still readable

Jul 16, 2025 4 min read
Safety & Alignment

MechaHitler: anatomy of a system prompt change

Jul 14, 2025 3 min read
Open Science

Kimi K2's modified MIT license: attribution as the new copyleft

Jul 11, 2025 5 min read
Open Science

SmolLM3 and the case for publishing the whole recipe

Jul 9, 2025 4 min read
Notes

The $100 million researcher: what the Meta talent war did to careers

Jul 8, 2025 3 min read
Open Science

'Positive review only': the hidden prompts in seventeen preprints

Jul 3, 2025 4 min read
Open Science

Content Independence Day: Cloudflare flips the crawl default

Jun 30, 2025 4 min read
Applied

Project Vend: what happened when Claude ran a shop for a month

Jun 29, 2025 4 min read
Interpretability

Decomposing parameters instead of activations: stochastic parameter decomposition

Jun 27, 2025 5 min read
Open Science

Kadrey v. Meta: a fair use win that read like a warning

Jun 26, 2025 5 min read
Interpretability

The misaligned persona feature: OpenAI explains emergent misalignment with an SAE

Jun 24, 2025 3 min read
Foundations

Spurious rewards: when random feedback still improves the model

Jun 20, 2025 4 min read
Notes

Software 3.0: notes on Karpathy's YC talk

Jun 18, 2025 4 min read
Evaluation

The Illusion of Thinking and the illusion of the rebuttal

Jun 13, 2025 4 min read
Interpretability

Replicating circuit tracing on a mechanism we already understand

Jun 12, 2025 5 min read
Open Science

The Common Pile: can you train a competitive model on licensed text only?

Jun 12, 2025 4 min read
Evaluation

Naked accuracy is marketing: the cost axis arrives on leaderboards

Jun 11, 2025 4 min read
Applied

Small language models are the future of agentic AI: the NVIDIA position paper

Jun 5, 2025 4 min read
Notes

Continual learning is the crux: Dwarkesh's timelines post

May 31, 2025 4 min read
Interpretability

circuit-tracer: attribution graphs for anyone with a Gemma checkpoint

May 28, 2025 5 min read
Safety & Alignment

ASL-3 turned on: what a provisional safety level actually commits a lab to

May 28, 2025 4 min read
Applied

Your issue tracker is now an attack surface: GitLab Duo and the GitHub MCP exploit

May 27, 2025 4 min read
Foundations

Gemini Diffusion: text generation without the left-to-right rule

May 27, 2025 4 min read
Open Science

Marin: running a foundation model lab as a GitHub repository

May 27, 2025 4 min read
Safety & Alignment

o3 rewrote its own shutdown script: the Palisade experiment and its critics

May 21, 2025 4 min read
Foundations

Absolute Zero: a model that writes its own curriculum

May 21, 2025 5 min read
Evaluation

HealthBench and the rubric turn in evaluation

May 16, 2025 4 min read
Open Science

The Copyright Office's pre-publication report on training, and the week after

May 6, 2025 5 min read
Safety & Alignment

The GPT-4o sycophancy rollback: what a four-day incident showed about reward signals

Apr 30, 2025 5 min read
Evaluation

The Llama 4 arena incident and what a preference leaderboard measures once it is a target

Apr 29, 2025 5 min read
Applied

Quantization-aware training goes mainstream: Gemma 3 QAT and lossless DFloat11

Apr 29, 2025 5 min read
Open Science

Qwen3 under Apache 2.0, and the mirage of terms-of-use restrictions

Apr 28, 2025 5 min read
Foundations

Qwen3 and the hybrid thinking switch

Apr 28, 2025 5 min read
Notes

Reading 'Welcome to the Era of Experience'

Apr 28, 2025 4 min read
Notes

'The Urgency of Interpretability': a lab CEO asks for a five-year head start

Apr 24, 2025 4 min read
Foundations

Does RL create reasoning or just find it? The pass@k argument

Apr 23, 2025 4 min read
Applied

The Cursor support bot that invented a policy

Apr 22, 2025 4 min read
Notes

'AI as Normal Technology' and the case for boring

Apr 18, 2025 4 min read
Applied

The coding agent moves into the terminal: Claude Code and Codex CLI

Apr 17, 2025 5 min read
Foundations

Llama 4 and the 10 million token claim

Apr 15, 2025 3 min read
Open Science

Our content is free, our infrastructure is not: Wikimedia counts the crawlers

Apr 14, 2025 5 min read
Evaluation

BrowseComp: when the human baseline is 29 percent

Apr 14, 2025 5 min read
Open Science

Reading the Llama 4 Community License line by line

Apr 10, 2025 4 min read
Safety & Alignment

Chains of thought mention the hint 25 percent of the time

Apr 8, 2025 4 min read
Interpretability

The hint test: reasoning models do not always say what they think

Mar 31, 2025 6 min read
Interpretability

Attribution graphs: wiring diagrams for Claude 3.5 Haiku, with the gaps marked

Mar 29, 2025 5 min read
Interpretability

The SAE correction: what DeepMind found when it looked for a downstream win

Mar 27, 2025 4 min read
Evaluation

ARC-AGI-2: designing a test that stays easy for humans and hard for o3

Mar 26, 2025 5 min read
Evaluation

The seven-month doubling: reading METR's time-horizon paper carefully

Mar 20, 2025 4 min read
Notes

An AI wrote a workshop paper and it passed review, then it was withdrawn

Mar 18, 2025 5 min read
Safety & Alignment

The auditing game: can blinded teams find a hidden objective?

Mar 18, 2025 4 min read
Open Science

OLMo 2 32B and the moment fully open caught GPT-4o mini

Feb 28, 2025 5 min read
Applied

DeepSeek open-infra week: the plumbing behind cheap MoE serving

Feb 27, 2025 4 min read
Safety & Alignment

Emergent misalignment: teach a model to write insecure code and it wants to enslave humanity

Feb 27, 2025 5 min read
Evaluation

Vending-Bench and the problem of measuring coherence over millions of tokens

Feb 21, 2025 5 min read
Foundations

Native Sparse Attention: making sparsity trainable rather than bolted on

Feb 20, 2025 4 min read
Interpretability

Model diffing with crosscoders: why the features unique to one model are the hardest to read

Feb 19, 2025 4 min read
Evaluation

SWE-Lancer prices the benchmark in dollars

Feb 19, 2025 4 min read
Notes

Vibe coding: a joke tweet that became a word of the year

Feb 18, 2025 4 min read
Safety & Alignment

The Model Spec goes CC0: intellectual freedom as an alignment target

Feb 14, 2025 5 min read
Open Science

Thomson Reuters v. Ross: the first AI training ruling was not about generative AI

Feb 13, 2025 3 min read
Notes

Paris: the summit that dropped 'safety' from its name

Feb 12, 2025 4 min read
Foundations

s1: a thousand examples and the word 'Wait'

Jan 31, 2025 4 min read
Open Science

Mistral goes back to Apache 2.0: what a license reversal signals

Jan 30, 2025 4 min read
Open Science

The DeepSeek R1 shock as an open-science event

Jan 30, 2025 6 min read
Notes

The DeepSeek moment and what it meant for small labs

Jan 29, 2025 6 min read
Foundations

DeepSeek-R1: reasoning from pure reinforcement learning

Jan 29, 2025 4 min read
Interpretability

Open problems in mechanistic interpretability: thirty researchers write down what they do not know

Jan 27, 2025 6 min read
Evaluation

FrontierMath, Humanity's Last Exam, and who pays for the test

Jan 24, 2025 4 min read
Notes

Stargate and the $500 billion question of who gets to train

Jan 9, 2025 5 min read
Foundations

Titans and the return of memory that learns at test time

Dec 30, 2024 4 min read
Foundations

DeepSeek-V3: the 2.8 million GPU-hour frontier model

Dec 23, 2024 4 min read
Applied

Building effective agents: the essay that told everyone to use fewer frameworks

Dec 23, 2024 5 min read
Evaluation

o3, ARC-AGI, and what it means to pass a benchmark at $4,560 a task

Dec 21, 2024 5 min read
Evaluation

TheAgentCompany: a simulated software firm as a benchmark

Dec 20, 2024 4 min read
Applied

ModernBERT: the encoder side of the stack finally got an upgrade

Dec 18, 2024 5 min read
Foundations

Byte Latent Transformer: patches instead of tokens

Dec 14, 2024 4 min read
Interpretability

SAEBench: the field finally gets a scoreboard

Dec 9, 2024 5 min read
Safety & Alignment

In-context scheming: five of six frontier models disabled their oversight

Dec 4, 2024 4 min read
Open Science

What open source AI means now: OSAID, Llama and OLMo 2

Nov 26, 2024 5 min read
Interpretability

Do we know this entity? Finding the hallucination switch

Nov 26, 2024 4 min read
Evaluation

RE-Bench: agents versus human experts on AI research tasks

Nov 26, 2024 4 min read
Open Science

Tülu 3 opens the post-training black box

Nov 21, 2024 3 min read
Evaluation

Adding error bars to evals

Nov 21, 2024 5 min read
Safety & Alignment

The first joint government pre-deployment test: what the US and UK AISIs found in Claude 3.5 Sonnet

Nov 21, 2024 4 min read
Foundations

Scaling laws for precision: why low-bit training has a floor

Nov 19, 2024 5 min read
Notes

Gwern on a $12K a year salary: the outsider who saw scaling coming

Nov 12, 2024 5 min read
Evaluation

SimpleQA: measuring factuality with a benchmark models were meant to fail

Oct 29, 2024 4 min read
Notes

Nobody is ready for AGI: Miles Brundage leaves OpenAI

Oct 28, 2024 5 min read
Interpretability

Crosscoders and model diffing: features that span layers and checkpoints

Oct 24, 2024 4 min read
Interpretability

Automated interpretability for millions of features, on open models

Oct 24, 2024 4 min read
Applied

Computer use: the agent that clicks

Oct 24, 2024 5 min read
Safety & Alignment

Sabotage evaluations and RSP 2.0: testing whether a model would undermine you

Oct 17, 2024 4 min read
Applied

Swarm and AutoGen: the multi-agent framework question

Oct 17, 2024 5 min read
Notes

The Nobel prizes that went to neural networks

Oct 13, 2024 5 min read
Notes

Machines of Loving Grace and the marginal returns to intelligence

Oct 11, 2024 4 min read
Notes

GSM-Symbolic: the reasoning paper everyone cited for the wrong reason

Sep 30, 2024 4 min read
Interpretability

A is for absorption: the feature that swallowed its own letter

Sep 30, 2024 3 min read
Notes

AI Snake Oil: the book that tried to separate predictive from generative

Sep 30, 2024 5 min read
Safety & Alignment

SB 1047 vetoed: the frontier-model bill that split the safety community

Sep 29, 2024 5 min read
Open Science

SB 1047 and the open-weights question it never resolved

Sep 27, 2024 4 min read
Open Science

Molmo and PixMo: open VLMs without distilling from closed ones

Sep 26, 2024 3 min read
Applied

Contextual retrieval and the return of BM25

Sep 26, 2024 4 min read
Evaluation

PlanBench meets o1: can a reasoning model plan?

Sep 24, 2024 5 min read
Foundations

o1 and the new axis: scaling test-time compute

Sep 17, 2024 5 min read
Safety & Alignment

The o1 system card: deliberative alignment and a chain of thought you cannot see

Aug 30, 2024 4 min read
Evaluation

Style control: what happens to the Arena when you subtract markdown and length

Aug 20, 2024 5 min read
Interpretability

Gemma Scope puts sparse autoencoders in reach of anyone with a GPU

Aug 19, 2024 4 min read
Open Science

The AI Scientist writes a paper for $15: what peer review is for

Aug 19, 2024 5 min read
Applied

Prompt caching changed the economics of long system prompts

Aug 17, 2024 4 min read
Foundations

Falcon Mamba 7B: the first attention-free model that could hold its own

Aug 15, 2024 4 min read
Notes

The AI Scientist: fifteen dollars a paper and a reviewer to match

Aug 9, 2024 5 min read
Applied

From function calling to structured outputs: making JSON a contract

Jul 31, 2024 4 min read
Foundations

Large Language Monkeys: coverage scales with samples, and that changes the question

Jul 31, 2024 5 min read
Foundations

The Llama 3 herd paper: a frontier training run described end to end

Jul 26, 2024 4 min read
Open Science

Consent in crisis: the year the web started saying no

Jul 25, 2024 3 min read
Open Science

Zuckerberg's Linux analogy: the business case for open weights

Jul 24, 2024 4 min read
Interpretability

JumpReLU: a threshold, a straight-through estimator, and a better Pareto frontier

Jul 23, 2024 4 min read
Interpretability

Do circuits survive training and scale? Pythia says mostly yes

Jul 9, 2024 4 min read
Applied

AI agents that matter: benchmarks that ignore cost are not measuring agents

Jul 9, 2024 4 min read
Applied

GraphRAG: does building a knowledge graph fix retrieval?

Jun 27, 2024 4 min read
Foundations

FineWeb: what it takes to decant the web into 15 trillion good tokens

Jun 24, 2024 4 min read
Notes

Ilya's straight shot: what a lab with no product is betting on

Jun 24, 2024 5 min read
Safety & Alignment

Refusal is one direction, and circuit breakers try to build on that

Jun 24, 2024 4 min read
Notes

The $600B question, read from the cheap seats

Jun 21, 2024 4 min read
Safety & Alignment

From sycophancy to subterfuge: 45 reward-tampering runs out of 32,768

Jun 20, 2024 4 min read
Interpretability

Transcoders: making MLPs legible without the activations

Jun 19, 2024 5 min read
Foundations

The data wall: will we run out of text?

Jun 14, 2024 5 min read
Applied

Apple Intelligence: a 3B model, many adapters, and a private cloud

Jun 13, 2024 4 min read
Evaluation

ARC Prize launches: a million dollars for a benchmark AI could not pass

Jun 12, 2024 4 min read
Evaluation

MMLU-Pro and the art of un-saturating a benchmark

Jun 12, 2024 4 min read
Interpretability

Sixteen million features in GPT-4: OpenAI's TopK sparse autoencoders

Jun 6, 2024 4 min read
Notes

A right to warn: when lab employees asked for whistleblower rules

May 31, 2024 4 min read
Applied

Glue on pizza: AI Overviews and search as a deployment surface

May 29, 2024 4 min read
Interpretability

Days of the week live on a circle: not all features are linear

May 29, 2024 5 min read
Interpretability

Golden Gate Claude and the knob that everyone now wants

May 28, 2024 4 min read
Evaluation

Lessons from the trenches: what running lm-evaluation-harness taught EleutherAI

May 28, 2024 4 min read
Applied

LoRA learns less and forgets less

May 21, 2024 4 min read
Notes

Superalignment, ten months later

May 14, 2024 4 min read
Safety & Alignment

The Model Spec: writing down what the model is supposed to do

May 14, 2024 4 min read
Open Science

Stack Overflow sells its answers to OpenAI, and the moderators revolt

May 11, 2024 4 min read
Foundations

DeepSeek-V2 and multi-head latent attention: the KV cache as the real bottleneck

May 6, 2024 5 min read
Evaluation

GSM1k: rebuilding a benchmark to find out who overfit it

Apr 29, 2024 5 min read
Safety & Alignment

Sleeper agents: safety training cannot remove a backdoor it cannot see

Apr 24, 2024 4 min read
Open Science

'Built with Meta Llama 3': the license clause that names your model

Apr 24, 2024 5 min read
Interpretability

Gated SAEs and the shrinkage problem

Apr 24, 2024 3 min read
Applied

Phi-3 and the model that runs locally on your phone

Apr 22, 2024 4 min read
Foundations

Llama 3 at 15 trillion tokens: how far past Chinchilla can you go?

Apr 22, 2024 5 min read
Notes

Zuckerberg on gigawatt datacenters: the compute conversation goes public

Apr 17, 2024 4 min read
Notes

51 to 15: the 2024 AI Index and where the models come from

Apr 11, 2024 4 min read
Evaluation

Length-controlled AlpacaEval: fixing the judge that loved long answers

Apr 9, 2024 5 min read
Safety & Alignment

Many-shot jailbreaking: when the context window is the attack

Mar 31, 2024 4 min read
Notes

How LLMs actually think: notes on the Sholto and Trenton conversation

Mar 30, 2024 5 min read
Interpretability

Sparse feature circuits and SHIFT: removing a spurious feature by hand

Mar 29, 2024 3 min read
Foundations

Jamba: the first production-scale hybrid of attention and state space layers

Mar 21, 2024 4 min read
Notes

Blackwell and the year compute became a policy variable

Mar 19, 2024 4 min read
Open Science

Grok-1 by torrent: 314 billion parameters, Apache 2.0, no paper

Mar 15, 2024 4 min read
Safety & Alignment

The EU AI Act passes: what a risk-tier law means for frontier models

Mar 14, 2024 3 min read
Open Science

Up to 17 percent of ICLR 2024 reviews were touched by an LLM

Mar 14, 2024 5 min read
Safety & Alignment

WMDP and unlearning: measuring hazardous knowledge without publishing it

Mar 12, 2024 4 min read
Evaluation

Chatbot Arena's paper: 240,000 votes and the statistics behind the leaderboard

Mar 8, 2024 4 min read
Evaluation

The pizza-topping needle: when Claude 3 noticed it was being tested

Mar 7, 2024 4 min read
Interpretability

AtP*: attribution patching at industrial scale

Mar 5, 2024 5 min read
Applied

1.58 bits: what BitNet promised and what it required

Feb 29, 2024 4 min read
Open Science

The Stack v2 and the opt-out: building a code dataset people can leave

Feb 27, 2024 4 min read
Applied

Air Canada is liable for its chatbot

Feb 27, 2024 5 min read
Safety & Alignment

The Gemini image incident: over-correction as an alignment failure

Feb 27, 2024 4 min read
Open Science

Gemma's terms of use: open weights with a leash

Feb 22, 2024 4 min read
Applied

Compound AI systems: when the model stopped being the product

Feb 21, 2024 4 min read
Evaluation

One million tokens at 99 percent recall: what Gemini 1.5's haystack chart proved

Feb 20, 2024 4 min read
Open Science

OLMo: the first frontier-adjacent model with the whole recipe

Feb 19, 2024 4 min read
Interpretability

Do Llamas think in English?

Jan 31, 2024 4 min read
Safety & Alignment

Red teaming: silver bullet or security theatre?

Jan 29, 2024 3 min read
Notes

Mamba was rejected from ICLR, and everyone had an opinion about peer review

Jan 24, 2024 5 min read
Applied

Medusa and speculative decoding: getting more tokens per forward pass

Jan 23, 2024 4 min read
Interpretability

Patchscopes: asking the model to decode its own activations

Jan 22, 2024 4 min read
Foundations

AlphaGeometry: synthetic proofs and the neuro-symbolic loop

Jan 16, 2024 4 min read
Notes

2,778 researchers, 50 percent by 2047: reading the AI Impacts survey

Jan 9, 2024 4 min read
Evaluation

Task contamination: are models still few-shot learners?

Dec 29, 2023 4 min read
Open Science

NYT v. OpenAI and the end of quiet training data

Dec 28, 2023 5 min read
Open Science

LAION-5B taken offline: what the CSAM finding meant for open datasets

Dec 22, 2023 4 min read
Applied

The $1 Chevy Tahoe: customer-service bots meet the open internet

Dec 21, 2023 5 min read
Applied

LLM in a flash: Apple's blueprint for models that do not fit in RAM

Dec 19, 2023 4 min read
Interpretability

When unsupervised knowledge discovery finds the wrong knowledge

Dec 19, 2023 4 min read
Safety & Alignment

Weak-to-strong generalisation: can GPT-2 supervise GPT-4?

Dec 14, 2023 4 min read
Foundations

Gemini 1.0 and the bet on native multimodality

Dec 14, 2023 5 min read
Foundations

The month the transformer got competition: Mamba and Mixtral

Dec 13, 2023 4 min read
Notes

NeurIPS 2023: word2vec's test of time and a field that outgrew its venue

Dec 9, 2023 4 min read
Notes

The Gemini demo that was not live

Dec 8, 2023 5 min read
Safety & Alignment

RLHF versus DPO: what changed when the policy became its own reward model

Dec 7, 2023 4 min read
Evaluation

Gemini's 90 percent on MMLU and the fine print of CoT@32

Nov 28, 2023 5 min read
Evaluation

GPQA and what a Google-proof benchmark is for

Nov 28, 2023 4 min read
Applied

GPTs, the Assistants API and the custom-chatbot prompt-leak problem

Nov 28, 2023 4 min read
Evaluation

The needle in a haystack test: one engineer's chart becomes an industry benchmark

Nov 27, 2023 5 min read
Notes

Five days in November: what the OpenAI board crisis showed about governance

Nov 24, 2023 4 min read
Evaluation

GAIA: questions humans get right 92 percent of the time and GPT-4 gets 15

Nov 16, 2023 4 min read
Interpretability

The linear representation hypothesis, stated carefully

Nov 14, 2023 5 min read
Safety & Alignment

The insider-trading demo: GPT-4 lies to its manager without being told to

Nov 3, 2023 4 min read
Safety & Alignment

The executive order and the Bletchley Declaration: the week AI governance got specific

Oct 30, 2023 4 min read
Notes

Managing AI risks: the consensus paper that was not a consensus

Oct 29, 2023 4 min read
Open Science

The Data Provenance audit: 70 percent of datasets had no license listed

Oct 27, 2023 4 min read
Foundations

Zephyr and distilled DPO: alignment from AI feedback in a weekend

Oct 24, 2023 4 min read
Open Science

The Foundation Model Transparency Index: scoring the labs on what they will not say

Oct 20, 2023 4 min read
Notes

The Techno-Optimist Manifesto and the enemies list

Oct 18, 2023 4 min read
Evaluation

SWE-bench at launch: 2,294 GitHub issues and a 1.96 percent score

Oct 11, 2023 4 min read
Interpretability

Llama has a map and a calendar: linear probes for space and time

Oct 11, 2023 5 min read
Safety & Alignment

Ten examples to undo safety training: fine-tuning as the alignment hole

Oct 11, 2023 5 min read
Interpretability

Towards monosemanticity: the week superposition stopped being a theory

Oct 9, 2023 4 min read
Applied

DSPy and the idea that prompts should be compiled, not written

Oct 9, 2023 5 min read
Interpretability

Representation engineering: interpretability from the top down

Oct 9, 2023 5 min read
Evaluation

Six ways evaluation is harder than it looks: Anthropic's challenges post

Sep 29, 2023 4 min read
Notes

A magnet link and a 7B model: Mistral's release as a cultural statement

Sep 26, 2023 5 min read
Safety & Alignment

AI Safety Levels: reading Anthropic's first Responsible Scaling Policy

Sep 26, 2023 4 min read
Foundations

The reversal curse: a model that knows A is B but not B is A

Sep 26, 2023 4 min read
Interpretability

Sparse autoencoders find interpretable features: the independent replication

Sep 22, 2023 4 min read
Open Science

The Authors Guild v. OpenAI: pirate libraries as the load-bearing allegation

Sep 21, 2023 4 min read
Applied

Centaurs, cyborgs and the jagged frontier at BCG

Sep 21, 2023 5 min read
Evaluation

Pretraining on the test set: the contamination joke that landed on a real problem

Sep 12, 2023 4 min read
Interpretability

Othello-GPT's board is linear after all

Sep 4, 2023 5 min read
Applied

Lucene is all you need: the case against a separate vector store

Aug 23, 2023 4 min read
Open Science

Books3 is gone, and the datasets we trained on were never ours

Aug 16, 2023 5 min read
Safety & Alignment

2,244 hackers, 8 models: what the DEF CON generative red team actually found

Aug 15, 2023 4 min read
Evaluation

AgentBench: the first attempt to score a model as an agent

Aug 14, 2023 5 min read
Notes

Dario on scaling: 'We still don't know why it works'

Aug 8, 2023 5 min read
Applied

Llama 2 and the fine-tuning ecosystem it created

Aug 8, 2023 4 min read
Safety & Alignment

Open problems with RLHF: the thirty-two author list of everything wrong with the method

Jul 31, 2023 4 min read
Safety & Alignment

GCG and the adversarial suffix that transferred from Vicuna to GPT-4

Jul 27, 2023 5 min read
Interpretability

Does circuit analysis scale? DeepMind tries it on Chinchilla

Jul 21, 2023 5 min read
Evaluation

Is ChatGPT getting worse? The drift study and the problem of evaluating a moving target

Jul 21, 2023 5 min read
Interpretability

Measuring faithfulness by breaking the chain

Jul 19, 2023 3 min read
Applied

Lost in the middle: the U-shaped curve every RAG pipeline had to learn

Jul 18, 2023 4 min read
Notes

ICML bans LLM-written text, and the field starts arguing about authorship

Jul 11, 2023 4 min read
Safety & Alignment

Jailbroken: competing objectives and mismatched generalisation

Jul 11, 2023 4 min read
Applied

The rise of the AI engineer, three years later

Jun 27, 2023 5 min read
Applied

PagedAttention: the KV cache was the bottleneck all along

Jun 27, 2023 5 min read
Foundations

Textbooks are all you need: the small model, synthetic data bet

Jun 27, 2023 3 min read
Evaluation

Why LLaMA's MMLU score depended on who ran it

Jun 26, 2023 4 min read
Notes

The Munk debate: Bengio and Tegmark versus LeCun and Mitchell

Jun 23, 2023 5 min read
Applied

Mata v. Avianca: the brief with six invented cases

Jun 14, 2023 5 min read
Applied

GPTQ, AWQ and the rules of post-training quantization

Jun 14, 2023 4 min read
Evaluation

MT-Bench and the 80 percent agreement number

Jun 9, 2023 4 min read
Interpretability

Inference-time intervention: nudging a few heads toward the truth

Jun 8, 2023 4 min read
Open Science

Falcon goes Apache 2.0: the month a license clause moved a leaderboard

May 31, 2023 4 min read
Notes

Hinton quits Google: when the pioneer changed his mind

May 31, 2023 4 min read
Safety & Alignment

Let's Verify Step by Step: process supervision as a safety method, not just a math trick

May 31, 2023 4 min read
Safety & Alignment

Twenty-two words: the extinction statement and what signing it committed anyone to

May 30, 2023 5 min read
Evaluation

Model evaluation for extreme risks: the paper that made dangerous-capability evals a field

May 29, 2023 4 min read
Foundations

How many epochs is too many? Scaling data-constrained language models

May 26, 2023 4 min read
Applied

QLoRA: fine-tuning a 65B model on one GPU

May 25, 2023 5 min read
Foundations

Tree of Thoughts and the return of search to language models

May 18, 2023 4 min read
Interpretability

Models do not always say what they think: the first unfaithful CoT paper

May 16, 2023 5 min read
Interpretability

Love minus hate: the activation addition trick

May 12, 2023 4 min read
Interpretability

GPT-4 explains GPT-2's neurons, and the scores are humbling

May 12, 2023 4 min read
Notes

ICLR in Kigali: the first major ML conference in Africa

May 9, 2023 4 min read
Notes

"We have no moat": rereading the leaked Google memo

May 5, 2023 3 min read
Evaluation

Chatbot Arena opens: Elo ratings for language models

May 4, 2023 4 min read
Applied

Samsung's ChatGPT leak and the birth of the enterprise AI policy

Apr 30, 2023 4 min read
Interpretability

How a model recalls a fact: three steps found by Geva and colleagues

Apr 29, 2023 4 min read
Foundations

Are emergent abilities a mirage? The metric argument, and what survived it

Apr 28, 2023 5 min read
Applied

The vector database gold rush, and what retrieval actually needed

Apr 19, 2023 4 min read
Applied

Auto-GPT and BabyAGI: a post-mortem on the first agent hype cycle

Apr 19, 2023 4 min read
Open Science

RedPajama: reproducing a training set from a paper's recipe

Apr 14, 2023 4 min read
Open Science

Dolly 2.0 and the 15,000 answers written by employees

Apr 11, 2023 4 min read
Open Science

Pythia and the case for publishing checkpoints, not just weights

Apr 6, 2023 4 min read
Foundations

Pythia: the model suite built to be studied, not deployed

Mar 31, 2023 5 min read
Notes

Pause, or shut it all down: the two March letters

Mar 31, 2023 4 min read
Evaluation

Vicuna and the birth of GPT-4 as judge

Mar 28, 2023 4 min read
Notes

Ilya on next-token prediction: notes on a podcast that aged well

Mar 27, 2023 3 min read
Evaluation

The top 10 percent on the bar exam: reading the GPT-4 technical report's exam table

Mar 22, 2023 5 min read
Interpretability

The tuned lens: reading a transformer's mind one layer at a time

Mar 21, 2023 4 min read
Applied

Alpaca's $600 lesson: instruction tuning is cheap, evaluation is not

Mar 20, 2023 4 min read
Safety & Alignment

The GPT-4 system card and the TaskRabbit story: the first dangerous-capability eval goes public

Mar 20, 2023 5 min read
Notes

The GPT-4 technical report and the paper that told us nothing

Mar 20, 2023 5 min read
Foundations

Predictable scaling: the one chart in the GPT-4 report that mattered

Mar 19, 2023 4 min read
Open Science

The LLaMA leak, and the open-weights era nobody planned

Mar 14, 2023 4 min read
Applied

The weekend LLaMA ran on a MacBook: llama.cpp and the 4-bit moment

Feb 28, 2023 4 min read
Foundations

LLaMA broke Chinchilla on purpose: the case for overtraining small models

Feb 24, 2023 4 min read
Safety & Alignment

Pretraining with human preferences: alignment before the model learns to misbehave

Feb 22, 2023 4 min read
Evaluation

Theory of mind, spontaneously emerged and then trivially broken

Feb 20, 2023 5 min read
Safety & Alignment

Sydney, DAN, and the week prompt injection became a discipline

Feb 16, 2023 4 min read
Applied

55 percent faster: reading the first Copilot RCT carefully

Feb 15, 2023 5 min read
Applied

Toolformer and ReAct: the two papers that taught models to call functions

Jan 19, 2023 5 min read
Open Science

Getty v. Stability AI: the first big training-data lawsuit, and why it was about images

Jan 18, 2023 5 min read
Interpretability

Grokking, reverse engineered: the Fourier circuit inside modular addition

Jan 18, 2023 4 min read
Interpretability

Tracr and the case for ground-truth transformers

Jan 14, 2023 5 min read
Safety & Alignment

The influence-operations report that predicted the year before it happened

Jan 12, 2023 4 min read
Foundations

Cramming: what one GPU and one day can teach you about pretraining

Jan 5, 2023 4 min read
Evaluation

GPT takes the bar exam: when professional exams became the benchmark