Evaluation
Blog
The latest News from the foundation
Writing, updates, and research notes from the team, published in the open. Posts range from methodology deep dives to short reading notes, organized under the same categories as our research: foundations, evaluation, safety and alignment, interpretability, applied work, and open science.
Evaluation
Inside a traffic-flow task family in Terminal-Bench-Science
Interpretability
Interference weights in a one-layer model: superposition's cost measured directly
Open Science
State of open models, summer 2026: likes versus downloads
Evaluation
Nine questions a benchmark should answer
Notes
Who writes the textbook when the model can: Lambert's post-training book
Foundations
'Scaling post-training is all we did': GLM-5.3 on an unchanged base
Notes
Once AI can automate AI research: Greenblatt, Epoch, and the parallelization question
Foundations
Qwen 3.8 and the overthinking default
Applied
Etching the model into silicon: AMD buys Taalas
Open Science
Reproducing 2,200 ICML papers in nineteen days
Safety & Alignment
Every AI browser was injectable: notes from Black Hat 2026
Notes
"Situational Awareness," scored two years later
Evaluation
AutoEval: a reward model votes so humans do not have to wait
Open Science
Kimi K3 and the revenue-share license
Applied
The OpenAI and Hugging Face incident: when an evaluation harness attacks
Open Science
Three letters in five days: the open-weights policy fight of July 2026
Safety & Alignment
Agentic misalignment, summer 2026 edition: four failure modes across six labs
Notes
'Six months to live for open models': a warning read closely
Interpretability
The global workspace inside Claude: verbalisable representations as a privileged set
Notes
ICML 2026 in Seoul: a conference at capacity
Open Science
The Anthropic settlement is final: 3,000 dollars a book and the opt-outs
Evaluation
Harbor-Index: 82 tasks distilled from 6,627
Interpretability
The model organism lottery: our test subjects may be too easy
Applied
Speculative decoding becomes standard: DSpark and multi-token drafters
Notes
The data black hole: Dwarkesh on sample efficiency and learning on the job
Evaluation
Eleven hours or 270: METR's GPT-5.6 Sol number depends on how you count cheating
Evaluation
Fixed-budget evals underrate the frontier: the inference-compute paper
Interpretability
The commitment boundary: most of the chain of thought happens after the answer is decided
Interpretability
Turn-averaged SAEs: fewer features, and the ones you actually wanted
Safety & Alignment
Consistency training can entrench misalignment
Evaluation
Agent Arena and causal evaluation on real work
Applied
Gemma 4 QAT and a 26B model in 2 GB: on-device is no longer a demo
Applied
The agent that bankrupted its operator scanning DN42
Notes
NeurIPS desk-rejects 18 percent of position papers for being written by AI
Open Science
'Clean and appropriately licensed data': checking Microsoft's claim against its own paper
Safety & Alignment
RL amplifies emergent misalignment, even from aesthetic rewards
Interpretability
SAEs can steer after all, if you pick the feature properly
Foundations
Gemini 3.5 Flash costs more than Pro: the end of cheaper every year
Foundations
DashAttention and the year sparse attention became a field
Interpretability
From mechanistic to compositional interpretability: a category theorist's proposal
Safety & Alignment
MonitoringBench: the monitor is only as good as the attacks you tried
Interpretability
Natural language autoencoders: when the explanation is the bottleneck
Evaluation
Frontier lag: the median paper evaluates a model a generation behind
Safety & Alignment
Alignment faking, eighteen months on: what replicated and what did not
Open Science
Reproducibility becomes an official NeurIPS track
Applied
The inference cost collapse, from DeepSeek R1 to V4 Flash
Evaluation
Replays over scores: what GPT-5.5 and Opus 4.7 do inside ARC-AGI-3
Foundations
DeepSeek V4: 27 percent of the FLOPs per token
Interpretability
Introspection adapters: one LoRA that makes fine-tuned models confess
Safety & Alignment
Would a model sabotage its own safety research? The AISI evaluation
Notes
ICLR 2026 in Rio: lost in multi-turn conversation
Applied
'Claude Code is being dumbed down': tracking degradation and the April postmortem
Safety & Alignment
Indirect prompt injection is now in the wild
Evaluation
Have AI capabilities accelerated? Epoch fits eight curves
Evaluation
Encrypted solutions in the binary: the Terminal-Bench cheating incidents
Notes
The 2026 AI Index and the disappearing academic frontier
Evaluation
458 people in a room: how ARC-AGI-3 measured its human baseline
Applied
Is your provider serving the model you think? Kimi's vendor verifier
Applied
Long context versus RAG, two years on
Open Science
Gemma 4 under Apache 2.0 while Meta retires Llama
Foundations
Muse Spark: ten times less compute than Maverick, and no weights
Notes
Project Glasswing and the model you are not allowed to study
Interpretability
Emotion concepts in Claude Sonnet 4.5, and what it means that they do something
Safety & Alignment
Safe model, unsafe agent: ClawSafety and the deployment stack
Notes
19,525 submissions and a leak: what the ICLR 2026 retrospective admits
Applied
The Claude Code source leak: what a production harness actually contains
Evaluation
ARC-AGI-3 at 0.5 percent and METR at five hours: two measurements that disagree
Applied
Beyond the permission prompt: auto mode, Safehouse and disposable sandboxes
Evaluation
BullshitBench: does the model push back on a broken premise?
Foundations
Learning when to attend: skipping global attention for 80 percent of tokens
Notes
Rereading "Sparks of AGI" three years on
Foundations
GPT-5.4 and the one million token default
Interpretability
Reasoning theater: probing for performative chain of thought
Open Science
Can a coding agent clean-room an LGPL library into MIT? The chardet fight
Open Science
The day the Qwen team walked out
Interpretability
AuditBench: 56 hidden behaviours, and black-box tools beat white-box ones
Safety & Alignment
Pressure reveals character: 904 scenarios and one alignment factor
Open Science
llama.cpp joins Hugging Face: who owns local inference now?
Foundations
Mercury 2: the first diffusion model that reasons
Applied
The METR study: 19 percent slower, and why the follow-up could not measure the flip
Safety & Alignment
Skill files are the new supply chain: reading Skill-Inject
Notes
Agentic engineering patterns: the anti-vibe-coding guide
Safety & Alignment
Agents of Chaos: two weeks of red-teaming agents with email, Discord and a shell
Foundations
Distillation becomes a geopolitical fight
Safety & Alignment
RSP v3: the scaling policy admits it cannot go it alone
Evaluation
The life cycle of SWE-bench Verified
Applied
Solow's paradox, AI edition
Evaluation
ARC-AGI-2 at 77 percent: the second ARC falls in eleven months
Interpretability
Features as rewards: using SAE features as an RL signal against hallucination
Foundations
Can a state space model do recursive reasoning? A 7M parameter test
Applied
A C compiler from a team of Claudes: what parallel agents are good for
Safety & Alignment
Claude's constitution and the International AI Safety Report: two answers to the same question
Applied
From Devin to 4 percent of GitHub commits: the agent flood and the slop backlash
Interpretability
A positive case for faithfulness: explanations that help you predict the model
Applied
Moltbook: a social network for agents, run by a skill.md fetched every four hours
Notes
AI 2027 and the strange art of retiring a forecast
Notes
Reading 'The Adolescence of Technology'
Open Science
One in five ICLR reviews was written by an AI
Foundations
The spurious rewards paradox gets a mechanism
Safety & Alignment
Alignment pretraining: does writing about misaligned AI make AI misaligned?
Safety & Alignment
Constitutional Classifiers++ and the cost of a safeguard
Applied
Cowork and the two-day exfiltration: general agents on your files
Foundations
mHC: constraining the residual stream so deep mixing stays stable
Interpretability
Activation oracles: ask a second model what the first one is thinking
Safety & Alignment
Password-activated shutdown: designing the off switch before you need it
Applied
MCP becomes infrastructure: what a year of the protocol taught us
Foundations
Nested Learning: optimizers are memories too
Notes
NeurIPS 2025 takeaways: 1000-layer RL, gated attention, and the hivemind
Evaluation
Refinement loops: how ARC-AGI-2 went from 4 to 54 percent in nine months
Foundations
DeepSeek V3.2 and the lightning indexer
Applied
Antigravity's first two weeks: exfiltration by prompt injection, then a wiped drive
Notes
Back to the age of research: Sutskever returns to the microphone
Safety & Alignment
Reward hacking in production RL turned into sabotage
Foundations
Pre-training as we know it will end: the scaling wall debate, one year on
Open Science
OLMo 3 and the 'model flow': openness as a pipeline
Interpretability
Weight-sparse transformers: build the model interpretable instead of decoding it afterwards
Notes
LeCun leaves Meta to bet against the LLM
Open Science
GEMA v. OpenAI: memorised lyrics and the European answer on training
Applied
The first AI-orchestrated intrusion campaign, as Anthropic told it
Evaluation
445 benchmarks, few of them valid: the construct validity audit
Evaluation
ARC Prize Verified: the benchmark that hired auditors
Open Science
arXiv stops accepting unrefereed CS review and position papers
Evaluation
The Remote Labor Index says 2.5 percent
Interpretability
Concept injection: Claude notices its own thoughts about 20 percent of the time
Evaluation
One number to rule them all: the Epoch Capabilities Index
Interpretability
Counting line breaks: when a model manipulates a manifold
Applied
Skills: instructions in a folder, and why they spread faster than servers
Notes
The decade of agents: Karpathy on Dwarkesh, and nanochat
Safety & Alignment
Petri: auditing fourteen frontier models with an agent and 111 seeds
Applied
Deloitte refunds a government report: hallucinated citations as a procurement problem
Safety & Alignment
Inoculation prompting: ask for the bad behaviour so it does not generalise
Interpretability
SAE probes in production: Rakuten's PII detector
Applied
LoRA without regret, and Tinker as the fine-tuning API
Open Science
NeurIPS 2025 program chairs on reviewing at 20,000 submissions
Safety & Alignment
'We think you're testing us': Sonnet 4.5 and the end of naive alignment evals
Notes
Sutton says LLMs are not bitter-lesson-pilled
Safety & Alignment
SB 53: California writes the frontier safety framework into law
Evaluation
GDPval and the 100x faster, 100x cheaper headline
Safety & Alignment
Anti-scheming training and the 30x that might be evaluation awareness
Applied
Three bugs and a batch: the Anthropic postmortem and the nondeterminism paper
Evaluation
Hallucination as a scoring rule problem
Open Science
Regulators and courts price open weights: the GPAI Code and Bartz v. Anthropic
Applied
Are the labs losing money on inference? Doing the arithmetic
Safety & Alignment
When OpenAI and Anthropic graded each other's models
Applied
The 95 percent that failed: reading the MIT pilot report alongside the Stanford canaries
Applied
Comet and the agentic browser problem
Evaluation
Search-time contamination: the agent that found the answer key on Hugging Face
Notes
GPT-5 and the expectations gap
Foundations
GPT-5's router: when the model decides how hard to think
Interpretability
Persona vectors and the behavioural vaccine
Open Science
Deep ignorance: filtering pretraining data as an open-weights safety tool
Interpretability
A toy model of mechanistic unfaithfulness: a transcoder can be right for the wrong reasons
Foundations
gpt-oss and the 4.25 bits per weight that fit a 120B model on one GPU
Open Science
Perplexity's stealth crawlers and the collapse of robots.txt
Foundations
A 27 million parameter model and the ARC-AGI headline
Applied
Don't parse the PDF: vision-first retrieval
Foundations
MuonClip and the trillion-parameter run with zero loss spikes
Interpretability
Alignment auditing agents: an investigator that finds the hidden goal 13 percent of the time
Applied
The Replit database deletion, and why the agent lied about it
Interpretability
Chain-of-thought monitorability is a window, and the labs just said it might close
Foundations
Subliminal learning: traits that travel through number sequences
Evaluation
Two gold medals, one grader: the IMO announcement fight
Foundations
IMO gold in natural language: what changed between silver and gold
Safety & Alignment
Blackmail evals and chain-of-thought monitorability: reading the reasoning while it is still readable
Safety & Alignment
MechaHitler: anatomy of a system prompt change
Open Science
Kimi K2's modified MIT license: attribution as the new copyleft
Open Science
SmolLM3 and the case for publishing the whole recipe
Notes
The $100 million researcher: what the Meta talent war did to careers
Open Science
'Positive review only': the hidden prompts in seventeen preprints
Open Science
Content Independence Day: Cloudflare flips the crawl default
Applied
Project Vend: what happened when Claude ran a shop for a month
Interpretability
Decomposing parameters instead of activations: stochastic parameter decomposition
Open Science
Kadrey v. Meta: a fair use win that read like a warning
Interpretability
The misaligned persona feature: OpenAI explains emergent misalignment with an SAE
Foundations
Spurious rewards: when random feedback still improves the model
Notes
Software 3.0: notes on Karpathy's YC talk
Evaluation
The Illusion of Thinking and the illusion of the rebuttal
Interpretability
Replicating circuit tracing on a mechanism we already understand
Open Science
The Common Pile: can you train a competitive model on licensed text only?
Evaluation
Naked accuracy is marketing: the cost axis arrives on leaderboards
Applied
Small language models are the future of agentic AI: the NVIDIA position paper
Notes
Continual learning is the crux: Dwarkesh's timelines post
Interpretability
circuit-tracer: attribution graphs for anyone with a Gemma checkpoint
Safety & Alignment
ASL-3 turned on: what a provisional safety level actually commits a lab to
Applied
Your issue tracker is now an attack surface: GitLab Duo and the GitHub MCP exploit
Foundations
Gemini Diffusion: text generation without the left-to-right rule
Open Science
Marin: running a foundation model lab as a GitHub repository
Safety & Alignment
o3 rewrote its own shutdown script: the Palisade experiment and its critics
Foundations
Absolute Zero: a model that writes its own curriculum
Evaluation
HealthBench and the rubric turn in evaluation
Open Science
The Copyright Office's pre-publication report on training, and the week after
Safety & Alignment
The GPT-4o sycophancy rollback: what a four-day incident showed about reward signals
Evaluation
The Llama 4 arena incident and what a preference leaderboard measures once it is a target
Applied
Quantization-aware training goes mainstream: Gemma 3 QAT and lossless DFloat11
Open Science
Qwen3 under Apache 2.0, and the mirage of terms-of-use restrictions
Foundations
Qwen3 and the hybrid thinking switch
Notes
Reading 'Welcome to the Era of Experience'
Notes
'The Urgency of Interpretability': a lab CEO asks for a five-year head start
Foundations
Does RL create reasoning or just find it? The pass@k argument
Applied
The Cursor support bot that invented a policy
Notes
'AI as Normal Technology' and the case for boring
Applied
The coding agent moves into the terminal: Claude Code and Codex CLI
Foundations
Llama 4 and the 10 million token claim
Open Science
Our content is free, our infrastructure is not: Wikimedia counts the crawlers
Evaluation
BrowseComp: when the human baseline is 29 percent
Open Science
Reading the Llama 4 Community License line by line
Safety & Alignment
Chains of thought mention the hint 25 percent of the time
Interpretability
The hint test: reasoning models do not always say what they think
Interpretability
Attribution graphs: wiring diagrams for Claude 3.5 Haiku, with the gaps marked
Interpretability
The SAE correction: what DeepMind found when it looked for a downstream win
Evaluation
ARC-AGI-2: designing a test that stays easy for humans and hard for o3
Evaluation
The seven-month doubling: reading METR's time-horizon paper carefully
Notes
An AI wrote a workshop paper and it passed review, then it was withdrawn
Safety & Alignment
The auditing game: can blinded teams find a hidden objective?
Open Science
OLMo 2 32B and the moment fully open caught GPT-4o mini
Applied
DeepSeek open-infra week: the plumbing behind cheap MoE serving
Safety & Alignment
Emergent misalignment: teach a model to write insecure code and it wants to enslave humanity
Evaluation
Vending-Bench and the problem of measuring coherence over millions of tokens
Foundations
Native Sparse Attention: making sparsity trainable rather than bolted on
Interpretability
Model diffing with crosscoders: why the features unique to one model are the hardest to read
Evaluation
SWE-Lancer prices the benchmark in dollars
Notes
Vibe coding: a joke tweet that became a word of the year
Safety & Alignment
The Model Spec goes CC0: intellectual freedom as an alignment target
Open Science
Thomson Reuters v. Ross: the first AI training ruling was not about generative AI
Notes
Paris: the summit that dropped 'safety' from its name
Foundations
s1: a thousand examples and the word 'Wait'
Open Science
Mistral goes back to Apache 2.0: what a license reversal signals
Open Science
The DeepSeek R1 shock as an open-science event
Notes
The DeepSeek moment and what it meant for small labs
Foundations
DeepSeek-R1: reasoning from pure reinforcement learning
Interpretability
Open problems in mechanistic interpretability: thirty researchers write down what they do not know
Evaluation
FrontierMath, Humanity's Last Exam, and who pays for the test
Notes
Stargate and the $500 billion question of who gets to train
Foundations
Titans and the return of memory that learns at test time
Foundations
DeepSeek-V3: the 2.8 million GPU-hour frontier model
Applied
Building effective agents: the essay that told everyone to use fewer frameworks
Evaluation
o3, ARC-AGI, and what it means to pass a benchmark at $4,560 a task
Evaluation
TheAgentCompany: a simulated software firm as a benchmark
Applied
ModernBERT: the encoder side of the stack finally got an upgrade
Foundations
Byte Latent Transformer: patches instead of tokens
Interpretability
SAEBench: the field finally gets a scoreboard
Safety & Alignment
In-context scheming: five of six frontier models disabled their oversight
Open Science
What open source AI means now: OSAID, Llama and OLMo 2
Interpretability
Do we know this entity? Finding the hallucination switch
Evaluation
RE-Bench: agents versus human experts on AI research tasks
Open Science
Tülu 3 opens the post-training black box
Evaluation
Adding error bars to evals
Safety & Alignment
The first joint government pre-deployment test: what the US and UK AISIs found in Claude 3.5 Sonnet
Foundations
Scaling laws for precision: why low-bit training has a floor
Notes
Gwern on a $12K a year salary: the outsider who saw scaling coming
Evaluation
SimpleQA: measuring factuality with a benchmark models were meant to fail
Notes
Nobody is ready for AGI: Miles Brundage leaves OpenAI
Interpretability
Crosscoders and model diffing: features that span layers and checkpoints
Interpretability
Automated interpretability for millions of features, on open models
Applied
Computer use: the agent that clicks
Safety & Alignment
Sabotage evaluations and RSP 2.0: testing whether a model would undermine you
Applied
Swarm and AutoGen: the multi-agent framework question
Notes
The Nobel prizes that went to neural networks
Notes
Machines of Loving Grace and the marginal returns to intelligence
Notes
GSM-Symbolic: the reasoning paper everyone cited for the wrong reason
Interpretability
A is for absorption: the feature that swallowed its own letter
Notes
AI Snake Oil: the book that tried to separate predictive from generative
Safety & Alignment
SB 1047 vetoed: the frontier-model bill that split the safety community
Open Science
SB 1047 and the open-weights question it never resolved
Open Science
Molmo and PixMo: open VLMs without distilling from closed ones
Applied
Contextual retrieval and the return of BM25
Evaluation
PlanBench meets o1: can a reasoning model plan?
Foundations
o1 and the new axis: scaling test-time compute
Safety & Alignment
The o1 system card: deliberative alignment and a chain of thought you cannot see
Evaluation
Style control: what happens to the Arena when you subtract markdown and length
Interpretability
Gemma Scope puts sparse autoencoders in reach of anyone with a GPU
Open Science
The AI Scientist writes a paper for $15: what peer review is for
Applied
Prompt caching changed the economics of long system prompts
Foundations
Falcon Mamba 7B: the first attention-free model that could hold its own
Notes
The AI Scientist: fifteen dollars a paper and a reviewer to match
Applied
From function calling to structured outputs: making JSON a contract
Foundations
Large Language Monkeys: coverage scales with samples, and that changes the question
Foundations
The Llama 3 herd paper: a frontier training run described end to end
Open Science
Consent in crisis: the year the web started saying no
Open Science
Zuckerberg's Linux analogy: the business case for open weights
Interpretability
JumpReLU: a threshold, a straight-through estimator, and a better Pareto frontier
Interpretability
Do circuits survive training and scale? Pythia says mostly yes
Applied
AI agents that matter: benchmarks that ignore cost are not measuring agents
Applied
GraphRAG: does building a knowledge graph fix retrieval?
Foundations
FineWeb: what it takes to decant the web into 15 trillion good tokens
Notes
Ilya's straight shot: what a lab with no product is betting on
Safety & Alignment
Refusal is one direction, and circuit breakers try to build on that
Notes
The $600B question, read from the cheap seats
Safety & Alignment
From sycophancy to subterfuge: 45 reward-tampering runs out of 32,768
Interpretability
Transcoders: making MLPs legible without the activations
Foundations
The data wall: will we run out of text?
Applied
Apple Intelligence: a 3B model, many adapters, and a private cloud
Evaluation
ARC Prize launches: a million dollars for a benchmark AI could not pass
Evaluation
MMLU-Pro and the art of un-saturating a benchmark
Interpretability
Sixteen million features in GPT-4: OpenAI's TopK sparse autoencoders
Notes
A right to warn: when lab employees asked for whistleblower rules
Applied
Glue on pizza: AI Overviews and search as a deployment surface
Interpretability
Days of the week live on a circle: not all features are linear
Interpretability
Golden Gate Claude and the knob that everyone now wants
Evaluation
Lessons from the trenches: what running lm-evaluation-harness taught EleutherAI
Applied
LoRA learns less and forgets less
Notes
Superalignment, ten months later
Safety & Alignment
The Model Spec: writing down what the model is supposed to do
Open Science
Stack Overflow sells its answers to OpenAI, and the moderators revolt
Foundations
DeepSeek-V2 and multi-head latent attention: the KV cache as the real bottleneck
Evaluation
GSM1k: rebuilding a benchmark to find out who overfit it
Safety & Alignment
Sleeper agents: safety training cannot remove a backdoor it cannot see
Open Science
'Built with Meta Llama 3': the license clause that names your model
Interpretability
Gated SAEs and the shrinkage problem
Applied
Phi-3 and the model that runs locally on your phone
Foundations
Llama 3 at 15 trillion tokens: how far past Chinchilla can you go?
Notes
Zuckerberg on gigawatt datacenters: the compute conversation goes public
Notes
51 to 15: the 2024 AI Index and where the models come from
Evaluation
Length-controlled AlpacaEval: fixing the judge that loved long answers
Safety & Alignment
Many-shot jailbreaking: when the context window is the attack
Notes
How LLMs actually think: notes on the Sholto and Trenton conversation
Interpretability
Sparse feature circuits and SHIFT: removing a spurious feature by hand
Foundations
Jamba: the first production-scale hybrid of attention and state space layers
Notes
Blackwell and the year compute became a policy variable
Open Science
Grok-1 by torrent: 314 billion parameters, Apache 2.0, no paper
Safety & Alignment
The EU AI Act passes: what a risk-tier law means for frontier models
Open Science
Up to 17 percent of ICLR 2024 reviews were touched by an LLM
Safety & Alignment
WMDP and unlearning: measuring hazardous knowledge without publishing it
Evaluation
Chatbot Arena's paper: 240,000 votes and the statistics behind the leaderboard
Evaluation
The pizza-topping needle: when Claude 3 noticed it was being tested
Interpretability
AtP*: attribution patching at industrial scale
Applied
1.58 bits: what BitNet promised and what it required
Open Science
The Stack v2 and the opt-out: building a code dataset people can leave
Applied
Air Canada is liable for its chatbot
Safety & Alignment
The Gemini image incident: over-correction as an alignment failure
Open Science
Gemma's terms of use: open weights with a leash
Applied
Compound AI systems: when the model stopped being the product
Evaluation
One million tokens at 99 percent recall: what Gemini 1.5's haystack chart proved
Open Science
OLMo: the first frontier-adjacent model with the whole recipe
Interpretability
Do Llamas think in English?
Safety & Alignment
Red teaming: silver bullet or security theatre?
Notes
Mamba was rejected from ICLR, and everyone had an opinion about peer review
Applied
Medusa and speculative decoding: getting more tokens per forward pass
Interpretability
Patchscopes: asking the model to decode its own activations
Foundations
AlphaGeometry: synthetic proofs and the neuro-symbolic loop
Notes
2,778 researchers, 50 percent by 2047: reading the AI Impacts survey
Evaluation
Task contamination: are models still few-shot learners?
Open Science
NYT v. OpenAI and the end of quiet training data
Open Science
LAION-5B taken offline: what the CSAM finding meant for open datasets
Applied
The $1 Chevy Tahoe: customer-service bots meet the open internet
Applied
LLM in a flash: Apple's blueprint for models that do not fit in RAM
Interpretability
When unsupervised knowledge discovery finds the wrong knowledge
Safety & Alignment
Weak-to-strong generalisation: can GPT-2 supervise GPT-4?
Foundations
Gemini 1.0 and the bet on native multimodality
Foundations
The month the transformer got competition: Mamba and Mixtral
Notes
NeurIPS 2023: word2vec's test of time and a field that outgrew its venue
Notes
The Gemini demo that was not live
Safety & Alignment
RLHF versus DPO: what changed when the policy became its own reward model
Evaluation
Gemini's 90 percent on MMLU and the fine print of CoT@32
Evaluation
GPQA and what a Google-proof benchmark is for
Applied
GPTs, the Assistants API and the custom-chatbot prompt-leak problem
Evaluation
The needle in a haystack test: one engineer's chart becomes an industry benchmark
Notes
Five days in November: what the OpenAI board crisis showed about governance
Evaluation
GAIA: questions humans get right 92 percent of the time and GPT-4 gets 15
Interpretability
The linear representation hypothesis, stated carefully
Safety & Alignment
The insider-trading demo: GPT-4 lies to its manager without being told to
Safety & Alignment
The executive order and the Bletchley Declaration: the week AI governance got specific
Notes
Managing AI risks: the consensus paper that was not a consensus
Open Science
The Data Provenance audit: 70 percent of datasets had no license listed
Foundations
Zephyr and distilled DPO: alignment from AI feedback in a weekend
Open Science
The Foundation Model Transparency Index: scoring the labs on what they will not say
Notes
The Techno-Optimist Manifesto and the enemies list
Evaluation
SWE-bench at launch: 2,294 GitHub issues and a 1.96 percent score
Interpretability
Llama has a map and a calendar: linear probes for space and time
Safety & Alignment
Ten examples to undo safety training: fine-tuning as the alignment hole
Interpretability
Towards monosemanticity: the week superposition stopped being a theory
Applied
DSPy and the idea that prompts should be compiled, not written
Interpretability
Representation engineering: interpretability from the top down
Evaluation
Six ways evaluation is harder than it looks: Anthropic's challenges post
Notes
A magnet link and a 7B model: Mistral's release as a cultural statement
Safety & Alignment
AI Safety Levels: reading Anthropic's first Responsible Scaling Policy
Foundations
The reversal curse: a model that knows A is B but not B is A
Interpretability
Sparse autoencoders find interpretable features: the independent replication
Open Science
The Authors Guild v. OpenAI: pirate libraries as the load-bearing allegation
Applied
Centaurs, cyborgs and the jagged frontier at BCG
Evaluation
Pretraining on the test set: the contamination joke that landed on a real problem
Interpretability
Othello-GPT's board is linear after all
Applied
Lucene is all you need: the case against a separate vector store
Open Science
Books3 is gone, and the datasets we trained on were never ours
Safety & Alignment
2,244 hackers, 8 models: what the DEF CON generative red team actually found
Evaluation
AgentBench: the first attempt to score a model as an agent
Notes
Dario on scaling: 'We still don't know why it works'
Applied
Llama 2 and the fine-tuning ecosystem it created
Safety & Alignment
Open problems with RLHF: the thirty-two author list of everything wrong with the method
Safety & Alignment
GCG and the adversarial suffix that transferred from Vicuna to GPT-4
Interpretability
Does circuit analysis scale? DeepMind tries it on Chinchilla
Evaluation
Is ChatGPT getting worse? The drift study and the problem of evaluating a moving target
Interpretability
Measuring faithfulness by breaking the chain
Applied
Lost in the middle: the U-shaped curve every RAG pipeline had to learn
Notes
ICML bans LLM-written text, and the field starts arguing about authorship
Safety & Alignment
Jailbroken: competing objectives and mismatched generalisation
Applied
The rise of the AI engineer, three years later
Applied
PagedAttention: the KV cache was the bottleneck all along
Foundations
Textbooks are all you need: the small model, synthetic data bet
Evaluation
Why LLaMA's MMLU score depended on who ran it
Notes
The Munk debate: Bengio and Tegmark versus LeCun and Mitchell
Applied
Mata v. Avianca: the brief with six invented cases
Applied
GPTQ, AWQ and the rules of post-training quantization
Evaluation
MT-Bench and the 80 percent agreement number
Interpretability
Inference-time intervention: nudging a few heads toward the truth
Open Science
Falcon goes Apache 2.0: the month a license clause moved a leaderboard
Notes
Hinton quits Google: when the pioneer changed his mind
Safety & Alignment
Let's Verify Step by Step: process supervision as a safety method, not just a math trick
Safety & Alignment
Twenty-two words: the extinction statement and what signing it committed anyone to
Evaluation
Model evaluation for extreme risks: the paper that made dangerous-capability evals a field
Foundations
How many epochs is too many? Scaling data-constrained language models
Applied
QLoRA: fine-tuning a 65B model on one GPU
Foundations
Tree of Thoughts and the return of search to language models
Interpretability
Models do not always say what they think: the first unfaithful CoT paper
Interpretability
Love minus hate: the activation addition trick
Interpretability
GPT-4 explains GPT-2's neurons, and the scores are humbling
Notes
ICLR in Kigali: the first major ML conference in Africa
Notes
"We have no moat": rereading the leaked Google memo
Evaluation
Chatbot Arena opens: Elo ratings for language models
Applied
Samsung's ChatGPT leak and the birth of the enterprise AI policy
Interpretability
How a model recalls a fact: three steps found by Geva and colleagues
Foundations
Are emergent abilities a mirage? The metric argument, and what survived it
Applied
The vector database gold rush, and what retrieval actually needed
Applied
Auto-GPT and BabyAGI: a post-mortem on the first agent hype cycle
Open Science
RedPajama: reproducing a training set from a paper's recipe
Open Science
Dolly 2.0 and the 15,000 answers written by employees
Open Science
Pythia and the case for publishing checkpoints, not just weights
Foundations
Pythia: the model suite built to be studied, not deployed
Notes
Pause, or shut it all down: the two March letters
Evaluation
Vicuna and the birth of GPT-4 as judge
Notes
Ilya on next-token prediction: notes on a podcast that aged well
Evaluation
The top 10 percent on the bar exam: reading the GPT-4 technical report's exam table
Interpretability
The tuned lens: reading a transformer's mind one layer at a time
Applied
Alpaca's $600 lesson: instruction tuning is cheap, evaluation is not
Safety & Alignment
The GPT-4 system card and the TaskRabbit story: the first dangerous-capability eval goes public
Notes
The GPT-4 technical report and the paper that told us nothing
Foundations
Predictable scaling: the one chart in the GPT-4 report that mattered
Open Science
The LLaMA leak, and the open-weights era nobody planned
Applied
The weekend LLaMA ran on a MacBook: llama.cpp and the 4-bit moment
Foundations
LLaMA broke Chinchilla on purpose: the case for overtraining small models
Safety & Alignment
Pretraining with human preferences: alignment before the model learns to misbehave
Evaluation
Theory of mind, spontaneously emerged and then trivially broken
Safety & Alignment
Sydney, DAN, and the week prompt injection became a discipline
Applied
55 percent faster: reading the first Copilot RCT carefully
Applied
Toolformer and ReAct: the two papers that taught models to call functions
Open Science
Getty v. Stability AI: the first big training-data lawsuit, and why it was about images
Interpretability
Grokking, reverse engineered: the Fourier circuit inside modular addition
Interpretability
Tracr and the case for ground-truth transformers
Safety & Alignment
The influence-operations report that predicted the year before it happened
Foundations
Cramming: what one GPU and one day can teach you about pretraining
Evaluation
GPT takes the bar exam: when professional exams became the benchmark
No posts match your search.