Can a coding agent clean-room an LGPL library into MIT? The chardet fight
chardet 7.0.0 was rewritten with Claude and relicensed from LGPL to MIT. The original author says the maintainer had no right to do that. An explainer on clean-room doctrine, what a plagiarism score can and cannot show, and where a model trained on the original code fits into the argument.
What happened
chardet is the Python character encoding detector that sits in the dependency tree of a very large fraction of Python projects. Mark Pilgrim wrote it in 2006 as a port of Mozilla's universal charset detection code and released it under the LGPL. Dan Blanchard has maintained it since 2012. This week Blanchard shipped version 7.0.0, described as a complete rewrite, built with Claude, under the MIT licence.
On March 4 Pilgrim, who has been out of public life since 2011, opened an issue titled 'No right to relicense this project'. His argument fits in a paragraph. LGPL code that is modified must stay LGPL. Calling the new version a complete rewrite does not matter, because the people who wrote it had ample exposure to the original, so it is not a clean-room implementation. And putting a code generator in the loop does not create any rights that did not exist before. He asked for the licence to be reverted.
What clean-room doctrine actually requires
Clean-room reimplementation describes a process. The classic form has two teams separated by a wall. One team reads the original and writes a specification containing only ideas, interfaces and behaviour. The other team, which has never seen the original, implements from the specification. The wall is evidence. If the second team could not have copied, then whatever they produced is independent, and independent work is not a derivative work no matter how similar it looks.
Blanchard concedes in his reply that no such wall existed. He has maintained the codebase for over a decade and he designed the new one. His argument is that the wall is a means to an end, and the end is a work that does not derive from the original, which he says he can demonstrate by measurement instead of by process. That is the crux of the dispute. Pilgrim is arguing about process. Blanchard is arguing about outcome.
What the plagiarism scores show and do not show
Blanchard ran JPlag 6.3.0 across every major release tag. JPlag tokenises Python into syntactic elements, discards names, comments and formatting, and looks for long matching token sequences, so renaming variables does not fool it. His table is the most useful thing in the thread. Consecutive releases in the old lineage score between 80 and 91 percent average similarity. Version 6.0.0, which reworked the single-byte detectors, scores only 3.3 percent average similarity to 5.2.0 but 80 percent maximum, because entire files were carried forward. Version 7.0.0 against 6.0.0 scores 0.04 percent average and 1.29 percent maximum. Against the original 1.1 release it scores 0.64 percent maximum.
Read carefully, this shows that no file in 7.0.0 structurally resembles any file in any prior release, and Blanchard says the matched tokens are things like argparse boilerplate and import blocks. That is strong evidence against literal or near-literal copying. It is weaker evidence on the question Pilgrim is actually raising, because a derivative work under copyright can share structure, sequence and organisation without sharing token sequences, and a token-level tool measures neither of those.
Several commenters raised a further point that the JPlag numbers cannot address at all. One of the planning documents committed during the rewrite instructs the agent to fetch the 6.0.0 charsets.py file from the repository and use it as the authoritative reference for which encodings belong to which era. Blanchard's process description says he started in an empty repository with no access to the old source tree. Both can be true, since the file was fetched by URL rather than from a local checkout, but it means the separation was not total, and the defenders of the rewrite have mostly argued that the file in question is a table of mappings rather than expressive code.
What the model changes
Here is where the argument becomes new. Blanchard says he instructed Claude not to base anything on LGPL or GPL code and then reviewed, tested and iterated on every part of the result. He did not write the code by hand. Critics in the thread make two objections that a traditional clean room never had to answer.
The first is about training. chardet is old, widely mirrored and almost certainly in the training data of any large code model. Telling a model that has already read the original not to copy it is, as one commenter put it, asking the person who saw the code multiple times nicely not to reproduce what they memorised. In the classic clean room the second team's ignorance is what makes the process work. A model cannot be ignorant of chardet. Whether that matters legally is not settled, and the thread contains confident assertions in both directions from people who are not lawyers.
The second objection is about authorship. If the code was generated rather than written, one commenter argues, then under current United States copyright guidance Blanchard contributed nothing copyrightable, and the question of what licence he can attach to it gets stranger rather than simpler. Simon Willison, who has followed this closely, says he leans towards the rewrite being legitimate while finding the arguments on both sides credible. Richard Fontana, who co-authored GPLv3, is reported as seeing no clear basis for a licence violation while saying he would not have chosen MIT. Informed opinion is divided, and nobody in the thread has the authority to decide.
Why he wanted MIT, and what we take from it
Blanchard's stated motive is prosaic. There was talk a decade ago of moving chardet into the Python standard library, which requires a permissive licence, and he has long wanted more contributors than an LGPL project attracts. Nothing in the thread suggests a commercial angle, which leaves the licence question open while making the dispute a case about doctrine rather than about bad faith.
What we take from it is that the coding agent has made a previously expensive thing cheap, and the doctrine that governed the expensive version assumed it would stay expensive. A clean-room rewrite used to need two teams and months. It now needs one maintainer and a weekend, with a design document and a similarity report as the evidence of independence. The measurement half of that is genuinely good practice and we would like to see it become standard for any AI-assisted rewrite. The process half, the wall, is the part that has quietly disappeared, and nobody has yet said what replaces it. Until a court or a standards body does, we would treat a relicensed rewrite of code the model has surely seen as an open question rather than a settled one, and we would not depend on the new licence for anything that matters.
Sources
From the foundation