Open weights are not open source: a view from a country with no TDM exception
In early August, Meta released the weights of Muse Glimmer 30B under a genuine Apache 2.0 license, dropping the "community license" format that the Open Source Initiative had rejected for the Llama family. In the same repository sits a separate file: USAGE_POLICY.md. Stefano Maffulli, who led the OSI through the definition of open source AI, asked in public what that document is, legally. "Lawyers?", he wrote. This piece is, among other things, an answer. It is also a report from a part of the world that runs on the open source commons and rarely sits at the tables where its rules get written.
Start with the distinction that orders everything, stated carefully: publishing a model's weights leaves the user roughly where a proprietary binary leaves them. The analogy is functional, not ontological. Source and object code are the same work in two formats; weights are not the training data in another format but a lossy statistical transformation of it. What carries over is the user's position, not the artifact's nature. The classic definition of source code, repeated in every open source license, is the preferred form of the work for making modifications. For a model, that form is not the weights. It is the training data (or auditable information about it), the training code, and the full recipe: hyperparameters, filtering, the decisions that shaped it. With weights you can run the model and fine-tune it; you cannot reproduce it, audit it in depth, or study why it does what it does. That is exactly the position of a proprietary software user holding a free executable. Useful. Not freedom.
The Open Source AI Definition (OSAID 1.0, October 2024) translates the four classic freedoms into three components that must be available under open terms: data information detailed enough for a skilled person to build a substantially equivalent system; the complete code for training, data processing and inference; and the parameters. Two details in the text deserve a lawyer's attention. The definition expressly allows these components to be licensed under reciprocal terms, so copyleft survives the jump to AI. And it states that "open source weights" means weights shipped together with the data information and the code that produced them. Weights alone cannot carry the name, whatever their license.
The definition is contested: the Free Software Foundation and the Software Freedom Conservancy objected to the compromise of describing training data rather than releasing it, and the debate is healthy. But a public yardstick is a yardstick, and the validation exercise that accompanied it named names. Llama 2, Grok, Phi-2 and Mixtral do not pass, for missing components or incompatible legal terms. Pythia, OLMo, Amber, CrystalCoder and T5 pass. BLOOM, Starcoder2 and Falcon would pass with their legal terms corrected. Not certifications, the OSI is careful to say. But the map is drawn: compliant models exist, and their relative obscurity next to the non-compliant ones shows that the label is contested by marketing, not by precision. The OSI itself lists fighting openwashing as one of the three reasons the definition exists.
Now the frontier moved. Muse Glimmer's license is clean: no user caps, no field-of-use restrictions. The restrictions migrated to the usage policy. Legally, everything turns on incorporation. Nothing in Apache 2.0 conditions the grant on that policy, and a side document without incorporation language is a statement of wishes: no assent, no consideration, incapable of cutting back rights already granted. If it were incorporated as a condition, the license would stop being Apache 2.0 in substance, because field-of-use and user restrictions fail criteria 5 and 6 of the Open Source Definition. That is the Llama defect in new packaging.
The interesting wrinkle is what happens in the practical middle case, where a click-through at download creates real assent. Section 2 of Apache 2.0 grants every downstream recipient a license directly from the contributors. Whoever receives the weights from the first downloader does not derive their rights from the platform contract that person accepted; they take them from the license itself. The policy binds the clicker and evaporates at the first redistribution. It cannot travel with the weights. The principle is not an Apache quirk: the Open Weight Definition codifies it as a criterion of openness, requiring that rights reach everyone the weights are redistributed to without executing an additional license. A restriction regime that demands individual adhesion at every download is, by definition, the opposite of an open license.
So what is the usage policy for? My hypothesis is functional: its audience is regulators, not licensees. In the AI Act era, a published acceptable-use policy is exhibit A of responsible release, the most visible and least costly piece of a provider's compliance file. Compliance signaling, not intellectual property. There is nothing wrong with signaling; the problem starts when regulators or courts attribute to it an obligational force its structure cannot support. Meanwhile, the model still ships without training data or pipeline. The license improved; the components did not. Under OSAID it remains open weights. A license is a necessary condition of openness. It was never a sufficient one.
The definitional field itself is now crowded, and the crowd is informative. OSAID draws a binary line, which is what you need to certify and to exempt. The Linux Foundation's Model Openness Framework sorts releases into tiers, which is what you need to compare. The Columbia Convening framework, published in Communications of the ACM, describes openness as a gradient across the whole stack and deliberately refuses to fix requirements, which is what you need to analyze. Three philosophies, three functions. The analytical map and the legal border are not the same thing, and confusing them is the hurried regulator's characteristic mistake.
Three things look different from Santiago. First, the layer underneath the openness debate. Most of Latin America has no text-and-data-mining exception and no fair use doctrine; Chile has neither. Training a model on local data sits in a legal gray zone that no "open" framework resolves. The activity does not stop; it relocates, or it gets governed by contract, with all the asymmetry that implies. The absence of law is also industrial policy, just somebody else's.
Second, where courts are silent, licensing discipline matters more, not less. There is no Latin American case law on software derivative works. None. The contract is the law between the parties, and the licensor's public, early and consistent interpretation of ambiguous terms can create more certainty than decades of judicial silence. Linus Torvalds' three-line note in the kernel's COPYING file, declaring that programs using system calls are not derivative works, is the canonical example: no legal force, and an entire industry built on it. The usage policy is the mirror image: a unilateral document that expands nothing and pretends to restrict what it cannot.
Third, the region's legislators are importing the AI Act's open source exemption at translation speed, without the interpretive apparatus behind it. If the legal category of "open" ends up covering licenses like Llama's, the exemption subsidizes openwashing; if it demands OSAID, almost no commercial model qualifies. The definition is the policy. Whoever drafts the definition decides. My recommendation for regional drafters is simple to state and hard to execute: define by reference to standards maintained by specialized third parties, knowing that choosing which one is itself a policy decision; exempt by risk rather than by label, because safety depends on deployment context and not on the isolated model; and legislate no obligation that no agency can enforce.
Update, September 19. On September 16, the Ninth Circuit affirmed the dismissal of the open source programmers' suit against GitHub, Microsoft and OpenAI over Copilot and Codex: the largest lawsuit ever brought over training on open source code, with more than $9 billion claimed. The claim rested on section 1202(b) of the DMCA, which bars removing copyright management information (author, notice, license terms) from an existing work. The court found Copilot did not remove that information: it generated new code that never carried it. No copy, no removal.
Two things most coverage missed. Plaintiffs had also argued the violation happened earlier, when license notices were stripped from the code before it went into training. That theory, the only one that actually touches training, was never decided on the merits. It was forfeited: at the first hearing, counsel was asked point blank whether training on stripped code breached the licenses' attribution requirement and answered "Perhaps it doesn't." The panel held the theory abandoned, not wrong. It stays fully open, for a better-pled case. And the breach-of-contract claims, resting on the license terms themselves (GPL, MIT, Apache), are still alive in the district court. That is a different doctrine, license as contract rather than license as copyright condition, the same architecture this piece already leans on for the usage-policy argument above. The most-cited AI ruling of the year has not yet reached the question everyone assumes it settled.
Update, September 21. Creative Commons just admitted the limit of its own invention. In a reflection on these years, the organization that best understands open licensing in the world says two things: that copyright alone cannot give it the balance it needs, and that AI labs treat any voluntary condition, a "yes, if..." or a "no, unless...", as a flat no, without even negotiating it. That is why, through CC Signals, they built a preference layer explicitly separate from the license, with no pretense of binding anyone, and why they are now exploring, still experimentally and with nothing built yet, a legal tool distinct from copyright to condition bulk access to data collections. The diagnosis matches this piece's own: a signal nobody is obligated to follow looks more like a wish than a rule.
This is a compressed English version of the seventh entry of Código Abierto, a ten-part Spanish series on the law of open source, written from Latin America. The full entry, with primary sources linked, is in Spanish. A longer academic treatment is in preparation.