An open-source AI system, under the Open Source Initiative's Open Source AI Definition 1.0 released October 28, 2024, must give users the freedoms to use, study, modify, and share it — and that requires code, model parameters, and detailed data information, not just downloadable weights (institutional, OSI). Most models marketed as "open" today stop at the weights.
What does open-source AI actually require?
The Open Source AI Definition, version 1.0, states that the preferred form for making modifications to a machine-learning system must include three elements: data information, code, and parameters. Each element is defined in unusually concrete terms. Data information means enough detail about training data for a skilled person to build a substantially equivalent system, including provenance, scope, labeling procedures, and processing methods.
Code means the complete source code used to train and run the system, from data filtering to training arguments and inference. Parameters means the model weights, potentially including checkpoints from intermediate training stages. The definition is explicit that missing any element breaks the claim, because a user without all three cannot meaningfully exercise the four freedoms. The full text is on the OSI's Open Source AI Definition page.
This matters because the label "open source" carries legal and commercial weight. A company building a product on a genuinely open-source model can fork it, retrain it, audit it, and move it between providers. A company building on a weights-only release is renting access to a frozen artifact under someone else's contract.
What is the difference between open weights and open source?
Open weights means the model's parameters can be downloaded and run locally. That is genuinely useful — it enables local inference, fine-tuning, and offline deployment — but it is a subset of open source. The gap is everything the definition asks for beyond the weights: the training code, the data description, and the rights to redistribute modifications freely.
Licenses are the other half of the gap. The Apache License 2.0, one of the most widely used permissive licenses in software, grants reuse, reproduction, modification, and distribution, and adds an explicit patent grant from contributors. A model released under Apache 2.0 with its code and data information can plausibly satisfy an open-source standard. A model released under a custom contract with usage conditions cannot, however friendly the marketing reads.
The practical test is simple: could a competent third party rebuild a substantially equivalent model from what was published? If the answer is no, the release is open weights, not open source.
The distinction is not a judgment about usefulness. Open-weight models power a large share of local inference, research reproduction, and fine-tuning work, precisely because downloadable parameters lower the barrier to entry. The point is narrower and practical: open weights carry obligations the downloader does not control. If the publisher changes hosting terms, withdraws old versions, or adds usage conditions in a later release, downstream projects inherit every one of those moves. Open source exists as a legal category to prevent exactly that inheritance.
Which licenses do the major model makers actually use?
The big labs mostly publish their own license text rather than adopting a standard one. Meta's Llama 4 Community License Agreement, with an effective date of April 5, 2025 as stated in the agreement, grants a non-exclusive, worldwide, non-transferable, royalty-free right to use, reproduce, distribute, and modify the "Llama Materials" — with conditions attached.
| License | What it grants | Who uses it |
|---|---|---|
| Apache License 2.0 | Reuse, modification, distribution, explicit patent grant | Community and independent model projects |
| Llama 4 Community License | Broad reuse rights with named conditions in a custom contract | Meta's Llama 4 family |
| OSI Open Source AI Definition | An evaluation standard, not a license: code, parameters, data information | Used to judge whether a release is open source |
The full terms are on the Llama 4 license page. The naming itself is informative: a "Community License" is a bespoke agreement, and it is not on the OSI's list of approved licenses. That does not make it a bad deal for every user — it means the rights need to be read as a contract before anything is built on top.
What changed when the definition arrived?
Before October 2024, "open-source AI" had no agreed meaning, which made every comparison an argument about vibes. The definition settled the vocabulary. It distinguishes an AI model — which it describes as consisting of the model architecture and model parameters, together with the inference code needed to run it — from the broader system around it. That distinction lets a license be judged component by component rather than all-or-nothing.
The immediate effect was on disclosure. After the definition shipped, model cards and release notes started being read against a concrete checklist: is the architecture documented, are the weights complete, is the training code present, is the data information sufficient to rebuild. Labs that had described releases as "open source" began adding qualifiers — "open weights," "open research" — a small rhetorical retreat that itself signals the definition did its job.
The second-order effect is on procurement. Public bodies and large enterprises now cite the definition in vendor questionnaires, asking not "is your model open?" but "which of the three elements do you publish, and under what terms?" A question that used to be answered with marketing now has to be answered with a table, and tables are harder to blur. For an independent publication covering the space, the definition is the neutral yardstick that neither the open-weight labs nor the closed labs control.
Why do labs release weights but not data?
Training data is where most open-weight releases fail the open-source test. Labs cite three recurring reasons: proprietary data pipelines are a competitive moat; much training data is licensed or scraped under terms that forbid republication; and disclosing datasets creates legal exposure. The result is that the data-information element of the definition is almost never satisfied by frontier labs.
There is also a commercial logic to "open-washing" — borrowing the reputation of open source while keeping the moat. A weights release buys community adoption, security research, and ecosystem lock-in at relatively low cost. Full openness would hand competitors the entire recipe. Reporting on model releases should therefore treat the word "open" as a claim to be checked, not a fact to be repeated.
How should a team evaluate an "open" model claim?
A structured check takes under an hour and prevents expensive surprises:
- Read the license text itself, not the announcement — conditions usually live in the early sections.
- Confirm the training and fine-tuning code actually ships, not only inference code.
- Look for data information: provenance, scope, filtering methodology — the element the definition demands.
- Check what parameters are published, including whether intermediate checkpoints are available.
- Map license conditions against your product — user-count caps and acceptable-use clauses decide the answer.
The open-source model landscape is best read as a spectrum: permissively licensed models with full code at one end, weights-only releases under custom contracts at the other. The definition's value is that it turns a marketing argument into a checklist — and most current releases fail the checklist on the same item, data.

