Between April and June 2026, three of the largest open-weight families shipped releases on Hugging Face: DeepSeek-V4-Flash, at 284 billion total parameters with a one-million-token context, under an MIT license (documented); Z.ai's GLM-5.2, at 753 billion, MIT-licensed; and Moonshot AI's Kimi K2.7 Code, at one trillion parameters, under a Modified MIT license. All figures come from the makers' model cards.
What exactly shipped between April and June 2026?
DeepSeek's V4-Flash arrived first as a preview, listed as DeepSeek-V4-Flash with 284B parameters (13B activated), supporting a context length of one million tokens under an MIT license. The card documents a Hybrid Attention Architecture combining CSA and HCA for long-context efficiency, Manifold-Constrained Hyper-Connections for stable signal propagation, the Muon optimizer during training, and pre-training on more than 32 trillion tokens. A two-stage post-training pipeline used GRPO reinforcement learning and on-policy distillation.
Z.ai's GLM-5.2 followed in June, positioned as its latest flagship for long-horizon tasks. The card claims a substantial capability leap over GLM-5.1 and, for the first time in the family, a solid one-million-token context — plus a proposed technique called IndexShare that reuses the same indexer across every four sparse attention layers, reducing per-token FLOPs by 2.9 times at that context length (documented). Moonshot's Kimi K2.7 Code, also June, is a coding-focused agentic model built on Kimi K2.6, cutting thinking-token usage by roughly 30 percent compared with its predecessor, per its card (company-claimed).
How do the three releases compare on paper?
Read as a row of specifications — the only honest way to read them before independent benchmarking — the three cards line up like this:
| Model | Total / active params | License | Context | Card-documented emphasis |
|---|---|---|---|---|
| DeepSeek-V4-Flash | 284B / 13B | MIT | 1M tokens | Hybrid attention architecture; 32T+ pretraining tokens; KV cache compression |
| GLM-5.2 | 753B / not stated | MIT | 1M tokens | Long-horizon tasks; IndexShare cuts per-token FLOPs 2.9x at 1M context |
| Kimi K2.7 Code | 1T / 32B | Modified MIT | 256K tokens | Agentic coding; ~30% fewer thinking tokens vs K2.6; native INT4 |
How do the licenses differ — and why does it matter?
Two of the three are plain MIT. DeepSeek's card lists an MIT license, and Z.ai describes GLM-5.2 as an MIT open-source license — no regional limits, technical access without borders (company language). For downstream users this is the permissive end of the spectrum: commercial use, modification and redistribution with attribution, without negotiated agreements.
Moonshot's choice is the instructive exception. The Kimi K2.7 Code card states that both the code repository and the model weights are released under the Modified MIT License — a custom variant rather than a standard open-source license. Modified MIT terms in this family have historically attached conditions to certain commercial service deployments, so the practical lesson generalizes: open-weight does not automatically mean unconditionally licensed, and the license file, not the marketing label, is the contract.
Why does mixture-of-experts dominate these releases?
All three cards share one architecture decision: sparse mixture-of-experts, where only a fraction of the network fires per token. DeepSeek activates 13B of 284B parameters; Kimi activates 32B of roughly one trillion. The economics are direct — memory cost scales with total parameters, but compute cost scales with the activated slice, so a huge model can run at the inference cost of a small one.
That trade is exactly what makes tera-parameter open weights practical. A dense one-trillion-parameter model would be unusable for almost everyone outside a hyperscaler; a sparse one with 32B active runs on serious but obtainable hardware. The pattern also explains the context-length race — long horizons multiply the value of cheap per-token compute, which is where all three makers have aimed their engineering, from hybrid attention to indexer sharing.
The engineering details on the cards are the evidence that long context is now a memory problem first. DeepSeek's V4.1-Flash follow-up paper, listed by the organization, is titled around pushing the limits of KV cache compression; GLM-5.2's IndexShare explicitly attacks per-token compute at the one-million-token mark; Kimi ships native INT4 quantization so a trillion-parameter weight set fits in roughly 595 gigabytes rather than two terabytes. Different tricks, same bottleneck — the cache and the memory bus, not the arithmetic.
Sparsity also quietly changes who can serve these models commercially. Because activated parameters are small, batched serving on a well-provisioned cluster yields usable throughput at prices competitive with closed APIs — which is precisely the business several inference providers have built on top of open-weight releases, and why the open-weight tier disciplines closed-model pricing even for customers who never download a weight file.
What does it take to actually run one?
The Kimi card is unusually concrete about deployment: the model uses native INT4 quantization, ships as roughly 595 GB of shards, and is deployable via vLLM, SGLang or KTransformers, with a pinned transformers version range (documented). That is a realistic picture of the floor: even a heavily quantized trillion-parameter model is a multi-GPU server commitment, not a workstation toy. Smaller sparse models are correspondingly kinder — but the open-weight tier's center of gravity has clearly moved upmarket in memory terms even as compute-per-token falls.
What do open weights still not give you?
Weights are the artifact, not the institution. None of the cards ship the training data, the full alignment pipeline, or any service guarantee — the user carries hosting, safety filtering and update cadence. Benchmark numbers on the cards are maker-reported until an independent evaluator publishes on the same tests; this analysis deliberately reports card claims as card claims. What the spring 2026 wave does establish is narrower and still significant: at every tier from 284B to 1T, permissively licensed frontier-adjacent weights with million-token context are now a downloadable commodity.

