Skip to content

The open-weight year: sovereign AI became a procurement option

Four frontier-adjacent models shipped with open weights between February and April 2026, under licences you can actually use. What that means for a NZ organisation with health, iwi or public-sector data, and the three questions that decide whether to run one yourself.

Between February and April 2026, Alibaba, Mistral, Google and DeepSeek each shipped a model with downloadable weights under a licence that lets you run it commercially, on your own hardware. On the independent index we use, the gap between the best open model and the best closed one is now six points, down from about thirteen a year earlier. For a NZ organisation holding health records, iwi data or public-sector information, that changes the decision. Sovereign AI used to mean accepting a much weaker model in exchange for control. That trade is now a lot shorter.

What you need to know

  • Four capable models arrived with open weights in three months. Qwen3.5 (16 February), Mistral Small 4 (16 March), Gemma 4 (2 April) and DeepSeek V4 (24 April), all under Apache 2.0 or MIT licences.
  • The licences are the real news. Gemma 4 dropped Google's custom terms for Apache 2.0, the licence Qwen and Mistral already ship under.
  • The gap to the frontier is small and measured. Artificial Analysis put the best open-weight models six points behind GPT-5.5 on its Intelligence Index at the end of April 2026.
  • Open weights move the cost, they do not remove it. You carry the hardware, the ops, the security evidence and the evaluation work a hosted vendor used to carry for you.
  • A hosted frontier model still wins on the hardest tasks. The closed models are ahead at the top of the range, and you pay for that in data leaving your control.

6 points

Gap from the best open-weight model to GPT-5.5 on the Artificial Analysis Intelligence Index, down from about 13 a year earlier

Source: Artificial Analysis, April 2026

80.6%

DeepSeek V4-Pro on SWE-bench Verified, on DeepSeek's own model card, released under MIT with 1.6T total and 49B active parameters

Source: DeepSeek, April 2026

4x H100

Minimum hardware Mistral states for running Mistral Small 4, a 119B-parameter Apache 2.0 model with 6B active

Source: Mistral AI, March 2026

What actually shipped, and under what terms

The models matter less than the licences, so start there. Qwen3.5-397B-A17B landed on 16 February 2026 under Apache 2.0: 397 billion parameters in total, 17 billion active per token. Mistral Small 4 followed on 16 March, also Apache 2.0, at 119 billion total and 6 billion active, one model covering reasoning, vision and coding. Google released Gemma 4 on 2 April in five sizes from about 2 billion effective parameters to a 31 billion dense model, covering more than 140 languages, and moved the family from its own Gemma Terms of Use to Apache 2.0. DeepSeek published V4 as a preview on 24 April under MIT: a 1.6 trillion parameter V4-Pro with 49 billion active, a 284 billion V4-Flash with 13 billion active, and a 1 million token context on both.

Apache 2.0 and MIT are the licences your legal team already accepts for the rest of your software stack: no user cap, no acceptable-use policy the publisher can rewrite. Gemma's move is the tell. Google gave up its custom terms because the Chinese labs had made permissive licensing the price of being taken seriously.

On quality, the honest position is "close, and measured". Artificial Analysis had Kimi K2.6 and MiMo V2.5 Pro tied as the leading open-weight models at 54 on 30 April, DeepSeek V4-Pro at 52, against GPT-5.5 at 60. A year earlier the best open model sat about 13 points behind the best closed one. DeepSeek's own card reports 80.6% on SWE-bench Verified for V4-Pro. That is the publisher's number: a claim to test, not a fact to plan on.

What open weights cost you

A hosted API hides three costs that come back to you the moment you self-host. The first is hardware and operations. Mistral states a minimum of four H100s to run Small 4. DeepSeek V4-Pro at 1.6 trillion parameters is a GPU cluster, not a server. Gemma 4's 31B dense model is the smallest of the four families' top models, which is why it is the realistic starting point for most NZ organisations rather than the model with the best headline score. Whatever you pick, someone has to patch, monitor and upgrade it, and keep it running when a card fails on a Sunday.

The second cost is evaluation. A public benchmark tells you nothing about your triage notes, your grant applications, your te reo content or the way your board wants a risk summary written. With open weights that judgement is yours. You need a small evaluation set built from your real work, scored the same way every time, before you pick a model and again before you swap one.

The third cost is the security evidence. Running the model inside your boundary answers the sovereignty question, but the audit will still ask who can access the weights, where prompts are logged, how environments are separated, and what happens when an output is wrong. The paperwork is the same whether the model is open or closed; with open weights it is your paperwork. The Trust Centre lists what we document for a deployment.

The three questions that decide it

What class of data is in the prompt? If the answer is public information or your own marketing copy, a hosted model with a no-training clause is fine and cheaper. If the answer is health records, whānau data, Māori data with iwi obligations attached, or anything a public-sector procurement team will ask you to locate, open weights are now a real option rather than a compromise. Residency is where the data sits. Sovereignty is who can compel it. An open model on hardware you own in Aotearoa is the one configuration where the second answer is "nobody offshore".

Can you run your own evals? If you cannot describe ten real tasks from your organisation and say what a good answer looks like, you are not ready to pick a model, open or closed. Build that set first. It is yours whichever model wins next quarter.

Who carries the operations? Three real options. Your own team runs it, which suits organisations that already run infrastructure. A partner runs it inside your boundary, which is what RIVER's Sovereign AI deployment is: hardware in your building, your chosen models deployed and tuned, operated by you or by RIVER. Or the data class does not justify it and you take a hosted model with NZ hosting and a no-training configuration. All three are defensible. Picking one by default is not.

The maths of a language model does not change when you download the weights. What changes is that the failure modes become yours to find. I have watched a model that scores brilliantly on a public benchmark fall apart on a clinical summary because the domain was thin in its training data. You will not learn that from a leaderboard. You learn it by running your own work through it, measuring, and being willing to say the popular model is wrong for you.

Dr Vincent RussellMachine Learning (AI) Engineer

What to actually do

Classify the data before you shop for a model. Sort your candidate use cases by what is in the prompt, not by how exciting the demo is. The data class picks the deployment model, and that narrows the models to a handful.

Build the evaluation set this quarter. Ten to fifty real tasks, a plain description of a good answer for each, scored the same way every time. Run Gemma 4 31B against it as the small baseline and a hosted frontier model as the ceiling. The distance between those two numbers on your own work is the only gap that matters.

Decide who carries the operations, in writing. Name the team or the partner, the hardware, the upgrade cadence and the evidence pack the audit will ask for. If nobody can be named, you have not chosen sovereign AI yet, you have chosen a download.