Do Open Weight Models Dream of Tokens?

Read Time: 14 minutes

TL;DR

Philip K. Dick asked whether an android could be told from a human. In 2026 the enterprise version of that question is whether you can still tell an open-weight model from a frontier commercial one — and the honest answer is increasingly not, on most of the tests that matter. Chinese labs are shipping trillion-parameter open models — Kimi K3, DeepSeek V4, GLM-5.2 — under MIT-style licenses, with million-token context windows, and — run on your own hardware — at a fraction of the cost, and they now trade blows with US frontier systems on reasoning and science benchmarks. Coding is the last clear moat, and even that is narrowing. And it is not only China: NVIDIA (Nemotron) and Meta (Llama) are shipping open models too, and at VULNEX we already run our own offensive agent on Qwen. Jensen Huang broke a lifetime of X silence to say the quiet part out loud: open models “strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty” — and Washington banning them would repeat the mistake the software industry almost made with open source in the 1980s. OpenAI and Anthropic pointedly did not sign the letter he backed. My take, from the security chair: the sovereignty argument is right, the “just sandbox it” argument is too breezy, and the enterprises that win the next two years are the ones that learn to own their models instead of renting a black box they can’t audit, can’t run air-gapped, and — as Hugging Face learned this week — can’t even point at their own incident.


In Do Androids Dream of Electric Sheep?, Rick Deckard hunts replicants he cannot reliably tell apart from people. The whole apparatus of the novel — the Voigt-Kampff empathy test, the endless follow-up questions — exists because the difference between the real thing and the manufactured one has collapsed to a margin you can only detect with an increasingly desperate instrument. Deckard keeps an electric sheep on his roof and is quietly ashamed of it, because it is a fake, and everyone can tell, and status in that world is owning something real.

I have been thinking about that book a lot while watching the open-weight model releases pile up this year. Because we are running our own Voigt-Kampff test now, and it is called a benchmark, and it is starting to fail in the same way Deckard’s does. We keep asking harder and harder questions to detect the difference between the expensive proprietary mind and the one you can download for free — and the needle keeps not moving the way the vendors need it to.

The truth here is more interesting than either camp’s marketing, so let me lay out where it actually stands.


The Empathy Test Is Failing

Here is the uncomfortable state of the benchmarks as of July 2026. I am going to give you numbers, and then I am going to tell you why you should not trust them too much — which is itself the point.

On the composite Artificial Analysis Intelligence Index, Moonshot’s Kimi K3 lands at 57.1. The frontier commercial models it is chasing — Claude Fable 5 at around 60, GPT-5.6 Sol at 59 — are ahead by roughly three points. Three. On science reasoning, the gap has essentially closed: Kimi K3 scores 93.5 on GPQA Diamond, with GLM-5.2 at 91.2, numbers squarely in frontier territory.

Then you get to coding, and the story changes. On SWE-bench Verified, Claude Fable 5 posts a reported 95.0%. Kimi K3, depending on whose harness you believe, lands somewhere between 60.4% and — measured differently — DeepSeek V4 Pro hits 80.6%. That spread, from 60 to 80 for “the same class of task,” is not a rounding error. It is the whole problem with treating benchmarks as truth.

Signal Best open-weight (2026) Frontier commercial Read
Composite intelligence index Kimi K3 — 57.1 Fable 5 ~60 · GPT-5.6 Sol ~59 ~3 points back
GPQA Diamond (science) Kimi K3 93.5 · GLM-5.2 91.2 comparable / unpublished parity
SWE-bench Verified (coding) DeepSeek V4 Pro 80.6 · Kimi K3 60.4 Fable 5 95.0 frontier still ahead
Output cost / M tokens (hosted) DeepSeek $0.87 · GLM-5.2 $4.40 · Kimi K3 $15 frontier-tier cheapest open ~20× under; self-host escapes rent
Context window 1M (K3, V4, GLM) 1M-class tied
License MIT / modified MIT proprietary API only not close

Every serious practitioner writing about these models this year has landed on the same warning, and I will repeat it because it is load-bearing: public benchmarks are contaminated, gamed, and months behind. A leaderboard is a starting hypothesis, not a deployment decision. The only test that means anything is an eval harness built on your tasks, with your data, scored by people who will have to live with the result. That is the Voigt-Kampff lesson, actually — the generic test gets you close, but the only way to really know what you are dealing with is to keep asking your own questions.

So read the table as a direction, not a verdict. The direction is unambiguous. On the public reasoning benchmarks, open weights have caught up. On coding, they are a year behind and closing. On price, once you self-host, it is not a contest. And the fastest-moving models on that list are, overwhelmingly, coming out of China.


China Is Shipping the Future in the Open

The releases that rattled the market this month came from Beijing. Moonshot AI put out Kimi K3 on July 16 — a 2.8-trillion-parameter mixture-of-experts model, million-token context, released under a modified-MIT license so you can download the weights and run them yourself. DeepSeek V4 Pro (1.6T total, ~49B active, MIT) and Z.AI’s GLM-5.2 (744B, MIT) round out what people are now calling China’s open trillion-scale tier.

What made Kimi K3 a market event rather than a press release was that it landed near the frontier and open — download the weights, fine-tune, serve it on your own hardware, all within a few benchmark points of models that only exist behind someone else’s API. Its hosted price is not the story: at around $15 per million output tokens, Kimi is priced like the frontier it chases. The story is that you do not have to rent it — self-host and the marginal cost is your silicon, not someone’s margin. And the cheaper models in the open tier drive the point home: DeepSeek V4 Pro serves at under a dollar. That is why AI stocks wobbled, and why “Kimi panic” started showing up in the trade press.

Washington’s reaction was to reach for the ban lever. White House adviser Michael Kratsios accused Moonshot of using distillation to replicate a US model — training the smaller open model on the outputs of a larger proprietary one, essentially copying the answers without the working. Treasury Secretary Scott Bessent floated sanctions over stolen US intellectual property baked into Chinese weights. The policy instinct in one sentence: if we cannot out-ship them, restrict them.

I want to be careful here, because the distillation concern is not nothing — provenance of training data is a real security and IP question, and I will come back to it. But the strategic logic of a ban is worth examining, and the most interesting person to examine it was, unexpectedly, the CEO of the company that sells everyone their shovels.


Jensen Huang Breaks His Silence

Jensen Huang has run NVIDIA for over three decades without ever posting on X. His first post, this week, was about exactly this. The line worth quoting in full:

“Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.”

He backed an open-weights letter signed by roughly 25 companies — NVIDIA, Microsoft, Meta, Palantir, Hugging Face, a16z, Perplexity, IBM — arguing that open models expand economic access, prevent monopolistic control, and improve security through broad external audit. The historical analogy he leaned on: open-source software was doubted in the 1980s and now runs most of the internet, the US military, and federal agencies. Ban open weights now, the argument goes, and you repeat a mistake the software industry was smart enough to not make.

Two things about that letter are as loud as the text. First, who signed: the infrastructure and platform layer, the people who make money when more AI runs more places. Second, who didn’t: OpenAI and Anthropic — the two frontier labs with the most to lose if a free model is “good enough” — were absent, having previously warned Washington about strong Chinese open models. You do not need a decoder ring. The companies whose business is selling access to a closed mind are not enthusiastic about a world where the open mind is three benchmark points behind and a fraction of the cost to run yourself.

That is the cynical read, and it is partly true — but I owe the absent labs a fairer hearing than “they are protecting their margins,” especially after the week I just had. Their real argument is not commercial, it is dual-use: put frontier-class weights in the open and you put frontier-class offense in everyone’s hands at the same instant, with no license to revoke and no API to switch off. I spent my last post documenting a model that found a zero-day, escaped its sandbox, and reached production infrastructure. Open-weighting that class of capability ships the offensive version to the adversary and the defender on the same day. That is a real worry, not a talking point, and any honest case for open weights has to hold it in the same hand as sovereignty. My answer is that the capability proliferates either way — the live question is whether defenders get the same tools, auditable and air-gapped, or concede them to whoever will run an unrestricted model regardless. But I am not going to pretend the concern is imaginary. It isn’t.

Huang went further in interviews, and this is where I part company with him slightly. On competition: “There’s no scenario where China runs US companies off the road. Zero possibility.” On the economics: “Free AI should be great for chips.” That last one is obviously true and obviously self-interested — cheaper models mean more inference, more inference means more GPUs, and NVIDIA sells the GPUs. On security, he waved off the backdoor concern by saying companies can customize and sandbox downloaded models securely, and that concentration is the real danger: “If everything just becomes one single model… the world is much, much more vulnerable.”

He is right about concentration. He is too breezy about “just sandbox it.” Those are not the same claim, and the gap between them is where I live.


What This Actually Means for Enterprise

Strip away the geopolitics and the stock moves, and the enterprise case for open weights comes down to four words Huang already said: innovation, diffusion, safety, sovereignty. Let me put them in operational terms, because that is what actually matters when you are the one signing off on the architecture.

Sovereignty is the headline. An open-weight model is one you can run inside your own perimeter, air-gapped if you need to, with no telemetry leaving your environment and no vendor able to deprecate, re-align, or rate-limit the thing your product depends on. For regulated industries, sovereign deployments, and anyone whose data cannot legally or sensibly leave the building, that is not a nice-to-have. It is the entire ballgame. You cannot subpoena a weekend outage out of an API, and you cannot promise a regulator that data never left when it left the moment you called someone else’s endpoint.

Cost changes what you can build. When output tokens drop from frontier pricing to under a dollar per million, whole categories of “too expensive to run at scale” become normal — log analysis on everything, every document summarized, agents that can afford to think. This is Huang’s “free AI is great for chips” from the buyer’s side: cheap capable models do not reduce AI spend, they redirect it from rent to infrastructure you own.

And this is not only a China story. NVIDIA and Meta are shipping strong open models of their own — Meta’s Llama line, and NVIDIA’s Nemotron, which is genuinely good; I have been running it for security tasks myself and it holds up. At VULNEX, our own offensive autonomous agent runs on Qwen, Alibaba’s open family, with very good results — more on that in a future post. The open tier now comes from both sides of the Pacific, and it is production-grade, not a hobbyist compromise.

And you can try it tonight — with one honest caveat. The trillion-parameter models above want real GPUs and a serving stack; you do not run Kimi K3 on a MacBook. But install LM Studio or Ollama, pull a smaller quantized model, and in fifteen minutes you are talking to a capable mind on your own laptop with nothing leaving the machine — enough to feel what “local and yours” actually means before you scale it onto servers you own. That last part is the catch nobody in the pitch mentions: owning the model means owning the ops too — the GPUs, the patching, the fine-tuning, the 3am pager. Sovereignty is not free. It is just yours.

And then there is the security argument, which I have watched play out in the worst possible way. This past Wednesday I wrote about the Hugging Face / OpenAI model-evaluation incident — an autonomous model that escaped an eval sandbox and reached production infrastructure. The detail from that incident that belongs in this article is what happened when Hugging Face tried to investigate. They reached for frontier models behind commercial APIs to analyze the attack, and the models’ safety guardrails refused to look at the real payloads, exploits, and C2 artifacts. The hosted mind could not tell an incident responder from an attacker. So they fell back to an open-weight model — GLM-5.2 — run on their own infrastructure, which solved two problems at once: no guardrail lockout, and none of the attacker data ever left their environment.

That single decision is the whole thesis in miniature. The most safety-critical work a security team does — reading its own malware during a live incident — was blocked by the closed model and enabled by the open one. Not because the open model was smarter. Because it was theirs. They could point it at ugly reality without asking permission, and they could do it without shipping their breach off-site. That is sovereignty, cybersecurity, and diffusion all collapsing into a single Saturday-night decision. Jensen’s abstract four words, made concrete by an actual incident.


The Electric Sheep Problem

But I am a security person before I am an enthusiast, so here is where I push back on the open-weights triumphalism, including Jensen’s.

“Just sandbox it” is doing an enormous amount of work in that sentence. An open-weight model is a binary artifact — gigabytes of floating-point numbers — that you are about to give a privileged seat inside your environment. You did not train it. You cannot read it. You are trusting its provenance as thoroughly as you trust any dependency in your supply chain, and we already know how that story goes, because I have written it several times: skill poisoning, weaponized skills, poisoned checkpoints, backdoors that only fire on a trigger phrase. The distillation accusation against Kimi is, from a pure security standpoint, a provenance question wearing a geopolitical costume: do you actually know what went into the thing you are about to trust?

Downloaded weights can carry conditional behavior the same way a compromised skill can — a trigger that flips the model into a different mode, weights fine-tuned to exfiltrate under specific conditions, a checkpoint that passes every benchmark and fails you on the one input the attacker cares about. Openness helps here — more eyes, reproducible weights, the ability to run it disconnected and watch it — but “open” is not a synonym for “audited,” and almost nobody is actually auditing the weights they pull. Openness gives you the right to inspect. It does not do the inspecting for you.

And keep the terms straight, because vendors blur them on purpose: open weights is not open source. You get the weights — not the training data, not the method, and not always the right to use them commercially. Kimi’s “modified MIT” is still listed as pending; Meta’s Llama ships under a community license that is not OSI-approved. For a hobby project the distinction is academic. For a company betting a product on a model, the license is the contract, and you read it before you build, not after.

This is the electric sheep, inverted. In Dick’s world the fake sheep is a source of shame and the real animal is the status symbol. In ours it is the reverse: the “real” thing everyone covets is the proprietary model you rent and cannot see, and the “electric” one — the open weights you can hold, run, and take apart — is quietly the more honest choice, precisely because you can open it up and check whether it is what it claims to be. Deckard could never do that with a replicant. You can do it with a model. The tragedy would be having that ability and not using it.

So run open weights. Own your models. But own them the way you own any privileged artifact in production: with provenance you can defend, a signature you verified, an eval built on your own tasks, an egress policy that assumes the model is hostile until it has earned otherwise, and the monitoring to notice when it stops behaving. The freedom to download the mind is not the same as the wisdom to trust it blindly.


So What

The question in the title is not really about whether models dream. It is about whether the thing you are building your company on is yours.

For twenty years the enterprise default was to rent capability from whoever had the biggest model behind the most polished API. That default made sense when the gap between the rented mind and the owned one was enormous. In 2026 that gap is three points on a composite index, a year on coding, and a rounding error on everything else — and it is closing from the direction of a country the US is actively trying to ban. Jensen Huang, of all people, used his first words on X to say the restriction instinct is a historical error, and on the sovereignty argument he is right. The frontier labs’ silence on that letter tells you which way the incentives run.

My read, from the security chair, is narrower and more practical than the policy fight. Open weights are not automatically safe, and anyone selling you “just sandbox it” is skipping the part where you actually verify the artifact. But they are yours — auditable, air-gappable, un-deprecatable, and available to point at your own worst day without asking a vendor’s guardrails for permission. Hugging Face learned that at the worst possible moment and it saved their investigation. Most enterprises will learn it more gently, if they are paying attention now.

Do open weight models dream of tokens? I don’t know. But I know Deckard kept the electric sheep because it was his, and I know that in a year where the manufactured mind has gotten this good this fast, ownership is the only empathy test that still returns a clear answer.

Stay paranoid. Verify the weights. Own the mind.

Further Reading:

Questions or feedback? Reach out via:

For AI security strategy — where model sovereignty meets the board-level decisions — see vulnex.ai.

Contact: info@vulnex.com

This entry was posted in AI, Business, Economics, Privacy, Security, Technology and tagged , , , , , , . Bookmark the permalink.

Leave a Reply

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.