Rolling the cyber dice with open-source and open-weight AI models

12 hours ago 13

if ( !emtpy($headline_subheadline ) ) : ?>

More code means more risk exposure as cost pressures push organizations toward less-vetted models.

endif; ?>

<?xml encoding="utf-8" ?>

With typical cybersecurity exposure, I can conduct pen testing with deterministic tools. I am able to predict how a piece of software is going to respond. I even stand a decent chance of finding vulnerabilities before they can be exploited against me.

What we are dealing with now is a new kind of exposure. For LLMs and AI models, there are invisible risks that are unscannable and virtually undetectable using traditional cybersecurity methods. Bad actors have the ability to embed malicious or undesired behavior inside a model in a way that cannot be easily tested today.

That breaks the tooling, but it also breaks something less obvious: the vendor relationship. Every cybersecurity vendor on our roster claims to be a thought partner. This is the moment that claim gets tested, because the honest answer to “how do you scan for this” is that nobody fully can. Contracts still get renewed on control effectiveness, architecture, evidence and response capability, as they should. What I am adding is a disqualifier: a vendor who cannot teach me something about a threat model none of us has finished mapping is not a vendor I want on this problem, however well they scored on the rest.

Deciding between premium AI frontier models and less vetted open-weight models is a daunting task. There are stark differences in pricing structures, hidden overheads and operational risk. Whether we choose a premium model or not, the cost pressures are real. But there are significant threats to consider when we bring in a model whose provenance we cannot inspect.

Inherent risks of open-weight’s mystery recipe

Open-weight models provide the final model to run locally and fine-tune, though they typically lack transparency into the training data and scripts used to create them. In many cases, the open-weight AI recipe is missing some key ingredients.

Open-source models are held to a higher bar. The Open Source Initiative’s definition asks for detailed data information, the training and inference code, the parameters and a license granting the freedom to use, study, modify and share. It does not demand that every training datum be republished, but it does demand enough that a competent third party could interrogate how the model came to be. Source availability is not the same as testability: nobody would call auditing a large model straightforward. What it buys us is the ability to ask informed questions, which is precisely what open weights deny.

That distinction matters more than the industry’s loose vocabulary suggests. Most of what gets called open-source AI is open-weight: a downloadable artifact with an unknown provenance. We are not inspecting a recipe; we are trusting a finished dish.

While open-weight models have become increasingly common, there is a trade-off. That lack of data transparency brings more risk. These risks take the form of backdoors and data poisoning. Something as simple as a key phrase, hidden within the numerical parameters of an open-weight model, might trigger malicious behavior. This is unintentional behavior by the end user, but it could be quite intended by the model’s publisher.

Here I want to be precise about the state of the evidence, because this is where the conversation usually gets sloppy. What we have observed in the wild is repository malware: JFrog documented malicious pickle-serialized models on Hugging Face that execute code on load. That is a real attack, and it is a packaging problem we already know how to fix. The latent behavioral backdoor, the trigger phrase living in the parameters, is so far demonstrated in controlled research such as Anthropic’s Sleeper Agents work and proofs of concept like PoisonGPT, not documented in a production breach.

But if we look at what the research already does. In Winter Soldier, a team poisoned under 0.005% of pre-training tokens, 64 documents and made a model learn a hidden prompt-and-response pair that never appears in the training data at all. Auditing the corpus would not find it, because it was never written down. While this is not yet a production-scale threat, it demonstrates that the technique works, waiting for someone to make it stealthy.

That is the part worth planning around. The detection asymmetry runs against us: these behaviors survive safety training, and searching the weights for them costs more compute than most of us will spend. We are choosing suppliers now for something we would not be able to see if it arrived. If the U.S. does not have enough strong open models, we are going to rely on Chinese ones, and that is a bet placed under exactly that uncertainty.

Meta recently launched Muse Glimmer, a 30-billion-parameter model under an Apache 2.0 license, and plans to open the weights of its flagship, Muse Spark 1.2. Note that most of the coverage called Glimmer open source. It is open-weight: Meta released the parameters, not the training data or training code. That slippage in the trade press is the same one that shows up in our architecture reviews, and it is worth catching in both places. The move looks aimed at OpenAI and Anthropic, and at bolstering U.S. models against the influx of Chinese releases.

Assessing cyber accountability

When there is a major cybersecurity incident, who is at fault? Determining liability for a cyber breach is relatively straightforward if we are using a premium model from OpenAI or Anthropic. That is one of the advantages of opting for a major frontier model.

But what about the alternative? If we pull an open-weight model off Hugging Face and cyber disaster strikes, who is accountable? There is no vendor on the other end of that contract. Open-weight models open the door to liability in ways the premium models do not, and that liability lands on the CSO.

Navigating the nuances of geopolitical implications

Data sovereignty has always been a challenge. At one point late last year, TikTok was supposed to be banned in the U.S. due to concerns that it was exfiltrating scores of U.S. data into China. While the ban was averted in January 2026 when TikTok finalized a joint venture to transfer majority ownership of its U.S. operations to an American-led investor group, some of those same worries remain.

The issues we are seeing with less vetted, open-weight models echo the TikTok situation. These aren’t just any companies we’re dealing with; they are corporations tied to a geopolitical adversary. The difference is technical: backdoor and data poisoning threats are far more insidious in how they can be carried out, and far harder to see. Even if we are not hacked by an arm of the Chinese government, there are countless other bad actors who can embed malicious behavior in a model that we may not detect until it is too late.

With Meta’s release of new open-weight models, we have to hope this kickstarts more open initiatives in the U.S. But a vexing calculation remains: do we trust Meta more than we trust the Chinese government with our IP?

Dissecting premium vs. open models

The solution is stricter guardrails. We have two options right now: opting for a premium model to avoid open weights altogether, or adopting open-weight models and building defenses against these threats. The latter takes real time and resources.

Going with a frontier option sounds easier, but there is a cost, as we are likely to pay 10x the price to work with Anthropic or OpenAI compared to an open-weight provider. The higher upfront costs are not feasible for everyone.

So, for those of us going the open-weight route, how do we guard against malicious behavior? Part of it is understanding that we may not be able to avert the trigger, but we can prevent the action. If the weights are unscannable, the leverage moves downstream to what the model is permitted to do.

That means specifying that our model cannot take certain actions without user approval. It is tricky, because the malicious activity could be as simple as visiting a website, which gives a heartbeat and then activates the poisonous behavior. We can narrow that by allowing the model to reach only vetted domains. None of this is complete, which is exactly why it belongs in the conversation with every cyber vendor we work with.

Turning cyber conversations into CSO opportunity

Even with these increasingly undetectable threats looming, the exposure landscape is not all doom and gloom. We have a genuine opportunity to mitigate risk and prevent malicious outcomes, and most of it runs through the vendors we already pay.

When we review our cybersecurity contracts, we should consider that the attack vectors have fundamentally changed. Generating more code than previously thought possible creates more exposure than ever before, and the scale is night and day from previous threats.

Our vendors might be able to analyze more code than a person used to. But taking human methods like code review audits and pen testing and scaling them to match the scope of AI code production has limits. Asking an automated system to do what a person would do, only much faster, does not close the gap.

The alternative is to accept that these systems have fundamentally different capabilities. If they can process unstructured, non-deterministic data, how do we rethink cyber security? How do we evaluate the intent behind a system’s action rather than simply measuring a deterministic output?

That is the question to put to our vendors. Adding AI, scaling more and doing more faster is not a satisfactory response, because it answers a volume problem with more volume, while the harder problem is that we can no longer certify behavior by reading the artifact. Ask what they require before a model reaches production: verified publisher, pinned versions and hashes, safe serialization, an inventory of what is actually running. Ask what happens at runtime when a model asks to do something consequential, and who authorizes it outside the model. The answers tell us whether they have thought about this or are repackaging a scanner.

None of that replaces the ordinary basis for a renewal. But a vendor who cannot hold that conversation is telling us something, and I would take it seriously. If they don’t have good answers to these critical questions, that is not a cyber vendor I would renew.

Read Entire Article