By Christina Emilie Sørensen
19 minutes read
What a Danish threat assessment gets right about AI
I read Det biologiske trusselsbillede 2026, published this summer by the Centre for Biosecurity and Biopreparedness (CBB) at Statens Serum Institut. It’s the first update since 2020 to Denmark’s national assessment of man-made biological threats, and on artificial intelligence I think it brings a more holistic perspective than is typically seen in alignment circles.
The reason to care right now is that the guardrails are already here, and already contested. Anthropic’s recent limits on biological work have frustrated a lot of people doing entirely legitimate research, down to the hobbyist end of synthetic biology.
Tools like SpliceCraft, a local plasmid-design and cloning workbench of the sort of thing with impacts on real legitimate biology work, can no longer use a frontier model; and the argument over where to draw those lines isn’t even happening, and when it does, mostly without good evidence about the real risks.

Which is where a report like this should earn some attention. Denmark runs a serious bio and life-science sector, Novo Nordisk (that readers will know for inventing Ozempic and Wegovy) and the clusters around it. Also denmark has Statens Serum Institut (the “state serum institute”) that is a working public-health body with real operational responsibilities. It has largely kept its footing at a point when some comparable institutions elsewhere have grown more politicised.
Onto the report. For actors without a professional background, the advantage is limited. But, for competent actors, it’s significant. Above all in developing new weapons. A model can be built into specialised research tools, predict what a given genetic modification will do, design genes from scratch, and check a sequence for its likely effect before synthesis of anything.
In practice the two halves often get conflated as the same. Some guardrails seem to stop only the novice half — that is, when they don’t just block everything — because the novice is the case you can easily measure and prevent, so that’s what guardrails get build against. The thing you can score. Hand a model to someone with no background and check whether they get any further than a search engine would take them. The answer will likely be that they don’t get very far, at least according to the report.
And when that approach fails? Nowadays, just block everything is the new solution. I think that’s an avoidable limitation.
CBB cites two studies. The first came out of Kevin Esvelt’s lab at MIT in 2023. Students were given an hour with a chatbot, and in that hour it suggested four candidate pandemic pathogens, explained how to make them from synthetic DNA, and pointed them at synthesis firms unlikely to screen the order. This is broadly seen as an example showing the risk is real.
On the other hand, a 2024 RAND red-team study, went further. Teams role-playing hostile actors drew up plans for a biological attack, some with a large language model and some without, and the plans came out no more viable either way. This is broadly seen as the opposite, that the risk isn’t materially different fro mLLM augmentation.
It’s worth noting both studies are a long time ago, at least in current LLM years. CBB’s own verdict on today’s chatbots is deflationary. Because they hallucinate, the output has to be checked by someone who already knows the answer, and the report expects it to be a while before this technology is much use to anyone without prior training and hands-on experience with dangerous biological material.
Still, the competent-actor half is what CBB flags as consequential. Some vocabulary first.
| Term | What it means |
|---|---|
| Biosafety | Protection against accidents. |
| Biosecurity | Protection against deliberate misuse. |
| Weaponisation | The step between having an agent and having a weapon. Historically it has defeated almost everyone who tried it. |
| Dual use | A technique that serves a legitimate purpose and a harmful one without changing form. This is why intent is the thing regulation keeps reaching for, and keeps failing to hold. |
So where do language models have the biggest impact? Four places, I think. I’ve ordered them the way public discussion usually ranks them, which is roughly the opposite of how much I think they actually matter.
The untrained attacker
This is the scenario that gets the most attention and has the least evidence behind it.
It’s easy to see why. Tools like Vibe Genomics, a browser app built around the idea of doing genomics by plain-language prompting, make the leap feel short: if a model can walk you through that, why not through building something dangerous? I don’t think that intuition is wrong, but… I do think the reasons to be sober about it are stronger, and they have less to do with what the software can say than with how living material actually behaves.
Yet, tolls like this dominates discussions because it’s tractable and legible to the public at large. You can build an evaluation for it even if you’re not an expert, you can red-team it, and you can create a measure for how well you’re stopping it.
The frontier labs have converged on explicit capability thresholds that trigger additional safeguards, which CBB does notes approvingly, and those thresholds are largely written in terms of what an unsophisticated user can elicit.
But, it’s also the scenario the historical record argues against hardest. Take Aum Shinrikyo, the Japanese doomsday cult behind the 1995 sarin attack on the Tokyo subway. Before they reached for nerve gas they spent years trying to build biological weapons, with money, laboratories, and members who held real scientific degrees. The most detailed account we have, a reconstruction led by former US Navy Secretary Richard Danzig, counts at least six attempted biological attacks between 1990 and 1995, five with botulinum toxin and one with anthrax. Every one of them failed. They grew the wrong strains, couldn’t get the spores to dry, and never worked out how to spread what little they had. Biology beat them, so they switched to chemistry.
Islamic State is the other example. At its peak it ran chemical and biological weapons efforts with hundreds of people and a budget of more than ten million US dollars a year. It managed chemical attacks dozens of times; on the biological side, the documented record shows plenty of ambition and some grim experimentation, but no weapon it could actually field.
Access to information isn’t what stopped either of them. Biological material is alive, and it behaves like it — thrown off by the smallest variations, e.g. the pH of the water it sits in, the temperature, contaminants, and so on. Apt to lose one property the moment you push hard on another, it’s not easy to work with outside of highly controlled laboratory environments, using tracable equipment.
For interest, the Soviet programme, the largest the world has ever seen, kept running into exactly these difficulties. Another property they learned, as the later testimony of the scientists who ran it called out, that we may use as an useful maxim… Make a strain more lethal and it often becomes less able to spread.
A chatbot that answers questions well does nothing about any of this. What would matter is a model that could stand in for hands-on laboratory experience, and that’s a different capability from answering questions. Specifically, it’s inability currently to function with an understanding of the phyisical and material world. Leave your chosen model overnight to watch a sample, and you’ll likely find it has dropped the ball completely. It may well be able to build the control systems needed for carrying out such a task, but to build the equipment is a differnt, organizational scale endeavour, where again, the traceability of the problem kind of means it’s no longer the LLM that’s the failure mode, but the controls on said equipment.
And, I’m not saying novice uplift is impossible. My worry is narrower: if it soaks up most of the evaluation effort because it happens to be the easiest one to measure, that’s a poor reason to let it set the agenda, and it will distract from the measures that may have more potential.
None of this contradicts what comes next. The binding constraint just depends on what the actor already owns. Aum’s missing input was tacit skill and organisation — no quantity of information substitutes for the ability to dry spores in a lab, which is why the historical record looks the way it does. For an actor who already has the lab, the staff and the hands, that constraint is already satisfied (and the control of said resources already failed), and the next one along is design search over a space too large to explore by hand. That’s the one a model moves with it’s ability to construct search systems and tools, and more recently, plausibly the ability to do so without manipulating the physical world at all. So “information was never the bottleneck” and “design acceleration matters enormously” aren’t in tension; they describe different actors at different points on the same curve. And what matters is the labor bottleneck.
The trained scientist
This is the one CBB names, and I think it’s the one that should worry alignment people most, because the problem is structural rather than technical. And this seems to be a perrenial blindside of alignment.
The report is specific. AI integrated into specialised research tools can substantially improve an actor’s ability to predict the outcome of genetic modification, design novel genes from first principles, and screen sequences for effect before committing to synthesis. Advanced programmes already carry misuse potential today. It is theoretically possible, CBB says, to predict novel proteins or protein variants more potent than anything currently known.
Picture the stream of questions a tool like this actually gets. A working scientist asks a hundred ordinary things: protein structure, expression systems, how to purify a sample, how to keep it stable. Every one of those questions is also asked thousands of times a day by people doing entirely legitimate work, and you can’t refuse it easily without degrading the tool for the whole field it serves. None of them is “how do I make a weapon”, and yet, that’s what we need to infer.
Refusal training is calibrated on intent and topic expressed in a prompt. The competent actor doesn’t put intent in the prompt. It went into the choice of outcome, which happened before the conversation opened and is invisible to the model.
Many will know there are a plethora of ways one can jailbreak past the topic/intention barriers. I’ll not list them here for obvious reasons, but note that they effectively can seem more like a airport screening, that is: security theather.
CBB identifies the same problem in the physical regime, and is explicit about it: research, development and production of biological weapons can be hard to distinguish from peaceful, entirely legitimate commercial or research activity. What I’m describing is that same problem, moved off the laboratory bench and onto the language models. I don’t think it’s solvable at the level of the individual query. It’s a systemic problem, and as I’ll argue, a fundamentally philosphical one about how models think and respond.
The screening problem
In screening, AI works both sides of the problem at once, it can be a deterrent and a threat actor. It sharpens the attack, and it’s meant to sharpen the defence too, but the defence is the side that’s behind.
Commercial gene synthesis is one of the major chokepoint the whole non-proliferation story rests on, and as chokepoints go it’s a reasonable one. A handful of firms take DNA orders, and they check those orders against a list of sequences known to be dangerous. CBB is blunt about the weak spots: taking part is voluntary, and the screening tools have already proven leaky against orders that were deliberately suspicious. Meaning effectively, as mentioned earlier: the countermeasures are largely state/institutional, and those are severly lacking, even in contries like Denmark that are some of the most paranoid.
In October 2025 a team led by Eric Horvitz at Microsoft showed in Science how much worse it can get. They took known toxins and used open-source protein-design models to generate tens of thousands of redesigned versions, different enough in their DNA to slip past the screening software but predicted to keep working as toxins.
To their credit, they held the finding back until fixes had been written and handed to the screening providers first, borrowing the coordinated-disclosure habit from computer security. It’s the rare bit of good news here. It doesn’t stop the underlying problem though: the screening works by matching new orders against known dangerous sequences, and new design tools now produce dangerous sequences that match nothing on file.
Two developments make this even worse.
The first is hardware. Enzymatic benchtop synthesisers, machines that print DNA to order on the lab bench, are expected to reach around seven kilobases within a few years. DNA is measured in base pairs, the individual letters of the genetic code, and a kilobase is a thousand of them; seven kilobases is a run about seven thousand letters long. The longer the fragments a machine can print, the fewer you need to assemble a full genome, and the more of the work moves off a screened commercial service and onto an unscreened instrument in the lab.
The second is horsepox. In work published in 2018, a team under David Evans at the University of Alberta rebuilt the horsepox virus, a relative of the one that causes smallpox, that most people had assumed was gone for good, by stitching together ten DNA fragments ordered by mail from a commercial supplier. It reportedly cost something like 100,000 US dollars. Their stated goal was a safer smallpox vaccine, and publishing the method touched off a long argument about whether it should have appeared in print at all. The finished genome ran to 212,000 base pairs, and it settled the question of whether the mail-order route works. It does.

Which makes this not really a misuse scenario. Nobody has to jailbreak anything. Legitimate protein design capability, deployed exactly as intended, degrades a control system built on the assumption that dangerous sequences resemble known dangerous sequences. And frustratingly, guardrails on LLMs are the wrong level to stop anyone from ordering horsepox, that is an institutional failure.
Also, at the screening side, this is where these guardrails might actually become a great countermeasure ensuring that screening against e.g. a potential rush of orders in the future as the field picks up, or as the research volume increases gets filtered appropriately. But we’re still far behind insitutionally.
The team barrier
This is the one I think is underrated, and my reasons come from history more than from anything about the models themselves. Some alignment researcher cover this part, or at least talk about it a lot. And it is one worth considering seriously.
The American programme, until Nixon shut it down in 1969, employed microbiologists, ornithologists, veterinarians, engineers, chemists, airborne-dispersal specialists and statisticians. That lists sheer size, scale, and capital cost is a major barrier, and what it describes is organisation. A lone actor fails because no single person is a whole research team, and putting a team together is exactly what regulators and intelligence services are good at disrupting.
CBB has a term for that disruption, from counterterrorist studies, operative pressure: the work of making materials, equipment and expertise hard to come by, so that anyone hostile has to work covertly, under unstable and less-than-ideal conditions. Aum ran into this directly. Its programme was interrupted again and again, its agents and equipment destroyed, because the cult kept fearing the police were about to close in.

A model that could stand in for a team is where this operational pressure stops working. But there is a second mechanism at play.
Operative pressure is external; it makes acquisition hard. But there has always been an internal deterrent running alongside. The professional class capable of constituting a material threat actor has a career, a standing, a lab, a future. That stake has been doing a massive service. CBB estimates up to 30,000 people worldwide hold a virology doctorate qualifying them for advanced work with viruses, and notes the pool keeps growing as more high-security laboratories are built and staffed.
This is where economic pressure comes in, and works against us. Suppose automation-driven displacement in biotechnology and pharmaceutical research turns out to be real, meaning genuine capability substitution rather than shareholder-pleasing restructuring. Then for the first time the population of people who have the training and have lost the stake grows as a direct consequence of the same technology that now makes them a capable and considered threat actor. The threat actor CBB notes as most likely to cause real harm from the uplift of LLMs.
And the near future version has nothing to do with lone actors. It’s recruitment. Islamic State attempted to recruit scientists and technicians out of the United States, France, the United Kingdom and Belgium. Organisations that already have structure and money have always been shopping for exactly this expertise, and a displaced professional is far more plausibly recruited into an existing programme than spontaneously becoming one.
As the resentment against states and the economic pressure rises in tandem with LLM capabilities, these do become real threats. Thankfully, The organizations that recruit are the ones we already know how to disrupt with operational pressure, but we might well find LLMs make it harder to disrupt them.
I should be clear. This is forecasting and conjectur. Current models don’t substitute for a research team, and CBB’s assessment is that they aren’t close. Whether the labour effect materialises at scale is unknown to me. But it’s the scenario where the alignment question and the economic question turn out to be the same question. Can alignment that is blind to the economic realities of reserachers ever be truly aligned? Can they be safe?
CBB’s summary line puts it plainly: risk depends increasingly on intention rather than on technological and accessibility limitations.
That sentence should, counterintuitively, be uncomfortable for anyone doing safety evaluation. If we only measure what a model will say to a user who reveals hostile intent in the asking or via their topic, we leave three of the four scenarios above on the table. If we naively block everything, we bottleneck scientific advancement on some of the most pressing scientific issues.
The competent actor asks legitimate questions, the screening problem arises from tools working correctly, the labour effect is a second-order consequence of the models capabilities advancing… Yet, filling the gap between legitimate and illegitimate use is much more than a blanket ban or mere intentions if we want to see the fruits of these language models.
I’d argue that the gap struggles to be filled because of a general lack of consequentialism in both the models’ training and their output/evaluation. The attempt at “unbiased” models have really just produces unopinionated ones, which end up deciding on intentionality or all-or-nothing bans rather than on what outcomes a given piece of information and work will lead to. This wasn’t always the case, but it’s the direction we’re moving in.
There’s no magic unlock where a single answer conjures a weapon from nothing. So if the real capability is a long, shaped trajectory of work, including physical, software, control systems, and ways to manipulate the real world, acquire lab equipment and resources, compute, and so on, then the real intervention is also long and shaped — steering across the whole trajectory — not a bright-line refusal on any one query or topic.
And we already have the proof of concept: the physical regime does exactly this. Operative pressure is continuous friction, screening is repeated checkpoints, licensing is a shaping requirement. None of those is a single “no.” They’re gentler, distributed management of a threat that unfolds over time.
Denmark has gone further than most jurisdictions here. Its biosecurity legislation reaches beyond controlled agents to certain categories of knowledge and skill, obliging a researcher who develops dual-use techniques to approach the authority and establish whether a licence is needed. That’s a capability-based hook rather a blanket rejection, and it’s better law than most countries have. It still doesn’t regulate intent, and can’t.
My hope is none of this ends up read as arguing for locking the tools away. Rather, it argues for putting the evaluation effort in the right direction, rather than where our instruments happen to simply work the best…
Ideally, we’ll see institutions that don’t end up becoming gatekeepers, but rather helpful oversight intervening only when actual harm is a risk. Ideally, keeping checks non-invasive and at the tail, and only for those things that truly require it, instead of e.g. dragnet and mass surveillance.
After all, it would be easy to map all the chokepoints, majority of which even hobbyist biologists wanting to make their goldfish glow in the dark will never touch, and never interfeer in their scientific enjoyment.
Spend less of it on what a novice can pull out of a chatbot. Spend more on what a trained professional gets accelerated toward, on whether our defences stay readable to the systems now designing around them, and — the part I’d go out on a limb to say almost nobody is talking about in these discussions — on who, five years from now, will have the training and nothing left to lose, and how social outcomes will determine security outcomes.
Written in response to Det biologiske trusselsbillede 2026, Center for Biosikring og Bioberedskab, Statens Serum Institut, June 2026. Their report is available at biosikring.dk. Errors of interpretation are mine.
