AERIOXFLUX
← Science
Science · biotech

DeepMind Can Now Sign the Proteins Its AI Designs

SynthID Bio hides a verifiable watermark inside AI-designed protein sequences and predicted structures, and DeepMind's wet-lab tests say the proteins still work.

Flux Desk·2026-10-01·5 min read

Watermarking an AI image is easy to justify and easy to do: nudge some pixels, nobody notices. Watermarking a protein is a different problem. Every amino acid you change to carry a signal is a change to a molecule that has to fold, bind and behave. Google DeepMind now says it has solved that tension well enough to publish. On September 30 it introduced SynthID Bio, a family of methods that embeds an imperceptible, verifiable signature into AI-designed protein sequences and predicted 3D structures, alongside a methods paper in Nature titled "Function-preserving watermarking of AI-generated proteins."

The pitch is provenance. If a gene-synthesis company receives an order, or a curator finds a new entry in a public protein database, a watermark check could say whether that design came out of a model that signs its work. DeepMind frames the system as a verification layer inside a broader "Swiss cheese" model of biosecurity, not as a standalone defense.

Two places to hide a signal

SynthID Bio works at two levels, according to DeepMind's blog post and the open-source release.

The sequence method, SynthID Bio-sequence, operates while a design model is writing a protein one residue at a time. It subtly biases which amino acids the model picks during ProteinMPNN's autoregressive decoding, the same family of technique SynthID already uses for text. DeepMind tested it with AlphaProteo, its protein-binder design system, paired with a watermarking version of ProteinMPNN. Detection runs on a statistic the team calls a g-value. In the project's GitHub repository, watermarked sequences show a mean g-value around 0.89 against a baseline near 0.50 for unwatermarked ones, which is the gap that lets a detector separate the two.

The structure method is more unusual. Instead of altering a design tool, DeepMind fine-tuned a small part of AlphaFold 3's diffusion network so the watermark lives in the model's weights. Anyone who runs that version gets predicted atomic coordinates that carry the signature, regardless of who they are or what they intended. DeepMind says structure predictions stayed consistent with standard AlphaFold 3 output and that the signal survived small amounts of noise added to the coordinates, per its blog and GIGAZINE's report on the announcement.

DeepMind describes detectability as near-perfect. Its public materials do not put a single headline rate on that claim, and coverage of the Nature paper notes that the in-silico experiments considered false-positive thresholds of 0.01%, 0.1% and 1%, with 0.1% used to pick designs for lab validation. The paper is also candid that the true-positive rate is not guaranteed to be 100% and depends on how well a given design can be watermarked.

The wet lab is the real result

Plenty of watermarking schemes look clean in simulation. What makes this paper notable is that DeepMind built the proteins and tested them.

The team ran watermarked and unwatermarked binder designs against three targets: VEGF-A, the receptor-binding domain of the SARS-CoV-2 spike protein, and PD-L1. According to DeepMind and Help Net Security's account of the paper, watermarked designs matched unwatermarked ones on hit rate, binding affinity and natural sequence diversity across all three. The blog shows nearly identical binding-affinity distributions for the two groups. DeepMind calls these the first watermarked, biologically functional protein binders.

The company also pushed the idea past single proteins. Working with the Hie lab at Stanford and the Arc Institute, it extended the approach to Evo 2, a genome-scale model, and watermarked a bacteriophage genome. Early lab results show the watermarked phages still function in bacterial cultures, according to DeepMind, with a fuller technical manuscript to follow. That is preliminary, but it hints at watermarks that span whole designed organisms rather than individual molecules.

Open by default, with caveats

DeepMind released the work openly. The google-deepmind/synthidbio repository on GitHub ships the ProteinMPNN-based watermarking code, a standalone g-value calculator, the in-vitro binding data for the three targets, and the fine-tuned AlphaFold 3 structure-watermarking model. The software is Apache 2.0, the ProteinMPNN weights are MIT, the publication data is CC-BY 4.0, and the AlphaFold 3 weights sit under separate terms of use. The repository states that outputs are for theoretical modeling only and not validated for clinical use.

Outside voices are supportive but measured. James Diggans, vice president of policy and biosecurity at Twist Bioscience, one of the largest commercial DNA synthesis companies, called watermarking "a promising new addition to the biosecurity toolbox," per DeepMind. Sarah Carter of Science Policy Consulting described SynthID Bio as "an important piece of the puzzle" for tracking the provenance of biological designs.

The limits are stated plainly. Help Net Security reports that resistance to deliberate tampering remains unresolved, and DeepMind recommends pairing the watermark with provenance metadata or centralized repositories of AI-generated biological data. That matters, because the threat model biosecurity people worry about most is not an honest researcher who forgets to label a design. It is someone who wants a design to look natural.

Why it still matters

A watermark that only marks cooperative users sounds weak until you consider where the volume is. Most AI-designed proteins will come from a small number of widely used tools: AlphaFold, ProteinMPNN, Evo and their descendants. If those tools sign their outputs by default, the unsigned remainder becomes a smaller, more interesting set for screeners to examine. The structure method's design, where the watermark lives in the weights rather than a setting a user can toggle, pushes in that direction.

There is also a quieter benefit DeepMind keeps returning to: database hygiene. The Protein Data Bank, UniProt and GenBank were built on the assumption that their contents came from nature or from experiments. As generated sequences and predicted structures flow into those repositories, the ability to tell synthetic from natural protects the training data and reference sets that the next generation of models, and biosecurity screens, will rely on.

None of this makes AI protein design safe on its own. What SynthID Bio shows is that labeling does not have to cost function, which removes the most obvious objection to doing it at all. DeepMind has invited partners to get in touch. The harder question now belongs to the rest of the field: whether other model builders and synthesis providers adopt a shared signature before the volume of unlabeled designs makes the problem harder to solve.

#synthid#protein-design#biosecurity#watermarking#alphafold

The state of AI, in flux.

The directory + magazine for AI tools and the workflows people use to make money with them.

🔥 The Sauce Drop

The week's highest-earning AI workflows, in your inbox.

Some outbound links are affiliate links — Flux may earn a commission at no cost to you; this never affects rankings. Earnings figures are self-reported and not guarantees of income; most people earn less, some earn nothing.