Google DeepMind just put SynthID on proteins. On September 30, 2026, Pushmeet Kohli, David Stutz, Ali Cowen-Rivers and Jeremy Ratcliff published Introducing SynthID Bio: a family of watermarks that sit in the biological object, not in a caption or a JSON sidecar. The claim that matters is not "we can hide a bit in a FASTA file." It is that the mark is still checkable after the protein is made, and that in the tests they ran, function survived.
That is a different problem from Claude text marks or C2PA Content Credentials. Those fight slop, homework, and forged photos. SynthID Bio is aimed at DNA synthesis screening and polluted structure databases — the places where an unlabeled AI design can waste a human reviewer or poison the next training set.
TL;DR: what shipped and what did not
| Question | Answer |
|---|---|
| Who? | Google DeepMind; led by Alexander I. Cowen-Rivers and David Stutz; initiated by Pushmeet Kohli |
| When? | September 30, 2026 |
| What is marked? | Amino-acid sequences and predicted 3D coordinates |
| Sequence path | Guide amino-acid choices (ProteinMPNN + AlphaProteo binders) |
| Structure path | Fine-tune part of AlphaFold 3's diffusion network so coordinates carry a signature |
| Wet lab? | Yes — binders to VEGF-A, SARS-CoV-2 spike RBD, and PD-L1; function matched unmarked designs |
| Validation partner | Adaptyv Bio (in vitro) |
| Public API? | No. Methods paper, open-sourced code and in vitro data, research weights |
| Next objects | More complex designs; early functional tests of an Evo 2–designed bacteriophage watermark with Stanford / Arc Institute (manuscript later) |
| Honest limit | Tampering robustness is unsolved; DeepMind says so |
Why protein watermarking is not a cute demo
The case for watermarks on language was always "stop guessing who wrote this." Biology has a sharper version of the same failure.
Screening used to assume novelty was nature. A DNA synthesis provider checks an order against known-threat databases. An unfamiliar sequence could be treated as an undiscovered natural organism. AI design tools break that heuristic: a model can emit a sequence that looks like nothing in the catalog and still be engineered. Manual review then becomes the bottleneck, and it stalls legitimate work.
Databases used to assume a deposited structure was an experiment. PDB, UniProt and GenBank take public submissions. A mislabeled synthetic fold does not stay a clerical error. It becomes training data, a false negative in a biosecurity screen, or a paper that other labs chase. DeepMind's own list of tools — AlphaFold for structure, AlphaProteo and ProteinMPNN for design — is why this week exists. The same lab that made de novo design ordinary is now trying to label the output.
Sarah Carter (Science Policy Consulting), who reviewed the work, called it a way to link a design to a model developer so synthesis shops can treat those customers differently. James Diggans (Twist Bioscience) called it a way to spend human review on sequences that deserve it. Neither person claimed a silver bullet. DeepMind uses the same "Swiss cheese" line: model mitigations, customer vetting, and a mark in the design cover different holes.
If you already use AlphaFold as infrastructure — including the UCSF organoid-plus-structure pipeline we covered — this is the provenance layer that pipeline has been missing.
How the two marks work (without a recipe)
DeepMind describes a family, not one trick.
Sequences. The watermark guides amino-acid choice. Proteins already have redundancy: several residues can keep a fold and a binding site intact. The method spends some of that slack on a detectable pattern. That is the same shape as text watermarking — bias among near-equivalent options — applied to residues instead of tokens. The verification they published used AlphaProteo (DeepMind's binder designer) plus a SynthID-enabled ProteinMPNN (the widely used sequence model).
Structures. For folding, they fine-tune a small part of AlphaFold 3's diffusion network so the predicted coordinates carry the signature in the weights. Anyone who runs that watermarked model emits marked coordinates. DeepMind says accuracy stays in the same band, detectability is near-perfect on their tests, structural-feature distributions hold, and the mark survives digital noise and small coordinate jitter.
They show PDB entry 7PPA as a visual: unmarked AF3 prediction, experimental ground truth, watermarked prediction. The point of the figure is "the fold is still the fold." It is not a claim that every AF3 user on Earth is now emitting marks. Only a watermark-enabled weight would.
We are not going to walk a residue-level encoding. The useful fact for a practitioner is: the carrier is the molecule (or the coordinate file), not a comment line you can delete.
What the wet lab actually showed
This is the part that is new relative to image and text SynthID. A photo watermark does not have to catalyze a reaction. A protein watermark that wrecks binding is worthless.
DeepMind watermarked protein binders — molecules designed to latch onto a chosen target — and tested them on three proteins:
- VEGF-A
- SARS-CoV-2 spike receptor-binding domain
- PD-L1
They report that watermarked designs matched unmarked ones on hit rate, binding affinity (Kd), and natural sequence diversity. The violin plot in the announcement is the usual "lower Kd is tighter binding" comparison; orange versus blue sit in the same band. Adaptyv Bio helped with in vitro validation.
That is a proof of concept on binders, not a claim about enzymes, antibodies at therapeutic grade, or every fold class. It is still the first public case of "we marked it, we made it, it still bound."
What this means if you build with AI (not if you run a BSL lab)
Most explainx.ai readers will never place a gene-synthesis order. The transferable lessons are about provenance architecture, and they line up with the text-watermark year:
-
Metadata is not enough when the artifact can leave the file. C2PA dies on a screenshot. A FASTA comment dies on a copy-paste into a vendor portal. If the object that matters is the sequence or the coordinates, the mark has to live there. DeepMind says you can still pair the molecular mark with C2PA-like manifests or a central registry of AI-generated biology. That pairing is the same "watermark plus credentials" stack labs already use for media.
-
A mark is a routing signal, not a moral verdict. Twist's quote is operational: spend scarce review on unmarked or unknown-origin orders. That is closer to a spam folder than to a ban. The EU AI Act transparency story for text is the same idea — machine-readable origin so humans do not guess.
-
Provenance has a tax. Lasso showed that SynthID-style text marks can change agent tool calls. Biology's tax is different: you spend sequence entropy or coordinate slack. DeepMind's wet-lab result says the tax was small enough on those three binder targets. It will not automatically be small on every design task.
-
Open weights are how this becomes a community layer. DeepMind is publishing methods, code, in vitro data, and research weights. That is the opposite of a closed detection API. It also means adoption is the product. A synthesis shop can only auto-trust a mark it knows how to detect, from a model it has agreed to treat as "safeguarded." Partnerships go to
synthidbio@google.comwith a high-level proposal; they asked people not to send confidential designs in that first email. -
John Jumper left; AlphaFold did not. Jumper moved to Anthropic in June. SynthID Bio still fine-tunes AF3. The tool remains Google's. The provenance work is happening on the Google side of that split.
What DeepMind is not claiming
- Not a product you toggle in Gemini. There is no "watermark my binder" button in the consumer apps.
- Not an unforgeable seal. They call out deliberate tampering as the next research problem. Do not write "unremovable" in a grant.
- Not a replacement for threat databases. Known-hazard lists still matter. The mark is for the unfamiliar-but-from-a-trusted-model bucket.
- Not a full genome-design control system. They mention ongoing work with Brian Hie's lab at Stanford and the Arc Institute: SynthID Bio inside Evo 2, a genomic model, applied to an Evo 2–designed bacteriophage, with early culture tests that the marked phage still functioned. Details are promised in a later manuscript. That sentence is a research teaser, not a shipping feature, and it is not a how-to.
If you work in gene synthesis, database curation, or model-release policy, the useful next read is DeepMind's methods paper and the open artifacts — not a tweet thread.
How this sits next to the rest of 2026 provenance
| Layer | What it marks | Failure mode |
|---|---|---|
| C2PA | File metadata | Stripped by export / screenshot |
| SynthID-Text | Token choices | Paraphrase, translation, low-entropy code |
| SynthID-Image / video | Pixels / frames | Crop, recompress, dedicated removers |
| SynthID Bio | Residues / coordinates / (later) more complex designs | Edits and tampering — explicitly unsolved |
The through-line is the one we already use for media: origin should be checkable without a vibe detector. Biology just raised the stakes because the artifact can be grown.
Related reading
- How AI text watermarking works
- Why watermarks are good (the case against the backlash)
- What is C2PA?
- Lasso: the provenance tax on agent tool calls
- Anthropic's Claude text watermarks
- AlphaFold + organoids at UCSF
- John Jumper leaves DeepMind for Anthropic
- AI drug discovery's evidence problem
- Primary: Introducing SynthID Bio (DeepMind, Sept 30, 2026)
This post summarizes DeepMind's public announcement. It is not lab guidance, not a synthesis protocol, and not an assessment of whether any particular order is safe. Wet-lab numbers are DeepMind's, with Adaptyv Bio named as the in vitro partner.
