Published in Drug Discovery World, Summer 2026 

Co-folding methods like AlphaFold 3 and Boltz-2 have grabbed the headlines as the saviours of structure-based design, promising to unify binding-site recognition, pose prediction, protein flexibility and affinity estimation. The poses look chemically reasonable, they score well on public benchmarks, and now claim accuracy comparable to the best physical simulations without the usual cost and experimental effort. 

But a plausible bound pose is not evidence that a model understands molecular recognition, and it’s important to look beyond the headlines. Studies have shown that model performance can decline as targets move away from the training data. In other cases, models continue placing ligands in near-identical poses even after changes to the protein that should weakened or abolished binding.  

In this article for Drug Discovery World, Nathan Brown (Director of Science) and Matthew Segall (CEO) draw on the “mirage effect” from recent multimodal AI research, where a model confidently describes an image it was never shown. They ask whether co-folding models risk a similar illusion: producing a statistically plausible structure without reasoning about the interactions that actually drive binding.  

The argument isn’t that co-folding should be dismissed, ts outputs should be treated as hypotheses to test rather than answers to trust. 

What you’ll take away 

  • Why the Protein Data Bank reflects what scientists could solve and chose to publish, and what that means for models trained on it 
  • What recent benchmarks show about performance on familiar orthosteric sites compared with allosteric, cryptic and novel targets, where discovery value is highest 
  • Why perturbation testing matters, and what it tells us when a predicted pose barely changes after modifications that should abolish binding 
  • How strong scores on public benchmarks can fail to hold up in blinded, real-world medicinal chemistry projects 
  • The difference between predictive usefulness and mechanistic accuracy, and why a model can be valuable in one mode but unreliable in the other 
  • What better benchmarking should look like, and where co-folding is most likely to earn its place in the computational toolbox today 

Explore other content