1. Biomolecular Structure Prediction Tools

This page of BioMoDes lists state-of-the-art and emerging tools for Biomolecular Structure Prediction.
1.1. Protein (Monomer) Structure Prediction
2024 (Click to collapse/expand)
-
OpenFold: A fully open-source and trainable PyTorch reimplementation of AF2 with training code and data. OpenFold, trained from scratch, matched AF2 in accuracy and introduced some technical modifications that offered improved speed and memory efficiency over AF2.
Paper published: May 14, 2024
Paper | Code (GitHub) | Data (Open Data) | Documentation | Colab Notebook -
AlphaFold 3: The latest version of the AlphaFold model that is capable of predicting, with high accuracy, the structures of complexes containing nearly all molecular types present in the PDB (protein, DNA, RNA, small-molecule ligands, ions, and modified residues).
Paper published: April 29, 2024
Paper 1 | Webserver -
RaptorX-Single: A single sequence protein structure prediction method that integrates protein language model and a Evoformer-based structure generation module. RaptorX-Single outperforms MSA-based (AF2) and MSA-free methods in predicting structures of antibodies (fine-tuned), structures of orphan proteins, and effects of single mutations. RaptorX-Single also runs faster than MSA-based AF2.
Paper published: March 20, 2024
Paper Code (GitHub) -
Evo: A long-context foundation model that generalizes across the central dogma of biology: DNA, RNA, and proteins. Evo is a 7 billion parameter model trained to generate DNA sequences and is capable of prediction and generative tasks, from molecules to whole genomes.
Preprint posted: March 06, 2024
Preprint | Code (GitHub) | Code (PyPI) | Blog | Playground | Colab Notebook -
AlphaFind: A web-based search engine for finding structures similar to a given query in the entire proteome. AlphaFind queries the AlphaFold DB to find similar structures for any given input.
Preprint posted Feb 18, 2024
Preprint | Code (GitHub) | Wiki | Webserver -
ALBATROSS: A tool and webserver that combines sequence design, large-scale coarse-grained simulations and deep learning for the prediction of conformational properties (ensemble) of intrinsically disordered proteins.
Paper published: Jan 31, 2024
Paper 1 | Paper 2 | Preprint | Colab Notebooks | Code (GitHub) | Webserver -
DMFold: A tool that integrates large genomic and metagenomics sequence databases for improved protein structure prediction.
Paper published: Jan 02, 2024
Paper | Code (DMFold) | Code (DeepMSA2) | Webserver (DMFold) | Webserver (DeepMSA2)
2023
-
Chroma: A diffusion model for protein design developed by Generate:Biomedicines. Chroma is a programmable generative model that can directly sample novel protein structures and sequences. Some of its capabilities include: complexes, symmetries, substructures, shapes, “semantics” and even natural-language prompts.
Paper published: Nov 15, 2023
Paper | Preprint | Code (GitHub) | Colab Notebooks -
AFCluster: An AlphaFold2-based method to predict multiple biologically-relevant conformations of protein structures.
Paper published: Nov 13, 2023
Paper | Code (GitHub) | Colab Notebook
Note: There is a preprint challenging the claims in the AF-Cluster paper. You can read it here. -
ESMFold: A method to predict atomic-level protein structure from sequence using a large protein language model. ESMFold is nearly as accurate as alignment-based methods and considerably faster, enabling the construction of the ESM Metagenomic Atlas, a database containing more than 617 million metagenomic protein sequences with one-third being high confidence predictions.
Paper published: March 16, 2023
Paper | Preprint | Code (GitHub) | Webserver | ESM Metagenomic Atlas
1.2. Protein Complex Structure Prediction
2026 (Click to collapse/expand)
-
ESMFold2: A structure-prediction head from Biohub (Alex Rives’ team) built on top of ESMC, a protein language “world model” trained on 2.8 billion protein sequences. It predicts high-resolution, all-atom 3D structures for proteins, protein-protein/antibody-antigen complexes, and other biomolecules (small molecules, DNA, RNA, modified amino acids) directly from sequence, with optional MSA input for challenging targets, and can run in fast single-sequence mode. ESMFold2 matches or exceeds AlphaFold3 across diverse evaluation datasets and leads on DockQ pass-rate for protein-protein and antibody-antigen complexes on FoldBench; the release also includes ESM Atlas, a database of over 1 billion predicted structures.
Preprint posted: May 27, 2026
Paper | Code (GitHub) | Model (HuggingFace) -
OpenDDE: An open-source, all-atom foundation model for biomolecular structure prediction from Aureka Research, intended as the structural core of a broader AI drug-discovery platform. It expands residue tokens into seven categories of structural tokens (backbone, side chain, nucleic-acid backbone/base, ligand/atom) refined via pair-conditioned attention, adds a differentiable shape-complementarity objective for interface geometry, and unifies structure prediction and conditional design as variants of the same diffusion denoising problem. OpenDDE posts the highest reported ranked DockQ success rate among tested models on three antibody-antigen benchmarks (PXMeter-AB, FoldBench-AB, 2026ARK-AB), outperforming AlphaFold3 and ESMFold2, though the large gap to oracle (best-of-sample) performance exposes confidence-ranking as the main current bottleneck; design capability is architecturally supported but not yet benchmarked.
Preprint posted: July 20, 2026
Technical Report | Code (GitHub) | Model (HuggingFace) | Homepage
2025 (Click to collapse/expand)
-
Boltz-2: A next-generation biomolecular foundation model from MIT/Valence Labs/Recursion that extends Boltz-1 with binding affinity prediction alongside structure prediction, adding an affinity module, improved controllability, GPU optimizations, and training on a large collection of synthetic and molecular dynamics data. On the FEP+ (OpenFE) affinity benchmark, Boltz-2 achieves a Pearson correlation of 0.62 — comparable to physics-based FEP — while running over 1,000x faster, and it outperformed all submitted methods in the CASP16 affinity challenge across 140 complexes.
Preprint posted: June 18, 2025
Preprint | Code (GitHub) | Webserver -
Protenix: An open-source, trainable PyTorch reproduction of AlphaFold3 from ByteDance, built for high-accuracy prediction of protein, ligand, and nucleic-acid complex structures, with a modular framework supporting full training and inference (custom CUDA kernels, BF16 training). Across benchmarks including PoseBusters V2, low-homology PDB sets, and CASP15 RNA, Protenix achieves state-of-the-art performance in protein-ligand, protein-protein, and protein-nucleic-acid structure prediction.
Preprint posted: January 11, 2025
Preprint | Code (GitHub)
2024
-
AlphaPulldown2: A major upgrade of the AlphaPulldown pipeline for high-throughput protein-protein interaction screening and complex structural modeling, from the Kosinski and Schwede labs. It is now orchestrated via Snakemake for full workflow automation and checkpointed resumption, adds modular backend support beyond AlphaFold2-Multimer (UniFold, AlphaLink2, extensible to OpenFold), integration of cross-linking mass-spec (XL-MS) restraints via AlphaLink2, LZMA2-compressed feature storage (97.7% size reduction — 55 GB vs. 2.4 TB for 20,581 human proteins), and ModelCIF-formatted outputs for FAIR compliance, plus a public repository of precomputed MSA/template features for 14 model-organism proteomes.
Preprint posted: November 28, 2024
Paper published: March 14, 2025
Paper | Preprint | Code (GitHub) -
Boltz-1: The first fully open-source biomolecular structure prediction model to match AlphaFold3-level performance, from the MIT Jameel Clinic. Boltz-1 extends and reproduces the AlphaFold3 technical report — architecture, data curation, training, and inference — and performs comparably to AlphaFold3 and Chai-1 on benchmarks including CASP15. An inference-time technique, Boltz-steering (Boltz-1x), resolves physical/steric issues in predicted structures while maintaining accuracy.
Preprint posted: November 20, 2024
Preprint | Code (GitHub) -
Chai-1: A multi-modal foundation model from Chai Discovery for biomolecular structure prediction, enabling unified prediction of proteins, small molecules, DNA, RNA, and glycosylation at state-of-the-art accuracy across a variety of benchmarks. By default the model generates five sample predictions using embeddings without MSAs or templates, with an option for automatic MSA generation via the ColabFold MMseqs2 server.
Preprint posted: October 11, 2024
Preprint | Code (GitHub) -
Umol: A deep learning method for all-atom protein-ligand complex structure prediction from protein sequence and ligand SMILES string.
Paper published: May 28, 2024
Paper | Code (GitHub) | Colab Notebook -
OpenFold: A fully open-source and trainable PyTorch reimplementation of AF2 with training code and data. OpenFold, trained from scratch, matched AF2 in accuracy and introduced some technical modifications that offered improved speed and memory efficiency over AF2.
Paper published: May 14, 2024
Paper | Preprint | Code (GitHub) | Data (Open Data) | Documentation | Colab Notebook -
AlphaFold 3: The latest version of the AlphaFold model that is capable of predicting, with high accuracy, the structures of complexes containing nearly all molecular types present in the PDB (protein, DNA, RNA, small-molecule ligands, ions, and modified residues).
Paper published: April 29, 2024
Paper 1 | Webserver -
FABind+: An improved version of FABind, for the prediction of protein-ligand binding based on pocket prediction and docking.
Preprint posted: March 29, 2024
Preprint | Code (GitHub) -
RosettaFold All-Atom: A network capable of predicting the structures of all atoms of a biological unit, including proteins, nucleic acids, small molecules, metals, covalent modifications (covalently modified proteins). In other words, RF-AA can generate accurate models for complexes of proteins with other protein and non-protein molecules. It’s basically a “predict-every-(bio)molecule” tool. RF-AA also provides error estimates of its predictions.
Paper published: March 07, 2024
Paper | Code (GitHub) -
Evo: A long-context foundation model that generalizes across the central dogma of biology: DNA, RNA, and proteins. Evo is a 7 billion parameter model trained to generate DNA sequences and is capable of prediction and generative tasks, from molecules to whole genomes.
Preprint posted: March 06, 2024
Preprint | Code (GitHub) | Code (PyPI) | Blog | Playground | Colab Notebook -
GlycoSHIELD, GlycoALPHAFOLD, GlycoTRAJ, GlycoSASA,…: A set of tools and webserver for modeling glycoprotein morphology and structural dynamics.
Paper published: Feb 29, 2024
Paper | Code (GitLab) | Webserver -
PocketGen: A method for generating full-atom ligand-binding pockets to design small molecule-binding proteins. PocketGen uses a co-design strategy that, given the ligand and the scaffold, simultaneously designs the sequence and structure of the protein pocket.
Preprint posted: Feb 28, 2024
Preprint | Code (GitHub) | Blog -
DiffDock/DiffDock-L: A diffusion generative model for blind molecular docking.
Preprint posted: Feb 11, 2023 | Preprint posted: 28 Feb 2024
Preprint 1, DiffDock | Preprint 2, DiffDock-L | Code (GitHub) | Demo/DiffDock-Web (HuggingFace) -
NeuralPLexer: A tool for predicting protein–ligand complex structures using protein sequence and ligand molecular graph inputs.
Paper published: Feb 12, 2024
Paper | Code (GitHub) | Code (Code Ocean) -
CombFold: A combinatorial and hierarchical assembly algorithm combined with AlphaFold2 for predicting structures of large protein assemblies.
Paper published: Feb 07, 2024
Paper | Code (GitHub) | Code (Code Ocean) | Colab Notebook -
DynamicBind: A generative model and webserver for predicting ligand-specific protein-ligand complex structure. DynamicBind is a “dynamic docking” method that attempts to overcome the limitations of traditional docking and MD simulation.
Paper published: Feb 05, 2024
Paper | Code (GitHub) | Webserver -
DMFold-Multimer: The protein structure prediction tool that, by integrating large genomic and metagenomics sequence databases, outperformed 86 other methods in the complex modeling section of CASP15 .
Paper published: Jan 02, 2024
Paper | Code (DMFold-Multimer) | Code (DeepMSA2-Multimer) | Webserver (DMFold-Multimer) | Webserver (DeepMSA2-Multimer)
2023
-
FragFold: An AlphaFold2-based method for high-throughput prediction of peptide binding to protein targets. It’s a method that was employed for high-throughput computational discovery of inhibitory protein fragments.
Preprint posted: Dec 20, 2023
Preprint | Code (GitHub) -
Chroma: Chroma, a diffusion model for protein design developed by Generate:Biomedicines. Chroma is a programmable generative model that can directly sample novel protein structures and sequences. Some of its capabilities include: complexes, symmetries, substructures, shapes, “semantics” and even natural-language prompts.
Paper published: Nov 15, 2023
Paper | Preprint | Code (GitHub) | Colab Notebooks -
RoseTTAFold2: A redesigned, more efficient version of the three-track RoseTTAFold architecture that incorporates frame-aligned point error, additional recycling during training, and distillation from AlphaFold2 predictions. RoseTTAFold2 reaches AlphaFold2-level accuracy on monomers and AlphaFold2-Multimer-level accuracy on complexes, with better computational scaling to large proteins and assemblies.
Preprint posted: May 25, 2023
Preprint -
LightDock: A protein-protein, protein-peptide, and protein-nucleic acid flexible docking framework based on the Glowworm Swarm Optimization (GSO) algorithm. LightDock is not just a protocol but a framework that accepts multiple user-selected scoring functions and force-fields.
Paper published: May 04, 2023
Paper | Code (GitHub) - Home | Code (GitHub) - Python Implementation | Code (GitHub) - Rust Implementation | Webserver | Homepage
2022
-
AlphaFill: A tool that, using sequence and structure similarity, FILLS the gap in the AF2 protein structure database of all known protein sequences. AlphaFill populates AF2 structure models with relevant small-molecule ligands and cofactors by, essentially, transplanting ligands found in experimentally determined structures of homologous proteins.
Paper published: Nov 24, 2022
Paper | Code (GitHub) | Webserver -
AlphaFold2-Multimer: AlphaFold2 retrained for the prediction of protein-protein complexes.
Preprint posted: March 10, 2022
Preprint | Code (GitHub)
2021
-
AlphaFold2: A deep learning system that predicts a protein’s 3D structure from its amino acid sequence with near-experimental accuracy, using an attention-based Evoformer architecture over multiple sequence alignments and structural templates. AlphaFold2’s CASP14 performance was widely recognized as a landmark solution to the protein structure prediction problem and underlies the AlphaFold Protein Structure Database.
Paper published: July 15, 2021
Paper | Code (GitHub) | Colab Notebook
1.3 Antibody Structure Prediction
2024
-
GeoAB: A method for computational design and optimization (affinity maturation) of antibody. GeoAB utilizes a co-design strategy, predicting the structure of a CDR and optimized 1D sequences for structure.
Preprint posted: May 17, 2024
Preprint | Code (GitHub) -
H3-OPT: A model for predicting antibody structures based on AlphaFold2 and a pre-trained protein language model.
Preprint posted: Mar 14, 2024
Preprint | Code (GitHub) -
DeepSP: A deep learning method to predict the stability of monoclonal antibodies from sequence.
Preprint posted: Mar 03, 2024
Preprint | Code (GitHub) -
tFold-Ab and tFold-Ag: Methods for antibody and antibody-antigen complex modelling and design by Tencent.
Preprint posted: Feb 08, 2024
Preprint | Code (GitHub)
1.4. RNA Structure Prediction
2024
- RhoFold+: An end-to-end, fully automated deep learning pipeline for predicting 3D structures of single-chain RNAs directly from sequence. It integrates RNA-FM, an RNA language model pretrained on ~23.7 million non-coding RNA sequences, with an AlphaFold-style structure module adapted for RNA, and also predicts secondary structure and inter-helical angles as empirically checkable intermediate outputs. In retrospective evaluation on RNA-Puzzles and CASP15 natural RNA targets, RhoFold+ outperformed existing computational methods and matched or exceeded human expert prediction groups.
Paper published: November 21, 2024
Paper | Code (GitHub) | Webserver - AutoRNA: A method for the prediction of RNA 3D structure using VAE.
Preprint posted: June 27, 2024
Preprint | Code (GitHub) - RNADiffFold: A generative model for RNA secondary structure prediction that leverages neural networks from RNA-FM and UFold for feature extraction.
Preprint posted: June 02, 2024
Preprint | Code (GitHub) - Evo: A long-context foundation model that generalizes across the central dogma of biology: DNA, RNA, and proteins. Evo is a 7 billion parameter model trained to generate DNA sequences and is capable of prediction and generative tasks, from molecules to whole genomes.
Preprint posted: March 06, 2024
Preprint | Code (GitHub) | Code (PyPI) | Blog | Playground | Colab Notebook
1.5. Protein Conformational Ensemble Prediction
2024 (Click to collapse/expand)
-
P2DFlow: An SE(3) flow-matching generative model for predicting protein conformational ensembles, trained on MD simulation data (ATLAS dataset). It introduces a physically-motivated prior for the flow process and an extra dimension encoding ensemble-membership/intermediate-state identity, guiding generation toward physically realistic ensemble distributions rather than a single static structure. P2DFlow outperforms baseline generative ensemble methods AlphaFlow and STR2STR across accuracy, diversity, and dynamics-fidelity metrics, correctly capturing fluctuations observed in both crystal structures and MD simulations of held-out test proteins.
Preprint posted: November 26, 2024
Preprint | Paper | Code (GitHub) -
AlphaFlow/ESMFlow: Fine-tuned versions of AlphaFold/ESMFold and retrained on MD simulation ensembles. AlphaFlow generates protein conformational ensembles, including experimental ensembles (as in the PDB), and molecular dynamics simulation ensembles.
Preprint posted Feb 7, 2024
Preprint | Code (GitHub) -
ALBATROSS: A tool and webserver that combines sequence design, large-scale coarse-grained simulations and deep learning for the prediction of conformational properties (ensemble) of intrinsically disordered proteins.
Paper published: Jan 31, 2024
Paper 1 | Paper 2 | Preprint | Colab Notebooks | Code (GitHub) | Webserver
I try my best to make the information on this website as accurate as possible.
If you find any errors in the contents of this page or any other page on this website,
I would greatly appreciate that you kindly get in touch with me at
contact[at]abeebyekeen[dot]com.
If you are interested in joining my free weekly “BioMoDes and Top Reads” newsletter, please subscribe below.