AlphaFold DB
AI-Predicted 3D Macromolecular Protein Structures at Scale
DeepMind and EMBL-EBI's revolutionary structural biology database spanning virtually all cataloged proteins in UniProt.
Technical Overview & Biological Significance
AlphaFold DB provides open access to over 214 million protein structure predictions generated by the AlphaFold 2 and AlphaFold 3 deep learning systems. It covers the entire human proteome and key model organisms, solving the 50-year-old protein folding challenge and accelerating structural biology, drug target identification, and enzymology.
The Biological Challenge Solved
Experimental determination of protein 3D structures via X-ray crystallography, Cryo-EM, or NMR is labor-intensive, expensive, and can take years per protein. Millions of known genomic sequences lacked corresponding structural models, obstructing mechanistic insights into disease mutations and drug binding.
Key Computational Capabilities
Proteome-Wide Coverage
Contains atomic-resolution models for over 214 million proteins across 1 million species.
Per-Residue Confidence (pLDDT)
Every amino acid residue is annotated with a local confidence score (0-100), accurately identifying well-folded domains and intrinsically disordered regions (IDRs).
Predicted Aligned Error (PAE)
Visual matrix detailing relative domain orientation confidence, essential for multi-domain proteins.
Direct PDB & mmCIF Downloads
Instant download of coordinate files ready for PyMOL, ChimeraX, and molecular docking workflows.
Executable Commands & API Usage
Retrieve AlphaFold atomic coordinates using standard curl or Python requests:
# Download predicted structure PDB for Human Hemoglobin Beta (P68871)
curl -O https://alphafold.ebi.ac.uk/files/AF-P68871-F1-model_v4.pdb
# Or query API for metadata and PAE JSON
curl -s https://alphafold.ebi.ac.uk/api/prediction/P68871 | jq .Input Specifications & Inference Results
UniProt ID: P68871 (HBB - Hemoglobin subunit beta) Sequence: VHLTPEEKSAVTALWGKVNVDEVGGEALGRLLVVYPWTQRFFE...
Users search by UniProt ID or paste amino acid sequences to locate precomputed models.
ATOM 1 N VAL A 1 15.234 24.128 10.450 1.00 94.20 N ATOM 2 CA VAL A 1 16.102 25.291 10.120 1.00 95.10 C ATOM 3 C VAL A 1 15.340 26.540 9.780 1.00 93.80 C ATOM 4 O VAL A 1 14.120 26.510 9.650 1.00 92.40 O
Calibrated quantitative scores ready for downstream analysis.
Biological Result Interpretation Guide
How researchers interpret confidence thresholds, fold-change values, and functional impact:
Side-chain orientations and backbone geometry suitable for structure-based drug design and active-site analysis.
Reliable secondary structure topology (alpha-helices and beta-sheets).
Should not be interpreted as an unstructured error; typically reflects biologically flexible or natively unfolded loops in physiological solution.
AlphaFold DB vs. Traditional Computational Methods
| Dimension | AlphaFold DB (This Tool) | Traditional Pipelines |
|---|---|---|
| Accuracy (CASP GDT) | GDT_TS > 90 (Near-experimental) | Homology modeling GDT 60-70 |
| Reliance on Templates | Learns co-evolution without requiring homologous PDB | Fails if no homologous template (>30% identity) exists |
| Database Scale | 214+ Million 3D structures ready in seconds | ~200,000 experimental structures in PDB |
Frequently Asked Questions (FAQ)
Can I use AlphaFold structures for molecular docking?
Yes, regions with pLDDT > 90 are widely used in computational drug screening and virtual docking campaigns with AutoDock Vina, Schrödinger Glide, or OpenEye.
What does low pLDDT (<50) mean in AlphaFold?
A low pLDDT score often indicates an intrinsically disordered protein (IDP) segment that lacks a single fixed 3D conformation in isolation.