AlphaFold DB
AI-Predicted 3D Macromolecular Protein Structures at Scale
DeepMind and EMBL-EBI's revolutionary structural biology database spanning virtually all cataloged proteins in UniProt.
技术架构与背景概述
AlphaFold DB provides open access to over 214 million protein structure predictions generated by the AlphaFold 2 and AlphaFold 3 deep learning systems. It covers the entire human proteome and key model organisms, solving the 50-year-old protein folding challenge and accelerating structural biology, drug target identification, and enzymology.
攻克的核心生物学挑战
Experimental determination of protein 3D structures via X-ray crystallography, Cryo-EM, or NMR is labor-intensive, expensive, and can take years per protein. Millions of known genomic sequences lacked corresponding structural models, obstructing mechanistic insights into disease mutations and drug binding.
核心能力与技术亮点
Proteome-Wide Coverage
Contains atomic-resolution models for over 214 million proteins across 1 million species.
Per-Residue Confidence (pLDDT)
Every amino acid residue is annotated with a local confidence score (0-100), accurately identifying well-folded domains and intrinsically disordered regions (IDRs).
Predicted Aligned Error (PAE)
Visual matrix detailing relative domain orientation confidence, essential for multi-domain proteins.
Direct PDB & mmCIF Downloads
Instant download of coordinate files ready for PyMOL, ChimeraX, and molecular docking workflows.
实战调用命令与代码示例
Retrieve AlphaFold atomic coordinates using standard curl or Python requests:
# Download predicted structure PDB for Human Hemoglobin Beta (P68871)
curl -O https://alphafold.ebi.ac.uk/files/AF-P68871-F1-model_v4.pdb
# Or query API for metadata and PAE JSON
curl -s https://alphafold.ebi.ac.uk/api/prediction/P68871 | jq .输入参数与模型输出数据看板
UniProt ID: P68871 (HBB - Hemoglobin subunit beta) Sequence: VHLTPEEKSAVTALWGKVNVDEVGGEALGRLLVVYPWTQRFFE...
Users search by UniProt ID or paste amino acid sequences to locate precomputed models.
ATOM 1 N VAL A 1 15.234 24.128 10.450 1.00 94.20 N ATOM 2 CA VAL A 1 16.102 25.291 10.120 1.00 95.10 C ATOM 3 C VAL A 1 15.340 26.540 9.780 1.00 93.80 C ATOM 4 O VAL A 1 14.120 26.510 9.650 1.00 92.40 O
数值经归一化处理,直接对应下游生物表型预测。
预测结果生物学解读指南
科研人员如何理解预测数值、判别致病阈值与分子调控机制:
Side-chain orientations and backbone geometry suitable for structure-based drug design and active-site analysis.
Reliable secondary structure topology (alpha-helices and beta-sheets).
Should not be interpreted as an unstructured error; typically reflects biologically flexible or natively unfolded loops in physiological solution.
AlphaFold DB vs. 传统分析工具横向对比
| 对比维度 | AlphaFold DB (This Tool) | 传统方法 |
|---|---|---|
| Accuracy (CASP GDT) | GDT_TS > 90 (Near-experimental) | Homology modeling GDT 60-70 |
| Reliance on Templates | Learns co-evolution without requiring homologous PDB | Fails if no homologous template (>30% identity) exists |
| Database Scale | 214+ Million 3D structures ready in seconds | ~200,000 experimental structures in PDB |
常见问题与专家解答 (FAQ)
Can I use AlphaFold structures for molecular docking?
Yes, regions with pLDDT > 90 are widely used in computational drug screening and virtual docking campaigns with AutoDock Vina, Schrödinger Glide, or OpenEye.
What does low pLDDT (<50) mean in AlphaFold?
A low pLDDT score often indicates an intrinsically disordered protein (IDP) segment that lacks a single fixed 3D conformation in isolation.