AlphaGenome
Unified Deep Learning for Non-Coding Variant Effects & Epigenetic Predictions
Google DeepMind's frontier AI model mapping genetic variations directly to molecular phenotypes across cell types and tissues.
Technical Overview & Biological Significance
AlphaGenome is a breakthrough deep learning model from Google DeepMind engineered to decode the 98% non-coding genome. While traditional variant effect predictors rely on evolutionary conservation or hand-engineered biochemical annotations, AlphaGenome models regulatory DNA sequence-to-function end-to-end. Given a raw DNA sequence or a genomic variant (SNV/Indel), it predicts changes in gene expression (RNA-seq), chromatin accessibility (DNase/ATAC-seq), histone modifications (ChIP-seq), and transcription factor binding profiles across hundreds of human cell lines and tissues.
The Biological Challenge Solved
Over 90% of disease-associated variants identified by Genome-Wide Association Studies (GWAS) reside in non-coding regulatory elements (promoters, enhancers, silencers). Discerning true causal variants from passenger linkage disequilibrium (LD) blocks and predicting their quantitative effect on downstream gene expression has long been a key bottleneck in translational genomics.
Key Computational Capabilities
Single-Nucleotide Sensitivity in 1Mb Context
Processes up to 1 million base pairs of surrounding genomic sequence, capturing long-range enhancer-promoter looping interactions that dictate gene transcription.
Multi-Modal Epigenetic Predictions
Simultaneously predicts RNA-seq abundance, DNase-I hypersensitive sites, and active histone marks (H3K27ac, H3K4me3) in a unified tensor representation.
Tissue & Cell-Type Specificity
Distinguishes cell-type specific regulatory wiring, such as hepatocytes vs. T-cells, enabling precise interpretation of tissue-restricted autoimmune or metabolic traits.
Saturation Mutagenesis Scanning
Capable of in silico in-silico saturation mutagenesis (ISM), scanning every single base in a locus to build predictive regulatory maps.
Executable Commands & API Usage
Query AlphaGenome using 1-based genomic coordinates (chr:pos:ref>alt) on GRCh38 to compute tissue-specific variant disruption:
import requests
import json
# Define the target genomic variant (e.g. non-coding regulatory SNP)
# Coordinate format: chr:pos:ref>alt in GRCh38
variant_id = "chr7:117559590:A>G"
payload = {
"variant": variant_id,
"genome_build": "GRCh38",
"target_assays": ["RNA-seq", "DNASE", "H3K27ac"],
"tissues": ["lung_epithelium", "whole_blood", "liver"],
"window_size_bp": 100000
}
response = requests.post(
"https://api.alphagenome.org/v1/predict_variant",
headers={"Authorization": "Bearer YOUR_RESEARCH_API_KEY"},
json=payload
)
predictions = response.json()
print("Predicted expression log2FC:", predictions["gene_effects"][0]["log2_fold_change"])Input Specifications & Inference Results
chr7:117559590:A>G Reference Allele: A Alternate Allele: G Locus: CFTR upstream enhancer region
Input specifies the chromosome, nucleotide position, reference base, and observed alternate allele.
{
"variant": "chr7:117559590:A>G",
"summary_score": 0.884,
"significance_tier": "High Regulatory Impact",
"gene_effects": [
{
"gene_symbol": "CFTR",
"ensembl_id": "ENSG00000001626",
"tissue": "lung_epithelium",
"log2_fold_change": -1.42,
"p_value": 0.00018,
"motif_disrupted": "FOXA1_pioneer_factor"
},
{
"gene_symbol": "CFTR",
"tissue": "whole_blood",
"log2_fold_change": -0.04,
"p_value": 0.62,
"motif_disrupted": "none"
}
],
"epigenetics": {
"chromatin_accessibility_delta": -0.68,
"H3K27ac_signal_delta": -0.81
}
}Calibrated quantitative scores ready for downstream analysis.
Biological Result Interpretation Guide
How researchers interpret confidence thresholds, fold-change values, and functional impact:
Indicates the alternate allele disrupts a critical positive transcription factor binding motif in the enhancer, drastically attenuating target gene transcription.
The variant transitions the chromatin state from open (accessible to RNA Pol II and pioneer factors) to nucleosome-dense/closed chromatin.
Because FOXA1 pioneer activity is lung-epithelial specific, the variant exerts zero effect in whole blood cells, explaining why patient blood transcriptomes appear normal.
AlphaGenome vs. Traditional Computational Methods
| Dimension | AlphaGenome (This Tool) | Traditional Pipelines |
|---|---|---|
| Input Modality | Raw DNA sequence (up to 1Mb context) | Conservation scores (PhyloP, PhastCons) + Overlap beds |
| Expression Effect | Predicts directional log2 fold change per tissue | Binary score or rank (pathogenic / benign) |
| Long-Range Enhancers | Fully modeled via attention transformer | Proximal promoter biases, blind to distal loops |
| Cell-type Specificity | Hundreds of specific human tissues/cells | Static pan-human metric (e.g. CADD score) |
Frequently Asked Questions (FAQ)
What is AlphaGenome and how does it differ from AlphaFold?
While AlphaFold predicts the 3D atomic structure of folded proteins from amino acid sequences, AlphaGenome operates on DNA regulatory sequences to predict gene transcription, epigenetic signals, and non-coding variant consequences.
What genomic coordinate format does AlphaGenome require?
AlphaGenome standardly uses 1-based GRCh38 coordinates in the format `chr[N]:[pos]:[ref]>[alt]`. It also accepts raw FASTA sequences for synthetic promoter and enhancer engineering.
Can AlphaGenome prioritize non-coding GWAS hits?
Yes, fine-mapping GWAS loci is a primary use case. By scoring all variants within an LD block, AlphaGenome highlights which specific nucleotide substitution causes functional transcriptional disruption.