The human genome is made up of roughly three billion bases, or letters, of DNA.Credit: Yuichiro Chino/GettyThe human genome is an easy place to get lost.
An artificial-intelligence-generated ‘atlas’ of the human genome, unveiled1 today by Google DeepMind, aims to guide scientists through our biological code.
One of the most common types of variation in the human genome is changes to individual DNA nucleotides, or letters.
The AlphaGenome Atlas charts the effects of nine billion single-letter changes in the human genome — every possible mutation of this kind — using predictions generated by the AlphaGenome AI model2, released last year by DeepMind in London.
To create the AlphaGenome Atlas, DeepMind computed predictions for each of the three possible nucleotide changes for every DNA letter in the human genome — one petabyte’s worth of data.
The human genome is made up of roughly three billion bases, or letters, of DNA.Credit: Yuichiro Chino/Getty
The human genome is an easy place to get lost. Only 2% of its three billion letters encode proteins, and the rest is diabolically hard to decipher. An artificial-intelligence-generated ‘atlas’ of the human genome, unveiled1 today by Google DeepMind, aims to guide scientists through our biological code.
One of the most common types of variation in the human genome is changes to individual DNA nucleotides, or letters. These substitutions contribute to differences between people in factors including disease risk; some rare single-letter changes can directly cause disease.
The AlphaGenome Atlas charts the effects of nine billion single-letter changes in the human genome — every possible mutation of this kind — using predictions generated by the AlphaGenome AI model2, released last year by DeepMind in London. The atlas is freely available for non-commercial use.
The tool could help researchers to draw links between genetic variants and rare, unexplained diseases and uncover hidden mechanisms underlying common illnesses and biological traits, say researchers. It might even reveal some of the rules by which DNA sequences control gene activity.
But it won’t replace experiments or, in the case of diagnosing disease, accounting for differences specific to individuals, says Martin Kircher, a bioinformatician at the Max Delbrück Centre for Molecular Medicine in Berlin. “This is a useful and generous way to scale up access to a strong model.”
Instant access
Since AlphaGenome’s release, around 9,000 researchers have accessed the model’s predictions through an automated programming interface (API), says Dhavi Hariharan, a DeepMind product manager. But doing so requires writing software code — a barrier for some biologists, she says.
To create the AlphaGenome Atlas, DeepMind computed predictions for each of the three possible nucleotide changes for every DNA letter in the human genome — one petabyte’s worth of data. It also captures more than 100 million short insertions or deletions observed in human genomes. The effort was inspired by DeepMind’s AlphaFold database of more than 200 million protein-structure predictions, which has been accessed by millions of users, according to the company.
“If you remove the friction, you also increase the curiosity for people to dive in,” says Žiga Avsec, who leads the AlphaGenome team. “Instant access is something that feels magical.”
The atlas, like AlphaGenome, provides thousands of predictions about the potential effects of a variant, from how it affects the tissues in which a nearby gene is expressed to the shape of folded-up DNA, called chromatin. But a repeated ask from API users, says Hariharan, was simplicity: “Can you give me a single score that’ll help me understand — do I care about this variant, do I dig deeper?”
DeepMind’s AlphaGenome Atlas provides predictions about different potential effects of genetic variants.Credit: Google DeepMind
For the atlas, Avsec’s team developed a metric of a variant’s predicted biological effects, the AlphaGenome Variant Impact (AVI) score. This number reliably discerned disease-causing mutations from harmless changes in a clinical genomics database, DeepMind and academic researchers report in the preprint describing the results released today1. The AVI score and other predictions in the atlas also helped a team at the Broad Institute of MIT and Harvard in Cambridge, Massachusetts, to prioritize a non-coding variant as a possible cause of a severe epilepsy in a person.
Rare-disease researchers have usually relied on less computationally demanding models to interpret such variants because applying models such as AlphaGenome across the entire human genome is unfeasible for most researchers, say Mafalda Dias and Jonathan Frazer, computational biologists at the Centre for Genomic Regulation in Barcelona, Spain, in an e-mail to Nature. “By removing that computational barrier, the Atlas should be a valuable resource.”
DNA motifs
Avsec is especially excited about using the atlas to try to uncover the function of short stretches of DNA, called motifs, that are found throughout the genome. Some help to control the production of messenger RNA (mRNA), which carries instructions for encoding individual proteins. Others attract transcription-factor proteins that regulate the expression of individual genes, but the picture is incomplete.
Using atlas predictions as a guide, the researchers mapped thousands of DNA motifs across the human genome and then inferred the roles that those motifs could have in different cell types — be it activating genes, repressing them or altering DNA accessibility. The work “gives us a searchable dictionary for non-coding DNA”, Julia Zeitlinger, a molecular biologist at the Stowers Institute for Medical Research in Kansas City, Missouri, and preprint co-author, said at a press briefing.