Loss-of-function (LoF) mutations in human protein-coding genes are known to play a major role in severe diseases such as cystic fibrosis and muscular dystrophy, and have also recently been shown to influence the risk of complex diseases such as type 1diabetes and Crohn's disease. We have recently conducted the largest systematic survey to date of human LoF variants, as part of the 1000 Genomes Project, which has confirmed the value of LoF variants for human disease studies and also identified key chalenges for the detection and interpretation of these variants. We propose to overcome these challenges by constructing robust, accurate tools for the annotation, characterization and high-throughput genotyping of LoF variants. Firstly, we will develop an integrated informatic pipeline (Annotation of LoF Transcripts, ALoFT) for the identification and filtering of all classes of LoF variant, including single nucleotide substitutions (SNPs), insertions and deletions. Secondly, we will exploit data from RNA sequencing experiments and disease mutation databases to create more accurate predictive models of the efects of genetic variants on gene expression and splicing, and of their probability of disease causation. Finally, we will apply the ALoFT pipeline and the predictive models described above to over 30,000 human exomes and genomes sequenced as part of other NIH-funded projects, using the resulting annotation as the basis for a publicly accessible database of validated LoF variants, dbLoF. We will use our functional annotation to develop a weighted association test and apply this to the discovery of novel disease risk variants in these sequenced individuals. In addition, we will use the catalogue of LoF variants identified in these samples to design a custom genotyping array permitting rapid, cost-effective interrogation of the majority of common human LoF variants in human cohorts, allowing the phenotypic effects of these variants to be assessed in separately funded association studies. This study will provide powerful tools for discovering and characterizing natural loss-of-function variants, and for exploring their potential association with human disease risk.

Public Health Relevance

Genetic variants that cause the complete loss of function (LoF) of human protein-coding genes are known to play a major role in severe human disease, but are also highly susceptible to sequencing and annotation artifacts. We propose the development of a suite of analytical tools for the accurate identification and filtering of LoF variants, guided and validated using RNA sequencing data and databases of known severe disease mutations. We will apply these tools to large-scale human genome sequence data, generating a high-quality catalogue of LoF variants in the human population, and guiding the design of a genotyping array for further studies assessing the effects of these variants on human phenotypes and disease risk.

National Institute of Health (NIH)
National Institute of General Medical Sciences (NIGMS)
Research Project (R01)
Project #
Application #
Study Section
Special Emphasis Panel (ZGM1-GDB-7 (CP))
Program Officer
Krasnewich, Donna M
Project Start
Project End
Budget Start
Budget End
Support Year
Fiscal Year
Total Cost
Indirect Cost
Massachusetts General Hospital
United States
Zip Code
Minikel, Eric Vallabh; Vallabh, Sonia M; Lek, Monkol et al. (2016) Quantifying prion disease penetrance using large population control cohorts. Sci Transl Med 8:322ra9
Narasimhan, Vagheesh M; Hunt, Karen A; Mason, Dan et al. (2016) Health and population effects of rare gene knockouts in adult humans with related parents. Science 352:474-7
Lek, Monkol; Karczewski, Konrad J; Minikel, Eric V et al. (2016) Analysis of protein-coding genetic variation in 60,706 humans. Nature 536:285-91
Minikel, Eric Vallabh; MacArthur, Daniel G (2016) Publicly Available Data Provide Evidence against NR1H3 R415Q Causing Multiple Sclerosis. Neuron 92:336-338
Hogarth, Marshall W; Garton, Fleur C; Houweling, Peter J et al. (2016) Analysis of the ACTN3 heterozygous genotype suggests that α-actinin-3 controls sarcomeric composition and muscle function in a dose-dependent fashion. Hum Mol Genet 25:866-77
Zou, James; Valiant, Gregory; Valiant, Paul et al. (2016) Quantifying unobserved protein-coding variants in human populations provides a roadmap for large-scale sequencing projects. Nat Commun 7:13293
GTEx Consortium (2015) Human genomics. The Genotype-Tissue Expression (GTEx) pilot analysis: multitissue gene regulation in humans. Science 348:648-60
Rivas, Manuel A; Pirinen, Matti; Conrad, Donald F et al. (2015) Human genomics. Effect of predicted protein-truncating genetic variants on the human transcriptome. Science 348:666-9
Sergouniotis, Panagiotis I; Chakarova, Christina; Murphy, Cian et al. (2014) Biallelic variants in TTLL5, encoding a tubulin glutamylase, cause retinal dystrophy. Am J Hum Genet 94:760-9
Lappalainen, Tuuli; Sammeth, Michael; Friedländer, Marc R et al. (2013) Transcriptome and genome sequencing uncovers functional variation in humans. Nature 501:506-11