The aim of the GENCODE consortium is to annotate all evidence-based gene features in the human genome at a high accuracy, including protein-coding loci with alternatively splices variants, non-coding loci and pseudogenes. With this proposal we aim to extend GENCODE to the mouse genome and use the comparison of corresponding human and mouse loci to improve both sets of annotation. Despite the tremendous progress of current GENCODE production project, and the current outstanding quality of at least the protein-coding gene set, a complete annotation of all human genes is far from complete. For example, it has recently become clear that the number of non-coding RNA genes is far greater than previously supposed. It is also recognized that there are still substantial numbers of alternative transcripts still to be discovered from transcriptomics studies of additional cell types.
Our first aim i s therefore to continue to improve the coverage and accuracy of the GENCODE human gene set.
Our second aim i s to apply to the mouse genome the same annotation approaches as we have applied to human to generate the human GENCODE gene set. To achieve both goals we will integrate computational approaches, expert manual annotation and targeted experimental approaches as we have done for human. We will also use comparative approaches to use the resulting mouse annotation to inform and improve the human GENCODE gene set. A comprehensive knowledge of the location and structure of genes in the human genome is central to our understanding of human biology and the mechanisms of disease. Similarly for mouse, a comprehensive high quality gene set will aid in the design of experiments and the interpretation of the effects of gene knockouts and resulting phenotypes. Also, since mouse is used as a model of human, knowledge of its genes and their relationship to human genes will help inform human gene function. The outputs of regular releases of GENCODE gene sets will therefore be of benefit to the entire community of human and mouse researchers.

Public Health Relevance

A comprehensive knowledge of the location and structure of genes in the human genome is central to our understanding of human biology and the mechanisms of disease. Since mouse is used as a model of human, knowledge of its genes and their relationship to human genes also helps inform human gene function.

Agency
National Institute of Health (NIH)
Institute
National Human Genome Research Institute (NHGRI)
Type
Biotechnology Resource Cooperative Agreements (U41)
Project #
1U41HG007234-01
Application #
8503762
Study Section
Special Emphasis Panel (ZHG1-HGR-M (J2))
Program Officer
Feingold, Elise A
Project Start
2013-04-01
Project End
2017-03-31
Budget Start
2013-04-01
Budget End
2014-03-31
Support Year
1
Fiscal Year
2013
Total Cost
$2,582,206
Indirect Cost
$172,682
Name
Sanger Institute
Department
Type
DUNS #
346013253
City
Cambridge
State
Country
United Kingdom
Zip Code
Pujar, Shashikant; O'Leary, Nuala A; Farrell, Catherine M et al. (2018) Consensus coding sequence (CCDS) database: a standardized set of human and mouse protein-coding regions supported by expert curation. Nucleic Acids Res 46:D221-D228
Newman, Victoria; Moore, Benjamin; Sparrow, Helen et al. (2018) The Ensembl Genome Browser: Strategies for Accessing Eukaryotic Genome Data. Methods Mol Biol 1757:115-139
Wang, Junling; Pejaver, Vikas Rao; Dann, Geoffrey P et al. (2018) Target site specificity and in vivo complexity of the mammalian arginylome. Sci Rep 8:16177
Zerbino, Daniel R; Achuthan, Premanand; Akanni, Wasiu et al. (2018) Ensembl 2018. Nucleic Acids Res 46:D754-D761
Casper, Jonathan; Zweig, Ann S; Villarreal, Chris et al. (2018) The UCSC Genome Browser database: 2018 update. Nucleic Acids Res 46:D762-D769
Rodriguez, Jose Manuel; Rodriguez-Rivas, Juan; Di Domenico, Tomás et al. (2018) APPRIS 2017: principal isoforms for multiple gene sets. Nucleic Acids Res 46:D213-D217
Loughran, Gary; Jungreis, Irwin; Tzani, Ioanna et al. (2018) Stop codon readthrough generates a C-terminally extended variant of the human vitamin D receptor with reduced calcitriol response. J Biol Chem 293:4434-4444
Tardaguila, Manuel; de la Fuente, Lorena; Marti, Cristina et al. (2018) SQANTI: extensive characterization of long-read transcript sequences for quality control in full-length transcriptome identification and quantification. Genome Res :
Jain, Miten; Koren, Sergey; Miga, Karen H et al. (2018) Nanopore sequencing and assembly of a human genome with ultra-long reads. Nat Biotechnol 36:338-345
Kolmogorov, Mikhail; Armstrong, Joel; Raney, Brian J et al. (2018) Chromosome assembly of large and complex genomes using multiple references. Genome Res 28:1720-1732

Showing the most recent 10 out of 88 publications