The overarching goal of TCGA is to change the practice of cancer nnedicine and improve patient survival through cancer genomics. A key deliverable is to enable access and use of complex multi- dimensional genomic data for downstream studies. We propose to operate a GDAC-A center with the leadership, expertise and infrastructure required to develop an analysis pipeline that will generate pre- defined integrative analyses and interpretations that are tailor-designed for hypothesis-testing by basic, translational and clinical investigators. Our team consists of experts in cancer biology, genomics and bioinformatics with a track record of leadership in TCGA. The analytical tools and pipeline structure are based on our extensive TCGA experiences and designed to optimally achieve its goals. This pipeline will be built using the GenePattern bioinformatic workflow environment - a flexible and modular architecture that is caBIG and caGRID compliant, maintained in the well-established, robust and secure IT infrastructure at the Broad Institute and can be operated 24/7 as a Production Pipeline. Leveraging this well-established resource, we will pursue the following specific aims.
Aim 1. We will define caBIG compliant data format for all input and output files. To further enhance standardization, we propose two additions to the standard data structure defined in the Pilot Project (Levels 1-4). Level 0 will define specific versions of all reference databases used in the analyses and Level 5 will capture disease-level findings that incorporate prior knowledge.
Aim 2. We will design analysis modules to consolidate data from all components of TCGA and to perform integrative analyses. Results will be submitted to DCC in caBIG compliant output files accompanied by human-readable reports containing text summaries, tables and figures in a format understandable to scientists of diverse disciplines, similar to the Results Section of a publication. In addition, we are committed to continuous technical and analytical improvement of the pipeline, particularly in supporting the transition to next-generation sequencing platforms.
Aim 3. We will implement this high-throughput analysis pipeline in an industrial-level production mode with rigorous quality control, leveraging the Broad's infrastructural support and extensive experiences in running and maintaining high-throughput computational pipelines.

Public Health Relevance

This GDAC-A center will deliver, in a high-throughput and reliable manner, integrative analyses of TCGA data to bridge the gap between TCGA data generation and their use in biomedical research and eventual translation into the clinic. This is a key deliverable of TCGA, thus this effort is highly relevant.

National Institute of Health (NIH)
National Cancer Institute (NCI)
Resource-Related Research Projects--Cooperative Agreements (U24)
Project #
Application #
Study Section
Special Emphasis Panel (ZCA1-SRLB-U (O1))
Program Officer
Yang, Liming
Project Start
Project End
Budget Start
Budget End
Support Year
Fiscal Year
Total Cost
Indirect Cost
Broad Institute, Inc.
United States
Zip Code
Cancer Genome Atlas Research Network; Linehan, W Marston; Spellman, Paul T et al. (2016) Comprehensive Molecular Characterization of Papillary Renal-Cell Carcinoma. N Engl J Med 374:135-45
Brenan, Lisa; Andreev, Aleksandr; Cohen, Ofir et al. (2016) Phenotypic Characterization of a Comprehensive Set of MAPK1/ERK2 Missense Mutants. Cell Rep 17:1171-1183
Haradhvala, Nicholas J; Polak, Paz; Stojanov, Petar et al. (2016) Mutational Strand Asymmetries in Cancer Genomes Reveal Mechanisms of DNA Damage and Repair. Cell 164:538-49
Zheng, Siyuan; Cherniack, Andrew D; Dewal, Ninad et al. (2016) Comprehensive Pan-Genomic Characterization of Adrenocortical Carcinoma. Cancer Cell 29:723-36
Black, Joshua C; Zhang, Hailei; Kim, Jaegil et al. (2016) Regulation of Transient Site-specific Copy Gain by MicroRNA. J Biol Chem 291:4862-71
Camargo, M Constanza; Bowlby, Reanne; Chu, Andy et al. (2016) Validation and calibration of next-generation sequencing to identify Epstein-Barr virus-positive gastric cancer in The Cancer Genome Atlas. Gastric Cancer 19:676-81
Campbell, Joshua D; Alexandrov, Anton; Kim, Jaegil et al. (2016) Distinct patterns of somatic genome alterations in lung adenocarcinomas and squamous cell carcinomas. Nat Genet 48:607-16
Ceccarelli, Michele; Barthel, Floris P; Malta, Tathiane M et al. (2016) Molecular Profiling Reveals Biologically Discrete Subsets and Pathways of Progression in Diffuse Glioma. Cell 164:550-63
Kim, Jaegil; Mouw, Kent W; Polak, Paz et al. (2016) Somatic ERCC2 mutations are associated with a distinct genomic signature in urothelial tumors. Nat Genet 48:600-6
Shukla, Sachet A; Rooney, Michael S; Rajasagi, Mohini et al. (2015) Comprehensive analysis of cancer-associated somatic mutations in class I HLA genes. Nat Biotechnol 33:1152-8

Showing the most recent 10 out of 51 publications