The Data Compilation Core (Core B) will develop and maintain a central resource of analysis-ready, annotated and documented data sets from clinical trials and related studies to be utilized by the investigators of the program. These data sets will be used to evaluate the methods developed in this program as well as to demonstrate the software developed in the Computational Resource Core (Core C). The primary source of the data will be the clinical trials and related studies of the Cancer and Leukemia Group B (CALGB), one of the major NCI-sponsored cancer cooperative groups. In addition, data from cancer research studies conducted at two large NCI-designated Comprehensive Cancer Centers, the Lineberger Comprehensive Cancer Center at UNC and the Duke Comprehensive Cancer Center, will also be utilized. This is a major advantage for the program in that the data sets provided can be exceptionally well annotated and documented, with the direct involvement of clinical and statistical scientists who were involved in the primary design and analysis of the studies.

Public Health Relevance

A major disadvantage of using public data sets is that the investigator is often unable to understand the clinical and molecular data as the data are provided without appropriate documentation. Indeed, it is not possible to carry out a thorough statistical analysis of data from clinical trials without taking into account and understanding the design of the study, the specifics of the data collection process, the history of the study and the medical issues. This core will address these issues by providing analysis-ready data sets with extensive annotation and documentation.

National Institute of Health (NIH)
National Cancer Institute (NCI)
Research Program Projects (P01)
Project #
Application #
Study Section
Special Emphasis Panel (ZCA1)
Project Start
Project End
Budget Start
Budget End
Support Year
Fiscal Year
Total Cost
Indirect Cost
University of North Carolina Chapel Hill
Chapel Hill
United States
Zip Code
Acharya, Chaitanya R; McCarthy, Janice M; Owzar, Kouros et al. (2016) Exploiting expression patterns across multiple tissues to map expression quantitative trait loci. BMC Bioinformatics 17:257
Laber, Eric B; Zhao, Ying-Qi; Regh, Todd et al. (2016) Using pilot data to size a two-arm randomized trial to find a nearly optimal personalized treatment strategy. Stat Med 35:1245-56
Li, Zhiguo; Owzar, Kouros (2016) Fitting Cox Models with Doubly Censored Data Using Spline-Based Sieve Marginal Likelihood. Scand Stat Theory Appl 43:476-486
Wang, Xiaofei; Berry, Mark F (2016) Risk calculators are useful but.... J Thorac Cardiovasc Surg 151:706-7
Wang, Xuefeng; Chen, Mengjie; Yu, Xiaoqing et al. (2016) Global copy number profiling of cancer genomes. Bioinformatics 32:926-8
Ivanova, Anastasia; Wang, Yunfei; Foster, Matthew C (2016) The rapid enrollment design for Phase I clinical trials. Stat Med 35:2516-24
Zhang, Daowen; Sun, Jie Lena; Pieper, Karen (2016) Bivariate Mixed Effects Analysis of Clustered Data with Large Cluster Sizes. Stat Biosci 8:220-233
Schifano, Elizabeth D; Wu, Jing; Wang, Chun et al. (2016) Online Updating of Statistical Inference in the Big Data Setting. Technometrics 58:393-403
Lizotte, Daniel J; Laber, Eric B (2016) Multi-Objective Markov Decision Processes for Data-Driven Decision Support. J Mach Learn Res 17:
Minsker, Stanislav; Zhao, Ying-Qi; Cheng, Guang (2016) Active Clinical Trials for Personalized Medicine. J Am Stat Assoc 111:875-887

Showing the most recent 10 out of 378 publications