This work aims to identify protein functional determinants and to compare them across the proteome to predict protein function. The approach is predicated on a phylogenomic algorithm, the Evolutionary Trace (ET), that identifies key functional residues in proteins;and on ET Annotation (ETA) algorithms, which extract from ET analysis 3D templates, describing the composition and conformation of key residues involved in binding or in catalysis, and then search in other structures for geometric matches to these 3D templates that suggest a common function. Preliminary data have extensively validated ET, both computationally and through experiments, and ETA has become a useful tool to annotate function on structural genomics proteins. Both methods, however, can still gain in sensitivity, specificity and scalability. To do so we propose in Aim 1, first, to improve the ET identification of key functional residues, by optimizing the selection of the input sequences and by a new measure of residue functional importance, and, second, to refine the selection of 3D templates.
In Aim 2, we propose a new network-based annotation diffusion method to compare all 3D template matches at once and to add in functional information from other sources, such as from proteins without known structure.
Aim 3 is experimental and it will test our predictions through mutations and assays on proteins of direct medical interest including one that controls drug resistance in bacteria and another that is a marker of drug resistance in malaria. In the long term, these results should help to focus protein engineering and drug design to the most functionally and therapeutically relevant parts of a protein, and, most broadly, link the massive and exponentially growing amounts of raw sequence and structure data to biological function and its molecular basis.

Public Health Relevance

Modern biology excels at producing volumes of basic information on the composition of our genes and on the structure of the proteins that they encode. However, much of this potentially useful information lies fallow and does not contribute to our understanding of the basic biology of disease or to the development of new drugs and treatments. The reason is that it remains difficult to know what these new genes actually do, and how they do it. This work develops computational methods to answer both questions. In so doing it should help identify the function of novel protein and help connect them to pathological processes. For example, to test some of our tools and predictions, we will experimentally study two proteins of medical interest, one that orchestrates drug resistance in bacteria, and another that marks drug resistance in malaria.

Agency
National Institute of Health (NIH)
Institute
National Institute of General Medical Sciences (NIGMS)
Type
Research Project (R01)
Project #
5R01GM079656-07
Application #
8537933
Study Section
Special Emphasis Panel (ZRG1-BCMB-B (03))
Program Officer
Wehrle, Janna P
Project Start
2007-04-01
Project End
2015-08-31
Budget Start
2013-09-01
Budget End
2014-08-31
Support Year
7
Fiscal Year
2013
Total Cost
$377,204
Indirect Cost
$136,179
Name
Baylor College of Medicine
Department
Genetics
Type
Schools of Medicine
DUNS #
051113330
City
Houston
State
TX
Country
United States
Zip Code
77030
Chun, Yun Shin; Passot, Guillaume; Yamashita, Suguru et al. (2017) Deleterious Effect of RAS and Evolutionary High-risk TP53 Double Mutation in Colorectal Liver Metastases. Ann Surg :
Gallion, Jonathan; Koire, Amanda; Katsonis, Panagiotis et al. (2017) Predicting phenotype from genotype: Improving accuracy through more robust experimental and computational modeling. Hum Mutat 38:569-580
Wilson, Stephen J; Wilkins, Angela D; Lin, Chih-Hsu et al. (2017) DISCOVERY OF FUNCTIONAL AND DISEASE PATHWAYS BY COMMUNITY DETECTION IN PROTEIN-PROTEIN INTERACTION NETWORKS. Pac Symp Biocomput 22:336-347
Koire, Amanda; Kim, Young Won; Wang, Jarey et al. (2017) Codon-level co-occurrences of germline variants and somatic mutations in cancer are rare but often lead to incorrect variant annotation and underestimated impact prediction. PLoS One 12:e0174766
Katsonis, Panagiotis; Lichtarge, Olivier (2017) Objective assessment of the evolutionary action equation for the fitness effect of missense mutations across CAGI-blinded contests. Hum Mutat 38:1072-1084
Xu, Qifang; Tang, Qingling; Katsonis, Panagiotis et al. (2017) Benchmarking predictions of allostery in liver pyruvate kinase in CAGI4. Hum Mutat 38:1123-1131
Schönegge, Anne-Marie; Gallion, Jonathan; Picard, Louis-Philippe et al. (2017) Evolutionary action and structural basis of the allosteric switch controlling ?2AR functional selectivity. Nat Commun 8:2169
Gallion, Jonathan; Wilkins, Angela D; Lichtarge, Olivier (2017) HUMAN KINASES DISPLAY MUTATIONAL HOTSPOTS AT COGNATE POSITIONS WITHIN CANCER. Pac Symp Biocomput 22:414-425
Cancer Genome Atlas Research Network. Electronic address: wheeler@bcm.edu; Cancer Genome Atlas Research Network (2017) Comprehensive and Integrative Genomic Characterization of Hepatocellular Carcinoma. Cell 169:1327-1341.e23
Lua, Rhonald C; Wilson, Stephen J; Konecki, Daniel M et al. (2016) UET: a database of evolutionarily-predicted functional determinants of protein sequences that cluster as functional sites in protein structures. Nucleic Acids Res 44:D308-12

Showing the most recent 10 out of 60 publications