PubMed Health⌕ Search

Biomedical subjects

Michelle L Green

Publications and source records attributed to Michelle L Green.

3 recordsLinked to original sources

Computational prediction of human metabolic pathways from the complete human genome.

BACKGROUND: We present a computational pathway analysis of the human genome that assigns enzymes encoded therein to predicted metabolic pathways. Pathway assignments place genes in their larger biological context, and are a necessary first step toward quantitative modeling of metabolism. RESULTS: Our analysis assigns 2,709 human enzymes to 896 bioreactions; 622 of the enzymes are assigned roles in 135 predicted metabolic pathways. The predicted pathways closely match the known nutritional requirements of humans. This analysis identifies probable omissions in the human genome annotation in the form of 203 pathway holes (missing enzymes within the predicted pathways). We have identified putative genes to fill 25 of these holes. The predicted human metabolic map is described by a Pathway/Genome Database called HumanCyc, which is available at http://HumanCyc.org/. We describe the generation of HumanCyc, and present an analysis of the human metabolic map. For example, we compare the predicted human metabolic pathway complement to the pathways of Escherichia coli and Arabidopsis thaliana and identify 35 pathways that are shared among all three organisms. CONCLUSIONS: Our analysis elucidates a significant portion of the human metabolic map, and also indicates probable unidentified genes in the genome. HumanCyc provides a genome-based view of human nutrition that associates the essential dietary requirements of humans with a set of metabolic pathways whose existence is supported by the human genome. The database places many human genes in a pathway context, thereby facilitating analysis of gene expression, proteomics, and metabolomics datasets through a publicly available online tool called the Omics Viewer.

Arabidopsis↗

A Bayesian method for identifying missing enzymes in predicted metabolic pathway databases.

BACKGROUND: The PathoLogic program constructs Pathway/Genome databases by using a genome's annotation to predict the set of metabolic pathways present in an organism. PathoLogic determines the set of reactions composing those pathways from the enzymes annotated in the organism's genome. Most annotation efforts fail to assign function to 40-60% of sequences. In addition, large numbers of sequences may have non-specific annotations (e.g., thiolase family protein). Pathway holes occur when a genome appears to lack the enzymes needed to catalyze reactions in a pathway. If a protein has not been assigned a specific function during the annotation process, any reaction catalyzed by that protein will appear as a missing enzyme or pathway hole in a Pathway/Genome database. RESULTS: We have developed a method that efficiently combines homology and pathway-based evidence to identify candidates for filling pathway holes in Pathway/Genome databases. Our program not only identifies potential candidate sequences for pathway holes, but combines data from multiple, heterogeneous sources to assess the likelihood that a candidate has the required function. Our algorithm emulates the manual sequence annotation process, considering not only evidence from homology searches, but also considering evidence from genomic context (i.e., is the gene part of an operon?) and functional context (e.g., are there functionally-related genes nearby in the genome?) to determine the posterior belief that a candidate has the required function. The method can be applied across an entire metabolic pathway network and is generally applicable to any pathway database. The program uses a set of sequences encoding the required activity in other genomes to identify candidate proteins in the genome of interest, and then evaluates each candidate by using a simple Bayes classifier to determine the probability that the candidate has the desired function. We achieved 71% precision at a probability threshold of 0.9 during cross-validation using known reactions in computationally-predicted pathway databases. After applying our method to 513 pathway holes in 333 pathways from three Pathway/Genome databases, we increased the number of complete pathways by 42%. We made putative assignments to 46% of the holes, including annotation of 17 sequences of previously unknown function. CONCLUSIONS: Our pathway hole filler can be used not only to increase the utility of Pathway/Genome databases to both experimental and computational researchers, but also to improve predictions of protein function.

Amino Acid Oxidoreductases↗

A multidomain TIGR/olfactomedin protein family with conserved structural similarity in the N-terminal region and conserved motifs in the C-terminal region.

Based on the similarity between the TIGR (trabecular-meshwork inducible glucocorticoid response) (also known as myocilin) and olfactomedin protein families identified throughout the length of the TIGR protein, we have identified more distantly related proteins to determine the elements essential to the function/structure of the TIGR and olfactomedin proteins. Using a sequence walk method and the Shotgun program, we have identified a family including 31 olfactomedin domain-containing sequences. Multiple sequence alignments and secondary structure analyses were used to identify conserved sequence elements. Pairwise identity in the olfactomedin domain ranges from 8 to 64%, with an average pairwise identity of 24%. The N-terminal regions of the proteins fall into two subgroups, one including the TIGR and olfactomedin families and another group of apparently unrelated domains. The TIGR and olfactomedin sequences display conserved motifs including a residual leucine zipper region and maintain a similar secondary structure throughout the N-terminal region. The correlation between conserved elements and disease-associated mutations and apparent polymorphisms in human TIGR was also examined to evaluate the apparent importance of conserved residues to the function/structure of TIGR. Several residues have been identified as essential to the function and/or structure of the human TIGR protein based on their degree of conservation across the family and their implication in the pathogenesis of primary open-angle glaucoma. Additionally, we have identified a group of chitinase sequences containing several of the highly conserved motifs present in the C-terminal region of the olfactomedin domain-containing sequences.

Algorithms↗