PubMed Health⌕ Search

Biomedical subjects

M Huynen

Publications and source records attributed to M Huynen.

14 recordsLinked to original sources

Modularity in the gain and loss of genes: applications for function prediction.

Genes that are clustered on multiple genomes and are likely to functionally interact tend to be gained or lost together during genome evolution. Here, we demonstrate that exceptions to this pattern indicate relatively distant functional interactions between the encoded proteins. Hence, this can be used to divide predicted clusters of functionally interacting proteins into sub-clusters, and as such, to refine the prediction of their function and functional interactions.

Bacterial Proteins↗

Re-annotating the Mycoplasma pneumoniae genome sequence: adding value, function and reading frames.

Four years after the original sequence submission, we have re-annotated the genome of Mycoplasma pneumoniae to incorporate novel data. The total number of ORFss has been increased from 677 to 688 (10 new proteins were predicted in intergenic regions, two further were newly identified by mass spectrometry and one protein ORF was dismissed) and the number of RNAs from 39 to 42 genes. For 19 of the now 35 tRNAs and for six other functional RNAs the exact genome positions were re-annotated and two new tRNA(Leu) and a small 200 nt RNA were identified. Sixteen protein reading frames were extended and eight shortened. For each ORF a consistent annotation vocabulary has been introduced. Annotation reasoning, annotation categories and comparisons to other published data on M.pneumoniae functional assignments are given. Experimental evidence includes 2-dimensional gel electrophoresis in combination with mass spectrometry as well as gene expression data from this study. Compared to the original annotation, we increased the number of proteins with predicted functional features from 349 to 458. The increase includes 36 new predictions and 73 protein assignments confirmed by the published literature. Furthermore, there are 23 reductions and 30 additions with respect to the previous annotation. mRNA expression data support transcription of 184 of the functionally unassigned reading frames.

Amino Acid Sequence↗

Exploitation of gene context.

Recently, a number of techniques have been proposed that use completely sequenced genomes for the function prediction of individual proteins encoded therein. They use the fusion of genes, their conserved location in operons or merely their co-occurrence in genomes to predict the existence of functional interactions between the proteins they encode. This type of information complements functional features that are predicted by classical homology-based search techniques.

Animals↗

Predicting protein function by genomic context: quantitative evaluation and qualitative inferences.

Various new methods have been proposed to predict functional interactions between proteins based on the genomic context of their genes. The types of genomic context that they use are Type I: the fusion of genes; Type II: the conservation of gene-order or co-occurrence of genes in potential operons; and Type III: the co-occurrence of genes across genomes (phylogenetic profiles). Here we compare these types for their coverage, their correlations with various types of functional interaction, and their overlap with homology-based function assignment. We apply the methods to Mycoplasma genitalium, the standard benchmarking genome in computational and experimental genomics. Quantitatively, conservation of gene order is the technique with the highest coverage, applying to 37% of the genes. By combining gene order conservation with gene fusion (6%), the co-occurrence of genes in operons in absence of gene order conservation (8%), and the co-occurrence of genes across genomes (11%), significant context information can be obtained for 50% of the genes (the categories overlap). Qualitatively, we observe that the functional interactions between genes are stronger as the requirements for physical neighborhood on the genome are more stringent, while the fraction of potential false positives decreases. Moreover, only in cases in which gene order is conserved in a substantial fraction of the genomes, in this case six out of twenty-five, does a single type of functional interaction (physical interaction) clearly dominate (>80%). In other cases, complementary function information from homology searches, which is available for most of the genes with significant genomic context, is essential to predict the type of interaction. Using a combination of genomic context and homology searches, new functional features can be predicted for 10% of M. genitalium genes.

Bacterial Proteins↗

Pathway alignment: application to the comparative analysis of glycolytic enzymes.

Comparative analysis of metabolic pathways in different genomes yields important information on their evolution, on pharmacological targets and on biotechnological applications. In this study on glycolysis, three alternative ways of comparing biochemical pathways are combined: (1) analysis and comparison of biochemical data, (2) pathway analysis based on the concept of elementary modes, and (3) a comparative genome analysis of 17 completely sequenced genomes. The analysis reveals a surprising plasticity of the glycolytic pathway. Isoenzymes in different species are identified and compared; deviations from the textbook standard are detailed. Several potential pharmacological targets and by-passes (such as the Entner-Doudoroff pathway) to glycolysis are examined and compared in the different species. Archaean, bacterial and parasite specific adaptations are identified and described.

Enzymes↗

Neutral evolution of mutational robustness.

We introduce and analyze a general model of a population evolving over a network of selectively neutral genotypes. We show that the population's limit distribution on the neutral network is solely determined by the network topology and given by the principal eigenvector of the network's adjacency matrix. Moreover, the average number of neutral mutant neighbors per individual is given by the matrix spectral radius. These results quantify the extent to which populations evolve mutational robustness-the insensitivity of the phenotype to mutations-and thus reduce genetic load. Because the average neutrality is independent of evolutionary parameters-such as mutation rate, population size, and selective advantage-one can infer global statistics of neutral network topology by using simple population data available from in vitro or in vivo evolution. Populations evolving on neutral networks of RNA secondary structures show excellent agreement with our theoretical predictions.

Evolution, Molecular↗

A pyrimidine-rich exonic splicing suppressor binds multiple RNA splicing factors and inhibits spliceosome assembly.

The bovine papillomavirus type 1 (BPV-1) exonic splicing suppressor (ESS) is juxtaposed immediately downstream of BPV-1 splicing enhancer 1 and negatively modulates selection of a suboptimal 3' splice site at nucleotide 3225. The present study demonstrates that this pyrimidine-rich ESS inhibits utilization of upstream 3' splice sites by blocking early steps in spliceosome assembly. Analysis of the proteins that bind to the ESS showed that the U-rich 5' region binds U2AF65 and polypyrimidine tract binding protein, the C-rich central part binds 35- and 54-55-kDa serine/arginine-rich (SR) proteins, and the AG-rich 3' end binds alternative splicing factor/splicing factor 2. Mutational and functional studies indicated that the most critical region of the ESS maps to the central C-rich core (GGCUCCCCC). This core sequence, along with additional nonspecific downstream nucleotides, is sufficient for partial suppression of spliceosome assembly and splicing of BPV-1 pre-mRNAs. The inhibition of splicing by the ESS can be partially relieved by excess purified HeLa SR proteins, suggesting that the ESS suppresses pre-mRNA splicing by interfering with normal bridging and recruitment activities of SR proteins.

Alternative Splicing↗

Predicting function: from genes to genomes and back.

Predicting function from sequence using computational tools is a highly complicated procedure that is generally done for each gene individually. This review focuses on the added value that is provided by completely sequenced genomes in function prediction. Various levels of sequence annotation and function prediction are discussed, ranging from genomic sequence to that of complex cellular processes. Protein function is currently best described in the context of molecular interactions. In the near future it will be possible to predict protein function in the context of higher order processes such as the regulation of gene expression, metabolic pathways and signalling cascades. The analysis of such higher levels of function description uses, besides the information from completely sequenced genomes, also the additional information from proteomics and expression data. The final goal will be to elucidate the mapping between genotype and phenotype.

Bacterial Proteins↗

Homology-based fold predictions for Mycoplasma genitalium proteins.

Homology search techniques based on the iterative PSI-BLAST method in combination with various filters for low sequence complexity are applied to assign folds to all Mycoplasma genitalium proteins. The resulting procedure (implemented as a web server) is able to predict at least one domain in 37% of these proteins automatically, with an estimated accuracy higher than 98%. Taking structural features such as coiled coil or transmembrane regions aside, folds can be assigned to more than half of the globular proteins in a bacterium just by iterative sequence comparison.

Bacterial Proteins↗

Differential genome analysis applied to the species-specific features of Helicobacter pylori.

We introduce a simple and rapid strategy to identify genes that are responsible for species-specific phenotypes. The genome of a species that has a specific phenotype is compared with at least one, closely related, species that lacks this phenotype. Homologous genes that are shared among the species compared are identified and discarded from the list of candidates for species-specific genes. The process is automated and rapidly yields a small subset of the genome that likely contains genes responsible for the species-specific features. Functions are assigned to the genes, and dubious annotations are filtered out. Information is extracted not only from the presence of genes, but also from their absence with respect to known phenotypes. We have applied the technique to identify a set of species-specific genes in Helicobacter pylori by comparing it with its closest relatives for which complete genome sequences are available, Haemophilus influenzae and Escherichia coli. Of the genes of this set for which functional features can be obtained, a large fraction (63%, 123 proteins) is (potentially) involved in H. pylori's interaction with its host. We hypothesize that a family of outer membrane proteins is critical for the ability of H. pylori to colonize host cells in highly acidic environments.

Amino Acid Sequence↗

Conservation of gene order: a fingerprint of proteins that physically interact.

A systematic comparison of nine bacterial and archaeal genomes reveals a low level of gene-order (and operon architecture) conservation. Nevertheless, a number of gene pairs are conserved. The proteins encoded by conserved gene pairs appear to interact physically. This observation can therefore be used to predict functions of, and interactions between, prokaryotic gene products.

Archaeal Proteins↗

Assessing the reliability of RNA folding using statistical mechanics.

We have analyzed the base-pairing probability distributions of 16 S and 16 S-like, and 23 S and 23 S-like ribosomal RNAs of Archaea, Bacteria, chloroplasts, mitochondria and Eukarya, as predicted by the partition function approach for RNA folding introduced by McCaskill. A quantitative analysis of the reliability of RNA folding is done by comparing the base-pairing probability distributions with the structures predicted by comparative sequence analysis (comparative structures). We distinguish two factors that show a relationship to the reliability of RNA minimum free energy structure. The first factor is the dominance of one particular base-pair or the absence of base-pairing for a given base within the base-pairing probability distribution (BPPD). We characterize the BPPD per base, including the probability of not base-pairing, by its Shannon entropy (S). The S value indicates the uncertainty about the base-pairing of a base: low S values result from BPPDs that are strongly dominated by a single base-pair or by the absence of base-pairing. We show that bases with low S values have a relatively high probability that their minimum free energy (MFE) structure corresponds to the comparative structure. The BPPDs of prokaryotes that live at high temperatures (thermophilic Archaea and Bacteria) have, calculated at 37 degrees C, lower S values than the BPPDs of prokaryotes that live at lower temperatures (mesophilic and psychrophilic Archaea and Bacteria). This reflects an adaptation of the ribosomal RNAs to the environmental temperature. A second factor that is important to consider with regard to the reliability of MFE structure folding is a variable degree of applicability of the thermodynamic model of RNA folding for different groups of RNAs. Here we show that among the bases that show low S values, the Archaea and Bacteria have similar, high probabilities (0.96 and 0.94 in 16 S and 0.93 and 0.91 in 23 S, respectively) that the MFE structure corresponds to the comparative structure. These probabilities are lower in the chloroplasts (16 S 0.91, 23 S 0.79), mitochondria (16 S-like 0.89, 23 S-like 0.69) and Eukarya (18 S 0.81, 28 S 0.86).

Computer Simulation↗