PubMed HealthSearch

Biomedical subjects

M A Huynen

Publications and source records attributed to M A Huynen.

12 recordsLinked to original sources

Genome phylogeny based on gene content.

Species phylogenies derived from comparisons of single genes are rarely consistent with each other, due to horizontal gene transfer, unrecognized paralogy and highly variable rates of evolution. The advent of completely sequenced genomes allows the construction of a phylogeny that is less sensitive to such inconsistencies and more representative of whole-genomes than are single-gene trees. Here, we present a distance-based phylogeny constructed on the basis of gene content, rather than on sequence identity, of 13 completely sequenced genomes of unicellular species. The similarity between two species is defined as the number of genes that they have in common divided by their total number of genes. In this type of phylogenetic analysis, evolutionary distance can be interpreted in terms of evolutionary events such as the acquisition and loss of genes, whereas the underlying properties (the gene content) can be interpreted in terms of function. As such, it takes a position intermediate to phylogenies based on single genes and phylogenies based on phenotypic characteristics. Although our comprehensive genome phylogeny is independent of phylogenies based on the level of sequence identity of individual genes, it correlates with the standard reference of prokarytic phylogeny based on sequence similarity of 16s rRNA. Thus, shared gene content between genomes is quantitatively determined by phylogeny, rather than by phenotype, and horizontal gene transfer has only a limited role in determining the gene content of genomes.

Archaea

Automatic detection of conserved RNA structure elements in complete RNA virus genomes.

We propose a new method for detecting conserved RNA secondary structures in a family of related RNA sequences. Our method is based on a combination of thermodynamic structure prediction and phylogenetic comparison. In contrast to purely phylogenetic methods, our algorithm can be used for small data sets of approximately 10 sequences, efficiently exploiting the information contained in the sequence variability. The procedure constructs a prediction only for those parts of sequences that are consistent with a single conserved structure. Our implementation produces reasonable consensus structures without user interference. As an example we have analysed the complete HIV-1 and hepatitis C virus (HCV) genomes as well as the small segment of hantavirus. Our method confirms the known structures in HIV-1 and predicts previously unknown conserved RNA secondary structures in HCV.

Algorithms

Measuring genome evolution.

The determination of complete genome sequences provides us with an opportunity to describe and analyze evolution at the comprehensive level of genomes. Here we compare nine genomes with respect to their protein coding genes at two levels: (i) we compare genomes as "bags of genes" and measure the fraction of orthologs shared between genomes and (ii) we quantify correlations between genes with respect to their relative positions in genomes. Distances between the genomes are related to their divergence times, measured as the number of amino acid substitutions per site in a set of 34 orthologous genes that are shared among all the genomes compared. We establish a hierarchy of rates at which genomes have changed during evolution. Protein sequence identity is the most conserved, followed by the complement of genes within the genome. Next is the degree of conservation of the order of genes, whereas gene regulation appears to evolve at the highest rate. Finally, we show that some genomes are more highly organized than others: they show a higher degree of the clustering of genes that have orthologs in other genomes.

Animals

The frequency distribution of gene family sizes in complete genomes.

We compare the frequency distribution of gene family sizes in the complete genomes of six bacteria (Escherichia coli, Haemophilus influenzae, Helicobacter pylori, Mycoplasma genitalium, Mycoplasma pneumoniae, and Synechocystis sp. PCC6803), two Archaea (Methanococcus jannaschii and Methanobacterium thermoautotrophicum), one eukaryote (Saccharomyces cerevisiae), the vaccinia virus, and the bacteriophage T4. The sizes of the gene families versus their frequencies show power-law distributions that tend to become flatter (have a larger exponent) as the number of genes in the genome increases. Power-law distributions generally occur as the limit distribution of a multiplicative stochastic process with a boundary constraint. We discuss various models that can account for a multiplicative process determining the sizes of gene families in the genome. In particular, we argue that, in order to explain the observed distributions, gene families have to behave in a coherent fashion within the genome; i.e., the probabilities of duplications of genes within a gene family are not independent of each other. Likewise, the probabilities of deletions of genes within a gene family are not independent of each other.

Algorithms

Smoothness within ruggedness: the role of neutrality in adaptation.

RNA secondary structure folding algorithms predict the existence of connected networks of RNA sequences with identical structure. On such networks, evolving populations split into subpopulations, which diffuse independently in sequence space. This demands a distinction between two mutation thresholds: one at which genotypic information is lost and one at which phenotypic information is lost. In between, diffusion enables the search of vast areas in genotype space while still preserving the dominant phenotype. By this dynamic the success of phenotypic adaptation becomes much less sensitive to the initial conditions in genotype space.

Adaptation, Biological

Exploring phenotype space through neutral evolution.

RNA secondary-structure folding algorithms predict the existence of connected networks of RNA sequences with identical secondary structures. Fitness landscapes that are based on the mapping between RNA sequence and RNA secondary structure hence have many neutral paths. A neutral walk on these fitness landscapes gives access to a virtually unlimited number of secondary structures that are a single point mutation from the neutral path. This shows that neutral evolution explores phenotype space and can play a role in adaptation.

Algorithms

Base pairing probabilities in a complete HIV-1 RNA.

We have calculated the base pair probability distribution for the secondary structure of a full length HIV-1 genome using the partition function approach introduced by McCaskill (1990). By analyzing the full distribution of base pair probabilities instead of a restricted number of secondary structures, we gain more complete and reliable information about the secondary structure of HIV-1. We introduce methods that condense the information in the probability distribution to one value per nucleotide in the sequence. Using these methods we represent the secondary structure as a weighted average of the base pair probabilities, and we can identify interesting secondary structures that have relatively well-defined base pairing. The results show high probabilities for the known secondary structures at the 5'-end of the molecule that have been predicted on the basis of biochemical data. The Rev response element (RRE) appears as a distinct element in the secondary structure. It has a meta-stable domain at the high affinity site for the binding of Rev. The overall structure decomposes into fairly small independent structures in the first 4,000 bases of the molecule. The remaining 5,000 bases (excluding the terminal repeat) form a single, large structure, on top of which the RRE is located.

Algorithms

Pattern generation in molecular evolution: exploitation of the variation in RNA landscapes.

Evolution of RNA secondary structure is studied using simulation techniques and statistical analysis of fitness landscapes. The transition from RNA sequence to RNA secondary structure leads to fitness landscapes that have local variations in their "ruggedness." Evolution exploits these variations. In stable environments it moves the quasispecies toward relatively "flat" peaks, where not only the master sequence but also its mutants have a high fitness. In a rapidly changing environment, the situation is reversed; evolution moves the quasispecies to a region where the correlation between secondary structures of "neighboring" RNA sequences is relatively low. In selection for simple secondary structures the movement toward flat peaks leads to pattern generation in the RNA sequences. Patterns are generated at the level of polynucleotide frequencies and the distribution of purines and pyrimidines. The patterns increase the modularity of the sequence. They thereby prevent the formation of alternative secondary structures after mutations. The movement of the quasispecies toward relatively rugged parts of the landscape results in pattern generation at the level of the RNA secondary structure. The base-pairing frequency of the sequences increases. The patterns that are generated in the RNA sequences and the RNA secondary structures are not directly selected for and can be regarded as a side effect of the evolutionary dynamics of the system.

Animals

Multiple coding and the evolutionary properties of RNA secondary structure.

This article evaluates evolutionary properties of the transition from RNA primary sequence to RNA secondary structure. It focuses on the restrictions that the conservation of a protein code in an RNA sequence puts on its potential to evolve towards a specific secondary structure. Restricting the mutations to those that do not affect the coding for a protein restricts both the accessibility and the connectivity of the sequence space. The accessibility is restricted because only certain point mutations are allowed. The connectivity is restricted because no insertions and deletions are allowed. Simulating an evolutionary search process for a specific secondary structure shows that (i) the reduction of allowable point mutations allows for adaptation to some large-scale topology, but strongly reduces the possibility of small-scale adaptations, (ii) the abolition of insertions and deletions has very little effect on the results of the search process. During the evolutionary search process for a secondary structure with a specific topology and a high frequency of base-pairing the quasispecies moves into a subspace in which the similarity between secondary structures of neighboring sequences is relatively high. Increased similarity between second structures of neighboring sequences is also found in the Rev responsive element (RRE) in the lentiviruses Caprine arthritis-encephalitis virus and Visna virus. In these viruses a biased nucleotide frequency in the RRE region suggests that selection for the RRE RNA secondary structure affects the amino acid sequence of the env gene. Our results show a variation in the ruggedness of fitness landscapes which are based on a high degree of epistatic interactions. Fitness landscapes play an essential role, not only in biotic evolution, but also in all kinds of optimization processes (Hill Climbing, Simulated Annealing, Genetic Algorithms, etc). Variation in their ruggedness should therefore be taken into account in the analysis of these processes.

Amino Acid Sequence

Equal G and C contents in histone genes indicate selection pressures on mRNA secondary structure.

Protein-specific versus taxon-specific patterns of nucleotide frequencies were studied in histone genes. The third positions of codons have a (well-known) taxon-specific G+C level and a histone type-specific G/C ratio. This ratio counterbalances the G/C ratio in the first and second positions so that the overall G and C levels in the coding region become approximately equal. The compensation of the G/C ratio indicates a selection pressure at the mRNA level rather than a selection pressure or mutation bias at the DNA level or a selection pressure on codon usage. The structure of histone mRNAs is compatible with the hypothesis that the G/C compensation is due to selection pressures on mRNA secondary structure. Nevertheless, no specific motifs seem to have been selected, and the free energy of the secondary structures is only slightly lower than that expected on the basis of nucleotide frequencies.

Animals