PubMed HealthSearch

Biomedical subjects

M Gouy

Publications and source records attributed to M Gouy.

At least 19 recordsLinked to original sources

Inferring phylogenies from DNA sequences of unequal base compositions.

A new method for computing evolutionary distances between DNA sequences is proposed. Contrasting with classical methods, the underlying model does not assume that sequence base compositions (A, C, G, and T contents) are at equilibrium, thus allowing unequal base compositions among compared sequences. This makes the method more efficient than the usual ones in recovering phylogenetic trees from sequence data when base composition is heterogeneous within the data set, as we show by using both simulated and empirical data. When applied to small-subunit ribosomal RNA sequences from several prokaryotic or eukaryotic organisms, this method provides evidence for an early divergence of the microsporidian Vairimorpha necatrix in the eukaryotic lineage.

Algorithms

Isolation and characterization of a cDNA encoding a chicken actin-like protein.

We report the isolation and characterization of a chicken cDNA which putatively encodes an actin-like protein (chACTL). This 394-amino-acid (aa) polypeptide shares sequence homology (81, 70 and 67% identical aa, respectively) with three actin-related proteins (ARP) described for Drosophila melanogaster (ARP14D), Caenorhabditis elegans (ACTL) and Saccharomyces cerevisiae (ACT2). At least six chACTL transcripts were detected in different tissues during chick embryogenesis. Sequence analysis suggests that at least three groups of ARP have been evolutionarily conserved.

Actins

NRSub: a non-redundant data base for the Bacillus subtilis genome.

We have organized the DNA sequences of Bacillus subtillis from the EMBL collection to build the NRSub data base. This data base is free from duplications and all detected overlapping sequences are merged into contigs. Data on gene mapping and codon usage are also included. NRSub is publically available through anonymous FTP in flat file format or structured on the form of an ACNUC data base. Under this format, it is possible to use NRSub with the retrieval program Query--win. This program integrates a graphical interface and may be installed on any kind of UNX computer under X Window and on which the Vibrant and Motif libraries are available.

Bacillus subtilis

HOVERGEN: a database of homologous vertebrate genes.

Comparison of homologous genes is a major step for many studies related to genome structure, function or evolution. Similarity search programs easily find genes homologous to a given sequence. However, only very tedious manual procedures allow the retrieval of all sets of homologous genes sequenced for a given set of species. Moreover, this search often generates errors due to the complexity of data to be managed simultaneously: phylogenetic trees, alignments, taxonomy, sequences and related information. HOVERGEN helps to solve these problems by integrating all this information. HOVERGEN corresponds to GenBank sequences from all vertebrate species, with some data corrected, clarified, or completed, notably to address the problem of redundancy. Coding sequences have been classified in gene families. Protein multiple alignments and phylogenetic trees have been calculated for each family. Sequences and related information have been structured in an ACNUC database which permits complex selections. A graphical interface has been developed to visualize and edit trees. Genes are displayed in color, according to their taxonomy. Users have directly access to all information attached to sequences and to multiple alignments simply by clicking on genes. This graphical tool gives thus a rapid and simple access to all data necessary to interpret homology relationships between genes. HOVERGEN allows the user to easily select sets of homologous vertebrate genes, and thus is particularly useful for comparative sequence analysis, or molecular evolution studies.

Animals

Molecular phylogeny of Eubacteria: a new multiple tree analysis method applied to 15 sequence data sets questions the monophyly of gram-positive bacteria.

Phylogenetic relationships between the major eubacterial phyla were studied using the sequences of 15 homologous bacterial genes. Neither the classical concatenation strategy nor a new multiple tree analysis method involving statistical tests of the inferred phylogenetic relationships provided any (solid) conclusions about eubacterial phylogeny; no pairs of eubacterial phyla proved to be closer to each other in the 15 reconstructed trees than would be expected for trees with random topologies. The phylogeny of Eubacteria therefore appears to be tightly bush-like. Moreover, results from both concatenation and multiple tree analysis raise doubts concerning the monophyly of the so-called Gram-positive bacteria phylum, since the monophyly hypothesis is no more strongly supported by data than its alternatives. It is noteworthy that the structural bases for the Gram-positive phenotype are not incompatible with the hypothesis of independent emergence of this character at two different times.

Bacteria

Phylogenetic position of foraminifera inferred from LSU rRNA gene sequences.

A 5'-terminal region of 1600-1800 base pairs was amplified, cloned, and sequenced in the large subunit rDNA (LSU rDNA) of four species of foraminifera. These sequences were compared with the homologous regions of 16 eukaryotic taxa in order to establish the phylogenetic position of foraminifera. Analysis of 610 unambiguously aligned bases shows that foraminifera branch closely to plasmodial and cellular slime molds in the middle of the eukaryotic tree--that is, much earlier than suggested by the fossil record. These data, the first DNA sequences reported for foraminifera, will help analyze this class of protists and the early evolution of eukaryotes.

Animals

Molecular phylogenetic analysis of Nitrobacter spp.

The phylogeny of bacteria belonging to the genus Nitrobacter was investigated by sequencing the whole 16S rRNA gene. The average level of similarity for the three Nitrobacter strains examined was high (99.2%), and the similarity level between Nitrobacter winogradskyi and Nitrobacter sp. strain LL, which represent two different genomic species, was even higher (99.6%). When all of the Nitrobacter strains and their phylogenetic neighbors Bradyrhizobium and Rhodopseudomonas species were considered, the average similarity level was 98.1%. When complete sequences were used, Nitrobacter hamburgensis clustered with the two other Nitrobacter strains, while this was not the case when partial sequences were used. The two Rhodopseudomonas palustris strains examined exhibited a low similarity level (97.6%) and were not clustered.

Base Sequence

Heterogeneity of hepatitis C virus genotypes in France.

The genotypes of French hepatitis C virus (HCV) isolates were investigated by amplification of a domain from the non-structural region 3 (NS3) using nested PCR, followed by hybridization with two genotype-specific probes, F1 (HCV type I-specific) and F2 (HCV type II-specific). Among 119 HCV RNA-positive sera, 91% of samples were NS3 PCR positive. Most samples (83.2%) hybridized with one or the other probe only, whereas a few samples (4.2%) hybridized with both F1 and F2 probes (HB). A small percentage (3.4%) of samples appeared unable to hybridize with either probe (HN). For some of these samples (HB1, HB2, HN1, HN2, HN3, HN4), part of the NS3, core and envelope regions were sequenced and the corresponding deduced consensus sequences were compared with those of prototype isolates of the four HCV genotypes (types I to IV). A phylogenetic tree was constructed to illustrate the relationship between these isolates. The results obtained showed that (i) HN4 appears to be more closely related to type III than to type IV HCV genotypes, which suggests that in France there may exist additional although minor genotypes besides the two major types, F1 and F2. (ii) HB1, HB2, HN1, HN2 and probably HN3 belong to the type II HCV genotype. The association between sequence diversity and putative biological difference for isolates within the same genotype remains to be elucidated.

Africa, Northern

Molecular phylogeny of the symbiotic actinomycetes of the genus Frankia matches host-plant infection processes.

Nucleotide sequences of approximately 213 bp of the nif H-D intergene and the beginning of nifD were determined for symbiotic Frankia isolates from the major host-infectivity groups. This region of the nif operon is variable enough to classify most infective Frankia strains at the species level. Phylogenetic inferences from these sequences are in agreement with the 16S rRNA-derived phylogeny of the genus and, thus, are in favor of an intrageneric evolution of nif genes by orthology. Phylogenetic lineages derived from combined nifH-D intergene and partial nifD and 16S rRNA sequences are supported for at least 93% of bootstrap replicates and are useful for investigating evolutionary relationships of the genus and symbiotic properties of this microorganism. The genus Frankia is divided into two major phylogenetic clusters that match with the separation of species according to the mechanism of infection of actinorhizal plants. One cluster groups species strictly adapted to the mechanism of root hair infection (RHI), and the other groups species adapted to the mechanism of direct intercellular penetration. In the RHI cluster, the species infective on Casuarina plants appears to have emerged from strains infective on Alnus. The concordance between the symbiotic properties and the molecular phylogeny of Frankia strains indicates a major role for the host plant in the evolution and speciation of the genus Frankia.

Actinomyces

Insect muscle actins differ distinctly from invertebrate and vertebrate cytoplasmic actins.

Invertebrate actins resemble vertebrate cytoplasmic actins, and the distinction between muscle and cytoplasmic actins in invertebrates is not well established as for vertebrate actins. However, Bombyx and Drosophila have actin genes specifically expressed in muscles. To investigate if the distinction between muscle and cytoplasmic actins evidenced by gene expression analysis is related to the sequence of corresponding genes, we compare the sequences of actin genes of these two insect species and of other Metazoa. We find that insect muscle actins form a family of related proteins characterized by about 10 muscle-specific amino acids. Insect muscle actins have clearly diverged from cytoplasmic actins and form a monophyletic group emerging from a cluster of closely related proteins including insect and vertebrate cytoplasmic actins and actins of mollusc, cestode, and nematode. We propose that muscle-specific actin genes have appeared independently at least twice during the evolution of animals: insect muscle actin genes have emerged from an ancestral cytoplasmic actin gene within the arthropod phylum, whereas vertebrate muscle actin genes evolved within the chordate lineage as previously described.

Actins

Evolution of the primate beta-globin gene region: nucleotide sequence of the delta-beta-globin intergenic region of gorilla and phylogenetic relationships between African apes and man.

A 6.0-kb DNA fragment from Gorilla gorilla including the 5' part of the beta-globin gene and about 4.5 kb of its upstream flanking region was cloned and sequenced. The sequence was compared to the human, chimpanzee, and macaque delta-beta intergenic region. This analysis reveals four tandemly repeated sequences (RS), at the same location in the four species, showing a variable number of repeats generating both intraspecific (polymorphism) and interspecific variability. These tandem arrays delimit five regions of unique sequence called IG for intergenic. The divergence for these IG sequences is 1.85 +/- 0.22% between human and gorilla, which is not significantly different from the value estimated in the same region between chimpanzee and human (1.62 +/- 0.21%). The CpG and TpA dinucleotides are avoided. CpGs evolve faster than other sequence sites but do not confuse phylogenetic inferences by producing parallel mutations in different lineages. About 75% of CpG doublets have become TpG or CpA since the common ancestor, in agreement with the methylation/deamination pattern. Comparison of this intergenic region gives information on branching order within Hominoidea. Parsimony and distance-based methods when applied to the delta-beta intergenic region provide evidence (although not statistically significant) that human and chimpanzee are more closely related to each other than to gorilla. CpG sites are indeed rich in information by carrying substitutions along the short internal branch. Combining these results with those on the psi eta-delta intergenic region, shows in a statistically significant way that chimpanzee is the closest relative of human.

Animals

Nucleotide sequence of nifD from Frankia alni strain ArI3: phylogenetic inferences.

The complete nucleotide sequence of the nifD gene encoding the alpha subunit of component I of nitrogenase from Frankia alni strain ArI3 was determined. The coding region is 1,458 bp in length and encodes a polypeptide of 486 residues with a predicted molecular weight of 53,500. Phylogenetic inferences with 12 complete published nifD sequences were drawn using a variety of approaches. Frankia nifD clusters with proteobacteria rather than with Clostridium pasteurianum, the other Gram-positive bacterium studied. Extant eubacterial nif genes seem to have at least three distinct evolutionary origins as a result of ancient gene duplications. Within the Gram-positive bacterial phylum, functional nif genes descend from different duplicates.

Actinomycetales

Molecular phylogeny of Rodentia, Lagomorpha, Primates, Artiodactyla, and Carnivora and molecular clocks.

Phylogenetic analysis of DNA sequences from primates, rodents, lagomorphs, artiodactyls, carnivores, and birds strongly suggests that the order Rodentia is an outgroup to the other four mammalian orders and that Artiodactyla and Carnivora belong to a superordinal clade. Further, there is strong evidence against the Glires concept, which unites Lagomorpha and Rodentia. The radiation among Lagomorpha, Primates, and Artiodactyla--Carnivora is very bush-like, but there is some evidence that Lagomorpha has branched off first. Thus, the branching sequence for these five orders of mammals seems to be Rodentia, Lagomorpha, Primates, Artiodactyla, and Carnivora. The branching date for Rodentia could be as early as 100 million years ago. The rate of nucleotide substitution in the rodent lineage is shown to be at least 1.5 times higher than those in the other four mammalian lineages.

Animals

Phylogenetic analysis based on rRNA sequences supports the archaebacterial rather than the eocyte tree.

How many primary lineages of life exist and what are their evolutionary relationships? These are fundamental but highly controversial issues. Woese and co-workers propose that archaebacteria, eubacteria and eukaryotes are the three primary lines of descent and their relationships can be represented by Fig. 1a (the 'archaebacterial tree') if one neglects the root of the tree. In contrast, Lake claims that archaebacteria are paraphyletic, and he groups eocytes (extremely thermophilic, sulphur-dependent bacteria) with eukaryotes, and halobacteria with eubacteria (the 'eocyte tree', Fig. 1b). Lake's view has gained considerable support as a result of an analysis of small subunit ribosomal RNA sequence data by a new approach, the evolutionary parsimony method. Here we report that analysis of small subunit data by the neighbour-joining and maximum parasimony methods favours the archaebacterial tree and that computer simulations using either the archaebacterial or the eocyte tree as a model tree show that the probability of recovering the model tree is very high (greater than 90 per cent) for both the neighbour-joining and maximum parsimony methods but is relatively low for the evolutionary parsimony method. Moreover, analysis of large subunit rRNA sequences by all three methods strongly favours the archaebacterial tree.

Archaea

Date of the monocot-dicot divergence estimated from chloroplast DNA sequence data.

The divergence between monocots and dicots represents a major event in higher plant evolution, yet the date of its occurrence remains unknown because of the scarcity of relevant fossils. We have estimated this date by reconstructing phylogenetic trees from chloroplast DNA sequences, using two independent approaches: the rate of synonymous nucleotide substitution was calibrated from the divergence of maize, wheat, and rice, whereas the rate of nonsynonymous substitution was calibrated from the divergence of angiosperms and bryophytes. Both methods lead to an estimate of the monocot-dicot divergence at 200 million years (Myr) ago (with an uncertainty of about 40 Myr). This estimate is also supported by analyses of the nuclear genes encoding large and small subunit ribosomal RNAs. These results imply that the angiosperm lineage emerged in Jurassic-Triassic time, which considerably predates its appearance in the fossil record (approximately 120 Myr ago). We estimate the divergence between cycads and angiosperms to be approximately 340 Myr, which can be taken as an upper bound for the age of angiosperms.

Animals

Molecular phylogeny of the kingdoms Animalia, Plantae, and Fungi.

The branching order of the kingdoms Animalia, Plantae, and Fungi has been a controversial issue. Using the transformed distance method and the maximum parsimony method, we investigated this problem by comparing the sequences of several kinds of macromolecules in organisms spanning all three kingdoms. The analysis was based on the large-subunit and small-subunit ribosomal RNAs, 10 isoacceptor transfer RNA families, and six highly conserved proteins. All three sets of sequences support the same phylogenetic tree: plants and animals are sibling kingdoms that have diverged more recently than the fungi. The ribosomal RNA and protein data sets are large enough so that in both cases the inferred phylogeny is statistically significant. The present report appears to be the first to provide statistically conclusive molecular evidence for the phylogeny of the three kingdoms. The determination of this phylogeny will help us to understand the evolution of various molecular, cellular, and developmental characters shared by any two of the three kingdoms. Noting that the large-subunit rRNA sequences have evolved at similar rates in the three kingdoms, we estimated the ratio of the time since the animal-plant split to the time since the fungal divergence to be 0.90.

Animal Population Groups

RPA190, the gene coding for the largest subunit of yeast RNA polymerase A.

Yeast RNA polymerases are being extensively studied at the gene level. The entire gene encoding the largest subunit of RNA polymerase A, A190, was isolated and characterized in detail. Southern hybridization and gene disruption experiments showed that the RPA190 gene is unique in the haploid yeast genome and essential for cell viability. Nuclease S1 mapping was used to identify mRNA 5' and 3' termini. RPA190 encodes a polypeptide chain of 186,270 daltons in a large uninterrupted reading frame. A dot matrix comparison of the deduced amino acid sequence of subunit A190 with Escherichia coli beta' and cognate subunits B220 and C160 from yeast RNA polymerases B and C showed a conserved pattern of homology regions (I-VI). A potential DNA-binding site (zinc-binding motif) is conserved in the N-terminal region I. Remarkably, the A190 subunit does not harbor the heptapeptide repeated sequence present in the B220 subunit. The sequence of the A190 subunit diverges from B220 and C160 by the presence of two hydrophilic domains inserted between homology regions I and II, and V and VI. From their codon usage and third base pyrimidine bias, RNA polymerase genes RPA190, RPB220, RPC160, and RPC40 fall among yeast genes expressed at an average level. The RPA190 5'-flanking region contains features present in other polymerase genes that might function in regulation.

Amino Acid Sequence