PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Explosive lineage-specific expansion of the orphan nuclear receptor HNF4 in nematodes.

The nuclear receptor superfamily expanded in at least two episodes: one early in metazoan evolution, the second within the vertebrate lineage. An exception to this pattern is the genome of the nematode Caenorhabditis elegans, which encodes more than 270 nuclear receptors, most of them highly divergent. We generated 128 cDNA sequences for 76 C. elegans nuclear receptors, confirming that these are active genes. Among these numerous receptors are 13 orthologues of nuclear receptors found in arthropods and/or vertebrates. We show that the supplementary nuclear receptors (supnrs) originated from an explosive burst of duplications of a unique orphan receptor, HNF4. This origin has specific implications for the role of ligand binding in the function and evolution of the nematode supplementary nuclear receptors. Moreover, the supplementary nuclear receptors include a group of very rapidly evolving genes found primarily on chromosome V. We propose a model of lineage-specific duplications from a chromosome on which duplication and substitution rates are highly increased. Our results provide a framework to study nuclear receptors in nematodes, as well as to consider the functional and evolutionary consequences of lineage-specific duplications.

Animals↗

The origin and evolution of operons: the piecewise building of the proteobacterial histidine operon.

The structure and organization of 470 histidine biosynthetic genes from 47 different proteobacteria were combined with phylogenetic inference to investigate the mechanisms responsible for assembly of the his pathway and the origin of his operons. Data obtained in this work showed that a wide variety of different organization strategies of his gene arrays exist and that some his genes or entire his operons are likely to have been horizontally transferred between bacteria of the same or different proteobacterial branches. We propose a "piecewise" model for the origin and evolution of proteobacterial his operons, according to which the initially scattered his genes of the ancestor of proteobacteria coded for monofunctional enzymes (except possibly for hisD) and underwent a stepwise compacting process that reached its culmination in some gamma-proteobacteria. The initial step of operon buildup was the formation of the his "core," a cluster consisting of four genes (hisBHAF) whose products interconnect histidine biosynthesis to both de novo synthesis of purine metabolism and that occurred in the common ancestor of the alpha/beta/gamma branches, possibly after its separation from the epsilon one. The following step was the formation of three mini-operons (hisGDC, hisBHAF, hisIE) transcribed from independent promoters, that very likely occurred in the ancestor of the beta/gamma-branch, after its separation from the alpha one. Then the three mini-operons joined together to give a compact operon. In most gamma-proteobacteria the two fusions involving the gene pairs hisN-B and hisI-E occurred. Finally the gamma-proteobacterial his operon was horizontally transferred to other proteobacteria, such as Campylobacter jejuni. The biological significance of clustering of his genes is also discussed.

Databases, Nucleic Acid↗

Molecular evolution of duplicated ray finned fish HoxA clusters: increased synonymous substitution rate and asymmetrical co-divergence of coding and non-coding sequences.

In this study the molecular evolution of duplicated HoxA genes in zebrafish and fugu has been investigated. All 18 duplicated HoxA genes studied have a higher non-synonymous substitution rate than the corresponding genes in either bichir or paddlefish, where these genes are not duplicated. The higher rate of evolution is not due solely to a higher non-synonymous-to-synonymous rate ratio but to an increase in both the non-synonymous as well as the synonymous substitution rate. The synonymous rate increase can be explained by a change in base composition, codon usage, or mutation rate. We found no changes in nucleotide composition or codon bias. Thus, we suggest that the HoxA genes may experience an increased mutation rate following cluster duplication. In the non-Hox nuclear gene RAG1 only an increase in non-synonymous substitutions could be detected, suggesting that the increased mutation rate is specific to duplicated Hox clusters and might be related to the structural instability of Hox clusters following duplication. The divergence among paralog genes tends to be asymmetric, with one paralog diverging faster than the other. In fugu, all b-paralogs diverge faster than the a-paralogs, while in zebrafish Hoxa-13a diverges faster. This asymmetry corresponds to the asymmetry in the divergence rate of conserved non-coding sequences, i.e., putative cis-regulatory elements. These results suggest that the 5' HoxA genes in the same cluster belong to a co-evolutionary unit in which genes have a tendency to diverge together.

Animals↗

Phylogenetic analysis of eukaryotic thiolases suggests multiple proteobacterial origins.

Eukaryotic thiolases are essential enzymes located in three different compartments (peroxisome, mitochondrion, and cytosol) that can display catabolic or anabolic functions. They are responsible for the thiolytic cleavage of oxidized acyl-CoA (thiolase I; EC 2.3.1.16) and the synthesis or degradation of acetoacetyl-CoA (thiolase II; EC 2.3.1.9). Phylogenetic analysis of eukaryotic thiolase sequences showed that they form six distinct clusters, one of them highly divergent, which are in good correlation with their class and subcellular location. When analyzed together with a representative sample of prokaryotic thiolases, all eukaryotic thiolase groups emerged close to proteobacterial sequences. Metazoan cytosolic thiolase II was related to alpha-proteobacterial sequences, suggesting a mitochondrial origin. Unexpectedly, cytosolic thiolases from green plants and fungi as well as at least one member of all eukaryotic peroxisomal and mitochondrial thiolases had delta-proteobacteria as closest relatives. Our analysis suggests that these eukaryotic peroxisomal and mitochondrial thiolases may have been acquired from delta-proteobacteria prior to the ancestor of all known eukaryotes.

Acetyl-CoA C-Acyltransferase↗

Cytosine methylation is not the major factor inducing CpG dinucleotide deficiency in bacterial genomes.

CpG dinucleotide deficiency has been found in viruses, mitochondria, prokaryotes, and eukaryotes. The consensual explanation is that it is due to deamination of methylated cytosines, as established for vertebrate and plants. However, we still do not know whether C5 cytosine methylation is also the major cause of CpG deficiency in bacteria. By combining annotation and experimental data identifying the presence of C5 cytosine methyltransferases with analysis of CpG relative abundance in 67 bacterial species, we found that CpG relative abundance in most bacterial genomes that have cytosine C5 methyltransferases tends to be in the normal range (observed/expected values between 0.82 and 1.21). In contrast, many bacterial species likely to be lacking C5 cytosine methylation showed CpG deficiency. Furthermore, when comparing genomes with one another, TpG and CpA relative abundances were found to be independent from CpG relative abundance. This contrasted with intragenome analyses, where C3pG1 relative abundance (the subscripts refer to position of a nucleotide in a codon) was found to be generally positively correlated with T3pG1 relative abundances when plotted against GC content in protein coding sequences (CDSs). This suggests the existence of alternative mechanisms contributing to CpG deficiency in bacteria.

Bacteria↗

New aspects on lanosterol 14alpha-demethylase and cytochrome P450 evolution: lanosterol/cycloartenol diversification and lateral transfer.

Sterol 14alpha-demethylase (CYP51) is a member of the cytochrome P450 superfamily, widely found in animals, fungi, and plants but present in few prokaryotic groups. CYP51 is currently believed to be the ancestral cytochrome P450 that has been transferred from prokaryotes to eukaryotic kingdoms. We propose an alternate view of CYP51 evolution that has an impact on understanding the evolution of the entire CYP superfamily. Two hundred forty-nine bacterial and four archaeal CYP sequences have been aligned and a bacterial CYP tree designed, showing a separation of two branches. Prokaryotic CYP51s cluster to the minor branch, together with other eukaryote-like CYPs. Mycobacterial and methylococcal CYP51s cluster together (100% bootstrap probability), while Streptomyces CYP51 remains on a distant branch. A CYP51 phylogenetic tree has been constructed from 44 sequences resulting in a ((plant, bacteria),(animal, fungi)) topology (100% bootstrap probability). This is in accordance with the lanosterol/cycloartenol diversification of sterol biosynthesis. The lanosterol branch (nonphotosynthetic lineage) follows the previously proposed topology of animal and fungal orthologues (100% bootstrap probability), while plant and D. discoideum CYP51s belong to the cycloartenol branch (photosynthetic lineage), all in accordance with biochemical data. Bacterial CYP51s cluster within the cycloartenol branch (69% bootstrap probability), which is indicative of a lateral gene transfer of a plant CYP51 to the methylococcal/mycobacterial progenitor, suggesting further that bacterial CYP51s are not the oldest CYP genes. Lateral gene transfer is likely far more important than hitherto thought in the development of the diversified CYP superfamily. Consequently, bacterial CYPs may represent a mixture of genes with prokaryotic and eukaryotic origin.

Bacteria↗

Many independent origins of trans splicing of a plant mitochondrial group II intron.

We examined the cis- vs. trans-splicing status of the mitochondrial group II intron nad1i728 in 439 species (427 genera) of land plants, using both Southern hybridization results (for 416 species) and intron sequence data from the literature. A total of 164 species (157 genera), all angiosperms, was found to have a trans-spliced form of the intron. Using a multigene land plant phylogeny, we infer that the intron underwent a transition from cis to trans splicing 15 times among the sampled angiosperms. In 10 cases, the intron was fractured between its 5' end and the intron-encoded matR gene, while in the other 5 cases the fracture occurred between matR and the 3' end of the intron. The 15 intron fractures took place at different time depths during the evolution of angiosperms, with those in Nymphaeales, Austrobaileyales, Chloranthaceae, and eumonocots occurring early in angiosperm evolution and those in Syringodium filiforme, Hydrocharis morsus- ranae, Najas, and Erodium relatively recently. The trans-splicing events uncovered in Austrobaileyales, eumonocots, Polygonales, Caryophyllales, Sapindales, and core Rosales reinforce the naturalness of these major clades of angiosperms, some of which have been identified solely on the basis of recent DNA sequence analyses.

Blotting, Southern↗

Gene conversion and functional divergence in the beta-globin gene family.

Different models of gene family evolution have been proposed to explain the mechanism whereby gene copies created by gene duplications are maintained and diverge in function. Ohta proposed a model which predicts a burst of nonsynonymous substitutions following gene duplication and the preservation of duplicates through positive selection. An alternative model, the duplication-degeneration-complementation (DDC) model, does not explicitly require the action of positive Darwinian selection for the maintenance of duplicated gene copies, although purifying selection is assumed to continue to act on both copies. A potential outcome of the DDC model is heterogeneity in purifying selection among the gene copies, due to partitioning of subfunctions which complement each other. By using the d(N)/ d(S) (omega) rate ratio to measure selection pressure, we can distinguish between these two very different evolutionary scenarios. In this study we investigated these scenarios in the beta-globin family of genes, a textbook example of evolution by gene duplication. We assembled a comprehensive dataset of 72 vertebrate beta-globin sequences. The estimated phylogeny suggested multiple gene duplication and gene conversion events. By using different programs to detect recombination, we confirmed several cases of gene conversion and detected two new cases. We tested evolutionary scenarios derived from Ohta's model and the DDC model by examining selective pressures along lineages in a phylogeny of beta-globin genes in eutherian mammals. We did not find significant evidence for an increase in the omega ratio following major duplication events in this family. However, one exception to this pattern was the duplication of gamma-globin in simian primates, after which a few sites were identified to be under positive selection. Overall, our results suggest that following gene duplications, paralogous copies of beta-globin genes evolved under a nonepisodic process of functional divergence.

Animals↗

Isochore structures in the genome of the plant Arabidopsis thaliana.

Arabidopsis thaliana is an important model system for the study of plant biology. We have analyzed the complete genome sequences of Arabidopsis by using a newly developed windowless method for the GC content computation, the cumulative GC profile. It is shown that the Arabidopsis genome is organized into a mosaic structure of isochores. All the centromeric regions are located in GC-rich isochores, called centromere-isochores, which are characterized by a high GC content but low gene and T-DNA insertion densities. This characteristic distinguishes centromere-isochores from the other class of GC-rich isochores, called GC-isochores, which have high gene and T-DNA insertion densities. Consequently, 15 isochores have been identified, i.e., 7 AT-isochores, 3 GC-isochores, and 5 centromere-isochores. The genes in centromere-isochores, which have the highest GC content, have much shorter intron lengths and lower intron numbers, compared to those of the other two types. There is also considerable difference in the numbers and lengths of transposable elements (TEs) between AT and GC-isochores, i.e., the TE number (length) of AT-isochores is 6.3 (7.3) times that of GC-isochores. It is generally believed that TEs are accumulated in the regions surrounding the centromeres. However, within these TE-rich regions, there are regions of extremely low TE numbers (TE deserts), which correspond to the positions of centromere-isochores. In addition, a heterochromatic knob is located at the boundary of an AT-isochore. Furthermore, we show that the differences in GC content among isochores are mainly due to the GC content variation of introns, the third codon positions and intergenic regions.

Analysis of Variance↗

Position-associated GC asymmetry of gene duplicates.

It is well known that repositioning of a gene often exerts a strong impact on its own expression and whole development. Here we report the results of genome-wide analyses suggesting that repositioning may also radically change the evolutionary fate of gene duplicates. As an indicator of these changes, we used the GC content of gene pairs which originated by duplication. This indicator turned out to be duplicate-asymmetric, which means that genes in a pair differ significantly in GC content despite their apparent origin from a common ancestor. Such an asymmetry necessarily implies that after duplication two originally identical genes mutated in opposite directions-toward GC-rich and GC-poor content, respectively. In mammalian genomes, this trend is definitely associated with presumably methylated hypermutable CpG sites, and in a typical GC-asymmetric gene pair, its two member genes are embedded in GC-contrasting isochores. However, we unexpectedly found similar significant GC asymmetry in fish, fly, worm, and yeast. This means that neither methylation alone nor methylation in combination with isochores can be counted as a primary cause of the GC asymmetry; rather they represent specific realizations of some universal principle of genome evolution. Remarkably, genes from pairs with the greatest GC asymmetry tend to be on different chromosomes, suggesting that the mutational difference between gene duplicates is associated with translocation of a new gene to a different place in the genome, whereas GC symmetric pairs demonstrate the opposite tendency. A recently emerged extra gene copy is usually on the same chromosome as is its parent but quickly, by 0.05 substitution per synonymous site, either has perished or occupies a different chromosome. During this earliest posttranslocation period, the ratio of nonsynonymous/synonymous base substitutions is unusually high, suggesting a rapid adaptive evolution of novel functions. In a general context of evolution by gene duplication, our interpretation of this position-dependent GC asymmetry between duplicated genes is that evolution of redundant genes toward a new function has often been associated with their very early, postduplication repositioning in the genome, with a concomitant abrupt change in epigenetic control of tissue/stage-specific expression and an increase in the mutation rate. Of eight eukaryotic genomes studied, the most distinguished in this respect is the human genome.

Animals↗

Biological soil crusts of sand dunes in Cape Cod National Seashore, Massachusetts, USA.

Biological soil crusts cover hundreds of hectares of sand dunes at the northern tip of Cape Cod National Seashore (Massachusetts, USA). Although the presence of crusts in this habitat has long been recognized, neither the organisms nor their ecological roles have been described. In this study, we report on the microbial community composition of crusts from this region and describe several of their physical and chemical attributes that bear on their environmental role. Microscopic and molecular analyses revealed that eukaryotic green algae belonging to the genera Klebsormidium or Geminella formed the bulk of the material sampled. Phylogenetic reconstruction of partial 16S rDNA sequences obtained from denaturing gradient gel electrophoresis (DGGE) fingerprints also revealed the presence of bacterial populations related to the subclass of the Proteobacteria, the newly described phylum Geothrix/ Holophaga/ Acidobacterium, the Cytophaga/ Flavobacterium/ Bacteroides group, and spirochetes. The presence of these crusts had significant effects on the hydric properties and nutrient status of the natural substrate. Although biological soil crusts are known to occur in dune environments around the world, this study enhances our knowledge of their geographic distribution and suggests a potential ecological role for crust communities in this landscape.

Bacteria↗

Reconsidering the human immunoglobulin heavy-chain locus: 1. An evaluation of the expressed human IGHD gene repertoire.

We have used a bioinformatics approach to evaluate the completeness and functionality of the reported human immunoglobulin heavy-chain IGHD gene repertoire. Using the hidden Markov-model-based iHMMune-align program, 1,080 relatively unmutated heavy-chain sequences were aligned against the reported repertoire. These alignments were compared with alignments to 1,639 more highly mutated sequences. Comparisons of the frequencies of gene utilization in the two databases, and analysis of features of aligned IGHD gene segments, including their length, the frequency with which they appear to mutate, and the frequency with which specific mutations were seen, were used to determine the reliability of alignments to the less commonly seen IGHD genes. Analysis demonstrates that IGHD4-23 and IGHD5-24, which have been reported to be open reading frames of uncertain functionality, are represented in the expressed gene repertoire; however, the functionality of IGHD6-25 must be questioned. Sequence similarities make the unequivocal identification of members of the IGHD1 gene family problematic, although all genes except IGHD1-14*01 appear to be functional. On the other hand, reported allelic variants of IGHD2-2 and of the IGHD3 gene family appear to be nonfunctional, very rare, or nonexistent. Analysis also suggests that the reported repertoire is relatively complete, although one new putative polymorphism (IGHD3-10*p03) was identified. This study therefore confirms a surprising lack of diversity in the available IGHD gene repertoire, and restriction of the germline sequence databases to the functional set described here will substantially improve the accuracy of IGHD gene alignments and therefore the accuracy of analysis of the V-D-J junction.

Alleles↗

Identification and characterization of upstream open reading frames (uORF) in the 5' untranslated regions (UTR) of genes in Saccharomyces cerevisiae.

We have taken advantage of recently sequenced hemiascomycete fungal genomes to computationally identify additional genes potentially regulated by upstream open reading frames (uORFs). Our approach is based on the observation that the structure, including the uORFs, of the post-transcriptionally uORF regulated Saccharomyces cerevisiae genes GCN4 and CPA1 is conserved in related species. Thirty-eight candidate genes for which uORFs were found in multiple species were identified and tested. We determined by 5' RACE that 15 of these 38 genes are transcribed. Most of these 15 genes have only a single uORF in their 5' UTR, and the length of these uORFs range from 3 to 24 codons. We cloned seven full-length UTR sequences into a luciferase (LUC) reporter system. Luciferase activity and mRNA level were compared between the wild-type UTR construct and a construct where the uORF start codon was mutated. The translational efficiency index (TEI) of each construct was calculated to test the possible regulatory function on translational level. We hypothesize that uORFs in the UTR of RPC11, TPK1, FOL1, WSC3, and MKK1 may have translational regulatory roles while uORFs in the 5' UTR of ECM7 and IMD4 have little effect on translation under the conditions tested.

5' Untranslated Regions↗

Generation of an oligonucleotide array for analysis of gene expression in Chlamydomonas reinhardtii.

The availability of genome sequences makes it possible to develop microarrays that can be used for profiling gene expression over developmental time, as organisms respond to environmental challenges, and for comparison between wild-type and mutant strains under various conditions. The desired characteristics of microarrays (intense signals, hybridization specificity and extensive coverage of the transcriptome) were not fully met by the previous Chlamydomonas reinhardtii microarray: probes derived from cDNA sequences (approximately 300 bp) were prone to some nonspecific cross-hybridization and coverage of the transcriptome was only approximately 20%. The near completion of the C. reinhardtii nuclear genome sequence and the availability of extensive cDNA information have made it feasible to improve upon these aspects. After developing a protocol for selecting a high-quality unigene set representing all known expressed sequences, oligonucleotides were designed and a microarray with approximately 10,000 unique array elements (approximately 70 bp) covering 87% of the known transcriptome was developed. This microarray will enable researchers to generate a global view of gene expression in C. reinhardtii. Furthermore, the detailed description of the protocol for selecting a unigene set and the design of oligonucleotides may be of interest for laboratories interested in developing microarrays for organisms whose genome sequences are not yet completed (but are nearing completion).

Animals↗

Analysis of bovine mammary gland EST and functional annotation of the Bos taurus gene index.

Functional genomic studies of the mammary gland require an appropriate collection of cDNA sequences to assess gene expression patterns from the different developmental and operational states of underlying cell types. To better capture the range of gene expression, a normalized cDNA library was constructed from pooled bovine mammary tissues, and 23,202 expressed sequence tags (EST) were produced and deposited into GenBank. Assembly of these EST with sequences in the Bos taurus Gene Index (BtGI) helped to form 5751 of the current 23,883 tentative consensus (TC) sequences. The majority (87%) of these 5751 assemblies contained only one to three mammary-derived EST. In contrast, 18% of the mammary EST assembled with TC sequences corresponding to 12 genes. These results suggest library normalization was only partially effective, because the reduction in EST for genes abundantly transcribed during lactation could be attributed to pooling. For better assessment of novel content in the mammary library and to add to existing annotation of all bovine sequence elements, gene ontology assignments, and comparative sequence analyses against human genome sequence, human and rodent gene indices, and an index of orthologous alignments of genes across eukaryotes (TOGA) were performed, and results were added to existing BtGI annotation. Over 35,000 of the bovine elements significantly matched human genome sequence, and the positions of some alignments (3%) were unique relative to those using human expressed sequences. Because 3445 TC sequences had no significant match with any data set, mammary-derived cDNA clones representing 23 of these elements were analyzed further for expression and novelty. Only one clone met criteria suggesting the corresponding gene was a divergent ortholog or expressed sequence unique to cattle. These results demonstrate that bovine sequence expression data serve as a resource for characterizing mammalian transcriptomes and identifying those genes potentially unique to ruminants.

Animals↗

An interactive bovine in silico SNP database (IBISS).

An interactive bovine in silico SNP (IBISS) database has been created through the clustering and aligning of bovine EST and mRNA sequences. Approximately 324,000 EST and mRNA sequences were clustered to produce 29,965 clusters (producing 48,679 consensus sequences) and 48,565 singletons. A SNP screening regime was placed on variations detected in the multiple sequence alignment files to determine which SNPs are more likely to be real rather than sequencing errors. A small subset of predicted SNPs was validated on a diverse set of bovine DNA samples using PCR amplification and sequencing. Fifty percent of the predicted SNPs in the "putative >1" category were polymorphic in the population sampled. The IBISS database represents more than just a SNP database; it is also a genomic database containing uniformly annotated predicted gene mRNA and protein sequences, gene structure, and genomic organization information.

Animals↗

Characterization of a centromeric marker on mouse chromosome 11 and its introgression in a domesticus/musculus hybrid zone.

It has been proposed that the distribution of Robertsonian chromosome fusions and the Chromosome 11 Nucleolar Organizer Region (NOR) in the Danish hybrid zone between M. m. musculus and M. m. domesticus stems from centromeric incompatibilities between the two subspecies. To test this hypothesis, we identified and characterized a diagnostic subspecific marker closely linked to the centromere on mouse Chromosome 11. Using an allele-specific PCR assay, we investigated the introgression pattern of this centromere in a large sample of mice from a North-South transect of the hybrid zone in Jutland. Domesticus alleles were found to introgress far away from the center of the zone on the musculus side. These results suggest there is no incompatibility between the domesticus centromere of Chromosome 11 in the musculus genomic background.

Animals↗