PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Evolution of tRNA-like sequences and genome variability.

Transfer RNA (tRNA)-like sequences were searched for in the nine basic taxonomic divisions of GenBank-121 (viruses, phages, bacteria, plants, invertebrates, vertebrates, rodents, mammals, and primates) by an original program package implementing a dynamic profile alignment approach for the genetic texts' analysis, in using 22 profiles of tRNAs of different isotypes. In total, 175,901 previously unknown tRNA-like sequences were revealed. The locations of the tRNA-likes were considered over the regions whose functional meaning is described by standard Feature Keys in GenBank. Many regions containing the tRNA-like sequences were recognized as known repeats. A mode of distribution of the tRNA-like sequences in a genome was proposed as expansion in a content of the various transposable elements. An analysis of the integrity of RNA polymerase III inner promoters in the tRNA-like sequences over the GenBank divisions has shown a high possibility of generating new copies of short interspersed nuclear element (SINE) repeats in all divisions, excepting primates. The numerous tRNA-likes found in the regions of RNA polymerase II promoters have suggested an adaptation of RNA polymerase III promoter to a binding of RNA polymerase II.

Algorithms↗

Some microsatellites may act as novel polymorphic cis-regulatory elements through transcription factor binding.

Although microsatellites with functional effects have been described, generally, these repeats are considered as "junk" DNA in the same way as other repetitive sequences. Our aim was to investigate if certain microsatellites can have a functional role as cis-regulatory elements. A database was created of all short tandem repeats, from 2 to 10 bases, located in the first 10-kb 5' of the transcription start sites of all annotated genes of the human genome. Of 114 microsatellites selected based on their size and location in the promoter, 51 were found to be polymorphic. Using electrophoretic mobility shift assay (EMSA), we studied five repetitive motifs and three displayed specific protein binding which were found in 12 of the polymorphic microsatellites. An interesting microsatellite is the CTC/GAG repeat which, as double-stranded (DS) DNA, bound specificity protein 1 (SP1) with high affinity, formed triplexes in vitro and displayed differences in SP1 binding and triplex formation capacity for repeats with distinct numbers of repeat units. Interestingly, the polypyrimidine strand of the repeat (CTC) bound other proteins such as polypyrimidine tract-binding protein 1 (PTBP1) as single-stranded (SS) DNA, and a model with two alternative DNA conformations is proposed for these repeats. Distinct protein binding to DS DNA was also observed for different numbers of AAACA and AAAAT repeats. Our results suggest that certain microsatellites may act as cis-regulatory elements, controlling gene expression through transcription factor binding and/or secondary DNA structure formation. Due to their high polymorphism and abundance, they might represent an important source of quantitative genetic variation.

Base Sequence↗

Genometric analyses of the organization of circular chromosomes: a universal pressure determines the direction of ribosomal RNA genes transcription relative to chromosome replication.

Selective pressures related to gene function and chromosomal architecture are acting on genome sequences and can be revealed, for instance, by appropriate genometric methods. Cumulative nucleotide skew analyses, i.e., GC, TA, and ORF orientation skews, predict the location of the origin of DNA replication for 88 out of 100 completely sequenced bacterial chromosomes. These methods appear fully reliable for proteobacteria, Gram-positives, and spirochetes as well as for euryarchaeotes. Based on this genome architecture information, coorientation analyses reveal that in prokaryotes, ribosomal RNA (rRNA) genes encoding the small and large ribosomal subunits are all transcribed in the same direction as DNA replication; that is, they are located along the leading strand. This result offers a simple and reliable method for circumscribing the region containing the origin of the DNA replication and reveals a strong selective pressure acting on the orientation of rRNA genes similar to the weaker one acting on the orientation of ORFs. Rate of coorientation of transfer RNA (tRNA) genes with DNA replication appears to be taxon-specific. Analyzing nucleotide biases such as GC and TA skews of genes and plotting one against the other reveals a taxonomic clusterization of species. All ribosomal RNA genes are enriched in Gs and depleted in Cs, the only so far known exception being the rRNA genes of deuterostomian mitochondria. However, this exception can be explained by the fact that in the chromosome of the human mitochondrion, the model of the deuterostomian organelle genome, DNA replication, and rRNA transcription proceed in opposite directions. A general rule is deduced from prokaryotic and mitochondrial genomes: ribosomal RNA genes that are transcribed in the same direction as the DNA replication are enriched in Gs, and those transcribed in the opposite direction are depleted in Gs.

Base Composition↗

Multiple novel transcription initiation sites for NRG1.

The large neuregulin 1 gene (NRG1) has been mapped to a 1.125 Mb region on chromosome 8p11-21. Three major forms of NRG1 (types I-III), all with distinct amino-termini encoded by unique 5'-exons, have been described. We report here the discovery of nine novel NRG1 exons, including six alternative 5'-exons, increasing the number of potential promoters in NRG1 from three to nine. The novel transcripts of NRG1 described here use the novel 5'-exons which are either coding or non-coding. The functional relevance of the predicted proteins they encode has not been evaluated. Three of the novel 5'-exons are well conserved in syntenic rat and mouse sequences; they encode proteins with novel amino-termini, here termed types IV-VI. NRG1 plays a central role in neural development and is most likely involved in regulation of synaptic plasticity, or how the brain responds or adapts to the environment. The unusually complex gene structure may facilitate spatial and temporal regulation of NRG1 expression, fine-tune NRG1 protein function at different stages during development of the nervous system, and adapt responses to the environment in the adult brain.

5' Flanking Region↗

Rapid turnover and species-specificity of vomeronasal pheromone receptor genes in mice and rats.

Pheromones are used by individuals of the same species to elicit behavioral or physiological changes, and they are perceived primarily by the vomeronasal organ (VNO) in terrestrial vertebrates. VNO pheromone receptors are encoded by the V1r and V2r gene superfamilies in mammals. A comparison of the V1r and V2r repertoires between closely related species can provide significant insights into the evolutionary genetic mechanisms responsible for species-specific pheromone communications. A total of 137 putatively functional V1r genes of 12 families were previously identified from the mouse genome. We report the identification of 95 putatively functional V1r genes from the draft rat genome sequence. These genes map primarily to four blocks in two chromosomes. The rat V1r genes can be phylogenetically grouped into 10 families, which are shared with mouse, and 2 new families, which are rat-specific. Even in many shared families, gene numbers differ between the two species, apparently due to frequent gene duplication and pseudogenization after the separation of the two species. Molecular dating suggests that most of the rat V1r families emerged before or during the radiation of mammalian orders, but many duplications within families occurred as recently as in the past 10 million years (MY). Our results show that the evolution of the V1r repertoire is characterized by exceptionally fast gene turnover via gains and losses of individual genes, suggesting rapid and substantial changes in pheromone communication between species.

Animals↗

Different age distribution patterns of human, nematode, and Arabidopsis duplicate genes.

We studied the age distribution of duplicate genes in each of four eukaryotic genomes: human, Arabidopsis thaliana, Caenorhabditis elegans, and Drosophila melanogaster. The four distributions differ greatly from each other, contrary to the previous proposal of a universal L-shaped distribution in all eukaryotic genomes studied. Indeed, only the distribution in humans is L-shaped. The distribution in Arabidopsis is consistent with the hypothesis of an ancient genome duplication with no recent burst of duplication events, while the distribution in C. elegans is nearly uniform. We also applied a nonparametric method to the human distribution to show that the rate of loss of duplicate genes decreases over time, contrary to the proposal of an exponential decay. One possible explanation of the decreasing rate of loss of duplicate genes over time could be rapid functional divergence between duplicate genes, providing an advantage for the retention of both duplicates.

Algorithms↗

Mutation and selection on the anticodon of tRNA genes in vertebrate mitochondrial genomes.

The H-strand of vertebrate mitochondrial DNA is left single-stranded for hours during the slow DNA replication. This facilitates C-->U mutations on the H-strand (and consequently G-->A mutations on the L-strand) via spontaneous deamination which occurs much more frequently on single-stranded than on double-stranded DNA. For the 12 coding sequences (CDS) collinear with the L-strand, NNY synonymous codon families (where N stands for any of the four nucleotides and Y stands for either C or U) end mostly with C, and NNR and NNN codon families (where R stands for either A or G) end mostly with A. For the lone ND6 gene on the other strand, the codon bias is the opposite, with NNY codon families ending mostly with U and NNR and NNN codon families ending mostly with G. These patterns are consistent with the strand-specific mutation bias. The codon usage biased towards C-ending and A-ending in the 12 CDS sequences affects the codon-anticodon adaptation. The wobble site of the anticodon is always G for NNY codon families dominated by C-ending codons and U for NNR and NNN codon families dominated by A-ending codons. The only, but consistent, exception is the anticodon of tRNA-Met which consistently has a 5'-CAU-3' anticodon base-pairing with the AUG codon (the translation initiation codon) instead of the more frequent AUA. The observed CAU anticodon (matching AUG) would increase the rate of translation initiation but would reduce the rate of peptide elongation because most methionine codons are AUA, whereas the unobserved UAU anticodon (matching AUA) would increase the elongation rate at the cost of translation initiation rate. The consistent CAU anticodon in tRNA-Met suggests the importance of maximizing the rate of translation initiation.

Animals↗

Estimation of ancestral gene set of bilaterian animals and its implication to dynamic change of gene content in bilaterian evolution.

To understand the process of bilaterian evolution, we estimated ancestral gene sets at the split of plant-animal-fungi and the divergence of bilaterian animals and from 1,236,790 non-redundant genes. We, then, examined how the numbers of the gene clusters have changed since the split. As a result, we estimated the numbers of gene clusters in the ancestral gene sets of plant-animal-fungi and bilaterian animals to be at least 2469 and 6577, respectively. Thus, we found a 2.7-fold increase in the number of gene clusters during the period from the evolutionary split of plant-animal-fungi to the divergence of bilaterian animals. Moreover, when we compared these numbers of ancestral gene clusters with those of extant animals such as the nematode, fly, mouse and human, we found that the extant bilaterian animals have retained more than 3500 gene clusters of the ancestral gene set, and have lost more than 1600 gene clusters. It suggests that these processes of genomic diversification provided bilaterian animals with molecular basis for species diversity.

Animals↗

The silk moth Bombyx mori U1 and U2 snRNA variants are differentially expressed.

Five U1 and eight U2 isoforms of the silk moth Bombyx mori exhibiting internal nucleotide differences have been previously identified and characterized in various tissues and developmental stages. In this investigation, it is demonstrated that the levels of some snRNA variants differ in egg and silk gland tissue and change during development. Qualitative and quantitative differences in the U1 and U2 variant populations were observed at three developmental points (early, middle and late) of the silk gland (SG) during the fifth instar larval stage of the silk moth. Statistical analyses of the various isoform populations across the fifth instar larval and egg stages show significant differences for some of the U1 and U2 variants. The representation of variant sequences in expressed U1 and U2 sequences (RT-PCR libraries) and in a whole-genome shotgun (WGS) assembly database was confirmed. In addition, conserved elements in the promoter 5'-flanking region of the U1 and U2 variants were identified in the WGS.

Animals↗

Gene splice sites correlate with nucleosome positions.

Gene sequences in the vicinity of splice sites are found to possess dinucleotide periodicities, especially RR and YY, with the period close to the pitch of nucleosome DNA. This confirms previously reported findings about preferential positioning of splice junctions within the nucleosomes. The RR and YY dinucleotides oscillate counter-phase, i.e., their respective preferred positions are shifted about half-period from one another, as it was observed earlier for AA and TT dinucleotides. Species specificity of nucleosome positioning DNA pattern is indicated by the predominant use of the periodical GG(CC) dinucleotides in human and mouse genes, as opposed to predominant AA(TT) dinucleotides in Arabidopsis and C. elegans.

Alternative Splicing↗

Expressed sequence tags from the laboratory-grown miniature tomato (Lycopersicon esculentum) cultivar Micro-Tom and mining for single nucleotide polymorphisms and insertions/deletions in tomato cultivars.

Laboratory-grown miniature tomato (Lycopersicon esculentum) cultivar Micro-Tom has attracted attention as a host for functional genomics research. In this study, we generated 35,824 expressed sequence tags (ESTs) from leaves and fruits of Micro-Tom. The ESTs comprised 10,287 unigenes (5007 contigs and 5280 singletons), including 1858 novel tomato unigenes. Of the 18 unigenes that shared strong homology with tobacco chloroplast genome sequences, one unigene was likely derived from polyadenylated transcripts of the atpH gene. Interestingly, ESTs for vacuolar invertase, pectate lyase and alcohol acyl transferase were underrepresented in the Micro-Tom data set. From all of the ESTs, we mined 2039 candidate single nucleotide polymorphisms (SNPs) and 121 candidate insertions and deletions (indels) based on homology with four tomato inbred lines, E6203, R11-13, Rio Grande PtoR and R11-12, and a wild relative, L. pennellii TA56, for which sequence data was publicly available with more than 5000 entries. Direct genome sequencing of several SNP or indel sites in Micro-Tom and L. esculentum E6203 suggested that more than 69% of the candidate sites were truly polymorphic, making them useful for the preparation of DNA markers.

DNA, Complementary↗

An investigation of the variation in the transition bias among various animal mitochondrial DNA.

The transition:transversion ratio (ts/tv) is known to be very high in human mitochondrial DNA, but we have little information about this ratio in other species. Here we investigate the transition bias in animal mitochondrial DNA using single nucleotide polymorphism data at four-fold degenerate sites. We investigate this pattern of polymorphism in the cytochrome b gene (cyt-b) in 70 species using a total of 1823 mutations. We show that most species show a bias towards transitions but that the ratio varies significantly between species. There is little evidence for variation within orders or genera and between closely related species such as the great apes. The majority of the variation appears to be at a higher phylogenetic levels: between orders and classes. We test whether the variation in ts/tv ratio could be due to variation in the metabolic rate by considering whether the ratio is correlated to base composition. We find no evidence that the metabolic rate affects the ts/tv ratio. We also investigate the relative frequencies of C to T or T to C (C<-->T) mutations and A to G or G to A (A<-->G) mutations. We show that overall they occur at significantly different frequencies, and that there is significant variation in their relative frequency between species and between classes. We find no evidence in support of the hypothesis that this variation could be due to different metabolic rates.

Animals↗

Automatic gene collection system for genome-scale overview of G-protein coupled receptors in eukaryotes.

We have developed an automatic system for identifying GPCR (G-protein coupled receptor) genes from various kinds of genomes, which is finally deposited in the SEVENS database (http://sevens.cbrc.jp/), by integrating such software as a gene finder, a sequence alignment tool, a motif and domain assignment tool, and a transmembrane helix predictor. SEVENS enables us to perform a genome-scale overview of the "GPCR universe" using sequences that are identified with high accuracy (99.4% sensitivity and 96.6% specificity). Using this system, we surveyed the complete genomes of 7 eukaryotes and 224 prokaryotes, and found that there are 4 to 1016 GPCR genes in the 7 eukaryotes, and only a total of 16 GPCR genes in all the prokaryotes. Our preliminary results indicate that 11 subfamilies of the Class A family, the Class 2(B) family, the Class 3(C) family and the fz/smo family are commonly found among human, fly, and nematode genomes. We also analyzed the chromosomal locations of the GPCR genes with the Kolmogorov-Smirnov test, and found that species-specific families, such as olfactory, taste, and chemokine receptors in human and nematode chemoreceptor in worm, tend to form clusters extensively, whereas no significant clusters were detected in fly and plant genomes.

Animals↗

Characterization and prediction of alternative splice sites.

Human alternative isoform, cryptic, skipped, and constitutive splice sites from the ALTEXTRON database were analysed regarding splice site strength, composition, GC content, position and binding site strength of polypyrimidine tract and branch site. Several features were identified which distinguish alternative isoform and cryptic splice sites, but not skipped splice sites from constitutive ones. These include splice site strength, introns GC content, U2AF35 binding site score, and oligonucleotide frequencies. For the predictive classification of splice sites, pattern recognition models for different splicing factor binding sites and oligonucleotide frequency models (OFMs) were combined using backpropagation networks. 67.45% of acceptor sites and 71.23% of donor sites are correctly classified by networks trained for classification of constitutive and alternative isoform/cryptic splice sites. A web-application for the prediction of alternative splice sites is available at http://es.embnet.org/~mwang/assp.html .

Alternative Splicing↗

Human chromosome 21/Down syndrome gene function and pathway database.

Down syndrome, trisomy of human chromosome 21, is the most common genetic cause of intellectual disability. Correlating the increased expression, due to gene dosage, of the >300 genes encoded by chromosome 21 with specific phenotypic features is a goal that becomes more feasible with the increasing availability of large scale functional, expression and evolutionary data. These data are dispersed among diverse databases, and the variety of formats and locations, plus their often rapid growth, makes access and assimilation a daunting task. To aid the Down syndrome and chromosome 21 community, and researchers interested in the study of any chromosome 21 gene or ortholog, we are developing a comprehensive chromosome 21-specific database with the goals of (i) data consolidation, (ii) accuracy and completeness through expert curation, and (iii) facilitation of novel hypothesis generation. Here we describe the current status of data collection and the immediate future plans for this first human chromosome-specific database.

Base Sequence↗

EST-based identification of genes expressed in the liver of adult seabass (Dicentrarchus labrax, L.).

The scarcity of the genomic resources for some fish species, in spite of their commercial interest, could retard the positive effects that modern biotechnology can offer to aquaculture industry. Then an effort should be made to reduce, as far as it concerns genomic resources, the gap that separates farming species from "model organisms". In this paper, we present an EST project in which we performed single pass sequencing on 1229 randomly selected clones from a sea bass cDNA library. The sequences are deposited in the NCBI database with the following accession numbers: from , from , from and from . EST cataloguing and profiling of seabass will set the basis for functional genomic research in this species, but will also serve for comparative and environmental genomics, for the identification of polymorphic markers useful, for example, to survey the disease resistance of fish, for the discovery of new molecular markers of exposure and for the production of micro- and macro-arrays.

Animals↗

Penaeus monodon gene discovery project: the generation of an EST collection and establishment of a database.

A large-scale expressed sequence tag (EST) sequencing project was undertaken for the purpose of gene discovery in the black tiger shrimp Penaeus monodon. Initially, 15 cDNA libraries were constructed from different tissues (eyestalk, hepatopancrease, haematopoietic tissue, haemocyte, lymphoid organ, and ovary) of shrimp, reared under normal or stress conditions, to identify tissue-specific genes and genes responding to infection and heat stress. A total of 10,100 clones were analyzed by single-pass sequencing from the 5' end. Clustering and assembling of these ESTs resulted in a total of 4845 unique sequences with 917 overlapping contigs and 3928 singletons. The redundancy of each cDNA library ranged from 13.4% to 61.3% with an overall redundancy of 61.1%. About half of these ESTs (2365 clones, 48.8%) showed significant homology (BLASTX, e-values <10(-4)) to known genes. A high proportion of P. monodon ESTs was most similar to the predicted protein sequences from various organisms, e.g. Homo sapiens (9%), Mus musculus (7%), Drosophila (6%), Gallus sp.(6%), and Anopheles (5%). Only 6% showed the highest similarity to other known genes from shrimp due to the limited sequence entries of the species in the public database. Several tissue-specific transcripts were identified as well as the candidate genes that may be implicated in the immune response. In addition, bioinformatic mining of microsatellites from the P. monodon ESTs identified 997 unique microsatellite containing ESTs in which 74 loci resided within the genes of known functions. Consequently, the P. monodon EST database was established. The EST sequence data and the BLAST results were stored and made available through a web-accessible database (). This EST database provides a useful resource for gene identification and functional genomic studies of shrimp.

Animals↗