PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

The presence of a haloarchaeal type tyrosyl-tRNA synthetase marks the opisthokonts as monophyletic.

Lateral gene transfer plays an important role in the evolution of life. Events of ancient gene transfer can transmit genetic novelties to descendent lineages and subsequently shape their genetic systems. We here present the analyses of the gene encoding tyrosyl-tRNA synthetase (tyrRS), which reveal two eukaryotic tyrRS lineages, one including the opisthokonts and the other the remaining eukaryotes. The different origins of tyrRS lineages between the opisthokonts and the remaining eukaryotes indicate a likely case of ancient lateral gene transfer of tyrRS from an archaeon to the opisthokonts, which lends further support for the monophyly of the latter group. Ancient paralogy followed by differential gene loss is an alternative, albeit less parsimonious explanation for the distribution of the two eukaryotic tyrRS types. In either case, the presence of a haloarchaeal tyrRS type in the opisthokonts marks this group as monophyletic. This finding also points to the potential utility of ancient gene transfer events as molecular markers for major organismal lineages.

Amino Acid Sequence↗

Database resources of the National Center for Biotechnology Information.

In addition to maintaining the GenBank nucleic acid sequence database, the National Center for Biotechnology Information (NCBI) provides data analysis and retrieval resources that operate on the data in GenBank and a variety of other biological data made available through NCBI's Web site. NCBI data retrieval resources include Entrez, PubMed, LocusLink and the Taxonomy Browser. Data analysis resources include BLAST, Electronic PCR, OrfFinder, RefSeq, UniGene, HomoloGene, Database of Single Nucleotide Polymorphisms (dbSNP), Human Genome Sequencing, Human MapViewer, GeneMap'99, Human-Mouse Homology Map, Cancer Chromosome Aberration Project (CCAP), Entrez Genomes, Clusters of Orthologous Groups (COGs) database, Retroviral Genotyping Tools, Cancer Genome Anatomy Project (CGAP), SAGEmap, Gene Expression Omnibus (GEO), Online Mendelian Inheri-tance in Man (OMIM), the Molecular Modeling Database (MMDB) and the Conserved Domain Database (CDD). Augmenting many of the Web applications are custom implementations of the BLAST program optimized to search specialized data sets. All of the resources can be accessed through the NCBI home page at: http://www.ncbi.nlm.nih. gov.

Animals↗

Database resources of the National Center for Biotechnology.

In addition to maintaining the GenBank(R) nucleic acid sequence database, the National Center for Biotechnology Information (NCBI) provides data analysis and retrieval resources for the data in GenBank and other biological data made available through NCBI's Web site. NCBI resources include Entrez, PubMed, PubMed Central (PMC), LocusLink, the NCBITaxonomy Browser, BLAST, BLAST Link (BLink), Electronic PCR (e-PCR), Open Reading Frame (ORF) Finder, References Sequence (RefSeq), UniGene, HomoloGene, ProtEST, Database of Single Nucleotide Polymorphisms (dbSNP), Human/Mouse Homology Map, Cancer Chromosome Aberration Project (CCAP), Entrez Genomes and related tools, the Map Viewer, Model Maker (MM), Evidence Viewer (EV), Clusters of Orthologous Groups (COGs) database, Retroviral Genotyping Tools, SAGEmap, Gene Expression Omnibus (GEO), Online Mendelian Inheritance in Man (OMIM), the Molecular Modeling Database (MMDB), the Conserved Domain Database (CDD), and the Conserved Domain Architecture Retrieval Tool (CDART). Augmenting many of the Web applications are custom implementations of the BLAST program optimized to search specialized data sets. All of the resources can be accessed through the NCBI home page at: http://www.ncbi.nlm.nih.gov.

Animals↗

Database resources of the National Center for Biotechnology Information.

In addition to maintaining the GenBank nucleic acid sequence database, the National Center for Biotechnology Information (NCBI) provides data retrieval systems and computational resources for the analysis of data in GenBank and other biological data made available through NCBI's website. NCBI resources include Entrez, Entrez Programming Utilities, PubMed, PubMed Central, Entrez Gene, the NCBI Taxonomy Browser, BLAST, BLAST Link (BLink), Electronic PCR, OrfFinder, Spidey, RefSeq, UniGene, HomoloGene, ProtEST, dbMHC, dbSNP, Cancer Chromosomes, Entrez Genomes and related tools, the Map Viewer, Model Maker, Evidence Viewer, Clusters of Orthologous Groups (COGs), Retroviral Genotyping Tools, HIV-1/Human Protein Interaction Database, SAGEmap, Gene Expression Omnibus (GEO), Online Mendelian Inheritance in Man (OMIM), the Molecular Modeling Database (MMDB), the Conserved Domain Database (CDD) and the Conserved Domain Architecture Retrieval Tool (CDART). Augmenting many of the Web applications are custom implementations of the BLAST program optimized to search specialized datasets. All of the resources can be accessed through the NCBI home page at http://www.ncbi.nlm.nih.gov.

Amino Acid Sequence↗

Evidence for the presence of a cellulase gene in the last common ancestor of bilaterian animals.

Until recently, the textbook view of cellulose hydrolysis in animals was that gut-resident symbiotic organisms such as bacteria or unicellular eukaryotes are responsible for the cellulases produced. This view has been challenged by the characterization and sequencing of endogenous cellulase genes from some invertebrate animals, including plant-parasitic nematodes, arthropods and a mollusc. Most of these genes are completely unrelated in terms of sequence, and their evolutionary origins remain unclear. In the case of plant-parasitic nematodes, it has been suggested that their ancestor obtained a cellulase gene via horizontal gene transfer from a prokaryote, and similar suggestions have been made about a cellulase gene recently discovered in a sea squirt. To improve understanding about the evolution of animal cellulases, we searched for all known types of these enzymes in GenBank, and performed phylogenetic comparisons. Low phylogenetic resolution was found among most of the sequences examined, however, positional identity in the introns of cellulase genes from a termite, a sea squirt and an abalone provided compelling evidence that a similar gene was present in the last common ancestor of protostomes and deuterostomes. In a different enzyme family, cellulases from beetles and plant-parasitic nematodes were found to cluster together. This result questions the idea of lateral gene transfer into the ancestors of the latter, although statistical tests did not allow this possibility to be ruled out. Overall, our results suggest that at least one family of endogenous cellulases may be more widespread in animals than previously thought.

Amino Acid Sequence↗

A molecular compendium of genes expressed in multiple myeloma.

We have created a molecular resource of genes expressed in primary malignant plasma cells using a combination of cDNA library construction, 5' end single-pass sequencing, bioinformatics, and microarray analysis. In total, we identified 9732 nonredundant expressed genes. This dataset is available as the Myeloma Gene Index (www.uhnres.utoronto.ca/akstewart_lab).Predictably, the sequenced profile of myeloma cDNAs mirrored the known function of immunoglobulin-producing, high-respiratory rate, low-cycling, terminally differentiated plasma cells. Nevertheless, approximately 10% of myeloma-expressed sequences matched only entries in the database of Expressed Sequence Tags (dbEST) or the high-throughput genomic sequence (htgs) database. Numerous novel genes of potential biologic significance were identified. We therefore spotted 4300 sequenced cDNAs on glass slides creating a myeloma-enriched microarray. Several of the most highly expressed genes identified by sequencing, such as a novel putative disulfide isomerase (MGC3178), tumor rejection antigen TRA1, heat shock 70-kDa protein 5, and annexin A2, were also differentially expressed between myeloma and B lymphoma cell lines using this myeloma-enriched microarray. Furthermore, a defined subset of 34 up-regulated and 18 down-regulated genes on the array were able to differentiate myeloma from nonmyeloma cell lines. These not only include genes involved in B-cell biology such as syndecan, BCMA, PIM2, MUM1/IRF4, and XBP1, but also novel uncharacterized genes matching sequences only in the public databases. In summary, our expressed gene catalog and myeloma-enriched microarray contains numerous genes of unknown function and may complement other commercially available arrays in defining the molecular portrait of this hematopoietic malignancy. GenBank Accession numbers include BF169967-BF176369, BF185966-BF185969, and BF177280-BF177455.

Amino Acid Sequence↗

MRP8, a new member of ABC transporter superfamily, identified by EST database mining and gene prediction program, is highly expressed in breast cancer.

BACKGROUND: With the completion of the human draft genome sequence, efforts are now devoted to identifying new genes. We have developed a computer-based strategy that utilizes the EST database to identify new genes that could be targets for the immunotherapy of cancer or could be involved in the multistep process of cancer. MATERIALS AND METHODS: Utilizing our computer-based screening strategy, we identified a cluster of expressed sequence tags (ESTs) that are highly expressed in breast cancer. Northern blot and reverse transcriptase polymerase chain reaction (RT-PCR) analyses demonstrated the tissue specificity of the computer-generated cluster and comparison with the human genome sequence assisted in isolating a full-length cDNA clone. RESULTS: We identified a new gene that is highly expressed in breast cancer. This gene is expressed at moderate levels in normal breast and testis and at very low levels in liver, brain, and placenta. The gene has two major transcripts of 4.5 kb and 4.1 kb. The 4.5-kb transcript is very abundant in breast cancer, and has an open reading frame of 1382 amino acids. The predicted protein sequence of the 4.5-kb transcript reveals that it has high homology with MRP5, a member of multidrug resistant-associated protein family (MRP). There are seven reported members in the MRP family; we designate this gene as MRP8 (ABCC11). The 4.5-kb MRP8 transcript consists of 31 exons and is located in a genomic region of over 80.4 kb on chromosome 16q12.1. The smaller 4.1-kb transcript of MRP8 is found in testis and may initiate within intron 6 of the gene. CONCLUSION: The selective expression of MRP8 (ABCC11), a new member of ATP-binding cassette transporter superfamily could be a molecular target for the treatment of breast cancer.

ATP-Binding Cassette Transporters↗

Identification and analysis of Arabidopsis expressed sequence tags characteristic of non-coding RNAs.

Sequencing of the Arabidopsis genome has led to the identification of thousands of new putative genes based on the predicted proteins they encode. Genes encoding tRNAs, ribosomal RNAs, and small nucleolar RNAs have also been annotated; however, a potentially important class of genes has largely escaped previous annotation efforts. These genes correspond to RNAs that lack significant open reading frames and encode RNA as their final product. Accumulating evidence indicates that such "non-coding RNAs" (ncRNAs) can play critical roles in a wide range of cellular processes, including chromosomal silencing, transcriptional regulation, developmental control, and responses to stress. Approximately 15 putative Arabidopsis ncRNAs have been reported in the literature or have been annotated. Although several have homologs in other plant species, all appear to be plant specific, with the exception of signal recognition particle RNA. Conversely, none of the ncRNAs reported from yeast or animal systems have homologs in Arabidopsis or other plants. To identify additional genes that are likely to encode ncRNAs, we used computational tools to filter protein-coding genes from genes corresponding to 20,000 expressed sequence tag clones. Using this strategy, we identified 19 clones with characteristics of ncRNAs, nine putative peptide-coding RNAs with open reading frames smaller than 100 amino acids, and 11 that could not be differentiated between the two categories. Again, none of these clones had homologs outside the plant kingdom, suggesting that most Arabidopsis ncRNAs are likely plant specific. These data indicate that ncRNAs represent a significant and underdeveloped aspect of Arabidopsis genomics that deserves further study.

Algorithms↗

Database and structural characterization of intermolecular interactions in nucleic acid and protein complex.

Protein-nucleic acid interactions and manners in which they control cellular communication are interesting subjects. Structural interaction data based on three-dimensional structure provide the valuable information for understanding these interactions. The database cataloging the interaction motifs of nucleic acid moieties has been developed. DNA/RNA molecules with the specific three-dimensional structure, express the specific structural and biological functions. The polymorphic nature of recognition interactions and/or motifs in protein-nucleic acid complex are generated from a variety of week forces, such as hydrogen bonds and hydrophobic interactions. The geometrical parameters about hydrogen bond and base stacking would provide the tolerant aspect for these characterizations. The tentative database including these geometrical parameters with the species and physical properties of the surrounding amino acids has been constructed. The user can obtain the selected list with several descriptors for the structure definition and sequence properties.

Databases, Nucleic Acid↗

A common philosophy and FORTRAN 77 software package for implementing and searching sequence databases.

I present a common philosophy for implementing the EMBL and GENBANK (BBN-Los Alamos) nucleic acid sequence databases, as well as the National Biological Foundation (Dayhoff) protein sequence database. The associated FORTRAN 77 fully transportable software package includes: 1) modules for implementing each of these databases from the initial magnetic tape file, 2) modules performing a fast mnemonic access, 3) modules performing key-string access and allowing the definition of user-specific database subsets, 4) a common probe searching module allowing the stacking of multiple combined search requests over the databases. This software is particularly suitable for 32-bit mini/microcomputers but would eventually run on 16-bit computers.

Amino Acid Sequence↗

Database resources of the National Center for Biotechnology Information.

In addition to maintaining the GenBank(R) nucleic acid sequence database, the National Center for Biotechnology Information (NCBI) provides data analysis and retrieval and resources that operate on the data in GenBank and a variety of other biological data made available through NCBI's Web site. NCBI data retrieval resources include Entrez, PubMed, LocusLink and the Taxonomy Browser. Data analysis resources include BLAST, Electronic PCR, OrfFinder, RefSeq, UniGene, Database of Single Nucleotide Polymorphisms (dbSNP), Human Genome Sequencing pages, GeneMap'99, Davis Human-Mouse Homology Map, Cancer Chromosome Aberration Project (CCAP) pages, Entrez Genomes, Clusters of Orthologous Groups (COGs) database, Retroviral Genotyping Tools, Cancer Genome Anatomy Project (CGAP) pages, SAGEmap, Online Mendelian Inheritance in Man (OMIM) and the Molecular Modeling Database (MMDB). Augmenting many of the Web applications are custom implementations of the BLAST program optimized to search specialized data sets. All of the resources can be accessed through the NCBI home page at: http://www.ncbi.nlm.nih. gov

Animals↗

Knowledge discovery in GenBank.

We describe various methods designed to discover knowledge in the GenBank nucleic acid sequence database. Using a grammatical model of gene structure, we create a parse tree of a gene using features listed in the FEATURE TABLE. The parse tree infers features that are not explicitly listed, but which follow from the listed features. This method discovers 30% more introns and 40% more exons when applied to a globin gene subset of GenBank. Parse tree construction also entails resolving ambiguity and inconsistency within a FEATURE TABLE. We transform the parse tree into an augmented FEATURE TABLE that represents inferred gene structure explicitly and unambiguously, thereby greatly improving the utility of the FEATURE TABLE to researchers. We then describe various analogical reasoning techniques designed to exploit the homologous nature of genes. We build a classification hierarchy that reflects the evolutionary relationship between genes. Descriptive grammars of gene classes are then induced from the instance grammars of genes. Case based reasoning techniques use these abstract gene class descriptions to predict the presence and location of regulatory features not listed in the FEATURE TABLE. A cross-validation test shows a success rate of 87% on a globin gene subset of GenBank.

Algorithms↗

A mechanism for maintaining an up-to-date GenBank database via Usenet.

In this paper, we describe an automated system for distributing updates to the GenBank nucleic acid sequence database, using the Usenet news system as the underlying transport mechanism. Our system allows new loci to be distributed as soon as the sequences are available, over existing networks, using existing Usenet software and infrastructure currently available on a wide range of computer systems.

Amino Acid Sequence↗

Database resources of the National Center for Biotechnology Information: 2002 update.

In addition to maintaining the GenBank nucleic acid sequence database, the National Center for Biotechnology Information (NCBI) provides data analysis and retrieval resources that operate on the data in GenBank and a variety of other biological data made available through NCBI's web site. NCBI data retrieval resources include Entrez, PubMed, LocusLink and the Taxonomy Browser. Data analysis resources include BLAST, Electronic PCR, OrfFinder, RefSeq, UniGene, HomoloGene, Database of Single Nucleotide Polymorphisms (dbSNP), Human Genome Sequencing, Human MapViewer, Human inverted exclamation markVMouse Homology Map, Cancer Chromosome Aberration Project (CCAP), Entrez Genomes, Clusters of Orthologous Groups (COGs) database, Retroviral Genotyping Tools, SAGEmap, Gene Expression Omnibus (GEO), Online Mendelian Inheritance in Man (OMIM), the Molecular Modeling Database (MMDB) and the Conserved Domain Database (CDD). Augmenting many of the web applications are custom implementations of the BLAST program optimized to search specialized data sets. All of the resources can be accessed through the NCBI home page at http://www.ncbi.nlm.nih.gov.

Amino Acid Sequence↗

Database resources of the National Center for Biotechnology Information: update.

In addition to maintaining the GenBank(R) nucleic acid sequence database, the National Center for Biotechnology Information (NCBI) provides data analysis and retrieval resources for the data in GenBank and other biological data made available through NCBI's website. NCBI resources include Entrez, PubMed, PubMed Central, LocusLink, the NCBI Taxonomy Browser, BLAST, BLAST Link (BLink), Electronic PCR, OrfFinder, Spidey, RefSeq, UniGene, HomoloGene, ProtEST, dbMHC, dbSNP, Cancer Chromosome Aberration Project (CCAP), Entrez Genomes and related tools, the Map Viewer, Model Maker, Evidence Viewer, Clusters of Orthologous Groups (COGs) database, Retroviral Genotyping Tools, SARS Coronavirus Resource, SAGEmap, Gene Expression Omnibus (GEO), Online Mendelian Inheritance in Man (OMIM), the Molecular Modeling Database (MMDB), the Conserved Domain Database (CDD) and the Conserved Domain Architecture Retrieval Tool (CDART). Augmenting many of the web applications are custom implementations of the BLAST program optimized to search specialized data sets. All of the resources can be accessed through the NCBI home page at: http://www.ncbi.nlm.nih.gov.

Animals↗