PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

An infrastructure for comparative genomics to functionally characterize genes and proteins.

Current genome projects are resulting in a flood of sequence data. The interpretation of these sequences is lagging, and optimized data analysis strategies need to be developed. Much can be learned from comparing different genomes, as genomes of distant organisms may still encode proteins with high sequence similarity. The order of genes (co linearity) in genomes may also be conserved to some extend. We have employed both these observations to create a multi-functional, computational analysis system (genomeSCOUT) which allows for rapid identification and functional characterization of genes and proteins through genome comparison. With a number of independent algorithms, information about different levels of protein homology (concerning e.g. paralogs, orthologs and clusters of orthologous groups, COGs) and gene order is collected and stored in several value added databases. These databases are then used for interactive comparison of genomes and subsequent analysis. The application is based on the well established data integration system SRS. This ensures (1) fast handling of large genomic data sets, (2) straightforward access to a multitude of biological databases, (3) unique linking functions between these databases, (4) highly efficient collection of information on genes and proteins, and 5. fully integrated and user friendly graphical representations of search results. This application can be used for projects as diverse as the correct annotation of genomes, the optimization of (micro) organisms for industrial production, or the identification of drug targets.

Computational Biology↗

Biological sequence compression algorithms.

Today, more and more DNA sequences are becoming available. The information about DNA sequences are stored in molecular biology databases. The size and importance of these databases will be bigger and bigger in the future, therefore this information must be stored or communicated efficiently. Furthermore, sequence compression can be used to define similarities between biological sequences. The standard compression algorithms such as gzip or compress cannot compress DNA sequences, but only expand them in size. On the other hand, CTW (Context Tree Weighting Method) can compress DNA sequences less than two bits per symbol. These algorithms do not use special structures of biological sequences. Two characteristic structures of DNA sequences are known. One is called palindromes or reverse complements and the other structure is approximate repeats. Several specific algorithms for DNA sequences that use these structures can compress them less than two bits per symbol. In this paper, we improve the CTW so that characteristic structures of DNA sequences are available. Before encoding the next symbol, the algorithm searches an approximate repeat and palindrome using hash and dynamic programming. If there is a palindrome or an approximate repeat with enough length then our algorithm represents it with length and distance. By using this preprocessing, a new program achieves a little higher compression ratio than that of existing DNA-oriented compression algorithms. We also describe new compression algorithm for protein sequences.

Algorithms↗

Sequence analysis of HIV-1 insertion sites in peripheral blood lymphocytes.

An essential component of the HIV-1 life cycle involves insertion in the genome of an infected cell. The site of HIV-1 integration has the potential to disrupt a gene and perturb a normal cellular function. To begin to address whether disease pathogenesis may correlate with the site of insertion, flanking cellular sequences at these HIV integrated regions were directly amplified from peripheral blood mononuclear cells DNA from a broad range of infected individuals using an inverse polymerase chain reaction strategy. Amplified flanking regions were sequenced and examined for similarity to the nucleic acid database. In this group of analyzed samples, the HIV-1 provirus was inserted within non-coding regions throughout the genome of the infected host, in which 7/14 sites were positioned in close proximity to different Alu repetitive elements while 2/14 sites were located within intron sequences. Insertions were also detected at sites without a specific gene designation but not within short tandem repetitive sequences, telomeres or centromeric repeat regions. Altogether, it is expected that this approach will yield new information on sites of integration by HIV-1 that may be associated with the pathogenic manifestations of disease progression.

Base Sequence↗

AliBaba2: context specific identification of transcription factor binding sites.

Currently, prediction of transcription factor binding sites is widely done using matrices collected from literature. This leads to several problems. We cannot actively control the conservation of the matrices, we cannot systematically use all binding sites available, we do not know which sites were used and which were discarded in matrix construction, we cannot compare and evaluate matrices easily, we cannot detect redundancy and we cannot control sensitivity and specificity. So we are lacking control during the identification process. In this paper a method to overcome these problems is proposed. It is assumed that each binding site has an unknown context which determines its sequence. This leads to the idea of constructing specific matrices for each sequence we are analysing. To do so we have to regard identification of binding sites as a general process, starting at a dataset of known binding sites and ending with the identification of a potential new binding site. In this paper such a process is presented. Besides overcoming the mentioned problems, the implementation also reaches a significantly higher accuracy than current approaches. Evaluations are done analysing all binding sites of TRANSFAC 3.5 public. The resulting tool AliBaba2 is available at http://wwwiti.cs.uni-magdeburg.de/grabe/alibaba2.

Algorithms↗

Resistance gene analogues of Arabidopsis thaliana: recognition by structure.

Following completion of Arabidopsis thaliana sequencing projects, multiple resistance gene analogues (RGAs) have been identified. In this work a review of the current state of knowledge available in protein databases and scientific articles is presented. Putative resistance genes were identified by using BLAST searches as well as HMM fingerprints (the latter to infer existence of characteristic domains). The representation of all five classes of putative resistance genes in Col-0 ecotype was examined, along with the statistics on RGAs present on all five chromosomes of Arabidopsis thaliana.

Arabidopsis↗

[HBV C gene mutation in the transmission from father to infant].

OBJECTIVE: Hepatitis B virus (HBV) DNA was detected from infants whose mothers were negative for all HBV markers and the fathers were HBV carrier, the homology of HBV sequence of fathers and fetus was high, and HBV mutations concentrated on some points, and the transmission of HBV from father to fetus was also identified in some reports. The present study aimed to study HBV transmission from father to infant. METHODS: The study enrolled 16 pairs of fathers who were HBV carriers and infants whose mothers were negative for HBV markers. The infants had evidences for intrauterine HBV infection. The five HBV serum markers HBsAg, HBeAg, anti-HBe, anti-HBs, and anti-HBc were detected with ELISA. The positive results for HBsAg and/or HBeAg were regarded as markers of HBV infection. Amplification of HBV DNA was done using a nested PCR method. The first amplification was carried out using primer C1 (nt 2394-2370), and primer C3 (nt 1730-1754). The second amplification was carried out using primer C2 (nt 1955-1974) and primer C6 (nt 2348-2330). Both primers were designed to amplify the part of sequence coding for the hepatitis B C antigen. The size of the amplified fragment obtained by the nested PCR was expected to be 394 bp. The PCR products were electrophoresed on 1.5% agarose gels, which were then stained with ethidium bromide and observed with ultraviolet transillumination. When 394 bp specific band was detectable, the sample was designated positive. Then the positive samples were identified by dot blot. The second PCR products were extracted by phenol-chloroform and 70% ethanol precipitation, then resuspended in TE buffer (pH8.0), and used as the template for cloning. The template was connected into pGEM-T vector by ligase. The ligated products were cloned into fresh competent JM109 cells, and incubated for 90 minutes at 37 degrees C on roller drum. Finally several dilutions were plated on plates containing ampicillin, X-Gal and IPTG, and incubated at 37 degrees C overnight. The white colony on plates was used for identification by the nested PCR with the above primers. When the 394 bp band was detectable by electrophoresis of PCR products in 1.5% agarose gels, the colony was designated positive; a positive colony was incubated in LB medium for 8 to 12 hrs, then plasmid was extracted using the Wizard Plus SV Minipreps DNA Purification System Kit (Promega). The purified plasmid was sent to Beijing Saibaisheng Company for sequencing. The homology of HBV C nt 2022-2301 sequence was compared between fathers and infants. RESULTS: The homology of HBV C nt 2022-2301 sequence were 99% - 100% in 16 pairs of fathers and infants. The results were referred to the published sequence of HBV adw/adr clones, and the nucleic acid databases were searched for homology by using BLAST tool on Internet. HBV of the sixteen pairs of father/infant was closely related to the Japan strain (Genebank accession number AF121249), but there were still 17 more mutations at nucleotide positions 2029, 2034, 2044, 2059, 2078, 2095, 2104, 2154, 2161, 2169, 2189, 2201, 2233, 2251, 2284, 2288, 2293. Moreover the mutations at positions 2189, 2288 resulted in the substitution of the encoded amino acid (corresponding to amino acid positions 97 and 130, respectively), the other mutations at the position were nonphenotypic. The mutation of 2189, 2288 nucleotide of HBV C gene caused 97, 130 amino acid substitution for isoleucine to leucine and proline to threonine. The mutation of 2189, 2288 nucleotide of HBV C gene were detected in 6 (37.5%) of 16 pairs of fathers and infants. CONCLUSION: The HBV transmission from father to infants did exist. The main HBV C gene mutation strains also existed in the transmission.

Adult↗

Sequence protein alignment with composition new evolutions (SPACne): a program for the identification of polypeptides using amino acid composition. A user friendly modification of SPAC.

SPAC (sequence protein alignment with composition) is a software that retrieve from protein or nucleic acid databases, the sequences corresponding to a protein or peptide whose only amino acid composition and molecular weight are known. By accurately matching a DNA or a protein sequence to candidate protein or peptide fragment, this software may be used as a fast and cheap method for protein characterization. This paper describes a modification of the SPAC software, SPAC new evolutions (SPACne), which enables a more efficient and user friendly method of protein identification using amino acids composition. SPACne is available online at the web site: http://bioweb.pasteur.fr/seqanal/interfaces/spacne.html.

Algorithms↗

Implications of compositionality in the gene ontology for its curation and usage.

In this paper we argue that a richer underlying representational model for the Gene Ontology that captures the implicit compositional structure of GO terms could have a positive impact on two activities crucial to the success of GO: ontology curation and database annotation. We show that many of the new terms added to GO in a one-year span appear to be compositional variations of other terms. We found that 90.2% of the 3,652 new terms added between July 2003 and July 2004 exhibited characteristics of compositionality. We also examine annotations available from the GO Consortium website that are either manually curated or automatically generated. We found that 74.5% and 63.2% of GO terms are seldom, if ever, used in manual and automatic annotations, respectively. We show that there are features that tend to distinguish terms that are used from those that are not. In order to characterize the effect of compositionality on the combinatorial properties of GO, we employ finite state automata that represent sets of GO terms. This representational tool demonstrates how ontologies can grow very fast, and also shows that small conceptual changes can directly result in a large number of changes to the terminology. We argue that the curation and annotation findings we report are influenced by the combinatorial properties that present themselves in an ontology that does not have a model that properly captures the compositional structure of its terms.

Computational Biology↗

Subfamily hmms in functional genomics.

The limitations of homology-based methods for prediction of protein molecular function are well known; differences in domain structure, gene duplication events and errors in existing database annotations complicate this process. In this paper we present a method to detect and model protein subfamilies, which can be used in high-throughput, genome-scale phylogenomic inference of protein function. We demonstrate the method on a set of nine PFAM families, and show that subfamily HMMs provide greater separation of homologs and non-homologs than is possible with a single HMM for each family. We also show that subfamily HMMs can be used for functional classification with a very low expected error rate. The BETE method for identifying functional subfamilies is illustrated on a set of serotonin receptors.

Animals↗

[Methods for improving the quality of prediction in the process of automatic annotating A4].

Modifications of the previously described adaptive algorithm of automatic annotating A4 have been considered, which make it possible to improve the quality of prediction. First, the direct use of the so-called basis statistics eta refines the quality of prediction compared with the previously proposed statistics gamma. Second, the quality is improved if not all similar sequences found but only part of them are used. This decreases the noising of the data, which in turn improves the quality of prognosis.

Algorithms↗

Molecular analysis of spontaneous nephrotropic anti-laminin antibodies in an autoimmune MRL-lpr/lpr mouse.

To explore the genetic relationship between anti-laminin and anti-DNA autoantibodies (autoAb), VH gene and gene family expression were determined among autoAb derived from an individual 6-mo-old MRL-lpr/lpr mouse. Whereas 85% of the anti-DNA Ig were identified by one of two VH family probes, 7183 and VHJ558, none of the anti-laminin antibodies (Ab) examined were recognized by these probes. Subsequent V region sequence analysis of three of the anti-laminin Ab revealed that they in fact utilized a J558 VH gene (VH50). Furthermore, FR2 and CDR2 oligonucleotide probes complementary to VH50 recognized multiple anti-laminin Ab by Northern blot analysis; the FR2 probe recognized two control anti-DNA Ab, but neither probe recognized anti-DNA Ab from the same mouse. Polymerase chain reaction amplification of MRL-lpr/lpr genomic liver DNA using primers generated from VH50 and Vk50 sequences indicated that all three anti-laminin Ig have a single replacement mutation in both their VH and Vk genes. Search of the nucleic acid databases revealed that both germline VH and Vk genes are expressed unmutated by murine lupus anti-dsDNA autoAb, previously sequenced in other laboratories. Sequence comparisons suggest that differences in anti-DNA and anti-laminin reactivity may be dependent upon somatically generated differences in the CDR3 regions of the H and L chains. The results indicate that lupus anti-laminin Ab can arise from distinct B cell populations but express the same unmutated germline V region genes as lupus anti-dsDNA autoAb. They further raise the possibility that these distinct B cell populations may be activated and expanded either: independently, by distinct Ig receptor ligands such as the Ag, laminin and DNA; or simultaneously, by a common ligand such as an anti-Id recognizing a common V region epitope.

Amino Acid Sequence↗

Molecular evolution of hydantoinases.

The complete amino acid sequence of the hydantoinase from Arthrobacter aurescens DSM 3745 has been derived by automated Edman degradation. This is the first ever reported amino acid sequence of a non-ATP-dependent hydantoinase, which hydrolyzes 5'-monosubstituted hydantoin derivatives L-selectively. A homology search performed in protein and nucleic acid databases retrieved only distantly related proteins. All of these are members of the recently described protein superfamily of amidohydrolases related to ureases (Holm and Sander, Proteins 28: 72-82, 1997). Phylogenetic analysis revealed that the novel hydantoinase forms a new branch separate from other hydantoin cleaving enzymes like dihydropyrimidinases (EC 3.5.2.2) and allantoinases (EC 3.5.2.5). Our results suggests that the enzymes of this protein superfamily have evolved from a common ancestor and therefore are the product of divergent evolution. We show further that the enclosed gene families developed very early in evolution, probably prior to the formation of the three domains, Archaea, Eukarya and Bacteria. Hydantoinases related to ATP-dependent N-methylhydantoinases (EC 3.5.2.14) or 5-oxoprolinases (EC 3.5.2.9) do not belong to this superfamily.

Amidohydrolases↗

Thermodynamic database for protein-nucleic acid interactions (ProNIT).

MOTIVATION: Protein-nucleic acid interactions are fundamental to the regulation of gene expression. In order to elucidate the molecular mechanism of protein-nucleic acid recognition and analyze the gene regulation network, not only structural data but also quantitative binding data are necessary. Although there are structural databases for proteins and nucleic acids, there exists no database for their experimental binding data. Thus, we have developed a Thermodynamic Database for Protein-Nucleic Acid Interactions (ProNIT). RESULTS: We have collected experimentally observed binding data from the literature. ProNIT contains several important thermodynamic data for protein-nucleic acid binding, such as dissociation constant (K(d)), association constant (K(a)), Gibbs free energy change (DeltaG), enthalpy change (DeltaH), heat capacity change (DeltaC(p)), experimental conditions, structural information of proteins, nucleic acids and the complex, and literature information. These data are integrated into a relational database system together with structural and functional information to provide flexible searching facilities by using combinations of various terms and parameters. A www interface allows users to search for data based on various conditions, with different display and sorting options, and to visualize molecular structures and their interactions. AVAILABILITY: ProNIT is freely accessible at the URL http://www.rtc.riken.go.jp/jouhou/pronit/pronit.html.

Amino Acid Sequence↗

Database resources of the National Center for Biotechnology Information.

In addition to maintaining the GenBank nucleic acid sequence database, the National Center for Biotechnology Information (NCBI) provides analysis and retrieval resources for the data in GenBank and other biological data made available through NCBI's Web site. NCBI resources include Entrez, the Entrez Programming Utilities, MyNCBI, PubMed, PubMed Central, Entrez Gene, the NCBI Taxonomy Browser, BLAST, BLAST Link (BLink), Electronic PCR, OrfFinder, Spidey, Splign, RefSeq, UniGene, HomoloGene, ProtEST, dbMHC, dbSNP, Cancer Chromosomes, Entrez Genomes and related tools, the Map Viewer, Model Maker, Evidence Viewer, Clusters of Orthologous Groups, Retroviral Genotyping Tools, HIV-1, Human Protein Interaction Database, SAGEmap, Gene Expression Omnibus, Entrez Probe, GENSAT, Online Mendelian Inheritance in Man, Online Mendelian Inheritance in Animals, the Molecular Modeling Database, the Conserved Domain Database, the Conserved Domain Architecture Retrieval Tool and the PubChem suite of small molecule databases. Augmenting many of the Web applications are custom implementations of the BLAST program optimized to search specialized datasets. All of the resources can be accessed through the NCBI home page at: http://www.ncbi.nlm.nih.gov.

Databases, Genetic↗

Database resources of the National Center for Biotechnology Information.

In addition to maintaining the GenBank nucleic acid sequence database, the National Center for Biotechnology Information (NCBI) provides analysis and retrieval resources for the data in GenBank and other biological data made available through NCBI's Web site. NCBI resources include Entrez, the Entrez Programming Utilities, My NCBI, PubMed, PubMed Central, Entrez Gene, the NCBI Taxonomy Browser, BLAST, BLAST Link(BLink), Electronic PCR, OrfFinder, Spidey, Splign, RefSeq, UniGene, HomoloGene, ProtEST, dbMHC, dbSNP, Cancer Chromosomes, Entrez Genome, Genome Project and related tools, the Trace and Assembly Archives, the Map Viewer, Model Maker, Evidence Viewer, Clusters of Orthologous Groups (COGs), Viral Genotyping Tools, Influenza Viral Resources, HIV-1/Human Protein Interaction Database, Gene Expression Omnibus (GEO), Entrez Probe, GENSAT, Online Mendelian Inheritance in Man (OMIM), Online Mendelian Inheritance in Animals (OMIA), the Molecular Modeling Database (MMDB), the Conserved Domain Database (CDD), the Conserved Domain Architecture Retrieval Tool (CDART) and the PubChem suite of small molecule databases. Augmenting many of the Web applications are custom implementations of the BLAST program optimized to search specialized data sets. These resources can be accessed through the NCBI home page at www.ncbi.nlm.nih.gov.

Animals↗

Molecular Biocomputing Suite: a word processor add-in for the analysis and manipulation of nucleic acid and protein sequence data.

In all fields of molecular biology, researchers are increasingly challenged by experiments planned and evaluated on the basis of nucleic acid and protein sequence data generally retrieved from public databases. Despite the wide spectrum of available Web-based software tools for sequence analysis, the routine use of these tools has disadvantages, particularly because of the elaborate and heterogeneous ways of data input, output, and storage. Here we present a Visual Basic-encoded Microsoft Word Add-In, the Molecular BioComputing Suite (MBCS), available at the BioTechniques Software Library (www.BioTechniques.com). The MBCS software aims to manage and expedite a wide range of sequence analyses and manipulations using an integrated text editor environment including menu-guided commands. Its independence of sequence formats enables MBCS to be used as a pivotal application between other software tools for sequence analysis, manipulation, annotation, and editing.

Amino Acid Sequence↗