PubMed HealthSearch

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Databases in molecular biology: a CODATA task group at work.

A certain concern exists that the exponential growth of nucleic acid and protein sequence data will saturate the channels of data acquisition, distribution and utilization on the one hand and, on the other hand, that even the actual resources are still not fully and easily accessible to any bench scientist. Despite the stake of the scientific community at large in the fundamental data collected in this field, there has been in past years only a modest effort to discuss the common problems at an international level. Three international meetings were organized in 1987 on this subject: the annual meeting of CODATA Task Group on Coordination of Protein Sequence Data Banks (Nice, France, January 1987), the EMBL/NIH Workshop concerned primarily with nucleic acid databases (Heidelberg, FRG, February 1987) and the CODATA Workshop on Nucleic Acid and Protein Sequencing Data (Gaithersburg, USA, May 1987).

Amino Acid Sequence

Statistical analysis of nucleotide sequences.

In order to scan nucleic acid databases for potentially relevant but as yet unknown signals, we have developed an improved statistical model for pattern analysis of nucleic acid sequences by modifying previous methods based on Markov chains. We demonstrate the importance of selecting the appropriate parameters in order for the method to function at all. The model allows the simultaneous analysis of several short sequences with unequal base frequencies and Markov order k not equal to 0 as is usually the case in databases. As a test of these modifications, we show that in E. coli sequences there is a bias against palindromic hexamers which correspond to known restriction enzyme recognition sites.

Base Sequence

A systematic method for studying the spatial distribution of water molecules around nucleic acid bases.

A new method to analyze the distribution of water molecules around the bases in DNA is presented. This method relies on the notion of a "hydrated building block," which represents the joint observed hydration around all bases of a particular type, in structures of a particular conformation type. The hydrated building blocks were constructed using atomic coordinates from 40 structures contained in the Nucleic Acid Database. Pseudoelectron densities were calculated for water molecules in each hydrated building block using standard crystallographic procedures. The electron densities were fitted to obtain "average building blocks," which represent bases with waters only at average or probable positions. Both types of building blocks were used to construct models of hydrated DNA oligomers. The essential features of the solvent structure around d(CGCGAATTCGCG)2 in the B form and d(CGCGCG)2 in the Z form were reproduced.

Adenine

Increment of efficiency in the identification of noble genes by colony hybridization assay. (Screening of low redundant clones from human fetal cDNA library).

For the rapid identification of noble genes in a specific tissue by computer analysis from the cDNA sequences determined by single-pass cDNA sequencing, clone redundancy was one of the major obstacles. To facilitate the efficiency in identification of noble genes, it was necessary to reduce the number of clones to be sequenced by eliminating the redundant clones for a rapid analysis. In order to increase the probability of isolating noble sequences from the cDNA clones of human fetal liver tissue origin, colony hybridization assay was adopted and redundant clones were efficiently removed. Four cDNA clones highly redundant in the human fetal liver cDNA libraries including alpha-globin, gamma-globin, serum albumin and H19 RNA sequences were selected as the probes. Two hundreds and sixty two cDNA clones were randomly selected and tested with the probes for hybridization properties. The identity of each cDNA clone giving positive or negative signals in the hybridization assay was determined by DNA homology search with the nucleic acid databases. Among the 76 clones giving positive signals, 57 clones (75%) were found to be identical to the probe sequences and could be eliminated by colony hybridization assay before neucleotide sequencing.

Cloning, Molecular

CORGEN: a FORTRAN-77 generator of standard and non-standard DNA helices from the sequence.

An analytical procedure CORGEN generates a variety of DNA double-stranded structures from user-supplied sequence using a nucleic acid database incorporated into a standard FORTRAN-77 program. Alternatively, the cylindrical polar coordinates of DNA components may be supplied from the external table. An algorithm that performs intercalation sites in DNA is described. This procedure can be used to generate complexes of antibiotics with DNA. Non-standard DNA structures can be built by alternating the global helical twist and global helical rise in the regular DNA helix. The procedures described can be used for computer generation of a variety of non-standard DNA structures which can be subjected to molecular mechanics and dynamics simulations.

Base Sequence

Water molecules in DNA recognition II: a molecular dynamics view of the structure and hydration of the trp operator.

The structure and hydration of the DNA duplex d-(AGCGTACTAGTACGCT)2 corresponding to the trp operator fragment used in the crystal structure of the half site complex (PDB entry 1TRR) was studied by a 1.4 ns molecular dynamics simulation in water. The simulation, starting from a B-DNA conformation, used a non-bonded cutoff of 1.4 nm with a reaction field correction and resulted in a stable trajectory. The average DNA conformation obtained was closer to the ones found in the crystal structures of the complexes (PDB entries 1TRO and 1TRR) than to the crystal structure of unbound trp operator (Nucleic Acid Database entry BDJ061). The DNA hydration was characterized in terms of hydrogen bond percentages and corresponding residence times. The residence times of water molecules within 0.35 nm of the DNA non-exchangeable protons were calculated for comparison with NMR measurements of intermolecular water-DNA NOEs and nuclear magnetic relaxation dispersion measurements. No significant difference was found between major and minor groove hydration. The DNA donors and acceptors were hydrogen bonded to water molecules for 77(+/-19)% of the time on average. The average residence time of the hydrogen bonded water molecules was 11(+/-11) ps with a maximum of 223 ps. When all water molecules within NOE distance (0.35 nm) of non-exchangeable protons were considered, the average residence times increased to an average of 100(+/-4) ps and a maximum of 608 ps. These results agree with the experimental NMR results of Sunnerhagen et al. which did not show any evidence for water molecules bound with more than 1 ns residence time on the DNA surface. The exchange of hydration water from the DNA occurred in the major groove primarily through direct exchange with the bulk solvent, while access to and from the minor groove frequently proceeded via pathways involving ribose O3' and O4' and phosphate O2P oxygen atoms. The most common water diffusion pathways in the minor groove were perpendicular to the groove direction. In general, water molecules visited only a limited number of sites in the DNA grooves before exiting. The hydrogen bonding sites, where hydrogen bonds could be formed with donor and acceptor groups of the DNA, were filled with water molecules with an average B-factor value of 0.58 mn2. No special values were observed at any of the sites, where water molecules were observed both in the trp repressor/operator co-crystals and in the crystal structure of unbound DNA.

Bacterial Proteins

Identification of phosphorylated proteins from thrombin-activated human platelets isolated by two-dimensional gel electrophoresis by electrospray ionization-tandem mass spectrometry (ESI-MS/MS) and liquid chromatography-electrospray ionization-mass spectrometry (LC-ESI-MS).

Two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) is a powerful tool to separate complex protein mixtures including whole cell lysates. In combination with immunoblotting techniques or radioactive labeling techniques it is a fast and convenient way to demonstrate the presence of certain proteins or protein modifications. With the development of extremely sensitive analytical techniques such as matrix-assisted laser desorption/ionization-mass spectrometry (MALDI-MS) or electrospray ionization (ESI)-MS, it has become possible to use 2-D gels not only as an analytical but also as a preparative tool. Starting with a number of spots excised from 2-D gels, a protein can be identified using different strategies involving enzymatic cleavage of the protein in the gel matrix, elution of the resulting peptides and analysis of these peptides by mass spectrometry. The obtained peptide mass fingerprint or fragment ion spectra from peptides can be used to screen protein or nucleic acid databases in order to identify the protein. We have used the techniques described above to identify proteins from human platelets which change their phosphorylation state following activation of platelets by thrombin. Platelets were radioactively labeled with [32P]orthophosphate and stimulated. Several protein spots in the observed range of 10-80 kDa and an isoelectric point of 3-10 showed a significant increase or decrease in phosphorylation. We present the results from the investigation of a spot group representing different isoforms and phosphorylation states of myosin light chain.

Amino Acid Sequence

Striking evolutionary conservation of a cis-element related to nuclear receptor target sites and present in TR2 orphan receptor genes.

A systematic scanning of nucleic acid databases for DNA elements made of combinations of RGGTCA nuclear receptor half sites, has revealed that identical 19 nucleotide-long motifs composed of two inverted RGGTCA sites with a spacing of 7 nucleotides (IR7), are present upstream of the regions coding for the human TR2 and of the sea urchin SpSHR2 orphan receptors. We have developed an experimental strategy based on PCR, to check if this IR7 could correspond to an unusually long cis-element, conserved along evolution and regulating the TR2 genes. We found that indeed IR7 is present in the 5' untranslated region of TR2 genes from all species tested, including Xenopus, rainbow trout, zebrafish and mouse. The exact conservation throughout the animal kingdom of such a long, non repetitive and non coding genomic region, highly suggests that it should ensure important biological functions. In addition, this work has allowed the identification of a new, non coding, upstream exon in the mouse TR2 gene present in testicular TR2 mRNAs.

Animals

Soluble acid invertase determines the hexose-to-sucrose ratio in cold-stored potato tubers.

Cold storage of potato (Solanum tuberosum L.) tubers is known to cause accumulation of reducing sugars. Hexose accumulation has been shown to be cultivar-dependent and proposed to be the result of sucrose hydrolysis via invertase. To study whether hexose accumulation is indeed related to the amount of invertase activities, two different approaches were used: (i) neutral and acidic invertase activities as well as soluble sugars were measured in cold-stored tubers of 24 potato cultivars differing in the cold-induced accumulation of reducing sugars and (ii) antisense potato plants with reduced soluble acid invertase activities were created and the soluble sugar accumulation in cold-stored tubers was studied. The cold-induced hexose accumulation in tubers from the different potato cultivars varied strongly (up to eightfold). Large differences were also detected with respect to soluble acid (50-fold) and neutral (5-fold) invertase activities among the different cultivars. Although there was almost no correlation between the total amount of invertase activity and the accumulation of reducing sugars there was a striking correlation between the hexose/sucrose ratio and the extractable soluble invertase activity. To exclude the possibility that other cultivar-specific features could account for the obtained results, the antisense approach was used to decrease the amount of soluble acid invertase activity in a uniform genetic background. To this end the cDNA of a cold-inducible soluble acid invertase (EMBL nucleic-acid database accession no. X70368) was cloned from the cultivar Desirée, and transgenic potato plants were created expressing this cDNA in the antisense orientation under control of the constitutive 35S cauliflower mosaic virus promotor. Analysis of the harvested and cold-stored tubers showed that inhibition of the soluble acid invertase activity leads to a decreased hexose and an increased sucrose content compared with controls. As was already found for the different potato cultivars the hexose/sucrose ratio decreased with decreasing invertase activities but the total amount of soluble sugars did not significantly change. From these data we conclude that invertases do not control the total amount of soluble sugars in cold-stored potato tubers but are involved in the regulation of the ratio of hexose to sucrose.

Cold Temperature

Genes selectively expressed in the infectious (metacyclic) stage of Leishmania major promastigotes encode a potential basic-zipper structural motif.

Complementary DNA clones representing transcripts selectively expressed in the non-dividing, infective (metacyclic) stage of Leishmania major promastigotes (MP) were identified by differential and subtractive screening. The majority of the selected clones hybridized on Northern blots to a set of transcripts highly expressed by MP, but to a much lower extent in proliferating and stationary-phase attenuated promastigotes. Stationary, but not log-phase cultures, of each of 5 L. major strains showing a potential for differentiation to metacyclics, expressed these transcripts (MAT-1; MP-associated transcripts). From sequence analysis of full-length cDNA clones corresponding to the predominating MAT-1 species, an open reading frame encoding a 139 aa polypeptide (15.4 kDa) was predicted and supported by immunoprecipitation by kala-azar sera of reticulocyte extract translation products using in vitro transcribed RNA. Although no significant primary sequence homology to database nucleic acid and protein sequences was found, the sequence displays similarities to the basic-zipper families of transcription regulatory proteins.

Amino Acid Sequence

Amino acid sequences of lysozymes newly purified from invertebrates imply wide distribution of a novel class in the lysozyme family.

Lysozymes were purified from three invertebrates: a marine bivalve, a marine conch, and an earthworm. The purified lysozymes all showed a similar molecular weight of 13 kDa on SDS/PAGE. Their N-terminal sequences up to the 33rd residue determined here were apparently homologous among them; in addition, they had a homology with a partial sequence of a starfish lysozyme which had been reported before. The complete sequence of the bivalve lysozyme was determined by peptide mapping and subsequent sequence analysis. This was composed of 123 amino acids including as many as 14 cysteine residues and did not show a clear homology with the known types of lysozymes. However, the homology search of this protein on the protein or nucleic acid database revealed two homologous proteins. One of them was a gene product, CELF22 A3.6 of C. elegans, which was a functionally unknown protein. The other was an isopeptidase of a medicinal leech, named destabilase. Thus, a new type of lysozyme found in at least four species across the three classes of the invertebrates demonstrates a novel class of protein/lysozyme family in invertebrates. The bivalve lysozyme, first characterized here, showed extremely high protein stability and hen lysozyme-like enzymatic features.

Amino Acid Sequence

eccDNABase: A Comprehensive and High-Quality Database for Extrachromosomal Circular DNA.

Extrachromosomal circular DNA (eccDNA) refers to small, circular DNA molecules that originate from chromosomal sequences and are prevalent across nearly all eukaryotic organisms. In humans, eccDNAs are widely distributed in normal tissues, cancerous tissues, and body fluids, where they play important roles in tumorigenesis and are often associated with poor clinical outcomes. Given their biological and clinical significance, a well-integrated and high-quality database is essential for advancing eccDNA-related research. To address this need, we developed eccDNABase, a comprehensive and curated resource for browsing, searching, and analyzing eccDNAs across multiple species. The database systematically catalogs eccDNA-disease associations from diverse tissues and organisms. Currently, eccDNABase contains 1,875,452 eccDNA-disease associations, encompassing 8,398 ecDNA entries across nine species, 63 diseases, and healthy individuals. Each entry provides detailed information, including eccDNA ID, type, chromosomal localization, species, tissue or cell line source, disease name and Disease Ontology ID, overlap length and percentage with genes, oncogene overlap, detection method, and links to literature and source databases. Given its extensive and curated datasets, eccDNABase serves as a valuable resource for both basic and translational research, offering deeper insights into the role of eccDNA in health and disease. The database is publicly accessible at http://cgga.org.cn/eccDNABase/.

Humans

Hydrogen bond geometry in DNA-minor groove binding drug complexes.

The geometry of the hydrogen bonding interaction between DNA and minor-groove binding drugs has been analyzed from a sample of 22 crystal structures of DNA-drug complexes, retrieved from the Nucleic Acid Database. Seventy-seven interactions between the drugs and acceptor groups in the nucleotide bases can be classified as hydrogen bonds. Their geometry departs significantly from linearity since, in most instances, the interactions can be described as three-center or multiple hydrogen bonds. Results also show that there is no preference for hydrogen bonds involving positively charged groups in the drugs. Relationships between hydrogen bond geometry and positioning of the drug along the minor groove are also discussed. The information presented may be useful in the design of new specific minor groove binding drugs.

Anti-Bacterial Agents

Molecular cloning of two novel transmembrane ligands for Eph-related kinases (LERKS) that are related to LERK-2.

A search of the nucleic acid database of expressed sequence tags (ESTs) revealed several partial cDNA sequences that could encode proteins homologous to the ligands for Eph-related kinases (LERKs). Oligonucleotides designed from the ESTs were used to probe a human brain cDNA library and obtain overlapping clones that encoded two different novel LERKS (NLERK-1 and NLERK-2). NLERK-1 and NLERK-2 are most closely related to human LERK-2/Elk-ligand and they form a subclass of LERKs that contain a transmembrane domain and a conserved cytoplasmic domain. Full-length NLERK-1 was expressed as a glycosylated membrane protein in COS cells and was not secreted into the medium. Full-length NLERK-2 was similarly expressed in COS cells but both membrane-bound and a truncated, proteolytically-released form were detected. Engineered forms of both NLERK-1 and NLERK-2 lacking transmembrane and cytoplasmic domains were also expressed in COS cells and each was detected in the extracellular medium.

Amino Acid Sequence

Molecular analysis of spontaneous nephrotropic anti-laminin antibodies in an autoimmune MRL-lpr/lpr mouse.

To explore the genetic relationship between anti-laminin and anti-DNA autoantibodies (autoAb), VH gene and gene family expression were determined among autoAb derived from an individual 6-mo-old MRL-lpr/lpr mouse. Whereas 85% of the anti-DNA Ig were identified by one of two VH family probes, 7183 and VHJ558, none of the anti-laminin antibodies (Ab) examined were recognized by these probes. Subsequent V region sequence analysis of three of the anti-laminin Ab revealed that they in fact utilized a J558 VH gene (VH50). Furthermore, FR2 and CDR2 oligonucleotide probes complementary to VH50 recognized multiple anti-laminin Ab by Northern blot analysis; the FR2 probe recognized two control anti-DNA Ab, but neither probe recognized anti-DNA Ab from the same mouse. Polymerase chain reaction amplification of MRL-lpr/lpr genomic liver DNA using primers generated from VH50 and Vk50 sequences indicated that all three anti-laminin Ig have a single replacement mutation in both their VH and Vk genes. Search of the nucleic acid databases revealed that both germline VH and Vk genes are expressed unmutated by murine lupus anti-dsDNA autoAb, previously sequenced in other laboratories. Sequence comparisons suggest that differences in anti-DNA and anti-laminin reactivity may be dependent upon somatically generated differences in the CDR3 regions of the H and L chains. The results indicate that lupus anti-laminin Ab can arise from distinct B cell populations but express the same unmutated germline V region genes as lupus anti-dsDNA autoAb. They further raise the possibility that these distinct B cell populations may be activated and expanded either: independently, by distinct Ig receptor ligands such as the Ag, laminin and DNA; or simultaneously, by a common ligand such as an anti-Id recognizing a common V region epitope.

Amino Acid Sequence

Molecular evolution of hydantoinases.

The complete amino acid sequence of the hydantoinase from Arthrobacter aurescens DSM 3745 has been derived by automated Edman degradation. This is the first ever reported amino acid sequence of a non-ATP-dependent hydantoinase, which hydrolyzes 5'-monosubstituted hydantoin derivatives L-selectively. A homology search performed in protein and nucleic acid databases retrieved only distantly related proteins. All of these are members of the recently described protein superfamily of amidohydrolases related to ureases (Holm and Sander, Proteins 28: 72-82, 1997). Phylogenetic analysis revealed that the novel hydantoinase forms a new branch separate from other hydantoin cleaving enzymes like dihydropyrimidinases (EC 3.5.2.2) and allantoinases (EC 3.5.2.5). Our results suggests that the enzymes of this protein superfamily have evolved from a common ancestor and therefore are the product of divergent evolution. We show further that the enclosed gene families developed very early in evolution, probably prior to the formation of the three domains, Archaea, Eukarya and Bacteria. Hydantoinases related to ATP-dependent N-methylhydantoinases (EC 3.5.2.14) or 5-oxoprolinases (EC 3.5.2.9) do not belong to this superfamily.

Amidohydrolases

Prediction of gene expression specificity by promoter sequence patterns.

We present here a heuristic method toward predicting the expression specificity in the transcriptional process, which is known to be regulated in large part by promoter sequences, by observing the appearance of conserved sequence patterns in a group of known promoters, such as for housekeeping or tissue-specific genes. Statistically conserved patterns were automatically extracted from a set of unaligned sequences up to 200 bp upstream of the transcription initiation site, by a standard procedure using the Markov chain and binomial distribution models. Furthermore, to obtain signal sequences of optimal lengths we devised a method that combines the multiple alignment and the analysis of the information content (or relative entropy). Groups of related promoters were compiled from the EPD eukaryotic promoter database and the EMBL nucleic acid sequence database. Each promoter was examined for its specificity by linear discriminant analysis to test the validity of the extracted patterns. Our method could correctly discriminate 77.6% of the housekeeping gene promoters and 62.9% of the liver promoters.

Algorithms