PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Linkage of amino acid variation and evolution of human immunodeficiency virus type 1 gp120 envelope glycoprotein (subtype B) with usage of the second receptor.

To clarify the relationship between the amino acid variations of the gp120 of human immunodeficiency virus type 1 (HIV-1) and the chemokine receptors that are used as the second receptor for HIV, we evaluated amino acid site variation of gp120 between the X4 strains (use CXCR4) and the R5 strains (use CCR5) from 21 sequences of subtype B. Our analysis showed that residues 306 and 322 in the V3 loop and residue 440 in the C4 region were associated with usage of the second receptor. The polymorphism at residue 440 is clearly associated with the usage of the second receptor: The amino acid at position 440 was a basic amino acid in the R5 strains, and a nonbasic and smaller amino acid in the X4 strains, while the V3 loop of the X4 strains was more basic than that of the R5 strains. This suggests that residue 440 in the C4 region, which is close to the V3 loop in the three-dimensional structure, is critical in determining which second receptor is used. Analysis of codon frequency suggests that, in almost all cases, the difference at residue 440 between basic amino acids in the R5 strains and nonbasic amino acids in the X4 strains could be due to a single nucleotide change. These findings predict that the evolutionary changes in amino acid residue 440 may be correlated with evolutionary changes in the V3 loop. One possibility is that a change in electric charge at residue 440 compensates for a change in electric charge in the V3 loop. The amino acid polymorphism at position 440 can be useful to predict the cell tropism of a strain of HIV-1 subtype B.

Amino Acid Sequence↗

Loss of protein structure stability as a major causative factor in monogenic disease.

The most common cause of monogenic disease is a single base DNA variant resulting in an amino acid substitution. In a previous study, we observed that a high fraction of these substitutions appear to result in reduction of stability of the corresponding protein structure. We have now investigated this phenomenon more fully. A set of structural effects, such as reduction in hydrophobic area, overpacking, backbone strain, and loss of electrostatic interactions, is used to represent the impact of single residue mutations on protein stability. A support vector machine (SVM) was trained on a set of mutations causative of disease, and a control set of non-disease causing mutations. In jack-knifed testing, the method identifies 74% of disease mutations, with a false positive rate of 15%. Evaluation of a set of in vitro mutagenesis data with the SVM established that the majority of disease mutations affect protein stability by 1 to 3 kcal/mol. The method's effective distinction between disease and non-disease variants, strongly supports the hypothesis that loss of protein stability is a major factor contributing to monogenic disease.

Amino Acid Substitution↗

Modeling amino acid substitution patterns in orthologous and paralogous genes.

We study to what degree patterns of amino acid substitution vary between genes using two models of protein-coding gene evolution. The first divides the amino acids into groups, with one substitution rate for pairs of residues in the same group and a second for those in differing groups. Unlike previous applications of this model, the groups themselves are estimated from data by simulated annealing. The second model makes substitution rates a function of the physical and chemical similarity between two residues. Because we model the evolution of coding DNA sequences as opposed to protein sequences, artifacts arising from the differing numbers of nucleotide substitutions required to bring about various amino acid substitutions are avoided. Using 10 alignments of related sequences (five of orthologous genes and five gene families), we do find differences in substitution patterns. We also find that, although patterns of amino acid substitution vary temporally within the history of a gene, variation is not greater in paralogous than in orthologous genes. Improved understanding of such gene-specific variation in substitution patterns may have implications for applications such as sequence alignment and phylogenetic inference.

Actins↗

Matching peptide mass spectra to EST and genomic DNA databases.

The use of mass spectrometry data to search molecular sequence databases is a well-established method for protein identification. The technique can be extended to searching raw genomic sequences, providing experimental confirmation or correction of predicted coding sequences, and has the potential to identify novel genes and elucidate splicing patterns.

Amino Acid Sequence↗

Prediction of gene expression specificity by promoter sequence patterns.

We present here a heuristic method toward predicting the expression specificity in the transcriptional process, which is known to be regulated in large part by promoter sequences, by observing the appearance of conserved sequence patterns in a group of known promoters, such as for housekeeping or tissue-specific genes. Statistically conserved patterns were automatically extracted from a set of unaligned sequences up to 200 bp upstream of the transcription initiation site, by a standard procedure using the Markov chain and binomial distribution models. Furthermore, to obtain signal sequences of optimal lengths we devised a method that combines the multiple alignment and the analysis of the information content (or relative entropy). Groups of related promoters were compiled from the EPD eukaryotic promoter database and the EMBL nucleic acid sequence database. Each promoter was examined for its specificity by linear discriminant analysis to test the validity of the extracted patterns. Our method could correctly discriminate 77.6% of the housekeeping gene promoters and 62.9% of the liver promoters.

Algorithms↗

ProTherm and ProNIT: thermodynamic databases for proteins and protein-nucleic acid interactions.

ProTherm and ProNIT are two thermodynamic databases that contain experimentally determined thermodynamic parameters of protein stability and protein-nucleic acid interactions, respectively. The current versions of both the databases have considerably increased the total number of entries and enhanced search interface with added new fields, improved search, display and sorting options. As on September 2005, ProTherm release 5.0 contains 17,113 entries from 771 proteins, retrieved from 1497 scientific articles (approximately 20% increase in data from the previous version). ProNIT release 2.0 contains 4900 entries from 273 research articles, representing 158 proteins. Both databases can be queried using WWW interfaces. Both quick search and advanced search are provided on this web page to facilitate easy retrieval and display of the data from these databases. ProTherm is freely available online at http://gibk26.bse.kyutech.ac.jp/jouhou/Protherm/protherm.html and ProNIT at http://gibk26.bse.kyutech.ac.jp/jouhou/pronit/pronit.html.

DNA↗

The Molecular Biology Database Collection: 2005 update.

The Nucleic Acids Research Molecular Biology Database Collection is a public online resource that lists the databases described in this and previous issues of Nucleic Acids Research together with other databases of value to the biologist and available throughout the world. All databases included in this Collection are freely available to the public. The 2005 update includes 719 databases, 171 more than the 2004 one. The databases are organized in a hierarchical classification that simplifies the process of finding the right database for any given task. The growing number of databases related to immunology, plant and organelle research have been accommodated by separating them into three new categories. The database summaries provide brief descriptions of the databases, contact details, appropriate references and acknowledgements. The online summaries also serve as a venue for the maintainers of each database to introduce database updates and other improvements in the scope and tools. These updates are particularly important for those databases that have not been described in print in the recent past. The database list and summaries are available online at the Nucleic Acids Research web site, http://nar.oupjournals.org/.

Allergy and Immunology↗

Fluorescence energy transfer monitored competitive equilibria of nucleic acids: applications in thermodynamics and screening.

Precise thermodynamic characterization of nucleic acid complex stability is required to understand a variety of biologically significant events as well as to exploit the specific recognition capabilities of nucleic acids in biotechnology, diagnostics, and therapeutics. The development of a database of nucleic acid thermodynamics with sufficient precision to foster further developments in these areas requires new and improved measurement techniques. The combination of a competitive equilibrium titration with fluorescence energy transfer based detection provides a method for precise measurement of differences in free energy values for nucleic acid duplexes that far exceeds in precision those accessible via conventional methods. The method can be applied to detect and to characterize any deviation in a nucleic acid that alters duplex stability. Such deviations include, but are not limited to, mismatches; single nucleotide polymorphisms (SNP); chemically modified nucleotide bases, sugars or phosphates; and conformational anomalies or folding motifs, such as, loops or hairpins.

Binding, Competitive↗

Analysis of the genomic organization of the human cationic amino acid transporters CAT-1, CAT-2 and CAT-4.

By screening nucleotide databases, sequences containing the complete genes of the human cationic amino acid transporters (hCATs) 1, 2 and 4 were identified. Analysis of the genomic organization revealed that hCAT-2 consists of 12 translated exons and most likely of 2 untranslated exons. The splice variants hCAT-2A and hCAT-2B use exon 7 and 6, respectively. The hCAT-2 gene structure is closely related to the structure of hCAT-1, suggesting that they belong to a common gene family. hCAT-4 consists of only 4 translated exons and 3 short introns. Exons of identical size and highly homologous to exon 3 of hCAT-4 are present in hCAT-1 and hCAT-2.

Amino Acid Transport System y+↗

Post genomic analysis of permeases from the amino acid/auxin family in protozoan parasites.

The "amino acid/auxin permeases" is probably the most represented family of transporters in the Trypanosoma cruzi genome. Using a high-throughput searching routine and preliminary data from the T. cruzi genome project, more than 15,000 sequences were iteratively assembled into contigs, and 60 open reading frames corresponding to different putative amino acid transporters, clustered in 12 groups, were detected and characterized in silico. T. cruzi genomic organization of such sequences showed that these putative amino acid transporter genes are in an unusually large number and arranged in repeat clusters comprising about 0.2% of the genome. These data suggest that the family has evolved following tandem duplication events and constitutes a novel family of variable proteins in protozoan organisms. The mRNA expression of the predicted genes was demonstrated in infective and non-infective parasite forms. Orthologous sequences were also identified in other unicellular parasites such as Leishmania spp., Plasmodium spp., and Trypanosoma brucei.

Amino Acid Sequence↗

Negative selection on neutralization epitopes of poliovirus surface proteins: implications for prediction of candidate epitopes for immunization.

For development of effective vaccines against viruses, it is of importance to choose appropriate epitopes as the target for immunization. These epitopes should eventually be determined experimentally, but it would be helpful if we could predict candidate epitopes computationally because it accelerates the entire process. To predict candidate epitopes for immunization, it is of great interest to characterize the target epitopes of poliovirus vaccine, which has empirically proven to be the most effective among all vaccines available. Here I show that almost all amino acid sites of poliovirus surface proteins VP1, VP2, and VP3 including neutralization epitopes are negatively selected and no site is under positive selection. These results, together with those obtained in previous studies, indicate that vaccines directed against epitopes, which consist of negatively selected sites protect vaccinees more effectively than those directed against epitopes which contain positively selected sites. These observations suggest that candidate epitopes for immunization are predicted by the molecular evolutionary analysis of viral protein (and its coding nucleotide) sequences, as the epitopes which consist exclusively of negatively selected amino acid sites.

Amino Acids↗

In search of the prototype of nitric oxide synthase.

Recent identification of the prokaryotic genes related to the catalytic oxygenase domain of mammalian nitric oxide synthase (NOS) has led to speculations on the origins of the NO signaling network. NOS activity in eukaryotes relies on the concerted action of the oxygenase domain with an electron-donating reductase domain that is fused to it. A fused reductase domain is, however, absent in prokaryotes. Consequently, we searched bacterial genomes for homologs of the reductase domain and identified candidate genes. On the basis of genomic sequence and protein structural analysis, we here propose that sulfite reductase flavoprotein is a prototype of the mammalian NOS reductase domain and a complementing interaction partner of the bacterial NOS oxygenase protein.

Amino Acid Sequence↗

Evidence for Golgi bodies in proposed 'Golgi-lacking' lineages.

Golgi bodies are nearly ubiquitous in eukaryotic cells. The apparent lack of such structures in certain eukaryotic lineages might be taken to mean that these protists evolved prior to the acquisition of the Golgi, and it raises questions of how these organisms function in the absence of this crucial organelle. Here, we report gene sequences from five proposed 'Golgi-lacking' organisms (Giardia intestinalis, Spironucleus barkhanus, Entamoeba histolytica, Naegleria gruberi and Mastigamoeba balamuthi). BLAST and phylogenetic analyses show these genes to be homologous to those encoding components of the retromer, coatomer and adaptin complexes, all of which have Golgi-related functions in mammals and yeast. This is, to our knowledge, the first molecular evidence for Golgi bodies in two major eukaryotic lineages (the pelobionts and heteroloboseids). This substantiates the suggestion that there are no extant primitively 'Golgi-lacking' lineages, and that this apparatus was present in the last common eukaryotic ancestor, but has been altered beyond recognition several times.

Amino Acid Sequence↗

Ancestral organization of the MHC revealed in the amphibian Xenopus.

With the advent of the Xenopus tropicalis genome project, we analyzed scaffolds containing MHC genes. On eight scaffolds encompassing 3.65 Mbp, 122 MHC genes were found of which 110 genes were annotated. Expressed sequence tag database screening showed that most of these genes are expressed. In the extended class II and class III regions the genomic organization, excluding several block inversions, is remarkably similar to that of the human MHC. Genes in the human extended class I region are also well conserved in Xenopus, excluding the class I genes themselves. As expected from previous work on the Xenopus MHC, the single classical class I gene is tightly linked to immunoproteasome and transporter genes, defining the true class I region, present in all nonmammalian jawed vertebrates studied to date. Surprisingly, the immunoproteasome gene PSMB10 is found in the class III region rather than in the class I region, likely reflecting the ancestral condition. Xenopus DMalpha, DMbeta, and C2 genes were identified, which are not present or not clearly identifiable in the genomes of any teleosts. Of great interest are novel V-type Ig superfamily (Igsf) genes in the class III region, some of which have inhibitory motifs (ITIM) in their cytoplasmic domains. Our analysis indicates that the vertebrate MHC experienced a vigorous rearrangement in the bony fish and bird lineages, and a translocation and expansion of the class I genes in the mammalian lineage. Thus, the amphibian MHC is the most evolutionary conserved MHC so far analyzed.

Amino Acid Sequence↗

Actin binding LIM protein 3 (abLIM3).

LIM domain proteins were demonstrated to play key roles in various biological processes such as embryonic development, cell lineage determination, and cancer differentiation. Actin binding LIM protein 1 (abLIM1) was reported to be localized in a genomic region often deleted in human cancers and suggested to be involved in axon guidance. Recently, existence of a second family member was reported, actin binding LIM protein 2. By means of computational biology and comparative genomics, we now characterized an additional, third member of the actin binding LIM protein subgroup, actin binding LIM protein 3 (abLIM3). The human mRNA sequence was previously annotated as differentially regulated in hepatoblastoma compared to normal livers. Conservation of key structural features of abLIM1 and abLIM2, four LIM domains and a VHD domain, suggested comparable biological function of abLIM3 as a linker between actin cytoskeleton and cell signaling pathways. AbLIM3 was found to be conserved in vertebrates, as orthologous sequences were characterized for mouse, fish, and frog. In addition, we report the existence of abLIM2 orthologs in fish and frog, suggesting a similar degree of evolutionary conservation. The intracellular localization of the abLIM3 protein was predicted to be nuclear by means of Reinhardt's neural network and the k-nearest neighbor algorithm. The corresponding abLIM3 gene was localized to chromosome 5q32 and spanned 119 kb, organized in 24 exons. An RT-PCR based expression profile available from the human unidentified gene-encoded (HUGE) database demonstrated highest expression for abLIM3 in heart, lung, liver, and brain/cerebellum accompanied by lower expression in multiple other tissues. Furthermore, abLIM3 was expressed in fetal liver, CNS, and spinal cord.

Amino Acid Sequence↗

DoD: Database of Databases--updated molecular biology databases.

Database of Databases (DoD) is a collection of molecular biology databases extracted from Nucleic Acids Research, 2005 Database issue. DoD is constructed using javascript and html code. The 14 categories of 719 databases are provided with a search option, keyword help and database description linked to respective pages. Keyword help lists the search strings used to perform an individual search against categorized databases. Database description provides brief information about the database with a link to full-text article and main web page. DoD is available online and can be accessed at http://www.progenebio.in/DoD/DoD.htm.

Computational Biology↗

Fast computer search for similar DNA sequences.

An extremely fast method of searching a nucleic acid sequence database against a probe sequence is described. The method is based on the detection of deviation from expected number and deviation from random spatial distribution of sub-sequences which are unique within a sequence, and shared between that sequence and the probe. On an IBM 3081 computer, total search of an encoded form of the EMBL nucleic acid sequence database with a 1 kbase probe sequence is completed in a few seconds. Previous best methods for a similar task required a few minutes.

Animals↗