PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Harnessing the cellular immune system to the gene-prediction cart.

Prediction of genes and verification of their bona fide expression in the cell are major challenges of the post-genomic era. Here, we demonstrate how information from the apparently unrelated field of cellular immunology can be recruited for these challenging tasks. The cellular immune system presents short peptides that are the degradation products of both foreign and self-proteins expressed in the cell. We carried out a comprehensive search comparing these peptides to all accumulated human sequence data. Our findings illustrate how these 'presented self-peptides' are informative for the identification of new genes, for hypothetical gene verification, for verifying gene expression at the protein level and for supporting splice junctions.

Antigen Presentation↗

Increasing molecular diversity of secreted phospholipases A(2) and their receptors and binding proteins.

Secreted phospholipases A(2) (sPLA(2)s) form a large family of structurally related enzymes which are widespread in nature. Snake venoms are known for decades to contain a tremendous molecular diversity of sPLA(2)s which can exert a myriad of toxic and pharmacological effects. Recent studies indicate that mammalian cells also express a variety of sPLA(2)s with ten distinct members identified so far, in addition to the various other intracellular PLA(2)s. Furthermore, scanning of nucleic acid databases fueled by the different genome projects indicates that several sPLA(2)s are also present in invertebrate animals like Drosophila melanogaster as well as in plants. All of these sPLA(2)s catalyze the hydrolysis of glycerophospholipids at the sn-2 position to release free fatty acids and lysophospholipids, and thus could be important for the biosynthesis of biologically active lipid mediators. However, the recent identification of a variety of membrane and soluble proteins that bind to sPLA(2)s suggests that the sPLA(2) enzymes could also function as high affinity ligands. So far, most of the binding data have been accumulated with venom sPLA(2)s and group IB and IIA mammalian sPLA(2)s. Collectively, venom sPLA(2)s have been shown to bind to membrane and soluble mammalian proteins of the C-type lectin superfamily (M-type sPLA(2) receptor and lung surfactant proteins), to pentraxin and reticulocalbin proteins, to factor Xa and to N-type receptors. Venom sPLA(2)s also associate with three distinct types of sPLA(2) inhibitors purified from snake serum that belong to the C-type lectin superfamily, to the three-finger protein superfamily and to proteins containing leucine-rich repeats. On the other hand, mammalian group IB and IIA sPLA(2)s can bind to the M-type receptor, and group IIA sPLA(2)s can associate with lung surfactant proteins, factor Xa and proteoglycans including glypican and decorin, a mammalian protein containing a leucine-rich repeat.

Amino Acid Sequence↗

Characterization of binding sites of eukaryotic transcription factors.

To explore the nature of eukaryotic transcription factor (TF) binding sites and determine how they differ from surrounding DNA sequences, we examined four features associated with DNA-binding sites: G+C content, pattern complexity, palindromic structure, and Markov sequence ordering. Our analysis of the regulatory motifs obtained from the TRANSFAC database, using yeast intergenic sequences as background, revealed that these four features show variable enrichment in motif sequences. For example, motif sequences were more likely to have palindromic structure than were background sequences. In addition, these features were tightly localized to the regulatory motifs, indicating that they are a property of the motif sequences themselves and are not shared by the general promoter "environment" in which the regulatory motifs reside. By breaking down the motif sequences according to the TF classes to which they bind, more specific associations were identified. Finally, we found that some correlations, such as G+C content enrichment, were species-specific, while others, such as complexity enrichment, were universal across the species examined. The quantitative analysis provided here should increase our understanding of protein-DNA interactions and also help facilitate the discovery of regulatory motifs through bioinformatics.

Animals↗

Righting the wrongs.

Explore the source record for details and available documents.

Databases, Nucleic Acid↗

Amino acid sequences of lysozymes newly purified from invertebrates imply wide distribution of a novel class in the lysozyme family.

Lysozymes were purified from three invertebrates: a marine bivalve, a marine conch, and an earthworm. The purified lysozymes all showed a similar molecular weight of 13 kDa on SDS/PAGE. Their N-terminal sequences up to the 33rd residue determined here were apparently homologous among them; in addition, they had a homology with a partial sequence of a starfish lysozyme which had been reported before. The complete sequence of the bivalve lysozyme was determined by peptide mapping and subsequent sequence analysis. This was composed of 123 amino acids including as many as 14 cysteine residues and did not show a clear homology with the known types of lysozymes. However, the homology search of this protein on the protein or nucleic acid database revealed two homologous proteins. One of them was a gene product, CELF22 A3.6 of C. elegans, which was a functionally unknown protein. The other was an isopeptidase of a medicinal leech, named destabilase. Thus, a new type of lysozyme found in at least four species across the three classes of the invertebrates demonstrates a novel class of protein/lysozyme family in invertebrates. The bivalve lysozyme, first characterized here, showed extremely high protein stability and hen lysozyme-like enzymatic features.

Amino Acid Sequence↗

Protein ranking: from local to global structure in the protein similarity network.

Biologists regularly search databases of DNA or protein sequences for evolutionary or functional relationships to a given query sequence. We describe a ranking algorithm that exploits the entire network structure of similarity relationships among proteins in a sequence database by performing a diffusion operation on a precomputed, weighted network. The resulting ranking algorithm, evaluated by using a human-curated database of protein structures, is efficient and provides significantly better rankings than a local network search algorithm such as psi-blast.

Algorithms↗

Evolutionary dynamics of olfactory receptor genes in fishes and tetrapods.

Olfaction, which is an important physiological function for the survival of mammals, is controlled by a large multigene family of olfactory receptor (OR) genes. Fishes also have this gene family, but the number of genes is known to be substantially smaller than in mammals. To understand the evolutionary dynamics of OR genes, we conducted a phylogenetic analysis of all functional genes identified from the genome sequences of zebrafish, pufferfish, frogs, chickens, humans, and mice. The results suggested that the most recent common ancestor between fishes and tetrapods had at least nine ancestral OR genes, and all OR genes identified were classified into nine groups, each of which originated from one ancestral gene. Eight of the nine group genes are still observed in current fish species, whereas only two group genes were found from mammalian genomes, showing that the OR gene family in fishes is much more diverse than in mammals. In mammals, however, one group of genes, gamma, expanded enormously, containing approximately 90% of the entire gene family. Interestingly, the gene groups observed in mammals or birds are nearly absent in fishes. The OR gene repertoire in frogs is as diverse as that in fishes, but the expansion of group gamma genes also occurred, indicating that the frog OR gene family has both mammal- and fish-like characters. All of these observations can be explained by the environmental change that organisms have experienced from the time of the common ancestor of all vertebrates to the present.

Animals↗

Phylogenomic analysis of type I polyketide synthase genes in pathogenic and saprobic ascomycetes.

Fungal type I polyketides (PKs) are synthesized by PK synthases (PKSs) and include well known secondary metabolites such as the anticholesterol drug lovastatin and the potent natural carcinogen aflatoxin. Other type I PKs are known to be virulence factors for some plant pathogens and pigments such as melanin. In this study, a phylogenomic approach was used to investigate the origin and diversity of fungal genes encoding putative PKSs that are predicted to synthesize type I PKs. The resulting genealogy, constructed by using the highly conserved PKS ketosynthase (KS) domain, indicated that: (i). Species within subphylum Pezizomycotina (phylum Ascomycota) but not early diverging ascomycetes, like Saccharomyces cerevisiae (Saccharomycotina) or Schizosaccharomyces pombe (Taphrinomycotina), had large numbers (7-25) of PKS genes. (ii). Bacteria and fungi had separate groups of PKS genes; the few exceptions are the likely result of horizontal gene transfer from bacteria to various sublineages of fungi. (iii). The bulk of genes encoding fungal PKSs fell into eight groups. Four groups were predicted to synthesize variously reduced PKs, and four groups were predicted to make unreduced PKs. (iv). Species within different classes of Pezizomycotina shared the same groups of PKS genes. (v). Different fungal genomes shared few putative orthologous PKS genes, even between closely related genomes in the same class or genus. (vi) The discontinuous distributions of orthologous PKSs among fungal species can be explained by gene duplication, divergence, and gene loss; horizontal gene transfer among fungi does not need to be invoked.

Ascomycota↗

The Leishmania major RNA polymerase II largest subunit lacks a carboxy-terminus heptad repeat structure and its encoding gene is linked with the calreticulin gene.

The gene encoding the RNA polymerase II largest subunit (RPOIILS) has been isolated and sequenced from the kinetoplastid protozoan, Leishmania (Leishmania) major. The RPOIILS gene was shown to be present as a single copy and is composed of an uninterrupted open reading frame of 4.99 kb, specifying a protein 1663 aa in length with a predicted molecular mass of approximately 185 kDa. The carboxy terminus domain (CTD) of the RPOIILS from L. (L.) major, typical of the more evolutionary primitive protozoa, lacked a heptad repeat structure which is present in higher eukaryotes and some other protozoan phyla. Comparison of the predicted aa composition of the CTD from a diverse range of eukaryotic species revealed the abundance of Ser and Pro residues as the only discernible evolutionary conservative feature. A putative ATG start codon for an additional expressed sequence was located 1.1 kb downstream of the L. (L.) major RPOIILS gene stop codon. Nucleic acid database searches revealed the identity of this gene as that encoding the calcium binding protein calreticulin (CLT). The close proximity of the RPOIILS and CLT genes in L. (L.) major raises the possibility that these genes are transcribed as part of the same polycistronic unit.

Amino Acid Sequence↗

A novel thiocationic liposomal formulation of antisense oligonucleotides with activity against Mycobacterium tuberculosis.

This study describes the development of a novel thiocationic (OBEHYTOP) lipid-based formulation of phosphorothioate antisense oligonucleotides (PAOs) showing inhibitory activity against mycobacterium tuberculosis (mTB) as measured by an in vitro BACTEC 460TB assay. PAOs were designed based on sequences complementary to essential regions of the mycobacterial genome from published nucleic acid databases in GenBank. These included the superoxide dismutase sod A gene (TBS3), catalase-peroxidase katG gene (TBK1, TBK10), RNA polymerase beta-subunit rpo B gene (TBR5) and diaminopimelate decarboxylase lys A gene (TBL5). The effect of PAOs (TBS3, K1, K10, R5 and L5) alone on mTB was not significant compared with the no-drug control over a period of exposure of 150 h (ranges of -11.8 to +23.58% at 72 h; 15.26 to +25.82% at 96 h and -5.51 to +24.00% at 150 h). Liposomal formulations (10:5:2 OBEHYTOP:oleic acid:vitamin D3) of PAOs resulted in statistically significant (p < 0.05 in all cases) inhibition (ranges of -51.45 to -63.00% at 72 h; -56.75 to -67.96% at 96 h; -51.45 to -60.26% at 150 h) compared with PAOs alone, thiocationic liposomal control and liposomal components. Positive controls of streptomycin and isoniazid used at their minimum inhibitory concentrations of 2.00 and 0.10 microM, respectively, resulted in average % inhibition values of -94% and -97.36%, respectively, indicating that these thiocationic lipid-formulated PAOs showed inhibitory activity directed against mTB in vitro.

Analysis of Variance↗

Data merging for integrated microarray and proteomic analysis.

The functioning of even a simple biological system is much more complicated than the sum of its genes, proteins and metabolites. A premise of systems biology is that molecular profiling will facilitate the discovery and characterization of important disease pathways. However, as multiple levels of effector pathway regulation appear to be the norm rather than the exception, a significant challenge presented by high-throughput genomics and proteomics technologies is the extraction of the biological implications of complex data. Thus, integration of heterogeneous types of data generated from diverse global technology platforms represents the first challenge in developing the necessary foundational databases needed for predictive modelling of cell and tissue responses. Given the apparent difficulty in defining the correspondence between gene expression and protein abundance measured in several systems to date, how do we make sense of these data and design the next experiment? In this review, we highlight current approaches and challenges associated with integration and analysis of heterogeneous data sets, focusing on global analysis obtained from high-throughput technologies.

Animals↗

Exploiting big biology: integrating large-scale biological data for function inference.

The amount of data produced by molecular biologists is growing at an exponential rate. Some of the fastest growing sets of data are measurements of gene expression, comparable in quantity only to gene sequences and the vast biological literature. Both gene expression data and sequence data offer hints as to the functions of thousands of newly discovered genes, but neither give complete answers. Therefore, much effort is being focused on integrating these large data sets and combining them with all available functional data to draw inferences about the functions of uncharacterised genes. This review discusses the most pertinent functional data for genome-wide functional inference and describes several methods by which these disparate data types are being integrated.

Data Collection↗

Public services from the European Bioinformatics Institute.

The European Bioinformatics Institute (EBI) provides numerous free-of-charge, publicly available bioinformatics services that can be divided into the following categories: ftp downloads; data submissions processing and biological database production; access to query; analysis and retrieval systems and tools; user support; training and education and industry support through EBI's SME program. These services are all available at the website. It is imperative that EBI's data as well as the tools to analyse it efficiently are made available in a free and unambiguous way to the scientific community. An important part of the EBI's mission is to make this happen in a fast, reliable and efficient manner. This paper serves as a brief introduction to each of these services.

Computational Biology↗