PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

Distinguishing protein-coding from non-coding RNAs through support vector machines.

RIKEN's FANTOM project has revealed many previously unknown coding sequences, as well as an unexpected degree of variation in transcripts resulting from alternative promoter usage and splicing. Ever more transcripts that do not code for proteins have been identified by transcriptome studies, in general. Increasing evidence points to the important cellular roles of such non-coding RNAs (ncRNAs). The distinction of protein-coding RNA transcripts from ncRNA transcripts is therefore an important problem in understanding the transcriptome and carrying out its annotation. Very few in silico methods have specifically addressed this problem. Here, we introduce CONC (for "coding or non-coding"), a novel method based on support vector machines that classifies transcripts according to features they would have if they were coding for proteins. These features include peptide length, amino acid composition, predicted secondary structure content, predicted percentage of exposed residues, compositional entropy, number of homologs from database searches, and alignment entropy. Nucleotide frequencies are also incorporated into the method. Confirmed coding cDNAs for eukaryotic proteins from the Swiss-Prot database constituted the set of true positives, ncRNAs from RNAdb and NONCODE the true negatives. Ten-fold cross-validation suggested that CONC distinguished coding RNAs from ncRNAs at about 97% specificity and 98% sensitivity. Applied to 102,801 mouse cDNAs from the FANTOM3 dataset, our method reliably identified over 14,000 ncRNAs and estimated the total number of ncRNAs to be about 28,000.

Animals↗

eMelanoBase: an online locus-specific variant database for familial melanoma.

A proportion of melanoma-prone individuals in both familial and non-familial contexts has been shown to carry inactivating mutations in either CDKN2A or, rarely, CDK4. CDKN2A is a complex locus that encodes two unrelated proteins from alternately spliced transcripts that are read in different frames. The alpha transcript (exons 1alpha, 2, and 3) produces the p16INK4A cyclin-dependent kinase inhibitor, while the beta transcript (exons 1beta and 2) is translated as p14ARF, a stabilizing factor of p53 levels through binding to MDM2. Mutations in exon 2 can impair both polypeptides and insertions and deletions in exons 1alpha, 1beta, and 2, which can theoretically generate p16INK4A-p14ARF fusion proteins. No online database currently takes into account all the consequences of these genotypes, a situation compounded by some problematic previous annotations of CDKN2A-related sequences and descriptions of their mutations. As an initiative of the international Melanoma Genetics Consortium, we have therefore established a database of germline variants observed in all loci implicated in familial melanoma susceptibility. Such a comprehensive, publicly accessible database is an essential foundation for research on melanoma susceptibility and its clinical application. Our database serves two types of data as defined by HUGO. The core dataset includes the nucleotide variants on the genomic and transcript levels, amino acid variants, and citation. The ancillary dataset includes keyword description of events at the transcription and translation levels and epidemiological data. The application that handles users' queries was designed in the model-view-controller architecture and was implemented in Java. The object-relational database schema was deduced using functional dependency analysis. We hereby present our first functional prototype of eMelanoBase. The service is accessible via the URL www.wmi.usyd.edu.au:8080/melanoma.html.

Computer Security↗

Applications of computational algorithm tools to identify functional SNPs in cytokine genes.

Understanding the functions of single nucleotide polymorphisms (SNPs) can greatly help to understand the genetics of the human phenotype variation and especially the genetic basis of human complex diseases. However, how to identify functional SNPs from a pool containing both functional and neutral SNPs is challenging. In this study, we analyzed the genetic variations that can alter the expression and function of a group of cytokine proteins using computational tools. As a result, we extracted 4552 SNPs from 45 cytokine proteins from SNPper database. Of particular interest, 828 SNPs were in the 5'UTR region, 961 SNPs were in the 3' UTR region, and 85 SNPs were non-synonymous SNPs (nsSNPs), which cause amino acid change. Evolutionary conservation analysis using the SIFT tool suggested that 8 nsSNPs may disrupt the protein function. Protein structure analysis using the PolyPhen tool suggested that 5 nsSNPs might alter protein structure. Binding motif analysis using the UTResource tool suggested that 27 SNPs in 5' or 3'UTR might change protein expression levels. Our study demonstrates the presence of naturally occurring genetic variations in the cytokine proteins that may affect their expressions and functions with possible roles in complex human disease, such as immune diseases.

Algorithms↗

Temperature and the expression of myogenic regulatory factors (MRFs) and myosin heavy chain isoforms during embryogenesis in the common carp Cyprinus carpio L.

Embryos of the common carp, Cyprinus carpio L., were reared from fertilization of the eggs to inflation of the swim bladder in the larval stage at 18 and 25 degrees C. cRNA probes were used to detect transcripts of the myogenic regulatory factors MyoD, Myf-5 and myogenin, and five myosin heavy chain (MyHC) isoforms during development. The genes encoding Myf-5 and MyoD were switched on first in the unsegmented mesoderm, followed by myogenin as the somites developed. Myf-5 and MyoD transcripts were initially limited to the adaxial cells, but Myf-5 expression spread laterally into the presomitic mesoderm before somite formation. Two distinct bands of staining could be seen corresponding to the cellular fields of the forming somites, but as each furrow delineated, Myf-5 mRNA levels declined. Upon somite formation, MyoD expression spread laterally to encompass the full somite width. Expression of the myogenin gene was also switched on during somite formation, and expression of both transcripts persisted until the somites became chevron-shaped. Expression of MyoD was then downregulated shortly before myogenin. The expression patterns of the carp myogenic regulatory factor (MRF) genes most-closely resembled that seen in the zebrafish rather than the rainbow trout (where expression of MyoD remains restricted to the adaxial domain of the somite for a prolonged period) or the herring (where expression of MyoD persists longer than that of myogenin). Expression of two embryonic forms of MyHC began simultaneously at the 25-30 somite stage and continued until approximately two weeks post-hatch. However, the three adult isoforms of fast muscle MyHC were not detected in any stage examined, emphasizing a developmental gap that must be filled by other, as yet uncharacterised, MyHC isoform(s). No differences in the timing of expression of any mRNA transcripts were seen between temperature groups. A phylogenetic analysis of the MRFs was conducted using all available full-length amino acid sequences. A neighbour-joining tree indicated that all four members evolved from a common ancestral gene, which first duplicated into two lineages, each of which underwent a further duplication to produce Myf-5 and MyoD, and myogenin and MRF4. Parologous copies of MyoD from trout and Xenopus clustered closely together within clades, indicating recent duplications. By contrast, MyoD paralogues from gilthead seabream were more divergent, indicating a more-ancient duplication.

Animals↗

Splicing profile based protein categorization between human and mouse genomes by use of the DDBJ Web services.

In one scenario of gene evolution, exon shuffling plays a fundamental role in increasing gene diversity. This paper is an appraisal of the biological relevance of categorising proteins by their splicing profiles (exon-intron structures). The central question is whether protein function is more correlated with splicing profiles than sequence similarity, or not. To approach this question, a splicing profile similarity (SPS) index, which measures relative exon length discrepancy, was devised. Arbitrary human proteins were compared, in terms of SPS and amino acid sequence similarity, to their 1) mouse orthologues and 2) human paralogues, which epitomise functional equivalence and non-equivalence, respectively, to methodically elucidate the global relationship between a) biological function, b) splicing profile similarity, and c) sequence similarity. Protein function is more correlated with splicing profile similarity than sequence similarity as demonstrated by the fact that human-mouse orthologues (HMOs) display significantly higher splicing profile similarity than do human-human paralogues (HHPs), despite the mutual sequence similarity between these two categories. This finding indicates that splicing profile-based protein categorisation is biologically meaningful.

Alternative Splicing↗

Latent periodicity of 21 bases typical for MCP II gene is widely present in various bacterial genes.

The existence of a typical latent periodicity of 21 bases from the Tar chemoreceptor gene of Escherichia coli (E. coli) (MCP II) in the bacterial genes has been investigated in this work. Among 583 annotated bacterial genes and ORFs in the GenBank, in which the typical periodicity has been found, the chemoreceptors' genes constituted the most numerous group (18.5%). This typical latent periodicity of 21 bases has been revealed in many different genes of regulatory proteins, DNA polymerases, reductases, kinases and others. The numbers in such gene groups varied from 1 to 4% of the total analyzed genes. The 2D-structures analysis of the amino acid residues, which have been translated from the genes' regions with 21 bases periodicity, has shown that, though the enrichment of alpha-helical structures in such sequences is kept in all cases, it is seen that the latent periodicity of 21 bases is a very sensitively tuned basis, allowing the translated residues to smoothly change from one conformation to another. Interesting results have been obtained for 16S rRNAs genes of proteobacteria. Short sequences-determinants have been revealed in the genes, which select beta and gamma proteobacteria with an accuracy of above 90%.

Bacterial Proteins↗

Methylation and repeats in silent and nonsense mutations of p53.

All exonic CG sequences in p53 are methylated; this epigenetic modification is correlated with frequent G:C-->A:T transitions in p53. Recent reports reveal the presence in p53 of non-CG methylation in CC and CCC sequences, complementary to sites of selective guanosine adduct formation (GG and GGG), and the association of genetic instability with methylation at repetitive sequences. We presently investigated the distribution of methylation sites and repetitive elements in silent and nonsense p53 mutations (2051) among the IARC's TP53 somatic mutation database for exons 5-8. Silent mutations are nonrandom, but mostly involve G:C-->A:T transitions (62%); in particular C-->T mutations (39% of all silent mutations) are mostly correlated with CC and CCC sequences, while G-->A mutations with GG sequences. Sequence analysis of all non-G:C-->A:T silent mutations reveals the frequent formation of new methylation sites (CG), new CCC and GGG sequences in the resulting sequence, refinement of symmetry elements at interrupted microsatellite-like sequences and formation of small repeats (55.3%). The G:C-->A:T silent mutations characterize cancers associated with cigarette smoking (e.g. bladder or lung and bronchus cancer versus colorectal cancer); on the contrary, non-G:C-->A:T silent mutations have similar frequencies in most cancers. Nonsense mutations in exons 5-8, all resulting in mutants lacking amino acids 307-393, which are crucial for p53 activity, were also analyzed. The frequency of nonsense mutations is higher at methylated sites or repeats 1-2 nucleotides removed from methylation sites. Frameshift mutations are also more frequent at repeated sequences. The frequent G:C-->A:T silent mutations could indicate that CC and CCC sequences of exons 5-8 are occasionally targets of non-CpG methylation of cytosine. This process of de novo methylation in the presence of microsatellite-like sequences and small repeats might influence the genetic stability of a variety of genes.

Base Sequence↗

Phylogenetic analysis, genome evolution and the rate of gene gain in the Herpesviridae.

We used complete sequence data from 30 complete Herpesviridae genomes to investigate phylogenetic relationships and patterns of genome evolution. The approach was to identify orthologous gene clusters among taxa and to generate a genomic matrix of gene content. We identified 17 genes with homologs in all 30 taxa and concatenated a subset of 10 of these genes for phylogenetic inference. We also constructed phylogenetic trees on the basis of gene content data. The amino acid and gene content phylogenies were largely concordant, but the amino acid data had much higher internal support. We mapped gene gain events onto the phylogenetic tree by assuming that genes were gained only once during the evolution of herpesviruses. Thirty genes were inferred to be present in the ancestor of all herpesvirus, a number smaller than previously hypothesized. Few genes of recent origin within herpesviruses could be identified as originating from transfer between virus and vertebrate hosts. Inferred rates of gene gain were heterogeneous, with both taxonomic and temporal biases. Nonetheless, the average rate of gene gain was approximately 3.5 x 10(-7) genes gained per year, which is an order of magnitude higher than the nucleotide mutation rate for these large DNA viruses.

Databases, Nucleic Acid↗

Whole proteome prokaryote phylogeny without sequence alignment: a K-string composition approach.

A systematic way of inferring evolutionary relatedness of microbial organisms from the oligopeptide content, i.e., frequency of amino acid K-strings in their complete proteomes, is proposed. The new method circumvents the ambiguity of choosing the genes for phylogenetic reconstruction and avoids the necessity of aligning sequences of essentially different length and gene content. The only "parameter" in the method is the length K of the oligopeptides, which serves to tune the "resolution power" of the method. The topology of the trees converges with K increasing. Applied to a total of 109 organisms, including 16 Archaea, 87 Bacteria, and 6 Eukarya, it yields an unrooted tree that agrees with the biologists' "tree of life" based on SSU rRNA comparison in a majority of basic branchings, and especially, in all lower taxa.

Algorithms↗

Monomorphism of human cytochrome c.

Cytochrome c (Cyt c) has key roles in both mitochondrial electron transfer and apoptosis onset and is therefore likely undergoing a strong selective pressure against amino acid variation. Nevertheless, a phylogenetically fast amino acid replacement rate in the Cyt c of species of the anthropoid primate lineage was recently reported. We therefore looked for the presence of nonsynonymous single nucleotide polymorphisms (nsSNPs) in the human Cyt c (HGNC approved gene symbol: CYCS), which, given its cellular constraints, could have important functional consequences, and found a large number of putative nsSNPs reported in the dbSNP database. We then subjected these putative SNPs to experimental validation by sequencing the Cyt c gene in a panel of 95 individuals assumed as a standard reference of the human population diversity. Surprisingly, none of the putative SNPs survived experimental validation. We conclude that non-rare allelic variants of the Cyt c protein are absent in the human populations analyzed in this study.

Alleles↗

cfMethDB: A Comprehensive cfDNA Methylation Data Resource for Cancer Biomarkers.

Cancer is a major global health threat, and early detection is crucial for improving patient outcomes. DNA methylation in circulating cell-free DNA (cfDNA) has emerged as a promising biomarker for non-invasive cancer diagnosis. However, the integration and utilization of existing cfDNA methylation data have been limited, hindering comprehensive research efforts, particularly in the discovery of cfDNA methylation biomarkers. To address this challenge, we introduced cfMethDB, a comprehensive database dedicated to cfDNA methylation in cancer that encompasses 4828 publicly available datasets. Through standardized analysis, we identified 1,048,770 differentially methylated cytosines (DMCs) as candidate biomarkers across seven cancer types. With cfMethDB, we not only identified known cfDNA methylation biomarkers, but also discovered several genes, such as ZIC4, that could be novel biomarkers. Moreover, cfMethDB offers a suite of user-friendly tools, including biomarker evaluation, pan-cancer search, and end motif analysis. We hope that cfMethDB will serve as a valuable platform for the discovery of novel cancer cfDNA methylation biomarkers and facilitate cancer research and clinical applications. cfMethDB is publicly available at https://cfmethdb.hzau.edu.cn/home.

Humans↗

Genomic analysis of influenza A viruses, including avian flu (H5N1) strains.

This study was designed to conduct genomic analysis in two steps, such as the overall relative synonymous codon usage (RSCU) analysis of the five virus species in the orthomyxoviridae family, and more intensive pattern analysis of the four subtypes of influenza A virus (H1N1, H2N2, H3N2, and H5N1) which were isolated from human population. All the subtypes were categorized by their isolated regions, including Asia, Europe, and Africa, and most of the synonymous codon usage patterns were analyzed by correspondence analysis (CA). As a result, influenza A virus showed the lowest synonymous codon usage bias among the virus species of the orthomyxoviridae family, and influenza B and influenza C virus were followed, while suggesting that influenza A virus might have an advantage in transmitting across the species barrier due to their low codon usage bias. The ENC values of the host-specific HA and NA genes represented their different HA and NA types very well, and this reveals that each influenza A virus subtype uses different codon usage patterns as well as the amino acid compositions. In NP, PA and PB2 genes, most of the virus subtypes showed similar RSCU patterns except for H5N1 and H3N2 (A/HK/1774/1999) subtypes which were suspected to be transmitted across the species barrier, from avian and porcine species to human beings, respectively. This distinguishable synonymous codon usage patterns in non-human origin viruses might be useful in determining the origin of influenza A viruses in genomic levels as well as the serological tests. In this study, all the process, including extracting sequences from GenBank flat file and calculating codon usage values, was conducted by Java codes, and these bioinformatics-related methods may be useful in predicting the evolutionary patterns of pandemic viruses.

Africa↗

Identification of Ugandan HIV type 1 variants with unique patterns of recombination in pol involving subtypes A and D.

Most HIV-1 infections in Uganda are caused by subtypes A and D. The prevalence of recombination and the sites of specific breakpoints between these subtypes have not been reported. HIV-1 pol sequences encoding protease (amino acids 1-99) and reverse transcriptase (amino acids 1-324) from 102 pregnant Ugandan women were analyzed by the Recombinant Identification Program, SimPlot, and examination of phylogenetically informative sites to identify sites of recombination between sequence segments belonging to different subtypes. Thirteen percent (13 of 102) of the pol sequences contained strong evidence of recombination between subtypes A and D. At least nine different patterns of recombination were observed. Five women infected with a recombinant virus transmitted the recombinant virus perinatally. In this population-based study, intersubtype recombinants were common. The large number of different types of pol recombinants identified suggests that recombination occurs readily in the pol region. Perinatal transmission of the recombinant viruses demonstrates their evolutionary stability.

Databases, Nucleic Acid↗

Defining genes in the genome of the hyperthermophilic archaeon Pyrococcus furiosus: implications for all microbial genomes.

The original genome annotation of the hyperthermophilic archaeon Pyrococcus furiosus contained 2,065 open reading frames (ORFs). The genome was subsequently automatically annotated in two public databases by the Institute for Genomic Research (TIGR) and the National Center for Biotechnology Information (NCBI). Remarkably, more than 500 of the originally annotated ORFs differ in size in the two databases, many very significantly. For example, more than 170 of the predicted proteins differ at their N termini by more than 25 amino acids. Similar discrepancies were observed in the TIGR and NCBI databases with the other archaeal and bacterial genomes examined. In addition, the two databases contain 60 (NCBI) and 221 (TIGR) ORFs not present in the original annotation of P. furiosus. In the present study we have experimentally assessed the validity of 88 previously unannotated ORFs. Transcriptional analyses showed that 11 of 61 ORFs examined were expressed in P. furiosus when grown at either 95 or 72 degrees C. In addition, 7 of 54 ORFs examined yielded heat-stable recombinant proteins when they were expressed in Escherichia coli, although only one of the seven ORFs was expressed in P. furiosus under the growth conditions tested. It is concluded that the P. furiosus genome contains at least 17 ORFs not previously recognized in the original annotation. This study serves to highlight the discrepancies in the public databases and the problems of accurately defining the number and sizes of ORFs within any microbial genome.

Archaeal Proteins↗

Cloning, mapping and association study with carcass traits of the porcine SDHD gene.

A 1320-bp cDNA containing the full coding region of the porcine succinate dehydrogenase complex, subunit D (SDHD) gene was obtained by random sequencing of clones from a Chinese Tongcheng pig 55-day fetal longissimus dorsi muscle cDNA library. Analysis of the SDHD gene across the INRA-University of Minnesota porcine radiation hybrid panel indicated close linkage with microsatellite marker SW2401, located on SSC9p21. The open reading frame of this cDNA covers 480 bp and encodes 159 amino acids. The deduced porcine amino acid sequence showed greater similarity with human and bovine protein sequences than with those from mouse and rat. The BLAST analysis of the porcine SDHD to NCBI identified Unigene Cluster Ssc.2586. Possible single nucleotide polymorphisms (SNP) were identified by alignment of expressed sequence tags in the cluster. The polymerase chain reaction (PCR) single strand conformation polymorphism, sequencing, and PCR restriction fragment length polymorphism were used to confirm and detect a synonymous polymorphic MboI site within the open-reading frame. Allele frequencies of this SNP were investigated in two commercial and five Chinese local pig breeds. These five Chinese breeds had very high frequencies for one allele, whereas frequencies of both alleles were intermediate in Large White and Duroc. An association analysis suggested that different SDHD genotypes have significant differences in loin-muscle area (P < 0.01).

Animals↗

Specific identification and molecular typing analysis of Lactobacillus johnsonii by using PCR-based methods and pulsed-field gel electrophoresis.

A fast and reliable Multiplex-PCR assay was established to identify the species Lactobacillus johnsonii. Two opposing rRNA gene-targeted primers have been designed for this specific PCR detection. Specificity was verified with DNA samples isolated from different lactic acid bacteria. Out of 47 Lactobacillus strains isolated from different environments, 16 were identified as L. johnsonii by PCR. The same set of strains was investigated with five alternative molecular typing methods: enterobacterial repetitive intergenic consensus PCR (ERIC-PCR), repetitive extragenic palindromic PCR (REP-PCR), amplified fragment length polymorphism, single triplicate arbitrarily primed PCR, and pulsed-field gel electrophoresis in order to compare the discriminatory power of these methods. The reported data strongly support the highly significant heterogeneity among all L. johnsonii isolates, potentially linked to their origin of isolation. The use of species-specific primers as well as rapid and highly powerful PCR-based molecular typing tools (namely ERIC- and REP-PCR techniques) should be respectively envisaged for identifying, differentiating and monitoring L. johnsonii strains from various environmental samples, for product monitoring, for species tracing in clinical studies as well as bacterial profiling of various microecological or gastrointestinal environments.

Bacterial Typing Techniques↗

Alternatively and constitutively spliced exons are subject to different evolutionary forces.

There has been a controversy on whether alternatively spliced exons (ASEs) evolve faster than constitutively spliced exons (CSEs). Although it has been noted that ASEs are subject to weaker selective constraints than CSEs, so they evolve faster, there have also been studies that indicated slower evolution in ASEs than in CSEs. In this study, we retrieve more than 5,000 human-mouse orthologous exons and calculate the synonymous (KS) and nonsynonymous (KA) substitution rates in these exons. Our results show that ASEs have higher KA values and higher KA/KS ratios than CSEs, indicating faster amino acid-level evolution in ASEs. The faster evolution may be in part due to weaker selective constraints. It is also possible that the faster rate is in part due to faster functional evolution in ASEs. On the other hand, the majority of ASEs have lower KS values than CSEs. With reference to the substitution rate in introns, we show that the KS values in ASEs are close to the neutral substitution rate, whereas the synonymous substitution rate in CSEs has likely been accelerated. The elevated synonymous rate in CSEs is not related to CpG dinucleotides or low-complexity regions of protein but may be weakly related to codon usage bias. The overall trends of higher KA and lower KS in ASEs than in CSEs are also observed in human-rat and mouse-rat comparisons. Therefore, our observations hold for mammals of different molecular clocks.

Animals↗

Construction and use of a computerized DNA fingerprint database for lactic acid bacteria from silage.

Efficient selection of new silage inoculant strains from a collection of over 10,000 isolates of lactic acid bacteria (LAB) requires excellent strain discrimination. Toward that end, we constructed a GelCompar II database of DNA fingerprint patterns of ethidium bromide-stained EcoRI fragments of total LAB DNA separated by conventional agarose gel electrophoresis. We found that the total DNA patterns were strain-specific; 56/60 American Type Culture Collection strains of 33 species of LAB could be distinguished. Enterococcus faecium strains ATCC19434 and ATCC35667 had identical total DNA patterns and RiboPrints. Lactobacillus rhamnosus strains ATCC7469 and ATCC27773 also had identical total DNA patterns, but different RiboPrints. EcoRI RiboPrint patterns could distinguish only about 9/23 Lactobacillus plantarum strains and about 6/10 Lactobacillus buchneri strains, whereas all 33 strains could be distinguished by EcoRI total DNA patterns. Despite gel-to-gel variation, new DNA patterns can be readily grouped with existing patterns using GelCompar II. The database contains large homogenous clusters of L. plantarum, E. faecium, L. buchneri, Lactobacillus brevis and Pediococcus species that can be used for tentative taxonomic assignment. We routinely use the DNA fingerprint database to identify and characterize new strains, eliminate duplicate isolates and for quality control of inoculant product strains. The GelCompar II database has been in continuous use for 7 years and contains more than 3600 patterns representing approximately 700 unique patterns from over 300 gels and is the largest computerized DNA fingerprint database for LAB yet reported.

DNA Fingerprinting↗