PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

The structure analysis and antigenicity study of the N protein of SARS-CoV.

The Coronaviridae family is characterized by a nucleocapsid that is composed of the genome RNA molecule in combination with the nucleoprotein (N protein) within a virion. The most striking physiochemical feature of the N protein of SARS-CoV is that it is a typical basic protein with a high predicted pI and high hydrophilicity, which is consistent with its function of binding to the ribophosphate backbone of the RNA molecule. The predicted high extent of phosphorylation of the N protein on multiple candidate phosphorylation sites demonstrates that it would be related to important functions, such as RNA-binding and localization to the nucleolus of host cells. Subsequent study shows that there is an SR-rich region in the N protein and this region might be involved in the protein-protein interaction. The abundant antigenic sites predicted in the N protein, as well as experimental evidence with synthesized polypeptides, indicate that the N protein is one of the major antigens of the SARS-CoV. Compared with other viral structural proteins, the low variation rate of the N protein with regards to its size suggests its importance to the survival of the virus.

Amino Acid Motifs↗

Object-oriented knowledge bases for the analysis of prokaryotic and eukaryotic genomes.

The amount of biological sequences introduced in the general collections, and the growing complexity of the biological knowledge require the construction of models to formalize this knowledge and particularly the relationships between several data types. Two examples of such situations are presented here, they result from the biological research lead in our team in the field of molecular evolution. ColiGene is a modelling of E. coli genetics devoted to the analysis of relationships between genomic sequences and gene expressivity. MultiMap implements a new formalization of genome maps allowing manipulation of "maps of maps" in two species. Application of ColiGene and MultiMap are not restricted to molecular evolution and, for instance, MultiMap offers new capabilities for infering data on a genome from knowledge on another species. This could be essential for many mapping projects (human, mouse but also other mammals like pig). Development and implementation of those models have been done using an object-oriented knowledge base management system (SHIRKA) interfaced with a dedicated genomic data base management system (ACNUC). Graphical interfaces have been designed to give an environment similar to the biological representations used by biologists.

Animals↗

Hypothesis-driven approach to predict transcriptional units from gene expression data.

MOTIVATION: A major issue in computational biology is the reconstruction of functional relationships among genes, for example the definition of regulatory or biochemical pathways. One step towards this aim is the elucidation of transcriptional units, which are characterized by co-responding changes in mRNA expression levels. These units of genes will allow the generation of hypotheses about respective functional interrelationships. Thus, the focus of analysis currently moves from well-established functional assignment through comparison of protein and DNA sequences towards analysis of transcriptional co-response. Tools that allow deducing common control of gene expression have the potential to complement and extend routine BLAST comparisons, because gene function may be inferred from common transcriptional control. RESULTS: We present a co-clustering strategy of genome sequence information and gene expression data, which was applied to identify transcriptional units within diverse compendia of expression profiles. The phenomenon of prokaryotic operons was selected as an ideal test case to generate well-founded hypotheses about transcriptional units. The existence of overlapping and ambiguous operon definitions allowed the investigation of constitutive and conditional expression of transcriptional units in independent gene expression experiments of Escherichia coli. Our approach allowed identification of operons with high accuracy. Furthermore, both constitutive mRNA co-response as well as conditional differences became apparent. Thus, we were able to generate insight into the possible biological relevance of gene co-response. We conclude that the suggested strategy will be amenable in general to the identification of transcriptional units beyond the chosen example of E.coli operons. AVAILABILITY: The analyses of E.coli transcript data presented here are available upon request or at http://csbdb.mpimp-golm.mpg.de/

Algorithms↗

Mitochondrial DNA analysis of horses recovered from a frozen tomb (Berel site, Kazakhstan, 3rd Century BC).

Sequence polymorphism of the mitochondrial DNA D-loop was used to determine the genetic diversity of horses recovered from a Scythian princely tomb dating from the beginning of the 3rd century BC. Eight haplotypes were found among the 13 ancient horse samples tested. Phylogenetical analysis showed that these ancient horse's sequences, along with two Yakut ones, were distributed throughout the tree defined by modern horses' sequences and are closely related to them. No clear geographical affiliation of the specimens studied was thus determined. Our work, among others, supports the very ancient origin of the matrilines in horses.

Animals↗

Cloning and sequence analysis of the rat augmenter of liver regeneration (ALR) gene: expression of biologically active recombinant ALR and demonstration of tissue distribution.

A full-length cDNA clone encoding a purified augmenter of liver regeneration (ALR) factor prepared from the cytosol of weanling rat livers was isolated. The 1.2-kb cDNA included a 299-bp 5' untranslated region, a 375-bp coding region, and a 550-bp 3' untranslated region. It encoded a protein consisting of 125 amino acids. The molecular weight of ALR calculated from the cDNA was 15,081, which is consistent with the size estimated by SDS/PAGE under reducing conditions. The molecular weight of the purified native ALR estimated by SDS/PAGE under nonreducing conditions was approximately 30,000; thus ALR apparently has a homodimeric structure. The recombinant ALR produced by expression of the cDNA in COS cells was tested in vivo in the canine Eck fistula model and found to have potency equivalent to the purified native ALR. The 125-aa sequence deduced from the rat ALR cDNA shows 50% homology to the amino acid sequence of the gene for oxidative phosphorylation and vegetative growth in the yeast Saccharomyces cerevisiae.

Amino Acid Sequence↗

Determination and analysis of the complete nucleotide sequence of human herpesvirus.

Human herpesvirus 7 (HHV-7) is a recently isolated betaherpesvirus that is prevalent in the human population, with primary infection usually occurring in early childhood. HHV-7 is related to human herpesvirus 6 (HHV-6) in terms of both biological and, from limited prior DNA sequence analysis, genetic criteria. However, extensive analysis of the HHV-7 genome has not been reported, and the precise phylogenetic relationship of HHV-7 to the other human betaherpesviruses HHV-6 and human cytomegalovirus has not been determined. Here I report on the determination and analysis of the complete DNA sequence of HHV-7 strain JI. The data establish that the close biological relationship of HHV-6 and HHV-7 is reflected at the genetic level, where there is a very high degree of conservation of genetic content and encoded amino acid sequences. The data also delineate loci of divergence between the HHV-6 and HHV-7 genomes, which occur at the genome terminal in the region of the terminal direct-repeat elements and within limited regions of the unique component. Of potential significance with respect to biological and evolutionary divergence of HHV-6 and HHV-7 are notable structural differences in putative transcriptional regulatory genes specified by the direct-repeat and immediate-early region A loci of these viruses and the absence of an equivalent of the HHV-6 adeno-associated virus type 2 rep gene homolog in HHV-7.

Amino Acid Sequence↗

Decoding the decoding region: analysis of eukaryotic release factor (eRF1) stop codon-binding residues.

Peptide synthesis in eukaryotes terminates when eukaryotic release factor 1 (eRF1) binds to an mRNA stop codon and occupies the ribosomal A site. Domain 1 of the eRF1 protein has been implicated in stop codon recognition in a number of experimental studies. In order to further pinpoint the residues of this protein involved in stop codon recognition, we sequenced and compared eRF1 genes from a variety of ciliated protozoan species. We then performed a series of computational analyses to evaluate the conservation, accessibility, and structural environment of each amino acid located in domain 1. With this new dataset and methodology, we were able to identify eight specific amino acid sites important for stop codon recognition and also to propose a set of cooperative paired substitutions that may underlie stop codon reassignment. Our results are more consistent with current experimental data than previously described models.

Amino Acid Sequence↗

Alignments anchored on genomic landmarks can aid in the identification of regulatory elements.

MOTIVATION: The transcription start site (TSS) has been located for an increasing number of genes across several organisms. Statistical tests have shown that some cis-acting regulatory elements have positional preferences with respect to the TSS, but few strategies have emerged for locating elements by their positional preferences. This paper elaborates such a strategy. First, we align promoter regions without gaps, anchoring the alignment on each promoter's TSS. Second, we apply a novel word-specific mask. Third, we apply a clustering test related to gapless BLAST statistics. The test examines whether any specific word is placed unusually consistently with respect to the TSS. Finally, our program A-GLAM, an extension of the GLAM program, uses significant word positions as new 'anchors' to realign the sequences. A Gibbs sampling algorithm then locates putative cis-acting regulatory elements. Usually, Gibbs sampling requires a preliminary masking step, to avoid convergence onto a dominant but uninteresting signal from a DNA repeat. However, since the positional anchors focus A-GLAM on the motif of interest, masking DNA repeats during Gibbs sampling becomes unnecessary. RESULTS: In a set of human DNA sequences with experimentally characterized TSSs, the placement of 791 octonucleotide words was unusually consistent (multiple test corrected P < 0.05). Alignments anchored on these words sometimes located statistically significant motifs inaccessible to GLAM or AlignACE. AVAILABILITY: The A-GLAM program and a list of statistically significant words are available at ftp://ftp.ncbi.nih.gov/pub/spouge/papers/archive/AGLAM/.

Amino Acid Motifs↗

A quantization method based on threshold optimization for microarray short time series.

BACKGROUND: Reconstructing regulatory networks from gene expression profiles is a challenging problem of functional genomics. In microarray studies the number of samples is often very limited compared to the number of genes, thus the use of discrete data may help reducing the probability of finding random associations between genes. RESULTS: A quantization method, based on a model of the experimental error and on a significance level able to compromise between false positive and false negative classifications, is presented, which can be used as a preliminary step in discrete reverse engineering methods. The method is tested on continuous synthetic data with two discrete reverse engineering methods: Reveal and Dynamic Bayesian Networks. CONCLUSION: The quantization method, evaluated in comparison with two standard methods, 5% threshold based on experimental error and rank sorting, improves the ability of Reveal and Dynamic Bayesian Networks to identify relations among genes.

Algorithms↗

Logos: a modular bayesian model for de novo motif detection.

The complexity of the global organization and internal structure of motifs in higher eukaryotic organisms raises significant challenges for motif detection techniques. To achieve successful de novo motif detection, it is necessary to model the complex dependencies within and among motifs and to incorporate biological prior knowledge. In this paper, we present LOGOS, an integrated LOcal and GlObal motif Sequence model for biopolymer sequences, which provides a principled framework for developing, modularizing, extending and computing expressive motif models for complex biopolymer sequence analysis. LOGOS consists of two interacting submodels: HMDM, a local alignment model capturing biological prior knowledge and positional dependency within the motif local structure; and HMM, a global motif distribution model modeling frequencies and dependencies of motif occurrences. Model parameters can be fit using training motifs within an empirical Bayesian framework. A variational EM algorithm is developed for de novo motif detection. LOGOS improves over existing models that ignore biological priors and dependencies in motif structures and motif occurrences, and demonstrates superior performance on both semi-realistic test data and cis-regulatory sequences from yeast and Drosophila genomes with regard to sensitivity, specificity, flexibility and extensibility.

Algorithms↗

Molecular characterizations of oxytetracycline resistant bacteria and their resistance genes from mariculture waters of China.

Oxytetracycline-resistant bacteria were isolated from a mariculture farm in China, and accounted for 32.23% and 5.63% of the total culturable microbes of the sea cucumber and the sea urchin rearing waters respectively. Marine vibrios, especially strains related to Vibrio splendidus or V. tasmaniensis, were the most abundant resistant isolates. For oxytetracycline resistance, tet(A), tet(B) and tet(D) genes were detected in both sea cucumber and sea urchin rearing ponds. The dominant resistance type for V. tasmaniensis-like strains was the combination of both tet(A) and tet(B) genes, while the major resistance type for V. splendidus-like strains was a single tet(D) gene. Most of the sea cucumber tet-positive isolates harbored a chloramphenicol-resistance gene, either cat IV or cat II, while only a few sea urchin tet-positive isolates harbored a cat gene, actually cat IV. The coexistence of tet and cat genes in the strains isolated from the mariculture farm studied was helpful in explaining some of the multi-resistance mechanisms.

Aquaculture↗

Partial structure of a large canine cholecystokinin (CCK58): amino acid sequence.

A cholecystokinin molecule larger than any previously chemically characterized was purified from canine proximal small intestine mucosa. The purification procedure consisted of sequential steps of affinity chromatography, gel filtration, and high pressure liquid chromatography. Activity was detected and quantitated by radioimmunoassay with an antibody that recognized the carboxyl terminal sequence of porcine cholecystokinin. Microsequencing of the purified peptide revealed an amino terminal nonadecapeptide sequence (AQKVNSGEPRAHLGALLAR) not present in known cholecystokinin molecules followed by a nonadecapeptide sequence (YIQQARKAPSGRMSVIKNL) that corresponds exactly to the amino terminal sequence of porcine cholecystokinin 39 except for reversed positions of a Met and a Val residue. Based on the sequence analysis, immunoreactivity, and presence of biological activity in two bioassay systems, this peptide, tentatively named cholecystokinin 58, may be a biosynthetic precursor of the smaller forms previously characterized in gastrointestinal and brain tissues.

Amino Acid Sequence↗

Shortest triplet clustering: reconstructing large phylogenies using representative sets.

BACKGROUND: Understanding the evolutionary relationships among species based on their genetic information is one of the primary objectives in phylogenetic analysis. Reconstructing phylogenies for large data sets is still a challenging task in Bioinformatics. RESULTS: We propose a new distance-based clustering method, the shortest triplet clustering algorithm (STC), to reconstruct phylogenies. The main idea is the introduction of a natural definition of so-called k-representative sets. Based on k-representative sets, shortest triplets are reconstructed and serve as building blocks for the STC algorithm to agglomerate sequences for tree reconstruction in O(n2) time for n sequences. Simulations show that STC gives better topological accuracy than other tested methods that also build a first starting tree. STC appears as a very good method to start the tree reconstruction. However, all tested methods give similar results if balanced nearest neighbor interchange (BNNI) is applied as a post-processing step. BNNI leads to an improvement in all instances. The program is available at http://www.bi.uni-duesseldorf.de/software/stc/. CONCLUSION: The results demonstrate that the new approach efficiently reconstructs phylogenies for large data sets. We found that BNNI boosts the topological accuracy of all methods including STC, therefore, one should use BNNI as a post-processing step to get better topological accuracy.

Algorithms↗

A multistep bioinformatic approach detects putative regulatory elements in gene promoters.

BACKGROUND: Searching for approximate patterns in large promoter sequences frequently produces an exceedingly high numbers of results. Our aim was to exploit biological knowledge for definition of a sheltered search space and of appropriate search parameters, in order to develop a method for identification of a tractable number of sequence motifs. RESULTS: Novel software (COOP) was developed for extraction of sequence motifs, based on clustering of exact or approximate patterns according to the frequency of their overlapping occurrences. Genomic sequences of 1 Kb upstream of 91 genes differentially expressed and/or encoding proteins with relevant function in adult human retina were analyzed. Methodology and results were tested by analysing 1,000 groups of putatively unrelated sequences, randomly selected among 17,156 human gene promoters. When applied to a sample of human promoters, the method identified 279 putative motifs frequently occurring in retina promoters sequences. Most of them are localized in the proximal portion of promoters, less variable in central region than in lateral regions and similar to known regulatory sequences. COOP software and reference manual are freely available upon request to the Authors. CONCLUSION: The approach described in this paper seems effective for identifying a tractable number of sequence motifs with putative regulatory role.

Algorithms↗

Oligo kernels for datamining on biological sequences: a case study on prokaryotic translation initiation sites.

BACKGROUND: Kernel-based learning algorithms are among the most advanced machine learning methods and have been successfully applied to a variety of sequence classification tasks within the field of bioinformatics. Conventional kernels utilized so far do not provide an easy interpretation of the learnt representations in terms of positional and compositional variability of the underlying biological signals. RESULTS: We propose a kernel-based approach to datamining on biological sequences. With our method it is possible to model and analyze positional variability of oligomers of any length in a natural way. On one hand this is achieved by mapping the sequences to an intuitive but high-dimensional feature space, well-suited for interpretation of the learnt models. On the other hand, by means of the kernel trick we can provide a general learning algorithm for that high-dimensional representation because all required statistics can be computed without performing an explicit feature space mapping of the sequences. By introducing a kernel parameter that controls the degree of position-dependency, our feature space representation can be tailored to the characteristics of the biological problem at hand. A regularized learning scheme enables application even to biological problems for which only small sets of example sequences are available. Our approach includes a visualization method for transparent representation of characteristic sequence features. Thereby importance of features can be measured in terms of discriminative strength with respect to classification of the underlying sequences. To demonstrate and validate our concept on a biochemically well-defined case, we analyze E. coli translation initiation sites in order to show that we can find biologically relevant signals. For that case, our results clearly show that the Shine-Dalgarno sequence is the most important signal upstream a start codon. The variability in position and composition we found for that signal is in accordance with previous biological knowledge. We also find evidence for signals downstream of the start codon, previously introduced as transcriptional enhancers. These signals are mainly characterized by occurrences of adenine in a region of about 4 nucleotides next to the start codon. CONCLUSIONS: We showed that the oligo kernel can provide a valuable tool for the analysis of relevant signals in biological sequences. In the case of translation initiation sites we could clearly deduce the most discriminative motifs and their positional variation from example sequences. Attractive features of our approach are its flexibility with respect to oligomer length and position conservation. By means of these two parameters oligo kernels can easily be adapted to different biological problems.

Algorithms↗

The phylogenetic distribution of metazoan microRNAs: insights into evolutionary complexity and constraint.

How complex body plans evolved in animals such as fruit flies and vertebrates, as compared to the relatively simple jellyfish and sponges, is not known, given the similarity of developmental genetic repertoires shared by all these taxa. Here, we show that a core set of 18 microRNAs (miRNAs), non-coding RNA molecules that negatively regulate the expression of protein-coding genes, are found only in protostomes and deuterostomes and not in sponges or cnidarians. Because many of these miRNAs are expressed in specific tissues and/or organs, miRNA-mediated regulation could have played a fundamental evolutionary role in the origins of organs such as brain and heart--structures not found in cnidarians or sponges--and thus contributed greatly to the evolution of complex body plans. Furthermore, the continuous acquisition and fixation of miRNAs in various animal groups strongly correlates both with the hierarchy of metazoan relationships and with the non-random origination of metazoan morphological innovations through geologic time.

Animals↗

Identifying projected clusters from gene expression profiles.

In microarray gene expression data, clusters may hide in certain subspaces. For example, a set of co-regulated genes may have similar expression patterns in only a subset of the samples in which certain regulating factors are present. Their expression patterns could be dissimilar when measuring in the full input space. Traditional clustering algorithms that make use of such similarity measurements may fail to identify the clusters. In recent years a number of algorithms have been proposed to identify this kind of projected clusters, but many of them rely on some critical parameters whose proper values are hard for users to determine. In this paper, a new algorithm that dynamically adjusts its internal thresholds is proposed. It has a low dependency on user parameters while allowing users to input some domain knowledge should they be available. Experimental results show that the algorithm is capable of identifying some interesting projected clusters.

Algorithms↗