PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Comparative analysis of selected genes from Diachasmimorpha longicaudata entomopoxvirus and other poxviruses.

The Diachasmimorpha longicaudata entomopoxvirus (DlEPV) is the first symbiotic EPV described from a parasitic wasp. The DlEPV is introduced into the tephritid fruit fly larval host along with the wasp egg at oviposition. We sequenced a shotgun genomic library of the DlEPV DNA and analyzed and compared the predicted protein sequences of eight ORFs with those of selected poxviruses and other organisms. BlastP searches showed that five of these are homologous to poxvirus putative proteins such as metalloprotease, a putative membrane protein, late transcription factor-3, virion surface protein, and poly (A) polymerase (PAP) regulatory small subunit. Three of these are similar to those of other organisms such as the gamma-glutamyltransferase (GGT) of Arabidopsis thaliana, eukaryotic initiation factor 4A (eIF4A) of Caenorhabditis briggsae and lambda phage integrase (lambda-Int) of Enterococcus faecium. Transcription motifs for early (TGA,A/T,XXXXA) or late (TAAATG, TAAT, or TAAAT) gene expression conserved in poxviruses were identified with those ORFs. Phylogenetic analysis of multiple alignments of five ORFs and 20 poxvirus homologous sequences and of a concatenate of multiple alignments suggested that DlEPV probably diverged from the ancestral node between the fowlpox virus and the genus B, lepidopteran and orthopteran EPVs, to which Amsacta moorei and Melanoplus sanguinipes EPV, respectively, belong. The DlEPV putative GGT, eIF4A, and lambda-Int contained many conserved domains that typified these proteins. These homologues may be involved in either viral pathogenicity or enhancing parasitism via the gamma-glutamyl cycle and compensation of eIF4A levels in the parasitized fly, or via the integration of a portion of the viral genome into the wasp and/or parasitized fly.

Amino Acid Motifs↗

A fast and efficient method for isolation of the BAC end.

As a new developmental vector system, the bacterial artificial chromosome (BAC) has been used widely in constructing genomic libraries and in generating transgenic animals. Isolation of the BAC insert end is useful to analyze the BAC clone. Here, we describe a fast and efficient method to obtain the BAC end by ligating the BAC fragments digested with Not I and another selected restriction enzyme into universal cloning vector, followed by determining the correct clones with HindIII digestion. Further DNA sequencing analysis verified the results mentioned above.

Chromosomes, Artificial, Bacterial↗

Occurrence and characterization of actinobacteria and thermoactinomycetes isolated from pulp and board samples containing recycled fibres.

The aim of this study was to characterize the actinobacterial population present in pulps and boards containing recycled fibres. A total of 107 isolates was identified on the basis of their pigmentation, morphological properties, fatty acid profiles and growth temperature. Of the wet pulp and water sample isolates (n=87), 74.7% belonged to the genus Streptomyces, 17.2% to Nocardiopsis and 8.0% to thermoactinomycetes, whereas all the board sample isolates (n=20) were thermoactinomycetes. The identification of 53 isolates was continued by molecular methods. Partial 16S rDNA sequencing and automated ribotyping divided the Streptomyces isolates (n=31) into 14 different taxa. The most common streptomycetes were the mesophilic S. albidoflavus and moderately thermophilic S. thermocarboxydus. The Nocardiopsis isolates (n=11) belonged to six different taxa, whereas the thermoactinomycetes were mainly members of the species Laceyella sacchari (formerly Thermoactinomyces sacchari). The results indicated the probable presence of one or more new species within each of these genera. Obviously, the drying stage used in the board making processes had eliminated all members of the species Streptomyces and Nocardiopsis present in the wet recycled fibre pulp samples. Only the thermotolerant endospores of L. sacchari were still present in the final products. The potential of automated ribotyping for identifying actinobacteria was indicated, as soon as comprehensive identification libraries became available.

Actinobacteria↗

Characterizing the regulatory logic of transcriptional control at the DNA sequence level by ensembles of thermodynamic models.

MOTIVATION: Understanding how the genome encodes the regulatory logic of transcription is a main challenge of the post-genomic era, and can be overcome with the aid of customized computational tools. RESULTS: We report an automated framework for analyzing an ensemble of fits to data of a thermodynamics-based sequence-level model for transcriptional regulation. The fits are clustered accordingly with their intrinsic regulatory logic. A multiscale analysis enables visualization of quantitative features resulting from the deconvolution of the regulatory profile provided by multiple transcription factors interacting with the locus of a gene. Quantitative experimental data on reporters driven by the whole locus of the even-skipped gene in the blastoderm of Drosophila embryos was used for validating our approach. A few clusters of highly active DNA binding sites within the enhancers collectively modulate even-skipped gene transcription. Analysis of variable enhancers' length shows the importance of bound protein-protein interactions for transcriptional regulation. The interplay between activation and quenching enables function conservation of enhancers despite length variations. AVAILABILITY AND IMPLEMENTATION: The transcription factor level data used for performing the reported study is accessible in the input files in Zenodo and GitHub as well the full code. Additional data from formerly FlyEx database will be available under request.

Thermodynamics↗

MILANO--custom annotation of microarray results using automatic literature searches.

BACKGROUND: High-throughput genomic research tools are becoming standard in the biologist's toolbox. After processing the genomic data with one of the many available statistical algorithms to identify statistically significant genes, these genes need to be further analyzed for biological significance in light of all the existing knowledge. Literature mining--the process of representing literature data in a fashion that is easy to relate to genomic data--is one solution to this problem. RESULTS: We present a web-based tool, MILANO (Microarray Literature-based Annotation), that allows annotation of lists of genes derived from microarray results by user defined terms. Our annotation strategy is based on counting the number of literature co-occurrences of each gene on the list with a user defined term. This strategy allows the customization of the annotation procedure and thus overcomes one of the major limitations of the functional annotations usually provided with microarray results. MILANO expands the gene names to include all their informative synonyms while filtering out gene symbols that are likely to be less informative as literature searching terms. MILANO supports searching two literature databases: GeneRIF and Medline (through PubMed), allowing retrieval of both quick and comprehensive results. We demonstrate MILANO's ability to improve microarray analysis by analyzing a list of 150 genes that were affected by p53 overproduction. This analysis reveals that MILANO enables immediate identification of known p53 target genes on this list and assists in sorting the list into genes known to be involved in p53 related pathways, apoptosis and cell cycle arrest. CONCLUSIONS: MILANO provides a useful tool for the automatic custom annotation of microarray results which is based on all the available literature. MILANO has two major advances over similar tools: the ability to expand gene names to include all their informative synonyms while removing synonyms that are not informative and access to the GeneRIF database which provides short summaries of curated articles relevant to known genes. MILANO is available at http://milano.md.huji.ac.il.

Algorithms↗

Metric for measuring the effectiveness of clustering of DNA microarray expression.

BACKGROUND: The recent advancement of microarray technology with lower noise and better affordability makes it possible to determine expression of several thousand genes simultaneously. The differentially expressed genes are filtered first and then clustered based on the expression profiles of the genes. A large number of clustering algorithms and distance measuring matrices are proposed in the literature. The popular ones among them include hierarchal clustering and k-means clustering. These algorithms have often used the Euclidian distance or Pearson correlation distance. The biologists or the practitioners are often confused as to which algorithm to use since there is no clear winner among algorithms or among distance measuring metrics. Several validation indices have been proposed in the literature and these are based directly or indirectly on distances; hence a method that uses any of these indices does not relate to any biological features such as biological processes or molecular functions. RESULTS: In this paper we have proposed a metric to measure the effectiveness of clustering algorithms of genes by computing inter-cluster cohesiveness and as well as the intra-cluster separation with respect to biological features such as biological processes or molecular functions. We have applied this metric to the clusters on the data set that we have created as part of a larger study to determine the cancer suppressive mechanism of a class of chemicals called retinoids.We have considered hierarchal and k-means clustering with Euclidian and Pearson correlation distances. Our results show that genes of similar expression profiles are more likely to be closely related to biological processes than they are to molecular functions. The findings have been supported by many works in the area of gene clustering. CONCLUSION: The best clustering algorithm of genes must achieve cohesiveness within a cluster with respect to some biological features, and as well as maximum separation between clusters in terms of the distribution of genes of a behavioral group across clusters. We claim that our proposed metric is novel in this respect and that it provides a measure of both inter and intra cluster cohesiveness. Best of all, computation of the proposed metric is easy and it provides a single quantitative value, which makes comparison of different algorithms easier. The maximum cluster cohesiveness and the maximum intra-cluster separation are indicated by the metric when its value is 0.We have demonstrated the metric by applying it to a data set with gene behavioral groupings such as biological process and molecular functions. The metric can be easily extended to other features of a gene such as DNA binding sites and protein-protein interactions of the gene product, special features of the intron-exon structure, promoter characteristics, etc. The metric can also be used in other domains that use two different parametric spaces; one for clustering and the other one for measuring the effectiveness.

Algorithms↗

Network theory to understand microarray studies of complex diseases.

Complex diseases, such as allergy, diabetes and obesity depend on altered interactions between multiple genes, rather than changes in a single causal gene. DNA microarray studies of a complex disease often implicate hundreds of genes in the pathogenesis. This indicates that many different mechanisms and pathways are involved. How can we understand such complexity? How can hypotheses be formulated and tested? One approach is to organize the data in network models and to analyze these in a top-down manner. Globally, networks in nature are often characterized by a small number of highly connected nodes, while the majority of nodes have few connections. The highly connected nodes serve as hubs that affect many other nodes. Such hubs have key roles in the network. In yeast cells, for example, deletion of highly connected proteins is associated with increased lethality, compared to deletion of less connected proteins. This suggests the biological relevance of networks. Moving down in the network structure, there may be sub-networks or modules with specific functions. These modules may be further dissected to analyze individual nodes. In the context of DNA microarray studies of complex diseases, gene-interaction networks may contain modules of co-regulated or interacting genes that have distinct biological functions. Such modules may be linked to specific gene polymorphisms, transcription factors, cellular functions and disease mechanisms. Genes that are reliably active only in the context of their modules can be considered markers for the activity of the modules and may thus be promising candidates for biomarkers or therapeutic targets. This review aims to give an introduction to network theory and how it can be applied to microarray studies of complex diseases.

Humans↗

Definition of the tempo of sequence diversity across an alignment and automatic identification of sequence motifs: Application to protein homologous families and superfamilies.

It is often possible to identify sequence motifs that characterize a protein family in terms of its fold and/or function from aligned protein sequences. Such motifs can be used to search for new family members. Partitioning of sequence alignments into regions of similar amino acid variability is usually done by hand. Here, I present a completely automatic method for this purpose: one that is guaranteed to produce globally optimal solutions at all levels of partition granularity. The method is used to compare the tempo of sequence diversity across reliable three-dimensional (3D) structure-based alignments of 209 protein families (HOMSTRAD) and that for 69 superfamilies (CAMPASS). (The mean alignment length for HOMSTRAD and CAMPASS are very similar.) Surprisingly, the optimal segmentation distributions for the closely related proteins and distantly related ones are found to be very similar. Also, optimal segmentation identifies an unusual protein superfamily. Finally, protein 3D structure clues from the tempo of sequence diversity across alignments are examined. The method is general, and could be applied to any area of comparative biological sequence and 3D structure analysis where the constraint of the inherent linear organization of the data imposes an ordering on the set of objects to be clustered.

Amino Acid Motifs↗

Discrete profile comparison using information bottleneck.

Sequence homologs are an important source of information about proteins. Amino acid profiles, representing the position-specific mutation probabilities found in profiles, are a richer encoding of biological sequences than the individual sequences themselves. However, profile comparisons are an order of magnitude slower than sequence comparisons, making profiles impractical for large datasets. Also, because they are such a rich representation, profiles are difficult to visualize. To address these problems, we describe a method to map probabilistic profiles to a discrete alphabet while preserving most of the information in the profiles. We find an informationally optimal discretization using the Information Bottleneck approach (IB). We observe that an 80-character IB alphabet captures nearly 90% of the amino acid occurrence information found in profiles, compared to the consensus sequence's 78%. Distant homolog search with IB sequences is 88% as sensitive as with profiles compared to 61% with consensus sequences (AUC scores 0.73, 0.83, and 0.51, respectively), but like simple sequence comparison, is 30 times faster. Discrete IB encoding can therefore expand the range of sequence problems to which profile information can be applied to include batch queries over large databases like SwissProt, which were previously computationally infeasible.

Algorithms↗

Analysis of the hli gene family in marine and freshwater cyanobacteria.

Certain cyanobacteria thrive in natural habitats in which light intensities can reach 2000 micromol photon m(-2) s(-1) and nutrient levels are extremely low. Recently, a family of genes designated hli was demonstrated to be important for survival of cyanobacteria during exposure to high light. In this study we have identified members of the hli gene family in seven cyanobacterial genomes, including those of a marine cyanobacterium adapted to high-light growth in surface waters of the open ocean (Prochlorococcus sp. strain Med4), three marine cyanobacteria adapted to growth in moderate- or low-light (Prochlorococcus sp. strain MIT9313, Prochlorococcus marinus SS120, and Synechococcus WH8102), and three freshwater strains (the unicellular Synechocystis sp. strain PCC6803 and the filamentous species Nostoc punctiforme strain ATCC29133 and Anabaena sp. [Nostoc] strain PCC7120). The high-light-adapted Prochlorococcus Med4 has the smallest genome (1.7 Mb), yet it has more than twice as many hli genes as any of the other six cyanobacterial species, some of which appear to have arisen from recent duplication events. Based on cluster analysis, some groups of hli genes appear to be specific to either marine or freshwater cyanobacteria. This information is discussed with respect to the role of hli genes in the acclimation of cyanobacteria to high light, and the possible relationships among members of this diverse gene family.

Amino Acid Sequence↗

Accelerated bisulfite-deamination of cytosine in the genomic sequencing procedure for DNA methylation analysis.

Understanding the biological consequences of DNA methylation is a current focus of intensive studies. A standard method for analyzing the methylation at position 5 of cytosines in genomic DNA involves chemical modification of the DNA with bisulfite, followed by PCR amplification and sequencing. Bisulfite deaminates cytosine, but it deaminates 5-methylcytosine only very slowly, thereby allowing determination of the methylated sites. The determination is usually performed using sodium bisulfite solutions of 3-5 M concentration with an incubation period of 12-16 hr at 50 degrees C. We demonstrate here that this deamination can be speeded up significantly by increasing the bisulfite concentration and the temperature with which the reaction is performed. In an experiment, in which denatured DNA was treated with 9 M bisulfite for 10 min at pH 5.4 and 90 degrees C, deamination of cytosines occurred to an extent of 99.6%, while 5-methylcytosine residues in the DNA were deaminated at less than 10%. Using a plasmid DNA fragment, we observed that the DNA can serve as a template for PCR amplification after the bisulfite treatment. This new procedure is expected to offer an improved genomic sequencing method, leading to the promotion of research on understanding the biological and medical significance of DNA methylation.

Animals↗

Sequence analysis with the Kestrel SIMD parallel processor.

Computer aided sequence analysis is a critical aspect of current biological research. Sequence information from the genome sequencing projects fills databases so quickly that humans cannot examine it all. Hence there is a heavy reliance on computer algorithms to point out the few important nuggets for human examination. Sequence search algorithms range from simple to complex, as does the representation of the biological data. Typically though, simple algorithms are used on the simplest of data representations because of the large computational demands of anything more complex. This leads to missed hits because the simple search techniques are often not sufficiently sensitive. Here we describe the implementation of several sensitive sequence analysis algorithms on the Kestrel parallel processor, a single-instruction multiple-data (SIMD) processor developed and built at UCSC. Performance of the Smith-Waterman and Hidden Markov Model algorithms, with both Viterbi and Expectation Maximization methods ranges from 6 to 20 times faster than standard computers.

Algorithms↗

An image-processing approach to dotplots: an X-Window-based program for interactive analysis of dotplots derived from sequence and structural data.

We present an approach to the study of the relationships between biological sequences and structures applying image analysis methods to dotplots. We introduce a set of analytical tools based on different types of digital image-processing filters that are new within the context of dotplots. We have reformulated some of the usual approaches in dotplot analysis as mathematical operations on images within the framework of mathematical morphology. An X-Window-based implementation of this new approach has been developed and is available by anonymous FTP.

Data Interpretation, Statistical↗

The BioTools Suite. A comprehensive suite of platform-independent bioinformatics tools.

The BioTools Suite is a set of three comprehensive, platform-independent software packages (PepTool, GeneTool, and ChromaTool) developed for sequence assembly and analysis. In addition to supporting a large number of standard bioinformatics functions, these programs also incorporate a number of useful innovations including uniform graphical-user interface (GUI) design, direct internet connectivity, a novel approach to feature annotation, and a variety of enhanced algorithms for large scale proteome and genome analysis. This article describes the key features, recent changes, and general operation of all three programs.

Computational Biology↗

Computational evidence for hundreds of non-conserved plant microRNAs.

BACKGROUND: MicroRNAs (miRNA) are small (20-25 nt) non-coding RNA molecules that regulate gene expression through interaction with mRNA in plants and metazoans. A few hundred miRNAs are known or predicted, and most of those are evolutionarily conserved. In general plant miRNA are different from their animal counterpart: most plant miRNAs show near perfect complementarity to their targets. Exploiting this complementarity we have developed a method for identification plant miRNAs that does not rely on phylogenetic conservation. RESULTS: Using the presumed targets for the known miRNA as positive controls, we list and filter all segments of the genome of length approximately 20 that are complementary to a target mRNA-transcript. From the positive control we recover 41 (of 92 possible) of the already known miRNA-genes (representing 14 of 16 families) with only four false positives. Applying the procedure to find possible new miRNAs targeting any annotated mRNA, we predict of 592 new miRNA genes, many of which are not conserved in other plant genomes. A subset of our predicted miRNAs is additionally supported by having more than one target that are not homologues. CONCLUSION: These results indicate that it is possible to reliably predict miRNA-genes without using genome comparisons. Furthermore it suggests that the number of plant miRNAs have been underestimated and points to the existence of recently evolved miRNAs in Arabidopsis.

Animals↗

Phylogenetic supermatrix analysis of GenBank sequences from 2228 papilionoid legumes.

A comprehensive phylogeny of papilionoid legumes was inferred from sequences of 2228 taxa in GenBank release 147. A semiautomated analysis pipeline was constructed to download, parse, assemble, align, combine, and build trees from a pool of 11,881 sequences. Initial steps included all-against-all BLAST similarity searches coupled with assembly, using a novel strategy for building length-homogeneous primary sequence clusters. This was followed by a combination of global and local alignment protocols to build larger secondary clusters of locally aligned sequences, thus taking into account the dramatic differences in length of the heterogeneous coding and noncoding sequence data present in GenBank. Next, clusters were checked for the presence of duplicate genes and other potentially misleading sequences and examined for combinability with other clusters on the basis of taxon overlap. Finally, two supermatrices were constructed: a "sparse" matrix based on the primary clusters alone (1794 taxa x 53,977 characters), and a somewhat more "dense" matrix based on the secondary clusters (2228 taxa x 33,168 characters). Both matrices were very sparse, with 95% of their cells containing gaps or question marks. These were subjected to extensive heuristic parsimony analyses using deterministic and stochastic heuristics, including bootstrap analyses. A "reduced consensus" bootstrap analysis was also performed to detect cryptic signal in a subtree of the data set corresponding to a "backbone" phylogeny proposed in previous studies. Overall, the dense supermatrix appeared to provide much more satisfying results, indicated by better resolution of the bootstrap tree, excellent agreement with the backbone papilionoid tree in the reduced bootstrap consensus analysis, few problematic large polytomies in the strict consensus, and less fragmentation of conventionally recognized genera. Nevertheless, at lower taxonomic levels several problems were identified and diagnosed. A large number of methodological issues in supermatrix construction at this scale are discussed, including detection of annotation errors in GenBank sequences; the shortage of effective algorithms and software for local multiple sequence alignment; the difficulty of overcoming effects of fragmentation of data into nearly disjoint blocks in sparse supermatrices; and the lack of informative tools to assess confidence limits in very large trees.

Algorithms↗