PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Molecular community analysis of microbial diversity.

New technologies that avoid the need for either gene amplification (e.g. microarrays) or nucleic acid extraction (e.g. in situ PCR) have recently been implemented in microbial ecology. Together with new approaches for culturing microorganisms and an increased understanding of the biases of molecular methods, these techniques form the most exciting advances in this field during the past year.

Ecosystem↗

Impact of viral infection on the gene expression profiles of proliferating normal human peripheral blood mononuclear cells infected with HIV type 1 RF.

Exploiting the power of high-density gene arrays, the simultaneous expression analysis of 5600 cellular genes was executed on proliferating peripheral blood mononuclear cells (PBMCs) from three normal human donors that were infected in vitro with the T cell tropic laboratory strain of HIV-1, RF. Profiles of expressed genes were assessed at 1, 12, 24, 48, and 72 hr postinfection and compared with those of matched uninfected PBMCs. Viral infection resulted in an overall increase in the number of genes expressed with peaks of expression at 1, 12, and 48 hr postinfection. Functional clustering of genes whose expression level in infected PBMCs varied by 2-fold or greater from levels in the controls indicated that cellular activation markers, proteins associated with immune cell function and with transcription and translation, exhibited increased expression subsequent to viral infection. Gene families exhibiting a decline in gene expression were confined to the 72 hr time point and included genes associated with catabolism and a subset of genes involved with cell signaling and synthetic pathways. Self-organizing map (SOM) cluster analysis identified temporal patterns of coordinated gene expression in infected PBMCs including genes associated with the immune response, the cytoskeleton, and ribosomal subunit structural proteins required for protein synthesis.

Cell Division↗

An overview of toxicogenomics.

Toxicogenomics is a rapidly developing discipline that promises to aid scientists in understanding the molecular and cellular effects of chemicals in biological systems. This field encompasses global assessment of biological effects using technologies such as DNA microarrays or high throughput NMR and protein expression analysis. This review provides an overview of advancing multiple approaches (genomic, proteomic, metabonomic) that may extend our understanding of toxicology and highlights the importance of coupling such approaches with classical toxicity studies.

Animals↗

A hybrid clustering approach to recognition of protein families in 114 microbial genomes.

BACKGROUND: Grouping proteins into sequence-based clusters is a fundamental step in many bioinformatic analyses (e.g., homology-based prediction of structure or function). Standard clustering methods such as single-linkage clustering capture a history of cluster topologies as a function of threshold, but in practice their usefulness is limited because unrelated sequences join clusters before biologically meaningful families are fully constituted, e.g. as the result of matches to so-called promiscuous domains. Use of the Markov Cluster algorithm avoids this non-specificity, but does not preserve topological or threshold information about protein families. RESULTS: We describe a hybrid approach to sequence-based clustering of proteins that combines the advantages of standard and Markov clustering. We have implemented this hybrid approach over a relational database environment, and describe its application to clustering a large subset of PDB, and to 328577 proteins from 114 fully sequenced microbial genomes. To demonstrate utility with difficult problems, we show that hybrid clustering allows us to constitute the paralogous family of ATP synthase F1 rotary motor subunits into a single, biologically interpretable hierarchical grouping that was not accessible using either single-linkage or Markov clustering alone. We describe validation of this method by hybrid clustering of PDB and mapping SCOP families and domains onto the resulting clusters. CONCLUSION: Hybrid (Markov followed by single-linkage) clustering combines the advantages of the Markov Cluster algorithm (avoidance of non-specific clusters resulting from matches to promiscuous domains) and single-linkage clustering (preservation of topological information as a function of threshold). Within the individual Markov clusters, single-linkage clustering is a more-precise instrument, discerning sub-clusters of biological relevance. Our hybrid approach thus provides a computationally efficient approach to the automated recognition of protein families for phylogenomic analysis.

Algorithms↗

Clustering protein sequence and structure space with infinite Gaussian mixture models.

We describe a novel approach to the problem of automatically clustering protein sequences and discovering protein families, subfamilies etc., based on the theory of infinite Gaussian mixtures models. This method allows the data itself to dictate how many mixture components are required to model it, and provides a measure of the probability that two proteins belong to the same cluster. We illustrate our methods with application to three data sets: globin sequences, globin sequences with known three-dimensional structures and G-protein coupled receptor sequences. The consistency of the clusters indicate that our method is producing biologically meaningful results, which provide a very good indication of the underlying families and subfamilies. With the inclusion of secondary structure and residue solvent accessibility information, we obtain a classification of sequences of known structure which both reflects and extends their SCOP classifications. A supplementray web site containing larger versions of the figures is available at http://public.kgi.edu/approximately wid/PSB04/index.html

Amino Acid Sequence↗

An OspA serotyping system for Borrelia burgdorferi based on reactivity with monoclonal antibodies and OspA sequence analysis.

A total of 136 Borrelia burgdorferi sensu latu strains from various biological sources (ticks, human skin, and cerebrospinal fluid) and geographical sources (Europe and North America) were investigated by Western blot (immunoblot) with eight monoclonal antibodies against different epitopes of the outer surface protein A (OspA). On the basis of the differential reactivities of these monoclonal antibodies, seven OspA serotypes were defined. As determined by 16S rRNA sequence analysis, these serotypes correlated well with recently delineated genospecies: serotype 1 corresponds to B. burgdorferi sensu strictu, serotype 2 corresponds to group VS461, and serotypes 3 to 7 correspond to Borrelia garinii sp. nov. (G. Baranton, D. Postic, I. Saint Girons, P. Boerlin, J.-C. Piffaretti, M. Assous, and P. A. D. Grimont, Int. J. Syst. Bacteriol. 42:378-383, 1992). Antigenic differences were confirmed by partial sequence analysis of OspA of representatives of each serotype. Comparative sequence analysis suggested that serotype 5 OspA resulted from genetic recombination of serotype 4 and 6 ospA genes. Serotype 2 (group VS461) was most prevalent among European skin isolates (49 of 62 isolates). Among all B. garinii strains included in this study, serotype 6 was most frequently found in ticks and only rarely in human skin and cerebrospinal fluid, whereas serotypes 4 and 5 were isolated from patients but never from ticks. Our data suggest different pathogenic potentials and organotropisms of distinct OspA serotypes and raise the question of true antigenic variation among B. garinii strains.

Amino Acid Sequence↗

Single molecule techniques for biomedicine and pharmacology.

The present review gives a short summary on techniques useful for single molecule research, describes experiments on in vitro single molecule detection and reactions of single molecules and finally reports on the behavior of single molecules and single virus particles in living cells. One experiment on single molecule enzyme kinetics of lactate dehydrogenase, an enzyme used in the diagnosis of heart attacks and one experiment on restriction analysis of individual DNa molecules are described in some detail. Where it is possible, the relevance to pharmacology and biomedicine is emphasized, often as a perspective or suggestion for experiments, since in this young field of science a not too large variety of experiments have indeed already been devoted directly to drug action.

Animals↗

The microglial gene regulatory network activated by interferon-gamma.

We have analysed the microglial pathway stimulated by interferon-gamma (IFN-gamma) using an in silico approach employing a database of eukaryotic molecular interactions and a microarray dataset validated by quantitative real-time PCR (qRT-PCR). Following IFN-gamma stimulation, production of neuroprotective factors by microglia was found to be reduced while caspase 1 and serping1 which are involved in cell death cascades are up-regulated suggesting a safeguarding mechanism. Extracellular matrix interactions and intracellular protein degradation are altered in concert with these changes. The regulatory network of IFN-gamma responsive microglial genes is outlined in detail and differentially expressed genes are mapped to their respective cellular compartments. A pathway approach to the analysis of microarray data is advocated since overlaying pathway and actual expression data as shown here greatly facilitates understanding the biological meaning of a gene regulatory network. In addition, genes of similar function that are differentially regulated are less likely to be false positives than single unrelated genes.

Animals↗

Automated prediction of domain boundaries in CASP6 targets using Ginzu and RosettaDOM.

Domain boundary prediction is an important step in both experimental and computational protein structure characterization. We have developed two fully automated domain parsing methods: the first, Ginzu, which we have described previously, utilizes information from homologous sequences and structures, while the second, RosettaDOM, which has not been described previously, uses only information in the query sequence. Ginzu iteratively assigns domains by homology to structures and sequence families using successively less confident methods. RosettaDOM uses the Rosetta de novo structure prediction method to build three-dimensional models, and then applies Taylor's structure based domain assignment method to parse the models into domains. Domain boundaries observed repeatedly in the models are predicted to be domain boundaries for the protein. Interestingly, RosettaDOM produced quite good domain predictions for proteins of a size typically considered to be beyond the reach of de novo structure prediction methods. For remote fold recognition targets and new folds, both Ginzu and RosettaDOM produced promising results, and in some cases where one method failed to detect the correct domain boundary, it was correctly identified by the other method. We describe here the successes and failures using both methods, and address the possibility of incorporating both protocols into an improved hybrid method.

Algorithms↗

Characterization of terminal-repeat retrotransposon in miniature (TRIM) in Brassica relatives.

We have newly identified five Terminal-repeat retrotransposon in miniature (TRIM) families, four from Brassica and one from Arabidopsis. A total of 146 elements, including three Arabidopsis families reported before, are extracted from genomics data of Brassica and Arabidopsis, and these are grouped into eight distinct lineages, Br1 to Br4 derived from Brassica and At1 to At4 derived from Arabidopsis. Based on the occurrence of TRIM elements in 434 Mb of B. oleracea shotgun sequences and 96 Mb of B. rapa BAC end sequences, total number of TRIM members of Br1, Br2, Br3, and Br4 families are roughly estimated to be present in 660 and 530 copies in B. oleracea and B. rapa genomes, respectively. Studies on insertion site polymorphisms of four elements across taxa in the tribe Brassiceae infer the taxonomic lineage and dating of the insertion time. Active roles of the TRIM elements for evolution of the duplicated genes are inferred in the highly replicated Brassica genome.

Base Sequence↗

Direct and indirect transcriptional targets of DAF-16.

Several genes involved in the determination of life span have been identified by mutation in the free-living soil nematode Caenorhabditis elegans. One of the key pathways studied in the context of life span is the DAF-2 pathway. The daf-2 gene is homologous to the insulin and insulin-like growth factor 1 receptor families. A downstream gene, daf-16, encodes a protein that is homologous to the forkhead transcription factor. A study by McElwee, Bubb, and Thomas, published in the current issue of Aging Cell, used genome-scale gene expression analysis to search for genes that are differentially expressed between long-lived daf-2(e1370) and short-lived daf-16(m27);daf-2(e1370) animals. In doing so, they identified candidate direct and indirect targets of DAF-16. In this Perspective, I discuss the results of this study.

Aging↗

Computational prediction of microRNA genes in silkworm genome.

MicroRNAs (miRNAs) constitute a novel, extensive class of small RNAs (approximately 21 nucleotides), and play important gene-regulation roles during growth and development in various organisms. Here we conducted a homology search to identify homologs of previously validated miRNAs from silkworm genome. We identified 24 potential miRNA genes, and gave each of them a name according to the common criteria. Interestingly, we found that a great number of newly identified miRNAs were conserved in silkworm and Drosophila, and family alignment revealed that miRNA families might possess single nucleotide polymorphisms. miRNA gene clusters and possible functions of complement miRNA pairs are discussed.

Animals↗

ADGO: analysis of differentially expressed gene sets using composite GO annotation.

MOTIVATION: Genes are typically expressed in modular manners in biological processes. Recent studies reflect such features in analyzing gene expression patterns by directly scoring gene sets. Gene annotations have been used to define the gene sets, which have served to reveal specific biological themes from expression data. However, current annotations have limited analytical power, because they are classified by single categories providing only unary information for the gene sets. RESULTS: Here we propose a method for discovering composite biological themes from expression data. We intersected two annotated gene sets from different categories of Gene Ontology (GO). We then scored the expression changes of all the single and intersected sets. In this way, we were able to uncover, for example, a gene set with the molecular function F and the cellular component C that showed significant expression change, while the changes in individual gene sets were not significant. We provided an exemplary analysis for HIV-1 immune response. In addition, we tested the method on 20 public datasets where we found many 'filtered' composite terms the number of which reached approximately 34% (a strong criterion, 5% significance) of the number of significant unary terms on average. By using composite annotation, we can derive new and improved information about disease and biological processes from expression data. AVAILABILITY: We provide a web application (ADGO: http://array.kobic.re.kr/ADGO) for the analysis of differentially expressed gene sets with composite GO annotations. The user can analyze Affymetrix and dual channel array (spotted cDNA and spotted oligo microarray) data for four species: human, mouse, rat and yeast. CONTACT: chu@kribb.re.kr SUPPLEMENTARY INFORMATION: http://array.kobic.re.kr/ADGO.

Algorithms↗

Analysis and functional classification of transcripts from the nematode Meloidogyne incognita.

BACKGROUND: Plant parasitic nematodes are major pathogens of most crops. Molecular characterization of these species as well as the development of new techniques for control can benefit from genomic approaches. As an entrée to characterizing plant parasitic nematode genomes, we analyzed 5,700 expressed sequence tags (ESTs) from second-stage larvae (L2) of the root-knot nematode Meloidogyne incognita. RESULTS: From these, 1,625 EST clusters were formed and classified by function using the Gene Ontology (GO) hierarchy and the Kyoto KEGG database. L2 larvae, which represent the infective stage of the life cycle before plant invasion, express a diverse array of ligand-binding proteins and abundant cytoskeletal proteins. L2 are structurally similar to Caenorhabditis elegans dauer larva and the presence of transcripts encoding glyoxylate pathway enzymes in the M. incognita clusters suggests that root-knot nematode larvae metabolize lipid stores while in search of a host. Homology to other species was observed in 79% of translated cluster sequences, with the C. elegans genome providing more information than any other source. In addition to identifying putative nematode-specific and Tylenchida-specific genes, sequencing revealed previously uncharacterized horizontal gene transfer candidates in Meloidogyne with high identity to rhizobacterial genes including homologs of nodL acetyltransferase and novel cellulases. CONCLUSIONS: With sequencing from plant parasitic nematodes accelerating, the approaches to transcript characterization described here can be applied to more extensive datasets and also provide a foundation for more complex genome analyses.

Animals↗

Alfresco--a workbench for comparative genomic sequence analysis.

Comparative analysis of genomic sequences provides a powerful tool for identifying regions of potential biologic function; by comparing corresponding regions of genomes from suitable species, protein coding or regulatory regions can be identified by their homology. This requires the use of several specific types of computational analysis tools. Many programs exist for these types of analysis; not many exist for overall view/control of the results, which is necessary for large-scale genomic sequence analysis. Using Java, we have developed a new visualization tool that allows effective comparative genome sequence analysis. The program handles a pair of sequences from putatively homologous regions in different species. Results from various different existing external analysis programs, such as database searching, gene prediction, repeat masking, and alignment programs, are visualized and used to find corresponding functional sequence domains in the two sequences. The user interacts with the program through a graphic display of the genome regions, in which an independently scrollable and zoomable symbolic representation of the sequences is shown. As an example, the analysis of two unannotated orthologous genomic sequences from human and mouse containing parts of the UTY locus is presented.

Algorithms↗

Dynamic covariation between gene expression and proteome characteristics.

BACKGROUND: Cells react to changing intra- and extracellular signals by dynamically modulating complex biochemical networks. Cellular responses to extracellular signals lead to changes in gene and protein expression. Since the majority of genes encode proteins, we investigated possible correlations between protein parameters and gene expression patterns to identify proteome-wide characteristics indicative of trends common to expressed proteins. RESULTS: Numerous bioinformatics methods were used to filter and merge information regarding gene and protein annotations. A new statistical time point-oriented analysis was developed for the study of dynamic correlations in large time series data. The method was applied to investigate microarray datasets for different cell types, organisms and processes, including human B and T cell stimulation, Drosophila melanogaster life span, and Saccharomyces cerevisiae cell cycle. CONCLUSION: We show that the properties of proteins synthesized correlate dynamically with the gene expression profile, indicating that not only is the actual identity and function of expressed proteins important for cellular responses but that several physicochemical and other protein properties correlate with gene expression as well. Gene expression correlates strongly with amino acid composition, composition- and sequence-derived variables, functional, structural, localization and gene ontology parameters. Thus, our results suggest that a dynamic relationship exists between proteome properties and gene expression in many biological systems, and therefore this relationship is fundamental to understanding cellular mechanisms in health and disease.

Animals↗

BADASP: predicting functional specificity in protein families using ancestral sequences.

SUMMARY: Burst After Duplication with Ancestral Sequence Predictions (BADASP) is a software package for identifying sites that may confer subfamily-specific biological functions in protein families following functional divergence of duplicated proteins. A given protein phylogeny is grouped into subfamilies based on orthology/paralogy relationships and/or user definitions. Ancestral sequences are then predicted from the sequence alignment and the functional specificity is calculated using variants of the Burst After Duplication method, which tests for radical amino acid substitutions following gene duplications that are subsequently conserved. Statistics are output along with subfamily groupings and ancestral sequences for an easy analysis with other packages. AVAILABILITY: BADASP is freely available from http://www.bioinformatics.rcsi.ie/~redwards/badasp/

Algorithms↗

Genome-wide location analysis: insights on transcriptional regulation.

Gene expression analysis of microarray data can provide a global view of the transcriptome of a cell or specific tissue type, revealing important information about the kinds of signaling pathways, genes and protein classifications that are active. However, transcript profiles alone do not reveal how expression levels are controlled or which transcription factors (TFs) are responsible. Establishing transcriptional regulatory networks requires knowledge of TFs bound to promoter, enhancer and repressor elements. Accessibility of these sites and an additional level of control are mediated by chromatin and DNA modifications. Genome-wide location analysis is a tool for identifying protein-DNA interaction sites on a genomic scale. Applications of this tool are proving invaluable in determining in vivo target genes of TFs, epigenetic marks and cis-regulatory elements. Here, we will discuss how advances have been made in each of these categories and how this has helped to elucidate regulatory networks and control mechanisms.

Animals↗