PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Study of coordinative gene expression at the biological process level.

MOTIVATION: Cellular processes are not isolated groups of events. Nevertheless, in most microarray analyses, they tend to be treated as standalone units. To shed light on how various parts of the interlocked biological processes are coordinated at the transcription level, there is a need to study the between-unit expressional relationship directly. RESULTS: We approach this issue by constructing an index of correlation function to convey the global pattern of coexpression between genes from one process and genes from the entire genome. Processes with similar signatures are then identified and projected to a process-to-process association graph. This top-down method allows for detailed gene-level analysis between linked processes to follow up. Using the cell-cycle gene-expression profiles for Saccharomyces cerevisiae, we report well-organized networks of biological processes that would be difficult to find otherwise. Using another dataset, we report a sharply different network structure featuring cellular responses under environmental stress. SUPPLEMENTARY INFORMATION: http://kiefer.stat.ucla.edu/lap2/download/KL_supplement.pdf.

Algorithms↗

Data clustering in life sciences.

Clustering has a wide range of applications in life sciences and over the years has been used in many areas ranging from the analysis of clinical information, phylogeny, genomics, and proteomics. The primary goal of this article is to provide an overview of the various issues involved in clustering large biological datasets, describe the merits and underlying assumptions of some of the commonly used clustering approaches, and provide insights on how to cluster datasets arising in various areas within life sciences. We also provide a brief introduction to CLUTO, a general purpose toolkit for clustering various datasets, with an emphasis on its applications to problems and analysis requirements within life sciences.

Algorithms↗

Loss of expression of ZAC/LOT1 in squamous cell carcinomas of head and neck.

BACKGROUND: ZAC/Lot1 is a previously identified candidate tumor suppressor gene. The gene maps to the human chromosome 6q24-q25, a region frequently deleted in squamous cell carcinomas of the head and neck and other solid tumors. METHODS: We have used a model of head and neck squamous cell carcinoma (HNSCC) and cell lines to analyze the role of the candidate tumor suppressor gene ZAC/Lot1 in oral carcinogenesis. We analyzed the expression in 11 cell lines, and we performed loss of heterozygosity (LOH)- and sequence analyses in 51 primary tumors. RESULTS: Three (27.3%) of 11 cell lines showed a distinctly reduced expression of ZAC/Lot1 compared with expression levels of the gene in the normal oral mucosa. In addition, we analyzed 51 primary squamous cell carcinomas of the head and neck for LOH with seven microsatellite markers flanking ZAC/Lot1. We detected an average LOH rate of 31.4% in the region of interest. Sequence analysis revealed no mutations for the ZAC/Lot1 coding exons, including the exon/intron boundaries. CONCLUSIONS: These data could suggest a minimal role for ZAC/Lot1 in a subgroup of HNSCC tumors.

Adult↗

Nutrient-gene interactions: a single nutrient and hundreds of target genes.

Based on the effects of a selective experimental zinc deficiency in a rodent model we explore the use of transcriptome profiling for assessing nutrient-gene interactions in the liver at the molecular and cellular levels. Zinc deficiency caused pleiotropic alterations in mRNA/protein levels of hundreds of genes. In the context of observed metabolic alterations in hepatic metabolism, possible mechanisms are discussed for how a low zinc status may be sensed and transmitted into changes in various metabolic pathways. However, it also becomes obvious that analysis of such complex nutrient-gene interactions beyond the descriptional level is a real challenge for systems biology.

Animals↗

Identification of gene expression patterns using planned linear contrasts.

BACKGROUND: In gene networks, the timing of significant changes in the expression level of each gene may be the most critical information in time course expression profiles. With the same timing of the initial change, genes which share similar patterns of expression for any number of sampling intervals from the beginning should be considered co-expressed at certain level(s) in the gene networks. In addition, multiple testing problems are complicated in experiments with multi-level treatments when thousands of genes are involved. RESULTS: To address these issues, we first performed an ANOVA F test to identify significantly regulated genes. The Benjamini and Hochberg (BH) procedure of controlling false discovery rate (FDR) at 5% was applied to the P values of the F test. We then categorized the genes with a significant F test into 4 classes based on the timing of their initial responses by sequentially testing a complete set of orthogonal contrasts, the reverse Helmert series. For genes within each class, specific sequences of contrasts were performed to characterize their general 'fluctuation' shapes of expression along the subsequent sampling time points. To be consistent with the BH procedure, each contrast was examined using a stepwise Studentized Maximum Modulus test to control the gene based maximum family-wise error rate (MFWER) at the level of alphanew determined by the BH procedure. We demonstrated our method on the analysis of microarray data from murine olfactory sensory epithelia at five different time points after target ablation. CONCLUSION: In this manuscript, we used planned linear contrasts to analyze time-course microarray experiments. This analysis allowed us to characterize gene expression patterns based on the temporal order in the data, the timing of a gene's initial response, and the general shapes of gene expression patterns along the subsequent sampling time points. Our method is particularly suitable for analysis of microarray experiments in which it is often difficult to take sufficiently frequent measurements and/or the sampling intervals are non-uniform.

Algorithms↗

Identification of candidate maternal-effect genes through comparison of multiple microarray data sets.

Transcriptional profiling by microarray hybridization has become a standard method to analyze global gene expression and has resulted in the availability of enormous amounts of experimental data. Given the number of different microarray platforms currently in use, it is critical to determine how reproducible results are from one platform to another. Additional variability may also arise from tissue collection and protocol differences among laboratories. In an effort to identify genes whose maternal mRNA pools are critical during preimplantation development, we compared published results of three independent studies of the mouse preimplantation embryo transcriptome, each performed in a different laboratory using different microarray platforms. We searched the combined data set for genes whose expression patterns were consistent among the three experiments. Querying for presence or absence at single developmental windows indicates that between 52% and 60% of genes are in agreement among the three experiments. Searching for expression patterns across three developmental windows (oocyte + 1-cell, 2- through 8-cell, and blastocyst stage) revealed approximately 33% agreement among the three experiments, although the majority of these genes were either always present or always absent. Using this approach, we identified 51 genes with a predicted expression pattern of maternal RNA only (not present during 2-cell through 8-cell or at the blastocyst stage). RT-PCR validation indicates 37 (72%) of these candidates have the microarray-predicted expression pattern and represent candidate maternal-effect genes. Based on our analysis, we conclude that data mining microarray experiments in this way greatly enhances candidate gene expression pattern accuracy.

Animals↗

A cDNA microarray approach to decipher sunflower (Helianthus annuus) responses to the necrotrophic fungus Phoma macdonaldii.

To identify the genes involved in the partial resistance of sunflower (Helianthus annuus) to the necrotrophic fungus Phoma macdonaldii, we developed a 1000-element cDNA microarray containing carefully chosen genes putatively involved in primary metabolic pathways, signal transduction and biotic stress responses. A two-pass general linear model was used to normalize the data and then to detect differentially expressed genes. This method allowed us to identify 38 genes differentially expressed among genotypes, treatments and times, mainly belonging to plant defense, signaling pathways and amino acid metabolism. Based on a set of genes whose differential expression was highly significant, we propose a model in which negative regulation of a dual-specificity MAPK phosphatase could be implicated in sunflower defense mechanisms against the pathogen. The resulting activation of the MAP kinase cascade could subsequently trigger defense responses (e.g. thaumatin biosynthesis and phenylalanine ammonia lyase activation), under the control of transcription factors belonging to MYB and WRKY families. Concurrently, the activation of protein phosphatase 2A (PP2A), which is implicated in cell death inhibition, could limit pathogen development. The results reported here provide a valuable first step towards the understanding and analysis of the P. macdonaldii-sunflower interaction.

Ascomycota↗

Analysis of the wheat endosperm transcriptome.

Among the cereals, wheat is the most widely grown geographically and is part of the staple diet in much of the world. Understanding how the cereal endosperm develops and functions will help generate better tools to manipulate grain qualities important to end-users. We used a genomics approach to identify and characterize genes that are expressed in the wheat endosperm. We analyzed the 17,949 publicly available wheat endosperm EST sequences to identify genes involved in the biological processes that occur within this tissue. Clustering and assembly of the ESTs resulted in the identification of 6,187 tentative unique genes, 2,358 of which formed contigs and 3,829 remained as singletons. A BLAST similarity search against the NCBI non-redundant sequence database revealed abundant messages for storage proteins, putative defense proteins, and proteins involved in starch and sucrose metabolism. The level of abundance of the putatively identified genes reflects the physiology of the developing endosperm. Half of the identified genes have unknown functions. Approximately 61% of the endosperm ESTs has been tentatively mapped in the hexaploid wheat genome. Using microarrays for global RNA profiling, we identified endosperm genes that are specifically up regulated in the developing grain.

Expressed Sequence Tags↗

Pathway Miner: extracting gene association networks from molecular pathways for predicting the biological significance of gene expression microarray data.

UNLABELLED: We have developed a web-based system (Pathway Miner) for visualizing gene expression profiles in the context of biological pathways. Pathway Miner catalogs genes based on their role in metabolic, cellular and regulatory pathways. A Fisher exact test is provided as an option to rank pathways. The genes are mapped onto pathways and gene product association networks are extracted for genes that co-occur in pathways. The networks can be filtered for analysis based on user-selected options. AVAILABILITY: Pathway Miner is a freely available web accessible tool at http://www.biorag.org/pathway.html

Algorithms↗

Theoretical and computational studies of the glucose signaling pathways in yeast using global gene expression data.

We have combined DNA microarray experiments with novel computational methods as a means of defining the topology of a biological signal transduction pathway. By DNA microarray techniques, we previously acquired data on expression over time of all genes in the yeast Saccharomyces following addition of glucose to wild-type cells and to cells mutated in one or more components of the Ras signaling network. In addition, we examined the time course of expression following activation of components of the Ras signaling network in the absence of glucose addition. In this current study, we have applied a novel theoretical and computational framework to these data to identify the network topology of the glucose signaling pathway in yeast and the role of Ras components in that network. The computational approach involves clustering genes by expression pattern, postulating a signaling network topology superstructure that includes all possible component interconnections and then evaluating the feasibility of the superstructure interconnections by optimization methods using Mixed Integer Linear Programming techniques. This approach is the first rigorous mathematical framework for addressing the biological network topology issue, and the novel formulation features the introduction of discrete variables for the connectivity and logical expressions that connect the experimental observations to the network structure. This analysis yields a topology for the glucose signaling pathway that is consistent with, and an extension of, known biological interactions in glucose signaling.

Computer Simulation↗

Operomics: molecular analysis of tissues from DNA to RNA to protein.

The identification of coding sequences in a number of species, including human in the near future, has ushered in the post-genome era. In this era, technologies are becoming available that allow the profiling of tissues and cell populations at the genomic, transcriptomic and proteomic levels. The molecular analysis of tissues at all three levels has been referred to as operomics. This review covers some basic technologies for operomics and their application to some lymphoid disorders. It is proposed that no one type of analysis is fully informative and that information that can be derived from the different compartments encompassed in operomics is complementary. Prospects for introducing such profiling technologies into the clinical laboratory will depend on their robustness, their user friendliness and the clinical relevance of the added information they provide, which cannot be captured through other technologies in use in the clinical laboratory.

Genomics↗

Fundamentals of experimental design for cDNA microarrays.

Microarray technology is now widely available and is being applied to address increasingly complex scientific questions. Consequently, there is a greater demand for statistical assessment of the conclusions drawn from microarray experiments. This review discusses fundamental issues of how to design an experiment to ensure that the resulting data are amenable to statistical analysis. The discussion focuses on two-color spotted cDNA microarrays, but many of the same issues apply to single-color gene-expression assays as well.

Animals↗

Phylogenetic tree-building.

Cladistic analysis is an approach to phylogeny reconstruction that groups taxa in such a way that those with historically more-recent ancestors form groups nested within groups of taxa with more-distant ancestors. This nested set of taxa can be represented as a branching diagram or tree (a cladogram), which is an hypothesis of the evolutionary history of the taxa. The analysis is performed by searching for nested groups of shared derived character states. These shared derived character states define monophyletic groups of taxa (clades), which include all of the descendants of the most recent common ancestor. If all of the characters for a set of taxa are congruent, then reconstructing the phylogenetic tree is unproblematic. However, most real data sets contain incongruent characters, and consequently a wide range of tree-building methods has been developed. These methods differ in a variety of characteristics, and they may produce topologically distinct trees for a single data set. None of the currently-available methods are simultaneously efficient, powerful, consistent and robust, and thus there is no single ideal method. However, many of them appear to perform well under a wide range of conditions, with the exception of the UPGMA method and the Invariants method.

Algorithms↗

Molecular Biocomputing Suite: a word processor add-in for the analysis and manipulation of nucleic acid and protein sequence data.

In all fields of molecular biology, researchers are increasingly challenged by experiments planned and evaluated on the basis of nucleic acid and protein sequence data generally retrieved from public databases. Despite the wide spectrum of available Web-based software tools for sequence analysis, the routine use of these tools has disadvantages, particularly because of the elaborate and heterogeneous ways of data input, output, and storage. Here we present a Visual Basic-encoded Microsoft Word Add-In, the Molecular BioComputing Suite (MBCS), available at the BioTechniques Software Library (www.BioTechniques.com). The MBCS software aims to manage and expedite a wide range of sequence analyses and manipulations using an integrated text editor environment including menu-guided commands. Its independence of sequence formats enables MBCS to be used as a pivotal application between other software tools for sequence analysis, manipulation, annotation, and editing.

Amino Acid Sequence↗

Sequence analysis of the nucleocapsid protein gene of human coronavirus 229E.

Human coronaviruses are important human pathogens and have also been implicated in multiple sclerosis. To further understand the molecular biology of human coronavirus 229E (HCV-229E), molecular cloning and sequence analysis of the viral RNA have been initiated. Following established protocols, the 3'-terminal 1732 nucleotides of the genome were sequenced. A large open reading frame encodes a 389 amino acid protein of 43,366 Da, which is presumably the nucleocapsid protein. The predicted protein is similar in size, chemical properties, and amino acid sequence to the nucleocapsid proteins of other coronaviruses. This is especially evident when the sequence is compared with that of the antigenically related porcine transmissible gastroenteritis virus (TGEV), with which a region of 46% amino acid sequence homology was found. Hydropathy profiles revealed the existence of several conserved domains which could have functional significance. An intergenic consensus sequence precedes the 5'-end of the proposed nucleocapsid protein gene. The consensus sequence is present in other coronaviruses and has been proposed as the site of binding of the leader sequence for mRNA transcriptional start. This region was also examined by primer extension analysis of mRNAs, which identified a 60-nucleotide leader sequence. The 3'-noncoding region of the genome contains an 11-nucleotide sequence, which is relatively conserved throughout the Coronavirus family and lends support to the theory that this region is important for the replication of negative-strand RNA.

Amino Acid Sequence↗

Pathways to the analysis of microarray data.

The development of microarray technology allows the simultaneous measurement of the expression of many thousands of genes. The information gained offers an unprecedented opportunity to fully characterize biological processes. However, this challenge will only be successful if new tools for the efficient integration and interpretation of large datasets are available. One of these tools, pathway analysis, involves looking for consistent but subtle changes in gene expression by incorporating either pathway or functional annotations. We review several methods of pathway analysis and compare the performance of three, the binomial distribution, z scores, and gene set enrichment analysis, on two microarray datasets. Pathway analysis is a promising tool to identify the mechanisms that underlie diseases, adaptive physiological compensatory responses and new avenues for investigation.

Algorithms↗