PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Global gene expression analyses of hematopoietic stem cell-like cell lines with inducible Lhx2 expression.

BACKGROUND: Expression of the LIM-homeobox gene Lhx2 in murine hematopoietic cells allows for the generation of hematopoietic stem cell (HSC)-like cell lines. To address the molecular basis of Lhx2 function, we generated HSC-like cell lines where Lhx2 expression is regulated by a tet-on system and hence dependent on the presence of doxycyclin (dox). These cell lines efficiently down-regulate Lhx2 expression upon dox withdrawal leading to a rapid differentiation into various myeloid cell types. RESULTS: Global gene expression of these cell lines cultured in dox was compared to different time points after dox withdrawal using microarray technology. We identified 267 differentially expressed genes. The majority of the genes overlapping with HSC-specific databases were those down-regulated after turning off Lhx2 expression and a majority of the genes overlapping with those defined as late progenitor-specific genes were the up-regulated genes, suggesting that these cell lines represent a relevant model system for normal HSCs also at the level of global gene expression. Moreover, in situ hybridisations of several genes down-regulated after dox withdrawal showed overlapping expression patterns with Lhx2 in various tissues during embryonic development. CONCLUSION: Global gene expression analysis of HSC-like cell lines with inducible Lhx2 expression has identified genes putatively linked to self-renewal/differentiation of HSCs, and function of Lhx2 in organ development and stem/progenitor cells of non-hematopoietic origin.

Animals↗

Local correlation of expression profiles with gene annotations--proof of concept for a general conciliatory method.

MOTIVATION: Integrated analysis of expression data and gene ontology annotations is a prime example of biological data that need co-explanatory interpretation. This particular application is used to validate a new method for integrated analysis of varied biological information. RESULTS: The proposed method consists of determining local correlation coefficients and the corresponding P-values calculated per biological entity. This measure considers the combined intensity and significance of the agreement or disagreement, between two data sources about the same biological entity. The method is applied to the integrated analysis of gene expression and annotation of two gene sets, one from yeast and other from mouse. The potential of the method to generate accurate mechanistic hypothesis is also demonstrated. Specially, negative correlation results pose a new kind of biological hypothesis. Method performance was compared with annotation enrichment methods, and optimal conditions for the superiority of local correlation results are discussed.

Algorithms↗

Reverse taxonomy: an approach towards determining the diversity of meiobenthic organisms based on ribosomal RNA signature sequences.

Organisms living in or on the sediment layer of water bodies constitute the benthos fauna, which is known to harbour a large number of species of diverse taxonomic groups. The benthos plays a significant role in the nutrient cycle and it is, therefore, of high ecological relevance. Here, we have explored a DNA-taxonomic approach to access the meiobenthic organismic diversity, by focusing on obtaining signature sequences from a part of the large ribosomal subunit rRNA (28S), the D3-D5 region. To obtain a broad representation of taxa, benthos samples were taken from 12 lakes in Germany, representing different ecological conditions. In a first approach, we have extracted whole DNA from these samples, amplified the respective fragment by PCR, cloned the fragments and sequenced individual clones. However, we found a relatively large number of recombinant clones that must be considered PCR artefacts. In a second approach we have, therefore, directly sequenced PCR fragments that were obtained from DNA extracts of randomly picked individual organisms. In total, we have obtained 264 new unique sequences, which can be readily placed into taxon groups, based on phylogenetic comparison with currently available database sequences. The group with the highest taxon abundance were nematodes and protozoa, followed by chironomids. However, we find also that we have by far not exhausted the diversity of organisms in the samples. Still, our data provide a framework within which a meiobenthos DNA signature sequence database can be constructed, that will allow to develop the necessary techniques for studying taxon diversity in the context of ecological analysis. Since many taxa in our analysis are initially only identified via their signature sequences, but not yet their morphology, we propose to call this approach 'reverse taxonomy'.

Animals↗

Human endogenous retroviral element k10 (HERV-K10): chromosomal localization by somatic hybrid mapping and fluorescence in situ hybridization.

The human endogenous retrovirus K10 (HERV-K10) was mapped to human chromosomes using HERV-K10 specific PCR primers on a somatic hybrid mapping panel. A non-random chromosomal location was demonstrated with PCR signals on chromosomes 1, 3, 4, 5, 6, 7, 10, 11, 12, 14, 15, 19, 20, 21, 22 and Y. There was a lack of PCR products on the other chromosomes, even after hybridization with a HERV-K10 specific probe. To further localize the HERV-K10 sequence we used fluorescence in situ hybridization. Chromosomes 1, 3, 6, 7, 10, 11, 12 and 22 were found to contain several HERV-K10 sequences in different regions. The presence of several integration sites on some chromosomes is consistent with previous studies demonstrating 30-50 copies of the HERV-K10 sequence per haploid genome. The mapping information reported in this study will assist the analysis of the biological significance of the HERV-K10 sequence.

Base Sequence↗

Stage-specific alterations of the genome, transcriptome, and proteome during colorectal carcinogenesis.

To identify sequential alterations of the genome, transcriptome, and proteome during colorectal cancer progression, we have analyzed tissue samples from 36 patients, including the complete mucosa-adenoma-carcinoma sequence from 8 patients. Comparative genomic hybridization (CGH) revealed patterns of stage specific, recurrent genomic imbalances. Gene expression analysis on 9K cDNA arrays identified 58 genes differentially expressed between normal mucosa and adenoma, 116 genes between adenoma and carcinoma, and 158 genes between primary carcinoma and liver metastasis (P < 0.001). Parallel analysis of our samples by CGH and expression profiling revealed a direct correlation of chromosomal copy number changes with chromosome-specific average gene expression levels. Protein expression was analyzed by two-dimensional gel electrophoresis and subsequent mass spectrometry. Although there was no direct match of differentially expressed proteins and genes, the majority of them belonged to identical pathways or networks. In conclusion, increasing genomic instability and a recurrent pattern of chromosomal imbalances as well as specific gene and protein expression changes correlate with distinct stages of colorectal cancer progression. Chromosomal aneuploidies directly affect average resident gene expression levels, thereby contributing to a massive deregulation of the cellular transcriptome. The identification of novel genes and proteins might deliver molecular targets for diagnostic and therapeutic interventions.

Adult↗

Techniques for sample preparation including methods for concentrating peptide samples.

In the current era of proteomics two main analytical techniques are employed for protein identification. By far the fastest and most sensitive procedure for protein identification employs biological mass spectrometry, while de novo sequence analysis by classical Edman degradation is currently diminishing. In order to achieve the highest sensitivity for both techniques, great demands need to be put on sample preparation. In this paper we review three different aspects of protein sample preparation. Firstly, we discuss the use of polyacrylamide or agarose gel systems in which, during electrophoresis, proteins present in multiple primary gel pieces are eluted and simultaneously concentrated in a small secondary gel volume, whereby the overall sensitivity of Edman sequencing can be greatly increased. In a second chapter we review automation strategies occurring in the protein field which allow the automatic handling of multiple protein spots at the same time. In this context, we describe the use of auto-sampling techniques for further mass spectrometric studies and protein digestion robots allowing the simultaneous preparation of tens of gel-separated proteins. Finally we discuss various strategies for the preparation of biological peptide samples such as protein digests for both matrix-assisted laser desorption ionisation and electrospray ionisation mass spectrometry.

Automation↗

Metabolic and transcriptional patterns accompanying glutamine depletion and repletion in mouse hepatoma cells: a model for physiological regulatory networks.

An important objective in postgenomic biology is to link gene expression to function by developing physiological networks that include data from the genomic and functional levels. Here, we develop a model for the analysis of time-dependent changes in metabolites, fluxes, and gene expression in a hepatic model system. The experimental framework chosen was modulation of extracellular glutamine in confluent cultures of mouse Hepa1-6 cells. The importance of glutamine has been demonstrated previously in mammalian cell culture by precipitating metabolic shifts with glutamine depletion and repletion. Our protocol removed glutamine from the medium for 24 h and returned it for a second 24 h. Flux assays of glycolysis, the tricarboxylic acid (TCA) cycle, and lipogenesis were used at specified intervals. All of these fluxes declined in the absence of glutamine and were restored when glutamine was repleted. Isotopomer spectral analysis identified glucose and glutamine as equal sources of lipogenic carbon. Metabolite measurements of organic acids and amino acids indicated that most metabolites changed in parallel with the fluxes. Experiments with actinomycin D indicated that de novo mRNA synthesis was required for observed flux changes during the depletion/repletion of glutamine. Analysis of gene expression data from DNA microarrays revealed that many more genes were anticorrelated with the glycolytic flux and glutamine level than were correlated with these indicators. In conclusion, this model may be useful as a prototype physiological regulatory network where gene expression profiles are analyzed in concert with changes in cell function.

Animals↗

Gene co-expression network topology provides a framework for molecular characterization of cellular state.

MOTIVATION: Gene expression data have become an instrumental resource in describing the molecular state associated with various cellular phenotypes and responses to environmental perturbations. The utility of expression profiling has been demonstrated in partitioning clinical states, predicting the class of unknown samples and in assigning putative functional roles to previously uncharacterized genes based on profile similarity. However, gene expression profiling has had only limited success in identifying therapeutic targets. This is partly due to the fact that current methods based on fold-change focus only on single genes in isolation, and thus cannot convey causal information. In this paper, we present a technique for analysis of expression data in a graph-theoretic framework that relies on associations between genes. We describe the global organization of these networks and biological correlates of their structure. We go on to present a novel technique for the molecular characterization of disparate cellular states that adds a new dimension to the fold-based methods and conclude with an example application to a human medulloblastoma dataset. RESULTS: We have shown that expression networks generated from large model-organism expression datasets are scale-free and that the average clustering coefficient of these networks is several orders of magnitude higher than would be expected for similarly sized scale-free networks, suggesting an inherent hierarchical modularity similar to that previously identified in other biological networks. Furthermore, we have shown that these properties are robust with respect to the parameters of network construction. We have demonstrated an enrichment of genes having lethal knockout phenotypes in the high-degree (i.e. hub) nodes in networks generated from aggregate condition datasets; using process-focused Saccharomyces cerivisiae datasets we have demonstrated additional high-degree enrichments of condition-specific genes encoding proteins known to be involved in or important for the processes interrogated by the microarrays. These results demonstrate the utility of network analysis applied to expression data in identifying genes that are regulated in a state-specific manner. We concluded by showing that a sample application to a human clinical dataset prominently identified a known therapeutic target. AVAILABILITY: Software implementing the methods for network generation presented in this paper is available for academic use by request from the authors in the form of compiled linux binary executables.

Algorithms↗

CSB.DB: a comprehensive systems-biology database.

SUMMARY: The open access comprehensive systems-biology database (CSB.DB) presents the results of bio-statistical analyses on gene expression data in association with additional biochemical and physiological knowledge. The main aim of this database platform is to provide tools that support insight into life's complexity pyramid with a special focus on the integration of data from transcript and metabolite profiling experiments. The central part of CSB.DB, which we describe in this applications note, is a set of co-response databases that currently focus on the three key model organisms, Escherichia coli, Saccharomyces cerevisiae and Arabidopsis thaliana. CSB.DB gives easy access to the results of large-scale co-response analyses, which are currently based exclusively on the publicly available compendia of transcript profiles. By scanning for the best co-responses among changing transcript levels, CSB.DB allows to infer hypotheses on the functional interaction of genes. These hypotheses are novel and not accessible through analysis of sequence homology. The database enables the search for pairs of genes and larger units of genes, which are under common transcriptional control. In addition, statistical tools are offered to the user, which allow validation and comparison of those co-responses that were discovered by gene queries performed on the currently available set of pre-selectable datasets. AVAILABILITY: All co-response databases can be accessed through the CSB.DB Web server (http://csbdb.mpimp-golm.mpg.de/).

Database Management Systems↗

Identification and characterization of an anterior fat body protein in an insect.

We purified a novel protein with a molecular mass of 34 kDa from the Sarcophaga larval fat body. This protein, named AFP (anterior fat body protein), was restricted almost exclusively to the anterior fat body. The AFP content decreased after pupation on disintegration of the fat body tissue. cDNA analysis revealed that this protein consists of 306 amino acid residues and exhibits significant structural similarity with mammalian regucalcin (senescence marker protein-30), a calcium-binding liver protein. However, AFP did not seem to exhibit strong affinity with calcium. These results suggested that a seemingly uniform fat body tissue exhibits a regional difference in its function along the anterior-posterior axis.

Amino Acid Sequence↗

Global analysis of HuR-regulated gene expression in colon cancer systems of reducing complexity.

HuR, a protein that binds to target mRNAs and can enhance their stability and translation, is increasingly recognized as a pivotal regulator of gene expression during cell division and tumorigenesis. We sought to identify collections of HuR-regulated mRNAs in colon cancer cells by systematic, cDNA array-based assessment of gene expression in three systems of varying complexity. First, comparison of gene expression profiles among tumors with different HuR abundance revealed highly divergent gene expression patterns, and virtually no changes in previously reported HuR target mRNAs. Assessment of gene expression patterns in a second system of reduced complexity, cultured colon cancer cells expressing different HuR levels, rendered more conserved sets of HuR-regulated mRNAs. However, the definitive identification of direct HuR target mRNAs required a third system of still lower complexity, wherein HuR-RNA complexes immunoprecipitated from colon cancer cells were subject to cDNA array hybridization to elucidate the endogenous HuR-bound mRNAs. Comparison of the transcript sets identified in each system revealed a strikingly limited overlap in HuR-regulated mRNAs. The data derived from this systematic analysis of HuR-regulated genes highlight the value of low-complexity, biochemical characterization of protein-RNA interactions. More importantly, however, the data underscore the broad usefulness of integrated approaches comprising systems of low complexity (protein-nucleic acid) and high complexity (cells, tumors) to comprehensively elucidate the gene regulatory events that underlie biological processes.

Animals↗

A microarray data analysis framework for postmortem tissues.

This paper will give a complete methodological approach to the processing of oligonucleotide microarray data from postmortem tissue, particularly brain matter. Attention will be drawn to each of the important stages in the process; specifically the quality control, gene expression value calculation, multiple hypothesis testing and correlation analyses. We shall initially discuss the theoretical foundations of each individual method and subsequently apply the ensemble to a sample data set to illustrate and visualise important points.

Algorithms↗

Glucocorticoid receptor-induced MAPK phosphatase-1 (MPK-1) expression inhibits paclitaxel-associated MAPK activation and contributes to breast cancer cell survival.

Glucocorticoid receptor (GR) activation has recently been shown to inhibit apoptosis in breast epithelial cells. We have previously described a group of genes that is rapidly up-regulated in these cells following dexamethasone (Dex) treatment. In an effort to dissect the mechanisms of GR-mediated breast epithelial cell survival, we now examine the molecular events downstream of GR activation. Here we show that GR activation leads to both the rapid induction of MAPK phosphatase-1 (MKP-1) mRNA and its sustained expression. Induction of the MKP-1 protein in the MCF10A-Myc and MDA-MB-231 breast epithelial cell lines was also seen. Paclitaxel treatment resulted in MAPK activation and apoptosis of MDA-MB-231 breast cancer cells, and both processes were inhibited by Dex pretreatment. Furthermore, induction of MKP-1 correlated with the inhibition of extracellular signal-regulated kinase (ERK1/2) and c-Jun N-terminal kinase (JNK) activity, whereas p38 activity was minimally affected. Blocking Dex-induced MKP-1 induction using small interfering RNA increased ERK1/2 and JNK phosphorylation and decreased cell survival. ERK1/2 and JNK inactivation was associated with Ets-like transcription factor-1 (ELK-1) dephosphorylation. To explore the gene expression changes that occur downstream of ELK-1 dephosphorylation, we used a combination of temporal gene expression data and promoter element analyses. This approach revealed a previously unrecognized transcriptional target of ELK-1, the human tissue plasminogen activator (tPA). We verified the predicted ELK-1--> tPA transcriptional regulatory relationship using a luciferase reporter assay. We conclude that GR-mediated MAPK inactivation contributes to cell survival and that the potential transcriptional targets of this inhibition can be identified from large scale gene array analysis.

Amino Acid Motifs↗

Protein profiles associated with survival in lung adenocarcinoma.

Morphologic assessment of lung tumors is informative but insufficient to adequately predict patient outcome. We previously identified transcriptional profiles that predict patient survival, and here we identify proteins associated with patient survival in lung adenocarcinoma. A total of 682 individual protein spots were quantified in 90 lung adenocarcinomas by using quantitative two-dimensional polyacrylamide gel electrophoresis analysis. A leave-one-out cross-validation procedure using the top 20 survival-associated proteins identified by Cox modeling indicated that protein profiles as a whole can predict survival in stage I tumor patients (P = 0.01). Thirty-three of 46 survival-associated proteins were identified by using mass spectrometry. Expression of 12 candidate proteins was confirmed as tumor-derived with immunohistochemical analysis and tissue microarrays. Oligonucleotide microarray results from both the same tumors and from an independent study showed mRNAs associated with survival for 11 of 27 encoded genes. Combined analysis of protein and mRNA data revealed 11 components of the glycolysis pathway as associated with poor survival. Among these candidates, phosphoglycerate kinase 1 was associated with survival in the protein study, in both mRNA studies and in an independent validation set of 117 adenocarcinomas and squamous lung tumors using tissue microarrays. Elevated levels of phosphoglycerate kinase 1 in the serum were also significantly correlated with poor outcome in a validation set of 107 patients with lung adenocarcinomas using ELISA analysis. These studies identify new prognostic biomarkers and indicate that protein expression profiles can predict the outcome of patients with early-stage lung cancer.

Adenocarcinoma↗

Novel phylogenetic assignment database for terminal-restriction fragment length polymorphism analysis of human colonic microbiota.

Various molecular-biological approaches using the 16S rRNA gene sequence have been used for the analysis of human colonic microbiota. Terminal- restriction fragment length polymorphism (T-RFLP) analysis is suitable for a rapid comparison of complex bacterial communities. Terminal-restriction fragment (T-RF) length can be calculated from a known sequence, thus one can predict bacterial species on the basis of their T-RF length by this analysis. The aim of this study was to build a phylogenetic assignment database for T-RFLP analysis of human colonic microbiota (PAD-HCM), and to demonstrate the effectiveness of PAD-HCM compared with the results of 16S rRNA gene clone library analysis. PAD-HCM was completed to include 342 sequence data obtained using four restriction enzymes. Approximately 80% of the total clones detected by 16S rRNA gene clone library analysis were the same bacterial species or phylotypes as those assigned from T-RF using PAD-HCM. Moreover, large T-RFs consisted of common species or phylotypes detected by both analytical methods. All pseudo-T-RFs identified by mung bean nuclease digestion could not be assigned to a bacterial species or phylotype, and this finding shows that pseudo-T-RFs can also be predicted using PAD-HCM. We conclude that PAD-HCM built in this study enables the prediction of T-RFs at the species level including difficult-to-culture bacteria, and that it is very useful for the T-RFLP analysis of human colonic microbiota.

Adult↗

Hydrogen sulfide induces serum-independent cell cycle entry in nontransformed rat intestinal epithelial cells.

Hydrogen sulfide (H2S), produced by commensal sulfate-reducing bacteria, is an environmental insult that potentially contributes to chronic intestinal epithelial disorders. We tested the hypothesis that exposure of nontransformed intestinal epithelial cells (IEC-18) to the reducing agent sodium hydrogen sulfide (NaHS) activates molecular pathways that underlie epithelial hyperplasia, a phenotype common to both ulcerative colitis (UC) and colorectal cancer. Exposure of IEC-18 cells to NaHS rapidly increased the NADPH/NADP ratio, reduced the intracellular redox environment, and inhibited mitochondrial respiratory activity. The addition of 0.2-5 mM NaHS for 4 h increased the IEC-18 proliferative cell fraction (P<0.05), as evidenced by analysis of the cell cycle and proliferating cell nuclear antigen expression, while apoptosis occurred only at the highest concentration of NaHS. Thirty minutes of NaHS exposure increased (P<0.05) c-Jun mRNA concentrations, consistent with the observed activation of mitogen activated protein kinases (MAPK). Microarray analysis confirmed an increase (P<0.05) in MAPK-mediated proliferative activity, likely reflecting the reduced redox environment of NaHS-treated cells. These data identify functional pathways by which H2S may initiate epithelial dysregulation and thereby contribute to UC or colorectal cancer. Thus, it becomes crucial to understand how genetic background may affect epithelial responsiveness to this bacterial-derived environmental insult.

Animals↗

Phylogeography of western Pacific Leucetta 'chagosensis' (Porifera: Calcarea) from ribosomal DNA sequences: implications for population history and conservation of the Great Barrier Reef World Heritage Area (Australia).

Leucetta 'chagosensis' is a widespread calcareous sponge, occurring in shaded habitats of Indo-Pacific coral reefs. In this study we explore relationships among 19 ribosomal DNA sequence types (the ITS1-5.8S-ITS2 region plus flanking gene sequences) found among 54 individuals from 28 locations throughout the western Pacific, with focus on the Great Barrier Reef (GBR). Maximum parsimony analysis revealed phylogeographical structuring into four major clades (although not highly supported by bootstrap analysis) corresponding to the northern/central GBR with Guam and Taiwan, the southern GBR and subtropical regions south to Brisbane, Vanuatu and Indonesia. Subsequent nested clade analysis (NCA) confirmed this structure with a probability of > 95%. After NCA of geographical distances, a pattern of range expansion from the internal Indonesian clade was inferred at the total cladogram level, as the Indonesian clade was found to be the internal and therefore oldest clade. Two distinct clades were found on the GBR, which narrowly overlap geographically in a line approximately from the Whitsunday Islands to the northern Swain Reefs. At various clade levels, NCA inferred that the northern GBR clade was influenced by past fragmentation and contiguous range expansion events, presumably during/after sea level low stands in the Pleistocene, after which the northern GBR might have been recolonized from the Queensland Plateau in the Coral Sea. The southern GBR clade is most closely related to subtropical L. 'chagosensis', and we infer that the southern GBR probably was recolonized from there after sea level low stands, based on our NCA results and supported by oceanographic data. Our results have important implications for conservation and management of the GBR, as they highlight the importance of marginal transition zones in the generation and maintenance of species rich zones, such as the Great Barrier Reef World Heritage Area.

Animals↗

Sequence-based identification of microbial pathogens: a reconsideration of Koch's postulates.

Over 100 years ago, Robert Koch introduced his ideas about how to prove a causal relationship between a microorganism and a disease. Koch's postulates created a scientific standard for causal evidence that established the credibility of microbes as pathogens and led to the development of modern microbiology. In more recent times, Koch's postulates have evolved to accommodate a broader understanding of the host-parasite relationship as well as experimental advances. Techniques such as in situ hybridization, PCR, and representational difference analysis reveal previously uncharacterized, fastidious or uncultivated, microbial pathogens that resist the application of Koch's original postulates, but they also provide new approaches for proving disease causation. In particular, the increasing reliance on sequence-based methods for microbial identification requires a reassessment of the original postulates and the rationale that guided Koch and later revisionists. Recent investigations of Whipple's disease, human ehrlichiosis, hepatitis C, hantavirus pulmonary syndrome, and Kaposi's sarcoma illustrate some of these issues. A set of molecular guidelines for establishing disease causation with sequence-based technology is proposed, and the importance of the scientific concordance of evidence in supporting causal associations is emphasized.

Base Sequence↗