PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 685 records · Page 38Linked to original sources

Implied alignment: a synapomorphy-based multiple-sequence alignment method and its use in cladogram search.

A method to align sequence data based on parsimonious synapomorphy schemes generated by direct optimization (DO; earlier termed optimization alignment) is proposed. DO directly diagnoses sequence data on cladograms without an intervening multiple-alignment step, thereby creating topology-specific, dynamic homology statements. Hence, no multiple-alignment is required to generate cladograms. Unlike general and globally optimal multiple-alignment procedures, the method described here, implied alignment (IA), takes these dynamic homologies and traces them back through a single cladogram, linking the unaligned sequence positions in the terminal taxa via DO transformation series. These "lines of correspondence" link ancestor-descendent states and, when displayed as linearly arrayed columns without hypothetical ancestors, are largely indistinguishable from standard multiple alignment. Since this method is based on synapomorphy, the treatment of certain classes of insertion-deletion (indel) events may be different from that of other alignment procedures. As with all alignment methods, results are dependent on parameter assumptions such as indel cost and transversion:transition ratios. Such an IA could be used as a basis for phylogenetic search, but this would be questionable since the homologies derived from the implied alignment depend on its natal cladogram and any variance, between DO and IA + Search, due to heuristic approach. The utility of this procedure in heuristic cladogram searches using DO and the improvement of heuristic cladogram cost calculations are discussed.

Animals↗

Detection of periodicity in eukaryotic genomes on the basis of power spectrum analysis.

In the present study, we identified periodic patterns in nucleotide sequence, and characterized nucleotide sequences that confer periodicities to Arabidopsis thaliana and Drosophila melanogaster on the basis of a power spectrum method and frequency of nucleotide sequences. To assign regions that contribute to each periodicity we calculated periodic nucleotide distributions by a parameter proposed in the paper. In A. thaliana, we obtained three periodicities (248 bp-, 167 bp-, and 126 bp) in chromosome 3, three peaks (174 bp-, 88 bp-, and 59 bp-period) in chromosome 4, and four periodicities (356 bp, 174 bp, 88 bp, and 59 bp) in chromosome 5. These are relation to ORF that consists of Gly-rich amino acid sequences including histone protein that consists of Gly-, Ser-, and Ala-rich amino acids residues. For D. melanogaster genome we found that G or C spectral curves have flat region at middle frequency range from f = 10(-4) to 10(-5) (corresponding to cyclic size 1 kb-5 kb), which may be associated with randomness of base sequence composition. This property has not been observed in Saccharomyces cerevisiae, Caenorhabditis elegans, and Homo sapiens yet.

Animals↗

Mechanisms of indomethacin-induced alterations in the choline phospholipid metabolism of breast cancer cells.

Human mammary epithelial cells (HMECs) exhibit an increase in phosphocholine (PC) and total choline-containing compounds, as well as a switch from high glycerophosphocholine (GPC)/low PC to low GPC/high PC, with progression to malignant phenotype. The treatment of human breast cancer cells with a nonsteroidal anti-inflammatory agent, indomethacin, reverted the high PC/low GPC pattern to a low PC/high GPC pattern indicative of a less malignant phenotype, supported by decreased invasion. Here, we have characterized mechanisms underlying indomethacin-induced alterations in choline membrane metabolism in malignant breast cancer cells and nonmalignant HMECs labeled with [1,2-13C]choline using 1H and 13C magnetic resonance spectroscopy. Microarray gene expression analysis was performed to understand the molecular mechanisms underlying these changes. In breast cancer cells, indomethacin treatment activated phospholipases that, combined with an increased choline phospholipid biosynthesis, led to increased GPC and decreased PC levels. However, in nonmalignant HMECs, activation of the anabolic pathway alone was detected following indomethacin treatment. Following indomethacin treatment in breast cancer cells, several candidate genes, such as interleukin 8, NGFB, CSF2, RHOB, EDN1, and JUNB, were differentially expressed, which may have contributed to changes in choline metabolism through secondary effects or signaling cascades leading to changes in enzyme activity.

Anti-Inflammatory Agents, Non-Steroidal↗

Three-dimensional cluster analysis identifies interfaces and functional residue clusters in proteins.

Three-dimensional cluster analysis offers a method for the prediction of functional residue clusters in proteins. This method requires a representative structure and a multiple sequence alignment as input data. Individual residues are represented in terms of regional alignments that reflect both their structural environment and their evolutionary variation, as defined by the alignment of homologous sequences. From the overall (global) and the residue-specific (regional) alignments, we calculate the global and regional similarity matrices, containing scores for all pairwise sequence comparisons in the respective alignments. Comparing the matrices yields two scores for each residue. The regional conservation score (C(R)(x)) defines the conservation of each residue x and its neighbors in 3D space relative to the protein as a whole. The similarity deviation score (S(x)) detects residue clusters with sequence similarities that deviate from the similarities suggested by the full-length sequences. We evaluated 3D cluster analysis on a set of 35 families of proteins with available cocrystal structures, showing small ligand interfaces, nucleic acid interfaces and two types of protein-protein interfaces (transient and stable). We present two examples in detail: fructose-1,6-bisphosphate aldolase and the mitogen-activated protein kinase ERK2. We found that the regional conservation score (C(R)(x)) identifies functional residue clusters better than a scoring scheme that does not take 3D information into account. C(R)(x) is particularly useful for the prediction of poorly conserved, transient protein-protein interfaces. Many of the proteins studied contained residue clusters with elevated similarity deviation scores. These residue clusters correlate with specificity-conferring regions: 3D cluster analysis therefore represents an easily applied method for the prediction of functionally relevant spatial clusters of residues in proteins.

Adenosine Triphosphate↗

Microarray analysis of peroxisome proliferator-activated receptor-gamma induced changes in gene expression in macrophages.

We used a combination of expression microarray and Northern blot analyses to identify target genes for peroxisome proliferator-activated receptor (PPAR) gamma in RAW264.7 macrophages. PPARgamma natural ligand 15-deoxy-Delta(12,14) prostaglandin and synthetic ligands ciglitazone and rosiglitazone increased the expression of scavenger receptor CD36 and ATP-binding cassette transporter A1, as well as adipophilin (a lipid droplet coating protein involved in intracellular lipid storage and transport), calpain (a protease implicated in ABCA1 protein degradation), and ADAM8 (a disintegrin and metalloprotease protein involved in cell adhesion). These findings are relevant to understanding the effect of PPARgamma activation on gene expression and cognate pathways in macrophages.

Animals↗

Bayesian sparse hidden components analysis for transcription regulation networks.

MOTIVATION: In systems like Escherichia Coli, the abundance of sequence information, gene expression array studies and small scale experiments allows one to reconstruct the regulatory network and to quantify the effects of transcription factors on gene expression. However, this goal can only be achieved if all information sources are used in concert. RESULTS: Our method integrates literature information, DNA sequences and expression arrays. A set of relevant transcription factors is defined on the basis of literature. Sequence data are used to identify potential target genes and the results are used to define a prior distribution on the topology of the regulatory network. A Bayesian hidden component model for the expression array data allows us to identify which of the potential binding sites are actually used by the regulatory proteins in the studied cell conditions, the strength of their control, and their activation profile in a series of experiments. We apply our methodology to 35 expression studies in E.Coli with convincing results. AVAILABILITY: www.genetics.ucla.edu/labs/sabatti/software.html SUPPLEMENTARY INFORMATION: The supplementary material are available at Bioinformatics online.

Algorithms↗

Genomic profiling of acquired resistance to apoptosis in cells derived from human atherosclerotic lesions: potential role of STATs, cyclinD1, BAD, and Bcl-XL.

Current theories suggest that atherosclerosis, plaque rupture, stroke, and restenosis after angioplasty may involve defective apoptotic mechanisms in vascular cells. Prior work has demonstrated that cells from human atherosclerotic lesions, and cells from the aorta of aged rats, exhibit functional resistance to apoptosis induced by TGF-beta and glucocorticoids. The present studies demonstrate that human lesion-derived cells (LDC) are also resistant to apoptosis induced by fas ligation compared to cells derived from the adjacent media, and that in vitro expansion of LDC causes acquired resistance to apoptosis. Microarray profiling of fas-resistant versus sensitive cells identified a set of genes including STATs, caspase 1, cyclin D1, Bcl-xL, VDAC2, and BAD. The STAT proteins have been implicated in resistance to apoptosis, potentially via their ability to modulate caspase 1 (ICE), Bcl-xL, and cyclin D1 expression. Western blot analysis of sensitive and resistant LDC clonal lines confirmed increases in cyclin D1, STAT6, Bcl-xL, and BAD, with decreased expression of caspase 1. Thus, transcript profiling has identified a potential pathway of apoptotic regulation in subsets of lesion cells. The resistant phenotype may contribute to plaque stability and excessive vascular repair, while sensitive cells may be involved in plaque rupture and infarction. The data suggests both genetic interventions and novel small-molecule inhibitors that may be effective modulators of apoptosis in atherosclerosis, angina, and in-stent restenosis.

Apoptosis↗

Nucleotide sequence analysis of natural and combinatorial anti-PDC-E2 antibodies in patients with primary biliary cirrhosis. Recapitulating immune selection with molecular biology.

We have analyzed at the nucleotide level the variable region gene sequences of five human mAbs and five recombinant Fab fragments derived from the mesenteric lymph nodes of patients with primary biliary cirrhosis. Both mAbs and Fabs were monospecific for dihydrolipoamide acetyltransferase, the E2 subunit of the pyruvate dehydrogenase complex, which has been shown to be the major autoantigen of primary biliary cirrhosis. We found that although the mAbs, mainly of the IgM isotype, were encoded by a diverse array of VH and VL gene segments either as direct copies of germline genes or somatically mutated, the recombinant IgG Fabs expressed clonally related heavy chains displaying a high number of somatic mutations that very likely occurred in the context of Ag selection. Combinatorial pairing of clonally related heavy chains with highly homologous light chains suggests that the IgG anti-pyruvate dehydrogenase complex repertoire of primary biliary cirrhosis patients is the result of the clonal expansion of a restricted set of B cells.

Amino Acid Sequence↗

Retinoblastoma tumor suppressor targets dNTP metabolism to regulate DNA replication.

The retinoblastoma tumor suppressor, RB, is a negative regulator of the cell cycle that is inactivated in the majority of human tumors. Cell cycle inhibition elicited by RB has been attributed to the attenuation of CDK2 activity. Although ectopic cyclins partially overcome RB-mediated S-phase arrest at the replication fork, DNA replication remains inhibited and cells fail to progress to G(2) phase. These data suggest that RB regulates an additional execution point in S phase. We observed that constitutively active RB attenuates the expression of specific dNTP synthetic enzymes: dihydrofolate reductase, ribonucleotide reductase (RNR) subunits R1/R2, and thymidylate synthase (TS). Activation of endogenous RB and related proteins by p16ink4a yielded similar effects on enzyme expression. Conversely, targeted disruption of RB resulted in increased metabolic protein levels (dihydrofolate reductase, TS, RNR-R2) and conferred resistance to the effect of TS or RNR inhibitors that diminish available dNTPs. Analysis of dNTP pools during RB-mediated cell cycle arrest revealed significant depletion, concurrent with the loss of TS and RNR protein. Importantly, the effect of active RB on cell cycle position and available dNTPs was comparable to that observed with specific antimetabolites. Together, these results show that RB-mediated transcriptional repression attenuates available dNTP pools to control S-phase progression. Thus, RB employs both canonical cyclin-dependent kinase/cyclin regulation and metabolic regulation as a means to limit proliferation, underscoring its potency in tumor suppression.

Adenoviridae↗

Solanum nigrum: a model ecological expression system and its tools.

Plants respond to environmental stresses through a series of complicated phenotypic responses, which can be understood only with field studies because other organisms must be recruited for their function. If ecologists are to fully participate in the genomics revolution and if molecular biologists are to understand adaptive phenotypic responses, native plant ecological expression systems that offer both molecular tools and interesting natural histories are needed. Here, we present Solanum nigrum L., a Solanaceous relative of potato and tomato for which many genomic tools are being developed, as a model plant ecological expression system. To facilitate manipulative ecological studies with S. nigrum, we describe: (i) an Agrobacterium-based transformation system and illustrate its utility with an example of the antisense expression of RuBPCase, as verified by Southern gel blot analysis and real-time quantitative PCR; (ii) a 789-oligonucleotide microarray and illustrate its utility with hybridizations of herbivore-elicited plants, and verify responses with RNA gel blot analysis and real-time quantitative PCR; (iii) analyses of secondary metabolites that function as direct (proteinase inhibitor activity) and indirect (herbivore-induced volatile organic compounds) defences; and (iv) growth and fitness-estimates for plants grown under field conditions. Using these tools, we demonstrate that attack from flea beetles elicits: (i) a large transcriptional change consistent with elicitation of both jasmonate and salicylate signalling; and (ii) increases in proteinase inhibitor transcripts and activity, and volatile organic compound release. Both flea beetle attack and jasmonate elicitation increased proteinase inhibitors and jasmonate elicitation decreased fitness in field-grown plants. Hence, proteinase inhibitors and jasmonate-signalling are targets for manipulative studies.

Animals↗

Transcriptome analysis of bud burst in sessile oak (Quercus petraea).

Expression patterns of hundreds of transcripts in apical buds were monitored during bud flushing in sessile oak (Quercus petraea), in order to identify genes differentially expressed between the quiescent and active stage of bud development. Different transcriptomic techniques combining the construction of suppression subtractive hybridization (SSH) libraries and the monitoring of gene expression using macroarray and real-time reverse transcriptase polymerase chain reaction (RT-PCR) were performed to dissect bud burst, with a special emphasis on the onset of the process. We generated 801 expressed sequence tags (ESTs) derived from six developmental stages of bud burst. Macroarray experiment revealed a total of 233 unique transcripts exhibiting differential expression during the process, and a putative function was assigned to 65% of them. Cell rescue/defense-, metabolism-, protein synthesis-, cell cycle- and transcription-related transcripts were among the most regulated genes. Macroarray and real-time RT-PCR showed that several genes exhibited contrasted expressions between quiescent and swelling buds, such as a putative homologue of the transcription factor DAG2 (Dof Affecting Germination 2), previously reported to be involved in the control of seed germination in Arabidopsis thaliana. These differentially expressed genes constitute relevant candidates for signaling pathway of bud burst in trees.

Cluster Analysis↗

Identification of metabolic units induced by environmental signals.

MOTIVATION: Biological cells continually need to adapt the activity levels of metabolic functions to changes in their living environment. Although genome-wide transcriptional data have been gathered in a large variety of environmental conditions, the connections between the expression response to external changes and the induction or repression of specific metabolic functions have not been investigated at the genome scale. RESULTS: We present here a correlation-based analysis for identifying the expression response of genes involved in metabolism to specific external signals, and apply it to analyze the transcriptional response of Saccharomyces cerevisiae to different stress conditions. We show that this approach leads to new insights about the specificity of the genomic response to given environmental changes, and allows us to identify genes that are particularly sensitive to a unique condition. We then integrate these signal-induced expression data with structural data of the yeast metabolic network and analyze the topological properties of the induced or repressed subnetworks. They reveal significant discrepancies from random networks, and in particular exhibit a high connectivity, allowing them to be mapped back to complete metabolic routes.

Adaptation, Physiological↗

Intrinsic heterogeneity in adipose tissue of fat-specific insulin receptor knock-out mice is associated with differences in patterns of gene expression.

Mice with a fat-specific insulin receptor knock-out (FIRKO) have reduced adipose tissue mass, are protected against obesity, and have an extended life span. White adipose tissue of FIRKO mice is also characterized by a polarization into two major populations of adipocytes, one small (<50 microm) and one large (>100 microm), which differ with regard to basal triglyceride synthesis and lipolysis, as well as in the expression of fatty acid synthase, sterol regulatory element-binding protein 1c, and CCAAT/enhancer-binding protein alpha (C/EBP-alpha). Gene expression analysis using RNA isolated from large and small adipocytes of FIRKO and control (IR lox/lox) mice was performed on oligonucleotide microarrays. Of the 12,488 genes/expressed sequence tags represented, 111 genes were expressed differentially in the four populations of adipocytes at the p < 0.001 level. These alterations exhibited 10 defined patterns and occurred in response to two distinct regulatory effects. 63 genes were identified as changed in expression depending primarily upon adipocyte size, including C/EBP-alpha, C/EBP-delta, superoxide dismutase 3, and the platelet-derived growth factor receptor. 48 genes were regulated primarily by impairment of insulin signaling, including transforming growth factor beta, interferon gamma, insulin-like growth factor I receptor, activating transcription factor 3, aldehyde dehydrogenase 2, and protein kinase Cdelta. These data suggest an intrinsic heterogeneity of adipocytes with differences in gene expression related to adipocyte size and insulin signaling.

Adipocytes↗

Physiological function as regulation of large transcriptional programs: the cellular response to genotoxic stress.

The responses to ionizing radiation and other genotoxic environmental stresses are complex and are regulated by a number of overlapping molecular pathways. One such stress signaling pathway involves p53, which regulates the expression of over 100 genes already identified. It is also becoming increasingly apparent that the pattern of stress gene expression has some cell type specificity. It may be possible to exploit these differences in stress gene responsiveness as molecular markers through the use of a combined informatics and functional genomics approach. The techniques of microarray analysis potentially offer the opportunity to monitor changes in gene expression across the entire set of expressed genes in a cell or organism. As an initial step in the development of a functional genomics approach to stress gene analysis, we have recently demonstrated the utility of cDNA microarray hybridization to measure radiation-stress gene responses and identified a number of previously unknown radiation-regulated genes. The responses of some of these genes to DNA-damaging agents vary widely in cell lines from different tissues of origin and different genetic backgrounds. While this again highlights the importance of a cellular context to genotoxic stress responses, it also raises the prospect of expression-profiling of cell lines, tissues, and tumors. Such profiles may have a predictive value if they can define regions of 'expression space' that correlate with important endpoints, such as response to cancer therapy regimens, or identification of exposures to environmental toxins.

Animals↗

Determination of strongly overlapping signaling activity from microarray data.

BACKGROUND: As numerous diseases involve errors in signal transduction, modern therapeutics often target proteins involved in cellular signaling. Interpretation of the activity of signaling pathways during disease development or therapeutic intervention would assist in drug development, design of therapy, and target identification. Microarrays provide a global measure of cellular response, however linking these responses to signaling pathways requires an analytic approach tuned to the underlying biology. An ongoing issue in pattern recognition in microarrays has been how to determine the number of patterns (or clusters) to use for data interpretation, and this is a critical issue as measures of statistical significance in gene ontology or pathways rely on proper separation of genes into groups. RESULTS: Here we introduce a method relying on gene annotation coupled to decompositional analysis of global gene expression data that allows us to estimate specific activity on strongly coupled signaling pathways and, in some cases, activity of specific signaling proteins. We demonstrate the technique using the Rosetta yeast deletion mutant data set, decompositional analysis by Bayesian Decomposition, and annotation analysis using ClutrFree. We determined from measurements of gene persistence in patterns across multiple potential dimensionalities that 15 basis vectors provides the correct dimensionality for interpreting the data. Using gene ontology and data on gene regulation in the Saccharomyces Genome Database, we identified the transcriptional signatures of several cellular processes in yeast, including cell wall creation, ribosomal disruption, chemical blocking of protein synthesis, and, critically, individual signatures of the strongly coupled mating and filamentation pathways. CONCLUSION: This works demonstrates that microarray data can provide downstream indicators of pathway activity either through use of gene ontology or transcription factor databases. This can be used to investigate the specificity and success of targeted therapeutics as well as to elucidate signaling activity in normal and disease processes.

Algorithms↗

Transcriptional profiling of human hematopoiesis during in vitro lineage-specific differentiation.

To better understand the transcriptional program that a ccompanies orderly lineage-specific hematopoietic differentiation, we performed serial oligonucleotide microarray analysis of human normal CD34+ bone marrow cells during lineage-specific differentiation. CD34+ bone marrow cells isolated from healthy individuals were selectively stimulated in vitro with the cytokines erythropoietin (EPO), thrombopoietin (TPO), granulocyte colony-stimulating factor (G-CSF), and granulocyte macrophage colony-stimulating factor (GM-CSF). Cells from each of the lineages were harvested after 4, 7, and 11 days of culture for expression profiling. Gene expression was analyzed by oligonucleotide microarrays (HG-U133A; Affymetrix, Santa Clara, CA). Experiments were done in triplicates. We identified 258 genes that are consistently upregulated or downregulated during the course of lineage-specific differentiation within each specific lineage (horizontal change). In addition, we identified 52 genes that contributed to a specific expression profile, yielding a genetic signature specific for successive stages of differentiation within each of the three lineages. Analysis of horizontal changes selected 21 continuously upregulated genes for EPO-induced differentiation (including GTPase activator proteins RAP1GA1 and ARHGAP8, which regulate small Rho GTPases), 21 for G-CSF-induced/GM-CSF-induced differentiation, and 91 for TPO-induced differentiation (including DLK1, of which the role in normal hematopoiesis is not defined). During the lineage-specific differentiation, 58 (erythropoiesis), 30 (granulopoiesis), and 37 (thrombopoiesis) genes were significantly downregulated, respectively. The expression of selected genes was confirmed by real-time polymerase chain reaction. Our data encompass the first extensive transcriptional profile of human hematopoiesis during in vitro lineage-specific differentiation.

Antigens, CD34↗

Quantification of the effect of vaccination on transmission of avian influenza (H7N7) in chickens.

Recent outbreaks of highly pathogenic avian influenza (HPAI) viruses in poultry and their threatening zoonotic consequences emphasize the need for effective control measures. Although vaccination of poultry against avian influenza provides a potentially attractive control measure, little is known about the effect of vaccination on epidemiologically relevant parameters, such as transmissibility and the infectious period. We used transmission experiments to study the effect of vaccination on the transmission characteristics of HPAI A/Chicken/Netherlands/03 H7N7 in chickens. In the experiments, a number of infected and uninfected chickens is housed together and the infection chain is monitored by virus isolation and serology. Analysis is based on a stochastic susceptible, latently infected, infectious, recovered (SEIR) epidemic model. We found that vaccination is able to reduce the transmission level to such an extent that a major outbreak is prevented, important variables being the type of vaccine (H7N1 or H7N3) and the moment of challenge after vaccination. Two weeks after vaccination, both vaccines completely block transmission. One week after vaccination, the H7N1 vaccine is better than the H7N3 vaccine at reducing the spread of the H7N7 virus. We discuss the implications of these findings for the use of vaccination programs in poultry and the value of transmission experiments in the process of choosing vaccine.

Animals↗

A primer on gene expression and microarrays for machine learning researchers.

Data originating from biomedical experiments has provided machine learning researchers with an important source of motivation for developing and evaluating new algorithms. A new wave of algorithmic development has been initiated with the publication of gene expression data derived from microarrays. Microarray data analysis is particularly challenging given the large number of measurements (typically in the order of thousands) that are reported for relatively few samples (typically in the order of dozens). Many data sets are now available on the web. It is important that machine learning researchers understand how data are obtained and which assumptions are necessary in the analysis. Microarray data have the potential to cause significant impact in machine learning research, not just as a rich and realistic source of cases for testing new algorithms, as has been the UCI machine learning repository in the past decades, but also as a main motivation for their development. In this article, we briefly review the biology underlying microarrays, the process of obtaining gene expression measurements, and the rationale behind the common types of analyses involved in a microarray experiment. We outline the main challenges and reiterate critical considerations regarding the construction of supervised learning models that use this type of data. The goal of this article is to familiarize machine learning researchers with data originated from gene expression microarrays.

Algorithms↗