PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatics analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Annotation of genes involved in glycerolipid biosynthesis in Chlamydomonas reinhardtii: discovery of the betaine lipid synthase BTA1Cr.

Lipid metabolism in flowering plants has been intensely studied, and knowledge regarding the identities of genes encoding components of the major fatty acid and membrane lipid biosynthetic pathways is very extensive. We now present an in silico analysis of fatty acid and glycerolipid metabolism in an algal model, enabled by the recent availability of expressed sequence tag and genomic sequences of Chlamydomonas reinhardtii. Genes encoding proteins involved in membrane biogenesis were predicted on the basis of similarity to proteins with confirmed functions and were organized so as to reconstruct the major pathways of glycerolipid synthesis in Chlamydomonas. This analysis accounts for the majority of genes predicted to encode enzymes involved in anabolic reactions of membrane lipid biosynthesis and compares and contrasts these pathways in Chlamydomonas and flowering plants. As an important result of the bioinformatics analysis, we identified and isolated the C. reinhardtii BTA1 (BTA1Cr) gene and analyzed the bifunctional protein that it encodes; we predicted this protein to be sufficient for the synthesis of the betaine lipid diacylglyceryl-N,N,N-trimethylhomoserine (DGTS), a major membrane component in Chlamydomonas. Heterologous expression of BTA1Cr led to DGTS accumulation in Escherichia coli, which normally lacks this lipid, and allowed in vitro analysis of the enzymatic properties of BTA1Cr. In contrast, in the bacterium Rhodobacter sphaeroides, two separate proteins, BtaARs and BtaBRs, are required for the biosynthesis of DGTS. Site-directed mutagenesis of the active sites of the two domains of BTA1Cr allowed us to study their activities separately, demonstrating directly their functional homology to the bacterial orthologs BtaARs and BtaBRs.

Algal Proteins↗

[Cloning of an expressed sequence tag with restriction display polymerase chain reaction].

OBJECTIVE: To isolate gene fragments from SH-SY5Y cells by way of restriction display polymerase chain reaction (RD-PCR). METHODS: Total mRNA was extracted from SH-SY5Y cells followed by synthesis of the single-strand cDNA with Oligo (dT18) as the anchored primer, and the second strand was synthesized by nick translation. The double strands were cleft with restriction enzyme Sau3A I and the fragments ligated with a universal adapter to be amplified with the universal primers and selected primers. The products were then ligated into the pMD18-T vector and sequenced. RESULTS: One of the sequenced clones was retrieved in the National Center for Biotechnology Information (NCBI) databases with Blast program. The results showed that the sequence possessed great similarity to one fragment of the 17th chromosome in the genome. Sequence analysis with GenScan software indicated that the EST might be one section of an unknown gene. CONCLUSION: RD-PCR provides simple and efficient approach for isolating EST from cells, and cDNA clone sequencing combined with bioinformatics analysis may be helpful in identifying new genes.

Base Sequence↗

Identification and analysis of the mouse basic/Helix-Loop-Helix transcription factor family.

The basic/Helix-Loop-Helix (bHLH) proteins are a family of transcription factors that regulates a variety of biological processes. Based on a previously defined consensus motif, we identified the complete set of bHLH protein family from the mouse proteome databases and carried out a series of bioinformatics analysis. As results, 124 mouse bHLH proteins were identified in this study, and 28 of them were additional bHLH proteins beyond the previous report. These 124 mouse bHLH proteins were classified into groups from A to F by the nomenclature and phylogenetic analysis. Statistic analysis of the Gene Ontology annotation of these proteins showed that the bHLH proteins tend to perform functions related to cell differentiation and development. Gene function enrichment analysis among six groups illuminated that the proteins in certain group tend to have special biology functions, so that the molecular function of the uncharacterized proteins in groups could be inferred.

Amino Acid Sequence↗

2DDB - a bioinformatics solution for analysis of quantitative proteomics data.

BACKGROUND: We present 2DDB, a bioinformatics solution for storage, integration and analysis of quantitative proteomics data. As the data complexity and the rate with which it is produced increases in the proteomics field, the need for flexible analysis software increases. RESULTS: 2DDB is based on a core data model describing fundamentals such as experiment description and identified proteins. The extended data models are built on top of the core data model to capture more specific aspects of the data. A number of public databases and bioinformatical tools have been integrated giving the user access to large amounts of relevant data. A statistical and graphical package, R, is used for statistical and graphical analysis. The current implementation handles quantitative data from 2D gel electrophoresis and multidimensional liquid chromatography/mass spectrometry experiments. CONCLUSION: The software has successfully been employed in a number of projects ranging from quantitative liquid-chromatography-mass spectrometry based analysis of transforming growth factor-beta stimulated fi-broblasts to 2D gel electrophoresis/mass spectrometry analysis of biopsies from human cervix. The software is available for download at SourceForge.

Computational Biology↗

Long-term depression activates transcription of immediate early transcription factor genes: involvement of serum response factor/Elk-1.

Long-term depression (LTD) is one of the paradigms used in vivo or ex vivo for studying memory formation. In order to identify genes with potential relevance for memory formation we used mouse organotypic hippocampal slice cultures in which chemical LTD was induced by applications of 3,5-dihydroxyphenylglycine (DHPG). The induction of chemical LTD was robust, as monitored electrophysiologically. Gene expression analysis after chemical LTD induction was performed using cDNA microarrays containing >7,000 probes. The DHPG-induced expression of immediate early genes (c-fos, junB, egr1 and nr4a1) was subsequently verified by TaqMan polymerase chain reaction. Bioinformatic analysis suggested a common regulator element [serum response factor (SRF)/Elk-1 binding sites] within the promoter region of these genes. Indeed, here we could show a DHPG-dependent binding of SRF at the SRF response element (SRE) site within the promoter region of c-fos and junB. However, SRF binding to egr1 promoter sites was constitutive. The phosphorylation of the ternary complex factor Elk-1 and its localization in the nucleus of hippocampal neurones after DHPG treatment was shown by immunofluorescence using a phosphospecific antibody. We suggest that LTD leads to SRF/Elk-1-regulated gene expression of immediate early transcription factors, which could in turn promote a second broader wave of gene expression.

Animals↗

Conserved features of type III secretion.

Type III secretion systems (TTSSs) are essential mediators of the interaction of many Gram-negative bacteria with human, animal or plant hosts. Extensive sequence and functional similarities exist between components of TTSS from bacteria as diverse as animal and plant pathogens. Recent crystal structure determinations of TTSS proteins reveal extensive structural homologies and novel structural motifs and provide a basis on which protein interaction networks start to be drawn within the TTSSs, that are consistent with and help rationalize genetic and biochemical data. Such studies, along with electron microscopy, also established common architectural design and function among the TTSSs of plant and mammalian pathogens, as well as between the TTSS injectisome and the flagellum. Recent comparative genomic analysis, bioinformatic genome mining and genome-wide functional screening have revealed an unsuspected number of newly discovered effectors, especially in plant pathogens and uncovered a wider distribution of TTSS in pathogenic, symbiotic and commensal bacteria. Functional proteomics and analysis further reveals common themes in TTSS effector functions across phylogenetic host and pathogen boundaries. Based on advances in TTSS biology, new diagnostics, crop protection and drug development applications, as well as new cell biology research tools are beginning to emerge.

Amino Acid Sequence↗

Functional assignment of the 20 S proteasome from Trypanosoma brucei using mass spectrometry and new bioinformatics approaches.

As experimental technologies for characterization of proteomes emerge, bioinformatic analysis of the data becomes essential. Separation and identification technologies currently based on two-dimensional gels/mass spectrometry provide the inherent analytical power required. This strategy involves protein spot digestion and accurate mass mapping together with computational interrogation of available data bases for protein functional identification. When either no exact match is found or when the possible matches only partially account for molecular weights actually observed, peptide sequencing by tandem mass spectrometry has emerged as the methodology of choice to provide the basic additional information required. To evaluate the capabilities of bioinformatics methods employed for identifying homologs of a protein of interest, we attempted to identify the major proteins from the 20 S proteasome of Trypanosoma brucei using sequence information determined using mass spectrometry. The results suggest that neither the traditional query engines, BLAST and FASTA, nor specialized software developed for analysis of sequence information obtained by mass spectrometry are able to identify even closely related sequences at statistically significant scores. To address this deficit, new bioinformatics approaches were developed for concomitant use of the multiple fragments of short sequence typically available from methods of tandem mass spectrometry. These approaches rely on the occurrence of congruence across searches of multiple fragments from a single protein. This method resulted in sharply better statistical significance values for correct hits in the data base output relative to that achieved for independent searches using single sequence fragments.

Algorithms↗

Detection of circulating cancer cells with K-ras oncogene using membrane array.

K-ras oncogene is frequently found in human cancers and thus may serve as a potential diagnostic marker for cancer cells in circulation. So far, there is no reliable method for detecting cancer cells with K-ras oncogene in peripheral blood. The objective of this study was to develop a diagnostic membrane array using activated K-ras oncogene-associated molecules as detection targets. In our previous study, cDNA microarray analysis showed that there were 94 genes differentially expressed in K-ras mutant stably transfected adrenocortical cells. In the present study, we obtained 22 up-regulated genes in the closest relation to K-ras oncogene through bioinformatic analysis. At first, we carried out membrane array analysis by using in vitro culture cells. We demonstrated that this diagnostic technique was feasible and highly sensitive. A number as low as 5 cancer cells bearing K-ras oncogene in 1 ml of blood could be distinctively detected. Then, we collected blood specimens from 76 cancer patients. Direct sequencing analysis of these 76 samples showed that K-ras mutation was present in 43 patients with mutation sites mainly at codons 13, 15 and 61, which have been commonly established to be activated sites. We subsequently analyzed these 76 specimens with our diagnostic membrane array. Thirty-nine specimens were detected as positive for activated K-ras oncogene. Eighty percent (12/15) of mutations occurred at codon 13, 72.7% (8/11) at codon 61, and 88.9% (8/9) at codon 15 were accurately detected by our diagnostic membrane. Finally, through a series of biostatistical analyses, the sensitivity, specificity and accuracy of the diagnostic membrane array were 83.7, 90.9 and 86.8%, respectively. These findings suggest that the K-ras oncogene membrane array has a great potential for further investigation and clinical application.

Computational Biology↗

Exploring drug action on Mycobacterium tuberculosis using affymetrix oligonucleotide genechips.

DNA microarrays have rapidly emerged as an important tool for Mycobacterium tuberculosis research. While the microarray approach has generated valuable information, a recent survey has found a lack of correlation among the microarray data produced by different laboratories on related issues, raising a concern about the credibility of research findings. The Affymetrix oligonucleotide array has been shown to be more reliable for interrogating changes in gene expression than other platforms. However, this type of array system has not been applied to the pharmacogenomic study of M. tuberculosis. The goal here was to explore the strength of the Affymetrix array system for monitoring drug-induced gene expression in M. tuberculosis, compare with other related studies, and conduct cross-platform analysis. The genome-wide gene expression profiles of M. tuberculosis in response to drug treatments including INH (isoniazid) and ethionamide were obtained using the Affymetrix array system. Up-regulated or down-regulated genes were identified through bioinformatic analysis of the microarray data derived from the hybridization of RNA samples and gene probes. Based on the Affymetrix system, our method identified all drug-induced genes reported in the original reference work as well as some other genes that have not been recognized previously under the same drug treatment. For instance, the Affymetrix system revealed that Rv2524c (fas) was induced by both INH and ethionamide under the given levels of concentration, as suggested by most of the probe sets implementing this gene sequence. This finding is contradictory to previous observations that the expression of fas is not changed by INH treatment. This example illustrates that the determination of expression change for certain genes is probe-dependent, and the appropriate use of multiple probe-set representation is an advantage with the Affymetrix system. Our data also suggest that whereas the up-regulated gene expression pattern reflects the drug's mode of action, the down-regulated pattern is largely non-specific. According to our analysis, the Affymetrix array system is a reliable tool for studying the pharmacogenomics of M. tuberculosis and lends itself well in the research and development of anti-TB drugs.

Antitubercular Agents↗

Gene expression profiling of depression and suicide in human prefrontal cortex.

Mood disorders are a major cause of disability. Etiology includes genetic and environmental factors, but the responsible genes have yet to be identified. Using DNA microarrays, we have conducted a large-scale gene expression analysis, in two regions of the human prefrontal cortex from post-mortem matched groups of subjects with major depression who had died by suicide, and control subjects who died from other causes and were free from psychiatric disorders. Bioinformatic analysis was used to investigate molecular and cellular pathways potentially involved in depression and suicidal behavior. We tested several hypotheses of disease pathology and of their putative molecular impact, including changes in single genes, the existence of subgroups of patients or disease subtypes, or the possibility of common biological pathways being affected in the disease process. Within the analytical limits of this relatively large genomic study, we found no evidence for molecular differences that correlated with depression and suicide, suggesting a pathology that is below the detection level of current genomic approaches, or that is either localized to other brain areas, or more associated with post-transcriptional effects and/or changes in protein levels or functions, rather than altered transcriptome in the prefrontal cortex.

Adult↗

Complete genomic sequence of the virulent Salmonella bacteriophage SP6.

We report the complete genome sequence of enterobacteriophage SP6, which infects Salmonella enterica serovar Typhimurium. The genome contains 43,769 bp, including a 174-bp direct terminal repeat. The gene content and organization clearly place SP6 in the coliphage T7 group of phages, but there is approximately 5 kb at the right end of the genome that is not present in other members of the group, and the homologues of T7 genes 1.3 through 3 appear to have undergone an unusual reorganization. Sequence analysis identified 10 putative promoters for the SP6-encoded RNA polymerase and seven putative rho-independent terminators. The terminator following the gene encoding the major capsid subunit has a termination efficiency of about 50% with the SP6-encoded RNA polymerase. Phylogenetic analysis of phages related to SP6 provided clear evidence for horizontal exchange of sequences in the ancestry of these phages and clearly demarcated exchange boundaries; one of the recombination joints lies within the coding region for a phage exonuclease. Bioinformatic analysis of the SP6 sequence strongly suggested that DNA replication occurs in large part through a bidirectional mechanism, possibly with circular intermediates.

Amino Acid Sequence↗

The planetary biology of cytochrome P450 aromatases.

BACKGROUND: Joining a model for the molecular evolution of a protein family to the paleontological and geological records (geobiology), and then to the chemical structures of substrates, products, and protein folds, is emerging as a broad strategy for generating hypotheses concerning function in a post-genomic world. This strategy expands systems biology to a planetary context, necessary for a notion of fitness to underlie (as it must) any discussion of function within a biomolecular system. RESULTS: Here, we report an example of such an expansion, where tools from planetary biology were used to analyze three genes from the pig Sus scrofa that encode cytochrome P450 aromatases-enzymes that convert androgens into estrogens. The evolutionary history of the vertebrate aromatase gene family was reconstructed. Transition redundant exchange silent substitution metrics were used to interpolate dates for the divergence of family members, the paleontological record was consulted to identify changes in physiology that correlated in time with the change in molecular behavior, and new aromatase sequences from peccary were obtained. Metrics that detect changing function in proteins were then applied, including KA/KS values and those that exploit structural biology. These identified specific amino acid replacements that were associated with changing substrate and product specificity during the time of presumed adaptive change. The combined analysis suggests that aromatase paralogs arose in pigs as a result of selection for Suoidea with larger litters than their ancestors, and permitted the Suoidea to survive the global climatic trauma that began in the Eocene. CONCLUSIONS: This combination of bioinformatics analysis, molecular evolution, paleontology, cladistics, global climatology, structural biology, and organic chemistry serves as a paradigm in planetary biology. As the geological, paleontological, and genomic records improve, this approach should become widely useful to make systems biology statements about high-level function for biomolecular systems.

Amino Acid Sequence↗

Combined application of behavior genetics and microarray analysis to identify regional expression themes and gene-behavior associations.

In this report we link candidate genes to complex behavioral phenotypes by using a behavior genetics approach. Gene expression signatures were generated for the prefrontal cortex, ventral striatum, temporal lobe, periaqueductal gray, and cerebellum in eight inbred strains from priority group A of the Mouse Phenome Project. Bioinformatic analysis of regionally enriched genes that were conserved across all strains revealed both functional and structural specialization of particular brain regions. For example, genes encoding proteins with demonstrated anti-apoptotic function were over-represented in the cerebellum, whereas genes coding for proteins associated with learning and memory were enriched in the ventral striatum, as defined by the Expression Analysis Systematic Explorer (EASE) application. Association of regional gene expression with behavioral phenotypes was exploited to identify candidate behavioral genes. Phenotypes that were investigated included anxiety, drug-naive and ethanol-induced distance traveled across a grid floor, and seizure susceptibility. Several genes within the glutamatergic signaling pathway (i.e., NMDA/glutamate receptor subunit 2C, calmodulin, solute carrier family 1 member 2, and glutamine synthetase) were identified in a phenotype-dependent and region-specific manner. In addition to supporting evidence in the literature, many of the genes that were identified could be mapped in silico to surrogate behavior-related quantitative trait loci. The approaches and data set described herein serve as a valuable resource to investigate the genetic underpinning of complex behaviors.

Alcoholism↗

Transforming growth factor-beta-regulated gene transcription and protein expression in human GFAP-negative lamina cribrosa cells.

Primary open-angle glaucoma (POAG) is a progressive optic neuropathy, which is a major cause of worldwide visual impairment and blindness. Pathological hallmarks of the glaucomatous optic nerve head (ONH) include retinal ganglion cell axon loss and extracellular matrix (ECM) remodeling of the lamina cribrosa layer. Transforming growth factor-beta (TGF-beta) is an important pro-fibrotic modulator of ECM metabolism, whose levels are elevated in human POAG lamina cribrosa tissue compared with non-glaucomatous controls. We hypothesize that in POAG, lamina cribrosa (LC) glial cells respond to elevated TGF-beta, producing a remodeled ONH ECM. Using Affymetrix microarrays, we report the first study examining the effect of TGF-beta1 on global gene expression profiles in glial fibrillary acidic acid (GFAP)-negative LC glial cells in vitro. Prominent among the differentially expressed genes were those with established fibrogenic potential, including CTGF, collagen I, elastin, thrombospondin, decorin, biglycan, and fibromodulin. Independent TaqMan and Sybr Green quantitative PCR analysis significantly validated genes involved in regulation of cell proliferation (platelet-derived growth factor [PDGF-alpha]), angiogenesis (vascular endothelial growth factor [VEGF]), ECM accumulation and degradation (CTGF, IL-11, and ADAMT-S5), and growth factor binding (ESM-1). Bioinformatic analysis of the ESM-1 promoter identified putative Smad and Runx transcription factor binding sites, and luciferase assays confirmed that TGF-beta1 drives transcription of the ESM-1 gene. TGF-beta1 induces expression and release of ECM components in LC cells, which may be important in regulating matrix remodeling in the lamina cribrosa. In disease states such as POAG, the LC cell may represent an important pro-fibrotic cell type and an attractive target for novel therapeutic strategies.

Cells, Cultured↗

Evolution of NIN-like proteins in Arabidopsis, rice, and Lotus japonicus.

Genetic studies in Lotus japonicus and pea have identified Nin as a core symbiotic gene required for establishing symbiosis between legumes and nitrogen fixing bacteria collectively called Rhizobium. Sequencing of additional Lotus cDNAs combined with analysis of genome sequences from Arabidopsis and rice reveals that Nin homologues in all three species constitute small gene families. In total, the Arabidopsis and rice genomes encode nine and three NIN-like proteins (NLPs), respectively. We present here a bioinformatics analysis and prediction of NLP evolution. On a genome scale we show that in Arabidopsis, this family has evolved through segmental duplication rather than through tandem amplification. Alignment of all predicted NLP protein sequences shows a composition with six conserved modules. In addition, Lotus and pea NLPs contain segments that might characterize NIN proteins of legumes and be of importance for their function in symbiosis. The most conserved region in NLPs, the RWP-RK domain, has secondary structure predictions consistent with DNA binding properties. This motif is shared by several other small proteins in both Arabidopsis and rice. In rice, the RWP-RK domain sequences have diversified significantly more than in Arabidopsis. Database searches reveal that, apart from its presence in Arabidopsis and rice, the motif is also found in the algae Chlamydomonas and in the slime mold Dictyostelium discoideum. Thus, the origin of this putative DNA binding region seems to predate the fungus-plant divide.

Amino Acid Sequence↗

Heterochromatic genes in Drosophila: a comparative analysis of two genes.

Centromeric heterochromatin comprises approximately 30% of the Drosophila melanogaster genome, forming a transcriptionally repressive environment that silences euchromatic genes juxtaposed nearby. Surprisingly, there are genes naturally resident in heterochromatin, which appear to require this environment for optimal activity. Here we report an evolutionary analysis of two genes, Dbp80 and RpL15, which are adjacent in proximal 3L heterochromatin of D. melanogaster. DmDbp80 is typical of previously described heterochromatic genes: large, with repetitive sequences in its many introns. In contrast, DmRpL15 is uncharacteristically small. The orthologs of these genes were examined in D. pseudoobscura and D. virilis. In situ hybridization and whole-genome assembly analysis show that these genes are adjacent, but not centromeric in the genome of D. pseudoobscura, while they are located on different chromosomal elements in D. virilis. Dbp80 gene organization differs dramatically among these species, while RpL15 structure is conserved. A bioinformatic analysis in five additional Drosophila species demonstrates active repositioning of these genes both within and between chromosomal elements. This study shows that Dbp80 and RpL15 can function in contrasting chromatin contexts on an evolutionary timescale. The complex history of these genes also provides unique insight into the dynamic nature of genome evolution.

Amino Acid Sequence↗

Proteomic identification of a novel protein regulated in CA1 and CA3 hippocampal regions during intermittent hypoxia.

The CA1 and CA3 regions of the hippocampus markedly differ in their susceptibility to hypoxia in general, and more particularly to the intermittent hypoxia (IH) that characterizes sleep apnea. We used proteomic analysis to build a database of proteins expressed in normoxic CA1 and CA3. The current hippocampus protein database identifies 106 proteins. A hypothetical protein with accession number AK006737 (gimid R:12839969) was strongly upregulated in the CA1, but not CA3 hippocampal region. Bioinformatic analysis revealed that the unknown protein contained a high stringency protein kinase e binding site. Domain analysis demonstrated the presence of a conserved sequence indicative of macrophage scavenger receptors. Using proteomic analysis we have previously demonstrated that acute (6 h) IH-mediated CA1 injury results from complex interactions between pathways involving increased metabolism, induction of stress-induced proteins and apoptosis, and ultimately disruption of structural proteins and cell integrity. The current findings identify a hypothetical protein that may play a key role in the response of CA1 to IH. These findings provide initial insights into mechanisms underlying differences in susceptibility to hypoxia in neural tissue and demonstrate how proteomic analysis can be used to generate new hypotheses, which define neuronal adaptation to IH.

Animals↗

A distal enhancer in the interferon-gamma (IFN-gamma) locus revealed by genome sequence comparison.

Large-scale cross-species DNA sequence comparison has become a powerful tool to identify conserved cis-regulatory modules of genes. However, bioinformatic analysis alone cannot reveal how an evolutionarily conserved region regulates gene expression: whether it functions as an enhancer, silencer, or insulator; whether its function is cell-type restricted; and whether biologically relevant transcription factors bind to the element. Here we combine bioinformatics with wet-lab techniques to illustrate a general and systematic method of identifying functional conserved regulatory regions of genes. We applied this approach to the interferon-gamma (IFN-gamma) gene. Comparison of human and mouse IFN-gamma reveals a highly conserved non-coding sequence located approximately 5 kb 5' of the transcription start site. This region coincides with constitutive and inducible DNase I hypersensitivity sites present in IFN-gamma-producing Th1 cells but not in Th2 cells that do not produce IFN-gamma. Histone methylation at the 5' conserved non-coding sequences indicates a more accessible chromatin structure in Th1 cells compared with Th2 cells. This element binds two transcription factors known to be essential for IFN-gamma expression: nuclear factor of activated T cells, an inducible transcription factor, and T-box protein expressed in T cells, a cell lineage-restricted transcription factor. Together, these findings identify a highly conserved distal enhancer in the IFN-gamma cytokine locus and validate our approach as a successful method to detect cis-regulatory elements.

Animals↗