Genomics. DOE hits potholes on the road to systems biology.
Explore the source record for details and available documents.
SEARCH · PubMed Health
Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Both cDNA microarray and spectroscopic data provide indirect information about the chemical compounds present in the biological tissue under consideration. In this paper simple univariate and bivariate measures are used to investigate correlations between both types of high dimensional analyses. A large dataset of 42 hemp samples on which 3456 cDNA clones and 351 NIR wavelengths have been measured, was analyzed using graphical representations. For this purpose we propose clustered correlation and clustered discrimination images. Large, tissue-related differences are seen to dominate the cDNA-NIR correlation structure but smaller, more difficult to detect, variety-related differences can be found at specific cDNA clone/NIR wavelength combinations.
BACKGROUND: The gene expression profiles of most human tissues have been studied by determining the transcriptome of whole tissue homogenates. Due to the solid composition of tissues it is difficult to study the transcriptomes of individual cell types that compose a tissue. To overcome the problem of heterogeneity we have developed a method to isolate individual cell types from whole tissue that are a source of RNA suitable for transcriptome profiling. RESULTS: Using monoclonal antibodies specific for basal (integrin beta4), luminal secretory (dipeptidyl peptidase IV), stromal fibromuscular (integrin alpha 1), and endothelial (PECAM-1) cells, respectively, we separated the cell types of the prostate with magnetic cell sorting (MACS). Gene expression of MACS-sorted cell populations was assessed with Affymetrix GeneChips. Analysis of the data provided insight into gene expression patterns at the level of individual cell populations in the prostate. CONCLUSION: In this study, we have determined the transcriptome profile of a solid tissue at the level of individual cell types. Our data will be useful for studying prostate development and cancer progression in the context of single cell populations within the organ.
Recent microarray studies of mouse and human osteoblast differentiation in vitro have identified novel transcription factors that may be important in the establishment and maintenance of differentiation. These findings help unravel the pattern of gene-expression changes that underly the complex process of bone formation.
With the increasing number of whole genome sequences available, genomic research has shifted toward the annotation of functional elements and transcribed regions. Thus, the related field of transcriptome research requires accurate methods for the profiling of genes that are not biased by known sequence information, and that also allow for the identification of promoter regions. Starting with serial analysis of gene expression (SAGE), methods making use of short sequencing tags have greatly contributed to transcriptome studies. Here we review recent developments in the use of short sequencing tags in expression profiling, gene discovery and genome annotation. These tags are obtained from the 5' end of mRNAs, both terminal ends of mRNAs, or genomic regions. The 5' end-specific tags, with their ability to identify transcripts along with their transcriptional start sites, will be of particular interest for gene network studies and may become one of the most important approaches in systems biology.
We report a new set of nine primer pairs specifically developed for amplification of Brassica plastid SSR markers. The wide utility of these markers is demonstrated for haplotype identification and detection of polymorphism in B. napus, B. nigra, B. oleracea, B. rapa and in related genera Arabidopsis, Camelina, Raphanus and Sinapis. Eleven gene regions (ndhB-rps7 spacer, rbcL-accD spacer, rpl16 intron, rps16 intron, atpB-rbcL spacer, trnE-trnT spacer, trnL intron, trnL-trnF spacer, trnM-atpE spacer, trnR-rpoC2 spacer, ycf3-psaA spacer) were sequenced from a range of Brassica and related genera for SSR detection and primer design. Other sequences were obtained from GenBank/EMBL. Eight out of nine selected SSR loci showed polymorphism when amplified using the new primers and a combined analysis detected variation within and between Brassica species, with the number of alleles detected per locus ranging from 5 (loci MF-6, MF-1) to 11 (locus MF-7). The combined SSR data were used in a neighbour-joining analysis (SMM, D (DM) distances) to group the samples based on the presence and absence of alleles. The analysis was generally able to separate plastid types into taxon-specific groups. Multi-allelic haplotypes were plotted onto the neighbour joining tree. A total number of 28 haplotypes were detected and these differentiated 22 of the 41 accessions screened from all other accessions. None of these haplotypes was shared by more than one species and some were not characteristic of their predicted type. We interpret our results with respect to taxon differentiation, hybridisation and introgression patterns relating to the 'Triangle of U'.
UNLABELLED: Many bioinformatics problems can be tackled from a fresh angle offered by the network perspective. Directly inspired by metabolic network structural studies, we propose an improved gene clustering approach for inferring gene signaling pathways from gene microarray data. Based on the construction of co-expression networks that consists of both significantly linear and non-linear gene associations together with controlled biological and statistical significance, our approach tends to group functionally related genes into tight clusters despite their expression dissimilarities. We illustrate our approach and compare it to the traditional clustering approaches on a yeast galactose metabolism dataset and a retinal gene expression dataset. Our approach greatly outperforms the traditional approach in rediscovering the relatively well known galactose metabolism pathway in yeast and in clustering genes of the photoreceptor differentiation pathway. AVAILABILITY: The clustering method has been implemented in an R package "GeneNT" that is freely available from: http://www.cran.org.
MOTIVATION: Secondary-Structure Guided Superposition tool (SSGS) is a permissive secondary structure-based algorithm for matching of protein structures and in particular their fragments. The algorithm was developed towards protein structure prediction via fragment assembly. RESULTS: In a fragment-based structural prediction scheme, a protein sequence is cut into building blocks (BBs). The BBs are assembled to predict their relative 3D arrangement. Finally, the assemblies are refined. To implement this prediction scheme, a clustered structural library representing sequence patterns for protein fragments is essential. To create a library, BBs generated by cutting proteins from the PDB are compared and structurally similar BBs are clustered. To allow structural comparison and clustering of the BBs, which are often relatively short with flexible loops, we have devised SSGS. SSGS maintains high similarity between cluster members and is highly efficient. When it comes to comparing BBs for clustering purposes, the algorithm obtains better results than other, non-secondary structure guided protein superimposition algorithms.
The success achieved for protein structure prediction of loop regions with insertions and deletions by knowledge-based methods depends on the quality of the underlying information, i.e. a fragment data bank as complete as possible is needed. However, the greater the number of proteins contributing to the data base the more redundant information is included, which leads to structurally similar proposals in loop predictions and to longer times for extracting fragments. So it is not only necessary to increase the number of proteins for building the loop data base but also to cluster the resulting fragments according to their structural similarities in order to remove redundancy. Here, a new, non-redundant fragment data bank is described, which is based on all proteins in the Brookhaven Protein Data Bank (release 7/95) with a resolution > or = 2.0 A and which can be updated easily by including new information from structures to be solved in the future. In the clustering process presented, the resulting clusters are optimized in several cycles until self-consistency. In this way all redundant information is removed without loosing any significantly different fragments. Finally the resulting fragment data bank is analysed with respect to its completeness.
Pseudomonas aeruginosa strains from the chronic lung infections of cystic fibrosis (CF) patients are phenotypically and genotypically diverse. Using strain PAO1 whole genome DNA microarrays, we assessed the genomic variation in P. aeruginosa strains isolated from young children with CF (6 months to 8 years of age) as well as from the environment. Eighty-nine to 97% of the PAO1 open reading frames were detected in 20 strains by microarray analysis, while subsets of 38 gene islands were absent or divergent. No specific pattern of genome mosaicism defined strains associated with CF. Many mosaic regions were distinguished by their low G + C content; their inclusion of phage related or pyocin genes; or by their linkage to a vgr gene or a tRNA gene. Microarray and phenotypic analysis of sequential isolates from individual patients revealed two deletions of greater than 100 kbp formed during evolution in the lung. The gene loss in these sequential isolates raises the possibility that acquisition of pyomelanin production and loss of pyoverdin uptake each may be of adaptive significance. Further characterization of P. aeruginosa diversity within the airways of individual CF patients may reveal common adaptations, perhaps mediated by gene loss, that suggest new opportunities for therapy.
BACKGROUND: Time series microarray experiments are widely used to study dynamical biological processes. Due to the cost of microarray experiments, and also in some cases the limited availability of biological material, about 80% of microarray time series experiments are short (3-8 time points). Previously short time series gene expression data has been mainly analyzed using more general gene expression analysis tools not designed for the unique challenges and opportunities inherent in short time series gene expression data. RESULTS: We introduce the Short Time-series Expression Miner (STEM) the first software program specifically designed for the analysis of short time series microarray gene expression data. STEM implements unique methods to cluster, compare, and visualize such data. STEM also supports efficient and statistically rigorous biological interpretations of short time series data through its integration with the Gene Ontology. CONCLUSION: The unique algorithms STEM implements to cluster and compare short time series gene expression data combined with its visualization capabilities and integration with the Gene Ontology should make STEM useful in the analysis of data from a significant portion of all microarray studies. STEM is available for download for free to academic and non-profit users at http://www.cs.cmu.edu/~jernst/stem.
BACKGROUND: Gene panels represent a widely used strategy for genetic testing in a vast range of Mendelian disorders. While this approach aids reliable bioinformatic detection of short coding variants, it often fails to detect many larger variants. Recent studies have recommended the adoption of pangenome references (as opposed to linear reference genomes like GRCh38) to augment detection of large variants from targeted sequencing, potentially providing diagnostic laboratories with the possibility to streamline diagnostic work-ups and reduce costs. METHODS: Here, we analyze 1969 cardiomyopathy cases and 1805 controls sequenced with the Illumina Trusight Cardio panel using a pangenome-based workflow (GRAF) and five conventional orthogonal methodologies (GATK HaplotypeCaller, GATK-gCNV, ExomeDepth, Manta and Lumpy-SV) to detect variants ≥ 20 bp in size. RESULTS: Following lab-based variant validation by means of PCR and Sanger sequencing, we show that GRAF conjugates higher precision and recall (F1 score 0.86) compared with other methods (F1 0-0.57) in detecting potentially pathogenic variants ≥ 20 bp from short-read panel data. Results were complemented by a comparison of the tools' performance in detecting ground truth variants on reference sample HG002 from Genome In A Bottle, which confirmed GRAF to outperform other tools also on exome sequencing (F1 0.97 vs. 0-0.94). Notably, in the HG002 benchmark dataset, GRAF also showed slightly improved performance compared to GATK HaplotypeCaller in the identification of small variants (1-19 bp; F1 0.975 vs. 0.968). CONCLUSIONS: Our results indicate that pangenome-based workflows aid improved detection of large variants from targeted sequencing data in the clinical context and suggest that they may contribute to more unified variant detection frameworks for all-size genetic variants in the future.
Microarray technology is associated with many sources of experimental uncertainty. In this review we discuss a number of approaches for dealing with this uncertainty in the processing of data from microarray experiments. We focus here on the analysis of high-density oligonucleotide arrays, such as the popular Affymetrix GeneChip array, which contain multiple probes for each target. This set of probes can be used to determine an estimate for the target concentration and can also be used to determine the experimental uncertainty associated with this measurement. This measurement uncertainty can then be propagated through the downstream analysis using probabilistic methods. We give examples showing how these credibility intervals can be used to help identify differential expression, to combine information from replicated experiments and to improve the performance of principal component analysis.
The transforming growth factor beta (TGF-beta) family of growth modulators play critical roles in tissue development and maintenance. Recent data suggest that individual TGF-beta isoforms (TGF-beta 1, -beta 2 and -beta 3) have overlapping yet distinct biological actions and target cell specificities, both in developing and adult tissues. The TGF-beta 3 isoform was purified to homogeneity from both natural and recombinant sources and characterized by laser desorption mass spectrometry, by protein sequencing, by amino acid analysis and by biological activity. TGF-beta 3 was the major TGF-beta isoform in umbilical cord (230 ng/g), and was physically and biologically indistinguishable from recombinant TGF-beta 3 and from the tumor growth inhibitory (TGI) protein found in umbilical cord. Immunohistochemistry using antipeptide TGF-beta 3 specific antibody showed TGF-beta 3 localization in perivascular smooth muscle.
A cysteine specific cleavage reaction was used for the preparation of biologically active peptides from recombinant fusion proteins. The fusion protein through cysteine was prepared by a recombinant DNA technology and then treated with cyanylating reagents such as 2-nitro-5-thiocyanatobenzoic acid (NTCB) and 1-cyano-4-(dimethylamino) pyridinium tetrafluoroborate (DMAP-CN) to release the desired product. As an example, we have selected a glucagon-like peptide 1 (7-37) (termed insulinotropin). We constructed an expression vector for a fusion protein in which insulinotropin and human basic fibroblast growth factor (hbFGF) mutein (abbreviated as CS 23) are connected by cysteine and then expressed it in Escherichia coli cells. The fusion protein, after refolding, was purified by heparin affinity chromatography, since CS23 has a strong affinity for heparin. The affinity-purified fusion protein was treated with NTCB or DMAP-CN to give crude insulinotropin, which was then purified by reversed phase (rp) high-performance liquid chromatography (HPLC). From various criteria such as amino acid analysis, amino acid sequence and the biological activity, the purified material obtained was found to be methionylated insulinotropin (Met-insulinotropin) with full activity. The specificity and simplicity of the present method make it versatile and convenient for the preparation of biologically active peptides.
A report on the meeting 'Unravelling Nature's Networks: from Microarray and Proteomic Analysis to Systems Biology', Sheffield, UK, 21-22 July 2003.
I describe the current version of the sequence analysis package developed at the MRC Laboratory of Molecular Biology, which has come to be known as the "Staden Package." The package covers most of the standard sequence analysis tasks such as restriction site searching, translation, pattern searching, comparison, gene finding, and secondary structure prediction, and provides powerful tools for DNA sequence determination. Currently the programs are only available for computers running the UNIX operating system. Detailed information about the package is available from our WWW site: http:@www.mrc-lmb.cam.ac.uk/pubseq/.
Many different biological questions are routinely studied using transcriptional profiling on microarrays. A wide range of approaches are available for gleaning insights from the data obtained from such experiments. The appropriate choice of data-analysis technique depends both on the data and on the goals of the experiment. This review summarizes some of the common themes in microarray data analysis, including detection of differential expression, clustering, and predicting sample characteristics. Several approaches to each problem, and their relative merits, are discussed and key areas for additional research highlighted.