PubMed Health⌕ Search

Biomedical subjects

Andrew F Siegel

Publications and source records attributed to Andrew F Siegel.

10 recordsLinked to original sources

Genetic mapping at 3-kilobase resolution reveals inositol 1,4,5-triphosphate receptor 3 as a risk factor for type 1 diabetes in Sweden.

We mapped the genetic influences for type 1 diabetes (T1D), using 2,360 single-nucleotide polymorphism (SNP) markers in the 4.4-Mb human major histocompatibility complex (MHC) locus and the adjacent 493 kb centromeric to the MHC, initially in a survey of 363 Swedish T1D cases and controls. We confirmed prior studies showing association with T1D in the MHC, most significantly near HLA-DR/DQ. In the region centromeric to the MHC, we identified a peak of association within the inositol 1,4,5-triphosphate receptor 3 gene (ITPR3; formerly IP3R3). The most significant single SNP in this region was at the center of the ITPR3 peak of association (P=1.7 x 10(-4) for the survey study). For validation, we typed an additional 761 Swedish individuals. The P value for association computed from all 1,124 individuals was 1.30 x 10(-6) (recessive odds ratio 2.5; 95% confidence interval [CI] 1.7-3.9). The estimated population-attributable risk of 21.6% (95% CI 10.0%-31.0%) suggests that variation within ITPR3 reflects an important contribution to T1D in Sweden. Two-locus regression analysis supports an influence of ITPR3 variation on T1D that is distinct from that of any MHC class II gene.

Adolescent↗

A third approach to gene prediction suggests thousands of additional human transcribed regions.

The identification and characterization of the complete ensemble of genes is a main goal of deciphering the digital information stored in the human genome. Many algorithms for computational gene prediction have been described, ultimately derived from two basic concepts: (1) modeling gene structure and (2) recognizing sequence similarity. Successful hybrid methods combining these two concepts have also been developed. We present a third orthogonal approach to gene prediction, based on detecting the genomic signatures of transcription, accumulated over evolutionary time. We discuss four algorithms based on this third concept: Greens and CHOWDER, which quantify mutational strand biases caused by transcription-coupled DNA repair, and ROAST and PASTA, which are based on strand-specific selection against polyadenylation signals. We combined these algorithms into an integrated method called FEAST, which we used to predict the location and orientation of thousands of putative transcription units not overlapping known genes. Many of the newly predicted transcriptional units do not appear to code for proteins. The new algorithms are particularly apt at detecting genes with long introns and lacking sequence conservation. They therefore complement existing gene prediction methods and will help identify functional transcripts within many apparent "genomic deserts."

Algorithms↗

The mobile nucleoporin Nup2p and chromatin-bound Prp20p function in endogenous NPC-mediated transcriptional control.

Nuclear pore complexes (NPCs) govern macromolecular transport between the nucleus and cytoplasm and serve as key positional markers within the nucleus. Several protein components of yeast NPCs have been implicated in the epigenetic control of gene expression. Among these, Nup2p is unique as it transiently associates with NPCs and, when artificially tethered to DNA, can prevent the spread of transcriptional activation or repression between flanking genes, a function termed boundary activity. To understand this function of Nup2p, we investigated the interactions of Nup2p with other proteins and with DNA using immunopurifications coupled with mass spectrometry and microarray analyses. These data combined with functional assays of boundary activity and epigenetic variegation suggest that Nup2p and the Ran guanylyl-nucleotide exchange factor, Prp20p, interact at specific chromatin regions and enable the NPC to play an active role in chromatin organization by facilitating the transition of chromatin between activity states.

Active Transport, Cell Nucleus↗

A data integration methodology for systems biology.

Different experimental technologies measure different aspects of a system and to differing depth and breadth. High-throughput assays have inherently high false-positive and false-negative rates. Moreover, each technology includes systematic biases of a different nature. These differences make network reconstruction from multiple data sets difficult and error-prone. Additionally, because of the rapid rate of progress in biotechnology, there is usually no curated exemplar data set from which one might estimate data integration parameters. To address these concerns, we have developed data integration methods that can handle multiple data sets differing in statistical power, type, size, and network coverage without requiring a curated training data set. Our methodology is general in purpose and may be applied to integrate data from any existing and future technologies. Here we outline our methods and then demonstrate their performance by applying them to simulated data sets. The results show that these methods select true-positive data elements much more accurately than classical approaches. In an accompanying companion paper, we demonstrate the applicability of our approach to biological data. We have integrated our methodology into a free open source software package named POINTILLIST.

Informatics↗

A data integration methodology for systems biology: experimental verification.

The integration of data from multiple global assays is essential to understanding dynamic spatiotemporal interactions within cells. In a companion paper, we reported a data integration methodology, designated Pointillist, that can handle multiple data types from technologies with different noise characteristics. Here we demonstrate its application to the integration of 18 data sets relating to galactose utilization in yeast. These data include global changes in mRNA and protein abundance, genome-wide protein-DNA interaction data, database information, and computational predictions of protein-DNA and protein-protein interactions. We divided the integration task to determine three network components: key system elements (genes and proteins), protein-protein interactions, and protein-DNA interactions. Results indicate that the reconstructed network efficiently focuses on and recapitulates the known biology of galactose utilization. It also provided new insights, some of which were verified experimentally. The methodology described here, addresses a critical need across all domains of molecular and cell biology, to effectively integrate large and disparate data sets.

Chromatin Immunoprecipitation↗

Reverse engineering galactose regulation in yeast through model selection.

We examine the application of statistical model selection methods to reverse-engineering the control of galactose utilization in yeast from DNA microarray experiment data. In these experiments, relationships among gene expression values are revealed through modifications of galactose sugar level and genetic perturbations through knockouts. For each gene variable, we select predictors using a variety of methods, taking into account the variance in each measurement. These methods include maximization of log-likelihood with Cp, AIC, and BIC penalties, bootstrap and cross-validation error estimation, and coefficient shrinkage via the Lasso.

Journal Article↗

Control of yeast filamentous-form growth by modules in an integrated molecular network.

On solid growth media with limiting nitrogen source, diploid budding-yeast cells differentiate from the yeast form to a filamentous, adhesive, and invasive form. Genomic profiles of mRNA levels in Saccharomyces cerevisiae yeast-form and filamentous-form cells were compared. Disparate data types, including genes implicated by expression change, filamentation genes known previously through a phenotype, protein-protein interaction data, and protein-metabolite interaction data were integrated as the nodes and edges of a filamentation-network graph. Application of a network-clustering method revealed 47 clusters in the data. The correspondence of the clusters to modules is supported by significant coordinated expression change among cluster co-member genes, and the quantitative identification of collective functions controlling cell properties. The modular abstraction of the filamentation network enables the association of filamentous-form cell properties with the activation or repression of specific biological processes, and suggests hypotheses. A module-derived hypothesis was tested. It was found that the 26S proteasome regulates filamentous-form growth.

Cell Cycle↗

Initial proteome analysis of model microorganism Haemophilus influenzae strain Rd KW20.

The proteome of Haemophilus influenzae strain Rd KW20 was analyzed by liquid chromatography (LC) coupled with ion trap tandem mass spectrometry (MS/MS). This approach does not require a gel electrophoresis step and provides a rapidly developed snapshot of the proteome. In order to gain insight into the central metabolism of H. influenzae, cells were grown microaerobically and anaerobically in a rich medium and soluble and membrane proteins of strain Rd KW20 were proteolyzed with trypsin and directly examined by LC-MS/MS. Several different experimental and computational approaches were utilized to optimize the proteome coverage and to ensure statistically valid protein identification. Approximately 25% of all predicted proteins (open reading frames) of H. influenzae strain Rd KW20 were identified with high confidence, as their component peptides were unambiguously assigned to tandem mass spectra. Approximately 80% of the predicted ribosomal proteins were identified with high confidence, compared to the 33% of the predicted ribosomal proteins detected by previous two-dimensional gel electrophoresis studies. The results obtained in this study are generally consistent with those obtained from computational genome analysis, two-dimensional gel electrophoresis, and whole-genome transposon mutagenesis studies. At least 15 genes originally annotated as conserved hypothetical were found to encode expressed proteins. Two more proteins, previously annotated as predicted coding regions, were detected with high confidence; these proteins also have close homologs in related bacteria. The direct proteomics approach to studying protein expression in vivo reported here is a powerful method that is applicable to proteome analysis of any (micro)organism.

Aerobiosis↗

Spectral analysis of distributions: finding periodic components in eukaryotic enzyme length data.

We introduce the spectral analysis of distributions (SAD), a method for detecting and evaluating possible periodicity in experimental data distributions (histograms) of arbitrary shape. SAD determines whether a given empirical distribution contains a periodic component. We also propose a system of probabilistic mixture distributions to model a histogram consisting of a smooth background together with peaks at periodic intervals, with each peak corresponding to a fixed number of subunits added together. This mixture distribution model allows us to estimate the parameters of the data and to test the statistical significance of the estimated peaks. The analysis is applied to the length distribution of eukaryotic enzymes.

Algorithms↗

Discovering regulatory and signalling circuits in molecular interaction networks.

MOTIVATION: In model organisms such as yeast, large databases of protein-protein and protein-DNA interactions have become an extremely important resource for the study of protein function, evolution, and gene regulatory dynamics. In this paper we demonstrate that by integrating these interactions with widely-available mRNA expression data, it is possible to generate concrete hypotheses for the underlying mechanisms governing the observed changes in gene expression. To perform this integration systematically and at large scale, we introduce an approach for screening a molecular interaction network to identify active subnetworks, i.e., connected regions of the network that show significant changes in expression over particular subsets of conditions. The method we present here combines a rigorous statistical measure for scoring subnetworks with a search algorithm for identifying subnetworks with high score. RESULTS: We evaluated our procedure on a small network of 332 genes and 362 interactions and a large network of 4160 genes containing all 7462 protein-protein and protein-DNA interactions in the yeast public databases. In the case of the small network, we identified five significant subnetworks that covered 41 out of 77 (53%) of all significant changes in expression. Both network analyses returned several top-scoring subnetworks with good correspondence to known regulatory mechanisms in the literature. These results demonstrate how large-scale genomic approaches may be used to uncover signalling and regulatory pathways in a systematic, integrative fashion.

Algorithms↗