PubMed Health⌕ Search

Biomedical subjects

Ingo Ruczinski

Publications and source records attributed to Ingo Ruczinski.

16 recordsLinked to original sources

Trio-based GWAS reveals loci associated with different forms of isolated cleft lip.

Orofacial clefts (OFCs) are the most common craniofacial birth defect and comprise a diverse group of traits with complex and heterogeneous etiologies. Genetic studies of OFCs typically approach this diversity by stratifying cases into broad diagnostic classes, including cleft lip (CL), cleft palate (CP), and cleft lip with palate (CLP). Although this strategy has yielded important insights into OFC risk, it ignores the phenotypic heterogeneity within each subtype. CL exhibits marked phenotypic variability, involving differences in alveolar involvement, laterality, and sidedness that may reflect distinct etiologies. Given this phenotypic diversity within CL, we assembled a multi-ancestry cohort of 837 nonsyndromic CL case-parent trios with whole-genome sequencing and detailed phenotyping. We performed genome-wide association scans (GWAS) via transmission disequilibrium tests for CL overall and for 14 CL subtypes defined by involvement of the alveolus (with and without), laterality (uni- and bilateral), and sidedness (left and right). We identified four genome-wide significant loci. Two loci, IRF6 and 8q24.21, were both detected in the overall CL GWAS. PLCB1/PLCB4 and MAFB were detected in GWASs of alveolar cleft involvement and CL left sidedness, respectively. These subtype-specific associations were followed by case-only comparisons that reflect the presence or absence of alveolus cleft or left-sided bias of CL to confirm the specificity of the association signal to the particular subtype. Our results provide evidence of within-class CL subtype-specific genetic links for loci previously discussed in the context of primary OFC classes and demonstrate the value of granular OFC subtype characterization to capture trait-specific associations.

Alveolus Cleft↗

Fundamentals of FAIR biomedical data analyses in the cloud using custom pipelines.

As the biomedical data ecosystem increasingly embraces the findable, accessible, interoperable, and reusable (FAIR) data principles to publish multimodal datasets to the cloud, opportunities for cloud-based research continue to expand. Besides the potential for accelerated and diverse biomedical discovery that comes from a harmonized data ecosystem, the cloud also presents a shift away from the standard practice of duplicating data to computational clusters or local computers for analysis. However, despite these benefits, researcher migration to the cloud has lagged, in part due to insufficient educational resources to train biomedical scientists on cloud infrastructure. There exists a conceptual lack especially around the crafting of custom analytic pipelines that require software not pre-installed by cloud analysis platforms. We here present three fundamental concepts necessary for custom pipeline creation in the cloud. These overarching concepts are workflow and cloud provider agnostic, extending the utility of this education to serve as a foundation for any computational analysis running any dataset in any biomedical cloud platform. We illustrate these concepts using one of our own custom analyses, a study using the case-parent trio design to detect sex-specific genetic effects on orofacial cleft (OFC) risk, which we crafted in the biomedical cloud analysis platform CAVATICA.

Cloud Computing↗

SNPchip: R classes and methods for SNP array data.

UNLABELLED: High-density single nucleotide polymorphism microarrays (SNP chips) provide information on a subject's genome, such as copy number and genotype (heterozygosity/homozygosity) at a SNP. While fluorescence in situ hybridization and karyotyping reveal many abnormalities, SNP chips provide a higher resolution map of the human genome that can be used to detect, e.g., aneuploidies, microdeletions, microduplications and loss of heterozygosity (LOH). As a variety of diseases are linked to such chromosomal abnormalities, SNP chips promise new insights for these diseases by aiding in the discovery of such regions, and may suggest targets for intervention. The R package SNPchip contains classes and methods useful for storing, visualizing and analyzing high density SNP data. Originally developed from the SNPscan web-tool, SNPchip utilizes S4 classes and extends other open source R tools available at Bioconductor. This has numerous advantages, including the ability to build statistical models for SNP-level data that operate on instances of the class, and to communicate with other R packages that add additional functionality. AVAILABILITY: The package is available from the Bioconductor web page at www.bioconductor.org. SUPPLEMENTARY INFORMATION: The supplementary material as described in this article (case studies, installation guidelines and R code) is available from http://biostat.jhsph.edu/~iruczins/publications/sm/

Models, Statistical↗

Analysis and visualization of chromosomal abnormalities in SNP data with SNPscan.

BACKGROUND: A variety of diseases are caused by chromosomal abnormalities such as aneuploidies (having an abnormal number of chromosomes), microdeletions, microduplications, and uniparental disomy. High density single nucleotide polymorphism (SNP) microarrays provide information on chromosomal copy number changes, as well as genotype (heterozygosity and homozygosity). SNP array studies generate multiple types of data for each SNP site, some with more than 100,000 SNPs represented on each array. The identification of different classes of anomalies within SNP data has been challenging. RESULTS: We have developed SNPscan, a web-accessible tool to analyze and visualize high density SNP data. It enables researchers (1) to visually and quantitatively assess the quality of user-generated SNP data relative to a benchmark data set derived from a control population, (2) to display SNP intensity and allelic call data in order to detect chromosomal copy number anomalies (duplications and deletions), (3) to display uniparental isodisomy based on loss of heterozygosity (LOH) across genomic regions, (4) to compare paired samples (e.g. tumor and normal), and (5) to generate a file type for viewing SNP data in the University of California, Santa Cruz (UCSC) Human Genome Browser. SNPscan accepts data exported from Affymetrix Copy Number Analysis Tool as its input. We validated SNPscan using data generated from patients with known deletions, duplications, and uniparental disomy. We also inspected previously generated SNP data from 90 apparently normal individuals from the Centre d'Etude du Polymorphisme Humain (CEPH) collection, and identified three cases of uniparental isodisomy, four females having an apparently mosaic X chromosome, two mislabelled SNP data sets, and one microdeletion on chromosome 2 with mosaicism from an apparently normal female. These previously unrecognized abnormalities were all detected using SNPscan. The microdeletion was independently confirmed by fluorescence in situ hybridization, and a region of homozygosity in a UPD case was confirmed by sequencing of genomic DNA. CONCLUSION: SNPscan is useful to identify chromosomal abnormalities based on SNP intensity (such as chromosomal copy number changes) and heterozygosity data (including regions of LOH and some cases of UPD). The program and source code are available at the SNPscan website http://pevsnerlab.kennedykrieger.org/snpscan.htm.

Base Sequence↗

Imputation methods to improve inference in SNP association studies.

Missing single nucleotide polymorphisms (SNPs) are quite common in genetic association studies. Subjects with missing SNPs are often discarded in analyses, which may seriously undermine the inference of SNP-disease association. In this article, we develop two haplotype-based imputation approaches and one tree-based imputation approach for association studies. The emphasis is to evaluate the impact of imputation on parameter estimation, compared to the standard practice of ignoring missing data. Haplotype-based approaches build on haplotype reconstruction by the expectation-maximization (EM) algorithm or a weighted EM (WEM) algorithm, depending on whether case-control status is taken into account. The tree-based approach uses a Gibbs sampler to iteratively sample from a full conditional distribution, which is obtained from the classification and regression tree (CART) algorithm. We employ a standard multiple imputation procedure to account for the uncertainty of imputation. We apply the methods to simulated data as well as a case-control study on developmental dyslexia. Our results suggest that imputation generally improves efficiency over the standard practice of ignoring missing data. The tree-based approach performs comparably well as haplotype-based approaches, but the former has a computational advantage. The WEM approach yields the smallest bias at a price of increased variance.

Algorithms↗

On the precision of experimentally determined protein folding rates and phi-values.

Phi-values, a relatively direct probe of transition-state structure, are an important benchmark in both experimental and theoretical studies of protein folding. Recently, however, significant controversy has emerged regarding the reliability with which phi-values can be determined experimentally: Because phi is a ratio of differences between experimental observables it is extremely sensitive to errors in those observations when the differences are small. Here we address this issue directly by performing blind, replicate measurements in three laboratories. By monitoring within- and between-laboratory variability, we have determined the precision with which folding rates and phi-values are measured using generally accepted laboratory practices and under conditions typical of our laboratories. We find that, unless the change in free energy associated with the probing mutation is quite large, the precision of phi-values is relatively poor when determined using rates extrapolated to the absence of denaturant. In contrast, when we employ rates estimated at nonzero denaturant concentrations or assume that the slopes of the chevron arms (mf and mu) are invariant upon mutation, the precision of our estimates of phi is significantly improved. Nevertheless, the reproducibility we thus obtain still compares poorly with the confidence intervals typically reported in the literature. This discrepancy appears to arise due to differences in how precision is calculated, the dependence of precision on the number of data points employed in defining a chevron, and interlaboratory sources of variability that may have been largely ignored in the prior literature.

Fluorometry↗

Methods for the accurate estimation of confidence intervals on protein folding phi-values.

Phi-values provide an important benchmark for the comparison of experimental protein folding studies to computer simulations and theories of the folding process. Despite the growing importance of phi measurements, however, formulas to quantify the precision with which phi is measured have seen little significant discussion. Moreover, a commonly employed method for the determination of standard errors on phi estimates assumes that estimates of the changes in free energy of the transition and folded states are independent. Here we demonstrate that this assumption is usually incorrect and that this typically leads to the underestimation of phi precision. We derive an analytical expression for the precision of phi estimates (assuming linear chevron behavior) that explicitly takes this dependence into account. We also describe an alternative method that implicitly corrects for the effect. By simulating experimental chevron data, we show that both methods accurately estimate phi confidence intervals. We also explore the effects of the commonly employed techniques of calculating phi from kinetics estimated at non-zero denaturant concentrations and via the assumption of parallel chevron arms. We find that these approaches can produce significantly different estimates for phi (again, even for truly linear chevron behavior), indicating that they are not equivalent, interchangeable measures of transition state structure. Lastly, we describe a Web-based implementation of the above algorithms for general use by the protein folding community.

Algorithms↗

Associations of classic Kaposi sarcoma with common variants in genes that modulate host immunity.

Classic Kaposi sarcoma (CKS) is an inflammatory-mediated neoplasm primarily caused by Kaposi sarcoma-associated herpesvirus (KSHV). Kaposi sarcoma lesions are characterized, in part, by the presence of proinflammatory cytokines and growth factors thought to regulate KSHV replication and CKS pathogenesis. Using genomic DNA extracted from 133 CKS cases and 172 KSHV-latent nuclear antigen-positive, population-based controls in Italy without HIV infection, we examined the risk of CKS associated with 28 common genetic variants in 14 immune-modulating genes. Haplotypes were estimated for IL1A, IL1B, IL4, IL8, IL8RB, IL10, IL12A, IL13, and TNF. Compared with controls, CKS risk was decreased with 1235T/-1010G-containing diplotypes of IL8RB (odds ratio, 0.49; 95% confidence interval, 0.30-0.78; P = 0.003), whereas risk was increased with diplotypes of IL13 containing the promoter region variant 98A (rs20541, alias +130; odds ratio, 1.88; 95% confidence interval, 1.15-3.08; P = 0.01) when considered in multivariate analysis. Risk estimates did not substantially vary by age, sex, incident disease, or disease burden. Our data provide preliminary evidence for variants in immune-modulating genes that could influence the risk of CKS. Among KSHV-seropositive Italians, CKS risk was associated with diplotypes of IL8RB and IL13, supporting laboratory evidence of immune-mediated pathogenesis.

Adult↗

Primary and secondary transcriptional effects in the developing human Down syndrome brain and heart.

BACKGROUND: Down syndrome, caused by trisomic chromosome 21, is the leading genetic cause of mental retardation. Recent studies demonstrated that dosage-dependent increases in chromosome 21 gene expression occur in trisomy 21. However, it is unclear whether the entire transcriptome is disrupted, or whether there is a more restricted increase in the expression of those genes assigned to chromosome 21. Also, the statistical significance of differentially expressed genes in human Down syndrome tissues has not been reported. RESULTS: We measured levels of transcripts in human fetal cerebellum and heart tissues using DNA microarrays and demonstrated a dosage-dependent increase in transcription across different tissue/cell types as a result of trisomy 21. Moreover, by having a larger sample size, combining the data from four different tissue and cell types, and using an ANOVA approach, we identified individual genes with significantly altered expression in trisomy 21, some of which showed this dysregulation in a tissue-specific manner. We validated our microarray data by over 5,600 quantitative real-time PCRs on 28 genes assigned to chromosome 21 and other chromosomes. Gene expression values from chromosome 21, but not from other chromosomes, accurately classified trisomy 21 from euploid samples. Our data also indicated functional groups that might be perturbed in trisomy 21. CONCLUSIONS: In Down syndrome, there is a primary transcriptional effect of disruption of chromosome 21 gene expression, without a pervasive secondary effect on the remaining transcriptome. The identification of dysregulated genes and pathways suggests molecular changes that may underlie the Down syndrome phenotypes.

Astrocytes↗

Site-specific dimensions across a highly denatured protein; a single molecule study.

Do highly denatured proteins adopt random coil configurations? Here, we address this question by measuring residue-to-residue separations across the denatured FynSH3 domain. Using single-molecule Forster resonance energy transfer techniques, we have collected transfer efficiency probability distributions for dye-labeled, denatured protein. Applying maximum likelihood analysis to the interpretation of these distributions, we have determined the through-space distance between five residue pairs in the protein's guanidine hydrochloride-unfolded and trifluoroethanol-unfolded states. We find that, while the dimensions of the guanidine hydrochloride -unfolded molecule generally coincide with the dimensions predicted for a random coil ensemble, potentially statistically significant deviations from random coil behavior are also evident. These small, site-specific deviations may provide a means of reconciling earlier, scattering-based evidence for the random coil nature of the unfolded state with more site-specific spectroscopic evidence suggesting residual structure. We have also studied the unfolded ensemble populated in 50% trifluoroethanol, a denaturant that induces a highly helical unfolded state. We find that the size and shape of the unfolded ensemble under these conditions is effectively indistinguishable from that populated in guanidinium hydrochloride solutions, suggesting that the gross structure of the denatured state is, perhaps surprisingly, independent of the chemistry of the cosolvent.

Escherichia coli Proteins↗

Protein folding: defining a "standard" set of experimental conditions and a preliminary kinetic data set of two-state proteins.

Recent years have seen the publication of both empirical and theoretical relationships predicting the rates with which proteins fold. Our ability to test and refine these relationships has been limited, however, by a variety of difficulties associated with the comparison of folding and unfolding rates, thermodynamics, and structure across diverse sets of proteins. These difficulties include the wide, potentially confounding range of experimental conditions and methods employed to date and the difficulty of obtaining correct and complete sequence and structural details for the characterized constructs. The lack of a single approach to data analysis and error estimation, or even of a common set of units and reporting standards, further hinders comparative studies of folding. In an effort to overcome these problems, we define here a "consensus" set of experimental conditions (25 degrees C at pH 7.0, 50 mM buffer), data analysis methods, and data reporting standards that we hope will provide a benchmark for experimental studies. We take the first step in this initiative by describing the folding kinetics of 30 apparently two-state proteins or protein domains under the consensus conditions. The goal of our efforts is to set uniform standards for the experimental community and to initiate an accumulating, self-consistent data set that will aid ongoing efforts to understand the folding process.

Biochemistry↗

Identifying interacting SNPs using Monte Carlo logic regression.

Interactions are frequently at the center of interest in single-nucleotide polymorphism (SNP) association studies. When interacting SNPs are in the same gene or in genes that are close in sequence, such interactions may suggest which haplotypes are associated with a disease. Interactions between unrelated SNPs may suggest genetic pathways. Unfortunately, data sets are often still too small to definitively determine whether interactions between SNPs occur. Also, competing sets of interactions could often be of equal interest. Here we propose Monte Carlo logic regression, an exploratory tool that combines Markov chain Monte Carlo and logic regression, an adaptive regression methodology that attempts to construct predictors as Boolean combinations of binary covariates such as SNPs. The goal of Monte Carlo logic regression is to generate a collection of (interactions of) SNPs that may be associated with a disease outcome, and that warrant further investigation. As such, the models that are fitted in the Markov chain are not combined into a single model, as is often done in Bayesian model averaging procedures. Instead, the most frequently occurring patterns in these models are tabulated. The method is applied to a study of heart disease with 779 participants and 89 SNPs. A simulation study is carried out to investigate the performance of the Monte Carlo logic regression approach.

Haplotypes↗

Random-coil behavior and the dimensions of chemically unfolded proteins.

Spectroscopic studies have identified a number of proteins that appear to retain significant residual structure under even strongly denaturing conditions. Intrinsic viscosity, hydrodynamic radii, and small-angle x-ray scattering studies, in contrast, indicate that the dimensions of most chemically denatured proteins scale with polypeptide length by means of the power-law relationship expected for random-coil behavior. Here we further explore this discrepancy by expanding the length range of characterized denatured-state radii of gyration (R(G)) and by reexamining proteins that reportedly do not fit the expected dimensional scaling. We find that only 2 of 28 crosslink-free, prosthetic-group-free, chemically denatured polypeptides deviate significantly from a power-law relationship with polymer length. The R(G) of the remaining 26 polypeptides, which range from 16 to 549 residues, are well fitted (r(2) = 0.988) by a power-law relationship with a best-fit exponent, 0.598 +/- 0.028, coinciding closely with the 0.588 predicted for an excluded volume random coil. Therefore, it appears that the mean dimensions of the large majority of chemically denatured proteins are effectively indistinguishable from the mean dimensions of a random-coil ensemble.

Guanidine↗

Distributions of beta sheets in proteins with application to structure prediction.

We recently developed the Rosetta algorithm for ab initio protein structure prediction, which generates protein structures from fragment libraries using simulated annealing. The scoring function in this algorithm favors the assembly of strands into sheets. However, it does not discriminate between different sheet motifs. After generating many structures using Rosetta, we found that the folding algorithm predominantly generates very local structures. We surveyed the distribution of beta-sheet motifs with two edge strands (open sheets) in a large set of non-homologous proteins. We investigated how much of that distribution can be accounted for by rules previously published in the literature, and developed a filter and a scoring method that enables us to improve protein structure prediction for beta-sheet proteins. Proteins 2002;48:85-97.

Algorithms↗

Residues participating in the protein folding nucleus do not exhibit preferential evolutionary conservation.

To what extent does natural selection act to optimize the details of protein folding kinetics? In an effort to address this question, the relationship between an amino acid's evolutionary conservation and its role in protein folding kinetics has been investigated intensively. Despite this effort, no consensus has been reached regarding the degree to which residues involved in native-like transition state structure (the folding nucleus) are conserved. Here we report the results of an exhaustive, systematic study of sequence conservation among residues known to participate in the experimentally (Phi-value) defined folding nuclei of all of the appropriately characterized proteins reported to date. We observe no significant evidence that these residues exhibit any anomalous sequence conservation. We do observe, however, a significant bias in the existing kinetic data: the mean sequence conservation of the residues that have been the subject of kinetic characterization is greater than the mean sequence conservation of all residues in 13 of 14 proteins studied. This systematic experimental bias gives rise to the previous observation that the median conservation of residues reported to participate in the folding nucleus is greater than the median conservation of all of the residues in a protein. When this bias is corrected (by comparing, for example, the conservation of residues known to participate in the folding nucleus with that of other, kinetically characterized residues) the previously reported preferential conservation is effectively eliminated. In contrast to well-established theoretical expectations, both poorly and highly conserved residues are apparently equally likely to participate in the protein-folding nucleus.

Bias↗

Contact order and ab initio protein structure prediction.

Although much of the motivation for experimental studies of protein folding is to obtain insights for improving protein structure prediction, there has been relatively little connection between experimental protein folding studies and computational structural prediction work in recent years. In the present study, we show that the relationship between protein folding rates and the contact order (CO) of the native structure has implications for ab initio protein structure prediction. Rosetta ab initio folding simulations produce a dearth of high CO structures and an excess of low CO structures, as expected if the computer simulations mimic to some extent the actual folding process. Consistent with this, the majority of failures in ab initio prediction in the CASP4 (critical assessment of structure prediction) experiment involved high CO structures likely to fold much more slowly than the lower CO structures for which reasonable predictions were made. This bias against high CO structures can be partially alleviated by performing large numbers of additional simulations, selecting out the higher CO structures, and eliminating the very low CO structures; this leads to a modest improvement in prediction quality. More significant improvements in predictions for proteins with complex topologies may be possible following significant increases in high-performance computing power, which will be required for thoroughly sampling high CO conformations (high CO proteins can take six orders of magnitude longer to fold than low CO proteins). Importantly for such a strategy, simulations performed for high CO structures converge much less strongly than those for low CO structures, and hence, lack of simulation convergence can indicate the need for improved sampling of high CO conformations. The parallels between Rosetta simulations and folding in vivo may extend to misfolding: The very low CO structures that accumulate in Rosetta simulations consist primarily of local up-down beta-sheets that may resemble precursors to amyloid formation.

Algorithms↗