PubMed Health⌕ Search

PubMed · 17204461

SNPchip: R classes and methods for SNP array data.

Abstract

UNLABELLED: High-density single nucleotide polymorphism microarrays (SNP chips) provide information on a subject's genome, such as copy number and genotype (heterozygosity/homozygosity) at a SNP. While fluorescence in situ hybridization and karyotyping reveal many abnormalities, SNP chips provide a higher resolution map of the human genome that can be used to detect, e.g., aneuploidies, microdeletions, microduplications and loss of heterozygosity (LOH). As a variety of diseases are linked to such chromosomal abnormalities, SNP chips promise new insights for these diseases by aiding in the discovery of such regions, and may suggest targets for intervention. The R package SNPchip contains classes and methods useful for storing, visualizing and analyzing high density SNP data. Originally developed from the SNPscan web-tool, SNPchip utilizes S4 classes and extends other open source R tools available at Bioconductor. This has numerous advantages, including the ability to build statistical models for SNP-level data that operate on instances of the class, and to communicate with other R packages that add additional functionality. AVAILABILITY: The package is available from the Bioconductor web page at www.bioconductor.org. SUPPLEMENTARY INFORMATION: The supplementary material as described in this article (case studies, installation guidelines and R code) is available from http://biostat.jhsph.edu/~iruczins/publications/sm/

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Robert B Scharpf, Jason C Ting, Jonathan Pevsner, Ingo Ruczinski. 2007-01-04. SNPchip: R classes and methods for SNP array data.. https://doi.org/10.1093/bioinformatics%2Fbtl638

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Penalized Cumulative Probability Model for a Continuous Outcome Subject to Detection Limits.

Mixed-type outcome data occur when the outcome variable's distribution is a mixture of both continuous and discrete ordinal variables. Such mixed-type outcomes are common in biomedical, psychological, and the health sciences, particularly for variables having either a detection or quantitation limit. When interest lies in identifying a combination of genomic features associated with a mixed-type outcome, any method used would require a variable selection strategy for high-dimensional data. Unfortunately, few variable selection methods exist for modeling a mixed-type outcome when the covariate space is high dimensional. This study develops a high-dimensional penalized cumulative probability model (CPM), to allow for the identification of genomic features associated with mixed-type outcome of interest. We demonstrated how such model may be estimated using the iterative penalization procedure-the generalized monotone incremental forward stagewise (GMIFS) algorithm. The Model-X knockoffs procedure was combined with the estimation algorithm to control the false discovery rates (FDR) when performing variable selection. Through extensive simulation studies, our penalized CPM was shown to outperform alternative methods in terms of controlled variable selection performance by achieving high statistical power with the FDR being controlled at the target level. We demonstrate the utility of our method by applying it to predict estimated glomeruli filtration rate (eGFR) in kidney transplant recipients at 24 months post-transplant using baseline gene expression data as predictors. Our CPM model identified five genes associated with this mixed-type outcome which have important links to renal disease, which may provide prognostic guidance for kidney transplantation recipients.

Models, Statistical↗

Using SAS to conduct nonparametric residual bootstrap multilevel modeling with a small number of groups.

In multilevel modeling, researchers often encounter data with a relatively small number of units at the higher levels. As a result, of this and/or non-normality of the residuals, model parameter estimates, particularly the variance components and standard errors of parameter estimates at the group level, may be biased, thus the corresponding statistical inferences may not be trustworthy. This problem can be addressed by using bootstrap methods to estimate the standard errors of the parameter estimates for significance testing. This study illustrates how to use statistical analysis system (SAS) to conduct nonparametric residual bootstrap multilevel modeling. Specific SAS programs for such modeling are provided.

Models, Statistical↗

'A finite size pencil beam for IMRT dose optimization'--a simpler analytical function for the finite size pencil beam kernel.

A simple and finite-termed analytical function for the finite size pencil beam kernel was constructed. The dose cross-profile of a semi-infinite field with field edge at x = 0 can be well fitted by the Boltzmann function. The pencil beam cross-profile of width 2x(0) can be obtained as the difference between two semi-infinite fields shifted by 2x(0). If the profile is centred about x = 0, it can derive from P(x + x(0)) - P(x - x(0)). The penumbra influence can be taken by the penumbra tuning factor f. The parameters A(1), A(2), A(3), A(4), f can be obtained by fitting depth-dose curves and cross-profiles for a set of square fields. The two-dimensional dose distribution F(x, y, x(0), y(0), A(1), A(2), A(3), A(4), f(1), f(2)) of a pencil beam of width (2x(0), 2y(0)) is defined by multiplication of two independent one-dimensional profiles.

Models, Statistical↗