PubMed Health⌕ Search

Biomedical subjects

Bobbie-Jo M Webb-Robertson

Publications and source records attributed to Bobbie-Jo M Webb-Robertson.

7 recordsLinked to original sources

A benchmarking study of feature screening approaches across type 1 diabetes omics studies classification settings.

In recent years, high dimensional omics analyses have become more commonplace for investigating complex biological systems. Typically, these studies attempt to identify key biomolecules associated with a particular biological process. Often, machine learning (ML) is used to identify these biomolecules, typically by learning which biomolecules are highly predictive of a treatment, biological outcome, or phenotype. A major challenge of applying ML to high throughput omics is overcoming noise when sample size is limited and unbalanced with respect to tens of thousands of biomolecules measured. Thus, feature selection (the process of reducing the number of predictors) is both a critical and common step in the ML analysis pipeline. While much attention has been given to embedding and wrapping techniques for feature selection in the omics space, filter-based methods for model-free feature selection have appealing theoretical properties. This manuscript evaluates sure screening, a class of filter-based feature selection methods which provide analytical guarantees for true feature set retention. Here, we cover existing feature screening methods based on the sure screening principal, available software, methods to improve feature screening, and contextualize feature screening in the larger discussion of feature selection for omics data analysis. Additionally, a suite of model-free sure screening approaches is applied and compared for several omics biomedical applications in a ML classification context. We identified BcorSIS as the most effective and computationally efficient screening method across various omics datasets, consistently outperforming others like CSIS and DCSIS in runtime.

Humans↗

Genetic variants in the complement system and their potential link in the aetiology of type 1 diabetes.

Type 1 diabetes is an autoimmune disease in which one's own immune system destroys insulin-secreting beta cells in the pancreas. This process results in life-long dependence on exogenous insulin for survival. Both genetic and environmental factors play a role in disease initiation, progression, and ultimate clinical diagnosis of type 1 diabetes. This review will provide background on the natural history of type 1 diabetes and the role of genetic factors involved in the complement system, as several recent studies have identified changes in levels of these proteins as the disease evolves from pre-clinical through to clinically apparent disease.

Humans↗

Integration of Infant Metabolite, Genetic, and Islet Autoimmunity Signatures to Predict Type 1 Diabetes by Age 6 Years.

CONTEXT: Biomarkers that can accurately predict risk of type 1 diabetes (T1D) in genetically predisposed children can facilitate interventions to delay or prevent the disease. OBJECTIVE: This work aimed to determine if a combination of genetic, immunologic, and metabolic features, measured at infancy, can be used to predict the likelihood that a child will develop T1D by age 6 years. METHODS: Newborns with human leukocyte antigen (HLA) typing were enrolled in the prospective birth cohort of The Environmental Determinants of Diabetes in the Young (TEDDY). TEDDY ascertained children in Finland, Germany, Sweden, and the United States. TEDDY children were either from the general population or from families with T1D with an HLA genotype associated with T1D specific to TEDDY eligibility criteria. From the TEDDY cohort there were 702 children will all data sources measured at ages 3, 6, and 9 months, 11.4% of whom progressed to T1D by age 6 years. The main outcome measure was a diagnosis of T1D as diagnosed by American Diabetes Association criteria. RESULTS: Machine learning-based feature selection yielded classifiers based on disparate demographic, immunologic, genetic, and metabolite features. The accuracy of the model using all available data evaluated by the area under a receiver operating characteristic curve is 0.84. Reducing to only 3- and 9-month measurements did not reduce the area under the curve significantly. Metabolomics had the largest value when evaluating the accuracy at a low false-positive rate. CONCLUSION: The metabolite features identified as important for progression to T1D by age 6 years point to altered sugar metabolism in infancy. Integrating this information with classic risk factors improves prediction of the progression to T1D in early childhood.

Autoantibodies↗

Normalization approaches for removing systematic biases associated with mass spectrometry and label-free proteomics.

Central tendency, linear regression, locally weighted regression, and quantile techniques were investigated for normalization of peptide abundance measurements obtained from high-throughput liquid chromatography-Fourier transform ion cyclotron resonance mass spectrometry (LC-FTICR MS). Arbitrary abundances of peptides were obtained from three sample sets, including a standard protein sample, two Deinococcus radiodurans samples taken from different growth phases, and two mouse striatum samples from control and methamphetamine-stressed mice (strain C57BL/6). The selected normalization techniques were evaluated in both the absence and presence of biological variability by estimating extraneous variability prior to and following normalization. Prior to normalization, replicate runs from each sample set were observed to be statistically different, while following normalization replicate runs were no longer statistically different. Although all techniques reduced systematic bias to some degree, assigned ranks among the techniques revealed that for most LC-FTICR-MS analyses linear regression normalization ranked either first or second. However, the lack of a definitive trend among the techniques suggested the need for additional investigation into adapting normalization approaches for label-free proteomics. Nevertheless, this study serves as an important step for evaluating approaches that address systematic biases related to relative quantification and label-free proteomics.

Animals↗

A study of spectral integration and normalization in NMR-based metabonomic analyses.

Metabonomics involves the quantitation of the dynamic multivariate metabolic response of an organism to a pathological event or genetic modification [J.K. Nicholson, J.C. Lindon, E. Holmes, Xenobiotica 29 (1999) 1181-1189]. The analysis of these data involves the use of appropriate multivariate statistical methods; Principal Component Analysis (PCA) has been documented as a valuable pattern recognition technique for 1H NMR spectral data [J.T. Brindle, H. Antti, E. Holmes, G. Tranter, J.K. Nicholson, H.W. Bethell, S. Clarke, P.M. Schofield, E. McKilligin, D.E. Mosedale, D.J. Grainger, Nat. Med. 8 (2002) 1439-1444; B.C. Potts, A.J. Deese, G.J. Stevens, M.D. Reily, D.G. Robertson, J. Theiss, J. Pharm. Biomed. Anal. 26 (2001) 463-476; D.G. Robertson, M.D. Reily, R.E. Sigler, D.F. Wells, D.A. Paterson, T.K. Braden, Toxicol. Sci. 57 (2000) 326-337; L.C. Robosky, D.G. Robertson, J.D. Baker, S. Rane, M.D. Reily, Comb. Chem. High Throughput Screen. 5 (2002) 651-662]. Prior to PCA the raw data is typically processed through four steps; (1) baseline correction, (2) endogenous peak removal, (3) integration over spectral regions to reduce the number of variables, and (4) normalization. The effect of the size of spectral integration regions and normalization has not been well studied. The variability structure and classification accuracy on two distinctly different datasets are assessed via PCA and a leave-one-out cross-validation approach under two normalization approaches and an array of spectral integration regions. The first dataset consists of urine from 15 male Wistar-Hannover rats dosed with ANIT measured at five time points, mimicking drug-induced cholangiolitic hepatitis [D.G. Robertson, M.D. Reily, R.E. Sigler, D.F. Wells, D.A. Paterson, T.K. Braden, Toxicol. Sci. 57 (2000) 326-337; J.P. Shockcor, E. Holmes, Curr. Top. Med. Chem. 2 (2002) 35-51; N.J. Waters, E. Holmes, A. Williams, C.J. Waterfield, R.D. Farrant, J.K. Nicholson, Chem. Res. Toxicol. 14 (2001) 1401-1412]. The second data is serum samples from young male C57BL/6 mice subjected to instillation of pancreatic elastase producing emphysema type symptoms [C. Kuhn, S.Y. Yu, M. Chraplyvy, H.E. Linder, R.M. Senior, Lab. Invest. 34 (1976) 372-380; C. Kuhn, R.M. Senior, Lung 155 (1978) 185-197]. This study indicates that independent of the normalization method the classification accuracy achieved from metabonomic studies is not highly sensitive to the size of the spectral integration region. Additionally, both datasets scaled to mean zero and unity variance (auto-scaled) have higher variability within classification accuracy over spectral integration window widths than data scaled to the total intensity of the spectrum. Of the top 10 latent variables for the ANIT dataset the auto-scale normalization has standard deviations larger than the total-scale in seven cases. In the case of the elastase all standard deviations are larger for the auto-scaling.

Animals↗

Comparison of probability and likelihood models for peptide identification from tandem mass spectrometry data.

We evaluate statistical models used in two-hypothesis tests for identifying peptides from tandem mass spectrometry data. The null hypothesis H(0), that a peptide matches a spectrum by chance, requires information on the probability of by-chance matches between peptide fragments and peaks in the spectrum. Likewise, the alternate hypothesis H(A), that the spectrum is due to a particular peptide, requires probabilities that the peptide fragments would indeed be observed if it was the causative agent. We compare models for these probabilities by determining the identification rates produced by the models using an independent data set. The initial models use different probabilities depending on fragment ion type, but uniform probabilities for each ion type across all of the labile bonds along the backbone. More sophisticated models for probabilities under both H(A) and H(0) are introduced that do not assume uniform probabilities for each ion type. In addition, the performance of these models using a standard likelihood model is compared to an information theory approach derived from the likelihood model. Also, a simple but effective model for incorporating peak intensities is described. Finally, a support-vector machine is used to discriminate between correct and incorrect identifications based on multiple characteristics of the scoring functions. The results are shown to reduce the misidentification rate significantly when compared to a benchmark cross-correlation based approach.

Databases, Protein↗

Enabling proteomics discovery through visual analysis. The peptide permutation and protein prediction tool.

Proteins play a key role in cellular processes, making proteomics central to understanding systems biology. MS techniques provide a means to observe entire proteomes at a global level. Yet, high-throughput MS proteomics techniques generate data faster than it can currently be analyzed. The success of proteomics depends on high-throughput experimental techniques coupled with sophisticated visual analysis and data-mining methods. Visual analysis has been applied successfully in a number of fields plagued with huge, complex data sets and will likely be an important tool in proteomics discovery. PQuad, a novel visualization of MS proteomics data, provides powerful analysis capabilities that support a number of proteomic data applications. In particular, PQuad supports differential proteomics by simplifying the comparison of peptide sets from different experimental conditions as well as different protein identification or confidence scoring techniques. Finally, PQuad supports data validation and quality control by providing a variety of resolutions for huge amounts of data to reveal errors undetected by other methods.

Algorithms↗