PubMed Health⌕ Search

Biomedical subjects

Soumyaroop Bhattacharya

Publications and source records attributed to Soumyaroop Bhattacharya.

6 recordsLinked to original sources

Epithelial cell PPAR[gamma] contributes to normal lung maturation.

Peroxisome proliferator-activated receptor (PPAR)-gamma is a member of the nuclear hormone receptor superfamily that can promote cellular differentiation and organ development. PPARgamma expression has been reported in a number of pulmonary cell types, including inflammatory, mesenchymal, and epithelial cells. We find that PPARgamma is prominently expressed in the airway epithelium in the mouse lung. In an effort to define the physiological role of PPARgamma within the lung, we have ablated PPARgamma using a novel line of mice capable of specifically targeting the airway epithelium. Airway epithelial cell PPARgamma-targeted mice display enlarged airspaces resulting from insufficient postnatal lung maturation. The increase in airspace size is accompanied by alterations in lung physiology, including increased lung volumes and decreased tissue resistance. Genome-wide expression profiling reveals a reduction in structural extracellular matrix (ECM) gene expression in conditionally targeted mice, suggesting a disruption in epithelial-mesenchymal interactions necessary for the establishment of normal lung structure. Expression profiling of airway epithelial cells isolated from conditionally targeted mice indicates PPARgamma regulates genes encoding known PPARgamma targets, additional lipid metabolism enzymes, and markers of cellular differentiation. These data reveal airway epithelial cell PPARgamma is necessary for normal lung structure and function.

Animals↗

Transformation of expression intensities across generations of Affymetrix microarrays using sequence matching and regression modeling.

The utility of previously generated microarray data is severely limited owing to small study size, leading to under-powered analysis, and failure of replication. Multiplicity of platforms and various sources of systematic noise limit the ability to compile existing data from similar studies. We present a model for transformation of data across different generations of Affymetrix arrays, developed using previously published datasets describing technical replicates performed with two generations of arrays. The transformation is based upon a probe set-specific regression model, generated from replicate measurements across platforms, performed using correlation coefficients. The model, when applied to the expression intensities of 5069 shared, sequence-matched probe sets in three different generations of Affymetrix Human oligonucleotide arrays, showed significant improvement in inter generation correlations between sample-wide means and individual probe set pairs. The approach was further validated by an observed reduction in Euclidean distance between signal intensities across generations for the predicted values. Finally, application of the model to independent, but related datasets resulted in improved clustering of samples based upon their biological, as opposed to technical, attributes. Our results suggest that this transformation method is a valuable tool for integrating microarray datasets from different generations of arrays.

Algorithms↗

A classification-based machine learning approach for the analysis of genome-wide expression data.

Three important areas of data analysis for global gene expression analysis are class discovery, class prediction, and finding dysregulated genes (biomarkers). The clinical application of microarray data will require marker genes whose expression patterns are sufficiently well understood to allow accurate predictions on disease subclass membership. Commonly used methods of analysis include hierarchical clustering algorithms, t-, F-, and Z-tests, and machine learning approaches. We describe an approach called the maximum difference subset (MDSS) algorithm that combines classification algorithms, classical statistics, and elements of machine learning and provides a coherent framework. By integrating prediction accuracy, the MDSS algorithm learns the critical threshold of statistical significance (the alpha or P-value), eliminating the arbitrariness of setting a threshold of statistical significance and minimizing the effect of the normality assumptions. To reduce the false positive rate and to increase external validity of the predictive gene set, a jackknife step is used. This step identifies and removes genes in the initial MDSS with low combined predictive utility. The overall MDSS provides a prediction that is less dependent on an arbitrary study design (sample inclusion or exclusion) and should thus have high external validity. We demonstrate that this approach, unlike other published methods, identifies biomarkers capable of predicting the outcome of anthracycline-cytarabine chemotherapy in cases of acute myeloid leukemia. By incorporating two criteria-statistical significance and predictive utility-the approach learns the significance level relevant for a given data set. The MDSS approach can be used with any test and classifier operator pair.

Acute Disease↗

Overcoming confounded controls in the analysis of gene expression data from microarray experiments.

A potential limitation of data from microarray experiments exists when improper control samples are used. In cancer research, comparisons of tumour expression profiles to those from normal samples is challenging due to tissue heterogeneity (mixed cell populations). A specific example exists in a published colon cancer dataset, in which tissue heterogeneity was reported among the normal samples. In this paper, we show how to overcome or avoid the problem of using normal samples that do not derive from the same tissue of origin as the tumour. We advocate an exploratory unsupervised bootstrap analysis that can reveal unexpected and undesired, but strongly supported, clusters of samples that reflect tissue differences instead of tumour versus normal differences. All of the algorithms used in the analysis, including the maximum difference subset algorithm, unsupervised bootstrap analysis, pooled variance t-test for finding differentially expressed genes and the jackknife to reduce false positives, are incorporated into our online Gene Expression Data Analyzer ( http:// bioinformatics.upmc.edu/GE2/GEDA.html ).

Algorithms↗