PubMed HealthSearch

Biomedical subjects

Christopher Carpenter

Publications and source records attributed to Christopher Carpenter.

3 recordsLinked to original sources

scBaseCount: An AI agent-curated, standardized, auto-updated single-cell data repository.

Single-cell RNA sequencing has transformed cell biology by enabling precise transcriptomic measurements of individual cells. The Sequence Read Archive (SRA) is the largest public repository of sequencing reads, yet much of it remains underutilized due to unstandardized metadata. Here, we introduce scBaseCount, a database that leverages an AI agent to automate discovery and metadata extraction and standardize data processing. Built by mining all 10x Genomics datasets, scBaseCount is the largest public repository of single-cell gene expression data, comprising over 502 million cells across 27 organisms and 75 tissues. It offers an unbiased view of the data landscape within the SRA and enables the training of more performant computational models through access to broader phenotypic diversity. Uniform processing enables measurement of both intronic and exonic reads and non-coding gene expression and improves alignment across experiments. Moreover, scBaseCount provides a blueprint for how AI can be leveraged to autonomously curate biological data repositories.

Single-Cell Analysis

Tahoe-100M: Mapping drug-induced molecular phenotypes at single-cell resolution.

We present Tahoe-100M, a giga-scale single-cell perturbation atlas comprising 100 million transcriptomes from 50 diverse cancer cell lines treated with 1,100 drug-dose conditions. This parallel profiling of thousands of perturbations at single-cell resolution with minimal batch effects is enabled by the Mosaic platform, which multiplexes genetically distinct cell models into balanced "cell villages." Beyond cataloging transcriptomic shifts, Tahoe-100M systematically quantifies cellular phenotypes, including proliferation, cytotoxicity, lineage-specific vulnerabilities, and cell-cycle changes. It captures population-level transcriptomic heterogeneity, characterizing whether drug responses drive cells toward divergent fates or convergent states. Pathway-based signatures define drug-induced expression programs, classify mechanisms of action, reveal off-target activities, and expose adaptive stress responses associated with resistance. By unifying cellular and molecular readouts, this broadly applicable perturbation atlas advances our ability to model gene regulation, drug response, and network dynamics. Its public release enables the training of AI frameworks to advance predictive models of cell behavior.

Humans

Predicting cellular responses to perturbation across diverse contexts with State.

While machine learning models offer potential for predicting transcriptomic effects of perturbation, they currently struggle to generalize across cellular contexts. Here, we introduce State, a machine learning model that predicts perturbation effects while accounting for cellular heterogeneity within and across experiments. State is trained using single-cell gene expression data to predict perturbation effects across sets of cells. State improved discrimination of effects on large datasets by more than 30% and identified differentially expressed genes across genetic, signaling, and chemical perturbations with significantly improved accuracy compared with baselines. Its cell embeddings trained on observational data from 167 million cells enable the identification of strong perturbations in cellular contexts where no perturbations were observed during training. We further introduce Cell-Eval, a comprehensive evaluation framework that can be used to evaluate future models. Overall, the performance and flexibility of State set the stage for scaling the development of AI models of cell state.

Machine Learning