PubMed Health⌕ Search

Biomedical subjects

Atul J Butte

Publications and source records attributed to Atul J Butte.

16 recordsLinked to original sources

Integrating expert knowledge into large language models improves performance for psychiatric reasoning and diagnosis.

BACKGROUND AND METHODS: The authors sought to evaluate the performance of common large language models (LLMs) in psychiatric diagnosis, and the impact of integrating expert-derived reasoning on their performance. Clinical case vignettes and associated diagnoses were retrieved from the DSM-5-TR Clinical Cases book. Diagnostic decision trees were retrieved from the DSM-5-TR Handbook of Differential Diagnosis and refined for LLM use. Three LLMs were prompted to provide diagnosis candidates for the vignettes either by directly prompting or using the decision trees. These candidates and diagnostic categories were compared against the correct diagnoses. The positive predictive value (PPV), sensitivity, and F1 statistic were used to measure performance. RESULTS: When directly prompted to predict diagnoses, the best LLM by F1 statistic (gpt-4o) had sensitivity of 76.7 % and PPV of 40.4 %. When making use of the refined decision trees, PPV was significantly increased (65.3 %) without a significant reduction in sensitivity (70.9 %). Across all experiments, the use of the decision trees statistically significantly increased the PPV, significantly increased the F1 statistic in 5/6 experiments, and significantly reduced sensitivity in 4/6 experiments. DISCUSSION: When used to predict psychiatric diagnoses from case vignettes, direct prompting of the LLMs yielded most true positive diagnoses but had significant overdiagnosis. Integrating expert-derived reasoning into the process using decision trees improved LLM performance (as measured by F1 statistic), primarily by suppressing overdiagnosis with a lower-magnitude negative impact on sensitivity. This suggests that the integration of clinical expert-derived reasoning could improve the performance of LLM-based tools in the behavioral health setting.

Humans↗

Creation and implications of a phenome-genome network.

Although gene and protein measurements are increasing in quantity and comprehensiveness, they do not characterize a sample's entire phenotype in an environmental or experimental context. Here we comprehensively consider associations between components of phenotype, genotype and environment to identify genes that may govern phenotype and responses to the environment. Context from the annotations of gene expression data sets in the Gene Expression Omnibus is represented using the Unified Medical Language System, a compendium of biomedical vocabularies with nearly 1-million concepts. After showing how data sets can be clustered by annotative concepts, we find a network of relations between phenotypic, disease, environmental and experimental contexts as well as genes with differential expression associated with these concepts. We identify novel genes related to concepts such as aging. Comprehensively identifying genes related to phenotype and environment is a step toward the Human Phenome Project.

Aging↗

Multiplexed protein array platforms for analysis of autoimmune diseases.

Several proteomics platforms have emerged in the past decade that show great promise for filling in the many gaps that remain from earlier studies of the genome and from the sequencing of the human genome itself. This review describes applications of proteomics technologies to the study of autoimmune diseases. We focus largely on biased technology platforms that are capable of analyzing a large panel of known analytes, as opposed to techniques such as two-dimensional gel electrophoresis (2DIGE) or mass spectroscopy that represent unbiased approaches (as reviewed in 1). At present, the main analytes that can be systematically studied in autoimmunity include autoantibodies, cytokines and chemokines, components of signaling pathways, and cell-surface receptors. We review the most commonly used platforms for such studies, citing important discoveries and limitations that exist. We conclude by reviewing advances in biomedical informatics that will eventually allow the human proteome to be deciphered.

Autoantibodies↗

Systematic survey reveals general applicability of "guilt-by-association" within gene coexpression networks.

BACKGROUND: Biological processes are carried out by coordinated modules of interacting molecules. As clustering methods demonstrate that genes with similar expression display increased likelihood of being associated with a common functional module, networks of coexpressed genes provide one framework for assigning gene function. This has informed the guilt-by-association (GBA) heuristic, widely invoked in functional genomics. Yet although the idea of GBA is accepted, the breadth of GBA applicability is uncertain. RESULTS: We developed methods to systematically explore the breadth of GBA across a large and varied corpus of expression data to answer the following question: To what extent is the GBA heuristic broadly applicable to the transcriptome and conversely how broadly is GBA captured by a priori knowledge represented in the Gene Ontology (GO)? Our study provides an investigation of the functional organization of five coexpression networks using data from three mammalian organisms. Our method calculates a probabilistic score between each gene and each Gene Ontology category that reflects coexpression enrichment of a GO module. For each GO category we use Receiver Operating Curves to assess whether these probabilistic scores reflect GBA. This methodology applied to five different coexpression networks demonstrates that the signature of guilt-by-association is ubiquitous and reproducible and that the GBA heuristic is broadly applicable across the population of nine hundred Gene Ontology categories. We also demonstrate the existence of highly reproducible patterns of coexpression between some pairs of GO categories. CONCLUSION: We conclude that GBA has universal value and that transcriptional control may be more modular than previously realized. Our analyses also suggest that methodologies combining coexpression measurements across multiple genes in a biologically-defined module can aid in characterizing gene function or in characterizing whether pairs of functions operate together.

Animals↗

Prediction of preadipocyte differentiation by gene expression reveals role of insulin receptor substrates and necdin.

The insulin/IGF-1 (insulin-like growth factor 1) signalling pathway promotes adipocyte differentiation via complex signalling networks. Here, using microarray analysis of brown preadipocytes that are derived from wild-type and insulin receptor substrate (Irs) knockout animals that exhibit progressively impaired differentiation, we define 374 genes/expressed-sequence tags whose expression in preadipocytes correlates with the ultimate ability of the cells to differentiate. Many of these genes, including preadipocyte factor-1 (Pref-1) and multiple members of the Wnt signalling pathway, are related to early adipogenic events. Necdin is also markedly increased in Irs knockout cells that cannot differentiate, and knockdown of necdin restores brown adipogenesis with downregulation of Pref-1 and Wnt10a expression. Insulin receptor substrate proteins regulate a necdin-E2F4 interaction that represses peroxisome-proliferator-activated receptor gamma (PPARgamma) transcription via a cyclic AMP response element binding protein (CREB)-dependent pathway. Together these define a key signalling network that is involved in brown preadipocyte determination.

Adipocytes↗

A computational model to define the molecular causes of type 2 diabetes mellitus.

BACKGROUND: Metabolic abnormalities associated with type 2 diabetes mellitus (DM2) are caused in part by inadequate insulin action and resulting changes in gene expression in the skeletal muscle. Two recent, independent studies of human skeletal muscle biopsies from ethnically diverse DM2 patients have identified coordinated reductions in the expression of the oxidative phosphorylation (OXPHOS) genes. Whether these reductions are a consequence or a cause of impaired insulin sensitivity remains an open question. METHODS: To address this question and to define the underlying molecular causes consistent with the expression changes reported in the muscle studies, we created a large-scale computable model to analyze the molecular actions and effects of insulin on muscle gene expression. The model enables computer-aided reasoning using over 210,000 molecular relationships assembled from the DM2 literature. RESULTS: We integrated the data from these muscle biopsy studies into the model and used computer-aided causal reasoning to discover mechanisms that can link alterations in OXPHOS genes to decreases in glucose transport, insulin signaling, and risk factors associated to post-transplant diabetes mellitus. CONCLUSIONS: The emerging hypotheses describe biologic effects in DM2 and offer important cues for molecular targeted therapy.

Algorithms↗

Genome-wide analysis of host responses to the Pseudomonas aeruginosa type III secretion system yields synergistic effects.

The type III secretion system (TTSS) is a dedicated bacterial pathogen protein targeting system that directly affects host cell signalling and response pathways. Our goal was to identify host responses to the Pseudomonas aeruginosa effectors, introduced into target cells utilizing the TTSS. We carried out expression profiling of a human lung pneumocyte cell line A549 exposed to isogenic mutants of P. aeruginosa PAK lacking individual or a combination of TTSS components. We then devised a data analysis method to isolate the key responses to specific secreted bacterial effector proteins as well as components of the TTSS machinery. Individually, the effector proteins elicited host responses consistent with their known functions, many of which were cell cycle-related. However, our analysis has shown that the effector proteins elicit a distinct host transcriptional response when present in combination, suggesting a synergistic effect. Furthermore, the pattern of host transcriptional responses is consistent with the pore forming ability of the TTSS needle complex. This study shows that the individual components of the TTSS define an integrated system and that a systems biology approach is required to fully understand the complex interplay between pathogen and host.

Bacterial Proteins↗

Conserved mechanisms across development and tumorigenesis revealed by a mouse development perspective of human cancers.

Identification of common mechanisms underlying organ development and primary tumor formation should yield new insights into tumor biology and facilitate the generation of relevant cancer models. We have developed a novel method to project the gene expression profiles of medulloblastomas (MBs)--human cerebellar tumors--onto a mouse cerebellar development sequence: postnatal days 1-60 (P1-P60). Genomically, human medulloblastomas were closest to mouse P1-P10 cerebella, and normal human cerebella were closest to mouse P30-P60 cerebella. Furthermore, metastatic MBs were highly associated with mouse P5 cerebella, suggesting that a clinically distinct subset of tumors is identifiable by molecular similarity to a precise developmental stage. Genewise, down- and up-regulated MB genes segregate to late and early stages of development, respectively. Comparable results for human lung cancer vis-a-vis the developing mouse lung suggest the generalizability of this multiscalar developmental perspective on tumor biology. Our findings indicate both a recapitulation of tissue-specific developmental programs in diverse solid tumors and the utility of tumor characterization on the developmental time axis for identifying novel aspects of clinical and biological behavior.

Animals↗

Quantifying the relationship between co-expression, co-regulation and gene function.

BACKGROUND: It is thought that genes with similar patterns of mRNA expression and genes with similar functions are likely to be regulated via the same mechanisms. It has been difficult to quantitatively test these hypotheses on a large scale because there has been no general way of determining whether genes share a common regulatory mechanism. Here we use data from a recent genome wide binding analysis in combination with mRNA expression data and existing functional annotations to quantify the likelihood that genes with varying degrees of similarity in mRNA expression profile or function will be bound by a common transcription factor. RESULTS: Genes with strongly correlated mRNA expression profiles are more likely to have their promoter regions bound by a common transcription factor. This effect is present only at relatively high levels of expression similarity. In order for two genes to have a greater than 50% chance of sharing a common transcription factor binder, the correlation between their expression profiles (across the 611 microarrays used in our study) must be greater than 0.84. Genes with similar functional annotations are also more likely to be bound by a common transcription factor. Combining mRNA expression data with functional annotation results in a better predictive model than using either data source alone. CONCLUSIONS: We demonstrate how mRNA expression data and functional annotations can be used together to estimate the probability that genes share a common regulatory mechanism. Existing microarray data and known functional annotations are sufficient to identify only a relatively small percentage of co-regulated genes.

Gene Expression Profiling↗

Genome-scale expression profiling of Hutchinson-Gilford progeria syndrome reveals widespread transcriptional misregulation leading to mesodermal/mesenchymal defects and accelerated atherosclerosis.

Hutchinson-Gilford progeria syndrome (HGPS) is a rare genetic disease with widespread phenotypic features resembling premature aging. HGPS was recently shown to be caused by dominant mutations in the LMNA gene, resulting in the in-frame deletion of 50 amino acids near the carboxyl terminus of the encoded lamin A protein. Children with this disease typically succumb to myocardial infarction or stroke caused by severe atherosclerosis at an average age of 13 years. To elucidate further the molecular pathogenesis of this disease, we compared the gene expression patterns of three HGPS fibroblast cell strains heterozygous for the LMNA mutation with three normal, age-matched cell strains. We defined a set of 361 genes (1.1% of the approximately 33,000 genes analysed) that showed at least a 2-fold, statistically significant change. The most prominent categories encode transcription factors and extracellular matrix proteins, many of which are known to function in the tissues severely affected in HGPS. The most affected gene, MEOX2/GAX, is a homeobox transcription factor implicated as a negative regulator of mesodermal tissue proliferation. Thus, at the gene expression level, HGPS shows the hallmarks of a developmental disorder affecting mesodermal and mesenchymal cell lineages. The identification of a large number of genes implicated in atherosclerosis is especially valuable, because it provides clues to pathological processes that can now be investigated in HGPS patients or animal models.

Adolescent↗

Coordinated reduction of genes of oxidative metabolism in humans with insulin resistance and diabetes: Potential role of PGC1 and NRF1.

Type 2 diabetes mellitus (DM) is characterized by insulin resistance and pancreatic beta cell dysfunction. In high-risk subjects, the earliest detectable abnormality is insulin resistance in skeletal muscle. Impaired insulin-mediated signaling, gene expression, glycogen synthesis, and accumulation of intramyocellular triglycerides have all been linked with insulin resistance, but no specific defect responsible for insulin resistance and DM has been identified in humans. To identify genes potentially important in the pathogenesis of DM, we analyzed gene expression in skeletal muscle from healthy metabolically characterized nondiabetic (family history negative and positive for DM) and diabetic Mexican-American subjects. We demonstrate that insulin resistance and DM associate with reduced expression of multiple nuclear respiratory factor-1 (NRF-1)-dependent genes encoding key enzymes in oxidative metabolism and mitochondrial function. Although NRF-1 expression is decreased only in diabetic subjects, expression of both PPAR gamma coactivator 1-alpha and-beta (PGC1-alpha/PPARGC1 and PGC1-beta/PERC), coactivators of NRF-1 and PPAR gamma-dependent transcription, is decreased in both diabetic subjects and family history-positive nondiabetic subjects. Decreased PGC1 expression may be responsible for decreased expression of NRF-dependent genes, leading to the metabolic disturbances characteristic of insulin resistance and DM.

Adult↗

Reproducibility of gene expression across generations of Affymetrix microarrays.

BACKGROUND: The development of large-scale gene expression profiling technologies is rapidly changing the norms of biological investigation. But the rapid pace of change itself presents challenges. Commercial microarrays are regularly modified to incorporate new genes and improved target sequences. Although the ability to compare datasets across generations is crucial for any long-term research project, to date no means to allow such comparisons have been developed. In this study the reproducibility of gene expression levels across two generations of Affymetrix GeneChips (HuGeneFL and HG-U95A) was measured. RESULTS: Correlation coefficients were computed for gene expression values across chip generations based on different measures of similarity. Comparing the absolute calls assigned to the individual probe sets across the generations found them to be largely unchanged. CONCLUSION: We show that experimental replicates are highly reproducible, but that reproducibility across generations depends on the degree of similarity of the probe sets and the expression level of the corresponding transcript.

Calibration↗

PGAGENE: integrating quantitative gene-specific results from the NHLBI programs for genomic applications.

SUMMARY: PGAGENE is a web-based gene-specific genomic data search engine, which allows users to search over 5.9 million pieces of collective genetic and genomic data from the NHLBI supported Programs for Genomic Applications. This data includes microarray measurements, SNPs, and mutations, and data may be found using symbols, parts of gene names or products, Affymetrix probe IDs, GenBank accession numbers, UniGene IDs, dbSNP IDs, and others. The PGAGENE indexing agent periodically maps all publicly available gene-specific PGA data onto LocusLink using dynamically generated cross-referencing tables.

Base Sequence↗

Computerized recruiting for clinical trials in real time.

STUDY OBJECTIVE: Success of prospective studies, particularly in the emergency department, often depends on immediate identification of eligible patients to ensure timely sample collection and initiation of study interventions. We report use of a real-time automated notification system to identify potential patients for a clinical trial at the time of ED registration on the basis of information routinely collected. We hypothesize that the automated notification system improves the rate of investigator notification. METHODS: We performed a prospective comparison of the notification rate by the automated notification system compared with that by ED clinicians. RESULTS: In the 11 months before use of the automated notification system, the investigator was notified by ED staff for 56% of 61 potentially eligible patients. During 10 months of using the automated notification system, the investigator was paged by the automated notification system for 84% of 49 potentially eligible patients. CONCLUSION: The automated notification system improves study investigator notification. Use requires online linked registration, a database, and paging systems. The automated notification system is a potentially valuable tool in the recruitment of patients for clinical trials.

Clinical Trials as Topic↗

Comparing expression profiles of genes with similar promoter regions.

MOTIVATION: Gene regulatory elements are often predicted by seeking common sequences in the promoter regions of genes that are clustered together based on their expression profiles. We consider the problem in the opposite direction: we seek to find the genes that have similar promoter regions and determine the extent to which these genes have similar expression profiles. RESULTS: We use the data sets from experiments on Saccharomyces cerevisiae. Our similarity measure for the promoter regions is based on the set of common mapped or putative transcription factor binding sites and other regulatory elements in the upstream region of the genes, as contained in the Saccharomyces cerevisiae Promoter Database. We pair up the genes with high similarity scores and compare their expression levels in time-course experiment data. We find that genes with similar promoter regions on the average have significantly higher correlation, but it can vary widely depending on the genes. This confirms that the presence of similar regulatory elements often does not correspond to similarity in expression profiles and indicates that finding transcription factor binding sites or other regulatory elements starting with the expression patterns may be limited in many cases. Regardless of the correlation, the degree to which the profiles agree under different experimental conditions can be examined to derive hypotheses concerning the role of common regulatory elements. Overall, we find that considering the relationship between the promoter regions and the expression profiles starting with the regulatory elements is a difficult but useful process that can provide valuable insights.

Databases, Nucleic Acid↗

Analysis of matched mRNA measurements from two different microarray technologies.

MOTIVATION: [corrected] The existence of several technologies for measuring gene expression makes the question of cross-technology agreement of measurements an important issue. Cross-platform utilization of data from different technologies has the potential to reduce the need to duplicate experiments but requires corresponding measurements to be comparable. METHODS: A comparison of mRNA measurements of 2895 sequence-matched genes in 56 cell lines from the standard panel of 60 cancer cell lines from the National Cancer Institute (NCI 60) was carried out by calculating correlation between matched measurements and calculating concordance between cluster from two high-throughput DNA microarray technologies, Stanford type cDNA microarrays and Affymetrix oligonucleotide microarrays. RESULTS: In general, corresponding measurements from the two platforms showed poor correlation. Clusters of genes and cell lines were discordant between the two technologies, suggesting that relative intra-technology relationships were not preserved. GC-content, sequence length, average signal intensity, and an estimator of cross-hybridization were found to be associated with the degree of correlation. This suggests gene-specific, or more correctly probe-specific, factors influencing measurements differently in the two platforms, implying a poor prognosis for a broad utilization of gene expression measurements across platforms.

Cluster Analysis↗