PubMed HealthSearch

SEARCH · PubMed Health

Results for “EHR”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Robust replication of associations across patient-mediated and provider-sourced EHR data in the All of Us research program.

The All of Us Research Program is assembling a nationwide cohort with electronic health record (EHR) resources through two complementary pathways: healthcare provider organization (HPO)-sourced EHRs and patient-mediated EHR (PME) contributed through patient portal linkages. The comparative research utility of these two data sources has not been systematically evaluated. Here, we compared PME and HPO EHRs with respect to disease prevalence, phenotype-phenotype associations, and replication of established genotype-phenotype associations using data from 19,703 PME and 373,887 HPO participants. We benchmarked disease prevalence against national estimates, conducted phenome-wide association studies for 10 commonly studied diseases, and tested replication of more than 5000 established genotype-phenotype associations across multiple ancestral groups. Disease prevalence was consistently lower in PME than in HPO, although prevalence of most diseases in both cohorts exceeded national estimates. Both data sources reproduced known phenotype-phenotype associations and showed moderate-to-strong concordance in effect sizes across the phenome. The overall genotype-phenotype replication rate was 49.1% (5399/10,999) in HPO and 5.9% (381/6482) in PME across ancestral groups, with effect sizes strongly correlated among well-powered associations (R&#x2009;=&#x2009;0.84, P&#x2009;<&#x2009;0.001). To disentangle the impact of sample size from data quality, we performed 1:1 propensity score matching. After matching, the replication gap in genotype-phenotype associations narrowed from 8.3-fold to 1.3-fold, with equivalent replication rates among adequately powered associations and strongly concordant effect sizes; comorbidity patterns were also consistent across all 10 diseases tested. These findings demonstrate that both data sources are valuable for clinical and genomic research and can inform other cohorts integrating provider-derived and patient-mediated EHRs.

Computational biology and bioinformatics

Characterizing trends in clinical genetic testing: A single-center analysis of EHR data from 1.8 million patients over two decades.

A lack of structural data in electronic health records (EHRs) makes assessing the impact of genetic testing on clinical practice challenging. We extracted clinical genetic tests from the EHRs of more than 1.8 million patients seen at Vanderbilt University Medical Center from 2002 to 2022. With these data, we quantified the use of clinical genetic testing in healthcare and described how testing patterns and results changed over time. We assessed trends in types of genetic tests, tracked usage across medical specialties, and introduced a new measure, the genetically attributable fraction (GAF), to quantify the proportion of observed phenotypes attributable to a genetic diagnosis over time. We identified 104,392 tests and 19,032 molecularly confirmed diagnoses. The proportion of patients with genetic testing in their EHRs increased from 1.0% in 2002 to 6.1% in 2022, and testing became more comprehensive with the growing use of multi-gene panels. The number of unique diseases diagnosed with genetic testing increased from 51 in 2002 to 509 in 2022, and there was a rise in the number of variants of uncertain significance. The phenome-wide GAF for 6,505,620 diagnoses made in 2022 was 0.46%, and the GAF was greater than 5% for 74 phenotypes, including pancreatic insufficiency (67%), chorea (64%), atrial septal defect (24%), microcephaly (17%), paraganglioma (17%), and ovarian cancer (6.8%). Our study provides a comprehensive quantification of the increasing role of genetic testing at a major academic medical institution and demonstrates its growing utility in explaining the observed medical phenome.

Humans

An EHR-based framework for modeling growth curves and constructing growth centile charts for genetic disorders.

Growth modeling is central to human genetics, as deviations from typical growth can signal an underlying disorder. In this cohort study, we developed a generalizable framework for generating growth charts across genetic conditions using electronic health records (EHR). Leveraging 22 years of longitudinal EHR data from 452,470 patients across 15 genetic conditions and unaffected individuals, we generated sex- and condition-specific growth charts using Generalized Additive Models for Location, Scale, and Shape, and quantified differences in size, timing, and intensity using SuperImposition by Translation and Rotation (SITAR). SITAR-derived growth parameters showed strong concordance with established annotations in OMIM and Orphanet, and identified previously unreported growth patterns. We stratified cystic fibrosis by CFTR functional class and observed greater growth impairment in individuals with homozygous minimal-function variants compared to those with residual function. This framework provides a generalizable approach for leveraging EHR data to refine genotype-phenotype relationships and enable continuous updating of growth charts across genetic conditions.

Journal Article

A doubly robust framework for addressing outcome-dependent selection bias in multi-cohort EHR studies.

Selection bias can hinder accurate estimation of association parameters in binary disease risk models using non-probability samples like electronic health records (EHRs). The issue is compounded when participants are recruited from multiple clinics/centers with varying selection mechanisms that may depend on the disease/outcome of interest. Traditional inverse-probability-weighted (IPW) methods, based on constructed parametric selection models, often struggle with misspecifications when selection mechanisms vary across cohorts. This paper introduces a new Joint Augmented Inverse Probability Weighted (JAIPW) method, which integrates individual-level data from multiple cohorts collected under potentially outcome-dependent selection mechanisms, with data from an external probability sample. JAIPW offers double robustness by incorporating a flexible auxiliary score model to address potential misspecifications in the selection models. We outline the asymptotic properties of the JAIPW estimator, and our simulations reveal that JAIPW achieves up to 6 times lower relative bias and 5 times lower root mean square error (RMSE) compared to the best performing joint IPW methods under scenarios with misspecified selection models. Applying JAIPW to the Michigan Genomics Initiative (MGI), a multi-clinic EHR-linked biobank, combined with external national probability samples, resulted in cancer-sex association estimates closely aligned with national benchmark estimates. We also analyzed the association between cancer and polygenic risk scores (PRS) in MGI to illustrate a situation where the exposure variable is not measured in the external probability sample.

Selection Bias

Privacy-Enhancing Sequential Learning under Heterogeneous Selection Bias in Multi-Site EHR Data.

OBJECTIVE: To develop privacy-enhancing statistical methods for estimation of binary disease risk model association parameters across multiple electronic health record (EHR) sites with heterogeneous selection mechanisms, without sharing raw individual-level data. We illustrate their utility through a cross-biobank analysis of smoking and 97 cancer subtypes using data from the NIH All of Us (AOU) and the Michigan Genomics Initiative (MGI). MATERIALS AND METHODS: Large-scale biobanks often follow heterogeneous recruitment strategies and store data in separate cloud-based platforms, making centralized algorithms infeasible. To address this, we propose two decentralized sequential estimators namely, Sequential Pseudo-likelihood (SPL) and Sequential Augmented Inverse Probability Weighting (SAIPW) that leverage external population-level information to adjust for selection bias, with valid variance estimation. SAIPW additionally protects against misspecification of the selection model using flexible machine learning based auxiliary outcome models. We compare SPL and SAIPW with the existing Sequential Unweighted (SUW) estimator and with centralized and meta learning extensions of IPW and AIPW in simulations under both correctly specified and misspecified selection mechanisms. We apply the methods to harmonized data from MGI ( n = 50,935) and AOU ( n = 241,563) to estimate smoking-cancer associations. RESULTS: In simulations, SUW exhibited substantial bias and poor coverage. SPL and SAIPW yielded unbiased estimates with valid coverage probabilities under correct model specification, with SAIPW remaining robust under selection model misspecification. Both approaches showed no notable efficiency loss relative to centralized methods. Meta-learning methods were efficient for large sites but failed in settings with small cohort sizes and rare outcome prevalence. In real-data analysis, strong associations were consistently identified between smoking and cancers of the lung, bladder, and larynx, aligning with established epidemiological evidence. CONCLUSION: Our framework enables valid, privacy-enhancing inference across EHR cohorts with heterogeneous selection, supporting scalable, decentralized research using real-world data.

Journal Article

RESCUE: An end-to-end multi-agent LLM system for proactive rare-disease patient screening in the EHR.

BACKGROUND: Rare diseases affect a significant portion of the global population, yet patients often endure a lengthy diagnostic odyssey, frequently missing the opportunity for timely diagnoses with exome or genome sequencing (ES/GS). Existing informatics tools often rely on pre-identified patients or rigid, institution-specific rule sets, failing to address the broader operational question of clinical utility and feasibility. METHODS: We introduce RESCUE (Rare Disease Detection and Escalation Support via a Learning Health System), an end-to-end, multi-agent LLM-powered workflow designed for proactive rare-disease diagnosis across the entire electronic health record (EHR). RESCUE utilizes a team of specialized agents including Ontology, Modeling, Screening, and Review, to automate the screening process to identify candidates for diagnostic testing based on their clinical features. The Ontology Agent classifies clinical data into a four-tier genetic-evidence taxonomy; the Modeling Agent builds a positive-unlabeled (PU) XGBoost classifier to identify potential cases; the Screening Agent applies these models across the EHR population; and the Review Agent evaluates candidates by sampling clinical notes to ensure medical necessity and operational feasibility for genomic testing. RESULTS: Using electronic medical record data from a pediatric hospital, our retrospective evaluation on a holdout set (n=12,591) demonstrates strong discrimination between patients who received diagnostic genomic testing and those who did not (AUC 0.808). Of nearly 500,000 patients in the institutional base, 175,842 met inclusion criteria for screening; among these, RESCUE-flagged candidates were 7.4-fold more likely to receive subsequent genomic assessments compared to controls. Blinded manual chart reviews confirmed that RESCUE identifies previously missed, medically appropriate patients for ES/GS with 80% precision, while simultaneously accounting for prior testing history. CONCLUSIONS: By decoupling expert roles into modular agents, RESCUE offers a flexible, scalable, and adaptable framework for screening patients for rare-disease diagnostic genomic testing. This approach overcomes the limitations of traditional rule-based methods and provides a reproducible, agentic pathway to reduce diagnostic delays and improve patient care at an institutional scale.

Journal Article

Unsupervised characterization of 100,272 EHR patients identifies high-risk groups and comorbidities linked to premature aging.

Electronic health records (EHRs) contain extensive multidimensional patient data, presenting challenges for the discovery of novel and meaningful clinical patterns. Unsupervised clustering of high-dimensional clinical data holds great potential for identifying novel clinical patterns. Here, we performed unsupervised clustering and characterized 100,272 patients in the Electronic Medical Records and GEnomics (eMERGE) Network. We identified 70 clusters defined by distinct comorbidity patterns. Meanwhile, age and sex are also strongly associated with patient stratification, influencing phenotype prevalence and onset time. Notably, phenotype onset time accurately predicted chronological age and was significantly associated with overall mortality risk. Besides age and sex, we assessed the contribution of genetic variation to phenotype development and observed evidence of cross-phenotype associations influencing cluster membership and comorbidity patterns. However, the role of genetics recedes during aging. We also identified several high-risk clusters with elevated Charlson Comorbidity Index (CCI) scores and validated these findings in an independent cohort. Further analysis of these clusters revealed phenotypes linked to premature aging and highlighted a survival selection among older participants in observational studies. Overall, this study enables phenome-wide unsupervised patient stratification for multimorbidity discovery in largely unannotated clinical data, offering valuable insights into patient stratification, comorbidity analysis, aging, and health outcomes.

Journal Article

Integrating explainable AI with multiomics systems biology and EHR data mining for personalized drug repurposing in Alzheimer's disease.

Alzheimer's disease (AD) is characterized by region- and patient-specific molecular heterogeneity, which hinders therapeutic design. In this study, we introduce PRISM-ML (PRecision-medicine using Interpretable Systems and Multiomics with Machine Learning), an open-source integrated analysis pipeline that combines interpretable machine learning with systems biology and electronic health record (EHR) data mining to elucidate the molecular diversity of AD and predict promising drug repurposing opportunities. First, we integrated and harmonized transcriptomic (bulk RNA-seq) and genomic (genome-wide association study) data from 2105 brain samples, each with matched data from the same individual (1363 AD patients, 742 controls; nine tissues), sourced from three independent studies. Random forest classifiers with SHapley Additive exPlanations (SHAP) identified patient-specific biomarkers; unsupervised clustering resolved 36 molecularly distinct "subtissues" (clusters of samples); and gene-gene co-expression networks prioritized 262 high-centrality bottleneck genes as putative regulators of dysregulated pathways. Next, knowledge graph-based drug repurposing predicted six FDA-approved drugs that simultaneously target multiple bottleneck genes and multiple AD-relevant pathways. Notably, in a large U.S. de-identified insurance-claims database (n = 364733), exposure to promethazine, one of the candidate drugs, was associated with a 57-62 % lower incidence of AD versus an active antihistamine comparator (adjusted hazard ratio 0.38; inverse-probability weighted 0.43; both p < 0.001), providing real-world support for its repurposing potential. In summary, PRISM-ML, as an explainable multi-omics analysis pipeline, is readily transferable to other complex diseases, advancing precision medicine.

Computational Biology

Nucleolar localization of rRNA coding sequences in Prorocentrum micans Ehr. (dinomastigote, kingdom Protoctist) by in situ hybridization.

To define the molecular mechanisms of ribosome biogenesis and to find out in which nucleolar compartment transcription of rDNA occurs, we have performed in situ hybridization (ISH) of RNase-treated cryosections using biotinylated rRNA coding sequences as a probe and the eukaryotic dinoflagellate nucleolar system as a model. Recent data from ISH of eukaryotic ribosomal genes by electron microscopy (EM) has so far failed to establish a consensus which clearly defines the function of the three compartments of the nucleolus. Dinomastigote protoctists are the only known eukaryotes whose chromatin is totally devoid of nucleosomes. Their chromosomes remain permanently condensed during the entire cell cycle and active nucleoli arise from an unwound part of some of the otherwise compact chromosomes. In this work, DNA-DNA hybrids were detected either by fluorescent avidin or by indirect immunogold staining procedures in EM; this is the first use of cryosections to detect hybrids in EM not only in the nucleolus sensu lato but also in a dinomastigote cell. Coding sequences of ribosomal genes were detected both in the periphery of the nucleolar organizer region (NOR), which corresponds to the unwound part of the nucleolar chromosome, and in the proximal part of the fibrillo-granular (FG) region. These results suggest that the rRNA gene transcription predominantly occurs at the periphery of the NOR where the coding sequences are located. A predictive model summarizes and allows discussions and comparisons with other eukaryotes in which nucleolar mechanisms were previously studied. This leads to the conclusion that dinoflagellate cells constitute an excellent model for the study of the functional structure of the eukaryotic nucleolus.(ABSTRACT TRUNCATED AT 250 WORDS)

Animals

Feasibility of implementation, diagnostic accuracy, and end-user impact of an electronic health record (EHR)-based ureteral stent tracking tool in a pediatric population.

INTRODUCTION & OBJECTIVES: Ureteral stent tracking systems have reduced stent retention in adults, but their accuracy and impact in pediatrics have been minimally explored. With low event rates in children, such tools may yield high false positives, raising questions on balancing event prevention with provider burden. We aimed to evaluate the feasibility, diagnostic accuracy, and end-user impact of an Electronic Surveillance Tool for Evaluating Nephroureteral stent Tracking (eSTENT) at our institution. STUDY DESIGN: eSTENT, implemented in 1/2024, flags ureteral stents at risk for retention based on implant documentation, expected explant date, and explant documentation. Monthly reports are generated for stents missing explant documentation. We retrospectively evaluated the diagnostic performance of eSTENT from 1/2024-8/2025 at our pediatric hospital. A usability survey including a validated 1-7 implementation score (higher = easier implementation) was distributed to pediatric urologists and operating room nurses. RESULTS: Of 172 cases with ureteral stent placement, eSTENT flagged 28 events (16%) in 24 patients. Of these, 26 represented documentation gaps where explant had been appropriate. Two flags had no documentation of explant, representing near miss events that were identified. No retained stents occurred, consistent with high sensitivity and modest specificity. There were no flags in the last 6 months of the study period. Survey response rate was 100% for surgeons and 55% for nurses. Before eSTENT, stents were not routinely tracked. All surgeons and 93% of nurses reported no added burden, despite occasional misidentification of retained stents. Three surgeons found eSTENT beneficial, four were neutral, and free-text responses generally cited eSTENT's "fail safe" nature as positive. Nurses suggested improvements, including user support and integrated documentation reminders. The average implementation score among both groups was 6/7, indicating easy adoption. DISCUSSION: While the impact of stent tracking tools in adult literature has been positive, our study emphasizes the feasibility of broader adoption at a pediatric hospital. Integration of eSTENT may avoid the potentially devastating consequences of a retained stent. Prioritizing sensitivity over specificity appears acceptable for a "never event" in patient safety. Our study is limited by the retrospective nature of data collection and survey bias. CONCLUSIONS: Though no stents were retained in the study period, eSTENT appropriately flagged two cases without added burden to most end-users. Further optimization is warranted, but adoption in pediatric centers may enhance care reliability.

Humans

Observations on the ultrastructure of the choanoflagellate Codosiga botrytis (Ehr.) Saville-Kent with special reference to the flagellar apparatus.

The ultrastructure of the choanoflagellate Codosiga botrytis is described with particular reference to the flagellar appendages, the flagellar rootlet system, the transition zone, the basal body and accessory centrioles, and the stalk. The controversial early reports of flagellar appendages in this species have been confirmed and they have been detected in 2 further species, Salpingoeca frequentissima and Monosiga sp. The appendages consist of a delicate bilateral vane 2 mum wide on either side of the axis, composed of extremely fine overlapping or interwoven fibrils. The flagellar root system consists of a large number of radiating microtubules associated with bands of electron-dense material near the basal body; striated roots are absent. The microtubules extend from several separate foci, those in any one group originating near a composite electron-dense band, and for a distance of 300 nm from the basal body they are separated by blocks of interstitial material. The flagellar basal body forms one of a diplosome pair of centrioles. The triplet microtubules of the accessory centriole are embedded in amorphous electron-dense material and the whole is enveloped in a sheath of similar appearance. The existence of a third centriole close to the diplosome pair is also reported. The relatively complex structure of the flagellar transitional zone is described. The stalk is composed of a core of circular lacunae, which may or may not contain finger-like protoplasmic extensions of the posterior end of the cell, surrounded by a continuation of the sheath material which encloses the remainder of the protoplast. In the stalk only there is a further closely sheathing layer about 15 nm thick which is regularly striated, the spacing of the striations in shadowcast material and sections being about 3 times that measured by negative staining. The structure of choanoflagellates differs widely from that of the algal class Chrysophyceae, the group in which they are included in some classifications, and from the remainder of the algae; they do not appear to have a place in either the algae or the plant kingdom. The structure of Codosiga botrytis is briefly compared with that of sponge choanocytes and collared cells in the Metazoa and some of the possible phylogenetic implications of this are indicated.

Animals

Algorithms for the identification of prevalent diabetes in the All of Us Research Program validated using polygenic scores.

The All of Us Research Program (AoU) is an initiative designed to gather a comprehensive and diverse dataset from at least one million individuals across the USA. This longitudinal cohort study aims to advance research by providing a rich resource of genetic and phenotypic information, enabling powerful studies on the epidemiology and genetics of human diseases. One critical challenge to maximizing its use is the development of accurate algorithms that can efficiently and accurately identify well-defined disease and disease-free participants for case-control studies. This study aimed to develop and validate type 1 (T1D) and type 2 diabetes (T2D) algorithms in the AoU cohort, using electronic health record (EHR) and survey data. Building on existing algorithms and using diagnosis codes, medications, laboratory results, and survey data, we developed and implemented algorithms for identifying prevalent cases of type 1 and type 2 diabetes. The first set of algorithms used only EHR data (EHR-only), and the second set used a combination of EHR and survey data (EHR+). A universal algorithm was also developed to identify individuals without diabetes. The performance of each algorithm was evaluated by testing its association with polygenic scores (PSs) for type 1 and type 2 diabetes. We demonstrated the feasibility and utility of using AoU EHR and survey data to employ diabetes algorithms. For T1D, the EHR-only algorithm showed a stronger association with T1D-PS compared to the EHR&#x2009;+&#x2009;algorithm (DeLong p-value&#x2009;=&#x2009;3&#x2009;&#xd7;&#x2009;10-5). For T2D, the EHR&#x2009;+&#x2009;algorithm outperformed both the EHR-only and the existing T2D definition provided in the AoU Phenotyping Library (DeLong p-values&#x2009;=&#x2009;0.03 and 1&#x2009;&#xd7;&#x2009;10-4, respectively), identifying 25.79% and 22.57% more cases, respectively, and providing an improved association with T2D PS. We provide a new validated type 1 diabetes definition and an improved type 2 diabetes definition in AoU, which are freely available for diabetes research in the AoU. These algorithms ensure consistency of diabetes definitions in the cohort, facilitating high-quality diabetes research.

Humans

Surface markers on human activated T lymphocytes. III. High-affinity E-rosette receptors.

High-affinity E-rosette receptor (EhR) is suggested to be an activation marker of human T lymphocytes. The present study has been undertaken to estimate the dependence of EhR expression upon the T cell activation. To this aim the expression of EhR on T cells stimulated with PHA was examined. Nonsynchronized and arrested in G1 phase of cell cycle cultures of T lymphocytes were employed. The kinetics of expression of either de novo induced or reexpressed EhR was estimated by using two T cell subsets (originally EhR- and EhR+). In accordance with opinion of others, our results have shown that EhR is an activation marker of human T cells. Moreover, we have found that EhR is the marker of early and late activation stage because: its expression is induced in G1 phase of cell cycle, it continues to increase with the time of mitogen stimulation and the maximal level of EhR expression coincides with the time of the maximal RNA and DNA synthesis. Both induced de novo and reexpressed Eh receptors display the same expression kinetics.

Antigens, Differentiation

Menopause in the All of Us Research Program: a descriptive summary of electronic health record and survey response across sociodemographic characteristics.

OBJECTIVES: Menopause is a significant physiological transition with implications for health outcomes (eg, cardiometabolic disease), yet gaps remain in understanding this transition, including how menopause timing and type influence health outcomes. Large-scale cohort studies in midlife (age=40-60) females, including the All of Us Research Program (AoURP), provide opportunities to study menopause across diverse populations and data modalities. We characterized menopause-related data in AoURP, focusing on age distributions and concordance between electronic health record (EHR) diagnosis codes and survey responses. METHODS: We analyzed menopause-related surveys, EHR diagnostic codes, and genomic data among ~396,000 AoURP female participants. We summarized menopause-related variables across data sources, evaluated overlap between survey, EHR, and genomic data sets, and described age distributions overall and across sociodemographic characteristics. RESULTS: Among ~396,000 females, survey responses captured ~193,000 menopause observations, nearly seven times more than EHR diagnoses (~28,000), suggesting under-ascertainment in EHR data. Nearly all females (~99%) with an EHR menopause diagnosis reported menopause in the survey. Approximately 22,000 participants had overlapping menopause-related EHR, survey, and genomic data. Survey age patterns matched expectations, with participants predominantly <40 years reporting premenopausal status and those >60 years reporting postmenopausal status. A small subset with age >70 years (N&#x2248;1,700; 4%) reported no menopause, suggesting response or recall bias. EHR menopause codes were concentrated after age 45 years, with a notable spike at age 65. Modest differences in survey-based menopause age distributions were observed across sociodemographic characteristics (eg, race and ancestry). CONCLUSIONS: These findings inform sampling strategies, power calculations, phenotype definition, and study design for menopause research using AoURP data.

Age

Heart rate as a predictor of first-trimester spontaneous abortion after ultrasound-proven viability.

To discover whether first-trimester spontaneous abortion can be predicted by embryonic heart rate (EHR), we performed a cross-sectional study during the first trimester of pregnancy using high-frequency transvaginal sonography combined with pulsed Doppler. Heart rate was measured in 603 embryos; of these, 580 continued beyond 13 weeks' gestation and 23 ended in first-trimester spontaneous abortion. Based on the continuing pregnancies, we constructed and compared EHR nomograms relating to gestational age, mean diameter of the gestational sac, and crown-rump length (CRL). Embryonic heart rate correlated best with CRL (r = 0.87), and the correlation was best described by a second-degree polynomial regression equation. The mean EHR increased progressively from 110 beats per minute (bpm) at CRL of 3-4 mm to 171-178 bpm at CRL 15-32 mm. At CRL greater than 32 mm, the EHR remained stable at a mean of 170 bpm. The EHRs of the 23 embryos that spontaneously aborted in the first trimester were evaluated according to these nomograms. In 15 cases, the EHR fell outside the 95% confidence interval for CRL (sensitivity 65%), but it was within normal limits in eight (false-negative rate 35%). In ten embryos for which pregnancy continued beyond 13 weeks, the EHR fell outside the 95% confidence interval (specificity 98%, false-positive rate 2%). Our findings suggest that EHR measurements in early pregnancy may be useful in the prediction of first-trimester spontaneous abortion after ultrasound-proven viability.

Abortion, Spontaneous

Negative descriptors in electronic health records of patients with diabetes.

BACKGROUND: Negative descriptors in electronic health records (EHR) contribute to worse health outcomes; studies show they are also more prevalent in EHRs of women and racial minorities and affect downstream research biases. Similar and unique patterns of negative descriptors may also exist in the records of blind patients, including those with diabetic retinopathy. Diabetic retinopathy is a preventable but leading cause of blindness in the US that is disproportionally high among women and racial and ethnic minorities. METHODS: Using EHR from a large medical center, we created "matched" cohorts of patients with a type 2 diabetes-only diagnosis and patients with a diagnosis of diabetic retinopathy. We identified previously used and new, disability and patient-related negative descriptors and assessed patterns of biased language in the EHR, comparing patients by retinopathy diagnosis (yes/no), and changes in patterns of language usage pre- and post- the retinopathy diagnosis. We also assessed differences between patients with type 2 diabetes at the intersection of blindness (ie, retinopathy diagnosis) and self-reported gender and race and ethnicity marginalization. RESULTS: The EHRs of patients with diabetic retinopathy were significantly more likely than those of patients with diabetes-only diagnoses to contain biased language, across queried negative descriptors. The biasing language was consistently more prevalent in EHRs of patients with diabetic retinopathy identifying as women, Black/African Americans and Hispanic compared to White men and more likely to occur following patients' retinopathy diagnosis. CONCLUSIONS: Our study indicates the presence of both disability- and intersectional biases in EHRs. We discuss findings' implications and suggest steps to address them.

Humans

Real-World Treatment Patterns and Clinical Outcomes After First-Line Therapy in Patients with KRAS G12C-Mutant Advanced Non-Small-Cell Lung Cancer in the United States.

BACKGROUND: Approximately 13% of NSCLC cases have KRAS G12C mutations. As therapeutic strategies targeting KRAS G12C-mutant NSCLC evolve, it is important to understand clinical presentation and current outcomes for these patients. METHODS: This retrospective study used data from two US nationwide databases, an electronic health records (EHR) database and a clinico-genomic database (CGDB) of EHR data linked to data from comprehensive genomic profiling tests. Eligible patients had advanced NSCLC, initiated first-line therapy from August 2018 to December 2022, and had KRAS test results. Clinicopathologic characteristics, treatments, real-world progression-free survival (rwPFS), and overall survival (OS) were analyzed. RESULTS: There were 1227 patients with KRAS G12C-mutant NSCLC in the EHR database and 447 in the CGDB. First-line regimen was platinum-based chemotherapy plus pembrolizumab for 46% and pembrolizumab monotherapy for 20%. Less than 40% of patients received second-line therapy. Median (95% CI) OS for KRAS G12C-mutant NSCLC patients in the EHR was 17.0 (15.2-18.9) months. Variables significantly associated with shorter OS included PD-L1 <1%, brain metastases, STK11 co-mutation, and poor performance status. Patients treated with platinum-based chemotherapy plus pembrolizumab had median rwPFS of 5.3 (4.5-7.3) months and OS of 12.8 (11.1-17.3) months in the CGDB; median OS was 15.6 (12.5-18.6) months in the EHR. Patients with PD-L1 &#x2265; 50% treated with pembrolizumab monotherapy had median rwPFS of 4.6 (3.0-15.6) months and OS of 20.4 (10.3-38.5) months in the CGDB; median OS was 22.1 (18.7-30.7) in the EHR. CONCLUSIONS: These data provide a real-world benchmark of outcomes for patients with KRAS G12C-mutant NSCLC receiving the current standard of care and indicate an unmet need for more effective first-line therapies.

KRAS G12C