PubMed HealthSearch

SEARCH · PubMed Health

Results for “Datasets as Topic”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

43 records · Page 3Linked to original sources

Meta-analysis models with group structure for pleiotropy detection at gene and variant level using summary statistics from multiple datasets.

Genome-wide association studies (GWASs) have highlighted the importance of pleiotropy in human diseases, where one gene can impact 2 or more unrelated traits. Examining shared genetic risk factors across multiple diseases can enhance our understanding of these conditions by pinpointing new genes and biological pathways involved. Furthermore, with an increasing wealth of GWAS summary statistics available to the scientific community, leveraging these findings across multiple phenotypes could unveil novel pleiotropic associations. Existing selection methods examine pleiotropic associations one by one at a scale of either the genetic variant or the gene, and thus cannot consider all the genetic information at the same time. To address this limitation, we propose a new approach called MPSG (Meta-analysis model adapted for Pleiotropy Selection with Group structure). This method performs a penalized multivariate meta-analysis method adapted for pleiotropy and takes into account the group structure information nested in the data to select relevant variants and genes (or pathways) from all the genetic information. To do so, we implemented an alternating direction method of multipliers algorithm. We compared the performance of the method with other benchmark meta-analysis approaches such as GCPBayes, PLACO, and ASSET by considering as inputs different kinds of summary statistics. We provide an application of our method to the identification of potential pleiotropic genes between breast and thyroid cancers.

Humans

Methods for defining equity-stratifying variables: a systematic review of validation studies.

BACKGROUND AND OBJECTIVE: Disease burden is often disproportionally higher among those who are socially disadvantaged by factors defined in the PROGRESS-Plus framework (ie, Place of residence, Race/ethnicity/culture/language, Occupation, Gender/sex, Religion, Education, Socioeconomic status, and Social capital, with "Plus" covering features like age and disability). The accuracy and applicability of case definitions to identify these variables from administrative and clinical health data are unknown. We conducted a systematic review to explore how equity-stratifying variables, as categorized by the PROGRESS-Plus framework, have been defined and validated in epidemiologic studies using administrative health, population-level, or electronic health record (EHR) data. METHODS: Medline, EMBASE, CINAHL, Web of Science, and Google Scholar were searched from the inception of the databases to 2024 for validation studies of equity-stratifying variables in adults using administrative health datasets, health registries, or EHR data. Titles and abstracts, followed by relevant full-text articles, were screened in duplicate by two reviewers for eligibility. The data sources utilized, algorithms employed, and their associated performance measures were extracted and synthesized from included studies. Given substantial heterogeneity in study design, equity-stratifying variable definition, and performance metrics, meta-analysis was not possible. RESULTS: Of the 9099 unique citations screened, 188 full texts were reviewed and 116 were included in this review. Most studies were published between 2019 and 2024 (n = 64, 55%) and were validation studies of race/ethnicity definitions that used race/ethnicity codes or surname list algorithms (n = 66, 57%). No studies examined religion. Regarding the reported performance measure estimates, the race/ethnicity/culture/language equity-stratifying variables category had the largest variability across sensitivity, positive predictive value (PPV), and Cohen's Kappa. Occupation validation studies had the lowest variation in sensitivity and PPV. CONCLUSION: Despite an increasing number of publications reporting on the validation of equity-stratifying variables relevant to the PROGRESS-Plus framework, performance measures varied widely across studies. The significant heterogeneity in equity-stratifying variable definitions and methods used to validate them support the need for further rigorous validation of equity-stratifying variables in administrative and clinical health data. PLAIN LANGUAGE SUMMARY: Disease burden is often higher in people who experience financial hardships, lower level of education, discrimination due to race/ethnicity, and unstable housing. These social factors can be considered health equity factors and are important for understanding health inequalities. Health researchers often use large datasets, such as hospital or electronic health records (EHRs), to study these health equity factors. However, it is not clear how accurately these data sources capture information about people's social circumstances and how these factors are defined. In this study, we reviewed existing research to understand how health equity factors have been defined across health data sources and how accurate they are at measuring aspects of health equity and social disadvantage. Of the more than 9000 studies we identified, we included 116 that met our criteria for this systematic review. Most included studies focused on identifying race and ethnicity, often using codes or surname-based methods. We found that the accuracy of these methods varied widely across studies, meaning results may not always be reliable or comparable. Overall, our findings show that there are inconsistencies in how social factors are defined and measured in health data. This makes it difficult to fully understand and address health inequalities using routinely collected health data. More work is needed to develop and validate better quality and more consistent methods for capturing these important social factors.

Humans

Morphometrical analysis in ulcerative colitis with dysplasia and carcinoma.

Semi-automatic image analysis was used to assess the epithelium in ulcerative colitis with dysplasia and carcinoma. There were three main sources of variation within the dataset: (1) nuclear size, nuclear cytoplasmic ratio and nuclear stratification; (2) the variation of nuclear size; and (3) nuclear shape and polarity. Discriminant analysis chose the mean nuclear cytoplasmic ratio % and the coefficient of variation of nucleus to cell apex distance to derive a scoring system which completely separated normal mucosa (n = 20) and carcinoma (n = 30). The classification rule allocated all high grade dysplasia to the tumour category. Scores for regeneration and low grade dysplasia overlapped with each other and the normal and tumour groups. Scatter plots of the two discriminating variables showed good separation of regeneration and high grade dysplasia, and a degree of overlap with low grade dysplasia. The scatter plots allowed identification of overlapping and misallocated cases, requiring review of their histology and redesignation of the diagnosis in five cases. This study confirms quantitatively the visual criteria used in grading mucosal changes and their trend from regeneration through dysplasia to carcinoma. It underlines the necessity of assessing not only cytological but also architectural and inflammatory components when diagnosing regeneration and low grade dysplasia. Mucosal morphometry may be of use in confirming high grade dysplasia which is an indication for colectomy.

Adenocarcinoma

Classification of normal colorectal mucosa and adenocarcinoma by morphometry.

Semi-automatic image analysis was used to make a morphometrical assessment of 15 nuclear and cellular variables in normal (n = 20) and malignant (n = 30) colorectal epithelium. Principal components analysis on the matrix of correlations between variables identified four main sources of variation within the dataset. These were, in decreasing order of importance: (1) nuclear size, nuclear cytoplasmic ratio and nuclear position within the cell; (2) the variability of nuclear size; (3) nuclear elongation and polarity; (4) nuclear shape and its variation. Discriminant analysis was conducted between histologically normal mucosa (n = 10) and adenocarcinoma in ulcerative colitis (n = 20). Using stepwise variable selection, the mean nuclear cytoplasmic ratio (normal, mean 20.4 (s.d. +/- 2.0); tumour, mean 39.7 (s.d. +/- 7.0)) and the coefficient of variation of nucleus to cell apex distance (normal, mean 19.2 (s.d. +/- 7.5); tumour, mean 47.8 (s.d. +/- 9.1)) were chosen as discriminating features. They were used to derive a discriminant function which gave perfect discrimination between the two groups. Scatter plots of these two variables confirmed complete separation of normal mucosa from adenocarcinoma and provided a simple method of applying the discriminant function. Discriminatory performance did not deteriorate when the function was applied to further normals (n = 10) and adenocarcinoma (n = 10). This study highlights the descriptive differences between normal and malignant colorectal epithelium and shows that case allocation may be made to these two lesion categories using a morphometrically-derived classification rule.

Adenocarcinoma

Eligibility of real-world patients for aspirin primary prevention trials in cardiovascular disease.

BACKGROUND: Evidence for the net benefit of aspirin for primary prevention of cardiovascular disease (CVD) is finely balanced, leading to variation in guideline recommendations internationally. External validity of randomised clinical trial (RCT) evidence may therefore be of particular importance. The aim of this study is to characterise real-world patients according to their eligibility for guideline-cited aspirin RCTs for primary CVD prevention. METHODS: Eligibility criteria from 14 RCTs were applied to a linked primary care/hospital discharge dataset of people&#x2009;&#x2265;&#x2009;40 years without CVD. Proportions eligible for each trial were calculated, and characteristics of eligible and ineligible patients compared for each trial, including Cox regression analysis of event rates for major adverse cardiovascular events (MACE), major bleeding events, and non-cardiovascular mortality. RESULTS: Of 570,211 included patients (300,500 [52.7%] women, 336,877 [59%]&#x2009;<&#x2009;60 years), the median proportion ineligible for 14 RCTs was 90.7% (range 42.5-99.4%) and 24.0% of patients were ineligible for all RCTs. On average, trial-ineligible populations were younger (median age trial-ineligible 57.8 vs trial-eligible 62.6 years, p&#x2009;=&#x2009;0.008) and a lower proportion had hypertension (23.9% vs 50.9%, p&#x2009;=&#x2009;0.004), diabetes (6.4% vs 11.5%, p&#x2009;=&#x2009;0.015), or a regular statin prescription (11.8% vs 26.7%, p&#x2009;=&#x2009;0.001). Trial-ineligible populations had a higher hazard of MACE compared to trial-eligible in four RCTs and lower in ten (hazard ratio [HR] range across all RCTs 0.45 [95%CI 0.40-0.51] to 2.78 [95%CI 2.61-2.96]). Hazards of bleeding events in the trial-ineligible were lower than the trial-eligible in eight RCTs and higher in four (HR range across all RCTs 0.63 [95%CI, 0.59-0.66] to 1.69 [95%CI, 1.53-1.86]), and time-varying hazards of non-CVD death were consistently lower in four RCTs and higher in five (HR range across all RCTs and time points 0.29 [95%CI 0.24-0.36] to 11.42 [95%CI 9.91-13.17]). CONCLUSIONS: Compared with trial-ineligible populations within the same age and sex strata, RCTs recruited people of varying CVD risk but often excluded people at high risk of bleeding or non-CVD death, highlighting that many trials may overestimate the net benefit of aspirin for primary prevention.

Humans

Vegetable and fruit consumption and cancer risk.

The relationship between cancer risk and frequency of consumption of green vegetables and fruit has been analyzed using data from an integrated series of case-control studies conducted in northern Italy between 1983 and 1990. The overall dataset included the following histologically confirmed cancers: oral cavity and pharynx, 119; oesophagus, 294; stomach, 564; colon, 673; rectum, 406; liver, 258; gall-bladder, 41; pancreas, 303; larynx, 149; breast, 2,860; endometrium, 567; ovary, 742; prostate, 107; bladder, 365; kidney, 147; thyroid, 120; Hodgkin's disease, 72; non-Hodgkin lymphomas, 173; myelomas, 117; and a total of 6,147 controls admitted to hospital for acute non-neoplastic conditions, unrelated to long-term dietary modifications. Multivariate relative risks (RR) for subsequent tertiles of vegetable and fruit consumption were derived after allowance for age, sex, area of residence, education and smoking. For vegetables, there was a consistent pattern of protection for all epithelial cancers, with RRs in the upper tertile ranging from 0.2 for oesophagus, liver and larynx to 0.7 for breast. All the trends in risk were in the same direction and significant for all carcinomas except gall-bladder. In contrast, no protection was afforded by high vegetable consumption against non-epithelial lymphoid neoplasms. With reference to fruit, strong inverse relationships were observed for cancers of the upper digestive and respiratory tract, with RRs in the upper tertile between 0.2 and 0.3 for oral cavity and pharynx, oesophagus and larynx relative to the lowest tertile. The lower the location of the tumour in the digestive tract, the weaker appeared to be the protection afforded. Significant inverse relationships were observed for liver, pancreas, prostate and urinary sites, but not for rectum, breast and female genital cancers or thyroid. No relationship emerged for lymphomas and myelomas. Even in the absence of a clear biological interpretation, the consistency and strength of the patterns observed indicate that, in this population, frequent green vegetable intake is associated with a substantial reduction of risk for several common epithelial cancers, and that fruit intake has a favourable effect, especially on upper digestive cancers and, probably, also on urinary tract neoplasms.

Adult

Artificial intelligence-based tumour infiltrating lymphocyte quantification in patients with triple-negative breast cancer: an independent validation study.

BACKGROUND: Tumour-infiltrating lymphocytes (TILs) are a robust prognostic marker in patients with triple-negative breast cancer. Artificial intelligence (AI)-derived computational tools assessing TILs could improve efficiency, but require independent validation against clinical outcomes. We aimed to compare the prognostic performance of AI-derived TIL scores with pathologist-scored TILs in a large, prospectively collected dataset pooled from randomised controlled trials. METHODS: CATALINA was an independent, external validation study using prospectively collected long-term clinical outcome data pooled from seven randomised clinical trials conducted at multiple sites. We independently evaluated two previously validated AI pipelines that generate five computationally assessed tumour-infiltrating lymphocyte (cTIL) scores by masked, independent deployment of locked models. cTIL scores were correlated with the mean of the pathologist-scored stromal TILs (sTILs) in 220 digitised haematoxylin and eosin whole slide images in a cohort of patients with early-stage triple-negative or HER-2 positive breast cancer, previously scored by trained pathologists in a TIL-reproducibility study. Prognostic performance was assessed in a separate cohort of patients with early triple-negative breast cancer pooled from seven prospective, randomised adjuvant trials. Multivariable Cox regression models adjusted for clinicopathological factors and study heterogeneity assessed associations of cTIL score and sTIL score with invasive disease-free survival, distant disease-free survival, and overall survival. 5-year discrimination was estimated using time-dependent area under the receiver operating characteristic curve (AUC). FINDINGS: Individual data were collated from 1759 patients, of whom 1356 had complete clinicopathological data, pathologist sTIL scores, and cTIL scores available. Modest correlation (r 0&#xb7;375-0&#xb7;473) was observed between cTIL scores and the mean pathologist sTIL score. Both sTIL and cTIL were independently associated with 5-year invasive disease-free survival, distant disease-free survival, and overall survival after adjustment for clinicopathological factors (hazard ratio for invasive disease-free survival was 0&#xb7;73 [95% CI 0&#xb7;66-0&#xb7;82]; q<0&#xb7;0001, distant disease-free survival was 0&#xb7;70 [0&#xb7;61-0&#xb7;79]; q<0&#xb7;0001, and overall survival was 0&#xb7;72 [0&#xb7;63-0&#xb7;82]; q<0&#xb7;0001 for sTIL scores and 0&#xb7;80 [0&#xb7;73-0&#xb7;89]; q<0&#xb7;0001, 0&#xb7;77 [0&#xb7;69-0&#xb7;86]; q<0&#xb7;0001, and 0&#xb7;79 [0&#xb7;70-0&#xb7;88]; q=0&#xb7;0002, respectively, for percentage_lymphocyte scores). In models adjusted for clinicopathological variables and sTIL score, cTIL score did not maintain a statistically significant prognostic association. Both sTIL and cTIL scores improved the 5-year AUC over clinicopathological variables alone, while cTIL score did not significantly further improve AUC when combined with clinicopathological variables and sTIL score. INTERPRETATION: Two cTIL models deployed entirely without retraining or modification provided statistically significant prognostic information and improved risk discrimination compared with clinicopathological variables alone in this large, platform-based, independent validation study. Although cTIL score did not incrementally improve prognostication compared with models combining clinicopathological variables with sTIL score, these findings support the application of cTILs as a reproducible prognostic biomarker, particularly in settings where routine or widespread pathologist assessment is unavailable. FUNDING: Breast Cancer Research Foundation (USA).

Humans