PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Random Forest”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Extreme climatic events drive consistent and predictable shifts in soil antibiotic resistance genes.

Antimicrobial resistance (AMR) is a growing One Health challenge, and as climate warming intensifies extreme events, it remains unclear how these disturbances affect soil antibiotic resistance genes (ARGs). Here we analyzed the data from a controlled experiment using soils from 30 grassland sites across ten European countries, which simulated drought, flooding, freeze-thaw, and heatwaves to explore ARG dynamics. Overall, ARGs exhibited relatively small but highly consistent shifts across treatments. Heatwaves caused the strongest reductions in ARG abundance and in their linkages with mobile genetic elements (MGEs), a pattern that may reflect a hypothesized metabolic-genetic trade-off, in which microbial investment may shift from core metabolism toward stress signaling and structural maintenance. ARG dynamics during and after disturbance were governed by distinct soil physicochemical properties, with temperature and nutrient status determining acute responses, whereas soil moisture and seasonal variability in temperature and precipitation shaped longer-term legacy effects. Cross-validated random-forest models showed positive predictive performance for Bray-Curtis-based compositional responses within the environmental range represented by the 30 grassland sites. Our findings enhance the understanding of how soil ARGs respond to extreme climatic events and provide a step toward predicting extreme-event impacts on soil resistomes with relevance to One Health.

Soil Microbiology↗

Improved classification of mass spectrometry database search results using newer machine learning approaches.

Manual analysis of mass spectrometry data is a current bottleneck in high throughput proteomics. In particular, the need to manually validate the results of mass spectrometry database searching algorithms can be prohibitively time-consuming. Development of software tools that attempt to quantify the confidence in the assignment of a protein or peptide identity to a mass spectrum is an area of active interest. We sought to extend work in this area by investigating the potential of recent machine learning algorithms to improve the accuracy of these approaches and as a flexible framework for accommodating new data features. Specifically we demonstrated the ability of boosting and random forest approaches to improve the discrimination of true hits from false positive identifications in the results of mass spectrometry database search engines compared with thresholding and other machine learning approaches. We accommodated additional attributes obtainable from database search results, including a factor addressing proton mobility. Performance was evaluated using publically available electrospray data and a new collection of MALDI data generated from purified human reference proteins.

Amino Acid Sequence↗

Discovery of novel diagnostic biomarkers of hepatocellular carcinoma associated with immune infiltration.

OBJECTIVE: Diagnosis of hepatocellular carcinoma (HCC) remains challenging for clinicians. Machine learning approaches and big data analyses are viable strategies for identifying HCC diagnostic markers. MATERIALS AND METHODS: In this study, we downloaded mRNA expression profiles of HCC from the GEO database and used random forest and machine learning algorithms, such as least absolute shrinkage and selection operator, to screen for reliable diagnostic genes. Disease Ontology, Kyoto Encyclopedia of Genes and Genomes (KEGG) and Gene Set Enrichment Analysis enrichment analyses were performed to explore differential gene functions and disease pathways. CIBERSORT was performed to calculate the immune cell infiltration of HCC and the correlation between diagnostic genes and immune cells. Cell experiments were performed to evaluate the function of R-spondin 3 (RSPO3) in HCC cells. Immunohistochemical staining was used to evaluate the protein expression of CD138, CD206 and iNOS. RESULTS: The results indicated that extracellular matrix protein 1 (ECM1), Niemann-Pick C1-Like 1 (NPC1L1) and RSPO3 were down-regulated in HCC compared with the normal group (p&#x2009;<&#x2009;0.05), which was validated in clinical tissue samples. Moreover, ECM1, NPC1L1 and RSPO3 had high diagnostic values (AUC > 0.75) for HCC in both training and test groups. Immuno-infiltration analysis revealed that ECM1 and RSPO3 were highly positively correlated with neutrophil and macrophage M2 levels, whereas they were negatively correlated with Tregs. RSPO3-si affected cell proliferation and apoptosis in HCC. Furthermore, RSPO3 exhibited a positive correlation with tumour progression, the proportion of plasma cells and M2 macrophages in mice, while showing a negative association with M1 macrophages. CONCLUSION: The present study identified ECM1, NPC1L1 and RSPO3 as new diagnostic biomarkers for HCC based on normal and diseased samples from HCC, meanwhile the pro-oncogenic function of RSPO3 and its regulation on immune infiltration have been confirmed.

Carcinoma, Hepatocellular↗

Dynamic lysine acetylation and succinylation of platelet proteins regulates platelet storage lesion: mechanistic insights from multi-omics.

OBJECTIVES: Platelet storage lesion (PSL) severely impairs platelet function during storage, presenting a major hurdle in transfusion medicine; however, the dynamic interplay between global proteomic changes and post-translational modifications (PTMs) underlying these functional deteriorations remains insufficiently characterized. Here, we report the first comprehensive multi-omics analysis integrating global proteomics, acetylomics, and succinylomics to dissect the molecular dynamics during platelet storage. METHODS: We performed quantification of global proteomics, acetylome and succinylome based on TMT-labeled LC-MS/MS analysis, combined with antibody-affinity enrichment and purification. Dynamic molecular changes and functional transformation of platelet were also characterized under proper conditions stored for 1, 3, 5, 7&#x2009;days, respectively. RESULTS: We systematically characterized 3,609 proteins, 1,308 acetylation sites, and 1,947 succinylation sites across multiple storage time points (D1, D3, D5, D7). We distinct temporal patterns of post-translational modifications, with succinylation showing more extensive coverage than acetylation in platelets. Pathway enrichment analysis revealed extensive metabolic reprogramming involving complement activation, energy metabolism, and cellular detoxification processes. The identification of specific motif patterns provided mechanistic insights into the functional specificity of these modifications. Random forest machine learning identified 20 core regulatory proteins representing critical nodes in PSL development. Furthermore, we employed real - time quantitative polymerase chain reaction (RT - QPCR) to measure the expression levels of key genes related to platelet function and PTM - associated pathways. CONCLUSION: By mapping the interplay between proteomic abundance shifts and PTM dynamics, this study provides a multidimensional understanding of PSL, establishing a foundational framework for optimizing storage protocols and enhancing transfusion safety.

Blood Platelets↗

Comparison of methods for chemical-compound affinity prediction.

The selection of effective features from various descriptors of chemical compounds and the exploitation of the most appropriate classifier is a momentous issue in improving overall accuracies of virtual screening of chemical compounds. In this article, the performance of various feature-selection methods and various classifiers of chemical compound-protein binding affinities are compared by using six series of compounds: cytochrome P450 2C9 inhibitors, multi-drug-resistance reversal compounds, estrogen receptor ligands, inhibitors of human ether-a-go-go-related genes, and ligands of serotonin receptor 5HT1A and 5HT2A. As a result, it was found that the genetic algorithm was superior to the other feature-selection methods, and its combination with Random Forests and Adaboosts or Baggings gave almost the same performance as support-vector machines and was superior to the other classifiers. The precision and recall of these methods were almost the same or ascendant to those of previous work. The automatically selected descriptors for each protein-compound affinity prediction were plausible and would be informative to interpret the resulting model.

Algorithms↗

Blood-based DNA methylation markers for autism spectrum disorder identification using machine learning.

BACKGROUND: Autism spectrum disorder (ASD) is a complex neurodevelopmental disorder lacking objective biomarkers for early diagnosis. DNA methylation is a promising epigenetic marker, and machine learning offers a data-driven classification approach. However, few studies have examined whole-blood, genome-wide DNA methylation profiles for ASD diagnosis in school-aged children. METHODS: We analyzed genome-wide DNA methylation data from GEO dataset GSE113967, including 52 children with ASD and 48 typically developing (TD) controls. Differentially methylated positions (DMPs) were identified, and feature selection was performed using support vector machine-recursive feature elimination with cross-validation (SVM-RFECV). Classification models were developed using random forest (RF), extreme gradient boosting (XGBoost), and decision tree (DT) classifiers. A nomogram visualized feature contributions. RESULTS: A total of 138 DMPs differentiated ASD from TD children. Eleven CpG sites selected by SVM-RFECV formed the basis for model construction. RF and XGBoost achieved the highest accuracy (75%), with DT reaching 70%. Functional annotation indicated enrichment in cell adhesion and immune-related pathways. CONCLUSIONS: This exploratory study demonstrates the feasibility of integrating peripheral blood DNA methylation data with machine learning to distinguish children with ASD. While limited by sample size and moderate accuracy, this study provides methodological insights into the feasibility of integrating epigenetic and computational approaches for ASD-related biomarker exploration.

Humans↗

MicroRNAs signatures in small extracellular vesicles for psychological resilience in young adults using machine learning.

AIMS: Psychological resilience refers to an individual's capacity to adapt to adverse events. MicroRNAs (miRNAs) play a crucial role in regulating post-transcriptional processes, while small extracellular vesicles (sEVs) act as transport vehicles. This study aimed to employ genome-wide profiling to identify and validate differences in the expression of resilience-associated sEV-miRNAs between low resilience (LR) and high resilience (HR) in young adults. METHODS: Eighty participants were divided into LR or HR based on the Connor - Davidson Resilience Scale (CD-RISC). The expression levels of the target sEV-miRNAs in LR and HR were compared and analyzed. RESULTS: Expression analyses demonstrated significant differences in let-7b, miR-151b, miR-335, and miR-193a between LR and HR (p&#x2009;<&#x2009;0.01), with let-7b showing the highest discriminative ability. The AUC values for each sEV-miRNA ranged from 0.74 to 0.94, based on logistic regression and three machine learning models: random forest, support vector machine, and eXtreme gradient boosting. Based on leave-one-out cross-validation in different models, the combined four sEV-miRNAs demonstrated strong performance for detecting LR (AUC&#x2009;=&#x2009;0.87-0.90). Sex-specific differences were also observed, with female participants showing more pronounced resilience signatures in targeted sEV-miRNAs. CONCLUSIONS: These findings suggest that sEV-miRNAs hold potential as biomarkers for psychological resilience in young adults.

Humans↗

Integrating explainable artificial intelligence with multiomics systems biology and electronic health record data mining for personalized drug repurposing in Alzheimer's disease.

Alzheimer's disease (AD) is characterized by region- and patient-specific molecular heterogeneity, which hinders therapeutic design. In this study, we introduce PRISM-ML (PRecision-medicine using Interpretable Systems and Multiomics with Machine Learning), an open-source integrated analysis pipeline that combines interpretable machine learning with systems biology and electronic health records data mining to elucidate the molecular diversity of AD and predict promising drug repurposing opportunities. First, we integrated and harmonized transcriptomic (bulk RNA-seq) and genomic (genome-wide association study) data from 2105 brain samples, each with matched data from the same individual (1363 AD patients, 742 controls; 9 tissues), sourced from three independent studies. Random forest classifiers with SHapley Additive exPlanations identified patient-specific biomarkers; unsupervised clustering resolved 36 molecularly distinct subtissues (defined as clusters of samples within a brain tissue that share a specific expression pattern); and gene-gene coexpression networks prioritized 262 high-centrality bottleneck genes as putative regulators of dysregulated pathways. Next, knowledge graph-based drug repurposing predicted six Food and Drug Administration (FDA)-approved drugs that simultaneously target multiple bottleneck genes and multiple AD-relevant pathways. Notably, in a large US de-identified insurance-claims database (n&#x2009;=&#x2009;364&#xa0;733), exposure to promethazine, one of the candidate drugs, was associated with a 57%-62% lower incidence of AD versus an active antihistamine comparator (adjusted hazard ratio 0.38; inverse-probability weighted 0.43; both P&#x2009;<&#x2009;.001), providing real-world support for its repurposing potential. In summary, PRISM-ML, as an explainable multiomics analysis pipeline, is readily transferable to other complex diseases, advancing precision medicine.

Alzheimer Disease↗

A leakage-aware genomic prediction pipeline for meropenem resistance in Klebsiella pneumoniae using transformer-based resistome representation learning.

MOTIVATION: Antimicrobial resistance (AMR) in Klebsiella pneumoniae, particularly to carbapenems such as meropenem, is a major global health problem. Machine learning is increasingly used to predict resistance from genomic markers; however, many models fail to capture high-level gene-gene interactions and may exhibit inflated performance due to lineage-biased prediction. Existing genomic prediction models largely rely on flat feature representations that fail to capture epistatic gene interactions, and commonly suffer from inflated performance estimates due to phylogenetic data leakage. To address these limitations simultaneously, a leakage-aware hybrid TabTransformer-CatBoost pipeline was developed, combining self-attention-based resistome representation learning with gradient boosting classification under clade-aware data partitioning. A self-attention encoder converts sparse gene presence-absence profiles into contextualized latent embeddings, which are subsequently classified using gradient boosting to capture lineage-aware AMR patterns. RESULTS: The proposed architecture outperformed classical baselines including Logistic Regression, Random Forest, XGBoost, and optimized CatBoost models. Internal accuracy reached 92.59% for the Chained Hybrid configuration (area under the receiver operating characteristic curve, AUROC = 0.8670, F1&#x2009;=&#x2009;0.8537). Performance gains primarily originated from the embedding stage, as confirmed by ablation analysis. External validation across independent multinational cohorts (n&#x2009;=&#x2009;305) demonstrated generalizability (AUROC = 0.8105; F1&#x2009;=&#x2009;0.7552). Permutation testing produced near-zero Matthews Correlation Coefficient (MCC)&#x2009;=&#x2009;0.0091, indicating predictions reflect genuine biological signal rather than noise. These results establish attention-based genomic embedding with gradient boosting as a scalable, interpretable, and leakage-aware framework for clinical AMR prediction. AVAILABILITY AND IMPLEMENTATION: The source code for the TabTransformer-CatBoost framework, including preprocessing pipelines and pre-trained embeddings, is available at https://github.com/SibelKervanci/kp-meropenem-tabtransformer.

Journal Article↗

plinkQC: an integrated tool for ancestry inference, sample selection, and quality control in population genetics.

MOTIVATION: Population genetic analyses rely on high quality datasets that pass rigorous controls for sample and marker quality. Many analyses also require additional processing including identification of ancestry and sample relatedness. A software package that addresses all these common, yet crucial tasks is missing. RESULTS: We have developed plinkQC, an R/CRAN package that combines these functionalities into a single software package with detailed vignettes for example applications. plinkQC determines the ancestry of study samples via a pre-trained random forest classifier that reaches 98% performance accuracy with just 5% of marker overlap between reference and user data. To obtain the maximal set of unrelated study samples, we developed a graph-based pruning method, taking both relationship estimates and sample quality into account. We demonstrate optimal sample selection on the 1000 Genomes project, where we retain an additional 71 samples compared to publicly available exclusion lists. Finally, plinkQC bundles these results together with per-individual and per-marker quality control checks into three simple functions and returns both the quality controlled dataset and quality control report about each step of the analysis. AVAILABILITY AND IMPLEMENTATION: plinkQC is available as an R/CRAN package. The documentation and code are available on github: https://meyer-lab-cshl.github.io/plinkQC/ and https://github.com/meyer-lab-cshl/plinkQC_manuscript.

Software↗

Comparison of statistical methods for classification of ovarian cancer using mass spectrometry data.

MOTIVATION: Novel methods, both molecular and statistical, are urgently needed to take advantage of recent advances in biotechnology and the human genome project for disease diagnosis and prognosis. Mass spectrometry (MS) holds great promise for biomarker identification and genome-wide protein profiling. It has been demonstrated in the literature that biomarkers can be identified to distinguish normal individuals from cancer patients using MS data. Such progress is especially exciting for the detection of early-stage ovarian cancer patients. Although various statistical methods have been utilized to identify biomarkers from MS data, there has been no systematic comparison among these approaches in their relative ability to analyze MS data. RESULTS: We compare the performance of several classes of statistical methods for the classification of cancer based on MS spectra. These methods include: linear discriminant analysis, quadratic discriminant analysis, k-nearest neighbor classifier, bagging and boosting classification trees, support vector machine, and random forest (RF). The methods are applied to ovarian cancer and control serum samples from the National Ovarian Cancer Early Detection Program clinic at Northwestern University Hospital. We found that RF outperforms other methods in the analysis of MS data.

Algorithms↗

Modeling the relationship between LVAD support time and gene expression changes in the human heart by penalized partial least squares.

MOTIVATION: Heart failure affects more than 20 million people in the world. Heart transplantation is the most effective therapy, but the number of eligible patients far outweighs the number of available donor hearts. The left mechanical ventricular assist device (LVAD) has been developed as a successful substitution therapy that aids the failing ventricle while a patient is waiting for the donor heart. We obtained genomics data from paired human heart samples harvested at the time of LVAD implant and explant. The heart failure patients in our study were supported by the LVAD for various periods of time. The goal of this study is to model the relationship between the time of LVAD support and gene expression changes. RESULTS: To serve the purpose, we propose a novel penalized partial least squares (PPLS) method to build a regression model. Compared with partial least squares and Breiman's random forest method, PPLS gives the best prediction results for the LVAD data.

Adaptation, Physiological↗

Standardization and denoising algorithms for mass spectra to classify whole-organism bacterial specimens.

MOTIVATION: Application of mass spectrometry in proteomics is a breakthrough in high-throughput analyses. Early applications have focused on protein expression profiles to differentiate among various types of tissue samples (e.g. normal versus tumor). Here our goal is to use mass spectra to differentiate bacterial species using whole-organism samples. The raw spectra are similar to spectra of tissue samples, raising some of the same statistical issues (e.g. non-uniform baselines and higher noise associated with higher baseline), but are substantially noisier. As a result, new preprocessing procedures are required before these spectra can be used for statistical classification. RESULTS: In this study, we introduce novel preprocessing steps that can be used with any mass spectra. These comprise a standardization step and a denoising step. The noise level for each spectrum is determined using only data from that spectrum. Only spectral features that exceed a threshold defined by the noise level are subsequently used for classification. Using this approach, we trained the Random Forest program to classify 240 mass spectra into four bacterial types. The method resulted in zero prediction errors in the training samples and in two test datasets having 240 and 300 spectra, respectively.

Algorithms↗

Prediction of the phenotypic effects of non-synonymous single nucleotide polymorphisms using structural and evolutionary information.

MOTIVATION: There has been great expectation that the knowledge of an individual's genotype will provide a basis for assessing susceptibility to diseases and designing individualized therapy. Non-synonymous single nucleotide polymorphisms (nsSNPs) that lead to an amino acid change in the protein product are of particular interest because they account for nearly half of the known genetic variations related to human inherited diseases. To facilitate the identification of disease-associated nsSNPs from a large number of neutral nsSNPs, it is important to develop computational tools to predict the phenotypic effects of nsSNPs. RESULTS: We prepared a training set based on the variant phenotypic annotation of the Swiss-Prot database and focused our analysis on nsSNPs having homologous 3D structures. Structural environment parameters derived from the 3D homologous structure as well as evolutionary information derived from the multiple sequence alignment were used as predictors. Two machine learning methods, support vector machine and random forest, were trained and evaluated. We compared the performance of our method with that of the SIFT algorithm, which is one of the best predictive methods to date. An unbiased evaluation study shows that for nsSNPs with sufficient evolutionary information (with not <10 homologous sequences), the performance of our method is comparable with the SIFT algorithm, while for nsSNPs with insufficient evolutionary information (<10 homologous sequences), our method outperforms the SIFT algorithm significantly. These findings indicate that incorporating structural information is critical to achieving good prediction accuracy when sufficient evolutionary information is not available. AVAILABILITY: The codes and curated dataset are available at http://compbio.utmem.edu/snp/dataset/

Algorithms↗

Accurate identification of abnormal ploidy using an artificial intelligence model in preimplantation genetic testing.

STUDY QUESTION: Can ultra-low-coverage whole-genome sequencing (ulc-WGS) accurately identify abnormal ploidy during preimplantation genetic testing (PGT)? SUMMARY ANSWER: The artificial intelligence (AI)-based PGT-Plus model demonstrates high accuracy in ploidy detection, offering a cost-effective solution that enhances clinical utility of PGT. WHAT IS KNOWN ALREADY: The predominant PGT for aneuploidy can identify chromosomal aneuploidies but cannot determine ploidy status. Transferring embryos with ploidy abnormalities can result in miscarriage and molar pregnancy. On the other hand, in ART, fertilization is assessed by morphological pronuclear assessment at the zygote stage. However, it has a low specificity in the prediction of abnormal ploidy status and embryos deemed abnormally fertilized can yield healthy pregnancies. Accurately identified abnormal ploidy in PGT-A can resolve current limitations and expand the utility range of PGT-A. Several studies have identified ploidy abnormalities; however, they were mainly based on single-nucleotide polymorphism (SNP) arrays or needed to combine additional targeted-next-generation sequencing (NGS) information. Studies based on ulc-WGS remain scarce. STUDY DESIGN SIZE DURATION: The study consisted of two stages: methodology establishment and validation. An AI model, named PGT-Plus, was developed using 653 samples with known ploidy status, which was further validated using 792 different ploidy status samples. In the clinical application stage, the approach was used to analyse the ploidy status of 19&#x2009;103 normally fertilized PGT blastocysts and 140 single pronucleus (1PN)-derived blastocysts collected between May 2022 and December 2023. All blastocysts were tested using trophectoderm biopsy and NGS. PARTICIPANTS/MATERIALS SETTING METHODS: The methodology is based on the ulc-WGS data. First, based on samples with known ploidy status: the heterozygosity rate of high-frequency biallelic SNPs, the likelihood ratio (LLR) of alleles was calculated under different assumptions ('both parental homologs' [BPH] from a single parent, 'single parental homolog' [SPH] from each parent, disomy, and monosomy) by leveraging allele frequencies and linkage disequilibrium (LD) measured in the 1000 genomes project database. Twenty-three continuous candidate features derived from heterozygosity rates and LLRs of chromosomes or selected windows were included to establish the ploidy prediction AI model. Gini importance analysis and multicollinearity mitigation was performed for feature selection, then the performance of Random Forest (RF), Support Vector Machine (SVM), and Logistic Regression for modelling was compared. Subsequently, the parameter optimization was performed based on the RF model. Ploidy constitution concordance was evaluated in known ploidy status samples. The frequency of abnormal ploidy in normal fertilized PGT blastocysts and 1PN-derived blastocysts (including conventional IVF and ICSI) was evaluated. MAIN RESULTS AND THE ROLE OF CHANCE: Eleven features were collected for model architecture compared to SVM and Logistic Regression; RF achieved superior performance for ploidy detection. The AI model achieved an AUC of 1 for genome-wide-uniparental diploidy (GW-UPD), 1 for triploidy, and 0.99 for diploidy. For the 792 validation samples, 99.5% of samples were successfully detected using the AI model, and the model showed 100% accuracy for ploidy classification. In the clinical application stage, out of 19&#x2009;103 PGT samples, 19&#x2009;069 were successfully analysed using the model, with 110 (0.57%) identified as having abnormal ploidy embryos. Among these, 12.7% (14/110) were identified as GW-UPD, and 87.3% (96/110) were triploid. Among 5563 diploid blastocysts transferred, 3478 clinical pregnancies were achieved. Subsequent ploidy analysis was performed for 217 spontaneous abortion and 935 prenatal diagnostic samples, and no abnormal ploidy was identified. Furthermore, of the 140 1PN embryos tested, 40 (28.6%) exhibited GW-UPD, 3 (2.1%) exhibited triploidy, and 97 (69.3%) were determined to be biparental and normally fertilized. Among the 97 biparental embryos, 46 were diploid, 11 were mosaic, and 40 were aneuploid. In terms of the insemination pattern, the percentage of abnormal ploidy in ICSI was significantly higher than in conventional IVF (P&#x2009;<&#x2009;0.01, 37.1% vs. 2.9%, respectively). With full informed consent, 20 patients without euploidy from normal fertilization chose 1PN-derived biparental and diploid blastocysts to transfer, resulting in 10 clinical pregnancies and 9 ongoing pregnancies. LARGE-SCALE DATA: N/A. LIMITATIONS REASONS FOR CAUTION: Some rare ploidy abnormalities, such as polyploidy with an equal number of identical sets of chromosomes and ploidy mosaicism cannot be accurately identified. Moreover, the origin of abnormal ploidy was not identified due to the unavailability of DNA from both parents. WIDER IMPLICATIONS OF THE FINDINGS: The PGT-Plus AI model provides a ploidy evaluation method based on the conventional PGT-A data and integrates directly into standard PGT-A workflows. Clinical utility results suggest that the model is a valuable tool for identifying embryos with abnormal ploidy in PGT-A and rescuing normal diploid embryos from abnormally fertilized embryos. These findings demonstrate that PGT-Plus significantly enhances the diagnostic accuracy of PGT. STUDY FUNDING/COMPETING INTERESTS: This study was supported by grants from Major Scientific Program of CITIC Group (No. 2023ZXKYB34100, to Ge.L.), Hunan Provincial Grant for Innovative Province Construction (2019SK4012), Hunan Xiangjiang New District (Changsha High-tech Zone) key core technology research project in 2023, and Science Foundation of Hunan Province (Grant 2023JJ30422). All authors declared no conflicts of interest..

artificial intelligence↗

nsSNPAnalyzer: identifying disease-associated nonsynonymous single nucleotide polymorphisms.

Nonsynonymous single nucleotide polymorphisms (nsSNPs) are prevalent in genomes and are closely associated with inherited diseases. To facilitate identifying disease-associated nsSNPs from a large number of neutral nsSNPs, it is important to develop computational tools to predict the nsSNP's phenotypic effect (disease-associated versus neutral). nsSNPAnalyzer, a web-based software developed for this purpose, extracts structural and evolutionary information from a query nsSNP and uses a machine learning method called Random Forest to predict the nsSNP's phenotypic effect. nsSNPAnalyzer server is available at http://snpanalyzer.utmem.edu/.

Algorithms↗

Identifying JAK2 and ANXA5 as Key Genes Linking Obstructive Sleep Apnea and Oxidative Stress via Machine Learning and Multilayer Transcriptomic Integration With Functional Validation.

Obstructive sleep apnea (OSA) is a common and severe sleep disorder closely associated with oxidative stress (OS). This study aims to identify and validate potential OS-related genes associated with OSA through bioinformatics methods. We successfully identified OS-related differentially expressed genes (OS-DEGs) by combining the limma test, weighted correlation network analysis (WGCNA), and OS-related genes from the GeneCards database. Key genes and potential biological roles were further identified using Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG), enrichment analysis, protein-protein interaction (PPI) network analysis, Lasso regression analysis, random forest algorithm, and support vector machine recursive feature elimination (SVM-RFE) method. Evaluate and validate the accuracy of key genes through receiver operating characteristic (ROC) curve analysis. The human single-cell RNA sequencing (scRNA-seq) dataset is used for cell classification annotation, analysis of key gene single-cell expression profiles, and virtual gene knockout experiments based on the scTenifoldKnk algorithm. Integrating scRNA-seq sequencing, pseudotime trajectory inference, cell-cell communication analysis, and bulk immune infiltration deconvolution reveals monocyte subtype remodeling in OSA. Finally, the expression levels of key genes in clinical samples were validated using real-time quantitative PCR (RT-qPCR) and Western blotting. A total of 57 common DEGs, indicating significant enrichment in OS, inflammation, and tumor pathways, particularly prominent in the immunometabolism pathway. By integrating DEGs, WGCNA, PPI results, and machine learning methods, key genes Janus kinase 2 (JAK2) and ANXA5 were screened out. JAK2 was significantly upregulated under disease conditions, while ANXA5 was significantly downregulated. ROC curve exhibited high accuracy (area under the curve [AUC] >&#x2009;0.85). Human scRNA-seq analysis revealed that key genes were predominantly highly expressed in monocytes. Virtual knockout experiments demonstrated that these key genes play a crucial role in regulating immune responses and inflammatory reactions. PPI networks and enrichment analysis verified that downstream genes S100P, ALOX5AP, PROK2, and PADI4 may collaboratively participate in immune response and inflammation regulation. Finally, clinical sample experiment further validated the results of bioinformatics analysis. This study provides new research insights for the diagnosis, mechanism research, and treatment development of OSA in the future by integrating multilayer transcriptomic and machine learning techniques.

Humans↗

Evaluation of the Neurobehavioral Functioning Inventory as a depression screening tool after traumatic brain injury.

OBJECTIVE: To examine the utility of the Neurobehavioral Functioning Inventory (NFI) for diagnosing depression in a rehabilitation setting. DESIGN: In a prospective study, a structured clinical interview (Structured Clinical Interview for DSM-IV-TR) was used to identify DSM-IV-defined major depressive disorder (MDD) symptoms among patients with traumatic brain injury (TBI). NFI Depression scale items were compared with DSM-IV diagnosis obtained by the Structured Clinical Interview for DSM-IV Axis I Disorders. SETTING: Outpatient neuropsychology clinic at a university hospital, private outpatient physical medicine and rehabilitation clinic, and a long-term specialized living assistance program. PARTICIPANTS: Participants consisted of 78 patients with TBI who were at least 3 months postinjury and 18 years of age or older. MAIN OUTCOME MEASURES: Structured Clinical Interview for DSM-IV Axis I Disorders and the NFI. RESULTS: Psychiatric diagnostic interview with the Structured Clinical Interview for DSM-IV Axis I Disorders indicated that 50% of patients with TBI in our sample had at least one of the following in their lifetime: MDD, MDD due to general medical condition, dysthymia, or adjustment disorder with depressed mood. Thirty percent met diagnostic criteria for current MDD with or without general medical condition. Analyses of the NFI items revealed that individuals with depression endorsed greater levels of problems than did those without depression on 14 of the 32 items related to the DSM-IV symptom domains for depression (P < .00156 with Bonferroni correction). In predicting the diagnosis of depression using individual NFI items, the classification rate based on the Random Forests estimate was 83%. CONCLUSION: Findings indicate that the NFI items differentiated between depressed and nondepressed patients with TBI. Imposing minimal burden on patients and staff, the NFI appears to have good predictive value in diagnosing major depression. In clinical practice and research, the NFI is a potentially valuable screening tool for identifying major depression in persons with TBI.

Adolescent↗