PubMed HealthSearch

SEARCH · PubMed Health

Results for “representation learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

CAKR: commutative algebra k-mer representations for genomics.

Despite the availability of various sequence analysis models, comparative genomic analysis remains a challenge in genomics, genetics, and phylogenetics. Commutative algebra, a fundamental tool in algebraic geometry and number theory, has rarely been used in data and biological sciences. In this study, we introduce commutative algebra k-mer representations as a nonlinear algebraic framework for analyzing genomic sequences. This representation bridges commutative algebra, algebraic topology, combinatorics, and machine learning to establish a mathematical framework for comparative genomic analysis. We evaluate its effectiveness on three tasks including genetic variant classification, phylogenetic tree reconstruction, and viral classification, typically requiring alignment-based, alignment-free, and machine-learning approaches, respectively. In this work, we show that commutative algebra k-mer representations outperform five state-of-the-art sequence analysis methods across twelve primary datasets, with two additional supplementary fragment-placement benchmarks, especially in viral classification, and maintain relatively stable predictive accuracy as dataset size increases, underscoring scalability and robustness.

Genomics

Classification of crossed immunoelectrophoretic patterns using digital image processing and artificial neural networks.

A method is presented which makes it possible to present crossed immunoelectrophoretic patterns to an artificial neural network. The electrophoretic patterns are presented for the artificial neural network as three-dimensional vectors and it is shown that it is possible with this representation to train the network to learn the patterns and classify them. It was found that the ability to generalize was substantially increased by the addition of noise to the input patterns during training. Furthermore, the addition of noise decreased the number of presentations needed to reach the predetermined error level. The trained neural network was able to classify all distorted patterns correctly within an error range of 1%.

Image Processing, Computer-Assisted

Maladaptive anticipatory saccades in schizophrenia.

We compared the saccades made by 8 neuroleptic-treated and 7 drug-free schizophrenic inpatients with those made by 11 normal controls during two eye movement tasks. The first task was designed to elicit visually guided but not internally guided saccades. The second task was designed so that optimal performance required saccades be guided on the basis of an internal representation of target behavior. During the first task, schizophrenics made visually guided saccades that were as accurate as those made by control, but both drug-free and neuroleptic-treated schizophrenics made intrusive saccades at a significantly higher rate than control subjects. Most of these maladaptive saccades appeared to be premature attempts to anticipate target jump. During the second eye movement task, which for optimal performance required use of an internal representation to guide eye movements, most patients learned to anticipate target jump as well as controls. However, neuroleptic-treated patients made significantly smaller adaptive anticipatory saccades than either drug-free schizophrenic patients or normal subjects. These finding are discussed as they relate to the prefrontal cortex-basal ganglia circuits involved in the regulation of behavior by representational knowledge and the idea that the abnormal anticipatory saccades we observed represent a failure in the sensorimotor gating of information derived from internal representations.

Adult

MO-GCAN: multi-omics integration based on graph convolutional and attention networks.

MOTIVATION: Cancer subtypes play a critical role in disease progression, prognosis, and treatment, making their detection essential for tailoring precision medicine. Studies have shown that multi-omics integration outperforms single-omics approaches in cancer subtyping tasks. However, due to the high-dimensionality of multi-omics data, many existing studies either fail to capture the correlation between true labels and learned features, or lack sufficient capacity to model complex biological representations. These limitations hinder the full potential of leveraging the rich and complementary information embedded in multi-omics datasets. RESULT: We propose a framework that leverages supervised feature learning and classification based on a graph-based learning approach with attention mechanism for cancer subtyping. More specifically, we train graph convolutional network models on each omics dataset to extract latent representations, which are then concatenated to form a comprehensive multi-omics feature embedding. We further develop sample fusion network based on the omics-specific graphs, incorporating the derived features and feeding them into a graph attention model for subtype classification. This two-stage multi-omics framework is applied to eight cancer types, with performance evaluated in terms of test accuracy, training time, macro-averaged precision, recall, and F-score. Experimental results show that the proposed method outperforms state-of-the-art approaches across various cancer types. Additionally, we provide empirical evidence supporting the hypothesis that retaining a limited number of high-confidence edges and utilizing enriched embeddings from intermediate graph neural network layers can improve predictive performance. AVAILABILITY AND IMPLEMENTATION: Data and the code are available at https://github.com/YD-00/MO-GCAN-Updated.git.

Neoplasms

Stimulus representation: a subprocess of imprinting and conditioning.

We suggest a way to reconcile imprinting and associative learning that respects the real differences between the two phenomena but helps to recognize underlying commonalities. Rather than treating each type of learning as the manifestation of a unitary mechanism, we approach learning as a combination of separate subprocesses. Exploration of the literature regarding one of these subprocesses, namely, that governing the representation of stimuli, revealed striking similarities between imprinting and conditioning. These similarities suggest predictions for fresh experimental work that will help to uncover the general rules by which combinations of stimulus features are represented in memory.

Animals

Tomtom-lite: accelerating Tomtom enables large-scale and real-time motif similarity scoring.

SUMMARY: Pairwise sequence similarity is a core operation in genomic analysis, yet most attention has been given to sequences made up of discrete characters. With the growing prevalence of machine learning, calculating similarities for sequences of continuous representations, e.g. frequency-based position-weight matrices (PWMs) and attribution-based contribution-weight matrices, is taking on newfound importance. Tomtom has previously been proposed as an algorithm for identifying pairs of PWMs whose similarity is statistically significant, but the implementation remains inefficient for both real-time and large-scale analysis. Accordingly, we have re-implemented Tomtom as a numba-accelerated Python function that is natively multi-threaded, avoids cache misses, more efficiently caches intermediate values, and uses approximations at compute bottlenecks. Here, we provide a detailed description of the original Tomtom method and present results demonstrating that our re-implementation can achieve over a 1000-fold speedup compared with the original tool on reasonable tasks. AVAILABILITY AND IMPLEMENTATION: Our implementation of Tomtom is freely available as a Python package at https://github.com/jmschrei/memesuite-lite, which can be downloaded via pip install memelite or at https://zenodo.org/records/17008952.

Software

A full review of online education resources available on antifungal stewardship.

BACKGROUND AND OBJECTIVES: Antifungal resistance represents an increasing global threat, driven by the rising burden of fungal disease. Antifungal stewardship (AFS) is a critical component of broader antimicrobial resistance (AMR) efforts, but education in this area remains less established than antibacterial stewardship initiatives. The scope and characteristics of the current landscape of online AFS resources have not yet been systematically described. To identify and evaluate online educational resources focused on fungal disease management and AFS, and assess their accessibility, format, educational design and implementation focus. METHODS: A structured search of internet search engines, distribution platforms and organizational websites was conducted to identify English-language web-based resources related to fungal disease management and stewardship. Resources were evaluated using predefined criteria including access model, format, length, educational design, interactivity and AFS content. An overall educational value score (1-10) was assigned. RESULTS: Twenty-three educational resources were identified. Most were delivered as online unfacilitated courses (11, 48%) and were short (<4&#x2005;h) (12, 52%). Most focused on guidelines and syndromic management (18, 78%) and targeted doctors and/or nurses/midwives (22, 96%). Limited interactivity was reported in nine (39%) courses. Five courses (22%) had either a substantial or comprehensive focus on AFS. CONCLUSIONS: Online AFS educational resources are available and support awareness and knowledge development. However, they remain relatively few in number. Greater emphasis on implementation-focused learning, behaviour change components and broader global representation may enhance their impact.

Journal Article

Nonverbal visual short-term memory as a function of age and dimensionality in learning-disabled children.

A serial recognition task was used to compare performance of 2 learning disability age groups with 2- and 3-dimensional representations of nonlabeled 8-point random shapes. Age-related increases in short-term memory (STM) performance for both dimensions were found. No significant differences were found between 2- or 3-dimensional stimuli. Contrary to reports of STM performance with normal children, learning-disabled children showed no primacy effect for the 2-dimensional treatment, and second choices were not consistently correct when the first choice was incorrect, These findings were interpreted according to Flavell's notions of mediational inefficiencies.

Age Factors

Cognitive mechanisms of face processing.

Evidence from natural and induced errors of face recognition, from the effects of different cues on resolving errors, and from the latencies to make different decisions about seen faces, all suggest that familiar face recognition involves a fixed, invariant sequence of stages. To recognize a familiar face, a perceptual description of a seen face must first activate a long-standing representation of the appearance of the face of the familiar person. 'Semantic' knowledge about such things as the person's occupation and personality are accessed next, followed, in the final stage, by the name. Certain factors affect the ease of familiar face recognition. Faces seen in the recent past are recognized more readily (repetition priming), as are distinctive faces, and faces preceded by those of related individuals (associative priming). Our knowledge of these phenomena is reviewed for the light it can shed upon the mechanisms of face recognition. Four aspects of face recognition--graded similarity effects and part-to-whole completion in repetition priming, prototype extraction with simultaneous retention of information about individual exemplars, and distinctiveness effects in classification and identification--are proposed as being compatible with distributed memory accounts of cognitive representations.

Association Learning

Leveraging Interradiomic Feature Relationships for Enhanced Prediction of Distant Metastasis and Characterization of Heterogeneity in Head and Neck Cancer.

PURPOSE: Distant metastasis remains a major cause of treatment failure in head and neck (HN) cancer, highlighting the need for more accurate early risk stratification. This study developed and validated a deep radiomics framework to characterize tumor heterogeneity from pretreatment computed tomography (CT) images and improve prediction of distant metastasis-free survival (DMFS). METHODS AND MATERIALS: This multicenter study included 3421 patients with HN cancer from 4 cohorts across 12 institutions. Radiomics features were extracted from primary tumors and transformed into OmicsMaps, a structured representation that spatially organizes interfeature relationships to facilitate learning of complex prognostic patterns. A convolutional neural network was trained to derive prognostic signatures, which were integrated with key clinical variables to construct an OmicsMap-clinical fusion model for patient risk stratification. Model performance was assessed using the concordance index (C-index) and time-dependent area under the receiver operating characteristic curve (AUC) in the CT Images from Large Head and Neck Cohort (RADCURE), HEAD-NECK-RADIOMICS-HN1 (HN1), and Head-Neck-Positron Emission Tomography-Computed Tomography (HN-PET-CT) cohorts. Radiogenomic analyses using RNA-seq data were conducted in the Cancer Genome Atlas Head-Neck Squamous Cell Carcinoma (TCGA-HNSC) cohort to investigate biological characteristics associated with the imaging-defined risk groups. RESULTS: The OmicsMap achieved C-index values of 0.742, 0.768, and 0.671 in the RADCURE, HN1, and HN-PET-CT cohorts, outperforming the conventional radiomics approach by 5.40%-6.37%. Incorporating clinical variables further improved generalizability, yielding a C-index of 0.864 (HN1) and 0.730 (HN-PET-CT), with time-dependent AUC of 0.727-0.895. The fusion model consistently stratified patients into distinct high- and low-risk groups for both DMFS and overall survival across cohorts (P <.01). Radiogenomic analyses revealed enrichment of immune-related pathways in the low-risk group, whereas the high-risk group exhibited a more aggressive phenotype enriched for proliferation, hypoxia, and epithelial-mesenchymal transition pathways, along with a fibrosis-prone tumor microenvironment characterized by extracellular matrix remodeling. CONCLUSIONS: Modeling interradiomic feature relationships using the OmicsMap representation substantially improves CT-based prediction of DMFS and characterization of tumor heterogeneity in HN cancer, supporting precision risk stratification in clinical oncology.

Journal Article

A transcription factor regulatory atlas for activity inference and perturbation prediction.

Inferring transcription factor (TF) activity from transcriptomes and predicting transcriptome-wide responses to TF perturbations remain challenging, in part because available TF-mRNA resources often face a trade-off between precision and coverage and typically lack signed regulatory information. Here, we present TFActProfiler, a TF-mRNA resource and computational framework that learns signed, quantitative TF-mRNA regulatory coefficients by integrating heterogeneous prior evidence (ChIP-based, motif-based, and curated TF-mRNA annotations) with large-scale bulk and single-cell RNA-seq atlases. TFActProfiler contains 2&#x2009;606&#x2009;176 signed TF-mRNA interactions and improves TF activity inference in TF knockdown benchmarks relative to widely used regulon resources while retaining broad TF and target coverage. In addition, because the same learned regulatory coefficients can be used to model downstream transcriptional effects, TFActProfiler enables prediction of transcriptome-wide gene expression responses to TF knockdown without training on task-matched perturbation data. When perturbation datasets are available, TFActProfiler can be further refined to achieve performance comparable to state-of-the-art machine-learning baselines. By providing a direction-aware representation of TF-mRNA regulation for both activity inference and perturbation-response modeling, TFActProfiler supports systematic dissection of gene regulatory programs across diverse cellular contexts.

Transcription Factors

Exercise Therapy in Down Syndrome: A Systematic Review and Meta-Analysis Focused on Muscle Strength, Redox Balance, and Inflammatory Profile.

OBJECTIVE: This study systematically reviewed and meta-analyzed randomized and quasi-randomized controlled trials investigating the impact of exercise therapy on muscle strength, redox balance, and inflammatory profile in individuals with Down syndrome. DESIGN: Systematic review and meta-analysis. DATA SOURCES: Cochrane Central Register of Controlled Trials, MEDLINE, CINAHL, SPORTDiscus, EMBASE, and PEDro. ELIGIBILITY CRITERIA FOR SELECTING STUDIES: Randomized and quasi-randomized controlled trials exploring exercise therapy effects on muscle strength and redox balance in individuals with Down syndrome. Although no initial restrictions on age, gender, or health condition were applied during the search process, all included studies focused on adult participants (>18 yr old). No language restrictions were applied, and the search covered the period from 1970 to 2021. RESULTS: We assessed the abstract of 1964 studies. Of the 46 studies meeting the inclusion criteria for the period 2004-2021, 32 focused on muscle strength, and 14 examined redox balance and inflammation. A total of 1611 participants with a mean age of 27 yr were included. This review confirmed that different exercise modalities are prone to improve muscle strength (random effect (95% confidence interval): 0.66, 0.54 to 0.78), redox balance and inflammatory profile (random effect (95% confidence interval): -1.04, -1.31 to -0.76) in this population. The multimodel inference suggested that the frequency of training (times per week) might play a significant role in the main effect. Unsupervised machine learning algorithms displayed a pattern-based graphic representation to assess heterogeneity. CONCLUSIONS: Exercise training demonstrated a positive impact on muscle strength in adults with Down syndrome. The review provides valuable insights into the effects of exercise therapy on individuals with Down syndrome, emphasizing the need for tailored training prescriptions.

Humans

scFANCL: Dual contrastive learning with false-negative correction at cell level for single-cell RNA-seq clustering.

BACKGROUND: Single-cell RNA sequencing (scRNA-seq) enables cellular characterization at single-cell resolution. However, its high dimensionality, sparsity, and noise make clustering challenging. Approaches utilizing contrastive learning and data augmentation have been introduced to improve representation quality for scRNA-seq clustering. In particular, dual contrastive frameworks combining instance- and cluster-level objectives can capture both cell-cell similarities and inter-cluster variations. However, existing dual contrastive frameworks focus primarily on discrete cluster boundaries, neglecting the biological continuity inherent in scRNA-seq data. METHODS: We propose scFANCL, a dual contrastive framework designed to capture biological continuity in scRNA data. Rather than treating all non-augmented samples as negatives, scFANCL applies a cosine-similarity-based threshold to exclude cells of the same type from the negative pool, preserving continuous transcriptional relationships among them while maintaining inter-cluster separation. RESULTS: Extensive experiments across seven publicly available scRNA-seq datasets demonstrated that scFANCL achieves competitive clustering performance compared with existing baseline methods, consistently yielding high ARI and NMI scores across datasets of varying size and complexity. Ablation studies further confirmed the contribution of the false negative filtering component, showing measurable improvements over variants without filtering. Downstream analyses further suggest that the learned embeddings may reflect biologically meaningful transcriptional transitions, including continuous differentiation trajectories within related cell types. The source code is available at https://github.com/mjuailab/scFANCL . CONCLUSIONS: scFANCL addresses a key limitation of conventional contrastive learning by applying a cosine-similarity-based threshold to exclude cells of the same type from the negative pool, thereby preserving biological continuity within cell types while maintaining inter-cluster separation. Evaluations across seven benchmark scRNA-seq datasets demonstrate competitive clustering performance, with learned embeddings capturing biologically meaningful transcriptional structure and characteristics of rare cell populations.

Clustering Algorithms

An introduction to model-based imaging.

The purpose of this paper is to clarify the distinction between the recognition of form, i.e. pattern recognition, and the interpretation of visual scenes, i.e. image understanding. Pattern recognition is part of image understanding, but the latter also includes cognitive tasks such as learning and inference. The key to developing image-understanding systems is to concentrate on the representation and use of models. This paper is a brief outline of the components of a model-based image-understanding system. First, the notions of iconic, categorical and symbolic knowledge are described. Although they appear to be disparate, the common notion is that the image understanding is based on recognizing concepts and not recognizing form. Next, the notion of a concept is defined, followed by representation techniques and control strategies for using concepts. Last, an example is given of an image-understanding system that learns to recognize concepts such as radiographic projections of teeth in panoramic radiographs.

Expert Systems

CrossAttOmics: multiomics data integration with cross-attention.

MOTIVATION: Advances in high throughput technologies enabled large access to various types of omics. Each omics provides a partial view of the underlying biological process. Integrating multiple omics layers would help have a more accurate diagnosis. However, the complexity of omics data requires approaches that can capture complex relationships. One way to accomplish this is by exploiting the known regulatory links between the different omics, which could help in constructing a better multimodal representation. RESULTS: In this article, we propose CrossAttOmics, a new deep-learning architecture based on the cross-attention mechanism for multiomics integration. Each modality is projected in a lower dimensional space with its specific encoder. Interactions between modalities with known regulatory links are computed in the feature representation space with cross-attention. The results of different experiments carried out in this article show that our model can accurately predict the types of cancer by exploiting the interactions between multiple modalities. CrossAttOmics outperforms other methods when there are few paired training examples. Our approach can be combined with attribution methods like LRP to identify which interactions are the most important. AVAILABILITY AND IMPLEMENTATION: The code is available at https://github.com/Sanofi-Public/CrossAttOmics and https://doi.org/10.5281/zenodo.15065928. TCGA data can be downloaded from the Genomic Data Commons Data Portal. CCLE data can be downloaded from the depmap portal.

Humans

Predicting enhancer-promoter interactions using a stacking-based ensemble strategy.

MOTIVATION: Enhancer-promoter interactions (EPIs) are essential for gene regulation and disease progression. Recent studies have shown that distal enhancers can regulate target genes through interactions with nearby promoters, providing important insights into transcriptional regulation mechanisms. Although high-throughput experimental techniques have enabled large-scale identification of EPIs, these methods are often costly and time-consuming. In addition, existing computational approaches still face challenges in effectively integrating heterogeneous feature representations from different cell lines. RESULTS: We propose a stacked ensemble framework for EPI prediction that integrates feature representations from diverse cell line datasets using multiple machine learning algorithms. The extracted complementary patterns are further combined by an XGBoost classifier to improve robustness against overfitting. Experiments on six independent datasets show that the proposed method achieves superior accuracy and generalization compared with existing EPI prediction models, with an average AUROC of 0.909 while maintaining computational efficiency. AVAILABILITY: The source code and its archived release are available at GitHub and Zenodo. The Zenodo archive provides a versioned snapshot of the repository: https://zenodo.org/records/19952998.

Promoter Regions, Genetic