PubMed HealthSearch

SEARCH · PubMed Health

Results for “MixUp”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

3 recordsLinked to original sources

Hypernetwork-guided fusion with intra-class MixUp for breast cancer subtyping.

Accurate breast cancer subtyping guides treatment selection, yet histopathology captures morphology without molecular state, while genomic profiling captures molecular signatures without spatial context. Existing fusion methods rely on concatenation, or on attention applied only after each modality is encoded independently. This work identifies a scale-dependent asymmetry in the direction of cross-modal conditioning: the direction that performs best under limited samples is not the one that holds at scale, and the reversal is traced to the capacity of the modulation pathway rather than to the fusion principle. The comparison is carried out within a hypernetwork-guided framework in which an auxiliary network maps one modality to conditioning parameters that modulate the other's feature representation, shaping features at the parametric level rather than the decision stage; modulation is patient-specific rather than patch-specific. Both directions are instantiated-gene-to-image (HyperG2I) and image-to-gene (HyperI2G) - and trained under a label-aware MixUp strategy that interpolates within-class samples across both modalities, preserving the hard binary labels clinical decisions require. The framework is evaluated on two paired TCGA-BRCA cohorts-one limited-sample, one independently assembled at scale-under a single protocol spanning two whole-slide representations, multiple visual backbones, and both conditioning directions. On the limited-sample cohort, gene-to-image conditioning at its optimal augmentation setting exceeds early fusion and both unimodal baselines, giving the highest recall on the aggressive Basal/HER2 class of any configuration evaluated, and an ablation favours intra-class over inter-class mixing. At scale this ordering does not hold: image-to-gene conditioning sustains its performance whereas gene-to-image does not, recovering only partially under the full tissue bag and isolating the capacity of the modulation pathway as the binding constraint. Direction and capacity of cross-modal conditioning, rather than fusion depth alone, therefore govern how such frameworks scale.

Breast Neoplasms

DNABERT-S: pioneering species differentiation with species-aware DNA embeddings.

SUMMARY: We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e. DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 28 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. AVAILABILITY AND IMPLEMENTATION: Model, codes, and data are publically available at https://github.com/MAGICS-LAB/DNABERT_S.

Sequence Analysis, DNA

DNABERT-S: Pioneering Species Differentiation with Species-Aware DNA Embeddings.

We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e., DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 23 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. Model, codes, and data is publicly available at https://github.com/MAGlCS-LAB/DNABERT_S.

Journal Article