PubMed HealthSearch

SEARCH · PubMed Health

Results for “Mixture effects”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

The Association of Prenatal Dietary Factors with Child Autism Diagnosis and Autism-Related Traits Using a Mixtures Approach: Results from the Environmental Influences on Child Health Outcomes Cohort.

BACKGROUND: Previous research on the role of maternal diet in relation to autism has focused on examining individual nutrient associations. Few studies have examined associations with multiple nutrients using mixtures approaches, which may better reflect true exposure scenarios. OBJECTIVES: This study aims to examine associations of nutrient mixtures with children's autism diagnosis and trait scores within a large, diverse population. METHODS: Participants were drawn from the United States Environmental influences on Child Health Outcomes (ECHO) consortium. Maternal prenatal diet was reported via validated food frequency questionnaires. Children's autism-related traits were measured using the Social Responsiveness Scale (SRS) and autism diagnoses were from parent reports of physician diagnosis. Bayesian kernel machine regression was used to examine the overall mixture effect and interactions between a set of 5 primary nutrients (folate, vitamin D, omega 3 and omega 6 fatty acids, and iron), adjusted for potential confounders, in relationship to child outcomes. Secondary analyses were conducted in a subset of cohorts with an expanded set of 14 nutrients. Traditional linear and logistic regression models were also analyzed for comparison of results to mixture models. RESULTS: A total of 2614 participants drawn from 7 ECHO cohorts were included in primary analysis. Mixture analyses suggested that increasing the overall 5-nutrient mixture was associated with lower SRS scores. Individual U-shaped associations and bivariate interactions between folate and omega 3 fatty acids were suggested. In the subset included in the secondary analyses of the 14-nutrient mixture, a modest inverse trend remained, but individual nutrient associations were altered, with vitamin D demonstrating higher relative importance than other nutrients. Strong associations with autism diagnosis were not observed. CONCLUSIONS: In this large sample, we found evidence for combined nutrient effects with broader autism-related traits. Because results for individual nutrients were sensitive to mixture components, replication of combined associations between nutrients and autism-related outcomes is needed.

Humans

Biocontrol effect of a solid-state fermentation-derived extract mixture of Trichoderma asperellum on sunflower Sclerotinia rot and associated host defense responses.

Sclerotinia disease is a destructive fungal disease of sunflowers, soybeans, and other economically important crops, causing substantial yield loss and quality deterioration. Long-term reliance on dose-dependent broad-spectrum fungicides is constrained by resistance risks and potential environmental burdens, creating tension with the sustainability goal of "reducing pesticide use while improving efficacy." Here, we explore a Trichoderma spp.-based microbial disease management strategy. Whole-genome sequencing of Trichoderma asperellum TCS007 isolated from Antarctic marine sediments, coupled with genome mining, predicted diverse biosynthetic gene clusters putatively associated with siderophores, polyketides, nonribosomal peptides, and terpenoids; the corresponding metabolites are not chemically confirmed and require further validation. Using a solid-state fermentation workflow, we prepared a fermentation-derived extract mixture (TCS007-SSF-Ex). In vitro assays showed dose-dependent inhibition of Sclerotinia sclerotiorum by TCS007-SSF-Ex (EC50 = 1.252 mg/L), and microscopy revealed cellular damage-consistent changes, including organelle disruption and plasmolysis. Pathogen transcriptomic and metabolism-related analyses indicated broad perturbations in organelle biogenesis and metabolic processes, with significant alterations in pathways associated with succinate, D-glucose, and phenylacetate; these results are consistent with growth inhibition and reduced pathogenicity, but specific molecular targets and causal links remain to be validated. In vivo, under certain application conditions, triple applications increased APX activity (+492.5%) and β-1,3-glucanase activity (+419.6%). Collectively, this work supports a "pathogen suppression-host defense induction" framework and facilitates subsequent identification of active components and mechanistic validation.IMPORTANCESclerotinia diseases cause recurrent and economically important losses in oilseed crops, while long-term fungicide use is constrained by resistance risks and environmental burdens. Trichoderma-based biocontrol is a promising complementary strategy, yet evidence supporting metabolite-containing Trichoderma-derived preparations as immune elicitors remains less consolidated than that for living inoculants, and scalable production routes are still needed. Here, we examine an Antarctic marine sediment-derived strain, Trichoderma asperellum TCS007, and a solid-state fermentation (SSF)-derived extract mixture (TCS007-SSF-Ex) produced via solid-state fermentation. We combine in vitro antifungal assays, pathogen ultrastructural observations, and correlative omics analyses with in vivo measurements of sunflower defense enzymes (APX and β-1,3-glucanase) to evaluate a "pathogen suppression-host defense induction" framework. Our findings support the potential of SSF-derived Trichoderma metabolite mixtures for greener management of Sclerotinia disease and provide a foundation for future chemical identification of active components and mechanistic validation.

Ascomycota

Unveiling non-small cell lung cancer treatment effect heterogeneity: a comparative analysis of statistical methods.

BACKGROUND: For patients with advanced non-small cell lung cancer lacking targetable genomic alterations, the impact of clinicogenomic characteristics on the effectiveness of combining chemotherapy with immunotherapy is unclear. METHODS: We evaluated 4 statistical methods for detecting heterogeneous treatment effects related to clinical factors, including programmed death-ligand 1 expression, tumor mutation burden, and stage at diagnosis, using the American Association for Cancer Research Project Genomics Evidence Neoplasia Exchange BioPharma Collaborative dataset supplemented with institutional data collected under the same data curation model. A 2-sided P value of no more than .05 was used to denote statistical significance for all analyses. RESULTS: The mixture model revealed 2 latent subgroups: in one subgroup, there was no meaningful treatment effect, with average progression-free survival (PFS) only 5% longer with immunotherapy alone (95% confidence interval [CI] = -19% to 35%); in the second subgroup, immunotherapy alone was associated with a 35% decrease in average PFS (95% CI = -59% to 2%), corresponding to a ratio in treatment effects of 1.62 (95% CI = 1.02 to 2.57). There was a marginal association between lower tumor mutation burden levels and membership in the subgroup with improved PFS following receipt of chemoimmunotherapy. The causal survival forest highlighted the importance of tumor mutation burden (variable importance ranking: 1) and programmed death-ligand 1 (variable importance ranking: 3) when assessing heterogeneity. In contrast, the accelerated failure time and Cox proportional hazards models did not detect any statistically significant heterogeneous treatment effects. In simulations, the mixture model identified heterogeneous treatment effects more frequently than other methods, especially with weak covariate relationships, demonstrating its utility for informing personalized treatment approaches. CONCLUSIONS: The application of novel statistical methods to large scale clinico-genomic databases offers an opportunity to more accurately identify heterogeneous treatment effects in some settings as compared to traditional statistical methods. Applying such methods to the AACR Project GENIE BPC non-small cell lung cancer data indicated a potential association between decreasing tumor mutation burden and improved outcomes with chemoimmunotherapy as compared to immunotherapy alone.

Humans

From molecular responses to environmental monitoring: advances and translational gaps in omics approaches in fish environmental toxicology.

Fish occupy a central position in aquatic ecosystems and serve as important bioindicators for environmental monitoring, as well as powerful translational models for understanding toxic mechanisms conserved across higher vertebrates. In recent years, omics techniques have proven to be powerful tools to address complex environmental questions that conventional toxicology methods cannot answer. Despite this potential, a critical translational gap remains between molecular findings and their use in ecological risk assessment frameworks. This review critically synthesizes advances across omics techniques including epigenomics, transcriptomics, metabolomics and proteomics and their integration. Special emphasis is placed on methodological considerations and practical aspects of these techniques in fish environmental toxicology and environmental monitoring. Evidence from single-omics studies suggests conserved biomarker signatures across species while characterizing complex phenomena like non-monotonic dose-response relationships, mixture toxicity and transgenerational and stereoselective effects with implications for population level monitoring. Multi-omics studies, especially those involving triple omics, further enhance mechanistic resolution by reconstructing adverse outcome pathways. We further evaluate using case studies when additional molecular layers provide critical insight and when they offer limited advantage, a strategic distinction with direct implications in environmental monitoring programmes. Finally, current limitations and future directions that will ultimately bridge the translational gap and hold promise for advancing mechanistic ecotoxicology and predictive environmental monitoring are discussed.

Animals

Heterogeneity Analysis of Associations Involving the Large-Scale Online MindCrowd Survey Memory Test.

INTRODUCTION: Alzheimer's disease and related disorders (ADRDs), as well as general age-related cognitive decline, are known to be multifactorial with heterogeneous etiologies. Identifying and accommodating heterogeneity in any one ADRD-related data set can be pursued using different analytical techniques, each with different assumptions or purposes. For example, whereas a great deal of research has explored clustering individuals or variables that exhibit greater similarity in some way, little research has explored evidence for heterogeneity in the relationships between relevant outcomes, such as performance on a memory test, and risk factors such as environmental exposures, behaviors, or genetic factors among individuals. METHODS: We explored evidence of heterogeneity in the relationships between ability on a memory test, specifically the paired associate learning (PAL) test, and multiple social and demographic risk factors using the large MindCrowd study database (n > 90,000 individuals). We focused on mixtures of regression models but compared models assuming many interaction effects among independent variables as well as random effects. RESULTS: We ultimately find substantial evidence for heterogeneity and offer an intuitive explanation for it involving individual motivation for participating in the MindCrowd study. Basically, we argue that our mixture of regression model analysis results suggest that a smaller group of individuals (∼16%) likely participated in the MindCrowd study out of a concern for their cognitive abilities as they exhibit stronger and statistically significant negative associations between age, number of medications they are on, some ancestries, and the number correct on the PAL test. They also exhibit stronger positive associations between education and PAL test results in a dose-dependent manner suggesting that a "cognitive reserve" associated with greater education could benefit them. Analysis models assuming interaction terms and random effects suggested that other forms of heterogeneity in the relationships between variables exist in the data set, but their results do not carry with them the same intuitive explanation that the results of the mixture model analyses do. CONCLUSION: We find evidence for heterogeneity in the relationships between social and demographic variables and PAL test results in the large MindCrowd study database. This heterogeneity is likely due to individuals with and without concerns for their cognitive abilities participating in the study. We also find other types of evidence in the data set. Our results should motivate caution in the use of large epidemiological study or survey-oriented data sets to build predictive models of clinical or subclinical pathologies without exploring or accommodating heterogeneity. Our results also suggest that one should include questions about motivation to participate in large epidemiological studies since different motivations may impact important relationships between independent and dependent variables.

Humans

Associations between multiple essential trace metal concentrations and risk of hyperuricemia: insights from a central Chinese population.

Previous studies have indicated that levels of individual essential trace metals are related to hyperuricemia (HUA), but evidence on their combined effects is limited. To address this gap, the associations of individual and joint levels of 12 essential trace metals (manganese, selenium, nickel, chromium, cobalt, tin, iron, molybdenum, zinc, strontium, vanadium, and copper) with the risk of HUA were investigated in 2,021 adults recruited from Hunan Province, China. Inductively coupled plasma mass spectrometry (ICP-MS) was employed to determine urinary metal concentrations. Logistic regression, Bayesian kernel machine regression (BKMR), and quantile g calculation (Qgcomp) were applied to evaluate the associations of single and mixture metal concentrations with HUA. Of the participants, 516 (25.53%) were diagnosed with HUA. Inverse associations were found between vanadium, chromium, manganese, iron, cobalt, selenium, strontium, and molybdenum levels and HUA, with ORs ranging from 0.63 to 0.91. Conversely, a positive association was observed between zinc concentration and HUA [OR (95% CI): 1.17 (1.01, 1.37)]. Both BKMR and Qgcomp models showed a negative overall effect of essential trace metals on HUA risk, with strontium (- 43.6%) and vanadium (- 27.8%) being the main contributors. In addition, formal interaction tests revealed significant effect modification by age for tin and by BMI for zinc. In conclusion, the levels of essential trace metals were linked to a decreased risk of HUA, and these associations were modified by age and BMI only for specific metals.

Humans

Urinary Metal Levels, Cognitive Test Performance, and Dementia in the Multi-Ethnic Study of Atherosclerosis.

IMPORTANCE: Metals are established neurotoxicants, but evidence of their association with cognitive performance at low chronic exposure levels is limited. OBJECTIVE: To investigate the association of urinary metal levels, individually and as a mixture, with cognitive tests and dementia diagnosis, including effect modification by apolipoprotein ε4 allele (APOE4). DESIGN, SETTING, AND PARTICIPANTS: The multicenter prospective cohort Multi-Ethnic Study of Atherosclerosis (MESA) was started from July 2000 to August 2002, with follow-up through 2018. A total of 6303 MESA participants were included. Data analysis was performed from October 12, 2023, to June 13, 2024. EXPOSURE: Urine samples were collected at baseline (2000-2002), and arsenic, cadmium, cobalt, copper, lead, manganese, tungsten, uranium, and zinc levels were measured in 2020-2022. MAIN OUTCOMES AND MEASURES: Digit Symbol Coding (DSC) (n = 3819) (possible score range, 0-133), Cognitive Abilities Screening Instrument (CASI) (n = 3918) (possible score range, 0-100), and Digit Span (DS) (n = 4176) (possible score range, 0-30) cognitive tests were administered in 2010-2012; higher scores of each test indicate increasing levels of positive response. RESULTS: A total of 6303 participants were followed up for dementia diagnosis through 2018. The median age at baseline was 60 (IQR, 53-70) years, and 3303 participants (52.4%) were female. The median cognitive scores were 51 (IQR, 38-64) for DSC, 90 (IQR, 84-95) for CASI, and 15 (IQR, 12-18) for DS. There were 559 cases of dementia through the follow-up period. Inverse associations with DSC were identified: mean differences in z scores per IQR increase in metal levels were -0.03 (95% CI, -0.07 to 0.00) for arsenic, -0.05 (95% CI, -0.09 to -0.004) for cobalt, -0.05 (95% CI, -0.07 to -0.02) for copper, -0.04 (95% CI, -0.08 to -0.001) for uranium, and -0.03 (95% CI, -0.06 to -0.01) for zinc. Among 1058 APOE4 carriers, manganese was also inversely associated with DSC. The joint mean difference of DSC comparing percentile 95th with the 25th of the 9-metal mixture was -0.30 (95% CI, -0.47 to -0.14) for APOE4 carriers and -0.10 (95% CI, -0.19 to -0.01) for noncarriers. Arsenic, cadmium, cobalt, copper, tungsten, uranium, and zinc were individually associated with dementia, with hazard ratios per IQR of metal ranging from 1.15 (95% CI, 1.03-1.29) for tungsten to 1.46 (95% CI, 1.06-2.02) for uranium. The joint hazard ratio of dementia comparing percentiles 95th with the 25th of the 9-metal mixture was 1.71 (95% CI, 1.24-3.89), with no significant difference by APOE4 status. CONCLUSIONS AND RELEVANCE: In this study, participants with higher concentrations of metals in their urine, compared with those with lower concentrations, had worse performance on cognitive tests and greater likelihood of developing dementia. The findings of this multicenter multiethnic cohort study might inform screening and potential interventions for prevention of dementia based on individuals' metal exposure levels and genetic profiles.

Humans

A flexible framework for robust and efficient Mendelian randomization with debiasing.

Mendelian randomization (MR) has been widely used to infer causal relationships between exposures and outcomes in epidemiological studies. However, classical MR assumptions can be violated when genetic variants are associated with outcomes through pathways other than the exposure, leading to uncorrelated and/or correlated pleiotropy. Additionally, measurement error arising from the inherent uncertainty in summary statistics obtained from large-scale genome-wide association studies can introduce bias into the causal effect estimate. To address these issues, we develop a debiased mixture inverse variance weighting ($\mathsf{dmIVW}$) method with three major advantages. First, it is capable of simultaneously handling various types of pleiotropy and eliminating the bias caused by uncertainty. Second, it can guard against distortion caused by invalid genetic variants while effectively harnessing their information. Third, our unified framework facilitates a fair comparison and combination of a series of submodels, encompassing several popular MR methods as special cases. Through real data applications, the effectiveness and robustness of $\mathsf{dmIVW}$ in estimating the causal effects of risk factors on common diseases are demonstrated.

Mendelian Randomization Analysis

Rapid and exceptionally small-scale adaptation of the alpine plant Cardamine resedifolia to mining-contaminated soils in multi-stress condition.

The mechanisms by which plants tolerate soil contamination have been studied in details in controlled laboratory conditions, but they still remain largely unexplored in natural conditions where mixtures of contaminants are present in soils and their effects might interact with other environmental variables. This is especially true in high-altitude alpine environments, where abiotic stress is naturally heightened, but which so far have received little attention in environmental pollution studies. As we were interested in the tolerance mechanisms at play on very fine spatiotemporal scales for alpine plants growing under multi-stress conditions, we chose Cardamine resedifolia as our biological model. This plant is indeed frequently found in areas contaminated by Trace Metals and Metalloids and Polycyclic Aromatic Hydrocarbons in high elevation. We studied populations from former copper, silver-lead, and coal mines in alpine environments, along with populations growing on nearby reference soils. We measured genetic variability within populations as well as genetic differentiation between them, and tested for local adaptation to soil contamination using reciprocal transplants. Population pairs showing signs of local adaptation were then examined using genome scans to identify genes potentially under selection. We found high levels of genetic differentiation between populations growing on contaminated and reference soils a few dozen meters apart. In most cases local adaptation was detected, especially in former copper mines. Genome scans identified genes involved in metal stress management as potentially being under selection. This study provides evidence for rapid adaptation to human-induced pollution in alpine plants at remarkably small spatial scales. It offers new insights into the short-term ecological and evolutionary consequences of mining activities in alpine ecosystems, particularly in relation to substrate-driven differentiation.

Alpine plants

Sex and tissue resolved co-expression networks reveal a female placental-brain axis protective against prenatal PCB exposure.

BACKGROUND: Neurodevelopmental disorders have a strong male bias that is poorly understood. The placenta provides molecular information about environmental interactions with genetics (including biological sex) that shape developmental processes in the brain. We investigate placental-brain transcriptional responses in an established mouse model of prenatal exposure to a human-relevant mixture of polychlorinated biphenyls (PCBs). RESULTS: To understand sex, tissue, and dosage effects in embryonic (E18) brain and placenta RNAseq data, we use weighted gene correlation network analysis (WGCNA) to create gene networks that could be compared across sex or tissue. WGCNA reveals that expression within most correlated gene networks is significantly and strongly associated with PCB exposure, but frequently in opposite directions between male-female and placenta-brain comparisons. In WGCNA and differentially expressed gene analyses, more transcriptional changes are observed in male brain than placenta, but the reverse is seen in females. Furthermore, female X-inactive specific transcript (Xist) levels correlate with sex-specific and non-monotonic PCB dose response, suggesting an X-linked protective epigenetic mechanism. The transcriptomic effects of low-dose PCB exposure are significantly opposed by dietary folic acid supplementation across both sexes but are strongest in female placentas. PCB and folic acid interacting gene networks are enriched in metabolic pathways involved in energy usage and translation, with female-specific protective effects enriched in PPAR, thermogenesis, glycerolipid, and O-glycan biosynthesis, as opposed to toxicant responses in male brain. CONCLUSIONS: A female protective effect in response to prenatal PCB exposure appears to be mediated by dose-dependent sex differences in transcriptional modulation of placental metabolic pathways.

Female

CodonMoE: DNA language models for codon-dependent mRNA prediction.

MOTIVATION: Genomic language models (gLMs) face a fundamental efficiency challenge: one must either maintain separate specialized models for each biological modality (DNA and RNA) or develop large multimodal architectures. Both approaches impose significant computational burdens-modality-specific models require redundant infrastructure despite inherent biological connections, while multi-modal architectures demand increased parameter counts and extensive cross-modality pretraining. RESULTS: To address this limitation, we introduce CodonMoE (Adaptive Mixture of Codon Reformative Experts), a lightweight adapter that transforms DNA language models into effective RNA analyzers without RNA-specific pretraining. Our theoretical analysis establishes CodonMoE as a universal approximator at the codon level, capable of mapping arbitrary functions from codon sequences to codon-dependent RNA properties given sufficient expert capacity. Across four RNA prediction tasks spanning stability, expression, and regulation, DNA models augmented with CodonMoE significantly outperform their unmodified counterparts, with the HyenaDNA+CodonMoE series achieving state-of-the-art results using 80% fewer parameters than specialized RNA models. By maintaining sub-quadratic complexity while achieving superior performance, our approach provides a principled path toward unifying genomic language modeling, leveraging more abundant DNA data and reducing computational overhead while preserving modality-specific performance advantages. AVAILABILITY AND IMPLEMENTATION: Source code for the method and to reproduce the results is available at https://github.com/Kingsford-Group/CodonMoE.

Codon

Phthalates and sex steroid hormones across the perimenopausal period: A longitudinal analysis of the Midlife Women's Health Study.

BACKGROUND: The menopausal transition involves significant sex hormone changes. Environmental chemicals, such as urinary phthalate metabolites, are associated with sex hormone levels in cross-sectional studies. Few studies have assessed longitudinal associations between urinary phthalate metabolite concentrations and sex hormone levels during menopausal transition. METHODS: Pre- and perimenopausal women from the Midlife Women's Health Study (MWHS) (n = 751) contributed data at up to 4 annual study visits. We quantified 9 individual urinary phthalate metabolites and 5 summary measures (e.g., phthalates in plastics (∑Plastic)), using pooled annual urine samples. We measured serum estradiol, testosterone, and progesterone collected at each study visit, unrelated to menstrual cycling. Linear mixed-effects models and hierarchical Bayesian kernel machine regression analyses evaluated adjusted associations between individual and phthalate mixtures with sex steroid hormones longitudinally. RESULTS: We observed associations between increased concentrations of certain phthalate metabolites and lower testosterone and higher sub-ovulatory progesterone levels, e.g., doubling of monoethyl phthalate (MEP), monobenzyl phthalate (MBzP), di-2-ethylhexyl phthalate (∑DEHP) metabolites, ∑Plastic, and ∑Phthalates concentrations were associated with lower testosterone (e.g., for ∑DEHP: -4.51%; 95% CI: -6.72%, -2.26%). For each doubling of MEP, certain DEHP metabolites, and summary measures, we observed higher mean sub-ovulatory progesterone (e.g., ∑AA (metabolites with anti-androgenic activity): 6.88%; 95% CI: 1.94%, 12.1%). Higher levels of the overall time-varying phthalate mixture were associated with lower estradiol and higher progesterone levels, especially for 2nd year exposures. CONCLUSIONS: Phthalates were longitudinally associated with sex hormone levels during the menopausal transition. Future research should assess such associations and potential health impacts during this understudied period.

Humans

Characterization of oxidative status in maize protoplasts under temperature and saline-alkali stresses.

BACKGROUND: Protoplasts have emerged as a powerful model system in plant functional genomics, offering significant utility in functional gene analysis, protein interaction studies, and transient expression platforms for gene editing. Despite their versatility, inherent limitations restrict their broader application, highlighting the need for systematic investigations into their responses to abiotic stressors, such as temperature fluctuations and saline-alkali conditions (200 mM saline mixture: 170mM NaCl and 30mM Na2CO3, pH = 9.1). RESULTS: In this study, we comprehensively examined the effects of varying temperatures and saline-alkali stress on the integrity, viability, and reactive oxygen species (ROS) metabolism of maize protoplasts. Key markers of oxidative stress-including ROS accumulation, lipid peroxidation (measured as malondialdehyde, MDA), antioxidant enzyme activity (superoxide dismutase, SOD), and hydrogen peroxide (H2O2) levels-were quantified to assess the oxidative stress response. Protoplasts maintained at 4 °C demonstrated enhanced stability and antioxidant capacity, preserving cell viability and endogenous protein integrity for up to 16 h. Conversely, exposure to 37 °C significantly compromised protoplast viability, while incubation at 28 °C exerted minimal effects within 16 h. CONCLUSIONS: Our study investigated the effects of various temperature stresses and salt-alkali stress on maize protoplasts. The results demonstrated that both temperature and salt-alkali stress significantly impacted protoplast production, viability, and the expression of endogenous proteins. These findings not only characterize the redox response of maize protoplasts, but also provide guidance for protoplast isolation and other procedures: 4 °C is suitable for short-term maintenance, 25-28 °C for routine functional assays, and 37 °C should be avoided. These findings provide valuable insights into the stress responses of protoplasts and establish a foundation for future research aimed at improving plant stress tolerance through protoplast-based techniques.

Zea mays

DNABERT-S: pioneering species differentiation with species-aware DNA embeddings.

SUMMARY: We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e. DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 28 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. AVAILABILITY AND IMPLEMENTATION: Model, codes, and data are publically available at https://github.com/MAGICS-LAB/DNABERT_S.

Sequence Analysis, DNA

DNABERT-S: Pioneering Species Differentiation with Species-Aware DNA Embeddings.

We introduce DNABERT-S, a tailored genome model that develops species-aware embeddings to naturally cluster and segregate DNA sequences of different species in the embedding space. Differentiating species from genomic sequences (i.e., DNA and RNA) is vital yet challenging, since many real-world species remain uncharacterized, lacking known genomes for reference. Embedding-based methods are therefore used to differentiate species in an unsupervised manner. DNABERT-S builds upon a pre-trained genome foundation model named DNABERT-2. To encourage effective embeddings to error-prone long-read DNA sequences, we introduce Manifold Instance Mixup (MI-Mix), a contrastive objective that mixes the hidden representations of DNA sequences at randomly selected layers and trains the model to recognize and differentiate these mixed proportions at the output layer. We further enhance it with the proposed Curriculum Contrastive Learning (C2LR) strategy. Empirical results on 23 diverse datasets show DNABERT-S's effectiveness, especially in realistic label-scarce scenarios. For example, it identifies twice more species from a mixture of unlabeled genomic sequences, doubles the Adjusted Rand Index (ARI) in species clustering, and outperforms the top baseline's performance in 10-shot species classification with just a 2-shot training. Model, codes, and data is publicly available at https://github.com/MAGlCS-LAB/DNABERT_S.

Journal Article

Systematic evaluation of one-dimensional-to-two-dimensional near-infrared spectroscopy transformations with deep learning for quantifying coconut sap adulteration.

Near-infrared (NIR) spectroscopy have limitations when combined with deep learning (DL) algorithms because they rely on low-dimensional datasets. Therefore, we investigated the potential of transforming one-dimensional (1D) NIR spectra into two-dimensional (2D) spectrograms using synchronous and asynchronous techniques and the continuous wavelet transform (CWT) and their effectiveness by integrating with DL for detecting adulteration in coconut sap. NIR spectra (12,500-4000 cm-1) were collected from binary mixtures (0%-100%;w/w). The performance of all DL (convolutional neural networks-CNN, AlexNet and ResNet) models was compared with that of partial least squares (PLS). The models were ranked in the mentioned order based on their performances: 2D-CWT > 2D-asynchronous > 2D-synchronous > 1D/2D-PLS. The important features of the best model can be explained and visualized using gradient-weighted-class-activation-mapping. The findings highlight that the 1D-to-2D NIR data transformation combined with DL is a highly robust approach because it addresses the feature representation gap in NIR data and effectively captures the spatial-spectral correlations.

Spectroscopy, Near-Infrared

The genetic basis of adaptation to copper pollution in Drosophila melanogaster.

Introduction: Heavy metal pollutants can have long lasting negative impacts on ecosystem health and can shape the evolution of species. The persistent and ubiquitous nature of heavy metal pollution provides an opportunity to characterize the genetic mechanisms that contribute to metal resistance in natural populations. Methods: We examined variation in resistance to copper, a common heavy metal contaminant, using wild collections of the model organism Drosophila melanogaster. Flies were collected from multiple sites that varied in copper contamination risk. We characterized phenotypic variation in copper resistance within and among populations using bulked segregant analysis to identify regions of the genome that contribute to copper resistance. Results and Discussion: Copper resistance varied among wild populations with a clear correspondence between resistance level and historical exposure to copper. We identified 288 SNPs distributed across the genome associated with copper resistance. Many SNPs had population-specific effects, but some had consistent effects on copper resistance in all populations. Significant SNPs map to several novel candidate genes involved in refolding disrupted proteins, energy production, and mitochondrial function. We also identified one SNP with consistent effects on copper resistance in all populations near CG11825, a gene involved in copper homeostasis and copper resistance. We compared the genetic signatures of copper resistance in the wild-derived populations to genetic control of copper resistance in the Drosophila Synthetic Population Resource (DSPR) and the Drosophila Genetic Reference Panel (DGRP), two copper-naïve laboratory populations. In addition to CG11825, which was identified as a candidate gene in the wild-derived populations and previously in the DSPR, there was modest overlap of copper-associated SNPs between the wild-derived populations and laboratory populations. Thirty-one SNPs associated with copper resistance in wild-derived populations fell within regions of the genome that were associated with copper resistance in the DSPR in a prior study. Collectively, our results demonstrate that the genetic control of copper resistance is highly polygenic, and that several loci can be clearly linked to genes involved in heavy metal toxicity response. The mixture of parallel and population-specific SNPs points to a complex interplay between genetic background and the selection regime that modifies the effects of genetic variation on copper resistance.

Drosophila

The curious case of sporadic nematode susceptibility in "Tifguard" peanut (Arachis hypogaea): seed mixture or genetic instability?

The Runner-type peanut (Arachis hypogaea L.) cultivar "Tifguard" carries an introgressed chromosomal segment on chromosome A09 from A. cardenasii that confers resistance to root-knot nematode (RKN). Despite this, a proportion of "Tifguard" plants show RKN symptoms, which could plausibly be attributed to seed mixture or outcrossing. However, recent work has shown that cultivated peanut exhibits surprisingly frequent large-scale chromosomal instability (1% to 5%); suggesting that resistance loss could arise from spontaneous structural genomic change. To test these possibilities, we grew foundation seed in an RKN-infested field and collected symptomatic and asymptomatic plants. Lineages derived by single-seed descent were genotyped using the Axiom Arachis 48K SNP array v2 and whole-genome sequencing. Symptomatic lineages lacked the A. cardenasii introgression on chromosome A09 and instead carried the complete endogenous A. hypogaea A09 region at the expected dosage. There was no evidence of large-scale homoeologous exchange, deletion, or other genomic instability affecting this chromosome. Most susceptible plants were closely related to resistant "Tifguard" but lacked the A09 introgression, with a smaller proportion assignable to known nematode-susceptible cultivars, implicating seed mixture with a possible contribution from cross-pollination rather than genomic instability. Because resistance depends on a single major-effect segment, rare events have disproportionate phenotypic impact, placing high demands on genetic purity. For important traits conferred by major loci, marker-based testing across seed-increase stages could verify trait retention directly, and is increasingly practical as marker costs decline.

Arachis