PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “High-throughput sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

A streamlined protocol for small-scale protoplast generation and CRISPR/Cpf1-mediated genome editing in Fusarium oxysporum.

Fusarium oxysporum is a significant threat to agriculture and One Health, requiring advanced molecular tools for functional genomic analyses and biological control agent development. Existing gene-editing methods are hampered by costly protoplast preparation protocols and by CRISPR-Cas9 limitations, such as restricted protospacer adjacent motif (PAM) sequences and complex guide RNA requirements. We engineered an efficient CRISPR/Cpf1 system that overcomes these issues through three main innovations: small-scale protoplast generation using filter column-based methods that greatly reduce enzyme consumption while simplifying workflows, a CRISPR/Cpf1 system with shorter guide RNA design and staggered DNA cleavage to promote homologous recombination, and minimal homology arm strategies that significantly decrease cloning complexity. Extensive validation confirms successful gene targeting with molecular verification and functional analysis via standardized pathogenicity assays. This integrated platform offers affordable, accessible tools for systematic F. oxysporum research, enhancing fundamental understanding of plant-pathogen interactions and supporting high-throughput screening vital for agricultural biotechnology and biological agent development.

CRISPR/Cpf1↗

Genome-wide analysis of Enterococcus faecalis genes that facilitate interspecies competition with Lactobacillus crispatus.

Enterococci are opportunistic pathogens notorious for causing a variety of infections. While both Enterococcus faecalis and Lactobacillus crispatus are commensal residents of the vaginal tract, the molecular mechanisms that enable E. faecalis to take advantage of a vaginal biome with lower counts of lactobacilli to colonize the vaginal tract and induce aerobic vaginitis remain unknown. Here, we show that L. crispatus eradicates E. faecalis in a contact-independent manner. Using transposon sequencing to identify E. faecalis OG1RF transposon (Tn) mutants that are either under-represented or over-represented when co-cultured with L. crispatus, we found that Tn mutants with disruption in the dltABCD operon, that encodes the proteins responsible for the D-alanylation of teichoic acids, and OG1RF_11697 encoding for an uncharacterized hypothetical protein are more susceptible to killing by L. crispatus. Inversely, Tn mutants with disruption in ldh1, which encodes for L-lactate dehydrogenase, are more resistant to L. crispatus killing. Using the Galleria mellonella infection model, we show that co-injection of L. crispatus with E. faecalis OG1RF enhances larvae survival while this L. crispatus-mediated protection was lost in larvae co-infected with either L. crispatus and E. faecalisΔldh1 or Δldh1Δldh2 strains. Last, using RNA sequencing to identify E. faecalis genes that are differently expressed in the presence of L. crispatus, we found major changes in the expression of genes associated with glycerophospholipid metabolism, central metabolism, and general stress responses. The findings in this study provide insights into how E. faecalis mitigate assaults by L. crispatus.IMPORTANCEEnterococcus faecalis is an opportunistic pathogen notorious for causing a multitude of infections. As vaginal commensals, E. faecalis must interact with Lactobacillus crispatus, but how E. faecalis overcomes or mitigate assaults by L. crispatus killing remains unknown. We show that L. crispatus eradicates E. faecalis temporally in a contact-independent manner. Using high-throughput molecular approaches, we identified genetic determinants that enable E. faecalis to compete with L. crispatus. This study represents an important first step for the identification of adaptive genetic traits required for enterococci to tolerate assaults by lactobacilli.

Enterococcus faecalis↗

Enhanced identification of key bacterial motility genes via a cross-species genomic hybrid feature machine learning approach.

Efficient and accurate identification of functional genes is critical to biological research, yet traditional single-species approaches are often limited by low efficiency. Previously, we established a novel method for identifying key genes using cross-species protein domain features and machine learning. However, the high multiplicity of gene members associated with specific domains creates a substantial workload for subsequent experimental validation. To address this, this study proposes an enhanced approach that integrates EggNOG-based protein sequence annotation with domain analysis. Unannotated sequences are subsequently analyzed for protein domains, generating a comprehensive "direct gene annotation plus domain" hybrid feature matrix. While the hybrid matrix model yielded comparable predictive accuracy, it significantly enhanced feature resolution: the top 50 predicted features were all known motility-related genes or domains. Furthermore, among the top 100 ranked features, 58 are confirmed to be directly related to motility based on experimental evidence. Although strict genus-level control still yielded 51 confirmed features, excessive taxonomic restriction drastically reduces the number of training genomes, which may paradoxically impair identification efficiency. These results demonstrate that the new method effectively reduces the subsequent experimental workload and enables high-throughput identification of functional genes in a single analysis. With accuracy and efficiency far exceeding those of existing single-species identification methods, it provides a highly efficient solution for mining key genes underlying other complex bacterial phenotypes.

Machine Learning↗

2-Mercaptoethanol/DMSO Workflow Enables Highly Reproducible Quantitative Proteomics.

Proteomics provides a systematic and high-throughput approach to comprehensively characterize protein networks, enabling insights into cellular functions and disease mechanisms. Carbamidomethylation using iodoacetamide (IAA), a common method for cysteine alkylation, is known to cause nonspecific modifications that increase spectral complexity in mass spectrometry and reduce quantitative accuracy. Here, we established a reproducibility-focused 2-mercaptoethanol (2-ME)/dimethyl sulfoxide (DMSO) workflow and systematically evaluated its quantitative performance at the proteome-wide level. Mouse liver proteomes were processed using either 2-ME/DMSO or conventional IAA treatment, followed by liquid chromatography-tandem mass spectrometry (LC-MS/MS) analysis. The optimized 2-ME treatment increased the number of cysteine-modified peptides by 1.6- to 1.9-fold. Although total protein identifications were comparable, 77% of proteins exhibited improved sequence coverage with the optimized 2-ME treatment. Quantitative reproducibility was also enhanced, with the peptide quantified CV ≤ 20% increasing from 61.4% with IAA treatment to 86.1% with 2-ME treatment, and protein quantified CV ≤ 20% increasing from 80.6% with IAA treatment to 93.5% with 2-ME treatment. Application of this new workflow to ovarian clear cell carcinoma reliably detected cisplatin-induced alterations. The 2-ME/DMSO workflow offers a simple and highly reproducible proteomics strategy for accurate quantitative proteomics.

Animals↗

Multi-Omics and Integrative Analytics in Natural Products Discovery.

Natural products (NPs) have long been an essential source of new bioactive compounds for drug discovery; however, traditional methods for screening and isolating these compounds can be slow and often yield diminishing returns. Fortunately, advanced multi-omics and computational approaches present powerful solutions to these challenges. This review highlights innovative methodologies that integrate metabolomics, genomics, transcriptomics, and proteomics with bioinformatics and analytical chemistry to accelerate NP discovery. For instance, untargeted metabolomics platforms like high-resolution liquid chromatography-tandem mass spectrometry (LC-MS/MS) and Global Natural Products Social (GNPS) molecular networking allow for comprehensive profiling of new compounds, while targeted isotope-labeling strategies enhance this process. Additionally, genome and metagenome mining tools such as antibiotics and secondary metabolite analysis shell (antiSMASH), Deep Biosynthetic Gene Cluster (DeepBGC), and Pipeline for Reconstructing Integrated Syntheses of Metabolites (PRISM) quickly identify biosynthetic gene clusters (BGCs) in both cultured and uncultured organisms, often using heterologous expression to validate products. Transcriptomic analyses, including RNA sequencing (RNA-seq), co-expression networks, and fluxomics, help clarify how pathways are regulated, while quantitative proteomics techniques like tandem mass tags/isobaric tags for relative and absolute quantitation (TMT/iTRAQ) and label-free methods, along with chemoproteomics approaches such as cellular thermal shift assay and thermal proteome profiling (TPP), uncover molecular targets and their mechanisms of action. This review also places significant emphasis on the role of artificial intelligence (AI) and machine learning (ML) in integrating multi-omics data, spanning activities from constructing gene-metabolite correlation networks to leveraging knowledge graphs and graph neural networks for data fusion and functional prediction. Finally, this review concludes by discussing the synergistic benefits of multi-omics for natural-product discovery, addressing current technical challenges, and exploring future directions toward high-throughput, intelligent data integration for next-generation NP research.

Biological Products↗

A time-resolved single-cell roadmap of the logic driving anterior neural crest diversification from neural border to migration stages.

Neural crest cells exemplify cellular diversification from a multipotent progenitor population. However, the full sequence of early molecular choices orchestrating the emergence of neural crest heterogeneity from the embryonic ectoderm remains elusive. Gene-regulatory-networks (GRN) govern early development and cell specification toward definitive neural crest. Here, we combine ultradense single-cell transcriptomes with machine-learning and large-scale transcriptomic and epigenomic experimental validation of selected trajectories, to provide the general principles and highlight specific features of the GRN underlying neural crest fate diversification from induction to early migration stages using Xenopus frog embryos as a model. During gastrulation, a transient neural border zone state precedes the choice between neural crest and placodes which includes multiple converging gene programs. During neurulation, transcription factor connectome, and bifurcation analyses demonstrate the early emergence of neural crest fates at the neural plate stage, alongside an unbiased multipotent-like lineage persisting until epithelial-mesenchymal transition stage. We also decipher circuits driving cranial and vagal neural crest formation and provide a broadly applicable high-throughput validation strategy for investigating single-cell transcriptomes in vertebrate GRNs in development, evolution, and disease.

Animals↗

Comparative cellular analysis of motor cortex in human, marmoset and mouse.

The primary motor cortex (M1) is essential for voluntary fine-motor control and is functionally conserved across mammals1. Here, using high-throughput transcriptomic and epigenomic profiling of more than 450,000 single nuclei in humans, marmoset monkeys and mice, we demonstrate a broadly conserved cellular makeup of this region, with similarities that mirror evolutionary distance and are consistent between the transcriptome and epigenome. The core conserved molecular identities of neuronal and non-neuronal cell types allow us to generate a cross-species consensus classification of cell types, and to infer conserved properties of cell types across species. Despite the overall conservation, however, many species-dependent specializations are apparent, including differences in cell-type proportions, gene expression, DNA methylation and chromatin state. Few cell-type marker genes are conserved across species, revealing a short list of candidate genes and regulatory mechanisms that are responsible for conserved features of homologous cell types, such as the GABAergic chandelier cells. This consensus transcriptomic classification allows us to use patch-seq (a combination of whole-cell patch-clamp recordings, RNA sequencing and morphological characterization) to identify corticospinal Betz cells from layer 5 in non-human primates and humans, and to characterize their highly specialized physiology and anatomy. These findings highlight the robust molecular underpinnings of cell-type diversity in M1 across mammals, and point to the genes and regulatory pathways responsible for the functional identity of cell types and their species-specific adaptations.

Animals↗

The ecology and evolution of microbial immune systems: a look on the wild vibrio side.

Natural populations of vibrio beyond the well-studied pandemic strains of Vibrio cholerae, provide a powerful model for investigating the eco-evolutionary dynamics of microbial immune systems. Their genetic diversity, ecological versatility, ease of culturability and the availability of time-series data enable detailed studies of phage-host interactions in natural contexts. This review synthesizes recent advances in vibriophage research, highlighting key findings and emerging tools. High-throughput assays and genomic tools have offered new perspectives on phage specificity, host range and the evolutionary pressures shaping these interactions. Theoretical frameworks, such as arms race and fluctuating selection dynamics, are informed by empirical data from vibrio-phage systems, with time-series sampling providing crucial insights into their temporal and spatial dynamics. A major finding is the role of mobile genetic elements (MGEs) in encoding bacterial defence systems, which shape phage-host coevolution. Discoveries like the phage satellite PICMI illustrate how MGEs facilitate the transfer of antiviral systems, influencing ecological and evolutionary dynamics. The paradox of generalist vibriophages, rare despite their broad host ranges, is also explored. By integrating experimental approaches with field observations, vibriophage research advances microbial ecology and informs sustainable applications in aquaculture and phage therapy, reinforcing vibrios as a versatile model system.This article is part of the discussion meeting issue 'The ecology and evolution of bacterial immune systems'.

Bacteriophages↗

Identification and characterization of ectopic chromosomal amplifications in acute myeloid leukemia cell limes using high-throughput chromosome conformation capture screening.

Despite advanced molecular diagnostics, improving outcomes for refractory acute myeloid leukemia (AML) remains challenging. Although many cancer-related genes are identified, their molecular mechanisms are not fully elucidated. Amplification is a mechanism of cancer-associated gene activation, and ectopic gene amplification may have particularly high pathological significance. However, research on ectopically amplified cancer-associated genes in leukemia remains limited. Here, we evaluated the usefulness of high-throughput chromosomal conformation capture (Hi-C) as a screening method for ectopic gene amplification and assessed whether ectopic amplification of cancer-associated genes may represent a general phenomenon in AML. We screened the U-937 and NB-4 cell lines using in situ Hi-C. Regions appearing as "high-intensity bands" in Hi-C contact maps were identified and validated using fluorescence in situ hybridization (FISH). Additionally, copy number variation analysis was performed using whole-genome sequencing (WGS) to extract cancer-associated genes with ectopic amplification. In the U-937, three genomic regions showing "high-intensity bands" were identified and confirmed as ectopic amplifications-including PDCD1LG2 (PD-L2), CD274 (PD-L1), and JAK2; that is, four copies were detected by WGS, and amplification signals were observed by FISH. In the NB-4, four such regions were detected, including MYC and KRAS, with expression level of 498 transcripts per million (TPM) and 34 TPM, respectively. Copy number variation analysis further identified multiple cancer-associated genes with ectopic amplification. Overall, these findings demonstrate the presence of ectopic amplification of cancer-associated genes in AML cell lines and support the usefulness of Hi-C as a screening method for detecting such genomic alterations.

Acute myeloid leukemia↗

Genome-wide analysis suggests a differential microRNA signature associated with normal and diabetic human corneal limbus.

Small non-coding RNAs, in particular microRNAs (miRNAs), regulate fine-tuning of gene expression and can impact a wide range of biological processes. However, their roles in normal and diseased limbal epithelial stem cells (LESC) remain unknown. Using deep sequencing analysis, we investigated miRNA expression profiles in central and limbal regions of normal and diabetic human corneas. We identified differentially expressed miRNAs in limbus vs. central cornea in normal and diabetic (DM) corneas including both type 1 (T1DM/IDDM) and type 2 (T2DM/NIDDM) diabetes. Some miRNAs such as miR-10b that was upregulated in limbus vs. central cornea and in diabetic vs. normal limbus also showed significant increase in T1DM vs. T2DM limbus. Overexpression of miR-10b increased Ki-67 staining in human organ-cultured corneas and proliferation rate in cultured corneal epithelial cells. MiR-10b transfected human organ-cultured corneas showed downregulation of PAX6 and DKK1 and upregulation of keratin 17 protein expression levels. In summary, we report for the first time differential miRNA signatures of T1DM and T2DM corneal limbus harboring LESC and show that miR-10b could be involved in the LESC maintenance and/or their early differentiation. Furthermore, miR-10b upregulation may be an important mechanism of corneal diabetic alterations especially in the T1DM patients.

Adult↗

Nanobodies: From High-Throughput Identification to Therapeutic Development.

The camelid single-domain antibody fragment, commonly referred to as a nanobody, achieves the targeting power of conventional monoclonal antibodies (mAbs) at only a fraction of their size. Isolated from camelid species (including llamas, alpacas, and camels), their small size at ∼15 kDa, low structural complexity, and high stability compared with conventional antibodies have propelled nanobody technology into the limelight of biologic development. Nanobodies are proving themselves to be a potent complement to traditional mAb therapies, showing success in the treatment of, for example, autoimmune diseases and cancer, and more recently as therapeutic options to treat infectious diseases caused by rapidly evolving biological targets such as the SARS-CoV-2 virus. This review highlights the benefits of applying a proteomic approach to identify diverse nanobody sequences against a single antigen. This proteomic approach coupled with conventional yeast/phage display methods enables the production of highly diverse repertoires of nanobodies able to bind the vast epitope landscape of an antigen, with epitope sampling surpassing that of mAbs. Additionally, we aim to highlight recent findings illuminating the structural attributes of nanobodies that make them particularly amenable to comprehensive antigen sampling and to synergistic activity-underscoring the powerful advantage of acquiring a large, diverse nanobody repertoire against a single antigen. Lastly, we highlight the efforts being made in the clinical development of nanobodies, which have great potential as powerful diagnostic reagents and treatment options, especially when targeting infectious disease agents.

Animals↗

A 29-plex MOL-PCR assay for simultaneous detection of selected major, non-typing, and accessory virulence genes in Clostridium perfringens.

Clostridium perfringens is an important pathogen of humans and animals, responsible for a broad spectrum of diseases mediated by diverse toxins and virulence factors. Precise and extended toxin-gene profiling is valuable for strain characterization and molecular epidemiological surveillance. Here, we describe the development of a 29-plex Multiple Oligonucleotide Ligation PCR (MOL-PCR) assay that enables the simultaneous detection of a large and important panel of 27 C. perfringens toxin-related genes - covering major typing toxins as well as an extended panel of non-typing and accessory virulence genes - thus moving beyond the classical toxinotyping framework. The assay was evaluated in comparison with six multiplex qPCR assays. In both systems, the gene encoding the Clostridium perfringens-specific serine O-acetyltransferase (EpsC) was used as a molecular marker for species confirmation, and an internal amplification control was included to detect potentially false-negative results. Analytical specificity testing confirmed exclusive amplification in C. perfringens and sequencing confirmed the toxin-gene profiles of reference strains. Comparative analysis of 72 reference and field isolates (1,944 data points) demonstrated complete concordance for 637 positive detections, yielding 100% positive agreement and 99.7% negative agreement relative to the comparative qPCR method. The limit of detection was 100 fg/µl (approx. 3 × 101 genome equivalents; GE) for qPCR and 1  pg/µl (approx. 3 × 102 GE) for MOL-PCR. Despite its high multiplex level, MOL-PCR showed high agreement with qPCR. The developed MOL-PCR method provides a rapid, high-throughput, and cost-effective tool for expanded toxin-gene profiling of C. perfringens isolates targeting major typing toxins and selected non-typing and accessory virulence genes. Therefore, it may support advanced toxin-gene characterization, molecular epidemiology, and One Health-oriented surveillance of evolving virulence landscapes.

Clostridium perfringens↗

A pluripotent stem cell atlas of multilineage differentiation.

Human pluripotent stem cells offer a scalable platform to study genetic and signalling mechanisms governing cell lineage decisions during differentiation. Genome-wide and single-cell transcriptomics technologies likewise offer high-throughput analysis of heterogeneous cell differentiation states. While in vivo development has been extensively characterised using these technologies, there remains a need for comprehensive single-cell transcriptomic profiling of stem cell differentiation from pluripotency. Understanding gene expression changes governing differentiation in vitro is key to developing high fidelity differentiation protocols and understanding fundamental mechanisms of development. We generated a single-cell RNA sequencing time course to study the role of developmental signalling pathways on multilineage diversification from pluripotency in vitro. The combined dataset of over 60,000 cells spans cell types from a time course of differentiation across all germ layers, ranging from gastrulation cell states to progenitor and committed cell types. These data provide a diverse benchmarking reference point to compare against in vivo development and advance understanding of signalling regulation of differentiation, providing insights into protocol development, drug screening, and regenerative medicine applications.

Pluripotent Stem Cells↗

HI-FEVER: a Nextflow pipeline for the high-throughput discovery and annotation of endogenous viral elements.

SUMMARY: Endogenous viral elements (EVEs) offer valuable insights into virus and host evolution, but their detection remains computationally and biologically challenging. We present HI-FEVER, a user-friendly Nextflow pipeline for the discovery of EVEs in eukaryotic host genomes. HI-FEVER is highly parallelizable and customizable, ensuring computational efficiency while allowing researchers to fine-tune parameters to their specific needs. Its output provides a comprehensive analysis of discovered EVEs, including detailed annotations which can provide evolutionary insights. HI-FEVER scales seamlessly to handle millions of viral protein queries across multiple host genomes on both laptops and high-performance computing nodes. AVAILABILITY AND IMPLEMENTATION: The HI-FEVER source code is available on GitHub at https://github.com/PaleovirologyLab/hi-fever. Minimal reference databases, test datasets and benchmarking results are hosted on the Open Science Framework at https://osf.io/y357r. A detailed wiki is available at https://github.com/PaleovirologyLab/hi-fever/wiki, including usage instructions, parameter descriptions, and guidance on interpreting outputs. The pipeline includes a Pixi environment compatible with Conda and Apptainer containerization, and Docker images. HI-FEVER has been tested on Linux, Windows (via WSL2), and macOS (Intel and ARM64).

Software↗

Advances in the diagnosis and classification of B-ALL: comparative insights from updated guidelines.

Accurate molecular classification is essential for diagnosis, risk stratification, and treatment selection in B-cell lymphoblastic leukemia (B-ALL). In this study, we performed a comprehensive, real-world reclassification of 1015 consecutively diagnosed B-ALL patients using the fifth edition of the World Health Organization Classification of Haematolymphoid Tumours (WHO-HAEM5) and the International Consensus Classification (ICC). An integrative genomic strategy that combined whole transcriptome sequencing, fusion detection, mutational analysis, and cytogenetics enabled reclassification according to both the WHO-HAEM5 and ICC frameworks, thereby substantially reducing the proportion of unclassifiable B-ALL from 41.9% (2016 WHO revision [WHO-HAEM4R]) to 15.9% (WHO-HAEM5) and 11.9% (ICC). Distinct clinical and prognostic features were identified across newly defined subtypes. Multivariable analysis confirmed that this genomic classification is a robust, independent predictor of survival after adjusting for age, minimal residual disease status, and transplant intervention. Specifically, HLF-rearranged and MEF2D-rearranged B-ALL conferred a persistently poor prognosis across all age groups despite allogeneic hematopoietic stem cell transplantation, highlighting an urgent need for novel therapeutic strategies. Gene expression profiling resolved cryptic subtypes, including ETV6::RUNX1-like, ZNF384-rearranged-like, and BCR::ABL1-like B-ALL, and uncovered diagnostic ambiguity in patients with concurrent lesions. In addition, we report emerging high-risk groups, including IDH1/2- and ZEB2 Q1072-mutated B-ALL, that may warrant recognition as distinct molecular entities. Our findings demonstrate the clinical use of integrative transcriptomic profiling in refining B-ALL taxonomy in guiding risk-adapted therapies and informing future revisions of diagnostic standards. This study supports the incorporation of high-throughput molecular diagnostics into routine leukemia classification and precision treatment planning.

Humans↗

[Applications and Challenges of Deep Learning in Human Genome Research].

In recent years, the advent of high-throughput omics technologies has fueled an explosive growth in human genomic data. Uncovering the latent functions within this vast data has become a significant challenge in functional genomics research. While traditional statistical methods have proved successful for analyzing smaller-scale datasets in the past, they exhibit clear limitations in analytical efficiency and integrating multi-dimensional data, struggling to meet the escalating demands of contemporary genomic analysis. The introduction of deep learning (DL) technologies offers a novel paradigm for this field. This review systematically examines the advances in applying deep learning to human genomics research. Studies demonstrate that when ample labeled data is available, discriminative DL computational methods-such as Convolutional Neural Networks (CNNs) and Long Short-Term Memory networks (LSTMs)-achieve high accuracy and efficiency in genomic variant discovery tasks. Furthermore, generative DL methods, particularly Large Language Models (LLMs) leveraging self-supervised pre-training strategies, effectively integrate complex genomic information and exhibit superior performance in functional genomic sequence annotation and gene regulation studies. This review also explores the application of LLMs in multi-omics data integration and prediction. Looking ahead, the continued accumulation of long-read sequencing and high-dimensional data is expected to enable DL technologies to integrate increasingly complex and heterogeneous genomic information, playing an increasingly crucial role in human genomics research.

Deep Learning↗

Development of a PCR-based technique for genotyping UGT1A1 gene and distribution of rs3064744 alleles in the Russian population.

BACKGROUND: Accurate determination of tandem thymine-adenine (TA) repeat numbers in the UGT1A1 promoter region (rs3064744) is essential for diagnosing Gilbert's syndrome and personalizing therapy with toxic agents like irinotecan and atazanavir. However, traditional polymerase chain reaction (PCR) assays face severe limitations due to the AT-rich sequence and overlapping melting temperatures (Tm) of the highly homologous 7TA and 8TA alleles. In this context, melting curve analysis (MCA) employing fluorophore-quencher systems has emerged as a promising alternative. The purpose of this study was to develop a novel genotyping approach combining optimized aPCR-MCA analysis with an automated classifier to overcome the limitations posed by the differentiation of highly homologous alleles and to demonstrate its practical application, providing the distribution of rs3064744 genotypes across four regional cohorts of the Russian population. METHODS: A specialized Dual Head 1D-convolutional neural network (1D-CNN) ensemble with Test-Time Augmentation (TTA) was developed. The model was trained and internally validated on 1,620 engineered plasmid samples, and independently evaluated on an external clinical test set of 440 unique patient genomic DNA specimens. Real-time PCR was performed on CFX96 and DTprime platforms. Additionally, population-wide screening was conducted on 997 archival clinical samples from Moscow, Sakha (Yakutia), Dagestan, and Rostov regions. RESULTS: While 5TA and 6TA alleles were easily separated, absolute Tm distributions of 7TA and 8TA alleles overlapped significantly, and non-uniform Tm shifts of 0.8 °C-1.4 °C occurred across platforms. Conventional absolute Tm thresholding was therefore inadequate. By assessing relative morphological curve divergence against co-amplified 7TA/7TA and 7TA/8TA reference anchors, the 1D-CNN ensemble neutralized instrument noise. It achieved 100% accuracy on internal validation and 100% concordance (440/440) with clinical reference pyrosequencing. Population screening revealed that Dagestan, Yakutia, and Rostov cohorts closely align with the European population. Rare 5TA and 8TA alleles were detected at low frequencies in Yakutia and Moscow. CONCLUSION: Combining LNA-modified aPCR-MCA with a comparative 1D-CNN model successfully circumvents thermodynamic limitations and eliminates human operator bias. This integrated system offers an accessible, high-throughput, and clinically valid solution for routine UGT1A1 pharmacogenetic testing.

1D-CNN↗

Development of a 10K breeder-friendly SNP chip for faba bean.

INTRODUCTION: Faba bean breeding and genomics have seen steady progress in recent years, supported by genome sequences and high-density genotyping platforms. These tools have been valuable for trait mapping, diversity assessment, and genomic research, but they have limited routine use in breeding programs due to their relatively high cost. Recent progress in establishing an optimized, cost-efficient genotyping-by-sequencing protocol tailored to the large and complex faba bean genome has created the foundation for a more accessible genotyping solution. METHODS: Using this approach, we explored the genetic diversity of faba bean germplasm from various panels, providing a comprehensive representation of the crop's genetic landscape. From this dataset, we identified and selected a high-quality set of informative SNP markers that are evenly distributed across the genome. Building on these resources, we designed a breeder-friendly 10K SNP chip. RESULTS: The 10K SNP chip delivers high accuracy, broad genomic coverage, and affordability. The chip was validated across diverse germplasm panels, demonstrating strong clustering performance, high reproducibility, and applicability to breeding-relevant germplasm. DISCUSSION: This platform offers a cost-effective alternative to higher-density arrays, enabling its integration into genomic selection, marker-assisted breeding, and diversity monitoring, ultimately supporting accelerated genetic gain and the delivery of improved varieties to farmers.

SNP chip↗