PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “High-throughput sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

High-throughput recovery of integron cassettes for gene discovery screens.

Integrons capture functional genes in mobile genetic elements called integron cassettes, which represent an untapped source of genes of biotechnological interest. Here we present two tools, cassette gatherer and cassette hunter, that enable high-throughput establishment of gene libraries either from genetically tractable strains or directly from DNA. We re-engineered a class 1 integron into counterselection markers on a plasmid or on the chromosome of a naturally competent Vibrio cholerae, which enabled capture of single cassettes in a sequence- and function-independent manner. When applied to Vibrio strains and genomic libraries, our tools recovered hundreds of single cassettes per assay with more than 99% specificity. We further subjected the library of cassettes generated by the hunter and gatherer tools to screens against phages ICP2 and T4, and identified nine phage-defence systems, including five previously undescribed. These tools enable rapid and large-scale recovery of integron cassettes that could be leveraged for functional gene discovery.

Journal Article↗

Herd-level heterogeneity of antimicrobial resistance in commensal Escherichia coli: A nationwide high-throughput survey of Australian pig herds.

Antimicrobial resistance in commensal Escherichia coli provides a useful indicator for overall antimicrobial resistance burden. We applied this approach to assess antimicrobial resistance within and between commercial pig herds across Australia. A high-throughput robotic workflow was used to isolate 2730 E. coli colonies from rectal contents collected in 2022 from healthy slaughter pigs (n = 300) representing 30 herds (∼70% of national production). Up to 94 isolates per herd underwent antimicrobial susceptibility testing using the Robotic Antimicrobial Susceptibility Platform. Isolate- and herd-level antimicrobial resistance indices were calculated, weighting antimicrobials by their human health importance. Resistance to first-line agents was widespread: ampicillin 77% and tetracycline 79%. By contrast, resistance to critically important antimicrobials was rare (ciprofloxacin 0.11%; extended-spectrum cephalosporins 0.04%), and no clinical resistance to carbapenems or colistin was detected. Overall, 56.9% of isolates were multi-class resistant. Herd-level antimicrobial resistance within indices ranged from 1.51 to 5.76, revealing substantial between-herd heterogeneity. Three herds carried critically important antimicrobials-resistant isolates that would likely have been missed using conventional, lower-density sampling approaches. Whole-genome sequencing identified fluoroquinolone-resistant isolates belonging to ST10 and ST69 (both qnrS1), and ST744 (Quinolone Resistance Determining Region mutations plus blaCTX-M-27). By testing approximately tenfold more isolates than conventional surveys, we uncovered considerable antimicrobial resistance with heterogeneity within and between animals and herds, including farm-specific variability. This expanded sampling also enabled detection of critically important antimicrobial resistance at very low prevalence. In conclusion, high-throughput, high-density testing offers a practical early-warning system and herd-level benchmark to inform surveillance and targeted interventions.

Animals↗

PyEvoMotion: a Python tool for population-based time-course analysis of genome evolution.

SUMMARY: We present PyEvoMotion, an open-source Python tool for inferring molecular clock models with time-dependent Gaussian noise from high-throughput genomic datasets. PyEvoMotion features a command-line interface and a modular architecture, allowing seamless integration into larger bioinformatic pipelines. The tool supports customizable filtering, temporal discretization definition, and mutation classification, making it adaptable to diverse research needs. While traditional phylogenetic methods may encounter computational challenges with large datasets, PyEvoMotion can process thousands to millions of sequences to compute statistical parameters associated with a stochastic differential equation model, thereby weighting the genetic variation within the population. Using viral genomic data, we demonstrate its capability to infer evolutionary rates and detect non-Brownian evolutionary motions with subdiffusive behavior. PyEvoMotion shows potential to provide overlooked insights into genome evolution in different contexts. AVAILABILITY AND IMPLEMENTATION: The open source software is available on GitHub at https://github.com/luksgrin/PyEvoMotion and on SourceForge at https://sourceforge.net/projects/pyevomotion.

Software↗

An Amplicon Panel for High-Throughput and Low-Cost Genotyping of Yesso Scallop Mizuhopecten yessoensis.

The Yesso scallop Mizuhopecten yessoensis was imported from Japan to western Canada in the late 1980s to establish an economically viable scallop aquaculture industry. Since this time, the industry in Canada has operated with existing genetic diversity within the broodstock, which is considerably limited relative to wild populations. The sector has not been able to realise its full potential in part due to idiopathic hatchery failures and farm stock collapses due to disease outbreaks associated with the intracellular bacterial pathogen Francisella halioticida. To support Yesso scallop production and breeding, here we generate a low-density, genotyping-by-sequencing amplicon panel using single nucleotide polymorphism (SNP) markers that are evenly spaced across the M. yessoensis genome and that show high heterozygosity in Canada and Japan. The panel can also exploit the high genetic polymorphism of the M. yessoensis genome, with de novo SNP calling identifying over 2,500 high quality SNPs within the 579 sequenced amplicons. We demonstrate the utility and versatility of this new genotyping tool for breeding applications including parentage assignment, low density family-based genome-wide association study, trait heritability evaluation to determine potential for genomic selection, and species differentiation (against the weathervane scallop Patinopecten caurinus). We did not find any genomic regions significantly associated with F. halioticida resistance but did identify potential for genomic selection. We could separate the two species based on genotypes, and did not see evidence of a past M. yessoensis x P. caurinus hybridization event within the M. yessoensis breeding population at Vancouver Island University. This low-cost genotyping panel is expected to accelerate selective breeding improvements for M. yessoensis in Canada and elsewhere.

Animals↗

Mapping ovarian cellular and molecular landscape across the lifespan of women: a scoping review.

BACKGROUND: With growing interest in ART, fertility preservation, and postmenopausal health of women, reproductive medicine is increasingly focused on characterizing oocytes and ovarian tissue composition, as well as understanding the molecular mechanisms that guide ovarian function throughout its lifecycle. High-throughput omics technologies have enabled the characterization of different molecular layers, leading to substantial advances in our understanding of their complex dynamics. However, not all molecular aspects are studied equally, and studies examining the same modalities often show inconsistencies, underscoring the need for data standardization and highlighting the potential for using transformative artificial intelligence and machine-learning (AI/ML) methods for ovary studies. OBJECTIVE AND RATIONALE: This study aims to evaluate how multi-omic studies have advanced our understanding of the ovarian lifecycle from fetal development to postmenopause. We systematically reviewed published studies that have investigated molecular/omic layers, including the genome, methylome, transcriptome, and proteome throughout ovarian development and aging. Our analysis identified key molecular and cellular patterns, highlighted inconsistencies across studies and addressed gaps in data analysis, interpretation, and reproducibility to guide future research. SEARCH METHODS: We conducted a systematic literature search of Medline (PubMed), Embase (Ovid), and Web of Science Core Collection (Clarivate) using a combination of controlled and free text terms for human ovary, oogenesis, folliculogenesis, ovary development and (epi)genome, transcriptome, proteome, and multi-omic mechanisms to find relevant articles published before August 2025. To focus the scope of the current review, studies of domesticated and farm animals, rodents and other model organisms, non-human primates, as well as those examining various human ovarian pathologies were excluded. OUTCOMES: The search identified 23 546 studies for screening, of which 637 full-text studies were assessed for eligibility. Subsequently, we extracted data from 121 studies. Most studies analyzed the transcriptome of oocytes, granulosa cells, and ovarian tissue from reproductive-age individuals (n = 91), with fewer studies examining samples from individuals of advanced reproductive age (n = 45) and fetal (n = 16) samples. Transcriptome analyses were most common (n = 103, 85%), followed by proteome (n = 19, 16%) and epigenome (n = 14, 12%) studies. We found substantial variation in how studies defined and reported participants' groups as well as in their sequencing technologies and data analysis methods, with a lack of standardized reporting of background clinical information, data analysis methods, and pipeline details. The key findings underscore the prevailing consensus on genes defining major ovarian cell types and their roles throughout the ovarian lifespan, from prenatal development to postmenopausal transformation. This review highlighted the underrepresentation of certain patient groups, particularly prepubertal and peri-/postmenopausal individuals, among researched populations, due to obvious clinical and ethical reasons. WIDER IMPLICATIONS: This scoping review offers a comprehensive overview and benchmark of the current state of high-throughput omics-based research on ovarian cellular composition and molecular dynamics. To address these shortcomings, we propose general recommendations for multi-omics ovary studies and emphasize the necessity for more thorough multi-omic data integration by effectively applying novel AI/ML approaches. They can potentially improve the quality of multi-omics analyses at both single-cell and tissue levels despite limited sample sizes and enable integration of molecular profiling data with clinical and radiology datasets, enabling a more comprehensive understanding of ovarian biology. Such advancements can enhance reproducibility of research findings and guide future research to deepen our understanding of ovarian biology and ultimately support the development of medical technologies for better preserving fertility and alleviating infertility. REGISTRATION NUMBER: A protocol was published a priori on the Open Science Framework (https://osf.io/z38gb/).

Female↗

Exome sequencing and large-scale analysis of electronic medical record-linked biobank data identify candidate deafness genes.

INTRODUCTION: Rapid advances in whole-exome sequencing (WES) have enabled large-scale detection of pathogenic variants. Although hundreds of genes are implicated in hearing loss, up to half of inherited cases remain unsolved, limiting eligibility for gene therapy trials that require genetic diagnosis. Biobanks and electronic medical records (EMRs) offer opportunities to integrate genomic and clinical data at scale and expand the spectrum of hearing loss genes. Despite clinical value, EMRs often lack key information such as inheritance patterns, posing challenges for accurate interpretation. METHODS: WES was performed on DNA samples from 1038 hearing-impaired patients enrolled in the Maccabi Research and Innovation Center Tipa Biobank. Clinical data were extracted from EMRs. Audiograms were available for all cases, although data on age of onset, family history and mode of inheritance were mostly unavailable. We applied a scalable bioinformatics analysis strategy for high-throughput annotation, filtering and prioritisation of WES variants across more than 1000 patients, designed to accommodate incomplete and heterogeneous clinical records. RESULTS: Using this approach, 15% of cases were solved or potentially solved through known or novel variants in established deafness genes. Homozygous variants in novel candidate genes were identified in 3% of cases. Functional characterisation was performed for promising candidate genes to validate their role in the ear. CONCLUSION: These findings demonstrate that WES can determine disease aetiology in large, genetically heterogeneous populations, even in the context of incomplete clinical data. This approach supports large-scale genetic screening and provides a framework for identifying patients who may benefit from emerging gene-based therapies.

Genetic Testing↗

A streamlined protocol for small-scale protoplast generation and CRISPR/Cpf1-mediated genome editing in Fusarium oxysporum.

Fusarium oxysporum is a significant threat to agriculture and One Health, requiring advanced molecular tools for functional genomic analyses and biological control agent development. Existing gene-editing methods are hampered by costly protoplast preparation protocols and by CRISPR-Cas9 limitations, such as restricted protospacer adjacent motif (PAM) sequences and complex guide RNA requirements. We engineered an efficient CRISPR/Cpf1 system that overcomes these issues through three main innovations: small-scale protoplast generation using filter column-based methods that greatly reduce enzyme consumption while simplifying workflows, a CRISPR/Cpf1 system with shorter guide RNA design and staggered DNA cleavage to promote homologous recombination, and minimal homology arm strategies that significantly decrease cloning complexity. Extensive validation confirms successful gene targeting with molecular verification and functional analysis via standardized pathogenicity assays. This integrated platform offers affordable, accessible tools for systematic F. oxysporum research, enhancing fundamental understanding of plant-pathogen interactions and supporting high-throughput screening vital for agricultural biotechnology and biological agent development.

CRISPR/Cpf1↗

Enhanced identification of key bacterial motility genes via a cross-species genomic hybrid feature machine learning approach.

Efficient and accurate identification of functional genes is critical to biological research, yet traditional single-species approaches are often limited by low efficiency. Previously, we established a novel method for identifying key genes using cross-species protein domain features and machine learning. However, the high multiplicity of gene members associated with specific domains creates a substantial workload for subsequent experimental validation. To address this, this study proposes an enhanced approach that integrates EggNOG-based protein sequence annotation with domain analysis. Unannotated sequences are subsequently analyzed for protein domains, generating a comprehensive "direct gene annotation plus domain" hybrid feature matrix. While the hybrid matrix model yielded comparable predictive accuracy, it significantly enhanced feature resolution: the top 50 predicted features were all known motility-related genes or domains. Furthermore, among the top 100 ranked features, 58 are confirmed to be directly related to motility based on experimental evidence. Although strict genus-level control still yielded 51 confirmed features, excessive taxonomic restriction drastically reduces the number of training genomes, which may paradoxically impair identification efficiency. These results demonstrate that the new method effectively reduces the subsequent experimental workload and enables high-throughput identification of functional genes in a single analysis. With accuracy and efficiency far exceeding those of existing single-species identification methods, it provides a highly efficient solution for mining key genes underlying other complex bacterial phenotypes.

Machine Learning↗

2-Mercaptoethanol/DMSO Workflow Enables Highly Reproducible Quantitative Proteomics.

Proteomics provides a systematic and high-throughput approach to comprehensively characterize protein networks, enabling insights into cellular functions and disease mechanisms. Carbamidomethylation using iodoacetamide (IAA), a common method for cysteine alkylation, is known to cause nonspecific modifications that increase spectral complexity in mass spectrometry and reduce quantitative accuracy. Here, we established a reproducibility-focused 2-mercaptoethanol (2-ME)/dimethyl sulfoxide (DMSO) workflow and systematically evaluated its quantitative performance at the proteome-wide level. Mouse liver proteomes were processed using either 2-ME/DMSO or conventional IAA treatment, followed by liquid chromatography-tandem mass spectrometry (LC-MS/MS) analysis. The optimized 2-ME treatment increased the number of cysteine-modified peptides by 1.6- to 1.9-fold. Although total protein identifications were comparable, 77% of proteins exhibited improved sequence coverage with the optimized 2-ME treatment. Quantitative reproducibility was also enhanced, with the peptide quantified CV ≤ 20% increasing from 61.4% with IAA treatment to 86.1% with 2-ME treatment, and protein quantified CV ≤ 20% increasing from 80.6% with IAA treatment to 93.5% with 2-ME treatment. Application of this new workflow to ovarian clear cell carcinoma reliably detected cisplatin-induced alterations. The 2-ME/DMSO workflow offers a simple and highly reproducible proteomics strategy for accurate quantitative proteomics.

Animals↗

Multi-Omics and Integrative Analytics in Natural Products Discovery.

Natural products (NPs) have long been an essential source of new bioactive compounds for drug discovery; however, traditional methods for screening and isolating these compounds can be slow and often yield diminishing returns. Fortunately, advanced multi-omics and computational approaches present powerful solutions to these challenges. This review highlights innovative methodologies that integrate metabolomics, genomics, transcriptomics, and proteomics with bioinformatics and analytical chemistry to accelerate NP discovery. For instance, untargeted metabolomics platforms like high-resolution liquid chromatography-tandem mass spectrometry (LC-MS/MS) and Global Natural Products Social (GNPS) molecular networking allow for comprehensive profiling of new compounds, while targeted isotope-labeling strategies enhance this process. Additionally, genome and metagenome mining tools such as antibiotics and secondary metabolite analysis shell (antiSMASH), Deep Biosynthetic Gene Cluster (DeepBGC), and Pipeline for Reconstructing Integrated Syntheses of Metabolites (PRISM) quickly identify biosynthetic gene clusters (BGCs) in both cultured and uncultured organisms, often using heterologous expression to validate products. Transcriptomic analyses, including RNA sequencing (RNA-seq), co-expression networks, and fluxomics, help clarify how pathways are regulated, while quantitative proteomics techniques like tandem mass tags/isobaric tags for relative and absolute quantitation (TMT/iTRAQ) and label-free methods, along with chemoproteomics approaches such as cellular thermal shift assay and thermal proteome profiling (TPP), uncover molecular targets and their mechanisms of action. This review also places significant emphasis on the role of artificial intelligence (AI) and machine learning (ML) in integrating multi-omics data, spanning activities from constructing gene-metabolite correlation networks to leveraging knowledge graphs and graph neural networks for data fusion and functional prediction. Finally, this review concludes by discussing the synergistic benefits of multi-omics for natural-product discovery, addressing current technical challenges, and exploring future directions toward high-throughput, intelligent data integration for next-generation NP research.

Biological Products↗

The ecology and evolution of microbial immune systems: a look on the wild vibrio side.

Natural populations of vibrio beyond the well-studied pandemic strains of Vibrio cholerae, provide a powerful model for investigating the eco-evolutionary dynamics of microbial immune systems. Their genetic diversity, ecological versatility, ease of culturability and the availability of time-series data enable detailed studies of phage-host interactions in natural contexts. This review synthesizes recent advances in vibriophage research, highlighting key findings and emerging tools. High-throughput assays and genomic tools have offered new perspectives on phage specificity, host range and the evolutionary pressures shaping these interactions. Theoretical frameworks, such as arms race and fluctuating selection dynamics, are informed by empirical data from vibrio-phage systems, with time-series sampling providing crucial insights into their temporal and spatial dynamics. A major finding is the role of mobile genetic elements (MGEs) in encoding bacterial defence systems, which shape phage-host coevolution. Discoveries like the phage satellite PICMI illustrate how MGEs facilitate the transfer of antiviral systems, influencing ecological and evolutionary dynamics. The paradox of generalist vibriophages, rare despite their broad host ranges, is also explored. By integrating experimental approaches with field observations, vibriophage research advances microbial ecology and informs sustainable applications in aquaculture and phage therapy, reinforcing vibrios as a versatile model system.This article is part of the discussion meeting issue 'The ecology and evolution of bacterial immune systems'.

Bacteriophages↗

Identification and characterization of ectopic chromosomal amplifications in acute myeloid leukemia cell limes using high-throughput chromosome conformation capture screening.

Despite advanced molecular diagnostics, improving outcomes for refractory acute myeloid leukemia (AML) remains challenging. Although many cancer-related genes are identified, their molecular mechanisms are not fully elucidated. Amplification is a mechanism of cancer-associated gene activation, and ectopic gene amplification may have particularly high pathological significance. However, research on ectopically amplified cancer-associated genes in leukemia remains limited. Here, we evaluated the usefulness of high-throughput chromosomal conformation capture (Hi-C) as a screening method for ectopic gene amplification and assessed whether ectopic amplification of cancer-associated genes may represent a general phenomenon in AML. We screened the U-937 and NB-4 cell lines using in situ Hi-C. Regions appearing as "high-intensity bands" in Hi-C contact maps were identified and validated using fluorescence in situ hybridization (FISH). Additionally, copy number variation analysis was performed using whole-genome sequencing (WGS) to extract cancer-associated genes with ectopic amplification. In the U-937, three genomic regions showing "high-intensity bands" were identified and confirmed as ectopic amplifications-including PDCD1LG2 (PD-L2), CD274 (PD-L1), and JAK2; that is, four copies were detected by WGS, and amplification signals were observed by FISH. In the NB-4, four such regions were detected, including MYC and KRAS, with expression level of 498 transcripts per million (TPM) and 34 TPM, respectively. Copy number variation analysis further identified multiple cancer-associated genes with ectopic amplification. Overall, these findings demonstrate the presence of ectopic amplification of cancer-associated genes in AML cell lines and support the usefulness of Hi-C as a screening method for detecting such genomic alterations.

Acute myeloid leukemia↗

A pluripotent stem cell atlas of multilineage differentiation.

Human pluripotent stem cells offer a scalable platform to study genetic and signalling mechanisms governing cell lineage decisions during differentiation. Genome-wide and single-cell transcriptomics technologies likewise offer high-throughput analysis of heterogeneous cell differentiation states. While in vivo development has been extensively characterised using these technologies, there remains a need for comprehensive single-cell transcriptomic profiling of stem cell differentiation from pluripotency. Understanding gene expression changes governing differentiation in vitro is key to developing high fidelity differentiation protocols and understanding fundamental mechanisms of development. We generated a single-cell RNA sequencing time course to study the role of developmental signalling pathways on multilineage diversification from pluripotency in vitro. The combined dataset of over 60,000 cells spans cell types from a time course of differentiation across all germ layers, ranging from gastrulation cell states to progenitor and committed cell types. These data provide a diverse benchmarking reference point to compare against in vivo development and advance understanding of signalling regulation of differentiation, providing insights into protocol development, drug screening, and regenerative medicine applications.

Pluripotent Stem Cells↗

HI-FEVER: a Nextflow pipeline for the high-throughput discovery and annotation of endogenous viral elements.

SUMMARY: Endogenous viral elements (EVEs) offer valuable insights into virus and host evolution, but their detection remains computationally and biologically challenging. We present HI-FEVER, a user-friendly Nextflow pipeline for the discovery of EVEs in eukaryotic host genomes. HI-FEVER is highly parallelizable and customizable, ensuring computational efficiency while allowing researchers to fine-tune parameters to their specific needs. Its output provides a comprehensive analysis of discovered EVEs, including detailed annotations which can provide evolutionary insights. HI-FEVER scales seamlessly to handle millions of viral protein queries across multiple host genomes on both laptops and high-performance computing nodes. AVAILABILITY AND IMPLEMENTATION: The HI-FEVER source code is available on GitHub at https://github.com/PaleovirologyLab/hi-fever. Minimal reference databases, test datasets and benchmarking results are hosted on the Open Science Framework at https://osf.io/y357r. A detailed wiki is available at https://github.com/PaleovirologyLab/hi-fever/wiki, including usage instructions, parameter descriptions, and guidance on interpreting outputs. The pipeline includes a Pixi environment compatible with Conda and Apptainer containerization, and Docker images. HI-FEVER has been tested on Linux, Windows (via WSL2), and macOS (Intel and ARM64).

Software↗

Advances in the diagnosis and classification of B-ALL: comparative insights from updated guidelines.

Accurate molecular classification is essential for diagnosis, risk stratification, and treatment selection in B-cell lymphoblastic leukemia (B-ALL). In this study, we performed a comprehensive, real-world reclassification of 1015 consecutively diagnosed B-ALL patients using the fifth edition of the World Health Organization Classification of Haematolymphoid Tumours (WHO-HAEM5) and the International Consensus Classification (ICC). An integrative genomic strategy that combined whole transcriptome sequencing, fusion detection, mutational analysis, and cytogenetics enabled reclassification according to both the WHO-HAEM5 and ICC frameworks, thereby substantially reducing the proportion of unclassifiable B-ALL from 41.9% (2016 WHO revision [WHO-HAEM4R]) to 15.9% (WHO-HAEM5) and 11.9% (ICC). Distinct clinical and prognostic features were identified across newly defined subtypes. Multivariable analysis confirmed that this genomic classification is a robust, independent predictor of survival after adjusting for age, minimal residual disease status, and transplant intervention. Specifically, HLF-rearranged and MEF2D-rearranged B-ALL conferred a persistently poor prognosis across all age groups despite allogeneic hematopoietic stem cell transplantation, highlighting an urgent need for novel therapeutic strategies. Gene expression profiling resolved cryptic subtypes, including ETV6::RUNX1-like, ZNF384-rearranged-like, and BCR::ABL1-like B-ALL, and uncovered diagnostic ambiguity in patients with concurrent lesions. In addition, we report emerging high-risk groups, including IDH1/2- and ZEB2 Q1072-mutated B-ALL, that may warrant recognition as distinct molecular entities. Our findings demonstrate the clinical use of integrative transcriptomic profiling in refining B-ALL taxonomy in guiding risk-adapted therapies and informing future revisions of diagnostic standards. This study supports the incorporation of high-throughput molecular diagnostics into routine leukemia classification and precision treatment planning.

Humans↗

[Applications and Challenges of Deep Learning in Human Genome Research].

In recent years, the advent of high-throughput omics technologies has fueled an explosive growth in human genomic data. Uncovering the latent functions within this vast data has become a significant challenge in functional genomics research. While traditional statistical methods have proved successful for analyzing smaller-scale datasets in the past, they exhibit clear limitations in analytical efficiency and integrating multi-dimensional data, struggling to meet the escalating demands of contemporary genomic analysis. The introduction of deep learning (DL) technologies offers a novel paradigm for this field. This review systematically examines the advances in applying deep learning to human genomics research. Studies demonstrate that when ample labeled data is available, discriminative DL computational methods-such as Convolutional Neural Networks (CNNs) and Long Short-Term Memory networks (LSTMs)-achieve high accuracy and efficiency in genomic variant discovery tasks. Furthermore, generative DL methods, particularly Large Language Models (LLMs) leveraging self-supervised pre-training strategies, effectively integrate complex genomic information and exhibit superior performance in functional genomic sequence annotation and gene regulation studies. This review also explores the application of LLMs in multi-omics data integration and prediction. Looking ahead, the continued accumulation of long-read sequencing and high-dimensional data is expected to enable DL technologies to integrate increasingly complex and heterogeneous genomic information, playing an increasingly crucial role in human genomics research.

Deep Learning↗

Development of a PCR-based technique for genotyping UGT1A1 gene and distribution of rs3064744 alleles in the Russian population.

BACKGROUND: Accurate determination of tandem thymine-adenine (TA) repeat numbers in the UGT1A1 promoter region (rs3064744) is essential for diagnosing Gilbert's syndrome and personalizing therapy with toxic agents like irinotecan and atazanavir. However, traditional polymerase chain reaction (PCR) assays face severe limitations due to the AT-rich sequence and overlapping melting temperatures (Tm) of the highly homologous 7TA and 8TA alleles. In this context, melting curve analysis (MCA) employing fluorophore-quencher systems has emerged as a promising alternative. The purpose of this study was to develop a novel genotyping approach combining optimized aPCR-MCA analysis with an automated classifier to overcome the limitations posed by the differentiation of highly homologous alleles and to demonstrate its practical application, providing the distribution of rs3064744 genotypes across four regional cohorts of the Russian population. METHODS: A specialized Dual Head 1D-convolutional neural network (1D-CNN) ensemble with Test-Time Augmentation (TTA) was developed. The model was trained and internally validated on 1,620 engineered plasmid samples, and independently evaluated on an external clinical test set of 440 unique patient genomic DNA specimens. Real-time PCR was performed on CFX96 and DTprime platforms. Additionally, population-wide screening was conducted on 997 archival clinical samples from Moscow, Sakha (Yakutia), Dagestan, and Rostov regions. RESULTS: While 5TA and 6TA alleles were easily separated, absolute Tm distributions of 7TA and 8TA alleles overlapped significantly, and non-uniform Tm shifts of 0.8 °C-1.4 °C occurred across platforms. Conventional absolute Tm thresholding was therefore inadequate. By assessing relative morphological curve divergence against co-amplified 7TA/7TA and 7TA/8TA reference anchors, the 1D-CNN ensemble neutralized instrument noise. It achieved 100% accuracy on internal validation and 100% concordance (440/440) with clinical reference pyrosequencing. Population screening revealed that Dagestan, Yakutia, and Rostov cohorts closely align with the European population. Rare 5TA and 8TA alleles were detected at low frequencies in Yakutia and Moscow. CONCLUSION: Combining LNA-modified aPCR-MCA with a comparative 1D-CNN model successfully circumvents thermodynamic limitations and eliminates human operator bias. This integrated system offers an accessible, high-throughput, and clinically valid solution for routine UGT1A1 pharmacogenetic testing.

1D-CNN↗

Development of a 10K breeder-friendly SNP chip for faba bean.

INTRODUCTION: Faba bean breeding and genomics have seen steady progress in recent years, supported by genome sequences and high-density genotyping platforms. These tools have been valuable for trait mapping, diversity assessment, and genomic research, but they have limited routine use in breeding programs due to their relatively high cost. Recent progress in establishing an optimized, cost-efficient genotyping-by-sequencing protocol tailored to the large and complex faba bean genome has created the foundation for a more accessible genotyping solution. METHODS: Using this approach, we explored the genetic diversity of faba bean germplasm from various panels, providing a comprehensive representation of the crop's genetic landscape. From this dataset, we identified and selected a high-quality set of informative SNP markers that are evenly distributed across the genome. Building on these resources, we designed a breeder-friendly 10K SNP chip. RESULTS: The 10K SNP chip delivers high accuracy, broad genomic coverage, and affordability. The chip was validated across diverse germplasm panels, demonstrating strong clustering performance, high reproducibility, and applicability to breeding-relevant germplasm. DISCUSSION: This platform offers a cost-effective alternative to higher-density arrays, enabling its integration into genomic selection, marker-assisted breeding, and diversity monitoring, ultimately supporting accelerated genetic gain and the delivery of improved varieties to farmers.

SNP chip↗