PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “High-throughput sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Multichannel genomic recording of biological information with ENGRAM.

Molecular recording is an emerging paradigm for measuring biology over time. Enhancer-mediated genomic recording of activity in multiplex (ENGRAM) is a recently described synthetic biology circuit architecture that converts the transient activity of cis-regulatory elements (CREs) into stable genomic records that can be retrospectively recovered via DNA sequencing. Here we provide a step-by-step protocol for conducting ENGRAM experiments and analyzing the resulting data. We also describe key design considerations for ENGRAM recorders, summarize the strengths and limitations of ENGRAM, and highlight applications, including multiplex signal recording and high-throughput CRE screening. In contrast to other systems for DNA-based recording in mammalian systems, ENGRAM relies on prime editing-mediated insertions to record the activity of a given CRE, such that it is inherently multiplexable-for example, four-base-pair insertions can represent the activities of up to 256 distinct CREs. A further contrast lies with ENGRAM's compatibility with DNA Typewriter, which facilitates the capture of signal order. For users with basic skills in molecular biology, mammalian cell culture and DNA sequencing analysis, ENGRAM experiments can typically be completed within 5-6 weeks.

Genomics↗

Generation of spCAS9 expressing human mesenchymal stem cell line to study gene function during osteoblast differentiation.

Human bone marrow-derived stromal cells (hMSCs) are a great resource for studying how genes influence cell fate and differentiation into various cell types like osteoblasts, adipocytes, and chondrocytes, among other cell types. However, genetic manipulation of primary hMSCs has been challenging due to their short lifespan and cellular senescence after limited passaging. Their low and unstable transfection efficiency also complicates gene delivery or inactivation, hindering long-term functional studies. The limited lifespan has been effectively solved by immortalizing hMSCs with telomerase reverse transcriptase (hMSCs-TERT). The use of these cells is ideal for functional studies of osteoblast and adipocyte differentiation through genetic manipulation, providing a stable and reliable model. Here, we have engineered a stable CAS9 expressing hMSC-TERT cell line (hMSC-TERTCAS9) via lentiviral transduction. The constitutive expression of spCas9 enables efficient and reproducible gene editing. We demonstrate the potential of these hMSC-TERTCAS9 cells for generating gene disruptions using plasmid delivery of guide RNAs as a fast and efficient strategy for targeted genome editing. The edited cells can be sorted and expanded as single cells to obtain homogenous clonal cell lines with mono- as well as bi-allelic gene deletions, a crucial step for producing reliable experimental results. We further validate this cell line as a powerful tool for studying gene function during hMSC proliferation and differentiation, providing 3 distinct examples of its utility. Through the generation of indels, single-cell sorting, and clonal selection, we have efficiently inactivated the vitamin D receptor and created both larger (256 nucleotides) gene disruptions in Forkhead box protein O1 and precise removals of a small genomic sequence (73 nucleotides) coding for microRNA MIR675. This novel hMSC-TERTCAS9 cell line represents a significant advancement, offering a stable, efficient, and versatile platform for advanced genetic studies, high-throughput screening, and the creation of reliable cellular disease models.

CRISPR-Cas9↗

Plant cis-regulatory grammar: Decoding the multidimensional code of transcriptional regulation for programmable crop engineering.

Cis-regulatory elements (CREs) orchestrate the spatiotemporal precision of gene expression that underlies plant development, adaptation, and domestication. Decoding the cis-regulatory grammar of plant genomes remains a central challenge in modern biology, with profound implications for programmable crop engineering. Here, recent conceptual and technological advances are synthesized to reshape our understanding of plant CREs. This review first argues that CRE function is not only an intrinsic property of DNA sequence alone but also emerges from a multidimensional context, including chromatin accessibility, histone modifications, three-dimensional genome topology, and cell type-specific regulatory landscapes. Furthermore, the convergence of single-cell epigenomics, high-throughput functional assays, and CRISPR-based dissection has begun to unravel this contextual grammar, revealing the computational principles governing transcriptional regulation. Critically, we propose that artificial intelligence (AI) platforms are catalyzing an ongoing transition from descriptive discovery to predictive engineering, wherein these platforms outperform natural evolution in designing synthetic CREs. Finally, a roadmap is outlined toward a plant regulatory grammar foundation model, which will enable truly predictive engineering of gene expression when fine-tuned for specific tasks. Collectively, the integration of single-cell resolution maps, precise genome editing, AI-driven design, and regulatory-compliant delivery systems promises to transform our ability to reprogram plant gene regulation for next-generation agriculture, bridging the gap between foundational regulatory biology and tangible crop improvement.

artificial intelligence↗

Decoding sequence recognition code of nucleic acid-binding proteins of human-infecting DNA viruses.

Human-infecting DNA viruses remain major health threats, yet the DNA-recognition mechanisms of their nucleic acid-binding proteins (NBPs) are poorly understood. Here, we systematically profiled 103 viral NBPs from human-infecting DNA viruses, with three NBPs from non-human-infecting DNA viruses as controls, using high-throughput screening. This analysis identified diverse DNA-binding motifs and specificity modules, including convergent recognition of a conserved CCACC motif across phylogenetically distant viruses. Notably, viral NBP binding-site distributions varied with genome size, and several NBPs from small-genome viruses showed enrichment on mitochondrial DNA. Functional assays further supported their mitochondrial association and effects on mitochondrial membrane potential. By integrating an ivTRT-based ssDNA-SELEX workflow, we further found that ssDNA viral NBPs recognize dimer-like and inverted-repeat sequences with potential to form stem-loop structures. Collectively, this study constructs a comprehensive viral NBP DNA-recognition atlas, offering a fundamental resource for elucidating viral genome recognition mechanisms, virus-mitochondria interactions, and developing future antiviral strategies.

Letter↗

GHT-SELEX demonstrates unexpectedly high intrinsic sequence specificity and complex DNA binding of many human transcription factors.

There is ongoing debate regarding the degree to which transcription factors (TFs) independently specify genomic binding: TF binding motifs are typically short and degenerate, yielding many more binding site predictions than observed in cells. Here we present genomic high-throughput SELEX (GHT-SELEX)-a scalable method that surveys intrinsic binding of purified TFs to the fragmented, naked and unmodified genome. GHT-SELEX peaks for 179 diverse human TFs display surprisingly high overlap with chromatin immunoprecipitation sequencing peaks for the same TF. Comparable overlap can be obtained from motifs using appropriate analytical approaches. For C2H2 zinc finger (zf) proteins-the largest class of human TFs-GHT-SELEX shows that modular, alternative engagement of C2H2-zf domains is the norm, enabling several types of distinct target sites, and frequently involving internal duplication and divergence within the C2H2-zf array. Altogether, it is common for TFs to delineate a large fraction of in vivo genomic binding sites independently of other cellular factors.

Humans↗

High-throughput recovery of integron cassettes for gene discovery screens.

Integrons capture functional genes in mobile genetic elements called integron cassettes, which represent an untapped source of genes of biotechnological interest. Here we present two tools, cassette gatherer and cassette hunter, that enable high-throughput establishment of gene libraries either from genetically tractable strains or directly from DNA. We re-engineered a class 1 integron into counterselection markers on a plasmid or on the chromosome of a naturally competent Vibrio cholerae, which enabled capture of single cassettes in a sequence- and function-independent manner. When applied to Vibrio strains and genomic libraries, our tools recovered hundreds of single cassettes per assay with more than 99% specificity. We further subjected the library of cassettes generated by the hunter and gatherer tools to screens against phages ICP2 and T4, and identified nine phage-defence systems, including five previously undescribed. These tools enable rapid and large-scale recovery of integron cassettes that could be leveraged for functional gene discovery.

Journal Article↗

Herd-level heterogeneity of antimicrobial resistance in commensal Escherichia coli: A nationwide high-throughput survey of Australian pig herds.

Antimicrobial resistance in commensal Escherichia coli provides a useful indicator for overall antimicrobial resistance burden. We applied this approach to assess antimicrobial resistance within and between commercial pig herds across Australia. A high-throughput robotic workflow was used to isolate 2730 E. coli colonies from rectal contents collected in 2022 from healthy slaughter pigs (n = 300) representing 30 herds (∼70% of national production). Up to 94 isolates per herd underwent antimicrobial susceptibility testing using the Robotic Antimicrobial Susceptibility Platform. Isolate- and herd-level antimicrobial resistance indices were calculated, weighting antimicrobials by their human health importance. Resistance to first-line agents was widespread: ampicillin 77% and tetracycline 79%. By contrast, resistance to critically important antimicrobials was rare (ciprofloxacin 0.11%; extended-spectrum cephalosporins 0.04%), and no clinical resistance to carbapenems or colistin was detected. Overall, 56.9% of isolates were multi-class resistant. Herd-level antimicrobial resistance within indices ranged from 1.51 to 5.76, revealing substantial between-herd heterogeneity. Three herds carried critically important antimicrobials-resistant isolates that would likely have been missed using conventional, lower-density sampling approaches. Whole-genome sequencing identified fluoroquinolone-resistant isolates belonging to ST10 and ST69 (both qnrS1), and ST744 (Quinolone Resistance Determining Region mutations plus blaCTX-M-27). By testing approximately tenfold more isolates than conventional surveys, we uncovered considerable antimicrobial resistance with heterogeneity within and between animals and herds, including farm-specific variability. This expanded sampling also enabled detection of critically important antimicrobial resistance at very low prevalence. In conclusion, high-throughput, high-density testing offers a practical early-warning system and herd-level benchmark to inform surveillance and targeted interventions.

Animals↗

PyEvoMotion: a Python tool for population-based time-course analysis of genome evolution.

SUMMARY: We present PyEvoMotion, an open-source Python tool for inferring molecular clock models with time-dependent Gaussian noise from high-throughput genomic datasets. PyEvoMotion features a command-line interface and a modular architecture, allowing seamless integration into larger bioinformatic pipelines. The tool supports customizable filtering, temporal discretization definition, and mutation classification, making it adaptable to diverse research needs. While traditional phylogenetic methods may encounter computational challenges with large datasets, PyEvoMotion can process thousands to millions of sequences to compute statistical parameters associated with a stochastic differential equation model, thereby weighting the genetic variation within the population. Using viral genomic data, we demonstrate its capability to infer evolutionary rates and detect non-Brownian evolutionary motions with subdiffusive behavior. PyEvoMotion shows potential to provide overlooked insights into genome evolution in different contexts. AVAILABILITY AND IMPLEMENTATION: The open source software is available on GitHub at https://github.com/luksgrin/PyEvoMotion and on SourceForge at https://sourceforge.net/projects/pyevomotion.

Software↗

An Amplicon Panel for High-Throughput and Low-Cost Genotyping of Yesso Scallop Mizuhopecten yessoensis.

The Yesso scallop Mizuhopecten yessoensis was imported from Japan to western Canada in the late 1980s to establish an economically viable scallop aquaculture industry. Since this time, the industry in Canada has operated with existing genetic diversity within the broodstock, which is considerably limited relative to wild populations. The sector has not been able to realise its full potential in part due to idiopathic hatchery failures and farm stock collapses due to disease outbreaks associated with the intracellular bacterial pathogen Francisella halioticida. To support Yesso scallop production and breeding, here we generate a low-density, genotyping-by-sequencing amplicon panel using single nucleotide polymorphism (SNP) markers that are evenly spaced across the M. yessoensis genome and that show high heterozygosity in Canada and Japan. The panel can also exploit the high genetic polymorphism of the M. yessoensis genome, with de novo SNP calling identifying over 2,500 high quality SNPs within the 579 sequenced amplicons. We demonstrate the utility and versatility of this new genotyping tool for breeding applications including parentage assignment, low density family-based genome-wide association study, trait heritability evaluation to determine potential for genomic selection, and species differentiation (against the weathervane scallop Patinopecten caurinus). We did not find any genomic regions significantly associated with F. halioticida resistance but did identify potential for genomic selection. We could separate the two species based on genotypes, and did not see evidence of a past M. yessoensis x P. caurinus hybridization event within the M. yessoensis breeding population at Vancouver Island University. This low-cost genotyping panel is expected to accelerate selective breeding improvements for M. yessoensis in Canada and elsewhere.

Animals↗

Mapping ovarian cellular and molecular landscape across the lifespan of women: a scoping review.

BACKGROUND: With growing interest in ART, fertility preservation, and postmenopausal health of women, reproductive medicine is increasingly focused on characterizing oocytes and ovarian tissue composition, as well as understanding the molecular mechanisms that guide ovarian function throughout its lifecycle. High-throughput omics technologies have enabled the characterization of different molecular layers, leading to substantial advances in our understanding of their complex dynamics. However, not all molecular aspects are studied equally, and studies examining the same modalities often show inconsistencies, underscoring the need for data standardization and highlighting the potential for using transformative artificial intelligence and machine-learning (AI/ML) methods for ovary studies. OBJECTIVE AND RATIONALE: This study aims to evaluate how multi-omic studies have advanced our understanding of the ovarian lifecycle from fetal development to postmenopause. We systematically reviewed published studies that have investigated molecular/omic layers, including the genome, methylome, transcriptome, and proteome throughout ovarian development and aging. Our analysis identified key molecular and cellular patterns, highlighted inconsistencies across studies and addressed gaps in data analysis, interpretation, and reproducibility to guide future research. SEARCH METHODS: We conducted a systematic literature search of Medline (PubMed), Embase (Ovid), and Web of Science Core Collection (Clarivate) using a combination of controlled and free text terms for human ovary, oogenesis, folliculogenesis, ovary development and (epi)genome, transcriptome, proteome, and multi-omic mechanisms to find relevant articles published before August 2025. To focus the scope of the current review, studies of domesticated and farm animals, rodents and other model organisms, non-human primates, as well as those examining various human ovarian pathologies were excluded. OUTCOMES: The search identified 23 546 studies for screening, of which 637 full-text studies were assessed for eligibility. Subsequently, we extracted data from 121 studies. Most studies analyzed the transcriptome of oocytes, granulosa cells, and ovarian tissue from reproductive-age individuals (n = 91), with fewer studies examining samples from individuals of advanced reproductive age (n = 45) and fetal (n = 16) samples. Transcriptome analyses were most common (n = 103, 85%), followed by proteome (n = 19, 16%) and epigenome (n = 14, 12%) studies. We found substantial variation in how studies defined and reported participants' groups as well as in their sequencing technologies and data analysis methods, with a lack of standardized reporting of background clinical information, data analysis methods, and pipeline details. The key findings underscore the prevailing consensus on genes defining major ovarian cell types and their roles throughout the ovarian lifespan, from prenatal development to postmenopausal transformation. This review highlighted the underrepresentation of certain patient groups, particularly prepubertal and peri-/postmenopausal individuals, among researched populations, due to obvious clinical and ethical reasons. WIDER IMPLICATIONS: This scoping review offers a comprehensive overview and benchmark of the current state of high-throughput omics-based research on ovarian cellular composition and molecular dynamics. To address these shortcomings, we propose general recommendations for multi-omics ovary studies and emphasize the necessity for more thorough multi-omic data integration by effectively applying novel AI/ML approaches. They can potentially improve the quality of multi-omics analyses at both single-cell and tissue levels despite limited sample sizes and enable integration of molecular profiling data with clinical and radiology datasets, enabling a more comprehensive understanding of ovarian biology. Such advancements can enhance reproducibility of research findings and guide future research to deepen our understanding of ovarian biology and ultimately support the development of medical technologies for better preserving fertility and alleviating infertility. REGISTRATION NUMBER: A protocol was published a priori on the Open Science Framework (https://osf.io/z38gb/).

Female↗

Exome sequencing and large-scale analysis of electronic medical record-linked biobank data identify candidate deafness genes.

INTRODUCTION: Rapid advances in whole-exome sequencing (WES) have enabled large-scale detection of pathogenic variants. Although hundreds of genes are implicated in hearing loss, up to half of inherited cases remain unsolved, limiting eligibility for gene therapy trials that require genetic diagnosis. Biobanks and electronic medical records (EMRs) offer opportunities to integrate genomic and clinical data at scale and expand the spectrum of hearing loss genes. Despite clinical value, EMRs often lack key information such as inheritance patterns, posing challenges for accurate interpretation. METHODS: WES was performed on DNA samples from 1038 hearing-impaired patients enrolled in the Maccabi Research and Innovation Center Tipa Biobank. Clinical data were extracted from EMRs. Audiograms were available for all cases, although data on age of onset, family history and mode of inheritance were mostly unavailable. We applied a scalable bioinformatics analysis strategy for high-throughput annotation, filtering and prioritisation of WES variants across more than 1000 patients, designed to accommodate incomplete and heterogeneous clinical records. RESULTS: Using this approach, 15% of cases were solved or potentially solved through known or novel variants in established deafness genes. Homozygous variants in novel candidate genes were identified in 3% of cases. Functional characterisation was performed for promising candidate genes to validate their role in the ear. CONCLUSION: These findings demonstrate that WES can determine disease aetiology in large, genetically heterogeneous populations, even in the context of incomplete clinical data. This approach supports large-scale genetic screening and provides a framework for identifying patients who may benefit from emerging gene-based therapies.

Genetic Testing↗

A streamlined protocol for small-scale protoplast generation and CRISPR/Cpf1-mediated genome editing in Fusarium oxysporum.

Fusarium oxysporum is a significant threat to agriculture and One Health, requiring advanced molecular tools for functional genomic analyses and biological control agent development. Existing gene-editing methods are hampered by costly protoplast preparation protocols and by CRISPR-Cas9 limitations, such as restricted protospacer adjacent motif (PAM) sequences and complex guide RNA requirements. We engineered an efficient CRISPR/Cpf1 system that overcomes these issues through three main innovations: small-scale protoplast generation using filter column-based methods that greatly reduce enzyme consumption while simplifying workflows, a CRISPR/Cpf1 system with shorter guide RNA design and staggered DNA cleavage to promote homologous recombination, and minimal homology arm strategies that significantly decrease cloning complexity. Extensive validation confirms successful gene targeting with molecular verification and functional analysis via standardized pathogenicity assays. This integrated platform offers affordable, accessible tools for systematic F. oxysporum research, enhancing fundamental understanding of plant-pathogen interactions and supporting high-throughput screening vital for agricultural biotechnology and biological agent development.

CRISPR/Cpf1↗

Enhanced identification of key bacterial motility genes via a cross-species genomic hybrid feature machine learning approach.

Efficient and accurate identification of functional genes is critical to biological research, yet traditional single-species approaches are often limited by low efficiency. Previously, we established a novel method for identifying key genes using cross-species protein domain features and machine learning. However, the high multiplicity of gene members associated with specific domains creates a substantial workload for subsequent experimental validation. To address this, this study proposes an enhanced approach that integrates EggNOG-based protein sequence annotation with domain analysis. Unannotated sequences are subsequently analyzed for protein domains, generating a comprehensive "direct gene annotation plus domain" hybrid feature matrix. While the hybrid matrix model yielded comparable predictive accuracy, it significantly enhanced feature resolution: the top 50 predicted features were all known motility-related genes or domains. Furthermore, among the top 100 ranked features, 58 are confirmed to be directly related to motility based on experimental evidence. Although strict genus-level control still yielded 51 confirmed features, excessive taxonomic restriction drastically reduces the number of training genomes, which may paradoxically impair identification efficiency. These results demonstrate that the new method effectively reduces the subsequent experimental workload and enables high-throughput identification of functional genes in a single analysis. With accuracy and efficiency far exceeding those of existing single-species identification methods, it provides a highly efficient solution for mining key genes underlying other complex bacterial phenotypes.

Machine Learning↗

2-Mercaptoethanol/DMSO Workflow Enables Highly Reproducible Quantitative Proteomics.

Proteomics provides a systematic and high-throughput approach to comprehensively characterize protein networks, enabling insights into cellular functions and disease mechanisms. Carbamidomethylation using iodoacetamide (IAA), a common method for cysteine alkylation, is known to cause nonspecific modifications that increase spectral complexity in mass spectrometry and reduce quantitative accuracy. Here, we established a reproducibility-focused 2-mercaptoethanol (2-ME)/dimethyl sulfoxide (DMSO) workflow and systematically evaluated its quantitative performance at the proteome-wide level. Mouse liver proteomes were processed using either 2-ME/DMSO or conventional IAA treatment, followed by liquid chromatography-tandem mass spectrometry (LC-MS/MS) analysis. The optimized 2-ME treatment increased the number of cysteine-modified peptides by 1.6- to 1.9-fold. Although total protein identifications were comparable, 77% of proteins exhibited improved sequence coverage with the optimized 2-ME treatment. Quantitative reproducibility was also enhanced, with the peptide quantified CV ≤ 20% increasing from 61.4% with IAA treatment to 86.1% with 2-ME treatment, and protein quantified CV ≤ 20% increasing from 80.6% with IAA treatment to 93.5% with 2-ME treatment. Application of this new workflow to ovarian clear cell carcinoma reliably detected cisplatin-induced alterations. The 2-ME/DMSO workflow offers a simple and highly reproducible proteomics strategy for accurate quantitative proteomics.

Animals↗

Identification and characterization of ectopic chromosomal amplifications in acute myeloid leukemia cell limes using high-throughput chromosome conformation capture screening.

Despite advanced molecular diagnostics, improving outcomes for refractory acute myeloid leukemia (AML) remains challenging. Although many cancer-related genes are identified, their molecular mechanisms are not fully elucidated. Amplification is a mechanism of cancer-associated gene activation, and ectopic gene amplification may have particularly high pathological significance. However, research on ectopically amplified cancer-associated genes in leukemia remains limited. Here, we evaluated the usefulness of high-throughput chromosomal conformation capture (Hi-C) as a screening method for ectopic gene amplification and assessed whether ectopic amplification of cancer-associated genes may represent a general phenomenon in AML. We screened the U-937 and NB-4 cell lines using in situ Hi-C. Regions appearing as "high-intensity bands" in Hi-C contact maps were identified and validated using fluorescence in situ hybridization (FISH). Additionally, copy number variation analysis was performed using whole-genome sequencing (WGS) to extract cancer-associated genes with ectopic amplification. In the U-937, three genomic regions showing "high-intensity bands" were identified and confirmed as ectopic amplifications-including PDCD1LG2 (PD-L2), CD274 (PD-L1), and JAK2; that is, four copies were detected by WGS, and amplification signals were observed by FISH. In the NB-4, four such regions were detected, including MYC and KRAS, with expression level of 498 transcripts per million (TPM) and 34 TPM, respectively. Copy number variation analysis further identified multiple cancer-associated genes with ectopic amplification. Overall, these findings demonstrate the presence of ectopic amplification of cancer-associated genes in AML cell lines and support the usefulness of Hi-C as a screening method for detecting such genomic alterations.

Acute myeloid leukemia↗

Advances in the diagnosis and classification of B-ALL: comparative insights from updated guidelines.

Accurate molecular classification is essential for diagnosis, risk stratification, and treatment selection in B-cell lymphoblastic leukemia (B-ALL). In this study, we performed a comprehensive, real-world reclassification of 1015 consecutively diagnosed B-ALL patients using the fifth edition of the World Health Organization Classification of Haematolymphoid Tumours (WHO-HAEM5) and the International Consensus Classification (ICC). An integrative genomic strategy that combined whole transcriptome sequencing, fusion detection, mutational analysis, and cytogenetics enabled reclassification according to both the WHO-HAEM5 and ICC frameworks, thereby substantially reducing the proportion of unclassifiable B-ALL from 41.9% (2016 WHO revision [WHO-HAEM4R]) to 15.9% (WHO-HAEM5) and 11.9% (ICC). Distinct clinical and prognostic features were identified across newly defined subtypes. Multivariable analysis confirmed that this genomic classification is a robust, independent predictor of survival after adjusting for age, minimal residual disease status, and transplant intervention. Specifically, HLF-rearranged and MEF2D-rearranged B-ALL conferred a persistently poor prognosis across all age groups despite allogeneic hematopoietic stem cell transplantation, highlighting an urgent need for novel therapeutic strategies. Gene expression profiling resolved cryptic subtypes, including ETV6::RUNX1-like, ZNF384-rearranged-like, and BCR::ABL1-like B-ALL, and uncovered diagnostic ambiguity in patients with concurrent lesions. In addition, we report emerging high-risk groups, including IDH1/2- and ZEB2 Q1072-mutated B-ALL, that may warrant recognition as distinct molecular entities. Our findings demonstrate the clinical use of integrative transcriptomic profiling in refining B-ALL taxonomy in guiding risk-adapted therapies and informing future revisions of diagnostic standards. This study supports the incorporation of high-throughput molecular diagnostics into routine leukemia classification and precision treatment planning.

Humans↗

[Applications and Challenges of Deep Learning in Human Genome Research].

In recent years, the advent of high-throughput omics technologies has fueled an explosive growth in human genomic data. Uncovering the latent functions within this vast data has become a significant challenge in functional genomics research. While traditional statistical methods have proved successful for analyzing smaller-scale datasets in the past, they exhibit clear limitations in analytical efficiency and integrating multi-dimensional data, struggling to meet the escalating demands of contemporary genomic analysis. The introduction of deep learning (DL) technologies offers a novel paradigm for this field. This review systematically examines the advances in applying deep learning to human genomics research. Studies demonstrate that when ample labeled data is available, discriminative DL computational methods-such as Convolutional Neural Networks (CNNs) and Long Short-Term Memory networks (LSTMs)-achieve high accuracy and efficiency in genomic variant discovery tasks. Furthermore, generative DL methods, particularly Large Language Models (LLMs) leveraging self-supervised pre-training strategies, effectively integrate complex genomic information and exhibit superior performance in functional genomic sequence annotation and gene regulation studies. This review also explores the application of LLMs in multi-omics data integration and prediction. Looking ahead, the continued accumulation of long-read sequencing and high-dimensional data is expected to enable DL technologies to integrate increasingly complex and heterogeneous genomic information, playing an increasingly crucial role in human genomics research.

Deep Learning↗

Development of a PCR-based technique for genotyping UGT1A1 gene and distribution of rs3064744 alleles in the Russian population.

BACKGROUND: Accurate determination of tandem thymine-adenine (TA) repeat numbers in the UGT1A1 promoter region (rs3064744) is essential for diagnosing Gilbert's syndrome and personalizing therapy with toxic agents like irinotecan and atazanavir. However, traditional polymerase chain reaction (PCR) assays face severe limitations due to the AT-rich sequence and overlapping melting temperatures (Tm) of the highly homologous 7TA and 8TA alleles. In this context, melting curve analysis (MCA) employing fluorophore-quencher systems has emerged as a promising alternative. The purpose of this study was to develop a novel genotyping approach combining optimized aPCR-MCA analysis with an automated classifier to overcome the limitations posed by the differentiation of highly homologous alleles and to demonstrate its practical application, providing the distribution of rs3064744 genotypes across four regional cohorts of the Russian population. METHODS: A specialized Dual Head 1D-convolutional neural network (1D-CNN) ensemble with Test-Time Augmentation (TTA) was developed. The model was trained and internally validated on 1,620 engineered plasmid samples, and independently evaluated on an external clinical test set of 440 unique patient genomic DNA specimens. Real-time PCR was performed on CFX96 and DTprime platforms. Additionally, population-wide screening was conducted on 997 archival clinical samples from Moscow, Sakha (Yakutia), Dagestan, and Rostov regions. RESULTS: While 5TA and 6TA alleles were easily separated, absolute Tm distributions of 7TA and 8TA alleles overlapped significantly, and non-uniform Tm shifts of 0.8 °C-1.4 °C occurred across platforms. Conventional absolute Tm thresholding was therefore inadequate. By assessing relative morphological curve divergence against co-amplified 7TA/7TA and 7TA/8TA reference anchors, the 1D-CNN ensemble neutralized instrument noise. It achieved 100% accuracy on internal validation and 100% concordance (440/440) with clinical reference pyrosequencing. Population screening revealed that Dagestan, Yakutia, and Rostov cohorts closely align with the European population. Rare 5TA and 8TA alleles were detected at low frequencies in Yakutia and Moscow. CONCLUSION: Combining LNA-modified aPCR-MCA with a comparative 1D-CNN model successfully circumvents thermodynamic limitations and eliminates human operator bias. This integrated system offers an accessible, high-throughput, and clinically valid solution for routine UGT1A1 pharmacogenetic testing.

1D-CNN↗