PubMed HealthSearch

SEARCH · PubMed Health

Results for “High-throughput sequencing data”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Bridging the gap between legacy polymerase chain reaction-based microsatellite data with high-throughput sequencing data for conservation genomics.

Microsatellites are powerful markers for tracking genetic variation in wildlife populations due to their high polymorphism and genome-wide abundance. While polymerase chain reaction (PCR)-based fragment size analysis has been the standard for genotyping microsatellites, high-throughput sequencing offers greater resolution and the opportunity to sync historical datasets with modern analyses. We evaluated how genotypes from whole-genome sequencing align with PCR data for 15 microsatellite loci in 11 North American brown bears (Ursus arctos). Brown bear populations in the 48 contiguous United States have declined from approximately 50,000 to fewer than 2,000 over the past decades. Their endangered status has prompted extensive research and genetic monitoring, yielding large, multiyear microsatellite datasets upon which future conservation efforts can build. We achieved an overall microsatellite genotype concordance rate of 94.5% comparing high-throughput sequencing results to PCR based-fragment size results. All discrepancies occurred at complex loci containing multiple insertions and/or deletions (indels). Physically linked indels or single nucleotide polymorphisms (SNPs) occurring within the loci were misinterpreted as independent insertions, underscoring the need for genotyping tools that incorporate phasing when genotyping. To evaluate coverage effects, we downsampled high-throughput sequence data from 30x to 2x. Concordance remained high at 20 to 30x but dropped sharply at 10x, with 5x and 2x having discordant genotypes or insufficient coverage for genotyping. Accurate genotyping required both sufficient depth and number of reads spanning the entire repeat regions. Our results show that short-read whole-genome sequencing can recover microsatellite genotypes with high accuracy when paired with careful variant interpretation. By aligning historical PCR datasets with modern sequencing data, we can preserve decades of genetic insight and strengthen long-term monitoring of at-risk populations.

Animals

Genetic determinants of gestational diabetes mellitus in thai pregnant women: role of GCKR, CDKAL1, TCF7L2, NEDD1, and CMIP variants.

BACKGROUND: Gestational diabetes mellitus (GDM) has a high global prevalence and arises from complex interactions between genetic predisposition and environmental factors. GDM is associated with metabolic disturbances and chronic low-grade inflammation, both of which contribute to its pathogenesis. This study aimed to investigate the association between GDM and 135 single-nucleotide polymorphisms (SNPs) across 20 genes related to metabolic traits. METHODS: In this case-control study, 152 pregnant women with GDM and 684 pregnant women with normal glucose tolerance (NGT) who underwent antenatal examination at Siriraj Hospital, Bangkok, were enrolled. Clinical data and blood samples were collected from all participants. Genomic DNA was isolated and subjected to whole-genome sequencing using the DNBSEQ-T7RS high-throughput sequencing platform. Genotype analyses were performed using R software, and haplotype analyses were conducted using the online SNPStats software. RESULTS: After adjusting for maternal age and pre-pregnancy body mass index, polymorphisms in TCF7L2 (rs34872471, rs7901695, rs4506565, rs7903146, rs12243326, and rs12255372), NEDD1 (rs10431408, rs11830756, rs249579, rs249585, and rs4762339), CMIP (rs2306115 and rs201681534), CDKAL1 (rs4710942), GCKR (rs2293572 and rs2293571), and GCK (rs5883890) were significantly associated with the risk of GDM. Haplotype analysis demonstrated that the TCF7L2 rs12243326-rs12255372 CA haplotype was associated with a decreased risk of GDM (OR = 0.44, 95% CI: 0.23-0.81), while the NEDD1 rs249579-rs249585-rs4762339 GGT haplotype was associated with an increased risk of GDM (OR = 1.40, 95% CI: 1.08-1.82). CONCLUSIONS: These findings suggest that genetic variations in TCF7L2, NEDD1, CMIP, CDKAL1, GCK, and GCKR contribute to GDM susceptibility in the Thai population.

Humans

Host-independent metagenomics reveal gut bacteria contribution to Delia antiqua growth by vitamin B6 provision.

Insect guts host a diverse and abundant array of microorganisms. These microbes improve host fitness by extensively involving in a range of crucial physiological processes, which have mainly been revealed by high-throughput sequencing, particularly metagenomics. However, it is almost impossible to make an accurate and complete distinction between the genetic functions of microbial symbionts and insect hosts without host genome data. By comparing metagenomic data from gut germ-free and nonaxenic larvae, we accurately identified the data belonging to the gut microbiome of the onion maggot Delia antiqua (Diptera: Anthomyiidae). Besides, a correlation between bacteria of the genus Wohlfahrtiimonas (Gammaproteobacteria: Pseudomonadaceae) and vitamin B6 metabolism was detected through collinearity analysis. Furthermore, in vitro tests confirmed that the gut bacterium Wohlfahrtiimonas larvae contributed to the growth of D. antiqua larvae via the independent synthesis of vitamin B6. This study provides a comprehensive view of the gut bacterial diversity in D. antiqua and reveals a functional profile that is strictly specific to the gut microbiota of this species. It has preliminarily revealed the functional differentiation between insect hosts and their symbiotic microorganisms. This study also offers a technical reference for the study of microbial symbiotic functions in other insect-microbe symbioses without host genomic data.

Animals

Long-read low-pass sequencing enhances variant detection in a peanut MAGIC population.

Accurate genotyping accelerates crop improvement, yet long-read sequencing remains underused in breeding due to cost. We present a scalable long-read low-pass (LRLP) sequencing framework for high-throughput variant discovery and trait mapping. Using PacBio HiFi reads in an allotetraploid peanut (Arachis hypogaea; AABB, 2n = 4x = 40) MAGIC population, we generated both LRLP and short-read low-pass (SRLP) data. At comparable depths, LRLP achieved substantially greater whole-genome and gene-space coverage than SRLP. Data were analyzed using both a single-reference genome and an 18-parent pangenome graph constructed with KhufuPan, a new tool for graph-based genotyping. Across analytical approaches, LRLP consistently identified more SNPs, indels (2-1,000 bp), and structural variants (>1 kb) than SRLP, improving genotype resolution and selection accuracy, particularly for large structural variants. By reducing cost barriers and increasing variant discovery in complex genomes, LRLP provides a practical path for deploying advanced genomics in under-resourced and orphan crops critical to global food security.

Arachis

Clinical and genetic features of Ph-negative myeloproliferative neoplasms with dual-driver gene positivity.

OBJECTIVES: To investigate the clinical laboratory characteristics and gene mutation features of dual-driver gene positivity in patients with Philadelphia chromosome-negative myeloproliferative neoplasm (Ph-negative MPN). METHODS: We conducted a retrospective analysis of clinical data and genetic test results from 203 newly diagnosed patients with Ph-negative MPN. Of these, 194 had single-driver gene positivity and 9 had dual-driver gene positivity. High-throughput sequencing was used to detect mutations in JAK2, CALR, and MPL. Clinical characteristics and gene mutation profiles were compared between the two patient groups. RESULTS: The incidence of dual-driver gene positivity was 4.4% (9/203), with the most common combinations being JAK2 with CALR (4 patients) and JAK2 with MPL (4 patients). Compared with the single-driver group, the dual-driver group had a significantly higher risk of bleeding [4.1% (8/194) vs. 33.3% (3/9), P = 0.008] and a higher proportion of uncommon mutations [3.6% (7/194) vs. 33.3% (3/9), P = 0.006]. No statistically significant differences were observed between the two groups regarding age, thrombosis incidence, splenomegaly, or routine blood test indicators. During follow-up, 1 patient in the dual-driver group died from cerebrovascular disease. No leukaemia transformation or disease-related deaths occurred among the remaining patients. DISCUSSION: The increased bleeding risk in dual-driver patients may be related to a higher proportion of CALR mutations, elevated platelet counts, and higher variant allele frequencies, though these findings require validation in larger cohorts due to the small sample size. The higher prevalence of uncommon mutations suggests a more complex mutational landscape in this subgroup. CONCLUSION: Patients with Ph-negative MPN and dual-driver gene positivity may have a higher risk of bleeding and a more complex gene mutation profile.

Humans

Bridging the airway microbiome and targeted therapy in bronchiectasis: multi-omics insights, endotypes and emerging therapies.

Bronchiectasis is a heterogeneous chronic airway disease primarily driven by persistent infection, microbial dysbiosis and dysregulated host immunity. While culture-based microbiology has historically informed clinical management, advances in high-throughput sequencing and multi-omic technologies have transformed our understanding of the airway ecosystem, revealing that disease activity is shaped not only by individual pathogens, but by complex and dynamic host-microbe interactions. Despite the breadth of descriptive microbiome data, translation into clinically actionable diagnostics or therapies has been limited. Importantly, cross-sectional correlations between microbiota and inflammation do not establish cause and effect, underscoring the need to embed host-microbiome profiling within both longitudinal and interventional therapeutic trials. In this review, we critically appraise current microbial and host multi-omics research in bronchiectasis, integrating microbiome studies with host inflammatory, proteomic and immunophenotyping data. We highlight themes emerging across cohorts, including low microbial diversity, pathogen dominance, loss of commensal networks and neutrophil-driven inflammation, and discuss how these features align with biological endotypes associated with exacerbations and treatment response. Drawing on lessons from host-directed therapeutic successes, we examine translational roadblocks limiting microbiome-guided care. We further review emerging microbiome-modulating strategies such as pathogen-specific biologics, bacteriophage therapy, live biotherapeutic products, biofilm-targeting adjuncts and precision antibiotic stewardship. Finally, we propose a roadmap toward microbiome-informed precision medicine through harmonised methodologies, integration of host and microbial biomarkers into clinical trials, and embedding multi-omics pipelines within large international registries. Collectively, these advances have the potential to shift bronchiectasis research and clinical management towards rationally designed, precision medicine-driven therapeutic strategies.

Humans

ChIP-seq profiling identifies diapause-regulated H3K27me3 targets in the fat body of Culex pipiens.

Culex pipiens, a principal vector of significant arboviruses, survives winter through diapause, a hormonally controlled inactive phase that enhances endurance under severe cold circumstances. Recent data suggests that epigenetic processes, namely histone post-translational modifications (hPTMs), play a crucial role in regulating seasonal dormancy. Prior studies from our laboratory indicated a decrease in the methylation of Histone 3 (H3K27me3) in diapausing fat body tissue, associated with elevated expression of the histone demethylase UTX. Nonetheless, the precise genomic areas impacted by these chromatin alterations remained unidentified. We used chromatin immunoprecipitation coupled with high-throughput sequencing (ChIP-seq) to delineate the genome-wide distribution of H3K27me3 across fat body chromatin in diapausing (D) and non-diapausing (ND) female Cx. pipiens. Notably, the higher signal at transcription start sites (TSSs) reflects localized redistribution rather than a global decrease, as diapausing fat bodies retain less H3K27me3 overall but concentrate it at promoters. To investigate the functional significance of these chromatin alterations, we confirmed a number of target loci via ChIP-qPCR and assessed gene expression with qRT-PCR. We identified many critical genes that were markedly increased in diapausing mosquitoes, exhibiting an inverse relation to H3K27me3 enrichment. Our data demonstrates different H3K27me3 chromatin landscapes between diapausing and non-diapausing Cx. pipiens, corroborating a hypothesis of selective, locus-specific repression in the non-diapause state and its targeted removal during diapause to permit activation of dormancy-associated genes. These results suggest that chromatin remodeling is a core driver of the diapause switch.

Animals

Performance comparison of rapid and native barcoding methods for Oxford Nanopore sequencing of Poliovirus Viral Protein 1 (VP1) amplicons.

Accurate and timely sequencing of poliovirus is critical for global eradication efforts, particularly for molecular epidemiology based on the typing region of the genome, viral protein 1 (VP1). While Oxford Nanopore Technologies (ONT) sequencing has expanded capabilities for poliovirus surveillance, the relative performance of different ONT library preparation methods, including ligation-based (Native Barcoding) and transposase-based (Rapid Barcoding) approaches, has not been systematically evaluated. In this study, we compared rapid barcoding and native barcoding workflows for sequencing VP1 amplicons from 17 type 2 poliovirus-positive samples, each processed in triplicate. Native barcoding generated significantly more sequencing output, producing approximately 2.3-fold greater total read yield than rapid barcoding, and demonstrated higher run-to-run reproducibility (R2 = 0.979-0.998 vs. 0.847-0.929, respectively; p&#x202f;<&#x202f;0.001). In addition, native barcoding generated 80% of the total yield achieved by rapid barcoding within approximately 7&#x202f;h, whereas rapid barcoding required approximately 40&#x202f;h to reach the same output. Despite these differences, both methods produced identical VP1 consensus sequences across all samples, with comparable read quality (median per-base Q-scores of approximately Q17-Q18). Rapid barcoding provided substantial practical advantages, reducing hands-on library preparation time (55 vs. 200&#x202f;min) and per-sample cost ($12.82 vs. $16.54), while simplifying workflow and reducing technical complexity. These findings indicate that sequencing yield may not be a determinant of downstream analytical outcomes for poliovirus VP1 ONT sequencing. Rapid barcoding therefore represents a cost-effective and efficient approach for routine poliovirus surveillance, whereas native barcoding remains advantageous in applications requiring rapid data generation or maximal sequencing depth.

Poliovirus

Strategies for mosaic variant calling in brain disorders.

The human brain is a genomic mosaic, where postzygotic mutations arising from embryogenesis to senescence drive diverse neurodevelopmental and neurodegenerative diseases. Because of numerous sequencing artifacts at ultralow variant allele frequencies (VAFs), detecting these variants remains a significant analytical challenge. This review focuses on single-nucleotide variants and small indels, summarizing current strategies for aligning sampling methods, including bulk, laser capture microdissection, and single-cell genomics, with the expected clonal architecture of the brain. It emphasizes that mosaic detection sensitivity is fundamentally constrained by sequencing depth, since even the most advanced algorithms cannot identify variants not physically represented in the sequencing library. The review further recommends the selection of variant calling algorithms based on validated VAF detection performance, matching tools like MuTect2 and MosaicForecast to their optimal performance ranges. Furthermore, we discuss how multitissue sampling, as emphasized by the SMaHT project, addresses the matched-control dilemma and supports accurate variant classification via cross-tissue VAF gradients. Integrating these established pipelines with multiomics modalities, including transcriptomic and epigenetic data, could advance the field toward a functional understanding of how the somatic genome impacts human brain health and disease.

Humans

Exome sequencing and large-scale analysis of electronic medical record-linked biobank data identify candidate deafness genes.

INTRODUCTION: Rapid advances in whole-exome sequencing (WES) have enabled large-scale detection of pathogenic variants. Although hundreds of genes are implicated in hearing loss, up to half of inherited cases remain unsolved, limiting eligibility for gene therapy trials that require genetic diagnosis. Biobanks and electronic medical records (EMRs) offer opportunities to integrate genomic and clinical data at scale and expand the spectrum of hearing loss genes. Despite clinical value, EMRs often lack key information such as inheritance patterns, posing challenges for accurate interpretation. METHODS: WES was performed on DNA samples from 1038 hearing-impaired patients enrolled in the Maccabi Research and Innovation Center Tipa Biobank. Clinical data were extracted from EMRs. Audiograms were available for all cases, although data on age of onset, family history and mode of inheritance were mostly unavailable. We applied a scalable bioinformatics analysis strategy for high-throughput annotation, filtering and prioritisation of WES variants across more than 1000 patients, designed to accommodate incomplete and heterogeneous clinical records. RESULTS: Using this approach, 15% of cases were solved or potentially solved through known or novel variants in established deafness genes. Homozygous variants in novel candidate genes were identified in 3% of cases. Functional characterisation was performed for promising candidate genes to validate their role in the ear. CONCLUSION: These findings demonstrate that WES can determine disease aetiology in large, genetically heterogeneous populations, even in the context of incomplete clinical data. This approach supports large-scale genetic screening and provides a framework for identifying patients who may benefit from emerging gene-based therapies.

Genetic Testing

Comparison of paralog identification methods and their impact on species tree topologies in target capture phylogenomics within the Sindora clade (Detarioideae: Leguminosae).

Target capture is a common method of generating high throughput DNA sequencing data for phylogenetic reconstruction of species relationships, for which single copy genes are usually most informative. However, a pervasive problem with target capture is that putatively single copy genes may in fact be paralogs resulting from gene duplication, which are problematic for phylogenetic inference because their evolutionary history may differ from the divergence history of species. Here, we use as a case study a target enrichment dataset of 88 species of Detarioideae (Leguminosae) with a focus on the Sindora clade to examine approaches for handling paralogs, including the built-in paralog handling functions in HybPiper and CAPTUS, plus subsequent steps using Putative Paralog Detection and the tree-based Yang & Smith orthology inference approach. We compare the paralogs flagged using these methods and verify their performance with BLAST mapping against a reference genome sequence of Sindora glabra, and then subsequently compare the species tree topologies produced across these methods. Our comparisons of paralogs flagged across the Sindora clade show that the Putative Paralog Detection pipeline was the most accurate in identifying paralogs in terms of its similarity to the BLAST mapping, followed by the built-in paralog identification function of CAPTUS. However, the results we recovered for the Detarioideae subfamily suggest that the largest differences in species tree topology resulted from the use of paralog-filtered alignments (such as with the Putative Paralog Detection pipeline and the Yang & Smith orthology inference approaches) rather than just by removing the sequences of identified paralogous genes. This was the true for HybPiper-assembled datasets but was not seen in CAPTUS-assembled datasets. In all comparisons, the topological differences caused by different paralog handling methods tended to be confined to clades where processes such as hybridisation and introgression are prevalent. Our study provides a roadmap to establish the best approach to identify, eliminate or separate paralogs in the absence of a chromosomally contiguous reference genome for a study group, and highlights the importance of careful data inspection and processing in addition to understanding the extent of paralogy and paralog characteristics (e.g. sequence divergence between copies) for their study group.

Phylogeny

Unveiling the Molecular Secrets of Seaweeds: A Comprehensive Review of Bioinformatics Applications in Algal Research.

Recent advances in high-throughput sequencing, bioinformatics, and multi-omics technologies have transformed seaweed research by overcoming long-standing challenges associated with complex genomes, diverse life cycles, and limited genomic resources. This review provides a comprehensive overview of bioinformatics approaches used to investigate seaweed genomics, transcriptomics, proteomics, metabolomics, microbiomes, and functional genomics, with emphasis on the computational tools and databases that support these analyses. Applications of bioinformatics in phylogenetics, drug discovery, microbiome characterization, and the development of biofuels, nutraceuticals, pharmaceuticals, and sustainable agriculture are also discussed. Particular attention is given to emerging strategies involving multi-omics integration, genome editing, artificial intelligence, machine learning, and synthetic biology that are reshaping seaweed research. The review further examines current challenges, including incomplete genomic resources, data standardization, and the need for experimental validation of computational predictions. Collectively, these advances highlight the growing role of bioinformatics in enabling systems-level understanding of seaweed biology and accelerating their translation into sustainable biotechnological and marine bioeconomy applications.

macroalgal genomics

Mitochondrial genome characteristics and phylogenetic analysis of Ramaria longispora.

This study, for the first time, assembled and annotated the complete mitochondrial genome of R.&#xa0;longispora using high-throughput sequencing technology. The genome is a circular molecule with a total length of 157,712&#x2009;bp and a GC content of 31.55%. It encodes 71 genes, including 15 core protein-coding genes (PCGs), 25 transfer RNA (tRNA) genes, 2 ribosomal RNA (rRNA) genes, 5 free-stranding open reading frames (ORFs), and 24 intronic ORFs. Among these, most free-stranding ORFs have unknown functions but include a DNA polymerase gene, while the intronic ORFs primarily encode LAGLIDADG and GIY-YIG endonucleases. The mitochondrial genome contains 39 introns. Phylogenetic analyses based on 15 core PCGs using Bayesian inference (BI) and maximum likelihood (ML) methods revealed that this R. longispora is most closely related to Ramaria flavescens and Ramaria ichnusensis. This study provides foundational data for mitochondrial genome research in the Ramaria genus and offers important references for taxonomic and evolutionary studies of this group.

Mitochondrial genome

Expanded detection of canine enteric viruses in UK dogs with diarrhoea.

Canine enteric viruses are an important cause of gastrointestinal disease in pet dogs worldwide. Routine diagnosis often relies on pathogen-specific PCR assays, which may fail to detect some viruses, particularly neglected pathogens or genetically divergent variants of established threats. This limits both clinical characterization of affected patients and broader understanding of disease ecology. To address these limitations, we applied metagenomics and a viral discovery bioinformatics pipeline to faecal samples from diarrhoeic dogs in the UK that had been submitted routinely for PCR-based diagnostic testing. Across 80 dogs, we identified 12 viruses known to infect canids, 9 of which have not previously been reported in UK dogs. Among these, several taxa with prior associations to gastrointestinal disease were identified, including canine sapovirus and canine minute virus. By contrast, for other viruses newly detected in the UK, including bufavirus and rotavirus C, clinical relevance in dogs remains unclear. Notably, an identified protoparvovirus fell within the same species as human-canine-associated parvovirus 1, a recently described lineage detected in both canine and human oropharyngeal samples. We also identified a canine parvovirus 2 strain that clustered with a predominantly wildlife-associated lineage, consistent with occasional exposure at the domestic-wildlife interface rather than established circulation in dogs. These two detections illustrate how genome-level surveillance can help prioritize viruses for targeted investigation of host range and transmission context. Overall, these data broaden the catalogue of viruses associated with diarrhoeic dogs in the UK and support periodic review of diagnostic targets informed by viral metagenomic surveillance, while highlighting the need for controlled studies to assess causality and clinical relevance.

Animals

Comprehensive Viral Detection and Profiling of Plasma Cell-Free RNA in Patients With Suspected Hemophagocytic Lymphohistiocytosis.

Hemophagocytic lymphohistiocytosis (HLH) is a severe, rapidly progressive disease. While viral infection is considered a common etiology of pediatric HLH, specific causative viruses other than the Epstein-Barr virus (EBV) have been rarely identified. This study utilized metagenomic next-generation sequencing (NGS) to identify potential causative pathogens in plasma samples from 17 pediatric patients with suspected HLH. Additionally, one case each of confirmed EBV- and cytomegalovirus (CMV)-associated HLH was analyzed for methodological validation. Plasma cell-free RNA (cfRNA) profiling was performed using NGS data to assess the host transcriptome response. Significant viral reads of human herpesvirus-6B, human herpesvirus-7, and Hubei reo-like virus (HRLV) 14 were detected using metagenomic NGS in one patient each. Plasma cfRNA profiles from five patients with viral infection (including EBV and CMV) were compared to those of 14 patients without viral infection. By comparing the two patient groups, 1053 differentially expressed genes were identified. The gene ontology (GO) term of "adaptive immune response" (GO: 0002250) was significantly enriched among upregulated genes in the virus-positive group. Furthermore, an isolated cluster consisting specifically of mitochondrial RNAs, was identified in the upregulated genes of the virus-positive group. Using metagenomic NGS, several candidate viral pathogens were identified in patients with suspected infection-related HLH. The viral genome of HRLV 14, previously undetected in human clinical samples, was identified in one patient. The results from plasma cfRNA profiling suggest that mitochondrial RNAs may reflect the underlying pathogenesis of virus-associated HLH and have potential utility as disease biomarkers.

Humans

Targeted sequencing reveals a distinct genetic alteration landscape in oral multiple primary squamous cell carcinomas.

OBJECTIVE: Oral multiple primary cancers (MPCs) are associated with poor clinical outcomes, yet their genomic characteristics remain insufficiently understood. DESIGN: Fifty-four formalin-fixed paraffin-embedded (FFPE) tumor samples from 30 patients with oral MPCs were analyzed using high-depth targeted sequencing of a customized 14-gene panel derived from prior whole-exome sequencing data. Detected alterations were analyzed after removal of synonymous mutations. RESULTS: Non-silent genomic alterations were identified in 59.3% (32/54) of samples, involving 19 patients. A total of 70 variant loci across 13 genes were detected. AKAP13 was the most frequently mutated gene at both the sample (22.2%, 12/54), with recurrent mutations observed across multiple patients. In contrast, TP53 mutations occurred at a substantially lower frequency (11.1%, 6/54). Marked inter- and intra-patient mutational heterogeneity was observed. CONCLUSIONS: FFPE-based targeted sequencing enabled an initial characterization of genomic alterations in oral MPCs. Recurrent alterations in AKAP13, GLI2, JMJD1C, and DNAH8, together with the relatively low frequency of TP53 alterations, identify candidate genomic features for further investigation and provide a basis for future studies of the molecular basis of oral MPCs.

Humans

Upscaling Genotyping by Amplicon Sequencing With GBAS-GUI.

Genotyping by amplicon sequencing (GBAS) is a relatively low-cost approach for generating genotypic data compared with established genomic methods, making it highly scalable and particularly suitable for large-scale genetic monitoring projects. However, most existing analytical pipelines are either marker-specific, insufficiently scalable, or lacking efficient data management systems for the long-term integration of genotypic information, limiting the full potential of GBAS. Here, we address this gap by introducing GBAS-GUI (https://github.com/sonnenbe-dot/GBAS-GUI), a pipeline capable of generating GBAS-based genotypic data for a wide variety of loci at scale. GBAS-GUI integrates a graphical user interface with multiple checkpoints to improve accessibility and robustness. It implements multiprocessing architecture and a relational database that links genotypic data with associated sample metadata to enhance scalability and data management. The pipeline further enables marker screening through automated calculation of polymorphism information content (PIC) and implements a strategy to recover homologous genotypic information from paralogous loci with non-overlapping amplicon length ranges. Using multiple empirical datasets, we demonstrate substantial improvements in processing speed, database management and handling artefacts related to co-amplification of unspecific regions and duplicates of the same genomic region. We further show that incorporating the full sequence information captured by an amplicon increases marker information content beyond what is achievable with length-based genotyping alone and expands the analytical versatility of GBAS. Overall, GBAS-GUI provides a robust, scalable and versatile framework that unlocks the potential of GBAS for large-scale population genetic and phylogeographic studies.

Genotyping Techniques

Lynch syndrome-associated urothelial carcinoma: clinical and molecular findings from a single-institution cohort.

Lynch syndrome-associated urothelial carcinoma (LS-UC) is a rare and undercharacterized clinical entity. While FGFR3 alterations are well described in sporadic urothelial carcinoma, their prevalence and clinical implications in LS-UC remain unclear. We aimed to provide a comprehensive clinical and molecular characterization of LS-UC. We conducted a retrospective single-center study including patients with Lynch syndrome (LS) and histologically confirmed urothelial carcinoma (UC). Clinical, pathological, treatment, and follow-up data were collected. Targeted next-generation sequencing was performed on available tumor samples to assess genomic alterations, with particular attention to FGFR3 mutations. A total of 27 patients with LS-UC were identified, with a predominance of upper urinary tract involvement (70%). Most tumors were diagnosed at an early stage and initially managed with local treatment. During a median follow-up of 92 months, 48% of patients experienced recurrence, with a median time to recurrence of 37 months. Recurrences were predominantly local and were mainly managed with additional surgical or intravesical treatments. No deaths were attributable to UC at last follow-up. Molecular analysis was feasible in 9 cases. FGFR3 mutations were detected in 67% of evaluable samples, with the recurrent p.Arg248Cys hotspot identified in 55% of cases. Additional alterations involved TP53, SWI/SNF complex genes, and PIK3CA, which co-occurred with FGFR3 p.Arg248Cys. No gene fusions were identified. This study expands the limited molecular and clinical evidence on Lynch syndrome-associated urothelial carcinoma. Beyond confirming the recurrent role of FGFR3 (notably p.Arg248Cys), our comprehensive multigene profiling enriches the current genomic knowledge for this rare population. Multi-center collaborative efforts remain essential to aggregate larger datasets and ultimately guide personalized patient management.

Humans