PubMed HealthSearch

SEARCH · PubMed Health

Results for “algorithms”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Algorithms for the identification of prevalent diabetes in the All of Us Research Program validated using polygenic scores.

The All of Us Research Program (AoU) is an initiative designed to gather a comprehensive and diverse dataset from at least one million individuals across the USA. This longitudinal cohort study aims to advance research by providing a rich resource of genetic and phenotypic information, enabling powerful studies on the epidemiology and genetics of human diseases. One critical challenge to maximizing its use is the development of accurate algorithms that can efficiently and accurately identify well-defined disease and disease-free participants for case-control studies. This study aimed to develop and validate type 1 (T1D) and type 2 diabetes (T2D) algorithms in the AoU cohort, using electronic health record (EHR) and survey data. Building on existing algorithms and using diagnosis codes, medications, laboratory results, and survey data, we developed and implemented algorithms for identifying prevalent cases of type 1 and type 2 diabetes. The first set of algorithms used only EHR data (EHR-only), and the second set used a combination of EHR and survey data (EHR+). A universal algorithm was also developed to identify individuals without diabetes. The performance of each algorithm was evaluated by testing its association with polygenic scores (PSs) for type 1 and type 2 diabetes. We demonstrated the feasibility and utility of using AoU EHR and survey data to employ diabetes algorithms. For T1D, the EHR-only algorithm showed a stronger association with T1D-PS compared to the EHR + algorithm (DeLong p-value = 3 × 10-5). For T2D, the EHR + algorithm outperformed both the EHR-only and the existing T2D definition provided in the AoU Phenotyping Library (DeLong p-values = 0.03 and 1 × 10-4, respectively), identifying 25.79% and 22.57% more cases, respectively, and providing an improved association with T2D PS. We provide a new validated type 1 diabetes definition and an improved type 2 diabetes definition in AoU, which are freely available for diabetes research in the AoU. These algorithms ensure consistency of diabetes definitions in the cohort, facilitating high-quality diabetes research.

Humans

An optimal algorithm for automatic genotype elimination.

In an effort to accelerate likelihood computations on pedigrees, Lange and Goradia defined a genotype-elimination algorithm that aims to identify those genotypes that need not be considered during the likelihood computation. For pedigrees without loops, they showed that their algorithm was optimal, in the sense that it identified all genotypes that lead to a Mendelian inconsistency. Their algorithm, however, is not optimal for pedigrees with loops, which continue to pose daunting computational challenges. We present here a simple extension of the Lange-Goradia algorithm that we prove is optimal on pedigrees with loops, and we give examples of how our new algorithm can be used to detect genotyping errors. We also introduce a more efficient and faster algorithm for carrying out the fundamental step in the Lange-Goradia algorithm-namely, genotype elimination within a nuclear family. Finally, we improve a common algorithm for computing the likelihood of a pedigree with multiple loops. This algorithm breaks each loop by duplicating a person in that loop and then carrying out a separate likelihood calculation for each vector of possible genotypes of the loop breakers. This algorithm, however, does unnecessary computations when the loop-breaker vector is inconsistent. In this paper we present a new recursive loop breaker-elimination algorithm that solves this problem and illustrate its effectiveness on a pedigree with six loops.

Algorithms

Quantifying and improving rheumatoid arthritis algorithm performance in biobank settings.

OBJECTIVE: To quantify and improve the performance of standard rheumatoid arthritis (RA) algorithms in a biobank setting. METHODS: This retrospective cohort study within the Mayo Clinic (MC) Biobank and MC Tapestry Study identified RA cases by presence of at least two RA codes OR positive anti-cyclic citrullinated peptide antibodies (CCP) plus disease-modifying anti-rheumatic drug (DMARD) prescription as of 7/18/2022. Rheumatology physicians manually verified all RA cases using RA criteria and/or rheumatology physician diagnosis plus DMARD use. All other biobank participants served as non-RA controls. We defined seropositivity as rheumatoid factor and/or anti-CCP positivity. We assessed rules-based and Electronic Medical Records and Genomics (eMERGE) RA algorithms using positive predictive value (PPV). Finally, we developed a novel RA algorithm using a LASSO-based machine learning approach with five-fold cross validation. RESULTS: We identified 1,316 confirmed RA cases (968 MC Biobank, 348 Tapestry, 70 % seropositive) and 82,123 non-RA controls (mean age 65, 61 % female). The PPV of 3 RA codes was 43 %, codes plus DMARD was 54 %, and codes plus DMARD plus seropositivity was 85 %. The PPV of eMERGE was 77 %. Available in the MC Biobank, self-reported RA (PPV 10 %) only minimally improved algorithm performance (PPV from 83 % to 85 %), whereas family history of RA (PPV 3 %) worsened performance. At 90 % PPV, the novel RA algorithm incorporating key variables such as anti-CCP and DMARD use increased sensitivity by 4-11 % compared to eMERGE. CONCLUSION: Rules-based and eMERGE RA algorithms had worse performance in biobank than administrative settings. Our novel RA algorithm outperformed these standard algorithms.

Humans

Singletrack: an algorithm for improving memory consumption and performance of gap-affine sequence alignment.

MOTIVATION: Advances in DNA sequencing have outpaced advances in computation, making sequence alignment a major bottleneck in genome data analyses. Classical dynamic programming (DP) algorithms are particularly memory-intensive, especially when computing gap-affine and dual gap-affine alignments. Existing strategies to reduce memory consumption often sacrifice speed or alignment accuracy. RESULTS: We present Singletrack, an efficient algorithm for backtrace gap-affine and dual gap-affine alignments that requires storing a single DP matrix while preserving optimal alignment results. Compared to classical DP algorithms, Singletrack removes the need to store additional matrices (i.e. 2 for gap-affine and 4 for dual gap-affine), significantly reducing memory consumption and, in turn, reducing pressure on the memory hierarchy and improving overall performance. Most importantly, Singletrack is a general backtrace method compatible with state-of-the-art DP-based algorithms and heuristics, such as the Suzuki-Kasahara (SK) and the Wavefront Alignment (WFA) algorithms. We demonstrate that Singletrack reduces memory consumption for both SK and WFA algorithms, lowering SK usage by 2× and 4× and WFA usage by 3× and 5× for gap-affine and dual gap-affine alignments, respectively. Moreover, replacing KSW2's memory-reduction technique with Singletrack accelerates its SK implementation by up to 1.4× at the cost of doubling memory consumption, while Singletrack increases the performance of the WFA implementation in WFA2-lib by 1.2-2.1×. Compared to the efficient linear-memory BiWFA algorithm, the Singletrack-accelerated version of WFA trades a practical increase in memory usage for up to 5.2× higher performance. AVAILABILITY AND IMPLEMENTATION: The Singletrack implementations presented in this work are available on Zenodo (DOI: 10.5281/zenodo.18770585) and GitHub (https://github.com/LorienLV/singletrack).

Algorithms

Future promise, current clinical ambiguity: a systematic review of machine learning algorithm outputs predicting risk of cardiovascular disease.

OBJECTIVE: To examine whether the outputs of machine learning algorithms designed to predict risk of cardiovascular disease (CVD) address known deficiencies of the Framingham Risk Score (FRS) and improve risk estimates. METHODS: For this critical review, Medline, Embase and IEEE were searched from inception to 1 January 2025. Included were studies describing machine learning algorithms designed to specifically compare output of cardiovascular risk assessment with the FRS. Commentaries, letters, unpublished work or non-peer-reviewed papers were excluded.Following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, two reviewers screened titles and abstracts independently, then populated a purpose-built data extraction form. A subsequent qualitative thematic analysis focused on algorithms' strengths, added value, potential harms, unintended consequences and equity implications.The main outcome assessed was whether, among healthy adults, the algorithm improved CVD risk prediction relative to the FRS. RESULTS: Of 707 studies retrieved, 29 met inclusion criteria. 23 reported improved predictive ability relative to the FRS. Most datasets and/or medical records used included sociodemographic predictors of CVD not included among FRS inputs. Some added costly diagnostic tests like CT angiography to FRS screening indicators. When they were defined, inputs and outcomes such as hypertension or myocardial infarction did not always adhere to FRS values. Statistical significance was generally taken as a proxy for clinical significance. Some algorithms overestimated the number at risk compared with the FRS without discussing whether that larger proportion might be at risk of overdiagnosis rather than CVD, while a few decreased the proportion found to be at risk. CONCLUSIONS: Use of artificial intelligence to improve accuracy of risk assessment for CVD demonstrates the technological capacity to merge known sociodemographic predictors with biologic variables and examine non-linear interactions among these. Still needed to achieve patient benefit is clinical insight, adherence to screening principles and cost-benefit assessment of inputs selected.

Humans

dGAMLSS: an exact, distributed algorithm to fit Generalized Additive Models for Location, Scale, and Shape for privacy-preserving population reference charts.

MOTIVATION: There is growing interest in estimating population reference ranges across age and sex to better identify atypical clinically-relevant measurements throughout the lifespan. For this task, the World Health Organization recommends using Generalized Additive Models for Location, Scale, and Shape (GAMLSS), which can model non-linear growth trajectories under complex distributions that address the heterogeneity in human populations.Fitting GAMLSS models requires large, generalizable sample sizes, especially for accurate estimation of extreme quantiles, but obtaining such multi-site data can be challenging due to privacy concerns and practical considerations. In settings where patient data cannot be shared, privacy-preserving distributed algorithms for federated learning can be used, but no such algorithm exists for GAMLSS. RESULTS: We propose distributed GAMLSS (dGAMLSS), a distributed algorithm that can fit GAMLSS models across multiple sites without sharing patient-level data. This includes specific considerations for the fitting of smooth functions at varying levels of communication efficiency. We demonstrate the effectiveness of dGAMLSS in constructing population reference charts across clinical, genomics, and neuroimaging settings and show that dGAMLSS is able to reproduce pooled reference charts and inference down to numerical differences. AVAILABILITY AND IMPLEMENTATION: An R package providing examples of the dGAMLSS algorithm, as well as functions for sharing and aggregating site-specific parameters, is available at https://github.com/hufengling/dGAMLSS.

Algorithms

Algorithms to reconstruct past indels: The deletion-only parsimony problem.

Ancestral sequence reconstruction is an important task in bioinformatics, with applications ranging from protein engineering to the study of genome evolution. When sequences can only undergo substitutions, optimal reconstructions can be efficiently computed using well-known algorithms. However, accounting for indels in ancestral reconstructions is much harder. First, for biologically-relevant problem formulations, no polynomial-time exact algorithms are available. Second, multiple reconstructions are often equally parsimonious or likely, making it crucial to correctly display uncertainty in the results. Here, we consider a parsimony approach where only deletions are allowed, while addressing the aforementioned limitations. First, we describe an exact algorithm to obtain all the optimal solutions. The algorithm runs in polynomial time if only one solution is sought. Second, we show that all possible optimal reconstructions for a fixed node can be represented using a graph computable in polynomial time. While previous studies have proposed graph-based representations of ancestral reconstructions, this result is the first to offer a solid mathematical justification for this approach. Finally we provide arguments for the relevance of the deletion-only case for the general case.

Algorithms

Streamlining Diagnosis of Bardet-Biedl Syndrome: New Diagnostic Algorithm With Updated Criteria.

Considerable advances have been made in our understanding of Bardet-Biedl syndrome (BBS), particularly in its core clinical features and molecular genetics, warranting an update to the existing diagnostic criteria framework. Using a rigorous, evidence-based, and consensus-driven process, a multidisciplinary group of international experts and patient-led organizations developed an updated diagnostic algorithm. This algorithm provides practical, updated guidance for clinicians, including a pathway for accurately incorporating genetic findings into the diagnostic process. We recommend that a clinical diagnosis requires either 4 major criteria or 3 major and 2 minor criteria. Revised major criteria are retinal dystrophy, obesity (or overweight in individuals <&#x2009;2&#x2009;years old), congenital anomalies of the kidney and urinary tract or chronic kidney disease, hypogonadism/genital anomalies, neurodevelopmental/neurocognitive manifestations, and postaxial polydactyly. The diagnosis can also be established with a positive genetic testing result in patients exhibiting &#x2265;&#x2009;1 major criterion, provided that genetic findings should be interpreted in the context of the patient's clinical presentation, age, family history, and overlap with related ciliopathies. These consensus criteria offer a simple algorithm incorporating updated definitions for major and minor criteria and genetic testing to support a timely and accurate diagnosis of patients with BBS, inform genetic counseling, and potentially facilitate earlier access to treatment. Trial Registration: CRIBBS Registry; ClinicalTrials.gov: NCT02329210.

Humans

Generating three-dimensional genome structures with a variational quantum algorithm.

Chromosome conformation capture experiments have revealed the underlying spatial interactions that govern three-dimensional (3D) genome organization and topology. Detecting 3D contacts between genomic loci considerably enhances our understanding of fundamental regulatory processes. Modeling 3D structures from experimental contact matrices can further contextualize the relationship between 3D genome organization and regulation. While classical algorithms have been successful in reconstructing genomic conformations, we investigate the prospect of quantum computation to aid in modeling the conformational space. In this context, we propose a novel variational quantum algorithm (VQA) to model the distribution of 3D genomic structures from experimental contact data. Through rigorous evaluations, we demonstrate the capability of our algorithm to sample ensembles of viable 3D conformations that agree well with experimental and simulated contact data. Furthermore, we extend our methodology to model the conformational space of a single cell or a population of cells. In the advent of sufficient quantum utility, the insights gained from this study can serve as a foundation for investigating high-resolution, large-scale ensembles of genomic conformations through generative VQAs.

Algorithms

Optimising parent selection in plant breeding: comparing metaheuristic algorithms for genotype building.

Stacking desirable haplotypes across the genome to develop superior genotypes has been implemented in several crop species. A major challenge in Optimal Haplotype Selection is identifying a set of parents that collectively contain all desirable haplotypes, a complex combinatorial problem with countless possibilities. In this study, we evaluated the performance of metaheuristic search algorithms (MSAs)-genetic algorithm (GA), differential evolution (DE), particle swarm optimisation (PSO), and simulated annealing (SA) for optimising parent selection under two genotype building (GB) objectives: Optimal Haplotype Selection (OHS) and Optimal Population Value (OPV). Using a diverse wheat population of 583 lines genotyped for 29,972 SNPs, forming 7645 haplotype blocks and phenotyped for stripe rust scores, we assessed each algorithm's performance across fitness optimisation, convergence speed, and computational efficiency. GA consistently achieved high fitness and rapid convergence, while DE showed robustness but required longer runtime and careful tuning. PSO performed well under the OHS criterion but was less effective for OPV. SA, although computationally lighter, was less consistent in finding optimal solutions. Simulation over 100 breeding cycles showed that OHS outperformed both OPV and GEBV-based selection in long-term genetic gain and diversity retention. OHS maintained heterozygosity and additive variance, which are key for sustainable improvement, while GEBV selection led to early allele fixation. Our findings underscore the potential of GB strategies that prioritise the collective performance of parent sets rather than individual ranking to enhance selection outcomes in genomic-assisted breeding programmes.

Plant Breeding

Targeted next-generation sequencing for drug-resistant tuberculosis diagnosis: implementation considerations for bacterial load, regimen selection and diagnostic algorithm placement.

INTRODUCTION: Early and accurate diagnosis of drug-resistant tuberculosis (DR-TB) is essential for improving treatment outcomes. Phenotypic drug susceptibility testing (pDST) is comprehensive but slow, while rapid molecular assays provide resistance information for a limited number of drugs. Targeted next-generation sequencing (tNGS) offers the potential for broad and rapid resistance detection, but its integration into diagnostic algorithms has been hindered by uncertainty about its placement within existing workflows. METHODS: This study evaluated the extent to which two tNGS solutions-Deeplex Myc-TB (GenoScreen) and TB Drug Resistance Test (Oxford Nanopore Technologies, ONT)-provided interpretable drug resistance results that could inform regimen design, in comparison to other WHO-recommended molecular assays and pDST. Data were collected from three high-burden DR-TB settings under the Seq&Treat study. Sequencing success rates and drug resistance detection were analysed based on: (1) the initial Xpert MTB/RIF result (very low, low, medium, high), (2) resistance results for drugs in WHO-recommended regimens and (3) performance relative to other WHO-endorsed assays. The potential impact of different algorithms on the estimates was also considered. Key factors influencing successful tNGS adoption within diagnostic pathways were identified, leveraging insights from the Seq&Treat diagnostic accuracy study. RESULTS: Sequencing success rates were 88.5% (GenoScreen) and 93.1% (ONT) across 763 samples. While tNGS provided complete resistance data for 73%-86% of drugs in recommended regimens, pDST achieved 92%-93%. Both tNGS solutions matched or exceeded the sensitivity of WHO-recommended molecular assays. CONCLUSIONS: This study highlights the critical role of tNGS as a centralised tool for comprehensive drug resistance testing to inform DR-TB treatment decisions following initial screening assays. By complementing existing molecular tests with tNGS, diagnostic workflows can be optimised to ensure timely and comprehensive resistance detection. These findings support policy updates to integrate tNGS into global TB diagnostic algorithms. TRIAL REGISTRATION NUMBER: NCT04239326.

Humans

MWENA: a novel sample re-weighting-based algorithm for disease classification and data interpretation using extracellular vesicles omics data.

BACKGROUND AND OBJECTIVE: Extracellular vesicles (EVs), considered as a form of liquid biopsy, have gained significant attention in recent years due to their stability and the preservation of disease markers. Research studies underscore the clinical significance of molecules found in EVs, highlighting their role as communicative mediators between cells. However, analyzing this data is challenging due to noisy measurements, having far more variables than samples, and some groups (e.g., disease subtypes or experimental conditions) having much less data than others. We therefore develop an algorithm to address aforementioned challenges for the classification of imbalanced EVs omics data. METHODS AND RESULTS: We propose the EV Meta-Weight Elastic Net Algorithm (MWENA), which utilizes logistic regression with elastic net regularization for the classification and identification of EV signatures, effectively addressing the challenges posed by high-dimensional small sample sizes. To mitigate issues related to class imbalance and high noise levels, MWENA incorporates an automatic sample re-weighting function, which uses a meta-net to adaptively learn generalizable patterns directly from the data itself. We validate the MWENA algorithm on both simulated data and EVs omics data, covering six classification tasks that involve four different types of diseases (pancreatic ductal adenocarcinoma, interstitial lung diseases, colorectal cancer, and ovarian cancer) and three clinical scenarios (disease diagnosis, disease-stage screening, and disease-subtype classification). Compared to other machine learning methods, MWENA demonstrates superiority in identifying small class samples and achieves the highest scores in both sensitivity and G-means. Biological analysis is also performed to further explore the significance of selected signatures as biological markers and their roles in disease mechanisms. CONCLUSIONS: We anticipate that our proposed approach will take a modest step in harnessing EV omics data to discover biomarkers, aiding researchers in gaining a comprehensive understanding of biological processes.

Extracellular Vesicles

Gut microbiota-derived metabolites target C5AR1/KDM2A/HCAR3 axis in inflammatory bowel disease: a multi-machine learning algorithms and molecular docking study.

BACKGROUND: Inflammatory bowel disease (IBD) is a chronic recurrent disorder. Gut microbiota-derived metabolites regulate intestinal homeostasis, but their molecular mechanisms in IBD remain unclear. Current studies lack systematic "microbiota-metabolite-target" network mining with multi-method validation. This study integrates network pharmacology, three machine learning algorithms, and molecular docking to construct this regulatory network in IBD. METHODS: Transcriptome data were obtained from the Gene Expression Omnibus (GEO) database. Differentially expressed genes (DEGs) were identified using limma (p < 0.05, |log2FC| > 0.5). Weighted gene co-expression network analysis (WGCNA) with an optimal soft threshold of &#x3b2; = 7 was performed to identify key module genes. Candidate genes were obtained by intersecting DEGs, gut microbiota-associated genes from the gutMGene database, and WGCNA module genes. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses were conducted to explore the functional roles of candidate genes. Core genes were identified using three machine learning algorithms (LASSO, Boruta, and SVM-RFE), followed by protein-protein interaction (PPI) network analysis. Molecular docking was performed to assess the binding affinities between hub proteins and gut microbiota-derived metabolites. RESULTS: A total of 885 DEGs were identified between the IBD and control groups, including 463 upregulated and 422 downregulated genes. WGCNA identified 280 key module genes from the purple and yellow modules. The intersection of DEGs, gut microbiota-associated genes, and WGCNA module genes yielded 19 core candidate genes. PPI network analysis combined with three machine learning algorithms jointly identified C5AR1, KDM2A, and HCAR3 as core hub genes. ROC curve analysis demonstrated that all three hub genes achieved AUC values greater than 0.7 in both the training and validation sets, indicating excellent diagnostic performance for IBD. Enrichment analysis revealed significant associations with the TNF, NF-&#x3ba;B, and IL-17 signaling pathways. Molecular docking confirmed stable binding of C5AR1 with 1,3-Diphenylpropan-2-Ol (-7.87 &#xb1; 0.83 kcal&#xb7;mol-&#xb9;) and HCAR3 with 3-Indolepropionic Acid (-6.35 &#xb1; 0.70 kcal&#xb7;mol-&#xb9;), both below -5.0 kcal&#xb7;mol-&#xb9;. CONCLUSION: This study first constructs a "gut microbiota-metabolite-hub gene" axis in IBD, providing a computational framework for microbiota-targeted precision therapy, and identifying C5AR1/KDM2A/HCAR3 as computationally predicted diagnostic biomarkers and 1,3-Diphenylpropan-2-Ol/3-Indolepropionic Acid as candidate intervention molecules that warrant further experimental validation.

Molecular Docking Simulation

Tumor-immune partitioning and clustering algorithm for identifying tumor-immune cell spatial interaction signatures within the tumor microenvironment.

BACKGROUND: Growing evidence supports the importance of characterizing the organizational patterns of various cellular constituents in the tumor microenvironment in precision oncology. Most existing data on immune cell infiltrates in tumors, which are based on immune cell counts or nearest neighbor-type analyses, have failed to fully capture the cellular organization and heterogeneity. METHODS: We introduce a computational algorithm, termed Tumor-Immune Partitioning and Clustering (TIPC), that jointly measures immune cell partitioning between tumor epithelial and stromal areas and immune cell clustering versus dispersion. As proof-of-principle, we applied TIPC to a prospective cohort incident tumor biobank containing 931 colorectal carcinoma cases. TIPC identified tumor subtypes with unique spatial patterns between tumor cells and T lymphocytes linked to certain molecular pathologic and prognostic features. T lymphocyte identification and phenotyping were achieved using multiplexed (multispectral) immunofluorescence. In a separate hepatocellular carcinoma cohort, we replaced the stromal component with specific immune cell types-CXCR3+CD68+ or CD8+-to profile their spatial relationships with CXCL9+CD68+ cells. RESULTS: Six unsupervised TIPC subtypes based on T lymphocyte distribution patterns were identified, comprising two cold and four hot subtypes. Three of the four hot subtypes were associated with significantly longer colorectal cancer (CRC)-specific survival compared to a reference cold subtype. Our analysis showed that variations in T-cell densities among the TIPC subtypes did not strictly correlate with prognostic benefits, underscoring the prognostic significance of immune cell spatial patterns. Additionally, TIPC revealed two spatially distinct and cell density-specific subtypes among microsatellite instability-high colorectal cancers, indicating its potential to upgrade tumor subtyping. TIPC was also applied to additional immune cell types, eosinophils and neutrophils, identified using morphology and supervised machine learning; here two tumor subtypes with similarly low densities, namely 'cold, tumor-rich' and 'cold, stroma-rich', exhibited differential prognostic associations. Lastly, we validated our methods and results using The Cancer Genome Atlas colon and rectal adenocarcinoma data (n = 570). Moreover, applying TIPC to hepatocellular carcinoma cases (n = 27) highlighted critical cell interactions like CXCL9-CXCR3 and CXCL9-CD8. CONCLUSIONS: Unsupervised discoveries of microgeometric tissue organizational patterns and novel tumor subtypes using the TIPC algorithm can deepen our understanding of the tumor immune microenvironment and likely inform precision cancer immunotherapy.

Humans

HiCForecast: dynamic network optical flow estimation algorithm for spatiotemporal Hi-C data forecasting.

MOTIVATION: The exploration of the 3D organization of DNA within the nucleus in relation to various stages of cellular development has led to experiments generating spatiotemporal Hi-C data. However, there is limited spatiotemporal Hi-C data for many organisms, impeding the study of 3D genome dynamics. To overcome this limitation and advance our understanding of genome organization, it is crucial to develop methods for forecasting Hi-C data at future time points from existing timeseries Hi-C data. RESULT: In this work, we designed a novel framework named HiCForecast, adopting a dynamic voxel flow algorithm to forecast future spatiotemporal Hi-C data. We evaluated how well our method generalizes forecasting data across different species and systems, ensuring performance in homogeneous, heterogeneous, and general contexts. Using both computational and biological evaluation metrics, our results show that HiCForecast outperforms the current state-of-the-art algorithm, emerging as an efficient and powerful tool for forecasting future spatiotemporal Hi-C datasets. AVAILABILITY AND IMPLEMENTATION: HiCForecast is publicly available at https://github.com/OluwadareLab/HiCForecast.

Algorithms

MarkerMatch: a proximity-based probe-matching algorithm for joint analysis of copy-number variants from different genotyping arrays.

MOTIVATION: Copy-number variants (CNVs) are a form of genetic structural variation with increasing importance in complex human disorders. Both DNA sequencing and microarray data can be used to detect CNVs, which can be used in genetic association tests. Unlike genotypes, CNV detection in microarrays requires the use of observed intensity signals at each probe, which limits the imputability for analyses that span multiple array types. Thus far, a consensus set of probes (those present on all arrays) has been used to circumvent the problem of differing array-specific sensitivities. This has led to excessive reduction in overall sensitivity since arrays can have an undesirably low probe overlap. To overcome this limitation, we developed MarkerMatch, a proximity-based algorithm that matches probes across different genotyping microarrays to maximize the number of probes considered in the CNV calling algorithm, thereby increasing the resolution and sensitivity while preserving precision. RESULTS: By analyzing CNV calls from 4906 individuals genotyped across three different arrays, we show that the MarkerMatch approach improves sensitivity by increasing the density of probes available for CNV calling while maintaining precision or improving it relative to the current practice (e.g. use of consensus probes only). We further demonstrate that MarkerMatch matches the CNV detection from current practice in terms of F1 score and PPV for larger CNVs. We also optimize MarkerMatch parameters, DMAX and Method, and find an optimal DMAX setting at 10&#x2009;kb, with no clear optimal candidate based on Method, indicating that parameters for this metric should be determined on a use case basis. AVAILABILITY: The R package for MarkerMatch is available at: https://github.com/FranjoIM/MarkerMatch. The code used for analysis and implementation is available at: https://doi.org/10.5281/zenodo.18460979. The live notebook is available at https://fivankovic.notion.site/2026-markermatch.

DNA Copy Number Variations

MarkerMatch: A Proximity-Based Probe-Matching Algorithm for Joint Analysis of Copy-Number Variants from Different Genotyping Arrays.

MOTIVATION: Copy-number variants (CNVs) are a form of genetic structural variation with increasing importance in complex human disorders. Both DNA sequencing and microarray data can be used to call CNVs, which can be used in association tests, such as association between CNV number and disease status. Unlike genotypes, CNV detection in microarrays requires the use of observed intensity signals at each probe, which limits the imputability for analyses that span multiple array types. Thus far, a consensus set of probes (the intersection encompassing the probes that occur in common on all arrays) has been used to circumvent the problem of differing array-specific sensitivities. This has, however, led to excessive reduction in overall sensitivity of CNV calls as arrays can have an undesirably low overlap of probe sets. To overcome this limitation, we developed MarkerMatch, a proximity-based algorithm that matches probes across different genotyping microarrays to maximize the number of probes considered in the CNV calling algorithm, thereby increasing the resolution and sensitivity while preserving precision. RESULTS: By analyzing CNV calls from 4,906 individuals genotyped across three different arrays (Global Screening Array, Omni2.5 array, and Omni Express Exome array), we show that the MarkerMatch approach improves sensitivity by increasing the density of probes available for CNV calling while maintaining precision or improving it relative to the current practice (e.g., use of consensus probes only). We further demonstrate that MarkerMatch exceeds the output from current practice in terms of F1 score, Fowlkes-Mallows index, and Jaccard index. We also optimize MarkerMatch parameters, D MAX and Method, and find an optimal D MAX setting at 10kb, with no clear optimal candidate based on Method, indicating that parameters for this metric should be determined on a use case basis.

Journal Article

Parallel algorithms for phylogenetic inference under a structured coalescent approximation.

While advances in molecular epidemiology and computational modeling have enhanced our capacity to track pathogen evolution, the accurate reconstruction of spatiotemporal transmission dynamics remains essential for developing epidemic preparedness frameworks and implementing outbreak response measures. Structured coalescent models offer a phylogeographic framework by restricting lineage coalescence events to geographically proximate host populations. Although the Bayesian structured coalescent approximation (BASTA) provides a tractable approach, contemporary phylogeographic analyses involving dozens of geographic localities and hundreds to thousands of viral genomes substantially exceed the computational capacity of existing implementations. The BASTA likelihood scales cubically with deme count and quadratically with sequence count due to matrix exponentiation and pairwise coalescent probability calculations. Here, we introduce a comprehensive algorithmic restructuring of the structured coalescent likelihood that eliminates redundancies, optimizes memory access, and exposes parallelization opportunities. Our approach reorganizes computations along three dimensions: (i) independent calculation of deme-transition probability matrices across time intervals; (ii) simultaneous evaluation of partial likelihood vectors within temporal slices; and (iii) concurrent aggregation of coalescent probabilities. Algorithmic restructuring cuts average coalescent likelihood computation by 7-8 fold, and parallelization further boosts performance to 10-26 fold, enabling joint phylogeographic analyses of dengue virus across 10 South American countries and H5N1 avian influenza across 20 Eurasian regions to finish in a fraction of prior time. This computational efficiency also enables comparison between backward-in-time structured coalescent approximations and forward-in-time phylogeographic methods, revealing that the former provides appropriately conservative posterior estimates, particularly at intermediate phylogenetic depths. We integrate our implementation into the popular BEAST X and BEAGLE software packages, with an accompanying interface in BEAUti X to easily set up the analyses, providing researchers with an accessible and scalable tool for real-time phylogeographic surveillance of rapidly evolving pathogens.

Journal Article