PubMed HealthSearch

SEARCH · PubMed Health

Results for “coalescent tree inferences”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

5 recordsLinked to original sources

Comparing ARG Inference Methods Under Transmission of Reproductive Success: Tree Imbalance Matters.

Inferring coalescent trees from genomic data has become a major subject in population genetics, particularly with the recent advances in tree sequence reconstruction methods. However, it remains unclear how well these methods perform for imbalanced genealogies. Such imbalances can arise from processes such as cultural transmission of reproductive success (CTRS) or positive selection. Using simulated genomic data, we benchmarked three major software packages, SINGER, Relate, and tsinfer, by comparing the imbalance of reconstructed trees by these methods with that of the true simulated trees, for three indices that quantify this imbalance. The three methods performed well under scenarios yielding balanced trees. However, their accuracy declined as imbalance increased. Performances also varied with mutation rate, recombination rate, and sample size. This study opens possibilities for applying these methods to infer CTRS or positive selection in large-scale genomic datasets, using simulation-based inference such as approximate Bayesian computation.

Models, Genetic

Phlag: scalable detection of genomics regions with unexplained phylogenetic heterogeneity.

MOTIVATION: Phylogenetic analyses of entire genomes (phylogenomics) have revealed abundant heterogeneity of evolutionary histories. While much has been done to model this heterogeneity and to infer species trees despite it, the current toolkit has a limitation. Most methods assume that gene trees across the genome differ but are all sampled from the same distribution, defined by models such as the multi-species coalescent (MSC), and parametrized consistently across the genome. Empirical data strongly suggest this assumption is often violated because the species tree, its parameters, or the process generating the gene trees can all change across the genome. Errors in the data can further compound this heterogeneity. RESULTS: To address this challenge, we define the problem of detecting what segments of the genome are inconsistent with a putative species tree, even after allowing discordance according to MSC. We model gene trees not as a set, but rather as a series (a realization of a stochastic process) along genomic positions. We propose a Hidden Markov Model (HMM) approach applied to quartet statistics measured from gene trees and tie the model to MSC using simulations. The combined use of these three ideas leads to a scalable method called Phlag. On simulated and real data, we show that Phlag can detect many cases of change in underlying evolutionary processes, including reduced recombination rates, population size changes, and admixture, all using the same algorithm. AVAILABILITY AND IMPLEMENTATION: Phlag is available at github.com/bo1929/phlag. All results and scripts can be found at github.com/bo1929/shared.phlag.

Phylogeny

Phlag: Scalable detection of genomics regions with unexplained phylogenetic heterogeneity.

MOTIVATION: Phylogenetic analyses of entire genomes (phylogenomics) have revealed abundant heterogeneity of evolutionary histories. While much has been done to model this heterogeneity and to infer species trees despite it, the current toolkit has a limitation. Most methods assume that gene trees across the genome differ but are all sampled from the same distribution , defined by models such as the multi-species coalescent (MSC), and parametrized consistently across the genome. Empirical data strongly suggest this assumption is often violated because the species tree, its parameters, or the process generating the gene trees can all change across the genome. Errors in the data can further compound this heterogeneity. RESULTS: To address this challenge, we define the problem of detecting what segments of the genome are inconsistent with a putative species tree, even after allowing discordance according to MSC. We model gene trees not as a set, but rather as a series (a realization of a stochastic process) along genomic positions. We propose a Hidden Markov Model (HMM) approach applied to quartet statistics measured from gene trees and tie the model to MSC using simulations. The combined use of these three ideas leads to a scalable method called Phlag. On simulated and real data, we show that Phlag can detect many cases of change in underlying evolutionary processes, including reduced recombination rates, population size changes, and admixture, all using the same algorithm. AVAILABILITY AND IMPLEMENTATION: Phlag is available at github.com/bo1929/phlag . All results and scripts can be found at github.com/bo1929/shared.phlag .

Journal Article

Nuclear single-copy orthologous genes as phylogenomic markers for resolving the closely related firefly genera Pteroptyx, Medeopteryx, and Trisinuata (Coleoptera: Lampyridae: Luciolinae).

Fireflies (Lampyridae) are bioluminescent beetles with broad ecological roles across temperate and tropical ecosystems, occupying diverse habitats including forests, wetlands, grasslands, mangroves, and riverine systems. The subfamily Luciolinae is primarily distributed across Asia and the Indo-Pacific. Phylogenetic relationships among three closely related Luciolinae genera - Medeopteryx, Pteroptyx, and Trisinuata - remain unresolved using mitochondrial genome data alone. This study used nuclear genome data to resolve relationships among these genera and identify a lighter-weight nuclear marker panel for expanding taxon sampling. Draft genomes were reconstructed for fifteen firefly species, eight from the focal genera, and analyzed with five published firefly genomes. Using BUSCO and OrthoFinder, 1,011 nuclear single-copy orthologs (SCOs) were identified for phylogenomic inference. Discordance between concatenation- and coalescence-based phylogenies indicated incomplete lineage sorting (ILS). The coalescence-based phylogeny recoveredPteroptyxas monophyletic and sister to a (Medeopteryx,Trisinuata) clade, with Trisinuata nested within a non-monophyletic Medeopteryx; however, quartet support at the base of Pteroptyx, particularly at Pt. valida, was low.Filtering for compositional homogeneity, clock-likeness, and species-tree concordance yielded 103 SCOs with a significantly higher proportion of parsimony-informative sites than non-selected loci, retaining the backbone topology with higher gene concordance support at scored clades, while ILS-driven discordance at Pt. valida persists - confirming that the reduced panel retains phylogenetic resolving power for future taxon sampling. These findings demonstrate a practical framework for using nuclear SCOs to resolve close phylogenetic relationships within Luciolinae. Future work should expand taxon sampling - especially forTrisinuata - alongside long-read assemblies, for a more robust phylogenomic framework.

Fireflies

Beyond Level-1: Identifiability of a Class of Galled Tree-Child Networks.

Inference of phylogenetic networks is of increasing interest in the genomic era. However, the extent to which phylogenetic networks are identifiable from various types of data remains poorly understood, despite its crucial role in justifying methods. This work obtains strong identifiability results for large sub-classes of galled tree-child semidirected networks. Some of the conditions our proofs require, such as the identifiability of a network's tree of blobs or the circular order of 4 taxa around a cycle in a level-1 network, are already known to hold for many data types. We show that all these conditions hold for quartet concordance factor data under various gene tree models, yielding the strongest results from 2 or more samples per taxon. Although the network classes we consider have topological restrictions, they include non-planar networks of any level and are substantially more general than level-1 networks - the only class previously known to enjoy identifiability from many data types. Our work establishes a route for proving future identifiability results for tree-child galled networks from data types other than quartet concordance factors, by checking that explicit conditions are met.

Mathematical Concepts