PubMed HealthSearch

SEARCH · PubMed Health

Results for “phylogenetic network”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

SNaQ.jl: Improved scalability for level-1 phylogenetic network inference.

MOTIVATION: Phylogenetic networks represent complex biological scenarios that are overlooked in trees, such as hybridization and horizontal gene transfer. Although numerous methods have been developed for phylogenetic network inference, their scalability is severely limited by the computational demands of likelihood optimization and the vastness of network space. Composite (or pseudo-) likelihood approaches like SNaQ have improved computational tractability for network inference, but they remain inadequate for datasets of sizes routinely handled by tree inference methods. RESULTS: Here, we introduce SNaQ.jl, a new standalone Julia package with the composite likelihood inference originally implemented within PhyloNetworks.jl as well as new scalability features that enhance computational efficiency through (i) parallelization of quartet likelihood calculations during composite likelihood computation, (ii) weighted random selection of quartets, and (iii) probabilistic decision-making during network search. Through a simulation study and empirical data analysis, we show that this new version of SNaQ.jl (version 1.1) improves average runtimes by up to 499% on average with no change in function parameters or method accuracy. AVAILABILITY AND IMPLEMENTATION: SNaQ.jl is a new open source Julia package available at https://github.com/JuliaPhylo/SNaQ.jl.

Phylogeny

A measure of the denseness of a phylogenetic network.

The concept of phylogenetic denseness bears critically on the accuracy of evolutionary pathways inferred from experimentally sequenced proteins isolated from extant species. In this paper I develop an objective measure, rho, of denseness to supplement previous intuitive concepts and which permits one to use this concept in comparing the quality of different evolutionary reconstructions. This measure is used to examine several published phylogenetic trees: insulin, alpha-hemoglobin, beta-hemoglobin, myoglobin, cytochrome c, and the parvalbumin family. The paper emphasizes 1) the importance of denseness in accurately estimating the number of nucleotide replacements which separate homologous sequences when this estimation is made by the method of parsimony, 2) the value of this concept in assessing the quality of those estimates, and 3) the use of this concept as a biologically practical heuristic method for identifying poorly studied regions in a phylogenetic tree, whether or not the tree was obtained by the parsimony method.

Biological Evolution

PaNDA: Efficient Optimization of Phylogenetic Diversity in Networks.

Phylogenetic diversity (PD) plays an important role in biodiversity, conservation, and evolutionary studies by measuring the diversity of a set of taxa based on their phylogenetic relationships. In phylogenetic trees, a subset of k taxa with maximum PD can be found by a simple and efficient greedy algorithm. However, this algorithmic tractability is lost when considering phylogenetic networks, which incorporate reticulate evolutionary events such as hybridization and horizontal gene transfer. To address this challenge, we introduce PaNDA (Phylogenetic Network Diversity Algorithms), the first software package and interactive graphical user-interface for exploring, visualizing, and maximizing diversity in phylogenetic networks. PaNDA includes a novel algorithm to find a subset of k taxa with maximum diversity, running in polynomial time for networks of bounded scanwidth, a measure of tree-likeness of a network that grows slower than the well-known level measure. This algorithm considers the variant of PD on networks in which the branch lengths of all paths from the root to the selected taxa contribute towards their diversity. We demonstrate the scalability of this algorithm on simulated networks, successfully analyzing level-15 networks with up to 200 taxa in seconds. We also provide a proof-of-concept analysis using a phylogenetic network on Xiphophorus species, illustrating how the tool can support diversity studies based on real genomic data. The software is easily installable and freely available at https://github.com/nholtgrefe/panda. Additionally, we extend the definition of PD to semi-directed phylogenetic networks, which are mixed graphs increasingly used in phylogenetic analysis to model uncertainty of the root location. We prove that finding a subset of k taxa with maximum diversity remains NP-hard on semi-directed networks, but do present a polynomial-time algorithm for networks with bounded level.

network

Episode clustering in phylogenetic networks.

MOTIVATION: The classical duplication episode clustering (EC) model introduced by Guigó et al. in the 1990s provides a foundational approach for inferring genomic duplication events crucial to understanding genome evolution. This model clusters single gene duplications from a collection of gene trees at locations in the species tree to minimize the total number of such locations, called duplication episodes. However, it does not capture reticulate evolutionary histories. RESULTS: Here, we introduce NetEC, a novel extension of this problem to phylogenetic networks. To solve NetEC, we first develop a polynomial-time dynamic programming (DP) algorithm for testing whether a given set of network nodes can serve as episode locations. We then propose a main inference algorithm that utilizes this DP component to optimize the episode count; while the feasibility test runs in polynomial time, the full optimization has exponential worst-case complexity, and an optional heuristic mode is provided for larger instances. We also propose an extended episode analysis procedure that identifies additional genomic duplication candidates below reticulation nodes, complementing the main algorithm by resolving potential upward clustering of duplications induced by reticulation. We evaluate our method on simulated data and on an empirical Pandanales dataset comprising over 29 000 gene trees, demonstrating exact and accurate inference of genomic duplication events even in the presence of multiple reticulations. AVAILABILITY AND IMPLEMENTATION: All experiments were conducted using the NetEC tool (https://github.com/ppgorecki/netec), with all input data, scripts, and parameter settings for reproduction available in the same repository.

Phylogeny

Beyond Level-1: Identifiability of a Class of Galled Tree-Child Networks.

Inference of phylogenetic networks is of increasing interest in the genomic era. However, the extent to which phylogenetic networks are identifiable from various types of data remains poorly understood, despite its crucial role in justifying methods. This work obtains strong identifiability results for large sub-classes of galled tree-child semidirected networks. Some of the conditions our proofs require, such as the identifiability of a network's tree of blobs or the circular order of 4 taxa around a cycle in a level-1 network, are already known to hold for many data types. We show that all these conditions hold for quartet concordance factor data under various gene tree models, yielding the strongest results from 2 or more samples per taxon. Although the network classes we consider have topological restrictions, they include non-planar networks of any level and are substantially more general than level-1 networks - the only class previously known to enjoy identifiability from many data types. Our work establishes a route for proving future identifiability results for tree-child galled networks from data types other than quartet concordance factors, by checking that explicit conditions are met.

Mathematical Concepts

Phylotranscriptomics Allows Distinguishing Major Gene Flow Events from Incomplete Lineage Sorting in Rapidly Diversifying Mimetic Orchids (Genus Ophrys).

Ophrys orchids (or bee orchids) provide an outstanding example of a plant adaptive radiation. Over the last 5 million years, this genus has diversified into hundreds of taxa as a result of its unconventional pollination strategy, known as "sexual swindling". However, the rapid and substantial diversification of this genus, combined with its capacity for hybridization and large genome size, poses significant challenges in addressing its systematics. We used phylotranscriptomics as a genome complexity reduction technique to infer the phylogenetic relationships among Ophrys main lineages. More than seven thousand gene trees enabled us to determine the relative contributions of gene flow and incomplete lineage sorting (ILS) in Ophrys evolution. First, we propose a new phylogenetic hypothesis for the genus with an unprecedented resolution that largely confirms the relationships between the main Ophrys lineages, but also provides new insights within each subgenera. By combining phylogenetic network inference with introgression analyzes based on gene tree topologies and branch lengths, we then show that the numerous phylogenetic incongruences among gene tree topologies result from a pervasive background of ILS, over which stand out several well-supported, ancient and potentially adaptive gene flow events between lineages. These major gene flow events provide a new perspective on the evolution of the Ophrys genus and its pollination, questioning previous hypotheses inferred without considering its reticulate evolution, and providing a better understanding of discrepancies observed among previous phylogenetic studies of the genus.

Orchidaceae

Mitochondrial DNA clones and matriarchal phylogeny within and among geographic populations of the pocket gopher, Geomys pinetis.

Restriction endonuclease assay of mitochondria DNA (mtDNA) and standard starch-gel electrophoresis of proteins encoded by nuclear genes have been used to analyze phylogenetic relatedness among a large number of pocket gophers (Geomys pinetis) collected throughout the range of the species. The restriction analysis clearly distinguishes two populations within the species, an eastern and a western form, which differ by at least 3% in mtDNA sequence. Qualitative comparisons of the restriction phenotypes can also be used to identify mtDNA "clones" within each form. The mtDNA clones interconnect in a phylogenetic network which represents an estimate of matriarchal phylogeny for G. pinetis. Although the protein electrophoretic data also differentiate the eastern and western forms, the data are of limited usefulness in establishing relationships among more local subpopulations. The comparison between these two data sets suggests that restriction analysis of mtDNA is probably unequalled by other techniques currently available for determining phylogenetic relationships among conspecific organisms.

Albumins

Cytonuclear conflict and reticulate evolution in the Morelloid clade (Solanum, Solanaceae): Insights from genome skimming and network Phylogenomics.

The Morelloid clade (black nightshades) is one of the most strongly supported clades within the megadiverse Solanum genus. It comprises 76 globally distributed, non-spiny herbaceous and suffrutescent species. While often erroneously considered poisonous weeds, several species are economically important as orphan crops. The clade is closely related to tomato and potato but, due to a lack of focused breeding efforts, remains a putative reservoir of genetic diversity for crop improvement. Despite this potential, we lack fundamental knowledge on the evolution of the Morelloid clade. The group includes polyploid species with unknown parental origins-likely reflecting reticulate processes such as hybridization, introgression, and associated backcrossing events. Prior analyses have been unable to disentangle these processes, leaving the mechanisms underlying reticulate evolution in the Morelloid clade poorly understood. Here, we use genome skimming to produce a well-supported maximum likelihood plastid phylogeny from complete circularized plastomes and a coalescent-based species tree from combined Angiosperms353 and conserved ortholog set nuclear markers. Our dataset, composed of previously published data and deep genome skimming from herbarium samples, spans 26 Morelloid species. To investigate phylogenetic discordance, we used a nuclear phylogenetic network, multispecies coalescent simulations, a fused rooted nuclear chloroplast tree, and quantification of nuclear gene tree concordance. We show that incongruence between nuclear and plastid trees is pervasive and cannot be explained by incomplete lineage sorting alone. Instead, our results demonstrate that events consistent with repeated chloroplast capture have shaped the reticulate evolutionary history of the clade, especially among African polyploid and Pan-American diploid lineages.

Phylogeny

Historical introgression as a driver of diversification of diploid Picris (Compositae) in the Mediterranean Basin.

The Mediterranean Basin is recognized as one of the world's most prominent biodiversity hotspots, where past climatic changes have driven range shifts, secondary contact between populations, and gene exchange. This study investigates the impact of historical introgression on the diversification of diploid members of the genus Picris (Compositae). Using nuclear and plastid genome data obtained through the Hyb-Seq approach, we assess whether introgression contributed to the evolution of the Mediterranean Picris, potentially giving rise to multiple regional endemics. We also test whether introgression was associated with the transfer of traits such as life strategy and fruit morphology, which are involved in habitat-specific adaptation. Phylogenetic network analysis revealed two major introgression events that shaped evolutionary trajectories within the genus. The earliest and most complex events involved the Turkish endemic P. campylocarpa, which hybridized with the most recent common ancestor (MRCA) of the P. cyprica-P. pauciflora lineage and with the MRCA of the B1 subclade, comprising the P. hieracioides group and the P. scaberrima-P. strigosa lineage. The latter introgression preceded shifts from iteroparity to semelparity and from heterocarpy to homocarpy, ruling out an adaptive introgression origin for these traits. Nevertheless, all detected historical introgression events contributed to the diversification of diploid Picris taxa.

Diploidy

Mitochondrial DNA control-region and coding-region data highlight geographically structured diversity and post-domestication population dynamics in worldwide donkeys.

Donkeys (Equus asinus) have been used extensively in agriculture and transportations since their domestication, ca. 5000-7000 years ago, but the increased mechanization of the last century has largely spoiled their role as burden animals, particularly in developed countries. Consequently, donkey breeds and population sizes have been declining for decades, and the diversity contributed by autochthonous gene pools has been eroded. Here, we examined coding-region data extracted from 164 complete mitogenomes and 1392 donkey mitochondrial DNA (mtDNA) control-region sequences to (i) assess worldwide diversity, (ii) evaluate geographical patterns of variation, and (iii) provide a new nomenclature of mtDNA haplogroups. The topology of the Maximum Parsimony tree confirmed the two previously identified major clades, i.e. Clades 1 and 2, but also highlighted the occurrence of a deep-diverging lineage within Clade 2 that left a marginal trace in modern donkeys. Thanks to the identification of stable and highly diagnostic coding-region mutational motifs, the two lineages were renamed as haplogroup A and haplogroup B, respectively, to harmonize clade nomenclature with the standard currently adopted for other livestock species. Control-region diversity and population expansion metrics varied considerably between geographical areas but confirmed North-eastern Africa as the likely domestication center. The patterns of geographical distribution of variation analyzed through phylogenetic networks and AMOVA confirmed the co-occurrence of both haplogroups in all sampled populations, while differences at the regional level point to the joint effects of demography, past human migrations and trade following the spread of donkeys out of the domestication center. Despite the strong decline that donkey populations have undergone for decades in many areas of the world, the sizeable mtDNA variability we scored, and the possible identification of a new early radiating lineage further stress the need for an extensive and large-scale characterization of donkey nuclear genome diversity to identify hotspots of variation and aid the conservation of local breeds worldwide.

Animals

IQ-NET: fast and accurate quartet phylogenetic inference using deep learning trained on empirical DNA alignments.

Phylogenetic inference is fundamental to modern biology, with many applications including evolutionary biology, epidemiology, and comparative genomics. While maximum likelihood and Bayesian methods remain the gold standard for phylogenetic analysis, they rely on simplifying assumptions and are computationally intensive. Recent machine learning approaches for phylogenetics offer speed advantages, but have several limitations: exclusive reliance on simulated data for training, inadequate handling of gaps, and sensitivity to input sequence order. Here, we introduce IQ-NET (Intelligent Quartet NETwork), a deep learning framework that solves these limitations to infer four-taxon trees. IQ-NET estimates both tree topology and branch lengths directly from gapped alignments. IQ-NET outperforms existing machine learning methods in terms of accuracy, and obtained a 24-fold speedup compared with the widely used maximum likelihood software, IQ-TREE. We finally introduce a pipeline using IQ-NET and the ASTRAL software to reconstruct a larger species tree, i.e., with more than four taxa.

Empirical data training

Polyploidy Arithmetic.

Polyploidy occurs in plants and animals, and is an important force in speciation and genome evolution. The main focus of this paper is the following fundamental question that was recently posed by Huber and Maher: Given the ploidy numbers of a collection of extant species, or their ploidy profile, what is the smallest number of hybridizations needed in any evolutionary history for these species to completely represent these numbers? In this paper, we shall show that this question can be rephrased in terms of addition chains and the closely related addition sequences, which have been studied for over a century in mathematics and computer science. These are sequences of natural numbers that start with 1, so that each number in the sequence larger than 1 is the sum of two other numbers arising earlier in the sequence. In our first main result, we show that finding the smallest number of hybridization events to explain a ploidy profile, or the hybrid number, is equivalent to solving the so-called addition sequence problem. This immediately implies that computing the hybridization number is computationally intractable. Even so, it also leads to new connections to representing polyploid evolution using networks. More specifically, in our second main result we show that ploidy profiles representable by tree-child networks are exactly the addition chains, implying a polynomial-time algorithm for identifying these profiles. We then consider beaded tree-child networks, which permit the representation of autopolyploidy events, and in our third main result we provide a greedy polynomial-time algorithm to decide whether a given profile can be realized by such a network. We expect that our results can be leveraged in future work through, for example, making use of known algorithms for computing short addition sequences to give bounds for the hybrid number, and in guiding network reconstruction for polyploid species.

Polyploidy

Logical Exploration of Cinnamoyl-Containing Nonribosomal Peptides via Metabologenomic Targeting and Regulator Overexpression.

A targeted method for discovering cinnamoyl-containing nonribosomal peptides (CCNPs), a unique class of bioactive compounds, was devised by using cinnamoyl isomerase, a key enzyme in the biosynthesis of the cinnamoyl moiety, as a genome mining probe. A total of 39 hit strains were obtained, including 35 from polymerase chain reaction-based screening of the in-house bacterial library (2.5% of 1400 strains) targeting the cinnamoyl isomerase-encoding gene and 4 from the genome mining of online databases. Sequence similarity networking and phylogenetic analyses of the isomerase amplicons (∼530 bp) classified the CCNPs into three major substructure-based groups (Z-, E-, and M-type CCNPs) and revealed distinct clade-structure relationships (13 clades). To overcome the challenge of silent biosynthetic gene clusters, we activated these clusters by overexpressing conserved cluster-situated LuxR regulators combined with extensive culture optimization. CCNP production was metabolomically detected in the bacterial extracts by using the characteristic UV absorption and MS/MS fragments of cinnamoyl moieties. CCNP production was observed in 20 of the 39 hit strains, resulting in the isolation of 6 new CCNPs, including oxy-skyllamycin B (2), gwanacinnamycin (3), and luxocinnamycins A-D (4-7), with high structural novelty. Their structures were elucidated using comprehensive spectroscopic analyses and multiple-step chemical derivatizations, and the putative biosynthetic pathways were bioinformatically proposed. Gwanacinnamycin (3) exhibited significant antimycobacterial activity, whereas luxocinnamycin A (4) displayed moderate antiproliferative activity against stomach cancer cells. Our findings highlight a targeted metabologenomic approach combined with transcriptional regulator overexpression as a logical and efficient platform for the discovery of bioactive compounds from nature.

Peptides

Genome-wide SNP data support species boundaries in sympatric Polylepis Ruiz & Pav. (Rosaceae) species from Bolivia and Ecuador.

Species delimitation in the South American genus Polylepis is notoriously challenging due to high morphological similarity and phenotypic plasticity, likely driven by hybridization and gene flow. Previous phylogenetic studies suggested that genetic structure aligns more strongly with geography than with taxonomy, questioning existing species concepts and hampering conservation efforts. We used double-digest RAD sequencing (ddRADseq) to generate genome-wide SNP data for 11 Polylepis species sampled across multiple localities in Bolivia and Ecuador. Population genetic analyses, phylogenetic inference, and network approaches were combined to assess whether genetic structure aligns more closely with taxonomy or geography. Morphologically defined species formed largely cohesive genetic lineages across regions, with species identity explaining substantially more genetic variation than locality. While localized admixture and reticulation were detected among closely related taxa, widespread species showed strong genetic cohesion and clear separation from congeners. Our results indicate that the sampled Polylepis species from Bolivia and Ecuador maintain distinct genetic identities despite localized signals consistent with gene flow. This genome-wide support for current taxonomy highlights Polylepis as a valuable model for studying speciation under gene flow and indicates that multiple geographic sampling will be essential in reconstructing a robust phylogeny of the genus, with important implications for conservation planning in Andean montane forests.

Bolivia

GraphyloVar: predicting the impact of non-coding variants using a multi-species sequence model.

MOTIVATION: Understanding the functional impact of genetic variants is a key problem for precision medicine. Tools like CADD, PhyloP, and PhastCons are useful, but they often look at each position in the genome in isolation. This means they can miss important information from the evolutionary history that connects different species. In this paper, we extend our previous model, Graphylo, to predict the effects of variants. Our new model, GraphyloVar, is built to directly utilize the phylogenetic tree that relates the species. RESULTS: GraphyloVar is a deep learning model that considers both DNA sequence and evolutionary patterns from many species. It uses two main components: Graph Convolutional Networks (GCNs) to process the phylogenetic tree, and Transformer encoders to extract features from the DNA sequences. Pre-trained to predict population-level allele frequencies on the TOPMed whole-genome sequencing cohort, GraphyloVar achieves an AUROC of 0.6246 zero-shot on &#x223c;149M held-out variants, and an ensemble with CADD reaches 0.6442 (+0.020, P<10-15). Fine-tuned GraphyloVar achieves the highest AUROC across all 13 MPRA benchmark datasets. By integrating deep learning with explicit phylogenetic input, GraphyloVar offers a powerful and complementary approach to variant effect prediction that utilizes the full evolutionary history from many species to better identify and prioritize important non-coding variants. AVAILABILITY AND IMPLEMENTATION: Code and datasets are available at https://github.com/DongjoonLim/GraphyloVar under DOI: 10.5281/zenodo.20616818.

Phylogeny

Genome-wide characterization of heat shock protein genes reveals thermal stress-responsive candidates in Litopenaeus vannamei.

Heat shock proteins (HSPs) are conserved molecular chaperones involved in protein folding, refolding, aggregation prevention, and degradation of damaged proteins. However, the genomic organization and thermal responsiveness of HSP genes in the Pacific white shrimp (Litopenaeus vannamei) remain incompletely understood. Here, we performed a genome-wide analysis of the HSP gene family and examined its phylogenetic relationships, structural features, duplication patterns, sequence variation, interaction networks, and transcriptional responses to acute heat stress. A total of 34 HSP genes were identified and classified into the HSP90, HSP70, HSP40/DNAJ, HSP60, and small HSP families. Phylogenetic, motif, gene structure, synteny, and subcellular localization analyses revealed evolutionary conservation and structural diversification among family members. Three duplicated gene pairs were identified, comprising two segmental duplications and one tandem duplication. All pairs exhibited Ka/Ks ratios below 1, consistent with purifying selection of varying strength. Sequence analysis identified 295 nonsynonymous single-nucleotide polymorphisms, of which 12 were consistently predicted to be deleterious by multiple algorithms. Protein-protein interaction analysis indicated enrichment of protein-folding and cellular stress-response functions. RT-qPCR analysis showed significant induction of HSPA4, HSP90AA1, TRAP1, BiP, and DNAJA1 after 6, 12, and 24&#xa0;h of exposure to 34&#xa0;&#xb0;C, whereas DNAJC3 was significantly induced only at 12&#xa0;h. All six genes reached their highest transcript abundance at 12&#xa0;h. These findings may provide a genomic framework for HSP genes in L. vannamei and identify candidate genes and variants associated with thermal stress responses.

Animals

Distribution of vasotocin- and mesotocin-like immunoreactivities in the brain of the South African clawed frog Xenopus-laevis.

In order to obtain more insight into primitive and derived conditions of neuropeptidergic systems in vertebrates, in particular amphibians, we have studied immunohistochemically the distribution of vasotocin (AVT) and mesotocin (MST) neuronal elements in the brain of the South African clawed frog Xenopus laevis. Apart from a well-developed hypothalamohypophysial system, the antibodies revealed the existence of extrahypothalamic AVT- and MST-immunoreactive cell groups as well as extensive extrahypothalamic networks of immunoreactive fibres, thus confirming the phylogenetic constancy of this condition in vertebrates. The wide distribution of AVT- and MST-immunoreactive fibres throughout the brains of amphibians suggests that the two neuropeptidergic systems are involved not only in hypothalamohypophysial interactions, but also, as in mammals, in a variety of other brain functions. In particular, the relationship of AVT- and MST-immunoreactive fibres with catecholaminergic cell bodies was noted. The present study has underscored once again that considerable differences in relative densities of AVT- and MST-immunoreactive fibres occur between species, even within a single order of vertebrates.

Animals

Phylogenetic Methods Meet Deep Learning.

Deep learning (DL) has been widely used in various scientific fields, but its integration into phylogenetics has been slower, primarily due to the complex nature of phylogenetic data. The studies that apply DL to sequencing data often limit analyses to four-taxon trees. Many of these studies serve as "proof of principle" and perform similarly to traditional phylogeny reconstruction methods. New ways of using training data, such as encoding with compact bijective ladderized vectors or transformers, enable the handling of much larger trees and genomic data sets. This short perspective focuses on the application of DL in phylogenetics, introducing prevalent DL architectures. We highlight potential problems in the field by discussing the risks of using simulation-based training data and emphasize the importance of reproducibility and robustness in computational estimates. Finally, we explore promising research areas, including the combination of phylogenetics and population genetics in DL, the analysis of neighbor dependencies, and the potential to significantly reduce computational cost compared to traditional methods. This perspective illustrates the potential of DL in complementing traditional phylogeny reconstruction methods and aiding the advancement of phylogenetic analysis, especially in performing computationally demanding tasks such as model selection or estimating branch support values.

Humans