PubMed HealthSearch

SEARCH · PubMed Health

Results for “Lineage tree reconstruction”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Bayesian inference of lineage trees by joint analysis of single-cell multimodal lineage-tracing data with BiLinT.

The advent of single-cell lineage-tracing technologies has enabled the simultaneous profiling of gene expression and lineage barcodes. However, accurate, high-resolution reconstruction of cell lineage trees remains challenging because most existing approaches treat these modalities separately and therefore fail to fully exploit their complementary information. Here we present BiLinT, a Bayesian framework that jointly models multimodal single-cell lineage-tracing data for lineage tree reconstruction. BiLinT integrates barcode evolution (a continuous-time Markov chain) with gene expression dynamics (an Ornstein-Uhlenbeck process) within a unified probabilistic model. Across synthetic and real data sets, BiLinT provides accurate lineage-tree reconstruction and reveals differentiation-associated clonal structure and developmental fate biases.

Journal Article

Paired Single-Cell Transcriptome and DNA Barcode Detection in Zebrafish Using ScarTrace.

ScarTrace is a CRISPR/Cas9-based genetic lineage tracing method that allows for uniquely barcoding the DNA of single cells at a target GFP sequence during developing zebrafish embryos. Single cells from barcoded adult zebrafish can be isolated from various tissues (e.g., marrow, brain, eyes, fins), and their transcriptome and barcode sequences are captured by single-cell cDNA amplification and genomic DNA nested PCR, respectively. Computationally, cell type and barcode identification permit clone tracing and lineage tree reconstruction of tissues to unravel fate decisions during embryogenesis.

Animals

Statistical method for estimating the standard errors of branch lengths in a phylogenetic tree reconstructed without assuming equal rates of nucleotide substitution among different lineages.

A statistical method is developed for estimating the standard errors of branch lengths in a phylogenetic tree reconstructed without assuming equal rates of nucleotide substitution among different lineages. This method can be easily used for testing whether the length of an interior branch in a reconstructed tree is positive, i.e., whether the topology of the tree is correct. Computer simulations indicate that this method is appropriate for a statistical test. As an example, this method is applied to phylogenetic trees reconstructed for the four hominoid species: human, chimpanzee, gorilla, and orangutan. The results obtained show that the present method provides a powerful statistical test.

Animals

LINNAEUS: Simultaneous Single-Cell Lineage Tracing and Cell Type Identification.

A key goal of biology is to understand the origin of the many cell types that can be observed during diverse processes such as development, regeneration, and disease. Single-cell RNA-sequencing (scRNA-seq) is commonly used to identify cell types in a tissue or organ. However, organizing the resulting taxonomy of cell types into lineage trees to understand the origins of cell states and relationships between cells remains challenging. Here we present LINNAEUS (Spanjaard et al, Nat Biotechnol 36:469-473. https://doi.org/10.1038/nbt.4124 , 2018; Hu et al, Nat Genet 54:1227-1237. https://doi.org/10.1038/s41588-022-01129-5 , 2022) (LINeage tracing by Nuclease-Activated Editing of Ubiquitous Sequences)-a strategy for simultaneous lineage tracing and transcriptome profiling in thousands of single cells. By combining scRNA-seq with computational analysis of lineage barcodes, generated by genome editing of transgenic reporter genes, LINNAEUS can be used to reconstruct organism-wide single-cell lineage trees. LINNAEUS provides a systematic approach for tracing the origin of novel cell types, or known cell types under different conditions.

Single-Cell Analysis

LAML-Pro: joint maximum likelihood inference of cell genotypes and cell lineage trees.

MOTIVATION: Recent dynamic lineage tracing technologies use genome editing to induce heritable mutations, or edits, that accumulate across successive cell divisions. These edits are measured using single-cell sequencing or imaging, providing data to reconstruct cell lineages at single-cell resolution. Current computational approaches to infer cell lineage trees, or phylogenies, from these data perform two separate steps: (i) Identify each cell's edits (genotype) from the raw sequencing or imaging data; (ii) Infer a cell lineage tree from the cell genotypes. However, genotyping cells is an inexact process and genotype errors can yield an inaccurate lineage tree. For example, using fluorescence based-imaging to measure edits results in a high fraction (≈25%-50%) of uncertain or erroneous genotypes. RESULTS: We introduce Lineage Analysis via Maximum Likelihood with PRobabilistic Observations (LAML-Pro), an algorithm that jointly infers cell genotypes and a cell lineage tree. LAML-Pro is based on the Probabilistic Mixed-type Missing Observation (PMMO) model, which we derive to describe both the genome editing and genotype observation processes. LAML-Pro constructs lineage trees from thousands of cells in under an hour by leveraging the sparsity of transitions under the PMMO model. On simulated data, we demonstrate that LAML-Pro corrects genotype errors and infers substantially more accurate trees than existing methods which are vulnerable to genotype errors. Applied to data from two recent imaging-based lineage tracing systems, LAML-Pro reduces genotype errors by 5-fold and produces more spatially coherent lineage trees compared to existing methods. AVAILABILITY AND IMPLEMENTATION: LAML-Pro is implemented in C++ and is available as both a command-line interface and as a Python library at: github.com/raphael-group/LAML-Pro.

Cell Lineage

LAML-Pro: Joint Maximum Likelihood Inference of Cell Genotypes and Cell Lineage Trees.

MOTIVATION: Recent dynamic lineage tracing technologies use genome editing to induce heritable mutations, or edits, that accumulate across successive cell divisions. These edits are measured using single-cell sequencing or imaging, providing data to reconstruct cell lineages at single-cell resolution. Current computational approaches to infer cell lineage trees, or phylogenies, from these data perform two separate steps: (1) Identify each cell's edits (genotype) from the raw sequencing or imaging data; (2) Infer a cell lineage tree from the cell genotypes. However, genotyping cells is an inexact process and genotype errors can yield an inaccurate lineage tree. For example, using fluorescence based-imaging to measure edits results in a high fraction (≈ 25-50%) of uncertain or erroneous genotypes. RESULTS: We introduce Lineage Analysis via Maximum Likelihood with PRobabilistic Observations (LAML-Pro), an algorithm that jointly infers cell genotypes and a cell lineage tree. LAML-Pro is based on the Probabilistic Mixed-type Missing Observation (PMMO) model, which we derive to describe both the genome editing and genotype observation processes. LAML-Pro constructs lineage trees from thousands of cells in under an hour by leveraging the sparsity of transitions under the PMMO model. On simulated data, we demonstrate that LAML-Pro corrects genotype errors and infers substantially more accurate trees than existing methods which are vulnerable to genotype errors. Applied to data from two recent imaging-based lineage tracing systems, LAML-Pro reduces genotype errors by 5-fold and produces more spatially coherent lineage trees compared to existing methods. AVAILABILITY AND IMPLEMENTATION: LAML-Pro is freely available at: github.com/raphael-group/LAML-Pro.

Journal Article

Introgression shapes the genomic conflict landscape of Malus, providing evidence for a reticulate backbone in a woody crop lineage.

Phylogenomic discordance is widespread across plants, but its evolutionary significance is often obscured when conflict is treated primarily as analytical noise rather than as evidence of underlying processes. In woody lineages in particular, incomplete lineage sorting, introgression, and genome duplication can interact over long timescales to produce complex genomic histories that are not adequately summarized by a strictly bifurcating tree. Here, we use Malus as a model woody genus to investigate how these processes structure conflict across a genus-scale, accession-based phylogenomic framework. Using broad taxon sampling, hundreds of nuclear loci, plastid genomes, and genome-wide SNP summaries, we reconstruct a robust nuclear backbone for sampled Malus lineages and evaluate where discordance is concentrated and which processes best explain it. Nuclear analyses resolve eight major clades, whereas conflict is non-random and localized to recurrent hotspots rather than evenly distributed across the tree. Cytonuclear discordance is similarly concentrated, especially around Clade H, represented by sampled accessions of M. tschonoskii, where localized plastid-nuclear disagreement is consistent with candidate plastid capture or organellar introgression. Multiple complementary analyses further indicate that the strongest conflict is not explained by ILS alone, but instead reflects lineage-structured introgression, while polyploid complexes represent additional localized sources of evolutionary complexity. Together, these results provide evidence for a reticulate genomic backbone in Malus and show how integrating nuclear, plastid, and genome-wide conflict analyses can help distinguish background discordance from process-specific signals in woody plant radiations. Several lineage-level reticulation hypotheses identified here should now be tested with broader population-level sampling and curated reference accessions.

Malus

Genetically distinct within-host subpopulations of hepatitis C virus persist after Direct-Acting Antiviral treatment failure.

Analysis of viral genetic data has previously revealed distinct within-host population structures in both untreated and interferon-treated chronic hepatitis C virus (HCV) infections. While multiple subpopulations persisted during the infection, each subpopulation was observed only intermittently. However, it was unknown whether similar patterns were also present after Direct-Acting Antiviral (DAA) treatment, where viral populations were often assumed to go through narrow bottlenecks. Here we tested for the maintenance of population structure after DAA treatment failure, and whether there were different evolutionary rates along distinct lineages where they were observed. We analysed whole-genome next-generation sequencing data generated from a randomised study using DAAs (the BOSON study). We focused on samples collected from patients (N=84) who did not achieve sustained virological response (i.e., treatment failure) and had sequenced virus from multiple timepoints. Given the short-read nature of the data, we used a number of methods to identify distinct within-host lineages including tracking concordance in intra-host nucleotide variant (iSNV) frequencies, applying sequenced-based and tree-based clustering algorithms to sliding windows along the genome, and haplotype reconstruction. Distinct viral subpopulations were maintained among a high proportion of individuals post DAA treatment failure. Using maximum likelihood modelling and model comparison, we found an overdispersion of viral evolutionary rates among individuals, and significant differences in evolutionary rates between lineages within individuals. These results suggest the virus is compartmentalised within individuals, with the varying evolutionary rates due to different viral replication rates and/or different selection pressures. We endorse lineage awareness in future analyses of HCV evolution and infections to avoid conflating patterns from distinct lineages, and to recognise the likely existence of unsampled subpopulations.

Humans

Phylogenomics and evolution of the Lauraceae based on targeted capture data.

The family Lauraceae, a hyper-diverse magnoliid family comprising approximately 63 genera and over 3,000 species, plays a key ecological role in tropical and subtropical forests. Yet deep relationships among its nine tribes remain unresolved, likely due to limited sampling and complex evolutionary processes such as incomplete lineage sorting (ILS) and gene flow. To address these challenges, we generated datasets of 255 single-copy nuclear genes and chloroplast genomes using a newly designed Lauraceae-specific probe set, achieving the most comprehensive genus-level sampling (84%) to date. Phylogenomic analyses reconstructed a robust nuclear tree, which resolved the Neocinnamomeae as sister to the Caryodaphnopsideae and revealed pronounced gene tree conflict and pervasive cytonuclear discordance. To investigate the evolutionary processes underlying these patterns, comprehensive analyses were conducted. The results indicate that conflicting nuclear gene trees reflect the combined effects of ILS, gene tree estimation error, and gene flow, with ILS dominating across the core Lauraceae, whereas cytonuclear discordance is primarily driven by extensive gene flow. Diversification analyses further indicate that episodes of rapid lineage accumulation coincide with major gene flow events, suggesting a potential role of gene flow in the diversification of Lauraceae. Overall, this study provides a robust nuclear phylogenomic framework for Lauraceae and demonstrates that gene flow had profound effects on its evolutionary history, shedding light on the contribution of gene flow to the diversification of hyper-diverse tropical plant lineages.

Cytonuclear discordance

Structural genomics sheds light on protein functions and remote homologs across the insect tree of life.

Protein structure bridges the sequence-function relationship, enabling deep exploration of biological processes across diverse organisms. Insects, the most diverse animal lineage, accounting for over 50% of all described animal species, provide an exceptional system for exploring sequence-structure-function relationships. Here, we reconstructed a comprehensive and well-resolved phylogeny of 4854 insects, spanning all orders. Leveraging this framework, we created an atlas of 13.29 million predicted protein structures from 824 representative species, including 11.63 million newly predicted structures. Structural clustering revealed that proteins with divergent sequences but similar structures could be effectively grouped together. Structural similarity searches against proteins with well-characterized functions yielded annotations for 7.61 million insect proteins, including up to 14% of previously unannotated proteins. We further identified 750 million remote homologs between insect proteins, many of which trace back to ancient branches of the insect phylogeny. Remarkably, despite extensive sequence divergence, cGAS-like receptors (cGLRs) were structurally conserved across all 824 insects. Experimental assays demonstrated that these structurally identified cGLRs play a crucial role in antiviral defense in the yellow fever mosquito. Our findings highlight the significance of structural genomics for understanding protein function and evolution across the tree of life.

Animals

Phylogenetic position of the enigmatic starfish family Podosphaerasteridae (Asteroidea, Valvatida) with a morphological observation of the skeletal structure by micro-CT.

Background The genus Podosphaeraster comprises seven species, characterised by a distinctive spherical body, all currently known from the seabed at depths below approximately 70 m. Its peculiar morphology has made its phylogenetic placement a subject of ongoing debate. It was initially suggested to be placed in Sphaerasteridae, the same family as fossil species. However, subsequent detailed skeletal analyses of the fossil forms suggested that this similarity was likely due to convergent evolution. Recent molecular analyses have revealed that Valvatida, in which the genus is currently placed, is likely a large polyphyletic group, leaving its taxonomic position still unresolved. New information A detailed examination of the internal skeletal structure of Podosphaeraster toyoshiomaruae, collected from the seas around Japan, was conducted using micro-focus X-ray computed tomography. Concurrently, shotgun sequencing was performed to identify key molecular markers for recent asteroid phylogeny. Additionally, shotgun sequencing determined the complete mitochondrial genome. Despite conservative evolution amongst asteroidean mitochondrial genomes, a translocation of the COX2 gene was revealed, representing the first discovery of a major protein-coding gene translocation within Asteroidea. In the phylogenetic tree, P. toyoshiomaruae was positioned as the most basal lineage within Valvatida. However, the statistical support for this placement was low, potentially due to the long-branch attraction caused by the excessively rapid evolutionary rate. Images reconstructed by micro-CT confirmed the presence of calcified reinforcement in the mesentery and showed its detailed structure for the first time. The mesentery skeleton was found to connect to the V-plate and five pairs of plates, including three kinds of marginal plates. This suggests that the marginal plates of this species may not be homologous with those of other asteroids. Although varying degrees of marginal plate reduction are shared with the order Velatida, we consider this to be a case of convergent evolution. Our phylogenetic analysis indicates a close relationship between P. toyoshiomaruae and Poraniidae (and other Valvatida), all of which possess marginal plates differentiated to varying extents, the homology of which remains uncertain.

Asteroidea

Molecular phylogeny of the subgenus Sophophora of Drosophila derived from large subunit of ribosomal RNA sequences.

RNA sequencing has been used to assess the relationships among species of the subgenus Sophophora of the genus Drosophila. Two divergent domains, D1 and D2, of the large ribosomal RNA (28S), totalling 550 nucleotides have been sequenced using the rRNA direct sequencing method. A tree has been reconstructed from the neighbor-joining algorithm and the confidence intervals were evaluated by the bootstrap procedure. Results have shown that the branching of the willistoni and saltans groups of the subgenus Sophophora is very ancient and probably predates that of the subgenus Drosophila. The other groups and subgroups of Sophophora are clustered in three main lineages: 1) the melanogaster and oriental subgroups; 2) the montium subgroup; 3) the ananassae subgroup of the melanogaster group clustered with the fima and obscura groups. Thus, in comparison with our results, several taxa of various ranks appear paraphyletic (the genus Drosophila, the subgenus Sophophora and the melanogaster group). Our biochemical phylogeny is only in partial agreement with the pattern of Throckmorton's radiations as well as with classical taxonomy, both based on morphological data.

Animals

In genomes we trust: Assessing genomic reliability within the family Nectriaceae.

Reliable evolutionary inference increasingly depends on public genome resources, and the effects of uneven assembly quality, incomplete metadata, and biased taxonomic sampling remain poorly quantified. Using the species-rich fungal lineage Nectriaceae as a model system, we analysed 1530 genome sequence assemblies to assess metadata completeness, sampling representation, and genome quality. One-third of the assemblies lacked essential metadata, sequencing was heavily skewed toward a few agriculturally important lineages, and sampling of many genera was limited or nonexistent. BUSCO and QUAST metrics revealed substantial heterogeneity in assembly quality, with widespread fragmentation and numerous assemblies falling outside expected quality thresholds. From 763 single-copy orthologs identified in 576 higher-quality genomes, we reconstructed a phylogenomic backbone and quantified gene- and site-level concordance across the tree. Although major clades were broadly recovered, extensive gene-tree discordance and a polyphyletic Fusarium nisikadoi species complex revealed unresolved boundaries and conflict among loci. These results show how data quality, incomplete sampling, and discordant genomic histories can constrain phylogenomic resolution, and provide a general framework for improving comparative genomic resources and large-scale evolutionary inference.

Gene-tree discordance

Whole genome-based reclassification of the genus Metabacillus: Proposal for five novel genera, Chryseobacillus gen. nov., Cohnibacillus gen. nov., Salimetabacillus gen. nov., Pantoeobacillus gen. nov., and Lutimetabacillus gen. nov. and the description of one novel bacterial species, Chryseobacillus diguaensis sp. nov. isolated from soil in the Digua reservoir.

Comprehensive phylogenomic and comparative genomic analyses were conducted to clarify the taxonomic boundaries of the genus Metabacillus. Phylogenetic trees reconstructed from a set of single-copy orthologous proteins (SCOPs) revealed that the genus, as currently defined, is polyphyletic. The type species of the genus Metabacillus and its closest relatives formed a consistent clade, herein designated as Metabacillus sensu stricto. The remaining species were grouped into three well-supported clades: Kandeliae, Indicus, and Mangrovi, and two single-taxon lineages: M. arenae and M. lacus. The phylogenomic delineation found in these divergent taxa was corroborated by either inconsistent distribution patterns or the absence of previously defined conserved signature indels (CSIs) specific to Metabacillus. Genomic metrics, including Average Nucleotide Identity (ANI), Average Amino acid Identity (AAI), and digital DNA-DNA hybridization (dDDH) further supported the taxonomic delineation proposed here. The observed genomic divergence was mirrored by phenotypic differences, including variations in GC content ranges. Based on this polyphasic evidence, we propose the reclassification of the genus Metabacillus taxa into five novel genera: Chryseobacillus gen. nov. (encompassing the Kandeliae clade), Cohnibacillus gen. nov. (M. lacus), Salimetabacillus gen. nov. (M. arenae), Pantoeobacillus gen. nov. (Indicus clade), and Lutimetabacillus gen. nov. (Mangrovi clade). The core lineage is retained as Metabacillus sensu stricto, for which an emended description of the genus Metabacillus is also provided. A novel bacterial strain, designated as MAU-250T, was isolated from a soil sample collected on the shore of an artificial reservoir in the Andean foothills of the Maule Region in central Chile. Public metagenome screening supported a low-abundance taxon with broad ecological adaptability, preferentially associated with soil habitats. A polyphasic analysis based on phenotypic traits and genomic distances (78.0% ANIb and 19.8% dDDH against its closest relative) also supported its designation as a novel species, for which the name Chryseobacillus diguaensis sp. nov. is proposed. The type strain is MAU-250T (=RGM 3146T = IMI 507634T).

Phylogeny

Genome-wide barriers to gene flow reveal the genetic basis of viviparity evolution in a lizard.

Viviparity (live-bearing) is a major evolutionary transition repeatedly linked with ecological and evolutionary diversification throughout vertebrates. Live-bearing reproduction entails a novel suite of phenotypes and life history traits, but the genetic processes by which such a reproductive innovation evolves are unknown. Remarkable among amniotes, the common lizard (Zootoca vivipara) has extant oviparous (egg-laying) and viviparous lineages and a to-date unresolved history of parity mode emergence. This species represents an ideal model to reconstruct the evolutionary and genetic mechanisms of how viviparity arises. By analyzing whole genomes of individuals from across the species' distribution, we robustly show that viviparity evolved once. However, gene flow from oviparous to viviparous populations is found to be long-term and extensive, causing pronounced gene tree discordance. We inferred signals of selection for viviparity in many independent regions across the genome, and these were recruited over considerable time. Genomic barriers to gene flow between oviparity and viviparity were found genome wide. These are enriched for regions under selection for parity mode and for genes known to be involved in pregnancy and parturition in squamates and mammals. Further implicating their functional role in viviparity, we show that genes in genomic regions under selection and resisting gene flow are more highly expressed in the uterus of viviparous lizards during pregnancy. Our study demonstrates that viviparity in an amniote evolved by selection in the face of gene flow and primarily by the genome-wide accumulation of functional regulatory variants. These results reveal how complex adaptive innovations can arise and be maintained.

Animals

Identification of a novel HIV-1 circulating recombinant form (CRF209_cpx) and its descendant unique recombinant form (URF) CRF209_cpx/B among MSM in Guangdong, southern China.

BACKGROUND: The epidemic of human immunodeficiency virus type 1 (HIV-1) continues to pose a significant global health challenge, with increasing genetic diversity. The co-circulation of multiple subtypes among the local population facilitates the emergence of unique or circulating recombinant forms (URFs or CRFs). In China, the predominant strains include CRF07_BC, CRF01_AE, CRF55_01B, and subtype B. This study characterizes a novel CRF209_cpx and its descendant recombinant CRF209_cpx/B among men who have sex with men (MSM) in Guangdong, southern China. METHODS: Individuals infected with URFs with similar genetic characteristics were recruited during routine surveillance of pretreatment drug resistance. Near full-length genomes (NFLGs) were amplified with two overlapping fragments using a serial dilution nested PCR approach after reverse transcription. We used SimPlot and IQ-TREE softwares to conduct recombination analyses and phylogenetic inferences. Time-scaled maximum clade credibility (MCC) phylogenetic trees were reconstructed using BEAST software to estimate evolutionary origins. Genotypic drug resistance mutations were interpreted via the Stanford HIV Database, and coreceptor usage was predicted using geno2pheno coreceptor 2.5 and the HIVcoPRED tool. RESULTS: Four NFLG sequences were obtained and identified as a novel CRF209_cpx, generated by recombination among CRF01_AE, CRF07_BC and subtype B. Phylogenetic analyses revealed that all the parental segments clustered with lineages prevalent among MSM in China. Bayesian evolutionary analysis estimated that the most recent common ancestor (tMRCA) of CRF209_cpx to have evolved between 2011 and 2013. The fifth strain was identified as a URF recombined from nascent CRF209_cpx and B. No transmitted drug resistance mutation was detected in these five sequences. The four CRF209_cpx sequences primarily utilized the CXCR4 coreceptor, while the URF exhibited R5/X4 dual tropism. CONCLUSIONS: The emergence of the complex CRF209_cpx and novel URF of CRF209_cpx/B highlights the active HIV-1 epidemic within the MSM population in Guangdong, underscoring the necessity for enhanced molecular surveillance and precise public health intervention in this key population.

HIV-1

The archaeal roots of eukaryotic life.

Resolving the biological and geological events that led to the origin of eukaryotes is an ongoing challenge in biology. A major step in the evolution of complex cellular life was the merger between an ancestral host cell and a bacterium (that became the mitochondrion) some two billion years ago. Recently, metagenomics has enabled the reconstruction of a broad diversity of genomes, referred to as the Asgard Archaea. The Asgards are monophyletic with eukaryotes on the tree of life. Asgards have an array of genes, previously thought exclusive to eukaryotes, involved in cellular trafficking, the ubiquitin system, endosomal sorting, and cytoskeleton formation, with growing evidence demonstrating the functions of these proteins mirror those in eukaryotes. This gene repertoire suggests that these Archaea are descendants of the archaeal host from which eukaryotes evolved. Increased sampling has revealed that Asgard lineages are metabolically versatile and play key roles in various ecosystems and uncovered evolutionary transitions between Archaea and eukaryotes, such as innovations in eukaryotic defense systems. The positioning of eukaryotes in the Asgards is debated, but eukaryotes appear to branch within the Heimdallarchaeia. Lineages within this group, particularly Hodarchaeales and Kariarchaeaceae, contain a broad repertoire of eukaryote-like traits, including high-energy yielding metabolisms. Observing and studying Asgard interactions with bacterial descendants of mitochondria in a modern setting will transform our understanding of the origin of complex cellular life.

Archaea

Relative efficiencies of the maximum-parsimony and distance-matrix methods of phylogeny construction for restriction data.

The relative efficiencies of the maximum-parsimony (MP), UPGMA, and neighbor-joining (NJ) methods in obtaining the correct tree (topology) for restriction-site and restriction-fragment data were studied by computer simulation. In this simulation, six DNA sequences of 16,000 nucleotides were assumed to evolve following a given model tree. The recognition sequences of 20 different six-base restriction enzymes were used to identify the restriction sites of the DNA sequences generated. The restriction-site data and restriction-fragment data thus obtained were used to reconstruct a phylogenetic tree, and the tree obtained was compared with the model tree. This process was repeated 300 times. The results obtained indicate that when the rate of nucleotide substitution is constant the probability of obtaining the correct tree (Pc) is generally higher in the NJ method than in the MP method. However, if we use the average topological deviation from the model tree (dT) as the criterion of comparison, the NJ and MP methods are nearly equally efficient. When the rate of nucleotide substitution varies with evolutionary lineage, the NJ method is better than the MP method, whether Pc or dT is used as the criterion of comparison. With 500 nucleotides and when the number of nucleotide substitutions per site was very small, restriction-site data were, contrary to our expectation, more useful than sequence data. Restriction-fragment data were less useful than restriction-site data, except when the sequence divergence was very small. UPGMA seems to be useful only when the rate of nucleotide substitution is constant and sequence divergence is high.

Computer Simulation