PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “bootstrap”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

New aspects on lanosterol 14alpha-demethylase and cytochrome P450 evolution: lanosterol/cycloartenol diversification and lateral transfer.

Sterol 14alpha-demethylase (CYP51) is a member of the cytochrome P450 superfamily, widely found in animals, fungi, and plants but present in few prokaryotic groups. CYP51 is currently believed to be the ancestral cytochrome P450 that has been transferred from prokaryotes to eukaryotic kingdoms. We propose an alternate view of CYP51 evolution that has an impact on understanding the evolution of the entire CYP superfamily. Two hundred forty-nine bacterial and four archaeal CYP sequences have been aligned and a bacterial CYP tree designed, showing a separation of two branches. Prokaryotic CYP51s cluster to the minor branch, together with other eukaryote-like CYPs. Mycobacterial and methylococcal CYP51s cluster together (100% bootstrap probability), while Streptomyces CYP51 remains on a distant branch. A CYP51 phylogenetic tree has been constructed from 44 sequences resulting in a ((plant, bacteria),(animal, fungi)) topology (100% bootstrap probability). This is in accordance with the lanosterol/cycloartenol diversification of sterol biosynthesis. The lanosterol branch (nonphotosynthetic lineage) follows the previously proposed topology of animal and fungal orthologues (100% bootstrap probability), while plant and D. discoideum CYP51s belong to the cycloartenol branch (photosynthetic lineage), all in accordance with biochemical data. Bacterial CYP51s cluster within the cycloartenol branch (69% bootstrap probability), which is indicative of a lateral gene transfer of a plant CYP51 to the methylococcal/mycobacterial progenitor, suggesting further that bacterial CYP51s are not the oldest CYP genes. Lateral gene transfer is likely far more important than hitherto thought in the development of the diversified CYP superfamily. Consequently, bacterial CYPs may represent a mixture of genes with prokaryotic and eukaryotic origin.

Bacteria↗

Population pharmacokinetic modeling and model validation of a spicamycin derivative, KRN5500, in phase 1 study.

PURPOSE: KRN5500, a novel spicamycin derivative, shows the greatest activity against a human tumor xenograft model and the highest therapeutic index among spicamycin derivatives. KRN5500 is currently under clinical development in Japan and the United States. The objective of this study was to develop a population pharmacokinetic model that describes the KRN5500 plasma concentration versus time data. METHODS: Data were collected from 18 patients entered in a phase 1 study. These patients received KRN5500 3-21 mg/m2 as a 2-h infusion. A total of 219 concentration measurements were available. The data were analyzed using the nonlinear mixed effect model (NONMEM) program. In addition, the basic and final population pharmacokinetic models were evaluated using bootstrapping resampling. RESULTS: The basic model selected was a two-compartment model with a combination of additive and constant coefficient of variation error models. The basic model fitted well not only the original data, but also 100 bootstrap replicates generated from the original data set. With regard to the effect of covariates selected by generalized additive modeling analysis, gender (SEX) and performance status were found to be possible determinants of the volume of central compartment by NONMEM analysis. The final regression model for V1 was V1 = theta V1 (1--SEX x theta SEX), where V1 is the typical population value of the volume of central compartment, and SEX = 0 if the patient is male, otherwise SEX = 1. The final model was fitted to the 200 bootstrapped samples. The mean parameter estimates were within 15% of those obtained with the original data set. CONCLUSIONS: The KRN5500 plasma concentration versus time data obtained from the phase 1 study were well described by the population pharmacokinetic model. Further evaluation by bootstrapping showed that the population pharmacokinetic model was stable.

Adult↗

Multivariable modeling of radiotherapy outcomes, including dose-volume and clinical factors.

PURPOSE: The probability of a specific radiotherapy outcome is typically a complex, unknown function of dosimetric and clinical factors. Current models are usually oversimplified. We describe alternative methods for building multivariable dose-response models. METHODS: Representative data sets of esophagitis and xerostomia are used. We use a logistic regression framework to approximate the treatment-response function. Bootstrap replications are performed to explore variable selection stability. To guard against under/overfitting, we compare several analytical and data-driven methods for model-order estimation. Spearman's coefficient is used to evaluate performance robustness. Novel graphical displays of variable cross correlations and bootstrap selection are demonstrated. RESULTS: Bootstrap variable selection techniques improve model building by reducing sample size effects and unveiling variable cross correlations. Inference by resampling and Bayesian approaches produced generally consistent guidance for model order estimation. The optimal esophagitis model consisted of 5 dosimetric/clinical variables. Although the xerostomia model could be improved by combining clinical and dose-volume factors, the improvement would be small. CONCLUSIONS: Prediction of treatment response can be improved by mixing clinical and dose-volume factors. Graphical tools can mitigate the inherent complexity of multivariable modeling. Bootstrap-based variable selection analysis increases the reliability of reported models. Statistical inference methods combined with Spearman's coefficient provide an efficient approach to estimating optimal model order.

Carcinoma, Non-Small-Cell Lung↗

A combined bootstrap/histogram analysis approach for computing a lateralization index from neuroimaging data.

Cerebral hemispheric specialization has traditionally been described using a lateralization index (LI). Such an index, however, shows a very severe threshold dependency and is prone to be influenced by statistical outliers. Reliability of this index thus has been inherently weak, and the assessment of this reliability is as yet not possible as methods to detect such outliers are not available. Here, we propose a new approach to calculating a lateralization index on functional magnetic resonance imaging data by combining a bootstrap procedure with a histogram analysis approach. Synthetic and real functional magnetic resonance imaging data was used to assess performance of our approach. Using a bootstrap algorithm, 10,000 indices are iteratively calculated at different thresholds, yielding a robust mean, maximum and minimum LI and thus allowing to attach a confidence interval to a given index. Taking thresholds into account, an overall weighted bootstrapped lateralization index is calculated. Additional histogram analyses of these bootstrapped values allow to judge reliability and the influence of outliers within the data. We conclude that the proposed methods yield a robust and specific lateralization index, sensitively detect outliers and allow to assess the underlying data quality.

Brain↗

Uncertainties in model-based outcome predictions for treatment planning.

PURPOSE: Model-based treatment-plan-specific outcome predictions (such as normal tissue complication probability [NTCP] or the relative reduction in salivary function) are typically presented without reference to underlying uncertainties. We provide a method to assess the reliability of treatment-plan-specific dose-volume outcome model predictions. METHODS AND MATERIALS: A practical method is proposed for evaluating model prediction based on the original input data together with bootstrap-based estimates of parameter uncertainties. The general framework is applicable to continuous variable predictions (e.g., prediction of long-term salivary function) and dichotomous variable predictions (e.g., tumor control probability [TCP] or NTCP). Using bootstrap resampling, a histogram of the likelihood of alternative parameter values is generated. For a given patient and treatment plan we generate a histogram of alternative model results by computing the model predicted outcome for each parameter set in the bootstrap list. Residual uncertainty ("noise") is accounted for by adding a random component to the computed outcome values. The residual noise distribution is estimated from the original fit between model predictions and patient data. RESULTS: The method is demonstrated using a continuous-endpoint model to predict long-term salivary function for head-and-neck cancer patients. Histograms represent the probabilities for the level of posttreatment salivary function based on the input clinical data, the salivary function model, and the three-dimensional dose distribution. For some patients there is significant uncertainty in the prediction of xerostomia, whereas for other patients the predictions are expected to be more reliable. In contrast, TCP and NTCP endpoints are dichotomous, and parameter uncertainties should be folded directly into the estimated probabilities, thereby improving the accuracy of the estimates. Using bootstrap parameter estimates, competing treatment plans can be ranked based on the probability that one plan is superior to another. Thus, reliability of plan ranking could also be assessed. CONCLUSIONS: A comprehensive framework for incorporating uncertainties into treatment-plan-specific outcome predictions is described. Uncertainty histograms for continuous variable endpoint models provide a straightforward method for visual review of the reliability of outcome predictions for each treatment plan.

Humans↗

External validation is necessary in prediction research: a clinical example.

BACKGROUND AND OBJECTIVES: Prediction models tend to perform better on data on which the model was constructed than on new data. This difference in performance is an indication of the optimism in the apparent performance in the derivation set. For internal model validation, bootstrapping methods are recommended to provide bias-corrected estimates of model performance. Results are often accepted without sufficient regard to the importance of external validation. This report illustrates the limitations of internal validation to determine generalizability of a diagnostic prediction model to future settings. METHODS: A prediction model for the presence of serious bacterial infections in children with fever without source was derived and validated internally using bootstrap resampling techniques. Subsequently, the model was validated externally. RESULTS: In the derivation set (n=376), nine predictors were identified. The apparent area under the receiver operating characteristic curve (95% confidence interval) of the model was 0.83 (0.78-0.87) and 0.76 (0.67-0.85) after bootstrap correction. In the validation set (n=179) the performance was 0.57 (0.47-0.67). CONCLUSION: For relatively small data sets, internal validation of prediction models by bootstrap techniques may not be sufficient and indicative for the model's performance in future patients. External validation is essential before implementing prediction models in clinical practice.

Bacterial Infections↗

Confidence intervals for the receiver operating characteristic area in studies with small samples.

RATIONALE AND OBJECTIVES: The authors performed this study to address two practical questions. First, how large does the sample size need to be for confidence intervals (CIs) based on the usual asymptotic methods to be appropriate? Second, when the sample size is smaller than this threshold, what alternative method of CI construction should be used? MATERIALS AND METHODS: The authors performed a Monte Carlo simulation study where 95% CIs were constructed for the receiver operating characteristic (ROC) area and for the difference between two ROC areas for rating and continuous test results--for ROC areas of moderate and high accuracy--by using both parametric and nonparametric estimation methods. Alternative methods evaluated included several bootstrap CIs and CIs with the Student t distribution. RESULTS: For the difference between two ROC areas, CIs based on the asymptotic theory provided adequate coverage even when the sample size was very small (20 patients). In contrast, for a single ROC area, the asymptotic methods do not provide adequate CI coverage for small samples; for ROC areas of high accuracy, the sample size must be large (more than 200 patients) for the asymptotic methods to be applicable. The recommended alternative (bootstrap percentile, bootstrap t, or bootstrap bias-corrected accelerated method) depends on the estimation approach, format of the test results, and ROC area. CONCLUSION: Currently, there is not a single best alternative for constructing CIs for a single ROC area for small samples.

Confidence Intervals↗

Single-pass attenuated total reflection Fourier transform infrared spectroscopy for the prediction of protein secondary structure.

Principal component regression (PCR) was applied to a spectral library of proteins in H2O solution acquired by single-pass attenuated total reflectance (ATR) Fourier transform infrared (FT-IR) spectroscopy. PCR was used to predict the secondary structure content, principally alpha-helical and the beta-sheet content, of proteins within a spectral library. Quantitation of protein secondary structure content was performed as a proof of principle that use of single-pass ATR-FT-IR is an appropriate method for protein secondary structure analysis. The ATR-FT-IR method permits acquisition of the entire spectral range from 700 to 3900 cm(-1) without significant interference from water bands. An "inside model space" bootstrap and a genetic algorithm (GA) were used to improve prediction results. Specifically, the bootstrap was utilized to increase the number of replicates for adequate training and validation of the PCR model. The GA was used to optimize PCR parameters, particularly wavenumber selection. The use of the bootstrap allowed for adequate representation of variability in the amide A, amide B, and C-H stretching regions due to differing levels of sample hydration. Implementation of the bootstrap improved the robustness of the PCR models significantly; however, the use of a GA only slightly improved prediction results. Two spectral libraries are presented where one was better suited for beta-sheet content prediction and the other for alpha-helix content prediction. The GA-optimized PCR method for alpha-helix content prediction utilized 120 wavenumbers within the amide I, II, A, B, and IV and the C-H stretching regions and 18 factors. For beta-sheet content predictions, 580 wavenumbers within the amide I, II, A, and B and the C-H stretching regions and 18 factors were used. The validation results using these two methods yielded an average absolute error of 1.7% for alpha-helix content prediction and an average absolute error of 2.3% for beta-sheet content prediction. After the PCR models were developed and validated, they were used to predict the alpha-helix and beta-sheet content of two unknowns, casein and immunoglobulin G.

Multivariate Analysis↗

V4 region of small subunit rDNA indicates polyphyly of the Fellodistomidae (Digenea) which is supported by morphology and life-cycle data.

There is no morphological synapomorphy for the disparate digeneans, the Fellodistomidae Nicoll, 1909. Although all known life-cycles of the group include bivalves as first intermediate hosts, there is no convincing morphological synapomorphy that can be used to unite the group. Sequences from the V4 region of small subunit (18S) rRNA genes were used to infer phylogenetic relationships among 13 species of Fellodistomidae from four subfamilies and eight species from seven other digenean families: Bivesiculidae; Brachylaimidae; Bucephalidae; Gorgoderidae; Gymnophallidae; Opecoelidae; and Zoogonidae. Outgroup comparison was made initially with an aspidogastrean. Various species from the other digenean families were used as outgroups in subsequent analyses. Three methods of analysis indicated polyphyly of the Fellodistomidae and at least two independent radiations of the subfamilies, such that they were more closely associated with other digeneans than to each other. The Tandanicolinae was monophyletic (100% bootstrap support) and was weakly associated with the Gymnophallidae (< 50-55% bootstrap support). Monophyly of the Baccigerinae was supported with 78-87% bootstrap support, and monophyly of the Zoogonidae + Baccigerinae received 77-86% support. The remaining fellodistomid species, Fellodistomum fellis, F. agnotum and Coomera brayi (Fellodistominae) plus Proctoeces maculatus and Complexobursa sp. (Proctoecinae), formed a separate clade with 74-92% bootstrap support. On the basis of molecular, morphological and life-cycle evidence, the subfamilies Baccigerinae and Tandanicolinae are removed from the Fellodistomidae and promoted to familial status. The Baccigerinae is promoted under the senior synonym Faustulidae Poche, 1926, and the Echinobrevicecinae Dronen, Blend & McEachran, 1994 is synonymised with the Faustulidae. Consequently, species that were formerly in the Fellodistomidae are now distributed in three families: Felldistomidae; Faustulidae (syn. Baccigerinae Yamaguti, 1954); and Tandanicolidae Johnston, 1927. We infer that the use of bivalves as intermediate hosts by this broad range of families indicates multiple host-switching events within the radiation of the Digenea.

Animals↗

In vitro dissolution profile comparison--statistics and analysis of the similarity factor, f2.

PURPOSE: To describe the properties of the similarity factor (f2) as a measure for assessing the similarity of two dissolution profiles. Discuss the statistical properties of the estimate based on sample means. METHODS: The f2 metrics and the decision rule is evaluated using examples of dissolution profiles. The confidence interval is calculated using bootstrapping method. The bias of the estimate using sample mean dissolution is evaluated. RESULTS: 1. f2 values were found to be sensitive to number of sample points, after the dissolution plateau has been reached. 2. The statistical evaluation of f2 could be made using 90% confidence interval approach. 3. The statistical distribution of f2 metrics could be simulated using 'Bootstrap' method. A relatively robust distribution could be obtained after more than 500 'Bootstraps'. 4. A statistical 'bias correction' was found to reduce the bias. CONCLUSIONS: The similarity factor f2 is a simple measure for the comparison of two dissolution profiles. But the commonly used similarity factor estimate f2 is a biased and conservative estimate of f2. The bootstrap approach is a useful tool to simulate the confidence interval.

Chemistry, Pharmaceutical↗

Overcredibility of molecular phylogenies obtained by Bayesian phylogenetics.

Bayesian phylogenetics has recently been proposed as a powerful method for inferring molecular phylogenies, and it has been reported that the mammalian and some plant phylogenies were resolved by using this method. The statistical confidence of interior branches as judged by posterior probabilities in Bayesian analysis is generally higher than that as judged by bootstrap probabilities in maximum likelihood analysis, and this difference has been interpreted as an indication that bootstrap support may be too conservative. However, it is possible that the posterior probabilities are too high or too liberal instead. Here, we show by computer simulation that posterior probabilities in Bayesian analysis can be excessively liberal when concatenated gene sequences are used, whereas bootstrap probabilities in neighbor-joining and maximum likelihood analyses are generally slightly conservative. These results indicate that bootstrap probabilities are more suitable for assessing the reliability of phylogenetic trees than posterior probabilities and that the mammalian and plant phylogenies may not have been fully resolved.

Amino Acid Substitution↗

Phylogenetic supermatrix analysis of GenBank sequences from 2228 papilionoid legumes.

A comprehensive phylogeny of papilionoid legumes was inferred from sequences of 2228 taxa in GenBank release 147. A semiautomated analysis pipeline was constructed to download, parse, assemble, align, combine, and build trees from a pool of 11,881 sequences. Initial steps included all-against-all BLAST similarity searches coupled with assembly, using a novel strategy for building length-homogeneous primary sequence clusters. This was followed by a combination of global and local alignment protocols to build larger secondary clusters of locally aligned sequences, thus taking into account the dramatic differences in length of the heterogeneous coding and noncoding sequence data present in GenBank. Next, clusters were checked for the presence of duplicate genes and other potentially misleading sequences and examined for combinability with other clusters on the basis of taxon overlap. Finally, two supermatrices were constructed: a "sparse" matrix based on the primary clusters alone (1794 taxa x 53,977 characters), and a somewhat more "dense" matrix based on the secondary clusters (2228 taxa x 33,168 characters). Both matrices were very sparse, with 95% of their cells containing gaps or question marks. These were subjected to extensive heuristic parsimony analyses using deterministic and stochastic heuristics, including bootstrap analyses. A "reduced consensus" bootstrap analysis was also performed to detect cryptic signal in a subtree of the data set corresponding to a "backbone" phylogeny proposed in previous studies. Overall, the dense supermatrix appeared to provide much more satisfying results, indicated by better resolution of the bootstrap tree, excellent agreement with the backbone papilionoid tree in the reduced bootstrap consensus analysis, few problematic large polytomies in the strict consensus, and less fragmentation of conventionally recognized genera. Nevertheless, at lower taxonomic levels several problems were identified and diagnosed. A large number of methodological issues in supermatrix construction at this scale are discussed, including detection of annotation errors in GenBank sequences; the shortage of effective algorithms and software for local multiple sequence alignment; the difficulty of overcoming effects of fragmentation of data into nearly disjoint blocks in sparse supermatrices; and the lack of informative tools to assess confidence limits in very large trees.

Algorithms↗

Superior feature-set ranking for small samples using bolstered error estimation.

MOTIVATION: Ranking feature sets is a key issue for classification, for instance, phenotype classification based on gene expression. Since ranking is often based on error estimation, and error estimators suffer to differing degrees of imprecision in small-sample settings, it is important to choose a computationally feasible error estimator that yields good feature-set ranking. RESULTS: This paper examines the feature-ranking performance of several kinds of error estimators: resubstitution, cross-validation, bootstrap and bolstered error estimation. It does so for three classification rules: linear discriminant analysis, three-nearest-neighbor classification and classification trees. Two measures of performance are considered. One counts the number of the truly best feature sets appearing among the best feature sets discovered by the error estimator and the other computes the mean absolute error between the top ranks of the truly best feature sets and their ranks as given by the error estimator. Our results indicate that bolstering is superior to bootstrap, and bootstrap is better than cross-validation, for discovering top-performing feature sets for classification when using small samples. A key issue is that bolstered error estimation is tens of times faster than bootstrap, and faster than cross-validation, and is therefore feasible for feature-set ranking when the number of feature sets is extremely large.

Algorithms↗

Multivariate time-to-event models for studies of recurrent childhood diseases.

BACKGROUND: Standard Poisson regression analyses of infectious disease incidence may not be optimal when covariates change over time. When a longitudinal study is of sufficient duration so that multiple disease episodes per person may occur, within-subject correlation may be present and require special statistical consideration. Seasonality of disease incidence is a common feature of field studies of respiratory and diarrhoeal diseases in children. Accurate analysis may require keeping the children aligned with respect to the same baseline hazard function. These data features often are ignored in the analysis of such studies. METHODS: Methods for accounting for multiple events are discussed. We propose the use of a counting process model to retain the alignment with respect to season. This model is supplemented by a simple bootstrap procedure to account for the within-child correlation. RESULTS: The bootstrap technique is illustrated with data from a randomized trial of the effects of vitamin A supplementation on childhood morbidity. Standard errors for the main effect of vitamin A are larger using the bootstrap, indicating the presence of positive correlation. CONCLUSIONS: Time-to-event models can be useful in the analysis of studies of childhood morbidity, especially for situations with seasonality, waning of effect, or multiple events. The bootstrap provides an appropriate, straightforward method for handling within-child correlation in such settings.

Biometry↗

Tie trees generated by distance methods of phylogenetic reconstruction.

In examining genetic data in recent publications, Backeljau et al. showed cases in which two or more different trees (tie trees) were constructed from a single data set for the neighbor-joining (NJ) method and the unweighted pair group method with arithmetic mean (UPGMA). However, it is still unclear how often and under what conditions tie trees are generated. Therefore, I examined these problems by computer simulation. Examination of cases in which tie trees occur shows that tie trees can appear when no substitutions occur along some interior branch(es) on a tree. However, even when some substitutions occur along interior branches, tie trees can appear by chance if parallel or backward substitutions occur at some sites. The simulation results showed that tie trees occur relatively frequently for sequences with low divergence levels or with small numbers of sites. For such data, UPGMA sometimes produced tie trees quite frequently, whereas tie trees for the NJ method were generally rare. In the simulation, bootstrap values for clusters (tie clusters) that differed among tie trees were mostly low (< 60%). With a small probability, relatively high bootstrap values (at most 70%-80%) appeared for tie clusters. The bias of the bootstrap values caused by an input order of sequence can be avoided if one of the different paths in the cycles of making an NJ or UPGMA tree is chosen at random in each bootstrap replication.

Computer Simulation↗

Confidence intervals for the mean of diagnostic test charge data containing zeros.

In this paper, we consider the problem of interval estimation for the mean of diagnostic test charges. Diagnostic test charge data may contain zero values, and the nonzero values can often be modeled by a log-normal distribution. Under such a model, we propose three different interval estimation procedures: a percentile-t bootstrap interval based on sufficient statistics and two likelihood-based confidence intervals. For theoretical properties, we show that the two likelihood-based one-sided confidence intervals are only first-order accurate and that the bootstrap-based one-sided confidence interval is second-order accurate. For two-sided confidence intervals, all three proposed methods are second-order accurate. A simulation study in finite-sample sizes suggests all three proposed intervals outperform a widely used minimum variance unbiased estimator (MVUE)-based interval except for the case of one-sided lower end-point intervals when the skewness is very small. Among the proposed one-sided intervals, the bootstrap interval has the best coverage accuracy. For the two-sided intervals, when the sample size is small, the bootstrap method still yields the best coverage accuracy unless the skewness is very small, in which case the bias-corrected ML method has the best accuracy. When the sample size is large, all three proposed intervals have similar coverage accuracy. Finally, we analyze with the proposed methods one real example assessing diagnostic test charges among older adults with depression.

Biometry↗

Estimating genetic correlations in natural populations in the absence of pedigree information: accuracy and precision of the Lynch method.

Usually, genetic correlations are estimated from breeding designs in the laboratory or greenhouse. However, estimates of the genetic correlation for natural populations are lacking, mostly because pedigrees of wild individuals are rarely known. Recently Lynch (1999) proposed a formula to estimate the genetic correlation in the absence of data on pedigree. This method has been shown to be particularly accurate provided a large sample size and a minimum (20%) proportion of relatives. Lynch (1999) proposed the use of the bootstrap to estimate standard errors associated with genetic correlations, but did not test the reliability of such a method. We tested the bootstrap and showed the jackknife can provide valid estimates of the genetic correlation calculated with the Lynch formula. The occurrence of undefined estimates, combined with the high number of replicates involved in the bootstrap, means there is a high probability of obtaining a biased upward, incomplete bootstrap, even when there is a high fraction of related pairs in a sample. It is easier to obtain complete jackknife estimates for which all the pseudovalues have been defined. We therefore recommend the use of the jackknife to estimate the genetic correlation with the Lynch formula. Provided data can be collected for more than two individuals at each location, we propose a group sampling method that produces low standard errors associated with the jackknife, even when there is a low fraction of relatives in a sample.

Analysis of Variance↗

Phylogenetic analyses of Chlamydia psittaci strains from birds based on 16S rRNA gene sequence.

The nucleotide sequences of 16S ribosomal DNA (rDNA) were determined for 39 strains of Chlamydia psittaci (34 from birds and 5 from mammals) and for 4 Chlamydia pecorum strains. The sequences were compared phylogenetically with the gene sequences of nine Chlamydia strains (covering four species of the genus) retrieved from nucleotide databases. In the neighbor-joining tree, C. psittaci strains were more closely related to each other than to the other Chlamydia species, although a feline pneumonitis strain was distinct (983 to 98.6% similarity to other strains) and appeared to form the deepest subline within the species of C. psittaci (bootstrap value, 99%). The other strains of C. psittaci exhibiting similarity values of more than 99% were branched into several subgroups. Two pigeon strains and one turkey strain formed a distinct clade recovered in 97% of the bootstrapped trees. The other pigeon strains seemed to be distinct from the strains from psittacine birds, with 88% of bootstrap value. In the cluster of psittacine strains, three parakeet strains and an ovine abortion strain exhibited a specific association (level of sequence similarity, 99.9% or more; bootstrap value, 95%). These suggest that at least four groups of strains exist within the species C. psittaci. The 16S rDNA sequence is a valuable phylogenetic marker for the taxonomy of chlamydiae, and its analysis is a reliable tool for identification of the organisms.

Abortion, Veterinary↗