PubMed Health⌕ Search

Biomedical subjects

D C Hoyle

Publications and source records attributed to D C Hoyle.

4 recordsLinked to original sources

Principal-component-analysis eigenvalue spectra from data with symmetry-breaking structure.

Principal component analysis (PCA) is a ubiquitous method of multivariate statistics that focuses on the eigenvalues lambda and eigenvectors of the sample covariance matrix of a data set. We consider p, N-dimensional data vectors xi drawn from a distribution with covariance matrix C. We use the replica method to evaluate the expected eigenvalue distribution rho(lambda) as N--> infinity with p=alphaN for some fixed alpha. In contrast to existing studies we consider the case where C contains a number of symmetry-breaking directions, so that the sample data set contains some definite structure. Explicitly we set C=sigma2I+sigma(2)Sigma(S)(m=1)A(m)B(m)B(T)(m), with A(m)>0 for all m. We find that the bulk of the eigenvalues are distributed as for the case when the elements of xi are independent and identically distributed. With increasing alpha a series of phase transitions are observed, at alpha=A(-2)(m), m=1,2,..., S, each time a single delta function, delta(lambda-lambda(u)(A(m))), separates from the upper edge of the bulk distribution, where lambda(u)(A)=sigma(2)[1+A][1+(alphaA)(-1)]. We confirm the results of the replica analysis by studying the Stieltjes transform of rho(lambda). This suggests that the results obtained from the replica analysis are universal, irrespective of the distribution from which xi is drawn, provided the fourth moment of each element of xi exists.

Fourier Analysis↗

Transcriptome profiling of a Saccharomyces cerevisiae mutant with a constitutively activated Ras/cAMP pathway.

Often changes in gene expression levels have been considered significant only when above/below some arbitrarily chosen threshold. We investigated the effect of applying a purely statistical approach to microarray analysis and demonstrated that small changes in gene expression have biological significance. Whole genome microarray analysis of a pde2Delta mutant, constructed in the Saccharomyces cerevisiae reference strain FY23, revealed altered expression of approximately 11% of protein encoding genes. The mutant, characterized by constitutive activation of the Ras/cAMP pathway, has increased sensitivity to stress, reduced ability to assimilate nonfermentable carbon sources, and some cell wall integrity defects. Applying the Munich Information Centre for Protein Sequences (MIPS) functional categories revealed increased expression of genes related to ribosome biogenesis and downregulation of genes in the cell rescue, defense, cell death and aging category, suggesting a decreased response to stress conditions. A reduced level of gene expression in the unfolded protein response pathway (UPR) was observed. Cell wall genes whose expression was affected by this mutation were also identified. Several of the cAMP-responsive orphan genes, upon further investigation, revealed cell wall functions; others had previously unidentified phenotypes assigned to them. This investigation provides a statistical global transcriptome analysis of the cellular response to constitutive activation of the Ras/cAMP pathway.

Cell Wall↗

Factors affecting the errors in the estimation of evolutionary distances between sequences.

Phylogenetic methods that use matrices of pairwise distances between sequences (e.g., neighbor joining) will only give accurate results when the initial estimates of the pairwise distances are accurate. For many different models of sequence evolution, analytical formulae are known that give estimates of the distance between two sequences as a function of the observed numbers of substitutions of various classes. These are often of a form that we call "log transform formulae". Errors in these distance estimates become larger as the time t since divergence of the two sequences increases. For long times, the log transform formulae can sometimes give divergent distance estimates when applied to finite sequences. We show that these errors become significant when t approximately 1/2 |lambda(max)|(-1) logN, where lambda(max) is the eigenvalue of the substitution rate matrix with the largest absolute value and N is the sequence length. Various likelihood-based methods have been proposed to estimate the values of parameters in rate matrices. If rate matrix parameters are known with reasonable accuracy, it is possible to use the maximum likelihood method to estimate evolutionary distances while keeping the rate parameters fixed. We show that errors in distances estimated in this way only become significant when t approximately 1/2 |lambda(1)|(-1) logN, where lambda(1) is the eigenvalue of the substitution rate matrix with the smallest nonzero absolute value. The accuracy of likelihood-based distance estimates is therefore much higher than those based on log transform formulae, particularly in cases where there is a large range of timescales involved in the rate matrix (e.g., when the ratio of transition to transversion rates is large). We discuss several practical ways of estimating the rate matrix parameters before distance calculation and hence of increasing the accuracy of distance estimates.

Evolution, Molecular↗

RNA sequence evolution with secondary structure constraints: comparison of substitution rate models using maximum-likelihood methods.

We test models for the evolution of helical regions of RNA sequences, where the base pairing constraint leads to correlated compensatory substitutions occurring on either side of the pair. These models are of three types: 6-state models include only the four Watson-Crick pairs plus GU and UG; 7-state models include a single mismatch state that combines all of the 10 possible mismatches; 16-state models treat all mismatch states separately. We analyzed a set of eubacterial ribosomal RNA sequences with a well-established phylogenetic tree structure. For each model, the maximum-likelihood values of the parameters were obtained. The models were compared using the Akaike information criterion, the likelihood-ratio test, and Cox's test. With a high significance level, models that permit a nonzero rate of double substitutions performed better than those that assume zero double substitution rate. Some models assume symmetry between GC and CG, between AU and UA, and between GU and UG. Models that relaxed this symmetry assumption performed slightly better, but the tests did not all agree on the significance level. The most general time-reversible model significantly outperformed any of the simplifications. We consider the relative merits of all these models for molecular phylogenetics.

Base Pairing↗