PubMed HealthSearch

Biomedical subjects

B E Blaisdell

Publications and source records attributed to B E Blaisdell.

17 recordsLinked to original sources

Methods and algorithms for statistical analysis of protein sequences.

We describe several protein sequence statistics designed to evaluate distinctive attributes of residue content and arrangement in primary structure. Considered are global compositional biases, local clustering of different residue types (e.g., charged residues, hydrophobic residues, Ser/Thr), long runs of charged or uncharged residues, periodic patterns, counts and distribution of homooligopeptides, and unusual spacings between particular residue types. The computer program SAPS (statistical analysis of protein sequences) calculates all the statistics for any individual protein sequence input and is available for the UNIX environment through electronic mail on request to V.B. (volker/genomic@stanford.edu).

Algorithms

Quantile distributions of amino acid usage in protein classes.

A comparative study of the compositional properties of various protein sets from both cellular and viral organisms is presented. Invariants and contrasts of amino acid usages have been discerned for different protein function classes and for different species using robust statistical methods based on quantile distributions and stochastic ordering relationships. In addition, a quantitative criterion to assess amino acid compositional extremes relative to a reference protein set is proposed and applied. Invariants of amino acid usage relate mainly to the central range of quantile distributions, whereas contrasts occur mainly in the tails of the distributions, especially contrasts between eukaryote and prokaryote species. Influences from genomic constraint are evident, for example, in the arginine:lysine ratios and the usage frequencies of residues encoded by G + C-rich versus A + T-rich codon types. The structurally similar amino acids, glutamate versus aspartate and phenylalanine versus tyrosine, show stochastic dominance relationships for most species protein sets favoring glutamate and phenylalanine respectively. The quantile distribution of hydrophobic amino acid usages in prokaryote data dominates the corresponding quantile distribution in human data. In contrast, glutamate, cysteine, proline and serine usages in human proteins dominate the corresponding quantile distributions in Escherichia coli. E. coli dominates human in the use of basic residues, but no dominance ordering applies to acidic residues. The discussion centers on commonalities and anomalies of the amino acid compositional spectrum in relation to species, function, cellular localization, biochemical and steric attributes, complexity of the amino acid biosynthetic pathway, amino acid relative abundances and founder effects.

Amino Acids

An efficient algorithm for identifying matches with errors in multiple long molecular sequences.

An efficient algorithm is described for finding matches, repeats and other word relations, allowing for errors, in large data sets of long molecular sequences. The algorithm entails hashing on fixed-size words in conjunction with the use of a linked list connecting all occurrences of the same word. The average memory and run time requirement both increase almost linearly with the total sequence length. Some results of the program's performance on a database of Escherichia coli DNA sequences are presented.

Algorithms

Very long charge runs in systemic lupus erythematosus-associated autoantigens.

Systemic lupus erythematosus and other chronic systemic autoimmune diseases are associated with circulating autoantibodies reactive with a limited set of mostly nuclear proteins. Using rigorous statistical methods we have identified segments of highly significant charge concentration in the majority of the characteristic nuclear and cytoplasmic autoantigens. Extremely long runs of charged residues, including some sequences of greater than 20 consecutive charged residues (purely acidic or mixed basic and acidic), occur in about a third of these proteins, whereas equivalent runs are found in less than 3% of other mammalian proteins. The other sequences have less extreme charge clusters, the type and location of which are often conserved between several otherwise nonsimilar antigens. We propose that supercharged surfaces render the targeted host proteins strongly immunogenic and that antinuclear antibody profiles might result from chronic exposure to intracellular contents, possibly in conjunction with crossreactive viral products. The limited number of potential systemic autoantigens may partly be due to the rarity of requisite charge properties.

Amino Acid Sequence

Average values of a dissimilarity measure not requiring sequence alignment are twice the averages of conventional mismatch counts requiring sequence alignment for a variety of computer-generated model systems.

A measure of sequence similarity, dt, not requiring prior sequence alignment gave correct results for a variety of computer-generated model sequences without and with gaps for all degrees of substitution, s. Measure d was the squared Euclidean distance between vectors of counts of t-tuplets of characters in the two sequences. In models without gaps and without Needleman-Wunsch alignment, average d was very closely equal to twice average conventional mismatch counts, m. In these models one of each of the conditions on the Jukes-Cantor model was violated in turn: (1) both descendant lineages receive the same number of substitutions, (2) all sites are equally likely to be substituted, (3) all different replacement characters are equally likely to be chosen, and (4) all original characters are equally likely to be substituted. In Jukes-Cantor models with gaps Needleman-Wunsch alignment was necessarily performed, a procedure that generally produced incorrect values of m. For these models average d was found to be very closely equal to twice the average m estimated from the known value of s using the inverted Jukes-Cantor formula.

Algorithms

Evidence for selective evolution in codon usage in conserved amino acid segments of human alphaherpesvirus proteins.

The genomes of human viruses herpes simplex 1 (HSV1) and varicella zoster (VZV), although similar in biology, largely concordant in gene order, and identical in many amino acid segments, differ widely in their genomic G + C (abbreviated S) content, which is high in HSV1 (68%) and low in VZV (46%). This paper analyzes several striking codon usage contrasts. The S difference in coding regions is dramatically large in codon site 3, S3, about 42%. The large difference in S3 is maintained at the same level in a subset of closely similar genes and even in corresponding identical amino acid blocks. A similar difference in S levels in silent site 1 (S1) is found in leucine and arginine. The difference in S3 levels occurs in every gene and in every multicodon amino acid form. The S difference also exists in amino acid usage, with HSV1 using significantly more codon types SSN, while VZV uses more codon types WWN (where W stands for A or T). The nonoverlapping and narrow histograms of S3 gene frequencies in both viruses suggest that the difference has arisen and been maintained by a process of selective rather than nonselective effects. This is in sharp contrast to the relatively large variance seen for highly similar genes in the human versus yeast analysis. Interpretations and hypotheses to explain the HSV1 vs VZV codon usage disparity relate to virus-host interactions, to the role of viral genes in DNA metabolism, to availability of molecular resources (molecular Gause exclusion principle), and to differences in genomic structure.

Amino Acids

Contrasts in codon usage of latent versus productive genes of Epstein-Barr virus: data and hypotheses.

Epstein-Barr virus (EBV) has two different modes of existence: latent and productive. There are eight known genes expressed during latency (and hardly at all during the productive phase) and about 70 other ("productive") genes. It is shown that the EBV genes known to be expressed during latency display codon usage strikingly different from that of genes that are expressed during lytic growth. In particular, the percentage of S3 (G or C in codon site 3) is persistently lower (about 20%) in all latent genes than in nonlatent genes. Moreover, S3 is lower in each multicodon amino acid form. Also, the percentage of S in silent codon sites 1 of leucine and arginine is lower in latent than in nonlatent genes. The largest absolute differences in amino acid usage between latent and nonlatent genes emphasize codon types SSN and WWN (W means nucleotide A or T and N is any nucleotide). Two principal explanations to account for the EBV latent versus productive gene codon disparity are proposed. Latent genes have codon usage substantially different from that of host cell genes to minimize the deleterious consequences to the host of viral gene expression during latency. (Productive genes are not so constrained.) It is also proposed that the latency genes of EBV were acquired recently by the viral genome. Evidence and arguments for these proposals are presented.

Amino Acid Sequence

A method to identify distinctive charge configurations in protein sequences, with application to human herpesvirus polypeptides.

Charge interactions are of great importance for protein function and structure, and for a variety of cellular and biochemical processes. We present a systematic approach to the detection of distinctive clusters, runs and periodic patterns of charged residues in a protein sequence. Criteria and formulae are set forth to assess statistical significance of these charge configurations. For the 80-odd proteins potentially encoded by the Epstein-Barr virus, only the major nuclear antigens of the latent state and the transactivator of the lytic cycle contain separated charge clusters of opposite sign as well as periodic charge patterns. From our studies of the polypeptides of the human herpesviruses and of a broad collection of human and other viral protein sequences, distinctive charge configurations appear to be associated with viral capsid and core proteins (positive clusters or runs, mostly at the carboxyl terminus), with many viral glycoproteins and membrane-associated proteins (negative charge clusters), and with transactivators and transforming proteins (multiple charge structures). The statistics developed in this paper apply more generally to other than charge properties of a protein and should aid in the evaluation of a large variety of sequence features.

Cytomegalovirus

Effectiveness of measures requiring and not requiring prior sequence alignment for estimating the dissimilarity of natural sequences.

Various measures of sequence dissimilarity have been evaluated by how well the additive least squares estimation of edges (branch lengths) of an unrooted evolutionary tree fit the observed pairwise dissimilarity measures and by how consistent the trees are for different data sets derived from the same set of sequences. This evaluation provided sensitive discrimination among dissimilarity measures and among possible trees. Dissimilarity measures not requiring prior sequence alignment did about as well as did the traditional mismatch counts requiring prior sequence alignment. Application of Jukes-Cantor correction to singlet mismatch counts worsened the results. Measures not requiring alignment had the advantage of being applicable to sequences too different to be critically alignable. Two different measures of pairwise dissimilarity not requiring alignment have been used: (1) multiplet distribution distance (MDD), the square of the Euclidean distance between vectors of the fractions of base signlets (or doublets, or triplets, or ...) in the respective sequences, and (2) complements of long words (CLW), the count of bases not occurring in significantly long common words. MDD was applicable to sequences more different than was CLW (noncoding), but the latter often gave better results where both measures were available (coding). MDD results were improved by using longer mutliplets and, if the sequences were coding, by using the larger amino acid and codon alphabets rather than the nucleotide alphabet. The additive least squares method could be used to provide a reasonable consensus of different trees for the same set of species (or related genes).

Animals

Average values of a dissimilarity measure not requiring sequence alignment are twice the averages of conventional mismatch counts requiring sequence alignment for a computer-generated model system.

Three measures of sequence dissimilarity have been compared on a computer-generated model system in which substitutions in random sequences were made at randomly selected sites and the replacement character was chosen at random from the set of characters different from the original occupant of the site. The three measures were the conventional mismatch count between aligned sequences (AMC = m) and two measures not requiring prior sequence alignment. The latter two measures were the squared Euclidean distance between vectors of counts of t-tuples (t = 1-6) of characters in the two sequences (multiplet distribution distances or MDD = d) and counts of characters not covered by word structures of statistically significant length common to the two sequences (common long words or CLW = SIB, SIS, or SAB). Average MDD distances were found to be two times average mismatch counts in the simulated sequences for all values of t from 1 to 6 and all degrees of substitution from one per sequence to so many as to produce, effectively, random sequences. This simple relation held independently of sequence length and of sequence composition. The relation was confirmed by exact results on small model systems and by formal asymptotic results in the limit of so few substitutions that no double hits occur and in the limit of two random sequences. The coefficient of variation for MDD distances was greater than that for mismatch counts for singlets but both measures approached the same low value for sextets. Needleman-Wunsch alignment produced incorrect mismatch counts at higher degrees of substitution. The model satisfied the conditions for the derivation of the Jukes-Cantor asymptotic adjustment, but its application produced increasingly bad results with increasing degrees of substitution in accord with earlier results on model and natural sequences. This fact was a consequence of the increase with increasing degrees of substitution of the sensitivity of the adjustment to error in the observations. Average CLW distances for a variety of common word structures were more or less parallel to MDD distances for appropriately long t-tuples. These results on model systems supported the validity of the two dissimilarity measures not requiring sequence alignment that was found in earlier work on natural sequences (Blaisdell 1989).

Base Sequence

Distinctive charge configurations in proteins of the Epstein-Barr virus and possible functions.

The protein products of several open reading frames (ORFs) of the Epstein-Barr virus (EBV) are remarkable in their distribution of charged residues. The nuclear antigen proteins EBNA1-EBNA4 of the EBV latent state contain separate significant clusters of charge of each sign. They (excepting EBNA4) also feature distinctive periodic charge patterns [e.g., (+, O)8, (O, -, -)7] and significant tandem repeats. None of the other ORFs (about 80) of the genome possess the conjunction of these properties. Only the protein encoded from BMLF1, the first immediate early transactivator protein, contains significant multiple charge clusters and periodic charge patterns. All proteins that contain significant repeats also contain at least one significant charge cluster of a single sign. These include EBNA5 and LYDMA produced during latency and BZLF1, whose expression terminates latency and initiates productive growth. It is reasonable to conclude that these aggregate significant charge configurations and repeats are important functionally for the latent existence and for the initiation of the lytic cycle and may be characteristic of these conditions. We discuss how large multimeric protein structures bound together by clusters of unlike charge may provide a mechanism for regulation of the expression of these proteins.

Amino Acid Sequence

A model for the development of the tandem repeat units in the EBV ori-P region and a discussion of their possible function.

This paper presents an analysis of the repeat units of the ori-P region of the Epstein-Barr virus (EBV) genome. These repeat units are well-conserved palindromes. The pattern of these repeats, their lengths, phases, and the distribution of the relatively few substitutions are explained by a scenario that gives a reasonable course for the evolutionary development of the pattern. The scenario suggests a model for the production of an initiating 3/2 palindrome from a moderately lengthy sequence. The palindromic units are then multiplied in judicious combinations by mechanisms of unequal crossing-over events associated with some point substitutions and a few instances of slippage replication. The potential secondary structures of the two separated tandem palindromic repeat regions in ori-P are contrasted. Possible modes of binding of Epstein-Barr nuclear antigen (EBNA) 1 protein to these hairpins are discussed. A number of possibilities for the origin and development of the ori-P region in relation to viral and cellular function are considered.

Base Sequence

A measure of the similarity of sets of sequences not requiring sequence alignment.

Determination of first- and second-order Markov chain homogeneity of sets of nuclear eukaryotic DNA sequences, both coding and noncoding, finds similarities imperceptible to the standard Needleman-Wunsch base matching or dot-matrix algorithms. These measures of the similarities of the distributions of adjacent pairs or triplets are in agreement with accepted evolutionary-tree topologies. Hierarchical clustering of the distributions of doublets of 30 miscellaneous coding sequences gives clusters in reasonable agreement with accepted biological classifications. In addition to similarity by homology, there is also observed similarity of disparate genes in the same organism--for example, all three disparate yeast genes (two enzymes and actin) form a well-distinguished cluster.

Animals

A method of estimating from two aligned present-day DNA sequences their ancestral composition and subsequent rates of substitution, possibly different in the two lineages, corrected for multiple and parallel substitutions at the same site.

The course of evolutionary change in DNA sequences has been modeled as a Markov process. The Markov process was represented by discrete time matrix methods. The parameters of the Markov transition matrices were estimated by least-squares direct-search optimization of the fit of the calculated divergence matrix to that observed for two aligned sequences. The Markov process corrected for multiple and parallel substitutions of bases at the same site. The method avoided the incorrect assumption of all previously described methods that the divergence between two present-day sequences is twice the divergence of either from the common and unknown ancestral sequence. The three previous methods were shown to be equivalent. The present method also avoided the undesirable assumptions that sequence composition has not changed with time and that the substitution rates in the two descendant lineages were the same. It permitted simultaneous estimation of ancestral sequence composition and, if applicable, of different substitution rates for the two descendant lineages, provided the total number of estimated parameters was less than 16. Properties of the Markov chain were discussed. It was proved for symmetric substitution matrices that all elements of the equilibrium divergence matrix equal 1/16, and that the total difference in the divergence matrix at epoch k equals the total change in the common substitution matrix at epoch 2k for all values of k. It was shown how to resolve an ambiguity in the assignment of two different substitution rates to the two descendant lineages when four or more similar sequences are available. The method was applied to the divergence matrix for codon site 3 for the mouse and rabbit beta-globins. This observed divergence matrix was significantly asymmetric and required at least two different substitution rates. This result could be achieved only by using different asymmetric substitution matrices for the two lineages.

Animals

Effect of dietary ascorbic acid on the incidence of spontaneous mammary tumors in RIII mice.

A study of the effect of different amounts of L-ascorbic acid (vitamin C), between 0.076% and 8.3%, contained in the food has been carried out with ten groups of RIII mice (seven ascorbic acid and three control groups), with 50 mice in each group. With an increase in the amount of ascorbic acid there is a highly significant decrease in the first-order rate constant for appearance of the first spontaneous mammary tumor after the lag time to detection by palpation. There is also an increase in the lag time. The mean body weight and mean food intake were not significantly different for the seven ascorbic acid groups. Striking differences were observed between the 0.076% ascorbic acid and the control groups (which synthesize the vitamin): smaller food intake, decreased lag time, and increased rate constant of appearance of the first mammary tumor. This comparison cannot be made experimentally for guinea pigs and primates because the control groups would develop scurvy.

Animals

Automated metabolic profiling of organic acids in human urine. II. Analysis of urine samples from "healthy" adults, sick children, and children with neuroblastoma.

Normalized median, minimum, and maximum values (analytical concentration factors) are given for 134 organic acids in urine of nine adult control subjects, five juvenile control subjects, and five children with neuroblastoma. The organic acids, separated by anion-exchange chromatography, were analyzed by a gas chromatograph-mass spectrometer-computer system. Sixty substances in this fraction are positively identified, and, of these, mean absolute concentrations are listed for 20. An additional 81 substances, sought but not found by this method, and 16 other substances found in a subset of these urines by another analytical method, are also listed. Measured retention indices on 5% OV-17 and a selected discriminating ion are given for each of the total of 231 compounds. Results are compared for the three groups of subjects, and the value of normalizing the data is discussed.

Adolescent