PubMed Health⌕ Search

Biomedical subjects

T L Bailey

Publications and source records attributed to T L Bailey.

28 records · Page 2Linked to original sources

Score distributions for simultaneous matching to multiple motifs.

Several computer algorithms now exist for discovering multiple motifs (expressed as weight matrices) that characterize a family of protein sequences known to be homologous. This paper describes a method for performing similarity searches of protein sequence databases using such a group of motifs. By simultaneously using all the motifs that characterize a protein family, the sensitivity and specificity of the database search are increased. We define the p-value for a target sequence to be the probability of a random sequence of the same length scoring as well or better in comparison to all the motifs that characterize the family. (The p-value of a database search can be determined from this value and the size of the database.) We show that estimating the distribution of single motif scores by a Gaussian extreme value distribution is insufficiently accurate to provide a useful estimate of the p-value, but that this deficiency can be corrected by reestimating the parameters of the underlying Gaussian distribution from observed scores for comparison of a given motif and sequence database. These parameters are used to calculate a "reduced variate" which has a Gumbel limiting distribution. Multiple motif scores are combined to give a single p-value by using the sum of the reduced variates for the motif scores as the test statistic. We give a computationally efficient approximation to the distribution of the sum of independent Gumbel random variables and verify experimentally that it closely approximates the distribution of the test statistic. Experiments on pseudorandom sequences show that the approximated p-values are conservative, so the significance of high scores in database searches will not be overstated. Experiments with real protein sequences and motifs identified by the MEME algorithm show that determining an overall p-value based on the combination of multiple motifs gives significantly better database search results than using p-values of single motifs.

Models, Chemical↗

Meta-MEME: motif-based hidden Markov models of protein families.

MOTIVATION: Modeling families of related biological sequences using Hidden Markov models (HMMs), although increasingly widespread, faces at least one major problem: because of the complexity of these mathematical models, they require a relatively large training set in order to accurately recognize a given family. For families in which there are few known sequences, a standard linear HMM contains too many parameters to be trained adequately. RESULTS: This work attempts to solve that problem by generating smaller HMMs which precisely model only the conserved regions of the family. These HMMs are constructed from motif models generated by the EM algorithm using the MEME software. Because motif-based HMMs have relatively few parameters, they can be trained using smaller data sets. Studies of short chain alcohol dehydrogenases and 4Fe-4S ferredoxins support the claim that motif-based HMMs exhibit increased sensitivity and selectivity in database searches, especially when training sets contain few sequences.

Alcohol Dehydrogenase↗

Testicular shape and its relationship to sperm production in mature Holstein bulls.

This study was conducted to determine the relationship between testicular shape, scrotal circumference (SC) and sperm production. Twenty-seven mature Holstein bulls were evaluated subjectively and objectively for testicular shape as indicated by testicular length and width, then placed in 1 of 3 groups. Group 1 contained 17 bulls with a normal ovoid testicular shape and a length to width ratio of 1.61:1 +/- 0.01 (SEM). Group 2 was composed of 4 bulls with a long, slender testicular shape and a length to width ratio of 1.95:1 +/- 0.06 (SEM). Group 3 was comprised of 6 bulls with spheroid-shaped testicles and a length to width ratio of 1.3:1 +/- 0.03 (SEM). All the groups were statistically different for length to width ratios (P < 0.05). Length measurements from cranial to caudal pole of the testis proper were also different between groups (P < 0.05). Width or testicular diameter was different between Group 2 and Group 3 at P < 0.05; however, there was no difference between Group 1 and Group 2 or between Group 1 and Group 3. Predicted volumes and weights of testicles were not significantly different between groups. Scrotal circumference measurements were significantly different between groups (P < 0.05). Group 1 had an average SC of 43.07 +/- 0.36 cm (SEM), Group 2 of 39.33 +/- 1.18 cm (SEM) and Group 3 of 46.22 +/- 0.69 cm (SEM). Sperm production for a twice daily, 2-day-per-week collection schedule revealed a statistically significant difference for sperm output. A total of 2742 ejaculates was evaluated. A total of 1818 ejaculates was evaluated in Group 1, 440 ejaculates in Group 2 and 484 ejaculates in Group 3. The mean spermatozoal harvest per day for Group 1 bulls was 13.62 +/- 0.09 x 10(9) (SEM). Group 2 bulls with the longer-shaped testicles produced 14.82 +/- 0.18 x 10(9) (SEM) spermatozoa per day, and Group 3 bulls, with the more rounded testicle shape and the significantly larger SC produced 11.72 +/- 0.64 x 10(9)(SEM) sperm cells per day. All 3 groups were statistically different at the P = 0.05 level. The results suggest that prediction of sperm production may be dependent on factors other than SC, testicular volume, or weight. Testicular shape may influence sperm output in mature Holstein bulls.

Journal Article↗

ParaMEME: a parallel implementation and a web interface for a DNA and protein motif discovery tool.

Many advanced software tools fail to reach a wide audience because they require specialized hardware, installation expertise, or an abundance of CPU cycles. The worldwide web offers a new opportunity for distributing such systems. One such program, MEME, discovers repeated patterns, called motifs, in sets of DNA or protein sequences. This tool is now available to biologists over the worldwide web, using an asynchronous, single-program multiple-data version of the program called ParaMEME that runs on an Intel Paragon XP/S parallel computer at the San Diego Super-computer Center. ParaMEME scales gracefully to 64 nodes on the Paragon with efficiencies > 72% for large data sets. The worldwide web interface to ParaMEME accepts a set of sequences interactively from a user, submits the sequences to the Paragon for analysis, and e-mails the results back to the user. ParaMEME is available for free public use at http://@www.sdsc.edu/CompSci/Biomed/ MEME.

Algorithms↗

The megaprior heuristic for discovering protein sequence patterns.

Several computer algorithms for discovering patterns in groups of protein sequences are in use that are based on fitting the parameters of a statistical model to a group of related sequences. These include hidden Markov model (HMM) algorithms for multiple sequence alignment, and the MEME and Gibbs sampler algorithms for discovering motifs. These algorithms are sometimes prone to producing models that are incorrect because two or more patients have been combined. The statistical model produced in this situation is a convex combination (weighted average) of two or more different models. This paper presents a solution to the problem of convex combinations in the form of a heuristic based on using extremely low variance Dirichlet mixture priors as part of the statistical model. This heuristic, which we call the megaprior heuristic, increase the strength (i.e., decreases the variance) of the prior in proportion to the size of the sequence dataset. This causes each column in the final model to strongly resemble the mean of a single component of the prior, regardless of the size of the dataset. We describe the cause of the convex combination problem, analyze it mathematically, motivate and describe the implementation of the megaprior heuristic, and show how it can effectively eliminate the problem of convex combinations in protein sequence pattern discovery.

Algorithms↗

The value of prior knowledge in discovering motifs with MEME.

MEME is a tool for discovering motifs in sets of protein or DNA sequences. This paper describes several extensions to MEME which increase its ability to find motifs in a totally unsupervised fashion, but which also allow it to benefit when prior knowledge is available. When no background knowledge is asserted. MEME obtains increased robustness from a method for determining motif widths automatically, and from probabilistic models that allow motifs to be absent in some input sequences. On the other hand, MEME can exploit prior knowledge about a motif being present in all input sequences, about the length of a motif and whether it is a palindrome, and (using Dirichlet mixtures) about expected patterns in individual motif positions. Extensive experiments are reported which support the claim that MEME benefits from, but does not require, background knowledge. The experiments use seven previously studied DNA and protein sequence families and 75 of the protein families documented in the Prosite database of sites and patterns, Release 11.1.

Algorithms↗

Fitting a mixture model by expectation maximization to discover motifs in biopolymers.

The algorithm described in this paper discovers one or more motifs in a collection of DNA or protein sequences by using the technique of expectation maximization to fit a two-component finite mixture model to the set of sequences. Multiple motifs are found by fitting a mixture model to the data, probabilistically erasing the occurrences of the motif thus found, and repeating the process to find successive motifs. The algorithm requires only a set of unaligned sequences and a number specifying the width of the motifs as input. It returns a model of each motif and a threshold which together can be used as a Bayes-optimal classifier for searching for occurrences of the motif in other databases. The algorithm estimates how many times each motif occurs in each sequence in the dataset and outputs an alignment of the occurrences of the motif. The algorithm is capable of discovering several different motifs with differing numbers of occurrences in a single dataset.

Algorithms↗

Activation of human peripheral blood T lymphocytes by pharmacological induction of protein-tyrosine phosphorylation.

Protein-tyrosine kinase and protein-tyrosine phosphatase (PTPase) activities are essential for T-cell antigen receptor-mediated signaling. To assess the functional consequences of alteration of the levels of tyrosine phosphorylation in normal human T cells, the effects of vanadate and hydrogen peroxide were studied. In combination, these agents induced tyrosine phosphorylation of cellular substrates, elevated cytosolic free calcium, and induced interleukin 2 receptor (IL-2R) alpha chain expression but not IL-2 secretion. However, anti-CD28 antibody in combination with vanadate and hydrogen peroxide induced IL-2 secretion, consistent with the requirement for a costimulatory signal in the induction of this gene. The effects of vanadate and hydrogen peroxide were enhanced in the absence of the T-cell PTPase, CD45. Thus, acute pharmacologic manipulation of the level of tyrosine phosphorylation in normal T cells correlates with partial, but not full, activation of these cells; in concert with a costimulatory signal provided by perturbation of the CD28 molecule, the complete program of activation is initiated. These agents should prove useful in dissecting signaling pathways involved in the regulation of genes critical to the immune response.

Animals↗

Expression of v-src in a murine T-cell hybridoma results in constitutive T-cell receptor phosphorylation and interleukin 2 production.

Ligand binding to the T-cell antigen receptor results in phosphatidylinositol hydrolysis and the resultant activation of protein kinase C, as well as the activation of a receptor-coupled protein-tyrosine kinase. As a model for tyrosine kinase activation in T cells, we used retroviral gene transfer to express the v-src oncogene in an antigen-specific murine T-cell hybridoma. Clones that expressed v-src mRNA demonstrated constitutive tyrosine phosphorylation of several cellular substrates, including the zeta chain of the T-cell receptor, and constitutive interleukin 2 production. Thus, expression of a constitutively active protein-tyrosine kinase such as pp60v-src appears to be sufficient to induce the expression of at least one gene critical to the process of T-cell activation.

Animals↗

In vitro and xenogenous capacitation-like changes of fresh, cooled, and cryopreserved stallion sperm as assessed by a chlortetracycline stain.

Like the human female, the mare experiences reproductive tract pathology that may sometimes be circumvented by the use of assisted reproductive technologies (ARTs). One such technology, gamete intrafallopian transfer (GIFT), may be used in mares that exhibit ovulatory, oviductal, or uterine abnormalities that limit the use of common ARTs, such as embryo transfer. Homologous GIFT has been successfully performed in the horse; however, the logistics, costs, and associated risks of surgically transferring gametes to the oviducts of a recipient mare are considerably high. Use of a less costly species in a heterologous or xenogenous procedure would therefore be beneficial. This study represents the preliminary investigation into the use of sheep as recipients for xenogenous GIFT procedures using equine gametes. We investigated the capacitation response of fresh, cooled, or frozen stallion sperm after 1) in vivo incubation in the reproductive tract of estrous and anestrous ewes as well as 2) in vitro incubation in a modified Krebs/ Ringer extender at 37 degreesC with and without the addition of heparin at 10 IU/mL for up to 8 hours. A chlortetracycline (CTC) fluorescent stain was used to assess the capacitation response of sperm. Findings indicated that oviductal fluid samples recovered from estrous ewes had significantly higher numbers of sperm exhibiting capacitation-like staining patterns when compared to samples recovered from anestrous ewes (P < .05). Fresh semen yielded higher capacitation-like staining patterns after in vivo incubation than did frozen-thawed or cooled samples. A transition from majority CTC unreacted sperm to majority CTC non-acrosome intact sperm was demonstrated for both in vivo and in vitro studies. In vitro incubation of stallion sperm with heparin did not result in an increased capacitation-like staining response over time when compared with nonheparinized samples. Results from this study suggest that xenogenous capacitation of stallion sperm may occur in the estrous ewe.

Anestrus↗