PubMed Health⌕ Search

Biomedical subjects

E Eskin

Publications and source records attributed to E Eskin.

4 recordsLinked to original sources

Identification of additional variants within the human dopamine transporter gene provides further evidence for an association with bipolar disorder in two independent samples.

The dopamine transporter (DAT) is the site of action of stimulants, and variations in the human DAT gene (DAT1) have been associated with susceptibility to several psychiatric disorders including attention deficit hyperactivity disorder (ADHD) and bipolar disorder. We have previously reported the association of bipolar disorder to novel SNPs in the 3' end of DAT1. We now report the identification of 20 additional SNPs in DAT1 for a total of 63 variants. We also report evidence for association to bipolar disorder in a second independent sample of families. Eight newly identified SNPs and 14 previously identified SNPs were analyzed in two independent samples of 50 and 70 families each using the transmission disequilibrium test. Two of the eight new SNPs, one in intron 8 and one in intron 13, were found to be moderately associated with bipolar disorder, each in one of the two independent samples. Analysis of haplotypes comprised of all 22 SNPs in sliding windows of five adjacent SNPs revealed an association to the region near introns 7 and 8 in both samples (empirical P-values 0.002 and 0.001, respectively, for the same window). The haplotype block structure observed in the gene in our previous study was confirmed in this sample with greater resolution allowing for discrimination of a third haplotype block in the middle of the gene. Together, these data are consistent with the presence of multiple variants in DAT1 that convey susceptibility to bipolar disorder.

Bipolar Disorder↗

Combining text mining and sequence analysis to discover protein functional regions.

Recently presented protein sequence classification models can identify relevant regions of the sequence. This observation has many potential applications to detecting functional regions of proteins. However, identifying such sequence regions automatically is difficult in practice, as relatively few types of information have enough annotated sequences to perform this analysis. Our approach addresses this data scarcity problem by combining text and sequence analysis. First, we train a text classifier over the explicit textual annotations available for some of the sequences in the dataset, and use the trained classifier to predict the class for the rest of the unlabeled sequences. We then train a joint sequence text classifier over the text contained in the functional annotations of the sequences, and the actual sequences in this larger, automatically extended dataset. Finally, we project the classifier onto the original sequences to determine the relevant regions of the sequences. We demonstrate the effectiveness of our approach by predicting protein sub-cellular localization and determining localization specific functional regions of these proteins.

Algorithms↗

Using mixtures of common ancestors for estimating the probabilities of discrete events in biological sequences.

Accurately estimating probabilities from observations is important for probabilistic-based approaches to problems in computational biology. In this paper we present a biologically-motivated method for estimating probability distributions over discrete alphabets from observations using a mixture model of common ancestors. The method is an extension of substitution matrix-based probability estimation methods. In contrast to previous such methods, our method has a simple Bayesian interpretation and has the advantage over Dirichlet mixtures that it is both effective and simple to compute for large alphabets. The method is applied to estimate amino acid probabilities based on observed counts in an alignment and is shown to perform comparably to previous methods. The method is also applied to estimate probability distributions over protein families and improves protein classification accuracy.

Amino Acid Sequence↗

Protein family classification using sparse Markov transducers.

In this paper we present a method for classifying proteins into families using sparse Markov transducers (SMTs). Sparse Markov transducers, similar to probabilistic suffix trees, estimate a probability distribution conditioned on an input sequence. SMTs generalize probabilistic suffix trees by allowing for wild-cards in the conditioning sequences. Because substitutions of amino acids are common in protein families, incorporating wildcards into the model significantly improves classification performance. We present two models for building protein family classifiers using SMTs. We also present efficient data structures to improve the memory usage of the models. We evaluate SMTs by building protein family classifiers using the Pfam database and compare our results to previously published results.

Algorithms↗