PubMed Health⌕ Search

Biomedical subjects

Anton Yuryev

Publications and source records attributed to Anton Yuryev.

10 recordsLinked to original sources

Automatic pathway building in biological association networks.

BACKGROUND: Scientific literature is a source of the most reliable and comprehensive knowledge about molecular interaction networks. Formalization of this knowledge is necessary for computational analysis and is achieved by automatic fact extraction using various text-mining algorithms. Most of these techniques suffer from high false positive rates and redundancy of the extracted information. The extracted facts form a large network with no pathways defined. RESULTS: We describe the methodology for automatic curation of Biological Association Networks (BANs) derived by a natural language processing technology called Medscan. The curated data is used for automatic pathway reconstruction. The algorithm for the reconstruction of signaling pathways is also described and validated by comparison with manually curated pathways and tissue-specific gene expression profiles. CONCLUSION: Biological Association Networks extracted by MedScan technology contain sufficient information for constructing thousands of mammalian signaling pathways for multiple tissues. The automatically curated MedScan data is adequate for automatic generation of good quality signaling networks. The automatically generated Regulome pathways and manually curated pathways used for their validation are available free in the ResNetCore database from Ariadne Genomics, Inc. 1. The pathways can be viewed and analyzed through the use of a free demo version of PathwayStudio software. The Medscan technology is also available for evaluation using the free demo version of PathwayStudio software.

Databases, Bibliographic↗

Experimental validation of data mined single nucleotide polymorphisms from several databases and consecutive dbSNP builds.

Rapid development in the annotation of human genetic variation has increased the numbers of single nucleotide polymorphisms (SNPs) in candidate genes by several orders of magnitude. The selection of both useful target SNPs for disease-gene association studies and SNPs associated with the treatment response is therefore an increasingly challenging task. We describe a workflow for selecting SNPs based on their putative function and frequency in candidate genes extracted from PubMed resources. The annotation of each SNP and its frequency in a Caucasian population was assessed in several databases. Approximately 4000 SNPs were identified from an initial 233 candidate genes. In a case study, we performed actual genotyping of 1030 of these SNPs in 213 genes and obtained 710 successfully genotyped SNPs. Using the flow-chart outlined here, only 87 SNPs were monomorphic (approximately 12%). This study reports the frequency of SNPs in a Caucasian population, selected in silico, using a candidate gene approach and validated by actually genotyping 193 individuals. The selected genotypes represent a valuable set of verified candidate SNPs for pharmacogenetic studies in Caucasian populations.

Breast Neoplasms↗

Targeting transcription factors in cell regulation.

The author compares transcription factors (TFs) with 13 other protein classes from the point of their ability to regulate various cellular processes. The comparison is performed using data from the ResNet 4.0 database containing molecular interactions extracted from scientific literature. The author introduces two quantitative characteristics for evaluating the ability of every protein to regulate a cell process. Using these measures, he evaluates the efficiency of TFs and other protein classes to regulate the biological process pathways. It was found that TFs, on average, are not the best class for regulating an individual cell process. They have lower regulatory specificity, (i.e., single TF tends to regulate many different biological processes). TFs also tend to be placed downstream in the biological process pathways, being a target of the regulatory relation more often than being a regulator. Possible implications of these findings for drug development are discussed.

Cell Physiological Phenomena↗

Binding properties and evolution of homodimers in protein-protein interaction networks.

We demonstrate that protein-protein interaction networks in several eukaryotic organisms contain significantly more self-interacting proteins than expected if such homodimers randomly appeared in the course of the evolution. We also show that on average homodimers have twice as many interaction partners than non-self-interacting proteins. More specifically, the likelihood of a protein to physically interact with itself was found to be proportional to the total number of its binding partners. These properties of dimers are in agreement with a phenomenological model, in which individual proteins differ from each other by the degree of their 'stickiness' or general propensity toward interaction with other proteins including oneself. A duplication of self-interacting proteins creates a pair of paralogous proteins interacting with each other. We show that such pairs occur more frequently than could be explained by pure chance alone. Similar to homodimers, proteins involved in heterodimers with their paralogs on average have twice as many interacting partners than the rest of the network. The likelihood of a pair of paralogous proteins to interact with each other was also shown to decrease with their sequence similarity. This points to the conclusion that most of interactions between paralogs are inherited from ancestral homodimeric proteins, rather than established de novo after duplication. We finally discuss possible implications of our empirical observations from functional and evolutionary standpoints.

Animals↗

Primer design and marker clustering for multiplex SNP-IT primer extension genotyping assay using statistical modeling.

MOTIVATION: The optimization of the primer design is critical for the development of high-throughput SNP genotyping methods. Recently developed statistical models of the SNP-IT primer extension genotyping reaction allow further improvement of primer quality for the assay. RESULTS: Here we describe how the statistical models can be used to improve primer design for the assay. We also show how to optimize clustering of the SNP markers into multiplex panels using statistical model for multiplex SNP-IT. The primer set failure probability calculated by a model is used as a minimization function for both primer selection and primers clustering. Three clustering algorithms for the multiplex genotyping SNP-IT assay are described and their relative performance is evaluated. We also describe the approaches to improve the speed of primer design and clustering calculations when using the statistical models. Our clustering decreases the average failure probability of the marker set by 7-25%. The experimental marker failure rate in the multiplex reaction was reduced dramatically and success rate can be achieved as high as 96%. AVAILABILITY: The primer design using statistical models is freely available from www.autoprimer.com.

Algorithms↗

Auto-validation of fluorescent primer extension genotyping assay using signal clustering and neural networks.

BACKGROUND: SNP genotyping typically incorporates a review step to ensure that the genotype calls for a particular SNP are correct. For high-throughput genotyping, such as that provided by the GenomeLab SNPstream instrument from Beckman Coulter, Inc., the manual review used for low-volume genotyping becomes a major bottleneck. The work reported here describes the application of a neural network to automate the review of results. RESULTS: We describe an approach to reviewing the quality of primer extension 2-color fluorescent reactions by clustering optical signals obtained from multiple samples and a single reaction set-up. The method evaluates the quality of the signal clusters from the genotyping results. We developed 64 scores to measure the geometry and position of the signal clusters. The expected signal distribution was represented by a distribution of a 64-component parametric vector obtained by training the two-layer neural network onto a set of 10,968 manually reviewed 2D plots containing the signal clusters. CONCLUSION: The neural network approach described in this paper may be used with results from the GenomeLab SNPstream instrument for high-throughput SNP genotyping. The overall correlation with manual revision was 0.844. The approach can be applied to a quality review of results from other high-throughput fluorescent-based biochemical assays in a high-throughput mode.

Artificial Intelligence↗

A simple and practical dictionary-based approach for identification of proteins in Medline abstracts.

OBJECTIVE: The aim of this study was to develop a practical and efficient protein identification system for biomedical corpora. DESIGN: The developed system, called ProtScan, utilizes a carefully constructed dictionary of mammalian proteins in conjunction with a specialized tokenization algorithm to identify and tag protein name occurrences in biomedical texts and also takes advantage of Medline "Name-of-Substance" (NOS) annotation. The dictionaries for ProtScan were constructed in a semi-automatic way from various public-domain sequence databases followed by an intensive expert curation step. MEASUREMENTS: The recall and precision of the system have been determined using 1000 randomly selected and hand-tagged Medline abstracts. RESULTS: The developed system is capable of identifying protein occurrences in Medline abstracts with a 98% precision and 88% recall. It was also found to be capable of processing approximately 300 abstracts per second. Without utilization of NOS annotation, precision and recall were found to be 98.5% and 84%, respectively. CONCLUSION: The developed system appears to be well suited for protein-based Medline indexing and can help to improve biomedical information retrieval. Further approaches to ProtScan's recall improvement also are discussed.

Abstracting and Indexing↗

Extracting human protein interactions from MEDLINE using a full-sentence parser.

MOTIVATION: The living cell is a complex machine that depends on the proper functioning of its numerous parts, including proteins. Understanding protein functions and how they modify and regulate each other is the next great challenge for life-sciences researchers. The collective knowledge about protein functions and pathways is scattered throughout numerous publications in scientific journals. Bringing the relevant information together becomes a bottleneck in a research and discovery process. The volume of such information grows exponentially, which renders manual curation impractical. As a viable alternative, automated literature processing tools could be employed to extract and organize biological data into a knowledge base, making it amenable to computational analysis and data mining. RESULTS: We present MedScan, a completely automated natural language processing-based information extraction system. We have used MedScan to extract 2976 interactions between human proteins from MEDLINE abstracts dated after 1988. The precision of the extracted information was found to be 91%. Comparison with the existing protein interaction databases BIND and DIP revealed that 96% of extracted information is novel. The recall rate of MedScan was found to be 21%. Additional experiments with MedScan suggest that MEDLINE is a unique source of diverse protein function information, which can be extracted in a completely automated way with a reasonably high precision. Further directions of the MedScan technology improvement are discussed. AVAILABILITY: MedScan is available for commercial licensing from Ariadne Genomics, Inc.

Abstracting and Indexing↗

Novel raf kinase protein-protein interactions found by an exhaustive yeast two-hybrid analysis.

We have performed an exhaustive unbiased yeast two-hybrid analysis to identify interaction partners of two human Raf kinase isoforms, A-Raf and C-Raf, using their N-terminal regulatory domain as "bait." A total of 20 different human proteins were found to interact with Raf isoforms. Several of these interactions were novel and an extensive bioinformatics evaluation was performed for each. The novel putative interactions include a signalosome component, TOPK/PBK kinase, and two new putative protein phosphatases. The cysteine-rich zinc-binding domain (CRD) of Raf was found to interact with all 20 proteins and to achieve isoform-specific interactions. Since similar putative CRDs are present in a variety of protein serine-threonine kinases, the data suggest that the CRD may function as a major protein-protein interaction domain of these kinases. We propose possible functional consequences of these novel Raf interactions.

Amino Acid Sequence↗

Predicting the success of primer extension genotyping assays using statistical modeling.

Using an empirical panel of more than 20 000 single base primer extension (SNP-IT) assays we have developed a set of statistical scores for evaluating and rank ordering various parameters of the SNP-IT reaction to facilitate high-throughput assay primer design with improved likelihood of success. Each score predicts either signal magnitude from primer extension or signal noise caused by mispriming of primers and structure of the PCR product. All scores have been shown to correlate with the success/failure rate of the SNP-IT reaction, based on analysis of assay results. A logistic regression analysis was applied to combine all scored parameters into one measure predicting the overall success/failure rate of a given SNP marker. Three training sets for different types of SNP-IT reaction, each containing about 22 000 SNP markers, were used to assign weights to each score and optimize the prediction of the combined measure. c-Statistics of 0.69, 0.77 and 0.72 were achieved for three training sets. This new statistical prediction can be used to improve primer design for the SNP-IT reaction and evaluate the probability of genotyping success for a given SNP based on analysis of the surrounding genomic sequence.

DNA Primers↗