Biomedical language processing: what's beyond PubMed?
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to Lawrence Hunter.
Explore the source record for details and available documents.
BACKGROUND: Our approach to Task 1A was inspired by Tanabe and Wilbur's ABGene system. Like Tanabe and Wilbur, we approached the problem as one of part-of-speech tagging, adding a GENE tag to the standard tag set. Where their system uses the Brill tagger, we used TnT, the Trigrams 'n' Tags HMM-based part-of-speech tagger. Based on careful error analysis, we implemented a set of post-processing rules to correct both false positives and false negatives. We participated in both the open and the closed divisions; for the open division, we made use of data from NCBI. RESULTS: Our base system without post-processing achieved a precision and recall of 68.0% and 77.2%, respectively, giving an F-measure of 72.3%. The full system with post-processing achieved a precision and recall of 80.3% and 80.5% giving an F-measure of 80.4%. We achieved a slight improvement (F-measure = 80.9%) by employing a dictionary-based post-processing step for the open division. We placed third in both the open and the closed division. CONCLUSION: Our results show that a part-of-speech tagger can be augmented with post-processing rules resulting in an entity identification system that competes well with other approaches.
Explore the source record for details and available documents.
The present studies extend recent findings that mice null for the alpha(2A) adrenergic receptor (alpha(2A) AR KO mice) lack suppression of exogenous secretagogue-stimulated insulin secretion in response to alpha(2) AR agonists by evaluating the endogenous secretagogue, glucose, ex vivo, and providing in vivo data that baseline insulin levels are elevated and baseline glucose levels are decreased in alpha(2A) AR KO mice. These latter findings reveal that the alpha(2A) AR subtype regulates glucose-stimulated insulin release in response to endogenous catecholamines in vivo. The changes in alpha(2A) AR responsiveness and resultant changes in insulin/glucose homeostasis encouraged us to utilize proteomics strategies to identify possible alpha(2A) AR downstream signaling molecules or other resultant changes due to perturbation of alpha(2A) AR expression. Although agonist stimulation of islets from wild type (WT) mice did not significantly alter islet protein profiles, several proteins were enriched in islets from alpha(2A) AR KO mice when compared with those from WT mice, including an enzyme participating in insulin protein processing. The present studies document the important role of the alpha(2A) AR subtype in tonic suppression of insulin release in response to endogenous catecholamines as well as exogenous alpha(2) agonists and provide insights into pleiotropic changes that result from loss of alpha(2A) AR expression and tonic suppression of insulin release.
In this paper we argue that a richer underlying representational model for the Gene Ontology that captures the implicit compositional structure of GO terms could have a positive impact on two activities crucial to the success of GO: ontology curation and database annotation. We show that many of the new terms added to GO in a one-year span appear to be compositional variations of other terms. We found that 90.2% of the 3,652 new terms added between July 2003 and July 2004 exhibited characteristics of compositionality. We also examine annotations available from the GO Consortium website that are either manually curated or automatically generated. We found that 74.5% and 63.2% of GO terms are seldom, if ever, used in manual and automatic annotations, respectively. We show that there are features that tend to distinguish terms that are used from those that are not. In order to characterize the effect of compositionality on the combinatorial properties of GO, we employ finite state automata that represent sets of GO terms. This representational tool demonstrates how ontologies can grow very fast, and also shows that small conceptual changes can directly result in a large number of changes to the terminology. We argue that the curation and annotation findings we report are influenced by the combinatorial properties that present themselves in an ontology that does not have a model that properly captures the compositional structure of its terms.
Two-dimensional gel electrophoresis (2-DE) was used to separate protein samples solubilized from the nucleus accumbens and hippocampus of alcohol-naïve, adult, male inbred alcohol-preferring (iP) and alcohol-nonpreferring (iNP) rats. Several protein spots were excised from the gel, destained, digested with trypsin, and analyzed by mass spectrometry. In the hippocampus, 1629 protein spots were matched to the reference pattern, and in the nucleus accumbens, 1390 protein spots were matched. Approximately 70 proteins were identified in both regions. In the hippocampus, only 8 of the 1629 matched protein spots differed in abundance between the iP and iNP rats. In the nucleus accumbens, 32 of the 1390 matched protein spots differed in abundance between the iP and iNP rats. In the hippocampus, the abundances of all 8 proteins were higher in the iNP than iP rat. In the nucleus accumbens, the abundances of 31 of 32 proteins were higher in the iNP than iP rat. In the hippocampus, only 2 of the 8 proteins that differed could be identified, whereas in the nucleus accumbens 21 of the 32 proteins that differed were identified. Higher abundances of cellular retinoic acid-binding protein 1 and a calmodulin-dependent protein kinase (both of which are involved in cellular signaling pathways) were found in both regions of the iNP than iP rat. In the nucleus accumbens, additional differences in the abundances of proteins involved in (i) metabolism (e.g., calpain, parkin, glucokinase, apolipoprotein E, sorbitol dehydrogenase), (ii) cyto-skeletal and intracellular protein transport (e.g., beta-actin), (iii) molecular chaperoning (e.g., grp 78, hsc70, hsc 60, grp75, prohibitin), (iv) cellular signaling pathways (e.g., protein kinase C-binding protein), (v) synaptic function (e.g., complexin I, gamma-enolase, syndapin IIbb), (vi) reduction of oxidative stress (thioredoxin peroxidase), and (vii) growth and differentiation (hippocampal cholinergic neurostimulating peptide) were found. The results of this study indicate that selective breeding for disparate alcohol drinking behaviors produced innate alterations in the expression of several proteins that could influence neuronal function within the nucleus accumbens and hippocampus.
OBJECTIVE: We sought to identify genes with differential expression in cerebral cavernous malformations (CCMs), arteriovenous malformations (AVMs), and control superficial temporal arteries (STAs) and to confirm differential expression of genes previously implicated in the pathobiology of these lesions. METHODS: Total ribonucleic acid was isolated from four CCM, four AVM, and three STA surgical specimens and used to quantify lesion-specific messenger ribonucleic acid expression levels on human gene arrays. Data were analyzed with the use of two separate methodologies: gene discovery and confirmation analysis. RESULTS: The gene discovery method identified 42 genes that were significantly up-regulated and 36 genes that were significantly down-regulated in CCMs as compared with AVMs and STAs (P = 0.006). Similarly, 48 genes were significantly up-regulated and 59 genes were significantly down-regulated in AVMs as compared with CCMs and STAs (P = 0.006). The confirmation analysis showed significant differential expression (P < 0.05) in 11 of 15 genes (angiogenesis factors, receptors, and structural proteins) that previously had been reported to be expressed differentially in CCMs and AVMs in immunohistochemical analysis. CONCLUSION: We identify numerous genes that are differentially expressed in CCMs and AVMs and correlate expression with the immunohistochemistry of genes implicated in cerebrovascular malformations. In future efforts, we will aim to confirm candidate genes specifically related to the pathobiology of cerebrovascular malformations and determine their biological systems and mechanistic relevance.
A response to Life sentences: Ontology recapitulates philology by Sydney Brenner, Genome Biology 2002, 3:comment1006.1-1006.2.
UNLABELLED: One of the challenges to the effective utilization of cDNA microarray analysis in mouse models of oncogenesis is the choice of a critical set of probes that are informative for human disease. Given the thousands of genes with a potential role in human oncogenesis and the hundreds of thousands of mouse sequences available for use as probes, selection of an informative set of mouse probes can be an overwhelming task. We have developed a web based sequence mining tool using DataBase Independent (DBI) Perl to annotate publicly available sequences. The Mouse Oncochip Design Tool uses the Mouse Genome Database (MGD) developed and maintained by the Jackson Laboratories for mouse DNA sequences. There are over 380 000 sequences in their database. The output list has been ordered to present the genes more likely to be informative in a mouse model of human cancer using a candidate set of oncogenes to order the list. Mouse sequences that represent genes that are homologous with a member of a human oncogene set are listed first. In addition it provides a set of links for information on clone source gene function. CONTACT: http://nciarray.nci.nih.gov/cgi-bin/me/mouse_design.cgi