PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Natural Language Processing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Tailored gene array databases: applications in mechanistic toxicology.

MOTIVATION: The development of an annotated global database suitable for a wide range of investigations is a challenging and labor-intensive task. Thus, the development of databases tailored for specific applications remains necessary. For example, in the field of toxicology, no annotated gene array databases are now available that may assist in the correlation of changes in gene activity to cellular functions and processes associated with the toxic response. RESULTS: As an example of a tailored annotated database, an attempt was made to systematize available biological information on genes present on the Affymetrix Rat Toxicology U34 GeneChip, with a focus on how the gene products relate to liver cells and their response to chemical toxins. The information collected was imbedded in a local relational database to analyze data obtained in toxicological gene array experiments with hydrazine-exposed hepatocytes. The advantages and benefits of the tailored database in the biological interpretation of the results are demonstrated.

Abstracting and Indexing↗

Initial large-scale exploration of protein-protein interactions in human brain.

Study of protein interaction networks is crucial to post-genomic systems biology. Aided by high-throughput screening technologies, biologists are rapidly accumulating protein-protein interaction data. Using a random yeast two-hybrid (R2H) process, we have performed large-scale yeast two-hybrid searches with approximately fifty thousand random human brain cDNA bait fragments against a human brain cDNA prey fragment library. From these searches, we have identified 13,656 unique protein-protein interaction pairs involving 4,473 distinct known human loci. In this paper, we have performed our initial characterization of the protein interaction network in human brain tissue. We have classified and characterized all identified interactions based on Gene Ontology (GO) annotation of interacting loci. We have also described the "scale-free" topological structure of the network.

Brain↗

PDQ Wizard: automated prioritization and characterization of gene and protein lists using biomedical literature.

SUMMARY: PDQ Wizard automates the process of interrogating biomedical references using large lists of genes, proteins or free text. Using the principle of linkage through co-citation biologists can mine PubMed with these proteins or genes to identify relationships within a biological field of interest. In addition, PDQ Wizard provides novel features to define more specific relationships, highlight key publications describing those activities and relationships, and enhance protein queries. PDQ Wizard also outputs a metric that can be used for prioritization of genes and proteins for further research. AVAILABILITY: PDQ Wizard is freely available from http://www.gti.ed.ac.uk/pdqwizard/.

Abstracting and Indexing↗

MeKE: discovering the functions of gene products from biomedical literature via sentence alignment.

MOTIVATION: Research on roles of gene products in cells is accumulating and changing rapidly, but most of the results are still reported in text form and are not directly accessible by computers. To expedite the progress of functional bioinformatics, it is, therefore, important to efficiently process large amounts of biomedical literature and transform the knowledge extracted into a structured format usable by biologists and medical researchers. Our aim was to develop an intelligent text-mining system that will extract from biomedical documents knowledge about the functions of gene products and thus facilitate computing with function. RESULTS: We have developed an ontology-based text-mining system to efficiently extract from biomedical literature knowledge about the functions of gene products. We also propose methods of sentence alignment and sentence classification to discover the functions of gene products discussed in digital texts. AVAILABILITY: http://ismp.csie.ncku.edu.tw/~yuhc/meke/

Biomedical Research↗

An ontology for a Robot Scientist.

MOTIVATION: A Robot Scientist is a physically implemented robotic system that can automatically carry out cycles of scientific experimentation. We are commissioning a new Robot Scientist designed to investigate gene function in S. cerevisiae. This Robot Scientist will be capable of initiating >1,000 experiments, and making >200,000 observations a day. Robot Scientists provide a unique test bed for the development of methodologies for the curation and annotation of scientific experiments: because the experiments are conceived and executed automatically by computer, it is possible to completely capture and digitally curate all aspects of the scientific process. This new ability brings with it significant technical challenges. To meet these we apply an ontology driven approach to the representation of all the Robot Scientist's data and metadata. RESULTS: We demonstrate the utility of developing an ontology for our new Robot Scientist. This ontology is based on a general ontology of experiments. The ontology aids the curation and annotating of the experimental data and metadata, and the equipment metadata, and supports the design of database systems to hold the data and metadata. AVAILABILITY: EXPO in XML and OWL formats is at: http://sourceforge.net/projects/expo/. All materials about the Robot Scientist project are available at: http://www.aber.ac.uk/compsci/Research/bio/robotsci/.

Artificial Intelligence↗

Use of morphological analysis in protein name recognition.

Protein name recognition aims to detect each and every protein names appearing in a PubMed abstract. The task is not simple, as the graphic word boundary (space separator) assumed in conventional preprocessing does not necessarily coincide with the protein name boundary. Such boundary disagreement caused by tokenization ambiguity has usually been ignored in conventional preprocessing of general English. In this paper, we argue that boundary disagreement poses serious limitations in biomedical English text processing, not to mention protein name recognition. Our key idea for dealing with the boundary disagreement is to apply techniques used in Japanese morphological analysis where there are no word boundaries. Having evaluated the proposed method with GENIA corpus 3.02, we obtain F-measure of 69.01 on a strict criterion and 79.32 on a relaxed criterion. The result is comparable to other published work in protein name recognition, without resorting to manually prepared ad hoc feature engineering. Further, compared to the conventional preprocessing, the use of morphological analysis as preprocessing improves the performance of protein name recognition and reduces the execution time.

Abstracting and Indexing↗

GOstat: find statistically overrepresented Gene Ontologies within a group of genes.

SUMMARY: Modern experimental techniques, as for example DNA microarrays, as a result usually produce a long list of genes, which are potentially interesting in the analyzed process. In order to gain biological understanding from this type of data, it is necessary to analyze the functional annotations of all genes in this list. The Gene-Ontology (GO) database provides a useful tool to annotate and analyze the functions of a large number of genes. Here, we introduce a tool that utilizes this information to obtain an understanding of which annotations are typical for the analyzed list of genes. This program automatically obtains the GO annotations from a database and generates statistics of which annotations are overrepresented in the analyzed list of genes. This results in a list of GO terms sorted by their specificity. AVAILABILITY: Our program GOstat is accessible via the Internet at http://gostat.wehi.edu.au

Abstracting and Indexing↗

Constructing biological networks through combined literature mining and microarray analysis: a LMMA approach.

MOTIVATION: Network reconstruction of biological entities is very important for understanding biological processes and the organizational principles of biological systems. This work focuses on integrating both the literatures and microarray gene-expression data, and a combined literature mining and microarray analysis (LMMA) approach is developed to construct gene networks of a specific biological system. RESULTS: In the LMMA approach, a global network is first constructed using the literature-based co-occurrence method. It is then refined using microarray data through a multivariate selection procedure. An application of LMMA to the angiogenesis is presented. Our result shows that the LMMA-based network is more reliable than the co-occurrence-based network in dealing with multiple levels of KEGG gene, KEGG Orthology and pathway. AVAILABILITY: The LMMA program is available upon request.

Abstracting and Indexing↗

Mandarin and English single word processing studied with functional magnetic resonance imaging.

The cortical organization of language in bilinguals remains disputed. We studied 24 right-handed fluent bilinguals: 15 exposed to both Mandarin and English before the age of 6 years; and nine exposed to Mandarin in early childhood but English only after the age of 12 years. Blood oxygen level-dependent contrast functional magnetic resonance imaging was performed while subjects performed cued word generation in each language. Fixation was the control task. In both languages, activations were present in the prefrontal, temporal, and parietal regions, and the supplementary motor area. Activations in the prefrontal region were compared by (1) locating peak activations and (2) counting the number of voxels that exceeded a statistical threshold. Although there were differences in the magnitude of activation between the pair of languages, no subject showed significant differences in peak-location or hemispheric asymmetry of activations in the prefrontal language areas. Early and late bilinguals showed a similar pattern of overlapping activations. There are no significant differences in the cortical areas activated for both Mandarin and English at the single word level, irrespective of age of acquisition of either language.

Adolescent↗

A study of biomedical concept identification: MetaMap vs. people.

Although huge amounts of unstructured text are available as a rich source of biomedical knowledge, to process this unstructured knowledge requires tools that identify concepts from free-form text. MetaMap is one tool that system developers in biomedicine have commonly used for such a task, but few have studied how well it accomplishes this task in general. In this paper, we report on a study that compares MetaMap's performance against that of six people. Such studies are challenging because the task is inherently subjective and establishing consensus is difficult. Nonetheless, for those concepts that subjects generally agreed on, MetaMap was able to identify most concepts, if they were represented in the UMLS. However, MetaMap identified many other concepts that peo-ple did not. We also report on our analysis of the types of failures that MetaMap exhibited as well as trends in the way people chose to identify concepts.

Abstracting and Indexing↗

Literature mining and database annotation of protein phosphorylation using a rule-based system.

MOTIVATION: A large volume of experimental data on protein phosphorylation is buried in the fast-growing PubMed literature. While of great value, such information is limited in databases owing to the laborious process of literature-based curation. Computational literature mining holds promise to facilitate database curation. RESULTS: A rule-based system, RLIMS-P (Rule-based LIterature Mining System for Protein Phosphorylation), was used to extract protein phosphorylation information from MEDLINE abstracts. An annotation-tagged literature corpus developed at PIR was used to evaluate the system for finding phosphorylation papers and extracting phosphorylation objects (kinases, substrates and sites) from abstracts. RLIMS-P achieved a precision and recall of 91.4 and 96.4% for paper retrieval, and of 97.9 and 88.0% for extraction of substrates and sites. Coupling the high recall for paper retrieval and high precision for information extraction, RLIMS-P facilitates literature mining and database annotation of protein phosphorylation.

Abstracting and Indexing↗

Web content management by self-organization.

We present a new method for content management and knowledge discovery using a topology-preserving neural network. The method, termed topological organization of content (TOC), can generate a taxonomy of topics from a set of unannotated, unstructured documents. The TOC consists of a hierarchy of self-organizing growing chains (GCs), each of which can develop independently in terms of size and topics. The dynamic development process is validated continuously using a proposed entropy-based Bayesian information criterion (BIC). Each chain meeting the criterion spans child chains, with reduced vocabularies and increased specializations. This results in a topological tree hierarchy, which can be browsed like a table of contents directory or web portal. A brief review is given on existing methods for document clustering and organization, and clustering validation measures. The proposed approach has been tested and compared with several existing methods on real world web page datasets. The results have clearly demonstrated the advantages and efficiency in content organization of the proposed method in terms of computational cost and representation. The TOC can be easily adapted for large-scale applications. The topology provides a unique, additional feature for retrieving related topics and confining the search space.

Abstracting and Indexing↗

Evaluation of text data mining for database curation: lessons learned from the KDD Challenge Cup.

MOTIVATION: The biological literature is a major repository of knowledge. Many biological databases draw much of their content from a careful curation of this literature. However, as the volume of literature increases, the burden of curation increases. Text mining may provide useful tools to assist in the curation process. To date, the lack of standards has made it impossible to determine whether text mining techniques are sufficiently mature to be useful. RESULTS: We report on a Challenge Evaluation task that we created for the Knowledge Discovery and Data Mining (KDD) Challenge Cup. We provided a training corpus of 862 articles consisting of journal articles curated in FlyBase, along with the associated lists of genes and gene products, as well as the relevant data fields from FlyBase. For the test, we provided a corpus of 213 new ('blind') articles; the 18 participating groups provided systems that flagged articles for curation, based on whether the article contained experimental evidence for gene expression products. We report on the evaluation results and describe the techniques used by the top performing groups.

Abstracting and Indexing↗

Knowledge representation and retrieval using conceptual graphs and free text document self-organisation techniques.

Hospitals generate and store a large amount of clinical data each year, a significant portion of which is in free text format. Conventional database storage and retrieval algorithms are incapable of effectively processing free text medical data. The rich information and knowledge buried in healthcare records are unavailable for clinical decision-making. We examined a number of techniques for structuring and processing free text documents to effective and efficient for information retrieval and knowledge discovery. One critical success criterion is that the complexity of the techniques must be polynomial both in space and time for them to be able to cope with very large databases. We used conceptual graphs (CG) to capture the structure and semantic information/knowledge contained within the free text medical documents. Ordering and self-organising techniques (lattice techniques and knowledge space) were used to improve organisation of concepts from standard medical nomenclatures and large sets of free text medical documents. Pair-wise union of CG was performed to identify the common generalisation structure and a lattice structure of these CG documents. A combination of all three techniques allowed us to organise a set of 9000 discharge summaries into a generalisation hierarchy that supported efficient and rich information/knowledge retrieval.

Classification↗

Computerized extraction of coded findings from free-text radiologic reports. Work in progress.

A computerized data acquisition tool, the special purpose radiology understanding system (SPRUS), has been implemented as a module in the Health Evaluation through Logical Processing Hospital Information System. This tool uses semantic information from a diagnostic expert system to parse free-text radiology reports and to extract and encode both the findings and the radiologists' interpretations. These coded findings and interpretations are then stored in a clinical data base. The system recognizes both radiologic findings and diagnostic interpretations. Initial tests showed a true-positive rate of 87% for radiographic findings and a bad data rate of 5%. Diagnostic interpretations are recognized at a rate of 95% with a bad data rate of 6%. Testing suggests that these rates can be improved through enhancements to the system's thesaurus and the computerized medical knowledge that drives it. This system holds promise as a tool to obtain coded radiologic data for research, medical audit, and patient care.

Artificial Intelligence↗

Evolution of web site design: implications for medical education on the Internet.

Since its inception, the world wide web (WWW) has possessed the potential for becoming a 'watershed' medium for conveying complex, structured information across vast temporal and geographical barriers. In 1995, the MedWorld project (http:(/)/medworld.stanford.edu) was created at the Stanford University School of Medicine in an effort to innovate and explore the design process of creating WWW applications specifically for medical education. Until recently, the evolution of WWW applications has been mainly driven by technological advances in client-server technology, enabling or translating traditional modes of collaborative medical education (e.g. voice, presence, print, motion) into WWW devices and applications. Many of these applications, while technologically advanced, lack focused development of interface and interactivity design, which may enhance learning experiences. WWW applications which incorporate design innovation in parity with advances in client-server technology have been termed, 'third generation' web sites and have the potential to improve the quality of WWW applications designed for medical education. This work describes how the MedWorld project has created a 'third generation' WWW application by utilizing innovation in information, interface and interactivity design to create innovative WWW technology for the medical education arena.

Communications Media↗

German adaptations of ICD-10.

The introduction of the ICD-10, published by WHO 1992-94 in English and by DIMDI 1994/95 in German, is a very slow process. Some states introduced ICD-10 for the preparation of statistics of mortality, but only few use it for morbidity. ICD-9 is in Germany only in hospitals still in use. Much effort was put into the improvement of the official ICD-10 and the development of additional aids for more simple and better encoding of diagnoses. Thus a revision especially for ambulatory health care (ICD-10-SGBV with the incorporated ICD-10-Basisschlüssel) and a collection of German terms and expressions of diagnoses that are not at all part of the official ICD-10 (ICD-10-Diagnosenthesaurus) were published. Three years ago a conversion table ICD-9/-10 was developed which can now be harmonised with WHO's Translator. The experiences with all these instruments are satisfying. The development of methods for automatic encoding of free-text phrases of diagnoses has now been started.

Disease↗

Finding the evidence for protein-protein interactions from PubMed abstracts.

MOTIVATION: Protein-protein interactions play critical roles in biological processes, and many biologists try to find or to predict crucial information concerning these interactions. Before verifying interactions in biological laboratory work, validating them from previous research is necessary. Although many efforts have been made to create databases that store verified information in a structured form, much interaction information still remains as unstructured text. As the amount of new publications has increased rapidly, a large amount of research has sought to extract interactions from the text automatically. However, there remain various difficulties associated with the process of applying automatically generated results into manually annotated databases. For interactions that are not found in manually stored databases, researchers attempt to search for abstracts or full papers. RESULTS: As a result of a search for two proteins, PubMed frequently returns hundreds of abstracts. In this paper, a method is introduced that validates protein-protein interactions from PubMed abstracts. A query is generated from two given proteins automatically and abstracts are then collected from PubMed. Following this, target proteins and their synonyms are recognized and their interaction information is extracted from the collection. It was found that 67.37% of the interactions from DIP-PPI corpus were found from the PubMed abstracts and 87.37% of interactions were found from the given full texts. AVAILABILITY: Contact authors.

Abstracting and Indexing↗