PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Natural Language Processing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

GAPSCORE: finding gene and protein names one word at a time.

MOTIVATION: New high-throughput technologies have accelerated the accumulation of knowledge about genes and proteins. However, much knowledge is still stored as written natural language text. Therefore, we have developed a new method, GAPSCORE, to identify gene and protein names in text. GAPSCORE scores words based on a statistical model of gene names that quantifies their appearance, morphology and context. RESULTS: We evaluated GAPSCORE against the Yapex data set and achieved an F-score of 82.5% (83.3% recall, 81.5% precision) for partial matches and 57.6% (58.5% recall, 56.7% precision) for exact matches. Since the method is statistical, users can choose score cutoffs that adjust the performance according to their needs. AVAILABILITY: GAPSCORE is available at http://bionlp.stanford.edu/gapscore/

Abstracting and Indexing↗

AnaGram: protein function assignment.

SUMMARY: AnaGram is a web service for protein function assignment based on identity detection of small significant fragments (protomotifs) that can act as modular pieces in peptide construction. The system is able to assign function by finding correlations between protomotifs and functional annotations contained in SWISS-PROT and Medline databases. In addition, function ontologies are used for hierarchical organization of the predicted functions. Extensive tests have been carried out to evaluate the accuracy and performance of the system. AVAILABILITY: http://jaguar.genetica.uma.es/anagram.htm

Algorithms↗

Knowledge discovery by automated identification and ranking of implicit relationships.

MOTIVATION: New relationships are often implicit from existing information, but the amount and growth of published literature limits the scope of analysis an individual can accomplish. Our goal was to develop and test a computational method to identify relationships within scientific reports, such that large sets of relationships between unrelated items could be sought out and statistically ranked for their potential relevance as a set. RESULTS: We first construct a network of tentative relationships between 'objects' of biomedical research interest (e.g. genes, diseases, phenotypes, chemicals) by identifying their co-occurrences within all electronically available MEDLINE records. Relationships shared by two unrelated objects are then ranked against a random network model to estimate the statistical significance of any given grouping. When compared against known relationships, we find that this ranking correlates with both the probability and frequency of object co-occurrence, demonstrating the method is well suited to discover novel relationships based upon existing shared relationships. To test this, we identified compounds whose shared relationships predicted they might affect the development and/or progression of cardiac hypertrophy. When laboratory tests were performed in a rodent model, chlorpromazine was found to reduce the progression of cardiac hypertrophy.

Abstracting and Indexing↗

SaRAD: a Simple and Robust Abbreviation Dictionary.

MOTIVATION: Due to recent interest in the use of textual material to augment traditional experiments it has become necessary to automatically cluster, classify and filter natural language information. RESULTS: The Simple and Robust Abbreviation Dictionary (SaRAD) provides an easy to implement, high performance tool for the construction of a biomedical symbol dictionary. The algorithms, applied to the MEDLINE document set, result in a high quality dictionary and toolset to disambiguate abbreviation symbols automatically.

Abbreviations as Topic↗

Predicting subcellular localization of proteins using machine-learned classifiers.

MOTIVATION: Identifying the destination or localization of proteins is key to understanding their function and facilitating their purification. A number of existing computational prediction methods are based on sequence analysis. However, these methods are limited in scope, accuracy and most particularly breadth of coverage. Rather than using sequence information alone, we have explored the use of database text annotations from homologs and machine learning to substantially improve the prediction of subcellular location. RESULTS: We have constructed five machine-learning classifiers for predicting subcellular localization of proteins from animals, plants, fungi, Gram-negative bacteria and Gram-positive bacteria, which are 81% accurate for fungi and 92-94% accurate for the other four categories. These are the most accurate subcellular predictors across the widest set of organisms ever published. Our predictors are part of the Proteome Analyst web-service.

Algorithms↗

Automated extraction of mutation data from the literature: application of MuteXt to G protein-coupled receptors and nuclear hormone receptors.

MOTIVATION: The amount of genomic and proteomic data that is published daily in the scientific literature is outstripping the ability of experimental scientists to stay current. Reviews, the traditional medium for collating published observations, are also unable to keep pace. For some specific classes of information (e.g. sequences and protein structures), obligatory data deposition policies have helped. However, a great deal of other valuable information is spread throughout the literature hindering coherent access. We are involved in the Molecular Class-Specific Information System (MCSIS) project, a collaborative effort to design and automate the maintenance of protein family databases. The first two databases, the GPCRDB and NucleaRDB, are focused on G protein-coupled receptors (GPCRs) and nuclear hormone receptors (NRs), respectively. The main aim of the MCSIS project is to gather heterogeneous data from across a variety of electronic and literature sources in order to draw new inferences about the target protein families. RESULTS: We present a computational method that identifies and extracts mutation data from the scientific literature. We focused on the extraction of single point mutations for the GPCR and NR superfamilies. After validation by plausibility filters, the mutation data is integrated into the corresponding MCSIS where it is combined with structural and sequence information already stored in these databases. We extracted and validated 2736 true point mutations from 914 articles on GPCRs and 785 true point mutations from 1094 articles on NRs. The current version of our automated extraction algorithm identifies 49.3% of the GPCR point mutations with a specificity of 87.9%, and 64.5% of the NR point mutations with a specificity of 85.8%. MuteXt routinely analyzes 100 electronic articles in approximately 1 h.

Algorithms↗

Ontologizing gene-expression microarray data: characterizing clusters with Gene Ontology.

An XML-based Java application is described that provides a function-oriented overview of the results of cluster analysis of gene-expression microarray data based on Gene Ontology terms and associations. The application generates one HTML page with listings of the frequencies of explicit and implicit Gene Ontology annotations for each cluster, and separate, linked pages with listings of explicit annotations for each gene in a cluster.

Cluster Analysis↗

CVD: the intestinal crypt/villus in situ hybridization database.

UNLABELLED: The intestinal crypt/villus in situ hybridization database (CVD) query interface is a web-based tool to search for genes with similar relative expression patterns along the crypt/villus axis of the mammalian intestine. The CVD is an online database holding information for relative gene expression patterns in the mammalian intestine and is based on the scoring of in situ hybridization experiments reported in the literature. CVD contains expression data for 88 different genes collected from 156 different in situ hybridization profiles. The web-based query interface allows execution of both single gene queries and pattern searches. The query results provide links to the most relevant public gene databases. AVAILABILITY: http://pc113.imbg.ku.dk/ps/

Animals↗

NetAffx Gene Ontology Mining Tool: a visual approach for microarray data analysis.

SUMMARY: The NetAffx Gene Ontology (GO) Mining Tool is a web-based, interactive tool that permits traversal of the GO graph in the context of microarray data. It accepts a list of Affymetrix probe sets and renders a GO graph as a heat map colored according to significance measurements. The rendered graph is interactive, with nodes linked to public web sites and to lists of the relevant probe sets. The GO Mining Tool provides visualization combining biological annotation with expression data, encompassing thousands of genes in one interactive view. AVAILABILITY: GO Mining Tool is freely available at http://www.affymetrix.com/analysis/query/go_analysis.affx

Abstracting and Indexing↗

PINdb: a database of nuclear protein complexes from human and yeast.

SUMMARY: Proteins Interacting in the Nucleus database (PINdb) is a database of protein complexes purified from the nucleus of human and yeast cells. It is compiled from the published literature and existing databases. Currently, PINdb contains mostly protein complexes that may be involved in gene transcription. To facilitate comparative analyses and identification of protein complexes, the compositional information is integrated with standardized gene nomenclature, annotation and protein sequences from public databases. The PINdb web interface provides a number of tools for (1) comparison of protein complexes, (2) search for a protein complex by its published name or by a partial list of its components and (3) browsing specific subsets or a functional classification of the complexes. Availablity: http://pin.mskcc.org

Abstracting and Indexing↗

'Harvester': a fast meta search engine of human protein resources.

SUMMARY: We have developed a Web-based tool named 'Harvester' that bulk-collects bioinformatic data on human proteins from various databases and prediction servers. The information on every single protein is assembled on a single HTML page as a combination of database screen-shots and plain text. A full text meta search engine, similar to Google trade mark, allows screening of the whole genome proteome for current protein functions and predictions in a few seconds. With Harvester it is now possible to compare and check the quality of different database entries and prediction algorithms on a single page. A feedback forum allows users to comment on Harvester and to report database inconsistencies. AVAILABILITY: The service is freely available to the academic community at http://harvester.embl.de.

Database Management Systems↗

A tool for gene expression based PubMed search through combining data sources.

UNLABELLED: We present a new tool for the semi-automated querying of PubMed using a batch of tens to thousands of GenBank accession numbers or UniGene cluster ids. By combining information from UniGene and SWISS-PROT, microGENIE obtains information on the biological relevance of expressed genes, as identified by micro-array experiments, with minimal user intervention and time investment. AVAILABILITY: microGENIE is freely available from http://www.cs.vu.nl/microgenie SUPPLEMENTARY INFORMATION: The web site above supplies examples of input and output files.

Database Management Systems↗

Ligand Depot: a data warehouse for ligands bound to macromolecules.

UNLABELLED: Ligand Depot is an integrated data resource for finding information about small molecules bound to proteins and nucleic acids. The initial release (version 1.0, November, 2003) focuses on providing chemical and structural information for small molecules found as part of the structures deposited in the Protein Data Bank. Ligand Depot accepts keyword-based queries and also provides a graphical interface for performing chemical substructure searches. A wide variety of web resources that contain information on small molecules may also be accessed through Ligand Depot. AVAILABILITY: Ligand Depot is available at http://ligand-depot.rutgers.edu/. Version 1.0 supports multiple operating systems including Windows, Unix, Linux and the Macintosh operating system. The current drawing tool works in Internet Explorer, Netscape and Mozilla on Windows, Unix and Linux.

Binding Sites↗

MedPost: a part-of-speech tagger for bioMedical text.

SUMMARY: We present a part-of-speech tagger that achieves over 97% accuracy on MEDLINE citations. AVAILABILITY: Software, documentation and a corpus of 5700 manually tagged sentences are available at ftp://ftp.ncbi.nlm.nih.gov/pub/lsmith/MedPost/medpost.tar.gz

Abstracting and Indexing↗

AliasServer: a web server to handle multiple aliases used to refer to proteins.

UNLABELLED: AliasServer provides services that facilitate the assembly of data or datasets that make use of different identifiers for refering to the same protein. This resource relies on a database which contains, for a given organism, a non-redundant list of protein sequences associated with a set of aliases. AVAILABILITY: AliasServer is available as an interactive Web server at http://cbi.labri.fr/outils/alias/ and as a web service using a SOAP interface. The complete tool, including sources and data, is available for local installations upon request. SUPPLEMENTARY INFORMATION: Technical documentation is available at http://cbi.labri.fr/outils/alias/asdoc.pdf

Algorithms↗

Pedro: a configurable data entry tool for XML.

UNLABELLED: Pedro is a Java application that dynamically generates data entry forms for data models expressed in XML Schema, producing XML data files that validate against this schema. The software uses an intuitive tree-based navigation system, can supply context-sensitive help to users and features a sophisticated interface for populating data fields with terms from controlled vocabularies. The software also has the ability to import records from tab delimited text files and features various validation routines. AVAILABILITY: The application, source code, example models from several domains and tutorials can be downloaded from http://pedro.man.ac.uk/.

Computer Graphics↗

Efficient combination of multiple word models for improved sequence comparison.

MOTIVATION: Studies of efficient and sensitive sequence comparison methods are driven by a need to find homologous regions of weak similarity between large genomes. RESULTS: We describe an improved method for finding similar regions between two sets of DNA sequences. The new method generalizes existing methods by locating word matches between sequences under two or more word models and extending word matches into high-scoring segment pairs (HSPs). The method is implemented as a computer program named DDS2. Experimental results show that DDS2 can find more HSPs by using several word models than by using one word model. AVAILABILITY: The DDS2 program is freely available for academic use in binary code form at http://bioinformatics.iastate.edu/aat/align/align.html and in source code form from the corresponding author.

Algorithms↗

Thermodynamics of enzyme-catalyzed reactions--a database for quantitative biochemistry.

UNLABELLED: The Thermodynamics of Enzyme-catalyzed Reactions Database (TECRDB) is a comprehensive collection of thermodynamic data on enzyme-catalyzed reactions. The data, which consist of apparent equilibrium constants and calorimetrically determined molar enthalpies of reaction, are the primary experimental results obtained from thermodynamic studies of biochemical reactions. The results from approximately 1000 published papers containing data on approximately 400 different enzyme-catalyzed reactions constitute the essential information in the database. The information is managed using Oracle and is available on the Web. AVAILABILITY: http://xpdb.nist.gov/enzyme_thermodynamics/

Biochemistry↗