PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Natural Language Processing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 757 records · Page 42Linked to original sources

Implementations of clinical functional magnetic resonance imaging using character-based paradigms for the prediction of Chinese language dominance.

PURPOSE: Recently, functional MRI (fMRI) using word generation (WG) tasks has been shown to be effective for mapping the Chinese language-related brain areas. In clinical applications, however, patients' performance cannot be easily monitored during WG tasks. In this study, we evaluated the feasibility of a word choice (WC) paradigm in the clinical setting and compared the results with those from WG tasks. METHOD: Intrasubject comparisons of fMRI with both WG and WC paradigms were performed on six normal human subjects and two tumor patients. Subject responses in the WC paradigm, based on semantic judgments, were recorded. Activation strength, extent, and laterality were evaluated and compared. RESULTS: Our results showed that fMRI with the WC paradigm evoked weaker neuronal activation than that with the WG paradigm in Chinese language-related brain areas. It was sufficient to reveal language laterality for clinical use, however. In addition, it resulted in less nonlanguage-specific brain activation. CONCLUSION: Results from the patient data demonstrated strong evidence for the necessity of incorporating response monitoring during fMRI studies, which suggested that fMRI with the WC paradigm is more appropriate to be implemented for the prediction of Chinese language dominance in clinical environments.

Adult↗

Word recognition software use in a busy orthopaedic practice.

For the past 3 years, our orthopaedic office has used word recognition programs to expedite medical record reporting, operative dictation, and letter writing. The use of these programs has resulted in cost savings, increased efficiency, and an improvement in the quality and the thoroughness of our medical records. The methods whereby we were able to incorporate a word recognition program into our operative dictations and our office notes are discussed and explained. Advantages and disadvantages of different programs were evaluated and discussed. We have been able to use word recognition programs in the daily function of our orthopaedic office although not without some difficulties. These programs have resulted in cost savings and time savings and provided medicolegal protection because of the thoroughness of the documentation. The use of templates allows efficiency comparable with conventional dictation. Word recognition programs can be used effectively in busy orthopaedic offices when combined with note templates.

Computer User Training↗

Working words: real-life lexicon of North American workers.

OBJECTIVE: This study describes a new computer methodology for analyzing workers' free text work descriptions. METHODS: Computerized lexical analysis was applied to work descriptions of participants in the Lung Health Study, a smoking-cessation study in persons with early chronic obstructive pulmonary disease. Text was parsed and analyzed as single term roots and pairs of roots commonly occurring together. RESULTS: The frequencies of terms reflect the work of a population; our subjects' most frequently used terms included "sale, office, service, business, engine[er], secretary, construct, driv[e], comput[e], teach, truck." Standard classification schemes (NAICS and SOC) and textbooks use terms inconsistent with those of actual workers. Many common empirical terms imply both industry and job information content, although traditional coding schemes separate industry and job title. CONCLUSIONS: Formal analyses of language may facilitate communication, identify translation priorities, and allow automated work coding.

Female↗

The emergence of communication in evolutionary robots.

Evolutionary robotics is a biologically inspired approach to robotics that is advantageous to studying the evolution of communication. A new model for the emergence of communication is developed and tested through various simulation experiments. In the first simulation, the emergence of simple signalling behaviour is studied. This is used to investigate the inter-relationships between communication abilities, namely linguistic production and comprehension, and other behavioural skills. The model supports the hypothesis that the ability to form categories from direct interaction with an environment constitutes the grounds for subsequent evolution of communication and language. In the second simulation, evolutionary robots are used to study the emergence of simple syntactic categories, e.g. action names (verbs). Comparisons between the two simulations indicate that the signalling lexicon emerged in the first simulation follows the evolutionary pattern of nouns, as observed in related models on the evolution of syntactic categories. Results also support the language-origin hypothesis on the fact that nouns precede verbs in both phylogenesis and ontogenesis. Further extensions of this new evolutionary robotic model for testing hypotheses on language origins are also discussed.

Adaptation, Physiological↗

Solvable null model for the distribution of word frequencies.

Zipf's law asserts that in all natural languages the frequency of a word is inversely proportional to its rank. The significance, if any, of this result for language remains a mystery. Here we examine a null hypothesis for the distribution of word frequencies, a so-called discourse-triggered word choice model, which is based on the assumption that the more a word is used, the more likely it is to be used again. We argue that this model is equivalent to the neutral infinite-alleles model of population genetics and so the degeneracy of the different words composing a sample of text is given by the celebrated Ewens sampling formula [Theor. Pop. Biol. 3, 87 (1972)]], which we show to produce an exponential distribution of word frequencies.

Algorithms↗

Euclidean distance between syntactically linked words.

We study the Euclidean distance between syntactically linked words in sentences. The average distance is significantly small and is a very slowly growing function of sentence length. We consider two nonexcluding hypotheses: (a) the average distance is minimized and (b) the average distance is constrained. Support for (a) comes from the significantly small average distance real sentences achieve. The strength of the minimization hypothesis decreases with the length of the sentence. Support for (b) comes from the very slow growth of the average distance versus sentence length. Furthermore, (b) predicts, under ideal conditions, an exponential distribution of the distance between linked words, a trend that can be identified in real sentences.

Artificial Intelligence↗

Development of speechreading supplements based on automatic speech recognition.

In manual-cued speech (MCS) a speaker produces hand gestures to resolve ambiguities among speech elements that are often confused by speechreaders. The shape of the hand distinguishes among consonants; the position of the hand relative to the face distinguishes among vowels. Experienced receivers of MCS achieve nearly perfect reception of everyday connected speech. MCS has been taught to very young deaf children and greatly facilitates language learning, communication, and general education. This manuscript describes a system that can produce a form of cued speech automatically in real time and reports on its evaluation by trained receivers of MCS. Cues are derived by a hidden markov models (HMM)-based speaker-dependent phonetic speech recognizer that uses context-dependent phone models and are presented visually by superimposing animated handshapes on the face of the talker. The benefit provided by these cues strongly depends on articulation of hand movements and on precise synchronization of the actions of the hands and the face. Using the system reported here, experienced cue receivers can recognize roughly two-thirds of the keywords in cued low-context sentences correctly, compared to roughly one-third by speechreading alone (SA). The practical significance of these improvements is to support fairly normal rates of reception of conversational speech, a task that is often difficult via SA.

Adult↗

Comparison of two schemes for automatic keyword extraction from MEDLINE for functional gene clustering.

One of the key challenges of microarray studies is to derive biological insights from the unprecedented quatities of data on gene-expression patterns. Clustering genes by functional keyword association can provide direct information about the nature of the functional links among genes within the derived clusters. However, the quality of the keyword lists extracted from biomedical literature for each gene significantly affects the clustering results. We extracted keywords from MEDLINE that describes the most prominent functions of the genes, and used the resulting weights of the keywords as feature vectors for gene clustering. By analyzing the resulting cluster quality, we compared two keyword weighting schemes: normalized z-score and term frequency-inverse document frequency (TFIDF). The best combination of background comparison set, stop list and stemming algorithm was selected based on precision and recall metrics. In a test set of four known gene groups, a hierarchical algorithm correctly assigned 25 of 26 genes to the appropriate clusters based on keywords extracted by the TDFIDF weighting scheme, but only 23 og 26 with the z-score method. To evaluate the effectiveness of the weighting schemes for keyword extraction for gene clusters from microarray profiles, 44 yeast genes that are differentially expressed during the cell cycle were used as a second test set. Using established measures of cluster quality, the results produced from TFIDF-weighted keywords had higher purity, lower entropy, and higher mutual information than those produced from normalized z-score weighted keywords. The optimized algorithms should be useful for sorting genes from microarray lists into functionally discrete clusters.

Artificial Intelligence↗

AZuRE, a scalable system for automated term disambiguation of gene and protein names.

Researchers, hindered by a lack of standard gene and protein-naming conventions, endure long, sometimes fruitless, literature searches. A system is described which is able to automatically assign gene names to their LocusLink ID (LLID) in previously unseen MEDLINE abstracts. The system is based on supervised learning and builds a model for each LLID. The training sets for all LLIDs are extracted automatically from MEDLINE references in the LocusLink and SwissProt databases. A validation was done of the performance for all 20,546 human genes with LLIDs. Of these, 7,344 produced good quality models (F-measure > 0.7, nearly 60% of which were > 0.9) and 13,202 did not, mainly due to insufficient numbers of known document references. A hand validation of MEDLINE documents for a set of 66 genes agreed well with the system's internal accuracy assessment. It is concluded that it is possible to achieve high quality gene disambiguation using scaleable automated techniques.

Databases, Protein↗

Comparative analysis of gene sets in the Gene Ontology space under the multiple hypothesis testing framework.

The Gene Ontology (GO) resource can be used as a powerful tool to uncover the properties shared among, and specific to, a list of genes produced by high-throughput functional genomics studies, such as microarray studies. In the comparative analysis of several gene lists, researchers maybe interested in knowing which GO terms are enriched in one list of genes but relatively depleted in another. Statistical tests such as Fisher's exact test or Chi-square test can be performed to search for such GO terms. However, because multiple GO terms are tested simultaneously, individual p-values from individual tests do not serve as good indicators for picking GO terms. Furthermore, these multiple tests are highly correlated, usual multiple testing procedures that work under an independence assumption are not applicable. In this paper we introduce a procedure, based on False Discovery Rate (FDR), to treat this correlated multiple testing problem. This procedure calculates a moderately conserved estimator of q-value for every GO term. We identify the GO terms with q-values that satisfy a desired level as the significant GO terms. This procedure has been implemented into the GoSurfer software. GoSurfer is a windows based graphical data mining tool. It is freely available at http://www.gosurfer.org.

Algorithms↗

Clustering genes using gene expression and text literature data.

Clustering of gene expression data is a standard technique used to identify closely related genes. In this paper, we develop a new clustering algorithm, MSC (Multi-Source Clustering), to perform exploratory analysis using two or more diverse sources of data. In particular, we investigate the problem of improving the clustering by integrating information obtained from gene expression data with knowledge extracted from biomedical text literature. In each iteration of algorithm MSC, an EM-type procedure is employed to bootstrap the model obtained from one data source by starting with the cluster assignments obtained in the previous iteration using the other data sources. Upon convergence, the two individual models are used to construct the final cluster assignment. We compare the results of algorithm MSC for two data sources with the results obtained when the clustering is applied on the two sources of data separately. We also compare it with that obtained using the feature level integration method that performs the clustering after simply concatenating the features obtained from the two data sources. We show that the z-scores of the clustering results from MSC are better than that from the other methods. To evaluate our clusters better, function enrichment results are presented using terms from the Gene Ontology database. Finally, by investigating the success of motif detection programs that use the clusters, we show that our approach integrating gene expression data and text data reveals clusters that are biologically more meaningful than those identified using gene expression data alone.

Artificial Intelligence↗

Investigation into biomedical literature classification using support vector machines.

Specific topic search in the PubMed Database, one of the most important information resources for scientific community, presents a big challenge to the users. The researcher typically formulates boolean queries followed by scanning the retrieved records for relevance, which is very time consuming and error prone. We applied Support Vector Machines (SVM) for automatic retrieval of PubMed articles related to Human genome epidemiological research at CDC (Center for disease Control and Prevention). In this paper, we discuss various investigations into biomedical literature classification and analyze the effect of various issues related to the choice of keywords, training sets, kernel functions and parameters for the SVM technique. We report on the various factors above to show that SVM is a viable technique for automatic classification of biomedical literature into topics of interest such as epidemiology, cancer, birth defects etc. In all our experiments, we achieved high values of PPV, sensitivity and specificity.

Abstracting and Indexing↗

On the building of binary spelling interfaces for augmentative communication.

A criterion for design of optimal binary spelling interfaces (SI)--the average expectation of number of writing steps (trials) required to write one letter--is presented and discussed. This criterion is relevant for practical usage of any menu-oriented alternative communication system, mechanical device or brain-computer-interface, when a user (typically, a patient with devastating neuromuscular handicap) can not create an intended single binary response with an absolute reliability. An algorithm for building of a corresponding binary tree is developed and evaluated. This algorithm is efficient, when selection probabilities have essentially different values (the worst case).

Algorithms↗