PubMed Health⌕ Search

Biomedical subjects

Qing T Zeng

Publications and source records attributed to Qing T Zeng.

12 recordsLinked to original sources

Extracting principal diagnosis, co-morbidity and smoking status for asthma research: evaluation of a natural language processing system.

BACKGROUND: The text descriptions in electronic medical records are a rich source of information. We have developed a Health Information Text Extraction (HITEx) tool and used it to extract key findings for a research study on airways disease. METHODS: The principal diagnosis, co-morbidity and smoking status extracted by HITEx from a set of 150 discharge summaries were compared to an expert-generated gold standard. RESULTS: The accuracy of HITEx was 82% for principal diagnosis, 87% for co-morbidity, and 90% for smoking status extraction, when cases labeled "Insufficient Data" by the gold standard were excluded. CONCLUSION: We consider the results promising, given the complexity of the discharge summaries and the extraction tasks.

Asthma↗

Feature-guided clustering of multi-dimensional flow cytometry datasets.

BACKGROUND: Flow cytometry produces large multi-dimensional datasets of the physical and molecular characteristics of individual cells. The objective of this study was to simplify the cytometry datasets by arranging or clustering "objects" (cells) into a smaller number of relatively homogeneous groups (clusters) on the basis of interobject similarities and dissimilarities. RESULTS: The algorithm was designed to be driven by histogram features; that is, the relevant single parameter histogram features were used to guide multidimensional k-means clustering without an a priori estimate of cluster number. To test this approach, we simulated cell-derived datasets using protein-coated microspheres (artificial "cells"). The microspheres were constructed to provide 119 populations in 40 samples. The feature-guided (FG) approach accurately identified 100% of the predetermined cluster combinations. In contrast, an approach based on the partition index (PI) cluster validity measure accurately identified 83.2% of the clusters. Direct comparisons of the two methods indicated that the FG method was significantly more accurate than PI in identifying both the number of clusters and the number of objects within the clusters (p<.0001). CONCLUSION: We conclude that parameter feature analysis can be used to effectively guide k-means clustering of flow cytometry datasets.

Algorithms↗

Exploring and developing consumer health vocabularies.

Laypersons ("consumers") often have difficulty finding, understanding, and acting on health information due to gaps in their domain knowledge. Ideally, consumer health vocabularies (CHVs) would reflect the different ways consumers express and think about health topics, helping to bridge this vocabulary gap. However, despite the recent research on mismatches between consumer and professional language (e.g., lexical, semantic, and explanatory), there have been few systematic efforts to develop and evaluate CHVs. This paper presents the point of view that CHV development is practical and necessary for extending research on informatics-based tools to facilitate consumer health information seeking, retrieval, and understanding. In support of the view, we briefly describe a distributed, bottom-up approach for (1) exploring the relationship between common consumer health expressions and professional concepts and (2) developing an open-access, preliminary (draft) "first-generation" CHV. While recognizing the limitations of the approach (e.g., not addressing psychosocial and cultural factors), we suggest that such exploratory research and development will yield insights into the nature of consumer health expressions and assist developers in creating tools and applications to support consumer health information seeking.

Health↗

Assisting consumer health information retrieval with query recommendations.

OBJECTIVE: Health information retrieval (HIR) on the Internet has become an important practice for millions of people, many of whom have problems forming effective queries. We have developed and evaluated a tool to assist people in health-related query formation. DESIGN: We developed the Health Information Query Assistant (HIQuA) system. The system suggests alternative/additional query terms related to the user's initial query that can be used as building blocks to construct a better, more specific query. The recommended terms are selected according to their semantic distance from the original query, which is calculated on the basis of concept co-occurrences in medical literature and log data as well as semantic relations in medical vocabularies. MEASUREMENTS: An evaluation of the HIQuA system was conducted and a total of 213 subjects participated in the study. The subjects were randomized into 2 groups. One group was given query recommendations and the other was not. Each subject performed HIR for both a predefined and a self-defined task. RESULTS: The study showed that providing HIQuA recommendations resulted in statistically significantly higher rates of successful queries (odds ratio = 1.66, 95% confidence interval = 1.16-2.38), although no statistically significant impact on user satisfaction or the users' ability to accomplish the predefined retrieval task was found. CONCLUSION: Providing semantic-distance-based query recommendations can help consumers with query formation during HIR.

Adult↗

Identifying consumer-friendly display (CFD) names for health concepts.

We have developed a systematic methodology using corpus-based text analysis followed by human review to assign "consumer-friendly display (CFD) names" to medical concepts from the National Library of Medicine (NLM) Unified Medical Language System (UMLS) Metathesaurus. Using NLM MedlinePlus queries as a corpus of consumer expressions and a collaborative Web-based tool to facilitate review, we analyzed 425 frequently occurring concepts. As a preliminary test of our method, we evaluated 34 ana-lyzed concepts and their CFD names, using a questionnaire modeled on standard reading assessments. The initial results that consumers (n=10) are more likely to understand and recognize CFD names than alternate labels suggest that the approach is useful in the development of consumer health vocabularies for displaying understandable health information.

Community Participation↗

Reformulation of consumer health queries with professional terminology: a pilot study.

BACKGROUND: The Internet is becoming an increasingly important resource for health-information seekers. However, consumers often do not use effective search strategies. Query reformulation is one potential intervention to improve the effectiveness of consumer searches. OBJECTIVE: We endeavored to answer the research question: "Does reformulating original consumer queries with preferred terminology from the Unified Medical Language System (UMLS) Metathesaurus lead to better search returns?" METHODS: Consumer-generated queries with known goals (n=16) that could be mapped to UMLS Metathesaurus terminology were used as test samples. Reformulated queries were generated by replacing user terms with Metathesaurus-preferred synonyms (n=18). Searches (n=36) were performed using both a consumer information site and a general search engine. Top 30 precision was used as a performance indicator to compare the performance of the original and reformulated queries. RESULTS: Forty-two percent of the searches utilizing reformulated queries yielded better search returns than their associated original queries, 19% yielded worse results, and the results for the remaining 39% did not change. We identified ambiguous lay terms, expansion of acronyms, and arcane professional terms as causes for changes in performance. CONCLUSIONS: We noted a trend towards increased precision when providing substitutions for lay terms, abbreviations, and acronyms. We have found qualitative evidence that reformulating queries with professional terminology may be a promising strategy to improve consumer health-information searches, although we caution that automated reformulation could in fact worsen search performance when the terminology is ill-fitted or arcane.

Consumer Behavior↗

Positive attitudes and failed queries: an exploration of the conundrums of consumer health information retrieval.

Several studies have found that consumers report a high level of satisfaction with the Internet as a health information resource. Belied by this positive attitude, however, are other studies reporting that consumers were often unsuccessful in searching for health information. In this paper, we present an interview and observation study in which we asked health consumers to search for health information on the Internet after first stating their search goals. Upon the conclusion of the session they were asked to evaluate their searches. We found that many consumers were unable to find satisfactory information when performing a specific query, while in general the group viewed health information retrieval (HIR) on the Internet in a positive light. We analyzed the observed search sessions to determine what factors accounted for the failure of specific searches and positive attitudes, and also discussed potential informatics solutions.

Adult↗

GLIF3: a representation format for sharable computer-interpretable clinical practice guidelines.

The Guideline Interchange Format (GLIF) is a model for representation of sharable computer-interpretable guidelines. The current version of GLIF (GLIF3) is a substantial update and enhancement of the model since the previous version (GLIF2). GLIF3 enables encoding of a guideline at three levels: a conceptual flowchart, a computable specification that can be verified for logical consistency and completeness, and an implementable specification that is intended to be incorporated into particular institutional information systems. The representation has been tested on a wide variety of guidelines that are typical of the range of guidelines in clinical use. It builds upon GLIF2 by adding several constructs that enable interpretation of encoded guidelines in computer-based decision-support systems. GLIF3 leverages standards being developed in Health Level 7 in order to allow integration of guidelines with clinical information systems. The GLIF3 specification consists of an extensible object-oriented model and a structured syntax based on the resource description framework (RDF). Empirical validation of the ability to generate appropriate recommendations using GLIF3 has been tested by executing encoded guidelines against actual patient data. GLIF3 is accordingly ready for broader experimentation and prototype use by organizations that wish to evaluate its ability to capture the logic of clinical guidelines, to implement them in clinical systems, and thereby to provide integrated decision support to assist clinicians.

Artificial Intelligence↗

Relationships among different subjective measurements of consumer health information retrieval performance.

BACKGROUND: Millions of consumers perform health information retrieval (HIR) online. To better understand the consumers' perspective on HIR performance, we conducted an observation and interview study of 97 health information consumers. METHODS: Consumers were asked to perform HIR tasks and we recorded their view regarding performance using several differ-ent subjective measurements: finding the desired information, usefulness of the information found, satisfaction with the information, and intention to continue searching. Statistical analysis was applied to verify if the multiple subjective measurements were redundant. RESULT: The measurements ranged from slight agreement to no agreement among them. A number of reasons were identified for this lack of agreement. CONCLUSION: Although related, the four subjective measurements of HIR performance are distinct from each other and carried different useful information

Consumer Behavior↗

Coverage of patient safety terms in the UMLS metathesaurus.

The integration and large-scale analyses of medical error databases would be greatly facilitated by the use of a standard terminology. We investigated the availability in the UMLS metathesaurus of concepts that are required for coding patient safety data. Terms from three proprietary patient safety terminologies were mapped to the concepts in UMLS by an automated mapping program developed by us. From these candidate mappings, the concept that matched its corresponding term was selected manually. The reliability of the mapping procedure was verified by manually searching for terms in the UMLS Knowledge Source Server. Matching concepts in UMLS were identified for less than 27% of the terms in the study dataset. The matching rates of terms that describe the type of error and the causes of errors were even lower. The lack of such terms in the existing standard terminologies underscores the need for development of a standard patient safety terminology.

Humans↗

A technique to improve the spelling suggestion rank in medical queries.

Correct spelling is crucial for online search engines to function well, and health information is highly sought after online. We propose a technique for increasing the effectiveness of spell-checking tools for use with medical queries. Our results show a marked improvement in the ranking of the correct term within the suggestion list returned by the spelling correction tool, as well as a lessening of the drawbacks associated with using larger dictionaries.

Dictionaries, Medical as Topic↗

Visual representation of cell subpopulation from flow cytometry data.

Flow cytometric systems are useful for protein identification and expression analysis, especially characterizing particular lineage or sublineage of cells. We clustered flow cytometry data of bone marrow cells into subpopulations using a clustering algorithm with its physical characteristics (cell size and cell granularity) and different molecular composition (cell reactivity with monoclonal antibodies). To display the cell subpopulations, we created a colored map according to the mean of 5 flow cytometry parameters based on a cluster. Such a map can reveal subpopulation properties that are not evident in the widely used scatter plot.

Algorithms↗