PubMed Health⌕ Search

Biomedical subjects

Ryan McDonald

Publications and source records attributed to Ryan McDonald.

4 recordsLinked to original sources

An automated procedure to identify biomedical articles that contain cancer-associated gene variants.

The proliferation of biomedical literature makes it increasingly difficult for researchers to find and manage relevant information. However, identifying research articles containing mutation data, a requisite first step in integrating large and complex mutation data sets, is currently tedious, time-consuming and imprecise. More effective mechanisms for identifying articles containing mutation information would be beneficial both for the curation of mutation databases and for individual researchers. We developed an automated method that uses information extraction, classifier, and relevance ranking techniques to determine the likelihood of MEDLINE abstracts containing information regarding genomic variation data suitable for inclusion in mutation databases. We targeted the CDKN2A (p16) gene and the procedure for document identification currently used by CDKN2A Database curators as a measure of feasibility. A set of abstracts was manually identified from a MEDLINE search as potentially containing specific CDKN2A mutation events. A subset of these abstracts was used as a training set for a maximum entropy classifier to identify text features distinguishing "relevant" from "not relevant" abstracts. Each document was represented as a set of indicative word, word pair, and entity tagger-derived genomic variation features. When applied to a test set of 200 candidate abstracts, the classifier predicted 88 articles as being relevant; of these, 29 of 32 manuscripts in which manual curation found CDKN2A sequence variants were positively predicted. Thus, the set of potentially useful articles that a manual curator would have to review was reduced by 56%, maintaining 91% recall (sensitivity) and more than doubling precision (positive predictive value). Subsequent expansion of the training set to 494 articles yielded similar precision and recall rates, and comparison of the original and expanded trials demonstrated that the average precision improved with the larger data set. Our results show that automated systems can effectively identify article subsets relevant to a given task and may prove to be powerful tools for the broader research community. This procedure can be readily adapted to any or all genes, organisms, or sets of documents.

Computational Biology↗

Automatically annotating documents with normalized gene lists.

BACKGROUND: Document gene normalization is the problem of creating a list of unique identifiers for genes that are mentioned within a document. Automating this process has many potential applications in both information extraction and database curation systems. Here we present two separate solutions to this problem. The first is primarily based on standard pattern matching and information extraction techniques. The second and more novel solution uses a statistical classifier to recognize valid gene matches from a list of known gene synonyms. RESULTS: We compare the results of the two systems, analyze their merits and argue that the classification based system is preferable for many reasons including performance, simplicity and robustness. Our best systems attain a balanced precision and recall in the range of 74%-92%, depending on the organism.

Animals↗

Identifying gene and protein mentions in text using conditional random fields.

BACKGROUND: We present a model for tagging gene and protein mentions from text using the probabilistic sequence tagging framework of conditional random fields (CRFs). Conditional random fields model the probability P(t/o) of a tag sequence given an observation sequence directly, and have previously been employed successfully for other tagging tasks. The mechanics of CRFs and their relationship to maximum entropy are discussed in detail. RESULTS: We employ a diverse feature set containing standard orthographic features combined with expert features in the form of gene and biological term lexicons to achieve a precision of 86.4% and recall of 78.7%. An analysis of the contribution of the various features of the model is provided.

Genes↗

The use of TaqMan PCR assay for detection of Bordetella pertussis infection from clinical specimens.

OBJECTIVE: The routine clinical laboratory detection of Bordetella pertussis is through culture, which can require 5 to 7 days for the bacteria to grow. Using a polymerase chain reaction (PCR) assay can shorten this detection time while increasing the sensitivity of detection with similar specificity. This study compared culture with TaqMan PCR for detection of B pertussis in clinical specimens and the turnaround time for each assay during the pertussis season. MATERIALS AND METHODS: Nasopharyngeal swabs in Regan-Lowe transport media were collected from 1556 persons who had symptoms of whooping cough or who had had contact with infected persons; the swabs were submitted for B pertussis detection during the pertussis season. A single nasopharyngeal swab from each patient was submitted for both culture and TaqMan PCR detection. Upon receipt of the specimens, the swabs were inoculated onto Regan-Lowe agar for culture and incubated for up to 7 days. The same swab was processed for PCR detection using TaqMan PCR assay. A second nested PCR was used on positive specimens for resolution purposes. The TaqMan PCR assay was performed 3 to 5 days a week, whereas the culture was performed 6 days a week. All specimens were processed on the same day or earliest possible working day for TaqMan or culture, and specimens queued for resolution by nested PCR were batched. RESULTS: There were a total of 275 PCR positives and 28 culture positives. After resolution with the second nested PCR, the sensitivity, specificity, positive predictive value, and negative predictive value were 100%, 97.4%, 87.6%, and 100% for TaqMan PCR and 11.6%, 100%, 100%, and 85.7% for culture. The average turnaround time for positive culture was 5.1 days, and the average turnaround time for PCR was 2.3 days. CONCLUSION: The TaqMan PCR assay has superior sensitivity and shorter turnaround time over culture because it can be finished within one working day. Furthermore, the same swab can be used for culture of the bacteria for antibiotic susceptibility testing. The early detection of pertussis using TaqMan PCR assay allows early intervention on the spread of the disease and the ability to culture the bacteria from the same swab, thereby eliminating the need for a second swab and allowing for detection of antibiotic resistance.

Bordetella pertussis↗