PubMed Health⌕ Search

Biomedical subjects

J D Wren

Publications and source records attributed to J D Wren.

3 recordsLinked to original sources

Heuristics for identification of acronym-definition patterns within text: towards an automated construction of comprehensive acronym-definition dictionaries.

OBJECTIVES: To develop an automated, accurate and scalable method by which acronym-definition pairs can be identified within text. Its primary advantage is in enabling information processing methods to resolve author-defined acronyms, but it also allows an automated creation of a reference work on acronym definitions. This has several advantages over manual or semi-automated methods, besides time and effort saved, such as enabling identification of relative frequencies for alternate acronyms and definitions as well as spelling, phrasing and hyphenation variants for a unique acronym-definition pair. It also aids users in identifying acronym/definition variants present in the literature that may not necessarily be in biomedical databases. METHODS: A set of heuristics to accurately locate and identify the boundaries of acronym-definition pairs was developed and refined in terms of precision and recall on subsets of MEDLINE records. These training sets were gradually increased in size and heuristics re-evaluated to ensure scalability. RESULTS: Our final set of Acronym Resolving General Heuristics (ARGH) had a sample-based estimated rate of 96.5 +/- 0.4% precision and 93.0 +/- 2.7% recall when tested on over 12 million MEDLINE records, identifying more than 174,000 unique acronyms and their 737,000 associated definitions. CONCLUSIONS: We estimate that as much as 36% of the acronyms in MEDLINE are associated with more than one definition and, conversely, up to 10% of definitions are associated with more than one acronym. The number of unique acronyms in MEDLINE is increasing at a rate of approximately 11,000 per year, while the number of definitions associated with them is growing at approximately four times that rate. Access to the ARGH database is available online at http://lethargy.swmed.edu/ARGH/argh.asp. The heuristic module and database are available upon request.

Abbreviations as Topic↗

Searching for microsatellite mutations in coding regions in lung, breast, ovarian and colorectal cancers.

RepX represents a new informatics approach to probe the UniGene database for potentially polymorphic repeat sequences in the open reading frame (ORF) of genes, 56% of which were found to be actually polymorphic. We now have performed mutational analysis of 17 such sites in genes not found to be polymorphic (<0.03 frequency) in a large panel of human cancer genomic DNAs derived from 31 lung, 21 breast, seven ovarian, 21 (13 microsatellite instability (MSI)+ and eight MSI-) colorectal cancer cell lines. In the lung, breast and ovarian tumor DNAs we found no mutations (<0.03-0.04 rate of tumor associated open reading frame mutations) in these sequences. By contrast, 18 MSI+ colorectal cancers (13 cancer cell lines and five primary tumors) with mismatch repair defects exhibited six mutations in three of the 17 genes (SREBP-2, TAN-1, GR6) (P<0.000003 compared to all other cancers tested). We conclude that coding region microsatellite alterations are rare in lung, breast, ovarian carcinomas and MSI (-) colorectal cancers, but are relatively frequent in MSI (+) colorectal cancers with mismatch repair deficits.

Base Pair Mismatch↗

Repeat polymorphisms within gene regions: phenotypic and evolutionary implications.

We have developed an algorithm that predicted 11,265 potentially polymorphic tandem repeats within transcribed sequences. We estimate that 22% (2,207/9,717) of the annotated clusters within UniGene contain at least one potentially polymorphic locus. Our predictions were tested by allelotyping a panel of approximately 30 individuals for 5% of these regions, confirming polymorphism for more than half the loci tested. Our study indicates that tandem-repeat polymorphisms in genes are more common than is generally believed. Approximately 8% of these loci are within coding sequences and, if polymorphic, would result in frameshifts. Our catalogue of putative polymorphic repeats within transcribed sequences comprises a large set of potentially phenotypic or disease-causing loci. In addition, from the anomalous character of the repetitive sequences within unannotated clusters, we also conclude that the UniGene cluster count substantially overestimates the number of genes in the human genome. We hypothesize that polymorphisms in repeated sequences occur with some baseline distribution, on the basis of repeat homogeneity, size, and sequence composition, and that deviations from that distribution are indicative of the nature of selection pressure at that locus. We find evidence of selective maintenance of the ability of some genes to respond very rapidly, perhaps even on intragenerational timescales, to fluctuating selective pressures.

Algorithms↗