PubMed Health⌕ Search

Biomedical subjects

George Hripcsak

Publications and source records attributed to George Hripcsak.

17 recordsLinked to original sources

Inherited Susceptibility to Urinary Tract Infections from Kidney Papilla to Bladder.

Urinary tract infections (UTIs) are traditionally viewed as environmentally driven, yet their inherited susceptibility remains largely unexplored. We conducted a cross-biobank genome-wide association study of recurrent UTIs in 1,860,836 individuals (213,869 cases and 1,646,967 controls). We identified 36 genetic susceptibility loci and performed tissue-based multi-omic mapping to prioritize candidate causal genes. UTI risk alleles preferentially modulated epithelial gene expression in kidney and bladder, converging on urinary epithelia structure and function. PSCA, encoding a secreted epithelial surface protein, emerged as the strongest candidate under genetic control; the gene product is constitutively secreted into the urine from kidney papilla and bladder epithelia, binds uropathogenic E. coli, and inhibits bacterial growth in vitro. Our findings define the polygenic architecture of UTIs and highlight the critical role of uroepithelial surface defenses, providing a new framework for host-directed, non-antibiotic interventions.

Journal Article↗

Genome-wide association meta-regression identifies stem cell lineage orchestration as a key driver of acne risk.

Over 85% of the population experience acne at some point in their lives, with its severity spanning a quantitative spectrum, from mild, transient outbreaks to more persistent, severe forms of the condition. Moderate to severe disease poses a substantial global burden arising from both the physical and psychological impacts of this highly visible condition. The analytical approach taken in this study aimed to address the impact of variation in the dichotomisation of acne case control status, driven by ascertainment and study design, on effect size estimates across independent genetic association studies of acne. Through a fixed intercept meta-regression framework, we combined evidence genome-wide for association with acne across studies in which case-control status had been ascertained in different settings, allowing for different severity threshold definitions. Across a combined sample of 73,997 cases and 1,103,940 controls of European, South Asian and African American ancestry we identify genetic variation at 165 genomic loci that influence acne risk. There is evidence for both shared and ancestry specific components to the genetic susceptibility to acne and for sex differences in the magnitude of effect of risk alleles at three loci. We observe that common genetic variation explains 13.4% of acne heritability on the liability scale. Consistent with the hypothesis that genetic risk primarily operates at the level of individual pilosebaceous units, a polygenic score derived from this case-control study of acne susceptibility is associated with both self-reported and clinically assessed acne severity in adolescence, further strengthening the link between genetic risk and disease severity. Prioritisation of causal genes at the identified acne risk loci, provides genetic validation of the targets of established and emerging acne therapies, including retinoid treatments. The identified acne risk loci are enriched for genes encoding downstream effectors of RXRA signalling, including SOX9 and components of the WNT and p53 pathways. Illustrating that the control of stem cell lineage plasticity and cellular fate are important mechanisms through which genetic variation influences acne susceptibility within the pilosebaceous unit.

Journal Article↗

The role of domain knowledge in automating medical text report classification.

OBJECTIVE: To analyze the effect of expert knowledge on the inductive learning process in creating classifiers for medical text reports. DESIGN: The authors converted medical text reports to a structured form through natural language processing. They then inductively created classifiers for medical text reports using varying degrees and types of expert knowledge and different inductive learning algorithms. The authors measured performance of the different classifiers as well as the costs to induce classifiers and acquire expert knowledge. MEASUREMENTS: The measurements used were classifier performance, training-set size efficiency, and classifier creation cost. RESULTS: Expert knowledge was shown to be the most significant factor affecting inductive learning performance, outweighing differences in learning algorithms. The use of expert knowledge can affect comparisons between learning algorithms. This expert knowledge may be obtained and represented separately as knowledge about the clinical task or about the data representation used. The benefit of the expert knowledge is more than that of inductive learning itself, with less cost to obtain. CONCLUSION: For medical text report classification, expert knowledge acquisition is more significant to performance and more cost-effective to obtain than knowledge discovery. Building classifiers should therefore focus more on acquiring knowledge from experts than trying to learn this knowledge inductively.

Algorithms↗

Measuring agreement in medical informatics reliability studies.

Agreement measures are used frequently in reliability studies that involve categorical data. Simple measures like observed agreement and specific agreement can reveal a good deal about the sample. Chance-corrected agreement in the form of the kappa statistic is used frequently based on its correspondence to an intraclass correlation coefficient and the ease of calculating it, but its magnitude depends on the tasks and categories in the experiment. It is helpful to separate the components of disagreement when the goal is to improve the reliability of an instrument or of the raters. Approaches based on modeling the decision making process can be helpful here, including tetrachoric correlation, polychoric correlation, latent trait models, and latent class models. Decision making models can also be used to better understand the behavior of different agreement metrics. For example, if the observed prevalence of responses in one of two available categories is low, then there is insufficient information in the sample to judge raters' ability to discriminate cases, and kappa may underestimate the true agreement and observed agreement may overestimate it.

Decision Making↗

Of truth and pathways: chasing bits of information through myriads of articles.

Knowledge on interactions between molecules in living cells is indispensable for theoretical analysis and practical applications in modern genomics and molecular biology. Building such networks relies on the assumption that the correct molecular interactions are known or can be identified by reading a few research articles. However, this assumption does not necessarily hold, as truth is rather an emerging property based on many potentially conflicting facts. This paper explores the processes of knowledge generation and publishing in the molecular biology literature using modelling and analysis of real molecular interaction data. The data analysed in this article were automatically extracted from 50000 research articles in molecular biology using a computer system called GeneWays containing a natural language processing module. The paper indicates that truthfulness of statements is associated in the minds of scientists with the relative importance (connectedness) of substances under study, revealing a potential selection bias in the reporting of research results. Aiming at understanding the statistical properties of the life cycle of biological facts reported in research articles, we formulate a stochastic model describing generation and propagation of knowledge about molecular interactions through scientific publications. We hope that in the future such a model can be useful for automatically producing consensus views of molecular interaction data.

Algorithms↗

Use of natural language processing to translate clinical information from a database of 889,921 chest radiographic reports.

PURPOSE: To evaluate translation of chest radiographic reports by using natural language processing and to compare the findings with those in the literature. MATERIALS AND METHODS: A natural language processor coded 10 years of narrative chest radiographic reports from an urban academic medical center. Coding for 150 reports was compared with manual coding. Frequencies and co-occurrences of 24 clinical conditions (diseases, abnormalities, and clinical states) were estimated. The ratio of right to left lung mass, association of pleural effusion with other conditions, and frequency of bullet and stab wounds were compared with independent observations. The sensitivity and specificity of the system's pneumothorax coding were compared with those of manual financial coding. RESULTS: The system coded 889,921 reports on 251,186 patients. On the basis of manual coding of 150 reports, the processor's sensitivity (0.81) and specificity (0.99) were comparable to those previously reported for natural language processing and for expert coders. The frequencies of the selected conditions ranged from 0.22 for pleural effusion to 0.0004 for tension pneumothorax. The database confirmed earlier observations that lung cancer occurs in a 3:2 right-to-left ratio. The association of pleural effusion with other conditions mirrored that in the literature. Bullet and stab wounds decreased during 10 years at a rate consistent with crime statistics. A review of pneumothorax cases showed that the database (sensitivity, 1.00; specificity, 0.996) was more accurate than financial discharge coding (sensitivity, 0.17; P =.002; specificity, 0.996; not significant). CONCLUSION: Internal and external validation in this study confirmed the accuracy of natural language processing for translating chest radiographic narrative reports into a large database of information.

Databases, Factual↗

A comparison of the Charlson comorbidities derived from medical language processing and administrative data.

The objective of this study was to develop a medical language processing (MLP) system, which consisted of MedLEE and a set of inference rules, to identify 19 Charlson comorbidities from discharge summaries and chest x-ray reports. We used 233 cases to learn the patterns that were indicative of comorbidities for developing the inference rules. We then used an independent data set of 3,662 pneumonia patients to identify comorbidities by MLP compared with administrative data (ICD-9 codes). A stratified random sample of 190 records from disagreement cases was manually reviewed. The sensitivity, specificity, and accuracy for the MLP system/ICD-9 codes in this testing set were 0.84/0.16, 0.70/0.30, and 0.77/0.23 respectively. Thirteen of the 19 comorbidities studied were underreported in the administrative data. The kappa values ranged from 0.19 for peptic ulcer to 0.70 for lymphoma. We conclude that comorbidities derived from natural language processing of medical records can improve ICD-9-based approaches.

Adult↗

Representing nested semantic information in a linear string of text using XML.

XML has been widely adopted as an important data interchange language. The structure of XML enables sharing of data elements with variable degrees of nesting as long as the elements are grouped in a strict tree-like fashion. This requirement potentially restricts the usefulness of XML for marking up written text, which often includes features that do not properly nest within other features. We encountered this problem while marking up medical text with structured semantic information from a Natural Language Processor. Traditional approaches to this problem separate the structured information from the actual text mark up. This paper introduces an alternative solution, which tightly integrates the semantic structure with the text. The resulting XML markup preserves the linearity of the medical texts and can therefore be easily expanded with additional types of information.

Programming Languages↗

The effect of sample size and disease prevalence on supervised machine learning of narrative data.

This paper examines the independent effects of outcome prevalence and training sample sizes on inductive learning performance. We trained 3 inductive learning algorithms (MC4, IB, and Naïve-Bayes) on 60 simulated datasets of parsed radiology text reports labeled with 6 disease states. Data sets were constructed to define positive outcome states at 4 prevalence rates (1, 5, 10, 25, and 50%) in training set sizes of 200 and 2,000 cases. We found that the effect of outcome prevalence is significant when outcome classes drop below 10% of cases. The effect appeared independent of sample size, induction algorithm used, or class label. Work is needed to identify methods of improving classifier performance when output classes are rare.

Algorithms↗

The sublanguage of cross-coverage.

At Columbia-Presbyterian Medical Center, free-text "Signout" notes are typed into the electronic record by clinicians for the purpose of cross-coverage. We plan to "unlock" information about adverse events contained in these notes in a subsequent project using Natural Language Processing (NLP). To better understand the requirements for parsing, Signout notes were compared to other common medical notes (ambulatory clinic notes and discharge summaries) on a series of quantitative metrics. They are shorter (mean length 59.25 words vs. 144.11 and 340.85 for ambulatory and discharge notes respectively) and use more abbreviations (26.88% vs. 20.07% and 3.57%). Despite being terser, Signout notes use less ambiguous abbreviations (8.34% vs. 9.09% and 18.02%). Differences were found using Relative Entropy and Squared Chi-square Distance in a novel fashion to compare these medical corpora. Signout notes appear to constitute a unique sublanguage of medicine. The implications for parsing free-text cross-coverage notes into coded medical data are discussed.

Linguistics↗

Reference standards, judges, and comparison subjects: roles for experts in evaluating system performance.

Medical informatics systems are often designed to perform at the level of human experts. Evaluation of the performance of these systems is often constrained by lack of reference standards, either because the appropriate response is not known or because no simple appropriate response exists. Even when performance can be assessed, it is not always clear whether the performance is sufficient or reasonable. These challenges can be addressed if an evaluator enlists the help of clinical domain experts. 1) The experts can carry out the same tasks as the system, and then their responses can be combined to generate a reference standard. 2)The experts can judge the appropriateness of system output directly. 3) The experts can serve as comparison subjects with which the system can be compared. These are separate roles that have different implications for study design, metrics, and issues of reliability and validity. Diagrams help delineate the roles of experts in complex study designs.

Evaluation Studies as Topic↗

Columbia University's Informatics for Diabetes Education and Telemedicine (IDEATel) project: technical implementation.

The Columbia University Informatics for Diabetes Education and Telemedicine IDEATel) project is a four-year demonstration project funded by the Centers for Medicare and Medicaid Services with the overall goal of evaluating the feasibility, acceptability, effectiveness, and cost-effectiveness of telemedicine. The focal point of the intervention is the home telemedicine unit (HTU), which provides four functions: synchronous videoconferencing over standard telephone lines, electronic transmission for fingerstick glucose and blood pressure readings, secure Web-based messaging and clinical data review, and access to Web-based educational materials. The HTU must be usable by elderly patients with no prior computer experience. Providing these functions through the HTU requires tight integration of six components: the HTU itself, case management software, a clinical information system, Web-based educational material, data security, and networking and telecommunications. These six components were integrated through a variety of interfaces, providing a system that works well for patients and providers. With more than 400 HTUs installed, IDEATel has demonstrated the feasibility of large-scale home telemedicine.

Case Management↗

Columbia University's Informatics for Diabetes Education and Telemedicine (IDEATel) Project: rationale and design.

The Columbia University Informatics for Diabetes Education and Telemedicine (IDEATel) Project is a four-year demonstration project funded by the Centers for Medicare and Medicaid Services with the overall goals of evaluating the feasibility, acceptability, effectiveness, and cost-effectiveness of telemedicine in the management of older patients with diabetes. The study is designed as a randomized controlled trial and is being conducted by a state-wide consortium in New York. Eligibility requires that participants have diabetes, are Medicare beneficiaries, and reside in federally designated medically underserved areas. A total of 1,500 participants will be randomized, half in New York City and half in other areas of the state. Intervention participants receive a home telemedicine unit that provides synchronous videoconferencing with a project-based nurse, electronic transmission of home fingerstick glucose and blood pressure data, and Web access to a project Web site. End points include glycosylated hemoglobin, blood pressure, and lipid levels; patient satisfaction; health care service utilization; and costs. The project is intended to provide data to help inform regulatory and reimbursement policies for electronically delivered health care services.

Case Management↗

Mapping abbreviations to full forms in biomedical articles.

OBJECTIVE: To develop methods that automatically map abbreviations to their full forms in biomedical articles. METHODS: The authors developed two methods of mapping defined and undefined abbreviations (defined abbreviations are paired with their full forms in the articles, whereas undefined ones are not). For defined abbreviations, they developed a set of pattern-matching rules to map an abbreviation to its full form and implemented the rules into a software program, AbbRE (for "abbreviation recognition and extraction"). Using the opinions of domain experts as a reference standard, they evaluated the recall and precision of AbbRE for defined abbreviations in ten biomedical articles randomly selected from the ten most frequently cited medical and biological journals. They also measured the percentage of undefined abbreviations in the same set of articles, and they investigated whether they could map undefined abbreviations to any of four public abbreviation databases (GenBank LocusLink, SWISSPROT, LRABR of the UMLS Specialist Lexicon, and BioABACUS). RESULTS: AbbRE had an average 0.70 recall and 0.95 precision for the defined abbreviations. The authors found that an average of 25 percent of abbreviations were defined in biomedical articles and that of a randomly selected subset of undefined abbreviations, 68 percent could be mapped to any of four abbreviation databases. They also found that many abbreviations are ambiguous (i.e., they map to more than one full form in abbreviation databases). CONCLUSION: AbbRE is efficient for mapping defined abbreviations. To couple AbbRE with abbreviation databases for the mapping of undefined abbreviations, not only exhaustive abbreviation databases but also a method to resolve the ambiguity of abbreviations in the databases are needed.

Abbreviations as Topic↗

Design and analysis of controlled trials in naturally clustered environments: implications for medical informatics.

In medical informatics research, study questions frequently involve individuals who are grouped into clusters. For example, an intervention may be aimed at a clinician (who treats a cluster of patients) with the intention of improving the health of individual patients. Correlation among individuals within a cluster can lead to incorrect estimates of the sample size required to detect an effect and inappropriate estimates of the confidence intervals and the statistical significance of the intervention effects. Contamination, which is the spread of the effect of an intervention or control treatment to the opposite group, often occurs between individuals within clusters. It leads to an attenuation of the effect of the intervention and reduced power to detect a difference. If individuals are randomized in a clinical trial (individual-randomized trial), then correlation must be taken into account in the analysis, and the sample size may need to be increased to compensate for contamination. Randomizing clusters rather than individuals (cluster-randomized trials) can eliminate contamination and may be preferred for logistical reasons. Cluster-randomized trials are generally less efficient than individual-randomized trials, so the tradeoffs must be assessed. Correlation must be taken into account in the analysis and in the sample-size calculations for cluster-randomized trials.

Cluster Analysis↗

Detecting adverse events using information technology.

CONTEXT: Although patient safety is a major problem, most health care organizations rely on spontaneous reporting, which detects only a small minority of adverse events. As a result, problems with safety have remained hidden. Chart review can detect adverse events in research settings, but it is too expensive for routine use. Information technology techniques can detect some adverse events in a timely and cost-effective way, in some cases early enough to prevent patient harm. OBJECTIVE: To review methodologies of detecting adverse events using information technology, reports of studies that used these techniques to detect adverse events, and study results for specific types of adverse events. DESIGN: Structured review. METHODOLOGY: English-language studies that reported using information technology to detect adverse events were identified using standard techniques. Only studies that contained original data were included. MAIN OUTCOME MEASURES: Adverse events, with specific focus on nosocomial infections, adverse drug events, and injurious falls. RESULTS: Tools such as event monitoring and natural language processing can inexpensively detect certain types of adverse events in clinical databases. These approaches already work well for some types of adverse events, including adverse drug events and nosocomial infections, and are in routine use in a few hospitals. In addition, it appears likely that these techniques will be adaptable in ways that allow detection of a broad array of adverse events, especially as more medical information becomes computerized. CONCLUSION: Computerized detection of adverse events will soon be practical on a widespread basis.

Accidental Falls↗