PubMed Health⌕ Search

Biomedical subjects

Lucila Ohno-Machado

Publications and source records attributed to Lucila Ohno-Machado.

At least 19 recordsLinked to original sources

Privacy-Enhancing Sequential Learning under Heterogeneous Selection Bias in Multi-Site EHR Data.

OBJECTIVE: To develop privacy-enhancing statistical methods for estimation of binary disease risk model association parameters across multiple electronic health record (EHR) sites with heterogeneous selection mechanisms, without sharing raw individual-level data. We illustrate their utility through a cross-biobank analysis of smoking and 97 cancer subtypes using data from the NIH All of Us (AOU) and the Michigan Genomics Initiative (MGI). MATERIALS AND METHODS: Large-scale biobanks often follow heterogeneous recruitment strategies and store data in separate cloud-based platforms, making centralized algorithms infeasible. To address this, we propose two decentralized sequential estimators namely, Sequential Pseudo-likelihood (SPL) and Sequential Augmented Inverse Probability Weighting (SAIPW) that leverage external population-level information to adjust for selection bias, with valid variance estimation. SAIPW additionally protects against misspecification of the selection model using flexible machine learning based auxiliary outcome models. We compare SPL and SAIPW with the existing Sequential Unweighted (SUW) estimator and with centralized and meta learning extensions of IPW and AIPW in simulations under both correctly specified and misspecified selection mechanisms. We apply the methods to harmonized data from MGI ( n = 50,935) and AOU ( n = 241,563) to estimate smoking-cancer associations. RESULTS: In simulations, SUW exhibited substantial bias and poor coverage. SPL and SAIPW yielded unbiased estimates with valid coverage probabilities under correct model specification, with SAIPW remaining robust under selection model misspecification. Both approaches showed no notable efficiency loss relative to centralized methods. Meta-learning methods were efficient for large sites but failed in settings with small cohort sizes and rare outcome prevalence. In real-data analysis, strong associations were consistently identified between smoking and cancers of the lung, bladder, and larynx, aligning with established epidemiological evidence. CONCLUSION: Our framework enables valid, privacy-enhancing inference across EHR cohorts with heterogeneous selection, supporting scalable, decentralized research using real-world data.

Journal Article↗

Concordance of SARS-CoV-2 Antibody Results during a Period of Low Prevalence.

Accurate, highly specific immunoassays for severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) are needed to evaluate seroprevalence. This study investigated the concordance of results across four immunoassays targeting different antigens for sera collected at the beginning of the SARS-CoV-2 pandemic in the United States. Specimens from All of Us participants contributed between January and March 2020 were tested using the Abbott Architect SARS-CoV-2 IgG (immunoglobulin G) assay (Abbott) and the EuroImmun SARS-CoV-2 enzyme-linked immunosorbent assay (ELISA) (EI). Participants with discordant results, participants with concordant positive results, and a subset of concordant negative results by Abbott and EI were also tested using the Roche Elecsys anti-SARS-CoV-2 (IgG) test (Roche) and the Ortho-Clinical Diagnostics Vitros anti-SARS-CoV-2 IgG test (Ortho). The agreement and 95% confidence intervals were estimated for paired assay combinations. SARS-CoV-2 antibody concentrations were quantified for specimens with at least two positive results across four immunoassays. Among the 24,079 participants, the percent agreement for the Abbott and EI assays was 98.8% (95% confidence interval, 98.7%, 99%). Of the 490 participants who were also tested by Ortho and Roche, the probability-weighted percentage of agreement (95% confidence interval) between Ortho and Roche was 98.4% (97.9%, 98.9%), that between EI and Ortho was 98.5% (92.9%, 99.9%), that between Abbott and Roche was 98.9% (90.3%, 100.0%), that between EI and Roche was 98.9% (98.6%, 100.0%), and that between Abbott and Ortho was 98.4% (91.2%, 100.0%). Among the 32 participants who were positive by at least 2 immunoassays, 21 had quantifiable anti-SARS-CoV-2 antibody concentrations by research assays. The results across immunoassays revealed concordance during a period of low prevalence. However, the frequency of false positivity during a period of low prevalence supports the use of two sequentially performed tests for unvaccinated individuals who are seropositive by the first test. IMPORTANCE What is the agreement of commercial SARS-CoV-2 immunoglobulin G (IgG) assays during a time of low coronavirus disease 2019 (COVID-19) prevalence and no vaccine availability? Serological tests produced concordant results in a time of low SARS-CoV-2 prevalence and no vaccine availability, driven largely by the proportion of samples that were negative by two immunoassays. The CDC recommends two sequential tests for positivity for future pandemic preparedness. In a subset analysis, quantified antinucleocapsid and antispike SARS-CoV-2 IgG antibodies do not suggest the need to specify the antigen targets of the sequential assays in the CDC's recommendation because false positivity varied as much between assays targeting the same antigen as it did between assays targeting different antigens.

Humans↗

Genomic analysis of mouse retinal development.

The vertebrate retina is comprised of seven major cell types that are generated in overlapping but well-defined intervals. To identify genes that might regulate retinal development, gene expression in the developing retina was profiled at multiple time points using serial analysis of gene expression (SAGE). The expression patterns of 1,051 genes that showed developmentally dynamic expression by SAGE were investigated using in situ hybridization. A molecular atlas of gene expression in the developing and mature retina was thereby constructed, along with a taxonomic classification of developmental gene expression patterns. Genes were identified that label both temporal and spatial subsets of mitotic progenitor cells. For each developing and mature major retinal cell type, genes selectively expressed in that cell type were identified. The gene expression profiles of retinal Müller glia and mitotic progenitor cells were found to be highly similar, suggesting that Müller glia might serve to produce multiple retinal cell types under the right conditions. In addition, multiple transcripts that were evolutionarily conserved that did not appear to encode open reading frames of more than 100 amino acids in length ("noncoding RNAs") were found to be dynamically and specifically expressed in developing and mature retinal cell types. Finally, many photoreceptor-enriched genes that mapped to chromosomal intervals containing retinal disease genes were identified. These data serve as a starting point for functional investigations of the roles of these genes in retinal development and physiology.

Animals↗

Prediction of mortality in an Indian intensive care unit. Comparison between APACHE II and artificial neural networks.

OBJECTIVE: To compare hospital outcome prediction using an artificial neural network model, built on an Indian data set, with the APACHE II (Acute Physiology and Chronic Health Evaluation II) logistic regression model. DESIGN: Analysis of a database containing prospectively collected data. SETTING: Medical-neurological ICU of a university hospital in Mumbai, India. SUBJECTS: Two thousand sixty-two consecutive admissions between 1996 and 1998. INTERVENTIONS: None. MEASUREMENTS AND RESULTS: The 22 variables used to obtain day-1 APACHE II score and risk of death were recorded. Data from 1,962 patients were used to train the neural network using a back-propagation algorithm. Data from the remaining 1,000 patients were used for testing this model and comparing it with APACHE II. There were 337 deaths in these 1,000 patients; APACHE II predicted 246 deaths while the neural network predicted 336 deaths. Calibration, assessed by the Hosmer-Lemeshow statistic, was better with the neural network (H=22.4) than with APACHE II (H=123.5) and so was discrimination (area under receiver operating characteristic curve =0.87 versus 0.77, p=0.002). Analysis of information gain due to each of the 22 variables revealed that the neural network could predict outcome using only 15 variables. A new model using these 15 variables predicted 335 deaths, had calibration (H=27.7) and discrimination (area under receiver operating characteristic curve =0.88) which was comparable to the 22-variable model (p=0.87) and superior to the APACHE II equation (p<0.001). CONCLUSION: Artificial neural networks, trained on Indian patient data, used fewer variables and yet outperformed the APACHE II system in predicting hospital outcome.

APACHE↗

Multivariate selection of genetic markers in diagnostic classification.

Analysis of gene expression data obtained from microarrays presents a new set of challenges to machine learning modeling. In this domain, in which the number of variables far exceeds the number of cases, identifying relevant genes or groups of genes that are good markers for a particular classification is as important as achieving good classification performance. Although several machine learning algorithms have been proposed to address the latter, identification of gene markers has not been systematically pursued. In this article, we investigate several algorithms for selecting gene markers for classification. We test these algorithms using logistic regression, as this is a simple and efficient supervised learning algorithm. We demonstrate, using 10 different data sets, that a conditionally univariate algorithm constitutes a viable choice if a researcher is interested in quickly determining a set of gene expression levels that can serve as markers for disease. We show that the classification performance of logistic regression is not very different from that of more sophisticated algorithms that have been applied in previous studies, and that the gene selection in the logistic regression algorithm is reasonable in both cases. Furthermore, the algorithm is simple, its theoretical basis is well established, and our user-friendly implementation is now freely available on the internet, serving as a benchmarking tool for the development of new algorithms.

Algorithms↗

Diagnostic accuracy of chest X-rays acquired using a digital camera for low-cost teleradiology.

Store-and-forward telemedicine, using e-mail to send clinical data and digital images, offers a low-cost alternative for physicians in developing countries to obtain second opinions from specialists. To explore the potential usefulness of this technique, 91 chest X-ray images were photographed using a digital camera and a view box. Four independent readers (three radiologists and one pulmonologist) read two types of digital (JPEG and JPEG2000) and original film images and indicated their confidence in the presence of eight features known to be radiological indicators of tuberculosis (TB). The results were compared to a "gold standard" established by two different radiologists, and assessed using receiver operating characteristic (ROC) curve analysis. There was no statistical difference in the overall performance between the readings from the original films and both types of digital images. The size of JPEG2000 images was approximately 120KB, making this technique feasible for slow internet connections. Our preliminary results show the potential usefulness of this technique particularly for tuberculosis and lung disease, but further studies are required to refine its potential.

Data Compression↗

Protecting patient privacy by quantifiable control of disclosures in disseminated databases.

One of the fundamental rights of patients is to have their privacy protected by health care organizations, so that information that can be used to identify a particular individual is not used to reveal sensitive patient data such as diagnoses, reasons for ordering tests, test results, etc. A common practice is to remove sensitive data from databases that are disseminated to the public, but this can make the disseminated database useless for important public health purposes. If the degree of anonymity of a disseminated data set could be measured, it would be possible to design algorithms that can assure that the desired level of confidentiality is achieved. Privacy protection in disseminated databases can be facilitated by the use of special ambiguation algorithms. Most of these algorithms are aimed at making one individual indistinguishable from one or more of his peers. However, even in databases considered "anonymous", it may still be possible to obtain sensitive information about some individuals or groups of individuals with the use of pattern recognition algorithms. In this article, we study the problem of determining the degree of ambiguation in disseminated databases and discuss its implications in the development and testing of "anonymization" algorithms.

Algorithms↗

A primer on gene expression and microarrays for machine learning researchers.

Data originating from biomedical experiments has provided machine learning researchers with an important source of motivation for developing and evaluating new algorithms. A new wave of algorithmic development has been initiated with the publication of gene expression data derived from microarrays. Microarray data analysis is particularly challenging given the large number of measurements (typically in the order of thousands) that are reported for relatively few samples (typically in the order of dozens). Many data sets are now available on the web. It is important that machine learning researchers understand how data are obtained and which assumptions are necessary in the analysis. Microarray data have the potential to cause significant impact in machine learning research, not just as a rich and realistic source of cases for testing new algorithms, as has been the UCI machine learning repository in the past decades, but also as a main motivation for their development. In this article, we briefly review the biology underlying microarrays, the process of obtaining gene expression measurements, and the rationale behind the common types of analyses involved in a microarray experiment. We outline the main challenges and reiterate critical considerations regarding the construction of supervised learning models that use this type of data. The goal of this article is to familiarize machine learning researchers with data originated from gene expression microarrays.

Algorithms↗

A greedy algorithm for supervised discretization.

We present a greedy algorithm for supervised discretization using a metric defined on the space of partitions of a set of objects. This proposed technique is useful for preparing the data for classifiers that require nominal attributes. Experimental work on decision trees and naïve Bayes classifiers confirm the efficacy of the proposed algorithm.

Algorithms↗

Deciphering gene expression profiles generated from DNA microarrays and their applications in oral medicine.

Genome-wide monitoring of gene expression profiles using DNA microarrays provides a unique approach to exploring the biological processes underlying oral diseases and disorders by providing a comprehensive survey of a cell's or tissue's transcriptional mapping. This revolutionary technology allows for the simultaneous assessment of the transcription levels of tens of thousands of genes, and of their relative expression between normal and diseased cells. As microarray data analysis evolves, there is a widespread hope that microarrays will significantly impact our ability to explore the genetic changes associated with disease etiology and development, ultimately leading to the discovery of new biomarkers for disease diagnosis and prognosis prediction as well as new therapeutic tools. The goal of this manuscript is to review 2 of the most commonly used microarray technologies, provide an overview of data analyses involved in a typical microarray experiment, and comment upon the application of microarrays to oral medicine.

Biomarkers↗

The Goodman-Kruskal coefficient and its applications in genetic diagnosis of cancer.

Increasing interest in new pattern recognition methods has been motivated by bioinformatics research. The analysis of gene expression data originated from microarrays constitutes an important application area for classification algorithms and illustrates the need for identifying important predictors. We show that the Goodman-Kruskal coefficient can be used for constructing minimal classifiers for tabular data, and we give an algorithm that can construct such classifiers.

Algorithms↗

International training in health informatics: a Brazilian experience.

Technology is transforming not only the practice of health-care but also professional training and educational models. Developing countries, such as Brazil, are increasingly suffering from a severe shortage of health informatics specialists. Training of professionals in this field is expensive, and there is a limited supply of high-quality teaching resources available. We envision that training in health informatics can be better achieved if cultural and technological barriers are anticipated and the training program is prepared accordingly. We describe our four-year experience of a Brazil/USA training program and discuss lessons learned during its implementation. Eleven onsite courses, one seminar, and two conferences were developed under this unique initiative, which made possible the collaboration among different countries and distinguished leaders in the field of medical informatics.

Brazil↗

A neural network-based similarity index for clustering DNA microarray data.

A common approach to the analysis of gene expression data is to define clusters of genes that have similar expression. A critical step in cluster analysis is the determination of similarity between the expression levels of two genes. We introduce a neural network-based similarity index as a non-linear similarity index and compare the results with other proximity measures for Saccharomyces cerevisiae gene expression data. We show that the clusters obtained using Euclidean distance, correlation coefficients, and mutual information were not significantly different. The clusters formed with the neural network-based index were more in agreement with those defined by functional categories and common regulatory motifs.

Algorithms↗

An Epicurean learning approach to gene-expression data classification.

We investigate the use of perceptrons for classification of microarray data where we use two datasets that were published in [Nat. Med. 7 (6) (2001) 673] and [Science 286 (1999) 531]. The classification problem studied by Khan et al. is related to the diagnosis of small round blue cell tumours (SRBCT) of childhood which are difficult to classify both clinically and via routine histology. Golub et al. study acute myeloid leukemia (AML) and acute lymphoblastic leukemia (ALL). We used a simulated annealing-based method in learning a system of perceptrons, each obtained by resampling of the training set. Our results are comparable to those of Khan et al. and Golub et al., indicating that there is a role for perceptrons in the classification of tumours based on gene-expression data. We also show that it is critical to perform feature selection in this type of models, i.e. we propose a method for identifying genes that might be significant for the particular tumour types. For SRBCTs, zero error on test data has been obtained for only 13 out of 2308 genes; for the ALL/AML problem, we have zero error for 9 out of 7129 genes that are used for the classification procedure. Furthermore, we provide evidence that Epicurean-style learning and simulated annealing-based search are both essential for obtaining the best classification results.

Algorithms↗

No-reflow is an independent predictor of death and myocardial infarction after percutaneous coronary intervention.

BACKGROUND: No-reflow occurring during percutaneous coronary intervention (PCI) has been associated with poor inhospital outcomes. The objectives of this analysis were to evaluate the occurrence of no-reflow as an independent predictor of adverse events and to determine whether treatment with intracoronary vasodilator therapy affected clinical outcomes. METHODS: We prospectively collected data from 4264 consecutive patients undergoing PCI, identifying those with no-reflow, and analyzed their treatments and clinical outcomes. RESULTS: No-reflow was identified in 135 of 4264 patients (3.2%). Baseline demographics were comparable, but patients with no-reflow were more likely to have acute myocardial infarction, unstable angina, and cardiogenic shock and to have undergone saphenous vein graft interventions. No-reflow was highly predictive of postprocedural myocardial infarction (17.7% vs 3.5% in patients without no-reflow, P <.001) and death (7.4% vs 2.0%, P <.001) and remained a strong independent predictor of death or myocardial infarction after multivariate analysis (odds ratio 3.6, P <.001). The administration of intracoronary verapamil, sodium nitroprusside, or both was not associated with a reduction in the rate of death or myocardial infarction (adjusted odds ratio of death or myocardial infarction 1.04, P =.945 for nitroprusside; and adjusted odds ratio of death or myocardial infarction 0.94, P =.91 for verapamil), despite an improvement in angiographic flow rates for patients treated with sodium nitroprusside. CONCLUSIONS: No-reflow is a strong independent predictor of inhospital mortality and postprocedural myocardial infarction. Administration of verapamil or sodium nitroprusside was not associated with improved inhospital outcomes in patients with no-reflow, although anterograde flow rates were improved in patients treated with sodium nitroprusside.

Aged↗

Microarrays and clinical dentistry.

BACKGROUND: The Human Genome Project, or HGP, has inspired a great deal of exciting biology recently by enabling the development of new technologies that will be essential for understanding the different types of abnormalities in diseases related to the oral cavity. LITERATURE REVIEWED: The authors review current literature pertaining to the advanced microarray technologies arising from the HGP and how they can contribute to dentistry. This technology has become a standard tool for monitoring activities of genes at both academic and pharmaceutical research institutions. RESULTS: With the availability of the DNA sequences for the entire human genome, attention now is focused on understanding various diseases at the genome level. Deciphering the molecular behavior of genetically encoded proteins is crucial to obtaining a more comprehensive picture of disease processes. Important progress has been made using microarrays, which have been shown to be effective in identifying gene expression patterns and variations that correlate with cellular development, physiology and function. Arrays can be used to classify tissue samples accurately based on molecular profiles and to select candidate genes related to a number of cancers, including oral cancer. This type of oral genetic approach will aid in the understanding of disease progression, thus improving diagnosis and treatment for patients. CLINICAL IMPLICATIONS: Microarrays hold much promise for the analysis of diseases in the oral cavity. As the technology evolves, dentists may see these tools as screening tests for better managing patients' dental care.

Anti-Bacterial Agents↗

Classification and identification of genes associated with oral cancer based on gene expression profiles. A preliminary study.

Oral squamous cell carcinoma (OSCC) is an aggressive malignancy. The five-year survival rate remains largely unchanged for the past 40 years. Early diagnosis has been shown to correlate with increased survival, based on cytologic changes. In order to improve our current treatment strategies for OSCC, it is necessary to understand the genetic and molecular networks underlying this disease. In this preliminary study, we illustrate the application of DNA microarrays to study OSCC. Using computational and statistical algorithms, we were able to differentiate (or classify) "cancer" and "normal" samples based on the behavior of the gene expression profiles. We found 651 genes to be associated with cancer. This article describes a preliminary study of current developments from the Human Genome Project (HGP) and its application to OSCC.

Carcinoma, Squamous Cell↗