PubMed Health⌕ Search

Biomedical subjects

Yves A Lussier

Publications and source records attributed to Yves A Lussier.

At least 19 recordsLinked to original sources

Worldwide Innovative Network Consortium: Building a Common Global Cancer Database.

This review shares the ongoing work of the global Worldwide Innovative Network (WIN) Consortium for Precision Medicine to synthesize emerging cancer treatment data and to define the requirements for a common global cancer database that can truly support precision oncology. We performed a narrative review of emerging cancer treatment data, molecular profiling technologies, and existing clinicogenomic databases, focusing on how tumors are characterized, how subgroups are defined, and how demographic, lifestyle, and environmental factors are captured. The growth in molecular profiling technologies and the development of new targeted therapies are transforming cancer care. Tumors, regardless of tissue origin, are increasingly defined as composites of multiple, often rare, subgroups, each with distinct biology and likely response to specific therapies, based on multidimensional profiling of the tumor and its microenvironment. The solution lies in building vast databases that capture racial and ethnic diversity, reflected in genomic data, as well as diet and lifestyle factors that may have epigenetic impact on gene expression and post-translational modifications. A truly inclusive and informative data set must reflect global diversity, and there are multiple examples of demography-dependent differences in genomic signals. With members caring for and studying patients with cancer across five continents, WIN is actively exploring pathways to create a global cancer database, rich in clinical and molecular detail, granular enough for precise analysis, and large enough to power artificial intelligence-driven insights, provided appropriate data quality, validation, and governance frameworks are in place. This review surveys the current landscape and outlines practical paths forward to achieve this goal.

Humans↗

Computational approaches to phenotyping: high-throughput phenomics.

The recent completion of the Human Genome Project has made possible a high-throughput "systems approach" for accelerating the elucidation of molecular underpinnings of human diseases, and subsequent derivation of molecular-based strategies to more effectively prevent, diagnose, and treat these diseases. Although altered phenotypes are among the most reliable manifestations of altered gene functions, research using systematic analysis of phenotype relationships to study human biology is still in its infancy. This article focuses on the emerging field of high-throughput phenotyping (HTP) phenomics research, which aims to capitalize on novel high-throughput computation and informatics technology developments to derive genomewide molecular networks of genotype-phenotype associations, or "phenomic associations." The HTP phenomics research field faces the challenge of technological research and development to generate novel tools in computation and informatics that will allow researchers to amass, access, integrate, organize, and manage phenotypic databases across species and enable genomewide analysis to associate phenotypic information with genomic data at different scales of biology. Key state-of-the-art technological advancements critical for HTP phenomics research are covered in this review. In particular, we highlight the power of computational approaches to conduct large-scale phenomics studies.

Computational Biology↗

Integration of curated databases to identify genotype-phenotype associations.

BACKGROUND: The ability to rapidly characterize an unknown microorganism is critical in both responding to infectious disease and biodefense. To do this, we need some way of anticipating an organism's phenotype based on the molecules encoded by its genome. However, the link between molecular composition (i.e. genotype) and phenotype for microbes is not obvious. While there have been several studies that address this challenge, none have yet proposed a large-scale method integrating curated biological information. Here we utilize a systematic approach to discover genotype-phenotype associations that combines phenotypic information from a biomedical informatics database, GIDEON, with the molecular information contained in National Center for Biotechnology Information's Clusters of Orthologous Groups database (NCBI COGs). RESULTS: Integrating the information in the two databases, we are able to correlate the presence or absence of a given protein in a microbe with its phenotype as measured by certain morphological characteristics or survival in a particular growth media. With a 0.8 correlation score threshold, 66% of the associations found were confirmed by the literature and at a 0.9 correlation threshold, 86% were positively verified. CONCLUSION: Our results suggest possible phenotypic manifestations for proteins biochemically associated with sugar metabolism and electron transport. Moreover, we believe our approach can be extended to linking pathogenic phenotypes with functionally related proteins.

Computational Biology↗

An integrative genomic approach to uncover molecular mechanisms of prokaryotic traits.

With mounting availability of genomic and phenotypic databases, data integration and mining become increasingly challenging. While efforts have been put forward to analyze prokaryotic phenotypes, current computational technologies either lack high throughput capacity for genomic scale analysis, or are limited in their capability to integrate and mine data across different scales of biology. Consequently, simultaneous analysis of associations among genomes, phenotypes, and gene functions is prohibited. Here, we developed a high throughput computational approach, and demonstrated for the first time the feasibility of integrating large quantities of prokaryotic phenotypes along with genomic datasets for mining across multiple scales of biology (protein domains, pathways, molecular functions, and cellular processes). Applying this method over 59 fully sequenced prokaryotic species, we identified genetic basis and molecular mechanisms underlying the phenotypes in bacteria. We identified 3,711 significant correlations between 1,499 distinct Pfam and 63 phenotypes, with 2,650 correlations and 1,061 anti-correlations. Manual evaluation of a random sample of these significant correlations showed a minimal precision of 30% (95% confidence interval: 20%-42%; n = 50). We stratified the most significant 478 predictions and subjected 100 to manual evaluation, of which 60 were corroborated in the literature. We furthermore unveiled 10 significant correlations between phenotypes and KEGG pathways, eight of which were corroborated in the evaluation, and 309 significant correlations between phenotypes and 166 GO concepts evaluated using a random sample (minimal precision = 72%; 95% confidence interval: 60%-80%; n = 50). Additionally, we conducted a novel large-scale phenomic visualization analysis to provide insight into the modular nature of common molecular mechanisms spanning multiple biological scales and reused by related phenotypes (metaphenotypes). We propose that this method elucidates which classes of molecular mechanisms are associated with phenotypes or metaphenotypes and holds promise in facilitating a computable systems biology approach to genomic and biomedical research.

Algorithms↗

Natural language processing and visualization in the molecular imaging domain.

Molecular imaging is at the crossroads of genomic sciences and medical imaging. Information within the molecular imaging literature could be used to link to genomic and imaging information resources and to organize and index images in a way that is potentially useful to researchers. A number of natural language processing (NLP) systems are available to automatically extract information from genomic literature. One existing NLP system, known as BioMedLEE, automatically extracts biological information consisting of biomolecular substances and phenotypic data. This paper focuses on the adaptation, evaluation, and application of BioMedLEE to the molecular imaging domain. In order to adapt BioMedLEE for this domain, we extend an existing molecular imaging terminology and incorporate it into BioMedLEE. BioMedLEE's performance is assessed with a formal evaluation study. The system's performance, measured as recall and precision, is 0.74 (95% CI: [.70-.76]) and 0.70 (95% CI [.63-.76]), respectively. We adapt a JAVA viewer known as PGviewer for the simultaneous visualization of images with NLP extracted information.

Animals↗

Bio-Ontology and text: bridging the modeling gap.

MOTIVATION: Natural language processing (NLP) techniques are increasingly being used in biology to automate the capture of new biological discoveries in text, which are being reported at a rapid rate. Yet, information represented in NLP data structures is classically very different from information organized with ontologies as found in model organisms or genetic databases. To facilitate the computational reuse and integration of information buried in unstructured text with that of genetic databases, we propose and evaluate a translational schema that represents a comprehensive set of phenotypic and genetic entities, as well as their closely related biomedical entities and relations as expressed in natural language. In addition, the schema connects different scales of biological information, and provides mappings from the textual information to existing ontologies, which are essential in biology for integration, organization, dissemination and knowledge management of heterogeneous phenotypic information. A common comprehensive representation for otherwise heterogeneous phenotypic and genetic datasets, such as the one proposed, is critical for advancing systems biology because it enables acquisition and reuse of unprecedented volumes of diverse types of knowledge and information from text. RESULTS: A novel representational schema, PGschema, was developed that enables translation of phenotypic, genetic and their closely related information found in textual narratives to a well-defined data structure comprising phenotypic and genetic concepts from established ontologies along with modifiers and relationships. Evaluation for coverage of a selected set of entities showed that 90% of the information could be represented (95% confidence interval: 86-93%; n = 268). Moreover, PGschema can be expressed automatically in an XML format using natural language techniques to process the text. To our knowledge, we are providing the first evaluation of a translational schema for NLP that contains declarative knowledge about genes and their associated biomedical data (e.g. phenotypes). AVAILABILITY: http://zellig.cpmc.columbia.edu/PGschema

Abstracting and Indexing↗

Terminology model discovery using natural language processing and visualization techniques.

Medical terminologies are important for unambiguous encoding and exchange of clinical information. The traditional manual method of developing terminology models is time-consuming and limited in the number of phrases that a human developer can examine. In this paper, we present an automated method for developing medical terminology models based on natural language processing (NLP) and information visualization techniques. Surgical pathology reports were selected as the testing corpus for developing a pathology procedure terminology model. The use of a general NLP processor for the medical domain, MedLEE, provides an automated method for acquiring semantic structures from a free text corpus and sheds light on a new high-throughput method of medical terminology model development. The use of an information visualization technique supports the summarization and visualization of the large quantity of semantic structures generated from medical documents. We believe that a general method based on NLP and information visualization will facilitate the modeling of medical terminologies.

Automation↗

Genestrace: phenomic knowledge discovery via structured terminology.

The era of applied genomic medicine is quickly approaching accompanied by the increasing availability of detailed genetic information. Understanding the genetic etiology behind complex, multi-gene diseases remains an important challenge. In order to uncover the putative genetic etiology of complex diseases, we designed a method that explores the relationships between two major terminological and ontological resources: the Unified Medical Language System (UMLS) and the Gene Ontology (GO). The UMLS has a mainly clinical emphasis; Gene Ontology has become the standard for biological annotations of genes and gene products. Using statistical and semantic relationships within and between the two resources, we are able to infer relationships between disease concepts in the UMLS and gene products annotated using GO and its associated databases. We validated our inferences by comparing them to the known gene-disease relationships, as defined in the Online Mendelian Inheritance in Man's morbidmap (OMIM). The proof-of-concept methods presented here are unique in that they bypass the ambiguity of the direct extraction of gene or disease term from MEDLINE. Additionally, our methods provide direct links to clinically significant diseases through established terminologies or ontologies. The preliminary results presented here indicate the potential utility of exploiting the existing, manually curated relationships in biomedical resources as a tool for the discovery of potentially valuable new gene-disease relationships.

Computational Biology↗

Partitioning knowledge bases between advanced notification and clinical decision support systems.

Due to the varying rates of change of ephemeral administrative and enduring clinical knowledge in decision support systems (DSSs), the functional partition of knowledge base (KB) components can lead to more efficient and cost-effective system implementation and maintenance. Our prototype loosely couples a clinical event monitor developed by Columbia University Medical Center (CUMC) with a secure notification service proxy developed by IBM Research to form a novel and complex clinical event communication service.

Computer Security↗

Visualizing information across multidimensional post-genomic structured and textual databases.

MOTIVATION: Visualizing relationships among biological information to facilitate understanding is crucial to biological research during the post-genomic era. Although different systems have been developed to view gene-phenotype relationships for specific databases, very few have been designed specifically as a general flexible tool for visualizing multidimensional genotypic and phenotypic information together. Our goal is to develop a method for visualizing multidimensional genotypic and phenotypic information and a model that unifies different biological databases in order to present the integrated knowledge using a uniform interface. RESULTS: We developed a novel, flexible and generalizable visualization tool, called PhenoGenesviewer (PGviewer), which in this paper was used to display gene-phenotype relationships from a human-curated database (OMIM) and from an automatic method using a Natural Language Processing tool called BioMedLEE. Data obtained from multiple databases were first integrated into a uniform structure and then organized by PGviewer. PGviewer provides a flexible query interface that allows dynamic selection and ordering of any desired dimension in the databases. Based on users' queries, results can be visualized using hierarchical expandable trees that present views specified by users according to their research interests. We believe that this method, which allows users to dynamically organize and visualize multiple dimensions, is a potentially powerful and promising tool that should substantially facilitate biological research. AVAILABILITY: PhenogenesViewer as well as its support and tutorial are available at http://www.dbmi.columbia.edu/pgviewer/ CONTACT: Lussier@dbmi.columbia.edu.

Computer Graphics↗

A tool for abstracting relevant classes of concepts: the Common Ancestry Summarizer.

Controlled Medical Terminologies (CMTs) are indispensable tools for present medical information systems and medical informatics researches. Concept-oriented architecture and multiple-hierarchy have been accepted as two of the desiderata of contemporary CMTs. A common problem for informaticians working with class-based methods for terminologies is to find the Most Relevant Common Ancestor (MRCA) for a group of concepts, the common ancestor which has the closest semantic distance to all the given concepts. Finding an ancestor concept to summarize a group of concepts is required to map concepts from the knowledgebase to those of appropriate granularity in CMTs. However, manually exploring the hierarchical relationships and determining the MRCA are daunting tasks due to the massive size, the multiple hierarchies and in some case the cycles of the hierarchical networks of terminologies. In our study, we developed a web-based visualization tool, the Common Ancestry Summarizer (CAS), that graphically displays the hierarchical relations of multiple concepts within CMTs, and also assist users to identify the MRCA of these concepts by identifying the Lowest Common Ancestors (LCAs). The CAS has been shown useful in organizing classes for the controlled terms used in clinical guidelines (e.g. "the bottom-up method").

Algorithms↗

Automating terminological networks to link heterogeneous biomedical databases.

As cross-disciplinary research escalates, researchers are facing the challenge of linking disparate biomedical databases that have been developed without common indexes. Manually indexing these large-scale databases is laborious and often impractical. Solutions involving mediating terminologies have been proposed, but coordination of terms from the databases of interest to these mediating terminologies is also laborious, and regular synchronization between indexes is an additional problem. In this study we describe a novel method of linking heterogeneous databases using terminology networks constructed with automated mapping methods. Linkage was established between two disparate biomedical databases (SNOMED-CT and HDG), using two relevant intermediating databases (UMLS and OMIM). One gold standard of 514 distinct matches is used as proof-of-principle. In conclusion, as hypothesized, 1) Manually curated pathways provide high precision, but offer low recall, 2) the automated terminology pathways can significantly increase recall at acceptable precision. Taken together, our conclusion may suggest the combined manual and automated terminology networks could offer recall and precision in an incremental manner

Abstracting and Indexing↗

Mining OMIM for insight into complex diseases.

Understanding clinical phenotypes through their corresponding genotypes is one of the principal goals of genetic research. Though achieving this goal is relatively simple with single gene syndromes, more complex diseases often consist of varied clinical phenotypes that may be the result of interactions among multiple genetic loci. Microarray technology has brought the phenotype -genotype relationship to the molecular level, using differently behaving cancers, for example, as the basis for comparing patterns of gene expression. With this feasibility study, we attempted to use similar methods of analysis at the clinical level, in order to evaluate our hypothesis that the clustering of clinical phenotypes would provide information that would be useful in elucidating their underlying genotypes. Because of its breadth of content and detailed descriptions, we used OMIM as our source material for phenotypic and genetic information. After processing the source material, we then performed self-organizing map and hierarchical clustering analysis on representative diseases by phenotypic category. Through pre-determined queries over this analysis, we made two findings of potential clinical significance, one concerning diabetes and another concerning progressive neurologic diseases. Our methods provide a formal approach to analyzing phenotypes among diverse diseases, and may help indicate fruitful areas for further research into their underlying genetic causes.

Cluster Analysis↗

Putting data integration into practice: using biomedical terminologies to add structure to existing data sources.

A major purpose of biomedical terminologies is to provide uniform concept representation, allowing for improved methods of analysis of biomedical information. While this goal is being realized in bioinformatics, with the emergence of the Gene Ontology as a standard, there is still no real standard for the representation of clinical concepts. As discoveries in biology and clinical medicine move from parallel to intersecting paths, standardized representation will become more important. A large portion of significant data, however, is mainly represented as free text, upon which conducting computer-based inferencing is nearly impossible. In order to test our hypothesis that existing biomedical terminologies, specifically the UMLS Metathesaurus and SNOMED CT, could be used as templates to implement semantic and logical relationships over free text data that is important both clinically and biologically, we chose to analyze OMIM (Online Mendelian Inheritance in Man). After finding OMIM entries' conceptual equivalents in each respective terminology, we extracted the semantic relationships that were present and evaluated a subset of them for semantic, logical, and biological legitimacy. Our study reveals the possibility of putting the knowledge present in biomedical terminologies to its intended use, with potentially clinically significant consequences.

Databases, Genetic↗

Adapting current Arden Syntax knowledge for an object oriented event monitor.

Arden Syntax for Medical Logic Module (MLM)1 was designed for writing and sharing task-specific health knowledge in 1989. Several researchers have developed frameworks to improve the sharability and adaptability of Arden Syntax MLMs, an issue known as "curly braces" problem. Karadimas et al proposed an Arden Syntax MLM-based decision support system that uses an object oriented model and the dynamic linking features of the Java platform.2 Peleg et al proposed creating a Guideline Expression Language (GEL) based on Arden Syntax's logic grammar.3 The New York Presbyterian Hospital (NYPH) has a collection of about 200 MLMs. In a process of adapting the current MLMs for an object-oriented event monitor, we identified two problems that may influence the "curly braces" one: (1) the query expressions within the curly braces of Arden Syntax used in our institution are cryptic to the physicians, institutional dependent and written ineffectively (unpublished results), and (2) the events are coded individually within a curly braces, resulting sometimes in a large number of events - up to 200.

Decision Support Systems, Clinical↗

A "systematics" tool for medical terminologies.

Finding the hierarchical relations amongst multiple terms within medical terminologies that support multiple parents to a term is a common task, especially for trainees and knowledge engineers implementing or maintaining medical logic modules or guidelines. Examples of such terminologies include the UMLS and the Medical Entity Dictionary (MED). In addition, the task of identifying and discriminating amongst some common ancestors to a list of terms is a recurrent theme. This is also a common concern in the science of classification (systematics). Some nearest common ancestors have distinct valuable properties for classification and simplification of lists. Although there exist some visualized navigating and editing tools for the UMLS and the MED, they are browsers that show a large number of unrelated and irrelevant relationships to the task at hand. While algorithms have been well studied in computer science to solve such a problem over semantic networks and trees, to our knowledge, they have not been used with a visualization tool in biomedicine. We developed a visualized tool that graphically displays the hierarchical relations of multiple terms, and helps identifying the nearest common ancestors of these terms.

Classification↗