PubMed Health⌕ Search

Biomedical subjects

Lifeng Chen

Publications and source records attributed to Lifeng Chen.

10 recordsLinked to original sources

An electrostatic repulsion model of centromere organisation.

During cell division, chromosomes reorganise into compact bodies in which centromeres localise precisely at the chromatin surface1-4 to enable kinetochore-microtubule interactions essential for genome segregation5-8. The physical principles guiding this centromere positioning remain unknown. Here, we reveal that human core centromeres are directed to the chromatin surface by repulsion of centromere-associated proteins - independent of condensin-mediated loop extrusion and microtubule engagement. Using cellular perturbations, biochemical reconstitution, and multiscale molecular dynamics simulations, we show that chromatin surface localisation emerges from repulsion between condensed chromatin and both the kinetochore and the highly negatively charged centromere protein, CENP-B. Together, these elements form a centromeric region composed of two domains with opposing affinities, one favouring integration within the mitotic chromosome and the other favouring exposure to the surrounding cytoplasm, thereby driving surface positioning. Tethering synthetic negatively charged proteins to chromatin was sufficient to recapitulate this surface localisation in cells and in vitro, indicating that electrostatic repulsion is a key determinant of surface localisation. These findings demonstrate that centromere layering is not hardwired by chromatin folding patterns but instead emerges from phase separation in chromatin. Our work uncovers electrostatic polarity as a general and programmable mechanism to spatially organise chromatin.

Journal Article↗

Cryptosporidium parvum: identification of a new surface adhesion protein on sporozoite and oocyst by screening of a phage-display cDNA library.

Cryptosporidium parvum is a significant cause of diarrheal disease worldwide. The specific molecules that mediate C. parvum-host interaction and the molecular mechanisms involved in the pathogenesis are unknown. In this study we described a novel phage display method to identify surface adhesion proteins of C. parvum. A cDNA library of the sporozoite and oocyst stages of C. parvum expressed on the surface of T7 phage was screened with intestinal epithelial cells (IECs) from the newborn Cryptosporidium-free Holstein calves. Proteins that selectively and specifically bound to IECs were then enriched using a multi-step panning procedure. Two proteins of C. parvum were selected, one was previously reported (p23), which was an important surface adhesion protein; the other was a novel surface adherence protein (CP12). Sequence analysis showed that CP12 has a N-terminal signal peptide, a transmembrane region, a N-glycosylation site, a casein kinase II phosphorylation site and two N-myristoylation sites. Immunofluorescence assay (IFA) using antibody specific for rCP12 demonstrated that the antibody can specifically bind the surface of sporozoite and oocyst, especially apical region of sporozoite. The surface localization of CP12 and its involvement in the host-parasite interaction suggest that it may serve as an effective target for specific preventive and therapeutic measures for cryptosporidiosis.

Amino Acid Sequence↗

Natural language processing and visualization in the molecular imaging domain.

Molecular imaging is at the crossroads of genomic sciences and medical imaging. Information within the molecular imaging literature could be used to link to genomic and imaging information resources and to organize and index images in a way that is potentially useful to researchers. A number of natural language processing (NLP) systems are available to automatically extract information from genomic literature. One existing NLP system, known as BioMedLEE, automatically extracts biological information consisting of biomolecular substances and phenotypic data. This paper focuses on the adaptation, evaluation, and application of BioMedLEE to the molecular imaging domain. In order to adapt BioMedLEE for this domain, we extend an existing molecular imaging terminology and incorporate it into BioMedLEE. BioMedLEE's performance is assessed with a formal evaluation study. The system's performance, measured as recall and precision, is 0.74 (95% CI: [.70-.76]) and 0.70 (95% CI [.63-.76]), respectively. We adapt a JAVA viewer known as PGviewer for the simultaneous visualization of images with NLP extracted information.

Animals↗

Inhibition of krr1 gene expression in Giardia canis by a virus-mediated hammerhead ribozyme.

Giardia, a most primitive eukaryote, infects several species including human and it is a major agent of waterborne outbreak of diarrhea. It has been difficult to employ standard genetic methods in the study of Giardia, but the RNA virus-based transfection system has been developed and used for the genetic manipulation. KRR1 protein is responsible for ribosome biosynthesis in Giardia. In this study, cDNA encoding hammerhead ribozyme flanked with various lengths of antisense Krr1 RNA were cloned into a viral vector pGCV634/GFP/GCV2174 derived from the genome of Giardia canis virus (GCV). RNA transcripts of the plasmids showed high cleavage activities on Krr1 mRNA in vitro. They were electroporated into GCV-infected G. canis trophozoites and Krr1 mRNA level was decreased by 72% with the ribozyme KRzS and 86% with the ribozyme KRzL, while the control ribozyme TRzS showed no effect on the level of Krr1 mRNA. The two hammerhead ribozyme transfected cells grew slowly, their internal structures got blurred and the cells were deformed. These results indicated that GCV could be useful tool for gene manipulation of G. canis.

Animals↗

Identifying metabolic enzymes with multiple types of association evidence.

BACKGROUND: Existing large-scale metabolic models of sequenced organisms commonly include enzymatic functions which can not be attributed to any gene in that organism. Existing computational strategies for identifying such missing genes rely primarily on sequence homology to known enzyme-encoding genes. RESULTS: We present a novel method for identifying genes encoding for a specific metabolic function based on a local structure of metabolic network and multiple types of functional association evidence, including clustering of genes on the chromosome, similarity of phylogenetic profiles, gene expression, protein fusion events and others. Using E. coli and S. cerevisiae metabolic networks, we illustrate predictive ability of each individual type of association evidence and show that significantly better predictions can be obtained based on the combination of all data. In this way our method is able to predict 60% of enzyme-encoding genes of E. coli metabolism within the top 10 (out of 3551) candidates for their enzymatic function, and as a top candidate within 43% of the cases. CONCLUSION: We illustrate that a combination of genome context and other functional association evidence is effective in predicting genes encoding metabolic enzymes. Our approach does not rely on direct sequence homology to known enzyme-encoding genes, and can be used in conjunction with traditional homology-based metabolic reconstruction methods. The method can also be used to target orphan metabolic activities.

Energy Metabolism↗

Predicting genes for orphan metabolic activities using phylogenetic profiles.

Homology-based methods fail to assign genes to many metabolic activities present in sequenced organisms. To suggest genes for these orphan activities we developed a novel method that efficiently combines local structure of a metabolic network with phylogenetic profiles. We validated our method using known metabolic genes in Saccharomyces cerevisiae and Escherichia coli. We show that our method should be easily transferable to other organisms, and that it is robust to errors in incomplete metabolic networks.

Databases, Nucleic Acid↗

Giardia lamblia: stable expression of green fluorescent protein mediated by giardiavirus.

Giardia lamblia, an early diverging eukaryote that infects several species including humans and a major agent of water-borne diarrhea throughout the world, can be infected with a double-stranded RNA virus, giardiavirus (GLV). A chimeric GLV cDNA and green fluorescent protein (GFP) according to the cis-acting signals of the GLV genome required for expression of foreign gene was constructed and its in vitro transcript was electroporated into GLV-infected G. lamblia trophozoites, GFP was expressed transiently. pGDH5/NEO/GLV was constructed by combining the neomycin resistance cassette in which the neomycin phosphotransferase gene was flanked by Giardia glutamate dehydrogenase (GDH) uncoding regions and the transcription cassette in which the chimera of GLV cDNA and GFP was located downstream from GDH gene promoter on a single plasmid. This plasmid was electroporated into G. lamblia and the transfectants persistently expressed GFP under G418 selection. This stable transfection system should provide a valuable tool for genetic study of G. lamblia.

Animals↗

Gene name ambiguity of eukaryotic nomenclatures.

MOTIVATION: With more and more scientific literature published online, the effective management and reuse of this knowledge has become problematic. Natural language processing (NLP) may be a potential solution by extracting, structuring and organizing biomedical information in online literature in a timely manner. One essential task is to recognize and identify genomic entities in text. 'Recognition' can be accomplished using pattern matching and machine learning. But for 'identification' these techniques are not adequate. In order to identify genomic entities, NLP needs a comprehensive resource that specifies and classifies genomic entities as they occur in text and that associates them with normalized terms and also unique identifiers so that the extracted entities are well defined. Online organism databases are an excellent resource to create such a lexical resource. However, gene name ambiguity is a serious problem because it affects the appropriate identification of gene entities. In this paper, we explore the extent of the problem and suggest ways to address it. RESULTS: We obtained gene information from 21 organisms and quantified naming ambiguities within species, across species, with English words and with medical terms. When the case (of letters) was retained, official symbols displayed negligible intra-species ambiguity (0.02%) and modest ambiguities with general English words (0.57%) and medical terms (1.01%). In contrast, the across-species ambiguity was high (14.20%). The inclusion of gene synonyms increased intra-species ambiguity substantially and full names contributed greatly to gene-medical-term ambiguity. A comprehensive lexical resource that covers gene information for the 21 organisms was then created and used to identify gene names by using a straightforward string matching program to process 45,000 abstracts associated with the mouse model organism while ignoring case and gene names that were also English words. We found that 85.1% of correctly retrieved mouse genes were ambiguous with other gene names. When gene names that were also English words were included, 233% additional 'gene' instances were retrieved, most of which were false positives. We also found that authors prefer to use synonyms (74.7%) to official symbols (17.7%) or full names (7.6%) in their publications. CONTACT: lifeng.chen@dbmi.columbia.edu

Abstracting and Indexing↗

Translational accuracy during exponential, postdiauxic, and stationary growth phases in Saccharomyces cerevisiae.

When the yeast Saccharomyces cerevisiae shifts from rapid growth on glucose to slow growth on ethanol, it undergoes profound changes in cellular metabolism, including the destruction of most of the translational machinery. We have examined the effect of this metabolic change, termed the diauxic shift, on the frequency of translational errors. Recoding sites are mRNA sequences that increase the frequency of translational errors, providing a convenient reporter of translational accuracy. We found that the diauxic shift causes no overall change in translational accuracy but does cause a strong reduction in the frequency of one type of programmed error: Ty +1 frameshifting. Genetic data suggest that this effect may be due to changes in the relative amounts of tRNA participating in translation elongation. We discuss possible implications for expression strategies that use recoding.

Codon↗

Extracting phenotypic information from the literature via natural language processing.

In recent years, the amount of biomedical knowledge has been increasing exponentially. Several Natural Language Processing (NLP) systems have been developed to help researchers extract, encode and organize new information automatically from textual literature or narrative reports. Some of these systems focus on extracting biological entities or molecular interactions while others retrieve and encode clinical information. To exploit gene functions in the post-genome era, it is necessary to extract phenotypic information automatically from the literature as well. However, few NLP projects have focused on this. We present the development of a system called BioMedLEE that extracts a broad variety of phenotypic information from the biomedical literature. The system was developed by adapting MedLEE, an existing clinical information extraction NLP engine. A feasibility evaluation study of BioMedLEE was performed using 300 randomly chosen journal titles. Results showed that experts achieved an average precision rate of 65.4%, (95%CI: [58.0%, 72.8%]) and a recall rate of 73.0%, (95%CI: [66.2%, 80.0%]). BioMedLEE had 64.0% precision and 77.1% recall respectively, according to expert agreements.

Databases as Topic↗