PubMed Health⌕ Search

Biomedical subjects

Xionglei He

Publications and source records attributed to Xionglei He.

9 recordsLinked to original sources

Genome-wide profiling the integration patterns with T7-PCR.

Integration of exogenous gene fragments into the host genomes is a widely used and powerful method for studying gene functions, advancing molecular breeding, and conducting gene therapy. Accurately identifying the integration sites is essential for ensuring both the safety and efficacy of genome engineering efforts. However, current mapping techniques are constrained by high costs and a low signal-to-noise ratio. In this study, we developed an innovative tool for mapping integration sites, leveraging T7 polymerase-mediated in vitro transcription (T7-IVT) to capture the junction fragments surrounding integration loci. This approach converts genomic flanking sequences into RNA, enabling the simultaneous enrichment of junction fragments and the elimination of background genomic DNA, thereby significantly enhancing the signal-to-noise ratio. We have validated the efficiency of this method, named T7-PCR, across yeast, plant, and human cells under diverse integration scenarios. T7-PCR outperforms current next-generation sequencing (NGS)-based mapping strategies in terms of efficiency and accuracy, with minimal positional effects. This method is highly applicable for high-throughput transgene screening and also supports the development of next-generation tools for targeted integration of large fragments.

Humans↗

Toward a molecular understanding of pleiotropy.

Pleiotropy refers to the observation of a single gene influencing multiple phenotypic traits. Although pleiotropy is a common phenomenon with broad implications, its molecular basis is unclear. Using functional genomic data of the yeast Saccharomyces cerevisiae, here we show that, compared with genes of low pleiotropy, highly pleiotropic genes participate in more biological processes through distribution of the protein products in more cellular components and involvement in more protein-protein interactions. However, the two groups of genes do not differ in the number of molecular functions or the number of protein domains per gene. Thus, pleiotropy is generally caused by a single molecular function involved in multiple biological processes. We also provide genomewide evidence that the evolutionary conservation of genes and gene sequences positively correlates with the level of gene pleiotropy.

Evolution, Molecular↗

Why do hubs tend to be essential in protein networks?

The protein-protein interaction (PPI) network has a small number of highly connected protein nodes (known as hubs) and many poorly connected nodes. Genome-wide studies show that deletion of a hub protein is more likely to be lethal than deletion of a non-hub protein, a phenomenon known as the centrality-lethality rule. This rule is widely believed to reflect the special importance of hubs in organizing the network, which in turn suggests the biological significance of network architectures, a key notion of systems biology. Despite the popularity of this explanation, the underlying cause of the centrality-lethality rule has never been critically examined. We here propose the concept of essential PPIs, which are PPIs that are indispensable for the survival or reproduction of an organism. Our network analysis suggests that the centrality-lethality rule is unrelated to the network architecture, but is explained by the simple fact that hubs have large numbers of PPIs, therefore high probabilities of engaging in essential PPIs. We estimate that approximately 3% of PPIs are essential in the yeast, accounting for approximately 43% of essential genes. As expected, essential PPIs are evolutionarily more conserved than nonessential PPIs. Considering the role of essential PPIs in determining gene essentiality, we find the yeast PPI network functionally more robust than random networks, yet far less robust than the potential optimum. These and other findings provide new perspectives on the biological relevance of network structure and robustness.

Binding Sites↗

Transcriptional reprogramming and backup between duplicate genes: is it a genomewide phenomenon?

Deleting a duplicate gene often results in a less severe phenotype than deleting a singleton gene, a phenomenon commonly attributed to functional compensation among duplicates. However, duplicate genes rapidly diverge in expression patterns after duplication, making functional compensation less probable for ancient duplicates. Case studies suggested that a gene may provide compensation by altering its expression upon removal of its duplicate copy. On the basis of this observation and a genomic analysis, it was recently proposed that transcriptional reprogramming and backup among duplicates is a genomewide phenomenon in the yeast Saccharomyces cerevisiae. Here we reanalyze the yeast data and show that the high dispensability of duplicate genes with low expression similarity is a consequence of expression similarity and gene dispensability, each being correlated with a third factor, the number of protein interactions per gene. There is little evidence supporting widespread functional compensation of divergently expressed duplicate genes by transcriptional reprogramming.

Genes, Duplicate↗

Higher duplicability of less important genes in yeast genomes.

Gene duplication plays an important role in evolution because it is the primary source of new genes. Many recent studies showed that gene duplicability varies considerably among genes. Several considerations led us to hypothesize that less important genes have higher rates of successful duplications, where gene importance is measured by the fitness reduction caused by the deletion of the gene. Here, we test this hypothesis by comparing the importance of two groups of singleton genes in the yeast Saccharomyces cerevisiae (Sce). Group S genes did not duplicate in four other yeast species examined, whereas group D experienced duplication in these species. Consistent with our hypothesis, we found group D genes to be less important than group S genes. Specifically, 17% of group D genes are essential in Sce, compared to 28% for group S. Furthermore, deleting a group D gene in Sce reduces the fitness by 24% on average, compared to 38% for group S. Our subsequent analysis showed that less important genes have more cis-regulatory motifs, which could lead to a higher chance of subfunctionalization of duplicate genes and result in an enhanced rate of gene retention. Less important genes may also have weaker dosage imbalance effects and cause fewer genetic perturbations when duplicated. Regardless of the cause, our observation indicates that the previous finding of a less severe fitness consequence of deleting a duplicate gene than deleting a singleton gene is at least in part due to the fact that duplicate genes are intrinsically less important than singleton genes and suggests that the contribution of duplicate genes to genetic robustness has been overestimated.

Computational Biology↗

Gene complexity and gene duplicability.

Eukaryotic genes are on average more complex than prokaryotic genes in terms of expression regulation, protein length, and protein-domain structure [1-5]. Eukaryotes are also known to have a higher rate of gene duplication than prokaryotes do [6, 7]. Because gene duplication is the primary source of new genes [], the average gene complexity in a genome may have been increased by gene duplication if complex genes are preferentially duplicated. Here, we test this "gene complexity and gene duplicability" hypothesis with yeast genomic data. We show that, on average, duplicate genes from either whole-genome or individual-gene duplication have longer protein sequences, more functional domains, and more cis-regulatory motifs than singleton genes. This phenomenon is not a by-product of previously known mechanisms, such as protein function [10-13], evolutionary rate [14, 15], dosage [11], and dosage balance [16], that influence gene duplicability. Rather, it appears to have resulted from the sub-neo-functionalization process in duplicate-gene evolution [11]. Under this process, complex genes are more likely to be retained after duplication because they are prone to subfunctionalization, and gene complexity is regained via subsequent neofunctionalization. Thus, gene duplication increases both gene number and gene complexity, two important factors in the origin of genomic and organismal complexity.

Chromatin Immunoprecipitation↗

Significant impact of protein dispensability on the instantaneous rate of protein evolution.

The neutral theory of molecular evolution predicts that important proteins evolve more slowly than unimportant ones. High-throughput gene-knockout experiments in model organisms have provided information on the dispensability, and therefore importance, of thousands of proteins in a genome. However, previous studies of the correlation between protein dispensability and evolutionary rate were equivocal, and it has been proposed that the observed correlation is due to the covariation with the level of gene expression or is limited to duplicate genes. We here analyzed the gene dispensability data of the yeast Saccharomyces cerevisiae and estimated protein evolutionary rates by comparing S. cerevisiae with nine species of varying degrees of divergence from S. cerevisiae. The correlation between gene dispensability and evolutionary rate, although low, is highly significant, even when the gene expression level is controlled for or when duplicate genes are excluded. Our results thus support the hypothesis of lower evolution rates for more important proteins, a widely used principle in the daily practice of molecular biology. When the evolutionary rate is estimated from closely related species, the ratio between the mean rate of nonessential proteins to that of essential proteins is 1.4. This ratio declines to 1.1 when the evolutionary rate is estimated from distantly related species, suggesting that the importance of a protein may change in evolution, so the dispensability data obtained from a model organism only predicts a short-term rate of protein evolution. A comparison of the fitness contributions of orthologous genes in yeast and nematode supports this conclusion.

Animals↗

Rapid subfunctionalization accompanied by prolonged and substantial neofunctionalization in duplicate gene evolution.

Gene duplication is the primary source of new genes. Duplicate genes that are stably preserved in genomes usually have divergent functions. The general rules governing the functional divergence, however, are not well understood and are controversial. The neofunctionalization (NF) hypothesis asserts that after duplication one daughter gene retains the ancestral function while the other acquires new functions. In contrast, the subfunctionalization (SF) hypothesis argues that duplicate genes experience degenerate mutations that reduce their joint levels and patterns of activity to that of the single ancestral gene. We here show that neither NF nor SF alone adequately explains the genome-wide patterns of yeast protein interaction and human gene expression for duplicate genes. Instead, our analysis reveals rapid SF, accompanied by prolonged and substantial NF in a large proportion of duplicate genes, suggesting a new model termed subneofunctionalization (SNF). Our results demonstrate that enormous numbers of new functions have originated via gene duplication.

Amino Acid Sequence↗

A genome sequence of novel SARS-CoV isolates: the genotype, GD-Ins29, leads to a hypothesis of viral transmission in South China.

We report a complete genomic sequence of rare isolates (minor genotype) of the SARS-CoV from SARS patients in Guangdong, China, where the first few cases emerged. The most striking discovery from the isolate is an extra 29-nucleotide sequence located at the nucleotide positions between 27,863 and 27,864 (referred to the complete sequence of BJ01) within an overlapped region composed of BGI-PUP5 (BGI-postulated uncharacterized protein 5) and BGI-PUP6 upstream of the N (nucleocapsid) protein. The discovery of this minor genotype, GD-Ins29, suggests a significant genetic event and differentiates it from the previously reported genotype, the dominant form among all sequenced SARS-CoV isolates. A 17-nt segment of this extra sequence is identical to a segment of the same size in two human mRNA sequences that may interfere with viral replication and transcription in the cytosol of the infected cells. It provides a new avenue for the exploration of the virus-host interaction in viral evolution, host pathogenesis, and vaccine development.

Base Sequence↗