PubMed Health⌕ Search

Biomedical subjects

Xuegong Zhang

Publications and source records attributed to Xuegong Zhang.

At least 19 recordsLinked to original sources

Identifying fate-determining transcription factors with single-cell omics.

Single-cell sequencing enables the systematic discovery of cell fate-determining transcription factors (TFs), or key TFs, that define cellular identity or drive cell state transitions. A wide range of computational methods have been developed for this goal, but they differ substantially in the input data and the biological questions they address. In this article, we systematically review computational approaches for key TF identification and organize them from three perspectives: whether they identify TFs defining cell state identity or driving state transitions, whether transitions are modeled as discrete or continuous processes, and whether TFs act individually or combinatorially. We summarize key features and application scenarios of relevant methods to guide tool selection and discuss emerging trends in this field toward programmable and active control of cell fate.

Transcription Factors↗

Primary transcripts and expressions of mammal intergenic microRNAs detected by mapping ESTs to their flanking sequences.

MicroRNAs (miRNAs) are a class of approximately 22-nt small RNAs that regulate posttranscriptional gene expression. Thousands of expressed sequence tags (ESTs) have been identified by using upstream 2500-nt and downstream 4000-nt flanking sequences to BLAST in the dbEST database. The cotranscription of the miRNAs and their flanking sequences covered by the matched ESTs is verified by RT-PCR. It directly reveals that a large portion of mammalian intergenic miRNAs are first transcribed as long primary transcripts (pri-miRNAs). Also, the transcripts' ranges of tens of pri-miRNAs are predicted by the EST-extension method. We then extracted the tissue-specific expression information from the annotations of the matched ESTs and established the expression profile of the studied miRNAs for tens of tissues. This provided a new way to establish the expression profiles of miRNAs. Results show that the human brain, lung, liver, and eye and the mouse brain, eye, and mammary gland are tissues in which enriched numbers of miRNAs are expressed.

3' Flanking Region↗

A new method for detecting human recombination hotspots and its applications to the HapMap ENCODE data.

Computational detection of recombination hotspots from population polymorphism data is important both for understanding the nature of recombination and for applications such as association studies. We propose a new method for this task based on a multiple-hotspot model and an (approximate) log-likelihood ratio test. A truncated, weighted pairwise log-likelihood is introduced and applied to the calculation of the log-likelihood ratio, and a forward-selection procedure is adopted to search for the optimal hotspot predictions. The method shows a relatively high power with a low false-positive rate in detecting multiple hotspots in simulation data and has a performance comparable to the best results of leading computational methods in experimental data for which recombination hotspots have been characterized by sperm-typing experiments. The method can be applied to both phased and unphased data directly, with a very fast computational speed. We applied the method to the 10 500-kb regions of the HapMap ENCODE data and found 172 hotspots among the three populations, with average hotspot width of 2.4 kb. By comparisons with the simulation data, we found some evidence that hotspots are not all identical across populations. The correlations between detected hotspots and several genomic characteristics were examined. In particular, we observed that DNaseI-hypersensitive sites are enriched in hotspots, suggesting the existence of human beta hotspots similar to those found in yeast.

Base Composition↗

Predicting methylation status of CpG islands in the human brain.

MOTIVATION: Over 50% of human genes contain CpG islands in their 5'-regions. Methylation patterns of CpG islands are involved in tissue-specific gene expression and regulation. Mis-epigenetic silencing associated with aberrant CpG island methylation is one mechanism leading to the loss of tumor suppressor functions in cancer cells. Large-scale experimental detection of DNA methylation is still both labor-intensive and time-consuming. Therefore, it is necessary to develop in silico approaches for predicting methylation status of CpG islands. RESULTS: Based on a recent genome-scale dataset of DNA methylation in human brain tissues, we developed a classifier called MethCGI for predicting methylation status of CpG islands using a support vector machine (SVM). Nucleotide sequence contents as well as transcription factor binding sites (TFBSs) are used as features for the classification. The method achieves specificity of 84.65% and sensitivity of 84.32% on the brain data, and can also correctly predict about two-third of the data from other tissues reported in the MethDB database. AVAILABILITY: An online predictor based on MethCGI is available at http://166.111.201.7/MethCGI.html CONTACT: mzhang@cshl.edu SUPPLEMENTARY INFORMATION: Supplementary data available at Bioinformatics online and http://166.111.201.7/help.html.

Algorithms↗

Recursive SVM feature selection and sample classification for mass-spectrometry and microarray data.

BACKGROUND: Like microarray-based investigations, high-throughput proteomics techniques require machine learning algorithms to identify biomarkers that are informative for biological classification problems. Feature selection and classification algorithms need to be robust to noise and outliers in the data. RESULTS: We developed a recursive support vector machine (R-SVM) algorithm to select important genes/biomarkers for the classification of noisy data. We compared its performance to a similar, state-of-the-art method (SVM recursive feature elimination or SVM-RFE), paying special attention to the ability of recovering the true informative genes/biomarkers and the robustness to outliers in the data. Simulation experiments show that a 5%- approximately 20% improvement over SVM-RFE can be achieved regard to these properties. The SVM-based methods are also compared with a conventional univariate method and their respective strengths and weaknesses are discussed. R-SVM was applied to two sets of SELDI-TOF-MS proteomics data, one from a human breast cancer study and the other from a study on rat liver cirrhosis. Important biomarkers found by the algorithm were validated by follow-up biological experiments. CONCLUSION: The proposed R-SVM method is suitable for analyzing noisy high-throughput proteomics and microarray data and it outperforms SVM-RFE in the robustness to noise and in the ability to recover informative features. The multivariate SVM-based method outperforms the univariate method in the classification performance, but univariate methods can reveal more of the differentially expressed features especially when there are correlations between the features.

Algorithms↗

The effect of GeneChip gene definitions on the microarray study of cancers.

The Affymetrix GeneChip is a popular microarray platform for genome-wide expression profiling and has been widely used in functional genomics especially in the classification of cancers. Due to the updating of genome data, much of the genome information with which the chips were designed is out-of-date and it has been reported that many of the genes/transcripts on the chips differ from their original definition when mapping the probes to the new genome information. Dai et al. have reported that the updated definition can cause as much as 30-50% discrepancy in the genes selected as differentially expressed on a heart tissue expression profiling dataset. Understanding the nature of this difference is therefore very important for the utilization of the data. In this work, with a large cancer dataset as an example, we compared two major definitions and investigated their effects on classification, clustering, discovery of differentially expressed genes and gene-set-based analysis. Results show that the two definitions agree well on clustering and classification results but genes and gene sets discovered as differentially expressed or enriched can be very different. Discoveries based on the Affymetrix definition can cover most of those based on the new definition, but tend to have more false positives.

Base Sequence↗

Symptom combinations associated with outcome and therapeutic effects in a cohort of cases with SARS.

Severe acute respiratory syndrome (SARS) is an infectious disease and some of its symptoms were clinically indistinguishable of those from similar diseases. This study aimed to find the symptom combinations associated with adverse outcome and the therapeutic effects in a cohort of patients with probable SARS retrospectively. In 2003, 123 SARS cases in Beijing were subjected to a strictly western medicine (WM) treatment, or a combined treatment (WM plus Herba houttuyniae injection, addition of individualized herbal treatments when necessary), of which 115 were followed till death or discharge; 8 were transferred and lost to follow-up. In both treatment groups, clinical manifestations were evaluated daily; development of signs and symptoms, and their possible relationship with outcome, were assessed. The relationships between these sign/symptom complexes and outcome under two treatment protocols were evaluated and differences were noted. Dynamic symptom combinations, dividing into the early, the medium-term and the durational symptom clusters, were identified as likely being related to the adverse outcomes of SARS (p < 0.05, p < 0.01). Compared with a strictly WM treatment, the combined treatment resulted in a longer hospital stay (p = 0.028), a non-statistically significant mortality rate decrease (combined treatment: 9.6% versus WM: 11.1%), and a significant improvement of arthralgia and myalgia (p < 0.05) in the early symptom cluster. Additionally, the combined protocol improved arterial oxyhemoglobin saturation significantly at day 22 (p < 0.05). In conclusion, the progress and outcome of SARS may be associated with specific temporal patterns of development in combination of several non-specific signs and symptom complexes, which are also helpful for evaluating the therapeutic effects on SARS patients.

Adolescent↗

Embryonics: a path to artificial life?

Electronic systems, no matter how clever and intelligent they are, cannot yet demonstrate the reliability that biological systems can. Perhaps we can learn from these processes, which have developed through millions of years of evolution, in our pursuit of highly reliable systems. This article discusses how such systems, inspired by biological principles, might be built using simple embryonic cells. We illustrate how they can monitor their own functional integrity in order to protect themselves from internal failure or from hostile environmental effects and how faults caused by DNA mutation or cell death can be repaired and thus full system functionality restored.

Animals↗

Classification of real and pseudo microRNA precursors using local structure-sequence features and support vector machine.

BACKGROUND: MicroRNAs (miRNAs) are a group of short (approximately 22 nt) non-coding RNAs that play important regulatory roles. MiRNA precursors (pre-miRNAs) are characterized by their hairpin structures. However, a large amount of similar hairpins can be folded in many genomes. Almost all current methods for computational prediction of miRNAs use comparative genomic approaches to identify putative pre-miRNAs from candidate hairpins. Ab initio method for distinguishing pre-miRNAs from sequence segments with pre-miRNA-like hairpin structures is lacking. Being able to classify real vs. pseudo pre-miRNAs is important both for understanding of the nature of miRNAs and for developing ab initio prediction methods that can discovery new miRNAs without known homology. RESULTS: A set of novel features of local contiguous structure-sequence information is proposed for distinguishing the hairpins of real pre-miRNAs and pseudo pre-miRNAs. Support vector machine (SVM) is applied on these features to classify real vs. pseudo pre-miRNAs, achieving about 90% accuracy on human data. Remarkably, the SVM classifier built on human data can correctly identify up to 90% of the pre-miRNAs from other species, including plants and virus, without utilizing any comparative genomics information. CONCLUSION: The local structure-sequence features reflect discriminative and conserved characteristics of miRNAs, and the successful ab initio classification of real and pseudo pre-miRNAs opens a new approach for discovering new miRNAs.

Animals↗

Multi-locus penetrance variance analysis method for association study in complex diseases.

Common heritable diseases often result from the action of several different genes, each of which contributes to the total observed variability in the disease trait. Traditional single-locus association approaches rely heavily on the marginal effects of single-locus and tend to ignore the multigenic nature of complex diseases. The increasing request for localizing genes underlying traits in multi-gene diseases has led to the development of some statistical methods. In this study, we develop a multi-locus analysis method - multi-locus penetrance variance analysis (MPVA), and conduct systematical simulation studies to evaluate its performance. Our results show that compared with other multi-locus methods, MPVA has some advantage in detecting complicated interactions under different epistatic models, and its performance is stable and robust.

Analysis of Variance↗

The effect of U1 snRNA binding free energy on the selection of 5' splice sites.

The importance of U1 snRNA binding free energy in the regulation of alternative splicing has been studied in some genes with site-directed mutagenesis. Here we report a large-scale analysis of its impact on 5' splice site (5'ss) selection in human genome. The results show that free energy exerts different effects on alternative 5'ss choice in different situations and -8.1 kcal/mol is a threshold. When both free energies of two competing 5'ss are larger than -8.1 kcal/mol, the 5'ss with lower free energy is more frequently used. However, in other pairs of 5'ss, lower-free-energy 5'ss does not seem to be favored and even the other 5'ss is used more frequently, which suggests that very low binding free energy would impair splicing. Some observations hold true only for those alternative 5' splicing with short alternative exons (<50nt), which implies a complex mechanism of 5'ss selection involving both U1 snRNA binding free energy and regulatory factors.

Alternative Splicing↗

MicroRNA identification based on sequence and structure alignment.

MOTIVATION: MicroRNAs (miRNA) are approximately 22 nt long non-coding RNAs that are derived from larger hairpin RNA precursors and play important regulatory roles in both animals and plants. The short length of the miRNA sequences and relatively low conservation of pre-miRNA sequences restrict the conventional sequence-alignment-based methods to finding only relatively close homologs. On the other hand, it has been reported that miRNA genes are more conserved in the secondary structure rather than in primary sequences. Therefore, secondary structural features should be more fully exploited in the homologue search for new miRNA genes. RESULTS: In this paper, we present a novel genome-wide computational approach to detect miRNAs in animals based on both sequence and structure alignment. Experiments show this approach has higher sensitivity and comparable specificity than other reported homologue searching methods. We applied this method on Anopheles gambiae and detected 59 new miRNA genes. AVAILABILITY: This program is available at http://bioinfo.au.tsinghua.edu.cn/miralign. SUPPLEMENTARY INFORMATION: Supplementary information is available at http://bioinfo.au.tsinghua.edu.cn/miralign/supplementary.htm.

Algorithms↗

htSNPer1.0: software for haplotype block partition and htSNPs selection.

BACKGROUND: There is recently great interest in haplotype block structure and haplotype tagging SNPs (htSNPs) in the human genome for its implication on htSNPs-based association mapping strategy for complex disease. Different definitions have been used to characterize the haplotype block structure in the human genome, and several different performance criteria and algorithms have been suggested on htSNPs selection. RESULTS: A heuristic algorithm, generalized branch-and-bound algorithm, is applied to the searching of minimal set of haplotype tagging SNPs (htSNPs) according to different htSNPs performance criteria. We develop a software htSNPer1.0 to implement the algorithm, and integrate three htSNPs performance criteria and four haplotype block definitions for haplotype block partitioning. It is a software with powerful Graphical User Interface (GUI), which can be used to characterize the haplotype block structure and select htSNPs in the candidate gene or interested genomic regions. It can find the global optimization with only a fraction of the computing time consumed by exhaustive searching algorithm. CONCLUSION: htSNPer1.0 allows molecular geneticists to perform haplotype block analysis and htSNPs selection using different definitions and performance criteria. The software is a powerful tool for those focusing on association mapping based on strategy of haplotype block and htSNPs.

Algorithms↗

Multi-locus association study of schizophrenia susceptibility genes with a posterior probability method.

Schizophrenia is a serious neuropsychiatric illness affecting about 1% of the world's population. It is considered a complex inheritance disorder. A number of genes are involved in combination in the etiology of the disorder. Evidence implicates the altered dopaminergic transmission in schizophrenia. In the present study, in order to identify susceptibility genes for schizophrenia in dopaminergic metabolism, we analyzed 59 single nucleotide polymorphisms (SNPs) in 24 genes of the dopaminergic pathway among 82 unrelated patients with schizophrenia and 108 matched normal controls. Considering that traditional single-locus association studies ignore the multigenic nature of complex diseases and do not take into account possible interactions between susceptibility genes, we proposed a multi-locus analysis method, using the posterior probability of morbidity as a measure of absolute disease risk for a multi-locus genotype combination, and developed an algorithm based on perturbation and average to detect the susceptibility multi-locus genotype combinations, as well as to repress noise and avoid false positive results at our best. A three-locus SNP genotype combination involved in the interactions of COMT and ALDH3B1 genes was detected to be significantly susceptible to schizophrenia.

Aldehyde Dehydrogenase↗

Evidence and characteristics of putative human alpha recombination hotspots.

Understanding recombination rate variation is very important for studying genome diversity and evolution, and for investigation of phenotypic association and genetic diseases. Recombination hotspots have been observed in many species and are well studied in yeast. Recent study demonstrated that recombination hotspots are also a ubiquitous feature of the human genome. But the nature of human hotspots remains largely unknown. We have developed and validated a novel computational method for testing the existence of hotspots as well as for localizing them with either unphased or phased genotyping data. To study the characteristics of hotspots within or close to genes, we scanned for unusually high levels of recombination using the European population samples in the SeattleSNPs database, and found evidence for the existence of human alpha hotspots similar to those of yeast. This type of hotspots, found at promoter regions, accounts for about half of the total detected and appears to depend on some specific transcription factor binding sites (such as CGCCCCCGC). These characteristics can explain the observed weak correlation between hotspots and GC-content, and their variation may contribute to the diversity of hotspot distribution among different individuals and species. These long-sought putative human alpha recombination hotspots should deserve further experimental investigations.

Binding Sites↗

The effect of haplotype-block definitions on inference of haplotype-block structure and htSNPs selection.

It has been recently suggested that the human genome is organized as a series of haplotype blocks, and efforts to create a genome-wide haplotype map are already underway. Several computational algorithms have been proposed to partition the genome. However, little is known about their behaviors in relation to the haplotype-block partitioning and haplotype-tagging SNPs selection. Here, we present a systematic comparison of three classes of haplotype-block partition definitions, a diversity-based method, a linkage-disequilibrium (LD)-based method, and a recombination-based method. The data used were derived from a coalescent simulation under both a uniform recombination model and one that assumes recombination hotspots. There were considerable differences in haplotype information loss in the measure of entropy when the partition methods were compared under different population-genetics scenarios. Under both recombination models, the results from the LD-based definition and the recombination-based definition were more similar to each other than were the results from the diversity-based definition. This work demonstrates that when undertaking haplotype-based association mapping, the choice of haplotype-block definition and SNP selection requires careful consideration.

Algorithms↗

Molecular classification of liver cirrhosis in a rat model by proteomics and bioinformatics.

Liver cirrhosis is a worldwide health problem. Reliable, noninvasive methods for early detection of liver cirrhosis are not available. Using a three-step approach, we classified sera from rats with liver cirrhosis following different treatment insults. The approach consisted of: (i) protein profiling using surface-enhanced laser desorption/ionization (SELDI) technology; (ii) selection of a statistically significant serum biomarker set using machine learning algorithms; and (iii) identification of selected serum biomarkers by peptide sequencing. We generated serum protein profiles from three groups of rats: (i) normal (n=8), (ii) thioacetamide-induced liver cirrhosis (n=22), and (iii) bile duct ligation-induced liver fibrosis (n=5) using a weak cation exchanger surface. Profiling data were further analyzed by a recursive support vector machine algorithm to select a panel of statistically significant biomarkers for class prediction. Sensitivity and specificity of classification using the selected protein marker set were higher than 92%. A consistently down-regulated 3495 Da protein in cirrhosis samples was one of the selected significant biomarkers. This 3495 Da protein was purified on-chip and trypsin digested. Further structural characterization of this biomarkers candidate was done by using cross-platform matrix-assisted laser desorption/ionization mass spectrometry (MALDI-MS) peptide mass fingerprinting (PMF) and matrix-assisted laser desorption/ionization time of flight/time of flight (MALDI-TOF/TOF) tandem mass spectrometry (MS/MS). Combined data from PMF and MS/MS spectra of two tryptic peptides suggested that this 3495 Da protein shared homology to a histidine-rich glycoprotein. These results demonstrated a novel approach to discovery of new biomarkers for early detection of liver cirrhosis and classification of liver diseases.

Algorithms↗