PubMed Health⌕ Search

Biomedical subjects

Yi-Xue Li

Publications and source records attributed to Yi-Xue Li.

At least 19 recordsLinked to original sources

Hierarchical modularity of nested bow-ties in metabolic networks.

BACKGROUND: The exploration of the structural topology and the organizing principles of genome-based large-scale metabolic networks is essential for studying possible relations between structure and functionality of metabolic networks. Topological analysis of graph models has often been applied to study the structural characteristics of complex metabolic networks. RESULTS: In this work, metabolic networks of 75 organisms were investigated from a topological point of view. Network decomposition of three microbes (Escherichia coli, Aeropyrum pernix and Saccharomyces cerevisiae) shows that almost all of the sub-networks exhibit a highly modularized bow-tie topological pattern similar to that of the global metabolic networks. Moreover, these small bow-ties are hierarchically nested into larger ones and collectively integrated into a large metabolic network, and important features of this modularity are not observed in the random shuffled network. In addition, such a bow-tie pattern appears to be present in certain chemically isolated functional modules and spatially separated modules including carbohydrate metabolism, cytosol and mitochondrion respectively. CONCLUSION: The highly modularized bow-tie pattern is present at different levels and scales, and in different chemical and spatial modules of metabolic networks, which is likely the result of the evolutionary process rather than a random accident. Identification and analysis of such a pattern is helpful for understanding the design principles and facilitate the modelling of metabolic networks.

Algorithms↗

Dynamic analysis of optimality in myocardial energy metabolism under normal and ischemic conditions.

To better understand the dynamic regulation of optimality in metabolic networks under perturbed conditions, we reconstruct the energetic-metabolic network in mammalian myocardia using dynamic flux balance analysis (DFBA). Additionally, we modified the optimal objective from the maximization of ATP production to the minimal fluctuation of the profile of metabolite concentration under ischemic conditions, extending the hypothesis of original minimization of metabolic adjustment to create a composite modeling approach called M-DFBA. The simulation results are more consistent with experimental data than are those of the DFBA model, particularly the retentive predominant contribution of fatty acid to oxidative ATP synthesis, the exact mechanism of which has not been elucidated and seems to be unpredictable by the DFBA model. These results suggest that the systemic states of metabolic networks do not always remain optimal, but may become suboptimal when a transient perturbation occurs. This finding supports the relevance of our hypothesis and could contribute to the further exploration of the underlying mechanism of dynamic regulation in metabolic networks.

Computer Simulation↗

Systematic analysis of head-to-head gene organization: evolutionary conservation and potential biological relevance.

Several "head-to-head" (or "bidirectional") gene pairs have been studied in individual experiments, but genome-wide analysis of this gene organization, especially in terms of transcriptional correlation and functional association, is still insufficient. We conducted a systematic investigation of head-to-head gene organization focusing on structural features, evolutionary conservation, expression correlation and functional association. Of the present 1,262, 1,071, and 491 head-to-head pairs identified in human, mouse, and rat genomes, respectively, pairs with 1- to 400-base pair distance between transcription start sites form the majority (62.36%, 64.15%, and 55.19% for human, mouse, and rat,respectively) of each dataset, and the largest group is always the one with a transcription start site distance of 101 to 200 base pairs. The phylogenetic analysis among Fugu, chicken, and human indicates a negative selection on the separation of head-to-head genes across vertebrate evolution, and thus the ancestral existence of this gene organization. The expression analysis shows that most of the human head-to-head genes are significantly correlated,and the correlation could be positive, negative, or alternative depending on the experimental conditions. Finally, head to-head genes statistically tend to perform similar functions, and gene pairs associated with the significant cofunctions seem to have stronger expression correlations. The findings indicate that the head-to-head gene organization is ancient and conserved, which subjects functionally related genes to correlated transcriptional regulation and thus provides an exquisite mechanism of transcriptional regulation based on gene organization. These results have significantly expanded the knowledge about head-to-head gene organization. Supplementary materials for this study are available at http://www.scbit.org/h2h.

Animals↗

Combining gene expression profiles and protein-protein interaction data to infer gene functions.

The ever-increasing flow of gene expression profiles and protein-protein interactions has catalyzed many computational approaches for inference of gene functions. Despite all the efforts, there is still room for improvement, for the information enriched in each biological data source has not been exploited to its fullness. A composite method is proposed for classifying unannotated genes based on expression data and protein-protein interaction (PPI) data, which extracts information from both data sources in novel ways. With the noise nature of expression data taken into consideration, importance is attached to the consensus expression patterns of gene classes instead of the actual expression profiles of individual genes, thus characterizing the composite method with enhanced robustness against microarray data variation. With regard to the PPI network, the traditional clear-cut binary attitude towards inter- and intra-functional interactions is abandoned, whereas a more objective perspective into the PPI network structure is formed through incorporating the varied function-function interaction probabilities into the algorithm. The composite method was implemented in two numerical experiments, where its improvement over single-data-source based methods was observed and the superiority of the novel data handling operations was discussed.

Algorithms↗

In silico discovery of human natural antisense transcripts.

BACKGROUND: Several high-throughput searches for potential natural antisense transcripts (NATs) have been performed recently, but most of the reports were focused on cis type. A thorough in silico analysis of human transcripts will help expand our knowledge of NATs. RESULTS: We have identified 568 NATs from human RefSeq RNA sequences. Among them, 403 NATs are reported for the first time, and at least 157 novel NATs are trans type. According to the pairing region of a sense and antisense RNA pair, hNATs are divided into 6 classes, of which about 87% involve 5' or 3' UTR sequences, supporting the regulatory role of UTRs. Among a total of 535 NAT pairs related with splice variants, 77.4% (414/535) have their pairing regions affected or completely eliminated by alternative splicing, suggesting significant relationship of alternative splicing and antisense-directed regulation. The extensive occurrence of splice variants in hNATs and other multiple pairing patterns results in a one-to-many relationship, allowing the formation of complex regulation networks. Based on microarray data from Stanford Microarray Database, two hNAT pairs were found to display significant inverse expression patterns before and after insulin injection. CONCLUSION: NATs might carry out more extensive and complex functions than previously thought. Combined with endogenous micro RNAs, hNATs could be regarded as a special group of transcripts contributing to the complex regulation networks.

Algorithms↗

Comparisons of graph-structure clustering methods for gene expression data.

Although many numerical clustering algorithms have been applied to gene expression data analysis, the essential step is still biological interpretation by manual inspection. The correlation between genetic co-regulation and affiliation to a common biological process is what biologists expect. Here, we introduce some clustering algorithms that are based on graph structure constituted by biological knowledge. After applying a widely used dataset, we compared the result clusters of two of these algorithms in terms of the homogeneity of clusters and coherence of annotation and matching ratio. The results show that the clusters of knowledge-guided analysis are the kernel parts of the clusters of Gene Ontology (GO)-Cluster software, which contains the genes that are most expression correlative and most consistent with biological functions. Moreover, knowledge-guided analysis seems much more applicable than GO-Cluster in a larger dataset.

Algorithms↗

EMMA: an efficient massive mapping algorithm using improved approximate mapping filtering.

Efficient massive mapping algorithm (EMMA), an algorithm on efficiently mapping massive cDNAs onto genomic sequences, has recently been developed. The process of mapping massive cDNAs onto genomic sequences has been improved using more approximate mapping filtering based on an enhanced suffix array coupled with a pruned fast hash table, algorithms of block alignment extensions, and k-longest paths. When compared with the classical BLAT software in this field, the computing of EMMA ranges from two to forty-one times faster under similar prediction precisions.

Algorithms↗

Network analysis of the protein chain tertiary structures of heterocomplexes.

In this paper, the tertiary structures of protein chains of heterocomplexes were mapped to 2D networks; based on the mapping approach, statistical properties of these networks were systematically studied. Firstly, our experimental results confirmed that the networks derived from protein structures possess small-world properties. Secondly, an interesting relationship between network average degree and the network size was discovered, which was quantified as an empirical function enabling us to estimate the number of residue contacts of the protein chains accurately. Thirdly, by analyzing the average clustering coefficient for nodes having the same degree in the network, it was found that the architectures of the networks and protein structures analyzed are hierarchically organized. Finally, network motifs were detected in the networks which are believed to determine the family or superfamily the networks belong to. The study of protein structures with the new perspective might shed some light on understanding the underlying laws of evolution, function and structures of proteins, and therefore would be complementary to other currently existing methods.

Models, Molecular↗

KDE Bioscience: platform for bioinformatics analysis workflows.

Bioinformatics is a dynamic research area in which a large number of algorithms and programs have been developed rapidly and independently without much consideration so far of the need for standardization. The lack of such common standards combined with unfriendly interfaces make it difficult for biologists to learn how to use these tools and to translate the data formats from one to another. Consequently, the construction of an integrative bioinformatics platform to facilitate biologists' research is an urgent and challenging task. KDE Bioscience is a java-based software platform that collects a variety of bioinformatics tools and provides a workflow mechanism to integrate them. Nucleotide and protein sequences from local flat files, web sites, and relational databases can be entered, annotated, and aligned. Several home-made or 3rd-party viewers are built-in to provide visualization of annotations or alignments. KDE Bioscience can also be deployed in client-server mode where simultaneous execution of the same workflow is supported for multiple users. Moreover, workflows can be published as web pages that can be executed from a web browser. The power of KDE Bioscience comes from the integrated algorithms and data sources. With its generic workflow mechanism other novel calculations and simulations can be integrated to augment the current sequence analysis functions. Because of this flexible and extensible architecture, KDE Bioscience makes an ideal integrated informatics environment for future bioinformatics or systems biology research.

Biological Science Disciplines↗

Cross-host evolution of severe acute respiratory syndrome coronavirus in palm civet and human.

The genomic sequences of severe acute respiratory syndrome coronaviruses from human and palm civet of the 2003/2004 outbreak in the city of Guangzhou, China, were nearly identical. Phylogenetic analysis suggested an independent viral invasion from animal to human in this new episode. Combining all existing data but excluding singletons, we identified 202 single-nucleotide variations. Among them, 17 are polymorphic in palm civets only. The ratio of nonsynonymous/synonymous nucleotide substitution in palm civets collected 1 yr apart from different geographic locations is very high, suggesting a rapid evolving process of viral proteins in civet as well, much like their adaptation in the human host in the early 2002-2003 epidemic. Major genetic variations in some critical genes, particularly the Spike gene, seemed essential for the transition from animal-to-human transmission to human-to-human transmission, which eventually caused the first severe acute respiratory syndrome outbreak of 2002/2003.

Amino Acid Substitution↗

MPSS: an integrated database system for surveying a set of proteins.

SUMMARY: We design and implement an integrated database system called 'multi-protein survey system' (MPSS), which provides a platform to retrieve information about many proteins at a time. This system integrates several important and widely used databases including SwissProt, TrEMBL, PDB and InterPro, plus useful references such as GO and KEGG to other databases. Users may submit a group of protein IDs, entry names, SwissProt/TrEMBL accession numbers or GenBank GIs through MPSS' web interface, and obtain protein annotation information from public databases and pre-computed molecular properties speedily. MPSS can also supply comprehensive information about query proteins, including 3D structures, domains, pathway, gene ontology and visual presentation of mapping to the GO tree and KEGG pathway, to provide an up-to-date view of available knowledge with regard to the structures and molecular functions of proteins under study. AVAILABILITY: MPSS is freely accessible at http://www.scbit.org/mpss/

Database Management Systems↗

Prediction of protein secondary structure using improved two-level neural network architecture.

In this paper we propose constructing an improved two-level neural network to predict protein secondary structure. Firstly, we code the whole protein composition information as the inputs to the first-level network besides the evolutionary information. Secondly, we calculate the reliability score for each residue position based on the output of the first-level network, and the role of the second-level network is to take full advantage of the residues with a higher reliability score to impact the neighboring residues with a lower one for improving the whole prediction accuracy. Thirdly, considering it is indeed a problem that the target protein can be lost in the multiple sequence alignment we propose to code single sequence into the second-level network. The experimental results show that our proposed method can efficiently improve the prediction accuracy.

Algorithms↗

Computational methods for protein-protein interaction and their application.

Protein-protein interactions play a central role in numerous processes in cell and are one of the main research fields in current functional proteomics. The increase of finished genomic sequences has greatly stimulated the progress for detecting the functions of the genes and their encoded proteins. As complementary ways to the high through-put experimental methods, various methods of bioinformatics have been developed for the study of the protein-protein interaction. These methods range from the sequence homology-based to the genomic-context based. Recently, it tends to integrate the data from different methods to build the protein-protein interaction network, and to predict the protein function from the analysis of the network structure. Efforts are ongoing to improve these methods and to search for novel aspects in genomes that could be exploited for function prediction. This review highlights the recent advances of the bioinformatics methods in protein-protein interaction researches. In the end, the application of the protein-protein interaction has also been discussed.

Computational Biology↗

[Analysis and application of SNP and haplotype in the human genome].

Single nucleotide polymorphism (SNP) is the most common type of genetic variant in human genome. Haplotype, defined as a specific set of alleles observed on a single chromosome, or a part of a chromosome,has been an integral part of human genetics for decades. The goal of the international HapMap project is to determine the common patterns of DNA sequence variation and find the Tag SNPs representing all SNPs in the human genome. Some studies demonstrated that the analyses of haplotype defined by the grouping and interaction of several variants rather than any individual SNP correlated with complex phenotypes. Here, we describe the definitions of SNPs, genotype, haplotype and some information of the HapMap project. In this review, we summarize the current three haplotype-inference methods, including Clark' method, EM algorithm and Byes approach, and the different defining methods for haplotype block, as well as the methods for choosing tag SNPs and association studies of complex diseases using haplotype. The major public SNP databases and applications of SNPs and haplotype in common complex diseases and drug response are also introduced in the paper.

Algorithms↗

A high-throughput approach for subcellular proteome: identification of rat liver proteins using subcellular fractionation coupled with two-dimensional liquid chromatography tandem mass spectrometry and bioinformatic analysis.

Four fractions from rat liver (a crude mitochondria (CM) and cytosol (C) fraction obtained with differential centrifugation, a purified mitochondrial (PM) fraction obtained with nycodenz density gradient centrifugation, and a total liver (TL) fraction) were analyzed with two-dimensional liquid chromatography tandem mass spectrometry analysis. A total of 564 rat proteins were identified and were bioinformatically annotated according to their physicochemical characteristics and functions. While most extreme alkaline ribosomal proteins were identified in the TL fraction, the C fraction mainly included neutral enzymes and the PM fraction enriched alkaline proteins and proteins with electron transfer activity or oxygen binding activity. Such characteristics were more apparent in proteins identified only in the TL, C, or PM fraction. The Swiss-Prot annotation and the bioinformatic prediction results proved that the C and PM fractions had enriched cytoplasmic or mitochondrial proteins, respectively. Combination usage of subcellular fractionation with two-dimensional liquid chromatography tandem mass spectrometry was proved to be a high-throughput, sensitive, and effective analytical approach for subcellular proteomics research. Using such a strategy, we have constructed the largest proteome database to date for rat liver (564 rat proteins) and its cytosol (222 rat proteins) and mitochondrial fractions (227 rat proteins). Moreover, the 352 proteins with Swiss-Prot subcellular location annotation in the 564 identified proteins were used as an actual subcellular proteome dataset to evaluate the widely used bioinformatics tools such as PSORT, TargetP, TMHMM, and GRAVY.

Animals↗

Scoring hidden Markov models to discriminate beta-barrel membrane proteins.

A new method is presented for identification of beta-barrel membrane proteins. It is based on a hidden Markov model (HMM) with an architecture obeying these proteins' construction principles. Once the HMM is trained, log-odds score relative to a null model is used to discriminate beta-barrel membrane proteins from other proteins. The method achieves only 10% false positive and false negative rates in a six-fold cross-validation procedure. The results compare favorably with existing methods. This method is proposed to be a valuable tool to quickly scan proteomes of entirely sequenced organisms for beta-barrel membrane proteins.

Algorithms↗

Semantic search among heterogeneous biological databases based on gene ontology.

Semantic search is a key issue in integration of heterogeneous biological databases. In this paper, we present a methodology for implementing semantic search in BioDW, an integrated biological data warehouse. Two tables are presented: the DB2GO table to correlate Gene Ontology (GO) annotated entries from BioDW data sources with GO, and the semantic similarity table to record similarity scores derived from any pair of GO terms. Based on the two tables, multifarious ways for semantic search are provided and the corresponding entries in heterogeneous biological databases in semantic terms can be expediently searched.

Database Management Systems↗