PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Elucidating the Mechanism of Xiaoqinglong Decoction in Chronic Urticaria Treatment: An Integrated Approach of Network Pharmacology, Bioinformatics Analysis, Molecular Docking, and Molecular Dynamics Simulations.

INTRODUCTION: Xiaoqinglong Decoction (XQLD) is a traditional Chinese medicinal formula commonly used to treat chronic urticaria (CU). However, its underlying therapeutic mechanisms remain incompletely characterized. This study employed an integrated approach combining network pharmacology, bioinformatics, molecular docking, and molecular dynamics simulations to identify the active components, potential targets, and related signaling pathways involved in XQLD's therapeutic action against CU, thereby providing a mechanistic foundation for its clinical application. METHODS: The active components of XQLD and their corresponding targets were identified using the Traditional Chinese Medicine Systems Pharmacology (TCMSP) database. CU-related targets were retrieved from the OMIM and GeneCards databases. Subsequently, core components and targets were determined via protein-protein interaction (PPI) network analysis and component-target-pathway network construction. Topological analyses were performed using Cytoscape software to prioritize core nodes within these networks. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were conducted via the DAVID database to identify enriched biological processes and signaling pathways. Molecular docking was performed to evaluate binding interactions between key components and core targets, while molecular dynamics (MD) simulations were employed to assess the stability of the component-target complexes with the lowest binding energy. Finally, CU-related targets of XQLD were validated using datasets from the Gene Expression Omnibus (GEO) database. RESULTS: A total of 135 active components and 249 potential targets of XQLD were identified, alongside 1,711 CU-related targets. Core components, such as quercetin, kaempferol, beta-sitosterol, naringenin, stigmasterol, and luteolin, exhibited high degree values in the constructed networks. The core targets identified included AKT1, TNF, IL6, TP53, PTGS2, CASP3, BCL2, ESR1, PPARG, and MAPK3. GO and KEGG pathway enrichment analyses revealed the PI3K-Akt signaling pathway as a central regulatory mechanism. Molecular docking studies demonstrated strong binding affinities between active components and core targets, with the stigmasterol-AKT1 complex exhibiting the lowest binding energy (-11.4 kcal/mol) and high stability in MD simulations. Validation using GEO datasets identified 12 core genes shared between CU-related targets and XQLD-associated targets, including PTGS2 and IL6, which were also prioritized as core targets in the network pharmacology analyses. DISCUSSION: This study comprehensively integrates multidisciplinary approaches to clarify the potential molecular mechanisms of XQLD in treating CU, highlighting its multitarget and multipathway synergistic effects. Molecular docking and dynamics simulations confirm the stable interaction between stigmasterol and the core target AKT1. Additionally, GEO dataset analysis verifies the pathogenic relevance of targets such as PTGS2 and IL6, significantly enhancing the credibility of our findings. These results provide a modern scientific basis for the traditional therapeutic effects of XQLD on CU and have important implications for developing multitarget treatments for this condition. However, this study mainly relies on database mining and computational simulations. Further in vitro and in vivo experimental validations are needed to confirm the predicted component-target-pathway interactions. CONCLUSION: This study identifies the active components, potential targets, and pathways through which XQLD exerts therapeutic effects on CU. These findings provide a theoretical foundation for further mechanistic studies and support their clinical application in the treatment of CU.

Molecular Docking Simulation↗

Integrated Bioinformatics Analysis Revealing that the NSDHL Gene Might Be Associated with the Progression of Western HFD/SW-Induced Hepatocellular Carcinoma.

BACKGROUND AND OBJECTIVE: Hepatocellular carcinoma (HCC) remains a significant global health concern. However, the etiology and pathogenesis of HCC have yet to be fully elucidated. Previous studies have indicated a close association between obesity and the occurrence and progression of HCC. The objective of this study was to employ bioinformatics strategies in order to explore key genes associated with the clinical diagnosis and prognosis of HCC induced by a Western high-fat diet and sugar water (HFD/SW). MATERIALS AND METHODS: We obtained the expression profile chip data GSE197884 from the Gene Expression Omnibus (GEO) database. Subsequently, “DESeq” and “Limma” R packages were employed to identify differentially expressed genes (DEGs) while constructing a co-expressed gene network using weighted gene co-expression analysis (WGCNA). Functional enrichment analyses were then carried out, followed by the construction of a protein-protein interaction (PPI) network to uncover core genes. The core genes were confirmed through data retrieved from The Cancer Genome Atlas (TCGA) database in order to determine their status as hub genes. Finally, survival and tumor immune infiltration analyses were performed to unveil the prognostic significance of these hub genes. RESULTS: In total, 126 intersection targets were retrieved through the Venn diagram. Gene ontology (GO) enrichment and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analyses revealed that the DEGs were primarily related to the proliferation and apoptosis of HCC cells, the digestion and metabolism of liver cells, the HCC tumor microenvironment, and immune response. The PPI network analysis identified 11 core targets, among which seven hub genes, including NSDHL, MVK, SQLW, GCAT, ALAS2, GLDC, and AGXT, were obtained after TCGA database validation. Furthermore, it was found that NSDHL was closely associated with the clinical diagnosis and prognosis of HCC induced by HFD/SW and also affected the cellular immune infiltration in the HCC tumor microenvironment. CONCLUSION: The present study demonstrated a significantly elevated expression of NSDHL in HCC tissues, suggesting its potential as a specific biomarker for precise clinical diagnosis and prognosis assessment of HCC induced by HFD/SW.

Computational Biology↗

Genome-Wide Identification and Bioinformatics Analysis of the FAD Gene Family in Walnut (Juglans regia L.).

Fatty acid desaturase (FAD) is a core catalytic enzyme in plants for the synthesis of unsaturated fatty acids, profoundly affecting plant growth, development, and adaptability to various environmental stresses. The walnut (Juglans regia L.) is an important woody oil tree species, and its kernel is rich in unsaturated fatty acids. Systematic identification of the walnut FAD gene family and analysis of its function are of great significance for revealing the molecular mechanisms underlying unsaturated fatty acid metabolism in the walnut. Based on walnut whole-genome data, this study used homology alignment and hidden Markov model search methods to identify the JrFAD gene family members. Subsequently, a variety of bioinformatics tools were used to systematically analyze their structural characteristics, evolutionary expansion mechanism, expression regulation, and function. A total of 21 JrFAD gene family members were identified and classified into five subfamilies. The family genes were unevenly distributed on nine chromosomes. WGD/segmental duplication was the main expansion method, and the duplicated gene pairs experienced strong purification selection. The family gene promoter sequence is rich in regulatory elements that respond to light, plant hormones, and various stresses. The expression pattern analysis showed that JrFAD3.1 and JrFAD2.3 showed high expression specifically during the rapid accumulation of walnut kernel oil. This study clarified the composition and evolutionary characteristics of the FAD gene family in the walnut, which provides useful information for in-depth analyses of its functional mechanism in the regulation of lipid metabolism, and also identified potential candidate gene resources for the genetic improvement of walnut varieties with high amounts of unsaturated fatty acids.

Juglans↗

Comparative bioinformatic analysis of complete proteomes and protein parameters for cross-species identification in proteomics.

Peptide mass fingerprinting (PMF) remains the most amenable technique for protein identification in proteomics, using mass spectrometry as the primary analytical technique coupled with bioinformatics. This relies on the presence of the amino acid sequence of the protein in the current databanks. Despite this, it is desirable to be able to use the technique for organisms whose genomes are not yet fully sequenced and apply cross-species protein identification. In this study, we have re-examined the feasibility of such approaches by considering the extent of protein similarity between genome sequences using a data set of 29 complete bacterial and two eukaryotic genomes. A range of protein and peptide features are considered, including protein isoelectric focussing point, protein mass, and amino acid conservation. The effectiveness of PMF approaches has then been tested with a series of computer simulations with varying peptide number and mass accuracy for several cross-species tests. The results show that PMF alone is unsuitable in general for divergent species jumps, or when protein similarity is less than 70% identity. Despite this, there exists a considerable enrichment above random of tryptic peptide conservation and PMF promises to remain useful when combined with other data than just peptide masses for cross-species protein identification.

Algorithms↗

Comparative proteome bioinformatics: identification of a whole complement of putative protein tyrosine kinases in the model flowering plant Arabidopsis thaliana.

Phosphorylation by protein tyrosine kinases is crucial to the control of growth and development of multicellular eukaryotes, including humans, and it also seems to play an important role in multicellular prokaryotes. A plant tyrosine-specific kinase has not been identified yet; hence, plants have been suggested to share with unicellular eukaryote yeast a tyrosine phosphorylation system where a limited number of stress proteins are tyrosyl-phosphorylated only by a few dual-specificity (serine/threonine and tyrosine) kinases. However, preliminary evidence obtained so far suggests that tyrosine phosphorylation in plants depends on the developmental conditions. Since sequencing of the genome of the model flowering plant Arabidopsis thaliana has been recently completed, we have performed a bioinformatic screening of the whole Arabidopsis proteome to identify a model complement of bona fide protein tyrosine kinases. In silico analyses suggest that < 4% of Arabidopsis kinases are tyrosine-specific kinases, whose gene expression has been assessed by a preliminary polymerase chain reaction screening of an Arabidopsis cDNA library. Finally, immunological evidence confirms that the number of Arabidopsis proteins specifically phosphorylated on tyrosine residues is much higher than in yeast.

Amino Acid Sequence↗

Candidate genes for nicotine dependence via linkage, epistasis, and bioinformatics.

Many smoking-related phenotypes are substantially heritable. One genome scan of nicotine dependence (ND) has been published and several others are in progress and should be completed in the next 5 years. The goal of this hypothesis-generating study was two-fold. First, we present further analyses of our genome scan data for ND published by Straub et al. [1999: Mol Psychiatry 4:129-144] (PMID: 10208445). Second, we used the method described by Cox et al. [1999: Nat Genet 21:213-215] (PMID: 9988276) to search for epistatic loci across the markers used in the genome scan. The overall results of the genome scan nearly reached the rigorous Lander and Kruglyak [1995: Nat Genet 11:241-247] criteria for "significant" linkage with the best findings on chromosomes 10 and 2. We then looked for correspondence between genes located in the 10 regions implicated in affected sibling pair (ASP) and epistatic linkage analyses with a list of genes suggested by microarray studies of experimental nicotine exposure and candidate genes from the literature. We found correspondence between linkage and microarray/candidate gene studies for genes involved with the mitogen-activated protein kinase (MAPK) signaling system, nuclear factor kappa B (NFKB) complex, neuropeptide Y (NPY) neurotransmission, a nicotinic receptor subunit (CHRNA2), the vesicular monoamine transporter (SLC18A2), genes in pathways implicated in human anxiety (HTR7, TDO2, and the endozepine-related protein precursor, DKFZP434A2417), and the micro 1-opioid receptor (OPRM1). Although the hypotheses resulting from these linkage and bioinformatic analyses are plausible and intriguing, their ultimate worth depends on replication in additional linkage samples and in future experimental studies.

Chromosome Mapping↗

Ontologies: Formalising biological knowledge for bioinformatics.

An ontology is a domain of knowledge structured through formal rules so that it can be interpreted and used by computers. Ontologies are becoming increasingly important in bioinformatics because they can be linked to the information in databases and their knowledge then used to query the databases. Typical examples in current use are the Gene Ontology, which incorporates much of our knowledge about gene products, and ontologies of developmental anatomy, which, for example, facilitate tissue-based queries to gene expression databases both textually and spatially. This article considers the production, formulation and types of bio-ontologies together with the reasons why they are so useful.

Animals↗

Serial analysis of gene expression (SAGE): unraveling the bioinformatics tools.

Serial analysis of gene expression (SAGE) is a powerful technique that can be used for global analysis of gene expression. Its chief advantage over other methods is that it does not require prior knowledge of the genes of interest and provides qualitative and quantitative data of potentially every transcribed sequence in a particular cell or tissue type. This is a technique of expression profiling, which permits simultaneous, comparative and quantitative analysis of gene-specific, 9- to 13-basepair sequences. These short sequences, called SAGE tags, are linked together for efficient sequencing. The sequencing data are then analyzed to identify each gene expressed in the cell and the levels at which each gene is expressed. The main benefit of SAGE includes the digital output and the identification of novel genes. In this review, we present an outline of the method, various bioinformatics methods for data analysis and general applications of this important technology.

Animals↗

The physics and bioinformatics of binding and folding-an energy landscape perspective.

It has been recognized in the last few years that unstructured proteins play an important role in biological organisms, often participating in signal transduction, transcriptional regulation, and a variety of other regulatory activities. Various hypotheses have been put forward for the ubiquity of the unfolded state; rapid turnover, faster or more specific binding kinetics, multifunctionality may all possibly explain apparent ubiquitousness of unfolded proteins in eukaryotic cells. In this paper we extend the energy landscape theory of protein folding to construct an analytical model of how binding and folding are coupled thermodynamically when the energy landscape is partially rugged. To deduce the parameters that enter the theory, which is based on Generalized Random Energy Model, we have analyzed in a bioinformatic sense a large structural database of more than 500 protein complexes. We find that Miyazawa-Jernigan contact potential shows similar energy gaps for folding for both hydrophobic and hydrophilic proteins, but that for binding contacts hydrophobic interfaces turn out to be funneled while hydrophilic ones are antifunneled. This suggests evolution has found a mechanism for avoiding frustration between folding and binding by making use of indirect water-mediated interactions. By juxtaposing the monomeric protein folding free energy profile in the protein complex database with another database consisting of only well-folded monomers, we estimate that at least 15% of monomers in the former database are unfolded in the absence of partner protein interface interactions. When employing the parameters characteristic of these unfolded monomers to construct binding/folding phase diagrams, we find that these monomers would indeed fold if sufficiently stabilizing binding contacts, consistent with that fold, are formed.

Computational Biology↗

Aligning experimental design with bioinformatics analysis to meet discovery research objectives.

The utility of genomic technology and bioinformatic analytical support to provide new and needed insight into the molecular basis of disease, development, and diversity continues to grow as more research model systems and populations are investigated. Yet deriving results that meet a specific set of research objectives requires aligning or coordinating the design of the experiment, the laboratory techniques, and the data analysis. The following paragraphs describe several important interdependent factors that need to be considered to generate high quality data from the microarray platform. These factors include aligning oligonucleotide probe design with the sample labeling strategy if oligonucleotide probes are employed, recognizing that compromises are inherent in different sample procurement methods, normalizing 2-color microarray raw data, and distinguishing the difference between gene clustering and sample clustering. These factors do not represent an exhaustive list of technical variables in microarray-based research, but this list highlights those variables that span both experimental execution and data analysis.

Computational Biology↗

Glycosylation patterns of human chorionic gonadotropin revealed by liquid chromatography-mass spectrometry and bioinformatics.

Due to their extensive structural heterogeneity, the elucidation of glycosylation patterns in glycoproteins such as the subunits of human chorionic gonadotropin (hCG), hCG-alpha, and hCG-beta, remains one of the most challenging problems in the proteomic analysis of post-translational modifications. In consequence, glycosylation is usually studied after decomposition of the intact proteins to the proteolytic peptide level. However, by this approach all information about the combination of the different glycopeptides in the intact protein is lost. In this study we have, therefore, attempted to combine the results of glycan identification after tryptic digestion with molecular mass measurements on the native starting material of the new first WHO Reference Reagents (RR) for hCG-alpha (99/720) and hCG-beta (99/650). Despite the extremely high number of possible combinations of the glycans identified in the tryptic peptides by HPLC-MS (>1000 for hCG-alpha and >10 000 for hCG-beta), the mass spectra of intact hCG-alpha and hCG-beta revealed only a limited number of glycoforms present in hCG preparations from pools of pregnancy urines. Peak annotations for hCG-alpha were performed with the help of a bioinformatic algorithm that generated a database containing all possible modifications of the proteins, including modifications possibly introduced during sample preparation such as oxidation or truncation, for subsequent searches for combinations fitting the mass difference between the polypeptide backbone and the measured molecular masses. Fourteen different glycoforms of hCG-alpha, containing biantennary, partly sialylized hybrid-type glycans, including methionine-oxidized and N-terminally truncated forms, were identified. Mass spectra of high quality were also obtained for hCG-beta, however, a database search mass accuracy of +/-5 Da was insufficient to unambiguously assign the possible combinations of post-translational modifications. In summary, mass spectrometric fingerprints of intact molecules were shown to be highly useful for the characterization of glycosylation patterns of different hCG preparations such as the new first WHO RR for immunoassays and could be the first step in establishing biophysical reference methods for hCG and related molecules.

Chorionic Gonadotropin↗

Bioinformatic analysis of protein structure-function relationships: case study of leukocyte elastase (ELA2) missense mutations.

Cyclic and congenital neutropenia are caused by mutations in the human neutrophil elastase (HNE) gene (ELA2), leading to an immunodeficiency characterized by decreased or oscillating levels of neutrophils in the blood. The HNE mutations presumably cause loss of enzyme activity, consequently leading to compromised immune system function. To understand the structural basis for the disease, we implemented methods from bioinformatics to analyze all the known HNE missense mutations at both the sequence and structural level. Our results demonstrate that the 32 different mutations have diverse effects on HNE structure and function, affecting structural disorder and aggregation tendencies, stability maintaining contacts, and electrostatic properties. A large proportion of the mutations are located at conserved amino acids, which are usually essential in determining protein structure and function. The majority of the disease-causing HNE missense mutations lead to major structural changes and loss of stability in the protein. A few mutations also affect functional residues, leading into decreased catalytic activity or altered ligand binding. Our analysis reveals the putative effects of all known missense mutations in HNE, thus allowing the structural basis of cyclic and congenital neutropenia to be elucidated. We have employed and analyzed a set of some 30 different methods for predicting the effects of amino acid substitutions. We present results and experience from the analysis of the applicability of these methods in the analysis of numerous genes, proteins, and diseases to reveal protein structure-function relationships and disease genotype-phenotype correlations.

Antigens, Surface↗

Structural bioinformatic approaches to understand cross-reactivity.

Cross-reactivity of allergens results from the presence of antibody-accessible conserved surface structures. These can best be studied when allergens have been structurally defined by X-ray crystallography or another structure determination method. When this is not the case, mimotope technology provides a useful alternative for elucidating antibody-binding sites on allergens. Structural bioinformatic approaches have been used to study the cross-reactivity of inhalant allergens with labile food allergens (Bet v 1 family) as well as the cross-reactivity between stable food allergens such as members of the nonspecific lipid transfer protein family. It was found that the degree of similarity of the structures correlated with the observed IgE cross-reactivities. However, IgE cross-reactivity between structurally unrelated allergens has not been demonstrated to date.

Allergens↗

Evaluation of available IgE-binding epitope data and its utility in bioinformatics.

This paper reviews the role played by IgE-binding epitopes in eliciting clinical symptoms, the types of IgE-binding epitopes in allergenic proteins, the methods used to identify IgE-binding epitopes, and the availability of IgE-binding epitopes in allergenic sources. Finally, bioinformatics methods to assess protein allergenicity using knowledge of IgE-binding epitopes are discussed.

Allergens↗

PhosphoSite: A bioinformatics resource dedicated to physiological protein phosphorylation.

PhosphoSite is a curated, web-based bioinformatics resource dedicated to physiologic sites of protein phosphorylation in human and mouse. PhosphoSite is populated with information derived from published literature as well as high-throughput discovery programs. PhosphoSite provides information about the phosphorylated residue and its surrounding sequence, orthologous sites in other species, location of the site within known domains and motifs, and relevant literature references. Links are also provided to a number of external resources for protein sequences, structure, post-translational modifications and signaling pathways, as well as sources of phospho-specific antibodies and probes. As the amount of information in the underlying knowledgebase expands, users will be able to systematically search for the kinases, phosphatases, ligands, treatments, and receptors that have been shown to regulate the phosphorylation status of the sites, and pathways in which the phosphorylation sites function. As it develops into a comprehensive resource of known in vivo phosphorylation sites, we expect that PhosphoSite will be a valuable tool for researchers seeking to understand the role of intracellular signaling pathways in a wide variety of biological processes.

Animals↗

Proteome informatics II: bioinformatics for comparative proteomics.

The present review attempts to cover the most recent initiatives directed towards representing, storing, displaying and processing protein-related data suited to undertake "comparative proteomics" studies. Data interpretation is brought into focus. Efforts invested into analysing and interpreting experimental data increasingly express the need for adding meaning. This trend is perceptible in work dedicated to determining ontologies, modelling interaction networks, etc. In parallel, technical advances in computer science are spurred by the development of the Web and the growing need to channel and understand massive volumes of data. Biology benefits from these advances as an application of choice for many generic solutions. Some examples of bioinformatics solutions are discussed and directions for on-going and future work conclude the review.

Algorithms↗

Integrative proteomics: structure, function, and interaction report on the 3rd joint meeting of the British Society for Proteome Research and the European Bioinformatics Institute, July 2006.

This report summarizes the highlights of the recent British Society for Proteome Research (BSPR) meeting jointly organized with the European Bioinformatics Institute (EBI) which was held at the Wellcome Trust Genome Campus, Hinxton, Cambridge, UK in July 2006. This was the third annual scientific meeting organized by the BSPR and EBI and the theme of this years meeting was Integrative Proteomics: Structure, function and interaction. A wealth of local and overseas speakers were invited to discuss both their own work and specific challenges present in modern day proteomic based experiments.

Computational Biology↗

Subunit E of mitochondrial ATP synthase: a bioinformatic analysis reveals a phosphopeptide binding motif supporting a multifunctional regulatory role and identifies a related human brain protein with the same motif.

The mitochondrial adenosine triphosphate (ATP) synthase is located in the inner membrane and consists of at least 16 subunit types in animals, one of which is subunit e, the function of which is not clearly defined. A highly homologous protein is located in the nucleus and named progesterone receptor binding protein (RBF), to designate its role in this organelle. In addition, the expression level of subunit e in mammalian cells fluctuates greatly and is induced by certain carcinogens and elevated in liver cancers. Because these previous observations suggested to us that subunit e may play multifunctional regulatory roles, we employed a bioinformatic approach to test this view. First, from sequence alignment studies, secondary structure analyses, and basic local alignment search tool (BLAST) searches, we concluded that mitochondrial subunit e and the homologous nuclear protein RBF are most likely the same protein. Second, we examined the known sequence and structure of one of the most common multifunctional cell regulatory proteins, the 14-3-3 protein, involved in phosphopeptide binding, and deduced that it has an apparent binding motif (-KX(6)R---RY-). Third, from careful examination of the conserved residues within all subunit e sequences in the database, we discovered that this protein has a comparable binding motif (-RY---KX(6)R-). Finally, in a BLAST search for additional homologs of subunit e, we found a human brain protein, KIAA1578, the C-terminal 30 amino acids of which are identical to those of human subunit e. This protein also contains a potential phosphopeptide binding motif. In summary, these studies provide support for the view that subunit e is a multifunctional cell regulator involved in cell signaling, and implicate the involvement of the KIAA1578 protein in cell signaling as well. These studies suggest also that, while functioning as a subunit of mitochondrial ATP synthases, subunit e may help regulate these complexes by binding to phosphopeptides within one or more of the other subunit types.

14-3-3 Proteins↗