PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Bioinformatics: use in bacterial vaccine discovery.

Bioinformatics has now become a common laboratory name for groups studying genomic sequences. It is composed of many different, yet interrelated scientific fields such as genomics, proteomics, and transcriptional profiling. The availability of complete genomic sequences, especially prokaryotic organisms, allows one to rapidly identify, analyze, and clone genes of interest. For bacterial vaccine discovery, one can "mine" the genomic sequence for potential surface targets using various algorithms, characterize these gene targets, and produce primers for cloning, all before one enters the wet laboratory. This review will focus on various genomic mining tools/algorithms available for predicting open reading frames and their associated annotation (if known), physical and functional characterization, and cellular localization. Finally, examples are given of how all of this is being used for the identification of potential bacterial vaccine candidates.

Animals↗

Structural bioinformatics: methods, concepts and applications to blood coagulation proteins.

Structural and theoretical analyses of proteins are central to the understanding of complex molecular mechanisms and are fundamental to the drug discovery process. Computational techniques yield useful insights into an ever-wider range of biomolecular systems. Protein three-dimensional structures and molecular functions can be predicted in some circumstances, while experimental structures can be analyzed in depth via such computational approaches. Non-covalent binding of biomolecules can be understood by considering structural, thermodynamic and kinetic issues, and theoretical simulations of such events can be attempted. The central role of electrostatic interactions with regard to protein function, structure and stability has been investigated and some electrostatic properties can be modeled theoretically. Computer methods thus help to prioritize, design, analyze and rationalize biochemical experiments. Cardiovascular diseases and associated blood coagulation disorders are leading causes of death worldwide. Blood coagulation involves more than 30 proteins that interact specifically with various degrees of affinity. Many of these molecules can also bind transiently to phospholipid surfaces. Numerous point mutations in the genes of coagulation proteins and regulators have been identified. Understanding the coagulation cascade, its regulation and the impact of mutations is required for the development of new therapies and diagnostic tools. In this review, we describe concepts and methods pertaining to the field of structural bioinformatics. We provide examples of applications of these approaches to blood coagulation proteins and show that such studies can give insights about molecular mechanisms contributing to cardiovascular disease susceptibility.

Blood Coagulation↗

Bridging the gap between structural bioinformatics and receptor research: the membrane-embedded, ligand-gated, P2X glycoprotein receptor.

MOTIVATION: No details on P2X receptor architecture had been known at the atomic resolution level. Using comparative homology-based molecular modelling and threading, it was attempted to predict the three-dimensional structure of P2X receptors. This prediction could not be carried out, however, because important properties of the P2X family differ considerably from that of the potential template proteins. This paper reviews an alternative approach consisting of three research fields: bioinformatics, structural modelling, and a variety of the results of biological experiments. MODEL: Starting point is the amino acid sequence. Using the sequential data, the first step is a secondary structure prediction. The resulting secondary structure is converted into a three-dimensional geometry. Then, the secondary and tertiary structures are optimized by using the quantum chemistry RHF/3-21G minimal basic set and the all-atom molecular mechanics AMBER96 force field. The fold of the membrane-embedded protein is simulated by a suitable dielectricum. The structure is refined using a conjugate gradient minimizer (Fletcher-Reeves modification of the Polak-Ribiere method). The results of the geometry optimization were checked by a Ramanchandran plot, rotamer analysis, all-atom contact dots, and the C(beta) deviation. As additional tools for the model building, multiple alignment analysis and comparative sequence-function analysis were used. The approach is exemplified on the membrane-embedded, ligand-gated P2X3 receptor subunit, a monovalent-bivalent cation channel-forming glycoprotein that is activated by extracellular adenosine 5'-triphosphate. From these results, a topology of the pore-forming motif of the P2X3 receptor subunit was proposed. It is believed that a fully functional P2X channel requires a precise coupling between (i) two distinct peptide modules, an extracellularly occurring ATP-binding module and a pore module that includes a long transmembrane and short intracellular part, (ii) an interaction surface with membranes, and (iii) hydrogen bonding forces of the residues and hydrated cations. Furthermore, this paper demonstrates the role of quantitative structure-activity relationships (QSARs) in P2X research (calcium ion permeability of the wild-type and after site-directed mutagenesis of the rat P2X2 receptor protein, KN-62 analogs as competitive antagonists of the human P2X7 receptor). EXPERIMENTAL PROOFS: The predictions are experimentally testable and may provide an additional interpretation of experimental observations published in literature. In particular, there is the good agreement of the geometry optimized P2X3 structure with experimentally proposed P2X receptor models obtained by neurophysiological, biochemical, pharmacological, and mutation experiments. Although the rat P2X3 receptor subunit is more complex (397 amino acids) than the KcsA protein (160 amino acids), the overall folds of the peptide backbone atoms are similar. LIMITATIONS: To avoid semantic confusion, it should be noted that "prediction" is defined in a probabilistic sense. Matches to generic rules do not mean "this is true" but rather "this might be true". Only biological and chemical knowledge can determine whether or not these predictions are meaningful. Thus, the results from the computational tools are probabilistic predictions and subject to further experimental verification. AVAILABILITY: The geometry optimized P2X3 receptor subunit is freely available for academic researchers on e-mail request (PDB format).

Amino Acid Sequence↗

Reproducible research: a bioinformatics case study.

While scientific research and the methodologies involved have gone through substantial technological evolution the technology involved in the publication of the results of these endeavors has remained relatively stagnant. Publication is largely done in the same manner today as it was fifty years ago. Many journals have adopted electronic formats, however, their orientation and style is little different from a printed document. The documents tend to be static and take little advantage of computational resources that might be available. Recent work, Gentleman and Temple Lang (2003), suggests a methodology and basic infrastructure that can be used to publish documents in a substantially different way. Their approach is suitable for the publication of papers whose message relies on computation. Stated quite simply, Gentleman and Temple Lang (2003) propose a paradigm where documents are mixtures of code and text. Such documents may be self-contained or they may be a component of a compendium which provides the infrastructure needed to provide access to data and supporting software. These documents, or compendiums, can be processed in a number of different ways. One transformation will be to replace the code with its output -- thereby providing the familiar, but limited, static document. In this paper we apply these concepts to a seminal paper in bioinformatics, namely The Molecular Classification of Cancer, Golub et al (1999). The authors of that paper have generously provided data and other information that have allowed us to largely reproduce their results. Rather than reproduce this paper exactly we demonstrate that such a reproduction is possible and instead concentrate on demonstrating the usefulness of the compendium concept itself.

Journal Article↗

Functional bioinformatics: the cellular response database.

Biological Scientists function in an increasingly data rich environment. The emerging field of bioinformatics is attempting to insure that this flow of information can be structured to support the generation of significant biological hypothesis and ultimately new knowledge. To date, most of the current databases have focused on protein and nucleic acid sequence information as the principle type of data stored for further interpretation. In this paper, we describe the Cellular Response Database. This database stores functional information regarding the changes of cellular gene expression associated with various stimuli, and supports queries linking cell types, expressed genes, and inducers. The database is designed to support information-intensive queries to aid in the determination of biological function, and is flexible enough to allow the storage of a broad range of experimental data such as cytotoxicity data, immunoassays of target gene protein expression, and others.

Cell Physiological Phenomena↗

Bioinformatics approaches for detecting gene-gene and gene-environment interactions in studies of human disease.

Neurological and mental disorders occur often, with approximately 450 million people suffering from them worldwide. Like most other common diseases, neurological disorders are hypothesized to be highly complex, with interactions among genes and risk factors playing a major role in the process. In recent years it has become obvious that for common diseases there may be more complex interactions among genes with and without strong independent main effects. These effects are more difficult to detect using traditional methodologies. In this manuscript the author introduces the concept of epistasis and the challenges associated with detecting it. Next, she briefly mentions a number of bioinformatics approaches that have been developed to deal with this issue. Multifactor dimensionality reduction is a methodology developed specifically to deal with the challenge of detecting interaction effects in the absence of statistically detectable main effects in studies of common disorders, such as Alzheimer disease or brain cancer. Finally, the author describes the future directions for this technique and related methodologies.

Computational Biology↗

Spoligologos: a bioinformatic approach to displaying and analyzing Mycobacterium tuberculosis data.

Spacer oligonucleotide (spoligotyping) analysis is a rapid polymerase chain reaction-based method of DNA fingerprinting the Mycobacterium tuberculosis complex. We examined spoligotype data using a bioinformatic tool (sequence logo analysis) to elucidate undisclosed phylogenetic relationships and gain insights into the global dissemination of strains of tuberculosis. Logo analysis of spoligotyping data provides a simple way to describe a fingerprint signature and may be useful in categorizing unique spoligotypes patterns as they are discovered. Large databases of DNA fingerprint information, such as those from the U.S. National Tuberculosis Genotyping and Surveillance Network and the European Concerted Action on Tuberculosis, contain information on thousands of strains from diverse regions. The description of related spoligotypes has depended on exhaustive listings of the individual spoligotyping patterns. Logo analysis may become another useful graphic method of visualizing and presenting spoligotyping clusters from these databases.

Bacterial Typing Techniques↗

Exploring the Role of HSD17B2 in Colorectal Cancer Through Bioinformatic Analysis: Preliminary Insights for Prognostic Evaluation.

Colorectal cancer (CRC) is the third most commonly diagnosed cancer and the second leading cause of cancer-related mortality worldwide. Although screening has reduced CRC in older adults, cases in younger individuals are rising, highlighting the need for early biomarkers. Emerging research highlights the role of estrogen metabolism in CRC progression, with enzymes such as hydroxysteroid (17-beta) dehydrogenase (HSD17B) being increasingly implicated. In this study, we performed a bioinformatics analysis using publicly available datasets, including The Cancer Genome Atlas Colon Adenocarcinoma (TCGA-COAD) cohort and two independent Gene Expression Omnibus (GEO) cohorts (GSE40967 and GSE41258), to investigate the role of HSD17B enzymes in CRC. Our results suggest that HSD17B2 is frequently downregulated in precancerous lesions and early-stage CRC, which may contribute to elevated estradiol levels and a tumor-promoting microenvironment. In advanced stages, higher HSD17B2 expression levels are associated with poorer survival outcomes in retrospective cohorts. Other HSD17B enzymes also exhibit significant expression changes, further complicating the hormonal landscape of CRC. In addition, estrone, traditionally considered a weaker estrogen, emerges as a potential driver of CRC progression. Our in-silico analyses indicate that HSD17B2 and HSD17B11 warrant further investigation as candidate biomarkers for distinguishing CRC from benign and precancerous conditions, with the combination showing strong discriminatory power in Receiver Operating Characteristic (ROC) analyses. Overall, these findings highlight the potential role of estrogen metabolism in CRC and suggest that HSD17B enzymes may hold value as candidate prognostic and diagnostic indicators, though their clinical utility remains hypothetical at this stage. Experimental and clinical validation is strictly required to confirm these in silico observations and to clarify their mechanisms in CRC.

Humans↗

[The optimal combination of serum tumor markers with bioinformatics in diagnosis of colorectal carcinoma].

OBJECTIVE: To identify the optimal combination of serum tumor markers with bioinformatics in diagnosis of colorectal cancer. METHODS: The serum levels of CEA, AFP, NSE, CA199, CA242, CA724, CA211 and TPA were detected in 128 patients with colorectal carcinoma and 113 health subjects. The serum tumor markers were evaluated with the area under curves. The optimal combination of serum tumor markers was selected and the diagnostic model with artificial neural network was established. RESULTS: CEA, CA199, CA242, CA211, CA724 were selected for the optimal combination and the artificial neural network was built. The model was evaluated by a 5-cross validation approach. The model had a specificity of 95%, sensitivity of 83% and positive predictive value of 95% in diagnosis of colorectal carcinoma. CONCLUSION: The combination of optimal serum tumor markers has a high sensitivity and specificity in diagnosis of colorectal carcinoma.

Adult↗

Bioinformatics and Quantitative Real-Time Polymerase Chain Reaction Analysis of SUCNR1 and GPR37L1 in Schizophrenia.

Schizophrenia is a severe, complex, and multifactorial mental disorder involving numerous genetic susceptibility elements, leading to substantial disability, morbidity, and mortality. Despite significant progress in understanding its pathophysiology and etiology, specific diagnostic biomarkers for schizophrenia remain elusive. This study aimed to identify candidate molecular markers associated with schizophrenia. An integrated bioinformatics analysis was performed on the public microarray dataset GSE54913. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analyses revealed that the most significantly enriched GO terms were related to channel activity, including passive transmembrane transporter activity, ion channel activity, gated channel activity, and substrate-specific channel activity. The top five enriched KEGG pathways were insulin secretion, cAMP signaling pathway, nucleotide excision repair, TNF signaling pathway, and glutathione metabolism. Validation was conducted using quantitative real-time polymerase chain reaction (qRT-PCR) on an independent sample set from Wuhan Rongjun Youfu Hospital. The qRT-PCR results were largely consistent with the microarray analysis (Pearson r = 0.89, 95% CI: 0.66-0.97). Protein-protein interaction (PPI) network analysis identified two hub genes, SUCNR1 and GPR37L1, which were significantly associated with the GO term 'ion channel activity' and enriched in the KEGG pathway 'insulin secretion'. Furthermore, SUCNR1 expression showed a negative correlation with verbal memory scores (r = -0.54, P = 0.015), whereas GPR37L1 expression showed a positive correlation (r = 0.59, P = 0.0034). These findings suggest that altered SUCNR1 and GPR37L1 expression may be associated with schizophrenia and may represent candidate molecular markers for further investigation.

Humans↗

Pan-cancer Bioinformatics Analysis Combined with Colon Cancer Experimental Validation: A Study on TMED3 as a Diagnostic and Prognostic Biomarker.

Transmembrane Emp24 Protein Transport Domain 3 (TMED3), a member of the p24 protein family, has been implicated in tumor proliferation, invasion, and migration. This study aimed to evaluate the expression patterns, prognostic significance, immune associations, and potential biological functions of TMED3 across multiple cancer types using pan-cancer bioinformatics analysis combined with immunohistochemical (IHC) validation in colon cancer. Multiomics datasets from The Cancer Genome Atlas, Genotype-Tissue Expression, UALCAN, Human Protein Atlas, and cBioPortal databases were analyzed to investigate TMED3 expression and genetic alterations in pan-cancer. Immunohistochemistry was performed to evaluate TMED3 protein expression in colon cancer tissues. Kaplan-Meier survival analysis and Cox regression analysis were used to assess the prognostic value of TMED3. Spearman correlation analysis was conducted to evaluate the associations of TMED3 with tumor mutational burden, microsatellite instability (MSI), immune cell infiltration, and immune checkpoints. Gene Set Enrichment Analysis was performed to investigate potential biological pathways associated with TMED3 in colon cancer. TMED3 expression was elevated in most tumor types and was associated with unfavorable overall survival and disease-specific survival in adrenocortical carcinoma, colon adenocarcinoma, and uveal melanoma. The greatest frequency of TMED3 genetic alterations was identified in mesothelioma, with amplification representing the predominant alteration type. In addition, TMED3 expression showed significant correlations with tumor mutational burden and microsatellite instability in kidney renal clear cell carcinoma, stomach adenocarcinoma, and uterine corpus endometrial carcinoma. TMED3 expression was also associated with immune infiltration and immune checkpoint expression in several tumors. IHC analysis demonstrated increased TMED3 expression in colon cancer tissues compared with normal colon tissues and showed an association with T stage. Functional enrichment analysis identified pathways related to ribosome, antigen processing and presentation, oxidative phosphorylation, and pentose phosphate. These findings indicate that TMED3 may represent a promising biomarker for the diagnosis and prognostic evaluation of colon cancer as well as other tumor types.

Humans↗

Coupling in silico and in vitro analysis of peptide-MHC binding: a bioinformatic approach enabling prediction of superbinding peptides and anchorless epitopes.

The ability to define and manipulate the interaction of peptides with MHC molecules has immense immunological utility, with applications in epitope identification, vaccine design, and immunomodulation. However, the methods currently available for prediction of peptide-MHC binding are far from ideal. We recently described the application of a bioinformatic prediction method based on quantitative structure-affinity relationship methods to peptide-MHC binding. In this study we demonstrate the predictivity and utility of this approach. We determined the binding affinities of a set of 90 nonamer peptides for the MHC class I allele HLA-A*0201 using an in-house, FACS-based, MHC stabilization assay, and from these data we derived an additive quantitative structure-affinity relationship model for peptide interaction with the HLA-A*0201 molecule. Using this model we then designed a series of high affinity HLA-A2-binding peptides. Experimental analysis revealed that all these peptides showed high binding affinities to the HLA-A*0201 molecule, significantly higher than the highest previously recorded. In addition, by the use of systematic substitution at principal anchor positions 2 and 9, we showed that high binding peptides are tolerant to a wide range of nonpreferred amino acids. Our results support a model in which the affinity of peptide binding to MHC is determined by the interactions of amino acids at multiple positions with the MHC molecule and may be enhanced by enthalpic cooperativity between these component interactions.

Amino Acid Sequence↗

Identifiying human MHC supertypes using bioinformatic methods.

Classification of MHC molecules into supertypes in terms of peptide-binding specificities is an important issue, with direct implications for the development of epitope-based vaccines with wide population coverage. In view of extremely high MHC polymorphism (948 class I and 633 class II HLA alleles) the experimental solution of this task is presently impossible. In this study, we describe a bioinformatics strategy for classifying MHC molecules into supertypes using information drawn solely from three-dimensional protein structure. Two chemometric techniques-hierarchical clustering and principal component analysis-were used independently on a set of 783 HLA class I molecules to identify supertypes based on structural similarities and molecular interaction fields calculated for the peptide binding site. Eight supertypes were defined: A2, A3, A24, B7, B27, B44, C1, and C4. The two techniques gave 77% consensus, i.e., 605 HLA class I alleles were classified in the same supertype by both methods. The proposed strategy allowed "supertype fingerprints" to be identified. Thus, the A2 supertype fingerprint is Tyr(9)/Phe(9), Arg(97), and His(114) or Tyr(116); the A3-Tyr(9)/Phe(9)/Ser(9), Ile(97)/Met(97) and Glu(114) or Asp(116); the A24-Ser(9) and Met(97); the B7-Asn(63) and Leu(81); the B27-Glu(63) and Leu(81); for B44-Ala(81); the C1-Ser(77); and the C4-Asn(77).

Alleles↗

Autophagy and related processes in trypanosomatids: insights from genomic and bioinformatic analyses.

The targeting in eukaryotic cells of cellular components to the lysosome or vacuole for degradation is called autophagy. Not only cytoplasmic macromolecules and bulk cytoplasm are subject to this process; entire organelles such as peroxisomes can be degraded. Autophagy of peroxisomes is called pexophagy. Unpublished evidence suggests that the analogous processing of glycosomes in the protozoan kinetoplastids occurs. Taking advantage of the (near-) complete status of three trypanosomatid genomes, a census of components of autophagy and related processes has been undertaken in these organisms. Simple database searches were supplemented by more advanced analyses where necessary. At most, only half of the components characterized in yeasts are present in trypanosomatids suggesting an unexpectedly streamlined version of autophagy occurs in these organisms. The cytoplasm-to-vacuole targeting (Cvt) system for delivery of proteins to the vacuole seems entirely absent in trypanosomatids. The accuracy of the census is supported by the coordinated absence of functionally linked components such as the conjugation system involving ATG12, ATG5, ATG10 and ATG16 that acts at the step of vesicle expansion and completion. Overall, the results are consistent with a scenario of taxon-specific addition of components to a minimal core, a hypothesis that should be readily testable by further genomic surveys allied to laboratory experiments. A bioinformatics analysis of the trypanosomatidal proteins was carried out, highlighting the paucity of information available regarding their structures and enabling prioritization of targets for future structural biology work.

Animals↗

p53 gain-of-function: tumor biology and bioinformatics come together.

p53 is typically viewed as a tumor suppressor. However, many missense somatic and germline mutations in the p53 gene cause gain-of-function whereby p53 acquires novel biochemical activities, such as the ability to transactivate transcription of new genes or to mediate new regulatory protein-protein interactions. Several recent studies show that at least some gain-of-function mutations of p53 are biologically relevant leading to a change in the tumor phenotype. Independent bioinformatic analysis of somatic mutation spectra of the p53 gene yields three lines of evidence supporting the notion that gain-of-function could be the prevalent mode of p53 evolution in tumors. (1) The hotspots in the p53 gene show signs of intensive positive selection. (2) The hotspots are located primarily in functionally important motifs of the DNA-binding domain of p53 which are highly conserved in interspecies evolution. (3) The spectra of hotspots significantly differ among various tumor types and the germline (Li-Fraumeni syndrome); in addition to the hotspots shared by the germline and some of the tumors, many are tumor-specific. The latter observation suggests an unexpected level of complexity of p53 evolution in tumors, with distinct novel function gained in different tumors.

Animals↗

Bioinformatics in the post-genome era.

Recent years saw a dramatic increase in genomic and proteomic data in public archives. Now with the complete genome sequences of human and other species in hand, detailed analyses of the genome sequences will undoubtedly improve our understanding of biological systems and at the same time require sophisticated bioinformatic tools. Here we review what computational challenges are ahead and what are the new exciting developments in this exciting field.

Animals↗

Physiological sub-typing of cold and freezing injury in Triticum turgidum subspecies with bioinformatic and expression characterization of glutathione reductase.

BACKGROUND: This study examined how different subspecies of Triticum turgidum (T. durum, T. polonicum, T. turanicum) respond to cold and freezing, assessing their water status, stress responses, and antioxidant system, with particular focus on the structure and function of glutathione reductase (TtGR). METHODS: TtGR genes were first identified from the T. turgidum genome using publicly available genomic resources such as Ensembl Plants. Promoter regions (~2 kb upstream) were analyzed to identify cis-regulatory elements using PlantCARE. Gene classification was performed based on predicted subcellular localization and conserved domain features. Plants were subjected to cold acclimation and freezing treatments, and physiological, biochemical, and enzymatic parameters were measured. RESULTS: Bioinformatics analyses identified four TtGR genes in the T. turgidum genome. The genes in two groups: cytosolic (Class I) and chloroplastic (Class II). Gene structure analysis showed a conserved exon-intron organization, while motif analysis confirmed the presence of Nicotinamide Adenine Dinucleotide Phosphate (NADPH)-binding and redox-active domains across all TtGR proteins. Several regulatory sequences in the promoters are involved in cold (DRE), abscisic acid (ABRE), and stress (STRE) responses, indicating that TtGR genes are dynamically regulated in response to environmental changes. Physiological analyses showed that freezing treatment reduces leaf water content in all genotypes, leading to turgor loss, hydrogen peroxide (H2O2) accumulation, and increased malondealdehyte (MDA) levels. However, tolerance mechanisms addressing water stress and membrane damage differ among genotypes. At the biochemical level, activation of the antioxidant defense system occurs in all genotypes. T. turanicum displays strong defense by significantly increasing enzyme activities, ensuring that the ascorbate-glutathione cycle continues under stress. By contrast, T. polonicum, although showing increased overall enzyme activities, experiences a dramatic drop in glutathione reductase (GR) activity at freezing temperatures, which restricts reduced glutathione (GSH) regeneration and creates a functional bottleneck in the antioxidant cycle. T. durum fails to sustain enzyme activities over the stress period, leading to an intermediate-sensitive response. Thus, whereas T. turanicum effectively maintains antioxidant function during freezing, T. polonicum and T. durum exhibit less efficient stress responses, either through enzymatic bottlenecks or a lack of sustained defense. CONCLUSIONS: One of the most striking findings of this study is the observed dissociation between TtGR gene expression levels and enzyme activities. Low temperature limits the link between transcription and enzyme function. The primary determinant of low-temperature tolerance in T. turgidum subspecies is the sustainability of GR enzyme activity and GSH regeneration under freezing conditions.

Triticum↗