PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Support vector machines for novel class detection in Bioinformatics.

Novelty detection techniques might be a promising way of dealing with high-dimensional classification problems in Bioinformatics. We present preliminary results of the use of a one-class support vector machine approach to detect novel classes in two Bioinformatics databases. The results are compatible with theory and inspire further investigation.

Artificial Intelligence↗

[From bioinformatics to systems biology: account of the 12th international conference on intelligent systems in molecular biology].

The paper reviews the 12th International Conference on Intelligent Systems for Molecular Biology/Third European Conference on Computational Biology 2004 that was held in Glasgow, UK, during July 31-August 4. A number of talks, papers and software demos from the conference in bioinformatics, genomics, proteomics, transcriptomics and systems biology are described. Recent applications of liquid chromatography - tandem mass spectrometry, comparative genomics and DNA microarrays are given along with the discussion of bioinformatics curricular in higher education.

Computational Biology↗

Combining bioinformatics resources for the structural modelling of eukaryotic metabolic networks.

The architecture of the cellular metabolic network is almost completely available from several databases. This has paved the way for computational analyses. Whereas kinetic modelling is still restrained to small metabolic sub-systems for which enzyme-kinetic details are known, so-called structural modelling techniques can be applied to complete metabolic networks even if the kinetics and regulation of the underlying enzymes is still unknown. Structural modelling requires detailed information on the presence of metabolic enzymes in a specific cell type of interest and the thermodynamics of the reactions, determining their direction under cellular conditions. If compartments are distinguished the sub-cellular compartmentation of reactions and enzymes and the membrane transporters exchanging metabolites between cellular compartments must be included. All this information cannot be taken from a single data base but has to be compiled from various Bioinformatics resources. Here we present an approach towards the organization of Bioinformatics data that enables the flux-balance analysis of comprehensive compartmentalized metabolic networks of eukaryotic cells with special focus on human hepatocytes.

Aspartate Aminotransferases↗

[The application of bioinformatics in the research of alternative splicing].

Alternative splicing, a fundamental and important regulatory mechanism in eukaryotes, allows one pre-mRNA to be processed into many different mature forms within a cell, each of which can have distinct functions. As alternative splicing is associated with human diseases, the study of alternative splicing becomes quite important. Bioinformatics is a new subject for the study of alternative splicing, especially for its regulatory mechanism, prediction and origin. Of course, bioinformatics must be combined with experimental research so as to clarify these aspects of alternative splicing. This paper reviewed the recent research progress in this field in the hope to gain a deeper understanding of eukaryotic gene expression regulation.

Alternative Splicing↗

[Bioinformatic analysis of adenoma-normal mucosa SSH library of colon].

We established a colonic adenoma-normal mucosa suppressive subtraction hybridization (SSH) library in 1999. In this study, we wanted to explore the expression profile of all candidate genes in this library. We developed an EST pipeline which contained two in-house software packages, nucleic acid analytical software and GetUni. The nucleic acid analytical software, an integrator of the universal bioinformatics tools including phred, phd2fasta, cross_match, repeatmasker and blast2.0, can blast sequences of differential clones with the downloaded non-redundant nucleotide (NR) database. GetUni can cluster these NR sequences into Unigene via matching with the downloaded Homo Sapiens UniGene database. Sixty-two candidate genes in A-N library were obtained via the high throughput automatic gene expression bioinformatics pipeline. Gene Ontology online analysis revealed that ribosome genes and immunity-regulating genes were the two most common categories in the KEGG or Biocarta Pathway. We also detected the expression of 2 genes with highest hits, Reg4 and FAM46A, by semi-quantitative RT-PCR. Both genes were up-regulated in 10 or 9 out of 10 adenomas in comparison with the paired normal mucosa, respectively. The candidate genes in A-N library would be of great significance in disclosing the molecular mechanism underlying in colonic adenoma initiation and progression.

Adenoma↗

[Bioinformatics and GenEnv database in biological risk management].

Identification and molecular typing of environmental isolates by molecular techniques requires knowledge of the genetic characteristics of the microbe species being examined. The introduction of automated sequences has greatly speeded up the entire sequencing process as well as improved the accuracy of the collected information. Bioinformatics tools have become indispensable not only for setting up research studies, but also for storing, organizing and managing enormous quantities of sequencing data. Despite its great advantages, the use of bioinformatics is hindered by difficulties in learning how to use its software tools. The GenEnv database was developed to provide operators involved in biological risk management with a user-friendly tool for sequence analysis. Presently, there are over 20.000 sequence records, and over 9000 bacterial species represented in the database. The initial gene set comprises rDNA16S, rpoB, gyrB. The system allows sequence-driven microbe identification as well as the development of study protocols for research on specific microbe species. Nucleotide sequences are represented graphically. The GenEnv database was designed as a tool for public health operators but also offers wide prospects for scientific research.

Computational Biology↗

Identification through bioinformatics of two new macrophage proinflammatory human chemokines: MIP-3alpha and MIP-3beta.

An increasing number of proinflammatory peptides, known as chemokines, are constantly being described and characterized. Because of their proven biologic functions in allergy, AIDS and, in general, inflammatory processes, these proteins have recently gained more attention. In this study we report the identification through bioinformatics of two new human chemokines: MIP-3alpha and MIP-3beta. Both of them belong to the beta- or CC chemokine family. Expression studies indicate that MIP-3alpha is predominantly expressed in lymph nodes, appendix, PBL, fetal liver, fetal lung and several cell lines. However, MIP-3beta expression is restricted to lymph nodes, thymus and appendix. Interestingly enough, both chemokines manifested a pattern of expression strongly regulated by IL-10. In contrast with other CC chemokines, MIP-3beta maps to chromosome 9. Here we show the importance of bioinformatics to discover new molecules with possible therapeutic effects and regulatory functions.

Base Sequence↗

Bioinformatics of cellular signalling.

The completion of the human genome sequencing provides a unique opportunity to understand the complex functioning of cells in terms of myriad biochemical pathways. Of special significance are pathways involved in cellular signalling. Understanding how signal transduction occurs in cells is of paramount importance to medicine and pharmacology. The major steps involved in deciphering signalling pathways are: (a) identifying the molecules involved in signalling; (b) figuring out who talks to whom, i.e. deciphering molecular interactions in a context specific manner; (c) obtaining the spatiotemporal location of the signalling events; (d) reconstructing signalling modules and networks evoked in specific response to input; (e) correlating the signalling response to different cellular inputs; and (f) deciphering cross-talk between signalling modules in response to single and multiple inputs. High-throughput experimental investigations offer the promise of providing data pertaining to the above steps. A major challenge, then, is the organization of this data into knowledge in the form of hypothesis, models and context-specific understanding. The Alliance for Cellular Signaling (AfCS) is a multi-institution, multidisciplinary project and its primary objective is to utilize a multitude of high throughput approaches to obtain context-specific knowledge of cellular response to input. It is anticipated that the AfCS experimental data in combination with curated gene and protein annotations, available from public repositories, will serve as a basis for reconstruction of signalling networks. It will then be possible to model the networks mathematically to obtain quantitative measures of cellular response. In this paper we describe some of the bioinformatics strategies employed in the AfCS.

Animals↗

Bioinformatics for the Structural Genomics of Poxviruses.

Poxviruses are large, complex viruses, and their host species are widespread across the tree of life. As a result, the bioinformatics analysis of their genomes can be complex. Here we show how a few helpful tools and strategies can be used to inform the analysis, leading to a better understanding of the structural properties of poxvirus genomes and to a more accurate quality control of, or comparison between, assembled sequences.

Poxviridae↗

A Comprehensive Bioinformatics Approach to Analysis of Variants: Variant Calling, Annotation, and Prioritization.

Next-Generation Sequencing (NGS), also known as high-throughput sequencing technologies, has enabled rapid and efficient sequencing of large amounts of DNA and RNA. These technologies have revolutionized the field of genomics, transcriptomics, and proteomics and have been widely used in cancer research, leading to advances in clinical diagnosis and treatment. Improvements in the NGS technologies enabled millions of fragments to be sequenced simultaneously in a time- and cost-effective manner and resulted in large amount of genomic data which require efficient analysis methods. Analysis of the genomic data requires both efficient computer resources and bioinformatics approaches. This chapter details a comprehensive computational approach and analysis steps for genomic data analysis.

Computational Biology↗

pSTRminer: integrated bioinformatic software for genome-wide identification and population-scale evaluation of polymorphic short tandem repeats.

Animal forensic genetics plays a critical role in criminal investigations by providing crucial evidence through domestic animal individualization and wildlife species identification. While human forensic genetics benefits from standardized short tandem repeats (STR) genotyping systems, animal forensic applications encounter significant challenges, including the limited availability of validated STR markers, the prevalence of error-prone dinucleotide STRs (di-STRs), and insufficient integration of population data. To address these challenges, we developed pSTRminer, an integrated bioinformatic tool that automates genome-wide STR mining and polymorphism evaluation. By applying pSTRminer to domestic cattle (Bos taurus), we identified 775,444 STRs de novo from the reference genome and genotyped them using whole-genome sequencing data from 60 Chinese and 111 African cattle to evaluate polymorphism across diverse genetic backgrounds. This led to the development of the cattle STR database (CSDB), comprising loci with a genotyping success rate&#x2009;&#x2265;&#x2009;40% and polymorphism information content (PIC)&#x2009;&#x2265;&#x2009;0.5. Experimental validation of 30 randomly selected tetranucleotide STRs (tetra-STRs) and 33 di-STRs via next-generation sequencing in a local Chinese cattle population (n&#x2009;=&#x2009;145) confirmed marker reliability. Although tetra-STRs had lower average polymorphism levels, they exhibited significantly lower stutter ratios (p&#x2009;<&#x2009;0.05), providing a viable path for identifying discriminative markers with fewer artifacts. Systematic screening revealed that certain tetra-STRs could surpass di-STRs in polymorphism. In conclusion, pSTRminer provides a scalable framework for developing standardized STR panels, facilitating the identification of robust and informative markers in forensic applications.

Bioinformatic software↗

The bioinformatics approach to identifying pathogenic variants for colorectal cancer (CRC).

Colorectal cancer (CRC) is the third most prevalent cancer globally, accounting for 9.6% of newly diagnosed cases and 9.3% of cancer-related deaths. It develops from the uncontrolled proliferation of glandular cells in the colon and rectum and is categorized into three primary types: sporadic, hereditary, and colitis-associated. While genetic susceptibility is a key factor in CRC pathogenesis, identifying high-impact pathogenic variants remains a significant challenge. This study integrates bioinformatics and population genetics approaches to identify CRC-associated single-nucleotide polymorphisms (SNPs) with potential clinical significance. CRC-associated SNPs were extracted from the Genome-Wide Association Studies (GWAS) Catalog, functionally annotated via HaploReg, and validated via Ensembl. In addition, expression quantitative trait locus (eQTL) data from the GTEx database were used to assess the effects of these variants on gene expression across human tissues. Our analysis identified three high-priority SNPs (rs9379084, rs3184504, and rs11557154) associated with the RREB1, ATXN2, SH2B3, and DCAF12 genes, which exhibited marked allele frequency differences among populations. These findings suggest potential biomarkers for CRC risk assessment and highlight the importance of genetic screening across diverse populations.

Bioinformatics↗

Analysis of differentially expressed genes in schizophrenia based on bioinformatics and corresponding mRNA expression levels.

OBJECTIVE: This study aimed to use bioinformatics analysis to identify differentially expressed genes (DEGs) involved in the pathogenesis of schizophrenia and validate their mRNA expression levels through real-time quantitative PCR (qPCR). MATERIAL/METHODS: Datasets from the publicly available Gene Expression Omnibus (GEO) database were analyzed using R software to identify DEGs. Functional enrichment analyses, including Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways, were conducted. A protein-protein interaction (PPI) network was constructed using Cytoscape software to identify key genes with notable expression changes. The expression levels of these key genes were subsequently validated in schizophrenia patients using qPCR to assess potential susceptibility genes. RESULTS: In total, 813 DEGs were identified, with six key genes highlighted through GO analysis and PPI network screening. Among these, HDAC1, UBA52, and FYN demonstrated statistically significant differences in mRNA expression between schizophrenia patients and healthy controls (P&#xa0;<&#xa0;0.05). CONCLUSIONS: This study identified several DEGs potentially linked to the pathogenesis of schizophrenia, suggesting that HDAC1, UBA52, and FYN could serve as candidate susceptibility genes and diagnostic biomarkers. These findings provide new insights and directions for future schizophrenia research.

Humans↗

Integrated bioinformatics analysis reveals cross-talking hub genes and therapeutic agents between sepsis and acute myocardial infarction.

BACKGROUND: Sepsis and acute myocardial infarction (AMI) are two significant diseases that may share overlapping etiological mechanisms. This study aims to systematically identify core genes common to both conditions and to explore their potential as therapeutic targets and drug candidates through an integrative analysis of clinical data and bioinformatics. METHODS: The AMI dataset was obtained from the GEO database, and RNA sequencing data were collected from blood samples of patients with sepsis at our hospital. Common genes were identified using differential expression gene analysis (DEG) and weighted gene co-expression network analysis (WGCNA). Functional enrichment analyses, including Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway analysis, were performed. A protein-protein interaction (PPI) network was constructed, and hub genes were identified using the MCC/Degree algorithm. Diagnostic value was assessed via receiver operating characteristic curve analysis. Immune infiltration patterns, single-cell sequencing data, and molecular docking simulations were employed to evaluate immune relevance and identify potential therapeutic compounds. RESULTS: A total of 417 genes were identified between sepsis and AMI, with enrichment analysis revealing significant involvement in inflammatory responses. Three hub genes-JAK2, MYD88, and TIMP1-were selected for further investigation. ROC curves confirmed their strong diagnostic performance for both diseases. Immune infiltration analysis showed that these core genes were significantly correlated with the infiltration levels of various immune cell types. Molecular docking indicated that quercetin exhibited stable binding affinity with the proteins encoded by these genes. qPCR validation further confirmed the upregulation of these three genes, supporting the anti-inflammatory effects of quercetin as a potential targeted therapy. CONCLUSION: JAK2, MYD88, and TIMP1 were identified as shared core genes in sepsis and AMI. These genes not only serve as potential diagnostic biomarkers but also offer novel targets for developing common therapeutic strategies for both conditions. Furthermore, quercetin emerges as a promising candidate for targeted treatment.

Humans↗

Phage bioinformatics tools: a review of computational approaches for bacteriophage research.

Rising clinical interest in phage therapy and the exponential growth of metagenomic sequence catalogues have driven a rapid expansion of bacteriophage bioinformatics. More than 80 dedicated tools, mostly published since 2020, now span identification, assembly, annotation, taxonomy, lifestyle prediction, defence-system detection, and host prediction. Aimed at experienced practitioners and developers, this review synthesizes the field through the lens of three successive computational paradigms: sequence homology, bounded by database completeness; machine learning, constrained by labelled training data; and foundation models, which now achieve Matthews correlation coefficients above 0.95 in identification tasks and, through structure-informed prediction, raise functional annotation to over half of phage genes. Furthermore, we map the upstream components, namely, gene callers, homology engines, protein language models, and structural search tools, that underpin most downstream pipelines, exposing shared infrastructure and ecosystem-level fragility when dependencies change. To translate this into practice, we propose web-based and command-line reference workflows calibrated to user expertise and sample types. Finally, we set an agenda for the next wave of tool development. Roughly half of phage genes still resist functional annotation despite structural methods; no broadly generalizable strain-level host predictor exists for phage therapy; varying true-positive rates (0%-97%) underscore the absence of standardized community benchmarks analogous to Critical Assessment of Structure Prediction or Critical Assessment of Metagenome Interpretation. As generative genome models begin designing synthetic phages, progress will depend less on producing standalone tools than on rigorous evaluation, interoperable infrastructure, and clinically meaningful prediction targets.

Computational Biology↗

Comparative performance of portable DNA extraction protocols and bioinformatics workflows for rapid detection of gram-negative bacteria and antimicrobial resistance using Oxford Nanopore sequencing.

Oxford Nanopore Technology (ONT) enables rapid, portable pathogen identification and antimicrobial resistance (AMR) detection, but the reliability of downstream genomic analyses is highly dependent on DNA extraction quality, particularly in resource-limited settings. This study comparatively evaluated four portable bacterial DNA extraction protocols derived from three commercial kits to determine their impact on nanopore sequencing performance, bioinformatics workflow completion, and field deployability. Six gram-negative bacterial isolates (Escherichia coli, n = 4; Pseudomonas sp., n = 1; and Salmonella sp., n = 1) were processed using four extraction protocols: SwiftX DNA, SwiftX DNA with proteinase K (ProtK), SwiftX ParaBact, and NucleoSpin Microbial. Twenty-four resulting DNA extracts were sequenced on a single multiplexed MinION R10.4.1 flow cell. Sequencing data were analyzed using validated Galaxy-based generic and species-specific pipelines. Workflow completion was defined as successful progression through quality control, assembly, virulence, plasmid, and AMR detection modules. DNA purity varied substantially by extraction protocol and was strongly associated with successful workflow completion (Kruskal-Wallis, P = 0.0006). Accordingly, NucleoSpin Microbial achieved 100% workflow completion, and SwiftX ParaBact achieved 83%, while both SwiftX DNA-based protocols failed to complete full workflows. Importantly, key AMR genes required to classify isolates as multidrug-resistant were consistently detected using both NucleoSpin Microbial and SwiftX ParaBact extractions. However, NucleoSpin Microbial assemblies showed significantly higher contiguity and enabled a broader, more complete detection of virulence factors, pathogenicity islands, plasmid replicons, and accessory AMR genes, reflecting enhanced genomic resolution.IMPORTANCERapid whole-genome sequencing is increasingly used to detect antimicrobial resistance and guide public health responses, but its reliability depends strongly on how bacterial DNA is extracted. In this study, we have shown that DNA extraction method choice has a major impact on Oxford Nanopore sequencing performance across clinically relevant gram-negative bacteria. While silica column-based extraction maximized genomic completeness and analytical depth, paramagnetic bead-based reverse purification offered superior portability with sufficient resolution for frontline AMR surveillance. These findings highlight a practical trade-off between field deployability and high-resolution genomic characterization in low-resource settings.

DNA extraction↗

Identification of potential biomarkers and mechanisms for keloid disorder based on comprehensive bioinformatics analysis and machine learning algorithms.

BACKGROUND: Keloid disorder (KD) encompasses a spectrum of fibroproliferative dermal conditions, the pathogenesis remains complex and incompletely understood. This study sought to identify biomarkers and potential therapeutic targets for KD through an integrative bioinformatics approach and machine learning analysis of RNA sequencing data. METHODS: RNA sequencing was performed on skin tissue samples from 13 patients with KD and 14 healthy controls. Using weighted gene co-expression network analysis and differential expression analysis revealed differentially expressed key module genes, and the CytoHubba plugin identified candidate genes. Subsequently analyzed using least absolute shrinkage and selection operator (LASSO) and support vector machine recursive feature elimination (SVM-RFE) methods to pinpoint feature genes associated with KD. Following this, biomarkers were determined through expression level validation, enrichment analysis, and immune infiltration analysis. RESULTS: A total of 420 differentially expressed key module genes were identified, and the top 10 genes with DMNC values were selected as candidate genes. Five feature genes were selected through LASSO and SVM-RFE, with NID2, MFAP2, COL8A1, and P4HA3 showing significant expression differences between KD and control samples, along with consistent expression patterns across datasets, identified as potential biomarkers. These four biomarkers were proved to possess high diagnostic potential, and they were found to exhibit significant positive correlations with one another. Functional enrichment analysis indicated that the primary KEGG pathways associated with these biomarkers included "steroid hormone biosynthesis" and "cytokine-cytokine receptor interaction." Moreover, immune infiltration analysis revealed that the four biomarkers were negatively correlated with type 17 T helper cells and positively correlated with 15 immune cell types, including activated B cells and central memory CD4 T cells. CONCLUSION: In conclusion, NID2, MFAP2, COL8A1, and P4HA3 were identified as key biomarkers for KD, offering new avenues for more targeted and effective diagnostic and therapeutic strategies for managing this condition.

Humans↗

Identification of mitochondrial energy metabolism-related candidate genes UQCR10 and NDUFA6 in pediatric tetralogy of fallot: an exploratory bioinformatics study.

BACKGROUND: Tetralogy of Fallot (TOF) is one of the most common cyanotic congenital heart diseases in infants and young children. Its molecular basis remains incompletely understood. This study aimed to identify mitochondrial energy metabolism-related candidate genes associated with pediatric TOF using public heart tissue transcriptomic datasets from the GEO database. METHODS: Datasets GSE146218 and GSE217772 were downloaded and merged, followed by batch-effect correction. Differential expression analysis was performed to identify differentially expressed genes (DEGs). Functional enrichment analysis, weighted gene co-expression network analysis (WGCNA), and protein-protein interaction (PPI) network analysis were used to prioritize candidate genes. The Comparative Toxicogenomics Database (CTD) was used as an exploratory literature-based tool to summarize gene-disease associations. RESULTS: A total of 960 DEGs were identified. Functional enrichment analyses showed that these genes were mainly enriched in mitochondrial energy metabolism-related pathways, including oxidative phosphorylation and the mitochondrial respiratory chain. WGCNA and PPI network analyses further prioritized UQCR10 and NDUFA6 as candidate genes, and both genes showed increased expression in TOF heart tissue samples. CTD analysis suggested literature-based associations between these genes and cardiovascular or developmental disease-related terms. CONCLUSION: This exploratory bioinformatics study identified UQCR10 and NDUFA6 as mitochondrial energy metabolism-related candidate genes upregulated in pediatric TOF heart tissue. These findings suggest that mitochondrial respiratory chain-related transcriptional alterations may be involved in TOF-associated myocardial remodeling or stress responses. Further experimental and clinical validation is required to confirm their biological relevance.

Humans↗