PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Chemistry in bioinformatics.

Chemical information is now seen as critical for most areas of life sciences. But unlike Bioinformatics, where data is openly available and freely re-usable, most chemical information is closed and cannot be re-distributed without permission. This has led to a failure to adopt modern informatics and software techniques and therefore paucity of chemistry in bioinformatics. New technology, however, offers the hope of making chemical data (compounds and properties) free during the authoring process. We argue that the technology is already available; we require a collective agreement to enhance publication protocols.

Access to Information↗

Atlas - a data warehouse for integrative bioinformatics.

BACKGROUND: We present a biological data warehouse called Atlas that locally stores and integrates biological sequences, molecular interactions, homology information, functional annotations of genes, and biological ontologies. The goal of the system is to provide data, as well as a software infrastructure for bioinformatics research and development. DESCRIPTION: The Atlas system is based on relational data models that we developed for each of the source data types. Data stored within these relational models are managed through Structured Query Language (SQL) calls that are implemented in a set of Application Programming Interfaces (APIs). The APIs include three languages: C++, Java, and Perl. The methods in these API libraries are used to construct a set of loader applications, which parse and load the source datasets into the Atlas database, and a set of toolbox applications which facilitate data retrieval. Atlas stores and integrates local instances of GenBank, RefSeq, UniProt, Human Protein Reference Database (HPRD), Biomolecular Interaction Network Database (BIND), Database of Interacting Proteins (DIP), Molecular Interactions Database (MINT), IntAct, NCBI Taxonomy, Gene Ontology (GO), Online Mendelian Inheritance in Man (OMIM), LocusLink, Entrez Gene and HomoloGene. The retrieval APIs and toolbox applications are critical components that offer end-users flexible, easy, integrated access to this data. We present use cases that use Atlas to integrate these sources for genome annotation, inference of molecular interactions across species, and gene-disease associations. CONCLUSION: The Atlas biological data warehouse serves as data infrastructure for bioinformatics research and development. It forms the backbone of the research activities in our laboratory and facilitates the integration of disparate, heterogeneous biological sources of data enabling new scientific inferences. Atlas achieves integration of diverse data sets at two levels. First, Atlas stores data of similar types using common data models, enforcing the relationships between data types. Second, integration is achieved through a combination of APIs, ontology, and tools. The Atlas software is freely available under the GNU General Public License at: http://bioinformatics.ubc.ca/atlas/

Computational Biology↗

Bioinformatics approaches for cross-species liver cancer analysis based on microarray gene expression profiling.

BACKGROUND: The completion of the sequencing of human, mouse and rat genomes and knowledge of cross-species gene homologies enables studies of differential gene expression in animal models. These types of studies have the potential to greatly enhance our understanding of diseases such as liver cancer in humans. Genes co-expressed across multiple species are most likely to have conserved functions. We have used various bioinformatics approaches to examine microarray expression profiles from liver neoplasms that arise in albumin-SV40 transgenic rats to elucidate genes, chromosome aberrations and pathways that might be associated with human liver cancer. RESULTS: In this study, we first identified 2223 differentially expressed genes by comparing gene expression profiles for two control, two adenoma and two carcinoma samples using an F-test. These genes were subsequently mapped to the rat chromosomes using a novel visualization tool, the Chromosome Plot. Using the same plot, we further mapped the significant genes to orthologous chromosomal locations in human and mouse. Many genes expressed in rat 1q that are amplified in rat liver cancer map to the human chromosomes 10, 11 and 19 and to the mouse chromosomes 7, 17 and 19, which have been implicated in studies of human and mouse liver cancer. Using Comparative Genomics Microarray Analysis (CGMA), we identified regions of potential aberrations in human. Lastly, a pathway analysis was conducted to predict altered human pathways based on statistical analysis and extrapolation from the rat data. All of the identified pathways have been known to be important in the etiology of human liver cancer, including cell cycle control, cell growth and differentiation, apoptosis, transcriptional regulation, and protein metabolism. CONCLUSION: The study demonstrates that the hepatic gene expression profiles from the albumin-SV40 transgenic rat model revealed genes, pathways and chromosome alterations consistent with experimental and clinical research in human liver cancer. The bioinformatics tools presented in this paper are essential for cross species extrapolation and mapping of microarray data, its analysis and interpretation.

Animals↗

A microarray study of MPP+-treated PC12 Cells: Mechanisms of toxicity (MOT) analysis using bioinformatics tools.

BACKGROUND: This paper describes a microarray study including data quality control, data analysis and the analysis of the mechanism of toxicity (MOT) induced by 1-methyl-4-phenylpyridinium (MPP+) in a rat adrenal pheochromocytoma cell line (PC12 cells) using bioinformatics tools. MPP+ depletes dopamine content and elicits cell death in PC12 cells. However, the mechanism of MPP+-induced neurotoxicity is still unclear. RESULTS: In this study, Agilent rat oligo 22K microarrays were used to examine alterations in gene expression of PC12 cells after 500 muM MPP+ treatment. Relative gene expression of control and treated cells represented by spot intensities on the array chips was analyzed using bioinformatics tools. Raw data from each array were input into the NCTR ArrayTrack database, and normalized using a Lowess normalization method. Data quality was monitored in ArrayTrack. The means of the averaged log ratio of the paired samples were used to identify the fold changes of gene expression in PC12 cells after MPP+ treatment. Our data showed that 106 genes and ESTs (Expressed Sequence Tags) were changed 2-fold and above with MPP+ treatment; among these, 75 genes had gene symbols and 59 genes had known functions according to the Agilent gene Refguide and ArrayTrack-linked gene library. The mechanism of MPP+-induced toxicity in PC12 cells was analyzed based on their genes functions, biological process, pathways and previous published literatures. CONCLUSION: Multiple pathways were suggested to be involved in the mechanism of MPP+-induced toxicity, including oxidative stress, DNA and protein damage, cell cycling arrest, and apoptosis.

1-Methyl-4-phenylpyridinium↗

2DDB - a bioinformatics solution for analysis of quantitative proteomics data.

BACKGROUND: We present 2DDB, a bioinformatics solution for storage, integration and analysis of quantitative proteomics data. As the data complexity and the rate with which it is produced increases in the proteomics field, the need for flexible analysis software increases. RESULTS: 2DDB is based on a core data model describing fundamentals such as experiment description and identified proteins. The extended data models are built on top of the core data model to capture more specific aspects of the data. A number of public databases and bioinformatical tools have been integrated giving the user access to large amounts of relevant data. A statistical and graphical package, R, is used for statistical and graphical analysis. The current implementation handles quantitative data from 2D gel electrophoresis and multidimensional liquid chromatography/mass spectrometry experiments. CONCLUSION: The software has successfully been employed in a number of projects ranging from quantitative liquid-chromatography-mass spectrometry based analysis of transforming growth factor-beta stimulated fi-broblasts to 2D gel electrophoresis/mass spectrometry analysis of biopsies from human cervix. The software is available for download at SourceForge.

Computational Biology↗

BioAfrica's HIV-1 proteomics resource: combining protein data with bioinformatics tools.

Most Internet online resources for investigating HIV biology contain either bioinformatics tools, protein information or sequence data. The objective of this study was to develop a comprehensive online proteomics resource that integrates bioinformatics with the latest information on HIV-1 protein structure, gene expression, post-transcriptional/post-translational modification, functional activity, and protein-macromolecule interactions. The BioAfrica HIV-1 Proteomics Resource http://bioafrica.mrc.ac.za/proteomics/index.html is a website that contains detailed information about the HIV-1 proteome and protease cleavage sites, as well as data-mining tools that can be used to manipulate and query protein sequence data, a BLAST tool for initiating structural analyses of HIV-1 proteins, and a proteomics tools directory. The Proteome section contains extensive data on each of 19 HIV-1 proteins, including their functional properties, a sample analysis of HIV-1HXB2, structural models and links to other online resources. The HIV-1 Protease Cleavage Sites section provides information on the position, subtype variation and genetic evolution of Gag, Gag-Pol and Nef cleavage sites. The HIV-1 Protein Data-mining Tool includes a set of 27 group M (subtypes A through K) reference sequences that can be used to assess the influence of genetic variation on immunological and functional domains of the protein. The BLAST Structure Tool identifies proteins with similar, experimentally determined topologies, and the Tools Directory provides a categorized list of websites and relevant software programs. This combined database and software repository is designed to facilitate the capture, retrieval and analysis of HIV-1 protein data, and to convert it into clinically useful information relating to the pathogenesis, transmission and therapeutic response of different HIV-1 variants. The HIV-1 Proteomics Resource is readily accessible through the BioAfrica website at: http://bioafrica.mrc.ac.za/proteomics/index.html.

Africa↗

Proteomic and bioinformatic analysis of epithelial tight junction reveals an unexpected cluster of synaptic molecules.

BACKGROUND: Zonula occludens, also known as the tight junction, is a specialized cell-cell interaction characterized by membrane "kisses" between epithelial cells. A cytoplasmic plaque of approximately 100 nm corresponding to a meshwork of densely packed proteins underlies the tight junction membrane domain. Due to its enormous size and difficulties in obtaining a biochemically pure fraction, the molecular composition of the tight junction remains largely unknown. RESULTS: A novel biochemical purification protocol has been developed to isolate tight junction protein complexes from cultured human epithelial cells. After identification of proteins by mass spectroscopy and fingerprint analysis, candidate proteins are scored and assessed individually. A simple algorithm has been devised to incorporate transmembrane domains and protein modification sites for scoring membrane proteins. Using this new scoring system, a total of 912 proteins have been identified. These 912 hits are analyzed using a bioinformatics approach to bin the hits in 4 categories: configuration, molecular function, cellular function, and specialized process. Prominent clusters of proteins related to the cytoskeleton, cell adhesion, and vesicular traffic have been identified. Weaker clusters of proteins associated with cell growth, cell migration, translation, and transcription are also found. However, the strongest clusters belong to synaptic proteins and signaling molecules. Localization studies of key components of synaptic transmission have confirmed the presence of both presynaptic and postsynaptic proteins at the tight junction domain. To correlate proteomics data with structure, the tight junction has been examined using electron microscopy. This has revealed many novel structures including end-on cytoskeletal attachments, vesicles fusing/budding at the tight junction membrane domain, secreted substances encased between the tight junction kisses, endocytosis of tight junction double membranes, satellite Golgi apparatus and associated vesicular structures. A working model of the tight junction consisting of multiple functions and sub-domains has been generated using the proteomics and structural data. CONCLUSION: This study provides an unbiased proteomics and bioinformatics approach to elucidate novel functions of the tight junction. The approach has revealed an unexpected cluster associating with synaptic function. This surprising finding suggests that the tight junction may be a novel epithelial synapse for cell-cell communication. REVIEWERS: This article was reviewed by Gáspár Jékely, Etienne Joly and Neil Smalheiser.

Journal Article↗

Bioconductor: open software development for computational biology and bioinformatics.

The Bioconductor project is an initiative for the collaborative creation of extensible software for computational biology and bioinformatics. The goals of the project include: fostering collaborative development and widespread use of innovative software, reducing barriers to entry into interdisciplinary scientific research, and promoting the achievement of remote reproducibility of research results. We describe details of our aims and methods, identify current challenges, compare Bioconductor to other open bioinformatics projects, and provide working examples.

Computational Biology↗

Triple primary synchronous liver cancer in one patient: the first case report and origin speculation through bioinformatics.

INTRODUCTION: A diagnosis of multiple primary liver tumors is extremely rare. Preoperative diagnosis based on imaging findings is difficult. Moreover, the clinical benefits of treatment strategies for multiple liver cancers remain unclear. Here, we report a case of three synchronous primary liver tumors with three distinct pathological types-hepatocellular carcinoma (HCC), intrahepatic cholangiocarcinoma (ICC), and combined hepatocellular-cholangiocarcinoma (cHCC‑CCA)-in a single patient. Bioinformatics analysis supported at least two clonal origins, with cHCC‑CCA and ICC sharing a common lineage based on identical HBV integration sites. CASE PRESENTATION: A 63-year-old female with a history of hepatitis B for several years presented with three lesions in hepatic segment VIII. Multiphase magnetic resonance imaging with gadolinium ethoxybenzyl diethylenetriaminepentaacetic acid revealed a diagnosis of multiple lesions, namely, cHCC‑CCA, with multiple intrahepatic metastases. The AFP level was normal, while the CA 19 - 9 level was mildly elevated (normal range ≤ 30.00 U/ml). Hepatectomy was performed, and postoperative assessment confirmed that the large lesion was cHCC‑CCA. However, the small lesions close to the large lesion were HCC and ICC. Gene testing revealed distinct mutational profiles among the three tumors. Similar gene mutations were detected in cHCC‑CCA and ICC. We also found that gene fragments of hepatitis B virus-C (HBV-C) were inserted into the genomes of ICC and cHCC‑CCA rather than that of HCC. The genomic integration site of HBV-C in cHCC‑CCA and ICC was the same. CONCLUSION: We report an extremely rare case of three synchronous primary liver tumors with three distinct pathological types (HCC, ICC, and cHCC‑CCA) in a single patient. Bioinformatics analysis supported at least two clonal origins, with cHCC‑CCA and ICC sharing a common lineage based on identical HBV integration sites. Hepatectomy represents a potential radical strategy for the treatment of multiple PLCs.

Humans↗

Proceedings: the Applications of Bioinformatics in Cancer Detection Workshop.

The Division of Cancer Prevention of the National Cancer Institute sponsored and organized the Applications of Bioinformatics in Cancer Detection Workshop on August 6-7, 2002. The goal of the workshop was to evaluate the state of the science of bioinformatics and determine how it may be used to assist early cancer detection, risk identification, risk assessment, and risk reduction. This paper summarizes the proceedings of this conference and points out future directions for research.

Computational Biology↗

[Interface between bioinformatics and docking study].

We describe the prospects of bioinformatics for drug discovery and discuss the current status, problems, and future direction of the interface between bioinformatics and docking studies. We also describe our recent work on sequence and structure analysis using the guanidino-modifying enzymes superfamily as a good example.

Binding Sites↗

[GeneChip system from a bioinformatical point of view].

GeneChip (Affymetrix, Inc., USA) employs a specific method for spotting DNA probes on chips, which is different from any other DNA chips, and can complete the whole process from sample preparation to data construction and analysis. The GeneChip system can be applied to both gene expression analysis and genomic mutation analysis, which would play an important role in human genome analysis in the future. Techniques for data construction ("wet" experimental techniques), which are the major components in the GeneChip system, are generally established as routine work in the first screening process in most laboratories worldwide. The most important point would be how we exchange experimental data produced by researchers and gene/genome information available both on the public and the commercial bases so that we reduce useful information on gene expression. Recently, the center of the research has been shifting to computing technology for data processing ("bioinformatics"). This article separately deals with gene expression analysis and genomic analysis, with emphasis on bioinformatics. We describe the data on gene expression screening, the gene targeting process, the analysis of genomic DNA mutations using the P53 probe array, and the HuSNP mapping assays, by presenting our experimental examples.

Computational Biology↗

Application of bioinformatics for DNA microarray data to bioscience, bioengineering and medical fields.

In the 1990s, DNA microarray or DNA chip as a novel biological experimental technology was developed, which enables the comprehensive measurement of the expression levels of hundreds of genes, simultaneously. Using this technique, a comprehensive understanding of the cell can be achieved. However, because even simple life forms, such as microorganisms, have more than a thousand kinds of genes, the data from a DNA chip cannot be analyzed without statistical and informational technology. Bioinformatics is the interdisciplinary research field integrating molecular biology with informatics, and it is expected to have a huge impact on the bioscientific, bioengineering and medical fields. There are many techniques in bioinformatics for the analysis of DNA microarray data; however, these are mainly divided into fold-change analysis, clustering, classification, genetic network analysis, and simulation. In this review, these techniques are briefly explained by using some examples.

Animals↗

Bioinformatics for whole-genome shotgun sequencing of microbial communities.

The application of whole-genome shotgun sequencing to microbial communities represents a major development in metagenomics, the study of uncultured microbes via the tools of modern genomic analysis. In the past year, whole-genome shotgun sequencing projects of prokaryotic communities from an acid mine biofilm, the Sargasso Sea, Minnesota farm soil, three deep-sea whale falls, and deep-sea sediments have been reported, adding to previously published work on viral communities from marine and fecal samples. The interpretation of this new kind of data poses a wide variety of exciting and difficult bioinformatics problems. The aim of this review is to introduce the bioinformatics community to this emerging field by surveying existing techniques and promising new approaches for several of the most interesting of these computational problems.

Journal Article↗

'PePApipe': A complete bioinformatics analysis pipeline for African Swine Fever Virus genome.

African Swine Fever Virus (ASFV) is of high concern in porcine livestock across the world due to both the high mortality rates and the trade restrictions imposed on affected regions. The viral genome is large and complex, and genomic analysis is essential for tracing its origin and evolution. Although several bioinformatics tools exist for genome assembly and analysis, no single platform integrates all necessary steps in an accessible and systematic way. In this study the authors developed 'PePApipe', a custom-built, user-friendly pipeline that enables rapid, complete, and efficient ASFV genome analysis. It is specifically designed for laboratory professionals with limited bioinformatics experience, requiring only basic command-line knowledge. Starting from raw sequencing data, PePApipe integrates thirteen software tools into one automated workflow, covering quality control and pre-processing of raw reads, de novo genome assembly and variant calling. Programmed in Python, it can be executed locally through bash scripts, or using a Slurm protocol for batch processing of multiple samples. The main outputs are the ASFV consensus genome sequence and a file listing its putative variants compared to the selected reference genome. PePApipe classifies generated files into structured folders and produces intermediate files that can be used as inputs for further or parallel analyses; users can also enable or disable specific steps in each particular case. This pipeline is adaptable and complementary to downstream steps such as viral genome annotation or genome visualization. By consolidating all stages of viral genome analysis into a single automated workflow, PePApipe reduces the likelihood of user error, and enhances reproducibility and efficiency. This user-friendly pipeline facilitates the transition from sequencing to assembly and downstream analysis of viral genomes, ensuring a fast and reliable response to molecular analysis demands. Finally, the pipeline can be easily adapted to the study of other viral species, expanding its application in infectious diseases surveillance.

African Swine Fever Virus↗

Correlation assessment of SARS-CoV-2 variants and their subvariants present in clinical and wastewater samples in Oregon, USA (February 7, 2021 - February 26, 2022) using the Freyja bioinformatics approach.

BACKGROUND: Wastewater surveillance is a valuable tool for monitoring SARS-CoV-2 at the community level. As the virus diversified into many variants and subvariants that share overlapping mutations, resolving them accurately from wastewater becomes a key bioinformatic challenge. OBJECTIVES AND AIMS: This study evaluated two distinct bioinformatic approaches, multilocus sequence typing (MLST) and Freyja, for identifying SARS-CoV-2 variants and subvariants in Oregon wastewater samples collected from February 2021 to February 2022. METHODS: The MLST approach identified SARS-CoV-2 variants using unique mutations curated from clinical samples. In contrast, the Freyja approach resolved variant and subvariant abundances using genome wide mutation profiles weighted by sequencing depth. In this study, the variant and subvariants relative abundances produced by both approaches were compared against those observed in clinical surveillance data. RESULTS: Both approaches identified SARS-CoV-2 variants at relative abundances that agreed closely with those observed in clinical surveillance data. However, only the Freyja approach identified over 200 Delta subvariants, divided into three clades (21A, 21I and 21J) and two levels (Level 1 and 2) based on Pango subvariants. Delta subvariants showed strong agreement at Level 1 subvariants (rs = 0.892-0.944), while agreement at Level 2 subvariants was inconsistent (rs = 0.324-0.903). CONCLUSIONS: The Freyja approach provided enhanced resolution of SARS-CoV-2 variants and subvariants in wastewater, at abundances that agreed with clinical surveillance. This added resolution is a critical advantage for public health surveillance as SARS-CoV-2 continues to evolve and share mutations across variants and subvariants.

Oregon↗

Leveraging bioinformatics approaches for drug repositioning in space radiation protection.

The health effects of space radiation, primarily Galactic Cosmic Rays (GCRs), on humans remain largely unknown, with potential cardiovascular consequences posing a significant threat to astronauts on long-duration spaceflight missions. Currently, there are no established pharmacological countermeasures for GCR exposure. Drug repositioning offers a promising strategy to accelerate pharmaceutical research in space medicine. This study leverages existing bioinformatics techniques to identify and prioritize potential drug candidates associated with proteomic perturbations following simulated GCR exposure using previously published murine cardiac proteomic data. A protein-protein interaction (PPI) network was constructed using the top differentially expressed proteins (DEPs) from murine heart tissue following exposure to 5-ion GCRs as seed nodes, focusing on experimentally supported interactions. Network topology, Markov clustering, and functional enrichment analyses were used to characterize biologically relevant proteins and pathways. Drug-protein interactions were predicted using Drugst.One and mapped to PPI clusters of interest to identify candidate drugs. Selected drug-macromolecule interactions were further explored using CB-Dock2 molecular docking and short-duration molecular dynamics simulations as hypothesis-generating structural assessments. Analysis of a key PPI network cluster consisting of several ATP synthase proteins identified 23 unique drug candidates. These analyses demonstrate a systematic approach for leveraging bioinformatics techniques to identify candidate molecular targets and generate pharmacological hypotheses in the context of space radiation countermeasures. Ultimately, this strategy introduces a hypothesis-generating framework for the prioritization of potential drug candidates for future computational characterization and experimental investigation against spaceflight stressors.

Animals↗