PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatic software”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

NemaFootPrinter: a web based software for the identification of conserved non-coding genome sequence regions between C. elegans and C. briggsae.

BACKGROUND: NemaFootPrinter (Nematode Transcription Factor Scan Through Philogenetic Footprinting) is a web-based software for interactive identification of conserved, non-exonic DNA segments in the genomes of C. elegans and C. briggsae. It has been implemented according to the following project specifications:a) Automated identification of orthologous gene pairs. b) Interactive selection of the boundaries of the genes to be compared. c) Pairwise sequence comparison with a range of different methods. d) Identification of putative transcription factor binding sites on conserved, non-exonic DNA segments. RESULTS: Starting from a C. elegans or C. briggsae gene name or identifier, the software identifies the putative ortholog (if any), based on information derived from public nematode genome annotation databases. The investigator can then retrieve the genome DNA sequences of the two orthologous genes; visualize graphically the genes' intron/exon structure and the surrounding DNA regions; select, through an interactive graphical user interface, subsequences of the two gene regions. Using a bioinformatics toolbox (Blast2seq, Dotmatcher, Ssearch and connection to the rVista database) the investigator is able at the end of the procedure to identify and analyze significant sequences similarities, detecting the presence of transcription factor binding sites corresponding to the conserved segments. The software automatically masks exons. DISCUSSION: This software is intended as a practical and intuitive tool for the researchers interested in the identification of non-exonic conserved sequence segments between C. elegans and C. briggsae. These sequences may contain regulatory transcriptional elements since they are conserved between two related, but rapidly evolving genomes. This software also highlights the power of genome annotation databases when they are conceived as an open resource and the possibilities offered by seamless integration of different web services via the http protocol. AVAILABILITY: The program is freely available at http://bio.ifom-firc.it/NTFootPrinter.

Animals↗

Bioinformatic analyses and validated experiments reveal an aging hallmark gene set and protective miR of coronary artery disease.

To investigate how aging hallmarks exert roles in the age-related disease of coronary artery disease (CAD). R software and the GEO2R online tool identified differentially expressed genes (DEGs) and differentially expressed microRNAs (DEMis) in CAD microarray datasets from the Gene Expression Omnibus. Genes common to target genes of DEMis, DEGs, and an aging gene list from Human Aging Genomic Resources were then identified and analyzed for protein-protein interactions and functional and pathway enrichment. An miR-mRNA network was constructed using Cytoscape. Receiver operating characteristic curve analysis assessed the diagnostic utility of DEMis in CAD. The expression of two DEMis from a CAD cohort was employed to validate the findings. An aging hallmark gene set, comprising 18 genes, was delineated, with the hub gene TP53 established through protein-protein interaction and microRNA-mRNA networks. Within the microRNA-mRNA network, two DEMis (hsa-miR-423-5p and hsa-miR-564) potentially regulated TP53, rendering them potential CAD biomarkers, as indicated by their area under the curves (AUC) surpassing 0.6. Validation experiments corroborated an AUC of 0.7002 for hsa-miR-423-5p and 0.7261 for hsa-miR-564, highlighting its protective association with CAD. Combining hsa-miR-423-5p, hsa-miR-564, total cholesterol (TC), high-density lipoprotein-cholesterol (HDL-C), low-density lipoprotein-cholesterol (LDL-C), white blood cells (WBC) achieved an area under the receiver operating characteristics curve of 0.783. A CAD-associated gene set was identified, with TP53 as the central hub. Hsa-miR-564 emerged as a potential protective factor against CAD.

Humans↗

Global computing for bioinformatics.

Global computing, the collaboration of idle PCs via the Internet in a SETI@home style, emerges as a new way of massive parallel multiprocessing with potentially enormous CPU power. Its relations to the broader, fast-moving field of Grid computing are discussed without attempting a review of the latter. This review (i) includes a short table of milestones in global computing history, (ii) lists opportunities global computing offers for bioinformatics, (iii) describes the structure of problems well suited for such an approach, (iv) analyses the anatomy of successful projects and (v) points to existing software frameworks. Finally, an evaluation of the various costs shows that global computing indeed has merit, if the problem to be solved is already coded appropriately and a suitable global computing framework can be found. Then, either significant amounts of computing power can be recruited from the general public, or--if employed in an enterprise-wide Intranet for security reasons--idle desktop PCs can substitute for an expensive dedicated cluster.

Computational Biology↗

Bioinformatic analysis of exon repetition, exon scrambling and trans-splicing in humans.

MOTIVATION: Using bioinformatic approaches we aimed to characterize poorly understood abnormalities in splicing known as exon scrambling, exon repetition and trans-splicing. RESULTS: We developed a software package that allows large-scale comparison of all human expressed sequence tags (EST) sequences to the entire set of human gene sequences. Among 5,992,495 EST sequences, 401 cases of exon repetition and 416 cases of exon scrambling were found. The vast majority of identified ESTs contain fragments rather than full-length repeated or scrambled exons. Their structures suggest that the scrambled or repeated exon fragments may have arisen in the process of cDNA cloning and not from splicing abnormalities. Nevertheless, we found 11 cases of full-length exon repetition showing that this phenomenon is real yet very rare. In searching for examples of trans-splicing, we looked only at reproducible events where at least two independent ESTs represent the same putative trans-splicing event. We found 15 ESTs representing five types of putative trans-splicing. However, all 15 cases were derived from human malignant tissues and could have resulted from genomic rearrangements. Our results provide support for a very rare but physiological occurrence of exon repetition, but suggest that apparent exon scrambling and trans-splicing result, respectively, from in vitro artifact and gene-level abnormalities. AVAILABILITY: Exon-Intron Database (EID) is available at http://www.meduohio.edu/bioinfo/eid. Programs are available at http://www.meduohio.edu/bioinfo/software.html. The Laboratory website is available at http://www.meduohio.edu/medicine/fedorov SUPPLEMENTARY INFORMATION: Supplementary file is available at http://www.meduohio.edu/bioinfo/software.html.

Algorithms↗

The PRINTS protein fingerprint database in its fifth year.

PRINTS is a database of protein family 'fingerprints' offering a diagnostic resource for newly-determined sequences. By contrast with PROSITE, which uses single consensus expressions to characterise particular families, PRINTS exploits groups of motifs to build characteristic signatures. These signatures offer improved diagnostic reliability by virtue of the mutual context provided by motif neighbours. To date, 800 fingerprints have been constructed and stored in PRINTS. The current version, 17.0, encodes approximately 4500 motifs, covering a range of globular and membrane proteins, modular polypeptides, and so on. The database is accessible via the UCL Bioinformatics World Wide Web (WWW) Server at http://www. biochem.ucl.ac.uk/bsm/dbbrowser/ . We have recently enhanced the usefulness of PRINTS by making available new, intuitive search software. This allows both individual query sequence and bulk data submission, permitting easy analysis of single sequences or complete genomes. Preliminary results indicate that use of the PRINTS system is able to assign additional functions not found by other methods, and hence offers a useful adjunct to current genome analysis protocols.

Animals↗

[Screening and analysis of coding SNPs of HLA-DQA1 gene involved in susceptibility for cervical cancer].

BACKGROUND & OBJECTIVE: Polymorphisms of human leukocyte antigen (HLA) gene play an important role in the development of cervical cancer. This study was to screen single nucleotide polymorphisms (SNPs) of HLA-DQA1 gene involved in susceptibility of cervical cancer by a bioinformatics approach, and analyze their correlations to abnormal gene functions. METHODS: SNPs of HLA-DQA1 were screened from a public database dbSNP by SNPper software, and relevant FASTA subsequences were also obtained from dbSNP. PARSESNP software was used to analyze cSNPs. RESULTS: Two SNPs, rs9272693 and rs9272703, which may induce mis-sense mutation, were identified in codon region of HLA-DQA1 gene. A PSSM difference>10 was used to predict deleterious mutation. CONCLUSIONS: SNPper software in combination with PARSESNP software could be used to analyze SNPs of HLA-DQA1 gene and select the variants in a conserved region, and it provides an evaluation criterion. But the results need to be verified in cervical cancer patients and control populations.

Databases as Topic↗

Graphical tools for comparative genome analysis.

Visualization of data is important for many data-rich disciplines. In biology, where data sets are becoming larger and more complex, graphical analysis is felt to be ever more pertinent. Although some patterns and trends in data sets may only be determined by sophisticated computational analysis, viewing data by eye can provide us with an extraordinary amount of information in an instant. Recent advances in bioinformatic technologies allow us to link graphical tools to data sources with ease, so we can visualize our data sets dynamically. Here, an overview of graphical software tools for comparative genome analysis is given, showing that a range of simple tools can provide us with a powerful view of the differences and similarities between genomes.

Animals↗

Bioinformatics of cellular signalling.

The completion of the human genome sequencing provides a unique opportunity to understand the complex functioning of cells in terms of myriad biochemical pathways. Of special significance are pathways involved in cellular signalling. Understanding how signal transduction occurs in cells is of paramount importance to medicine and pharmacology. The major steps involved in deciphering signalling pathways are: (a) identifying the molecules involved in signalling; (b) figuring out who talks to whom, i.e. deciphering molecular interactions in a context specific manner; (c) obtaining the spatiotemporal location of the signalling events; (d) reconstructing signalling modules and networks evoked in specific response to input; (e) correlating the signalling response to different cellular inputs; and (f) deciphering cross-talk between signalling modules in response to single and multiple inputs. High-throughput experimental investigations offer the promise of providing data pertaining to the above steps. A major challenge, then, is the organization of this data into knowledge in the form of hypothesis, models and context-specific understanding. The Alliance for Cellular Signaling (AfCS) is a multi-institution, multidisciplinary project and its primary objective is to utilize a multitude of high throughput approaches to obtain context-specific knowledge of cellular response to input. It is anticipated that the AfCS experimental data in combination with curated gene and protein annotations, available from public repositories, will serve as a basis for reconstruction of signalling networks. It will then be possible to model the networks mathematically to obtain quantitative measures of cellular response. In this paper we describe some of the bioinformatics strategies employed in the AfCS.

Animals↗

The Gabriella Miller Kids First Data Resource for genomic research in pediatric cancer and congenital anomalies.

Nine-year-old brain tumor patient Gabriella Miller challenged members of Congress to "stop talking and start doing" when providing federal funding for research into cures for pediatric cancer and congenital anomalies. Though she ultimately lost her life to that cancer, her advocacy efforts resulted in the 2014 Gabriella Miller Kids First Research Act, launching the Gabriella Miller Kids First Pediatric Research Program at the National Institutes of Health (NIH). The overarching goal of the Gabriella Miller Kids First Pediatric Research Program is to help researchers uncover new insights into the biology of childhood cancer and congenital anomalies. Following the signing of the Gabriella Miller Kids First Research Act 2.0 in January 2025, the program has been extended at NIH through 2028 to advance the groundwork laid in the program's first ten years. The Gabriella Miller Kids First Data Resource Center has since honored her legacy by building a comprehensive data resource for genomic research into pediatric conditions. Data from more than 30,000 participants annotated with demographic and clinical information related to their diagnoses have been released for secondary research and analysis using the center's web-based platforms. This paper analyzes the outcomes of the initiative and highlights breakthroughs made by the larger research community resulting from the availability of this data resource. We explore the future expansion of the data resource to include new modalities and tools for supporting life-saving research for children like Gabriella Miller.

Humans↗

An integrated biomedical knowledge extraction and analysis platform: using federated search and document clustering technology.

High content screening (HCS) requires time-consuming and often complex iterative information retrieval and assessment approaches to optimally conduct drug discovery programs and biomedical research. Pre- and post-HCS experimentation both require the retrieval of information from public as well as proprietary literature in addition to structured information assets such as compound libraries and projects databases. Unfortunately, this information is typically scattered across a plethora of proprietary bioinformatics tools and databases and public domain sources. Consequently, single search requests must be presented to each information repository, forcing the results to be manually integrated for a meaningful result set. Furthermore, these bioinformatics tools and data repositories are becoming increasingly complex to use; typically they fail to allow for more natural query interfaces. Vivisimo has developed an enterprise software platform to bridge disparate silos of information. The platform automatically categorizes search results into descriptive folders without the use of taxonomies to drive the categorization. A new approach to information retrieval for HCS experimentation is proposed.

Biomedical Research↗

EcoCyc: a comprehensive database resource for Escherichia coli.

The EcoCyc database (http://EcoCyc.org/) is a comprehensive source of information on the biology of the prototypical model organism Escherichia coli K12. The mission for EcoCyc is to contain both computable descriptions of, and detailed comments describing, all genes, proteins, pathways and molecular interactions in E.coli. Through ongoing manual curation, extensive information such as summary comments, regulatory information, literature citations and evidence types has been extracted from 8862 publications and added to Version 8.5 of the EcoCyc database. The EcoCyc database can be accessed through a World Wide Web interface, while the downloadable Pathway Tools software and data files enable computational exploration of the data and provide enhanced querying capabilities that web interfaces cannot support. For example, EcoCyc contains carefully curated information that can be used as training sets for bioinformatics prediction of entities such as promoters, operons, genetic networks, transcription factor binding sites, metabolic pathways, functionally related genes, protein complexes and protein-ligand interactions.

Computational Biology↗

Recent developments in biological sequence databases.

Biological sequence databases are currently being re-engineered to make them more efficient and easier to use. This re-engineering is also providing an infrastructure to make it easier to interrogate and integrate data from different sources. The net result of this effort should be a great improvement in the power and availability of bioinformatics resources to the general biology community.

Databases, Factual↗

ORBIT: an integrated environment for user-customized bioinformatics tools.

MOTIVATION: There are a large number of computational programs freely available to bioinformaticians via a client/server, web-based environment. However, the client interface to these tools (typically an html form page) cannot be customized from the client side as it is created by the service provider. The form page is usually generic enough to cater for a wide range of users. However, this implies that a user cannot set as 'default' advanced program parameters on the form or even customize the interface to his/her specific requirements or preferences. Currently, there is a lack of end-user interface environments that can be modified by the user when accessing computer programs available on a remote server running on an intranet or over the Internet. RESULTS: We have implemented a client/server system called ORBIT (Online Researcher's Bioinformatics Interface Tools) where individual clients can have interfaces created and customized to command-line-driven, server-side programs. Thus, Internet-based interfaces can be tailored to a user's specific bioinformatic needs. As interfaces are created on the client machine independent of the server, there can be different interfaces to the same server-side program to cater for different parameter settings. The interface customization is relatively quick (between 10 and 60 min) and all client interfaces are integrated into a single modular environment which will run on any computer platform supporting Java. The system has been developed to allow for a number of future enhancements and features. ORBIT represents an important advance in the way researchers gain access to bioinformatics tools on the Internet.

Computational Biology↗

VisPan: real-time visualisation of multiplex amplicon-based sequencing panels for rapid syndromic surveillance and pathogen detection.

MOTIVATION: Infectious diseases persist as a major global public health challenge. Diverse factors, including climate change, globalization, deforestation, human-animal interactions, lifestyle choices, and various biological factors, can contribute to their emergence and reemergence. Rapid detection and characterization of (re)emerging pathogens are therefore critical for effective outbreak management and for enhancing our understanding of epidemics by monitoring the transmission, spread, evolution, and genomics of pathogens. In this context, next-generation sequencing technologies (NGS), particularly long-read platforms such as Oxford Nanopore Technologies (ONT), have opened new avenues for real-time pathogen monitoring. However, the bioinformatics bottleneck remains a challenge, emphasizing the need for efficient, accessible, and user-friendly analysis tools. RESULTS: Here, we present a tool adapted from the RAMPART software that enables real-time data visualisation of multiplex PCR syndromic panels combined with Oxford Nanopore sequencing. This real-time analysis enables rapid pathogen detection, from raw data acquisition to taxonomic assignment, within minutes. The interface offers dynamic visual tracking of the sequencing run and amplicon coverage, facilitating immediate insights during diagnostic workflows. Validation experiments confirmed the system's reliability, accurately identifying all pathogens present in complex clinical or environmental samples. This tool provides an integrated, user-friendly solution for genomic pathogen surveillance in field or clinical settings.

Software↗

Protecting innovation in bioinformatics and in-silico biology.

Commercial success or failure of innovation in bioinformatics and in-silico biology requires the appropriate use of legal tools for protecting and exploiting intellectual property. These tools include patents, copyrights, trademarks, design rights, and limiting information in the form of 'trade secrets'. Potentially patentable components of bioinformatics programmes include lines of code, algorithms, data content, data structure and user interfaces. In both the US and the European Union, copyright protection is granted for software as a literary work, and most other major industrial countries have adopted similar rules. Nonetheless, the grant of software patents remains controversial and is being challenged in some countries. Current debate extends to aspects such as whether patents can claim not only the apparatus and methods but also the data signals and/or products, such as a CD-ROM, on which the programme is stored. The patentability of substances discovered using in-silico methods is a separate debate that is unlikely to be resolved in the near future.

Computational Biology↗

A system architecture for genomic data analysis.

MOTIVATION: Most of diseases are caused by a set of gene defects, which occur in a complex association. The association scheme of expressed genes can be modelled by genetic networks. Genetic networks are efficiently facilities to understand the dynamic of pathogenic processes by modelling molecular reality of cell conditions. In this sense a genetic network consists of first, a set of genes of specified cells, tissues or species and second, causal relations between these genes determining the functional condition of the biological system, i. e. under disease. A relation between two genes will exist if they both are directly or indirectly associated with disease [8]. Our goal is to characterize diseases (especially autoimmune diseases like chronic pancreatitis CP, multiple sclerosis MS, rheumatoid arthritis RA) by genetic networks generated by a computer system. We want to introduce this practice as a bioinformatic approach for finding targets.

Genetic Predisposition to Disease↗

Challenges in integrating biological data sources.

Scientific data of importance to biologists reside in a number of different data sources, such as GenBank, GSDB, SWISS-PROT, EMBL, and OMIM, among many others. Some of these data sources are conventional databases implemented using database management systems (DBMSs) and others are structured files maintained in a number of different formats (e.g., ASN.1 and ACE). In addition, software packages such as sequence analysis packages (e.g., BLAST and FASTA) produce data and can therefore be viewed as data sources. To counter the increasing dispersion and heterogeneity of data, different approaches to integrating these data sources are appearing throughout the bioinformatics community. This paper surveys the technical challenges to integration, classifies the approaches, and critiques the available tools and methodologies.

Chromosomes, Artificial, Yeast↗