PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatic software”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Evolving from bioinformatics in-the-small to bioinformatics in-the-large.

We argue the significance of a fundamental shift in bioinformatics, from in-the-small to in-the-large. Adopting a large-scale perspective is a way to manage the problems endemic to the world of the small-constellations of incompatible tools for which the effort required to assemble an integrated system exceeds the perceived benefit of the integration. Where bioinformatics in-the-small is about data and tools, bioinformatics in-the-large is about metadata and dependencies. Dependencies represent the complexities of large-scale integration, including the requirements and assumptions governing the composition of tools. The popular make utility is a very effective system for defining and maintaining simple dependencies, and it offers a number of insights about the essence of bioinformatics in-the-large. Keeping an in-the-large perspective has been very useful to us in large bioinformatics projects. We give two fairly different examples, and extract lessons from them showing how it has helped. These examples both suggest the benefit of explicitly defining and managing knowledge flows and knowledge maps (which represent metadata regarding types, flows, and dependencies), and also suggest approaches for developing bioinformatics database systems. Generally, we argue that large-scale engineering principles can be successfully adapted from disciplines such as software engineering and data management, and that having an in-the-large perspective will be a key advantage in the next phase of bioinformatics development.

Computational Biology↗

IL-1beta, but not BMP-7 leads to a dramatic change in the gene expression pattern of human adult articular chondrocytes--portraying the gene expression pattern in two donors.

Anabolic and catabolic cytokines and growth factors such as BMP-7 and IL-1beta play a central role in controlling the balance between degradation and repair of normal and (osteo)arthritic articular cartilage matrix. In this report, we investigated the response of articular chondrocytes to these factors IL-1beta and BMP-7 in terms of changes in gene expression levels. Large scale analysis was performed on primary human adult articular chondrocytes isolated from two human, independent donors cultured in alginate beads (non-stimulated and stimulated with IL-1beta and BMP-7 for 48 h) using Affymetrix gene chips (oligo-arrays). Biostatistical and bioinformatic evaluation of gene expression pattern was performed using the Resolver software (Rosetta). Part of the results were confirmed using real-time PCR. IL-1beta modulated significantly 909 out of 3459 genes detectable, whereas BMP-7 influenced only 36 out of 3440. BMP-7 induced mainly anabolic activation of chondrocytes including classical target genes such as collagen type II and aggrecan, while IL-1beta, both, significantly modulated the gene expression levels of numerous genes; namely, IL-1beta down-regulated the expression of anabolic genes and induced catabolic genes and mediators. Our data indicate that BMP-7 has only a limited effect on differentiated cells, whereas IL-1beta causes a dramatic change in gene expression pattern, i.e. induced or repressed much more genes. This presumably reflects the fact that BMP-7 signaling is effected via one pathway only (i.e. Smad-pathway) whereas IL-1beta is able to signal via a broad variety of intracellular signaling cascades involving the JNK, p38, NFkB and Erk pathways and even influencing BMP signaling.

Aged↗

A strategy for assembling the maize (Zea mays L.) genome.

UNLABELLED: Because the bulk of the maize (Zea mays L.) genome consists of repetitive sequences, sequencing efforts are being targeted to its 'gene-rich' fraction. Traditional assembly programs are inadequate for this approach because they are optimized for a uniform sampling of the genome and inherently lack the ability to differentiate highly similar paralogs. RESULTS: We report the development of bioinformatics tools for the accurate assembly of the maize genome. This software, which is based on innovative parallel algorithms to ensure scalability, assembled 730,974 genomic survey sequences fragments in 4 h using 64 Pentium III 1.26 GHz processors of a commodity cluster. Algorithmic innovations are used to reduce the number of pairwise alignments significantly without sacrificing quality. Clone pair information was used to estimate the error rate for improved differentiation of polymorphisms versus sequencing errors. The assembly was also used to evaluate the effectiveness of various filtering strategies and thereby provide information that can be used to focus subsequent sequencing efforts.

Algorithms↗

Mining sequence annotation databanks for association patterns.

MOTIVATION: Millions of protein sequences currently being deposited to sequence databanks will never be annotated manually. Similarity-based annotation generated by automatic software pipelines unavoidably contains spurious assignments due to the imperfection of bioinformatics methods. Examples of such annotation errors include over- and underpredictions caused by the use of fixed recognition thresholds and incorrect annotations caused by transitivity based information transfer to unrelated proteins or transfer of errors already accumulated in databases. One of the most difficult and timely challenges in bioinformatics is the development of intelligent systems aimed at improving the quality of automatically generated annotation. A possible approach to this problem is to detect anomalies in annotation items based on association rule mining. RESULTS: We present the first large-scale analysis of association rules derived from two large protein annotation databases-Swiss-Prot and PEDANT-and reveal novel, previously unknown tendencies of rule strength distributions. Most of the rules are either very strong or very weak, with rules in the medium strength range being relatively infrequent. Based on dynamics of error correction in subsequent Swiss-Prot releases and on our own manual analysis we demonstrate that exceptions from strong rules are, indeed, significantly enriched in annotation errors and can be used to automatically flag them. We identify different strength dependencies of rules derived from different fields in Swiss-Prot. A compositional breakdown of association rules generated from PEDANT in terms of their constituent items indicates that most of the errors that can be corrected are related to gene functional roles. Swiss-Prot errors are usually caused by under-annotation owing to its conservative approach, whereas automatically generated PEDANT annotation suffers from over-annotation. AVAILABILITY: All data generated in this study are available for download and browsing at http://pedant.gsf.de/ARIA/index.htm.

Conserved Sequence↗

[Bioinformatics and GenEnv database in biological risk management].

Identification and molecular typing of environmental isolates by molecular techniques requires knowledge of the genetic characteristics of the microbe species being examined. The introduction of automated sequences has greatly speeded up the entire sequencing process as well as improved the accuracy of the collected information. Bioinformatics tools have become indispensable not only for setting up research studies, but also for storing, organizing and managing enormous quantities of sequencing data. Despite its great advantages, the use of bioinformatics is hindered by difficulties in learning how to use its software tools. The GenEnv database was developed to provide operators involved in biological risk management with a user-friendly tool for sequence analysis. Presently, there are over 20.000 sequence records, and over 9000 bacterial species represented in the database. The initial gene set comprises rDNA16S, rpoB, gyrB. The system allows sequence-driven microbe identification as well as the development of study protocols for research on specific microbe species. Nucleotide sequences are represented graphically. The GenEnv database was designed as a tool for public health operators but also offers wide prospects for scientific research.

Computational Biology↗

[Cloning and bioinformatics of human REV3 gene promoter region and its response to carcinogen N-methyl-N'-nitro-N-nitrosoguanidine].

OBJECTIVE: To understand the up regulatory mechanism of human REV3 gene induced by the chemical carcinogen N-methyl-N'-nitro-N-nitrosoguanidine (MNNG). METHODS: Bioinformatic analysis of human REV3 gene promoter region was based on BLAST alignment, promoter prediction software and recognition of transcriptional factor binding sites. Cloning of human REV3 gene promoter region was performed by nested PCR. Response of human REV3 gene promoter to the chemical carcinogen MNNG was measured by transient transfection assay based on the dual luciferase reporter assay system. RESULT: Bioinformatic analysis showed that human REV3 gene promoter region was located on chromosome 6 PAC clone RP3-415N12, and that the hypothetical promoter region contained promoter sequences, rich CpG islands, and putative recognition sites for several transcriptional factors, including AP-1/c-Jun/c-Fos, AP-2, STAT, CREBP, and NF-kappaB. Reconstructed reporter plasmid pGL3- 2582 was established by inserting 2582 nucleotides from the promoter region into the luciferase reporter vector pGL3-Basic. Transient transfection assay showed the hypothetical REV3 promoter region had promoter function, and it responded to MNNG treatment (P<0.01). CONCLUSION: Human mutator REV3 gene promoter region has been successfully cloned. The response of REV3 promoter region to MNNG suggests that REV3 gene can be regulated at transcriptional level under conditions of genotoxic stress.

Base Sequence↗

[Application of bioinformatics in researches of industrial biocatalysis].

Industrial biocatalysis is currently attracting much attention to rebuild or substitute traditional producing process of chemicals and drugs. One of key focuses in industrial biocatalysis is biocatalyst, which is usually one kind of microbial enzyme. In the recent, new technologies of bioinformatics have played and will continue to play more and more significant roles in researches of industrial biocatalysis in response to the waves of genomic revolution. One of the key applications of bioinformatics in biocatalysis is the discovery and identification of the new biocatalyst through advanced DNA and protein sequence search, comparison and analyses in Internet database using different algorithm and software. The unknown genes of microbial enzymes can also be simply harvested by primer design on the basis of bioinformatics analyses. The other key applications of bioinformatics in biocatalysis are the modification and improvement of existing industrial biocatalyst. In this aspect, bioinformatics is of great importance in both rational design and directed evolution of microbial enzymes. Based on the successful prediction of tertiary structures of enzymes using the tool of bioinformatics, the undermentioned experiments, i.e. site-directed mutagenesis, fusion protein construction, DNA family shuffling and saturation mutagenesis, etc, are usually of very high efficiency. On all accounts, bioinformatics will be an essential tool for either biologist or biological engineer in the future researches of industrial biocatalysis, due to its significant function in guiding and quickening the step of discovery and/or improvement of novel biocatalysts.

Biocatalysis↗

VIJB: a companion of the JBROWSE genome browser for the visually impaired people.

MOTIVATION: The availability of touch-sensitive and haptic devices has been a keystone development for the inclusion of visually impaired people (VIPs) in modern, highly digitized work environments. Braille displays have proven efficient and versatile enough to parse large and complex text files, making bioinformatics and text-heavy programming accessible to VIPs. However, the complex graphical objects -combining numerous datasets- typically generated during data integration remain challenging, even with the aid of descriptive AI. This is particularly true in functional genomics. Here, we present VIJB, a simple application that displays the multilayered output of the JBROWSE genome browser on a Braille reader, enabling VIPs to fully participate in data integration in functional genomics. AVAILABILITY AND IMPLEMENTATION: VIJB is programmed in Python and relies on the scientific library NumPy, the braillegraph and pyBigWig libraries, and the TABIX software. The architecture is summarized in Supplementary Material 1, available as supplementary data at Bioinformatics online. VIJB is available for download at the GitHub repository https://GitHub.com/NiBuMNHN/VIJB and is licenced under the GPL 3.0.

Persons with Visual Disabilities↗

A Comprehensive Bioinformatics Approach to Analysis of Variants: Variant Calling, Annotation, and Prioritization.

Next-Generation Sequencing (NGS), also known as high-throughput sequencing technologies, has enabled rapid and efficient sequencing of large amounts of DNA and RNA. These technologies have revolutionized the field of genomics, transcriptomics, and proteomics and have been widely used in cancer research, leading to advances in clinical diagnosis and treatment. Improvements in the NGS technologies enabled millions of fragments to be sequenced simultaneously in a time- and cost-effective manner and resulted in large amount of genomic data which require efficient analysis methods. Analysis of the genomic data requires both efficient computer resources and bioinformatics approaches. This chapter details a comprehensive computational approach and analysis steps for genomic data analysis.

Computational Biology↗

Bioinformatics: harvesting information for plant and crop science.

Bioinformatics is an integral aspect of plant and crop science research. Developments in data management and analytical software are reviewed with an emphasis on applications in functional genomics. This includes information resources for Arabidopsis and crop species, and tools available for analysis and visualisation of comparative genomic data. Approaches used to explore relationships between plant genes and expressed sequences are compared, including use of ontologies. The impact of bioinformatics in forward and reverse genetics is described, together with the potential from data mining. The role of bioinformatics is explored in the wider context of plant and crop science.

Algorithms↗

Computer applications in biomolecular sciences. Part 2: bioinformatics and genome projects.

This article defines and describes some of the basics of bioinformatics and projects aimed at sequencing entire genomes. Emphasis is placed on some of the ways in which the primary structures of nucleic acids and proteins may be investigated and analysed to gain meaningful biological information using computers and appropriate software. The importance of the world wide net and access to it is given prominence, particularly in bioinformatics research and teaching.

Journal Article↗

Chemical cross-linking and mass spectrometry for mapping three-dimensional structures of proteins and protein complexes.

Chemical cross-linking of proteins, an established method in protein chemistry, has gained renewed interest in combination with mass spectrometric analysis of the reaction products for elucidating low-resolution three-dimensional protein structures and interacting sequences in protein complexes. The identification of the large number of cross-linking sites from the complex mixtures generated by chemical cross-linking, however, remains a challenging task. This review describes the most popular cross-linking reagents for protein structure analysis and gives an overview of the strategies employing intra- or intermolecular chemical cross-linking and mass spectrometry. The various approaches described in the literature to facilitate detection of cross-linking products and also computer software for data analysis are reviewed. Cross-linking techniques combined with mass spectrometry and bioinformatic methods have the potential to provide the basis for an efficient structural characterization of proteins and protein complexes.

Binding Sites↗

X-HUSAR, an X-based graphical interface for the analysis of genomic sequences.

Management and analysis of nucleotide and protein sequence and structure data constitute a traditional area of bioinformatics. Since the analytical programs are frequently developed by researchers, rather than software engineers, they tend to suffer from idiosyncratic and non-ergonomic man-machine interfaces. We report on HUSAR, our 140+ collection of third-party, as well as in-house developed or adapted, sequence manipulation and analysis tools, well integrated into the UNIX operating system environment and accessible via consistent menu-aware interface. Most of the HUSAR programs can be completely specified by UNIX command-line options; they can thus be run in batches or combined into pipes. Adding such a program into the HUSAR environment is almost a 'plug-and-play' exercise. HUSAR has been recently complemented with a graphical client interface, X-HUSAR, to support users on UNIX platforms with X11 windowing systems. The whole X-HUSAR interface is based on a single generic program, COMLIGEN, and a number of specific configuration files. COMLIGEN interprets those files and renders appropriate windows, menus, and other interactive elements, which help the end user in selecting application programs and specifying their options. Efforts of extending both HUSAR and X-HUSAR are roughly linear to the size of the collection.

Animals↗

The path to enlightenment: making sense of genomic and proteomic information.

Whereas genomics describes the study of genome, mainly represented by its gene expression on the DNA or RNA level, the term proteomics denotes the study of the proteome, which is the protein complement encoded by the genome. In recent years, the number of proteomic experiments increased tremendously. While all fields of proteomics have made major technological advances, the biggest step was seen in bioinformatics. Biological information management relies on sequence and structure databases and powerful software tools to translate experimental results into meaningful biological hypotheses and answers. In this resource article, I provide a collection of databases and software available on the Internet that are useful to interpret genomic and proteomic data. The article is a toolbox for researchers who have genomic or proteomic datasets and need to put their findings into a biological context.

Computational Biology↗

Jemboss reloaded.

Bioinformatics tools are freely available from websites all over the world. Often they are presented as web services, although there are many tools for download and use on a local machine. This tutorial section looks at Jemboss, a Java-based graphical user interface (GUI) for the EMBOSS bioinformatics suite, which combines the advantages of both web service and downloaded software.

Computational Biology↗

A streamlined approach to high-throughput proteomics.

Proteomics has rapidly become an important tool for life science research, allowing the integrated analysis of global protein expression from a single experiment. To accommodate the complexity and dynamic nature of any proteome, researchers must use a combination of disparate protein biochemistry techniques, often a highly involved and time-consuming process. Whilst highly sophisticated, individual technologies for each step in studying a proteome are available, true high-throughput proteomics that provides a high degree of reproducibility and sensitivity has been difficult to achieve. The development of high-throughput proteomic platforms, encompassing all aspects of proteome analysis and integrated with genomics and bioinformatics technology, therefore represents a crucial step for the advancement of proteomics research. ProteomIQ (Proteome Systems) is the first fully integrated, start-to-finish proteomics platform to enter the market. Sample preparation and tracking, centralized data acquisition and instrument control, and direct interfacing with genomics and bioinformatics databases are combined into a single suite of integrated hardware and software tools, facilitating high reproducibility and rapid turnaround times. This review will highlight some features of ProteomIQ, with particular emphasis on the analysis of proteins separated by 2D polyacrylamide gel electrophoresis.

Automation↗

[Cloning and characterization of three novel genes encoding transmembrane proteins of Schistosoma japonicum].

OBJECTIVE: To clone and analyze novel antigen molecules of Schistosoma japonicum (Sj), and to provide effective vaccine candidate antigens against schistosomiasis japonica. METHODS: Sj adult cDNA library was screened using sera of mice infected with Trichinella spiralis (Ts) and the inserts of positive clones were specifically amplified by PCR. The positive clones were sequenced and the sequence data were analyzed using Nucleotide BLAST software of NCBI and Expert Protein Analysis System of Swiss Institute of Bioinformatics. RESULTS: Nine positive clones were obtained after three rounds of immunoscreening. The size of these inserts ranged from 0.6 kb to 2.1 kb. Among five novel genes, Sj-Ts1, Sj-Ts3 and Sj-Ts5 (GenBank accession number: AY005816, AF299080 and AY024352, respectively) encode trans-membrane proteins with 83, 83 and 233 amino acids, respectively. Sj-Ts1 protein predicted contains one possible trans-membrance helix, one N-myristoylation site, two phosphorylation sites for protein kinase C and one for tyrosine kinase, Sj-Ts3 protein contains two possible transmembrance helices and one casein kinase II phosphorylation site, whereas Sj-Ts5 protein has five possible transmembrance helices, one N-glycosylation site, one N-myristoylation site, two phosphorylation sites for cAMP- and cGMP-dependent protein kinase and four for protein kinase C and one for casein kinase II. CONCLUSION: Three novel genes encoding three transmembrane proteins might be developed as new vaccine candidates against Sj infection.

Amino Acid Sequence↗

[Immunoscreening of Schistosoma japonicum adult worm cDNA library with sera vaccinated with cercaria antigen and analysis of novel genes].

OBJECTIVE: To explore antigens possessing common immunogenicity with Schistosoma japonicum (Sj) cercariae antigens, and to find out new candidate antigens for schistosomiasis diagnosis and vaccine. METHODS: Sj adult cDNA library was screened using sera from rabbits vaccinated with Sj cercariae antigen, the inserts of positive clones were amplified by PCR, all positive clones were sequenced and the data were analysed using Nucleotide BLAST software of NCBI and Expert Protein Analysis System of Swiss Institute of Bioinformatics. RESULTS: Thirteen positive clones were obtained after three rounds of immunoscreening, and all amplified by PCR. Among four novel genes, SjCAI, SjCA, SjCAI-2 and SjCAI-3 (GenBank accession number: AF495883, AF515834, AY118086 and AY129303, respectively) encoded proteins with 353, 161, 137 and 72 animo acids respectively. Sj CAI protein contained six DNA-binding zinc fingers and showed some homology to gastrula zinc finger protein XLCGF48.2; proteins encoded by SjCA, SjCAI-2 and SjCAI-3 respectively contained N-glycosylation sites and phosphorylation sites. CONCLUSION: Novel genes were obtained by immunoscreening Sj adult cDNA library.

Amino Acid Sequence↗