PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bioinformatics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Molecular cloning and bioinformatic analysis of SPATA4 gene.

Full-length cDNA sequences of four novel SPATA4 genes in chimpanzee, cow, chicken and ascidian were identified by bioinformatic analysis using mouse or human SPATA4 cDNA fragment as electronic probe. All these genes have 6 exons and have similar protein molecular weight and do not localize in sex chromosome. The mouse SPATA4 sequence is identified as significantly changed in cryptorchidism, which shares no significant homology with any known protein in swissprot databases except for the homologous genes in various vertebrates. Our searching results showed that all SPATA4 proteins have a putative conserved domain DUF1042. The percentages of putative SPATA4 protein sequence identity ranging from 30 % to 99 %. The high similarity was also found in 1 kb promoter regions of human, mouse and rat SPATA4 gene. The similarities of the sequences upstream of SPATA4 promoter also have a high proportion. The results of searching SymAtlas (http://symatlas.gnf.org/SymAtlas/) showed that human SPATA4 has a high expression in testis, especially in testis interstitial, leydig cell, seminiferous tubule and germ cell. Mouse SPATA4 was observed exclusively in adult mouse testis and almost no signal was detected in other tissues. The pI values of the protein are negative, ranging from 9.44 to 10.15. The subcellular location of the protein is usually in the nucleus. And the signal peptide possibilities for SPATA4 are always zero. Using the SNPs data in NCBI, we found 33 SNPs in human SPATA4 gene genomic DNA region, with the distribution of 29 SNPs in the introns. CpG island searching gives the data about CpG island, which shows that the regions of the CpG island have a high similarity with each other, though the length of the CpG island is different from each other. This research is a fundamental work in the fields of the bioinformational analysis, and also put forward a new way for the bioinformatic analysis of other genes.

Amino Acid Sequence↗

Functional Annotation Routines Used by ABRF Bioinformatics Core Facilities - Observations, Comparisons, and Considerations.

The functional annotation of gene lists is a common analysis routine required for most genomics experiments, and bioinformatics core facilities must support these analyses. In contrast to methods such as the quantitation of RNA-Seq reads or differential expression analysis, our research group noted a lack of consensus in our preferred approaches to functional annotation. To investigate this observation, we selected 4 experiments that represent a range of experimental designs encountered by our cores and analyzed those data with 6 tools used by members of the Association of Biomolecular Resource Facilities (ABRF) Genomic Bioinformatics Research Group (GBIRG). To facilitate comparisons between tools, we focused on a single biological result for each experiment. These results were represented by a gene set, and we analyzed these gene sets with each tool considered in our study to map the result to the annotation categories presented by each tool. In most cases, each tool produces data that would facilitate identification of the selected biological result for each experiment. For the exceptions, Fisher's exact test parameters could be adjusted to detect the result. Because Fisher's exact test is used by many functional annotation tools, we investigated input parameters and demonstrate that, while background set size is unlikely to have a significant impact on the results, the numbers of differentially expressed genes in an annotation category and the total number of differentially expressed genes under consideration are both critical parameters that may need to be modified during analyses. In addition, we note that differences in the annotation categories tested by each tool, as well as the composition of those categories, can have a significant impact on results.

Computational Biology↗

Integration of bioInformatics tools at the National University of Singapore (NUS).

In the past decade "Big Science" such as the Genome Project has generated an enormous amount of data in the life sciences. Concurrently, the synergy of this project with existing research has quickened the pace of biological discovery. But the major drawback that is beginning to be felt worldwide is the primitive level of organisation in the data accumulated. Without a proper framework or knowledge scaffold to hang and interconnect the various bits of data and information, the national knowledge-to-data ratio is declining rapidly. We are trying to serve a solution to this enigma by providing a World Wide Web (WWW) interface to Biosoftware and at the same time have come up with a database integration tool that can query heterogeneous, geographically scattered and disparate databases simultaneously. In this report we will talk about BioInformatics in general with specific reference to BioInformatics Centre (BIC) at the National University of Singapore.

Computational Biology↗

Recent advances in bioinformatics in the medical research environment and applications to the study of skin diseases.

BACKGROUND: The computer has become increasingly intertwined in society for the past 30 years. Within the academic health science centre, there is an increasing need for researchers to become skilled at using the Internet as a mechanism for the retrieval of scientific results and the underlying data. The discipline of bioinformatics, which uses computer technology to provide answers to biological questions, has been expanding in scope and utility for the past decade. Increasing numbers of research groups have been investing in bioinformatics infrastructure to aid in the research process. These continuing investments have led to the establishment for the first time of a supercomputing facility within a hospital. Such computational power is being used for the mapping of genes and the study of human disease. OBJECTIVE: A discussion of the increasing role of computational biology in the research environment of the clinician scientist is presented here. CONCLUSIONS: Though the investment in a supercomputer may not be possible in most research settings, several less expensive alternatives relying on existing desktop computers can provide supercomputer-like performance within nearly any environment.

Computational Biology↗

Bioinformatics in medical practice: what is necessary for a hospital?

Building bioinformatic facilities for a university hospital is pretty similar to using standardized building blocks to construct a house. Starting with the intention to built a dwelling house, a factory or just a shelter the architect draws a construction plan and determines the material to be used. In general, the building is then constructed by the workmen following exactly the plan. However, for particular reasons, minor alterations may be needed to improve the construction of the building. Here we use the metaphor of constructing a "bio-informatics building" to describe the steps needed to support the daily tasks of a university hospital medical microbiology department which uses genomic methods quite extensively for pathogen identification. Today the Giessen "bioinformatics building" is not yet complete but we have been able to lay solid foundations and erect the ground floor which is functional already. Using a combination of standard tools, internet accessible genomic databases and some own software tools we can support genome sequencing from the raw sequence to pathogen identification.

Computational Biology↗

Bioinformatics issues for automating the annotation of genomic sequences.

The rapid explosion in the amount of biological data being generated worldwide is surpassing efforts to manage analysis of the data. As part of an ongoing project to automate and manage bioinformatics analysis, the authors have designed and implemented a simple automated annotation system, which is described in this paper. The system is applied to existing GenBank/DDBJ/EMBL entries and compared with existing annotations to illustrate not only potential errors but also that they are generally not up-to-date, as a result of new versions of analysis tools and updates of genomic repositories. We highlight the important Bioinformatics issues of storage and management of information to ensure data and results are kept up-to-date in light of new information becoming available. Surprisingly, from just four database entries, a significant number of new features were found. We describe the results as well as identify important issues that need to be addressed in order to automate the re-analysis/re-annotation of genomic sequences within a reasonable timeframe.

Computational Biology↗

Grouping and identification of sequence tags (GRIST): bioinformatics tools for the NEIBank database.

NEIBank is a project to develop and organize genomics and bioinformatics resources for the eye. As part of this effort, tools have been developed for bioinformatics analysis and web based display of data from expressed sequence tag (EST) analyses. EST sequences are identified and formed into groups or clusters representing related transcripts from the same gene. This is carried out by a rules-based procedure called GRIST (GRouping and Identification of Sequence Tags) that uses sequence match parameters derived from BLAST programs. Linked procedures are used to eliminate non-mRNA contaminants. All data are assembled in a relational database and assembled for display as web pages with annotations and links to other informatics resources. Genome projects generate huge amounts of data that need to be classified and organized to become easily accessible to the research community. GRIST provides a useful tool for assembling and displaying the results of EST analyses. The NEIBank web site contains a growing set of pages cataloging the known transcriptional repertoire of eye tissues, derived from new NEIBank cDNA libraries and from eye-related data deposited in the dbEST section of GenBank.

Animals↗

Bioinformatics-based discovery of a novel factor with apparent specificity to colon cancer.

In a previous study, a data mining tool called Digital Differential Display (DDD) from the Cancer Genome Anatomy Project (CGAP) was used to predict solid tumor- and organ-specific genes from the expressed sequence tag (EST) database. To validate the use of bioinformatics approaches in gene discovery, one of the ESTs, which was predicted to be colon tumor-specific, was chosen for further study. Reverse Transcriptase-Polymerase Chain Reaction (RT-PCR) analysis of matched sets of cDNAs from normal and colon tumor tissues indicated that the EST was specifically expressed in the majority of colon tumors. Expression was also detected in early adenomas. Among other normal tissues, EST expression was detected only in the small intestine. The colon tumor specificity of this EST was inferred from the lack of expression in carcinomas of the breast, lung, ovary, pancreas and prostate. To validate the computational prediction of specificity, a full-length cDNA encompassing the entire open reading frame was cloned and, in view of its apparent specificity to the colon tumors, this gene was termed Colon Carcinoma Related Gene (CCRG). CCRG encodes a novel cysteine-rich motif and a putative signal peptide sequence. Supernatant from COS cells transfected with the CCRG expression vector stimulated proliferation of colon cancer cells. Immunoreactive CCRG was also detected in the paraffin sections of colon tumor samples. CCRG belongs to a new class of growth factors and may be important in the diagnosis and treatment of colon cancers. Identification of CCRG using bioinformatics approaches validates gene discovery using computational approaches.

Amino Acid Sequence↗

The Role of Bioinformatics in Identifying and Defining Potential Drug Targets.

The Second Annual Cutting Edge Approaches to Drug Design meeting, held March 13, 2002, in London, United Kingdom, was organized by the Royal Society of Chemistry's Molecular Modelling Group and the Structural Biology Group of the Biochemical Society. The emphasis was switched from last year's focus, the advances in structural biology and molecular modeling, to the role of bioinformatics in identifying and defining potential drug targets and the application of rational, computationally based techniques in library design and lead generation. Topics included the impact of bioinformatics on drug discovery, advances in chemical lead discovery and pharmacophore-based approaches to drug design. (c) 2002 Prous Science. All rights reserved.

Journal Article↗

Identification of probable genomic packaging signal sequence from SARS-CoV genome by bioinformatics analysis.

AIM: To predict the probable genomic packaging signal of SARS-CoV by bioinformatics analysis. The derived packaging signal may be used to design antisense RNA and RNA interfere (RNAi) drugs treating SARS. METHODS: Based on the studies about the genomic packaging signals of MHV and BCoV, especially the information about primary and secondary structures, the putative genomic packaging signal of SARS-CoV were analyzed by using bioinformatic tools. Multi-alignment for the genomic sequences was performed among SARS-CoV, MHV, BCoV, PEDV and HCoV 229E. Secondary structures of RNA sequences were also predicted for the identification of the possible genomic packaging signals. Meanwhile, the N and M proteins of all five viruses were analyzed to study the evolutionary relationship with genomic packaging signals. RESULTS: The putative genomic packaging signal of SARS-CoV locates at the 3' end of ORF1b near that of MHV and BCoV, where is the most variable region of this gene. The RNA secondary structure of SARS-CoV genomic packaging signal is very similar to that of MHV and BCoV. The same result was also obtained in studying the genomic packaging signals of PEDV and HCoV 229E. Further more, the genomic sequence multi-alignment indicated that the locations of packaging signals of SARS-CoV, PEDV, and HCoV overlaped each other. It seems that the mutation rate of packaging signal sequences is much higher than the N protein, while only subtle variations for the M protein. CONCLUSIONS: The probable genomic packaging signal of SARS-CoV is analogous to that of MHV and BCoV, with the corresponding secondary RNA structure locating at the similar region of ORF1b. The positions where genomic packaging signals exist have suffered rounds of mutations, which may influence the primary structures of the N and M proteins consequently.

Amino Acid Sequence↗

Small envelope protein E of SARS: cloning, expression, purification, CD determination, and bioinformatics analysis.

AIM: To obtain the pure sample of SARS small envelope E protein (SARS E protein), study its properties and analyze its possible functions. METHODS: The plasmid of SARS E protein was constructed by the polymerase chain reaction (PCR), and the protein was expressed in the E coli strain. The secondary structure feature of the protein was determined by circular dichroism (CD) technique. The possible functions of this protein were annotated by bioinformatics methods, and its possible three-dimensional model was constructed by molecular modeling. RESULTS: The pure sample of SARS E protein was obtained. The secondary structure feature derived from CD determination is similar to that from the secondary structure prediction. Bioinformatics analysis indicated that the key residues of SARS E protein were much conserved compared to the E proteins of other coronaviruses. In particular, the primary amino acid sequence of SARS E protein is much more similar to that of murine hepatitis virus (MHV) and other mammal coronaviruses. The transmembrane (TM) segment of the SARS E protein is relatively more conserved in the whole protein than other regions. CONCLUSION: The success of expressing the SARS E protein is a good starting point for investigating the structure and functions of this protein and SARS coronavirus itself as well. The SARS E protein may fold in water solution in a similar way as it in membrane-water mixed environment. It is possible that beta-sheet I of the SARS E protein interacts with the membrane surface via hydrogen bonding, this beta-sheet may uncoil to a random structure in water solution.

Circular Dichroism↗

A library of efficient bioinformatics algorithms.

In this paper we review some of the existing projects available in the bioinformatics field for facilitating the development of programs, but for which minimising the running time is not of primary importance. We point out the advantages of open source libraries for such tasks and we discuss some of the open source licenses available. Finally, we present the project ALiBio, which is aimed at facilitating the development of efficient programs in bioinformatics.

Algorithms↗

Artificial intelligence techniques for bioinformatics.

This review provides an overview of the ways in which techniques from artificial intelligence (AI) can be usefully employed in bioinformatics, both for modelling biological data and for making new discoveries. The paper covers three techniques: symbolic machine learning approaches (nearest neighbour and identification tree techniques), artificial neural networks and genetic algorithms. Each technique is introduced and supported with examples taken from the bioinformatics literature. These examples include folding prediction, viral protease cleavage prediction, classification, multiple sequence alignment and microarray gene expression analysis.

Algorithms↗

Bioinformatics: towards new directions for public health.

OBJECTIVES: Epidemiologists are reformulating their classical approaches to diseases by considering various issues associated to "omics" areas and technologies. Traditional differences between epidemiology and genetics include background, training, terminologies, study designs and others. Public health and epidemiology are increasingly looking forward to using methodologies and informatics tools, facilitated by the Bioinformatics community, for managing genomic information. Our aim is to describe which are the most important implications related with the increasing use of genomic information for public health practice, research and education. To review the contribution of bioinformatics to these issues, in terms of providing the methods and tools needed for processing genetic information from pathogens and patients. To analyze the research challenges in biomedical informatics related with the need of integration of clinical, environmental and genetic data and the new scenarios arisen in public health. METHODS: Review of the literature, Internet resources and material and reports generated by internal and external research projects. RESULTS: New developments are needed to advance in the study of the interactions between environmental agents and genetic factors involved in the development of diseases. The use of biomarkers, biobanks, and integrated genomic/clinical databases poses serious challenges for informaticians in order to extract useful information and knowledge for public health, biomedical research and healthcare. CONCLUSIONS: From an informatics perspective, integrated medical/biological ontologies and new semantic-based models for managing information provide new challenges for research in areas such as genetic epidemiology and the "omics" disciplines, among others. In this regard, there are various ethical, privacy, informed consent and social implications, that should be carefully addressed by researchers, practitioners and policy makers.

Computational Biology↗

[Bioinformatic analysis of Vibrio parahaemolyticus thermolabile hemolysin gene].

OBJECTIVE: To carry out bioinformatic analysis of Vibrio parahaemolyticus thermolabile hemolysin gene (tlh) obtained by PCR amplification. METHODS: The tlh gene amplified by PCR was cloned into the vector pET32a(+) and sequenced, followed by analysis of the biological information by with presenting the sequences to the websites of bioinformatics on the Internet. RESULTS AND CONCLUSIONS: The sequenced tlh gene (named tlh14-90) was entered into GenBank with the accession number of AY289609. Tlh14-90 has a length of 1 257 bp with both start and stop codons, having 99% homology with the tlh gene of WP1. Tlh14-90 is predicted to encode a protein containing 418 amino acids (named TLH14-90, with the molecular formula of C(2131)H(3184)N(548)O(649)S(16), molecular weight of 47 392.9, and the theoretical PI of 4.92). This protein consists of 4 Cys and the contents of Ala, Leu and Asn are 11.0%, 7.4% and 7.2%, respectively, having a hydrophobic parameter of -34 calculated using kappa-D method. Predicted as a alpha/protein for its secondary structure, TLH14-90 has not been identified for its tertiary structure.

Bacterial Proteins↗

[Gene screening of dermal papilla cells in the state of aggregative growth and full-length cloning of HSPC016 gene for bioinformatics].

OBJECTIVE: To screen the genes of dermal papilla cells (DPC) related to the property of aggregative growth, and clone the full-length cDNA of differential HSPC016 gene for functional analysis. METHODS: DPC were collected from the hair of an individual aged 18 approximately 30 6 hours after the death. The complete papillae of hair were isolated and then cultured. The total DNA was extracted. Suppression subtractive hybridization-polymerase chain reaction was employed to screen out the genes differentially expressed in the DPC under the state of aggregative growth pattern in vitro. Then, rapid amplification of cDNA ends (RACE) technique was used to amplify the full-length cDNA of HSPC016 in the DPC, and bioinformatic methods were used to analyze their possible function. RESULTS: A subtractive library of human DPC was set up, and some up-regulated and down-regulated genes in the DPCs were screened out successfully. HSPC016 was identified to be a gene of 400 bp cDNA. Bioinformatic analyses and databases searching on the Internet indicated that this gene was mapped on chromosome 3 q21.31 and included an open reading frame with 195 bp coding for an expected 64aa soluble protein. The putative protein belonged to PD053992 family and was homologous to T2FA gene in domain. CONCLUSION: The establishment of subtractive library of DPC provides a solid foundation for further screening the genes related to aggregative growth and analyzing their regulatory mechanisms in DPC. Further more, HSPC016 in DPC may act as a subunit of a functional complex and play a role on transcriptional regulation within nucleus.

Adult↗

Bioinformatics meets clinical informatics.

The field of bioinformatics has exploded over the past decade. Hopes have run high for the impact on preventive, diagnostic, and therapeutic capabilities of genomics and proteomics. As time has progressed, so has our understanding of this field. Although the mapping of the human genome will certainly have an impact on health care, it is a complex web to unweave. Addressing simpler "Single Nucleotide Polymorphisms" (SNPs) is not new, however, the complexity and importance of polygenic disorders and the greater role of the far more complex field of proteomics has become more clear. Proteomics operates much closer to the actual cellular level of human structure and proteins are very sensitive markers of health. Because the proteome, however, is so much more complex than the genome, and changes with time and environmental factors, mapping it and using the data in direct care delivery is even harder than for the genome. For these reasons of complexity, the expected utopia of a single gene chip or protein chip capable of analyzing an individual's genetic make-up and producing a cornucopia of useful diagnostic information appears still a distant hope. When, and if, this happens, perhaps a genetic profile of each individual will be stored with their medical record; however, in the mean time, this type of information is unlikely to prove highly useful on a broad scale. To address the more complex "polygenic" diseases and those related to protein variations, other tools will be developed in the shorter term. "Top-down" analysis of populations and diseases is likely to produce earlier wins in this area. Detailed computer-generated models will map a wide array of human and environmental factors that indicate the presence of a disease or the relative impact of a particular treatment. These models may point to an underlying genomic or proteomic cause, for which genomic or proteomic testing or therapies could then be applied for confirmation and/or treatment. These types of diagnostic and therapeutic requirements are most likely to be introduced into clinical practice through traditional forms of clinical practice guidelines and clinical decision support tools. The opportunities created by bioinformatics are enormous, however, many challenges and a great deal of additional research lay ahead before this research bears fruit widely at the care delivery level.

Computational Biology↗

[Cloning and bioinformatic analysis of a novel gene HCR2 up-regulated in aorta of hypercoagulable rat].

OBJECTIVE: The aim of this study was to clone the full-length cDNA of HCR2 up-regulated in aorta of hypercoagulable rat and make the relevant bioinformatic analysis. METHODS: Rapid amplification of cDNA end (RACE) and nested PCR technique was used to amplify the full-length cDNA of HCR2 from an EST (GenBank accession number BQ901227) which was significantly up-regulated in hypercoagulable rat. RESULTS: The full-length cDNA of HCR2 has been obtained (GenBank accession number AY234417). Semi-quantitative RT-PCR demonstrated that HCR2 is up-regulated in aorta of hypercoagulable rat. Bioinformatic analysis showed that HCR2 gene is located in rat chromosome 4q11. The full-length cDNA of HCR2 is 1275 bp, coding for a 78 amino acids polypeptide with a theoretical molecular weight of 8841.7 and isoelectric point of 8. 59. This protein has no signal peptide and transmembrane sequence. It may be a nuclear protein. It shares no significant homology with any known protein. CONCLUSION: The successful cloning of HCR2 has laid a solid foundation for further study of its function and its possible role in the occurrence and development of hypercoagulability.

Amino Acid Sequence↗