PubMed Health⌕ Search

Biomedical subjects

Coral del Val

Publications and source records attributed to Coral del Val.

6 recordsLinked to original sources

CAFTAN: a tool for fast mapping, and quality assessment of cDNAs.

BACKGROUND: The German cDNA Consortium has been cloning full length cDNAs and continued with their exploitation in protein localization experiments and cellular assays. However, the efficient use of large cDNA resources requires the development of strategies that are capable of a speedy selection of truly useful cDNAs from biological and experimental noise. To this end we have developed a new high-throughput analysis tool, CAFTAN, which simplifies these efforts and thus fills the gap between large-scale cDNA collections and their systematic annotation and application in functional genomics. RESULTS: CAFTAN is built around the mapping of cDNAs to the genome assembly, and the subsequent analysis of their genomic context. It uses sequence features like the presence and type of PolyA signals, inner and flanking repeats, the GC-content, splice site types, etc. All these features are evaluated in individual tests and classify cDNAs according to their sequence quality and likelihood to have been generated from fully processed mRNAs. Additionally, CAFTAN compares the coordinates of mapped cDNAs with the genomic coordinates of reference sets from public available resources (e.g., VEGA, ENSEMBL). This provides detailed information about overlapping exons and the structural classification of cDNAs with respect to the reference set of splice variants. The evaluation of CAFTAN showed that is able to correctly classify more than 85% of 5950 selected "known protein-coding" VEGA cDNAs as high quality multi- or single-exon. It identified as good 80.6 % of the single exon cDNAs and 85 % of the multiple exon cDNAs. The program is written in Perl and in a modular way, allowing the adoption of this strategy to other tasks like EST-annotation, or to extend it by adding new classification rules and new organism databases as they become available. We think that it is a very useful program for the annotation and research of unfinished genomes. CONCLUSION: CAFTAN is a high-throughput sequence analysis tool, which performs a fast and reliable quality prediction of cDNAs. Several thousands of cDNAs can be analyzed in a short time, giving the curator/scientist a first quick overview about the quality and the already existing annotation of a set of cDNAs. It supports the rejection of low quality cDNAs and helps in the selection of likely novel splice variants, and/or completely novel transcripts for new experiments.

Chromosome Mapping↗

GOPET: a tool for automated predictions of Gene Ontology terms.

BACKGROUND: Vast progress in sequencing projects has called for annotation on a large scale. A Number of methods have been developed to address this challenging task. These methods, however, either apply to specific subsets, or their predictions are not formalised, or they do not provide precise confidence values for their predictions. DESCRIPTION: We recently established a learning system for automated annotation, trained with a broad variety of different organisms to predict the standardised annotation terms from Gene Ontology (GO). Now, this method has been made available to the public via our web-service GOPET (Gene Ontology term Prediction and Evaluation Tool). It supplies annotation for sequences of any organism. For each predicted term an appropriate confidence value is provided. The basic method had been developed for predicting molecular function GO-terms. It is now expanded to predict biological process terms. This web service is available via http://genius.embnet.dkfz-heidelberg.de/menu/biounit/open-husar CONCLUSION: Our web service gives experimental researchers as well as the bioinformatics community a valuable sequence annotation device. Additionally, GOPET also provides less significant annotation data which may serve as an extended discovery platform for the user.

Artificial Intelligence↗

The LIFEdb database in 2006.

LIFEdb (http://www.LIFEdb.de) integrates data from large-scale functional genomics assays and manual cDNA annotation with bioinformatics gene expression and protein analysis. New features of LIFEdb include (i) an updated user interface with enhanced query capabilities, (ii) a configurable output table and the option to download search results in XML, (iii) the integration of data from cell-based screening assays addressing the influence of protein-overexpression on cell proliferation and (iv) the display of the relative expression ('Electronic Northern') of the genes under investigation using curated gene expression ontology information. LIFEdb enables researchers to systematically select and characterize genes and proteins of interest, and presents data and information via its user-friendly web-based interface.

Cell Proliferation↗

Lentiviral vector integration sites in human NOD/SCID repopulating cells.

BACKGROUND: Recent observations of insertional mutagenesis in preclinical and clinical settings emphasize the relevance of investigating comprehensively the spectrum of integration sites targeted by specific vectors. METHODS: We followed the engraftment of lentivirally transduced human cord blood (CB) progenitor cells after transplantation into NOD/SCID mice using a self-inactivating HIV-1-derived vector expressing the enhanced green fluorescent protein (EGFP). RESULTS: The mean of transduction of CD34(+) CB cells was 41%, as deduced from the percentage of EGFP(+) cells before transplantation. At 3 weeks post-transplantation, the average of EGFP(+) cells in the human cell population was 65 +/- 8%, and increased to 75 +/- 10% at 12 weeks post-transplantation. In order to determine the proviral integration sites in human NOD/SCID repopulating cells (SRCs) we used the ligation-mediated polymerase chain reaction (LM-PCR) technique. Sixty-eight percent of the integrations were found to be located in RefSeq genes, most of them in intron regions. Twenty percent of these integrations occurred within a distance of 10 kb from the transcription start site; a percentage that is significantly lower compared to that observed in cells transduced by gammaretroviral vectors. Sixty-two percent of integrations occurred in genes with a biological function in cell metabolism, and four integrations were located in genes with a role in tumorigenesis. CONCLUSIONS: These investigations indicate that integration of lentiviral vectors in human repopulating cells capable of engrafting NOD/SCID mice preferentially occur in coding regions of the human genome. Nevertheless, the clustering of integrations at the transcriptional start is not as high as that observed for gammaretroviral vectors.

Animals↗

High-throughput protein analysis integrating bioinformatics and experimental assays.

The wealth of transcript information that has been made publicly available in recent years requires the development of high-throughput functional genomics and proteomics approaches for its analysis. Such approaches need suitable data integration procedures and a high level of automation in order to gain maximum benefit from the results generated. We have designed an automatic pipeline to analyse annotated open reading frames (ORFs) stemming from full-length cDNAs produced mainly by the German cDNA Consortium. The ORFs are cloned into expression vectors for use in large-scale assays such as the determination of subcellular protein localization or kinase reaction specificity. Additionally, all identified ORFs undergo exhaustive bioinformatic analysis such as similarity searches, protein domain architecture determination and prediction of physicochemical characteristics and secondary structure, using a wide variety of bioinformatic methods in combination with the most up-to-date public databases (e.g. PRINTS, BLOCKS, INTERPRO, PROSITE SWISSPROT). Data from experimental results and from the bioinformatic analysis are integrated and stored in a relational database (MS SQL-Server), which makes it possible for researchers to find answers to biological questions easily, thereby speeding up the selection of targets for further analysis. The designed pipeline constitutes a new automatic approach to obtaining and administrating relevant biological data from high-throughput investigations of cDNAs in order to systematically identify and characterize novel genes, as well as to comprehensively describe the function of the encoded proteins.

Automation↗

PATH: a task for the inference of phylogenies.

UNLABELLED: Phylogenetic Analysis Task in Husar (PATH) is a task for the inference of phylogenies. It executes three phylogenetic methods and automatically chooses the evolutionary model for each set of data. The output of the tasks shows the consensus trees together with full results obtained from all executed methods. AVAILABILITY: PATH is available at the German EMBnet node after registration via www at http://genome.dkfz-heidelberg.de

Databases, Nucleic Acid↗