PubMed Health⌕ Search

Biomedical subjects

Richard Baldarelli

Publications and source records attributed to Richard Baldarelli.

3 recordsLinked to original sources

CDS annotation in full-length cDNA sequence.

The identification of coding sequences (CDS) is an important step in the functional annotation of genes. CDS prediction for mammalian genes from genomic sequence is complicated by the vast abundance of intergenic sequence in the genome, and provides little information about how different parts of potential CDS regions are expressed. In contrast, mammalian gene CDS prediction from cDNA sequence offers obvious advantages, yet encounters a different set of complexities when performed on high-throughput cDNA (HTC) sequences, such as the set of 60,770 cDNAs isolated from full-length enriched libraries of the FANTOM2 project. We developed a CDS annotation strategy that uses a variety of different CDS prediction programs to annotate the CDS regions of FANTOM2 cDNAs. These include rsCDS, which uses sequence similarity to known proteins; ProCrest; Longest-ORF and Truncated-ORF, which are ab initio based predictors; and finally, DECODER and NCBI CDS predictor, which use a combination of both principles. Aided by graphical displays of these CDS prediction results in the context of other sequence similarity results for each cDNA, FANTOM2 CDS inspection by curators and follow-up quality control procedures resulted in high quality CDS predictions for a total of 14,345 FANTOM2 clones.

Animals↗

Development and evaluation of an automated annotation pipeline and cDNA annotation system.

Manual curation has long been held to be the "gold standard" for functional annotation of DNA sequence. Our experience with the annotation of more than 20,000 full-length cDNA sequences revealed problems with this approach, including inaccurate and inconsistent assignment of gene names, as well as many good assignments that were difficult to reproduce using only computational methods. For the FANTOM2 annotation of more than 60,000 cDNA clones, we developed a number of methods and tools to circumvent some of these problems, including an automated annotation pipeline that provides high-quality preliminary annotation for each sequence by introducing an "uninformative filter" that eliminates uninformative annotations, controlled vocabularies to accurately reflect both the functional assignments and the evidence supporting them, and a highly refined, Web-based manual annotation tool that allows users to view a wide array of sequence analyses and to assign gene names and putative functions using a consistent nomenclature. The ultimate utility of our approach is reflected in the low rate of reassignment of automated assignments by manual curation. Based on these results, we propose a new standard for large-scale annotation, in which the initial automated annotations are manually investigated and then computational methods are iteratively modified and improved based on the results of manual curation.

Animals↗

Mammalian septins nomenclature.

There are 10 known mammalian septin genes, some of which produce multiple splice variants. The current nomenclature for the genes and gene products is very confusing, with several different names having been given to the same gene product and distinct names given to splice variants of the same gene. Moreover, some names are based on those of yeast or Drosophila septins that are not the closest homologues. Therefore, we suggest that the mammalian septin field adopt a common nomenclature system, based on that adopted by the Mouse Genomic Nomenclature Committee and accepted by the Human Genome Organization Gene Nomenclature Committee. The human and mouse septin genes will be named SEPT1-SEPT10 and Sept1-Sept10, respectively. Splice variants will be designated by an underscore followed by a lowercase "v" and a number, e.g., SEPT4_v1.

Alternative Splicing↗