PubMed Health⌕ Search

PubMed · 16551601

[Construction of standard human transcript dataset based on RefSeq and human genome sequence database].

Abstract

The NCBI Reference Sequence (RefSeq) database aimed to provide a biologically non-redundant collection of DNA, RNA, and protein sequences and to promote the research on genes and proteins of human beings and other species. However, because of widely distributed polymorphisms and different quality control of experiments in individual laboratories, there are potential problems need to be identified in the RefSeq database. Regarding which, we herein define the concept, standard transcript, based on the Central Dogmas of Biology that each standard transcript should be perfectly mapped to the standard genomic DNA sequence at the exon level. A large scale analysis for mapping all of the RefSeq records of human being (2005-4-18) to the officially released human genome sequence database (2005-4-20) was further performed using BLAT, Sim4 and a homemade program, EIparser, which was especially designed for this purpose. The standard transcripts based on the RefSeq database were obtained according to the alignment with standard human genome database. There are 9,771 RefSeq records of human being labeled with "NM_" and "NR_" could be perfectly mapped to human genome sequences, while other 10,943 records could be considered as standard transcripts after reasonable revision by comparing with the genome sequences according to all of the three methods. Moreover, the left 203 unrevisable records and 2,676 inconsistent records reported by the above programs could not be considered as standard transcripts and should be checked critically before using because of potential errors in them. Our study has thus provided a reference standard dataset of human beings with high quality for further bioinformatic and experimental analysis such as polymorphism and mutation of human genes. The reference standard dataset based on above criteria could be retrieved from http://biocompute.bmi.ac.cn/transcriptome/index.htm.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Zhi-Feng Li, Yu-Jian Li, Dong-Sheng Zhao, Xing-Yi Hang, Zheng-Zhi Wang, Zhi-Gang Luo, Cheng-Gang Zhang. 2006. [Construction of standard human transcript dataset based on RefSeq and human genome sequence database].. https://pubmed.ncbi.nlm.nih.gov/16551601/

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

CamK-DB: A k-mer MinHash fingerprint database for reference-free genotyping of Camellia accessions.

Tea (Camellia sinensis L.), a major global economic crop in Asia, poses challenges for genetic identification because its highly heterozygous, repetitive genome reduces the efficacy of conventional single-nucleotide polymorphism (SNP) and microsatellite markers, and interspecific hybridization further complicates the situation. To address these issues, CamK-DB was developed as a reference-free Camellia fingerprinting database built on MIKE MinHash sketches. We curated 418 candidate resequencing datasets, and built a database using standardized 5× genome-coverage fingerprints. Each accession is stored as a MIKE. jac fingerprint generated with k = 21 and recommended sketch/pre_cnt = 2000. CamK-DB provides a command-line interface for data management and a custom C++ query engine that computes top-10 matches using Jaccard similarity, complemented by a QT-based graphical interface for interactive analysis. This resource offers a robust and scalable framework for precise and routine germplasm identification, genomic phylogenetic inference, and strategic breeding program design. CamK-DB (database and code) is publicly available at https://github.com/sc-zhang/CamK-DB. CamK-DB binaries are provided for Windows 10/11 and Linux (x86_64, glibc ≥ 2.27).

Databases, Genetic↗

The Saccharomyces Genome Database-a history of ideas and accomplishments, 1994-2026.

The Saccharomyces Genome Database (SGD) is one of the longest-running and most consequential biological databases in the world. Founded in the early 1990s at Stanford University under the visionary leadership of David Botstein and developed under the long-term technical direction of J. Michael Cherry, SGD has served for more than three decades not only as the authoritative knowledge center for the budding yeast Saccharomyces cerevisiae, but also as the source for much of the fundamentals of eukaryotic biology. This history traces the arc of a remarkable intellectual and scientific project: beginning with the challenge of building the very first integrated eukaryotic genome database and evolving across 30 years into a global knowledge hub for genetics, functional genomics, and human disease research. The history is organized chronologically, with each section highlighting the central ideas, technical developments, and concrete accomplishments of that period.

Databases, Genetic↗

Gencube: centralized retrieval and integration of multi-omics resources from leading databases.

MOTIVATION: The volume of multi-omics data for diverse species is growing at an unprecedented rate, with new genome assemblies, related annotations, and high-throughput sequencing resources being submitted daily to various genomic data repositories. In response to this data influx, both existing and new databases are establishing optimized hierarchical structures to manage the vast amount of information. However, the lack of accessible command-line tools, combined with the functional limitations and unintuitive design of existing options, presents significant challenges for researchers. This gap underscores a critical need for a tool that enables streamlined retrieval and integration of omics data across these diverse repositories. RESULTS: We have developed Gencube, a command-line tool that enables centralized retrieval and integration of a comprehensive set of six different data types-genome assemblies, gene sets, annotations, sequences, comparative genomic data, and NGS-based omics resources-from various leading databases. AVAILABILITY AND IMPLEMENTATION: Gencube is a free and open-source tool, with its code available on GitHub: https://github.com/snu-cdrc/gencube and also archived on Zenodo: https://doi.org/10.5281/zenodo.14607649.

Databases, Genetic↗