PubMed Health⌕ Search

PubMed · 15333141

POSA: perl objects for DNA sequencing data analysis.

Abstract

BACKGROUND: Capillary DNA sequencing machines allow the generation of vast amounts of data with little hands-on time. With this expansion of data generation, there is a growing need for automated data processing. Most available software solutions, however, still require user intervention or provide modules that need advanced informatics skills to allow implementation in pipelines. RESULTS: Here we present POSA, a pair of new perl objects that describe DNA sequence traces and Phrap contig assemblies in detail. Methods included in POSA include basecalling with quality scores (by Phred), contig assembly (by Phrap), generation of primer3 input and automated SNP annotation (by PolyPhred). Although easily implemented by users with only limited programming experience, these objects considerabily reduce hands-on analysis time compared to using the Staden package for extracting sequence information from raw sequencing files and for SNP discovery. CONCLUSIONS: The POSA objects allow a flexible and easy design, implementation and usage of perl-based pipelines to handle and analyze DNA sequencing data, while requiring only minor programming skills.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jan A Aerts, Bart J Jungerius, Martien A M Groenen. 2004-08-27. POSA: perl objects for DNA sequencing data analysis.. https://doi.org/10.1186/1471-2164-5-60

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

EST analysis in Ginkgo biloba: an assessment of conserved developmental regulators and gymnosperm specific genes.

BACKGROUND: Ginkgo biloba L. is the only surviving member of one of the oldest living seed plant groups with medicinal, spiritual and horticultural importance worldwide. As an evolutionary relic, it displays many characters found in the early, extinct seed plants and extant cycads. To establish a molecular base to understand the evolution of seeds and pollen, we created a cDNA library and EST dataset from the reproductive structures of male (microsporangiate), female (megasporangiate), and vegetative organs (leaves) of Ginkgo biloba. RESULTS: RNA from newly emerged male and female reproductive organs and immature leaves was used to create three distinct cDNA libraries from which 6,434 ESTs were generated. These 6,434 ESTs from Ginkgo biloba were clustered into 3,830 unigenes. A comparison of our Ginkgo unigene set against the fully annotated genomes of rice and Arabidopsis, and all available ESTs in Genbank revealed that 256 Ginkgo unigenes match only genes among the gymnosperms and non-seed plants--many with multiple matches to genes in non-angiosperm plants. Conversely, another group of unigenes in Gingko had highly significant homology to transcription factors in angiosperms involved in development, including MADS box genes as well as post-transcriptional regulators. Several of the conserved developmental genes found in Ginkgo had top BLAST homology to cycad genes. We also note here the presence of ESTs in G. biloba similar to genes that to date have only been found in gymnosperms and an additional 22 Ginkgo genes common only to genes from cycads. CONCLUSION: Our analysis of an EST dataset from G. biloba revealed genes potentially unique to gymnosperms. Many of these genes showed homology to fully sequenced clones from our cycad EST dataset found in common only with gymnosperms. Other Ginkgo ESTs are similar to developmental regulators in higher plants. This work sets the stage for future studies on Ginkgo to better understand seed and pollen evolution, and to resolve the ambiguous phylogenetic relationship of G. biloba among the gymnosperms.

Contig Mapping↗

Whole-genome shotgun optical mapping of Rhodospirillum rubrum.

Rhodospirillum rubrum is a phototrophic purple nonsulfur bacterium known for its unique and well-studied nitrogen fixation and carbon monoxide oxidation systems and as a source of hydrogen and biodegradable plastic production. To better understand this organism and to facilitate assembly of its sequence, three whole-genome restriction endonuclease maps (XbaI, NheI, and HindIII) of R. rubrum strain ATCC 11170 were created by optical mapping. Optical mapping is a system for creating whole-genome ordered restriction endonuclease maps from randomly sheared genomic DNA molecules extracted from cells. During the sequence finishing process, all three optical maps confirmed a putative error in sequence assembly, while the HindIII map acted as a scaffold for high-resolution alignment with sequence contigs spanning the whole genome. In addition to highlighting optical mapping's role in the assembly and confirmation of genome sequence, this work underscores the unique niche in resolution occupied by the optical mapping system. With a resolution ranging from 6.5 kb (previously published) to 45 kb (reported here), optical mapping advances a "molecular cytogenetics" approach to solving problems in genomic analysis.

Contig Mapping↗

Genomic analysis of a region encompassing QRfs1 and QRfs2: genes that underlie soybean resistance to sudden death syndrome.

Candidate genes were identified for two loci, QRfs2 providing resistance to the leaf scorch called soybean (Glycine max (L.) Merr.) sudden death syndrome (SDS) and QRfs1 providing resistance to root infection by the causal pathogen Fusarium solani f.sp. glycines. The 7.5 +/- 0.5 cM region of chromosome 18 (linkage group G) was shown to encompass a cluster of resistance loci using recombination events from 4 near-isogenic line populations and 9 DNA markers. The DNA markers anchored 9 physical map contigs (7 are shown on the soybean Gbrowse, 2 are unpublished), 45 BAC end sequences (41 in Gbrowse), and contiguous DNA sequences of 315, 127, and 110 kbp. Gene density was high at 1 gene per 7 kbp only around the already sequenced regions. Three to 4 gene-rich islands were inferred to be distributed across the entire 7.5 cM or 3.5 Mbp showing that genes are clustered in the soybean genome. Candidate resistance genes were identified and a molecular basis for interactions among the disease resistance genes in the cluster inferred.

Contig Mapping↗