PubMed Health⌕ Search

Biomedical subjects

Nikola Stojanovic

Publications and source records attributed to Nikola Stojanovic.

3 recordsLinked to original sources

Identifying multiple alignment regions satisfying simple formulas and patterns.

MOTIVATION: When studying multiple alignments of genomic sequences one frequently aims to locate and count regions which satisfy a set of constraints. These regions may be putatively functional, but researchers may also be interested in quantifying the frequency of occurrences of certain patterns. RESULTS: We have developed a program that applies simple formulas and pattern specifications to multiple alignments, reporting the positions and counts of conforming regions. As an example, we have navigated a 15-species alignment of the CAV2-CAV1 region and outlined some findings regarding PPARgamma binding sites. AVAILABILITY: Our software and the accompanying documentation can be obtained at no charge by contacting the authors. It can also be accessed at http://ranger.uta.edu/~nick/compgen

Algorithms↗

Computational methods for the analysis of differential conservation in groups of similar DNA sequences.

Multiple sequence alignments are a powerful tool for identifying the regions of DNA which have been constrained in evolutionary divergence, presumably due to their functional role. However, such constraints rarely manifest themselves as perfect conservation of a site clearly standing out in its broader environment, as they reflect the species-specific differences in proteins, as well as the ability of some proteins to interact with multiple variants of their binding sequence. In this paper we explore the use of alignment column uncertainty as an aid in locating differential phylogenetic footprints, which refer to the sites in DNA where groups of related species exhibit sequence conservation, but where the pattern may vary between the groups. We use efficient, linear-time algorithms to locate such sites. We have performed a study of the mammalian CAV2-CAV1 gene region using our software, and we conclude with several observations concerning the differential conservation and the use of computational methods for its detection. The software developed for this project is available, free of charge, by contacting the author.

Animals↗

Identification of mixups among DNA sequencing plates.

MOTIVATION: During the process of high-throughput genome sequencing there are opportunities for mixups of reagents and data associated with particular projects. The sequencing templates or sequence data generated for an assembly may become contaminated with reagents or sequences from another project, resulting in poorer quality and inaccurate assemblies. RESULTS: We have developed a system to assess sequence assemblies and monitor for laboratory mixups. We describe several methods for testing the consistency of assemblies and resolving mixed ones. We use statistical tests to evaluate the distribution of sequencing reads from different plates into contigs, and a graph-based approach to resolve situations where data has been inappropriately combined. While these methods have been designed for use in a high-throughput DNA sequencing environment processing thousands of clones, they can be applied in any situation where distinct sequencing projects are performed at redundant coverage.

Algorithms↗