PubMed Health⌕ Search

Biomedical subjects

Anup Som

Publications and source records attributed to Anup Som.

2 recordsLinked to original sources

Theoretical foundation to estimate the relative efficiencies of the Jukes-Cantor+gamma model and the Jukes-Cantor model in obtaining the correct phylogenetic tree.

This paper deals with the theoretical foundation to estimate the relative efficiency (probability of inferring the true tree) of different nucleotide substitution models. A novel theoretical approach has been developed to estimate the relative efficiency of nucleotide substitution models based on the neighbor-relation method. The theory was developed directly by using the four-point condition. Initially theoretical formulas for the Jukes-Cantor+gamma (JC+Gamma) model and the Jukes-Cantor (JC) model were developed to estimate their relative efficiencies. Theoretical formulas were used on several model topologies for both models to obtain a true tree. Extensive simulations were performed on the same set of trees to test the strength of the theoretical approach. Simulation results demonstrated a good agreement with results obtained by the theoretical estimation. Overall the theoretical foundation for estimating model efficiencies is very accurate.

Algorithms↗

Word organization in coding DNA: a mathematical model.

This article deals with the relationship between vocabulary (total number of distinct oligomers or "words") and text-length (total number of oligomers or "words") for a coding DNA sequence (CDS). For natural human languages, Heaps established a mathematical formula known as Heaps' law, which relates vocabulary to text-length. Our analysis shows that Heaps' law fails to model this relationship for CDSs. Here we develop a mathematical model to establish the relationship between the number of type of words (vocabulary) and the number of words sampled (text-length) for CDSs, when non-overlapping nucleotide strings with the same length are treated as words. We use tangent-hyperbolic function, which captures the saturation property of vocabulary. Based on the parameters of the model, we formulate a mathematical equation, known as "equation of word organization", whose parameters essentially indicate that nucleotide organization of coding sequences are different from one another. We also compare the word organization of CDSs with the random word distribution and conclude that a CDS is neither similar to a natural human language nor to a random one. Moreover, these sequences have their unique nucleotide organization and it is completely structured for specific biological functioning.

Base Composition↗