PubMed Health⌕ Search

Biomedical subjects

Subhash C Basak

Publications and source records attributed to Subhash C Basak.

5 recordsLinked to original sources

Novel map descriptors for characterization of toxic effects in proteomics maps.

We consider a novel numerical characterization of proteomics maps based on the construction of a graph obtained by connecting all protein spots in a proteomics map that are at distance equal to, or smaller than, a critical distance D(c). We refer to the so constructed graph as a cluster graph and we calculate four associated characteristic matrices, previously considered in the literature: (1) the Euclidean-distance matrix ED; (2) the neighborhood-distance matrix ND; (3) the path-distance matrix based on the shortest paths between connected spots PD; and (4) the quotient matrix Q, the elements of which are given as the quotient of the corresponding elements of ED and ND matrices. Numerical descriptors for proteomics maps include in particular the leading eigenvalue of the Q matrix and the family of associated "higher order" matrices defined as powers of Q. These map descriptors show considerable sensitivity to perturbations of proteomics maps by toxicants.

Animals↗

A comparative study of proteomics maps using graph theoretical biodescriptors.

This paper reports the development of new methods for mathematical characterization of effects of different toxic agents on the cellular proteome. We describe numerical characterization of proteomics maps based on mathematical invariants. A graph is first associated with a proteomics map by considering partial ordering of spots on 2-D gels by ordering proteins with respect to the mass and the charge, the two properties by which proteins are separated. The graph is then embedded over the map, and several graph theoretical invariants have been constructed. In particular we consider invariants that can be extracted from the Euclidean distance-adjacency matrix of the embedded graph, in which only Euclidean distances between adjacent vertices of a graph are considered. The approach is illustrated using proteomics patterns of normal liver cells of rats and those derived from liver cells of animals exposed to four peroxisome proliferators. In contrast to direct comparison of spot abundance our approach incorporates information on spots locations. The difference between the two approaches is that in the first case only changes in abundances are considered as a measure of perturbation of the proteome map, but in the second case not only the charge but also the mass of proteins are used for ordering protein spots.

Animals↗

Prediction of cellular toxicity of halocarbons from computed chemodescriptors: a hierarchical QSAR approach.

A hierarchical quantitative structure-activity relationship (HiQSAR) approach was used to estimate toxicity and genetic toxicity for a set of 55 halocarbons using computed chemodescriptors. The descriptors consisted of topostructural (TS), topochemical (TC), geometrical, semiempirical (AM1) quantum chemical, and ab initio (STO-3G, 6-31G(d), 6-311G, 6-311G(d), and aug-cc-pVTZ) quantum chemical indices. For the two toxicity endpoints investigated, ARR and D(37), the TC indices gave the best cross-validated R(2) values. The 3-D indices also performed either as well as or slightly superior to the TC indices. For the four categories of quantum chemical indices used for the development of predictive models, the AM1 parameters gave the worst performance, and the most advanced ab initio (B3LYP/aug-CC-pVTZ) parameters gave the best results when used alone. This was also the case when the quantum chemical indices were used in the hierarchical QSAR approach for both of the toxicity endpoints, ARR and D(37). The models resulting from HiQSAR are of sufficiently good quality to estimate toxicity of halocarbons from structure.

Aspergillus niger↗

QSAR modeling of flotation collectors using principal components extracted from topological indices.

Several topological indices were calculated for substituted-cupferrons that were tested as collectors for the froth flotation of uranium. The principal component analysis (PCA) was used for data reduction. Seven principal components (PC) were found to account for 98.6% of the variance among the computed indices. The principal components thus extracted were used in stepwise regression analyses to construct regression models for the prediction of separation efficiencies (Es) of the collectors. A two-parameter model with a correlation coefficient of 0.889 and a three-parameter model with a correlation coefficient of 0.913 were formed. PCs were found to be better than partition coefficient to form regression equations, and inclusion of an electronic parameter such as Hammett sigma or quantum mechanically derived electronic charges on the chelating atoms did not improve the correlation coefficient significantly. The method was extended to model the separation efficiencies of mercaptobenzothiazoles (MBT) and aminothiophenols (ATP) used in the flotation of lead and zinc ores, respectively. Five principal components were found to explain 99% of the data variability in each series. A three-parameter equation with correlation coefficient of 0.985 and a two-parameter equation with correlation coefficient of 0.926 were obtained for MBT and ATP, respectively. The amenability of separation efficiencies of chelating collectors to QSAR modeling using PCs based on topological indices might lead to the selection of collectors for synthesis and testing from a virtual database.

Journal Article↗

Assessing model fit by cross-validation.

When QSAR models are fitted, it is important to validate any fitted model-to check that it is plausible that its predictions will carry over to fresh data not used in the model fitting exercise. There are two standard ways of doing this-using a separate hold-out test sample and the computationally much more burdensome leave-one-out cross-validation in which the entire pool of available compounds is used both to fit the model and to assess its validity. We show by theoretical argument and empiric study of a large QSAR data set that when the available sample size is small-in the dozens or scores rather than the hundreds, holding a portion of it back for testing is wasteful, and that it is much better to use cross-validation, but ensure that this is done properly.

Journal Article↗