PubMed Health⌕ Search

Biomedical subjects

Nenad Trinajstić

Publications and source records attributed to Nenad Trinajstić.

7 recordsLinked to original sources

Toxicity of aliphatic ethers: a comparative study.

The CROMRsel procedure was used to model the toxicity of aliphatic ethers against mice. The best model obtained is based on three molecular descriptors and is a better model than other QSAR models from the literature. The only comparable model is one by Ren, based on four descriptors.

Animals↗

On reformulated Zagreb indices.

Zagreb indices were reformulated in terms of the edge-degrees instead of the vertex-degrees as the original Zagreb indices. Three types of Zagreb indices were considered: original, modified and variable Zagreb indices. It is found that the optimum exponent of the variable reformulated Zagreb M2 index (v = -1/2) is identical with the exponent of the vertex-connectivity index, which is the most used topological index in QSPR and QSAR. The close relationship between the graph and its line graph is used to relate the original and reformulated indices.

Algorithms↗

Nanotubes: number of Kekulé structures and aromaticity.

Carbon nanotubes (CNTs) are composed of cylindrical graphite sheets consisting of sp(2) carbons. Due to their structure CNTs are considered to be aromatic systems. In this work the number of Kekulé structures (K) in "armchair" CNTs was estimated by using the transfer matrix technique. All Kekulé structures of the cyclic variants of naphthalene and benzo[c]phenanthrene have been generated and the basic patterns have been obtained. From this information the elements of the transfer matrix were derived. The results obtained indicate that K (and the resonance energy) is greater if tubulenes are extended in the vertical than in the horizontal direction. Tubulenes are therefore more stabile than cyclic strips. An illustration, obtained by using scanning probe microscope, has been attached to affirm the existence of thin CNTs.

Journal Article↗

Toward generating simpler QSAR models: nonlinear multivariate regression versus several neural network ensembles and some related methods.

In this study we want to test whether a simple modeling procedure used in the field of QSAR/QSPR can produce simple models that will be, at the same time, as accurate as robust Neural Network Ensemble (NNE) ones. We present results of application of two procedures for generating/selecting simple linear and nonlinear multiregression (MR) models: (1) method for selecting the best possible MR models (named as CROMRsel) and (2) Genetic Function Approximation (GFA) method from the Cerius2 program package. The obtained MR models are strictly compared with several NNE models. For the comparison we selected four QSAR data sets previously studied by NNE (Tetko et al. J. Chem. Inf. Comput. Sci. 1996, 36, 794-803. Kovalishyn et al. J. Chem. Inf. Comput. Sci. 1998, 38, 651-659.): (1) 51 benzodiazepine derivatives, (2) 37 carboquinone derivatives, (3) 74 pyrimidines, and (4) 31 antimycin analogues. These data sets were parameterized with 7, 6, 27, and 53 descriptors, respectively. Modeled properties were anti-pentylenetetrazole activity, antileukemic activity, inhibition constants to dihydrofolate reductase from MB1428 E. coli, and antifilarial activity, respectively. Nonlinearities were introduced into the MR models through 2-fold and/or 3-fold cross-products of initial (linear) descriptors. Then, using the CROMRsel and GFA programs (J. Chem. Inf. Comput. Sci. 1999, 39, 121-132) the sets of I (I < or = 8, in this paper) the best descriptors (according to the fit and leave-one-out correlation coefficients) were selected for multiregression models. Two classes of models were obtained: (1) linear or nonlinear MR models which were generated starting from the complete set of descriptors, and (2) nonlinear MR models which were generated starting from the same set of descriptors that was used in the NNE modeling. In addition, the descriptor selection method from CROMRsel was compared with the GFA method included in the QSAR module of the Cerius2 program. For each data set it has been found that the MR models have better cross-validated statistical parameters than the corresponding NNE models and that CROMRsel selects somewhat better MR models than the GFA method. MR models are also much simpler than NNEs, which is the important surprising fact, and, additionally, express calculated dependencies in a functional form. Moreover, MR models were shown to be better than all other models obtained by different methods on the same data sets ("old" multivariate regressions, functional-link-net models, back-propagation neural networks, genetic algorithm, and partial least squares models). This study also indicated that the robust NNE models cannot generate good models when applied on small data sets, suggesting that it is perhaps better to apply robust methods (like NNE ones) on larger data sets.

Journal Article↗

Atomic walk counts of negative order.

Atomic walk counts (awc's) of order k (k > or = 1) are the number of all possible walks of length k which start at a specified vertex (atom) i and end at any vertex j separated by m (0 < or = m < or = k) edges from vertex i. The sum of atomic walk counts of order k is the molecular walk count (mwc) of order k. The concept of atomic and molecular walk counts was extended to zero and negative orders by using a backward algorithm based on the usual procedure used to obtain the values of mwc's. The procedure can also be used in cases in which the adjacency matrix A related to the actual structure is singular and therefore A(-1) does not exist. awc's and mwc's of negative order may assume noninteger and even negative values. If matrix A is singular, atomic walk counts of zero order may not be equal to one.

Journal Article↗

Use of variable selection in modeling the secondary structural content of proteins from their composition of amino acid residues.

The possibility of prediction of protein secondary structure content from composition of their amino acid residues can help in bridging the gap between proteins of known primary sequence having an unknown secondary structure. Almost all recently published models for understanding the relationship between composition (frequency of occurrence) of amino acid residues and secondary structure content of proteins involved composition of all 20 amino acid residues. However, it is well-known that many amino acid residues are mutually similar according to their physicochemical properties (hydrophobicity, hydrophilicity, charge, size, etc.). Because of that, we were motivated to investigate the possibility of reduction of the total number of terms (frequencies of amino acid residues) in the models for describing the relation between the composition of amino acid residues and the percentage of residues belonging to alpha, beta, and coil secondary structure. For this purpose, the CROMRsel algorithm (J. Chem. Inf. Comput. Sci. 1999, 39, 121-132) for selection of a small subset of the most important variables/descriptors into the multiregression (MR) models, i.e., frequency of occurrence of amino acid residues in proteins, was used. Analysis was performed on a data set containing 475 proteins, taken from Proteins 1996, 25, 157-168. A complete data set was partitioned into a 317-protein training set and 158-protein test set. The best possible linear models containing I=1, ..., 20 frequencies were selected among all 20 frequencies of occurrence of amino acid residues on the 317-protein training set, and were used for performing prediction of the corresponding percentage of secondary structure content on the 158-protein test set. For the 317-protein data set the best selected concise models for the alpha, beta, and coil secondary structure contain only 9, 5, and 8 frequencies, respectively. Selected concise models are of the same or better fitted, cross-validated, and predictive statistical parameters than the models containing all 20 frequencies. Additionally, for each I (I=1, ...., 20) 30 the best possible random models were selected. In each case, the best possible real models are much better than each of the best possible random models, showing clearly that there is no risk of a chance correlation (what one could expect due to the application of an exhaustive search for the best model having I frequencies among all 20!/I!(20-I)! possible models). Finally, the best selected models on the complete 475-protein data set for the alpha, beta, and coil secondary structure contain only 7, 4, and 7 frequencies of amino acid residues, respectively. These models are much simpler and have better fitted and cross-validated errors than the corresponding models from the literature, that were obtained without using a procedure for selection of the most important frequencies of amino acid residues in proteins.

Amino Acids↗

Random walks and chemical graph theory.

Simple random walks probabilistically grown step by step on a graph are distinguished from walk enumerations and associated equipoise random walks. Substructure characteristics and graph invariants correspondingly defined for the two types of random walks are then also distinct, though there often are analogous relations. It is noted that the connectivity index as well as some resistance-distance-related invariants make natural appearances among the invariants defined from the simple random walks.

Journal Article↗