PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “collinearity analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

QSTR with extended topochemical atom indices. 2. Fish toxicity of substituted benzenes.

Considering the importance of quantitative structure-toxicity relationship (QSTR) studies in the field of aquatic toxicology from the viewpoint of ecological safety assessment, fish toxicity of various benzene derivatives has been modeled by the multiple regression technique using recently introduced extended topochemical atom (ETA) indices. The toxicity data have also been modeled using other selected topological descriptors and physicochemical variables, and the best ETA model has been compared to the non-ETA ones. Principal component factor analysis was used as the data preprocessing step to reduce the dimensionality of the data matrix and identify the important variables that are devoid of collinearities. All-possible-subsets regression was also applied on the parameters to cross-check the variable selection for the best model. Multiple linear regression analyses show that the best non-ETA model involves 1chi, ALogP98, and LUMO (energy) as predictor variables and the quality of the relation is as follows: n = 92, Q2 = 0.718, Ra2 = 0.730, R2 = 0.738, R = 0.859, F = 82.8 (df 3, 88), s = 0.340. On the other hand, the best ETA model has the following quality: n = 92, Q2 = 0.865, Ra2 = 0.876, R2 = 0.885, R = 0.941, F = 92.6 (df 7, 84), s = 0.230. The ETA relations showed positive contributions of molecular bulk (size), chloro and hydroxy substitutions in the benzene ring, and the simultaneous presence of methyl and nitro substitutions to the toxicity. Further, the presence of fluoro and ether functionality, amino or nitro functionality in an otherwise unsubstituted ring, and nitro functionality that is ortho to a chloro substituent decreases toxicity. An attempt to use non-ETA descriptors along with ETA ones did not improve the quality in comparison to the best ETA model. Interestingly, the ETA model developed presently for the fish toxicity is better than the previously reported models on the same data set. Thus, it appears that ETA descriptors have significant potential in QSAR/QSPR/QSTR studies, which warrants extensive evaluation.

Journal Article↗

Syntenic relationships between Medicago truncatula and Arabidopsis reveal extensive divergence of genome organization.

Arabidopsis and Medicago truncatula represent sister clades within the dicot subclass Rosidae. We used genetic map-based and bacterial artificial chromosome sequence-based approaches to estimate the level of synteny between the genomes of these model plant species. Mapping of 82 tentative orthologous gene pairs reveals a lack of extended macrosynteny between the two genomes, although marker collinearity is frequently observed over small genetic intervals. Divergence estimates based on non-synonymous nucleotide substitutions suggest that a majority of the genes under analysis have experienced duplication in Arabidopsis subsequent to divergence of the two genomes, potentially confounding synteny analysis. Moreover, in cases of localized synteny, genetically linked loci in M. truncatula often share multiple points of synteny with Arabidopsis; this latter observation is consistent with the large number of segmental duplications that compose the Arabidopsis genome. More detailed analysis, based on complete sequencing and annotation of three M. truncatula bacterial artificial chromosome contigs suggests that the two genomes are related by networks of microsynteny that are often highly degenerate. In some cases, the erosion of microsynteny could be ascribed to the selective gene loss from duplicated loci, whereas in other cases, it is due to the absence of close homologs of M. truncatula genes in Arabidopsis.

Arabidopsis↗

Analysis of tryptophan and tyrosine in cerebrospinal fluid by capillary electrophoresis and "ball lens" UV-pulsed laser-induced fluorescence detection.

For the purpose of this study, we used a "ball lens" UV laser-induced fluorescence (LIF) detector comprising a pulsed laser and a collinear optical arrangement. The fluorescence signal is induced by a pulsed laser and detected by a photomultiplier tube. When coupling the high-frequency pulsed laser to the LIF detector we used, the electronics which is designed for continuous wavelength (CW) lasers, "viewed" the laser as a continuous source. Despite this mismatch between the laser and the "ball lens" UV LIF detector, the sensitivity we obtained with tryptophan is comparable to the one obtained with the best "laboratory-made" detector described in the literature which used a CW UV laser. Limits of detection of 0.15 nM for tryptophan and 50 nM for tyrosine were estimated. As an application of this technology, we studied tryptophan and tyrosine in cerebrospinal fluids (CSFs). The analysis is very simple and works on very small samples (5 microl). It consists of using a 10 mM 3-cyclohexylamino-1-propanesulfonic acid, 15 mM sodium tetraborate, pH 9.2 buffer and injecting CSF diluted 20 times in water prior to injection. 5-Hydroxyindoleacetic acid was used as an internal standard. The separation is completed in less than 12 min. The capillary electrophoresis method which we chose is rapid, resolutive and allows accurate measurements. Recovery experiments in CSFs show recoveries between 97 and 102%. We investigated 14 different CSFs from patients who suffered from neurological disorders. Most of the concentrations vary in a range of 1.7 to 3.7 microM for Trp and 6.6 to 13.7 microM for Tyr, which is in the range observed in the literature. One patient who suffers from Huntington disease had a higher concentration of Tyr at 17.3 microM.

Buffers↗

Identification and transcriptional analysis of the homologues of the herpes simplex virus type 1 UL41 to UL51 genes in the genome of nononcogenic Marek's disease virus serotype 2.

Studies on Marek's disease virus serotype 2 (MDV2) are important for understanding the natural nonpathogenic phenotypes of MDV. We determined the 16770 bp nucleotide sequence of the MDV2 genome located in the right part the of unique long region. The analysis revealed 12 complete open reading frames (ORFs) with high amino acid sequence identities to the gene products of other alphaherpesviruses. The MDV2 ORFs were arranged collinearly with the prototype sequence of herpes simplex virus type 1 ranging from the UL41 to UL51 genes. Except for the MDV2 UL41 gene, all of the identified genes were confirmed to be transcribed with 3'-coterminal mRNAs and/or a unique transcript in the virus-infected cells. Transcriptional patterns for the regions of the MDV2 UL48 to UL49.5 genes were notably different from the similar area of MDV serotype 1.

Animals↗

Isolation and characterization of a chimpanzee alphaherpesvirus.

Although both beta- and gammaherpesviruses indigenous to great-ape species have been isolated, to date all alphaherpesviruses isolated from apes have proven to be human viruses [herpes simplex virus types 1 (HSV1) and 2 (HSV2) or varicella-zoster virus]. If the alphaherpesviruses have co-evolved with their host species, some if not all ape species should harbour their own alphaherpesviruses. Here, the isolation and characterization of an alphaherpesvirus from a chimpanzee (ChHV) are described. Sequencing of a number of genes throughout the ChHV genome indicates that it is collinear with that of HSV. Phylogenetic analyses place ChHV in a clade with HSV1 and HSV2, the alphaherpesviruses of Old World monkeys comprising a separate clade. Analysis of reactivity patterns of HSV2-immune human sera and ChHV-immune chimpanzee sera by competition ELISA support this relationship. Phylogenetic analyses also place ChHV rather than HSV1 as the closest relative of HSV2.

Alphaherpesvirinae↗

Pleiotropic quantitative trait loci contribute to population divergence in traits associated with life-history variation in Mimulus guttatus.

Evolutionary biologists seek to understand the genetic basis for multivariate phenotypic divergence. We constructed an F2 mapping population (N = 539) between two distinct populations of Mimulus guttatus. We measured 20 floral, vegetative, and life-history characters on parents and F1 and F2 hybrids in a common garden experiment. We employed multitrait composite interval mapping to determine the number, effect, and degree of pleiotropy in quantitative trait loci (QTL) affecting divergence in floral, vegetative, and life-history characters. We detected 16 QTL affecting floral traits; 7 affecting vegetative traits; and 5 affecting selected floral, vegetative, and life-history traits. Floral and vegetative traits are clearly polygenic. We detected a few major QTL, with all remaining QTL of small effect. Most detected QTL are pleiotropic, implying that the evolutionary shift between these annual and perennial populations is constrained. We also compared the genetic architecture controlling floral trait divergence both within (our intraspecific study) and between species, on the basis of a previously published analysis of M. guttatus and M. nasutus. Eleven of our 16 floral QTL map to approximately the same location in the interspecific map based on shared, collinear markers, implying that there may be a shared genetic basis for floral divergence within and among species of Mimulus.

Flowers↗

Isolation of cDNA and genomic clones encoding human pro-alpha 1 (III) collagen. Partial characterization of the 3' end region of the gene.

A cDNA library constructed from human fibroblast poly(A+) RNA was screened for the identification of chimeric molecules bearing collagen-specific sequences. Analysis of three of the resulting positive clones showed that they encoded for the COOH-terminal propeptide region of the human Type III collagen. In addition, three overlapping clones covering more than 21 kilobases of the Type III gene were isolated from Charon 4A libraries of human genomic fragments. Identity between these and the cDNA clones was obtained by direct DNA sequencing. Establishment of the exon/intron arrangement of the Type III gene was obtained by electron microscopic analysis in conjunction with sequencing of selected genomic regions. Sequence comparison with other collagen genes confirmed some evolutionary features of this important family of proteins. Finally, the collinearity of two mRNA transcripts with 3' noncoding region length polymorphism was established.

Amino Acid Sequence↗

Statistical and deterministic approaches to designing transformations of electrocardiographic leads.

Two different approaches can be used to investigate the relationships among electrocardiographic leads: a statistical one, based on the analysis of recorded electrocardiograms (ECGs), and a deterministic one, based on physical principles that govern the current flow in irregularly shaped volume conductors such as the human body. The purpose of this study was to compare these two approaches. For the statistical investigation, the data set consisted of 120-lead ECGs recorded in a population including normal subjects (n = 290), post-myocardial-infarction patients (n = 497), patients with a history of ventricular tachycardia but no evidence of a previous myocardial infarction (n = 105), and patients with a single-vessel coronary artery disease who underwent coronary angioplasty (n = 91). Lead transformations of interest were obtained by fitting the multiple-regression model to this data set by the least-squares method. For the deterministic investigation, we used a boundary-element model of the human torso to simulate body-surface potentials in response to three orthogonal unit dipoles placed consecutively at 1,239 ventricular source locations, and the resulting body-surface potential distributions (instead of the recorded ECGs) were then fitted by the multiple-regression model. The results suggest that the lead transformations should be preferably designed by statistical analysis of recorded ECGs. Regression models with a small number of predictors (eg, those based on three ECG leads) are the most reliable; those using more predictors are fraught with the danger of collinearity when predictors are highly correlated (as occurs in the standard 12-lead ECG). Model-derived deterministic transformations are compatible with statistically derived ones, provided that the distributed character of the cardiac sources is taken into account. We conclude that statistical associations among electrocardiographic leads can be reliably quantified in sufficiently large and diverse databases of recorded data; the causality of these associations can be supported by appropriate deterministic models based on the laws of physics.

Angioplasty, Balloon, Coronary↗

Analyzing intramolecular dynamics by fast Lyapunov indicators.

We report an analysis of intramolecular dynamics of the highly excited planar carbonyl sulfide below and at the dissociation threshold by the fast Lyapunov indicator method. By mapping out the variety of dynamical regimes in the phase space of this molecule, we obtain the degree of regularity of the system versus its energy. We combine this stability analysis with a periodic orbit search, which yields a family of elliptic periodic orbits in the regular part of phase space and a family of in-phase collinear hyperbolic orbits associated with the chaotic regime.

Journal Article↗

Rice cDNAs as a model for expressed genes of plants.

Large-scale rice cDNA analysis has produced a huge amount of nucleotide sequence information for expressed genes in rice. The genes of cDNA clones putatively identified by similarity search were originally found in many different organisms. However, genes identified at a higher confidence level were found in plants, especially in monocots. This means the sequence information produced in random cloning of rice cDNA is useful for the study of other Gramineae. Further, assigned gene names of cDNAs mapped on linkage group 6 were grouped by their original species. The functions of gene products for 51% of mapped cDNAs were assigned and 67% of them were known in plants. The map information obtained by linking the position of a cDNA locus and its assigned gene function is indispensable for elucidating collinearity of genes among plant genomes.

Base Sequence↗

Primary structure of the alcelaphine herpesvirus 1 genome.

Alcelaphine herpesvirus 1 (AHV-1) causes wildebeest-associated malignant catarrhal fever, a lymphoproliferative syndrome in ungulate species other than the natural host. Based on biological properties and limited structural data, it has been classified as a member of the genus Rhadinovirus of the subfamily Gammaherpes-virinae. Here, we report on cloning and structural analysis of the complete genome of AHV-1 C500. The low GC content DNA (L-DNA) region of the genome consists of 130,608 bp with low (46.17%) GC content and marked suppression of CpG dinucleotide frequency. Like in herpesvirus saimiri, the prototype of the rhadinoviruses, the L-DNA is flanked by approximately 20 to 25 GC-rich (71.83%) high GC content DNA (H-DNA) repeats of 1,113 to 1,118 nucleotides. The analysis of the L-DNA sequence revealed 70 open reading frames (ORFs), 61 of which showed homology to other herpesviruses. The conserved ORFs are arranged in four blocks collinear to other Rhadinovirus genomes. These gene blocks are flanked by nonconserved regions containing ORFs without similarities to known herpesvirus genes. Notably, a spliced reading frame with a coding capacity for a 199-amino-acid protein is located in a position homologous to the transforming genes of herpesvirus saimiri at the left end of the L-DNA. A gene with homology to the semaphorin family is located adjacent to this. Despite common biological and epidemiological properties, AHV-1 differs significantly from herpesvirus saimiri with regard to cell homologous genes, probably using a different set of effector proteins to achieve a similar T-lymphocyte-transforming phenotype.

Amino Acid Sequence↗

Analysis of genetic marker-phenotype relationships by jack-knifed partial least squares regression (PLSR).

The utility of a relatively new multivariate method, bi-linear modelling by cross-validated partial least squares regression (PLSR), was investigated in the analysis of QTL. The distinguishing feature of PLSR is to reveal reliable covariance structures in data of different types with regard to the same set objects. Two matrices X (here: genetic markers) and Y (here: phenotypes) are interactively decomposed into latent variables (PLS components, or PCs) in a way which facilitates statistically reliable and graphically interpretable model building. Natural collinearities between input variables are utilized actively to stabilise the modelling, instead of being treated as a statistical problem. The importance of cross-validation/jack-knifing as an intuitively appealing way to avoid overfitting, is emphasized. Two datasets from chromosomal mapping studies of different complexity were chosen for illustration (QTL for tomato yield and for oat heading date). Results from PLSR analysis were compared to published results and to results using the package PLABQTL in these data sets. In all cases PLSR gave at least similar explained validation variances as the reported studies. An attractive feature is that PLSR allows the analysis of several traits/replicates in one analysis, and the direct visual identification of individuals with desirable marker genotypes. It is suggested that PLSR may be useful in structural and functional genomics and in marker assisted selection, particularly in cases with limited number of objects.

Crosses, Genetic↗

Advantages and limitations of metaanalytic regressions of clinical trials data.

OBJECTIVE: To focus on methodology of metaregression and demonstrate to clinicians, through 2 published examples, some strengths and limitations. EXAMPLE 1 METHODS: Metaanalysis of data from 20 years of randomized trials of lidocaine prophylaxis in preventing primary ventricular fibrillation (VF) in myocardial infarction used separate data for control and active-treatment groups to model the risk of VF as a function of year of publication and other study characteristics. EXAMPLE 1 RESULTS: Collinearity between pairs of predictor variables can lead to difficulty in interpreting logistic regression models. EXAMPLE 2 METHODS: Metaanalysis of data from 7 trials (323 patients) measured treatment effect of immunosuppressive therapy for acute Crohn's disease by response-rate difference (RD, experimental minus control group). EXAMPLE 2 RESULTS: Weighted least-squares regression models of the RD suggested an association between RD and various study characteristics, but collinearity again led to difficulty in interpretation of multiple regression results. DISCUSSION: Warning signs of colinearity include: large pairwise correlations between predictor variables, large changes in coefficients caused by the addition or deletion of other variables, and extremely large SEs for coefficients. Suggestions for coping with collinearity include: removing redundant variables from the model, reducing reliance on interpretation of coefficients for confounding variables, forming one or more summary variables, centering the data, collecting more data, and using more sophisticated regression methods. CONCLUSIONS: Metaanalysis can explore variations in as well as summarize results of randomized trials. Although metaregression has advantages, study characteristics are often strongly associated with each other, leading to collinearity.

Acute Disease↗

DNA sequence and transcriptional analysis of the glycoprotein M gene of murine cytomegalovirus.

We have characterized the gene encoding the murine cytomegalovirus (MCMV) homologue of the human cytomegalovirus (HCMV) UL100 open reading frame (ORF) that encodes the HCMV glycoprotein M (gM) molecule. It was identified based on its collinearity with MCMV homologues of the HCMV UL99, UL102, UL103 and UL104 ORFs which lie in the HindIII G fragment of the K181 strain of MCMV. Sequencing of a 2.3 kb EcoRI-BamHI subfragment of the EcoRI G fragment adjacent to the EcoRI A fragment revealed the presence of the complete MCMV gM ORF and two incomplete ORFs, which corresponded to homologues of HCMV UL99 and UL102. The MCMV gM ORF consists of 1059 nucleotides and is expressed as a 1.2 kb transcript at late times post-infection. To precisely characterize the gM transcript, the 5' and 3' ends were mapped. It was found that the transcript initiates at nucleotides 740 or 745, and that the site of polyadenylation at nucleotide 1961 occurs downstream of the second potential polyadenylation signal located at nucleotide 1934. Based on these findings the MCMV gM is predicted to consist of 353 residues and when compared with HCMV gM has a 47% level of identity. Of great interest is the finding that the MCMV gM amino acid sequence is completely conserved among six isolates of MCMV that had been shown to exhibit considerable variation both in the MCMV glycoprotein B and the immediate-early 1 gene-encoded pp89 molecule. Thus, this glycoprotein appears to be antigenically conserved.

Amino Acid Sequence↗

Use of psychometric techniques in the analysis of epidemiologic data.

PURPOSE: This article demonstrates techniques for developing reliable multi-item scales for analysis of complex public health data. METHODS: Information from a questionnaire designed to evaluate the acceptability and efficacy of the female condom as a method for STD/HIV prevention was summarized using psychometric analysis. 1159 high-risk women attending STD clinics participated in this study. Questionnaire items were designed to measure nine domains of predictors of condom use. RESULTS: Principal components analysis was employed to reduce the number of potential predictors. Reliability of the multiple-item scales was assessed using Cronbach's alpha. Pearson's correlation coefficients were calculated to evaluate collinearity among multi-item scales. Approximately half (51%) of the questionnaire items that were analyzed were retained in the final scales. Data reduction procedures identified several multi-item scales with acceptable reliability (Cronbach's alpha >0.70). The correlation coefficients between scales was never >.5, suggesting that there was little collinearity among the scales. CONCLUSIONS: When focused on multiple partially interdependent determinants of an outcome, data reduction decreases the number of independent variables to be evaluated, ensures they have adequate reliability, maximizes strength of their association with outcomes, and reduces collinearity among predictors.

Adolescent↗

Strong phylogenetic signal from chloroplast genomes of three Barringtonia species provides the first genomic resources for their conservation.

BACKGROUND: The genus Barringtonia (Lecythidaceae) is a vital component of tropical coastal forests and mangrove ecosystems. Among its members, B. racemosa and B. fusicarpa are classified as Endangered and Vulnerable, respectively, due to habitat degradation and anthropogenic pressures, underscoring the urgent need for genetic studies to guide conservation. Chloroplast (cp.) genomes serve as essential resources for phylogenetic reconstruction and conservation genetics. However, the scarcity of cp. genome data for Barringtonia has limited comprehensive evolutionary and conservation-oriented investigations. RESULTS: We assembled and annotated the first complete cp. genomes of B. racemosa, B. fusicarpa, and B. acutangula. All three genomes exhibit the typical quadripartite structure, ranging from 158,959 bp (B. racemosa) to 159,837 bp (B. acutangula), and contain 132 genes (87 protein-coding, 37 tRNA, 8 rRNA) with a GC content of 36.68%-36.86%. Collinearity and IR boundary analyses revealed high structural conservation without large-scale rearrangements. Interspecific sequence-level variations were detected in simple sequence repeats (SSRs) and long repeats. Nucleotide diversity (π) analysis identified highly polymorphic regions, including rpl20 (π = 0.080), rpoA (π = 0.064), rps3 (π = 0.063), and ndhF (π = 0.060), which represent promising molecular markers for population genetics within the genus. Codon-based selection analyses (Ka/Ks) showed that all protein-coding genes are under strong purifying selection (mean Ka/Ks 0.32-0.37), with no evidence of positive selection. Pairwise genetic distances (p-distances) among Barringtonia species are extremely low (mean 0.0046), while distances to the related genus Bertholletia are ~ 6-fold higher, supporting their generic distinction. CONCLUSIONS: Phylogenetic analysis robustly supports Barringtonia as a monophyletic clade (bootstrap = 100%), with B. racemosa and B. fusicarpa forming a sister lineage to B. acutangula. This study provides the first high-quality cp. genome resources for the two threatened Barringtonia species, revealing strong structural and sequence conservation but no direct chloroplast genomic correlates of endangerment. The identified polymorphic regions and repeat markers lay a foundation for future population genetics, phylogeographic studies, and conservation-oriented genetic management of these ecologically important coastal plants.

Genome, Chloroplast↗

Statistical estimation of parameters in a disease transmission model: analysis of a Cryptosporidium outbreak.

Population dynamic models, commonly used tools in the study of epidemics and other complex population processes, are implicit non-linear mathematical equations. Inference based on such models can be difficult due to the problems associated with high dimensional parameters that may be non-identified and complex likelihood functions that are difficult to maximize. To address a problem of non-identifiability due to collinearity of parameter estimates in a mathematical model of the 1993 Milwaukee Cryptosporidium parvum outbreak, we examined the utility of a constrained profile likelihood approach. This method was used to study two parameters of interest from the mathematical model: (i). the rate of secondary transmission; (ii). the proportional increase in primary transmission due to water treatment failure. The estimated values of these parameters were shown to depend strongly on poorly understood aspects of Cryptosporidium epidemiology such as asymptomatic proportion and the population immune status. Our analysis demonstrated that the combination of a disease transmission model and a constrained profile likelihood procedure provides an effective approach for inference and estimation of important parameters regulating infectious disease outbreaks.

Animals↗

An evaluation of proposed frameworks for grouping polychlorinated biphenyl (PCB) congener data into meaningful analytic units.

BACKGROUND: Polychlorinated biphenyls (PCBs) have been associated with a variety of health outcomes. Enhanced laboratory techniques can provide a relatively large number of individual PCB congeners for investigation. However, to date there are no established frameworks for grouping a large number of PCB congeners into meaningful analytic units. METHODS: In a case-control study of serum PCB levels on breast cancer risk, measured levels of 56 PCB congener peaks were available for analysis. We considered several approaches for grouping these compounds based on 1) chlorination, 2) factor analysis, 3) enzyme induction, 4) enzyme induction and occurrence, and 5) enzyme induction, occurrence, and other toxicological aspects. The utility of a framework was based on the mechanism of biologic actions within each framework, lack of collinearity among congener groups, and frequency of detection of PCB congener groups in measured serum levels of 192 healthy postmenopausal women. RESULTS: Most participants had detectable levels for the proposed PCB congeners groups, using degree of chlorination as a grouping framework. In addition, the previously proposed grouping approach based on enzyme induction, occurrence, and other toxicological aspects was an applicable alternative to the crude approach of grouping by degree of chlorination. Grouping these congeners with respect to P450 enzyme induction activity, and the previously proposed framework based on enzyme induction and occurrence, did not fit these data as well, because only a small proportion of participants had detectable levels for the congener groups with the greatest toxicological potential. Statistical grouping did not result in an interpretable and meaningful clustering of these exposures. CONCLUSIONS: In these data, grouping with respect to degree of chlorination and the previously proposed framework based on enzyme induction, occurrence, and other toxicological aspects were the most useful approaches to reducing a large number of PCB congeners into meaningful analytic units. Factors affecting the utility of the proposed grouping frameworks are discussed.

Aged↗