PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “collinearity analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Excision from tRNA genes of a large chromosomal region, carrying avrPphB, associated with race change in the bean pathogen, Pseudomonas syringae pv. phaseolicola.

Pseudomonas syringae pv. phaseolicola (Pph) race 4 strain 1302A carries avirulence gene avrPphB. Strain RJ3, a sectoral variant from a 1302A culture, exhibited an extended host range in cultivars of bean and soybean resulting from the absence of avrPphB from the RJ3 chromosome. Complementation of RJ3 with avrPphB restored the race 4 phenotype. Both strains showed similar in planta growth in susceptible bean cultivars. Analysis of RJ3 indicated loss of > 40 kb of DNA surrounding avrPphB. Collinearity of the two genomes was determined for the left and right junctions of the deleted avrPphB region; the left junction is approximately 19 kb and the right junction > 20 kb from avrPphB in 1302A. Sequencing revealed that the region containing avrPphB was inserted into a tRNALYS gene, which was re-formed at the right junction in strain 1302A. A putative lysine tRNA pseudogene (PsitRNALYS) was found at the left junction of the insertion. All tRNA genes were in identical orientation in the chromosome. Genes near the left junction exhibited predicted protein homologies with gene products associated with a virulence locus of the periodontal pathogen Actinobacillus actinomycetemcomitans. Specific oligonucleotide primers that differentiate 1302A from RJ3 were designed and used to demonstrate that avrPphB was located in different regions of the chromosome in other strains of Pph. Deletion of a large region of the chromosome containing an avirulence gene represents a new route to race change in Pph.

Base Sequence↗

Prognosis for Wilms' tumor patients with nonmetastatic disease at diagnosis--results of the second National Wilms' Tumor Study.

Multivariate statistical methods were used to study prognosis for 632 patients entered on the second National Wilms' Tumor Study who had nonmetastatic, unilateral disease at diagnosis. Separate analyses were conducted for each of four endpoints: abdominal recurrence, distant metastasis, relapse without regard to site, and death. The two most important predictors for metastasis and general relapse were an unfavorable (anaplastic or sarcomatous) histology and the presence of microscopically confirmed disease in the regional lymph nodes. Operative spillage of tumor increased the rates of abdominal recurrence and death, even after accounting for histology and lymph node effects. The presence of a tumor thrombus in the renal vein or IVC increased the risk of metastasis, and intrarenal vascular invasion was associated with general relapse after accounting for histology, lymph nodes, and spillage. However, these latter associations were weaker, and some uncertainty remains regarding the true prognostic import of such findings due to a high degree of collinearity among variables. By contrast to the results of a similar data analysis for the first National Wilms' Tumor Study, specimen weight had no bearing on outcome, and the effects of age at diagnosis were entirely explained by the association of age with other more critical factors.

Adolescent↗

An appraisal of multivariable logistic models in the pulmonary and critical care literature.

OBJECTIVE: Multivariable modeling techniques are appearing in today's medical literature with increasing frequency. Improper reporting of these statistical models can potentially make the results of a study inaccurate, misleading, or difficult to interpret. We performed a manual literature search of five international pulmonary and critical care journals to determine the accuracy in the reporting of logistic regression modeling strategies. DESIGN: We examined all of the published manuscripts for 12 potential limitations in the reporting of important statistical methodologies over a 6-month period from July 1, 2000, until December 31, 2000. RESULTS: Of the 81 articles that included multivariable logistic regression analyses, only 65% (53 analyses) properly reported the coding classification of pertinent independent variables that were included in the final model. An odds ratio and confidence interval were reported for the independent variables included in the final model for 79% (64 analyses) and 74% (60 analyses), respectively. Only 12% (10 articles) referenced whether interaction terms or effect modifications were examined, 1% (1 article) reported testing for collinearity, and only 16% (13 articles) included a goodness-of-fit analysis of the logistic model. The type of statistical package was reported in 69% (56 articles). Finally, approximately 39% of the articles (22 of 57) may have overfit the logistic regression model, leading to potentially unreliable regression coefficients and odds ratios. CONCLUSIONS: Our results indicate that the reporting of multivariable logistic regression analyses in the pulmonary and critical care literature is often incomplete, therefore making it difficult for the reader to accurately interpret the manuscript. We recommend the implementation of adequate guidelines that will lead to overall improvements in the reporting and possibly to the conducting of multivariable analyses in the pulmonary medicine and critical care medicine literature.

Critical Care↗

Nonlinear modeling and adaptive monitoring with fuzzy and multivariate statistical methods in biological wastewater treatment plants.

A new approach to nonlinear modeling and adaptive monitoring using fuzzy principal component regression (FPCR) is proposed and then applied to a real wastewater treatment plant (WWTP) data set. First, principal component analysis (PCA) is used to reduce the dimensionality of data and to remove collinearity. Second, the adaptive credibilistic fuzzy-c-means method is used to appropriately monitor diverse operating conditions based on the PCA score values. Then a new adaptive discrimination monitoring method is proposed to distinguish between a large process change and a simple fault. Third, a FPCR method is proposed, where the Takagi-Sugeno-Kang (TSK) fuzzy model is employed to model the relation between the PCA score values and the target output to avoid the over-fitting problem with original variables. Here, the rule bases, the centers and the widths of TSK fuzzy model are found by heuristic methods. The proposed FPCR method is applied to predict the output variable, the reduction of chemical oxygen demand in the full-scale WWTP. The result shows that it has the ability to model the nonlinear process and multiple operating conditions and is able to identify various operating regions and discriminate between a sustained fault and a simple fault (or abnormalities) occurring within the process data.

Algorithms↗

Characterization of a putative Spodoptera exigua multicapsid nucleopolyhedrovirus helicase gene.

Putative baculovirus helicases have been implicated as playing an important role in viral DNA replication and host specificity. The Spodoptera exigua multicapsid nucleopolyhedrovirus (SeMNPV) helicase is therefore of interest since the virus only infects the beet army worm. Sequence analysis of the SeMNPV lef5-p39 (mu 46.5-55.1) region, which is collinear with the 39K-lef5 area in Autographa californica MNPV (AcMNPV), revealed an open reading frame (ORF) of 3666 bp potentially encoding a protein with a molecular mass of 143 kDa. This protein had considerable amino acid sequence similarity (58%) to AcMNPV p143, including seven conserved motifs characteristic of helicases. In cultured insect cells, this SeMNPV ORF is expressed from 4 to 12 h postinfection and its major transcript of 4 kb starts 11 to 12 nt upstream of the putative translational initiation site (ATG). To study their possible role in the specificity of baculovirus DNA replication, the putative AcMNPV and SeMNPV helicase genes were tested for their ability to replicate homologous regions (hrs; putative origins of DNA replication) in a transient DNA replication assay in insect cells. All viral cis- and trans-acting factors were provided as plasmids using either Achr2 or Sehr1 as the DNA replication origin. SeMNPV p143 could not substitute for AcMNPV p143 in the transient assays supplemented with either hr. Similar results were obtained when the SeMNPV and AcMNPV ie1 genes were exchanged. None of the essential AcMNPV trans-acting factors could be complemented by SeMNPV infections to support DNA replication of hrs. These data suggest a specific interaction between baculovirus DNA replication factors to form the replisome and/or between the replisome and the origin of DNA replication.

Amino Acid Sequence↗

Predictors of dropout and remission in family therapy for adolescent anorexia nervosa in a randomized clinical trial.

OBJECTIVE: The purpose of this study is to explore the predictors of dropout and remission in the treatment of adolescent anorexia nervosa (AN) using family therapy. METHOD: Data derived from a randomized clinical trial comparing short and long term family therapy for adolescents with AN were used. A rotated component analysis was employed to reduce the number of variables and to address problems of collinearity and multiple testing. Dropout was defined as participating in less than 80% of the assigned therapy. Participants were classified as remitted if they obtained an ideal body weight greater than 95% and a global eating disorder Examination score within two standard deviations of community norms at the end of 12 months. RESULTS: Co-morbid psychiatric disorder and being randomized to longer treatment predicted greater dropout. The presence of co-morbid psychiatric disorder, being older, and problematic family behaviors led to lower rates of remission. A reduction of child behavioral symptoms, a decline in problematic family behaviors, and early weight gain were all within treatment changes that increased the chance of remission. CONCLUSION: Co-morbid psychiatric disorder, family behaviors, and early response to treatment are important factors when predicting dropout and remission in family therapy for adolescent AN.

Adolescent↗

Spin-polarized scanning tunneling microscopy: insight into magnetism from nanostructures to atomic scale spin structures.

The system of Fe on W(001) is investigated using spin-integrated as well as spin-resolved scanning tunneling microscopy (STM). This study ranges from three-dimensional Fe islands down to the Fe monolayer and different growth modes are observed related to the preparation temperature. With scanning tunneling spectroscopy (STS), a layer-dependent electronic structure is observed that can easily be used to assign the local coverage to the investigated sample areas. Spin-resolved measurements of the ferromagnetic layers in the pseudomorphic regime immediately reveal the fourfold magnetic in-plane anisotropy. A direct comparison of the observed arrangement of the domains of the exposed layers shows a rotation of the easy axis from the fourth to the third monolayer and a collinear magnetic alignment of third and second monolayer. This is confirmed by the quantitative analysis of the layer-resolved intensities of differential tunneling conductance. The first monolayer does not show a magnetic component parallel to the surface but has a perpendicular anisotropy. For this layer, measurements with an applied magnetic field prove a c(2x2) antiferromagnetic structure, i.e., a checkerboard arrangement of spins.

Anisotropy↗

D matrix analysis of the Renner-Teller effect: an accurate three-state diabatization for NH2.

Some time ago we published our first article on the Renner-Teller (RT) model to treat the electronic interaction for a triatomic molecule [J. Chem. Phys. 124, 081106 (2006)]. The main purpose of that Communication was to suggest considering the RT phenomenon as a topological effect, just like the Jahn-Teller phenomenon. However, whereas in the first publication we just summarized a few basic features to support that idea, here in the present article, we extend the topological approach and show that all the expected features that characterize a three (multi) state RT-type'3 system of a triatomic molecule can be studied and analyzed within the framework of that approach. This, among other things, enables us to employ the topological D matrix [Phys. Rev. A 62, 032506 (2000)] to determine, a priori, under what conditions a three-state system can be diabatized. The theoretical presentation is accompanied by a detailed numerical study as carried out for the HNH system. The D-matrix analysis shows that the two original electronic states 2A1 and 2B1 (evolving from the collinear degenerate Pi doublet), frequently used to study this Renner-Teller-type system, are insufficient for diabatization. This is true, in particular, for the stable ground-state configurations of the HNH molecule. However, by including just one additional electronic state--a B state (originating from a collinear Sigma state)--it is found that a rigorous, meaningful three-state diabatization can be carried out for large regions of configuration space, particularly for those, near the stable configuration of NH2. This opens the way for an accurate study of this important molecule even where the electronic angular momentum deviates significantly from an integer value.

Journal Article↗

Nucleotide and predicted amino acid sequences of Marek's disease virus homologues of herpes simplex virus major tegument proteins.

The DNA sequence of an 8.4 kbp BamHI-EcoRI fragment of Marek's disease virus (MDV) strain GA was determined. Three of the predicted polypeptides are homologous to UL47, UL48 and UL49 encoding the major tegument proteins of herpes simplex virus type 1 (HSV-1), and four are homologous to HSV-1 UL45, UL46, UL49.5 and UL50. These seven genes are found in the long unique region of the MDV genome and are collinear with homologues in HSV-1 and varicella-zoster virus (VZV). Northern blot analysis revealed different transcriptional patterns from those of HSV-1 and VZV. MDV homologues of UL49.5, UL49 and UL47 lack a poly(A) signal immediately downstream of their coding regions. Amino acid conservation between MDV and HSV-1, and between MDV and VZV is as high as that between HSV-1 and VZV. The MDV homologue of UL48 shows 60% similarity to its HSV-1 counterpart. Amino acid sequence comparison reveals that the MDV homologue of UL48 lacks an acidic carboxyl terminus. This homologue, like the VZV homologue of UL48, may be involved in the trans-activation of immediate early genes and may function as an important component of the structural proteins.

Amino Acid Sequence↗

Murine herpesvirus 68 is genetically related to the gammaherpesviruses Epstein-Barr virus and herpesvirus saimiri.

Short nucleotide sequence analysis of seven restriction fragments of murine herpesvirus 68 (MHV-68) DNA has been undertaken and used to determine the overall genome organization and relatedness of this virus to other well characterized representatives of the alpha-, beta- and gammaherpesvirus subgroups. Nine genes have been identified which encode amino acid sequences with greater similarity to proteins of the gammaherpesvirus Epstein-Barr virus (EBV) than to the homologous products of the alphaherpesviruses varicella-zoster virus and herpes simples virus type 1 or the betaherpesvirus human cytomegalovirus. In addition, the genome organization of MHV-68 is shown to have an overall collinearity with that of the gammaherpesviruses EBV and herpesvirus saimiri. In common with these viruses, dinucleotide frequency analysis of MHV-68 coding sequences reveals a marked reduction in CpG dinucleotide frequency thus implicating a dividing cell population as the site of latency in vivo.

Amino Acid Sequence↗

Problems of correlations between explanatory variables in multiple regression analyses in the dental literature.

Multivariable analysis is a widely used statistical methodology for investigating associations amongst clinical variables. However, the problems of collinearity and multicollinearity, which can give rise to spurious results, have in the past frequently been disregarded in dental research. This article illustrates and explains the problems which may be encountered, in the hope of increasing awareness and understanding of these issues, thereby improving the quality of the statistical analyses undertaken in dental research. Three examples from different clinical dental specialties are used to demonstrate how to diagnose the problem of collinearity/multicollinearity in multiple regression analyses and to illustrate how collinearity/multicollinearity can seriously distort the model development process. Lack of awareness of these problems can give rise to misleading results and erroneous interpretations. Multivariable analysis is a useful tool for dental research, though only if its users thoroughly understand the assumptions and limitations of these methods. It would benefit evidence-based dentistry enormously if researchers were more aware of both the complexities involved in multiple regression when using these methods and of the need for expert statistical consultation in developing study design and selecting appropriate statistical methodologies.

Data Interpretation, Statistical↗

Estimation and inference in pharmacokinetic models: the effectiveness of model reformulation and resampling methods for functions of parameters.

It is well known that high parameter estimate correlations and asymptotic variance estimates can cause estimation and inference problems in the analysis of pharmacokinetic models. In this paper we show that analysis of three important functions of pharmacokinetic parameters, the half-life, mean residence time, and the area under the curve, can sometimes be greatly improved by reformulating the model to address collinearity and by using the bootstrap to form confidence intervals. The resultant estimators can be more accurate than the original ones, and resultant confidence intervals can be narrower. Of the three measures, the half-life estimator is much better behaved than the estimators of mean residence time and area under the curve under collinearity, suggesting that it (or measures like it) should be used more often.

Analysis of Variance↗

Prediction of the concentration of chlorophyll-a for Liuhai urban lakes in Beijing City.

The weekly water quality monitor data of Liuhai lakes between April 2003 and November 2004 in Beijing City were used as an example to build an artificial neural networks (ANN) model and a multi-varieties regression model respectively for predicting the fresh water algae bloom. The different predicted abilities of the two methods in Liuhai lakes were compared. A principle analysis method was first used to select the input variables of the models to avoid the phenomenon of collinearity in the data. The results showed that the input variables for the artificial neural networks were T, TP, transparency(SD), DO, chlorophyll-a (Chl-a), pH and the output variable was Chl-a. A three layer Levenberg-Marguardt feed forward learning algorithm in ANN was used to model the eutrophication process of Liuhai lakes. 20 nodes in hidden layer and 1 node of output for the ANN model had been optimized by trial and error method. A sensitivity analysis of the input variables was performed to evaluate their relative significance in determining the predicted values. The correlation coefficient between predicted value and observed value in all data and in test data were 0.717 and 0.816 respectively in the artificial neural networks. The stepwise regression method was used to simulate the linear relation between Chl-a and temperature, of which the correlation coefficient was 0.213. By comparing the results of the two models, it was found that neural network models were able to simulate non-linear behavior in the water eutrophication process of Liuhai lakes reasonably and could successfully estimate some extreme values from calibration and test data sets.

China↗

Genome sequence of an enhancin gene-rich nucleopolyhedrovirus (NPV) from Agrotis segetum: collinearity with Spodoptera exigua multiple NPV.

The genome sequence of a Polish isolate of Agrotis segetum nucleopolyhedrovirus (AgseNPV-A) was determined and analysed. The circular genome is composed of 147,544 bp and has a G+C content of 45.7 mol%. It contains 153 putative, non-overlapping open reading frames (ORFs) encoding predicted proteins of more than 50 aa, together making up 89.8 % of the genome. The remaining 10.2 % of the DNA constitutes non-coding regions and homologous-repeat regions. One hundred and forty-three AgseNPV-A ORFs are homologues of previously reported baculovirus gene sequences. There are ten unique ORFs and they account for 3 % of the genome in total. All 62 lepidopteran baculovirus genes, including the 29 core baculovirus genes, were found in the AgseNPV-A genome. The gene content and gene order of AgseNPV-A are most similar to those of Spodoptera exigua (Se) multiple NPV and their shared homologous genes are 100 % collinear. Three putative enhancin genes were identified in the AgseNPV-A genome. In phylogenetic analysis, the AgseNPV-A enhancins form a cluster separated from enhancins of the Mamestra species NPVs.

Animals↗

Diagnosing and dealing with multicollinearity.

The purpose of this article was to increase nurse researchers' awareness of the effects of collinear data in developing theoretical models for nursing practice. Collinear data distort the true value of the estimates generated from ordinary least-squares analysis. Theoretical models developed to provide the underpinnings of nursing practice need not be abandoned, however, because they fail to produce consistent estimates over repeated applications. It is also important to realize that multicollinearity is a data problem, not a problem associated with misspecification of a theorectical model. An investigator must first be aware of the problem, and then it is possible to develop an educated solution based on the degree of multicollinearity, theoretical considerations, and sources of error associated with alternative, biased, least-square regression techniques. Decisions based on theoretical and statistical considerations will further the development of theory-based nursing practice.

Humans↗

inGeno--an integrated genome and ortholog viewer for improved genome to genome comparisons.

BACKGROUND: Systematic genome comparisons are an important tool to reveal gene functions, pathogenic features, metabolic pathways and genome evolution in the era of post-genomics. Furthermore, such comparisons provide important clues for vaccines and drug development. Existing genome comparison software often lacks accurate information on orthologs, the function of similar genes identified and genome-wide reports and lists on specific functions. All these features and further analyses are provided here in the context of a modular software tool "inGeno" written in Java with Biojava subroutines. RESULTS: InGeno provides a user-friendly interactive visualization platform for sequence comparisons (comprehensive reciprocal protein--protein comparisons) between complete genome sequences and all associated annotations and features. The comparison data can be acquired from several different sequence analysis programs in flexible formats. Automatic dot-plot analysis includes output reduction, filtering, ortholog testing and linear regression, followed by smart clustering (local collinear blocks; LCBs) to reveal similar genome regions. Further, the system provides genome alignment and visualization editor, collinear relationships and strain-specific islands. Specific annotations and functions are parsed, recognized, clustered, logically concatenated and visualized and summarized in reports. CONCLUSION: As shown in this study, inGeno can be applied to study and compare in particular prokaryotic genomes against each other (gram positive and negative as well as close and more distantly related species) and has been proven to be sensitive and accurate. This modular software is user-friendly and easily accommodates new routines to meet specific user-defined requirements.

Base Sequence↗

Molecular and Physiological Insights into CAT- and SOD-Associated Redox Homeostasis Under Salt Stress in Artemisia argyi.

Soil salinity disrupts redox homeostasis and limits plant growth and development. Although catalase (CAT) and superoxide dismutase (SOD) are key enzymatic antioxidants, the CAT and SOD gene families have not been characterized in Artemisia argyi (A. argyi), a species of medicinal and ecological importance. While SOD and CAT serve as the primary enzymatic scavengers for reactive oxygen species (ROS) detoxification, their genomic architecture and stress-responsive regulatory networks in A. argyi have remained uncharacterized. In this study, we conducted the first comprehensive genome-wide analysis of these gene families in A. argyi, identifying 22 structurally conserved members (8 AarCATs and 14 AarSODs). Collinearity and synteny analyses revealed strict lineage-specific evolutionary conservation, while tertiary protein modeling and subcellular localization illustrated a highly organized multi-organelle defense compartmentalization. High salinity (up to 200 mM NaCl) reduced the stomatal conductance and net photosynthetic rate. Salt stress reduced growth and increased osmoprotectant and antioxidant accumulation in A. argyi. Furthermore, histochemical staining using nitroblue tetrazolium (NBT) and 3,3'-Diaminobenzidine (DAB) provided comprehensive evidence of significant accumulation of ROS in leaves, which indicates the intense oxidative stress triggered by ionic stress. Tissue-specific analysis revealed that AarCAT1, AarCSD1, and AarFSD2 were 3.9-, 7.9-, and 12.7-fold higher in leaves than in roots, respectively. Under stress, AarCAT6 and AarCSD1 were strongly repressed in leaves by ~50% and ~46-70%, respectively, whereas AarMSD2 and AarMSD3 were significantly induced in roots by ~2.2- and ~1.8-fold. These distinct expression patterns suggest their potential involvement in tissue-specific stress adaptation and ROS homeostasis. These findings uncover the evolutionary and physiological basis of salt tolerance in A. argyi, providing genetic targets for climate-resilient breeding.

Artemisia↗

Long-range comparison of human and mouse SCL loci: localized regions of sensitivity to restriction endonucleases correspond precisely with peaks of conserved noncoding sequences.

Long-range comparative sequence analysis provides a powerful strategy for identifying conserved regulatory elements. The stem cell leukemia (SCL) gene encodes a bHLH transcription factor with a pivotal role in hemopoiesis and vasculogenesis, and it displays a highly conserved expression pattern. We present here a detailed sequence comparison of 193 kb of the human SCL locus to 234 kb of the mouse SCL locus. Four new genes have been identified together with an ancient mitochondrial insertion in the human locus. The SCL gene is flanked upstream by the SIL gene and downstream by the MAP17 gene in both species, but the gene order is not collinear downstream from MAP17. To facilitate rapid identification of candidate regulatory elements, we have developed a new sequence analysis tool (SynPlot) that automates the graphical display of large-scale sequence alignments. Unlike existing programs, SynPlot can display the locus features of more than one sequence, thereby indicating the position of homology peaks relative to the structure of all sequences in the alignment. In addition, high-resolution analysis of the chromatin structure of the mouse SCL gene permitted the accurate positioning of localized zones accessible to restriction endonucleases. Zones known to be associated with functional regulatory regions were found to correspond precisely with peaks of human/mouse homology, thus demonstrating that long-range human/mouse sequence comparisons allow accurate prediction of the extent of accessible DNA associated with active regulatory regions.

Animals↗