PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

The three-dimensional structure of human transaldolase.

The crystal structure of human transaldolase has been determined to 2.45 A resolution. The enzyme folds into an alpha/beta barrel structure and is thus similar in structure to other class I aldolases. Structure-based sequence alignment of available sequences of the transaldolase subfamily reveals that eight active site residues are invariant in the whole subfamily. Other invariant residues are mainly involved in the formation of the hydrophobic core of the enzyme. Noteworthy is a hydrophobic cluster consisting of five invariant residues. Human transaldolase has been implicated as an autoantigen in multiple sclerosis and four immunodominant peptide segments are located at the surface of the enzyme, accessible to autoantibodies.

Amino Acid Sequence↗

Sequence heterogeneity of the small subunit ribosomal RNA genes among blastocystis isolates.

Genes encoding small subunit ribosomal RNA (SSUrRNA) of 16 Blastocystis isolates from humans and other animals were amplified by the polymerase chain reaction, and the corresponding fragments were cloned and sequenced. Alignment of these sequences with the previously reported ones indicated the presence of 7 different sequence patterns in the highly variable regions of the small subunit ribosomal RNA. Phylogenetic reconstruction analysis using Proteromonas lacertae as the outgroup clearly demonstrated that the 7 groups with the different sequence patterns are separated to form independent clades, 5 of which consisted of the Blastocystis isolates from both humans (B. hominis) and other animals. The presence of 3 higher order clades was also clearly supported in the phylogenetic tree. However, a relationship among the 4 groups including these 3 higher order clades was not settled with statistical confidence. The remarkable heterogeneity of small subunit ribosomal RNAs among different Blastocystis isolates found in this study confirmed, with sequence-based evidence, that these organisms are genetically highly divergent in spite of their morphological identity. The highly variable small subunit ribosomal RNA regions among the distinct groups will provide useful information for the development of group-specific diagnostic primers.

Animals↗

Signature of quaternary structure in the sequences of legume lectins.

Legume lectins exhibit a wide variety of oligomerization and sugar specificity while retaining the characteristic jelly-roll tertiary fold. An attempt has been made here to find whether this diversity is reflected in their primary structures by constructing phylogenetic trees. Dendrograms based on sequence alignment showed clustering related to the oligomeric nature of legume lectins. Though the clustering primarily follows the oligomeric states, it also appears to correlate with different sugar specificities indicating an interdependence of these two properties. Analysis of the structure-based alignment and the alignment of the sequences of the carbohydrate-binding loops alone also revealed the same features. By a close examination of the interfaces of the various oligomers it was also possible, in some cases, to pinpoint a few key residues responsible for the stabilization of the interfaces.

Amino Acid Sequence↗

Characterisation of the QM gene of Trypanosoma brucei.

The QM protein has been reported to have roles in both tumour suppression and transcription factor regulation in vertebrate cells, and in ribosome stability in both yeast and mammals. The present study isolated the QM gene of Trypanosoma brucei and determined its sequence. Alignment with QM sequences from Saccharomyces cerevisiae, Arabidopsis thaliana, Drosophila melanogaster and Homo sapiens revealed greater than 60% identity. Southern blot analysis revealed multiple copies of QM within the trypanosome genome. An epitope tag was inserted into the C-terminus of the T. brucei QM and the protein expressed under inducible control in procyclic form trypanosomes. Immune fluorescence microscopy revealed co-localisation with the GPI:protein transamidase component, GPI8, a distribution indicative of ribosome association with the rough endoplasmic reticulum.

Amino Acid Sequence↗

Improved diagnostic PCR assay for Actinobacillus pleuropneumoniae based on the nucleotide sequence of an outer membrane lipoprotein.

The gene (omlA) coding for an outer membrane protein of Actinobacillus pleuropneumoniae serotypes 1 and 5 has been described earlier and has formed the basis for development of a specific PCR assay. The corresponding regions of all 12 A. pleuropneumoniae reference strains of biovar 1 were sequenced. Alignment of the sequences revealed conserved terminal and variable middle regions, which divided the reference strains into four distinct groups. Primers were selected from the conserved 5' and 3' termini of the gene. A 950-bp amplicon was obtained from each of 102 tested field isolates of A. pleuropneumoniae obtained from lungs. Their identity was verified by sequencing approximately 500 bp of the amplification product from 50 of the A. pleuropneumoniae isolates, which all showed the expected DNA sequence characteristic of the serotype. To test the specificity of the reaction, 23 other bacterial species related to A. pleuropneumoniae or isolated from pigs were assayed. They were all found negative in the PCR, as were tonsil cultures from 50 pigs of an A. pleuropneumoniae-negative herd. The sensitivity assessed by agarose gel analysis of the PCR product was 10(2) CFU/PCR test tube. The specificity and sensitivity of this PCR compared to those of culture suggest the use of this PCR for routine identification of A. pleuropneumoniae.

Actinobacillus Infections↗

Evidence of three new members of malignant catarrhal fever virus group in muskox (Ovibos moschatus), Nubian ibex (Capra nubiana), and gemsbok (Oryx gazella).

Six members of the malignant catarrhal fever (MCF) virus group of ruminant rhadinoviruses have been identified to date. Four of these viruses are clearly associated with clinical disease: alcelaphine herpesvirus 1 (AlHV-1) carried by wildebeest (Connochaetes spp.); ovine herpesvirus 2 (OvHV-2), ubiquitous in domestic sheep; caprine herpesvirus 2 (CpHV-2), endemic in domestic goats; and the virus of unknown origin found causing classic MCF in white-tailed deer (Odocoileus virginianus; MCFV-WTD). Using serology and polymerase chain reaction with (degenerate primers targeting a portion of the herpesviral DNA polymerase gene, evidence of three previously unrecognized rhadinoviruses in the MCF virus group was found in muskox (Ovibos moschatus), Nubian ibex (Capra nubiana), and gemsbok (South African oryx, Oryx gazella), respectively. Base on sequence alignment, the viral sequence in the muskox is most closely related to MCFV-WTD (81.5% sequence identity) and that in the Nubian ibex is closest to CpHV-2 (89.3% identity). The viral sequence in the gemsbok is most closely related to AlHV-1 (85.1% identity). No evidence of disease association with these viruses has been found.

Amino Acid Sequence↗

Characterization of GBV-C infection in HIV-1 infected patients.

BACKGROUND: GB virus C, a positive-stranded RNA virus, is classified in the family Flaviviridae. It is currently believed that persistent infection occurs in 25-50% of infected individuals, however, it still remains an "orphan" virus in search of a role in human pathology. Molecular epidemiological studies have demonstrated that GBV-C infection is present in about 1-1.4% of the healthy population in developed countries, that it shares routes of transmission with HIV and HCV and that the prevalence of GBV-C in these populations is higher than in blood donors. On the basis of the sequence variation among the isolates, GBV-C is classified into at least four major genotypes. Preliminary evidence has suggested that GBV-C is a lymphotropic virus that replicates mainly in the spleen and bone marrow. Recently, several reports have investigated the possible beneficial effect of GBV-C co-infection on HIV disease progression to AIDS, reduced mortality in HIV infected individuals and lower HIV viral loads, not leading to a definitive conclusion yet. AIM: To investigate the role of GBV virus C co-infection in two different subsets of HIV-infected patients, and to evaluate the prevalence of GBV-C genotypes in Northern Italy. METHODS: A total of 86 HIV positive patients were examined for GBV-C viremia (years after HIV sera conversion: 12 +/- 5). Control population (Group A): 46 patients (mean age 42 years) with <200CD4/ml during the observation period. Longterm non progressor population (Group B): 40 patients, (mean age 40 years) with >500 CD4/ml for at least 8 years and never treated with HAART. After extraction of viral RNA from plasma samples, amplification of a highly conserved region of 5'UTR was performed by nested RT-PCR. All positive samples were genotyped by sequencing, alignment with published sequences and phylogenetic analysis. CD4 cell count, HIV plasma levels were also evaluated. RESULTS: 9 out of 46 (19.56%) in Group A and 15 out of 40 (37.5%) in Group B had detectable GBV-C viremia (p=0.064, OR 2.47, percent confidence interval 0.94 to 6.51). No statistical difference was observed when disease stage was evaluated between the two groups. In Group B, after regression analysis for CD4 cell count decrease over the period observed, no significant difference was detected between GBV-C positive and negative patients. No significant difference was observed in Group B in HIV viremia and CD4 cell count at time of GBV-C detection between GBV-C infected patients and GBV-C negative patients. All Italian patients were genotype 2, the only African patient carried GBV-C genotype 1. CONCLUSIONS: Although previous results suggest that GBV-C virus may be a favorable marker for long term non progression of HIV disease, whether it plays a direct anti-HIV role or just takes advantage of non progessors' higher CD4 cell count to replicate more efficiently, still remains to be answered. Follow up of untreated patients and further evaluation of virological interactions, between the viruses and the host immune system, will be helpful to shed some light on these observations, offering new prognostic and eventually therapeutical tools for the management of HIV patients.

Adult↗

Molecular modeling of the rabbit colonic (HKalpha2a) H+, K+ ATPase.

A model of the HKalpha2a subunit of the rabbit colonic H+, K+ ATPase has been generated using the crystal structure of the Ca(+2) ATPase as a template. A pairwise sequence alignment of the deduced primary sequences of the two proteins demonstrated that they share 29% amino acid sequence identity and 47% similarity. Using O (version 7) the model of HKalpha2a was constructed by interactively mutating, deleting, and inserting the amino acids that differed between the pairwise sequence alignment of the Ca(+2) ATPase and HKalpha2a. Insertions and deletions in the HKalpha2a sequence occur in apparent extra-membraneous loop regions. The HKalpha2a model was energy minimized and globally refined to a level comparable to that of the Ca(+2) ATPase structure using CNS. The charge distribution over the surface of HKalpha2a was evaluated in GRASP and possible secondary structure elements of HKalpha2a were visualized in BOBSCRIPT. Conservation and placement of residues that may be involved in ouabain binding by the H+, K+ ATPase were considered and a putative location for the beta subunit was postulated within the structure.

Amino Acid Sequence↗

[Alignment of DNA sequences of ompL1 genes of insert fragment of recombinant plasmid, pDC38 of L. interrogans serovar lai and L. kirschneri].

In a previous study a genomic library of L. interrogans serovar lai was constructed by the present authors. Hybridization analysis (In situ, dot blot, Southern blot) with the DNA fragment containing OmpL1 (alpha-32P labeled) was performed. One of positive clones designated pDC38, was analyzed with 9 restriction enzymes (EcoRI, Bam HI, Hind III, Bgl, XbalI, ScaI, KpnI, PstI, Dra II). DNA hybridization was applied to analyze the homology of the recombinant fragment of ompL1 with the DNA of 18 strains of L. interrogans. The results showed that the homology of fragments of ompL1 were present in pathogenic Leptospira strains, but they were not in the non-pathogenic Leptospira biflexa strain Patoc I, Leptonema illini strain 3055. Therefore, Dr. David A. Haake performed the sequencing of pDC38. The results showed that pDC38 contained two inserts of 2.7 kb and 3.0 kb. The 3.0 insert contained a complete copy of the ompL1 gene. Dr. David A. Haake amplified the ompL1 gene using PCR primers specific for the ends of the gene and cloned the amplicon into pBluescript KS for sequencing. The alignment of complete DNA sequences of ompL1 of L. kirschneri and L. interrogans serovar lai showed the similarity of nucleotide sequences was 90%, variation was 10%. The derived amino acid sequence showed there was a high degree of amino acid sequence homology with ompL1 of L. kirschneri. These findings indicated that ompL1 gene was one of the important outer membrane protein genes in the L. interrogans serovar lai and using the probe pDC38 might provide a good tool for classification and identification of Leptospira.

Amino Acid Sequence↗

Modeling residue usage in aligned protein sequences via maximum likelihood.

A computational method is presented for characterizing residue usage, i.e., site-specific residue frequencies, in aligned protein sequences. The method obtains frequency estimates that maximize the likelihood of the sequences in a simple model for sequence evolution, given a tree or a set of candidate trees computed by other methods. These maximum-likelihood frequencies constitute a profile of the sequences, and thus the method offers a rigorous alternative to sequence weighting for constructing such a profile. The ability of this method to discard misleading phylogenetic effects allows the biochemical propensities of different positions in a sequence to be more clearly observed and interpreted.

Amino Acids↗

Pairwise alignment incorporating dipeptide covariation.

MOTIVATION: Standard algorithms for pairwise protein sequence alignment make the simplifying assumption that amino acid substitutions at neighboring sites are uncorrelated. This assumption allows implementation of fast algorithms for pairwise sequence alignment, but it ignores information that could conceivably increase the power of remote homolog detection. We examine the validity of this assumption by constructing extended substitution matrices that encapsulate the observed correlations between neighboring sites, by developing an efficient and rigorous algorithm for pairwise protein sequence alignment that incorporates these local substitution correlations and by assessing the ability of this algorithm to detect remote homologies. RESULTS: Our analysis indicates that local correlations between substitutions are not strong on the average. Furthermore, incorporating local substitution correlations into pairwise alignment did not lead to a statistically significant improvement in remote homology detection. Therefore, the standard assumption that individual residues within protein sequences evolve independently of neighboring positions appears to be an efficient and appropriate approximation.

Algorithms↗

Discovery of novel conserved peptide domains by ortholog comparison within plant multi-protein families.

Assigning individual functions to the proteins encoded by the genome of the dicotyledonous reference species Arabidopsis thaliana is one of the major challenges in current plant molecular biology. Frequently, Arabidopsis protein families are biocomputationally analyzed by multiple amino acid sequence alignments of the respective family members for detection of conserved peptide motifs that might be of functional relevance. Mere sequence alignment of paralogous sequences may obscure amino acid patches that are highly conserved amongst orthologs and thus potentially relevant for isoform-specific protein function(s). Here I exemplarily illustrate this potential pitfall by amino acid sequence alignments of the heptahelical MLO proteins using either the suite of 15 isoforms (paralogs) encoded by the Arabidopsis genome or a collection of 13 ortholog sequences derived from a set of both monocotyledonous and dicotyledonous plant species. The findings are corroborated by an analogous analysis of the distinct plant multi-protein family of CONSTANS-like transcription regulators. The data reveal that the generally higher sequence similarity of orthologs versus paralogs is not uniformly distributed among the amino acid positions of the orthologs but at least partially clustered in distinct sites/domains, suggesting conservation of isoform-specific functional modules across taxa.

Amino Acid Sequence↗

Computational gene prediction using multiple sources of evidence.

This article describes a computational method to construct gene models by using evidence generated from a diverse set of sources, including those typical of a genome annotation pipeline. The program, called Combiner, takes as input a genomic sequence and the locations of gene predictions from ab initio gene finders, protein sequence alignments, expressed sequence tag and cDNA alignments, splice site predictions, and other evidence. Three different algorithms for combining evidence in the Combiner were implemented and tested on 1783 confirmed genes in Arabidopsis thaliana. Our results show that combining gene prediction evidence consistently outperforms even the best individual gene finder and, in some cases, can produce dramatic improvements in sensitivity and specificity.

Arabidopsis↗

Sequence protein alignment with composition new evolutions (SPACne): a program for the identification of polypeptides using amino acid composition. A user friendly modification of SPAC.

SPAC (sequence protein alignment with composition) is a software that retrieve from protein or nucleic acid databases, the sequences corresponding to a protein or peptide whose only amino acid composition and molecular weight are known. By accurately matching a DNA or a protein sequence to candidate protein or peptide fragment, this software may be used as a fast and cheap method for protein characterization. This paper describes a modification of the SPAC software, SPAC new evolutions (SPACne), which enables a more efficient and user friendly method of protein identification using amino acids composition. SPACne is available online at the web site: http://bioweb.pasteur.fr/seqanal/interfaces/spacne.html.

Algorithms↗

Using multiple interdependency to separate functional from phylogenetic correlations in protein alignments.

MOTIVATION: Multiple sequence alignments of homologous proteins are useful for inferring their phylogenetic history and to reveal functionally important regions in the proteins. Functional constraints may lead to co-variation of two or more amino acids in the sequence, such that a substitution at one site is accompanied by compensatory substitutions at another site. It is not sufficient to find the statistical correlations between sites in the alignment because these may be the result of several undetermined causes. In particular, phylogenetic clustering will lead to many strong correlations. RESULTS: A procedure is developed to detect statistical correlations stemming from functional interaction by removing the strong phylogenetic signal that leads to the correlations of each site with many others in the sequence. Our method relies upon the accuracy of the alignment but it does not require any assumptions about the phylogeny or the substitution process. The effectiveness of the method was verified using computer simulations and then applied to predict functional interactions between amino acids in the Pfam database of alignments.

Algorithms↗

SUPERFAMILY: HMMs representing all proteins of known structure. SCOP sequence searches, alignments and genome assignments.

The SUPERFAMILY database contains a library of hidden Markov models representing all proteins of known structure. The database is based on the SCOP 'superfamily' level of protein domain classification which groups together the most distantly related proteins which have a common evolutionary ancestor. There is a public server at http://supfam.org which provides three services: sequence searching, multiple alignments to sequences of known structure, and structural assignments to all complete genomes. Given an amino acid or nucleotide query sequence the server will return the domain architecture and SCOP classification. The server produces alignments of the query sequences with sequences of known structure, and includes multiple alignments of genome and PDB sequences. The structural assignments are carried out on all complete genomes (currently 59) covering approximately half of the soluble protein domains. The assignments, superfamily breakdown and statistics on them are available from the server. The database is currently used by this group and others for genome annotation, structural genomics, gene prediction and domain-based genomic studies.

Amino Acid Sequence↗

TreeDomViewer: a tool for the visualization of phylogeny and protein domain structure.

Phylogenetic analysis and examination of protein domains allow accurate genome annotation and are invaluable to study proteins and protein complex evolution. However, two sequences can be homologous without sharing statistically significant amino acid or nucleotide identity, presenting a challenging bioinformatics problem. We present TreeDomViewer, a visualization tool available as a web-based interface that combines phylogenetic tree description, multiple sequence alignment and InterProScan data of sequences and generates a phylogenetic tree projecting the corresponding protein domain information onto the multiple sequence alignment. Thereby it makes use of existing domain prediction tools such as InterProScan. TreeDomViewer adopts an evolutionary perspective on how domain structure of two or more sequences can be aligned and compared, to subsequently infer the function of an unknown homolog. This provides insight into the function assignment of, in terms of amino acid substitution, very divergent but yet closely related family members. Our tool produces an interactive scalar vector graphics image that provides orthological relationship and domain content of proteins of interest at one glance. In addition, PDF, JPEG or PNG formatted output is also provided. These features make TreeDomViewer a valuable addition to the annotation pipeline of unknown genes or gene products. TreeDomViewer is available at http://www.bioinformatics.nl/tools/treedom/.

Computer Graphics↗