PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “alignment chaining method”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 451 records · Page 25Linked to original sources

The primary structure of high density apolipoprotein-glutamine-I.

The major protein constituent of human plasma high density lipoproteins has been isolated and its complete amino-acid sequence determined. The protein, designated apolipoprotein-glutamine-I by the presence of carboxyl-terminal glutamine, is a single polypeptide chain of 245 amino-acid residues, including three residues of methionine. The protein is devoid of cysteine, cystine, and isoleucine. Cleavage of apolipoprotein-glutamine-I with cyanogen bromide yields four fragments with 94, 90, 36, and 25 amino acids. The amino-acid sequence of each fragment was determined by conventional methods, with proteolytic digestion with trypsin, chymotrypsin, and thermolysin. The alignment of the cyanogen bromide fragments was determined by the isolation of the methionine-containing tryptic peptides from apolipoprotein-glutamine-I. Inspection of the sequence of apolipoprotein-glutamine-I suggests an interesting distribution of amino acids that may account for its helical structure and its ability to bind and transport lipid.

Amino Acid Sequence↗

DMAPS: a database of multiple alignments for protein structures.

The database of multiple alignments for protein structures (DMAPS) provides instant access to pre-computed multiple structure alignments for all protein structure families in the Protein Data Bank (PDB). Protein structure families have been obtained from four distinct classification methods including SCOP, CATH, ENZYME and CE, and multiple structure alignments have been built for all families containing at least three members, using CE-MC software. Currently, multiple structure alignments are available for 3050 SCOP-, 3087 CATH-, 664 ENZYME- and 1707 CE-based families. A web-based query system has been developed to retrieve multiple alignments for these families using the PDB chain ID of any member of a family. Multiple alignments can be viewed or downloaded in six different formats, including JOY/html, TEXT, FASTA, PDB (superimposed coordinates), JOY/postscript and JOY/rtf. DMAPS is accessible online at http://bioinformatics.albany.edu/~dmaps.

Databases, Protein↗

Structural studies on the coat protein of alfalfa mosaic virus. The complete primary structure.

The complete amino acid sequence of the coat protein of alfalfa mosaic virus (strain 425) is reported. Sequence determinations were mainly performed on peptides obtained from fragmentation by cyanogen bromide and trypsin. Both manual and automatic sequence methods were used. Some refinements of the solid-phase Edman degradation were introduced. The final alignment of the peptides was established by means of alternative cleavage methods, such as limited tryptic digestion of intact virus particles, tryptic digestion after blockage of lysine residues and chymotryptic digestion. The coat protein consists of 220 amino acid residues corresponding to a molecular weight of 24252. A remarkable clustering of basic residues occurs in the N-terminal part of the protein chain. Several internal hydrophobic clusters and a strongly acidic site at the C-terminus can be observed. Two regions of sequence homology (12 residues) were found. Some features of the secondary structure are predicted.

Amides↗

Secondary structure prediction and unrefined tertiary structure prediction for cyclin A, B, and D.

We present heuristic-based predictions of the secondary and tertiary structures of cyclins A, B, and D, representatives of the cyclin superfamily. The list of suggested constraints for tertiary structure assembly was left unrefined in order to submit this report before an announced crystal structure for cyclin A becomes available. To predict these constraints, a master sequence alignment over 270 positions of cyclin types A, B, and D was adjusted based on individual secondary structure predictions for each type. We used new heuristics for predicting aromatic residues at protein-protein interfaces and to identify sequentially distinct regions in the protein chain that cluster in the folded structure. The boundaries of two conjectured domains in the cyclin fold were predicted based on experimental data in the literature. The domain that is important for interaction of the cyclins with cyclin-dependent kinases (CDKs) is predicted to contain six helices; the second domain in the consensus model contains both helices and a beta-sheet that is formed by sequentially distant regions in the protein chain. A plausible phosphorylation site is identified. This work represents a blinded test of the method for prediction of secondary and, to a lesser extent, tertiary structure from a set of homologous protein sequences. Evaluation of our predictions will become possible with the publication of the announced crystal structure.

Amino Acid Sequence↗

Differentiation of Campylobacter coli, Campylobacter jejuni, Campylobacter lari, and Campylobacter upsaliensis by a multiplex PCR developed from the nucleotide sequence of the lipid A gene lpxA.

We describe a multiplex PCR assay to identify and discriminate between isolates of Campylobacter coli, Campylobacter jejuni, Campylobacter lari, and Campylobacter upsaliensis. The C. jejuni isolate F38011 lpxA gene, encoding a UDP-N-acetylglucosamine acyltransferase, was identified by sequence analysis of an expression plasmid that restored wild-type lipopolysaccharide levels in Escherichia coli strain SM105 [lpxA(Ts)]. With oligonucleotide primers developed to the C. jejuni lpxA gene, nearly full-length lpxA amplicons were amplified from an additional 11 isolates of C. jejuni, 20 isolates of C. coli, 16 isolates of C. lari, and five isolates of C. upsaliensis. The nucleotide sequence of each amplicon was determined, and sequence alignment revealed a high level of species discrimination. Oligonucleotide primers were constructed to exploit species differences, and a multiplex PCR assay was developed to positively identify isolates of C. coli, C. jejuni, C. lari, and C. upsaliensis. We characterized an additional set of 41 thermotolerant isolates by partial nucleotide sequence analysis to further demonstrate the uniqueness of each species-specific region. The multiplex PCR assay was validated with 105 genetically defined isolates of C. coli, C. jejuni, C. lari, and C. upsaliensis, 34 strains representing 12 additional Campylobacter species, and 24 strains representing 19 non-Campylobacter species. Application of the multiplex PCR method to whole-cell lysates obtained from 108 clinical and environmental thermotolerant Campylobacter isolates resulted in 100% correlation with biochemical typing methods.

Acyltransferases↗

Peptide sequences binding to MHC class I proteins.

Motifs for peptides which bind specifically to the human class I major histocompatibility complex molecules HLA-A2 and B7 were determined by sequence analysis of class I-bound peptides selected from a random synthetic library of nonamers. Thirteen individual peptides were sequenced for HLA-A2, twelve individual and nine pooled peptides were sequenced for HLA-B7. Analysis of sequence alignment implicated four peptide positions in potential contact with the class I HLA-A2 molecule and three positions for the HLA-B7 molecule. The results demonstrate that a synthetic peptide library can be used to identify allele-specific motifs for class I molecules, providing information comparable to the results obtained from sequencing endogenous peptides. This method utilizes denatured class I heavy chains, and similar results were obtained using a class I protein purified from mammalian cells or by expression in Escherichia coli. This method has the potential to detect peptides which may not be generated physiologically, but due to their binding properties, may be valuable to predict or engineer immunomodulatory T cell epitopes.

Amino Acid Sequence↗

Identification and characterization of subfamily-specific signatures in a large protein superfamily by a hidden Markov model approach.

BACKGROUND: Most profile and motif databases strive to classify protein sequences into a broad spectrum of protein families. The next step of such database studies should include the development of classification systems capable of distinguishing between subfamilies within a structurally and functionally diverse superfamily. This would be helpful in elucidating sequence-structure-function relationships of proteins. RESULTS: Here, we present a method to diagnose sequences into subfamilies by employing hidden Markov models (HMMs) to find windows of residues that are distinct among subfamilies (called signatures). The method starts with a multiple sequence alignment (MSA) of the subfamily. Then, we build a HMM database representing all sliding windows of the MSA of a fixed size. Finally, we construct a HMM histogram of the matches of each sliding window in the entire superfamily. To illustrate the efficacy of the method, we have applied the analysis to find subfamily signatures in two well-studied superfamilies: the cadherin and the EF-hand protein superfamilies. As a corollary, the HMM histograms of the analyzed subfamilies revealed information about their Ca2+ binding sites and loops. CONCLUSIONS: The method is used to create HMM databases to diagnose subfamilies of protein superfamilies that complement broad profile and motif databases such as BLOCKS, PROSITE, Pfam, SMART, PRINTS and InterPro.

Binding Sites↗

Discovering new genes with advanced homology detection.

Most genome annotation protocols combine ab initio predictions with transcription and homology analyses to produce reliable gene predictions but they often fail to detect many actual genes. Alternative approaches involving more sensitive homology recognition methods are playing an increasingly important role in the next stage of gene discovery. The hunt for new genes is far from over.

Database Management Systems↗

Rapid identification of clinically relevant Nocardia species to genus level by 16S rRNA gene PCR.

Two regions of the gene coding for 16S rRNA in Nocardia species were selected as genus-specific primer sequences for a PCR assay. The PCR protocol was tested with 60 strains of clinically relevant Nocardia isolates and type strains. It gave positive results for all strains tested. Conversely, the PCR assay was negative for all tested species belonging to the most closely related genera, including Dietzia, Gordona, Mycobacterium, Rhodococcus, Streptomyces, and Tsukamurella. Besides, unlike the latter group of isolates, all Nocardia strains exhibited one MlnI recognition site but no SacI restriction site. This assay offers a specific and rapid alternative to chemotaxonomic methods for the identification of Nocardia spp. isolated from pathogenic samples.

Bacteriological Techniques↗

Exopolysaccharide-associated protein sorting in environmental organisms: the PEP-CTERM/EpsH system. Application of a novel phylogenetic profiling heuristic.

BACKGROUND: Protein translocation to the proper cellular destination may be guided by various classes of sorting signals recognizable in the primary sequence. Detection in some genomes, but not others, may reveal sorting system components by comparison of the phylogenetic profile of the class of sorting signal to that of various protein families. RESULTS: We describe a short C-terminal homology domain, sporadically distributed in bacteria, with several key characteristics of protein sorting signals. The domain includes a near-invariant motif Pro-Glu-Pro (PEP). This possible recognition or processing site is followed by a predicted transmembrane helix and a cluster rich in basic amino acids. We designate this domain PEP-CTERM. It tends to occur multiple times in a genome if it occurs at all, with a median count of eight instances; Verrucomicrobium spinosum has sixty-five. PEP-CTERM-containing proteins generally contain an N-terminal signal peptide and exhibit high diversity and little homology to known proteins. All bacteria with PEP-CTERM have both an outer membrane and exopolysaccharide (EPS) production genes. By a simple heuristic for screening phylogenetic profiles in the absence of pre-formed protein families, we discovered that a homolog of the membrane protein EpsH (exopolysaccharide locus protein H) occurs in a species when PEP-CTERM domains are found. The EpsH family contains invariant residues consistent with a transpeptidase function. Most PEP-CTERM proteins are encoded by single-gene operons preceded by large intergenic regions. In the Proteobacteria, most of these upstream regions share a DNA sequence, a probable cis-regulatory site that contains a sigma-54 binding motif. The phylogenetic profile for this DNA sequence exactly matches that of three proteins: a sigma-54-interacting response regulator (PrsR), a transmembrane histidine kinase (PrsK), and a TPR protein (PrsT). CONCLUSION: These findings are consistent with the hypothesis that PEP-CTERM and EpsH form a protein export sorting system, analogous to the LPXTG/sortase system of Gram-positive bacteria, and correlated to EPS expression. It occurs preferentially in bacteria from sediments, soils, and biofilms. The novel method that led to these findings, partial phylogenetic profiling, requires neither global sequence clustering nor arbitrary similarity cutoffs and appears to be a rapid, effective alternative to other profiling methods.

Amino Acid Motifs↗

Detection and identification of avian, duck, and goose reoviruses by RT-PCR: goose and duck reoviruses are part of the same genogroup in the genus Orthoreovirus.

A reverse transcription-polymerase chain reaction (RT-PCR) procedure for the detection of avian, duck, and goose reovirus (ARV, DRV, and GRV) RNA from cell culture supernatant and clinical samples was established. Based on multiple sequence alignment, a pair of degenerate primers was selected and synthesized. The amplified, cloned, and sequenced 598-base-pair products from the sigmaA-encoding gene fragment from 16 isolates (ranging over 30 years) indicated that the primer regions were well conserved. The sensitivity of this method was determined to be 10(-2) PFU. The specificity of the RT-PCR method was determined by testing specimens containing avian influenza A viruses, Newcastle disease virus, and infectious bronchitis virus, all of which yielded negative results with no discernible background. The efficiency of the system for detection of ARV, DRV, and GRV directly in 71/83 clinical samples was confirmed. The nucleotide sequence analysis indicated that DRV and GRV isolated from China in different locales and years were closely related, showing 97.4-100% homology to each other, but with only 86.7-88.5% identity to DRV 89026. The nucleotide and amino acid sequence identities in the amplified sigmaA-encoding gene were 74.2-78.4% and 86.9-92.0%, respectively, between duck/goose and chicken species. Phylogenetic analysis indicated that GRV and DRV aggregated into the same specified genogroup within subgroup II of the genus Orthoreovirus and are more closely related to ARV than to Nelson Bay virus. Overall, this study developed a sensitive and specific technique for the identification ARV, DRV, and GRV, and sequencing analysis has enhanced our understanding of the evolutionary relationship between ARV, DRV, and GRV.

Animals↗

Interpretation of oligonucleotide mass spectra for determination of sequence using electrospray ionization and tandem mass spectrometry.

Procedures are described for interpretation of mass spectra from collision-induced dissociation of polycharged oligonucleotides produced by electrospray ionization. The method is intended for rapid sequencing of oligonucleotides of completely unknown structure at approximately the 15-mer level and below, from DNA or RNA. Identification of sequence-relevant ions that are produced from extensive fragmentation in the quadrupole collision cell are based primarily on (1) recognition of 3'- and 5'- terminal residues as initial steps in mass ladder propagation, (2) alignment of overlapping nucleotide chains that have been constructed independently from each terminus, and (3) use of experimentally measured molecular mass in rejection of incorrect sequence candidates. Algorithms for sequence derivation are embodied in a computer program that requires < 2s for execution. The interpretation procedures are demonstrated for sequence location of simple forms of modification in the base and sugar. The potential for direct sequencing of components of mixtures is shown using an unresolved fraction of unknown oligonucleotides from ribosomal RNA.

Algorithms↗

Reconstructing evolutionary trees from DNA and protein sequences: paralinear distances.

The reconstruction of phylogenetic trees from DNA and protein sequences is confounded by unequal rate effects. These effects can group rapidly evolving taxa with other rapidly evolving taxa, whether or not they are genealogically related. All algorithms are sensitive to these effects whenever the assumptions on which they are based are not met. The algorithm presented here, called paralinear distances, is valid for a much broader class of substitution processes than previous algorithms and is accordingly less affected by unequal rate effects. It may be used with all nucleic acid, protein, or other sequences, provided that their evolution may be modeled as a succession of Markov processes. The properties of the method have been proven both analytically and by computer simulations. Like all other methods, paralinear distances can fail when sequences are misaligned or when site-to-site sequence variation of rates is extensive. To examine the usefulness of paralinear distances, the "origin of the eukaryotes" has been investigated by the analysis of elongation factor Tu sequences with a variety of sequence alignments. It has been found that the order in which sequences are pairwise aligned strongly determines the topology which is reconstructed by paralinear distances (as it does for all other reconstruction methods tested). When the parts of the alignment that are unaffected by alignment order are analyzed, paralinear distances strongly select the eocyte topology. This provides evidence that the eocyte prokaryotes are the closest prokaryotic relatives of the eukaryotes.

Algorithms↗

Statistical significance of probabilistic sequence alignment and related local hidden Markov models.

The score statistics of probabilistic gapped local alignment of random sequences is investigated both analytically and numerically. The full probabilistic algorithm (e.g., the "local" version of maximum-likelihood or hidden Markov model method) is found to have anomalous statistics. A modified "semi-probabilistic" alignment consisting of a hybrid of Smith-Waterman and probabilistic alignment is then proposed and studied in detail. It is predicted that the score statistics of the hybrid algorithm is of the Gumbel universal form, with the key Gumbel parameter lambda taking on a fixed asymptotic value for a wide variety of scoring systems and parameters. A simple recipe for the computation of the "relative entropy," and from it the finite size correction to lambda, is also given. These predictions compare well with direct numerical simulations for sequences of lengths between 100 and 1,000 examined using various PAM substitution scores and affine gap functions. The sensitivity of the hybrid method in the detection of sequence homology is also studied using correlated sequences generated from toy mutation models. It is found to be comparable to that of the Smith-Waterman alignment and significantly better than the Viterbi version of the probabilistic alignment.

Algorithms↗

A hidden Markov model for progressive multiple alignment.

MOTIVATION: Progressive algorithms are widely used heuristics for the production of alignments among multiple nucleic-acid or protein sequences. Probabilistic approaches providing measures of global and/or local reliability of individual solutions would constitute valuable developments. RESULTS: We present here a new method for multiple sequence alignment that combines an HMM approach, a progressive alignment algorithm, and a probabilistic evolution model describing the character substitution process. Our method works by iterating pairwise alignments according to a guide tree and defining each ancestral sequence from the pairwise alignment of its child nodes, thus, progressively constructing a multiple alignment. Our method allows for the computation of each column minimum posterior probability and we show that this value correlates with the correctness of the result, hence, providing an efficient mean by which unreliably aligned columns can be filtered out from a multiple alignment.

Algorithms↗

Prediction of protein interdomain linker regions by a hidden Markov model.

MOTIVATION: Our aim was to predict protein interdomain linker regions using sequence alone, without requiring known homology. Identifying linker regions will delineate domain boundaries, and can be used to computationally dissect proteins into domains prior to clustering them into families. We developed a hidden Markov model of linker/non-linker sequence regions using a linker index derived from amino acid propensity. We employed an efficient Bayesian estimation of the model using Markov Chain Monte Carlo, Gibbs sampling in particular, to simulate parameters from the posteriors. Our model recognizes sequence data to be continuous rather than categorical, and generates a probabilistic output. RESULTS: We applied our method to a dataset of protein sequences in which domains and interdomain linkers had been delineated using the Pfam-A database. The prediction results are superior to a simpler method that also uses linker index.

Algorithms↗

Contact-based sequence alignment.

This paper introduces the novel method of contact-based protein sequence alignment, where structural information in the form of contact mutation probabilities is incorporated into an alignment routine using contact-mutation matrices (CAO: Contact Accepted mutatiOn). The contact-based alignment routine optimizes the score of matched contacts, which involves four (two per contact) instead of two residues per match in pairwise alignments. The first contact refers to a real side-chain contact in a template sequence with known structure, and the second contact is the equivalent putative contact of a homologous query sequence with unknown structure. An algorithm has been devised to perform a pairwise sequence alignment based on contact information. The contact scores were combined with PAM-type (Point Accepted Mutation) substitution scores after parameterization of gap penalties and score weights by means of a genetic algorithm. We show that owing to the structural information contained in the CAO matrices, significantly improved alignments of distantly related sequences can be obtained. This has allowed us to annotate eight putative Drosophila IGF sequences. Contact-based sequence alignment should therefore prove useful in comparative modelling and fold recognition.

Algorithms↗