PubMed HealthSearch

SEARCH · PubMed Health

Results for “alignment chaining method”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Sequence of 18S rDNA of actinorhizal Alnus glutinosa (Betulaceae).

The small subunit ribosomal DNA for a woody actinorhizal, Alnus glutinosa, was isolated by the PCR method. Amplification products were cloned into the Bluescript SK- vector. Full sequence, 1698 bp, was obtained with NS1 to NS8 primers. Sequence alignments were made by UWGCG sequence data analysis computer programs. 18S rDNA sequence of A. glutinosa was compared to analogous segments of four other angiosperms, tomato, rice, maize and soybean. Sequence homologies are discussed and application for the technique is suggested.

Base Sequence

A simple method to generate non-trivial alternate alignments of protein sequences.

A major problem in sequence alignments based on the standard dynamic programming method is that the optimal path does not necessarily yield the best equivalencing of residues assessed by structural or functional criteria. An algorithm is presented that finds suboptimal alignments of protein sequences by a simple modification to the standard dynamic programming method. The standard pairwise weight matrix elements are modified in order to penalize, but not eliminate, the equivalencing of residues obtained from previous alignments. The algorithm thereby yields a limited set of alternate alignments that can differ considerably from the optimal. The approach is benchmarked on the alignments of immunoglobulin domains. Without a prior knowledge of the optimal choice of gap penalty, one of the suboptimal alignments is shown to be more accurate than the optimal.

Algorithms

Autolytic fragmentation of complement components C3 and C4 and its relationship to covalent binding activity.

The autolytic cleavage reaction of C3 and C4 and the covalent binding reaction of these proteins, are both aspects of the reactivity of an activated thiolester within these proteins. Autolytic cleavage occurs by internal nucleophilic attack on one face of the planar thiolester, while the covalent binding reaction of the activated proteins follows exposure of the opposite face of the thiolester to attack by external nucleophiles. Although the autolytic cleavage reaction does not occur under physiological conditions, the study of this phenomenon has provided valuable evidence in support of the mechanisms postulated for the physiological covalent binding reactions. The ease with which autolysis can be induced and observed in C3, C4, and alpha 2 M has provided a valuable method for detecting the active forms of these proteins in circumstances where other assays are impracticable, as, for example, in the examination of the uptake of active C3 by lymphocytes. Autolytic cleavage has also been used by Karp and colleagues to produce fragments used in characterizing genetic and biosynthetic variants of mouse C4 and the mouse protein Slp, which is structurally similar to C4. Gross structural comparisons made among C3, C4, and alpha 2 M on the basis of alignment of the autolytic cleavage sites and the protease-activation sites in these proteins were useful in predicting how the alpha-, beta-, and gamma-chains of C4, or the alpha- and beta-chains of C3, were aligned in the single polypeptide chain pro-forms of these proteins. The beta-alpha-gamma alignment deduced for C4 was also found by Goldberger and Colten. Similar alignments of cleavage sites have been used as a basis for evolutionary comparisons of complement proteins and alpha 2 M from species other than man. Although autolytic cleavage has been described only for C3, C4, alpha 2 M, and Slp, it is likely that other proteins will be found that exhibit this phenomenon. A possible candidate is pregnancy-associated plasma protein A (PAPP-A) which resembles alpha 2 M in many respects. The autolytic cleavage reaction will serve as a useful indicator in the detection of other proteins that undergo covalent binding by the mechanism discussed above.

Complement C3

Rat major acute-phase protein: biosynthesis and characterization of cDNA clone.

The major acute-phase protein (alpha 1-MAP) of rat serum is induced in response to inflammation. This induction may be attributed to a corresponding increase in the level of translatable mRNA for the protein. Using in vitro and in vivo systems, various biosynthetic processing intermediates of this glycoprotein have been isolated. alpha 1-MAP is translated in a rabbit reticulocyte system as a preprotein with an amino-terminal signal peptide and an apparent molecular weight of 51,000. Translation of rough microsomes yields a product with a mass of 57,000 Da, representing the core glycosylated form of alpha 1-MAP. Cotranslational glycosylation appears to occur in a stepwise fashion, since three glycosylated forms of alpha 1-MAP (51,000, 54,000, and 57,000 Da) were detected in polysome translations; these products were digested by endoglycosidase H to a 48,000-Da protein. Two intracellular forms of alpha 1-MAP were observed in vivo, a 57,000-Da (core carbohydrate sidechains) and a 66,000-Da protein (mature complex carbohydrate side-chains); the latter was the only component secreted into the culture medium. To extend our studies on this protein, a cDNA clone specific for alpha 1-MAP was isolated. The recombinant was positively identified by hybrid selection procedures and contains a 1.55-kb insert. Partial radiosequence analysis of the primary translation product indicated the distribution of Leu, Ile, Cys, and Met in the amino-terminal region of this protein. To relate the location of these amino acids with the nucleotide sequence, cDNA was analyzed by the method of Maxam and Gilbert. These results indicate that the cDNA insert contains the 3' poly(A) tail, and alignment of the 5' end of the cDNA with the available amino acid sequence of the primary translation product corroborated that the insert encodes the entire alpha 1-MAP protein except for the first four amino acids of the signal peptide.

Amino Acid Sequence

Genetic analysis of the polymorphism of the human apolipoprotein E using automated solid-phase sequencing.

A direct sequencing approach has been used to analyze the polymorphism in the human apolipoprotein E gene. A method is described, in which the DNA is amplified by the polymerase chain reaction, immobilized, and sequenced by a semi-automatic procedure adaptable to clinical diagnosis. The three alleles of the apolipoprotein E gene, which differ from each other by two nucleotide substitutions and which influence serum cholesterol levels, were analyzed. The solid-phase method was able to resolve the correct nucleotide sequence in samples from both homozygous and heterozygous individuals. No cloning steps are needed and the immobilization and separation of the DNA is accomplished using magnetic beads.

Apolipoproteins E

A multiple sequence alignment algorithm for homologous proteins using secondary structure information and optionally keying alignments to functionally important sites.

The programs described herein function as part of a suite of programs designed for pairwise alignment, multiple alignment, generation of randomized sequences, production of alignment scores and a sorting routine for analysis of the alignments produced. The sequence alignment programs penalize gaps (absences of residues) within regions of protein secondary structure and have the added option of 'fingerprinting' structurally or functionally important protein-residues. The multiple alignment program is based upon the sequence alignment method of Needleman and Wunsch and the multiple alignment extension of Barton and Sternberg. Our application includes the feature of optionally weighting active site, monomer--monomer, ligand contact or other important template residues to bias the alignment toward matching these residues. A sum-score for the alignments is introduced, which is independent of gap penalties. This score more adequately reflects the character of the alignments for a given scoring matrix than the gap-penalty-dependent total score described previously in the literature. In addition, individual amino acid similarity scores at each residue position in the alignments are printed with the alignment output to enable immediate quantitative assessment of homology at key sections of the aligned chains.

Algorithms

Gene finding in the chicken genome.

BACKGROUND: Despite the continuous production of genome sequence for a number of organisms, reliable, comprehensive, and cost effective gene prediction remains problematic. This is particularly true for genomes for which there is not a large collection of known gene sequences, such as the recently published chicken genome. We used the chicken sequence to test comparative and homology-based gene-finding methods followed by experimental validation as an effective genome annotation method. RESULTS: We performed experimental evaluation by RT-PCR of three different computational gene finders, Ensembl, SGP2 and TWINSCAN, applied to the chicken genome. A Venn diagram was computed and each component of it was evaluated. The results showed that de novo comparative methods can identify up to about 700 chicken genes with no previous evidence of expression, and can correctly extend about 40% of homology-based predictions at the 5' end. CONCLUSIONS: De novo comparative gene prediction followed by experimental verification is effective at enhancing the annotation of the newly sequenced genomes provided by standard homology-based methods.

Animals

An efficient and reliable method for cloning PCR-amplification products: a survey of point mutations in integrin cDNA.

A highly efficient, non-labor-intensive method for cloning DNA fragments produced by PCR amplification was used to carry out a rapid survey of potential point mutations in integrin alpha 6 cDNA from 17 different cell-type sources. The method includes glass powder purification of the PCR reaction mixture, followed by simultaneous treatment with T4 polynucleotide kinase and DNA polymerase I, and another glass powder purification. Sequences from multiple subclones of each cell type were readily generated, aligned and checked for mismatches. Several commonly used alternative procedures were compared for cloning efficiency and size-fidelity of inserted DNA fragments.

Amino Acid Sequence

Protein structure prediction.

Current methods developed for predicting protein structure are reviewed. The most widely used algorithms of Chou and Fasman and Garnier et al for predicting secondary structure are compared to the most recent ones including sequence similarity methods, neural network, pattern recognition or joint prediction methods. The best of these methods correctly predict 63-65% of the residues in the database with cross-validation for 3 conformations, helix, beta strand and coli with a standard deviation of 6-8% per protein. However, when a homologous protein is already in the database, the accuracy of prediction by the similarity peptide method of Levin and Garnier reaches about 90%. Some conclusions can be drawn on the mechanism of protein folding. As all the prediction methods only use the local sequence for prediction (+/- 8 residues maximum) one can infer that 65% of the conformation of a residue is dictated on average by the local sequence, the rest is brought by the folding. The best predicted proteins or peptide segments are those for which the folding has less effect on the conformation. Presently, prediction of tertiary structure is only of practical use when the structure of a homologous protein is already known. Amino acid alignment to define residues of equivalent spatial position is critical for modelling of the protein. We showed for serine proteases that secondary structure prediction can help to define a better alignment. Non-homologous segments of the polypeptide chain, such as loops, libraries of known loops and/or energy minimization with various force fields, are used without yet giving satisfactory solutions. An example of modelling by homology, aided by secondary structure prediction on 2 regulatory proteins, Fnr and FixK is presented.

Algorithms

The serine proteinase chain of human complement component C1s. Cyanogen bromide cleavage and N-terminal sequences of the fragments.

Human complement component C1s was purified from fresh blood by conventional methods of precipitation and chromatography. The single-chain zymogen form was activated by treatment with C1r. Reduction and carboxymethylation then allowed the light chain and heavy chain to be separated on DEAE-Sepharose CL-6B in 8 M-urea. Liquid-phase sequencing of the light chain determined 50 residues from the N-terminus. CNBr-cleavage fragments of the light chain were separated by high-pressure liquid chromatography on gel-permeation and reverse-phase columns. N-Terminal sequencing of these fragments determined the order of a further 138 residues, giving a total of 188 residues or about 75% of the light chain. Seven of these eight sequences could be readily aligned with the amino acid sequences of other serine proteinases. The typical serine proteinase active-site residues are clearly conserved in C1s, and the specificity-related side chain of the substrate-binding pocket is aspartic acid, as in trypsin, consistent with the proteolytic action of C1s on C4 at an arginine residue. Somewhat surprisingly, when the C1s sequence is compared with that of complement subcomponent C1r, the percentage difference (59%) is approximately the same as that found between the other mammalian serine proteinases (56-71%).

Amino Acid Sequence

Cloning, sequencing and homologies of the cbh-1 (exoglucanase) gene of Humicola grisea var. thermoidea.

Studies on the enzymes of the cellulase complex of the thermophilic fungus Humicola grisea var. thermoidea are described. A genomic library was constructed in the phage vector EMBL 4, and from this library two clones were isolated using as a probe the cloned cbh-1 (exoglucanase, EC 3.2.1.91) gene of Phanerochaete chrysosporium, a cellulolytic basidiomycete fungus. These clones were analysed by restriction mapping and Southern blotting, and one of them (lambda 3) was sub-cloned into the M13 phage vectors mp18 and mp19. The gene sequence was determined by the dideoxy chain-termination method. Sequence comparison with the equivalent genes from P. chrysosporium and Trichoderma reesei was made: in terms of primary sequence there is about 60% homology between the three species. Secondary structure prediction of the H. grisea sequence was also computed.

Amino Acid Sequence

Amino acid substitutions in structurally related proteins. A pattern recognition approach. Determination of a new and efficient scoring matrix.

Amino acid substitutions in evolutionarily related proteins have been studied from a structural point of view. We consider here that an amino acid al in a protein p1 has been replaced by the amino acid a2 in the structurally similar protein p2 if, after superposition of the p1 and p2 structures, the a1 and a2 C alpha atoms are no more than 1.2 A apart. Thirty-two proteins, grouped in 11 classes, have been analysed by this method. This produced 2860 amino acid pairs (substitutions), which were analysed by multi-dimensional statistical methods. The main results are as follows: (1) according to the observed exchangeability of amino acid side-chains, only four groups (strong clusters) could be delineated; (i) Ile and Val, (ii) Leu and Met, (iii) Lys, Arg and Gln, and (iv) Tyr and Phe. The other residues could not be classified. (2) The matrix of distances between amino acids, or scoring matrix, determined from this study, is different from any other published matrix. (3) Except for the distance matrices based on the chemical properties of amino acid side-chains, which can be grouped together, all other published matrices are different from one another. (4) The distance matrix determined in this study seems to be very efficient for aligning distantly related protein sequences.

Amino Acid Sequence

A method of estimating from two aligned present-day DNA sequences their ancestral composition and subsequent rates of substitution, possibly different in the two lineages, corrected for multiple and parallel substitutions at the same site.

The course of evolutionary change in DNA sequences has been modeled as a Markov process. The Markov process was represented by discrete time matrix methods. The parameters of the Markov transition matrices were estimated by least-squares direct-search optimization of the fit of the calculated divergence matrix to that observed for two aligned sequences. The Markov process corrected for multiple and parallel substitutions of bases at the same site. The method avoided the incorrect assumption of all previously described methods that the divergence between two present-day sequences is twice the divergence of either from the common and unknown ancestral sequence. The three previous methods were shown to be equivalent. The present method also avoided the undesirable assumptions that sequence composition has not changed with time and that the substitution rates in the two descendant lineages were the same. It permitted simultaneous estimation of ancestral sequence composition and, if applicable, of different substitution rates for the two descendant lineages, provided the total number of estimated parameters was less than 16. Properties of the Markov chain were discussed. It was proved for symmetric substitution matrices that all elements of the equilibrium divergence matrix equal 1/16, and that the total difference in the divergence matrix at epoch k equals the total change in the common substitution matrix at epoch 2k for all values of k. It was shown how to resolve an ambiguity in the assignment of two different substitution rates to the two descendant lineages when four or more similar sequences are available. The method was applied to the divergence matrix for codon site 3 for the mouse and rabbit beta-globins. This observed divergence matrix was significantly asymmetric and required at least two different substitution rates. This result could be achieved only by using different asymmetric substitution matrices for the two lineages.

Animals

Designated primers targeted canine TP53 gene hotspot regions.

BACKGROUND: Tumor protein 53 gene (TP53) is a critical factor that controls different cell activities such as cell cycle, DNA repair mechanism, autophagy, apoptosis, and metabolism. The TP53 gene is the most commonly mutated gene, especially in the 4-8 exons region. This mutation enhances the development of many abnormalities, such as the initiation of different types of cancer. AIM: The main objective of this study was to design and evaluate the efficacy of three different primer sets that targeted the TP53 gene at the hotspot regions. METHODS: To do that, twelve blood samples were collected from dogs belonging to the German Shepherd breed/K9 aged between 8-12 years. Then, the DNA extraction and polymerase chain reaction (PCR) took place by using the three primer sets, which were designed using SnapGene. The primer sets, namely, first primer, the second and the third targeted exons 5-9 located in the canine TP53 gene. In the following step, all the PCR products were sent for Sanger sequencing and then phylogenetic analysis. RESULTS: Our findings indicated that the first primer set consistently showed higher amplification signal efficiency and reduced dimer formation compared with the second and third primer sets, respectively, with a 60ºC annealing temperature. In addition, all the sequenced samples aligned with the reference canine TP53 gene in the phylogenetic tree. CONCLUSION: This study offered the best TP53 primer design that targeted the hotspot regions of the canine TP53 gene for researchers who are interested in targeting such regions in this gene.

Animals

A simple and rapid method for HLA-DQA1 genotyping by polymerase chain reaction-single strand conformation polymorphism and restriction enzyme cleavage analysis.

A simple and rapid method for identification of alleles at the human leucocyte antigen (HLA)-DQA1 locus is described. The polymorphic second exon of the HLA-DQA1 locus was amplified by the polymerase chain reaction (PCR) method. The amplified DNA was analyzed by single-strand conformation polymorphism (SSCP) and restriction enzyme cleavage assay. Using this method, the eight known DQA1 alleles could be distinguished from each other. This paper suggests that the method can be used for quick genotyping of DQA1 alleles, but detecting point mutations at various positions in a fragment as well as new HLA-DQA1 genotypes should also be possible.

Alleles

The heterogeneity of the polymeric intracellular hemoglobin of Glycera dibranchiata and the cDNA-derived amino acid sequence of one component.

The erythrocytes of the marine polychaete Glycera dibranchiata contain a number of different, single-chain hemoglobins, some of which self-associate into a 'polymeric' fraction. An oligodeoxynucleotide probe was synthesized based on partial amino acid sequences determined by chemical methods, and used to screen a cDNA library constructed from the poly(A+)mRNA of Glycera erythrocytes (Simons, P.C. and Satterlee, J.D. (1989) Biochemistry 28, 8525-8530). The longest positive inserts found were sequenced using the dideoxy nucleotide chain termination method. One complete clone was obtained: clone 5A, 816 bases long, contained 59 bases of 5'-untranslated RNA, an open reading frame of 441 bases coding for 147 amino acids and a 3'-untranslated region of 316 bases. The derived amino acid sequence of Glycera globin P1 was in agreement with the partial amino acid sequences obtained by chemical methods. Three additional inserts obtained in the screening were also sequenced: the inferred amino acid sequences proved to be partial globin sequences which were different from each other and from the sequence of P1. Thus, the 'polymeric' fraction of the intracellular hemoglobin of Glycera probably consists of at least four different globin chains much like the 'monomeric' fraction. Comparison of the 'polymeric' sequence with the two known 'monomeric' sequences, M-II and M-IV, shows that they share 54 identical residues. At 74 positions, the identical residues in M-II and M-IV differ from the corresponding residue in P1, including at E-7, where P1 has a distal His, in contrast to Leu in M-II and M-IV. The alignment of Bashford et al. ((1987) J. Mol. Biol. 196, 199-216) and their templates were used to examine the principal differences between the two types of Glycera globin sequences. They appear to consist of uncommon surface amino acid residues at positions C6 (Phe vs. Ala), E10 (Val vs. Lys), E17 (Lys vs. Val), G1 (Arg vs. Lys), G10 (Met vs. Ala) and H5 (Arg vs. Lys). One or more of these residues could be responsible for the self-association exhibited by the 'polymeric' Glycera globins.

Amino Acid Sequence

Molecular organization of Junin virus S RNA: complete nucleotide sequence, relationship with other members of the Arenaviridae and unusual secondary structures.

In this study, overlapping cDNA clones covering the entire S RNA molecule of Junin virus, an arenavirus that causes Argentine haemorrhagic fever, were generated. The complete sequence of this 3400 nucleotide RNA was determined using the dideoxynucleotide chain termination method. The nucleocapsid protein (N) and the glycoprotein precursor (GPC) genes were identified as two non-overlapping open reading frames of opposite polarity, encoding primary translation products of 564 and 481 amino acids, respectively. Intracellular processing of the latter yields the glycoproteins found in the viral envelope. Comparison of the Junin virus N protein with the homologous proteins of other arenaviruses indicated that amino acid sequences are conserved, the identity ranging from 46 to 76%. The N-terminal half of GPC exhibits an even higher degree of conservation (54 to 82%), whereas the C-terminal half is less conserved (21 to 50%). In all comparisons the highest level of amino acid sequence identity was seen when Junin virus and Tacaribe virus sequences were aligned. The nucleotide sequence at the 5' end of Junin virus S RNA is not identical to that determined of the other sequenced arenaviruses. However, it is complementary to the 3'-terminal sequences and may form a very stable panhandle structure (delta G-242.7 kJ/mol) involving the complete non-coding regions upstream from both the N and GPC genes. In addition, a distinct secondary structure was identified in the intergenic region, downstream from the coding sequences; Junin virus S RNA shows a potential secondary structure consisting of two hairpin loops (delta G -163.2 and -239.3 kJ/mol) instead of the single hairpin loop that is usually found in other arenaviruses. The analysis of the arenavirus S RNA nucleotide sequences and their encoded products is discussed in relation to structure and function.

Amino Acid Sequence

HLA-DR typing by PCR amplification with sequence-specific primers (PCR-SSP) in 2 hours: an alternative to serological DR typing in clinical practice including donor-recipient matching in cadaveric transplantation.

In most PCR-based tissue typing techniques the PCR amplification is followed by a post-amplification specificity step. In typing by PCR amplification with sequence-specific primers (PCR-SSP), typing specificity is part of the amplification step, which makes the technique almost as fast as serological tissue typing. In the present study primers were designed for DR "low-resolution" typing by PCR-SSP, i.e. identifying polymorphism corresponding to the serologically defined series DR1-DRw18. This resolution was achieved by performing 19 PCR reactions per individual, 17 for assigning DR1-DRw18 and 2 for the DRw52 and DRw53 superspecificities. Thirty cell lines and 121 individuals were typed by the DR "low-resolution" PCR-SSP technique, TaqI DRB-DQA-DQB RFLP analysis and serology. The concordance between PCR-SSP typing and RFLP analysis was 100%. The reproducibility was 100% in 40 samples typed on two separate occasions. No false-positive or false-negative typing results were obtained. All homozygous and heterozygous combinations of DR1-DRw18 could be distinguished. Amplification patterns segregated according to dominant Mendelian inheritance. DNA preparation, PCR amplification and post-amplification processing, including gel detection, documentation and interpretation, were performed in 2 hours. In conclusion, PCR-SSP is an accurate typing technique with high sensitivity, specificity and reproducibility. The method is rapid and inexpensive. DR "low-resolution" typing by the PCR-SSP technique is ideally suited for analyzing small numbers of samples simultaneously and is an alternative to serological DR typing in routine clinical practice including donor-recipient matching in cadaveric transplantations.

Alleles