PubMed HealthSearch

SEARCH · PubMed Health

Results for “Multiple sequence alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Complement components C1r/C1s, bone morphogenic protein 1 and Xenopus laevis developmentally regulated protein UVS.2 share common repeats.

Property patterns were constructed, based on an alignment of related domains in human complement subcomponents C1r and C1s as well as in the sea urchin protein uEGF. This kind of consensus pattern was able to identify similar domains in a human bone morphogenic protein, in a Xenopus laevis embryonal protein involved in dorsoanterior development and in a calcium-dependent serine protease secreted from malignant hamster embryo fibroblast cells. Because of the high level of overall sequence homology this protease may be the hamsters' equivalent of the human complement subcomponent C1s. The resulting multiple alignment of all studied domains suggests functionally and structurally important regions.

Amino Acid Sequence

An efficient and reliable method for cloning PCR-amplification products: a survey of point mutations in integrin cDNA.

A highly efficient, non-labor-intensive method for cloning DNA fragments produced by PCR amplification was used to carry out a rapid survey of potential point mutations in integrin alpha 6 cDNA from 17 different cell-type sources. The method includes glass powder purification of the PCR reaction mixture, followed by simultaneous treatment with T4 polynucleotide kinase and DNA polymerase I, and another glass powder purification. Sequences from multiple subclones of each cell type were readily generated, aligned and checked for mismatches. Several commonly used alternative procedures were compared for cloning efficiency and size-fidelity of inserted DNA fragments.

Amino Acid Sequence

A new method that simultaneously aligns and reconstructs ancestral sequences for any number of homologous sequences, when the phylogeny is given.

Among the fundamental problems in molecular evolution and in the analysis of homologous sequences are alignment, phylogeny reconstruction, and the reconstruction of ancestral sequences. This paper presents a fast, combined solution to these problems. The new algorithm gives an approximation to the minimal history in terms of a distance function on sequences. The distance function on sequences is a minimal weighted path length constructed from substitutions and insertions-deletions of segments of any length. Substitutions are weighted with an arbitrary metric on the set of nucleotides or amino acids, and indels are weighted with a gap penalty function of the form gk = a + (bxk), where k is the length of the indel and a and b are two positive numbers. A novel feature is the introduction of the concept of sequence graphs and a generalization of the traditional dynamic sequence comparison algorithm to the comparison of sequence graphs. Sequence graphs ease several computational problems. They are used to represent large sets of sequences that can then be compared simultaneously. Furthermore, they allow the handling of multiple, equally good, alignments, where previous methods were forced to make arbitrary choices. A program written in C implemented this method; it was tested first on 22 5S RNA sequences.

Algorithms

Sequence analysis of firefly luciferase family reveals a conservative sequence motif.

A conservative sequence motif was extracted from an alignment of the firefly luciferase family. AngR derived from a pathogenic bacterium and acetyl-CoA synthetases derived from two ascomycete fungi were identified as members of the firefly luciferase family by a homology search with the motif and other sequence comparison analyses. The motif sequence shares several characteristics with the phosphate-binding sites of phosphoproteins and nucleotide-binding proteins. A multiple alignment and an unrooted phylogenetic tree were constructed for the investigation of evolutionary relationships within the firefly luciferase family.

Amino Acid Sequence

PATMAT: a searching and extraction program for sequence, pattern and block queries and databases.

A program has been developed that provides molecular biologists with multiple tools for searching databases, yet uses a very simple interface. PATMAT can use protein or (translated) DNA sequences, patterns or blocks of aligned proteins as queries of databases consisting of amino acid or nucleotide sequences, patterns or blocks. The ability to search databases of blocks by 'on-the-fly' conversion to scoring matrices provides a new tool for detection and evaluation of distant relationships. PATMAT uses a pull-down, menu-driven interface to carry out its multiple searching, extraction and viewing functions. Each query or database type is recognized, reported, and the appropriate search carried out, with matches and alignments reported in windows as they occur. Any of the high scoring matches can be exported to a file, viewed and recalled as a query using only a few keystrokes or mouse selections. Searches of multiple database files are carried out by user selection within a window. PATMAT runs under DOS; the searching engine also runs under UNIX.

Amino Acid Sequence

Paramyosin gene (unc-15) of Caenorhabditis elegans. Molecular cloning, nucleotide sequence and models for thick filament structure.

Paramyosin is a major structural component of thick filaments isolated from many invertebrate muscles. The Caenorhabditis elegans paramyosin gene (unc-15) was identified by screening with specific antibodies an "exon-expression" library containing lacZ/nematode gene fusions. Short probes recovered from the library were used to identify bacteriophage lambda and cosmid clones that encompass the entire paramyosin (unc-15) gene. From these clones, numerous subclones containing epitopes reacting with anti-paramyosin sera were obtained, providing strong evidence that the initial cloned fragment was, in fact, derived from the structural gene for paramyosin. The complete nucleotide sequence of a 12 x 10(3) base-pair region spanning the gene was obtained. The gene is composed of ten short exons encoding a protein of 866 [corrected] amino acid residues. Paramyosin is highly similar to residues 267 to 1089 of myosin heavy chain rods. For most of its length, paramyosin appears to form an alpha-helical coiled-coil and shows the expected heptad repeat of hydrophobic amino acid residues and the 28-residue repeat of charged amino acids characteristic of myosin heavy chain rods. However, paramyosin differs from myosin in having non-helical extensions at both the N and C termini and an additional "skip" residue that interrupts the 28-residue repeat. The distribution of charges along the length of the paramyosin rod is also significantly different from that of myosin heavy chain rods. Potential charge-mediated interactions between paramyosin rods and between paramyosin and myosin rods were calculated using a model successfully applied previously to the analysis of the myosin rod sequences. Myosin rods aligned in parallel show optimal charge-charge interactions at multiples of 98 residue staggers (i.e. at axial displacements of multiples of 143 A). Paramyosin rods, in contrast, appear to interact optimally at parallel staggers of 493 residues (i.e. at axial displacements of 720 A) but show only weak interaction peaks at 98 or 296 residues. Similar calculations suggest optimal interactions between paramyosin molecules and myosin rods and in their anti-parallel alignments. The implications of these results for the structure of the bare zone and the assembly of nematode thick filaments are discussed.

Animals

Cooperation of transposable elements to endow global networks of initiators of hybrid assembly pathways of endogenous multiprotein complexes.

Mechanisms governing initiation steps of the assembly of endogenous multi-protein complexes (EMC) remain incompletely understood. Here, multiple lines of observations are reported describing the function-aligned initiation sequence of hybrid assembly pathways (HAP) of EMC. The first step of HAP-guided chain reactions of protein-protein interactions (PPI) of EMC assemblies constitutes the creation of cell type-specific pools of hetero and homo dimers. The molecular anatomy of HAP was elucidated by defining qualitative and quantitative characteristics of protein binding to a compendium of 200,393 distinct genomic regulatory elements (GRE), including 49,667 sequences representing control sets of genomic loci as well as 150,726 GRE of different evolutionary origins. The consensus sequence of HAP actions consists of: a) Initiation on genomic DNA of the formation of metastable hetero- and homodimers of EMCs' protein constituents; b) Release of dimers from DNA templates for delivery to the EMC assembly compartments; c) Assembly of defined EMC by sequential on demand addition of proteins to preformed dimers serving as attractors of EMC-specific ensembles of monomers. Chromosome-naïve DNA scaffolds facilitating creation of intracellular dimer pools engage networks of ~700 transcription factors (TFs), 534 of which manifest region-specific patterns of significantly enriched expression in 1358 brain regions. HAP initiators appear to operate within nucleosome-depleted islands of transposable elements (TE) - derived sequences within heterochromatin. PPI assembly lines of EMCs operate in 2 concurrent modes: TF-TF PPI cascade and PPI HUB protein cascade. Regardless of the number of DNA-bound initiator TFs (ranging from one to 716 TFs), both modes of operations reached the equilibrium at the PPI constituents saturation levels of ~245 proteins for TF-TF PPI modes and of ~351 proteins for PPI HUB protein modes. Distinct panels of DNA-bound initiator TFs and proteins of PPI cascade ensembles are enriched in either defined sets of neuroanatomical structures (TF-TF mode) or among structural-functional constituents of synapses (HUB proteins mode). Thus, these bifurcated cascades appear biologically congruent: TF-TF constituents map to transcriptional signatures of hundreds of brain regions, whereas HUB constituents map to synaptogenesis and synaptic structures, suggesting the unified logic of genomic functions coordinating region identity and connectivity. Evidence-supported examples of default operations of PPI-guided assemblies of hetero- and homodimers of Yamanaka factors, neurogenesis constituents, and protein components of postsynaptic density of excitatory and inhibitory synaptogenesis are reported with detailed analytical focus on human Claustrum. The foundational set of observations reported in this contribution should facilitate experimental and theoretical explorations of TE-seeded genomic codes for initiators of PPI chain reactions of protein dimerization creating pools of attractors to guide and accelerate the EMC assemblies.

Humans

Sequence homology and absence of mRNA defines a possible pseudogene member of the Trypanosoma cruzi gp85/sialidase multigene family.

A genomic clone, pTt21, containing DNA apparently transcribed specifically in Trypanosoma cruzi trypomastigotes, was obtained by differentially screening a genomic library with trypomastigote and epimastigote cDNA. This 3444-bp clone contained open reading frames at each end, separated by a 1.8-kb non-coding region. The translated polypeptide from the 3' open reading frame (ORF2) of 1037 bp had 25-30% identity with 5 recently published T. cruzi gp85/sialidase sequences, and 20-25% identity with bacterial sialidases. Rabbit antiserum raised against an Escherichia coli fusion protein derived from the 5' open reading frame (ORF1) identified a surface antigen of 160 kDa, specifically expressed in trypomastigotes. A probe containing the first 211 bp from ORF1 was used to obtain a complete copy (c1821) of a gene that was closely related to ORF1, and encoded another member of the gp85/sialidase family. c1821 encodes a protein of 897 amino acids, but assignment of the N-terminus of the polypeptide was not possible. The 5'-most start codon is an unfavourable context to act as a translation initiator, it does not align with the initiator methionines of other gp85/sialidase sequences, nor is it followed by a signal peptide sequence characteristically found in other gp85/sialidase sequences. Although homology with the 5' ends of other gp85/sialidase sequences decays towards the 5' end of c1821, alignment of c1821 with 4 other gp85/sialidases indicated that the coding sequence should extend upstream at least 160 amino acids. In this region of c1821 there are multiple stop codons in each frame. The presence of the stop codons, the alignment data and our inability to amplify reverse transcribed mRNA using four internal primers, suggest that c1821 may not be present as a mature mRNA and is a pseudogene. Comparison of the apparently non-repetitive 3' coding domain of c1821 with the corresponding repetitive domains of two other members of the gp85/sialidase family revealed a high degree of similarity in nucleotide but not in amino acid sequence, and c1821 may thus represent an evolutionary intermediate between sub-families of the gp85/sialidase superfamily.

Amino Acid Sequence

Structure and expression of the human apolipoprotein A-IV gene.

We have isolated the human apolipoprotein (apo) A-IV gene from a cosmid library and determined its complete nucleotide sequence. The gene contains three exons of 162, 127, and 1180 nucleotides separated by two introns of 357 and 777 nucleotides. A sequence polymorphism has been identified in the 3' noncoding portion of the third exon. The human apoA-IV gene lacks an intron in the area encoding the 5' nontranslated region of its mRNA, which distinguishes it from all the other human apolipoprotein genes whose sequences are known. Comparison matrix analysis of the human apoA-IV gene sequence revealed evidence for an ancestral 11-nucleotide repeat unit that spans the third exon. These repeated sequences are much more highly conserved than those present in either rat apoA-IV or in any other human apolipoprotein. Optimal alignments of the 5' flanking regions of the rat and human apoA-IV genes disclosed multiple deletions in the rat sequence as well as a highly conserved region of 90 nucleotides (90% sequence identity) located within 170 nucleotides of the start site of transcription. The 5' flanking regions of the human and rat apoA-IV genes were ligated to the bacterial chloramphenicol acetyltransferase gene, then transfected into different cultured cells. The apoA-IV gene sequences elicited preferential expression of chloramphenicol acetyltransferase activity when introduced into intestinally derived Caco-2 cells and liver-derived Hep-G2 cells, consistent with the tissue specificity of the native gene. Analysis of deletion mutants of the human apoA-IV 5' flanking region indicated that regions from -293 to -233 and from -127 to -60 upstream of the transcription start site contain sequences required for maximum gene expression. These findings on the structure and expression of rat and human apoA-IV should prove useful in studying the control of the apoA-IV gene.

Amino Acid Sequence

DNA sequence comparisons of the human, mouse, and rabbit immunoglobulin kappa gene.

A comparative analysis between human, mouse, and rabbit immunoglobulin (Ig) kappa-gene DNA sequences is presented. New formulas for determining the expected length and variance of the longest block identity (a succession of matching nucleotides) between multiple random sequences are given and are used to establish statistical criteria for ascertaining the significance of block identities shared in r out of s sequences. The statistically significant block identities within and between the Ig-kappa-gene sequences are ascertained, and alignment maps based on these similarities are constructed. The human and rabbit sequences (especially in the noncoding regions) and the human and mouse sequences (on the coding regions) show a similarity much stronger than that between the mouse and rabbit sequences. The existence of several highly significant shared oligonucleotides occurring in alignment with each other or with respect to the J- and C-gene segments suggests a configuration of multiple control sites. Discussion and interpretations of the form and distribution of the block identities are given.

Animals

Primary structure of carboxypeptidase T: delineation of functionally relevant features in Zn-carboxypeptidase family.

The primary structure of carboxypeptidase T--a Zn-dependent extracellular enzyme of Thermoactinomyces vulgaris--was determined from the cloned cpT gene nucleotide sequence and compared to Zn-carboxypeptidases from various organisms. The compilation and analysis of multiple alignment accompanied by consideration of available tertiary structure data have shown that in the overall spatial structure and active site arrangement CpT is similar to other enzymes constituting the Zn-carboxypeptidase family. Nine of 16 amino acid residues found to be strictly invariant are presumably located close to the active site. The preservation of His69, Glu72, Asn144, Arg145, His196, Tyr248, and Glu270 identified previously as essential catalytic site participants implicates basically the same catalytic mechanism in the Zn-carboxypeptidase family. It is proposed that Pro205 and Asp256 should play an important role in proper S1'-pocket spatial arrangement. The comparative analysis of amino acid variations in S1'-pocket enabled us to reveal structural determinants of the Zn-carboxypeptidase primary specificity. The relatively reduced size of the pocket and negative charge of Asp253 are supposed to contribute correspondingly to A- and B-type substrate preferences of carboxypeptidase T endowed with dual primary specificity.

Amino Acid Sequence

Definition of general topological equivalence in protein structures. A procedure involving comparison of properties and relationships through simulated annealing and dynamic programming.

A protein is defined as an indexed string of elements at each level in the hierarchy of protein structure: sequence, secondary structure, super-secondary structure, etc. The elements, for example, residues or secondary structure segments such as helices or beta-strands, are associated with a series of properties and can be involved in a number of relationships with other elements. Element-by-element dissimilarity matrices are then computed and used in the alignment procedure based on the sequence alignment algorithm of Needleman & Wunsch, expanded by the simulated annealing technique to take into account relationships as well as properties. The utility of this method for exploring the variability of various aspects of protein structure and for comparing distantly related proteins is demonstrated by multiple alignment of serine proteinases, aspartic proteinase lobes and globins.

Amino Acid Sequence

Hierarchical method to align large numbers of biological sequences.

The method presented here is intended as a compromise between finding a good overall alignment and the time taken to do so. Many multiple alignment algorithms spend an excessively large amount of effort trying to find the best global alignment. This time is often ill spent because the results of the standard dynamic programming alignment algorithm are dominated by the choice of gap penalty and the form of the score matrix, both of which have a poor theoretical foundation. Nonetheless, it is important that savings in time do not compromise the quality of the alignment. By using the consensus sequence approach, this danger is largely avoided as the conserved features of the sequences are quickly identified and preserved through further cycles. In the alignment of existing alignments, which is one of the more novel aspects of the method, each alignment was treated as an averaged consensus sequence with gaps making no contribution. This gives rise to the advantageous property that gaps will have a greater propensity to be inserted where there are already gaps and is equivalent to a local change in the gap penalty. This type of behavior represents a transition away from the homogeneous scoring schemes used in aligning two sequences toward a scoring scheme that depends on position in the sequence. The alignment of consensus sequences thus forms a bridge between simple pair alignment and the alignment of discrete patterns in which sequence features and allowed gap locations are exaggerated. To complete this transition the program described above has been integrated into the earlier pattern matching (template) program. Such templates can reliably locate sequence similarities that are too weak or scattered to be found by the more standard alignment methods and should therefore produce a further condensation of the sequence data bank. Only by continually extending our knowledge of the relationships between sequences to increasingly distant similarities can we hope to avoid being overwhelmed by the increasing amount of data.

Algorithms

Sequence-directed mutagenesis: evidence from a phylogenetic history of human alpha-interferon genes.

We have studied the potential contribution of template-dependent events to genetic variation in mammals by examining the sequence alterations that have occurred in the recent evolution of human interferon genes. Fifteen members of the human alpha-interferon gene family were aligned, and a phylogenetic history was inferred. Many multiple events are inferred to have occurred in the evolution of the interferon genes and for the majority of these local DNA sequences were present that were capable of serving as templates for their occurrence. We conclude that the DNA sequence has the potential to explain many of the inferred spontaneous events and to explain complex alterations to sequences--i.e., the joint occurrence of base substitutions and insertions/deletions. Thus, such a mechanism would often cause multiple sequence changes as a result of a single mutational event and would provide additional genetic variation for evolution. Sequence-directed mutations would depend upon the local DNA sequences and, hence, would not be random at the DNA level.

Base Sequence

tRNA-rRNA sequence homologies: evidence for an ancient modular format shared by tRNAs and rRNAs.

Homologies between tRNAs and rRNAs are identified in searches using various combinations of Escherichia coli, yeast, Halobacterium volcanii and bovine mitochondrial sequences. As in previously reported comparisons, the homologies are too frequent and long to be attributed to coincidence, and similar frequencies from inter- and intraspecies comparisons preclude evolutionary convergence as an explanation. In contrast to the earlier studies, patterns in the positioning of the homologies are now described. Graphing the positions of the homologies along orthogonal axes that represent numbers of bases in tRNA and rRNA shows recurring patterns in the alignments. Preferred spacings of integral multiples of 9 bases are found, suggesting a periodicity in the ancestral structure from which the tRNAs and rRNAs were derived. The periodicity also suggests persistence of a modular format in both classes of molecules that survived changes in sequence that occurred during evolution. A model is proposed for the generation of the ancestral molecule and the early evolution of the coding mechanism. Elongation by self-priming and self-templating gave a hairpin with a 9 base stem. Two additional cycles gave a 70-80 base tRNA-like structure. Additional cycles yielded a tandem repeat of this unit, roughly equivalent in size to the combined rRNAs of prokaryotes. The larger RNA would contain the information and materials for generating the smaller RNAs. It is proposed that multiple recombination among such molecules gave composite structures, presumed progenitors of today's t- and rRNAs. The distribution of the conserved domains among today's species argues for the existence of the ancestral molecule prior to divergence of lines leading to the various kingdoms. Their presence in the different nucleic acids suggests the existence of a nucleic acid with multiple functions prior to partitioning of these functions among the nucleic acids that exist today. The occurrence of overlaps, overlays and consensus alignments among the homologies provides the means for identifying contiguous and neighboring conserved regions and holds promise for the reconstruction of the sequence of an ancestral molecule.

Animals

Multiple alignment using simulated annealing: branch point definition in human mRNA splicing.

A method for the simultaneous alignment of a very large number of sequences using simulated annealing is presented. The total running time of the algorithm does not depend explicitly on the number of sequences treated. The method has been used for the simultaneous alignment of 1462 human intron sequences upstream of the intron-exon boundary. The consensus sequence of the aligned set together with a calculation of the Shannon information clearly shows that several sequence motives are conserved: (i) a previously undetected guanosine rich region, (ii) the branch point and (iii) the polypyrimidine tract. The nucleotide frequencies at each position of the branch point consensus sequence qualitatively reproduce the frequencies of the experimentally determined branch points.

Algorithms