PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Pairwise alignment of the DNA sequence using hypercomplex number representation.

A new set of DNA base-nucleic acid codes and their hypercomplex number representation have been introduced for taking the probability of each nucleotide into full account. A new scoring system has been proposed to suit the hypercomplex number representation of the DNA base-nucleic acid codes and incorporated with the method of dot matrix analysis and various algorithms of sequence alignment. The problem of DNA sequence alignment can be processed in a rather similar way to pairwise alignment of the protein sequence.

Base Sequence↗

Simultaneous and multivariate alignment of protein sequences: correspondence between physicochemical profiles and structurally conserved regions (SCR).

A general protein sequence alignment methodology for detecting a priori unknown common structural and functional regions is described. The method proposed in this paper is based on two basic requirements for a meaningful alignment. First, each sequence or segment of a sequence is characterized by a multivariate physicochemical profile. Second, the alignment is performed by considering all the sequences simultaneously, and the algorithm detects those regions that form a set of similar profiles. In order to test the structural meaning of the alignment obtained from the sequences, quantitative comparisons are performed with structurally conserved regions (SCR) determined from the X-ray structures of three serine proteases. Results suggest that the limits of the SCR may be predicted from the similarities between the physicochemical profiles of the sequences. The procedures are not completely automated. The final step requires a visual screening of alternative pathways in order to determine an optimal alignment.

Algorithms↗

Modelling of binding sites of the nicotinic acetylcholine receptor and their relation to models of the whole receptor.

Models for the acetylcholine (ACh)-binding site of the nicotinic acetylcholine receptor (nAChR) are proposed. These models have been developed by using the concept of the ligand-gated ion-channel (LGIC) superfamily of receptors that have evolved from a common ancestor. An initial component of the binding site was identified as a highly conserved 15-residue stretch of primary structure in the N-terminal extracellular region of all known LGIC subunits, based on aligned sequence data of LGICs. This subregion, termed the Cys-loop, was modelled as an amphiphilic beta-hairpin and we propose that it forms a major determinant of the binding cleft for agonists. This initial, partial binding-site model has been extended to include residues biochemically identified as spatially adjacent to the binding cleft. A recently developed technique for rapidly scanning the known protein structural database for 'non-homologous similarity' using just sequence information identified the known structure of the enzyme pyrophosphatase (PPase) as a candidate scaffold for the N-terminal domain of the nAChR. This similarity was investigated further using sequence alignments. A framework model of the full N-terminal domain in which the position of the Cys-loop and other binding-site determinants, as well as the main immunogenic region (MIR), have been mapped on to the PPase structure.

Amino Acid Sequence↗

An Eulerian path approach to global multiple alignment for DNA sequences.

With the rapid increase in the dataset of genome sequences, the multiple sequence alignment problem is increasingly important and frequently involves the alignment of a large number of sequences. Many heuristic algorithms have been proposed to improve the speed of computation and the quality of alignment. We introduce a novel approach that is fundamentally different from all currently available methods. Our motivation comes from the Eulerian method for fragment assembly in DNA sequencing that transforms all DNA fragments into a de Bruijn graph and then reduces sequence assembly to a Eulerian path problem. The paper focuses on global multiple alignment of DNA sequences, where entire sequences are aligned into one configuration. Our main result is an algorithm with almost linear computational speed with respect to the total size (number of letters) of sequences to be aligned. Five hundred simulated sequences (averaging 500 bases per sequence and as low as 70% pairwise identity) have been aligned within three minutes on a personal computer, and the quality of alignment is satisfactory. As a result, accurate and simultaneous alignment of thousands of long sequences within a reasonable amount of time becomes possible. Data from an Arabidopsis sequencing project is used to demonstrate the performance.

Algorithms↗

Purification, characterization, and primary structure of four depressant insect-selective neurotoxin analogs from scorpion (Buthus sindicus) venom.

Four depressant insect-selective neurotoxin analogs (termed Bs-dprIT1 to 4) from the venom of the scorpion Buthus sindicus were purified to homogeneity in a single step using reverse-phase HPLC. The molecular masses of the purified toxins were 6820.9, 6892.4, 6714.7, and 6657.1 Da, respectively, as determined by mass spectrometry. These long-chain neurotoxins were potent against insects with half lethal dose values of 67, 81, 103, and 78 ng/100 mg larva and 138, 160, 163, and 142 ng/100 mg cockroach, respectively, but were not lethal to mice even at the highest applied dose of 10 microg/20 g mouse. When injected into blowfly larvae (Sarcophaga falculata), Bs-dprIT1 to 4 induced classical manifestations of depressant toxins, i.e., a slow depressant flaccid paralysis. The primary structures of Bs-dprIT 1 to 4 revealed high sequence homology (60-75%) with other depressant insect toxins isolated from scorpion venoms. Despite the high sequence conservation, Bs-dprIT1 to 4 showed some remarkable features such as (i) the presence of methionine (Met(6) in Bs-dprIT1 and Met(24) in Bs-dprIT2 to 4) and histidine (His(53) and His(57) in Bs-dprIT1) residues, i.e., amino acid residues that are uncommon to this type of toxin; (ii) the substitution of two highly conserved tryptophan residues (Trp43 --> Ala and Trp53 --> His) in the sequence of Bs-dprIT1; and (iii) the occurrence of more positively charged amino acid residues at the C-terminal end than in other depressant insect toxins. Multiple sequence alignment, sequence analysis, sequence-based structure prediction, and 3D homology modeling studies revealed a protein fold and secondary structural elements similar to those of other scorpion toxins affecting sodium channel activation. The electrostatic potential calculated on the surface of the predicted 3D model of Bs-dprIT1 revealed a significant positive patch in the region of the toxin that is supposed to bind to the sodium channel.

Amino Acid Sequence↗

The important role of residue F268 in ligand binding by LXRbeta.

Liver X receptors (LXRs) are nuclear receptors that regulate the metabolism of cholesterol and bile acids. Despite information on the specificity of their natural ligands, oxysterols, relatively little is known about the ligand binding site in LXRs. The helix 3 region in the ligand binding domain (LBD) of peroxisome proliferator-activated receptors (PPARs) has been implicated in ligand entry. Sequence alignment of LXRs, farnesoid X receptor (FXR), and PPARs identified the corresponding helix 3 region in the LXRbeta LBD. Residues F268 and T272, which are conserved in all the aligned sequences and only in LXRs and FXR, respectively, were replaced with alanine. The effects of these mutations on ligand binding and receptor activation were examined using an in vitro ligand binding assay and a cell based reporter assay, respectively. The LXRbeta mutant F268A did not bind ligand. In contrast, conversion of T272 to alanine has no effect on ligand binding. By transiently expressing a chimeric receptor containing Escherichia coli tetracycline repressor (TetR) and LXRbeta LBD and a reporter with a TetR binding site, we show that mutant F268A lost the ability to activate transcription of the reporter, whereas mutant T272A still has an activity similar to that of the wild-type LXRbeta. These data, consistent with the findings in the in vitro ligand binding assay and our 3D modeling, are the first study that identifies a residue critical for ligand binding in LXRbeta.

Amino Acid Sequence↗

Libraries of hybrid proteins from distantly related sequences.

We introduce a method for sequence homology-independent protein recombination (SHIPREC) that can create libraries of single-crossover hybrids of unrelated or distantly related proteins. The method maintains the proper sequence alignment between the parents and introduces crossovers mainly at structurally related sites distributed over the aligned sequences. We used SHIPREC to create a library of interspecies hybrids of a membrane-associated human cytochrome P450 (1A2) and the heme domain of a soluble bacterial P450 (BM3). By fusing the hybrid gene library to the gene for chloramphenicol acetyl transferase (CAT), we were able to select for soluble and properly folded protein variants. Screening for 1A2 activity (deethylation of 7-ethoxyresorufin) identified two functional P450 hybrids that were more soluble in the bacterial cytoplasm than the wild-type 1A2 enzyme.

Amino Acid Sequence↗

A special-purpose processor for gene sequence analysis.

Advances in computational biology have occurred primarily in the areas of software and algorithm development; new designs of hardware to support biological computing are extremely scarce. This is due, we believe, to the presence of a non-trivial knowledge gap between molecular biologists and computer designers. The existence of this gap is unfortunate, as it has long been known that for certain problems, special-purpose computers can achieve significant cost/performance gains over general-purpose machines. We describe one such computer here: a custom accelerator for gene sequence analysis. The accelerator implements a version of the Needleman-Wunsch algorithm for nucleotide sequence alignment. Sequence lengths are constrained only by available memory; the product of sequence lengths in the current implementation can be up to 2(22). The machine is implemented as two NuBus boards connected to a Mac IIf/x, using a mixture of TTL and FPGA technology clocked at 10 MHz. The boards are completely functional, and yield a 15-fold performance improvement over an unassisted host.

Algorithms↗

CVTree: a phylogenetic tree reconstruction tool based on whole genomes.

Composition Vector Tree (CVTree) implements a systematic method of inferring evolutionary relatedness of microbial organisms from the oligopeptide content of their complete proteomes (http://cvtree.cbi.pku.edu.cn). Since the first bacterial genomes were sequenced in 1995 there have been several attempts to infer prokaryote phylogeny from complete genomes. Most of them depend on sequence alignment directly or indirectly and, in some cases, need fine-tuning and adjustment. The composition vector method circumvents the ambiguity of choosing the genes for phylogenetic reconstruction and avoids the necessity of aligning sequences of essentially different length and gene content. This new method does not contain 'free' parameter and 'fine-tuning'. A bootstrap test for a phylogenetic tree of 139 organisms has shown the stability of the branchings, which support the small subunit ribosomal RNA (SSU rRNA) tree of life in its overall structure and in many details. It may provide a quick reference in prokaryote phylogenetics whenever the proteome of an organism is available, a situation that will become commonplace in the near future.

Algorithms↗

A 3D model of the delta opioid receptor and ligand-receptor complexes.

A model for the 3D structure of the transmembrane domain of the delta opioid receptor was predicted from the sequence divergence analysis of 42 sequences of G-protein coupled peptide hormone receptors belonging to the opioid, somatostatin and angiotensin receptor families. No template was used in the prediction steps, which include multiple sequence alignment, calculation of a variability profile of the aligned sequences, use of the variability profile to identify the boundaries of transmembrane regions, prediction of their secondary structure, optimization of the packing shape in a helix bundle, prediction of side chain conformations and structural refinement. The general shape of the model is similar to that of the low resolution rhodopsin structure in that the TM3 and TM7 helices are most buried in the bundle and the TM1 and TM4 helices are most exposed to the lipid phase. An initial assessment of this model was made by determining to what extent a binding site identified using four structurally disparate high affinity delta opioid ligands was consistent with known mutational studies. With the assumption that the protonated amine nitrogen, a feature common to all delta opioid ligands, interacts with the highly conserved Asp127 in TM3, a pocket was found that satisfied the criteria of complementarity to the requirements for receptor recognition for these four diverse ligands, two delta selective antagonists (the fused ring naltrindole and the peptide Tyr-Tic-Phe-Phe-NH2) and the two agonists lofentanil and BW373U86 deduced from previous studies of the ligands alone. These ligands could be accommodated in a similar region of the receptor. The receptor binding site identified in the optimized complexes contained many residues in positions known to affect ligand binding in G-protein coupled receptors. These results also allowed identification of key residues as candidates for point mutations for further assessment and refinement of this model as well as preliminary indications of the requirements for recognition of this receptor.

Amino Acid Sequence↗

Structure-based function inference using protein family-specific fingerprints.

We describe a method to assign a protein structure to a functional family using family-specific fingerprints. Fingerprints represent amino acid packing patterns that occur in most members of a family but are rare in the background, a nonredundant subset of PDB; their information is additional to sequence alignments, sequence patterns, structural superposition, and active-site templates. Fingerprints were derived for 120 families in SCOP using Frequent Subgraph Mining. For a new structure, all occurrences of these family-specific fingerprints may be found by a fast algorithm for subgraph isomorphism; the structure can then be assigned to a family with a confidence value derived from the number of fingerprints found and their distribution in background proteins. In validation experiments, we infer the function of new members added to SCOP families and we discriminate between structurally similar, but functionally divergent TIM barrel families. We then apply our method to predict function for several structural genomics proteins, including orphan structures. Some predictions have been corroborated by other computational methods and some validated by subsequent functional characterization.

Bacterial Proteins↗

Domain structure and three-dimensional model of a group II intron-encoded reverse transcriptase.

Group II intron-encoded proteins (IEPs) have both reverse transcriptase (RT) activity, which functions in intron mobility, and maturase activity, which promotes RNA splicing by stabilizing the catalytically active RNA structure. The LtrA protein encoded by the Lactococcus lactis Ll.LtrB group II intron contains an N-terminal RT domain, with conserved sequence motifs RT1 to 7 found in the fingers and palm of retroviral RTs; domain X, associated with maturase activity; and C-terminal DNA-binding and DNA endonuclease domains. Here, partial proteolysis of LtrA with trypsin and Arg-C shows major cleavage sites in RT1, and between the RT and X domains. Group II intron and related non-LTR retroelement RTs contain an N-terminal extension and several insertions relative to retroviral RTs, some with conserved features implying functional importance. Sequence alignments, secondary-structure predictions, and hydrophobicity profiles suggest that domain X is related structurally to the thumb of retroviral RTs. Three-dimensional models of LtrA constructed by "threading" the aligned sequence on X-ray crystal structures of HIV-1 RT (1) account for the proteolytic cleavage sites; (2) suggest a template-primer binding track analogous to that of HIV-1 RT; and (3) show that conserved regions in splicing-competent LtrA variants include regions of the RT and X (thumb) domains in and around the template-primer binding track, distal regions of the fingers, and patches on the protein's back surface. These regions potentially comprise an extended RNA-binding surface that interacts with different regions of the intron for RNA splicing and reverse transcription.

Amino Acid Sequence↗

[An introduction of several programs used in genomic analysis].

Genomics is a novel subject that has been developed accompanying with the progress of human genome project. Genomics deals with the chemistry component, structure organization and evolution of genome at global level. As genomics associated with huge data, bioinformatics plays an important role in these processes of data production, data management and data mining. At present, many reliable programs have been used in genomic research successfully, which are usually accessible and downloaded freely. We address here the principles of some programs used wildly in genomics such as sequence alignment, sequence assembly, repeat identification and gene prediction, which are exemplified with typical programs respectively.

English Abstract↗

The Jalview Java alignment editor.

Multiple sequence alignment remains a crucial method for understanding the function of groups of related nucleic acid and protein sequences. However, it is known that automatic multiple sequence alignments can often be improved by manual editing. Therefore, tools are needed to view and edit multiple sequence alignments. Due to growth in the sequence databases, multiple sequence alignments can often be large and difficult to view efficiently. The Jalview Java alignment editor is presented here, which enables fast viewing and editing of large multiple sequence alignments.

Algorithms↗

A weighting system and algorithm for aligning many phylogenetically related sequences.

Most multiple sequence alignment programs explicitly or implicity try to optimize some score associated with the resulting alignment. Although the sum-of-pairs score is currently most widely used, it is inappropriate when the phylogenetic relationships among the sequences to be aligned are not evenly distributed, since the contributions of densely populated groups dominate those of minor members. This paper proposes an iterative multiple sequence alignment method which optimizes a weighted sum-of-pairs score, in which the weights given to individual sequence pairs are adjusted to compensate for the biased contributions. A simple method that rapidly calculates such a set of weights for a given phylogenetic tree is presented. The multiple sequence alignment is refined through partitioning and realignment restricted to the edges of the tree. Under this restriction, profile-based fast and rigorous group-to-group alignment is achieved at each iteration, rendering the overall computational cost virtually identical to that using an unweighted score. Consistency of nearly 90% was attained between structural and sequence alignments of multiple divergent globins, confirming the effectiveness of this strategy in improving the quality of multiple sequence alignment.

Algorithms↗

RevTrans: Multiple alignment of coding DNA from aligned amino acid sequences.

The simple fact that proteins are built from 20 amino acids while DNA only contains four different bases, means that the 'signal-to-noise ratio' in protein sequence alignments is much better than in alignments of DNA. Besides this information-theoretical advantage, protein alignments also benefit from the information that is implicit in empirical substitution matrices such as BLOSUM-62. Taken together with the generally higher rate of synonymous mutations over non-synonymous ones, this means that the phylogenetic signal disappears much more rapidly from DNA sequences than from the encoded proteins. It is therefore preferable to align coding DNA at the amino acid level and it is for this purpose we have constructed the program RevTrans. RevTrans constructs a multiple DNA alignment by: (i) translating the DNA; (ii) aligning the resulting peptide sequences; and (iii) building a multiple DNA alignment by 'reverse translation' of the aligned protein sequences. In the resulting DNA alignment, gaps occur in groups of three corresponding to entire codons, and analogous codon positions are therefore always lined up. These features are useful when constructing multiple DNA alignments for phylogenetic analysis. RevTrans also accepts user-provided protein alignments for greater control of the alignment process. The RevTrans web server is freely available at http://www.cbs.dtu.dk/services/RevTrans/.

Amino Acid Substitution↗

A simple method to generate non-trivial alternate alignments of protein sequences.

A major problem in sequence alignments based on the standard dynamic programming method is that the optimal path does not necessarily yield the best equivalencing of residues assessed by structural or functional criteria. An algorithm is presented that finds suboptimal alignments of protein sequences by a simple modification to the standard dynamic programming method. The standard pairwise weight matrix elements are modified in order to penalize, but not eliminate, the equivalencing of residues obtained from previous alignments. The algorithm thereby yields a limited set of alternate alignments that can differ considerably from the optimal. The approach is benchmarked on the alignments of immunoglobulin domains. Without a prior knowledge of the optimal choice of gap penalty, one of the suboptimal alignments is shown to be more accurate than the optimal.

Algorithms↗

Evaluation and improvements in the automatic alignment of protein sequences.

The accuracy of protein sequence alignment obtained by applying a commonly used global sequence comparison algorithm is assessed. Alignments based on the superposition of the three-dimensional structures are used as a standard for testing the automatic, sequence-based methods. Alignments obtained from the global comparison of five pairs of homologous protein sequences studied gave 54% agreement overall for residues in secondary structures. The inclusion of information about the secondary structure of one of the proteins in order to limit the number of gaps inserted in regions of secondary structure, improved this figure to 68%. A similarity score of greater than six standard deviation units suggests that an alignment which is greater than 75% correct within secondary structural regions can be obtained automatically for the pair of sequences.

Algorithms↗