PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “alignment chaining method”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Estimation of reversible substitution matrices from multiple pairs of sequences.

We present a method for estimating the most general reversible substitution matrix corresponding to a given collection of pairwise aligned DNA sequences. This matrix can then be used to calculate evolutionary distances between pairs of sequences in the collection. If only two sequences are considered, our method is equivalent to that of Lanave et al. (1984). The main novelty of our approach is in combining data from different sequence pairs. We describe a weighting method for pairs of taxa related by a known tree that results in uniform weights for all branches. Our method for estimating the rate matrix results in fast execution times, even on large data sets, and does not require knowledge of the phylogenetic relationships among sequences. In a test case on a primate pseudogene, the matrix we arrived at resembles one obtained using maximum likelihood, and the resulting distance measure is shown to have better linearity than is obtained in a less general model.

Algorithms↗

Fragment ranking in modelling of protein structure. Conformationally constrained environmental amino acid substitution tables.

Conformationally constrained environment-dependent amino acid residue substitution tables have been constructed from a database comprising 33 homologous families of protein sequences aligned on the basis of their three-dimensional structures. Residues are allotted to one of 216 (or 54) classes of combinations of structural features. These include nine main-chain conformation classes, three classes of side-chain accessibility and eight (or two) classes of side-chain involvement in three types of hydrogen bond. Seven different main-chain conformational classes outside of regions of regular structure were identified in an analysis of the distributions of phi-psi torsion angles in 84 high-resolution crystallographic structures. Residue substitutions at equivalent positions in the structural alignments are included where the main-chain conformational class is conserved. Frequency data in the form of 216 (or 54) environment specific (20 x 20 residue type) matrices are then converted to probabilities. Two smoothing regimes incorporating entropy-driven weights were applied to the set of 54 tables. Predicted residue substitutions have been generated for individual residue positions in beta-hairpins and the hypervariable regions of the immunoglobulins. These have been compared with the observed sequence variation at the same positions using rank correlation methods. Measurements of chi 2 distances demonstrate the considerable improvement in predictive power at key residue positions identified from interactive graphics studies when compared to the Dayhoff MDM250 mutation matrix. An illustrative example is given of an application of the method in the ranking of loop fragments in model building studies of structurally variable regions in two subtilisins. A combined template scoring procedure is found to be 26-fold more discriminatory than the Dayhoff matrix. The success rate is approximately 85%.

Amino Acid Sequence↗

Hyaluronan synthase expression in bovine eyes.

PURPOSE: Hyaluronan (HA), a high-molecular-weight linear glycosaminoglycan, is a component of the extracellular matrix (ECM). It is expressed in eyes and plays important roles in many biologic processes, including cell migration, proliferation, and differentiation. Hyaluronan is produced by HA synthase (HAS), which has three isoforms: HAS1, HAS2, and HAS3. In this study, the HAS expression in the anterior segment of bovine eyes was investigated to determine the significance of HA in eyes. METHODS: To obtain bovine HAS probes, degenerate oligonucleotide primers, based on well-conserved amino acid sequences including the catalytic region of each HAS isoform, were used for reverse transcription-polymerase chain reaction to amplify mRNA from bovine corneal endothelial cells (BCECs). Hyaluronan synthase-1 expression in the anterior segment of bovine eyes at the protein level was investigated by immunohistochemistry. RESULTS: All three HAS isoforms were expressed in BCECs at the mRNA level. Amplified cDNA fragments of HAS1, HAS2, and HAS3 from BCECs can be aligned to human counterparts, showing similarities of 100%, 97.3%, and 100%, respectively, at the amino acid level. Hyaluronan synthase 1 was expressed at the protein level in corneal epithelium, keratocyte, corneal endothelium, conjunctival epithelium, ciliary epithelium, capillary endothelium, and trabecular meshwork. CONCLUSIONS: Hyaluronan synthase isoforms were expressed in the ocular anterior segment and are speculated to be involved in HA production in situ.

Amino Acid Sequence↗

Molecular phylogenetic study of Theileria sp. (Thung Song) based on the thymidylate synthetase gene.

Theileria type Thung Song is an indigenous hemoparasite of dairy cattle from the south of Thailand. It has previously been classified using the analysis of a comparative set of small subunit ribosomal RNA nucleotide sequences. However, the classification of this parasite is still questionable since the Theileria type Thung Song was located as the intermediate parasite between pathogenic and benign groups. We use the thymidylate synthetase gene (TS) as an alternative for the rapid molecular phylogenetic tree construction of benign Theileria type Thung Song, Theileria sergenti and Theileria buffeli with Babesia bovis as an out-group. The partial nucleotide sequences were determined using PCR, cloning and dideoxy sequencing. The TS nucleotide sequence data were aligned and analyzed by distance and maximum likelihood methods to construct the phylogenetic trees. Bootstrap analysis was used to test the strength of the different phylogenetic reconstructions. All tree-building methods gave similar results. This study shows that T. sergenti and T. buffeli are closely related whereas Theileria type Thung Song is more distantly related.

Animals↗

COACH: profile-profile alignment of protein families using hidden Markov models.

MOTIVATION: Alignments of two multiple-sequence alignments, or statistical models of such alignments (profiles), have important applications in computational biology. The increased amount of information in a profile versus a single sequence can lead to more accurate alignments and more sensitive homolog detection in database searches. Several profile-profile alignment methods have been proposed and have been shown to improve sensitivity and alignment quality compared with sequence-sequence methods (such as BLAST) and profile-sequence methods (e.g. PSI-BLAST). Here we present a new approach to profile-profile alignment we call Comparison of Alignments by Constructing Hidden Markov Models (HMMs) (COACH). COACH aligns two multiple sequence alignments by constructing a profile HMM from one alignment and aligning the other to that HMM. RESULTS: We compare the alignment accuracy of COACH with two recently published methods: Yona and Levitt's prof_sim and Sadreyev and Grishin's COMPASS. On two sets of reference alignments selected from the FSSP database, we find that COACH is able, on average, to produce alignments giving the best coverage or the fewest errors, depending on the chosen parameter settings. AVAILABILITY: COACH is freely available from www.drive5.com/lobster

Algorithms↗

The use of differential display to isolate viral genomic sequence for rapid development of PCR-based detection methods. A test case using Taura syndrome virus.

The purpose of this study was to explore the efficacy of using differential display (DD) to isolate viral genomic sequence using tissues from infected organisms so that a PCR procedure to detect the pathogen may be developed rapidly. The model virus used was the Taura syndrome virus (TSV), a ssRNA virus that cause high rates of mortality at shrimp farms. Two random primers in combination with four anchored primers were used to isolate five cDNAs, ranging in size from 241 to 822 bp, that were differentially expressed in TSV-infected shrimp (Litopenaeus vannamei). PCR experiments revealed that four of the five encoded shrimp genes while the fifth was likely to be a TSV gene. Evidence that the putative TSV sequence is part of the TSV genome was obtained by the 97% sequence identity it shared with the published TSV genome. PCR primers were designed successfully using the differential display sequence to develop a RT-PCR-based method to detect TSV. Because differential display does not require physical isolation of the virus and only a small amount of infected sample is needed, the technique may be useful as a method to isolate nucleic acid sequences from emerging pathogens so that PCR primers for their detection may be developed rapidly.

Animals↗

Simple species identification of Trichinella isolates by amplification and sequencing of the 5S ribosomal DNA intergenic spacer region.

We developed a PCR-based assay using a single primer pair to amplify the 5S ribosomal DNA intergenic spacer region to identify Trichinella isolates. In our method, amplified products are directly sequenced on both strands and compared to GenBank sequences. Using this method, we were able to identify Trichinella spiralis, T. britovi and T. nativa. This method permits rapid species identification of Trichinella isolates; however, further evaluation is required before recommending this approach for routine use.

Animals↗

Nativelike topology assembly of small proteins using predicted restraints in Monte Carlo folding simulations.

By incorporating predicted secondary and tertiary restraints derived from multiple sequence alignments into ab initio folding simulations, it has been possible to assemble native-like tertiary structures for a test set of 19 nonhomologous proteins ranging from 29 to 100 residues in length and representing all secondary structural classes. Secondary structural restraints are provided by the PHD secondary structure prediction algorithm that incorporates multiple sequence information. Multiple sequence alignments also provide predicted tertiary restraints via a two-step process: First, seed side chain contacts are selected from a correlated mutation analysis, and then an inverse folding algorithm expands these seed contacts. The predicted secondary and tertiary restraints are incorporated into a lattice-based, reduced protein model for structure assembly and refinement. The resulting native-like topologies exhibit a coordinate root-mean-square deviation from native for the whole chain between 3.1 and 6.7 A, with values ranging from 2.6 to 4.1 A over approximately 80% of the structure. Overall, this study suggests that the use of restraints derived from multiple sequence alignments combined with a fold assembly algorithm is a promising approach to the prediction of the global topology of small proteins.

Algorithms↗

CONRAD: a method for identification of variable and conserved regions within proteins by scale-space filtering.

Advanced sequencing techniques allow rapid deduction of individual amino acid sequences of highly related proteins. Due to their quasi-species nature, viral genomes (e.g. HIV-1) represent one of the most common sources of related proteins. Another example of related proteins are immunoglobulins. Local differences in amino acid conservation are useful indicators of potential domain structures and immunological or functional epitopes prior to structural analysis of proteins. Although variability indices can be calculated by several methods, delineation of boundaries between sequence stretches with similar variability indices is left to the user. We use algorithmic scale-space filtering for delineation of conserved and variable sequence stretches within a protein which is performed on an algorithmic basis avoiding arbitrary assignments. Out method correctly identified variable regions for the human immunoglobulin lambda-chain V-regions (subgroup I). Prediction of the variable regions of the HIV-1 gp120 env protein was in agreement with empirical derived definitions. These examples indicate that our method is useful for the regional assignment of protein variability solely on the basis of amino acid sequences.

Algorithms↗

Detection of betanodavirus in juvenile barramundi, Lates calcarifer (Bloch), by antigen capture ELISA.

Betanodavirus infection of fish has been responsible for mass mortalities in aquaculture hatcheries worldwide. Betanodaviruses possess a bipartite single-stranded RNA genome consisting of the 3.1 kb RNA1 encoding an RNA-dependent RNA polymerase and the B2 protein, while the 1.4 kb RNA2 encodes the viral nucleocapsid protein, alpha. A panel of six monoclonal antibodies against the alpha protein of greasy grouper nervous necrosis virus (GGNNV) was developed for use in diagnostics. All antibodies reacted with native and recombinant alpha in immunoblot and indirect immunofluorescence assays. Each of the monoclonal antibodies reacted against discrete regions of the alpha protein, though none reacted with the extreme C-terminal region of the protein. One of the monoclonal antibodies, specific for the K151-T246 region of alpha, was used for the development of an antigen capture ELISA. In this assay we could detect 10(3)-10(4) TCID(50) units of virus derived from infected tissue culture supernatants. Head tissue extracts prepared from experimentally infected barramundi, Lates calcarifer, juveniles were assayed for GGNNV using the antigen capture assay and a clear increase in alpha antigen was detected from 5 to 15 days post-challenge. The assay thus represents a useful method for field-based detection of betanodavirus in fish hatcheries.

Amino Acid Sequence↗

Predicting protein subcellular locations using hierarchical ensemble of Bayesian classifiers based on Markov chains.

BACKGROUND: The subcellular location of a protein is closely related to its function. It would be worthwhile to develop a method to predict the subcellular location for a given protein when only the amino acid sequence of the protein is known. Although many efforts have been made to predict subcellular location from sequence information only, there is the need for further research to improve the accuracy of prediction. RESULTS: A novel method called HensBC is introduced to predict protein subcellular location. HensBC is a recursive algorithm which constructs a hierarchical ensemble of classifiers. The classifiers used are Bayesian classifiers based on Markov chain models. We tested our method on six various datasets; among them are Gram-negative bacteria dataset, data for discriminating outer membrane proteins and apoptosis proteins dataset. We observed that our method can predict the subcellular location with high accuracy. Another advantage of the proposed method is that it can improve the accuracy of the prediction of some classes with few sequences in training and is therefore useful for datasets with imbalanced distribution of classes. CONCLUSION: This study introduces an algorithm which uses only the primary sequence of a protein to predict its subcellular location. The proposed recursive scheme represents an interesting methodology for learning and combining classifiers. The method is computationally efficient and competitive with the previously reported approaches in terms of prediction accuracies as empirical results indicate. The code for the software is available upon request.

Amino Acid Sequence↗

Identification of Anopheles (Nyssorhynchus) albitarsis complex species (Diptera: Culicidae) using rDNA internal transcribed spacer 2-based polymerase chain reaction primes.

Anopheles (Nyssorhynchus) marajoara is a proven primary vector of malaria parasites in Northeast Brazil, and An. deaneorum is a suspected vector in Western Brazil. Both are members of the morphologically similar Albitarsis Complex, which also includes An. albitarsis and an undescribed species, An. albitarsis "B". These four species were recognized and can be identified using random amplified polymorphic DNA (RAPD) markers, but various other methodologies also point to multiple species under the name An. albitarsis. We describe here a technique for identification of these species employing polymerase chain reaction (PCR) primers based on ribosomal DNA internal transcribed spacer 2 (rDNA ITS2) sequence. Since this method is based on known sequence it is simpler than the sometimes problematical RAPD-PCR. Primers were tested on samples previously identified using RAPD markers with complete correlation.

Animals↗

Single-stranded conformational polymorphism analysis using automated capillary array electrophoresis apparatuses.

We describe a new environment of a single-stranded conformational polymorphism (SSCP) analysis using automated capillary array sequencers (e.g., ABI Prism 3100 and 3700). In this environment, electrophoretic conditions, settings for instrument management, and software for data analysis are adjusted for SSCP analysis. Highly reproducible results are obtained with this new system, and fragments with mutations and/or polymorphisms in different capillaries or different runs can be reliably detected. The relative peak heights between alleles are quantitative and reproducible between runs, and so allele frequencies of single nucleotide polymorphisms can be accurately estimated by a pooled DNA strategy. The method allows unattended, low-cost, and quantitative SSCP analysis using instruments that are widely accessible.

Alleles↗

Identification of Zucchini yellow mosaic potyvirus by RT-PCR and analysis of sequence variability.

A reverse transcription-polymerase chain reaction (RT-PCR) method was used to identify Zucchini yellow mosaic virus (ZYMV) in leaves of infected cucurbits. Oligonucleotide primers which annealed to regions in the nuclear inclusion body (NIb) and the coat protein (CP) genes, generated a 300-bp product from ZYMV and also from the closely related watermelon mosaic virus type 2 (WMV-2). However, no product was obtained from papaya ringspot potyvirus which also infects cucurbits. ZYMV and WMV-2 were differentiated using a third primer which was complementary to a sequence in the 3'-untranslated region; a 1186-bp amplified product was obtained for ZYMV only. Nucleotide sequence analysis of the 300-bp fragments of Australian ZYMV and WMV-2 strains revealed 93.7-100% sequence identity between ZYMV strains. Multiple sequence alignments indicated that the nucleotide sequence which codes for the N-terminus of the CP was 74-100% identical for different isolates of ZYMV. The Australian isolate of WMV-2 was 43-46% identical to all isolates of ZYMV and was 84.6% identical to a Florida isolate of WMV-2.

Amino Acid Sequence↗

A subfamily of acidic alpha-K(+) toxins.

Three homologous acidic peptides have been isolated from the venom of three different Parabuthus scorpion species, P. transvaalicus, P. villosus, and P. granulatus. Analysis of the primary sequences reveals that they structurally belong to subfamily 11 of short chain alpha-K(+)-blocking peptides (Tytgat, J., Chandy, K. G., Garcia, M. L., Gutman, G. A., Martin-Eauclaire, M. F., van der Walt, J. J., and Possani, L. D. (1999) Trends Pharmacol. Sci. 20, 444-447). These toxins are 36-37 amino acids in length and have six aligned cysteine residues, but they differ substantially from the other alpha-K(+) toxins because of the absence of the critical Lys(27) and their total overall negative charge. Parabutoxin 1 (PBTx1), which has been expressed by recombinant methods, has been submitted to functional characterization. Despite the lack of the Lys(27), this toxin blocks several Kv1-type channels heterologously expressed in Xenopus oocytes but with low affinities (micromolar range). Because a relationship between the biological activity and the acidic residue substitutions may exist, we set out to elucidate the relative impact of the acidic character of the toxin and the lack of the critical Lys(27) on the weak activity of PBTx1 toward Kv1 channels. To achieve this, a specific mutant named rPBTx1 T24F/V26K was made recombinantly and fully characterized on Kv1-type channels heterologously expressed in Xenopus oocytes. Analysis of rPBTx1 T24F/V26K displaying an affinity toward Kv1.2 and Kv1.3 channels in the nanomolar range shows the importance of the functional dyad above the acidic character of this toxin.

Amino Acid Sequence↗

Polymorphism of flagellin A gene in Helicobacter pylori.

AIM: To study the polymorphism of flagellin A genotype and its significance in Helicobacter pylori (H. pylori). METHODS: As the template, genome DNA was purified from six clinical isolates of H. pylori from outpatients, and the corresponding flagellin A fragments were amplified by polymerase chain reaction. All these products were sequenced. These sequences were compared with each other, and analyzed by software of FASTA program. RESULTS: Specific PCR products were amplified from all of these H. pylori isolates and no length divergence was found among them. Compared with each other, the highest ungapped identity is 99.10%, while the lowest is 94.65%. Using FASTA program, the alignments between query and library sequences derived from different H. pylori strains were higher than 90%. CONCLUSION: The nucleotide sequence of flagellin A in H. pylori is highly conservative with incident divergence. This information may be useful for gene diagnosis and further study on flagellar antigen phenotype.

Base Sequence↗

[Study of functional L1 retrotransposon in human type 2 diabetes susceptibility loci].

OBJECTIVE: To investigate the susceptibility gene of type 2 diabetes mellitus (T2DM) through a novel strategy. METHODS: Firstly, the common feature of the putative susceptibility genes in the reported susceptibility loci was searched by using NCBI BLAST, and a functional L1 retrotransposon in the loci was found. Secondly, the mRNA expression level of the functional L1 retrotransposon in 25 Han T2DM patients and 22 normal controls was investigated by reverse transcription-polymerase chain reaction, and statistical analysis was implemented in statistical package SPSS10.0. Thirdly, L1 retrotransponson genome mutation screening was performed via sequencing. RESULTS: Screening the human genome for the retrotransposon genome via alignment with the L1 genome using NCBI BLAST showed the functional L1 retrotransposons distribute on most chromosomes except for chromosomes 19, 21 and Y on which rare type 2 diabetes susceptibility loci were reported to reside, and their distribution sites are consistent with the locations of the reported candidate type 2 diabetes susceptibility loci. The mRNA expression level of the functional L1 retrotransposon in the T2DM patients was significantly lower than that in normal subjects (P<0.001). Nonsense mutations including deletion and/or point mutations were observed in all of the 6 T2DM patients tested, but no mutation was observed in all of the 4 normal controls tested. CONCLUSION: The functional L1 retrotransposon may be a candidate susceptibility gene of type 2 diabetes or a key regulator of the susceptibility genes, and it may be an ideal candidate biomarker for screening type 2 diabetes.

Adult↗

Comparison of various algorithms for recognizing short coding sequences of human genes.

MOTIVATION: Since the early 1980s of the twentieth century, there has been great progress in the development of computational gene-finding algorithms. Some problems, however, have not yet been solved currently. Recognizing short genes in prokaryotes and short exons in eukaryotes is one of such problems. The paper is devoted to assessing various algorithms, including those currently available and the new ones proposed here, in order to find the best algorithm to solve the issue. RESULTS: The databases consisting of phase-specific coding and non-coding sequences of human genes with length of 192, 162, 129, 108, 87, 63 and 42 bp, respectively, have been established. Based on the databases and a standard benchmark, 19 algorithms were evaluated, which include the methods of Markov models with orders of 1 through 5, codon usage, hexamer usage, codon preference, amino acid usage, codon prototype, Fourier transform and 8 Z curve methods with various numbers of parameters. Consequently, the Z curve methods with 69 and 189 parameters are the best ones among them, based on the databases constructed here. In addition to the highest recognition accuracy confirmed by 10-fold cross-validation tests, the Z curve methods are much simpler computationally than the second best one, the fifth-order Markov chain model, in which 12 288 parameters are used. We hope that the Z curve methods presented in this paper would be beneficial to the further development of gene-finding algorithms. AVAILABILITY: The programs of various Z curve methods are available on request.

Algorithms↗