PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Statistical significance in biological sequence analysis.

One of the major goals of computational sequence analysis is to find sequence similarities, which could serve as evidence of structural and functional conservation, as well as of evolutionary relations among the sequences. Since the degree of similarity is usually assessed by the sequence alignment score, it is necessary to know if a score is high enough to indicate a biologically interesting alignment. A powerful approach to defining score cutoffs is based on the evaluation of the statistical significance of alignments. The statistical significance of an alignment score is frequently assessed by its P-value, which is the probability that this score or a higher one can occur simply by chance, given the probabilistic models for the sequences. In this review we discuss the general role of P-value estimation in sequence analysis, and give a description of theoretical methods and computational approaches to the estimation of statistical signifiance for important classes of sequence analysis problems. In particular, we concentrate on the P-value estimation techniques for single sequence studies (both score-based and score-free), global and local pairwise sequence alignments, multiple alignments, sequence-to-profile alignments and alignments built with hidden Markov models. We anticipate that the review will be useful both to researchers professionally working in bioinformatics as well as to biomedical scientists interested in using contemporary methods of DNA and protein sequence analysis.

Computational Biology↗

Characterisation of polyoma late mRNA leader sequences by molecular cloning and DNA sequence analysis.

The leader sequences of two of the three polyoma virus late mRNAs were characterised by molecular cloning and DNA sequence analysis. A short single-stranded DNA fragment complementary to the 5' end of the body of mVP1 was used to prime cDNA synthesis, and double-stranded cDNA was inserted into a derivative of pAT 153. Analysis of fourteen mVP1 cDNAs and two mVP3 cDNAs allowed the precise determination of the leader-body joints, and demonstrated that the majority of polyoma late leader sequences consist of exact tandem repeats of a 57 nucleotide sequence present only once in the genomic DNA at 66-67 m.u. The sequences in the genomic DNA borderline this unit are typical of those found at RNA splice points. Leader sequences contained on average three to four repeat units. Nuclease S1 mapping of total late mRNA demonstrated that most mRNA 5' ends map heterogeneously in the 50 nucleotides 5' to the repeated sequence unit. The structures of the leader sequences strongly suggest that they are generated by appropriate splicing events from a tandemly repeated transcript of the entire circular viral genome.

Base Sequence↗

GEOMETRY: a software package for nucleotide sequence analysis using statistical geometry in sequence space.

GEOMETRY is a software package for the analysis of nucleotide sequences using the method of statistical geometry in sequence space. The package consists of programs performing estimation of the average geometry of sequence quartets, analysis of positional variability and computer simulation of parallel and tree-like sequence divergence with user-defined parameters. It provides an independent tool for evaluation of the reliability of conventional phylogenetic trees and calibration of the time of sequence divergence. GEOMETRY may be of interest for all scientists engaged in the study of molecular phylogeny. The package is available by anonymous FTP from ftp.bionet.nsk.su, directory /incoming/molevol/geom.exe, and will be available from EMBL file server (URL: http://@www.ebi.ac.uk)

Algorithms↗

Fluorescent labeling of cysteinyl residues to facilitate electrophoretic isolation of proteins suitable for amino-terminal sequence analysis.

A protein labeling procedure which enables detection of subpicomole quantities of proteins on sodium dodecyl sulfate (SDS)-polyacrylamide gels is described. Proteins are rendered fluorescent by reduction of disulfide bonds with dithiothreitol followed by alkylation with 5-N-[(iodoacetamidoethyl)amino]naphthalene-1-sulfonic acid (5-I-AEDANS) or 5-iodoacetamido-fluorescein. Labeling is performed prior to electrophoresis, thus eliminating the need for staining with dyes and destaining after electrophoresis. As little as 375 fmol (25 ng) of prelabeled bovine serum albumin can be readily visualized after electrophoresis. Bands are still visible after electrophoretic transfer to nitrocellulose. Simultaneous labeling of proteins in complex mixtures is possible using this technique. This includes cysteine containing proteins of disrupted Newcastle disease virus. The magnitudes of the molecular weight increases which occur upon labeling reflect the cysteine contents of proteins. The mode of chemical modification for the prelabeling procedure was chosen because of its compatibility with analytical techniques, such as amino acid analysis, peptide mapping, or sequence analysis, which may be applied to the protein after electroelution from SDS-acrylamide gels. It replaces the need for reduction and carboxymethylation prior to these analytical procedures. Protein-sequence analysis of prelabeled bovine serum albumin, including samples electroeluted from SDS-acrylamide gels, has justified the choice of this method to facilitate isolation of proteins for sequence analysis. Equivalent sequence data were obtained with reduced bovine serum albumin S-alkylated with iodoacetic acid or 5-I-AEDANS.

Amino Acid Sequence↗

DNA sequence analysis: a general, simple and rapid method for sequencing large oligodeoxyribonucleotide fragments by mapping.

Several electrophoretic and chromatographic systems have been investigated and compared for sequence analysis of oligodeoxyribonucleotides. Three systems were found to be useful for the separation of a series of sequential degradation products resulting from a labeled oligonucleotide: (I) 2-D electrophoresisdagger; (II) 2-D PEI-cellulose; and (III) 2-D homochromatography. System (III) proved generally most informative regardless of base composition and sequence. Furthermore, only in this system will the omission of an oligonucleotide in a series of oligonucleotides be self-evident from the two-dimensional map. The sequence of up to fifteen nucleotides can be determined solely by the characteristic mobility shifts of its sequential degradation products distributed on the two-dimensional map. With this method, ten nucleotides from the double-stranded region adjacent to the left-hand 3'-terminus and seven from the right-hand 3'-terminus of bacteriophage lambda DNA have been sequenced. Similarly, nine nucleotides from the double-stranded region adjacent to the left-hand 3'-terminus and five nucleotides from the right-hand terminus of bacteriophage phi80 DNA have also been sequenced. The advantages and disadvantages of each separation system with respect to sequence analysis are discussed.

Base Sequence↗

The tetracycline resistance determinants of RP1 and Tn1721: nucleotide sequence analysis.

Nucleotide sequences of the homologous tetracycline resistance (tet) determinants of plasmid RP1 and transposon Tn1721 have been determined. Two open reading frames of divergent polarity have been assigned to a regulatory gene (tetR) and a gene encoding a resistance protein (tetA). The intercistronic region contains appropriate regulatory and transcription signals. The tetR gene can code for a protein of 216 amino acids (deduced mol.wt. 23,288) and the tetA gene for a protein of 399 amino acids (deduced mol. wt. 42,205). Based on the deduced amino acid sequence, the tetA proteins of RP1/Tn1721 are 78% homologous with that of pBR322 and 45% homologous with that of Tn10. We conclude that a single tetA gene mediates resistance in each of these tet determinants.

Bacterial Proteins↗

Origin of tetracycline efflux proteins: conclusions from nucleotide sequence analysis.

The sequences of six tetracycline efflux proteins and three transport proteins which have some resemblance to them were compared. The tetracycline efflux proteins fall into three families: (i) those encoded by pBR322, RP1, and Tn10 (Escherichia coli); (ii) pT181 (Staphylococcus aureus) and pTHT15 (Bacillus subtilis); and (iii) tet347 (Streptomyces rimosus). There is global sequence homology within each of the first two families, but there is none between the families. The pT181/pTHT15 family shares close homology with the N-terminal half of the methylenomycin A efflux protein (Streptomyces coelicor), while tet347 resembles the C-terminal half. Portions of the N-terminal half of the Tn10-encoded protein show significant resemblance to portions in the N-terminal half of the pT181/pTHT15 family, but this sometimes occurs among transport proteins which do not have a common substrate. Tetracycline efflux proteins, therefore, appear to have arisen on at least two, or possibly three, separate occasions, probably from other transport proteins.

Amino Acid Sequence↗

Phylogenetic analysis of rumen bacteria by comparative sequence analysis of cloned 16S rRNA genes.

Comparative DNA sequence analysis of 16S rRNA genes (rDNA) was undertaken to further our understanding of the make-up of bacterial communities in the rumen fluid of dairy cattle. Total DNA was extracted from the rumen fluid of 10 cattle fed haylage/corn silage/concentrate rations at two different times. Rumen samples were collected on two separate occasions from five cows each. In experiment 1, 31 cloned rDNA sequences were analysed. In experiment 2, DNA extractions were amplified using either 12 or 30 cycles of PCR in order to examine biases introduced during the reactions. A set of 53 sequences were analysed in experiment 2 from DNA amplified using 12 cycles and 49 sequences from PCR using 30 cycles. Sequences from the 5' end of 16S rRNA gene were compared with existing sequences in the Ribosomal Database Project. Clones from experiment 1 produced a data set in which 55% of the sequences were similar to low G+C Gram-positive bacteria related to the genus Clostridia, the majority of which were closely related to bacteria in Cluster XIV. Approximately 30% of the cloned sequences were related to bacteria in the Prevotella-Bacteroides group. Clones from experiment 2 produced a data set in which the majority of sequences were related to the Prevotella-Bacteroides group, regardless of the number of cycles of PCR. The remaining sequences clustered with members of the genus Clostridia. The majority of rDNA sequences analysed in this study represent novel rumen bacteria which have not yet been isolated.

Journal Article↗

Omiga: a PC-based sequence analysis tool.

Computer-based sequence analysis, notation, and manipulation are a necessity for all molecular biologists working with any but the most simple DNA sequences. As sequence data become increasingly available, tools that can be used to manipulate and annotate individual sequences and sequence elements will become an even more vital implement in the molecular biologist's arsenal. The Omiga DNA and Protein Sequence Analysis Software tool, version 2.0 provides an effective and comprehensive tool for the analysis of both nucleic acid and protein sequences that runs on a standard PC available in every molecular biology laboratory. Omiga allows the import of sequences in several common formats. Upon importing sequences and assigning them to various projects, Omiga allows the user to produce, analyze, and edit sequence alignments. Sequences may also be queried for the presence of restriction sites, sequence motifs, and other sequence features, all of which can be added into the notations accompanying each sequence. This newest version of Omiga also allows for sequencing and polymerase chain reaction (PCR) primer prediction, a functionality missing in earlier versions. Finally, Omiga allows rapid searches for putative coding regions, and Basic Local Alignment Search Tool (BLAST) queries against public databases at the National Center for Biotechnology Information (NCBI).

Humans↗

Identification of Bacillus anthracis by rpoB sequence analysis and multiplex PCR.

Comparative sequence analysis was performed upon Bacillus anthracis and its closest relatives, B. cereus and B. thuringiensis. Portions of rpoB DNA from 10 strains of B. anthracis, 16 of B. cereus, 10 of B. thuringiensis, 1 of B. mycoides, and 1 of B. megaterium were amplified and sequenced. The determined rpoB sequences (318 bp) of the 10 B. anthracis strains, including five Korean isolates, were identical to those of Ames, Florida, Kruger B, and Western NA strains. Strains of the "B. cereus group" were separated into two subgroups, in which the B. anthracis strains formed a separate clade in the phylogenetic tree. However, B. cereus and B. thuringiensis could not be differentiated. Sequence analysis confirmed the five Korean isolates as B. anthracis. Based on the rpoB sequences determined in the present study, multiplex PCR generating either B. anthracis-specific amplicons (359 and 208 bp) or cap DNA (291 bp) in a virulence plasmid could be used for the rapid differential detection and identification of virulent B. anthracis.

Anthrax↗

Chromosomal localization and sequence analysis of a human episomal sequence with in vitro differentiating activity.

The genomic fragment carrying the human activator of liver function, previously described as an episome capable of inducing differentiation upon transfection into a dedifferentiated rat hepatoma cell line, was mapped on human chromosome 12q24.2-12q24.3. This chromosomal location was indistinguishable by in situ hybridization from that of the gene coding for the hepatic transcription factor HNF1. The sequence of the integrated form of the episome as well as its flanking sequences show that it is rich in retroposons. It contains a human ribosomal protein L21 processed pseudogene, one truncated L1Hs sequence, and 10 Alu repeats, which belong to different subfamilies.

Adult↗

Sequence analysis for assessing potential allergenicity.

Sequence analysis plays an important role in assessing the potential allergenicity of proteins used in transgenic foods, particularly for proteins that have not previously been part of the food supply. Sequence comparisons are used to indicate potential unexpected cross reactivity to existing allergens and to assess the potential for developing new sensitivities. Although the concept of using sequence analysis is straightforward, implementing a bioinformatic analysis that is accurate and complete can be complex. Several factors need to be considered, including the design and content of the sequence database, the analysis strategy, and the criteria for evaluating the results.

Algorithms↗

Sequence analysis of adenovirus DNA: complete nucleotide sequence of the spliced 5' noncoding region of adenovirus 2 hexon messenger RNA.

The complete nucleotide sequence of the 5' noncoding region of the adenovirus 2 hexon messenger RNA has been established by sequence analysis of reverse transcripts. Such transcripts were generated by extension of specific single-stranded DNA primers with reverse transcriptase after hybridization to purified hexon mRNA. The total length of the 5' noncoding region was determined to be 240 nucleotides, of which the spliced tripartite leader sequence contributes 202 nucleotides including the terminal m7G. The sizes of the different segments of the tripartite leader were estimated by comparing the established mRNA sequence with the genomic sequences for the first and third leader segments, and were found to be 42 nucleotides for the first segment, 71 nucleotides for the second and 89 nucleotides for the third. The estimates are ambiguous, however, due to the presence of tandemly repeated sequences at both ends of the intervening sequence between the third leader segment and the body of the hexon mRNA. The sequence of the leader allows the formation of hydrogen-bonded interactions with the 3' end of 18S ribosomal RNA near the capped 5' end and also close to the initiator AUG.

Adenoviruses, Human↗

[Bacterial 16S rDNA sequence analysis of Siberian tiger faecal flora].

Bacterial 16S rDNA library of Siberian tiger was developed and 15 different clones were obtained using EcoR I and Hind III in restriction fragment length polymorphism analysis. DNA sequencing and similarity analysis showed that 10 clones matched corresponding Clostridium sequences, of which 6 sequences had over 99% similarity with Clostridium novyi type A, and 4 sequences had 97% similarity with Swine manure bacterium RT-18B, which identified as Peptostreptococcus spp. The other five 16S rDNA sequences had 94% - 95% similarity with Clostridium pascui, Clostridium tetani E88, Clostridium sp. 14505 Clostridium perfringens and Carnobacterium sp. R-7279 respectively.

Animals↗

The chemistry of protein sequence analysis.

N-terminal sequence analysis by Edman chemistry continues to play an important role in the structural analysis of proteins and peptides. Improvements in the sensitivity of the method have been achieved mainly at the level of increasing the sensitivity of the on-line analysis of PTH amino acids by RP-HPLC (reverse phase high performance chromatography). Using microbore columns (0.8-1.0 mm), it is possible to run standards at the 0.5-1.0 pmol level and to sequence samples in the 1-5 pmol range. Due to constraints in current chromatographic methods, it is unlikely that further improvements in sensitivity will be achieved by this approach alone. Although alternative Edman reagents, including fluorescent chemistries, have promised to increase the sensitivity of sequencing into the low femtomole range, none of the methods have progressed into routine usage. These reagents and chemistries are critically evaluated in this review, and the problems which have prevented their further development discussed. Instrumental constraints are also considered. It is concluded that the development of more sensitive methods requires further research into both the chemistry and the instrumentation, and that alternative separation and detection methods may also play a role.

Amino Acid Sequence↗

Direct sequence analysis of human herpesvirus 6 (HHV-6) sequences from infants and comparison of HHV-6 sequences from mother/infant pairs.

Direct sequence analysis of polymerase chain reaction-amplified DNA fragments from the large tegument protein (LTP) gene of human herpesvirus 6 (HHV-6) was performed with use of uncultured peripheral blood mononuclear cells (PBMCs) from four mother/infant pairs. In two cases, LTP gene sequences were identical in paired mother/infant specimens, thus suggesting that mother-to-infant transmission of HHV-6 may have occurred. The genetic stability of HHV-6 strains was confirmed by the fact that there was no difference between amplified DNA fragments from sequential PBMC samples from two of two infants analyzed. In contrast, a change in the amplified viral strain was detected in an infant who had reinfection with HHV-6 variant B (HHV-6B). Furthermore, HHV-6B strains concurrently amplified from saliva and PBMCs from an adult were found to be different. The data suggest that HHV-6 may be frequently transmitted from mother-to-infant and that reinfection with HHV-6B may occur.

Adult↗

East Asian mtDNA haplogroup determination in Koreans: haplogroup-level coding region SNP analysis and subhaplogroup-level control region sequence analysis.

The present study analyzed 21 coding region SNP markers and one deletion motif for the determination of East Asian mitochondrial DNA (mtDNA) haplogroups by designing three multiplex systems which apply single base extension methods. Using two multiplex systems, all 593 Korean mtDNAs were allocated into 15 haplogroups: M, D, D4, D5, G, M7, M8, M9, M10, M11, R, R9, B, A, and N9. As the D4 haplotypes occurred most frequently in Koreans, the third multiplex system was used to further define D4 subhaplogroups: D4a, D4b, D4e, D4g, D4h, and D4j. This method allowed the complementation of coding region information with control region mutation motifs and the resultant findings also suggest reliable control region mutation motifs for the assignment of East Asian mtDNA haplogroups. These three multiplex systems produce good results in degraded samples as they contain small PCR products (101-154 bp) for single base extension reactions. SNP scoring was performed in 101 old skeletal remains using these three systems to prove their utility in degraded samples. The sequence analysis of mtDNA control region with high incidence of haplogroup-specific mutations and the selective scoring of highly informative coding region SNPs using the three multiplex systems are useful tools for most applications involving East Asian mtDNA haplogroup determination and haplogroup-directed stringent quality control.

Asian People↗

A method for high-performance sequence analysis using polyvinylidene difluoride membranes with a biphasic reaction column sequencer.

Methods have been developed for high-sensitivity sequence analysis of proteins electroblotted onto polyvinylidene difluoride (PVDF) membranes using a Hewlett-Packard G1005A protein sequencer. This sequencer normally uses a biphasic (hydrophobic/hydrophilic) reaction column which was designed to accommodate loading and cleanup of samples from diverse solutions. However, the standard column, programs, and chemistry were not designed to accommodate PVDF, which has become a common sequencing support. In this study, a systematic evaluation of the suitability of this sequencer for analysis using PVDF bound samples was performed and included evaluation of: different wash and extraction solvents, multiple programming changes, two alternative formulations of coupling reagents, and the effect of direction for solvent and reagent deliveries. High-performance analysis of PVDF bound samples was achieved by: using a modified reaction column with an empty hydrophobic (top) half of the column module, program modifications for the reaction column and converter, substitution of ethyl acetate for the standard S2/3 extraction solvent and using prototype Version 2.0 formulations of the coupling reagents, R1 and R2. High-performance sequence analyses of experimental samples electroblotted from either 1D or 2D gels onto high-retention PVDF membranes were obtained with a 41-min cycle time, including experimental samples with initial coupling yields < 2 pmol. Routine sequencer performance was comparable to, or slightly better than, a conventional gas-phase sequencer which had been previously optimized by us for high-performance sequence analysis of electroblotted samples in the low pmol range.

Amino Acid Sequence↗