PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

PANAL: an integrated resource for Protein sequence ANALysis.

SUMMARY: We present PANAL, an integrated resource for protein sequence analysis. The tool allows the user to simultaneously search a protein sequence for motifs from several databases, and to view the result as an intuitive graphical summary.

Computational Biology↗

Sequence complexity for biological sequence analysis.

A new statistical model for DNA considers a sequence to be a mixture of regions with little structure and regions that are approximate repeats of other subsequences, i.e. instances of repeats do not need to match each other exactly. Both forward- and reverse-complementary repeats are allowed. The model has a small number of parameters which are fitted to the data. In general there are many explanations for a given sequence and how to compute the total probability of the data given the model is shown. Computer algorithms are described for these tasks. The model can be used to compute the information content of a sequence, either in total or base by base. This amounts to looking at sequences from a data-compression point of view and it is argued that this is a good way to tackle intelligent sequence analysis in general.

Algorithms↗

The use of amino acid sequence analysis in assessing evolution.

The thirteen year history of assessing evolution by amino acid sequence analysis has made apparent the limitations imposed upon this system by the finite nature of the characters. This finiteness exists on several levels and ultimately expresses itself as parallelism, back mutation and the retention of primitive characters in the sequences of proteins from present day species and the putative ancestral protein chains. Sequence analysis shares these problems with other molecular approaches, but because it is concerned both with the nucleotide substitutions in the genome and with the functional roles of proteins, it has unique advantages. For example, the large fluctuation in the rate of fixation of mutations in a protein's evolution can be detected and used to point out the unreliability of any molecular clock for estimating divergence dates. Moreover, when consideration is given to studies which assign functional significance to specific amino acid sites in a protein, changes in function during the descent of a protein can be appreciated and their significance correlated with organismal evolution.

Amino Acid Sequence↗

Solid-phase sequence analysis of polypeptides eluted from polyacrylamide gels. An aid to interpretation of DNA sequences exemplified by the Escherichia coli unc operon and bacteriophage lambda.

An approach to sequencing proteins by the solid-phase method combined with isolation of proteins and polypeptides by gel electrophoresis is described. Mixtures of proteins or polypeptides resulting from digests are fractionated in the presence of dodecylsulphate in polyacrylamide gels. They are detected with Coomassie blue, eluted, selectively reacted with porous glass derivatives and sequenced in their amino-terminal regions with the aid of a new microsequencer. Alternatively they can be analysed or digested with enzymes and fingerprinted. It is a relatively rapid method of purifying proteins for sequence analysis which we have used to provide partial protein sequence data to complement DNA sequences. Nine genes, four from the unc operon of Escherichia coli encoding the alpha, beta, gamma and epsilon subunits of ATP synthase and five for capsid proteins of bacteriophage lambda, have been identified by this method.

Adenosine Triphosphatases↗

Identification, transcription and sequence analysis of the Spodoptera littoralis nucleopolyhedrovirus (SpliNPV) DNA polymerase gene.

Sequence analysis of a 6.4 kb DNA region from the Spodoptera littoralis multinucleocapsid nucleopolyhedrovirus (SpliNPV) revealed a large open reading frame (ORF) encoding a predicted polypeptide of 998 amino acid (aa) residues with a molecular mass of 114.93 kDa, located between 47.2-52.3 m.u. on the SpliNPV genome. Comparative sequence analyses demonstrated that the ORF encodes a DNA polymerase gene (dnapol) that contains conserved exonuclease domains and DNA polymerase motifs found in many prokaryotic, eukaryotic, and viral replicative DNA polymerases. A second ORF, ORF138, located between the lef-3 and dnapol, encodes a 138 aa polypetide that is homologous to ORF66 of the Autographa californica MNPV (AcMNPV). SpliNPV DNA polymerase shares an overall aa sequence identity of 39% with that of AcMNPV. A 3.0 kb SpliNPV dnapol-specific transcript was detected initially at 2 hpi and became abundant 48 hpi by Northern blot analysis. The transcription initiation site was mapped to an NPV early promoter element, ACGT. 3' RACE demonstrated that the SpliNPV dnapol transcript terminated at the polyadenylation signal AATAAA. Sequence analysis suggested that the SpliNPV dnapol and the dnapol of the NPV of S. litura (SpltNPV) are closely related.

Amino Acid Sequence↗

ColiGene: object-centered representation for the study of E coli gene expressivity by sequence analysis.

ColiGene is an object-centered knowledge base for the study of gene expressivity in Escherichia coli by DNA sequence analysis. This system was developed with the knowledge base management system SHIRKA. Objects represented in ColiGene are biological structures such as genes or regulatory signals. They are organized in a hierarchical structure of classes, subclasses and instances. Navigation through the knowledge base and the building of queries are made using a graphical interface. The base is coupled with the data base ACNUC which structures a specialized collection of sequences: EcoSeq. Several tools are also associated to ColiGene, either for sequence analysis or for a more general purpose. Some biological results have been obtained using ColiGene which are summarized here.

Artificial Intelligence↗

Comparative DNA sequence analysis of mapped wheat ESTs reveals the complexity of genome relationships between rice and wheat.

The use of DNA sequence-based comparative genomics for evolutionary studies and for transferring information from model species to related large-genome species has revolutionized molecular genetics and breeding strategies for improving those crops. Comparative sequence analysis methods can be used to cross-reference genes between species maps, enhance the resolution of comparative maps, study patterns of gene evolution, identify conserved regions of the genomes, and facilitate interspecies gene cloning. In this study, 5,780 Triticeae ESTs that have been physically mapped using wheat ( Triticum aestivum L.) deletion lines and segregating populations were compared using NCBI BLASTN to the first draft of the public rice ( Oryza sativa L.) genome sequence data from 3,280 ordered BAC/PAC clones. A rice genome view of the homoeologous wheat genome locations based on sequence analysis shows general similarity to the previously published comparative maps based on Southern analysis of RFLP. For most rice chromosomes there is a preponderance of wheat genes from one or two wheat chromosomes. The physical locations of non-conserved regions were not consistent across rice chromosomes. Some wheat ESTs with multiple wheat genome locations are associated with the non-conserved regions of similarity between rice and wheat. The inverse view, showing the relationship between the wheat deletion map and rice genomic sequence, revealed the breakdown of gene content and order at the resolution conferred by the physical chromosome deletions in the wheat genome. An average of 35% of the putative single copy genes that were mapped to the most conserved bins matched rice chromosomes other than the one that was most similar. This suggests that there has been an abundance of rearrangements, insertions, deletions, and duplications eroding the wheat-rice genome relationship that may complicate the use of rice as a model for cross-species transfer of information in non-conserved regions.

Chromosome Mapping↗

Alginate lyase from Klebsiella pneumoniae, subsp. aerogenes: gene cloning, sequence analysis and high-level production in Escherichia coli.

The alyA gene, encoding a secreted guluronate-specific alginate lyase (Aly) from Klebsiella pneumoniae subsp. aerogenes type 25, has been cloned. DNA sequence analysis reveals two possible translation start sites for the precursor form of Aly and a long open reading frame (ORF) predicted to encode a 287-amino-acid (aa) mature form of Aly, in agreement with N-terminal aa sequence analysis of the protein. Aly has a calculated molecular mass of 31.4 kDa, in good agreement with SDS-PAGE analysis, and a calculated pI of 9.39. Comparison of the deduced aa sequence with a mannuronate-specific lyase from a marine bacterium reveals 19.3% identity and 28.8% similarity with a 9-aa conserved region close to the C terminus, probably of functional or structural significance. There is no obvious sequence similarity with pectate lyases which also catalyse a beta-elimination reaction. Heterologous expression of K. pneumoniae alyA in Escherichia coli yields 10 mg of Aly per litre of culture supernatant, apparently due to non-specific release from the periplasm.

Amino Acid Sequence↗

Development of real-time image sequence analysis for evaluating posture change and respiratory rate of a subject in bed.

An image sequence analysis technique was developed to evaluate posture change and respiratory rate of a subject in bed without any physical contact. Although the image sequence analysis requires many calculations, the system can perform them in real time. The system consisted of a CCD video camera and a PC equipped with a high-speed image processor. To evaluate the system, we tested it on five subjects at a nursing home. The system evaluated 99.4% of the movements of subjects during the total monitoring time (about 61 hours). The waveform was flat when the subject was out of view of the video camera. The system has the possibility of evaluating not only posture changes and respiratory rate. but also sleeping patterns.

Adult↗

Molecular cloning and sequence analysis of the penton base genes of type II avian adenoviruses.

We describe here the identification of the penton base gene of hemorrhagic enteritis virus (HEV), a type II avian adenovirus, in a 2477-base pair (bp)-EcoRI fragment of the viral DNA by sequence analysis. Identification is based on an extensive amino acid homology between the HEV-open reading frame and the penton base of a fowl adenovirus (FAV-10) and various human adenoviruses. The 1344 bp-penton base gene of HEV encodes a 448-amino acid polypeptide of molecular weight of 50,843 Da. The nucleotide sequences of penton base genes of HEV and marble spleen disease virus (MSDV) are identical. The HEV penton base lacks the RGD motif, present in most human adenoviruses (Ad2, Ad3, Ad4, and Ad 12) suggesting that HEV may not use alpha v integrins to gain entry into host cells. Further sequence analysis revealed the presence of a Leu-Asp-Val (LDV) motif in the HEV penton base amino acid sequence similar to most of the human adenoviruses. LDV motif on the fibronectin has been shown to interact with the alpha 4 beta 1 integrins on cells, which includes lymphocytes and monocytes. The presence of LDV motif in the penton base of HEV implicates the involvement of alpha 4 beta 1 integrins in the viral internalization into host cells.

Amino Acid Sequence↗

DNA sequence analysis by hybridization with oligonucleotide microchips: MALDI mass spectrometry identification of 5mers contiguously stacked to microchip oligonucleotides.

Matrix-assisted laser desorption ionization mass spectrometry (MALDI MS) has been applied to increase the informational output from DNA sequence analysis. It has been used to analyze DNA by hybridization with microarrays of gel-immobilized oligonucleotides extended with stacked 5mers. In model experiments, a 28 nt long DNA fragment was hybridized with 10 immobilized, overlapping 8mers. Then, in a second round of hybridization DNA-8mer duplexes were hybridized with a mixture of 10 5mers. The stability of the 5mer complex with DNA was increased to raise the melting temperature of the duplex by 10-15 degrees C as a result of stacking interaction with 8mers. Contiguous 13 bp duplexes containing an internal break were formed. MALDI MS identified one or, in some cases, two 5mers contiguously stacked to each DNA-8mer duplex formed on the microchip. Incorporating a mass label into 5mers optimized MALDI MS monitoring. This procedure enabled us to reconstitute the sequence of a model DNA fragment and identify polymorphic nucleotides. The application of MALDI MS identification of contiguously stacked 5mers to increase the length of DNA for sequence analysis is discussed.

Animals↗

Detection, cloning, and sequence analysis of an indigenous plasmid from cellulolytic clostridial strain MCF1.

Nucleotide sequence analysis of a 2451-bp plasmid (pMCF1) from a cellulolytic Clostridium revealed that the protein specified by the largest open reading frame (ORF1) was homologous to RepB of Clostridium butyricum plasmid pCB101. The data suggest that pMCF1 belongs to the pC194 family of rolling-circle replicating plasmids and the ORF1 protein functions as its replication protein.

Amino Acid Motifs↗

Recurrence time statistics: versatile tools for genomic DNA sequence analysis.

With the completion of the human and a few model organisms' genomes, and with the genomes of many other organisms waiting to be sequenced, it has become increasingly important to develop faster computational tools which are capable of easily identifying the structures and extracting features from DNA sequences. One of the more important structures in a DNA sequence is repeat-related. Often they have to be masked before protein coding regions along a DNA sequence are to be identified or redundant expressed sequence tags (ESTs) are to be sequenced. Here we report a novel recurrence time-based method for sequence analysis. The method can conveniently study all kinds of periodicity and exhaustively find all repeat-related features from a genomic DNA sequence. An efficient codon index is also derived from the recurrence time statistics, which has the salient features of being largely species-independent and working well on very short sequences. Efficient codon indices are key elements of successful gene finding algorithms, and are particularly useful for determining whether a suspected EST belongs to a coding or non-coding region. We illustrate the power of the method by studying the genomes of E. coli, the yeast S. cervisivae, the nematode worm C. elegans, and the human, Homo sapiens. Our method requires approximately 6 . N byte memory and a computational time of N log N to extract all the repeat-related and periodic or quasi-periodic features from a sequence of length N without any prior knowledge on the consensus sequence of those features, hence enables us to carry out sequence analysis on the whole genomic scale by a PC.

Algorithms↗

Bifidobacteria identification based on 16S rRNA and pyruvate kinase partial gene sequence analysis.

The lack of a simple and rapid identification system for Bifidobacterium species makes them difficult to use in industrial applications. To obtain valuable discriminating factor, we studied different strains, and human isolates by two molecular taxonomy methods. First method was based on chrono-differentiation. A metabolic gene (pyruvate kinase) was chosen to be used as a systematic discriminating factor. A comparison of about 40 pyruvate kinase protein sequences allowed us to synthesize two oligonucleotides that were able to amplify a fragment of this corresponding gene in our strains. Based on these partial pyruvate kinase gene sequences, several clusters could be identified. The second method used in this study was based on 16S rRNA sequences analysis. We compared sequences present in GenBank database, and this allowed to separate bifidobacteria species into different clusters. They were different from those obtained with partial pyruvate kinase gene sequences analysis. So, by combining both methods, we were able to identify our isolates, when only 10% of them could be strictly identified using the 16S rRNA method. Moreover, pyruvate kinase analysis allowed to differentiate very ambivalent groups such as B. animalis/B. lactis or B. infantis/B. longum, but created different clusters for B. infantis species group, questioning on the homogeneity of this species.

Journal Article↗

Comparative DNA sequence analysis of wheat and rice genomes.

The use of DNA sequence-based comparative genomics for evolutionary studies and for transferring information from model species to crop species has revolutionized molecular genetics and crop improvement strategies. This study compared 4485 expressed sequence tags (ESTs) that were physically mapped in wheat chromosome bins, to the public rice genome sequence data from 2251 ordered BAC/PAC clones using BLAST. A rice genome view of homologous wheat genome locations based on comparative sequence analysis revealed numerous chromosomal rearrangements that will significantly complicate the use of rice as a model for cross-species transfer of information in nonconserved regions.

Chromosome Mapping↗

Nucleotide sequence analysis of variola virus HindIII M, L, I genome fragments.

DNA of the variola major virus strain India-1967 in the region of HindIII M, L, I fragments has been sequenced. Analysis of this sequence of 18029 bp revealed 19 potential open reading frames (ORFs). Four proposed proteins (L2R, H9R, L5L, L6R) contain metal-binding domains. Comparison of the variola virus (VAR) and vaccinia virus strain Copenhagen (COP) sequences show that the main differences are between proteins L1R and I5R. L1R contains 6 additional amino acid residues on the C-terminus. The protein I5R of VAR contains three Ca2+ binding domains but this COP has deletions in 2 of the 3 established domains. Possible functions of the predicted viral polypeptides are discussed.

Amino Acid Sequence↗

Application of ITS sequence analysis, RAPD and AFLP fingerprinting in characterising the yeast genus Fellomyces.

Three molecular techniques, ITS sequence analysis, random amplified polymorphic DNA (RAPD) and amplified fragment length polymorphism (AFLP) were used to study phylogenetic and genotypic relationships among strains of the genus Fellomyces. In the analyses were included strains isolated predominantly from epiphytic lichens collected in Indonesia, China and Mexico. The polyphasic approach indicated that the Fellomyces isolates are genotypically heterogeneous and that lichens represent a specific environment for selection of large number of the sterigmatoconidia producing species. The phylogenetic and genotypic analysis confirmed the existence of 11 currently accepted Fellomyces species and indicated that several species may be the new representatives of the genus. The RAPD and AFLP analyses demonstrated a higher potential in distinguishing the Fellomyces strains than the ITS regions. Since the sequence analysis showed low or no divergence among several strains, both RAPD and AFLP fingerprinting indicated that the strains may be discriminated at the species level.

Basidiomycota↗

Independent segregation of the VP4 and the VP7 genes in bovine rotaviruses as confirmed by VP4 sequence analysis of G8 and G10 bovine rotavirus strains.

We determined the complete nucleotide sequences of the VP4 genes of five bovine rotavirus strains (A5, 61A, A44, B223 and KK3). The deduced VP4 amino acid sequence of strain A5 with G8 serotype specificity is 91.1% identical to that of strain NCDV of the serotype G6, and strain 61A with G10 serotype specificity has a VP4 amino acid sequence (95.2% identity) similar to that of the UK strain of the G6 serotype. In contrast, the VP4 amino acid sequences of strains A44, B223 and KK3 of the G10 serotype, isolated in Thailand, the U.S.A. and Japan respectively, have very similar sequences to each other, but less similarity (50 to 60%) to other group A rotavirus strains reported so far. Their VP4 genes are 2352 nucleotides in length and encode 772 amino acids, four amino acids fewer than the VP4 proteins of other animal rotaviruses. Thus, the presence of three different VP4 types (P types) and the independent segregation of G serotype and P type were postulated in bovine rotaviruses by VP4 sequence analysis.

Amino Acid Sequence↗