PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Nucleotide sequence analysis shows that Rhopalosiphum padi virus is a member of a novel group of insect-infecting RNA viruses.

Rhopalosiphum padi virus (RhPV) is an aphid virus that has been considered a member of the Picornaviridae based on physicochemical properties. The 10,011-nt polyadenylated RNA genome of RhPV was completely sequenced. Analysis of the sequence revealed the presence of two open reading frames (ORFs). The predicted amino acid sequence of ORF1, representing the first 6600 nt of the RhPV genome, showed significant similarity to the nonstructural proteins of several plant and animal RNA viruses. Direct sequence analysis of the RhPV capsid proteins showed that ORF2, which represents the last 2900 nt, encodes the three structural proteins (28, 29, and 30 kDa). The predicted amino acid sequence of ORF2 is very similar to the corresponding regions of Drosophila C virus, Plautia stali intestine virus, and to a partial sequence from the 3' end of the cricket paralysis virus genome. The site of initiation of protein synthesis for ORF2 could not be determined from the amino acid and nucleotide sequences. ORF1 is preceded by 579 nt of noncoding RNA and the two ORFs are separated by more than 500 nt of noncoding RNA. Like picornaviruses, these regions may function to facilitate the cap-independent initiation of translation of the two ORFs. These data suggest that RhPV, Drosophila C virus, Plautia stali intestine virus, and probably cricket paralysis virus are members of a unique group of small RNA viruses that infect primarily insects.

Amino Acid Sequence↗

Determination of enterovirus serotype inferred from sequence analysis of PCR products.

BACKGROUND: Enterovirus infections are common in neonates. Virus isolation is the only diagnostic method to confirm enterovirus serotype infections, however, is not always successful. OBJECTIVES: A new approach for the diagnosis of enterovirus infections was performed, using the reference strain inferred from sequence analysis of PCR products. STUDY DESIGN: Virus isolation, enterovirus RT-PCR and sequence analysis were performed from clinical samples or stored sera from two neonates with fever and rash. Neutralizing test (NT) antibodies against prototype reference virus were measured in paired sera. RESULTS: Virus isolation was negative in both patients but the enterovirus genome was amplified in the acute phase sera obtained from the two patients. From the results of sequence analysis of 109 nucleotides located in the 5'-noncoding of the conserved region of enteroviruses, a high homology to echovirus types 25 and 30 was found. More than a 4-fold increase in NT antibodies against reference viruses was demonstrated in the acute and convalescent phase sera. They were confirmed as echovirus type 25 and 30 infection, respectively. CONCLUSIONS: These virological examinations are practical and useful for clinical settings for a diagnosis of enterovirus infections because of an insufficient positive rate in virus isolation.

Adult↗

PANORAMA: an integrated Web-based sequence analysis tool and its role in gene discovery.

As the exponential growth of DNA sequence information in databases continues, the task of converting this deposited information into knowledge becomes more dependent on integrative sequence analysis and visualization tools. PANORAMA is an Internet-accessible software package that performs a variety of informatics analyses on a given DNA sequence and returns a visual and interactive representation of the results. Its design is modular, so that further sequence analysis tools can be integrated with minimal effort. The utility of PANORAMA is demonstrated in the analysis of 650 kb of human genomic DNA from chromosome region 3p21.3, a region of potential tumor suppressor genes involved in lung cancer, breast cancer, and other forms of cancer. PANORAMA aided in the discovery of genes and alternate splice forms of known exons, in the demarcation of intron-exon boundaries, and in the identification of promoter regions and polymorphisms, all of which contributed to a better understanding of the region. PANORAMA is available on the World Wide Web at http://atlas.swmed.edu.

Animals↗

Mass spectrometric analysis of permethylated glycosphingolipids I. Sequence analysis of two blood-group B active glycosphingolipids from human B erythrocyte membranes.

Two blood group B active glycosphingolipids (B-I and B-II) formerly isolated and purified from human B erythrocytes (16) were investigated by mass spectrometry after permethylation. B-I yielded fragments up to m/e 1266 and B-II up to m/e 1495, showing the sequence of six and seven carbohydrate residues respectively. In combination with additional experimental evidence (18) the glycosphingolipids are demonstrated to be a gal-[ fuc ]-gal-glcNAc-gal-glc-ceramide (B-I) and a gal-[ fuc ]-gal-glcNAc-gal-glcNAc-gal-glc-ceramide (B-II). Mass spectrometric evidence for the ceramide residues are also obtained indicating besides spingosine C24-,C24:1-, and C22-fatty acids as main constituents.

ABO Blood-Group System↗

The nucleotide sequence of the variable region in Trypanosoma brucei completes the sequence analysis of the maxicircle component of mitochondrial kinetoplast DNA.

The nucleotide sequence of two non-contiguous DNA fragments of 4.0 and 2.2 kb, respectively, of the kinetoplast maxicircle of Trypanosoma brucei brucei EATRO strain 427 has been determined, completing the sequence analysis of the so-called variable region (see also de Vries et al., 1988, Mol. Biochem. Parasitol. 27, 71-82). Analysis of the entire 8-kb variable region sequence revealed the presence of a 5.2-kb cluster of imperfect, tandemly repeated sequences, flanked by DNA of unique sequence. Both repetitive and unique DNA evolve rapidly, but comparison to the closely related strain EATRO 164 indicated that the repetitive cluster is more prone to sequence and size divergence. The variable region is transcribed into RNAs of varying lengths but appears to be devoid of genes encoding mitochondrial proteins or tRNAs, as judged from computer analysis. Moreover, genes that could encode guide RNAs involved in producing the known edited mitochondrial mRNA sequences are also absent. The repetitive DNA cluster within this region consists of 14 blocks each containing one 130 bp repeat and a variable number of 19 bp repeats. A duplicated sequence was identified (5'-GGGGTTGGTGT) which proved to be identical to the eleven 5'-terminal residues of the universal minicircle dodecamer involved in initiation of leading strand synthesis. This suggests a role for these sequences in the initiation of maxicircle DNA replication. With the data presented in this report, the nucleotide sequence analysis of the 23016 bp maxicircle of T. brucei brucei EATRO strain 427 has been completed.

Animals↗

Light-generated oligonucleotide arrays for rapid DNA sequence analysis.

In many areas of molecular biology there is a need to rapidly extract and analyze genetic information; however, current technologies for DNA sequence analysis are slow and labor intensive. We report here how modern photolithographic techniques can be used to facilitate sequence analysis by generating miniaturized arrays of densely packed oligonucleotide probes. These probe arrays, or DNA chips, can then be applied to parallel DNA hybridization analysis, directly yielding sequence information. In a preliminary experiment, a 1.28 x 1.28 cm array of 256 different octanucleotides was produced in 16 chemical reaction cycles, requiring 4 hr to complete. The hybridization pattern of fluorescently labeled oligonucleotide targets was then detected by epifluorescence microscopy. The fluorescence signals from complementary probes were 5-35 times stronger than those with single or double base-pair hybridization mismatches, demonstrating specificity in the identification of complementary sequences. This method should prove to be a powerful tool for rapid investigations in human genetics and diagnostics, pathogen detection, and DNA molecular recognition.

Base Sequence↗

Edman degradation sequence analysis of resin-bound peptides synthesized by 9-fluorenylmethoxycarbonyl chemistry.

The efficacy of Edman degradation sequence analysis for evaluating the synthetic efficiency of peptide-resin assembly by 9-fluorenylmethoxycarbonyl (Fmoc) solid-phase methodology has been studied. Prior researchers have described the use of solid-phase "preview" sequence analysis for peptides synthesized by tertiary-butyloxycarbonyl (Boc) chemistry, where benzyl-based side-chain protecting groups and peptide-resin linkers are stable to the conditions of Edman chemistry. We have successfully sequenced a variety of resin-bound peptides synthesized by Fmoc chemistry, where tertiary-butyl-based side-chain protecting groups and peptide-resin linkers are labile to the conditions of Edman chemistry. Crude peptides are liberated from trifluoroacetic acid-labile linkers during the first cycle of Edman degradation and subsequently "embedded" in membranes. For peptides up to 20 residues, embedded sequencing repetitive yields were comparable to those of solid-phase sequencing. Preview sequencing of resin-bound Fmoc-synthesized peptides proved to be advantageous compared to other analytical methods, in that synthetic failures were detected and quantitated at the point of occurrence, regardless of whether incomplete Fmoc deprotection or incomplete coupling was responsible, and without interference from by-products formed during peptide-resin cleavage. Quantitative ninhydrin analysis, which previously has been found to give false positive results due to removal of the Fmoc group by a combination of reagents and high temperature, gave false negative results in this study, most probably due to incomplete removal of the Fmoc group prior to coupling. Quantitative sequence analysis results were supported by high-performance liquid chromatographic, amino acid and electrospray mass spectrometric analyses of the crude and purified peptides.

Amino Acid Sequence↗

BLMT: statistical sequence analysis using N-grams.

UNLABELLED: Statistical analysis of amino acid and nucleotide sequences, especially sequence alignment, is one of the most commonly performed tasks in modern molecular biology. However, for many tasks in bioinformatics, the requirement for the features in an alignment to be consecutive is restrictive and "n-grams" (aka k-tuples) have been used as features instead. N-grams are usually short nucleotide or amino acid sequences of length n, but the unit for a gram may be chosen arbitrarily. The n-gram concept is borrowed from language technologies where n-grams of words form the fundamental units in statistical language models. Despite the demonstrated utility of n-gram statistics for the biology domain, there is currently no publicly accessible generic tool for the efficient calculation of such statistics. Most sequence analysis tools will disregard matches because of the lack of statistical significance in finding short sequences. This article presents the integrated Biological Language Modeling Toolkit (BLMT) that allows efficient calculation of n-gram statistics for arbitrary sequence datasets. AVAILABILITY: BLMT can be downloaded from http://www.cs.cmu.edu/~blmt/source and installed for standalone use on any Unix platform or Unix shell emulation such as Cygwin on the Windows platform. Specific tools and usage details are described in a "readme" file. The n-gram computations carried out by the BLMT are part of a broader set of tools borrowed from language technologies and modified for statistical analysis of biological sequences; these are available at http://flan.blm.cs.cmu.edu/.

Algorithms↗

Trypanosoma cruzi: sequence analysis of the variable region of kinetoplast minicircles.

The comparisons of 170 sequences of kinetoplast DNA minicircle hypervariable region obtained from 19 stocks of Trypanosoma cruzi and 2 stocks of Trypanosoma cruzi marenkellei showed that only 56% exhibited a significant homology one with other sequences. These sequences could be grouped into homology classes showing no significant sequence similarity with any other homology group. The 44% remaining sequences thus corresponded to unique sequences in our data set. In the DTU I ("Discrete Typing Units") 51% of the sequences were unique. In contrast, in the DTU IId, 87.5% of sequences were distributed into three classes. The results obtained for T. cruzi marinkellei, showed that all sequences were unique, without any similarity between them and T. cruzi sequences. Analysis of palindromes in all sequence sets show high frequency of the EcoRI site. Analysis of repetitive sequences suggested a common ancestral origin of the kDNA. The editing mechanism that occurs in kinetoplastidae is discussed.

Animals↗

Differentiation of Actinobacillus pleuropneumoniae strains by sequence analysis of 16S rDNA and ribosomal intergenic regions, and development of a species specific oligonucleotide for in situ detection.

The aims of this study were to characterize and determine intraspecies and interspecies relatedness of Actinobacillus pleuropneumoniae to Actinobacillus lignieresii and Actinobacillus suis by sequence analysis of the ribosomal operon and to find a species-specific area for in situ detection of A. pleuropneumoniae. Amplification and sequence analysis of the 16S-23S rDNA ribosomal intergenic sequence (RIS) from the three species showed the existence of two RIS's, differing by about 100 bp. Both sequences contained a region resembling the ribonuclease III cleavage site found in Escherichia coli. The smaller RIS contained a Glu-tRNA gene, and the larger one contained genes encoding Ile-tRNA and Ala-tRNA. These tRNA's showed a high sequence homology to the respective tRNA genes found in E. coli. Sequence analysis of the RIS's showed a high degree of genetic similarity of 24 strains of A. pleuropneumoniae. The larger RIS's were different between the 3 species tested. The sequence of the 16S ribosomal gene was determined for 8 serotypes of A. pleuropneumoniae. These sequences showed only minor base differences, indicating a close genetic relatedness of these serotypes within the species. An oligonucleotide DNA probe designed from the 16S rRNA gene sequence of A. pleuropneumoniae was specific for all strains of the target species and did not cross react with A. lignieresii, the closest known relative of A. pleuropneumoniae. This species-specific DNA probe labeled with fluorescein was used for in situ hybridization experiments to detect A. pleuropneumoniae in biopsies of diseased porcine lungs.

Actinobacillus Infections↗

Cartography of ribosomal proteins of the 30S subunit from the halophilic Haloarcula marismortui and complete sequence analysis of protein HS26.

By two-dimensional polyacrylamide gel electrophoresis of 30S ribosomal subunit proteins (S proteins) from Haloarcula marismortui we identified 27 distinct spots and analyzed all of them by protein sequence analysis. We demonstrated that protein HmaS2 (HS2) is encoded by the open reading frame orfMSG and has sequence similarities to the S2 ribosomal protein family. The proteins HmaS5 and HmaS14 were identified as spots HS7 and HS21/HS22, respectively. Protein HS4 was characterized by amino-terminal sequence analysis. The spot HS25 was recognized as an individual protein and also characterized by sequence analysis. Furthermore, the complete primary sequence of HS26 is reported, showing similarity only to eukaryotic ribosomal proteins. The sequence data of a further basic protein shows a high degree of similarity to ribosomal protein S12, therefore, it was designated HmaS12. Slightly different results compared to published sequence data were obtained for the protein HS12 and HmaS19. The putative 'ribosomal' protein HSH could not be localized in the two-dimensional pattern of the total 30S ribosomal subunit proteins of H. marismortui. Therefore, it seems to be unlikely that this protein is a real constituent of the H. marismortui ribosome.

Amino Acid Sequence↗

Rapid grouping of HIV-1 infection in subtypes A to E by V3 peptide serotyping and its relation to sequence analysis.

We developed a typing assay for HIV-1 using subtype specific peptides corresponding to the five major subtypes of HIV-1 (A to E). In eight patients serologically subtyped as A (n = 1), B (n = 3), C (n = 3) and E (n = 1) phyllogenetic analysis of sequenced V3 domain DNA completely correlated to the peptide serotyping. Out of 106 HIV-1 seropositive samples of a diverse geographical origin 88 (83%) could be subtyped by the peptide assay. Five were of subtype A, 33 of subtype B, 48 of subtype C, one of subtype D, and one was of subtype E. Swedish patients were mainly of HIV-1 subtype B and Ethiopian patients were mainly of subtype C, confirming the performance of the assay. Furthermore, subtype specific antibodies may persist up to nine years in HIV-1 infected patients though sera close to AIDS diagnosis may be difficult to type.

Africa↗

Characterization of the genomic organization of human carcinoembryonic antigen (CEA): comparison with other family members and sequence analysis of 5' controlling region.

A cosmid containing the entire coding region for human carcinoembryonic antigen has been isolated. Detailed analysis and sequencing have determined an organization comprising nine exons encoding amino acids and one for a 3' untranslated fragment. Comparison with other family members reveals a complex pattern of homology at the 3' end of the gene. The 5' noncoding region is rich in purine-rich motifs and possible enhancer elements and has a region with properties similar to those of HTF islands.

Amino Acid Sequence↗

The 'shortmer' approach to nucleic acid sequence analysis. I: Computer simulation of sequencing projects to find economical primer sets.

In principle it is most economical to sequence large DNA fragments consecutively ('primer walking'), provided there is an immediate supply of sequencing primers. To solve the problem of primer supply we previously suggested generating a bank of short oligonucleotide primers ('shortmers'). In every sequencing reaction shortmers would have to be selected from this bank that are suitable to hybridize adjacently on the sequencing template. After their ligation the shortmers would form a long, and hence more specific, primer in the subsequent sequencing reaction. In the present study a computer simulation of large sequencing projects revealed a reduced set of approximately 12,000 selected octanucleotides (out of all 65,536) retaining maximum priming flexibility and minimum redundant information on the simulated sequence analyses. Establishing routine protocols for nucleic acid sequencing following the shortmer approach will abolish the tightest bottleneck of the consecutive sequencing route (primer supply) and hence may render this general scheme more attractive than the shotgun sequencing scheme. A twofold (or more) speed-up of genome sequencing projects by the shortmer approach may be assumed.

Algorithms↗

Cloning and sequence analysis of a cDNA plasmid for one of the rat liver glutathione S-transferase subunits.

We describe the construction and characterization of a cDNA plasmid for one of the rat liver glutathione S-transferase subunits. Poly(A)-RNA isolated from rat livers was enriched for glutathione S-transferase mRNA activity and used as templates to synthesize double stranded cDNA. The double stranded cDNAs were annealed to pBR322 through terminal deoxynucleotidyl transferase generated GC-tails followed by transformation into E. coli. Several candidate clones were selected by colony hybridization using polynucleotide kinase labeled liver and testis poly(A)-RNA probes. These candidate clones were further characterized by hybrid-selected translation of mRNA followed by immunoprecipitation and SDS gel electrophoresis. The positive clone, pGTR112 was mapped with restriction endonuclease analysis and sequenced by the chemical method of Maxam and Gilbert. The largest upen reading frame contains 142 amino acids very rich in Arg and Lys residues. The C-terminal residue phenylalanine of this open reading frame is consistent with what was reported for one of the ligandin subunits by Bhargava et al., (J. Biol. Chem. 253, 4116-4119, 1978). Among the 352 nucleotides covered by both pGTR112 and pGST94 described by Kalinyak and Taylor (J. Biol. Chem. 257, 523-530, 1982), there are only 9 nucleotide differences resulting in four changes of amino acid sequences.

Amino Acid Sequence↗

WIT: integrated system for high-throughput genome sequence analysis and metabolic reconstruction.

The WIT (What Is There) (http://wit.mcs.anl.gov/WIT2/) system has been designed to support comparative analysis of sequenced genomes and to generate metabolic reconstructions based on chromosomal sequences and metabolic modules from the EMP/MPW family of databases. This system contains data derived from about 40 completed or nearly completed genomes. Sequence homologies, various ORF-clustering algorithms, relative gene positions on the chromosome and placement of gene products in metabolic pathways (metabolic reconstruction) can be used for the assignment of gene functions and for development of overviews of genomes within WIT. The integration of a large number of phylogenetically diverse genomes in WIT facilitates the understanding of the physiology of different organisms.

Databases, Factual↗