PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Fluorescence detection in automated DNA sequence analysis.

We have developed a method for the partial automation of DNA sequence analysis. Fluorescence detection of the DNA fragments is accomplished by means of a fluorophore covalently attached to the oligonucleotide primer used in enzymatic DNA sequence analysis. A different coloured fluorophore is used for each of the reactions specific for the bases A, C, G and T. The reaction mixtures are combined and co-electrophoresed down a single polyacrylamide gel tube, the separated fluorescent bands of DNA are detected near the bottom of the tube, and the sequence information is acquired directly by computer.

Automation↗

Sequence analysis of a 282-kilobase region surrounding the citrus Tristeza virus resistance gene (Ctv) locus in Poncirus trifoliata L. Raf.

Citrus tristeza virus (CTV) is the major virus pathogen causing significant economic damage to citrus worldwide, and a single dominant gene, Ctv, provides broad spectrum resistance to CTV in Poncirus trifoliata L. Raf. Ctv was physically mapped to a 282-kb region using a P. trifoliata bacterial artificial chromosome library. This region was completely sequenced to about 8x coverage using a shotgun sequencing strategy and primer walking for gap closure. Sequence analysis predicts 22 putative genes, two mutator-like transposons and eight retrotransposons. This sequence analysis also revealed some interesting features of this region of the P. trifoliata genome: a disease resistance gene cluster with seven members and eight retrotransposons clustered in a 125-kb gene-poor region. Comparative sequence analysis suggests that six genes in the Ctv region have significant sequence similarity with their orthologs in bacterial artificial chromosome clones F7H2 and F21T11 from Arabidopsis chromosome I. However, the analysis of gene colinearity between P. trifoliata and Arabidopsis indicates that Arabidopsis genome sequence information may be of limited use for positional gene cloning in P. trifoliata and citrus. Analysis of candidate genes for Ctv is also discussed.

Amino Acid Sequence↗

Molecular cloning and sequence analysis of the human parainfluenza 3 virus gene encoding the L protein.

The sequence of the gene encoding the L protein of the human parainfluenza 3 virus was determined by direct dideoxy sequence analysis of the genomic 50 S RNA and confirmed by molecular cloning and sequence analysis of recombinant clones. A series of three overlapping clones was generated by primer extension using genomic 50 S RNA as the template. These clones originate within the 5' end of the hemagglutinin-neuraminidase gene, span the entire L gene, and extend into the extracistronic 5' end of the viral RNA. The L gene extends 6755 nucleotides (inclusive of the putative transcription initiation and polyadenylation signal sequences) and encodes a protein consisting of 2233 amino acids (MW 255,812). There are 44 nucleotides downstream of the putative polyadenylation signal sequence which may represent a negative-strand leader. The complementary sequence of the extracistronic region is nearly identical to the 3' end of the viral RNA. Thirty-three of the first thirty-nine nucleotides of the 3' ends of the plus and minus strands are conserved. Comparison of amino acid sequence homology with other paramyxoviral L proteins indicates a high degree of sequence conservation with Sendai virus (62%) and Newcastle disease virus (28%). In addition, four smaller regions were identified which shared extensive homology with the L protein of vesicular stomatitis virus, a member of the Rhabdoviridae family.

Amino Acid Sequence↗

Sequence analysis by additive scales: DNA structure for sequences and repeats of all lengths.

MOTIVATION: DNA structure plays an important role in a variety of biological processes. Different di- and tri-nucleotide scales have been proposed to capture various aspects of DNA structure including base stacking energy, propeller twist angle, protein deformability, bendability, and position preference. Yet, a general framework for the computational analysis and prediction of DNA structure is still lacking. Such a framework should in particular address the following issues: (1) construction of sequences with extremal properties; (2) quantitative evaluation of sequences with respect to a given genomic background; (3) automatic extraction of extremal sequences and profiles from genomic databases; (4) distribution and asymptotic behavior as the length N of the sequences increases; and (5) complete analysis of correlations between scales. RESULTS: We develop a general framework for sequence analysis based on additive scales, structural or other, that addresses all these issues. We show how to construct extremal sequences and calibrate scores for automatic genomic and database extraction. We show that distributions rapidly converge to normality as Nincreases. Pairwise correlations between scales depend both on background distribution and sequence length and rapidly converge to an analytically predictable asymptotic value. For di- and tri-nucleotide scales, normal behavior and asymptotic correlation values are attained over a characteristic window length of about 10-15 bp. With a uniform background distribution, pairwise correlations between empirically-derived scales remain relatively small and roughly constant at all lengths, except for propeller twist and protein deformability which are positively correlated. There is a positive (resp. negative) correlation between dinucleotide base stacking (resp. propeller twist and protein deformability) and AT-content that increases in magnitude with length. The framework is applied to the analysis of various DNA tandem repeats. We derive exact expressions for counting the number of repeat unit classes at all lengths. Tandem repeats are likely to result from a variety of different mechanisms, a fraction of which is likely to depend on profiles characterized by extreme structural features.

Animals↗

[Sequence analysis of translocation t (X; 18) genomic breakpoints characterized in synovial sarcoma].

OBJECTIVE: To analyze the DNA sequence characteristics of translocation t (X; 18) genomic breakpoints and to study the mechanism underlying chromosomal translocation t (X; 18) in synovial sarcoma. METHODS: Two cases of synovial sarcoma were studied utilizing long-distance polymerase chain reaction (PCR) and sequence analysis to amplify the genomic DNA of translocation t (X; 18) breakpoints. RESULTS: Translocation t (X; 18) was detected in both cases, which generated SYT-SSX1 and SYT-SSX2 fusion gene respectively. Sequence analysis revealed that intron 10 of SYT was fused to the intron 4 of SSX1 or SSX2. Sequences highly homologous to consensus recognition motifs of translin were found adjacent to breakpoints in all three genes. Breakpoints in the three genes were close to or even at several palindromic oligomer sequences. The breaks in intron 4 of SSX1 and SSX2 were near an Alu sequence. No Alu or other repetitive sequences were found 500 bp upstream or downstream from the break in intron 10 of SYT. One topoisomerase II consensus site was found between the two breakpoints but with considerable distance from intron 10 of SYT. CONCLUSIONS: All three genes involved in synovial sarcomas contain characteristic sequence motifs in the breakpoint regions which may play an important role in the genesis of chromosomal translocation in synovial sarcoma.

Base Sequence↗

Comparative sequence analysis of 634 kb of the mouse chromosome 16 region of conserved synteny with the human velocardiofacial syndrome region on chromosome 22q11.2.

Mouse genomic DNA sequence extending 634 kb on proximal mouse chromosome 16 was compared to the corresponding human sequence from chromosome 22q11.2. Haploinsufficiency for this region results in velocardiofacial syndrome (VCFS) in humans. The mouse region is rearranged into three conserved blocks relative to human, but gene content and position are highly conserved within these blocks. Examination of the boundaries of one of these blocks suggested that the evolutionary chromosomal rearrangement occurred in the mouse lineage, resulting in inactivation of the mouse orthologue of ZNF74. Sequence analysis identified 21 genes and 15 ESTs. These include 2 novel genes, Srec2 and Cals2, and previously undescribed splice variants of several other genes. Exon discovery was carried out using GRAIL2, MZEF, or comparative analysis across 491 kb of conserved mouse and human sequence. Sequence comparison was highly effective, identifying every gene and nearly every exon without the high frequency of false-positive predictions seen when algorithmic methods were used alone. In combination, these procedures identified every gene with no false-positive predictions. Comparative sequence analysis also revealed regions of extensive conservation among noncoding sequences, accounting for 6% of the sequence. A library of such sequences has been established to form a resource for generalized studies of regulatory and structural elements.

Abnormalities, Multiple↗

Primary sequence analysis and representation techniques in carbohydrates.

Sequence similarity calculations of carbohydrates present several problems which must be addressed if a computer implementation is to be achieved. These problems range from the computational representation of the complex carbohydrate structure to the method by which the comparison of residue and linkage is to be made. This paper therefore discusses the form of this representation and how two or more carbohydrates can be meaningful compared. An example set of results using this approach is presented and discussed to illustrate how similarity comparison can show relationships between carbohydrates, features that are otherwise hidden by the sheer volume of data which must be considered.

Algorithms↗

The cloning and sequence analysis of the aspC and tyrB genes from Escherichia coli K12. Comparison of the primary structures of the aspartate aminotransferase and aromatic aminotransferase of E. coli with those of the pig aspartate aminotransferase isoenzymes.

In this paper we describe the cloning and sequence analysis of the tyrB and aspC genes from Escherichia coli K12, which encode the aromatic aminotransferase and aspartate aminotransferase respectively. The tyrB gene was isolated from a cosmid carrying the nearby dnaB gene, identified by its ability to complement a dnaB lesion. Deletion and linker insertion analysis located the tyrB gene to a 1.7-kilobase NruI-HindIII-digest fragment. Sequence analysis revealed a gene encoding a 43 000 Da polypeptide. The gene starts with a GTG codon and is closely followed by a structure resembling a rho independent terminator. The aspC gene was cloned by screening gene banks, prepared from a prototrophic E. coli K12 strain, for plasmids able to complement the aspC tyrB lesions in the aminotransferase-deficient strain HW225. Sub-cloning and deletion analysis located the aspC gene on a 1.8-kilobase HincII-StuI-digest fragment. Sequence analysis revealed the presence of a gene encoding a 43 000 Da protein, the sequence of which is identical with that previously obtained for the aspartate aminotransferase from E. coli B. Considerable overproduction of the two enzymes was demonstrated. We compared the deduced protein sequences with those of the pig mitochondrial and cytoplasmic aspartate aminotransferases. From the extensive homology observed we are able to propose that the two E. coli enzymes possess subunit structures, subunit interactions and coenzyme-binding and substrate-binding sites that are very similar both to each other and to those of the mammalian enzymes and therefore must also have very similar catalytic mechanisms. Comparison of the aspC and tyrB gene sequences reveals that they appear to have diverged as much as is possible within the constraints of functionality and codon usage.

Animals↗

[Cloning and sequence analysis of human uric acid transporter gene].

OBJECTIVE: To obtain full-length human urate transporter (hUAT) gene. METHODS: Primers was designed according to the sequence of hUAT reported in Genbank. The target fragments were obtained by reverse transcriptional (RT) PCR from the mRNA extracted from human renal tubular epithelial cell lines (HK-2), followed by cloning into pEGFP-C1 plasmid. The cloned fragment was subjected to restriction mapping and sequence analysis for confirmation. RESULTS AND CONCLUSION: hUAT gene with correct sequence is successfully cloned from HK-2 cells. The restriction maps and sequence analysis of the selected clones were consistent with the sequence reported in Genbank.

Base Sequence↗

A modular class-aware workflow for small RNA sequencing analysis using mouse sperm as a case study.

BACKGROUND: Small RNA sequencing analysis is challenging because RNA classes differ in biogenesis, sequence redundancy, genomic organization, and annotation reliability. Integrated workflows accommodating these constraints remain limited, particularly for fragment-level and cluster-level analysis. METHODS: We present a reproducible, containerized, class-aware workflow for small RNA sequencing analysis, using mouse sperm as a case study. The workflow combines standardized preprocessing with complementary annotation and quantification strategies for microRNAs (miRNAs), transfer RNA-derived small RNAs (tsRNAs), ribosomal RNA-derived small RNAs (rsRNAs), and PIWI-interacting RNA (piRNA)-enriched genomic clusters. Using sperm small RNA data from offspring of lipopolysaccharide (LPS)-exposed male mice, we compared integrated-reference mapping, multi-class annotation, fragment-level tsRNA profiling, and genome-based piRNA cluster analysis, with custom modules for locus-aware harmonization and condition-specific cluster analysis. RESULTS: Integrated-reference mapping aligned 88.17% of reads and retained 690 features after filtering. It identified 11 differentially expressed miRNAs between LPS and controls, while other classes showed limited signal. Fragment-level profiling improved tsRNA resolution. piRNA cluster analysis identified 958 control and 940 LPS clusters, with 18 control-specific and no LPS-specific clusters. CONCLUSION: This workflow supports transparent, reproducible, class-aware interpretation of small RNA sequencing data while emphasizing cautious interpretation of piRNA-enriched signals from total small RNA sequencing.

Small non-coding RNA analysis↗

Prediction of whole-genome DNA-DNA similarity, determination of G+C content and phylogenetic analysis within the family Pasteurellaceae by multilocus sequence analysis (MLSA).

Genome predictions based on selected genes would be a very welcome approach for taxonomic studies, including DNA-DNA similarity, G+C content and representative phylogeny of bacteria. At present, DNA-DNA hybridizations are still considered the gold standard in species descriptions. However, this method is time-consuming and troublesome, and datasets can vary significantly between experiments as well as between laboratories. For the same reasons, full matrix hybridizations are rarely performed, weakening the significance of the results obtained. The authors established a universal sequencing approach for the three genes recN, rpoA and thdF for the Pasteurellaceae, and determined if the sequences could be used for predicting DNA-DNA relatedness within the family. The sequence-based similarity values calculated using a previously published formula proved most useful for species and genus separation, indicating that this method provides better resolution and no experimental variation compared to hybridization. By this method, cross-comparisons within the family over species and genus borders easily become possible. The three genes also serve as an indicator of the genome G+C content of a species. A mean divergence of around 1 % was observed from the classical method, which in itself has poor reproducibility. Finally, the three genes can be used alone or in combination with already-established 16S rRNA, rpoB and infB gene-sequencing strategies in a multisequence-based phylogeny for the family Pasteurellaceae. It is proposed to use the three sequences as a taxonomic tool, replacing DNA-DNA hybridization.

Base Composition↗

Peptide and protein sequence analysis by electron transfer dissociation mass spectrometry.

Peptide sequence analysis using a combination of gas-phase ion/ion chemistry and tandem mass spectrometry (MS/MS) is demonstrated. Singly charged anthracene anions transfer an electron to multiply protonated peptides in a radio frequency quadrupole linear ion trap (QLT) and induce fragmentation of the peptide backbone along pathways that are analogous to those observed in electron capture dissociation. Modifications to the QLT that enable this ion/ion chemistry are presented, and automated acquisition of high-quality, single-scan electron transfer dissociation MS/MS spectra of phosphopeptides separated by nanoflow HPLC is described.

Amino Acid Sequence↗

The synthesis of oligonucleotides containing an aliphatic amino group at the 5' terminus: synthesis of fluorescent DNA primers for use in DNA sequence analysis.

A rapid and versatile method has been developed for the synthesis of oligonucleotides which contain an aliphatic amino group at their 5' terminus. This amino group reacts specifically with a variety of electrophiles, thereby allowing other chemical species to be attached to the oligonucleotide. This chemistry has been utilized to synthesize several fluorescent derivatives of an oligonucleotide primer used in DNA sequence analysis by the dideoxy (enzymatic) method. The modified primers are highly fluorescent and retain their ability to specifically prime DNA synthesis. The use of these fluorescent primers in DNA sequence analysis will enable DNA sequence analysis to be automated.

Base Sequence↗

High-sensitivity sequence analysis of peptides and proteins by 4-NN-dimethylaminoazobenzene 4'-isothiocyanate.

A manual high-sensitivity sequencing method is described, in which 4-NN-dimethylaminoazobenzene 4'-isothiocyanate is used for the stepwise degradation of amino acid residues from the peptides. The 4-NN-dimethylaminoazobenzene 4'-thiazolinones of amino acids that were released, after conversion into their thiohydantoin derivatives, were identified by t.l.c. on polyamide sheets. This new method is simple and sensitive, and requires only 2-10nmol of peptides or proteins for extended sequence analysis. The method was tested on the sequence analysis of a hexapeptide (Leu-Trp-Met-Arg-Phe-Ala), bradykinin, glucagon and native lysozyme. Results show that the proposed procedure is a sensitive method for the sequence determination of short peptides as well as for the partial sequence determination of intact proteins.

Amino Acid Sequence↗

SeqHepB: a sequence analysis program and relational database system for chronic hepatitis B.

SeqHepB is a combination of a HBV genome sequence analysis program and a relational database that houses data collected from multiple data sources. Registered users can access the sequence analysis component of SeqHepB online for rapid and detailed interrogation of HBV genomic sequences. Its main function is to determine the HBV genotype, identify key mutations associated with antiviral resistance, and identify clinically important HBV mutants. All information generated is uploaded into a database and integrated with patient medical records, pathology laboratory tests, and supplemental virology results such as in vitro drug cross-resistance values. Combined with structured query language (SQL) queries developed in the database, it is possible to extract and correlate clinical, virological, and in vitro phenotypic data rapidly and efficiently. An important component of SeqHepB is its ability to integrate mutations detected within the reverse transcriptase (RT) and locate them onto a three-dimensional (3D) model of the HBV RT that can be viewed at any angle with known antiviral drug molecules in the catalytic pocket of the enzyme. SeqHepB will enable virologists and physicians to individualise patient management, cope with the explosion of antiviral associated HBV mutations, and to conduct cross-sectional retrospective or prospective studies on HBV-infected individuals during therapy.

DNA Mutational Analysis↗

PowerBLAST: a new network BLAST application for interactive or automated sequence analysis and annotation.

As the rate of DNA sequencing increases, analysis by sequence similarity search will need to become much more efficient in terms of sensitivity, specificity, automation potential, and consistency in annotation. PowerBLAST was developed, in part, to address these problems. PowerBLAST includes a number of options for masking repetitive elements and low complexity subsequences. It also has the capacity to restrict the search to any level of NCBI's taxonomy index, thus supporting "comparative genomics" applications. Postprocessing of the BLAST output using the SIM series of algorithms produces optimal, gapped alignments, and multiple alignments when a region of the query sequence matches multiple database sequences. PowerBLAST is capable of processing sequences of any length because it divides long query sequences into overlapping fragments and then merges the results after searching. The results may be viewed graphically, as a textual representation, or as an HTML page with links to GenBank and Entrez. For matching database sequences, annotated features are superimposed on the aligned query sequence in the output, thus greatly increasing the ease of interpretation. Such features may be used for automated annotation of new sequence because PowerBLAST output in ASN.1 form may be "dragged and dropped" into NCBI's Sequin program for sequence annotation and submission. PowerBLAST is capable of analyzing and annotating a 100-kb query in 60 min on NCBI's BLAST server.

Amino Acid Sequence↗

Typing of Candida glabrata in clinical isolates by comparative sequence analysis of the cytochrome c oxidase subunit 2 gene distinguishes two clusters of strains associated with geographical sequence polymorphisms.

We tested whether comparative sequence analysis of the mitochondrion-encoded cytochrome c oxidase subunit 2 gene (COX2) could be used to distinguish intraspecific variants of Candida glabrata. Mitochondrial genes are suitable for investigation of close phylogenetic relationships because they evolve much faster than nuclear genes, which in general exhibit very limited intraspecific variation. For this survey we used 11 clinical isolates of C. glabrata from three different geographical locations in Brazil, 10 isolates from one location in the United States, 1 American Type Culture Collection strain as an internal control, and the published sequence of strain CBS 138. The complete coding region of COX2 was amplified from total cellular DNA, and both strands were sequenced twice for each strain. These sequences were aligned with published sequences from other fungi, and the numbers of substitutions and phylogenetic relationships were determined. Typing of these strains was done by using 17 substitutions, with 8 being nonsynonymous and 9 being synonymous. Also, cDNAs made from purified mitochondrial polyadenylated RNA were sequenced to confirm that our sequences correspond to the expressed copies and not nuclear pseudogenes and that a frameshift mutation exists in the 3' end of the coding region (position 673) relative to the Saccharomyces cerevisiae sequence and the previously published C. glabrata sequence. We estimated the average evolutionary rate of COX2 to be 11.4% sequence divergence/10(8) years and that phylogenetic relationships of yeasts based on these sequences are consistent with rRNA sequence data. Our analysis of COX2 sequences enables typing of C. glabrata strains based on 13 haplotypes and suggests that positions 51 and 519 indicate a geographical polymorphism that discriminates strains isolated in the United States and strains isolated in Brazil. This provides for the first time a means of typing of Candida strains that cause infections by use of direct sequence comparisons and the associated divergence estimates.

Bacterial Proteins↗

Molecular typing of Toxoplasma gondii strains by GRA6 gene sequence analysis.

The utility of sequence polymorphisms in the dense granule antigen GRA6 gene as typing markers for Toxoplasma gondii was investigated. The coding region of GRA6 was amplified, sequenced and compared for 30 Toxoplasma strains from eight different zymodemes (Z1-Z8). Sequence alignment identified nucleotide polymorphisms at 24 positions out of 690 bp, which correlated with murine-virulence. Types I, II, and III could be distinguished from each other on the basis of three, 10, and six variable positions, respectively. Two deletions of 15 bp and 3 bp existed in the avirulent (type II) strains. With one exception, all polymorphic positions resulted in amino acid substitutions, and the two gaps of 15 bp and 3 bp caused the deletion of six amino acids in type II strains. Intra-specific polymorphisms were also found in the virulent group. A high degree of sequence polymorphism correlating with the phenotypes of T. gondii strains points to the GRA6 gene being a good marker for strain characterisation and typing of the isolates of this apicomplexan. The large variety of amino acid changes supports the view that the GRA6 protein plays an important role in the antigenicity and pathogenicity of T. gondii. The existence of polymorphic restriction sites for endonuclease MseI was used to develop a PCR-RFLP method which could simply differentiate the three different groups (types I, II, III) of T. gondii.

Animals↗