PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Base Sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

High-resolution HLA typing for the DRB3/4/5 genes by sequence-based typing.

The high degree of polymorphism of the HLA genes at the nucleotide sequence level has proven sequence-based typing a major typing strategy. For DRB1 the allelic variability is predominantly present in the second exon and by DNA sequencing of exon 2 all hitherto known DRB1 alleles can be detected. For the associated genes DRB3, DRB4 and DRB5 the situation is slightly different. Allelic differences are not limited to exon 2 and the sequence of exon 3 and sometimes exon 4 is needed for complete subtyping. Oligonucleotides to amplify the exons needed for subtyping of DRB3, DRB4 and DRB5 were designed. Gene-specific products were generated to make simultaneous detection of alleles in heterozygous combinations possible. In this way 238 individuals were fully typed for their DRB3, 4 and 5 subtypes. Additional samples were typed for only one of the genes. All samples had been previously typed by PCR-SSP. Concordant typing results were obtained for all individuals tested. The DRB3 alleles typed for included *0101, *0201, *0202 and *0301, for DRB4 they were *01011, *0102 and *0103 and for DRB5 *0101, *0102, *0103, *0105, *0201, *0202 and *0203. All alleles were easily detected by the protocol described except for DRB5*0201. Sequencing of exon 3 and 4 of the DRB5*0201 allele showed this allele to be a sequencing error and the sequences obtained were identical to the exon 2, 3 and 4 sequences of DRB5*0202. Two new alleles were identified in the samples studied, DRB4*0105 and DRB3*0207. Sequence based typing has been recognized as a valuable tool for HLA typing of DRB1, DQB1 and DPB1 since several years. It is shown to be a superior typing method as well in the detection of the different DRB3, 4 and 5 subtypes.

Alleles↗

Assign 2.0: software for the analysis of Phred quality values for quality control of HLA sequencing-based typing.

As improvements to DNA sequencing technology have resulted in increasing the throughput of DNA sequencing, the bottleneck for high throughput DNA sequencing-based typing (SBT) has shifted to sequence analysis, genotyping and quality control (QC). Consistent high-quality DNA sequence is required in order to reduce manual verification and editing of sequence electropherograms. However, identifying systematic changes in quality is difficult to achieve without the aid of sophisticated sequence analysis programs dedicated to this purpose. We describe a computer software program called Assign 2.0, which integrates sequence QC analysis and genotyping in order to facilitate high-throughput SBT. Assign 2.0 performs an analysis of Phred quality values in order to produce quality scores for a sample and a sequencing run. This enables sample-to-sample and run-to-run QC monitoring and provides a mechanism for the comparison of sequence quality between various genes, various reagents and various protocols with the aim of improving the overall quality of DNA sequence data. This, in turn, will result in reducing sequence analysis as a bottleneck for high-throughput SBT.

Alleles↗

DynaPred: a structure and sequence based method for the prediction of MHC class I binding peptide sequences and conformations.

MOTIVATION: The binding of endogenous antigenic peptides to MHC class I molecules is an important step during the immunologic response of a host against a pathogen. Thus, various sequence- and structure-based prediction methods have been proposed for this purpose. The sequence-based methods are computationally efficient, but are hampered by the need of sufficient experimental data and do not provide a structural interpretation of their results. The structural methods are data-independent, but are quite time-consuming and thus not suited for screening of whole genomes. Here, we present a new method, which performs sequence-based prediction by incorporating information obtained from molecular modeling. This allows us to perform large databases screening and to provide structural information of the results. RESULTS: We developed a SVM-trained, quantitative matrix-based method for the prediction of MHC class I binding peptides, in which the features of the scoring matrix are energy terms retrieved from molecular dynamics simulations. At the same time we used the equilibrated structures obtained from the same simulations in a simple and efficient docking procedure. Our method consists of two steps: First, we predict potential binders from sequence data alone and second, we construct protein-peptide complexes for the predicted binders. So far, we tested our approach on the HLA-A0201 allele. We constructed two prediction models, using local, position-dependent (DynaPred(POS)) and global, position-independent (DynaPred) features. The former model outperformed the two sequence-based methods used in our evaluation; the latter shows a much higher generalizability towards other alleles than the position-dependent models. The constructed peptide structures can be refined within seconds to structures with an average backbone RMSD of 1.53 A from the corresponding experimental structures.

Algorithms↗

De novo sequencing, peptide composition analysis, and composition-based sequencing: a new strategy employing accurate mass determination by fourier transform ion cyclotron resonance mass spectrometry.

A new strategy is described for the determination of amino acid sequences of unknown peptides. Different from the well-known but often inefficient de novo sequencing approach, the new method is based on a two-step process. In the first step the amino acid composition of an unknown peptide is determined on the basis of accurate mass values of the peptide precursor ion and a small number of accurate fragment ion mass values, and, as in de novo sequencing, without employing protein database information or other pre-information. In the second step the sequence of the found amino acids of the peptide is determined by scoring the agreement between expected and observed fragment ion signals of the permuted sequences. It was found that the new approach is highly efficient if accurate mass values are available and that it easily outstrips common approaches of de novo sequencing being based on lower accuracies and detailed knowledge of fragmentation behavior. Simple permutation and calculation of all possible amino acid sequences, however, is only efficient if the composition is known or if possible compositions are at least reduced to a small list. The latter requires the highest possible instrumental mass accuracy, which is currently provided only by fourier transform ion cyclotron resonance mass spectrometry. The connection between mass accuracy and peptide composition variability is described and an example of peptide compositioning and composition-based sequencing is presented.

Amino Acid Sequence↗

Base sequence specificity of three 2-chloroethylnitrosoureas.

Chemical modifications of guanine are some of the most common results of interactions of DNA with many carcinogens and anti-cancer drugs, including nitrosoureas, nitrogen mustards, triazenes, polycyclic aromatics, and aflatoxins. The base sequence specificity for alkylation of guanines by three 2-chloroethylnitrosoureas has been determined. Guanines in the midst of a run of guanines are more susceptible than guanines in other base sequences. We have shown that certain 2-chloroethylnitrosoureas (BCNU, CCNU and methyl-CCNU) follow this same pattern. However, the quantitative degree of higher specificity for guanine with guanines as nearest neighbors depended on both the guanine position alkylated and the structure of the alkyl group attached. For example, when hydroxyethylation of runs of guanine occurred at N-7, a 6- to 11-fold increase of alkylation occurred compared to that found in the random base sequences of DNA, while hydroxyethylation at O-6 increased 1.2 to 3.5-fold and chloroethylation at N-7 was 2-to 4-fold higher than in DNA. Guanines with thymines on both the 3' and 5' sides were much less susceptible, most notably in N-7-hydroxyethylation and N-7-chloroethylation. Since guanine-rich regions are found in regulatory regions of the genome, knowledge concerning the effect of base sequence upon the production of each of the potential DNA lesions is vital to gaining an understanding of the roles of these lesions in the anti-tumor activity of a drug.

Alkylation↗

Expression and base sequence of the citrate synthase gene of Acinetobacter anitratum.

The sequence of 1895 base pairs of Acinetobacter anitratum genomic DNA, containing the structural gene for the allosteric citrate synthase of that Gram-negative bacterium, is presented. The sequence contains an open reading frame of 424 codons, the 5' end of which is the same as the N-terminal sequence of A. anitratum citrate synthase, less the initiator methionine. The inferred amino acid sequence of the enzyme is about 70% identical with that of citrate synthase from Escherichia coli, which like the A. anitratum enzyme is sensitive to allosteric inhibition by NADH. There is also a more distant homology with the nonallosteric citrate synthases of pig heart and yeast. The gene contains sequences that strongly resemble those found in E. coli promoters, an E. coli type of ribosomal binding site, and a hyphenated dyad sequence at the 3' end of the gene which resembles the rho-independent terminators found in some E. coli genes. The plasmid clone containing the A. anitratum citrate synthase gene pLJD1, originally isolated because it hybridized with the cloned E. coli citrate synthase gene under conditions of reduced stringency, produces large amounts of A. anitratum citrate synthase in an E. coli host which lacks citrate synthase. This work completes proof of the hypothesis that the three major kinds of citrate synthases are formed of similar subunits, although their functional properties are different.

Acinetobacter↗

Genomic sequence sampling: a strategy for high resolution sequence-based physical mapping of complex genomes.

We present a simple and efficient method for constructing high resolution physical maps of large regions of genomic DNA based upon sampled sequencing. The physical map is constructed by ordering high density cosmid contigs and determining a sequence fragment from each end of every clone. The resulting map, which contains 30-50% of the complete DNA sequence, allows the identification of many genes and makes possible PCR amplification of virtually any part of the genome. We apply this strategy to the automated analysis of the genome of the primitive eukaryote Giardia lamblia and evaluate its applicability to the physical mapping and DNA sequencing of the human genome.

Amino Acid Sequence↗

Nucleotide base sequence of vibrionaceae 5 S rRNA.

Nucleotide base sequences of 5 S rRNAs isolated from Vibrio vulnificus, Vibrio anguillarum, and Aeromonas hydrophila were determined. Comparisons among these and sequences of 5 S rRNAs from other species of Vibrionaceae provide information useful in the evaluation of the evolution of bacterial species.

Aeromonas↗

The advantage of functional prediction based on clustering of yeast genes and its correlation with non-sequence based classifications.

Sequence similarity is probably the most widely used tool to infer functional linkage between proteins. The fully sequenced, much researched, genome of Saccharomyces cerevisiae gives us on opportunity to compare and statistically quantify computational methods based on sequence similarity, which aim to detect such linkage. In addition, the amount of data regarding Saccharomyces Cerevisiae genes and proteins, which is not directly based on sequence is rapidly increasing. Consequently, it allows investigation of the connections and correlation between classification based on these types of data and that based solely on sequence similarity. In this work we start with a simple clustering algorithm to cluster genes based on the BLAST E-score of their similarity. We analyze how well one can infer function from these clusters and for how many of the genes that are currently unknown one can suggest a prediction. Given these parameters, we show that even a simple algorithm achieves better results than simply considering the BLAST output of matching genes. In the second part of the paper, we show that there is a highly significant correlation (p-value < 10(-4) for the vast majority of the experiments) between the aforementioned clusters and other types of classifications. Namely, we show that a pair of genes being clustered together is correlated with these genes having similar expression patterns in DNA array experiments and with the encoded proteins being involved in protein-protein interactions. Although this correlation is highly significant, it is, of course, not strong enough to be, by itself, a tool for predicting co-regulation of genes or interaction of proteins. We discuss possible explanations for this correlation. Furthermore, the statistical evaluation of these results should be considered when developing tools that are aimed at making such predictions.

Algorithms↗

A sequence-based integrated map of chromosome 22.

The near-completion of the sequence for chromosome 22q revolutionizes map integration. We describe a sequence-based integrated map containing 968 loci including 516 known or predicted gene sequences, 317 STSs not included in these sequences, and 135 nonexpressed multinucleotide polymorphisms. The published sequence spans 34.6 Mb, inclusive of gaps estimated to total 1.1 Mb, compared with a top-down estimate of 43 Mb. This discrepancy is discussed, but will not be resolved until more of the genome is analyzed. The radiation hybrid map has 5% error in order and 34% error in location exceeding 1 Mb. The utility of a composite location based on evidence other than sequence is limited to regions not yet sequenced. A genetic map conditional on sequence order was constructed from pairwise lods. Its length of 74.8 cM in males and 80.2 cM in females is slightly less than the previous estimate not constrained by sequence order. Five recombination hot spots are detected, with differences in location between the sexes. Male recombination correlates with repetitive DNA, whereas female recombination does not. It remains to be seen whether this is true for other human chromosomes. An algorithm to improve the fit of cytogenetic bands sequence location reduces the discrepancies in cytogenetic assignment from 61 to 38. This sequence-based integrated map is represented in the genetic location database (LDB2000), which is available at http://cedar.genetics.soton.ac.uk/public_html/LDB2000.html.

Base Sequence↗

Typing single-nucleotide polymorphisms using a gel-based sequencer: a new data analysis tool and suggestions for improved efficiency.

Single-nucleotide polymorphisms (SNPs) are increasingly used as genetic markers. Although a high number of SNP-genotyping techniques have been described, most techniques still have low throughput or require major investments. For laboratories that have access to an automated sequencer, a single-base extension (SBE) assay can be implemented using the ABI SNaPshot trade mark kit. Here we present a modified protocol comprising multiplex template generation, multiplex SBE reaction, and multiplex sample analysis on a gel-based sequencer such as the ABI 377. These sequencers run on a Macintosh platform, but on this platform the software available for analysis of data from the ABI 377 has limitations. First, analysis of the size standard included with the kit is not facilitated. Therefore a new size standard was designed. Second, using Genotyper (ABI), the analysis of the data is very tedious and time consuming. To enable automated batch analysis of 96 samples, with 10 SNPs each, we developed SNPtyper. This is a spreadsheet-based tool that uses the data from Genotyper and offers the user a convenient interface to set parameters required for correct allele calling. In conclusion, the method described will enable any lab having access to an ABI sequencer to genotype up to 1000 SNPs per day for a single experimenter, without investing in new equipment.

Computational Biology↗

Base sequence effects in double helical DNA. I. Potential energy estimates of local base morphology.

A series of potential energy calculations have been carried out to estimate base sequence dependent structural differences in B-DNA. Attention has been focused on the simplest dimeric fragments that can be used to build long chains, computing the energy as a function of the orientation and displacement of the 16 possible base pair combinations within the double helix. Calculations have been performed, for simplicity, on free base pairs rather than complete nucleotide units. Conformational preferences and relative flexibilities are reported for various combinations of the roll, tilt, twist, lateral displacement, and propeller twist of individual residues. The predictions are compared with relevant experimental measures of conformation and flexibility, where available. The energy surfaces are found to fit into two distinct categories, some dimer duplexes preferring to bend in a symmetric fashion and others in a skewed manner. The effects of common chemical substitutions (uracil for thymine, 5-methyl cytosine for cytosine, and hypoxanthine for guanine) on the preferred arrangements of neighboring residues are also examined, and the interactions of the sugar-phosphate backbone are included in selected cases. As a first approximation, long range interactions between more distant neighbors, which may affect the local chain configuration, are ignored. A rotational isomeric state scheme is developed to describe the average configurations of individual dimers and is used to develop a static picture of overall double helical structure. The ability of the energetic scheme to account for documented examples of intrinsic B-DNA curvature is presented, and some new predictions of sequence directed chain bending are offered.

Base Composition↗

Structure-function inferences based on molecular modeling, sequence-based methods and biological data analysis of snake venom lectins.

Lectins are a structurally and functionally diverse group of proteins from different sources, capable to recognize and bind specifically carbohydrates. Several snake venoms contain calcium-dependent true lectins (SVLs) that recognize galactose. Herein, in order to enlighten some of the structure-function relationships of snake venom lectins (SVLs), we constructed theoretical models for 10 SVLs based on the Crotalus atrox lectin (CaL), the only SVL crystal structure available, and compared with other animal and plant lectins, and C-type lectin-like proteins (CLPs) that do not bind carbohydrates. Although these are theoretical structures, we could identify some SVL features, including: (i) a singular intrachain disulfide bond (Cys(38)-Cys(133)) that is not present in CLPs; (ii) a significant reorientation (39-41A) of the 80's loop position that folds back to the globular domain, assists the carbohydrate recognition domain (CRD), and orients the dimer formation, even in BfL-1 and BfL-2, which did not present the Cys(86) interchain; (iii) a CRD presenting a negative and concave surface that allows the interaction with the specific saccharide hydroxyl groups and calcium ion; (iv) the role of water molecules in some interchain interactions, similar to other animal and plant lectins; and (v) the inability of forming oligomers in contrast to CaL and some CLPs, such as convulxin.

Amino Acid Sequence↗

Typing of Mycoplasma pneumoniae by nucleic acid sequence-based amplification, NASBA.

Nucleic acid sequence-based amplification, NASBA, is an isothermal amplification technique for nucleic acids and was used for typing a collection of 24 Mycoplasma pneumoniae strains. A set of primers was chosen from the 16S rRNA sequence alignment of Mycoplasma species. The nucleotide sequences of the (-)RNA amplicons were determined for M. pneumoniae strains M15/83 (type 1) and FH (type 2), and revealed a one-point difference at the 16S rRNA level between the two types. Based on this result, two type-specific probes were constructed. The probes were hybridized in solution with the amplified nucleic acids of 24 M. pneumoniae strains in an enzyme-linked gel assay (ELGA). The results obtained by NASBA-based typing are in agreement with the classification of the 24 M. pneumoniae strains into two types by other typing methods, confirming the reliability of this technique.

Bacterial Typing Techniques↗

Comparison of the base-sequence complexities of polysomal and nuclear RNAs in growing Friend erythroleukemia cells.

The base-sequence complexities of polysomal poly(A)+ RNA, nuclear poly(A)+ RNA, and total nuclear RNA from Friend erythroleukemia cells in logarithmic phase of growth were determined by measuring the proportion of labeled unique mouse DNA sequences which formed hybrids when incubated with a vast excess of each RNA. It was estimated that the RNAs had been transcribed from 1.8, 7.6, and 8.3%, respectively, of the haploid mouse genome. Although these estimates for polysomal and nuclear poly(A)+ RNAs were 2.5 times greater than those previously determined by analysis of the kinetics of the hybridization reactions between the RNAs and the cDNAs transcribed from them, they confirmed that the base-sequence complexity of nuclear poly(A)+ RNA in these cells is at least four times greater than that of polysomal poly(A)+ RNA.

Base Sequence↗