PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequencing libraries”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

Libraries of random-sequence polypeptides produced with high yield as carboxy-terminal fusions with ubiquitin.

Libraries of random-sequence polypeptides have been shown to be valuable sources of novel molecules possessing a variety of useful biologic-like activities, some of which may hold promise as potential vaccines and therapeutics. Previous random peptide expression systems were limited to low levels of peptide production and often to short sequences. Here we describe a series of libraries designed for increased polypeptide length. Cloned as carboxy-terminal extensions of ubiquitin, the fusions were produced in E. coli at high levels, and were purified to homogeneity. The majority of the extension proteins examined could be cleaved from ubiquitin by treatment with a ubiquitin-fusion hydrolase. The libraries described here are appropriate sources of novel polypeptides with desired binding or catalytic function, as well as tools with which to examine inherent properties of proteins as a whole. Toward the latter goal, we have examined structural properties of random-sequence proteins purified from these libraries. Quite surprisingly, fluorescence emission spectra of intrinsic tryptophan residues in several purified fusion proteins, under native-like and denaturing conditions, often resemble those expected for folded and unfolded states, respectively. The results presented here detail an important expansion in the range of potential uses for random-sequence polypeptide libraries.

Amino Acids↗

Cloning of genes encoding redox proteins of known amino acid sequence from a library of the Desulfovibrio vulgaris (Hildenborough) genome.

A library of 900 recombinant phages has been constructed for the genome of Desulfovibrio vulgaris Hildenborough (1.7 x 10(6) bp) by cloning size-fractionated Sau3A fragments (15-20 kb) into the replacement vector lambda-2001. When a hydrogenase gene probe, a 4.7-kb SalI-EcoRI fragment of known nucleotide sequence, was used to screen the plaque lifted library, 23 positive clones were found, which together span 31 kb of D. vulgaris DNA. To facilitate the cloning of genes with oligodeoxynucleotides as probes, DNA was purified for all clones in the library and spotted on a 16 x 16-cm grid of nitrocellulose. This grid was incubated sequentially to identify lambda clones containing the gene for redox proteins of known amino acid sequence: cytochrome c3 (one 18-mer----four clones), flavodoxin (one 17-mer and one 26-mer----one clone) and rubredoxin (one 44-mer----21 clones). The four cyc-positive clones are also recognized by the rubredoxin oligodeoxynucleotide probe. Restriction mapping defines a 35-kb region of the D. vulgaris chromosome in which the rub and cyc loci are separated by 17.5 kb. The nucleotide sequence of the rubredoxin gene was determined and the deduced amino acid sequence found to agree with that determined in Bruschi [Biochim. Biophys. Acta 434 (1976) 4-17] with the exception of Thr-21 which is found to be encoded by GAC, an Asp codon. A plausible ribosome-binding site precedes the N-terminal initiator methionine residue. Rubredoxin does not have an N-terminal signal sequence which is in agreement with the cytoplasmic location of this redox protein.

Amino Acid Sequence↗

Library of sequence-specific radioimmunoassays for human chromogranin A.

BACKGROUND: Human chromogranin A (CgA) is an acidic protein widely expressed in neuroendocrine tissue and tumors. The extensive tissue- and tumor-specific cleavages of CgA at basic cleavage sites produce multiple peptides. METHODS: We have developed a library of RIAs specific for different epitopes, including the NH2 and COOH termini and three sequences adjacent to dibasic sites in the remaining part of CgA. RESULTS: The antisera raised against CgA(210-222) and CgA(340-348) required a free NH2 terminus for binding. All antisera displayed high titers, high indexes of heterogeneity ( approximately 1.0), and high binding affinities (Keff0 approximately 0.1 x 10(12) to 1.0 x 10(12) L/mol), implying that the RIAs were monospecific and sensitive. The concentration of CgA in different tissues varied with the assay used. Hence, in a carcinoid tumor the concentration varied from 0.5 to 34.0 nmol/g tissue depending on the specificity of the CgA assay. The lowest concentration in all tumors was measured with the assay specific for the NH2 terminus of CgA. This is consistent with the relatively low concentrations measured in plasma from carcinoid tumor patients by the N-terminal assay, whereas the assays using antisera raised against CgA(210-222) and CgA(340-348) measured increased concentrations. CONCLUSION: Only some CgA assays appear useful for diagnosis of neuroendocrine tumors, but the entire library is valuable for studies of the expression and processing of human CgA.

Adenoma↗

A genome-wide, end-sequenced 129Sv BAC library resource for targeting vector construction.

The majority of gene-targeting experiments in mice are performed in 129Sv-derived embryonic stem (ES) cell lines, which are generally considered to be more reliable at colonizing the germ line than ES cells derived from other strains. Gene targeting is reliant on homologous recombination of a targeting vector with the host ES cell genome. The efficiency of recombination is affected by many factors, including the isogenicity (H. te Riele et al., 1992, Proc. Natl. Acad. Sci. USA 89, 5128-5132) and the length of homologous sequence of the targeting vector and the location of the target locus. Here we describe the double-end sequencing and mapping of 84,507 bacterial artificial chromosomes (BACs) generated from AB2.2 ES cell DNA (129S7/SvEvBrd-Hprtb-m2). We have aligned these BACs against the mouse genome and displayed them on the Ensembl genome browser, DAS: 129S7/AB2.2. This library has an average insert size of 110.68 kb and average depth of genome coverage of 3.63- and 1.24-fold across the autosomes and sex chromosomes, respectively. Over 97% of the mouse genome and 99.1% of Ensembl genes are covered by clones from this library. This publicly available BAC resource can be used for the rapid construction of targeting vectors via recombineering. Furthermore, we show that targeting vectors containing DNA recombineered from this BAC library can be used to target genes efficiently in several 129-derived ES cell lines.

Animals↗

SRS--an indexing and retrieval tool for flat file data libraries.

SRS (Sequence Retrieval System) is an information indexing and retrieval system designed for libraries with a flat file format such as the EMBL nucleotide sequence databank, the SwissProt protein sequence databank or the Prosite library of protein subsequence consensus patterns. SRS supports the data structure of these libraries by providing special indices for implementing lists of subentities (e.g. feature tables) or hierarchically structured data-fields (e.g. taxonomic classification). A language (ODD) has been designed for the convenient specification of library format and organization, representation of individual data-fields within the system (design of indices) and structuring other data needed during retrieval. This ensures flexibility required for coping with different library formats, which are subject to continuous change. Queries and inspection of retrieved entries can be performed from a user interface with pull-down menus and windows. SRS supports various input and output formats but is particularly well adapted to the GCG programs.

Abstracting and Indexing↗

The occurrence of families of repetitive sequences in a library of cloned cDNA from human lymphocytes.

A library of cloned cDNAs representative of lymphocyte total poly(A)+ RNA was screened with total DNA probes at high clone density. 10% of the recombinants showed the presence of sequences which are repeated in the genome. Further analysis of six such isolated cDNA clones indicated that they contain different families of repetitive sequences with reiteration frequencies of between 150 and 45,000 copies per haploid genomes. Five of the six clones were found to contain single copy sequences as well as a repetitive sequence. cDNA clones containing repetitive sequences have been found to be derived from high, intermediate and low abundance classes of lymphocyte poly(A)+ RNA.

Cloning, Molecular↗

Temperature modifies gene expression in subcuticular epithelial cells of white spot syndrome virus-infected Litopenaeus vannamei.

Subtractive suppressive hybridization was used to identify differentially expressed genes in subcuticular tissues from white spot syndrome virus(WSSV)-infected shrimp kept at different temperatures. Subtractive libraries I and II contained genes expressed at 26 and 33 degrees C, respectively. Three hundred and seventy-nine insert positive clones were selected to confirm differential expression by dot-blot hybridization. Twenty-two clones from library I and eight from library II were sequenced. All sequences from Library I corresponded to white spot syndrome virus genes. From library II, five clones were homologous with previously reported expressed sequence tags of Litopenaeus vannamei, two had similarity with beta-actin and one transcript represented an unknown gene. Over-expression of VP15 in shrimp at 26 degrees C was further confirmed by real-time polymerase chain reaction (PCR), whereas beta-actin expression was similar in animals kept at both temperatures. Together, our results show that hyperthermia reduces the expression of WSSV genes on shrimp subcuticular epithelial cells.

Animals↗

Dynamic programming algorithms for biological sequence comparison.

Efficient dynamic programming algorithms are available for a broad class of protein and DNA sequence comparison problems. These algorithms require computer time proportional to the product of the lengths of the two sequences being compared [O(N2)] but require memory space proportional only to the sum of these lengths [O(N)]. Although the requirement for O(N2) time limits use of the algorithms to the largest computers when searching protein and DNA sequence databases, many other applications of these algorithms, such as calculation of distances for evolutionary trees and comparison of a new sequence to a library of sequence profiles, are well within the capabilities of desktop computers. In particular, the results of library searches with rapid searching programs, such as FASTA or BLAST, should be confirmed by performing a rigorous optimal alignment. Whereas rapid methods do not overlook significant sequence similarities, FASTA limits the number of gaps that can be inserted into an alignment, so that a rigorous alignment may extend the alignment substantially in some cases. BLAST does not allow gaps in the local regions that it reports; a calculation that allows gaps is very likely to extend the alignment substantially. Although a Monte Carlo evaluation of the statistical significance of a similarity score with a rigorous algorithm is much slower than the heuristic approach used by the RDF2 program, the dynamic programming approach should take less than 1 hr on a 386-based PC or desktop Unix workstation. For descriptive purposes, we have limited our discussion to methods for calculating similarity scores and distances that use gap penalties of the form g = rk. Nevertheless, programs for the more general case (g = q+rk) are readily available. Versions of these programs that run either on Unix workstations, IBM-PC class computers, or the Macintosh can be obtained from either of the authors.

Algorithms↗

Bovine beta-crystallin complementary DNA clones. Alternating proline/alanine sequence of beta B1 subunit originates from a repetitive DNA sequence.

A library of recombinant plasmids carrying complementary DNA sequences synthesized from bovine lens messenger RNAs was constructed. Clones coding for five different beta-crystallin subunits: beta B1, beta B3, beta Bp, beta s, beta A3 (and beta A1), were identified by means of hybridization selection, followed by one- and two-dimensional gel electrophoresis of the translational products. Under rather stringent conditions each of these clones hybridizes with its corresponding mRNA and does not show significant cross-hybridization with mRNAs coding for other beta-crystallins, except in the case of the homologous beta A3 and beta A1-crystallins. The beta A3 and beta A1 subunits seem to be encoded by one mRNA using two different AUG codons as start position for translation. We have also determined the nucleotide sequence of a beta B1-crystallin cDNA (pBL beta B1) which enabled us to deduce the complete amino acid sequence of the protein. The beta B1-crystallin, a characteristic component of the high molecular weight crystallin aggregate (beta H), is internally homologous both at DNA and protein level as has been reported for gamma- and other beta-crystallins. This is in agreement with the idea that these proteins had a common ancestral precursor gene that internally duplicated. The G + C content of the coding sequence of beta B1 is very high: 67% overall and even 84.2% for the first 170 nucleotides, due to a remarkable non-random codon usage. A proline/alanine repetition in the N-terminal domain of the protein is encoded by a repetitive "simple" DNA sequence.

Alanine↗

Structural organization and nucleotide sequence of mouse c-myb oncogene: activation in ABPL tumors is due to viral integration in an intron which results in the deletion of the 5' coding sequences.

Bacteriophage libraries of mouse DNA were screened for sequences homologous to the v-myb oncogene and two overlapping clones containing the v-myb related region were isolated. Restriction enzyme mapping, heteroduplex analysis and nucleotide sequence analysis revealed the presence of nine exons. Six of these exons are homologous to the v-myb region while the other three exons are derived from the 5' region which is deleted in the viral oncogene. The sequences downstream to the sixth v-myb exon are not included in the 17 kbp of DNA sequences analyzed in this study. Comparison of the structure of the normal c-myb clone with its rearranged couterpart present in plasmacytoid lymphosarcomas revealed that the rearrangements occur in this locus as a result of viral integration. Present studies demonstrate that such a viral insertion interrupts the c-myb coding region at a region identical to that observed in the generation of the v-myb gene of avian myeloblastosis virus and results in the synthesis of mRNAs that lack the same 5' coding region.

Amino Acid Sequence↗

Analysis of clones carrying repeated DNA sequences in two YAC libraries of Arabidopsis thaliana DNA.

YAC clones carrying repeated DNA sequences from the Arabidopsis thaliana genome have been characterized in two widely used Arabidopsis YAC libraries, the EG library and the EW library. Ribosomal, chloroplast and the paracentromeric repeat sequences are differentially represented in the two libraries. The coordinates of YAC clones hybridizing to these sequences are given. A high proportion of EG YAC clones were classified as containing chimaeric inserts because individual clones carried unique sequences and repetitive sequences originating from different locations in the genome. None of the EW YAC clones analysed were chimaeric in this way. YAC clones carrying tandemly repeated sequences, such as the paracentromeric or rDNA sequences, exhibited a high degree of instability. These observations need to be taken into account when using these libraries in the development of a physical map of the Arabidopsis genome and in chromosome walking experiments.

Arabidopsis↗

Direct cloning of specific genomic DNA sequences in plasmid libraries following fragment enrichment.

We describe a simple method to directly clone any DNA fragment for which a flanking restriction enzyme map is known. Genomic DNA is digested with multiple enzymes cutting outside the fragment to be cloned, selected by electroelution from an agarose gel, and cloned directly into a plasmid vector. It is only necessary to screen 10-1000 colonies and recombinant DNA is ready for immediate molecular analysis without further subcloning. The use of this technique is demonstrated for the cloning of a sequence from within the human alpha-globin complex that was previously shown to be "unclonable" in bacteriophage and cosmid vectors and which is a multiallelic general genetic marker, as well as both beta-globin alleles from an individual with beta-thalassaemia.

Alleles↗

Cloning and functional expression of glycosyltransferases from parasitic protozoans by heterologous complementation in yeast: the dolichol phosphate mannose synthase from Trypanosoma brucei brucei.

The gene for the enzyme dolichol phosphate mannose (Dol-P-Man) synthase from the parasitic protozoan Trypanosoma brucei brucei (T. brucei) was cloned by screening a T. brucei cDNA library and then sequenced. The library was constructed in a yeast expression vector and the positive clone was identified by complementation of a temperature-sensitive defect in the yeast strain DPM 1-6 [Orlean, Albright and Robbins (1988) J. Biol. Chem. 263, 17499-17507]. The insert of this clone displayed an open reading frame of 801 nucleotides coding for a putative protein of 267 amino acids. The deduced protein sequence showed an identity of 49% and a similarity of 69% with the published yeast sequence. Additional features of the T. brucei sequence are the presence of a putative signal sequence, a C-terminal transmembrane domain, a consensus sequence for phosphorylation by cAMP-dependent protein kinase and a stretch of five nucleotides immediately upstream from the putative initiation codon that could function as a prokaryotic ribosome binding site. A consensus sequence for dolichol binding (FI/VXF/YXXIPFXF/Y) found in the yeast protein could not be detected in the putative transmembrane domain of the T. brucei sequence. Biochemical characterization of the recombinant protein showed that it is functionally expressed in the yeast strain DPM 1-6 and Escherichia coli. In both constructs Dol-P-Man synthesis was shown in a cell-free system. Synthesis was stimulated by exogenous dolichol phosphate and inhibited by amphomycin. These results confirm that we have cloned the T. brucei Dol-P-Man synthase by heterologous complementation in yeast, an approach that might be applicable for other glycosyltransferases from various sources.

Amino Acid Sequence↗

Microsatellite markers from sugarcane (Saccharum spp.) ESTs cross transferable to erianthus and sorghum.

Analysis of a sugarcane (Saccharum spp.) EST (expressed sequence tag) library of 8678 sequences revealed approximately 250 microsatellite or simple sequence repeats (SSRs) sequences. A diversity of dinucleotide and trinucleotide SSR repeat motifs were present although most were of the (CGG)(n) trinucleotide motif. Primer sets were designed for 35 sequences and tested on five sugarcane genotypes. Twenty-one primer pairs produced a PCR product and 17 pairs were polymorphic. Primer pairs that produced polymorphisms were mainly located in the coding sequence with only a single pair located within the 5' untranslated region. No primer pairs producing a polymorphic product were found in the 3' untranslated region. The level of polymorphism (PIC value) in cultivars detected by these SSRs was low in sugarcane (0.23). However, a subset of these markers showed a significantly higher level of polymorphism when applied to progenitor and related genera (Erianthus sp. and Sorghum sp.). By contrast, SSRs isolated from sugarcane genomic libraries amplify more readily, show high levels of polymorphism within sugarcane with a higher PIC value (0.72) but do not transfer to related species or genera well.

Journal Article↗

Construction and utility of 10-kb libraries for efficient clone-gap closure for rice genome sequencing.

Rice is an important crop and a model system for monocot genomics, and is a target for whole genome sequencing by the International Rice Genome Sequencing Project (IRGSP). The IRGSP is using a clone by clone approach to sequence rice based on minimum tiles of BAC or PAC clones. For chromosomes 10 and 3 we are using an integrated physical map based on two fingerprinted and end-sequenced BAC libraries to identifying a minimum tiling path of clones. In this study we constructed and tested two rice genomic libraries with an average insert size of 10 kb (10-kb library) to support the gap closure and finishing phases of the rice genome sequencing project. The HaeIII library contains 166,752 clones covering approximately 4.6x rice genome equivalents with an average insert size of 10.5 kb. The Sau3AI library contains 138,960 clones covering 4.2x genome equivalents with an average insert size of 11.6 kb. Both libraries were gridded in duplicate onto 11 high-density filters in a 5 x 5 pattern to facilitate screening by hybridization. The libraries contain an unbiased coverage of the rice genome with less than 5% contamination by clones containing organelle DNA or no insert. An efficient method was developed, consisting of pooled overgo hybridization, the selection of 10-kb gap spanning clones using end sequences, transposon sequencing and utilization of in silico draft sequence, to close relatively small gaps between sequenced BAC clones. Using this method we were able to close a majority of the gaps (up to approximately 50 kb) identified during the finishing phase of chromosome-10 sequencing. This method represents a useful way to close clone gaps and thus to complete the entire rice genome.

Chromosomes, Artificial, Bacterial↗

Comparative sequence analysis of 634 kb of the mouse chromosome 16 region of conserved synteny with the human velocardiofacial syndrome region on chromosome 22q11.2.

Mouse genomic DNA sequence extending 634 kb on proximal mouse chromosome 16 was compared to the corresponding human sequence from chromosome 22q11.2. Haploinsufficiency for this region results in velocardiofacial syndrome (VCFS) in humans. The mouse region is rearranged into three conserved blocks relative to human, but gene content and position are highly conserved within these blocks. Examination of the boundaries of one of these blocks suggested that the evolutionary chromosomal rearrangement occurred in the mouse lineage, resulting in inactivation of the mouse orthologue of ZNF74. Sequence analysis identified 21 genes and 15 ESTs. These include 2 novel genes, Srec2 and Cals2, and previously undescribed splice variants of several other genes. Exon discovery was carried out using GRAIL2, MZEF, or comparative analysis across 491 kb of conserved mouse and human sequence. Sequence comparison was highly effective, identifying every gene and nearly every exon without the high frequency of false-positive predictions seen when algorithmic methods were used alone. In combination, these procedures identified every gene with no false-positive predictions. Comparative sequence analysis also revealed regions of extensive conservation among noncoding sequences, accounting for 6% of the sequence. A library of such sequences has been established to form a resource for generalized studies of regulatory and structural elements.

Abnormalities, Multiple↗

Antigen sequence- and library-based mapping of linear and discontinuous protein-protein-interaction sites by spot synthesis.

The knowledge (antigen-derived peptide scans)- and library (de novo)-based mapping of linear and discontinuous antibody epitopes as well as protein-protein contact sites in general by spot synthesis now is a well established technique. Due to its automation, this technique also promises great potential for applications in functional genomics. It should help to elucidate the complex network of interacting protein molecules involved in signal transduction events (Adam-klages et al. 1996; Hoffmüller et al. 1999). Although little chemistry is involved in the preparation of peptide scans or libraries and the synthesis procedure is relatively simple, the laboratories of immunologists or molecular biologists are often not equipped to perform spot synthesis. In this case scans or libraries can be purchased from commercial suppliers.

Antigen-Antibody Reactions↗