PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequencing libraries”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 901 records · Page 50Linked to original sources

Environmental whole-genome amplification to access microbial populations in contaminated sediments.

Low-biomass samples from nitrate and heavy metal contaminated soils yield DNA amounts that have limited use for direct, native analysis and screening. Multiple displacement amplification (MDA) using phi29 DNA polymerase was used to amplify whole genomes from environmental, contaminated, subsurface sediments. By first amplifying the genomic DNA (gDNA), biodiversity analysis and gDNA library construction of microbes found in contaminated soils were made possible. The MDA method was validated by analyzing amplified genome coverage from approximately five Escherichia coli cells, resulting in 99.2% genome coverage. The method was further validated by confirming overall representative species coverage and also an amplification bias when amplifying from a mix of eight known bacterial strains. We extracted DNA from samples with extremely low cell densities from a U.S. Department of Energy contaminated site. After amplification, small-subunit rRNA analysis revealed relatively even distribution of species across several major phyla. Clone libraries were constructed from the amplified gDNA, and a small subset of clones was used for shotgun sequencing. BLAST analysis of the library clone sequences showed that 64.9% of the sequences had significant similarities to known proteins, and "clusters of orthologous groups" (COG) analysis revealed that more than half of the sequences from each library contained sequence similarity to known proteins. The libraries can be readily screened for native genes or any target of interest. Whole-genome amplification of metagenomic DNA from very minute microbial sources, while introducing an amplification bias, will allow access to genomic information that was not previously accessible. The reported SSU rRNA sequences and library clone end sequences are listed with their respective GenBank accession numbers, DQ 404590 to DQ 404652, DQ 404654 to DQ 404938, and DX 385314 to DX 389173.

Bacillus Phages↗

The cab-m7 gene: a light-inducible, mesophyll-specific gene of maize.

Southern blot analysis has revealed the existence in maize of perhaps 12 members of the nuclear cab multigene family encoding the chlorophyll a- and b-binding proteins of the photosystem II light-harvesting complex. Hybridization with 3' probes derived from unsequenced cDNA clones showed that six members of this family differ from one another with respect to expression in mesophyll and/or bundle sheath cells and regulation by light. An additional member of this family, designated cab-m7, that encodes a 28 kDa primary translation product has now been identified. It has been cloned from a maize genomic library and sequenced to begin to define the bases for differences in the expression of these genes. This cab gene is shown to be strongly preferentially expressed in the mesophyll (vs. bundle sheath) cells of maize. Furthermore, the gene is photo-responsive; although small amounts of cab-m7 mRNA are present in etiolated leaves, the mRNA pool is 8-fold larger after six hours of illumination. DNA sequences upstream of the cab-m7 gene resemble those found in the 5'-flanking regions of some other plant genes.

Amino Acid Sequence↗

Molecular cloning and expression of a porcine chondrocyte nucleotide pyrophosphohydrolase.

The porcine 127-kDa nucleotide pyrophosphohydrolase (NTPPHase) had been previously purified from the conditioned culture media of porcine articular cartilage. Protein sequencing of an internal 61-kDa proteolytic fragment of NTPPHase (61-kDa NTPPHase) determined the 26 N-terminal amino acids. This sequence was used to amplify a DNA fragment, which was used as a probe to clone the gene encoding the 61-kDa NTPPHase from a porcine chondrocyte cDNA library. DNA sequence analysis showed the cDNA insert to be 2509 bp, corresponding to a predicted open reading frame (ORF) encoding 599 amino acids. The 26 N-terminal amino acids of the 61-kDa NTPPHase were located within the ORF immediately downstream of a putative protease recognition region, RRKRR. This is consistent with this cDNA insert representing an internal proteolytic fragment of the full length 127-kDa NTPPHase. BLAST and FASTA analysis confirmed that the deduced amino acid sequence of 61-kDa NTPPHase was unique and did not possess a high degree of homology to sequence in the non-redundant protein and nucleotide databases. Proteins that possess limited homology (< 17%) with the 61-kDa NTTPPHase include several prokaryotic and eukaryotic ATP pyrophosphate-lyases (adenylate cyclase). Northern blot analysis of porcine chondrocyte RNA showed that the DNA encoding the 61-kDa NTPPHase hybridized to a single 4.0-kb RNA transcript. This DNA probe also hybridized to a single species of human chondrocyte RNA. Expression of a 61-kDa protein was detected by coupled in-vitro transcription/translation. Western blot analysis of this in-vitro transcription/translation reaction detected a 61-kDa protein, using an antibody raised against the peptide sequence that was originally used to clone the 61-kDa NTPPHase. These data indicate the successful in-vitro cloning and expression of the porcine chondrocyte 61-kDa NTPPHase. Future studies that utilize the gene encoding the 61-kDa NTPPHase may allow the characterization of the role of NTPPHase in calcium pyrophosphate dihydrate (CPPD) crystal deposition disease.

Amino Acid Sequence↗

Oligonucleotide fingerprinting with (CAC)5: nonradioactive in-gel hybridization and isolation of individual hypervariable loci.

The first topic to be treated in this paper is the nonradioactive DNA fingerprinting by means of in-gel hybridization with digoxigenated (CAC)5. Besides the fact that time-consuming Southern blotting can be avoided, the dried agarose is an excellent matrix to produce background-free nonradioactive DNA fingerprints. There is no tendency of either the oligonucleotide probe or the antibody towards unspecific binding to the dried agarose. Prehybridization and blocking steps are therefore superfluous. Furthermore, we will discuss what effect the degree of crosslinking of the antibody-enzyme conjugates has. The second topic concerns the isolation and characterization of locus-specific probes from a human (CAC)5 fingerprint. The isolation and characterization of one variable probe, by screening complete genomic libraries, is described and discussed. This probe is compared to a hypervariable single-copy probe, isolated from a size-enriched genomic library. The sequence of the repeat flanking locus-specific probe is presented and a semi-specific, adaptor-mediated polymerase chain reaction was designed to amplify (CAC)n/(GTG)n flanking sequences.

Base Sequence↗

Temperature-sensitive variants of Saccharomyces cerevisiae iso-1-cytochrome c produced by random mutagenesis of codons 43 to 54.

In vitro random mutagenesis within the CYC1 gene from the yeast Saccharomyces cerevisiae was used to produce a library of mutants encompassing codons 43 to 54 of iso-1-cytochrome c. This region consists of an evolutionarily conserved structure within an evolutionarily diverse sequence. The library, on a low-copy-number yeast shuttle phagemid, was introduced into a yeast strain lacking cytochrome c. The ability of transformants harboring a functional cytochrome c to grow on the non-fermentable carbon source glycerol at 30 degrees C and 37 degrees C was used to determine the phenotype of nearly 1000 transformants. Approximately 90% of the missense mutants present in the library give rise to the wild-type phenotype, 7% result in the temperature-sensitive (Cycts) phenotype, and 3% give rise to the non-functional (Cyc-) phenotype. Phagemids from 20 Cycts and 30 Cyc- transformants were subjected to DNA sequence analysis. All the mutations occur within the targeted region. One-third of the mutants from Cyc- transformants and all the mutants from Cycts transformants are missense mutants. The remaining mutants from Cyc- transformants are nonsense or frame-shift mutants. Missense mutations within the codons for Gly45, Tyr46, Thr49, Asn52 or Ile53 alone are sufficient to produce temperature-sensitive behavior both in vivo and in the variant proteins. The deduced amino acid substitutions correlate remarkably well with side-chain dynamics, secondary structure and tertiary structure of the wild-type protein.

Base Sequence↗

A computer program for the estimation of protein and nucleic acid sequence diversity in random point mutagenesis libraries.

A computer program for the generation and analysis of in silico random point mutagenesis libraries is described. The program operates by mutagenizing an input nucleic acid sequence according to mutation parameters specified by the user for each sequence position and type of point mutation. The program can mimic almost any type of random mutagenesis library, including those produced via error-prone PCR (ep-PCR), mutator Escherichia coli strains, chemical mutagenesis, and doped or random oligonucleotide synthesis. The program analyzes the generated nucleic acid sequences and/or the associated protein library to produce several estimates of library diversity (number of unique sequences, point mutations, and single point mutants) and the rate of saturation of these diversities during experimental screening or selection of clones. This information allows one to select the optimal screen size for a given mutagenesis library, necessary to efficiently obtain a certain coverage of the sequence-space. The program also reports the abundance of each specific protein mutation at each sequence position, which is useful as a measure of the level and type of mutation bias in the library. Alternatively, one can use the program to evaluate the relative merits of preexisting libraries, or to examine various hypothetical mutation schemes to determine the optimal method for creating a library that serves the screen/selection of interest. Simulated libraries of at least 10(9) sequences are accessible by the numerical algorithm with currently available personal computers; an analytical algorithm is also available which can rapidly calculate a subset of the numerical statistics in libraries of arbitrarily large size. A multi-type double-strand stochastic model of ep-PCR is developed in an appendix to demonstrate the applicability of the algorithm to amplifying mutagenesis procedures. Estimators of DNA polymerase mutation-type-specific error rates are derived using the model. Analyses of an alpha-synuclein ep-PCR library and NNS synthetic oligonucleotide libraries are given as examples.

Algorithms↗

Construction of a swine YAC library allowing an efficient recovery of unique and centromeric repeated sequences.

A swine DNA genomic library was constructed in yeast artificial chromosome (YAC) using the pYAC4 vector and the AB1380 strain. The DNA prepared from two Large White males was partially digested with EcoRI and size selected after both digestion and ligation. The YAC library contained 33792 arrayed clones with an average size of 280 kb as estimated by analysis of 2% of the clones, thus representing a threefold coverage of the swine haploid genome. The library was organized in pools to facilitate the PCR screening. The complexity of the library was tested both for unique and centromeric repeated sequences. In all, 20 out of 22 primer sets allowed the characterization of one to six clones containing specific unique sequences. These sequences are known to be on Chromosomes (Chrs) 1, 2, 5, 6, 7, 8, 13, 14, 15, 17, and X. Eight additional clones carrying centromeric repeat units were also isolated with a single primer set. The sequencing of 37 distinct repeat units of about 340 bp subcloned from these eight YACs revealed high sequence diversity indicating the existence of numerous centromeric repeat unit subfamilies in swine. Furthermore, the analysis of the restriction patterns with selected enzymes suggested a higher order organization of the repeat units. According to preliminary FISH experiments on a small number of randomly chosen YACs and YACs carrying specific sequences, the chimerism appeared to be low. In addition, primed in situ labeling experiments favored the idea that the YACs with centromeric repeat sequences were derived from a subset of metacentric and submetacentric chromosomes.

Animals↗

Cloning and sequencing of a human cDNA encoding a putative transcription factor containing a bromodomain.

A 2985 bp cDNA was isolated from a Lambda Zap Express library and sequenced. The cDNA appeared to represent a previously unknown gene that encodes and acidic 757 amino acid protein containing a bromodomain, several potential sites for phosphorylation by casein kinase-II and small proline-rich segments. The results suggest that the encoded protein might be a novel transcription factor.

Amino Acid Sequence↗

Three rat preprotachykinin mRNAs encode the neuropeptides substance P and neurokinin A.

Synthetic oligonucleotides were used to screen a rat striatal cDNA library for sequences corresponding to the tachykinin peptides substance P and neurokinin A. The cDNA library was constructed from RNA isolated from the rostral portion of the rat corpus striatum, the site of striatonigral cell bodies. Two types of cDNAs were isolated and defined by restriction enzyme analysis and DNA sequencing to encode both substance P and neurokinin A. The two predicted preprotachykinin protein precursors (130 and 115 amino acids in length) differ from each other by a pentadecapeptide sequence between the two tachykinin sequences, and both precursors possess appropriate processing signals for substance P and neurokinin A production. The presence of a third preprotachykinin mRNA of minor abundance in rat striatum was established by S1 nuclease protection experiments. This mRNA encodes a preprotachykinin of 112 amino acids containing substance P but not neurokinin A. These three mRNAs are derived from one rat gene as a result of differential RNA processing; thus, this RNA processing pattern further increases the diversity of products that can be generated from the preprotachykinin gene.

Amino Acid Sequence↗

Cloning and characterization of a cDNA encoding a precursor for human uroguanylin.

Uroguanylin, a member of the guanylin peptide family, is an endogenous activator of intestinal guanylate cyclase (GC-C). A cDNA encoding a precursor for human uroguanylin was cloned from a human colon cDNA library and sequenced. The precursor was 112 amino acids long and included a signal peptide at the N-terminus and the human uroguanylin sequence at the C-terminus. RNA blot analysis and the reverse transcription-polymerase chain reaction (RT-PCR) showed that human uroguanylin mRNA is expressed in the stomach and intestine. Uroguanylin, as well as guanylin, may be a potent physiological regulator of intestinal fluid and electrolyte transport.

Amino Acid Sequence↗

Expression of GTP-binding protein alpha subunits in human thymocytes.

In this report, we investigate G protein alpha subunit diversity in human thymocytes, utilizing common properties shared by these genes and reverse transcription-polymerase chain reaction (RT-PCR). Sequence analysis of PCR amplified gene portions, indicate the presence of members from all four G-protein families that have been described thus far. The alpha subunit genes identified are: G alpha i1-3 and G alpha z but not G alpha o from the Gi family, G alpha s from the Gs family, G alpha 11, G alpha q, and G alpha 16 from the Gq family, and G alpha 12 and G alpha 13 from the G12 family. Also in this report we present the nucleotide and predicted amino acid sequences of the human G alpha 13 cloned from a thymocyte cDNA library. The sequence of the human G alpha 13 has not been previously reported. Comparison of this sequence with the reported murine G alpha 13 shows > 90% identity at the deduced amino acid sequence level. We conclude that thymocytes represent a useful experimental system for the study of G protein involvement in immune responses and lymphocyte development.

Amino Acid Sequence↗

The primary sequence and the subunit structure of mouse alpha-2-macroglobulin, deduced from protein sequencing of the isolated subunits and from molecular cloning of the cDNA.

Mouse plasma alpha-2-macroglobulin (m alpha 2M) was isolated and the N-terminal amino-acid sequences determined after separation of the 165-kDa and 35-kDa subunits. These sequences were compared to the protein sequence predicted by the cDNA, which was cloned from a mouse liver library and sequenced. From these data it is evident that both subunits are encoded by one mRNA of approximately 5 kb expressed predominantly in liver. The smaller subunit, with the N-terminal sequence DLSSSDLT, comprises the C-terminal 257 residues of m alpha 2M and is derived from a single-chain precursor probably by proteolytic processing at an arginine residue in the sequence PTRDLSS. Analysis of the predicted protein further showed all the salient features of a proteinase inhibitor of the macroglobulin family: a bait region that deviates from all known sequences in this family, a very conserved internal thiolester site and conserved cysteine residues and putative N-glycosylation sites. The synthesis of m alpha 2M in adult liver was demonstrated by Northern blotting and in fetal liver by in-situ hybridization. Transient transfection of COS cells with the cDNA under control of a viral promoter demonstrated the secretion and partial processing of m alpha 2M in the culture medium. In plasma the level of m alpha 2M was found to be stable as expected for the murine counterpart of human plasma alpha-2-macroglobulin. The possibilities of using the mouse as a genetic model to study this proteinase inhibitor in vivo are discussed.

Amino Acid Sequence↗

[Serial analysis of gene expression in the pituitary adenomas and para-tumor normal pituitary tissues].

OBJECTIVE: To observe the characteristics and difference of gene expression in the pituitary adenomas and para-tumor normal pituitary tissues. METHODS: Using serial analysis of gene expression (SAGE), two SAGE libraries were generated. Forty clones from each SAGE library were sequenced, and the results were analyzed by SAGE2000 software and compared with the SAGE map at NCBI. RESULTS: A total of 655 gene tags, representing 43 genes, were extracted from the 40 sequence files of the para-tumor normal pituitary tissues and 737 gene tags, representing 53 genes, were extracted from the 40 sequence files of the pituitary adenomas. Of these tags, 13 were not reported before. The genes related to pituitary hormone secretion and energy metabolism were highly expressed in the two kinds of tissues. Some growth factors and cytokines were also expressed, including those involved in the immunological system. But there were also much difference of gene expression in the two tissues. Thirty-one and five tags were only detected in para-tumor normal pituitary tissues and pituitary adenomas, respectively. CONCLUSIONS: Genes involved in hormones secretion and energy metabolism were highly expressed in the pituitary adenomas and para-tumor normal pituitary tissues. Many growth factors and cytokines were also expressed in pituitary. There was also much difference of gene expression in the two kinds of tissues. SAGE can be used not only in understanding the quantity information of gene expression, but also in finding new genes.

Adenoma↗

The iron-responsive element binding protein. Purification, cloning, and regulation in rat liver.

The iron-responsive element binding protein (IRE-BP) is a cytosolic protein that binds a highly conserved sequence in the untranslated regions of mRNAs involved in iron metabolism including ferritin, transferrin receptor, and erythroid 5-aminolevulinate acid synthase. This conserved sequence is termed the iron-responsive element and is necessary for the post-transcriptional regulation of these mRNAs by iron. The rat liver IRE-BP was purified to homogeneity by chromatographic methods and partial amino acid sequence was obtained. A cDNA was isolated from a rat liver cDNA library and sequenced. The amino acid sequence deduced from the cDNA sequence corresponds to a protein of 889 amino acids with a predicted molecular weight of 97.946. The NH2-terminal sequence obtained by Edman degradation matched the deduced amino acid sequence obtained from the cDNA, confirming the translational start site. Rat liver IRE-BP shares 95% identity with human IRE-BP and 98% identity with mouse IRE-BP indicating that the IRE-BPs have remained highly conserved during evolution. The 5'-untranslated region is at least 236 nucleotides and contains interesting structural features including two direct repeats, an inverted repeat, and three small open reading frames. The rat IRE-BP mRNA is approximately 3600 nucleotides and is expressed in a variety of rat tissues including liver, spleen, and gut. Over the course of 16 h following an intraperitoneal injection of iron in rats. IRE-BP RNA binding activity decreases to 50% of control levels. The decrease in IRE-BP RNA binding activity in extracts from iron-treated rats is reversible by pretreatment of the extracts with reducing agents. The steady-state levels of IRE-BP mRNA remain constant during iron treatment. These data suggest that the decrease in IRE-BP RNA binding activity by iron in rat liver is due to post-translational changes in the RNA binding affinity of the IRE-BP and not due a decrease in the transcription of the IRE-BP gene or to the destabilization of the IRE-BP mRNA.

Amino Acid Sequence↗

A genetic screen to identify sequences that mediate protein oligomerization in Escherichia coli.

Many proteins assemble as oligomeric complexes and in several cases a distinct domain mediates the interaction between the subunits. The identification of new oligomerization modules is relevant to comprehend both the architecture and the evolution of protein sequences and also for protein engineering applications. Using the bacteriophage lambda repressor dimerization assay, we searched Escherichia coli genomic libraries for sequences able to mediate protein oligomerization in vivo. We identified short peptides that can substitute very effectively the dimerizing domain of the repressor. Most of these peptides belong to open reading frames that are normally not expressed in the bacterial cell.

Amino Acid Sequence↗

Cloning and characterization of cDNA encoding a human arginyl-tRNA synthetase.

Arginyl-tRNA synthetase (ArgRS) plays a key role in protein synthesis as part of a multienzyme complex with a number of other aminoacyl-tRNA synthetase (aaRS) enzymes. We have isolated a full-length cDNA encoding ArgRS as part of a project on complementation of radiosensitivity in human cells with an Epstein-Barr Virus (EBV) vector-based human cDNA library. DNA sequence analysis identified an open reading frame of 1983 nucleotides with 87% homology to other mammalian ArgRS genes. The deduced amino acid (aa) sequence (661 aa) showed 87.7% identity to the Chinese hamster ovary (CHO) enzyme and 37.7% identity to the homologous Escherichia coli enzyme. Northern blot analysis revealed the presence of a single mRNA species of approx. 2.2 kb. The results described here demonstrate that ArgRS is highly conserved in mammalian cells and confirm the presence of a hydrophobic N-terminal region in the higher-molecular-weight complexed form of ArgRS.

Amino Acid Sequence↗

Mapping of RNA accessible sites by extension of random oligonucleotide libraries with reverse transcriptase.

A rapid and simple method for determining accessible sites in RNA that is independent of the length of target RNA and does not require RNA labeling is described. In this method, target RNA is allowed to hybridize with sequence-randomized libraries of DNA oligonucleotides linked to a common tag sequence at their 5'-end. Annealed oligonucleotides are extended with reverse transcriptase and the extended products are then amplified by using PCR with a primer corresponding to the tag sequence and a second primer specific to the target RNA sequence. We used the combination of both the lengths of the RT-PCR products and the location of the binding site of the RNA-specific primer to determine which regions of the RNA molecules were RNA extendible sites, that is, sites available for oligonucleotide binding and extension. We then employed this reverse transcription with the random oligonucleotide libraries (RT-ROL) method to determine the accessible sites on four mRNA targets, human activated ras (ha-ras), human intercellular adhesion molecule-1 (ICAM-1), rabbit beta-globin, and human interferon-gamma (IFN-gamma). Our results were concordant with those of other researchers who had used RNase H cleavage or hybridization with arrays of oligonucleotides to identify accessible sites on some of these targets. Further, we found good correlation between sites when we compared the location of extendible sites identified by RT-ROL with hybridization sites of effective antisense oligonucleotides on ICAM-1 mRNA in antisense inhibition studies. Finally, we discuss the relationship between RNA extendible sites and RNA accessibility.

Animals↗

Nucleotide sequence of rat calretinin cDNA.

A rat calretinin cDNA clone was selected by antibody screening of a lambda gt11 brain library. The sequence revealed remarkable nucleotide and amino acid homology with human calretinin (91.1% and 98.5%, respectively), with only four amino acid differences. A high degree of homology with chick calretinin was also observed (79.8% and 86.6%, respectively), with 36 amino acid differences. Although the role of this central nervous system protein has not been well characterized, the evolutionarily conserved calcium binding domains and connecting regions, in addition to the limited local changes observed between rat and chick primary structure, lead us to believe that calretinin interacts with other highly conserved constituents of brain cells. This calretinin cDNA clone provides a new probe for the analysis of a specific subset of neurons in the central nervous system. The probe will allow a more detailed analysis of calretinin regulation in the brain and will be useful for screening genomic libraries for the complete chromosomal gene. (GenBank accession No X66974).

Amino Acid Sequence↗