PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequencing libraries”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39Linked to original sources

Construction of cDNA libraries for highly efficient DNA sequencing from the 3' end of expressed genes.

The majority of expressed sequence tag (EST) sequences available today have been derived from the 5' ends of cDNA clones. Obtaining high-quality DNA sequences from the 3' ends of oligo(dT)-primed cDNA on a large scale has been difficult because of slippage of the DNA polymerase enzyme used in direct PCR and cycle sequencing. With the completion of whole genome sequencing for more and more organisms, mRNA 3'-UTR sequences can be particularly useful for clustering large numbers of ESTs for the effective discrimination of individual genes and gene families. We have identified a flaw in the widely used oligo(dT) primers for cDNA synthesis, and here we describe an improved priming approach to effectively synthesize cDNA devoid of homopolymeric nucleotide stretches from mRNA poly(A) tails to enable highly efficient and reliable DNA sequence determination from 3' mRNA ends. Using this method, we produced a rat lung cDNA library and successfully sequenced the 3' ends of 98% of all attempted clones.

Animals↗

Human prostate epithelial cell-type cDNA libraries and prostate expression patterns.

BACKGROUND: Transcriptome analysis is a powerful approach to uncovering genes responsible for diseases such as prostate cancer. Ideally, one would like to compare the transcriptomes of a cancer cell and its normal counterpart for differences. METHODS: Prostate luminal and basal epithelial cell types were isolated and cell-type-specific cDNA libraries were constructed. Sequence analysis of cDNA clones generated 505 luminal cell genes and 560 basal cell genes. These sequences were deposited in a public database for expression analysis. RESULTS: From these sequences, 119 unique luminal expressed sequence tags (ESTs) were extracted and assembled into a luminal-cell transcriptome set, while 154 basal ESTs were extracted and assembled into a basal-cell set. Interlibrary comparison was performed to determine representation of these sequences in cDNA libraries constructed from prostate tumors, PIN, cell lines. CONCLUSIONS: Our analysis showed that a significant number of epithelial cell genes were not represented in the various transcriptomes of prostate tissues, suggesting that they might be underrepresented in libraries generated from tissue containing multiple cell types. Although both luminal and basal cell types are epithelial, their transcriptomes are more divergent from each other than expected, underscoring their functional difference (secretory vs. nonsecretory). Tumor tissues show different expression of luminal and basal genes, with perhaps a trend towards expression of basal genes in advanced diseases.

DNA, Complementary↗

Adaptor-based uracil DNA glycosylase cloning simplifies shotgun library construction for large-scale sequencing.

An improved strategy for the preparation of libraries for the random sequencing of DNA is reported. The protocol is a modification of a previous adaptor-based strategy, and utilizes long (11 base) overhangs, which eliminates the unreliable step of vector-insert ligation. The random inserts are prepared by adaptor ligation, while the M13 vector is prepared as described for uracil DNA glycosylase (UDG) cloning of polymerase chain reaction (PCR) products, using PCR with uracil-containing primers, followed by UDG treatment to produce overhangs. This method has been found to reliably yield large numbers of clones. There is no background due to religation of the vector, and all clones contain inserts. In addition, the method is simple and suitable for export to other investigators. Libraries were constructed from cosmids containing human DNA and from human cDNAs in order to characterize a strategy for shotgun sequencing of multiple shorter fragments.

Base Sequence↗

Analysis of expressed sequence tags from a fetal human heart cDNA library.

Single-pass sequencing of randomly selected cDNA clones to generate expressed sequence tags (ESTs) has been widely used to identify novel genes and to study gene expression in a variety of tissues. We have generated 2244 ESTs from a human fetal heart library (GenBank Accession Nos. R30692-30774 and R56965-58824), which we present in this report. Of these, 51.7% showed no homology to known genes or were similar only to other ESTs, while 48.4% demonstrated homology to known transcripts. A total of 764 ESTs corresponding to known genes were used to study gene expression patterns in the fetal heart and to analyze differences in these patterns from those observed in the adult heart. These analyses demonstrate the utility of ESTs and sequence-tagged clones in comparative studies of gene expression in the cardiovascular system, and they reveal that differential gene expression underlies the structural and functional characteristics of the developing heart.

Adult↗

Alternatively spliced isoforms of the putative renal Na-K-Cl cotransporter are differentially distributed within the rabbit kidney.

We have used cDNA probes derived from the secretory form of the Na-K-Cl cotransporter to screen both cortical and medullary rabbit kidney cDNA libraries. A sequence of 4750 bases was identified from multiple clones. The DNA encodes a protein containing 1099 amino acids, which is 61% identical over its length to the secretory Na-K-Cl cotransporter from shark rectal gland. From analysis of amino acid hydropathy, we predict that this putative renal Na-K-Cl cotransporter has 12 transmembrane helices and large N- and C-terminal cytoplasmic regions. Two sites for N-linked glycosylation are predicted on an extracellular loop. Three potential sites for modulation by protein kinase A are in the C-terminal cytoplasmic domain. Most of the isolated renal cDNA clones were identical over all regions of overlap; however, there was a 96-bp region for which there were three different but homologous variants (A, B, and F). This region of divergence was identified as an alternatively spliced cassette exon since clones were identified that contained intronic DNA as well as consensus splice acceptor sites that bounded the region. Tissue Northern blot analysis revealed a broad band at approximately 5.1 kb that was unique to the kidney. High-stringency Northern blot analysis of cortical and medullary mRNA using antisense oligonucleotides synthesized over each of the three cassette exons revealed that the isoforms were differentially distributed within the kidney--B almost exclusively in cortex, F almost exclusively in medulla, and A about equally distributed.

Alternative Splicing↗

A genome-based resource for molecular cardiovascular medicine: toward a compendium of cardiovascular genes.

BACKGROUND: Large-scale partial sequencing of cDNA libraries to generate expressed sequence tags (ESTs) is an effective means of discovering novel genes and characterizing transcription patterns in different tissues. To catalogue the identities and expression levels of genes in the cardiovascular system, we initiated large-scale sequencing and analysis of human cardiac cDNA libraries. METHODS AND RESULTS: Using automated DNA sequencing, we generated 43,285 ESTs from human heart cDNA libraries. An additional 41,619 ESTs were retrieved from public databases, for a total of 84,904 ESTs representing more than 26 million nucleotides of raw cDNA sequence data from 13 independent cardiovascular system-based cDNA libraries. Of these, 55% matched to known genes in the Genbank/EMBL/DDBJ databases, 33% matched only to other ESTs, and 12% did not match to any known sequences (designated cardiovascular system-based ESTs, or CVbESTs). ESTs that matched to known genes were classified according to function, allowing for detection of differences in general transcription patterns between various tissues and developmental stages of the cardiovascular system. In silico Northern analysis of known gene matches identified widely expressed cardiovascular genes as well as genes putatively exhibiting greater tissue specificity or developmental stage specificity. More detailed analysis identified 48 genes potentially overexpressed in cardiac hypertrophy, at least 10 of which were previously documented as differentially expressed. Computer-based chromosomal localizations of 1048 cardiac ESTs were performed to further assist in the search for disease-related genes. CONCLUSIONS: These data represent the most extensive compilation of cardiovascular gene expression information to date. They further demonstrate the untapped potential of genome research for investigating questions related to cardiovascular biology and represent a first-generation genome-based resource for molecular cardiovascular medicine.

Blotting, Northern↗

Molecular characterization of the first heat shock protein 70 from a reef coral.

The branching coral Stylophora pistillata, one of the most abundant hermatypic corals along the coasts of the Red Sea, has been used for many years as a model species for coral biological studies. Here we characterize the first coral heat shock protein 70 gene (SP-HSP70), cloned from S. pistillata, to be used as a tool for studying coral stress response. The cloning was carried out by a combination of PCR methods using heterologous, degenerate HSP70-based primers, followed by plaque-lift screening of a genomic library. The sequenced clone (5212 bp), contains a complete 1953 bp, intronless open reading frame, and 5' and 3' flanking regions of 1,935 and 1,324 bp, respectively. TATA, CAAT, and ATF boxes as well as 11 putative heat shock elements were identified in the SP-HSP70 5' flanking region. A polyadenylation site was identified in the 3' flanking region. SP-HSP70 protein sequence resembles the cytosolic/nuclear HSP70 cluster. RT-PCR studies confirmed SP-HSP70 mRNA expression in corals grown within their normal physiological conditions. Furthermore, SP-HSP70 has been shown to belong to the coral genome and not to its symbiotic algae one, as revealed by SP-HSP70 PCR amplification, using purified algal and coral DNA templates.

Amino Acid Sequence↗

Identification of bacteria from a non-healing diabetic foot wound by 16 S rDNA sequencing.

Approximately 10-20% of diabetic foot wounds fail initial antibiotic treatment. It is generally believed that several bacterial species may be present in these types of wounds. Because some of these organisms cannot be easily cultured, proper identification is problematic and thus, appropriate treatment modalities cannot be applied. This report examined the bacterial flora present in a chronic diabetic foot wound that failed antibiotic treatment. A tissue sample was collected from the base of the wound and used for standard microbiological culturing. DNA from the sample was used to amplify bacterial 16 S rDNA gene sequences and a library of these sequences was made. The clones were placed into two major groups on the basis of their melting temperatures. Representatives of these groups were sequenced, and information was used to identify the bacteria present in the wound. The culture-based method identified a single anaerobic species, Bacteroides fragilis. The method employing rDNA sequencing identified B. fragilis as a dominant organism and Pseudomonas (Janthinobacterium) mephitica as a minor component. The results indicate that rDNA sequencing approach can be an important tool in the identification of bacteria from wounds.

Aged↗

Site-directed selection of oligonucleotide antagonists by competitive elution.

Oligonucleotide ligands that bind a protein or a small molecule of interest are readily isolated by in vitro selection and amplification of rare sequences from combinatorial libraries of sequence-randomized oligonucleotides (Gold et al., 1995). Classic systematic evolution of ligands by exponential enrichment (SELEX) protocols are affinity based (Tuerk and Gold, 1990), but because many problems and applications require antagonists, protocols for selecting inhibitors are both desirable and valuable. A widely applicable approach for isolating inhibitors is competitive elution with a molecule that binds the targeted molecule's active or binding site. We have used this approach to isolate antagonists of wheat germ agglutinin (WGA) from a library of 2'NH2-pyrimidine, 2'OH-purine oligonucleotides by elution with N N' N"-triacetylchitotriose, (GlcNAc)3. The highest affinity aptamers have equilibrium dissociation constants of 1 nM-20 nM for WGA, a 10(3)-10(4)-fold improvement relative to (GlcNAc)3, and unlike the carbohydrate, are highly specific. In addition to competing for binding with (GlcNAc)3, aptamers inhibit WGA-mediated agglutination of sheep erythrocytes, demonstrating that they are able to compete with natural ligands presented on the surfaces of cells. These results illustrate the feasibility of isolating high-affinity, high-specificity antagonists by competitive elution with low molecular weight, relatively low-affinity, and low-specificity small molecules.

Animals↗

The SBASE protein domain library, release 7.0: a collection of annotated protein sequence segments.

SBASE 7.0 is the seventh release of the SBASE protein domain library sequences that contains 237 937 annotated structural, functional, ligand-binding and topogenic segments of proteins, cross-referenced to all major sequence databases and sequence pattern collections. The entries are clustered into over 1811 groups and are provided with two WWW-based search facilities for on-line use. SBASE 7.0 is freely available by anonymous 'ftp' file transfer from ftp.icgeb. trieste.it. Automated searching of SBASE with BLAST can be carried out with the WWW servers http://www.icgeb.trieste.it/sbase/and http://sbase.abc.hu/sbase/

Amino Acid Sequence↗

Characterization of expressed sequence tags generated from skin cDNA clones of Equus caballus by single pass sequencing.

A cDNA library was built using RNA extracted from the skin tissue of an adult horse. The library was primed with oligo (dT) and sequences were directionally inserted in order to produce an expression library. The library has 5.8X 10(5) plaque forming units with 99.6% recombinant phage. The average insert size is 1.3 Kbp. Three hundred and thirteen expressed sequence tags (ESTs) were generated from sequencing of the 5 prime end of randomly selected skin cDNA clones. The ESTs were sequenced on an ABI 377 using Big-Dye chemistry. A similarity search was performed on each EST using the NCBI non-redundant protein database and 206 ESTs were putatively identified. Twenty six percent of the identified ESTs were redundant. The ESTs were categorized by function. The most frequently identified functional class was translational proteins.

Animals↗

Protocol for Duplex Sequencing of Mitochondrial DNA in Single Human Oocytes.

Oocytes are densely packed with mitochondria, the energy-producing organelles that contain their own genome, mitochondrial DNA (mtDNA). Each cell contains multiple copies of mtDNA, with copy number varying among tissue types. Oocytes possess the highest mtDNA copy number, containing hundreds of thousands of mtDNA molecules per cell. Because mitochondria are inherited exclusively through the maternal lineage, accurate detection of mtDNA variants is essential for studies of inheritance, aging, and disease. The presence of multiple mtDNA copies allows wild-type and mutant molecules to coexist within the same cell, a condition known as heteroplasmy, in which low-frequency and de novo variants may occur at frequencies below 1%. Conventional next-generation sequencing (NGS) lacks sufficient accuracy to reliably distinguish these rare variants from errors introduced during library preparation and sequencing. Here, we present a protocol for enriching mtDNA from single human oocytes using Exonuclease V to remove linear DNA, followed by duplex sequencing library preparation for highly accurate mtDNA analysis. This workflow enables error-corrected sequencing of individual oocytes, facilitating reliable detection of low-frequency mtDNA variants and analysis of heteroplasmy and de novo mutagenesis. The protocol provides a reproducible approach for investigating mitochondrial genome variation in single oocytes using Illumina-compatible sequencing platforms.

Humans↗

Localization of cloned mouse chromosome 7-specific DNA to lethal albino deletions.

Mouse chromosome X- and 7-specific DNA fragments have been isolated from a recombinant DNA library enriched for X(7) chromosome sequences. The library was enriched by flow sorting the X(7) chromosome, a derivation of the Cattanach translocation, prior to library construction. A DNA fragment was found to be located in a region deleted in newborn mice doubly heterozygous for the two albino deletions c3H and c6H in chromosome 7. These chromosome-specific DNA fragments will be useful for studying X inactivation spreading in the X-autosome translocation (T(X;7) 1 Ct) and for investigating the developmental effects of the lethal albino deletions.

Albinism↗

An empirical Bayesian significance test of cDNA library data.

Automated high-throughput sequencing of cDNA clones from numerous libraries has generated a wealth of information about both genome sequence and relative transcript abundances. A common statistical challenge in the analysis of library sequences is to infer whether there is differential expression for the same transcript under two different conditions, such as normal and diseased tissue. In contrast to the continuously variable intensity measurements from microarray experiments, data from cDNA library sequencing presents itself as a discrete count of the incidence of some clone or transcript in a finite sample. In this paper, we first propose a statistical model for data generated from cDNA library sequencing efforts. The model is based on the Poisson mixed with generalized inverse Gaussian (PGIG), introduced by Sichel (1971, 1975). PGIG has been used in modeling population abundance, ecological studies, word frequencies in publications, etc. Using data from the literature, we show that the proposed model provides a good fit to the observed data. Using this new model for cDNA library data, we developed an empirical Bayesian significance test (EBST) for inferring the statistical significance of differential gene expression from discrete data.

Bayes Theorem↗

Expressed sequence tags (ESTs) analysis of Acanthamoeba healyi.

Randomly selected 435 clones from Acanthamoeba healyi cDNA library were sequenced and a total of 387 expressed sequence tags (ESTs) had been generated. Based on the results of BLAST search, 130 clones (34.4%) were identified as the genes encoding surface proteins, enzymes for DNA, energy production or other metabolism, kinases and phosphatases, protease, proteins for signal transduction, structural and cytoskeletal proteins, cell cycle related proteins, transcription factors, transcription and translational machineries, and transporter proteins. Most of the genes (88.5%) are newly identified in the genus Acanthamoeba. Although 15 clones matched the genes of Acanthamoeba located in the public databases, twelve clones were actin gene which was the most frequently expressed gene in this study. These ESTs of Acanthamoeba would give valuable information to study the organism as a model system for biological investigations such as cytoskeleton or cell movement, signal transduction, transcriptional and translational regulations. These results would also provide clues to elucidate factors for pathogenesis in human granulomatous amoebic encephalitis or keratitis by Acanthamoeba.

Acanthamoeba↗

Altered sequence specificity identified from a library of DNA-binding small molecules.

BACKGROUND: The ability to target specific DNA sequences using small molecules has major implications for basic research and medicine. Previous studies revealed that a bis-intercalating molecule containing two 1,4,5,8-napthalenetetracarboxylic diimides separated by a lysine-tris-glycine linker binds to DNA cooperatively, in pairs, with a preference for G + C-rich sequences. Here we investigate the binding properties of a library of bis-intercalating molecules that have partially randomized peptide linkers. RESULTS: A library of bis-intercalating derivatives with varied peptide linkers was screened for sequence specificity using DNase I footprinting on a 231 base pair (bp) restriction fragment. The library mixtures produced footprints that were generally similar to the parent bis-intercalator, which bound within a 15 bp G + C-rich repeat above 125 nM. Nevertheless, subtle differences in cleavage enhancement bands followed by library deconvolution revealed a derivative with novel specificity. A lysine-tris-beta-alanine derivative was found to bind preferentially within a 19 bp palindrome, without substantial loss of affinity. CONCLUSIONS: Synthetically simple changes in the bis-intercalating compounds can produce derivatives with novel sequence specificity. The large size and symmetrical nature of the preferred binding sites suggest that cooperativity may be retained despite modified sequence specificity. Such findings, combined with structural data, could be used to develop versatile DNA ligands of modest molecular weight that target relatively long DNA sequences in a selective manner.

Base Sequence↗

Genomic resources for chicken.

The recent sequencing and draft assembly of a chicken genome has provided biologists with an invaluable research tool that complements a growing list of additional avian genomic resources. For many researchers, finding and using these resources is challenging, because information is presented through an increasing number of Web sites and browser navigation frequently requires specific knowledge and expertise. This primer provides an overview of online genomic resources for the chicken, including the Ensembl, UCSC, and NCBI annotated chicken genome browsers; expressed sequence tag and in situ hybridization databases; and sources for microarrays, cDNAs, and bacterial artificial chromosomes (BACs). Several short tutorials oriented toward the biologist with limited bioinformatics skills outline how to retrieve several types of commonly needed information and reagents.

Animals↗