PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequencing libraries”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Use of a Mycobacterium tuberculosis H37Rv bacterial artificial chromosome library for genome mapping, sequencing, and comparative genomics.

The bacterial artificial chromosome (BAC) cloning system is capable of stably propagating large, complex DNA inserts in Escherichia coli. As part of the Mycobacterium tuberculosis H37Rv genome sequencing project, a BAC library was constructed in the pBeloBAC11 vector and used for genome mapping, confirmation of sequence assembly, and sequencing. The library contains about 5,000 BAC clones, with inserts ranging in size from 25 to 104 kb, representing theoretically a 70-fold coverage of the M. tuberculosis genome (4.4 Mb). A total of 840 sequences from the T7 and SP6 termini of 420 BACs were determined and compared to those of a partial genomic database. These sequences showed excellent correlation between the estimated sizes and positions of the BAC clones and the sizes and positions of previously sequenced cosmids and the resulting contigs. Many BAC clones represent linking clones between sequenced cosmids, allowing full coverage of the H37Rv chromosome, and they are now being shotgun sequenced in the framework of the H37Rv sequencing project. Also, no chimeric, deleted, or rearranged BAC clones were detected, which was of major importance for the correct mapping and assembly of the H37Rv sequence. The minimal overlapping set contains 68 unique BAC clones and spans the whole H37Rv chromosome with the exception of a single gap of approximately 150 kb. As a postgenomic application, the canonical BAC set was used in a comparative study to reveal chromosomal polymorphisms between M. tuberculosis, M. bovis, and M. bovis BCG Pasteur, and a novel 12.7-kb segment present in M. tuberculosis but absent from M. bovis and M. bovis BCG was characterized. This region contains a set of genes whose products show low similarity to proteins involved in polysaccharide biosynthesis. The H37Rv BAC library therefore provides us with a powerful tool both for the generation and confirmation of sequence data as well as for comparative genomics and other postgenomic applications. It represents a major resource for present and future M. tuberculosis research projects.

Chromosome Mapping↗

Molecular cloning sequence and distribution of rat calspermin, a high affinity calmodulin-binding protein.

Calspermin is a heat-stable, acidic calmodulin-binding protein predominantly found in mammalian testis. The cDNA representing the rat form of this protein has been cloned from a rat testis lambda gt11 library. Sequence analysis of two overlapping clones revealed a 232-nucleotide 5'-nontranslated region, 510 nucleotides of open reading frame, a 148-nucleotide 3'-untranslated region, and a poly(A) tail. Authenticity of the clones was confirmed by comparison of a portion of the deduced amino acid sequence with the sequence of a tryptic peptide obtained from the rat testis protein. The lambda gt11 fusion protein was recognized by affinity purified antibodies to pig testis calspermin and bound 125I-calmodulin in a Ca2+-dependent manner. Calspermin cDNA encodes a 169-residue protein with a calculated Mr of 18,735. The putative calmodulin-binding domain is very close to the amino terminus of the protein. This region shows 46% identity with the calmodulin-binding region of rat brain Ca2+/calmodulin-dependent protein kinase II and 32% identity with the equivalent region of chicken smooth muscle myosin light chain kinase. The 5'-nontranslated region reveals significant homology with a portion of the catalytic region of the calmodulin-dependent protein kinase family. Calspermin contains a stretch of 17 contiguous glutamic acid residues in the central region of the molecule. Computer analysis predicts calspermin to be 81% alpha-helix and 14% random coil. Analysis of genomic DNA indicates calspermin to be the product of a unique gene. Northern blot analysis of rat testis RNA reveals a 1.1-kilobase mRNA. This RNA is restricted to testis among several rat tissues examined and could not be identified in total RNA isolated from testes of other mammals. Analysis of cells isolated from rat testis reveals calspermin mRNA to be predominantly expressed in postmeiotic cells indicating that it may be specific to haploid cells.

Amino Acid Sequence↗

Isolation and sequence of an interferon-tau-inducible, pregnancy- and bovine interferon-stimulated gene product 15 (ISG15)-specific, bovine ubiquitin-activating E1-like (UBE1L) enzyme.

Bovine (bov) interferon-stimulated gene product 15 (ISG15) is produced in the endometrium in response to conceptus-secreted interferon (IFN)-tau. ISG15 conjugates to endometrial proteins through an enzymatic pathway that is similar to ubiquitinylation. Ubiquitin-activating enzyme 1-like protein (UBE1L) initiates enzymatic conjugation by forming a thioester bond with ISG15, thus preparing it for transfer to the next series of enzymes. The bovUBE1L has not been described. We hypothesized that bovUBE1L was induced by pregnancy and IFN-tau in the endometrium. A 110-kDa protein was purified from bovine endometrial (BEND) cells based on affinity with recombinant (r) glutathione S-transferase (GST)-ISG15. This protein was digested in gel with trypsin. Seven peptides were purified using HPLC, sequenced using liquid chromatography-mass spectroscopy-mass spectroscopy and found to share 43-100% identity with human UBE1L. The full-length bovUBE1L cDNA was isolated from a BEND cell cDNA library, sequenced, and found to share 83% identity with human UBE1L cDNA. Northern blot revealed two mRNAs that were detected in greater (P<0.05) concentrations in endometrium from Day 17-21 pregnant versus nonpregnant cows. Western blots using antihuman UBE1L antibody revealed a similar pattern of pregnancy-associated expression of UBE1L protein in the uterus. The bovUBE1L mRNA was localized, using in situ hybridization, primarily to glandular and luminal epithelium, with more diffuse localization to stroma of the endometrium from pregnant cows. Because bovUBE1L was purified through its interaction with rGST-ISG15 and shares significant amino acid and cDNA sequence identity with human UBE1L, it is concluded that it mediates conjugation of ISG15 to uterine proteins in response to the developing and attaching conceptus.

Amino Acid Sequence↗

Characterization of Fsc1 cDNA for a mouse sperm fibrous sheath component.

The fibrous sheath is a major cytoskeletal structure in the principal piece of the mammalian sperm flagellum. We have cloned a cDNA and used it to characterize the expression of mRNA for a mouse sperm fibrous sheath protein. Peptides from a tryptic digest of fibrous sheath proteins were separated by HPLC and a 31 amino acid sequence was obtained from one of the peptides. Through the use of degenerate oligonucleotide polymerase chain reaction (PCR) primers predicted from this sequence, an 80-bp product was amplified from mouse testis first-strand cDNA. This was utilized as a probe to isolate a 2.9-kb cDNA clone from a mouse round spermatid cDNA library. Sequence analysis of the cDNA clone showed that it encodes a protein with an open reading frame of 849 amino acids and includes the original peptide sequence. The predicted protein has a molecular weight of 93,795 and contains 32 cysteine residues and 32 potential phosphorylation sites. It has no significant homology with other known cytoskeletal proteins. Northern blot analysis detected an mRNA of approximately 3 kb that was abundant in round spermatids of the mouse and in testes from six other mammalian species, but not in twelve somatic tissues from the mouse. In situ hybridization analysis indicated that the mRNA is first detected in step 1-6 spermatids, is most abundant in step 8-12 spermatids, and decreases in amount in step 13-15 spermatids, suggesting that expression of the mRNA occurs in the postmeiotic phase of spermatogenesis.(ABSTRACT TRUNCATED AT 250 WORDS)

Amino Acid Sequence↗

Structural and functional characterization of the USP11 deubiquitinating enzyme, which interacts with the RanGTP-associated protein RanBPM.

RanBPM is a RanGTP-binding protein required for correct nucleation of microtubules. To characterize the mechanism, we searched for RanBPM-binding proteins by using a yeast two-hybrid method and isolated a cDNA encoding the ubiquitin-specific protease USP11. The full-length cDNA of USP11 was cloned from a Jurkat cell library. Sequencing revealed that USP11 possesses Cys box, His box, Asp and KRF domains, which are highly conserved in many ubiquitin-specific proteases. By immunoblotting using HeLa cells, we concluded that 921-residue version of USP11 was the predominant form, and USP11 may be a ubiquitous protein in various human tissues. By immunofluorescence assay, USP11 primarily was localized in the nucleus of non-dividing cells, suggesting an association between USP11 and RanBPM in the nucleus. Furthermore, the association between USP11 and RanBPM in vivo was confirmed not only by yeast two-hybrid assay but also by co-immunoprecipitation assays using exogenously expressed USP11 and RanBPM. We next revealed proteasome-dependent degradation of RanBPM by pulse-chase analysis using proteasome inhibitors. In fact, ubiquitinated RanBPM was detected by both in vivo and in vitro ubiquitination assays. Finally, ubiquitin conjugation to RanBPM was inhibited in a dose-dependent manner by the addition of recombinant USP11. We conclude that RanBPM was the enzymic substrate for USP11 and was deubiquitinated specifically.

Adaptor Proteins, Signal Transducing↗

Purification, cDNA cloning and northern blot analysis of trehalase of pupal midgut of the silkworm, Bombyx mori.

Trehalase (alpha-glucoside-1-glucohydrolase, EC 3.2.1.28) was purified from silkworm pupal midgut to homogeneity by DEAE-Sepharose CL-6B and hydroxyapatite chromatography, and native gel electrophoresis. The enzyme had a molecular mass of 70 kDa. The N-terminal amino-acid sequence of the intact trehalase and its three fragments by V8 proteinase digestion was determined. Based on the amino-acid sequence, degenerate oligonucleotides were synthesized and used as primers in a polymerase chain reaction (PCR). Using a 0.8 kb PCR product as a hybridization probe, trehalase clones were isolated from the pupal midgut cDNA library. Sequence analysis revealed that the isolated trehalase cDNA contains 3103 nucleotides and comprises 579 amino acids, including a cleavable signal sequence and five potential N-glycosylation sites. Northern blot analysis clearly showed a 3.0 kb transcript in midgut, and Malpighian tubule, but not in fat body, silk gland, ovary, trachea, brain and suboesophageal ganglion.

Amino Acid Sequence↗

Volumetric DNA microscopy for mapping spatial transcriptomes in three dimensions.

The architecture and function of biological systems are inherently three-dimensional, yet most existing spatial transcriptomic technologies remain restricted to thin tissue sections, limiting their capacity to resolve cellular organization and microenvironments within intact tissue volumes. To address this limitation, we developed volumetric DNA microscopy, a scalable, optics-free approach for spatial transcriptome profiling directly within intact biological specimens. The method encodes spatial information into DNA molecules that form a dense intermolecular network in situ, enabling the reconstruction of three-dimensional spatial relationships through short-read sequencing and computational analysis. Here we detail the complete workflow including in situ cDNA synthesis, spatial encoding through DNA nanoball formation, dual-scale proximity bridging between neighboring nanoballs and spatial reconstruction via geodesic spectral embedding. Sequencing libraries can be generated within 7-8 d by a competent graduate-level molecular biologist, followed by standardized downstream computational analysis. Because the workflow requires only routine molecular biology reagents and a benchtop sequencer, volumetric DNA microscopy provides a versatile platform for exploring genetic and morphological features in intact tissues.

Spatial Transcriptomics↗

Alternative splicing of human glucose-6-phosphate dehydrogenase messenger RNA in different tissues.

Different forms of glucose-6-phosphate dehydrogenase (G-6-PD) have been described in different tissues. Moreover, the directly determined amino acid sequence amino end of the red cell enzyme does not exactly match the sequence deduced from cDNA isolated from HeLa cells or lymphoblasts. We have therefore investigated the sequence of cDNA from sperm, granulocytes, reticulocytes, brain, placenta, liver, lymphoblastoid cells, and cultured fibroblasts. A novel human cDNA, which has extra 138 bases coding 46 amino acids, was isolated from a lymphoblastoid cell library. Sequencing of genomic DNA amplified by the polymerase chain reaction (PCR) revealed that the extra sequence was derived from the 3'-end of intron 7 by alternative splicing. This longer form of mRNA was also detected in sperm and granulocytes. Sequence analysis using PCR-amplified cDNA revealed that the 5'-end of the coding sequence of G6PD mRNA in reticulocytes is identical to those in other tissues.

Amino Acid Sequence↗

Expressed sequence tags for the chicken genome from a normalized 10-day-old White Leghorn whole embryo cDNA library: 1. DNA sequence characterization and linkage analysis.

Expressed sequence tags (ESTs) provide a rapid and reliable method for gene discovery as well as a resource for the large-scale analysis of gene expression of known and unknown genes. Here we describe a normalized cDNA library developed from a 10-day-old White Leghorn chicken whole embryo. The utility of the library was evaluated by partial sequencing of 99 randomly selected insert-containing clones and the analysis of EST-targeted genomic regions for single nucleotide polymorphisms (SNPs) in the East Lansing chicken reference DNA mapping panel. Using stringent match criteria of percent identity of 80 or higher across a length of 50 or more bases, 46 ESTs matched database sequences including previously reported Gallus gallus genes. Thirty-seven of the 50 primer pairs developed from 50 unique ESTs amplified a single fragment. The size of the 37 amplicons ranged from 276 to 693 bp for a total of 17,508 and an average of 473. About 70% of the SNPs detected were either G-->A or C-->T transition. The number of SNPs detected within the amplicons from EST-targeted genomic regions ranged from 0 to 4 for a total of 65 and a frequency of about 1 every 470 bases. About 35% of the amplicons contained only 1 SNP, while 19% had 4 SNPs. Using the SNPs that were informative in the East Lansing reference panel, 17 ESTs were mapped on the East Lansing chicken genetic map. The ESTs described, as well as the nucleotide variants identified within the EST-targeted genomic regions, represent significant resources for genome analysis in the chicken.

Animals↗

Dual promoter structure of mouse and human fatty acid translocase/CD36 genes and unique transcriptional activation by peroxisome proliferator-activated receptor alpha and gamma ligands.

Fatty acid translocase (FAT)/CD36 is a glycoprotein involved in multiple membrane functions including uptake of long-chain fatty acids and oxidized low density lipoprotein. In mice, expression of the gene is regulated by peroxisome proliferator-activated receptor (PPAR) alpha in the liver and by PPAR gamma in the adipose tissues (Motojima, K., Passilly, P. P., Peters, J. M., Gonzalez, F. J., and Latruffe, N. (1998) J. Biol. Chem. 273, 16710-16714). However, the time course of PPAR alpha ligand-induced expression of FAT/CD36 in the liver, and also in the cultured hepatoma cells, is significantly slower than those of other PPAR alpha target genes. To study the molecular mechanism of the slow transcriptional activation of the gene by a PPAR ligand, we first cloned the 5' ends of the mRNA and then the mouse gene promoter region from a genomic bacterial artificial chromosome library. Sequencing analyses showed that transcription of the gene starts at two initiation sites 16 kb apart and splicing occurs alternatively, producing at least three mRNA species with different 5'-noncoding regions. The PPAR alpha ligand-responsive promoter in the liver was identified as the new upstream promoter where we found several possible binding sites for lipid metabolism-related transcriptional factors but not for PPAR. Neither promoter responded to a PPAR alpha ligand in the in vitro or in vivo reporter assays using cultured hepatoma cells and the liver of living mice. We also have cloned the human FAT/CD36 gene from a bacterial artificial chromosome library and identified a new independent promoter that is located 13 kb upstream of the previously reported promoter. Only the upstream promoter responded to PPAR alpha and PPAR gamma ligands in a cell type-specific manner. The absence of PPRE in the responding upstream promoter region, the delayed activation by the ligand, and the results of the reporter assays all suggested that transcriptional activation of the FAT/CD36 gene by PPAR ligands is indirectly dependent on PPAR.

5' Untranslated Regions↗

Rice bicoid-related cDNA sequence and its expression during early embryogenesis.

Bicoid is one of the important Drosophila maternal genes involved in the control of embryo polarity and larvae segmentation. To clone and characterize the rice bicoid-related genes, one cDNA clone, Rb24 (EMBL accession number: AJ2771380), was isolated by screening of rice unmature seed cDNA library. Sequence analysis indicates that Rb24 contains a putative amino acid sequence, which is homologous to unique 8 amino acids sequence within Drosophila bicoid homeodomain (50% identity, 75% similarity) and involves a lys-9 in putative helix 3. Northern blot analysis of rice RNA has shown that this sequence is expressed in a tissue-specific manner. The transcript was detected strongly in young panicles, but less in young leaves and roots. This results are further confirmed with paraffin section in situ hybridization. The signal is intensive in rice globular embryo and located at the apical tip of the embryo, then, along with the development of embryo, the signal is getting reduced and transfers into both sides of embryo. The existence of bicoid-related sequence in rice embryo and the similarity of polar distribution of bicoid and Rb24 mRNA in early embryo development may implicates a conserved maternal regulation mechanism of body axis presents in Drosophila and in rice.

Base Sequence↗

Cloning and sequencing of cDNAs encoding the human hepatocyte nuclear factor 4 indicate the presence of two isoforms in human liver.

Hepatocyte nuclear factor 4 (HNF-4) is a key transcription factor involved in the specific expression of many genes in liver and intestine. Sequences of cDNAs coding for HNF-4 have been established in rat and Drosophila melanogaster. Rat HNF-4 exhibits two isoforms which probably result from differential splicing. We have isolated HNF-4 cDNAs from an adult human cDNA library. Sequence analysis revealed that two HNF-4 isoforms are also present in human liver. The complete sequence of the longest human isoform has been established and compared to the rat HNF-4 amino-acid sequences.

Adult↗

Molecular cloning of a novel candidate G protein-coupled receptor from rat brain.

A PCR cloning strategy using primers designed from sequences selectively conserved among a cannabinoid receptor and two orphan receptors, was used to isolate novel G protein-coupled receptors. rCNL3, a 1.75 kb cDNA encoding a 363 amino acid protein, was isolated from a rat cerebral cortex library. Sequence analysis showed that rCNL3 possesses a number of structural characteristics of G protein-coupled receptors and has 61% amino acid identity (from transmembrane region one through the carboxyl-terminus) with two other candidate G protein-coupled receptors. Therefore, these three receptors may comprise a receptor subfamily with identical or closely related endogenous ligands. Northern and in situ hybridization experiments demonstrated that rCNL3 mRNA is expressed in the rat brain, with a prominent distribution in striatum.

Amino Acid Sequence↗

Analysis of 2166 clones from a human colorectal cancer cDNA library by partial sequencing.

Large scale sequencing of cDNAs from a tissue-specific library provides information on the functional phenotype of that tissue and the clones constitute a reservoir of biological markers. For these reasons, we have randomly sequenced 2166 clones of a cDNA library constructed with human colorectal cancer mRNAs. Database searches indicated that 1014 of the cDNAs represented known human genes or human homologs of other genes, 142 sequences corresponded to known ESTs, 119 sequences corresponded to 28S rRNA, repetitive sequences or poly(A) stretches only, and 891 corresponded to unknown transcripts representing the products of 740 new genes. Preliminary studies demonstrated that expression of some of them was altered in cancer. That cDNA collection is therefore a source of potential markers of colorectal cancer.

Animals↗

Identification of the 1.4 kb and 4.0 kb messages for the lipoprotein associated coagulation inhibitor and expression of the encoded protein.

Lipoprotein-Associated Coagulation Inhibitor (LACI) is a factor Xa dependent inhibitor of the factor VII(a)/Tissue Factor catalytic complex. Deduced from partial cDNA sequence, LACI's amino acid sequence has recently been reported. Northern blot analysis showed LACI cDNA hybridizes to RNAs of 1.4 and 4.0 kb in size. To complete the characterization of the LACI message(s), overlapping LACI cDNAs were isolated from a human endothelial cell library. Sequence analysis revealed the clones' inserts span 4023 bases of sequence, consisting of 381 bases of 5' untranslated sequence, an open reading frame of 912 bases, 2682 bases of 3' untranslated sequence and 48 bases of poly(A) sequence. In addition, a short 1.4 kb insert which encodes for LACI was found to contain 49 bases of 3' untranslated sequence and a 3' poly(A) tail. The 1.4 kb of sequence is contained in the 4.0 kb sequence, except for 14 bases of 5' sequence, suggesting that the LACI messages arise by the use of alternative termination and polyadenylation signals during processing. Northern blot analysis of RNA isolated from cells treated with actinomycin D showed both RNA species appear to be relatively stable. Using a bovine papilloma virus vector, LACI cDNA was transfected into mouse C127 fibroblasts. The recombinant LACI is recognized by polyclonal anti-LACI IgG, binds to factor Xa and inhibits VII(a)/Tissue Factor activity in a similar fashion as LACI purified from HepG2 cell conditioned media.

Amino Acid Sequence↗

The viroid and viroid-like RNA database.

The viroid and viroid-like RNA database is a compilation of all natural sequences published in journals or available from the GenBank and EMBL nucleotide sequence libraries. Several information regarding these RNA species such as the position of their self-catalytic domains and the open reading frame of the human hepatitits delta virus are provided. The database also includes a determination of the likely ancestral sequence of most species and a prediction of the most stable secondary structures of these sequences. This online database is available on the World Wide Web (http://www.callisto.si.usherb.ca/[symbol: see text]jpperra ). It should provide an excellent reference point for further phylogenetic and structure-function studies of these RNA species.

Computer Communication Networks↗

Rat and chicken s-rex/NSP mRNA: nucleotide sequence of main transcripts and expression of splice variants in rat tissues.

Two main transcripts of the s-rex/NSP gene are generated by different promoter usage and differential splicing in neuronal and endocrine tissues of higher vertebrates, suggesting that the encoded proteins function in neuroendocrine secretion. To know more about the structure, expression and evolution of this new gene, we have cloned full-length cDNAs for both 1.5 kb and 3.5 kb transcripts from rat and chicken brain cDNA libraries. Sequence analysis has revealed structures within the 3'-UTR that are conserved in these mRNAs and human NSP mRNA and that could be involved in specific compartmentalization of s-rex/NSP mRNA in neuronal cells. An additional transcript generated by differential splicing of internal exons has been cloned from a rat DRG library. Low levels of s-rex/NSP mRNAs have been detected in some non-neuroendocrine tissues, although substantial levels of a unique transcript have been found in rat tests. By RT-PCR analysis, other tissue-specific transcripts that are products of rare splicing events have been revealed.

Alternative Splicing↗

Compilation of small ribosomal subunit RNA structures.

The database on small ribosomal subunit RNA structure contained 1804 nucleotide sequences on April 23, 1993. This number comprises 365 eukaryotic, 65 archaeal, 1260 bacterial, 30 plastidial, and 84 mitochondrial sequences. These are stored in the form of an alignment in order to facilitate the use of the database as input for comparative studies on higher-order structure and for reconstruction of phylogenetic trees. The elements of the postulated secondary structure for each molecule are indicated by special symbols. The database is available on-line directly from the authors by ftp and can also be obtained from the EMBL nucleotide sequence library by electronic mail, ftp, and on CD ROM disk.

Animals↗