PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequencing libraries”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,081 records · Page 60Linked to original sources

Purification of a NF1-like DNA-binding protein from rat liver and cloning of the corresponding cDNA.

NF1-like proteins play a role in transcription of liver-specific genes. A DNA-binding protein, recognizing half of the canonical NF1 binding site (TGGCA) present on the human albumin and retinol-binding protein genes, has been purified from rat liver. Several peptides deriving from a tryptic digest of the purified protein were sequenced and the sequence was used to synthesize specific oligonucleotides. Two overlapping cDNA clones were obtained from a rat-liver cDNA library; their sequence reveals an open reading frame coding for 505 amino acids, including all the peptides sequenced from the purified protein. The DNA-binding domain, most likely located within the first 250 amino acids, is highly homologous to the sequence of CTF/NF1 purified from HeLa cells. Northern analysis reveals several mRNA species present in different combinations in various rat tissues.

Amino Acid Sequence↗

Cloning of a new human gene with short consensus repeats using the EST database.

The complement system, which provides many of the effector functions of humoral immunity and inflammation, is tightly regulated by various complement regulatory proteins. The most common structural feature of these proteins is a motif called short consensus repeat (SCR). In order to identify a new human complement regulatory protein, we performed a similarity search using SCR on the expressed sequence tag (EST) database and found a partial sequence of a new human gene. Using a probe containing this partial sequence, we obtained a full-length cDNA of this gene from a human umbilical vein endothelial cell (HUVEC) library. The sequencing reaction demonstrated an open reading frame of 1383 nucleotides coding for a 461 amino acid polypeptide with a deduced relative molecular mass of 51 000. Structural analysis showed that the protein has three SCRs with one transmembrane domain. A characteristic feature of these SCR was that they have six conserved cysteines per repeat instead of the usual four. Therefore, we named this cDNA THECY (three hexa-cysteine motifs). A six cysteine motif is a characteristic feature of selectins. We used northern blot analysis to show that a 2.0 kilobase (kb) transcript was ubiquitously present in most organs studied, and the mRNA was most abundant in the heart. In conclusion, we discovered a member of a new class of membrane-bound SCR-containing molecules using the EST database. Utilization of the EST database may be useful in the search for other new immunological proteins. The function of this gene remains to be elucidated.

Amino Acid Sequence↗

Weight matrix descriptions of four eukaryotic RNA polymerase II promoter elements derived from 502 unrelated promoter sequences.

Optimized weight matrices defining four major eukaryotic promoter elements, the TATA-box, cap signal, CCAAT-, and GC-box, are presented; they were derived by comparative sequence analysis of 502 unrelated RNA polymerase II promoter regions. The new TATA-box and cap signal descriptions differ in several respects from the only hitherto available base frequency Tables. The CCAAT-box matrix, obtained with no prior assumption but CCAAT being the core of the motif, reflects precisely the sequence specificity of the recently discovered nuclear factor NY-I/CP1 but does not include typical recognition sequences of two other purported CCAAT-binding proteins, CTF and CBP. The GC-box description is longer than the previously proposed consensus sequences but is consistent with Sp1 protein-DNA binding data. The notion of a CACCC element distinct from the GC-box seems not to be justified any longer in view of the new weight matrix. Unlike the two fixed-distance elements, neither the CCAAT- nor the GC-box occurs at significantly high frequency in the upstream regions of non-vertebrate genes. Preliminary attempts to predict promoters with the aid of the new signal descriptions were unexpectedly successful. The new TATA-box matrix locates eukaryotic transcription initiation sites as reliably as do the best currently available methods to map Escherichia coli promoters. This analysis was made possible by the recently established Eukaryotic Promoter Database (EPD) of the EMBL Nucleotide Sequence Data Library. In order to derive the weight matrices, a novel algorithm has been devised that is generally applicable to sequence motifs positionally correlated with a biologically defined position in the sequences. The signal must be sufficiently over-represented in a particular region relative to the given site, but need not be present in all members of the input sequence collection. The algorithm iteratively redefines the set of putative motif representatives from which a weight matrix is derived, so as to maximize a quantitative measure of local over-representation, an optimization criterion that naturally combines structural and positional constancy. A comprehensive description of the technique is presented in Methods and Data.

Algorithms↗

The murine genes Hox-5.1 and Hox-4.1 belong to the same HOX complex on chromosome 2.

Two different loci of Antennapedia-related homeobox-containing genes have been shown to map to mouse chromosome 2: the HOX-5 complex and the Hox-4.1 gene. These independently derived loci are likely to be parts of a single gene complex, although their close linkage has not yet been demonstrated. Since cosmid walks to extend the HOX-5 cluster and to potentially link the two loci were unsuccessful, we have used large restriction fragments separated by pulsed-field gel electrophoresis to demonstrate the linkage between probes from the HOX-5 region and sequences near Hox-4.1. To further define the distance between the two linked loci, we screened a NotI jumping library with sequences near the Hox-5.1 gene to obtain a marker within the region predicted to contain Hox-4.1. The jumping endpoint lies within genomic clones from a lambda phage walk extending from the 5' end of Hox-4.1, and thus provides clear evidence of linkage between the two Hox loci. Our results demonstrate that Hox-4.1 lies approximately 35 kb downstream of the Hox-5.1 gene and that the two loci do indeed thus constitute parts of the same HOX complex.

Animals↗

Impact of flooding on soil bacterial communities associated with poplar (Populus sp.) trees.

Soil bacterial communities were analyzed in different habitats (bulk soil, rhizosphere, rhizoplane) of poplar tree microcosms (Populus tremulaxP. alba) using cultivation-independent methods. The roots of poplar trees regularly experience flooded and anoxic conditions. Therefore, we also determined the effect of flooding on microbial communities in microcosm experiments. Total community DNA was extracted and bacterial 16S rRNA genes were amplified by PCR and analyzed by terminal restriction fragment length polymorphism (T-RFLP) analysis, cloning and sequencing. Clone libraries were created from all three habitats under both unflooded and flooded conditions resulting in a total of 281 sequences. Numbers of different sequences (<97% similarity) in the different habitats represented 16-55% of total bacterial species richness determined from the nonparametric richness estimator Chao1. According to the number of different terminal restriction fragments (T-RFs), all of the different habitats contained approximately 20 different operational taxonomic units (OTUs), except the flooded rhizoplane habitat whose community contained less OTUs. Results of cloning and T-RFLP analysis generally supported each other. Correspondence analysis of T-RFLP patterns showed that the bacterial communities were different in bulk soil, rhizosphere and rhizoplane and changed upon flooding. For example OTUs representing Bacillus sp. were highest in the unflooded bulk soil and rhizosphere. Sequences related to Aquaspirillum, in contrast, were predominant on the poplar roots and in the rhizosphere of flooded microcosms but were rarely found in the other habitats.

Bacteria↗

The sequence of palF, an environmental pH response gene in Aspergillus nidulans.

To molecularly characterize the influence of external pH in the secretion of enzymes by filamentous fungi, we have cloned and sequenced the palF gene of Aspergillus nidulans (An). An An wild-type (wt) chromosome VII-specific cosmid library was used to transform a palF15 mutant strain. Selection for complementation was done on medium containing beta-glycerolphosphate as the sole Pi source. Two cosmids were identified (W2G08 and W4G12) and further subcloning of one cosmid (W2G08) defined a 5-kb PstI genomic fragment, which fully complements the palF15 mutation. An internal fragment from the genomic clone recognized a single message of approx. 3.5 kb on Northern blot. cDNA clones were obtained from a lambda gt10 cDNA library and sequenced, showing a nucleotide (nt) sequence of 3311 base pairs (bp), with a 828 bp long 5'-untranslated region (UTR). The major open reading frame (ORF) identified in the sequence codes for a putative 775 amino acid (aa) protein, which shares some similarity with two putative ORF products of Saccharomyces cerevisiae (Sc).

Amino Acid Sequence↗

cDNA heterogeneity suggests structural variants related to the high-affinity IgE receptor.

The high-affinity IgE receptor present on mast cells and basophils is responsible for the IgE-mediated activation of these cells. The current model for this receptor depicts a four-subunit structure, alpha beta gamma 2. A cDNA for the alpha subunit was recently cloned and predicts a structure consisting of two homologous extracellular domains, a transmembrane segment, and a cytoplasmic tail. Using a synthetic oligonucleotide corresponding to the amino-terminal sequence of the alpha subunit, we identified a number of cDNA clones from a rat basophilic leukemia cell cDNA library. Nucleotide sequencing established four different forms of cDNA: one is nearly identical to the published cDNA; the second differs from the first in the 5' untranslated sequence; the other two forms use either one or the other of the 5'-end sequences as above and lack 163 base pairs in the region coding for the second extracellular domain. RNase protection analysis with radioactive RNA probes established the heterogeneity of rat basophilic leukemia cell mRNA with regard to both the 5' and the internal sequences. Our results suggest the existence of at least four different protein forms related to the alpha subunit of the high-affinity IgE receptor.

Amino Acid Sequence↗

Cloning and characterization of chicken IL-10 and its role in the immune response to Eimeria maxima.

We isolated the full-length chicken IL-10 (chIL-10) cDNA from an expressed sequence tag library derived from RNA from cecal tonsils of Eimeria tenella-infected chickens. It encodes a 178-aa polypeptide, with a predicted 162-aa mature peptide. Chicken IL-10 has 45 and 42% aa identity with human and murine IL-10, respectively. The structures of the chIL-10 gene and its promoter were determined by direct sequencing of a bacterial artificial chromosome containing chIL-10. The chIL-10 gene structure is similar to (five exons, four introns), but more compact than, that of its mammalian orthologues. The promoter is more similar to that of Fugu IL-10 than human IL-10. Chicken IL-10 mRNA expression was identified mainly in the bursa of Fabricius and cecal tonsils, with low levels of expression also seen in thymus, liver, and lung. Expression was also detected in PHA-activated thymocytes and LPS-stimulated monocyte-derived macrophages, with high expression in an LPS-stimulated macrophage cell line. Recombinant chIL-10 was produced and bioactivity demonstrated through IL-10-induced inhibition of IFN-gamma synthesis by mitogen-activated lymphocytes. We measured the expression of mRNA for chIL-10 and other signature cytokines in gut and spleen of resistant (line C.B12) and susceptible (line 15I) chickens during the course of an E. maxima infection. Susceptible chickens showed higher levels of chIL-10 mRNA expression in the spleen, both constitutively and after infection, and in the small intestine after infection than did resistant chickens. These data indicate a potential role for chIL-10 in changing the Th bias during infection with an intracellular protozoan, thereby contributing to susceptibility of line 15I chickens.

Amino Acid Sequence↗

Enhancer sequences from Arabidopsis thaliana obtained by library transformation of Nicotiana tabacum.

In this paper we report on the use of a bidirectional enhancer cloning vehicle to isolate and characterize new enhancer sequences from Arabidopsis thaliana. A library of A. thaliana genomic Sau3A segments was constructed in Escherichia coli in the binary plasmid enhancer cloning vehicle pROA97. The T-DNA based vector carries abbreviated TATA regions from the cauliflower mosaic virus 35S transcription unit upstream of two genes. The library was transferred via triparental mating into Agrobacterium tumefaciens. The neomycin phosphotransferase II gene was used for selection of kanamycin-resistant transformed tobacco callus cells. Approximately 1100 transgenic plants were regenerated and assayed for expression of the E. coli beta-glucuronidase (GUS) gene in leaves, stems, roots, or seeds. Plasmids carrying putative enhancer sequences were rescued from the genomes of transgenic plants and the cloned sequences were assayed for enhancer function in genetic selection experiments. Plants were regenerated from the kanamycin-resistant calli obtained in the secondary transformation experiments. Histochemical analysis of GUS activity in the leaf, stem, and root tissues of transgenic plants showed a variety of expression patterns. The DNA sequences are presented of five Arabidopsis segments which confer enhancer function.

Base Sequence↗

A dinucleotide mutation in dihydrodipicolinate synthase of Nicotiana sylvestris leads to lysine overproduction.

By applying a mutagenesis/selection procedure to obtain resistance to a lysine analog, S-(2-aminoethyl)L-cysteine (AEC), a lysine overproducing mutant in Nicotiana sylvestris was isolated. Amino acid analyses performed throughout plant development and of different organs of the N. sylvestris RAEC-1 mutant, revealed a developmental-dependent accumulation of free lysine. Lysine biosynthesis in the RAEC-1 mutant was enhanced due to a lysine feedback-desensitized dihydrodipicolinate synthase (DHDPS). Several molecular approaches were undertaken to identify the nucleotide change in the dhdps-r1 gene, the mutated gene coding for the lysine-desensitized enzyme. The enzyme was purified from wild-type plants for amino end microsequencing and 10 amino acids were identified. Using dicotyledon dhdps probes, a genomic fragment was cloned from an enriched library of DNA from the homozygote RAEC-1 mutant plant. A dhdps cDNA, putatively full-length, was isolated from a tobacco cDNA library. Nucleotide sequence analyses confirmed the presence of the previously identified amino end preceded by a chloroplast transit peptide sequence. Nucleotide sequence comparisons, enzymatic and immunological analyses revealed that the tobacco cDNA corresponds to a normal type of DHDPS, lysine feedback-regulated, and the genomic fragment to the mutated DHDPS, insensitive to lysine inhibition. Functional complementation of a DHDPS-deficient Escherichia coli strain was used as an expression system. Reconstruction between the cDNA and genomic fragment led to the production of a cDNA producing an insensitive form of DHDPS. Amino acid sequence comparisons pointed out, at position 104 from the first amino acid of the mature protein, the substitution of Asn to Ileu which corresponds to a dinucleotide mutation. This change is unique to the dhdps-r1 gene when compared with the wild-type sequence. The identification of the nucleotide and amino acid change of the lysine-desensitized DHDPS from RAEC-1 plant opens new perspectives for the improvement of the nutritional value of crops and possibly to develop a new plant selectable marker.

Amino Acid Sequence↗

Cloning and sequencing of CATR1.3, a human gene associated with tumorigenic conversion.

The human squamous cell carcinoma cell line SCC83-01-82 (SCC) contains mutations in both the H-ras and p53 genes, but it exhibits a nontumorigenic phenotype in nude mice. This cell line can be converted into a cell line with a tumorigenic phenotype, SCC83-01-82CA (CA), by treatment with the mutagen methyl methanesulfonate (MMS). This indicates that additional genetic events leading to expression of a cooperating tumor susceptibility gene(s) may be required for tumorigenicity. To identify the cooperating gene(s), an expression cDNA library was made from tumorigenic Ca cells. The library DNA was transfected into nontumorigenic SCC cells and the transfected SCC cells were then injected into nude mice for the selection of a tumorigenic phenotype. Tumors developed in 3 of the 18 mice after injection. Several new cell lines were established from these transfected cell-induced tumors and designated as CATR cells. Tumor histology and karyotype analysis of these cells indicated that they were of human epithelial cell origin. All the CATR cells have the library vector sequence integrated in their genome. Cell line CATR1 expressed a single message from the integrated library representing a 1.3-kb cDNA insert that was absent from untransfected SCC cells or MMS-converted CA cells. This 1.3-kb cDNA insert was cloned by PCR amplification of reverse-transcribed CATR1 total RNA and was designated CATR1.3. The nucleotide sequence of CATR1.3 encodes a peptide of 79 amino acids, has a long 3' untranslated region, and represents an unknown gene product that was associated with the tumorigenic conversion due to the transfected expression library.

Amino Acid Sequence↗

A modular assembly strategy for improving the substrate specificity of small catalytic peptides.

In contrast to large proteins, small peptide catalysts typically display limited specificity for small molecule substrates. This is presumably a result of the limited opportunities small peptides have to fold in a manner that provides for the formation of an isolated reaction vessel that effectively binds and sequesters substrates from bulk solvent while at the same time catalyzing their transformation. For the preparation of small peptide catalysts that possess improved substrate specificity, we have developed a modular assembly strategy that involves appending phage display-derived substrate binding-domain modules to catalytically active peptide domains. We demonstrate the potential of this strategy with the construction of a small 35-amino acid residue aldolase peptide with improved substrate specificity. The advantages of this approach are that it reduces the demand on the functionalization of the catalytic site and it is modular, therefore making its adaptation to a variety of specificities rapid. The modular assembly strategy studied here may present advantages over exhaustive searches of large random-sequence peptide libraries for peptides with singular function.

Amino Acid Sequence↗

Gene discovery within the planctomycete division of the domain Bacteria using sequence tags from genomic DNA libraries.

BACKGROUND: The planctomycetes comprise a distinct group of the domain Bacteria, forming a separate division by phylogenetic analysis. The organization of their cells into membrane-defined compartments including membrane-bounded nucleoids, their budding reproduction and complete absence of peptidoglycan distinguish them from most other Bacteria. A random sequencing approach was applied to the genomes of two planctomycete species, Gemmata obscuriglobus and Pirellula marina, to discover genes relevant to their cell biology and physiology. RESULTS: Genes with a wide variety of functions were identified in G. obscuriglobus and Pi. marina, including those of metabolism and biosynthesis, transport, regulation, translation and DNA replication, consistent with established phenotypic characters for these species. The genes sequenced were predominantly homologous to those in members of other divisions of the Bacteria, but there were also matches with nuclear genomic genes of the domain Eukarya, genes that may have appeared in the planctomycetes via horizontal gene transfer events. Significant among these matches are those with two genes atypical for Bacteria and with significant cell-biology implications - integrin alpha-V and inter-alpha-trypsin inhibitor protein - with homologs in G. obscuriglobus and Pi. marina respectively. CONCLUSIONS: The random-sequence-tag approach applied here to G. obscuriglobus and Pi. marina is the first report of gene recovery and analysis from members of the planctomycetes using genome-based methods. Gene homologs identified were predominantly similar to genes of Bacteria, but some significant best matches to genes from Eukarya suggest that lateral gene transfer events between domains may have involved this division at some time during its evolution.

Amino Acids↗

Human complement factor B: cDNA cloning, nucleotide sequencing, phenotypic conversion by site-directed mutagenesis and expression.

A full-length cDNA clone, BHL4-1, encoding factor B was isolated from a human liver cDNA library and sequenced in its entirety. It consists of 2388 bp which include a 5'-untranslated region of 40 bp, a single open reading frame, 2292 bp in length, and a 3'-untranslated region of 56 bp followed by a poly-A tail. The deduced amino acid sequence comprises 25 residues of a putative leader peptide and 739 residues of the mature polypeptide chain of the F allele of factor B. We constructed an S allele-like Q7R mutant of BHL4-1 by site-directed mutagenesis. Both the wild-type and mutant factor B cDNA were expressed transiently in a eukaryotic system. The specific hemolytic activities of the two recombinant factor B alleles and of native B were not significantly different from each other.

Alleles↗

The first nucleotide sequence of an archaeal elongation factor 1 beta gene.

An archaeal elongation factor 1 beta gene has been isolated for the first time from a Sulfolobus solfataricus genomic library. The sequenced clone (869 bp) contained two open reading frames, one coding for a protein made of 91 amino acid residues (SsEF-1 beta), the other one encoding a nonidentified product (ORF 115). The amino acid sequences of segments at the N- and C-terminal of the translated SsEF-1 beta were identical to those determined for the native protein. Northern and Southern analyses showed that the SsEF-1 beta gene is represented in S. solfataricus by a unique sequence. Compared to eubacterial or eukaryal corresponding genes the SsEF-1 beta is much shorter.

Amino Acid Sequence↗

Cloning and sequence analysis of the rat tumor necrosis factor-encoding genes.

We have isolated an approximately 22-kb TNF locus (encoding tumor necrosis factor) from a rat genomic library and sequenced the 7105-bp fragment that comprises the TNF-alpha and TNF-beta genes, including their flanking sequences. The two genes are tandemly arranged with TNF-beta 5' to TNF-alpha and separated by approximately 1.1 kb of intergenic space, and each gene consists of four exons and three introns, similar to those of the other species examined thus far. Comparison analysis showed that the rat TNF have high sequence homology with the mouse TNF (TNF-alpha, 86.5%; TNF-beta, 89.3%) and relatively low homology with the human, rabbit, and porcine TNF. The upstream sequence of rat TNF-alpha contains a number of sequence motifs implicated in the expression and regulation of eukaryotic genes, including binding sites for the transcription factors Sp-1, Ap-2, IFN.1 and NF-kappa B. The possible significance of potential regulatory sequence elements found in the rat TNF-alpha in the context of transcriptional regulatory mechanisms is discussed.

Amino Acid Sequence↗

Cloning, sequencing and overexpression of the leucine dehydrogenase gene from Bacillus cereus.

The L-leucine dehydrogenase gene from Bacillus cereus (DSM 626) was cloned from a partial genomic library and sequenced. The open reading frame has 1101 bp and codes for a protein of 39.9 kDa. The deduced amino acid sequence of the LeuDH from B. cereus shares 70-80% identity with LeuDH's from the thermophilic strains B. stearothermophilus and Thermoactinomyces intermedius. The active protein was overexpressed in Escherichia coli to yield approximately 30% of the total soluble protein.

Amino Acid Oxidoreductases↗

A novel porcine gene, alpha-1-antichymotrypsin 2 (SERPINA3-2): sequence, genomic organization, polymorphism and mapping.

A novel porcine gene, alpha-1-antichymotrypsin 2 (SERPINA3-2), a member of the serpin superfamily, was isolated from a porcine genomic library and sequenced. The genomic organization of the approximately 9.0 kb gene was determined on the basis of the porcine liver cDNA of SERPINA3-1 and SERPINA3-2, and comprises five exons and four introns. The coding sequence of SERPINA3-2 shares 86% identity with the paralogue, SERPINA3-1. Porcine SERPINA3-2 was found to be an orthologue of human SERPINA3 (71% identity of the coding sequences) and both genes have a similar genomic organization. Polymorphisms were found in intron 4 of the porcine gene using polymerase chain reaction-restriction fragment length polymorphism. The gene was mapped by linkage analysis and radiation hybrid mapping to the distal end of chromosome 7q, to the gene cluster of the protease inhibitors including PI1 (SERPINA1), PI2, PI3, PI4 (apparently paralogues of SERPINA3), and PO1A and PO1B. SERPINA3-2 is the first porcine serpin gene whose genomic organization has been determined.

Amino Acid Sequence↗