PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequencing libraries”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

In silico analysis of 2085 clones from a normalized rat vestibular periphery 3' cDNA library.

The inserts from 2400 cDNA clones isolated from a normalized Rattus norvegicus vestibular periphery cDNA library were sequenced and characterized. The Wackym-Soares vestibular 3' cDNA library was constructed from the saccular and utricular maculae, the ampullae of all three semicircular canals and Scarpa's ganglia containing the somata of the primary afferent neurons, microdissected from 104 male and female rats. The inserts from 2400 randomly selected clones were sequenced from the 5' end. Each sequence was analyzed using the BLAST algorithm compared to the Genbank nonredundant, rat genome, mouse genome and human genome databases to search for high homology alignments. Of the initial 2400 clones, 315 (13%) were found to be of poor quality and did not yield useful information, and therefore were eliminated from the analysis. Of the remaining 2085 sequences, 918 (44%) were found to represent 758 unique genes having useful annotations that were identified in databases within the public domain or in the published literature; these sequences were designated as known characterized sequences. 1141 sequences (55%) aligned with 1011 unique sequences had no useful annotations and were designated as known but uncharacterized sequences. Of the remaining 26 sequences (1%), 24 aligned with rat genomic sequences, but none matched previously described rat expressed sequence tags or mRNAs. No significant alignment to the rat or human genomic sequences could be found for the remaining 2 sequences. Of the 2085 sequences analyzed, 86% were singletons. The known, characterized sequences were analyzed with the FatiGO online data-mining tool (http://fatigo.bioinfo.cnio.es/) to identify level 5 biological process gene ontology (GO) terms for each alignment and to group alignments with similar or identical GO terms. Numerous genes were identified that have not been previously shown to be expressed in the vestibular system. Further characterization of the novel cDNA sequences may lead to the identification of genes with vestibular-specific functions. Continued analysis of the rat vestibular periphery transcriptome should provide new insights into vestibular function and generate new hypotheses. Physiological studies are necessary to further elucidate the roles of the identified genes and novel sequences in vestibular function.

Afferent Pathways↗

Sequence analysis of a rainbow trout cDNA library and creation of a gene index.

Expressed sequence tag (EST) projects have produced extremely valuable resources for identifying genes affecting phenotypes of interest. A large-scale EST sequencing project for rainbow trout was initiated to identify and functionally annotate as many unique transcripts as possible. Over 45,000 5' ESTs were obtained by sequencing clones from a single normalized library constructed using mRNA from six tissues. The production of this sequence data and creation of a rainbow trout Gene Index eliminating redundancy and providing annotation for these sequences will facilitate research in this species.

Animals↗

Gene expression in haemocytes of kuruma prawn, Penaeus japonicus, in response to infection with WSSV by EST approach.

Gene expression in haemocytes of the kuruma prawn (Penaeus japonicus) was investigated using an expressed sequence tag (EST) approach. Partial nucleotide sequences of cDNA library clones constructed from normal and white spot syndrome virus (WSSV)--infected P. japonicus haemocytes were determined. Of 635 clones obtained from the normal library, 284 (44.7%) significantly matched sequences in GenBank, and of 370 clones obtained from WSSV-infected library, 174 (47.0%) significantly matched sequences in the database. One hundred fifty-two deduced proteins were newly identified. Of these, 28 types were involved in biodefence. For the prophenoloxidase system, there are prophenoloxidase, coagulation factor G-beta chain precursor, factor D, Masquarade-like protease, transglutaminase (TGase), clottable protein and eight types of protease inhibitors (two types of antileukoproteinase, alpha-2-macroglobulin, chelonianin, elastase inhibitor, two types of Kazal inhibitor and Kunitz-type inhibitor). For antibacterial peptides, there are bactinecin 11, penaeidin-2 precursor and lysozyme c type. The others defence-related proteins are basophil leukocyte interleukin-3-regulated protein, natural killer enhancing factor (NK-EF), integral membrane protein (CD34+), ESM-1, Notch homologue and Drac homologue. For the adhesion proteins, there are beta-integrin, cell adhesion molecule (CAM) and three types of collagens. All ESTs representing protease inhibitors and tumour-related proteins were found only in the WSSV-infected library. Those encoding for apoptotic peptides were expressed at high levels in infected library. The putative defence proteins accounted for 2.7% of total ESTs in a normal shrimp library and 15.7% of the total ESTs in an infected library.

Amino Acid Sequence↗

Isolation of high-affinity GTP aptamers from partially structured RNA libraries.

Aptamers, RNA sequences that bind to target ligands, are typically isolated by in vitro selection from RNA libraries containing completely random sequences. To see whether higher-affinity aptamers can be isolated from partially structured RNA libraries, we selected for aptamers that bind GTP, starting from a mixture of fully random and partially structured libraries. Because stem-loops are common motifs in previously characterized aptamers, we designed the partially structured library to contain a centrally located stable stem-loop. We used an off-rate selection protocol designed to maximize the enrichment of high-affinity aptamers. The selection produced a surprisingly large number of distinct sequence motifs and secondary structures, including seven different aptamers with K(d)s ranging from 500 to 25 nanomolar. The engineered stem-loop was present in the three highest affinity aptamers, and in 12 of 13 independent isolates with a single consensus sequence, suggesting that its inclusion increased the abundance of high-affinity aptamers in the starting pool.

Base Sequence↗

cDNA cloning and characterization of PD1: a novel human testicular protein with different expressions in various testiculopathies.

A novel human cDNA sequence has been isolated from human testis cDNA library. This sequence, named PD1, reveals an open reading frame encoding a protein of 520 amino acids. A partial sequence similarity has been found with the RBM gene involved in the regulation of human spermatogenesis. Northern blot analysis for PD1 mRNA from several human tissues demonstrated two distinct transcripts of 2.7 (more abundant) and 4.0 kb and revealed that PD1 is expressed in testis and to a lesser extent also in spleen, thymus, and prostate. Immunohistochemical analysis of human testis showed that this protein is detected in the cytoplasm of Sertoli cells. Antibodies against a rhPD1 fragment were used for Western blot analysis, which confirmed the presence of a 60-kDa molecule in crude extract of human testicular cells obtained from fine-needle aspiration and showed different patterns in various testiculopathies, suggesting a role for such gene in human spermatogenesis.

Amino Acid Sequence↗

Structural analysis of the regulatory region of the human corticotropin releasing hormone gene.

A DNA fragment containing the human corticotropin releasing hormone (CRH) gene, along with 9 kb of upstream and 4 kb of downstream sequences, was isolated from a human genomic DNA library. Nucleotide sequence analysis of the proximal 918 nucleotides 5' flanking the putative major mRNA start site of the human gene and comparison to the 866 nucleotide long homologous ovine sequence, revealed that this region of the CRH gene consists of two distinct areas with different degrees of homology, varying from 72% to 94%. The putative functional features of the human sequence were identified. Many, but not all, features were conserved in the ovine sequence. The highly conserved nature of the regulatory region of this gene makes it a good candidate for tracing possible related genetic defects of the hypothalamic-pituitary-adrenal (HPA) axis.

Animals↗

Tissue specific methylation of human Y chromosomal DNA sequences.

This report describes two moderately repetitive human Y chromosomal DNA sequences isolated from a flow sorted Y chromosonal library. These sequences are present in XY male and XY female DNAs but absent in XX male and XX female DNAs. Genomic Southern blot analysis against DNAs isolated from different tissues showed tissue specific DNA methylation patterns. In contrast to the 2.1 kb Hae III repeats which are hypomethylated in sperm DNA, the moderately repetitive sequences used in this study are highly methylated in sperm, less methylated in blood and brain and least methylated in placental DNA.

Base Sequence↗

Gene expression in the gut of keratin-feeding clothes moths (Tineola) and keratin beetles (Trox) revealed by subtracted cDNA libraries.

Few lineages of insects are able to feed on keratin (hair, feathers) and it remains unknown which genes enable this metabolism and what is their evolutionary origin. We conducted a transcriptomic study of two keratin-feeding insects, the clothes moth Tineola bisselliella (Lepidoptera) and the keratin beetle Trox sp. (Coleoptera). Using subtracted cDNA libraries enriched for gut-expressed transcripts, a total of 672 clones sequenced per library resulted in > 150 tentative unique sequences for each species. Sequence similarity predictions identified 22.4% (Tineola) and 6.8% (Trox) of the ESTs as proteases, and mainly as serine proteases of the trypsin and chymotrypsin type, while lacking cysteine proteases. None of the sequences showed similarity to subtilisin type proteases that confers keratinolytic activities in prokaryotes and fungi. Neighbor-Joining trees grouped Tineola and Trox serine proteases near other lepidopteran and coleopteran sequences, respectively, but distant from each other. A few abundant ESTs had no database matches but their presence suggests a role specific to these keratin-feeding insects. While high expression of specific serine proteases appears linked to keratin digestion in both species, it remains to be established if their action requires additional enzymatic or physiological functions to initiate the degradation of the abundant cysteine bonds of keratins. These catabolic pathways are of great interest in the leather industry for the removal of hair, while proteinase inhibitors could prevent damage from clothes moths.

Animals↗

Nucleotide sequence of phospholipase A(2) gene expressed in snake pancreas reveals the molecular evolution of toxic phospholipase A(2) genes.

We have cloned two phospholipase A(2) (PLA(2)) DNA complementary to RNA that contained nucleotide sequences encoding pancreas loop by the reverse transcription-polymerase chain reaction cloning procedure using messenger RNA isolated from Laticauda semifasciata pancreas. Additionally, a gene clone encoding PLA(2) with the pancreatic loop sequence was isolated from a L. semifasciata genomic library. Subsequent sequence analysis revealed that PLA(2) clones encoding group IB" PLA(2). Comparative analysis of group IA and IB" PLA(2) genes revealed that the exon-intron organization is conserved in the genes of both groups. The invaded sequences in the second intron were very similar to those of the L. semifasciata group IA gene. This observation suggested that the integration of the invaded sequences occurred before the divergence of groups IA and IB" during the evolution of PLA(2) gene. The comparative analysis revealed that the arising of group IA PLA(2) occurred by the deletion and substitution of nucleotide sequences in exon III region during the process of accelerated evolution.

Amino Acid Sequence↗

Multiple forms of U2 snRNA coexist in the silk moth Bombyx mori.

Eight U2 snRNA variants were isolated from several Bombyx mori U2-specific RT-PCR libraries. U2 sequences and secondary structures were generated and examined in terms of potential RNA and protein interactions. Analysis indicated that nucleotide changes occurred in both stem/loop and single-stranded areas. Changes in the double stranded areas were either compensatory, single substitutions (e.g. C <--> U) or prevented the double-stranded formation of one or two base pairs. The polymorphisms were clustered in moderately conserved regions. Some of the changes observed generated stronger base pairing. Inter-species conserved protein or RNA-binding sites were relatively unaffected. No polymorphic sites were found in known functional sequences. Bombyx mori and Drosophila melanogaster U2 sequences are 95% and 70% similar at the 5'- and the 3'-ends of the molecule, respectively. Phylogenetic analysis of the U2 sequences demonstrates remarkable conservation across species.

Animals↗

Analysis of expressed sequence tags (ESTs) from the scaly green flagellate Scherffelia dubia Pascher emend. Melkonian et Preisig.

Partial sequencing of cDNA libraries to generate expressed sequence tags (ESTs) is an effective means of gene discovery, generation of molecular markers and characterization of transcription patterns. We have constructed an EST-database of the scaly green flagellate Scherffelia dubia (Chlorophyta) containing 361 sequences. cDNAs were obtained from interphase cells and from cells regenerating flagella. Analysis of the ESTs identified 138 EST-groups with significant similarity to known sequences. 134 EST-groups showed no significant similarity to any sequences in the databases. Most of the ESTs with similarity to known proteins are associated with typical interphase cell functions of a photosynthetic plant cell: assimilation of nutrients and biosynthesis of proteins. Others are related to the activation of the secretory pathway or the biogenesis of scales (e.g. kdo-synthase). Comparison of S. dubia ESTs with the genome of Arabidopsis thaliana and the EST database of Chlamydomonas reinhardtii revealed that S. dubia ESTs with similarity to known proteins were more similar to sequences in C. reinhardtii than to those of A. thaliana. Additionally, ESTs for guanylyl cyclase and cGMP phosphodiesterase are present in the two flagellates, but so far these gene products have not been found in embryophytes.

3' Untranslated Regions↗

Profile hidden Markov models.

The recent literature on profile hidden Markov model (profile HMM) methods and software is reviewed. Profile HMMs turn a multiple sequence alignment into a position-specific scoring system suitable for searching databases for remotely homologous sequences. Profile HMM analyses complement standard pairwise comparison methods for large-scale sequence analysis. Several software implementations and two large libraries of profile HMMs of common protein domains are available. HMM methods performed comparably to threading methods in the CASP2 structure prediction exercise.

Humans↗

Mouse cysteine-rich secretory protein 4 (CRISP4): a member of the Crisp family exclusively expressed in the epididymis in an androgen-dependent manner.

The final maturation of spermatozoa produced in the testis takes place during their passage through the epididymis. In this process, the proteins secreted into the epididymal lumen along with changes in the pH and salt composition of the epididymal fluid cause several biochemical changes and remodeling of the sperm plasma membrane. The Crisp family is a group of cysteine-rich secretory proteins that previously consisted of three members, one of which-CRISP1-is an epididymal protein shown to attach to the sperm surface in the epididymal lumen and to inhibit gamete membrane fusion. In the present paper, we introduce a new member of the Crisp protein family, CRISP4. The new gene was discovered through in silico analysis of the epididymal expressed sequence tag library deposited in the UniGene database. The peptide sequence of CRISP4 has a signal sequence suggesting that it is secreted into the epididymal lumen and might thus interact with sperm. Unlike the other members of the family, Crisp4 is located on chromosome 1 in a cluster of genes encoding for cysteine-rich proteins. Crisp4 is expressed in the mouse exclusively in epithelial cells of the epididymis in an androgen-dependent manner, and the expression of the gene starts at puberty along with the onset of sperm maturation. The identified murine CRISP4 peptide has high homology with human CRISP1, and the homology is higher than that between murine and human CRISP1, suggesting that CRISP4 represents the mouse counterpart of human CRISP1 and could have similar effects on sperm membrane as mouse and human CRISP1.

Amino Acid Sequence↗

Selection of phage-display peptides that bind specifically to the outer coat protein of Rice black streaked dwarf virus.

Several peptides that could bind specifically to the outer coat protein encoded by the S10 gene of Rice black streaked virus (RBSDV) were isolated from a phage-display random 12-mer peptide library. The sequence analysis showed that the amino acid motif (K)K**(*)P, the asterisk denoting any amino acid, might be the core sequence by which the peptides bind to the target protein. The peptide 1 that had a high affinity to RBSDV outer coat protein was synthesized by a chemical method and its fusion protein with glutathione-S-transferase (GST) was produced in an Escherichia coli expression system. The dot and Western blot analyses indicated that RBSDV could be detected with a high sensitivity in crude extracts of diseased plant leaves using a purified GST fusion protein. The circular dichroism (CD) spectroscopy revealed that the synthesized binding peptide but not a nonbinding peptide could bring about a marked change in the conformation of outer coat RBSDV protein. Since the protein functions only when it has correct conformation, the peptides binding specifically to it could possibly disturb the function of the virus outer coat protein and might be used to block the transmission pathway of the virus. Summing up, as these peptides showed a high specificity and sensitivity and diagnostic potential for RBSDV, they may represent the basis of a novel strategy for development of resistance to RBSDV.

Artificial Gene Fusion↗

Mutagenesis using trinucleotide beta-cyanoethyl phosphoramidites.

There is no easy way to selectively introduce mixtures of codon triplets into mutagenesis libraries. Solid-phase-supported DNA synthesis using successive coupling of mixtures of mononucleotides can be made to supply 32 codons, which gives redundancies in coding for 20 natural amino acids, as well as an often unwanted stop codon. Resin-splitting methods have been described, but the representation of all permutations is limited by mechanical factors for a large library, and the method is experimentally cumbersome. To demonstrate a third, improved method, the 3'-cyanoethyl phosphoramidite codon triplets dATA, dCTT, dATC, dATG and dAGC were made by solution-phase methods, with protecting groups fully compatible with modern automated phosphoramidite DNA synthesis chemistry. The reagents were then used to synthesize a 54-mer DNA fragment, wherein 15 internal base pairs were randomized by coupling a mixture of the five codons five times. The fragment was amplified as a cDNA pool, which was subcloned into a phagemid vector, and 16 randomly selected recombinants from this mini-library were sequenced. These clones showed random incorporation of the proper transcribed codon sequences at the correct location. Other functional tests involving the trinucleotide phosphoramidites showed modest (ca. 70%) coupling efficiencies and structural integrity of the DNA produced.

Amino Acid Sequence↗

Cystatin E is a novel human cysteine proteinase inhibitor with structural resemblance to family 2 cystatins.

A new member of the human cystatin superfamily, called cystatin E, has been found by expressed sequence tag (EST) sequencing in amniotic cell and fetal skin epithelial cell cDNA libraries. The sequence of a full-length amniotic cell cDNA clone contained an open reading frame encoding a putative 28-residue signal peptide and a mature protein of 121 amino acids, including four cysteine residues and motifs of importance for the inhibitory activity of Family 2 cystatins like cystatin C. Recombinant cystatin E was produced in a baculovirus expression system and isolated. An antiserum against the recombinant protein could be used for affinity purification of cystatin E from human urine, as confirmed by N-terminal sequencing. The mature recombinant protein processed by insect cells started at amino acid 4 (cystatin C numbering), and displayed reversible inhibition of papain and cathepsin B (Ki values of 0.39 and 32 nM, respectively), in competition with substrate. Cystatin E is thus a functional cysteine proteinase inhibitor despite relatively low amino acid sequence similarities with human cystatins (26-34% identity with sequences for the Family 2 cystatins C, D, S, SN, and SA; <30% with the Family 1 cystatins, A and B, and domains 2 and 3 of the Family 3 cystatin, kininogen). Unlike other human low Mr cystatins, cystatin E is a glycoprotein, carrying an N-linked carbohydrate chain at position 108. Northern blot analysis revealed that the cystatin E gene is expressed in most human tissues, with the highest mRNA amounts found in uterus and liver. A strikingly high incidence of cystatin E clones in cDNA libraries from fetal skin epithelium and amniotic membrane cells (>0.5% of clones sequenced) indicates a protective role of cystatin E during fetal development.

Amino Acid Sequence↗

[14 sequences from Chinese hamster genome preferentially binding to the nuclear matrix].

Fifteen sequences belonging to the Chinese hamster genome were isolated from a library of sequences preferentially binding to the nuclear matrix (matrix attachment regions, MAR), sequenced, and characterized. Fourteen of the 15 sequences (> 90%) bound to the nuclear matrix with affinities 2.5-60 times higher than those of control DNA fragments containing no MARs. One clone displayed a considerable homology to the ORF1 region of the mouse LINE repeat. Such MARs within LINE repeats may considerably alter the activities of some genes and the transcription status of chromatin domains upon the LINE repeat propagation in the genome over the course of evolution.

Amino Acid Sequence↗