PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequencing libraries”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 631 records · Page 35Linked to original sources

Characterization of new Schistosoma mansoni microsatellite loci in sequences obtained from public DNA databases and microsatellite enriched genomic libraries.

In the last decade microsatellites have become one of the most useful genetic markers used in a large number of organisms due to their abundance and high level of polymorphism. Microsatellites have been used for individual identification, paternity tests, forensic studies and population genetics. Data on microsatellite abundance comes preferentially from microsatellite enriched libraries and DNA sequence databases. We have conducted a search in GenBank of more than 16,000 Schistosoma mansoni ESTs and 42,000 BAC sequences. In addition, we obtained 300 sequences from CA and AT microsatellite enriched genomic libraries. The sequences were searched for simple repeats using the RepeatMasker software. Of 16,022 ESTs, we detected 481 (3%) sequences that contained 622 microsatellites (434 perfect, 164 imperfect and 24 compounds). Of the 481 ESTs, 194 were grouped in 63 clusters containing 2 to 15 ESTs per cluster. Polymorphisms were observed in 16 clusters. The 287 remaining ESTs were orphan sequences. Of the 42,017 BAC end sequences, 1,598 (3.8%) contained microsatellites (2,335 perfect, 287 imperfect and 79 compounds). The 1,598 BAC end sequences 80 were grouped into 17 clusters containing 3 to 17 BAC end sequences per cluster. Microsatellites were present in 67 out of 300 sequences from microsatellite enriched libraries (55 perfect, 38 imperfect and 15 compounds). From all of the observed loci 55 were selected for having the longest perfect repeats and flanking regions that allowed the design of primers for PCR amplification. Additionally we describe two new polymorphic microsatellite loci.

Animals↗

Peptide agonist of the thrombopoietin receptor as potent as the natural cytokine.

Two families of small peptides that bind to the human thrombopoietin receptor and compete with the binding of the natural ligand thrombopoietin (TPO) were identified from recombinant peptide libraries. The sequences of these peptides were not found in the primary sequence of TPO. Screening libraries of variants of one of these families under affinity-selective conditions yielded a 14-amino acid peptide (Ile-Glu-Gly-Pro-Thr-Leu-Arg-Gln-Trp-Leu-Ala-Ala-Arg-Ala) with high affinity (dissociation constant approximately 2 nanomolar) that stimulates the proliferation of a TPO-responsive Ba/F3 cell line with a median effective concentration (EC50) of 400 nanomolar. Dimerization of this peptide by a carboxyl-terminal linkage to a lysine branch produced a compound with an EC50 of 100 picomolar, which was equipotent to the 332-amino acid natural cytokine in cell-based assays. The peptide dimer also stimulated the in vitro proliferation and maturation of megakaryocytes from human bone marrow cells and promoted an increase in platelet count when administered to normal mice.

Amino Acid Sequence↗

Construction of a multi-functional cDNA library specific for mouse pancreatic islets and its application to microarray.

We have constructed a high-quality and multi-applicable cDNA library specific for mouse pancreatic islets. This is the first pancreatic islet cDNA library created using a recombination-based method, which can readily be converted into other applications including yeast two-hybrid and mammalian expression libraries. Based on sequence data of the library, we constructed a sequence database specific for mouse pancreatic islets. Among the 8882 non-redundant clones, 5799 were classified into specific functional categories using a classification system designed by the Gene Ontology Consortium, 10% of which were "molecular function unknown" genes. We also developed cDNA microarray membranes with 8108 non-redundant clones. Analyses of expression profiles of three different cell lines and of MIN6 cells with or without overexpression of transcription factor NeuroD1 established the usefulness and applicability of our microarrays. The mouse pancreatic islet cDNA library, sequence database, set of clones, and microarrays developed in this study should be useful resources for studies of pancreatic islets and related diseases including diabetes mellitus.

Animals↗

Expressed sequence tags for the chicken genome from a normalized, ten-day-old white leghorn whole embryo cDNA library. 2. Comparative DNA sequence analysis of guinea fowl, quail, and turkey genomes.

Accelerated efforts to develop a high-utility chicken genome map have resulted in the development of resources that may be useful for genetic analysis in other economically important poultry species. Here we describe a total of 26 comparative genomic DNA sequences (CGS) for the guinea fowl, Japanese quail, and domestic turkey developed using 10 primer pairs specific for 10 previously reported, unique, chicken expressed sequence tags (EST). The total length of CGS developed for each of the three species was 4,193, 4,597, and 6,057 bp in quail, turkey, and guinea fowl, respectively. About 70% of the CGS showed significant sequence similarity to reference database sequences, including the reference chicken EST and other avian and nonavian genes. A majority of the between-species comparisons of the CGS from all but two primer pairs were significant and ranged from 81 to 99%. The percentage similarity of the CGS appears to be a function of phylogenetic relatedness and was generally higher for comparisons between the chicken, quail, and turkey and lower between the guinea fowl and chicken, quail, or turkey. Maximum likelihood estimation of the phylogenetic relationships using CGS from two primer pairs also showed a closer relationship, as expected, among chicken, quail, and turkey than between guinea fowl and either chicken, quail, or turkey. Within the guinea fowl, quail, and turkey CGS developed, the total number of single nucleotide polymorphisms detected was 28, 17, and 14, respectively. Together, these resources represent tools that will facilitate genetic analysis of species that have been studied very little and our understanding of their genomes and genome evolution.

Animals↗

PLANET: a phage library analysis expert tool.

In recent years random peptide libraries displayed on filamentous phage have been widely used and new ideas and techniques are continuously developing in the field (1-5). Notwithstanding this growing interest in the technique and in its promising results, and the enormous increase in usage and scope, very little effort has been devoted to the implementation of software able to handle and analyze the growing number of phage library-derived sequences. In our laboratory, phage libraries are extensively used and peptide sequences are continuously produced, so that the need arose of creating a database (6) to collect all the experimental results in a format compatible with GCG sequence analysis packages (7). We present here the description of an XWindow-based software package named PLANET (Phage Library ANalysis Expert Tool) devoted to the maintenance and statistical analysis of the database.

Bacteriophages↗

Representation of DNA sequences in recombinant DNA libraries prepared by restriction enzyme partial digestion.

We present a theoretical study of the fraction of sequences incorporated in a recombinant DNA partial digest library as a function of the size of the library. The fraction incorporated depends on the degree of restriction enzyme partial digestion. If all restriction sites in the target DNA can be cleaved with the same rate, optimum incorporation of sequences is observed when the number average length of the digested DNA equals the desired average length of the cloned insert. Overdigestion severely reduces the fraction of sequences present in a sample of clones. Heterogeneity in restriction enzyme cleavage rates also reduces the fraction incorporated, and underdigestion improves sequence representation in the face of cleavage rate heterogeneity. Practical methods for determining the number average length of partially digested DNAs are also presented.

Base Composition↗

Gene discovery and expression profiling in porcine Peyer's patch.

Peyer's patches of the intestinal mucosa are essential for host defense and immune regulation in the enteric system. To better understand molecular mechanisms of Peyer's patch function, we have screened for differentially expressed genes specific to Peyer's patch. cDNA libraries were created from normal Peyer's patch, immune stimulated Peyer's patch, and pooled cDNA subtracted with fibroblast RNA. From the subtracted library, 3687 expressed sequence tags (ESTs), representing 2414 unique nucleotide sequences, were isolated, identified by BLAST searches against public databases, and spotted onto a microarray for gene expression profiling. Approximately 30% of these ESTs BLAST to genes of unknown function and 20% have no known homology in the public databases (novel genes). Of the novel genes, 70% are expressed in normal immune tissues by microarray analysis, suggesting that at least 371 of the unidentified EST sequences from the subtracted library are novel porcine genes and can now be further characterized to determine their function in the porcine Peyer's patch. We surmise that the products of these genes participate in biochemical and cellular functions related to the unique immunological and gastroenterological functions of the small intestine. The BLAST and gene ontology information for each of the subtracted library EST sequences, the normal and immune stimulated libraries, and the microarray are all valuable resources that will facilitate further examination of the biological function of porcine Peyer's patch tissue.

Animals↗

Utilisation of bacteriophage display libraries to identify peptide sequences recognised by equine herpesvirus type 1 specific equine sera.

Three filamentous phage random peptide display libraries were used in biopanning experiments with purified IgG from the serum of a gnotobiotic foal infected with equine herpesvirus-1 (EHV-1) to enrich for epitopes binding to anti-EHV-1 antibodies. The sequences of the amino acids displayed were aligned with protein sequences of EHV-1, thereby identifying a number of potential antibody binding regions. Presumptive epitopes were identified within the proteins encoded by genes 7 (DNA helicase/primase complex protein), 11 (tegument protein), 16 (glycoprotein C), 41 (integral membrane protein), 70 (glycoprotein G), 71 (envelope glycoprotein gp300), and 74 (glycoprotein E). Two groups of sequences, which aligned with either glycoprotein C (gC) or glycoprotein E (gE), identified type-specific epitopes which could be used to distinguish between sera from horses infected with either EHV-1 or EHV-4 in an ELISA using either the phage displaying the peptide or synthetic peptides as antigen. The gC epitope had been previously identified as an immunogenic region by conventional monoclonal antibody screening whereas the gE antibody binding region had not been previously identified. This demonstrates that screening of phage display peptide libraries with post-infection polyclonal sera is a suitable method for identifying diagnostic antigens for viral infections such as EHV-1.

Amino Acid Sequence↗

Mouse CD-RAP/MIA gene: structure, chromosomal localization, and expression in cartilage and chondrosarcoma.

A cDNA encoding a novel protein has been previously isolated from two independent sources: melanoma cell cultures and chondrocytes. The protein from human melanoma cell lines and tumors is called melanoma inhibitory activity (MIA) (Blesch et al. [1994] Cancer Res. 54:5695-5701) and the protein from primary bovine chondrocytes and cartilaginous tissues is called cartilage-derived retinoic acid-sensitive protein (CD-RAP) (Dietz and Sandell [1996] J. Biol. Chem. 271:3311-3316). In order to investigate the gene regulation and function of CD-RAP/MIA, the mouse gene locus was isolated and analyzed. Developmental expression was determined by in situ hybridization to mouse embryos. Expression was limited to cartilaginous tissues and was initiated with the advent of chondrogenesis, remaining abundant throughout development. The mouse gene was isolated and sequenced from a 129Sv library and sequenced directly from an additional strain, B6C3Fe. The mouse CD-RAP/MIA gene is 1.5 kbp and consists of four exons. The promoter sequence of the gene contains many potential regulatory domains including 8 basic helix-loop-helix protein-binding domains and an AT-rich domain, both motifs shown to be present in the cartilage-specific enhancer of the type II procollagen gene. Other potential cis-acting motifs include binding sites for GATA-1, NF-IL6, PEA3, w-elements, NF kappa B, Zeste and Sp1. The gene, called cdrap, was localized to the end of an arm of chromosome 7 at the same site as the transforming growth factor beta 1 (Tgf-beta 1) and the glucose phosphate isomerase 1 (Gpi 1) genes. Potential mouse mutants that mapped to the same region of chromosome 7 were identified. Two of the potential mutants with skeletal phenotypes were sequenced, pudgy (pu) and extra toes with spotting (XsJ); however, no mutations were found in the coding sequence. To determine whether CD-RAP/MIA is associated with tumors of cartilage, mRNAs from a variety of rodent tissues and cell lines were screened. Expression was detected in a rodent tumor, the Swarm rat chondrosarcoma and a chondrosarcoma cell line derived from it, but not in other tissues or tumors of non-cartilage origin. Immunolocalization revealed CD-RAP/MIA protein localized in cartilage only. These results show that the normal expression of CD-RAP/MIA is limited to cartilage; however, pathologically, it is expressed both in melanoma and chondrosarcoma. The restricted expression of CD-RAP/MIA may provide an opportunity to monitor cartilage metabolic activity as well as the tumor activity of melanoma and chondrosarcoma.

Amino Acid Sequence↗

The chemokine information source: identification and characterization of novel chemokines using the WorldWideWeb and expressed sequence tag databases.

The chemokine superfamily is a large group of more than 30 small proteins. Many of these were originally identified because of their role in the selective recruitment and activation of leukocytes during inflammation. More recently, some of the chemokine receptors and ligands have been implicated in the mechanism of viral infection for primate lentiviruses such as HIV-1. From the original identification of interleukin-8 (IL-8; the most studied member of the superfamily), the number of new family members has mushroomed over the last few years. Two events have dramatically altered the speed at which sequence information concerning novel chemokines has become available to the scientific community. First, many groups have been obtaining large amounts of sequence information from cDNA libraries by sequencing the clones at random, generating expressed sequence tags (ESTs). Although these ESTs are relatively short, typically less than 500 bases, this amount of sequence is usually sufficient to obtain the entire open reading frame for chemokines. Second, there has been a rapid growth in the use of the WorldWideWeb by bioinformatics groups. The Web was originally set up by the European Centre for Particle Physics (CERN) in Geneva as a method of transferring data between collaborating groups throughout the world. It has enabled biologists throughout the world to have almost instantaneous access both to the databases containing the EST sequences and to the automated tools that are required for searching such databases. With such methods we have been able to rapidly identify more than 10 new human chemokines from public domain databases. In addition to the known categories of chemokines, which are named C, CC, and CXC based on the spacings of N-terminal cysteine residues, we have been able to identify the first member of a novel chemokine subfamily, with a novel CXXXC cysteine spacing. Furthermore, we can subdivide the CC chemokines into monocyte chemotactic protein and macrophage inflammatory protein families based on their sequence identity levels, but also their clustering on the human genome, as identified on other Web sites. The rapid availability of all this data has reduced the amount of time spent on conventional gene identification, enabling us to move quickly on to trying to understand the biology and physiological relevance of these molecules. The novel chemokine sequences obtained and alignments with existing members of the superfamily are now contained within a Chemokine Information Source on an open access server, allowing further searching of chemokine sequences and increasing the availability of such data to the scientific community.

Amino Acid Sequence↗

Isolation and characterization of vascular smooth muscle cell growth promoting factor from bovine ovarian follicular fluid and its cDNA cloning from bovine and human ovary.

A protein possessing vascular smooth muscle cell (SMC) growth-promoting activity (VSGP) was purified from bovine ovarian follicular fluid. The purified protein showed a broad band on SDS-PAGE with an apparent molecular mass of 90-100 kDa. The purified protein was characterized by amino acid sequence analysis of its N-terminal and internal peptides. Based on the information of the peptide sequences, bovine ovarian cDNA library was screened and cDNA clones encoding the protein were isolated. Human homolog of the protein was also cloned from human ovarian cDNA library. Nucleotide sequence analysis revealed that bovine VSGP transcript has a 2421-bp open reading frame, which encodes a protein of 807 amino acid residues. A homology search indicated that bovine and human VSGP are counterparts of rat F-spondin, which has been previously identified as a promoter molecule of neurite extension in rat fetal floor plate. RNA blot analysis showed wide distribution of VSGP/F-spondin transcripts in fetal and adult human tissues. Especially the expression was highest in the adult human ovary. The purified bovine VSGP/F-spondin showed vascular SMC growth promoting activity with an ED(50) value of 10(-8) M. Together with these findings, we demonstrated here that VSGP/F-spondin is a major factor for vascular SMC proliferation in the ovary. In conclusion, our present study provides a distinct and important function of VSGP/F-spondin as a strong VSMC proliferation promoting factor, in addition to the previously proposed function in neuronal system, and also provides insight into mechanisms underlying vascular SMC proliferation during ovarian folliculogenesis.

Amino Acid Sequence↗

Cloning and functional characterization of Phaeodactylum tricornutum front-end desaturases involved in eicosapentaenoic acid biosynthesis.

Phaeodactylum tricornutum is an unicellular silica-less diatom in which eicosapentaenoic acid accumulates up to 30% of the total fatty acids. This marine diatom was used for cloning genes encoding fatty acid desaturases involved in eicosapentaenoic acid biosynthesis. Using a combination of PCR, mass sequencing and library screening, the coding sequences of two desaturases were identified. Both protein sequences contained a cytochrome b5 domain fused to the N-terminus and the three histidine clusters common to all front-end fatty acid desaturases. The full length clones were expressed in Saccharomyces cerevisiae and characterized as Delta5- and Delta6-fatty acid desaturases. The substrate specificity of each enzyme was determined and confirmed their involvement in eicosapentaenoic acid biosynthesis. Using both desaturases in combination with the Delta6-specific elongase from Physcomitrella patens, the biosynthetic pathways of arachidonic and eicosapentaenoic acid were reconstituted in yeast. These reconstitutions indicated that these two desaturases functioned in the omega3- and omega6-pathways, in good agreement with both routes coexisting in Phaeodactylum tricornutum. Interestingly, when the substrate selectivity of each enzyme was determined, both desaturases converted the omega3- and omega6-fatty acids with similar efficiencies, indicating that none of them was specific for either the omega3- or the omega6-pathway. To our knowledge, this is the first report describing the isolation and biochemical characterization of fatty acid desaturases from diatoms.

Amino Acid Sequence↗

A land plant-specific multigene family in the unicellular Mesostigma argues for its close relationship to Streptophyta.

The search for the unicellular relative of Streptophyta (i.e., land plants and their closest green algal relatives, the charophytes) started many years ago and remained centered around the scaly green flagellate, Mesostigma viride. To date, despite numerous studies, the phylogenetic position of Mesostigma is still debated and the nature of the unicellular ancestor of Streptophyta remains unknown. As molecular phylogenetic studies have produced conflicting results, we constructed a M. viride expressed sequence tags library and searched for sequences that are shared between M. viride and the Streptophyta (to the exclusion of the other green algal lineages--the Chlorophyta). Here, we report a multigene family that is restricted to Streptophyta and M. viride. The phylogenetic distribution of this complex character and its potential involvement in the evolution of an important land plant adaptive trait (i.e., three-dimensional tissues) argue that Mesostigma is a close unicellular relative of Streptophyta.

Algal Proteins↗

Compilation of 5S rRNA and 5S rRNA gene sequences.

The BERLIN RNA DATABANK as of December 31, 1987, contains a total of 509 sequences of 5S rRNAs or their genes, which is an increase of 45% over the last (1986) compilation (1). It covers sequences from 38 archaebacteria, 184 eubacteria, 14 plastids, 4 mitochondria, 258 eukaryotes and 11 eukaryotic pseudogenes. The BERLIN RNA DATABANK uses the format of the EMBL Nucleotide Sequence Data Library complemented by a Sequence Alignment (SA) field including secondary structure information as presented in this publication. The BERLIN RNA DATABANK is available on 360 or 1200 kb diskettes.

Animals↗

Compilation of 5S rRNA and 5S rRNA gene sequences.

The BERLIN RNA DATABANK as of December 31, 1989, contains a total of 667 sequences of 5S rRNAs or their genes, which is an increase of 114 new sequence entries over the last compilation (1). It covers sequences from 44 archaebacteria, 267 eubacteria, 20 plastids, 6 mitochondria, 319 eukaryotes and 11 eukaryotic pseudogenes. The hardcopy shows only the list of those organisms whose sequences have been determined. The BERLIN RNA DATABANK uses the format of the EMBL Nucleotide Sequence Data Library complemented by a Sequence Alignment (SA) field including secondary structure information.

Animals↗

Compilation of 5S rRNA and 5S rRNA gene sequences.

This is an update for the 5S rRNA sequences of the BERLIN RNA DATABANK last published in 1990 (1). The new entry consists of 25 eubacterial and 2 eukaryotic 5S rRNA sequences and 10 plant 5S rRNA pseudogenes (Table 1). Thus the BERLIN RNA DATABANK contains as of February 1, 1991 the 5S rRNA sequences of 44 archaebacteria, 292 eubacteria, 20 plastids, 6 mitochondria, 321 eukaryotes and 21 eukaryotic pseudogenes. The BERLIN RNA DATABANK uses the format of the EMBL Nucleotide Sequence Data Library complemented by a Sequence Alignment (SA) field including secondary structure information.

Bacteria↗

Compilation of 5S rRNA and 5S rRNA gene sequences.

The compilation of 5S rRNA and 5S rRNA gene nucleotide sequences as of 30 September 1996, contains a total of 1661 primary structures of 5S rRNAs or their genes, which is an increase of 928 new sequence entries over the last compilation. It covers sequences from 54 archaea, 449 eubacteria, 34 plastids, nine mitochondria and 430 eukaryotes. The databank uses the format of the EMBL Nucleotide Sequence Data Library complemented by a Sequence Alignment (SA) field including secondary structure information. The taxonomic classification of organisms was totally updated. Now the database is also available via anonymous FTP or WWW.

Base Sequence↗

Cloning of the authentic bovine gene encoding pepsinogen a and its expression in microbial cells.

Bovine pepsin is the second major proteolytic activity of rennet obtained from young calves and is the main protease when it is extracted from adult animals, and it is well recognized that the proteolytic specificity of this enzyme improves the sensory properties of cheese during maturation. Pepsin is synthesized as an inactive precursor, pepsinogen, which is autocatalytically activated at the pH of calf abomasum. A cDNA coding for bovine pepsin was assembled by fusing the cDNA fragments from two different bovine expressed sequence tag libraries to synthetic DNA sequences based on the previously described N-terminal sequence of pepsinogen. The sequence of this cDNA clearly differs from the previously described partial bovine pepsinogen sequences, which actually are rabbit pepsinogen sequences. By cloning this cDNA in different vectors we produced functional bovine pepsinogen in Escherichia coli and Saccharomyces cerevisiae. The recombinant pepsinogen is activated by low pH, and the resulting mature pepsin has milk-clotting activity. Moreover, the mature enzyme generates digestion profiles with alpha-, beta-, or kappa-casein indistinguishable from those obtained with a natural pepsin preparation. The potential applications of this recombinant enzyme include cheese making and bioactive peptide production. One remarkable advantage of the recombinant enzyme for food applications is that there is no risk of transmission of bovine spongiform encephalopathy.

Amino Acid Sequence↗