PubMed HealthSearch

SEARCH · PubMed Health

Results for “sequencing libraries”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Characterization of the short isoform of the growth hormone receptor synthesized by rat adipocytes.

Two mRNA transcripts that are believed to be alternately spliced products of the GH receptor gene have been reported in a variety of rat tissues. The smaller (1.2 kilobases) transcript was cloned from an adipocyte library, sequenced, and found to encode a protein identical to the soluble GH-binding protein (GHBP) in plasma. An assay that is specific for the short isoform of the GH receptor, often referred to as the GHBP, has been developed using a rabbit antiserum that recognizes the unique amino acid sequence at its carboxyl end. The assay depends upon immunoprecipitation of a complex consisting of [125I]human GH, the binding protein, antiserum, and protein-A cross-linked to agarose beads. To validate the assay, samples of rat plasma were analyzed and found to contain sufficient binding protein to bind 1.46 pmol (32 ng) GH/ml, with an affinity of 2.7 x 10(9) M-1. In adipocyte extracts, binding protein activity was sufficient to bind 61 fmol GH/g tissue, with an affinity of 2.3 x 10(9) M-1. The binding protein was found primarily in the particulate fraction of adipocytes, and it is estimated that adipocytes contain approximately 7000 copies of the binding protein/cell. Only 10% of the binding activity was present in the high speed supernatant of adipocyte homogenates, and soluble binding protein did not appear to be released into the incubation medium when adipocytes were incubated in vitro. A 50-kilodalton (kDa) 35S-labeled protein that may be a glycosylated form of the binding protein was immunoprecipitated from both the soluble and particulate fractions of adipocyte extracts by the antiserum, and addition of the synthetic peptide antigen blocked immunoprecipitation of this protein. A 150-kDa protein in the high speed supernatant fraction was also specifically immunoprecipitated by the antiserum. Although it is unlikely to be a glycosylated form of the binding protein, it may cross-react with the antiserum or perhaps be coprecipitated, because it interacts with the binding protein. In addition, 38- and 42-kDa bands were specifically immunoprecipitated from the detergent-treated particulate fraction of adipocyte extracts that were enriched for the binding protein by adsorption to immobilized GH. We conclude that 1) adipocytes synthesize the short isoform of the GH receptor, and that this protein is primarily associated with a membrane fraction of the cells; and 2) the GHBP expressed in adipocytes is not released into the incubation medium and differs in size from the GHBPs in rat plasma.(ABSTRACT TRUNCATED AT 400 WORDS)

Adipose Tissue

Uncovering the diagnostic potential of seminal fluid beyond fertility: cfDNA methylation analysis for the detection of clinically significant prostate cancer.

Research on the potential use of seminal fluid as a liquid biopsy for prostate cancer detection has been limited due to challenges associated with acquisition of this bodily fluid in clinical studies. Here we sought to expand on our previous analysis, which demonstrated high levels of prostate-derived cell free DNA (cfDNA) in seminal fluid in presumed healthy individuals, to a much larger cohort that included participants with prostate cancer. A total of 279 men scheduled for prostate biopsy were enrolled over 4 months across 12 sites. Prior to their biopsy, participants mailed a seminal fluid sample collected at home to the laboratory, from which cfDNA was extracted and underwent methylation analysis. Consistent with our earlier study in healthy individuals, we observed an abundance of high molecular weight (HMW) cfDNA in all samples. Tissue-of-origin deconvolution revealed that granulocytes and sperm were the principal contributors to seminal fluid cfDNA, while prostate-derived cfDNA was present at abundances readily detectable with current technologies. The nucleosomal fraction was very pronounced in some but not all samples and was determined to be correlated with the relative sperm signal. The sperm signal was also observed to be associated with an increase in small insert sizes (< 125 bp) in the sequenced libraries. Unsupervised clustering revealed two distinct populations driven by the abundance of sperm and granulocytes. Since summarizing at the genomic region level confounded tissues of different origins, fragment-level DNA methylation features were used to characterize and quantify the prostate cancer related signal, and features associated with clinically significant prostate cancer were identified. This study expands on our previous work to further characterize seminal fluid and highlights its potential as a promising liquid biopsy medium for the detection and monitoring of clinically significant prostate cancer.

Humans

Molecular cloning and expression of human trophoblast antigen FDO161G and its identification as 3 beta-hydroxy-5-ene steroid dehydrogenase.

The monoclonal antibody FDO161G reacts with a 43-kDa protein found in human extravillous trophoblast, syncytiotrophoblast, adrenal cortex, interstitial cells of the testis and ovarian follicle cumulus cells. cDNAs for this protein have been isolated from the lambda gt11 library, sequenced, and expressed in COS-7 cells. The protein was identified as 3 beta-hydroxy-5-ene steroid dehydrogenase (HSD). The sequence of the HSD protein raises questions about its association with cell membrane systems. The lack of reactivity of FDO161G with other tissues suggests that HSD has a limited tissue distribution and that other enzymes may exist in peripheral tissues, which can convert delta 5 3-hydroxysteroids to delta 4 3-ketosteroids.

3-Hydroxysteroid Dehydrogenases

Characterization of multiple cathepsin B mRNAs in murine B16a melanoma.

We have previously shown that the highly metastatic murine B16a melanoma expresses a high level of cathepsin B mRNA which is associated with three transcripts of 2.2, 4.0 and 5.0 kb, while in contrast only a single 2.2 kb cathepsin B RNA was detected in normal murine tissues. Using recombinant DNA techniques, cDNAs corresponding to these three transcripts have been isolated from a B16a melanoma cDNA library. Sequence analysis indicates that all three mRNA transcripts contain identical coding sequences for normal preprocathepsin B. However, the 4.0 and 5.0 kb transcripts contain unusually long extended 3' untranslated regions. These results suggest that the post-transcriptional processing pathway of the cathepsin B gene is modified in B16 melanomas. The results also indicate that the increased extracellular secretion of larger forms of cathepsin B by tumors is most likely due to post-translational mechanisms and does not involve alternative splicing or a coding mutation in the gene.

Amino Acid Sequence

A gene specifying subunit VIII of human cytochrome c oxidase is localized to chromosome 11 and is expressed in both muscle and non-muscle tissues.

Subunit VIII of mammalian cytochrome c oxidase (COX; EC 1.9.3.1) exists in at least two isoforms, because different but related polypeptides have been identified in COX isolated from liver and heart of both beef and pig. We have isolated a full length cDNA specifying subunit VIII of human COX from a human liver cDNA library. Sequences hybridizing to this cDNA are present at only one site, the COX8 locus, on human chromosome 11q12-q13. The deduced human polypeptide is 58% identical with COX VIII isolated from beef liver, but only 38% identical with COX VIII isolated from beef heart. Transcriptional analysis shows that an mRNA identical with the isolated cDNA is present in abundant amounts not only in human and monkey liver tissue, but in heart and skeletal muscle as well, tissues not known previously to contain this isoform. Since the only COX VIII subunit found in human heart agrees 100% with the polypeptide deduced from this coxVIII cDNA, it may well be that, in distinction to other mammals, only one form of COX VIII exists in primates.

Amino Acid Sequence

Osteogenesis imperfecta type IV. Detection of a point mutation in one alpha 1(I) collagen allele (COL1A1) by RNA/RNA hybrid analysis.

We have identified a point mutation in one alpha 1(I) collagen allele (COL1A1) of a child with the type IV osteogenesis imperfecta phenotype. When compared to parental and control samples, skin fibroblasts of the proband synthesized two populations of type I collagen molecules. One population was normal; the other was delayed in secretion and electrophoretic migration due to post-translational overmodification. Two-dimensional gel electrophoresis of the CNBr peptides demonstrated a gradient of overmodification beginning near the carboxyl-terminal CB peptides. This predicts that the mutation delaying helix formation is near the carboxyl-terminal end of one of the component chains of type I collagen. The mRNA of the patient was probed with overlapping antisense riboprobes to type I collagen cDNA. Cleavage of a mismatch in RNA/RNA hybrids of RNase A allowed the location of the mutation to a 225-base pair region of alpha 1(I) cDNA. The mismatch was not present in RNA/RNA hybrids from either parent. This region of both alpha 1(I) alleles of the patient was isolated by screening a lambda ZAP cDNA library. Sequence determination of both alleles demonstrated a single nucleotide change, G----A, resulting in the substitution of a serine for a glycine at amino acid residue 832. This point mutation occurs in the coding region for alpha 1(I) CB6 and is concordant with the protein data. The finding of a glycine substitution in an alpha 1(I) chain of a patient with the milder type IV osteogenesis imperfecta phenotype requires modification of current molecular models for types II and IV osteogenesis imperfecta.

Alleles

Cloning and analysis of cDNA clones for rat kidney alpha-spectrin.

We have isolated a 3922-base pair (bp) cDNA clone for rat nonerythroid alpha-spectrin from a rat kidney lambda gt11 cDNA library. Sequence analysis revealed that this cDNA contains an open reading frame of 3090 bp encoding for the C-terminal 1030 amino acid sequence of rat kidney alpha-spectrin. The 3'-untranslated region (including a 38-bp poly(A+) tail) contains an 832-bp sequence. A single mRNA of about 8 kilobase pairs was detected in rat liver, kidney, brain, heart, intestine, lung, testis, stomach, spleen, and muscle with varying abundances, which is consistent with and further confirms the presence of spectrins in nonerythroid tissues as demonstrated previously by immunoblot analysis. Southern blot analysis suggested that there is a single gene for nonerythroid alpha-spectrin. The derived amino acid sequence contains sequence from the spectrin 106-residue internal repeat 12 to the C terminus of rat kidney alpha-spectrin. Sequence comparison with human and chicken nonerythroid alpha-spectrin showed that nonerythroid alpha-spectrin is well conserved during evolution. The rat kidney alpha-spectrin sequence, when compared to rat brain alpha-spectrin, contains an extra 76-amino-acid sequence at the C terminus. Sequence comparison of all the internal repeats available revealed that the internal repeat 3, 4, 5, 6, 7, and 8 has highest sequence similarity with internal repeat 12, 13, 14, 15, 16, and 17, respectively. Therefore, internal repeats 3-8 and 12-17 are most likely derived from an ancestral gene through gene duplication, suggesting that the spectrin gene is derived from a half-spectrin gene by gene duplication and divergence during evolution.

Amino Acid Sequence

Cross-resistance and glutathione-S-transferase-pi levels among four human melanoma cell lines selected for alkylating agent resistance.

A panel of four cell sublines, each selected for resistance to a different antineoplastic agent, has been developed from a human malignant melanoma cell line G3361. Following repeated exposure to escalating doses of the drug of interest, cloned sublines were developed that are 9-fold resistant to cisplatin (G3361/CP), 11-fold resistant to 4-hydroxyperoxy-cyclophosphamide (4-HC) (G3361/HC), 4-fold resistant to carmustine (BCNU) (G3361/BCNU), and 4-fold resistant to melphalan (G3361/PAM). The cross-resistance of each resistant cell line was determined for cisplatin, BCNU, 4-HC, melphalan, carboplatin, nitrogen mustard, and Adriamycin. In general, the alkylating agent-resistant cell lines were specifically resistant to the drug used for selection with the exception of the G3361/CP line, which was greater than 10-fold resistant to the cisplatin analogue carboplatin, 4-fold resistant to 4-HC, and slightly (1.5-fold) resistant to melphalan, and the G3361/BCNU line, which was slightly (1.8-fold) resistant to melphalan. Collateral sensitivity of the G3361/CP, G3361/PAM, and G3361/4HC lines to killing by BCNU was also observed. Glutathione-S-transferase activity was elevated in each of the alkylating agent-resistant cell lines by 3- to 5-fold with chlorodinitrobenzene substrate. On Western blotting, the glutathione-S-transferase-pi (GST-pi) isoenzyme protein was elevated in the resistant cells by 3- to 5-fold. A complementary DNA (pTS4-10) coding for GST-pi has been cloned from a lambda gt11 library, sequenced, and used as a probe to determine the relative levels of GST-pi mRNA in the alkylating agent-resistant cell lines. GST-pi mRNA levels were elevated (8- to 15-fold) in the resistant cell lines, indicating that the GST-pi increases were mediated through an increase in mRNA levels. GST-pi elevations are a frequent event in cells selected for alkylating agent resistance, and in some cases, of multiple drug resistance. However, the lack of cross-resistance among cell lines selected for resistance to different alkylating agents, all of which have elevated GST-pi levels, indicates that increased levels of GST-pi cannot be the predominate mechanism for resistance to the tested drugs in these cell lines.

Alkylating Agents

Structure of the gene for mouse thymidylate synthase. Locations of introns and multiple transcriptional start sites.

We have isolated and analyzed the structure of the gene for thymidylate synthase from a 5-fluoro-2'-deoxyuridine-resistant 3T6 cell line that overproduces thymidylate synthase 50-fold by virtue of gene amplification. Three overlapping DNA segments containing the entire thymidylate synthase gene were identified in Charon 35 genomic libraries. Sequence analysis revealed that all of the coding regions were contained in a 12-kilobase segment of DNA. The gene has 6 introns ranging from 0.6 to 4.1 kilobases in length. The sequences at the 5' and 3' ends of each intron conformed to the consensus sequences for mammalian introns. S1 nuclease and primer extension assays showed that transcription of the thymidylate synthase gene initiates at several sites within a 66-nucleotide region. There are no TATAA or CCAAT sequences in the vicinity of the initiation sites. However, the region does contain DNA sequences that are known to stimulate binding of the transcription factors Sp1 and USF. These binding sites are adjacent to each other, suggesting that the two proteins may bind to the upstream region of the thymidylate synthase gene in a cooperative or competitive manner.

Animals

PROMOT: a FORTRAN program to scan protein sequences against a library of known motifs.

Information about the three-dimensional structure or function of a newly determined protein sequence can be obtained if the protein is found to contain a characterized motif or pattern of residues. Recently a database (PROSITE) has been established that contains 337 known motifs encoded as a list of allowed residue types at specific positions along the sequence. PROMOT is a FORTRAN computer program that takes a protein sequence and examines if it contains any of the motifs in PROSITE. The program also extends the definitions of patterns beyond those used in PROSITE to provide a simple, yet flexible, method to scan either a PROSITE or a user-defined pattern against a protein sequence database.

Amino Acid Sequence

Expressed sequence tags and chromosomal localization of cDNA clones from a subtracted retinal pigment epithelium library.

Expressed sequence tags (ESTs) provide useful molecular landmarks for physical mapping and identify the position of an expressed region in the genome. The use of subtracted cDNA libraries enriched for tissue-specific genes as a source of ESTs should reduce the repetitive isolation of constitutively expressed sequences. We report here the sequence tags from the 3'-end region of 58 new directionally cloned cDNAs from a subtracted human retinal pigment epithelium (RPE) cell line library. Eight of the cDNAs have been assigned to human chromosomes using PCR-based EST assays. Chromosomal mapping of subtracted RPE cDNA clones may also help in identifying candidate genes for inherited eye diseases.

Base Sequence

[Isolation and analysis of brain-specific sequences from cDNA libraries for various segments of the human brain].

The cDNA libraries in gt10 were constructed from total poly(A)+RNA of human forebrain cortex, cerebellar cortex and medulla oblongata. We selected the clones which gave hybridization signal with brain cDNA only, or gave no signal from these libraries. Expression pattern and structure of two brain-specific clones Hfb1 from forebrain library and Hmob3 from medulla oblongata library were analyzed in detail. Hfb1 hybridized to two different transcripts (about 5 and 2 kb) from frontal cortex, but to a single (longest) from cerebellum. Hfb1 sequence includes 958 nucleotides. Comparison of Hfb1 with the Gene Bank revealed no homology with the sequences present in the Bank. At 3'-end there is poly(A) tail of 24 bases, there is the AATCAA sequence 55 nucleotides upstream which probably serves as a polyadenylation signal. However, AATCAA directs polyadenylation in vitro with very low efficiency. We found no open reading frame in the clone and this is in agreement with the data indicating that brain-specific RNAs has extremely long 3'-untranslated regions. Hmob3 was partially sequences. We compared its primary structure with the sequences from the Gene Bank and revealed no homology. Hmob3 expresses in different parts of human brain and in sceletal muscle but does not express in other tissues.

Base Sequence

DeepGeSeq: deep learning library for genomic sequence modeling and analysis.

MOTIVATION: Deep learning methods have demonstrated significant potential in genomics, enabling broad applications such as sequence activity prediction, regulatory rule identification, and variant effect quantification. However, their widespread adoption is often hindered by the steep computational learning curve required for model construction, training, and downstream biological interpretation. Here, we introduce DeepGeSeq, a user-friendly Deep-learning library tailored for Genomic Sequence modeling and analysis. RESULTS: By integrating state-of-the-art architectural modules, DeepGeSeq streamlines the entire deep learning workflow, requiring minimal user input via a simple configuration file and an intuitive agentic skill. We comprehensively validate the efficacy of DeepGeSeq through diverse case studies, encompassing pipeline verification using synthetic datasets, the reproduction and application of established models, and model fine-tuning coupled with biological interpretation on user-defined data. Furthermore, we demonstrate DeepGeSeq's versatility in domain-specific applications, including single-cell ATAC-seq modeling for cell-type clustering, and MPRA data modeling coupled with in silico saturation mutagenesis to dissect cis-regulatory elements. Ultimately, DeepGeSeq bridges the gap between computational complexity and biological discovery, providing an accessible resource that facilitates the development and broad application of deep learning methods in genomics research. AVAILABILITY AND IMPLEMENTATION: https://github.com/JiaqiLi1024/DeepGeSeq.

Deep Learning

Cloning of genes encoding redox proteins of known amino acid sequence from a library of the Desulfovibrio vulgaris (Hildenborough) genome.

A library of 900 recombinant phages has been constructed for the genome of Desulfovibrio vulgaris Hildenborough (1.7 x 10(6) bp) by cloning size-fractionated Sau3A fragments (15-20 kb) into the replacement vector lambda-2001. When a hydrogenase gene probe, a 4.7-kb SalI-EcoRI fragment of known nucleotide sequence, was used to screen the plaque lifted library, 23 positive clones were found, which together span 31 kb of D. vulgaris DNA. To facilitate the cloning of genes with oligodeoxynucleotides as probes, DNA was purified for all clones in the library and spotted on a 16 x 16-cm grid of nitrocellulose. This grid was incubated sequentially to identify lambda clones containing the gene for redox proteins of known amino acid sequence: cytochrome c3 (one 18-mer----four clones), flavodoxin (one 17-mer and one 26-mer----one clone) and rubredoxin (one 44-mer----21 clones). The four cyc-positive clones are also recognized by the rubredoxin oligodeoxynucleotide probe. Restriction mapping defines a 35-kb region of the D. vulgaris chromosome in which the rub and cyc loci are separated by 17.5 kb. The nucleotide sequence of the rubredoxin gene was determined and the deduced amino acid sequence found to agree with that determined in Bruschi [Biochim. Biophys. Acta 434 (1976) 4-17] with the exception of Thr-21 which is found to be encoded by GAC, an Asp codon. A plausible ribosome-binding site precedes the N-terminal initiator methionine residue. Rubredoxin does not have an N-terminal signal sequence which is in agreement with the cytoplasmic location of this redox protein.

Amino Acid Sequence

The occurrence of families of repetitive sequences in a library of cloned cDNA from human lymphocytes.

A library of cloned cDNAs representative of lymphocyte total poly(A)+ RNA was screened with total DNA probes at high clone density. 10% of the recombinants showed the presence of sequences which are repeated in the genome. Further analysis of six such isolated cDNA clones indicated that they contain different families of repetitive sequences with reiteration frequencies of between 150 and 45,000 copies per haploid genomes. Five of the six clones were found to contain single copy sequences as well as a repetitive sequence. cDNA clones containing repetitive sequences have been found to be derived from high, intermediate and low abundance classes of lymphocyte poly(A)+ RNA.

Cloning, Molecular

Dynamic programming algorithms for biological sequence comparison.

Efficient dynamic programming algorithms are available for a broad class of protein and DNA sequence comparison problems. These algorithms require computer time proportional to the product of the lengths of the two sequences being compared [O(N2)] but require memory space proportional only to the sum of these lengths [O(N)]. Although the requirement for O(N2) time limits use of the algorithms to the largest computers when searching protein and DNA sequence databases, many other applications of these algorithms, such as calculation of distances for evolutionary trees and comparison of a new sequence to a library of sequence profiles, are well within the capabilities of desktop computers. In particular, the results of library searches with rapid searching programs, such as FASTA or BLAST, should be confirmed by performing a rigorous optimal alignment. Whereas rapid methods do not overlook significant sequence similarities, FASTA limits the number of gaps that can be inserted into an alignment, so that a rigorous alignment may extend the alignment substantially in some cases. BLAST does not allow gaps in the local regions that it reports; a calculation that allows gaps is very likely to extend the alignment substantially. Although a Monte Carlo evaluation of the statistical significance of a similarity score with a rigorous algorithm is much slower than the heuristic approach used by the RDF2 program, the dynamic programming approach should take less than 1 hr on a 386-based PC or desktop Unix workstation. For descriptive purposes, we have limited our discussion to methods for calculating similarity scores and distances that use gap penalties of the form g = rk. Nevertheless, programs for the more general case (g = q+rk) are readily available. Versions of these programs that run either on Unix workstations, IBM-PC class computers, or the Macintosh can be obtained from either of the authors.

Algorithms

Bovine beta-crystallin complementary DNA clones. Alternating proline/alanine sequence of beta B1 subunit originates from a repetitive DNA sequence.

A library of recombinant plasmids carrying complementary DNA sequences synthesized from bovine lens messenger RNAs was constructed. Clones coding for five different beta-crystallin subunits: beta B1, beta B3, beta Bp, beta s, beta A3 (and beta A1), were identified by means of hybridization selection, followed by one- and two-dimensional gel electrophoresis of the translational products. Under rather stringent conditions each of these clones hybridizes with its corresponding mRNA and does not show significant cross-hybridization with mRNAs coding for other beta-crystallins, except in the case of the homologous beta A3 and beta A1-crystallins. The beta A3 and beta A1 subunits seem to be encoded by one mRNA using two different AUG codons as start position for translation. We have also determined the nucleotide sequence of a beta B1-crystallin cDNA (pBL beta B1) which enabled us to deduce the complete amino acid sequence of the protein. The beta B1-crystallin, a characteristic component of the high molecular weight crystallin aggregate (beta H), is internally homologous both at DNA and protein level as has been reported for gamma- and other beta-crystallins. This is in agreement with the idea that these proteins had a common ancestral precursor gene that internally duplicated. The G + C content of the coding sequence of beta B1 is very high: 67% overall and even 84.2% for the first 170 nucleotides, due to a remarkable non-random codon usage. A proline/alanine repetition in the N-terminal domain of the protein is encoded by a repetitive "simple" DNA sequence.

Alanine