PubMed HealthSearch

SEARCH · PubMed Health

Results for “sequencing libraries”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Cloning and sequencing of a COUP transcription factor gene expressed in Xenopus embryos.

A cDNA clone encoding COUP transcription factor, a member of the steroid/thyroid receptor superfamily, has been isolated from a Xenopus neurula (stage 17 embryo) library. Sequencing of this clone reveals an open reading frame encoding a 397 amino acid protein. The amino acid sequence of Xenopus COUP has been compared with its human and Drosophila homologues showing that there are few similarities within the amino-terminal region, whereas the remainder of the protein, including the putative DNA and ligand binding domains, is very well conserved.

Amino Acid Sequence

Molecular cloning, sequence analysis and translation of proenkephalin mRNA from rat heart.

Proenkephalin mRNA is abundant in rat cardiac ventricles but surprisingly low levels of opioid peptides or precursor forms derived from proenkephalin are present in tissue extracts. Proenkephalin mRNA in rat heart was characterized at the molecular level with the use of cDNA sequencing, in vitro translation, and primer extension. Two positive proenkephalin cDNA clones were obtained by screening approx. 20,000 recombinant phages from a heart cDNA library. Sequence analysis of the cDNA clones indicated that the heart transcript was the same form as in rat brain, but differed from the germ cell-specific testis transcript that utilizes a different transcriptional start site. Heart proenkephalin cRNA translated efficiently, resulting in the synthesis of a 35 kDa protein that was immunoprecipitated by an antibody specific to the protein. The transcriptional initiation sites utilized in the heart were the same as in the brain, based on primer extension studies. These data suggest that the proenkephalin transcript found in abundance in rat heart is the same form as found in the brain, and differs from the testis-type transcript. We conclude that the scant level of proenkephalin-derived peptides in the heart is not due to an intrinsic inability of the proenkephalin transcript to translate.

Amino Acid Sequence

The human glutamate receptor cDNA GluR1: cloning, sequencing, expression and localization to chromosome 5.

The rat glutamate receptor is a 907 amino acid transmembrane protein. Using the rat GluR1 cDNA as a probe, we have isolated cDNA clones from a human hippocampal cDNA library. Sequence of a full length cDNA clone revealed 98.2% and 89.4% identity to the rat sequence at the amino acid and nucleotide levels respectively. The human cDNA clone detected an RNA transcript in human cerebral cortex, hippocampus and cerebellum, similar to that seen in rat. In situ hybridization experiments showed that human GluR1 mRNA is present in granule and pyramidal cells in the hippocampal formation and that there is no apparent difference of distribution between control patient and patient with Alzheimer's disease. Dot blot analysis of flow-sorted human chromosomes showed that the GluR1 gene maps to chromosome 5.

Aged

Homeobox containing genes in the nematode Caenorhabditis elegans.

We designed a unique 36-mer oligonucleotide probe, based on the most highly conserved amino acid sequences of Antennapedia-like homeodomains and the codon bias of Caenorhabditis elegans. This probe was then used to isolate four classes of genes from a C. elegans genomic library. Sequencing reveals that we have isolated three new homeobox genes, designated ceh-1, ceh-9 and ceh-10. The fourth homeobox gene, ceh-11, has recently been described by Schaller et al (Nucleic Acids Res. 18, 2033-2036). The amino acid sequence of ceh-1 is 87% similar to the honeybee H40 homeodomain, 85% similar to the Drosophila NK-1 homeodomain and 82% similar to the chicken CHox3 homeodomain. The sequence ceh-10 appears to be a member of the paired class of homeodomains. The other two sequences, ceh-9 and ceh-11, remain unclassified. Three of the four sequences have at least one intron within the homeobox region. Transcripts of ceh-10 and ceh-11 are present in embryonic RNA but are greatly diminished in later developmental stages. Three of the four new genes have been placed on the C. elegans genomic map.

Amino Acid Sequence

Organization of the rat UDP-glucuronosyltransferase, UDPGTr-2, gene and characterization of its promoter.

A lambda clone containing the entire gene and flanking sequences for a form of UDP-glucuronosyltransferase (UDPGTr-2) that glucuronidates testosterone and the foreign compounds, 4-hydroxybiphenyl and chloramphenicol, has been isolated from a rat liver genomic library. Sequence analysis of this clone revealed that the UDPGTr-2 gene is approximately 12 kilobase pairs in length and consists of six exons. All introns were found to interrupt protein coding regions of the gene. Four transcriptional start sites have been identified and are located 34, 35, 38, and 39 base pairs (bp) upstream of the translation initiation site. The 5'-flanking region of the gene contains a TATA-like sequence, CATAAA, 22 bp from the first transcription start site, potential AP-1 and v-MYB binding sites, and four sequence motifs that have been found in genes that are expressed predominantly in the liver. A 323-bp fragment encompassing these elements was fused upstream in both orientations to the coding sequence of placental alkaline phosphatase to assay promoter activity. Transient transfection of various cultured cell lines with the chimeric DNA demonstrated that this fragment, in the correct orientation, was able to function as an efficient promoter in the rat hepatoma cell lines Reuber H4-II-E and McA-RH7777. It was, however, inactive in hepatoma cell lines from two other species and in cell lines derived from other tissues. These results are consistent with the physiological expression of the rat UDPGTr-2 gene and suggest that the proximal 5'-flanking region of the gene may contain information which limits its expression to the liver.

Alkaline Phosphatase

Characterization of a highly polymorphic region 5' to JH in the human immunoglobulin heavy chain.

A cloned DNA segment 1.25 kilobases (kb) upstream from the joining segments of the human heavy chain immunoglobulin gene revealed extensive polymorphic variation at this locus, and the polymorphic pattern was stably transmitted to the next generation. Genomic restriction analysis showed that the polymorphism was caused by insertions/deletions within an MspI/BamHI fragment. Sequencing of one allele, 848 base pairs (bp) long, revealed eleven 50-base-pair tandem repeats. A second allele, 648 bp long, was cloned from a human genomic cosmid library, sequenced, and found to contain four fewer repeats than the first allele. A survey of 186 chromosomes from unrelated individuals of primarily northern European descent revealed at least six alleles.

Alleles

Cloning and characterization of a cDNA specific for bovine retina.

A retina-specific cDNA clone (pCR18) was selected from a bovine retinal cDNA library and characterized. The clone pCR18 consisted of 905 base pairs and hybridized to the mRNA of about 12S from the bovine retina, but not that from the brain or liver. The nucleotide sequence revealed a long open reading frame which encodes a 147 amino acid polypeptide of about 15,700 Da. No significant sequence homology with the predicted protein was found in the protein sequence library of about 3500. Messenger RNA which hybridized to pCR18 translated a polypeptide of about 19,000 Da in a reticulocyte translation system. Southern blot analysis indicated that the bovine genome contains a single copy of this gene. Furthermore, RNA dot analysis showed that the poly(A)+ RNA from the human retinoblastoma cell lines (Y79 and WERI) hybridized to pCR18, whose intensity was comparable to that of the bovine retina. In situ hybridization revealed that pCR18 was expressed mostly in some ganglion cells of the rat retina. The results suggest that cDNA clone (pCR18) encodes a protein specific for the retina and mRNA for pCR18 is mostly localized in the retinal ganglion cells and also expressed in the human retinoblastoma cells, although its function remains to be elucidated.

Animals

Molecular cloning sequence and distribution of rat calspermin, a high affinity calmodulin-binding protein.

Calspermin is a heat-stable, acidic calmodulin-binding protein predominantly found in mammalian testis. The cDNA representing the rat form of this protein has been cloned from a rat testis lambda gt11 library. Sequence analysis of two overlapping clones revealed a 232-nucleotide 5'-nontranslated region, 510 nucleotides of open reading frame, a 148-nucleotide 3'-untranslated region, and a poly(A) tail. Authenticity of the clones was confirmed by comparison of a portion of the deduced amino acid sequence with the sequence of a tryptic peptide obtained from the rat testis protein. The lambda gt11 fusion protein was recognized by affinity purified antibodies to pig testis calspermin and bound 125I-calmodulin in a Ca2+-dependent manner. Calspermin cDNA encodes a 169-residue protein with a calculated Mr of 18,735. The putative calmodulin-binding domain is very close to the amino terminus of the protein. This region shows 46% identity with the calmodulin-binding region of rat brain Ca2+/calmodulin-dependent protein kinase II and 32% identity with the equivalent region of chicken smooth muscle myosin light chain kinase. The 5'-nontranslated region reveals significant homology with a portion of the catalytic region of the calmodulin-dependent protein kinase family. Calspermin contains a stretch of 17 contiguous glutamic acid residues in the central region of the molecule. Computer analysis predicts calspermin to be 81% alpha-helix and 14% random coil. Analysis of genomic DNA indicates calspermin to be the product of a unique gene. Northern blot analysis of rat testis RNA reveals a 1.1-kilobase mRNA. This RNA is restricted to testis among several rat tissues examined and could not be identified in total RNA isolated from testes of other mammals. Analysis of cells isolated from rat testis reveals calspermin mRNA to be predominantly expressed in postmeiotic cells indicating that it may be specific to haploid cells.

Amino Acid Sequence

Volumetric DNA microscopy for mapping spatial transcriptomes in three dimensions.

The architecture and function of biological systems are inherently three-dimensional, yet most existing spatial transcriptomic technologies remain restricted to thin tissue sections, limiting their capacity to resolve cellular organization and microenvironments within intact tissue volumes. To address this limitation, we developed volumetric DNA microscopy, a scalable, optics-free approach for spatial transcriptome profiling directly within intact biological specimens. The method encodes spatial information into DNA molecules that form a dense intermolecular network in situ, enabling the reconstruction of three-dimensional spatial relationships through short-read sequencing and computational analysis. Here we detail the complete workflow including in situ cDNA synthesis, spatial encoding through DNA nanoball formation, dual-scale proximity bridging between neighboring nanoballs and spatial reconstruction via geodesic spectral embedding. Sequencing libraries can be generated within 7-8 d by a competent graduate-level molecular biologist, followed by standardized downstream computational analysis. Because the workflow requires only routine molecular biology reagents and a benchtop sequencer, volumetric DNA microscopy provides a versatile platform for exploring genetic and morphological features in intact tissues.

Spatial Transcriptomics

Alternative splicing of human glucose-6-phosphate dehydrogenase messenger RNA in different tissues.

Different forms of glucose-6-phosphate dehydrogenase (G-6-PD) have been described in different tissues. Moreover, the directly determined amino acid sequence amino end of the red cell enzyme does not exactly match the sequence deduced from cDNA isolated from HeLa cells or lymphoblasts. We have therefore investigated the sequence of cDNA from sperm, granulocytes, reticulocytes, brain, placenta, liver, lymphoblastoid cells, and cultured fibroblasts. A novel human cDNA, which has extra 138 bases coding 46 amino acids, was isolated from a lymphoblastoid cell library. Sequencing of genomic DNA amplified by the polymerase chain reaction (PCR) revealed that the extra sequence was derived from the 3'-end of intron 7 by alternative splicing. This longer form of mRNA was also detected in sperm and granulocytes. Sequence analysis using PCR-amplified cDNA revealed that the 5'-end of the coding sequence of G6PD mRNA in reticulocytes is identical to those in other tissues.

Amino Acid Sequence

Identification of the 1.4 kb and 4.0 kb messages for the lipoprotein associated coagulation inhibitor and expression of the encoded protein.

Lipoprotein-Associated Coagulation Inhibitor (LACI) is a factor Xa dependent inhibitor of the factor VII(a)/Tissue Factor catalytic complex. Deduced from partial cDNA sequence, LACI's amino acid sequence has recently been reported. Northern blot analysis showed LACI cDNA hybridizes to RNAs of 1.4 and 4.0 kb in size. To complete the characterization of the LACI message(s), overlapping LACI cDNAs were isolated from a human endothelial cell library. Sequence analysis revealed the clones' inserts span 4023 bases of sequence, consisting of 381 bases of 5' untranslated sequence, an open reading frame of 912 bases, 2682 bases of 3' untranslated sequence and 48 bases of poly(A) sequence. In addition, a short 1.4 kb insert which encodes for LACI was found to contain 49 bases of 3' untranslated sequence and a 3' poly(A) tail. The 1.4 kb of sequence is contained in the 4.0 kb sequence, except for 14 bases of 5' sequence, suggesting that the LACI messages arise by the use of alternative termination and polyadenylation signals during processing. Northern blot analysis of RNA isolated from cells treated with actinomycin D showed both RNA species appear to be relatively stable. Using a bovine papilloma virus vector, LACI cDNA was transfected into mouse C127 fibroblasts. The recombinant LACI is recognized by polyclonal anti-LACI IgG, binds to factor Xa and inhibits VII(a)/Tissue Factor activity in a similar fashion as LACI purified from HepG2 cell conditioned media.

Amino Acid Sequence

Brain-specific expression of transthyretin mRNA as revealed by cDNA cloning from brain.

cDNAs for rat transthyretin mRNA were cloned from a brain cDNA library. Sequencing analyses showed the presence of an additional 5' sequence that had not been reported for the liver mRNA corresponding to the flanking promoter region of the gene. This additional sequence was expressed only in the brain, suggesting the presence of a brain-specific promoter.

Animals

Comprehensive evaluation of new sequencer T20 and well-established T7 with 507 human samples.

The DNBSEQ-T20×2 (T20) sequencer, developed by MGI Tech, enables cost-effective human whole-genome sequencing (WGS) at 30× coverage for less than $100 per genome. Here, we evaluate the sequencing performance and data quality of the T20 platform by benchmarking it against the established DNBSEQ-T7 (T7) sequencer using 507 samples derived from blood (N = 75), stool (N = 242), and saliva (N = 190). The T20 exhibited lower sequencing quality metrics compared with the T7, with Q20 scores of 95.76%-95.83% and Q30 scores of 87.25%-87.40%, compared with 97.81%-97.93% and 93.26%-93.60%, respectively, for T7 data. Quality differences were more evident toward the end of reads, and PCR-free libraries sequenced on the T20 showed similar reductions in quality scores. The median empirical base error rate estimated from 102 ZymoBIOMICS samples was 0.33%. The T20 demonstrated comparable coverage uniformity to the T7 and showed high concordance in microbiome composition analysis, with a median Bray-Curtis dissimilarity of 0.02. Variant calling performance was highly consistent between the two platforms. Among variants with non-missing genotype calls on both platforms, 94.92% of SNPs and 87.20% of InDels showed concordant genotypes between T20 and T7. Overall, the T20 delivers reliable sequencing accuracy and reproducibility for large-scale genomic and microbiome studies, providing a cost-effective alternative for high-throughput sequencing applications.

Metagenomics

Physical characterization and sequence identification of the ovary maturating parsin. A new neurohormone purified from the nervous corpora cardiaca of the African locust (Locusta migratoria migratorioides.

A novel neurohormone, which anticipates ovarian maturation, was recently purified using liquid chromatography from the African locust nervous corpora cardiaca. Both its function and production by the pars intercerebralis of Locusta migratoria lead to its name, the ovary maturating parsin (Lom OMP). In this study, the Lom OMP was physically and chemically characterized. Its multiply charged ion spectrum was interpreted as two peaks of quite equal size having molecular masses of 6923.4 Da (major peak) and 6907.3 Da. The Lom OMP presented no periodic secondary structure according to the far ultraviolet circular dichroism spectrum obtained. It is composed of 65 amino acids and included a high concentration of alanine but is devoid of cysteine, isoleucine, methionine, lysine and threonine. The amino acid sequence indicated only one microheterogeneity, observed at position 26, consisted in the replacement of serine by alanine. The calculated Mr of the two acidic isoforms (calculated pHi = 4.87) were found to be in agreement with mass spectrometry measurements. When compared to the sequence libraries, the Lom OMP, the first insect gonadotropic neurohormone, was revealed as an unique protein.

Amino Acid Sequence

Complementary DNA clones of chicken proto-oncogene c-ets: sequence divergence from the viral oncogene v-ets.

The avian acute leukemia virus E26 induces erythroblastosis and myeloblastosis in chickens. The oncogene of this virus includes sequences derived from the cellular gene designated c-ets, which is normally expressed in lymphoid cells and whose product is a protein of apparent molecular weight ca. 54,000 daltons. Complementary DNA clones representing the major transcript of the chicken c-ets proto-oncogene were isolated from a spleen cell library. Sequence analysis of the cDNA revealed that it contains an open reading frame encoding a polypeptide of 441 amino acids with a molecular weight of 49,932 daltons. This open reading frame can be transcribed and translated in vitro into a 50 kd protein that is specifically immunoprecipitated with antiserum to the v-ets oncogene product. Within the central region of homology between c-ets and v-ets, there are only 5 nucleotide substitutions resulting in 4 amino acid changes. However, coding sequences at the 5' and 3' ends of the v-ets oncogene and the chicken c-ets cDNA differ from one another. These changes may be responsible for the differential functions of c-ets and v-ets in cells of different hematopoietic lineages and may account for the pathogenic properties of the v-ets oncogene.

Amino Acid Sequence

Tracking-seq: a universal off-target detection approach for CRISPR-Cas genome editing.

Tracking-seq is a highly sensitive method for genome-wide detection of off-target effects in cells edited with diverse genome editing modalities, including Cas9, cytosine base editors, adenine base editors and prime editors. Since most genome editors induce DNA repair pathways and generate single-stranded DNA (ssDNA) intermediates, Tracking-seq leverages this process by tracking replication protein A-a key protein that binds and protects ssDNA-to identify on-target and off-target events. Here we provide a detailed protocol for Tracking-seq, covering genome editing of cells, extraction of replication protein A-bound ssDNA, sequencing library construction and data analysis using our custom computational tool Offtracker. Tracking-seq is applicable to various genome editing scenarios with low cell input, delivering high-performance results. The entire workflow, from genome editing to data analysis, can be completed within 1-2 weeks, making it a rapid solution for assessing genome-wide off-target activity.

CRISPR-Cas Systems

Strategies for mosaic variant calling in brain disorders.

The human brain is a genomic mosaic, where postzygotic mutations arising from embryogenesis to senescence drive diverse neurodevelopmental and neurodegenerative diseases. Because of numerous sequencing artifacts at ultralow variant allele frequencies (VAFs), detecting these variants remains a significant analytical challenge. This review focuses on single-nucleotide variants and small indels, summarizing current strategies for aligning sampling methods, including bulk, laser capture microdissection, and single-cell genomics, with the expected clonal architecture of the brain. It emphasizes that mosaic detection sensitivity is fundamentally constrained by sequencing depth, since even the most advanced algorithms cannot identify variants not physically represented in the sequencing library. The review further recommends the selection of variant calling algorithms based on validated VAF detection performance, matching tools like MuTect2 and MosaicForecast to their optimal performance ranges. Furthermore, we discuss how multitissue sampling, as emphasized by the SMaHT project, addresses the matched-control dilemma and supports accurate variant classification via cross-tissue VAF gradients. Integrating these established pipelines with multiomics modalities, including transcriptomic and epigenetic data, could advance the field toward a functional understanding of how the somatic genome impacts human brain health and disease.

Humans

Construction of a prognostic model for gastric cancer based on immune infiltration and microenvironment, and exploration of MEF2C gene function.

BACKGROUND: Advanced gastric cancer (GC) exhibits a high recurrence rate and a dismal prognosis. Myocyte enhancer factor 2c (MEF2C) was found to contribute to the development of various types of cancer. Therefore, our aim is to develop a prognostic model that predicts the prognosis of GC patients and initially explore the role of MEF2C in immunotherapy for GC. METHODS: Transcriptome sequence data of GC was obtained from The Cancer Genome Atlas (TCGA), the Gene Expression Omnibus (GEO) and PRJEB25780 cohort for subsequent immune infiltration analysis, immune microenvironment analysis, consensus clustering analysis and feature selection for definition and classification of gene M and N. Principal component analysis (PCA) modeling was performed based on gene M and N for the calculation of immune checkpoint inhibitor (ICI) Score. Then, a Nomogram was constructed and evaluated for predicting the prognosis of GC patients, based on univariate and multivariate Cox regression. Functional enrichment analysis was performed to initially investigate the potential biological mechanisms. Through Genomics of Drug Sensitivity in Cancer (GDSC) dataset, the estimated IC50 values of several chemotherapeutic drugs were calculated. Tumor-related transcription factors (TFs) were retrieved from the Cistrome Cancer database and utilized our model to screen these TFs, and weighted correlation network analysis (WGCNA) was performed to identify transcription factors strongly associated with immunotherapy in GC. Finally, 10 patients with advanced GC were enrolled from Sun Yat-sen University Cancer Center, including paired tumor tissues, paracancerous tissues and peritoneal metastases, for preparing sequencing library, in order to perform external validation. RESULTS: Lower ICI Score was correlated with improved prognosis in both the training and validation cohorts. First, lower mutant-allele tumor heterogeneity (MATH) was associated with lower ICI Score, and those GC patients with lower MATH and lower ICI Score had the best prognosis. Second, regardless of the T or N staging, the low ICI Score group had significantly higher overall survival (OS) compared to the high ICI Score group. For its mechanisms, consistently, for Camptothecin, Doxorubicin, Mitomycin, Docetaxel, Cisplatin, Vinblastine, Sorafenib and Paclitaxel, all of the IC50 values were significantly lower in the low ICI Score group compared to the high ICI Score group. As a result, based on univariate and multivariate Cox regression, ICI Score was considered to be an independent prognostic factor for GC. And our Nomogram showed good agreement between predicted and actual probabilities. Based on CIBERSORT deconvolution analysis, there was difference of immune cell composition found between high and low ICI Score groups, probably affecting the efficacy of immunotherapy. Then, MEF2C, a tumor-related transcription factor, was screened out by WGCNA analysis. Higher MEF2C expression is significantly correlated with a worse OS. Moreover, its higher expression is also negatively correlated with tumor mutation burden (TMB) and microsatellite instability (MSI), but positively correlated with several immunosuppressive molecules, indicating MEF2C may exert its influence on tumor development by upregulating immunosuppressive molecules. Finally, based on transcriptome sequencing data on 10 paired tumor tissues from Sun Yat-sen University Cancer Center, MEF2C expression was significantly lower in paracancerous tissues compared to tumor tissues and peritoneal metastases, and it was also lower in tumor tissues compared to peritoneal metastases, indicating a potential positive association between MEF2C expression and tumor invasiveness. CONCLUSIONS: Our prognostic model can effectively predict outcomes and facilitate stratification GC patients, offering valuable insights for clinical decision-making. The identified transcription factor MEF2C can serve as a biomarker for assessing the efficacy of immunotherapy for GC.

Humans