PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

circASbase: A Comprehensive Database of Alternative Splicing Events in circRNAs.

Although extensive evidence has underscored the critical role of alternative splicing (AS) in generating mature circular RNA (circRNA) isoforms and augmenting their functional diversity, a significant gap remains in the availability of specialized databases housing circRNA AS events. To bridge this gap, we develop circASbase, a pioneering and comprehensive database that catalogs 452,129 AS events in 884,047 full-length circRNAs from 581 samples across 13 species, and provides rich annotations to facilitate understanding the splicing regulation of circRNA. Our findings reveal substantial differences between circRNAs and linear transcripts regarding the distribution and occurrence of AS events, highlighting the unique regulatory landscape of circRNAs. These special splicing events result in functional differences of circRNAs by affecting internal ribosome entry sites, N6-methyladenosine sites, open reading frames, protein features, microRNA targets, and more. In summary, circASbase not only meets the urgent need of the research community for data repositories, but also represents a significant advancement in our understanding of circRNA biology. With its user-friendly interfaces and web-based visualization tools, circASbase is poised to become an indispensable resource for researchers exploring the regulatory mechanisms and functional roles of AS events in circRNAs. This database will continuously drive new insights and discoveries in the field, setting the stage for further advancements in circRNA research. circASbase is freely available at http://reprod.njmu.edu.cn/cgi-bin/circASbase/.

Alternative Splicing↗

Characterization of a root-specific Arabidopsis terpene synthase responsible for the formation of the volatile monoterpene 1,8-cineole.

Arabidopsis is emerging as a model system to study the biochemistry, biological functions, and evolution of plant terpene secondary metabolism. It was previously shown that the Arabidopsis genome contains over 30 genes potentially encoding terpene synthases (TPSs). Here we report the characterization of a monoterpene synthase encoded by two identical, closely linked genes, At3g25820 and At3g25830. Transcripts of these genes were detected almost exclusively in roots. An At3g25820/At3g25830 cDNA was expressed in Escherichia coli, and the protein thus produced was shown to catalyze the formation of 10 volatile monoterpenes from geranyl diphosphate, with 1,8-cineole predominating. This protein was therefore designated AtTPS-Cin. The purified recombinant AtTPS-Cin displayed similar biochemical properties to other known monoterpene synthases, except for a relatively low K(m) value for geranyl diphosphate of 0.2 microm. At3g25820/At3g25830 promoter activity, measured with a beta-glucuronidase (GUS) reporter gene, was primarily found in the epidermis, cortex, and stele of mature primary and lateral roots, but not in the root meristem or the elongation zone. Although the products of AtTPS-Cin were not detected by direct extraction of plant tissue, the recent report of 1,8-cineole as an Arabidopsis root volatile (Steeghs M, Bais HP, de Gouw J, Goldan P, Kuster W, Northway M, Fall R, Vivanco JM [2004] Plant Physiol 135: 47-58) suggests that the enzyme products may be released into the rhizosphere rather than accumulated. Among Arabidopsis TPSs, AtTPS-Cin is most similar to the TPS encoded by At3g25810, a closely linked gene previously shown to be exclusively expressed in flowers. At3g25810 TPS catalyzes the formation of a set of monoterpenes that is very similar to those produced by AtTPS-Cin, but its major products are myrcene and (E)-beta-ocimene, and it does not form 1,8-cineole. These data demonstrate that divergence of organ expression pattern and product specificity are ongoing processes within the Arabidopsis TPS family.

Alkyl and Aryl Transferases↗

Novel transcription factors in human CD34 antigen-positive hematopoietic cells.

Transcription factors (TFs) and the regulatory proteins that control them play key roles in hematopoiesis, controlling basic processes of cell growth and differentiation; disruption of these processes may lead to leukemogenesis. Here we attempt to identify functionally novel and partially characterized TFs/regulatory proteins that are expressed in undifferentiated hematopoietic tissue. We surveyed our database of 15 970 genes/expressed sequence tags (ESTs) representing the normal human CD34(+) cells transcriptosome (http://westsun.hema.uic.edu/cd34.html), using the UniGene annotation text descriptor, to identify genes with motifs consistent with transcriptional regulators; 285 genes were identified. We also extracted the human homologues of the TFs reported in the murine stem cell database (SCdb; http://stemcell.princeton.edu/), selecting an additional 45 genes/ESTs. An exhaustive literature search of each of these 330 unique genes was performed to determine if any had been previously reported and to obtain additional characterizing information. Of the resulting gene list, 106 were considered to be potential TFs. Overall, the transcriptional regulator dataset consists of 165 novel or poorly characterized genes, including 25 that appeared to be TFs. Among these novel and poorly characterized genes are a cell growth regulatory with ring finger domain protein (CGR19, Hs.59106), an RB-associated CRAB repressor (RBAK, Hs.7222), a death-associated transcription factor 1 (DATF1, Hs.155313), and a p38-interacting protein (P38IP, Hs. 171185). The identification of these novel and partially characterized potential transcriptional regulators adds a wealth of information to understanding the molecular aspects of hematopoiesis and hematopoietic disorders.

Amino Acid Motifs↗

A multispecies comparison of the metazoan 3'-processing downstream elements and the CstF-64 RNA recognition motif.

BACKGROUND: The Cleavage Stimulation Factor (CstF) is a required protein complex for eukaryotic mRNA 3'-processing. CstF interacts with 3'-processing downstream elements (DSEs) through its 64-kDa subunit, CstF-64; however, the exact nature of this interaction has remained unclear. We used EST-to-genome alignments to identify and extract large sets of putative 3'-processing sites for mRNA from ten metazoan species, including Homo sapiens, Canis familiaris, Rattus norvegicus, Mus musculus, Gallus gallus, Danio rerio, Takifugu rubripes, Drosophila melanogaster, Anopheles gambiae, and Caenorhabditis elegans. In order to further delineate the details of the mRNA-protein interaction, we obtained and multiply aligned CstF-64 protein sequences from the same species. RESULTS: We characterized the sequence content and specific positioning of putative DSEs across the range of organisms studied. Our analysis characterized the downstream element (DSE) as two distinct parts - a proximal UG-rich element and a distal U-rich element. We find that while the U-rich element is largely conserved in all of the organisms studied, the UG-rich element is not. Multiple alignment of the CstF-64 RNA recognition motif revealed that, while it is highly conserved throughout metazoans, we can identify amino acid changes that correlate with observed variation in the sequence content and positioning of the DSEs. CONCLUSION: Our analysis confirms the early reports of separate U- and UG-rich DSEs. The correlated variations in protein sequence and mRNA binding sequences provide novel insights into the interactions between the precursor mRNA and the 3'-processing machinery.

Amino Acid Motifs↗

Comparative analysis of the Kekkon molecules, related members of the LIG superfamily.

Leucine-rich repeats (LRRs) and immunoglobulin (Ig) domains represent two of the most abundant sequence elements in metazoan proteomes. Despite this prevalence, comparatively few molecules containing both LRR and Ig (LIG) modules exist, and fewer still have been functionally defined. One LIG whose function has been investigated is the Drosophila protein Kekkon1 (Kek1). In vivo studies have demonstrated a role for Kek1 in Epidermal Growth Factor Receptor (EGFR) signaling and have suggested a role in neuronal pathfinding. Kek1 is the founding member of the Kek family, a group of six Drosophila transmembrane proteins that contain seven LRRs and a single Ig in their extracellular domains. While this arrangement of domains predicts a possible role as cell adhesion molecules (CAMs), to date little is known about the function or evolutionary relationship of these additional Kek molecules. Here we report that orthologs of Kek1, Kek2, Kek5, and Kek6 exist in the mosquito, Anopheles gambiae, and the honeybee, Apis mellifera, indicating that this family has been conserved for ~300 million years of evolutionary time. Comparative sequence analyses reveal remarkable identity among these orthologs, primarily in their extracellular regions. In contrast, the intracellular regions are more divergent, exhibiting only small pockets of conservation. In addition, we provide support for the general notion that these molecules may share common functions as CAMs, by demonstrating that Kek family members can form homotypic and heterotypic complexes.

Amino Acid Sequence↗

Identification of an AraC-like regulator gene required for induction of the 78-kDa ferrioxamine B receptor in Vibrio vulnificus.

We previously reported that the 78-kDa outer membrane receptor for ferrioxamine B is induced in iron-starved Vibrio vulnificus cells when desferrioxamine B was supplied exogenously. Based on its N-terminal amino acid sequence, a candidate gene for the ferrichrome B receptor was detected in the V. vulnificus CMCP6 genomic database. Here, two contiguous genes, named desR and desA, encoding a member of the AraC family of transcriptional activators and the ferrioxamine B receptor, respectively, were cloned from V. vulnificus M2799 and characterized. Primer extension analysis mapped the iron-regulated transcription initiation sites for desR and desA, and demonstrated involvement of desferrioxamine B in the induction of desA transcription. Insertion mutation of desR resulted in no production of DesA under iron-limiting conditions even in the presence of desferrioxamine B. The DesA production under the same conditions was restored to wild-type levels when the desR mutant was complemented with desR in trans. These results suggest that the desR gene is required for desferrioxamine B-inducible production of DesA in iron-starved cells.

Amino Acid Sequence↗

Molecular cloning and characterization of rat karyopherin alpha 1 gene: structure and expression.

Dopamine denervation in the striata of patients with Parkinson's disease (PD) leads to changes in neural plasticity. However, the mechanisms leading to the changes are still poorly understood. In an effort to study the molecular events in the denervated striatum, we identified and cloned rat karyopherin alpha 1 (KPNA1), a member of the importin/karyopherin alpha (KPNA) family. DNA sequence analysis revealed that the full-length cDNA, encoding rat KPNA1, was 4975 bp with a short 5'-untranslated region (UTR) of 70 bp, a putative coding sequence of 1617 bp, and an unusually long 3'-UTR of 3266 bp. The gene shared a high degree of similarity with its mouse and human homologs at both cDNA and protein levels. By computational analysis of its genomic sequence, the transcription unit was shown to span a 44-kb region and consist of 13 exons varying in size from 89 (6th exon) to 3454 bp (13th exon), and 12 introns varying in size from 0.3 to 8.9 kb. Reverse transcriptase-polymerase chain reaction (RT-PCR) analysis demonstrated that KPNA1 transcript existed in various adult tissues. Both Northern blot and semi-quantitative RT-PCR analysis showed that the expression level of KPNA1 mRNA was altered in the denervated striatum post-lesion in a time-dependent manner, reaching the maximum at 2 weeks post-lesion. Our results suggest involvement of KPNA1 in the striatal responses to denervation following 6-hydroxydopamine (6-OHDA)-induced lesion.

3' Flanking Region↗

The characterisation and functional analysis of the human glyoxalase-1 gene using methods of bioinformatics.

Methylglyoxal (MG), which forms MG-derived AGE, is elevated in diabetic subjects with vascular disease. Detoxification of MG occurs through the glyoxalase system incorporating glyoxalase-1 (GLO1) and glyoxalase-2. Perturbations of the glyoxalase-1 gene (GLO1) may result in vulnerability to vascular complications through alterations in AGE interactions. We used bioinformatics to predict the structure, function and genetic variation of GLO1. We identified a previously unreported exon. Seventy single nucleotide polymorphisms (SNPs) were identified bioinformatically. The amino acid substitution Ala 111 Glu was confirmed and predicted to be tolerant. Though no alternative splice variants were identified, novel multiple alternative transcription start sites and alternative 3' UTRs were demonstrated. Ubiquitous expression of GLO1 was confirmed. Conserved regulatory regions were predicted 5' to the transcription start site and in the distal promoter, and several predicted conserved transcription regulatory elements were suggested in the 5' UTR. This study of GLO1 demonstrates multiple sequence variants at DNA and mRNA levels, areas of sequence conservation and SNPs that are predicted to affect function. A differential ability of glyoxalase-1 to reduce the formation and subsequent interaction of AGEs may have a role in the structural and functional manifestations of diabetic vascular disease.

Amino Acid Sequence↗

Selection of peptides with affinity for the N-terminal domain of GATA-1: Identification of a potential interacting protein.

As most transcription factors, GATA-1 activities are mediated by interactions with multiple proteins. Those identified so far associate with the zinc-finger domain and/or surrounding sequences. In contrast, no proteins interacting with the N-terminal domain have been identified although several evidences suggest its involvement in the control of hematopoiesis. In an attempt to identify proteins that interact with the N-terminal transactivation domain of GATA-1, a random phage peptide library was screened with recombinant GATA-1 protein and the sequence of a selected peptide was used for database protein sequence retrieval. We selected a set of peptides sharing the core sequence phi-B((2-3))-nu((2-4)) (where phi, B, and nu represent hydrophobic, basic, and neutral residues, respectively). Using the sequence of the most represented peptide (pep5) as query, we retrieved the HIV accessory protein Nef. We show that Nef binds GATA-1 and GATA-3 in vitro in virtue of its sequence homology with pep5.

Amino Acid Sequence↗

Cysteine and tyrosine-rich 1 (CYYR1), a novel unpredicted gene on human chromosome 21 (21q21.2), encodes a cysteine and tyrosine-rich protein and defines a new family of highly conserved vertebrate-specific genes.

A novel human gene has been identified by in-depth bioinformatics analysis of chromosome 21 segment 40/105 (21q21.1), with no coding region predicted in any previous analysis. Brain-derived DNA complementary to RNA (cDNA) sequencing predicts a 154-amino acid product with no similarity to any known protein. The gene has been named cysteine and tyrosine-rich protein 1 gene (symbol cysteine and tyrosine-rich 1, CYYR1). The CYYR1 messenger RNA was found by Northern blot analysis in a broad range of tissues (two transcripts of 3.4 and 2.2 kb). The gene consists of four exons and spans about 107 kb, including a very large intron of 85.8 kb. Analysis of expressed sequence tags shows high CYYR1 expression in cells belonging to the amine precursor uptake and decarboxylation system. We also cloned the cDNA of the murine ortholog Cyyr1, which was mapped by a radiation hybrid panel on chromosome 16 within the region corresponding to that containing the respective human homolog on chromosome 21. Sequence and phylogenetic analysis led to identification of several genes encoding CYYR1 homologous proteins. The most prominent feature identified in the protein family is a central, unique cysteine and tyrosine-rich domain, which is strongly conserved from lower vertebrates (fishes) to humans but is absent in bacteria and invertebrates.

Amino Acid Sequence↗

An active DNA transposon family in rice.

The publication of draft sequences for the two subspecies of Oryza sativa (rice), japonica (cv. Nipponbare) and indica (cv. 93-11), provides a unique opportunity to study the dynamics of transposable elements in this important crop plant. Here we report the use of these sequences in a computational approach to identify the first active DNA transposons from rice and the first active miniature inverted-repeat transposable element (MITE) from any organism. A sequence classified as a Tourist-like MITE of 430 base pairs, called miniature Ping (mPing), was present in about 70 copies in Nipponbare and in about 14 copies in 93-11. These mPing elements, which are all nearly identical, transpose actively in an indica cell-culture line. Database searches identified a family of related transposase-encoding elements (called Pong), which also transpose actively in the same cells. Virtually all new insertions of mPing and Pong elements were into low-copy regions of the rice genome. Since the domestication of rice mPing MITEs have been amplified preferentially in cultivars adapted to environmental extremes-a situation that is reminiscent of the genomic shock theory for transposon activation.

Amino Acid Sequence↗

Automatic updating of the EMBL database via EMBNet.

The paper describes a procedure for updating the EMBL (European Molecular Biology Laboratory, Heidelberg) database of nucleic acid sequences and its indexes used by the University of Wisconsin Genetics Computer Group (GCG) software package, using updated entries for this database distributed via EMBNet. At present the procedure is being run on a MRC Clinical Research Centre's (CRC) SUN 4/280 server using SUNOS version 4.0.1 operating system.

Abstracting and Indexing↗

The evolution of homing endonuclease genes and group I introns in nuclear rDNA.

Group I introns are autonomous genetic elements that can catalyze their own excision from pre-RNA. Understanding how group I introns move in nuclear ribosomal (r)DNA remains an important question in evolutionary biology. Two models are invoked to explain group I intron movement. The first is termed homing and results from the action of an intron-encoded homing endonuclease that recognizes and cleaves an intronless allele at or near the intron insertion site. Alternatively, introns can be inserted into RNA through reverse splicing. Here, we present the sequences of two large group I introns from fungal nuclear rDNA, which both encode putative full-length homing endonuclease genes (HEGs). Five remnant HEGs in different fungal species are also reported. This brings the total number of known nuclear HEGs from 15 to 22. We determined the phylogeny of all known nuclear HEGs and their associated introns. We found evidence for intron-independent HEG invasion into both homologous and heterologous introns in often distantly related lineages, as well as the "switching" of HEGs between different intron peripheral loops and between sense and antisense strands of intron DNA. These results suggest that nuclear HEGs are frequently mobilized. HEG invasion appears, however, to be limited to existing introns in the same or neighboring sites. To study the intron-HEG relationship in more detail, the S943 group I intron in fungal small-subunit rDNA was used as a model system. The S943 HEG is shown to be widely distributed as functional, inactivated, or remnant ORFs in S943 introns.

Amino Acid Sequence↗

Cloning and functional expression of a novel human connexin-25 gene.

Gap junctions are intercellular, water-filled channels composed of transmembrane proteins called connexins, six of which are arranged radially and dock with six homologous proteins in an adjacent cell to form an approximate 16 A pore. Through this pore cell-to-cell transfer of small water-soluble molecules up to about 1000 daltons occurs along concentration gradients. Connexins comprise a multigene family that share consensus sequences in the trans-membrane domains and the first and second extracellular loops. Comparison of the protein sequences of known human connexins with the draft nucleotide sequence of the human genome revealed two clones from chromosome 6 which showed strong similarity to highly conserved connexin sequences. Detailed analysis revealed the presence of a 672 nt open reading frame in these clones, encoding a 223 amino acid polypeptide with a predicted molecular weight of about 25 kD. This is smaller than other known human connexins. The ORF of the potential connexin25 was amplified by semi-nested PCR using human genomic DNA as a template. To confirm that this new gene encodes a connexin, Cx25 was transfected into a gap junction deficient subclone of the human HeLa cell line. After selection of transformants, cells were microinjected with the fluorescent dye Lucifer yellow. Transfectants but not controls successfully transferred dye, demonstrating that this new gene encodes a functional connexin.

Amino Acid Sequence↗

[In silicon cloning of the human TECTB gene].

The coding sequence of the mouse Tectb and chick Tectb gene were subjected to Blastn searching against the human dbEST and Htgs in NCBI. One BAC clone sequence(GenBank: AL157786) was obtained, which shows high homology to the two genes. We predicted the exons and introns in the homologous region of AL157786 using GENSCAN, MZEF and Blast 2 sequence program, and then assembled the predicted exons into the coding sequence of the human TECTB. The open reading frame of human TECTB gene is 990 bp composed of ten exons, which encodes a protein of 329 amino acids. Human TECTB gene shows 88.1% identity in 990 bp overlap with that of the mouse Tectb gene and the predicted polypeptide shows 94.2% identity in 329 as with the mouse beta-tectorin. The TECTB gene was mapped to human chromosome 10q25 by electronic-PCR.

Amino Acid Sequence↗