PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

The NEIBank project for ocular genomics: data-mining gene expression in human and rodent eye tissues.

NEIBank is a project to gather and organize genomic resources for eye research. The first phase of this project covers the construction and sequence analysis of cDNA libraries from human and animal model eye tissues to develop an overview of the repertoire of genes expressed in the eye and a resource of cDNA clones for further studies. The sequence data are grouped and identified using the tools of bioinformatics and the results are displayed through a web site where they can be interrogated by keyword search, chromosome location, by Blast (sequence comparison) or by alignment on completed genomes. Many novel proteins and novel splice forms of known genes have already emerged from analysis of the accumulating data. This review provides an overview of the current state of the database for human eye tissues, with specific comparisons to some parallel data from mouse and rat, and with illustrative examples of the kinds of insights and discoveries these data can produce. One of the major themes that emerges is that at the molecular level human eye tissues have significant differences from those of rodents, encompassing species specific genes, alternative splice forms and great variation in levels of gene expression. These point to specific adaptations and mechanisms in the human eye and emphasize that care needs to be taken in the application of appropriate animal model systems.

Amino Acid Sequence↗

Overlap of the gene encoding the novel poly(ADP-ribose) polymerase Parp10 with the plectin 1 gene and common use of exon sequences.

We have recently identified PARP10 as a novel functional poly(ADP-ribose) polymerase. The gene encoding PARP10 is conserved in vertebrates but no orthologs were found in lower organisms. In addition to the poly(ADP-ribose) polymerase domain, PARP10 possesses several additional sequence motifs, including an RNA recognition motif and two ubiquitin interaction motifs. We characterized the murine genomic locus of the Parp10 gene. We noticed that 3' Parp10 sequences overlapped with the plectin 1 gene in a head-to-tail arrangement. Detailed analyses revealed that the two most 3' Parp10 exons (exons 10 and 11) are also used for plectin 1. While these two exons code for part of the poly(ADP-ribose) polymerase domain in Parp10, they are noncoding for plectin 1 due to the lack of appropriate start codons. Furthermore our findings suggest that at least one of the plectin 1 promoters is located within intron 9 of the Parp10 gene.

Amino Acid Sequence↗

Genomic annotation and expression analysis of the zebrafish Rho small GTPase family during development and bacterial infection.

The zebrafish genomic sequence database was analyzed for the presence of genes encoding members of the Rho small GTPases. The analysis shows the presence of 32 zebrafish Rho genes representing one or more homologs of the human RHOA, RND3, RHOF, RHOG, RHOH, RHOJ, RHOU, RHOV, CDC42, RAC1, RAC2, RAC3, RND1, RHOBTB1, RHOBTB2, RHOBTB3, and RHOT1 genes. By expression analysis using reverse transcriptase-PCR we show that at least 20 of the predicted zebrafish small GTPase genes are expressed in the adult stage. Interestingly, only 5 of these were found to be expressed at early embryonic stages, including rhoab, rhoad, cdc42a, cdc42c, and rac1a. We observed a strong upregulation of zebrafish rhogb expression after Mycobacterium marinum infection of adult fish. This complete annotation study provides a firm basis for the use of zebrafish as a model for analysis of Rho GTPase function in vertebrate development and the innate immune system.

Amino Acid Sequence↗

Positively selected amino acid sites in the entire coding region of hepatitis C virus subtype 1b.

To predict the amino acid sites important for the clearance of hepatitis C virus (HCV) subtype 1b in vivo, positively selected amino acid sites were detected by analyzing the sequence data collected from the international DNA databank. The rate of nonsynonymous substitutions per nonsynonymous site was compared with that of synonymous substitutions per synonymous site for each codon site in the entire coding region. As a result, 13 out of 3010 amino acid sites were found to be positively selected. Among the 13 positively selected amino acid sites, eight were located in the structural proteins and five were in the nonstructural proteins. Moreover, eight were located in B-cell epitopes and two were in T-cell epitopes. These observations suggest that both the antibody and the cytotoxic T lymphocyte are involved in the clearance of HCV subtype 1b in vivo. These positively selected amino acid sites represent candidate vaccination targets for HCV subtype 1b.

Amino Acids↗

Comparative genomic organization of the cbl genes.

The genomic organization of cbl genes from a variety of mammalian and non-mammalian species was determined by a combination of cloning and database searches. Humans and mice have three cbl genes (c-cbl,(1) cblb, and cblc) which show remarkable conservation of the intron/exon structure over the region of the genes which encode the highly conserved N-terminal region of the proteins including the RING finger. Searches of genomic, cDNA, and EST databases revealed that one or more cbl genes exist in chordates, insects, and worms. Comparison of the complexity and genomic organization of the cbl gene family and the predicted Cbl proteins from various species suggests that the three mammalian cbl genes arose by two duplications of an ancestral gene. The genomic organization of the cbl genes from various species provides insight into the evolution of the cbl gene family.

Amino Acid Sequence↗

Functional mapping of cannabinoid receptor homologs in mammals, other vertebrates, and invertebrates.

Over the past decade, several putative homologs of cannabinoid receptors (CBRs) have been identified by homology screening. Homology screening utilizes sequence alignment search engines to recognize homologs. We investigated these putative CBR homologs further by 'functional mapping' of their deduced amino acid sequences. The entire pharmacophore of a CBR has not yet been elucidated, but point-mutation studies have identified over 20 amino acid residues that impart CBR specificity for ligand recognition and/or signal transduction. Twenty point-mutation studies were used to construct a CBR functionality matrix. Sixteen putative CBR homologs were then mapped over the matrix. Several putative homologs did not hold up to this analysis: human GPR3, GPR6, GPR12, and Caenorhabditis elegans C02H7.2 expressed a series of crippling substitutions in the matrix, strongly suggesting they do not encode functional CBRs. Mapping the contested leech (Hirudo medicinalis) CBR sequence suggests that it encodes a functional CB1; it expresses fewer substitutions than the sea squirt (Ciona intestinalis) CB1 sequence. Mapping a putative CB2 ortholog in the puffer fish (Fugu rubripes T012234) suggests it may encode a CBR other than CB2. These findings are consistent with the lack of experimental data proving these putative CBRs have affinity for cannabinoid ligands. Matrix analysis also reveals that SR144528, a 'CB2-specific' synthetic antagonist, has affinity for non-mammalian CB1 receptors, and that L3.45 appears to be CB2-specific, its cognate in CB1 receptors is F3.45. In conclusion, functional mapping, utilizing point-mutation studies, may improve the specificity of homology screening performed by sequence alignment search engines.

Amino Acid Sequence↗

Compensation for nucleotide bias in a genome by representation as a discrete channel with noise.

MOTIVATION: Calculation of the information content of motifs in genomes highly biased in nucleotide composition is likely to lead to overestimates of the amount of useful information in the motif. Calculating relative information can compensate for biases, however the resulting information content is the amount seen by an observer and not by a macromolecule binding to the motif. The latter is needed to calculate the discriminatory power of the motif and to compare motifs between species. RESULTS: By treating a biased genome as a discrete channel with noise, in accordance with Shannon Information Theory, we were able to remove both 'Distortion' and 'Noise' from the motif and recover a more instructive biological 'signal.' A Java application, LogoPaint, was developed to remove nucleotide bias distortion and triplet frequency noise from motifs, calculate information content and present the motif as a logo. We demonstrate how this technique can 'unmask' motifs in the translation initiation regions of bacteria that are obscured by strong sequence biases. AVAILABILITY: LogoPaint is available to all users from the authors as an executable JAR file. Source code is available by arrangement.

Algorithms↗

A boosting approach for motif modeling using ChIP-chip data.

MOTIVATION: Building an accurate binding model for a transcription factor (TF) is essential to differentiate its true binding targets from those spurious ones. This is an important step toward understanding gene regulation. RESULTS: This paper describes a boosting approach to modeling TF-DNA binding. Different from the widely used weight matrix model, which predicts TF-DNA binding based on a linear combination of position-specific contributions, our approach builds a TF binding classifier by combining a set of weight matrix based classifiers, thus yielding a non-linear binding decision rule. The proposed approach was applied to the ChIP-chip data of Saccharomyces cerevisiae. When compared with the weight matrix method, our new approach showed significant improvements on the specificity in a majority of cases.

Algorithms↗

Adaptive diversification of bitter taste receptor genes in Mammalian evolution.

The diversity and evolution of bitter taste perception in mammals is not well understood. Recent discoveries of bitter taste receptor (T2R) genes provide an opportunity for a genetic approach to this question. We here report the identification of 10 and 30 putative T2R genes from the draft human and mouse genome sequences, respectively, in addition to the 23 and 6 previously known T2R genes from the two species. A phylogenetic analysis of the T2R genes suggests that they can be classified into three main groups, which are designated A, B, and C. Interestingly, while the one-to-one gene orthology between the human and mouse is common to group B and C genes, group A genes show a pattern of species- or lineage-specific duplication. It is possible that group B and C genes are necessary for detecting bitter tastants common to both humans and mice, whereas group A genes are used for species-specific bitter tastants. The analysis also reveals that phylogenetically closely related T2R genes are close in their chromosomal locations, demonstrating tandem gene duplication as the primary source of new T2Rs. For closely related paralogous genes, a rate of nonsynonymous nucleotide substitution significantly higher than the rate of synonymous substitution was observed in the extracellular regions of T2Rs, which are presumably involved in tastant-binding. This suggests the role of positive selection in the diversification of newly duplicated T2R genes. Because many natural poisonous substances are bitter, we conjecture that the mammalian T2R genes are under diversifying selection for the ability to recognize a diverse array of poisons that the organisms may encounter in exploring new habitats and diets.

Amino Acid Sequence↗

Overlapping translation of nucleic acid sequences for bioinformatics applications.

SUMMARY: An alternative method to TblastX has been developed. Nucleic acids in database and query sequences were translated into overlapping protein-like sequences (overlappingly translated sequences or OTSs) before searching with BlastP. Thus, each nucleic acid sequences is represented by a single 'protein like' sequence instead of three 'proteins' in different reading frames. The 3x3 comparison of TblastX is represented by a single comparison, giving faster results. Additional advantages are: (1) it can be more sensitive to detect weak sequence similarities than either blastN or TblastX; (2) codon redundancy is eliminated; (3) the sensitivity to single nucleotide polymorphism, mutation and sequencing errors is reduced; (4) it is insensitive to frame shifts. RESULTS: BlastP using OTS detected about two thirds of blastN and TblastX matches but discovered additional similarities. When blastN and TblastX against nucleic acids were compared to blastP against OTS, identical matches discovered by blastP were generally longer (602, respectively. 213 letters, p<0.01), had higher scores (748 respectively 460 bits, p<0.05) and lower E values (3.16E-20 vs. 1.17E+03, p<0.01) but the percentage identity was lower (25% respectively 61%, p<0.001). A qualitative evaluation with LALIGN showed an improvement of the visualization when OTS-s were used instead of nucleic acids. Many extensive sequence similarities became better visible, for example the repeating similarity between prion protein and human insulin gene micro-satellite, and the surprising similarity between the first part of prion protein coding region and the human pro-insulin (34.4% identity and additional 17.2% similarity through 238 residues, score >295 which is expected 4.6e-18 times by chance).

Amino Acid Sequence↗

The SSX gene family: characterization of 9 complete genes.

Human SSX genes comprise a gene family with 6 known members. SSX1, 2 and 4 have been found to be involved in the t(X;18) translocation characteristically found in all synovial sarcomas. Four (SSX1, 2, 4 and 5) are known to be expressed in a subset of tumors and testis, and anti-SSX antibodies have been found in sera from cancer patients. SSX antigens are thus typical cancer-testis (CT) antigens. To identify additional SSX family members, we isolated and characterized human genomic clones homologous to a prototype SSX cDNA. We also searched public databases for sequences similar to SSX. This identified 3 additional SSX genes, SSX7, 8, 9, and also completed the sequence of the formerly partially defined SSX6 gene. In addition to these novel SSX genes, several SSX pseudogenes were identified. With the exception of 1 pseudogene, all SSX genomic SSX sequences map to chromosome X. Among normal tissues, SSX7 mRNA was present only in testis, whereas SSX6, 8 and 9 were not detected in any normal tissue. SSX6 and 7 were expressed in 1 of 9 melanoma cell lines tested, whereas SSX8 and 9 expression was not detected in any tumor tissue or cell lines tested. SSX1, 2, 4 and 5 mRNA expression can be induced in cell lines by 5-aza-2-deoxycytidine or Trichostatin A. These agents also induce SSX6, but not SSX3, 7, 8 or 9 in the tumor cell lines tested, indicating that mechanisms other than methylation or histone acetylation may be responsible for the repressed state of some SSX genes.

Amino Acid Sequence↗

Reconstruction of ancestral protosplice sites.

Most of the eukaryotic protein-coding genes are interrupted by multiple introns. A substantial fraction of introns occupy the same position in orthologous genes from distant eukaryotes, such as plants and animals, and consequently are inferred to have been inherited from the common ancestor of these organisms. In contrast to these conserved introns, many other introns appear to have been gained during evolution of each major eukaryotic lineage. The mechanism(s) of insertion of new introns into genes remains unknown. Because the nucleotides that flank splice junctions are nonrandom, it has been proposed that introns are preferentially inserted into specific target sequences termed protosplice sites. However, it remains unclear whether the consensus nucleotides flanking the splice junctions are remnants of the original protosplice sites or if they evolved convergently after intron insertion. Here, we directly address the existence of protosplice sites by examining the context of introns inserted within codons that encode amino acids conserved in all eukaryotes and accordingly are not subject to selection for splicing efficiency. We show that introns are either predominantly inserted into specific protosplice sites, which have the consensus sequence (A/C)AG/Gt, or that they are inserted randomly but are preferentially fixed at such sites.

Amino Acid Sequence↗

Investigation of protein functions through data-mining on integrated human transcriptome database, H-Invitational database (H-InvDB).

H-Invitational Database (H-InvDB; ) is a human transcriptome database, containing integrative annotation of 41,118 full-length cDNA clones originated from 21,037 loci. H-InvDB is a product of the H-Invitational project, an international collaboration to systematically and functionally validate human genes by analysis of a unique set of high quality full-length cDNA clones using automatic annotation and human curation under unified criteria. Here, 19,574 proteins encoded by these cDNAs were classified into 11,709 function-known and 7865 function-unknown hypothetical proteins by similarity with protein databases and motif prediction (InterProScan). The proportion of "hypothetical proteins" in H-InvDB was as high as 40.4%. In this study, we thus conducted data-mining in H-InvDB with the aim of assigning advanced functional annotations to those hypothetical proteins. First, by data-mining in the H-InvDB version of GTOP, we identified 337 SCOP domains within 7865 H-Inv hypothetical proteins. Second, by data-mining of predicted subcellular localization by SOSUI and TMHMM in H-InvDB, we found 1032 transmembrane proteins within H-Inv hypothetical proteins. These results clearly demonstrate that structural prediction is effective for functional annotation of proteins with unknown functions. All the data in H-InvDB are shown in two main views, the cDNA view and the Locus view, and five auxiliary databases with web-based viewers; DiseaseInfo Viewer, H-ANGEL, Clustering Viewer, G-integra and TOPO Viewer; the data also are provided as flat files and XML files. The data consists of descriptions of their gene structures, novel alternative splicing isoforms, functional RNAs, functional domains, subcellular localizations, metabolic pathways, predictions of protein 3D structure, mapping of SNPs and microsatellite repeat motifs in relation with orphan diseases, gene expression profiling, and comparisons with mouse full-length cDNAs in the context of molecular evolution. This unique integrative platform for conducting in silico data-mining represents a substantial contribution to resources required for the exploration of human biology and pathology.

Amino Acid Sequence↗

BmiGI: a database of cDNAs expressed in Boophilus microplus, the tropical/southern cattle tick.

We used an expressed sequence tag approach to initiate a study of the genome of the southern cattle tick, Boophilus microplus. A normalized cDNA library was synthesized from pooled RNA purified from tick larvae which had been subjected to different treatments, including acaricide exposure, heat shock, cold shock, host odor, and infection with Babesia bovis. For the acaricide exposure experiments, we used several strains of ticks, which varied in their levels of susceptibility to pyrethroid, organophosphate and amitraz. We also included RNA purified from samples of eggs, nymphs and adult ticks and dissected tick organs. Plasmid DNA was prepared from 11,520 cDNA clones and both 5' and 3' sequencing performed on each clone. The sequence data was used to search public protein databases and a B. microplus gene index was constructed, consisting of 8270 unique sequences whose associated putative functional assignments, when available, can be viewed at the TIGR website (http://www.tigr.org/tdb/tgi). A number of novel sequences were identified which possessed significant sequence similarity to genes, which might be involved in resistance to acaricides.

Acetylcholinesterase↗

Reporting and representation of population descriptors in public RNA-seq databases.

Diverse and globally representative datasets are essential to genomic science and medicine. Here, we analyzed population descriptor metadata from RNA sequencing (RNA-seq) studies in two major public repositories: the Sequence Read Archive (SRA) and the Database of Genotypes and Phenotypes. We examined geographic and economic characteristics of institutions depositing the data and compared SRA-deposited descriptors to empirical estimates of genetic ancestry and to those reported in publications, analyzing trends over time. We found that 55% of RNA-seq samples were deposited by United States (US) institutions and 90% by institutions in high-income countries. Only 3% of SRA samples were associated with population descriptors, and among those with US Census terms, 69% were labeled as White. Among samples with continental descriptors, 56% were labeled as European. Our analyses emphasize widespread bias in the composition of public RNA-seq datasets and, more generally, a lack of consistent and careful reporting of population descriptors needing urgent improvement.

Humans↗

VIPL, a VIP36-like membrane protein with a putative function in the export of glycoproteins from the endoplasmic reticulum.

Subsets of glycoproteins are thought to require lectin-like membrane receptors for efficient export out of the endoplasmic reticulum (ER). To identify new members related to two previously characterized intracellular lectins ERGIC-53/p58 and VIP36, we carried out an extensive database search using the conserved carbohydrate recognition domain (CRD) as a search string. A gene, more closely related to VIP36 than to ERGIC-53/p58, and hence called VIPL (VIP36-Like), was identified. VIPL has been conserved through evolution from zebra fish to man. The 2.4-kb VIPL mRNA was widely expressed to varying levels in different tissues. Using an antiserum prepared against the CRD, the 32-kDa VIPL protein was detected in various cell lines. The single N-linked glycan of VIPL remained endoglycosidase H-sensitive during a 2-h pulse-chase, even when the protein was overexpressed or mutated to allow export to the plasma membrane. VIPL localized primarily to the ER and partly to the Golgi complex. Like VIP36, the cytoplasmic tail of VIPL terminates in the sequence KRFY, a motif characteristic for proteins recycling between the ER and ERGIC/cis-Golgi. Mutating the retrograde transport signal KR to AA resulted in transport of VIPL to the cell surface. Finally, knock-down of VIPL mRNA using siRNA significantly slowed down the secretion of two glycoproteins (M(R) 35 and 250 kDa) to the medium, suggesting that VIPL may also function as an ER export receptor.

Amino Acid Sequence↗

Conserved patterns of protein interaction in multiple species.

To elucidate cellular machinery on a global scale, we performed a multiple comparison of the recently available protein-protein interaction networks of Caenorhabditis elegans, Drosophila melanogaster, and Saccharomyces cerevisiae. This comparison integrated protein interaction and sequence information to reveal 71 network regions that were conserved across all three species and many exclusive to the metazoans. We used this conservation, and found statistically significant support for 4,645 previously undescribed protein functions and 2,609 previously undescribed protein interactions. We tested 60 interaction predictions for yeast by two-hybrid analysis, confirming approximately half of these. Significantly, many of the predicted functions and interactions would not have been identified from sequence similarity alone, demonstrating that network comparisons provide essential biological information beyond what is gleaned from the genome.

Amino Acid Sequence↗