PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Data protection in biomaterial banks for Parkinson's disease research: the model of GEPARD (Gene Bank Parkinson's Disease Germany).

Parkinson's disease (PD) is the second most common neurodegenerative disease. Although 10 gene loci have been identified to cause a Parkinsonian syndrome, these loci account only for a minority of PD patients. Large, systematic research programs are required to collect, store, and analyze DNA samples and clinical information to support further discovery of additional genetic components of PD or other movement disorders. Such programs facilitate research into the relationship between genotype and phenotype. The German Competence Network on Parkinson's disease (CNP) initiated the Gene Bank Parkinson's Disease Germany (GEPARD), providing an administrative and scientific infrastructure for the storage of DNA and clinical data that are electronically accessible and protective of patient rights. In this article, we offer guidance on how to establish a framework for a clinical genetic data and DNA bank, and describe GEPARD as a model that may be useful to other local, national, and international research groups developing similar programs.

Computer Security↗

Transcriptome of mouse uterus by serial analysis of gene expression (SAGE): comparison with skeletal muscle.

The aim of this study was to identify the transcriptome of the normal mouse uterus by Serial Analysis of Gene Expression method. mRNA was extracted from the uterus and also from the gastrocnemius muscle of mice. Short sequences (tags), each one usually corresponding to a distinct transcript, were isolated and concatemerized into long DNA molecules which were cloned and sequenced. We detected 44,484 tags for the uterus and 42,518 tags for the muscle, representing 14,543 and 14,958 potential transcript species, respectively. Seventy-five and sixty-nine genes were expressed at more than 0.1%, thus corresponding to 37 and 34% of the mRNA population detected in the respective tissues. In both cases, the most highly expressed genes are especially involved in muscle contraction, energy metabolism, and protein synthesis. Compared to skeletal muscle, some differentially expressed genes in the uterus are likely to correspond to its specific reproductive functions. The majority of these genes remain to be characterized. More than 70% of the different tags detected in the uterus did not match any sequence in the public databases and can represent novel or poorly identified genes. This study is the first quantitative description of the transcriptome of the uterus.

Animals↗

Protein expression levels of the Src activating protein AFAP are developmentally regulated in brain.

The Src family of nonreceptor tyrosine kinases plays an important role in modulating signals that affect growth cone extension, neuronal differentiation, and brain development. Recent reports indicate that the Src SH2/SH3 binding partner AFAP-110 has the capacity to modulate actin filament integrity as a cSrc activating protein and as an actin filament bundling protein. Both AFAP-110 and a brain specific isoform called AFAP-120 (collectively referred to as AFAP) exist at high levels in chick embryo brain. We sought to identify the localization of AFAP in mouse brain in order to identify its expression pattern and potential role as a cellular modulator of Src family kinase activity and actin filament integrity in the brain. In E16 mouse embryos, AFAP expression levels were very high and concentrated in the olfactory bulb, cortex, forebrain, cerebellum, and various peripheral sensory structures. In P3 mouse pups, overall expression was reduced compared to E16 embryos, and AFAP was found primarily in olfactory bulb, cortex, and cerebellum. AFAP expression levels were significantly reduced in adult mice, with high expression levels only detected in the olfactory bulb. Western blot analysis indicated that concentrated expression of AFAP correlates well with the AFAP-120 isoform, which appears to be a splice variant of AFAP-110. As the expression pattern of AFAP overlaps with the reported expression patterns of cSrc and Fyn, we hypothesize that AFAP is positioned to modulate signal transduction cascades that direct activation of these nonreceptor tyrosine kinases and concomitant cellular changes that occur in actin filaments during brain development.

Adaptor Proteins, Vesicular Transport↗

Identification of expressed sequence tags preferentially expressed in human placentas by in silico subtraction.

OBJECTIVES: To identify expressed sequence tag (EST) clusters preferentially expressed in placentas. METHODS: The National Center for Biotechnology's online UniGene database contains 14 placenta libraries. In silico (computer-based) subtraction compared placenta libraries against the remaining libraries to identify transcripts preferentially expressed in placentas. For known genes, placental expression or their use in prenatal diagnosis was then explored online using LocusLink and PubMed. RESULTS: Placentas preferentially expressed 475 EST clusters. Of these, 18 EST clusters with no known function were expressed exclusively in placentas. Of the remaining 457 EST clusters, 90 showed preferential placental expression by >/=25 times. Of these 90, literature searches on the 45 EST clusters with known functions showed 44 linked to placental physiology or proposed as markers for prenatal diagnosis [i.e. beta-hCG, pregnancy-specific glycoproteins, human placental lactogens, pregnancy-associated plasma protein A (PAPP-A)]. Selected genes with known function in pregnancy but whose preferential placental expression fell below the factor of 25 threshold were also identified. CONCLUSION: In silico subtraction identified 44 previously studied genes involved in placental physiology as well as 63 EST clusters preferentially expressed in placental tissue, which may serve as targets for future studies seeking novel markers for prenatal diagnosis or to better understand placental genetics.

Adult↗

Pharmacogenetics and pharmacoepidemiology.

This paper seeks to stimulate consideration of the short- and long-term possibilities raised by research into pharmacogenetics and pharmacoepidemiology, to identify potential tissue databanks linked to population data, and to summarize our ethical responsibilities. Short term, we might identify selected disorders or risk factors, such as for drug-associated serious events, and might find selected target populations. Long term, we might incorporate genomics techniques into population data and map genetic risk factors for diseases and drug responses. Tissue specimens that might be linked to clinical data lie in certain data resources: clinical trials, ad hoc cohorts, case registries, and cross-sectional and longitudinal population samples. Specific areas must be addressed in research using human tissue and genetic material, including the ethical threat to privacy. We must avoid misuse of genetic information by breaching confidentiality and putting at risk specific genetically identifiable groups. The National Bioethics Advisory Commission recently addressed the issue of serum and tissue data banks in a published report summarized in this paper. Social, methodologic, and ethical issues relate to any protocol using serum and tissue, and indicate a need for broad educational efforts. We must set guidelines as pharmacogenetics research becomes a dimension of pharmacoepidemiology.

Cohort Studies↗

Error-tolerant EST database searches by tandem mass spectrometry and multiTag software.

The MultiTag method (Sunyaev et al., Anal. Chem. 2003 15, 1307-1315) employs multiple error-tolerant searches with peptide sequence tags (Mann and Wilm, Anal. Chem. 1994, 66, 4390-4399) for the identification of proteins from organisms with unsequenced genomes. Here we demonstrate that the error-tolerant capabilities of MultiTag increased the number of peptide alignments and improved the confidence of identifications in an EST database. The MultiTag outperformed conventional database searching software that only utilizes stringent matching of tandem mass spectra to nucleotide sequences of ESTs.

Algorithms↗

Saccharomyces cerevisiae S288C genome annotation: a working hypothesis.

The S. cerevisiae genome is the most well-characterized eukaryotic genome and one of the simplest in terms of identifying open reading frames (ORFs), yet its primary annotation has been updated continually in the decade since its initial release in 1996 (Goffeau et al., 1996). The Saccharomyces Genome Database (SGD; www.yeastgenome.org) (Hirschman et al., 2006), the community-designated repository for this reference genome, strives to ensure that the S. cerevisiae annotation is as accurate and useful as possible. At SGD, the S. cerevisiae genome sequence and annotation are treated as a working hypothesis, which must be repeatedly tested and refined. In this paper, in celebration of the tenth anniversary of the completion of the S. cerevisiae genome sequence, we discuss the ways in which the S. cerevisiae sequence and annotation have changed, consider the multiple sources of experimental and comparative data on which these changes are based, and describe our methods for evaluating, incorporating and documenting these new data.

Base Sequence↗

Molecular cloning of novel mouse and human putative citrate lyase beta-subunit.

Using a fluorescent differential display (FDD) technique, a novel cDNA was identified by screening for gene expressed differentially between the Dunn osteosarcoma cell line and the LM8 cell line, an isolated variant of the Dunn cell line that has high metastatic potential to the lung. Molecular cloning of the cDNA revealed the clone has similarity to a bacterial fermentation enzyme, the citrate lyase beta-subunit (CL-beta). Northern blot and competitive reverse transcription-PCR (RT-PCR) analysis revealed up-regulation of the gene in the LM8 cell line. An RNA Master blot indicated that the mRNA encoding CL-beta is expressed abundantly in murine heart, liver, and kidney. A human expressed sequence tag (EST) database search suggested that a similar cDNA is expressed in humans. A gene with identical sequence is located on chromosome 13 in the genome database (Sanger centre, UK). These data suggest that a citrate fermentation pathway may exist in eukaryotes including mammals.

Animals↗

Search for genes positively selected during primate evolution by 5'-end-sequence screening of cynomolgus monkey cDNAs.

It is possible to assess positive selection by using the ratio of K(a) (nonsynonymous substitutions per plausible nonsynonymous sites) to K(s) (synonymous substitutions per plausible synonymous sites). We have searched candidate genes positively selected during primate evolution by using 5'-end sequences of 21,302 clones derived from cynomolgus monkey (Macaca fascicularis) brain cDNA libraries. Among these candidates, 10 genes that had not been shown by previous studies to undergo positive selection exhibited a K(a)/K(s) ratio > 1. Of the 10 candidate genes we found, 5 were included in the mitochondrial respiratory enzyme complexes, suggesting that these nuclear-encoded genes coevolved with mitochondrial-encoded genes, which have high mutation rates. The products of other candidate genes consisted of a cell-surface protein, a member of the lipocalin family, a nuclear transcription factor, and hypothetical proteins.

Animals↗

Many human genes are transcribed from the antisense promoter of L1 retrotransposon.

Human L1 retrotransposon has two transcription-regulatory regions: an internal or sense promoter driving transcription of the full-length L1, and an antisense promoter (ASP) driving transcription in the opposite direction into adjacent cellular sequences yielding chimeric transcripts. Both promoters are located in the 5'-untranslated region (5'-UTR) of L1. Chimeric transcripts derived from the L1 ASP are highly represented in expressed-sequence tag (EST) databases. Using a bioinformatics approach, we have characterized 10 chimeric ESTs (cESTs) derived from the EST division of GenBank. These cESTs contained 3' regions similar or identical to known cellular mRNA sequences. They were accurately spliced and preferentially expressed in tumor cell lines. Analysis of the hundreds of cESTs suggests that the L1 ASP-driven transcription is a common phenomenon not only for tumor cells but also for normal ones and may involve transcriptional interference or epigenetic control of different cellular genes.

5' Flanking Region↗

PipTools: a computational toolkit to annotate and analyze pairwise comparisons of genomic sequences.

Sequence conservation between species is useful both for locating coding regions of genes and for identifying functional noncoding segments. Hence interspecies alignment of genomic sequences is an important computational technique. However, its utility is limited without extensive annotation. We describe a suite of software tools, PipTools, and related programs that facilitate the annotation of genes and putative regulatory elements in pairwise alignments. The alignment server PipMaker uses the output of these tools to display detailed information needed to interpret alignments. These programs are provided in a portable format for use on common desktop computers and both the toolkit and the PipMaker server can be found at our Web site (http://bio.cse.psu.edu/). We illustrate the utility of the toolkit using annotation of a pairwise comparison of the mouse MHC class II and class III regions with orthologous human sequences and subsequently identify conserved, noncoding sequences that are DNase I hypersensitive sites in chromatin of mouse cells.

Animals↗

Key-string segmentation algorithm and higher-order repeat 16mer (54 copies) in human alpha satellite DNA in chromosome 7.

A new key-string segmentation algorithm for identification of alpha satellite DNAs and higher-order repeat (HOR) units was introduced and exemplified. Starting with an initial key string, we determine the dominant key string and HOR. Our key-string algorithm was used to scan the recent GenBank data for human alpha satellite DNA sequence AC017075.8 (193 277 bp) from the centromeric region of chromosome 7. The sequence was computationally segmented into one HOR domain (super-repeat domain) and two non-HOR domains. Dominant key-string GTTTCT provided segmentation in terms of alpha monomers. The HOR is tandemly repeated in 54 copies in the super-repeat (HOR) domain. Five insertions and three deletions in the HOR structure associated with a dominant key string were identified. Concensus HOR was constructed. Divergence of individual HOR copies from concensus amounts to 0.7% on the average, while divergence between 16 monomer variants within each HOR is on the average 20%. In the front and back domain, 199 monomer variants were identified that are not organized in HOR and diverge by 20-40%.

Algorithms↗

Sources of incongruence among mammalian mitochondrial sequences: COII, COIII, and ND6 genes are main contributors.

To investigate the origins of incongruence among mammalian mitochondrial protein-coding genes, we compiled a matrix that included 13 protein-coding-genes for 41 mammals from 14 different orders. This matrix was examined for congruence using different partitioning strategies. The incongruence length difference test showed significant incongruence among the 13 gene partitions used simultaneously, and the result was not affected by third codon or transversion weighting. In the pair-wise comparisons, significant incongruence was detected between NADH:ubiquinone oxidoreductase subunit 6 gene (ND6), cytochrome oxidase subunit II (COII), or cytochrome oxidase subunit III (COIII) gene partitioned individually against the rest of the genes. Omission of any of the 14 mammalian orders alone or in combinations from the matrix did not result in a statistically significant improvement of congruence, suggesting that taxonomic sampling will not improve congruence among the data sets. However, omission of the ND6, COII, and COIII significantly improved congruence in our data matrix. Possible origins of unusual phylogenetic properties of the three genes are discussed.

Animals↗

Molecular evolution of viral fusion and matrix protein genes and phylogenetic relationships among the Paramyxoviridae.

Phylogenetic relationships among the Paramyxoviridae, a broad family of viruses whose members cause devastating diseases of wildlife, livestock, and humans, were examined with both fusion (F) and matrix (M) protein-coding sequences. Neighbor-joining trees of F and M protein sequences showed that the Paramyxoviridae was divided into the two traditionally recognized subfamilies, the Paramyxovirinae and the Pneumovirinae. Within the Paramyxovirinae, the results also showed groups corresponding to three currently recognized genera: Respirovirus, Morbillivirus, and Rubulavirus. The relationships among the three genera of the Paramyxovirinae were resolved with M protein sequences and there was significant bootstrap support (100%) showing that members of the genus Respirovirus and the genus Morbillivirus were more closely related to each other than to members of the genus Rubulavirus. Both F and M phylogenies showed that Newcastle disease virus (NDV) was more closely related to the genus Rubulavirus than to the other two genera but were consistent with the proposal (B. S. Seal et al., 2000, Virus Res. 66, 1-11) that NDV be classified as a separate genus within the Paramyxovirinae. Both F and M phylogenies were also consistent with the proposal (L. Wang et al., 2000, J. Virol 74, 9972-9979) that Hendra virus be classified as a new genus closely related and basal to the genus Morbillivirus. Rinderpest was most closely related to measles and a more derived virus than to canine distemper virus, phocine distemper virus, or dolphin morbillivirus.

Databases, Nucleic Acid↗

The use of MassARRAY technology for high throughput genotyping.

This chapter will explore the role of mass spectrometry (MS) as a detection method for genotyping applications and will illustrate how MS evolved from an expert-user-technology to a routine laboratory method in biological sciences. The main focus will be time-of-flight (TOF) based devices and their use for analyzing single-nucleotide-polymorphisms (SNPs, pronounced snips). The first section will describe the evolution of the use of MS in the field of bioanalytical sciences and the protocols used during the early days of bioanalytical MALDI TOF mass spectrometry. The second section will provide an overview on intraspecies sequence diversity and the nature and importance of SNPs for the genomic sciences. This is followed by an exploration of the special and advantageous features of mass spectrometry as the key technology in modern bioanalytical sciences in the third chapter. Finally, the fourth section will describe the MassARRAY technology as an advanced system for automated high-throughput analysis of SNPs.

Databases, Nucleic Acid↗