PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Databases, Nucleic Acid”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 721 records · Page 40Linked to original sources

SNP identification, linkage disequilibrium, and haplotype analysis for a 200-kb genomic region in a Korean population.

Understanding patterns of linkage disequilibrium (LD) across genomes may facilitate association mapping studies to localize genetic variants influencing complex diseases, a recognition that led to the International Haplotype Mapping Project (HapMap). Divergent patterns of haplotype frequency and LD across global populations require that the HapMap database be supplemented with haplotype and LD data from additional populations. We conducted a pilot study of the LD and haplotype structure of a genomic region in a Korean population. A total of 165 SNPs were identified in a 200-kb region of 22q13.2 by direct sequencing. Unphased genotype data were generated for 76 SNPs in 90 unrelated Korean individuals. LD, haplotype diversity, and recombination rates were assessed in this region and compared with the HapMap database. The pattern of LD and haplotype frequencies of Korean samples showed a high degree of similarity with Japanese data. There was a strong correlation between high LD and low recombination frequency in this region. We found considerable similarities in local LD patterns between three Asian populations (Han Chinese, Japanese, and Korean) and the CEPH population. Haplotype frequencies were, however, significantly different between them. Our results should further the understanding of distinctive Korean genomic features and assist in designing appropriate association studies.

Asian People↗

ChickGCE: a novel germ cell EST database for studying the early developmental stage in chickens.

We established a database to study germ cells during the early developmental stage in the chicken. The ChickGCE database provides integrated expressed sequence tag (EST) data from chicken testis, ovary, embryonic gonads, and primordial germ cells. We gathered data on 10,294 ESTs from approximately 1000 embryonic gonads, and we experimentally determined 10,851 ESTs from primordial germ cells purified from 7955 embryonic gonads by magnetically activated cell sorting. The EST testis and ovary datasets were retrieved from the public database of The Institute for Genomic Research (TIGR). The EST data were clustered and assembled into unique sequences, contigs, and singletons. The ChickGCE database provides functional annotation, identification, and putative embryonic germ-cell-specific novel transcripts based on the Gene Ontology database, as well as statistical analyses of expression patterns and pair-wise comparisons of two types of tissue- and germ-cell-specific alternative splicing events in the chicken. The new database is accessible online and queries can be answered using several search options, including tissue database searches, keywords, clone IDs, expected values, and BLAST search scores.

Animals↗

Mutational spectrum in the recent human genome inferred by single nucleotide polymorphisms.

So far, there is no genome-wide estimation of the mutational spectrum in humans. In this study, we systematically examined the directionality of the point mutations and maintenance of GC content in the human genome using approximately 1.8 million high-quality human single nucleotide polymorphisms and their ancestral sequences in chimpanzees. The frequency of C-->T (G-->A) changes was the highest among all mutation types and the frequency of each type of transition was approximately fourfold that of each type of transversion. In intergenic regions, when the GC content increased, the frequency of changes from G or C increased. In exons, the frequency of G:C-->A:T was the highest among the genomic categories and contributed mainly by the frequent mutations at the CpG sites. In contrast, mutations at the CpG sites, or CpG-->TpG/CpA mutations, occurred less frequently in the CpG islands relative to intergenic regions with similar GC content. Our results suggest that the GC content is overall not in equilibrium in the human genome, with a trend toward shifting the human genome to be AT rich and shifting the GC content of a region to approach the genome average. Our results, which differ from previous estimates based on limited loci or on the rodent lineage, provide the first representative and reliable mutational spectrum in the recent human genome and categorized genomic regions.

Animals↗

Evolutionary dynamics of the ABCA chromosome 17q24 cluster genes in vertebrates.

ABCA is a subfamily of ATP-binding-cassette (ABC) transporter genes. In this subfamily, it was found that five ABCA genes cluster in a head-to-tail pattern in the human and mouse genomes, but only one was found in fish. To understand better the evolution of this cluster of genes, we screened 11 vertebrate genome sequences and newly identified 28 ABCA cluster genes. Comparative genomic analysis reveals that the ABCA5 gene is relatively evolutionarily conserved. In contrast, the repertoires of the other ABCA genes in this cluster diverge tremendously among species, which is due mainly to postspeciation duplications. In addition, maximum likelihood analysis reveals that positive selection is acting on the paralogous genes ABCA6 and Abca8a, suggesting that these two genes have possibly acquired new functions after duplication. Because most eukaryotic ABC proteins integrate into the cytoplasmic membrane and transport a wide range of substrates across it, we conjecture that newly duplicated ABCA cluster genes are under diversifying selection for the ability to recognize a diverse array of substrates.

ATP-Binding Cassette Transporters↗

hORFeome v3.1: a resource of human open reading frames representing over 10,000 human genes.

Complete sets of cloned protein-encoding open reading frames (ORFs), or ORFeomes, are essential tools for large-scale proteomics and systems biology studies. Here we describe human ORFeome version 3.1 (hORFeome v3.1), currently the largest publicly available resource of full-length human ORFs (available at ). Generated by Gateway recombinational cloning, this collection contains 12,212 ORFs, representing 10,214 human genes, and corresponds to a 51% expansion of the original hORFeome v1.1. An online human ORFeome database, hORFDB, was built and serves as the central repository for all cloned human ORFs (http://horfdb.dfci.harvard.edu). This expansion of the original ORFeome resource greatly increases the potential experimental search space for large-scale proteomics studies, which will lead to the generation of more comprehensive datasets.

Animals↗

Textmining in support of knowledge discovery for vaccine development.

Complete genome data of infectious microorganisms permit systematic computational sequence-based predictions and experimental testing of candidate vaccine epitopes. Both, predictions and the interpretation of experiments rely on existing information in the literature which is mostly manually extracted and curated. The growing amount of data and literature information has created a major bottleneck for the interpretation of results and maintenance of curated databases. The lack of suitable free-text information extraction, processing, and reporting tools prompted us to develop a knowledge discovery support system that enhances the understanding of immune response and vaccine development. The current prototype system, Gene expression/epitpopes/protein interaction (GEpi), focuses on molecular functions of HIV-infected T-cells and HIV epitope information, using textmining, and interrelation of biomolecular data from domain-specific databases with MEDLINE abstract-inferred information. Results showed that extraction and processing of molecular interaction, disease associations, and gene ontology-derived functional information generate intuitive knowledge reports that aid the interpretation of host-pathogen interaction. In contrast, epitope (word and sequence) information in MEDLINE abstracts is surprisingly sparse and often lacks necessary context information, such as HLA-restriction. Since the majority of epitope information is found in tables, figures, and legends of full-text articles, its extraction may not require sophisticated natural language processing techniques. Support of vaccine development through textmining requires therefore the timely development of domain-specific extraction rules for full-text articles, and a knowledge model for epitope-related information.

Animals↗

Molecular systematics of Salmonidae: combined nuclear data yields a robust phylogeny.

The phylogeny of salmonid fishes has been the focus of intensive study for many years, but some of the most important relationships within this group remain unclear. We used 269 Genbank sequences of mitochondrial DNA (from 16 genes) and nuclear DNA (from nine genes) to infer phylogenies for 30 species of salmonids. We used maximum parsimony and maximum likelihood to analyze each gene separately, the mtDNA data combined, the nuclear data combined, and all of the data together. The phylogeny with the best overall resolution and support from bootstrapping and Bayesian analyses was inferred from the combined nuclear DNA data set, for which the different genes reinforced and complemented one another to a considerable degree. Addition of the mitochondrial DNA degraded the phylogenetic signal, apparently as a result of saturation, hybridization, selection, or some combination of these processes. By the nuclear-DNA phylogeny: (1) (Hucho hucho, Brachymystax lenok) form the sister group to (Salmo, Salvelinus, Oncorhynchus, H. perryi); (2) Salmo is the sister-group to (Oncorhynchus, Salvelinus); (3) Salvelinus is the sister-group to Oncorhynchus; and (4) Oncorhynchus masou forms a monophyletic group with O. mykiss and O. clarki, with these three taxa constituting the sister-group to the five other Oncorhynchus species. Species-level relationships within Oncorhynchus and Salvelinus were well supported by bootstrap levels and Bayesian analyses. These findings have important implications for understanding the evolution of behavior, ecology and life-history in Salmonidae.

Animals↗

Phylogenetic relationships of South American lizards of the genus Stenocercus (Squamata: Iguania): A new approach using a general mixture model for gene sequence data.

The South American iguanian lizard genus Stenocercus includes 54 species occurring mostly in the Andes and adjacent lowland areas from northern Venezuela and Colombia to central Argentina at elevations of 0-4000m. Small taxon or character sampling has characterized all phylogenetic analyses of Stenocercus, which has long been recognized as sister taxon to the Tropidurus Group. In this study, we use mtDNA sequence data to perform phylogenetic analyses that include 32 species of Stenocercus and 12 outgroup taxa. Monophyly of this genus is strongly supported by maximum parsimony and Bayesian analyses. Evolutionary relationships within Stenocercus are further analyzed with a Bayesian implementation of a general mixture model, which accommodates variability in the pattern of evolution across sites. These analyses indicate a basal split of Stenocercus into two clades, one of which receives very strong statistical support. In addition, we test previous hypotheses using non-parametric and parametric statistical methods, and provide a phylogenetic classification for Stenocercus.

Animals↗

Phylogeny of the Callandrena subgenus of Andrena (Hymenoptera: Andrenidae) based on mitochondrial and nuclear DNA data: polyphyly and convergent evolution.

We propose a phylogenetic hypothesis of relationships within Callandrena, a North American subgenus of the bee genus Andrena, based on both mitochondrial and nuclear DNA sequences. Our data included 695 aligned base pairs comprising parts of the mitochondrial genes cytochrome oxidase subunits I and II and the intervening tRNA-leucine and 767 aligned base pairs of the F2 copy of the nuclear gene elongation factor-1alpha. We also suggest a preliminary hypothesis of relationships of the North American subgenera in the genus. Our analyses included 54 species of Callandrena, 42 species of Andrena representing 24 additional subgenera, and 11 outgroup species in the family Andrenidae. Parsimony analyses of each marker separately suggested that Callandrena was polyphyletic, with a combined analysis suggesting that there were at least two phylogenetically independent clades of bees with similar morphological features. Maximum likelihood and Bayesian analyses supported this conclusion, as did the non-parametric bootstrapping SOWH test. Convergence in morphological characters was likely due to their common use of members of Asteraceae as pollen hosts.

Animals↗

Molecular phylogeny of the Arctoidea (Carnivora): effect of missing data on supertree and supermatrix analyses of multiple gene data sets.

Phylogenetic relationships of 79 caniform carnivores were addressed based on four nuclear sequence-tagged sites (STS) and one nuclear exon, IRBP, using both supertree and supermatrix analyses. We recovered the three major arctoid lineages, Ursidae, Pinnipedia, and Musteloidea, as monophyletic, with Ursidae (bears) strongly supported as the basal arctoid lineage. Within Pinnipedia, Phocidae (true seals) were sister to the Otaroidea [Otariidae (fur seals and sea lions) and Odobenidae (walrus)]. Phocid subfamily and tribal designations were supported, but the otariid subfamily split between fur seals and sea lions was not. All family designations within Musteloidea were strongly supported: Mephitidae (skunks), Ailuridae (monotypic red panda), Mustelidae (weasels, badgers, otters), and Procyonidae (raccoons). A novel hypothesis for the position of the red panda was recovered, placing it as branching after Mephitidae and before Mustelidae+Procyonidae. Within Mustelidae, subfamily taxonomic changes are considered. This study represents the most comprehensive sampling to date of the Caniformia in a molecular study and contains the most complete molecular phylogeny for the Procyonidae. Our data set was also used in an empirical examination of the effect of missing data on both supertree and supermatrix analyses. Sequence for all genes in all taxa could not be obtained, so two variants of the data set with differing amounts of missing data were examined. The amount of missing data did not have a strong effect; instead, phylogenetic resolution was more dependent on the presence of sufficient informative characters. Supertree and supermatrix methods performed equivalently with incomplete data and were highly congruent; conflicts arose only in weakly supported areas, indicating that more informative characters are required to confidently resolve close species relationships.

Animals↗

Distinguishing gorilla mitochondrial sequences from nuclear integrations and PCR recombinants: guidelines for their diagnosis in complex sequence databases.

Nuclear integrations of mitochondrial DNA (Numts) are widespread in many taxa and if left undetected can confound phylogeny interpretation and bias estimates of mitochondrial DNA (mtDNA) diversity. This is particularly true in gorillas, where recent studies suggest multiple integrations of the first hypervariable (HV1) domain of the mitochondrial control region. Problems can also arise through the inadvertent incorporation of artifacts produced by in vitro recombination between sequence types during polymerase chain reaction amplification. This issue has attracted little attention yet could potentially exacerbate errors in databases already contaminated by Numts. Using a set of existing diagnostic tools, this study set out to systematically inventory Numts and PCR recombinants in a gorilla HV1 sequence database and address the degree to which existing public databases are contaminated. Phylogenetic analysis revealed three distinct gorilla HV1 Numt groups (I, II, and III) that could be readily differentiated from mtDNA sequences by Numt-specific diagnostic sites and sequence-based motifs. Several instances of genuine recombination were also identified by a suite of detection methods. The location of putative breakpoints was identified by eye and by likelihood analysis. Findings from this study reveal widespread nuclear contamination of gorilla HV1 GenBank databases and underline the importance of recognizing not only Numts but also PCR recombinant artifacts as potential sources of data contamination. Guidelines for the routine identification of Numts and in vitro recombinants are presented and should prove useful in the detection of similar artifacts in other species mtDNA databases.

Animals↗

Formation of new genes explains lower intron density in mammalian Rhodopsin G protein-coupled receptors.

Mammalian G protein-coupled receptor (GPCR) genes are characterised by a large proportion of intronless genes or a lower density of introns when compared with GPCRs of invertebrates. It is unclear which mechanisms have influenced intron density in this protein family, which is one of the largest in the mammalian genomes. We used a combination of Hidden Markov Models (HMM) and BLAST searches to establish the comprehensive repertoire of Rhodopsin GPCRs from seven species and performed overall alignments and phylogenetic analysis using the maximum parsimony method for over 1400 receptors in 12 subgroups. We identified 14 different Ancestral Receptor Groups (ARGs) that have members in both vertebrate and invertebrate species. We found that there exists a remarkable difference in the intron density among ancestral and new Rhodopsin GPCRs. The intron density among ARGs members was more than 3.5-fold higher than that within non-ARG members and more than 2-fold higher when considering only the 7TM region. This suggests that the new GPCR genes have been predominantly formed intronless while the ancestral receptors likely accumulated introns during their evolution. Many of the intron positions found in mammalian ARG receptor sequences were found to be present in orthologue invertebrate receptors suggesting that these intron positions are ancient. This analysis also revealed that one intron position is much more frequent than any other position and it is common for a number of phylogenetically different Rhodopsin GPCR groups. This intron position lies within a functionally important, conserved, DRY motif which may form a proto-splice site that could contribute to positional intron insertion. Moreover, we have found that other receptor motifs, similar to DRY, also contain introns between the second and third nucleotide of the arginine codon which also forms a proto-splice site. Our analysis presents compelling evidence that there was not a major loss of introns in mammalian GPCRs and formation of new GPCRs among mammals explains why these have fewer introns compared to invertebrate GPCRs. We also discuss and speculate about the possible role of different RNA- and DNA-based mechanisms of intron insertion and loss.

Animals↗

Analysis of orthologous gene expression between human pulmonary adenocarcinoma and a carcinogen-induced murine model.

Human adenocarcinoma (AC) is the most frequently diagnosed human lung cancer, and its absolute incidence is increasing dramatically. Compared to human lung AC, the A/J mouse-urethane model exhibits similar histological appearance and molecular changes. We examined the gene expression profiles of human and murine lung tissues (normal or AC) and compared the two species' datasets after aligning approximately 7500 orthologous genes. A list of 409 gene classifiers (P value <0.0001), common to both species (joint classifiers), showed significant, positive correlation in expression levels between the two species. A number of previously reported expression changes were recapitulated in both species, such as changes in glycolytic enzymes and cell-cycle proteins. Unexpectedly, joint classifiers in angiogenesis were uniformly down-regulated in tumor tissues. The eicosanoid pathway enzymes prostacyclin synthase (PGIS) and inducible prostaglandin E(2) synthase (PGES) were joint classifiers that showed opposite effects in lung AC (PGIS down-regulated; PGES up-regulated). Finally, tissue microarrays identified the same protein expression pattern for PGIS and PGES in 108 different non-small cell lung cancer biopsies, and the detection of PGIS had statistically significant prognostic value in patient survival. Thus, the A/J mouse-urethane model reflects significant molecular details of human lung AC, and comparison of changes in orthologous gene expression may provide novel insights into lung carcinogenesis.

Adenocarcinoma↗