PubMed Health⌕ Search

Biomedical subjects

Genomics

Find indexed PubMed genomics citations. Search gene expression, sequencing and genetic variation in titles, abstracts and supplied subjects, then open the PubMed record.

At least 703 records · Page 39Linked to original sources

Gene-based SNP discovery as part of the Japanese Millennium Genome Project: identification of 190,562 genetic variations in the human genome. Single-nucleotide polymorphism.

To construct an infrastructure for genome-wide association studies of common diseases or drug sensitivities, we have been systematically exploring common variants by resequencing genomic regions containing genes in DNA from 24 Japanese individuals. We have analyzed a total of 154 Mb, corresponding to approximately 5% of the human genome, and so far have identified 174,269 single-nucleotide polymorphisms and 16,293 insertion/deletion polymorphisms within gene regions, i.e., one polymorphism in 807 bp on average. Our data are freely available via our web site (http://snp.ims.u-tokyo.ac.jp) and will facilitate studies to identify genes associated with susceptibility to common diseases and genes involved in sensitivity to therapeutic drugs.

Gene Frequency↗

Complete mitochondrial genomes of eight cyclophyllidean tapeworms: genome pattern and phylogenetic analysis.

Cyclophyllidean tapeworms are widespread parasites of significant medical and veterinary importance. However, mitochondrial (mt) genomic resources for cyclophyllideans from China, particularly those recovered from wildlife hosts, remain comparatively limited. In this study, we sequenced and characterized the complete mt genomes of eight cyclophyllidean isolates collected from diverse wild and domestic hosts in China, including two Hymenolepis sp. isolates and two Raillietina sp. isolates from China, and four additional isolates of previously sequenced Taenia species. The circular mt genomes ranged from 13,387 to 14,021 bp in length, encoding 36 typical genes with variable non-coding regions. Comparative analysis revealed highly conserved gene composition and mostly conserved mt architecture, with localized rearrangement patterns detected among the cyclophyllidean lineages examined. In particular, all sampled Taeniidae exhibited a consistent trnL1-trnS2 arrangement, whereas the examined non-Taeniidae families showed the trnS2-trnL1 arrangement, confirming and extending, across additional wildlife-associated isolates, a previously proposed family-associated gene-order marker within Cyclophyllidea. Phylogenetic analyses based on concatenated amino acid sequences of the 12 protein-coding genes placed the eight isolates within their expected families, in topologies broadly consistent with previous mitogenomic studies. These data provide additional Chinese mitogenomic references, especially for underrepresented wildlife-associated isolates, and support family-associated gene-order patterns in Cyclophyllidea.

Animals↗

From genomics to post-genomics in Aspergillus.

Genome sequence data has recently become available for certain Aspergillus species. We consider the transition from genomics to a post-genomic era in Aspergillus, describing resources and methodologies available to underpin research efforts. Advances in our understanding of the fundamental biology of the Aspergilli, together with applications within the biotechnology and medical fields, are anticipated.

Aspergillosis↗

Evolutionary genomics of archaeal viruses: unique viral genomes in the third domain of life.

In terms of virion morphology, the known viruses of archaea fall into two distinct classes: viruses of mesophilic and moderately thermophilic Eueryarchaeota closely resemble head-and-tail bacteriophages whereas viruses of hyperthermophilic Crenarchaeota show a variety of unique morphotypes. In accord with this distinction, the sequenced genomes of euryarchaeal viruses encode many proteins homologous to bacteriophage capsid proteins. In contrast, initial analysis of the crenarchaeal viral genomes revealed no relationships with bacteriophages and, generally, very few proteins with detectable homologs. Here we describe a re-analysis of the proteins encoded by archaeal viruses, with an emphasis on comparative genomics of the unique viruses of Crenarchaeota. Detailed examination of conserved domains and motifs uncovered a significant number of previously unnoticed homologous relationships among the proteins of crenarchaeal viruses and between viral proteins and those from cellular life forms and allowed functional predictions for some of these conserved genes. A small pool of genes is shared by overlapping subsets of crenarchaeal viruses, in a general analogy with the metagenome structure of bacteriophages. The proteins encoded by the genes belonging to this pool include predicted transcription regulators, ATPases implicated in viral DNA replication and packaging, enzymes of DNA precursor metabolism, RNA modification enzymes, and glycosylases. In addition, each of the crenarchaeal viruses encodes several proteins with prokaryotic but not viral homologs, some of which, predictably, seem to have been scavenged from the crenarchaeal hosts, but others might have been acquired from bacteria. We conclude that crenarchaeal viruses are, in general, evolutionarily unrelated to other known viruses and, probably, evolved via independent accretion of genes derived from the hosts and, through more complex routes of horizontal gene transfer, from other prokaryotes.

Amino Acid Sequence↗

Comparative genome analysis of Campylobacter jejuni using whole genome DNA microarrays.

Whole genome DNA microarrays were constructed and used to investigate genomic diversity in 18 Campylobacter jejuni strains from diverse sources. New algorithms were developed that dynamically determine the boundary between the conserved and variable genes. Seven hypervariable plasticity regions (PR) were identified in the genome (PR1 to PR7) containing 136 genes (50%) of the variable gene pool. When comparisons were made with the sequenced strain NCTC11168, the number of absent or divergent genes ranged from 2.6% (40 genes) to 10.2% (163) and in total 16.3% (269) of the genes were variable. PR1 contains genes important in the utilisation of alternative electron acceptors for respiration and may confer a selective advantage to strains in restricted oxygen environments. PR2, 3 and 7 contain many outer membrane and periplasmic proteins and hypothetical proteins of unknown function that might be linked to phenotypic variation and adaptation to different ecological niches. PR4, 5 and 6 contain genes involved in the production and modification of antigenic surface structures.

Algorithms↗

Vertebrate genome sequencing: building a backbone for comparative genomics.

The human genome sequence provides a reference point from which we can compare ourselves with other organisms. Interspecies comparison is a powerful tool for inferring function from genomic sequence and could ultimately lead to the discovery of what makes humans unique. To date, most comparative sequencing has focused on pair-wise comparisons between human and a limited number of other vertebrates, such as mouse. Targeted approaches now exist for mapping and sequencing vertebrate bacterial artificial chromosomes (BACs) from numerous species, allowing rapid and detailed molecular and phylogenetic investigation of multi-megabase loci. Such targeted sequencing is complementary to current whole-genome sequencing projects, and would benefit greatly from the creation of BAC libraries from a diverse range of vertebrates.

Animals↗

Gene organization and sequence of the region containing the ribosomal protein genes RPL13A and RPS11 in the human genome and conserved features in the mouse genome.

We have determined the organization and sequence of the region containing two ribosomal protein (rp) genes in the human and mouse genomes. The two genes, human RPL13A and RPS11, and mouse Rpl13a and Rps11, are tandemly located in both genomes with an interval of only 4.6kb in the case of the human genes and 1.6kb in the case of the mouse genes. The human RPL13A and RPS11 are 4236bp and 3254bp in length and comprise eight and five exons respectively, whereas the mouse Rps11 is 1951bp long and has five exons. Structural comparison of these genes, including previously reported mouse Rpl13a, revealed a significant conservation of sequences in the promoter regions. Although most rp genes are dispersed throughout the human genome, the conserved features and adjacent localization indicate possible coordinate transcription of the two genes. Furthermore, we have found that four small nucleolar RNA (snoRNA) genes are located in the introns of the two rp genes, both human and mouse. U32, U33, and U34 snoRNAs are encoded in introns 2, 4, and 5 of RPL13A respectively, and U35 in the sixth intron of RPL13A and the third intron of RPS11. The same organization of these snoRNA genes was also observed in the case of the mouse genes.

Animals↗

Roles for genomic imprinting and the zygotic genome in placental development.

The placenta contains several types of feto-maternal interfaces where zygote-derived cells interact with maternal cells or maternal blood for the promotion of fetal growth and viability. The genetic factors regulating the interactions between different cell types within feto-maternal interfaces and the relative contributions of the maternal and zygotic genomes are poorly understood. Genomic imprinting, the epigenetic process responsible for parental origin-dependent functional differences between homologous chromosomes, has been proposed to contribute to these events. Previous studies showed that mouse conceptuses with an absence of imprinted differences between the two copies of chromosome 12 (upon paternal inheritance of both copies) die late in gestation and have a variety of defects, including placentomegaly. Here we examined the role of chromosome 12 imprinting in these placentae in more detail. We show that the spatial interactions between different cell types within feto-maternal interfaces are defective and identify abnormal behaviors in both zygote-derived and maternal cells that are attributed to the genome of the zygote but not the mother. These include compromised invasion of the maternal decidualized endometrium and the central maternal artery situated within it by zygote-derived trophoblast, abnormalities in the wall of the central maternal artery, and defects within the zygote-derived cellular layer of the labyrinth, which is in direct contact with maternal blood. These findings demonstrate multiple roles for chromosome 12 imprinting in the placenta that have not previously been associated with imprinting effects. They provide insights into the function of imprinting in placental development and have evolutionary and clinical implications.

Animals↗

Genomes OnLine Database (GOLD): a monitor of genome projects world-wide.

GOLD is a comprehensive resource for accessing information related to completed and ongoing genome projects world-wide. The database currently provides information on 350 genome projects, of which 48 have been completely sequenced and their analysis published. GOLD was created in 1997 and since April 2000 it has been licensed to Integrated Genomics. The database is freely available through the URL: http://igweb.integratedgenomics.com/GOLD/.

Animals↗

SUPFAM--a database of potential protein superfamily relationships derived by comparing sequence-based and structure-based families: implications for structural genomics and function annotation in genomes.

Members of a superfamily of proteins could result from divergent evolution of homologues with insignificant similarity in the amino acid sequences. A superfamily relationship is detected commonly after the three-dimensional structures of the proteins are determined using X-ray analysis or NMR. The SUPFAM database described here relates two homologous protein families in a multiple sequence alignment database of either known or unknown structure. The present release (1.1), which is the first version of the SUPFAM database, has been derived by analysing Pfam, which is one of the commonly used databases of multiple sequence alignments of homologous proteins. The first step in establishing SUPFAM is to relate Pfam families with the families in PALI, which is an alignment database of homologous proteins of known structure that is derived largely from SCOP. The second step involves relating Pfam families which could not be associated reliably with a protein superfamily of known structure. The profile matching procedure, IMPALA, has been used in these steps. The first step resulted in identification of 1280 Pfam families (out of 2697, i.e. 47%) which are related, either by close homologous connection to a SCOP family or by distant relationship to a SCOP family, potentially forming new superfamily connections. Using the profiles of 1417 Pfam families with apparently no structural information, an all-against-all comparison involving a sequence-profile match using IMPALA resulted in clustering of 67 homologous protein families of Pfam into 28 potential new superfamilies. Expansion of groups of related proteins of yet unknown structural information, as proposed in SUPFAM, should help in identifying 'priority proteins' for structure determination in structural genomics initiatives to expand the coverage of structural information in the protein sequence space. For example, we could assign 858 distinct Pfam domains in 2203 of the gene products in the genome of Mycobacterium tubercolosis. Fifty-one of these Pfam families of unknown structure could be clustered into 17 potentially new superfamilies forming good targets for structural genomics. SUPFAM database can be accessed at http://pauling.mbu.iisc.ernet.in/~supfam.

Animals↗

Genome-wide identification of replication origins in yeast by comparative genomics.

We discovered that sequences essential for replication origin function are frequently conserved in sensu stricto Saccharomyces species. Here we use analysis of phylogenetic conservation to identify replication origin sequences throughout the Saccharomyces cerevisiae genome at base pair resolution. Origin activity was confirmed for each of 228 predicted sites--representing 86% of apparent origin regions. This is the first study to determine the genome-wide location of replication origins at a resolution sufficient to identify the sequence elements bound by replication proteins. Our results demonstrate that phylogenetic conservation can be used to identify the origin sequences responsible for replicating a eukaryotic genome.

Base Sequence↗

Comparative analysis of chloroplast genomes: functional annotation, genome-based phylogeny, and deduced evolutionary patterns.

All protein sequences from 19 complete chloroplast genomes (cpDNA) have been studied using a new computational method able to analyze functional correlations among series of protein sequences contained in complete proteomes. First, all open reading frames (ORFs) from the cpDNAs, comprising a total of 2266 protein sequences, were compared against the 3168 proteins from Synechocystis PCC6803 complete genome to find functionally related orthologous proteins. Additionally, all cpDNA genomes were pairwise compared to find orthologous groups not present in cyanobacteria. Annotations in the cluster of othologous proteins database and CyanoBase were used as reference for the functional assignments. Following this protocol, new functional assignments were made for ORFs of unknown function and for ycfs (hypothetical chloroplast frames), which still lack a functional assignment. Using this information, a matrix of functional relationships was derived from profiles of the presence and/or absence of orthologous proteins; the matrix included 1837 proteins in 277 orthologous clusters. A factor analysis study of this matrix, followed by cluster analysis, allowed us to obtain accurate phylogenetic reconstructions and the detection of genes probably involved in speciation as phylogenetic correlates. Finally, by grouping common evolutionary patterns, we show that it is possible to determine functionally linked protein networks. This has allowed us to suggest putative associations for some unknown ORFs.

Bacterial Proteins↗

The T2T genome assembly of watershield (Brasenia schreberi) unveils genomic insights into aquatic adaptation.

Watershield (Brasenia schreberi), belonging to Cabombaceae within the order Nymphaeales, represents one of the early-diverged angiosperm lineages. This perennial floating leaf freshwater aquatic plant features submerged juvenile leaves enveloped in a thick layer of transparent gelatinous mucilage, aiding in its resistance to aquatic stress. However, the evolutionary history of the mechanisms underlying its specific phenotype remains unclear. In this study, we present the telomere-to-telomere level genome of B. schreberi, unveiling that it underwent two rounds of whole-genome duplications (WGDs) and a recent whole-genome triplication, with the most ancient WGD being shared by Nymphaeaceae. WGD and dispersed duplication significantly contributed to the expansion of gene families, which are primarily associated with environmental adaptation. Additionally, we discovered that mature leaves primarily conduct photosynthesis and may transport nutrients to underwater juvenile leaves for polysaccharide synthesis. We also identified an ancestral broad expression pattern of ABC genes, and the similar expression of anthocyanin biosynthesis genes across all flower organs resulted in entirely purple flowers. Our findings deepen the understanding of the evolution of this specific aquatic plant phenotypes.

Genome, Plant↗

Increased genome instability in Escherichia coli lon mutants: relation to emergence of multiple-antibiotic-resistant (Mar) mutants caused by insertion sequence elements and large tandem genomic amplifications.

Thirteen spontaneous multiple-antibiotic-resistant (Mar) mutants of Escherichia coli AG100 were isolated on Luria-Bertani (LB) agar in the presence of tetracycline (4 microg/ml). The phenotype was linked to insertion sequence (IS) insertions in marR or acrR or unstable large tandem genomic amplifications which included acrAB and which were bordered by IS3 or IS5 sequences. Five different lon mutations, not related to the Mar phenotype, were also found in 12 of the 13 mutants. Under specific selective conditions, most drug-resistant mutants appearing late on the selective plates evolved from a subpopulation of AG100 with lon mutations. That the lon locus was involved in the evolution to low levels of multidrug resistance was supported by the following findings: (i) AG100 grown in LB broth had an important spontaneous subpopulation (about 3.7x10(-4)) of lon::IS186 mutants, (ii) new lon mutants appeared during the selection on antibiotic-containing agar plates, (iii) lon mutants could slowly grow in the presence of low amounts (about 2x MIC of the wild type) of chloramphenicol or tetracycline, and (iv) a lon mutation conferred a mutator phenotype which increased IS transposition and genome rearrangements. The association between lon mutations and mutations causing the Mar phenotype was dependent on the medium (LB versus MacConkey medium) and the antibiotic used for the selection. A previously reported unstable amplifiable high-level resistance observed after the prolonged growth of Mar mutants in a low concentration of tetracycline or chloramphenicol can be explained by genomic amplification.

Anti-Bacterial Agents↗

Comparative genomics of insect-symbiotic bacteria: influence of host environment on microbial genome composition.

Commensal symbionts, thought to be intermediary amid obligate mutualists and facultative parasites, offer insight into forces driving the evolutionary transition into mutualism. Using macroarrays developed for a close relative, Escherichia coli, we utilized a heterologous array hybridization approach to infer the genomic compositions of a clade of bacteria that have recently established symbiotic associations: Sodalis glossinidius with the tsetse fly (Diptera, Glossina spp.) and Sitophilus oryzae primary endosymbiont (SOPE) with the rice weevil (Coleoptera, Sitophilus oryzae). Functional biologies within their hosts currently reflect different forms of symbiotic associations. Their hosts, members of distant insect taxa, occupy distinct ecological niches and have evolved to survive on restricted diets of blood for tsetse and cereal for the rice weevil. Comparison of genome contents between the two microbes indicates statistically significant differences in the retention of genes involved in carbon compound catabolism, energy metabolism, fatty acid metabolism, and transport. The greatest reductions have occurred in carbon catabolism, membrane proteins, and cell structure-related genes for Sodalis and in genes involved in cellular processes (i.e., adaptations towards cellular conditions) for SOPE. Modifications in metabolic pathways, in the form of functional losses complementing particularities in host physiology and ecology, may have occurred upon initial entry from a free-living to a symbiotic state. It is possible that these adaptations, streamlining genomes, act to make a free-living state no longer feasible for the harnessed microbe.

Animals↗

Vibrio cholerae phage K139: complete genome sequence and comparative genomics of related phages.

In this report, we characterize the complete genome sequence of the temperate phage K139, which morphologically belongs to the Myoviridae phage family (P2 and 186). The prophage genome consists of 33,106 bp, and the overall GC content is 48.9%. Forty-four open reading frames were identified. Homology analysis and motif search were used to assign possible functions for the genes, revealing a close relationship to P2-like phages. By Southern blot screening of a Vibrio cholerae strain collection, two highly K139-related phage sequences were detected in non-O1, non-O139 strains. Combinatorial PCR analysis revealed almost identical genome organizations. One region of variable gene content was identified and sequenced. Additionally, the tail fiber genes were analyzed, leading to the identification of putative host-specific sequence variations. Furthermore, a K139-encoded Dam methyltransferase was characterized.

Bacteriophages↗

GDR (Genome Database for Rosaceae): integrated web resources for Rosaceae genomics and genetics research.

BACKGROUND: Peach is being developed as a model organism for Rosaceae, an economically important family that includes fruits and ornamental plants such as apple, pear, strawberry, cherry, almond and rose. The genomics and genetics data of peach can play a significant role in the gene discovery and the genetic understanding of related species. The effective utilization of these peach resources, however, requires the development of an integrated and centralized database with associated analysis tools. DESCRIPTION: The Genome Database for Rosaceae (GDR) is a curated and integrated web-based relational database. GDR contains comprehensive data of the genetically anchored peach physical map, an annotated peach EST database, Rosaceae maps and markers and all publicly available Rosaceae sequences. Annotations of ESTs include contig assembly, putative function, simple sequence repeats, and anchored position to the peach physical map where applicable. Our integrated map viewer provides graphical interface to the genetic, transcriptome and physical mapping information. ESTs, BACs and markers can be queried by various categories and the search result sites are linked to the integrated map viewer or to the WebFPC physical map sites. In addition to browsing and querying the database, users can compare their sequences with the annotated GDR sequences via a dedicated sequence similarity server running either the BLAST or FASTA algorithm. To demonstrate the utility of the integrated and fully annotated database and analysis tools, we describe a case study where we anchored Rosaceae sequences to the peach physical and genetic map by sequence similarity. CONCLUSIONS: The GDR has been initiated to meet the major deficiency in Rosaceae genomics and genetics research, namely a centralized web database and bioinformatics tools for data storage, analysis and exchange. GDR can be accessed at http://www.genome.clemson.edu/gdr/.

Computer Graphics↗

Identifying genomic regions for fine-mapping using genome scan meta-analysis (GSMA) to identify the minimum regions of maximum significance (MRMS) across populations.

In order to detect linkage of the simulated complex disease Kofendrerd Personality Disorder across studies from multiple populations, we performed a genome scan meta-analysis (GSMA). Using the 7-cM microsatellite map, nonparametric multipoint linkage analyses were performed separately on each of the four simulated populations independently to determine p-values. The genome of each population was divided into 20-cM bin regions, and each bin was rank-ordered based on the most significant linkage p-value for that population in that region. The bin ranks were then averaged across all four studies to determine the most significant 20-cM regions over all studies. Statistical significance of the averaged bin ranks was determined from a normal distribution of randomly assigned rank averages. To narrow the region of interest for fine-mapping, the meta-analysis was repeated two additional times, with each of the 20-cM bins offset by 7 cM and 13 cM, respectively, creating regions of overlap with the original method. The 6-7 cM shared regions, where the highest averaged 20-cM bins from each of the three offsets overlap, designated the minimum region of maximum significance (MRMS). Application of the GSMA-MRMS method revealed genome wide significance (p-values refer to the average rank assigned to the bin) at regions including or adjacent to all of the simulated disease loci: chromosome 1 (p < 0.0001 for 160-167 cM, including D1), chromosome 3 (p-value < 0.0000001 for 287-294 cM, including D2), chromosome 5 (p-value < 0.001 for 0-7 cM, including D3), and chromosome 9 (p-value < 0.05 for 7-14 cM, the region adjacent to D4). This GSMA analysis approach demonstrates the power of linkage meta-analysis to detect multiple genes simultaneously for a complex disorder. The MRMS method enhances this powerful tool to focus on more localized regions of linkage.

Chromosomes, Human↗