PubMed Health⌕ Search

Biomedical subjects

Genomics

Find indexed PubMed genomics citations. Search gene expression, sequencing and genetic variation in titles, abstracts and supplied subjects, then open the PubMed record.

At least 559 records · Page 31Linked to original sources

The psychrophilic lifestyle as revealed by the genome sequence of Colwellia psychrerythraea 34H through genomic and proteomic analyses.

The completion of the 5,373,180-bp genome sequence of the marine psychrophilic bacterium Colwellia psychrerythraea 34H, a model for the study of life in permanently cold environments, reveals capabilities important to carbon and nutrient cycling, bioremediation, production of secondary metabolites, and cold-adapted enzymes. From a genomic perspective, cold adaptation is suggested in several broad categories involving changes to the cell membrane fluidity, uptake and synthesis of compounds conferring cryotolerance, and strategies to overcome temperature-dependent barriers to carbon uptake. Modeling of three-dimensional protein homology from bacteria representing a range of optimal growth temperatures suggests changes to proteome composition that may enhance enzyme effectiveness at low temperatures. Comparative genome analyses suggest that the psychrophilic lifestyle is most likely conferred not by a unique set of genes but by a collection of synergistic changes in overall genome content and amino acid composition.

Amino Acids↗

Genomics and the Human Genome Project: implications for psychiatry.

In the past decade the Human Genome Project has made extraordinary strides in understanding of fundamental human genetics. The complete human genetic sequence has been determined, and the chromosomal location of almost all human genes identified. Presently, a large international consortium, the HapMap Project, is working to identify a large portion of genetic variation in different human populations and the structure and relationship of these variants to each other. The Human Genome Project has approached human genetics on a scale not previously seen in biology. This has been made possible by dramatic advances in high throughput technology and bio-informatics. Tools such as gene chips and micro-arrays have spawned an entirely new strategy to examine the function and expression of genes in a massively parallel fashion. Together these tools have dramatically advanced our knowledge about the human genome. They promise powerful new approaches to complex genetic traits such as psychiatric illness. The goals and progress of the Human Genome Project and the technology involved are reviewed. The implications of this science for psychiatric genetics are discussed.

Computational Biology↗

SNPsFinder--a web-based application for genome-wide discovery of single nucleotide polymorphisms in microbial genomes.

UNLABELLED: Single nucleotide polymorphisms (SNPs) are the most abundant form of genetic variations in closely related microbial species, strains or isolates. Some SNPs confer selective advantages for microbial pathogens during infection and many others are powerful genetic markers for distinguishing closely related strains or isolates that could not be distinguished otherwise. To facilitate SNP discovery in microbial genomes, we have developed a web-based application, SNPsFinder, for genome-wide identification of SNPs. SNPsFinder takes multiple genome sequences as input to identify SNPs within homologous regions. It can also take contig sequences and sequence quality scores from ongoing sequencing projects for SNP prediction. SNPsFinder will use genome sequence annotation if available and map the predicted SNP regions to known genes or regions to assist further evaluation of the predicted SNPs for their functional significance. SNPsFinder can generate PCR primers for all predicted SNP regions according to user's input parameters to facilitate experimental validation. The results from SNPsFinder analysis are accessible through the World Wide Web. AVAILABILITY: The SNPsFinder program is available at http://snpsfinder.lanl.gov/. SUPPLEMENTARY INFORMATION: The user's manual is available at http://snpsfinder.lanl.gov/UsersManual/

Algorithms↗

The Genomic Threading Database: a comprehensive resource for structural annotations of the genomes from key organisms.

Currently, the Genomic Threading Database (GTD) contains structural assignments for the proteins encoded within the genomes of nine eukaryotes and 101 prokaryotes. Structural annotations are carried out using a modified version of GenTHREADER, a reliable fold recognition method. The Gen THREADER annotation jobs are distributed across multiple clusters of processors using grid technology and the predictions are deposited in a relational database accessible via a web interface at http://bioinf.cs.ucl.ac.uk/GTD. Using this system, up to 84% of proteins encoded within a genome can be confidently assigned to known folds with 72% of the residues aligned. On average in the GTD, 64% of proteins encoded within a genome are confidently assigned to known folds and 58% of the residues are aligned to structures.

Animals↗

Pseudomonas aeruginosa Genome Database and PseudoCAP: facilitating community-based, continually updated, genome annotation.

Using the Pseudomonas aeruginosa Genome Project as a test case, we have developed a database and submission system to facilitate a community-based approach to continually updated genome annotation (http://www.pseudomonas.com). Researchers submit proposed annotation updates through one of three web-based form options which are then subjected to review, and if accepted, entered into both the database and log file of updates with author acknowledgement. In addition, a coordinator continually reviews literature for suitable updates, as we have found such reviews to be the most efficient. Both the annotations database and updates-log database have Boolean search capability with the ability to sort results and download all data or search results as tab-delimited files. To complement this peer-reviewed genome annotation, we also provide a linked GBrowse view which displays alternate annotations. Additional tools and analyses are also integrated, including PseudoCyc, and knockout mutant information. We propose that this database system, with its focus on facilitating flexible queries of the data and providing access to both peer-reviewed annotations as well as alternate annotation information, may be a suitable model for other genome projects wishing to use a continually updated, community-based annotation approach. The source code is freely available under GNU General Public Licence.

DNA, Bacterial↗

SW-ARRAY: a dynamic programming solution for the identification of copy-number changes in genomic DNA using array comparative genome hybridization data.

Comparative genome hybridization (CGH) to DNA microarrays (array CGH) is a technique capable of detecting deletions and duplications in genomes at high resolution. However, array CGH studies of the human genome noting false negative and false positive results using large insert clones as probes have raised important concerns regarding the suitability of this approach for clinical diagnostic applications. Here, we adapt the Smith-Waterman dynamic-programming algorithm to provide a sensitive and robust analytic approach (SW-ARRAY) for detecting copy-number changes in array CGH data. In a blind series of hybridizations to arrays consisting of the entire tiling path for the terminal 2 Mb of human chromosome 16p, the method identified all monosomies between 267 and 1567 kb with a high degree of statistical significance and accurately located the boundaries of deletions in the range 267-1052 kb. The approach is unique in offering both a nonparametric segmentation procedure and a nonparametric test of significance. It is scalable and well-suited to high resolution whole genome array CGH studies that use array probes derived from large insert clones as well as PCR products and oligonucleotides.

Algorithms↗

Genome-wide comparison of differences in the integration sites of interspersed repeats between closely related genomes.

A technique for genome-wide detection of differences in the integration site positions of interspersed repeats in related genomes (DiffIR) is described. The technique is based on a whole- genome selective PCR amplification of the repeats' flanking regions followed by a differential hybridization screening of the arrayed library of the selected amplicons. The technique was successfully applied to the comparison of the integration sites in the human and chimpanzee genomes, allowing us to discover 11 new human-specific integrations of human endogenous retrovirus, K family (HML-2) long terminal repeats.

Animals↗

Comparative genomics and genome evolution in yeasts.

Yeasts provide a powerful model system for comparative genomics research. The availability of multiple complete genome sequences from different fungal groups--currently 18 hemiascomycetes, 8 euascomycetes and 4 basidiomycetes--enables us to gain a broad perspective on genome evolution. The sequenced genomes span a continuum of divergence levels ranging from multiple individuals within a species to species pairs with low levels of protein sequence identity and no conservation of gene order. One of the most interesting emerging areas is the growing number of events such as gene losses, gene displacements and gene relocations that can be attributed to the action of natural selection.

Centromere↗

High-resolution genomic profiling of chromosomal aberrations using Infinium whole-genome genotyping.

Array-CGH is a powerful tool for the detection of chromosomal aberrations. The introduction of high-density SNP genotyping technology to genomic profiling, termed SNP-CGH, represents a further advance, since simultaneous measurement of both signal intensity variations and changes in allelic composition makes it possible to detect both copy number changes and copy-neutral loss-of-heterozygosity (LOH) events. We demonstrate the utility of SNP-CGH with two Infinium whole-genome genotyping BeadChips, assaying 109,000 and 317,000 SNP loci, to detect chromosomal aberrations in samples bearing constitutional aberrations as well tumor samples at sub-100 kb effective resolution. Detected aberrations include homozygous deletions, hemizygous deletions, copy-neutral LOH, duplications, and amplifications. The statistical ability to detect common aberrations was modeled by analysis of an X chromosome titration model system, and sensitivity was modeled by titration of gDNA from a tumor cell with that of its paired normal cell line. Analysis was facilitated by using a genome browser that plots log ratios of normalized intensities and allelic ratios along the chromosomes. We developed two modes of SNP-CGH analysis, a single sample and a paired sample mode. The single sample mode computes log intensity ratios and allelic ratios by referencing to canonical genotype clusters generated from approximately 120 reference samples, whereas the paired sample mode uses a paired normal reference sample from the same individual. Finally, the two analysis modes are compared and contrasted for their utility in analyzing different types of input gDNA: low input amounts, fragmented gDNA, and Phi29 whole-genome pre-amplified DNA.

Cell Line, Tumor↗

Genome rearrangements in mammalian evolution: lessons from human and mouse genomes.

Although analysis of genome rearrangements was pioneered by Dobzhansky and Sturtevant 65 years ago, we still know very little about the rearrangement events that produced the existing varieties of genomic architectures. The genomic sequences of human and mouse provide evidence for a larger number of rearrangements than previously thought and shed some light on previously unknown features of mammalian evolution. In particular, they reveal that a large number of microrearrangements is required to explain the differences in draft human and mouse sequences. Here we describe a new algorithm for constructing synteny blocks, study arrangements of synteny blocks in human and mouse, derive a most parsimonious human-mouse rearrangement scenario, and provide evidence that intrachromosomal rearrangements are more frequent than interchromosomal rearrangements. Our analysis is based on the human-mouse breakpoint graph, which reveals related breakpoints and allows one to find a most parsimonious scenario. Because these graphs provide important insights into rearrangement scenarios, we introduce a new visualization tool that allows one to view breakpoint graphs superimposed with genomic dot-plots.

Algorithms↗

New evidence for genome-wide duplications at the origin of vertebrates using an amphioxus gene set and completed animal genomes.

The 2R hypothesis predicting two genome duplications at the origin of vertebrates is highly controversial. Studies published so far include limited sequence data from organisms close to the hypothesized genome duplications. Through the comparison of a gene catalog from amphioxus, the closest living invertebrate relative of vertebrates, to 3453 single-copy genes orthologous between Caenorhabditis elegans (C), Drosophila melanogaster (D), and Saccharomyces cerevisiae (Y), and to Ciona intestinalis ESTs, mouse, and human genes, we show with a large number of genes that the gene duplication activity is significantly higher after the separation of amphioxus and the vertebrate lineages, which we estimate at 650 million years (Myr). The majority of human orthologs of 195 CDY groups that could be dated by the molecular clock appear to be duplicated between 300 and 680 Myr with a mean at 488 million years ago (Mya). We detected 485 duplicated chromosomal segments in the human genome containing CDY orthologs, 331 of which are found duplicated in the mouse genome and within regions syntenic between human and mouse, indicating that these were generated earlier than the human-mouse split. Model based calculations of the codon substitution rate of the human genes included in these segments agree with the molecular clock duplication time-scale prediction. Our results favor at least one large duplication event at the origin of vertebrates, followed by smaller scale duplication closer to the bird-mammalian split.

Animals↗

The first-generation whole-genome radiation hybrid map in the horse identifies conserved segments in human and mouse genomes.

A first-generation radiation hybrid (RH) map of the equine (Equus caballus) genome was assembled using 92 horse x hamster hybrid cell lines and 730 equine markers. The map is the first comprehensive framework map of the horse that (1) incorporates type I as well as type II markers, (2) integrates synteny, cytogenetic, and meiotic maps into a consensus map, and (3) provides the most detailed genome-wide information to date on the organization and comparative status of the equine genome. The 730 loci (258 type I and 472 type II) included in the final map are clustered in 101 RH groups distributed over all equine autosomes and the X chromosome. The overall marker retention frequency in the panel is approximately 21%, and the possibility of adding any new marker to the map is approximately 90%. On average, the mapped markers are distributed every 19 cR (4 Mb) of the equine genome--a significant improvement in resolution over previous maps. With 69 new FISH assignments, a total of 253 cytogenetically mapped loci physically anchor the RH map to various chromosomal segments. Synteny assignments of 39 gene loci complemented the RH mapping of 27 genes. The results added 12 new loci to the horse gene map. Lastly, comparison of the assembly of 447 equine genes (256 linearly ordered RH-mapped and additional 191 FISH-mapped) with the location of draft sequences of their human and mouse orthologs provides the most extensive horse-human and horse-mouse comparative map to date. We expect that the foundation established through this map will significantly facilitate rapid targeted expansion of the horse gene map and consequently, mapping and positional cloning of genes governing traits significant to the equine industry.

Animals↗

Deductions about the number, organization, and evolution of genes in the tomato genome based on analysis of a large expressed sequence tag collection and selective genomic sequencing.

Analysis of a collection of 120,892 single-pass ESTs, derived from 26 different tomato cDNA libraries and reduced to a set of 27,274 unique consensus sequences (unigenes), revealed that 70% of the unigenes have identifiable homologs in the Arabidopsis genome. Genes corresponding to metabolism have remained most conserved between these two genomes, whereas genes encoding transcription factors are among the fastest evolving. The majority of the 10 largest conserved multigene families share similar copy numbers in tomato and Arabidopsis, suggesting that the multiplicity of these families may have occurred before the divergence of these two species. An exception to this multigene conservation was observed for the E8-like protein family, which is associated with fruit ripening and has higher copy number in tomato than in Arabidopsis. Finally, six BAC clones from different parts of the tomato genome were isolated, genetically mapped, sequenced, and annotated. The combined analysis of the EST database and these six sequenced BACs leads to the prediction that the tomato genome encodes approximately 35,000 genes, which are sequestered largely in euchromatic regions corresponding to less than one-quarter of the total DNA in the tomato nucleus.

Arabidopsis↗

Yeast genome sequencing: the power of comparative genomics.

For decades, unicellular yeasts have been general models to help understand the eukaryotic cell and also our own biology. Recently, over a dozen yeast genomes have been sequenced, providing the basis to resolve several complex biological questions. Analysis of the novel sequence data has shown that the minimum number of genes from each species that need to be compared to produce a reliable phylogeny is about 20. Yeast has also become an attractive model to study speciation in eukaryotes, especially to understand molecular mechanisms behind the establishment of reproductive isolation. Comparison of closely related species helps in gene annotation and to answer how many genes there really are within the genomes. Analysis of non-coding regions among closely related species has provided an example of how to determine novel gene regulatory sequences, which were previously difficult to analyse because they are short and degenerate and occupy different positions. Comparative genomics helps to understand the origin of yeasts and points out crucial molecular events in yeast evolutionary history, such as whole-genome duplication and horizontal gene transfer(s). In addition, the accumulating sequence data provide the background to use more yeast species in model studies, to combat pathogens and for efficient manipulation of industrial strains.

Biological Evolution↗

Analysis of the first genome fragment from the marine sponge-associated, novel candidate phylum Poribacteria by environmental genomics.

The novel candidate phylum Poribacteria is specifically associated with several marine demosponge genera. Because no representatives of Poribacteria have been cultivated, an environmental genomic approach was used to gain insights into genomic properties and possibly physiological/functional features of this elusive candidate division. In a large-insert library harbouring an estimated 1.1 Gb of microbial community DNA from Aplysina aerophoba, a Poribacteria-positive 16S rRNA gene locus was identified. Sequencing and sequence annotation of the 39 kb size insert revealed 27 open reading frames (ORFs) and two genes for stable RNAs. The fragment exhibited an overall G+C content of 50.5% and a coding density of 86.1%. The 16S rRNA gene was unlinked from a conventional rrn operon. Its flanking regions did not show any synteny to other 16S rRNA encoding loci from microorganisms with unlinked rrn operons. Two of the predicted hypothetical proteins were highly similar to homologues from Rhodopirellula baltica. Furthermore, a novel kind of molybdenum containing oxidoreductase was predicted as well as a series of eight ORFs encoding for unusual transporters, channel or pore forming proteins. This environmental genomics approach provides, for the first time, genomic and, by inference, functional information on the so far uncultivated, sponge-associated candidate division Poribacteria.

Animals↗

Genomic islands in the Corynebacterium efficiens genome.

Corynebacterium efficiens is a gram-positive nonpathogenic bacterium which can grow and produce glutamate at 40 degrees C or above. By using the cumulative GC profile method, we have identified four genomic islands which have many unifying genomic island-specific features in the C. efficiens genome. The presence of the gene encoding an aspartate kinase in a genomic island helps explain the unexpected low thermal stability of this enzyme; i.e., the adaptive mutations have not occurred extensively due to the recent horizontal gene transfer.

Aspartate Kinase↗

Complete genome sequence and comparative genomics of Shigella flexneri serotype 2a strain 2457T.

We determined the complete genome sequence of Shigella flexneri serotype 2a strain 2457T (4,599,354 bp). Shigella species cause >1 million deaths per year from dysentery and diarrhea and have a lifestyle that is markedly different from those of closely related bacteria, including Escherichia coli. The genome exhibits the backbone and island mosaic structure of E. coli pathogens, albeit with much less horizontally transferred DNA and lacking 357 genes present in E. coli. The strain is distinctive in its large complement of insertion sequences, with several genomic rearrangements mediated by insertion sequences, 12 cryptic prophages, 372 pseudogenes, and 195 S. flexneri-specific genes. The 2457T genome was also compared with that of a recently sequenced S. flexneri 2a strain, 301. Our data are consistent with Shigella being phylogenetically indistinguishable from E. coli. The S. flexneri-specific regions contain many genes that could encode proteins with roles in virulence. Analysis of these will reveal the genetic basis for aspects of this pathogenic organism's distinctive lifestyle that have yet to be explained.

Base Sequence↗

Design of a seven-genome Escherichia coli microarray for comparative genomic profiling.

We describe the design and evaluate the use of a high-density oligonucleotide microarray covering seven sequenced Escherichia coli genomes in addition to several sequenced E. coli plasmids, bacteriophages, pathogenicity islands, and virulence genes. Its utility is demonstrated for comparative genomic profiling of two unsequenced strains, O175:H16 D1 and O157:H7 3538 (Deltastx(2)::cat) as well as two well-known control strains, K-12 W3110 and O157:H7 EDL933. By using fluorescently labeled genomic DNA to query the microarrays and subsequently analyze common virulence genes and phage elements and perform whole-genome comparisons, we observed that O175:H16 D1 is a K-12-like strain and confirmed that its phi3538 (Deltastx(2)::cat) phage element originated from the E. coli 3538 (Deltastx(2)::cat) strain, with which it shares a substantial proportion of phage elements. Moreover, a number of genes involved in DNA transfer and recombination was identified in both new strains, providing a likely explanation for their capability to transfer phi3538 (Deltastx(2)::cat) between them. Analyses of control samples demonstrated that results using our custom-designed microarray were representative of the true biology, e.g., by confirming the presence of all known chromosomal phage elements as well as 98.8 and 97.7% of queried chromosomal genes for the two control strains. Finally, we demonstrate that use of spatial information, in terms of the physical chromosomal locations of probes, improves the analysis.

Coliphages↗