PubMed Health⌕ Search

Biomedical subjects

Genomics

Find indexed PubMed genomics citations. Search gene expression, sequencing and genetic variation in titles, abstracts and supplied subjects, then open the PubMed record.

At least 757 records · Page 42Linked to original sources

Mapping DNA-protein interactions in large genomes by sequence tag analysis of genomic enrichment.

Identifying the chromosomal targets of transcription factors is important for reconstructing the transcriptional regulatory networks underlying global gene expression programs. We have developed an unbiased genomic method called sequence tag analysis of genomic enrichment (STAGE) to identify the direct binding targets of transcription factors in vivo. STAGE is based on high-throughput sequencing of concatemerized tags derived from target DNA enriched by chromatin immunoprecipitation. We first used STAGE in yeast to confirm that RNA polymerase III genes are the most prominent targets of the TATA-box binding protein. We optimized the STAGE protocol and developed analysis methods to allow the identification of transcription factor targets in human cells. We used STAGE to identify several previously unknown binding targets of human transcription factor E2F4 that we independently validated by promoter-specific PCR and microarray hybridization. STAGE provides a means of identifying the chromosomal targets of DNA-associated proteins in any sequenced genome.

Cells, Cultured↗

The power and promise of population genomics: from genotyping to genome typing.

Population genomics has the potential to improve studies of evolutionary genetics, molecular ecology and conservation biology, by facilitating the identification of adaptive molecular variation and by improving the estimation of important parameters such as population size, migration rates and phylogenetic relationships. There has been much excitement in the recent literature about the identification of adaptive molecular variation using the population-genomic approach. However, the most useful contribution of the genomics model to population genetics will be improving inferences about population demography and evolutionary history.

Genetics, Population↗

The genome sequence of Blochmannia floridanus: comparative analysis of reduced genomes.

Bacterial symbioses are widespread among insects, probably being one of the key factors of their evolutionary success. We present the complete genome sequence of Blochmannia floridanus, the primary endosymbiont of carpenter ants. Although these ants feed on a complex diet, this symbiosis very likely has a nutritional basis: Blochmannia is able to supply nitrogen and sulfur compounds to the host while it takes advantage of the host metabolic machinery. Remarkably, these bacteria lack all known genes involved in replication initiation (dnaA, priA, and recA). The phylogenetic analysis of a set of conserved protein-coding genes shows that Bl. floridanus is phylogenetically related to Buchnera aphidicola and Wigglesworthia glossinidia, the other endosymbiotic bacteria whose complete genomes have been sequenced so far. Comparative analysis of the five known genomes from insect endosymbiotic bacteria reveals they share only 313 genes, a number that may be close to the minimum gene set necessary to sustain endosymbiotic life.

Animals↗

Jumping genes and shrinking genomes--probing the evolution of eukaryotic photosynthesis with genomics.

The advent of comparative genomics has revolutionized the study of the origin and evolution of eukaryotes and their organelles. Genomic analysis has revealed that the endosymbiosis that gave rise to plastids--the light-harvesting apparatus of photosynthetic eukaryotes--had a profound impact on the genetic composition of the host, far beyond the contribution of cyanobacterial genes for plastid-specific functions. Here I discuss recent advances in our appreciation of the mosaic nature of the eukaryotic nuclear genome, and the ongoing role endosymbiosis plays in shaping its content.

Eukaryota↗

Target selection for structural genomics: a single genome approach.

We describe our strategy for selecting targets for protein structure determination in context of structural genomics of a single genome. In the course of target selection, we have studied two of the smallest microbial genomes, Mycoplasma genitalium and Mycoplasma pneumoniae. To our surprise, we found that only 71 Mycoplasma genes or their orthologues can be considered as easy targets for high-throughput structural studies--far fewer than expected. We discuss the methods and criteria used for target selection and the reasons explaining rarity of easy targets. First, despite the common opinion that protein folds can be predicted for only 30-50% of genes, the number of "truly unknown" structures is less than one-third. Second, due to the different codon usage, two thirds of Mycoplasma proteins cannot be directly expressed in E. coli in high-throughput manner and require substitution by their homologues from other organisms. Third, membrane or large multi-domain proteins are difficult targets because of solubility and size issues and often require identification and structure determination of protein domains. Finally, we propose different approaches to address the difficult targets.

Cell Membrane↗

COmplete GENome Tracking (COGENT): a flexible data environment for computational genomics.

SUMMARY: We present a database of fully sequenced and published genomes to facilitate the re-distribution of data and ensure reproducibility of results in the field of computational genomics. For its design we have implemented an extremely simple yet powerful schema to allow linking of genome sequence data to other resources. AVAILABILITY: http://maine.ebi.ac.uk:8000/services/cogent/

Computational Biology↗

Using MoBIoS' scalable genome join to find conserved primer pair candidates between two genomes.

MOTIVATION: For the purpose of identifying evolutionary reticulation events in flowering plants, we determine a large number of paired, conserved DNA oligomers that may be used as primers to amplify orthologous DNA regions using the polymerase chain reaction (PCR). RESULTS: We develop an initial candidate set by comparing the Arabidopsis and rice genomes using MoBIoS (Molecular Biological Information System). MoBIoS is a metric-space database management system targeting life science data. Through the use of metric-space indexing techniques, two genomes can be compared in O(mlog n), where m and n are the lengths of the genomes, versus O(mn) for BLAST-based analysis. The filtering of low-complexity regions may also be accomplished by directly assessing the uniqueness of the region. We describe mSQL, a SQL extension being developed for MoBIoS that encapsulates the algorithmic details in a common database programming language, shielding end-users from esoteric programming. AVAILABILITY: Available upon request from authors.

Arabidopsis↗

Structural analysis of a Lotus japonicus genome. II. Sequence features and mapping of sixty-five TAC clones which cover the 6.5-mb regions of the genome.

Sixty-five TAC (transformation-competent artificial chromosomes) clones were selected from a genomic library of Lotus japonicus accession MG-20 based on the sequence information of expressed sequences tags (ESTs), cDNA and gene information, and their nucleotide sequences were determined. The average insert size of the TAC clone was approximately 100 kb, and the total length of the sequenced regions in this study is 6,556,100 bp. Together with the nucleotide sequences of 56 TAC clones previously reported, the regions sequenced so far total 12,029,295 bp. By comparison with the sequences in protein and EST databases and by analysis with computer programs for gene modeling, a total of 711 potential protein-encoding genes with known or predicted functions, 239 gene segments and 90 pseudogenes were identified in the newly sequenced regions. The average gene density assigned so far was 1 gene/9140 bp. The average length of the assigned genes was 2.6 kb, which is considerably larger than that assigned in the Arabidopsis thaliana genome (1.9 kb for 6451 genes). Introns were identified in approximately 73% of the potential genes, and the average number and length of the introns per gene were 3.4 and 377 bp, respectively. Simple sequence repeat length polymorphism (SSLP) or derived cleaved amplified polymorphic sequence (dCAPS) markers were generated based on the nucleotide sequences of the genomic clones obtained, and each clone was mapped onto the linkage map using the F2 mapping population derived from a cross of two accessions of L. japonicus, Gifu B-129 and Miyakojima MG-20. The sequence data, gene information and mapping information are available through the World Wide Web at http://www.kazusa.or.jp/lotus/.

Chromosome Mapping↗

Complete sequence of a sea lamprey (Petromyzon marinus) mitochondrial genome: early establishment of the vertebrate genome organization.

The complete nucleotide sequence of a sea lamprey (Petromyzon marinus) mitochondrial genome has been determined. The lamprey genome is 16,201 bp in length and contains genes for 13 proteins, two rRNAs, 22 tRNAs and two major noncoding regions. The order and transcriptional polarities of protein-coding genes are basically identical to those of other chordate mtDNAs, demonstrating that the common mitochondrial gene organization of vertebrates was established at an early stage of vertebrate evolution. The two major noncoding regions are separated by two tRNA genes. The first region probably functions as the control region because it contains distinctive conserved sequence blocks (CSB-II and III) common to other vertebrate control regions. The central conserved domain observed in other vertebrate control regions is not found in the lamprey, suggesting that it is a recently evolved functional domain in vertebrates. Noncoding segments are not found in the expected position of the origin of replication for the second strand, suggesting either that one of the tRNA genes has a dual function or that the second noncoding region may function as the second-strand origin. The base composition at the wobble positions of fourfold degenerate codon families is highly biased toward thymine (32.7%). Values of GC- and AT-skew are typical of vertebrate mitochondrial genomes.

Amino Acid Sequence↗

The Mouse Genome Database (MGD): a community resource. Status and enhancements. The Mouse Genome Informatics Group.

The Mouse Genome Database (MGD) is a comprehensive community database that integrates genetic, genomic and phenotypic information about the laboratory mouse. MGD provides detailed information about genes and genetic markers, elemental data from mapping experiments, descriptions of molecular segments including ESTs, probes, and cDNA clones, homology information between mouse and many other mammalian genomes, and phenotypic descriptions of gene mutations, gene function and mouse strains. All data are supported by citations. Interactive graphical displays of cytogenetic, genetic and physical maps are available. User support is provided through dedicated staff, bulletin boards, and user documentation. MGD can be accessed at http://www.informatics.jax.org

Animals↗

The TIGR rice genome annotation resource: annotating the rice genome and creating resources for plant biologists.

Rice is not only a major food staple for the world's population but it also is a model species for a major group of flowering plants, the monocotyledonous plants. Draft genomic sequence of two subspecies of rice, Oryza sativa spp. japonica and indica ssp. are publicly available. To provide the community with a resource to data-mine the rice genome, we have constructed an annotation resource for rice (http://www.tigr.org/tdb/e2k1/osa1/). In this resource, we have annotated the rice genome for gene content, identified motifs/domains within the predicted genes, constructed a rice repeat database, identified related sequences in other plant species, and identified syntenic sequences between rice and maize. All of the data is available through web-based interfaces, FTP downloads, and a Distributed Annotation System.

Chromosomes, Artificial↗

What's in the genome of a filamentous fungus? Analysis of the Neurospora genome sequence.

The German Neurospora Genome Project has assembled sequences from ordered cosmid and BAC clones of linkage groups II and V of the genome of Neurospora crassa in 13 and 12 contigs, respectively. Including additional sequences located on other linkage groups a total of 12 Mb were subjected to a manual gene extraction and annotation process. The genome comprises a small number of repetitive elements, a low degree of segmental duplications and very few paralogous genes. The analysis of the 3218 identified open reading frames provides a first overview of the protein equipment of a filamentous fungus. Significantly, N.crassa possesses a large variety of metabolic enzymes including a substantial number of enzymes involved in the degradation of complex substrates as well as secondary metabolism. While several of these enzymes are specific for filamentous fungi many are shared exclusively with prokaryotes.

Chromosome Mapping↗

Integr8 and Genome Reviews: integrated views of complete genomes and proteomes.

Integr8 is a new web portal for exploring the biology of organisms with completely deciphered genomes. For over 190 species, Integr8 provides access to general information, recent publications, and a detailed statistical overview of the genome and proteome of the organism. The preparation of this analysis is supported through Genome Reviews, a new database of bacterial and archaeal DNA sequences in which annotation has been upgraded (compared to the original submission) through the integration of data from many sources, including the EMBL Nucleotide Sequence Database, the UniProt Knowledgebase, InterPro, CluSTr, GOA and HOGENOM. Integr8 also allows the users to customize their own interactive analysis, and to download both customized and prepared datasets for their own use. Integr8 is available at http://www.ebi.ac.uk/integr8.

DNA, Archaeal↗

Multiple ribonuclease H-encoding genes in the Caenorhabditis elegans genome contrasts with the two typical ribonuclease H-encoding genes in the human genome.

Database searches of the Caenorhabditis elegans and human genomic DNA sequences revealed genes encoding ribonuclease H1 (RNase H1) and RNase H2 in each genome. The human genome contains a single copy of each gene, whereas C. elegans has four genes encoding RNase H1-related proteins and one gene for RNase H2. By analyzing the mRNAs produced from the C. elegans genes, examining the amino acid sequence of the predicted protein, and expressing the proteins in Esherichia coli we have identified two active RNase H1-like proteins. One is similar to other eukaryotic RNases H1, whereas the second RNase H (rnh-1.1) is unique. The rnh-1.0 gene is transcribed as a dicistronic message with three dsRNA-binding domains; the mature mRNA is transspliced with SL2 splice leader and contains only one dsRNA-binding domain. Formation of RNase H1 is further regulated by differential cis-splicing events. A single rnh-2 gene, encoding a protein similar to several other eukaryotic RNase H2L's, also has been examined. The diversity and enzymatic properties of RNase H homologues are other examples of expansion of protein families in C. elegans. The presence of two RNases H1 in C. elegans suggests that two enzymes are required in this rather simple organism to perform the functions that are accomplished by a single enzyme in more complex organisms. Phylogenetic analysis indicates that the active C. elegans RNases H1 are distantly related to one another and that the C. elegans RNase H1 is more closely related to the human RNase H1. The database searches also suggest that RNase H domains of LTR-retrotransposons in C. elegans are quite unrelated to cellular RNases H1, but numerous RNase H domains of human endogenous retroviruses are more closely related to cellular RNases H.

Amino Acid Sequence↗

Genomics and plant cells: application of genomics strategies to Arabidopsis cell biology.

In this review I seek to describe how the complete catalogue of plant genes and proteins, revealed by genome sequencing, can provide novel insights into cell biology. Many new analytical methods have been developed to digest the flood of genome sequence data, including analysis of the transcriptome, proteome and metabolites. High-throughput analysis of protein targeting and other methods will ascribe new information to proteins and create important links with other large datasets. To fulfil the potential revealed by this genomic information, many challenges have to be met. Among these are organizational changes needed to create common datasets accessible to all scientists, and bioinformatics solutions to capture and integrate diverse datasets. Once harnessed, these new strategies will irrevocably change the way we conduct plant science.

Arabidopsis↗

Genomic heterogeneity of rice dwarf phytoreovirus field isolates and nucleotide sequences of variants of genome segment 12.

Electrophoretic profiles of the dsRNAs of field isolates of rice dwarf virus (RDV) were compared with those of an isolate maintained at Hokkaido University (RDV-H). Unexpectedly, the genomic dsRNAs of most of the field isolates showed distinct electrophoretic mobility profiles. This was the case even among isolates from the same region. Genome segment 12 (S12) from some variants migrated faster than S12 from RDV-H. These RNAs were converted to full-length cDNAs and sequenced. S12 from all the variants had the same length of 1066 nucleotides with nucleotide sequence identities of 96 to 99%. Three open reading frames previously reported were present in all the variants, and the sequence identities were 95 to 99% for P12, 98 to 100% for P12OPa, and nearly 100% for P12OPb. A comparison of the nucleotide and amino acid sequences of the variants with sequences of the RDV-H and Akita isolate showed that there are two genomic types, one represented by RDV-H and the other by the Akita isolate.

Base Sequence↗

Completion of the Tula hantavirus genome sequence: properties of the L segment and heterogeneity found in the 3' termini of S and L genome RNAs.

In this study the L segment and the 5' and 3' termini of the S, M and L segments of the prototype Tula hantavirus (TUL) were sequenced, thus completing the first determination of the genome sequence of a hantavirus that has not been linked to any human disease. The TUL L segment comprises 6541 nt with one ORF of 6459 nt in the antigenome sense. This ORF potentially encodes a 2153 aa protein with a predicted molecular mass of 247 kDa. The amino acid sequence includes all the motifs conserved in RNA-dependent RNA polymerases. The 5' termini of all three genome RNAs (vRNAs) had the expected sequences conserved in hantaviruses. The 3' termini of M vRNAs were also conserved. However, the 3' termini of S and L vRNAs were heterogeneous as most of the sequenced 3' termini had either deletions of 1 to 22 nt or an extra 1 to 3 nt. No increase in the level of heterogeneity was seen in vRNAs of virions collected 3, 6, 9 and 12 days post-infection, suggesting that the heterogeneity already exists at the early stages of infection. The S and L vRNAs from infected cells had more truncated 3' termini than vRNAs from pelleted virus. Heterogeneity of the 3' termini of genome RNAs could decrease the efficiency of antigenome and mRNA syntheses and contribute to the slow growth observed for TUL and other hantaviruses in cell culture.

Genome, Viral↗

The genome of herpesvirus papio 2 is closely related to the genomes of human herpes simplex viruses.

Infection of baboons (Papio species) with herpesvirus papio 2 (HVP-2) produces a disease that is clinically similar to herpes simplex virus (HSV-1 and HSV-2) infection of humans. The development of a primate model of simplexvirus infection based on HVP-2 would provide a powerful resource to study virus biology and test vaccine strategies. In order to characterize the molecular biology of HVP-2 and justify further development of this model system we have constructed a physical map of the HVP-2 genome. The results of these studies have identified the presence of 26 reading frames that closely resemble HSV homologues. Furthermore, the HVP-2 genome shares a collinear arrangement with the genome of HSV. These studies further validate the development of the HVP-2 model as a surrogate system to study the biology of HSV infections.

Animals↗