PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “draft genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Strain-specific genomic regions of Ruminococcus flavefaciens FD-1 as revealed by combinatorial random-phase genome sequencing and suppressive subtractive hybridization.

Two closely related strains of the Gram-positive, cellulolytic ruminal bacterium Ruminococcus flavefaciens were compared at the genomic level by suppressive subtractive hybridization. The two strains investigated in this study differ by 1.94% in their respective 16S rDNA genes. Three hundred and eighty-four PCR-amplified products were cloned and then screened for their strain identity by dot blot hybridization. Based on redundancy percentages of the clones sequenced, 9.5% of the genome of the R. flavefaciens FD-1 strain is not present in the JM1 strain. The majority of identities of individual cloned subtracted products (642 bp average length) bore no relation to deposited sequences in GenBank (42% of the subtracted library), whereas of those with putative assigned functions 7% are loosely associated with fibre-degradation, 6% with insertion elements, transposons and phage-like ORFs, 5% with cell membrane associated proteins and 3% with signal transduction. Subtracted sequences were then supplemented with the draft (2 x coverage) genome sequence of R. flavefaciens FD-1 to indicate potential regions of rearrangement within the genome, including a novel insertion sequence.

Base Sequence↗

Much ado about bacteria-to-vertebrate lateral gene transfer.

When the International Human Genome Sequencing Consortium (IHGSC) published its draft of the human genome in February 2001, several genes were identified as possible bacteria-to-vertebrate transfers (BVTs). These genes were identified by their highly significant sequence similarity to bacterial genes in BLAST searches, and by their lack of matches among non-vertebrate eukaryote genes. Many were later rejected as BVTs by several methods, including recovery of probable orthologs from the genomes of incompletely sequenced eukaryotes. Whereas the BVT issue has received considerable attention, there has been no compilation of all potential BVTs considered to date, nor any proposal of a single comprehensive method for rigorously establishing the veracity of a putative BVT. In reviewing the work to date, we list all of the proteins examined and propose systematic tests to investigate whether a vertebrate gene proposed as a BVT is indeed of bacterial origin. We use the proposed strategy to test--and reject--one of the BVTs from the original IHGSC list.

Animals↗

Gene enrichment in plant genomic shotgun libraries.

The Arabidopsis genome (about 130 Mbp) has been completely sequenced; whereas a draft sequence of the rice genome (about 430 Mbp) is now available and the sequencing of this genome will be completed in the near future. The much larger genomes of several important crop species, such as wheat (about 16,000 Mbp) or maize (about 2500 Mbp), may not be fully sequenced with current technology. Instead, sequencing-analysis strategies are being developed to obtain sequencing and mapping information selectively for the genic fraction (gene space) of complex plant genomes.

DNA Methylation↗

Genomic resources for chicken.

The recent sequencing and draft assembly of a chicken genome has provided biologists with an invaluable research tool that complements a growing list of additional avian genomic resources. For many researchers, finding and using these resources is challenging, because information is presented through an increasing number of Web sites and browser navigation frequently requires specific knowledge and expertise. This primer provides an overview of online genomic resources for the chicken, including the Ensembl, UCSC, and NCBI annotated chicken genome browsers; expressed sequence tag and in situ hybridization databases; and sources for microarrays, cDNAs, and bacterial artificial chromosomes (BACs). Several short tutorials oriented toward the biologist with limited bioinformatics skills outline how to retrieve several types of commonly needed information and reagents.

Animals↗

Chromosome complement of the fungal plant pathogen Fusarium graminearum based on genetic and physical mapping and cytological observations.

A genetic map of the filamentous fungus Fusarium graminearum (teleomorph: Gibberella zeae) was constructed to both validate and augment the draft whole-genome sequence assembly of strain PH-1. A mapping population was created from a cross between mutants of the sequenced strain (PH-1, NRRL 31084, originally isolated from Michigan) and a field strain from Minnesota (00-676, NRRL 34097). A total of 111 ascospore progeny were analyzed for segregation at 235 loci. Genetic markers consisted of sequence-tagged sites, primarily detected as dCAPS or CAPS (n = 131) and VNTRs (n = 31), in addition to AFLPs (n = 66) and 7 other markers. While most markers exhibited Mendelian inheritance, segregation distortion was observed for 25 predominantly clustered markers. A linkage map was generated using the Kosambi mapping function, using a LOD threshold value of 3.5. Nine linkage groups were detected, covering 1234 cM and anchoring 99.83% of the draft sequence assembly. The nine linkage groups and the 22 anchored scaffolds from the sequence assembly could be assembled into four chromosomes, leaving only five smaller scaffolds (59,630 bp total) of the nuclear DNA unanchored. A chromosome number of four was confirmed by cytological karyotyping. Further analysis of the genetic map data identified variation in recombination rate in different genomic regions that often spanned several hundred kilobases.

Chromosome Segregation↗

Conserved noncoding sequences are reliable guides to regulatory elements.

A 'working draft' of the human genome sequence is now available. Comparisons with the sequences of mouse and other species will be a powerful approach to identifying functional segments of the noncoding regions, such as gene regulatory elements. However, the choice of a species for most effective comparison differs among various loci.

Animals↗

Complete nucleotide sequence of a P2 family lysogenic bacteriophage, varphiMhaA1-PHL101, from Mannheimia haemolytica serotype A1.

The 34,525 nucleotide sequence of a double-stranded DNA bacteriophage (phiMhaA1-PHL101) from Mannheimia haemolytica serotype A1 has been determined. The phage encodes 50 open reading frames. Twenty-three of the proteins are similar to proteins of the P2 family of phages. Other protein sequences are most similar to possible prophage sequences from the draft genome of Histophilus somni 2336. Fourteen open reading frames encode proteins with no known homolog. The P2 orthologues are collinear in phiMhaA1-PHL101, with the exception of the phage tail protein gene T, which maps in a unique location between the S and V genes. The phage ORFs can be arranged into 17 possible transcriptional units and many of the genes are predicted to be translationally coupled. Southern blot analysis revealed phiMhaA1-PHL101 sequences in other A1 isolates as well as in serotype A5, A6, A9, and A12 strains of M. haemolytica, but not in the related organisms, Mannheimia glucosida or Pasteurella trehalosi.

Bacteriophages↗

An assessment of the resistance gene analogues of Oryza sativa ssp. japonica: their presence and structure.

Rice is the first cereal genome of known draft sequence, and the finished sequence for it is now nearly complete. In this paper, we describe a preliminary analysis of known rice genes aimed to detect resistance gene analogues of known structural classes. Putative resistance genes were identified in a dual approach--by using BLASTP searches to identify candidate sequences and by using Hidden Markov Models to predict domain presence in the candidates. The set of proteins examined was obtained from the publicly available data of TIGR (The Institute for Genomic Research). 1744 distinct RGAs were identified, 597 of which belonged to the NBS-LRR class. Supplementary data (sequences and annotations) is available on the web site http:/gkoczyk.bioinfo.pl/CMBL.

Computational Biology↗

Ralstonia metallidurans, a bacterium specifically adapted to toxic metals: towards a catalogue of metal-responsive genes.

Ralstonia metallidurans, formerly known as Alcaligenes eutrophus and thereafter as Ralstonia eutropha, is a beta-Proteobacterium colonizing industrial sediments, soils or wastes with a high content of heavy metals. The type strain CH34 carries two large plasmids (pMOL28 and pMOL30) bearing a variety of genes for metal resistance. A chronological overview describes the progress made in the knowledge of the plasmid-borne metal resistance mechanisms, the genetics of R. metallidurans CH34 and its taxonomy, and the applications of this strain in the fields of environmental remediation and microbial ecology. Recently, the sequence draft of the genome of R. metallidurans has become available. This allowed a comparison of these preliminary data with the published genome data of the plant pathogen Ralstonia solanacearum, which harbors a megaplasmid (of 2.1 Mb) carrying some metal resistance genes that are similar to those found in R. metallidurans CH34. In addition, a first inventory of metal resistance genes and operons across these two organisms could be made. This inventory, which partly relied on the use of proteomic approaches, revealed the presence of numerous loci not only on the large plasmids pMOL28 and pMOL30 but also on the chromosome. It suggests that metal-resistant Ralstonia, through evolution, are particularly well adapted to the harsh environments typically created by extreme anthropogenic situations or biotopes.

Adenosine Triphosphatases↗

rSNP_Guide: an integrated database-tools system for studying SNPs and site-directed mutations in transcription factor binding sites.

Since the human genome was sequenced in draft, single nucleotide polymorphism (SNP) analysis has become one of the keynote fields of bioinformatics. We have developed an integrated database-tools system, rSNP_Guide (http://wwwmgs.bionet.nsc.ru/mgs/systems/rsnp/), devoted to prediction of transcription factor (TF) binding sites, alterations of which could be associated with disease phenotype. By inputting data on alterations in DNA sequence and in DNA binding pattern of an unknown TF, rSNP_Guide searches for a known TF with alterations in the recognition score calculated on the basis of TF site's sequence and consistent with the input alterations in DNA binding to the unknown TF. Our system has been tested on many relationships between known TF sites and diseases, as well as on site-directed mutagenesis data. Experimental verification of rSNP_Guide system was made on functionally important SNPs in human TDO2and mouse K-ras genes. Additional examples of analysis are reported involving variants in the human gammaA-globin (HBG1), hsp70(HSPA1A), and Factor IX (F9) gene promoters.

Animals↗

Molecular genetics of schizophrenia: past, present and future.

Schizophrenia is a severe neuropsychiatric disorder with a polygenic mode of inheritance which is also governed by non-genetic factors. Candidate genes identified on the basis of biochemical and pharmacological evidence are being tested for linkage and association studies. Neurotransmitters, especially dopamine and serotonin have been widely implicated in its etiology. Genome scan of all human chromosomes with closely spaced polymorphic markers is being used for linkage studies. The completion and availability of the first draft of Human Genome Sequence has provided a treasure-trove that can be utilized to gain insight into the so far inaccessible regions of the human genome. Significant technological advances for identification of single nucleo-tide polymorphisms (SNPs) and use of microarrays have further strengthened research methodologies for genetic analysis of complex traits. In this review, we summarize the evolution of schizophrenia genetics from the past to the present, current trends and future direction of research.

Anticipation, Genetic↗

Predicted ATP-binding cassette systems in the phytopathogenic mollicute Spiroplasma kunkelii.

Spiroplasma kunkelii is a cell wall-free, helical, and motile mycoplasma-like organism that causes corn stunt disease in maize. The bacterium has a compact genome with a gene set approaching the minimal complement necessary for cellular life and pathogenesis. A set of 21 ATP-binding cassette (ABC) domains was identified during the annotation of a draft S. kunkelii genome sequence. These 21 ABC domains are present in 18 predicted proteins, and are components of 16 functional systems, which account for 5% of the protein coding capacity of the S. kunkelii genome. Of the 16 systems, 11 are membrane-bound transporters, and two are cytosolic systems involved in DNA repair and the oxidative stress response; the genes for the remaining three hypothetical systems harbor nonsense and/or frameshift mutations, so their functional status is doubtful. Assembly of the 11 multicomponent transporters, and comparisons with other known systems permitted functional predictions for the S. kunkelii ABC transporter systems. These transporters convey a wide variety of substrates, and are critical for nutrient uptake, multidrug resistance, and perhaps virulence. Our findings provide a framework for functional characterization of the ABC systems in S. kunkelii.

ATP-Binding Cassette Transporters↗

A cluster of novel serotonin receptor 3-like genes on human chromosome 3.

The ligand-gated ion channel family includes receptors for serotonin (5-hydroxytryptamine, 5-HT), acetylcholine, GABA, and glutamate. Drugs targeting subtypes of these receptors have proven useful for the treatment of various neuropsychiatric and neurological disorders. To identify new ligand-gated ion channels as potential therapeutic targets, drafts of human genome sequence were interrogated. Portions of four novel genes homologous to 5-HT(3A) and 5-HT(3B) receptors were identified within human sequence databases. We named the genes 5-HT(3C1)-5-HT(3C4). Radiation hybrid (RH) mapping localized these genes to chromosome 3q27-28. All four genes shared similar intron-exon organizations and predicted protein secondary structure with 5-HT(3A) and 5-HT(3B). Orthologous genes were detected by Southern blotting in several species including dog, cow, and chicken, but not in rodents, suggesting that these novel genes are not present in rodents or are very poorly conserved. Two of the novel genes are predicted to be pseudogenes, but two other genes are transcribed and spliced to form appropriate open reading frames. The 5-HT(3C1) transcript is expressed almost exclusively in small intestine and colon, suggesting a possible role in the serotonin-responsiveness of the gut.

Alternative Splicing↗

The state of the art of mammalian promoter recognition.

The draft sequences of whole genomes are being published at an ever-increasing pace, thus providing access to the human genomic sequence and, more recently, the mouse sequence. Genomes of the invertebrates are also becoming available. Now that the genomic DNA of mammalian species is available, an old problem can be tackled with renewed vigour mammalian promoter prediction. Gene promoters have proved elusive for more than a decade, despite their pivotal role in gene regulation. Recently, however, several new developments have made it possible to make meaningful large-scale predictions. This paper reviews the methods used for the prediction of mammalian, mostly human, promoters.

Animals↗

Integrated analysis of the genome and the transcriptome by FANTOM.

The key to reliable annotation of a mammalian genome is broad characterisation of the transcriptional output, the transcriptome. FANTOM, the functional annotation of mouse cDNA, is a large-scale analysis of both the genome and the transcriptome of the mouse. In the early days of this work, the transcripts were characterised using our sophisticated methods. After the timely release of the first draft of mouse genome sequences, interesting information was obtained by its integration with these one-by-one annotations. Moreover, each transcript included its expression profile. Here, the two integrated annotation methods used by FANTOM are reviewed: one-by-one and categorised. One-by-one annotation refers to naming carried out based on well-known transcripts or its fragments using the top-down-style pipeline developed mostly by the FANTOM project. Categorised annotation, which refers to transcript grouping, not only helps naming of unknown transcripts, but will be the most utilised method for integration of the genome and the transcriptome from now on.

Abstracting and Indexing↗

TERMINUS--Telomeric End-Read Mining IN Unassembled Sequences.

UNLABELLED: TERMINUS is a set of tools to map telomeres on draft sequences of whole genome shotgun sequencing projects. It mines raw sequence reads (from a trace archive) for telomeric reads, assembles them into contigs representing individual chromosome ends and BLASTs the resulting consensus sequences against the genome assembly to identify telomere-proximal genomic contigs. Finally, it estimates the sizes of telomeric gaps and identifies clones for gap closure. TERMINUS is implemented as a set of Perl scripts that requires two sets of inputs: the NCBI Trace Archive files for a given genome project; and ancillary genome assembly information. Results are output in spreadsheets containing information that facilitates manual validation. AVAILABILITY: The TERMINUS package and supplementary information can be downloaded from http://www.genome.kbrin.uky.edu/fungi_tel/terminus/ CONTACT: farman@uky.edu.

Algorithms↗

Linkage disequilibrium between microsatellite markers extends beyond 1 cM on chromosome 20 in Finns.

Linkage disequilibrium (LD) is a proven tool for evaluating population structure and localizing genes for monogenic disorders. LD-based methods may also help localize genes for complex traits. We evaluated marker-marker LD using 43 microsatellite markers spanning chromosome 20 with an average density of 2.3 cM. We studied 837 individuals affected with type 2 diabetes and 386 mostly unaffected spouse controls. A test of homogeneity between the affected individuals and their spouses showed no difference, allowing the 1223 individuals to be analyzed together. Significant (P < 0.01) LD was observed using a likelihood ratio test in all (11/11) marker pairs within 1 cM, 78% (25/32) of pairs 1-3 cM apart, and 39% (7/18) of pairs 3-4 cM apart, but for only 12 of 842 pairs more than 4 cM apart. We used the human genome project working draft sequence to estimate kilobase (kb) intermarker distances, and observed highly significant LD (P < 10(-10)) for all six marker pairs up to 350 kb apart, although the correlation of LD with cM is slightly better than the correlation with megabases. These data suggest that microsatellites present at 1-cM density are sufficient to observe marker-marker LD in the Finnish population.

Alleles↗