PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “draft genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

The cell surface of Lactobacillus reuteri ATCC 55730 highlighted by identification of 126 extracellular proteins from the genome sequence.

Bioinformatical analyses of a draft genome sequence of the commensal bacterium Lactobacillus reuteri ATCC 55730 revealed 126 genes encoding putative extracellular proteins. The function, localization and distribution in bacterial species were predicted. Interestingly, few proteins possessed LPXTG motifs or C-terminal transmembrane anchors. Instead eight proteins were putatively anchored by GW repeats and several secreted proteins were likely to be re-associated to the surface. The majority of the extracellular proteins were widely distributed, i.e., found universally or in gram-positive bacteria, but 24 were only detected in L. reuteri. Further, the number of transporters was lower, while the number of enzyme was higher than in related species.

Amino Acid Motifs↗

C-type lectin-like domains in Fugu rubripes.

BACKGROUND: Members of the C-type lectin domain (CTLD) superfamily are metazoan proteins functionally important in glycoprotein metabolism, mechanisms of multicellular integration and immunity. Three genome-level studies on human, C. elegans and D. melanogaster reported previously demonstrated almost complete divergence among invertebrate and mammalian families of CTLD-containing proteins (CTLDcps). RESULTS: We have performed an analysis of CTLD family composition in Fugu rubripes using the draft genome sequence. The results show that all but two groups of CTLDcps identified in mammals are also found in fish, and that most of the groups have the same members as in mammals. We failed to detect representatives for CTLD groups V (NK cell receptors) and VII (lithostathine), while the DC-SIGN subgroup of group II is overrepresented in Fugu. Several new CTLD-containing genes, highly conserved between Fugu and human, were discovered using the Fugu genome sequence as a reference, including a CSPG family member and an SCP-domain-containing soluble protein. A distinct group of soluble dual-CTLD proteins has been identified, which may be the first reported CTLDcp group shared by invertebrates and vertebrates. We show that CTLDcp-encoding genes are selectively duplicated in Fugu, in a manner that suggests an ancient large-scale duplication event. We have verified 32 gene structures and predicted 63 new ones, and make our annotations available through a distributed annotation system (DAS) server http://anz.anu.edu.au:8080/Fugu_rubripes/ and their sequences as additional files with this paper. CONCLUSIONS: The vertebrate CTLDcp family was essentially formed early in vertebrate evolution and is completely different from the invertebrate families. Comparison of fish and mammalian genomes revealed three groups of CTLDcps and several new members of the known groups, which are highly conserved between fish and mammals, but were not identified in the study using only mammalian genomes. Despite limitations of the draft sequence, the Fugu rubripes genome is a powerful instrument for gene discovery and vertebrate evolutionary analysis. The composition of the CTLDcp superfamily in fish and mammals suggests that large-scale duplication events played an important role in the evolution of vertebrates.

Amino Acid Sequence↗

Integration of the cytogenetic map with the draft human genome sequence.

Chemically staining metaphase chromosomes resulting in an alternating dark and light banding pattern provide a tool by which abnormalities in chromosomes from diseased cells can be identified. The localization of these aberrations to a chromosomal region provides clues as to which gene or genes may contribute to a particular disease. With the sequencing of the human genome, it became critical to determine the positions of these cytogenetic bands within the sequence in order to take advantage of vast amount of information now anchored to the sequence, especially the locations of genes. The molecular basis of cytogenetic bands is not well understood, therefore their positions cannot be determined solely based on sequence information. We developed a dynamic programming algorithm that employs results from approximately 9500 fluorescence in situ hybridization experiments to approximate the locations of the 850 high-resolution bands in the June 2002 version of the draft human genome sequence. These band predictions support previously identified correlations between band stain intensity and certain structural characteristics of chromosomes, namely GC content, repeat structure content, CpG island density, gene density and degree of condensation.

Algorithms↗

High-density rat radiation hybrid maps containing over 24,000 SSLPs, genes, and ESTs provide a direct link to the rat genome sequence.

The laboratory rat is a major model organism for systems biology. To complement the cornucopia of physiological and pharmacological data generated in the rat, a large genomic toolset has been developed, culminating in the release of the rat draft genome sequence. The rat draft sequence used a variety of assembly packages, as well as data from the Radiation Hybrid (RH) map of the rat as part of their validation. As part of the Rat Genome Project, we have been building a high-density RH map to facilitate data integration from multiple maps and now to help validate the genome assembly. By incorporating vectors from our lab and several other labs, we have doubled the number of simple sequence length polymorphisms (SSLPs), genes, expressed sequence tags (ESTs), and sequence-tagged sites (STSs) compared to any other genome-wide rat map, a total of 24,437 elements. During the process, we also identified a novel approach for integrating the RH placement results from multiple maps. This new integrated RH map contains approximately 10 RH-mapped elements per Mb on the genome assembly, enabling the RH maps to serve as a scaffold for a variety of data visualization tools.

Animals↗

Locating sequence on FPC maps and selecting a minimal tiling path.

This study discusses three software tools, the first two aid in integrating sequence with an FPC physical map and the third automatically selects a minimal tiling path given genomic draft sequence and BAC end sequences. The first tool, FSD (FPC Simulated Digest), takes a sequenced clone and adds it back to the map based on a fingerprint generated by an in silico digest of the clone. This allows verification of sequenced clone positions and the integration of sequenced clones that were not originally part of the FPC map. The second tool, BSS (Blast Some Sequence), takes a query sequence and positions it on the map based on sequence associated with the clones in the map. BSS has multiple uses as follows: (1) When the query is a file of marker sequences, they can be added as electronic markers. (2) When the query is draft sequence, the results of BSS can be used to close gaps in a sequenced clone or the physical map. (3) When the query is a sequenced clone and the target is BAC end sequences, one may select the next clone for sequencing using both sequence comparison results and map location. (4) When the query is whole-genome draft sequence and the target is BAC end sequences, the results can be used to select many clones for a minimal tiling path at once. The third tool, pickMTP, automates the majority of this last usage of BSS. Results are presented using the rice FPC map, BAC end sequences, and whole-genome shotgun from Syngenta.

Chromosomes, Artificial, Bacterial↗

Phytophthora genome sequences uncover evolutionary origins and mechanisms of pathogenesis.

Draft genome sequences have been determined for the soybean pathogen Phytophthora sojae and the sudden oak death pathogen Phytophthora ramorum. Oömycetes such as these Phytophthora species share the kingdom Stramenopila with photosynthetic algae such as diatoms, and the presence of many Phytophthora genes of probable phototroph origin supports a photosynthetic ancestry for the stramenopiles. Comparison of the two species' genomes reveals a rapid expansion and diversification of many protein families associated with plant infection such as hydrolases, ABC transporters, protein toxins, proteinase inhibitors, and, in particular, a superfamily of 700 proteins with similarity to known oömycete avirulence genes.

Algal Proteins↗

Albidovulum molybdatiresistens sp. nov., a molybdate-resistant bacterium isolated from river water.

A Gram-stain-negative, aerobic, non-motile, catalase- and oxidase-positive, white rod-shaped strain, RF13T, was isolated from water samples of the Qingliang River in Fucheng County, Hebei Province, China, and was grown at 15-42 °C (optimum 35 °C), pH 6.0-8.0 (optimum pH 7), and 0-0.5% (w/v) NaCl (optimum concentration 0%). Phylogenetic analysis based on 16S rRNA gene sequences showed that strain RF13T belonged to the genus Albidovulum, with closest sequence similarity to Albidovulum salinarum MCCC 1K0602T (97.2%), Frigidibacter oleivorans CGMCC 1.3778T (97.2%), Allgaiera indica MCCC 1A01802T (96.8%), and Pseudothioclava arenosa KCTC 52190T (96.4%). The genome size of strain RF13T was 3.7 Mb, and the DNA G+C content was 64.6%. The DNA-DNA hybridisation value (dDDH), average nucleotide identity (ANI), and average amino acid identity (AAI) between strain RF13T and the reference strain were less than 20.0%, 78.8%, and 72.8%, respectively. Chemotaxonomic analysis revealed Summed feature 8 (48.4%) (C18:1 ω6c and/or C18:1 ω7c), C18:1 ω7c 11-methyl (22.1%), C18:0 3OH (7.9%), and C10:0 3OH (5.0%) as predominant fatty acids. The polar lipids consisted of phosphatidylglycerol, diphosphatidylglycerol, two unidentified aminolipids, two unidentified phospholipids, and three unidentified lipids. The predominant isoprenoid quinone was ubiquinone-10 (Q-10), and a small amount of Q-9 was also detected. In addition, strain RF13T exhibited a minimum inhibitory concentration (MIC) of 20 mM for molybdate in R2A broth medium and was capable of reducing molybdate to molybdenum blue. Based on the results of biochemical, physiological, phylogenetic, and chemotaxonomic analyses, combined with 16S rRNA gene sequence analyses and draft genome sequence comparisons, strain RF13T was considered to represent a novel species of the genus Albidovulum, and was therefore named Albidovulum molybdatiresistens sp. nov. The type strain was RF13T (= GDMCC 1.3414T= JCM 35643T).

Phylogeny↗

Integration of the rat recombination and EST maps in the rat genomic sequence and comparative mapping analysis with the mouse genome.

Inbred strains of the laboratory rat are widely used for identifying genetic regions involved in the control of complex quantitative phenotypes of biomedical importance. The draft genomic sequence of the rat now provides essential information for annotating rat quantitative trait locus (QTL) maps. Following the survey of unique rat microsatellite (11,585 including 1648 new markers) and EST (10,067) markers currently available, we have incorporated a selection of 7952 rat EST sequences in an improved version of the integrated linkage-radiation hybrid map of the rat containing 2058 microsatellite markers which provided over 10,000 potential anchor points between rat QTL and the genomic sequence of the rat. A total of 996 genetic positions were resolved (avg. spacing 1.77 cM) in a single large intercross and anchored in the rat genomic sequence (avg. spacing 1.62 Mb). Comparative genome maps between rat and mouse were constructed by successful computational alignment of 6108 mapped rat ESTs in the mouse genome. The integration of rat linkage maps in the draft genomic sequence of the rat and that of other species represents an essential step for translating rat QTL intervals into human chromosomal targets.

Animals↗

Complete mutation analysis panel of the 39 human HOX genes.

BACKGROUND: The HOX gene family consists of highly conserved transcription factors that specify the identity of the body segments along the anteroposterior axis of the embryo. Because the phenotypes of mice with targeted disruptions of Hox genes resemble some patterns of human malformations, mutations in HOX genes have been expected to be associated with a significant number of human malformations. Thus far, however, mutations have been documented in only three of the 39 human HOX genes (HOXD13, HOXA13, and HOXA11) partly because current knowledge on the complete coding sequence and genome structure is limited to only 20 of the 39 human HOX genes. METHODS: Taking advantage of the human and mouse draft genome sequences, we attempted to characterize the remaining 19 human HOX genes by bioinformatic analysis including phylogenetic footprinting, the probabilistic prediction method, and comparison of genomic sequences with the complete set of the human anonymous cDNA sequences. RESULTS: We were able to determine the full coding sequences of 19 HOX genes and their genome structure and successfully designed a complete set of PCR primers to amplify the entire coding region of each of the 39 HOX genes from genomic DNA. CONCLUSIONS: Our results indicate the usefulness of bioinformatic analysis of the draft genome sequences for clinically oriented research projects. It is hoped that the mutation panel provided here will serve as a launchpad for a new discourse on the genetic basis of human malformations.

Animals↗

Utilization of a zebra finch BAC library to determine the structure of an avian androgen receptor genomic region.

The zebra finch (Taeniopygia guttata) is an important model organism for studying behavior, neuroscience, avian biology, and evolution. To support the study of its genome, we constructed a BAC library (TG__Ba) using DNA from livers of females. The BAC library consists of 147,456 clones with 98% containing inserts of an average size of 134 kb and represents 15.5 haploid genome equivalents. By sequencing a whole BAC, a full-length androgen receptor open reading frame was identified, the first in an avian species. Comparison of BAC end sequences and the whole BAC sequence with the chicken genome draft sequence showed a high degree of conserved synteny between the zebra finch and the chicken genome.

Animals↗

PlasmoDB: the Plasmodium genome resource. An integrated database providing tools for accessing, analyzing and mapping expression and sequence data (both finished and unfinished).

PlasmoDB (http://PlasmoDB.org) is the official database of the Plasmodium falciparum genome sequencing consortium. This resource incorporates finished and draft genome sequence data and annotation emerging from Plasmodium sequencing projects. PlasmoDB currently houses information from five parasite species and provides tools for cross-species comparisons. Sequence information is also integrated with other genomic-scale data emerging from the Plasmodium research community, including gene expression analysis from EST, SAGE and microarray projects. The relational schemas used to build PlasmoDB [Genomics Unified Schema (GUS) and RNA Abundance Database (RAD)] employ a highly structured format to accommodate the diverse data types generated by sequence and expression projects. A variety of tools allow researchers to formulate complex, biologically based queries of the database. A version of the database is also available on CD-ROM (Plasmodium GenePlot), facilitating access to the data in situations where Internet access is difficult (e.g. by malaria researchers working in the field). The goal of PlasmoDB is to enhance utilization of the vast quantities of data emerging from genome-scale projects by the global malaria research community.

Animals↗

Short interspersed elements (SINEs) are a major source of canine genomic diversity.

SINEs are retrotransposons that have enjoyed remarkable reproductive success during the course of mammalian evolution, and have played a major role in shaping mammalian genomes. Previously, an analysis of survey-sequence data from an individual dog (a poodle) indicated that canine genomes harbor a high frequency of alleles that differ only by the absence or presence of a SINEC_Cf repeat. Comparison of this survey-sequence data with a draft genome sequence of a distinct dog (a boxer) has confirmed this prediction, and revealed the chromosomal coordinates for >10,000 loci that are bimorphic for SINEC_Cf insertions. Analysis of SINE insertion sites from the genomes of nine additional dogs indicates that 3%-5% are absent from either the poodle or boxer genome sequences--suggesting that an additional 10,000 bimorphic loci could be readily identified in the general dog population. We describe a methodology that can be used to identify these loci, and could be adapted to exploit these bimorphic loci for genotyping purposes. Approximately half of all annotated canine genes contain SINEC_Cf repeats, and these elements are occasionally transcribed. When transcribed in the antisense orientation, they provide splice acceptor sites that can result in incorporation of novel exons. The high frequency of bimorphic SINE insertions in the dog population is predicted to provide numerous examples of allele-specific transcription patterns that will be valuable for the study of differential gene expression among multiple dog breeds.

Animals↗

Molecular cloning and expression of mouse Wnt14, and structural comparison between mouse Wnt14-Wnt3a gene cluster and human WNT14-WNT3A gene cluster.

Glycoprotein WNTs play key roles in carcinogenesis and embryogenesis. Human WNT14 and WNT3A genes are clustered in human chromosome 1q42 region with an interval of about 58 kb. Here, mouse Wnt14 was isolated to compare the structure of human WNT14-WNT3A gene cluster with that of mouse Wnt14-Wnt3a gene cluster. Mouse Wnt14 showed 98.1% total-amino-acid identity with human WNT14, and 61.9% total-amino-acid identity with human WNT14B/WNT15. Mouse Wnt14 mRNA was expressed in adult brain, lung, skeletal muscle, heart, and 17-day embryo. Mouse Wnt14 and Wnt3a genes were clustered in head-to-head manner with an interval of about 16 kb. Exon-intron structures were well conserved between human WNT14-WNT3A gene cluster and mouse Wnt14-Wnt3a gene cluster. Capicua-related sequence and AK024248-related sequence were identified in the intergenic region of human Wnt14-Wnt3a gene cluster as well as in other human chromosomal loci, but not in that of mouse Wnt14-Wnt3a gene cluster. Capicua-related sequences were pseudogenes derived from Capicua gene on human chromosome 19q13. Capicua pseudogene and AK024248-related sequence were clustered in tail-to-tail manner with interval ranging from 2.2 to 11.0 kb. AK024248-related sequences in several human genome draft sequences were truncated in the 3'-portion compared with that in the intergenic region of human WNT14-WNT3A gene cluster. This is the first report on structural comparison of WNT gene clusters in human genome and in mouse genome.

Amino Acid Sequence↗

Cold adaptation in the Antarctic Archaeon Methanococcoides burtonii involves membrane lipid unsaturation.

Direct analysis of membrane lipids by liquid chromatography-electrospray mass spectrometry was used to demonstrate the role of unsaturation in ether lipids in the adaptation of Methanococcoides burtonii to low temperature. A proteomics approach using two-dimensional liquid chromatography-mass spectrometry was used to identify enzymes involved in lipid biosynthesis, and a pathway for lipid biosynthesis was reconstructed from the M. burtonii draft genome sequence. The major phospholipids were archaeol phosphatidylglycerol, archaeol phosphatidylinositol, hydroxyarchaeol phosphatidylglycerol, and hydroxyarchaeol phosphatidylinositol. All phospholipid classes contained a series of unsaturated analogues, with the degree of unsaturation dependent on phospholipid class. The proportion of unsaturated lipids from cells grown at 4 degrees C was significantly higher than for cells grown at 23 degrees C. 3-Hydroxy-3-methylglutaryl coenzyme A synthase, farnesyl diphosphate synthase, and geranylgeranyl diphosphate synthase were identified in the expressed proteome, and most genes involved in the mevalonate pathway and processes leading to the formation of phosphatidylinositol and phosphatidylglycerol were identified in the genome sequence. In addition, M. burtonii encodes CDP-inositol and CDP-glycerol transferases and a number of homologs of the plant geranylgeranyl reductase. It therefore appears that the unsaturation of lipids may be due to incomplete reduction of an archaeol precursor rather than to a desaturase mechanism. This study shows that cold adaptation in M. burtonii involves specific changes in membrane lipid unsaturation. It also demonstrates that global methods of analysis for lipids and proteomics linked to a draft genome sequence can be effectively combined to infer specific mechanisms of key biological processes.

Adaptation, Physiological↗

DBTSS: DataBase of human Transcriptional Start Sites and full-length cDNAs.

Although the information of cDNAs is indispensable for analyzing gene function, most of the cDNA sequences stored in current databases are imperfect in the sense that they lack the precise information of 5' end termini. To overcome this difficulty, we have developed the oligo-capping method to obtain full-length cDNAs, the information of which has been partly deposited in public databases. In this study, we further constructed human cDNA libraries enriched in clones containing the cap structure to systematically explore the 5' end structure of expressed genes. Of approximately 217 402 5' end sequences obtained, 111 382 have been matched to cDNA sequences of known genes (7889 genes) and are presented in our new database, DataBase of Transcriptional Start Sites (DBTSS; http://elmo.ims.u-tokyo.ac.jp/dbtss/). Sequence comparison between our entries and those of a reference sequence database, RefSeq, revealed that 4683 (34%) of RefSeq sequences should be extended towards the 5' ends. We also mapped each sequence on the human draft genome sequence to identify its transcriptional start site, which provides us with more detailed information on distribution patterns of transcriptional start sites and adjacent regulatory regions.

5' Flanking Region↗