PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “draft genome sequence”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

HUGE: a database for human large proteins identified in the Kazusa cDNA sequencing project.

We have been developing a HUGE database to summarize results from the sequence analysis of human novel large (>4 kb) cDNAs identified in the Kazusa cDNA sequencing project, systematically designated KIAA plus a four-digit number. HUGE currently contains nearly 2000 gene/protein characteristic tables harboring the results of the computer-assisted analysis of the cDNA and the predicted protein sequences together with those of expression profiling and chromosomal mapping. In the updated version of HUGE, we made it possible to compare each KIAA cDNA sequence with the corresponding entry in the human draft genome sequence that was published recently. Approximately 90% of KIAA cDNAs in HUGE can be localized along the human genome for at least half or more of the cDNA's length. Any nucleotide differences between the cDNA and the corresponding genomic sequences are also presented in detail. This new version of HUGE greatly helps us evaluate the completeness of cDNA clones and the accuracy of cDNA/genomic sequences. More interestingly, in some cases, the ability to compare cDNA with genomic sequences allows us to identify candidate sites of RNA editing. HUGE is available on the World Wide Web at http://www.kazusa.or.jp/huge.

Amino Acid Sequence↗

Succinate dehydrogenase functioning by a reverse redox loop mechanism and fumarate reductase in sulphate-reducing bacteria.

Sulphate- or sulphur-reducing bacteria with known or draft genome sequences (Desulfovibrio vulgaris, Desulfovibrio desulfuricans G20, Desulfobacterium autotrophicum [draft], Desulfotalea psychrophila and Geobacter sulfurreducens) all contain sdhCAB or frdCAB gene clusters encoding succinate : quinone oxidoreductases. frdD or sdhD genes are missing. The presence and function of succinate dehydrogenase versus fumarate reductase was studied. Desulfovibrio desulfuricans (strain Essex 6) grew by fumarate respiration or by fumarate disproportionation, and contained fumarate reductase activity. Desulfovibrio vulgaris lacked fumarate respiration and contained succinate dehydrogenase activity. Succinate oxidation by the menaquinone analogue 2,3-dimethyl-1,4-naphthoquinone depended on a proton potential, and the activity was lost after degradation of the proton potential. The membrane anchor SdhC contains four conserved His residues which are known as the ligands for two haem B residues. The properties are very similar to succinate dehydrogenase of the Gram-positive (menaquinone-containing) Bacillus subtilis, which uses a reverse redox loop mechanism in succinate : menaquinone reduction. It is concluded that succinate dehydrogenases from menaquinone-containing bacteria generally require a proton potential to drive the endergonic succinate oxidation. Sequence comparison shows that the SdhC subunit of this type lacks a Glu residue in transmembrane helix IV, which is part of the uncoupling E-pathway in most non-electrogenic FrdABC enzymes.

Amino Acid Sequence↗

Genome sequences of four bacterial strains isolated from the phyllosphere of Mangifera indica trees in the polluted tropical city of Medellín, Colombia.

Complete and draft genome sequences of four phyllosphere-associated bacterial strains (Microbacterium radiodurans, Brachybacterium rhamnosum, Sphingomonas citri, and Curtobacterium sp.) isolated from Mangifera indica leaves in polluted Medellín, Colombia, are presented. These resources enable future studies on plant-microbe interactions and phyllosphere microbial mediation of atmospheric pollutants under urban stress.

Mangifera indica↗

Development and characterization of a normalized canine retinal cDNA library for genomic and expression studies.

PURPOSE: Identification of causative mutations for retinal blinding disorders is often limited by restricted understanding of gene expression and underlying molecular mechanisms that trigger degenerative processes. This study was conducted to develop a catalog of canine retina-expressed genes that would provide a unique tool to investigate normal and altered function in the adult retina. Because of the conserved syntenies between the dog and human, this approach would identify new potential disease candidate genes for both species. METHODS: A canine normalized retinal cDNA library was produced and analyzed by using a modified PhredPhrap algorithm. Computerized annotation provided gene homology and chromosomal location for individual clones and contigs in a Web-accessible database. RESULTS: From 6316 cDNA clones, 3980 retinal expressed sequence tags (ESTs) were derived. Homology to the canine genome draft sequence was found for more than 99% of all ESTs, but only for 32% when compared with annotated canine cDNAs. Functional analysis suggests an enrichment of this library for genes involved with eye function and development, chaperone, or ribosomal functions when compared with mouse and human National Center for Biotechnology Information (NCBI) RefSeq entries. CONCLUSIONS: A combination of annotation approaches with ongoing mapping and expression studies provide functional data covering at least 27% to 30% of the currently proposed canine catalog of genes expressed in the retina. This is an essential first step toward establishing an integrated network for gene identification and expression patterns suitable for functional genetics, comparative genomics and evolutionary analysis of genes and gene families with respect to the developmental and degenerative processes of the retina.

Animals↗

EGassembler: online bioinformatics service for large-scale processing, clustering and assembling ESTs and genomic DNA fragments.

Expressed sequence tag (EST) sequencing has proven to be an economically feasible alternative for gene discovery in species lacking a draft genome sequence. Ongoing large-scale EST sequencing projects feel the need for bioinformatics tools to facilitate uniform EST handling. This brings about a renewed importance for a universal tool for processing and functional annotation of large sets of ESTs. EGassembler (http://egassembler.hgc.jp/) is a web server, which provides an automated as well as a user-customized analysis tool for cleaning, repeat masking, vector trimming, organelle masking, clustering and assembling of ESTs and genomic fragments. The web server is publicly available and provides the community a unique all-in-one online application web service for large-scale ESTs and genomic DNA clustering and assembling. Running on a Sun Fire 15K supercomputer, a significantly large volume of data can be processed in a short period of time. The results can be used to functionally annotate genes, to facilitate splice alignment analysis, to link the transcripts to genetic and physical maps, design microarray chips, to perform transcriptome analysis and to map to KEGG metabolic pathways. The service provides an excellent bioinformatics tool to research groups in wet-lab as well as an all-in-one-tool for sequence handling to bioinformatics researchers.

Computational Biology↗

Identification of promoter regions in the human genome by using a retroviral plasmid library-based functional reporter gene assay.

Attempts to identify regulatory sequences in the human genome have involved experimental and computational methods such as cross-species sequence comparisons and the detection of transcription factor binding-site motifs in coexpressed genes. Although these strategies provide information on which genomic regions are likely to be involved in gene regulation, they do not give information on their functions. We have developed a functional selection for promoter regions in the human genome that uses a retroviral plasmid library-based system. This approach enriches for and detects promoter function of isolated DNA fragments in an in vitro cell culture assay. By using this method, we have discovered likely promoters of known and predicted genes, as well as many other putative promoter regions based on the presence of features such as CpG islands. Comparison of sequences of 858 plasmid clones selected by this assay with the human genome draft sequence indicates that a significantly higher percentage of sequences align to the 500-bp segment upstream of the transcription start sites of known genes than would be expected from random genomic sequences. We also observed enrichment for putative promoter regions of genes predicted in at least two annotation databases and for clones overlapping with CpG islands. Functional validation of randomly selected clones enriched by this method showed that a large fraction of these putative promoters can drive the expression of a reporter gene in transient transfection experiments. This method promises to be a useful genome-wide function-based approach that can complement existing methods to look for promoters.

3T3 Cells↗

Up-regulation of Frizzled-7 (FZD7) in human gastric cancer.

Human Frizzled-7 (FZD7) and human FzE3, showing 98.8% nucleotide identity, encode almost identical WNT receptors with nine amino-acid substitutions. FzE3 is claimed to be expressed specifically in esophageal cancer. We determined the structure of the FZD7 gene and the FZD7 cDNA expressed in esophageal cancer. The FZD7 gene without intron and the FZD7 cDNAs isolated from esophageal cancer cell lines TE4 and TE5 were found to encode WNT receptor identical to FZD7, but not to FzE3. Nucleotide sequence of FzE3 was not identified on the human genome draft sequence. Thus, we could not obtain any data suggesting the existence of FzE3. Expression profile of FZD7 was also investigated. FZD7 was expressed throughout normal gastrointestinal tract, from esophagus to rectum. Among human esophageal and gastric cancer cell lines, expression level of FZD7 was relatively lower in esophageal cancer cell lines, and was highest in the gastric cancer cell line MKN7. FZD7 was up-regulated in one out of six cases of human primary gastric cancer. As over-expression of Frizzled-7 leads to activation of the WNT-beta-catenin-TCF pathway, up-regulation of FZD7 in human gastric cancer might play key roles in carcinogenesis through activation of the WNT-beta-catenin-TCF pathway.

Blotting, Northern↗

VMD: a community annotation database for oomycetes and microbial genomes.

The VBI Microbial Database (VMD) is a database system designed to host a range of microbial genome sequences. At present, the database contains genome sequence and annotation data of two plant pathogens Phytophthora sojae and Phytophthora ramorum. With the completion of the draft genome sequences of these pathogens in collaboration with the DOE Joint Genome Institute (JGI), we have created this resource to make the sequences publicly available. The genome sequences (95 MB for P.sojae and 65 MB for P.ramorum) were annotated with approximately 19,000 and approximately 16,000 gene models, respectively. We used two different statistical methods to validate these gene models, Fickett's and a log-likelihood method. Functional annotation of the gene models is based on results from BlastX and InterProScan screens. From the InterProScan results, we could assign putative functions to 17,694 genes in P.sojae and 14,700 genes in P.ramorum. We created an easy-to-use genome browser to view the genome sequence data, which opens to detailed annotation pages for each gene model. A community annotation interface is available for registered community members to add or edit annotations. There are approximately 1600 gene models for P.sojae and approximately 700 models for P.ramorum that have already been manually curated. A toolkit is provided as an additional resource for users to perform a variety of sequence analysis jobs. The database is publicly available at http://phytophthora.vbi.vt.edu/.

Databases, Nucleic Acid↗

Potential for retroposition by old Alu subfamilies.

Alu elements sharing sequence characteristics of the "old" subfamilies are thought to currently be retrotranspositionally inactive. We analyzed one of these old subfamilies of Alu elements, Sx, for sequence conservation relative to the consensus and the length of the "A-tail" as parameters to define the presence of potential Alu Sx source genes in the human genome. Sequence identity to the left half or the right half of the Alu Sx consensus sequence was evaluated for 4424 complete elements obtained from the human genome draft sequence. A small subset of Alu Sx left halves were found to be more conserved than any of the Alu Sx right halves. Selection for promoter function in active elements may explain the slightly higher conservation of the left half. In order to determine whether this sequence identity was the result of recent activity, or simply sequence conservation for older elements, PCR amplification of some of the loci containing Sx elements with conserved left/right halves from different primate genomes was carried out. Several of these Sx Alus were found to have amplified at a later evolutionary period (<35 mya) than expected based on previous studies of Sx elements. Analysis of "A-tail" length, a feature correlated with current retroposition activity, varied between Alu Sx element loci in different primates, where the length increased in specific Alu elements in the human genome. The presence of few conserved Alu Sx elements and the dynamic expansion/contraction of the A-tail suggests that some of these older subfamilies may still be active at very low levels or in a few individuals.

Alu Elements↗

The structure and evolution of centromeric transition regions within the human genome.

An understanding of how centromeric transition regions are organized is a critical aspect of chromosome structure and function; however, the sequence context of these regions has been difficult to resolve on the basis of the draft genome sequence. We present a detailed analysis of the structure and assembly of all human pericentromeric regions (5 megabases). Most chromosome arms (35 out of 43) show a gradient of dwindling transcriptional diversity accompanied by an increasing number of interchromosomal duplications in proximity to the centromere. At least 30% of the centromeric transition region structure originates from euchromatic gene-containing segments of DNA that were duplicatively transposed towards pericentromeric regions at a rate of six-seven events per million years during primate evolution. This process has led to the formation of a minimum of 28 new transcripts by exon exaptation and exon shuffling, many of which are primarily expressed in the testis. The distribution of these duplicated segments is nonrandom among pericentromeric regions, suggesting that some regions have served as preferential acceptors of euchromatic DNA.

Animals↗

RibAlign: a software tool and database for eubacterial phylogeny based on concatenated ribosomal protein subunits.

BACKGROUND: Until today, analysis of 16S ribosomal RNA (rRNA) sequences has been the de-facto gold standard for the assessment of phylogenetic relationships among prokaryotes. However, the branching order of the individual phlya is not well-resolved in 16S rRNA-based trees. In search of an improvement, new phylogenetic methods have been developed alongside with the growing availability of complete genome sequences. Unfortunately, only a few genes in prokaryotic genomes qualify as universal phylogenetic markers and almost all of them have a lower information content than the 16S rRNA gene. Therefore, emphasis has been placed on methods that are based on multiple genes or even entire genomes. The concatenation of ribosomal protein sequences is one method which has been ascribed an improved resolution. Since there is neither a comprehensive database for ribosomal protein sequences nor a tool that assists in sequence retrieval and generation of respective input files for phylogenetic reconstruction programs, RibAlign has been developed to fill this gap. RESULTS: RibAlign serves two purposes: First, it provides a fast and scalable database that has been specifically adapted to eubacterial ribosomal protein sequences and second, it provides sophisticated import and export capabilities. This includes semi-automatic extraction of ribosomal protein sequences from whole-genome GenBank and FASTA files as well as exporting aligned, concatenated and filtered sequence files that can directly be used in conjunction with the PHYLIP and MrBayes phylogenetic reconstruction programs. CONCLUSION: Up to now, phylogeny based on concatenated ribosomal protein sequences is hampered by the limited set of sequenced genomes and high computational requirements. However, hundreds of full and draft genome sequencing projects are on the way, and advances in cluster-computing and algorithms make phylogenetic reconstructions feasible even with large alignments of concatenated marker genes. RibAlign is a first step in this direction and may be particularly interesting to scientists involved in whole genome sequencing of representatives of new or sparsely studied eubacterial phyla. RibAlign is available at http://www.megx.net/ribalign.

Algorithms↗

Molecular cloning and characterization of human WNT7B.

WNT signaling molecules are implicated in carcinogenesis and embryogenesis. Only partial coding sequence of human WNT7B is reported so far, and human genome draft sequence corresponding to the WNT7B gene in human chromosome 22q13 region is not available at present. Here, we have cloned human WNT7B cDNAs, spanning the complete coding sequence, by using rapid amplification of cDNA ends (RACE) and cDNA-PCR. WNT7B encoded a 349-amino-acid polypeptide with three N-linked glycosylation sites and consensus amino-acid residues conserved among members of the WNT family. WNT7B showed 77.1% total-amino-acid identity with WNT7A. The 4.0-kb WNT7B was moderately expressed in fetal brain, weakly expressed in fetal lung and kidney, and faintly expressed in adult brain, lung and prostate. Expression levels of WNT7B mRNA in a lung cancer cell line A549, esophageal cancer cell lines TE2, TE3, TE4, TE5, TE6, TE7, TE10, TE12, a gastric cancer cell line TMK1, and pancreatic cancer cell lines BxPC-3, AsPC-1 and Hs766T were significantly higher than that in fetal kidney. In addition, WNT7B was up-regulated in 5 out of 10 cases of primary gastric cancer. These results strongly suggest that WNT7B might play important roles in various types of human cancer.

Amino Acid Sequence↗

Identification and expression pattern of Bmlark, a homolog of the Drosophila gene lark in Bombyx mori.

Two Bombyx mori isoforms of the gene lark, which is shown to play an important role in Drosophila circadian rhythms, were identified and named Bmlark-PA and Bmlark-PB, respectively. Bmlark-PA consists of 5 exons and encodes a protein of 343 amino acid residues which contains 3 functional domains: two RRM (RNA recognization motif) domains and an RTZF (retroviral-type zinc finger) and shares 72% identity with the Drosophila gene lark at the amino acid level. Bmlark-PB lacks the sequence between 118 and 791 nt of Bmlark-PA and codes for a protein of 68 amino acid residues, which contains no distinct functional domains. Alignments of the cDNAs of Bmlark to the genomic draft sequence of B. mori showed that the gene Bmlark had a single copy in the genome, suggesting that an alternative splicing mechanism occurs in the gene Bmlark. RT-PCR analysis indicated that Bmlark-PA was expressed only in late pupae and adult but Bmlark-PB was broadly expressed in many tissues and throughout the developmental stages from embryo to adult.

Alternative Splicing↗

Seeing chordate evolution through the Ciona genome sequence.

A draft sequence of the compact genome of the sea squirt Ciona intestinalis, a non-vertebrate chordate that diverged very early from other chordates, including vertebrates, illuminates how chordates originated and how vertebrate developmental innovations evolved.

Animals↗

Construction of new unencapsulated (rough) strains of Streptococcus pneumoniae.

To construct rough strains of Streptococcus pneumoniae in which the capsule locus was completely deleted, a genetic cassette to be used as a donor DNA in transformation was developed. The cassette contained an aphIII gene, conferring kanamycin resistance, flanked by segments of dexB and aliA. Since, in all strains of S. pneumoniae the capsule locus is between dexB and aliA, the DNA segments of these two genes allow insertion of a 1354-bp DNA fragment containing aphIII into the pneumococcal chromosome, determining the deletion of the whole capsule locus. The capsule locus was deleted from the classic type 2 and type 3 Avery's strains, from R6 (whose complete genome sequence is released) and Rx1 (the two most commonly used transformation recipients), from a type 3 clinical strain and type 19F clinical isolate G54 (whose draft genome sequence is annotated). The effect of capsule removal was tested in 4 isogenic pairs. In unencapsulated strains, growth rate increased up to 56% and transformation frequency increased up to 1075-fold. A correlation was observed between the increase in growth rate and an increase in transformation frequency.

Anti-Bacterial Agents↗

Evolution and origin of vomeronasal-type odorant receptor gene repertoire in fishes.

BACKGROUND: In teleost fishes that lack a vomeronasal organ, both main odorant receptors (ORs) and vomeronasal receptors family 2 (V2Rs) are expressed in the olfactory epithelium, and used for perception of water-soluble chemicals. In zebrafish, it is known that both ORs and V2Rs formed multigene families of about a hundred copies. Whereas the contribution of V2Rs in zebrafish to olfaction has been found to be substantially large, the composition and structure of the V2R gene family in other fishes are poorly known, compared with the OR gene family. RESULTS: To understand the evolutionary dynamics of V2R genes in fishes, V2R sequences in zebrafish, medaka, fugu, and spotted green pufferfish were identified from their draft genome sequences. There were remarkable differences in the number of intact V2R genes in different species. Most V2R genes in these fishes were tightly clustered in one or two specific chromosomal regions. Phylogenetic analysis revealed that the fish V2R family could be subdivided into 16 subfamilies that had diverged before the separation of the four fishes. Genes in two subfamilies in zebrafish and another subfamily in medaka increased in their number independently, suggesting species-specific evolution in olfaction. Interestingly, the arrangements of V2R genes in the gene clusters were highly conserved among species in the subfamily level. A genomic region of tetrapods corresponding to the region in fishes that contains the V2R cluster was found to have no V2R gene in any species. CONCLUSION: Our results have indicated that the evolutionary dynamics of fish V2Rs are characterized by rapid gene turnover and lineage-specific phylogenetic clustering. In addition, the present phylogenetic and comparative genome analyses have shown that the fish V2Rs have expanded after the divergence between teleost and tetrapod lineages. The present identification of the entire V2R repertoire in fishes would provide useful foundation to the future functional and evolutionary studies of fish V2R gene family.

Animals↗

Sequence-based, in situ detection of chromosomal abnormalities at high resolution.

We developed single copy probes from the draft genome sequence for fluorescence in situ hybridization (scFISH) which precisely delineate chromosome abnormalities at a resolution equivalent to genomic Southern analysis. This study illustrates how scFISH probes detect cryptic and subtle abnormalities and localize the sites of chromosome rearrangements. scFISH probes are substantially shorter than conventional recombinant DNA-derived probes, and C(o)t1 DNA is not required to suppress repetitive sequence hybridization. In this study, 74 single copy sequence probes (>1,500 bp) have been developed from >/=100 kb genomic intervals associated with either constitutional or acquired disorders. Applications of these probes include detection of congenital microdeletion syndromes on chromosomes 1, 4, 7, 15, 17, 22 and submicroscopic deletions involving the imprinting center on chromosome 15q11.2q13. We demonstrate how hybridization with multiple combinations of probes derived from the Smith-Magenis syndrome interval on chromosome 17 identified a patient with an atypical, proximal deletion breakpoint. A similar multi-probe hybridization strategy has also been used to delineate the translocation breakpoint region on chromosome 9 in chronic myelogenous leukemia. Probes have also been designed to hybridize to multiple cis paralogs, both enhancing the chromosomal target size and detecting chromosome rearrangements, for example, by splitting and separating a family of related sequences flanking an inversion breakpoint on chromosome 16 in acute myelogenous leukemia. These novel strategies for rapid and precise characterization of cytogenetic abnormalities are feasible because of the sequence-defined properties and dense euchromatic organization of single copy probes.

Adult↗

Genomic inventory and expression of Sox and Fox genes in the cnidarian Nematostella vectensis.

The Sox and Forkhead (Fox) gene families are comprised of transcription factors that play important roles in a variety of developmental processes, including germ layer specification, gastrulation, cell fate determination, and morphogenesis. Both the Sox and Fox gene families are divided into subgroups based on the amino acid sequence of their respective DNA-binding domains, the high-mobility group (HMG) box (Sox genes) or Forkhead domain (Fox genes). Utilizing the draft genome sequence of the cnidarian Nematostella vectensis, we examined the genomic complement of Sox and Fox genes in this organism to gain insight into the nature of these gene families in a basal metazoan. We identified 14 Sox genes and 15 Fox genes in Nematostella and conducted a Bayesian phylogenetic analysis comparing HMG box and Forkhead domain sequences from Nematostella with diverse taxa. We found that the majority of bilaterian Sox groups have clear Nematostella orthologs, while only a minority of Fox groups are represented, suggesting that the evolutionary pressures driving the diversification of these gene families may be distinct from one another. In addition, we examined the expression of a subset of these genes during development in Nematostella and found that some of these genes are expressed in patterns consistent with roles in germ layer specification and the regulation of cellular behaviors important for gastrulation. The diversity of expression patterns among members of these gene families in Nematostella reinforces the notion that despite their relatively simple morphology, cnidarians possess much of the molecular complexity observed in bilaterian taxa.

Amino Acid Sequence↗