PubMed HealthSearch

SEARCH · PubMed Health

Results for “complete genome”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

JC virus Type 1 has multiple subtypes: three new complete genomes.

The complete genomes of three new Type 1 strains of JC virus (JCV) from urine have been analysed. These were subtype 1A, subtype 1B and Type 4 as assigned from a short typing fragment in the VP1 gene. They differ from Mad1 (subtype 1A) by less than 1.0% of the DNA sequence. Based on its complete genome, the JCV Type 4 strain falls into a Type 1 subgroup. Type 4, with several Type 3-like sites in the short typing fragment, is a possible recombinant strain. The consensus of Type 1 DNA sequences is distinguished within the coding region from both Type 2 (strain GS/B) and five Type 3 (African and African American) strains at 64 sites. Most mutations are silent, but at 21 positions amino acid changes occur. Our findings define the subtypes of JCV Type 1 and support the validity of genotyping within the short VP1 fragment.

Base Sequence

Complete genome sequence of Methanobacterium thermoautotrophicum deltaH: functional analysis and comparative genomics.

The complete 1,751,377-bp sequence of the genome of the thermophilic archaeon Methanobacterium thermoautotrophicum deltaH has been determined by a whole-genome shotgun sequencing approach. A total of 1,855 open reading frames (ORFs) have been identified that appear to encode polypeptides, 844 (46%) of which have been assigned putative functions based on their similarities to database sequences with assigned functions. A total of 514 (28%) of the ORF-encoded polypeptides are related to sequences with unknown functions, and 496 (27%) have little or no homology to sequences in public databases. Comparisons with Eucarya-, Bacteria-, and Archaea-specific databases reveal that 1,013 of the putative gene products (54%) are most similar to polypeptide sequences described previously for other organisms in the domain Archaea. Comparisons with the Methanococcus jannaschii genome data underline the extensive divergence that has occurred between these two methanogens; only 352 (19%) of M. thermoautotrophicum ORFs encode sequences that are >50% identical to M. jannaschii polypeptides, and there is little conservation in the relative locations of orthologous genes. When the M. thermoautotrophicum ORFs are compared to sequences from only the eucaryal and bacterial domains, 786 (42%) are more similar to bacterial sequences and 241 (13%) are more similar to eucaryal sequences. The bacterial domain-like gene products include the majority of those predicted to be involved in cofactor and small molecule biosyntheses, intermediary metabolism, transport, nitrogen fixation, regulatory functions, and interactions with the environment. Most proteins predicted to be involved in DNA metabolism, transcription, and translation are more similar to eucaryal sequences. Gene structure and organization have features that are typical of the Bacteria, including genes that encode polypeptides closely related to eucaryal proteins. There are 24 polypeptides that could form two-component sensor kinase-response regulator systems and homologs of the bacterial Hsp70-response proteins DnaK and DnaJ, which are notably absent in M. jannaschii. DNA replication initiation and chromosome packaging in M. thermoautotrophicum are predicted to have eucaryal features, based on the presence of two Cdc6 homologs and three histones; however, the presence of an ftsZ gene indicates a bacterial type of cell division initiation. The DNA polymerases include an X-family repair type and an unusual archaeal B type formed by two separate polypeptides. The DNA-dependent RNA polymerase (RNAP) subunits A', A", B', B" and H are encoded in a typical archaeal RNAP operon, although a second A' subunit-encoding gene is present at a remote location. There are two rRNA operons, and 39 tRNA genes are dispersed around the genome, although most of these occur in clusters. Three of the tRNA genes have introns, including the tRNAPro (GGG) gene, which contains a second intron at an unprecedented location. There is no selenocysteinyl-tRNA gene nor evidence for classically organized IS elements, prophages, or plasmids. The genome contains one intein and two extended repeats (3.6 and 8.6 kb) that are members of a family with 18 representatives in the M. jannaschii genome.

Anaerobiosis

Reconstruction of amino acid biosynthesis pathways from the complete genome sequence.

The complete genome sequence of an organism contains information that has not been fully utilized in the current prediction methods of gene functions, which are based on piece-by-piece similarity searches of individual genes. We present here a method that utilizes a higher level information of molecular pathways to reconstruct a complete functional unit from a set of genes. Specifically, a genome-by-genome comparison is first made for identifying enzyme genes and assigning EC numbers, which is followed by the reconstruction of selected portions of the metabolic pathways by use of the reference biochemical knowledge. The completeness of the reconstructed pathway is an indicator of the correctness of the initial gene function assignment. This feature has become possible because of our efforts to computerize the current knowledge of metabolic pathways under the KEGG project. We found that the biosynthesis pathways of all 20 amino acids were completely reconstructed in Escherichia coli, Haemophilus influenzae, and Bacillus subtilis, and probably in Synechocystis and Saccharomyces cerevisiae as well, although it was necessary to assume wider substrate specificity for aspartate aminotransferases.

Amino Acids

A strategy for finding regions of similarity in complete genome sequences.

MOTIVATION: Complete genomic sequences will become available in the future. New methods to deal with very large sequences (sizes beyond 100 kb) efficiently are required. One of the main aims of such work is to increase our understanding of genome organization and evolution. This requires studies of the locations of regions of similarity. RESULTS: We present here a new tool, ASSIRC ('Accelerated Search for SImilarity Regions in Chromosomes'), for finding regions of similarity in genomic sequences. The method involves three steps: (i) identification of short exact chains of fixed size, called 'seeds', common to both sequences, using hashing functions; (ii) extension of these seeds into putative regions of similarity by a 'random walk' procedure; (iii) final selection of regions of similarity by assessing alignments of the putative sequences. We used simulations to estimate the proportion of regions of similarity not detected for particular region sizes, base identity proportions and seed sizes. This approach can be tailored to the user's specifications. We looked for regions of similarity between two yeast chromosomes (V and IX). The efficiency of the approach was compared to those of conventional programs BLAST and FASTA, by assessing CPU time required and the regions of similarity found for the same data set. AVAILABILITY: Source programs are freely available at the following address: ftp://ftp.biologie.ens. fr/pub/molbio/assirc.tar.gz CONTACT: vincens@biologie.ens.fr, hazout@urbb.jussieu.fr

Algorithms

Complete genome sequences of cellular life forms: glimpses of theoretical evolutionary genomics.

The availability of complete genome sequences of cellular life forms creates the opportunity to explore the functional content of the genomes and evolutionary relationships between them at a new qualitative level. With the advent of these sequences, the construction of a minimal gene set sufficient for sustaining cellular life and reconstruction of the genome of the last common ancestor of bacteria, eukaryotes, and archaea become realistic, albeit challenging, research projects. A version of the minimal gene set for modern-type cellular life derived by comparative analysis of two bacterial genomes, those of Haemophilus influenzae and Mycoplasma genitalium, consists of approximately 250 genes. A comparison of the protein sequences encoded in these genes with those of the proteins encoded in the complete yeast genome suggests that the last common ancestor of all extant life might have had an RNA genome.

Bacterial Proteins

Prediction of transcriptional regulatory sites in the complete genome sequence of Escherichia coli K-12.

MOTIVATION: As one of the best-characterized free-living organisms, Escherichia coli and its recently completed genomic sequence offer a special opportunity to exploit systematically the variety of regulatory data available in the literature in order to make a comprehensive set of regulatory predictions in the whole genome. RESULTS: The complete genome sequence of E.coli was analyzed for the binding of transcriptional regulators upstream of coding sequences. The biological information contained in RegulonDB (Huerta, A.M. et al., Nucleic Acids Res.,26,55-60, 1998) for 56 different transcriptional proteins was the support to implement a stringent strategy combining string search and weight matrices. We estimate that our search included representatives of 15-25% of the total number of regulatory binding proteins in E.coli. This search was performed on the set of 4288 putative regulatory regions, each 450 bp long. Within the regions with predicted sites, 89% are regulated by one protein and 81% involve only one site. These numbers are reasonably consistent with the distribution of experimental regulatory sites. Regulatory sites are found in 603 regions corresponding to 16% of operon regions and 10% of intra-operonic regions. Additional evidence gives stronger support to some of these predictions, including the position of the site, biological consistency with the function of the downstream gene, as well as genetic evidence for the regulatory interaction. The predictions described here were incorporated into the map presented in the paper describing the complete E.coli genome (Blattner,F.R. et al., Science, 277, 1453-1461, 1997). AVAILABILITY: The complete set of predictions in GenBank format is available at the url: http://www. cifn.unam.mx/Computational_Biology/E.coli-predictions CONTACT: ecoli-reg@cifn.unam.mx, collado@cifn.unam.mx

Bacterial Proteins

Identification of a ribonuclease H gene in both Mycoplasma genitalium and Mycoplasma pneumoniae by a new method for exhaustive identification of ORFs in the complete genome sequences.

Exhaustive identification of open reading frames in complete genome sequences is a difficult task. It is possible that important genes are missed. In our efforts to reanalyze the intergenic regions of Mycoplasma genitalium and Mycoplasma pneumoniae, we have newly identified a number of new open reading frames (ORFs) in both M. genitalium and M. pneumoniae. The most significant identification was that of a ribonuclease H enzyme in both species which until now has not been identified or assumed absent and interpreted as such. In this paper we discuss the biological importance of RNase H and its evolutionary implication. We also stress the usefulness of our method for identifying new ORFs by reanalyzing intergenic regions of existing ORFs in complete genome sequences.

Genome, Bacterial

Salmonella Pullorum strain SPullorum-YN-07 from dead embryos of Yanjin black-bone chickens: Complete genome with IncFII(S) and Col(pVC) plasmids and pathogenicity.

Salmonella Pullorum is a host-adapted pathogen that causes Pullorum disease in chickens and can be vertically transmitted via eggs, leading to embryonic mortality. The susceptibility and vertical transmission of S. Pullorum may vary among chicken breeds, yet genomic characterization of strains from dead embryos of indigenous breeds remains limited. This study isolated and characterized a Gram-negative short rod, designated Salmonella Pullorum strain SPullorum-YN-07, from dead embryos of Yanjin black-bone chickens, a native breed in Yunnan, China. The strain formed colorless colonies on MacConkey agar and red, non-H2S colonies on XLD agar, with biochemical reactions consistent with the genus Salmonella. Whole-genome sequencing using Illumina and PacBio platforms generated a complete genome consisting of one circular chromosome and four circular plasmids; plasmid replicon types IncFII(S) and Col(pVC) were identified in two of the plasmids. On the chromosome, a total of 340 virulence-associated genes were detected, including those involved in secretion systems, adhesion, motility, and immune modulation. Resistance gene analysis identified the acquired aminoglycoside resistance gene aac(6')-Iaa, alongside multiple intrinsic resistance determinants related to efflux pumps and target alteration. Multilocus sequence typing (MLST) assigned the strain to sequence type ST92, and core-genome phylogenetic analysis confirmed its clustering within the Salmonella Pullorum lineage. In a chick infection model, the strain induced depression, white diarrhea, and growth retardation, with clinical scores peaking at 10 days post-infection and a mortality rate of 10%. Bacterial colonization was highest in the cecum, and histopathological lesions were observed in the liver, spleen, and cecum. This study provides the first complete genomic characterization and pathogenicity assessment of an S. Pullorum strain isolated from dead embryos of Yanjin black-bone chickens, offering a foundation for understanding host-pathogen interactions in indigenous breeds and assessing cross-transmission risks to commercial poultry populations.

Complete genome

Helicobacter pylori unmasked--the complete genome sequence.

The publication of the complete genome sequence of Helicobacter pylori and the computer analysis of the genes it contains represents an enormous advance in H. pylori research. Besides providing an overview of H. pylori biology, the genome sequence will aid future research and radically alter the research process. In particular, it will enable a 'top down' approach whereby genes are investigated because of their similarity to genes of known function in other organisms. Other advances in biotechnology will allow all H. pylori genes to be studied simultaneously, proteins to be identified quickly, and potential drug targets to be evaluated efficiently. Progress in H. pylori research should now be rapid, and with the research interest and resources focused on it, H. pylori could become the paradigm for post-genomic research.

Genetics, Microbial

The complete genome of the hyperthermophilic bacterium Aquifex aeolicus.

Aquifex aeolicus was one of the earliest diverging, and is one of the most thermophilic, bacteria known. It can grow on hydrogen, oxygen, carbon dioxide, and mineral salts. The complex metabolic machinery needed for A. aeolicus to function as a chemolithoautotroph (an organism which uses an inorganic carbon source for biosynthesis and an inorganic chemical energy source) is encoded within a genome that is only one-third the size of the E. coli genome. Metabolic flexibility seems to be reduced as a result of the limited genome size. The use of oxygen (albeit at very low concentrations) as an electron acceptor is allowed by the presence of a complex respiratory apparatus. Although this organism grows at 95 degrees C, the extreme thermal limit of the Bacteria, only a few specific indications of thermophily are apparent from the genome. Here we describe the complete genome sequence of 1,551,335 base pairs of this evolutionarily and physiologically interesting organism.

Chromosome Mapping

Classification of all putative permeases and other membrane plurispanners of the major facilitator superfamily encoded by the complete genome of Saccharomyces cerevisiae.

On the basis of the complete genome sequence of the budding yeast Saccharomyces cerevisiae, a computer-aided analysis was carried out of all members of the major facilitator superfamily (MFS), which typically consists of permeases with 12 transmembrane spans. Analysis of all 5885 predicted open reading frames identified 186 potential MFS proteins. Binary sequence comparison made it possible to cluster 149 of them into 23 families. Putative permease functions could be assigned to 12 families, the largest including sugar, amino acid, and multidrug transport. Phylogenetic clustering of proteins allowed us to predict a possible permease function for a total of 119 proteins. Multiple sequence alignments were made for all families, and evolutionary trees were constructed for families with at least four members. The latter resulted in the identification of 21 subclusters with presumably tightly related permease function. No functional clues were predicted for a total of 41 clustered or unclustered proteins.

Computers

Complete genome sequence of Treponema pallidum, the syphilis spirochete.

The complete genome sequence of Treponema pallidum was determined and shown to be 1,138,006 base pairs containing 1041 predicted coding sequences (open reading frames). Systems for DNA replication, transcription, translation, and repair are intact, but catabolic and biosynthetic activities are minimized. The number of identifiable transporters is small, and no phosphoenolpyruvate:phosphotransferase carbohydrate transporters were found. Potential virulence factors include a family of 12 potential membrane proteins and several putative hemolysins. Comparison of the T. pallidum genome sequence with that of another pathogenic spirochete, Borrelia burgdorferi, the agent of Lyme disease, identified unique and common genes and substantiates the considerable diversity observed among pathogenic spirochetes.

Bacterial Proteins

The first complete genome sequence of Ammi majus latent virus from the new natural host culantro.

A potyvirus (isolate AMLV-CQ) infecting culantro (Eryngium foetidum L.) imported from Vietnam was identified by RT-PCR. The complete genome sequence of AMLV-CQ was determined to be 9,549 nucleotides in length. It contains a large open reading frame encoding a 3,082-amino-acid putative polyprotein, flanked by 5´ and 3´ untranslated regions (UTRs) of 77 and 226 nt, respectively. AMLV-CQ is closely related to five other completely sequenced potyviruses, sharing 68-69% nucleotide and 69-70% amino acid sequence identity. However, the coat protein (CP) gene shares 89% nucleotide and 93% amino acid sequence identity with that of a partially sequenced potyvirus, Ammi majus latent virus (isolate AMLV-WF17). These results suggest that AMLV-CQ and AMLV-WF17 are isolates of the same species. To our knowledge, this is the first report of a complete genome sequence of an AMLV isolate, and culantro was identified as a new natural host for this virus. In addition, a one-step RT-PCR assay was developed that provides a rapid, robust, and highly sensitive approach for the detection of AMLV.

Eryngium

Complete genomes from a xenic Dolichospermum flosaquae FBCC-A233 culture reveal genome-inferred metabolic asymmetry with associated bacteria.

Cyanobacteria form phycosphere communities with associated bacteria, but genome-resolved resources are needed to formulate testable hypotheses about their metabolic interactions. Here, we reconstructed three complete circular genomes from a unialgal xenic culture, including Dolichospermum flosaquae FBCC-A233 and two associated alphaproteobacterial genomes assigned to Sphingorhabdus sp. and Brevundimonas sp. Genome-wide read mapping and genome-quality assessment supported the three recovered genomes as high-quality circular reconstructions. Comparative genome analysis placed the cyanobacterial genome within the Dolichospermum flosaquae species cluster under the GTDB framework, while the associated bacterial genomes represented Sphingorhabdus sp. and a putative undescribed Brevundimonas species-level lineage. Genome architecture analysis indicated reduced genome size and gene content in Brevundimonas relative to genus-level references although additional metrics did not support a strong conclusion of classical genome streamlining. Selected KEGG module and KO-level reconstructions indicated genome-inferred metabolic asymmetries across the consortium. FBCC-A233 encoded photosynthesis- and nitrogen-related modules and a BioU-mediated de novo biotin biosynthesis route, whereas the associated bacteria lacked complete de novo biotin biosynthesis but retained biotin-dependent carboxylase genes. FBCC-A233 also encoded extensive anaerobic corrinoid biosynthesis potential; however, canonical DMB-containing cobalamin completion, cobamide identity, and complete transporter systems were not resolved. Together, these complete genomes provide a genome-resolved resource for investigating genome-inferred metabolic differentiation and ecological interactions in cyanobacteria-associated bacterial consortia.IMPORTANCEPhycosphere interactions between cyanobacteria and associated bacteria can shape aquatic microbial communities, but many proposed interactions remain difficult to evaluate without genome-resolved resources. This study provides three complete circular genomes from a unialgal xenic Dolichospermum flosaquae culture, capturing the cyanobacterium and two co-maintained bacterial associates. Our analysis identifies genome-inferred metabolic asymmetries, particularly in biotin- and cobamide-related pathways. D. flosaquae FBCC-A233 encoded candidate de novo biotin and corrinoid biosynthesis capacity, whereas the associated bacteria lacked complete de novo pathways but retained cofactor-dependent enzymes. These findings nominate cofactor-related dependencies as experimentally testable hypotheses while emphasizing unresolved uptake, export, cobamide identity, and growth-dependence mechanisms. The complete genomes and KO-level reconstructions generated here provide a resource for future studies of cyanobacteria-associated consortia.

Genome, Bacterial