PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Complete genome characterization”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Characterization of the complete genomic sequence of genotype II hepatitis A virus (CF53/Berne isolate).

The complete genomic sequence of hepatitis A virus (HAV) CF53/Berne strain was determined. Pairwise comparison with other complete HAV genomic sequences demonstrated that the CF53/Berne isolate is most closely related to the single genotype VII strain, SLF88. This close relationship was confirmed by phylogenetic analyses of different genomic regions, and was most pronounced within the capsid region. These data indicated that CF53/Berne and SLF88 isolates are related more closely to each other than are subtypes IA and IB. A histogram of the genetic differences between HAV strains revealed four separate peaks. The distance values for CF53/Berne and SLF88 isolates fell within the peak that contained strains of the same subtype, showing that they should be subtypes within a single genotype. The complete genomic data indicated that genotypes II and VII should be considered a single genotype, based upon the complete VP1 sequence, and it is proposed that the CF53/Berne isolate be classified as genotype IIA and strain SLF88 as genotype IIB. The CF53/Berne isolate is cell-adapted, and therefore its sequence was compared to that of two other strains adapted to cell culture, HM-175/7 grown in MK-5 and GBM grown in FRhK-4 cells. Mutations found at nucleotides 3889, 4087 and 4222 that were associated with HAV attenuation and cell adaptation in HM175/7 and GMB strains were not present in the CF53/Berne strain. Deletions found in the 5'UTR and P3A regions of the CF53/Berne isolate that are common to cell-adapted HAV isolates were identified, however.

Base Sequence↗

Analysis and characterization of the complete genome of tupaia (tree shrew) herpesvirus.

The tupaia herpesvirus (THV) was isolated from spontaneously degenerating tissue cultures of malignant lymphoma, lung, and spleen cell cultures of tree shrews (Tupaia spp.). The determination of the complete nucleotide sequence of the THV strain 2 genome resulted in a 195,857-bp-long, linear DNA molecule with a G+C content of 66.5%. The terminal regions of the THV genome and the loci of conserved viral genes were found to be G+C richer. Furthermore, no large repetitive DNA sequences could be identified. This is in agreement with the previous classification of THV as the prototype species of herpesvirus genome group F. The search for potential coding regions resulted in the identification of 158 open reading frames (ORFs) regularly distributed on both DNA strands. Seventy-six out of the 158 ORFs code for proteins that are significantly homologous to known herpesvirus proteins. The highest homologies found were to primate and rodent cytomegaloviruses. Biological properties, protein homologies, the arrangement of conserved viral genes, and phylogenetic analysis revealed that THV is a member of the subfamily Betaherpesvirinae. The evolutionary lineages of THV and the cytomegaloviruses seem to have branched off from a common ancestor. In addition, it was found that the arrangements of conserved genes of THV and murine cytomegalovirus strain Smith, both of which are not able to form genomic isomers, are colinear with two different human cytomegalovirus (HCMV) strain AD169 genomic isomers that differ from each other in the orientation of the long unique region. The biological properties and the high degree of relatedness of THV to the mammalian cytomegaloviruses allow the consideration of THV as a model system for investigation of HCMV pathogenicity.

Animals↗

Characterization of the complete genome of the Tupaia (tree shrew) adenovirus.

The members of the family Adenoviridae are widely spread among vertebrate host species and normally cause acute but innocuous infections. Special attention is focused on adenoviruses because of their ability to transform host cells, their possible application in vector technology, and their phylogeny. The primary structure of the genome of Tupaia adenovirus (TAV), which infects Tupaia spp. (tree shrew) was determined. Tree shrews are taxonomically assumed to be at the base of the phylogenetic tree of mammals and are frequently used as laboratory animals in neurological and behavior research. The TAV genome is 33,501 bp in length with a G+C content of 49.96% and has 166-bp inverted terminal repeats. Analysis of the complete nucleotide sequence resulted in the identification of 109 open reading frames (ORFs) with a coding capacity of at least 40 amino acid residues. Thirty-eight of them are predicted to encode viral proteins based on the presence of transcription and translation signals and sequence and positional conservation. Thirty viral ORFs were found to show significant similarities to known adenoviral genes, arranged into discrete early and late genome regions as they are known from mastadenoviruses. Analysis of the nucleotide content of the TAV genome revealed a significant CG dinucleotide depletion at the genome ends that suggests methylation of these genomic regions during the viral life cycle. Phylogenetic analysis of the viral gene products, including penton and hexon proteins, viral protease, terminal protein, protein VIII, DNA polymerase, protein IVa2, and 100,000-molecular-weight protein, revealed that the evolutionary lineage of TAV forms a separate branch within the phylogenetic tree of the Mastadenovirus genus.

Adenoviridae↗

Complete genome sequence and genomic characterization of the probiotic Limosilactobacillus reuteri PSC102.

BACKGROUND: Gut microbiota are potential sources of probiotics and play an essential role in maintaining intestinal health. Limosilactobacillus reuteri PSC102 (L. reuteri PSC102), which was isolated from the feces of healthy pigs, exhibited health-beneficial properties. AIM: We aimed to conduct a whole-genome sequencing analysis of L. reuteri PSC102 to determine its molecular characteristics as a probiotic strain. METHODS: Limosilactobacillus reuteri PSC102 cells were cultured in De Man-Rogosa-Sharpe medium, followed by DNA extraction for genomic analysis using the PacBio-Illumina sequencing platform. The EzBioCloud software was used to perform gene assembly, and the genes were interpreted by the National Center for Biotechnology Information (NCBI) and the Glimmer program. Core and pan-genomic analyses were performed to assess the extent of functional conservation in the genomic sequence. Moreover, the NCBI database and the Basic Local Alignment Search Tool software were used to identify antimicrobial resistance genes and virulence factors. RESULTS: Limosilactobacillus reuteri PSC102 consists of a single circular chromosome with 2,048,626 bp, a guanine- cytosine of 38.9%, 18 rRNA genes, and 69 tRNA genes. Among the 1,846 protein-coding sequences, genes associated with probiotic characteristics were identified, including genes involved in host-microbe interactions, stress tolerance, biogenesis, and defense mechanisms. Furthermore, the genome of L. reuteri PSC102 comprises 2,446 pan-genome and 1,222 core-genome orthologous gene clusters. A total of 74 unique genes were identified in L. reuteri PSC102 genome. These genes mostly encode proteins potentially involved in the transport and metabolism of amino acids and carbohydrates. Moreover, antibacterial resistance genes and virulence factors were absent in L. reuteri PSC102. CONCLUSION: The results of the molecular insight into L. reuteri PSC102 corroborates its use as a probiotic in humans and other animals.

Limosilactobacillus reuteri↗

Characterization of the complete genomic structure of the human versican gene and functional analysis of its promoter.

Versican is a modular proteoglycan involved in the control of cellular growth and differentiation. To understand versican gene regulation and transcriptional control, we have isolated genomic clones spanning the entire gene locus including 5'- and 3'-flanking sequences. Versican was encoded by 15 exons encompassing over 90 kilobase pairs of continuous DNA. The exon organization corresponded to the protein subdomains encoded by homologous proteins, with a remarkable conservation of exon size and intron phase. We discovered an additional exon just proximal to the glycosaminoglycan-binding region that was identical to a recently identified splice variant of versican (Dours-Zimmermann, M.T., and Zimmermann, D.R. (1994) J. Biol. Chem. 269, 32992-32998). The versican promoter harbored a typical TATA box located approximately 16 base pairs upstream of the transcription start site and binding sites for a number of transcription factors involved in regulated gene expression. This promoter was shown to be highly functional in transiently transfected cells of both mesenchymal and epithelial origin. Stepwise 5' deletions identified a strong enhancer element between -209 and -445 base pairs and a strong negative element between -445 and -632 base pairs. This study provides the molecular basis for discerning the transcriptional control of the versican gene and offers the opportunity to investigate genetic disorders linked to this important human gene.

Base Sequence↗

Vibrio cholerae phage K139: complete genome sequence and comparative genomics of related phages.

In this report, we characterize the complete genome sequence of the temperate phage K139, which morphologically belongs to the Myoviridae phage family (P2 and 186). The prophage genome consists of 33,106 bp, and the overall GC content is 48.9%. Forty-four open reading frames were identified. Homology analysis and motif search were used to assign possible functions for the genes, revealing a close relationship to P2-like phages. By Southern blot screening of a Vibrio cholerae strain collection, two highly K139-related phage sequences were detected in non-O1, non-O139 strains. Combinatorial PCR analysis revealed almost identical genome organizations. One region of variable gene content was identified and sequenced. Additionally, the tail fiber genes were analyzed, leading to the identification of putative host-specific sequence variations. Furthermore, a K139-encoded Dam methyltransferase was characterized.

Bacteriophages↗

Complete genome analysis and molecular characterization of Usutu virus that emerged in Austria in 2001: comparison with the South African strain SAAR-1776 and other flaviviruses.

Here we describe the complete genome sequences of two strains of Usutu virus (USUV), a mosquito-borne member of the genus Flavivirus in the Japanese encephalitis virus (JEV) serogroup. USUV was detected in Austria in 2001 causing a high mortality rate in blackbirds; the reference strain (SAAR-1776) was isolated in 1958 from mosquitoes in South Africa and has never been associated with avian mortality. The Austrian and South African isolates exhibited 97% nucleotide and 99% amino acid identity. Phylogenetic trees were constructed displaying the genetic relationships of USUV with other members of the genus Flavivirus. When comparing USUV with other JEV serogroup viruses, the closest lineage was Murray Valley encephalitis virus (nt: 73%, aa: 82%) followed by JEV (nt: 71%, aa: 81%) and West Nile virus (nt: 68%, aa: 75%). Comparison of the genomes showed that the conserved structural elements and putative enzyme motifs were homologous in the two USUV strains and the JEV serogroup. The factors that determine the severe clinical symptoms caused by the Austrian USUV strain in Eurasian blackbirds are discussed. We also offer a possible explanation for the origins and dispersal of USUV, JEV, and MVEV out of Africa.

Amino Acid Sequence↗

The complete mitochondrial genome sequence and characterization of single-nucleotide polymorphisms in the control region of the Asian seabass (Lates calcarifer).

We determined the complete mtDNA nucleotide sequence of Lates calcarifer using the shotgun sequencing method. The mitochondrial DNA (mtDNA) was 16,535 base pairs (bp) in length, and contained 13 protein coding genes, 22 transfer RNAs, 2 ribosomal RNAs, and one major noncoding control region (CR). The CR was unusually short at only 768 bp. A striking feature of the mitochondrial genome was the high G+C content (46.1%), which is among the highest in fish. The gene order was identical to that of a typical vertebrate. Phylogenetic analyses using concatenated amino acid sequences of 12 protein-coding genes of 30 fish species representing 14 suborders clearly showed Lates calcarifer was located in the cluster of fish species from the order Perciformes, supporting the traditional systematic classification. We characterized single-nucleotide polymorphisms (SNPs) in the CR by sequencing the complete CR of 25 individuals obtained from Australia and Singapore. A total of 68 SNPs were detected. Eighteen SNPs were fixed with alternative nucleotides in Australian and Singapore seabass, and these SNPs could be used for differentiating fish from the two countries.

Animals↗

RetrOryza: a database of the rice LTR-retrotransposons.

Long terminal repeat (LTR)-retrotransposons comprise a significant portion of the rice genome. Their complete characterization is thus necessary if the sequenced genome is to be annotated correctly. In addition, because LTR-retrotransposons can influence the expression of neighboring genes, the complete identification of these elements in the rice genome is essential in order to study their putative functional interactions with the plant genes. The aims of the database are to (i) Assemble a comprehensive dataset of LTR-retrotransposons that includes not only abundant elements, but also low copy number elements. (ii) Provide an interface to efficiently access the resources stored in the database. This interface should also allow the community to annotate these elements. (iii) Provide a means for identifying LTR-retrotransposons inserted near genes. Here we present the results, where 242 complete LTR-retrotransposons have been structurally and functionally annotated. A web interface to the database has been made available (http://www.retroryza.org/), through which the user can annotate a sequence or search for LTR-retrotransposons in the neighborhood of a gene of interest.

Databases, Nucleic Acid↗

Complete genome sequence and biological characterizations of a novel goose paramyxovirus-SF02 isolated in China.

A paramyxovirus designated as APMV-1 (NDV) isolate SF02 (abbre. as SF02) was recently isolated from goose in China. SF02 was identified as a member of Newcastle disease virus (NDV) genotype VII. NDV strains are generally pathogenic only for fowls, including chicken and pigeon, and not for waterfowls such as goose and duck, whereas SF02 is highly pathogenic for both fowls and waterfowls. In the present study the complete genome consisting of 15, 192 nucleotides of SF02 was sequenced. Genomes of SF02 and all known APMV-1, Strains contain 6 ORFs in the order of NP-P-M-F-HN-L, and that of SF02 had an extra 6 nts between NP and P genes. Moreover, an anti-sense ORF consisting of 549 nt at the 1960 to 1412 and deduced 182 amino acids was found in SF02. The SF02 genome shared 83% identity and its 6 ORFs 81.9-86.1% identities with the reference APMV-1 strains. The possible mechanism determining different host range and pathogenicity is discussed based on genetic analyses.

Amino Acid Sequence↗

The Friend virus genome: partial characterization of a complete DNA copy.

A complementary DNA probe has been prepared from the Friend murine erythroleukaemia virus complex released by Friend cells (FV cDNAD-) and Friend cells induced to differentiate (FV cDNAD+). Molecular hybridization analysis shows that: (a) FV cDNAD+ is close to being a complete copy of the virus genome and the distribution of sequences is uniform with respect to their distribution in the Friend virus genome. (b) Hybridization of 70S RNA from the cloned helper virus to the total FVc DNAD+ probe demonstrates that a large proportion of the cDNA is specific to the transforming spleen focus forming virus. (c) Hybridization of the probe to normal and transformed cell DNA shows that there are about seven Friend virus related genes in normal DNA and almost twice this amount in transformed cell DNA. A significant minor proportion (20%) of the cDNA probe anneals only to virus related sequences in the transformed cell DNA. (d) An analysis of the kinetics of annealing of the cDNA to an excess template RNA shows that the minimum base sequence complexity of the Friend virus complex is 4 x 10(6). (e) An analysis of the cross hybridization between FV cDNAD+ and 60 to 70S RNA isolated from virus released by uninduced and induced cells shows that the genome of the induced and uninduced Friend virus is almost identical.

Animals↗

Discovery of diverse anellovirus sequences in Thai human sequencing data.

UNLABELLED: Anelloviruses are part of the normal human viral flora. Although their diversity in humans has been investigated in many countries, and despite their initial detection in Thailand in 1999, knowledge of Thai anelloviruses remains very limited. This study analyzed 1,175 whole-genome sequencing data sets from Thai individuals to mine for potential anellovirus sequences. Our analyses detected anellovirus sequences in 149 data sets (12.68%), uncovering 434 partial anellovirus sequences and 77 complete genome sequences, characterized by the presence of terminal redundancy, complete orf1, and the conserved untranslated region upstream of the orf1 gene. Sequence analyses indicated that these viruses belong to seven genera, including Alphatorquevirus, Betatorquevirus, Gammatorquevirus, Hetorquevirus, Lamedtorquevirus, Samektorquevirus, and Yodtorquevirus. Notably, Hetorquevirus, Lamedtorquevirus, Samektorquevirus, and Yodtorquevirus had not previously been reported in Thailand. Phylogenetic analysis of ORF1 protein sequences showed that Thai anelloviruses form multiple phylogenetic clusters with non-Thai anelloviruses, indicating frequent cross-country transmission and multiple origins of the virus in Thailand. Furthermore, sequence similarity network analysis identified 33 potentially novel anellovirus species in our data set. Our findings greatly expand the knowledge of anellovirus diversity in Thailand and demonstrate the potential of human whole-genome sequencing data as a valuable resource for viral discovery. Lastly, we highlight and discuss some challenges with the use of the current pairwise sequence similarity-based classification scheme, in particular, how gaps can influence similarity calculation and potentially lead to inconsistencies with a phylogenetic-based classification scheme. IMPORTANCE: Anelloviruses are widespread in humans, yet their diversity remains poorly characterized in many regions, including Thailand. Here, we demonstrate that human sequencing data sets, originally generated without the intention for virome research, can be effectively mined for anellovirus sequences, including complete genomes. Our findings reveal a substantial number of previously unreported anelloviruses in Thailand, significantly expanding the known diversity of the virus. We also highlight potential limitations of the current anellovirus species classification scheme, which is based on pairwise orf1 sequence similarity analysis with a hard threshold cutoff at 69%. Our results reveal that the current scheme can sometimes yield taxonomic groupings that are inconsistent with phylogenetic relationships, particularly when significant alignment gaps are present. Overall, our results show that existing human sequencing data can be effectively repurposed for virus discovery research and suggest the need for more robust and phylogenetically informed classification frameworks as viral sequence databases continue to expand.

Humans↗

Complete nucleotide sequences and genome characterization of double-stranded RNA 1 and RNA 2 in the Raphanus sativus-root cv. Yidianhong [corrected].

Four distinct double-stranded (ds) RNA bands were extracted from leaves of Raphanus sativus-root cv. Yidianhong [corrected] with yellowing at the leaf edge in China. Purified viral particles of 28-30 nm in diameter contained dsRNA segments with the same number and mobility as these extracted directly from radish leaves. The two major dsRNA segments, namely RasR 1 and RasR 2, were 1866 and 1791 bp in length, respectively. Computer analysis predicted that they both contained a single open reading frame (ORF) on their plus-stranded RNA, putatively encoding a RNA dependent RNA polymerase and a capsid protein similar to that encoded by members of the family Partitiviridae. In addition, both RasR 1 and RasR 2 were highly conserved at the 5' untranslated regions (UTR) and had an adenosine-uracil rich stretch at the 3' UTR, with an identical terminal motif (5'-AAAAUAAAACC-3'). Taken together, these results suggest that the two major dsRNA segments constitute the genome of a partitivirus infecting radish.

3' Untranslated Regions↗

Isolation and identification of an enterovirus 77 recovered from a refugee child from Kosovo, and characterization of the complete virus genome.

The complete nucleotide sequence of an enterovirus 77 isolate is reported. The virus designated FR/CF496-99 (France/Clermont-Ferrand 496-1999) was recovered from the feces of a 4-year-old child hospitalized for Salmonella gastroenteritis. The virus was identified by a molecular typing assay based on the genomic sequence encoding the VP1 capsid protein. The phylogenetic analysis based on the VP1 sequence demonstrated that the enterovirus isolated in the child clustered with viruses included in the human enterovirus B species (HEV-B) and was most closely related to enterovirus 77. A sliding window analysis of the complete genome showed an overall nucleotide similarity >80% between the P3 genomic region of the FR/CF496-99 isolate and that of the echovirus 30 prototype strain. A comparative analysis based on partial 3D(pol) sequences showed that the FR/CF496-99 virus was more closely related to recent enteroviruses from different serotypes and different geographical areas than to the prototype strains collected in the 1950s. This suggests that, in this enterovirus, the 3D(pol) encoding sequence is of recent origin.

Capsid Proteins↗

TBL2, a novel transducin family member in the WBS deletion: characterization of the complete sequence, genomic structure, transcriptional variants and the mouse ortholog.

Williams-Beuren syndrome (WBS) is a developmental disorder with multi-system manifestations caused by haploinsufficiency for contiguous genes deleted in chromosome region 7q11.23. The size of the deletion is similar in most patients due to a genomic duplication that predisposes to unequal meiotic crossover events. While hemizygosity at the elastin locus is responsible for the cardiovascular features, the contribution of other genes to the WBS phenotype remains to be demonstrated. We have identified a novel gene, TBL2, in the common WBS deletion. TBL2 is expressed as a 2. 4-kb transcript predominantly in testis, skeletal muscle, heart and some endocrine tissues, with a larger approximately 5-kb transcript detected ubiquitously at lower levels. TBL2 encodes a protein with four putative WD40-repeats. An alternatively spliced transcript in TBL2 introduces a novel second exon with an in frame stop codon. This mRNA encodes a 75 amino acid protein with 43 amino acids identical to TBL2 at the N-terminus and no known functional domain. The mouse homolog, Tbl2, shows 84% sequence identity at the nucleotide level and 92% similarity at the amino acid level. Comparison of the mouse and human sequences identifies a conserved region that extends upstream of the previously published sequence with an initiation codon common to both species that adds 21 amino acids at the N-terminus. The Tbl2 gene has been mapped to mouse chromosome 5 in a region of conserved synteny with human 7q11.23. Since haploinsufficiency has been shown for other WD-repeat containing proteins, hemizygosity of TBL2 may contribute to some of the aspects of the complex WBS phenotype.

Amino Acid Sequence↗

5' Long serial analysis of gene expression (LongSAGE) and 3' LongSAGE for transcriptome characterization and genome annotation.

Complete genome annotation relies on precise identification of transcription units bounded by a transcription initiation site (TIS) and a polyadenylation site (PAS). To facilitate this process, we developed a set of two complementary methods, 5' Long serial analysis of gene expression (LS) and 3'LS. These analyses are based on the original SAGE and LS methods coupled with full-length cDNA cloning, and enable the high-throughput extraction of the first and the last 20 bp of each transcript. We demonstrate that the mapping of 5'LS and 3'LS tags to the genome allows the localization of TIS and PAS. By using 537 tag pairs mapping to the region of known genes, we confirmed that >90% of the tag pairs appropriately assigned to the first and last exons. Moreover, by using tag sequences as primers for RT-PCRs, we were able to recover putative full-length transcripts in 81% of the attempts. This large-scale generation of transcript terminal tags is at least 20-40 times more efficient than full-length cDNA cloning and sequencing in the identification of complete transcription units. The apparent precision and deep coverage makes 5'LS and 3'LS an advanced approach for genome annotation through whole-transcriptome characterization.

Animals↗

NPC1: Complete genomic sequence, mutation analysis, and characterization of haplotypes.

Niemann-Pick type C disease (NP-C) is a rare, autosomal recessive lipid storage disorder. At least 96% of all NP-C patients link to NPC1 which encodes for a lysosomally-targeted protein. We describe the complete genomic sequence of 57,052 kb corresponding to the transcribed region of human NPC1 including several exonic and intronic single nucleotide polymorphisms (SNPs). Sequencing of all exons, splice sites, and the promoter region of NPC1 in 12 unrelated Caucasian NP-C patients revealed nine novel and four known most likely disease-causing mutations. Ten unique mutations found only once in 24 disease alleles were observed in patients being compound heterozygous for two different mutations. Two of the three missense mutations identified more than once were observed in a total of four patients homozygous for the respective mutation along with homozygosity for the underlying haplotype. The patients were offspring of most likely nonconsanguineous couples. Based upon genotyping exonic SNPs c.2572A>G (I858V; g.45020A>G) and c.2793C>T (N931N; g.45686C>T) and segregation analysis we characterized the haplotype of all 24 NPC1 alleles and of 138 alleles of healthy Caucasian control subjects. All four permutations between the two SNPs were identified in the control alleles: 2572A-2793C (50%), 2572G-2793T (41%), 2572G-2793C (5%), and 2572A-2793T (4%). These data are suggestive for an ancestral intragenic recombination within a genomic fragment of <666 bp. While 17 of 24 NP-C alleles (71%) shared haplotype 2572G-2793T, this haplotype accounted for only 41% in the controls (p=0.007; 2-sided Fisher exact test) suggesting the possibility of an influence of the haplotypic background on expression of missense mutations in NPC1.

Carrier Proteins↗