PubMed HealthSearch

SEARCH · PubMed Health

Results for “comparison genomic”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

A-liner: linear alignment visualizer for genome comparisons.

SUMMARY: A-liner is a flexible command-line tool for linear visualization of genome-scale sequence alignments, supporting outputs from multiple aligners and integrated visualization of annotations, highlights, quantitative tracks, and coordinate scales. It is applicable to a wide range of organisms, from bacteria to large eukaryotic genomes, and facilitates efficient generation of publication-ready comparative genome visualizations. AVAILABILITY AND IMPLEMENTATION: The source code and example output files for a-liner are available in the GitHub repository: https://github.com/mokuno3430/a-liner. A-liner v1.1.0 has been archived on Zenodo at https://doi.org/10.5281/zenodo.19702001.

Software

Characterization of Dapalides D and E and Genomic Comparison of the Two Co-Occurring Dapalide-Producing Dapis spp.

Marine cyanobacteria are a rich source of diverse bioactive natural products, targeting proteins involved in many diseases. Here, we combined metagenomic analysis to enhance the structure elucidation process of two new cyclodepsipeptides named dapalides D (1) and E (2) from a collection of a cyanobacterial mat containing multiple Dapis species from Guam. Dapalides D/E are composed of 11 amino acids, including multiple identical units with different configurations. Enantioselective amino acid identification of the acid hydrolyzate established the identity of amino acids, including the configuration of α/β-stereogenic centers. Identification and analysis of the dapalides D/E biosynthetic gene cluster from a metagenome-assembled genome aided the elucidation of α-configuration and establishment of the order of individual building blocks, collectively revealing the total structure. Phylogenomic analysis indicates that the dapalides D/E producer belongs to Dapis sp. (Dapis sp. VPG23-80 MAG-2), which shares a 95.2% average nucleotide identity with Dapis sp. VPG23-80 MAG-1, the producer of dapalides A-C that cooccurs in the same assemblage. Dapalide D (1) showed moderate growth inhibitory activity against various cancer cell lines. This work expands the dapalide structure class and further highlights the use of combined chemical and metagenomic analyses for natural product structure elucidation.

Cyanobacteria

Molecular cloning and physical mapping of the genome of simian herpes B virus and comparison of genome organization with that of herpes simplex virus type 1.

The molecular structure of the genome of simian herpes B virus (SHBV) was determined by restriction endonuclease mapping studies. Genomic DNA was cleaved with restriction endonucleases BamHI and SalI into 41 and 58 fragments, respectively. Most of these fragments were cloned into the plasmid vector pACYC184; uncloned fragments were identified following isolation from agarose gels. Terminal fragments were identified by exonuclease digestion and radioactive end-labelling, and linkage of fragments was deduced by a combination of single and double digest experiments and cross-blot hybridizations. The genome is larger than that of herpes simplex virus type 1 (HSV-1), being approximately 165 kilobase pairs. Like that of HSV-1, the SHBV genome is composed of a long and a short unique region each flanked by inverted repeat sequences, which allow the unique regions to invert relative to one another, resulting in four possible isomeric arrangements of the molecule. Genome locations of several SHBV genes were compared with their HSV-1 homologues.

Animals

SURE-Pipe: a pipeline to compare genomes and extract shared and unique regions.

Identification of unique and shared genomic regions between organisms has substantial translational potential for the development of marker-based diagnostic assays and sequence homology-driven taxonomic classification. An automated pipeline capable of performing genome comparisons at both the intra- and inter-species levels with minimal computational requirements can significantly advance genome-driven translational research. Species-specific genomic regions are particularly valuable for sequence-based species identification and for developing DNA amplification- or hybridization-based diagnostic assays. Here, we present SURE-Pipe, an automated and flexible pipeline for genome comparison and extraction of unique and shared genomic regions (https://github.com/BPaul-bioinfoLAB/SURE-Pipe). Benchmarking of this pipeline using simulated datasets demonstrated high accuracy for shared and unique region identification. Using the pairwise genome comparison module, six genome pairs from diverse microorganisms were analysed, and identified the unique and shared regions. In addition, the multigenome comparison module was applied to 96 genomes representing 24 Bacillus species and identified species-specific genomic regions. These regions were highly conserved among four strains of a species (>98% sequence identity) and exhibit little to no similarity with other species. Species-specific primers designed for all 24 Bacillus species showed no off-target amplification in in-silico polymerase chain reaction analysis, indicating their specificity. Overall, SURE-Pipe provides a robust and multipurpose framework for comparative genomics, and the outcomes can be used for species identification and the development of genome-based diagnostic approaches.

Genome, Bacterial

The complete nucleotide sequence of pepper mottle virus genomic RNA: comparison of the encoded polyprotein with those of other sequenced potyviruses.

The complete nucleotide sequence of a pepper mottle virus isolate from California (PepMoV C) has been determined from cloned viral cDNAs. The PepMoV C genomic RNA is 9640 nucleotides excluding the poly(A) tail and contains a long open reading frame starting at nucleotide 168 and potentially encoding a polyprotein of 3068 amino acids. Comparison of the PepMoV C presumptive polyprotein with those of other sequenced members of the potyvirus group, including tobacco etch virus (TEV), tobacco vein mottling virus (TVMV), plum pox virus (PPV), and potato virus Y (PVY), allowed localization of putative protein cleavage sites. A similar analysis was used to determine the position of conserved viral protein-coding regions along the viral genomic RNA. These analyses confirm previous work indicating that genome organization is conserved among members of the genus Potyvirus. The localization of one PepMoV C gene product, the nuclear inclusion body protein a (NIa protein), was analyzed by expressing PepMoV cDNA deletion clones in bacteria and assaying for appearance of mature-sized coat protein, a cleavage product of the NIa protease. Comparative sequence analyses of the putative PepMoV polyprotein with those of TEV, TVMV, PPV, and PVY served to identify regions of the potyviral polyproteins which have diverged within the genus, as well as highly conserved protein features which may play an important functional role in the potyviral life cycle.

Amino Acid Sequence

No more than seven interruptions in the ovalbumin gene: comparison of genomic and double-stranded cDNA sequences.

We have determined the sequence of ovalbumin RNA (ov-mRNA) using a double-stranded cDNA (dscDNA) plasmid. We have also determined the sequence of the previously characterized exonic regions of the chicken ovalbumin gene. The comparison of these various sequences has shown that there are no additional interruptions in the mRNA-coding sequences above those 7 already characterized. There is only one single base discrepancy between the two mRNA sequences determined using the dscDNA or the genomic clones. This demonstrates the accuracy and reproducibility of the cloning and sequencing techniques. The ovalbumin mRNA sequence was found to be 1872 nucleotides in length, 13 nucleotides larger than the previous value reported by McReynolds et al. [Nature 273, 723-728 (1978)].

Animals

Detection of short tandem repeats in the cattle genome: a comparison of bioinformatic tools.

BACKGROUND: Short tandem repeats (STRs) are repetitive DNA sequences with 1–6 nucleotide repeat units, exhibiting high polymorphism due to varying repeat counts. STRs are more variable than SNPs and can cause genetic disorders. With population-scale cattle whole-genome sequencing data available, whole-genome STR identification has attracted new interest, but challenges remain due to the lack of standardized methods, sequencing data limitations, and the diversity of STR-calling tools. This study compared six STR-calling tools: HipSTR, GangSTR, and ExpansionHunter for short-read data, and Straglr, RepeatHMM, and LongTR for Oxford Nanopore (ONT) long-read data—using sequences from five Holstein cattle (two parent–offspring trios with a shared sire). This is the first cattle study to evaluate short- and long-read STR callers using both data types from the same animals. RESULTS: In short-read data, ExpansionHunter identified the highest number of polymorphic STRs (pSTRs) (327,690), followed by HipSTR (205,900) and GangSTR (110,680), with 93,023 loci detected by all three tools. In long-read data, LongTR detected 470,250 pSTRs, RepeatHMM 224,185, and Straglr 90,275, with only 33,253 loci shared among them. Mendelian consistency of STR genotypes in the trio offspring was high (> 0.8) for all short-read tools, with HipSTR and GangSTR highest at 0.98. LongTR was the only long-read tool with high consistency (0.88). Short-read tools also showed higher concordance in STR genotypes among themselves than was observed among long-read tools. However, long-read tools had a clear advantage in detecting large STRs. Relative to computational efficiency, HipSTR and GangSTR (short-reads), and LongTR (long-reads) required less memory and shorter runtimes than the other tools. CONCLUSIONS: Tool selection is critical for accurate whole-genome STR identification in cattle. For short-read data, HipSTR showed relatively high Mendelian consistency and concordance compared to the other tools, while ExpansionHunter was able to detect longer STRs but with lower Mendelian consistency. For long-read data, LongTR demonstrated higher consistency and computational efficiency relative to the other tools. Based on these results, HipSTR and LongTR are suggested as preferred options for short-read and ONT long-read datasets, respectively, in cattle STR analysis. These recommendations are based on the metrics observed in this study, and confirmatory analyses across additional breeds, larger sample sizes, and validated truth sets are encouraged.

Animals

RNA synthesis of vesicular stomatitis virus. VIII. Oligonucleotides of the structural genes and mRNA.

The single-stranded RNA genome of vesicular stomatitis virus (VSV, Indiana serotype, San Juan strain) yields approx. 75 RNase T1-resistant oligonucleotides ranging in size from 10 to 50 bases. Each of the five structural genes, isolated as duplex RNA molecules hybridized to complementary mRNA, contains two or more of these large oligonucleotides. One of the oligonucleotides is identified as part of the non-coding region near the 3' end of the genome. Comparison of these results with others indicate that the RNA sequence of VSV is apparently stable in the laboratory but not in the wild. RNase T1-resistant oligonucleotides are also shown for all five VSV mRN species. Whether the mRNA for these digestions are are isolated from duplex RNA molecules or as single-stranded RNA species, the oligonucleotide patterns for each mRNA are virtually identical, indicating that each mRNA is transcribed from contiguous sequences on the genome. Comparison with published oligonucleotide patterns obtained from other isolates of VSV or from VSV deletion mutants indicate that identity and changes in their genome structure can be correlated with specific structural genes.

Base Sequence

Pleomorphic Liposarcoma: Comprehensive Genomic Analysis of 39 Cases With Comparison to Other Genomically Complex Sarcomas.

Pleomorphic liposarcoma (PLPS) is an aggressive high-grade sarcoma that often shows diverse morphological features and can mimic high-grade undifferentiated pleomorphic sarcoma (UPS)/spindle cell sarcoma or myxofibrosarcoma (MFS), especially when pleomorphic lipoblasts are sparse. The molecular profile of PLPS is distinct from well differentiated/dedifferentiated liposarcoma and myxoid liposarcoma. In this study, we investigate 39 cases of PLPS by comprehensive genomic profiling, occurring in 32 patients with available molecular data. Cases were reviewed and morphologic parameters-lipoblastic component, UPS-like, and MFS-like areas were estimated. The genomic findings were collected and compared to UPS and MFS groups studied using the same platform. The cohort included 15 females and 17 males, with a median age of 56.5 (range, 34-78). The lower extremity (n = 17) was the most common site involved, followed by upper extremity (n = 5) and pelvis (n = 5). UPS-like and MFS-like patterns were the most common morphologic variants, ranging from 15% to 95% and 20% to 90%, respectively. TP53 (87%) and RB1 (51%) mutations and copy number alterations were the most common alterations seen, followed by ATRX (36%). Compared to UPS and MFS, TP53 and RB1 gene alterations were significantly more common in PLPS. Conversely, CDKN2A/B deletions were infrequent in PLPS. Survival analysis showed that MYC amplification was associated with significantly shorter overall survival in PLPS. Among histologic variants, CYSLTR2 alterations were found to be highest in cases with predominantly pleomorphic lipoblasts; additionally, strong correlations were found between gene alteration frequencies of MFS and MFS-like PLPS, and between UPS and UPS-like PLPS. RB1 allele-specific copy number analysis showed loss of heterozygosity in 82% of cases. Our cohort of PLPS showed a complex molecular landscape with distinct genetic alterations, histologic correlations, and clinical outcomes, highlighting its unique position among genomically complex sarcomas and providing insights that may inform future diagnostic and therapeutic approaches.

Humans

The nucleotide sequence of adenovirus type 11 early 3 region: comparison of genome type Ad11p and Ad11a.

The early 3 region (E3) of two strains (genome type Ad11p and Ad11a) of human adenovirus serotype 11, causing persistent urinary and acute respiratory illnesses, respectively, has been identified and partially sequenced. The sequenced E3 regions of Ad11p and Ad11a were 1980 and 1966 bp long and encoded three complete ORFs, 18.5, 20.3, 20.6k within the Ad11p genome and 18.5, 20.3, 20.2k within the Ad11a genome. The sequence analysis of the 18.5k gene product demonstrated that a transmembrane domain and a cytoplasmic domain of Ad11p, Ad11a, and Ad35 was identical. Ad11p and Ad35 were homologous in the signal sequence. There was one amino acid mismatch between Ad11p and Ad11a, represented by an alanine instead of a proline. The endoplasmic reticulum lumenal domain, which binds to class I MHC, was relatively conserved between Ad11p and Ad11a with the exception of Glu80 and Glu104 in Ad11p, which were replaced by Gln80 and Lys104 in Ad11a. Within the 20.2k protein of Ad11a, the amino acid sequence Thr-Thr-Ser-His was deleted from a position immediately upstream the transmembrane region of the Ad11p 20.6k protein. The 9.0k E3 open reading frame (ORF) of Ad3 was deleted in the genomes of Ad11p and Ad11a. It is noteworthy that Ad11p and Ad35 which both cause persistent infection of the urinary tract display a remarkable similarity in several ORFs of the E3 region.

Adenovirus E3 Proteins

Structure and function of herpesvirus genomes. I. comparison of five HSV-1 and two HSV-2 strains by cleavage their DNA with eco R I restriction endonuclease.

The restriction endonuclease Eco R I cleaves HSV-1 and HSV-2 DNA into specific fragments that can be resolved by agarose gel electrophoresis. Comparison of HSV-1 strains KOS, 14-012, MP, F, and CI 101, and HSV-2 strains 333 and 186, suggests that the DNAs from type 1 strains are similar but not identical, and that the type 2 strains differ greatly from type 1 strains. The molecular lengths of the fragments have been determined by electron microscopy and can be used to calibrate gel electrophoretic analyses of DNA fragments.

Agar

GenomeDecoder: inferring segmental duplications in highly repetitive genomic regions.

MOTIVATION: The emergence of the 'telomere-to-telomere' genomics brought the challenge of identifying segmental duplications (SDs) in complete genomes. It further opened a possibility for identifying the differences in SDs across individual human genomes and studying the SD evolution. These newly emerged challenges require algorithms for reconstructing SDs in the most complex genomic regions that evaded all previous attempts to analyze their architecture, such as rapidly evolving immunoglobulin loci. RESULTS: We describe the GenomeDecoder algorithm for inferring SDs and apply it to analyzing genomic architectures of various loci in primate genomes. Our analysis revealed that multiple duplications/deletions led to a rapid birth/death of immunoglobulin genes within the human population and large changes in genomic architecture of immunoglobulin loci across primate genomes. Comparison of immunoglobulin loci across primate genomes suggests that they are subjected to diversifying selection. AVAILABILITY AND IMPLEMENTATION: GenomeDecoder is available at https://github.com/ZhangZhenmiao/GenomeDecoder. The software version and test data used in this paper are uploaded to https://doi.org/10.5281/zenodo.14753844.

Humans

Strain-specific differences in Neisseria gonorrhoeae associated with the phase variable gene repertoire.

BACKGROUND: There are several differences associated with the behaviour of the four main experimental Neisseria gonorrhoeae strains, FA1090, FA19, MS11, and F62. Although there is data concerning the gene complements of these strains, the reasons for the behavioural differences are currently unknown. Phase variation is a mechanism that occurs commonly within the Neisseria spp. and leads to switching of genes ON and OFF. This mechanism may provide a means for strains to express different combinations of genes, and differences in the strain-specific repertoire of phase variable genes may underlie the strain differences. RESULTS: By genome comparison of the four publicly available neisserial genomes a revised list of 64 genes was created that have the potential to be phase variable in N. gonorrhoeae, excluding the opa and pilC genes. Amplification and sequencing of the repeat-containing regions of these genes allowed determination of the presence of the potentially unstable repeats and the ON/OFF expression state of these genes. 35 of the 64 genes show differences in the composition or length of the repeats, of which 28 are likely to be associated with phase variation. Two genes were expressed differentially between strains causing disseminated infection and uncomplicated gonorrhoea. Further study of one of these in a range of clinical isolates showed this association to be due to sample size and is not maintained in a larger sample. CONCLUSION: The results provide us with more evidence as to which genes identified through comparative genomics are indeed phase variable. The study indicates that there are large differences between these four N. gonorrhoeae strains in terms of gene expression during in vitro growth. It does not, however, identify any clear patterns by which previously reported behavioural differences can be correlated with the phase variable gene repertoire.

Bacterial Proteins

Highly Contiguous Is Not Chromosomally Accurate: Integrated Cytogenetic and Genomic Mapping in Two Turtle Genome.

High-quality genome assemblies are essential for robust research across biological and medical fields. Assembly errors can have far-reaching consequences for downstream analyses, including gene annotation and the inference of synteny. In contrast to the rapid growth of genomic data volume, there is a notable lag in the integration of chromosome-level assemblies with cytogenetic data. We conducted the first direct genome-to-genome comparison, integrating comparative chromosome painting, the alignment of chromosome-specific probes to available genome assemblies, and synteny-based comparison of independent chromosome-level assemblies of the loggerhead sea turtle (Caretta caretta, 2n = 56) and the red-eared slider (Trachemys scripta elegans, 2n = 50). Using two independent sets of flow-sorted chromosome-specific probes in cross-species hybridizations, together with the sequencing and mapping of chromosome-derived DNA libraries, we assigned assembled scaffolds to all physical chromosomes of both species. In C. caretta, chromosomal assignments and genome-wide synteny were fully consistent with the published assembly, except for the reduced sizes of two microchromosome scaffolds, which we attribute to under-representation of repetitive DNA. In contrast, in T. s. elegans, cytogenetic validation of the assemblies revealed a false rearrangement compared to a missed one. Our results show that even highly contiguous vertebrate genome assemblies can misrepresent chromosome structure. When cytogenetic analyses reveal such inaccuracies, updated reference genomes should be generated for widely studied species to enable accurate inference of karyotype evolution and downstream comparative genomic analyses.

FISH

Comparison of herpesviruses isolated from reindeer, goats, and cattle by restriction endonuclease analysis.

A genomic comparison of bovine herpesvirus 1 (BHV-1), caprine herpesvirus (CHV-2) and reindeer herpesvirus (RHV), was performed using 5 restriction endonucleases. Cross neutralization of these three herpesviruses showed that BHV-1 and CHV-2 had a relatively low degree of cross reaction with heterologous viruses. RHV showed a higher degree of such cross reactivity. The restriction endonuclease analyses showed that the migration patterns of the DNA segments were different for the three groups of herpesviruses. The enteric caprine strain could be differentiated from genital strains using BstE II and Hpa I. The genome size of reindeer herpesvirus was estimated to be approximately 86.8 x 10(6) Da (131.8 kbp), and indications of isomerization of this genome were found. It is concluded that reindeer herpesvirus is a distinct species within the family Herpesviridae.

Animals