PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genomics”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Trends between gene content and genome size in prokaryotic species with larger genomes.

Although the evolution process and ecological benefits of symbiotic species with small genomes are well understood, these issues remain poorly elucidated for free-living species with large genomes. We have compared 115 completed prokaryotic genomes by using the Clusters of Orthologous Groups database to determine whether there are changes with genome size in the proportion of the genome attributable to particular cellular processes, because this may reflect both cellular and ecological strategies associated with genome expansion. We found that large genomes are disproportionately enriched in regulation and secondary metabolism genes and depleted in protein translation, DNA replication, cell division, and nucleotide metabolism genes compared to medium- and small-sized genomes. Furthermore, large genomes do not accumulate noncoding DNA or hypothetical ORFs, because the portion of the genome devoted to these functions remained constant with genome size. Traits other than genome size or strain-specific processes are reflected by the dispersion around the mean for cell functions that showed no correlation with genome size. For example, Archaea had significantly more genes in energy production, coenzyme metabolism, and the poorly characterized category, and fewer in cell membrane biogenesis and carbohydrate metabolism than Bacteria. The trends we noted with genome size by using Clusters of Orthologous Groups were confirmed by our independent analysis with The Institute for Genomic Research's Comprehensive Microbial Resource and Kyoto Encyclopedia of Genes and Genomes' Orthology annotation databases. These trends suggest that larger genome-sized species may dominate in environments where resources are scarce but diverse and where there is little penalty for slow growth, such as soil.

Animals↗

A systematic method to identify genomic islands and its applications in analyzing the genomes of Corynebacterium glutamicum and Vibrio vulnificus CMCP6 chromosome I.

MOTIVATION: Some genomic islands contain horizontally transferred genes, which play critical roles in altering the genotypes and phenotypes of organisms, and horizontal gene transfer has been recognized as a universal event throughout bacterial evolution. A windowless method to display the distribution of genomic GC content, the cumulative GC profile, is proposed to identify genomic islands in genomes whose complete genome sequences are available. Two new indices are proposed to assess the codon usage bias and amino acid usage bias in genomic islands. RESULTS: A 211 kb genomic island (CGGI-1) has been identified in the genome of Corynebacterium glutamicum, and three genomic islands VVGI-1, VVGI-2 and VVGI-3, with lengths 167, 40 and 33 kb, respectively, have been identified in the genome of Vibrio vulnificus CMCP6 chromosome I. The CGGI-1 is flanked by two approximately 500 bp direct repeats, and utilizes a Val-tRNA as the integration site. For the VVGI-1 and VVGI-2, each has an integrase gene at 5' junction. All the identified genomic islands show unusual GC content, codon usage and amino acid usage, compared with the rest of the genomes. In addition, it is found that genomic islands are fairly homogenous in terms of GC content variation. An index, h, to quantify the homogeneity of GC content for genomic islands is proposed, and it is shown that h is less than 0.1 for all the genomic islands analyzed. The cumulative GC profile, as well as various indices to assess the codon usage bias, amino acid usage bias and homogeneity of the genomic islands, will be useful in the analysis of other genomes. AVAILABILITY: Programs used in this work and numerical results are available upon request.

Algorithms↗

Comprehensive genomic and computational insights into Brucella suis: pan-genome analysis, evolutionary perspectives, and in-silico vaccine design.

BACKGROUND: Brucella suis is a zoonotic intracellular pathogen responsible for brucellosis, mainly in swine and humans. Although numerous genome sequences are publicly available, an integrative genomic analysis combining pan-genome architecture, structural organization, evolutionary relationships, and vaccine-associated targets remains limited. RESULTS: In this study, we analyzed 91 publicly available B.suis genomes to characterize their pan-genome composition and genomic structure. The pan-genome exhibited an open configuration, indicating continued genomic diversification. A total of 2,146 core genes were identified, representing conserved functions essential for species maintenance, while the accessory genome reflected strain-level variability. Phylogenetic reconstruction based on single-copy orthologs revealed distinct evolutionary clades among the strains. A complementary phylogenetic analysis of pan-genome gene presence-absence patterns further supported clade differentiation and highlighted variation in accessory gene repertoires. Comparative synteny and genome structural analyses demonstrated largely conserved chromosomal organization with localized rearrangements across strains. Screening of the core proteome identified 64 putative antigenic proteins with predicted surface localization and immunogenic properties. Additionally, resistance-associated determinants related to tetracycline and doxycycline were detected in one genome within the dataset. CONCLUSIONS: This comprehensive genomic analysis defines the pan-genome structure, evolutionary relationships, and genome organization of B.suis. The integration of core and pan-genome-based phylogenies provides complementary insights into strain diversification, while the identified conserved antigenic candidates offer a foundation for future experimental validation and rational vaccine development strategies.

Genome, Bacterial↗

Genome-specific repetitive DNA and RAPD markers for genome identification in Elymus and Hordelymus.

We have developed RFLP and RAPD markers specific for the genomes involved in the evolution of Elymus species, i.e., the St, Y, H, P, and W genomes. Two P genome specific repetitive DNA sequences, pAgc1 (350 bp) and pAgc30 (458 bp), and three W genome specific sequences, pAuv3 (221 bp), pAuv7 (200 bp), and pAuv13 (207 bp), were isolated from the genomes of Agropyron cristatum and Australopyrum velutinum, respectively. Attempts to find Y genome specific sequences were not successful. Primary-structure analysis demonstrated that pAgc1 (P genome) and pAgc30 (P genome) share 81% similarity over a 227-bp stretch. The three W genome specific sequences were also highly homologous. Sequence comparison analysis revealed no homology to sequences in the EMBL-GenBank databases. Three to four genome-specific RAPD markers were found for each of the five genomes. Genome-specific bands were cloned and demonstrated to be mainly low-copy sequences present in various Triticeae species. The RFLP and RAPD markers obtained, together with the previously described H and St genome specific clones pHch2 and pP1Taq2.5 and the Ns genome specific RAPD markers were used to investigate the genomic composition of a few Elymus species and Hordelymus europaeus, whose genome formulas were unknown. Our results demonstrate that only three of eight Elymus species examined (the tetraploid species Elymus grandis and the hexaploid species Elymus caesifolius and Elymus borianus) really belong to Elymus.

Base Sequence↗

Interactions among genomic structure, function, and evolution revealed by comprehensive analysis of the Arabidopsis thaliana genome.

The genome in a higher organism consists of a number of types of nucleotide sequence-specialized components, with each having tens of thousands of members or elements. It is crucial for our understanding of how a genome as an entity is organized, functions, and evolves to determine how these components are organized in the genome and how they relate with each other; however, no such knowledge is available. Here, we report a comprehensive analysis of the organization and interaction of all 40 components constituting the genome of the plant model species, Arabidopsis thaliana, at the whole-genome and chromosome levels. The 40 components include (i) 6 genome structural components consisting of GC%, genes, retrotransposons, DNA transposons, simple repeats, and low complex repeats; (ii) 3 evolutionarily critical features consisting of recombination rate, nucleotide substitutions, and nucleotide insertions/deletions; and (iii) 31 categories of genes with different functions and numbers of functions. We show that the distributions of 39 of the 40 components of the genome (excepting GC%) deviate significantly from the random distribution model and different types of the genome components are significantly correlated. These results remained to be true even when the genomic regions, such as centromeric regions, where transposable and repeat elements are abundant were excluded from the analyses. These findings suggest that DNA molecules contained in the Arabidopsis genome are each organized and structured from their constituting components in an unambiguous manner and that different types of the components that constitute or characterize the genome interact. The analysis also showed that each chromosome consists of a similar set of the components at similar densities, suggesting that the unique organization and interaction pattern of the components in each chromosome may represent, at least in part, the identity of a chromosome or a genome at the genome level, thus partly accounting for the phenotypic variation among different species. The data also provide comprehensive and new insights into many phenomena significant in genome biology, with which we particularly discuss the variation of genetic recombination. The variation of genetic recombination rate along a chromosomal arm is shaped, not only by the distribution of simple repeats, retrotransposons, DNA transposons, and nucleotide substitutions, but also by the functions of genes contained, especially those with multiple functions, suggesting that variation of genetic recombination along a chromosomal arm is the result of interactions among the components constituting local genome structure, function, and evolution.

Arabidopsis↗

Importance of anchor genomes for any plant genome project.

Progress in agricultural and environmental technologies is hampered by a slower rate of gene discovery in plants than animals. The vast pool of genes in plants, however, will be an important resource for insertion of genes, via biotechnological procedures, into an array of plants, generating unique germ plasms not achievable by conventional breeding. It just became clear that genomes of grasses have evolved in a manner analogous to Lego blocks. Large chromosome segments have been reshuffled and stuffer pieces added between genes. Although some genomes have become very large, the genome with the fewest stuffer pieces, the rice genome, is the Rosetta Stone of all the bigger grass genomes. This means that sequencing the rice genome as anchor genome of the grasses will provide instantaneous access to the same genes in the same relative physical position in other grasses (e.g., corn and wheat), without the need to sequence each of these genomes independently. (i) The sequencing of the entire genome of rice as anchor genome for the grasses will accelerate plant gene discovery in many important crops (e.g., corn, wheat, and rice) by several orders of magnitudes and reduce research and development costs for government and industry at a faster pace. (ii) Costs for sequencing entire genomes have come down significantly. Because of its size, rice is only 12% of the human or the corn genome, and technology improvements by the human genome project are completely transferable, translating in another 50% reduction of the costs. (iii) The physical mapping of the rice genome by a group of Japanese researchers provides a jump start for sequencing the genome and forming an international consortium. Otherwise, other countries would do it alone and own proprietary positions.

Databases, Factual↗

Reference-Guided Chromosome-Scale Genome Assembly With Insights on Population Genomics of the Atlantic Goliath Grouper (Epinephelus itajara), Islas del Rosario, Colombia.

Epinephelus itajara, commonly known as the Atlantic Goliath grouper, is the largest species among the western North Atlantic groupers and is critically endangered. This species plays a crucial ecological, cultural, and economic role and has been the focus of captive breeding efforts at the Oceanario of the Rosario Islands, Colombia. However, despite its ecological and conservation importance, genomic resources and population genomic data for E. itajara remain scarce, particularly in the Colombian Caribbean. This study presents a reference-guided chromosome-scale genome assembly and an analysis of the population genomic structure of E. itajara using PacBio HiFi sequencing and Illumina technologies. The assembled genome has a total size of 1.12 Gb, with a contig N50 of 42.69 Mb and a scaffold N50 of 46.30 Mb. A total of 22,692 protein-coding genes were identified after masking 46% of the genome, which consists of repetitive elements. Comparative genomic analyses revealed a high degree of collinearity with closely related Epinephelus species and identified E. lanceolatus as the closest relative, supporting recent divergence and conserved genome architecture within the genus. Additionally, a population genomics analysis was conducted using 7706 high-quality SNPs to assess the genomic structure of captive populations. The results revealed four distinct genomic lineages, with moderate genetic differentiation among the sampled individuals. In the Colombian Caribbean, two unique lineages were identified, associated with the localities of Bahía Cispatá and Bahía Barbacoas, suggesting possible geographic isolation. These genomic resources provide valuable tools and new opportunities to better understand the genomic diversity, evolutionary history, and reproductive mechanisms of E. itajara. Moreover, they serve as a foundation for conservation strategies, including selective breeding programs aimed at increasing genomic diversity in captive populations and guiding restoration efforts in its natural habitat.

Epinephelus itajara↗

Whole genome sequencing and phylogenetic classification accelerate the implementation of respiratory syncytial virus genomic surveillance in Canada: a pilot study.

UNLABELLED: Whole genome sequencing (WGS) has emerged as a powerful tool to facilitate the study of existing and emerging infectious diseases. WGS-based genomic surveillance provides information on the genetic diversity and tracks the evolution of important viral pathogens, including respiratory syncytial virus (RSV). Multiplex tiling polymerase chain reaction (PCR) assays have been used to facilitate sequencing of a variety of pathogens in support of genomics-based surveillance initiatives. We developed, optimized, and implemented multiplex tiling PCR assays for RSVA and RSVB capable of generating near-complete genomes in the majority of contemporaneous specimens tested. A pilot data set comprising 52 RSVA and 37 RSVB genomes derived from Canadian clinical specimens during the 2022-2023 respiratory virus season was used to perform phylogenetic analyses using both near-complete genome and glycoprotein (G) sequences. Overall, the RSV phylogenetic tree built with whole genomes showed identical lineage clusters as compared to the G gene but was more discriminatory. Moreover, the availability of complete genomes enables the identification of a broader range of mutations. For instance, mutations identified in the fusion protein among Canadian isolates tested here, including S377N, K272M, S276N, S211N, S206I, and S209Q, could affect the efficacy of current vaccines or antiviral-based therapeutics. In conclusion, our work reinforces other recent studies demonstrating the utility of multiplex tiling PCR assays to facilitate high-throughput WGS of RSV, which is capable of supporting enhanced genomic surveillance initiatives, as well as the more comprehensive genomic analyses required to inform public health strategies for the development and usage of vaccines and antiviral drugs. IMPORTANCE: We present assays to efficiently sequence genomes of RSVA and RSVB. This enables researchers and public health agencies to acquire high-quality genomic data using rapid and cost-effective approaches. Genomic data-based comparative analysis can be used to conduct surveillance and monitor circulating isolates for efficacy of vaccines and antiviral therapeutics.

Humans↗

Genomics-enabled dissection of sea wheatgrass genome for advancing wheat genetic resources.

Wheat production is challenged by biotic and abiotic stresses. Alien gene transfer is an effective approach to tackle such challenges. We previously showed that sea wheatgrass (SWG; Thinopyrum junceiforme (2n = 2x = 28; J1J2) is an untapped resource possessing resistance to an array of pests and abiotic stress. However, the transfer of these important traits has been hindered by the lack of genomic resources and a clear picture of its genome constitution. Using multi-color genomic in situ hybridization, we distinguished the SWG sub-genomes and corroborated that the J1 sub-genome is closely related to the E genome of Th. elongatum and the J genome of Th. bessarabicum and the J2 sub-genome to the V genome of Dasypyrum villosum. Meanwhile, we developed a draft SWG genome assembly and 127 SWG-specific DNA markers covering the 14 SWG chromosomes. Screening a population of 466 BC2F1 and BC2F2 individuals, derived from backcrosses of wheat-SWG amphiploid to wheat, by the SWG-specific markers led to selection of 72 plants putatively carrying one or two SWG chromosomes. The genome painting analysis of the 72 plants eventually identified a set of 37 wheat-SWG chromosome addition lines covering all the 14 pairs of SWG chromosomes and two compensating Robertsonian translocations (RobTs). While the wheat-SWG chromosome addition lines and RobTs are invaluable genetic resources for wheat improvement via chromosome engineering, our results showed the power of genome-specific markers in combination with genome painting in dissection of a polyploid genome and implicated the origin of a group of important polyploid grasses.

Triticum↗

Measuring conservation of contiguous sets of autosomal markers on bovine and porcine genomes in relation to the map of the human genome.

Based on published information, we have identified 991 genes and gene-family clusters for cattle and 764 for pigs that have orthologues in the human genome. The relative linear locations of these genes on human sequence maps were used as "rulers" to annotate bovine and porcine genomes based on a CSAM (contiguous sets of autosomal markers) approach. A CSAM is an uninterrupted set of markers in one genome (primary genome; the human genome in this study) that is syntenic in the other genome (secondary genome; the bovine and porcine genomes in this study). The analysis revealed 81 conserved syntenies and 161 CSAMs between human and bovine autosomes and 50 conserved syntenies and 95 CSAMs between human and porcine autosomes. Using the human sequence map as a reference, these 991 and 764 markers could correlate 72 and 74% of the human genome with the bovine and porcine genomes, respectively. Based on the number of contiguous markers in each CSAM, we classified these CSAMs into five size groups as follows: singletons (one marker only), small (2-4 markers), medium (5-10 markers), large (11-20 markers), and very large (> 20 markers). Several bovine and porcine chromosomes appear to be represented as di-CSAM repeats in a tandem or dispersed way on human chromosomes. The number of potential CSAMs for which no markers are currently available were estimated to be 63 between human and bovine genomes and 18 between human and porcine genomes. These results provide basic guidelines for further gene and QTL mapping of the bovine and porcine genomes, as well as insight into the evolution of mammalian genomes.

Animals↗

Genome SEGE: a database for 'intronless' genes in eukaryotic genomes.

BACKGROUND: A number of completely sequenced eukaryotic genome data are available in the public domain. Eukaryotic genes are either 'intron containing' or 'intronless'. Eukaryotic 'intronless' genes are interesting datasets for comparative genomics and evolutionary studies. The SEGE database containing a collection of eukaryotic single exon genes is available. However, SEGE is derived using GenBank. The redundant, incomplete and heterogeneous qualities of GenBank data are a bottleneck for biological investigation in comparative genomics and evolutionary studies. Such studies often require representative gene sets from each genome and this is possible only by deriving specific datasets from completely sequenced genome data. Thus Genome SEGE, a database for 'intronless' genes in completely sequenced eukaryotic genomes, has been constructed. AVAILABILITY: http://sege.ntu.edu.sg/wester/intronless DESCRIPTION: Eukaryotic 'intronless' genes are extracted from nine completely sequenced genomes (four of which are unicellular and five of which are multi-cellular). The complete dataset is available for download. Data subsets are also available for 'intronless' pseudo-genes. The database provides information on the distribution of 'intronless' genes in different genomes together with their length distributions in each genome. Additionally, the search tool provides pre-computed PROSITE motifs for each sequence in the database with appropriate hyperlinks to InterPro. A search facility is also available through the web server. CONCLUSIONS: The unique features that distinguish Genome SEGE from SEGE is the service providing representative 'intronless' datasets for completely sequenced genomes. 'Intronless' gene sets available in this database will be of use for subsequent bio-computational analysis in comparative genomics and evolutionary studies. Such analysis may help to revisit the original genome data for re-examination and re-annotation.

Databases, Genetic↗

Genome-tools: a flexible package for genome sequence analysis.

Genome-tools is a Perl module, a set of programs, and a user interface that facilitates access to genome sequence information. The package is flexible, extensible, and designed to be accessible and useful to both nonprogrammers and programmers. Any relatively well-annotated genome available with standard GenBank genome files may be used with genome-tools. A simple Web-based front end permits searching any available genome with an intuitive interface. Flexible design choices also make it simple to handle revised versions of genome annotation files as they change. In addition, programmers can develop cross-genomic tools and analyses with minimal additional overhead by combining genome-tools modules with newly written modules. Genome-tools runs on any computer platform for which Perl is available, including Unix, Microsoft Windows, and Mac OS. By simplifying the access to large amounts of genomic data, genome-tools may be especially useful for molecular biologists looking at newly sequenced genomes, for which few informatics tools are available. The genome-tools Web interface is accessible at http://genome-tools.sourceforge.net, and the source code is available at http://sourceforge.net/projects/genome-tools.

Base Sequence↗

The Rise of Plant Pan-Genomes: From Genome Variation to Predictive Breeding.

Plant pan-genomics is entering a new phase beyond genome variation discovery, requiring a shift from cataloguing genomic diversity toward understanding how variation generates biological function and breeding value. Here, we propose that the future of plant pan-genomics will be shaped by three conceptual transitions. First, structural variation (SV), presence-absence variation (PAV), and haplotype diversity should be interpreted not merely as genomic differences, but as regulatory components that influence gene networks, chromatin organization, and complex traits. Second, the expansion from species-level pan-genomes to genus-level super pan-genomes provides an evolutionary framework for uncovering adaptive genetic modules preserved in wild relatives and overlooked during domestication. Third, integrating pan-genomes with pan-omics, three-dimensional genome analyses, and artificial intelligence will enable the transformation of genomic variation into predictive models for crop improvement. We further propose that the ultimate value of pan-genomes lies not in generating increasingly complete genome collections, but in establishing a mechanistic bridge between genome diversity, biological function, and breeding decisions. This transition will move crop improvement from empirical selection toward rational genome design, where evolutionary diversity can be systematically interpreted, predicted, and engineered.

Journal Article↗

The filarial genome project: analysis of the nuclear, mitochondrial and endosymbiont genomes of Brugia malayi.

The Filarial Genome Project (FGP) was initiated in 1994 under the auspices of the World Health Organisation. Brugia malayi was chosen as the model organism due to the availability of all life cycle stages for the construction of cDNA libraries. To date, over 20000 cDNA clones have been partially sequenced and submitted to the EST database (dbEST). These ESTs define approximately 7000 new Brugia genes. Analysis of the EST dataset provides useful information on the expression pattern of the most abundantly expressed Brugia genes. Some highly expressed genes have been identified that are expressed in all stages of the parasite's life cycle, while other highly expressed genes appear to be stage-specific. To elucidate the structure of the Brugia genome and to provide a basis for comparison to the Caenorhabditis elegans genome, the FGP is also constructing a physical map of the Brugia chromosomes and is sequencing genomic BAC clones. In addition to the nuclear genome, B. malayi possesses two other genomes: the mitochondrial genome and the genome of a bacterial endosymbiont. Eighty percent of the mitochondrial genome of B. malayi has been sequenced and is being compared to mitochondrial sequences of other nematodes. The bacterial endosymbiont genome found in B. malayi is closely related to the Wolbachia group of rickettsia-like bacteria that infects many insect species. A set of overlapping BAC clones is being assembled to cover the entire bacterial genome. Currently, half of the bacterial genome has been assembled into four contigs. A consortium has been established to sequence the entire genome of the Brugia endosymbiont. The sequence and mapping data provided by the FGP is being utilised by the nematode research community to develop a better understanding of the biology of filarial parasites and to identify new vaccine candidates and drug targets to aid the elimination of human filariasis.

Animals↗

Rapid genome evolution revealed by comparative sequence analysis of orthologous regions from four triticeae genomes.

Bread wheat (Triticum aestivum) is an allohexaploid species, consisting of three subgenomes (A, B, and D). To study the molecular evolution of these closely related genomes, we compared the sequence of a 307-kb physical contig covering the high molecular weight (HMW)-glutenin locus from the A genome of durum wheat (Triticum turgidum, AABB) with the orthologous regions from the B genome of the same wheat and the D genome of the diploid wheat Aegilops tauschii (Anderson et al., 2003; Kong et al., 2004). Although gene colinearity appears to be retained, four out of six genes including the two paralogous HMW-glutenin genes are disrupted in the orthologous region of the A genome. Mechanisms involved in gene disruption in the A genome include retroelement insertions, sequence deletions, and mutations causing in-frame stop codons in the coding sequences. Comparative sequence analysis also revealed that sequences in the colinear intergenic regions of these different genomes were generally not conserved. The rapid genome evolution in these regions is attributable mainly to the large number of retrotransposon insertions that occurred after the divergence of the three wheat genomes. Our comparative studies indicate that the B genome diverged prior to the separation of the A and D genomes. Furthermore, sequence comparison of two distinct types of allelic variations at the HMW-glutenin loci in the A genomes of different hexaploid wheat cultivars with the A genome locus of durum wheat indicates that hexaploid wheat may have more than one tetraploid ancestor.

Amino Acid Sequence↗

Alphavirus RNA genome repair and evolution: molecular characterization of infectious sindbis virus isolates lacking a known conserved motif at the 3' end of the genome.

The 3' nontranslated region of the genomes of Sindbis virus (SIN) and other alphaviruses carries several repeat sequence elements (RSEs) as well as a 19-nucleotide (nt) conserved sequence element (3'CSE). The 3'CSE and the adjoining poly(A) tail of the SIN genome are thought to act as viral promoters for negative-sense RNA synthesis and genome replication. Eight different SIN isolates that carry altered 3'CSEs were studied in detail to evaluate the role of the 3'CSE in genome replication. The salient findings of this study as it applies to SIN infection of BHK cells are as follows: i) the classical 19-nt 3'CSE of the SIN genome is not essential for genome replication, long-term stability, or packaging; ii) compensatory amino acid or nucleotide changes within the SIN genomes are not required to counteract base changes in the 3' terminal motifs of the SIN genome; iii) the 5' 1-kb regions of all SIN genomes, regardless of the differences in 3' terminal motifs, do not undergo any base changes even after 18 passages; iv) although extensive addition of AU-rich motifs occurs in the SIN genomes carrying defective 3'CSE, these are not essential for genome viability or function; and v) the newly added AU-rich motifs are composed predominantly of RSEs. These findings are consistent with the idea that the 3' terminal AU-rich motifs of the SIN genomes do not bind directly to the viral polymerase and that cellular proteins with broad AU-rich binding specificity may mediate this interaction. In addition to the classical 3'CSE, other RNA motifs located elsewhere in the SIN genome must play a major role in template selection by the SIN RNA polymerase.

3' Untranslated Regions↗

The sources of variation in the human genome and genome instability in human cancers.

The human genome is viewed as a stable collection of about 60,000-70,000 genes--a minority of protein--coding DNA sequences--dispersed in a large majority of noncoding DNA sequences--more than 90 per cent of the entire genome sequences. Some of these ubiquitous noncoding DNA sequences, metonymically called "parasitic DNA," "ballast DNA," "selfish DNA" or "extra DNA," especially, the repeated sequences tandemly organized, are not stable but vary with considerable frequency. Recently, the confused or inadequately known origin of native of pathological variations of these DNA sequences appears to be unravelled, with great implications in genome stability. The human chromosomes, the bearer of genome, store and carry it. Their structure is qualified to perform its fastidious functions. The chromosomal conformation, "with variable geometry," exposed to genetoxic action of different damaging factors and to torsional stress after their fast and repeated changes during mitosis. The exaggerate exceeding of the native variation of human genome in disease states, probably, generates genome instability. The chromosome fragility--the cellular phenotypic expression of these molecular instability--reflects the closely relations between the genome and its carrier. The pattern of DNA replication with asynchrony of different domains of "parcelled" genome and the results of replication, susceptible to be corrected by the action of DNA repair genes, render certain limited regions of genome more vulnerable to damaging. These "target" regions focused damaging effects and exhibit an increased susceptibility to breakage and recombination, often with chromosomal expression. The coincidence of these regions, frequently, with locations of many protooncogenes and sometimes, antioncogenes could be subsequently, starting points for a genuine chain of genomic events related to growth cell and cell division. Cancer multistage accumulation of various genomic disorders in a single cell tends to take advantage of discriminating situations of these regions, which themselves can generate other genetic disorders, involving its in carcinogenesis. The gene expression disorders or the genuine mutations of dominant protooncogenes and the recessive behaviour of antioncogenes explain the nature of human cancers--a mixture of inherited and somatically acquired gene disorders. They attest the recessive characteristic of human cell malignancy and emphasize the decisive role of cancer predisposition which operates in interaction with damaging environmental factors. Seemingly, the pivotal causes of genome instability originate from strange behaviour of certain repeated DNA sequences dispersed throughout the human genome. Perhaps they hold the key to the puzzle of cancer processes.

Chromosome Aberrations↗

A complete diploid human genome benchmark for personalized genomics.

Human genome resequencing typically involves mapping reads to a reference genome to call variants; however, this approach suffers from both technical and reference biases, leaving many duplicated and structurally polymorphic regions of the genome unmapped. Consequently, existing variant benchmarks, generated by the same methods, fail to assess these complex regions. To address this limitation, we present a telomere-to-telomere genome benchmark that achieves near-perfect accuracy (i.e. no detectable errors) across 99.4% of the complete, diploid HG002 genome. This benchmark adds 701.4 Mb of autosomal sequence and both sex chromosomes (216.8 Mb), totaling 15.3% of the genome that was absent from prior benchmarks. We also provide a diploid annotation of genes, transposable elements, segmental duplications, and satellite repeats, including 39,144 protein-coding genes across both haplotypes. To facilitate application of the benchmark, we developed tools for measuring the accuracy of sequencing reads, phased variant call sets, and genome assemblies against a diploid reference. Genome-wide analyses show that state-of-the-art de novo assembly methods resolve 2-7% more sequence and outperform variant calling accuracy by an order of magnitude, yielding just one error per 100 kb across 99.9% of the benchmark regions. Adoption of genome-based benchmarking is expected to accelerate the development of cost-effective methods for complete genome sequencing, expanding the reach of genomic medicine to the entire genome and enabling a new era of personalized genomics.

Journal Article↗