PubMed HealthSearch

SEARCH · PubMed Health

Results for “noncoding genome”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Towards efficient perturbation for the noncoding genome.

Deciphering the functionality of the noncoding genome, which includes important cis-regulatory elements (CREs) and transcribed noncoding RNA genes, remains technically challenging. Here, using massively parallel genetic screening, we systematically benchmark the performance of five representative loss-of-function perturbation tools, including single-guide RNA (gRNA) mediated SpCas9 cleavage or CRISPR interference, and paired gRNA (pgRNA) involved dual-SpCas9, Big Papi (paired SpCas9 and SaCas9) or dual-enAsCas12a fragment deletion methods, in decoding the roles of the noncoding genome. For targeting CREs such as enhancers, dual-SpCas9 outperforms other methods with superior efficiency in destroying functional genomic regions. For perturbing noncoding RNA genes, in addition to dual-SpCas9, other RNA-targeting methods such as RNA interference are recommended to discriminate transcript-dependent or -independent roles. A deep learning model, DeepDC, with an associated web server, is built to facilitate optimal dual-SpCas9 pgRNA design for efficiently deleting a genomic fragment. Together, our work provides practical guidance on selecting appropriate loss-of-function tools to resolve the functional complexity of the noncoding genome.

CRISPR-Cas Systems

Pervasive cryptic selection in the human noncoding genome.

The prevailing dogma in evolutionary genetics holds that mutations within sequences that are conserved across a phylogeny are deleterious in those species, and mutations outside are neutrally evolving. Indeed, such comparative genomic approaches have estimated that mutations in approximately 5% of the human genome experience negative selection. However, sites that have biological function in certain lineages but not in others, i.e. functional turnover, may violate this assumption since these sites may be invisible to comparative genomic approaches. Thus, the extent of such cryptic, or hidden, negative selection remains elusive. Here, we developed a statistical test to detect cryptic selection in human polymorphism data. Applying our approach to simulated data shows that cryptic selection shapes the site frequency spectrum (SFS) and the statistical detection power depends on the proportion of mutations experiencing cryptic selection, the amount of sequence tested, and the sample size. We applied our method to polymorphism data from the 1000 Genomes Project, comparing variants in putatively functional noncoding regions to those in putatively neutral regions. We detected pervasive signals of cryptic selection in putatively functional regions, even after filtering out the top 70% of conserved sites. Using simulations with varying levels of cryptic selection, we estimated the extent of genome-wide constraint in the human genome. Our approximation suggests that mutations in at least 7% of the human genome are under negative selection, which is greater than the estimates from conservation-based methods, and that many of these mutations have escaped detection by comparative genomic methods. In sum, our results highlight the evolutionary dynamic nature of the noncoding genome and suggest the need to account for functional turnover when identifying putatively neutral variants for evolutionary analyses.

Journal Article

Popcorn: prediction of short coding and noncoding genomic sequences in prokaryotes.

SUMMARY: The most challenging prokaryotic genes to identify often correspond to short ORFs (sORFs) encoding small proteins or to noncoding RNAs. RNA-seq experiments commonly evince small transcripts that do not correspond to annotated genes and are candidates for novel coding sORFs or small regulatory RNAs, but it can be difficult to accurately assess whether the numerous small transcripts are coding or not. We present Popcorn (PrOkaryotic Prediction of Coding OR Noncoding), a novel machine learning method for determining whether prokaryotic sequences are coding or noncoding. We find that Popcorn is effective in distinguishing coding from noncoding sequences, including coding sORFs and noncoding RNAs. AVAILABILITY AND IMPLEMENTATION: Freely available for use on the web at https://cs.wellesley.edu/∼btjaden/Popcorn. Source code available at https://github.com/btjaden/Popcorn and https://doi.org/10.5281/zenodo.15120075.

Open Reading Frames

Mechanisms underlying disease-causing variants in promoters and enhancers.

The study of human monogenic disorders has been a powerful tool for generating a deep understanding of protein function/dysfunction and for uncovering underlying biological mechanisms. Here we explore the insights that an expanding catalog of noncoding monogenic disease variants can provide into the functions of the noncoding genome. We focus on small genetic alterations (one to a few tens of base pairs) in cis-regulatory elements-promoters, enhancers and silencers-and their potential mechanisms of action, such as loss or gain of function. We discuss the challenges in determining pathogenicity for variants in the noncoding genome, discuss why there might be so few concrete examples and highlight the opportunities for advancing this area of human genetics by using experimental and machine-learning tools.

Journal Article

Alu Overexpression Leads to an Increased Double-Stranded RNA Signature in Dermatomyositis.

OBJECTIVE: Dermatomyositis is an autoimmune condition characterized by a high interferon signature of unknown etiology. Because coding sequences constitute <1.2% of our genomes, there is a need to explore the role of the noncoding genome in disease pathogenesis. Our genomes include roughly 1.2 million Alu elements occupying approximately 10% of the genome, which can form double-stranded (ds) RNA capable of triggering MDA5 leading to interferon production. METHODS: We aligned muscle biopsy RNA sequencing data to the telomere-to-telomere reference genome and quantified short interspersed elements including Alus. Because Alus have a propensity to form dsRNA and are the major targets of both adenosine deaminase RNA specific and MDA5, we quantified adenosine to inosine (A-to-I) RNA editing, which reflects dsRNA in vivo. RESULTS: Dermatomyositis muscle (n = 39) showed a global elevation in Alu expression (including inverted-repeat Alus with high potential to form dsRNA) as well as an increased expression of unique Alu elements (n = 557, q < 0.05) compared with healthy controls (n = 34), in a pattern not seen in other myositis types (n = 81). Most (75.3%) of these Alus originated from genomic regions outside genes. A cluster of the uniquely overexpressed Alus (n = 167) correlated with interferon-stimulated genes and markers of myositis activity. Additionally, we found a uniquely expanded Alu A-to-I editome in dermatomyositis, reflecting an increase in dsRNA. Edited Alus clustered on chromosome 19, which is known to have the highest concentration of dsRNA. CONCLUSION: We hypothesize that overexpressed Alus in dermatomyositis form endogenous dsRNA that exceeds the capacity of RNA editing enzymes and triggers dsRNA sensors leading to interferon production.

Humans

Uniform processing and analysis of IGVF massively parallel reporter assay data with MPRAsnakeflow.

As researchers and clinicians seek to identify human genomic alterations relevant to traits and disorders, identifying and aggregating evidence providing mechanistic support for associations between alterations and phenotypes remains challenging. In particular, the study of noncoding genomic variation remains a major challenge because of the lack of accurate functional annotation for activity in a given context and across alleles. Experimental evidence is critical for prioritizing and interpreting functional effects of genetic alterations. Massively parallel reporter assays (MPRAs) have emerged as a powerful high-throughput approach, enabling quantification of regulatory element activity and allelic effects, as well as systematic dissection of gene regulatory logic and variant effects across different contexts. However, the diversity of MPRA designs, lack of standardized formats, and many potential processing parameters hamper data integration, reproducibility, and meta-analyses across studies. To address these challenges, the Impact of Genomic Variation on Function (IGVF) Consortium established an MPRA focus group to develop community standards, including harmonized file formats, and robust analysis pipelines for a wide range of library types and experimental designs. Here, we present these formats and comprehensive computational tools, MPRAlib and MPRAsnakeflow, for uniform processing from raw sequencing reads to counts, processing, and visualization. Using diverse MPRA data sets, we investigated technical variability sources including barcode sequence bias, outlier barcodes, and delivery method (episomal vs. lentiviral). Our results establish best practices for MPRA data generation and analysis, facilitating robust, reproducible research and large-scale integration. The presented tools and standards are publicly available, providing a foundation for future collaborative efforts in regulatory genomics.

Humans

Analysis of multiple restriction fragment length polymorphisms of the gene for the human complement receptor type I. Duplication of genomic sequences occurs in association with a high molecular mass receptor allotype.

Human CR1 exhibits an unusual form of polymorphism in which allotypic variants differ in the molecular weight of their respective polypeptide chains. To address mechanisms involved in the generation of the CR1 allotypes, DNA from individuals having the F allotype (250,000 Mr), the S allotype (290,000 Mr), and the F' allotype (210,000 Mr) was digested by restriction enzymes, and Southern blots were hybridized with CR1 cDNA and genomic probes. With the use of Bam HI and Sac I, an additional restriction fragment was observed in 20 of 21 individuals having the S allotype with no associated loss of other restriction fragments. Southern blot analysis with a noncoding genomic probe derived from the S allotype-specific Bam HI fragment showed hybridization to this fragment and to two other fragments that were also present in FF individuals. Thus, an intervening sequence may be repeated twice in the F allele and three times in the S allele. A restriction fragment length polymorphism (RFLP) unique to two individuals expressing the F' allotype was seen with Eco RV, but the absence of persons homozygous for this rare allotype prevented further comparisons with the F and S allotypes. Analysis of the CR1 transcripts associated with the three CR1 allotypes indicated that these differed by 1.3-1.5 kb and had the same rank order as the corresponding allotypes. Taken together, these findings suggest that the S allele was generated from the F allele by the acquisition of additional sequences, the coding portion of which may correspond to a long homologous repeat of approximately 1.4 kb that has been identified in CR1 cDNA. We saw two other RFLPs with Hind III and Pvu II that were in linkage dysequilibrium with the Bam HI-Sac I RFLPs associated with the S allotype, and a third polymorphism was seen with Eco RI that was not in linkage dysequilibrium with the other polymorphisms. Thus, 10 commonly occurring CR1 alleles can be defined, making this locus a useful marker for the long arm of chromosome 1 to which the CR1 gene maps.

Alleles

Functional Annotation of the Major Histocompatibility Complex Locus.

The human major histocompatibility complex (MHC) locus has the greatest density of disease-associations in the human genome, including links to over 100 polygenic disorders. Its complex haplotype structure, rich gene density, and high degree of linkage disequilibrium combine to make deciphering the gene regulatory logic of the MHC locus extremely challenging. Employing complementary high-throughput CRISPR interference (CRISPRi) and activation (CRISPRa) epigenetic screens coupled with single-cell transcriptome profiling across three distinct human cell types, we identified hundreds of new connections between cis -regulatory elements (CREs) and their target genes in this locus. These CRE-gene links are largely cell type-specific and act as enhancers. Additionally, some CREs have complex features, including harboring both active and repressive histone marks, lacking chromatin accessibility, targeting multiple genes, or acting as silencers. Computational methods fail to predict a majority of these CRE-gene connections. These findings emphasize the potential for functional perturbation experiments to dissect complex loci and reveal shared and cell type-specific regulatory mechanisms relevant to genomics of complex diseases. Collectively, this study provides a unique resource for understanding the complex regulatory landscape within the MHC locus and supports the need for creating new models that encompass CRE-gene interactions, cell type-specific gene expression, and disease genetics in the noncoding genome.

Journal Article

The deletion of 41 proximal nucleotides reverts a poliovirus mutant containing a temperature-sensitive lesion in the 5' noncoding region of genomic RNA.

We generated a number of small deletions and insertions in the 5' noncoding region of an infectious cDNA copy of the poliovirus RNA genome. Transfection of these mutated cDNAs into COS-1 cells produced the following phenotypic categories: (i) wild-type mutations, (ii) lethal mutations, (iii) mutations exhibiting slow growth or low-titer properties, and (iv) temperature-sensitive (ts) mutations. The deletion of nucleotides 221 to 224 produced a ts virus, 220D1. Mutant 220D1 was found to have a dramatic reduction in growth, virus-specific protein and RNA synthesis, and the shutoff of host cell protein synthesis at 37 or 39 degrees C compared with 33 degrees C. Temperature shift experiments showed that the mutant viral RNA is not an effective template for protein or RNA synthesis at 39 degrees C and suggested a decreased stability of the 220D1 RNA at 39 degrees C. Selection for a non-ts revertant of 220D1 yielded the virus R2, which was no longer ts for growth or viral protein and RNA synthesis. Sequencing the 5' noncoding region of the genomic RNA from R2 revealed the deletion of 41 proximal nucleotides for an overall deletion of nucleotides 184 to 228. These data suggest that the deleted sequences are nonessential to the poliovirus life cycle during growth in HeLa cells. According to computer-predicted RNA secondary structures of the 5' noncoding region of poliovirus RNA, the R2 revertant virus has deleted an entire predicted stem-loop structure.

Animals

A poliovirus temperature-sensitive RNA synthesis mutant located in a noncoding region of the genome.

We have constructed an 8-base-pair insertion mutation in the 3' noncoding region of an infectious poliovirus cDNA clone that gives rise to a temperature-sensitive RNA synthesis mutant upon transfection into mammalian cells. The mutated cDNA was used to establish a cell line that releases the mutant poliovirus in a temperature-dependent fashion, representing a unique persistent viral infection. A poliovirus mutant mapping in the noncapsid region of the viral genome can be complemented in this cell line, implying that the cell line expresses viral proteins at the nonpermissive temperature.

Base Sequence

Restricted variability of a 17 nucleotide stretch within the 5'-noncoding region of poliovirus genome.

The outbreak of poliomyelitis in Finland in 1984 was caused by a wild strain of poliovirus 3 with uncommon molecular and antigenic properties. We prepared a synthetic oligonucleotide probe complementary to nucleotides 494-510 in the 5'-noncoding part of the genome of a representative strain of the outbreak. This short nucleotide stretch was found to be relatively well conserved within the outbreak and uncommon among 82 independent poliovirus isolates. It may thus be a useful marker for screening isolates to identify those requiring more detailed genetic comparison. The sequences of the corresponding region of the genome are known for 32 separate poliovirus strains and 3 coxsackie B virus strains and show 6 fully conserved nucleotides that could assume a constant hairpin-loop position in a hypothetical secondary structure of the RNA. This could explain the persistence of a particular 17 nucleotide sequence for 40 years in nature in this highly variable region of the poliovirus genome.

Animals

Construction of less neurovirulent polioviruses by introducing deletions into the 5' noncoding sequence of the genome.

Viral attenuation may be due to lowered efficiency of certain steps essential for viral multiplication. For the construction of less neurovirulent strains of poliovirus in vitro, we introduced deletions into the 5' noncoding sequence (742 nucleotides long) of the genomes of the Mahoney and Sabin 1 strains of poliovirus type 1 by using infectious cDNA clones of the virus strains. Plaque sizes shown by deletion mutants were used as a marker for rate of viral proliferation. Deletion mutants of both the strains thus constructed lacked a genome region of nucleotide positions 564 to 726. The sizes of plaques displayed by these deletion mutants were smaller than those by the respective parental viruses, although a phenotype referring to reproductive capacity at different temperatures (rct) of viruses was not affected by introduction of the deletion. Monkey neurovirulence tests were performed on the deletion mutants. The results clearly indicated that the deletion mutants had much less neurovirulence than with the corresponding parent viruses. Production of infectious particles and virus-specific protein synthesis in cells infected with the deletion mutants started later than in those infected with the parental viruses. The rate at which cytopathic effect progressed was also slower in cells infected with the mutants. Phenotypic stability of the deletion mutant for small-plaque phenotype and temperature sensitivity was investigated after passaging the mutant at an elevated temperature of 37.5 degrees C. Our data strongly suggested that the less neurovirulent phenotype introduced by the deletion is very stable during passaging of the virus.

Animals

Intracellular modifications induced by poliovirus reduce the requirement for structural motifs in the 5' noncoding region of the genome involved in internal initiation of protein synthesis.

A series of genetic deletions based partly on two RNA secondary structure models (M. A. Skinner, V. R. Racaniello, G. Dunn, J. Cooper, P. D. Minor, and J. W. Almond, J. Mol. Biol. 207:379-392, 1989; E. V. Pilipenko, V. M. Blinov, L. I. Romanova, A. N. Sinyakov, S. V. Maslova, and V. I. Agol, Virology 168:201-209, 1989) was made in the cDNA encoding the 5' noncoding region (5' NCR) of the poliovirus genome in order to study the sequences that direct the internal entry of ribosomes. The modified cDNAs were placed between two open reading frames in a single transcriptional unit and used to transfect cells in culture. Internal entry of ribosomes was detected by measuring translation from the second open reading frame in the bicistronic mRNA. When assayed alone, a large proportion of the poliovirus 5' NCR superstructure including several well-defined stem-loops was required for ribosome entry and efficient translation. However, in cells cotransfected with a complete infectious poliovirus cDNA, the requirement for the stem-loops in this large superstructure was reduced. The results suggest that virus infection modifies the cellular translational machinery, so that shortened forms of the 5' NCR are sufficient for cap-independent translation, and that the internal entry of ribosomes occurs by two distinct modes during the virus replication cycle.

DNA Mutational Analysis

Oxidation-reduction sensitive interaction of a cellular 50-kDa protein with an RNA hairpin in the 5' noncoding region of the poliovirus genome.

Genetic and biochemical analyses of the 5' noncoding region of poliovirus have indicated the importance of this region in both translation and amplification of the viral RNA. The role of the cellular machinery required for these events is just beginning to be revealed. Using an RNA gel retention assay, we have identified a cellular 50-kDa protein that forms a specific complex with a stable stem-loop structure present in the viral 5' noncoding region. The formation of the RNA-protein complex is dependent on the availability of free sulfhydryl groups in the protein. The possible involvement of this RNA-protein complex in the regulation of viral gene expression is discussed.

Base Sequence

Origin of noncoding DNA sequences: molecular fossils of genome evolution.

The total amount of noncoding sequences on chromosomes of contemporary organisms varies significantly from species to species. We propose a hypothesis for the origin of these noncoding sequences that assumes that (i) an approximately equal to 0.55-kilobase (kb)-long reading frame composed the primordial gene and (ii) a 20-kb-long single-stranded polynucleotide is the longest molecule (as a genome) that was polymerized at random and without a specific template in the primordial soup/cell. The statistical distribution of stop codons allows examination of the probability of generating reading frames of approximately equal to 0.55 kb in this primordial polynucleotide. This analysis reveals that with three stop codons, a run of at least 0.55-kb equivalent length of nonstop codons would occur in 4.6% of 20-kb-long polynucleotide molecules. We attempt to estimate the total amount of noncoding sequences that would be present on the chromosomes of contemporary species assuming that present-day chromosomes retain the prototype primordial genome structure. Theoretical estimates thus obtained for most eukaryotes do not differ significantly from those reported for these specific organisms, with only a few exceptions. Furthermore, analysis of possible stop-codon distributions suggests that life on earth would not exist, at least in its present form, had two or four stop codons been selected early in evolution.

Animals

Presence of poly(A) in a flavivirus: significant differences between the 3' noncoding regions of the genomic RNAs of tick-borne encephalitis virus strains.

A poly(A) tail was identified on the 3' end of the prototype tick-borne encephalitis (TBE) virus strain Neudoerfl. This is in contrast to the general lack of poly(A) in the genomic RNAs of mosquito-borne flaviviruses analyzed so far. Analysis of several closely related strains of TBE virus, however, revealed the existence of two different types of 3' noncoding (NC) regions. One type (represented by strain Neudoerfl) is only 114 nucleotides long and carries a 3'-terminal poly(A) structure. This was also found in several TBE virus strains isolated from different geographic regions over a period of almost 30 years. The other type (represented by strain Hypr) is 461 nucleotides long and not polyadenylated. The sequence homology between the two types of TBE virus 3' NC regions terminates at a specific position 81 nucleotides after the stop codon. The second type of 3' NC region more closely resembles the common flavivirus pattern, including the potential for the formation of a 3'-terminal hairpin structure. However, it lacks primary sequence elements that are conserved among other flavivirus genomes.

Animals

Base mutations in the terminal noncoding regions of the genome of vesicular stomatitis virus isolated from persistent infections of L cells.

The 3'-terminal regions of the genomes of vesicular stomatitis virus obtained from two long-term, independently initiated persistent infections of L cells were found to contain several sequence mutations. In contrast to the hypermutability displayed in the 5'-terminal regions of the genomes of viruses obtained from persistent infections of baby hamster kidney (BHK) cells (P. J. O'Hara, F. M. Horodyski, S. T. Nichol, and J. J. Holland, J. Virol. 49, 793-798, 1984), no 5' mutations were detected in viruses from L-cell carrier lines. The absence of detectable defective interfering (DI) particles in the L-cell carrier cultures may account for this difference. Plus-strand leader RNA made by the viruses from persistently infected L cells failed to accumulate from 5 to 8 hr postinfection unlike the accumulation noted for the leader RNA generated by wild-type VSV. Minus-strand leader RNA, on the other hand, accumulated at a similar or increased rate compared to wild type. The relationship of these observations to the processes of host shutoff, viral transcription, and replication are discussed.

Animals

Segment-specific and common nucleotide sequences in the noncoding regions of influenza B virus genome RNAs.

The nucleotide sequences of the 3' noncoding regions of all eight segments of influenza B virus RNA and the sequences of the 5' noncoding regions of segments 4-8 were determined in virus strains isolated over a period of 40 years. Nearly complete conservation of the noncoding sequences was found. Nine nucleotides at the 3' termini and 11 nucleotides at the 5' termini were common to all segments examined. In the region immediately adjacent to the common 3' terminal region, the nucleotides were specific for each segment and these segment-specific sequences were conserved in all strains examined. In each of the five segments in which both termini were examined, the segment-specific 3' sequences exhibited perfect inverted complementarity to a segment-specific sequence adjacent to the common 5' terminus. In addition, in the 3' noncoding region of RNA segments 1-3, which encode proteins involved in RNA synthesis, a single nucleotide substitution at position 10 was found that distinguishes these segments from segments 4-8. Comparison of these data with published reports has revealed that some of the features found in the noncoding regions of influenza B virus are also present in influenza A and C virus RNAs. In the RNAs of all three virus types, there is a segment-specific sequence of nucleotides near the 3' terminus that shows inverted complementarity to a sequence near the 5' terminus. This segment-specific sequence may play a role in the transcription of individual segments or in sorting of segments during virion assembly.

Animals