PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Genome Components”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Genomic scrap yard: how genomes utilize all that junk.

Interspersed repetitive sequences are major components of eukaryotic genomes. Repetitive elements comprise over 50% of the mammalian genome. Because the specific function of these elements remains to be defined and because of their unusual 'behaviour' in the genome, they are often quoted as a selfish or junk DNA. Our view of the entire phenomenon of repetitive elements has to now be revised in the light of data on their biology and evolution, especially in the light of what we know about the retroposons. I would like to argue that even if we cannot define the specific function of these elements, we still can show that they are not useless pieces of the genomes. The repetitive elements interact with the surrounding sequences and nearby genes. They may serve as recombination hot spots or acquire specific cellular functions such as RNA transcription control or even become part of protein coding regions. Finally, they provide very efficient mechanism for genomic shuffling. As such, repetitive elements should be called genomic scrap yard rather than junk DNA. Tables listing examples of recruited (exapted) transposable elements are available at http://www.ncbi.nlm.gov/Makalowski/ScrapYard/

Animals↗

Sequences of ten circular ssDNA components associated with the milk vetch dwarf virus genome.

Milk vetch dwarf virus (MDV) is a member of the proposed genus Nanovirus, and its genome is composed of multiple, circular ssDNA components of about 1 kb. We have cloned and sequenced ten ssDNA components and designated them MDV-C1 to C10. Each DNA component contains one potential major open reading frame, and contains a putative stem-loop structure in the non-coding region. Notably, four components (C1, C2, C3 and C10) encode distinct replication-associated (Rep) proteins of 33 kDa, which show only limited (42-57%) amino acid identity. The six other components encode proteins with calculated molecular masses ranging from 12.7 to 19.7 kDa. Comparison of the sequences with those of other nanoviruses reveals that MDV is closely related to faba bean necrotic yellows virus (FBNYV) and subterranean clover stunt virus (SCSV). Six putative MDV genome products, including one Rep and five non-Rep proteins, show high (70.4-90.9%) amino acid identity to the corresponding six FBNYV proteins, whereas two other Rep proteins encoded by MDV-C2 and C3 are 82.3% and 73.0% identical to those encoded by SCSV-C2 and C6, respectively. These results indicate that MDV, FBNYV and SCSV have diverged from a common origin, which had multiple Rep components. In addition, the putative proteins encoded by MDV-C4 and its homologues contain a consensus retinoblastoma-binding motif, suggesting that they may be involved in controlling the host cell cycle.

Amino Acid Sequence↗

Linkage of two distinct AT-rich minisatellites at multiple loci in the genome of Theileria parva.

Minisatellite tandem repeat elements are well known components of vertebrate genomes, but have not yet been extensively characterized in lower eukaryotes. We describe two unusual, AT-rich minisatellites of the protozoan parasite Theileria parva whose sequences are unrelated to the G/C-rich i minisatellite superfamily' of vertebrate and plant genomes. The T. parva tandem repeats, one with a conserved sequence T2-5ACACA (6-17 copies), and the other with a 6-bp core sequence of either ACTATA or TATACT associated with additional variable sequences in repeats of 10-17bp (3-7 copies), were closely linked at more than 20 sites in the T. parva genome, separated by 390, 510 and 660bp at three loci analysed in detail. Such linkage is without precedent in minisatellites so far analysed in other organisms. The minisatellite loci were widely dispersed on 13 out of 33 genomic SfiI fragments, on all four T. parva chromosomes and did not exhibit a telomeric bias in their distribution. Analysis of flanking sequences revealed no obvious conserved sequences between the five loci, or other multicopy repeat sequences outside the minisatellite regions. The T2-5 ACACA minisatellite was highly effective as a multilocus fingerprinting probe for discrimination of T. parva isolates. Analysis of two individual minisatellite loci revealed variation between the genomic DNAs of two T. parva isolates in the copy number of the constituent repeats within the array, similar to that typical of vertebrate minisatellites. 1998 Elsevier Science B.V.

Animals↗

PRISM-G: an interpretable privacy scoring framework for assessing risk in synthetic human genome data.

MOTIVATION: Synthetic genomic data promises broader data access, but unresolved privacy risks remain a major concern. Existing evaluations often rely on similarity-based metrics that measure proximity between real and synthetic genomes, overlooking additional mechanisms through which genomic information may leak. RESULTS: We introduce PRISM-G, a model-agnostic framework that quantifies privacy exposure in synthetic genomic data across three complementary components: proximity to real genomes in genetic-coordinate space, replay of familial or population-structure patterns, and trait-linked exposure through rare variants and membership-inference signals. These components are normalized and combined through a risk-averse aggregation into a single 0-100 PRISM-G score. By pairing PRISM-G with downstream utility metrics, the framework also enables analysis of privacy-utility trade-offs across generative models. We evaluated PRISM-G on synthetic cohorts generated by a generative adversarial network (GAN), a restricted Boltzmann machine (RBM), and a logic-based SAT solver (Genomator). Our results show that privacy vulnerabilities arise along different axes across models and marker densities, demonstrating that a single similarity-based metric is insufficient to characterize genomic privacy risk. AVAILABILITY AND IMPLEMENTATION: The source code of PRISM-G is available at https://github.com/alejocrojo09/prismg.

Humans↗

A thermal denaturation study of genomic DNAs from North American minnows (Cyprinidae: Teleostei).

Base compositions and differential melting rate profiles of genomic DNAs from twenty species of North American cyprinid fishes were generated via thermal denaturation. Base pair composition expressed as % GC values ranged among the twenty species from 36.1-41.3%. This range is considerably broader than that observed at comparable taxonomic levels in other vertebrate groups. Both the range and average difference in base pair composition between species in the diverse and rapidly evolving genus Notropis were considerably greater than those between species in other North American cyprinid genera. This may indicate that genomic changes at the level of base pair composition are frequent and possibly important events in cyprinid evolution. Compositional heterogeneity and asymmetry values among the twenty species were uniform and low, respectively, suggesting that most of the species lacked DNA components in their genomes which differed substantially from their main-band DNAs in base pair composition. The melting rate profiles revealed a prominent and distinct heavy or GC-rich DNA component in the genomes of three species belonging to the subgenus Cyprinella of Notropis. These and other data suggest that the heavy melting component may reflect a large, comparatively GC-rich family of highly repeated or satellite DNA sequences common to all three genomes.

Animals↗

The terminal a sequence of the herpes simplex virus genome contains the promoter of a gene located in the repeat sequences of the L component.

The herpes simplex virus DNA genome consists of two covalently linked components, L and S. The unique sequences of the L component are flanked by 9-kilobase-pair inverted repeat sequences ab and b'a', whereas those of the S component are flanked by 6.5-kilobase-pair inverted repeat sequences c'a' and ca. We report that the 500-base-pair a sequence contains the promoter-regulatory domain and the transcription initiation site of a diploid gene, the coding sequences of which are located in the b sequences of the inverted repeats of the L component. The chimeric gene constructed by fusion of the a sequence to the coding sequences of the thymidine kinase gene and recombined into the viral genome was regulated as a gamma 1 gene. The size of the protein predicted from its sequence is 358 amino acids; it was designated as infected cell protein (ICP) 34.5. Thus, the inverted repeats flanking the unique sequences of the L component contain two genes specifying ICP0 and ICP34.5, respectively. Moreover, in addition to the cis-acting sites for the inversion of L and S components relative to each other, for cleavage of unit length DNA molecules from head-to-tail concatemers, and for packaging of the DNA into capsids, the a sequence also contains the promoter-regulatory domain and transcription initiation sites of a gene.

Amino Acid Sequence↗

Organization of nucleotide sequences in the chicken genome.

The four major components of chicken DNA were prepared by density gradient centrifugation and characterized in several basic properties: relative amounts, dG + dC content, buoyant densities, compositional heterogeneity, and reassociation kinetics. While the relative amounts and the compositions of the major components of chicken DNA were similar to those found in mammalian genomes, their compositional heterogeneities were found to be narrower. The relative amounts of interspersed repeated and unique sequences were strikingly different in different components and also different from those found in the corresponding major components of mouse and human DNAs. If one takes into consideration that major DNA components (a) account for practically all of main-band DNA and (b) derive by preparative breakage from very long DNA segments of fairly homogeneous composition, the isochores, our findings indicate that the distribution of interspersed repeats is different in different chromosomal regions and is species-specific.

Animals↗

Unique and redundant roles for HOG MAPK pathway components as revealed by whole-genome expression analysis.

The Saccharomyces cerevisiae high osmolarity glycerol (HOG) mitogen-activated protein kinase pathway is required for osmoadaptation and contains two branches that activate a mitogen-activated protein kinase (Hog1) via a mitogen-activated protein kinase kinase (Pbs2). We have characterized the roles of common pathway components (Hog1 and Pbs2) and components in the two upstream branches (Ste11, Sho1, and Ssk1) in response to elevated osmolarity by using whole-genome expression profiling. Several new features of the HOG pathway were revealed. First, Hog1 functions during gene induction and repression, cross talk inhibition, and in governing the regulatory period. Second, the phenotypes of pbs2 and hog1 mutants are identical, indicating that the sole role of Pbs2 is to activate Hog1. Third, the existence of genes whose induction is dependent on Hog1 and Pbs2 but not on Ste11 and Ssk1 suggests that there are additional inputs into Pbs2 under our inducing conditions. Fourth, the two upstream pathway branches are not redundant: the Sln1-Ssk1 branch has a much more prominent role than the Sho1-Ste11 branch for activation of Pbs2 by modest osmolarity. Finally, the general stress response pathway and both branches of the HOG pathway all function at high osmolarity. These studies demonstrate that cells respond to increased osmolarity by using different signal transduction machinery under different conditions.

Gene Expression Profiling↗

Complete genome sequence of the marine, chemolithoautotrophic, ammonia-oxidizing bacterium Nitrosococcus oceani ATCC 19707.

The gammaproteobacterium Nitrosococcus oceani (ATCC 19707) is a gram-negative obligate chemolithoautotroph capable of extracting energy and reducing power from the oxidation of ammonia to nitrite. Sequencing and annotation of the genome revealed a single circular chromosome (3,481,691 bp; G+C content of 50.4%) and a plasmid (40,420 bp) that contain 3,052 and 41 candidate protein-encoding genes, respectively. The genes encoding proteins necessary for the function of known modes of lithotrophy and autotrophy were identified. Contrary to betaproteobacterial nitrifier genomes, the N. oceani genome contained two complete rrn operons. In contrast, only one copy of the genes needed to synthesize functional ammonia monooxygenase and hydroxylamine oxidoreductase, as well as the proteins that relay the extracted electrons to a terminal electron acceptor, were identified. The N. oceani genome contained genes for 13 complete two-component systems. The genome also contained all the genes needed to reconstruct complete central pathways, the tricarboxylic acid cycle, and the Embden-Meyerhof-Parnass and pentose phosphate pathways. The N. oceani genome contains the genes required to store and utilize energy from glycogen inclusion bodies and sucrose. Polyphosphate and pyrophosphate appear to be integrated in this bacterium's energy metabolism, stress tolerance, and ability to assimilate carbon via gluconeogenesis. One set of genes for type I ribulose-1,5-bisphosphate carboxylase/oxygenase was identified, while genes necessary for methanotrophy and for carboxysome formation were not identified. The N. oceani genome contains two copies each of the genes or operons necessary to assemble functional complexes I and IV as well as ATP synthase (one H(+)-dependent F(0)F(1) type, one Na(+)-dependent V type).

Adenosine Triphosphate↗

Numerous potentially functional but non-genic conserved sequences on human chromosome 21.

The use of comparative genomics to infer genome function relies on the understanding of how different components of the genome change over evolutionary time. The aim of such comparative analysis is to identify conserved, functionally transcribed sequences such as protein-coding genes and non-coding RNA genes, and other functional sequences such as regulatory regions, as well as other genomic features. Here, we have compared the entire human chromosome 21 with syntenic regions of the mouse genome, and have identified a large number of conserved blocks of unknown function. Although previous studies have made similar observations, it is unknown whether these conserved sequences are genes or not. Here we present an extensive experimental and computational analysis of human chromosome 21 in an effort to assign function to sequences conserved between human chromosome 21 (ref. 8) and the syntenic mouse regions. Our data support the presence of a large number of potentially functional non-genic sequences, probably regulatory and structural. The integration of the properties of the conserved components of human chromosome 21 to the rapidly accumulating functional data for this chromosome will improve considerably our understanding of the role of sequence conservation in mammalian genomes.

Animals↗

Genome organization in Halobacterium halobium: a 70 kb island of more (AT) rich DNA in the chromosome.

The more A + T rich fractionated component (FII DNA) of the Halobacterium halobium genome constitutes one third of the total DNA and upon isolation consists of covalently closed circular DNA (pHH1 and minor cccDNA) and nonsupercoiled sequences. We have investigated the physical organization of the non cccDNA in FII by a chromosome walk using one copy of the halobacterial insertion element ISH1 as a start point. This chromosome walk led to the isolation of 160 kb of chromosomal DNA containing 70 kb of FII DNA covalently linked to more G + C rich sequences (FI DNA). Copies of three previously characterized insertion elements (ISH1, ISH2, and ISH26) as well as at least 10 other repeated sequences are clustered within this chromosomal FII DNA "island". Unique sequences are found in the FI DNA flanking the FII DNA island as well as in 40 kb of FI DNA surrounding the bacterio-opsin gene. The presence of pHH1 in H. halobium and closely related species correlates with the occurrence of the characterized chromosomal FII DNA island. Halophilic purple membrane producing isolates YC81819-9, GN101, SB3 and GRA lack pHH1 and the 70 kb FII DNA, but contain all of the FI DNA sequences tested. We propose that pHH1 and this chromosomal FII DNA are characteristic genomic components of H. halobium and closely related species, and, that the 70 kb FII DNA might represent a large insertion in the chromosome of H. halobium and closely related species. The conservation of both FI and FII DNA sequences can be used for strain classification and determination of evolutionary relationships among halo-bacteria.

Base Sequence↗

Analysis by comparative genomic hybridization of epithelial and spindle cell components in sarcomatoid carcinoma and carcinosarcoma: histogenetic aspects.

Sarcomatoid carcinomas and carcinosarcomas are histologically malignant biphasic neoplasms with an epithelial and a spindle cell component. Both a polyclonal and a monoclonal origin have been postulated for these tumours, but the latter has been favoured. For carcinosarcoma, the stem cell from which the epithelial and mesenchymal components are derived is expected to be more immature than the epithelial stem cell from which different components of sarcomatoid carcinoma originate, since in the latter, immunohistochemical or ultrastructural epithelial characteristics are still detectable. In the present study, comparative genomic hybridization was used to test the hypothesis that both tumour components in sarcomatoid carcinoma have more chromosomal aberrations in common than those in carcinosarcoma. From three sarcomatoid carcinomas originating from the urinary bladder and two carcinosarcomas from the pharynx, the epithelial and spindle cell components were microdissected and analysed for their respective chromosomal aberrations, using comparative genomic hybridization. High-level homology was seen in chromosomal aberrations between the different components in each tumour. This level of homology was even higher in the carcinosarcomas (65 and 91 per cent) than in both sarcomatoid carcinomas (21-51 per cent). The different phenotypic components of both sarcomatoid carcinoma and carcinosarcoma show a large overlap of chromosomal aberrations, strongly suggesting a monoclonal origin for all of these tumours. These findings do not support the hypothesis that the divergence between epithelial and spindle cell components occurs at an earlier stage in carcinosarcomas than in sarcomatoid carcinoma.

Carcinoma↗

The major components of the mouse and human genomes. 1. Preparation, basic properties and compositional heterogeneity.

Main-band DNA from mammals and birds can be resolved by density gradient centrifugation techniques into three or four families of fragments of different dG + dC contents. These major DNA components are similar in their buoyant densities and relative amounts in all species tested and are observed in DNA preparations ranging in Mr from 2 X 10(6) to over 200 X 10(6). In the present work, the four major components of mouse and human DNAs were prepared and characterized in several basic properties: relative amounts, dG + dC contents, buoyant densities and compositional heterogeneity. The results obtained lead to the following conclusions: (a) the major DNA components of mouse and man form at least 85% and possibly the totality of the main bands of these DNAs; (b) they have very low compositional heterogeneities over a wide molecular weight range; (c) they derive from very large chromosomal DNA segments of fairly homogeneous base composition, for which the name 'isochores' is proposed. A comparison of the compositional heterogeneity of main-band DNAs from warm-blooded and cold-blooded vertebrates confirms our previous conclusion that these DNAs are characterized by a different sequence organization.

Animals↗

Agptools: a utility suite for editing genome assemblies.

SUMMARY: The AGP format is a tab-separated table format describing how components of a genome assembly fit together. A standard submission format for genome assemblies is a fasta file giving the sequence of contigs along with an AGP file showing how these components are assembled into larger pieces like scaffolds or chromosomes. For this reason, many scaffolding software pipelines output assemblies in this format. However, although many programs for assembling and scaffolding genomes read and write this format, there is currently no published software for making edits to AGP files when performing assembly curation. We present agptools, a suite of command-line programs that can perform common operations on AGP files, such as breaking and joining sequences, inverting pieces of assembly components, assembling contigs into larger sequences based on an AGP file, and transforming between coordinate systems of different assembly layouts. Additionally, agptools includes an API that writers of other software packages can use to read, write, and manipulate AGP files within their own programs. AVAILABILITY AND IMPLEMENTATION: Source code and binaries freely available for download at https://github.com/WarrenLab/agptools, implemented in Python and supported on all operating systems.

Software↗

SAG/ROC2/Rbx2/Hrt2, a component of SCF E3 ubiquitin ligase: genomic structure, a splicing variant, and two family pseudogenes.

We have recently cloned and characterized an evolutionarily conserved gene, Sensitive to Apoptosis Gene (SAG), which encodes a redox-sensitive antioxidant protein that protects cells from apoptosis induced by redox agents. The SAG protein was later found to be the second family member of ROC/Rbx/Hrt, a component of the Skp1-cullin-F box protein (SCF) E3 ubiquitin ligase, being required for yeast growth and capable of promoting cell growth during serum starvation. Here, we report the genomic structure of the SAG gene that consists of four exons and three introns. We also report the characterization of a SAG splicing variant (SAG-v), that contains an additional exon (exon 2; 264 bp) not present in wildtype SAG. The inclusion of exon 2 disrupts the SAG ORF and gives rise to a protein of 108 amino acids that contains the first 59 amino acids identical to SAG and a 49-amino acid novel sequence at the C terminus. The entire RING-finger domain of SAG was not translated because of several inframe stop codons within the exon 2. The SAG-v protein was expressed in multiple human tissues as well as cell lines, but at a much lower level than wildtype SAG. Unlike SAG, SAG-v was not able to rescue yeast cells from lethality in a ySAG knockout, nor did it bind to cullin-1 or have ligase activity, probably because of the lack of the RING-finger domain. Finally, we report the identification of two SAG family pseudogenes, SAGP1 and SAGP2, that share 36% or 47% sequence identity with ROC1/Rbx1/Hrt1 and 30% or 88% with SAG, respectively. Both genes are intronless with two inframe stop codons.

Alternative Splicing↗

The Mouse Genome Database (MGD): integrating biology with the genome.

The Mouse Genome Database (MGD) is one component of the Mouse Genome Informatics (MGI) system (http://www.informatics.jax.org), a community database resource for the laboratory mouse. MGD strives to provide a comprehensive knowledgebase about the mouse with experiments and data annotated from both literature and online sources. MGD curates and presents consensus and experimental data representations of genetic, genotype (sequence) and phenotype information including highly detailed reports about genes and gene products. Primary foci of integration are through representations of relationships between genes, sequences and phenotypes. MGD collaborates with other bioinformatics groups to curate a definitive set of information about the laboratory mouse and to build and implement the data and semantic standards that are essential for comparative genome analysis. Recent developments in MGD discussed here include an extensive integration of the mouse sequence data and substantial revisions in the presentation, query and visualization of sequence data.

Animals↗

Clonal origin of metastatic testicular teratomas.

PURPOSE: Testicular teratomas in adult patients are histologically diverse tumors that frequently coexist with other germ cell tumor (GCT) components. These mixed GCTs often metastasize to retroperitoneal lymph nodes where multiple GCT elements are frequently present in the same metastatic lesion. Neither the genetic relationships among the different components in metastatic lesions nor the relationships between primary and metastatic GCT components have been elucidated. EXPERIMENTAL DESIGN: We examined metastases from 31 patients who underwent primary retroperitoneal lymph node dissection for metastatic testicular GCT. All patients had metastatic mature teratoma with one or more other GCT components. This study included a total of 72 metastatic GCT components and 16 primary GCT components from 31 patients. Genomic DNA samples from each component were prepared from formalin-fixed, paraffin-embedded tissue sections using laser-assisted microdissection. Loss of heterozygosity (LOH) assays for seven microsatellite polymorphic markers on chromosomes 1p36 (D1S1646), 9p21 (D9S171 and IFNA), 9q21 (D9S303), 13q22-q31 (D13S317), 18q22 (D18S543), and 18q21 (D18S60) were done to assess clonality. RESULTS: Twenty-nine of 31 (94%) cases showed allelic loss in one or more components of the metastatic GCTs. Twenty-nine of 31 mature teratomas showed allelic loss in at least one of seven microsatellite polymorphic markers analyzed. The frequency of allelic loss in informative cases of metastatic mature teratoma was 27% (8 of 30) with D1S1646, 34% (10 of 29) with D9S171, 37% (10 of 27) with IFNA, 27% (8 of 30) with D9S303, 46% (13 of 28) with D13S317, 26% (7 of 27) with D18S543, and 36% (10 of 28) with D18S60. Completely concordant allelic loss patterns between the mature teratoma and all of the other metastatic GCT components were seen in 26 of 29 cases in which the mature teratoma component showed LOH. Nearly identical allelic loss patterns were seen in the three remaining cases. In six cases analyzed, LOH patterns of each metastatic component were compared with each GCT component of the primary testicular tumor. In all six cases, each primary and metastatic component showed an identical pattern of allelic loss. CONCLUSION: Our data support the common clonal origin of metastatic mature teratomas with other components of metastatic testicular GCTs and with each component of the primary tumor.

Adolescent↗

The genome-wide localization of Rsc9, a component of the RSC chromatin-remodeling complex, changes in response to stress.

The cellular response to environmental changes includes widespread modifications in gene expression. Here we report the identification and characterization of Rsc9, a member of the RSC chromatin-remodeling complex in yeast. The genome-wide localization of Rsc9 indicated a relationship between genes targeted by Rsc9 and genes regulated by stress; treatment with hydrogen peroxide or rapamycin, which inhibits TOR signaling, resulted in genome-wide changes in Rsc9 occupancy. We further show that Rsc9 is involved in both repression and activation of mRNAs regulated by TOR as well as the synthesis of rRNA. Our results illustrate the response of a chromatin-remodeling factor to signaling cascades and suggest that changes in the activity of chromatin-remodeling factors are reflected in changes in their localization in the genome.

Amino Acid Sequence↗