PubMed Health⌕ Search

Biomedical subjects

Genomics

Find indexed PubMed genomics citations. Search gene expression, sequencing and genetic variation in titles, abstracts and supplied subjects, then open the PubMed record.

At least 361 records · Page 20Linked to original sources

Vertebrate genome sequencing: building a backbone for comparative genomics.

The human genome sequence provides a reference point from which we can compare ourselves with other organisms. Interspecies comparison is a powerful tool for inferring function from genomic sequence and could ultimately lead to the discovery of what makes humans unique. To date, most comparative sequencing has focused on pair-wise comparisons between human and a limited number of other vertebrates, such as mouse. Targeted approaches now exist for mapping and sequencing vertebrate bacterial artificial chromosomes (BACs) from numerous species, allowing rapid and detailed molecular and phylogenetic investigation of multi-megabase loci. Such targeted sequencing is complementary to current whole-genome sequencing projects, and would benefit greatly from the creation of BAC libraries from a diverse range of vertebrates.

Animals↗

Gene organization and sequence of the region containing the ribosomal protein genes RPL13A and RPS11 in the human genome and conserved features in the mouse genome.

We have determined the organization and sequence of the region containing two ribosomal protein (rp) genes in the human and mouse genomes. The two genes, human RPL13A and RPS11, and mouse Rpl13a and Rps11, are tandemly located in both genomes with an interval of only 4.6kb in the case of the human genes and 1.6kb in the case of the mouse genes. The human RPL13A and RPS11 are 4236bp and 3254bp in length and comprise eight and five exons respectively, whereas the mouse Rps11 is 1951bp long and has five exons. Structural comparison of these genes, including previously reported mouse Rpl13a, revealed a significant conservation of sequences in the promoter regions. Although most rp genes are dispersed throughout the human genome, the conserved features and adjacent localization indicate possible coordinate transcription of the two genes. Furthermore, we have found that four small nucleolar RNA (snoRNA) genes are located in the introns of the two rp genes, both human and mouse. U32, U33, and U34 snoRNAs are encoded in introns 2, 4, and 5 of RPL13A respectively, and U35 in the sixth intron of RPL13A and the third intron of RPS11. The same organization of these snoRNA genes was also observed in the case of the mouse genes.

Animals↗

Roles for genomic imprinting and the zygotic genome in placental development.

The placenta contains several types of feto-maternal interfaces where zygote-derived cells interact with maternal cells or maternal blood for the promotion of fetal growth and viability. The genetic factors regulating the interactions between different cell types within feto-maternal interfaces and the relative contributions of the maternal and zygotic genomes are poorly understood. Genomic imprinting, the epigenetic process responsible for parental origin-dependent functional differences between homologous chromosomes, has been proposed to contribute to these events. Previous studies showed that mouse conceptuses with an absence of imprinted differences between the two copies of chromosome 12 (upon paternal inheritance of both copies) die late in gestation and have a variety of defects, including placentomegaly. Here we examined the role of chromosome 12 imprinting in these placentae in more detail. We show that the spatial interactions between different cell types within feto-maternal interfaces are defective and identify abnormal behaviors in both zygote-derived and maternal cells that are attributed to the genome of the zygote but not the mother. These include compromised invasion of the maternal decidualized endometrium and the central maternal artery situated within it by zygote-derived trophoblast, abnormalities in the wall of the central maternal artery, and defects within the zygote-derived cellular layer of the labyrinth, which is in direct contact with maternal blood. These findings demonstrate multiple roles for chromosome 12 imprinting in the placenta that have not previously been associated with imprinting effects. They provide insights into the function of imprinting in placental development and have evolutionary and clinical implications.

Animals↗

Genomes OnLine Database (GOLD): a monitor of genome projects world-wide.

GOLD is a comprehensive resource for accessing information related to completed and ongoing genome projects world-wide. The database currently provides information on 350 genome projects, of which 48 have been completely sequenced and their analysis published. GOLD was created in 1997 and since April 2000 it has been licensed to Integrated Genomics. The database is freely available through the URL: http://igweb.integratedgenomics.com/GOLD/.

Animals↗

SUPFAM--a database of potential protein superfamily relationships derived by comparing sequence-based and structure-based families: implications for structural genomics and function annotation in genomes.

Members of a superfamily of proteins could result from divergent evolution of homologues with insignificant similarity in the amino acid sequences. A superfamily relationship is detected commonly after the three-dimensional structures of the proteins are determined using X-ray analysis or NMR. The SUPFAM database described here relates two homologous protein families in a multiple sequence alignment database of either known or unknown structure. The present release (1.1), which is the first version of the SUPFAM database, has been derived by analysing Pfam, which is one of the commonly used databases of multiple sequence alignments of homologous proteins. The first step in establishing SUPFAM is to relate Pfam families with the families in PALI, which is an alignment database of homologous proteins of known structure that is derived largely from SCOP. The second step involves relating Pfam families which could not be associated reliably with a protein superfamily of known structure. The profile matching procedure, IMPALA, has been used in these steps. The first step resulted in identification of 1280 Pfam families (out of 2697, i.e. 47%) which are related, either by close homologous connection to a SCOP family or by distant relationship to a SCOP family, potentially forming new superfamily connections. Using the profiles of 1417 Pfam families with apparently no structural information, an all-against-all comparison involving a sequence-profile match using IMPALA resulted in clustering of 67 homologous protein families of Pfam into 28 potential new superfamilies. Expansion of groups of related proteins of yet unknown structural information, as proposed in SUPFAM, should help in identifying 'priority proteins' for structure determination in structural genomics initiatives to expand the coverage of structural information in the protein sequence space. For example, we could assign 858 distinct Pfam domains in 2203 of the gene products in the genome of Mycobacterium tubercolosis. Fifty-one of these Pfam families of unknown structure could be clustered into 17 potentially new superfamilies forming good targets for structural genomics. SUPFAM database can be accessed at http://pauling.mbu.iisc.ernet.in/~supfam.

Animals↗

Comparative analysis of chloroplast genomes: functional annotation, genome-based phylogeny, and deduced evolutionary patterns.

All protein sequences from 19 complete chloroplast genomes (cpDNA) have been studied using a new computational method able to analyze functional correlations among series of protein sequences contained in complete proteomes. First, all open reading frames (ORFs) from the cpDNAs, comprising a total of 2266 protein sequences, were compared against the 3168 proteins from Synechocystis PCC6803 complete genome to find functionally related orthologous proteins. Additionally, all cpDNA genomes were pairwise compared to find orthologous groups not present in cyanobacteria. Annotations in the cluster of othologous proteins database and CyanoBase were used as reference for the functional assignments. Following this protocol, new functional assignments were made for ORFs of unknown function and for ycfs (hypothetical chloroplast frames), which still lack a functional assignment. Using this information, a matrix of functional relationships was derived from profiles of the presence and/or absence of orthologous proteins; the matrix included 1837 proteins in 277 orthologous clusters. A factor analysis study of this matrix, followed by cluster analysis, allowed us to obtain accurate phylogenetic reconstructions and the detection of genes probably involved in speciation as phylogenetic correlates. Finally, by grouping common evolutionary patterns, we show that it is possible to determine functionally linked protein networks. This has allowed us to suggest putative associations for some unknown ORFs.

Bacterial Proteins↗

The T2T genome assembly of watershield (Brasenia schreberi) unveils genomic insights into aquatic adaptation.

Watershield (Brasenia schreberi), belonging to Cabombaceae within the order Nymphaeales, represents one of the early-diverged angiosperm lineages. This perennial floating leaf freshwater aquatic plant features submerged juvenile leaves enveloped in a thick layer of transparent gelatinous mucilage, aiding in its resistance to aquatic stress. However, the evolutionary history of the mechanisms underlying its specific phenotype remains unclear. In this study, we present the telomere-to-telomere level genome of B. schreberi, unveiling that it underwent two rounds of whole-genome duplications (WGDs) and a recent whole-genome triplication, with the most ancient WGD being shared by Nymphaeaceae. WGD and dispersed duplication significantly contributed to the expansion of gene families, which are primarily associated with environmental adaptation. Additionally, we discovered that mature leaves primarily conduct photosynthesis and may transport nutrients to underwater juvenile leaves for polysaccharide synthesis. We also identified an ancestral broad expression pattern of ABC genes, and the similar expression of anthocyanin biosynthesis genes across all flower organs resulted in entirely purple flowers. Our findings deepen the understanding of the evolution of this specific aquatic plant phenotypes.

Genome, Plant↗

Vibrio cholerae phage K139: complete genome sequence and comparative genomics of related phages.

In this report, we characterize the complete genome sequence of the temperate phage K139, which morphologically belongs to the Myoviridae phage family (P2 and 186). The prophage genome consists of 33,106 bp, and the overall GC content is 48.9%. Forty-four open reading frames were identified. Homology analysis and motif search were used to assign possible functions for the genes, revealing a close relationship to P2-like phages. By Southern blot screening of a Vibrio cholerae strain collection, two highly K139-related phage sequences were detected in non-O1, non-O139 strains. Combinatorial PCR analysis revealed almost identical genome organizations. One region of variable gene content was identified and sequenced. Additionally, the tail fiber genes were analyzed, leading to the identification of putative host-specific sequence variations. Furthermore, a K139-encoded Dam methyltransferase was characterized.

Bacteriophages↗

The human genome and comparative genomics: understanding human evolution, biology, and medicine.

The entire 2.9-billion-letter sequence (nucleotide base pairs) of the human genome is available as a resource for scientific discovery. Some of the findings from the completion of the human genome were expected, confirming knowledge anticipated by many years of research and analysis in both human and comparative genetics. Other findings were not expected. In either case, the availability of the human genome is likely to have significant implications on basic research, clinical investigation, and ultimately the practice of medicine.

Biological Evolution↗

A whole-genome linkage scan suggests several genomic regions potentially containing QTLs underlying the variation of stature.

Human height is a complex trait under the control of both genetic and environment factors. In order to identify genomic regions underlying the variation of stature, we performed a whole-genome linkage analysis on a sample of 53 human pedigrees containing 1,249 sib pairs, 1,098 grandparent-grandchildren pairs, 1,993 avuncular pairs, and 1,172 first-cousin pairs. Several genomic regions were suggested by our study to be linked with human height variation. These regions include 5q31 at 144 cM from pter on chromosome 5 (with a maximum LOD score of 2.14 in multipoint linkage analyses), Xp22 at the marker DXS1060, and Xq25 at DXS1001 on the X chromosome (with LOD scores of 1.95 and 1.91, respectively, in two-point linkage analyses). Noticeably, Xp22 happens to be the very region where a newly identified gene underlying idiopathic short stature, SHOX, maps. Based on our findings, further confirmation and fine-mapping studies are to be pursued on expanded samples and/or with denser markers for eventual identification of major functional genes involved in human height variation.

Body Height↗

Pleomorphic Liposarcoma: Comprehensive Genomic Analysis of 39 Cases With Comparison to Other Genomically Complex Sarcomas.

Pleomorphic liposarcoma (PLPS) is an aggressive high-grade sarcoma that often shows diverse morphological features and can mimic high-grade undifferentiated pleomorphic sarcoma (UPS)/spindle cell sarcoma or myxofibrosarcoma (MFS), especially when pleomorphic lipoblasts are sparse. The molecular profile of PLPS is distinct from well differentiated/dedifferentiated liposarcoma and myxoid liposarcoma. In this study, we investigate 39 cases of PLPS by comprehensive genomic profiling, occurring in 32 patients with available molecular data. Cases were reviewed and morphologic parameters-lipoblastic component, UPS-like, and MFS-like areas were estimated. The genomic findings were collected and compared to UPS and MFS groups studied using the same platform. The cohort included 15 females and 17 males, with a median age of 56.5 (range, 34-78). The lower extremity (n = 17) was the most common site involved, followed by upper extremity (n = 5) and pelvis (n = 5). UPS-like and MFS-like patterns were the most common morphologic variants, ranging from 15% to 95% and 20% to 90%, respectively. TP53 (87%) and RB1 (51%) mutations and copy number alterations were the most common alterations seen, followed by ATRX (36%). Compared to UPS and MFS, TP53 and RB1 gene alterations were significantly more common in PLPS. Conversely, CDKN2A/B deletions were infrequent in PLPS. Survival analysis showed that MYC amplification was associated with significantly shorter overall survival in PLPS. Among histologic variants, CYSLTR2 alterations were found to be highest in cases with predominantly pleomorphic lipoblasts; additionally, strong correlations were found between gene alteration frequencies of MFS and MFS-like PLPS, and between UPS and UPS-like PLPS. RB1 allele-specific copy number analysis showed loss of heterozygosity in 82% of cases. Our cohort of PLPS showed a complex molecular landscape with distinct genetic alterations, histologic correlations, and clinical outcomes, highlighting its unique position among genomically complex sarcomas and providing insights that may inform future diagnostic and therapeutic approaches.

Humans↗

Sampling the genomic pool of protein tyrosine kinase genes using the polymerase chain reaction with genomic DNA.

The polymerase chain reaction (PCR), with cDNA as template, has been widely used to identify members of protein families from many species. A major limitation of using cDNA in PCR is that detection of a family member is dependent on temporal and spatial patterns of gene expression. To circumvent this restriction, and in order to develop a technique that is broadly applicable we have tested the use of genomic DNA as PCR template to identify members of protein families in an expression-independent manner. This test involved amplification of DNA encoding protein tyrosine kinase (PTK) genes from the genomes of three animal species that are well known development models; namely, the mouse Mus musculus, the fruit fly Drosophila melanogaster, and the nematode worm Caenorhabditis elegans. Ten PTK genes were identified from the mouse, 13 from the fruit fly, and 13 from the nematode worm. Among these kinases were 13 members of the PTK family that had not been reported previously. Selected PTKs from this screen were shown to be expressed during development, demonstrating that the amplified fragments did not arise from pseudogenes. This approach will be useful for the identification of many novel members of gene families in organisms of agricultural, medical, developmental and evolutionary significance and for analysis of gene families from any species, or biological sample whose habitat precludes the isolation of mRNA. Furthermore, as a tool to hasten the discovery of members of gene families that are of particular interest, this method offers an opportunity to sample the genome for new members irrespective of their expression pattern.

Amino Acid Sequence↗

Enrichment for loci identical-by-descent between pairs of mouse or human genomes by genomic mismatch scanning.

Mapping genes that underlie complex genetic traits, including genes that determine susceptibility to common diseases, requires an efficient method for high-resolution genotyping. Single-nucleotide differences between pairs of allelic sequences from unrelated individuals occur approximately once in every kilobase. Genomic mismatch scanning (GMS), by analyzing numerous single-nucleotide polymorphisms in a single genome-wide step, offers a potentially powerful and efficient approach to linkage analysis. GMS, originally developed in a yeast system, is shown here to be applicable to the more complex mouse and human genomes.

Adenosine Triphosphatases↗

A whole-genome radiation hybrid map of the dog genome.

A whole genome radiation hybrid (RH) map of the canine genome was constructed by typing 400 markers, including 218 genes and 182 microsatellites, on a panel of 126 radiation hybrid cell lines. Fifty-seven RH groups have been determined with lod scores greater than 6, and 180 framework landmarks were ordered with odds greater than 1000:1. Average spacing between adjacent markers is 23 cR5000, an estimated physical distance of 3.8 Mb. Fourteen groups have been assigned to 9 of the canine chromosomes, and a comparison of RH and genetic groups allowed the successful bridging of both types of data on one map composed of 31 RH and 13 syntenic RH groups. Comparison of canine, human, mouse, and pig maps underlined regions of conserved synteny. This integrated map, covering an estimated 80% of the dog genome, should prove a powerful tool for localizing and identifiying genes implicated in pathological and phenotypical traits.

Animals↗

The probability of occurrence of oligomer motifs in the human genome and genomic microheterogeneity.

A previously published method for predicting the frequency of random occurrence of a completely specified DNA oligomer in a longer sequence dataset has been generalized to allow degeneracy in the oligomer sequence. With this enhancement, several datasets consisting of sequences from the human genome were searched for the occurrence of consensus binding sites for a set of 13 transcription factors. Although because of the biological significance of these sequences one might predict that they would occur more often than the random frequency, many of the consensus oligomers were found at lower than expected frequencies. Several (G+C)-rich oligomers were found to be moderately over-represented, but this could be accounted for, in part, by the occurrence of (G+C)-rich tracts in the human sequences. Regions very high in (G+C) were found to occur at much higher frequencies than expected in the human genome, and this severely limits the usefulness of this approach for predicting the frequency of (G+C)-rich oligomers. Unexpectedly, more than 1% of the human genome consists of tracts at least 28 bp in length with a (G+C) content greater than 85%.

Base Sequence↗

From chromosomal alterations to target genes for therapy: integrating cytogenetic and functional genomic views of the breast cancer genome.

A vast number of recurrent chromosomal alterations have been implicated in cancer development and progression. However, most of the genes involved in recurrent chromosomal alterations in solid tumors remain unknown, despite the recent substantial progress in genomic research and availability of high-throughput technologies. For example, it is now possible to quickly identify large numbers of differentially expressed genes in cancer specimens using cDNA microarrays. Integration of this "functional genomic view" of the cancer genome with the "cytogenetic view" could lead to the identification of genes playing a critical role in cancer development and progression. In this review, we illustrate how the combination of three different microarray technologies, cDNA, CGH, and tissue microarrays, makes it possible to directly identify genes involved in chromosomal rearrangements in cell line model systems and then rapidly explore their significance as potential diagnostic and therapeutic targets in human primary breast cancer progression.

Breast Neoplasms↗

The genome of the archaeal virus SIRV1 has features in common with genomes of eukaryal viruses.

The virus SIRV1 of the extremely thermophilic archaeon Sulfolobus has a double-stranded DNA genome similar in architecture to the genomes of eukaryal viruses of the families Poxviridae, Pycodnaviridae, and Asfarviridae: the two strands of the 32,301 bp long linear genome are covalently connected forming a continuous polynucleotide chain and 2029 kb long inverted repeats are present at the termini. Very likely it also shares with these viruses mechanisms of initiation of replication and resolution of replicative intermediates.

Base Sequence↗

Pulsed-field gel electrophoresis analysis of the genome of Rhodococcus fascians: genome size and linear and circular replicon composition in virulent and avirulent strains.

Total DNA of virulent and avirulent strains of Rhodococcus fascians was resolved by pulsed-field gel electrophoresis (PFGE) into a discrete number of fragments by digestion with the endonucleases AseI and DraI. Restriction endonucleases PacI, PmeI, and SwaI yielded no fragments upon digestion of R. fascians genome, and all the other tested endonucleases recognizing 6 bp released too many fragments. The genome size was 5.6 megabases for the type strain R. fascians DSM 20669, and 5.8 megabases for the virulent R. fascians D188 strain. However the genome size of R. fascians CECT 3001 (NRRL B15096) was 8.0 megabases. No linear chromosome in the megabase range was observed under pulse conditions in which Saccharomyces cerevisiae and Schizosaccharomyces pombe chromosomes were perfectly resolved, suggesting that the R. fascians chromosome is circular. A new linear plasmid pIRN640 of 640 kb was found in the avirulent R. fascians CECT 3001 that did not hybridize with a probe internal to the fas region of pFiD188 known to be involved in plant pathogenicity in the virulent strain R. fascians D188. Virulence was correlated in all strains tested with the presence of the fas region. The AseI and DraI bands corresponding to the extrachromosomal elements were identified providing the basis for a physical map of this organism.

DNA, Circular↗