PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Comparative genomic analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Comparative genomic analysis of the mouse and rat amylase multigene family.

The rat and mouse amylase gene families were characterized using sequence data from the UCSC genome assembly. We found that the rat genome contains one amylase-1 and two amylase-2 genes, lying close to one another on the same chromosome. Detailed analysis revealed at least six additional amylase pseudogenes in the rat genome in the region adjacent to the amylase-2 genes. In contrast, the mouse has one amylase-1 gene and five amylase-2 genes; the latter are tandemly and systematically arranged on the same chromosome and were generated by segmental duplication. Detailed analysis revealed that the mouse has two amylase pseudogenes, located 5' to the five amylase-2 segments. Thus, the amylase genes of mouse and rat tend to be amplified; the sequences of some of them are fixed while others have become pseudogenes during evolution. This is the second report of amylase genomic organization in mammals and the first in the rodents.

Amylases↗

Comparative genomic analysis of the MHC: the evolution of class I duplication blocks, diversity and complexity from shark to man.

The major histocompatibility complex (MHC) genomic region is composed of a group of linked genes involved functionally with the adaptive and innate immune systems. The class I and class II genes are intrinsic features of the MHC and have been found in all the jawed vertebrates studied so far. The MHC genomic regions of the human and the chicken (B locus) have been fully sequenced and mapped, and the mouse MHC sequence is almost finished. Information on the MHC genomic structures (size, complexity, genic and intergenic composition and organization, gene order and number) of other vertebrates is largely limited or nonexistent. Therefore, we are mapping, sequencing and analyzing the MHC genomic regions of different human haplotypes and at least eight nonhuman species. Here, we review our progress with these sequences and compare the human MHC structure with that of the nonhuman primates (chimpanzee and rhesus macaque), other mammals (pigs, mice and rats) and nonmammalian vertebrates such as birds (chicken and quail), bony fish (medaka, pufferfish and zebrafish) and cartilaginous fish (nurse shark). This comparison reveals a complex MHC structure for mammals and a relatively simpler design for nonmammalian animals with a hypothetical prototypic structure for the shark. In the mammalian MHC, there are two to five different class I duplication blocks embedded within a framework of conserved nonclass I and/or nonclass II genes. With a few exceptions, the class I framework genes are absent from the MHC of birds, bony fish and sharks. Comparative genomics of the MHC reveal a highly plastic region with major structural differences between the mammalian and nonmammalian vertebrates. Additional genomic data are needed on animals of the reptilia, crocodilia and marsupial classes to find the origins of the class I framework genes and examples of structures that may be intermediate between the simple and complex MHC organizations of birds and mammals, respectively.

Animals↗

Identifying sigma factors in Mycobacterium smegmatis by comparative genomic analysis.

Mycobacterium smegmatis is a saprophytic species that has been used for 15 years as a model to perform heterologous regulation and virulence studies of Mycobacterium tuberculosis. Members of the extracytoplasmic sigma factors family, which are required for adaptive responses to various environmental stresses, are responsible for some of the virulence traits of M. tuberculosis. A bioinformatic search on the genome of M. smegmatis has predicted the existence of 26 sigma factors, which is twice the number that are present in M. tuberculosis. A phylogenetic analysis has shown that despite this high number of sigma factors the orthologs of the genes sigC, sigI and sigK of M. tuberculosis are absent in the M. smegmatis genome. Several sigma factors are specific for M. smegmatis, with a special enrichment in the sigH and, to a lesser extent, in the sigJ and sigL subfamily, pinpointing the potential variability of the repertoire of adaptive response in this saprophytic species.

Adaptation, Physiological↗

Comparative genomic analysis of the Haloferax volcanii DS2 and Halobacterium salinarium GRB contig maps reveals extensive rearrangement.

Anonymous probes from the genome of Halobacterium salinarium GRB and 12 gene probes were hybridized to the cosmid clones representing the chromosome and plasmids of Halobacterium salinarium GRB and Haloferax volcanii DS2. The order of and pairwise distances between 35 loci uniquely cross-hybridizing to both chromosomes were analyzed in a search for conservation. No conservation between the genomes could be detected at the 15-kbp resolution used in this study. We found distinct sets of low-copy-number repeated sequences in the chromosome and plasmids of Halobacterium salinarium GRB, indicating some degree of partitioning between these replicons. We propose alternative courses for the evolution of the haloarchaeal genome: (i) that the majority of genomic differences that exist between genera came about at the inception of this group or (ii) that the differences have accumulated over the lifetime of the lineage. The strengths and limitations of investigating these models through comparative genomic studies are discussed.

Blotting, Southern↗

Prediction of functional modules based on comparative genome analysis and Gene Ontology application.

We present a computational method for the prediction of functional modules encoded in microbial genomes. In this work, we have also developed a formal measure to quantify the degree of consistency between the predicted and the known modules, and have carried out statistical significance analysis of consistency measures. We first evaluate the functional relationship between two genes from three different perspectives--phylogenetic profile analysis, gene neighborhood analysis and Gene Ontology assignments. We then combine the three different sources of information in the framework of Bayesian inference, and we use the combined information to measure the strength of gene functional relationship. Finally, we apply a threshold-based method to predict functional modules. By applying this method to Escherichia coli K12, we have predicted 185 functional modules. Our predictions are highly consistent with the previously known functional modules in E.coli. The application results have demonstrated that our approach is highly promising for the prediction of functional modules encoded in a microbial genome.

Bayes Theorem↗

Comparative genomic analysis of tumors: detection of DNA losses and amplification.

We demonstrate the use of representational difference analysis for cloning probes that detect DNA loss and amplification in tumors. Using DNA isolated from human tumor cell lines to drive hybridization against matched normal DNA, we were able to identify six genomic regions that are homozygously deleted in cultured cancer cells. When this method was applied in the reverse way, using normal DNA to drive hybridization against tumor cell DNA, we readily isolated probes detecting amplification. Representational difference analysis was also performed on DNAs derived from tumor biopsies, and we thereby discovered a probe detecting very frequent homozygous loss in colon cancer cell lines and located on chromosome 3p.

Animals↗

Assigning protein functions by comparative genome analysis: protein phylogenetic profiles.

Determining protein functions from genomic sequences is a central goal of bioinformatics. We present a method based on the assumption that proteins that function together in a pathway or structural complex are likely to evolve in a correlated fashion. During evolution, all such functionally linked proteins tend to be either preserved or eliminated in a new species. We describe this property of correlated evolution by characterizing each protein by its phylogenetic profile, a string that encodes the presence or absence of a protein in every known genome. We show that proteins having matching or similar profiles strongly tend to be functionally linked. This method of phylogenetic profiling allows us to predict the function of uncharacterized proteins.

Bacterial Proteins↗

Comparative genomics analysis of human sequence variation in the UGT1A gene cluster.

Common polymorphisms within the human UGT1A gene locus are associated with irinotecan and tranilast toxicity. To uncover additional functional variation across this gene cluster, cross-species sequence comparisons were performed. Evolutionarily conserved segments (a total of 47.1 kb) were re-sequenced in 24 African-American, 24 European-American, and 24 Asian individuals, and 381 segregating sites (including 123 singletons) were identified. Highly conserved coding sites were less likely to be polymorphic than diverged sites (P<0.0001) but this pattern was not observed at non-coding sites (P=0.1025). Among coding variants, the distribution of those computationally predicted to affect function was skewed toward low frequencies. Some alleles occurred at similar frequencies in each population; others had wide disparities. Although strong linkage disequilibrium was detected among the hepatically expressed genes, the degree of linkage disequilibrium varied among populations. These results suggest that rare functional gene variants and inter-population variability must be considered in the interpretation of association studies between UGT1A and drug metabolism/toxicity phenotypes.

Animals↗

BLOCK-based PCR markers to find gene family members in human and comparative genome analysis.

Degenerate primer pairs that include consensus sequences of evolutionary conserved portions of protein families (BLOCKs or ancient conserved regions) can be used to screen by polymerase chain reaction (PCR) for cognate cDNAs and YACs through much of phylogeny. Nine such primer pairs were developed, and five with sites on human chromosomes 7 or X were shown to identify YACs from chromosome-specific locations, including a candidate for a new zinc finger gene in Xq28. When linked to contig-based genomic maps, such BLOCK-based PCR assays may provide a route to recover the members and study the development of families containing up to 40% of genes, in genomes as diverse as humans, nematodes, and yeast.

Actins↗

Comparative genomic analysis of solvent extrusion pumps in Pseudomonas strains exhibiting different degrees of solvent tolerance.

Organic solvents are inherently toxic for microorganisms. Their effects depend not only on the nature of the compound, but also on the intrinsic tolerance of the bacterial species and strains. Three efflux pumps belonging to the RND (resistance-nodulation-cell division) family of multidrug extrusion pumps are the main factor involved in the high intrinsic tolerance to toluene of Pseudomonas putida DOT-T1E. We have analyzed the tolerance to toluene shocks [0.1% and 0.3% (v/v)] of a number of strains belonging to different species of the genus Pseudomonas upon growth in the absence and in the presence of sublethal concentrations of toluene. The strains can be grouped in three categories: (1) highly resistant strains, in which almost 100% of the cells precultured in the presence of sublethal concentrations of toluene withstood a 0.3% (v/v) toluene shock, (2) moderately resistant strains, in which only a fraction (10(-4)-1) of the cells withstood a 0.1% (v/v) toluene shock, but fewer than 1 in 10(7) cells survived a sudden 0.3% (v/v) toluene shock regardless of the growth conditions, and (3) sensitive strains, in which regardless of the growth conditions fewer than 10(-5) cells survived a 0.1% (v/v) toluene shock. We also studied the number and type of efflux pumps in different strains in comparison with the P. putida DOT-T1E strain.

Adaptation, Physiological↗

Comparative genomic analysis in the region of a major Plasmodium-refractoriness locus of Anopheles gambiae.

We have sequenced six overlapping clones from a library of bacterial artificial chromosome (BAC) clones derived from a laboratory strain of the mosquito, Anopheles gambiae, the major vector of human malaria in Africa. The resulting uninterrupted 528-kb sequence is from the 8C region of the mosquito 2R chromosome, at or very near the major refractoriness locus associated with melanotic encapsulation of parasites. This sequence represents the first extensive view of the mosquito genome structure encompassing 48 genes. Genomic comparison reveals that the majority of the orthologues are found in six microsyntenic clusters in Drosophila melanogaster. A BAC clone that is wholly contained within this region demonstrates the existence of a remarkable degree of local polymorphism in this species, which may prove important for its population structure and vectorial capacity.

Amino Acid Sequence↗

Comparative genomic analysis identifies an evolutionary shift of vomeronasal receptor gene repertoires in the vertebrate transition from water to land.

Two evolutionarily unrelated superfamilies of G-protein coupled receptors, V1Rs and V2Rs, bind pheromones and "ordinary" odorants to initiate vomeronasal chemical senses in vertebrates, which play important roles in many aspects of an organism's daily life such as mating, territoriality, and foraging. To study the macroevolution of vomeronasal sensitivity, we identified all V1R and V2R genes from the genome sequences of 11 vertebrates. Our analysis suggests the presence of multiple V1R and V2R genes in the common ancestor of teleost fish and tetrapods and reveals an exceptionally large among-species variation in the sizes of these gene repertoires. Interestingly, the ratio of the number of intact V1R genes to that of V2R genes increased by approximately 50-fold as land vertebrates evolved from aquatic vertebrates. A similar increase was found for the ratio of the number of class II odorant receptor (OR) genes to that of class I genes, but not in other vertebrate gene families. Because V1Rs and class II ORs have been suggested to bind to small airborne chemicals, whereas V2Rs and class I ORs recognize water-soluble molecules, these increases reflect a rare case of adaptation to terrestrial life at the gene family level. Several gene families known to function in concert with V2Rs in the mouse are absent outside rodents, indicating rapid changes of interactions between vomeronasal receptors and their molecular partners. Taken together, our results demonstrate the exceptional evolutionary fluidity of vomeronasal receptors, making them excellent targets for studying the molecular basis of physiological and behavioral diversity and adaptation.

Animals↗

Comparative genomic analysis of the eight-membered ring cystine knot-containing bone morphogenetic protein antagonists.

TGF-beta family proteins with a cystine knot motif serve as ligands for diverse families of plasma membrane receptors. Bone morphogenetic protein (BMP) antagonists represent a subgroup of these proteins, some of which bind BMPs and antagonize their actions during development and morphogenesis. Availability of completed genome sequences from diverse organisms allows bioinformatic analysis of the evolution of BMP antagonists and facilitates their classification. Using a regular expression algorithm (http://BioRegEx.stanford.edu), an exhaustive search of the human genome identified all cystine knot-containing BMP antagonists. Based on the size of the cystine ring, these proteins were divided into three subfamilies: CAN (eight-membered ring), twisted gastrulation (nine-membered ring), as well as chordin and noggin (10-membered ring). The CAN family can be divided further into four subgroups based on a conserved arrangement of additional cysteine residues-gremlin and PRDC, cerberus and coco, and DAN, together with USAG-1 and sclerostin. We searched for orthologs of human BMP antagonists in the genomes of model organisms and analyzed their phylogenetic relationship. New human paralogs were identified together with the verification of orthologous relationships of known genes. We also discuss the physiological roles of the CAN subfamily of BMP antagonists and the associated genetic defects. Based on the known three-dimensional structure of key cystine knot proteins, we postulated disulfide bondings for eight-membered ring BMP antagonists to predict their potential folding and dimerization.

Animals↗

Comparative genomic analysis of hyperthermophilic archaeal Fuselloviridae viruses.

The complete genome sequences of two Sulfolobus spindle-shaped viruses (SSVs) from acidic hot springs in Kamchatka (Russia) and Yellowstone National Park (United States) have been determined. These nonlytic temperate viruses were isolated from hyperthermophilic Sulfolobus hosts, and both viruses share the spindle-shaped morphology characteristic of the Fuselloviridae family. These two genomes, in combination with the previously determined SSV1 genome from Japan and the SSV2 genome from Iceland, have allowed us to carry out a phylogenetic comparison of these geographically distributed hyperthermal viruses. Each virus contains a circular double-stranded DNA genome of approximately 15 kbp with approximately 34 open reading frames (ORFs). These Fusellovirus ORFs show little or no similarity to genes in the public databases. In contrast, 18 ORFs are common to all four isolates and may represent the minimal gene set defining this viral group. In general, ORFs on one half of the genome are colinear and highly conserved, while ORFs on the other half are not. One shared ORF among all four genomes is an integrase of the tyrosine recombinase family. All four viral genomes integrate into their host tRNA genes. The specific tRNA gene used for integration varies, and one genome integrates into multiple loci. Several unique ORFs are found in the genome of each isolate.

Archaeal Viruses↗

[Comparative genomic analysis of vibrio cholerae El Tor preseventh and seventh pandemic strains isolated in various periods].

Genetic organization of 52 Vibrio cholerae El Tor biotype preseventh and seventh pandemic strains isolated in various periods was studied by PCR assay and DNA-DNA hybridization. It was established that the genome of most ancient of analyzed strains isolated from a diarrhea patient in 1910 was devoid of CTX and RS1 prophages, vibrio pathogenicity islands (VPI and VPI-2), and pandemic islands (VSP-1 and VSP-2) that contain key virulence genes. The appearance of pathogenic properties in cholera vibrios for the first time causing a local outbreak of cholera in 1937 is connected with the acquisition of VPI and CTX that carried genes tcpA and ctx-AB, respectively, which are responsible for the colonization of small intestine and encode the production of cholera toxin. The appearance of seventh pandemic agent for cholera was shown to correlate with the acquisition by its precursor of two additional blocks of genes VSP-1 and VSP-2. This finding strongly supports the involvement of these genes in formation of the pandemic potential in strains. Molecular typing methods allowed elucidation of differences in the genetic organization between prepandemic and pandemic strains. The detected variability of the genome of contemporary virulent strains may be a reason for the occurrence of etiological agent for cholera with new properties.

Base Sequence↗

Comparative genome analysis of the mouse imprinted gene impact and its nonimprinted human homolog IMPACT: toward the structural basis for species-specific imprinting.

Mouse Impact is a paternally expressed gene encoding an evolutionarily conserved protein of unknown function. Here we identified IMPACT, the human homolog of Impact, on chromosome 18q11. 2-12.1, a region syntenic to the mouse Impact locus. IMPACT was expressed biallelically in brain and in various tissues from two informative fetuses and in peripheral blood from an informative adult. To reveal the structural basis for the difference in allelic expression between the two species, we elucidated complete genome sequences for both mouse Impact ( approximately 38 kb) and human IMPACT ( approximately 30 kb). Sequence comparison revealed that the two genes share a well-conserved exon-intron organization but bear significantly different CpG islands. The mouse island lies in the first intron and contains characteristic tandem repeats. Furthermore, this island serves as a differentially methylated region (DMR) consisting of a hypermethylated maternal allele and an unmethylated paternal allele. Intriguingly, this intronic island is missing from the nonimprinted human IMPACT, whose sole CpG island spans the first exon, lacks any apparent repeats, and escapes methylation on both chromosomes. These results suggest that the intronic DMR plays a role in the imprinting of Impact.

Alleles↗

New findings on evolution of metal homeostasis genes: evidence from comparative genome analysis of bacteria and archaea.

In order to examine the natural history of metal homeostasis genes in prokaryotes, open reading frames with homology to characterized P(IB)-type ATPases from the genomes of 188 bacteria and 22 archaea were investigated. Major findings were as follows. First, a high diversity in N-terminal metal binding motifs was observed. These motifs were distributed throughout bacterial and archaeal lineages, suggesting multiple loss and acquisition events. Second, the CopA locus separated into two distinct phylogenetic clusters, CopA1, which contained ATPases with documented Cu(I) influx activity, and CopA2, which contained both efflux and influx transporters and spanned the entire diversity of the bacterial domain, suggesting that CopA2 is the ancestral locus. Finally, phylogentic incongruences between 16S rRNA and P(IB)-type ATPase gene trees identified at least 14 instances of lateral gene transfer (LGT) that had occurred among diverse microbes. Results from bootstrapped supported nodes indicated that (i) a majority of the transfers occurred among proteobacteria, most likely due to the phylogenetic relatedness of these organisms, and (ii) gram-positive bacteria with low moles percent G+C were often involved in instances of LGT. These results, together with our earlier work on the occurrence of LGT in subsurface bacteria (J. M. Coombs and T. Barkay, Appl. Environ. Microbiol. 70:1698-1707, 2004), indicate that LGT has had a minor role in the evolution of P(IB)-type ATPases, unlike other genes that specify survival in metal-stressed environments. This study demonstrates how examination of a specific locus across microbial genomes can contribute to the understanding of phenotypes that are critical to the interactions of microbes with their environment.

Adenosine Triphosphatases↗

Extrachromosomal gene amplification in acute myeloid leukemia; characterization by metaphase analysis, comparative genomic hybridization, and semi-quantitative PCR.

A case of acute myeloid leukemia (M-3) with complex karyotypic aberrations and double minute (dmin) chromosomes is presented. The patient had no history of prior exposure to mutagenic or carcinogenic agents or of other malignancies. She died from CNS involvement six weeks after the initial diagnosis. We used comparative genomic hybridization to identify the amplified sequences presumed to represent the dmin of the leukemic cells; the tumor/normal ratios indicated increased signal intensity at 8q24. This localization prompted investigation by semi-quantitative PCR that revealed amplification of the MYC oncogene. The extent of chromosome aberrations and the oncogene amplification, both linked with poor prognosis, may relate to the rapid course of this patient's disease.

Base Sequence↗