PubMed Health⌕ Search

Biomedical subjects

Rajinder Kaul

Publications and source records attributed to Rajinder Kaul.

16 recordsLinked to original sources

A comprehensive transposon mutant library of Francisella novicida, a bioweapon surrogate.

Francisella tularensis, the causative agent of tularemia, is one of the most infectious bacterial pathogens known and is a category A select agent. We created a sequence-defined, near-saturation transposon mutant library of F. tularensis novicida, a subspecies that causes a tularemia-like disease in rodents. The library consists of 16,508 unique insertions, an average of >9 insertions per gene, which is a coverage nearly twice that of the greatest previously achieved for any bacterial species. Insertions were recovered in 84% (1,490) of the predicted genes. To achieve high coverage, it was necessary to construct transposons carrying an endogenous Francisella promoter to drive expression of antibiotic resistance. An analysis of genes lacking (or with few) insertions identified nearly 400 candidate essential genes, most of which are likely to be required for growth on rich medium and which represent potential therapeutic targets. To facilitate genome-scale screening using the mutant collection, we assembled a sublibrary made up of two purified mutants per gene. The library provides a resource for virtually complete identification of genes involved in virulence and other nonessential processes.

Alleles↗

Potential source of Francisella tularensis live vaccine strain attenuation determined by genome comparison.

Francisella tularensis is a bacterial pathogen that causes the zoonotic disease tularemia and is important to biodefense. Currently, the only vaccine known to confer protection against tularemia is a specific live vaccine strain (designated LVS) derived from a virulent isolate of Francisella tularensis subsp. holarctica. The origin and source of attenuation of this strain are not known. To assist with the design of a defined live vaccine strain, we sought to determine the genetic basis of the attenuation of LVS. This analysis relied primarily on the comparison between the genome of LVS and Francisella tularensis holarctica strain FSC200, which differ by only 0.08% of their nucleotide sequences. Under the assumption that the attenuation was due to a loss of function(s), only coding regions were examined in this comparison. To complement this analysis, the coding regions of two slightly more distantly related Francisella tularensis strains were also compared against the LVS coding regions. Thirty-five genes show unique sequence variations predicted to alter the protein sequence in LVS compared to the other Francisella tularensis strains. Due to these polymorphisms, the functions of 15 of these genes are very likely lost or impaired. Seven of these genes were demonstrated to be under stronger selective constraints, suggesting that they are the most probable to be the source of LVS attenuation and useful for a newly defined vaccine.

Bacterial Vaccines↗

Genetic adaptation by Pseudomonas aeruginosa to the airways of cystic fibrosis patients.

In many human infections, hosts and pathogens coexist for years or decades. Important examples include HIV, herpes viruses, tuberculosis, leprosy, and malaria. With the exception of intensively studied viral infections such as HIV/AIDs, little is known about the extent to which the clonal expansion that occurs during long-term infection by pathogens involves important genetic adaptations. We report here a detailed, whole-genome analysis of one such infection, that of a cystic fibrosis (CF) patient by the opportunistic bacterial pathogen Pseudomonas aeruginosa. The bacteria underwent numerous genetic adaptations during 8 years of infection, as evidenced by a positive-selection signal across the genome and an overwhelming signal in specific genes, several of which are mutated during the course of most CF infections. Of particular interest is our finding that virulence factors that are required for the initiation of acute infections are often selected against during chronic infections. It is apparent that the genotypes of the P. aeruginosa strains present in advanced CF infections differ systematically from those of "wild-type" P. aeruginosa and that these differences may offer new opportunities for treatment of this chronic disease.

Adaptation, Physiological↗

The DNA sequence, annotation and analysis of human chromosome 3.

After the completion of a draft human genome sequence, the International Human Genome Sequencing Consortium has proceeded to finish and annotate each of the 24 chromosomes comprising the human genome. Here we describe the sequencing and analysis of human chromosome 3, one of the largest human chromosomes. Chromosome 3 comprises just four contigs, one of which currently represents the longest unbroken stretch of finished DNA sequence known so far. The chromosome is remarkable in having the lowest rate of segmental duplication in the genome. It also includes a chemokine receptor gene cluster as well as numerous loci involved in multiple human cancers such as the gene encoding FHIT, which contains the most common constitutive fragile site in the genome, FRA3B. Using genomic sequence from chimpanzee and rhesus macaque, we were able to characterize the breakpoints defining a large pericentric inversion that occurred some time after the split of Homininae from Ponginae, and propose an evolutionary history of the inversion.

Animals↗

High-throughput genotyping of intermediate-size structural variation.

The contribution of large-scale and intermediate-size structural variation (ISV) to human genetic disease and disease susceptibility is only beginning to be understood. The development of high-throughput genotyping technologies is one of the most critical aspects for future studies of linkage disequilibrium (LD) and disease association. Using a simple PCR-based method designed to assay the junctions of the breakpoints, we genotyped seven simple insertion and deletion polymorphisms ranging in size from 6.3 to 24.7 kb among 90 CEPH individuals. We then extended this analysis to a larger collection of samples (n=460) by application of an oligonucleotide extension-ligation genotyping assay. The analysis showed a high level of concordance ( approximately 99%) when compared with PCR/sequence-validated genotypes. Using the available HapMap data, we observed significant LD (r2=0.74-0.95) between each ISV and flanking single nucleotide polymorphisms, but this observation is likely to hold only for similar simple insertion/deletion events. The approach we describe may be used to characterize a large number of individuals in a cost-effective manner once the sequence organization of ISVs is known.

Cohort Studies↗

Foamy virus vector integration sites in normal human cells.

Foamy viruses (FVs) or spumaviruses are retroviruses that have been developed as vectors, but their integration patterns have not been described. We have performed a large-scale analysis of FV integration sites in unselected human fibroblasts (n = 1,008) and human CD34(+) hematopoietic cells (n = 1,821) by using a bacterial shuttle vector and a comparable analysis of lentiviral vector integration sites in CD34(+) cells (n = 1,331). FV vectors had a distinct integration profile relative to other types of retroviruses. They did not integrate preferentially within genes, despite a modest preference for integration near transcription start sites and a significant preference for CpG islands. The genomewide distribution of FV vector proviruses was nonrandom, with both clusters and gaps. Transcriptional profiling showed that gene expression had little influence on integration site selection. Our findings suggest that FV vectors may have desirable integration properties for gene therapy applications.

Antigens, CD34↗

Type IV pili-mediated secretion modulates Francisella virulence.

Francisella tularensis are the causative agent of the zoonotic disease, tularaemia. Among four F. tularensis subspecies, ssp. novicida (F. novicida) is pathogenic only for immunocompromised individuals, while all four subspecies are pathogenic for mice. This study utilized proteomic and bioinformatic approaches to identify seven F. novicida secreted proteins and the corresponding Type IV pilus (T4P) secretion system. The secreted proteins were predicted to encode two chitinases, a chitin binding protein, a protease (PepO), and a beta-glucosidase (BglX). The transcription of F. novicida pepO and bglX was regulated by the virulence regulator MglA. Intradermal infection of mice with F. novicida mutants defective in T4P secretion system or PepO resulted in enhanced F. novicida spread to systemic sites. Infection with F. novicida pepO mutants also resulted in increased neutrophil infiltration into the mouse airways. PepO is a zinc protease that is homologous to mammalian endothelin-converting enzyme ECE-1. Therefore, secretion of PepO likely results in increased production of endothelin and increased vasoconstriction at the infection site in skin that limits the F. novicida spread. Francisella human pathogenic strains contain a mutation in pepO predicted to abolish its secretion. Loss of PepO function may have contributed to evolution of highly virulent Francisellae.

Animals↗

Targeted, haplotype-resolved resequencing of long segments of the human genome.

Currently, challenges exist to acquire long-range (hundreds of kilobase pairs) phase-discriminated sequence across substantial numbers of individuals. We have developed a straightforward method for isolating and characterizing specific genomic regions in a haplospecific manner. Real-time PCR is carried out to STS content map and genotype pools of fosmid clones arrayed in 384-well microtiter plates. Single-nucleotide polymorphisms, microsatellite markers, and insertion-deletion polymorphisms are used to differentiate the target region into haplotype-specific tiling paths. DNA of clones from these tiling paths is retrieved from the library and either sequenced by standard shotgun methods or amplified in vitro and sequenced by a primer-based, directed method. This approach provides convenient access to complete, haplotype-resolved resequencing data from multiple individuals across tens to hundreds of thousands of basepairs. We illustrate its implementation with a detailed example of more than 400 kbp from the human CFTR region, across 15 individuals, and summarize our experience applying it to many other human loci.

Base Sequence↗

Fine-scale structural variation of the human genome.

Inversions, deletions and insertions are important mediators of disease and disease susceptibility. We systematically compared the human genome reference sequence with a second genome (represented by fosmid paired-end sequences) to detect intermediate-sized structural variants >8 kb in length. We identified 297 sites of structural variation: 139 insertions, 102 deletions and 56 inversion breakpoints. Using combined literature, sequence and experimental analyses, we validated 112 of the structural variants, including several that are of biomedical relevance. These data provide a fine-scale structural variation map of the human genome and the requisite sequence precision for subsequent genetic studies of human disease.

Base Pairing↗

Ancient haplotypes of the HLA Class II region.

Allelic variation in codons that specify amino acids that line the peptide-binding pockets of HLA's Class II antigen-presenting proteins is superimposed on strikingly few deeply diverged haplotypes. These haplotypes appear to have been evolving almost independently for tens of millions of years. By complete resequencing of 20 haplotypes across the approximately 100-kbp region that spans the HLA-DQA1, -DQB1, and -DRB1 genes, we provide a detailed view of the way in which the genome structure at this locus has been shaped by the interplay of selection, gene-gene interaction, and recombination.

Alleles↗

Evidence for diversifying selection at the pyoverdine locus of Pseudomonas aeruginosa.

Pyoverdine is the primary siderophore of the gram-negative bacterium Pseudomonas aeruginosa. The pyoverdine region was recently identified as the most divergent locus alignable between strains in the P. aeruginosa genome. Here we report the nucleotide sequence and analysis of more than 50 kb in the pyoverdine region from nine strains of P. aeruginosa. There are three divergent sequence types in the pyoverdine region, which correspond to the three structural types of pyoverdine. The pyoverdine outer membrane receptor fpvA may be driving diversity at the locus: it is the most divergent alignable gene in the region, is the only gene that showed substantial intratype variation that did not appear to be generated by recombination, and shows evidence of positive selection. The hypothetical membrane protein PA2403 also shows evidence of positive selection; residues on one side of the membrane after protein folding are under positive selection. R', previously identified as a type IV strain, is clearly derived from a type III strain via a 3.4-kb deletion which removes one amino acid from the pyoverdine side chain peptide. This deletion represents a natural modification of the product of a nonribosomal peptide synthetase enzyme, whose consequences are predictive from the DNA sequence. There is also linkage disequilibrium between the pyoverdine region and pvdY, a pyoverdine gene separated by 30 kb from the pyoverdine region. The pyoverdine region shows evidence of horizontal transfer; we propose that some alleles in the region were introduced from other soil bacteria and have been subsequently maintained by diversifying selection.

Bacterial Outer Membrane Proteins↗

Large-scale analysis of adeno-associated virus vector integration sites in normal human cells.

The integration sites of viral vectors used in human gene therapy can have important consequences for safety and efficacy. However, an extensive evaluation of adeno-associated virus (AAV) vector integration sites has not been completed, despite the ongoing use of AAV vectors in clinical trials. Here we have used a shuttle vector system to isolate and analyze 977 unique AAV vector-chromosome integration junctions from normal human fibroblasts and describe their genomic distribution. We found a significant preference for integrating within CpG islands and the first 1 kb of genes, but only a slight overall preference for transcribed sequences. Integration sites were clustered throughout the genome, including a major preference for integration in ribosomal DNA repeats, and 13 other hotspots that contained three or more proviruses within a 500-kb window. Both junctions were localized from 323 proviruses, allowing us to characterize the chromosomal deletions, insertions, and translocations associated with vector integration. These studies establish a profile of insertional mutagenesis for AAV vectors and provide unique insight into the chromosomal distribution of DNA strand breaks that may facilitate integration.

Chromosome Mapping↗

Comprehensive transposon mutant library of Pseudomonas aeruginosa.

We have developed technologies for creating saturating libraries of sequence-defined transposon insertion mutants in which each strain is maintained. Phenotypic analysis of such libraries should provide a virtually complete identification of nonessential genes required for any process for which a suitable screen can be devised. The approach was applied to Pseudomonas aeruginosa, an opportunistic pathogen with a 6.3-Mbp genome. The library that was generated consists of 30,100 sequence-defined mutants, corresponding to an average of five insertions per gene. About 12% of the predicted genes of this organism lacked insertions; many of these genes are likely to be essential for growth on rich media. Based on statistical analyses and bioinformatic comparison to known essential genes in E. coli, we estimate that the actual number of essential genes is 300-400. Screening the collection for strains defective in two defined multigenic processes (twitching motility and prototrophic growth) identified mutants corresponding to nearly all genes expected from earlier studies. Thus, phenotypic analysis of the collection may produce essentially complete lists of genes required for diverse biological activities. The transposons used to generate the mutant collection have added features that should facilitate downstream studies of gene expression, protein localization, epistasis, and chromosome engineering.

Escherichia coli↗

The DNA sequence of human chromosome 7.

Human chromosome 7 has historically received prominent attention in the human genetics community, primarily related to the search for the cystic fibrosis gene and the frequent cytogenetic changes associated with various forms of cancer. Here we present more than 153 million base pairs representing 99.4% of the euchromatic sequence of chromosome 7, the first metacentric chromosome completed so far. The sequence has excellent concordance with previously established physical and genetic maps, and it exhibits an unusual amount of segmentally duplicated sequence (8.2%), with marked differences between the two arms. Our initial analyses have identified 1,150 protein-coding genes, 605 of which have been confirmed by complementary DNA sequences, and an additional 941 pseudogenes. Of genes confirmed by transcript sequences, some are polymorphic for mutations that disrupt the reading frame.

Animals↗

Whole-genome sequence variation among multiple isolates of Pseudomonas aeruginosa.

Whole-genome shotgun sequencing was used to study the sequence variation of three Pseudomonas aeruginosa isolates, two from clonal infections of cystic fibrosis patients and one from an aquatic environment, relative to the genomic sequence of reference strain PAO1. The majority of the PAO1 genome is represented in these strains; however, at least three prominent islands of PAO1-specific sequence are apparent. Conversely, approximately 10% of the sequencing reads derived from each isolate fail to align with the PAO1 backbone. While average sequence variation among all strains is roughly 0.5%, regions of pronounced differences were evident in whole-genome scans of nucleotide diversity. We analyzed two such divergent loci, the pyoverdine and O-antigen biosynthesis regions, by complete resequencing. A thorough analysis of isolates collected over time from one of the cystic fibrosis patients revealed independent mutations resulting in the loss of O-antigen synthesis alternating with a mucoid phenotype. Overall, we conclude that most of the PAO1 genome represents a core P. aeruginosa backbone sequence while the strains addressed in this study possess additional genetic material that accounts for at least 10% of their genomes. Approximately half of these additional sequences are novel.

Adolescent↗

Genetic variation at the O-antigen biosynthetic locus in Pseudomonas aeruginosa.

The outer carbohydrate layer, or O antigen, of Pseudomonas aeruginosa varies markedly in different isolates of these bacteria, and at least 20 distinct O-antigen serotypes have been described. Previous studies have indicated that the major enzymes responsible for O-antigen synthesis are encoded in a cluster of genes that occupy a common genetic locus. We used targeted yeast recombinational cloning to isolate this locus from the 20 internationally recognized serotype strains. DNA sequencing of these isolated segments revealed that at least 11 highly divergent gene clusters occupy this region. Homology searches of the encoded protein products indicated that these gene clusters are likely to direct O-antigen biosynthesis. The O15 serotype strains lack functional gene clusters in the region analyzed, suggesting that O-antigen biosynthesis genes for this serotype are harbored in a different portion of the genome. The overall pattern underscores the plasticity of the P. aeruginosa genome, in which a specific site in a well-conserved genomic region can be occupied by any of numerous islands of functionally related DNA with diverse sequences.

Amino Acid Sequence↗