PubMed Health⌕ Search

Biomedical subjects

Andrew J Gentles

Publications and source records attributed to Andrew J Gentles.

9 recordsLinked to original sources

Annotation, submission and screening of repetitive elements in Repbase: RepbaseSubmitter and Censor.

BACKGROUND: Repbase is a reference database of eukaryotic repetitive DNA, which includes prototypic sequences of repeats and basic information described in annotations. Updating and maintenance of the database requires specialized tools, which we have created and made available for use with Repbase, and which may be useful as a template for other curated databases. RESULTS: We describe the software tools RepbaseSubmitter and Censor, which are designed to facilitate updating and screening the content of Repbase. RepbaseSubmitter is a java-based interface for formatting and annotating Repbase entries. It eliminates many common formatting errors, and automates actions such as calculation of sequence lengths and composition, thus facilitating curation of Repbase sequences. In addition, it has several features for predicting protein coding regions in sequences; searching and including Pubmed references in Repbase entries; and searching the NCBI taxonomy database for correct inclusion of species information and taxonomic position. Censor is a tool to rapidly identify repetitive elements by comparison to known repeats. It uses WU-BLAST for speed and sensitivity, and can conduct DNA-DNA, DNA-protein, or translated DNA-translated DNA searches of genomic sequence. Defragmented output includes a map of repeats present in the query sequence, with the options to report masked query sequence(s), repeat sequences found in the query, and alignments. CONCLUSION: Censor and RepbaseSubmitter are available as both web-based services and downloadable versions. They can be found at http://www.girinst.org/repbase/submission.html (RepbaseSubmitter) and http://www.girinst.org/censor/index.php (Censor).

Animals↗

Origin and diversification of minisatellites derived from human Alu sequences.

We analyze minisatellites derived from Alu fragments corresponding approximately to the first 44 bases of human Alu consensus sequences from different subfamilies. The origin of Alu-derived minisatellites appears to have been mediated by short flanking repeats, as first proposed by Haber and Louis [Haber, J.E., Louis, E.J., 1998. Minisatellite origins in yeast and humans. Genomics 48, 132-135.]. We also present evidence for base substitutions and deletions introduced to minisatellites by gene conversion with partially similar but unrelated flanking regions. Segments flanked by short direct repeats are relatively common in different regions of Alu and other repetitive sequences. Our analysis shows that they can be effectively used in comparative studies of the overall sequence context which may contribute to instability of DNA segments flanked by short direct repeats.

Alu Elements↗

Retroposition of processed pseudogenes: the impact of RNA stability and translational control.

Human processed pseudogenes are copies of cellular RNAs reverse transcribed and inserted into the nuclear genome by the enzymatic machinery of L1 (LINE1) non-LTR retrotransposons. Although it is generally accepted that germline expression is crucial for the heritable retroposition of cellular mRNAs, little is known about the influences of RNA stability, mRNA quality control and compartmentalization of translation on the retroposition of processed pseudogenes. We found that frequently retroposed human mRNAs are derived from stable transcripts with translation-competent functional reading frames that are resistant to nonsense-mediated RNA decay. They are preferentially translated on free cytoplasmic ribosomes and encode soluble proteins. Our results indicate that interactions between mRNAs and L1 proteins seem to occur at free cytoplasmic ribosomes.

Animals↗

Evolutionary diversity and potential recombinogenic role of integration targets of Non-LTR retrotransposons.

Short interspersed elements (SINEs) make up a significant fraction of total DNA in mammalian genomes, providing a rich substrate for chromosomal rearrangements by SINE-SINE recombinations. Proliferation of mammalian SINEs is mediated primarily by long interspersed element 1 (L1) non-long terminal repeat retrotransposons that preferentially integrate at DNA sequence targets with an average length of approximately 15 bp and containing conserved endonucleolytic nicking signals at both ends. We report that sequence variations in the first of the two nicking signals, represented by a 5'-TT-AAAA consensus sequence, affect the position of the second signal thus leading to target site duplications (TSDs) of different lengths. The length distribution of TSDs appears to be affected also by L1-encoded enzyme variants because targets with the same 5' nicking site can be of different average lengths in different mammalian species. Taking this into account, we reanalyzed the second nicking site and found that it is larger and includes more conserved sites than previously appreciated, with a consensus of 5'-ANTNTN-AA. We also studied potential involvement of the nicking sites in stimulating recombinations between SINEs. We determined that SINEs retaining TSDs with perfect 5'-TT-AAAA nicking sites appear to be lost relatively rapidly from the human and rat genomes and less rapidly from dog. We speculate that the introduction of DNA breaks induced by recurring endonucleolytic attacks at these sites, combined with the ubiquitousness of SINEs, may significantly promote recombination between repetitive elements, leading to the observed losses. At the same time, new L1 subfamilies may be selected for "incompatibility" with preexisting targets. This provides a possible driving force for the continual emergence of new L1 subfamilies which, in turn, may affect selection of L1-dependent SINE subfamilies.

Animals↗

Traffic of genetic information between segmental duplications flanking the typical 22q11.2 deletion in velo-cardio-facial syndrome/DiGeorge syndrome.

Velo-cardio-facial syndrome/DiGeorge syndrome results from unequal crossing-over events between two 240-kb low-copy repeats termed LCR22 (LCR22-2 and LCR22-4) on Chromosome 22q11.2, comprised of modules, each of which are >99% identical in sequence. To delineate regions in the LCR22s that might contain hotspots for 22q11.2 rearrangements, we scanned the interval for increased rates of recombination with the hypothesis that these regions might be more prone to breakage. We generated an algorithm to detect sites of altered recombination by searching for single nucleotide polymorphic positions in BAC clones from different libraries mapped to LCR22-2 and LCR22-4. This method distinguishes single nucleotide polymorphisms from paralogous sequence variants and complex polymorphic positions. Sites of shared polymorphism are considered potential sites of gene conversion or double cross-over between the two LCR22s. We found an inverse correlation between regions of paralogous sequence variants that are unique to a given position within one LCR22 and clusters of shared polymorphic sites, suggesting that these clusters depict altered recombination and not remnants of ancestral single nucleotide polymorphisms. We postulate that most shared polymorphic sites are products of past transfers of DNA information between the LCR22s, suggesting that frequent traffic of genetic material may induce genomic instability in the two LCR22s. We also found that gaps up to 1.5 kb long can be transferred between LCR22s.

Algorithms↗

Genome comparisons and analysis.

As we enter the post-genomic era, with the accelerating availability of complete genome sequences, new theoretical approaches and new experimental techniques, our ability to dissect cellular processes at the molecular level continues to expand. Recent advances include the application of RNA interference methods to characterize loss-of-function phenotype genes in higher eukaryotes, comparative analysis of the human and mouse genome sequences, and methods for reconciling contradictory phylogenetic reconstructions. New developments feed into the increasingly rich content of databases such as the COG database. The next phase of research will be increasingly dominated by efforts to integrate the deluge of data into our understanding of biological systems.

Amino Acid Sequence↗

Associations between human disease genes and overlapping gene groups and multiple amino acid runs.

Overlapping gene groups (OGGs) arise when exons of one gene are contained within the introns of another. Typically, the two overlapping genes are encoded on opposite DNA strands. OGGs are often associated with specific disease phenotypes. In this report, we identify genes with OGG architecture and genes encoding multiple long amino acid runs and examine their relations to diseases. OGGs appear to be susceptible to genomic rearrangements as happens commonly with the loci of the DiGeorge syndrome on human chromosome 22. We also examine the degree of conservation of OGGs between human and mouse. Our analyses suggest that (i) a high proportion of genes in OGG regions are disease-associated, (ii) genomic rearrangements are likely to occur within OGGs, possibly as a consequence of anomalous sequence features prevalent in these regions, and (iii) multiple amino acid runs are also frequently associated with pathologies.

Amino Acids↗

Genes, pseudogenes, and Alu sequence organization across human chromosomes 21 and 22.

Human chromosomes 21 and 22 (mainly the q-arms) were the first complete parts of the human genome released. Our analysis of genes, pseudogenes (Psig), and Alu repeats across these chromosomes include the following findings: The number of gene structures containing untranslated exons exceeds 25%; the terminal exon tends to be the largest among exons, whereas, the initial intron tends to be the largest among introns; single-exon gene length is approximately the mean gene exon number times the mean internal exon length; processed Psig lengths are on average approximately the same as single-exon gene length; and the G+C content and length of genes are uncorrelated. The counts and distribution of genes, Psig, and Alu sequences and G+C variation are evaluated with respect to clusters and overdispersions. Other assessments concern comparisons of intergenic lengths, properties of Psig sequences, and correlations between Alu and Psig sequences.

Alu Elements↗

Amino acid runs in eukaryotic proteomes and disease associations.

We present a comparative proteome analysis of the five complete eukaryotic genomes (human, Drosophila melanogaster, Caenorhabditis elegans, Saccharomyces cerevisiae, Arabidopsis thaliana), focusing on individual and multiple amino acid runs, charge and hydrophobic runs. We found that human proteins with multiple long runs are often associated with diseases; these include long glutamine runs that induce neurological disorders, various cancers, categories of leukemias (mostly involving chromosomal translocations), and an abundance of Ca(2 +) and K(+) channel proteins. Many human proteins with multiple runs function in development and/or transcription regulation and are Drosophila homeotic homologs. A large number of these proteins are expressed in the nervous system. More than 80% of Drosophila proteins with multiple runs seem to function in transcription regulation. The most frequent amino acid runs in Drosophila sequences occur for glutamine, alanine, and serine, whereas human sequences highlight glutamate, proline, and leucine. The most frequent runs in yeast are of serine, glutamine, and acidic residues. Compared with the other eukaryotic proteomes, amino acid runs are significantly more abundant in the fly. This finding might be interpreted in terms of innate differences in DNA-replication processes, repair mechanisms, DNA-modification systems, and mutational biases. There are striking differences in amino acid runs for glutamine, asparagine, and leucine among the five proteomes.

Animals↗