PubMed Health⌕ Search

Biomedical subjects

Alexei Fedorov

Publications and source records attributed to Alexei Fedorov.

15 recordsLinked to original sources

Advances in the Exon-Intron Database (EID).

Investigation of exon-intron gene structures is a non-trivial task due to enormous expansions of the eukaryotic genomes, great variety of gene forms, and the imperfectness in sequence data. A number of available informational systems on various gene characteristics complement each other and are indispensable for many genomic studies. Among them, the Exon-Intron Database (EID) is a good choice for large-scale computational examination of exon/intron structure and splicing. It has many internal filters that control for sequence quality, consistency of gene descriptions, accordance to standards, and possible errors. New innovations in EID are described. The collection of exons and introns has been extended beyond coding regions and current versions of EID contain data on untranslated regions of gene sequences as well. Intron-less genes are included as a special part of EID. For species with entirely sequenced genomes, species-specific databases have been generated. A novel Mammalian Orthologous Intron Database (MOID) has been introduced which includes the full set of introns that come from orthologous genes that have the same positions relative to the reading frames. Examples of statistical analyses of gene sequences using EID are provided. We present the latest data on our comparison of intron positions in 11,025 orthologous genes of human, mouse and rat, and find no convincing cases of intron gain. We discuss relevant data-quality issues of genomic databases. In particular, 5% of genes in genomic databases contain internal stop codons. This fact is due to a combination of biological reasons and also to errors in sequence annotations. The EID is freely available at www.meduohio.edu/bioinfo/eid/.

Base Sequence↗

Where is the difference between the genomes of humans and annelids?

The first systematic investigation of an annelid genome has revealed that the genes of the marine worm Platynereis dumerilii are more closely related to those of vertebrates than to those of insects or nematodes. For hundreds of millions of years vertebrates have preserved exon-intron structures descended from their last common ancestor with the annelids.

Animals↗

Bioinformatic analysis of exon repetition, exon scrambling and trans-splicing in humans.

MOTIVATION: Using bioinformatic approaches we aimed to characterize poorly understood abnormalities in splicing known as exon scrambling, exon repetition and trans-splicing. RESULTS: We developed a software package that allows large-scale comparison of all human expressed sequence tags (EST) sequences to the entire set of human gene sequences. Among 5,992,495 EST sequences, 401 cases of exon repetition and 416 cases of exon scrambling were found. The vast majority of identified ESTs contain fragments rather than full-length repeated or scrambled exons. Their structures suggest that the scrambled or repeated exon fragments may have arisen in the process of cDNA cloning and not from splicing abnormalities. Nevertheless, we found 11 cases of full-length exon repetition showing that this phenomenon is real yet very rare. In searching for examples of trans-splicing, we looked only at reproducible events where at least two independent ESTs represent the same putative trans-splicing event. We found 15 ESTs representing five types of putative trans-splicing. However, all 15 cases were derived from human malignant tissues and could have resulted from genomic rearrangements. Our results provide support for a very rare but physiological occurrence of exon repetition, but suggest that apparent exon scrambling and trans-splicing result, respectively, from in vitro artifact and gene-level abnormalities. AVAILABILITY: Exon-Intron Database (EID) is available at http://www.meduohio.edu/bioinfo/eid. Programs are available at http://www.meduohio.edu/bioinfo/software.html. The Laboratory website is available at http://www.meduohio.edu/medicine/fedorov SUPPLEMENTARY INFORMATION: Supplementary file is available at http://www.meduohio.edu/bioinfo/software.html.

Algorithms↗

Computer identification of snoRNA genes using a Mammalian Orthologous Intron Database.

Based on comparative genomics, we created a bioinformatic package for computer prediction of small nucleolar RNA (snoRNA) genes in mammalian introns. The core of our approach was the use of the Mammalian Orthologous Intron Database (MOID), which contains all known introns within the human, mouse and rat genomes. Introns from orthologous genes from these three species, that have the same position relative to the reading frame, are grouped in a special orthologous intron table. Our program SNO.pl searches for conserved snoRNA motifs within MOID and reports all cases when characteristic snoRNA-like structures are present in all three orthologous introns of human, mouse and rat sequences. Here we report an example of the SNO.pl usage for searching a particular pattern of conserved C/D-box snoRNA motifs (canonical C- and D-boxes and the 6 nt long terminal stem). In this computer analysis, we detected 57 triplets of snoRNA-like structures in three mammals. Among them were 15 triplets that represented known C/D-box snoRNA genes. Six triplets represented snoRNA genes that had only been partially characterized in the mouse genome. One case represented a novel snoRNA gene, and another three cases, putative snoRNAs. Our programs are publicly available and can be easily adapted and/or modified for searching any conserved motifs within mammalian introns.

Algorithms↗

What does the microsporidian E. cuniculi tell us about the origin of the eukaryotic cell?

The relationship among the three cellular domains Archaea, Bacteria, and Eukarya has become a central problem in unraveling the tree of life. This relationship can now be studied as the completely sequenced genomes of representatives of these cellular domains become available. We performed a bioinformatic investigation of the Encephalitozoon cuniculi proteome. E. cuniculi has the smallest sequenced eukaryotic genome, 2.9 megabases coding for 1997 proteins. The proteins of E. cuniculi were compared with a previously characterized set of eukaryotic signature proteins (ESPs). ESPs are found in a eukaryotic cell, whether from an animal, a plant, a fungus, or a protozoan, but are not found in the Archaea and the Bacteria. We demonstrated that 85% of the ESPs have significant sequence similarity to proteins in E. cuniculi. Hence, E. cuniculi, a minimal eukaryotic cell that has removed all inessential proteins, still preserves most of the ESPs that make it a member of the Eukarya. The locations and functions of these ESPs point to the earliest history of eukaryotes.

Animals↗

Introns: mighty elements from the RNA world.

The discovery of RNA-based enzymatic activity by Thomas Cech's and Sidney Altman's laboratories was a momentous event that led Walter Gilbert to the concept of an "RNA world"--a primitive ancient stage of life that existed before the appearance of DNA and protein molecules. A year later, Gilbert formulated "the exon theory of genes," which hypothesized that introns are very ancient genetic elements present at the earliest stages of life in the RNA world. This theory has been fiercely debated and still has vigorous supporters and opponents. In this communication, we explore peculiarities in the RNA-protein world and their effect on intron-exon structures. We demonstrate that these peculiarities, which exist in the absence of DNA, could shed light on introns' original functions as well as the important role they might have played in the origin of life. For ancient DNA-lacking cells, a crucial problem existed in distinguishing two distinct subsets of RNAs: those messenger molecules coding for proteins and those heritable genetic molecules complementary to messenger RNAs that propagate the genetic information through generations. We propose that ancient introns could act as markers of RNA subsets, directing them to different functions.

Animals↗

Mystery of intron gain.

For nearly 15 years, it has been widely believed that many introns were recently acquired by the genes of multicellular organisms. However, the mechanism of acquisition has yet to be described for a single animal intron. Here, we report a large-scale computational analysis of the human, Drosophila melanogaster, Caenorhabditis elegans, and Arabidopsis thaliana genomes. We divided 147,796 human intron sequences into batches of similar lengths and aligned them with each other. Different types of homologies between introns were found, but none showed evidence of simple intron transposition. Also, 106,902 plant, 39,624 Drosophila, and 6021 C. elegans introns were examined. No single case of homologous introns in nonhomologous genes was detected. Thus, we found no example of transposition of introns in the last 50 million years in humans, in 3 million years in Drosophila and C. elegans, or in 5 million years in Arabidopsis. Either new introns do not arise via transposition of other introns or intron transposition must have occurred so early in evolution that all traces of homology have been lost.

Animals↗

Large-scale comparison of intron positions in mammalian genes shows intron loss but no gain.

We compared intron-exon structures in 1,560 human-mouse orthologs and 360 mouse-rat orthologs. The origin of differences in intron positions between species was inferred by comparison with an outgroup, Fugu for human-mouse and human for mouse-rat. Among 10,020 intron positions in the human-mouse comparison, we found unequivocal evidence for five independent intron losses in the mouse lineage but no evidence for intron loss in humans or for intron gain in either lineage. Among 1,459 positions in rat-mouse comparisons, we found evidence for one loss in rat but neither loss in mouse nor gain in either lineage. In each case, the intron losses were exact, without change in the surrounding coding sequence, and involved introns that are extremely short, with an average of 200 bp, an order of magnitude shorter than the mammalian average. These results favor a model whereby introns are lost through gene conversion with intronless copies of the gene. In addition, the finding of widespread conservation of intron-exon structure, even over large evolutionary distances, suggests that comparative methods employing information about gene structures should be very successful in correctly predicting exon boundaries in genomic sequences.

Amino Acid Sequence↗

Phylogenetically older introns strongly correlate with module boundaries in ancient proteins.

The hypothesis that some (but not all) introns were used to construct ancient genes by exon shuffling of modules at the earliest stages of evolution is supported by the finding of an excess of phase-zero intron positions in the boundary regions of such modules in 276 ancient proteins (defined as common to eukaryotes and prokaryotes). Here we show further that as phase-zero intron positions are shared by distant taxa, and thus are truly phylogenetically ancient, their excess in the boundaries becomes greater, rising to an 80% excess if shared by four out of the five taxa: vertebrates, invertebrates, fungi, plants, and protists.

Computational Biology↗

Introns in gene evolution.

Introns are integral elements of eukaryotic genomes that perform various important functions and actively participate in gene evolution. We review six distinct roles of spliceosomal introns: (1) sources of non-coding RNA; (2) carriers of transcription regulatory elements; (3) actors in alternative and trans-splicing; (4) enhancers of meiotic crossing over within coding sequences; (5) substrates for exon shuffling; and (6) signals for mRNA export from the nucleus and nonsense-mediated decay. We consider transposable capacities of introns and the current state of the long-lasting debate on the 'early-or-late' origin of introns. Cumulative data on known types of contemporary exon shuffling and the estimation of the size of the underlying exon universe are also discussed. We argue that the processes central to introns-early (exon shuffling) and introns-late (intron insertion) theories are entirely compatible. Each has provided insight: the latter through elucidating the transposon capabilities of introns, and the former through understanding the importance of introns in genomic recombination leading to gene rearrangements and evolution.

Alternative Splicing↗

Large-scale comparison of intron positions among animal, plant, and fungal genes.

We purge large databases of animal, plant, and fungal intron-containing genes to a 20% similarity level and then identify the most similar animal-plant, animal-fungal, and plant-fungal protein pairs. We identify the introns in each BLAST 2.0 alignment and score matched intron positions and slid (near-matched, within six nucleotides) intron positions automatically. Overall we find that 10% of the animal introns match plant positions, and a further 7% are "slides." Fifteen percent of fungal introns match animal positions, and 13% match plant positions. Furthermore, the number of alignments with high numbers of matches deviates greatly from the Poisson expectation. The 30 animal-plant alignments with the highest matches (for which 44% of animal introns match plant positions) when aligned with fungal genes are also highly enriched for triple matches: 39% of the fungal introns match both animal and plant positions. This is strong evidence for ancestral introns predating the animal-plant-fungal divergence, and in complete opposition to any expectations based on random insertion. In examining the slid introns, we show that at least half are caused by imperfections in the alignments, and are most likely to be actual matches at common positions. Thus, our final estimates are that approximately equal 14% of animal introns match plant positions, and that approximately equal 17-18% of fungal introns match animal or plant positions, all of these being likely to be ancestral in the eukaryotes.

Amino Acid Sequence↗

The signal of ancient introns is obscured by intron density and homolog number.

In ancient genes whose products have known 3-dimensional structures, an excess of phase zero introns (those that lie between the codons) appear in the boundaries of modules, compact regions of the polypeptide chain. These excesses are highly significant and could support the hypothesis that ancient genes were assembled by exon shuffling involving compact modules. (Phase one and two introns, and many phase zero introns, appear to arise later.) However, as more genes, with larger numbers of homologs and intron positions, were examined, the effects became smaller, dropping from a 40% excess to an 8% excess as the number of intron positions increased from 570 to 3,328, even though the statistical significance remained strong. An interpretation of this behavior is that novel inserted positions appearing in homologs washed out the signal from a finite number of ancient positions. Here we show that this is likely to be the case. Analyses of intron positions restricted to those in genes for which relatively few intron positions from homologs are known, or to those in genes with a small number of known homologous gene structures, show a significant correlation of phase zero intron positions with the module structure, which weakens as the density of attributed intron positions or the number of homologs increases. These effects do not appear for phase one and phase two introns. This finding matches the expectation of the mixed model of intron origin, in which a fraction of phase zero introns are left from the assembly of the first genes, while other introns have been added in the course of evolution.

Evolution, Molecular↗

Regularities of context-dependent codon bias in eukaryotic genes.

Nucleotides surrounding a codon influence the choice of this particular codon from among the group of possible synonymous codons. The strongest influence on codon usage arises from the nucleotide immediately following the codon and is known as the N1 context. We studied the relative abundance of codons with N1 contexts in genes from four eukaryotes for which the entire genomes have been sequenced: Homo sapiens, Drosophila melanogaster, Caenorhabditis elegans and Arabidopsis thaliana. For all the studied organisms it was found that 90% of the codons have a statistically significant N1 context-dependent codon bias. The relative abundance of each codon with an N1 context was compared with the relative abundance of the same 4mer oligonucleotide in the whole genome. This comparison showed that in about half of all cases the context-dependent codon bias could not be explained by the sequence composition of the genome. Ranking statistics were applied to compare context-dependent codon biases for codons from different synonymous groups. We found regularities in N1 context-dependent codon bias with respect to the codon nucleotide composition. Codons with the same nucleotides in the second and third positions and the same N1 context have a statistically significant correlation of their relative abundances.

Animals↗

The origin of the eukaryotic cell: a genomic investigation.

We have collected a set of 347 proteins that are found in eukaryotic cells but have no significant homology to proteins in Archaea and Bacteria. We call these proteins eukaryotic signature proteins (ESPs). The dominant hypothesis for the formation of the eukaryotic cell is that it is a fusion of an archaeon with a bacterium. If this hypothesis is accepted then the three cellular domains, Eukarya, Archaea, and Bacteria, would collapse into two cellular domains. We have used the existence of this set of ESPs to test this hypothesis. The evidence of the ESPs implicates a third cell (chronocyte) in the formation of the eukaryotic cell. The chronocyte had a cytoskeleton that enabled it to engulf prokaryotic cells and a complex internal membrane system where lipids and proteins were synthesized. It also had a complex internal signaling system involving calcium ions, calmodulin, inositol phosphates, ubiquitin, cyclin, and GTP-binding proteins. The nucleus was formed when a number of archaea and bacteria were engulfed by a chronocyte. This formation of the nucleus would restore the three cellular domains as the Chronocyte was not a cell that belonged to the Archaea or to the Bacteria.

Animals↗

Do introns favor or avoid regions of amino acid conservation?

Are intron positions correlated with regions of high amino acid conservation? For a set of ancient conserved proteins, with intronless prokaryotic but intron-containing eukaryotic homologs, multiple sequence alignments identified residues invariant throughout evolution. Intron positions between codons show no preferences. However, introns lying after the first base of a codon prefer conserved regions, markedly in glycines. Because glycines are in excess in conserved regions, this behavior could reflect phase-one introns entering glycine residues randomly in the ancestral sequences. Examination of intron positions within codons of evolutionarily invariable amino acids showed that roughly 50% of these introns are bordered by guanines at both 5'- and 3'-ends, 25% have a G only before the intron, and 5% have a G only after the intron, whereas about 20% are bordered by nonguanine bases.

Alternative Splicing↗