PubMed Health⌕ Search

Biomedical subjects

W Makałowski

Publications and source records attributed to W Makałowski.

17 recordsLinked to original sources

The human genome structure and organization.

Genetic information of human is encoded in two genomes: nuclear and mitochondrial. Both of them reflect molecular evolution of human starting from the beginning of life (about 4.5 billion years ago) until the origin of Homo sapiens species about 100,000 years ago. From this reason human genome contains some features that are common for different groups of organisms and some features that are unique for Homo sapiens. 3.2 x 10(9) base pairs of human nuclear genome are packed into 23 chromosomes of different size. The smallest chromosome - 21st contains 5 x 10(7) base pairs while the biggest one -1st contains 2.63 x 10(8) base pairs. Despite the fact that the nucleotide sequence of all chromosomes is established, the organisation of nuclear genome put still questions: for example: the exact number of genes encoded by the human genome is still unknown giving estimations from 30 to 150 thousand genes. Coding sequences represent a few percent of human nuclear genome. The majority of the genome is represented by repetitiVe sequences (about 50%) and noncoding unique sequences. This part of the genome is frequently wrongly called "junk DNA". The distribution of genes on chromosomes is irregular, DNA fragments containing low percentage of GC pairs code lower number of genes than the fragments of high percentage of GC pairs.

Animals↗

Genomic scrap yard: how genomes utilize all that junk.

Interspersed repetitive sequences are major components of eukaryotic genomes. Repetitive elements comprise over 50% of the mammalian genome. Because the specific function of these elements remains to be defined and because of their unusual 'behaviour' in the genome, they are often quoted as a selfish or junk DNA. Our view of the entire phenomenon of repetitive elements has to now be revised in the light of data on their biology and evolution, especially in the light of what we know about the retroposons. I would like to argue that even if we cannot define the specific function of these elements, we still can show that they are not useless pieces of the genomes. The repetitive elements interact with the surrounding sequences and nearby genes. They may serve as recombination hot spots or acquire specific cellular functions such as RNA transcription control or even become part of protein coding regions. Finally, they provide very efficient mechanism for genomic shuffling. As such, repetitive elements should be called genomic scrap yard rather than junk DNA. Tables listing examples of recruited (exapted) transposable elements are available at http://www.ncbi.nlm.gov/Makalowski/ScrapYard/

Animals↗

Frequent human genomic DNA transduction driven by LINE-1 retrotransposition.

Human L1 retrotransposons can produce DNA transduction events in which unique DNA segments downstream of L1 elements are mobilized as part of aberrant retrotransposition events. That L1s are capable of carrying out such a reaction in tissue culture cells was elegantly demonstrated. Using bioinformatic approaches to analyze the structures of L1 element target site duplications and flanking sequence features, we provide evidence suggesting that approximately 15% of full-length L1 elements bear evidence of flanking DNA segment transduction. Extrapolating these findings to the 600,000 copies of L1 in the genome, we predict that the amount of DNA transduced by L1 represents approximately 1% of the genome, a fraction comparable with that occupied by exons.

3' Untranslated Regions↗

Human and nematode orthologs--lessons from the analysis of 1800 human genes and the proteome of Caenorhabditis elegans.

Recently, we have defined and analyzed over 1800 orthologous human and rodent genes. Here we extend this work to compare human and Caenorhabditis elegans coding sequences. 1880 human proteins were compared with about 20000 predicted nematode proteins presumably comprising nearly the complete proteome of C. elegans. We found that 44% of human/rodent orthologs have convincing nematode counterparts. On average, the amino acid similarity and identity between aligned human and C. elegans orthologous gene products are 69.3% and 49.1% respectively, and the nucleotide identity is 49.8%. Detailed investigation of our results suggests that some nematode gene predictions are incorrect, leading to erroneous pairing with human genes (e.g. calcineurin and polymerase II elongation factor III). Furthermore, other proteins (i.e. homologs of human ribosomal proteins S20 and L41, thymosin) are missing entirely from the nematode proteome, suggesting that it may not be complete. These results underscore the fact that metazoan gene prediction is a very challenging task and that most computer-predicted nematode genes require supporting evidence of their existence from comparative genomics and/or laboratory investigation.

Animals↗

Comparative analysis of 1196 orthologous mouse and human full-length mRNA and protein sequences.

A large set of mRNA and encoded protein sequences, from orthologous murine and human genes, was compiled to analyze statistical, biological, and evolutionary properties of coding and noncoding transcribed sequences. Protein sequence conservation varied between 36% and 100% identity, with an average value of 85%. The average degree of nucleotide sequence identity for the corresponding coding sequences was also approximately 85%, whereas 5' and 3' untranslated regions (UTRs) were less conserved, with aligned identities of 67% and 69%, respectively. For some mouse and human genes, nucleotide sequences are more highly conserved than the encoded protein sequences. A subset of 32 sequences, consisting of only mouse/human protein pairs for which the human sequence represents a positionally cloned disease gene, had properties very similar to the larger data set, suggesting that our data are representative of the genome as a whole. With respect to sequence conservation, two interesting outliers are the breast cancer (BRCAI) gene product and the testis-determining factor (SRY), both of which display among the lowest degrees of sequence identity. The occurrence of both introns and repetitive elements (e.g., Alu, Bl) in 5' and 3' UTRs was also studied. These results provide one benchmark for the "comparative genomics" of mice and humans, with practical implications for the cross-referencing of transcript maps. Also, they should prove useful in estimating the additional sampling diversity provided by mouse EST sequencing projects designed to complement the existing human cDNA collection.

Amino Acid Sequence↗

Alu sequences in the coding regions of mRNA: a source of protein variability.

Dispersion of repetitive sequence elements is a source of genetic variability that contributes to genome evolution. Alu elements, the most common dispersed repeats in the human genome, can cause genetic diseases by several mechanisms, including de novo Alu insertions and splicing of intragenic Alu elements into mRNA. Such mutations might contribute positively to protein evolution if they are advantageous or neutral. To test this hypothesis, we searched the literature and sequence databases for examples of protein-coding regions that contain Alu sequences: 17 Alu 'cassettes' inserted within 15 different coding sequences were found. In three instances, these events caused genetic diseases; the possible functional significance of the other Alu-containing mRNAs is discussed. Our analysis suggests that splice-mediated insertion of intronic elements is the major mechanism by which Alu segments are introduced into mRNAs.

Animals↗

Conserved signals in the 5' flanking region of eukaryotic nuclear tRNA genes.

The statistical analysis of 5' flanking regions of eukaryotic tRNA genes was done. The analysis of nucleotides in the sequence of fungi and invertebrates showed a high content of A and T in the flanking regions versus coding regions where G and C dominate. In contrast to these results in vertebrates sequences the preferences of any nucleotide in flanking regions was not observed. The analysis of tetrads showed five conserved signals: TTGT, (T/A)(T/A)ATA, A(C/T)(C/A)A in the tRNA genes of fungi, (A/T)TGA of invertebrates and (A/T)GAG of vertebrates. The analysis of 3' flanking regions did not show any conserved signals except well known poly-T tracks.

Animals↗

Nucleotide sequence of the mitochondrial 5S rRNA gene from lupine (Lupinus luteus).

A lupine mitochondrial clone containing 5S rRNA gene is characterized. The gene is located on the same strand as 18S rRNA and separated from it by 190 nucleotides. The intergenic region in different plants shows high degree of homology. In the case of lupine and soybean 43 nucleotides upstream of 5S rRNA gene exhibits 100% of homology. Comparisons of lupine 5S rRNA gene sequence with other plant mitochondrial 5S rRNA genes displays high degree of homology (from 89.8% to 95.8%).

Animals↗

Classical oncogenes and tumor suppressor genes: a comparative genomics perspective.

We have curated a reference set of cancer- related genes and reanalyzed their sequences in the light of molecular information and resources that have become available since they were first cloned. Homology studies were carried out for human oncogenes and tumor suppressors, compared with the complete proteome of the nematode, Caenorhabditis elegans, and partial proteomes of mouse and rat and the fruit fly, Drosophila melanogaster. Our results demonstrate that simple, semi-automated bioinformatics approaches to identifying putative functionally equivalent gene products in different organisms may often be misleading. An electronic supplement to this article provides an integrated view of our comparative genomics analysis as well as mapping data, physical cDNA resources and links to published literature and reviews, thus creating a "window" into the genomes of humans and other organisms for cancer biology.

Animals↗