PubMed Health⌕ Search

Biomedical subjects

Valer Gotea

Publications and source records attributed to Valer Gotea.

5 recordsLinked to original sources

iSoMAs: Finding isoform expression and somatic mutation associations in human cancers.

Aberrant alternative splicing, prevalent in cancer, impacts various cancer hallmarks involving proliferation, angiogenesis, and invasion. Splicing disruption often results from somatic point mutations rewiring functional pathways to support cancer cell survival. We introduce iSoMAs (iSoform expression and somatic Mutation Association), an efficient computational pipeline leveraging principal component analysis technique, to explore how somatic mutations influence transcriptome-wide gene expression at the isoform level. Applying iSoMAs to 33 cancer types comprising 9,738 tumor samples in The Cancer Genome Atlas, we identified 908 somatically mutated genes significantly associated with altered isoform expression across three or more cancer types. Mutations linked to differential isoform expression occurred through both cis- and trans-acting mechanisms, involving well-known oncogenes/suppressor genes, RNA binding protein and splicing factor genes. With wet-lab experiments, we verified direct association between TP53 mutations and differential isoform expression in cell cycle genes. Additional iSoMAs genes have been validated in the literature with independent cohorts and/or methods. Despite the complexity of cancer, iSoMAs attains computational efficiency via dimension reduction strategy and reveals critical associations between regulatory factors and transcriptional landscapes.

Humans↗

Spliceosomal small nuclear RNA genes in 11 insect genomes.

The removal of introns from the primary transcripts of protein-coding genes is accomplished by the spliceosome, a large macromolecular complex of which small nuclear RNAs (snRNAs) are crucial components. Following the recent sequencing of the honeybee (Apis mellifera) genome, we used various computational methods, ranging from sequence similarity search to RNA secondary structure prediction, to search for putative snRNA genes (including their promoters) and to examine their pattern of conservation among 11 available insect genomes (A. mellifera, Tribolium castaneum, Bombyx mori, Anopheles gambiae, Aedes aegypti, and six Drosophila species). We identified candidates for all nine spliceosomal snRNA genes in all the analyzed genomes. All the species contain a similar number of snRNA genes, with the exception of A. aegypti, whose genome contains more U1, U2, and U5 genes, and A. mellifera, whose genome contains fewer U2 and U5 genes. We found that snRNA genes are generally more closely related to homologs within the same genus than to those in other genera. Promoter regions for all spliceosomal snRNA genes within each insect species share similar sequence motifs that are likely to correspond to the PSEA (proximal sequence element A), the binding site for snRNA activating protein complex, but these promoter elements vary in sequence among the five insect families surveyed here. In contrast to the other insect species investigated, Dipteran genomes are characterized by a rapid evolution (or loss) of components of the U12 spliceosome and a striking loss of U12-type introns.

Animals↗

Do transposable elements really contribute to proteomes?

Recent studies indicate that the initial classification of transposable elements (TEs) as 'useless', 'selfish' or 'junk' pieces of DNA is not an accurate one. TEs seem to have complex regulatory functions and contribute to the coding regions of many genes. Because this contribution had been documented only at transcript level, we searched for evidence that would also support the translation of TE cassettes. Our findings suggest that the proportion of proteins with TE-encoded fragments (approximately 0.1%), although probably underestimated, is much less than what the data at transcript level suggest (approximately 4%). In all cases, the TE cassettes are derived from old TEs, consistent with the idea that incorporation (exaptation) of TE fragments into functional proteins requires long evolutionary periods. We therefore argue that functional proteins are unlikely to contain TE cassettes derived from young TEs, the role of which is probably limited to regulatory functions.

Animals↗

Transposable elements as a significant source of transcription regulating signals.

Transposable elements (TEs) are major components of eukaryotic genomes, contributing about 50% to the size of mammalian genomes. TEs serve as recombination hot spots and may acquire specific cellular functions, such as controlling protein translation and gene transcription. The latter is the subject of the analysis presented. We scanned TE sequences located in promoter regions of all annotated genes in the human genome for their content in potential transcription regulating signals. All investigated signals are likely to be over-represented in at least one TE class, which shows that TEs have an important potential to contribute to pre-transcriptional gene regulation, especially by moving transcriptional signals within the genome and thus potentially leading to new gene expression patterns. We also found that some TE classes are more likely than others to carry transcription regulating signals, which can explain why they have different retention rates in regions neighboring genes.

Base Sequence↗

Mastering seeds for genomic size nucleotide BLAST searches.

One of the most common activities in bioinformatics is the search for similar sequences. These searches are usually carried out with the help of programs from the NCBI BLAST family. As the majority of searches are routinely performed with default parameters, a question that should be addressed is how reliable the results obtained using the default parameter values are, i.e. what fraction of potential matches have been retrieved by these searches. Our primary focus is on the initial hit parameter, also known as the seed or word, used by the NCBI BLASTn, MegaBLAST and other similar programs in searches for similar nucleotide sequences. We show that the use of default values for the initial hit parameter can have a big negative impact on the proportion of potentially similar sequences that are retrieved. We also show how the hit probability of different seeds varies with the minimum length and similarity of sequences desired to be retrieved and describe methods that help in determining appropriate seeds. The experimental results described in this paper illustrate situations in which these methods are most applicable and also show the relationship between the various BLAST parameters.

Algorithms↗