PubMed Health⌕ Search

Biomedical subjects

Arthur Gruber

Publications and source records attributed to Arthur Gruber.

9 recordsLinked to original sources

TRAP: automated classification, quantification and annotation of tandemly repeated sequences.

TRAP, the Tandem Repeats Analysis Program, is a Perl program that provides a unified set of analyses for the selection, classification, quantification and automated annotation of tandemly repeated sequences. TRAP uses the results of the Tandem Repeats Finder program to perform a global analysis of the satellite content of DNA sequences, permitting researchers to easily assess the tandem repeat content for both individual sequences and whole genomes. The results can be generated in convenient formats such as HTML and comma-separated values. TRAP can also be used to automatically generate annotation data in the format of feature table and GFF files.

Algorithms↗

EGene: a configurable pipeline generation system for automated sequence analysis.

UNLABELLED: EGene is a generic, flexible and modular pipeline generation system that makes pipeline construction a modular job. EGene allows for third-party programs to be used and integrated according to the needs of distinct projects and without any previous programming or formal language experience being required. EGene comes with CoEd, a visual tool to facilitate pipeline construction and documentation. A series of components to build pipelines for sequence processing is provided. AVAILABILITY: http://www.lbm.fmvz.usp.br/egene/ CONTACT: alan@ime.usp.br; argruber@usp.br SUPPLEMENTARY INFORMATION: http://www.lbm.fmvz.usp.br/egene/

Chromosome Mapping↗

Identification and complete sequencing of novel human transcripts through the use of mouse orthologs and testis cDNA sequences.

The correct identification of all human genes, and their derived transcripts, has not yet been achieved, and it remains one of the major aims of the worldwide genomics community. Computational programs suggest the existence of 30,000 to 40,000 human genes. However, definitive gene identification can only be achieved by experimental approaches. We used two distinct methodologies, one based on the alignment of mouse orthologous sequences to the human genome, and another based on the construction of a high-quality human testis cDNA library, in an attempt to identify new human transcripts within the human genome sequence. We generated 47 complete human transcript sequences, comprising 27 unannotated and 20 annotated sequences. Eight of these transcripts are variants of previously known genes. These transcripts were characterized according to size, number of exons, and chromosomal localization, and a search for protein domains was undertaken based on their putative open reading frames. In silico expression analysis suggests that some of these transcripts are expressed at low levels and in a restricted set of tissues.

Amino Acid Sequence↗

Characterization of SCAR markers of Eimeria spp. of domestic fowl and construction of a public relational database (The Eimeria SCARdb).

This study reports the development and characterization of 151 sequence characterized amplified region (SCAR) markers for the seven Eimeria species that infect the domestic fowl. From this set, 84 markers are species-specific and 67 present partial specificity. The complete nucleotide sequence was derived for all markers, revealing the presence of micro- and minisatellite repetitive units in 22 SCARs, with up to five distinct repeat units being observed per marker. Only 15 markers showed significant hits in similarity searches against public sequence databases, thus confirming their anonymous and non-coding character. Finally, a relational database of the markers (the Eimeria SCARdb) was developed and made available on the Internet, providing a valuable resource of SCAR markers that can be useful for molecular diagnosis, and also for epizootiological, genetic variability and genome mapping studies.

Animals↗

A transcript finishing initiative for closing gaps in the human transcriptome.

We report the results of a transcript finishing initiative, undertaken for the purpose of identifying and characterizing novel human transcripts, in which RT-PCR was used to bridge gaps between paired EST clusters, mapped against the genomic sequence. Each pair of EST clusters selected for experimental validation was designated a transcript finishing unit (TFU). A total of 489 TFUs were selected for validation, and an overall efficiency of 43.1% was achieved. We generated a total of 59,975 bp of transcribed sequences organized into 432 exons, contributing to the definition of the structure of 211 human transcripts. The structure of several transcripts reported here was confirmed during the course of this project, through the generation of their corresponding full-length cDNA sequences. Nevertheless, for 21% of the validated TFUs, a full-length cDNA sequence is not yet available in public databases, and the structure of 69.2% of these TFUs was not correctly predicted by computer programs. The TF strategy provides a significant contribution to the definition of the complete catalog of human genes and transcripts, because it appears to be particularly useful for identification of low abundance transcripts expressed in a restricted set of tissues as well as for the delineation of gene boundaries and alternatively spliced isoforms.

Alternative Splicing↗

The generation and utilization of a cancer-oriented representation of the human transcriptome by using expressed sequence tags.

Whereas genome sequencing defines the genetic potential of an organism, transcript sequencing defines the utilization of this potential and links the genome with most areas of biology. To exploit the information within the human genome in the fight against cancer, we have deposited some two million expressed sequence tags (ESTs) from human tumors and their corresponding normal tissues in the public databases. The data currently define approximately 23,500 genes, of which only approximately 1,250 are still represented only by ESTs. Examination of the EST coverage of known cancer-related (CR) genes reveals that <1% do not have corresponding ESTs, indicating that the representation of genes associated with commonly studied tumors is high. The careful recording of the origin of all ESTs we have produced has enabled detailed definition of where the genes they represent are expressed in the human body. More than 100,000 ESTs are available for seven tissues, indicating a surprising variability of gene usage that has led to the discovery of a significant number of genes with restricted expression, and that may thus be therapeutically useful. The ESTs also reveal novel nonsynonymous germline variants (although the one-pass nature of the data necessitates careful validation) and many alternatively spliced transcripts. Although widely exploited by the scientific community, vindicating our totally open source policy, the EST data generated still provide extensive information that remains to be systematically explored, and that may further facilitate progress toward both the understanding and treatment of human cancers.

Chromosome Mapping↗

Pilot survey of expressed sequence tags (ESTs) from the asexual blood stages of Plasmodium vivax in human patients.

BACKGROUND: Plasmodium vivax is the most widely distributed human malaria, responsible for 70-80 million clinical cases each year and large socio-economical burdens for countries such as Brazil where it is the most prevalent species. Unfortunately, due to the impossibility of growing this parasite in continuous in vitro culture, research on P. vivax remains largely neglected. METHODS: A pilot survey of expressed sequence tags (ESTs) from the asexual blood stages of P. vivax was performed. To do so, 1,184 clones from a cDNA library constructed with parasites obtained from 10 different human patients in the Brazilian Amazon were sequenced. Sequences were automatedly processed to remove contaminants and low quality reads. A total of 806 sequences with an average length of 586 bp met such criteria and their clustering revealed 666 distinct events. The consensus sequence of each cluster and the unique sequences of the singlets were used in similarity searches against different databases that included P. vivax, Plasmodium falciparum, Plasmodium yoelii, Plasmodium knowlesi, Apicomplexa and the GenBank non-redundant database. An E-value of <10(-30) was used to define a significant database match. ESTs were manually assigned a gene ontology (GO) terminology RESULTS: A total of 769 ESTs could be assigned a putative identity based upon sequence similarity to known proteins in GenBank. Moreover, 292 ESTs were annotated and a GO terminology was assigned to 164 of them. CONCLUSION: These are the first ESTs reported for P. vivax and, as such, they represent a valuable resource to assist in the annotation of the P. vivax genome currently being sequenced. Moreover, since the GC-content of the P. vivax genome is strikingly different from that of P. falciparum, these ESTs will help in the validation of gene predictions for P. vivax and to create a gene index of this malaria parasite.

AT Rich Sequence↗

T cell epitope characterization in tandemly repetitive Trypanosoma cruzi B13 protein.

Proteins containing tandemly repetitive sequences are present in several immunodominant protein antigens in pathogenic protozoan parasites. The tandemly repetitive Trypanosoma cruzi B13 protein is recognized by IgG antibodies from 98% of Chagas' disease patients. Little is known about the molecular mechanisms that lead to the immunodominance of the repeated sequences, and there is limited information on T cell epitopes in such repetitive antigens. We finely characterized the T cell recognition of the tandemly repetitive, degenerate B13 protein by T cell lines, clones and PBMC from Chagas' disease cardiomyopathy (CCC), asymptomatic T. cruzi infected (ASY) and non-infected individuals (N). PBMC proliferative responses to recombinant B13 protein were restricted to individuals bearing HLA-DQA1*0501(DQ7), -DR1, and -DR2; B13 peptides bound to the same HLA molecules in binding assays. The HLA-DQ7-restricted minimal T cell epitope [FGQAAAG(D/E)KP] was identified with an overlapping combinatorial peptide library including all B13 sequence variants in T. cruzi Y strain B13 protein; the underlined small residues GQA were the major HLA contact residues. Among natural B13 15-mer variant peptides, molecular modeling showed that several variant positions were solvent (TCR)-exposed, and substitutions at exposed positions abolished recognition. While natural B13 variant peptide S15.9 seems to be the immunodominant epitope for Chagas' disease patients, S15.4 was preferentially recognized by CCC rather than ASY patients, which may be pathogenically relevant. This is the first thorough characterization of T cell epitopes of a tandemly repetitive protozoan antigen and may suggest a role for T cell help in the immunodominance of protozoan repetitive antigens.

Amino Acid Sequence↗