PubMed Health⌕ Search

Biomedical subjects

M A Andrade

Publications and source records attributed to M A Andrade.

At least 19 recordsLinked to original sources

Comparison of ARM and HEAT protein repeats.

ARM and HEAT motifs are tandemly repeated sequences of approximately 50 amino acid residues that occur in a wide variety of eukaryotic proteins. An exhaustive search of sequence databases detected new family members and revealed that at least 1 in 500 eukaryotic protein sequences contain such repeats. It also rendered the similarity between ARM and HEAT repeats, believed to be evolutionarily related, readily apparent. All the proteins identified in the database searches could be clustered by sequence similarity into four groups: canonical ARM-repeat proteins and three groups of the more divergent HEAT-repeat proteins. This allowed us to build improved sequence profiles for the automatic detection of repeat motifs. Inspection of these profiles indicated that the individual repeat motifs of all four classes share a common set of seven highly conserved hydrophobic residues, which in proteins of known three-dimensional structure are buried within or between repeats. However, the motifs differ at several specific residue positions, suggesting important structural or functional differences among the classes. Our results illustrate that ARM and HEAT-repeat proteins, while having a common phylogenetic origin, have since diverged significantly. We discuss evolutionary scenarios that could account for the great diversity of repeats observed.

Amino Acid Motifs↗

Simulation of plasticity in the adult visual cortex.

Retinal plasticity has been shown in the adult visual nervous system in mammals. Following a retinal lesion (scotoma) there is a reorganization of the cortical receptive field distribution: cortical neurons selective to visual stimuli in the area of the visual field corresponding to the retinal lesion, become selective to other parts of the visual field. In this work, we study this effect with a self-organizing neural network. In a first stage, the network reaches a pattern of connectivity that represents normal development of neuronal selectivity. The scotoma is simulated by perturbing accordingly the properties of a region of the input layer representing the retina. The system evolves to a new receptive field distribution mainly by means of the reorganization of the intra cortical connectivity. No major change of the geniculo cortical connectivity is detected. This may explain the surprisingly short time scale of the event.

Age Factors↗

XplorMed: a tool for exploring MEDLINE abstracts.

The most frequent access to the MEDLINE database of scientific abstracts is by keyword search. However, this is often not sufficient because although the user might find all the useful abstracts, these are buried in hundreds that are irrelevant. The exploratory tool XplorMed has been developed to analyse the result of any MEDLINE query. It suggests main groups of related topics and documents, sparing the user the need of reading all abstracts.

MEDLINE↗

A combination of the F-box motif and kelch repeats defines a large Arabidopsis family of F-box proteins.

In the sequences released by the Arabidopsis Genome Initiative (AGI), we have discovered a new large gene family (48 genes as of July 2000). A detailed computational and biochemical analysis of the predicted gene products reveals a novel family of plant F-box proteins, where the amino (N)-terminal F-box motif is followed by four kelch repeats and a characteristic carboxy-terminal domain. F-box proteins are an expanding family of eukaryotic proteins, which have been shown in some cases to be critical for the controlled degradation of cellular regulatory proteins via the ubiquitin pathway. The F-box motif of the At5g48990 gene product, a member of the family, was shown to be functionally active by its ability to mediate the in vitro interaction between At5g48990 and ASK1 proteins. F-box proteins specifically recruit the targets to be ubiquitinated, mainly through protein-protein interaction modules such as WD-40 domains or leucine-rich repeats (LRRs). The kelch repeats of the family described here form a potential protein-protein interaction domain, as molecular modelling of the kelch repeats according to the galactose oxidase crystal structure (the only solved structure containing kelch repeats) predicts a beta-propeller. The identification of this family of F-box proteins greatly expands the field of plant F-box proteins and suggests that controlled degradation of cellular proteins via the ubiquitin pathway could play a critical role in multiple plant cellular processes.

Amino Acid Motifs↗

Genome sequences and great expectations.

To assess how automatic function assignment will contribute to genome annotation in the next five years, we have performed an analysis of 31 available genome sequences. An emerging pattern is that function can be predicted for almost two-thirds of the 73,500 genes that were analyzed. Despite progress in computational biology, there will always be a great need for large-scale experimental determination of protein function.

Animals↗

Re-annotating the Mycoplasma pneumoniae genome sequence: adding value, function and reading frames.

Four years after the original sequence submission, we have re-annotated the genome of Mycoplasma pneumoniae to incorporate novel data. The total number of ORFss has been increased from 677 to 688 (10 new proteins were predicted in intergenic regions, two further were newly identified by mass spectrometry and one protein ORF was dismissed) and the number of RNAs from 39 to 42 genes. For 19 of the now 35 tRNAs and for six other functional RNAs the exact genome positions were re-annotated and two new tRNA(Leu) and a small 200 nt RNA were identified. Sixteen protein reading frames were extended and eight shortened. For each ORF a consistent annotation vocabulary has been introduced. Annotation reasoning, annotation categories and comparisons to other published data on M.pneumoniae functional assignments are given. Experimental evidence includes 2-dimensional gel electrophoresis in combination with mass spectrometry as well as gene expression data from this study. Compared to the original annotation, we increased the number of proteins with predicted functional features from 349 to 458. The increase includes 36 new predictions and 73 protein assignments confirmed by the published literature. Furthermore, there are 23 reductions and 30 additions with respect to the previous annotation. mRNA expression data support transcription of 184 of the functionally unassigned reading frames.

Amino Acid Sequence↗

Automated extraction of information in molecular biology.

We review data mining techniques in molecular biology, specifically those that extract information from the scientific literature itself. As more of the biological literature is published electronically, there is an opportunity, and even a need, to automatically summarize the literature in a customized way, for example by associating keywords to a topic. These keywords can be extracted from relevant publications. The process of keyword extraction can be automated and optimized to keep literature pointers automatically up-to-date or to filter relevant information from the literature. To illustrate these points, OMIM (Online Mendelian Inheritance in Man), a database of human inherited diseases, was linked to the literature and keywords were derived that covered distinct aspects such as genetic information on the one hand and disease-specific protein and phenotypic information on the other. They were used to extract information that is helpful for keeping entries about disease up-to-date.

Databases, Factual↗

Homology-based method for identification of protein repeats using statistical significance estimates.

Short protein repeats, frequently with a length between 20 and 40 residues, represent a significant fraction of known proteins. Many repeats appear to possess high amino acid substitution rates and thus recognition of repeat homologues is highly problematic. Even if the presence of a certain repeat family is known, the exact locations and the number of repetitive units often cannot be determined using current methods. We have devised an iterative algorithm based on optimal and sub-optimal score distributions from profile analysis that estimates the significance of all repeats that are detected in a single sequence. This procedure allows the identification of homologues at alignment scores lower than the highest optimal alignment score for non-homologous sequences. The method has been used to investigate the occurrence of eleven families of repeats in Saccharomyces cerevisiae, Caenorhabditis elegans and Homo sapiens accounting for 1055, 2205 and 2320 repeats, respectively. For these examples, the method is both more sensitive and more selective than conventional homology search procedures. The method allowed the detection in the SwissProt database of more than 2000 previously unrecognised repeats belonging to the 11 families. In addition, the method was used to merge several repeat families that previously were supposed to be distinct, indicating common phylogenetic origins for these families.

Algorithms↗

NAIL-Network Analysis Interface for Linking HMMER results.

SUMMARY: Network Analysis Interface for Linking HMMER results (NAIL) is a web-based tool for the analysis of results from a HMMER protein database-search. NAIL facilitates the selection of protein hits and the creation of an alignment, which can be used for a new sequence similarity search.

Databases, Factual↗

Functional classes in the three domains of life.

The evolutionary divergence among the three major domains of life can now be addressed through the first set of complete genomes from representative species. These model species from the three domains of life, Haemophilus influenzae for Bacteria, Saccharomyces cerevisiae for Eukarya, and Methanococcus jannaschii for Archaea, provide the basis for a universal functional classification and analysis. We have chosen 13 functional classes and three superclasses (ENERGY, COMMUNICATION and INFORMATION) as global descriptors of protein function. Compositional comparison of the three complete genomes reveals that functional classes are ubiquitous yet diverse in the three domains of life. Proteins related with ENERGY processes are generally represented in all three domains, while those related with COMMUNICATION represent the most distinctive functional feature of each single domain. Finally, functions related with INFORMATION processing (translation, transcription, and replication) show a complex behaviour. In Archaea, proteins in this superclass are related with proteins in either Eukarya or Bacteria, as recognized previously. The distribution of functional classes in the three domains accurately reflects the principal characteristics of cellular life forms.

Archaeal Proteins↗

High serum leptin levels in children with type 1 diabetes mellitus: contribution of age, BMI, pubertal development and metabolic status.

OBJECTIVE: Children with diabetes mellitus are prone to develop obesity and to experience a delay in onset of the pubertal process. In order to understand the role of leptin in these abnormalities, serum leptin levels were analysed in children with type 1 diabetes mellitus. SUBJECTS: Twenty diabetic girls, 23 diabetic boys and 66 healthy children (selected from a reference population of 706 normal children), age-, sex- and BMI-matched with diabetic patients, were studied. MEASURMENTS: Standing height, weight and BMI were determined in each child. Serum testosterone, oestradiol and leptin were measured by specific radioimmunoassays, and HBA1c by high performance liquid chromatography. RESULTS: Both diabetic girls and boys showed higher leptin levels than the normative healthy population and a group of age-, sex- and BMI-matched normal children. In an age-related analysis, leptin levels in diabetic girls rose from 7.4 +/- 1.2 and 8.1 +/- 2.1 microg/l for the 5-7.99 and 8-10.99 year groups, to 12.6 +/- 2.4 microg/l for the 11-13.99 year group, and to 15.6 +/- 4.0 microg/l in the 14-15.99 year group in parallel with body weight. Leptin concentrations were parallel but higher (P < 0.05) than those of healthy girls. Diabetic boys had lower leptin levels than girls and, in contrast with normal boys, did not show a drop after the 10-year period. Leptin levels were 4.9 +/- 2.2, 3.9 +/- 0.2, 5.5 +/- 0.6 and 5.1 +/- 0.9 microg/l for the 5-7.99, 8-10. 99, 11-13.99 and 14-15.99 year groups, respectively. When divided by pubertal stage, leptin levels in the prepuberty stage of diabetic girls (8.6 +/- 1.0 microg/l) were higher (P < 0.05) than those in the controls (4.1 +/- 0.4 microg/l). In overt puberty girls, leptin was higher (P < 0.05) for diabetic (15.9 +/- 2.9 microg/l) than for healthy girls (9.2 +/- 1.1 microg/l). In prepubertal boys, differences were observed in leptin levels (4.9 +/- 0.5 microg/l for diabetic boys and 3.4 +/- 0.6 microg/l for healthy boys). In the overt puberty stage, diabetic boys showed higher (P < 0.05) levels of leptin (5.2 +/- 0.7 microg/l) than the healthy matched controls (2.1 +/- 0.2 microg/l). A multiple step regression analysis in the diabetic children revealed no associations between leptin and other relevant variables such as glycosylated haemoglobin, daily insulin dose, or years of suffering from the disease. CONCLUSION: Serum leptin levels were higher in diabetic than in healthy children. These differences were not attributable to age, adiposity or stage of pubertal development, and were probably conditioned by the metabolic perturbation intrinsic to the diabetic state, or the chronic hyperinsulinemia.

Adolescent↗

Automated genome sequence analysis and annotation.

MOTIVATION: Large-scale genome projects generate a rapidly increasing number of sequences, most of them biochemically uncharacterized. Research in bioinformatics contributes to the development of methods for the computational characterization of these sequences. However, the installation and application of these methods require experience and are time consuming. RESULTS: We present here an automatic system for preliminary functional annotation of protein sequences that has been applied to the analysis of sets of sequences from complete genomes, both to refine overall performance and to make new discoveries comparable to those made by human experts. The GeneQuiz system includes a Web-based browser that allows examination of the evidence leading to an automatic annotation and offers additional information, views of the results, and links to biological databases that complement the automatic analysis. System structure and operating principles concerning the use of multiple sequence databases, underlying sequence analysis tools, lexical analyses of database annotations and decision criteria for functional assignments are detailed. The system makes automatic quality assessments of results based on prior experience with the underlying sequence analysis tools; overall error rates in functional assignment are estimated at 2.5-5% for cases annotated with highest reliability ('clear' cases). Sources of over-interpretation of results are discussed with proposals for improvement. A conservative definition for reporting 'new findings' that takes account of database maturity is presented along with examples of possible kinds of discoveries (new function, family and superfamily) made by the system. System performance in relation to sequence database coverage, database dynamics and database search methods is analysed, demonstrating the inherent advantages of an integrated automatic approach using multiple databases and search methods applied in an objective and repeatable manner. AVAILABILITY: The GeneQuiz system is publicly available for analysis of protein sequences through a Web server at http://www.sander.ebi.ac. uk/gqsrv/submit

Amino Acid Sequence↗

Immunohistochemical detection of parathyroid hormone-related protein in a cutaneous squamous cell carcinoma causing humoral hypercalcemia of malignancy.

Humoral hypercalcemia of malignancy is a cancer-related hypercalcemia caused by production of humoral factors by malignant cells in patients without bone metastases. Squamous cell carcinomas are the tumors most frequently associated with humoral hypercalcemia of malignancy, and parathyroid hormone-related protein is the main humoral factor implicated. In spite of the fact that normal keratinocytes produce parathyroid hormone-related protein, it is highly unusual for patients with squamous cell carcinomas of the skin to present with humoral hypercalcemia of malignancy. We present a well-documented case of cutaneous squamous cell carcinoma complicated by hypercalcemia in a patient with high levels of plasma parathyroid hormone-related protein and immunohistochemical evidence of high parathyroid hormone-related protein production by the tumoral cells.

Aged↗

Position-specific annotation of protein function based on multiple homologs.

I present in this work an algorithm for deriving protein functional annotations which are position-specific. The input is based on the results of a sequence similarity search of the query sequence against a sequence database. Strings of words are extracted from the descriptions of the proteins, and the correlation between proteins having the same descriptors and the amino acid conservation is used to compute a score that indicates which descriptor is likely to describe better the function of each particular residue. Analysis of the score curves and comparison of different functions allows an easy detection of parts of the sequence associated to different function. Different levels of functional specificity can be compared, allowing to choose the one that suits better the function of the protein. Immediate applications of this algorithm are, support for (automated) methods of protein functional annotation, and database coherence check.

Algorithms↗

Automatic extraction of biological information from scientific text: protein-protein interactions.

We describe the basic design of a system for automatic detection of protein-protein interactions extracted from scientific abstracts. By restricting the problem domain and imposing a number of strong assumptions which include pre-specified protein names and a limited set of verbs that represent actions, we show that it is possible to perform accurate information extraction. The performance of the system is evaluated with different cases of real-world interaction networks, including the Drosophila cell cycle control. The results obtained computationally are in good agreement with current biological knowledge and demonstrate the feasibility of developing a fully automated system able to describe networks of protein interactions with sufficient accuracy.

Algorithms↗

Updated catalogue of homologues to human disease-related proteins in the yeast genome.

The recent availability of the full Saccharomyces cerevisiae genome offers a perfect opportunity for revising the number of homologues to human disease-related proteins. We carried out automatic analysis of the complete S. cerevisiae genome and of the set of human disease-related proteins as identified in the SwissProt sequence data base. We identified 285 yeast proteins similar to 155 human disease-related proteins, including 239 possible cases of human-yeast direct functional equivalence (orthology). Of these, 40 cases are suggested as new, previously undiscovered relationships. Four of them are particularly interesting, since the yeast sequence is the most phylogenetically distant member of the protein family, including proteins related to diseases such as phenylketonuria, lupus erythematosus, Norum and fish eye disease and Wiskott-Aldrich syndrome.

Amino Acid Sequence↗