PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Screening of an annotated compound library for drug activity in a resistant myeloma cell line.

PURPOSE: Resistance to anticancer drugs is a major problem in chemotherapy. In order to identify drugs with selective cytotoxic activity in drug-resistant cancer cells, the annotated compound library LOPAC1280, containing compounds from 56 pharmacological classes, was screened in the myeloma cell line RPMI 8226 and its doxorubicin-resistant subline 8226/Dox40. METHODS: Cell survival was measured by the Fluorometric Microculture Cytotoxicity Assay. RESULTS: Selective cytotoxic activity in 8226/Dox40 was obtained for 33 compounds, with the most pronounced difference observed for the glucocorticoids. A microarray analysis of the cells showed a difference in mRNA-expression for the glucocorticoid receptor suggesting potential mechanisms for the difference in glucocorticoid sensitivity. In the presence of the glucocorticoid-receptor antagonist RU486, the sensitivity to the glucocorticoids was reduced and a similar effect level in RPMI 8226 and 8226/Dox40 was achieved. CONCLUSION: In conclusion, screening of mechanistically annotated compounds on drug-resistant cancer cells can identify compounds with selective activity and provide a basis for the development of novel treatments of drug-resistant malignancies.

Antineoplastic Agents↗

Analysis of bovine mammary gland EST and functional annotation of the Bos taurus gene index.

Functional genomic studies of the mammary gland require an appropriate collection of cDNA sequences to assess gene expression patterns from the different developmental and operational states of underlying cell types. To better capture the range of gene expression, a normalized cDNA library was constructed from pooled bovine mammary tissues, and 23,202 expressed sequence tags (EST) were produced and deposited into GenBank. Assembly of these EST with sequences in the Bos taurus Gene Index (BtGI) helped to form 5751 of the current 23,883 tentative consensus (TC) sequences. The majority (87%) of these 5751 assemblies contained only one to three mammary-derived EST. In contrast, 18% of the mammary EST assembled with TC sequences corresponding to 12 genes. These results suggest library normalization was only partially effective, because the reduction in EST for genes abundantly transcribed during lactation could be attributed to pooling. For better assessment of novel content in the mammary library and to add to existing annotation of all bovine sequence elements, gene ontology assignments, and comparative sequence analyses against human genome sequence, human and rodent gene indices, and an index of orthologous alignments of genes across eukaryotes (TOGA) were performed, and results were added to existing BtGI annotation. Over 35,000 of the bovine elements significantly matched human genome sequence, and the positions of some alignments (3%) were unique relative to those using human expressed sequences. Because 3445 TC sequences had no significant match with any data set, mammary-derived cDNA clones representing 23 of these elements were analyzed further for expression and novelty. Only one clone met criteria suggesting the corresponding gene was a divergent ortholog or expressed sequence unique to cattle. These results demonstrate that bovine sequence expression data serve as a resource for characterizing mammalian transcriptomes and identifying those genes potentially unique to ruminants.

Animals↗

Updating of transposable element annotations from large wheat genomic sequences reveals diverse activities and gene associations.

Triticeae species (including wheat, barley and rye) have huge and complex genomes due to polyploidization and a high content of transposable elements (TEs). TEs are known to play a major role in the structure and evolutionary dynamics of Triticeae genomes. During the last 5 years, substantial stretches of contiguous genomic sequence from various species of Triticeae have been generated, making it necessary to update and standardize TE annotations and nomenclature. In this study we propose standard procedures for these tasks, based on structure, nucleic acid and protein sequence homologies. We report statistical analyses of TE composition and distribution in large blocks of genomic sequences from wheat and barley. Altogether, 3.8 Mb of wheat sequence available in the databases was analyzed or re-analyzed, and compared with 1.3 Mb of re-annotated genomic sequences from barley. The wheat sequences were relatively gene-rich (one gene per 23.9 kb), although wheat gene-derived sequences represented only 7.8% (159 elements) of the total, while the remainder mainly comprised coding sequences found in TEs (54.7%, 751 elements). Class I elements [mainly long terminal repeat (LTR) retrotransposons] accounted for the major proportion of TEs, in terms of sequence length as well as element number (83.6% and 498, respectively). In addition, we show that the gene-rich sequences of wheat genome A seem to have a higher TE content than those of genomes B and D, or of barley gene-rich sequences. Moreover, among the various TE groups, MITEs were most often associated with genes: 43.1% of MITEs fell into this category. Finally, the TRIM and copia elements were shown to be the most active TEs in the wheat genome. The implications of these results for the evolution of diploid and polyploid wheat species are discussed.

DNA Transposable Elements↗

Frequency, type, distribution and annotation of simple sequence repeats in Rosaceae ESTs.

Genomic resources for peach, a model species for Rosaceae, are being developed to accelerate gene discovery in other Rosaceae species by comparative mapping. Simple sequence repeats (SSRs) are an important tool for comparative mapping because of their high polymorphism and transportability. To accelerate the development of SSR markers, we analyzed publicly available Rosaceae expressed sequence tags (ESTs) for SSRs. A total of 17,284 ESTs from almond, peach and rose were assembled into putatively non-redundant EST sets. For comparison, 179,099 ESTs from Arabidopsis were also used in the analysis. About 4% of the assembled ESTs contained SSRs in Rosaceae, which was higher than the 2.4% found in Arabidopsis. About half of the SSRs were found in the putative UTR, and the estimated average distance between SSRs in the UTR was 5.5 kb in rose, 5.1 kb in almond, 7 kb in peach and 13 kb in Arabidopsis. In the putative coding region, the estimated average distance was two to four times longer than in the UTR. Rosaceae ESTs containing SSRs were functionally annotated using the GenBank nr database and further classified using the gene ontology terms associated with the matching sequences in the SwissProt database. The detailed data including the sequences and annotation results are available from http://www.genome.clemson.edu/gdr/rosaceaessr/.

Arabidopsis↗

Descartes' fly: the geometry of genomic annotation.

The completion of the Drosophila melanogaster genome marks another significant milestone in the growth of sequence information. But it also contributes to the ever-widening gap between sequence information and biological knowledge. One important approach to reducing this gap is theoretical inference through computational technologies. Many computer programs have been designed to annotate genomic sequence information with biologically relevant information. Here, I suggest that all of these methods have a common structure in which the sequence fragments are "coordinated" by some method of description such as Hidden Markov models. The key to the algorithms lies in constructing the most efficient set of coordinates that allow extrapolation and interpolation from existing knowledge. Efficient extrapolation and interpolation are produced if the sequence fragments acquire a natural geometrical structure in the coordinated description. Finding such a coordinate frame is an inductive problem with no algorithmic solution. The greater part of the problem of genomic annotation lies in biological modeling of the data rather than in algorithmic improvements.

Animals↗

An annotation update via cDNA sequence analysis and comprehensive profiling of developmental, hormonal or environmental responsiveness of the Arabidopsis AP2/EREBP transcription factor gene family.

AP2/EREBP transcription factors (TFs) play functionally important roles in plant growth and development, especially in hormonal regulation and in response to environmental stress. Here we reported verification and correction of annotation through an exhaustive cDNA cloning and sequence analysis performed on 145 of 147 gene family members. A RACE analysis performed on genes with potential in-frame up-stream ATG codon resulted in identification of At2g28520 as an authentic AP2/EREBP member and corrected ORF annotations for three other members. A further phylogenetic analysis of this updated and likely complete family divided it into three major subfamilies. The expression patterns of the AP2/EREBP family members among the 11 organ or tissue types were examined using an oligo microarray and their hormonal and environmental responsiveness were further characterized using cDNA custom macroarrays. These detailed expression profile results provide strong support for a role for AP2/EREBP family members in development and in response to environmental stimuli, and a foundation for future functional analysis of this gene family.

Algorithms↗

Exploring the chemogenomic knowledge space with annotated chemical libraries.

The recent human genome initiatives have led to the discovery of a multitude of genes that are potentially associated with various pathologic conditions and, thus, have opened new horizons in drug discovery. Simultaneously, annotated chemical libraries have emerged as information-rich databases to integrate biological and chemical data. They can be useful for the discovery of new pharmaceutical leads, the validation of new biotargets and the determination of the structural basis of ligand selectivity within target families. Annotated libraries provide a strong information basis for computational design of target-directed combinatorial libraries, which are a key component of modern drug discovery. Today, the rational design of chemical libraries enhanced with chemogenomics data is a new area of progressive research.

Combinatorial Chemistry Techniques↗

Biological mechanism profiling using an annotated compound library.

We present a method for testing many biological mechanisms in cellular assays using an annotated library of 2036 small organic molecules. This annotated compound library represents a large-scale collection of compounds with diverse, experimentally confirmed biological mechanisms and effects. We found that this chemical library is (1) more structurally diverse than conventional, commercially available libraries, (2) enriched in active compounds in a tumor cell viability assay, and (3) capable of generating hypotheses regarding biological mechanisms underlying cellular processes. We elucidated biological mechanisms relevant to the antiproliferative activity of 85 compounds from this library that were selected using a high-throughput cell viability screen. We developed a novel automated scoring system for identifying statistically enriched mechanisms among such a subset of compounds. This scoring system can identify both previously known and potentially novel antiproliferative mechanisms.

Cell Line, Tumor↗

GeneLook: a novel ab initio gene identification system suitable for automated annotation of prokaryotic sequences.

With the rapid increases in the amounts of sequence data for prokaryotic genomes, it has become important to develop systems for automated and accurate genome annotation. We present herein a novel ab initio gene identification system, GeneLook, that predicts protein-coding open reading frames (ORFs) with high sensitivity and specificity with no prior knowledge of the sequence composition. The system predicts protein-coding ORFs in two stages, seed ORF selection and main prediction. In the selection of reliable seed ORFs containing at least 200 codons, GeneLook predicts translation start sites and operon structures through searches for ribosome-binding sites and a novel operon prediction algorithm. The codon and nucleotide frequencies of seed ORFs are then used to determine values for two new coding-potential parameters for identification of protein-coding ORFs of at least 34 codons and for another parameter that improves the prediction accuracy for GC-rich genomes. In the main prediction, GeneLook uses these parameters to identify the most likely genes of a given minimal length. We assessed the performance of GeneLook with two indices, sensitivity and specificity that are defined as true positives (TP)/(TP+false negatives) and TP/(TP+false positives), respectively. This system predicted protein-coding ORFs for Escherichia coli and Bacillus subtilis with sensitivities of 96.5% and 96.2%, respectively, and specificities of 96.9% and 96.1%, respectively. The system also identified 94.1% of annotated genes of the Pseudomonas aeruginosa genome, which is GC-rich, with high specificity (97.2%). Furthermore, GeneLook identified protein-coding ORFs with high accuracy from a wide variety of prokaryotic genomes.

Bacillus subtilis↗

Assisting medical annotation in Swiss-Prot using statistical classifiers.

Bio-medical knowledge bases are valuable resources for the research community. Original scientific publications are the main source used to annotate them. Medical annotation in Swiss-Prot is specifically targeted at finding and extracting data about human genetic diseases and polymorphisms. Curators have to scan through hundreds of publications to select the relevant ones. This workload can be greatly reduced by using bio-text mining techniques. Using a combination of natural language processing (NLP) techniques and statistical classifiers, we achieve recall points of up to 84% on the potentially interesting documents and a precision of more than 96% in detecting irrelevant documents. Careful analysis of the document pre-processing chain allows us to measure the impact of some steps on the overall result, as well as test different classifier configurations. The best combination was used to create a prototype of a search and classification tool that is currently tested by the database curators.

Databases, Protein↗

Distributed modules for text annotation and IE applied to the biomedical domain.

Biological databases contain facts from scientific literature that have been curated by hand to ensure high quality. Curation is time-consuming and can be supported by information extraction methods. We present a server software infrastructure which allows to easily plug in modules to identify biologically interesting pieces of text to be then presented in a web interface to the curator. There are modules which identify UniProt, UMLS and GO terminology, gene and protein names, mutations and protein-protein interactions. UniProt, UMLS and GO concepts are automatically linked to the original source. The module for mutations is based on syntax patterns and the one for protein-protein interactions relies on chunk parsing. All modules work as separate servers possibly distributed on different machines and can be combined into processing pipelines as necessary. Communication is based on XML annotated text streams, each server processing the XML elements it is designed for, and possibly adding more information in the form of XML annotation. The server and the underlying software are available to the public.

Abstracting and Indexing↗

Protozoan genomes: gene identification and annotation.

The draft sequence of several complete protozoan genomes is now available and genome projects are ongoing for a number of other species. Different strategies are being implemented to identify and annotate protein coding and RNA genes in these genomes, as well as study their genomic architecture. Since the genomes vary greatly in size, GC-content, nucleotide composition, and degree of repetitiveness, genome structure is often a factor in choosing the methodology utilised for annotation. In addition, the approach taken is dictated, to a greater or lesser extent, by the particular reasons for carrying out genome-wide analyses and the level of funding available for projects. Nevertheless, these projects have provided a plethora of material that will aid in understanding the biology and evolution of these parasites, as well as identifying new targets that can be used to design urgently required drug treatments for the diseases they cause.

Algorithms↗

A Bayesian network coding scheme for annotating biomedical information presented to genetic counseling clients.

We developed a Bayesian network coding scheme for annotating biomedical content in layperson-oriented clinical genetics documents. The coding scheme supports the representation of probabilistic and causal relationships among concepts in this domain, at a high enough level of abstraction to capture commonalities among genetic processes and their relationship to health. We are using the coding scheme to annotate a corpus of genetic counseling patient letters as part of the requirements analysis and knowledge acquisition phase of a natural language generation project. This paper describes the coding scheme and presents an evaluation of intercoder reliability for its tag set. In addition to giving examples of use of the coding scheme for analysis of discourse and linguistic features in this genre, we suggest other uses for it in analysis of layperson-oriented text and dialogue in medical communication.

Artificial Intelligence↗

MitoMorphy: an alignment and annotation tool for human mitochondrial DNA polymorphisms.

MitoMorphy uses a number of publicly available human mitochondrial DNA (mtDNA) sequences from different ethnic groups to compare and annotate the associated polymorphic data. It provides an integrated display of mtDNA sequence comparison, sequence variation, and annotation for 695 different mtDNA sequences from many different ethnic groups around the world.

Journal Article↗

Challenges in real-life emotion annotation and machine learning based detection.

Since the early studies of human behavior, emotion has attracted the interest of researchers in many disciplines of Neurosciences and Psychology. More recently, it is a growing field of research in computer science and machine learning. We are exploring how the expression of emotion is perceived by listeners and how to represent and automatically detect a subject's emotional state in speech. In contrast with most previous studies, conducted on artificial data with archetypal emotions, this paper addresses some of the challenges faced when studying real-life non-basic emotions. We present a new annotation scheme allowing the annotation of emotion mixtures. Our studies of real-life spoken dialogs from two call center services reveal the presence of many blended emotions, dependent on the dialog context. Several classification methods (SVM, decision trees) are compared to identify relevant emotional states from prosodic, disfluency and lexical cues extracted from the real-life spoken human-human interactions.

Artificial Intelligence↗

Genome-scale, biochemical annotation method based on the wheat germ cell-free protein synthesis system.

Since the complete genomic DNA sequencing of various species, attention has turned to the structural properties, and functional characteristics of proteins. Current cell-free protein expression systems from eukaryotes are capable of synthesizing proteins with high speed and accuracy; however, the yields are low due to their instability over time. This report reviews the high-throughput, genome-scale biochemical annotation method based on the cell-free system prepared from wheat embryos. We first briefly reviewed our highly efficient and robust wheat germ cell-free protein synthesis system, and then showed an application of the system for materialization and characterization of genetic information taking a cDNA library of protein kinase from Arabidopsis thaliana as an example. The procedure consists of: (1) fusion of the gene-of-interest to a purification-tag, amplified by the split-primer PCR method; (2) transcription and purification of mRNA; (3) cell-free protein synthesis in the bilayer system using 96-well titer plate; (4) affinity purification and activity measurement. We took 439 cDNAs encoding kinases among 1064 genes annotated so far, and they were translated in parallel into protein. Subsequent assay revealed 207 products having autophosphorylation activity. Furthermore, seven proteins out of 26 calcium-dependent protein kinase genes tested did phosphorylate a synthetic peptide substrate in the presence of calcium ion, demonstrating that the translation products, retained their substrate specificity. The information on biochemical function of gene products accumulated should revolutionize our understanding of biology and fundamentally alter the practice of medicine and influence other industries as well.

Arabidopsis↗

Helicopter emergency medical services transport outcomes literature: annotated review of articles published 2000-2003.

Helicopter emergency medical services (HEMS) and its possible association with outcomes improvement continues to be a subject of debate. As is the case with other scientific endeavors, debate over HEMS usefulness should be framed around an evidence-based assessment of the relevant literature. In an effort to facilitate the academic pursuit of assessment of HEMS utility, in late 2000 the National Association of EMS Physicians' Air Medical Committee prepared annotated bibliographies of the HEMS-related outcomes literature. As a result of that work, two review articles-one covering HEMS use in nontrauma and the other in trauma-published in 2002 in Prehospital Emergency Care surveyed HEMS outcomes-related literature published between 1980 and mid-2000. Given the broad interest in the earlier reviews, and the increasing rate of publication of HEMS studies, the current project was executed with the intent of updating the annotated HEMS outcomes-related bibliography, covering a three-year time interval (through 2003) since the prior reviews.

Air Ambulances↗

Metalloproteomics: high-throughput structural and functional annotation of proteins in structural genomics.

A high-throughput method for measuring transition metal content based on quantitation of X-ray fluorescence signals was used to analyze 654 proteins selected as targets by the New York Structural GenomiX Research Consortium. Over 10% showed the presence of transition metal atoms in stoichiometric amounts; these totals as well as the abundance distribution are similar to those of the Protein Data Bank. Bioinformatics analysis of the identified metalloproteins in most cases supported the metalloprotein annotation; identification of the conserved metal binding motif was also shown to be useful in verifying structural models of the proteins. Metalloproteomics provides a rapid structural and functional annotation for these sequences and is shown to be approximately 95% accurate in predicting the presence or absence of stoichiometric metal content. The project's goal is to assay at least 1 member from each Pfam family; approximately 500 Pfam families have been characterized with respect to transition metal content so far.

Binding Sites↗