PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Cell type annotation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Identification of a major protein on the cytosolic face of caveolae.

Cav-p60, a specific and ubiquitous caveolar protein, was immunoprecipitated from solubilized rat adipocyte plasma membranes and identified as similar to a GeneBank entry annotated mouse polymerase transcript release factor (PTRF) by MALDI-TOF and MS-MS of major fragments. Cloning and virtual translation of the corresponding rat adipocyte cDNA sequence revealed 98.7% identity with mouse PTRF. In vitro translation of this sequence produced a protein, which was recognized by antibodies to both cav-p60 and PTRF. EM gold labeling studies showed that a rabbit antiserum against murine PTRF immunolabeled caveolae specifically in adipocytes from both mouse and rat. In view of the reported function of the protein, which is exerted in the cell nucleus, its subcellular localization was investigated. We found that the protein could be purified by differential solubilization of a plasma membrane fraction followed by SDS-PAGE, and that the protein was as abundant as caveolin in this fraction. We were unable to detect the protein in cell nuclei by subcellular fractionation or fluorescence microscopy. The results show that in a large number of cell types, PTRF is essentially located to caveolae, and that each caveola harbors many copies of the protein. Consequently, we suggest the name Cavin for this protein.

Adipocytes↗

Formin homology 2 domains occur in multiple contexts in angiosperms.

BACKGROUND: Involvement of conservative molecular modules and cellular mechanisms in the widely diversified processes of eukaryotic cell morphogenesis leads to the intriguing question: how do similar proteins contribute to dissimilar morphogenetic outputs. Formins (FH2 proteins) play a central part in the control of actin organization and dynamics, providing a good example of evolutionarily versatile use of a conserved protein domain in the context of a variety of lineage-specific structural and signalling interactions. RESULTS: In order to identify possible plant-specific sequence features within the FH2 protein family, we performed a detailed analysis of angiosperm formin-related sequences available in public databases, with particular focus on the complete Arabidopsis genome and the nearly finished rice genome sequence. This has led to revision of the current annotation of half of the 22 Arabidopsis formin-related genes. Comparative analysis of the two plant genomes revealed a good conservation of the previously described two subfamilies of plant formins (Class I and Class II), as well as several subfamilies within them that appear to predate the separation of monocot and dicot plants. Moreover, a number of plant Class II formins share an additional conserved domain, related to the protein phosphatase/tensin/auxilin fold. However, considerable inter-species variability sets limits to generalization of any functional conclusions reached on a single species such as Arabidopsis. CONCLUSIONS: The plant-specific domain context of the conserved FH2 domain, as well as plant-specific features of the domain itself, may reflect distinct functional requirements in plant cells. The variability of formin structures found in plants far exceeds that known from both fungi and metazoans, suggesting a possible contribution of FH2 proteins in the evolution of the plant type of multicellularity.

Actins↗

Ovarian expression and function of neuropeptide systems in teleosts and anurans.

The hypothalamic-pituitary-gonadal axis regulates reproduction, sexual maturation, and spawning behaviours. Its evolutionary origins trace back to primitive jawless fish and has been well characterized in teleosts. Recent advances in multi-species genome sequencing, annotation, and experimental approaches for identifying and characterizing key regulators have advanced understanding of neuroendocrine regulation in teleost reproduction, reshaping existing models. Early studies in amphibians established that steroids are critical regulators of final oocyte maturation. Subsequent work in anurans revealed complex interactions among theca cells, follicular cells, and oocytes, supporting a three-cell model in which oocytes contribute to their own steroidogenic environment, challenging the traditional two-cell view of ovarian steroidogenesis. In teleosts, however, direct evidence that oocytes support steroid precursor delivery to theca and follicular cells is limited, and whether a comparable three-cell model applies remains an open hypothesis. Across both taxa, the roles of locally produced neuropeptides in coordinating interactions among theca cells, follicular cells, and oocytes remain largely uncharacterized. Here, we provide a short review of the localization and potential autocrine/paracrine functions of neuropeptides in teleost and amphibian ovaries and discuss existing knowledge gaps. We identify opportunities to leverage detailed localization studies that map neuropeptides to specific ovarian cell types and developmental stages, and discuss how integrating traditional and emerging experimental approaches can advance comparative studies in ovarian endocrinology. This work will improve our understanding of reproductive regulation in fishes and frogs, with applications in captive breeding, aquaculture, and endocrine disruption research.

Autocrine↗

VisANT: an online visualization and analysis tool for biological interaction data.

BACKGROUND: New techniques for determining relationships between biomolecules of all types--genes, proteins, noncoding DNA, metabolites and small molecules--are now making a substantial contribution to the widely discussed explosion of facts about the cell. The data generated by these techniques promote a picture of the cell as an interconnected information network, with molecular components linked with one another in topologies that can encode and represent many features of cellular function. This networked view of biology brings the potential for systematic understanding of living molecular systems. RESULTS: We present VisANT, an application for integrating biomolecular interaction data into a cohesive, graphical interface. This software features a multi-tiered architecture for data flexibility, separating back-end modules for data retrieval from a front-end visualization and analysis package. VisANT is a freely available, open-source tool for researchers, and offers an online interface for a large range of published data sets on biomolecular interactions, including those entered by users. This system is integrated with standard databases for organized annotation, including GenBank, KEGG and SwissProt. VisANT is a Java-based, platform-independent tool suitable for a wide range of biological applications, including studies of pathways, gene regulation and systems biology. CONCLUSION: VisANT has been developed to provide interactive visual mining of biological interaction data sets. The new software provides a general tool for mining and visualizing such data in the context of sequence, pathway, structure, and associated annotations. Interaction and predicted association data can be combined, overlaid, manipulated and analyzed using a variety of built-in functions. VisANT is available at http://visant.bu.edu.

Animals↗

Identification of a novel gene (HSN2) causing hereditary sensory and autonomic neuropathy type II through the Study of Canadian Genetic Isolates.

Hereditary sensory and autonomic neuropathy (HSAN) type II is an autosomal recessive disorder characterized by impairment of pain, temperature, and touch sensation owing to reduction or absence of peripheral sensory neurons. We identified two large pedigrees segregating the disorder in an isolated population living in Newfoundland and performed a 5-cM genome scan. Linkage analysis identified a locus mapping to 12p13.33 with a maximum LOD score of 8.4. Haplotype sharing defined a candidate interval of 1.06 Mb containing all or part of seven annotated genes, sequencing of which failed to detect causative mutations. Comparative genomics revealed a conserved ORF corresponding to a novel gene in which we found three different truncating mutations among five families including patients from rural Quebec and Nova Scotia. This gene, termed "HSN2," consists of a single exon located within intron 8 of the PRKWNK1 gene and is transcribed from the same strand. The HSN2 protein may play a role in the development and/or maintenance of peripheral sensory neurons or their supporting Schwann cells.

Amino Acid Sequence↗

Helicobacter pylori flagellar hook-filament transition is controlled by a FliK functional homolog encoded by the gene HP0906.

Helicobacter pylori is a human gastric pathogen which is dependent on motility for infection. The H. pylori genome encodes a near-complete complement of flagellar proteins compared to model enteric bacteria. One of the few flagellar genes not annotated in H. pylori is that encoding FliK, a hook length control protein whose absence leads to a polyhook phenotype in Salmonella enterica. We investigated the role of the H. pylori gene HP0906 in flagellar biogenesis because of linkage to other flagellar genes, because of its transcriptional regulation pattern, and because of the properties of an ortholog in Campylobacter jejuni (N. Kamal and C. W. Penn, unpublished data). A nonpolar mutation of HP0906 in strain CCUG 17874 was generated by insertion of a chloramphenicol resistance marker. Cells of the mutant were almost completely nonmotile but produced sheathed, undulating polyhook structures at the cell pole. Expression of HP0906 in a Salmonella fliK mutant restored motility, confirming that HP0906 is the H. pylori fliK gene. Mutation of HP0906 caused a dramatic reduction in H. pylori flagellin protein production and a significant increase in production of the hook protein FlgE. The HP0906 mutant showed increased transcription of the flgE and flaB genes relative to the wild type, down-regulation of flaA transcription, and no significant change in transcription of the flagellar intermediate class genes flgM, fliD, and flhA. We conclude that the H. pylori HP0906 gene product is the hook length control protein FliK and that its function is required for turning off the sigma(54) regulon during progression of the flagellar gene expression cascade.

Amino Acid Sequence↗

Immunoinformatics--the new kid in town.

The astounding diversity of immune system components (e.g. immunoglobulins, lymphocyte receptors, or cytokines) together with the complexity of the regulatory pathways and network-type interactions makes im munology a combinatorial science. Currently available data represent only a tiny fraction of possible situations and data continues to accrue at an exponential rate. Computational analysis has therefore become an essential element of immunology research with a main role of immunoinformatics being the management and analysis of immunological data. More advanced analyses of the immune system using computational models typically involve conversion of an immunological question to a computational problem, followed by solving of the computational problem and translation of these results into biologically meaningful answers. Major immunoinformatics developments include immunological databases, sequence analysis, structure modelling, mathematical modelling of the immune system, simulation of laboratory experiments, statistical support for immunological experimentation and immunogenomics. In this paper we describe the status and challenges within these sub-fields. We foresee the emergence of immunomics not only as a collective endeavour by researchers to decipher the sequences of T cell receptors, immunoglobulins, and other immune receptors, but also to functionally annotate the capacity of the immune system to interact with the whole array of selfand non-self entities, including genome-to-genome interactions.

Allergy and Immunology↗

A general model of G protein-coupled receptor sequences and its application to detect remote homologs.

G protein-coupled receptors (GPCRs) constitute a large superfamily involved in various types of signal transduction pathways triggered by hormones, odorants, peptides, proteins, and other types of ligands. The superfamily is so diverse that many members lack sequence similarity, although they all span the cell membrane seven times with an extracellular N and a cytosolic C terminus. We analyzed a divergent set of GPCRs and found distinct loop length patterns and differences in amino acid composition between cytosolic loops, extracellular loops, and membrane regions. We configured GPCRHMM, a hidden Markov model, to fit those features and trained it on a large dataset representing the entire superfamily. GPCRHMM was benchmarked to profile HMMs and generic transmembrane detectors on sets of known GPCRs and non-GPCRs. In a cross-validation procedure, profile HMMs produced an error rate nearly twice as high as GPCRHMM. In a sensitivity-selectivity test, GPCRHMM's sensitivity was about 15% higher than that of the best transmembrane predictors, at comparable false positive rates. We used GPCRHMM to search for novel members of the GPCR superfamily in five proteomes. All in all we detected 120 sequences that lacked annotation and are potentially novel GPCRs. Out of those 102 were found in Caenorhabditis elegans, four in human, and seven in mouse. Many predictions (65) belonged to Pfam domains of unknown function. GPCRHMM strongly rejected a family of arthropod-specific odorant receptors believed to be GPCRs. A detailed analysis showed that these sequences are indeed very different from other GPCRs. GPCRHMM is available at http://gpcrhmm.cgb.ki.se.

Amino Acids↗

Marked genomic differences characterize primary and secondary glioblastoma subtypes and identify two distinct molecular and clinical secondary glioblastoma entities.

Glioblastoma is classified into two subtypes on the basis of clinical history: "primary glioblastoma" arising de novo without detectable antecedent disease and "secondary glioblastoma" evolving from a low-grade astrocytoma. Despite their distinctive clinical courses, they arrive at an indistinguishable clinical and pathologic end point highlighted by widespread invasion and resistance to therapy and, as such, are managed clinically as if they are one disease entity. Because the life history of a cancer cell is often reflected in the pattern of genomic alterations, we sought to determine whether primary and secondary glioblastomas evolve through similar or different molecular pathogenetic routes. Clinically annotated primary and secondary glioblastoma samples were subjected to high-resolution copy number analysis using oligonucleotide-based array comparative genomic hybridization. Unsupervised classification using genomic nonnegative matrix factorization methods identified three distinct genomic subclasses. Whereas one corresponded to clinically defined primary glioblastomas, the remaining two stratified secondary glioblastoma into two genetically distinct cohorts. Thus, this global genomic analysis showed wide-scale differences between primary and secondary glioblastomas that were previously unappreciated, and has shown for the first time that secondary glioblastoma is heterogeneous in its molecular pathogenesis. Consistent with these findings, analysis of regional recurrent copy number alterations revealed many more events unique to these subclasses than shared. The pathobiological significance of these shared and subtype-specific copy number alterations is reinforced by their frequent occurrence, resident genes with clear links to cancer, recurrence in diverse cancer types, and apparent association with clinical outcome. We conclude that glioblastoma is composed of at least three distinct molecular subtypes, including novel subgroups of secondary glioblastoma, which may benefit from different therapeutic strategies.

Astrocytoma↗

Identification and phenotypic characterization of the cell-division protein CdpA.

Analysis of the automated computer annotation of the early draft phase genome of Lactobacillus acidophilus NCFM revealed the previously discovered S-layer gene slpA and an additional partial ORF with weak similarities to S-layer proteins. The entire gene was sequenced to reveal a 1799-bp gene coding for 599 amino acids with a calculated molecular mass of 64.8 kDa. No transcription or translation signals could be determined in close proximity to the 5'-region. However, a strong putative terminator with a free energy of -16.84 kcal/mol was identified directly downstream of the gene. A PSI-Blast analysis showed similarities to members of S-layer proteins, cell-wall associated proteinases and hexosyl-transferases. Calculation of an unrooted phylogenetic tree with other examples of S-layer proteins and proteinases placed the deduced protein separately from both groups. A derivative of L. acidophilus NCFM was constructed by targeted integration into the gene. SDS-PAGE analysis of non-covalently linked proteins of the cell wall of the mutant, compared to the wild type, revealed the loss of a cell-surface protein. Phenotypic analyses of the mutant revealed significant changes in cell morphology, altered responses to various environmental stresses, and lowered cell adhesion. Based on the in silico and functional analyses, we ascertained that this protein plays a role in cell-wall processing during the growth and cell-cell separation and designated the gene as cell-division protein, cdpA.

Bacterial Adhesion↗

Differential brain transcriptome of beta4 nAChR subunit-deficient mice: is it the effect of the null mutation or the background strain?

Studies using mice with beta4 nicotinic acetylcholine receptor (nAChR) subunit deficiency (beta4-/- mice) helped reveal the roles of this subunit in bradycardiac response to vagal stimulation, nicotine-induced seizure activity and anxiety. To identify genes that might be related to beta4-containing nAChRs activity, we compared the mRNA expression profiles of brains from beta4-/- and wild-type mice using Affymetrix U74Av2 microarray. Seventy-seven genes significantly differentiated between these two experimental groups. Of them, the two most downregulated were spastic paraplegia 21 (human) homolog (Spg21) and 6-pyruvoyl-tetrahydropterin synthase (Pts) genes. Since the targeted mutagenesis of the beta4 nAChR subunit was done by using two mouse strains, 129SvEv and C57BL/6J, it is possible that the genes closely linked to the mutated beta4 gene represent the 129SvEv allele and not the control C57BL/6J-driven allele. We examined this possibility by using public database and quantitative RT-PCR. The expression levels of Spg21 and Pts genes that, like the beta4 gene, are localized on mouse chromosome 9, as well as the expression levels of other genes located on this chromosome, were dependent on the mouse background strain. The 67 differentially expressed genes that are not located on chromosome 9 were further analyzed for overrepresented functional annotations and transcription regulatory elements compared with the entire microarray. Genes encoding for proteins involved in tyrosine phosphatase activity, calcium ion binding, cell growth and/or maintenance, and chromosome organization were overrepresented. Our data enhance the understanding of the molecular interactions involved in the beta4 nAChR subunit function. They also emphasize the need for careful interpretation of expression microarray studies done on genetically manipulated animals.

Animals↗

Identification of chicken transmembrane channel-like (TMC) genes: expression analysis in the cochlea.

Mutations of the human gene encoding transmembrane channel-like protein (TMC)1 cause dominant and recessive nonsyndromic hearing disorders, suggesting that this protein plays an important role in the inner ear. In this study, we cloned chicken Tmc2 (GgTmc2) from a cochlear cDNA library and we annotated four additional TMC family members: GgTmc1, GgTmc3, GgTmc6, and GgTmc7. All chicken TMCs possess the defining TMC signature motif and display high conservation of their genomic structure when compared with other vertebrate TMC genes. GgTmc1 is localized on the chicken sex chromosome Z at a locus that displays conserved synteny with the loci of mammalian orthologues residing on autosomes. In contrast, the locus of GgTmc2 does not exhibit conserved synteny with its mammalian orthologues. Because murine TMC1 and TMC2 are restrictively expressed in cochlear hair cells, we determined the expression of the chicken orthologues in the basilar papilla, the avian equivalent of the organ of Corti. While GgTmc2 was present throughout the basilar papilla and in other tissues, GgTmc1 transcript was detected specifically in the basal portion of the basilar papilla and was not detectable in any other tissue or organ studied. GgTmc3 and GgTmc6 were detectable in all organs analyzed. Antibody labeling revealed that GgTmc2 is predominantly associated with the lateral membranes of hair and supporting cells. The expression of GgTmc2 by both cell types was further confirmed by RT-PCR using isolated cells. This expression and subcellular localization of GgTmc2 is in agreement with the proposed potential role of this novel class of transmembrane proteins in ion transport.

Amino Acid Sequence↗

A high-throughput, near-saturating screen for type III effector genes from Pseudomonas syringae.

Pseudomonas syringae strains deliver variable numbers of type III effector proteins into plant cells during infection. These proteins are required for virulence, because strains incapable of delivering them are nonpathogenic. We implemented a whole-genome, high-throughput screen for identifying P. syringae type III effector genes. The screen relied on FACS and an arabinose-inducible hrpL sigma factor to automate the identification and cloning of HrpL-regulated genes. We determined whether candidate genes encode type III effector proteins by creating and testing full-length protein fusions to a reporter called Delta79AvrRpt2 that, when fused to known type III effector proteins, is translocated and elicits a hypersensitive response in leaves of Arabidopsis thaliana expressing the RPS2 plant disease resistance protein. Delta79AvrRpt2 is thus a marker for type III secretion system-dependent translocation, the most critical criterion for defining type III effector proteins. We describe our screen and the collection of type III effector proteins from two pathovars of P. syringae. This stringent functional criteria defined 29 type III proteins from P. syringae pv. tomato, and 19 from P. syringae pv. phaseolicola race 6. Our data provide full functional annotation of the hrpL-dependent type III effector suites from two sequenced P. syringae pathovars and show that type III effector protein suites are highly variable in this pathogen, presumably reflecting the evolutionary selection imposed by the various host plants.

Arabidopsis↗

Platelet-derived microvesicles induce differential gene expression in monocytic cells: a DNA microarray study.

Platelet-derived microvesicles (PMV) that are shed from the plasma membrane of activated platelets, expose various platelet-type antigens on their surface and are able to adhere to other blood cells and endothelial cells. There are several clinical conditions with markedly increased numbers of PMV, e.g. acute coronary syndrome, thrombotic microangiopathy and sepsis. To prove whether PMV may contribute to an inflammatory response we used DNA microarray technology to study the effect of PMV on gene expression in the prototypic monocytic cell line MonoMac 6 (MM6). PMV were generated by activating human platelets in plasma with collagen and subsequent removal of platelets and plasma by repeated centrifugation. MM6 were incubated for 2 h with PMV in a ratio corresponding to 75 platelets/cell, or saline as control. After RNA isolation, reverse transcription and fluorescence labelling, cDNA was hybridized on a medium density microarray comprising 5308 probes addressing 4868 transcripts of 4730 human genes relevant to inflammation, immune response and related processes. The formation of PMV-MM6 conjugates was associated with significant variations in gene expression, i.e. 93 genes were found to be differentially expressed (P < 0.001; q < 0.087). Among them, 47 genes with annotated transcripts and proteins were identified. Using Ingenuity Pathway Analysis, 37 of the differentially expressed genes were identified as parts of networks associated with functional pathways including cell-to-cell signalling, cellular growth and proliferation, regulation of gene expression and lipid metabolism. For sphingosine kinase-1 the increased expression could be confirmed exemplarily not only by RT-PCR but also on the enzyme activity level. The data indicate that PMV signal differential expression of inflammation-relevant genes in monocytic cells and may represent a novel link between hemostasis and inflammation.

Blood Platelets↗

Functional analysis of CHX21: a putative sodium transporter in Arabidopsis.

The functional role of CHX21, a member of the Arabidopsis thaliana CHX cation transporter family, has been investigated in plants growing under "ideal" conditions and in the presence of elevated NaCl levels. In public databases, AtCHX21 (At2g31910) is annotated as a putative Na+/H+ antiporter. In this study, Southern analysis was used to identify a genotype that contained a single transposon insertion within its genome; using PCR, this insertion was shown to be within the CHX21 locus. No CHX21 transcript was detectable in Atchx21 (mutant) plants using RT-PCR. In the absence of salt stress, Atchx21 showed significant quantitative differences from the wild type (AtCHX21) in development with respect to characters such as rosette width and flowering time. In the presence of 50 mM NaCl, (i) roots of Atchx21 elongated more slowly than the wild type, (ii) the leaf sap Na+ concentration was significantly lower in Atchx21 compared with the wild type, and (iii) the concentra) in the xylem was lower compared with the wild type. The concentration of Na+ exported from the leaf in the phloem was unchanged. Thus, loading of Na+ into the root xylem could explain changes in leaf concentration of Na+. This hypothesis was supported by immunolocalization which demonstrated that the AtCHX21 transporter could only be detected in root endodermal cells. Immunogold labelling of ultra-thin sections, followed by transmission electron microscopy, demonstrated the localization of the protein in the plasma membrane. The data demonstrate that the CHX21 transporter may play a role in regulation of xylem Na+ concentration and, consequently, Na+ accumulation in the leaf.

Arabidopsis↗

Comparison of CID spectra of singly charged polypeptide antibiotic precursor ions obtained by positive-ion vacuum MALDI IT/RTOF and TOF/RTOF, AP-MALDI-IT and ESI-IT mass spectrometry.

Various classes of polypeptide antibiotics, including blocked linear peptides (gramicidin D), side-chain-cyclized peptides (bacitracin, viomycin, capreomycin), side-chain-cyclized depsipeptides (virginiamycin S), real cyclic peptides (tyrocidin, gramcidin S) and side-chain-cyclized lipopeptides (polymyxin B and E, amfomycin), were investigated by low-energy collision induced dissociation (LE-CID) as well as high-energy CID (HE-CID). Ion trap (IT) based instruments with different desorption/ionization techniques such as electrospray ionization (ESI), atmospheric pressure matrix-assisted laser desorption/ionization (AP-MALDI) and vacuum MALDI (vMALDI) as well as a vMALDI-time-of-flight (TOF)/curved field-reflectron instrument fitted with a gas collision cell were used. For optimum comparability of data from different IT instruments, the CID conditions were standardized and only singly charged precursor ions were considered. Additionally, HE-CID data obtained from the TOF-based instrument were acquired and compared with LE-CID data from ITs. Major differences between trap-based and TOF-based CID data are that the latter data set lacks abundant additional loss of small neutrals (e.g. ammonia, water) but contains product ions down to the immonium-ion-type region, thereby allowing the detection of even single amino-acid (even unusual amino acids) substitutions. For several polypeptide antibiotics, mass spectrometric as well as tandem mass spectrometric data are shown and discussed for the first time, and some yet undescribed minor components are also reported. De novo sequencing of unusually linked minor components of (e.g. cyclic) polypeptides is practically impossible without knowledge of the exact structure and fragmentation behavior of the major components. Finally, the described standardized CID condition constitutes a basic prerequisite for creating a searchable, annotated MS(n)-database of bioactive compounds. The applied desorption/ionization techniques showed no significant influence on the type of product ions (neglecting relative abundances of product ions formed) observed, and therefore the type of analyzer connected with the CID process mainly determines the type of fragment ions.

Anti-Bacterial Agents↗

Comparative genomic analysis reveals independent expansion of a lineage-specific gene family in vertebrates: the class II cytokine receptors and their ligands in mammals and fish.

BACKGROUND: The high degree of sequence conservation between coding regions in fish and mammals can be exploited to identify genes in mammalian genomes by comparison with the sequence of similar genes in fish. Conversely, experimentally characterized mammalian genes may be used to annotate fish genomes. However, gene families that escape this principle include the rapidly diverging cytokines that regulate the immune system, and their receptors. A classic example is the class II helical cytokines (HCII) including type I, type II and lambda interferons, IL10 related cytokines (IL10, IL19, IL20, IL22, IL24 and IL26) and their receptors (HCRII). Despite the report of a near complete pufferfish (Takifugu rubripes) genome sequence, these genes remain undescribed in fish. RESULTS: We have used an original strategy based both on conserved amino acid sequence and gene structure to identify HCII and HCRII in the genome of another pufferfish, Tetraodon nigroviridis that is amenable to laboratory experiments. The 15 genes that were identified are highly divergent and include a single interferon molecule, three IL10 related cytokines and their potential receptors together with two Tissue Factor (TF). Some of these genes form tandem clusters on the Tetraodon genome. Their expression pattern was determined in different tissues. Most importantly, Tetraodon interferon was identified and we show that the recombinant protein can induce antiviral MX gene expression in Tetraodon primary kidney cells. Similar results were obtained in Zebrafish which has 7 MX genes. CONCLUSION: We propose a scheme for the evolution of HCII and their receptors during the radiation of bony vertebrates and suggest that the diversification that played an important role in the fine-tuning of the ancestral mechanism for host defense against infections probably followed different pathways in amniotes and fish.

Animals↗

Learning from the data: mining of large high-throughput screening databases.

High-throughput screening (HTS) campaigns in pharmaceutical companies have accumulated a large amount of data for several million compounds over a couple of hundred assays. Despite the general awareness that rich information is hidden inside the vast amount of data, little has been reported for a systematic data mining method that can reliably extract relevant knowledge of interest for chemists and biologists. We developed a data mining approach based on an algorithm called ontology-based pattern identification (OPI) and applied it to our in-house HTS database. We identified nearly 1500 scaffold families with statistically significant structure-HTS activity profile relationships. Among them, dozens of scaffolds were characterized as leading to artifactual results stemming from the screening technology employed, such as assay format and/or readout. Four types of compound scaffolds can be characterized based on this data mining effort: tumor cytotoxic, general toxic, potential reporter gene assay artifact, and target family specific. The OPI-based data mining approach can reliably identify compounds that are not only structurally similar but also share statistically significant biological activity profiles. Statistical tests such as Kruskal-Wallis test and analysis of variance (ANOVA) can then be applied to the discovered scaffolds for effective assignment of relevant biological information. The scaffolds identified by our HTS data mining efforts are an invaluable resource for designing SAR-robust diversity libraries, generating in silico biological annotations of compounds on a scaffold basis, and providing novel target family specific scaffolds for focused compound library design.

Algorithms↗