Gene ontology: tool for the unification of biology. The Gene Ontology Consortium.
Explore the source record for details and available documents.
SEARCH · PubMed Health
Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Many examples of the appearance of similar traits in different lineages are known during the evolution of organisms. However, the underlying genetic mechanisms have been elucidated in very few cases. Here, we provide a clear example of evolutionary parallelism, involving changes in the same genetic pathway, providing functional adaptation of RH1 pigments to deep-water habitats during the adaptive radiation of East African cichlid fishes. We determined the RH1 sequences from 233 individual cichlids. The reconstruction of cichlid RH1 pigments with 11-cis-retinal from 28 sequences showed that the absorption spectra of the pigments of nine species were shifted toward blue, tuned by two particular amino acid replacements. These blue-shifted RH1 pigments might have evolved as adaptations to the deep-water photic environment. Phylogenetic evidence indicates that one of the replacements, A292S, has evolved several times independently, inducing similar functional change. The parallel evolution of the same mutation at the same amino acid position suggests that the number of genetic changes underlying the appearance of similar traits in cichlid diversification may be fewer than previously expected.
A transforming N-ras gene has been cloned from acute myeloblastic leukemia bone marrow cells, in parallel with the N-ras gene derived from fibroblasts of the same patient. N-ras derived from fibroblasts lacked focus-forming activity in NIH/3T3 cells, indicating that gene activation in the leukemia cells must have occurred by a somatic event. Construction of chimeric molecules between the transforming and the normal N-ras genes and subsequent biological and sequence analysis of these constructs revealed that the transforming gene was altered by a point mutation changing amino acid 12 of the N-ras protein from glycine to aspartic acid.
The problem of identifying meaningful patterns (i.e., motifs) from biological data has been studied extensively due to its paramount importance. Three versions of this problem have been identified in the literature. One of these three problems is the planted (l, d)-motif problem. Several instances of this problem have been posed as a challenge. Numerous algorithms have been proposed in the literature that address this challenge. Many of these algorithms fall under the category of heuristic algorithms. In this paper we present algorithms for the planted (l, d)-motif problem that always find the correct answer(s). Our algorithms are very simple and are based on some ideas that are fundamentally different from the ones employed in the literature. We believe that the techniques we introduce in this paper will find independent applications.
Current research in the biosciences depends heavily on the effective exploitation of huge amounts of data. These are in disparate formats, remotely dispersed, and based on the different vocabularies of various disciplines. Furthermore, data are often stored or distributed using formats that leave implicit many important features relating to the structure and semantics of the data. Conceptual data modelling involves the development of implementation-independent models that capture and make explicit the principal structural properties of data. Entities such as a biopolymer or a reaction, and their relations, eg catalyses, can be formalised using a conceptual data model. Conceptual models are implementation-independent and can be transformed in systematic ways for implementation using different platforms, eg traditional database management systems. This paper describes the basics of the most widely used conceptual modelling notations, the ER (entity-relationship) model and the class diagrams of the UML (unified modelling language), and illustrates their use through several examples from bioinformatics. In particular, models are presented for protein structures and motifs, and for genomic sequences.
MOTIVATION: Accurate quantification of genotype uncertainty is pivotal in ensuring the reliability of genetic inferences drawn from NGS data. Genotype uncertainty is typically modeled using Genotype Likelihoods (GLs), which can help propagate measures of statistical uncertainty in base calls to downstream analyses. However, the effects of errors and biases in the estimation of GLs, introduced by biases in the original base call quality scores or the discretization of quality scores, as well as the choice of the GL model, remain under-explored. RESULTS: We present vcfgl, a versatile tool for simulating genotype likelihoods associated with simulated read data. It offers a framework for researchers to simulate and investigate the uncertainties and biases associated with the quantification of uncertainty, thereby facilitating a deeper understanding of their impacts on downstream analytical methods. Through simulations, we demonstrate the utility of vcfgl in benchmarking GL-based methods. The program can calculate GLs using various widely used genotype likelihood models and can simulate the errors in quality scores using a Beta distribution. It is compatible with modern simulators such as msprime and SLiM, and can output data in pileup, Variant Call Format (VCF)/BCF, and genomic VCF file formats, supporting a wide range of applications. The vcfgl program is freely available as an efficient and user-friendly software written in C/C++. AVAILABILITY AND IMPLEMENTATION: vcfgl is freely available at https://github.com/isinaltinkaya/vcfgl.
MOTIVATION: Protein language models (PLMs) are amongst the most exciting recent advances for characterizing protein sequences, and have enabled a diverse set of applications, including structure determination, functional property prediction, and mutation impact assessment, all from single protein sequences alone. State-of-the-art PLMs leverage transformer architectures originally developed for natural language processing, and are pre-trained on large protein databases to generate contextualized representations of individual amino acids. To harness the power of these PLMs to predict protein-level properties, these per-residue embeddings are typically "pooled" to fixed-size vectors that are further utilized in downstream prediction networks. Common pooling strategies include Cls-Pooling and Avg-Pooling, but neither of these approaches can capture the local substructures and long-range interactions observed in proteins. RESULTS: We propose the use of attention pooling, which can naturally capture these important features of proteins. To make the expensive attention operator (quadratic in the length of the input protein) feasible in practice, we introduce bag-of-mer pooling, or BoM-Pooling, a locality-aware hierarchical pooling technique that combines windowed average pooling with attention pooling. We empirically demonstrate that both full attention pooling and BoM-Pooling outperform previous pooling strategies on three important, diverse tasks: (i) predicting the activities of two proteins as they are varied; (ii) detecting remote homologs; and (iii) predicting signaling protein interactions with peptides. Overall, our work highlights the advantages of biologically inspired pooling techniques in protein sequence modeling and is a step toward more effective adaptations of language models in biological settings. AVAILABILITY AND IMPLEMENTATION: https://github.com/Singh-Lab/bom-pooling.
SUMMARY: Minimizer digestion is an increasingly common component of bioinformatics tools, including tools for de Bruijn graph assembly and sequence classification. We describe a new open source tool and library to facilitate efficient digestion of genomic sequences. It can produce digests based on the related ideas of minimizers, modimizers or syncmers. Digest uses efficient data structures, scales well to many threads, and produces digests with expected spacings between digested elements. AVAILABILITY AND IMPLEMENTATION: Digest is implemented in C++17 with a Python API, and is available open-source at https://github.com/VeryAmazed/digest. The python library is available on Bioconda. Rust bindings are available as a public crate at https://crates.io/crates/digest-rs.
MOTIVATION: Ribosome dynamics are vital in the process of protein expression. Current methods rely on ribosome profiling (Ribo-seq), RNA-seq profiles, and full genomic context. This restricts their use in de novo sequence design, like messenger RNA (mRNA) vaccines. Simulation-only approaches like the Totally Asymmetric Simple Exclusion Process (TASEP) oversimplify translation by focusing solely on codon elongation times. RESULTS: We present seq2ribo, a hybrid simulation and machine learning framework that predicts ribosome A-site locations using only an mRNA sequence as input. Our method first employs a novel structure-aware TASEP (sTASEP), which models translation using a comprehensive set of fitted parameters that include codon wait times and structural features, such as local angles, base-pairing, and discrete positional buckets. The ribosome locations generated by sTASEP are then processed by a polisher model, which learns to refine the simulated ribosome distributions. seq2ribo provides high-fidelity predictions of ribosome locations across diverse cell types (iPSC, HEK293, LCL, and RPE-1), significantly outperforming baselines. seq2ribo is the first method to achieve meaningful positional correlation with observed ribosome profiles from sequence alone, reaching transcript-level Pearson correlations up to 0.920 and within-transcript shape correlations up to 0.186, where all baselines yield near-zero values on these metrics. seq2ribo also reduces elementwise error by up to 37.7% relative to the sequence-only Translatomer baseline. By adding a task-specific head, seq2ribo achieves Pearson correlations up to 0.732 with experimental translation efficiency (TE) across several cell lines, and up to 0.903 with measured protein expression. By operating from sequence alone, seq2ribo provides a new tool for synthetic biology, enabling the rational design and optimization of mRNA sequences without the need for expression-level data or genomic context. AVAILABILITY: seq2ribo is available at https://github.com/Kingsford-Group/seq2ribo.
MOTIVATION: Understanding the functional impact of genetic variants is a key problem for precision medicine. Tools like CADD, PhyloP, and PhastCons are useful, but they often look at each position in the genome in isolation. This means they can miss important information from the evolutionary history that connects different species. In this paper, we extend our previous model, Graphylo, to predict the effects of variants. Our new model, GraphyloVar, is built to directly utilize the phylogenetic tree that relates the species. RESULTS: GraphyloVar is a deep learning model that considers both DNA sequence and evolutionary patterns from many species. It uses two main components: Graph Convolutional Networks (GCNs) to process the phylogenetic tree, and Transformer encoders to extract features from the DNA sequences. Pre-trained to predict population-level allele frequencies on the TOPMed whole-genome sequencing cohort, GraphyloVar achieves an AUROC of 0.6246 zero-shot on ∼149M held-out variants, and an ensemble with CADD reaches 0.6442 (+0.020, P<10-15). Fine-tuned GraphyloVar achieves the highest AUROC across all 13 MPRA benchmark datasets. By integrating deep learning with explicit phylogenetic input, GraphyloVar offers a powerful and complementary approach to variant effect prediction that utilizes the full evolutionary history from many species to better identify and prioritize important non-coding variants. AVAILABILITY AND IMPLEMENTATION: Code and datasets are available at https://github.com/DongjoonLim/GraphyloVar under DOI: 10.5281/zenodo.20616818.
We have developed an algorithm for designing multiple sequences of nucleic acids that have a uniform melting temperature between the sequence and its complement and that do not hybridize non-specifically with each other based on the minimum free energy (DeltaG (min)). Sequences that satisfy these constraints can be utilized in computations, various engineering applications such as microarrays, and nano-fabrications. Our algorithm is a random generate-and-test algorithm: it generates a candidate sequence randomly and tests whether the sequence satisfies the constraints. The novelty of our algorithm is that the filtering method uses a greedy search to calculate DeltaG (min). This effectively excludes inappropriate sequences before DeltaG (min) is calculated, thereby reducing computation time drastically when compared with an algorithm without the filtering. Experimental results in silico showed the superiority of the greedy search over the traditional approach based on the hamming distance. In addition, experimental results in vitro demonstrated that the experimental free energy (DeltaG (exp)) of 126 sequences correlated well with DeltaG (min) (|R| = 0.90) than with the hamming distance (|R| = 0.80). These results validate the rationality of a thermodynamic approach. We implemented our algorithm in a graphic user interface-based program written in Java.
Nevirapine-resistant variants were generated by serial passages in MT-2 cells in the presence of increasing drug concentrations. In passage 5, mutations V106A, Y181C and G190A were detected in the global population, associated with a 100-fold susceptibility decrease. Sequence analysis of biological clones obtained from passage 5 and subsequent passages showed that single mutants, detected in first passages, were progressively replaced in passage 15 by double mutants, correlating with a 500-fold increase in phenotypic resistance. Fitness determination of single mutants confirmed that, in the presence of nevirapine, every variant was more fit than wild-type with a fitness order Y181C>V106A>G190A>wild-type. Unexpectedly, in the absence of the drug, the Y181C resistant mutant was more fit than wild-type, with a fitness gradient Y181C>wild-type >G106A>or=V190A. Using a molecular clone in which the Y181C mutation was introduced by in vitro mutagenesis, the greater fitness of the Y181C mutant was confirmed in new competition cultures. These data exemplify the role of resistance mutations on virus phenotype but also on virus evolution leading, occasionally, to resistant variants fitter than the wild-type in the absence of the drug.
Flagellin from Bacillus thuringiensis subspecies alesti strain Bt75 was isolated both from the culture medium and from flagella. Two protein forms with molecular masses close to 32 kDa were obtained from flagella; one form was identical to the flagellin purified from the culture medium. The N-terminal amino acid sequences were identical for both forms. Two genes coding for flagellin have been identified in B. thuringiensis subsp. alesti. The flaB gene was cloned and sequenced in its entire length. The clone containing the flaA gene was incomplete. Both genes were expressed in the mid-exponential growth phase. The flaB gene was flanked by long (355 bp) direct repeats protruding into the coding region in both the N- and C-terminal parts of the gene. DNA sequences related to the flaB gene were found in most other B. thuringiensis subspecies, and in two of them, subsp. kurstaki and subsp. entomocidus, such sequences were present in multiple copies.
Explore the source record for details and available documents.
The importance of the cyanobacteria Prochlorococcus and Synechococcus in marine ecosystems in terms of abundance and primary production can be partially explained by ecotypic differentiation. Despite the dominance of eukaryotes within photosynthetic picoplankton in many areas a similar differentiation has never been evidenced for these organisms. Here we report distinct genetic [rDNA 18S and internal transcribed spacer (ITS) sequencing], karyotypic (pulsed-field gel electrophoresis), phenotypic (pigment composition) and physiological (light-limited growth rates) traits in 12 Ostreococcus strains (Prasinophyceae) isolated from various marine environments and depths, which suggest that the concept of ecotype could also be valid for eukaryotes. Internal transcribed spacer phylogeny grouped together four deep strains isolated between 90 m and 120 m depth from different geographical origins. Three deep strains displayed larger chromosomal bands, different chromosome hybridization patterns, and an additional chlorophyll (chl) c-like pigment. Furthermore, growth rates of deep strains show severe photo-inhibition at high light intensities, while surface strains do not grow at the lowest light intensities. These features strongly suggest distinct adaptation to environmental conditions encountered at surface and the bottom of the oceanic euphotic zone, reminiscent of that described in prokaryotes.
Early substantial evidence of the low susceptibility to beta-adrenoceptor antagonists of non alpha-adrenergic responses reducing gut motility and tone was reluctantly accepted as indicating a third beta-receptor subtype different from the beta 1 and beta 2. This applied likewise to lipolysis until new selective "lipolytic" beta-agonists poorly effective at established beta-receptors were introduced. Shortly afterwards these "lipolytic" as well as certain newer and even more selective beta-adrenoceptor agonists were shown to be potent inhibitors of intestinal motility. The latter are the "gut-specific" phenylethanolaminotetralins whose availability as pure isomers attested to the stringent stereochemical requirements for selectivity at non-beta 1, non-beta 2 beta-adrenoceptors. Acceptance of the functionally based concept of a beta 3-adrenoceptor was boosted on structural grounds by molecular biology studies. Sequence analysis indicated the existence in humans and rodents of genes coding for a third subtype of beta-receptor that, when expressed in transfected heterologous cells, had a pharmacological profile distinct from the previously established subtypes. Finally, aryloxypropanolaminotetralins have been prepared as the first selective antagonists of beta 3-adrenoceptors, thus providing unambiguous conclusive evidence of the distinctive functional features of those abundant in the rat colon. The therapeutic potential in gastroenterology of the newer compounds targetable on the beta 3-adrenoceptor is suggested by their potent intestinal action in vivo in animal models without any of the cardiovascular or other unwanted effects of conventional beta-adrenoceptor agonists and antagonists, and by the clinically confirmed importance of beta-adrenergic control of motor function throughout the alimentary canal. However, open questions include the incidence of species-related differences in beta 3-adrenoceptors, and as yet there are no data on gastrointestinal functions in humans under the influence of drugs designed to act selectively at these receptors.
Microorganisms populate every habitable environment on Earth and, through their metabolic activity, affect the chemistry and physical properties of their surroundings. They have done this for billions of years. Over the past decade, genetic, biochemical, and genomic approaches have allowed us to document the diversity of microbial life in geologic systems without cultivation, as well as to begin to elucidate their function. With expansion of culture-independent analyses of microbial communities, it will be possible to quantify gene activity at the species level. Genome-enabled biogeochemical modeling may provide an opportunity to determine how communities function, and how they shape and are shaped by their environments.