PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological sequence analysis”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

The interaction of boar sperm proacrosin with its natural substrate, the zona pellucida, and with polysulfated polysaccharides.

Boar sperm acrosin is an acrosomal protease with trypsin-like specificity, and it functions in fertilization by assisting sperm passage through the zona pellucida by limited hydrolysis of this extracellular matrix. In addition to a proteolytic active site domain, acrosin binds the zona pellucida at a separate binding domain that is lost during proacrosin autolysis. In this study, we quantitate the binding of proacrosin to the physiological substrate for acrosin, the zona pellucida, and to a non-substrate, the polysulfated polysaccharide fucoidan. Binding was analogous to sea urchin sperm bindin that binds egg jelly fucan and the vitelline envelope of sea urchin eggs. Proacrosin was found to bind to fucoidan and to the zona pellucida with binding affinities similar to bindin interaction with egg jelly fucan. These interactions were competitively inhibited by similar relative molecular mass polysulfated polymers. Since bindin and proacrosin have distinctly different amino acid sequences, their interaction with acidic sulfate esters demonstrates an example of convergent evolution wherein different macromolecules localized in analogous sperm compartments have the same biological function. From cDNA sequence analysis of proacrosin, this binding may be mediated through a consensus sequence for binding sulfated glycoconjugates. Proacrosin binding to the zona pellucida may serve as both a recognition or primary sperm receptor, as well as maintaining the sperm on the zona pellucida once the acrosome reaction has occurred.

Acrosin↗

A portable recalibration workflow for reference-based variant calling in non-human genomes.

A key computational step in reference-based variant calling is distinguishing true genetic variants from sequencing errors. Advanced tools and workflows have been developed to handle this by computational modelling of technical errors from the sequencing machines. However, these recalibration workflows have largely been evaluated for human data only and its exact applicability for non-human data remains unknown. Here, we conducted a systematic evaluation of variant calling on human, rice, sheep, and chickpea data, and found that existing workflows introduce unexpected statistical bias, thus leading to suboptimal variant calls for non-human data. To address this problem, we present simple guidelines for constructing a "pseudo-"database (pseudoDB) of genetic variants as a scalable and portable solution for recalibration and variant calling. With human data, our pseudoDB-based workflow performs comparably to existing dbSNP-based GATK3 workflows and those using DeepVariant, Strelka2, and FreeBayes. We extend this to other non-human genomes, namely cattle, brown bear, swan goose, African oil palm, Komodo dragon, and stevia, altogether resulting in the identification of up to 242.0% unique genetic variants. The majority of newly identified variants are within the non-coding regions, hinting at the rich diversity of genome regulation in the non-human population. Our pseudoDB-based workflow is agnostic to reference genomes and modular for easy integration with other computational workflows for human and non-human resequencing data.

Humans↗

Accelerating discoveries in the proteome and genome with MALDI TOF MS.

Recent developments in mass spectrometry (MS) provide scientists with an established analytical tool that addresses the demands for rapid, accurate and cost effective analyses of biomolecules. These advances clearly accelerate the rate and success of protein identification, genetic sequencing, determining biological variances and drug discovery. This review is intended to illustrate how matrix-assisted laser desorption/ionisation (MALDI) time-of-flight (TOF) MS is typically used to generate information about biomolecules under investigation. Additionally, examples will be used to describe the steps involved in preparing samples for MALDI TOF and obtaining answers through data management.

Animals↗

Multiple gene differential expression patterns in human ischemic liver: safe limit of warm ischemic time.

AIM: To investigate the multiple gene differential expression patterns in human ischemic liver and to produce the evidence about the hepatic ischemic safety time. METHODS: The responses of cells to hepatic ischemia and hypoxia at hepatic ischemia were analyzed by cDNA microarrary representing 4 000 different human genes containing 200 apoptotic correlative genes. RESULTS: There were lower or normal expression levels of apoptotic correlative genes during the periods of hepatic ischemia for 0-15 min, the maintenance homostatic genes were expressed significantly higher at the same time. But at the hepatic ischemia for 30 min, the expression levels of maintenance homeostatic genes were down-regulated, the expressions of many apoptotic correlative genes and nuclear transcription factors were activated and up-regulated. CONCLUSION: HIF-1, APAF-1, PCDC10, FBX5, DFF40, DFFA XIAP, survivin may be regarded as the signal genes to judge the degree of hepatic ischemic-hypoxic injure, and the apoptotic liver cell injury due to ischemia in different time limits. The safe limit of human hepatic warm ischemic time appears to be generally less then 30 min.

Down-Regulation↗

Atriopeptin biochemical pharmacology.

Several low-molecular-weight peptides that possess potent natriuretic, diuretic, and vascular smooth muscle relaxant activity have been isolated from atrial extracts. Elucidation of their structure indicates that they consist of a 17-membered ring of amino acids formed by a cystine disulfide bond and that they differ only in the composition of the amino and carboxy termini. The 24-amino-acid peptide atriopeptin (AP) III was selected as the reference compound for structure-activity studies. Amino-terminal amino acid extensions on APIII markedly increase the natriuretic-diuretic but not the renal vasodilatory response in anesthetized dogs, which suggests a heterogeneity of AP receptors in renal tubular and vascular tissues. Radioligand (125I-labeled APIII) binding studies with fresh rat kidney slices indicate that the primary renal sites of specific AP binding are in the glomerulus and in the papillary segment of the medulla, thus implicating these structures in the natriuretic-diuretic effect. Data obtained from radioimmunoassay, chromatographic migration, vasorelaxant biological activity, and peptide sequence analysis indicate that Ser-Leu-Arg-Arg-APIII is the major circulating form of low-molecular-weight atrial peptide present in rat plasma. Circulating APs fulfill many of the criteria for involvement in the endocrine regulation of fluid and electrolyte homeostasis.

Amino Acid Sequence↗

Molecular classification of cancer types from microarray data using the combination of genetic algorithms and support vector machines.

Simultaneous multiclass classification of tumor types is essential for future clinical implementations of microarray-based cancer diagnosis. In this study, we have combined genetic algorithms (GAs) and all paired support vector machines (SVMs) for multiclass cancer identification. The predictive features have been selected through iterative SVMs/GAs, and recursive feature elimination post-processing steps, leading to a very compact cancer-related predictive gene set. Leave-one-out cross-validations yielded accuracies of 87.93% for the eight-class and 85.19% for the fourteen-class cancer classifications, outperforming the results derived from previously published methods.

Algorithms↗

[Estimation of evolutionary distances between protein sequences].

Several estimates of the evolutionary distance between two homologous protein sequences were deduced, taking into account of the variation of replacement rate over amino acid sites. A maximum likelihood estimator was also presented with consideration of different probabilities of replacement among amino acids. Suggestions were made as to the application of these distance estimates to real sequence analysis.

Animals↗

CAKR: commutative algebra k-mer representations for genomics.

Despite the availability of various sequence analysis models, comparative genomic analysis remains a challenge in genomics, genetics, and phylogenetics. Commutative algebra, a fundamental tool in algebraic geometry and number theory, has rarely been used in data and biological sciences. In this study, we introduce commutative algebra k-mer representations as a nonlinear algebraic framework for analyzing genomic sequences. This representation bridges commutative algebra, algebraic topology, combinatorics, and machine learning to establish a mathematical framework for comparative genomic analysis. We evaluate its effectiveness on three tasks including genetic variant classification, phylogenetic tree reconstruction, and viral classification, typically requiring alignment-based, alignment-free, and machine-learning approaches, respectively. In this work, we show that commutative algebra k-mer representations outperform five state-of-the-art sequence analysis methods across twelve primary datasets, with two additional supplementary fragment-placement benchmarks, especially in viral classification, and maintain relatively stable predictive accuracy as dataset size increases, underscoring scalability and robustness.

Genomics↗

Integrated biclustering of heterogeneous genome-wide datasets for the inference of global regulatory networks.

BACKGROUND: The learning of global genetic regulatory networks from expression data is a severely under-constrained problem that is aided by reducing the dimensionality of the search space by means of clustering genes into putatively co-regulated groups, as opposed to those that are simply co-expressed. Be cause genes may be co-regulated only across a subset of all observed experimental conditions, biclustering (clustering of genes and conditions) is more appropriate than standard clustering. Co-regulated genes are also often functionally (physically, spatially, genetically, and/or evolutionarily) associated, and such a priori known or pre-computed associations can provide support for appropriately grouping genes. One important association is the presence of one or more common cis-regulatory motifs. In organisms where these motifs are not known, their de novo detection, integrated into the clustering algorithm, can help to guide the process towards more biologically parsimonious solutions. RESULTS: We have developed an algorithm, cMonkey, that detects putative co-regulated gene groupings by integrating the biclustering of gene expression data and various functional associations with the de novo detection of sequence motifs. CONCLUSION: We have applied this procedure to the archaeon Halobacterium NRC-1, as part of our efforts to decipher its regulatory network. In addition, we used cMonkey on public data for three organisms in the other two domains of life: Helicobacter pylori, Saccharomyces cerevisiae, and Escherichia coli. The biclusters detected by cMonkey both recapitulated known biology and enabled novel predictions (some for Halobacterium were subsequently confirmed in the laboratory). For example, it identified the bacteriorhodopsin regulon, assigned additional genes to this regulon with apparently unrelated function, and detected its known promoter motif. We have performed a thorough comparison of cMonkey results against other clustering methods, and find that cMonkey biclusters are more parsimonious with all available evidence for co-regulation.

Algorithms↗

Optimization of single-strand conformation polymorphism and sequence analysis of the mitochondrial control region in Pagellus bogaraveo (Sparidae, Teleostei): rationalized tools in fish population biology.

We report the isolation and sequencing of 400-550 base pairs (bp) of the mitochondrial DNA (mtDNA) control region of eight species of Sparidae (Perciformes, Teleostei). This sequence information allowed us to design specific primers to one of these species (Pagellus bogaraveo). The new set of primers was used to test a rationalized approach to study the mtDNA nucleotide variability at the intraspecific level. The single-strand conformation polymorphism (SSCP) technique was applied to detect sequence variation in two non-overlapping fragments of the control region of 32 individuals of P. bogaraveo. To assess the sensitivity of the method, the nucleotide sequence of the analysed region was determined for all the specimens. The results showed that, for one of the two fragments, SSCP analysis was able to detect 100% of the underlying genetic variability. In sharp contrast, nucleotide variation of the second DNA fragment was completely unresolved by SSCP under different experimental conditions. This suggests that the resolution power of SSCP is crucially dependent on the nature of the fragment subjected to the analysis; therefore, a preliminary test of the sensitivity of the method should be performed on each specific DNA fragment before starting a large-scale survey. A rationalized approach, combining the SSCP technique and a simplified sequencing procedure, is proposed for studying intraspecific polymorphism at the mtDNA control region in fish.

Animals↗

Validation of cytochrome b sequence analysis as a method of species identification.

One of the stages of dealing with biological material submitted to forensic laboratories is species identification. The aim of the present work was to validate and assess the possibility of applying sequence analysis of the region coding cytochrome b as a method of species identification in the field of forensic science. DNA originating from individuals from major phyla of vertebrates was isolated by the organic method from various specimens. Extracted DNA was subjected to PCR and direct cycle sequencing using a universal pair of primers. The validation process, performed according to TWGDAM recommendations, revealed that the technique is a very sensitive and reliable method of species identification allowing analysis of tiny amounts of material and also degraded material, and can be useful in the field of forensic genetics. The case example presented here, concerning the determination of species origin of biological evidence collected from fatal road accident, confirms that analysis can be carried out even when there is no reference sample, and the sequences obtained can be assessed through analysis of their similarity to sequences for cytochrome b present in DNA databases.

Animals↗

Biological and structural properties of MIP-1 alpha expressed in yeast.

The murine macrophage inflammatory proteins-1 alpha (MIP-1 alpha) and MIP-1 beta are distinct but closely related cytokines. Partially purified mixtures of the two proteins affect neutrophil function and cause local inflammation and fever. The particular properties of MIP-1 alpha have not been well studied, although it has been identified as being identical to an inhibitor of haemopoietic stem cell growth. We have expressed MIP-1 alpha in yeast cells and purified it to sequence homogeneity. Structural analysis of this biologically active material by circular dichroism and fluorescence spectroscopy confirms that MIP-1 alpha has a very similar secondary and tertiary structure to platelet factor 4 and interleukin 8 with which it shares limited sequence homology. The in-vitro stem cell inhibitory properties have been confirmed using a range of murine progenitor cells including purified bone marrow progenitor cells (FACS-1), the FDCP-mix A4 cell line, and spleen colony forming unit (CFU-S) populations. Plateau levels of inhibition of stem cell growth were achieved using concentrations of 0.15 micrograms/ml MIP-1 alpha. We have also demonstrated that MIP-1 alpha is active in vivo: 5 micrograms of MIP-1 alpha per mouse given as a bolus injection, protects stem cells from subsequent in-vitro killing by tritiated thymidine. MIP-1 alpha was also shown to enhance the proliferation of more committed progenitor granulocyte macrophage-colony forming cells (GM-CFC) in response to granulocyte macrophage-colony stimulating factor (GM-CSF).

Animals↗

A method for rapid similarity analysis of RNA secondary structures.

BACKGROUND: Owing to the rapid expansion of RNA structure databases in recent years, efficient methods for structure comparison are in demand for function prediction and evolutionary analysis. Usually, the similarity of RNA secondary structures is evaluated based on tree models and dynamic programming algorithms. We present here a new method for the similarity analysis of RNA secondary structures. RESULTS: Three sets of real data have been used as input for the example applications. Set I includes the structures from 5S rRNAs. Set II includes the secondary structures from RNase P and RNase MRP. Set III includes the structures from 16S rRNAs. Reasonable phylogenetic trees are derived for these three sets of data by using our method. Moreover, our program runs faster as compared to some existing ones. CONCLUSION: The famous Lempel-Ziv algorithm can efficiently extract the information on repeated patterns encoded in RNA secondary structures and makes our method an alternative to analyze the similarity of RNA secondary structures. This method will also be useful to researchers who are interested in evolutionary analysis.

Algorithms↗

libcov: a C++ bioinformatic library to manipulate protein structures, sequence alignments and phylogeny.

BACKGROUND: An increasing number of bioinformatics methods are considering the phylogenetic relationships between biological sequences. Implementing new methodologies using the maximum likelihood phylogenetic framework can be a time consuming task. RESULTS: The bioinformatics library libcov is a collection of C++ classes that provides a high and low-level interface to maximum likelihood phylogenetics, sequence analysis and a data structure for structural biological methods. libcov can be used to compute likelihoods, search tree topologies, estimate site rates, cluster sequences, manipulate tree structures and compare phylogenies for a broad selection of applications. CONCLUSION: Using this library, it is possible to rapidly prototype applications that use the sophistication of phylogenetic likelihoods without getting involved in a major software engineering project. libcov is thus a potentially valuable building block to develop in-house methodologies in the field of protein phylogenetics.

Algorithms↗

A weighted measure for the similarity analysis of DNA sequences.

Here we propose a weighted measure for the similarity analysis of DNA sequences. It is based on LZ complexity and (0,1) characteristic sequences of DNA sequences. This weighted measure enables biologists to extract similarity information from biological sequences according to their requirements. For example, by this weighted measure, one can obtain either the full similarity information or a similarity analysis from a given biological aspect. Moreover, the length of DNA sequence is not problematic. The application of the weighted measure to the similarity analysis of beta-globin genes from nine species shows its flexibility.

Animals↗

Upscaling Genotyping by Amplicon Sequencing With GBAS-GUI.

Genotyping by amplicon sequencing (GBAS) is a relatively low-cost approach for generating genotypic data compared with established genomic methods, making it highly scalable and particularly suitable for large-scale genetic monitoring projects. However, most existing analytical pipelines are either marker-specific, insufficiently scalable, or lacking efficient data management systems for the long-term integration of genotypic information, limiting the full potential of GBAS. Here, we address this gap by introducing GBAS-GUI (https://github.com/sonnenbe-dot/GBAS-GUI), a pipeline capable of generating GBAS-based genotypic data for a wide variety of loci at scale. GBAS-GUI integrates a graphical user interface with multiple checkpoints to improve accessibility and robustness. It implements multiprocessing architecture and a relational database that links genotypic data with associated sample metadata to enhance scalability and data management. The pipeline further enables marker screening through automated calculation of polymorphism information content (PIC) and implements a strategy to recover homologous genotypic information from paralogous loci with non-overlapping amplicon length ranges. Using multiple empirical datasets, we demonstrate substantial improvements in processing speed, database management and handling artefacts related to co-amplification of unspecific regions and duplicates of the same genomic region. We further show that incorporating the full sequence information captured by an amplicon increases marker information content beyond what is achievable with length-based genotyping alone and expands the analytical versatility of GBAS. Overall, GBAS-GUI provides a robust, scalable and versatile framework that unlocks the potential of GBAS for large-scale population genetic and phylogeographic studies.

Genotyping Techniques↗

A biological question and a balanced (orthogonal) design: the ingredients to efficiently analyze two-color microarrays with Confirmatory Factor Analysis.

BACKGROUND: Factor analysis (FA) has been widely applied in microarray studies as a data-reduction-tool without any a-priori assumption regarding associations between observed data and latent structure (Exploratory Factor Analysis).A disadvantage is that the representation of data in a reduced set of dimensions can be difficult to interpret, as biological contrasts do not necessarily coincide with single dimensions. However, FA can also be applied as an instrument to confirm what is expected on the basis of pre-established hypotheses (Confirmatory Factor Analysis, CFA). We show that with a hypothesis incorporated in a balanced (orthogonal) design, including 'SelfSelf' hybridizations, dye swaps and independent replications, FA can be used to identify the latent factors underlying the correlation structure among the observed two-color microarray data. An orthogonal design will reflect the principal components associated with each experimental factor. We applied CFA to a microarray study performed to investigate cisplatin resistance in four ovarian cancer cell lines, which only differ in their degree of cisplatin resistance. RESULTS: Two latent factors, coinciding with principal components, representing the differences in cisplatin resistance between the four ovarian cancer cell lines were easily identified. From these two factors 315 genes associated with cisplatin resistance were selected, 199 genes from the first factor (False Discovery Rate (FDR): 19%) and 152 (FDR: 24%) from the second factor, while both gene sets shared 36. The differential expression of 16 genes was validated with reverse transcription-polymerase chain reaction. CONCLUSION: Our results show that FA is an efficient method to analyze two-color microarray data provided that there is a pre-defined hypothesis reflected in an orthogonal design.

Antineoplastic Agents, Alkylating↗