PubMed HealthSearch

SEARCH · PubMed Health

Results for “Multiple Sequence Alignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Whole genome sequencing of unusual Hepatitis C virus subtypes and drug resistance analysis during direct-acting antiviral therapy in India.

INTRODUCTION AND OBJECTIVES: Pangenotypic direct-acting antivirals (DAA) are effective against highly prevalent Hepatitis C virus (HCV) subtypes, but have been clinically validated almost exclusively in high-income countries. Unusual HCV subtypes may carry natural polymorphisms, potentially impacting DAA susceptibility. We conducted full-genome characterization and resistance analysis of unusual HCV subtypes in patients receiving DAA treatment. PATIENTS AND METHODS: In this prospective hospital-based study, eligible patients were screened for anti-HCV antibodies and active infection was confirmed by diagnostic 5'NCR-based HCV RNA detection. Genotyping was performed by core region sequencing, and viral load quantified by real-time PCR. For whole genome sequencing, multiplex primers were designed using alignments of global reference sequences. Sequencing was carried out using the Oxford Nanopore Technology platform. Phylogenetic analysis used multiple sequence alignment and the HCV-GLUE resource for resistance-associated substitution (RAS) analysis. RESULTS: Predominant genotype was genotype 3 in 64.3% (n = 45); genotype 6 in 21.4% (n = 15); and genotype 1 in 14.2% (n = 10). Unusual HCV subtype 6xa was detected in two patients and showed no NS5A resistance mutations. One genotype 3b patient relapsed at 24 weeks post-DAA treatment completion and carried NS5A resistance-associated substitutions 30 K and 31 M both at baseline and at relapse, conferring high-level resistance to NS5A inhibitors. CONCLUSION: This is the first report from India of whole genome sequencing of HCV subtype 6xa. The identification of NS5A resistance mutations in the 3b relapse case underscores challenges for global HCV elimination strategies.

Humans

Suboptimal sequence alignment in molecular biology. Alignment with error analysis.

A molecular sequence alignment algorithm based on dynamic programming has been extended to allow the computation of all pairs of residues that can be part of optimal and suboptimal sequence alignments. The uncertainties inherent in sequence alignment can be displayed using a new form of dot plot. The method allows the qualitative assessment of whether or not two sequences are related, and can reveal what parts of the alignment are better determined than others. It also permits the computation of representative optimal and suboptimal alignments. The relation between alignment reliability and alignment parameters is discussed. Other applications are to cyclical permutations of sequences and the detection of self-similarity. An application to multiple sequence alignment is noted.

Algorithms

Amino acid sequence analysis of the annexin super-gene family of proteins.

The annexins are a widespread family of calcium-dependent membrane-binding proteins. No common function has been identified for the family and, until recently, no crystallographic data existed for an annexin. In this paper we draw together 22 available annexin sequences consisting of 88 similar repeat units, and apply the techniques of multiple sequence alignment, pattern matching, secondary structure prediction and conservation analysis to the characterisation of the molecules. The analysis clearly shows that the repeats cluster into four distinct families and that greatest variation occurs within the repeat 3 units. Multiple alignment of the 88 repeats shows amino acids with conserved physicochemical properties at 22 positions, with only Gly at position 23 being absolutely conserved in all repeats. Secondary structure prediction techniques identify five conserved helices in each repeat unit and patterns of conserved hydrophobic amino acids are consistent with one face of a helix packing against the protein core in predicted helices a, c, d, e. Helix b is generally hydrophobic in all repeats, but contains a striking pattern of repeat-specific residue conservation at position 31, with Arg in repeats 4 and Glu in repeats 2, but unconserved amino acids in repeats 1 and 3. This suggests repeats 2 and 4 may interact via a buried saltbridge. The loop between predicted helices a and b of repeat 3 shows features distinct from the equivalent loop in repeats 1, 2 and 4, suggesting an important structural and/or functional role for this region. No compelling evidence emerges from this study for uteroglobin and the annexins sharing similar tertiary structures, or for uteroglobin representing a derivative of a primordial one-repeat structure that underwent duplication to give the present day annexins. The analyses performed in this paper are re-evaluated in the Appendix, in the light of the recently published X-ray structure for human annexin V. The structure confirms most of the predictions and shows the power of techniques for the determination of tertiary structural information from the amino acid sequences of an aligned protein family.

Algorithms

An ATPase domain common to prokaryotic cell cycle proteins, sugar kinases, actin, and hsp70 heat shock proteins.

The functionally diverse actin, hexokinase, and hsp70 protein families have in common an ATPase domain of known three-dimensional structure. Optimal superposition of the three structures and alignment of many sequences in each of the three families has revealed a set of common conserved residues, distributed in five sequence motifs, which are involved in ATP binding and in a putative interdomain hinge. From the multiple sequence alignment in these motifs a pattern of amino acid properties required at each position is defined. The discriminatory power of the pattern is in part due to the use of several known three-dimensional structures and many sequences and in part to the "property" method of generalizing from observed amino acid frequencies to amino acid fitness at each sequence position. A sequence data base search with the pattern significantly matches sugar kinases, such as fuco-, glucono-, xylulo-, ribulo-, and glycerokinase, as well as the prokaryotic cell cycle proteins MreB, FtsA, and StbA. These are predicted to have subdomains with the same tertiary structure as the ATPase subdomains Ia and IIa of hexokinase, actin, and Hsc70, a very similar ATP binding pocket, and the capacity for interdomain hinge motion accompanying functional state changes. A common evolutionary origin for all of the proteins in this class is proposed.

Actins

Modeling Alternative Conformational States in CASP16.

The CASP16 Ensemble Prediction experiment assessed advances in methods for modeling proteins, nucleic acids, and their complexes in multiple conformational states. Targets included systems with experimental structures determined in two or three states, evaluated by direct comparison to experimental coordinates, as well as domain-linker-domain (D-L-D) targets assessed against statistical models from NMR and SAXS data. This paper focuses on the former class of multi-state targets. Ten ensembles were released as community challenges, including ligand-induced conformational changes, protein-DNA complexes, a trimeric protein, a stem-loop RNA, and multiple oligomeric states of a single RNA. For five targets, some groups produced reasonably accurate models of both reference states (best TM-score >0.75). However, with the exception of one protein-ligand complex (T1214), where an apo structure was available as a template, predictors generally failed to capture key structural details distinguishing the states. Overall, accuracy was significantly lower than for single-state targets in other CASP experiments. The most successful approaches generated multiple AlphaFold2 models using enhanced multiple sequence alignments and sampling protocols, followed by model quality based selection. While the AlphaFold3 server performed well on several targets, individual groups outperformed it in specific cases. By contrast, predictions for one protein-DNA complex, three RNA targets, and multiple oligomeric RNA states consistently fell short (TM-score <0.75). These results highlight both progress and persistent challenges in multi-state prediction. Despite recent advances, accurate modeling of conformational ensembles, particularly RNA and large multimeric assemblies, remains a critical frontier for structural biology.

AlphaFold2

Consensus patterns in DNA.

Matrices can provide realistic representations of protein/DNA specificity. In many cases simple mononucleotide-based matrices are adequate representations, but more complex matrices may be needed for other cases. Unlike simple consensus sequences, matrices allow for different penalties to be assessed for different changes to a binding site, a property that is essential for accurate description of a binding site pattern. When only a collection of binding site sequences is known, the best representation for the pattern is an information content formulation, based on both thermodynamic and statistical considerations. Quantitative data on relative binding affinities may be used to determine matrices that provide a best fit to the data. Matrix representations also provide an efficient method of aligning multiple sequences to identify binding site patterns that they have in common.

Base Sequence

Genetic diversity and recombination of&#xa0;NA-PRRSV field strains in Vietnam: Implications for vaccine efficacy.

Porcine reproductive and respiratory syndrome (PRRS) causes severe reproductive losses in pregnant sows and piglets, resulting in substantial economic impact on the swine industry worldwide. However, due to the significant genetic diversity and rapid evolutionary changes of the pathogen, continuous surveillance and detailed genetic analysis of circulating strains are essential. The current study aimed to evaluate the genetic diversity of the hypervariable (HV) region of non-structural protein 2 (nsp2) among North American PRRSV strains isolated from swine farms in Vietnam. Phylogenetic analysis and multiple sequence alignment were conducted to determine subtype classification and assess genetic variability. A total of 48 field isolates were obtained, of which 12.5% belonged to classical NA-PRRSV, 16.6% to NADC30-like and 70.9% to HP-PRRSV, primarily distributed across sublineages 1.4, 5.1, 8.7 and 8.9. Amino acid comparisons found multiple insertions, deletions and substitutions at various positions within the hypervariable region of nsp2. The study revealed substantial genetic variation in the HV region of nsp2 among NA-PRRSV field strains, largely associated with recombination and immune escape. These findings highlight epidemiological risks to vaccine efficacy and underscore the need for continuous molecular surveillance to support effective PRRSV control in Vietnam.

PRRSV

Development and epidemiological investigation of a TaqMan-based multiplex real-time quantitative PCR assay for simultaneous detection of five bovine viruses (BVDV, AKAV, BNoV, BEV, and BCoV).

INRODUCTION: Infectious diseases caused by bovine viral diarrhea virus (BVDV), Akabane virus (AKAV), bovine norovirus (BNoV), bovine enterovirus (BEV), and bovine coronavirus (BCoV) significantly threaten the cattle industry, resulting in substantial economic losses. These pathogens often present similar clinical signs, such as diarrhea, vomiting, and reproductive disorders in pregnant cattle, and frequent covert or mixed infections further complicate accurate diagnosis. Therefore, rapid, sensitive, and field&#x2011;deployable diagnostic methods are essential for effective disease surveillance and control in the cattle industry. METHODS: In this study, we report for the first time the establishment of a TaqMan&#x2011;based real&#x2011;time quantitative PCR (qPCR) assay that enables simultaneous detection of these five bovine viruses. Multiple sequence alignment of conserved genomic regions was performed, and virus&#x2011;specific primers and probes were designed and optimized using Beacon Designer 7 software. Subsequently, a TaqMan&#x2011;based multiplex real&#x2011;time qPCR assay was established for simultaneous detection of BVDV, AKAV, BNoV, BEV, and BCoV. The established detection method was applied to 200 clinical samples collected from 10 farms in multiple regions of Jilin Province. RESULTS: The results showed that the detection rates for BVDV, AKAV, BNoV, BEV, and BCoV were 33.50%, 0.50%, 4.50%, 7.50%, and 12.00%, respectively. Mixed infections were detected in 9 samples co&#x2011;infected with two of the five pathogens, with an overall mixed infection rate of 4.50%. Compared with conventional PCR, coincidence rates were 100% for BVDV, AKAV, BNoV, BEV, and BCoV. DISCUSSION: These findings indicate that the TaqMan multiplex real&#x2011;time qPCR assay developed here demonstrates favorable specificity, sensitivity, and reproducibility. This assay enables efficient detection and surveillance of bovine viruses, offering a reliable technical tool for the diagnosis and control of corresponding viral diseases in cattle.

Akabane virus (AKAV)

Molecular analysis of pleckstrin: the major protein kinase C substrate of platelets.

Activation of protein kinase C (PKC) in platelets causes the immediate phosphorylation of pleckstrin, an apparent Mr 40-47,000 protein previously called 40K or P47. Pleckstrin presumably plays an important but as yet unknown role in mediating cellular responses evoked by agonist-induced phosphoinositide turnover. We have cloned the cDNA for pleckstrin from the HL-60 human promyelocytic leukemia cell line by immunological screening of a lambda gt11 expression library (Tyers et al.: Nature 333:470-473, 1988) and now report further analysis of the pleckstrin sequence. Pleckstrin has a deduced Mr of 40,087 and is encoded by a 1,050-bp open reading frame which is preceded by a short open reading frame that terminates before the correct initiator methionine. A single polymorphic site was found in the coding region. An unusual pattern of sequence heterogeneity occurred about a poly(A) tract in the 3' untranslated region. The 3.0-kb pleckstrin mRNA induced upon differentiation of HL-60 cells apparently has heterogeneous 5' ends which undergo differential regulation during HL-60 cell maturation. Analysis by multiple sequence alignment with known PKC substrates identified a strong candidate site for phosphorylation by PKC and a potential Ca2+-binding EF-hand motif. No other similarities to proteins in current databases were found.

Amino Acid Sequence

The evolution of rhodopsins and neurotransmitter receptors.

Rhodopsins share a limited number of amino acid identities with a variety of other integral membrane proteins. Most of these proteins have seven putative transmembrane segments and are likely to play a role in transmembrane signaling. We have undertaken a systematic series of comparisons of primary and secondary structure in order to clarify the functional and evolutionary significance of these sequence similarities. On the basis of consistently high similarity scores, we find that the most internally consistent definition of the rhodopsin gene family would include vertebrate rhodopsins, alpha- and beta-adrenergic receptors, M1 and M2 muscarinic acetylcholine receptors, substance K receptors, and insect rhodopsins, while excluding bacteriorhodopsin, the mass human oncogene, vertebrate and insect nicotinic acetylcholine receptors, and the yeast STE2 and STE3 peptide receptors. The rhodopsin gene family is highly diverged at the primary sequence level but has maintained a conserved secondary structure, including a previously unidentified hierarchy of transmembrane segment hydrophobicity. We have developed new computer algorithms for progressive multiple sequence alignment and the analysis of local conservation of protein domains, and we have used these algorithms to examine the phylogeny of the rhodopsin gene family and the changing domains of sequence conservation. The results show striking differences and similarities in the conserved domains in each of the three main branches of the rhodopsin gene family, and indicate that color vision arose independently in the lines of descent leading to modern humans and fruit flies.

Algorithms

Structural and functional relationships of human DNA polymerases.

A continuing theme of our laboratory has been the understanding of human DNA polymerases at the structural level. We have purified DNA polymerases delta, epsilon and alpha from human placenta. Monoclonal antibodies to these polymerases were isolated and used as tools to study their immunochemical relationships. These studies have shown that while DNA polymerases delta, epsilon and alpha are discrete proteins, they must share common structural features by virtue of the ability of several of our monoclonal antibodies to exhibit cross-reactivity. A second approach we have taken is the molecular cloning of human DNA polymerase delta and epsilon. We have cloned the DNA polymerase delta cDNA, and this has allowed us to compare its primary structure to those of human polymerase alpha and other members of this polymerase family. Multiple sequence alignments have revealed that human DNA polymerase delta is also closely related to the herpes virus family of DNA polymerases. In situ hybridization has shown that the human DNA polymerase delta gene is localized to chromosome 19 q13.3-q13.4. In order to further determine the functional regions of the DNA polymerase delta structure we are currently expressing human pol delta in E. coli and baculovirus systems. Other work in our laboratory is directed toward examining the expression of DNA polymerase delta during the cell cycle.

Amino Acid Sequence

Molecular Cloning, Recombinant Expression, and In Silico Structural Analysis of Cu/Zn-Superoxide Dismutase from Trachyspermum ammi.

Superoxide dismutase (SOD) is an essential antioxidant metalloenzyme that is critical for the cellular defense against oxidative damage, as it scavenges superoxide radicals and maintains the redox status. Cytosolic Cu/Zn-SOD is particularly important in the regulation of oxidative stress among different isoforms in higher plants. While Cu/Zn-SODs from several plant species have been characterized, molecular information is limited for Trachyspermum ammi, a medicinally important member of a family Apiaceae with antioxidant potential.In the present study, an integrated molecular and in silico approach has been taken to clone and analyze a Cu/Zn type SOD gene from T. ammi to get insight into its structural and evolutionary characteristics. PCR amplification yielded an open reading frame of 456&#xa0;bp encoding a protein of 152 amino acids. Sequence analysis showed that plant Cu/Zn-SODs, especially those from Daucus carota, were highly similar to one another (about 90-95%).Multiple sequence alignment confirmed the presence of conserved catalytic motifs and metal-binding histidine residues, both of which are crucial for enzymatic function. Physicochemical analysis predicted the protein to be stable, hydrophilic and compatible with cytosolic localization. The analysis of secondary structure indicated a predominance of &#x3b2;-strands, consistent with the conserved &#x3b2;-barrel architecture of plant Cu/Zn-SODs.The three-dimensional structure was built by homology modeling using a closely related plant Cu/Zn-SOD template with high sequence identity. Structural validation demonstrated an acceptable stereochemical quality with 86.3% residues in the favored region of Ramachandran plot, satisfactory ERRAT and Verify3D scores, and a low RMSD value of 0.104&#xa0;&#xc5; on structural superimposition. Phylogenetic analysis placed the enzyme in the Apiaceae lineage, suggesting evolutionary conservation among related plant species. In conclusion, this study presents the first molecular and structural characterization of Cu/Zn-SOD from T. ammi and confirms the existence of a conserved structural framework typical of plant Cu/Zn-SODs. These results provide a basis for further studies concerning recombinant expression, enzymatic validation and potential relevance in antioxidant and plant stress biology.

Cloning, Molecular

Sequence and localization of human NASP: conservation of a Xenopus histone-binding protein.

In this study the sequence and localization of human testicular NASP (nuclear autoantigenic sperm protein) are reported. NASP cDNA contains 2561 nt encoding a protein of 787 amino acids. The open reading frame contains 2446 nt followed by an ochre stop codon (TAA) and 104 nucleotides of untranslated sequence containing a poly(A) addition signal 10 bases upstream of the poly(A) tail. Northern blot analysis of human testis poly(A) mRNA indicates a message of approximately 3.2 kb. Multiple sequence alignment (MSA) analysis of the encoded human NASP amino acid sequence with the sequence for the Xenopus histone-binding protein N1/N2 and the rabbit NASP amino acid sequence demonstrates that the human sequence and the Xenopus sequence have extensive amino acid homology upstream of the rabbit initiation codon. Significantly, there is an 85% identity between the human and the rabbit NASP sequences when the alignment starts at the N-terminal of the rabbit sequence and at amino acid 101 of the human sequence. The nuclear translocation signal found in N1/N2 and rabbit NASP is completely conserved in human NASP. The first histone-binding domain of Xenopus is 70% identical and 90% similar to the human NASP domain. The second histone-binding domain of Xenopus is 48% identical and 71% similar to the human NASP domain. MSA analysis of the three sequences generated an unrooted ancestral tree with two branches, indicating that fewer amino acid changes have occurred between the Xenopus and the human sequences than between the Xenopus and the rabbit sequences. In the human testis, NASP is localized predominantly in primary spermatocytes and round spermatids. Spermatogonia, Sertoli cells, Leydig cells, peritubular cells, and other somatic cells do not stain. Human spermatozoa contain NASP in the acrosomal region. Following the acrosome reaction, some NASP remains in the equatorial and postacrosomal regions. We propose that mammalian testes and sperm contain a histone-binding protein which may play a role in regulating the early events of spermatogenesis.

Amino Acid Sequence

Prediction of domain organisation and secondary structure of thyroid peroxidase, a human autoantigen involved in destructive thyroiditis.

Organ specific autoimmune diseases are relatively common immunological disorders in man which include thyroid autoimmune disease, insulin-dependent diabetes mellitus and myasthenia gravis. The target autoantigens in some of these diseases have recently been characterised. In thyroid autoimmune disease this includes the key enzyme, thyroid peroxidase (TPO), which is involved in the generation of thyroid hormone. Structural knowledge about autoantigens such as thyroid peroxidase will allow a greater understanding of the interaction between autoantigens and the aberrant immune response, and facilitate the development of strategies for antigen-specific therapeutic manipulation. We report here a prediction of the secondary structure of thyroid peroxidase, together with the results of circular dichroic spectroscopy of a homologous purified enzyme. A combination of 3 secondary structure prediction programs has been used, following multiple sequence alignment, and TPO has been found to consist mainly of alpha-helical conformation, with little beta-sheet present. This structure prediction, together with knowledge of the exon-intron boundaries allows a model for the domain organisation of the TPO molecule to be proposed.

Amino Acid Sequence

A structure-derived sequence pattern for the detection of type I copper binding domains in distantly related proteins.

A structure-based approach to the definition of sequence patterns characteristic of protein domains is presented by example. The approach requires a multiple sequence alignment of a family (or set of related families) as well as at least one three-dimensional structure. The pattern derived does not merely summarize the information in the known sequences but attempts to generalize the pattern specifications based on structural insight. In this example, the pattern-driven database search identified correctly most of the known type I copper-binding domains and detected the presence of a homologous domain in a previously unknown case (CopA protein). The significance of these results is discussed.

Amino Acid Sequence

Conservation analysis and structure prediction of the SH2 family of phosphotyrosine binding domains.

Src homology 2 (SH2) regions are short (approximately 100 amino acids), non-catalytic domains conserved among a wide variety of proteins involved in cytoplasmic signaling induced by growth factors. It is thought that SH2 domains play an important role in the intracellular response to growth factor stimulation by binding to phosphotyrosine containing proteins. In this paper we apply the techniques of multiple sequence alignment, secondary structure prediction and conservation analysis to 67 SH2 domain amino acid sequences. This combined approach predicts seven core secondary structure regions with the pattern beta-alpha-beta-beta-beta-beta-alpha, identifies those residues most likely to be buried in the hydrophobic core of the native SH2 domain, and highlights patterns of conservation indicative of secondary structural elements. Residues likely to be involved in phosphotyrosine binding are shown and orientations of the predicted secondary structures suggested which could enable such residues to cooperate in phosphate binding. We propose a consensus pattern that encapsulates the principal conserved features of the SH2 domains. Comparison of the proposed SH2 domain of akt to this pattern shows only 12/40 matches, suggesting that this domain may not exhibit SH2-like properties.

Amino Acid Sequence

Cloning and sequencing of cDNA clones encoding chicken lamins A and B1 and comparison of the primary structures of vertebrate A- and B-type lamins.

Nuclear lamins are intermediate-filament-type proteins forming a fibrillar meshwork underlying the inner nuclear membrane. The existence of multiple isoforms of lamin proteins in vertebrates is believed to reflect functional specializations during cell division and differentiation. Although biochemical criteria may be used to classify many lamin isoforms into A- and B-type subfamilies, the structural features distinguishing the members of these subfamilies remain to be characterized fully. Here, we report the complete primary structures of chicken lamins A and B1, as they are deduced from cloned cDNAs; in the accompanying paper we present the complete sequence of lamin B2, a second avian B-type lamin. Comparisons of the chicken lamin sequences with each other and with those of other lamins allow us to establish structural features that are common to members of both subfamilies. Conversely, multiple sequence alignments make it possible to identify a number of structural motifs that clearly differentiate B-type lamins from A-type lamins. With this information at hand, we attempt to correlate different biochemical properties of A- and B-type lamins with the presence or absence of specific sequence motifs.

Amino Acid Sequence

Analysis of insertions/deletions in protein structures.

An analysis of insertions and deletions (indels) occurring in a databank of multiple sequence alignments based on protein tertiary structure is reported. Indels prefer to be short (1 to 5 residues). The average intervening sequence length between them versus the percentage of residue identity in pairwise alignments shows an exponential behaviour, suggesting a stochastic process such that nearly every loop in an ancestral structure is a possible target for indels during evolution. The results also suggest a limit to the average size of indels accommodated by protein structures. The preferred indel conformations are reverse turn and coil as are the preferred conformations at the indel edges (N- and C-terminal sides). Interruptions in helices and strands were observed as very rare events.

Amino Acid Sequence