PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “sequence evolution”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,261 records · Page 70Linked to original sources

Preparation and partial structural characterization of alpha1T-glycoprotein from normal human plasma.

alpha1T-glycoprotein (alpha1T) was isolated from normal human plasma in the immunochemically homogeneous state. The partial amino acid sequence and carbohydrate chains of this glycoprotein were determined. To achieve this, the carboxymethylated alpha1T was analyzed by sequencing some of the lysylendoprotease, V8 protease, tryptic, and cyanogen bromide peptides as well as the N-terminal sequence of the protein. A large number of amino acid residues (460 amino acids) was determined by chemical procedure. The peptide sequences were compared with that of other proteins. A high degree of homology was found for proteins of the albumin family. Further, human alpha-albumin, a new member of this protein family, showed an amino acid sequence identical to that of alpha1T indicating that these two proteins are very similar in amino acid sequence and composition. These proteins are closely related to alpha-fetoprotein; however, five carbohydrate chains were found on alpha1T at Asn12, Asn88, Asn362, Asn381, and Asn467 as biantennary complex type chains and the chain on Asn362 possessed a rare consensus sequence of Asn-X-Cys. Thus, alpha1T distinguishes itself by possessing five N-glycans, a finding reported here for the first time for the ALB family.

Albumins↗

Using evolutionary trees in protein secondary structure prediction and other comparative sequence analyses.

Previously proposed methods for protein secondary structure prediction from multiple sequence alignments do not efficiently extract the evolutionary information that these alignments contain. The predictions of these methods are less accurate than they could be, because of their failure to consider explicitly the phylogenetic tree that relates aligned protein sequences. As an alternative, we present a hidden Markov model approach to secondary structure prediction that more fully uses the evolutionary information contained in protein sequence alignments. A representative example is presented, and three experiments are performed that illustrate how the appropriate representation of evolutionary relatedness can improve inferences. We explain why similar improvement can be expected in other secondary structure prediction methods and indeed any comparative sequence analysis method.

Amino Acid Sequence↗

Phylogenetic and epidemiologic analysis of the walleye dermal sarcoma virus.

Walleye dermal sarcoma virus (WDSV) is a newly described retrovirus that is etiologically associated with a multifocal skin tumor of a fish common in North America, the walleye. Tumor prevalence ranges from 27% of adult walleyes in a densely populated lake, Oneida Lake, New York, to 1% in less populated waters. Phylogenetic analysis of the surface (SU) domain of the WDSV envelope gene of isolates from different regions of North America showed that viral isolates formed distinct clusters according to their geographic origin, except viral isolates from Oneida Lake, which were also much more variable. Viral clones isolated from an individual tumor had identical nucleotide sequences. This finding is consistent with tumors developing from single infected dermal cells, and supports the etiological role of this virus in tumor development. Like in other retroviruses, the SU domain of the WDSV env gene was more variable than gag, and the ratio of nonsynonymous over synonymous mutations was comparable to that of the V3 loop of HIV-1. These findings indicate that WDSV SU is the object of strong selective immunologic pressures, like the SU domain of other retroviruses.

Base Sequence↗

Trematode and monogenean rRNA ITS2 secondary structures support a four-domain model.

The secondary structure of rRNA internal transcribed spacer 2 is important in the process of ribosomal biogenesis. Trematode ITS sequences are poorly conserved and difficult to align for phylogenetic comparisons above a family level. If a conserved secondary structure can be identified, it can be used to guide primary sequence alignments. ITS2 sequences from 39 species were compared. These species span four orders of trematodes (Echinostomiformes, Plagiorchiformes, Strigeiformes, and Paramphistomiformes) and one monogenean (Gyrodactyliformes). The sequences vary in length from 251 to 431 bases, with an average GC content of 48%. The monogenean sequence could not be aligned with confidence to the trematodes. Above the family level trematode sequences were alignable from the 5' end for 139 bases. Secondary structure foldings predicted a four-domain model. Three folding patterns were required for the apex of domain B. The folding pattern of domains C and D varies for each family. The structures display a high GC content within stems. Bases A and U are favored in unpaired regions and variable sites cluster. This produces a mosaic of conserved and variable regions with a structural conformation resistant to change. Two conserved strings were identified, one in domain B and the other in domain C. The first site can be aligned to a processing site identified in yeast and rat. The second site has been found in plants, and structural location appears to be important. A phylogenetic tree of the trematode sequences, aligned with the aid of secondary structures, distinguishes the four recognized orders.

Animals↗

tRNA creation by hairpin duplication.

Many studies have suggested that the modern cloverleaf structure of tRNA may have arisen through duplication of a primordial hairpin, but the timing of this duplication event has been unclear. Here we measure the level of sequence identity between the two halves of each of a large sample of tRNAs and compare this level to that of chimeric tRNAs constructed either within or between groups defined by phylogeny and/or specificity. We find that actual tRNAs have significantly more matches between the two halves than do random sequences that can form the tRNA structure, but there is no difference in the average level of matching between the two halves of an individual tRNA and the average level of matching between the two halves of the chimeric tRNAs in any of the sets we constructed. These results support the hypothesis that the modern tRNA cloverleaf arose from a single hairpin duplication prior to the divergence of modern tRNA specificities and the three domains of life.

Base Sequence↗

Phylogenetic analysis of evolutionary relationships of the planctomycete division of the domain bacteria based on amino acid sequences of elongation factor Tu.

Sequences from the tuf gene coding for the elongation factor EF-Tu were amplified and sequenced from the genomic DNA of Pirellula marina and Isosphaera pallida, two species of bacteria within the order Planctomycetales. A near-complete (1140-bp) sequence was obtained from Pi. marina and a partial (759-bp) sequence was obtained for I. pallida. Alignment of the deduced Pi. marina EF-Tu amino acid sequence against reference sequences demonstrated the presence of a unique 11-amino acid sequence motif not present in any other division of the domain Bacteria. Pi. marina shared the highest percentage amino acid sequence identity with I. pallida but showed only a low percentage identity with other members of the domain Bacteria. This is consistent with the concept of the planctomycetes as a unique division of the Bacteria. Neither primary sequence comparison of EF-Tu nor phylogenetic analysis supports any close relationship between planctomycetes and the chlamydiae, which has previously been postulated on the basis of 16S rRNA. Phylogenetic analysis of aligned EF-Tu amino acid sequences performed using distance, maximum-parsimony, and maximum-likelihood approaches yielded contradictory results with respect to the position of planctomycetes relative to other bacteria. It is hypothesized that long-branch attraction effects due to unequal evolutionary rates and mutational saturation effects may account for some of the contradictions.

Amino Acid Sequence↗

Diversity, distribution, and ancient taxonomic relationships within the TIR and non-TIR NBS-LRR resistance gene subfamilies.

Phylogenetic relationships among the NBS-LRR (nucleotide binding site-leucine-rich repeat) resistance gene homologues (RGHs) from 30 genera and nine families were evaluated relative to phylogenies for these taxa. More than 800 NBS-LRR RGHs were analyzed, primarily from Fabaceae, Brassicaceae, Poaceae, and Solanaceae species, but also from representatives of other angiosperm and gymnosperm families. Parsimony, maximum likelihood, and distance methods were used to classify these RGHs relative to previously observed gene subfamilies as well as within more closely related sequence clades. Grouping sequences using a distance cutoff of 250 PAM units (point accepted mutations per 100 residues) identified at least five ancient sequence clades with representatives from several plant families: the previously observed TIR gene subfamily and a minimum of four deep splits within the non-TIR gene subfamily. The deep splits in the non-TIR subfamily are also reflected in comparisons of amino acid substitution rates in various species and in ratios of nonsynonymous-to-synonymous nucleotide substitution rates ( K(A)/ K(S) values) in Arabidopsis thaliana. Lower K(A)/ K(S) values in the TIR than the non-TIR sequences suggest greater functional constraints in the TIR subfamily. At least three of the five identified ancient clades appear to predate the angiosperm-gymnosperm radiation. Monocot sequences are absent from the TIR subfamily, as observed in previous studies. In both subfamilies, clades with sequences separated by approximately 150 PAM units are family but not genus specific, providing a rough measure of minimum dates for the first diversification event within these clades. Within any one clade, particular taxa may be dramatically over- or underrepresented, suggesting preferential expansions or losses of certain RGH types within particular taxa and suggesting that no one species will provide models for all major sequence types in other taxa.

Amino Acid Sequence↗

Sarcocystis neurona major surface antigen gene 1 (SAG1) shows evidence of having evolved under positive selection pressure.

The major surface antigen gene 1 (SAG1) is conserved among members of Sarcocystidae and may play an important role in parasite pathogenesis. Additionally, generation and selection of different antigenic variants of SAG1 has the potential for inclusion in a subunit vaccine or in the development of a diagnostic assay. In this study, patterns of nucleotide polymorphism were used to test the hypothesis that natural selection promotes diversity in different parts of SAG1 of Sarcocystis neurona. Nucleotide and amino acid sequence analysis of SAG1 from multiple S. neurona isolates identified two alleles. Sequences were identical intra-allele and highly divergent inter-alleles. Also, phylogenetic reconstruction showed sequences clustering into two clades. Tajima's and Fu and Li's neutrality tests indicated that selection is more likely to be acting on SAG1. Moreover, a sliding window analysis based on the ratio of silent substitutions to amino acid replacements provided strong evidence that two short segments in the central and 3' domain of SAG1 have been under positive selection in the divergence of the two alleles, suggesting that it may be important for the evasion of host immune responses and would be a suitable target for vaccine development.

Animals↗

Divergence and conservation of the genomic RNAs of Taiwan and Hawaii strains of papaya ringspot potyvirus.

The complete nucleotide sequence of the genome of a Taiwan isolate of papaya ringspot potyvirus (PRSV YK) was determined from three overlapping cDNA clones and by direct RNA sequencing. Comparison was made with the reported Hawaii isolate of PRSV HA. Both genomes are 10,326 nucleotides long, excluding the poly(A)-tail. They encode a polyprotein of 3344 amino acids with a 5' leader of 85 nucleotides and a 3' non-translated region of 209 nucleotides. The two genomes share an overall nucleotide identity of 83.4% and an amino acid identity of 90.6%. The 3' non-translated regions show 92.3% identity. The first 23 nucleotides of the leaders are identical, while the remaining parts of the leaders only show 51.6% identity. The P1 protein genes of the two isolates are very different, with 70.9% nucleotide identity and 66.7% encoded amino acids identity. However, the other viral proteins of the two virus isolates are similar, with a 82.5-89.8% nucleotide identity of their genes and 91.2-97.6% amino acid identity, indicating that they are strains of the same potyvirus. Analysis of the ratios of nucleotide differences to the actual amino acid changes revealed that there are only 2.63 nucleotide changes for each amino acid change in the P1 protein, whereas for the other proteins 4.0-16.4 nucleotide changes are required for each amino acid replacement. The P1 protein has 58% of all the differences of polyprotein. The unusual variation in the leader sequences and the P1 proteins suggests that the two PRSV strains were derived from different evolutionary pathways in different geographic areas.

Amino Acid Sequence↗

Alcohol dehydrogenase of class III: consistent patterns of structural and functional conservation in relation to class I and other proteins.

Class III alcohol dehydrogenase from the lizard Uromastix hardwickii has been characterized. This non-mammalian, gnathostomatous vertebrate class III form allows correlations of structures and functions of this class, the traditional class I alcohol dehydrogenase, and other well-studied proteins. Catalytically, results show similar recoveries and activities of all vertebrate class III forms independent of source, similar activities also in invertebrates but in lower amounts, and considerably higher specific activities in microorganisms. Structurally, variability patterns are consistent throughout the vertebrate system with a ratio in accepted point mutations versus class I of 0.4. This ratio between different classes of a zinc enzyme is comparable to that between different heme proteins (cytochrome c and myoglobin), suggesting defined but non-identical functions also for the alcohol dehydrogenase classes.

Alcohol Dehydrogenase↗

Evolutionary relationship of hepatitis C, pesti-, flavi-, plantviruses, and newly discovered GB hepatitis agents.

Two flavivirus-like viruses, GB virus-A (GBV-A) and GB virus-B (GBV-B), were recently identified in the GB hepatitis agent, and are distinct from the hepatitis A to E viruses. The putative helicase domain of GBV-A and GBV-B was found to have amino acid sequence homology with hepatitis C virus (HCV), and distantly, is also related to pestiviruses, flaviviruses, and plant viruses. A phylogenetic tree construction showed that GBVs and HCV are closely related, and they are clustered with pestiviruses, flaviviruses and plant viruses in that order.

Amino Acid Sequence↗

Solubility of artificial proteins with random sequences.

A library of artificial random proteins of 141 amino acid residues of which 95 are random and which includes the 20 kinds of amino acids was prepared. Out of the 25 identified random proteins, 5 were soluble in the cell lysate, indicating that about 20% of the random proteins expressed in Escherichia coli are expected to be soluble. The soluble random proteins RP3-42 and RP3-45 and insoluble RP3-70 were purified. The solubility of the purified form is the same as that in the cell lysate.

Amino Acid Sequence↗

Neuroglobin, cytoglobin, and a novel, eye-specific globin from chicken.

Neuroglobin and cytoglobin are two recently discovered respiratory proteins of vertebrates. Here we report the first identification and expression analyses of these proteins in bird species. Neuroglobin from the domestic chicken Gallus gallus differs in approximately 30% from the mammalian proteins, but its genome structure shows the conservation of the B12.2, E11.0, and G7.0 intron positions. The chicken cytoglobin protein is shorter than the mammalian orthologs, from which it differs overall by approximately 25%, due to the absence of the C-terminal exon in the gene. Comparison of chicken and mammalian gene order shows that neuroglobin and cytoglobin are located on conserved syntenic chromosomal segments. While neuroglobin is expressed in the chicken's brain and eye, cytoglobin RNA was detected in all investigated tissues. In addition, a novel globin-type has been identified that is only expressed in the chicken's eye. The gene of this eye-globin contains the typical globin introns at B12.2 and G7.0. Phylogenetic analyses suggest that this globin is most closely related to the cytoglobin lineage. Although the function of this eye-globin remains presently uncertain, it adds an additional diversity to the vertebrate globin family.

Amino Acid Sequence↗

Cloning and characterization of the adipokinetic hormone receptor from the cockroach Periplaneta americana.

Cockroaches have long been used as insect models to investigate the actions of biologically active neuropeptides. Here, we describe the cloning and functional expression in Chinese hamster ovary cells of an adipokinetic hormone (AKH) G protein-coupled receptor from the cockroach Periplaneta americana. This receptor is only activated by various insect AKHs (we tested eight) and not by a library of 29 other insect or invertebrate neuropeptides and nine biogenic amines. Periplaneta has two intrinsic AKHs, Pea-AKH-1, and Pea-AKH-2. The Periplaneta AKH receptor is activated by low concentrations of both Pea-AKH-1 (EC50, 5 x 10(-9)M), and Pea-AKH-2 (EC50, 2 x 10(-9)M). Insects can be subdivided into two evolutionary lineages, holometabola (insects with a complete metamorphosis during development) and hemimetabola (incomplete metamorphosis). This paper describes the first AKH receptor from a hemimetabolous insect.

Amino Acid Sequence↗

Primary structure of a visual pigment in bullfrog green rods.

In frog retina there are special rod photoreceptor cells ('green rods') with physiological properties similar to those of typical vertebrate rods ('red rods'). A cDNA fragment encoding the putative green rod visual pigment was isolated from a retinal cDNA library of the bullfrog, Rana catesbeiana. Its deduced amino acid sequence has more than 65% identity with those of blue-sensitive cone pigments such as chicken blue and goldfish blue. Antisera raised against its C-terminal amino acid sequence recognized green rods. It is concluded that bullfrog green rods contain a visual pigment which is closely related to the blue-sensitive cone pigments of other non-mammalian vertebrates.

Amino Acid Sequence↗