PubMed Health⌕ Search

Biomedical subjects

Roy J Britten

Publications and source records attributed to Roy J Britten.

5 recordsLinked to original sources

Almost all human genes resulted from ancient duplication.

Results of protein sequence comparison at open criterion show a very large number of relationships that have, up to now, gone unreported. The relationships suggest many ancient events of gene duplication. It is well known that gene duplication has been a major process in the evolution of genomes. A collection of human genes that have known functions have been examined for a history of gene duplications detected by means of amino acid sequence similarity by using BLASTp with an expectation of two or less (open criterion). Because the collection of genes in build 35 includes sets of transcript variants, all genes of known function were collected, and only the longest transcription variant was included, yielding a 13,298-member library called KGMV (for known genes maximum variant). When all lengths of matches are accepted, >97% of human genes show significant matches to each other. Many form matches with a large number of other different proteins, showing that most genes are made up from parts of many others as a result of ancient events of duplication. To support the use of the open criterion, all of the members of the KGMV library were twice replaced with random protein sequences of the same length and average composition, and all were compared with each other with BLASTp at expectation two or less. The set of matches averaged 0.35% of that observed for the KGMV set of proteins.

Amino Acid Sequence↗

The majority of human genes have regions repeated in other human genes.

Amino acid sequence comparisons have been made between all of 25,193 human proteins with each of the others by using blast software (National Center for Biotechnology Information) and recording the results for regions that are significantly related in sequence, that is, have an expectation of <1 x 10(-3). The results are presented for each amino acid as the number of identical or similar amino acids matched in these aligned regions. This approach avoids summing or dealing directly with the different regions of any one protein that are often related to different numbers and types of other proteins. The results are presented graphically for a sample of 140 proteins. Relationships are not observed for 26.5% of the 12,728,866 amino acids. The average number of related amino acids is 36.5 for the majority (73.5%) that show relationships. The median number of recognized relationships is approximately 3 for all of the amino acids, and the maximum number is 718. The results demonstrate the overwhelming importance of gene regional duplication forming families of proteins with related domains and show the variety of the resulting patterns of relationship. The magnitude of the set of relationships leads to the conclusion that the principal process by which new gene functions arise has been by making use of preexisting genes.

Alu Elements↗

Coding sequences of functioning human genes derived entirely from mobile element sequences.

Among all of the many examples of mobile elements or "parasitic sequences" that affect the function of the human genome, this paper describes several examples of functioning genes whose sequences have been almost completely derived from mobile elements. There are many examples where the synthetic coding sequences of observed mRNA sequences are made up of mobile element sequences, to an extent of 80% or more of the length of the coding sequences. In the examples described here, the genes have named functions, and some of these functions have been studied. It appears that each of the functioning genes was originally formed from mobile elements and that in some process of molecular evolution a coding sequence was derived that could be translated into a protein that is of some importance to human biology. In one case (AD7C), the coding sequence is 99% made up of a cluster of Alu sequences. In another example, the gene BNIP3 coding sequence is 97% made up of sequences from an apparent human endogenous retrovirus. The Syncytin gene coding sequence appears to be made from an endogenous retrovirus envelope gene.

Genome, Human↗

Majority of divergence between closely related DNA samples is due to indels.

It was recently shown that indels are responsible for more than twice as many unmatched nucleotides as are base substitutions between samples of chimpanzee and human DNA. A larger sample has now been examined and the result is similar. The number of indels is approximately 1/12th of the number of base substitutions and the average length of the indels is 36 nt, including indels up to 10 kb. The ratio (R(u)) of unpaired nucleotides attributable to indels to those attributable to substitutions is 3.0 for this 2 million-nt chimp DNA sample compared with human. There is similar evidence of a large value of R(u) for sea urchins from the polymorphism of a sample of Strongylocentrotus purpuratus DNA (R(u) = 3-4). Other work indicates that similarly, per nucleotide affected, large differences are seen for indels in the DNA polymorphism of the plant Arabidopsis thaliana (R(u) = 51). For the insect Drosophila melanogaster a high value of R(u) (4.5) has been determined. For the nematode Caenorhabditis elegans the polymorphism data are incomplete but high values of R(u) are likely. Comparison of two strains of Escherichia coli O157:H7 shows a preponderance of indels. Because these six examples are from very distant systematic groups the implication is that in general, for alignments of closely related DNA, indels are responsible for many more unmatched nucleotides than are base substitutions. Human genetic evidence suggests that indels are a major source of gene defects, indicating that indels are a significant source of evolutionary change.

Animals↗

Divergence between samples of chimpanzee and human DNA sequences is 5%, counting indels.

Five chimpanzee bacterial artificial chromosome (BAC) sequences (described in GenBank) have been compared with the best matching regions of the human genome sequence to assay the amount and kind of DNA divergence. The conclusion is the old saw that we share 98.5% of our DNA sequence with chimpanzee is probably in error. For this sample, a better estimate would be that 95% of the base pairs are exactly shared between chimpanzee and human DNA. In this sample of 779 kb, the divergence due to base substitution is 1.4%, and there is an additional 3.4% difference due to the presence of indels. The gaps in alignment are present in about equal amounts in the chimp and human sequences. They occur equally in repeated and nonrepeated sequences, as detected by REPEATMASKER (http://ftp.genome.washington.edu/RM/RepeatMasker.html).

Animals↗