PubMed Health⌕ Search

Biomedical subjects

Wen-Hsiung Li

Publications and source records attributed to Wen-Hsiung Li.

At least 19 recordsLinked to original sources

Human-specific insertions and deletions inferred from mammalian genome sequences.

It has been suggested that insertions and deletions (indels) have contributed to the sequence divergence between the human and chimpanzee genomes more than do nucleotide changes (3% vs. 1.2%). However, although there have been studies of large indels between the two genomes, no systematic analysis of small indels (i.e., indels </= 100 bp) has been published. In this study, we first estimated that the false-positive rate of small indels inferred from human-chimpanzee pairwise sequence alignments is quite high, suggesting that the chimpanzee genome draft is not sufficiently accurate for our purpose. We have therefore inferred only human-specific indels using multiple sequence alignments of mammalian genomes. We identified >840,000 "small" indels, which affect >7000 UCSC-annotated human genes (>11,000 transcripts). These indels, however, amount to only approximately 0.21% sequence change in the human lineage for the regions compared, whereas in pseudogenes indels contribute to a sequence divergence of 1.40%, suggesting that most of the indels that occurred in genic regions have been eliminated. Functional analysis reveals that the genes whose coding exons have been affected by human-specific indels are enriched in transcription and translation regulatory activities but are underrepresented in catalytic and transporter activities, cellular and physiological processes, and extracellular region/matrix. This functional bias suggests that human-specific indels might have contributed to human unique traits by causing changes at the RNA and protein level.

Animals↗

Computational reconstruction of transcriptional regulatory modules of the yeast cell cycle.

BACKGROUND: A transcriptional regulatory module (TRM) is a set of genes that is regulated by a common set of transcription factors (TFs). By organizing the genome into TRMs, a living cell can coordinate the activities of many genes and carry out complex functions. Therefore, identifying TRMs is helpful for understanding gene regulation. RESULTS: Integrating gene expression and ChIP-chip data, we develop a method, called MOdule Finding Algorithm (MOFA), for reconstructing TRMs of the yeast cell cycle. MOFA identified 87 TRMs, which together contain 336 distinct genes regulated by 40 TFs. Using various kinds of data, we validated the biological relevance of the identified TRMs. Our analysis shows that different combinations of a fairly small number of TFs are responsible for regulating a large number of genes involved in different cell cycle phases and that there may exist crosstalk between the cell cycle and other cellular processes. MOFA is capable of finding many novel TF-target gene relationships and can determine whether a TF is an activator or/and a repressor. Finally, MOFA refines some clusters proposed by previous studies and provides a better understanding of how the complex expression program of the cell cycle is regulated. CONCLUSION: MOFA was developed to reconstruct TRMs of the yeast cell cycle. Many of these TRMs are in agreement with previous studies. Further, MOFA inferred many interesting modules and novel TF combinations. We believe that computational analysis of multiple types of data will be a powerful approach to studying complex biological systems when more and more genomic resources such as genome-wide protein activity data and protein-protein interaction data become available.

Algorithms↗

Protein complexity, gene duplicability and gene dispensability in the yeast genome.

Using functional genomic and protein structural data we studied the effects of protein complexity (here defined as the number of subunit types in a protein) on gene dispensability and gene duplicability. We found that in terms of gene duplicability the major distinction in protein complexity is between hetero-complexes, each of which includes at least two different types of subunits (polypeptides), and homo-complexes, which include monomers and complexes that consist of only subunits of one polypeptide type. However, gene dispensability decreases only gradually as the number of subunit types in a protein complex increases. These observations suggest that the dosage balance hypothesis can explain well gene duplicability of complex proteins, but cannot completely explain the difference in dispensabilities between hetero-complex subunits. It is likely that knocking out a gene coding for a hetero-complex subunit would disrupt the function of the whole complex, so that the deletion effect on fitness would increase with protein complexity. We also found that multi-domain polypeptide genes are less dispensable but more duplicable than single-domain polypeptide genes. Duplicate genes derived from the whole genome duplication event in yeast are more dispensable (except for ribosomal protein genes) than other duplicate genes. Further, we found that subunits of the same protein complex tend to have similar expression levels and similar effects of gene deletion on fitness. Finally, we estimated that in yeast the contribution of duplicate genes to genetic robustness against null mutation is approximately 9%, smaller than previously estimated. In yeast, protein complexity may serve as a better indicator of gene dispensability than do duplicate genes.

Computational Biology↗

Codon-usage bias versus gene conversion in the evolution of yeast duplicate genes.

Many Saccharomyces cerevisiae duplicate genes that were derived from an ancient whole-genome duplication (WGD) unexpectedly show a small synonymous divergence (K(S)), a higher sequence similarity to each other than to orthologues in Saccharomyces bayanus, or slow evolution compared with the orthologue in Kluyveromyces waltii, a non-WGD species. This decelerated evolution was attributed to gene conversion between duplicates. Using approximately 300 WGD gene pairs in four species and their orthologues in non-WGD species, we show that codon-usage bias and protein-sequence conservation are two important causes for decelerated evolution of duplicate genes, whereas gene conversion is effective only in the presence of strong codon-usage bias or protein-sequence conservation. Furthermore, we find that change in mutation pattern or in tDNA copy number changed codon-usage bias and increased the K(S) distance between K. waltii and S. cerevisiae. Intriguingly, some proteins showed fast evolution before the radiation of WGD species but little or no sequence divergence between orthologues and paralogues thereafter, indicating that functional conservation after the radiation may also be responsible for decelerated evolution in duplicates.

Codon↗

Radical amino acid change versus positive selection in the evolution of viral envelope proteins.

To detect positive selection in protein-coding sequence evolution, the ratio of the nonsynonymous to synonymous substitution rate (K(A)/K(S)) is commonly used. When this ratio is higher than 1, positive selection on nonsynonymous changes is considered to have occurred. However, the question of what kinds of amino acid change are likely to be involved in positive selection has not been well studied, though intuitively it seems that radical changes frequently occur in positively selected changes. To address this question, we examined chemically radical and conservative replacements in the evolution of hepatitis C virus (HCV) protein sequences. In the envelope region, 34 positively and 440 negatively selected sites were identified by the K(A)/K(S) ratio. Radical and conservative changes were compared between the two types of selected sites using two methods. First, the numbers of radical and conservative replacements were counted at the positively and negatively selected sites according to three kinds of chemical classifications. In all three classifications, the resulting ratios of the two numbers were not statistically different for the two types of selected sites (P>0.05). Second, the distribution of chemical changes was compared between the two types of selected sites using two kinds of chemical distances. The distributions of the two chemical distances were not statistically different for the two types of selected sites (P>0.05). These results indicate that the ratio of chemically radical and conservative changes is similar for positively and negatively selected sites in the envelope protein of HCV or, in other words, there is no correlation between radical change and positive selection in the evolution of this protein.

Amino Acid Substitution↗

Are GC-rich isochores vanishing in mammals?

Several studies of nucleotide substitution patterns in mammalian species suggested that GC-rich isochores might be vanishing in mammalian genomes. However, the number of genes and the number of genomes included in these studies might not have given a reliable broad view of the trend in GC change in mammals. It is therefore worth exploiting this issue with a broader coverage of mammalian genomes using a reliable approach, the maximum likelihood approach. We have applied two maximum likelihood methods to infer the ancestral GC contents of 176 mammalian genes from representative eutherian species and at least one marsupial species. Except for a large GC decrease in marsupial genes, we found no general decreasing trend in GC content in GC-rich genes or in other genes among eutherian mammals; indeed, the GC content of GC-rich genes appears to have increased in recent times in some genomes, e.g., the rabbit. For the large GC decrease in marsupials, it could be mainly due to the great reduction in chromosome number, which could lead to a large reduction in recombination rate and thus also a large reduction in the rate of gene conversion. Since many eutherian mammals still maintain a fairly large number of chromosomes, it is unlikely that GC-rich isochores are vanishing in these mammals.

Animals↗

Nucleotide variation and haplotype diversity in a 10-kb noncoding region in three continental human populations.

Noncoding regions are usually less subject to natural selection than coding regions and so may be more useful for studying human evolution. The recent surveys of worldwide DNA variation in four 10-kb noncoding regions revealed many interesting but also some incongruent patterns. Here we studied another 10-kb noncoding region, which is in 6p22. Sixty-six single-nucleotide polymorphisms were found among the 122 worldwide human sequences, resulting in 46 genotypes, from which 48 haplotypes were inferred. The distribution patterns of DNA variation, genotypes, and haplotypes suggest rapid population expansion in relatively recent times. The levels of polymorphism within human populations and divergence between humans and chimpanzees at this locus were generally similar to those for the other four noncoding regions. Fu and Li's tests rejected the neutrality assumption in the total sample and in the African sample but Tajima's test did not reject neutrality. A detailed examination of the contributions of various types of mutations to the parameters used in the neutrality tests clarified the discrepancy between these test results. The age estimates suggest a relatively young history in this region. Combining three autosomal noncoding regions, we estimated the long-term effective population size of humans to be 11,000 +/- 2800 using Tajima's estimator and 17,600 +/- 4700 using Watterson's estimator and the age of the most recent common ancestor to be 860,000 +/- 258,000 years ago.

Africa↗

Method for identifying transcription factor binding sites in yeast.

MOTIVATION: Identifying transcription factor binding sites (TFBSs) is helpful for understanding the mechanism of transcriptional regulation. The abundance and the diversity of genomic data provide an excellent opportunity for identifying TFBSs. Developing methods to integrate various types of data has become a major trend in this pursuit. RESULTS: We develop a TFBS identification method, TFBSfinder, which utilizes several data sources, including DNA sequences, phylogenetic information, microarray data and ChIP-chip data. For a TF, TFBSfinder rigorously selects a set of reliable target genes and a set of non-target genes (as a background set) to find overrepresented and conserved motifs in target genes. A new metric for measuring the degree of conservation at a binding site across species and methods for clustering motifs and for inferring position weight matrices are proposed. For synthetic data and yeast cell cycle TFs, TFBSfinder identifies motifs that are highly similar to known consensuses. Moreover, TFBSfinder outperforms well-known methods. AVAILABILITY: http://cg1.iis.sinica.edu.tw/~TFBSfinder/.

Algorithms↗

Reorganization of adjacent gene relationships in yeast genomes by whole-genome duplication and gene deletion.

In Saccharomyces, an ancient whole-genome duplication (WGD) and widespread duplicate gene deletion resulted in extensive reorganization of adjacent gene relationships. We have studied the evolution of adjacent gene pairs' identity, orientation, and spacing following whole-genome duplication and deletion (WGD-D) using comparative genomic analyses and simulations. Surveying adjacent gene organization across the Saccharomyces species complex, we find a genome-wide bias toward divergently and convergently transcribed gene pairs in all species but a reduction in this bias in the species that underwent WGD-D. Among neutral models of WGD-D, only single-gene deletion can produce the appropriate reduction in orientation bias and recapitulate the pattern of short, highly dispersed deletions we observe in Saccharomyces cerevisiae. To characterize the dynamics of WGD-D, we trace the conservation and creation of adjacent gene pairs along the S. cerevisiae lineage. We find that newly created adjacencies have a tandem orientation bias, while adjacencies conserved from prior to WGD-D have the same divergent-convergent bias as found in the species that diverged before WGD. We also find that adjacent gene pairs produced by WGD-D gained greater intergenic spacing but that this is reduced in the older adjacencies. Given this, and the preponderance of short deleted blocks, we argue that the deletion phase of WGD-D occurred primarily by small inactivating mutations followed by numerous small deletions. Newly created adjacent gene pairs also have an initial increase in mean log2 expression ratios and maximal expression levels, suggesting that increased intergenic spacing caused a genome-wide reduction in transcriptional interference.

Evolution, Molecular↗

Role of positive selection in the retention of duplicate genes in mammalian genomes.

The question of how duplicate genes are retained in a population remains controversial. The duplication-degeneration-complementation model, which involves no positive selection, stipulates a higher retention rate of duplicate genes in a small population than in a large one. This model has been accepted by many evolutionists. However, we found considerably more retentions and fewer losses of duplicate genes in the mouse genome than in the human genome, although the population size of rodents is in general larger than that of primates. Indeed, in nearly every interval of synonymous divergence between duplicate genes, the number of gene retentions in mouse is larger than that in human. Our findings suggest a more important role of positive selection in duplicate retention than duplication-degeneration-complementation. In addition, certain functional categories show a higher tendency of lineage-specific expansion than expected, suggesting lineage-specific selection or functional bias in retained duplicates.

Animals↗

Patterns of expansion and expression divergence in the plant polygalacturonase gene family.

BACKGROUND: Polygalacturonases (PGs) belong to a large gene family in plants and are believed to be responsible for various cell separation processes. PG activities have been shown to be associated with a wide range of plant developmental programs such as seed germination, organ abscission, pod and anther dehiscence, pollen grain maturation, fruit softening and decay, xylem cell formation, and pollen tube growth, thus illustrating divergent roles for members of this gene family. A close look at phylogenetic relationships among Arabidopsis and rice PGs accompanied by analysis of expression data provides an opportunity to address key questions on the evolution and functions of duplicate genes. RESULTS: We found that both tandem and whole-genome duplications contribute significantly to the expansion of this gene family but are associated with substantial gene losses. In addition, there are at least 21 PGs in the common ancestor of Arabidopsis and rice. We have also determined the relationships between Arabidopsis and rice PGs and their expression patterns in Arabidopsis to provide insights into the functional divergence between members of this gene family. By evaluating expression in five Arabidopsis tissues and during five stages of abscission, we found overlapping but distinct expression patterns for most of the different PGs. CONCLUSION: Expression data suggest specialized roles or subfunctionalization for each PG gene member. PGs derived from whole genome duplication tend to have more similar expression patterns than those derived from tandem duplications. Our findings suggest that PG duplicates underwent rapid expression divergence and that the mechanisms of duplication affect the divergence rate.

Arabidopsis↗

Alternatively and constitutively spliced exons are subject to different evolutionary forces.

There has been a controversy on whether alternatively spliced exons (ASEs) evolve faster than constitutively spliced exons (CSEs). Although it has been noted that ASEs are subject to weaker selective constraints than CSEs, so they evolve faster, there have also been studies that indicated slower evolution in ASEs than in CSEs. In this study, we retrieve more than 5,000 human-mouse orthologous exons and calculate the synonymous (KS) and nonsynonymous (KA) substitution rates in these exons. Our results show that ASEs have higher KA values and higher KA/KS ratios than CSEs, indicating faster amino acid-level evolution in ASEs. The faster evolution may be in part due to weaker selective constraints. It is also possible that the faster rate is in part due to faster functional evolution in ASEs. On the other hand, the majority of ASEs have lower KS values than CSEs. With reference to the substitution rate in introns, we show that the KS values in ASEs are close to the neutral substitution rate, whereas the synonymous substitution rate in CSEs has likely been accelerated. The elevated synonymous rate in CSEs is not related to CpG dinucleotides or low-complexity regions of protein but may be weakly related to codon usage bias. The overall trends of higher KA and lower KS in ASEs than in CSEs are also observed in human-rat and mouse-rat comparisons. Therefore, our observations hold for mammals of different molecular clocks.

Animals↗

Dynamic modeling of cis-regulatory circuits and gene expression prediction via cross-gene identification.

BACKGROUND: Gene expression programs depend on recognition of cis elements in promoter region of target genes by transcription factors (TFs), but how TFs regulate gene expression via recognition of cis elements is still not clear. To study this issue, we define the cis-regulatory circuit of a gene as a system that consists of its cis elements and the interactions among their recognizing TFs and develop a dynamic model to study the functional architecture and dynamics of the circuit. This is in contrast to traditional approaches where a cis-regulatory circuit is constructed by a mutagenesis or motif-deletion scheme. We estimate the regulatory functions of cis-regulatory circuits using microarray data. RESULTS: A novel cross-gene identification scheme is proposed to infer how multiple TFs coordinate to regulate gene transcription in the yeast cell cycle and to uncover hidden regulatory functions of a cis-regulatory circuit. Some advantages of this approach over most current methods are that it is based on data obtained from intact cis-regulatory circuits and that a dynamic model can quantitatively characterize the regulatory function of each TF and the interactions among the TFs. Our method may also be applicable to other genes if their expression profiles have been examined for a sufficiently long time. CONCLUSION: In this study, we have developed a dynamic model to reconstruct cis-regulatory circuits and a cross-gene identification scheme to estimate the regulatory functions of the TFs that control the regulation of the genes under study. We have applied this method to cell cycle genes because the available expression profiles for these genes are long enough. Our method not only can quantify the regulatory strengths and synergy of the TFs but also can predict the expression profile of any gene having a subset of the cis elements studied.

Cell Cycle↗

Evidence from opsin genes rejects nocturnality in ancestral primates.

It is firmly believed that ancestral primates were nocturnal, with nocturnality having been maintained in most prosimian lineages. Under this traditional view, the opsin genes in all nocturnal prosimians should have undergone similar degrees of functional relaxation and accumulated similar extents of deleterious mutations. This expectation is rejected by the short-wavelength (S) opsin gene sequences from 14 representative prosimians. We found severe defects of the S opsin gene only in lorisiforms, but no defect in five nocturnal and two diurnal lemur species and only minor defects in two tarsiers and two nocturnal lemurs. Further, the nonsynonymous-to-synonymous rate ratio of the S opsin gene is highest in the lorisiforms and varies among the other prosimian branches, indicating different time periods of functional relaxation among lineages. These observations suggest that the ancestral primates were diurnal or cathemeral and that nocturnality has evolved several times in the prosimians, first in the lorisiforms but much later in other lineages. This view is further supported by the distribution pattern of the middle-wavelength (M) and long-wavelength (L) opsin genes among prosimians.

Adaptation, Biological↗

Statistical methods for identifying yeast cell cycle transcription factors.

Knowing transcription factors (TFs) involved in the yeast cell cycle is helpful for understanding the regulation of yeast cell cycle genes. We therefore developed two methods for predicting (i) individual cell cycle TFs and (ii) synergistic TF pairs. The essential idea is that genes regulated by a cell cycle TF should have higher (lower, if it is a repressor) expression levels than genes not regulated by it during one or more phases of the cell cycle. This idea can also be used to identify synergistic interactions of TFs. Applying our methods to chromatin immunoprecipitation data and microarray data, we predict 50 cell cycle TFs and 80 synergistic TF pairs, including most known cell cycle TFs and synergistic TF pairs. Using these and published results, we describe the behaviors of 50 known or inferred cell cycle TFs in each cell cycle phase in terms of activation/repression and potential positive/negative interactions between TFs. In addition to the cell cycle, our methods are also applicable to other functions.

Cell Cycle↗

Expression divergence between duplicate genes.

A general picture of the role of expression divergence in the evolution of duplicate genes is emerging, thanks to the availability of completely sequenced genomes and functional genomic data, such as microarray data. It is now clear that expression divergence, regulatory-motif divergence and coding-sequence divergence all increase with the age of duplicate genes, although their exact interrelationships remain to be determined. It is also clear that gene duplication increases expression diversity and enables tissue or developmental specialization to evolve. However, the relative roles of subfunctionalization and neofunctionalization in the retention of duplicate genes remain to be clarified, especially for higher eukaryotes. In addition, the relationship between gene duplication and evolution of transcriptional regulatory networks is largely unexplored.

Evolution, Molecular↗

Protein function, connectivity, and duplicability in yeast.

Protein-protein interaction networks have evolved mainly through connectivity rewiring and gene duplication. However, how protein function influences these processes and how a network grows in time have not been well studied. Using protein-protein interaction data and genomic data from the budding yeast, we first examined whether there is a correlation between the age and connectivity of yeast proteins. A steady increase in connectivity with protein age is observed for yeast proteins except for those that can be traced back to Eubacteria. Second, we investigated whether protein connectivity and duplicability vary with gene function. We found a higher average duplicability for proteins interacting with external environments than for proteins localized within intracellular compartments. For example, proteins that function in the cell periphery (mainly transporters) show a high duplicability but are lowly connected. Conversely, proteins that function within the nucleus (e.g., transcription, RNA and DNA metabolisms, and ribosome biogenesis and assembly) are highly connected but have a low duplicability. Finally, we found a negative correlation between protein connectivity and duplicability.

Evolution, Molecular↗