Development of a multi-copy integration platform in Kluyveromyces marxianus enabled by a computational method for genome-wide identification of multi-copy integration loci.
Multi-copy integration is a core strategy for redirecting metabolic flux toward target compounds. However, its application has been hampered by the absence of methods for systematically identifying native multi-copy genomic loci. To overcome this, we developed a computational procedure for genome-wide identification of such loci. Theoretically, this method is potentially applicable to any genome-sequenced species as it only requires the genomic assembly of the target species as input. Applying the procedure to Kluyveromyces marxianus, we identified four groups of loci (KmCS1-4). Combining these loci-KmCS1-4 and the traditional 26S rDNA-with 14 markers with graded selection strengths, we established a versatile multi-copy integration toolkit comprising 70 plasmids. Each plasmid exhibits a unique integration pattern, collectively forming an integration profile. This profile serves as a manual, enabling users to select appropriate tools tailored to the expression requirements of rate-limiting enzymes in their pathways. Applying representative plasmids exhibiting low-, medium-, and high-copy integration patterns to lycopene biosynthesis modules resulted in lycopene titers of 3.5, 6.8 and 40.5 mg/L, corresponding to 2, 6 and 9 genomic copies, respectively, demonstrating a positive correlation between lycopene titers, genomic copy numbers and integration patterns, which highlights the versatility of the toolkit and its supporting manual. Our study not only provides a broadly applicable methodology for genome-wide identification of multi-copy loci, but also an efficient integration platform for K. marxianus.