SingingVoiceConversionviaSharedSpeakerSpaceandMin-Pooling
AdversariallyEnhancedFlowMatching
Abstract
Singing Voice Conversion (SVC) faces a trade-off between singer feature disentanglement and singing quality.
To address this, this paper proposes the MinFlow-SVC framework, a conditional flow matching-based SVC method
enhanced by min-pooling adversarial training. We adopt a KNN-based approach to map source singer features
into a shared singer space, removing source timbre to obtain content features like pitch, phonetics and
singing expression. To further enhance generation quality, we introduce a min-pooling adversarial training
strategy, which can detect and correct KNN-induced feature inconsistencies and improve conditional flow
matching generation quality with harmonic awareness. Experiments show our method outperforms existing
state-of-the-art SVC baselines in naturalness, intelligibility, timbre similarity and singing stability.