SingingVoiceConversionviaSharedSpeakerSpaceandMin-Pooling AdversariallyEnhancedFlowMatching

Abstract

Singing Voice Conversion (SVC) faces a trade-off between singer feature disentanglement and singing quality. To address this, this paper proposes the MinFlow-SVC framework, a conditional flow matching-based SVC method enhanced by min-pooling adversarial training. We adopt a KNN-based approach to map source singer features into a shared singer space, removing source timbre to obtain content features like pitch, phonetics and singing expression. To further enhance generation quality, we introduce a min-pooling adversarial training strategy, which can detect and correct KNN-induced feature inconsistencies and improve conditional flow matching generation quality with harmonic awareness. Experiments show our method outperforms existing state-of-the-art SVC baselines in naturalness, intelligibility, timbre similarity and singing stability.

Comparative experiment

source audio reference audio So-Vits-SVC DiffSVC NeuCoSVC MinFlow-SVC-1 MinFlow-SVC-5 MinFlow-SVC-10

Ablation study

source audio reference audio MinFlow-SVC-10 w/o Content Encoder w/o L(min-pooling) w/o L(adv) w/o L(adv-harm)