Skip to the content.

Personalized Electrolaryngeal Voice Conversion with a Single Pre-operative Utterance

Abstact

Electrolaryngeal voice conversion (ELVC) enhances the mechanical speech of laryngectomees to improve naturalness and intelligibility. While personalized ELVC typically requires extensive pre-operative recordings for supervised training, this paper investigates a one-shot scenario where only one natural utterance is available for timbre preservation. We evaluate two strategies using zero-shot VC models: (1) a cascaded approach, performing ELVC followed by secondary conversion to the target timbre; and (2) a pseudo-target strategy, using a zero-shot model to generate synthetic training data for a supervised ELVC model. Results indicate that the cascaded approach outperforms the pseudo-target strategy, remarkably improving speaker similarity and quality with minimal impact on intelligibility. We also demonstrate that direct application of zero-shot models to ELVC fails due to domain mismatch, confirming the necessity of our proposed framework for one-shot personalized voice restoration.

VC Approaches

Illustration of the implemented one-shot personalized ELVC frameworks and refinement

The Pseudo-target Strategy

flowchart

The Cascade Approach

flowchart

Supervised Refinement of the Feature-level Cascade Approach

flowchart

Objective evaluation

Fow Quality:

For speech intelligibility:

For speaker similarity:

Table 1: Average performance across four speaker pairs. SER indicates Syllable Error Rate (%). Lower SER is better; higher values are better for other metrics.

table_1

Table 2: Subjective evaluation results for S1 vs S2. MOS scores are reported as 95% confidence interval. Win/Same rate show the percentages for listener choices in the ABX test.

table_2

Audio samples

Category:

sample 1

Category System NL01 BNL02 NL08 NL06
Input EL (Unprocessed)
Input Target speech
Baselines Supervised ELVC
Baselines Direct Zero-Shot (Seed-VC)
Baselines Direct Zero-Shot (Vevo)
Baselines Direct Zero-Shot (FreeVC)
Pseudo-Target + Seed-VC
Pseudo-Target + Vevo
Pseudo-Target + FreeVC
Cascade (Waveform) + Seed-VC
Cascade (Waveform) + Vevo
Cascade (Waveform) + FreeVC
Cascade (Feature) + Seed-VC
Cascade (Feature) + Vevo
Cascade (Feature) + FreeVC
Refined Cascade (Feature) + Seed-VC

sample 2

Category System NL01 BNL02 NL08 NL06
Input EL (Unprocessed)
Input Target speech
Baselines Supervised ELVC
Baselines Direct Zero-Shot (Seed-VC)
Baselines Direct Zero-Shot (Vevo)
Baselines Direct Zero-Shot (FreeVC)
Pseudo-Target + Seed-VC
Pseudo-Target + Vevo
Pseudo-Target + FreeVC
Cascade (Waveform) + Seed-VC
Cascade (Waveform) + Vevo
Cascade (Waveform) + FreeVC
Cascade (Feature) + Seed-VC
Cascade (Feature) + Vevo
Cascade (Feature) + FreeVC
Refined Cascade (Feature) + Seed-VC

Reference

  1. Li, F.R., Hwang, H.T., Yen, M.C., Lo, M.T., Tsao, Y., Wang, H.M.: Improving exemplar-based electrolaryngeal speech voice conversion via robust content representations. In: Proc. APSIPA (2025) ↩

  2. Seed-VC GitHub: https://github.com/Plachtaa/seed-vc ↩

  3. FreeVC GitHub: https://github.com/OlaWod/FreeVC ↩

  4. Vevo GitHub: https://github.com/open-mmlab/Amphion/tree/main/models/vc/vevo ↩