Personalized Electrolaryngeal Voice Conversion with a Single Pre-operative Utterance
Abstact
Electrolaryngeal voice conversion (ELVC) enhances the mechanical speech of laryngectomees to improve naturalness and intelligibility. While personalized ELVC typically requires extensive pre-operative recordings for supervised training, this paper investigates a one-shot scenario where only one natural utterance is available for timbre preservation. We evaluate two strategies using zero-shot VC models: (1) a cascaded approach, performing ELVC followed by secondary conversion to the target timbre; and (2) a pseudo-target strategy, using a zero-shot model to generate synthetic training data for a supervised ELVC model. Results indicate that the cascaded approach outperforms the pseudo-target strategy, remarkably improving speaker similarity and quality with minimal impact on intelligibility. We also demonstrate that direct application of zero-shot models to ELVC fails due to domain mismatch, confirming the necessity of our proposed framework for one-shot personalized voice restoration.
VC Approaches
- LLE1: Locally Linear Embedding with the following features and their usage:
- Weight estimation and reconstruction:
- For Cascade (Waveform) and Pseudo-Target Strategy: features extracted from the 6th layer of WavLM Large
- For Cascade (Feature) and Refined Cascade (Feature): features extracted from either WavLM-Large, HuBERT, or Whisper-small
- Neighbor search and alignment: Chinese-HuBERT Large
- Weight estimation and reconstruction:
- Seed-VC2: : Seed-VC v1, a zero-shot diffusion-based VC model
- FreeVC3: A High-quality, text-free, one-shot VC model
- Vevo4: A zero-shot VC framework that decouples style, timbre, and linguistic content (we used only the flow-matching transformer component)
- ELXX/BYELXX: EL speaker id
- NLXX/BNLXX: NL speakers id (Male: NL01, NL08, BNL02; Female: NL06)
Illustration of the implemented one-shot personalized ELVC frameworks and refinement
The Pseudo-target Strategy

The Cascade Approach

Supervised Refinement of the Feature-level Cascade Approach

Objective evaluation
Fow Quality:
- SpeechBERTScore: objective metric for measuring semantic fidelity
- UTMOS: objective metric for predicted perceptual quality
For speech intelligibility:
- SER: syllable error rate of a given speech calculated using a self-trained syllable recognition system
For speaker similarity:
- Similarity: cosine similarity between the converted and reference utterances, calculated with a resemblyzer in the Amphion toolkit
Table 1: Average performance across four speaker pairs. SER indicates Syllable Error Rate (%). Lower SER is better; higher values are better for other metrics.

Table 2: Subjective evaluation results for S1 vs S2. MOS scores are reported as 95% confidence interval. Win/Same rate show the percentages for listener choices in the ABX test.

Audio samples
Category:
- source speech: speech from EL01, BYEL02, EL08, EL06
- target speech: speech from NL01, BNL02, NL08, NL06
- For the tables below, the rows are the systems implemented, and the columns are the target speakers used for conversion.
sample 1
- transcription: 他捐了很多衣物給災區 (ta juan le hen duo yi wu gei zai qu)
| Category | System | NL01 | BNL02 | NL08 | NL06 |
|---|---|---|---|---|---|
| Input | EL (Unprocessed) | ||||
| Input | Target speech | ||||
| Baselines | Supervised ELVC | ||||
| Baselines | Direct Zero-Shot (Seed-VC) | ||||
| Baselines | Direct Zero-Shot (Vevo) | ||||
| Baselines | Direct Zero-Shot (FreeVC) | ||||
| Pseudo-Target | + Seed-VC | ||||
| Pseudo-Target | + Vevo | ||||
| Pseudo-Target | + FreeVC | ||||
| Cascade (Waveform) | + Seed-VC | ||||
| Cascade (Waveform) | + Vevo | ||||
| Cascade (Waveform) | + FreeVC | ||||
| Cascade (Feature) | + Seed-VC | ||||
| Cascade (Feature) | + Vevo | ||||
| Cascade (Feature) | + FreeVC | ||||
| Refined Cascade (Feature) | + Seed-VC |
sample 2
- transcription: 電視報導那裡發生地震 (dian shi bao dao na li fa sheng di zhen)
| Category | System | NL01 | BNL02 | NL08 | NL06 |
|---|---|---|---|---|---|
| Input | EL (Unprocessed) | ||||
| Input | Target speech | ||||
| Baselines | Supervised ELVC | ||||
| Baselines | Direct Zero-Shot (Seed-VC) | ||||
| Baselines | Direct Zero-Shot (Vevo) | ||||
| Baselines | Direct Zero-Shot (FreeVC) | ||||
| Pseudo-Target | + Seed-VC | ||||
| Pseudo-Target | + Vevo | ||||
| Pseudo-Target | + FreeVC | ||||
| Cascade (Waveform) | + Seed-VC | ||||
| Cascade (Waveform) | + Vevo | ||||
| Cascade (Waveform) | + FreeVC | ||||
| Cascade (Feature) | + Seed-VC | ||||
| Cascade (Feature) | + Vevo | ||||
| Cascade (Feature) | + FreeVC | ||||
| Refined Cascade (Feature) | + Seed-VC |
Reference
-
Li, F.R., Hwang, H.T., Yen, M.C., Lo, M.T., Tsao, Y., Wang, H.M.: Improving exemplar-based electrolaryngeal speech voice conversion via robust content representations. In: Proc. APSIPA (2025) ↩
-
Seed-VC GitHub: https://github.com/Plachtaa/seed-vc ↩
-
FreeVC GitHub: https://github.com/OlaWod/FreeVC ↩
-
Vevo GitHub: https://github.com/open-mmlab/Amphion/tree/main/models/vc/vevo ↩