Abstract
The goal of speech anonymization is to modify the attributes of a speaker, such as age and gender, to hide their identity. While such modifications are effective in obfuscating the speakers' identity, they generally reduce the quality of the synthesized speech. In this study, we analyze the trade-off between anonymization and synthesis quality using controlled modifications of age and gender attributes and measure their effect using objective measures of speaker similarity and audio quality and linear regression. We validate these findings through listening tests of identity retention, and quality preferences. We find that anonymization performance increases with the magnitude of the modification, though at the expense of degradations in speech quality. However, we find there is an optimum range of modifications that balances identity suppression with naturalness.
Method
Given that our goal is speech anonymization in real-time, we used a parametric approach to modify specific speaker attributes, such as age and gender, within latent speaker embeddings. First, we applied principal component analysis (PCA) to the speaker embeddings, then we identified PC directions in the latent space that correlate with the desired attributes. For this purpose, we computed the Pearson correlation coefficient ρ of each PCA dimension with the speaker attributes (e.g., age).
Then, we used these correlation coefficients as weights to construct a composite direction Vage that captures the primary variance associated with age:
Vage = w1 PC1 + w2 PC2 + ⋯
where, wi = ρ(PCi, age). To modify the attribute in the embedding, we adjust the original embedding Z along this direction:
Z' = Z + λVage
where λ controls the extent of modification: positive λ increases the attribute (e.g., age), while negative λ decreases it. This method enables fine-grained control over age or gender within speaker embeddings by moving in attribute-correlated directions in latent space, as illustrated in the figure below. Note that, by identifying these latent vectors along the directions of highest variance in the data, the approach is robust to noise. Audio samples of the results of our approach to attribute editing are available in a footnote.
Notes
- Dataset (CMU-ARCTIC corpus): http://www.festvox.org/cmu_arctic/
- Synthesis Model: End-to-end Streaming model for Low-Latency Speech Anonymization [1] [Arxiv]
Audio Samples
Below are audio samples of the results of our approach to attribute editing in speaker embeddings.
| Speaker | Feminine -- | Original | Feminine ++ |
|---|---|---|---|
| RMS (Male) | |||
| BDL (Male) | |||
| SLT (Female) | |||
| CLB (Female) |
| Speaker | Age -- | Original | Age ++ |
|---|---|---|---|
| RMS (Male) | |||
| BDL (Male) | |||
| SLT (Female) | |||
| CLB (Female) |
References
[1] W. Quamer et al., "End-to-end streaming model for low-latency speech anonymization," in IEEE SLT 2024.