Voice Privacy using CycleGAN and Time-Scale Modification

Prajapati, Gauri P; Singh, Dipesh Kumar; Amin, Preet P; Patil, Hemant

Publication:
Voice Privacy using CycleGAN and Time-Scale Modification

dc.contributor.affiliation	DA-IICT, Gandhinagar
dc.contributor.author	Prajapati, Gauri P
dc.contributor.author	Singh, Dipesh Kumar
dc.contributor.author	Amin, Preet P
dc.contributor.author	Patil, Hemant
dc.contributor.researcher	Prajapati, Gauri P (201911058)
dc.contributor.researcher	Singh, Dipesh Kumar (201911057)
dc.contributor.researcher	Amin, Preet P (201801051)
dc.date.accessioned	2025-08-01T13:09:02Z
dc.date.issued	01-07-2022
dc.description.abstract	Extensive use of Intelligent Personal Assistants (IPA) and�biometrics�in our day-to-day life asks for�privacy preservation�while dealing with�personal data. To that effect, efforts have been made to preserve the personally identifiable characteristics from human voice using different speaker�anonymization�techniques. In this paper, we propose Cycle Consistent�Generative Adversarial Network�(CycleGAN) to modify (transform) the speaker�s gender as well as the other prosodic aspects using their Mel�cepstral coefficients�(MCEPs) and fundamental frequency (i.e.,�). For effective anonymization in the context of voice privacy, we propose two-level (i.e., double) anonymization, where first-level anonymization is done using CycleGAN, followed by second-level anonymization using time-scale modification. The speaker anonymization and intelligibility are measured objectively using the automatic speaker verification (ASV) and�automatic speech recognition�(ASR) experiments, respectively, on development and test sets of�Librispeech�and�VCTK�datasets. For CycleGAN-based anonymization, the average % EERs (% WERs) are 40.3% (8.89%) and 40.95% (9.37%) with original enrollments and anonymized trials of the development and�test datasets, respectively. The average % EERs (% WERs) for double anonymization are 46.19% (9.95%) and 44.76% (10.34%) with original enrollments and anonymized trials of the development and test datasets, respectively. For the voice privacy evaluation , the performance of�ASV system�is much important, when the enrollments and trials both are anonymized (called as A-A case), which is also briefly discussed in this work. The average % EERs for A-A case (test set) are 24.29% and 2.81% using CycleGAN-based anonymization and double anonymization, respectively. Objective evaluation for more advanced attack model (i.e., attacker having anonymized data) is also explored in this study. The performance reflected the robustness of proposed anonymization approach towards voice privacy. The�subjective tests�using 101 listeners and corresponding analysis of variance (ANOVA) and Tukey�Kramer-based Ad-hoc tests are also carried out in order to quote statistical significance of our results. The subjective test show that the CycleGAN and double anonymization approaches give better naturalness, intelligibility, and speaker dissimilarity than the state-of-the-art x-vector-based baseline system.
dc.identifier.citation	Gauri P. Prajapati, Dipesh Kumar Singh, Preet P. Amin, and Patil, Hemant A, "Voice Privacy using CycleGAN and Time-Scale Modification," Computer Speech & Language, Elsevier, 29 Jan. 2022, 101353. ISSN: 0885-2308, doi:10.1016/j.csl.2022.101353. [Available online]
dc.identifier.doi	10.1016/j.csl.2022.101353
dc.identifier.issn	1095-8363
dc.identifier.scopus	2-s2.0-85124037725
dc.identifier.uri	https://ir.daiict.ac.in/handle/dau.ir/1559
dc.identifier.wos	WOS:000820203600007
dc.language.iso	en
dc.publisher	Elsevier
dc.relation.ispartofseries	Vol. 74; No.
dc.source	Computer Speech & Language
dc.source.uri	https://www.sciencedirect.com/science/article/pii/S0885230822000031?via%3Dihub
dc.title	Voice Privacy using CycleGAN and Time-Scale Modification
dspace.entity.type	Publication
relation.isAuthorOfPublication	fdb7041b-280e-498b-b2ee-34f9bc351f4c
relation.isAuthorOfPublication.latestForDiscovery	fdb7041b-280e-498b-b2ee-34f9bc351f4c

Collections

Journal Article

Publication: Voice Privacy using CycleGAN and Time-Scale Modification

Files

Collections

Publication:
Voice Privacy using CycleGAN and Time-Scale Modification