Samples from a voice-conversion system evaluated across three conversion tasks, each converting one attribute of a recording while keeping the spoken content fixed. For every example: Source is the original recording, Target reference is a real recording carrying the target attribute, and Generated is the model’s conversion of the source content into that target. The italic line above each row is the source sentence, which the generated clip preserves.

Voice (timbre) conversion

Converting a source speaker's voice into a different target speaker's voice, content unchanged. Source and target clips are from LibriTTS.

SourceTarget referenceGenerated
“If I should not be in, wait for me.”
Speaker 2238 Speaker 5126 Speaker 5126
“What is his name?”
Speaker 2238 Speaker 5126 Speaker 5126
“In the old days she had dressed for her own sake to look pretty and be admired.”
Speaker 5126 Speaker 5322 Speaker 5322
“'God willing I'll find a way to repay you,' he said, finishing his wine.”
Speaker 5322 Speaker 8057 Speaker 8057
“Therefore it is necessary, by a final examination of their characteristics, to eliminate those features which are hostile to sociability.”
Speaker 8534 Speaker 6415 Speaker 6415
“If public utility had interfered, that forest-the only one for miles around-would still be standing.”
Speaker 8534 Speaker 8057 Speaker 8057
“So the government thinks.”
Speaker 8534 Speaker 6415 Speaker 6415
“There are thousands of documents, even official documents, to prove this, if necessary.”
Speaker 8534 Speaker 2238 Speaker 2238

Accent conversion

Converting a source speaker's regional accent into a different target accent, content unchanged. Source and target clips are VCTK recordings.

SourceTarget referenceGenerated
“She can scoop these things into three red bags, and we will go meet her Wednesday at the train station.”
English Irish Irish
“When a man looks for something beyond his reach, his friends say he is looking for the pot of gold at the end of the rainbow.”
English American American
“There is a solution, she believes.”
English Scottish Scottish
“The rainbow is a division of white light into many beautiful colors.”
Irish English English
“If the red of the second bow falls upon the green of the first, the result is to give a bow with an abnormally wide yellow band, since red and green light when mixed form yellow.”
Irish Scottish Scottish
“The difference in the rainbow depends considerably upon the size of the drops, and the width of the colored band increases as the size of the drops increases.”
American Scottish Scottish
“It is a simple equation.”
Scottish English English
“When a man looks for something beyond his reach, his friends say he is looking for the pot of gold at the end of the rainbow.”
Scottish Irish Irish

Style conversion

Converting a source speaking style (happy, sad, laughing, whisper, enunciated, default) into a different target style, content unchanged. Source and target clips are from Expresso.

SourceTarget referenceGenerated
“So, what was he really doing?”
Sad Whisper Whisper
“With exercising our first amendment rights?”
Happy Sad Sad
“With exercising our first amendment rights?”
Sad Enunciated Enunciated
“Yes, I have the doctor's appointment and checking the coffee on my to do list.”
Happy Default Default
“I was in Australia actually.”
Laughing Default Default
“We did a movie called Desperados for Netflix last year.”
Enunciated Laughing Laughing
“I'm nineteen and this is nineteen ninety eight.”
Sad Laughing Laughing
“The Norsemen considered the rainbow as a bridge over which the gods passed from earth to their home in the sky.”
Sad Enunciated Enunciated

Part of ongoing doctoral work on domain-robust voice conversion. See the project overview, or the cross-speaker pair corpora used to train the underlying diffusion voice-conversion model.