UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding

AI-generated keywords: UniCATS

AI-generated Key Points

The license of the paper does not allow us to build upon its content and the key points are generated using the paper metadata rather than the full article.

  • Authors Chenpeng Du, Yiwei Guo, Feiyu Shen, Zhijun Liu, Zheng Liang, Xie Chen, Shuai Wang, Hui Zhang, and Kai Yu introduce UniCATS framework for text-to-speech synthesis.
  • UniCATS addresses limitations of existing models like VALL-E and SPEAR-TTS in speech editing due to left-to-right generation constraints and limited audio quality of acoustic tokens.
  • UniCATS framework consists of two main components: CTX-txt2vec and CTX-vec2wav.
  • CTX-txt2vec utilizes contextual VQ-diffusion for predicting semantic tokens from input text.
  • CTX-vec2wav employs contextual vocoding to convert semantic tokens into waveforms while considering the acoustic context.
  • Experimental results show that CTX-vec2wav outperforms models like HifiGAN and AudioLM in speech resynthesis from semantic tokens.
  • UniCATS achieves state-of-the-art performance in speech continuation and editing tasks while leveraging contextual information.
Also access our AI generated: Comprehensive summary, Lay summary, Blog-like article; or ask questions about this paper to our AI assistant.

Authors: Chenpeng Du, Yiwei Guo, Feiyu Shen, Zhijun Liu, Zheng Liang, Xie Chen, Shuai Wang, Hui Zhang, Kai Yu

Accepted to AAAI 2024

Abstract: The utilization of discrete speech tokens, divided into semantic tokens and acoustic tokens, has been proven superior to traditional acoustic feature mel-spectrograms in terms of naturalness and robustness for text-to-speech (TTS) synthesis. Recent popular models, such as VALL-E and SPEAR-TTS, allow zero-shot speaker adaptation through auto-regressive (AR) continuation of acoustic tokens extracted from a short speech prompt. However, these AR models are restricted to generate speech only in a left-to-right direction, making them unsuitable for speech editing where both preceding and following contexts are provided. Furthermore, these models rely on acoustic tokens, which have audio quality limitations imposed by the performance of audio codec models. In this study, we propose a unified context-aware TTS framework called UniCATS, which is capable of both speech continuation and editing. UniCATS comprises two components, an acoustic model CTX-txt2vec and a vocoder CTX-vec2wav. CTX-txt2vec employs contextual VQ-diffusion to predict semantic tokens from the input text, enabling it to incorporate the semantic context and maintain seamless concatenation with the surrounding context. Following that, CTX-vec2wav utilizes contextual vocoding to convert these semantic tokens into waveforms, taking into consideration the acoustic context. Our experimental results demonstrate that CTX-vec2wav outperforms HifiGAN and AudioLM in terms of speech resynthesis from semantic tokens. Moreover, we show that UniCATS achieves state-of-the-art performance in both speech continuation and editing.

Submitted to arXiv on 13 Jun. 2023

Ask questions about this paper to our AI assistant

You can also chat with multiple papers at once here.

The license of the paper does not allow us to build upon its content and the AI assistant only knows about the paper metadata rather than the full article.

AI assistant instructions?

Results of the summarizing process for the arXiv paper: 2306.07547v6

This paper's license doesn't allow us to build upon its content and the summarizing process is here made with the paper's metadata rather than the article.

, , , , In their paper titled "UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding," authors Chenpeng Du, Yiwei Guo, Feiyu Shen, Zhijun Liu, Zheng Liang, Xie Chen, Shuai Wang, Hui Zhang, and Kai Yu introduce a novel approach to text-to-speech synthesis. They address the limitations of existing models such as VALL-E and SPEAR-TTS in terms of speech editing due to their left-to-right generation constraints and reliance on acoustic tokens with limited audio quality. The proposed UniCATS framework consists of two main components: CTX-txt2vec and CTX-vec2wav. <kw>CTX-txt2vec:</kw> This component utilizes contextual VQ-diffusion to predict semantic tokens from input text. This allows for seamless concatenation with surrounding context and incorporation of semantic information. <kw>CTX-vec2wav:</kw> On the other hand, this component employs contextual vocoding to convert these semantic tokens into waveforms while considering the acoustic context. Experimental results demonstrate that CTX-vec2wav outperforms existing models like HifiGAN and AudioLM in terms of speech resynthesis from semantic tokens. The authors' work has been accepted for presentation at AAAI 2024, showcasing the significance of their contributions to advancing text-to-speech technology. Their proposed UniCATS framework not only achieves state-of-the-art performance in both speech continuation and editing tasks but also addresses the limitations of existing models. With its unified approach and use of contextual information, UniCATS shows great potential for improving text-to-speech synthesis.
Created on 27 Jun. 2024

Assess the quality of the AI-generated content by voting

Score: 0

Why do we need votes?

Votes are used to determine whether we need to re-run our summarizing tools. If the count reaches -10, our tools can be restarted.

Similar papers summarized with our AI tools

Navigate through even more similar papers through a

tree representation

Look for similar papers (in beta version)

By clicking on the button above, our algorithm will scan all papers in our database to find the closest based on the contents of the full papers and not just on metadata. Please note that it only works for papers that we have generated summaries for and you can rerun it from time to time to get a more accurate result while our database grows.

Disclaimer: The AI-based summarization tool and virtual assistant provided on this website may not always provide accurate and complete summaries or responses. We encourage you to carefully review and evaluate the generated content to ensure its quality and relevance to your needs.