Speech Synthesis for AI Human
Deep Learning and Speech Processing based Projects
Deep Learning and Speech Processing based Projects
Natural Language Processing and Machine Learning based Projects
Data Preprocessing Tools based on Java and Python
Integrated Multimodal Prompt and Control for Emotional TTS
Published:
π Paper: Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation
Published:
π Paper: Glow: Generative Flow with Invertible 1x1 Convolutions
Published:
π Paper: AUTOVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss
Published:
π Paper: wav2vec: Unsupervised Pre-training for Speech Recognition
Published:
π Paper: vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
Published:
π Paper: Survey on Deep Neural Networks in Speech and Vision Systems
Published:
π Paper: Glow-TTS: A Generative Flow for Text-to-Speech Synthesis
π Summary: This paper introduces a flow-based model for TTS, improving robustness compared to Tacotron.
Published:
π Paper: wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
Published:
π Paper: Big Bird: Transformers for Longer Sequences
Published:
π Paper: PortaSpeech: Portable and High-Quality Generative Text-to-Speech
Published:
π Paper: YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for everyone
Published:
π Paper: Neural Grapheme-to-Phoneme Conversion with Pre-trained Grapheme Models
Published:
π Paper: Automatic Prosody Annotation with Pre-Trained Text-Speech Model
Published:
π Paper: Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Published:
π Paper: Collaborative Watermarking for Adversarial Speech Synthesis