Sitemap
A list of all the posts and pages found on the site. For you robots out there is an XML version available for digesting as well.
Pages
Posts
projects
Speech Synthesis for AI Human
Deep Learning and Speech Processing based Projects
Data Analysis Projects
Natural Language Processing and Machine Learning based Projects
Data Engineering
Data Preprocessing Tools based on Java and Python
research
Controllable Emotional TTS
Integrated Multimodal Prompt and Control for Emotional TTS
reviews
Conv-TasNet Paper Review
Published:
π Paper: Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation
Glow Paper Review
Published:
π Paper: Glow: Generative Flow with Invertible 1x1 Convolutions
AutoVC Paper Review
Published:
π Paper: AUTOVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss
wav2vec Paper Review
Published:
π Paper: wav2vec: Unsupervised Pre-training for Speech Recognition
VQ-wav2vec Paper Review
Published:
π Paper: vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
Survey on DNN in Speech and Vision Paper Review
Published:
π Paper: Survey on Deep Neural Networks in Speech and Vision Systems
Glow-TTS Paper Review
Published:
π Paper: Glow-TTS: A Generative Flow for Text-to-Speech Synthesis
π Summary: This paper introduces a flow-based model for TTS, improving robustness compared to Tacotron.
wav2vec 2.0 Paper Review
Published:
π Paper: wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
BigBird Paper Review
Published:
π Paper: Big Bird: Transformers for Longer Sequences
PortaSpeech Paper Review
Published:
π Paper: PortaSpeech: Portable and High-Quality Generative Text-to-Speech
YourTTS Paper Review
Published:
π Paper: YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot Voice Conversion for everyone
Neural Grapheme-to-Phoneme Paper Review
Published:
π Paper: Neural Grapheme-to-Phoneme Conversion with Pre-trained Grapheme Models
Automatic Prosody Annotation Paper Review
Published:
π Paper: Automatic Prosody Annotation with Pre-Trained Text-Speech Model
VALL-E Paper Review
Published:
π Paper: Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Watermarking for TTS Paper Review
Published:
π Paper: Collaborative Watermarking for Adversarial Speech Synthesis
