hackernews client

BruceWok

a month ago

LatentSync is a cutting-edge, open-source lip-synchronization framework powered by Audio-Conditioned Latent Diffusion Models. By integrating Whisper audio embeddings with advanced temporal alignment (TREPA), it transforms arbitrary audio and video inputs into photorealistic, high-resolution (512x512) talking head videos.

LatentSync1.6, an end-to-end lip-sync method

1 Comments

BruceWok