LongCat Video Avatar: Long Talking Avatar Video | Open Source | ComfyUI | Free & Unlimited
Also see ComfyUI Tutorial Series_ Ep01 - Introduction and Installation - PLAYLIST
and AI Filmmaking Part 1 _ Video gen with LTX 2.3_ You're doing it wrong.
and Claude Code + OpenRouter = Free UNLIMITED Coding (No RAM Needed)
Hi everyone! Today we’re diving into LongCat-Video — an exciting new open-source video generation model released by the Meituan LongCat Team. It’s a 13.6 billion-parameter foundational model that can generate videos from text, images, or even continue an existing video, all with one unified architecture.
Recently, the LongCat team introduced an important new direction called LongCat-Video-Avatar. This version focuses specifically on audio-driven character animation and expressive talking avatars.
LongCat-Video-Avatar is designed to generate natural and dynamic character motion driven directly by audio. Instead of only matching lip movements.
The model supports several native generation modes, including audio and text to video, audio, text, and image to video, as well as video continuation. It also works seamlessly with both single-stream and multi-stream audio inputs.
This means you can animate a single virtual presenter, create multi-character conversations, or drive complex scenes using layered audio, all within the same unified model.
One of the biggest strengths of LongCat-Video-Avatar is the stability of character identity. Even with longer audio clips, the avatar remains visually consistent while maintaining smooth and natural motion.
When combined with tools like ComfyUI, LongCat-Video-Avatar becomes a powerful part of a flexible AI video workflow. You can start from a single image, guide the style and personality with text, and let the audio bring the character to life.
LongCat-Video-Avatar GitHub Link:
🔗https://huggingface.co/meituan-longcat/LongCat-Video-Avatar
LongCat-Video GitHub Link:
🔗 https://github.com/meituan-longcat/LongCat-Video
Thank Kijai for following LongCat Video Avatar On ComfyUI:
🔗 Kijai’s ComfyUI Wan Wrapper (ComfyUI-WanVideoWrapper) :
https://github.com/kijai/ComfyUI-WanVideoWrapper
ComfyUI Workflow :
🔗 https://github.com/kijai/ComfyUI-WanVideoWrapper/blob/main/example\_workflows/LongCatAvatar\_audio\_image\_to\_video\_example\_01.json
✨Video You Might Enjoy:
🎬 WAN Animate :
https://youtu.be/8J7sP8JATMA?si=Qx27MlHQqGBkjkN-
🎬 Grok Imagine :
https://youtu.be/361qVMGn-cs?si=hVdSPqrlReGk0pd\_
Transcript
0:00 · Hi everyone. Today we're diving into LongCat video, an exciting new open-source video generation model released by the Mtoan Longat team. It's a 13.6 billion parameter foundational model that can generate videos from text, images, or even continue an existing video, all with one unified architecture. Recently, the Longat team introduced an important new direction called Longat video avatar. This version focuses specifically on audioddriven character animation and expressive talking avatars.
0:29 · Longat video avatar is designed to generate natural and dynamic character motion driven directly by audio. Instead of only matching lip movements, the model understands speech rhythm, emotion, and timing. Today, we'd like to show you comfy UI using the quantized model. First, you need to load the required model. Each model comes with a clear description and a direct download link, making it easy to locate and install the correct files without confusion. Next, upload your reference image.
0:58 · This image defines the appearance of the character and serves as the visual starting point for the animation.
1:06 · Simply upload your image here to proceed. After that, provide your audio input. The audio acts as the primary driving signal for the animation controlling lip movement, facial expressions, and motion.
1:19 · Please note that audio input requires additional supporting models. Detailed notes are provided in the interface to help you identify which audio models are needed and where to download them. To enhance the animation, you can increase the audio scale parameter for a stronger motion effect. Higher values result in more expressive facial movement and clearer synchronization with speech. The audio CFG option enables an extra model pass to further improve lip sync accuracy and facial detail.
1:46 · The recommended range is between three and five, which usually produces natural and stable results. The model uses an audio stride of two, meaning the output video runs at 16 frames per second while the input audio runs at 32 frames pers. If the output runs at 24 frames pers, then the input audio runs at 48 frames pers.
2:12 · Next, you will find the prompt box.
2:14 · Here, you can enter a text prompt to control the character style, emotion, or behavior. This allows you to fine-tune how the avatar looks and moves throughout the video. You can also adjust the video size by selecting your desired resolution. This gives you flexibility whether you want a portrait layout, landscape format, or a custom output size. Once everything is configured, simply process the video.
2:39 · After generation is complete, the final result will be displayed directly, allowing you to preview and review the animated video immediately. Next, you will see the extend section. This section allows you to add additional video segments to your project. By default, two extra segments are available, but you can easily add more if needed. To do this, simply copy the extend section and replicate the same node connections as the previous setup.
3:05 · Once connected, you can continue adding more video segments, enabling your avatar to appear in longer or multiple clips while maintaining consistent motion and appearance. Once everything is ready, upload your image and audio file. The default frame rate is 16 frames per second, so change it to 81 frames per second. Then click run to generate the video and view the final result. You can further enhance the animation by adding motion cues.
3:29 · We introduce longcat video, a foundational video generation model with 13.6B parameters, delivering strong performance across text to video, image to video, [music] and video continuation generation tasks or a moving background in the text prompt to make the result more dynamic. Please note that performance depends heavily on your system memory.
3:52 · With higher RAM, such as 24 GB, generating a single video segment may take around 5 minutes, while a 20- secondond video could take approximately 20 minutes. Systems with 12 GB of RAM may struggle significantly during generation. To sum it up, Long Cat Video Avatar represents a major step forward in open-source audioddriven avatar generation. It combines video, audio, text, and image inputs into a single unified system designed for expressive and natural character animation.
4:22 · If you're interested in AI avatars, next generation video tools, or building your own virtual presenters, LongCat Video Avatar is definitely a project to keep an eye on.