Skip to main content
Trevor.Dennis
Community Expert
Community Expert
July 6, 2026
Answered

Lip Sync New Video to Existing Audio.

  • July 6, 2026
  • 4 replies
  • 59 views

I have an audio track of my brother singing, and a 720P still image of him as a cartoon.

I would like to generate video which animates the image and lip syncs the cartoon with the existing audio.

I initially got this warning

 

But eventually got the process started with this prompt:

Sync the audio reference to the first frame reference image.  Animate the reference image to be playing the guitar, and lip sync the guitar player singing the song. . Maintain the original zoom ratio throughout. The audio to be unchanged. The facial expressions to match the lyrics of the song.

Unfortunately Firefly generated new audio, although it did a lovely job of syncing it to the visuals.

I used Seedance 2.0

The above 720P sized image as the First Image.

A 343Kb mp3 of the first 15 seconds of the song.

I have attached the resulting mp4.

Is it possible to do what I describe above?

Is there a secret to laying out the prompt?

    Correct answer Trevor.Dennis

    I’ve now done a ton of research, and this is the best result I found.  It uses video models outside of Firefly, and is quite complicated.  It is also quite expensive cost around US$200 in credits to create a 4 to 5 minute video with good lip sync

     

    Other worthy mentions are Jack Vs Ai   His latest upload (as I type this) being very interesting.  None of this is inside Firefly AFAIK.  The first link is using SeeDance 2.0 but my impression is that the partner model inside Firefly does not have all of the options of the stand alone model. 

    ElevenLabs deserves a mention, but like most of the lip sync models its models do not generate video lip synced to an uploaded audio track, but rather use the audio reference to learn the voice. You then have to provide a script in the prompt and the Ai will generate a talking head based on a first frame or image reference, and generate lip synced video that adheres to the references. 

     Free AI Voice Generator & Voice Agents Platform | ElevenLabs 

    This 15 second clip was generated using SeeDance 2.0 in Firefly and used 1250 credits.  The music is entirely random, but does sound more or less like my brother’s voice, and the lyrics are roughly aligned with those in the audio reference.

     

    Tarun Saini just to let you know that I would dearly love to be able to generate video based on a first frame reference, and that is lip synced to an uploaded audio track.  

    4 replies

    Trevor.Dennis
    Community Expert
    Trevor.DennisCommunity ExpertAuthorCorrect answer
    Community Expert
    July 7, 2026

    I’ve now done a ton of research, and this is the best result I found.  It uses video models outside of Firefly, and is quite complicated.  It is also quite expensive cost around US$200 in credits to create a 4 to 5 minute video with good lip sync

     

    Other worthy mentions are Jack Vs Ai   His latest upload (as I type this) being very interesting.  None of this is inside Firefly AFAIK.  The first link is using SeeDance 2.0 but my impression is that the partner model inside Firefly does not have all of the options of the stand alone model. 

    ElevenLabs deserves a mention, but like most of the lip sync models its models do not generate video lip synced to an uploaded audio track, but rather use the audio reference to learn the voice. You then have to provide a script in the prompt and the Ai will generate a talking head based on a first frame or image reference, and generate lip synced video that adheres to the references. 

     Free AI Voice Generator & Voice Agents Platform | ElevenLabs 

    This 15 second clip was generated using SeeDance 2.0 in Firefly and used 1250 credits.  The music is entirely random, but does sound more or less like my brother’s voice, and the lyrics are roughly aligned with those in the audio reference.

     

    Tarun Saini just to let you know that I would dearly love to be able to generate video based on a first frame reference, and that is lip synced to an uploaded audio track.  

    Community Expert
    July 6, 2026

    The video looks great ​@Trevor.Dennis.

    When I first started using Firefly, I posted a feature request for synching to audio. I haven’t tested this lately. Still good to see some progress has been made with your example even if though it’s not achieving exactly what you’re after.

    Sorry no tips. Just a like for what you’re doing and advocating for.

    Trevor.Dennis
    Community Expert
    Community Expert
    July 6, 2026

    @DeanUtian what I’m wondering Dean, is what the purpose of the audio reference is?  SeeDance 2.0 completely ignored it.  Kling 3.0 Omni has some sort of Audio reference, so I might try that.  

    Do you know if the Firefly Ai Assistant is what counts as the help documentation?  It’s very good, but I’m wondering if I could find out more with a deep dive into thje help pages if they exist.  

    I’m sending you a PM, so check your messages.

    Trevor.Dennis
    Community Expert
    Community Expert
    July 6, 2026

    OK, I have read what Google’s Ai summarizer says about it.  I have full Creative Cloud and Firefly Pro subscriptions and 8000 credits a month, so it looks doable.  Just a lot more complicated than I was hoping.  I’d still welcome any tips, even if just links to helpful videos.

    Thanks.