Time synced Text to speech with captions
It would be nice to be able to generate a voiceover from captions in order to replace poor quality voice recordings and improve consistency of videos. A part of my job is to generate voiceovers using outside tools in order to create instructional videos recorded by different people on different quality microphones, so that the final videos have a consistent audio quality regardless of who recorded it. It’s a slow task, and it occurred to me that given Adobe’s recent focus on AI QoL upgrades, this might be a good feature to add. Especially because we already have the audio timing created by the captions.
