Clip maker: Long-form to short-form (with & without speech)
I’d love to see a long-form to short-form video clipping feature added to Adobe Firefly.
The feature would allow users to upload a longer video and automatically generate multiple short-form clips optimised for platforms such as Instagram Reels, TikTok, YouTube Shorts and Facebook.
Ideally, there would be two clipping modes:
1. Speech-Based Clipping
Firefly analyses dialogue, captions, key phrases and topic changes to identify the strongest moments within a longer video. It could automatically create short clips with suggested hooks, captions and different duration options.
2. Visual / Non-Speech Clipping
This is particularly important for content that doesn’t contain dialogue, such as ASMR, beauty, product videos, craft content, tutorials, cooking, cinematic footage, behind-the-scenes videos and aesthetic content.
Instead of relying on speech, Firefly could analyse:
-
Visual changes and scene transitions
-
Camera movement
-
Subject movement and actions
-
Interesting or satisfying moments
-
Changes in composition
-
Audio peaks, sound effects and music beats
-
Repetitive or low-interest sections that could be removed
-
Natural start and end points for short-form clips
For example, I regularly work with ASMR content where there may be no dialogue at all. I’d like to be able to upload a 4–10 minute video and have Firefly identify several visually or sonically engaging moments and turn them into 15, 30 or 60-second clips automatically.
It would also be useful to have:
-
Automatic 9:16 reframing and subject tracking
-
Multiple clip suggestions from one video
-
Adjustable clip length
-
“More like this” regeneration for a selected clip
-
Platform presets for Reels, Shorts and TikTok
-
Automatic captions when speech is present
-
Optional music/audio beat detection
-
Hook or opening-frame suggestions
-
Ability to manually refine the AI-selected in/out points
-
Batch export of generated clips
This would make Firefly much more useful as an end-to-end AI video workflow, particularly for creators and brands producing large amounts of social content.
A visual-first clipping mode would also differentiate it from many existing AI clipping tools, which currently depend heavily on podcasts, interviews and talking-head content.
