CapFlow
Captions written, timed and burned in without the clip leaving the phone.
Pick a clip and CapFlow transcribes it with a speech model that runs on the device, groups the words into readable caption lines, times them against the audio and renders them into the exported video. There is no upload step, so there is no clip sitting on somebody else's server.
Captioning tools almost all work the same way: upload the clip, wait, download it back. That is a reasonable trade for accuracy and a bad one for anything you would rather not hand over. CapFlow does the transcription on the phone instead.
Timing is the hard part
A transcript is not captions. The words arrive with timestamps, and turning those into something readable means deciding where lines break, how many words can sit on screen at once, and how long a line has to stay up to be read at all. That grouping is its own module with its own tests, separate from the recognition.
Burned in, not overlaid
The captions are rendered into the exported video. What you post is what you saw — there is no companion subtitle file to lose and no player that decides not to show them.
Capabilities
On-device recognition
The speech model is fetched once and runs locally from then on. The video never leaves the device.
Readable line grouping
Word timings are grouped into caption lines sized for the frame, so phrases are not split across cards.
Rendered into the export
Captions are drawn into the output video rather than shipped as a separate track.
Lifetime unlock
One non-consumable purchase removes the length limit and the ads. No subscription.
Screens
What it looks like in use.
Select any screen to view it full size. Use the arrow keys to move between screens and Escape to close.
Choose a clip from the library; nothing is uploaded. Caption lines grouped from word timings, editable before export. Captions rendered into the exported video file.