Auto Caption Burner
Burn word-by-word animated captions into any video. Five styles, no watermark. Transcript and word timings included. $0.04 per 30-second block.
How it works
- 1Open it on Apify
Hit Run on Apify — it opens the tool in the cloud, no install.
- 2Set the inputs
Adjust
videoUrl,preset,language(sensible defaults are pre-filled). - 3Click Run
The tool runs on Apify’s cloud and collects the data for you.
- 4Export the results
Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.
Pricing
$0.04 per caption block (30s) = $40 per 1,000
| You are charged for | When | Price |
|---|---|---|
| Caption block (30s) | Per 30 seconds of captioned video. | $0.04 |
Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-08-08, and they are what you are actually charged.
Inputs
| Field | What it does | Type |
|---|---|---|
videoUrl | Public direct URL to the source video (.mp4 / .mov / .webm). You host it (S3/CDN/Drive direct link), no scraping is performed. | string |
preset | Visual style of the burned captions. | string |
language | ISO-639-1 code of the spoken language (e.g. en, es, fr). Leave 'auto' to auto-detect. | string |
allCaps | Force captions to uppercase (overrides the preset). | boolean |
fontSize | Override the preset font size. | integer |
primaryColor | Hex color of inactive words, e.g. #FFFFFF. | string |
highlightColor | Hex color of the currently-spoken word, e.g. #FFE000. | string |
position | Where captions sit on the frame. | string |
wordsPerGroup | How many words show on screen at once. | integer |
marginV | Distance from the chosen edge, at 1920px height. | integer |
crf | x264 CRF. Lower = higher quality/larger file. 18 = visually lossless, 23 = smaller. | integer |
openaiApiKey | Your OpenAI key for Whisper transcription. Kept private. Leave empty only if the actor owner has configured a shared key. | string |
transcriptionBaseUrl | OpenAI-compatible base URL for transcription. Default https://api.openai.com/v1. Use to point at a proxy. | string |
transcriptionModel | Whisper model name. Default whisper-1. | string |
What you get
A structured dataset — each result includes fields like:
okpresetwordCountdurationSecondsprocessingSecondsoutputExport every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.
Related tools in AI Video & Content Studio
Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.
Storyboard Video Generator
Turn images or a story into a Ken Burns slideshow video. Pan-and-zoom motion, optional audio, 9:16, 16:9 or 1:1 output. $0.015 per rendered second.
Story to Script Rewriter
Turn a story, article or Reddit post into a short-form script. You get a hook, tight narration, a title and extra hooks. $0.02 per script.
Subtitle Translator
Translate SRT and VTT subtitles into many languages in one run, or transcribe a video first. Timings stay exact. $0.05 per language, flat rate.
AI Thumbnail Generator
Generate video thumbnails with AI. A close-up face plus a bold hook headline. Sizes 9:16, 16:9 and 1:1. For YouTube, Shorts, Reels and ads.
Social Metadata Generator
Write titles, captions, hashtags, SEO tags and a pinned comment. For YouTube, TikTok, Reels, Shorts and X. $15.00 per 1,000 packs ($0.015 each).
Hook & Virality Scorer
Score any title, hook or script for viral potential. You get 0-100 with a grade, a breakdown and rewrite tips. No AI key. $3.00 per 1,000 hooks.
Where this tool sits
- Categories
- AI Video & Content Studio
Auto Caption Burner: word-by-word animated subtitles burned into your video
Send a direct link to a video you host. You get back an MP4 with captions burned into the picture, word by word, in one of five styles, plus an SRT file and the word timings as JSON.
Transcription runs on your own OpenAI key, which you paste into the input. That matters before you start: the captions cost what is listed below, and the Whisper transcription is billed to you by OpenAI separately. Leave the key out and the run writes a free sample row instead of doing any work.
| Input | A direct URL to a video you host, plus your OpenAI key |
| Output | A captioned MP4 in the key-value store, with an SRT and word timings |
| Ceiling | Around 13 minutes of audio, where transcription stops accepting the file |
| Account needed | Your own OpenAI key |
| Price | $0.04 per 30 seconds of video, flat on every plan |
🎬 What Auto Caption Burner does
It downloads your file, pulls the audio out, sends it to Whisper for word-level timings, then builds karaoke-style captions and re-encodes the video with them burned in. Burned in means the words are part of the picture, so they survive every upload and every repost. Nothing is watermarked.
Five presets, and they are real style differences rather than colour swaps: hormozi (white with a yellow pop on the spoken word), beast (white with green), tiktok (Bebas with red), clean (subtle) and karaoke (yellow and cyan). You can override the font size, both colours, the position, how many words sit on screen and how far from the edge they sit, all without leaving the preset.
Captions are built at your source dimensions, so a 1080x1920 vertical comes back 1080x1920. Nothing is cropped, letterboxed or re-framed.
You also get the SRT and a word-timing JSON, which is the part people underestimate. If you want to recut the video later, or feed the timings into another tool, that file has every word with its start and end.
📥 What you give it
{
"videoUrl": "https://cdn.example.com/my-short.mp4",
"openaiApiKey": "sk-...",
"preset": "hormozi",
"language": "en",
"allCaps": true,
"wordsPerGroup": 3
}
| Field | What the run uses if you leave it | What it is |
|---|---|---|
videoUrl | box starts with a placeholder | A public direct link to your own .mp4, .mov or .webm. S3, a CDN or a direct Drive link. Nothing is scraped. |
openaiApiKey | none | Your OpenAI key, stored as a secret. Without it the run returns a sample row and stops. |
preset | hormozi | hormozi, beast, tiktok, clean or karaoke. |
language | auto | An ISO-639-1 code like en, es, fr. auto lets Whisper work it out. |
allCaps | off | Force uppercase, overriding the preset. |
fontSize | the preset's | 24 to 200 px, measured at 1080 width. |
primaryColor, highlightColor | the preset's | Hex, like #FFFFFF and #FFE000. The highlight is the word being spoken. |
position | the preset's | top, center or bottom. |
wordsPerGroup | the preset's | 1 to 8 words on screen at once. |
marginV | the preset's | Distance from the chosen edge in pixels, measured at 1920 height. |
crf | 18 | Output quality. 18 is visually lossless and large, 23 is smaller. |
transcriptionBaseUrl | OpenAI's own | An OpenAI-compatible endpoint of your own, if you use one. |
transcriptionModel | whisper-1 | The Whisper model name. |
📤 What you get back
One dataset row per run, and the files themselves in the run's key-value store.
This is a real row from a run with no key supplied, which is the free sample the actor writes instead of working. On a real run the output keys are filled in and transcript holds your own text:
{
"ok": true,
"_sample": true,
"_notice": "SAMPLE OUTPUT - provide a public \"videoUrl\" and \"openaiApiKey\" for a real captioned video.",
"reason": "missing_api_key",
"videoUrl": "https://example.com/my-short.mp4",
"preset": "hormozi",
"sourceDimensions": { "width": 1080, "height": 1920 },
"durationSeconds": 12,
"wordCount": 18,
"transcript": "This sample row shows the output shape for automated checks and preview runs.",
"output": { "mp4Key": null, "srtKey": null, "wordsKey": null, "mp4Url": null, "srtUrl": null },
"processingSeconds": 0
}
| Field | What it is |
|---|---|
output.mp4Key, output.mp4Url | The captioned video in the key-value store, and a link to it. The key is captioned-<timestamp>.mp4. |
output.srtKey, output.srtUrl | The subtitle file, grouped five words to a cue, as captions-<timestamp>.srt. |
output.wordsKey | words-<timestamp>.json, holding every word with its start and end. Read it from the key-value store by that key. |
transcript | The whole spoken text as one string. |
wordCount | How many words Whisper returned. Zero words fails the run rather than handing you a silent video. |
durationSeconds | The source duration, rounded. This is what the caption blocks are counted from. |
sourceDimensions | The video's own width and height, which the output keeps. |
processingSeconds | How long the run took end to end. |
🧾 Reading the output
Two kinds of row can land in your dataset.
| Row | How to spot it | Charged |
|---|---|---|
| A finished video | no _sample key, and output.mp4Key filled in | yes |
| The sample row | _sample: true and a reason | no |
The sample row's reason says which piece was missing: missing_video_url or missing_api_key. Both are answered before anything is downloaded, so a run started to see what the actor does does no work and produces no caption blocks.
Anything that genuinely breaks fails the run rather than handing you a half-captioned file, and no caption blocks are charged for it. The captioned video is written to the store and the row pushed before billing, so if billing ever has a problem you still keep the video and the log says so.
▶️ How to run it
1. Open Auto Caption Burner and click Try for free. 2. Paste a direct link to your video into Video URL. It has to be a file link, not a page. 3. Paste your key into OpenAI API key. It is stored as a secret. 4. Pick a Caption style preset, and set ALL CAPS or Words per line if you want them. 5. Click Start, then open the Storage tab and download the captioned-*.mp4.
💰 How much does it cost?
$0.04 per 30 seconds of video. Flat on every Apify plan, no volume tiers. A 45-second short is two blocks, a three-minute video is six.
The count comes from the source duration, so a video is priced by its length, not by how many words it happens to contain. The free sample row is not charged, and a run that fails before the video is finished produces no caption blocks. Whisper transcription is billed to you by OpenAI on your own key and does not appear here.
💡 What people use it for
- Captioning shorts and reels so they still read with the sound off, which is most of the feed.
- Putting burned captions on ads, where an SRT sidecar file is not an option.
- Turning podcast and interview clips into social cuts, with the word timings kept for recutting.
- Running the same clip through two presets to see which holds attention better.
- Getting an SRT out of a video that never had one, as a side effect of captioning it.
🚧 What it does not do
- It does not fetch videos from social platforms. You give it a direct link to a file you host.
- Long videos fail. Transcription sends the whole audio track in one request and stops accepting
it somewhere around 13 minutes. Cut longer material into pieces first.
- No translation. Captions come out in the language that was spoken. Translate the SRT afterwards
if you need another language.
- No speaker labels and no diarisation. It is one stream of words.
- No reframing, no cropping, no resizing. The output keeps your source dimensions.
- No editing of the transcript before burning. What Whisper heard is what gets burned in.
- It needs your own OpenAI key. There is no shared key on this actor.
- Re-encoding takes real time on a long video, because the whole file is encoded, not just the
part with captions.
🧭 Which video tool do you need?
| If you want | Use |
|---|---|
| Captions burned into the picture | This one |
| Just the transcript and the timings, no video | Video Audio Transcriber |
| An existing subtitle file translated | Subtitle Translator |
| A long video cut into short clips | AI Viral Clip Cutter |
| A thumbnail for the finished video | Viral Thumbnail Generator |
❓ Questions people ask
Where do I put the video? Anywhere that serves a direct file link: S3, a CDN, a direct Drive link. The actor downloads it, it does not scrape a page.
Why do I need my own OpenAI key? Transcription is what makes word-by-word captions possible, and it runs on your key so the usage and the cost stay yours.
Will the captions be on top of the video or in a separate file? Both. Burned into the MP4, and an SRT beside it if you want to edit the wording later.
Can I change the font? You can change size, colours, position, spacing from the edge and words per line. The typefaces are the ones each preset ships with.
What happens with a silent video? Whisper returns no words and the run fails rather than handing you an unchanged file. No caption blocks are charged.
How long can the video be? Keep it under about 13 minutes of audio. Beyond that the transcription request is too large and the run will not finish.
🆘 If something breaks
Open the Issues tab on the actor page. Send the run ID and the video URL if it is one you can share. The log names the step that failed: download, probe, transcribe or burn.