Word Level Timestamps for Karaoke and TikTok Captions
Per-word start and end times alongside SRT and VTT files, for captions that pop word by word. Needs your own OpenAI key to return a real transcript.
What this run is set to
The settings saved on this example. Every one of them is editable once you open it on Apify — these are a starting point, not a limit.
| Setting | Value | What it does |
|---|---|---|
Media URL | https://cdn.example.com/reels/clip-music.mp4 | Public URL to a video or audio file (mp4, mov, mp3, wav, m4a, webm). |
Language | en | Spoken language ISO code, or 'auto' to detect. |
Include word timestamps | on | Return per-word start/end times (great for karaoke captions). |
Output files | srt, vtt | Which subtitle/text files to also produce: srt, vtt, txt. |
Pricing
$0.02 per transcribed minute = $20 per 1,000
| You are charged for | When | Price |
|---|---|---|
| Transcribed minute | Per minute of audio transcribed. | $0.02 |
Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-08-08, and they are what you are actually charged.
What you get
A structured dataset — each result includes fields like:
_demo_noticedurationSecondslanguagesegmentCountsegmentssourceUrlsrtKeytextvttKeywordCountsrtUrlvttUrlExport every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.
More use cases for Video & Audio Transcriber — Word-Level + SRT/VTT
Built on Video & Audio Transcriber — Word-Level + SRT/VTT — every input, output field, price and the full how-to are on the tool page.