Video & Audio Transcriber — Word-Level + SRT/VTT
Transcribe any video or audio URL with word-level timestamps. Download SRT, VTT and TXT. The language is found for you. $0.02 per transcribed minute.
How it works
- 1Open it on Apify
Hit Run on Apify — it opens the tool in the cloud, no install.
- 2Set the inputs
Adjust
mediaUrl,mediaUrls,language(sensible defaults are pre-filled). - 3Click Run
The tool runs on Apify’s cloud and collects the data for you.
- 4Export the results
Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.
Pricing
$0.02 per transcribed minute = $20 per 1,000
| You are charged for | When | Price |
|---|---|---|
| Transcribed minute | Per minute of audio transcribed. | $0.02 |
Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-08-08, and they are what you are actually charged.
Inputs
| Field | What it does | Type |
|---|---|---|
mediaUrl | Public URL to a video or audio file (mp4, mov, mp3, wav, m4a, webm). Use this for a single file, or mediaUrls for a batch. | string |
mediaUrls | Transcribe several files in one run, one dataset row per URL. Each item is a public video/audio URL. | array |
language | Spoken language ISO code, or 'auto' to detect. | string |
wordTimestamps | Return per-word start/end times (great for karaoke captions). | boolean |
outputFormats | Which subtitle/text files to also produce: srt, vtt, txt. | array |
openaiApiKey | Your OpenAI (Whisper) key. Kept private. | string |
model | Transcription model. Default whisper-1. | string |
baseUrl | OpenAI-compatible base URL. Default https://api.openai.com/v1. | string |
What you get
A structured dataset — each result includes fields like:
_demo_noticedurationSecondslanguagesegmentCountsegmentssourceUrlsrtKeytextvttKeywordCountsrtUrlvttUrlExport every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.
3 ready-to-run use cases
MP4 to SRT: Subtitle File From a Video URL
Point it at a public MP4 and get a timed SRT back, language detected automatically. Needs your own OpenAI key; without one you get a sample row.
Transcribe MP3 to Text: Podcast Episode Transcript
An episode URL in, clean text out, split into timed segments for show notes. Needs your own OpenAI key; without one you get a labelled sample row.
Word Level Timestamps for Karaoke and TikTok Captions
Per-word start and end times alongside SRT and VTT files, for captions that pop word by word. Needs your own OpenAI key to return a real transcript.
Related tools in YouTube & Creator Tools
Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.
AI Character Reference Bank
Make a consistent AI character reference set. You get portrait, 3/4, full body, side and expression sheets, plus a character-bible prompt.
YouTube Channel Scraper
Scrape videos from any YouTube channel, no login or API key. Rows carry titles, views, dates, duration, thumbnails.
YouTube Comments Scraper
Scrape YouTube comments with no login or key. Get text, likes, reply counts, author handle, verified badge and timestamp. $0.40 per 1,000.
YouTube Playlist Scraper
Export every video in a YouTube playlist. Get title, watch URL, views, age, runtime and thumbnail. No API key or OAuth. $0.40 per 1,000 videos.
YouTube Transcript Scraper
YouTube transcripts from links or a keyword search: text, timed segments, SRT, VTT and video details. No API key. $0.80 per 1,000.
YouTube Trending Scraper
Get the most-viewed YouTube videos for any topic, no login or key. Rows carry views, titles, channels, duration, thumbnails. $0.40 per 1,000 videos.
Where this tool sits
Video & Audio Transcriber: transcripts with word-level timing, from any media URL
Give it a public link to a video or an audio file and get back the full text, sentence-level segments, and a start and end time for every single word. It also writes ready-made .srt, .vtt and .txt files into the run's storage.
You bring your own OpenAI key, so the transcription itself is billed to you by OpenAI on top of what you pay here.
| Input | One public media URL, or a list of them. mp4, mov, webm, mp3, wav, m4a |
| Output | One row per file, with the transcript, segments, words and file links |
| Ceiling | About 5 GB per file at the default memory. Long recordings are the real limit, see below |
| Account needed | Your own OpenAI API key |
| Price | $0.02 per transcribed minute, flat on every plan |
🎙️ What Video & Audio Transcriber does
It downloads the file you point it at, pulls the speech out as compressed mono audio, and sends that for transcription. What comes back is the plain text, the segments with their timings, and the word list with a start and end time on each word. That word list is what karaoke-style captions need, and most transcript tools do not hand it over.
Give it mediaUrls instead of mediaUrl and it walks the list, writing one row per file. A file that fails does not stop the ones after it.
The language is detected for you unless you name it. Naming it is usually a little more accurate on short or noisy clips.
📥 What you give it
{
"mediaUrl": "https://example.com/podcast.mp3",
"language": "auto",
"wordTimestamps": true,
"outputFormats": ["srt", "vtt", "txt"],
"openaiApiKey": "sk-..."
}
| Field | Default | What it is |
|---|---|---|
mediaUrl | none | A public, direct link to one video or audio file. |
mediaUrls | none | A list, for a batch. One dataset row per URL. You can use this instead of mediaUrl or alongside it. |
language | auto | ISO code of the spoken language, or auto to detect it. |
wordTimestamps | true | Adds the per-word words array to the row. Turn it off for a much smaller row on long files. |
outputFormats | ["srt", "vtt"] when the field is absent, box starts at srt, vtt, txt | Which files to write into the run's storage. |
openaiApiKey | none | Your own key. Marked secret, so it is not stored with the run input. |
model | whisper-1 | Any transcription model your key can reach. |
baseUrl | https://api.openai.com/v1 | Point it at any OpenAI-compatible endpoint. |
The URL has to be the file itself, not a page that plays it. A link ending in .mp4 or .mp3 works; a video page does not.
📤 What you get back
A real row, with the long arrays and URLs cut short:
{
"ok": true,
"sourceUrl": "https://example.com/podcast.mp3",
"language": "en",
"text": "Welcome back to the show. Today we cover the deep ocean.",
"wordCount": 11,
"segmentCount": 2,
"durationSeconds": 8,
"segments": [
{ "start": 0, "end": 4, "text": "Welcome back to the show." },
{ "start": 4, "end": 8, "text": "Today we cover the deep ocean." }
],
"words": [{ "word": "Welcome", "start": 0, "end": 0.4 }, "..."],
"srtKey": "transcript-1757178142775-0.srt",
"srtUrl": "https://api.apify.com/v2/key-value-stores/.../records/transcript-...",
"vttUrl": "https://api.apify.com/v2/key-value-stores/.../records/transcript-..."
}
| Field | What it is |
|---|---|
words | One entry per word with its own start and end, in seconds. Present only when wordTimestamps is on. |
segments | Sentence-sized cues. This is what the .srt and .vtt files are built from. |
durationSeconds | Where the last speech segment ends, rounded. This is the length that gets billed, not the file's own runtime. |
text | The whole transcript as one string. |
srtUrl, vttUrl, txtUrl | Download links, present only for the formats you asked for. |
sourceUrl | Which input URL this row belongs to. On a batch this is the only way to tell rows apart. |
🧾 Reading the output
Three kinds of row can land in your dataset.
| Row | How to spot it | Charged a minute |
|---|---|---|
| A transcript | ok: true and a sourceUrl | yes |
| A failed file | ok: false and an error string | no |
| The sample row | _demo: true | no |
Check ok before you count rows. A failed file still writes a row, so a batch of ten can show ten rows with only six transcripts among them. The error field on those rows says what went wrong in plain words: a download that failed, a file over the size limit, or no speech in the audio.
The default table view hides sourceUrl and error. Switch the dataset to All fields, or export as JSON, when you are working through a batch.
You get the _demo sample row instead of real work when no media URL was given, or no key was.
▶️ How to run it
1. Open Video & Audio Transcriber and click Try for free. 2. Paste a direct file link into Media URL, or a list into Media URLs (batch). 3. Put your key into OpenAI API key (BYO). 4. Leave Include word timestamps on if you want karaoke captions, then click Start. 5. Read the rows in the dataset, or open the run's Key-value store tab for the subtitle files.
💰 How much does it cost?
$0.02 per transcribed minute. Flat on every Apify plan, no volume tiers. A 12 minute clip is 12 minutes, rounded up to the next whole minute, with one minute as the smallest charge per file.
The length billed is the speech in the file, measured to where the last segment ends. A file that fails and a sample row are not billed any minutes. Your OpenAI key is charged separately by OpenAI for the transcription itself.
💡 What people use it for
- Word-timed captions for Shorts and Reels, where each word pops as it is spoken.
- Turning a podcast back catalogue into searchable text, one run per batch of episodes.
- Pulling quotes out of recorded interviews with a timestamp you can jump to.
- Getting a
.vtttrack ready to upload alongside a video for accessibility.
🚧 What it does not do
- It does not pay for the model. Transcription runs on your own key and shows up on your OpenAI
bill.
- It does not split long recordings. The audio goes up as one upload, and transcription
endpoints commonly stop at 25 MB. At the bitrate used here that lands somewhere around 50 minutes of speech, so a full hour is likely to be refused. Cut long files before sending them.
- No speaker labels. You get the words and the timings, not who said them.
- It does not resolve a video page. Give it the file URL, not the page the player sits on.
- The run finishes even when every file failed. Failures are rows, not a failed run, so always
read ok.
- A run stops at one hour and a single download stops after 15 minutes, whichever comes first.
- Rows get large. An hour of speech with word timestamps is a heavy JSON row. Turn
wordTimestamps off when you only need the text.
- Accuracy is the model's. Accents, crosstalk and background noise affect it, and naming the
language usually helps more than changing anything else here.
🧭 Which audio tool do you need?
| If you want | Use |
|---|---|
| A transcript with word-level timings | This one |
| That transcript translated into other languages | Subtitle Translator |
| Captions burned into the picture | Auto Caption Burner |
| A video re-voiced in another language | AI Video Dubber |
| Text turned into a spoken audio file | AI Text-to-Speech Voiceover |
❓ Questions people ask
What counts as a minute? The speech in the file, measured to the end of the last segment and rounded up. Silence at the end of a recording does not add to it.
Can I transcribe several files at once? Yes. Put them in Media URLs (batch) and you get one row per file.
Why is durationSeconds shorter than my file? Because it measures speech, not runtime. A clip with a long musical outro ends its last segment well before the file does.
Can I use a different provider? Yes. Set baseUrl to any OpenAI-compatible endpoint and model to whatever it serves.
Why did one file in my batch come back empty? Read its error field. Most often the link was not a direct file link, or the audio had no speech in it.
Is it legal to transcribe this? Transcribing media you own or have the right to use is normally fine. Recordings of people carry personal data, which GDPR and similar laws cover, so have a reason for holding it. Apify's write-up on scraping and the law is a reasonable starting point, and we are not lawyers.
🆘 If something breaks
Open the Issues tab on the actor page. Send the run ID and the media URL you used. The error field on the failing row usually names the problem on its own.