Request a tool
All toolsAutomationsGuidesMCP serverRequest a toolPlatformsCategories
Video & Audio Transcriber — Word-Level + SRT/VTT icon

Video & Audio Transcriber — Word-Level + SRT/VTT

Transcribe any video or audio URL with word-level timestamps. Download SRT, VTT and TXT. The language is found for you. $0.02 per transcribed minute.

5 from 1 review on Apify 35 runs on Apify $0.02 per transcribed minute ($20 / 1,000)
Run this in the cloudRun on Apify →

YouTube & Creator Tools

How it works

  1. 1
    Open it on Apify

    Hit Run on Apify — it opens the tool in the cloud, no install.

  2. 2
    Set the inputs

    Adjust mediaUrl, mediaUrls, language (sensible defaults are pre-filled).

  3. 3
    Click Run

    The tool runs on Apify’s cloud and collects the data for you.

  4. 4
    Export the results

    Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.

Pricing

$0.02 per transcribed minute = $20 per 1,000

You are charged forWhenPrice
Transcribed minutePer minute of audio transcribed.$0.02

Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-08-08, and they are what you are actually charged.

Inputs

FieldWhat it doesType
mediaUrlPublic URL to a video or audio file (mp4, mov, mp3, wav, m4a, webm). Use this for a single file, or mediaUrls for a batch.string
mediaUrlsTranscribe several files in one run, one dataset row per URL. Each item is a public video/audio URL.array
languageSpoken language ISO code, or 'auto' to detect.string
wordTimestampsReturn per-word start/end times (great for karaoke captions).boolean
outputFormatsWhich subtitle/text files to also produce: srt, vtt, txt.array
openaiApiKeyYour OpenAI (Whisper) key. Kept private.string
modelTranscription model. Default whisper-1.string
baseUrlOpenAI-compatible base URL. Default https://api.openai.com/v1.string

What you get

A structured dataset — each result includes fields like:

_demo_noticedurationSecondslanguagesegmentCountsegmentssourceUrlsrtKeytextvttKeywordCountsrtUrlvttUrl

Export every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.

3 ready-to-run use cases

MP4 to SRT: Subtitle File From a Video URL

Point it at a public MP4 and get a timed SRT back, language detected automatically. Needs your own OpenAI key; without one you get a sample row.

Transcribe MP3 to Text: Podcast Episode Transcript

An episode URL in, clean text out, split into timed segments for show notes. Needs your own OpenAI key; without one you get a labelled sample row.

Word Level Timestamps for Karaoke and TikTok Captions

Per-word start and end times alongside SRT and VTT files, for captions that pop word by word. Needs your own OpenAI key to return a real transcript.

Related tools in YouTube & Creator Tools

Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.

AI Character Reference Bank iconYouTube & Creator Tools

AI Character Reference Bank

Make a consistent AI character reference set. You get portrait, 3/4, full body, side and expression sheets, plus a character-bible prompt.

2 use cases

YouTube Channel Scraper iconYouTube & Creator Tools

YouTube Channel Scraper

Scrape videos from any YouTube channel, no login or API key. Rows carry titles, views, dates, duration, thumbnails.

Ready to run — no setup

YouTube Comments Scraper iconYouTube & Creator Tools

YouTube Comments Scraper

Scrape YouTube comments with no login or key. Get text, likes, reply counts, author handle, verified badge and timestamp. $0.40 per 1,000.

Ready to run — no setup

YouTube Playlist Scraper iconYouTube & Creator Tools

YouTube Playlist Scraper

Export every video in a YouTube playlist. Get title, watch URL, views, age, runtime and thumbnail. No API key or OAuth. $0.40 per 1,000 videos.

2 use cases

YouTube Transcript Scraper iconYouTube & Creator Tools

YouTube Transcript Scraper

YouTube transcripts from links or a keyword search: text, timed segments, SRT, VTT and video details. No API key. $0.80 per 1,000.

2 use cases

YouTube Trending Scraper iconYouTube & Creator Tools

YouTube Trending Scraper

Get the most-viewed YouTube videos for any topic, no login or key. Rows carry views, titles, channels, duration, thumbnails. $0.40 per 1,000 videos.

Ready to run — no setup

See all YouTube & Creator Tools →

Video & Audio Transcriber: transcripts with word-level timing, from any media URL

Give it a public link to a video or an audio file and get back the full text, sentence-level segments, and a start and end time for every single word. It also writes ready-made .srt, .vtt and .txt files into the run's storage.

You bring your own OpenAI key, so the transcription itself is billed to you by OpenAI on top of what you pay here.

InputOne public media URL, or a list of them. mp4, mov, webm, mp3, wav, m4a
OutputOne row per file, with the transcript, segments, words and file links
CeilingAbout 5 GB per file at the default memory. Long recordings are the real limit, see below
Account neededYour own OpenAI API key
Price$0.02 per transcribed minute, flat on every plan

🎙️ What Video & Audio Transcriber does

It downloads the file you point it at, pulls the speech out as compressed mono audio, and sends that for transcription. What comes back is the plain text, the segments with their timings, and the word list with a start and end time on each word. That word list is what karaoke-style captions need, and most transcript tools do not hand it over.

Give it mediaUrls instead of mediaUrl and it walks the list, writing one row per file. A file that fails does not stop the ones after it.

The language is detected for you unless you name it. Naming it is usually a little more accurate on short or noisy clips.

📥 What you give it

{
  "mediaUrl": "https://example.com/podcast.mp3",
  "language": "auto",
  "wordTimestamps": true,
  "outputFormats": ["srt", "vtt", "txt"],
  "openaiApiKey": "sk-..."
}
FieldDefaultWhat it is
mediaUrlnoneA public, direct link to one video or audio file.
mediaUrlsnoneA list, for a batch. One dataset row per URL. You can use this instead of mediaUrl or alongside it.
languageautoISO code of the spoken language, or auto to detect it.
wordTimestampstrueAdds the per-word words array to the row. Turn it off for a much smaller row on long files.
outputFormats["srt", "vtt"] when the field is absent, box starts at srt, vtt, txtWhich files to write into the run's storage.
openaiApiKeynoneYour own key. Marked secret, so it is not stored with the run input.
modelwhisper-1Any transcription model your key can reach.
baseUrlhttps://api.openai.com/v1Point it at any OpenAI-compatible endpoint.

The URL has to be the file itself, not a page that plays it. A link ending in .mp4 or .mp3 works; a video page does not.

📤 What you get back

A real row, with the long arrays and URLs cut short:

{
  "ok": true,
  "sourceUrl": "https://example.com/podcast.mp3",
  "language": "en",
  "text": "Welcome back to the show. Today we cover the deep ocean.",
  "wordCount": 11,
  "segmentCount": 2,
  "durationSeconds": 8,
  "segments": [
    { "start": 0, "end": 4, "text": "Welcome back to the show." },
    { "start": 4, "end": 8, "text": "Today we cover the deep ocean." }
  ],
  "words": [{ "word": "Welcome", "start": 0, "end": 0.4 }, "..."],
  "srtKey": "transcript-1757178142775-0.srt",
  "srtUrl": "https://api.apify.com/v2/key-value-stores/.../records/transcript-...",
  "vttUrl": "https://api.apify.com/v2/key-value-stores/.../records/transcript-..."
}
FieldWhat it is
wordsOne entry per word with its own start and end, in seconds. Present only when wordTimestamps is on.
segmentsSentence-sized cues. This is what the .srt and .vtt files are built from.
durationSecondsWhere the last speech segment ends, rounded. This is the length that gets billed, not the file's own runtime.
textThe whole transcript as one string.
srtUrl, vttUrl, txtUrlDownload links, present only for the formats you asked for.
sourceUrlWhich input URL this row belongs to. On a batch this is the only way to tell rows apart.

🧾 Reading the output

Three kinds of row can land in your dataset.

RowHow to spot itCharged a minute
A transcriptok: true and a sourceUrlyes
A failed fileok: false and an error stringno
The sample row_demo: trueno

Check ok before you count rows. A failed file still writes a row, so a batch of ten can show ten rows with only six transcripts among them. The error field on those rows says what went wrong in plain words: a download that failed, a file over the size limit, or no speech in the audio.

The default table view hides sourceUrl and error. Switch the dataset to All fields, or export as JSON, when you are working through a batch.

You get the _demo sample row instead of real work when no media URL was given, or no key was.

▶️ How to run it

1. Open Video & Audio Transcriber and click Try for free. 2. Paste a direct file link into Media URL, or a list into Media URLs (batch). 3. Put your key into OpenAI API key (BYO). 4. Leave Include word timestamps on if you want karaoke captions, then click Start. 5. Read the rows in the dataset, or open the run's Key-value store tab for the subtitle files.

💰 How much does it cost?

$0.02 per transcribed minute. Flat on every Apify plan, no volume tiers. A 12 minute clip is 12 minutes, rounded up to the next whole minute, with one minute as the smallest charge per file.

The length billed is the speech in the file, measured to where the last segment ends. A file that fails and a sample row are not billed any minutes. Your OpenAI key is charged separately by OpenAI for the transcription itself.

💡 What people use it for

  • Word-timed captions for Shorts and Reels, where each word pops as it is spoken.
  • Turning a podcast back catalogue into searchable text, one run per batch of episodes.
  • Pulling quotes out of recorded interviews with a timestamp you can jump to.
  • Getting a .vtt track ready to upload alongside a video for accessibility.

🚧 What it does not do

  • It does not pay for the model. Transcription runs on your own key and shows up on your OpenAI

bill.

  • It does not split long recordings. The audio goes up as one upload, and transcription

endpoints commonly stop at 25 MB. At the bitrate used here that lands somewhere around 50 minutes of speech, so a full hour is likely to be refused. Cut long files before sending them.

  • No speaker labels. You get the words and the timings, not who said them.
  • It does not resolve a video page. Give it the file URL, not the page the player sits on.
  • The run finishes even when every file failed. Failures are rows, not a failed run, so always

read ok.

  • A run stops at one hour and a single download stops after 15 minutes, whichever comes first.
  • Rows get large. An hour of speech with word timestamps is a heavy JSON row. Turn

wordTimestamps off when you only need the text.

  • Accuracy is the model's. Accents, crosstalk and background noise affect it, and naming the

language usually helps more than changing anything else here.

🧭 Which audio tool do you need?

If you wantUse
A transcript with word-level timingsThis one
That transcript translated into other languagesSubtitle Translator
Captions burned into the pictureAuto Caption Burner
A video re-voiced in another languageAI Video Dubber
Text turned into a spoken audio fileAI Text-to-Speech Voiceover

❓ Questions people ask

What counts as a minute? The speech in the file, measured to the end of the last segment and rounded up. Silence at the end of a recording does not add to it.

Can I transcribe several files at once? Yes. Put them in Media URLs (batch) and you get one row per file.

Why is durationSeconds shorter than my file? Because it measures speech, not runtime. A clip with a long musical outro ends its last segment well before the file does.

Can I use a different provider? Yes. Set baseUrl to any OpenAI-compatible endpoint and model to whatever it serves.

Why did one file in my batch come back empty? Read its error field. Most often the link was not a direct file link, or the audio had no speech in it.

Is it legal to transcribe this? Transcribing media you own or have the right to use is normally fine. Recordings of people carry personal data, which GDPR and similar laws cover, so have a reason for holding it. Apify's write-up on scraping and the law is a reasonable starting point, and we are not lawyers.

🆘 If something breaks

Open the Issues tab on the actor page. Send the run ID and the media URL you used. The error field on the failing row usually names the problem on its own.