Request a tool
All toolsAutomationsGuidesMCP serverRequest a toolPlatformsCategories
Instagram Transcript Scraper icon

Instagram Transcript Scraper

Get the spoken transcript of any public Instagram reel or video. Full text, timed segments, SRT and WebVTT subtitles, language and duration.

267 runs on Apify $0.0025 per transcript ($2.5 / 1,000)
Run this in the cloudRun on Apify →

Developer & Research Tools

How it works

  1. 1
    Open it on Apify

    Hit Run on Apify — it opens the tool in the cloud, no install.

  2. 2
    Set the inputs

    Adjust postUrls, transcriptionApiKey, transcriptionBaseUrl (sensible defaults are pre-filled).

  3. 3
    Click Run

    The tool runs on Apify’s cloud and collects the data for you.

  4. 4
    Export the results

    Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.

Pricing

$0.0025 per transcript = $2.5 per 1,000

You are charged forWhenPrice
Transcript scrapedOne Instagram reel or video transcript delivered as a dataset row. Flat rate regardless of how long the video is. Posts with no speech, photo posts, dead links and private accounts are never charged.$0.0025

Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-08-16, and they are what you are actually charged.

Inputs

FieldWhat it doesType
postUrlsThe posts you want transcribed. A full link (https://www.instagram.com/reel/DcBsBkPOtmB/), a /p/ or /tv/ link, an instagram.com/share/ link, or just the bare shortcode - all work, and you can mix them in one list. Up to 1,000 per run.array
transcriptionApiKeyThe key for YOUR speech-to-text account - OpenAI, Groq, Deepgram's OpenAI-compatible endpoint, an audio-capable model behind an OpenAI-compatible gateway, or your own self-hosted Whisper server. Instagram does not publish transcripts, so the audio has to be listened to by something; using your key means you pay your provider's rate directly instead of a marked-up per-minute fee here. The key is used for the requests and nothing else - it is never stored, logged or written to the dataset.string
transcriptionBaseUrlThe OpenAI-compatible base URL your key belongs to. Leave as is for OpenAI. Examples: https://api.groq.com/openai/v1, https://api.deepgram.com/v1/openai, or your own server.string
transcriptionModelThe model name your provider expects. whisper-1 for OpenAI, whisper-large-v3-turbo for Groq, or the name of any audio-capable model your gateway exposes.string
maxItemsHow many transcripts to return at most. Keep it low while you are testing - you pay per transcript.integer
languageA language code such as es, fr, de or ja, passed to your speech-to-text provider as a hint. Leave empty to let it detect the language itself, which is usually right.string
maxDurationSecondsVideos longer than this are skipped with a free note rather than transcribed. Long videos cost you real money at your speech-to-text provider, so the limit is here to make that a decision rather than a surprise. Default 900 seconds (15 minutes).integer
transcriptionModeLeave on auto. auto tries the /audio/transcriptions endpoint first (the one that returns per-cue timestamps) and falls back to /chat/completions with an audio part if your gateway does not implement it. Force one shape with transcriptions or chat.string
transcriptionChatModelOnly used when the run falls back to /chat/completions and your audio-capable chat model has a different name from the transcription model above.string
concurrencyHow many posts to work through in parallel. Default 5, maximum 15. Lower it if your speech-to-text provider rate-limits you.integer
sessionCookiesLeave this empty. This is not the speech-to-text key, it is an Instagram cookie, and it is optional. Posts are looked up without your account. If you are transcribing a long list and posts come back blocked, paste your own Instagram cookie here, and the run uses it for the posts that need one. In Chrome: F12 → Application → Cookies → instagram.com. The `sessionid` cookie is the one that matters; `csrftoken` alongside it is better. Paste it as `sessionid=...; csrftoken=...`, one entry per account. Cookies are used for this run's requests and nothing else, never stored, never logged, never written to the dataset.array
proxyUrlsLeave this empty for a normal run. Fill it in only if you want the traffic to leave through proxy servers you already pay for, one URL per line, in the form http://user:pass@host:port.array

What you get

A structured dataset — each result includes fields like:

shortcodeauthorUsernamecaptiondurationSecondslanguagesegmentCountwordCounttexturlcreatedAtlikeCount

Export every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.

Related tools in Developer & Research Tools

Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.

GitHub Scraper iconDeveloper & Research Tools

GitHub Scraper

Search GitHub repos and users: stars, forks, language, topics, licence, plus user bio, company and followers. No token needed. $0.90 per 1,000 rows.

18 use cases

Stack Overflow / Stack Exchange Scraper iconDeveloper & Research Tools

Stack Overflow / Stack Exchange Scraper

Search Stack Overflow and Stack Exchange by keyword or tag. Score, answer count, views, reputation and body text. $2 per 1,000 questions.

2 use cases

Package Registry Scraper (npm + PyPI) iconDeveloper & Research Tools

Package Registry Scraper (npm + PyPI)

Get npm and PyPI package metadata as JSON. Version, license, author, repo, keywords and npm monthly downloads. $2 per 1,000 packages.

2 use cases

arXiv Scraper iconDeveloper & Research Tools

arXiv Scraper

Search arXiv papers by title, author, abstract or category. Get full abstracts, authors, categories, DOI, dates and PDF links. $2 per 1,000 papers.

2 use cases

OpenAlex Scholarly Works Scraper iconDeveloper & Research Tools

OpenAlex Scholarly Works Scraper

Search 250M+ OpenAlex papers with no API key. Get titles, authors, venue, year, citations, DOI, OA links and full abstracts. $2.00 per 1,000 papers.

2 use cases

Crossref Scholarly Works Scraper iconDeveloper & Research Tools

Crossref Scholarly Works Scraper

Search 150M+ papers on Crossref: DOI, title, authors, journal, publisher, date, citations and abstract. No API key. $1.00 per 1,000 works.

2 use cases

See all Developer & Research Tools →

Instagram Transcript Scraper: the spoken words out of any public reel or video

Paste reel or video links and get what was actually said, as plain text plus timed segments, SRT and WebVTT subtitles. Each row also carries the creator's caption, the hashtags, the author and the engagement counts, so the transcript arrives with its context attached.

The part to sort out before your first run: Instagram publishes no transcript of its own, so the audio has to be listened to by a speech-to-text service. You bring your own key, in transcriptionApiKey, and you pay your provider directly at their rate. Without a key the run returns one labelled sample row and nothing else.

InputReel, post or IGTV links, share links, or bare shortcodes
OutputOne row per transcript
Ceiling1,000 links per run
Account neededNo Instagram account. Your own speech-to-text key
Price$2.50 per 1,000 transcripts, flat on every plan, plus your provider's own charge

🔍 What Instagram Transcript Scraper does

It resolves each link, pulls the audio, sends it to the endpoint you named, and writes the result as a row. Works with OpenAI, Groq, Deepgram's OpenAI-compatible endpoint, any audio-capable model behind an OpenAI-compatible gateway, or your own self-hosted Whisper server.

Two shapes of endpoint exist and the run tries the better one first. /audio/transcriptions gives you per-cue timings, which is what makes segments, srt and vtt possible. A gateway that only implements /chat/completions still gives you the text, with srt and vtt coming back null. transcriptSource on every row tells you which you got.

No Instagram account needed. The cookie field is optional and separate from your speech-to-text key.

Your key is used for the requests and nothing else. It is never written to the dataset, never logged, and error text from your provider is scrubbed of anything key-shaped before it reaches a diagnostic row.

📥 What you give it

{
  "postUrls": ["https://www.instagram.com/reel/DcBsBkPOtmB/"],
  "transcriptionApiKey": "<your key>",
  "transcriptionBaseUrl": "https://api.openai.com/v1",
  "transcriptionModel": "whisper-1",
  "maxItems": 10
}
FieldDefaultWhat it is
postUrlsnoneFull links, /p/ and /tv/ links, instagram.com/share/ links, or bare shortcodes. Mix them freely, up to 1,000 per run.
transcriptionApiKeynoneYour speech-to-text key. Stored by Apify as a secret field.
transcriptionBaseUrlnone, the form starts at https://api.openai.com/v1The OpenAI-compatible base URL your key belongs to.
transcriptionModelnone, the form starts at whisper-1whisper-large-v3-turbo for Groq, or whatever your gateway exposes.
maxItemsnone, the form starts at 10Transcripts at most, 1 to 5,000. Keep it low while testing.
languagenoneA hint like es, fr, ja. Leave it empty and your provider detects the language, which is usually right.
maxDurationSeconds900, per the field's own helpSkip videos longer than this rather than sending them off to be transcribed.
transcriptionModenone, leave it on autoauto, transcriptions or chat. Force a shape only if auto guesses wrong for your gateway.
transcriptionChatModelnoneOnly for the chat fallback, when your audio-capable chat model has a different name.
concurrency5, per the field's own helpPosts worked through at once, up to 15. Lower it if your provider rate-limits you.
sessionCookiesnoneOptional Instagram cookie, not the speech-to-text key. For posts Instagram will not hand to a signed-out reader.
proxyUrlsnoneOptional. Servers of your own you want the run to use.

Start a run with no links and no key and you get one labelled sample row showing the output shape. That row is not charged.

📤 What you get back

Every run on this account so far has been a shape check rather than a paid transcription, so what follows is the actor's own labelled sample row from a recent run, trimmed, with one emoji cut out of the caption. A real transcript row has exactly this shape with _sample absent.

{
  "ok": true,
  "_sample": true,
  "recordType": "transcript",
  "shortcode": "DcBsBkPOtmB",
  "url": "https://www.instagram.com/reel/DcBsBkPOtmB/",
  "authorUsername": "example_creator",
  "authorIsVerified": true,
  "caption": "The one habit that changed my mornings ... #habits #morningroutine",
  "hashtags": ["habits", "morningroutine"],
  "durationSeconds": 41.6,
  "createdAt": "2026-08-02T15:11:04.000Z",
  "productType": "clips",
  "language": "en",
  "text": "Every one of them. You can't skip. It has to be about your long-term legacy.",
  "wordCount": 38,
  "characterCount": 196,
  "segmentCount": 3,
  "segments": [
    {"start": 0, "end": 2.8, "startTime": "00:00:00.000", "endTime": "00:00:02.800", "text": "Every one of them. You can't skip."}
  ],
  "srt": "1\n00:00:00,000 --> 00:00:02,800\nEvery one of them. You can't skip.",
  "vtt": "WEBVTT\n\n00:00:00.000 --> 00:00:02.800\nEvery one of them. You can't skip.",
  "transcriptSource": "speech-to-text (timestamped)",
  "audioSource": "audio track",
  "audioBytes": 472064,
  "viewCount": 1840000,
  "likeCount": 92100,
  "commentCount": 431,
  "musicTitle": "original sound",
  "scrapedAt": "2026-09-16T14:07:27.343Z"
}
FieldWhat it is
textThe whole transcript as one string.
segmentsPer-cue timings in seconds and as hh:mm:ss.mmm. Empty when your endpoint does not return them.
srt, vttSubtitle files, ready to save. Both null without timings.
captionThe creator's own caption, not the transcript. The two are separate on purpose.
transcriptSourcespeech-to-text (timestamped) or speech-to-text, so you can tell at a glance whether you got timings.
audioSourceaudio track when a separate audio stream existed, video track when the audio had to be taken out of the video.
viewCountUsually null. Instagram does not publish it here. likeCount and commentCount do arrive.
shortcodeUnique per post. Use it as your dedupe key across runs.

🧾 Reading the output

Three kinds of row land in your dataset, and only the first kind is billed.

RowHow to spot itBilled
A transcriptrecordType: "transcript" with text in ityes
The sample row_sample: true, and only when the run had no links or no keyno
A diagnostic_diagnostic: true and an errorCodeno

A row is only billed when the transcript text is not empty. Rows also carry a charged flag, which is how the actor labels real rows against samples and diagnostics.

CodeWhat it means
NOT_A_VIDEOA photo or a carousel. There is no audio to read.
NO_SPEECHThe audio was read and nobody spoke. Music-only reels land here.
TOO_LONGLonger than maxDurationSeconds, skipped before transcription.
TRANSCRIPTION_FAILEDYour endpoint refused the request. A wrong key or model name is the usual cause.
NO_MEDIA_URLInstagram served the post without a playable media link.
NOT_FOUNDDeleted, private, or a share link that does not resolve.
BAD_INPUTNot a recognisable link or shortcode.
BLOCKEDInstagram stopped answering for that post this time.
NO_RESULTSToo many posts in a row failed, so the run stopped rather than keep trying.
TIME_BUDGETThe run ran out of time before reaching that post.
PROXY_INPUT_ADJUSTEDSomething in proxyUrls could not be used as written.
NETWORK, CHARGE_ERROR, UNEXPECTED_ERRORSomething that did not fit the list above.

No diagnostic row is billed.

▶️ How to run it

1. Open Instagram Transcript Scraper and click Try for free. 2. Paste your links into Instagram reels or videos. 3. Put your key in Your speech-to-text API key, and set the endpoint and model if you are not using OpenAI. 4. Set Maximum transcripts to 1 for the first run, so a wrong key costs you one attempt. 5. Click Start, then download the dataset as JSON, CSV or Excel.

💰 How much does it cost?

$2.50 per 1,000 transcripts. Flat on every Apify plan, no volume tiers.

On top of that you pay your own speech-to-text provider, at their rate, on your own account. That is the whole reason the key field exists.

A row is billed only when there is transcript text in it. A photo, a silent reel, a video skipped for length, a failed call to your endpoint: none of those are billed here, though a failed call may still count against your provider.

💡 What people use it for

  • Turning a saved folder of reels into searchable text, so you can grep six months of a niche.
  • Pulling the hook from the first three seconds of a set of reels to see what actually opens them.
  • Building subtitle files for reels you have the rights to repost.
  • Feeding spoken claims into a review process, where the caption says one thing and the voiceover

says another.

🚧 What it does not do

  • It will not work without your own speech-to-text key. No key, no transcript.
  • Accuracy is your provider's, not ours. Accents, background music and crosstalk hurt, and a

different model gives a different result on the same audio.

  • No timings from a chat-shaped gateway. You get the text, and srt and vtt come back empty.
  • No photos or carousels. They have no audio.
  • Audio over 24 MB is refused rather than truncated.
  • maxDurationSeconds is a guard, not a guarantee. Instagram does not always publish a duration

to a signed-out reader, and when it does not, the length check has nothing to go on.

  • No private accounts.
  • No speaker labels and no translation. One language in, the same language out.

🧭 Which Instagram scraper do you need?

If you wantUse
The spoken words inside a reel or videoThis one
The reel's own caption, plays and likesInstagram Reel Scraper
One post or reel you have the link forInstagram Post Scraper
Reels found by keywordInstagram Reels Search
Comments under a reelInstagram Comments Scraper

❓ Questions people ask

Which key do I need? Any OpenAI-compatible speech-to-text key. OpenAI with whisper-1 is the default path. Groq, Deepgram's OpenAI-compatible endpoint and a self-hosted Whisper server all work.

Is my key safe? It is a secret field on Apify, used for the requests and nothing else. It is never written to the dataset or the log, and provider error text is scrubbed before it is shown.

Why did every post fail? Almost always a wrong key, endpoint or model name. Run one link first and read the TRANSCRIPTION_FAILED row, which carries the reason.

Do I need an Instagram account? No. The cookie field is optional and only for posts Instagram will not show a signed-out reader.

Can I get subtitles I can upload? Yes, when your endpoint returns timings. Save the srt or vtt field straight to a file.

Is scraping Instagram legal? This reads public posts only, never private accounts or messages. Results can still contain personal data, which GDPR and similar laws cover, so have a reason for collecting it. Apify's write-up on the legality of web scraping is a good starting point, and we are not lawyers.

🆘 If something breaks

Open the Issues tab on the actor page. Send the link you used and the run ID, and never your key. The errorCode on the diagnostic row usually explains it on its own.