Instagram Transcript Scraper
Get the spoken transcript of any public Instagram reel or video. Full text, timed segments, SRT and WebVTT subtitles, language and duration.
How it works
- 1Open it on Apify
Hit Run on Apify — it opens the tool in the cloud, no install.
- 2Set the inputs
Adjust
postUrls,transcriptionApiKey,transcriptionBaseUrl(sensible defaults are pre-filled). - 3Click Run
The tool runs on Apify’s cloud and collects the data for you.
- 4Export the results
Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.
Pricing
$0.0025 per transcript = $2.5 per 1,000
| You are charged for | When | Price |
|---|---|---|
| Transcript scraped | One Instagram reel or video transcript delivered as a dataset row. Flat rate regardless of how long the video is. Posts with no speech, photo posts, dead links and private accounts are never charged. | $0.0025 |
Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-08-16, and they are what you are actually charged.
Inputs
| Field | What it does | Type |
|---|---|---|
postUrls | The posts you want transcribed. A full link (https://www.instagram.com/reel/DcBsBkPOtmB/), a /p/ or /tv/ link, an instagram.com/share/ link, or just the bare shortcode - all work, and you can mix them in one list. Up to 1,000 per run. | array |
transcriptionApiKey | The key for YOUR speech-to-text account - OpenAI, Groq, Deepgram's OpenAI-compatible endpoint, an audio-capable model behind an OpenAI-compatible gateway, or your own self-hosted Whisper server. Instagram does not publish transcripts, so the audio has to be listened to by something; using your key means you pay your provider's rate directly instead of a marked-up per-minute fee here. The key is used for the requests and nothing else - it is never stored, logged or written to the dataset. | string |
transcriptionBaseUrl | The OpenAI-compatible base URL your key belongs to. Leave as is for OpenAI. Examples: https://api.groq.com/openai/v1, https://api.deepgram.com/v1/openai, or your own server. | string |
transcriptionModel | The model name your provider expects. whisper-1 for OpenAI, whisper-large-v3-turbo for Groq, or the name of any audio-capable model your gateway exposes. | string |
maxItems | How many transcripts to return at most. Keep it low while you are testing - you pay per transcript. | integer |
language | A language code such as es, fr, de or ja, passed to your speech-to-text provider as a hint. Leave empty to let it detect the language itself, which is usually right. | string |
maxDurationSeconds | Videos longer than this are skipped with a free note rather than transcribed. Long videos cost you real money at your speech-to-text provider, so the limit is here to make that a decision rather than a surprise. Default 900 seconds (15 minutes). | integer |
transcriptionMode | Leave on auto. auto tries the /audio/transcriptions endpoint first (the one that returns per-cue timestamps) and falls back to /chat/completions with an audio part if your gateway does not implement it. Force one shape with transcriptions or chat. | string |
transcriptionChatModel | Only used when the run falls back to /chat/completions and your audio-capable chat model has a different name from the transcription model above. | string |
concurrency | How many posts to work through in parallel. Default 5, maximum 15. Lower it if your speech-to-text provider rate-limits you. | integer |
sessionCookies | Leave this empty. This is not the speech-to-text key, it is an Instagram cookie, and it is optional. Posts are looked up without your account. If you are transcribing a long list and posts come back blocked, paste your own Instagram cookie here, and the run uses it for the posts that need one. In Chrome: F12 → Application → Cookies → instagram.com. The `sessionid` cookie is the one that matters; `csrftoken` alongside it is better. Paste it as `sessionid=...; csrftoken=...`, one entry per account. Cookies are used for this run's requests and nothing else, never stored, never logged, never written to the dataset. | array |
proxyUrls | Leave this empty for a normal run. Fill it in only if you want the traffic to leave through proxy servers you already pay for, one URL per line, in the form http://user:pass@host:port. | array |
What you get
A structured dataset — each result includes fields like:
shortcodeauthorUsernamecaptiondurationSecondslanguagesegmentCountwordCounttexturlcreatedAtlikeCountExport every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.
Related tools in Developer & Research Tools
Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.
GitHub Scraper
Search GitHub repos and users: stars, forks, language, topics, licence, plus user bio, company and followers. No token needed. $0.90 per 1,000 rows.
Stack Overflow / Stack Exchange Scraper
Search Stack Overflow and Stack Exchange by keyword or tag. Score, answer count, views, reputation and body text. $2 per 1,000 questions.
Package Registry Scraper (npm + PyPI)
Get npm and PyPI package metadata as JSON. Version, license, author, repo, keywords and npm monthly downloads. $2 per 1,000 packages.
arXiv Scraper
Search arXiv papers by title, author, abstract or category. Get full abstracts, authors, categories, DOI, dates and PDF links. $2 per 1,000 papers.
OpenAlex Scholarly Works Scraper
Search 250M+ OpenAlex papers with no API key. Get titles, authors, venue, year, citations, DOI, OA links and full abstracts. $2.00 per 1,000 papers.
Crossref Scholarly Works Scraper
Search 150M+ papers on Crossref: DOI, title, authors, journal, publisher, date, citations and abstract. No API key. $1.00 per 1,000 works.
Where this tool sits
- Categories
- Developer & Research Tools
Instagram Transcript Scraper: the spoken words out of any public reel or video
Paste reel or video links and get what was actually said, as plain text plus timed segments, SRT and WebVTT subtitles. Each row also carries the creator's caption, the hashtags, the author and the engagement counts, so the transcript arrives with its context attached.
The part to sort out before your first run: Instagram publishes no transcript of its own, so the audio has to be listened to by a speech-to-text service. You bring your own key, in transcriptionApiKey, and you pay your provider directly at their rate. Without a key the run returns one labelled sample row and nothing else.
| Input | Reel, post or IGTV links, share links, or bare shortcodes |
| Output | One row per transcript |
| Ceiling | 1,000 links per run |
| Account needed | No Instagram account. Your own speech-to-text key |
| Price | $2.50 per 1,000 transcripts, flat on every plan, plus your provider's own charge |
🔍 What Instagram Transcript Scraper does
It resolves each link, pulls the audio, sends it to the endpoint you named, and writes the result as a row. Works with OpenAI, Groq, Deepgram's OpenAI-compatible endpoint, any audio-capable model behind an OpenAI-compatible gateway, or your own self-hosted Whisper server.
Two shapes of endpoint exist and the run tries the better one first. /audio/transcriptions gives you per-cue timings, which is what makes segments, srt and vtt possible. A gateway that only implements /chat/completions still gives you the text, with srt and vtt coming back null. transcriptSource on every row tells you which you got.
No Instagram account needed. The cookie field is optional and separate from your speech-to-text key.
Your key is used for the requests and nothing else. It is never written to the dataset, never logged, and error text from your provider is scrubbed of anything key-shaped before it reaches a diagnostic row.
📥 What you give it
{
"postUrls": ["https://www.instagram.com/reel/DcBsBkPOtmB/"],
"transcriptionApiKey": "<your key>",
"transcriptionBaseUrl": "https://api.openai.com/v1",
"transcriptionModel": "whisper-1",
"maxItems": 10
}
| Field | Default | What it is |
|---|---|---|
postUrls | none | Full links, /p/ and /tv/ links, instagram.com/share/ links, or bare shortcodes. Mix them freely, up to 1,000 per run. |
transcriptionApiKey | none | Your speech-to-text key. Stored by Apify as a secret field. |
transcriptionBaseUrl | none, the form starts at https://api.openai.com/v1 | The OpenAI-compatible base URL your key belongs to. |
transcriptionModel | none, the form starts at whisper-1 | whisper-large-v3-turbo for Groq, or whatever your gateway exposes. |
maxItems | none, the form starts at 10 | Transcripts at most, 1 to 5,000. Keep it low while testing. |
language | none | A hint like es, fr, ja. Leave it empty and your provider detects the language, which is usually right. |
maxDurationSeconds | 900, per the field's own help | Skip videos longer than this rather than sending them off to be transcribed. |
transcriptionMode | none, leave it on auto | auto, transcriptions or chat. Force a shape only if auto guesses wrong for your gateway. |
transcriptionChatModel | none | Only for the chat fallback, when your audio-capable chat model has a different name. |
concurrency | 5, per the field's own help | Posts worked through at once, up to 15. Lower it if your provider rate-limits you. |
sessionCookies | none | Optional Instagram cookie, not the speech-to-text key. For posts Instagram will not hand to a signed-out reader. |
proxyUrls | none | Optional. Servers of your own you want the run to use. |
Start a run with no links and no key and you get one labelled sample row showing the output shape. That row is not charged.
📤 What you get back
Every run on this account so far has been a shape check rather than a paid transcription, so what follows is the actor's own labelled sample row from a recent run, trimmed, with one emoji cut out of the caption. A real transcript row has exactly this shape with _sample absent.
{
"ok": true,
"_sample": true,
"recordType": "transcript",
"shortcode": "DcBsBkPOtmB",
"url": "https://www.instagram.com/reel/DcBsBkPOtmB/",
"authorUsername": "example_creator",
"authorIsVerified": true,
"caption": "The one habit that changed my mornings ... #habits #morningroutine",
"hashtags": ["habits", "morningroutine"],
"durationSeconds": 41.6,
"createdAt": "2026-08-02T15:11:04.000Z",
"productType": "clips",
"language": "en",
"text": "Every one of them. You can't skip. It has to be about your long-term legacy.",
"wordCount": 38,
"characterCount": 196,
"segmentCount": 3,
"segments": [
{"start": 0, "end": 2.8, "startTime": "00:00:00.000", "endTime": "00:00:02.800", "text": "Every one of them. You can't skip."}
],
"srt": "1\n00:00:00,000 --> 00:00:02,800\nEvery one of them. You can't skip.",
"vtt": "WEBVTT\n\n00:00:00.000 --> 00:00:02.800\nEvery one of them. You can't skip.",
"transcriptSource": "speech-to-text (timestamped)",
"audioSource": "audio track",
"audioBytes": 472064,
"viewCount": 1840000,
"likeCount": 92100,
"commentCount": 431,
"musicTitle": "original sound",
"scrapedAt": "2026-09-16T14:07:27.343Z"
}
| Field | What it is |
|---|---|
text | The whole transcript as one string. |
segments | Per-cue timings in seconds and as hh:mm:ss.mmm. Empty when your endpoint does not return them. |
srt, vtt | Subtitle files, ready to save. Both null without timings. |
caption | The creator's own caption, not the transcript. The two are separate on purpose. |
transcriptSource | speech-to-text (timestamped) or speech-to-text, so you can tell at a glance whether you got timings. |
audioSource | audio track when a separate audio stream existed, video track when the audio had to be taken out of the video. |
viewCount | Usually null. Instagram does not publish it here. likeCount and commentCount do arrive. |
shortcode | Unique per post. Use it as your dedupe key across runs. |
🧾 Reading the output
Three kinds of row land in your dataset, and only the first kind is billed.
| Row | How to spot it | Billed |
|---|---|---|
| A transcript | recordType: "transcript" with text in it | yes |
| The sample row | _sample: true, and only when the run had no links or no key | no |
| A diagnostic | _diagnostic: true and an errorCode | no |
A row is only billed when the transcript text is not empty. Rows also carry a charged flag, which is how the actor labels real rows against samples and diagnostics.
| Code | What it means |
|---|---|
NOT_A_VIDEO | A photo or a carousel. There is no audio to read. |
NO_SPEECH | The audio was read and nobody spoke. Music-only reels land here. |
TOO_LONG | Longer than maxDurationSeconds, skipped before transcription. |
TRANSCRIPTION_FAILED | Your endpoint refused the request. A wrong key or model name is the usual cause. |
NO_MEDIA_URL | Instagram served the post without a playable media link. |
NOT_FOUND | Deleted, private, or a share link that does not resolve. |
BAD_INPUT | Not a recognisable link or shortcode. |
BLOCKED | Instagram stopped answering for that post this time. |
NO_RESULTS | Too many posts in a row failed, so the run stopped rather than keep trying. |
TIME_BUDGET | The run ran out of time before reaching that post. |
PROXY_INPUT_ADJUSTED | Something in proxyUrls could not be used as written. |
NETWORK, CHARGE_ERROR, UNEXPECTED_ERROR | Something that did not fit the list above. |
No diagnostic row is billed.
▶️ How to run it
1. Open Instagram Transcript Scraper and click Try for free. 2. Paste your links into Instagram reels or videos. 3. Put your key in Your speech-to-text API key, and set the endpoint and model if you are not using OpenAI. 4. Set Maximum transcripts to 1 for the first run, so a wrong key costs you one attempt. 5. Click Start, then download the dataset as JSON, CSV or Excel.
💰 How much does it cost?
$2.50 per 1,000 transcripts. Flat on every Apify plan, no volume tiers.
On top of that you pay your own speech-to-text provider, at their rate, on your own account. That is the whole reason the key field exists.
A row is billed only when there is transcript text in it. A photo, a silent reel, a video skipped for length, a failed call to your endpoint: none of those are billed here, though a failed call may still count against your provider.
💡 What people use it for
- Turning a saved folder of reels into searchable text, so you can grep six months of a niche.
- Pulling the hook from the first three seconds of a set of reels to see what actually opens them.
- Building subtitle files for reels you have the rights to repost.
- Feeding spoken claims into a review process, where the caption says one thing and the voiceover
says another.
🚧 What it does not do
- It will not work without your own speech-to-text key. No key, no transcript.
- Accuracy is your provider's, not ours. Accents, background music and crosstalk hurt, and a
different model gives a different result on the same audio.
- No timings from a chat-shaped gateway. You get the text, and
srtandvttcome back empty. - No photos or carousels. They have no audio.
- Audio over 24 MB is refused rather than truncated.
maxDurationSecondsis a guard, not a guarantee. Instagram does not always publish a duration
to a signed-out reader, and when it does not, the length check has nothing to go on.
- No private accounts.
- No speaker labels and no translation. One language in, the same language out.
🧭 Which Instagram scraper do you need?
| If you want | Use |
|---|---|
| The spoken words inside a reel or video | This one |
| The reel's own caption, plays and likes | Instagram Reel Scraper |
| One post or reel you have the link for | Instagram Post Scraper |
| Reels found by keyword | Instagram Reels Search |
| Comments under a reel | Instagram Comments Scraper |
❓ Questions people ask
Which key do I need? Any OpenAI-compatible speech-to-text key. OpenAI with whisper-1 is the default path. Groq, Deepgram's OpenAI-compatible endpoint and a self-hosted Whisper server all work.
Is my key safe? It is a secret field on Apify, used for the requests and nothing else. It is never written to the dataset or the log, and provider error text is scrubbed before it is shown.
Why did every post fail? Almost always a wrong key, endpoint or model name. Run one link first and read the TRANSCRIPTION_FAILED row, which carries the reason.
Do I need an Instagram account? No. The cookie field is optional and only for posts Instagram will not show a signed-out reader.
Can I get subtitles I can upload? Yes, when your endpoint returns timings. Save the srt or vtt field straight to a file.
Is scraping Instagram legal? This reads public posts only, never private accounts or messages. Results can still contain personal data, which GDPR and similar laws cover, so have a reason for collecting it. Apify's write-up on the legality of web scraping is a good starting point, and we are not lawyers.
🆘 If something breaks
Open the Issues tab on the actor page. Send the link you used and the run ID, and never your key. The errorCode on the diagnostic row usually explains it on its own.