AI Text-to-Speech Voiceover
Turn any script into an AI voiceover file: MP3, WAV, Opus or AAC. Pick the voice, speed and format. $40.00 per 1,000 voiceovers ($0.04 each).
How it works
- 1Open it on Apify
Hit Run on Apify — it opens the tool in the cloud, no install.
- 2Set the inputs
Adjust
text,texts,voice(sensible defaults are pre-filled). - 3Click Run
The tool runs on Apify’s cloud and collects the data for you.
- 4Export the results
Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.
Pricing
$0.04 per voiceover = $40 per 1,000
| You are charged for | When | Price |
|---|---|---|
| Voiceover generated | One completed voiceover stored and returned in the dataset. | $0.04 |
Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-09-30, and they are what you are actually charged.
Inputs
| Field | What it does | Type |
|---|---|---|
text | The text to convert to speech. Long scripts are chunked and stitched automatically. | string |
texts | Array of strings OR objects (uses script/scriptText/text/narration). One audio file per item. | array |
voice | AI voice. | string |
model | tts-1 (fast) or tts-1-hd (higher quality). | string |
format | Output audio format. | string |
speed | Playback speed 0.25–4.0 (1.0 = normal). | string |
openaiApiKey | Your OpenAI key (TTS). Kept private. | string |
baseUrl | OpenAI-compatible base URL. Default https://api.openai.com/v1. | string |
What you get
A structured dataset — each result includes fields like:
_demo_noticeaudioKeyaudioUrlcharacterschunksdurationSecondsformatindexmodeltextPreviewvoiceExport every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.
3 ready-to-run use cases
AI Voice Generator for YouTube Narration Scripts
Paste a script, pick a voice, get an MP3 to drop under B-roll. You supply your own OpenAI key. Without one it returns a labelled sample row instead of audio.
IVR Recording from Text - Phone Menu Prompts as AAC
Turn press one for sales into an AAC file your phone system can play. Needs your own OpenAI key. Speed sits at 0.95 so menu options land clearly.
Batch Text to Speech - One Audio File per Line
Give it a list of strings and each comes back as its own file, which suits app prompts and UI sounds. Bring your own OpenAI key, no key returns a sample.
Related tools in YouTube & Creator Tools
Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.
Video & Audio Transcriber — Word-Level + SRT/VTT
Transcribe any video or audio URL with word-level timestamps. Download SRT, VTT and TXT. The language is found for you. $0.02 per transcribed minute.
AI Character Reference Bank
Make a consistent AI character reference set. You get portrait, 3/4, full body, side and expression sheets, plus a character-bible prompt.
YouTube Channel Scraper
Scrape videos from any YouTube channel, no login or API key. Rows carry titles, views, dates, duration, thumbnails.
YouTube Comments Scraper
Scrape YouTube comments with no login or key. Get text, likes, reply counts, author handle, verified badge and timestamp. $0.40 per 1,000.
YouTube Playlist Scraper
Export every video in a YouTube playlist. Get title, watch URL, views, age, runtime and thumbnail. No API key or OAuth. $0.40 per 1,000 videos.
YouTube Transcript Scraper
YouTube transcripts from links or a keyword search: text, timed segments, SRT, VTT and video details. No API key. $0.80 per 1,000.
Where this tool sits
AI Text-to-Speech Voiceover: turn a script into an audio file you can download
Paste a script, or a whole list of them, and you get one audio file per item in MP3, WAV, Opus or AAC. Each file lands in the run's key-value store with a direct link, next to a row telling you the voice, the length and how many characters were read.
The part to know before you start: this runs on your own OpenAI key. You paste the key in, OpenAI bills you for the synthesis, and without a key the run hands back a free demo record instead of audio.
| Input | A script in text, or a list of them in texts |
| Output | One audio file and one row per item |
| Ceiling | No item cap. Long scripts are split and rejoined, so the run timeout is the real limit |
| Account needed | Your own OpenAI API key |
| Price | $0.04 per voiceover, flat on every plan |
🔊 What AI Text-to-Speech Voiceover does
It reads your text out loud and hands you the file. Six voices, four formats, and a speed dial from 0.25 to 4.0.
A long script does not need splitting by hand. Anything over about 3,500 characters is cut at sentence boundaries, each piece is synthesised on its own, and the pieces are joined back into a single file without re-encoding. The row tells you how many pieces went into it, in chunks, so you can see when that happened.
texts is the batch door. Give it an array of strings and you get one file per string. Give it an array of objects and it reads whichever of script, scriptText, text or narration is present. An object with none of those four keys is dropped quietly, which is worth knowing if a batch comes back shorter than you sent.
Items are done one after another rather than all at once, so a long list takes a while.
📥 What you give it
{
"text": "Welcome back to the channel. Today we are looking at one of the strangest mysteries of the deep ocean.",
"voice": "onyx",
"format": "mp3",
"speed": "1.0"
}
| Field | Default | What it is |
|---|---|---|
text | none | One script, as a single string. The form opens with an example sentence filled in, which is a starting point rather than a default. |
texts | none | Batch mode. An array of strings, or of objects keyed script, scriptText, text or narration. One audio file per item. |
voice | onyx | One of alloy, echo, fable, onyx, nova, shimmer. onyx is the deep male one, nova and shimmer are female. |
model | tts-1 | tts-1 is quick, tts-1-hd sounds better and costs you more on your own OpenAI bill. |
format | mp3 | mp3, wav, opus or aac. |
speed | 1.0 | Playback speed between 0.25 and 4.0. It is a text field, so type a number: anything that is not one quietly becomes 1.0. |
openaiApiKey | none | Your own OpenAI key. Marked secret, so it is not written into the run's visible input. |
baseUrl | none | Advanced. Any OpenAI-compatible /audio/speech host. Left empty it uses https://api.openai.com/v1. |
Run it with no key at all and you get a single labelled demo record in the key-value store, so you can see the shape before you wire anything up. Give it a key but no text and no texts and the run fails, because there is nothing to read.
📤 What you get back
One dataset row per finished voiceover, and the audio itself in the run's key-value store under a key like voiceover-1-1757000000000.mp3.
| Field | What it is |
|---|---|
ok | true on a delivered voiceover. |
index | Which item this was, counting from 1. |
voice, model, format | What actually got used, after defaults are applied. |
characters | How long the input text was. |
chunks | How many separate pieces were synthesised and joined. 1 means it fitted in one. |
durationSeconds | Measured off the finished file. It reads 0 when that measurement failed, not when the audio is empty. |
audioKey | The key-value store key the file is saved under. |
audioUrl | A direct link to the file. |
textPreview | The first 160 characters of what was read, so you can tell rows apart. |
No example row is printed here. No run on record has produced one with a working customer key, and a plausible row typed out by hand is worse than none.
🧾 Reading the output
Three kinds of record come out of a run, and only one of them is in the dataset.
| Record | Where it lands | Billed |
|---|---|---|
| A finished voiceover | A dataset row, ok: true | yes |
| The keyless demo | Key-value store, SAMPLE_AND_NOTICES | no |
| An item that failed | Key-value store, SAMPLE_AND_NOTICES, ok: false | no |
So everything in the dataset is audio you can use. One failed item does not stop the batch: it is recorded in the store and the run moves to the next item.
Worth knowing, because it is the thing that confuses people: if every item fails, the run still finishes green with an empty dataset. The reasons are all sitting in SAMPLE_AND_NOTICES in the key-value store. An empty dataset on this actor means look there, not that your text was empty.
The Overview table shows index, voice, format, durationSeconds, characters, audioUrl and textPreview. model and chunks are real fields but are not among the columns, so open the row itself or download the JSON if you want them.
▶️ How to run it
1. Open AI Text-to-Speech Voiceover and click Try for free to see the demo record. 2. Paste your OpenAI key into OpenAI API key (BYO). 3. Put your script in Text / script, or a list of scripts in Texts (batch). 4. Pick the voice and format, then click Start. 5. Take the files from the run's Storage tab, or follow audioUrl from each dataset row.
💰 How much does it cost?
$0.04 per voiceover. Flat on every Apify plan, no volume tiers. Fifty narration blocks in one batch run come to $2.00.
You pay per delivered file. The keyless demo record and any item that failed are not charged. OpenAI bills your own key separately for the synthesis itself, and tts-1-hd costs you more there than tts-1.
💡 What people use it for
- Narration for faceless video channels, where a week of scripts goes in as one
textsbatch. - Audiobook and long-article reading, leaning on the automatic splitting so nothing has to be cut
up by hand.
- IVR and phone-menu prompts, where WAV is usually the format the phone system wants.
- Trying the same script in several voices before committing to one.
🚧 What it does not do
- No voice cloning, and no custom voices. Six voices, listed above, and that is the list.
- No SSML. There are no tags for pauses, emphasis or pronunciation. Punctuation and
speedare
the only controls over delivery.
- A rejected key stops the batch, rather than retrying its way down your whole list.
durationSecondscan read0when the finished file could not be measured. The audio is
still there.
- Joined
opusandaacfiles are less proven than MP3. Long scripts in those two formats are
stitched the same way, but MP3 is the one with playback evidence behind it. Use MP3 or WAV if a long script has to be right first time.
- No video, no captions, no music bed. This is the audio track only.
🧭 Which AI media actor do you need?
| If you want | Use |
|---|---|
| A script read out loud as an audio file | This one |
| A finished vertical video with script, voice and captions | AI Faceless Video Generator |
| An existing video re-voiced in another language | AI Video Dubber |
| Scene images for a story before you voice it | AI Storyboard Generator |
| Word-by-word captions burned onto a video you already have | Auto Caption Burner |
❓ Questions people ask
Do I need my own OpenAI key? Yes. There is no shared key behind this, and without one you get the demo record rather than audio.
Is my key safe here? It is a secret field, so it is not stored in the run's visible input. Use one you can revoke, as with anything you paste into a tool you did not write.
How long can a script be? There is no fixed character ceiling. Long text is split at sentence boundaries and joined back together, so the practical limit is how long the run is allowed to take.
Can I use something other than OpenAI? If it speaks the same /audio/speech shape, point baseUrl at it. Anything else will not work.
Why is my dataset empty when the run went green? Every item failed. The reasons are in SAMPLE_AND_NOTICES in the run's key-value store.
Can I run it on a schedule? Yes, through Apify's scheduler, or start it from the API and read the dataset when it finishes.
🆘 If something breaks
Open the Issues tab on the actor page. Include the run ID and which format and voice you used, and check SAMPLE_AND_NOTICES in the run's key-value store first, since the per-item reason is usually already written there.