Reddit Text Cleaner — TTS-Ready Narration
Turn raw Reddit posts into clean TTS narration. Strips markdown, links and edit stamps, splits into sentences. No AI key. $0.20 per 1,000 texts.
How it works
- 1Open it on Apify
Hit Run on Apify — it opens the tool in the cloud, no install.
- 2Set the inputs
Adjust
text,texts,expandAbbreviations(sensible defaults are pre-filled). - 3Click Run
The tool runs on Apify’s cloud and collects the data for you.
- 4Export the results
Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.
Pricing
$0.0002 per text cleaned = $0.2 per 1,000
| You are charged for | When | Price |
|---|---|---|
| Text cleaned | One cleaned/TTS-ready text item. | $0.0002 |
Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-08-04, and they are what you are actually charged.
Inputs
| Field | What it does | Type |
|---|---|---|
text | A single block of text to clean (e.g. a Reddit post body). | string |
texts | An array of strings OR post objects (uses scriptText/narration/selftext/body/text). Lets you pipe the Reddit Scraper's output straight in. | array |
expandAbbreviations | Expand Reddit/internet abbreviations for TTS (AITA → Am I the asshole, MIL → mother-in-law, IMO → in my opinion…). | boolean |
profanityMode | keep = leave as-is · soft = swap for mild words (great for monetization-safe TTS) · censor = f*** · remove = delete. | string |
wpm | Words-per-minute used to estimate read time. | integer |
What you get
A structured dataset — each result includes fields like:
charCountcleanedhookScoreoriginalreadTimeSecondssentenceCountttsSegmentswordCountExport every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.
2 ready-to-run use cases
Reddit Text to Speech Prep: Clean One Post for Narration
Strips markdown and links, spells out AITA and TL;DR, then hands back sentence-sized segments and a read time in seconds. You bring the voice.
Reddit Stories for YouTube: Clean a Batch for Voiceover
Feed an array of scraped posts and every body comes back narration-ready and split into sentences, with a word count and read time per story.
Related tools in Social Media Scrapers
Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.
TikTok Comments Scraper
Scrape TikTok comments from any public video: text, likes, reply count, author and timestamp. No login. $0.25 per 1,000 comments.
TikTok Hashtag Scraper
Scrape TikTok hashtag videos: plays, likes, comments, shares, author and sound. No login or API key. $0.25 per 1,000 videos.
TikTok Sound Scraper
Scrape the TikTok videos using a sound: views, likes, comments, shares, caption, author and MP4 links. $0.25 per 1,000 videos.
TikTok Followers and Following Scraper
Get a public TikTok account's following list with no login. Followers need your own cookie. Handles, bios, avatars, regions. $0.35 per 1,000 profiles.
Instagram Comments Scraper
Scrape comments from any public Instagram post or reel: text, author, likes, replies and timestamp. Nothing to log into. $0.40 per 1,000 rows.
Twitter (X) Search Scraper
Scrape X (Twitter) search by keyword or advanced query. Get the full text, engagement, media and author. No login needed from you. $0.40 per 1,000 tweets.
Where this tool sits
- Categories
- Social Media Scrapers
- Platforms
Reddit Text Cleaner: turn a raw Reddit post into narration a voice can read
Paste a post body, or pipe a whole batch of them in, and you get back text a text-to-speech step can read straight off. Markdown gone, link syntax gone, the "Edit: thanks for the gold" trailer gone, and the result split into one segment per sentence.
It runs on fixed rules, not a model, so there is no key to supply and nothing to pay a model provider. That also means it is literal: it does what its lists say and nothing cleverer, and the places where that bites are written out below rather than left for you to find.
| Input | One block of text, or an array of texts or post objects |
| Output | One row per text |
| Ceiling | No fixed limit, one row per text you send |
| Account needed | None |
| Price | $0.20 per 1,000 texts |
🧹 What Reddit Text Cleaner does
Four passes over each text. It strips markdown and link syntax down to plain words. It expands Reddit shorthand so a voice does not spell it out letter by letter, so AITA becomes "Am I the asshole" and MIL becomes "mother-in-law". It handles swearing the way you ask, from leaving it alone to removing it. Then it splits the result into sentences.
Each row also carries the numbers you need to decide whether a story fits a clip: word count, character count, sentence count, an estimated read time at your chosen speed, and a hook score out of 100 worked out from the opening sentence.
The texts array takes post objects as well as plain strings. It looks for scriptText, narration, selftext, body or text on each object, which means the output of the Reddit scraper linked below drops in without reshaping.
📥 What you give it
{
"texts": [
"AITA for leaving? **So** here's the _story_. Check [this](https://x.com).",
"TIFU by replying all. Edit: yes I know."
],
"profanityMode": "soft",
"wpm": 150
}
| Field | Default | What it is |
|---|---|---|
text | none | A single block of text. Fine for one post. |
texts | none | An array. Strings, or post objects carrying any of the fields named above. |
expandAbbreviations | true | Spells out Reddit and internet shorthand. Read the limits before leaving it on. |
profanityMode | keep | keep, soft for mild swaps, censor for f***, or remove. |
wpm | 150 | Narration speed used to estimate read time. |
Give it one or the other. A run with both text and texts empty fails rather than finishing with an empty dataset.
📤 What you get back
A real row from a real run:
{
"ok": true,
"original": "AITA for leaving? **So** here's the _story_. Check [this](https://x.com).\n\nEdit: thanks for the awards! TL;DR: I left.",
"cleaned": "Am I the asshole for leaving? So here's the story. Check this.",
"ttsSegments": [
"Am I the asshole for leaving?",
"So here's the story.",
"Check this."
],
"sentenceCount": 3,
"wordCount": 12,
"charCount": 62,
"readTimeSeconds": 4.8,
"hookScore": 64
}
| Field | What it is |
|---|---|
cleaned | The narration text. This is the thing you feed to a voice. |
ttsSegments | The same text as an array of sentences, for per-line timing or captions. |
original | What you sent, kept for reference and cut off at 4,000 characters. cleaned is never cut. |
wordCount, charCount, sentenceCount | Measured on the cleaned text. |
readTimeSeconds | Word count against your wpm. Set wpm to 0 and this comes back null. |
hookScore | 0 to 100, scored on the first sentence. Useful for sorting a batch before you pick. |
🧾 Reading the output
Every row is a cleaned text and every row is charged. ok is always true, there are no sample rows and no diagnostic rows, and nothing is filtered out silently.
Two things do get dropped, and it is worth knowing which:
| What you sent | What happens |
|---|---|
An object in texts with none of the known text fields on it | Skipped, no row written |
| Nothing at all, in either field | The run fails instead of writing an empty dataset |
If the row count does not match what you sent, the first line is why.
▶️ How to run it
1. Open Reddit Text Cleaner and click Try for free. 2. Paste one post into Text, or put your batch into Texts. 3. Pick a Profanity handling mode. soft is the one to use for ad-safe voiceover. 4. Click Start. 5. Take cleaned or ttsSegments from the dataset, or read them from the API.
💰 How much does it cost?
$0.20 per 1,000 texts, which is $0.0002 each. Flat on every Apify plan, no volume tiers.
One text in, one row out, one charge. There is no model behind this and no key to add, so there is no separate provider bill on top.
💡 What people use it for
- Prepping Reddit stories for a faceless video pipeline, where markdown read aloud ruins the take.
- Making narration ad-safe in bulk with
soft, instead of editing swearing by hand. - Sorting a pile of candidate stories by
hookScoreandreadTimeSecondsbefore picking the three
worth shooting.
- Turning long posts into per-sentence segments so captions line up with the audio.
🚧 What it does not do
- It does not remove emoji. Invisible zero-width characters go, visible emoji stay exactly where
they were. Strip them yourself if your voice reads them out.
- Abbreviation expansion is blunt. It matches shorthand without caring about case, so "5 mil"
comes out as "5 mother-in-law" and a name like Sil becomes "sister-in-law". If your text has ordinary words that look like Reddit shorthand, turn expandAbbreviations off.
- An
Edit:,ETA:orPS:label cuts the rest. Everything from the first one of those to the
end of the text is removed, which is right for a thank-you trailer and wrong when the story keeps going after it.
- Sentence splitting is mechanical. An abbreviation with a full stop in it, like "Mr.", starts a
new segment.
- English only, and one fixed word list. No language detection, and you cannot supply your own
expansions or swaps.
- It does not fetch anything. Give it text. To pull the posts in the first place, use the
scrapers linked below.
🧭 Which Reddit actor do you need?
| If you want | Use |
|---|---|
| Text you already have, cleaned up for narration | This one |
| The posts themselves, from a subreddit or a profile | Reddit Scraper |
| Posts matching a keyword | Reddit Search Scraper |
| Reddit videos as one MP4 with the sound | Reddit Video Scraper |
| A post rewritten into a hook and a script | Story to Script Rewriter |
❓ Questions people ask
Does this call an AI model? No. It is rules and word lists, which is why there is no key field and why the output is the same every time you send the same text.
Can I feed it the Reddit Scraper's output directly? Yes. Put the rows into texts as they are. It reads scriptText, narration, selftext, body or text off each object.
Why is my read time wrong? It is word count against wpm, nothing more. Set wpm to the speed your voice actually reads at. A wpm of 0 gives you null rather than a number.
Is original the whole post? Only up to 4,000 characters. The cleaned field keeps the lot, so use that if length matters.
What does hookScore measure? How strong the first sentence looks as an opener, scored 0 to 100. It is a sorting aid for a batch, not a verdict.
🆘 If something breaks
Open the Issues tab on the actor page. Paste the text that came out wrong along with the run ID, and that is usually enough to see it.