Reddit Scraper
Scrape Reddit posts with no login or API key. Rows carry titles, selftext, scores, authors, comment counts and permalinks. $0.50 per 1,000 rows, flat.
How it works
- 1Open it on Apify
Hit Run on Apify — it opens the tool in the cloud, no install.
- 2Set the inputs
Adjust
sources,subreddits,users(sensible defaults are pre-filled). - 3Click Run
The tool runs on Apify’s cloud and collects the data for you.
- 4Export the results
Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.
Pricing
$0.0005 per Reddit post = $0.5 per 1,000
| You are charged for | When | Price |
|---|---|---|
| Reddit post scraped | One Reddit post in the dataset. | $0.0005 |
Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-08-04, and they are what you are actually charged.
Inputs
| Field | What it does | Type |
|---|---|---|
sources | Any mix of subreddits and Reddit user profiles. Accepts: 'askreddit', 'r/tifu', 'u/GallowBoob', 'user/spez', or a full reddit.com URL. Works for ANY subreddit or profile, not limited to AITA. | array |
subreddits | Optional alternative to 'sources', plain subreddit names. Merged with 'sources'. | array |
users | Optional. Reddit usernames to scrape submissions from (without u/). Merged with 'sources'. | array |
method | rss = no key/login needed (recommended, default). oauth = use Reddit app creds below (adds score/comments). json = legacy anonymous API (often blocked). | string |
sort | top, hot (trending), new (latest), rising, or controversial. | string |
time | Time window for 'top' and 'controversial' sorts (ignored for hot/new/rising). E.g. top of all time, top this year, top this week. | string |
maxPostsPerSubreddit | How many qualifying stories to return per subreddit. | integer |
minScore | Skip posts below this upvote count. Only has an effect with the oauth method, the one that gets upvote counts from Reddit. With the default rss method there are no counts, so this changes nothing. | integer |
minWords | Skip stories shorter than this (too thin for a short). | integer |
maxWords | Skip stories longer than this (won't fit a 30–90s short). | integer |
minHookScore | Filter out weak openers. 0 = keep all. | integer |
requireStory | Keep only self/text-post stories (skip link/image posts) and apply the word-count fit. Off by default = return ALL posts. Turn on for faceless story videos. | boolean |
includeNsfw | Include posts marked over-18. Off by default. Only has an effect with the oauth method. The default rss method cannot tell which posts are over-18, so it never leaves one out, and every row it writes says over18: false. | boolean |
cleanText | Strip markdown/links/edit-stamps and produce TTS-ready sentences. Turn off for raw text. | boolean |
rawMode | Output the complete raw Reddit post object instead of the shaped record. | boolean |
dedupeAcrossRuns | Remember post IDs between runs so you never get the same story twice. | boolean |
commentLimit | Fetch up to this many top comments per post (0 = none). Reliable when you add Reddit app credentials below; anonymous Reddit usually blocks comment access. | integer |
redditClientId | Optional. The default RSS method needs no login. Add a free 'script' app's client ID (reddit.com/prefs/apps) only if you want upvote score + comment counts (not exposed via RSS). | string |
redditClientSecret | The secret for your Reddit app (reddit.com/prefs/apps). | string |
What you get
A structured dataset — each result includes fields like:
authorcreatedUtcfetchedAtfitsShortflairhookScoreidisStorynarrationnumCommentsover18postTypereadTimeSecondsscoresubreddittitleselftextfetchedCommentCounturlExport every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.
20 ready-to-run use cases
Reddit Dataset Builder for LLM Fine-Tuning and RAG
Upvoted question threads from r/askscience, r/AskHistorians, ELI5 and r/AskEngineers, cleaned up and downloadable as JSON or CSV. Change the subs to yours.
Reddit Relationship Stories for YouTube and TikTok
Top weekly r/relationship_advice, r/MaliciousCompliance and r/pettyrevenge threads, trimmed by length and scored on how strong the opening line is.
Reddit Horror Stories from r/nosleep for Narration
Top r/nosleep and r/shortscarystories posts, filtered on upvotes and length, markdown cleaned out for narration. NSFW posts are left in on this preset.
Reddit Stock Sentiment: WSB, r/stocks and r/options
Hot posts and top comments from r/wallstreetbets, r/stocks and r/options, deduped daily. It hands you the raw text; it does not score sentiment for you.
Reddit Crypto Sentiment Data from r/CryptoCurrency
Daily hot threads and comments across r/CryptoCurrency, r/CryptoMarkets, r/ethfinance and r/solana. Raw text for your own model. No scoring is built in.
Reddit Trending Topics - Rising Posts in Niche Subs
Rising threads from r/LocalLLaMA, r/Biohackers, r/EVs and r/PassiveHouse, deduped run to run so you only ever see what is new. Swap in the subs you follow.
Reddit Top Posts of All Time from Any Subreddit
The all-time top threads of any subreddit as a downloadable dataset. Preset is r/dataisbeautiful, r/changemyview and r/science; put your own in the box.
AskReddit Stories for TikTok and YouTube Shorts
Top weekly r/AskReddit, r/tifu and r/confession answers, capped at 400 words and filtered on the opening line. Text cleaned up for text to speech.
Reddit Job Postings Monitor - r/forhire and r/hiring
New posts from r/forhire, r/hiring, r/freelance and r/jobbit, newest first and deduped between runs so a gig only reaches you once. No login needed.
Reddit Lead Generation - r/SaaS and r/marketing Feed
Fresh r/SaaS, r/marketing, r/Entrepreneur and r/smallbusiness threads with their replies, deduped daily. You still have to read them and pick the good ones.
Reddit Monitoring Tool for Brand Mentions in Tech Subs
New posts and top comments from the subs you list, deduped daily. There is no keyword filter: you get the fresh threads and search them for your name.
Reddit Market Research - Complaints and Feature Gaps
Upvoted r/SaaS, r/Notion, r/ProductManagement and r/webdev threads with up to 50 replies each, where people say what a tool gets wrong. Text posts only.
Reddit Best Alternatives Threads - Why Buyers Switch
Year-top r/SaaS, r/selfhosted, r/webapps and r/software threads about what people moved to and why, with up to 40 replies each. Text posts, cleaned up.
Reddit Keyword Research - Questions People Really Ask
Top upvoted question threads from any niche subreddit with their replies, in the words readers use. Preset covers finance, fitness and gardening.
Reddit Product Recommendations from r/BuyItForLife
Year-top threads from r/BuyItForLife, r/HeadphoneAdvice, r/MechanicalKeyboards and r/SkincareAddiction with up to 50 replies, where people name what they own.
Reddit Supplement Reviews - r/Supplements, r/Nootropics
Experience write-ups from r/Supplements, r/Nootropics, r/Biohackers and r/AdvancedFitness, text posts only, up to 40 replies each. Nothing is scored for you.
Reddit AITA Stories for YouTube Shorts and TikTok
Top self-post AITA threads, filtered by word count so they fit a 30-90s short, with the markdown stripped for text to speech. Point it at any subreddit.
Reddit WallStreetBets Scraper: Top Posts This Week
The highest-voted r/wallstreetbets threads of the past week, so you can see which tickers the sub is loud about. Swap in any subreddit or time range.
Reddit User Post History - Scrape Any Username
One account's submissions, newest first, from a u/ username or a profile URL. No login. This preset returns the posts only, not their comment replies.
Reddit Comment Scraper - Top Threads and Their Replies
Top r/AskReddit questions with their highest replies. Comments need a free Reddit app client ID and secret; anonymous Reddit usually blocks them.
Related tools in Social Media Scrapers
Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.
Reddit Posts Search Scraper
Search Reddit by keyword: title, post text, author, subreddit, permalink and dates. Add a free Reddit app for scores. $0.891 per 1,000 posts.
Reddit Video Scraper & Downloader
Download v.redd.it videos from any subreddit or Reddit user, with the audio joined back in. Get post details and dedupe. $0.03 per video record.
Telegram Channel Scraper
Scrape public Telegram channels with no bot token and no API key. Get post text, date, views, author, media links and subscriber count. $1.50 per 1,000.
Pinterest Search Scraper
Search Pinterest by keyword: pin title, description, image URL, destination link, pinner and board. No login. $1.60 per 1,000 pins.
Bluesky Scraper
Scrape Bluesky posts and profiles by handle, no login. Archive keyword search needs your own app password. $2.00 per 1,000 posts.
Mastodon Scraper
Scrape public Mastodon posts by hashtag, account or timeline: text, author, followers, engagement, media and tags. $1.50 per 1,000.
Where this tool sits
- Categories
- Social Media Scrapers
- Platforms
Reddit Scraper: posts and full post text from any subreddit or user, no login
Give it a list of subreddits or user profiles and it returns their posts as rows. Title, author, the whole body text, permalink, timestamp, and a cleaned narration version of the body if you want one.
One thing to know before you start, because it changes what the rows look like. The keyless path reads the public feed, and that feed does not carry vote data. score and numComments arrive as 0, upvoteRatio and flair as null, on every row. That is the feed staying quiet, not the post having no votes. If you need the real numbers, add your own free Reddit app credentials and set method to oauth.
| Input | Subreddit names, user profiles, or reddit.com URLs |
| Output | One row per post |
| Ceiling | 100 posts per source, per run |
| Account needed | None by default. Your own Reddit app only if you want scores and comments |
| Price | $0.50 per 1,000 rows, flat on every plan |
🔍 What Reddit Scraper does
You hand it a mixed list. askreddit, r/tifu, u/GallowBoob, a full reddit.com link: all four shapes work and you can put them in the same list. It reads one page per source, in the sort order you picked, and writes one row per post.
Every row carries the full selftext, not a truncated preview. With cleanText left on you also get a narration field with the markdown, links and edit stamps taken out, and ttsSegments holding that text split into sentences, which is the shape a text-to-speech step wants.
dedupeAcrossRuns is on by default. It remembers the last 5,000 post ids it delivered, so a daily run on the same subreddit gives you what is new rather than the same ten posts again.
📥 What you give it
{
"sources": ["r/tifu", "r/AskReddit", "u/GallowBoob"],
"sort": "top",
"time": "week",
"maxPostsPerSubreddit": 25
}
| Field | Default | What it is |
|---|---|---|
sources | none | Subreddits and user profiles, mixed freely. Give it at least one. |
subreddits, users | none | Older aliases. Anything here is merged into sources. |
method | rss | rss needs no login. oauth uses your credentials and fills in scores and comments. json is a legacy route that is usually blocked. |
sort | top | top, hot, new, rising or controversial. |
time | day | The window for top and controversial. Ignored by the other sorts. |
maxPostsPerSubreddit | 10 | Posts per source. 1 to 100. This is your spend cap. |
cleanText | true | Produces the narration text and sentence segments. Turn it off for raw body text. |
requireStory | false | Text posts only, and applies the word-count fit. Off means all post types. |
minWords, maxWords | 30, 800 | Only applied when requireStory is on. |
minHookScore | 0 | Drops posts whose opener scores below this. 0 keeps everything. |
minScore | 0 | Minimum upvotes. Only does anything on the oauth path. |
includeNsfw | false | Only does anything on the oauth path. See the limits below. |
commentLimit | 0 | Top comments per post. Needs oauth to work in practice. |
dedupeAcrossRuns | true | Skips post ids this actor already delivered to you. |
rawMode | false | Returns Reddit's own post object instead of the shaped row. |
redditClientId, redditClientSecret | none | Your own free script app from reddit.com/prefs/apps. |
proxyConfiguration | Apify default | Leave it alone unless the run has to leave through your own servers. Your own URLs are used exactly as given. |
Give it at least one source. Leave sources empty and the run does not stop: it falls back to r/amitheasshole, fetches those posts and bills you for them.
📤 What you get back
A real row from a real run, on the keyless path:
{
"id": "1wf18de",
"subreddit": "AskReddit",
"sourceType": "subreddit",
"sourceName": "AskReddit",
"sort": "top",
"timeRange": "day",
"title": "What are some creepy or unbelievable facts you know about 9/11?",
"url": "https://www.reddit.com/r/AskReddit/comments/1wf18de/what_are_some_creepy...",
"author": "DeltaBravo_",
"score": 0,
"upvoteRatio": null,
"numComments": 0,
"createdUtc": 1789285288,
"over18": false,
"isStory": false,
"flair": null,
"selftext": "",
"postType": "link",
"scriptText": "What are some creepy or unbelievable facts you know about 9/11?",
"narration": "",
"ttsSegments": [],
"wordCount": 11,
"readTimeSeconds": 4.4,
"hookScore": 56,
"fitsShort": false,
"fetchedAt": "2026-09-14T05:46:11.426Z"
}
| Field | What it is |
|---|---|
id | Reddit's post id without the t3_ prefix. Use it as your dedupe key. |
url | Always the Reddit permalink. On a link post it is not the external address. |
postType | text or link. selftext is empty on a link post, as above. |
scriptText | Title plus body, cleaned when cleanText is on. narration is the body alone. |
hookScore | 0 to 100, worked out from the opening line. fitsShort is the word-count verdict. |
readTimeSeconds | Word count divided by 2.5, so a rough read-aloud estimate. |
comments | Only present when commentLimit is above 0. Each entry has author, body, score and depth. |
🧾 Reading the output
Every row in the dataset is a post, and every post row is charged. This actor writes no sample rows and no diagnostic rows, so there is nothing to filter out.
Which fields actually carry a value depends on the path you ran:
| Field | Keyless (rss) | With your credentials (oauth) |
|---|---|---|
title, author, url, selftext, createdUtc | filled | filled |
score, numComments | always 0 | the real counts |
upvoteRatio, flair | always null | filled when Reddit has them |
over18 | always false, whatever the post is | the real flag |
comments | usually empty | filled up to commentLimit |
A source that is private, banned or misspelled writes a warning into the run log and produces no row. If every source fails that way the run still finishes green with an empty dataset, so check the row count rather than the run status.
▶️ How to run it
1. Open Reddit Scraper and click Try for free. 2. Put your subreddits and profiles into Sources, one per line. 3. Pick a Sort and, for top or controversial, a Time range. 4. Set Max stories per subreddit. Start at 10 while you look at the row shape. 5. Click Start, then download the dataset as JSON, CSV or Excel, or read it from the API.
💰 How much does it cost?
$0.50 per 1,000 rows. Flat on every Apify plan, no volume tiers.
You pay per row delivered. A source that comes back empty produces no rows, and a post that dedupeAcrossRuns already gave you is never delivered again, so neither adds anything to your bill.
💡 What people use it for
- Pulling story posts for short-form video, which is what
narration,ttsSegmentsandfitsShort
exist for.
- Watching a handful of subreddits daily and only reading what is new, with dedupe doing the work.
- Building a text corpus where the full body matters and a preview snippet would be useless.
- Tracking what one prolific poster is putting out, by pointing it at
u/<name>instead of a sub.
🚧 What it does not do
- No keyword search. It reads sources you name. Searching Reddit is a separate actor, linked
below.
- No vote data without your own credentials. On the keyless path
scoreandnumCommentsare
written as 0 rather than left empty, which is easy to misread as a post nobody touched.
minScoreandincludeNsfwdo nothing on the keyless path. There is no score to filter on,
and over-18 posts come through labelled over18: false. Run oauth if either filter matters.
- Credentials alone are not enough. Set
methodtooauthas well, or the run stays keyless
and ignores them.
- One page per source. 100 posts is the hard ceiling for a single source in a single run.
- Link posts hide the destination.
urlis the Reddit permalink, never the article it points at. rawModeskips comments. Ask for Reddit's own object and you get that, without the comments
array.
- Comments are unreliable while keyless. Set
commentLimitwithoutoauthand they usually come
back empty.
🧭 Which Reddit actor do you need?
| If you want | Use |
|---|---|
| Posts from subreddits or profiles you name | This one |
| Posts matching a keyword, across Reddit or inside one sub | Reddit Search Scraper |
| Reddit videos as a single MP4 with the sound in it | Reddit Video Scraper |
| Text you already have, cleaned up for narration | Reddit Text Cleaner |
❓ Questions people ask
Do I need a Reddit account? Not for the default run. The credential fields are optional and only buy you scores, comment counts and comment bodies.
Why is every score 0? The keyless feed does not publish vote counts, and the row writes 0 rather than leaving the field empty. Add your own app credentials and set method to oauth.
My second run returned nothing. Is it broken? Probably not. dedupeAcrossRuns is on, so a run that finds only posts you already have delivers nothing. Turn it off to see everything again.
Can I get more than 100 posts from one subreddit? Not in one run. Split the work across sorts and time windows, or schedule it and let dedupe collect over days.
Is scraping Reddit legal? This reads public posts only, never private messages or accounts. The rows can still hold personal data, which GDPR and similar laws cover, so have a reason for keeping it. Apify's write-up on the legality of web scraping is a sensible starting point, and we are not lawyers.
Can I run it on a schedule? Yes. Use Apify's scheduler, or start it from the API and read the dataset when it finishes.
🆘 If something breaks
Open the Issues tab on the actor page. Include the sources you asked for and the run ID, and the log tells us the rest.