Substack Publication Scraper
Scrape any Substack publication archive. One row per post: title, link, date, author, word count, comments, reactions and paywall status.
How it works
- 1Open it on Apify
Hit Run on Apify — it opens the tool in the cloud, no install.
- 2Set the inputs
Adjust
publications,maxPostsPerPublication,sort(sensible defaults are pre-filled). - 3Click Run
The tool runs on Apify’s cloud and collects the data for you.
- 4Export the results
Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.
Pricing
$0.00026 per post = $0.26 per 1,000
| You are charged for | When | Price |
|---|---|---|
| Post scraped | One Substack post. Publications that do not exist are never charged. | $0.00026 |
Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-09-20, and they are what you are actually charged.
Inputs
| Field | What it does | Type |
|---|---|---|
publications | One publication per line, up to 20. A handle on its own works (astralcodexten), and so does a full address. If the publication uses its own domain rather than a substack.com one, give that domain - a publication on its own domain only answers on that domain. | array |
maxPostsPerPublication | How many posts to take from each publication, newest first by default. Hard ceiling 2,000. A publication with fewer posts than this simply returns fewer rows. Keep it low while you are testing - you pay per post. | integer |
sort | Which posts you get when you ask for fewer than the whole archive. "new" is the archive in reverse date order. "top" is the publication's most popular. "community" is its discussion threads first. | string |
proxyUrls | Leave this empty for a normal run. Fill it in only if you want the traffic to leave through proxy servers you already pay for, one URL per line, in the form http://user:pass@host:port. | array |
What you get
A structured dataset — each result includes fields like:
publicationtitleurlpostDateaudienceisPaywalledsubtitledescriptionpreviewTextwordCountcommentCountreactionCountrestackCountauthorsExport every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.
Related tools in Developer & Research Tools
Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.
GitHub Scraper
Search GitHub repos and users: stars, forks, language, topics, licence, plus user bio, company and followers. No token needed. $0.90 per 1,000 rows.
Stack Overflow / Stack Exchange Scraper
Search Stack Overflow and Stack Exchange by keyword or tag. Score, answer count, views, reputation and body text. $2 per 1,000 questions.
Package Registry Scraper (npm + PyPI)
Get npm and PyPI package metadata as JSON. Version, license, author, repo, keywords and npm monthly downloads. $2 per 1,000 packages.
arXiv Scraper
Search arXiv papers by title, author, abstract or category. Get full abstracts, authors, categories, DOI, dates and PDF links. $2 per 1,000 papers.
OpenAlex Scholarly Works Scraper
Search 250M+ OpenAlex papers with no API key. Get titles, authors, venue, year, citations, DOI, OA links and full abstracts. $2.00 per 1,000 papers.
Crossref Scholarly Works Scraper
Search 150M+ papers on Crossref: DOI, title, authors, journal, publisher, date, citations and abstract. No API key. $1.00 per 1,000 works.
Where this tool sits
- Categories
- Developer & Research Tools
Substack Publication Scraper: every post in an archive, one row each
Give it a Substack handle like astralcodexten and get the publication's archive back as rows. Title, link, date, author, word count, comments, reactions, restacks, and whether the post sits behind the paywall.
Read this before you buy: no row carries the article text. A publication's public archive does not publish post bodies, for free posts either, so there is nothing to hand you. What you do get is audience, which is the real free-or-paid flag, and Substack's own short preview on about half the posts.
| Input | Substack handles or domains, up to 20 per run |
| Output | One row per post |
| Ceiling | 2,000 posts per publication |
| Account needed | None, no cookies, no browser |
| Price | $0.26 per 1,000 posts, flat on every plan |
📰 What Substack Publication Scraper does
It reads a publication's own archive and walks it to the depth you ask for. A handle on its own is enough. If the publication moved onto its own domain, give that domain: a publication on its own domain only answers there, and its old substack.com address will hand back an empty archive rather than an error.
You can point it at 20 publications in one run, and each gets its own post budget, so one prolific newsletter cannot eat the whole run.
A word on what arriving means. Three different wrong answers all come back as HTTP 200 here: a publication that has moved returns an empty list, one of them redirects to a profile page and answers with a web page, and a site that is not Substack at all serves its own page from the same address. Each of those is checked and turned into an uncharged row that names what to do about it. Nothing is billed for a response that was not an archive.
📥 What you give it
{
"publications": ["astralcodexten", "www.bigtechnology.com"],
"maxPostsPerPublication": 50,
"sort": "new"
}
| Field | What the run uses if you leave it | What it is |
|---|---|---|
publications | box starts at astralcodexten and www.bigtechnology.com | One per line, up to 20. A bare handle works, so does a full address. Give the publication's own domain if it has one. |
maxPostsPerPublication | 50 | Posts to take from each publication, 1 to 2,000. A publication with fewer simply returns fewer. This is also your spending cap. |
sort | new | new is the archive in reverse date order, top is the publication's most popular, community puts discussion threads first. |
proxyUrls | none | Optional. Your own servers, one URL per line as http://user:pass@host:port. Leave it empty for a normal run. |
📤 What you get back
A real row from a recent run:
{
"ok": true,
"charged": true,
"recordType": "post",
"publication": "astralcodexten.substack.com",
"title": "Your Book Review: This Is Going To Hurt",
"url": "https://www.astralcodexten.com/p/your-book-review-this-is-going-to",
"postDate": "2026-09-18T20:53:58.344Z",
"audience": "everyone",
"isPaywalled": false,
"subtitle": "Finalist #9 in the Book Review Contest",
"description": "Finalist #9 in the Book Review Contest",
"previewText": null,
"wordCount": 12639,
"commentCount": 295,
"reactionCount": 152,
"restackCount": 7,
"authors": "Scott Alexander",
"postType": "newsletter",
"slug": "your-book-review-this-is-going-to",
"postId": "216363681",
"coverImage": "https://substack-post-media.s3.amazonaws.com/public/images/916daf3c-...png",
"language": "en",
"scrapedAt": "2026-09-21T02:09:44.003Z"
}
| Field | What it is |
|---|---|
postId | Substack's own id. Stable, so use it to dedupe across runs. |
audience | everyone or only_paid, straight from Substack. isPaywalled is that same fact as a boolean. |
previewText | Substack's own short preview, present on roughly half of all posts and null on the rest. It is not the article. |
wordCount | Substack's count for the full post, including the part you cannot read. |
authors | Every byline joined with commas, or null when the post has none. |
postType | newsletter, podcast, thread and so on, as Substack labels it. |
restackCount | Substack restacks, its own share count. |
url | The canonical link, which is the publication's own domain when it has one. |
🧾 Reading the output
Three kinds of row can land in your dataset, and recordType tells them apart.
| Row | How to spot it | Charged |
|---|---|---|
| A post | recordType: "post" | yes |
| The sample row | _sample: true, recordType: "sample" | no |
| A diagnostic | _diagnostic: true and an errorCode | no |
The charged field is a label on the row, not a receipt. It is stamped when the row is built, which is before billing happens, so treat it as the same signal as recordType: it separates real posts from the sample and diagnostic rows, nothing more.
| Code | What it means |
|---|---|
NOT_FOUND | No Substack publication answers at that address. Check the spelling, or use the publication's own domain. |
NO_RESULTS | The address answered but returned no posts, which is what a publication that has moved to its own domain does. |
TIME_BUDGET | The run ran out of time before reaching that publication. Everything already delivered is complete. |
RATE_LIMITED | Substack asked the run to slow down and it gave up on that item. Re-run it. |
BLOCKED | Substack refused that request. Re-run it, or supply your own servers in proxyUrls. |
SERVER_ERROR | Substack returned a server error. Usually passes on its own. |
NETWORK | Substack was unreachable. Re-run it. |
PROXY_INPUT_ADJUSTED | A network setting in your input was not usable, so the run used its own. |
CHARGE_ERROR | Billing could not be recorded for some rows. Rare, and worth telling us about. |
One bad publication never stops the others, and a diagnostic row never fails the run.
▶️ How to run it
1. Open Substack Publication Scraper and click Try for free. 2. Put one handle per line into Substack publications. 3. Set Posts per publication. Start around 10 while you are looking at the output shape. 4. Pick an Order if reverse date order is not what you want. 5. Click Start, then download the dataset as JSON, CSV or Excel, or read it from the Apify API.
💰 How much does it cost?
$0.26 per 1,000 posts. Flat on every Apify plan, no volume tiers.
You pay per post row delivered. Duplicates, the sample row, diagnostic rows and a publication that returns nothing are all free.
💡 What people use it for
- Tracking which posts a newsletter puts behind the paywall and which it leaves open, using
audience over time.
- Ranking a publication's back catalogue by
reactionCountorcommentCountto see what actually
landed.
- Pulling 20 newsletters in one field into a single sheet to compare posting cadence.
- Watching a competitor's publishing rhythm on a schedule, deduped on
postId.
🚧 What it does not do
- No article text. The archive does not publish post bodies.
previewTextis Substack's own
snippet and it is null about half the time.
- No subscriber counts, no revenue, no email list size. None of that is public.
- No comments themselves, only how many there are.
- A publication on its own domain only answers on that domain. Its old substack.com address
returns an empty archive, and you will get a row saying so.
topandcommunityare Substack's own orderings, applied before paging, so asking for 20 of
them gives you the 20 it serves first.
- No Substack Notes, no podcasts beyond the post record, and no cross-publication search.
- Rows are a snapshot. Reaction and comment counts move, and
scrapedAtrecords when they were
read.
- Up to 20 publications per run. Split a longer list across runs.
🧭 Which news scraper do you need?
| If you want | Use |
|---|---|
| A Substack publication's archive | This one |
| News by keyword, with the publisher's real link | Google News Scraper |
| Worldwide coverage with no API key | GDELT News Scraper |
| Hacker News stories and comments | Hacker News Scraper |
❓ Questions people ask
Do I need a Substack account? No. No account, no cookies, no browser.
Can I get the post text? No, and nothing else can either from a public archive. Substack does not publish bodies there, for free posts or paid ones.
How do I know if a post is paywalled? isPaywalled, or audience if you want Substack's own wording. Every row has it.
What do I put in for a publication on its own domain? That domain. www.bigtechnology.com, not the old substack.com address, which answers empty.
Can I watch several newsletters at once? Yes, up to 20 per run, each with its own post budget.
Is scraping Substack archives legal? These are public archive pages. Rows carry author names, which is personal data under GDPR, so have a reason for collecting it. Apify's write-up on scraping and the law is a sensible starting point, and we are not lawyers.
🆘 If something breaks
Open the Issues tab on the actor page. Send the run ID and the publication you used. The errorCode on the diagnostic row usually names the problem on its own.