Mastodon Scraper
Scrape public Mastodon posts by hashtag, account or timeline: text, author, followers, engagement, media and tags. $1.50 per 1,000.
How it works
- 1Open it on Apify
Hit Run on Apify — it opens the tool in the cloud, no install.
- 2Set the inputs
Adjust
instance,mode,query(sensible defaults are pre-filled). - 3Click Run
The tool runs on Apify’s cloud and collects the data for you.
- 4Export the results
Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.
Pricing
$0.0015 per post = $1.5 per 1,000
| You are charged for | When | Price |
|---|---|---|
| Post returned | Charged per toot returned. | $0.0015 |
Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-06-13, and they are what you are actually charged.
Inputs
| Field | What it does | Type |
|---|---|---|
instance | The Mastodon instance host to scrape, without the scheme (e.g. "mastodon.social", "fosstodon.org", "mas.to"). The actor calls that instance's public REST API. | string |
mode | What to scrape: "hashtag" timeline (set query to the hashtag), "account" toots (set query to the @handle), or the "public" / federated timeline of the instance. Note: in account mode, boosted/reblogged rows show the reblogger as author (not the original poster). Some instances (e.g. mastodon.social) require auth for the public timeline and will return a BLOCKED diagnostic in public mode. | string |
query | For hashtag mode: the hashtag to fetch, with or without the leading # (e.g. "opensource"). For account mode: the handle, with or without @ (e.g. "Mastodon" or "[email protected]"). Ignored in public mode. | string |
local | Public mode only. If on, returns only posts originating on this instance (local timeline). If off, returns the federated timeline (posts from across the fediverse). Ignored in hashtag and account modes. | boolean |
maxItems | Maximum number of posts to return. The actor pages through the timeline (40 per request) until it reaches this limit or runs out of posts. | integer |
notionConnector | Optional. Write each post as a page into your Notion when the run finishes. Authorize a Notion connector once in Settings → API & Integrations → MCP connectors, then pick it here. Leave empty to skip (default), results are always saved to the dataset regardless. | string |
notionParentId | Optional. The Notion data source ID of the database to write into (only used if a Notion connector is set). Leave empty to create the pages privately in your workspace instead. | string |
What you get
A structured dataset — each result includes fields like:
authorauthorFollowersauthorNamecreatedAtdetailsfavouritesCountidinstancelanguagelocalmediaUrlsmodequeryreblogsCounttexturlExport every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.
2 ready-to-run use cases
Mastodon Timeline of Any Public Account, as JSON
Pull a public mastodon timeline: post text, date, language, tags, reply, reblog and favourite counts, plus the author's follower count. No app registration.
Mastodon Hashtag Feed: #infosec on infosec.exchange
A live mastodon hashtag feed for #infosec on infosec.exchange. CVE chatter and breach reports with post text, author, date, reblog and reply counts.
Related tools in Social Media Scrapers
Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.
Lemmy Scraper
Scrape public Lemmy posts from any instance: title, body, author, score, votes, comments and permalink. Pick the host. $1.50 per 1,000 posts.
Reddit Text Cleaner — TTS-Ready Narration
Turn raw Reddit posts into clean TTS narration. Strips markdown, links and edit stamps, splits into sentences. No AI key. $0.20 per 1,000 texts.
TikTok Comments Scraper
Scrape TikTok comments from any public video: text, likes, reply count, author and timestamp. No login. $0.25 per 1,000 comments.
TikTok Hashtag Scraper
Scrape TikTok hashtag videos: plays, likes, comments, shares, author and sound. No login or API key. $0.25 per 1,000 videos.
TikTok Sound Scraper
Scrape the TikTok videos using a sound: views, likes, comments, shares, caption, author and MP4 links. $0.25 per 1,000 videos.
TikTok Followers and Following Scraper
Get a public TikTok account's following list with no login. Followers need your own cookie. Handles, bios, avatars, regions. $0.35 per 1,000 profiles.
Where this tool sits
- Categories
- Social Media Scrapers
- Platforms
- Mastodon
Mastodon Scraper: public posts by hashtag, account or timeline, no login
Point it at any Mastodon server and pull public posts: the text with the HTML stripped out, who posted it, their follower count, replies, boosts and favourites, any media links, the language and the hashtags.
The thing to understand before you plan anything: a Mastodon server only holds the posts that reached it. Ask two servers for the same hashtag and you get two different sets, neither of them complete. There is no central index to ask instead, so if coverage matters, run it against two or three servers and merge.
| Input | A server host, plus a hashtag or a handle |
| Output | One row per post |
| Ceiling | 2,000 posts per run |
| Account needed | None, and no login anywhere |
| Price | $1.50 per 1,000 posts, flat on every plan |
🔍 What Mastodon Scraper does
One server per run, one of three modes.
Hashtag reads that server's timeline for a tag. Put opensource or #opensource in query, either works.
Account reads one account's posts. Mastodon for a local account, or [email protected] for someone elsewhere, which the server you asked resolves for you.
Public reads the server's own timeline. With local off you get the federated view, everything that server has seen. With local on you get only posts written there, which is the better one for watching a single community.
One caveat on that last mode: large servers, mastodon.social among them, require a login for their public timeline. Ask anyway and you get a BLOCKED row and nothing else. Hashtag and account modes still work on those servers.
📥 What you give it
{
"instance": "fosstodon.org",
"mode": "hashtag",
"query": "opensource",
"maxItems": 200
}
| Field | Default | What it is |
|---|---|---|
instance | mastodon.social | The server host, without https://. fosstodon.org, mas.to, infosec.exchange. |
mode | hashtag | hashtag, account or public. |
query | none | The hashtag or the handle. Leading # and @ are stripped for you. Ignored in public mode, required in the other two. The Console shows opensource as an example, but that is a prefill, so an API call has to send its own. |
local | off | Public mode only. On means posts written on that server, off means everything it has federated. |
maxItems | 100 | 1 to 2,000. Paging is automatic, forty at a time. |
notionConnector | none | Optional. Writes every delivered post into your Notion. Authorise a connector once under Settings, API and Integrations, MCP connectors, then pick it here. |
notionParentId | none | Optional. The Notion data source id to write into. Leave it empty and the pages are created privately in your workspace. |
proxyConfiguration | off | Optional network setting. Off is right for a normal run. |
📤 What you get back
A real row from a recent run:
{
"ok": true,
"id": "117267800308722144",
"text": "Linuxiac: Navidrome 0.64 Music Server & Streamer Adds Experimental Jellyfin Music API Support\nhttps://linuxiac.com/navidrome-0-64-adds-experimental-jellyfin-music-api-support/\n#linux #opensource #tech",
"author": "kiltedtux",
"authorName": "FuzzyFeeds",
"authorFollowers": 26,
"createdAt": "2026-09-14T05:43:48.744Z",
"repliesCount": 0,
"reblogsCount": 0,
"favouritesCount": 0,
"language": "en",
"mediaUrls": [],
"tags": ["linux", "opensource", "tech"],
"url": "https://mastodon.social/@kiltedtux/117267800308722144"
}
| Field | What it is |
|---|---|
text | The post with HTML tags removed and entities decoded. Links and hashtags stay as text. |
author, authorName | The handle and the display name. On a boosted post, both belong to the person who passed it on, not the original poster. |
repliesCount, reblogsCount, favouritesCount | The numbers that server knows about, which are lower than the true totals across the network. 0 also means absent. |
mediaUrls | Direct links to attached images and video, or an empty list. |
tags | Hashtags Mastodon parsed out of the post. |
url | The post's public web address. |
Boosted posts are the one thing to watch. The row credits the account that passed the post on, and its engagement numbers come from that wrapper rather than from the original, so those counts read low. Filter them out by comparing author against the handle you asked for.
🧾 Reading the output
Posts carry ok: true. Anything with ok: false carries an errorCode and is not charged.
| Code | What it means |
|---|---|
BAD_INPUT | An unknown mode, or a hashtag or account run with no query. |
NOT_FOUND | That server does not know the account. Check the handle, and add @server for a remote one. |
NO_RESULTS | The timeline exists and is empty as far as this server is concerned. |
BLOCKED | That server wants a login for what you asked. Common on public timelines of the big servers. |
NETWORK | The host was unreachable, or the name was misspelt. |
The Console's Overview table shows author, text, favourites, boosts, date and URL, and renders diagnostic rows as blank lines. Download the dataset as JSON, CSV or Excel for the tags, media and follower counts.
▶️ How to run it
1. Open Mastodon Scraper and click Try for free. 2. Type the server into Instance, without https://. 3. Pick a Mode, then fill Query with the hashtag or handle. 4. Set Max posts, then click Start. 5. Download the dataset as JSON, CSV, Excel or XML.
💰 How much does it cost?
$1.50 per 1,000 posts. Flat on every Apify plan, no volume tiers, and the same in all three modes.
You pay per post row. Diagnostic rows are free, and a hashtag or timeline that turns up empty does not bill you for results.
💡 What people use it for
- Watching a brand or project name as a hashtag across two or three servers and merging on
id. - Archiving an account's public posts before a server shuts down, which happens often enough to
matter.
- Building a plain-text corpus from a niche server's local timeline, with
localswitched on. - Finding who in a community actually gets boosted, by sorting a tag pull on
reblogsCount.
🚧 What it does not do
- No replies and no threads. You get the post, not the conversation under it.
- One server per run. Covering the network means several runs and a merge on your side.
- No search across servers. Mastodon has no central index, and neither does this.
- Public timelines are often gated. Expect
BLOCKEDon the largest servers in public mode. - No private, followers-only or direct posts, and no login that could reach them.
- A failure partway through a long run loses that run's posts. You get the diagnostic row
instead. Smaller maxItems values survive a shaky server better.
- Counts are one server's view, always at or below the real total.
- Rows are a snapshot. Posts get edited, deleted and defederated.
🧭 Which social scraper do you need?
| If you want | Use |
|---|---|
| Mastodon posts by hashtag, account or timeline | This one |
| Bluesky posts and profiles | Bluesky Scraper |
| Lemmy communities and comments | Lemmy Scraper |
| Reddit posts and comment trees | Reddit Scraper |
| X posts from a search query | X Search Scraper |
❓ Questions people ask
Which server should I use? The one closest to the topic. A niche server's local timeline is richer for its subject than a giant server's federated firehose.
Why did the same hashtag give me different posts yesterday? Because the server's view changes as posts federate in. That is Mastodon working as designed, not a gap here.
Can I read someone on another server? Yes. Use [email protected] in account mode and the server you asked will resolve it.
Why is the public timeline empty or blocked? That server requires a login for it. Use hashtag or account mode, or try a smaller server.
How do I drop boosts? Keep only rows where author matches the handle you asked for.
Is this legal? These are public posts on public timelines, served by an API built to be read. They are still written by identifiable people, so GDPR and similar laws apply and you want a reason for collecting them. Apify's write-up on scraping and the law is a good starting point, and we are not lawyers.
🆘 If something breaks
Open the Issues tab on the actor page. Send the server, the mode, the query and the run id. The errorCode on the diagnostic row usually names the problem by itself.