Lemmy Scraper
Scrape public Lemmy posts from any instance: title, body, author, score, votes, comments and permalink. Pick the host. $1.50 per 1,000 posts.
How it works
- 1Open it on Apify
Hit Run on Apify — it opens the tool in the cloud, no install.
- 2Set the inputs
Adjust
instance,mode,query(sensible defaults are pre-filled). - 3Click Run
The tool runs on Apify’s cloud and collects the data for you.
- 4Export the results
Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.
Pricing
$0.0015 per post = $1.5 per 1,000
| You are charged for | When | Price |
|---|---|---|
| Post returned | Charged per post returned. | $0.0015 |
Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-06-13, and they are what you are actually charged.
Inputs
| Field | What it does | Type |
|---|---|---|
instance | The Lemmy instance host to scrape (bare domain, no https://). Examples: lemmy.world, lemmy.ml, beehaw.org, sh.itjust.works. | string |
mode | What to scrape: "feed" = the instance front-page feed; "community" = a single community (put its name in Query); "search" = search posts by keyword (put the term in Query). | string |
query | For mode "community": the community name, e.g. "technology" or "[email protected]". For mode "search": the keywords to search for, e.g. "linux". Ignored in mode "feed". | string |
sort | How to sort posts. Hot/Active rank by recent engagement; New is chronological; the Top* options rank by score within a time window. | string |
maxItems | Maximum number of posts to return. The actor paginates (50 per request) until it reaches this many or runs out of posts. | integer |
notionConnector | Optional. Write each post as a page into your Notion when the run finishes. Authorize a Notion connector once in Settings → API & Integrations → MCP connectors, then pick it here. Leave empty to skip (default), results are always saved to the dataset regardless. | string |
notionParentId | Optional. The Notion data source ID of the database to write into (only used if a Notion connector is set). Leave empty to create the pages privately in your workspace instead. | string |
What you get
A structured dataset — each result includes fields like:
authorauthorActorbodycommentscommunitycommunityTitledownvotesidnsfwpostUrlpublishedscorethumbnailtitleExport every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.
2 ready-to-run use cases
Lemmy World Front Page as JSON: Hot Posts, Scores
The front page as structured rows: title, link, author, community, score, upvotes, downvotes and comment count. Sort by Hot, New or Top. No login.
Lemmy Search: Every Post Mentioning Your Keyword
Searches lemmy.world for a keyword and returns each post with title, link, body, author, community, score and comment count. Newest first.
Related tools in Social Media Scrapers
Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.
Reddit Text Cleaner — TTS-Ready Narration
Turn raw Reddit posts into clean TTS narration. Strips markdown, links and edit stamps, splits into sentences. No AI key. $0.20 per 1,000 texts.
TikTok Comments Scraper
Scrape TikTok comments from any public video: text, likes, reply count, author and timestamp. No login. $0.25 per 1,000 comments.
TikTok Hashtag Scraper
Scrape TikTok hashtag videos: plays, likes, comments, shares, author and sound. No login or API key. $0.25 per 1,000 videos.
TikTok Sound Scraper
Scrape the TikTok videos using a sound: views, likes, comments, shares, caption, author and MP4 links. $0.25 per 1,000 videos.
TikTok Followers and Following Scraper
Get a public TikTok account's following list with no login. Followers need your own cookie. Handles, bios, avatars, regions. $0.35 per 1,000 profiles.
Instagram Comments Scraper
Scrape comments from any public Instagram post or reel: text, author, likes, replies and timestamp. Nothing to log into. $0.40 per 1,000 rows.
Where this tool sits
- Categories
- Social Media Scrapers
- Platforms
- Lemmy
Lemmy Scraper: posts from any instance, by feed, community or keyword
Give it a Lemmy host and a mode, and it reads the public API and hands back posts. Title, the markdown body, the external link, author, community, score, upvotes, downvotes, comment count and the thread permalink, one row each.
The host is the first choice you make here, not an afterthought. Every Lemmy server holds only what it has federated in, so the same community asked on two hosts comes back different, and a small server has seen less of the network than a large one.
| Input | One instance host, plus a community name or a search term |
| Output | One row per post |
| Ceiling | 1,000 posts per run |
| Account needed | None, and no API key |
| Price | $1.50 per 1,000 posts, flat on every plan |
🔍 What Lemmy Scraper does
Three modes off the same host. feed reads the instance front page. community reads one community, either bare like technology or fully qualified like [email protected] when the community lives on another server. search runs the instance's own keyword search.
It pages 50 posts at a time until it has the number you asked for or the instance runs out, deduplicates on post id inside the run, and stops early when a page comes back short.
One host per run. If you want three instances, that is three runs.
📥 What you give it
{
"instance": "lemmy.world",
"mode": "community",
"query": "technology",
"sort": "TopWeek",
"maxItems": 200
}
| Field | Default | What it is |
|---|---|---|
instance | lemmy.world | Bare host, no https://. lemmy.ml, beehaw.org and sh.itjust.works work the same way. |
mode | feed | feed, community or search. |
query | none | The community name in community mode, the keywords in search mode, ignored in feed. The Console box starts at technology; an API call has to send its own. |
sort | Hot | Hot, Active, New, TopDay, TopWeek, TopMonth or TopAll. A value it does not recognise falls back to Hot and says so in the log. |
maxItems | 100 | 1 to 1,000. |
notionConnector | none | Optional. Writes one Notion page per post when the run finishes. Authorise the connector once under Settings, API & Integrations, MCP connectors. |
notionParentId | none | Optional. The Notion data source to write into. Leave it empty and the pages land privately in your workspace. |
proxyConfiguration | off | Optional network settings. A normal run does not need them. |
query is required in community and search mode. Leave it out and the run returns a BAD_INPUT row instead of posts.
📤 What you get back
A real row from a recent run, trimmed where the image URLs run long:
{
"ok": true,
"id": 51894672,
"title": "*69",
"url": "https://slrpnk.net/pictrs/image/b3fc49f0-cf9a-...jpeg",
"body": null,
"author": "Track_Shovel",
"authorActor": "https://slrpnk.net/u/Track_Shovel",
"community": "lemmyshitpost",
"communityTitle": "Lemmy Shitpost",
"score": 36,
"comments": 2,
"upvotes": 37,
"downvotes": 1,
"nsfw": false,
"thumbnail": "https://lemmy.world/pictrs/image/7149d061-8120-...jpeg",
"published": "2026-09-14T04:48:43.513395Z",
"postUrl": "https://lemmy.world/post/51894672"
}
| Field | What it is |
|---|---|
url and postUrl | Two different links, and the difference matters. url is the external page the post points at, and is null on a text post. postUrl is the Lemmy thread itself and is always there. |
body | The post text as markdown, so bold stays bold. Stray HTML is cleaned out. null when the post is just a link. |
authorActor | The federated user URL. That is the identity that stays stable across instances, not author. |
comments | The count. The comment text itself is not part of this actor. |
downvotes | 0 on instances that switch downvotes off, which several do as a moderation choice. That is a server setting, not unanimous approval. |
id | Stable. Use it as your deduplication key across runs. |
🧾 Reading the output
Two kinds of row land in your dataset.
| Row | How to spot it | Charged |
|---|---|---|
| A post | ok: true and a postUrl | yes |
| A diagnostic | ok: false and an errorCode | no |
| Code | What it means |
|---|---|
BAD_INPUT | The host is not a valid hostname, the mode is unknown, or query is missing in community or search mode. |
NOT_FOUND | The host answered with something that is not Lemmy's API, or that community is not on it. Check the spelling, or use the fully qualified name@host form. |
NO_RESULTS | Host and community were both fine, there was simply nothing to return. |
▶️ How to run it
1. Open Lemmy Scraper and click Try for free. 2. Leave Lemmy instance at lemmy.world, or type another host. 3. Pick a Mode, and fill in Query unless you chose the front-page feed. 4. Set Max posts. Start around 20 to see the shape of the output, then click Start. 5. Download the dataset as JSON, CSV or Excel, or read it from the Apify API.
💰 How much does it cost?
$1.50 per 1,000 posts. Flat on every Apify plan, no volume tiers.
You pay for posts delivered. Duplicates dropped inside the run and diagnostic rows are not charged, and a search that finds nothing costs you nothing.
💡 What people use it for
- Watching one community on a schedule with
sortset toNew, and diffing each run against the
one before it.
- Reading what actually got traction in a niche community last month with
TopMonth, wherescore
next to comments separates what people upvoted from what they argued about.
- Keyword monitoring, with realistic expectations. Lemmy is small, so you get genuine mentions and
not many of them.
- Filing a community's posts into Notion as pages, straight off the run.
🚧 What it does not do
- No comment text. You get
commentsas a count and nothing else. - One host per run. Comparing instances means one run each.
- You see only what the host has federated in. A server that has defederated from another will
not show you its posts, and a small server has simply seen fewer of them.
- Search ranking belongs to the instance, and this actor does not re-order it.
- 1,000 posts per run is the hard ceiling, whatever the community holds.
- If a page fails partway through a long run, the run ends with a diagnostic row rather than the
posts it had already collected. Re-run it, ideally with a smaller maxItems.
- Public posts only. No logins, nothing private, nothing already removed.
🧭 Which social scraper do you need?
| If you want | Use |
|---|---|
| Lemmy posts from any instance | This one |
| Mastodon hashtags and timelines | Mastodon Scraper |
| Bluesky posts and profiles | Bluesky Scraper |
| Reddit posts by keyword | Reddit Search Scraper |
| Hacker News stories and comments | Hacker News Scraper |
❓ Questions people ask
Do I need a Lemmy account? No. The API this reads is public and there is nothing to sign up for.
Which instance should I pick? The one whose community you actually want. For a community hosted elsewhere, use the qualified name@host form so the server you are asking knows where to look.
Why did the same search give different answers on two hosts? Because each holds a different slice of the network. That is federation working normally, not a fault in the run.
Can I schedule it? Yes, like any Apify actor. sort on New with a small maxItems makes a cheap hourly watch.
Can I get the comments too? Not from this one. Only the count comes back.
Is scraping Lemmy posts legal? These are public posts on public instances. They still contain personal data, which GDPR and similar laws cover, so have a reason for collecting it. Apify's write-up on scraping and the law is a good place to start, and we are not lawyers.
🆘 If something breaks
Open the Issues tab on the actor page. Send the instance, the mode and the run ID. The errorCode on the diagnostic row usually names the problem on its own.