Bluesky Scraper
Scrape Bluesky posts and profiles by handle, no login. Archive keyword search needs your own app password. $2.00 per 1,000 posts.
How it works
- 1Open it on Apify
Hit Run on Apify — it opens the tool in the cloud, no install.
- 2Set the inputs
Adjust
searchQuery,searchMode,blueskyIdentifier(sensible defaults are pre-filled). - 3Click Run
The tool runs on Apify’s cloud and collects the data for you.
- 4Export the results
Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.
Pricing
$0.002 per post = $2 per 1,000
| You are charged for | When | Price |
|---|---|---|
| Post returned | Charged per post returned. | $0.002 |
Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-06-12, and they are what you are actually charged.
Inputs
| Field | What it does | Type |
|---|---|---|
searchQuery | Keyword(s) to match in Bluesky posts (e.g. "artificial intelligence"). Quoted "phrases" stay whole and several words must all appear. With a Bluesky handle + app password below this searches the full post archive; without one it captures matching posts live from Bluesky's public post stream, because Bluesky no longer serves archive search to signed-out clients. Leave empty to scrape specific authors instead. | string |
searchMode | How to answer a search query. "Auto" uses the full archive when you supply an app password and captures live posts when you do not. "Live capture" always watches the public post stream. "Archive only" never substitutes live capture: without credentials it reports that archive search is unavailable and returns nothing (uncharged). | string |
blueskyIdentifier | Your Bluesky handle or email, e.g. yourname.bsky.social. Needed ONLY to search the full post archive - Bluesky returns 403 for signed-out archive search. Author handles, profiles and live capture need no login. | string |
blueskyAppPassword | An app password from bsky.app -> Settings -> Privacy and security -> App passwords. Never your account password. Needed ONLY to search the full post archive. | string |
authorHandles | Bluesky handles to scrape (e.g. bsky.app, jay.bsky.team). For each handle the actor returns the author's profile plus their recent posts. The leading @ is optional. Leave empty if using a Search query instead. | array |
language | Keep only posts their author tagged with this language. Works on author handles and on the topic-feed lane. Posts with no language tag are left out, since there is nothing to check them against, and some accounts never set one. Left-out posts are not charged. Each handle and each topic feed is read at most 1,000 posts deep, so a language an account rarely uses can return fewer posts than you asked for. Archive search and live capture do not take a language: a keyword that would use either is refused before anything is read. | string |
liveWatchSecs | How long live capture watches Bluesky's public post stream before it stops. It also stops early once it has collected "maxItems" matches. Common keywords fill up in seconds; a rare keyword needs a longer window. | integer |
maxItems | Maximum number of posts to return per search query or per author handle. Pagination follows the API cursor until this limit is reached. | integer |
notionConnector | Optional. Write each post as a page into your Notion when the run finishes. Authorize a Notion connector once in Settings → API & Integrations → MCP connectors, then pick it here. Leave empty to skip (default), results are always saved to the dataset regardless. | string |
notionParentId | Optional. The Notion data source ID of the database to write into (only used if a Notion connector is set). Leave empty to create the pages privately in your workspace instead. | string |
What you get
A structured dataset — each result includes fields like:
authorHandlesdetailssearchQueryauthorHandletextlikeCountrepostCountcreatedAtpostUrlExport every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.
5 ready-to-run use cases
Bluesky Search for Brand Mentions and Engagement
Mentions of a brand with each post's author, likes and reposts. Keyless runs read the live stream, so add an app password to reach older posts.
Bluesky API: Batch Export Posts from Many Handles
Feed a list of handles, get one profile row and the recent posts for each, with like, repost and reply counts. Works with no login.
Bluesky Firehose Capture: Keyword Posts for NLP
Watches Bluesky's public stream and keeps posts matching your keyword. Add an app password to search the archive instead; without one it is live only.
Bluesky Scraper: Posts by Keyword or Hashtag
Without a Bluesky app password this watches the live post stream and keeps what matches, so it finds new posts rather than searching the back catalogue.
Bluesky Follower Count and Post History for a Handle
One row for the profile with follower, following and post counts, then the account's recent posts with likes, reposts and replies. No login needed.
Related tools in Social Media Scrapers
Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.
Mastodon Scraper
Scrape public Mastodon posts by hashtag, account or timeline: text, author, followers, engagement, media and tags. $1.50 per 1,000.
Lemmy Scraper
Scrape public Lemmy posts from any instance: title, body, author, score, votes, comments and permalink. Pick the host. $1.50 per 1,000 posts.
Reddit Text Cleaner — TTS-Ready Narration
Turn raw Reddit posts into clean TTS narration. Strips markdown, links and edit stamps, splits into sentences. No AI key. $0.20 per 1,000 texts.
TikTok Comments Scraper
Scrape TikTok comments from any public video: text, likes, reply count, author and timestamp. No login. $0.25 per 1,000 comments.
TikTok Hashtag Scraper
Scrape TikTok hashtag videos: plays, likes, comments, shares, author and sound. No login or API key. $0.25 per 1,000 videos.
TikTok Sound Scraper
Scrape the TikTok videos using a sound: views, likes, comments, shares, caption, author and MP4 links. $0.25 per 1,000 videos.
Where this tool sits
- Categories
- Social Media Scrapers
- Platforms
- Bluesky
Bluesky Scraper: posts and profiles by handle, plus three ways to follow a keyword
Give it Bluesky handles and it returns each author's profile and their recent posts, with no login at all. Give it a keyword instead and it returns matching posts with engagement counts and links.
Keyword search is the part worth reading before you buy. Bluesky closed its full archive search to signed-out clients, so a keyword run without your own app password gets posts from public topic feeds or from a live capture window, not the whole history.
| Input | Bluesky handles, or a keyword, or both |
| Output | One row per post, plus one profile row per handle |
| Ceiling | 1,000 posts per handle or per query |
| Account needed | None for handles. Your own app password for full archive search |
| Price | $2.00 per 1,000 posts, flat on every plan |
🦋 What Bluesky Scraper does
Point it at handles and it is simple: one profile row with the follower counts and bio, then that author's recent posts, each with text, timestamps and the like, repost, reply and quote counts. Nothing to sign in to.
A keyword is where it forks, and source on every row tells you which lane produced it:
source | What it is | Needs |
|---|---|---|
archive_search | The real archive search, back through history | Your own app password |
topic_feed | Recent posts from public Bluesky topic feeds whose own name or description matches your keyword | Nothing |
live_capture | Bluesky's public post stream watched for a set number of seconds, keeping the posts that match | Nothing |
On auto it takes the archive when you signed in, topic feeds when a feed matches your keyword, and live capture when neither applies. Every keyless keyword run also writes one free notice row saying exactly which lane ran and why.
📥 What you give it
{
"searchQuery": "artificial intelligence",
"searchMode": "auto",
"authorHandles": ["bsky.app"],
"maxItems": 100,
"liveWatchSecs": 60
}
| Field | Default | What it is |
|---|---|---|
searchQuery | box starts at artificial intelligence | Keywords to match. Quoted phrases stay whole, and several bare words must all appear. |
searchMode | auto | auto, live or archive. archive refuses to substitute another lane: with no credentials it says so and returns nothing. |
authorHandles | [], box starts at bsky.app | Handles to scrape, like jay.bsky.team. A leading @ is fine. |
language | any | Keep only posts tagged with this language, like es for Spanish. Works on author handles and the topic-feed lane. Posts in other languages and posts with no language tag are left out, uncharged. Each handle and each feed is read at most 1,000 posts deep, so a language an account rarely uses can return fewer than maxItems. A keyword that would go to the archive or to live capture is refused with a free BAD_INPUT row. |
blueskyIdentifier | none | Your own handle or email. Only used to reach the archive search. |
blueskyAppPassword | none | An app password from bsky.app, Settings, Privacy and security, App passwords. Never your account password. Marked secret. |
maxItems | 100 | Per query and per handle, up to 1,000. Three handles at 100 is up to 300 rows. |
liveWatchSecs | 60 | How long live capture watches the stream, 10 to 900 seconds. It stops early once maxItems is reached. |
notionConnector | none | Optional. Write every post into your own Notion as well as the dataset. |
notionParentId | none | Optional. The Notion data source ID to write into. |
proxyConfiguration | off | Optional, and off by default because a normal run does not need it. |
A rejected app password does not stop the run. It logs a warning and falls back to the keyless lanes, so check source on the rows if you expected archive results.
📤 What you get back
A real post row from a recent keyless run, with the text cut short:
{
"ok": true,
"type": "post",
"uri": "at://did:plc:thtnfvzp2gvyeke24vvrlh4k/app.bsky.feed.post/3mvhejxpvac2g",
"postUrl": "https://bsky.app/profile/hustlrbase.bsky.social/post/3mvhejxpvac2g",
"authorHandle": "hustlrbase.bsky.social",
"authorName": "Hustle with Elisa",
"authorDid": "did:plc:thtnfvzp2gvyeke24vvrlh4k",
"text": "“Overwhelmed by Managing Multiple Client Accounts?” ... #ClientManagement #Automation #AI",
"createdAt": "2026-09-14T05:22:45.479Z",
"likeCount": 0,
"repostCount": 0,
"replyCount": 0,
"quoteCount": 0,
"langs": [],
"feedName": "Artificial Intelligence",
"feedUri": "at://did:plc:ugbenc43qu4v3hemephm2cry/app.bsky.feed.generator/aaamqhwmov3y4",
"feedCreator": "raefmeeuwisse.bsky.social",
"matchesQueryText": false,
"source": "topic_feed",
"matchedQuery": "artificial intelligence"
}
| Field | What it is |
|---|---|
matchesQueryText | Whether your keyword appears in the post text itself. On the topic_feed lane it is often false, because the feed matched your keyword and the individual post is just on that feed. Filter on it when you need a literal match. |
source | archive_search, topic_feed or live_capture. Always check this before drawing a conclusion about coverage. |
uri | The post's permanent at:// identifier. Stable, so use it to dedupe across runs. |
feedName, feedUri, feedCreator | Which topic feed a topic_feed row came from, and who runs it. |
capturedDuringSecs | On live_capture rows, how long the watch window was. |
queriedHandle | On author-mode rows, which handle you asked for. |
langs | Language tags the author's client set. Often empty, because many clients set none. |
A profile row carries type: "profile" with handle, displayName, description, followersCount, followsCount, postsCount, avatar, banner, createdAt and profileUrl.
🧾 Reading the output
Four kinds of row share the dataset.
| Row | How to spot it | Charged as a post |
|---|---|---|
| A post | type: "post" | yes |
| A profile | type: "profile" | no |
| A notice | _notice: true | no |
| A diagnostic | ok: false and an errorCode | no |
| The no-input sample | _sample: true | no |
Filter on type, not on the presence of text. Notice and profile rows share the dataset with posts, and the default table view renders them as partly blank lines. With a language set, the topic-feed notice row also counts the posts read, left out, and left out for having no tag.
| Code | What it means |
|---|---|
BAD_INPUT | You asked for archive mode with no app password, or set a language on a keyword that would go to the archive or live capture. Nothing was read. |
NO_RESULTS | The lane ran and matched nothing, that handle returned neither a profile nor posts, or none of the posts read was in your language. The row says how many were read. |
NOT_FOUND | That handle does not resolve. Check the spelling, including the .bsky.social part. |
RATE_LIMITED | Bluesky asked for a slower pace. Re-run with a smaller maxItems. |
SERVER_ERROR | Bluesky answered with a server error. Usually passes on its own. |
BLOCKED | Bluesky would not serve that request this time. |
NETWORK | Bluesky could not be reached, including when a supplied credential was refused. |
▶️ How to run it
1. Open Bluesky Scraper and click Try for free. 2. For authors, put handles into Author handles and clear Search query. 3. For a keyword, type it into Search query. For the full archive, add your handle and an app password from bsky.app. 4. Set Max items per query/author, then click Start. 5. Read the notice row first if you ran a keyword without credentials. It says which lane answered.
💰 How much does it cost?
$2.00 per 1,000 posts, which is $0.002 each. Flat on every Apify plan, no volume tiers.
Posts are what you pay for. Profile rows, notice rows, the no-input sample and diagnostic rows are all free, duplicates and posts left out by language are dropped before they reach you, and a run that returns no posts costs you nothing.
One thing to know before you budget: on the topic_feed lane you pay for every post the feed returned, including ones where matchesQueryText is false. If you only want literal matches, sign in for the archive or use live capture.
💡 What people use it for
- Watching a handful of accounts daily, deduped on
uri, to see what they posted and how it landed. - Catching mentions of a product name as they happen, with
searchModeonliveand a longer
window.
- Pulling a whole topic feed's recent posts to see who is active in a niche.
- Backfilling a keyword properly with an app password, then keeping it current with short live runs.
🚧 What it does not do
- No archive search without your own app password. Bluesky refuses signed-out archive search,
and no setting here changes that.
- Live capture only sees the window. It watches the public stream for the seconds you allow and
returns the matches inside it. Nothing older exists to return, and a rare keyword can produce zero.
- Fresh live posts look unpopular. A post captured seconds after it was written has no likes yet,
so the counts are near zero by nature rather than by fault.
- Author timelines are recent posts only, capped at 1,000 per handle. It is not a full account
export.
- Topic feeds are somebody else's editorial choice. A feed decides what goes on it, so those
rows reflect the feed, not a search.
- A language goes by the tag the author's app set. Plenty of accounts, news outlets among them,
post with no tag at all, and those posts are left out once you pick a language. Archive search and live capture do not take one.
- No threads, no replies as a tree, no follower lists, no direct messages.
- No account is created for you. The app password stays yours, and archive search runs through
your own account.
🧭 Which social scraper do you need?
| If you want | Use |
|---|---|
| Bluesky posts and profiles | This one |
| The same idea on Mastodon | Mastodon Scraper |
| Threads posts and profiles | Threads Scraper |
| Keyword search on X | Twitter Search Scraper |
| Public Telegram channel posts | Telegram Channel Scraper |
❓ Questions people ask
Do I need a Bluesky account? Only for archive keyword search. Handles, profiles, timelines, topic feeds and live capture all work signed out.
What is an app password? A separate password Bluesky issues for tools, under Settings, Privacy and security. You can revoke it any time, and it is not your account password.
Why did my keyword return posts that do not mention it? You were on the topic-feed lane. Check source and matchesQueryText on the rows, and read the notice row.
Why did live capture return nothing? Nobody posted a match during the window. Raise liveWatchSecs, or sign in and use the archive.
Can I scrape several authors in one run? Yes. Each handle gets its own profile row and its own maxItems worth of posts.
Is scraping Bluesky legal? These are public posts on a public network. They are still personal data, which GDPR and similar laws cover, so have a reason for collecting it. Apify's write-up on scraping and the law is a good starting point, and we are not lawyers.
🆘 If something breaks
Open the Issues tab on the actor page. Send the run ID, the handle or keyword, and the source value on the rows you got. The errorCode on the diagnostic row usually names the problem on its own.