Stack Overflow / Stack Exchange Scraper
Search Stack Overflow and Stack Exchange by keyword or tag. Score, answer count, views, reputation and body text. $2 per 1,000 questions.
How it works
- 1Open it on Apify
Hit Run on Apify — it opens the tool in the cloud, no install.
- 2Set the inputs
Adjust
query,site,tags(sensible defaults are pre-filled). - 3Click Run
The tool runs on Apify’s cloud and collects the data for you.
- 4Export the results
Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.
Pricing
$0.002 per question = $2 per 1,000
| You are charged for | When | Price |
|---|---|---|
| Question returned | Charged per question returned. | $0.002 |
Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-06-13, and they are what you are actually charged.
Inputs
| Field | What it does | Type |
|---|---|---|
query | Keywords to search for in question titles and bodies (e.g. "async await", "git rebase conflict"). Can be left empty if you provide one or more Tags instead. | string |
site | Which Stack Exchange network site to search. | string |
tags | Comma-separated tags to filter by, e.g. "javascript,promise" or "python,pandas". Optional. A question must carry ALL listed tags. You can search by tags alone with an empty query. | string |
sortBy | Ordering of results. "votes" = highest score first, "relevance" = best keyword match, "creation" = newest, "activity" = most recently active. | string |
minScore | Only return questions with at least this score (net votes). Leave empty for no minimum. Works only with Sort by = Votes: Stack Exchange applies a minimum to whatever the results are sorted by, so with another order the run refuses it before searching. | integer |
includeAcceptedAnswer | Adds the answer the asker accepted to each question row: its score, author and body as plain text. Questions with no accepted answer get null. Same price per question. Other answers and comments are not read. | boolean |
maxItems | Maximum number of questions to return. The actor paginates the API (100 per page) until this many are collected or there are no more results. | integer |
notionConnector | Optional. Write each question as a page into your Notion when the run finishes. Authorize a Notion connector once in Settings → API & Integrations → MCP connectors, then pick it here. Leave empty to skip (default), results are always saved to the dataset regardless. | string |
notionParentId | Optional. The Notion data source ID of the database to write into (only used if a Notion connector is set). Leave empty to create the pages privately in your workspace instead. | string |
What you get
A structured dataset — each result includes fields like:
answerCountbodycreatedAtisAnsweredownerNameownerReputationquestionIdscoretagstitleurlviewCountExport every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.
2 ready-to-run use cases
Python Pandas Questions on Stack Overflow by Votes
Highest-scoring questions carrying both the python and pandas tags, with score, views, answer count and the question body as plain text.
Stack Overflow API Search for async/await Questions
Keyword search through the official Stack Exchange API, ranked by relevance, with score, views, tags and body. Questions only, no answer text.
Related tools in Developer & Research Tools
Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.
Package Registry Scraper (npm + PyPI)
Get npm and PyPI package metadata as JSON. Version, license, author, repo, keywords and npm monthly downloads. $2 per 1,000 packages.
arXiv Scraper
Search arXiv papers by title, author, abstract or category. Get full abstracts, authors, categories, DOI, dates and PDF links. $2 per 1,000 papers.
OpenAlex Scholarly Works Scraper
Search 250M+ OpenAlex papers with no API key. Get titles, authors, venue, year, citations, DOI, OA links and full abstracts. $2.00 per 1,000 papers.
Crossref Scholarly Works Scraper
Search 150M+ papers on Crossref: DOI, title, authors, journal, publisher, date, citations and abstract. No API key. $1.00 per 1,000 works.
Wikipedia Scraper
Search Wikipedia by keyword or by exact title. Get the intro text, the full article, thumbnails and categories. Any language. $1.00 per 1,000 pages.
Internet Archive Scraper
Search Internet Archive (archive.org) for books, audio, film and web items. Title, creator, year, downloads, subjects and URL. $2.00 per 1,000 items.
Where this tool sits
- Categories
- Developer & Research Tools
- Platforms
- Stack Overflow
Stack Overflow Scraper: questions, tags, scores and the question body
Search Stack Overflow and five sibling Stack Exchange sites by keywords, by tags, or by both. Each row is one question: the title, the score, how many answers it has, the view count, its tag list, who asked it and their reputation, and the whole question body as plain text.
Questions first. Turn on includeAcceptedAnswer and each row also carries the accepted answer as plain text. The other answers and the comments are not in the output, so if you need whole threads, this is the wrong purchase however cheap it is.
| Input | Keywords, tags, or both |
| Output | One row per question |
| Ceiling | 1,000 questions per run |
| Account needed | None, and no Stack Exchange key |
| Price | $2.00 per 1,000 questions, flat on every plan |
🔍 What Stack Overflow Scraper does
It runs one search and pages through it until it has as many questions as you asked for.
query matches words in titles and bodies. tags filters, and it is an AND: put javascript,promise and you get only questions carrying both, not either. You can use tags on their own with no keywords, which is the cleanest way to pull the top questions in a tag.
sortBy decides the order, and it changes what a small run gives you. votes returns the long-standing classics. creation returns what was asked most recently. activity returns what people are still arguing about. relevance leans on the keyword match. On votes you can also set minScore to keep only questions at or above a score.
Six sites are available: Stack Overflow, Server Fault, Super User, Ask Ubuntu, MathOverflow and Software Engineering. One site per run.
With includeAcceptedAnswer on, each question also carries the answer its asker accepted: score, author and the body as plain text. Many questions never get one, and those say so with null.
📥 What you give it
{
"query": "async await",
"tags": "javascript,promise",
"site": "stackoverflow",
"sortBy": "votes",
"maxItems": 50
}
| Field | Default | What it is |
|---|---|---|
query | none | Keywords matched against titles and bodies. The Console shows async await as an example, but that is a prefill, so an API call has to send its own. |
tags | none | Comma separated, and every tag must be present on the question. The Console prefills javascript, which again does not travel to an API call. |
site | stackoverflow | One of stackoverflow, serverfault, superuser, askubuntu, mathoverflow, softwareengineering. |
sortBy | votes | votes, relevance, creation or activity. Always highest or newest first. |
minScore | none | Only questions with at least this score. Stack Exchange applies it, so lower-scored questions are never fetched or billed. Works with sortBy set to votes only; with another order the run refuses it, because the site would read the number as a date. |
includeAcceptedAnswer | false | Adds each question's accepted answer to its row, at the same price per question. Works with every site and sort. |
maxItems | 50 | 1 to 1,000. Paging is automatic, a hundred at a time. |
notionConnector | none | Optional. Writes every delivered question into your Notion. Authorise a connector once under Settings, API and Integrations, MCP connectors, then pick it here. |
notionParentId | none | Optional. The Notion data source id to write into. Leave it empty and the pages are created privately in your workspace. |
proxyConfiguration | off | Optional network setting. Off is right for a normal run. |
Send a query, tags, or both. Leaving both empty is not refused, and what comes back then is whatever the site decides to hand over, charged like any other rows. It is the one input mistake here that costs money.
📤 What you get back
A real row from a recent run, with the body cut short here for length:
{
"ok": true,
"questionId": 37576685,
"title": "Using async/await with a forEach loop",
"url": "https://stackoverflow.com/questions/37576685/using-async-await-with-a-foreach-loop",
"score": 3353,
"answerCount": 35,
"viewCount": 2481345,
"isAnswered": true,
"tags": ["javascript", "node.js", "promise", "async-await", "ecmascript-2017"],
"ownerName": "Saad",
"ownerReputation": 55029,
"createdAt": "2016-06-01T18:55:58.000Z",
"body": "Are there any issues with using async/await in a forEach loop? I'm trying to loop through an array of files and await on the contents of each file. ..."
}
| Field | What it is |
|---|---|
body | The question, converted from HTML to plain text. Code blocks survive as text, indentation mostly does not. |
isAnswered | Stack Exchange's own flag: true when an answer was accepted or has a positive score, so it can be true with no accepted answer. Not the same as answerCount being above zero. |
acceptedAnswerId | Only with includeAcceptedAnswer. The accepted answer's id, or null when the question has none. |
acceptedAnswer | Only with includeAcceptedAnswer. {answerId, score, ownerName, ownerReputation, createdAt, body}, with body as plain text. null when there is no accepted answer, or when it could not be read in this run (then acceptedAnswerId is still set). |
ownerName, ownerReputation | Both null when the account was deleted or the post is anonymous. |
score | Net votes, which can be negative. |
title | Passed through as the API sends it, so an HTML entity like " can survive into this field. Only body is decoded. |
tags | The question's full tag list, which is usually wider than the tags you filtered on. |
🧾 Reading the output
Questions carry ok: true. Anything with ok: false carries an errorCode and is not charged.
| Code | What it means |
|---|---|
NO_RESULTS | The search ran and matched nothing. Usually a tag that does not exist, or two tags no single question carries. |
BAD_INPUT | minScore was set with a sortBy other than votes. Nothing was searched; the row's note says what to change. |
RATE_LIMITED | Stack Exchange asked the run to stop for the day. Whatever was already collected is in the dataset and complete as far as it goes. |
NETWORK | Stack Exchange was unreachable or answered with something unusable. Re-run it. |
A RATE_LIMITED row after some questions is not a failed run. It means the run delivered what it could and stopped rather than spinning. The same goes for an acceptedAnswerId with an empty acceptedAnswer: that answer could not be read in this run, usually because the day's allowance ran out. The run log says how many were read.
The Console's Overview table hides tags, isAnswered and body. Download the dataset as JSON, CSV or Excel, or switch the table to all fields.
▶️ How to run it
1. Open Stack Overflow Scraper and click Try for free. 2. Type into Search query, or fill Tags instead, or both. 3. Pick a Stack Exchange site if you want one of the other five. 4. Choose a Sort by order and set Max questions, then click Start. 5. Download the dataset as JSON, CSV, Excel or XML.
💰 How much does it cost?
$2.00 per 1,000 questions. Flat on every Apify plan, no volume tiers.
You pay per question row. Duplicates removed inside a run and every diagnostic row are not charged, and a search that matches nothing does not bill for results. Questions below minScore are never fetched, so they cost nothing either. The accepted answer adds nothing to the price.
💡 What people use it for
- Finding the questions a product's documentation should have answered, by searching its name and
reading what people are stuck on.
- Comparing tag volume between two libraries before committing to one.
- Building an evaluation set for a coding model out of plain-text question bodies.
- Watching a tag weekly with
sortByoncreationto see what problems are new.
🚧 What it does not do
- The accepted answer only, and only when you ask for it. Other answers and comments are not
read. answerCount tells you how many answers there are.
- One site per run. Searching Stack Overflow and Ask Ubuntu together means two runs.
- No date range input. Sort by
creationand filter oncreatedAtafterwards. - Only six of the network's sites, the ones in the list above.
- Tags are AND, never OR. Three tags usually returns far less than you expect.
- There is a daily ceiling on how much anyone can pull without a key. When it is reached the run
stops and says so instead of failing. Accepted answers count against it too, about one extra read per hundred questions.
- Rows are a snapshot. Scores, views and answer counts move.
🧭 Which developer data scraper do you need?
| If you want | Use |
|---|---|
| Stack Exchange questions, tags and scores | This one |
| Repositories, stars and issues from GitHub | GitHub Scraper |
| npm and PyPI package metadata and download counts | Package Registry Scraper |
| Developer articles and their tags | DEV Community Scraper |
❓ Questions people ask
Do I need a Stack Exchange API key? No.
Can I search tags without keywords? Yes, and it is the better way to pull a tag's top questions. Leave the query empty and fill tags.
Why did two tags return almost nothing? Because both have to be on the same question. Drop one.
Can I get the answers? The accepted one, yes: turn on includeAcceptedAnswer. The others are not read, so take url from the row and open the question.
Why does my title contain "? The title is passed through as the API sends it. Decode it on your side, or use body, which is already plain text.
Is this legal? Stack Exchange publishes this through a public API meant to be read, and the content is under a Creative Commons licence that asks for attribution. Apify's write-up on scraping and the law is a good starting point, and we are not lawyers.
🆘 If something breaks
Open the Issues tab on the actor page. Send the query, the tags and the run id. The errorCode on the diagnostic row usually names the problem by itself.