Request a tool
All toolsAutomationsGuidesMCP serverRequest a toolPlatformsCategories
Stack Overflow / Stack Exchange Scraper icon

Stack Overflow / Stack Exchange Scraper

Search Stack Overflow and Stack Exchange by keyword or tag. Score, answer count, views, reputation and body text. $2 per 1,000 questions.

100 runs on Apify $0.002 per question ($2 / 1,000)
Run this in the cloudRun on Apify →

Developer & Research Tools

How it works

  1. 1
    Open it on Apify

    Hit Run on Apify — it opens the tool in the cloud, no install.

  2. 2
    Set the inputs

    Adjust query, site, tags (sensible defaults are pre-filled).

  3. 3
    Click Run

    The tool runs on Apify’s cloud and collects the data for you.

  4. 4
    Export the results

    Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.

Pricing

$0.002 per question = $2 per 1,000

You are charged forWhenPrice
Question returnedCharged per question returned.$0.002

Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-06-13, and they are what you are actually charged.

Inputs

FieldWhat it doesType
queryKeywords to search for in question titles and bodies (e.g. "async await", "git rebase conflict"). Can be left empty if you provide one or more Tags instead.string
siteWhich Stack Exchange network site to search.string
tagsComma-separated tags to filter by, e.g. "javascript,promise" or "python,pandas". Optional. A question must carry ALL listed tags. You can search by tags alone with an empty query.string
sortByOrdering of results. "votes" = highest score first, "relevance" = best keyword match, "creation" = newest, "activity" = most recently active.string
minScoreOnly return questions with at least this score (net votes). Leave empty for no minimum. Works only with Sort by = Votes: Stack Exchange applies a minimum to whatever the results are sorted by, so with another order the run refuses it before searching.integer
includeAcceptedAnswerAdds the answer the asker accepted to each question row: its score, author and body as plain text. Questions with no accepted answer get null. Same price per question. Other answers and comments are not read.boolean
maxItemsMaximum number of questions to return. The actor paginates the API (100 per page) until this many are collected or there are no more results.integer
notionConnectorOptional. Write each question as a page into your Notion when the run finishes. Authorize a Notion connector once in Settings → API & Integrations → MCP connectors, then pick it here. Leave empty to skip (default), results are always saved to the dataset regardless.string
notionParentIdOptional. The Notion data source ID of the database to write into (only used if a Notion connector is set). Leave empty to create the pages privately in your workspace instead.string

What you get

A structured dataset — each result includes fields like:

answerCountbodycreatedAtisAnsweredownerNameownerReputationquestionIdscoretagstitleurlviewCount

Export every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.

2 ready-to-run use cases

Python Pandas Questions on Stack Overflow by Votes

Highest-scoring questions carrying both the python and pandas tags, with score, views, answer count and the question body as plain text.

Stack Overflow API Search for async/await Questions

Keyword search through the official Stack Exchange API, ranked by relevance, with score, views, tags and body. Questions only, no answer text.

Related tools in Developer & Research Tools

Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.

Package Registry Scraper (npm + PyPI) iconDeveloper & Research Tools

Package Registry Scraper (npm + PyPI)

Get npm and PyPI package metadata as JSON. Version, license, author, repo, keywords and npm monthly downloads. $2 per 1,000 packages.

2 use cases

arXiv Scraper iconDeveloper & Research Tools

arXiv Scraper

Search arXiv papers by title, author, abstract or category. Get full abstracts, authors, categories, DOI, dates and PDF links. $2 per 1,000 papers.

2 use cases

OpenAlex Scholarly Works Scraper iconDeveloper & Research Tools

OpenAlex Scholarly Works Scraper

Search 250M+ OpenAlex papers with no API key. Get titles, authors, venue, year, citations, DOI, OA links and full abstracts. $2.00 per 1,000 papers.

2 use cases

Crossref Scholarly Works Scraper iconDeveloper & Research Tools

Crossref Scholarly Works Scraper

Search 150M+ papers on Crossref: DOI, title, authors, journal, publisher, date, citations and abstract. No API key. $1.00 per 1,000 works.

2 use cases

Wikipedia Scraper iconDeveloper & Research Tools

Wikipedia Scraper

Search Wikipedia by keyword or by exact title. Get the intro text, the full article, thumbnails and categories. Any language. $1.00 per 1,000 pages.

3 use cases

Internet Archive Scraper iconDeveloper & Research Tools

Internet Archive Scraper

Search Internet Archive (archive.org) for books, audio, film and web items. Title, creator, year, downloads, subjects and URL. $2.00 per 1,000 items.

2 use cases

See all Developer & Research Tools →

Stack Overflow Scraper: questions, tags, scores and the question body

Search Stack Overflow and five sibling Stack Exchange sites by keywords, by tags, or by both. Each row is one question: the title, the score, how many answers it has, the view count, its tag list, who asked it and their reputation, and the whole question body as plain text.

Questions first. Turn on includeAcceptedAnswer and each row also carries the accepted answer as plain text. The other answers and the comments are not in the output, so if you need whole threads, this is the wrong purchase however cheap it is.

InputKeywords, tags, or both
OutputOne row per question
Ceiling1,000 questions per run
Account neededNone, and no Stack Exchange key
Price$2.00 per 1,000 questions, flat on every plan

🔍 What Stack Overflow Scraper does

It runs one search and pages through it until it has as many questions as you asked for.

query matches words in titles and bodies. tags filters, and it is an AND: put javascript,promise and you get only questions carrying both, not either. You can use tags on their own with no keywords, which is the cleanest way to pull the top questions in a tag.

sortBy decides the order, and it changes what a small run gives you. votes returns the long-standing classics. creation returns what was asked most recently. activity returns what people are still arguing about. relevance leans on the keyword match. On votes you can also set minScore to keep only questions at or above a score.

Six sites are available: Stack Overflow, Server Fault, Super User, Ask Ubuntu, MathOverflow and Software Engineering. One site per run.

With includeAcceptedAnswer on, each question also carries the answer its asker accepted: score, author and the body as plain text. Many questions never get one, and those say so with null.

📥 What you give it

{
  "query": "async await",
  "tags": "javascript,promise",
  "site": "stackoverflow",
  "sortBy": "votes",
  "maxItems": 50
}
FieldDefaultWhat it is
querynoneKeywords matched against titles and bodies. The Console shows async await as an example, but that is a prefill, so an API call has to send its own.
tagsnoneComma separated, and every tag must be present on the question. The Console prefills javascript, which again does not travel to an API call.
sitestackoverflowOne of stackoverflow, serverfault, superuser, askubuntu, mathoverflow, softwareengineering.
sortByvotesvotes, relevance, creation or activity. Always highest or newest first.
minScorenoneOnly questions with at least this score. Stack Exchange applies it, so lower-scored questions are never fetched or billed. Works with sortBy set to votes only; with another order the run refuses it, because the site would read the number as a date.
includeAcceptedAnswerfalseAdds each question's accepted answer to its row, at the same price per question. Works with every site and sort.
maxItems501 to 1,000. Paging is automatic, a hundred at a time.
notionConnectornoneOptional. Writes every delivered question into your Notion. Authorise a connector once under Settings, API and Integrations, MCP connectors, then pick it here.
notionParentIdnoneOptional. The Notion data source id to write into. Leave it empty and the pages are created privately in your workspace.
proxyConfigurationoffOptional network setting. Off is right for a normal run.

Send a query, tags, or both. Leaving both empty is not refused, and what comes back then is whatever the site decides to hand over, charged like any other rows. It is the one input mistake here that costs money.

📤 What you get back

A real row from a recent run, with the body cut short here for length:

{
  "ok": true,
  "questionId": 37576685,
  "title": "Using async/await with a forEach loop",
  "url": "https://stackoverflow.com/questions/37576685/using-async-await-with-a-foreach-loop",
  "score": 3353,
  "answerCount": 35,
  "viewCount": 2481345,
  "isAnswered": true,
  "tags": ["javascript", "node.js", "promise", "async-await", "ecmascript-2017"],
  "ownerName": "Saad",
  "ownerReputation": 55029,
  "createdAt": "2016-06-01T18:55:58.000Z",
  "body": "Are there any issues with using async/await in a forEach loop? I'm trying to loop through an array of files and await on the contents of each file. ..."
}
FieldWhat it is
bodyThe question, converted from HTML to plain text. Code blocks survive as text, indentation mostly does not.
isAnsweredStack Exchange's own flag: true when an answer was accepted or has a positive score, so it can be true with no accepted answer. Not the same as answerCount being above zero.
acceptedAnswerIdOnly with includeAcceptedAnswer. The accepted answer's id, or null when the question has none.
acceptedAnswerOnly with includeAcceptedAnswer. {answerId, score, ownerName, ownerReputation, createdAt, body}, with body as plain text. null when there is no accepted answer, or when it could not be read in this run (then acceptedAnswerId is still set).
ownerName, ownerReputationBoth null when the account was deleted or the post is anonymous.
scoreNet votes, which can be negative.
titlePassed through as the API sends it, so an HTML entity like " can survive into this field. Only body is decoded.
tagsThe question's full tag list, which is usually wider than the tags you filtered on.

🧾 Reading the output

Questions carry ok: true. Anything with ok: false carries an errorCode and is not charged.

CodeWhat it means
NO_RESULTSThe search ran and matched nothing. Usually a tag that does not exist, or two tags no single question carries.
BAD_INPUTminScore was set with a sortBy other than votes. Nothing was searched; the row's note says what to change.
RATE_LIMITEDStack Exchange asked the run to stop for the day. Whatever was already collected is in the dataset and complete as far as it goes.
NETWORKStack Exchange was unreachable or answered with something unusable. Re-run it.

A RATE_LIMITED row after some questions is not a failed run. It means the run delivered what it could and stopped rather than spinning. The same goes for an acceptedAnswerId with an empty acceptedAnswer: that answer could not be read in this run, usually because the day's allowance ran out. The run log says how many were read.

The Console's Overview table hides tags, isAnswered and body. Download the dataset as JSON, CSV or Excel, or switch the table to all fields.

▶️ How to run it

1. Open Stack Overflow Scraper and click Try for free. 2. Type into Search query, or fill Tags instead, or both. 3. Pick a Stack Exchange site if you want one of the other five. 4. Choose a Sort by order and set Max questions, then click Start. 5. Download the dataset as JSON, CSV, Excel or XML.

💰 How much does it cost?

$2.00 per 1,000 questions. Flat on every Apify plan, no volume tiers.

You pay per question row. Duplicates removed inside a run and every diagnostic row are not charged, and a search that matches nothing does not bill for results. Questions below minScore are never fetched, so they cost nothing either. The accepted answer adds nothing to the price.

💡 What people use it for

  • Finding the questions a product's documentation should have answered, by searching its name and

reading what people are stuck on.

  • Comparing tag volume between two libraries before committing to one.
  • Building an evaluation set for a coding model out of plain-text question bodies.
  • Watching a tag weekly with sortBy on creation to see what problems are new.

🚧 What it does not do

  • The accepted answer only, and only when you ask for it. Other answers and comments are not

read. answerCount tells you how many answers there are.

  • One site per run. Searching Stack Overflow and Ask Ubuntu together means two runs.
  • No date range input. Sort by creation and filter on createdAt afterwards.
  • Only six of the network's sites, the ones in the list above.
  • Tags are AND, never OR. Three tags usually returns far less than you expect.
  • There is a daily ceiling on how much anyone can pull without a key. When it is reached the run

stops and says so instead of failing. Accepted answers count against it too, about one extra read per hundred questions.

  • Rows are a snapshot. Scores, views and answer counts move.

🧭 Which developer data scraper do you need?

If you wantUse
Stack Exchange questions, tags and scoresThis one
Repositories, stars and issues from GitHubGitHub Scraper
npm and PyPI package metadata and download countsPackage Registry Scraper
Developer articles and their tagsDEV Community Scraper

❓ Questions people ask

Do I need a Stack Exchange API key? No.

Can I search tags without keywords? Yes, and it is the better way to pull a tag's top questions. Leave the query empty and fill tags.

Why did two tags return almost nothing? Because both have to be on the same question. Drop one.

Can I get the answers? The accepted one, yes: turn on includeAcceptedAnswer. The others are not read, so take url from the row and open the question.

Why does my title contain "? The title is passed through as the API sends it. Decode it on your side, or use body, which is already plain text.

Is this legal? Stack Exchange publishes this through a public API meant to be read, and the content is under a Creative Commons licence that asks for attribution. Apify's write-up on scraping and the law is a good starting point, and we are not lawyers.

🆘 If something breaks

Open the Issues tab on the actor page. Send the query, the tags and the run id. The errorCode on the diagnostic row usually names the problem by itself.