Website Intelligence Crawler
Crawl a public website into clean text, Markdown and chunks ready for embedding. No API key, no LLM. Follows robots.txt. $0.60 per 1,000 pages.
How it works
- 1Open it on Apify
Hit Run on Apify — it opens the tool in the cloud, no install.
- 2Set the inputs
Adjust
startUrls,urls,maxPages(sensible defaults are pre-filled). - 3Click Run
The tool runs on Apify’s cloud and collects the data for you.
- 4Export the results
Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.
Pricing
$0.0006 per page crawled = $0.6 per 1,000
| You are charged for | When | Price |
|---|---|---|
| Page crawled | Charged once per successfully crawled public page returned as a real row. Sample and diagnostic rows are never charged. | $0.0006 |
Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-08-10, and they are what you are actually charged.
Inputs
| Field | What it does | Type |
|---|---|---|
startUrls | Public HTTP(S) URLs to crawl. Values are normalized and deduplicated. | array |
urls | Additional public HTTP(S) URLs. Same-site links are followed from each seed. | array |
maxPages | Maximum successful or diagnostic page records per seed URL. | integer |
maxDepth | Maximum same-site link depth from each seed. Zero fetches only the seed page. | integer |
maxConcurrency | Maximum requests processed concurrently for each seed site. | integer |
useRobotsTxt | Honor public robots.txt disallow and crawl-delay rules. | boolean |
query | Whitespace-separated terms used for deterministic ranking. No AI or API key is used. | string |
chunkSizeChars | Target maximum characters in each deterministic content chunk. | integer |
chunkOverlapChars | Characters repeated between adjacent chunks to preserve context. | integer |
outputFormat | Keep clean text, Markdown, or both representations in each page record. | string |
requestTimeoutSecs | Maximum time allowed for each HTTP request. | integer |
maxResponseSizeKb | Maximum response body size accepted from one page in kilobytes. | integer |
fallbackToProxy | Retry blocked direct requests through the configured Apify Proxy. | boolean |
What you get
A structured dataset — each result includes fields like:
urltitledescriptiondepthstatuswordCountlinkschunkserrorCodeerrorok_samplediagnostics_noticerequestedUrlseedUrlsiteKeylanguagecanonicalUrltextmarkdownqueryoutputFormatExport every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.
Related tools in Developer & Research Tools
Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.
GitHub Scraper
Search GitHub repos and users: stars, forks, language, topics, licence, plus user bio, company and followers. No token needed. $0.90 per 1,000 rows.
Stack Overflow / Stack Exchange Scraper
Search Stack Overflow and Stack Exchange by keyword or tag. Score, answer count, views, reputation and body text. $2 per 1,000 questions.
Package Registry Scraper (npm + PyPI)
Get npm and PyPI package metadata as JSON. Version, license, author, repo, keywords and npm monthly downloads. $2 per 1,000 packages.
arXiv Scraper
Search arXiv papers by title, author, abstract or category. Get full abstracts, authors, categories, DOI, dates and PDF links. $2 per 1,000 papers.
OpenAlex Scholarly Works Scraper
Search 250M+ OpenAlex papers with no API key. Get titles, authors, venue, year, citations, DOI, OA links and full abstracts. $2.00 per 1,000 papers.
Crossref Scholarly Works Scraper
Search 150M+ papers on Crossref: DOI, title, authors, journal, publisher, date, citations and abstract. No API key. $1.00 per 1,000 works.
Where this tool sits
- Categories
- Developer & Research Tools
Website Intelligence Crawler: a public site as clean text, Markdown and chunks for embedding
Point it at a URL and it crawls the same-site pages and hands back each one as readable text, as Markdown, and as chunks already sized for an embedding model. Title, description, language, word count, the links it found and the status code ride along on every row.
No API key, no model call, nothing to authorise. The chunking and the relevance ranking are both plain arithmetic, so the same input gives the same output every time.
The one limit to understand first: there is no browser here. It reads the HTML the server sends. A page that builds its content in JavaScript after loading comes back with very little text, and no setting changes that. For ordinary documentation, blogs, marketing sites and knowledge bases, which is what most crawls are, the HTML is the content.
| Input | Public HTTP or HTTPS URLs to start from |
| Output | One row per page, with text, Markdown and chunks |
| Ceiling | 100 pages per starting URL, 5 levels of same-site links |
| Account needed | None |
| Price | $0.60 per 1,000 pages, which is $0.0006 each, flat on every plan |
🕸️ What Website Intelligence Crawler does
From each URL you give it, it follows same-site links down to the depth you set, and writes a record per page. It stays on the site it started on, so an outbound link to somebody else's domain is recorded in links and not followed.
The text is the readable part of the page, with navigation, scripts and boilerplate stripped. The Markdown keeps the heading structure, paragraphs and lists, which is the version you want if the pages are going into a document store.
Chunks are cut to the size you ask for with a little overlap so a sentence is not sliced in half. Each one carries its character count and an estimated token count, so you can budget an embedding run before you start it.
Give it a query and every chunk gets a relevanceScore against those words. It is deterministic term matching, not a model, and it is there to let you keep the 20 chunks that matter instead of all 400.
robots.txt is honoured by default, including crawl delay.
📥 What you give it
{
"startUrls": [{"url": "https://docs.example.com/"}],
"maxPages": 50,
"maxDepth": 2,
"outputFormat": "both",
"chunkSizeChars": 1200,
"query": "pricing plans billing"
}
| Field | Default | What it is |
|---|---|---|
startUrls | none | The pages to start from. Normalised and deduplicated. |
urls | none | More starting URLs, as plain strings, if that shape suits you better. |
maxPages | 10 | Records per starting URL, 1 to 100. Failed pages count too. |
maxDepth | 1 | Same-site link depth, 1 to 5. Setting it to 0 behaves the same as 1, so use maxPages: 1 when you want the seed page alone. |
maxConcurrency | 3 | Requests in flight per site, 1 to 10. Leave it low on small servers. |
useRobotsTxt | true | Honour Disallow rules and crawl delay. It does not interpret Allow exceptions, so a site that relies on those may be crawled less than it permits. |
query | none | Words to score chunks against. No model, no key. |
chunkSizeChars | 1200 | Target characters per chunk, 300 to 5,000. |
chunkOverlapChars | 120 | Characters repeated between neighbouring chunks, up to 800. Zero behaves as 120. |
outputFormat | both | json keeps the text, markdown keeps the Markdown, both keeps each. |
requestTimeoutSecs | 15 | Per request, 3 to 60. |
maxResponseSizeKb | 2048 | Biggest page body accepted, 128 to 8,192 KB. |
fallbackToProxy | false | Retry a refused page once by another route. Direct is always tried first. |
proxyConfiguration | none | Your own servers, for that retry, if you have them. |
Leave the URLs empty and you get one sample row showing the shape, with nothing fetched and no rows charged.
📤 What you get back
A real row from a real run:
{
"ok": true,
"requestedUrl": "https://example.com/",
"url": "https://example.com/",
"seedUrl": "https://example.com/",
"siteKey": "example.com",
"depth": 0,
"status": 200,
"title": "Example Domain",
"description": null,
"language": "en",
"canonicalUrl": "https://example.com/",
"text": "Example DomainThis domain is for use in documentation examples without needing permission. Avoid use in operations.Learn more",
"markdown": "# Example Domain\n\nThis domain is for use in documentation examples without needing permission. Avoid use in operations.\n\nLearn more",
"links": [],
"wordCount": 17,
"contentType": "text/html",
"chunks": [
{
"chunkIndex": 0,
"text": "Example DomainThis domain is for use in documentation examples without needing permission. Avoid use in operations.Learn more",
"charCount": 125,
"estimatedTokens": 32,
"relevanceScore": null
}
],
"query": null,
"outputFormat": "both",
"fetchedAt": "2026-09-06T17:08:09.247Z"
}
| Field | What it is |
|---|---|
requestedUrl, url, canonicalUrl | What was asked for, what answered after redirects, and what the page says its own address is. |
seedUrl, siteKey, depth | Which starting URL this page came from, the site it belongs to, and how many links away it was. |
text | The readable content, boilerplate removed. Absent when outputFormat is markdown. |
markdown | Headings, paragraphs and lists. Absent when outputFormat is json. A page built entirely of tables falls back to the plain text. |
links | Same-site links found on the page, with their anchor text, up to 250. |
chunks | chunkIndex, text, charCount, estimatedTokens, and relevanceScore when you gave a query. |
wordCount, language, status | Useful for filtering out thin pages and redirect landing pages before you embed anything. |
contentType | Always reads text/html, because HTML is the only thing kept. |
🧾 Reading the output
Page rows carry ok: true. Failures carry diagnostics: true with an errorCode, and are never charged.
| Row | How to spot it | Charged |
|---|---|---|
| A page | ok is true | yes |
| The sample row | _sample is true | no |
| A failed page | diagnostics is true, and errorCode says why | no |
| Code | What it means |
|---|---|
BAD_INPUT | Not a public HTTP or HTTPS address. Credentials in the URL and other schemes land here. |
ROBOTS_DISALLOWED | The site's own robots.txt asks crawlers not to read that path. |
UNSUPPORTED_CONTENT | A PDF, an image, JSON, anything that is not an HTML page. |
HTTP_ERROR | The server answered with an error status. |
NETWORK | The server could not be reached, or answered badly. |
TIMEOUT | Slower than requestTimeoutSecs. |
RESPONSE_TOO_LARGE | Bigger than maxResponseSizeKb. |
SSRF_BLOCKED | The address points somewhere private, so it was refused. |
One quirk of the table view: a failed row has no url, and the address it tried sits in requestedUrl, which the view does not show. Open the row or export the dataset to see which page it was.
▶️ How to run it
1. Open Website Intelligence Crawler and click Try for free. 2. Put your address into Start URLs. 3. Set Maximum pages and Maximum link depth. Ten pages at depth 1 is a good look before you commit to a whole site. 4. If the crawl is feeding a search or a RAG index, set Chunk size to whatever your embedding model likes. 5. Click Start, then download the dataset as JSON, CSV or Excel, or read it from the API.
💰 How much does it cost?
$0.60 per 1,000 pages, which is $0.0006 each. Flat on every Apify plan, no volume tiers.
You pay per page that came back. The sample row is not charged, and neither is a page that failed, was disallowed by robots.txt, was too large, or turned out to be a PDF. Those still arrive as rows so you know they happened.
Fifty pages is three cents. A hundred-page site crawled weekly is a quarter a month.
💡 What people use it for
- The dull half of a RAG pipeline: crawl the docs, get chunks with token counts, embed them.
- Keeping a support bot's knowledge current by re-crawling the help centre on a schedule.
- Content audits.
wordCount,title,descriptionandstatusacross every page, in one export. - Pulling a competitor's whole marketing site as Markdown to read properly.
- Finding the pages that actually mention a term, using
queryandrelevanceScore.
🚧 What it does not do
- No browser. JavaScript-rendered content is not there to read.
- Same site only. Links to other domains are recorded, not followed.
- HTML only. PDFs, Word files and images come back as unsupported rather than as text.
- Nothing behind a login. No cookies, no sign-in, no interactive challenges.
- 100 pages per starting URL. For a bigger site, give it several starting URLs, one per section.
- No screenshots. The screenshot generator linked below does that.
- No summarising or tagging. No model runs here, which is the point: nothing is invented and
nothing is rephrased.
- Depth 0 is not honoured. It behaves as depth 1.
🧭 Which site tool do you need?
| If you want | Use |
|---|---|
| A site as text, Markdown and chunks | This one |
| Pictures of the pages instead | Website Screenshot Generator |
| Repositories and developer profiles | GitHub Scraper |
| Questions and answers from Stack Overflow | Stack Overflow Scraper |
❓ Questions people ask
Does it run JavaScript? No. It reads the HTML the server sends, which is the whole content on most documentation and marketing sites.
Will it hammer my site? It keeps to robots.txt, including crawl delay, and you can drop maxConcurrency to 1.
What chunk size should I use? Whatever your embedding model is happiest with. 1,200 characters with 120 of overlap is a reasonable default, and estimatedTokens on each chunk tells you where you landed.
Can I crawl one page only? Set maxPages to 1. Depth 0 does not do it.
Do I get both text and Markdown? With outputFormat: both, yes, on the same row. Pick one if the rows are getting large.
Is crawling a public site legal? It reads public pages and honours robots.txt. Site terms still apply to what you do with the content afterwards. Apify's write-up on the legality of web scraping is a good starting point, and we are not lawyers.
🆘 If something breaks
Open the Issues tab on the actor page. Send the starting URL and the run ID. If a diagnostic row landed, its errorCode and requestedUrl usually name the reason already.