Wayback Machine Scraper
List every Wayback Machine snapshot of a page, a path or a whole domain: capture time, status, type and archive link, with date and status filters. No key.
How it works
- 1Open it on Apify
Hit Run on Apify — it opens the tool in the cloud, no install.
- 2Set the inputs
Adjust
urls,matchType,from(sensible defaults are pre-filled). - 3Click Run
The tool runs on Apify’s cloud and collects the data for you.
- 4Export the results
Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.
Pricing
$0.00013 per capture = $0.13 per 1,000
| You are charged for | When | Price |
|---|---|---|
| Capture | One archived capture returned in the dataset. Sample rows, notes and repeated captures are not charged. | $0.00013 |
Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-10-03, and they are what you are actually charged.
Inputs
| Field | What it does | Type |
|---|---|---|
urls | The pages or sites to look up, one per line. A full link and a bare address find the same captures: https://www.nasa.gov/ and nasa.gov are one page to the archive. End an address with /* for everything under that path, or start it with *. for a whole domain and its subdomains. Up to 500 addresses a run. | array |
matchType | Exact is that one page. A path takes every page whose address starts with it. A host takes everything on that host, and a domain adds every subdomain. A /* or *. typed into an entry wins over this for that entry. | string |
from | Only captures taken on or after this date, in UTC. Write 2015, 2015-06 or 2015-06-01. A date in any other format is refused before the run starts. Leave it empty to start at the first capture. | string |
to | Only captures taken up to and including this date, in UTC. 2019 means through the last day of 2019. Leave it empty to run to the latest capture. | string |
statusCodes | Only captures that got these answers when the archive took them, such as 200 or 404, or a whole class written 3xx. Leave it empty for all. Captures the archive stored as unchanged since the one before carry no status, so any status filter leaves them out. | array |
mimeTypes | Only captures of these content types, written in full: text/html, application/pdf, or a family such as image/*. A bare "html" matches nothing. Leave it empty for all. | array |
collapse | Every capture lists them all. Only when the content changed drops a capture identical to the one before it. One per day, month or year keeps the first capture of each period, page by page. One per address lists every archived address once with its first capture, which is the quick way to list every page a site ever had. | string |
maxCapturesPerUrl | The most captures from one address, or from one path, host or domain. A busy home page has hundreds of thousands, so this is your main spending cap. | integer |
maxCaptures | The most captures in the whole run, across all your entries. Each capture returned is one charge. | integer |
newestFirst | Start from the latest capture and work back. Only for exact pages. With a thinning option on, it keeps the latest capture of each period instead of the first. | boolean |
What you get
A structured dataset — each result includes fields like:
inputUrlcapturedAtstatusCodemimeTypeoriginalUrlarchiveUrllengthdigesterrorCodemessageExport every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.
Related tools in Developer & Research Tools
Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.
GitHub Scraper
Search GitHub repos and users: stars, forks, language, topics, licence, plus user bio, company and followers. No token needed. $0.90 per 1,000 rows.
Stack Overflow / Stack Exchange Scraper
Search Stack Overflow and Stack Exchange by keyword or tag. Score, answer count, views, reputation and body text. $2 per 1,000 questions.
Package Registry Scraper (npm + PyPI)
Get npm and PyPI package metadata as JSON. Version, license, author, repo, keywords and npm monthly downloads. $2 per 1,000 packages.
arXiv Scraper
Search arXiv papers by title, author, abstract or category. Get full abstracts, authors, categories, DOI, dates and PDF links. $2 per 1,000 papers.
OpenAlex Scholarly Works Scraper
Search 250M+ OpenAlex papers with no API key. Get titles, authors, venue, year, citations, DOI, OA links and full abstracts. $2.00 per 1,000 papers.
Crossref Scholarly Works Scraper
Search 150M+ papers on Crossref: DOI, title, authors, journal, publisher, date, citations and abstract. No API key. $1.00 per 1,000 works.
Where this tool sits
- Categories
- Developer & Research Tools
Give it pages or whole sites and get every capture the Wayback Machine holds of them, one row per capture: when it was taken, the HTTP status, the content type, the size and a link to the archived copy. It works on one page, everything under a path, a whole host, or a domain with its subdomains, with date, status and content type filters.
It lists captures and does not download them, so each row is a link to the copy rather than the copy itself. And a few sites are excluded from the archive or keep their capture list private, which no setting gets around.
| Input | Web pages or sites: one page, a path, a host or a whole domain |
| Output | One row per capture: capture time, the URL as captured, status, content type, digest, stored size, archive links |
| Ceiling | 100,000 captures per entry, 1,000,000 per run, 500 entries |
| Speed | 3.5 to 5 seconds per page or site looked up, with or without captures, so a full list of 500 takes about forty minutes. Give a long list a longer timeout |
| Account needed | None from you |
| Price | $0.13 per 1,000 captures, flat on every plan. The free plan's $5 a month covers about 38,000 |
🔍 What Wayback Machine Scraper does
For each page or site it reads the Wayback Machine's capture index, the list behind the archive's own calendar pages, and turns each entry into a row. A busy home page has hundreds of thousands of captures, so the run reads them a few thousand at a time, oldest first, until it has the number you asked for or the list ends.
The archive applies your dates and filters, and every row is checked again before it is kept. A capture outside your dates, with another status or another content type, is never charged.
To make a long history readable, keep one capture per day, month or year, keep only the captures where the content changed, or list each archived URL once. On a path or a domain this is done page by page, so one page's capture never hides another's.
The run keeps to a couple of dozen requests a minute, well inside what the archive asks of automated tools. If the archive asks for a pause anyway, the run stops there and lists the entries it did not reach.
📋 What data you get from each Wayback Machine capture
| What you get | Field |
|---|---|
| The page or site you gave, and how it was looked up | inputUrl, matchType |
| The URL as the archive captured it, and the archive's key for it | originalUrl, urlKey |
| When the capture was taken | timestamp, capturedAt |
| The HTTP status and the content type the archive got | statusCode, mimeType |
| A fingerprint of the content, and the stored size | digest, length |
| The capture in the archive's viewer, and as first served | archiveUrl, rawArchiveUrl |
| When the row was read | scrapedAt |
▶️ How to scrape the Wayback Machine
1. Open Wayback Machine Scraper and click Try for free. 2. Paste your pages or sites into Web addresses, one per line. 3. Pick What each address covers, then add dates or filters if you want them. 4. Set Captures per address and click Start. 5. Download the dataset as JSON, CSV or Excel, or read it from the Apify API.
💰 How much does it cost to scrape the Wayback Machine?
$0.13 per 1,000 captures. Flat on every Apify plan, no volume tiers. On the free plan, the $5 Apify gives you each month covers about 38,000 captures.
You pay per capture delivered. The sample row, diagnostic rows, captures the checks drop and repeats are not charged, and a page with no captures adds nothing. Captures per address and Captures in total are the two caps on what one run can spend.
📥 What you give it
{
"urls": ["https://www.nasa.gov/", "apify.com/store/*"],
"matchType": "exact",
"from": "2015",
"to": "2020-06",
"statusCodes": ["200"],
"mimeTypes": ["text/html"],
"collapse": "month",
"maxCapturesPerUrl": 1000,
"maxCaptures": 100000,
"newestFirst": false
}
| Field | Default | What it is |
|---|---|---|
urls | none, the form starts with https://www.nasa.gov/ | The pages or sites, one per line, up to 500. nasa.gov and https://www.nasa.gov/ are the same page to the archive. End one with /* for everything under that path, or start it with *. for a domain and its subdomains. |
matchType | exact | exact for that page, prefix for everything under the path, host for every page on the host, domain for the host and its subdomains. A /* or *. typed into an entry wins over this. |
from, to | none | Dates in UTC, written 2015, 2015-06 or 2015-06-01. to covers the whole period, so 2019 runs to the last second of 2019. A date in any other format is refused before the run starts; an impossible one, like 2021-02-30, stops the run before it looks anything up. |
statusCodes | all | HTTP codes such as 200 or 404, or a class such as 3xx. |
mimeTypes | all | Content types in full, such as text/html or application/pdf, or a family such as image/*. A bare html matches nothing. |
collapse | none, the form starts on year | digest drops a capture identical to the one before it. day, month and year keep the first capture of each period. url lists each archived URL once, with its first capture. |
maxCapturesPerUrl | 1000, the form starts at 100 | The most captures from one page, path, host or domain, up to 100,000. |
maxCaptures | 100000 | The most captures in the whole run, up to 1,000,000. |
newestFirst | false | Start from the latest capture and work back. Exact pages only. With thinning on, it keeps the latest capture of each period rather than the first. |
A status filter leaves out unchanged captures. When the archive finds the same content again it often stores a revisit record, which has no status code and the type warc/revisit. Without a filter those rows are part of the history. Ask for any status and they drop out.
📤 What you get back
A real row from a run on 3 October 2026: nasa.gov's first capture.
{
"recordType": "capture",
"inputUrl": "https://www.nasa.gov/",
"matchType": "exact",
"originalUrl": "http://www.nasa.gov:80/",
"urlKey": "gov,nasa)/",
"timestamp": "19961231235847",
"capturedAt": "1996-12-31T23:58:47Z",
"statusCode": 200,
"mimeType": "text/html",
"digest": "MGIGF4GRGGF5GKV6VNCBAXOE3OR5BTZC",
"length": 1811,
"archiveUrl": "https://web.archive.org/web/19961231235847/http://www.nasa.gov:80/",
"rawArchiveUrl": "https://web.archive.org/web/19961231235847id_/http://www.nasa.gov:80/",
"scrapedAt": "2026-10-03T05:21:54.874Z"
}
| Field | How to read it |
|---|---|
inputUrl, matchType | The entry as you typed it and how it was looked up, so rows group back to your list. |
originalUrl | The URL as the archive captured it, with scheme, port and query string. |
urlKey | The archive's own key for the page. Two spellings of one page share it. |
timestamp, capturedAt | When the capture was taken: the archive's 14-digit stamp, and the same moment in UTC. |
statusCode | The HTTP status the archive got. null on revisit records. |
mimeType | The content type it got, or warc/revisit. |
digest | A fingerprint of the content. Equal digests mean identical content. |
length | The size of the stored record in bytes, compressed. It is not the size of the page. |
rawArchiveUrl | The same capture as it was first served, without the archive's toolbar or rewritten links. |
🧾 Reading the output
| Row | How to spot it | Billed |
|---|---|---|
| A capture | recordType is capture | yes |
| The sample | recordType is sample, with _sample: true. Only when no pages were given | no |
| A diagnostic | recordType is diagnostic, with _diagnostic: true and an errorCode | no |
Keep the rows whose recordType is capture to get the captures alone.
| Code | What it means |
|---|---|
BAD_INPUT | An entry that isn't a usable web page or site, or a setting the actor can't read. The message says which. |
NO_CAPTURES | The archive has nothing for that page or site, or nothing that matches your dates and filters. |
EXCLUDED | The site is excluded from the Wayback Machine, so its captures can't be listed. |
RESTRICTED | The Wayback Machine keeps this site's capture list private. Seen on theguardian.com. |
NOT_ANSWERED | The archive didn't answer for that page or site this time. Try it again later, and narrow a big domain with dates or a path. |
REFUSED | The archive turned the lookup down. Rare, and worth telling us about. |
PARTIAL | A later page of a long list couldn't be read. The captures before it are in the dataset. |
NONE_MATCHED | The archive listed captures, but none of them matched what was asked. |
NOT_LOOKED_UP | The run stopped before it reached that entry: a cap, the time limit, or a pause the archive asked for. |
MAX_CHARGE_TOO_LOW | The maximum charge set for the run doesn't cover one capture, so nothing was looked up. |
One page can appear twice in the same second, once over http and once over https. The archive holds both, so both are listed. The run report (RUN_REPORT in the key-value store) gives the outcome for each entry, including whether it stopped at your cap.
💡 What people use it for
- Working out when a page changed, by listing only the captures where its content did.
- Rebuilding redirects after a migration, from a list of every page the old site ever had.
- Checking a domain's past before buying it: when it was first captured, when it went quiet, what it
served in between.
- Finding the captures around a date you have to cite, with links that open the page as it was.
Checking a domain before you buy it, in three steps:
1. Run this actor on the domain's home page with collapse set to year, to see when it was first captured and when the captures stop. 2. Put the same domain into the domains field of Domain Inspector for today's registrar and expiry date, in rdap.registrar and rdap.expiresAt. 3. Open the archiveUrl of the last few captures before it went quiet to see what it was serving.
🚧 What it does not do
- It lists captures and does not fetch them. No HTML, text or files come back. The links do.
- Only what the archive holds. A page it never captured, or captured under a spelling you didn't
give, won't be there.
- Thinning a big path or domain can stop short. The archive can only thin a site as one long
list, so the actor reads the captures itself and keeps one per page per period. It stops after reading 25 captures for each one you asked for, 200,000 at most, and the run report says when.
- Newest first is for exact pages only.
- Excluded sites stay excluded. Some owners have asked the archive to hide their site, and those
come back as EXCLUDED. A few large sites have their capture lists kept private by the archive, and those come back as RESTRICTED.
lengthis the stored size, compressed, not the size of the page.- A whole domain with a filter that matches little can be too big to answer. The archive gives up
on a query after a minute, and such an entry comes back NOT_ANSWERED. Narrow it with dates or a path.
🧭 Which archive scraper do you need?
| If you want | Use |
|---|---|
| The captures of a page or a site in the Wayback Machine | This one |
| Books, audio, film and other items held on archive.org | Internet Archive Scraper |
| A site's pages as they read today, as clean text | Website Intelligence Crawler |
| DNS, registration and certificate details for a domain | Domain Inspector |
| A screenshot of a page as it looks now | Website Screenshot Generator |
❓ Questions people ask
Do I need an archive.org account or a key?
No. The capture index is public.
How do I see a page as it looked on a given day?
Find the row with the date you want and open its archiveUrl. For the original bytes without the archive's toolbar, use rawArchiveUrl.
Why do some captures have no status code?
They are revisit records: the archive saw the same content again and stored a pointer to the earlier copy. Their mimeType is warc/revisit.
What does excluded mean?
The site's owner asked the Wayback Machine not to show it. The archive answers the same way for everyone.
Can I list every page a website ever had?
Yes. Give the domain, set matchType to domain and collapse to url, which lists each archived URL once.
Can I call it from code or connect it to an AI assistant?
Yes. The API tab has ready-made code for Python, JavaScript and the command line. For Claude, ChatGPT or another MCP client, connect https://mcp.apify.com/?tools=fetch-actor-details,dami_studio/wayback-machine-scraper. Either way the run happens on your Apify account at the same price.
Is scraping the Wayback Machine legal?
The capture index is public and this reads it the way the archive's own calendar does. What you may do with an archived page depends on the page. Apify's write-up on the legality of web scraping is a good starting point, and we are not lawyers.
🆘 If something breaks
Open the Issues tab on the actor page. Send the run ID and the pages or sites you used. The errorCode on a diagnostic row usually names the problem on its own.