Request a tool
All toolsAutomationsGuidesMCP serverRequest a toolPlatformsCategories
Wayback Machine Scraper icon

Wayback Machine Scraper

List every Wayback Machine snapshot of a page, a path or a whole domain: capture time, status, type and archive link, with date and status filters. No key.

20 runs on Apify $0.00013 per capture ($0.13 / 1,000)
Run this in the cloudRun on Apify →

Developer & Research Tools

How it works

  1. 1
    Open it on Apify

    Hit Run on Apify — it opens the tool in the cloud, no install.

  2. 2
    Set the inputs

    Adjust urls, matchType, from (sensible defaults are pre-filled).

  3. 3
    Click Run

    The tool runs on Apify’s cloud and collects the data for you.

  4. 4
    Export the results

    Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.

Pricing

$0.00013 per capture = $0.13 per 1,000

You are charged forWhenPrice
CaptureOne archived capture returned in the dataset. Sample rows, notes and repeated captures are not charged.$0.00013

Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-10-03, and they are what you are actually charged.

Inputs

FieldWhat it doesType
urlsThe pages or sites to look up, one per line. A full link and a bare address find the same captures: https://www.nasa.gov/ and nasa.gov are one page to the archive. End an address with /* for everything under that path, or start it with *. for a whole domain and its subdomains. Up to 500 addresses a run.array
matchTypeExact is that one page. A path takes every page whose address starts with it. A host takes everything on that host, and a domain adds every subdomain. A /* or *. typed into an entry wins over this for that entry.string
fromOnly captures taken on or after this date, in UTC. Write 2015, 2015-06 or 2015-06-01. A date in any other format is refused before the run starts. Leave it empty to start at the first capture.string
toOnly captures taken up to and including this date, in UTC. 2019 means through the last day of 2019. Leave it empty to run to the latest capture.string
statusCodesOnly captures that got these answers when the archive took them, such as 200 or 404, or a whole class written 3xx. Leave it empty for all. Captures the archive stored as unchanged since the one before carry no status, so any status filter leaves them out.array
mimeTypesOnly captures of these content types, written in full: text/html, application/pdf, or a family such as image/*. A bare "html" matches nothing. Leave it empty for all.array
collapseEvery capture lists them all. Only when the content changed drops a capture identical to the one before it. One per day, month or year keeps the first capture of each period, page by page. One per address lists every archived address once with its first capture, which is the quick way to list every page a site ever had.string
maxCapturesPerUrlThe most captures from one address, or from one path, host or domain. A busy home page has hundreds of thousands, so this is your main spending cap.integer
maxCapturesThe most captures in the whole run, across all your entries. Each capture returned is one charge.integer
newestFirstStart from the latest capture and work back. Only for exact pages. With a thinning option on, it keeps the latest capture of each period instead of the first.boolean

What you get

A structured dataset — each result includes fields like:

inputUrlcapturedAtstatusCodemimeTypeoriginalUrlarchiveUrllengthdigesterrorCodemessage

Export every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.

Related tools in Developer & Research Tools

Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.

GitHub Scraper iconDeveloper & Research Tools

GitHub Scraper

Search GitHub repos and users: stars, forks, language, topics, licence, plus user bio, company and followers. No token needed. $0.90 per 1,000 rows.

18 use cases

Stack Overflow / Stack Exchange Scraper iconDeveloper & Research Tools

Stack Overflow / Stack Exchange Scraper

Search Stack Overflow and Stack Exchange by keyword or tag. Score, answer count, views, reputation and body text. $2 per 1,000 questions.

2 use cases

Package Registry Scraper (npm + PyPI) iconDeveloper & Research Tools

Package Registry Scraper (npm + PyPI)

Get npm and PyPI package metadata as JSON. Version, license, author, repo, keywords and npm monthly downloads. $2 per 1,000 packages.

2 use cases

arXiv Scraper iconDeveloper & Research Tools

arXiv Scraper

Search arXiv papers by title, author, abstract or category. Get full abstracts, authors, categories, DOI, dates and PDF links. $2 per 1,000 papers.

2 use cases

OpenAlex Scholarly Works Scraper iconDeveloper & Research Tools

OpenAlex Scholarly Works Scraper

Search 250M+ OpenAlex papers with no API key. Get titles, authors, venue, year, citations, DOI, OA links and full abstracts. $2.00 per 1,000 papers.

2 use cases

Crossref Scholarly Works Scraper iconDeveloper & Research Tools

Crossref Scholarly Works Scraper

Search 150M+ papers on Crossref: DOI, title, authors, journal, publisher, date, citations and abstract. No API key. $1.00 per 1,000 works.

2 use cases

See all Developer & Research Tools →

Give it pages or whole sites and get every capture the Wayback Machine holds of them, one row per capture: when it was taken, the HTTP status, the content type, the size and a link to the archived copy. It works on one page, everything under a path, a whole host, or a domain with its subdomains, with date, status and content type filters.

It lists captures and does not download them, so each row is a link to the copy rather than the copy itself. And a few sites are excluded from the archive or keep their capture list private, which no setting gets around.

InputWeb pages or sites: one page, a path, a host or a whole domain
OutputOne row per capture: capture time, the URL as captured, status, content type, digest, stored size, archive links
Ceiling100,000 captures per entry, 1,000,000 per run, 500 entries
Speed3.5 to 5 seconds per page or site looked up, with or without captures, so a full list of 500 takes about forty minutes. Give a long list a longer timeout
Account neededNone from you
Price$0.13 per 1,000 captures, flat on every plan. The free plan's $5 a month covers about 38,000

🔍 What Wayback Machine Scraper does

For each page or site it reads the Wayback Machine's capture index, the list behind the archive's own calendar pages, and turns each entry into a row. A busy home page has hundreds of thousands of captures, so the run reads them a few thousand at a time, oldest first, until it has the number you asked for or the list ends.

The archive applies your dates and filters, and every row is checked again before it is kept. A capture outside your dates, with another status or another content type, is never charged.

To make a long history readable, keep one capture per day, month or year, keep only the captures where the content changed, or list each archived URL once. On a path or a domain this is done page by page, so one page's capture never hides another's.

The run keeps to a couple of dozen requests a minute, well inside what the archive asks of automated tools. If the archive asks for a pause anyway, the run stops there and lists the entries it did not reach.

📋 What data you get from each Wayback Machine capture

What you getField
The page or site you gave, and how it was looked upinputUrl, matchType
The URL as the archive captured it, and the archive's key for itoriginalUrl, urlKey
When the capture was takentimestamp, capturedAt
The HTTP status and the content type the archive gotstatusCode, mimeType
A fingerprint of the content, and the stored sizedigest, length
The capture in the archive's viewer, and as first servedarchiveUrl, rawArchiveUrl
When the row was readscrapedAt

▶️ How to scrape the Wayback Machine

1. Open Wayback Machine Scraper and click Try for free. 2. Paste your pages or sites into Web addresses, one per line. 3. Pick What each address covers, then add dates or filters if you want them. 4. Set Captures per address and click Start. 5. Download the dataset as JSON, CSV or Excel, or read it from the Apify API.

💰 How much does it cost to scrape the Wayback Machine?

$0.13 per 1,000 captures. Flat on every Apify plan, no volume tiers. On the free plan, the $5 Apify gives you each month covers about 38,000 captures.

You pay per capture delivered. The sample row, diagnostic rows, captures the checks drop and repeats are not charged, and a page with no captures adds nothing. Captures per address and Captures in total are the two caps on what one run can spend.

📥 What you give it

{
  "urls": ["https://www.nasa.gov/", "apify.com/store/*"],
  "matchType": "exact",
  "from": "2015",
  "to": "2020-06",
  "statusCodes": ["200"],
  "mimeTypes": ["text/html"],
  "collapse": "month",
  "maxCapturesPerUrl": 1000,
  "maxCaptures": 100000,
  "newestFirst": false
}
FieldDefaultWhat it is
urlsnone, the form starts with https://www.nasa.gov/The pages or sites, one per line, up to 500. nasa.gov and https://www.nasa.gov/ are the same page to the archive. End one with /* for everything under that path, or start it with *. for a domain and its subdomains.
matchTypeexactexact for that page, prefix for everything under the path, host for every page on the host, domain for the host and its subdomains. A /* or *. typed into an entry wins over this.
from, tononeDates in UTC, written 2015, 2015-06 or 2015-06-01. to covers the whole period, so 2019 runs to the last second of 2019. A date in any other format is refused before the run starts; an impossible one, like 2021-02-30, stops the run before it looks anything up.
statusCodesallHTTP codes such as 200 or 404, or a class such as 3xx.
mimeTypesallContent types in full, such as text/html or application/pdf, or a family such as image/*. A bare html matches nothing.
collapsenone, the form starts on yeardigest drops a capture identical to the one before it. day, month and year keep the first capture of each period. url lists each archived URL once, with its first capture.
maxCapturesPerUrl1000, the form starts at 100The most captures from one page, path, host or domain, up to 100,000.
maxCaptures100000The most captures in the whole run, up to 1,000,000.
newestFirstfalseStart from the latest capture and work back. Exact pages only. With thinning on, it keeps the latest capture of each period rather than the first.

A status filter leaves out unchanged captures. When the archive finds the same content again it often stores a revisit record, which has no status code and the type warc/revisit. Without a filter those rows are part of the history. Ask for any status and they drop out.

📤 What you get back

A real row from a run on 3 October 2026: nasa.gov's first capture.

{
  "recordType": "capture",
  "inputUrl": "https://www.nasa.gov/",
  "matchType": "exact",
  "originalUrl": "http://www.nasa.gov:80/",
  "urlKey": "gov,nasa)/",
  "timestamp": "19961231235847",
  "capturedAt": "1996-12-31T23:58:47Z",
  "statusCode": 200,
  "mimeType": "text/html",
  "digest": "MGIGF4GRGGF5GKV6VNCBAXOE3OR5BTZC",
  "length": 1811,
  "archiveUrl": "https://web.archive.org/web/19961231235847/http://www.nasa.gov:80/",
  "rawArchiveUrl": "https://web.archive.org/web/19961231235847id_/http://www.nasa.gov:80/",
  "scrapedAt": "2026-10-03T05:21:54.874Z"
}
FieldHow to read it
inputUrl, matchTypeThe entry as you typed it and how it was looked up, so rows group back to your list.
originalUrlThe URL as the archive captured it, with scheme, port and query string.
urlKeyThe archive's own key for the page. Two spellings of one page share it.
timestamp, capturedAtWhen the capture was taken: the archive's 14-digit stamp, and the same moment in UTC.
statusCodeThe HTTP status the archive got. null on revisit records.
mimeTypeThe content type it got, or warc/revisit.
digestA fingerprint of the content. Equal digests mean identical content.
lengthThe size of the stored record in bytes, compressed. It is not the size of the page.
rawArchiveUrlThe same capture as it was first served, without the archive's toolbar or rewritten links.

🧾 Reading the output

RowHow to spot itBilled
A capturerecordType is captureyes
The samplerecordType is sample, with _sample: true. Only when no pages were givenno
A diagnosticrecordType is diagnostic, with _diagnostic: true and an errorCodeno

Keep the rows whose recordType is capture to get the captures alone.

CodeWhat it means
BAD_INPUTAn entry that isn't a usable web page or site, or a setting the actor can't read. The message says which.
NO_CAPTURESThe archive has nothing for that page or site, or nothing that matches your dates and filters.
EXCLUDEDThe site is excluded from the Wayback Machine, so its captures can't be listed.
RESTRICTEDThe Wayback Machine keeps this site's capture list private. Seen on theguardian.com.
NOT_ANSWEREDThe archive didn't answer for that page or site this time. Try it again later, and narrow a big domain with dates or a path.
REFUSEDThe archive turned the lookup down. Rare, and worth telling us about.
PARTIALA later page of a long list couldn't be read. The captures before it are in the dataset.
NONE_MATCHEDThe archive listed captures, but none of them matched what was asked.
NOT_LOOKED_UPThe run stopped before it reached that entry: a cap, the time limit, or a pause the archive asked for.
MAX_CHARGE_TOO_LOWThe maximum charge set for the run doesn't cover one capture, so nothing was looked up.

One page can appear twice in the same second, once over http and once over https. The archive holds both, so both are listed. The run report (RUN_REPORT in the key-value store) gives the outcome for each entry, including whether it stopped at your cap.

💡 What people use it for

  • Working out when a page changed, by listing only the captures where its content did.
  • Rebuilding redirects after a migration, from a list of every page the old site ever had.
  • Checking a domain's past before buying it: when it was first captured, when it went quiet, what it

served in between.

  • Finding the captures around a date you have to cite, with links that open the page as it was.

Checking a domain before you buy it, in three steps:

1. Run this actor on the domain's home page with collapse set to year, to see when it was first captured and when the captures stop. 2. Put the same domain into the domains field of Domain Inspector for today's registrar and expiry date, in rdap.registrar and rdap.expiresAt. 3. Open the archiveUrl of the last few captures before it went quiet to see what it was serving.

🚧 What it does not do

  • It lists captures and does not fetch them. No HTML, text or files come back. The links do.
  • Only what the archive holds. A page it never captured, or captured under a spelling you didn't

give, won't be there.

  • Thinning a big path or domain can stop short. The archive can only thin a site as one long

list, so the actor reads the captures itself and keeps one per page per period. It stops after reading 25 captures for each one you asked for, 200,000 at most, and the run report says when.

  • Newest first is for exact pages only.
  • Excluded sites stay excluded. Some owners have asked the archive to hide their site, and those

come back as EXCLUDED. A few large sites have their capture lists kept private by the archive, and those come back as RESTRICTED.

  • length is the stored size, compressed, not the size of the page.
  • A whole domain with a filter that matches little can be too big to answer. The archive gives up

on a query after a minute, and such an entry comes back NOT_ANSWERED. Narrow it with dates or a path.

🧭 Which archive scraper do you need?

If you wantUse
The captures of a page or a site in the Wayback MachineThis one
Books, audio, film and other items held on archive.orgInternet Archive Scraper
A site's pages as they read today, as clean textWebsite Intelligence Crawler
DNS, registration and certificate details for a domainDomain Inspector
A screenshot of a page as it looks nowWebsite Screenshot Generator

❓ Questions people ask

Do I need an archive.org account or a key?

No. The capture index is public.

How do I see a page as it looked on a given day?

Find the row with the date you want and open its archiveUrl. For the original bytes without the archive's toolbar, use rawArchiveUrl.

Why do some captures have no status code?

They are revisit records: the archive saw the same content again and stored a pointer to the earlier copy. Their mimeType is warc/revisit.

What does excluded mean?

The site's owner asked the Wayback Machine not to show it. The archive answers the same way for everyone.

Can I list every page a website ever had?

Yes. Give the domain, set matchType to domain and collapse to url, which lists each archived URL once.

Can I call it from code or connect it to an AI assistant?

Yes. The API tab has ready-made code for Python, JavaScript and the command line. For Claude, ChatGPT or another MCP client, connect https://mcp.apify.com/?tools=fetch-actor-details,dami_studio/wayback-machine-scraper. Either way the run happens on your Apify account at the same price.

Is scraping the Wayback Machine legal?

The capture index is public and this reads it the way the archive's own calendar does. What you may do with an archived page depends on the page. Apify's write-up on the legality of web scraping is a good starting point, and we are not lawyers.

🆘 If something breaks

Open the Issues tab on the actor page. Send the run ID and the pages or sites you used. The errorCode on a diagnostic row usually names the problem on its own.