Request a tool
All toolsAutomationsGuidesMCP serverRequest a toolPlatformsCategories
Internet Archive Scraper icon

Internet Archive Scraper

Search Internet Archive (archive.org) for books, audio, film and web items. Title, creator, year, downloads, subjects and URL. $2.00 per 1,000 items.

161 runs on Apify $0.002 per item ($2 / 1,000)
Run this in the cloudRun on Apify →

Developer & Research Tools

How it works

  1. 1
    Open it on Apify

    Hit Run on Apify — it opens the tool in the cloud, no install.

  2. 2
    Set the inputs

    Adjust query, mediaType, sort (sensible defaults are pre-filled).

  3. 3
    Click Run

    The tool runs on Apify’s cloud and collects the data for you.

  4. 4
    Export the results

    Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.

Pricing

$0.002 per item = $2 per 1,000

You are charged forWhenPrice
Item returnedCharged per archive item returned.$0.002

Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-06-13, and they are what you are actually charged.

Inputs

FieldWhat it doesType
queryKeywords to search the Internet Archive for (e.g. "nasa apollo", "jazz"). Supports Lucene operators used by archive.org, e.g. "title:(grateful dead) AND year:[1977 TO 1980]". Required.string
mediaTypeRestrict results to one media type, or leave empty for any. texts = books/documents, audio = music/recordings, movies = video/film, software, image, web (archived sites), data, collection.string
sortOrder of results. downloads = most-downloaded first, date = newest item date first, publicdate = most recently added to archive.org first, relevance = the archive's default relevance ranking.string
maxItemsMaximum number of unique items to return. The actor paginates 100 per request until this many items are collected or the result set is exhausted.integer
notionConnectorOptional. Write each item as a page into your Notion when the run finishes. Authorize a Notion connector once in Settings → API & Integrations → MCP connectors, then pick it here. Leave empty to skip (default), results are always saved to the dataset regardless.string
notionParentIdOptional. The Notion data source ID of the database to write into (only used if a Notion connector is set). Leave empty to create the pages privately in your workspace instead.string

What you get

A structured dataset — each result includes fields like:

creatordatedescriptiondownloadsidentifiermediaTypepublicdatesubjectstitleurlyear

Export every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.

2 ready-to-run use cases

Download Books From Archive.org: Search Texts to JSON

Public-domain books by keyword with title, item link and download count. Author and year come from whoever uploaded, so about a quarter lack them.

Internet Archive Search, Newest Uploads First

Any archive.org topic sorted by the date it was added, with title, upload date and item link. Handy for watching a subject for fresh material.

Related tools in Developer & Research Tools

Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.

Hacker News Scraper iconDeveloper & Research Tools

Hacker News Scraper

Search HN stories, Show HN, Ask HN and comments, or pull the front page. Get points, author, comment counts and links. $1 per 1,000 items.

3 use cases

Research MCP Server — 15 Tools for AI Agents iconDeveloper & Research Tools

Research MCP Server — 15 Tools for AI Agents

One MCP endpoint gives your AI agent fifteen live research tools. Papers, code, news, SEC filings, packages and crypto prices.

4 use cases

DEV.to Scraper iconDeveloper & Research Tools

DEV.to Scraper

Scrape DEV.to articles by tag, author or sort. Fields: title, URL, tags, reactions, comments, reading time, cover image, full body. $2.00 per 1,000 articles.

Ready to run — no setup

Wikidata Scraper iconDeveloper & Research Tools

Wikidata Scraper

Search Wikidata for people, companies, places, books or films. Get the ID, the name, other names, a description and the Wikipedia link. $0.20 per 1,000.

Ready to run — no setup

Domain Inspector iconDeveloper & Research Tools

Domain Inspector

Check many domains at once. Get DNS, WHOIS registrar and expiry, TLS dates, redirects, security headers, robots and tech. $1.50 per 1,000.

Ready to run — no setup

GitHub Scraper iconDeveloper & Research Tools

GitHub Scraper

Search GitHub repos and users: stars, forks, language, topics, licence, plus user bio, company and followers. No token needed. $0.90 per 1,000 rows.

18 use cases

See all Developer & Research Tools →

Internet Archive Scraper: search archive.org for books, audio, film and archived web

Search archive.org the way its own advanced search works, and get the results as rows: the identifier, the title, the creator, the year, the download count, the subject tags and the link. Filter to one media type, sort by downloads, by date or by what was added most recently.

The metadata is written by whoever uploaded the item, so it is uneven. Plenty of items have no creator, no year and no description, and that is the archive rather than a gap here. The sample row further down is exactly that case.

InputSearch keywords, or a Lucene query
OutputOne row per archive.org item
Ceiling10,000 items per run
Account neededNone, and no API key
Price$2.00 per 1,000 items, flat on every plan

🔍 What Internet Archive Scraper does

Plain keywords work. So does the archive's own query syntax, which is worth knowing if you are doing anything precise: title:(grateful dead) AND year:[1977 TO 1980] does what it looks like it does.

Pick a media type to stay inside books, audio, film, software, images, archived web pages, datasets or collections, or leave it empty and get everything. Sorting is by downloads, by the item's own date, by when it was added to the archive, or by the archive's relevance ranking.

The run pages a hundred at a time until it has the number of unique items you asked for or the result set runs out.

📥 What you give it

{
  "query": "title:(apollo 11) AND year:[1969 TO 1972]",
  "mediaType": "movies",
  "sort": "downloads",
  "maxItems": 500
}
FieldDefaultWhat it is
querybox starts at nasa apolloKeywords, or an archive.org Lucene query. Required.
mediaTypeanytexts, audio, movies, software, image, web, data or collection. Empty means any.
sortdownloadsdownloads, date for the item's own date, publicdate for recently added, or relevance.
maxItems100How many unique items to return, up to 10,000.
notionConnectornoneOptional. Writes each item into your Notion once the run finishes.
notionParentIdnoneOptional. The Notion data source to write into.
proxyConfigurationoffOptional network settings. Off by default, and a normal run does not need it.

An OR covers your whole query. When you pick a media type, a query with an OR in it is bracketed before the type is added, so jazz OR blues with audio returns jazz and blues recordings, all of them audio.

📤 What you get back

A real row from a recent run:

{
  "ok": true,
  "identifier": "apolloaudiocollection",
  "title": "Apollo",
  "creator": null,
  "year": null,
  "date": null,
  "mediaType": "collection",
  "downloads": 1259223,
  "subjects": [],
  "description": null,
  "publicdate": "2010-12-06T19:02:28Z",
  "url": "https://archive.org/details/apolloaudiocollection"
}
FieldWhat it is
identifierThe archive's permanent key for the item, and the part of the URL that matters. Use it to dedupe and to fetch the item's files elsewhere.
creator, year, datenull whenever the uploader left them blank, which is often.
mediaTypeWhich kind of item it is. collection means a container of other items, not a single thing you can play or read.
downloadsThe archive's count, and 0 when the field was missing altogether.
subjectsThe uploader's tags, as an array. Frequently empty.
descriptionThe first 500 characters as the uploader wrote it, HTML and all, so strip it before displaying it.
publicdateWhen the item was added to archive.org, which is not the same as when it was made.

🧾 Reading the output

Two kinds of row land in your dataset, and ok tells them apart.

RowHow to spot itCharged
An itemok: true and an identifieryes
A diagnosticok: false and an errorCodeno

The overview table in the Apify console shows the item columns only, so a diagnostic row looks blank there. Switch to the JSON or All fields view to read it.

CodeWhat it means
BAD_INPUTNo query, or a media type that is not one of the listed values.
NO_RESULTSThe search worked and nothing matched. The numFound field on the row shows the archive's own count.
RATE_LIMITEDarchive.org asked for a slower pace than the run could keep. Try a smaller run.
SERVER_ERRORarchive.org answered 5xx. Usually passes.
BLOCKEDarchive.org refused the request. Re-run it.
NETWORKarchive.org was unreachable, or answered with something that was not the JSON it promised. The details field says which.

▶️ How to run it

1. Open Internet Archive Scraper and click Try for free. 2. Type keywords into Search query, or paste an archive.org query if you have one. 3. Pick a Media type unless you want everything. 4. Set Max items, choose a Sort by, then click Start. 5. Download the dataset as JSON, CSV or Excel, or read it from the Apify API.

💰 How much does it cost?

$2.00 per 1,000 items. Flat on every Apify plan, no volume tiers.

You pay per item delivered. Items that appear twice across pages are dropped before they are counted, diagnostic rows are not charged, and a search that matches nothing is not charged.

💡 What people use it for

  • Finding every archived recording of a band or a broadcaster, sorted by what people actually play.
  • Building a reading list of public domain books on a subject, with the identifiers to fetch them.
  • Checking whether a site was captured, using the web media type.
  • Tracking what a collection has gained recently, with sort on publicdate and a schedule.
  • Handing a model a list of source material with links, rather than a page of search results.

🚧 What it does not do

  • Metadata and a link, never the files. No books, audio, video or page captures are downloaded.
  • It does not read Wayback Machine snapshots. Archived sites appear as items under the web

media type, not as captures of a URL on a given date.

  • Uploader metadata is patchy. Missing creators, years and descriptions are normal.
  • downloads cannot tell you zero from missing. Both read 0.
  • Collections are containers. A collection row is a shelf, not an item, and it shows up whenever

you leave the media type empty.

  • description keeps the uploader's HTML and stops at 500 characters, which can cut mid-tag.
  • 10,000 items per run, and archive.org's own result set can end sooner.
  • No file lists, sizes or formats. Take the identifier to the archive's item page for those.

🧭 Which archive scraper do you need?

If you wantUse
Anything held on archive.org, by keywordThis one
Book metadata, ISBNs and coversBooks Scraper
Reader ratings and reviews for booksGoodreads Books Scraper
Scholarly papers with abstracts and citationsOpenAlex Scraper
Wikipedia articles as clean textWikipedia Scraper

❓ Questions people ask

Do I need an archive.org account? No. This reads the public search that anyone can use.

Can I download the books or the audio? Not from this actor. Take the identifier from each row and fetch the item's files from archive.org directly.

Why do some rows have no creator or year? Because nobody typed them in when the item was uploaded. The archive accepts what it is given.

Can I use the archive's advanced query syntax? Yes, the query goes through as written. Field searches and date ranges both work.

Is this legal? The search index is public and this reads it the way a browser does. What you may then do with a given item depends on that item's own rights, which vary enormously across the archive. Apify's write-up on scraping and the law is a good starting point, and we are not lawyers.

🆘 If something breaks

Open the Issues tab on the actor page. Send the run ID and the query you used. The errorCode on the diagnostic row usually names the problem on its own.