Internet Archive Scraper
Search Internet Archive (archive.org) for books, audio, film and web items. Title, creator, year, downloads, subjects and URL. $2.00 per 1,000 items.
How it works
- 1Open it on Apify
Hit Run on Apify — it opens the tool in the cloud, no install.
- 2Set the inputs
Adjust
query,mediaType,sort(sensible defaults are pre-filled). - 3Click Run
The tool runs on Apify’s cloud and collects the data for you.
- 4Export the results
Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.
Pricing
$0.002 per item = $2 per 1,000
| You are charged for | When | Price |
|---|---|---|
| Item returned | Charged per archive item returned. | $0.002 |
Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-06-13, and they are what you are actually charged.
Inputs
| Field | What it does | Type |
|---|---|---|
query | Keywords to search the Internet Archive for (e.g. "nasa apollo", "jazz"). Supports Lucene operators used by archive.org, e.g. "title:(grateful dead) AND year:[1977 TO 1980]". Required. | string |
mediaType | Restrict results to one media type, or leave empty for any. texts = books/documents, audio = music/recordings, movies = video/film, software, image, web (archived sites), data, collection. | string |
sort | Order of results. downloads = most-downloaded first, date = newest item date first, publicdate = most recently added to archive.org first, relevance = the archive's default relevance ranking. | string |
maxItems | Maximum number of unique items to return. The actor paginates 100 per request until this many items are collected or the result set is exhausted. | integer |
notionConnector | Optional. Write each item as a page into your Notion when the run finishes. Authorize a Notion connector once in Settings → API & Integrations → MCP connectors, then pick it here. Leave empty to skip (default), results are always saved to the dataset regardless. | string |
notionParentId | Optional. The Notion data source ID of the database to write into (only used if a Notion connector is set). Leave empty to create the pages privately in your workspace instead. | string |
What you get
A structured dataset — each result includes fields like:
creatordatedescriptiondownloadsidentifiermediaTypepublicdatesubjectstitleurlyearExport every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.
2 ready-to-run use cases
Download Books From Archive.org: Search Texts to JSON
Public-domain books by keyword with title, item link and download count. Author and year come from whoever uploaded, so about a quarter lack them.
Internet Archive Search, Newest Uploads First
Any archive.org topic sorted by the date it was added, with title, upload date and item link. Handy for watching a subject for fresh material.
Related tools in Developer & Research Tools
Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.
Hacker News Scraper
Search HN stories, Show HN, Ask HN and comments, or pull the front page. Get points, author, comment counts and links. $1 per 1,000 items.
Research MCP Server — 15 Tools for AI Agents
One MCP endpoint gives your AI agent fifteen live research tools. Papers, code, news, SEC filings, packages and crypto prices.
DEV.to Scraper
Scrape DEV.to articles by tag, author or sort. Fields: title, URL, tags, reactions, comments, reading time, cover image, full body. $2.00 per 1,000 articles.
Wikidata Scraper
Search Wikidata for people, companies, places, books or films. Get the ID, the name, other names, a description and the Wikipedia link. $0.20 per 1,000.
Domain Inspector
Check many domains at once. Get DNS, WHOIS registrar and expiry, TLS dates, redirects, security headers, robots and tech. $1.50 per 1,000.
GitHub Scraper
Search GitHub repos and users: stars, forks, language, topics, licence, plus user bio, company and followers. No token needed. $0.90 per 1,000 rows.
Where this tool sits
- Categories
- Developer & Research Tools
- Platforms
- Academic & Research
Internet Archive Scraper: search archive.org for books, audio, film and archived web
Search archive.org the way its own advanced search works, and get the results as rows: the identifier, the title, the creator, the year, the download count, the subject tags and the link. Filter to one media type, sort by downloads, by date or by what was added most recently.
The metadata is written by whoever uploaded the item, so it is uneven. Plenty of items have no creator, no year and no description, and that is the archive rather than a gap here. The sample row further down is exactly that case.
| Input | Search keywords, or a Lucene query |
| Output | One row per archive.org item |
| Ceiling | 10,000 items per run |
| Account needed | None, and no API key |
| Price | $2.00 per 1,000 items, flat on every plan |
🔍 What Internet Archive Scraper does
Plain keywords work. So does the archive's own query syntax, which is worth knowing if you are doing anything precise: title:(grateful dead) AND year:[1977 TO 1980] does what it looks like it does.
Pick a media type to stay inside books, audio, film, software, images, archived web pages, datasets or collections, or leave it empty and get everything. Sorting is by downloads, by the item's own date, by when it was added to the archive, or by the archive's relevance ranking.
The run pages a hundred at a time until it has the number of unique items you asked for or the result set runs out.
📥 What you give it
{
"query": "title:(apollo 11) AND year:[1969 TO 1972]",
"mediaType": "movies",
"sort": "downloads",
"maxItems": 500
}
| Field | Default | What it is |
|---|---|---|
query | box starts at nasa apollo | Keywords, or an archive.org Lucene query. Required. |
mediaType | any | texts, audio, movies, software, image, web, data or collection. Empty means any. |
sort | downloads | downloads, date for the item's own date, publicdate for recently added, or relevance. |
maxItems | 100 | How many unique items to return, up to 10,000. |
notionConnector | none | Optional. Writes each item into your Notion once the run finishes. |
notionParentId | none | Optional. The Notion data source to write into. |
proxyConfiguration | off | Optional network settings. Off by default, and a normal run does not need it. |
An OR covers your whole query. When you pick a media type, a query with an OR in it is bracketed before the type is added, so jazz OR blues with audio returns jazz and blues recordings, all of them audio.
📤 What you get back
A real row from a recent run:
{
"ok": true,
"identifier": "apolloaudiocollection",
"title": "Apollo",
"creator": null,
"year": null,
"date": null,
"mediaType": "collection",
"downloads": 1259223,
"subjects": [],
"description": null,
"publicdate": "2010-12-06T19:02:28Z",
"url": "https://archive.org/details/apolloaudiocollection"
}
| Field | What it is |
|---|---|
identifier | The archive's permanent key for the item, and the part of the URL that matters. Use it to dedupe and to fetch the item's files elsewhere. |
creator, year, date | null whenever the uploader left them blank, which is often. |
mediaType | Which kind of item it is. collection means a container of other items, not a single thing you can play or read. |
downloads | The archive's count, and 0 when the field was missing altogether. |
subjects | The uploader's tags, as an array. Frequently empty. |
description | The first 500 characters as the uploader wrote it, HTML and all, so strip it before displaying it. |
publicdate | When the item was added to archive.org, which is not the same as when it was made. |
🧾 Reading the output
Two kinds of row land in your dataset, and ok tells them apart.
| Row | How to spot it | Charged |
|---|---|---|
| An item | ok: true and an identifier | yes |
| A diagnostic | ok: false and an errorCode | no |
The overview table in the Apify console shows the item columns only, so a diagnostic row looks blank there. Switch to the JSON or All fields view to read it.
| Code | What it means |
|---|---|
BAD_INPUT | No query, or a media type that is not one of the listed values. |
NO_RESULTS | The search worked and nothing matched. The numFound field on the row shows the archive's own count. |
RATE_LIMITED | archive.org asked for a slower pace than the run could keep. Try a smaller run. |
SERVER_ERROR | archive.org answered 5xx. Usually passes. |
BLOCKED | archive.org refused the request. Re-run it. |
NETWORK | archive.org was unreachable, or answered with something that was not the JSON it promised. The details field says which. |
▶️ How to run it
1. Open Internet Archive Scraper and click Try for free. 2. Type keywords into Search query, or paste an archive.org query if you have one. 3. Pick a Media type unless you want everything. 4. Set Max items, choose a Sort by, then click Start. 5. Download the dataset as JSON, CSV or Excel, or read it from the Apify API.
💰 How much does it cost?
$2.00 per 1,000 items. Flat on every Apify plan, no volume tiers.
You pay per item delivered. Items that appear twice across pages are dropped before they are counted, diagnostic rows are not charged, and a search that matches nothing is not charged.
💡 What people use it for
- Finding every archived recording of a band or a broadcaster, sorted by what people actually play.
- Building a reading list of public domain books on a subject, with the identifiers to fetch them.
- Checking whether a site was captured, using the
webmedia type. - Tracking what a collection has gained recently, with
sortonpublicdateand a schedule. - Handing a model a list of source material with links, rather than a page of search results.
🚧 What it does not do
- Metadata and a link, never the files. No books, audio, video or page captures are downloaded.
- It does not read Wayback Machine snapshots. Archived sites appear as items under the
web
media type, not as captures of a URL on a given date.
- Uploader metadata is patchy. Missing creators, years and descriptions are normal.
downloadscannot tell you zero from missing. Both read0.- Collections are containers. A
collectionrow is a shelf, not an item, and it shows up whenever
you leave the media type empty.
descriptionkeeps the uploader's HTML and stops at 500 characters, which can cut mid-tag.- 10,000 items per run, and archive.org's own result set can end sooner.
- No file lists, sizes or formats. Take the
identifierto the archive's item page for those.
🧭 Which archive scraper do you need?
| If you want | Use |
|---|---|
| Anything held on archive.org, by keyword | This one |
| Book metadata, ISBNs and covers | Books Scraper |
| Reader ratings and reviews for books | Goodreads Books Scraper |
| Scholarly papers with abstracts and citations | OpenAlex Scraper |
| Wikipedia articles as clean text | Wikipedia Scraper |
❓ Questions people ask
Do I need an archive.org account? No. This reads the public search that anyone can use.
Can I download the books or the audio? Not from this actor. Take the identifier from each row and fetch the item's files from archive.org directly.
Why do some rows have no creator or year? Because nobody typed them in when the item was uploaded. The archive accepts what it is given.
Can I use the archive's advanced query syntax? Yes, the query goes through as written. Field searches and date ranges both work.
Is this legal? The search index is public and this reads it the way a browser does. What you may then do with a given item depends on that item's own rights, which vary enormously across the archive. Apify's write-up on scraping and the law is a good starting point, and we are not lawyers.
🆘 If something breaks
Open the Issues tab on the actor page. Send the run ID and the query you used. The errorCode on the diagnostic row usually names the problem on its own.