arXiv Scraper
Search arXiv papers by title, author, abstract or category. Get full abstracts, authors, categories, DOI, dates and PDF links. $2 per 1,000 papers.
How it works
- 1Open it on Apify
Hit Run on Apify — it opens the tool in the cloud, no install.
- 2Set the inputs
Adjust
query,sortBy,maxItems(sensible defaults are pre-filled). - 3Click Run
The tool runs on Apify’s cloud and collects the data for you.
- 4Export the results
Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.
Pricing
$0.002 per paper = $2 per 1,000
| You are charged for | When | Price |
|---|---|---|
| Paper returned | Charged per paper returned. | $0.002 |
Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-06-13, and they are what you are actually charged.
Inputs
| Field | What it does | Type |
|---|---|---|
query | arXiv search query. Use arXiv field prefixes: all: (all fields), ti: (title), au: (author), abs: (abstract), cat: (category). Examples: "all:large language models", "cat:cs.CL", "ti:attention AND au:vaswani". Combine terms with AND / OR / ANDNOT. | string |
sortBy | How to order results. Relevance ranks by match quality; Submitted date sorts by original submission; Last updated date sorts by most recent revision. Newest/most-relevant first. | string |
maxItems | Maximum number of papers to return. The actor paginates 100 per request and pauses ~3s between pages to respect arXiv's rate guidance. arXiv hard-limits total reachable results to about 30000. | integer |
notionConnector | Optional. Write each paper as a page into your Notion when the run finishes, handy for building a literature-review database. Authorize a Notion connector once in Settings → API & Integrations → MCP connectors, then pick it here. Leave empty to skip (default), results are always saved to the dataset regardless. | string |
notionParentId | Optional. The Notion data source ID of the database to write papers into (only used if a Notion connector is set). Leave empty to create the pages privately in your workspace instead. | string |
What you get
A structured dataset — each result includes fields like:
absUrlabstractarxivIdauthorscategoriesdoipdfUrlprimaryCategorypublishedAttitleupdatedAtExport every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.
2 ready-to-run use cases
LLM Papers on arXiv: Abstracts, Authors and PDF Links
Keyword search across arXiv for language-model work, ranked by relevance, with the abstract, the author list and a direct PDF link on every row.
NLP Papers From arXiv cs.CL, Newest Submissions First
Sorted by submission date, so the top of the run is what went up today. Abstract, authors, categories and a PDF link on each paper. No DOI field.
Related tools in Developer & Research Tools
Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.
OpenAlex Scholarly Works Scraper
Search 250M+ OpenAlex papers with no API key. Get titles, authors, venue, year, citations, DOI, OA links and full abstracts. $2.00 per 1,000 papers.
Crossref Scholarly Works Scraper
Search 150M+ papers on Crossref: DOI, title, authors, journal, publisher, date, citations and abstract. No API key. $1.00 per 1,000 works.
Wikipedia Scraper
Search Wikipedia by keyword or by exact title. Get the intro text, the full article, thumbnails and categories. Any language. $1.00 per 1,000 pages.
Internet Archive Scraper
Search Internet Archive (archive.org) for books, audio, film and web items. Title, creator, year, downloads, subjects and URL. $2.00 per 1,000 items.
Hacker News Scraper
Search HN stories, Show HN, Ask HN and comments, or pull the front page. Get points, author, comment counts and links. $1 per 1,000 items.
Research MCP Server — 15 Tools for AI Agents
One MCP endpoint gives your AI agent fifteen live research tools. Papers, code, news, SEC filings, packages and crypto prices.
Where this tool sits
- Categories
- Developer & Research Tools
- Platforms
- Academic & Research
arXiv Scraper: papers, full abstracts and PDF links for any arXiv search
Write an arXiv query and get one row per paper: the id, title, the whole abstract, every author, the categories, both dates and direct links to the abstract page and the PDF. No account, no API key.
Metadata and links only. The PDF link is there for you to fetch yourself; this actor does not download or read the paper.
| Input | An arXiv search query using its own field prefixes |
| Output | One row per paper |
| Ceiling | 30,000 papers, which is arXiv's own depth limit on one query |
| Account needed | None, and no API key |
| Price | $2.00 per 1,000 papers, flat on every plan |
🔬 What arXiv Scraper does
It runs your query against arXiv and pages through the results, pausing between pages because arXiv asks callers to. Duplicate papers across pages are dropped by id.
The query uses arXiv's own syntax, which is worth two minutes of your time because it is what makes this precise:
| Prefix | Searches |
|---|---|
all: | every field |
ti: | the title |
au: | the authors |
abs: | the abstract |
cat: | the category, like cs.CL or math.AG |
Combine them with AND, OR and ANDNOT. ti:attention AND au:vaswani does exactly what it looks like. Sort by relevance, by submission date or by last revision, newest first in both date orders.
📥 What you give it
{
"query": "all:large language models",
"sortBy": "relevance",
"maxItems": 50
}
| Field | Default | What it is |
|---|---|---|
query | box starts at all:large language models | The arXiv query. Use the prefixes above; a bare phrase with no prefix is not what arXiv expects. |
sortBy | relevance | relevance, submittedDate or lastUpdatedDate. Both date orders are newest first. |
maxItems | 50 | Papers to return, up to 30,000. A value above that is clamped, because arXiv will not page deeper on one query. |
notionConnector | none | Optional. Write every paper into your own Notion as well as the dataset, which is a quick way to start a literature review. |
notionParentId | none | Optional. The Notion data source ID to write into. |
proxyConfiguration | off | Optional, and off by default because arXiv is public and a normal run does not need it. |
There is no date-range field. To narrow by time, sort on submittedDate and cut the rows on publishedAt yourself.
📤 What you get back
A real row from a recent run, with the abstract cut short:
{
"ok": true,
"charged": true,
"arxivId": "2609.13144",
"title": "Type Diversity Enables Transformers to Generalise Compositionally",
"abstract": "Compositional generalisation has been divided into lexical and structural generalisation. Previous work has found that structural generalisation is harder than lexical for Transformers...",
"authors": ["Anssi Moisio", "Mathias Creutz", "Mikko Kurimo"],
"primaryCategory": "cs.CL",
"categories": ["cs.CL"],
"publishedAt": "2026-09-11T00:00:00.000Z",
"updatedAt": "2026-09-11T00:00:00.000Z",
"doi": null,
"absUrl": "https://arxiv.org/abs/2609.13144",
"pdfUrl": "https://arxiv.org/pdf/2609.13144"
}
| Field | What it is |
|---|---|
arxivId | The id with any trailing version marker removed, so 2609.13144v2 arrives as 2609.13144. Stable, and the dedupe key inside a run. |
abstract | The whole abstract, with line breaks collapsed into single spaces. |
publishedAt, updatedAt | First submission and latest revision. On a paper never revised they are the same. |
doi | null on most papers. arXiv only has one when the author added it after journal publication. |
absUrl | The abstract page, pointing at the specific version the search returned. |
pdfUrl | A direct PDF link. Nothing is downloaded for you. |
primaryCategory | The one category the author filed it under. categories holds all of them, deduplicated. |
🧾 Reading the output
Two kinds of row can land in your dataset.
| Row | How to spot it | What it is |
|---|---|---|
| A paper | ok: true with charged: true and an arxivId | a result |
| A diagnostic | ok: false and an errorCode | not a result, and not charged |
charged: true is simply the marker every paper row carries. Use it, or ok, to tell results from diagnostics. Do not read it as a record of what your account was billed.
| Code | What it means |
|---|---|
BAD_INPUT | The query was empty. Give one in arXiv syntax. |
NO_RESULTS | The query ran and matched no papers. Loosen it, or check the prefix spelling. |
NOT_FOUND | arXiv answered 404 for the request. |
RATE_LIMITED | arXiv asked for a slower pace. Re-run with a smaller maxItems. |
SERVER_ERROR | arXiv answered with a server error. Usually passes on its own. |
BLOCKED | arXiv would not serve the search this time. |
NETWORK | arXiv could not be reached. |
Rows arrive at the end. Everything is collected before anything is written, so a large run shows an empty dataset until it finishes. Keep a first run small while you get the query right.
The default table view has no ok column, so a diagnostic row renders as one blank line. Switch to All fields or export as JSON.
▶️ How to run it
1. Open arXiv Scraper and click Try for free. 2. Type a query into Search query, for example cat:cs.CL AND abs:retrieval. 3. Pick a Sort by. Use submittedDate when you care about what is new. 4. Set Max papers. Start around 20 while you tune the query, then click Start. 5. Download the dataset as JSON, CSV or Excel, or read it from the Apify API.
💰 How much does it cost?
$2.00 per 1,000 papers, which is $0.002 each. Flat on every Apify plan, no volume tiers.
You pay per paper row delivered. Duplicates across pages are dropped before they reach you, diagnostic rows are not charged, and a query that matches nothing costs you nothing.
💡 What people use it for
- Watching a category like
cat:cs.CLon a schedule and pulling everything new since yesterday. - Building a literature list on a topic, with every abstract already in the row to read or
summarise.
- Following one author with
au:to see what they have put up and when it was last revised. - Collecting PDF links for a set of papers to fetch and process in your own pipeline.
🚧 What it does not do
- No full text. Abstracts and links. The PDF is yours to fetch from
pdfUrl. - No citation counts, no references, no related-paper graph. arXiv does not publish them.
- No peer-review status. A paper on arXiv may be a preprint, a submitted manuscript or a
published article, and nothing on the row tells you which.
doiis usuallynull. Most preprints have none, and arXiv does not go looking.- No date-range filter. Sort by date and trim the rows yourself.
- Sort order is always newest or most relevant first. Ascending is not offered.
- About 30,000 results per query, maximum. That is arXiv's depth limit, not a setting here.
Split a broad query into narrower ones rather than raising the number.
- The run finishes even when every request failed, so read
okrather than the run status.
🧭 Which research scraper do you need?
| If you want | Use |
|---|---|
| arXiv preprints with abstracts and PDF links | This one |
| Published works with citation counts and open-access links | OpenAlex Scraper |
| Publisher metadata for a DOI | Crossref Scraper |
| Encyclopedia articles and their categories | Wikipedia Scraper |
| Patent filings by inventor or assignee | Google Patents Search Scraper |
❓ Questions people ask
Do I need an arXiv account or an API key? No. arXiv's search is open to anyone.
Why did my search return nothing? Usually the syntax. A bare phrase with no prefix is not a valid arXiv query; start it with all: and try again.
Can I get the paper itself? Not from this actor. pdfUrl gives you a direct link to fetch.
Which date should I sort on? submittedDate for when work first appeared, lastUpdatedDate for what was revised recently. They differ a lot on papers that keep being updated.
Can I schedule it? Yes. A daily cat: query sorted by submittedDate, deduped on arxivId, keeps a category feed current.
Is scraping arXiv legal? arXiv publishes metadata for public reuse and asks callers to be polite about request rate, which this actor is. Individual papers carry their own licences, so check those before republishing text. Apify's write-up on scraping and the law is a good starting point, and we are not lawyers.
🆘 If something breaks
Open the Issues tab on the actor page. Send the run ID and the exact query string you used. The errorCode and hint on the diagnostic row usually name the problem on their own.