Request a tool
All toolsAutomationsGuidesMCP serverRequest a toolPlatformsCategories
arXiv Scraper icon

arXiv Scraper

Search arXiv papers by title, author, abstract or category. Get full abstracts, authors, categories, DOI, dates and PDF links. $2 per 1,000 papers.

5 from 1 review on Apify 169 runs on Apify $0.002 per paper ($2 / 1,000)
Run this in the cloudRun on Apify →

Developer & Research Tools

How it works

  1. 1
    Open it on Apify

    Hit Run on Apify — it opens the tool in the cloud, no install.

  2. 2
    Set the inputs

    Adjust query, sortBy, maxItems (sensible defaults are pre-filled).

  3. 3
    Click Run

    The tool runs on Apify’s cloud and collects the data for you.

  4. 4
    Export the results

    Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.

Pricing

$0.002 per paper = $2 per 1,000

You are charged forWhenPrice
Paper returnedCharged per paper returned.$0.002

Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-06-13, and they are what you are actually charged.

Inputs

FieldWhat it doesType
queryarXiv search query. Use arXiv field prefixes: all: (all fields), ti: (title), au: (author), abs: (abstract), cat: (category). Examples: "all:large language models", "cat:cs.CL", "ti:attention AND au:vaswani". Combine terms with AND / OR / ANDNOT.string
sortByHow to order results. Relevance ranks by match quality; Submitted date sorts by original submission; Last updated date sorts by most recent revision. Newest/most-relevant first.string
maxItemsMaximum number of papers to return. The actor paginates 100 per request and pauses ~3s between pages to respect arXiv's rate guidance. arXiv hard-limits total reachable results to about 30000.integer
notionConnectorOptional. Write each paper as a page into your Notion when the run finishes, handy for building a literature-review database. Authorize a Notion connector once in Settings → API & Integrations → MCP connectors, then pick it here. Leave empty to skip (default), results are always saved to the dataset regardless.string
notionParentIdOptional. The Notion data source ID of the database to write papers into (only used if a Notion connector is set). Leave empty to create the pages privately in your workspace instead.string

What you get

A structured dataset — each result includes fields like:

absUrlabstractarxivIdauthorscategoriesdoipdfUrlprimaryCategorypublishedAttitleupdatedAt

Export every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.

2 ready-to-run use cases

LLM Papers on arXiv: Abstracts, Authors and PDF Links

Keyword search across arXiv for language-model work, ranked by relevance, with the abstract, the author list and a direct PDF link on every row.

NLP Papers From arXiv cs.CL, Newest Submissions First

Sorted by submission date, so the top of the run is what went up today. Abstract, authors, categories and a PDF link on each paper. No DOI field.

Related tools in Developer & Research Tools

Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.

OpenAlex Scholarly Works Scraper iconDeveloper & Research Tools

OpenAlex Scholarly Works Scraper

Search 250M+ OpenAlex papers with no API key. Get titles, authors, venue, year, citations, DOI, OA links and full abstracts. $2.00 per 1,000 papers.

2 use cases

Crossref Scholarly Works Scraper iconDeveloper & Research Tools

Crossref Scholarly Works Scraper

Search 150M+ papers on Crossref: DOI, title, authors, journal, publisher, date, citations and abstract. No API key. $1.00 per 1,000 works.

2 use cases

Wikipedia Scraper iconDeveloper & Research Tools

Wikipedia Scraper

Search Wikipedia by keyword or by exact title. Get the intro text, the full article, thumbnails and categories. Any language. $1.00 per 1,000 pages.

3 use cases

Internet Archive Scraper iconDeveloper & Research Tools

Internet Archive Scraper

Search Internet Archive (archive.org) for books, audio, film and web items. Title, creator, year, downloads, subjects and URL. $2.00 per 1,000 items.

2 use cases

Hacker News Scraper iconDeveloper & Research Tools

Hacker News Scraper

Search HN stories, Show HN, Ask HN and comments, or pull the front page. Get points, author, comment counts and links. $1 per 1,000 items.

3 use cases

Research MCP Server — 15 Tools for AI Agents iconDeveloper & Research Tools

Research MCP Server — 15 Tools for AI Agents

One MCP endpoint gives your AI agent fifteen live research tools. Papers, code, news, SEC filings, packages and crypto prices.

4 use cases

See all Developer & Research Tools →

arXiv Scraper: papers, full abstracts and PDF links for any arXiv search

Write an arXiv query and get one row per paper: the id, title, the whole abstract, every author, the categories, both dates and direct links to the abstract page and the PDF. No account, no API key.

Metadata and links only. The PDF link is there for you to fetch yourself; this actor does not download or read the paper.

InputAn arXiv search query using its own field prefixes
OutputOne row per paper
Ceiling30,000 papers, which is arXiv's own depth limit on one query
Account neededNone, and no API key
Price$2.00 per 1,000 papers, flat on every plan

🔬 What arXiv Scraper does

It runs your query against arXiv and pages through the results, pausing between pages because arXiv asks callers to. Duplicate papers across pages are dropped by id.

The query uses arXiv's own syntax, which is worth two minutes of your time because it is what makes this precise:

PrefixSearches
all:every field
ti:the title
au:the authors
abs:the abstract
cat:the category, like cs.CL or math.AG

Combine them with AND, OR and ANDNOT. ti:attention AND au:vaswani does exactly what it looks like. Sort by relevance, by submission date or by last revision, newest first in both date orders.

📥 What you give it

{
  "query": "all:large language models",
  "sortBy": "relevance",
  "maxItems": 50
}
FieldDefaultWhat it is
querybox starts at all:large language modelsThe arXiv query. Use the prefixes above; a bare phrase with no prefix is not what arXiv expects.
sortByrelevancerelevance, submittedDate or lastUpdatedDate. Both date orders are newest first.
maxItems50Papers to return, up to 30,000. A value above that is clamped, because arXiv will not page deeper on one query.
notionConnectornoneOptional. Write every paper into your own Notion as well as the dataset, which is a quick way to start a literature review.
notionParentIdnoneOptional. The Notion data source ID to write into.
proxyConfigurationoffOptional, and off by default because arXiv is public and a normal run does not need it.

There is no date-range field. To narrow by time, sort on submittedDate and cut the rows on publishedAt yourself.

📤 What you get back

A real row from a recent run, with the abstract cut short:

{
  "ok": true,
  "charged": true,
  "arxivId": "2609.13144",
  "title": "Type Diversity Enables Transformers to Generalise Compositionally",
  "abstract": "Compositional generalisation has been divided into lexical and structural generalisation. Previous work has found that structural generalisation is harder than lexical for Transformers...",
  "authors": ["Anssi Moisio", "Mathias Creutz", "Mikko Kurimo"],
  "primaryCategory": "cs.CL",
  "categories": ["cs.CL"],
  "publishedAt": "2026-09-11T00:00:00.000Z",
  "updatedAt": "2026-09-11T00:00:00.000Z",
  "doi": null,
  "absUrl": "https://arxiv.org/abs/2609.13144",
  "pdfUrl": "https://arxiv.org/pdf/2609.13144"
}
FieldWhat it is
arxivIdThe id with any trailing version marker removed, so 2609.13144v2 arrives as 2609.13144. Stable, and the dedupe key inside a run.
abstractThe whole abstract, with line breaks collapsed into single spaces.
publishedAt, updatedAtFirst submission and latest revision. On a paper never revised they are the same.
doinull on most papers. arXiv only has one when the author added it after journal publication.
absUrlThe abstract page, pointing at the specific version the search returned.
pdfUrlA direct PDF link. Nothing is downloaded for you.
primaryCategoryThe one category the author filed it under. categories holds all of them, deduplicated.

🧾 Reading the output

Two kinds of row can land in your dataset.

RowHow to spot itWhat it is
A paperok: true with charged: true and an arxivIda result
A diagnosticok: false and an errorCodenot a result, and not charged

charged: true is simply the marker every paper row carries. Use it, or ok, to tell results from diagnostics. Do not read it as a record of what your account was billed.

CodeWhat it means
BAD_INPUTThe query was empty. Give one in arXiv syntax.
NO_RESULTSThe query ran and matched no papers. Loosen it, or check the prefix spelling.
NOT_FOUNDarXiv answered 404 for the request.
RATE_LIMITEDarXiv asked for a slower pace. Re-run with a smaller maxItems.
SERVER_ERRORarXiv answered with a server error. Usually passes on its own.
BLOCKEDarXiv would not serve the search this time.
NETWORKarXiv could not be reached.

Rows arrive at the end. Everything is collected before anything is written, so a large run shows an empty dataset until it finishes. Keep a first run small while you get the query right.

The default table view has no ok column, so a diagnostic row renders as one blank line. Switch to All fields or export as JSON.

▶️ How to run it

1. Open arXiv Scraper and click Try for free. 2. Type a query into Search query, for example cat:cs.CL AND abs:retrieval. 3. Pick a Sort by. Use submittedDate when you care about what is new. 4. Set Max papers. Start around 20 while you tune the query, then click Start. 5. Download the dataset as JSON, CSV or Excel, or read it from the Apify API.

💰 How much does it cost?

$2.00 per 1,000 papers, which is $0.002 each. Flat on every Apify plan, no volume tiers.

You pay per paper row delivered. Duplicates across pages are dropped before they reach you, diagnostic rows are not charged, and a query that matches nothing costs you nothing.

💡 What people use it for

  • Watching a category like cat:cs.CL on a schedule and pulling everything new since yesterday.
  • Building a literature list on a topic, with every abstract already in the row to read or

summarise.

  • Following one author with au: to see what they have put up and when it was last revised.
  • Collecting PDF links for a set of papers to fetch and process in your own pipeline.

🚧 What it does not do

  • No full text. Abstracts and links. The PDF is yours to fetch from pdfUrl.
  • No citation counts, no references, no related-paper graph. arXiv does not publish them.
  • No peer-review status. A paper on arXiv may be a preprint, a submitted manuscript or a

published article, and nothing on the row tells you which.

  • doi is usually null. Most preprints have none, and arXiv does not go looking.
  • No date-range filter. Sort by date and trim the rows yourself.
  • Sort order is always newest or most relevant first. Ascending is not offered.
  • About 30,000 results per query, maximum. That is arXiv's depth limit, not a setting here.

Split a broad query into narrower ones rather than raising the number.

  • The run finishes even when every request failed, so read ok rather than the run status.

🧭 Which research scraper do you need?

If you wantUse
arXiv preprints with abstracts and PDF linksThis one
Published works with citation counts and open-access linksOpenAlex Scraper
Publisher metadata for a DOICrossref Scraper
Encyclopedia articles and their categoriesWikipedia Scraper
Patent filings by inventor or assigneeGoogle Patents Search Scraper

❓ Questions people ask

Do I need an arXiv account or an API key? No. arXiv's search is open to anyone.

Why did my search return nothing? Usually the syntax. A bare phrase with no prefix is not a valid arXiv query; start it with all: and try again.

Can I get the paper itself? Not from this actor. pdfUrl gives you a direct link to fetch.

Which date should I sort on? submittedDate for when work first appeared, lastUpdatedDate for what was revised recently. They differ a lot on papers that keep being updated.

Can I schedule it? Yes. A daily cat: query sorted by submittedDate, deduped on arxivId, keeps a category feed current.

Is scraping arXiv legal? arXiv publishes metadata for public reuse and asks callers to be polite about request rate, which this actor is. Individual papers carry their own licences, so check those before republishing text. Apify's write-up on scraping and the law is a good starting point, and we are not lawyers.

🆘 If something breaks

Open the Issues tab on the actor page. Send the run ID and the exact query string you used. The errorCode and hint on the diagnostic row usually name the problem on their own.