Request a tool
All toolsAutomationsGuidesMCP serverRequest a toolPlatformsCategories
OpenAlex Scholarly Works Scraper icon

OpenAlex Scholarly Works Scraper

Search 250M+ OpenAlex papers with no API key. Get titles, authors, venue, year, citations, DOI, OA links and full abstracts. $2.00 per 1,000 papers.

5 from 1 review on Apify 162 runs on Apify $0.002 per work ($2 / 1,000)
Run this in the cloudRun on Apify →

Developer & Research Tools

How it works

  1. 1
    Open it on Apify

    Hit Run on Apify — it opens the tool in the cloud, no install.

  2. 2
    Set the inputs

    Adjust query, fromDate, sort (sensible defaults are pre-filled).

  3. 3
    Click Run

    The tool runs on Apify’s cloud and collects the data for you.

  4. 4
    Export the results

    Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.

Pricing

$0.002 per work = $2 per 1,000

You are charged forWhenPrice
Work returnedCharged per scholarly work returned.$0.002

Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-06-13, and they are what you are actually charged.

Inputs

FieldWhat it doesType
queryKeywords to search OpenAlex works for (title, abstract and fulltext are searched), e.g. "machine learning", "crispr gene editing". Required.string
fromDateOptional. Only return works published on or after this date (YYYY-MM-DD, e.g. 2023-01-01). Adds a from_publication_date filter.string
sortHow to order results: Relevance (best match for the query), Citations (most-cited first), or Date (newest first).string
filterOptional advanced filter passed straight to the OpenAlex API filter param. Comma-separated key:value pairs, e.g. "type:article,is_oa:true,from_publication_date:2023-01-01". See the OpenAlex docs for available filter keys. Merged with From publication date.string
maxItemsMaximum number of works to return. Cursor pagination fetches 50 per page until this many unique works are collected.integer
notionConnectorOptional. Write each result as a page into your Notion when the run finishes. Authorize a Notion connector once in Settings → API & Integrations → MCP connectors, then pick it here. Leave empty to skip (default), results are always saved to the dataset regardless.string
notionParentIdOptional. The Notion data source ID of the database to write into (only used if a Notion connector is set). Leave empty to create the pages privately in your workspace instead.string

What you get

A structured dataset — each result includes fields like:

abstractauthorscitationsconceptsdoiinstitutionsisOpenAccessoaUrlopenalexIdpublicationDatetitletypeurlvenueyear

Export every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.

2 ready-to-run use cases

CRISPR Papers from OpenAlex, Newest First

Gene-editing work published since 2024, sorted by date, with authors, journal, DOI and open-access link. Abstracts are rebuilt from OpenAlex inverted index.

Most Cited Deep Learning Papers, Ranked by Citations

The foundational reading list pulled from OpenAlex with authors, venue, year and citation count. Abstracts come through on most works, not quite all of them.

Related tools in Developer & Research Tools

Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.

Crossref Scholarly Works Scraper iconDeveloper & Research Tools

Crossref Scholarly Works Scraper

Search 150M+ papers on Crossref: DOI, title, authors, journal, publisher, date, citations and abstract. No API key. $1.00 per 1,000 works.

2 use cases

Wikipedia Scraper iconDeveloper & Research Tools

Wikipedia Scraper

Search Wikipedia by keyword or by exact title. Get the intro text, the full article, thumbnails and categories. Any language. $1.00 per 1,000 pages.

3 use cases

Internet Archive Scraper iconDeveloper & Research Tools

Internet Archive Scraper

Search Internet Archive (archive.org) for books, audio, film and web items. Title, creator, year, downloads, subjects and URL. $2.00 per 1,000 items.

2 use cases

Hacker News Scraper iconDeveloper & Research Tools

Hacker News Scraper

Search HN stories, Show HN, Ask HN and comments, or pull the front page. Get points, author, comment counts and links. $1 per 1,000 items.

3 use cases

Research MCP Server — 15 Tools for AI Agents iconDeveloper & Research Tools

Research MCP Server — 15 Tools for AI Agents

One MCP endpoint gives your AI agent fifteen live research tools. Papers, code, news, SEC filings, packages and crypto prices.

4 use cases

DEV.to Scraper iconDeveloper & Research Tools

DEV.to Scraper

Scrape DEV.to articles by tag, author or sort. Fields: title, URL, tags, reactions, comments, reading time, cover image, full body. $2.00 per 1,000 articles.

Ready to run — no setup

See all Developer & Research Tools →

OpenAlex Scraper: search papers and get the abstract as readable text

Search OpenAlex by keyword and get back the paper, the authors, the institutions behind them, the year, the citation count, the DOI, whether it is open access and where to read it. The abstract comes back as ordinary prose.

That last part is the work. OpenAlex does not store abstracts as text. It stores a word-position index, and this actor rebuilds the sentences from it. Papers that have no index have no abstract, and the field is null rather than a guess.

InputSearch keywords
OutputOne row per work
Ceiling10,000 works per run
Account neededNone, and no API key
Price$2.00 per 1,000 works, flat on every plan

🔍 What OpenAlex Scraper does

Your keywords go against OpenAlex's own search, which reads titles, abstracts and full text where it has them. The run then pages with a cursor, fifty at a time, until it has the number of unique works you asked for or OpenAlex runs out.

You can order by relevance, by citation count or by date, and set a published-on-or-after floor. For anything more specific there is a raw filter field that goes straight through to OpenAlex, so type:article,is_oa:true works exactly as their documentation describes it.

Nothing is written to your dataset until the whole set has been collected, so you get either the full result set or a diagnostic row explaining why not.

📥 What you give it

{
  "query": "protein folding",
  "sort": "citations",
  "fromDate": "2020-01-01",
  "maxItems": 500
}
FieldDefaultWhat it is
querybox starts at machine learningThe keywords to search for. Required.
fromDatenoneYYYY-MM-DD. Only works published on or after this date.
sortrelevancerelevance, citations for most cited, or date for newest first.
filternoneAn OpenAlex filter string, passed through untouched. Comma-separated key:value pairs such as type:article,is_oa:true.
maxItems100How many works to return, up to 10,000.
notionConnectornoneOptional. Writes each work into your Notion once the run finishes.
notionParentIdnoneOptional. The Notion data source to write into.
proxyConfigurationoffOptional network settings. Off by default, and a normal run does not need it.

filter is handed to OpenAlex as you typed it, and a mistake in it does not come back as an error. OpenAlex answers a malformed filter with an empty result set, so the run finishes on NO_RESULTS and you are left thinking your query matched nothing. Check the filter keys against OpenAlex's documentation before blaming the query.

If your filter already contains from_publication_date, the fromDate field is ignored. Set the date in one place, not both.

📤 What you get back

A real row from a recent run, with the author and institution lists cut short here for length:

{
  "ok": true,
  "openalexId": "https://openalex.org/W2101234009",
  "doi": "https://doi.org/10.48550/arxiv.1201.0490",
  "title": "Scikit-learn: Machine Learning in Python",
  "authors": ["Fabián Pedregosa", "Gaël Varoquaux", "Alexandre Gramfort", "..."],
  "institutions": ["Commissariat à l'Énergie Atomique et aux Énergies Alternatives", "..."],
  "year": 2012,
  "publicationDate": "2012-01-02",
  "type": "article",
  "venue": "ORBi (University of Liège)",
  "citations": 63984,
  "concepts": ["Python (programming language)", "Documentation", "Computer science", "..."],
  "isOpenAccess": true,
  "oaUrl": "https://orbi.uliege.be/handle/2268/225787",
  "abstract": "Scikit-learn is a Python module integrating a wide range of ... machine learning algorithms for medium-scale supervised and unsupervised problems.",
  "url": "https://openalex.org/W2101234009"
}
FieldWhat it is
openalexIdOpenAlex's permanent id. Use it as your key, since not every work has a DOI.
abstractRebuilt from OpenAlex's word-position index. null where there is no index to rebuild from.
venueOpenAlex's primary location for the work. That is often the journal, but it can be a repository, as on the row above.
citationsOpenAlex's cited-by count, and 0 when OpenAlex has no figure at all.
institutionsThe affiliations OpenAlex resolved for the authors, deduplicated. Sometimes messy, because the matching is theirs.
conceptsOpenAlex's own topic tags, not keywords the authors chose.
isOpenAccess, oaUrlWhether a free copy is known, and where it is. oaUrl is null when none is known.
typearticle, preprint, book-chapter, dataset and so on, in OpenAlex's vocabulary.

🧾 Reading the output

Two kinds of row land in your dataset, and ok tells them apart.

RowHow to spot itCharged
A workok: true and an openalexIdyes
A diagnosticok: false and an errorCodeno

The overview table in the Apify console shows the work columns only, so a diagnostic row looks blank there. Switch to the JSON or All fields view to read it.

CodeWhat it means
BAD_INPUTNo query was given.
NO_RESULTSThe request worked and nothing came back. A malformed filter also lands here.
RATE_LIMITEDOpenAlex asked for a slower pace than the run could keep. Try a smaller run.
SERVER_ERROROpenAlex answered 5xx. Usually passes.
BLOCKEDOpenAlex refused the request. Re-run it.
NETWORKOpenAlex was unreachable. Re-run it.

▶️ How to run it

1. Open OpenAlex Scraper and click Try for free. 2. Type your keywords into Search query. 3. Set Max works. Start around 50 to see the row shape. 4. Pick a Sort by, add a From publication date if you want, then click Start. 5. Download the dataset as JSON, CSV or Excel, or read it from the Apify API.

💰 How much does it cost?

$2.00 per 1,000 works. Flat on every Apify plan, no volume tiers.

You pay per work delivered. Works that appear twice across pages are dropped before they are counted, diagnostic rows are not charged, and a search that matches nothing is not charged.

💡 What people use it for

  • Pulling a field's most-cited work with the abstracts attached, in one file, for a reading list.
  • Feeding titles and abstracts to a model that has to summarise research it was not trained on.
  • Finding which institutions keep appearing on a topic, using the institutions field.
  • Filtering to open access with is_oa:true so every row has somewhere to read it for free.
  • Watching a topic on a schedule with sort on date, so each run surfaces what is new.

🚧 What it does not do

  • Keyword search over works only. No author, institution or funder lookup by entity.
  • No full text. You get the abstract and a link, not the paper.
  • Abstracts are missing where OpenAlex has no index for them, and that is common on older and

paywalled records.

  • citations cannot tell you zero from unknown. Both read 0.
  • A bad filter looks like an empty search, not like an error.
  • No end date field. Use the raw filter if you need an upper bound.
  • Institutions and concepts are OpenAlex's matching, and it is not perfect. Treat them as hints,

not as ground truth.

  • All or nothing. A failure partway through discards what had been collected, so you get a

diagnostic row rather than a partial set.

🧭 Which research scraper do you need?

If you wantUse
Works with abstracts, citations and institutionsThis one
Anything with a registered DOI, straight from CrossrefCrossref Scraper
Preprints with a PDF linkarXiv Scraper
Patents rather than papersGoogle Patents Search Scraper
Books, audio, film and archived web pagesInternet Archive Scraper

❓ Questions people ask

Do I need an API key? No. OpenAlex is open and the actor talks to it directly.

Why is the abstract null on some rows? OpenAlex only keeps a word-position index for a work if the publisher made one available. No index, no abstract to rebuild.

Why is venue a repository instead of a journal? OpenAlex picks one primary location per work, and for a paper with a preprint or a repository copy it sometimes picks that. The DOI still points at the published version.

Can I search by author or institution? Not as a mode. You can approximate it through the raw filter field using OpenAlex's own filter keys.

Why did my filtered search return nothing? Most often a filter key or value that OpenAlex does not recognise. It answers with an empty set rather than an error, so check the filter first.

Is this legal? OpenAlex publishes this data openly under a public domain dedication and invites this kind of use. Apify's write-up on scraping and the law is a good starting point, and we are not lawyers.

🆘 If something breaks

Open the Issues tab on the actor page. Send the run ID, the query and any filter you used. The errorCode on the diagnostic row usually names the problem on its own.