OpenAlex Scholarly Works Scraper
Search 250M+ OpenAlex papers with no API key. Get titles, authors, venue, year, citations, DOI, OA links and full abstracts. $2.00 per 1,000 papers.
How it works
- 1Open it on Apify
Hit Run on Apify — it opens the tool in the cloud, no install.
- 2Set the inputs
Adjust
query,fromDate,sort(sensible defaults are pre-filled). - 3Click Run
The tool runs on Apify’s cloud and collects the data for you.
- 4Export the results
Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.
Pricing
$0.002 per work = $2 per 1,000
| You are charged for | When | Price |
|---|---|---|
| Work returned | Charged per scholarly work returned. | $0.002 |
Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-06-13, and they are what you are actually charged.
Inputs
| Field | What it does | Type |
|---|---|---|
query | Keywords to search OpenAlex works for (title, abstract and fulltext are searched), e.g. "machine learning", "crispr gene editing". Required. | string |
fromDate | Optional. Only return works published on or after this date (YYYY-MM-DD, e.g. 2023-01-01). Adds a from_publication_date filter. | string |
sort | How to order results: Relevance (best match for the query), Citations (most-cited first), or Date (newest first). | string |
filter | Optional advanced filter passed straight to the OpenAlex API filter param. Comma-separated key:value pairs, e.g. "type:article,is_oa:true,from_publication_date:2023-01-01". See the OpenAlex docs for available filter keys. Merged with From publication date. | string |
maxItems | Maximum number of works to return. Cursor pagination fetches 50 per page until this many unique works are collected. | integer |
notionConnector | Optional. Write each result as a page into your Notion when the run finishes. Authorize a Notion connector once in Settings → API & Integrations → MCP connectors, then pick it here. Leave empty to skip (default), results are always saved to the dataset regardless. | string |
notionParentId | Optional. The Notion data source ID of the database to write into (only used if a Notion connector is set). Leave empty to create the pages privately in your workspace instead. | string |
What you get
A structured dataset — each result includes fields like:
abstractauthorscitationsconceptsdoiinstitutionsisOpenAccessoaUrlopenalexIdpublicationDatetitletypeurlvenueyearExport every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.
2 ready-to-run use cases
CRISPR Papers from OpenAlex, Newest First
Gene-editing work published since 2024, sorted by date, with authors, journal, DOI and open-access link. Abstracts are rebuilt from OpenAlex inverted index.
Most Cited Deep Learning Papers, Ranked by Citations
The foundational reading list pulled from OpenAlex with authors, venue, year and citation count. Abstracts come through on most works, not quite all of them.
Related tools in Developer & Research Tools
Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.
Crossref Scholarly Works Scraper
Search 150M+ papers on Crossref: DOI, title, authors, journal, publisher, date, citations and abstract. No API key. $1.00 per 1,000 works.
Wikipedia Scraper
Search Wikipedia by keyword or by exact title. Get the intro text, the full article, thumbnails and categories. Any language. $1.00 per 1,000 pages.
Internet Archive Scraper
Search Internet Archive (archive.org) for books, audio, film and web items. Title, creator, year, downloads, subjects and URL. $2.00 per 1,000 items.
Hacker News Scraper
Search HN stories, Show HN, Ask HN and comments, or pull the front page. Get points, author, comment counts and links. $1 per 1,000 items.
Research MCP Server — 15 Tools for AI Agents
One MCP endpoint gives your AI agent fifteen live research tools. Papers, code, news, SEC filings, packages and crypto prices.
DEV.to Scraper
Scrape DEV.to articles by tag, author or sort. Fields: title, URL, tags, reactions, comments, reading time, cover image, full body. $2.00 per 1,000 articles.
Where this tool sits
- Categories
- Developer & Research Tools
- Platforms
- Academic & Research
OpenAlex Scraper: search papers and get the abstract as readable text
Search OpenAlex by keyword and get back the paper, the authors, the institutions behind them, the year, the citation count, the DOI, whether it is open access and where to read it. The abstract comes back as ordinary prose.
That last part is the work. OpenAlex does not store abstracts as text. It stores a word-position index, and this actor rebuilds the sentences from it. Papers that have no index have no abstract, and the field is null rather than a guess.
| Input | Search keywords |
| Output | One row per work |
| Ceiling | 10,000 works per run |
| Account needed | None, and no API key |
| Price | $2.00 per 1,000 works, flat on every plan |
🔍 What OpenAlex Scraper does
Your keywords go against OpenAlex's own search, which reads titles, abstracts and full text where it has them. The run then pages with a cursor, fifty at a time, until it has the number of unique works you asked for or OpenAlex runs out.
You can order by relevance, by citation count or by date, and set a published-on-or-after floor. For anything more specific there is a raw filter field that goes straight through to OpenAlex, so type:article,is_oa:true works exactly as their documentation describes it.
Nothing is written to your dataset until the whole set has been collected, so you get either the full result set or a diagnostic row explaining why not.
📥 What you give it
{
"query": "protein folding",
"sort": "citations",
"fromDate": "2020-01-01",
"maxItems": 500
}
| Field | Default | What it is |
|---|---|---|
query | box starts at machine learning | The keywords to search for. Required. |
fromDate | none | YYYY-MM-DD. Only works published on or after this date. |
sort | relevance | relevance, citations for most cited, or date for newest first. |
filter | none | An OpenAlex filter string, passed through untouched. Comma-separated key:value pairs such as type:article,is_oa:true. |
maxItems | 100 | How many works to return, up to 10,000. |
notionConnector | none | Optional. Writes each work into your Notion once the run finishes. |
notionParentId | none | Optional. The Notion data source to write into. |
proxyConfiguration | off | Optional network settings. Off by default, and a normal run does not need it. |
filter is handed to OpenAlex as you typed it, and a mistake in it does not come back as an error. OpenAlex answers a malformed filter with an empty result set, so the run finishes on NO_RESULTS and you are left thinking your query matched nothing. Check the filter keys against OpenAlex's documentation before blaming the query.
If your filter already contains from_publication_date, the fromDate field is ignored. Set the date in one place, not both.
📤 What you get back
A real row from a recent run, with the author and institution lists cut short here for length:
{
"ok": true,
"openalexId": "https://openalex.org/W2101234009",
"doi": "https://doi.org/10.48550/arxiv.1201.0490",
"title": "Scikit-learn: Machine Learning in Python",
"authors": ["Fabián Pedregosa", "Gaël Varoquaux", "Alexandre Gramfort", "..."],
"institutions": ["Commissariat à l'Énergie Atomique et aux Énergies Alternatives", "..."],
"year": 2012,
"publicationDate": "2012-01-02",
"type": "article",
"venue": "ORBi (University of Liège)",
"citations": 63984,
"concepts": ["Python (programming language)", "Documentation", "Computer science", "..."],
"isOpenAccess": true,
"oaUrl": "https://orbi.uliege.be/handle/2268/225787",
"abstract": "Scikit-learn is a Python module integrating a wide range of ... machine learning algorithms for medium-scale supervised and unsupervised problems.",
"url": "https://openalex.org/W2101234009"
}
| Field | What it is |
|---|---|
openalexId | OpenAlex's permanent id. Use it as your key, since not every work has a DOI. |
abstract | Rebuilt from OpenAlex's word-position index. null where there is no index to rebuild from. |
venue | OpenAlex's primary location for the work. That is often the journal, but it can be a repository, as on the row above. |
citations | OpenAlex's cited-by count, and 0 when OpenAlex has no figure at all. |
institutions | The affiliations OpenAlex resolved for the authors, deduplicated. Sometimes messy, because the matching is theirs. |
concepts | OpenAlex's own topic tags, not keywords the authors chose. |
isOpenAccess, oaUrl | Whether a free copy is known, and where it is. oaUrl is null when none is known. |
type | article, preprint, book-chapter, dataset and so on, in OpenAlex's vocabulary. |
🧾 Reading the output
Two kinds of row land in your dataset, and ok tells them apart.
| Row | How to spot it | Charged |
|---|---|---|
| A work | ok: true and an openalexId | yes |
| A diagnostic | ok: false and an errorCode | no |
The overview table in the Apify console shows the work columns only, so a diagnostic row looks blank there. Switch to the JSON or All fields view to read it.
| Code | What it means |
|---|---|
BAD_INPUT | No query was given. |
NO_RESULTS | The request worked and nothing came back. A malformed filter also lands here. |
RATE_LIMITED | OpenAlex asked for a slower pace than the run could keep. Try a smaller run. |
SERVER_ERROR | OpenAlex answered 5xx. Usually passes. |
BLOCKED | OpenAlex refused the request. Re-run it. |
NETWORK | OpenAlex was unreachable. Re-run it. |
▶️ How to run it
1. Open OpenAlex Scraper and click Try for free. 2. Type your keywords into Search query. 3. Set Max works. Start around 50 to see the row shape. 4. Pick a Sort by, add a From publication date if you want, then click Start. 5. Download the dataset as JSON, CSV or Excel, or read it from the Apify API.
💰 How much does it cost?
$2.00 per 1,000 works. Flat on every Apify plan, no volume tiers.
You pay per work delivered. Works that appear twice across pages are dropped before they are counted, diagnostic rows are not charged, and a search that matches nothing is not charged.
💡 What people use it for
- Pulling a field's most-cited work with the abstracts attached, in one file, for a reading list.
- Feeding titles and abstracts to a model that has to summarise research it was not trained on.
- Finding which institutions keep appearing on a topic, using the
institutionsfield. - Filtering to open access with
is_oa:trueso every row has somewhere to read it for free. - Watching a topic on a schedule with
sortondate, so each run surfaces what is new.
🚧 What it does not do
- Keyword search over works only. No author, institution or funder lookup by entity.
- No full text. You get the abstract and a link, not the paper.
- Abstracts are missing where OpenAlex has no index for them, and that is common on older and
paywalled records.
citationscannot tell you zero from unknown. Both read0.- A bad
filterlooks like an empty search, not like an error. - No end date field. Use the raw filter if you need an upper bound.
- Institutions and concepts are OpenAlex's matching, and it is not perfect. Treat them as hints,
not as ground truth.
- All or nothing. A failure partway through discards what had been collected, so you get a
diagnostic row rather than a partial set.
🧭 Which research scraper do you need?
| If you want | Use |
|---|---|
| Works with abstracts, citations and institutions | This one |
| Anything with a registered DOI, straight from Crossref | Crossref Scraper |
| Preprints with a PDF link | arXiv Scraper |
| Patents rather than papers | Google Patents Search Scraper |
| Books, audio, film and archived web pages | Internet Archive Scraper |
❓ Questions people ask
Do I need an API key? No. OpenAlex is open and the actor talks to it directly.
Why is the abstract null on some rows? OpenAlex only keeps a word-position index for a work if the publisher made one available. No index, no abstract to rebuild.
Why is venue a repository instead of a journal? OpenAlex picks one primary location per work, and for a paper with a preprint or a repository copy it sometimes picks that. The DOI still points at the published version.
Can I search by author or institution? Not as a mode. You can approximate it through the raw filter field using OpenAlex's own filter keys.
Why did my filtered search return nothing? Most often a filter key or value that OpenAlex does not recognise. It answers with an empty set rather than an error, so check the filter first.
Is this legal? OpenAlex publishes this data openly under a public domain dedication and invites this kind of use. Apify's write-up on scraping and the law is a good starting point, and we are not lawyers.
🆘 If something breaks
Open the Issues tab on the actor page. Send the run ID, the query and any filter you used. The errorCode on the diagnostic row usually names the problem on its own.