Crossref Scholarly Works Scraper
Search 150M+ papers on Crossref: DOI, title, authors, journal, publisher, date, citations and abstract. No API key. $1.00 per 1,000 works.
How it works
- 1Open it on Apify
Hit Run on Apify — it opens the tool in the cloud, no install.
- 2Set the inputs
Adjust
query,fromDate,filterType(sensible defaults are pre-filled). - 3Click Run
The tool runs on Apify’s cloud and collects the data for you.
- 4Export the results
Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.
Pricing
$0.001 per work = $1 per 1,000
| You are charged for | When | Price |
|---|---|---|
| Work returned | Charged per scholarly work returned. | $0.001 |
Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-06-12, and they are what you are actually charged.
Inputs
| Field | What it does | Type |
|---|---|---|
query | Keywords to search Crossref for across titles, authors, abstracts, and metadata (e.g. "deep learning", "CRISPR gene editing", "climate change adaptation"). It can be left empty only when you give a journal ISSN below. | string |
fromDate | Only return works published on or after this date, in YYYY-MM-DD format (e.g. 2020-01-01). Leave empty for no date floor. | string |
filterType | Only return works of this Crossref type. Leave empty for all types. "journal-article" is the most common for research papers. | string |
issn | Only works from these journals, by ISSN, such as 2169-3536. Give several and you get works from any of them. With an ISSN set you can leave the search query empty and get the journal's whole list. | array |
sort | How to order results. "Relevance" matches the query best; "Most cited" surfaces influential papers; "Newest first" sorts by publication date descending and stops at 10,000 works. | string |
maxItems | Maximum number of scholarly works to return. Uses deep cursor pagination to fetch beyond 100 reliably. | integer |
notionConnector | Optional. Write each result as a page into your Notion when the run finishes. Authorize a Notion connector once in Settings → API & Integrations → MCP connectors, then pick it here. Leave empty to skip (default), results are always saved to the dataset regardless. | string |
notionParentId | Optional. The Notion data source ID of the database to write into (only used if a Notion connector is set). Leave empty to create the pages privately in your workspace instead. | string |
What you get
A structured dataset — each result includes fields like:
abstractauthorscitationsdoiissnjournalpublishedDatepublishersubjectstitletypeurlExport every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.
2 ready-to-run use cases
Find Most Cited Papers on a Topic, Ranked by Citations
Crossref works sorted by citation count, with DOI, title, journal and publisher. Authors are on most rows but not all, and abstracts on very few.
Crossref Metadata Search: Every Work Type for a Topic
Articles, books, datasets and preprints for one query, each with a DOI, title, publisher and citation count. Author lists come through where they exist.
Related tools in Developer & Research Tools
Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.
Wikipedia Scraper
Search Wikipedia by keyword or by exact title. Get the intro text, the full article, thumbnails and categories. Any language. $1.00 per 1,000 pages.
Internet Archive Scraper
Search Internet Archive (archive.org) for books, audio, film and web items. Title, creator, year, downloads, subjects and URL. $2.00 per 1,000 items.
Hacker News Scraper
Search HN stories, Show HN, Ask HN and comments, or pull the front page. Get points, author, comment counts and links. $1 per 1,000 items.
Research MCP Server — 15 Tools for AI Agents
One MCP endpoint gives your AI agent fifteen live research tools. Papers, code, news, SEC filings, packages and crypto prices.
DEV.to Scraper
Scrape DEV.to articles by tag, author or sort. Fields: title, URL, tags, reactions, comments, reading time, cover image, full body. $2.00 per 1,000 articles.
Wikidata Scraper
Search Wikidata for people, companies, places, books or films. Get the ID, the name, other names, a description and the Wikipedia link. $0.20 per 1,000.
Where this tool sits
- Categories
- Developer & Research Tools
- Platforms
- Academic & Research
Crossref Scraper: search scholarly works by keyword and get the DOI back
Type a research phrase and get back the works Crossref has a DOI for: journal articles, preprints, book chapters, datasets, dissertations. Every row carries the DOI, the title, the authors, the journal, the publisher, the publication date and the citation count.
The honest part first. Crossref holds what publishers deposited with it, and plenty of them deposit the bare minimum. Abstracts, subject terms and ISSNs are missing on a lot of older records, and the citation count reads 0 both when a work genuinely has none and when nobody has told Crossref.
| Input | Search keywords, journal ISSNs, or both |
| Output | One row per work |
| Ceiling | 20,000 works per run, 10,000 when sorted newest first |
| Account needed | None, and no API key |
| Price | $1.00 per 1,000 works, flat on every plan |
🔍 What Crossref Scraper does
It runs your keywords against Crossref's search, which reads across titles, authors, abstracts and the rest of the deposited metadata, then pages through the matches with a cursor so a large request keeps going past the first hundred.
You can narrow it to one work type or to particular journals by ISSN, set a published-on-or-after floor, and choose whether the results come back by relevance, by citation count or newest first. Works with no DOI are skipped, and a DOI that appears twice across pages is only delivered once.
Nothing is written to your dataset until the whole set has been collected, so you get either the full result set you asked for or a diagnostic row explaining why not.
📥 What you give it
{
"query": "CRISPR gene editing",
"filterType": "journal-article",
"fromDate": "2020-01-01",
"sort": "is-referenced-by-count",
"maxItems": 500
}
| Field | Default | What it is |
|---|---|---|
query | box starts at deep learning | The keywords to search for. Required unless you give an issn. |
filterType | all types | One Crossref type: journal-article, proceedings-article, book-chapter, book, posted-content for preprints, dataset, report, dissertation or monograph. |
issn | none | Journal ISSNs, such as 2169-3536, up to 50. You get works from any of them. With an ISSN set, query can be empty to get the journal's whole list. |
fromDate | none | YYYY-MM-DD. Only works published on or after this date. There is no matching end date. |
sort | relevance | relevance, is-referenced-by-count for most cited, or published for newest first, which stops at 10,000 works. |
maxItems | 100 | How many works to return, up to 20,000. |
notionConnector | none | Optional. Writes each work into your Notion once the run finishes. |
notionParentId | none | Optional. The Notion data source to write into. |
proxyConfiguration | off | Optional network settings. Off by default, and a normal run does not need it. |
fromDate has to be a real date shaped YYYY-MM-DD, and an ISSN has to be 8 characters like 0950-382X. Anything else stops the run with a BAD_INPUT row before a single search is made.
📤 What you get back
A real row from a recent run:
{
"ok": true,
"doi": "10.1067/mva.1993.50616",
"title": "Light reflection rheography: A simple noninvasive screening test for deep vein thrombosis",
"authors": ["Thomas W. Wakefield", "Subodh W. Arora", "David J. K. Lam", "..."],
"journal": "Journal of Vascular Surgery",
"publisher": "Elsevier BV",
"type": "journal-article",
"publishedDate": "1993-11",
"citations": 0,
"subjects": [],
"issn": ["0741-5214"],
"abstract": null,
"url": "https://doi.org/10.1067/mva.1993.50616"
}
The author list on that row is cut short here for length. A real row carries all six names.
| Field | What it is |
|---|---|
doi | Always present. Works without one are dropped before they reach you, so this is safe as a key. |
authors | Names as Given Family. An organisation shows up under its own name when there is no person. |
publishedDate | YYYY-MM-DD where the publisher deposited a full date, otherwise YYYY-MM or just YYYY. Older records are usually the short forms. |
citations | Crossref's referenced-by count, and 0 when Crossref has no figure at all. |
subjects, issn | Arrays, frequently empty. Both depend on what the publisher deposited. |
abstract | Plain text with the publisher's markup stripped out, or null. Missing far more often than present. |
url | Crossref's own link, falling back to https://doi.org/<DOI>. |
🧾 Reading the output
Two kinds of row land in your dataset, and ok tells them apart.
| Row | How to spot it | Charged |
|---|---|---|
| A work | ok: true and a doi | yes |
| A diagnostic | ok: false and an errorCode | no |
The overview table in the Apify console shows the work columns only, so a diagnostic row looks blank there. Switch to the JSON or All fields view to read it.
| Code | What it means |
|---|---|
BAD_INPUT | No query and no ISSN, a fromDate that is not a real YYYY-MM-DD date, an ISSN that is not one, or a search Crossref itself refused. The row says which. |
NO_RESULTS | The search worked and nothing matched. totalResults on the row shows what Crossref counted. |
RATE_LIMITED | Crossref asked for a slower pace than the run could keep. Try a smaller run. |
SERVER_ERROR | Crossref answered 5xx. Usually passes. |
BLOCKED | Crossref refused the request. Re-run it. |
NETWORK | Crossref was unreachable. Re-run it. |
▶️ How to run it
1. Open Crossref Scraper and click Try for free. 2. Type your keywords into Search query. 3. Set Max works. Start around 50 to see the row shape. 4. Pick a Work type filter and a Sort order if the defaults do not suit you, then click Start. 5. Download the dataset as JSON, CSV or Excel, or read it from the Apify API.
💰 How much does it cost?
$1.00 per 1,000 works. Flat on every Apify plan, no volume tiers.
You pay per work delivered. Duplicate DOIs and works without a DOI are dropped before they are counted, diagnostic rows are not charged, and a search that matches nothing is not charged.
💡 What people use it for
- Building a DOI list for a literature review, then fetching the papers themselves elsewhere.
- Sorting a field by citation count to see which work everything else is built on.
- Watching a topic on a schedule with
sortonpublished, so each run surfaces what is new. - Checking which publisher holds a set of papers before writing to them about access.
- Feeding titles and abstracts into a model that has to summarise a field it was not trained on.
🚧 What it does not do
- No lookup by DOI or by author. Keywords and journal ISSNs are the ways in.
- No end date.
fromDatesets a floor and nothing sets a ceiling. - Newest first stops at 10,000 works, and a few records carry publication dates years ahead,
which sort to the top. In one test the first rows were dated 2115 and 2106. Deep in that order, works sharing a date can repeat across pages; repeats are dropped and not charged, so a full run ends a little short (9,872 of 10,000 in that test).
- Abstracts are mostly missing. Crossref only has one if the publisher deposited it, and many
never do.
citationscannot tell you zero from unknown. Both read0.- No full text, and no PDF. You get the DOI and the link, not the paper.
- Paging stops on a short page. If Crossref returns fewer than 100 items in one page, the run
treats that as the end even when its own total says otherwise.
- All or nothing. A failure partway through discards what had been collected, so you get a
diagnostic row rather than a partial set.
- Crossref is publisher-deposited metadata, not a curated index. Quality varies by publisher and
by decade.
🧭 Which research scraper do you need?
| If you want | Use |
|---|---|
| Works with a registered DOI, from Crossref | This one |
| Open citation graphs, institutions and open-access status | OpenAlex Scraper |
| Preprints with a PDF link | arXiv Scraper |
| Books, audio, film and archived web pages | Internet Archive Scraper |
| Book metadata and ISBNs | Books Scraper |
❓ Questions people ask
Do I need an API key? No. Crossref publishes this openly and the actor talks to it directly.
Why are so many abstracts null? Crossref stores what the publisher sent. Most publishers never deposit an abstract, so the field is empty far more often than it is filled.
Can I look a paper up by its DOI? Not here. This searches by keyword and by journal ISSN.
Can I list what one journal published? Yes. Put its ISSN in issn and leave the query empty. A big journal holds more than one run delivers, so sort newest first or by citations to get the end you care about.
Why did I get fewer works than maxItems? Either Crossref has fewer matches, or a page came back short and paging stopped there. You are charged for what arrived.
Can I get the PDF? No. You get the DOI and the publisher's link, and access depends on the publisher.
Is this legal? Crossref publishes this metadata openly for exactly this kind of use. Apify's write-up on scraping and the law is a good starting point, and we are not lawyers.
🆘 If something breaks
Open the Issues tab on the actor page. Send the run ID and the query you used. The errorCode on the diagnostic row usually names the problem on its own.