Package Registry Scraper (npm + PyPI)
Get npm and PyPI package metadata as JSON. Version, license, author, repo, keywords and npm monthly downloads. $2 per 1,000 packages.
How it works
- 1Open it on Apify
Hit Run on Apify — it opens the tool in the cloud, no install.
- 2Set the inputs
Adjust
registry,searchQuery,packageNames(sensible defaults are pre-filled). - 3Click Run
The tool runs on Apify’s cloud and collects the data for you.
- 4Export the results
Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.
Pricing
$0.002 per package = $2 per 1,000
| You are charged for | When | Price |
|---|---|---|
| Package returned | Charged per package returned. | $0.002 |
Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-06-13, and they are what you are actually charged.
Inputs
| Field | What it does | Type |
|---|---|---|
registry | Which package registry to use. npm supports both keyword search and exact-name lookup; PyPI supports exact-name lookup only (it has no clean public search API). | string |
searchQuery | Keywords to search the npm registry for (e.g. "react state management"). npm only, ignored for PyPI. Leave empty if you are looking up exact package names instead. | string |
packageNames | Exact package names to look up directly. Works for BOTH registries. For PyPI this is the only supported mode (e.g. ["requests", "fastapi"]). For npm, scoped names like "@types/node" are supported. | array |
maxItems | Maximum number of packages to return from an npm search query. Only applies to npm search; ignored for exact-name lookups. | integer |
includeDownloads | Fetch last-month download counts for each npm package via the npm downloads API. npm only. PyPI does not expose a public download-count endpoint. Adds one request per package. Note: monthlyDownloads is null when this is off, for PyPI packages, or if the downloads API call fails for a given package (a warning is logged in that case). | boolean |
notionConnector | Optional. Write each package as a page into your Notion when the run finishes. Authorize a Notion connector once in Settings → API & Integrations → MCP connectors, then pick it here. Leave empty to skip (default), results are always saved to the dataset regardless. | string |
notionParentId | Optional. The Notion data source ID of the database to write into (only used if a Notion connector is set). Leave empty to create the pages privately in your workspace instead. | string |
What you get
A structured dataset — each result includes fields like:
authordescriptionhomepagekeywordslicensemonthlyDownloadsnameregistryrepositoryscoreurlversionExport every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.
2 ready-to-run use cases
npm Package Search with Monthly Download Counts
npm package search by keyword: name, version, description, repo link and monthly downloads per hit. Search hits carry no license field. Name lookups do.
npm License Checker for a List of Dependencies
Feed the names from your package.json and get the license and source repo back for each npm package. A compliance pass without installing anything.
Ready-made automation using this tool
A finished workflow you can import into n8n and run — this tool does the data-gathering inside it. Free, and yours to change.
Know when a package you depend on ships a new version
Watches the npm or PyPI packages you list and tells you when one releases — but only the kinds of release you asked for. Set it to major versions only and you hear about breaking changes without twenty patch notifications a week.
Related tools in Developer & Research Tools
Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.
arXiv Scraper
Search arXiv papers by title, author, abstract or category. Get full abstracts, authors, categories, DOI, dates and PDF links. $2 per 1,000 papers.
OpenAlex Scholarly Works Scraper
Search 250M+ OpenAlex papers with no API key. Get titles, authors, venue, year, citations, DOI, OA links and full abstracts. $2.00 per 1,000 papers.
Crossref Scholarly Works Scraper
Search 150M+ papers on Crossref: DOI, title, authors, journal, publisher, date, citations and abstract. No API key. $1.00 per 1,000 works.
Wikipedia Scraper
Search Wikipedia by keyword or by exact title. Get the intro text, the full article, thumbnails and categories. Any language. $1.00 per 1,000 pages.
Internet Archive Scraper
Search Internet Archive (archive.org) for books, audio, film and web items. Title, creator, year, downloads, subjects and URL. $2.00 per 1,000 items.
Hacker News Scraper
Search HN stories, Show HN, Ask HN and comments, or pull the front page. Get points, author, comment counts and links. $1 per 1,000 items.
Where this tool sits
- Categories
- Developer & Research Tools
- Platforms
- Academic & Research
Package Registry Scraper: npm and PyPI metadata in one row shape
Look up packages by exact name on npm or PyPI, or search npm by keywords, and get one row each: name, latest version, description, author, homepage, repository, licence, keywords and the project page URL. npm rows can also carry last month's download count.
The two registries are not symmetrical, and it matters before you plan a run. npm does keyword search and exact names. PyPI does exact names only, because it publishes no clean search API for anyone to call.
| Input | Exact package names, or npm keywords |
| Output | One row per package |
| Ceiling | 1,000 packages per npm search. Name lookups are not capped |
| Account needed | None, and no registry key |
| Price | $2.00 per 1,000 packages, flat on every plan |
🔍 What Package Registry Scraper does
Two ways in, and you can use both at once on npm.
Exact names. Put a list in packageNames and each one is looked up directly. Scoped npm names like @types/node work. A name that does not exist gets its own NOT_FOUND row, uncharged, and the run carries on with the rest.
npm keyword search. Put words in searchQuery and the registry's own search decides what matches, ordered by its relevance score, up to your maxItems.
There is a real difference between the two that the field list hides: a search row has no licence and usually no repository link. npm's search index does not carry them. If you are doing a licence audit, feed the names in through packageNames and look them up properly.
includeDownloads is on by default and adds last month's download count to npm rows. PyPI has no public endpoint for that, so PyPI rows never carry it.
📥 What you give it
{
"registry": "npm",
"packageNames": ["react", "@types/node"],
"includeDownloads": true
}
| Field | Default | What it is |
|---|---|---|
registry | npm | npm or pypi. |
packageNames | none | Exact names. Works on both registries, and is the only mode PyPI supports. |
searchQuery | none | Keywords, npm only. Ignored on PyPI unless you sent no names, in which case you get BAD_INPUT. The Console shows react state management as an example, but that is a prefill, so an API call has to send its own. |
maxItems | 50 | 1 to 1,000. Caps the npm search only. It does nothing to a list of names. |
includeDownloads | on | Last month's downloads for npm packages. Costs one extra lookup per package. |
notionConnector | none | Optional. Writes every delivered package into your Notion. Authorise a connector once under Settings, API and Integrations, MCP connectors, then pick it here. |
notionParentId | none | Optional. The Notion data source id to write into. Leave it empty and the pages are created privately in your workspace. |
proxyConfiguration | off | Optional network setting. Off is right for a normal run. |
packageNames has no ceiling. Paste a 4,000-line requirements.txt and you get up to 4,000 charged rows, because maxItems is not watching that path. Trim the list to what you need.
📤 What you get back
A real row from a recent npm search:
{
"ok": true,
"registry": "npm",
"name": "unstated-next",
"version": "1.1.0",
"description": "200 bytes to never think about React state management libraries ever again",
"author": "thejameskyle",
"homepage": null,
"repository": null,
"license": null,
"keywords": [],
"score": 364.87637,
"url": "https://www.npmjs.com/package/unstated-next",
"monthlyDownloads": 330524
}
That row shows the search-path gap plainly: license, repository, homepage and keywords are all empty, because the search index does not carry them. Look the same package up by name and they are filled in.
| Field | What it is |
|---|---|
version | The latest published version. npm's latest tag, PyPI's current release. |
license | The declared licence string. null on every npm search row. |
repository | Normalised into a clickable https URL rather than a git address. |
score | npm's own search relevance. Present on search rows, null on npm name lookups, and the key is absent from PyPI rows entirely. |
monthlyDownloads | npm only. null if the count could not be read, and the key is absent when includeDownloads is off or the row is from PyPI. |
author | On a search row this can be the publishing account's handle rather than the person named in the package file. |
🧾 Reading the output
Packages carry ok: true. Anything with ok: false carries an errorCode and is not charged. One run can contain both, which is normal when a name list has a typo in it.
| Code | What it means |
|---|---|
BAD_INPUT | Neither names nor a query, or a PyPI run with a searchQuery and no names. |
NOT_FOUND | That one name is not in the registry. Check the spelling, the scope, and whether it was unpublished. |
NO_RESULTS | Nothing at all came back. An npm search that matched no packages ends here. |
NETWORK | The registry was unreachable or answered with something unusable. Re-run it. |
The Console's Overview table hides license and repository, which are the two columns a licence audit actually needs. Download the dataset as JSON, CSV or Excel, or switch the table to all fields.
▶️ How to run it
1. Open Package Registry Scraper and click Try for free. 2. Pick a Registry. 3. Paste names into Package names, one per line. For an npm search, type into Search query instead. 4. Leave Include monthly downloads on unless you do not want it, then click Start. 5. Download the dataset as JSON, CSV, Excel or XML.
💰 How much does it cost?
$2.00 per 1,000 packages. Flat on every Apify plan, no volume tiers, and the same on both registries.
You pay per package row. The same name twice in one run is charged once. NOT_FOUND names, a search that matched nothing and every diagnostic row are not charged. Turning downloads on does not change the price.
💡 What people use it for
- Auditing the licences behind a
package.jsonor arequirements.txtby feeding the names in and
reading the license column.
- Comparing two libraries side by side on version, downloads and last release before choosing one.
- Watching a dependency on a schedule and firing an alert when
versionmoves. - Sizing a niche: search npm for the keywords, then rank by
monthlyDownloadsto see what people
actually install.
🚧 What it does not do
- No PyPI search. Exact names only on that side.
- No PyPI download counts. There is no public endpoint to read them from.
- No licence on npm search rows. Use
packageNameswhen the licence is the point. - No dependency trees and no version history. One row is the current state of one package.
- No download trend.
monthlyDownloadsis last month's total, a single number, not a series. - No GitHub stars or issues.
repositorygives you the link to go and look. - Rows are a snapshot. Versions and download counts move daily.
🧭 Which developer data scraper do you need?
| If you want | Use |
|---|---|
| npm and PyPI package metadata and downloads | This one |
| Repositories, stars and issues from GitHub | GitHub Scraper |
| Stack Exchange questions, tags and scores | Stack Overflow Scraper |
| Developer articles and their tags | DEV Community Scraper |
❓ Questions people ask
Can I search PyPI by keyword? No. PyPI has no clean public search API, so names are the only route. Find the name elsewhere and look it up here.
Why is license empty? You came in through the npm search. Search results do not carry it. Take the names from that run and look them up in a second run.
Can I mix a search and a name list? On npm, yes. Both run and the results are deduplicated.
Does maxItems protect me from a huge name list? No. It caps the search only. Trim the list yourself.
What happens to a name that does not exist? It comes back as its own NOT_FOUND row, you are not charged for it, and the other names still run.
Is this legal? Both registries publish this through open APIs meant to be read, and it is package metadata rather than personal data. Apify's write-up on scraping and the law is a good starting point, and we are not lawyers.
🆘 If something breaks
Open the Issues tab on the actor page. Send the registry, the names or query you used and the run id. The errorCode on the diagnostic row usually names the problem by itself.