Request a tool
All toolsAutomationsGuidesMCP serverRequest a toolPlatformsCategories
Package Registry Scraper (npm + PyPI) icon

Package Registry Scraper (npm + PyPI)

Get npm and PyPI package metadata as JSON. Version, license, author, repo, keywords and npm monthly downloads. $2 per 1,000 packages.

278 runs on Apify $0.002 per package ($2 / 1,000)
Run this in the cloudRun on Apify →

Developer & Research Tools

How it works

  1. 1
    Open it on Apify

    Hit Run on Apify — it opens the tool in the cloud, no install.

  2. 2
    Set the inputs

    Adjust registry, searchQuery, packageNames (sensible defaults are pre-filled).

  3. 3
    Click Run

    The tool runs on Apify’s cloud and collects the data for you.

  4. 4
    Export the results

    Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.

Pricing

$0.002 per package = $2 per 1,000

You are charged forWhenPrice
Package returnedCharged per package returned.$0.002

Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-06-13, and they are what you are actually charged.

Inputs

FieldWhat it doesType
registryWhich package registry to use. npm supports both keyword search and exact-name lookup; PyPI supports exact-name lookup only (it has no clean public search API).string
searchQueryKeywords to search the npm registry for (e.g. "react state management"). npm only, ignored for PyPI. Leave empty if you are looking up exact package names instead.string
packageNamesExact package names to look up directly. Works for BOTH registries. For PyPI this is the only supported mode (e.g. ["requests", "fastapi"]). For npm, scoped names like "@types/node" are supported.array
maxItemsMaximum number of packages to return from an npm search query. Only applies to npm search; ignored for exact-name lookups.integer
includeDownloadsFetch last-month download counts for each npm package via the npm downloads API. npm only. PyPI does not expose a public download-count endpoint. Adds one request per package. Note: monthlyDownloads is null when this is off, for PyPI packages, or if the downloads API call fails for a given package (a warning is logged in that case).boolean
notionConnectorOptional. Write each package as a page into your Notion when the run finishes. Authorize a Notion connector once in Settings → API & Integrations → MCP connectors, then pick it here. Leave empty to skip (default), results are always saved to the dataset regardless.string
notionParentIdOptional. The Notion data source ID of the database to write into (only used if a Notion connector is set). Leave empty to create the pages privately in your workspace instead.string

What you get

A structured dataset — each result includes fields like:

authordescriptionhomepagekeywordslicensemonthlyDownloadsnameregistryrepositoryscoreurlversion

Export every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.

2 ready-to-run use cases

npm Package Search with Monthly Download Counts

npm package search by keyword: name, version, description, repo link and monthly downloads per hit. Search hits carry no license field. Name lookups do.

npm License Checker for a List of Dependencies

Feed the names from your package.json and get the license and source repo back for each npm package. A compliance pass without installing anything.

Ready-made automation using this tool

A finished workflow you can import into n8n and run — this tool does the data-gathering inside it. Free, and yours to change.

Know when a package you depend on ships a new version

Watches the npm or PyPI packages you list and tells you when one releases — but only the kinds of release you asked for. Set it to major versions only and you hear about breaking changes without twenty patch notifications a week.

All automations →

Related tools in Developer & Research Tools

Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.

arXiv Scraper iconDeveloper & Research Tools

arXiv Scraper

Search arXiv papers by title, author, abstract or category. Get full abstracts, authors, categories, DOI, dates and PDF links. $2 per 1,000 papers.

2 use cases

OpenAlex Scholarly Works Scraper iconDeveloper & Research Tools

OpenAlex Scholarly Works Scraper

Search 250M+ OpenAlex papers with no API key. Get titles, authors, venue, year, citations, DOI, OA links and full abstracts. $2.00 per 1,000 papers.

2 use cases

Crossref Scholarly Works Scraper iconDeveloper & Research Tools

Crossref Scholarly Works Scraper

Search 150M+ papers on Crossref: DOI, title, authors, journal, publisher, date, citations and abstract. No API key. $1.00 per 1,000 works.

2 use cases

Wikipedia Scraper iconDeveloper & Research Tools

Wikipedia Scraper

Search Wikipedia by keyword or by exact title. Get the intro text, the full article, thumbnails and categories. Any language. $1.00 per 1,000 pages.

3 use cases

Internet Archive Scraper iconDeveloper & Research Tools

Internet Archive Scraper

Search Internet Archive (archive.org) for books, audio, film and web items. Title, creator, year, downloads, subjects and URL. $2.00 per 1,000 items.

2 use cases

Hacker News Scraper iconDeveloper & Research Tools

Hacker News Scraper

Search HN stories, Show HN, Ask HN and comments, or pull the front page. Get points, author, comment counts and links. $1 per 1,000 items.

3 use cases

See all Developer & Research Tools →

Package Registry Scraper: npm and PyPI metadata in one row shape

Look up packages by exact name on npm or PyPI, or search npm by keywords, and get one row each: name, latest version, description, author, homepage, repository, licence, keywords and the project page URL. npm rows can also carry last month's download count.

The two registries are not symmetrical, and it matters before you plan a run. npm does keyword search and exact names. PyPI does exact names only, because it publishes no clean search API for anyone to call.

InputExact package names, or npm keywords
OutputOne row per package
Ceiling1,000 packages per npm search. Name lookups are not capped
Account neededNone, and no registry key
Price$2.00 per 1,000 packages, flat on every plan

🔍 What Package Registry Scraper does

Two ways in, and you can use both at once on npm.

Exact names. Put a list in packageNames and each one is looked up directly. Scoped npm names like @types/node work. A name that does not exist gets its own NOT_FOUND row, uncharged, and the run carries on with the rest.

npm keyword search. Put words in searchQuery and the registry's own search decides what matches, ordered by its relevance score, up to your maxItems.

There is a real difference between the two that the field list hides: a search row has no licence and usually no repository link. npm's search index does not carry them. If you are doing a licence audit, feed the names in through packageNames and look them up properly.

includeDownloads is on by default and adds last month's download count to npm rows. PyPI has no public endpoint for that, so PyPI rows never carry it.

📥 What you give it

{
  "registry": "npm",
  "packageNames": ["react", "@types/node"],
  "includeDownloads": true
}
FieldDefaultWhat it is
registrynpmnpm or pypi.
packageNamesnoneExact names. Works on both registries, and is the only mode PyPI supports.
searchQuerynoneKeywords, npm only. Ignored on PyPI unless you sent no names, in which case you get BAD_INPUT. The Console shows react state management as an example, but that is a prefill, so an API call has to send its own.
maxItems501 to 1,000. Caps the npm search only. It does nothing to a list of names.
includeDownloadsonLast month's downloads for npm packages. Costs one extra lookup per package.
notionConnectornoneOptional. Writes every delivered package into your Notion. Authorise a connector once under Settings, API and Integrations, MCP connectors, then pick it here.
notionParentIdnoneOptional. The Notion data source id to write into. Leave it empty and the pages are created privately in your workspace.
proxyConfigurationoffOptional network setting. Off is right for a normal run.

packageNames has no ceiling. Paste a 4,000-line requirements.txt and you get up to 4,000 charged rows, because maxItems is not watching that path. Trim the list to what you need.

📤 What you get back

A real row from a recent npm search:

{
  "ok": true,
  "registry": "npm",
  "name": "unstated-next",
  "version": "1.1.0",
  "description": "200 bytes to never think about React state management libraries ever again",
  "author": "thejameskyle",
  "homepage": null,
  "repository": null,
  "license": null,
  "keywords": [],
  "score": 364.87637,
  "url": "https://www.npmjs.com/package/unstated-next",
  "monthlyDownloads": 330524
}

That row shows the search-path gap plainly: license, repository, homepage and keywords are all empty, because the search index does not carry them. Look the same package up by name and they are filled in.

FieldWhat it is
versionThe latest published version. npm's latest tag, PyPI's current release.
licenseThe declared licence string. null on every npm search row.
repositoryNormalised into a clickable https URL rather than a git address.
scorenpm's own search relevance. Present on search rows, null on npm name lookups, and the key is absent from PyPI rows entirely.
monthlyDownloadsnpm only. null if the count could not be read, and the key is absent when includeDownloads is off or the row is from PyPI.
authorOn a search row this can be the publishing account's handle rather than the person named in the package file.

🧾 Reading the output

Packages carry ok: true. Anything with ok: false carries an errorCode and is not charged. One run can contain both, which is normal when a name list has a typo in it.

CodeWhat it means
BAD_INPUTNeither names nor a query, or a PyPI run with a searchQuery and no names.
NOT_FOUNDThat one name is not in the registry. Check the spelling, the scope, and whether it was unpublished.
NO_RESULTSNothing at all came back. An npm search that matched no packages ends here.
NETWORKThe registry was unreachable or answered with something unusable. Re-run it.

The Console's Overview table hides license and repository, which are the two columns a licence audit actually needs. Download the dataset as JSON, CSV or Excel, or switch the table to all fields.

▶️ How to run it

1. Open Package Registry Scraper and click Try for free. 2. Pick a Registry. 3. Paste names into Package names, one per line. For an npm search, type into Search query instead. 4. Leave Include monthly downloads on unless you do not want it, then click Start. 5. Download the dataset as JSON, CSV, Excel or XML.

💰 How much does it cost?

$2.00 per 1,000 packages. Flat on every Apify plan, no volume tiers, and the same on both registries.

You pay per package row. The same name twice in one run is charged once. NOT_FOUND names, a search that matched nothing and every diagnostic row are not charged. Turning downloads on does not change the price.

💡 What people use it for

  • Auditing the licences behind a package.json or a requirements.txt by feeding the names in and

reading the license column.

  • Comparing two libraries side by side on version, downloads and last release before choosing one.
  • Watching a dependency on a schedule and firing an alert when version moves.
  • Sizing a niche: search npm for the keywords, then rank by monthlyDownloads to see what people

actually install.

🚧 What it does not do

  • No PyPI search. Exact names only on that side.
  • No PyPI download counts. There is no public endpoint to read them from.
  • No licence on npm search rows. Use packageNames when the licence is the point.
  • No dependency trees and no version history. One row is the current state of one package.
  • No download trend. monthlyDownloads is last month's total, a single number, not a series.
  • No GitHub stars or issues. repository gives you the link to go and look.
  • Rows are a snapshot. Versions and download counts move daily.

🧭 Which developer data scraper do you need?

If you wantUse
npm and PyPI package metadata and downloadsThis one
Repositories, stars and issues from GitHubGitHub Scraper
Stack Exchange questions, tags and scoresStack Overflow Scraper
Developer articles and their tagsDEV Community Scraper

❓ Questions people ask

Can I search PyPI by keyword? No. PyPI has no clean public search API, so names are the only route. Find the name elsewhere and look it up here.

Why is license empty? You came in through the npm search. Search results do not carry it. Take the names from that run and look them up in a second run.

Can I mix a search and a name list? On npm, yes. Both run and the results are deduplicated.

Does maxItems protect me from a huge name list? No. It caps the search only. Trim the list yourself.

What happens to a name that does not exist? It comes back as its own NOT_FOUND row, you are not charged for it, and the other names still run.

Is this legal? Both registries publish this through open APIs meant to be read, and it is package metadata rather than personal data. Apify's write-up on scraping and the law is a good starting point, and we are not lawyers.

🆘 If something breaks

Open the Issues tab on the actor page. Send the registry, the names or query you used and the run id. The errorCode on the diagnostic row usually names the problem by itself.