Request a tool
All toolsAutomationsGuidesMCP serverRequest a toolPlatformsCategories
GitHub Scraper icon

GitHub Scraper

Search GitHub repos and users: stars, forks, language, topics, licence, plus user bio, company and followers. No token needed. $0.90 per 1,000 rows.

5 from 2 reviews on Apify 4,356 runs on Apify $0.0009 per item ($0.9 / 1,000)
Run this in the cloudRun on Apify →

Developer & Research Tools

How it works

  1. 1
    Open it on Apify

    Hit Run on Apify — it opens the tool in the cloud, no install.

  2. 2
    Set the inputs

    Adjust query, type, sort (sensible defaults are pre-filled).

  3. 3
    Click Run

    The tool runs on Apify’s cloud and collects the data for you.

  4. 4
    Export the results

    Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.

Pricing

$0.0009 per item = $0.9 per 1,000

You are charged forWhenPrice
Item returnedCharged per repo or user returned.$0.0009

Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-08-19, and they are what you are actually charged.

Inputs

FieldWhat it doesType
queryGitHub search syntax. For repositories: "language:python stars:>1000 machine learning", "topic:cli created:>2023-01-01". For users: "location:berlin followers:>500", "fullstack in:bio". See GitHub's search docs for all qualifiers.string
typeWhat to search for: repositories or users/organizations. Users are additionally enriched with profile details (name, bio, company, location, followers, public repos).string
sortHow to sort repository results (applies to repository searches; user searches use GitHub's relevance ranking). Best match is GitHub's default relevance score.string
maxItemsMaximum number of repositories/users to return. GitHub's Search API caps at 1000 results per query (10 pages of 100), so split large jobs by qualifier (e.g. star range or date range).integer
notionConnectorOptional. Write each item as a page into your Notion when the run finishes. Authorize a Notion connector once in Settings → API & Integrations → MCP connectors, then pick it here. Leave empty to skip (default), results are always saved to the dataset regardless.string
notionParentIdOptional. The Notion data source ID of the database to write into (only used if a Notion connector is set). Leave empty to create the pages privately in your workspace instead.string
githubTokenOptional GitHub personal access token. Strongly recommended: without it you get only 60 requests/hr and 10 searches/min; with it you get 5000 requests/hr and 30 searches/min, so larger jobs finish. No special scopes needed for public data. Kept private.string

What you get

A structured dataset — each result includes fields like:

createdAtdefaultBranchdescriptiondetailsforksfullNamehomepageidlanguagelicenseloginnameopenIssuesownerstarsurl

Export every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.

18 ready-to-run use cases

AI Agent GitHub Repos Created in 2026

Repos tagged ai-agents, created since January, already past 100 stars, ranked by stars. The created date is on every row so you can see the pace.

npm Dependency Audit: Repos Idle Since Mid-2023

Popular npm-tagged repos with no push since June 2023, archived ones excluded, ranked by stars. Idle is not the same as unsafe, but it is where to look.

Awesome Self-Hosted Apps Ranked by GitHub Stars

Repos tagged self-hosted above 2000 stars, archived ones excluded, with stars, language and last push. Raw material for a curated list.

React Native GitHub Repos, 1k+ Stars and Still Active

React Native repos above 1000 stars with a commit this year, so abandoned projects drop out. Stars and last-push date sit on every row.

GitHub Top Repositories for Solidity Smart Contracts

Solidity repos tagged smart-contracts above 500 stars, with forks, licence and last push. Repo metadata; contract source is not pulled.

GitHub Repository Search: MIT-Licensed Go Repos

Non-fork Go repos under MIT, ranked by stars, with the licence on every row. Metadata for assembling a corpus; no source code is downloaded.

MCP Servers on GitHub Created in 2026

Repos tagged mcp-server, created since January 2026, above 20 stars, newest commit first. Worth a look if you track what agent tooling shows up.

GitHub Organization Repositories Ranked by Stars

Swap org:stripe for any organisation and get its repos by stars, with language, forks and topics. Public repos only, as you would expect.

GitHub Trending Repos in Data Engineering, 1k+ Stars

Repos tagged data-engineering, created since 2025, already past 1000 stars. Created date and star count together show how fast they got there.

Rust Developers on GitHub in San Francisco

Accounts writing Rust with San Francisco in their location field and 50+ followers. Name, bio, blog and repo count. No email addresses.

FastAPI GitHub Repos and Who Maintains Them

The most-starred FastAPI projects with stars, forks, topics and the owner login on each row. Owner is the account holding the repo, not a contributor list.

Good First Issues: Beginner-Friendly Python Repos

Python repos tagged good-first-issues with 200+ stars, newest commit first. Rows are repositories, so the individual issues are not pulled.

GitHub Profile Scraper for TypeScript Developers

Accounts writing TypeScript with 100+ followers and 20+ public repos, with name, bio, blog and repo count. No emails, and no filter by employer.

GitHub Security Repositories in JavaScript, by Commit

JavaScript repos matching security and vulnerability terms, most recent commits first, with stars and pushed date. Repo metadata, not scan results.

GitHub Repos by Stars: Python Web Frameworks

Python repos over 5000 stars matching web framework, with forks, licence and last-updated date. It is a keyword search, so a few near-misses come through.

LLM GitHub Repos Sorted by Latest Commit

Repos tagged llm with 1000+ stars, most recently pushed first. Shows which large-model projects are still shipping and which have gone quiet.

Most Forked Repos in Machine Learning on GitHub

Ranked by forks rather than stars, which surfaces what people actually build on. Forks, stars, language and last push on each row.

GitHub Developer Search: Go Devs Based in Berlin

Go developers whose GitHub location says Berlin, 50+ followers, with name, bio, blog and repo count. Company is blank on most accounts.

Related tools in Developer & Research Tools

Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.

Stack Overflow / Stack Exchange Scraper iconDeveloper & Research Tools

Stack Overflow / Stack Exchange Scraper

Search Stack Overflow and Stack Exchange by keyword or tag. Score, answer count, views, reputation and body text. $2 per 1,000 questions.

2 use cases

Package Registry Scraper (npm + PyPI) iconDeveloper & Research Tools

Package Registry Scraper (npm + PyPI)

Get npm and PyPI package metadata as JSON. Version, license, author, repo, keywords and npm monthly downloads. $2 per 1,000 packages.

2 use cases

arXiv Scraper iconDeveloper & Research Tools

arXiv Scraper

Search arXiv papers by title, author, abstract or category. Get full abstracts, authors, categories, DOI, dates and PDF links. $2 per 1,000 papers.

2 use cases

OpenAlex Scholarly Works Scraper iconDeveloper & Research Tools

OpenAlex Scholarly Works Scraper

Search 250M+ OpenAlex papers with no API key. Get titles, authors, venue, year, citations, DOI, OA links and full abstracts. $2.00 per 1,000 papers.

2 use cases

Crossref Scholarly Works Scraper iconDeveloper & Research Tools

Crossref Scholarly Works Scraper

Search 150M+ papers on Crossref: DOI, title, authors, journal, publisher, date, citations and abstract. No API key. $1.00 per 1,000 works.

2 use cases

Wikipedia Scraper iconDeveloper & Research Tools

Wikipedia Scraper

Search Wikipedia by keyword or by exact title. Get the intro text, the full article, thumbnails and categories. Any language. $1.00 per 1,000 pages.

3 use cases

See all Developer & Research Tools →

GitHub Scraper: repositories and users from a GitHub search query

Type the query you would type into GitHub's own search box, language:rust stars:>10000, and you get one row per result. For repositories that is stars, forks, open issues, language, topics, licence and the three dates. For users it is the profile: name, bio, company, location, followers and public repo count.

The cap is GitHub's, not ours: any search stops at 1,000 results, however many matches it reports. Split a big job by star range or by date and run the pieces.

Worth knowing before you trust an empty result. GitHub answers a query with a broken qualifier, and a request with a bad token, in a way this actor currently reports as "no matches". So if a search you expect results from comes back with nothing, check the qualifier spelling and the token before concluding the answer is really empty.

InputOne GitHub search query, in GitHub's own syntax
OutputOne row per repository or user
Ceiling1,000 results per query, which is GitHub's own limit
Account neededNone. Your own GitHub token is optional and lifts the rate limits a long way
Price$0.90 per 1,000 rows, which is $0.0009 each, flat on every plan

🐙 What GitHub Scraper does

It runs your query against GitHub's search, pages through the results 100 at a time, and writes a row for each one. Repository searches sort by stars, forks, recent update or best match. User searches use GitHub's own relevance ranking and ignore the sort setting, because GitHub does.

Every qualifier GitHub documents works, since the query goes through as you wrote it: topic:cli created:>2023-01-01, location:berlin followers:>500, fullstack in:bio.

A user row is the search result plus a second call for the profile, so you get the bio, company and follower count rather than just a login.

Duplicates across pages are dropped before anything is written, so a repository appearing twice in GitHub's paging does not cost you twice.

Add your own GitHub token and the run gets GitHub's signed-in allowance, 5,000 requests an hour and 30 searches a minute, instead of the small signed-out one. No special scopes: public data only. Bigger jobs need it to finish.

📥 What you give it

{
  "query": "language:python stars:>5000 web framework",
  "type": "repositories",
  "sort": "stars",
  "maxItems": 200
}
FieldDefaultWhat it is
querynoneGitHub search syntax, exactly as the site takes it. The form arrives with an example in it.
typerepositoriesrepositories or users. Users and organisations both come back under users.
sortstarsstars, forks, updated or best-match. Repository searches only.
maxItems1001 to 1,000. GitHub stops at 1,000 for any single query.
githubTokennoneYour own personal access token, marked secret. No scopes needed for public data. Recommended for anything past a small run.
notionConnector, notionParentIdnoneOptional. Also write each row into your connected Notion workspace.
proxyConfigurationnoneOptional and usually unnecessary.

📤 What you get back

A real repository row from a real run:

{
  "ok": true,
  "fullName": "fastapi/fastapi",
  "name": "fastapi",
  "owner": "fastapi",
  "url": "https://github.com/fastapi/fastapi",
  "description": "FastAPI framework, high performance, easy to learn, fast to code, ready for production",
  "stars": 102502,
  "forks": 9921,
  "openIssues": 82,
  "language": "Python",
  "topics": ["api", "async", "asyncio", "fastapi", "framework", "json", "openapi", "python", "rest", "swagger"],
  "license": "MIT",
  "homepage": "https://fastapi.tiangolo.com/",
  "defaultBranch": "master",
  "createdAt": "2018-12-08T08:21:47Z",
  "updatedAt": "2026-09-21T12:45:53Z",
  "pushedAt": "2026-09-18T21:24:37Z"
}

Repository rows:

FieldWhat it is
fullName, name, owner, urlThe repo and who owns it. fullName is your key.
stars, forks, openIssuesNumbers at read time, not running totals.
language, topics, licenseThe main language GitHub detects, the topic tags, and the licence as its SPDX id.
homepage, defaultBranchAs set on the repo. Null when there is none.
createdAt, updatedAt, pushedAtpushedAt is the one that tells you whether a project is alive. updatedAt moves on a star, too.

User rows, from a type: users run:

FieldWhat it is
login, url, type, idThe account. type separates a person from an organisation.
name, bio, company, location, blogProfile text, as filled in. Plenty of accounts leave these empty.
followers, publicRepos, createdAtFollower count, public repo count, and when the account was made.

🧾 Reading the output

Data rows carry ok: true. Notes carry ok: false and an errorCode, and are never charged.

RowHow to spot itCharged
A repository or userok is trueyes
A diagnosticok is false, and errorCode says what happenedno
CodeWhat it means
BAD_INPUTThe query was empty.
NO_RESULTSGitHub returned no matches. Check the qualifier spelling and your token before believing it, since a rejected query currently arrives here too.
RATE_LIMITEDGitHub's limit was hit. rateLimitResetsAt says when it clears. Add a token, or ask for fewer items.
BLOCKEDGitHub refused the request. Usually its short-term limit on bursts, so wait a minute and re-run.
NOT_FOUNDThe thing being asked for is not there.
SERVER_ERRORGitHub's own error. Worth one re-run.
NETWORKGitHub was unreachable or answered badly.

The default table view is built for repositories, so a users run looks blank in it. Open a row, or export the dataset, and the profile fields are all there.

▶️ How to run it

1. Open GitHub Scraper and click Try for free. 2. Put your query into Search query. Build it on GitHub first if you are not sure of it, then paste it across. 3. Pick Search type: repositories or users. 4. Set Max items, up to 1,000. For anything past a hundred or so, paste a GitHub token. 5. Click Start, then download the dataset as JSON, CSV or Excel, or read it from the API.

💰 How much does it cost?

$0.90 per 1,000 rows, which is $0.0009 each. Flat on every Apify plan, no volume tiers.

You pay per result row, repository or user. Diagnostics are not charged, and neither are duplicates that GitHub's paging returned twice, since they are dropped before anything is written.

A thousand repositories, which is the most any single query can give you, is ninety cents.

💡 What people use it for

  • Finding maintainers to talk to: location:berlin language:go followers:>200 as a user search.
  • Tracking a topic over time. The same query weekly, keyed on fullName, shows what is growing.
  • Competitive work on open source: who forked what, which projects stopped being pushed to.
  • Licence audits across a topic, using the SPDX id on every row.
  • Feeding a Notion database of interesting repos straight from the run.

🚧 What it does not do

  • 1,000 results per query. GitHub's cap. Split by stars:1000..5000, by created: window, or

by language.

  • No repo contents. No README, no file tree, no issues, no pull requests, no contributor list.
  • No private data. Public repositories and public profiles only, whatever token you paste.
  • sort does nothing on a user search. GitHub ranks those itself.
  • A rejected query reads as an empty one. A misspelt qualifier or a bad token comes back as "no

matches" rather than as an error.

  • A rate limit part-way through loses that run's rows. You get the note rather than a partial

page, so use a token for long jobs.

  • The odd user row comes back thin, with only the login and URL, when GitHub's profile call did

not answer for that account.

  • Counts are a snapshot. Stars move.

🧭 Which developer data scraper do you need?

If you wantUse
GitHub repositories or users from a searchThis one
Questions, tags and scores from Stack OverflowStack Overflow Scraper
npm and PyPI package metadata and downloadsnpm + PyPI Package Scraper
Structured facts and entity claimsWikidata Scraper

❓ Questions people ask

Do I need a GitHub token? Not for a small run. For anything bigger, yes, because the signed-out allowance is tight and a run that hits it stops.

Is my token safe? It goes in a secret input field, it is only used against GitHub's own API, and it never appears in the dataset or the log. A read-only token with no scopes is enough.

Why did my query return nothing? Check the qualifier first. stars:>1000 works, star:>1000 does not, and a rejected query looks like an empty one here.

How do I get more than 1,000 repos? Run several queries with narrower ranges. Star bands and creation-date windows split most jobs cleanly.

Can I search organisations? Yes, they come back in a users search with type telling you which is which.

Is scraping GitHub legal? This uses GitHub's own public search API and returns public data. Profiles are personal data under GDPR and similar laws, so have a reason for holding them. Apify's write-up on the legality of web scraping is a good starting point, and we are not lawyers.

🆘 If something breaks

Open the Issues tab on the actor page. Send the exact query and the run ID. If a diagnostic row landed, its errorCode and hint usually name the reason already. Never paste your token into an issue.