GitHub Scraper
Search GitHub repos and users: stars, forks, language, topics, licence, plus user bio, company and followers. No token needed. $0.90 per 1,000 rows.
How it works
- 1Open it on Apify
Hit Run on Apify — it opens the tool in the cloud, no install.
- 2Set the inputs
Adjust
query,type,sort(sensible defaults are pre-filled). - 3Click Run
The tool runs on Apify’s cloud and collects the data for you.
- 4Export the results
Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.
Pricing
$0.0009 per item = $0.9 per 1,000
| You are charged for | When | Price |
|---|---|---|
| Item returned | Charged per repo or user returned. | $0.0009 |
Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-08-19, and they are what you are actually charged.
Inputs
| Field | What it does | Type |
|---|---|---|
query | GitHub search syntax. For repositories: "language:python stars:>1000 machine learning", "topic:cli created:>2023-01-01". For users: "location:berlin followers:>500", "fullstack in:bio". See GitHub's search docs for all qualifiers. | string |
type | What to search for: repositories or users/organizations. Users are additionally enriched with profile details (name, bio, company, location, followers, public repos). | string |
sort | How to sort repository results (applies to repository searches; user searches use GitHub's relevance ranking). Best match is GitHub's default relevance score. | string |
maxItems | Maximum number of repositories/users to return. GitHub's Search API caps at 1000 results per query (10 pages of 100), so split large jobs by qualifier (e.g. star range or date range). | integer |
notionConnector | Optional. Write each item as a page into your Notion when the run finishes. Authorize a Notion connector once in Settings → API & Integrations → MCP connectors, then pick it here. Leave empty to skip (default), results are always saved to the dataset regardless. | string |
notionParentId | Optional. The Notion data source ID of the database to write into (only used if a Notion connector is set). Leave empty to create the pages privately in your workspace instead. | string |
githubToken | Optional GitHub personal access token. Strongly recommended: without it you get only 60 requests/hr and 10 searches/min; with it you get 5000 requests/hr and 30 searches/min, so larger jobs finish. No special scopes needed for public data. Kept private. | string |
What you get
A structured dataset — each result includes fields like:
createdAtdefaultBranchdescriptiondetailsforksfullNamehomepageidlanguagelicenseloginnameopenIssuesownerstarsurlExport every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.
18 ready-to-run use cases
AI Agent GitHub Repos Created in 2026
Repos tagged ai-agents, created since January, already past 100 stars, ranked by stars. The created date is on every row so you can see the pace.
npm Dependency Audit: Repos Idle Since Mid-2023
Popular npm-tagged repos with no push since June 2023, archived ones excluded, ranked by stars. Idle is not the same as unsafe, but it is where to look.
Awesome Self-Hosted Apps Ranked by GitHub Stars
Repos tagged self-hosted above 2000 stars, archived ones excluded, with stars, language and last push. Raw material for a curated list.
React Native GitHub Repos, 1k+ Stars and Still Active
React Native repos above 1000 stars with a commit this year, so abandoned projects drop out. Stars and last-push date sit on every row.
GitHub Top Repositories for Solidity Smart Contracts
Solidity repos tagged smart-contracts above 500 stars, with forks, licence and last push. Repo metadata; contract source is not pulled.
GitHub Repository Search: MIT-Licensed Go Repos
Non-fork Go repos under MIT, ranked by stars, with the licence on every row. Metadata for assembling a corpus; no source code is downloaded.
MCP Servers on GitHub Created in 2026
Repos tagged mcp-server, created since January 2026, above 20 stars, newest commit first. Worth a look if you track what agent tooling shows up.
GitHub Organization Repositories Ranked by Stars
Swap org:stripe for any organisation and get its repos by stars, with language, forks and topics. Public repos only, as you would expect.
GitHub Trending Repos in Data Engineering, 1k+ Stars
Repos tagged data-engineering, created since 2025, already past 1000 stars. Created date and star count together show how fast they got there.
Rust Developers on GitHub in San Francisco
Accounts writing Rust with San Francisco in their location field and 50+ followers. Name, bio, blog and repo count. No email addresses.
FastAPI GitHub Repos and Who Maintains Them
The most-starred FastAPI projects with stars, forks, topics and the owner login on each row. Owner is the account holding the repo, not a contributor list.
Good First Issues: Beginner-Friendly Python Repos
Python repos tagged good-first-issues with 200+ stars, newest commit first. Rows are repositories, so the individual issues are not pulled.
GitHub Profile Scraper for TypeScript Developers
Accounts writing TypeScript with 100+ followers and 20+ public repos, with name, bio, blog and repo count. No emails, and no filter by employer.
GitHub Security Repositories in JavaScript, by Commit
JavaScript repos matching security and vulnerability terms, most recent commits first, with stars and pushed date. Repo metadata, not scan results.
GitHub Repos by Stars: Python Web Frameworks
Python repos over 5000 stars matching web framework, with forks, licence and last-updated date. It is a keyword search, so a few near-misses come through.
LLM GitHub Repos Sorted by Latest Commit
Repos tagged llm with 1000+ stars, most recently pushed first. Shows which large-model projects are still shipping and which have gone quiet.
Most Forked Repos in Machine Learning on GitHub
Ranked by forks rather than stars, which surfaces what people actually build on. Forks, stars, language and last push on each row.
GitHub Developer Search: Go Devs Based in Berlin
Go developers whose GitHub location says Berlin, 50+ followers, with name, bio, blog and repo count. Company is blank on most accounts.
Related tools in Developer & Research Tools
Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.
Stack Overflow / Stack Exchange Scraper
Search Stack Overflow and Stack Exchange by keyword or tag. Score, answer count, views, reputation and body text. $2 per 1,000 questions.
Package Registry Scraper (npm + PyPI)
Get npm and PyPI package metadata as JSON. Version, license, author, repo, keywords and npm monthly downloads. $2 per 1,000 packages.
arXiv Scraper
Search arXiv papers by title, author, abstract or category. Get full abstracts, authors, categories, DOI, dates and PDF links. $2 per 1,000 papers.
OpenAlex Scholarly Works Scraper
Search 250M+ OpenAlex papers with no API key. Get titles, authors, venue, year, citations, DOI, OA links and full abstracts. $2.00 per 1,000 papers.
Crossref Scholarly Works Scraper
Search 150M+ papers on Crossref: DOI, title, authors, journal, publisher, date, citations and abstract. No API key. $1.00 per 1,000 works.
Wikipedia Scraper
Search Wikipedia by keyword or by exact title. Get the intro text, the full article, thumbnails and categories. Any language. $1.00 per 1,000 pages.
Where this tool sits
- Categories
- Developer & Research Tools
- Platforms
- GitHub
GitHub Scraper: repositories and users from a GitHub search query
Type the query you would type into GitHub's own search box, language:rust stars:>10000, and you get one row per result. For repositories that is stars, forks, open issues, language, topics, licence and the three dates. For users it is the profile: name, bio, company, location, followers and public repo count.
The cap is GitHub's, not ours: any search stops at 1,000 results, however many matches it reports. Split a big job by star range or by date and run the pieces.
Worth knowing before you trust an empty result. GitHub answers a query with a broken qualifier, and a request with a bad token, in a way this actor currently reports as "no matches". So if a search you expect results from comes back with nothing, check the qualifier spelling and the token before concluding the answer is really empty.
| Input | One GitHub search query, in GitHub's own syntax |
| Output | One row per repository or user |
| Ceiling | 1,000 results per query, which is GitHub's own limit |
| Account needed | None. Your own GitHub token is optional and lifts the rate limits a long way |
| Price | $0.90 per 1,000 rows, which is $0.0009 each, flat on every plan |
🐙 What GitHub Scraper does
It runs your query against GitHub's search, pages through the results 100 at a time, and writes a row for each one. Repository searches sort by stars, forks, recent update or best match. User searches use GitHub's own relevance ranking and ignore the sort setting, because GitHub does.
Every qualifier GitHub documents works, since the query goes through as you wrote it: topic:cli created:>2023-01-01, location:berlin followers:>500, fullstack in:bio.
A user row is the search result plus a second call for the profile, so you get the bio, company and follower count rather than just a login.
Duplicates across pages are dropped before anything is written, so a repository appearing twice in GitHub's paging does not cost you twice.
Add your own GitHub token and the run gets GitHub's signed-in allowance, 5,000 requests an hour and 30 searches a minute, instead of the small signed-out one. No special scopes: public data only. Bigger jobs need it to finish.
📥 What you give it
{
"query": "language:python stars:>5000 web framework",
"type": "repositories",
"sort": "stars",
"maxItems": 200
}
| Field | Default | What it is |
|---|---|---|
query | none | GitHub search syntax, exactly as the site takes it. The form arrives with an example in it. |
type | repositories | repositories or users. Users and organisations both come back under users. |
sort | stars | stars, forks, updated or best-match. Repository searches only. |
maxItems | 100 | 1 to 1,000. GitHub stops at 1,000 for any single query. |
githubToken | none | Your own personal access token, marked secret. No scopes needed for public data. Recommended for anything past a small run. |
notionConnector, notionParentId | none | Optional. Also write each row into your connected Notion workspace. |
proxyConfiguration | none | Optional and usually unnecessary. |
📤 What you get back
A real repository row from a real run:
{
"ok": true,
"fullName": "fastapi/fastapi",
"name": "fastapi",
"owner": "fastapi",
"url": "https://github.com/fastapi/fastapi",
"description": "FastAPI framework, high performance, easy to learn, fast to code, ready for production",
"stars": 102502,
"forks": 9921,
"openIssues": 82,
"language": "Python",
"topics": ["api", "async", "asyncio", "fastapi", "framework", "json", "openapi", "python", "rest", "swagger"],
"license": "MIT",
"homepage": "https://fastapi.tiangolo.com/",
"defaultBranch": "master",
"createdAt": "2018-12-08T08:21:47Z",
"updatedAt": "2026-09-21T12:45:53Z",
"pushedAt": "2026-09-18T21:24:37Z"
}
Repository rows:
| Field | What it is |
|---|---|
fullName, name, owner, url | The repo and who owns it. fullName is your key. |
stars, forks, openIssues | Numbers at read time, not running totals. |
language, topics, license | The main language GitHub detects, the topic tags, and the licence as its SPDX id. |
homepage, defaultBranch | As set on the repo. Null when there is none. |
createdAt, updatedAt, pushedAt | pushedAt is the one that tells you whether a project is alive. updatedAt moves on a star, too. |
User rows, from a type: users run:
| Field | What it is |
|---|---|
login, url, type, id | The account. type separates a person from an organisation. |
name, bio, company, location, blog | Profile text, as filled in. Plenty of accounts leave these empty. |
followers, publicRepos, createdAt | Follower count, public repo count, and when the account was made. |
🧾 Reading the output
Data rows carry ok: true. Notes carry ok: false and an errorCode, and are never charged.
| Row | How to spot it | Charged |
|---|---|---|
| A repository or user | ok is true | yes |
| A diagnostic | ok is false, and errorCode says what happened | no |
| Code | What it means |
|---|---|
BAD_INPUT | The query was empty. |
NO_RESULTS | GitHub returned no matches. Check the qualifier spelling and your token before believing it, since a rejected query currently arrives here too. |
RATE_LIMITED | GitHub's limit was hit. rateLimitResetsAt says when it clears. Add a token, or ask for fewer items. |
BLOCKED | GitHub refused the request. Usually its short-term limit on bursts, so wait a minute and re-run. |
NOT_FOUND | The thing being asked for is not there. |
SERVER_ERROR | GitHub's own error. Worth one re-run. |
NETWORK | GitHub was unreachable or answered badly. |
The default table view is built for repositories, so a users run looks blank in it. Open a row, or export the dataset, and the profile fields are all there.
▶️ How to run it
1. Open GitHub Scraper and click Try for free. 2. Put your query into Search query. Build it on GitHub first if you are not sure of it, then paste it across. 3. Pick Search type: repositories or users. 4. Set Max items, up to 1,000. For anything past a hundred or so, paste a GitHub token. 5. Click Start, then download the dataset as JSON, CSV or Excel, or read it from the API.
💰 How much does it cost?
$0.90 per 1,000 rows, which is $0.0009 each. Flat on every Apify plan, no volume tiers.
You pay per result row, repository or user. Diagnostics are not charged, and neither are duplicates that GitHub's paging returned twice, since they are dropped before anything is written.
A thousand repositories, which is the most any single query can give you, is ninety cents.
💡 What people use it for
- Finding maintainers to talk to:
location:berlin language:go followers:>200as a user search. - Tracking a topic over time. The same query weekly, keyed on
fullName, shows what is growing. - Competitive work on open source: who forked what, which projects stopped being pushed to.
- Licence audits across a topic, using the SPDX id on every row.
- Feeding a Notion database of interesting repos straight from the run.
🚧 What it does not do
- 1,000 results per query. GitHub's cap. Split by
stars:1000..5000, bycreated:window, or
by language.
- No repo contents. No README, no file tree, no issues, no pull requests, no contributor list.
- No private data. Public repositories and public profiles only, whatever token you paste.
sortdoes nothing on a user search. GitHub ranks those itself.- A rejected query reads as an empty one. A misspelt qualifier or a bad token comes back as "no
matches" rather than as an error.
- A rate limit part-way through loses that run's rows. You get the note rather than a partial
page, so use a token for long jobs.
- The odd user row comes back thin, with only the login and URL, when GitHub's profile call did
not answer for that account.
- Counts are a snapshot. Stars move.
🧭 Which developer data scraper do you need?
| If you want | Use |
|---|---|
| GitHub repositories or users from a search | This one |
| Questions, tags and scores from Stack Overflow | Stack Overflow Scraper |
| npm and PyPI package metadata and downloads | npm + PyPI Package Scraper |
| Structured facts and entity claims | Wikidata Scraper |
❓ Questions people ask
Do I need a GitHub token? Not for a small run. For anything bigger, yes, because the signed-out allowance is tight and a run that hits it stops.
Is my token safe? It goes in a secret input field, it is only used against GitHub's own API, and it never appears in the dataset or the log. A read-only token with no scopes is enough.
Why did my query return nothing? Check the qualifier first. stars:>1000 works, star:>1000 does not, and a rejected query looks like an empty one here.
How do I get more than 1,000 repos? Run several queries with narrower ranges. Star bands and creation-date windows split most jobs cleanly.
Can I search organisations? Yes, they come back in a users search with type telling you which is which.
Is scraping GitHub legal? This uses GitHub's own public search API and returns public data. Profiles are personal data under GDPR and similar laws, so have a reason for holding them. Apify's write-up on the legality of web scraping is a good starting point, and we are not lawyers.
🆘 If something breaks
Open the Issues tab on the actor page. Send the exact query and the run ID. If a diagnostic row landed, its errorCode and hint usually name the reason already. Never paste your token into an issue.