Y Combinator Companies Scraper
Scrape the YC startup directory by batch, industry, region, tag or status. Get the pitch, team size, founded year, website and founders. $0.0009 each.
How it works
- 1Open it on Apify
Hit Run on Apify — it opens the tool in the cloud, no install.
- 2Set the inputs
Adjust
batches,industries,regions(sensible defaults are pre-filled). - 3Click Run
The tool runs on Apify’s cloud and collects the data for you.
- 4Export the results
Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.
Pricing
$0.0009 per company = $0.9 per 1,000
| You are charged for | When | Price |
|---|---|---|
| Company scraped | Each Y Combinator company written to the dataset. Sample and diagnostic rows are never charged. | $0.0009 |
Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-08-16, and they are what you are actually charged.
Inputs
| Field | What it does | Type |
|---|---|---|
batches | YC batches to pull, written either short (W24, S23, F24, P25) or long (Winter 2024, Summer 2023). Leave empty to search across every batch. Each batch is queried separately, so ten batches and 100 rows gives you ten companies from each. | array |
industries | Filter by the industry labels the directory itself uses, for example B2B, Consumer, Healthcare, Fintech, Industrials, Education, Infrastructure, Marketing, Security, Climate. Several values are combined with OR. | array |
regions | Filter by the region labels the directory uses, for example United States of America, Europe, India, Latin America, Remote, Canada, Africa, Southeast Asia. Several values are combined with OR. | array |
tags | Filter by the free-form tags on a company profile, for example Artificial Intelligence, SaaS, Developer Tools, Marketplace, Generative AI, Open Source, Health Tech. Several values are combined with OR. | array |
statuses | Keep only companies with these statuses. Leave empty for all four. | array |
stages | Keep only companies at these stages, as the directory labels them. Leave empty for both. The directory's search can't filter on stage, so the run reads its results and keeps the companies that match; the others are never charged and their profile pages aren't read. A stage filter can return fewer companies than you asked for: with one set, a run reads at most ten times Maximum companies from the directory, never fewer than 1,000 and never more than 10,000. It applies to searches, not to companies you name in Company profile URLs. | array |
searchTerms | Free-text search across company names, pitches and descriptions, for example "robotics", "climate", "payments in Africa". Each term is searched separately and can be combined with the filters above. | array |
companyUrls | Look up specific companies instead of searching. Paste full profile URLs (https://www.ycombinator.com/companies/doordash) or bare slugs (doordash). Up to 500 per run. | array |
allCompanies | Sweep the entire directory instead of filtering. Combine it with a high row limit; the run walks batch by batch, newest first. Remember you pay per company. | boolean |
isHiring | Keep only companies flagged as hiring on their directory card. | boolean |
topCompaniesOnly | Keep only the companies the directory marks as top companies. | boolean |
nonprofitOnly | Keep only the non-profit organisations in the directory. | boolean |
includeFounders | On by default. Reads each company's own profile page as well, which is what fills in founder names and titles, the year the company was founded, its city and country, and its LinkedIn, X, Facebook, Crunchbase and GitHub links. Turn it off for a faster, lighter run when you only need the directory card. The price per company is the same either way. | boolean |
maxItems | Total number of companies to return in this run, shared evenly across your batches and search terms. Default 20, hard ceiling 7,000 (the directory holds about 6,200). Keep it low while you are testing - you pay per company. | integer |
proxyUrls | Leave this empty for a normal run. Fill it in only if you want the traffic to leave through proxy servers you already pay for, one URL per line, in the form http://user:pass@host:port. | array |
What you get
A structured dataset — each result includes fields like:
nameoneLinerbatchstatusteamSizeyearFoundedlocationwebsiteycProfileUrlfounderNamesindustrytagslongDescriptionExport every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.
1 ready-to-run use cases
Y Combinator Companies That Are Hiring
List YC companies that are hiring, by batch: name, description, location, website and founders.
Related tools in Developer & Research Tools
Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.
GitHub Scraper
Search GitHub repos and users: stars, forks, language, topics, licence, plus user bio, company and followers. No token needed. $0.90 per 1,000 rows.
Stack Overflow / Stack Exchange Scraper
Search Stack Overflow and Stack Exchange by keyword or tag. Score, answer count, views, reputation and body text. $2 per 1,000 questions.
Package Registry Scraper (npm + PyPI)
Get npm and PyPI package metadata as JSON. Version, license, author, repo, keywords and npm monthly downloads. $2 per 1,000 packages.
arXiv Scraper
Search arXiv papers by title, author, abstract or category. Get full abstracts, authors, categories, DOI, dates and PDF links. $2 per 1,000 papers.
OpenAlex Scholarly Works Scraper
Search 250M+ OpenAlex papers with no API key. Get titles, authors, venue, year, citations, DOI, OA links and full abstracts. $2.00 per 1,000 papers.
Crossref Scholarly Works Scraper
Search 150M+ papers on Crossref: DOI, title, authors, journal, publisher, date, citations and abstract. No API key. $1.00 per 1,000 works.
Where this tool sits
- Categories
- Developer & Research Tools
Y Combinator Companies Scraper: search the YC startup directory by batch, industry or tag
Pick a batch like W24, an industry, a region, a tag or a status, or paste company profile URLs. You get one row per company: the pitch, the long description, batch, status, team size, founded year, location, website, tags and the founders with their LinkedIn and X links.
The awkward bit is the founded year. It comes off the company's own profile page, and a fair number of profiles simply do not carry one, so yearFounded arrives null rather than estimated.
| Input | Batches, industries, regions, tags, statuses, stages, search terms or profile URLs |
| Output | One row per company |
| Ceiling | 7,000 companies per run |
| Account needed | None |
| Price | $0.90 per 1,000 companies, flat on every plan |
🏢 What Y Combinator Companies Scraper does
The YC directory is a searchable index of every company that has been through the programme, and each company also has its own profile page. This reads both and joins them into one row.
The directory card gives you the name, one-liner, batch, status, team size, industries, regions, tags, logo and website. The profile page adds the founders with their titles and bios, the year the company was founded, the city and country, and links to LinkedIn, X, Facebook, Crunchbase and GitHub. Every row has the same keys whichever way you searched.
Filters combine the way you would expect. Values inside one list are OR'd together, and different lists narrow each other. Ask for W24 and S23 with industries: ["B2B"] and you get B2B companies from either batch.
📥 What you give it
Every field is optional. Run it with the input empty and you get one labelled sample row back.
{
"batches": ["W24", "S23"],
"industries": ["B2B"],
"statuses": ["Active"],
"maxItems": 200
}
| Field | Default | What it is |
|---|---|---|
batches | none | Short (W24, S23, F24, P25) or long (Winter 2024). Up to 60. Each is queried separately, so ten batches and 100 rows gives ten from each. |
industries | none | The directory's own labels: B2B, Consumer, Healthcare, Fintech and so on. Up to 60. |
regions | none | The directory's region labels, like United States of America, Europe, India, Remote. Up to 60. |
tags | none | The free-form profile tags, like Generative AI, Marketplace, Open Source. Up to 60. |
statuses | all four | Active, Acquired, Public, Inactive. |
stages | both | Early, Growth. The directory's search can't filter on stage, so the run reads its results and keeps the companies that match, reading at most ten times maxItems entries a run (1,000 to 10,000). A rare stage can return fewer companies than maxItems, with a free READ_LIMIT row saying so. Searches only, not companyUrls. |
searchTerms | none | Free text across names, pitches and descriptions. Each term is searched on its own. |
companyUrls | none | Up to 500 profile URLs or bare slugs, to look companies up instead of searching. |
allCompanies | off | Sweep the whole directory, batch by batch, newest first. Pair it with a high row limit. |
isHiring | not applied | Three-state. See the note below. |
topCompaniesOnly | off | Only companies the directory marks as top companies. |
nonprofitOnly | off | Only the non-profits. |
includeFounders | true | Reads each profile page too, which is what fills in the founders, founded year, city, country and social links. Turn it off for a lighter run. The price is the same either way. |
maxItems | 20 | Total rows for the run, 1 to 7,000, shared across your batches and terms. |
proxyUrls | none | Leave it empty for a normal run. It is there for callers who want traffic to leave through servers they already pay for, as http://user:pass@host:port. |
isHiring has three states, not two. Leave it out and no hiring filter is applied. Set it true and you get companies flagged as hiring. Set it to false and you get only the companies that are not hiring, which is a legitimate query and almost never what somebody means. The Console writes a literal false into the input once you have ticked the box and unticked it again, so if a run comes back full of companies that are not hiring, remove the field rather than unticking it.
📤 What you get back
A real row, from run NctVtHlEJQFnIqU35:
{
"ok": true, "charged": true, "recordType": "company", "query": "Winter 2024",
"name": "Indemni", "slug": "indemni",
"oneLiner": "Cargo Theft and Fraud Prevention Platform",
"longDescription": "We are building a safer supply chain. Cargo Theft has been increasing yearly, ...",
"batch": "Winter 2024", "batchCode": "W24", "status": "Active",
"teamSize": 7, "yearFounded": 2024,
"location": "San Francisco, CA, USA", "city": "San Francisco", "country": "US",
"regions": ["United States of America", "America / Canada", "Remote", "Partly Remote"],
"website": "http://www.indemni.com",
"ycProfileUrl": "https://www.ycombinator.com/companies/indemni",
"industry": "B2B", "industries": ["B2B", "Supply Chain and Logistics"],
"subindustry": "B2B -> Supply Chain and Logistics",
"tags": ["Identity", "Logistics", "Supply Chain", "Fraud Prevention", "Fraud Detection"],
"founders": [
{ "name": "Omar Draz", "title": "Founder",
"linkedinUrl": "https://linkedin.com/in/odraz",
"twitterUrl": "https://twitter.com/oamdraz",
"bio": "Ex-DoorDash, worked on Fraud, Growth and Logistics. Currently building!" }
],
"founderNames": "Omar Draz", "founderCount": 1,
"linkedinUrl": "https://www.linkedin.com/company/100487698/admin",
"twitterUrl": null, "facebookUrl": null, "crunchbaseUrl": null, "githubUrl": null,
"logoUrl": "https://bookface-images.s3.amazonaws.com/small_logos/a4f88ce8a0...png",
"formerNames": ["Alacrity"],
"isHiring": false, "nonprofit": false, "topCompany": false, "stage": "Early",
"launchedAt": "2024-02-15T20:40:36.000Z", "companyId": "29533",
"scrapedAt": "2026-09-21T01:52:29.035Z"
}
longDescription and logoUrl are cut short above. A real row carries both in full.
| Field | What it is |
|---|---|
companyId | The directory's own id, as a string. Use it as your dedupe key. |
query | Which batch, term or filter produced this row, so you can tell where a result came from when you searched several things at once. |
founders | Objects with name, title, linkedinUrl, twitterUrl and bio. founderNames is the same names flattened for a spreadsheet. |
formerNames | Names the company has traded under before. Useful when your own list is out of date. |
batch, batchCode | Winter 2024 and W24. Both are there so you do not have to convert. |
stage | The directory's own label, Early or Growth. The stages filter reads it. |
launchedAt | When the directory record was created, in ISO UTC. Not the founding date. |
🧾 Reading the output
Three kinds of row.
| Row | How to spot it | Billed |
|---|---|---|
| A company | recordType: "company" | yes |
| The sample row | _sample: true, and only when the input named no filter and no URL | no |
| A diagnostic | _diagnostic: true and ok: false | no |
charged is set when the row is built, a moment before the charge goes out, so read it as "this is a real row" rather than as a receipt. A failed charge adds a CHARGE_ERROR diagnostic at the end.
The Console's default table hides the diagnostic columns, so those rows look blank there. Read them from the JSON.
| Code | What it means |
|---|---|
NO_RESULTS | The filter ran and matched nothing, or nothing at the stage you picked. An unrecognised batch string does this. |
READ_LIMIT | With stages set, the run read as many directory entries as it reads for one and stopped with what matched. |
BAD_INPUT | A value in stages is not Early or Growth. Nothing was searched. |
NOT_FOUND | A slug in companyUrls is not in the directory. |
RATE_LIMITED, BLOCKED | The directory pushed back. Ask for less in one run. |
NETWORK | A request could not be completed. Re-run it. |
TIME_BUDGET | The run ran out of time before that slice was read. |
PROXY_INPUT_ADJUSTED | Something in your proxyUrls was not usable and was adjusted. |
▶️ How to run it
1. Open Y Combinator Companies Scraper and click Try for free. 2. Put batches in Batches, or terms in Search terms, or slugs in Company profile URLs. 3. Narrow it with Industries, Regions, Tags or Company status if you want. 4. Set Maximum companies, keeping it low on the first run. 5. Click Start, then download the dataset as JSON, CSV or Excel, or read it from the Apify API.
💰 How much does it cost?
$0.90 per 1,000 companies. The same rate on every Apify plan, with no volume tiers.
One charge per company row. The same company returned by two of your filters is only charged once. Sample rows and diagnostic rows are never charged, and neither are companies a stage filter leaves out. Reading each company's profile page for the founders does not change the rate: it is the same per company with includeFounders on or off.
💡 What people use it for
- Pulling a whole batch the week it is announced, for a newsletter or a tracker.
- Building a founder outreach list, since
founderscarries LinkedIn and X links directly. - Filtering the directory for a thesis: a region, an industry and
Activestatus together. - Checking which YC companies in a space are hiring, then pairing it with a jobs scraper.
- Refreshing a stale list, using
formerNamesto catch companies that renamed.
🚧 What it does not do
- Companies only. No YC jobs, no Launch YC posts, no news, no funding rounds.
- No founder emails, and no contact details of any kind. Names and social links are what the
directory publishes.
yearFounded,city,countryand the social links come from the profile page, so they are
null when includeFounders is off.
- A profile page that fails to load still ships its row, and still bills, with
foundersempty
and those profile-only fields null. Nothing on the row separates "could not read" from "not published", so re-run any slug that comes back suspiciously bare.
- One filter with no batches tops out at 1,000 matches, because that is as far as the
directory's own search will page. Split the request by batch to get past it.
- 60 values per filter list, 500 URLs, 7,000 rows in a run. Anything past that is dropped.
- Team size and status are whatever the directory says today. They are not audited and they go
stale when a company does not update its own record.
🧭 Which company scraper do you need?
| If you want | Use |
|---|---|
| The Y Combinator startup directory | This one |
| New product launches and their makers | Product Hunt Scraper |
| Company pages on LinkedIn | LinkedIn Companies Scraper |
| Newly incorporated UK companies | Companies House New Companies Scraper |
| Who a startup is hiring | Wellfound Jobs Scraper |
❓ Questions people ask
Do I need a YC login or an API key? No. Nothing to paste in and nothing to renew.
How do I get one whole batch? Put the batch code in batches and set maxItems above the batch size. A recent batch runs to a few hundred companies.
Can I get the entire directory? Tick allCompanies and raise maxItems. It walks batch by batch, newest first, and you pay per company, so decide the number before you start it.
Why did one filter stop at 1,000? Because the directory's search pages that far and no further. Add batches and the run splits the query per batch, which gets around it.
What is query for? It names the batch, term or filter each row came from, which is the only way to attribute results when you ran several searches in one go.
Is this legal? The directory and the profile pages are public. Rows carry founders' names and links, which data-protection law covers, so have a reason for collecting them. Apify's write-up on scraping and the law is a fair start, and we are not lawyers.
🆘 If something breaks
Open the Issues tab on the actor page. Send the run ID and the input you used. A diagnostic row's errorCode and the query field on it usually point straight at the filter that went wrong.