Request a tool
All toolsAutomationsGuidesMCP serverRequest a toolPlatformsCategories
Y Combinator Companies Scraper icon

Y Combinator Companies Scraper

Scrape the YC startup directory by batch, industry, region, tag or status. Get the pitch, team size, founded year, website and founders. $0.0009 each.

122 runs on Apify $0.0009 per company ($0.9 / 1,000)
Run this in the cloudRun on Apify →

Developer & Research Tools

How it works

  1. 1
    Open it on Apify

    Hit Run on Apify — it opens the tool in the cloud, no install.

  2. 2
    Set the inputs

    Adjust batches, industries, regions (sensible defaults are pre-filled).

  3. 3
    Click Run

    The tool runs on Apify’s cloud and collects the data for you.

  4. 4
    Export the results

    Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.

Pricing

$0.0009 per company = $0.9 per 1,000

You are charged forWhenPrice
Company scrapedEach Y Combinator company written to the dataset. Sample and diagnostic rows are never charged.$0.0009

Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-08-16, and they are what you are actually charged.

Inputs

FieldWhat it doesType
batchesYC batches to pull, written either short (W24, S23, F24, P25) or long (Winter 2024, Summer 2023). Leave empty to search across every batch. Each batch is queried separately, so ten batches and 100 rows gives you ten companies from each.array
industriesFilter by the industry labels the directory itself uses, for example B2B, Consumer, Healthcare, Fintech, Industrials, Education, Infrastructure, Marketing, Security, Climate. Several values are combined with OR.array
regionsFilter by the region labels the directory uses, for example United States of America, Europe, India, Latin America, Remote, Canada, Africa, Southeast Asia. Several values are combined with OR.array
tagsFilter by the free-form tags on a company profile, for example Artificial Intelligence, SaaS, Developer Tools, Marketplace, Generative AI, Open Source, Health Tech. Several values are combined with OR.array
statusesKeep only companies with these statuses. Leave empty for all four.array
stagesKeep only companies at these stages, as the directory labels them. Leave empty for both. The directory's search can't filter on stage, so the run reads its results and keeps the companies that match; the others are never charged and their profile pages aren't read. A stage filter can return fewer companies than you asked for: with one set, a run reads at most ten times Maximum companies from the directory, never fewer than 1,000 and never more than 10,000. It applies to searches, not to companies you name in Company profile URLs.array
searchTermsFree-text search across company names, pitches and descriptions, for example "robotics", "climate", "payments in Africa". Each term is searched separately and can be combined with the filters above.array
companyUrlsLook up specific companies instead of searching. Paste full profile URLs (https://www.ycombinator.com/companies/doordash) or bare slugs (doordash). Up to 500 per run.array
allCompaniesSweep the entire directory instead of filtering. Combine it with a high row limit; the run walks batch by batch, newest first. Remember you pay per company.boolean
isHiringKeep only companies flagged as hiring on their directory card.boolean
topCompaniesOnlyKeep only the companies the directory marks as top companies.boolean
nonprofitOnlyKeep only the non-profit organisations in the directory.boolean
includeFoundersOn by default. Reads each company's own profile page as well, which is what fills in founder names and titles, the year the company was founded, its city and country, and its LinkedIn, X, Facebook, Crunchbase and GitHub links. Turn it off for a faster, lighter run when you only need the directory card. The price per company is the same either way.boolean
maxItemsTotal number of companies to return in this run, shared evenly across your batches and search terms. Default 20, hard ceiling 7,000 (the directory holds about 6,200). Keep it low while you are testing - you pay per company.integer
proxyUrlsLeave this empty for a normal run. Fill it in only if you want the traffic to leave through proxy servers you already pay for, one URL per line, in the form http://user:pass@host:port.array

What you get

A structured dataset — each result includes fields like:

nameoneLinerbatchstatusteamSizeyearFoundedlocationwebsiteycProfileUrlfounderNamesindustrytagslongDescription

Export every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.

1 ready-to-run use cases

Y Combinator Companies That Are Hiring

List YC companies that are hiring, by batch: name, description, location, website and founders.

Related tools in Developer & Research Tools

Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.

GitHub Scraper iconDeveloper & Research Tools

GitHub Scraper

Search GitHub repos and users: stars, forks, language, topics, licence, plus user bio, company and followers. No token needed. $0.90 per 1,000 rows.

18 use cases

Stack Overflow / Stack Exchange Scraper iconDeveloper & Research Tools

Stack Overflow / Stack Exchange Scraper

Search Stack Overflow and Stack Exchange by keyword or tag. Score, answer count, views, reputation and body text. $2 per 1,000 questions.

2 use cases

Package Registry Scraper (npm + PyPI) iconDeveloper & Research Tools

Package Registry Scraper (npm + PyPI)

Get npm and PyPI package metadata as JSON. Version, license, author, repo, keywords and npm monthly downloads. $2 per 1,000 packages.

2 use cases

arXiv Scraper iconDeveloper & Research Tools

arXiv Scraper

Search arXiv papers by title, author, abstract or category. Get full abstracts, authors, categories, DOI, dates and PDF links. $2 per 1,000 papers.

2 use cases

OpenAlex Scholarly Works Scraper iconDeveloper & Research Tools

OpenAlex Scholarly Works Scraper

Search 250M+ OpenAlex papers with no API key. Get titles, authors, venue, year, citations, DOI, OA links and full abstracts. $2.00 per 1,000 papers.

2 use cases

Crossref Scholarly Works Scraper iconDeveloper & Research Tools

Crossref Scholarly Works Scraper

Search 150M+ papers on Crossref: DOI, title, authors, journal, publisher, date, citations and abstract. No API key. $1.00 per 1,000 works.

2 use cases

See all Developer & Research Tools →

Y Combinator Companies Scraper: search the YC startup directory by batch, industry or tag

Pick a batch like W24, an industry, a region, a tag or a status, or paste company profile URLs. You get one row per company: the pitch, the long description, batch, status, team size, founded year, location, website, tags and the founders with their LinkedIn and X links.

The awkward bit is the founded year. It comes off the company's own profile page, and a fair number of profiles simply do not carry one, so yearFounded arrives null rather than estimated.

InputBatches, industries, regions, tags, statuses, stages, search terms or profile URLs
OutputOne row per company
Ceiling7,000 companies per run
Account neededNone
Price$0.90 per 1,000 companies, flat on every plan

🏢 What Y Combinator Companies Scraper does

The YC directory is a searchable index of every company that has been through the programme, and each company also has its own profile page. This reads both and joins them into one row.

The directory card gives you the name, one-liner, batch, status, team size, industries, regions, tags, logo and website. The profile page adds the founders with their titles and bios, the year the company was founded, the city and country, and links to LinkedIn, X, Facebook, Crunchbase and GitHub. Every row has the same keys whichever way you searched.

Filters combine the way you would expect. Values inside one list are OR'd together, and different lists narrow each other. Ask for W24 and S23 with industries: ["B2B"] and you get B2B companies from either batch.

📥 What you give it

Every field is optional. Run it with the input empty and you get one labelled sample row back.

{
  "batches": ["W24", "S23"],
  "industries": ["B2B"],
  "statuses": ["Active"],
  "maxItems": 200
}
FieldDefaultWhat it is
batchesnoneShort (W24, S23, F24, P25) or long (Winter 2024). Up to 60. Each is queried separately, so ten batches and 100 rows gives ten from each.
industriesnoneThe directory's own labels: B2B, Consumer, Healthcare, Fintech and so on. Up to 60.
regionsnoneThe directory's region labels, like United States of America, Europe, India, Remote. Up to 60.
tagsnoneThe free-form profile tags, like Generative AI, Marketplace, Open Source. Up to 60.
statusesall fourActive, Acquired, Public, Inactive.
stagesbothEarly, Growth. The directory's search can't filter on stage, so the run reads its results and keeps the companies that match, reading at most ten times maxItems entries a run (1,000 to 10,000). A rare stage can return fewer companies than maxItems, with a free READ_LIMIT row saying so. Searches only, not companyUrls.
searchTermsnoneFree text across names, pitches and descriptions. Each term is searched on its own.
companyUrlsnoneUp to 500 profile URLs or bare slugs, to look companies up instead of searching.
allCompaniesoffSweep the whole directory, batch by batch, newest first. Pair it with a high row limit.
isHiringnot appliedThree-state. See the note below.
topCompaniesOnlyoffOnly companies the directory marks as top companies.
nonprofitOnlyoffOnly the non-profits.
includeFounderstrueReads each profile page too, which is what fills in the founders, founded year, city, country and social links. Turn it off for a lighter run. The price is the same either way.
maxItems20Total rows for the run, 1 to 7,000, shared across your batches and terms.
proxyUrlsnoneLeave it empty for a normal run. It is there for callers who want traffic to leave through servers they already pay for, as http://user:pass@host:port.

isHiring has three states, not two. Leave it out and no hiring filter is applied. Set it true and you get companies flagged as hiring. Set it to false and you get only the companies that are not hiring, which is a legitimate query and almost never what somebody means. The Console writes a literal false into the input once you have ticked the box and unticked it again, so if a run comes back full of companies that are not hiring, remove the field rather than unticking it.

📤 What you get back

A real row, from run NctVtHlEJQFnIqU35:

{
  "ok": true, "charged": true, "recordType": "company", "query": "Winter 2024",
  "name": "Indemni", "slug": "indemni",
  "oneLiner": "Cargo Theft and Fraud Prevention Platform",
  "longDescription": "We are building a safer supply chain. Cargo Theft has been increasing yearly, ...",
  "batch": "Winter 2024", "batchCode": "W24", "status": "Active",
  "teamSize": 7, "yearFounded": 2024,
  "location": "San Francisco, CA, USA", "city": "San Francisco", "country": "US",
  "regions": ["United States of America", "America / Canada", "Remote", "Partly Remote"],
  "website": "http://www.indemni.com",
  "ycProfileUrl": "https://www.ycombinator.com/companies/indemni",
  "industry": "B2B", "industries": ["B2B", "Supply Chain and Logistics"],
  "subindustry": "B2B -> Supply Chain and Logistics",
  "tags": ["Identity", "Logistics", "Supply Chain", "Fraud Prevention", "Fraud Detection"],
  "founders": [
    { "name": "Omar Draz", "title": "Founder",
      "linkedinUrl": "https://linkedin.com/in/odraz",
      "twitterUrl": "https://twitter.com/oamdraz",
      "bio": "Ex-DoorDash, worked on Fraud, Growth and Logistics. Currently building!" }
  ],
  "founderNames": "Omar Draz", "founderCount": 1,
  "linkedinUrl": "https://www.linkedin.com/company/100487698/admin",
  "twitterUrl": null, "facebookUrl": null, "crunchbaseUrl": null, "githubUrl": null,
  "logoUrl": "https://bookface-images.s3.amazonaws.com/small_logos/a4f88ce8a0...png",
  "formerNames": ["Alacrity"],
  "isHiring": false, "nonprofit": false, "topCompany": false, "stage": "Early",
  "launchedAt": "2024-02-15T20:40:36.000Z", "companyId": "29533",
  "scrapedAt": "2026-09-21T01:52:29.035Z"
}

longDescription and logoUrl are cut short above. A real row carries both in full.

FieldWhat it is
companyIdThe directory's own id, as a string. Use it as your dedupe key.
queryWhich batch, term or filter produced this row, so you can tell where a result came from when you searched several things at once.
foundersObjects with name, title, linkedinUrl, twitterUrl and bio. founderNames is the same names flattened for a spreadsheet.
formerNamesNames the company has traded under before. Useful when your own list is out of date.
batch, batchCodeWinter 2024 and W24. Both are there so you do not have to convert.
stageThe directory's own label, Early or Growth. The stages filter reads it.
launchedAtWhen the directory record was created, in ISO UTC. Not the founding date.

🧾 Reading the output

Three kinds of row.

RowHow to spot itBilled
A companyrecordType: "company"yes
The sample row_sample: true, and only when the input named no filter and no URLno
A diagnostic_diagnostic: true and ok: falseno

charged is set when the row is built, a moment before the charge goes out, so read it as "this is a real row" rather than as a receipt. A failed charge adds a CHARGE_ERROR diagnostic at the end.

The Console's default table hides the diagnostic columns, so those rows look blank there. Read them from the JSON.

CodeWhat it means
NO_RESULTSThe filter ran and matched nothing, or nothing at the stage you picked. An unrecognised batch string does this.
READ_LIMITWith stages set, the run read as many directory entries as it reads for one and stopped with what matched.
BAD_INPUTA value in stages is not Early or Growth. Nothing was searched.
NOT_FOUNDA slug in companyUrls is not in the directory.
RATE_LIMITED, BLOCKEDThe directory pushed back. Ask for less in one run.
NETWORKA request could not be completed. Re-run it.
TIME_BUDGETThe run ran out of time before that slice was read.
PROXY_INPUT_ADJUSTEDSomething in your proxyUrls was not usable and was adjusted.

▶️ How to run it

1. Open Y Combinator Companies Scraper and click Try for free. 2. Put batches in Batches, or terms in Search terms, or slugs in Company profile URLs. 3. Narrow it with Industries, Regions, Tags or Company status if you want. 4. Set Maximum companies, keeping it low on the first run. 5. Click Start, then download the dataset as JSON, CSV or Excel, or read it from the Apify API.

💰 How much does it cost?

$0.90 per 1,000 companies. The same rate on every Apify plan, with no volume tiers.

One charge per company row. The same company returned by two of your filters is only charged once. Sample rows and diagnostic rows are never charged, and neither are companies a stage filter leaves out. Reading each company's profile page for the founders does not change the rate: it is the same per company with includeFounders on or off.

💡 What people use it for

  • Pulling a whole batch the week it is announced, for a newsletter or a tracker.
  • Building a founder outreach list, since founders carries LinkedIn and X links directly.
  • Filtering the directory for a thesis: a region, an industry and Active status together.
  • Checking which YC companies in a space are hiring, then pairing it with a jobs scraper.
  • Refreshing a stale list, using formerNames to catch companies that renamed.

🚧 What it does not do

  • Companies only. No YC jobs, no Launch YC posts, no news, no funding rounds.
  • No founder emails, and no contact details of any kind. Names and social links are what the

directory publishes.

  • yearFounded, city, country and the social links come from the profile page, so they are

null when includeFounders is off.

  • A profile page that fails to load still ships its row, and still bills, with founders empty

and those profile-only fields null. Nothing on the row separates "could not read" from "not published", so re-run any slug that comes back suspiciously bare.

  • One filter with no batches tops out at 1,000 matches, because that is as far as the

directory's own search will page. Split the request by batch to get past it.

  • 60 values per filter list, 500 URLs, 7,000 rows in a run. Anything past that is dropped.
  • Team size and status are whatever the directory says today. They are not audited and they go

stale when a company does not update its own record.

🧭 Which company scraper do you need?

If you wantUse
The Y Combinator startup directoryThis one
New product launches and their makersProduct Hunt Scraper
Company pages on LinkedInLinkedIn Companies Scraper
Newly incorporated UK companiesCompanies House New Companies Scraper
Who a startup is hiringWellfound Jobs Scraper

❓ Questions people ask

Do I need a YC login or an API key? No. Nothing to paste in and nothing to renew.

How do I get one whole batch? Put the batch code in batches and set maxItems above the batch size. A recent batch runs to a few hundred companies.

Can I get the entire directory? Tick allCompanies and raise maxItems. It walks batch by batch, newest first, and you pay per company, so decide the number before you start it.

Why did one filter stop at 1,000? Because the directory's search pages that far and no further. Add batches and the run splits the query per batch, which gets around it.

What is query for? It names the batch, term or filter each row came from, which is the only way to attribute results when you ran several searches in one go.

Is this legal? The directory and the profile pages are public. Rows carry founders' names and links, which data-protection law covers, so have a reason for collecting them. Apify's write-up on scraping and the law is a fair start, and we are not lawyers.

🆘 If something breaks

Open the Issues tab on the actor page. Send the run ID and the input you used. A diagnostic row's errorCode and the query field on it usually point straight at the filter that went wrong.