Request a tool
All toolsAutomationsGuidesMCP serverRequest a toolPlatformsCategories
Shein Catalog Scraper icon

Shein Catalog Scraper

Scrape Shein's public product catalogue: product IDs, full titles, product and image URLs and last-changed dates, across thirteen storefronts. No prices.

54 runs on Apify $0.0002 per catalogue entry ($0.2 / 1,000)
Run this in the cloudRun on Apify →

Developer & Research Tools

How it works

  1. 1
    Open it on Apify

    Hit Run on Apify — it opens the tool in the cloud, no install.

  2. 2
    Set the inputs

    Adjust region, maxItems, contentType (sensible defaults are pre-filled).

  3. 3
    Click Run

    The tool runs on Apify’s cloud and collects the data for you.

  4. 4
    Export the results

    Download as JSON, CSV or Excel, or pipe straight into your app, Google Sheets, or an AI agent.

Pricing

$0.0002 per catalogue entry = $0.2 per 1,000

You are charged forWhenPrice
Catalogue entryOne product entry from Shein's public catalogue: id, full title, URL, image and the date the listing was last stamped. Prices are not included. Sample and diagnostic rows are free.$0.0002

Pay-per-event pricing: you are billed per result, not per subscription. Billing is handled by Apify on your own account. These are the live Apify store prices, in effect since 2026-09-20, and they are what you are actually charged.

Inputs

FieldWhat it doesType
regionWhich Shein storefront to read. These thirteen publish a catalogue that can be read; the other nineteen storefronts checked do not, so there is nothing to add to this list. Each storefront keeps its own ids, its own URLs, and titles in its own language.string
maxItemsHow many rows to return. The run stops the moment it has this many and cancels the download in the middle of the file, so asking for 200 genuinely transfers about 200 rows' worth of bytes rather than the whole 12,000-entry file. Keep it low while you are testing - you pay per row.integer
contentTypeProducts is the catalogue itself, and the only one with titles and images. Category pages and store pages are the other two families of URL in the same index: far smaller, and carrying only a URL and a date, so their title is derived from the URL and their image is empty.string
titleContainsCase-insensitive. A row is kept if its title or its URL contains any one of these. Up to 20 terms. Leave it empty to take everything.array
updatedSinceA date as YYYY-MM-DD. READ THIS BEFORE YOU SET IT: Shein re-stamps the whole catalogue at once, not each product when it changes. Measured on 2026-09-20: two files read end to end were 12,000 of 12,000 on a single date, all 440 files in the index carried that date, and samples from 24 files across the whole catalogue matched. So this is a gate on when Shein last regenerated the catalogue, not a list of what changed - set it to yesterday and you may get everything, set it to tomorrow and you will get nothing. To find genuine changes, pull the ids and diff them against your last pull. If a run drops every row to this filter it tells you which dates it actually saw.string
maxSitemapFilesShein splits the catalogue into files of 12,000 entries each. This caps how many one run opens, which is the real bound on how long a narrow filter can hunt before giving up, and on what a fruitless hunt costs. With no filter set one file is 12,000 rows, so nine files cover the 100,000-row limit. To look deeper than twenty files, move the start file forward and run again. The run stops at whichever it reaches first, this or the row limit, and the log says which.integer
startFileWhere in the catalogue to begin. The files do not overlap - file 1 and file 221 were checked and share none of their 12,000 ids - so this is how you walk the whole catalogue across several runs without repeating yourself. Note the file your last row came from and start the next run there. A number past the end is refused before anything is downloaded, and the message tells you how many files that storefront actually has today.integer
proxyUrlsLeave this empty for a normal run. Fill it in only if you want the traffic to leave through proxy servers you already pay for, one URL per line, in the form http://user:pass@host:port.array

What you get

A structured dataset — each result includes fields like:

regioncontentTypeproductIdtitleurlimageUrllastmodchangefreqpricesourceSitemapscrapedAt

Export every run as JSON, CSV or Excel, or send it to your app, a database, Google Sheets, or an AI agent.

Related tools in Developer & Research Tools

Other ready-to-run tools in the same category — all pay-per-use on the Apify cloud.

GitHub Scraper iconDeveloper & Research Tools

GitHub Scraper

Search GitHub repos and users: stars, forks, language, topics, licence, plus user bio, company and followers. No token needed. $0.90 per 1,000 rows.

18 use cases

Stack Overflow / Stack Exchange Scraper iconDeveloper & Research Tools

Stack Overflow / Stack Exchange Scraper

Search Stack Overflow and Stack Exchange by keyword or tag. Score, answer count, views, reputation and body text. $2 per 1,000 questions.

2 use cases

Package Registry Scraper (npm + PyPI) iconDeveloper & Research Tools

Package Registry Scraper (npm + PyPI)

Get npm and PyPI package metadata as JSON. Version, license, author, repo, keywords and npm monthly downloads. $2 per 1,000 packages.

2 use cases

arXiv Scraper iconDeveloper & Research Tools

arXiv Scraper

Search arXiv papers by title, author, abstract or category. Get full abstracts, authors, categories, DOI, dates and PDF links. $2 per 1,000 papers.

2 use cases

OpenAlex Scholarly Works Scraper iconDeveloper & Research Tools

OpenAlex Scholarly Works Scraper

Search 250M+ OpenAlex papers with no API key. Get titles, authors, venue, year, citations, DOI, OA links and full abstracts. $2.00 per 1,000 papers.

2 use cases

Crossref Scholarly Works Scraper iconDeveloper & Research Tools

Crossref Scholarly Works Scraper

Search 150M+ papers on Crossref: DOI, title, authors, journal, publisher, date, citations and abstract. No API key. $1.00 per 1,000 works.

2 use cases

See all Developer & Research Tools →

Shein Catalog Scraper: product ids, titles and images for 13 storefronts, no prices

Pick a Shein storefront and you get its public catalogue as rows: product id, full title, product link, image link and Shein's own date stamp. Germany alone lists 5,271,135 products, and one run can take up to 100,000 of them. There are no prices, no stock and no sizes in it, on any storefront, and no setting that adds them.

InputA storefront, plus optional title words and a date
OutputOne row per catalogue entry
Ceiling100,000 rows or 20 catalogue files per run, whichever comes first
Account neededNone
Price$0.20 per 1,000 rows, flat on every plan

🔍 What Shein Catalog Scraper does

It reads the catalogue Shein publishes for each storefront and hands it over one entry per row. If you need Shein prices, this is the wrong purchase however cheap the rows are. If you need the list itself, every product id and title on a storefront or every listing whose title says "cargo", this is what it does.

Shein splits a storefront's catalogue into files of 12,000 entries. The files do not overlap (file 1 and file 221 were checked and share none of their ids), so startFile lets you walk the whole catalogue across several runs without repeating yourself. Every row names the file it came from.

The run stops the moment it has the rows you asked for, even halfway through a file. A 12,000-row run took nine seconds in testing.

Filters work on entries as they are read, because there is no search behind a catalogue file. A narrow filter therefore reads a lot and delivers a little. That costs time, not money, since you only pay for rows delivered, and maxSitemapFiles puts a limit on how far it hunts.

updatedSince is not a change feed. Shein re-stamps the whole catalogue at once, not each product when it changes. On 2026-09-20 two files read end to end were 12,000 of 12,000 on one date, and every file in the index carried it. So the date says when Shein last regenerated the catalogue. Set it to yesterday and you may get everything; set it to tomorrow and you get nothing. A run can also come back with today's date and yesterday's side by side, though inside one file the date is always the same. To see what really changed, pull the ids and compare them with your last pull.

📥 What you give it

{
  "region": "de",
  "maxItems": 200,
  "titleContains": ["cargo"],
  "maxSitemapFiles": 10
}
FieldIf you leave it outWhat it is
regiona free sample rowThe storefront: de, uk, it, es, pt, nl, pl, fr, ro, at, se, ch or tr. gb works for the UK. Each has its own ids, links and titles in its own language.
maxItems200Rows to return, 1 to 100,000. Keep it low while testing, since you pay per row.
contentTypeproductsproducts is the catalogue, and the only type with titles and images. categories and stores are category and store pages, carrying just a link and a date.
titleContainsno filterUp to 20 words. A row is kept if its title or link contains any one of them, ignoring case.
updatedSinceno filterA date as YYYY-MM-DD. Keeps entries stamped on or after it. Read the warning above first.
maxSitemapFiles10How many catalogue files one run may open, 1 to 20. With no filter, nine files cover the 100,000-row limit.
startFile1Where in the catalogue to begin, 1 to 500. A number past the end is refused before anything is read, and the note says how many files that storefront has today.
proxyUrlsnot usedYour own servers, one per line as http://user:pass@host:port, if you want the run to go out through them. Leave it empty otherwise.

📤 What you get back

A real row from a real run on the German storefront:

{
  "ok": true,
  "recordType": "catalogEntry",
  "region": "de",
  "contentType": "products",
  "productId": "87740328",
  "title": "Suitcases",
  "url": "https://de.shein.com/Suitcases-p-87740328.html",
  "imageUrl": "https://img.ltwebstatic.com/v4/j/spmp/2025/05/23/c8/174797332031d410969cc879219609176ffc61bd82_square.jpg",
  "lastmod": "2026-09-29",
  "changefreq": "daily",
  "sourceSitemap": "sitemap-products-1.xml",
  "scrapedAt": "2026-09-30T02:32:29.470Z",
  "price": null
}
FieldWhat it is
productIdShein's id, taken from the link. Null on category and store rows.
titleThe full title as Shein publishes it. On category and store rows it is read off the link, because Shein gives those no title.
url, imageUrlThe product page and one image. imageUrl is empty on category and store rows.
lastmodShein's date stamp, YYYY-MM-DD. One date for the whole catalogue, not a per-product change date.
changefreqShein's own hint about how often the entry changes.
sourceSitemapThe catalogue file the row came from, so you know where to resume.
priceAlways null. It is there so nobody has to wonder whether a missing key means they did something wrong.
region, contentType, scrapedAtThe storefront, the type of entry, and when the run read it.

🧾 Reading the output

RowHow to spot itCharged
A catalogue entryrecordType: "catalogEntry"yes
The sample_sample: true, recordType: "sample"no
A note_diagnostic: true, recordType: "diagnostic", with an errorCodeno

To keep only real rows, filter on recordType equal to catalogEntry. Filtering on ok: true is not enough, because the sample row carries it too. The sample appears only when a run has no storefront chosen, and shows the shape of a row rather than live data.

Each note carries an errorCode, a short error and a details line saying what happened.

errorCodeWhat it means
BAD_INPUTA storefront this cannot read, or a startFile past the end. The note lists what is available.
NO_RESULTSNothing to deliver. Your filters matched nothing (the note lists the catalogue's dates), the storefront has no files of that type today, or a file was empty.
NOT_FOUNDThat storefront's catalogue was not there on this run.
BLOCKED, RATE_LIMITED, NETWORKA catalogue file, or the storefront's index, did not come back or stopped part way. Rows already returned stay in the dataset, and after a failed file the run moves on to the next one.
UNEXPECTED_ERRORSomething broke. Rows already written are kept.
PROXY_INPUT_ADJUSTEDA connection setting sent through the API is not offered here and was replaced.

▶️ How to run it

1. Open Shein Catalog Scraper and click Try for free. 2. Choose a Storefront. 3. Set Maximum rows. Start small, 200 is plenty for a first look. 4. To narrow it down, add words under Only titles containing one of these words. 5. Click Start, then download the dataset as JSON, CSV or Excel.

💰 How much does it cost?

$0.20 per 1,000 rows, flat on every Apify plan, with no volume tiers. One catalogue file's worth of rows, 12,000 of them, comes to $2.40.

Not charged: the sample row, every note, rows your filters drop, and any file that fails or comes back empty. A run that matches nothing charges for no rows at all. If you set a maximum cost for the run, it stops there and says so, rather than going over.

💡 What people use it for

  • Keeping a list of every product id on a storefront and comparing it week to week. Ids that appear

are new listings, ids that vanish are gone.

  • Pulling every listing whose title mentions a word, like cargo or linen, across a storefront.
  • Counting how often a style, material or fit turns up in product titles.
  • Collecting product links and images for a set of listings you then check by hand.

🚧 What it does not do

  • No prices. Not on any row, any setting or any storefront.
  • No stock, sizes, colours, ratings, reviews or seller data either. The catalogue says what

exists, not what it costs or whether it is in stock.

  • Thirteen storefronts, mostly European. There is no US, Mexican, Brazilian, Indian, Canadian,

Australian or Japanese storefront in the list, and no setting that adds one.

  • lastmod cannot tell you what changed. It is one date stamped across the whole catalogue.
  • No translation. Titles come in the storefront's own language.
  • One run is one storefront. Four storefronts means four runs.
  • An entry is not proof a product is for sale. A listing leaving the catalogue is the only sign

you get that it went away.

🧭 Which fashion and retail scraper do you need?

If you wantUse
Shein product ids, titles, links and imagesThis one
Zara products with prices, sizes and per-size stockZara Scraper
A Shopify store's whole catalogue, with prices and variantsShopify Products Scraper
TikTok Shop products with prices and units soldTikTok Shop Scraper
AliExpress products by search term, with price and ordersAliExpress Products Scraper

❓ Questions people ask

Can I get Shein prices from this? No. Not with a setting, a different storefront or your own servers. The catalogue carries no price.

How many products are there? Germany lists 440 files, 5,271,135 products. The UK has slightly more and Turkey the fewest, at 307 files. The 13 storefronts hold about 60 million between them.

How do I pull a whole storefront? Run it repeatedly, moving startFile forward. The sourceSitemap on your last row tells you how far the previous run got.

How do I find what changed since last week? Compare this week's ids with last week's. updatedSince cannot do it, for the reason above.

Do I need an account or an API key? No. The catalogue is public.

Is this legal? It reads a catalogue Shein publishes openly, with no login. Apify's write-up on the legality of web scraping is a good starting point. We are not lawyers.

🆘 If something breaks

Open the Issues tab on the actor page and send the input you used and the run ID. The note rows in the dataset usually name the file and the reason already.