companies of all sizes














































Extracting product data from PDF and Excel price lists
Wholesale catalogs are often published as PDF or Excel files, not web pages. Here's how Scrapliy turns those into tracked products.
Updated 2026-05-01 · Wholesale & Dropshipping · 7 min read
Step-by-step walkthrough
How PDF catalogs are parsed
Every line of extracted PDF text is checked for a detected price ($, €, £, ₺, or a currency code)—that's the gate against titles, footers, and page numbers, none of which carry one. SKU codes, quantities, and the product name are recovered from what's left on the line.
How Excel/CSV catalogs are parsed
Header names are matched against common conventions (Product Name/Title, Price/Wholesale Price/Cost, SKU/Item Code, Qty/Stock) and both a name and a price column need to be identified before any rows are ingested.
Where it goes
Parsed rows resolve into the same Entity Engine as every other source, so a supplier's PDF price list shows up in price history and cross-source conflicts exactly like a crawled page or merchant feed.
Ready to set this up for your store?
Add a site, pick a crawl schedule, and Scrapliy handles discovery, extraction, and change detection from there.
Related: Wholesale & Dropshipping → · Compare tools · Blog · All guides