Demo project
Catalog parser: a website into CSV and XLSX
Crawls a catalog by its robots.txt rules, collects product cards, exports CSV and XLSX.
development · Node.js · cheerio · ExcelJS

The task
“Get the data off this site into a spreadsheet” is one of the most common freelance requests. A client has a catalog - their own, a partner’s, a supplier’s - and needs it as a table they can actually work in: check prices, add a margin, build a price list, upload to a marketplace. Done by hand it is hours of copy-paste, plus the price typos that only surface after the file has been sent on.
This stand shows the whole mechanism end to end. The source is my own demo shop, PrintCraft, served locally as static files: the company is invented, the catalog structure is real. No third-party site is touched, and that is deliberate. A robots.txt on its own does not settle the legal side of scraping, but reading it and obeying it is the first thing you do on a paid job, and it is done here: the rules are parsed before the first request to the catalog.
The solution
A run is four steps: rules, crawl, extraction, export.
Before the first request the parser reads robots.txt. It picks the group that applies - matched on its own user agent, or the catch-all User-agent: * - then reads the disallow and allow rules and the Crawl-delay. When two rules collide the longer path wins, and on equal length allow wins. Blocked paths are skipped with the exact rule named in the log: in the run of 24 August one address was dropped by Disallow: /cart/. Between requests it waits, taking whichever is longer, the site’s delay or its own minimum. It never leaves the host it started on: links to other domains go into a counter and are never requested.
Each product page yields eight fields: name, category from the breadcrumbs, price, availability, material, available colors, a short description and the page path. A card with no name or no price does not reach the table - it is logged with the list of missing fields and the run carries on.
One run writes two files. catalog.csv is UTF-8 with a BOM and semicolons: Cyrillic survives intact, and where the system list separator is a semicolon Excel opens the file on a double click, no import wizard. For other locales the separator is a --sep flag. catalog.xlsx is a single catalog sheet with a frozen header row, an autofilter, set column widths and wrapped text in the description column.
Details that are easy to miss
- In XLSX the price is a number, not text. Format
# ##0, aligned right. The column sums and sorts straight away; otherwise the table would need cleaning by hand, which eats the time the export just saved. - Availability is not invented. The demo shop has no stock flag, so the parser judges by the state of the “Add to cart” button and by explicit wording on the page, and records what it judged by: button active, button disabled, page text, no order button on the page. In that last case the cell reads “no data”, not “out of stock”. You can see which rows to trust.
- A network failure does not kill the run. Responses 408, 425, 429 and 5xx, timeouts and dropped connections are retried: by default three attempts per address, which is two retries, waiting 400 ms then 800 ms. After that the refusal is logged and the crawl moves on to the next address. The attempt count and the base pause are the
--attemptsand--backoffflags. - Failures are tested against a live server, not a stub. A clean pass over the demo shop never produces a retry, so a separate script exercises the failure path: it starts a server that genuinely answers 503, genuinely drops the socket and genuinely stays silent past the timeout. The same script reads the written files back from disk: the BOM, a separator escaped inside a value, the cell type of the price, the frozen header.
- Every run leaves a trail.
journal.jsonlrecords line by line: the robots rules, every page with its status and response time, every product found, both exports and the totals. It stores file names only, never absolute paths, so the log can be handed to a client as it is.
The numbers
The run of 24 August, with the log sitting next to the code. 13 site pages crawled, 9 of them product cards, and all 9 reached the table. One address skipped by Disallow: /cart/, 0 retries, 0 unanswered requests. The run took 12.3 seconds, 12 of which were the wait required by Crawl-delay: 1; the requests themselves took 128 milliseconds. The ceiling here is not the code, it is politeness to the site.
Screenshots


