The breadth of the catalog required parallel acquisition and a storage strategy that could work with millions of listings without turning every item into a separate read or write.
02 / How I approached it
How I approached it
Parallel parsers combined account and proxy management, throttling and retry strategies across the category scan. I organized normalization and Elasticsearch processing around bulk writes and bulk reads, keeping storage work batched as the collected dataset grew. The resulting data supported analysis of changes and trends.
System flow
Collect
Normalize
Index in batches
Analyze changes
What that means in practice
Architecture & decisions
01
Historical category-wide scope
The project scanned Amazon categories across the catalog, with books deliberately excluded. Its historical workload was on the order of millions of listings. That breadth made acquisition, normalization and storage design one connected problem: category data had to become a consistent dataset that could be analysed in batches.
02
Acquisition under external constraints
The parsers had to account for blocking, multiple accounts, proxies and failed requests. Parallel processing, throttling and retries made these conditions part of the acquisition design. Normalization gave downstream indexing and analysis a consistent representation despite differences in the collected source data.
03
Bulk on both sides of storage
Elasticsearch was used for both bulk writes and bulk reads. Batching mattered on both sides: collected records were indexed together and analysis consumed groups of records, avoiding a storage workload dominated by individual requests. This decision shaped the data path as much as the choice of search engine itself.
04
Trend analysis and a bounded product lifetime
The goal of collection was to identify changes and trends in product data. The service operated for roughly five months before priorities moved to other work. The archived project brings together the full engineering chain: parallel acquisition, normalization, bulk storage access and analysis of a large historical dataset.
03 / The outcome
The outcome
The service ran for about five months at a historical scale on the order of millions of listings. It supported category-wide trend analysis before the team changed focus and the project was archived.