Back to workWork / Amazon Data / Trends

Amazon Data / Trends

Large-scale Amazon scraping and trend analysis

Data acquisition and processing architecture
01 / The challenge

The challenge

The breadth of the catalog required parallel acquisition and a storage strategy that could work with millions of listings without turning every item into a separate read or write.

02 / How I approached it

How I approached it

Parallel parsers combined account and proxy management, throttling and retry strategies across the category scan. I organized normalization and Elasticsearch processing around bulk writes and bulk reads, keeping storage work batched as the collected dataset grew. The resulting data supported analysis of changes and trends.

System flow
  1. Collect

  2. Normalize

  3. Index in batches

  4. Analyze changes

What that means in practice

Architecture & decisions

01

Historical category-wide scope

The project scanned Amazon categories across the catalog, with books deliberately excluded. Its historical workload was on the order of millions of listings. That breadth made acquisition, normalization and storage design one connected problem: category data had to become a consistent dataset that could be analysed in batches.

02

Acquisition under external constraints

The parsers had to account for blocking, multiple accounts, proxies and failed requests. Parallel processing, throttling and retries made these conditions part of the acquisition design. Normalization gave downstream indexing and analysis a consistent representation despite differences in the collected source data.

03

Bulk on both sides of storage

Elasticsearch was used for both bulk writes and bulk reads. Batching mattered on both sides: collected records were indexed together and analysis consumed groups of records, avoiding a storage workload dominated by individual requests. This decision shaped the data path as much as the choice of search engine itself.

04

Trend analysis and a bounded product lifetime

The goal of collection was to identify changes and trends in product data. The service operated for roughly five months before priorities moved to other work. The archived project brings together the full engineering chain: parallel acquisition, normalization, bulk storage access and analysis of a large historical dataset.

03 / The outcome

The outcome

The service ran for about five months at a historical scale on the order of millions of listings. It supported category-wide trend analysis before the team changed focus and the project was archived.

04 / Tools & technologies

Tools & technologies

Continue exploring

Liqsale / Daily Snipes

Contact

Complex challenge? Let’s make it work.

Architecture, technical leadership, or a product that needs a solid foundation.

Start a conversation

A role, a product, or a technical challenge. Share a little context.

To use the form, enable JavaScript. You can also write to me directly on Telegram.

Write on Telegram