Methodology & Pipeline
A comprehensive overview of the data ingestion, entity extraction, forecasting metrics, and academic bibliography underpinning the FairTrade AI project.
01. Executive Overview
Shrinkflation represents a modern hidden inflation mechanism where consumer packaging volume or net quantity is downsized while the product's Maximum Retail Price (MRP) remains unchanged. Standard economic measurement structures, such as India's Consumer Price Index (CPI), track the nominal cost of commodities but fail to account for weight contraction. As a result, standard indexes completely overlook a silent cost expansion.
FairTrade AI bridges this information deficit by introducing an automated longitudinal audit framework. By comparing years of archived web snapshot timelines against active e-commerce platforms, the pipeline reconstructs historical product registries and exposes hidden quantity shifts.
02. Data Ingestion Pipeline
To reconstruct product weight timelines spanning 2021 to 2026, the data acquisition layer queries the Wayback Machine CDX API for historical product URL captures across Amazon.in, BigBasket, Blinkit, and Flipkart.
- Query index URLs to locate all snapshots logged over 5 years.
- Filter snapshots to isolate stable structural captures (removing bad redirects, blank screens, and HTTP error pages).
- Methodically render snapshots using headless browsers (Playwright/Selenium) to parse late-loaded javascript DOM assets.
03. Adaptive Inference Layer
Reconstructed e-commerce pages present "temporal noise" because website code designs shift over five-year timelines. Traditional regex or strict CSS selector selectors fail when rendering old HTML layouts.
FairTrade AI implements an Adaptive Extraction layer. The system first crawls standard semantic data objects (JSON-LD schemas, markup meta-tags). If these sources are missing or incomplete, the engine calls a Google Gemini 2.5 Flash fallback loop, passing unstructured raw HTML segments. The LLM extracts the true weight and MRP entities based on contextual descriptors, regardless of web structural changes.
04. Mathematical Metrics
A. Shrinkflation Score Formula
Calculated by scaling the physical weight reduction ratio against a severity constant of 500:
* A package contraction of 10% translates to a score of 50 (Medium Risk threshold). Contractions of 20% or more clamp to the maximum score of 100 (High Risk worst offender).
B. Quantity-Adjusted Purchasing Power (QAPP)
Represents the true unit value paid per 100 units of physical size:
C. ML Trend Forecasts
A linear regression least-squares check maps weight timelines to years, calculating package drift slopes. The trendline projects future weights 24 months forward:
05. Caveats & Limitations
Users auditing timeline outputs should evaluate these statistical limitations:
- Archival Gaps: The Wayback Machine does not log daily page snapshots. Gaps in timestamps may delay identifying the exact date of package downsizes.
- Inference Confidence: Gemini extraction queries carry a 0.0 to 1.0 confidence score based on markup clarity. Low-confidence parsed data points are excluded from production timelines.
- Market Speculation: Projections represent statistical regression extrapolations and do not forecast future retail raw material costs.
06. Bibliography & References
Kumar, A. (2022). "Shrinkflation: A Kind of Hidden Inflation." Journal of Advances and Scholarly Researches in Allied Education, Vol. 19, Issue 4.
Oxera Agenda. (2016). "Shrinkflation! A bite missing?" Oxera Compelling Economics.
Kuligin, L. (2026). "Layout-Aware Text Extraction Using Heuristic Segmentation and LLM-Based Refinement." Defensive Publications Series.
Etzioni, O. et al. (2008). "Open Information Extraction from the Web." Communications of the ACM, Vol. 51, No. 12.
Baeza-Yates, R. et al. (2011). "The 1st Temporal Web Analytics Workshop." Proceedings of WWW 2011.