Neighborhood assortment data: why averages mislead
Walk two independent stores a mile apart and read the shelves. Same format, similar square footage, noticeably different products. Neighborhood assortment data captures what averages erase: real stores stock what their shoppers actually ask for, so the "average store" described in a market report is a store that doesn't exist. Averages summarize. They don't describe.
Why the average store is a fiction
Try a two-store market as a thought experiment. One store stocks a specialty beverage; its neighbor doesn't. The market average reports the item in half of stores, a statement that's true of neither shelf. Scale that up across hundreds of items and thousands of stores, and the average assortment becomes a blend that no single store carries.
Plan against that blend and you're stocking for a shopper who doesn't exist. Sales projections inherit the same problem: an item's average performance across all doors mixes stores where it sells briskly with stores that never carried it at all.
The average isn't wrong, exactly. It's an answer to a different question, something closer to "what does the whole market look like from very far away." Useful for sizing. Useless for describing any particular shelf.
Distribution math makes the fiction concrete. An item selling well in the subset of stores that carry it can look mediocre when averaged across a market full of stores that never stocked it. The average blends a performance question with a distribution question and answers neither cleanly.
How is assortment actually decided in independent stores?
By the person closest to the demand. An independent owner hears requests at the counter, watches what sells out by Friday, and sees what gathers dust. Distributor catalogs, cooler space, and counter real estate set the constraints; daily observation sets the choices. Assortment in these stores is accumulated local knowledge, adjusted continuously.
That's skilled curation, and it deserves to be described that way. For a brand, store-level scan data is the closest available view of that curation at scale: thousands of independent decisions, visible as what's actually stocked and selling, door by door.
The turnover runs faster than outsiders expect. A request at the counter can become an order by the next distributor visit, and an item that stops moving can lose its shelf within weeks. That responsiveness is what makes store-level data valuable: the shelves are a running record of decisions made close to demand.
What's the responsible way to analyze neighborhood variation?
Work at store level first and aggregate second. Measure an item's rate of sale where it's actually stocked, sometimes called velocity where distributed, before averaging across doors that never carried it. Let clusters of stores with similar observed assortment and sales patterns emerge from the data itself.
The line to hold: describe shelves and transactions, not people. Skip demographic inference entirely. Assuming what "a neighborhood like that" buys substitutes stereotype for observation, and it's less accurate than simply reading what each store sells. Observed demand beats inferred demand every time.
Timing matters as much as geography in this work. Assortment patterns shift as stores respond to their shoppers, so a cluster built on last year's data describes last year's shelves. Re-run the clustering on a regular cadence and treat archetypes as living summaries rather than fixed labels. None of this is exotic; it's just resisting the urge to summarize before you've observed.
Published analysis should also aggregate for privacy, reporting patterns across groups of stores rather than exposing any single door. That's standard practice at NRS Insights, where channel-level reporting is built from store-level records that stay confidential.
Frequently asked questions
What does "velocity where distributed" mean?
It's an item's rate of sale counted only in the stores that actually carry it, instead of averaged across every store in a market. It answers how the product performs where it exists, which is the fair test before deciding whether it has earned space on more shelves.
Why not use demographic data to predict assortment?
It substitutes assumptions for observation, and assumptions about neighborhoods age badly and can drift into stereotype. Store-level scan data shows what shoppers in each store actually buy, which is more accurate and more respectful. The shelf and the register are better evidence than any inference.
How do brands act on assortment variation without managing every door?
Most teams cluster stores by observed assortment and velocity patterns, build a handful of store archetypes, and tailor plans by archetype rather than door by door. The archetypes come from sales behavior in the data, not from assumptions about who lives nearby.
Averages have their place; the report archive shows how channel-level reads and store-level truth can coexist in one methodology.