The Agentic Category Manager: assortment, incrementality, and shelf-space decisions in FMCG
A retailer is preparing its annual yogurt review. The category holds 186 SKUs, growth has slowed, availability is inconsistent, store teams say the range is too complex, supply wants fewer products, the private-label team wants more space, and suppliers have proposed six innovations. The buyer asks for 20 delistings. The first analysis looks straightforward: rank every SKU by sales, remove the slowest, give the space to the fastest. The bottom 20 SKUs are only 4% of category sales — simplify the shelf, keep 96% of the revenue.
The ranking hides almost everything that matters. One of those slow SKUs is the only lactose-free family pack in the range. Another is the preferred product of a small, highly loyal shopper group. A third has weak national velocity and performs strongly in urban stores. A fourth transfers almost no demand to the rest of the assortment when it disappears. A fifth sells little because it is frequently out of stock — a supply failure being read as a demand signal. A sixth looks productive only because it holds six facings at eye level. Meanwhile, near the top of the ranking sits a SKU so substitutable that removing it would barely dent the category, and another that mostly cannibalizes a more profitable pack from the same manufacturer.
So the real question is not which products sell the least. It is which combination of products, attributes, brands, pack sizes, prices, and shelf positions best serves shopper needs while maximizing sustainable category value under limited physical and operational capacity. A shelf is a constrained commercial system: every added SKU consumes width, depth, inventory, working capital, replenishment effort, warehouse space, master-data maintenance, planogram complexity, and shopper attention — while every removed SKU risks lost category sales, lost loyalty, store switching, unmet needs, weaker perceived variety, supplier tension, and lost differentiation. Variety against simplicity, choice against clarity, incrementality against duplication, localization against operational scale, shelf productivity against shopper loyalty.
The greatest value of agentic AI in category management is not automatic assortment selection. It is a governed decision layer that connects shopper choice, SKU incrementality, demand transfer, shelf capacity, retailer and manufacturer economics, operational feasibility, and human judgment.
The objective is not to fit as many products onto a shelf as possible. Nor is it to remove as many as possible. It is to build the smallest assortment and clearest shelf that satisfies the relevant shopper needs, protects category loyalty, delivers operationally feasible availability, and produces sustainable value for the retailer and the category.
What category management really manages
Category management treats a group of related products as a strategic business unit rather than a collection of independently managed suppliers: define a shopper-relevant category, assign its role, assess performance and shopper structure, select strategies, then coordinate assortment, price, promotion, merchandising, space, distribution, and availability. An agentic system does not replace that process — it makes it more continuous, more evidence-driven, more granular, easier to localize and simulate, and far more connected to implementation and learning.
It begins with a boundary. A manufacturer defines “our yogurt portfolio”; the retailer may need “chilled cultured dairy”; the shopper is thinking “breakfast for my family” or “a high-protein snack.” Boundaries decide which products count as substitutes, complements, competitors, and adjacent opportunities — a narrow definition misses substitution, a broad one becomes operationally unusable. The agent must keep the operational category distinct from the shopper need space, because a shopper may move between Greek yogurt, protein pudding, cottage cheese, a breakfast drink, and a snack bar depending on the mission.
Then comes the category role, which governs everything downstream. A destination category helps shoppers choose the retailer and may justify greater variety, deeper segmentation, innovation, and regional differentiation. A routine category needs dependable range, competitive price, strong availability, and operational efficiency. Seasonal categories expand and contract. Convenience categories prioritize simplicity, core coverage, and fast choice. The same assortment logic should not be applied to all four — and the agent should never optimize an assortment without knowing which role it is optimizing for. Strategy (penetration, transaction, margin, loyalty, excitement, frequency, premiumization) precedes tactics (assortment, space, price, promotion, merchandising, availability); a tactic without a strategy is just local optimization.
This is also why category captaincy is delicate. A supplier appointed as category partner brings genuine expertise and analytical resources, and also influence over decisions affecting its competitors. EU vertical-agreement guidance recognizes that category-management agreements can create efficiencies while warning that they can disadvantage competing suppliers or facilitate inappropriate exchange of sensitive information. Any agent operating in that seat needs transparent objectives, category-level evidence, neutral item rules, traceable recommendations, retailer ownership of the decision, legal controls, and hard separation of sensitive information.
The shelf is a choice architecture
A shelf does not merely hold inventory; it structures decisions. Shoppers do not evaluate every SKU — they follow rules. One shopper goes need → format → brand → flavour → pack size → price. Another starts from a dietary requirement. Another starts from the promotion. The shopper decision tree describes the hierarchy of attributes that structures choice — need state, user, occasion, format, brand, benefit, flavour, ingredient, dietary requirement, pack size, price tier, sustainability, origin — and the order is what makes it actionable. If the first split in yogurt is dairy versus plant-based, removing the only plant-based family pack creates a structural gap. If the first split is adult versus children, that same pack plays a different role entirely. A sales ranking cannot see either.
Critically, the shopper tree is not the product hierarchy. A manufacturer thinks brand → sub-brand → flavour → pack; a shopper may think need → format → dietary requirement → price → brand. Behavioural trees are inferred from switching, baskets, loyalty data, substitution, purchase sequences, panel data, and attribute similarity; stated trees come from asking shoppers directly. Both are useful; they frequently disagree. And substitution is asymmetric: shoppers denied the premium organic SKU may accept the mainstream organic one, while shoppers denied the mainstream one may not trade up. The agent must preserve directionality rather than collapsing substitution into a symmetric similarity score.
From the tree comes need-state coverage— affordable family breakfast needs an entry family pack, a high-protein snack needs a protein single-serve, a children’s lunchbox needs a small multipack, a lactose-free household needs a lactose-free core pack. A need does not require many SKUs; it requires enough credible choice. And that is the line between meaningful variety (dairy, plant-based, lactose-free, high-protein) and duplication (five nearly identical strawberry 500 g mainstream yogurts). The goal of rationalization is to remove the second without touching the first. Note too that perceived variety is not SKU count — shoppers read variety through brands, formats, flavours, price tiers, pack sizes, benefits, and visual organization, and the research on assortment size is genuinely context-dependent: larger ranges can improve perceived choice in some settings and create cognitive effort, uncertainty, and choice avoidance in others. A shelf can hold exactly the right products and still underperform because it cannot be navigated.
Velocity is not incrementality
Assortment planning decides which products are listed, retained, expanded, localized, reduced, replaced, or delisted — at market, retailer, banner, format, cluster, store, and channel level. Macro assortment sets the broad range for a market or format; micro assortment tailors it to cluster, demographics, mission, space, region, climate, and competition. Formally it is a constrained optimization: choose a set of SKUs to maximize category value subject to shelf space, inventory, operations, shopper coverage, strategy, and supplier commitments. Which means category value must be defined explicitly — gross profit, contribution, sales, shopper retention, inventory return, basket contribution, or a stated combination.
The measurement ladder climbs away from raw sales. Sales rank merely records performance under current conditions — distribution, space, price, promotion, availability, location, store mix, listing age, marketing, seasonality. Velocity normalizes by exposure (units per store per week, revenue per facing, units per available store-week) and must itself be adjusted: for distribution (€10m across 1,000 stores may be weaker than €3m across 200), for availability (a product should never be delisted because the company failed to keep it on shelf), and for promotion (separate base from promotional velocity and full-price productivity). Margin velocity — gross profit per store-week, contribution per facing-week — and inventory productivity (GMROI, turn, days of supply, waste) complete the commercial picture. None of them answers the actual question.
The four quadrants each imply a different action. High velocity, high incrementality: protect, and consider more space. High velocity, low incrementality: still valuable, but the category can tolerate fewer facings, replacement, or a harder negotiation. Low velocity, high incrementality: the most misread cell — these serve a niche, protect loyalty, cover a dietary need, or prevent store switching, and they usually deserve targeted distribution rather than national delisting. Low velocity, low incrementality: the genuine rationalization candidates — after checking launch stage, supply issues, strategic commitments, seasonality, and data quality. A richer model adds margin, loyalty, space productivity, and operational complexity; the discipline is to keep the dimensions visible rather than collapsing them into one opaque score.
| Removed SKU | Own brand | Other brand | Private label | Category lost | Store lost |
|---|---|---|---|---|---|
| SKU A | 40% | 25% | 15% | 10% | 10% |
| SKU B | 10% | 15% | 20% | 15% | 40% |
| SKU C | 70% | 15% | 10% | 3% | 2% |
SKU B sells less than SKU A and matters more to the retailer, because 40% of its demand walks out of the store. Transfer estimates come from household panel switching, loyalty behaviour, out-of-stock events, past delistings, store variation, assortment experiments, choice models (multinomial, nested and mixed logit, latent class, hierarchical Bayesian, Markov substitution, ML demand models), surveys, attribute similarity, and basket data — and the language model should call those services, never reproduce them. Two cautions carry most of the risk. Out-of-stock substitution is not delisting transfer: shoppers who expect an item to return behave differently from shoppers who find it permanently gone. And similarity is not transfer: two strawberry yogurts at the same size and price can serve completely different shoppers when one is a children’s product and the other is high-protein.
Two effects sit outside the substitution frame. Complementarity and basket halo: pasta and sauce, coffee and creamer, razors and blades, pet food and treats — removing a product can reduce basket value well beyond its own sales, though correlation needs care, since popular items appear in large baskets partly because of who buys them. And loyalty: a low-sales SKU can carry high customer importance for high-value households, frequent shoppers, families, or dietary-need shoppers. Research on grocery assortment reduction indicates that retaining shoppers’ preferred items materially reduces store-switching risk, which makes exclusive-buyer rates, repeat intensity, and alternative acceptance real inputs to the decision. A product’s true value is direct margin plus retained category demand plus retained store trips plus basket contribution — hard to estimate precisely, and worse to ignore entirely.
Every SKU should have a role
- Traffic driver. Attracts shoppers or protects store choice.
- Core volume item. Dependable high velocity, the backbone of the range.
- Margin contributor. Generates disproportionate category profit.
- Need-state protector. Covers a specific shopper requirement nothing else serves.
- Variety driver and innovation. Creates meaningful differentiation or serves an emerging need.
- Price-image item and premiumizer. One supports perceived category value; the other trades shoppers up.
- Basket builder, seasonal, regional. Drives complementary purchases, a defined period, or a specific geography.
- Tail and duplicate. A narrow niche — or little unique demand and no strategic role at all.
Roles then structure the range itself: a core assortment that is broadly relevant, consistently available, and operationally reliable (national core, format core, cluster core), an optional assortment activated by local demand, store size, demographics, mission, region, or channel, and a tail that can live online, in selected stores, seasonally, on demand, or on a marketplace — because omnichannel retail lets digital breadth absorb what the physical shelf cannot. Each product then receives one of five actions: add (fills a need gap, adds incrementality, serves emerging demand, improves coverage, differentiates), retain (the role still holds), expand (strong velocity and incrementality, but underdistributed or underfaced), reduce (still necessary, but with excess distribution or facings relative to its localized demand), or delist.
The reverse decision needs equal discipline. A listing should be judged on need gap, differentiation, expected categoryincrementality, own-portfolio cannibalization, retailer economics, required space, supply readiness, and launch support — not on forecast sales alone. The innovation hurdle is four questions: how much demand is incremental, which existing products lose, which need is newly served, and why the retailer’s category grows. New items also need incubation spaceand a defined review window — launch date, distribution ramp, support period, evaluation date, success measures, exit criteria — so the system can tell a mature underperformer from a new item still ramping, a seasonal product, and a supply-constrained innovation. Protection without review is how weak items become permanent. Private label follows the same logic: entry, mainstream, premium, specialist, or exclusive-innovation roles, judged not by “which has the higher margin rate” but by which combination of national brands (traffic, innovation, marketing investment, trust, category recruitment) and own label (margin, value perception, differentiation, control) creates the best total proposition.
One range rarely fits every store
Stores differ by size, location, demographics, income, household size, culture, mission, competition, channel, tourism, climate, and local preference — so a nationally optimized assortment is frequently locally wrong. Store clustering groups stores with similar demand using format, space, category sales, demographics, mission, geography, shopper segments, category mix, price sensitivity, brand preference, and online penetration. Behavioural variables (actual category purchases, brand preference, price sensitivity, pack mix, trip mission) generally predict assortment demand better than demographics — though demographics help explain why a cluster behaves as it does.
A workable structure is modular: national core + format module + cluster module + local exceptions + seasonal module. That balances relevance against operational scale, which matters because localization is not free — every additional local range adds planning complexity, replenishment complexity, inventory fragmentation, planogram work, and implementation risk. Two further disciplines keep it honest. Cluster stability: clusters that churn every review make supply unstable and learning impossible, so monitor whether they remain behaviourally meaningful. And local exceptions must be reason-coded, time-bound, evaluated, and visible — otherwise store-level requests quietly dismantle the category strategy.
Assortment and space are one decision
The range determines what needs space; the space determines how the range performs. Optimizing them sequentially misses the interaction — academic work modelling assortment selection and shelf-space allocation jointly, with substitution, stochastic demand, and space elasticity, found integrated models outperforming sequential ones in empirical applications. Space itself is measured in linear centimetres, horizontal and vertical facings, shelf area, cubic capacity, peg positions, and display slots. A facing — one visible front presentation — buys visibility, capacity, fewer stockouts, and stronger brand blocking, at the cost of scarce width. Capacity is units per facing in depth × facings × stacking; divide by daily sales and you have shelf days of supply, which is where merchandising meets replenishment.
Space elasticity — percentage change in demand ÷ percentage change in space — varies by category, brand, position, store, fixture, baseline facings, package design, and replenishment, and it operates through two mechanisms that should be separated wherever possible: a visibility effect (more shoppers notice and select the product) and a capacity effect (the product simply stays in stock longer). Minimum facings may be required for visibility, handling, case-pack fit, shelf-ready packaging, or block integrity; beyond some point, extra facings add almost nothing. Position compounds it: eye level, hand level, top, bottom, aisle entrance, endcap, checkout, and secondary display each change visibility, accessibility, and premium perception, while horizontal versus vertical blocking and adjacency (pasta near sauce, formula by age stage, cleaning by task, yogurt by need) determine whether shoppers can navigate at all.
All of it lands in the planogram— which products sit where, with how many facings, in what orientation, against real fixture dimensions. That requires GTINs, product images, width, height, depth, orientation, nesting, case configuration, minimum facings, sales, margin, inventory, shelf dimensions, adjacency rules, and category strategy; GS1’s product-image specification defines planogram imagery and dimensions precisely so space-planning software can replicate products and capacity accurately. Shelf-ready packaging must be represented as it actually arrives, or the plan will not survive contact with the store.
Finally, a shelf plan that ignores operations fails in the aisle. Backroom space, case-pack quantity, delivery frequency, replenishment labour, shelf capacity, and demand volatility are all binding constraints — research treats assortment, space, and replenishment as one connected planning problem. And on-shelf availability is not system inventory: stock can be in the backroom, in the wrong location, blocked, miscounted, unreplenished, or missing from the digital shelf. Phantom inventory — the system says available, the shopper cannot find it — quietly corrupts every velocity number downstream. A fast SKU with too few facings will stock out repeatedly, so more space can create value with no visibility effect at all; a slow SKU may hold many days of supply on a single facing, which is an acceptable cost if it protects a unique need. In fresh and chilled categories the trade-off tightens further: more space increases waste, less space increases stockouts.
Endless aisle is not costless
Online assortment escapes the physical shelf limit but not the real constraints: shopper attention, search ranking, navigation, fulfilment, warehouse capacity, content quality, and availability. A sprawling digital range creates choice overload, poor findability, low-quality listings, duplicate products, fulfilment complexity, and search dilution. Digital facings are search ranking, sponsored placement, filters, recommendations, category navigation, images, ratings, and availability. A product can hold a physical core role, an online-only role, click-and-collect, selected-store, or marketplace status — and when an item leaves the store but stays online, some shoppers migrate, some substitute in-store, and some leave, so channel migration belongs in the model. One trap deserves naming: poor digital content — title, images, dimensions, ingredients, nutrition, allergens, claims, search terms — looks exactly like weak demand in the data.
The agentic category-management workflow
Category business brief
-> deterministic intake + entity validation
-> Agentic Category Manager
-> shopper, item, assortment, financial,
space, inventory, and market tools
-> choice, transfer, and optimization engines
-> planogram and replenishment services
-> independent category evaluator
-> retailer and supplier review
-> authorized implementation
-> store and digital execution monitoring
-> post-reset measurement
-> governed learningThe agent earns its place on the uncertain parts: resolving category boundaries, identifying the attributes that structure choice, retrieving fragmented evidence, selecting the right incrementality model, detecting gaps, investigating conflicting data, generating localized scenarios, explaining transfer, coordinating space and replenishment analysis, and preparing retailer-ready recommendations when constraints shift. Governed software owns sales aggregation, financial calculation, substitution and choice models, assortment and space optimization, planogram rendering, inventory maths, authorization, workflow state, execution, and audit. Humans own category definition, retailer strategy, category role, value judgments, supplier negotiation, legal interpretation, final listing and delisting authority, strategic exceptions, and accountability.
The twenty-three stages
- 01Category brief. Retailer, market, category, objective, role, formats, horizon, space constraints, strategic priorities, deadline.
- 02Category-definition check. Is the boundary shopper-relevant? Which substitutes sit outside it? Which complements matter? Is the decision national, format, or cluster-level?
- 03Entity resolution. Category, subcategory, segment, brand, SKU, GTIN, retailer, banner, store, cluster, planogram, fixture.
- 04Evidence plan. Category and SKU economics, distribution, availability, promotion, decision tree, incrementality, transfer, loyalty, attributes, basket relationships, dimensions, replenishment, market gaps, innovation pipeline.
- 05Data validation. Missing GTINs, hierarchy errors, duplicated products, wrong dimensions, inactive items, incomplete store data, false zeroes, out-of-stock and promotion distortion, mismatched periods.
- 06Category assessment. Growth by segment, brand, private label, channel, store, shopper group, price tier, format, pack size, innovation.
- 07Shopper structure. Decision tree, need states, switching relationships, loyalty groups, store missions, key attributes.
- 08Item diagnosis. For every SKU: sales, margin, velocity, distribution, availability, promotion dependence, incrementality, transfer, loyalty, need-state role, space, inventory, complexity.
- 09Gap and duplication analysis. Underserved needs, missing price tiers, pack roles and dietary requirements; redundant products, excessive flavours, overdistributed tail, underdistributed winners.
- 10Store clustering. Create or retrieve clusters, then validate stability, demand differences, operational feasibility, and space profiles.
- 11Assortment scenario generation. Current range, bottom-seller removal, low-incrementality removal, loyalty-protected, innovation-led, localized tail, expanded core, private-label growth, maximum simplification.
- 12Demand-transfer simulation. Retained, transferred, and lost category demand; store-switching risk; manufacturer portfolio and private-label effects; basket effects.
- 13Financial evaluation. Category revenue, gross profit, contribution, inventory, working capital, waste, operational cost, supplier portfolio effects.
- 14Space allocation. Listed products, facings, positions, category blocks, capacity, replenishment frequency, adjacency.
- 15Planogram simulation. Physical fit, visual clarity, category flow, brand blocks, capacity, minimum facings, accessibility.
- 16Operational feasibility. Distribution centre, case packs, store replenishment, backroom, master data, shelf labels, reset labour, product availability, launch readiness.
- 17Independent evaluation. Objective alignment, neutrality, shopper coverage, correct transfer model, financial integrity, space feasibility, data authority, policy compliance, disclosed uncertainty.
- 18Recommendation packet. Diagnosis, assortment actions, shelf plan, cluster differences, expected impact, risks, implementation, measurement plan.
- 19Human decision. Retailer stakeholders decide; supplier teams review what concerns them.
- 20Implementation preparation. Listing and delisting files, planograms, distribution instructions, supplier notices, store reset tasks, replenishment parameters, digital changes, evaluation design.
- 21Execution monitoring. Listing, distribution, planogram compliance, availability, store exceptions, digital content, shelf labels, stock.
- 22Post-reset analysis. Category sales and profit, actual transfer, availability, shopper retention, store switching, inventory, supplier performance, innovation.
- 23Learning. Validate transfer assumptions, cluster differences, space response, delisting impact, innovation incrementality, and implementation problems.
The agent’s twenty-seven tools
- Resolution and diagnosis: resolve_category_entities (category, subcategory, segment, need state, product and retailer hierarchy), get_category_scorecard (sales, volume, profit, growth, penetration, frequency, distribution, availability, inventory, promotion, share), get_item_performance (sales, margin, velocity, distribution, availability, promotion dependence, lifecycle).
- Shopper evidence: get_shopper_decision_tree (attributes, hierarchy, segment differences, source, confidence, effective period), get_customer_item_importance (loyal households, exclusive buyers, repeat, trip importance, basket).
- Incrementality: estimate_item_incrementality (retailer and manufacturer incrementality, confidence, model version, evidence), estimate_demand_transfer (destination probabilities, not a single number).
- Range structure: identify_assortment_gaps (missing needs, attributes, price points, pack roles, brands, occasions), identify_duplicate_items (attribute similarity, switching, incrementality, price and shopper overlap), create_store_clusters (approved methods only).
- Simulation and economics: simulate_assortment (adds, retains, delists, distribution, cluster, constraints → sales, profit, transfer, loyalty risk, inventory, confidence), calculate_retailer_category_pnl and calculate_manufacturer_portfolio_impact — deterministic services that keep the two sets of economics visibly separate.
- Space: get_product_dimensions (width, height, depth, orientation, case pack, shelf-ready packaging, GTIN), calculate_shelf_capacity, estimate_space_response (capacity effect, visibility effect, confidence, applicable range), optimize_assortment_and_space (jointly, under substitution and strategy), generate_planogram (artifact, placement, facings, shelf usage, constraint warnings).
- Operations: check_replenishment_feasibility (shelf days of supply, frequency, case-pack fit, backroom needs, labour risk), check_on_shelf_availability (availability, out-of-stock risk, phantom-inventory warning, lost-sales estimate), get_innovation_pipeline (readiness, target shopper, need gap, forecast, listing status).
- Governance and execution: retrieve_category_policy (authority, supplier-data rules, private-label rules, legal controls, approval), create_category_decision_packet, create_implementation_tasks (listing, delisting, planogram, data, supply, store tasks).
- Monitoring and learning: monitor_reset, run_post_reset_analysis, propose_category_learning — a reviewable candidate that never writes memory on its own.
Weak: "SKU A should be removed."
Strong: sku_id: SKU-A
recommended_action: localize
current_distribution: 92% recommended: 34%
velocity_percentile: 28
retailer_incrementality: 0.61
manufacturer_incrementality: 0.73
loyal_households_at_risk: 8,240
space_released_cm: 18
expected_category_sales_change: +0.2%
confidence: mediumDurable state, governed learning
A range review runs for months across many hands, so the workflow needs explicit state — brief, definition review, assessment, evidence gathering, data issue, shopper analysis, assortment diagnosis, scenario generation, space planning, scenario review, awaiting retailer decision, awaiting supplier validation, implementation planning, scheduled, reset in progress, monitoring, post-reset analysis, completed, cancelled. A new assortment version is required when SKU actions, distribution, cluster logic, shelf space, prices, innovation timing, or product dimensions change; the planogram binds to a specific assortment version, dimension set, cluster, effective date, fixture, and approval. Without that binding, stores execute a plan that no longer matches the decision anyone approved.
Memory is selective. Worth keeping once validated: stable demand-transfer patterns, cluster-specific needs, historical delisting outcomes, shelf-response estimates, recurring availability issues, retailer strategy, approved category definitions, proven incubation periods, and validated override outcomes. Never automatic: supplier claims, a single buyer comment, unverified competitor strategy, a temporary supply failure treated as a permanent demand signal, speculative shopper interpretation, or sensitive future pricing and promotion information. The learning record captures the decision, retailer and cluster, items added and removed, space changes, predicted versus actual transfer, predicted versus actual category impact, availability, shopper and operational effects, root cause, reusable lesson, applicability, and reviewer — because a reset that underperformed because it was never properly executed teaches nothing about demand.
Decision rights, neutrality, governance
Ten functions keep their crafts. The retail category manager or buyer owns the category role, retailer strategy, the assortment and shelf-space decisions, supplier negotiation, and final accountability. The manufacturer category manager owns category insight, shopper analysis, the supplier portfolio recommendation, and fact-based collaboration. Space planning owns fixtures, facings, layout, and planogram standards; merchandising owns category flow, visual strategy, and adjacency; shopper insights owns missions, decision trees, segmentation, and loyalty; data science owns incrementality, transfer models, optimization, validation, and experiments; supply chain owns availability, replenishment, case packs, and backroom feasibility; store operations owns reset execution and local feedback; finance owns definitions and category profit; legal and compliance owns category-captain boundaries, supplier information, and competition law.
| Decision | Agent | Human | Software |
|---|---|---|---|
| Define category | Analyse | Decide | Store taxonomy |
| Build shopper tree | Coordinate | Validate | Model |
| Estimate incrementality | Explain | Challenge | Estimate |
| Generate assortment | Coordinate | Set constraints | Optimize |
| Select delisting | Recommend | Retailer decides | Validate |
| Allocate space | Recommend | Space team approves | Optimize |
| Build planogram | Prepare | Approve | Render |
| Implement reset | Coordinate | Authorize | Execute tasks |
| Store learning | Propose | Validate | Save |
Neutrality is the governance question unique to this workflow. A manufacturer-owned agent must not quietly maximize its own sales, its own share of shelf, or competitor delistings while the stated objective is retailer category growth — so every run should declare its primary and secondary objectives, whose economics are included, the hard constraints, and the decision owner. Retailer and manufacturer economics should be displayed side by side, conflicts included: “this scenario improves retailer category profit, reduces Manufacturer A sales, and increases Manufacturer B and private label.” Category-captain controls follow: the retailer owns the decision, item rules are neutral, competing products are evaluated consistently, recommendations carry provenance, access to sensitive data is limited, an independent retailer review exists, legal approves the boundaries, and future competitor plans are never used. Every delisting must be reproducible from approved criteria and open to challenge, with an appeal path — corrected data, a supply explanation, innovation evidence, a contractual constraint — that is logged and evaluated rather than absorbed informally.
Evaluating the Agentic Category Manager
The failure surface spans the whole decision: wrong hierarchy, wrong category boundary, sales-ranking bias, inaccurate transfer, poor loyalty estimates, missing private label, invalid dimensions, an impossible planogram, ignored replenishment, a biased supplier recommendation, or an unexecuted reset. Component tests cover entity resolution, hierarchy, decision-tree retrieval, item performance, incrementality, transfer, financial calculation, dimensions, capacity, and policy. Historical replay runs past resets on planning-date information only and compares the proposed assortment with the actual decision and outcome — while accepting that the unchosen assortment was never observed, so causal models, store variation, test-control designs, and expert review must carry the counterfactual. Prospective tests in matched stores or clusters remain the strongest evidence, and should measure well beyond immediate sales.
Trajectory tests require the agent to resolve the category and item hierarchy, adjust for distribution and availability, estimate incrementality and transfer, check loyalty and need coverage, evaluate retailer economics, verify physical space and replenishment, and request approval — while prohibiting delisting by sales rank alone, favouring a supplier without evidence, using invalid dimensions, claiming exact transfer without support, and implementing before authorization. The metric families run wide: assortment-model (category sales and profit prediction, transfer and lost-demand accuracy, ranking accuracy, interval coverage, cluster performance), item-decision (delisting success, false delisting rate, successful additions, retained low-value items, innovation survival, distribution-expansion and localization value), shelf (revenue and profit per metre, contribution per facing, on-shelf availability, planogram compliance, replenishment labour, waste), shopper (penetration, incidence, repeat, loyalty, store switching, perceived variety, need-state coverage), retailer and manufacturer outcomes, process (cycle time, analyst hours, data exceptions, decision reversals, implementation delay, store exceptions), and safety (biased recommendations, prohibited information use, cross-retailer leakage, unsupported delistings, approval bypass).
The outcome that matters is shopper-adjusted category contribution per unit of scarce retail and operational capacity — not average sales per SKU.
Forty SKUs, four metres, two innovations
An illustrative chilled-yogurt category: 40 SKUs across four metres of shelf — mainstream family (10 SKUs, 35% of sales, 30% of space), single serve (8 / 20% / 18%), children’s (7 / 16% / 18%), high protein (6 / 14% / 14%), lactose free (4 / 7% / 8%), plant based (5 / 8% / 12%). The retailer wants less complexity and space for two innovations. The sales-ranking proposal removes two lactose-free SKUs, two plant-based SKUs, and two children’s flavours: six SKUs out, 58 centimetres released, 96% of sales retained.
The agentic diagnosis reorders the list entirely. Lactose-free L2 ranks 35th of 40 with low velocity — and high retailer incrementality, a high exclusive-buyer rate, only 42% transfer to the remaining range, and material store-loss risk: retain, in high-demand clusters. Plant-based P4 ranks 38th with low incrementality and 81% transfer to P1 and P2: delist nationally. Children’s C7 ranks 34th, but availability is 71% — its adjusted velocity is medium and its underperformance is partly supply-driven: fix availability before judging it. Mainstream M8 ranks 9th with high velocity, low incrementality, 88% transfer to M3 and M4, and five facings: reduce distribution or facings. The decision tree explains why the naive list was dangerous — dietary requirement is an earlysplit, so cutting coverage there carries disproportionate risk — and four store clusters (urban high plant-based at 3.5 m, family suburban at 4.5 m, value-led at 3.0 m, large destination at 6.0 m) mean “national delist” is rarely the only option.
| Scenario | Category sales | Category profit | Loyalty risk |
|---|---|---|---|
| 1 · Sales-rank rationalization | −0.8% | +0.3% | High in clusters A and B |
| 2 · Incrementality-led rationalization | +0.7% | +1.4% | Low |
| 3 · Maximum simplification (30 SKUs) | −2.5% | +0.4% | High in urban and dietary segments |
Scenario 2 wins on both axes at once. It delists four genuinely substitutable items, localizes L2 to the clusters that want it, expands the high-protein core where it performs, adds a high-protein family pack and an affordable plant-based core pack, cuts M3 from five facings to three, and raises C2 from two to three because it keeps stocking out. National SKU reduction is 7.5% while the average physical assortment falls 12% — smaller shelves, not a smaller catalogue. Space shifts rather than shrinks: mainstream family from 120 to 106 cm, children’s 72 → 76, high protein 56 → 68, lactose free 32 → 34, with the total still 400 cm. The retailer forecast is +0.7% sales, +1.4% gross profit, −8% inventory units, and +2.5 points of availability. Some suppliers gain and some lose — which is stated openly rather than buried. The main uncertainty is M8’s transfer, so the recommendation is a matched-store test before national rollout: 20 test and 20 control stores, an eight-week pre-period and a twelve-week post-period, measuring sales, profit, transfer, availability, store switching, and basket.
Implementation and data readiness
- 01Phase 0 — align the category standard. Definition, role, objectives, financial metrics, incrementality definition, product hierarchy, decision rights, legal boundaries.
- 02Phase 1 — data foundation. Retailer POS, product master, store hierarchy, distribution, inventory, promotion, margin, loyalty or panel, product dimensions, planograms.
- 03Phase 2 — analytical foundation. Decision trees, incrementality and transfer models, loyalty metrics, clustering, assortment simulation, space optimization, replenishment model.
- 04Phase 3 — diagnostic agent. Explains category performance, item roles, gaps, duplication, availability, and space productivity. No live listing decisions.
- 05Phase 4 — scenario copilot. Assortment alternatives, transfer simulations, shelf scenarios, decision packets; humans fully responsible.
- 06Phase 5 — shadow recommendation. Runs alongside the existing range review; compare recommendations, decision quality, cycle time, and outcomes.
- 07Phase 6 — advisory deployment. Category managers use it in live reviews; measure trust, overrides, corrections, time, and decision quality.
- 08Phase 7 — approval-based implementation. The agent prepares delisting files, distribution changes, planograms, and store tasks; humans approve execution.
- 09Phase 8 — continuous monitoring. Emerging gaps, tail growth, duplication, availability problems, planogram drift, innovation performance — detected between reviews, not at them.
A strong pilot takes one retailer, one category, one market, 30–100 SKUs, several store formats, stable POS, accurate dimensions, an engaged category owner, and a manageable reset cycle. Avoid starting with the full store, thousands of SKUs, an incomplete item hierarchy, missing dimensions, unclear retailer authority, or a highly seasonal category with thin history. Success criteria belong in writing beforehand: zero unauthorized assortment changes, accurate hierarchy resolution, demand-transfer evidence available for priority SKUs, materially shorter review cycles, category sales and profit at agreed thresholds, availability that does not deteriorate, and no material loyalty segment unintentionally removed.
The minimum viable data is SKU sales, units, price, margin, store distribution, promotion, product and store hierarchy, item status, product dimensions, and shelf space; it gets materially stronger with loyalty households, panel, baskets, out-of-stock records, planogram compliance, store-level inventory, missions, attributes, competitor distribution, digital availability, images, replenishment, and waste. Three foundations carry disproportionate weight. Product attributes — brand, format, flavour, pack size, price tier, benefit, ingredient, dietary requirement, user, occasion, sustainability, origin — are what decision trees and gap analysis are built from. Physical dimensions — width, height, depth, orientation, tray, case, nesting, peg, unit of measure, standardized to GS1 — are what makes capacity real. And assortment status must distinguish active, new, incubation, seasonal, discontinued, temporarily unavailable, supply constrained, online only, and cluster-specific, with retailer weeks, reset dates, listing dates, launches, promotions, stockouts, and planogram effective dates aligned on one calendar.
Twenty failure modes
- 01Delisting the lowest sellers. Sales rank treated as incrementality.
- 02Confusing supplier and retailer incrementality. Unique to the supplier, easily substituted in the category.
- 03Ignoring loyalty. A niche product quietly protects high-value shoppers.
- 04One national assortment. Local needs disappear into an average.
- 05Excessive localization. Operational complexity outweighs the local benefit.
- 06Optimizing assortment before space. The chosen range cannot be merchandised.
- 07Optimizing space before assortment. More facings for a structurally weak range.
- 08Ignoring availability. Supply-constrained products look unproductive.
- 09Treating out-of-stock substitution as delisting transfer. Shoppers expecting a return behave differently.
- 10Treating attribute similarity as substitution. Similar products serve different needs.
- 11Ignoring the digital shelf. Physical delisting silently becomes channel migration.
- 12Ignoring case packs and replenishment. A beautiful planogram that cannot be filled.
- 13Overprotecting innovation. Weak new items stay indefinitely.
- 14Killing innovation too early. Judged before distribution and awareness mature.
- 15Optimizing manufacturer share of shelf. The recommendation loses category neutrality.
- 16Hidden private-label bias. Expansion without accounting for traffic loss.
- 17False precision in transfer rates. Weak evidence presented as an exact percentage.
- 18No test plan. A national reset before any causal validation.
- 19The agent executes delistings. Distribution changed without authorized retailer approval.
- 20No post-reset learning. The next review repeats this year’s assumptions.
The CATEGORY Method and maturity model
- 01Clarify the category, role, and objective. Shopper need, boundary, role, retailer strategy, horizon, success metrics, constraints.
- 02Analyse shopper choice and category performance. Decision trees, missions, segments, loyalty, price tiers, category growth, store differences.
- 03Tag every item with an evidence-based role. Core, traffic, margin, need-state protector, variety, innovation, duplicate, tail.
- 04Estimate incrementality, transfer, and loyalty risk. Retailer and supplier incrementality, substitution, category loss, store-switching risk, basket effects, uncertainty.
- 05Generate localized assortment and shelf scenarios. National core, formats, clusters, adds, delists, space, facings, planograms.
- 06Optimize shopper, retailer, manufacturer, and operational value. Sales, profit, loyalty, supplier effects, availability, inventory, replenishment, complexity.
- 07Route decisions through governance and implementation. Retailer authority, supplier contribution, legal controls, approval, listing, planogram, store execution, monitoring.
- 08Yield validated learning. Actual transfer, category impact, shopper and space response, availability, implementation quality, reusable lessons.
| Level | What it adds | Characteristics |
|---|---|---|
| 0 · Sales-ranked assortment | A ranking | National range, spreadsheets, bottom-seller delisting, little shopper evidence |
| 1 · Performance-based assortment | Normalization | Sales, velocity, margin, distribution, basic segmentation |
| 2 · Shopper and incrementality management | Causality | Decision trees, loyalty, demand transfer, need coverage, store clusters |
| 3 · Integrated assortment and space | Joint optimization | Facings, planograms, replenishment, availability, range and space solved together |
| 4 · Agentic Category Manager | Orchestration | Adaptive evidence gathering, scenario orchestration, localized recommendations, neutral governance, evaluation |
| 5 · Continuous category operating system | Closed loop | Always-on assortment health, continuous transfer learning, exception detection, physical and digital integration |
Part XX in brief — the practitioner templates. The method ships as seven working documents: the category brief (retailer, definition, role, objectives, formats, SKU count, available space, known issues, deadline); the SKU role card (segment, brand, pack, price tier, role, sales, margin, velocity, distribution, availability, both incrementality figures, transfer destinations, loyalty risk, need-state coverage, space, recommended action); the delisting card and innovation listing card (the two directions of the same decision, each with confidence, scope, and — for innovation — incubation and exit criteria); the assortment scenario card (adds, retains, localizations, delists, sales and profit impact, supplier impact, transfer, loyalty risk, inventory, space released, availability, complexity, P10/P50/P90, required approval); the shelf-space card (dimensions, current and proposed facings, capacity, daily sales, shelf days of supply, revenue and profit per facing, incremental facing value, replenishment frequency, position); the decision packet; and the post-reset learning card. If these artifacts do not exist, the range review was not designed — it was negotiated.
Frequently asked questions
Does the agent decide which products are listed?
No. It should not hold autonomous authority over listing, delisting, or shelf allocation. It analyses, simulates, recommends, prepares, and monitors under retailer governance; the retailer decides.
Why can’t we just delist the slowest sellers?
Because a slow seller may protect a unique need, a loyal shopper group, or store traffic — while a fast seller may be readily substituted by products already on the shelf. Rank measures current conditions; incrementality measures what would actually be lost.
What is the difference between retailer and manufacturer incrementality?
Retailer incrementality measures unique contribution to the retailer’s category; manufacturer incrementality measures unique contribution to the supplier’s portfolio. A product can be highly incremental for one and entirely substitutable for the other, which is exactly why both belong in the packet.
What is demand transfer?
An estimate of where shoppers go when an item is unavailable or removed: another pack in the same brand, another brand, private label, another category, another retailer, or no purchase at all.
Should high-selling products always receive more facings?
No. The question is what the next facing adds — considering incremental demand, availability, capacity, category strategy, and what else that space could do.
What is on-shelf availability?
Whether the product is available where the shopper expects to find and buy it — not whether inventory exists somewhere in the store. Backroom stock, misplacement, and phantom inventory all read as weak demand.
How should new products be evaluated?
On need gap, differentiation, expected category incrementality, cannibalization, retailer economics, and operational readiness — with a defined incubation period, success measures, and exit criteria agreed before launch.
How should private label be treated?
As one or more strategic category propositions, not as a high-margin replacement for national brands. The right question is which combination creates the best total category and shopper proposition.
Can a manufacturer-owned agent act as a neutral category manager?
Only if objectives, evidence, item rules, data access, and retailer decision rights are transparent and governed. Category captaincy carries real legal risk — disadvantaging competing suppliers, or inappropriate exchange of commercially sensitive information — and qualified counsel should define the operating boundaries.
What is the biggest AI mistake in category management?
Automating sales-ranked delisting while ignoring shopper needs, incrementality, demand transfer, loyalty, shelf capacity, availability, replenishment, and retailer authority.
Conclusion
Category management looks simple product by product: every item has sales, margin, distribution, and a position on the shelf. But a category is not a collection of independent items — it is a system of shopper choices. Add a SKU and demand shifts. Remove one and demand transfers. Give more space and both visibility and availability change. Cut the range and choice, loyalty, inventory, and store operations all move at once. Expand private label and the price architecture changes. Localize and relevance improves while complexity grows. Which is why the decision cannot be compressed into rank, remove the bottom, add the new products.
An assortment platform can calculate scenarios. A space-planning tool can build a planogram. A loyalty system can identify customer importance. A transfer model can estimate substitution. The Agentic Category Manager connects them into one governed process: it resolves products, stores, clusters, and fixtures; separates velocity from incrementality; recognizes that a high seller can be substitutable and a slow seller can protect something strategically important; estimates where demand will move; generates localized alternatives; links range to space; checks replenishment and availability; exposes conflicts between retailer and manufacturer outcomes; prepares decisions for authorized people; verifies that the reset was actually implemented; and converts real transfer and shopper response into learning. The CATEGORY Method walks the route: clarify, analyse, tag, estimate, generate, optimize, route, yield.
The practitioner’s time then moves where it belongs. Less of it spent joining datasets, rebuilding item tables, copying planograms, and reconciling contradictory reports; more of it spent on category strategy, shopper needs, retailer trade-offs, supplier collaboration, innovation, judgment, and implementation. That is not the automation of category management.
The defining question is not which SKUs deserve to remain on shelf. It is which assortment and shelf architecture best serves the category’s shoppers while using limited retail and operational capacity to create sustainable category value.
Want help choosing the right architecture for your process?
We map where agents create leverage in FMCG operations, then build and ship the ones that pay back. One call to pressure-test your highest-leverage use case.