The Perfect Store agent: from retail execution photos to corrective action
A field sales representative enters a supermarket on Tuesday morning. The visit checklist runs to 42 questions. In the soft-drinks aisle, one of the company’s hero SKUs is missing from the shelf — while the retailer’s inventory system shows 18 units in stock. Another SKU has one facing instead of the agreed four. A promotional pack is present, but its price label is not. The company’s display stands near the entrance with two competitor products inside it. A new flavour that should have launched three days ago is absent.
The representative photographs the shelf. The retail-execution app analyses the images and returns: overall Perfect Store compliance: 73%, with one out-of-stock, two planogram deviations, one missing price label, one display issue, and one missing new product. This looks useful. It answers none of the questions that matter next.
Is the hero SKU truly out of stock, or sitting in the backroom? Is the inventory record simply wrong? Did the store fail to replenish, or was the product never delivered, or was the retailer’s distribution centre short? Is the missing new product a listing problem, a master-data problem, or a supply delay? Can the representative fix any of this right now, or does a store employee need to act, or should this become an order, an escalation to the key account manager, or a supply investigation? Is the missing price label a store failure or evidence that the promotion was never activated properly? How much revenue is lost per day if nothing changes? Which of the five problems deserves the 25 minutes actually available? How will anyone confirm the correction happened? And after the representative leaves — who ensures the shelf is still compliant tomorrow?
This is the real Perfect Store problem. Detecting a shelf gap is not resolving it. Measuring execution is not improving execution. Taking a photograph is not creating commercial value. The traditional process stops too early: visit store → complete survey → take photo → calculate compliance → publish dashboard. The business receives a retrospective score. The shopper still faces an empty shelf.
The highest-value application of agentic AI in retail execution is not automatic shelf auditing. It is a closed-loop system that converts observed store conditions into prioritized, evidence-based, assigned, and verified corrective actions.
The objective is not the highest possible compliance score. It is to ensure the products, prices, promotions, and placements the business designed are genuinely available to the shopper — and to correct the most valuable gaps while action is still possible.
What a Perfect Store really is
A Perfect Store is an outlet in which the right products are available, visible, correctly priced, correctly positioned, properly promoted, easy to shop, and supported by the operational conditions required to sustain that execution. Enterprise systems often compress this to “the right product, in stock, at the right time, at the right price,” but a mature framework reaches further: assortment, distribution, on-shelf availability, share of shelf, facings, position, brand blocking, planogram compliance, promotional availability and price and display, secondary placement, POS materials, retailer media, freshness, competitor context, digital availability, and store-staff engagement.
Critically, “perfect” should never mean every standard always achieved. A programme should define the few execution conditions that matter, for a particular store type, during a particular period, for a particular commercial strategy. A neighbourhood store and a hypermarket cannot be judged against the same picture. A promotional week and an ordinary week should not share standards. A launch store and a maintenance store need different actions. The Perfect Store is contextual, and the Picture of Success is what makes that context observable — converting strategy into store conditions: 12 mandatory SKUs, all hero SKUs available, minimum 18 facings, core brand at eye level, brand block intact, price within the approved retailer file, promotional SKU and label and display active, one approved display during the event, current campaign material installed. It must be store-specific enough to be meaningful, standardized enough to measure, achievable, commercially justified, and operationally observable. A planogram is one component of it, not the whole.
Three perspectives have to align. The manufacturer wants distribution, availability, visibility, promotion execution, share of shelf, launch support, displays, and incremental orders. The retailer wants category sales, shopper experience, availability, shelf productivity, planogram integrity, price accuracy, inventory health, labour efficiency, freshness, and safety. The shopper wants the expected product, available now, at a clear price, easy to find, in good condition, with the promised offer. Execution is the commercial last mile — strategy, customer agreement, supply, delivery, shelf execution, shopper purchase — and everything upstream is worth nothing if the last step fails.
Segments, standards, and the score
Start from strategy, not from the metrics that are easiest to collect. Objectives — availability, promotional execution, innovation support, share of shelf, price compliance, brand visibility, order value, field productivity, retailer collaboration — should be explicitly prioritized, then translated into standards that vary by store archetype. A strategic hypermarket may carry the complete assortment, high share of shelf, multiple displays, and frequent monitoring. An urban supermarket with limited space needs a focused range, hero-SKU availability, and compact promotional execution. A neighbourhood store needs a must-stock list, availability, price, display, and order capture. A discounter needs the contracted assortment, availability, the correct channel-specific pack, and agreed price positioning. A pharmacy or specialist channel needs core assortment, a benefit block, staff education, and premium visibility.
Within a store, standards should sort into three tiers. Mandatory conditions are required to qualify at all — hero SKUs available, promotional price correct, safety requirements met. Weighted conditions contribute to the score — share of shelf, display quality, POS material. Opportunity conditions go beyond minimum execution — add a facing, secure a secondary placement, list an incremental SKU. That tiering then drives a per-store objective hierarchy the representative can actually act on: primary, fix hero-SKU availability; secondary, install launch material; opportunity, increase the order on an understocked range.
The metric set itself splits into leading execution measures (availability, facings, price, promotion compliance, display, planogram, visit completion) and lagging business measures (sales, revenue, share, distribution, promotion uplift, customer profitability) — and the programme’s credibility rests on connecting the two. Twenty core metrics recur across FMCG programmes: assortment compliance (with present defined precisely, since contractually listed, active in retailer master data, delivered, in the backroom, and on the shelf are five different states), numeric and weighted distribution, on-shelf availability, out-of-stock, void, share of shelf, facings and facing compliance, planogram compliance, brand blocking, shelf position, price compliance, promotion compliance, display compliance, secondary placement, POS-material compliance, product condition, competitor context, order opportunity, and visit execution.
Two definitional points repay attention. On-shelf availability — as ECR frames it — is whether the product is available for sale to the shopper in the expected place at the time they want to buy it, which is a materially stronger claim than system inventory > 0. And point-in-time audits miss intra-day stockouts, so a time-weighted view (available SKU-time ÷ expected SKU-time) tells a truer story than an observation-level percentage. On facing compliance, cap the score where appropriate: six facings against a four-facing target should not read as 150% compliant. And a promotion can be technically active and commercially dead, because promotion compliance decomposes into correct product, dates, price, mechanic, stores, feature, display, availability, and materials — with availability being the one that most often fails silently.
The store truth problem
No single data source represents the shelf. A photograph shows the visible shelf, facings, gaps, placement, price labels, promotional materials, and displays — and cannot prove true store inventory, backroom stock, delivery status, the reason for an absence, future replenishment, correct retailer listing, or the checkout price. POS shows transaction history and velocity, but zero sales can mean no demand, a stockout, a store closure, missing data, or a product that was never ranged. Perpetual inventory can be wrong through shrink, receiving errors, mis-scans, damage, waste, misplacement, counting errors, or bad transfers — research on retail availability has repeatedly shown shelf conditions and inventory records diverging materially. The planogram may be current, outdated, store-specific, format-level, incorrectly implemented, or inconsistent with the retailer’s live assortment. The promotion calendar shows the plan; the store may have different dates, delayed materials, incomplete price activation, no stock, or a local exception. Store employees know things no system holds — and their statements should be recorded with source and confidence, not treated as fact.
So store truth is an evidence graph, and the agent’s job is to combine sources rather than trust one. A useful hierarchy runs direct observation → authoritative system record → corroborated operational evidence → human statement → inference, and contradiction should be surfaced, never smoothed over. When the photo shows an empty shelf, inventory shows 18 units, and the store employee says there is nothing in the backroom, the possibilities are inventory inaccuracy, misplaced stock, unprocessed damage, or a recognition error — and the correct response is not a confident diagnosis but a physical stock check.
Photography and computer vision
The image is an operational input, and poor capture guarantees poor analysis. Standardize angle, distance, lighting, orientation, overlap, coverage, resolution, location, and timestamp; guide the capture workflow (stand at the required distance, align with the shelf, capture left, centre, right, confirm quality, submit); and record metadata — image ID, store, aisle, fixture, category, representative, timestamp, permitted geolocation, device, capture sequence, expected planogram. Validate before analysing: blur, exposure, glare, occlusion, angle, coverage, resolution, duplicates, wrong category, wrong store. Where quality fails, guided recapture while the representative is still standing there is worth more than a perfect diagnosis after they leave. Large fixtures need several images stitched into a virtual shelf, or realogram — a 2025 large-scale planogram-compliance study built exactly this pipeline across thousands of convenience stores. And because photographs can inadvertently capture shoppers, employees, receipts, and restricted areas, capture guidance, face blurring, data minimization, retention rules, access controls, and retailer agreements are part of the design rather than an afterthought.
The vision pipeline then runs image-quality validation → shelf detection → segmentation → stitching → product detection → classification → facing estimation → price-label detection and OCR → planogram alignment → compliance classification → confidence and review. Each stage has characteristic difficulties. Fine-grained classification is hard because FMCG packs differ only by small text, colour, a flavour marker, pack count, or a promotional flash — regular versus sugar-free, 500 ml versus 550 ml, old versus new artwork. Packaging changes constantly through redesign, seasonal artwork, sustainability claims, promotions, regulation, and language, so a catalogue that is not maintained will misclassify current products. Occlusion is endemic: products behind shelf lips, in trays, blocked by labels, stacked, turned sideways. Facing estimation struggles with partial facings, stacking, nesting, gaps, and shelf-ready packaging. OCR fights small text, glare, electronic labels, oblique angles, currency formats, promotion stickers, and multiple prices. And empty-shelf detection needs planogram, assortment, and inventory context, because a gap may be an intentional space, a discontinued product, a reset, shallow shelf depth, a product sold down, or simply a recognition failure.
From detection to root cause
Observations convert into a structured gap taxonomy: availability (empty shelf, partial stockout, hero or promotional SKU absent, display empty), distribution (expected SKU not listed, retailer system inactive, store exception, lost distribution, launch not activated), planogram (wrong sequence or shelf, insufficient facings, broken brand block, unauthorized product), pricing (wrong price, missing label, promotion not activated, expired label, checkout mismatch), promotion, display (missing, empty, competitor contamination, wrong location, damaged, expired campaign), product condition, and data (planogram unavailable, product unrecognized, stale inventory feed, missing price file, wrong retailer mapping). Each points to a different owner — which is the entire reason the taxonomy exists.
The practical classification behind that tree spans store execution (backroom stock, replenishment not completed, misplacement, shelf reset, damaged stock removed, staff capacity, label confusion), store ordering (no order, order too low, wrong reorder parameter, suspended order, minimum order, case-pack issues), retailer distribution centre (DC stockout, allocation, picking error, delivery delay, short shipment), manufacturer supply (production shortfall, forecast miss, customer allocation, transport failure, launch delay, packaging shortage), master data (inactive product, wrong GTIN, incorrect store range, price-file failure, launch-date mismatch), and promotion (uplift underestimated, coverage beyond plan, display sold faster, incomplete inventory loading, early activation).
Weak: Root cause: store staff failed to replenish.
Strong: most_likely_cause: backroom replenishment failure
evidence: 18 units recorded in store
24 units delivered yesterday
zero shelf facings detected
zero POS sales since 10:00
confidence: medium
required_confirmation: physical backroom checkPrioritizing and correcting
The representative may have 25 minutes, so every gap cannot receive equal attention. Priority combines commercial impact (expected lost sales and contribution, promotion and launch risk, retailer penalty, strategic-product impact), actionability (can the rep fix it now, can a store employee act today, or does it need upstream supply action), persistence risk (will it recur — inventory inaccuracy, wrong reorder parameter, chronic low facings, systemic delivery problems), confidence (confirmed, probable, possible), strategic importance, and effort. Hard rules escalate regardless of score: a hero SKU out of stock during a promotion, a safety or expiry issue, a major promotional price mismatch, a launch SKU missing in a priority store, an unexecuted display contract, a severe inventory-record discrepancy. Suppression rules work the other way, keeping non-applicable standards, low-confidence low-value observations, already-assigned issues, discontinued products, planned local exceptions, and duplicate detections out of the field team’s way.
What the representative should actually receive is a sequenced list — check the backroom for the hero SKU, ask the store manager to correct the promotional price, photograph the corrected display, open a launch-listing investigation, record the competitor display — not a compliance score of 73%. The corrective-action library behind that list spans immediate representative actions (replenish, rotate, correct orientation, reorganize the brand block, install POS material, refill the display, remove expired material, create an order, capture proof, speak to the manager), store-staff actions (bring stock from the backroom, correct the label, investigate inventory, activate the promotion, implement the planogram, process receiving, count stock), supervisor actions, key account actions (resolve listings, challenge distribution, correct the price file, enforce promotional agreements), supply-chain actions, master-data actions, and data-science actions (investigate repeated false positives, update the recognition model, recalibrate OOS detection).
Visit value and collaboration
Not every store deserves the same visit frequency. Traditional routing by fixed frequency, geography, historic route, and account class is simple and allocates effort poorly. Dynamic prioritization surfaces stores with a probable OOS, declining sales, an active promotion, a launch, unresolved actions, low compliance, high opportunity, or repeated inventory issues — scored as expected visit value = probability of a material issue × value recoverable × probability of successful correction − visit cost. Fixed frequencies need not disappear; they provide stable coverage while event-driven visits handle high-value exceptions. Once priority stores are identified, routing optimizes travel time, store windows, visit duration, capacity, territory, and skills — because some actions need merchandising skill, others negotiation, technical setup, inventory investigation, or relationship authority. One caution carries through: a store-health score that is as opaque as the compliance score it replaced is no improvement. Show the drivers.
And most execution issues cannot be solved unilaterally. Manufacturers generally do not control shelf replenishment, retailer inventory records, checkout pricing, store staffing, planogram implementation, or store ordering — so the system has to support collaboration. A retailer-ready packet carries store, SKU, timestamp, image, expected condition, observed condition, inventory evidence, sales evidence, and the requested action. The framing matters enormously: not “store staff failed,” but “the shelf is empty while inventory records indicate 18 units; a physical check and replenishment are requested.” ECR frames on-shelf availability explicitly as a joint retailer–manufacturer priority, and a joint workflow — detect, validate, create the retailer task, check stock, replenish or correct the record, confirm the outcome, review recurring causes together — needs an agreed SLA covering issue type, response owner, response time, evidence standard, escalation, data access, and closure. Shared execution evidence must never become a channel for exchanging future competitor prices, future promotions, confidential supplier strategy, or retailer-sensitive data.
The agentic Perfect Store
Field mobile app or shelf camera
-> image-quality service
-> computer-vision recognition
-> realogram and observation layer
-> Perfect Store Agent
-> planogram, assortment, price, promotion, POS,
inventory, order, delivery, and customer tools
-> root-cause diagnosis
-> action prioritization
-> human or governed execution
-> proof of correction
-> sales and compliance monitoring
-> validated learningThe division of labour is unusually clean here. Computer vision inspects images, detects products, classifies SKUs, counts facings, detects gaps and labels, aligns to the planogram, and returns structured observations. The language-model agent interprets multiple observations, selects tools, reconciles conflicting evidence, diagnoses likely causes, identifies missing information, prioritizes actions, explains recommendations, routes tasks, monitors closure, and summarizes recurring patterns. Deterministic software owns calculations, KPI scoring, order quantities, route optimization, authorization, inventory updates, order creation, task state, audit, and financial impact. Humans own physical correction, retailer interaction, commercial negotiation, ambiguous validation, high-impact decisions, approval, and accountability.
Two contracts hold the seams together. The observation contract returns store, fixture, image, SKU, observation type, actual and expected facings, position, confidence, bounding box, and timestamp. The diagnostic contract returns issue, likely root cause, evidence, alternative causes, commercial impact, actionability, recommended action, owner, deadline, and verification. One orchestrator with specialist tools is usually enough to start; specialists — image quality, OSA diagnostician, promotion compliance, visit priority, retailer collaboration, evaluator — earn their place only where specialization, permissions, or parallelism create value.
The seventeen stages
- 01Define the applicable Picture of Success. Retailer, store, format, category, period, promotion, planogram version, applicable KPIs.
- 02Prepare the visit. Open actions, last compliance, recent sales, inventory risk, promotion, launch, visit priority.
- 03Capture evidence. Shelf and display photos, price labels, survey answers, store notes, remote cameras.
- 04Validate evidence quality. Request recapture, request manual validation, or mark unavailable — before anything downstream runs.
- 05Produce structured observations. SKU absent, facings below target, promotion label missing, display contaminated, price mismatch.
- 06Consolidate related observations. An empty shelf, an empty display, stopped POS sales, and positive inventory may be one root cause, not four issues.
- 07Retrieve contextual evidence. Inventory, POS, orders, delivery, planogram, price, promotion, assortment, visit history.
- 08Diagnose root cause. Primary hypothesis, alternatives, evidence, confidence, information still required.
- 09Estimate value at risk. Expected lost sales and contribution, promotion and launch exposure, retailer importance.
- 10Determine actionability. Representative authority, store-employee availability, stock in store, order window, escalation route, deadline.
- 11Prioritize corrective actions. A sequenced in-store plan that fits the actual visit window.
- 12Execute or request action. Refill, reorder, correct display, install material, request price correction, open an investigation, escalate.
- 13Capture proof. After image, task confirmation, system status, order ID, retailer acknowledgement.
- 14Verify closure. The agent or a deterministic evaluator confirms the expected condition actually exists.
- 15Monitor persistence. Do sales resume, does the shelf stay available, does the issue recur, is inventory corrected, does delivery arrive?
- 16Measure outcome. Expected versus actual sales, compliance, time to closure, cost, repeat rate.
- 17Learn. Effective corrective actions, recurring root causes, store patterns, model failures, visit-policy improvements.
The agent’s twenty-seven tools
- Context and standards: resolve_store_context (retailer, banner, store, format, cluster, territory, sales potential, applicable programme), get_picture_of_success (required assortment, facings, position, prices, promotions, displays, POS materials, KPI weights), get_current_planogram (ID, version, effective dates, fixture, sequence, facings, positions).
- Vision: validate_image_quality (blur, exposure, angle, coverage, recapture instruction), analyze_shelf_image, stitch_shelf_images (the realogram), compare_realogram_to_planogram (missing, misplaced, facing and position deviations, unknowns), read_price_labels (SKU candidate, observed price, promotion text, date, confidence), recognize_display_and_posm (type, campaign, products, contamination, stock status, material condition).
- Store context: get_store_inventory (on hand, sellable, backroom where available, reserved, last count, record confidence), get_store_pos_sales (recent sales, zero-sales pattern, velocity, promotion response), get_store_orders, get_delivery_status, get_assortment_status (authorized, active, expected start, local exception, discontinuation), get_price_and_promotion_status.
- Analysis: estimate_lost_sales (through approved demand and substitution models), diagnose_osa_issue (combining vision, inventory, POS, delivery, assortment, order), calculate_action_priority (a deterministic scoring service), recommend_order_quantity (stock, sales, forecast, case pack, shelf capacity, order cycle, promotion).
- Action: create_store_action, create_retailer_action_request (under an approved workflow), create_sales_order_draft (a draft — it must not commit the retailer without authority), escalate_to_key_account (an evidence packet).
- Closure and learning: capture_after_image (linked to the original issue), verify_corrective_action, monitor_store_outcome, propose_execution_learning — a reviewable candidate, never an automatic write.
Weak: "Store has an availability issue."
Strong: case_id: PS-20418 store_id: NL-AMS-428
sku_id: SKU-184
issue: probable shelf OOS
observation: 0 facings detected (expected 4)
system_inventory: 18 units
last_delivery: 24 units received yesterday
pos_sales: zero since 10:04
estimated_sales_at_risk: EUR 286 per day
likely_root_cause: backroom replenishment or
inventory-record inaccuracy
recommended_action: physical stock check; refill if
stock exists, else inventory correction
owner: field representative + store manager
priority: critical
verification: after photograph + POS resumptionCases, memory, and decision rights
Three state machines run together. The case moves through detected, validation required, diagnosis in progress, action ready, assigned, in progress, waiting for retailer, waiting for supply, corrected, verification pending, verified, monitoring, reopened, closed. The visit moves through planned, en route, in store, evidence capture, actions in progress, completed, follow-up required. The action moves through created, accepted, started, blocked, completed, rejected, expired, verified — and that last state is what separates a closed-loop system from a task list. Memory holds repeated store inventory inaccuracies, recurring retailer price-file failures, store-specific replenishment patterns, effective correction types, image-recognition weaknesses, promotion execution patterns, and recurring planogram exceptions — scoped by store, retailer, format, product, category, representative, or market, and always with an effective date, confidence, evidence, review date, and deletion rule, because execution patterns change. Never stored automatically: unverified store-employee statements, temporary local exceptions, speculative competitor plans, a single incorrect detection, personal information from images, expired promotional terms, or blame-oriented narratives.
| Decision | Agent | Human | Software |
|---|---|---|---|
| Detect shelf deviation | Coordinate | Validate ambiguity | Recognize |
| Diagnose likely cause | Investigate | Confirm where needed | Retrieve evidence |
| Prioritize action | Recommend | Supervisor adjusts policy | Score |
| Refill shelf | Recommend | Rep or store acts | Record |
| Correct retailer price | Prepare request | Retailer decides | Update |
| Create order | Draft | Authorized person confirms | Submit |
| Escalate commercial issue | Prepare | KAM owns | Route |
| Close issue | Recommend | Owner confirms | Verify |
| Store learning | Propose | Expert validates | Save |
Twelve roles hold distinct ownership, from the field representative (evidence capture, authorized in-store action, the retailer conversation, proof of completion) and merchandiser (shelf, display, POS materials, facings) through supervisor, key account manager, trade marketing (the Picture of Success itself), category management, RGM, demand and supply planning, customer service, retailer store employees, data science, and master data. One point is easy to miss and expensive to get wrong: the system must know what the representative is actually permitted to do — touch stock, change facings, install materials, create an order, speak to store management, photograph, access the backroom, request a stock count — because authority varies by retailer and country, and assigning an action outside it damages the relationship the programme depends on.
Eight layers, not one accuracy number
A technically excellent vision model creates little value if actions are poorly prioritized, root causes are wrong, owners do not respond, corrections go unverified, or sales do not improve. So evaluation runs in layers: image quality (accepted images, recapture rate, blur, angle, coverage, stitching success), computer vision (shelf and product detection precision and recall, classification accuracy, facing-count error, price-label detection, OCR accuracy, planogram alignment, OOS detection), observation (is the structured claim correct), root cause (accuracy, confidence calibration, evidence completeness, appropriate requests for validation), priority (did high-value issues rank first, were low-value ones suppressed, was urgency and actionability right), action (recommendation and owner correctness, permission compliance, acceptance, completion, time to action), closure (proof quality, verification accuracy, recurrence, false closure, reopened cases), and business (OSA improvement, recovered sales, promotion execution, distribution, price compliance, incremental orders, field productivity, retailer satisfaction).
Three measurement disciplines separate a credible programme from a flattering one. Consequence-weighted evaluation: missing a hero-SKU out-of-stock is not equivalent to miscounting one facing on a tail SKU, so weight by sales, strategic importance, and action consequence — and watch both sides of the trade, since a low-precision system overwhelms field teams while a low-recall system quietly misses the value. Causal evaluation: improved compliance and improved sales are correlated for many reasons, so use matched stores, phased rollout, randomized visit assignment where possible, difference-in-differences, and controlled field experiments — a retail field experiment found targeted external audits reduced shelf out-of-stocks and inventory inaccuracies and increased sales in the studied setting, which is exactly the kind of evidence a compliance dashboard cannot produce. And lost-sales estimation must acknowledge that sales during an out-of-stock reveal nothing about original demand, because shoppers substitute, delay, switch stores, or abandon the purchase — and research shows that relying on inaccurate system inventory as the availability signal materially understates the loss. Add model-drift monitoring for packaging redesigns, new SKUs, new retailers, lighting and camera changes, new fixtures, and seasonal displays; add safety metrics for unauthorized actions, incorrect retailer communication, privacy breaches, cross-retailer leakage, false inventory corrections, and approval bypass.
The programme should be judged on commercial value recovered through verified corrections, relative to field effort, technology cost, and retailer friction — not on the number of photographs analysed.
Five observations, four cases, 24 minutes
A beverage manufacturer, a large urban supermarket, carbonated soft drinks, during a national summer promotion. The Picture of Success: 12 mandatory SKUs, the hero 1.5-litre SKU at 4 facings, the zero-sugar 1.5-litre at 3 facings, the promotional multipack on display and shelf, promotional price €4.99, a brand block of at least 120 cm, and the summer campaign header installed. The representative captures three shelf images, one display image, one label image. Vision reports: hero SKU 0 facings, zero-sugar 2 facings, multipack present, display present but 45% competitor products, observed promotional price €5.99, header present. Traditional output: Perfect Store compliance 68%.
The agent instead retrieves context per observation and produces four cases with different owners and different urgency. On the hero SKU: inventory 22 units, 30 delivered yesterday, zero POS sales for six hours, expected velocity 18 units per day — a probable backroom replenishment or inventory-record discrepancy, €96 per day at risk, critical, actioned by a physical backroom check, refill to four facings if stock exists, a stock-count request if not, and an after photograph. On the promotional price: the retailer price file says €4.99, the shelf label says €5.99, the event is active — a store shelf-label execution failure, also critical, risking shopper response, promotion claims, and a retailer complaint; the fix is an authorized store employee replacing the label, with checkout verification where the process permits. On the display: agreed 100% manufacturer products against 55% observed and low stock — not maintained after setup, high priority, restore and refill where permitted and discuss with the department manager. On zero-sugar facings: 7 units in stock against 12 daily sales, with 24 units arriving tomorrow — low stock with a delivery due, medium, so keep the two facings, place no emergency order, and verify after delivery.
| Case | Diagnosis | Priority | Owner |
|---|---|---|---|
| Hero SKU shelf OOS | Backroom replenishment or record error | Critical | Rep + store manager |
| Promotional price mismatch | Store shelf-label execution failure | Critical | Store employee |
| Display contamination | Display not maintained after setup | High | Rep + department manager |
| Zero-sugar facings below target | Low stock, delivery due tomorrow | Medium | Monitor only |
The representative finds 20 hero-SKU units in the backroom and refills the shelf. The store employee replaces the price label. The display is restored from backroom promotional stock. After-action vision confirms the hero SKU at 4 facings, the price at €4.99, and the display at 94% manufacturer products — cases 1 to 3 verified, case 4 left in monitoring. Over the following three days hero-SKU sales resume, promotional multipack sales rise, the zero-sugar delivery arrives, and no repeat out-of-stock appears. Illustratively: €1,140 of sales recovered during the promotion, 24 minutes of representative time, two retailer actions, three issues closed.
Then the part that compounds. The system notices this store has had three backroom-related hero-SKU shelf gaps in six weeks and proposes a structural cause — replenishment frequency or inventory process — with a joint retailer store-process review as the recommended action. That is the difference between a programme that refills the same shelf forever and one that eventually stops having to. The figures are illustrative; a production system must use validated sales, substitution, margin, inventory, and execution models.
Implementation and readiness
- 01Phase 0 — define the Picture of Success. Store segments, objectives, KPIs, targets, critical gates, actions, ownership, evidence. Before any AI.
- 02Phase 1 — standardize retail-execution data. Store and product hierarchy, must-stock lists, planograms, price files, promotion calendar, visit history, issue taxonomy.
- 03Phase 2 — improve image capture. Capture guidance, quality validation, metadata, recapture, retailer-specific workflows.
- 04Phase 3 — deploy vision in measurement mode. Recognize products, count facings, detect gaps, compare planograms — with humans validating uncertain results.
- 05Phase 4 — add contextual diagnosis. Connect POS, inventory, orders, deliveries, promotions, assortment.
- 06Phase 5 — add prioritized recommendations. Store correction, order, retailer request, escalation — or no action.
- 07Phase 6 — approval-based execution. The agent creates tasks, order drafts, and escalation packets; humans authorize material external actions.
- 08Phase 7 — closed-loop verification. After evidence, system confirmation, outcome monitoring.
- 09Phase 8 — dynamic visit planning. Prioritize by expected issue, recoverable value, actionability, route economics.
- 10Phase 9 — structural root-cause management. Use repeated issues to fix replenishment, inventory accuracy, forecasting, assortment, planograms, and retailer processes.
A strong first pilot takes one country, one retailer, one category, 50–200 stores, 20–100 SKUs, stable planograms, reliable images, POS and inventory access, an engaged field team, and clear corrective actions — staged as offline image validation → rep-assisted measurement → shadow diagnosis → action recommendation → approval-based routing → closed-loop correction → dynamic visit optimization. Avoid starting with every category and retailer, no current product images, an inaccurate store hierarchy, unclear retailer permissions, no corrective-action process, or only a compliance dashboard. Success criteria: vision accuracy meeting category-specific thresholds, critical-OOS recall above the approved minimum, false-action rate below threshold, less time on manual auditing, better corrective-action closure, improved OSA and promotion compliance, and no unauthorized retailer action.
The minimum viable data is store master, product master, GTIN, product images, expected assortment, planogram, visit schedule, promotion calendar, representative identity, and shelf photographs; it strengthens with POS, inventory, orders, deliveries, retailer price files, shelf capacity, customer agreements, display commitments, route data, store labour, and remote shelf sensors. Two requirements deserve emphasis because they break vision systems quietly. Product-master reference images must track promotional and redesigned packaging, or current products become unknown products. And the price taxonomy must distinguish expected regular price, expected promotional price, observed shelf price, checkout price, unit price, and effective dates — since price OCR reads the shelf label, which is not the transactional truth. Volatile fields — inventory, price, promotion, order, delivery, planogram, assortment — all need timestamps, because a diagnosis built on a stale feed is a confident wrong answer.
Twenty failure modes
- 01One Perfect Store for every outlet. Standards ignore store mission and capacity.
- 02Compliance score without commercial priority. A missing poster gets the attention a hero-SKU OOS deserved.
- 03Treating computer vision as the whole solution. The system detects gaps it cannot resolve.
- 04Poor image capture. Most recognition errors originate in the field process, not the model.
- 05Outdated product catalogue. New packaging becomes an unknown product.
- 06Image absence equals inventory OOS. Backroom stock and misplacement are ignored.
- 07Positive inventory equals availability. Phantom inventory stays hidden.
- 08Assuming the planogram is current. The reference itself is wrong.
- 09Treating price OCR as transactional truth. The checkout system may differ.
- 10Every detection becomes a task. Field teams are overwhelmed and start ignoring the list.
- 11Assigning actions outside representative authority. Retailer relationships take the damage.
- 12Closing actions by checkbox. No proof exists that anything changed.
- 13Immediate correction hiding structural failure. The shelf is refilled forever; the reorder parameter is never fixed.
- 14Scheduling visits only by fixed frequency. High-value exceptions wait their turn.
- 15Measuring the field team only by visits. People optimize quantity instead of commercial outcomes.
- 16Perfect Store as employee surveillance. Trust and adoption collapse together.
- 17Manufacturer score ignoring retailer value. Recommendations create conflict instead of collaboration.
- 18AI inventing root causes. Diagnosis without evidence is worse than no diagnosis.
- 19Submitting orders or messages without authorization. Commercial control is bypassed.
- 20No causal evaluation. Sales changes get attributed to the programme with no valid comparison.
The PERFECT Method and maturity model
- 01Pin down the Picture of Success. Store segment, shopper mission, commercial objective, execution standards, critical KPIs, applicable period.
- 02Establish reliable store evidence. Photographs, planograms, POS, inventory, prices, promotions, orders, deliveries — combined, not chosen between.
- 03Recognize and rank execution gaps. Availability, assortment, planogram, price, promotion, display, condition — prioritized by business consequence.
- 04Find the root cause. Store execution, inventory, ordering, delivery, supply, master data, commercial setup.
- 05Execute the right corrective action. Immediate correction, retailer task, order, escalation, or structural fix — assigned to the owner who can actually perform it.
- 06Confirm closure and commercial effect. Corrected shelf, corrected system, resumed sales, completed order, persistent outcome.
- 07Turn execution outcomes into learning. Visit policy, image models, replenishment, assortment, planograms, promotion planning, retailer collaboration.
| Level | What it adds | Characteristics |
|---|---|---|
| 0 · Manual store audit | A record | Paper or spreadsheet survey, manual photos, limited follow-up, retrospective reporting |
| 1 · Digitized retail execution | Structure | Mobile visits, standardized questionnaires, geolocation, photos, dashboards, task tracking |
| 2 · Computer-vision measurement | Scale | Product recognition, facings, OOS detection, planogram comparison, price and display recognition |
| 3 · Contextual execution management | Diagnosis | POS, inventory, promotion and order context, root-cause diagnosis, prioritized actions |
| 4 · Agentic Perfect Store | Closed loop | Adaptive investigation, action recommendations, dynamic visit priorities, retailer workflows, proof-based closure, evaluation |
| 5 · Continuous execution operating system | Prevention | Always-on shelf monitoring, predictive OSA risk, event-driven field activity, upstream root-cause correction, closed-loop commercial learning |
Part XXIII in brief — the practitioner templates. The method ships as nine working documents: the Picture of Success card (segment, category, period, objective, mandatory assortment, hero SKUs, facings, position, share-of-shelf target, price, promotion, display and POS requirements, critical gates); the store observation card (expected versus observed condition, image ID, vision confidence, human validation, evidence timestamp); the OSA diagnosis card (expected and actual facings, store and backroom inventory, recent POS, open order, last delivery, assortment and promotion status, most likely and alternative causes, confidence, next check); the corrective-action card (impact, priority, action, owner, authority required, deadline, expected result, verification method, escalation); the display and price-compliance cards; the visit-priority card (sales potential, active promotion, open critical issues, OSA risk, launch activity, last visit, expected recoverable value, travel cost, recommended date, required skill); the decision packet; and the post-action learning card (initial observation, hypothesis, action taken, time to correction, verification, sales and compliance outcome, recurrence, confirmed root cause, structural recommendation, reviewer). The last two fields are the ones that turn a visit programme into an operating system.
Frequently asked questions
Is the Perfect Store Agent the same as shelf image recognition?
No. Image recognition identifies visible store conditions. The agent combines those conditions with commercial and operational data to determine what should happen next, who should do it, and how closure will be proven.
Can a photograph prove that a store has no inventory?
No. It can show that a product appears absent from the visible shelf. Inventory, backroom, order, and delivery evidence are all required before the absence becomes a diagnosis.
What is the difference between inventory availability and on-shelf availability?
Inventory may exist in the store or in the system while the product remains unavailable to the shopper on the shelf. Phantom inventory — stock the system reports but nobody can find — is the gap between the two.
Should every store use the same Picture of Success?
No. Standards should vary by store format, size, mission, commercial importance, available space, and campaign period. A single universal standard guarantees that most stores are measured against conditions that do not apply to them.
How should issues be prioritized?
By value at risk, urgency, actionability, strategic importance, confidence, and effort — not by compliance severity alone. A missing poster and a missing hero SKU are not the same problem.
Should the agent automatically submit an order?
Only where the company and the retailer have explicitly authorized that workflow and the action stays inside approved limits. Otherwise it should produce a draft or a recommendation.
How should closure be verified?
Through after photographs, inventory correction, order confirmation, POS resumption, delivery status, or retailer acknowledgement — never through a self-reported checkbox.
Should fixed visit frequencies disappear?
Not necessarily. Fixed frequencies provide stable coverage and retailer predictability, while event-driven visits supplement them for high-value exceptions. The two are complementary.
What are the main privacy risks?
Store images may capture shoppers, employees, personal information, geolocation, retailer-sensitive operations, and commercially sensitive data. The programme also must not become a surveillance mechanism aimed at the field team.
What is the biggest AI mistake in retail execution?
Deploying image recognition that produces compliance scores without connecting observations to root causes, corrective actions, ownership, verification, and business outcomes.
Conclusion
The Perfect Store has traditionally been run as a compliance programme: define standards, visit stores, collect surveys and photographs, report whether execution met expectations. That creates visibility, and visibility is worth having. But visibility alone does not put the product back on the shelf, correct the price, activate the promotion, restore the display, fix the inventory record, change the reorder parameter, or tell anyone whether the same problem returns tomorrow.
Computer vision transforms the speed and scale at which a store can be observed. It recognizes products, facings, gaps, prices, displays, and planogram deviations. What it cannot independently know is whether stock is in the backroom, whether an order failed, whether a delivery is late, whether the planogram is current, whether the retailer authorized the assortment, which action carries the most value, and who has the authority to act. That is the gap the agentic layer fills: it connects visual evidence with inventory, sales, orders, deliveries, promotions, prices, assortment, customer agreements, and field workflows — turning a gap into a decision case. What is missing? Is the observation reliable? Why is this happening? How much does it matter? Can it be corrected now? Who should act? What proof confirms closure? What structural change prevents recurrence? The PERFECT Method walks the route: pin down, establish, recognize, find, execute, confirm, turn into learning.
The future of retail execution is not more photographs, more surveys, or another compliance dashboard. It is a continuously learning execution system in which computer vision observes, deterministic systems calculate and transact, agents diagnose and coordinate, field teams and retailers act, and evidence verifies the outcome.
The defining question is not whether AI can analyse retail-execution photographs. It is whether the organization can convert store observations into the few corrective actions most likely to improve availability, execution, and commercial performance — and confirm that those actions worked.
Want help choosing the right architecture for your process?
We map where agents create leverage in FMCG operations, then build and ship the ones that pay back. One call to pressure-test your highest-leverage use case.