Skip to content
TiMiNa
All playbooks
Playbook · Demand planning

Beyond forecast accuracy: building an agentic demand-planning exception manager for FMCG

By Misagh Akhondzad/28 min read
Demand planningFMCGForecastingAgentic workflows

A demand planner opens the system on Monday. The portfolio holds 2,800 SKUs, 14 distribution locations, 22 customer groups, and 26 forecast periods — hundreds of thousands of forecast combinations. The weekly statistical run has finished. It displays 1,942 exceptions.

Some concern low-volume items with almost no financial impact. Some are duplicate alerts for the same underlying event. Some exist only because percentage errors explode when demand approaches zero. Some are for products already discontinued. Some reflect temporary out-of-stocks rather than changing demand. Some are caused by a promotion that is already known and correctly modelled. Some are mathematically unusual and operationally irrelevant, because the supply plan is already fixed. And some involve serious revenue and service risk — buried hundreds of rows down the list.

So the planner exports to Excel and adds what the system did not know: customer importance, product priority, inventory, lead time, promotion status, financial value, business knowledge. They message the key account manager, ask marketing whether a campaign is still running, ask supply whether production can still change, and compare the current forecast against several previous versions. One large exception turns out to be a missing price file. Another is retailer forward buying. A third is a product that has been unavailable for three weeks. A fourth is genuine demand acceleration that could stock out two customers. By Wednesday they have reviewed 127 exceptions. The other 1,815 get no attention at all.

The organization later reports that forecast accuracy improved by 1.2 percentage points. It cannot answer which interventions created the improvement, which made things worse, which exceptions mattered financially, which signals arrived too late to act on, which errors led to lost sales and which to excess inventory, which deviations were caused by data rather than demand, which commercial inputs are repeatedly wrong, which overrides added value, which products should no longer receive manual review — or how much planner time was spent on noise.

The planner’s scarce resource is not forecast rows. It is decision-making attention.

This is the real demand-planning problem. It is not that forecasts are inaccurate — forecasts will always carry uncertainty. It is that attention is finite and organizations allocate it badly, treating every forecast combination as though it deserves equal review. A 100% error on a product selling two units a month matters less than a 7% bias on a high-revenue core item. A sudden forecast increase may be entirely correct because a promotion was approved. A statistically stable forecast may be commercially dangerous because the retailer has just announced a delisting. An error may be impossible to correct because the production lead time has passed. And a forecast can be highly accurate yet still produce poor inventory outcomes, because uncertainty, service cost, lead time, and asymmetric business risk were never in the frame.

Which reframes the purpose of exception management. It is not to identify everything unusual. It is to find the relatively small number of demand situations where investigation or intervention creates meaningful business value — preventing a stockout, avoiding excess inventory, reducing obsolescence, correcting a systematic bias, improving a promotion forecast, catching a data defect, challenging an unsupported commercial override, accelerating a new-product response, triggering supply replanning, or stopping repetitive low-value manual work. Most planning platforms already support statistical forecasts, causal factors, demand sensing, error calculations, thresholds, overrides, hierarchies, and exception tables. What they mostly do not do is tell the planner why a rule was crossed, whether it matters, whether it is still actionable, which function owns the response, what evidence is missing, which intervention is appropriate, and whether that intervention later added value.

Part I · Foundations

What demand planning actually plans

Demand planning is the disciplined creation, interpretation, and governance of forward-looking demand assumptions so the organization can make better decisions about inventory, capacity, purchasing, production, deployment, finance, and customer service. It is broader than forecasting: forecasting estimates demand mathematically, while demand planning determines which forward-looking signal the organization should use — selecting the forecast, interpreting the model, incorporating valid business intelligence, handling exceptions, achieving cross-functional alignment, recording assumptions, and measuring performance. A forecasting model can be technically excellent while the demand-planning process around it is poor.

Four distinctions cause most of the damage when they blur. Demand plan versus sales forecast: a sales forecast may carry targets and commercial ambition; a demand plan should represent unconstrained expected demand at a defined planning level under stated assumptions. Unconstrained demand versus constrained sales: if demand is 100,000, only 80,000 units are available, and history records 80,000, the model learns the wrong number — and the missing 20,000 were lost, backordered, delayed, substituted, or cancelled, which are five different futures. Sell-in versus sell-out: shipments diverge from consumption through inventory loading, forward buying, customer stock reduction, pipeline fill, ordering cycles, and service constraints. Orders versus consumption: an order reflects batching, minimum quantities, inventory policy, speculation, and lead-time changes as much as it reflects demand. In FMCG the best signal usually combines streams — shipments, orders, sell-out, POS, depletions, e-commerce, customer forecasts — with the process defining which stream is used, at which horizon, for which decision, and how they reconcile.

Then there is the hierarchy problem. FMCG demand exists simultaneously across product (portfolio → category → brand → SKU → pack), customer (market → channel → customer → banner → store), location (network → region → plant → DC → depot), and time (year → quarter → month → week → day), and one plan may need to be coherent across all four. Accuracy changes with aggregation — at lower levels, zeros, intermittency, random variation, promotion effects, and availability distortion all increase — so the lowest available level is not automatically the best planning level. The forecast should be evaluated at the grain of the decision it serves.

DecisionRelevant planning grainTypical lead time
Annual financial outlookBrand · market · month12 months
Material purchaseComponent · network · month16 weeks
Production capacityProduct family · plant · week6 weeks
Customer promotionSKU · customer · week4–8 weeks
DeploymentSKU · DC · week2 weeks
Store replenishmentSKU · store · day2 days
Match the planning grain to the decision

Two more concepts complete the frame. Lag is how far in advance a forecast was made — a one-week-ahead forecast may be excellent while the eight-week-ahead forecast that production actually needs is poor, so accuracy must be measured at the lags that match decision lead times. And planning time divides into frozen (production or purchasing committed), slushy (changes possible but costly), and liquid zones, which should drive different intervention rules entirely. An exception is only useful if it arrives before the decision becomes irreversible. Finally, a mature system preserves every forecast version — statistical, sensed, planner-adjusted, sales input, consensus, approved plan, financial — because overwriting the statistical forecast destroys the ability to learn anything.

Part II · The limits of accuracy

Why forecast accuracy is not enough

Forecast performance drives inventory, service, production, purchasing, labour, logistics, working capital, obsolescence, and financial planning. It is necessary. It is also not the outcome. A forecast can become more accurate while producing less business value: accuracy improves on low-value tail items while core-item bias worsens; the forecast improves after supply decisions are fixed; the point estimate sharpens while uncertainty stays badly calibrated; planners spend hours for a trivial gain; an override improves WAPE and increases stockout cost; a conservative forecast lowers average error and causes repeated shortages; a highly accurate aggregate forecast allocates demand incorrectly across locations. Recent field research on forecast overrides is explicit that forecast-performance improvement does not necessarily translate into improved profit.

The honest evaluation question is whether the forecast and the planning process improved the decisions they exist to support. That makes planning value a compound of forecast accuracy plus uncertainty calibration plus decision timeliness plus economic asymmetry plus operational actionability minus planner effort. Error costs are asymmetric: underforecasting creates lost sales, low service, penalties, expediting, missed launches; overforecasting creates inventory, write-offs, waste, working-capital cost, markdowns, warehouse congestion. For perishables, overforecasting is punishing; for high-margin critical items, underforecasting is worse. Which means the preferred plan may deliberately differ from the statistical mean, depending on margin, service target, shelf life, lead time, substitution, shortage cost, and inventory cost.

Part III · Metrics

Forecast metrics for practitioners

Error is actual minus forecast — under that convention, positive means underforecast. Organizations sometimes use the opposite sign. Define it explicitly and in writing, because half the arguments about bias in demand reviews are actually arguments about sign conventions. From there, each metric answers a different question and fails in a different way.

MetricWhat it tells youWhere it fails
ME · mean errorBias directionPositive and negative errors cancel; not an accuracy measure
MAEAverage error in original unitsNot comparable across products of different scale
RMSEError with large misses penalized moreScale-dependent; dominated by outliers
MAPEPercentage error, intuitive to explainUndefined or explosive at zero and near-zero demand
WAPEPortfolio error weighted by volumeHides strategic low-volume items; no direction
MASEPerformance against a naïve benchmarkNeeds a well-chosen benchmark; less intuitive
RMSSEScaled error penalizing large missesSame benchmark dependency; used in M5
BiasSystematic over- or underforecastingMust be read alongside an absolute measure
The practitioner's metric reference

Two disciplines matter more than the choice of metric. First, absolute accuracy and bias belong together: forecast A at 20% WAPE with 0% bias and forecast B at 20% WAPE with +18% bias are not the same forecast, and only one of them is systematically building inventory. A tracking signal comparing cumulative error to mean absolute deviation helps surface persistent bias, with calibrated thresholds. Second, weight error by consequence: economic forecast error — error × estimated unit consequence — is far more useful for prioritization than raw percentages, because it is the only version that ranks a 7% miss on a core SKU above a 100% miss on a tail item.

Then measure at the right lag (one-week, four-week, eight-week, thirteen-week, tied to the decisions each supports — SAP planning documentation supports error calculation at defined lags and over lead-time intervals precisely because error feeds safety stock) and the right aggregation (operational, managerial, financial — never claim success from high aggregate accuracy when lower-level allocation drives execution). For intermittent demand, MAPE is unsuitable and point accuracy alone is insufficient: evaluate occurrence and size separately, and lean on MASE, RMSSE, probability calibration, pinball loss, and the resulting stockout and inventory cost.

A point forecast is one summary of many possible futures

Two products with the same 10,000-unit point forecast can carry completely different risk — one plausibly 9,500–10,500, the other 5,000–18,000 — and they should not produce the same inventory decision or the same exception. A probabilistic forecast gives quantiles (P10 8,000 / P50 10,000 / P90 14,000) and prediction intervals, which is what safety stock, service levels, scenario planning, risk prioritization, capacity decisions, and financial ranges actually consume. Quantiles also handle asymmetry directly: high-service items plan to a higher quantile, perishables to a lower one. Judge the distribution on calibration (if the 80% interval contains actual demand half the time, uncertainty is understated) and sharpness (a calibrated but enormous interval is useless). And note that uncertainty generates its own exceptions: the interval widened sharply, downside risk grew, upside exceeds capacity, the interval crossed a service threshold, model confidence deteriorated — all while the point forecast never moved.

Part IV · Forecast Value Added

Does each step in the process earn its place?

Forecast Value Added measures whether a step in the forecasting process improves the forecast relative to the input it received: error before the step minus error after it. Positive means the step created value; negative means it destroyed some. Every step in a demand-planning chain — naïve forecast, statistical forecast, demand sensing, planner override, sales adjustment, consensus plan — consumes time, software, meetings, and organizational attention. FVA asks the question those meetings never ask of themselves.

28%Naïveforecast19%+9Statisticalforecast18%+1Demand-sensed21%-3Planner-adjusted23%-2Consensusforecastillustrative WAPE — lower is better; the label above each column is value added or destroyed
The FVA ladder: which process steps actually earn their place

Aggregate FVA hides the pattern that makes it actionable. Broken out by segment, the same process might show planner adjustments adding six points on promotions and three on new products while destroying two on core baseline and eight on the long tail — which is not an argument against planner judgment but a precise map of where it belongs. FVA can also be measured by planner, team, business unit, override reason, and horizon; it should be used for process improvement, never as a public league table, because punitive FVA simply teaches planners to hide adjustments or to avoid judgment they should be exercising.

Applied to the process itself, FVA becomes a design tool: evaluate model selection, demand sensing, planner review, sales review, management review, and the consensus meeting, and redesign or remove any step that repeatedly destroys value. Segments with consistently negative manual FVA graduate to no-touch forecasting, where the planner monitors only material exceptions — which is how attention gets reallocated to where it demonstrably pays.

Part V · Overrides

Disciplined human judgment

Planners override forecasts for good reasons — a future promotion, a customer listing or delisting, a price change, competitor action, a known order, a capacity transition, new distribution, supply recovery, one-time demand — and for weak ones: target pressure, “sales feels optimistic,” reaction to a single data point, anchoring on last month, avoiding a management challenge, compensating for an unrelated supply constraint, or simple habit. The most consequential confusion is the first: a target expresses ambition, a forecast expresses expectation, and the gap between them is exactly what action is supposed to close. Force them to match and the supply plan quietly becomes aspirational.

Research repeatedly finds asymmetric override behaviour — commercial participants adjust upward more readily, influenced by optimism, targets, incentives, planned initiatives, and fear of shortages. That does not make upward overrides wrong; it makes them worth evidence. Discipline comes from four requirements. An override needs a reason code and evidence, graded honestly as confirmed (a signed promotion, an approved listing, a customer order), probable (a reliable customer indication, a repeated signal), or speculative (a sales expectation, an unconfirmed rumour). It needs an explicit expiry — one week, an event window, several months, or a permanent baseline change — because temporary events with no decay permanently distort the baseline. It needs a defined allocation when made above SKU level, since the method used to push a category adjustment down to brand, SKU, customer, and location changes the answer. And it needs approval thresholds tied to value, supply consequence, size, horizon, product risk, and evidence quality — not every override, but the material ones.

Parts VI–VII · Exceptions

From raw alerts to decision cases

A demand-planning exception is a condition indicating that the plan, its drivers, its data, or its operational consequences require review. It is not necessarily an error — it is a signal that the normal flow may be insufficient. Every exception should answer three questions: What changed? Why does it matter? Can we still do something? Traditional systems answer only the first. They compute measure + threshold + hierarchy level = alert, which is useful and generates overload, because static thresholds treat a 20% change on a €10,000 product like a 20% change on a €10 million one, one root cause produces several alerts, the system knows nothing of margin, customer importance, supply flexibility, promotions, or lifecycle, the planner must investigate manually, the alert may arrive after the action window has closed, and nothing measures whether the intervention helped.

Raw alerts (four rows, one root cause):
  forecast change +35%
  MAPE above threshold
  customer order spike
  inventory risk flag

Consolidated case:
  Core detergent SKU at Retailer A shows a probable
  unmodelled promotion beginning Week 34. Forecast is
  28% below expected sell-out. Supply plan can still
  be changed for five days. Revenue at risk: EUR 420,000.
  Recommended: verify retailer promotion, then adjust
  event demand after commercial confirmation.
The same signal, before and after consolidation

The lifecycle of a real case runs detect → consolidate → diagnose → prioritize → investigate → decide → act → monitor → learn, and every case needs an owner, supporting functions, a decision deadline, a status, and a next action. Statuses should be explicit — new, triaged, investigating, waiting for input, recommendation ready, approved, actioned, monitoring, resolved, closed with no action, escalated — because “closed with no action” is a legitimate and frequently correct outcome that most alerting systems cannot even express.

The twelve exception families

  • Forecast performance. High error, high bias, deteriorating accuracy, poor FVA, model instability, widening uncertainty.
  • Demand change. Sudden increase or decline, structural break, trend acceleration, unusual order pattern, changed customer forecast.
  • Commercial events. Missing or moved promotion, listing, delisting, price change, customer win or loss, media campaign, competitor event.
  • Data quality. Missing actuals, duplicate demand, wrong price, broken hierarchy, missing promotion flag, invalid unit conversion, corrupted history, late feeds.
  • Availability and lost sales. Decline during a stockout, suppressed baseline, phantom decline, constrained sales mistaken for demand.
  • Lifecycle. Launch, ramp, maturity change, decline, discontinuation, replacement, cannibalization from innovation.
  • Model. Loss of fit, changed driver relationships, residual patterns, interval miscalibration, changed model selection, missing features.
  • Override. Large, unsupported, or repeated overrides; positive-bias patterns; overrides after freeze; deteriorating override FVA.
  • Hierarchy. SKU inconsistent with brand, customer allocations that do not reconcile, conflicting location totals, top-down versus bottom-up divergence.
  • Supply impact. Demand above capacity, inventory below risk threshold, projected excess, changes inside the frozen horizon, service and obsolescence risk.
  • Financial. Revenue deviation, margin risk, budget gap, working-capital impact, write-off exposure.
  • Process. Overdue approvals, missing commercial input, unresolved consensus gaps, repeated exceptions, stale assumptions, no owner.
Part VIII · Prioritization

Allocating scarce attention

Not every exception deserves human review. Priority should approximate expected decision value across seven dimensions. Business impact: revenue, margin, service, inventory, write-off, customer importance, strategic products, production disruption. Urgency: decision deadline, production and replenishment lead times, promotion start, customer commitments, shelf life. Actionability: can an intervention still change the outcome — production changeable (high), deployment changeable (medium), event already past (low)? Uncertainty relevance: attention pays where the outcome range is wide and the decision genuinely differs across scenarios; if every plausible outcome leads to the same action, further analysis is waste. Diagnosis confidence: low confidence may itself justify investigation. Strategic importance: top customer, launch, regulated or constrained product, hero brand, seasonal event. And repetition: a small recurring exception often signals a structural process failure — sales always overrides upward, one customer always sends late forecasts, one interface always fails.

In practice this becomes a bounded score (financial impact 0–5, service impact 0–5, urgency 0–5, actionability 0–5, strategic importance 0–3, evidence confidence 0–3) with two mechanisms sitting outside it. Hard escalation rules bypass scoring entirely: a probable stockout at a top customer, a major launch with a missing forecast, demand above capacity, a regulatory or safety issue, a data failure affecting a large portfolio, an unapproved override above a financial threshold. Suppression rules work the other way, silencing or grouping low-value items, duplicates, already-modelled promotions, discontinued items, non-actionable historical exceptions, noise inside tolerance, and expected seasonal moves. The result is an explicit attention budget — perhaps twenty critical cases for deep investigation, fifty for fast review, and an automated tier that runs no-touch unless a hard rule fires. That is a planning policy, not a report setting, and it should be reviewed like one.

Part IX · Diagnosis

Diagnose before you override

A forecast deviation can originate in any of eight layers, and the order of investigation matters enormously — because the first two masquerade convincingly as changed demand. If the data is invalid, nothing should be changed on the basis of it. If demand was constrained by a stockout, allocation, capacity, order rejection, delivery failure, or store closure, then observed sales are not demand, and lowering the baseline is precisely the wrong response.

1. Is the data valid?Fix the feed — do not touch the forecast2. Was demand constrained?Stockout: estimate lost sales3. Was there a known event?Promotion, price, listing, weather4. Did ordering behaviour change?Forward buying, not consumption5. Is it lifecycle-driven?Launch ramp, decline, cannibalization6. Did the market shift?Structural break: rebuild the baseline7. Did the model fail?Missing driver or stale parameters8. Otherwise it is noise — no action
Diagnose before you override: the ordered root-cause checks

The output of that walk is not a verdict but an evidence packet: most likely cause, alternative causes, evidence supporting each, evidence still missing, confidence, and the recommended next check. Two failure modes are worth naming. The agent must not conclude that weather caused demand growth merely because demand and temperature moved together — correlation requires approved causal or diagnostic methods before it becomes a stated cause. And model explainability is not operational diagnosis: “the promotion flag contributed +15%” explains the model, while “was the promotion correctly configured, executed, and reflected in sell-out?” explains the business. Both are useful; only the second tells anyone what to do.

Part X · Playbooks

The recurring FMCG exceptions

Promotion exceptions trigger when a promotion is not represented, uplift falls outside the expected range, dates move, events overlap, orders exceed the event forecast, or the post-promotion dip is unmodelled. The checks are retailer, SKU, mechanic, discount, feature and display, store coverage, trade terms, inventory, comparable events, and execution probability — and the correct action may be to add or update the event, adjust uplift, update coverage, request commercial confirmation, trigger a supply check, or do nothing because it is already modelled.

Customer-order spikes are the classic trap: the cause may be genuine sell-out acceleration, forward buying, inventory loading, a duplicated order, distribution expansion, a promotion, order batching, or panic ordering — so the evidence required is sell-out, customer inventory, historical ordering pattern, the promotion calendar, order status, and customer communication. Never convert an order spike into baseline demand automatically. The mirror image, a sudden decline, may be real demand loss — or a stockout, a delisting, a price increase, a competitor promotion, a data failure, store closures, or retailer inventory reduction.

Persistent bias in either direction is a structural signal, not an override opportunity: persistent underforecast points to recalibrating the baseline, adding a missing causal driver, or reviewing distribution; persistent overforecast points to optimistic sales adjustments, lost distribution, unmodelled cannibalization, launch overestimation, a target sitting inside the forecast, or an obsolete promotion assumption. In neither case should a permanent override be applied before the structural cause is understood.

New products need decomposition rather than accuracy targets: launch demand = pipeline fill + store inventory + shopper trial + repeat + promotional loading, and confusing those components makes every subsequent read wrong. The forecast should move through phases — a pre-launch analogue, an initial shipment forecast, early sell-out sensing, a distribution-adjusted forecast, and a repeat-informed steady state — updated continuously as distribution, rate of sale, trial, repeat, availability, media, and retailer execution arrive. End-of-life inverts the objective entirely: the goal is controlled depletion, substitution, write-off avoidance, customer communication, and transition to a replacement — history should stop being extrapolated.

Intermittent and long-tail demand should be governed by policy rather than reviewed item by item. Focus on occurrence probability, demand-size distribution, lead-time demand, and inventory consequence; manually reviewing every zero-to-positive change manufactures noise. ABC-XYZ segmentation (value against variability) is a reasonable starting frame — AX high-value and stable runs automated, AZ high-value and volatile gets close exception management, CX automates, CZ takes a simplified policy — but it is crude on its own and improves with lifecycle, margin, lead time, service, shelf life, strategic importance, and substitution.

Parts XI–XII · Sensing and consensus

Near-term signals and cross-functional inputs

Demand sensing updates the near-term view using recent sell-out, orders, inventory, distribution, promotions, weather, online activity, and short-term trends. It does not replace the mid-term plan; it improves near-term responsiveness, and it is only useful when its horizon aligns with replenishment, deployment, production flexibility, and customer ordering — a one-week signal cannot help a twelve-week manufacturing decision. Sensing generates its own exceptions (sensed demand diverging from consensus, near-term order acceleration, a sell-out drop, promotion execution variance, availability change) and its own risk: noise amplification, where the system overreacts to one-day spikes, order timing, data latency, promotions, or stockouts. Diagnose the signal before escalating it.

Consensus planningcombines the statistical model, the planner, sales, marketing, finance, customer collaboration, and leadership into one agreed plan — and consensus does not mean averaging opinions. A good process separates evidence from assumptions from targets from risks from commitments. Its failure modes are predictable: the target replaces expected demand, the loudest stakeholder wins, upward adjustments go unchallenged, source versions are not preserved, meeting time goes to low-value items, assumptions go undocumented, and adjustments never expire. The agent’s contribution is preparation and memory — surfacing material gaps, identifying conflicting assumptions, showing historical FVA by source, keeping forecast and target visibly separate, summarizing decisions, recording ownership, and tracking unresolved risks.

Sales forecast:        120,000
Statistical forecast:   95,000
Finance forecast:      100,000
Gap:                    25,000

Commercial evidence:   unconfirmed customer listing

Recommendation:  retain the 95,000 base plan,
                 add 25,000 as a conditional scenario,
                 escalate listing confirmation.
A consensus gap, handled properly
Parts XIII–XIV · The workflow

The agentic exception-management workflow

Demand-planning system
  -> exception detection services
  -> Agentic Exception Manager
  -> demand, commercial, inventory, supply,
     financial, data-quality, and model tools
  -> deterministic diagnostic + forecasting services
  -> planner recommendation
  -> human decision or bounded automation
  -> approved forecast update
  -> supply and inventory response
  -> outcome monitoring and FVA
The target architecture: a governed hybrid

The agent collects raw alerts, consolidates related ones, retrieves context, diagnoses likely causes, calculates business priority, identifies missing evidence, assigns owners, recommends action, explains uncertainty, monitors resolution, and measures intervention value. Governed software owns forecasting, error calculation, financial calculation, inventory projection, supply simulation, prioritization rules, policy enforcement, authorization, the forecast update itself, and audit. Humans own judgment under ambiguity, customer interpretation, commercial commitment, high-impact overrides, cross-functional trade-offs, escalation, and accountability. Functionally the system needs a detector, consolidator, diagnostician, prioritizer, investigator, recommender, workflow coordinator, evaluator, and learning manager — which may be one agent with tools or several specialized components. Start with one coordinated agent unless separation creates measurable value.

The sixteen stages

  1. 01Ingest forecast versions. Statistical, sensed, planner, sales input, consensus, prior versions, actuals.
  2. 02Generate raw signals. From error, bias, demand change, orders, promotions, inventory, supply, commercial events, data quality, model monitoring.
  3. 03Consolidate signals. Group by product, customer, location, event, root-cause hypothesis, and time window.
  4. 04Enrich business context. Revenue, margin, customer importance, lifecycle, lead time, inventory, service, capacity, promotion, FVA history.
  5. 05Check data integrity. Validate actuals, confirm units, check missing feeds, inspect availability, verify hierarchy — before any diagnosis.
  6. 06Diagnose. Likely causes, competing explanations, evidence, confidence.
  7. 07Calculate actionability. Time to decision, freeze status, supply flexibility, inventory, customer commitment.
  8. 08Estimate business impact. Revenue and margin at risk, inventory and service exposure, capacity implications, write-off risk.
  9. 09Prioritize. Critical, high, medium, low, suppressed.
  10. 10Recommend a response. Adjust forecast, create a temporary override, update an event driver, correct data, request confirmation, trigger a supply scenario, escalate, monitor — or take no action.
  11. 11Human review. Concise diagnosis, evidence, impact, recommendation, alternative, uncertainty, deadline.
  12. 12Apply the approved action. Through an authorized tool, never directly.
  13. 13Propagate implications. Supply planning, inventory planning, customer service, finance, commercial.
  14. 14Monitor outcome. Actual demand, inventory, service, supply response, forecast performance.
  15. 15Calculate FVA. Original forecast versus adjusted forecast versus actual versus decision outcome.
  16. 16Close and learn. Root cause, decision, outcome, lesson, automation candidate.
Part XV · Toolset

The agent’s twenty-two tools

  • Forecast and actuals: get_forecast_versions (statistical, sensed, planner, sales input, consensus, approved plan, with creation lag and version), get_actual_demand (sales, orders, shipments, sell-out, adjusted demand, lost-sales estimate).
  • Measurement: calculate_forecast_metrics (MAE, WAPE, MAPE where valid, MASE, RMSSE, bias, lag performance), calculate_fva (version against version), get_forecast_distribution (quantiles, intervals, calibration, uncertainty drivers).
  • Detection and integrity: detect_structural_change (level shifts, trend changes, volatility changes, residual drift), check_data_quality (missing data, duplicates, unit anomalies, hierarchy failures, late feeds, invalid values), check_availability_and_lost_sales (inventory, stockout periods, service, constrained demand, lost-sales range).
  • Commercial context: get_commercial_events (promotions, price changes, listings, delistings, marketing, customer events), get_customer_order_context (orders, cancellations, cadence, customer inventory, forward-buying signal), get_promotion_forecast (baseline, uplift, mechanic, coverage, event demand, post-event effect), get_market_drivers (approved category trend, weather signal, macro driver, competitive observation).
  • Lifecycle: get_product_lifecycle (launch, growth, mature, decline, end-of-life, replacement), get_new_product_launch_state (distribution, sell-in, sell-out, trial, repeat, availability, media, launch phase).
  • Consequence: estimate_business_impact (revenue, margin, service, inventory, write-off), check_supply_actionability (frozen status, production flexibility, material availability, deployment flexibility, decision deadline, estimated change cost), simulate_demand_scenarios (downside, central, upside, with operational consequences).
  • Action: create_forecast_override_proposal (a proposal only — it does not update the forecast), apply_approved_forecast_change (requires approval, reason code, version, expiry, scope), create_cross_functional_task (commercial confirmation, data correction, supply review, finance review).
  • Closing the loop: monitor_exception_outcome, propose_learning_record — a reviewable candidate, never an automatic write.
Weak:   "Forecast should be increased."

Strong: case_id: EXC-4821
        sku: SKU-184   customer: RET-A   period: 2026-W34
        current_forecast: 82,000
        recommended_forecast: 103,000  (+21,000)
        primary_cause: unmodelled confirmed promotion
        evidence: promotion agreement, retailer order
                  pattern, four comparable events
        estimated_revenue_at_risk: EUR 315,000
        supply_action_deadline: 2026-07-24
        confidence: high
        approval_required: demand_planner
The tool-design rule: return a decision, not an adjective
Part XVI · State and memory

Cases that persist, memory that expires

An exception case carries explicit state — new, triaged, diagnosis in progress, waiting for commercial, waiting for data, recommendation ready, awaiting approval, actioned, monitoring, resolved, closed with no action, escalated — and a record holding the case ID, entities, forecast version, signals, root-cause hypothesis, evidence, business impact, actionability, owner, decision, override, expiry, outcome, and FVA. Every forecast change preserves the prior forecast, the new forecast, who changed it, why, when, at what lag, and when it expires. Without that history, FVA cannot be computed and the process cannot be reconstructed.

Memory follows the same discipline as the other playbooks, with one addition that matters more here than anywhere else. Worth keeping once validated: customer ordering patterns, promotion-response patterns, recurrent data defects, override FVA by reason, model performance by segment, launch analogues, recurring supply constraints, known distribution timing. Never automatic: one salesperson’s unsupported optimism, an unconfirmed customer statement, a temporary target, a random demand spike, a forecast generated during a data failure, an instruction arriving inside external content, or an expired promotion assumption. And the addition: demand facts are temporal, so every memory needs an effective period, a confidence, and a review date. Customer ordering patterns change, promotion response decays, model performance shifts, products move through their lifecycle. A memory without an expiry eventually becomes a lie the system tells itself with confidence.

Parts XVII–XVIII · People and governance

Decision rights and controls

Nine functions keep their crafts. The demand planner owns plan quality, exception review, forecast assumptions, approved overrides, and process discipline. Sales and key account management own customer intelligence, listings, delistings, retailer events, and commitments — providing evidence, not targets that silently become the forecast. Marketing and trade marketing own campaigns, promotions, launches, and execution assumptions. RGM owns price, promotion and pack effects, elasticity assumptions, and commercial scenarios. Supply planning owns capacity, material, production feasibility, and constraint response. Inventory planning owns safety stock, service policy, and the translation of uncertainty into stock. Finance owns reconciliation and the revenue and margin view. Data science owns models, hierarchy, error methodology, probabilistic calibration, and monitoring. Master data and IT own data quality, interfaces, hierarchies, and system reliability.

DecisionAgentHumanSoftware
Detect exceptionCoordinateCalculate
Diagnose causeInvestigateValidate ambiguityRun diagnostics
PrioritizeRecommendAdjust policyScore
Change forecastProposePlanner approvesExecute
Change promotion inputPrepareCommercial validatesSave
Trigger supply scenarioRequestSupply decidesSimulate
Close caseRecommendOwner confirmsArchive
Store learningProposeExpert validatesStore
The decision-rights matrix

The governance rules are short and load-bearing. The agent does not own the demand plan — it investigates, recommends, and coordinates while authorized planners own changes. The system visibly preserves demand forecast, commercial target, financial target, and supply capability as four separate things, and never silently merges them. Override authority scales with percentage change, financial value, horizon, product class, supply impact, and reason. Inside the frozen horizon, demand changes may still be recorded, but supply changes require specific authority and the execution implications must be stated explicitly rather than assumed. Material changes require separation of duties: commercial provides input, the planner validates demand impact, supply assesses feasibility, and an authorized owner approves. Data access is scoped to relevant products, customers, commercial information, and financial fields. And the approval packet must carry the exception summary, current and proposed forecast, demand range, root cause, evidence, financial impact, inventory and service impact, supply actionability, alternatives, override duration, the required decision, and expiry.

Part XIX · Evaluation

Evaluating the exception manager

The agent can prioritize noise, miss a high-value issue, diagnose a stockout as demand decline, double-count a promotion, use stale commercial information, recommend a target as a forecast, create excessive overrides, ignore supply lead time, falsely claim FVA, or expose sensitive customer data. Evaluation therefore runs layer by layer. Component tests cover version retrieval, metrics, bias, FVA, event retrieval, lost-sales adjustment, lifecycle, business-impact calculation, supply actionability, and policy. Detection is measured on precision, recall, severe-exception recall, false alerts, duplicates, and the risk carried by suppression — and severe-exception recall is the one number that should never be traded for a cleaner queue. Prioritization asks whether the most valuable cases ranked highest and whether urgency, actionability, and suppression were right. Diagnosis is measured on root-cause accuracy, evidence completeness, calibrated confidence, and the usefulness of the suggested next step. Recommendation is measured on correct action and correct no-action, change magnitude, expiry, owner, and policy compliance.

Trajectory tests require the agent to check data quality, check availability, retrieve commercial events, estimate business impact, check actionability, and request authorization — while prohibiting overrides without approval, targets converted into demand, invented commercial events, ignored stockouts, and silent updates to a frozen plan. Outcome tests verify that the approved forecast was updated exactly once, the correct version preserved, supply implications routed, the case monitored, and FVA calculated afterwards. The metric families then span business (service level, lost sales, inventory, write-off, expedites, working capital, cost per exception), forecast (WAPE, MASE, RMSSE, bias, interval coverage, accuracy by lag, FVA), process (raw alerts versus consolidated cases, cases reviewed, resolution time, planner hours, exception recurrence, no-touch percentage), override (frequency, direction, FVA by reason and planner, expired and unsupported overrides), and safety (unauthorized changes, cross-customer leakage, unsupported commercial claims, target contamination, approval bypass, stale evidence).

The metric that captures the whole discipline: business value protected per planner hour spent.

Part XX · Worked example

A 132,000-unit order that is not demand

A national beverage company, one-litre family juice, a major grocery retailer, weeks 31–34. The statistical forecast is 100,000 units per week, the planner-adjusted forecast 104,000, the consensus 105,000. Then the signals arrive: retailer orders of 132,000 units for week 32, sell-out growth of +4%, retailer inventory up 22%, no promotion on record, supply availability of 120,000 units, and a production-change deadline five days out. The traditional exception reads: orders exceed forecast by 25.7%.

The agent retrieves historical retailer ordering, retailer inventory, sell-out, the promotion calendar, customer communication, comparable events, and supply lead time — and finds that the retailer has a two-week promotion scheduled, approved by email but never entered into TPM, and has begun loading inventory. Comparable events show sell-in rising 28–35%, sell-out rising 16–22%, and a post-event dip of 8–12%. So the order spike decomposes into promotion sell-out demand + retailer inventory loading + normal demand. The 132,000-unit order should not become the new baseline.

1 · No adjustment2 · Match the order3 · Event decomposition
Demand plan, W32105,000132,000121,000
Shipment expectation105,000132,000132,000 (held separately)
W33 post-event105,000~132,00095,000
Primary riskPromotion underforecast; ~88% serviceForward buying read as consumptionRequires promotion confirmation
ConsequenceLost salesPost-event overforecast, excess stockBaseline protected
Three ways to respond to the same signal

The recommendation is scenario 3, expressed as five concrete actions: add the confirmed promotion event; set unconstrained demand at 121,000 for week 32; maintain a separate shipment signal at 132,000; set the post-promotion forecast to 95,000 for week 33; trigger a supply scenario for 121,000 with customer-allocation review — and do not permanently adjust the baseline. The business case: €180,000 of revenue at risk with no action, 11,000 units of excess inventory risk if the order is treated as demand, five days to the production decision, priority critical. The planner approves the event adjustment, the account manager confirms the promotion, and supply approves an incremental run of 12,000 units, covering the remainder from stock.

The outcome: week 32 sell-out 119,000, week 32 shipments 131,000, week 33 sell-out 94,000. Statistical forecast error 19,000 units; consensus error 14,000; the agent-supported event forecast error 2,000 units. The intervention added forecast value — and, just as importantly, kept shipment loading and consumption apart. A weaker agent would have matched the customer order, raised the baseline permanently, ignored retailer inventory, omitted the post-promotion dip, called the spike a demand-sensing signal, or updated the forecast without approval. Every one of those is a plausible-looking action that reduces business value. The numbers are illustrative; a production system requires current company data and validated models.

Parts XXI–XXII · Roadmap and data

Implementation and data readiness

  1. 01Phase 0 — establish planning definitions. Demand, sales, orders, shipments, unconstrained demand, forecast versions, lag, accuracy, bias, FVA, override, exception.
  2. 02Phase 1 — preserve forecast history. Versions, creation dates, lags, adjustments, actuals, reason codes. Without historical versions, FVA cannot be measured at all.
  3. 03Phase 2 — clean the exception rules. Before adding an agent: remove duplicate alerts, segment thresholds, retire obsolete rules, define hard escalations, correct the hierarchy.
  4. 04Phase 3 — build business enrichment. Connect value, customer, lifecycle, inventory, lead time, promotion, supply, and service to the signal.
  5. 05Phase 4 — diagnostic copilot. The agent consolidates alerts, retrieves evidence, suggests causes, and prepares cases. Humans make every change.
  6. 06Phase 5 — recommendation mode. The agent recommends adjust, no action, monitor, or escalate.
  7. 07Phase 6 — approval-based forecast changes. The agent prepares structured changes, the planner approves, the tool executes.
  8. 08Phase 7 — FVA-driven no-touch planning. Segments with negative human FVA move toward automation; planners concentrate on high-value exceptions.
  9. 09Phase 8 — cross-functional orchestration. The agent routes promotion confirmation, data correction, supply scenarios, and financial review.
  10. 10Phase 9 — policy-bounded administration. Automatically close duplicates, correct approved data defects, expire temporary overrides, create tasks, update confirmed event metadata. Material demand changes stay governed.

A strong pilot takes one market, one planning team, one category, 200–1,000 SKUs, reliable forecast history, promotion data, inventory and supply context, and measurable planner workload. Avoid starting where forecast-version history is missing, demand definitions are unclear, there is no planner ownership, reason codes do not exist, actuals are incomplete, source systems are fragmented, or the scope is every product and market at once. Success criteria go in writing first: reduce raw alerts by 70% through consolidation and suppression while maintaining 100% recall on defined critical exceptions, halve exception-review time, improve planner FVA, reduce persistent bias, cut avoidable stockouts and excess inventory, and achieve zero unauthorized forecast changes.

The minimum viable data is actual demand, forecast versions with creation dates, product and customer hierarchy, location, promotion, price, distribution, inventory, supply lead time, planner overrides, and reason codes; it strengthens considerably with sell-out, customer inventory, availability, lost-sales estimates, marketing events, weather, competitor events, orders, shipment status, supply flexibility, margin, and service cost. Four foundations decide whether the rest works. Forecast-version history must carry forecast ID, version, creation date, target period, lag, source, value, override reason, and owner. Adjusted actuals must handle stockouts, constrained sales, returns, pipeline fill, one-time orders, and system errors — under governance, not planner discretion. An event taxonomy must standardize promotion, price change, listing, delisting, launch, media, distribution expansion, customer event, weather event, and supply recovery. And a reason-code taxonomymust be specific enough that “other” never becomes the dominant category — because the moment it does, override FVA analysis becomes unreadable.

Part XXIII · Reality checks

Twenty failure modes

  1. 01More alerts equal better control. The planner is overwhelmed and reviews 7% of them.
  2. 02Percentage-error thresholds on low volume. Noise dominates the queue.
  3. 03Prioritizing by accuracy alone. Financial and service impact never enter the ranking.
  4. 04Ignoring actionability. The alert arrives after the decision deadline.
  5. 05Treating sales as unconstrained demand. Stockouts quietly lower the forecast.
  6. 06Treating orders as consumption. Forward buying raises the baseline.
  7. 07Mixing target and forecast. Supply plans become aspirational.
  8. 08Unlimited overrides. Every forecast becomes a negotiation.
  9. 09No expiry on overrides. Temporary events permanently distort the baseline.
  10. 10Measuring only aggregate accuracy. Location and customer allocation failures stay hidden.
  11. 11Using MAPE universally. Zero and low demand manufacture misleading exceptions.
  12. 12Ignoring bias. Positive and negative errors cancel into false comfort.
  13. 13Ignoring uncertainty. A point forecast looks more certain than it is.
  14. 14Measuring FVA without preserving versions. The process cannot be reconstructed.
  15. 15Punitive FVA. Planners hide adjustments or stop exercising necessary judgment.
  16. 16The agent changes forecasts directly. Human authority and audit are bypassed.
  17. 17The language model performs statistical calculations. Use governed analytical services.
  18. 18The agent invents root causes. Diagnosis unmoored from evidence is worse than no diagnosis.
  19. 19Every exception becomes an override. The right response is often a data fix, an event update, a supply action, or nothing.
  20. 20No business-outcome measurement. Accuracy improves; service and inventory do not.
Parts XXIV–XXV · Framework

The EXCEPTION Method and maturity model

  1. 01Establish the planning decision. Demand signal, planning grain, horizon, lag, business decision, owner, cost of error.
  2. 02eXamine deviations across the full demand system. Forecast, actual, orders, sell-out, events, inventory, supply, data quality, uncertainty.
  3. 03Consolidate signals into decision cases. Group duplicate alerts, related entities, common events, shared root causes.
  4. 04Explain likely root causes. Data, availability, commercial event, customer ordering, lifecycle, market, model, noise.
  5. 05Prioritize by impact, urgency, and actionability. Business consequence, strategic importance, decision deadline, ability to intervene, evidence.
  6. 06Test interventions and alternatives. No action, forecast adjustment, event update, data correction, supply response, escalation.
  7. 07Institute governed action. Owner, reason, evidence, approval, version, expiry, audit.
  8. 08Observe the operational outcome. Actual demand, service, inventory, supply, financial effect, model behaviour.
  9. 09Normalize learning through FVA. Forecast impact, decision impact, override performance, root-cause recurrence, automation potential.
LevelWhat it addsCharacteristics
0 · Manual forecast maintenanceNothing systematicSpreadsheets, broad manual overrides, little forecast history, no exception discipline
1 · Statistical forecastingA modelSystem forecast, basic accuracy measurement, planner adjustments, fixed reports
2 · Exception-based planningFocusThresholds, exception tables, segmentation, reason codes, preserved forecast versions
3 · FVA-driven demand planningProofNaïve benchmarks, process-step FVA, override measurement, no-touch segments, bias governance
4 · Agentic Exception ManagerJudgment supportConsolidated cases, contextual diagnosis, business prioritization, cross-functional workflow, probabilistic risk, governed recommendations
5 · Continuous demand decision systemClosed loopAlways-on monitoring, adaptive exception policies, calibrated uncertainty, value-based attention, outcome optimization
Demand-planning maturity

Part XXVI in brief — the practitioner templates. The method ships as eight working documents: the exception card (entities, period, forecast version, current and previous forecast, latest signal, exception type, business impact, urgency, action deadline, likely root cause, evidence, missing evidence, recommended action, alternative, owner, confidence); the forecast-override card (scope, periods, original and proposed forecast, reason code, evidence, demand-or-target classification, business and supply impact, effective date, expiry, approver); the FVA scorecard (error at each process step, FVA by source, bias, business outcome, recommended process action); the root-cause card and exception-priority card (the two halves of triage); the new-product and promotion exception cards (the two recurring cases that need their own evidence sets); the decision packet; and the post-exception learning card (approved action, adjusted forecast, actual demand, forecast FVA, service and inventory and financial outcome, confirmed root cause, decision quality, reusable lesson, automation candidate, reviewer). The last field is the one that compounds: every resolved exception is a candidate for never needing a human again.

Frequently asked questions

Does this replace the forecasting model, or the planner?

Neither. It coordinates the work around statistical, machine-learning, causal, and probabilistic forecasting models, and it reduces low-value investigation so planners can concentrate on consequential decisions, uncertainty, and cross-functional coordination.

Can planner overrides make forecasts worse?

Yes. Research shows judgmental adjustments add value inconsistently, which is exactly why override performance should be measured by segment, reason, horizon, and planner rather than assumed.

Should all negative-FVA overrides be prohibited?

No. Some decisions improve operational outcomes while worsening one accuracy metric — buying inventory against an asymmetric service risk, for example. Evaluate forecast value and business value together.

Is MAPE a good metric?

It is useful for sufficiently large positive demand and behaves badly when actual demand is zero or near zero. WAPE is better for portfolios but hides strategic low-volume items; MASE compares against a naïve benchmark and works across differently scaled series. Use more than one, and always alongside bias.

Why measure accuracy at multiple lags?

Because different decisions need different lead times. A forecast that is accurate one week ahead may be unusable for an eight-week production decision — and the eight-week number is the one that determines whether the product exists.

Should customer orders directly replace the forecast?

Not automatically. Orders reflect forward buying, batching, inventory loading, and constrained behaviour as well as demand. Decompose the spike before adopting it.

How should stockouts be handled?

Observed sales must be adjusted or interpreted carefully, because availability constraints suppress recorded demand. Reducing a baseline because a product was unavailable is one of the most common self-inflicted forecasting errors in FMCG.

Which products belong on no-touch forecasting?

Stable products and segments where statistical models demonstrably outperform manual adjustment and where the risk is controlled — proven by measured FVA, not assumed from the segmentation matrix.

Should the agent ever change a forecast automatically?

Material forecast changes should require authorized planner approval. Bounded administrative actions — closing duplicates, expiring overrides, creating tasks — may be automated after strong evaluation and with policy controls.

What is the biggest AI mistake in demand planning?

Building another forecast generator or alert bot without addressing business impact, root-cause diagnosis, actionability, forecast versions, planner overrides, operational outcomes, and FVA.

Conclusion

Demand planning is usually run as a forecast-production process. The system generates numbers, planners adjust numbers, the organization measures error. But the purpose of demand planning is not to produce forecasts — it is to support decisions under uncertainty: how much to produce, what to buy, where to place inventory, which customers to prioritize, how to respond to a promotion, when to escalate risk, and how much uncertainty to carry. A forecast can be statistically sophisticated and operationally weak. A planner can improve one metric and reduce business value. An exception system can generate 1,942 alerts and still fail to surface the one case that mattered.

So the next evolution is not more forecast rows, more dashboards, more alerts, or more manual overrides. It is better allocation of attention. The Agentic Demand-Planning Exception Manager does not replace the statistical engine; it asks whether the engine is operating in the right context. It does not treat every change as a problem; it determines whether the change matters. It does not treat every exception as an override opportunity; it identifies whether the right response is a data correction, an event update, a forecast change, a supply action, an escalation, monitoring, or nothing at all. It does not assume observed sales equal demand — it checks availability. It does not assume customer orders equal consumption — it checks inventory and ordering behaviour. It does not assume planner judgment adds value — it measures FVA. And it does not optimize a single forecast metric; it connects forecast performance to service, inventory, margin, working capital, and planner effort. The EXCEPTION Method walks the route: establish, examine, consolidate, explain, prioritize, test, institute, observe, normalize.

The defining question is not how to make every forecast more accurate. It is which deviations deserve intervention, what is causing them, what action can still improve the outcome — and whether that intervention ultimately created value.

From guide to production

Want help choosing the right architecture for your process?

We map where agents create leverage in FMCG operations, then build and ship the ones that pay back. One call to pressure-test your highest-leverage use case.

All playbooks