Turn the symptom into a testable decision.
Weak new-product performance is not automatically a recommendation problem. First separate product choice, exposure, availability, price, channel and launch execution. The sequence below is designed to preserve definitions, expose alternative explanations and lead to a decision that can be validated.
- Define launch successChoose contribution margin, sell-through, conversion, repeat purchase or another outcome with a fixed horizon and return treatment.
- Construct comparable peersMatch launch month, market, category, price, channel, exposure and available inventory.
- Separate selection from executionMeasure whether the item was wrong or whether the right item lacked stock, traffic, content or placement.
- Model decision-time informationUse only attributes and signals available when the buyer made the assortment decision.
- Validate portfolio behaviorEvaluate category coverage, cannibalization, supplier concentration and tail risk—not only item-level accuracy.
Ask for the minimum data that can change the answer.
Begin with read-only access and a field-level purpose. Reconcile samples before scaling extraction, preserve event time and source provenance, and record missingness rather than silently filling it.
Validate the claim before changing the operation.
Run a decision-time holdout, then a shadow buying cycle
Train or tune the selection method on older launches and evaluate it on a later untouched period using only data available at selection time. Compare it with the current buyer process and simple baselines. If it holds up, run the next cycle in shadow mode: produce recommendations, record buyer overrides and measure the portfolio without changing commitments until error patterns are understood.
What makes the diagnosis look right and still fail.
- Predicting exposure instead of appealItems with more placement and stock look better even if selection quality is unchanged.
- Using future outcomes as featuresLater reviews, markdowns or realized stock leak the answer into the model.
- Ignoring missing-not-at-random dataProducts not selected have no sales outcome, so naive labels favor past choices.
- Optimizing units instead of economicsA high-volume recommendation can destroy margin or increase returns.
- Replacing buyer judgment too earlyOverrides provide valuable constraints and reveal data the system does not yet represent.
Primary and official references
These sources define the measurement, control or operating context. They do not replace validation on the company’s own data.
Questions enterprise teams ask first.
Is this a recommender system?
It may use recommendation methods, but the business decision is assortment selection under inventory, margin and range constraints.
How do we evaluate products that were never selected?
Use experiments, comparable market exposure, causal methods or staged shadow testing; never pretend the counterfactual is directly observed.
What does a useful benchmark compare?
Compare against the existing buying process, simple rules and relevant top-performing practices using the same decision horizon and constraints.
Can unstructured trend research be included?
Yes, but timestamp sources, preserve provenance and test whether the added signal improves an untouched holdout.