Articles

    Marketing Incrementality Case Assessment: Scoring Measurement Judgment

    August 29, 2026
    10 min read
    By Netpy Editorial Team
    Updated August 29, 2026

    The hiring signal sits in the counterfactual

    A candidate may define incrementality yet approve a campaign because its dashboard shows healthy ROAS. The hiring signal is whether they ask what would have happened without the spend.

    A marketing-incrementality case tests that judgment under incomplete reporting, deadline pressure, competing explanations for conversion movement, and no perfect experiment. Candidates must decide what can be claimed now, what needs testing, and how to explain it to a budget owner.

    Keep the work sample narrow: it is not a broad analytics screen, growth certification, or SQL interview. It shows whether a manager can protect investment decisions from attribution-shaped reasoning.

    A case packet that forces trade-offs

    Give every candidate the same packet and time limit, for example 75 minutes independently and 25 minutes of review. Include enough data to reward clear thinking, while leaving gaps a careful candidate must surface rather than assume.

    Use this scenario:

    A subscription meal-planning app ran paid social in six treatment metros for four weeks. Six comparison metros received no paid social. The growth team wants approval to extend the campaign nationally next month. Treatment metros were chosen because prior CPA looked similar to the comparison group; they were not randomly assigned.

    Metric Treatment, four weeks before Treatment, test weeks Comparison, four weeks before Comparison, test weeks
    Eligible site visits 60,000 62,000 58,000 59,000
    Trial starts 1,200 1,740 1,120 1,304
    Activated trials 720 1,020 690 760
    Paid social spend $0 $180,000 $0 $0

    State these constraints:

    • A national creator promotion ran during test weeks three and four, with unidentifiable metro-level reach.
    • Brand search rose in treatment metros, but platform reporting claims credit for some trials.
    • Trials become paid subscriptions after 14 days; the last two test weeks have not matured.
    • Data is weekly and metro-aggregated. User exposure, household location, profit margin, and competitor activity are unavailable.
    • Leadership wants a one-page decision note, next-test measurement plan, and a calculation showing what current figures can and cannot support.

    Ask candidates to list assumptions before calculating. Subtracting platform-attributed conversions from spend answers a media-reporting question, not an incrementality question.

    Name the decision owner and date so candidates do not propose an academically clean design that misses the operating problem. Here the VP of Growth must decide whether to release a national budget; a defensible conditional recommendation is better than demanding perfect evidence.

    Choosing a defensible counterfactual

    Score design fit, not loyalty to one method. A geo holdout may fit because delivery and outcomes are observed by metro, but it risks geographic spillover, unbalanced markets, small samples, and contamination from the national promotion.

    Design choice Suitable conditions Strong candidate specifies Failure mode to mention
    Matched geo holdout Spend controllable by market; conversions observed by market Eligibility, matching variables, random assignment within pairs, duration, exposed and holdout markets Cross-market shopping or national media weakens separation
    Switchback by time period Media turns on/off cleanly; demand is stable enough to model Alternating schedule, washout, seasonality controls, pre-set analysis window Promotions, payday cycles, and learning align with treatment
    User-level holdout Exposure and conversion link at user level Randomization unit, suppression, identity limits, contamination checks Platforms may not honor suppression; users may have multiple identities
    Observational estimate No experiment can run before the decision Explicit causal caveat, sensitivity checks, holdout-validation plan Correlation presented as causal lift

    The prompt should not require geo experimentation. A capable candidate may hold the national rollout, rebuild with paired randomization, and pre-register difference-in-differences; another may allow a limited rollout while reserving randomized holdout metros. Either can score well if it fits the constraints.

    Candidates can calculate a directional difference-in-differences estimate for trial starts:

    Incremental trial estimate = (1,740 - 1,200) - (1,304 - 1,120) = 356

    This is not proof. It assumes parallel trends: without paid social, treatment and comparison metros would have changed similarly. Non-random selection and the creator promotion make that uncertain; candidates should say so and explain how the next test reduces uncertainty.

    Four anchored scoring dimensions

    Use four 0-to-5 dimensions for 20 points. Anchors matter more than elegant labels; without them, interviewers reward confidence or caution differently and scores reflect panel preference.

    Dimension 0–1 points 2–3 points 4–5 points
    Design choice Treats attributed conversions or ROAS as the counterfactual Suggests control but leaves assignment, exposure, or timing vague Chooses feasible randomization or clearly limited observation; defines treatment, holdout, eligibility, contamination checks, decision window
    Assumptions and threats Ignores selection bias, maturation, spillover, or concurrent activity Names risk without showing its effect on the claim Ranks threats, identifies untestable assumptions, proposes mitigation or sensitivity check
    Analysis and metric logic Reports platform CPA or raw growth as incremental impact Calculates lift but mixes trials and paid-customer outcomes States counterfactual, checks denominator and cohort timing, separates leading indicators from mature revenue, shows formula limits
    Communication and decision quality Makes unqualified scale/stop call Gives caveats but no operational decision Gives concise recommendation, confidence boundaries, guardrail, and owned next measurement action

    A five is decision-ready reasoning without disguised uncertainty, not the longest deck. A three can reflect sound instincts but missing operational detail; a one is experimentation language resting on attribution reports.

    Set rules before interviews. For a senior marketing manager, require at least 14 points and no score below 3 in design choice or analysis. For measurement-program ownership, raise the experimental-design and stakeholder-communication bar. Do not use one pass mark for roles with different accountability.

    Two answers can both pass

    This response would score well:

    The current comparison is directional only. Trial starts rose by 540 in treatment metros and 184 in comparison metros, implying 356 incremental trial starts under a parallel-trends assumption. At $180,000 of spend, that is about $506 per estimated incremental trial start. I would not compare that figure with subscription CAC until the cohort matures and we know activation-to-paid conversion. The creator promotion and non-random market selection may bias the estimate. I recommend a six-week paired geo holdout before national expansion, randomizing one metro in each matched pair to media on or off. The primary outcome is activated trials; paid subscriptions at day 30 are secondary. Pause expansion if incremental activated-trial cost exceeds the margin-based threshold set by Finance.

    This can earn 18 or 19 points: it calculates the available estimate, labels assumptions, avoids turning trial data into a subscription claim, and gives leadership a next decision.

    A different answer can pass:

    Do not treat the current result as a scale signal. Comparison markets may not represent treated markets, and the creator promotion is an unknown co-intervention. Allow a capped rollout only in markets held outside the next experiment, then run a user-level suppression test if the platform can verify suppression and identity-match quality. If it cannot, use matched geo pairs. Report incremental activated trials per dollar and a confidence interval or uncertainty range agreed with the analytics team.

    This may score 16 or 17. It does not overwork the arithmetic, but identifies why redesign is needed and offers a conditional alternative. Penalizing it for not choosing geo holdout would test interviewer preference.

    Contrast a weak response: paid social generated 1,740 trials at a CPA of $103, so scale nationally. It earns 0 for design, 0 or 1 for assumptions, 1 for analysis, and 1 for communication: clear, but not causal reasoning.

    ROAS is not causal evidence

    Make red flags explicit in the assessor guide; otherwise interviewers may excuse a polished answer that reaches the desired conclusion.

    • Attributed conversions are called incremental without a holdout, counterfactual model, or causal assumption.
    • ROAS is used as incrementality proof. It is attributed revenue divided by spend, not evidence that revenue would not have arrived anyway.
    • Lift uses raw post-period growth while ignoring comparison, pre-period trend, or seasonality.
    • A method is chosen before checking controllable treatment and outcome data at the same unit.
    • Trial starts, activations, and paid subscriptions are treated as interchangeable.
    • The candidate demands a perfect experiment but cannot advise a bounded interim decision.
    • Statistical significance substitutes for credible design; a precise estimate with contaminated assignment can still mislead.

    A red flag need not mean automatic rejection. Candidates may recover when probed by identifying the flaw and revising the recommendation. Record revised reasoning separately from the initial answer to assess coachability without erasing first instinct.

    Calibrating raters before interviews

    Calibration is assessment design, not administration. Before live interviews, give raters the packet, rubric, and three anonymized sample responses. Each scores independently, writes one sentence of evidence per dimension, then compares notes in a short meeting.

    Discuss gaps by dimension. If one rater awards five for correct arithmetic and another awards two because selection bias is ignored, clarify that analysis includes causal interpretation, not calculation alone. Revise anchor wording rather than splitting the difference.

    Pre-set reconciliation: discuss total-score gaps above two points or dimension gaps of two or more. Keep a calibration log of disputed pattern, agreed anchor, and example language for later interviewers or changed roles.

    Do not let the hiring manager score first. Senior voices can turn a rubric into post-hoc justification; independent scoring preserves evidence for panel inspection.

    Use scores alongside role evidence

    The work sample should influence, not determine, the decision. A candidate may show strong incrementality judgment but lack stakeholder management, technical partnership, or channel knowledge. Conversely, a successful channel operator who cannot distinguish attribution from causal lift should not own budget experiments without support.

    Pair the case with a structured interview on decisions made, probing treatment unit, outcome metric, decision rule, and a result that contradicted expectations. It complements rather than replaces a credible growth-skills assessment across the role.

    Store anonymized scores and hiring outcomes, then review the rubric after a hiring cycle. Seek recurring disagreement, not a false claim that a small sample predicts job performance. The immediate goal is fairer evidence-based candidate comparison.

    Build the first packet this week

    Start with a real team decision, remove sensitive details, and add imperfections that force judgment. Write expected evidence for each rubric dimension before inviting candidates. Have two internal raters run the packet and defend different acceptable answers.

    The assessment should show more than that a candidate knows measurement terms: it should reveal whether they can protect a budget decision from an attractive but unsupported causal story.

    Related Articles