You need to decide whether the next dollar belongs in Search, Performance Max, Shopping, Demand Gen, Video, or App. Looking at campaign-level ROAS alone will not answer that question. Changing one part of the account can alter what the other campaigns capture, so the decision has to be evaluated at the portfolio level.
Google Campaign Mix Experiments gives you a way to compare complete campaign combinations rather than treating every campaign as an isolated unit. Used carefully, the beta can tell you whether a different mix produces a better business result. Used casually, it can produce a confident-looking answer to a badly framed question.
Start with the spending decision, not the campaign list
A useful mix experiment begins with a decision you could make after seeing the result. “Test Performance Max” is not a decision. “Determine whether moving budget from the current Search and Shopping mix into a Search and Performance Max mix improves conversion value at the same total budget” is.
Write your hypothesis in this form:
If we change [one portfolio variable] while holding [the important controls] constant, we expect [primary metric] to improve enough to justify [the account change].
Campaign mix experiment hypothesis template
The phrase “enough to justify” matters. A measurable difference is not automatically a commercially important difference. Before launch, define the smallest improvement that would cover the operational cost, additional complexity, or risk created by the proposed mix. That threshold is your materiality rule.
Choose one primary metric that matches the decision:
- ROAS fits a revenue-efficiency decision when your conversion values are dependable.
- CPA fits a cost-efficiency decision when the counted conversions have reasonably comparable business value.
- Conversions fits a volume decision when generating more qualified actions is the main objective.
- Conversion value fits a growth decision when total value matters more than efficiency alone.
Google supports reporting around ROAS, CPA, conversions, and conversion value. You can inspect all of them, but naming one primary metric in advance prevents a common analytical mistake: searching the results for whichever metric makes the preferred arm look best.
Key takeaways
- Frame the experiment as a portfolio-level business decision, not a request to identify the best individual campaign.
- Change one meaningful variable between arms and keep the other important conditions aligned.
- Keep total budgets comparable unless total spend is explicitly the variable under test.
- Avoid shared budgets and material account changes while the experiment is running.
- Preselect the primary metric, confidence interval, materiality rule, and minimum duration before looking at outcomes.
- Plan for at least six to eight weeks, but do not assume that duration alone guarantees a decisive result.
Build arms that isolate one portfolio variable

An experiment arm is one complete version of the campaign portfolio. The beta supports up to five arms, and the same campaign can appear in more than one arm. That flexibility is valuable because you can preserve the common parts of the account while changing only the element you need to evaluate.
More arms are not inherently better. Every additional arm creates another comparison and divides the available traffic. Use the fewest arms that can answer the decision. For many questions, a current-state control and one alternative are enough.
The framework covers Search, Performance Max, Shopping, Demand Gen, Video, and App campaigns. Hotels campaigns are excluded. That breadth lets you test a cross-channel plan, but it does not remove the need for a clean experimental contrast.
| Decision | What changes between arms | What should stay aligned |
|---|---|---|
| Channel budget allocation | The distribution of budget among campaign types | Total portfolio budget, measurement, and other material settings |
| Consolidation versus fragmentation | The number or structure of campaigns | Total budget, business objective, and the intended audience or inventory scope |
| Bidding strategy | The bidding approach being evaluated | Campaign mix, budget treatment, targeting, and measurement |
| Targeting option | The selected targeting treatment | Budgets, bidding, creative treatment, and the rest of the portfolio |
| Feature adoption | The feature is used in one arm and not the other | Everything not required to enable that feature |
Suppose you change campaign structure, bidding, targeting, and budget distribution in the same arm. A winning result tells you that the package performed differently, but not which change caused it. You also cannot tell whether one helpful change compensated for another harmful one. That may be acceptable when the package itself is the business decision, but it is a poor design when you need reusable knowledge.
Budget handling deserves particular care. If you want to test the mix, keep the total planned budget equal and change its internal allocation. If you want to test a higher total spend level, make total spend the sole intended difference. Do not quietly give the preferred arm both a different campaign combination and more money; the result will not distinguish the effect of mix from the effect of spend.
Traffic can be allocated among arms with splits starting at 1%, and reporting is adjusted to the smallest split so the comparison remains fair. Treat 1% as a configuration boundary, not a recommendation. A very small arm may receive too little information to resolve a commercially modest difference, especially when conversions are sparse. The better question is whether every arm can accumulate enough relevant outcomes during the planned window.
Protect the comparison for the full test window
A strong setup can still fail after launch. New promotions, tracking changes, creative replacements, altered conversion values, revised targets, and unplanned budget moves can all change the conditions under which the arms are being compared. If those interventions affect the arms differently, you no longer have the experiment you designed.
Plan to run a campaign mix experiment for at least six to eight weeks. This is a minimum operating window, not a promise of statistical certainty. An account with limited conversion volume or a small true difference may still produce a wide range of plausible outcomes after that period.
Before launch, complete a short preflight:
- Validate measurement. Confirm that the conversions and values feeding the primary metric represent the business outcome you intend to optimize. Fix tracking before the experiment, not during it.
- Check arm symmetry. Verify that the total budgets and non-tested settings are aligned wherever the hypothesis requires them to be.
- Remove shared-budget dependencies. Google advises avoiding shared budgets during these experiments. A shared budget can redistribute spend across campaigns and obscure the portfolio treatment you meant to test.
- List prohibited changes. Record which budgets, bidding settings, targets, campaign structures, features, and measurement rules must remain untouched.
- Record unavoidable events. If a promotion, inventory interruption, landing-page failure, or other business event occurs, document when it began, which campaigns it affected, and whether it compromised comparability.
- Set review dates. Monitor for broken delivery or measurement, but do not repeatedly judge the winner from early fluctuations.
- Define stop conditions. Separate genuine operational failures, such as broken tracking, from ordinary underperformance. A disappointing early result is not by itself evidence that the experiment is invalid.
The instruction to avoid significant changes does not mean ignoring a serious problem. If tracking fails or an arm cannot deliver as designed, protect the business and correct the problem. Then decide whether the comparison remains interpretable or needs to be restarted. The mistake is pretending that a materially altered test still answers the original hypothesis.
Keep a change log even when no restart is needed. Record the date, affected arms, reason, and expected impact of every intervention. When the result arrives several weeks later, that log will help you distinguish a real portfolio effect from a mid-test account event.
Read the portfolio result before diagnosing campaigns

The Experiment summary should answer the question you wrote before launch: did one complete mix improve the primary business metric enough to change your decision? Campaign-level reporting then helps you understand where the portfolio difference appeared. Reversing that order invites cherry-picking.
One campaign can improve while the portfolio remains flat or declines. Another campaign can look weaker while the total arm improves because the mix is capturing demand more efficiently as a whole. Campaign-level movement is diagnostic evidence; it is not a substitute for the arm-level result.
Google lets you view experiment reporting with 95%, 80%, or 70% confidence intervals. Choose the interval before reading the outcome. A more conservative interval demands stronger evidence and will generally produce a wider range. A lower interval accepts more uncertainty. Switching among them until a preferred arm appears convincing turns an analytical setting into a result-shopping tool.
Read the result through three separate lenses:
- Direction: Which arm currently appears better on the primary metric?
- Uncertainty: Does the interval leave room for a materially different conclusion, including a meaningful loss?
- Materiality: Is the likely difference large enough to justify the budget move, structural complexity, or operational burden?
Do not collapse those questions into a single winner label. A positive point estimate with a broad interval can still be inconclusive. A statistically clear but commercially tiny improvement may not justify rebuilding the account. An interval that includes little or no difference does not prove that the arms are identical; it means this run did not resolve the difference precisely enough under the selected standard.
Use the metric in the context of its inputs. ROAS and conversion value depend on the quality of the values assigned to conversions. CPA can look healthier when the mix generates cheaper but less valuable actions. Conversion volume can increase while efficiency deteriorates. These are not reasons to abandon a primary metric. They are reasons to make sure it represents the decision before the test begins and to use the other metrics as context rather than alternate finish lines.
Turn the finding into a controlled account decision
The result should lead to one of three actions: adopt the alternative, retain the current mix, or collect more evidence. Write the rule before launch so the post-test discussion is about evidence and tradeoffs rather than stakeholder preference.
- Adopt: The alternative improves the preselected primary metric, the uncertainty is acceptable under the chosen interval, and the effect exceeds your materiality threshold.
- Retain: The alternative is worse, creates an unacceptable downside, or fails to produce enough benefit to cover its complexity and cost.
- Collect more evidence: The plausible range includes outcomes that would lead to different business decisions. Treat this as unresolved, not as a tie and not as permission to select the preferred narrative.
If you adopt a winning mix, implement the treatment you actually tested. Adding new targeting, changing bids, moving the total budget, and restructuring campaigns during rollout creates a new package whose performance was never evaluated. Make the validated change first, observe it under normal account conditions, and treat later improvements as separate decisions.
If the result is inconclusive, do not automatically rerun the same design. First identify why the answer remained unclear. The true difference may be too small to matter, an arm may have received too little useful traffic, the primary outcome may be too sparse, or account changes may have weakened the comparison. Rerun only when you can improve the design or when resolving the decision is worth another full testing window.
A compact decision record makes the learning reusable. Save these fields with the result:
- The business decision and one-sentence hypothesis
- The campaigns and settings included in every arm
- The single intended difference between arms
- Total budget treatment and traffic allocation
- The primary metric and materiality threshold
- The preselected confidence interval
- The planned and actual run dates
- All material account or business events during the test
- The arm-level result and relevant campaign-level diagnosis
- The final decision, owner, and implementation boundary
Your best first use of Campaign Mix Experiments is the largest unresolved allocation decision that can still be isolated cleanly. Write the hypothesis, name the metric, and sketch the control and alternative on one page. If you cannot explain exactly what changes and what stays fixed, the experiment is not ready to launch.



























