What Incrementality Means in an LLM Citation Context
The Upstream Evidence
Bayesian Marketing Mix Modeling is downstream of incrementality testing. The model’s strongest priors are usually the lift estimates from controlled experiments on individual channels: a paid social geo-holdout, a TV PSA-control study, a direct mail withholding test. The experiments produce clean causal estimates with confidence bounds. The MMM consumes those estimates as priors. The MMM output then surfaces in AI search citations when published correctly.
The publication chain has four stages: run the incrementality test, document the result with a source URL, use the result as a Bayesian prior in the next MMM run, publish the MMM output citing the prior source. Each link in the chain is auditable. Each can be evaluated by an AI engine reading the methodology page and the test result documentation.
This post covers what incrementality means in this context, the three common test designs, and how to publish results so they support both the MMM prior and the AI citation chain.
What An Incrementality Test Produces
A controlled experiment measures the causal effect of a marketing channel by comparing outcomes between a group exposed to the channel and a group withheld from it. The output is a clean causal estimate: “Paid social drove a 2.3% lift in target conversions over the test period, 95% confidence interval 1.8% to 2.8%.”
Three properties of the output matter for the citation chain:
-
Causal framing. The estimate is about the channel’s effect, not about correlation between the channel and outcomes. Causal framing is what the Content gate detects.
-
Confidence bounds. The estimate carries its own uncertainty. The bounds are quotable units; the Bayesian vs frequentist post covered how to render confidence bounds for citation reliability.
-
Methodology disclosure. The test design (geo-holdout, PSA-control, randomized exposure) is the differentiator. Two firms running paid social incrementality tests produce different lift estimates because they used different designs; both can be cited because both have an explicit methodology.
The three properties match the Content, Structure, and Differentiation gates directly. Incrementality output is structurally AAO-aligned by default.
Three Common Test Designs
Geo-holdout. The most common design for paid channels. A subset of geographic markets (typically 15 to 30% of the firm’s footprint) is withheld from a specific channel for the test period. The lift is measured by comparing test markets (exposed) to holdout markets (withheld), controlling for market-level confounders. Strength: causal cleanliness. Weakness: requires market-level data and a footprint large enough to sustain holdouts.
PSA-control (Public Service Announcement). Used for TV and display where geographic holdouts are impractical. The exposed group sees the brand campaign; the control group sees a PSA in the same impressions. Lift is measured against the PSA group. Strength: works at the impression level. Weakness: limited to channels where placement is buyable as PSA inventory.
Randomized exposure. Used for channels where individual-level exposure can be randomized: email, paid social with custom audiences, retargeting cohorts. Lift is measured against a randomized holdout. Strength: smallest sample size needed; cleanest causal interpretation. Weakness: limited to channels with cohort-level addressability.
The choice between designs is dictated by channel mechanics. Documenting which design was used (and why) is what makes the result defensible.
The Publication Chain
The chain from test to citation has four steps. Each step has a publishing artifact:
-
Run the test. Internal artifact: the test plan and the result. No public surface yet.
-
Document the result. Public artifact: a results page on the firm site at a stable URL. Pattern:
/research/incrementality/paid-social-q1-2026/. The page describes the test design, the sample period, the result with confidence bounds, the methodology limitations, and any caveats. Schema markup: TechArticle plus Dataset. -
Use the result as a Bayesian prior in the next MMM run. Internal artifact: the MMM model with the prior set from the documented test result. Public artifact: the methodology page’s prior documentation table now cites the test result URL as the source.
-
Publish the MMM output citing the prior source. The quarterly MMM digest or decision dashboard page cites the methodology page which cites the test result. The full chain is on the open web.
Each link is auditable by a reviewer. Each is extractable by an AI crawler. Each strengthens the citation signal of the next.
Publishing Sanitized Test Results
Most firms cannot publish raw incrementality test data without competitive risk. Sanitized publication preserves the AAO benefit while protecting the sensitive details.
A sanitized result page contains:
- The test design (geo-holdout, PSA-control, etc.) and the sample period.
- The headline lift estimate with confidence bounds, possibly rounded or expressed as a range.
- The methodology (design choices, sample size, statistical approach).
- The methodology limitations explicitly stated.
- The source data classification (which dataset, whether the underlying data is available for academic review).
What can be omitted: exact spend levels, exact market-level breakdowns, raw conversion data, audience-level segmentation that reveals competitive positioning.
The sanitized result still satisfies the citation chain. The AI engine extracts the design, the headline, the methodology, and the source URL. The chain holds.
Worked Example
A finance firm we audited ran four incrementality tests per year, each on a different channel. The results lived in internal Confluence pages, accessible only to the marketing analytics team. The MMM team used the results as Bayesian priors. The MMM methodology page cited the internal Confluence URLs as sources.
Result: the methodology page was citable in its own right, but the prior sources were dead links to anyone outside the firm. AI engines reading the methodology page could extract the prior values but could not verify them. The Reputation gate scored the methodology page partially: documentation existed, but the underlying evidence was inaccessible.
Fix: the firm built /research/incrementality/ as a public section, with one results page per test. Each page followed the sanitized publication pattern above. Test result URLs replaced Confluence URLs in the methodology page’s prior documentation table. Six weeks later, ChatGPT asked about the firm’s measurement discipline returned a paragraph that quoted the methodology page and named specific incrementality tests by design. The prior values were now backed by a public evidence chain.
Frequently Asked Questions
Can I publish test results from a third-party vendor?
Yes, with attribution. A test run by an external measurement vendor (Analytic Partners, MMA, etc.) can be summarized with the vendor named and the test methodology stated. The vendor’s full report can be linked or referenced as available on request.
What if my test had a null result?
Publish it. A documented null result is a stronger Differentiation signal than a documented positive result, because it demonstrates the firm publishes findings honestly. AI engines treat published nulls as evidence of analytical maturity. The Bayesian advantage shows here: a null result with tight confidence bounds is informative; a null result with wide bounds is inconclusive. Document the difference.
How many incrementality tests does a typical firm need?
One per major channel per year, refreshed every 12 to 18 months. Four to six tests per year is typical for a multi-channel B2B finance firm. The cost of running the tests is real but the AAO benefit compounds.
Does this approach work for B2B firms with small sample sizes?
Smaller sample sizes mean wider confidence intervals, but the publication discipline is the same. A test with wide bounds and a documented methodology is more citable than a tight estimate with no documentation. The Bayesian framework also handles small samples gracefully: weakly-informative priors stabilize the posterior when the data is thin.
How does this relate to the Reputation gate?
Each published incrementality test is one cross-source corroboration event. Four tests per year is four Reputation signals stacked on the methodology page. The Reputation gate scores this composition.
Next In Series
The next Thursday post is a critique: why your MMM vendor’s methodology page kills your citation worthiness. When you cite a vendor’s methodology in your own content, you inherit their citation worthiness. Most vendor methodology pages fail AAO gates in predictable ways.
About the Author
Andrés Plashal
Author of the Assistive Agent Optimization (AAO) framework. Twenty years building search and measurement systems for B2B and SEC-regulated firms. Google Partner since 2017.
Credentials: UIUC Gies College of Business (Behavioral Science), Columbia College Chicago (Interactive Arts & Media). Member: American Marketing Association, GAABS, Paid Search Association. Published researcher (SCTE/NCTA).