Skip to content
Strategy6 min read

Google's Meridian GeoX Is Free: How to Run Your First Geo-Lift Test

Google has made Meridian GeoX, its open-source geo experiment tool, generally available. How to design a first geo-lift test, what changes in small markets such as Cyprus and the Gulf, and why a Google tool needs independent checks.

Share

Most advertisers now set budgets from numbers the ad platforms report about themselves. A geo experiment steps outside that loop: change media in some regions, leave the rest alone, and compare what happens to sales. Google has made a free tool for this generally available. It removes the software cost of a first test, but not the design work, the media cost, or the need to check a Google tool that will often be pointed at Google spend.

What Google released

Meridian GeoX left beta and became available worldwide in early September 2026, reported by PPC News Feed on September 8 and by PPC Land on September 9. Google's developer page describes it as an open-source incrementality solution for "transparent, cost-effective, and publisher-agnostic Geo experiments". It runs on its own, and its results can be converted into priors that calibrate Meridian, Google's open-source marketing mix model.

It supports holdback, go-dark and heavy-up designs, plus multi-cell studies that compare several treatments against one shared control group. The default method is time-based regression, with stratified sampling to assign regions to groups. Time-based regression models how test and control regions moved together before the test, then projects the control forward to estimate what would have happened otherwise.

Google's GeoX product manager claimed budget savings of more than 31% for large advertisers compared with other open-source tools, according to PPC Land, which notes that no sample, test period or comparison set was disclosed and that the figure has not been independently replicated. Treat it as a vendor claim, not a planning assumption.

Designing a first geo-lift test

Four decisions determine whether the result is worth anything, and all four are yours.

Choose the question, then the design

DesignWhat changes in the test regionsThe question it answersMain cost
HoldbackA new campaign runs everywhere except the holdout regionsDoes this new activity add sales at all?Sales forgone in the holdout regions during launch
Go-darkSpend on an existing channel stopsWhat would we lose if we cut this channel?Possible lost sales in the test regions for the test period
Heavy-upSpend rises in the test regionsDoes more budget still return more?Extra media spend with an uncertain return

For a first test, pick one channel where the budget is large and the doubt is real: the line item finance questions every quarter, or a brand search budget that may be buying clicks that would have come anyway.

Measure an outcome you own

The outcome should come from your own first-party data: e-commerce orders, qualified CRM leads, till sales. Google's documentation sets firm data rules: daily data only (weekly data is rejected), pre-test history of at least three times the planned test length, and at least two regions in each group.

Select regions that behave alike

Good test and control groups track each other closely before the test. Pull daily sales by region, ideally over a full seasonal cycle, and look for regions whose curves move together. Then check that every ad platform in the test can target and exclude the regions you chose. A postcode design fails if one platform only targets by city.

Size duration and budget with the design step

The library's code sets default design values of 0.1 for significance and 0.8 for statistical power, with a two-sided test. In plain terms, the test is sized to detect an effect of a given size most of the time, and the smaller that effect, the more regions, weeks or spend difference you need. If the smallest detectable lift is bigger than any lift you could plausibly expect, the test will tell you nothing, and it is cheaper to learn that on paper. Allow for conversion lag too: a product with a three-week consideration cycle needs a longer test window and read-out period.

Small markets: Cyprus, the Gulf and other thin geographies

Geo experiments work best with many comparable regions to split. In a small market the arithmetic gets harder quickly.

  • Few regions. The fewer regions there are, the harder it is to find a control group that tracks the test group, and the larger a lift must be before the test can see it.
  • One dominant city. Where one city carries most of the demand, nothing comparable can be matched against it. Leaving it out is often cleaner than forcing it into a group.
  • Spillover. In a compact country people shop, commute and see outdoor media across regional lines, which leaks the treatment into the control group and shrinks the measured effect.
  • Targeting granularity. Geographic targeting options differ by platform and market. Confirm what each supports before designing around it.
  • Seasons and events. Tourist seasons, public holidays and one-off events can move one region and not another. Keep them out of the test window or balance them across groups.

When there are too few regions, there are fallbacks: a longer test with a bigger spend difference, such as switching a channel off in the test regions rather than trimming it, or a test in a larger market that sells the same product through the same channel, with the learning carried across carefully. Neither is as clean as a well-powered geo test, and the report should say so.

Why a Google tool measuring Google spend needs checking

The methodology can be read and audited, which is more than most platform lift studies offer. But Google maintains the code, and the README says outside pull requests are very difficult to merge because the code is linked to Google internal systems and has to pass internal review. More importantly, most of the choices that can bias a geo test are not in the code but in which channel is tested, which outcome is counted, which regions are picked and when the test stops.

The branded-search signal is a good example. Google's Meridian documentation presents branded Google query volume as a common brand equity variable for its full-funnel model. Search Engine Journal cautions that a rise in branded searches does not by itself show that a campaign caused later sales, since seasonality, competitors, promotions and news coverage all move it. It is also data from the company selling the media. Five habits keep a geo programme honest:

  1. Fix the design, outcome metric and decision rule in writing before launch, and do not stop or extend the test on early results.
  2. Use sales or leads from your own systems as the outcome, never a platform's conversion count.
  3. Test non-Google channels the same way, so the programme is not only grading Google.
  4. Have someone who does not manage the Google account read the results, or re-run the analysis with a second method.
  5. Compare the geo result with your other evidence, such as the mix model and customer surveys, and investigate disagreements.

Where geo tests sit alongside modelling, surveys and first-party signals is covered in our guide to measuring marketing in an AI attribution world. A clean result earns its keep when you are rebalancing a paid media mix, because it shows what a channel adds rather than what it claims.

Sources

  • https://developers.google.com/meridian/geox
  • https://github.com/google/meridian-geox
  • https://ppcnewsfeed.com/ppc-news/2026-09/meridian-geox-launches-globally-causal-experiments/
  • https://ppc.land/googles-meridian-geox-exits-beta-claiming-31-cheaper-geo-experiments/
  • https://www.searchenginejournal.com/google-launches-meridian-geox-globally/589030/
  • https://developers.google.com/meridian/geox/data-validation-and-quality-checks
  • https://developers.google.com/meridian/docs/advanced-modeling/full-funnel

Have a challenge?
Let's turn it into growth.

Start a Project