Results & analytics

Reading A/B results

When a campaign runs as an A/B test, adfit shows a share of your matching visitors the original page and the rest the personalized one, then compares what each group did. This guide explains every number on the Performance page and, just as important, what it does not say.

Where to find your results

Open a site from Sites, then choose Performance under Results. You get one block per published campaign, and inside it one card per goal that counts in A/B results.

The Range and Show controls at the top change the numbers in the tiles and the daily trend. The verdict, the confidence and the sample size always belong to the current run, whatever range you pick.

Screenshot coming soon
The Performance page: a campaign block with its run status, and a goal card with the Original and Personalized tiles, lift, confidence and sample size.

How an A/B test works here

  • A campaign starts without an A/B test: every matching visitor sees the personalized version, and there is no original to compare against. To start one, open A/B settings in the campaign header and turn on Run as A/B test.
  • The holdback is the share of matching visitors who see your original page. You can pick 10%, 20%, 30%, 50%, and 50% is the recommended start.
  • Every campaign has its own A/B setting and holdback, changed under A/B settings in the campaign header. The site-wide switch A/B testing turns experiments off for the whole site.
  • A visitor stays in the same group for as long as adfit remembers them, so they do not flip between versions from page to page. See Snippet behavior for the persistence setting.
  • A lower holdback sends more visitors to the personalized page, but it makes results slower: a 10% holdback needs roughly three times the total visitors of a 50/50 split to reach the same certainty.

Reading a campaign block

The header of each campaign tells you the state of its experiment:

  • Running, Ended, Ended: plan changed, Page changed or A/B test off, followed by the run number, its start date and the holdback.
  • Run history lists earlier runs with the reason they ended (for example A/B settings changed or Targeting changed) and their result.
  • A/B settings and, where it applies, Restart experiment.

A run in the history can carry a note. Each one flags a comparison to read with a little care:

  • Holdback raised: returning visitors may already have seen the personalized page: you increased the share that sees the original, so some of those visitors had already seen the personalized page under the earlier setting.
  • Visitors before this run saw the personalized page: before this run the campaign showed the personalized version to everyone, so its returning visitors are not meeting it for the first time.
  • Site key rotated: returning visitors were re-randomized: after a new site key, returning visitors are assigned to a group again from scratch.

Changing the setup starts a new run

Changing the holdback, the targeting or the baseline ends the current run and starts a fresh one. The numbers collected so far move to the run history and do not mix with the new run.

Reading a goal card

  • Two tiles, Original and Personalized, each with its conversion rate, its visitors and its conversions.
  • The lift: how much better or worse the personalized version did, in percent, followed by Likely range (95%). The range matters more than the single number: if it spans zero, the difference could still be nothing.
  • Confidence with a mark at the threshold this read has to clear: 99.9% for an early read, 95% at the final read.
  • Sample: how far this run is along its planned sample size.
  • By page and By source for the numbers per page and per traffic source, and Daily trend for the day by day chart, which also has a Show as table view.
  • In the daily trend, days without any data show as a hatched band marked no data, and the lines break there. Show as table marks such a day as (no data).
  • Without an A/B test, a goal card shows only the Personalized tile, with no verdict, lift, confidence or sample, and its Daily trend, Show as table, By page and By source list only Personalized.

Cards for goals other than your primary one carry the badge Indicative: same calculation, but they do not decide the verdict. They never get a final read, so they show no Early read or Final read badge, and their threshold stays at 99.9%.

What confidence means

Confidence is not the chance of winning

Confidence is 100% minus the p value of a two sided test. It says how unlikely a difference this large would be if both versions performed the same. It is not the probability that the personalized version is better.

An example: at 97% confidence, a difference this big would show up in fewer than 3 out of 100 tests where the two versions are in truth exactly equal. That is a good reason to believe the difference is real. It is not a statement that the personalized version has a 97% chance of being the better page, and adfit never claims that.

Read the Likely range (95%) next to the lift together with the confidence. It shows the range of true effects that fit your data, and it is usually wider than people expect.

What each verdict means

  • Collecting data: not enough visitors or conversions yet. adfit shows what is still missing and no confidence at all.
  • Leaning personalized or Leaning original: one version is ahead, but it has not cleared the threshold for this read. This is a direction, not a result.
  • Personalized wins or Original wins: the difference cleared the threshold that applies: 99.9% for an early read, 95% at the final read.
  • No clear difference: the run reached its planned sample size without a significant difference. That is a real answer, not a failure: it means any effect is likely smaller than the run was built to detect.
  • Results unreliable: adfit hides the verdict because something is off with the data, for example an uneven visitor split or a changed page. The visitors, conversions and conversion rate of each version stay visible; lift and confidence are hidden, and the card says why.

Early read and final read

Next to the verdict of your primary goal, a badge shows Early read or Final read. Early read means there is no final read yet. While a run is still going, an early read can still flip. When a run has ended, the tooltip on the badge reads: Early read: this run ended before reaching a final read. Its early read can no longer flip, because a run that has ended takes in no new visitors or conversions.

A final read needs two things: the planned sample size is reached, and the test has run for at least 7 days with enough traffic, counted in whole weeks. A day has enough traffic when it has at least 10% of the visitors of an average day with visitors in this test. Days with less, for example while your ads are paused, and days excluded for unusual traffic do not count. adfit checks both every day and gives the final read on the first day both are true, not only at the end of a week. Final read means both conditions are met, and that read is the one adfit stands behind.

The two reads do not use the same bar, and that is the point. For an early read, adfit only calls a winner at 99.9% confidence. At the final read, 95% is enough. The mark on the confidence bar always shows the threshold that applies right now, so the bar and the verdict never tell you different things. Above the bar, the card shows both numbers side by side, for example Confidence 99.4% on the left and Threshold 99.9% on the right.

Checking every day is not free

Looking at a running test every day and acting on the first 95% you see makes false positives far more likely (roughly one in three under daily peeking). That is exactly what the higher early bar protects you from. Treat early reads as a signal, not a decision. The final read is the result.

After the final read, adfit freezes the verdict. Visitors keep arriving, the live numbers keep moving (they also include conversions that arrive late), but the frozen read stays the answer for that run.

If your page changes after the final read, that frozen result still describes the period it was measured in, and adfit flags the run as changed. Update the baseline and start a new run to measure the page you have now.

When adfit shows no confidence yet

Below 100 visitors per version, or fewer than about 10 conversions in total at a 50% holdback (more at lower holdbacks), no confidence is shown. Percentages at this size swing wildly, and a number that moves by ten points because two more people converted would only mislead you.

The card tells you exactly what is missing, for example 62 of 100 visitors per version and 4 of 10 conversions in total needed.

How long a test takes

adfit plans a sample size per run from your original's conversion rate and a minimum effect of 20%, and it never calls a read final before the test has run for at least 7 days with enough traffic, counted in whole weeks, so one strong weekday cannot decide your test.

  • The lower your conversion rate, the more visitors the run needs. At a 2% to 3% conversion rate and a 50/50 split, a run is a matter of weeks, not days.
  • A lower holdback stretches it further, because the original group grows slowly and it is the smaller group that limits the answer.
  • Until the original group has enough conversions, the target is an estimate and adfit labels it as one. After that it is frozen for the run, so the finish line does not move.

What the original group sees

Visitors in the original group get your page exactly as you built it: adfit applies no change to it. It still decides which group the visitor belongs to before anything on the page is touched, and it still counts that visitor, which is why the comparison has two sides at all.

To avoid a visible flicker, adfit briefly holds back the elements a campaign may change before it knows which version to show. That hold is capped by your safety reveal timeout, described on Snippet behavior, and the original content is revealed at the latest when the timeout is reached.

When a conversion counts

  • The visitor must have been assigned to the campaign first. A conversion from someone who never matched is not attributed.
  • adfit counts one conversion per visitor, goal and run, however often the goal fires.
  • Conversions are attributed while the experiment is running and within your persistence window (same page, 24 hours, or 30 days). Conversions after an experiment ended are not attributed.
  • A visitor who arrives through two different campaigns counts in both experiments.

Visitors adfit cannot recognize for long

Without a consent signal, adfit recognizes a visitor only within one UTC day. A visitor who returns the next day counts as a new visitor and may see the other version. This slightly understates the measured lift, it never inflates it.

If your site asks for tracking consent and you connect it, adfit can remember a visitor for 30 days instead. See Consent and visitor tracking.

Notices that pause or change a result

Uneven visitor split

If visitors are split noticeably differently from your holdback, the campaign header shows the note Uneven visitor split with the measured split and a Check installation link. While the run is going and the primary goal has no final read yet, the note reads, for example: Visitors split 44.0% / 56.0% instead of about 50% / 50%. adfit hides the verdict while the split is uneven. If it persists, check your snippet installation.

  • If the primary goal already has its final read, the split no longer changes that read. Verdicts without a final read stay hidden while the split is uneven.
  • When the run has ended, the note only names the split of that run and asks you to check your snippet installation. It says nothing about the verdict.

Where a verdict is hidden, the goal card keeps the visitors, conversions and conversion rate of each version and says why, for example: Lift and confidence are hidden because of an uneven visitor split.

If adfit finds the uneven split only after the final read, the final verdict stays, and the card of the primary goal adds: Uneven visitor split, not found at the final read. The verdict stays. Later numbers may be affected.

The original changed

If your page changes underneath a running experiment, the comparison no longer means what it says: some visitors saw one original, others saw another. adfit flags this as Original changed: seen by visitors (detected in real visits) or Original may have changed: server scan, and marks the run Page changed.

Update the baseline on the campaign's pages, then press Restart experiment. adfit never refreshes a baseline by itself, and that is deliberate: silently adopting a changed page in the middle of a run would rewrite the comparison you already collected.

If the run has already ended, the campaign header notes when the original changed and that the run's results since then are unreliable. It does not ask you to update the baseline or restart; Restart experiment appears only where a new run is possible.

If adfit finds the change only after the final read, the final verdict stays, and the card of the primary goal adds: Page change found after the final read. The verdict stays. Later numbers may be affected. If the visitor split is uneven as well, the card shows the split warning instead.

Unusual traffic on some days

When a day looks like automated or otherwise unusual traffic, adfit takes that day out of the results and notes Unusual traffic on 2 days excluded from the results. The experiment keeps running, and you do not need to restart it. Excluded days are dimmed in the daily trend.

Reduced coverage

On very large traffic spikes adfit stops counting new visitors for a while to protect the measurement. The card then notes that coverage is reduced. Your personalizations are unaffected.

Frozen results

  • Pageview limit reached: personalization pauses, so results freeze at the cap until next month or an upgrade. See Pageview usage.
  • This site is paused or Your subscription is inactive: results freeze until the site or the subscription is active again.
  • Measurement is off: adfit is not measuring on this site at all. Personalizations keep running.
  • A/B testing is off: every matching visitor gets the personalized version, so there is no original to compare against. adfit still counts personalized visitors and conversions.

Results and your pageview allowance

A pageview counts only when adfit decides what a visitor sees: a personalization is applied, or the original is shown to a visitor in an A/B test. All other traffic is free.

Visitors in the original group of a running A/B test are part of that: adfit decided what they see and measured them, so they use your monthly allowance just like personalized visitors do. Visits through a test link are the exception on the results side: they never appear in your results.

Results are indicative

adfit gives you a well defined statistical read: a planned sample size, a stated threshold for each read (99.9% for an early read, 95% at the final read), and a verdict that stops moving at the final read. It is not proof.

  • adfit filters bots, caps how many new visitors it counts and drops days that look manipulated, but no measurement on the open web is immune to someone deliberately sending fake traffic.
  • Seasonality, your own campaigns and changes elsewhere on your site all move conversion rates. A test measures the difference between two versions in the same period, not why the period looked the way it did.