You measure what a deploy did to user behavior by comparing what real users did in the sessions before it went live against what they did in the sessions after, on the flows you care about, with a significance test between the 2 windows. Not a dashboard glance. Not a replay someone watched. A before and after, on your own sessions, per release.

That is the whole method. The rest of this article is about why it is harder than it sounds, and what it looks like when it runs on every deploy without anyone doing it by hand.

The two halves of the loop

We build 2 things at UXSense, and they only make sense together. Drift checks every pull request against what users actually do, before merge. Releases measures what every deploy did to real user behavior, after. The boundary is simple: Drift checks pull requests before merge. Releases tells you what actually happened after you shipped.

Most of what we write is about Drift, because that is where the damage gets stopped. This piece is about the other half, because Drift's verdicts are only worth anything if something grades them. Every Drift prediction is stored against the release that ships it; the Release Impact Report for that deploy is what it is graded against. A PR checker that never learns whether it was right is just a bot with opinions.

What you are actually looking for

Behavioral drift is when a change breaks how users navigate your product without breaking the code. Users learn a product: which button, which order, which step comes next. A refactor moves the button, reorders the steps, or hides the section. Every path still technically works. But returning users hesitate, retry, backtrack, or abandon a step they have completed a hundred times, and nothing in CI can see it because nothing in CI knows what users do.

The symptoms are quiet. Hesitation before a click. A retry on a form. Time on a step creeping up. Power users stumbling in familiar places. Support questions rising with no clear cause. Dashboards stay green through all of it, because dashboards measure the code and the traffic, not the muscle memory.

So the question after a deploy is not "did it work." The question is "did this release change how users behave, and does it matter." That has 3 parts: what changed, who it changed for, and whether it was on purpose.

Why a replay summary is not a measurement

The obvious shortcut is session replay. You already have recordings in PostHog or Sentry. Watch a few after the deploy, see if anything looks off.

The problem is sampling. You watch 5 sessions out of a few thousand, and you pick the 5 that happened to be in front of you. A person doing this is anecdotal. A model doing this is anecdotal at scale: it reads a replay and writes a paragraph, and the paragraph is a guess about what the replay meant. It might be a good guess. It is not a number, and nobody can test it.

A measurement is different. It takes every relevant session in the before window and every relevant session in the after window, computes the same behavioral figures on each (completion, abandonment at each step, time per step, retries, backtracking), and runs a significance test on the difference. If the after window is not large enough to say anything, it says so. Every figure comes from sessions. No model writes a number.

That distinction is the whole product. A Release Impact Report reads the release to work out what was claimed, then checks the claim against what users did. The reading is the model's job. The figures are not.

What the report has to answer

For a deploy to be measured rather than described, the output has to answer a fixed set of questions, the same set every time:

  • What changed in user behavior, on which flows.
  • Which user groups were affected. Returning users and new users usually diverge, because only one group had muscle memory to break.
  • Where friction increased, and where it decreased. Releases improve things too, and a report that only finds problems is not measuring.
  • Whether the change was likely intentional. A redesign that moves a button on purpose should read differently from a refactor that moved it by accident.
  • How confident the system is, and why. At 30 sessions a report is an observational audit and should say so. The reliability label upgrades as sessions accumulate.
  • What to fix first, with one action for engineering, one for product, one for design.

If a tool cannot answer those 6 things per release, it is giving you an account of what happened, not a measurement of what changed.

Making it run on every deploy

The method is easy to do once. The value is in doing it every time, without anyone deciding to.

That means the release has to be detected, not declared. Publish a GitHub release or tag and it is picked up. Connect the Vercel or Render deploy webhook and every production deploy becomes a release, no tags required. Deduplication is automatic, one release per repo and tag, so a webhook retry never produces a duplicate report. For platforms we do not hook into yet, a name and a timestamp adds one by hand.

Each release starts the clock. Sessions before it are the baseline, sessions after it are the comparison, and the report ships once there is enough behavior to read. Your first report is a UX Health Audit of what is already broken, before you ship anything new, because the before window exists the moment you connect a source.

The source can be ours or one you already run. If PostHog or Sentry is recording, connect it and there is no new code on your site. [VERIFY: Sentry connection is listed in getting-started.md but has no dedicated help page; confirm it is live before this ships.] Connecting PostHog also gives you a backtest: what the report would have said about your last release, built from recordings you already have, no waiting.

FAQ

Do I need Drift to measure releases? No. Releases works standalone. Drift needs Releases to calibrate its verdicts; Releases needs nothing from Drift. They are better together, which is what the loop is for.

How many sessions before the first report? About 30 for a first signal. At that size the report is observational and says so. The reliability label upgrades at 75, 150, and 300 sessions. [VERIFY: thresholds from help/getting-started.md; confirm still current in app.]

What if I deploy continuously with no tags? Use the Vercel or Render deploy webhook. Every production deploy becomes a release and gets its own report.

Can it tell a deliberate redesign from an accident? It reads the release to work out what was claimed, then reports whether the observed behavior change matches the claim. The judgement about whether the change was likely intentional is in the report, with the figures that support it beside it.

Where to start

If you are already checking pull requests with Drift, every verdict is stored against the release that ships it, and the Release Impact Report for that deploy is what it is graded against. If you are not, start there: Drift is the front door, and the first report on your last deploy is waiting in the sessions you already have. Check your pull requests against real user behavior.

[VERIFY: automatic grading live? claims.md and launch-packet-2026-09-05.md conflict.]