Your CI has a check for types, a check for tests, and a check for lint. None of them knows which button on the page carries most of your checkout completions. Drift adds one more GitHub check, and that is the only thing it knows.
Here's how it works, what the 4 verdicts mean, and why a quiet check is the result you want most of the time.
What behavioral drift is
Behavioral drift is when a change breaks how users navigate a product without breaking the code. Users learn an interface: which button, which order, which step comes next. A refactor moves the button, renames the label, or reorders the steps. Tests pass because the path still technically works. Users hesitate, retry, backtrack, or abandon, because the path they learned is gone. Nothing in CI catches it, because nothing in CI knows what users do.
What the check scores against
When you install the UXSense GitHub App, every pull request is scored against your behavioral load map. That map is built from real recorded sessions, from whichever source you already have: the UXSense recorder, PostHog, or Sentry. It's a deterministic model of which elements your users depend on: how much interaction each one carries, whether it sits on the path to a goal completion, and how ritualized returning users' behavior around it is.
No model guesses a verdict. The model's one job is to read the PR description and work out what the change claims to do. Every number in the check comes from sessions, manifests, or replay runs.
The 4 verdicts
Within a minute of opening or updating a PR, a check named "UXSense behavioral drift" appears.
Clean. The PR touches nothing carrying meaningful user load. Most PRs should land here. A quiet check is a real result, not a miss.
Advisory. Worth a glance. 2 examples: the change touches a surface users rely on but the PR description doesn't mention it (unclaimed surface, the failure agent-written code is known for), or returning users interact with the element ritually and a change would break muscle memory without anyone noticing.
Warn. A high-load element, one on the path to a goal completion, was removed, renamed, or plausibly disturbed.
Block. A recorded real-user journey failed when replayed against the PR's preview deployment. This is the only failing verdict, and you can override it.
Clean and advisory verdicts stay in the checks tab. A comment on the PR appears only at warn or block, and it's one sticky comment, edited in place, not a thread. When something warrants attention the comment says which element, how much load it carries, and which goal path it sits in.
2 tiers of evidence, and only one can block
Drift's signals come in 2 kinds.
The probabilistic tier is the load map: what the PR changes, weighed against what users depend on. It produces advisory and warn. It doesn't block, because it's an inference about a diff, not an observation of a failure.
The deterministic tier is path replay: your users' actual recorded journeys, run against the PR's preview deploy. If a journey that worked before fails now, that's a block. This tier needs a preview deployment for the PR, on Vercel or Render, and it's part of the Team tier. Backtests on historical releases use the probabilistic tier only, because the previews for shipped releases no longer exist, so they cap at warn and the report says so.
That split is deliberate. A check that blocks on a guess gets muted. A check that blocks only when a real journey demonstrably broke keeps its authority.
Coverage, and why a quiet check isn't a missed one
Every check carries a coverage figure: the share of observed interactions that resolved to a stable element identity. Markup without stamps resolves by semantic anchors (role, accessible name, route), which covers a lot. Icon-only buttons and heavily dynamic text can fall through. Installing the UXSense build stamp takes coverage to near-total.
Low coverage makes checks quieter, never noisier. If Drift can't see an element, it doesn't guess about it. So a clean verdict at low coverage means "nothing I could see," and the figure tells you how much that was.
Graded against what actually shipped
Every Drift prediction is stored against the release that ships it, so the Release Impact Report for that deploy is what grades it. The check earns its verdicts from what happened after merge, not from a rule someone wrote once.
What the app can and can't do
The app asks for 4 permissions: checks (write) to post verdicts, pull requests (write) for the single sticky comment, contents (read) to resolve release tags and read changed-file lists, and deployments (read) to find preview deploys for replay. It never pushes code and never opens pull requests.
The check is informational by default. If you want block verdicts to gate merges, add the UXSense check to your branch protection rules. That choice stays with you.
Checks are unmetered on every tier, including free, and free is an ongoing tier rather than a trial. No verdict is withheld for billing.
FAQ
Does it work on every PR from day one? It needs sessions first. The load map is built from recorded user behavior, so a new project with no traffic has nothing to score against. [VERIFY: exact session count before the first Drift check fires; getting-started.md gives 30 sessions for the first report.]
Will it comment on every PR? No. Comments post only at warn or block. Clean and advisory verdicts live in the checks tab. One comment per PR, edited in place.
Can it block a merge? Only when a recorded user journey fails against the PR's preview deployment, which needs a preview deploy and the Team tier. Even then it's overridable, and it only gates merges if you add it to branch protection.
Do I need a new recorder? No. If you already have PostHog or Sentry session replay, connect it and the map builds from what's there.
Install the GitHub App and see your first check on the Drift page.