Free interactive tool

Lead Score Band Checker: Does Your Visitor Score Rank?

A visitor lead score is only useful if a higher band means a higher chance of the outcome you care about, such as an opportunity within 90 days. This checker takes the records and outcomes in each score band from your CRM, for the period you set the weights on and for a later holdout period, and tests three things: whether the top band converts clearly better than the rest, whether any band converts clearly worse than the band below it, and whether the top band held up after the period it was tuned on. The answer is on your own data, with the uncertainty shown.

The formula

conversion rate           p = outcomes ÷ records in the band
95% Wilson interval       (p + z²/2n ± z·√(p(1−p)/n + z²/4n²)) ÷ (1 + z²/n),   z = 1.96
lift                      p_band ÷ p_all bands
two-proportion z          z = (p1 − p2) ÷ √(p̄(1−p̄)(1/n1 + 1/n2)),   p̄ = pooled rate

checks (on the holdout period when you enter one)
  separation   top band vs all other bands:          clear when z ≥ 1.96
  ordering     each band vs the band above it:       inversion is clear when z ≥ 1.96
  shrinkage    top band, build period vs holdout:    clear drop when z ≥ 1.96

The Wilson interval (Wilson, 1927) is used instead of the simple “p ± 1.96 × standard error” because B2B bands often have a handful of outcomes, and at small counts the simple interval is unreliable and can even run below zero; Brown, Cai and DasGupta (2001) recommend Wilson for small samples. The z test is the pooled two-proportion test in the NIST/SEMATECH e-Handbook of Statistical Methods, section 7.3.3. The holdout period is the point of the tool: a score checked only on the months its weights were tuned on looks better than it is, because the weights fit that period’s noise as well as its signal (Hastie, Tibshirani and Friedman, chapter 7).

Run your numbers

List the bands from the lowest score to the highest. Build period: the months you set or last changed the weights on. Holdout period: later months the weights never saw. Leave the holdout columns at 0 if you have none yet.

Band 1 (lowest scores)
Band 2
Band 3
Band 4 (highest scores)

The score ranks accounts in the holdout period60+ converts at 9.1% against 1.8% for the other bands in the holdout period (z = 3.65), no band is clearly out of order, and the top band's drop from 15.0% to 9.1% is within noise (z = 0.97). Keep the weights; do not raise them on this evidence.
Checks ran on
Holdout period
All bands, build period
2.7%
All bands, holdout period
2.1%
Top band vs the rest
9.1% vs 1.8% · z = 3.65
Top band, build vs holdout
15.0% → 9.1% · z = 0.97
Bands out of order
None
Conversion rate per band, with the 95% Wilson interval and lift over all bands
BandBuild: rateBuild: liftHoldout: outcomesHoldout: rateHoldout: 95% intervalHoldout: lift
Below 201.0%0.4×9 of 8201.1%0.6% to 2.1%0.5×
20-393.0%1.1×8 of 3102.6%1.3% to 5.0%1.2×
40-596.7%2.5×6 of 1304.6%2.1% to 9.7%2.2×
60+15.0%5.6×5 of 559.1%3.9% to 19.6%4.3×

Every band has fewer than 10 outcomes in the period checked, so the intervals are wide and the z values are approximate. That is normal at B2B volumes: read the verdict as a direction, and add a month before you change weights on it.

Worked example (hypothetical numbers)

A team uses the four bands from the lead-scoring guide (below 20, 20-39, 40-59, 60+) and counts an outcome as an opportunity created within 90 days of the score (hypothetical numbers, the tool’s default input). They set the weights on Q2: 1,280 scored accounts, 34 opportunities. Q3 is the holdout: 1,315 accounts, 28 opportunities.

In Q2 the 60+ band converted at 15.0% (9 of 60), 5.6 times the Q2 average of 2.7%. In Q3 it converted at 9.1% (5 of 55), 4.3 times the Q3 average of 2.1%, against 1.8% for the other three bands combined (23 of 1,260). That gap gives z = 3.65, well above 1.96, so the top band is still clearly better. The bands are in order in Q3 (1.1%, 2.6%, 4.6%, 9.1%). The drop from 15.0% to 9.1% looks large, but with 60 and 55 accounts it gives z = 0.97: it could be noise.

Verdict: the score ranks accounts in the holdout. The team keeps the weights and does not raise the 60+ band’s routing priority on the Q2 figure; they plan on about 9%, not 15%. Every band has fewer than 10 opportunities in Q3, so the Q3 interval for 60+ runs from 3.9% to 19.6%, and they re-run the check after Q4.

Where do the inputs come from?

  • Bands. The score ranges your routing already uses, lowest first. If you store only the total score, group it in a CRM report. Keep the band edges the same in both periods.
  • Records. Accounts (or leads) that entered each band during the period, counted once, at the band they reached first or at their peak; pick one rule and state it. Take out customers, employees and suppressed records first, as in the exclusion rules.
  • Outcome. One event with a fixed window, for example an opportunity created or a meeting held within 90 days of entering the band. Wait until the window has closed for every record in the period, or recent records will look like misses.
  • Build and holdout periods. The build period is the one the current weights were set or last changed on. The holdout is a later period with no weight changes in it. If you changed weights last month, the holdout starts after that change.

What should you change when a check fails?

  • Top band not separated. Merge the top two bands, or route on the top band only as a review queue, not an alert, until a later period gives more outcomes. Adding signals to push more accounts into the top band can dilute it rather than fix it.
  • Bands out of order. Look at which signals move accounts between the two bands. A signal that adds points but lowers conversion, such as repeated blog visits or a noisy page event, needs a lower weight, a cap or a decay rule.
  • Top band shrank clearly. Part of the build result came from that period. Use fewer, coarser weights, and re-check on the next period before you change them again.
  • Score ranks. Set rep capacity for the band you route with the alert capacity planner, using the holdout rate, not the build rate, for the expected yield.

When is the check too simple?

  • Very small counts. The z test is a large-sample approximation. When a band has only a few outcomes, NIST’s handbook points to an exact test (Fisher’s) instead; treat the verdict as a direction and wait for more data.
  • Several checks at once. Each adjacent pair is a separate test, so with many bands one pair can cross 1.96 by chance. That is one reason to keep four or five bands, not ten.
  • Changes between periods. A new campaign, a new identification vendor or a tag outage changes who lands in each band. If the mix changed, check it first with the match-rate drop decomposer; the same traffic-mix logic applies to score bands.
  • Ranking is not causation. A band that converts well tells you where to look first. It does not show that the outreach caused the opportunity, which is a sourced vs influenced question.

Sources

  • Wilson, E. B. (1927). “Probable inference, the law of succession, and statistical inference.” Journal of the American Statistical Association 22(158), 209-212. doi:10.1080/01621459.1927.10502953
  • Brown, L. D., Cai, T. T., and DasGupta, A. (2001). “Interval estimation for a binomial proportion.” Statistical Science 16(2), 101-133. doi:10.1214/ss/1009213286
  • NIST/SEMATECH e-Handbook of Statistical Methods, section 7.3.3, “How can we determine whether two processes produce the same proportion of defectives?” itl.nist.gov (read 2026-09-27).
  • Hastie, T., Tibshirani, R., and Friedman, J. (2009). The Elements of Statistical Learning, 2nd ed., chapter 7, “Model Assessment and Selection.” hastie.su.domains