Blog Article

Sales skill has never had a real standard. So we built one.

For two years we scored sales conversations twice - once with our models, once with trained human raters. What we learned sent us back to rebuild our own measurement system.

Val Yaromenko
Jul 26, 2026
Big Sister AI V2 Team Skills view: six sales reps scored 0-100 overall with a per-skill grid covering Ice Break, Empathy Signal, Post-Sell, Agenda and Consent, Discovery Depth and Qualification Coverage

For two years we did something slightly obsessive: we scored sales conversations twice. Once with our models, once with trained human raters, side by side, across thousands of interactions.

We did it to find out whether the machine could keep up. What we actually learned was that our own measurement system was the weak link - and it sent us back to rebuild the whole thing.

What was broken

Our first system scored seven skills as binary checkboxes. Did the rep handle the objection: yes or no. Did they establish next steps: yes or no.

Run that across thousands of real conversations and the flaw becomes obvious. Almost nothing in a sales call is binary. A rep who acknowledges a budget objection and moves on has not "failed objection handling" - they've done part of it. A rep who reframes to total cost of ownership and confirms the customer accepted the reframe has done something categorically better. A checkbox flattens both into the same mark.

Worse, binary scoring hides progress. A rep improving steadily inside the same checkbox shows no movement at all, which makes the score useless for the thing managers most need it for: knowing whether coaching worked. It is the same failure mode behind most AI tools that sales teams adopt and quietly stop using - the tool produces output, but nothing in it tells you whether anything changed.

What we built instead

Big Sister V2 runs a different standard.

18 skills per conversation, each scored 1-5 against a rubric with written reference examples for every level. Not a checkbox - a graded judgment with a defined bar for each step.

Every score cites its evidence. Each skill score comes with the rationale and the exact timestamp in the transcript that earned it. A rep can see why. A manager can check. That is what makes a score arguable - and a score you can argue with is the only kind anyone trusts.

One 0-100 Sales Score per rep and per team. The 18 skill scores roll up into a single composite. The weighting is our IP, tuned against two years of paired human and AI scoring - but the inputs behind any score are fully visible, and you can always open the skill breakdown and the transcript underneath it.

An appeals process. Any rep can dispute a score. Over 10% of interactions are pulled for second-judge review. Every appeal, override and independent re-score becomes a labelled human judgment that refreshes the reference set the model learns from.

Big Sister AI Referee console resolving a disputed Objection Handling score: AI scorer 2, rep appeal 4, AI judge 4 with transcript evidence at 14:22, moving the call composite Sales Score from 62 to 71
An appeal in the Referee console. The AI scorer said 2, the rep argued 4, the judge ruled 4 and cited the moment at 14:22 that settles it. The call composite moves 62 to 71.

That last piece is the part most people miss. The standard is not static. It gets sharper every time a human disagrees with it.

How we know the new standard holds up

Rebuilding the scoring engine alongside the rubric meant we had to re-validate from scratch. So we ran a controlled comparison.

Three trained expert raters - people who assess sales calls professionally - independently scored the same set of meetings across the full skill rubric. Then we scored the same set with the production model.

The three humans agreed with each other at Krippendorff's alpha = 0.319.

For context: 0.667 is the conventional floor for drawing even tentative conclusions from rated data. Three professionals, same rubric, same conversations, landed at less than half of it. That is the actual state of sales assessment across this industry - not under-measured, unreliably measured. When one manager says "your discovery needs work," an equally qualified manager would often disagree.

Then we added our model to the panel as a fourth coder. Group reliability moved from 0.319 to 0.322 - a change of 0.003. Statistically, adding our scorer to a panel of three trained experts is indistinguishable from adding another trained expert.

Measured like for like, the model lands closer to any given expert (mean absolute error 1.049 on the 1-5 scale) than two experts land to each other (1.203), and it's closer than the human pair on 12 of 16 skills. The weakest is Opinion Safety, which is reverse-scored - a known calibration item we're tracking in the open.

What this does not say: that our AI beats human experts. It doesn't. Measured against the panel's averaged consensus, a human scores better than our model - we ran that comparison too and excluded it, because averaging three raters cancels their noise and makes an unfairly easy target. We'd rather report the honest version.

Methodology, stated plainly: the controlled comparison covered 20 meetings and 16 skills, 233 multi-rater units, pre-adjudication, with the rubric tuned on this same set. Small, and favorable to us on all three counts. It's why we don't yet use the word benchmark. Gold Batch 2 - new meetings the rubric has never seen, rated blind - is running now, and we'll publish it in September whichever way it falls.

Why it matters that a machine can hold the standard

A trained human rater is reliable in the way this data shows: roughly as reliable as another trained human, which is to say not very, and expensive either way. Managers review under 3% of their team's conversations.

A model that performs at the level of a trained expert can score 100% of them. Every call, every email, every week, against the same rubric, with the evidence attached. Consistency at coverage is the entire product - not brilliance on one call, but the same standard applied to all of them. That is also why we price per scored rep rather than per seat: the unit that matters is the person being measured, not the person logging in.

We also join scores to your CRM, so skill data sits next to deal data rather than in a separate tool. Whether month-one scores predict month-three close rates is a question we're actively measuring; our first cohort reads out in September, and we'll publish that too.

Where this goes

Every appeal sharpens the reference set. Every new customer adds conversations scored against the same standard. Eventually that supports something no single-tenant tool can produce: telling a team their objection handling sits in the 30th percentile for B2B SaaS - not just that it scored 34.

A standard is only worth the word if it holds across companies. That's what we're building toward.

Questions we get

What is a Sales Score?

A Sales Score is a single 0-100 number for a rep, a team, or an individual conversation. It is the composite of 18 individual skills, each scored 1-5 against a written rubric with reference examples for every level. Every skill score carries its rationale and the timestamp in the transcript that earned it, so the composite can always be opened up and checked.

How accurate is AI scoring of sales calls compared to human experts?

On our validation set, the model lands closer to any given expert rater (mean absolute error 1.049 on the 1-5 skill scale) than two expert raters land to each other (1.203), and closer on 12 of 16 skills. Added to a panel of three trained experts as a fourth coder, it moved group reliability from Krippendorff's alpha 0.319 to 0.322 - a difference of 0.003. It does not beat human experts; it performs like one, at full coverage instead of sampled coverage.

What does a Krippendorff's alpha of 0.319 mean?

Krippendorff's alpha measures how much independent raters agree beyond chance: 1.0 is perfect agreement, 0 is chance. The conventional floor for drawing even tentative conclusions from rated data is 0.667. Three trained sales-call raters scoring the same meetings against the same rubric reached 0.319 - less than half that floor. Sales assessment does not have a reliability problem so much as a reliability vacuum.

Can a sales rep dispute a Sales Score?

Yes. Any rep can appeal any skill score, and an independent AI judge re-scores the disputed skill against the transcript and cites its evidence. Separately, over 10% of interactions are pulled for second-judge review whether or not anyone appealed. Every appeal, override and independent re-score becomes a labelled human judgment that feeds back into the reference set the scorer is calibrated against.

Is Big Sister AI a replacement for my notetaker or my CRM?

No. A CRM is the system of record and a notetaker is the game tape. Big Sister is the referee - the governance layer that scores what happened and holds every rep to the same standard. Most customers run it alongside both.

Big Sister AI scores every sales conversation 0-100, with every score cited to the transcript. See it on your own calls.

Back to blog

See your team's
real scoreboard.

V2 opens to a limited first group. Tell us about your team and we score it first. One objective 0-100 Sales Score per rep, from real conversations.

We reach out personally when your spot opens. No spam. No pressure.

V2 launches to a limited first group. Join the waitlist →