Most sales scorecards fail before a single call gets graded because the criteria are unscorable: they describe a quality of the call rather than a moment inside it. "Built rapport." "Strong discovery." "Handled objections well." None of these can be verified by two different people listening to the same recording, so the scorecard does not measure the call; it measures the grader's mood.
The solution is a structural pattern: write every criterion as a trigger and an action. A trigger is the specific moment in the call that makes the criterion relevant. The action is the specific, observable behavior the rep is expected to produce once that moment arrives. "Built rapport" is a trait that requires a grader to invent evidence. "When the prospect raised the competitor, the rep asked what was driving the evaluation" is a moment a grader either heard or did not hear. The first is subjective; the second is verifiable.
This article provides a working method for enablement leads, RevOps, and sales-operations practitioners who own a scorecard and need consistent scoring across managers. It covers how to find the behaviors worth scoring, how to write them as trigger-and-action pairs, how to set score anchors that hold their meaning across graders, how to weight criteria so the scorecard reflects the actual sale, and why "not applicable" belongs on the rubric as a designed outcome rather than a gap someone forgot to fill in.
A scorecard exists to answer one question consistently: did this call move the deal, and did the rep execute the specific behaviors that make that happen. When a criterion is phrased as a trait (rapport, energy, confidence, "consultative approach"), two graders listening to the identical recording will land on different numbers because each is scoring an impression rather than an event. One manager hears warmth in a fast-talking rep; another hears a rep who did not let the prospect finish a sentence. Neither is wrong. The criterion gave them nothing to check against, so they each supplied their own definition.
The calibration problem is a writing problem. Alignment sessions do not fix a criterion that has no fixed referent in the call. The only way two graders reliably agree is if the criterion points at a moment in the transcript that either happened or did not, followed by an action that either occurred or did not. That is what trigger and action gives a scorecard: a shared object to evaluate, rather than a shared feeling to negotiate.
This also changes what a low score means. On a trait-based scorecard, a "2" on rapport is an opinion about the rep's personality. On a trigger-and-action scorecard, a missed criterion means one identifiable thing: the moment arrived and the expected action was not taken. That failure is coachable; a personality assessment is not.
A scoreable criterion has two parts, and both are required.
The trigger is a specific, identifiable event inside the call: a thing the prospect said or did, or a phase the call reached. "The prospect mentions pricing." "The prospect names a competitor." "The call reaches the first five minutes." "The prospect describes a problem in their own words." A trigger is checkable: a grader can point to the timestamp where it happened, or confirm that it never happened at all.
The action is the specific, observable output the rep is expected to produce once the trigger fires: the literal phrase the rep should say or the literal move the rep should make, not a category of good behavior. "Asked what was driving the evaluation." "Quantified the cost of the problem in dollars or hours." "Named the competitor's actual limitation rather than deflecting." "Asked for the next meeting before ending the call."
Put together: when [trigger], the rep [action]. That sentence structure is the core method. Everything else in this article is about applying it well: finding the right triggers, writing the right actions, setting anchors, weighting, and testing.
Compare the two forms directly:
Unscorable (trait)Scoreable (trigger and action)Built rapportWhen the call opened, asked a question about the prospect's role or team before pitchingHandled objections wellWhen the prospect raised a concern about price, addressed it before moving to a new topicStrong discoveryWhen the prospect described a problem, asked a follow-up that quantified its impactControlled the callWhen the prospect asked a question outside the rep's agenda, answered it and returned to the agenda within two exchangesCreated urgencyWhen the prospect described a deadline or trigger event, tied the proposed timeline to it explicitly
Each right-column entry can be checked against a transcript by any grader. Portability is the objective: a criterion that only the rubric's author can grade consistently is not a shared criterion; it is an unrepeatable judgment call.
Writing the sentence correctly is necessary but not sufficient. The trigger and the action both have to be real, specific, and tied to something that correlates with the outcome the organization cares about. Otherwise the scorecard is measuring the wrong thing precisely, which is worse than measuring nothing because it produces false confidence.
Start with call evidence. If you are deciding between platforms rather than writing criteria yet, our guide to tools that score sales calls objectively covers how the options differ. The common approach is to convene managers for a brainstorming session on "what good looks like." That produces trait language every time, because traits are what people reach for when they are describing a feeling rather than reviewing a transcript. Pull a set of calls that closed and a set that did not, from the same stage and the same segment, and listen for the specific moment where the two sets diverge. The observable divergence is the specific action winning reps took that losing reps did not take. Often it is one exchange: the winning rep asked a quantifying follow-up question when the prospect named a problem; the losing rep said "got it" and moved on.
This is where a scorecard earns its weight, because the behaviors worth scoring are not evenly distributed across a call. A competitive mention in discovery is a high-leverage trigger: what the rep does in the next thirty seconds determines whether the deal survives that competitor's involvement. A generic rapport-building moment in the first sixty seconds is lower leverage. If a scorecard treats both as equally worth a point, it is not reflecting the actual sale; it is reflecting an even split that was easier to design.
A workable test for whether a candidate criterion is scoreable: can two people, given only the transcript and no context on who the rep is, independently mark the same score. If the answer depends on tone, energy, or "how it felt," the criterion is not ready: it needs a more specific trigger and a more literal action. If the answer depends only on whether specific words or a specific move appear at a specific point, it is ready.
Once a set of trigger-and-action criteria exists, the instinct is to give each one equal weight and average them. That treats every moment in the call as equally consequential, which is not true for any real sales motion. A missed close attempt on a first meeting where a next step was never proposed does more damage to a deal than a missed rapport-opener. The weighting should be set by asking, for each criterion, how much it moved outcomes in the won/lost review, not by dividing 100 points evenly across however many criteria fit on a page.
A practical approach: group criteria by the call phase they belong to (opening, discovery, value articulation, objection handling, close) and set a phase-level weight based on where the deal review showed the calls diverged between won and lost. If most losses trace to weak objection handling and most wins share a strong quantified discovery, those two phases should carry more of the score than the opening or the close. Within a phase, weight the specific criteria the same way: the criterion tied to the highest-leverage trigger (a named competitor, a stated budget constraint, a stalling signal) carries more points than a criterion tied to a routine, low-stakes moment.
Weighting requires periodic revision. A pitch that changes after a launch, a new competitor that starts coming up in more deals, a shift in average deal size: each of these can change which behaviors are predictive, and a scorecard that never gets reweighted drifts out of relevance while still producing a number that looks precise.
A trigger-and-action criterion tells a grader what to look for. It does not, by itself, tell them how to score a partial or a poor attempt at the action, and that gap is where calibration breaks down even on well-written criteria. The solution is the same discipline applied one level deeper: write the score anchors as specific, quoted-style descriptions of what each score sounds like in the transcript, not as adjectives.
Take the criterion: when the prospect raises a competitor, the rep asks what is driving the evaluation.
Anchors written this way turn the rubric into something closer to a set of worked examples than a scale of adjectives. A grader is no longer deciding what "somewhat asked" means; they are matching the call to the closest anchor. Once anchors are set and calls start getting scored, the next question is what to do with the results, see our guide to turning AI call scores into 1:1 coaching plans. This is also what makes the rubric usable by someone who was not in the room when it was written (a new manager, a QA hire, or a peer reviewer grading a sample for calibration) because the meaning of a 3 does not depend on a grader's familiarity with internal scoring conventions.
A scorecard built on triggers has an outcome that a trait-based scorecard never has to confront: some criteria will not apply to a given call, because the trigger never fired. If the prospect never raises a competitor, the criterion "when the prospect raises a competitor, ask what is driving the evaluation" has nothing to score. This is the rubric working as designed.
The instinct to avoid is forcing a score anyway: marking it a "3" by default, or excluding the call from scoring altogether because "the criteria did not fit." Both moves corrupt the data. A default score treats a moment that never happened as though the rep handled it moderately well, which drags the aggregate score toward the middle and hides which reps are strong. Marking the call unscoreable throws away every other criterion that did fire on that call.
The correct handling is an N/A designation: the criterion sits out of that call's score entirely, and the call's final score is computed only from the criteria whose triggers fired. This has a second benefit beyond scoring accuracy: it turns the scorecard into a diagnostic of how often each moment occurs in the pipeline. A criterion that comes back N/A on nearly every call is either testing a rare event (which is worth documenting) or is written around a trigger too narrow to be useful. A criterion that never returns N/A is written so broadly that it fires on every call regardless of what happened, which is its own warning sign that the trigger is not specific enough.
The trigger-and-action pattern applies the same way at every phase of a call, but the triggers worth watching change as the call moves forward. A usable scorecard should have criteria distributed across the phases where behavior diverges between winning and losing calls, rather than clustering everything in one phase because it is the easiest to observe.
Opening. The trigger is the start of the call itself. The action worth scoring is not "built rapport" but something concrete: did the rep confirm the agenda and the time available, or ask a question about the prospect's role before moving into pitch. A call that opens without either is easy to spot and easy to fix.
Discovery. The highest-density phase for good triggers, because prospects volunteer specific things: a problem, a deadline, a stakeholder, a budget range. Each of those is a trigger with an obvious paired action: when the prospect names a problem, does the rep quantify it; when the prospect names a stakeholder, does the rep ask about that person's role in the decision; when the prospect mentions a timeline, does the rep tie it to something concrete rather than letting it pass.
Value articulation. The trigger is usually the rep's own transition into pitching, or a prospect question about capability. The action worth scoring is whether the rep connects the capability back to something the prospect said in discovery, rather than delivering a generic feature list. This is one of the few phases where the trigger is something the rep does rather than the prospect; the pattern does not require the prospect to initiate every trigger.
Objection handling. This is the highest-density phase for competitive and pricing triggers. "When the prospect raises price, does the rep address it directly before changing topics." "When the prospect names a competitor, does the rep ask what is driving the evaluation": the opening example of this article. These triggers are the highest-leverage moments on the call because how they are handled determines whether the deal survives past that meeting.

Close. The trigger is the end of the call approaching. The action is whether the rep proposes a specific next step with a date attached, rather than ending on "I will follow up" with no commitment. This is one of the easiest criteria to write and one of the most predictive, because a call that ends without a next step is a call that stalls.
A scorecard does not need a criterion in every one of these phases on every call type: a fifteen-minute qualification call and a technical deep-dive will emphasize different phases. The discipline is the same regardless: pick the phase, find the trigger that recurs in that phase, and pair it with an action that a rep can be coached toward.

Step 1: Pull the evidence before writing a single criterion. Gather a sample of won and lost calls from the same stage and segment. Listen for the specific moment, not the general impression, where the two groups diverge. Write down the trigger exactly as it occurs and the action the winning reps took, in their own words if possible.
Step 2: Draft each criterion as "when [trigger], the rep [action]." Resist any draft that reads as an adjective ("effectively," "clearly," "confidently"). If the draft cannot survive having those words deleted, the action is not specific enough yet.
Step 3: Stress-test the trigger for specificity. Ask whether two people, given only a transcript, would agree on whether the trigger occurred. "The prospect seemed hesitant" fails this test. "The prospect asked about the cancellation policy" passes it.
Step 4: Write anchors for 1, 3, and 5 as concrete descriptions of what each score sounds like, following the pattern above. Skip 2 and 4 initially; they can be filled in once the outer anchors are stable, and forcing five distinct anchors too early produces filler language.
Step 5: Set eligibility conditions explicitly. State which calls the criterion applies to (discovery calls, first meetings, calls where the competitor is named) so N/A is a designed outcome rather than a grader's improvisation.
Step 6: Weight by leverage, not by count. Assign weight based on how strongly the criterion's phase and trigger correlated with the win/loss split found in Step 1, not by dividing points evenly across however many criteria exist.
Step 7: Pilot on a fixed sample before rolling out. Give the draft rubric to two or three graders and have them independently score the same ten to fifteen calls without discussing them first. Compare scores criterion by criterion, not just on the final total: a scorecard can average out to agreement while individual criteria disagree wildly, which hides exactly the problem the pilot exists to find.
Step 8: Recalibrate where graders disagree. Wherever two graders land on different scores for the same criterion on the same call, the criterion is still carrying an unresolved trait. Go back to Step 2 for that criterion specifically: tighten the trigger, make the action more literal, or split the criterion in two if it was testing two different things at once.
Step 9: Re-run the pilot after each rewrite, on the same calls, until agreement holds. A rubric is ready to ship only after independent graders land on the same score for the same criterion without discussion.
A criterion returns N/A on almost every call. Either the trigger is rare (which is worth documenting, and possibly worth moving to a specialized scorecard used only when that trigger applies) or the trigger is written too narrowly. "When the prospect names this specific competitor by name" might need to widen to "when the prospect names any competing vendor."
A criterion never returns N/A. This indicates the trigger was written broadly enough to fire on every call regardless of content, which turns it into a trait criterion expressed in trigger language. Tighten the trigger to something that could plausibly not happen.
Graders agree on the score but for different reasons. This appears in review conversations rather than in numeric agreement, and is easy to miss because the pilot in Step 7 only checks whether scores match. Ask graders to state which moment in the call they scored against. If two graders point at different moments and still land on the same number, the criterion is ambiguous even though it looks calibrated.
The rubric grows every quarter and never shrinks. Every new initiative or pitch change adds a criterion; few processes prompt removal. A scorecard with more than roughly eight to ten scored criteria per call type means graders are spending more time on bookkeeping than listening, and the low-leverage criteria from Step 6's weighting exercise are candidates for removal, not just low weight.
A rep's score does not move even after coaching. If a rep was coached on a specific criterion and the next several calls still miss it, check whether the roleplay or coaching given afterward was anchored to the actual call and moment the rep missed, rather than a generic version of the skill. A criterion built on a real trigger requires correction tied to the specific instance: a rep who missed a competitor mention on a specific deal is better served practicing that exact scenario than a generic "objection handling" drill. This is the loop Hyperbound Perform closes on live calls: it scores every deal against the team's scorecards, surfaces the criteria the rep is fumbling, and feeds those misses back into targeted practice.

A trigger-and-action criterion tells a grader exactly when to look and what observable behavior to score. For example, "when the prospect mentions a competitor, the rep asks what is driving the evaluation." This removes personality judgments and makes two graders far more likely to score the same call the same way.
Different managers score the same call differently because most scorecards use trait-based language like "built rapport" or "strong discovery," which has no fixed referent in the transcript. A trigger-and-action scorecard addresses this by pointing each criterion at a specific moment and a specific observable action, so two graders are checking the same thing instead of comparing impressions.
Most scorecards should have between six and ten scored criteria per call type. That range covers the phases where behavior actually diverges between won and lost calls without turning grading into bookkeeping. More than ten usually means low-leverage or trait-like criteria have crept back in.
No. Eligibility should be set per criterion. A first meeting and a technical deep-dive do not share the same triggers, so each criterion should specify which call types it applies to. That prevents forced scores and artificial N/A results.
Enablement or RevOps usually pulls the won/lost call evidence, while a frontline sales manager sanity-checks whether each trigger and action reflects a real good call. Reps are valuable reviewers of whether an action is achievable in the moment, but they are not the primary authors.
Weight criteria by how strongly each trigger and phase predicted outcomes in the won/lost review, not by dividing points evenly. Objection handling and quantified discovery often carry more weight than opening rapport or a generic close because they are higher-leverage moments in the real sale.
N/A means the trigger for a criterion never fired, so that criterion is excluded from the call's score instead of being forced to a default number. This keeps the data honest and also shows how often each trigger occurs in the pipeline.
Recalibrate when the pitch, competitive environment, or average deal profile changes meaningfully, not on a fixed calendar. A scorecard tied to an outdated pitch scores reps against a sales motion that no longer exists.
Yes. Because each criterion names an exact trigger and observable action, coaching can start from a specific call timestamp and behavior rather than a vague impression. That makes feedback more actionable and easier for reps to apply.