Setter quality assurance: score DMs across accounts

Setter quality assurance is the routine that pulls a sample of real conversations every week, scores them against a written standard, and turns the gap into one change to the play. It is not a performance review and it is not a booking report. It looks at the only part of the work your team fully controls, which is what gets typed into the inbox before anyone books anything.
This is written for whoever runs several setters or several client accounts: agency owner, head of sales, ops manager. Past a certain size you stop being able to read everything, and the choice is no longer between reading all conversations and reading none. It is between reading a sample on purpose and reading whatever happens to land in front of you on a bad day.
TL;DR
- Score the conversation, not the outcome. Bookings are too rare and too confounded by lead source to coach on week by week.
- Sample on purpose: five conversations per setter per week, stratified, including at least one that was lost after the prospect replied.
- Ten criteria, of which three are pass or fail compliance gates that never get partial credit.
- One change per week, and it goes into the shared playbook. A fix that lives in a private message to one setter is not a fix.
- Across client accounts, keep one scorecard plus a short annex per account. Ten scorecards is how a quality programme dies.
Table of contents
- What setter quality assurance actually measures
- Why booking rates alone cannot coach a setter
- How many conversations to review, and which ones
- The ten criteria scorecard
- How to score without turning the review into a trial
- The weekly review, in forty five minutes
- What changes after a review
- Quality assurance across several client accounts
- The three numbers that tell you the programme works
- Known ways to ruin a quality programme
- Where SetScale fits
- FAQ
- Conclusion
What setter quality assurance actually measures
Three different jobs get confused under the same word, and separating them is most of the work.
Lead qualification asks whether the prospect deserves a call. That is a property of the prospect, and it belongs in a written grid your team applies the same way every time, which is the subject of the lead qualification framework for DM teams. Performance measurement asks what came out of the month: conversations handled, calls booked, calls held. Quality assurance asks a third question, and only that one: given the conversation this setter was handed, did they run the play correctly.
The distinction matters because the first two are properties of the market and the last is a property of your operation. A setter can run a flawless conversation and lose a prospect who was never going to buy, and can run a sloppy one and book a call because the prospect arrived ready. Coach on outcomes and you reward the second while punishing the first, and within a quarter your team has learned to work the easy threads and let the hard ones die quietly.
A scored conversation covers three layers. Compliance: did we respect the platform rules, the consent we actually have, and the promises the client allows. Craft: did we open on the right foot, ask one question at a time, listen to the answer. Commercial: did we ask for the call when it was earned, with a concrete slot rather than a vague invitation. If you are measuring anything the setter cannot change by typing differently tomorrow, you are measuring the wrong thing.
Why booking rates alone cannot coach a setter
Run the arithmetic on your own numbers before accepting this. A setter who handles thirty real conversations in a week and books four calls gives you four data points, spread across different traffic sources, client accounts and offers, two of which were prospects who asked for the call themselves. You cannot separate skill from luck at that sample size, and certainly not fast enough to correct a habit before it hardens. The same thirty conversations, sampled and scored, give a dense signal on the same person in the same week: if four of the five you read show the same missing step, that is a gap in the play rather than noise, and usually a gap for the whole team, since everyone works from the same instructions.
Outcome numbers also arrive late. A call booked this week is held next week and closed the week after, so the feedback loop on a bad habit runs a month, and the sales follow up process for setter teams has the same defect: by the time a lost thread shows up as a missing number, the conversation is cold and the lesson has no owner. Keep those numbers for the monthly report and for agency client reporting, where a client wants to know what their money produced. They simply cannot do the coaching job, and asking them to is how teams end up with a leaderboard instead of a standard.
How many conversations to review, and which ones
Random sampling sounds rigorous and produces a useless read. Most conversations in a social inbox end in the first two messages, so a purely random sample fills up with threads where nothing happened. Stratify instead.
| Slot | What to pull | Why it earns a place |
|---|---|---|
| 1 | A conversation that ended in a booked call | Confirms the play is being followed when it works, and catches promises made to get the booking |
| 2 | A conversation lost after the prospect replied at least twice | The single richest slot: the prospect was interested and something in the exchange lost them |
| 3 | A conversation still open after five days | Shows what your team does with ambiguity, which is where most silent losses live |
| 4 | A random pull from the week | The control. Without it you only ever read the interesting cases |
| 5 | One chosen by the setter | Their pick tells you what they think good looks like, which is worth as much as the score |
Five per setter per week is the working default for a team of three to eight setters. Above eight, hold the number and rotate: every setter gets a full read at least every second week, and nobody goes a month without one.
Across client accounts the constraint is different. Add one rule on top: every client account gets at least one scored conversation per month, whatever the setter rotation says. An account nobody has read in six weeks is where a tone drift or an unauthorised promise is currently growing, and it is the account whose renewal conversation will go badly.
Pull the sample yourself, or have the system pull it. Never ask a setter to send you five conversations to review: slot five exists so that self selection has a bounded place rather than an invisible one.
The ten criteria scorecard
Ten is the ceiling. Every criterion you add past ten costs reviewer attention and buys almost nothing, because raters stop reading and start ticking. The three gates are pass or fail: a conversation that fails a gate scores zero overall regardless of the rest, and it goes on the agenda whatever else the week contains.
| # | Criterion | What good looks like | Type |
|---|---|---|---|
| 1 | Consent and platform rules respected | The message was sent inside the rules the channel allows, on the basis of an interaction the prospect actually initiated | Gate |
| 2 | No promise outside the mandate | Nothing about price, deliverables, guarantees or timelines beyond what the client has written down as allowed | Gate |
| 3 | No sensitive claim | No health, income, legal or outcome claim, in any phrasing, including a softened one | Gate |
| 4 | First reply inside the response target | The first human or automated reply landed inside the target your team publishes | Scored |
| 5 | Opening matched the source | The first line references the ad, the story, the comment or the form the prospect came from | Scored |
| 6 | One question at a time | No stacked questions, no interrogation block, no message that asks three things and gets one answer | Scored |
| 7 | Qualification asked in order and recorded | The grid was followed and the answers were written to the record, not left in the thread | Scored |
| 8 | Objection handled once | Answered, then moved on. An objection argued three times is a lost thread with extra steps | Scored |
| 9 | Call proposed with a concrete slot | Two named times, not an invitation to check a calendar and come back | Scored |
| 10 | Handover note or closing reason written | The closer can open the record cold and know what was promised, or the thread is closed with a reason | Scored |
Criterion 4 depends on having published a target in the first place, which is the subject of the lead response time SLA for teams. Criterion 7 and criterion 10 both depend on the record actually existing where the closer will look for it, which is the whole argument of DM to CRM integration.
Write the "what good looks like" column in your own words before you use this. A criterion that two reviewers read differently is not a criterion, it is a mood.
How to score without turning the review into a trial
Use three points, not five. Zero for not done, one for partly done, two for done. Five point scales feel more precise and produce more disagreement, because the middle three values mean different things to different reviewers and nobody can defend a three against a four.
Calibrate monthly. Two reviewers score the same three conversations independently, then compare. When they disagree on a criterion, assume the criterion is badly written rather than that one reviewer is wrong: rewrite the definition, note the date, move on. Teams that skip calibration end up with scores comparable only within one reviewer, which quietly makes the whole history useless the day that person leaves.
Score the conversation, not the person. In the review itself, refer to the thread by its identifier and the account, not by the setter's name. This is not politeness theatre: it changes what the room talks about. Once a name is on the screen the discussion becomes about whether that person is good, and the actual question, which is whether the play is followed and whether the play is right, disappears.
Publish the aggregate, keep the individual private. Team scores per criterion belong on the wall, individual scores belong in the one to one alongside the conversations they came from. A public per person leaderboard is the fastest way to make setters cherry pick easy threads, which corrupts the sample and then the numbers.
The weekly review, in forty five minutes
The programme dies from length before it dies from anything else. Timebox it hard.
Ten minutes, silent reading. Everyone reads the same two conversations from the sample, cold, without commentary. No preamble, no context given by the person who chose them.
Twenty minutes, scoring and discussion. Score the two conversations against the ten criteria, out loud, criterion by criterion. Disagreement on a score is the useful part of the meeting: it exposes the criteria that are not written clearly enough. Note every disagreement, do not resolve it by seniority.
Fifteen minutes, the one change. Pick a single criterion where the sample is weakest and decide exactly what changes in the play. Write it down before the meeting ends. If nothing was written down, the meeting did not happen.
One change per week is deliberate. Teams that leave a review with six improvements apply zero, because six improvements is another word for no priority. The thirty day ramp for training appointment setters has the same constraint at a different scale: sequence beats volume, every time.
What changes after a review
The output of a review is an edit to a shared artefact, not a message to a person. That is the rule that separates a quality programme from a management habit.
If the sample shows setters stacking three questions in one message, the fix is a rewritten message template in the shared library and a line in the play, not a reminder to the two people who did it in the sample. Everyone else is doing it too, you simply did not read their threads this week.
If the sample shows the call being proposed as an open invitation instead of two named times, the fix is a snippet with the two named times and a rule about when to send it. If the sample shows objections argued into the ground, the fix is a written stop rule: answer once, offer the call, close the thread with a reason if the answer is no. The way threads get closed and re-opened is exactly what the sales follow up process governs, and the two documents should never contradict each other.
Keep a dated log of the changes. Six weeks later, when someone asks why the play says two named times, the answer should be a date and a sample rather than an opinion. It is also what a new setter reads in their first week, and it teaches more than the play itself, because it carries the reasoning.
Quality assurance across several client accounts
Agencies hit a specific wall here. Client A allows a price range to be mentioned, client B forbids it. Client C considers a lead qualified at a different threshold than client D. The temptation is one scorecard per client, and it is a trap: within two months nobody remembers which version is current and the reviews stop.
Keep one scorecard and add a one page annex per account. The annex holds only the things that genuinely differ, which in practice is three to five lines: what can be said about price, what counts as qualified here, what the client wants never to be promised, which channel is in scope, and the response target that account is sold on. Criteria 1, 2 and 3 stay identical in wording and change only in what the annex says is allowed.
Rotate deliberately across accounts rather than letting the busiest one absorb the whole sample. The account generating the most volume is not the account at the most risk. The agency playbook for managing multiple Instagram accounts covers the operational side of running several inboxes at once, and quality review is the part of it that has no technical shortcut.
Be careful with what you show the client. A quality score is an internal instrument, and putting it in a monthly report turns it into a number your team optimises for rather than learns from. Report outcomes to the client, keep the scores for the room where the play gets changed.
The three numbers that tell you the programme works
Coverage. The share of setters and of client accounts that got at least one scored conversation in the month. This is the only number that should be at one hundred percent, and the first one to slip when the week is busy. Track it weekly, not monthly, because a missed week is invisible until three of them have stacked.
Gate pass rate. The share of reviewed conversations that clear all three gates. Anything below one hundred percent is a standing agenda item, and a single failure on criterion 2 or 3 deserves more attention than a whole month of mediocre scores on criterion 6. Compliance failures are the only category here that can cost a client account rather than a booking.
Movement on the criterion you chose. After a change is made, that criterion should visibly improve in the following samples. If it does not, the change was wrong or it never reached the people doing the work, and both of those are worth knowing within two weeks rather than at the quarterly review.
Then, quarterly and separately, the lagging numbers: booked call rate, show rate, what the closers say about handovers. They belong to the setter and closer model and move slowly. Reading them weekly is how you talk yourself out of a good change after one unlucky fortnight.
Known ways to ruin a quality programme
It becomes a performance review. The moment scores feed pay or ranking, setters optimise the score rather than the conversation, and the sample stops describing reality. Keep the two systems apart and say so out loud.
The scorecard grows. Someone adds an eleventh criterion, then a fourteenth, and the review that used to take forty five minutes takes two hours and gets cancelled. Adding a criterion should require removing one.
Only booked conversations get reviewed. This is survivorship in its purest form, and it is the default when the sample is not pulled deliberately. Everything you learn will be about threads that already worked.
Compliance gets folded into the average. A conversation that promised something outside the mandate is not a seven out of ten: it is a failure with an owner and a fix, and averaging it away is how a client account gets lost politely. It is where the channel rules described in the WhatsApp Business API setup for teams and in what click to WhatsApp ads require get checked in a real thread.
Where SetScale fits
To be direct: the product is not open yet. No dashboard, no integration list, no client result to point at. A white label direction is planned for agencies, and a seven day free trial at opening.
It is being built around the idea in this article: the reviewable unit of a setting operation is the conversation, and a team running several closers needs the sample, the standard and the record in the same place rather than in three tools and a spreadsheet. That is AI setting infrastructure for teams and agencies.
If that is the piece missing from your operation, Join the waitlist and you will hear when it opens.
FAQ
We only have two setters. Is a formal quality programme worth it? Yes, and it is cheaper than it sounds: ten conversations a week, forty five minutes. Two setters is also the point where the play gets written for the first time, and reviewing conversations is how you find out what to write in it. The programme scales down better than it scales up.
Should setters see their own scores? They should see their own scores and the conversations behind them, in a one to one, never on a shared board. A score without the conversation attached is a verdict, and it teaches nothing.
Can the scoring be automated? Parts of it. Criterion 4 is a timestamp comparison, criterion 10 is a field that is either filled or empty. The judgement criteria, whether the objection was answered or the call was earned, still need a human reading a thread. Automate the mechanical ones to buy attention for the others, not to replace the review.
Does this apply to an AI setter as well as a human one? The scorecard applies unchanged, and it matters more, not less. An automated responder applies the same mistake to every conversation on every account simultaneously, so a criterion that fails in a sample of five is failing at volume. The gates in particular should be checked more often on automated flows.
Conclusion
Setter quality assurance is not a report and not a ranking. It is the weekly habit of reading a deliberately chosen sample, scoring it against ten agreed criteria, and changing exactly one thing in the shared play as a result. Teams that keep it for a quarter end up with a written standard they can hand to a new setter. Teams that measure only bookings end up with strong opinions and no evidence.
Start with the sample rule this week, before the scorecard is perfect. Five conversations per setter, the five slots in the table above, no scoring the first time, just reading. What is wrong will be obvious within two weeks. Then the three routes to appointment setting puts numbers against building it in house, outsourcing it, or tooling it. And if you would rather the sample, the standard and the record arrive together with the operation, Join the waitlist.