Short answer
To measure call centre quality, define what a good contact looks like based on what drives customer satisfaction, loyalty and business outcomes; turn that into a weighted scorecard; evaluate a random, representative sample of contacts (around 400 from a large population gives about ±5% accuracy at 95% confidence); calibrate evaluators regularly; and combine internal monitoring, customer feedback, mystery shopping and AI analysis. The more confidence both sides have in the measure, the more reasonable it becomes to link payments or service credits to it.
Key points
- Start from customer outcomes, not from a generic checklist.
- Sample randomly and representatively — you do not need to review everything.
- Calibrate often so different evaluators score the same way.
- Combine several lenses: monitoring, customer feedback, mystery shopping and AI.
Link quality to outcomes
By understanding the relationship between what happens in conversations and what customers do afterwards — whether they are satisfied, whether they call back, whether they stay or buy — you can define quality measures that genuinely improve the experience. Analyse call recordings alongside customer feedback and repeat-contact data to find the behaviours that matter most.
Build a weighted scorecard
- List the behaviours and outcomes that define a good contact.
- Weight each one by its impact on customers and on business risk.
- Separate critical failures (for example, a missing regulatory statement or incorrect information that harms a customer) from coaching points.
- Write clear guidance for each item, with examples, so evaluators score consistently.
Sample properly
The contacts you evaluate must represent all contacts: across times, days, agents, contact types and channels. They should be selected randomly and reviewed unobtrusively, so that agent behaviour is not affected by knowing which calls will be scored.
The size of the exercise need not be enormous. Statistical sampling means a random sample of about 385–400 contacts from a large population gives results within roughly ±5 percentage points at 95% confidence. If you want reliable results for each team or contact type, you need an adequate sample for each segment.
Use several lenses
| Method | Strength | Limitation |
|---|---|---|
| Internal monitoring | Detailed, coaching-focused | Can be subjective without calibration |
| Customer surveys (CSAT, CES, NPS) | The customer’s own view | Low response rates, biased samples |
| Mystery shopping | Tests the full journey from outside, including access and waiting | Scenario-based, not real customers |
| AI quality monitoring | Covers most or all interactions, spots trends and risks | Needs clear criteria and human calibration |
| Operational data | Repeat contacts, complaints, transfers | Indirect measure of quality |
Calibrate
Have client, provider and any independent evaluators score the same contacts and discuss differences. Track the variance between evaluators and aim to reduce it over time.
Use quality in outsourcing contracts
The more confidence both sides have in the objectivity and accuracy of a measure, the more willing they are to link it to payment, service credits or incentives. That is why independent or jointly governed quality measurement is often used in outsourcing relationships. See contracts and SLAs and our quality assurance services.
Frequently asked questions
What should a call quality scorecard include?
Typically: greeting and identification, understanding the need, accuracy of information, resolution, compliance statements, empathy and tone, ownership, and a clear close. Weight each element by its effect on customer outcomes and business risk, and treat critical compliance failures separately.
How often should quality calibration happen?
At least monthly between client and provider, and more often during transition or after process changes. Internal calibration within the quality team is often weekly.