How an 800-seat NBFC collections floor moved from a 2% sample to 100% call auditing
6 min read · 18 November 2025
A Delhi NCR lender ran compliance QA on two calls in a hundred. Quallix Audit reviewed all of them, and the first month surfaced 31 risk patterns the scorecard was never designed to catch.
The floor
The client is a non-banking financial company running an 800-seat in-house collections floor out of Delhi NCR, calling across five languages on early-bucket and hard-bucket recovery queues.
Compliance exposure on a floor this size is not theoretical. RBI fair-practice conduct rules govern what a recovery agent may say, when they may call and how they must identify themselves. A single pattern of coercive language, repeated across a shift, is a regulatory event — not an isolated mistake.
The problem: QA on two calls in a hundred
A nine-person QA team listened to roughly 2% of calls, chosen semi-randomly and weighted towards agents who had already been flagged. Every review took the reviewer the full length of the call plus form-filling time.
That sampling rate produced two failure modes. First, findings were arguable: an agent could always claim the sampled call was unrepresentative, and the floor manager had no evidence to counter that. Second, patterns that lived in the 98% were invisible. If a specific script line drifted into non-compliant phrasing across a team, nothing in the process would surface it until a customer complaint arrived.
The QA cycle itself ran two weeks behind the calls it reviewed, so coaching always addressed behaviour the agent had already repeated hundreds of times.
What changed with Quallix Audit
Quallix Audit was connected to the existing dialer, so no agent workflow changed. Every completed call is ingested, transcribed and diarised, then scored against a configurable rubric built from the client's own QA sheet plus the default RBI fair-practice conduct checks.
Scores are not the output on their own. Each finding is anchored to a timestamp and a quoted utterance, so a disputed score opens directly to the seconds of audio and the transcript line that produced it. Arguments about representativeness ended in the first week, because the sample was the population.
Risk patterns are clustered rather than reported call by call. Instead of 4,000 individual low scores, the compliance lead sees a ranked list: which phrasing, which queue, which shift, how many occurrences, trending up or down.
Results after the first quarter
Coverage went from about 2% of calls to 100%. The QA team stopped transcribing and form-filling and moved to reviewing flagged segments, which compressed the review cycle roughly six-fold — findings now reach a team lead the next morning rather than a fortnight later.
Thirty-one distinct risk patterns were flagged in the first month, of which eleven were phrasings that had never appeared on the manual scorecard at all. Two were fixed with a script change on the same day they were surfaced.
Outcome
The QA team did not shrink. It moved up the stack: from listening to calls to designing rubrics, arbitrating disputed findings and running targeted coaching against evidence.
In the client's words: they went from arguing about a 2% sample to reviewing every single call.