Every ticket and call audited. People spend their time only where it matters.
Your QA bots proved that AI can audit tickets. The QMS (Quality Management System) is the dedicated tool that makes it complete, reliable and trustworthy, for every campaign. This page explains what's coming, how it will work, and what it means for each role.
Everything shown uses anonymized or made-up sample data. Nothing here is live.
What the bots taught us
Your QA bots showed that AI can judge customer conversations. What held them back was everything around the judging: one long prompt was doing the auditing, the bookkeeping and the integrations all at once.
| What happened | What the QMS does instead |
|---|---|
| Failures were silent. An email connector showed “connected” and sent nothing for 19 days. | Stalls, failed runs, expiring connections and unconfirmed emails raise alerts. Nothing counts as “sent” without the email provider's confirmation. |
| Settings lived in prompts. Bot IDs, agent names, SLA windows and recipients were pasted into prompt text. | A new campaign is configuration, not prompt editing: helpdesk, scope, hours, roster, routing, all in one setup. |
| Mistakes were hard to undo. One false-positive auto-fail (a password shared per client procedure) needed an SOP edit. | Auto-fails wait for a person to confirm; an overturn can propose a rule exception so the rule stops firing on that case. |
Three principles
The AI judges. Code counts.
The AI reads the conversation (or listens to the call) and returns, for each rubric criterion, a verdict with the evidence it's based on. Everything else (fetching, scope, arithmetic, routing, alerts) is ordinary, predictable code.
People handle the exceptions.
The AI audits 100% of interactions. People see only what needs a human: auto-fails, safety issues, unclear cases, disputes, plus a weekly blind sample to keep the AI honest.
Nothing fails quietly.
If a run stalls, a connection expires or an email isn't confirmed, someone is told: that day, not three weeks later.
From conversation to coaching
1 · Collect
Pull new tickets and calls from Zendesk, Gorgias or Aircall. Keep only what's in scope; split messages by agent.
2 · Facts
Compute what doesn't need judgment: response times, SLA, hold and silence on calls, who spoke when.
3 · Judge
For each criterion: pass, deduct, or “can't tell”, with a quote or call moment as evidence.
4 · Score & check
95 minus deductions, plus WoW. Check every quote really is in the conversation.
5 · Route
Normal audits are accepted automatically. Exceptions go to an analyst; auto-fails and safety alert right away.
6 · Act
Analysts review exceptions. Team leads coach on patterns. Senior QA tunes the AI with calibration.
Who does what
| Who | Does | Never does |
|---|---|---|
| AI | Reads or listens; decides pass / deduct / can't tell per criterion; quotes the evidence; suggests a better reply. | Calculate scores, decide who gets an email, change your helpdesk. |
| Code | Fetching, filtering, deduplicating, timing and SLA maths, call metrics, scoring, routing, alerts, exports. | Judge quality. (That's what the AI and your people are for.) |
| People | Confirm or overturn auto-fails, review exceptions, score the blind sample, coach agents, own the rubric. | Re-check the thousands of routine audits by hand. |
The QMS only reads from your helpdesk and phone system. It never writes back to tickets, and it never contacts agents directly; team leads stay the channel.
A day in the life, by role
Each role sees its own handful of screens. Each film (one to four minutes) follows a fictional person through a day with the prototype. The names are made up; the roles are yours.
How the work flows
Each workflow is marked In the prototype (you can click through it today) or Planned (designed, not built yet). Dashed steps are planned.
How a score is calculated
The same rules your bots used
- Every audit starts at 95.
- Each criterion either passes or takes its full deduction. No partial credit, and no evidence means a deduction (unless it doesn't apply).
- WoW: +5 for exceptional, evidenced behaviour. The maximum is 100.
- Auto-fail sets the score to 0, but only once a person confirms it.
- If the AI can't tell on any criterion, there's no score; it goes to an analyst.
Example: active reading missed (−7), template placeholder left in (−5):
Why code does the maths
In the bots, the AI added up its own score, and sometimes got it wrong. The same ticket was once scored both 95 and 100. In the QMS the AI only gives verdicts; code calculates the score, the same way every time. If the AI's own total disagrees, that's flagged.
Every score traces to evidence
Open any audit and you see each criterion's verdict, the exact quote (or call moment) behind it, and “How it's calculated”. The quote is checked against the real conversation. If it can't be found, it's marked unverified.
How calls are audited
In the QMS, every call is audited from the recording, by the same rubric as tickets.
- Call facts from the phone system: duration, queue wait, hold, outcome, which agent. Exact, no listening needed.
- The recording plus a transcript that labels who spoke when. Zendesk Talk and Aircall record in mono (both voices on one track), like your analysts hear them, so the speaker labels come from the transcript.
- Metrics in code: talk share, dead air, and (where the labels are precise) response time and overtalk.
- The AI listens and judges each criterion, quoting what was said.
- Code finds the quote in the transcript, so every piece of evidence jumps to the exact moment. If it can't, it's marked unverified.
Illustration with made-up content. In the prototype, open a call and click the evidence chip.
How we know the AI is right: calibration
Nobody has to take the AI on faith. Every week, the system measures it against your analysts and shows where it agrees, where it's too lenient, and where it's too harsh, criterion by criterion.
1 · Sample
Each week, a few audits that were accepted automatically, across channels, bands and agents.
2 · Score blind
Your analyst scores each on the same rubric, without seeing the AI's verdict.
3 · Compare
Per criterion: agree, AI too lenient, AI too harsh.
4 · Settle
Who was right? The AI, the analyst, or is the rubric unclear? The settled answer also corrects that audit.
5 · Measure
Agreement per criterion over a rolling month, and how it changes with each instruction version.
6 · Fix & test
Improve the AI's instructions or the rubric wording. A new version only goes live if it does better on past settled cases.
Criteria earn trust one by one
A criterion is accepted automatically only once it agrees with your analysts often enough, against a target your QA team sets. Until then its verdicts keep going to people. Go-live can be gradual: no “big bang”.
Why scores won't swing after a change
When a prompt change goes straight to production, every score can shift overnight. In the QMS, every change to the instructions is tested against cases your people have already settled before it touches a real score.
All films
Narrated walkthroughs of the prototype. The narration is a synthetic voice, and all data is sample data.
Q&A
Straight answers, including where something isn't decided yet. If your question isn't here, send it to us. We'll add it.
From prototype to go-live
A proposed path. Each step starts only when you're comfortable with the previous one.
Prototype
Clickable screens on sample data, to agree on how it should work.
Shadow mode on a first campaign
The QMS audits everything alongside today's bot and your human audits. Only confirmed auto-fails and safety items are routed.
Calibration
Weekly blind samples until each criterion meets your agreement target.
Go-live & expand
The first campaign goes live, an Aircall campaign next, then the others. The old bots retire one by one.
What we've proposed: please confirm
| Exceptions-only review | The AI audits everything; analysts review exceptions plus a weekly blind sample. Everything else is accepted automatically. |
| Auto-fails need a person | An auto-fail stays off the agent's scorecard until an analyst confirms it. |
What we need from you
| Transcription add-ons | Do the pilot campaigns' accounts have Aircall AI Assist, or Zendesk Copilot or Zendesk QA? This decides how precise call metrics can be. |
| One test call per platform | A short staff-to-staff call (no customers) on Zendesk Talk and Aircall, so we can check the recording and transcript format. |
| Your questions | Anything this page doesn't answer. We'll add it to the Q&A. |
| Once the direction is agreed | The pilot campaign's rubric and SOP, the pilot contacts (QA analyst, team lead), and where team leads record coaching today. |
Using the prototype
Start with the Walkthrough
It opens on three short guided flows. Each step takes you to the right screen as the right role.
Switch roles
The menu at the top right changes who you are (QA Analyst, Senior QA, Team Lead, Trainer, Ops, Leadership), and the navigation changes with it.
Reset demo & Prototype notes
Reset demo undoes everything you clicked. Prototype notes shows what's sample data, what's a placeholder and what's still to be decided.