Guides
How Scorecard Design Works in Intryc: Criteria, Modes, and Weighting
Author
Published
Intryc
September 17, 2026
A QA scorecard in Intryc is the set of criteria - standard, critical, or soft - that AutoQA and human evaluators score every conversation against, each independently weighted and optionally linked to a knowledge-base article. It matters because a scorecard that can't tell a non-negotiable compliance check from a nice-to-have tone note either lets serious failures slip through at a high average score, or inflates fail rates on things that don't deserve to sink an evaluation. Scorecard design is for QA leads, compliance owners, and CX ops managers who need the scoring model to match how their business actually judges quality - not a fixed rubric a vendor built for someone else's support org.
I'm Alex Marantelos, co-founder and CEO of Intryc. Our scorecard model was shaped by Intryc's Head of QA, who previously served as Head of QA at Deel - so the criterion modes below come from someone who has had to design a scorecard a real compliance team would trust, not a generic rubric template. (Deel is also an Intryc customer today, running it in production across 1,000+ users - a separate fact from this team member's earlier role there.)
Why Does Scorecard Design Matter for QA and CX Leaders?
Most QA tools give you one dial: a weight. That's fine for grading tone, and it breaks the moment you need one check - a compliance disclosure, a data-protection step - to be non-negotiable while everything else on the same scorecard stays proportional. If "confirmed the customer's identity" and "used a friendly greeting" both just add points, a rep can skip the identity check, still hit a high score, and nobody's alerted until an auditor asks. That's the gap most QA scorecards don't close: they can weight importance, but they can't distinguish "important" from "must never fail."
Intryc's answer is your scorecard, your rules: unlimited scorecards and criteria, run in parallel per queue, channel, team, or agent type - human or AI. But "your rules" only means something if the rules have real mechanics behind them. Intryc scorecards support unlimited criteria with three modes - weighted, critical (one failure fails the evaluation, optionally at zero weight so it gates without inflating scores), and soft (feedback-only while a new standard is piloted) - with per-criterion weighting and knowledge-base-linked criteria so the AI grades against your policy.
That distinction matters most for the QA leads who are also accountable for compliance: a scorecard that scores 100% of conversations - or an intelligent, attribute-based sample, your choice, same scorecard, same 90% accuracy - guaranteed - is only as trustworthy as the criteria doing the judging.
How Does Scorecard Design Work in Intryc?
- Pick a criterion mode. Standard criteria contribute score times weight proportionally, like a normal rubric line. Critical criteria are gates: a full markdown (not a partial one) fails the entire evaluation regardless of every other criterion's score, and a critical criterion can carry zero weight - so passing it adds nothing to the score, but failing it still zeroes the evaluation. That's the pattern for compliance disclosures and data-protection steps that shouldn't inflate a score just for being present, but must never be missed. Soft criteria are scored and shown to the agent for coaching, but excluded from the total - the way to pilot a new standard before it counts.
- Set the weight, separately from the mode. Weighting runs 0-10 per criterion. Mode controls whether one failure is fatal; weight controls how much a passing or failing criterion moves the score. Treating "important" and "critical" as the same decision is the most common scorecard design mistake - it either inflates fail rates on things that aren't true deal-breakers, or lets a real deal-breaker hide inside a proportional average.
- Link criteria to your knowledge base. Criteria can point at specific KB articles, macros, and tools, so AutoQA is grading solution correctness against your actual policy - not tone alone, and not a generic "was the answer good" judgment call.
- Make criteria channel-aware. One quality dimension can carry a different bar for chat versus email versus voice without duplicating the criterion across three scorecards.
- Change criteria without rewriting history. When a criterion's instructions change, historical scores stay intact - they are not recalculated under the new criterion. A scorecard that silently re-grades the past on every edit can't be trusted as an audit trail; Intryc's can.
- Track accuracy per criterion, not per platform. Every criterion gets its own accuracy number, and the system separates two different problems that get lumped together everywhere else: an ambiguous instruction (fix the criterion) versus evaluators disagreeing with each other (run a calibration session) - two different diagnoses, two different fixes.
A practical example: a fintech compliance team sets "confirmed customer identity before discussing account details" as critical, weight zero - it never inflates the score, but a miss fails the evaluation outright and routes to compliance review. The same scorecard scores "empathetic tone" as standard, weight 4, and pilots a new "offered proactive account tip" check as soft for a month before deciding whether it counts. All three run in the same evaluation, on the same ticket, scored independently.
On conditional logic: if a vendor's demo shows scorecard questions that branch - "only ask B if A failed" - on their own live account, ask to see it work on a scorecard they didn't pre-build for the demo. That's a fair diligence question for any QA vendor, Intryc included, before you rely on it for a workflow.
Scorecard Design in Intryc vs. a Fixed QA Rubric
| Dimension | Intryc scorecard design | Fixed/legacy QA rubric |
|---|---|---|
| Criterion modes | Standard, critical, soft - each solving a different problem | Usually one mode: every checkbox just adds or subtracts points |
| Non-negotiable checks | Critical mode, settable to zero weight - fails the evaluation without inflating the score | Weighted like everything else, or handled as a manual override outside the tool |
| Weighting | 0-10 per criterion, independent of mode | Often fixed at rubric creation, one scale for the whole sheet |
| Policy grounding | Criteria link directly to KB articles, macros, tools | Evaluator judgment against a static written policy doc |
| Channel handling | One criterion, channel-aware sub-bars for chat/email/voice | Separate scorecards duplicated per channel, or no channel distinction |
| Editing history | Historical scores stay intact after a criterion changes | Scores can shift retroactively when a question is edited |
| Accuracy visibility | Per-criterion accuracy, split into instruction problems vs. evaluator disagreement | One undifferentiated accuracy or agreement number, if measured at all |
| Coverage | Every conversation - human and AI - or an intelligent sample, same scorecard | Typically a fixed random sample, same rubric either way |
The short version: a fixed rubric tells you what to check. A scorecard built with modes, weighting, and KB links tells you what to check, how much it should matter, whether it's a deal-breaker, and whether the AI is grading against your actual policy or just guessing at tone.
Scorecard Design Examples
- Fintech compliance gate. A critical, zero-weight criterion for identity verification or a required disclosure - never adds to the score, always fails the evaluation and routes to review if missed.
- Multi-channel support org. One "resolution clarity" criterion with channel-aware sub-bars, so chat's terser bar and voice's more conversational bar don't require two separate scorecards.
- Piloting a new standard. A soft criterion tracks a proposed new check - "offered a proactive next step" - for a month, visible to agents for feedback, before a QA lead decides whether to promote it to standard and give it a weight.
- KB-linked solution accuracy. A criterion pointed at a specific troubleshooting article so AutoQA checks the agent's answer against that article's steps, not a general sense of whether the reply "sounded right."
- Human and AI on the same scorecard. The same criteria set scores both a support rep and an AI agent's (Ada, Decagon, Intercom Fin) conversations under their own identities, so a QA lead can see whether the bot's criteria are holding steady or deteriorating against the same bar as the humans.
Scorecard and QA Tools
- Intryc - AI-native scorecards with three criterion modes, per-criterion weighting, KB-linking, and per-criterion accuracy tracking, scored on every conversation or an intelligent sample.
- MaestroQA - Legacy QA platform with a custom metric builder and conditional-logic scorecard options; evaluate its specific branching capability in a live demo on your own scorecard.
- Zendesk QA (Klaus) - Native-to-Zendesk scorecard tool, weighted criteria, tied closely to the Zendesk ticketing ecosystem.
- Scorebuddy - Configurable rubric builder with scoring groups and auto-fail options.
- Playvox - Workforce and QA scorecards aimed at contact-center operations.
- EvaluAgent - QA scorecards paired with calibration and coaching workflows.
- Observe.AI - Contact-center AI platform with QA scoring layered on voice analytics.
How to Get Started Designing a Scorecard in Intryc
- List your non-negotiables first. Identify the checks that must never be missed - compliance disclosures, data-protection steps, safety language - and mark them critical, at zero weight if they shouldn't also inflate the score.
- Group the rest by weight, not by gut feel. Decide which standard criteria matter most on a 0-10 scale before building the scorecard, so weighting reflects a real priority order rather than whatever got added last.
- Link criteria to your actual policy. Attach the relevant KB article or macro to any criterion that's really asking "did the agent follow procedure," so AutoQA is grading against your documentation, not inferring intent.
- Pilot new standards as soft criteria. Before a new check counts toward the score, run it soft for a few weeks and use the feedback to decide the right weight and mode.
- Watch per-criterion accuracy, not one number. Once evaluations run, check which criteria have low accuracy and whether the fix is an instruction rewrite or a calibration session with your evaluators.
Scorecard Design FAQ
What is scorecard design in Intryc? Scorecard design in Intryc is choosing which criteria evaluate a conversation, which of three modes each one runs in (standard, critical, or soft), how much weight it carries (0-10), and whether it's linked to a knowledge-base article - all configurable per scorecard, with unlimited scorecards and criteria per workspace.
What's the difference between a critical and a standard criterion? A standard criterion contributes score times weight proportionally to the total. A critical criterion is a gate: a full markdown on it fails the entire evaluation regardless of every other score, and it can be set to zero weight so it gates without adding points for simply being present - the right pattern for compliance checks that must never be missed but shouldn't inflate a score.
Does changing a scorecard question rewrite past scores? No. When a scorecard question changes, historical scores stay intact; they are not recalculated under the new criterion. That matters for anyone using QA scores as an audit trail over time.
Can Intryc scorecards branch or show conditional questions? Conditional or branching scorecard logic isn't something we document as a confirmed capability today. If that matters for your workflow, ask any vendor - Intryc included - to demonstrate the exact conditional logic you rely on, on your own scorecard, before you commit to it.
How does knowledge-base linking work on a criterion? A criterion can link directly to specific KB articles, macros, and tools, so AutoQA checks the agent's response against that documented policy or procedure, not just tone or general helpfulness.
Does Intryc score 100% of conversations or a sample? Both, your choice. Intryc evaluates 100% of conversations in real time, or an intelligent attribute-based sample targeted by risk, sentiment, agent, or ticket type - your choice, same scorecard, same 90% accuracy - guaranteed either way. That's a meaningful jump for most teams, since most QA programs today review less than 5% of conversations.
Can the same scorecard evaluate human agents and AI agents? Yes. Scorecards can run on the same criteria for both, and Intryc imports AI-agent conversations from Ada, Decagon, and Intercom Fin under the bot's own identity, so the chatbot is sampled, scored, and trended as its own agent - on the same scorecard as the humans, or its own.
Further Reading
- How Dispute Resolution Works in Intryc - what happens when an agent disputes a human or AI-scored evaluation, and how accepted disputes feed criterion accuracy.
- How Calibration Sessions Work in Intryc - blind scoring, a lead-reviewer benchmark, and the side-by-side comparison that separates AI disagreement from evaluator disagreement.
- AI QA Scoring: How Accuracy Is Measured - the accuracy formula behind the 90% accuracy - guaranteed commitment.
- Intryc vs. MaestroQA - a head-to-head on scorecard flexibility, dispute handling, and pricing model.
- What Is AI QA? - the category definition this page sits under.

