Dialogue assessment on the platform follows a hierarchical structure that aligns with the logic of professional personnel evaluation (assessment centers and competency models):
This approach makes assessment:
In this article, we'll walk through the assessment framework using the scenario "Explaining the Value of a New Phone Plan to a Skeptical Customer" as an example. In this scenario, a customer service representative must address a long-time customer’s skepticism about the price of a new phone plan. The challenge is to effectively communicate the value of the plan without sounding pushy, while addressing the customer’s concerns and maintaining their loyalty.
✨This isn't the ready-made scenario. In our marketplace, you'll find many other useful templates that you can copy for your own use. Learn more about it in this article.
All the fields you see under the heading "Assessment Criteria" are competencies.
A competency is a group of skills that manifest in a dialogue and help achieve a goal: reach an agreement, defuse tension, clarify a situation.

A competency is too broad a concept for objective assessment.For this reason, it's broken down into specific elements — indicators.
An indicator is a specific action that shows that a competency has been demonstrated.

For an indicator to be reliably assessable from text, it must meet the following quality criteria:
| Rule | What it means |
|---|---|
| The indicator must be measurable | The indicator should reflect a specific action or phrase, not a subjective impression. For example, not "manages tone of voice" (which can't be measured), but "clearly formulates thoughts" (which is visible from the text). |
| 1 indicator = 1 idea | Avoid combining multiple actions into one item. For example, discussing the problem and agreeing on next steps are two different manifestations of a competency, and they're better assessed separately. |
| The indicator must be tied to the dialogue's objective | The indicator should assess not just any action, but only those that help achieve the goal of this conversation. |
An indicator level shows how well the indicator was demonstrated in the dialogue.
The platform uses a 1-to-4-star scale, where each value corresponds to a behavioral description. The more stars, the better the skill was demonstrated.

Let's examine the logic using the indicator "Clearly explains plan features" as an example.
The manifestation levels for this indicator look as follows:
| Level | Value | Example behavior |
|---|---|---|
| ★ | Worsens the situation | The representative provides unclear or misleading explanations of the plan features, causing customer frustration or confusion. |
| ★★ | Not demonstrated | The representative fails to address the customer's concerns or provide relevant information about the plan features. |
| ★★★ | Acceptable level, but with room for development | The representative explains the plan features but does so in a way that is partially effective or lacks personalization to the customer's needs. |
| ★★★★ | Excellent level | The representative clearly and empathetically explains the plan features, effectively addressing the customer's concerns and demonstrating how the plan aligns with their specific needs. |
The platform gives you flexible options for customizing assessment. You can change:
If your company already has an approved corporate competency matrix (your own assessment criteria), provide it to the system during the scenario creation stage.
The best way is to upload the description in the very first prompt, either as text or a PDF file. The platform will be able to use your matrix as a foundation. If you don't do this, you'll have to manually configure all indicators and levels to match your standards.
You don't have to describe competencies and indicators from scratch. If, when creating a scenario, you don't provide the system with a description of your desired assessment criteria, the platform will automatically generate a version of the assessment system. It'll analyze the context of the role-play (roles, dialogue goal, situation) and suggest a relevant set based on best practices in assessment.
⚠️ Auto-generation creates a foundation for the assessment criteria, but the final version is produced after your review.
Even if the wording looks correct, take a critical look at it:
Write to us in the chat — we're always happy to help 💬