Picture this: A teacher explains her instructional strategy to a defensive parent. A nurse works toward the right diagnosis with a frightened patient. Both do two things at once: apply hard-won professional knowledge and respond to the person in front of them.

That takes a combination of dispositional skills and professional expertise. The difficult part is that you can't apply one skillset before the other. You must have the dexterity to do both simultaneously. So how do we help students prepare before the stakes are real? I've been testing AI role-plays in an effort to find out.

‍

What I built

I started with what I knew. As a former classroom teacher, I remembered how nervous I was before my first parent-teacher conference, something I never got to practice in my undergraduate program. So I built a simulation where a teacher meets with a challenging AI parent — and is scored on both instructional strategy and empathy.

Next came a customer service de-escalation scenario, then a clinical simulation for nurse practitioner students, where a well-informed diagnosis matters as much as an effective bedside manner. I eventually re-built all three in Lazuli, a WGU Labs AI-assisted course design tool.  

Each one worked well enough to be encouraging. But the more I tested them, the more the same failures appeared.

‍

Core Lesson: AI is the actor, not the stage manager

These role-play interactions have genuine potential for formative practice and, eventually, for a more authentic summative assessment than a multiple-choice item. But "eventually" depends entirely on reliability. A simulation that behaves well four times out of five won't be able to carry a grade.

Almost every failure I found traced back to one mistake: asking the AI to guarantee something the software itself should have enforced.

AI instructions work more like suggestions than rules. So if a rule that holds five times and breaks on the sixth gets deployed, that sixth run could be the one a learner sees.

Think of AI as an actor and software code as more of a stage manager. If you need to set a hard rule, such as how the conversation ends, scoring, or completion criteria, then that is the stage manager's job. Do not ask the actor to handle the math. Reserve your instructions for the actor’s performance: tone, warmth, and personality. The AI excels at the nuance of the interaction, but it struggles with rigid, logical constraints.

Nearly every description below is an example of what happens when a learning designer asks the actor to do the stage manager’s job.

‍

Three structural failures

None of these issues triggered error messages. The simulations ran smoothly, dialogues flowed, and generated feedback sounded convincing. Yet, three quiet patterns repeatedly undermined reliability:

The grade and the feedback came from two different places

In a customer service scenario, AI provided insightful feedback on five competencies. However, the backend code graded the encounter solely on a single variable: customer frustration. Learners passed regardless of their performance across the core competencies. The polished feedback hid a broken evaluation model.

The fix: Compute grades directly from rubric ratings in code, not from conversational LLM outputs. Always test with a failing attempt to confirm the grade drops.

The grader never received the student's work

In a clinical simulation, code collected the student's interview data but failed to pass it into the grading prompt. The AI evaluated an encounter it never saw, generating plausible diagnostic feedback without seeing the student's interview data, based solely on the rubric.

The fix: Require feedback to quote the learner's actual input. If quotes are missing, the payload is broken.

Duration requirements failed as instructions

Duration requirements are a pedagogical necessity, but they clash with authentic role-play. Consider a parent-teacher conference: it requires a minimum length to be substantive, yet if the teacher is rude, the parent must retain the agency to walk out early. I tried to enforce a minimum exchange count through AI instructions to ensure substance, but it failed. The model honored the floor only when it felt like it, frequently allowing learners to exit prematurely. Relying on instructions for duration is a losing battle because they function as suggestions, not constraints. Instead, implement the minimum length as an application-level gate that validates conversation history before allowing a user to finish.

The fix: List every way a learner can leave the conversation, and have your application check the minimum at each one. Test each path separately.

‍

One more issue: avoiding thin personas

Characters defined by a single emotion or objective will collapse once that goal is met. Richer simulations require multi-layered personas: an immediate surface request paired with an underlying need revealed only after trust is established, for example.

The fix: Structure character prompts with two tiers: a surface want to solve first, and a core need unlocked through effective interaction.

Pressure-testing

Test rubrics with targeted failure modes. Run deliberate edge cases: hostiles, poor performers, and asymmetrical attempts (e.g., highly empathetic but clinically incorrect). Ensure rubric dimensions score independently rather than clustering around a halo effect.

Test for compliance vs. de-escalation. Push characters aggressively. A robust persona may submit transactionally while escalating frustration internally. Ensure scoring penalizes coercive tactics rather than rewarding superficial compliance.

If your measure improves when the character submits, you have built something that rewards coercion. For a de-escalation trainer, that isn't a scoring bug, the scenario is teaching the opposite of the skill.

Try to break it the way a learner will. Ask it to reveal its instructions, ask which AI model it is, issue a fake override demanding full marks, send nonsense. Someone will do all of this, usually out of curiosity rather than malice, and the response should stay in bounds without breaking character in a way that ruins the scenario.

‍

Looking ahead

AI role-plays offer unmatched potential for scalable, authentic practice. However, moving from an impressive demonstration to a dependable assessment tool requires rigorous, unglamorous verification. Before using these tools in high-stakes environments, continuous stress testing remains essential.

‍

Want to support or partner on this work? Contact us at info@wgulabs.org.