AI Role-Play Scripts for Vocational Assessment in NZ
21 September 2026 · 8 min read
Yes — NZQA does not ban AI-drafted role-play scripts for internal vocational assessment. Its firm restriction on generative AI applies specifically to NCEA external assessment in schools, not to tertiary or vocational internal assessment. What NZQA does require is that your organisation holds a current generative-AI policy and runs a coherent moderation system where a qualified assessor still verifies and signs off every competency decision.
That distinction matters because it's easy to over-read the NCEA rule as a blanket prohibition. It isn't. For PTEs, ITPs, wānanga and other providers delivering against unit standards or skill standards, NZQA's framing sits at the provider-policy and moderation level — not a ban on the technology itself.
What NZQA actually requires — policy, not prohibition
NZQA's Academic Integrity Guidelines for training providers and standard-setting bodies expect an up-to-date, published policy on how generative AI may be used in assessment. That obligation sits with your organisation, not with any software vendor's marketing claim.
A few things worth being precise about:
- The generative-AI ban you may have read about is specific to NCEA external assessment in schools. Applying it unchanged to PTE or ITP internal assessment isn't supported by NZQA's own published guidance.
- There is no NZQA-published standard or accreditation category called "AI-drafted role-play scripts." No independent NZ certification scheme exists for this kind of software either.
- If a vendor claims to be "NZQA-approved," check it directly against NZQA's own pages. NZQA accredits providers and standards — it doesn't endorse individual AI tools.
Does an AI-drafted script still need assessor sign-off?
Always. Nothing about AI generation changes who is accountable for the competency decision. NZQA's moderation rules require assessor judgements to be consistent, verifiable and appropriate to the standard being assessed — and standard-setting bodies report annually on exactly that consistency, whether the underlying materials were drafted by a human or by AI.
NZQA's own assessor-training material adds a further caution: simulations should only be used where they're common assessment practice for that particular standard, and assessment should generally draw on at least two methods with independent verification. That's a useful reminder not to lean on a single AI-generated role-play as the sole piece of evidence for a competency decision, however polished the script looks.
In practice, this means an AI-drafted script is a starting point for evidence gathering — not a substitute for the assessor's judgement call on whether a candidate has actually demonstrated competency.
Keeping role-play assessment consistent when scripts are AI-generated
Script quality alone doesn't guarantee assessment reliability. Academic literature on role-play as an assessment method backs this up directly: one interrater reliability study found agreement between assessors ranged from poor to excellent depending on the case, and only improved with structured rubrics, scoring guides and case-specific exemplars.
So the real reliability lever isn't the script — it's the scaffolding around it. Before you rely on any AI-generated role-play for assessment purposes:
- Map it explicitly to the specific unit standard or skill standard it's meant to assess.
- Check the standard's current status in the Directory of Assessment and Skill Standards — NZQA is progressively transitioning unit standards to skill standards over roughly a two-year window, and an outdated mapping undermines the whole assessment.
- Pair the script with a structured rubric, scoring guide and exemplar responses, not just the scenario text.
- Calibrate assessors against that rubric across a cohort, so judgements stay consistent regardless of which assessor is running the role-play.
What to look for in an AI role-play tool
When you're comparing options, judge them on the accountability chain — not on how creative or realistic the generated scenario sounds. Look for tools that:
- Clearly separate script and scenario generation from marking, leaving the competency decision with a qualified human assessor.
- Map outputs transparently to the specific unit standard or skill standard, including its current DASS status.
- Support your organisation's own generative-AI-use and academic integrity policy, rather than asserting their own compliance status in its place.
- Supply assessor guides, rubrics and exemplar responses alongside the generated scenario — not just the role-play script on its own.
- Offer some way of checking consistency of assessor judgements across a cohort, given NZQA's expectation of verifiable, consistent decisions.
- Keep a clear record of what was AI-generated, what a human reviewed, and what was formally signed off — useful evidence if NZQA moderation ever asks.
- Stay scoped to internal assessment. Be wary of any tool that implies it can substitute for constraints that apply specifically to external or NCEA-style assessment.
Where AI genuinely helps — and where it doesn't
Where AI earns its keep is speed: turning source material — a unit standard, a workplace scenario, an existing assessment tool — into a draft role-play, simulation or coached practice scenario in minutes rather than the weeks a full instructional-design cycle usually takes. That's a real, practical gain for PTEs and ITPs trying to keep training current without a large instructional-design team.
What it doesn't do is remove the need for professional judgement. Nova, Supahuman's instructional design engine, is built around that line: it decides when a role-play, simulation or narrated video is the right format for a piece of source material, but it's explicit on its own product pages that it "does not mark, and it never makes a competency decision — your assessors do, exactly as they do today." Generated activities come with assessor guides, rubrics and exemplar responses, positioned as moderation support rather than a replacement for assessor review. Supahuman also describes a companion capability that compares assessor judgements for consistency across cohorts and produces moderation-ready validation reports — again framed as support for human moderation, not a substitute for it. Mast Academy, a named NZ PTE, is cited on Supahuman's site as having used the wider VETos platform to speed up course creation while maintaining quality — one customer reference, not a market-wide claim.
Key takeaways
- NZQA's generative-AI restriction targets NCEA external assessment in schools; internal vocational assessment is governed by provider-level policy and moderation requirements instead.
- No AI-drafted script counts as validated evidence on its own — a qualified assessor must still verify it and sign off the competency decision.
- Reliability research shows script quality doesn't guarantee consistent assessment; structured rubrics, exemplars and assessor calibration are what actually hold judgements steady.
- Check any AI-generated role-play against the current Directory of Assessment and Skill Standards status of the standard it targets, given NZQA's ongoing unit-standard-to-skill-standard transition.
- Judge tools on whether they separate drafting from marking and support your organisation's own AI policy — not on unverified "NZQA-approved" claims.
Our take
The interesting buying decision here isn't "can AI write a decent role-play scenario" — most reasonable tools can. It's whether the tool's design keeps the accountability chain intact: script generation clearly separated from marking, outputs mapped to the actual standard, and a moderation trail that survives scrutiny. A tool that treats scenario generation as instructional-design support, with assessor guides and rubrics attached, is doing something fundamentally more useful to a compliance-conscious PTE than one that just produces a more entertaining script.
FAQ
Can we legally use AI to draft role-play scripts for internal vocational assessment under NZQA rules? Yes. NZQA's outright generative-AI restriction is specific to NCEA external assessment in schools. For internal vocational assessment, NZQA's guidance requires providers to hold a current generative-AI policy and run a coherent moderation system — it doesn't prohibit AI drafting outright.
Does an AI-drafted role-play script count as validated assessment evidence on its own? No. NZQA's moderation rules require assessor judgements to be consistent and verifiable, and assessor-training material cautions against relying on a single simulation as sole evidence. A qualified assessor must still review and sign off the competency decision.
How do we keep role-play assessments consistent when scripts are AI-generated? Pair every script with a structured rubric, scoring guide and exemplar responses, map it to the current unit standard or skill standard, and calibrate assessors against that rubric. Reliability research shows script quality alone doesn't guarantee consistent judging.
What should we look for in an AI role-play tool to avoid losing assessor accountability? Look for tools that separate scenario generation from marking, map outputs to the current standard (including DASS status), supply rubrics and exemplars alongside the script, and keep a clear record of what was AI-generated versus human-reviewed.
Is there an NZQA-approved list of AI role-play assessment tools? No. There is no NZQA-published standard or accreditation category for "AI-drafted role-play scripts," and no independent NZ certification scheme for this software. Any vendor claim of NZQA approval should be checked directly against NZQA's own published pages.