← All resources

AI Assessment Writing for NZ Vocational Training

15 September 2026 · 9 min read

AI Assessment Writing for NZ Vocational Training

AI Assessment Writing for NZ Vocational Training

AI can speed up drafting vocational assessments in New Zealand, but on its own it hasn't proven reliable at passing national external moderation. Independent 2025 research found AI-generated baseline assessments failed Workforce Development Council moderation, while AI-adapted versions of already-approved assessments performed strongly — so the safer use is contextualising validated content, with a qualified assessor signing off every output.

That distinction — drafting versus adapting — is the whole ballgame for PTEs, ITPs and other tertiary education organisations (TEOs) shopping this category. Vendors will happily demo how fast their tool writes an assessment. Fewer will show you exactly where a human has to stop, check, and take accountability before that draft goes anywhere near a learner or a moderator.

What "AI assessment writing" actually means here

In the NZ vocational context, an AI assessment-writing tool typically does one or more of:

  • Drafting assessment tasks and instructions against a listed unit standard or skill standard.
  • Mapping generated content to elements, performance criteria and evidence requirements.
  • Adapting an existing, already-approved assessment for a different delivery mode (workplace, classroom, distance) or learner group (including ESOL and LLN learners).
  • Generating supporting material — assessor guides, marking guides, training resources — linked to the assessment.

None of this changes who is accountable. NZQA's guidance for tertiary providers doesn't ban AI in assessment or moderation; it expects TEOs to hold a current, coherent policy on acceptable use. That's a materially different position from NCEA's external-assessment rules, which sit in the schools sector and don't apply here.

Can AI-drafted assessments actually pass moderation?

This is the question every buyer should ask before anything else, and there's now a real NZ data point to answer it with.

Manukau Institute of Technology, working with the Construction and Infrastructure Centre of Vocational Excellence (ConCOVE Tūhura), tested AI-generated vocational assessments against national moderation in 2025. The baseline AI-drafted assessments failed — moderators flagged level inappropriateness and a misunderstanding of NZ vocational assessment principles. But AI-personalised versions of assessments that had already been through validation and approval performed well and drew strong feedback from expert reviewers.

The takeaway isn't that AI is unfit for this work. It's that AI is currently far more reliable at adapting something a human has already validated than at originating an assessment unsupervised. Any vendor claiming their tool writes moderation-ready assessments straight out of the box should be asked, specifically, what evidence supports that beyond their own demo.

Keeping your assessor accountable when AI drafts the content

TEOs holding consent to assess must keep meeting their Consent and Moderation Requirements (CMRs) against every standard on the Directory of Assessment and Skill Standards (DASS) — including when assessing or moderation staff change — and must take part in national external moderation run by Workforce Development Councils. None of that obligation moves onto a piece of software.

Flow diagram showing AI drafting an assessment through to assessor validation and national moderation

A workflow that keeps assessors accountable should make the human step visible, not incidental:

  1. AI drafts or adapts content against the loaded standard.
  2. A qualified assessor checks the draft against the standard's actual elements, performance criteria and evidence requirements — not just a general read-through.
  3. The assessor edits, approves, or rejects the draft and records that decision.
  4. Only assessor-approved material enters your validation and moderation-readiness records.
  5. National external moderation proceeds as normal, against human-owned, human-signed-off assessments.

If a vendor can't point to where step 2 and 3 sit in their product — separate from the AI generation step — that's a gap worth pressing on.

Handling the unit standard to skill standard transition

The DASS is progressively moving unit standards to skill standards. That's not a cosmetic relabelling — it changes what's current and assessable. An AI tool that doesn't track this accurately can quite easily draft or map an assessment against a standard that's already been superseded, which is a compliance problem, not just an inconvenience.

Worth asking any vendor directly: how does your system know which version of a standard is current, and what happens to assessments already built against a unit standard once it converts to a skill standard? A vague answer here is a red flag regardless of how polished the rest of the demo is.

Evaluation criteria: what to actually weigh

Set feature lists aside for a moment. These are the questions that separate a genuinely useful tool from a liability:

  • Standards accuracy — does it track current DASS data precisely enough to avoid mapping against a superseded standard?
  • Human sign-off design — is the assessor's validation step built into the workflow and visible, or is the AI draft presented as a finished decision?
  • Mapping built for checking — is the mapping to elements, performance criteria and evidence requirements presented in a form an assessor can actually verify, quickly?
  • Moderation-readiness support — does it help you produce validation-ready documentation without implying it replaces the assessor's or moderator's judgement?
  • Evidence over claims — has the vendor shown how their tool handles a real standard from your own portfolio, rather than only a generic demo?
  • Fit across delivery modes and learners — can it genuinely adapt for workplace, classroom and distance delivery, and for ESOL or LLN learners, without drifting from the standard?
  • Ongoing monitoring — does it flag DASS or CMR changes that affect assessments you've already built, so nothing quietly goes out of date?
Checklist of criteria for evaluating AI assessment-writing tools for NZ vocational training

What to ask a vendor before you sign

Ask them to walk through a real standard from your own portfolio, live, and show you:

  • Exactly where the standards mapping happens, and what it looks like for an assessor to check it.
  • How the tool behaves if that standard is mid-transition from a unit standard to a skill standard.
  • Where, precisely, human review and sign-off sit in the workflow — not described in a slide, but demonstrated in the product.
  • What happens to a rejected or edited draft — does it feed back into the system, or just disappear.

No AI assessment-generator has a published, independent NZQA audit of its output at the time of writing. That doesn't rule products out — it means vendor time-saved figures and outcome claims deserve testing against your own standards before you take them as fact.

Key takeaways

  • AI-drafted assessments have failed national moderation in independent NZ testing; AI-adapted versions of already-approved assessments have performed well — adaptation is the safer use case today.
  • NZQA expects TEOs to hold a clear policy on acceptable AI use in assessment and moderation, not to ban it outright.
  • Your CMRs, consent to assess, and national moderation obligations stay with your organisation and your assessors — no tool changes that.
  • Any tool must correctly track the unit standard to skill standard transition on the DASS or risk mapping against superseded standards.
  • Treat vendor time-saved and outcome claims as claims to verify on your own portfolio, not settled fact — there's no public NZQA audit of a specific product yet.

Our take

The category is genuinely useful, and the honest research so far backs a specific, narrower use than a lot of marketing suggests: AI is good at extending and adapting content a qualified assessor has already validated, and much shakier at originating assessments unsupervised. A buyer who treats this as "AI writes the assessment" is setting themselves up for a moderation problem. A buyer who treats it as "AI drafts, my assessor decides" is using the technology roughly where the evidence says it currently belongs.

We build VETos with that distinction in mind — mapping generated content against a standard's elements, performance criteria and evidence requirements in a form meant for an assessor to check before use, and tracking DASS changes including the unit-to-skill-standard shift so existing materials don't quietly fall out of date. That's a design choice, not a guarantee, and we'd rather you test it against your own standards than take our word for it.

FAQ

Can AI actually draft assessments that will pass NZQA national external moderation? Not reliably on its own. Independent 2025 research from Manukau Institute of Technology and ConCOVE Tūhura found AI-generated baseline assessments failed Workforce Development Council moderation, while AI-adapted versions of already-approved assessments performed well. Adaptation under human review, not unsupervised drafting, is the evidence-backed approach.

How do we keep an assessor accountable when AI is drafting content? By keeping a visible sign-off step: a qualified assessor checks each AI draft against the standard's elements, performance criteria and evidence requirements, then approves, edits or rejects it before it enters your validation or moderation records. Your CMRs and consent to assess obligations stay with your organisation regardless of the tool.

How does an AI tool handle the shift from unit standards to skill standards on the DASS? A reliable tool tracks which version of a standard is current on the DASS and flags or updates assessments as unit standards convert to skill standards. Ask any vendor to demonstrate this on a live standard from your own portfolio, not a generic example.

What should we ask vendors to see where human review sits in their workflow? Ask them to walk through standards mapping and the unit-to-skill-standard transition on one of your real standards, and to point to the exact screen or step where an assessor validates and signs off the output before it's used or submitted for moderation.

Does NZQA ban the use of AI in vocational assessment? No. NZQA's guidance for tertiary providers expects organisations to maintain a current, coherent policy on acceptable AI use in assessment and moderation, rather than prohibiting it outright. This is distinct from NCEA's external-assessment rules, which apply to the schools sector.

If you're evaluating this category, the most useful next step isn't another demo — it's handing a vendor one of your own standards and watching, in detail, where their software stops and your assessor's judgement takes over.

Share

See VETos on your own scope.

A 30-minute walkthrough — bring a unit of competency and watch a validation-ready draft take shape.

VETos is coming to the UK.

Join the early-adopter programme and help shape it for FE, ITPs and EPA.

Join the waitlist