Do not make every assessment a battle to prove whether a student used AI, and do not assume a return to paper is the only credible response. Instead, decide what the assessment must show, state the AI condition before students begin, and collect a small bundle of evidence : the final product, evidence of process and revision, and a short individual demonstration of understanding. Use supervised, no-AI work when the learning target itself is independent fluency; allow bounded AI use when the target includes critique, research practice, editing, or professional tool judgment; and make AI use an explicit object of assessment when students need to learn to use it responsibly. For a substantial task, a practical default is: an early checkpoint completed in class or otherwise supervised; a final product with sources and a short AI-use disclosure; a two- to four-minute conference, explanation, or comparable individual response; and a rubric that credits subject learning, reasoning, process, and responsible tool use—not “sounding human.” This makes authorship less central than demonstrated learning. It also gives students a legitimate way to use AI where it serves the learning goal, rather than teaching them that concealment is the skill that matters. Current assessment guidance for formal qualifications similarly emphasizes clear rules, acknowledgement, authentication of a student’s own work, and appropriate investigation—not detection alone. JCQ, AI Use in Assessments (2025) Bottom line: assess the learning process and the learner’s judgment as well as the polished product. Use AI detectors neither as a grade nor as the basis for an accusation.
[2][3][4][5]The assessment-design framework: Claim → Conditions → Evidence → Conversation → Judgment
Use this five-part sequence before creating a task. It works with paper, locked devices, shared documents, labs, performances, and projects.
| Step | Teacher decision | Useful question | Output for students |
|---|---|---|---|
| 1. Claim | Name the precise learning students must demonstrate. | Is the target independent recall, disciplinary reasoning, writing craft, collaboration, or responsible AI judgment? | A one-sentence learning target and success criteria. |
| 2. Conditions | Set the tool boundary in advance. | Is generative AI prohibited, bounded, or intentionally integrated? What accessibility tools remain allowed? | A visible AI-condition label on the task. |
| 3. Evidence | Choose two or three complementary traces of learning. | What would let a teacher see thinking, not just a polished answer? | Checkpoints, annotations, source notes, draft/revision evidence, a short performance, or a conference. |
| 4. Conversation | Plan a brief individual check proportional to the stakes. | Can the student explain a key decision, transfer the idea, or respond to a small change? | Two or three predictable prompts and an accessible alternative where needed. |
| 5. Judgment | Score the stated learning, then respond to gaps with more evidence. | Does the evidence support the claim? What is the next instructional step? | A rubric and a normal reassessment/repair path. |
Multiple kinds of evidence are already a sound assessment principle, not an AI-era novelty. Ontario’s K–12 framework, for example, describes evidence from observations, conversations, and student products. Ontario School Effectiveness Framework The U.S. Department of Education likewise frames formative assessment as evidence used to adapt teaching and learning. Artificial Intelligence and the Future of Teaching and Learning, p. 40
Choose the AI condition by learning target
Do not use one class-wide rule for every assignment. Mark each assessment with one of these conditions.
| AI condition | Use it when the assessment is meant to show… | Examples of permitted / prohibited use | Evidence to collect |
|---|---|---|---|
| AI off: independent demonstration | What a student can do independently right now. | No generative AI, translation generation, answer generation, or automated code generation. Approved accommodations and permitted calculators remain available. | Supervised response, performance, problem set, lab technique, or short viva. |
| AI bounded: assistance under declared limits | A student’s own reasoning and creation, with a limited support tool. | Allow idea generation, vocabulary alternatives, feedback on a student-written draft, or debugging hints after a student checkpoint. Prohibit AI-written paragraphs, unverified citations, or generated final answers. | Original checkpoint, final work, brief disclosure, and a question about a choice the student made. |
| AI integrated: responsible use is part of the target | How a student prompts, evaluates, verifies, revises, and takes responsibility for AI-assisted work. | Allow selected AI tools; require prompt/output excerpts or a concise log, source verification, and a critique of at least one error, bias, or limitation. | Decision log, verified sources, annotated AI output, final product, and individual defense. |
“AI off” is not “technology off.” If a student’s approved assistive technology is needed for access, remove only the tool capability that invalidates the particular learning claim; do not casually remove the accommodation. U.S. OCR notes that educational institutions must provide equal access to educational benefits and opportunities delivered digitally, and gives examples such as screen readers, captions, voice recognition, and keyboard access. U.S. Department of Education OCR, Technology Accessibility
Build an evidence bundle, not a surveillance system
For each major assessment, select the lightest combination that supports a valid judgment. A bundle need not mean recording every keystroke or collecting every chat.
A reliable default bundle for a major task
| Evidence | What it reveals | Keep it manageable |
|---|---|---|
| Checkpoint | Initial understanding before substantial outside help. | A 10-minute claim, outline, worked example, annotation, sketch, or code plan in class. |
| Process trail | How the student developed, tested, and revised ideas. | Collect one planning artifact plus two meaningful revisions or a lab/design log—not every intermediate file. |
| Final product | Application, communication, craft, and synthesis. | Score it against subject criteria, not against a guess about authorship. |
| AI-use disclosure | Whether tool use stayed within the stated condition and whether the student evaluated it. | Use a 4–6 line form; do not ask for account passwords or an entire private chat history. |
| Individual demonstration | Understanding, transfer, and ownership of decisions. | Use a 2–4 minute conference, a one-question exit response, a recorded explanation, or an accessible equivalent. |
Revision history can be useful context: it may show when a student made a major change and give the teacher a concrete place to ask for explanation. It is not proof of authorship. A file can have incomplete history, be shared, be pasted into gradually, or be altered; keyboard-level monitoring creates privacy, equity, and workload costs. Treat a revision trail as a feedback and conversation aid, never as a lie detector.
A short AI-use disclosure students can actually complete
Assessment AI condition: Off / Bounded / Integrated
Tool(s), if any:
What I asked the tool to do:
What I used, changed, or rejected:
How I checked facts, quotations, data, code, or sources:
I can explain the decisions in this submission: Yes / I need help identifying a decision to explain.
The disclosure should reward honest, accurate documentation. It should not reward using AI, nor punish a student who did not use it when its use was optional. When AI provides information or sources, students must verify and cite the underlying sources; a tool’s citations are leads, not evidence by themselves. JCQ guidance on acknowledgement and verification
Assessment patterns that work across subjects
These are patterns, not scripts. Adapt the tool condition to the standard and age group.
| Subject | Assessment design | AI condition | Individual evidence of learning |
|---|---|---|---|
| English / language arts | Students annotate a common passage and write a claim-and-evidence plan in class. They draft an argument, revise after feedback, and annotate two revisions that improved reasoning. | Bounded: AI may give feedback on an already student-written paragraph; students identify advice accepted and rejected. It may not generate the argument or evidence. | A 3-minute conference: “Why this claim? Which quotation carries the most weight? What would you change for a skeptical reader?” |
| History / social studies | Students investigate a local or contested historical question using a source set, then make a public-facing briefing or exhibit. | Bounded or integrated: AI may suggest inquiry questions or counterclaims, but students must locate and cite real sources, identify provenance, and verify every factual claim. | Student points to two source decisions and explains how one source complicated or changed the conclusion. |
| Mathematics | Students model a local data set or non-routine scenario, show representations and assumptions, then revise after a “what if the constraint changes?” prompt. | Off for a mastery check; integrated for a later task that compares a student solution with an AI method and diagnoses an error or hidden assumption. | A short worked explanation of one key step and a transfer problem with changed numbers or conditions. |
| Science | Teams conduct a lab, keep individual observation/data notes, analyse results, and write a claim–evidence–reasoning conclusion. | Bounded: AI may help improve clarity after data analysis, but cannot invent data, citations, safety guidance, or results. Students label any AI wording and check it against raw data. | Each student explains one data point, uncertainty, or procedural decision; the teacher samples rather than interrogating every line. |
| World language | Combine a prepared composition with a spontaneous interpersonal exchange and comprehension task. | Usually off for the spontaneous target-language demonstration; possibly bounded for vocabulary practice or feedback outside it, according to the learning target. Do not conflate permitted accessibility/translation support with cheating. | A short conversation, audio response, signed/augmentative response as appropriate, or another equivalent demonstration. |
| Computer science / career and technical education | Students ship a small feature, submit design notes and tests, then handle a live change request or debug a related defect. | Integrated: AI-assisted code may be allowed if students identify it, test it, explain it, and document security/licensing checks required by the course. Off for an individual fundamentals check. | Live demo plus “change this requirement” or explain a selected function and its test case. |
| Visual arts / design | Students create a portfolio with ideation, prototypes, critique notes, and a final rationale for an intended audience. | Integrated only when generative tools are part of the stated medium; disclose inputs, iterations, and borrowed/reference material. Off/bounded when the target is hand technique or original observational work. | A critique in which the student explains visual choices, revisions, and limitations of the medium. |
Authentic tasks are not automatically AI-resistant. Their value is that they create consequential decisions—local evidence, an audience, constraints, feedback, testing, and revision—that students can explain. Pair authenticity with an individual checkpoint.
A usable 100-point rubric
Use the same categories across a unit where possible, then change the descriptors for the discipline. The weight of “responsible AI use” can be zero when AI is off and higher when AI use itself is an outcome.
| Criterion | Points | What strong evidence looks like |
|---|---|---|
| Learning-target achievement | 40 | Demonstrates the specified knowledge or skill accurately and independently to the degree required by the task. |
| Reasoning, evidence, and disciplinary judgment | 25 | Makes defensible choices; uses reliable evidence/data/sources; explains assumptions, trade-offs, or uncertainty. |
| Process, iteration, and response to feedback | 20 | Shows purposeful planning, revision, testing, reflection, or improvement. Process evidence is relevant, not performative busywork. |
| Responsible tool use and provenance | 15 | Follows the stated AI condition; discloses relevant use accurately; verifies claims/sources; distinguishes the student’s judgment from tool output. In an AI-off task, this means following the condition—not proving a negative. |
Use four performance levels for each row:
| Level | Short descriptor |
|---|---|
| 4 – Secure | Accurate, well-supported, and independently explainable; process and tool decisions strengthen the work. |
| 3 – Proficient | Meets the target with sound reasoning; minor gaps do not undermine the core claim. |
| 2 – Developing | Partial understanding or weak evidence; needs a targeted revision or additional demonstration. |
| 1 – Beginning / insufficient evidence | The available evidence does not yet support a judgment; complete a structured reassessment or evidence-recovery step. |
Do not grade “human voice,” an expected writing style, compliance with a preferred drafting platform, or the absence of AI when AI was permitted. Grade the stated outcome and the quality of the evidence supporting it.
How to use oral defense without turning it into an accusation
An oral defense works best when it is routine, brief, and designed into the task—not triggered only when a teacher is suspicious. It can be an individual conference, a recorded explanation, a quick whiteboard response, or another comparable method.
For a project, ask the same small set of prompts of everyone or of a pre-announced rotating sample:
- “Show me the most important decision you made and the evidence behind it.”
- “What did you revise after feedback or testing, and why?”
- “If this source/data point/requirement changed, what would change in your conclusion or design?”
Publish the prompts or their categories beforehand. Give the student thinking time and an appropriate alternative mode of response. A student who communicates through AAC, needs processing time, has an anxiety-related accommodation, is still acquiring the classroom language, or has a different access need should not be forced into an impromptu spoken performance unrelated to the intended outcome. Flexible, barrier-aware ways to express learning align with the CAST UDL Guidelines 3.0, while approved accommodations and local policy remain controlling.
If evidence does not support a judgment, say that plainly: “I need another demonstration of this target.” Then offer a normal, documented reassessment: a supervised rewrite of one section, a new comparable problem, an annotated source explanation, or an accessible conference. Do not begin with “prove you did not use AI.”
Homework: make its role honest
Out-of-class work can still matter. What it cannot reliably do on its own is establish independent mastery of a skill that AI can complete.
| Homework purpose | Design and grading move |
|---|---|
| Practice and retrieval | Make it low stakes; grade completion, reflection, or selected work, and use it to plan instruction. Reassess independent mastery later in class. |
| Preparation for discussion or lab | Ask for annotations, questions, predictions, or an evidence map. Start the next class with a short individual application. |
| Revision and extension | Let students seek feedback within the stated AI condition, then require a revision note explaining what they changed and why. |
| AI literacy | Deliberately assign students to compare an AI response with course sources, identify errors/unsupported claims, and improve a prompt or response. Grade the critique and verification. |
| High-stakes individual claim | Do not rely only on homework. Add a supervised checkpoint and an individual demonstration. |
This is not a reason to give up on homework; it is a reason to stop treating a take-home polished product as the only evidence of individual learning.
Equity, accessibility, privacy, and workload guardrails
Equity and access
- Do not require students to have a personal paid AI subscription, a home device, high-speed internet, or an account that their family cannot create. Offer a comparable no-AI route whenever AI use is optional.
- Keep the cognitive target separate from the method of response when the method is not itself the target. A handwritten essay may measure handwriting stamina, speed, and fine-motor access as well as argument; a rapid oral defense may measure verbal fluency and anxiety tolerance as well as historical thinking.
- Preserve accommodations. Design equivalent evidence, not identical performance: typed, spoken, signed, captioned, screen-reader-compatible, supported, or extended-time responses may all be appropriate depending on the target and plan.
- Teach disclosure explicitly with exemplars. Students unfamiliar with academic conventions should not lose points because “use AI responsibly” was never made concrete.
Privacy and child-safety boundary
Use only school/district-approved tools and accounts. Do not have students put names, grades, disability information, behavioural records, private family details, confidential case material, or identifiable peer work into a public AI service. In U.S. schools, the Department of Education advises teachers to check whether an application is approved and notes conditions applicable when personally identifiable information from education records is disclosed under FERPA’s school-official exception. U.S. Department of Education, student-privacy FAQ
This is not legal advice. Consult the school’s privacy officer/IT team and local requirements before adopting a new service, especially for minors. UNESCO likewise calls for a human-centred, age-appropriate approach and protection of data privacy in education. UNESCO, Guidance for Generative AI in Education and Research
False positives and fair response
Do not use an AI-writing detector to assign a zero, reduce a grade, or initiate discipline. A peer-reviewed 2023 study of seven detectors found a mean false-positive rate of 61.3% for the study’s set of 91 TOEFL essays written by non-native English writers. Liang et al., “GPT detectors are biased against non-native English writers,” Patterns That result does not mean every detector, language, or task has the same error rate; it is enough to show why a detector score cannot establish misconduct fairly.
If a tool flags work, or the work seems inconsistent with available evidence, treat it as a prompt to gather neutral, learning-focused evidence. Ask for a normal conference or comparable additional demonstration, document the evidence and the student’s explanation, and follow the school’s process. Never try to infer guilt from accent, dialect, disability, multilingualism, neatness, vocabulary, typing patterns, or a “vibe.”
Teacher workload
The framework should replace some grading, not create an extra bureaucracy.
- Use an in-class checkpoint for all students, then sample 6–8 short conferences per period on a rotating schedule.
- Score one meaningful revision note rather than every draft.
- Use a common three-question defense bank for a whole unit.
- Ask teams to create a shared product but collect one individual evidence item per student.
- Reallocate points: grade fewer polished take-home products deeply; use low-stakes homework for practice and feedback.
- Keep the disclosure form short. Full chat logs are usually more work, more private, and less instructionally useful than a student’s concise account of decisions.
The risk-management principle is proportionality: select controls that fit the stakes and context, then revisit them as the tool and harms change. NIST, Generative AI Profile (AI 600-1)
A two-week implementation plan
Week 1: establish the routine
- Audit the next three assessments. For each, write the independent learning claim in one sentence. Flag any task whose only evidence is a take-home final product.
- Assign an AI condition. Label the task “AI off,” “AI bounded,” or “AI integrated,” and list permitted tools, prohibited uses, accommodations, and disclosure expectations.
- Add two evidence points. For the next major task, add an in-class plan/annotation/worked-example checkpoint and one brief individual response.
- Show examples. Walk through one acceptable disclosure, one inappropriate undisclosed use, and one AI-assisted response whose sources must be checked.
Week 2: refine without escalating surveillance
- Run the task and conference. Ask the same three questions of a rotating sample or everyone when the task is high stakes; offer accessible alternatives in advance.
- Calibrate the rubric. Mark three anonymized exemplars with colleagues or students. Check whether the rubric accidentally rewards polish over reasoning or disadvantages a response mode unrelated to the target.
- Use results for teaching. If students cannot explain their final product, reteach or reassess the specific concept; do not treat this automatically as a conduct case.
- Update one policy sentence. Keep what produced useful evidence. Remove a requirement that did not change instructional decisions.
Assignment language a teacher can adapt
Learning target: I can make and defend a [discipline-specific] claim using [specified evidence/skill].
AI condition: Bounded. You may use an approved AI tool only after completing the in-class evidence plan. You may ask for feedback on clarity or counterarguments. You may not ask it to write the final submission, invent sources/data, or replace your analysis. Verify every claim and cite the source you actually used.
Submit: your in-class plan, final product, two revision notes, and the short AI-use disclosure. Be ready to explain one decision and respond to one small change in the evidence.
Access: Use approved accommodations and contact me before the checkpoint if you need a comparable way to demonstrate the target.
Common failure modes—and better alternatives
| Failure mode | Why it fails | Better move |
|---|---|---|
| “Just run everything through a detector.” | False positives, easy evasion, and no direct evidence of understanding. It moves attention from learning to suspicion. | Design a checkpoint plus individual demonstration; use a detector score, if policy permits it at all, only as a reason to collect neutral additional evidence. |
| “Ban AI everywhere, forever.” | May preserve some independent measures but ignores AI literacy, creates inconsistent enforcement, and makes homework grades misleading. | Reserve AI-off conditions for genuine independent targets; teach bounded/integrated use in explicit tasks. |
| “AI is allowed” with no boundary. | Students cannot tell whether brainstorming, paraphrasing, translation, source-finding, coding, or final-answer generation is expected. | Label a condition and give examples of permitted, prohibited, and disclosure-required use. |
| Grade only the final product. | A strong product can mask weak understanding; a student who does real work may not get credit for judgement and iteration. | Score product, process, and explanation using a compact evidence bundle. |
| Turn oral defense into an interrogation. | It can be inequitable, humiliating, and impractical; it encourages concealment instead of learning. | Make short, predictable individual demonstrations routine and accessible for everyone. |
| Treat revision history as proof. | It can be incomplete or manipulated and pressures students into monitoring-heavy tools. | Use it as one voluntary/contextual process artifact; accept equivalent planning or revision evidence. |
| Disable accommodations to prevent AI. | It may change the construct being measured and deny access. | Preserve approved accommodations; restrict only the capability that invalidates the stated target. |
| Let AI cite itself. | AI-generated citations and explanations may be inaccurate or unsupported. | Require students to locate, verify, and cite the actual source or data. |
Viable alternatives when this framework does not fit
- Short, frequent supervised checks: Best for foundational fluency, computation, reading comprehension, and writing-to-time tasks. Paper is one option; a locked digital environment may be more accessible.
- Portfolio review: Best for long-form creative, research, design, or performance work. Use selected artifacts, reflection, and a final conference rather than every version.
- Open-resource oral or written application: Best when recall is less important than judgment. Permit sources and ask students to apply them to a novel case.
- Performance task with individual accountability: Best for labs, group projects, and career/technical courses. Observe teams, then collect a brief individual explanation or change request.
- AI critique task: Best when tool literacy is a stated outcome. Give all students the same AI output, then ask them to fact-check, annotate limitations, revise it, and explain their choices.
None of these makes deception impossible. The goal is a defensible, humane, instructionally useful judgment—one that does not depend on a teacher winning a technological cat-and-mouse game.
Limitations and what may change
The framework does not prove authorship, guarantee that AI use will be disclosed, or replace the safeguards required for formal examinations. It also requires time to teach routines and to align accommodations. Its central trade-off is deliberate: gather a little more direct evidence of understanding rather than spend more time trying to infer provenance from a finished file.
AI tools, student access, school policies, vendor terms, and assessment rules change quickly. Review the AI conditions, approved-tool list, data-sharing practice, and detector policy at least each term. The U.S. Department of Education’s AI report stresses that educational technology needs human-centred, equitable design and attention to context; it should be read as guidance, not as a product endorsement. U.S. Department of Education, Artificial Intelligence and the Future of Teaching and Learning
Evidence
Sources used for this answer.
Question signals show what people need. Primary documentation supports the answer. Both remain visible.
- 01Middle School/High School teachers - How are you assessing students when AI is inevitable?Reddit · question signal · checked 25 Aug 2026
- 02Ontario School Effectiveness Frameworkfiles.ontario.ca · primary evidence · checked 25 Aug 2026
- 03Artificial Intelligence and the Future of Teaching and Learning , p. 40U.S. Department of Education · primary evidence · checked 25 Aug 2026
- 04U.S. Department of Education OCR, Technology AccessibilityU.S. Department of Education · primary evidence · checked 25 Aug 2026
- 05JCQ guidance on acknowledgement and verificationjcq.org.uk · primary evidence · checked 25 Aug 2026
- 06CAST UDL Guidelines 3.0udlguidelines.cast.org · primary evidence · checked 25 Aug 2026
- 07U.S. Department of Education, student-privacy FAQstudentprivacy.ed.gov · primary evidence · checked 25 Aug 2026
- 08UNESCO, Guidance for Generative AI in Education and ResearchUNESCO · primary evidence · checked 25 Aug 2026
- 09Liang et al., “GPT detectors are biased against non-native English writers,” Patternsdoi.org · primary evidence · checked 25 Aug 2026
- 10NIST, Generative AI Profile (AI 600-1)doi.org · primary evidence · checked 25 Aug 2026