If AI Can Complete Assessments, What Is Assessment For?
A student turns in a polished essay explaining a historical event. The argument is clear, the paragraphs are organized, and the sources look credible. But how much of that work shows what the student understands? How much was shaped or produced by AI?
That question is becoming harder to answer. AI can now generate essays, solve problems, create presentations, and produce code. AI agents may soon be able to work through multi-step assignments with little direction from a student. When a system can produce the thing we usually grade, we need to ask what assessment is meant to show.
Assessment is often treated as a way to measure learning after it has happened. But assessment also shapes what students choose to learn and how they approach that learning. If a course rewards a finished product without asking how it was made or what the student can do with it, then students may focus on producing the product. AI makes that strategy faster, but it didn’t create the incentive.
The challenge is not simply to determine whether AI was involved. It is to decide what evidence would help us understand a student’s learning.
That evidence will differ by discipline and task. A final answer may matter in one setting. In another, the important evidence may include the choices a student made, how they used feedback, whether they can explain their reasoning, and how well they can apply what they learned in a new situation.
Research on assessment has long emphasized that students respond to what is assessed. In their review of assessment and learning, David Boud and Nancy Falchikov argued that assessment should prepare students to make judgments about their own work beyond the course, not simply measure performance for an instructor. That perspective becomes especially relevant when AI can help create work that looks complete: students need to develop the ability to judge the quality of work, including their own, and understand what they can responsibly stand behind. aieducationalresearch.com
A polished submission alone may not show whether a student can make that judgment. But neither does a brief oral explanation automatically provide a complete picture. No single assessment format can reveal everything we want to know. The task is to choose forms of evidence that fit the learning goal.
If students are learning to write, the final essay matters, but so might the development of a claim, the revision process, and the student’s ability to explain why they made particular choices. If they are learning to troubleshoot, the answer matters, but so does how they narrowed down the possibilities and responded when their first attempt failed. If they are learning to evaluate sources, the conclusion matters, but the evidence they considered and the reasons they trusted or rejected it matter too.
This does not mean that every assignment needs more steps or that faculty must monitor every stage of student work. It means we should be clear about what a task is intended to help students learn and what evidence would make that learning visible.
AI can be part of that process. Students might use it to generate a counterargument, explain a difficult concept, or offer feedback on a draft. Then they can show what they accepted, rejected, corrected, or changed and why. That kind of work can help students practice evaluating AI while also giving faculty a clearer view of their thinking.
The research on generative AI and assessment is still developing. A 2024 review of generative AI in higher education found that early research and practice were already raising questions about how assessment should adapt as AI becomes part of student work. The field has not settled on one solution, and no single assessment method will fit every course. The more useful question is whether our assessments give students meaningful opportunities to demonstrate the capabilities the course intends to develop. www.sciencedirect.com
Some of that evidence may come from a final product. Some may come from a student’s explanation, demonstration, reflection, or application of an idea in a different context. Together, those pieces can show more than any one artifact alone.
The purpose is not to prove that students worked without help. Learning has always involved tools, feedback, and collaboration. The purpose is to understand what students can do, what they have learned, and what they still need to develop.
If AI can complete an assessment, that does not make assessment pointless. It does ask us to examine whether we are assessing the capability we care about or only the product that used to stand in for it.
Next, I want to consider a closely related question: If AI can provide feedback quickly and continuously, why does human feedback still matter?
Continuing the Conversation
Series 1: AI Is Exposing Existing Problems ✓ Completed
Series 2: What We Do About It ✓ Completed
Series 3: Cultivating Human Thinking in an AI World ✓ Completed
Series 4: Learning Alongside AI ✓ Completed
Series 5: Preparing Students for an AI World ✓ Completed
Series 6: The Purpose of Education in an AI World
Current Post (2 of 8): If AI Can Complete Assessments, What Is Assessment For?
Next Up: If AI Can Provide Feedback, Why Does Human Feedback Still Matter?