How to Evaluate Students’ Work in the Age of AI: New Assessment Models and Grading Rubrics
The entry of artificial intelligence into the mainstream of education, thanks to tools such as ChatGPT, resembled an earthquake. In the teacher’s room, next to the discussion about lesson plans, there was a question full of doubts: “How can I check if this is my student’s authentic work?” Traditional methods, based on the evaluation of the final essay or paper, suddenly lost their raison d’être, and the feeling of helplessness became a common experience. Trying to ban AI altogether is like building a sand dam in the face of an impending tsunami – a technology that is and will be a fundamental part of our students’ worlds.
The fight against it is lost in advance and, more importantly, it is a fight against progress. The real pedagogical challenge lies not in unmasking, but in wise adaptation. We need to deeply revise our approach, asking ourselves a key question: what do we really want to evaluate? In this in-depth article, which is the essence of one of the chapters of my comprehensive guide “Artificial Intelligence for Schools and Teachers”, I will present you with a practical model for moving from evaluating a dead product to evaluating a living thought process. Instead of building walls, we will learn to build bridges by creating assessment criteria in the age of AI that promote integrity, develop the competencies of the future, and restore the meaning of our work.
Why the traditional approach is no longer outdated
Before we build a new system, we need to fully understand why the old one is in ruins. The first instinct of many institutions was to look for magical “AI detectors”. It’s a dead end. These tools are notoriously unreliable, generate false accusations, and will never keep up with the pace of development of language models. Relying on them is not only ineffective, but also pedagogically harmful.
- How to Evaluate Students' Work in the Age of AI: New Assessment Models and Grading Rubrics
- Why the traditional approach is no longer outdated
- Assessment Philosophy 2.0: Shifting Focus from Product to Process and Added Value
- Definition of the student's own contribution in the 21st century
- Construction of a modern assessment rubric – Descriptive evaluation model
- Can working with AI earn a six?
- From assessment to support: Formative assessment and feedback in a new dimension
- Change your assessment to shape the future
Choose a plan below.
The problem goes much deeper. If the definition of homework comes down to compiling and summarizing publicly available facts (e.g. “Describe the most important achievements of Nicolaus Copernicus”), then AI will not only complete it faster, but often at a higher language level than the student. Further commissioning and evaluation of such works becomes an educational fiction. Artificial intelligence in education is a merciless mirror that exposes the weakness of reproduction-based tasks, and forces us to design learning experiences that require critical thinking, creativity, and personal involvement.
Assessment Philosophy 2.0: Shifting Focus from Product to Process and Added Value
The most effective response to the challenges of AI is to strategically integrate it into the teaching process and treat it as a tool, the skillful use of which becomes a new, key subject of assessment. We no longer ask “Did the student use AI?”, but “How well and wisely did the student use AI?”. Our goal is to precisely evaluate AI works by analyzing the path taken by the student.
Definition of the student’s own contribution in the 21st century
The central concept that we need to redefine is how to evaluate a student’s own contribution. In the age of AI, “self-contribution” no longer means “creating everything from scratch in isolation.” The new definition is “added value” – everything that a student has done with raw material that they could obtain with the help of technology. It is a conscious, intellectual processing that gives the work a unique character.
During the assessment, we should focus on the following competencies, which are worth their weight in gold today:
- The Art of Prompt Engineering (Asking Questions): The ability to dialogue with AI precisely, formulate complex, multi-step commands to obtain valuable, rather than generic, answers. It is an analytical and strategic skill.
- Verification Competencies and Critical Analysis: Was the student just a passive recipient or an active verifier? We assess its ability to detect AI substantive errors (so-called hallucinations), identify algorithm biases, and confront the data obtained with reliable, external sources.
- Ability to Synthesize and Personalize: Was the student able to weave the information from the AI into the broader context of knowledge from lessons, personal experiences or other readings? Did he give the work a coherent, original narrative and structure, or is it just a glued together collection of paragraphs?
- Metacognitive Awareness: We assess the student’s ability to reflect on their own learning process. Can he precisely describe and justify at what moments and for what purpose he reached for AI support, and which elements are the result of his exclusive work?
Construction of a modern assessment rubric – Descriptive evaluation model
To move from philosophy to practice, we need a specific tool. Instead of a rigid table, I present a descriptive model that can be flexibly customized. Below, I break down what an extensive rubric for evaluating work with ChatGPT could look like, describing each level of advancement for each criterion.
Criterion 1: Verification of Information and Substantive Reliability
At the basic level (passing/passing), the student treats the AI’s answer as the ultimate truth. The work may contain factual errors and “hallucinations” copied directly from the model. The lack of any references to external sources indicates minimal commitment to verification.
At the advanced level (good/very good grade), the student demonstrates basic reliability. It checks key facts, dates and names, confronting them with several reliable sources (e.g. online encyclopedias, proven educational portals). Can reject obviously false information and attaches a bibliography confirming the verification process.
At the master level (excellent grade), the student approaches the material from AI with professional suspicion. Not only does it verify the facts, but it actively looks for potential biases in the AI narrative (e.g., an over-focus on the Western perspective). He/she is able to compare different sources, identify contradictions in them and consciously present them in his/her work. He demonstrates an in-depth understanding that AI is a tool, not an oracle.
Criterion 2: Own contribution — Synthesis, analysis and personalisation
The basic level is simple editing of AI-generated text. The student corrects stylistic errors, changes the order of sentences, but the core of the argument and the structure remain intact. The own contribution is cosmetic.
The extended level shows that the student treated the AI text as raw material. You can see a significant interference in the structure here – adding your own introductory paragraphs, enriching arguments with examples discussed in class or coming from your own observations. The work begins to take on an individual character.
The master level is a real synthesis. The student uses AI as a springboard to formulate their own original theses. Information from various sources (AI, manual, articles, own knowledge) is seamlessly integrated into a coherent, authorial narrative. The text is no longer “text with AI”, but a student’s work, in which AI has played the role of an intelligent research assistant.
Criterion 3: Process documentation and metacognitive reflection
At the basic level, the student attaches a short note to the work saying “I used ChatGPT” at best. It is a declaration that does not provide any informative value about the course of work.
At the advanced level, the student attaches snippets of their conversation with the AI, especially those key prompts that defined the direction of work. In a short description, he is able to indicate which parts of the text were created with significant support of technology and justify his decision.
At the master level, the student presents a fully transparent and annotated record of their interaction with the AI. He analyzes which of his questions were effective and which were not, and describes how his understanding of the subject evolved during his dialogue with the machine. It provides a critical reflection on the limitations of the tool he encountered and what he learned about the process of acquiring and processing knowledge itself. This is proof of the highest awareness and control over technology.
Can working with AI earn a six?
Many teachers intuitively feel that work that is not one hundred percent “independent” cannot be evaluated at the highest level. It’s a trap of old thinking. In the new paradigm, the answer to the question “can work with AI be at 6?” is unambiguous: yes, without the slightest doubt.
An excellent grade ceases to be a reward for a solitary restorative effort, and becomes an expression of appreciation for mastery in navigating a new, complex information ecosystem. A student in the AI era is a digital virtuoso who:
- It orchestrates dialogue with AI, asking precise, multi-level questions that go far beyond simple commands.
- He treats AI as an intellectual sparring partner, asking for counterarguments to his own theses, which deepens his analysis.
- He instantly switches between synthesis and verification, using the time saved not for laziness, but for deeper research of the subject.
- He creates work that is a unique added value – something that neither he himself in isolation nor AI alone would be able to create.
Such a six is of much greater value, because it assesses the competencies that will really determine the success of our students in the future.
From assessment to support: Formative assessment and feedback in a new dimension
The new approach to assessment opens up fantastic opportunities for AI-formative assessment. By analyzing the student’s work process (e.g. through an attached transcript of a conversation with AI), the teacher can provide incredibly precise and practical feedback. Such feedback for a student using AI is real coaching.
Instead of writing a vague “little specific,” the teacher might write, “I noticed that your first prompt was ‘describe impressionism.’ This is too broad a command, which is why the AI gave you a generic answer. Next time, try to phrase them like this: ‘Compare the painting technique of Monet and Renoir in the context of their approach to light, give three specific examples of paintings for each artist.’ You will see what quality of answers you get.” This is learning in practice.
Change your assessment to shape the future
Faced with the AI revolution, we have a choice: we can entrench ourselves in positions defending the old methods, wasting energy on an unequal fight, or we can meet it and use it as a catalyst for profound, positive change. Implementing transparent rules for evaluating AI papers is the most effective way to promote academic integrity and develop key competencies. By changing our rubrics, we are changing the definition of educational success.
The knowledge and strategies contained in this article are a solid foundation, but at the same time only a fragment of a broader roadmap that I have prepared for Polish teachers.
Are you ready to navigate education in the age of AI with full confidence and confidence?
I invite you to read the full, comprehensive version of my guide “Artificial Intelligence for Schools and Teachers”. It’s your must-have, complete AI guide for education that translates theory into everyday classroom practice.
Don’t let change get ahead of you. Become their conscious leader. Invest in competencies that will allow you to assess fairly, teach more effectively, and gain confidence in the new digital reality.