Marking homework with the help of AI

Teachers who are required to mark written assignments, particularly exercises with a standardised marking scheme.

The experiment focuses on the marking of French exam papers for the Brevet. The existence of a detailed national marking scheme and highly standardised assessment criteria provides a favourable framework: it reduces subjective bias and allows for the precise measurement of discrepancies between human marking and AI marking.

Positive aspects identified

  • The marking itself is very quick.
  • Marking follows the official marking scheme.
  • Generosity is shown when an idea is present but poorly expressed.
  • Clear justification is provided for the marks awarded for each question.

Logbook — a complementary approach

  • The teacher records their feedback; the AI transcribes it, summarises it and suggests a mark.
  • Pupils receive a detailed report via the digital learning platform.
  • Teachers using the system estimate that it saves 50% of their time.

Marking students’ work is time-consuming but helps teachers get to know their pupils better. In 2025, what can we realistically expect from AI for this task? This factsheet draws on feedback from Thibaud Hayette, a literature teacher who has trialled the marking of handwritten French exam papers for the 2024 Brevet, and also presents the Logbook tool, which has been trialled in several local education authorities.

 

Handwritten transcrip

Checking the HTR transcript takes a long time. When the handwriting is difficult to read, the AI generates its own text based on partially deciphered words.

Reintroduced errors

The AI automatically corrects pupils’ spelling mistakes, which must then be re-entered manually in order to be marked correctly.

A complex dictation to assess

Marking dictation requires numerous prompts to categorise errors correctly — even though this is a straightforward task for the teacher. The time-saving benefit is only significant when there are a large number of papers to mark.

IA more generous

Differences in marks were observed between human marking and AI marking, with the latter generally being more lenient.

Objectives

  • Compare AI and human marking to determine whether there is a genuine time saving.
  • Assess the objectivity of the corrections produced by the AI in line with the official marking scheme.

Example of implementation

Overview and objectives of the trial

Marking papers takes up a significant portion of teachers’ timetables (source: the Ministry): the average is reportedly 6.5 hours per week for humanities subjects. Whilst there are tools available to make this process easier (such as La Quizinière’s online multiple-choice question marking system, or a smartphone app that can identify a pupil’s correct answers on paper or a digital whiteboard simply by taking a photo), the use of artificial intelligence seems even more promising (the Annecy-based company Compilatio is currently trialling software designed for this purpose with volunteer teachers) as it will be able to process larger volumes of text, such as an essay, for example.

But does that mean we should entrust the marking of pupils’ work to a machine?
The aim of this experiment was to use actual work produced by secondary school pupils – taken from official exam papers collected during the 2024 French DNB exam – and to compare human marking with machine marking. Which is faster? Which is more objective? Which is more effective? And finally, is the technology sufficiently advanced to be relied upon? In short, who wins the ‘battle’ between man and machine?

To form as objective an opinion as possible, I obtained seven official exam papers, prior to marking [1]. These were seven papers contained in envelopes from seven different markers who were kind enough to take part in the experiment. I scanned them, page by page, in high resolution (600 dpi) using the photocopier at the examination centre, thereby obtaining PDF files (one file per page).

I then returned the papers to the teachers, and once they had been marked, I scanned them again, this time including the marks awarded, a few annotations and other comments, so that they could be compared at a later date.

Handwriting recognition is much better, but there is still plenty of room for improvement

The first hurdle is that exam papers, like most of pupils’ work, are handwritten. For those familiar with the OCR software [2] of yesteryear, the task at hand seemed daunting: converting a pupil’s handwriting—which is not always legible—into digital text [3]. If the marker has to retype everything, the time saved is obviously nil. But here again, artificial intelligence has enabled fantastic progress. I therefore used NotebookLM, one of the interfaces of the Gemini AI [4], repurposing it from its primary use, which is to process and summarise information contained in documents. It is possible to import different types of documents, and so I fed it a first, uncorrected copy page by page.

Here is the prompt I entered after uploading the document: “Can you convert the entire text into a digital version?”. The result then appears in the interface.

Conversion of handwritten text into digital text

It has to be said that the result looks impressive at first glance. However, on closer inspection, numerous alterations can be spotted, likely occurring when the AI attempts to find coherence where none exists, or simply when handwritten character recognition has been ineffective: the renumbering of questions (omitting the a. or b.), missing accents or errors involving grammatical homophones are corrected, for example (this would be particularly problematic for dictation: in such cases, one must request a faithful reproduction of the student’s text, retaining the errors), but sometimes this goes further in paragraphs that have been poorly ‘read’ by the AI. Thus, in the example below, the student’s use of ‘Firstly’ to introduce their first point likely led the AI to expect a ‘Secondly’ that does not appear. The student actually used ‘next’ to introduce the second argument, which is perfectly acceptable, or at least entirely understandable to a human marker. The AI’s rewriting is therefore no more effective, and is often far worse than the original: this can lead to significant biases in the marking, particularly when entire sentences are rewritten, as is again the case here. In its defence, AI sometimes encounters the same difficulties as the human eye when handwriting is difficult to read and the lines are not spaced out (which should be mandatory on official exam papers). Finally, it sometimes happens that the text ‘read’ bears absolutely no relation to what is written: AI abhors a vacuum and, based on a few fragments of sentences, is capable of inventing anything (this is a form of hallucination [5]).

A paragraph from the student’s essay:

Excerpt from a handwritten copy
The paragraph converted to digital text by AI:
‘3. No, the characters do not manage to communicate easily with one another. Firstly, the three officers waited a long time before they could speak to her; after an hour of waiting, she ran away and isolated herself. Secondly, Penanster and the others then realised that she was deaf, which caused problems with communication.”

So I had to check and reproduce exactly what the student had written on each exam paper, to ensure the AI could mark them as objectively as possible. But this is time-consuming, not to mention the time taken to scan the papers and then upload them to NotebookLM.

To conclude this section, the human eye remains indispensable, given the sheer number of changes made by AI—whether in its efforts to improve a text or due to its difficulty in recognising handwriting. Furthermore, the time-saving aspect is highly debatable, and may even work against AI, as everything must be double-checked before moving on to the next stage. Nevertheless, we can highlight the real progress made in terms of HTR [6], and with even more sophisticated tools—which are not yet accessible to the general public but do exist—we can hope that this will no longer pose a problem in the near future.

AI-assisted proofreading

After thoroughly proofreading the entire first draft and compiling it into a single PDF document, I used ChatGPT-4o (free version) to correct it.
Here is what I asked it to do:

Instruction générative de départ donnée à ChatGPT

So I uploaded both documents, and within seconds, the AI provided the initial findings in detail:

Excerpt from the AI-generated responses

That was when I began comparing human proofreading with machine proofreading.

a. Reading comprehension and interpretation skills, and grammar and language skills

Initial findings in terms of marks: there is a difference of just 0.5 marks between the machine’s and the examiner’s marks on the first 11 questions of the first paper (Reading Comprehension and Interpretation Skills, and Grammar and Language Skills). Here are the details of any discrepancies:

Comparison between AI and humans
A reminder of the question Points awarded by the AI Points awarded by the human marker Comments
Question 2: ‘What do they have in common? Two answers are required. (2 marks)’ 1 / 2 2 / 2 The examiner was a little generous compared to the marking criteria
Question 5b: ‘Complete this physical description of Marguerite with a character description by identifying two of her character traits. Justify each character trait by referring to the text. (4 marks)’ 2 / 4 4 / 4 The AI relied on the national marking scheme, which did not mention a character trait that was nevertheless clearly present in Marguerite’s portrait. The examiner therefore awarded the marks quite logically. The AI is not to blame, as it did not even have the source text.
Question 7: ‘Image. Do you think this poster is a good illustration of the text? Explain your answer using two arguments. Each argument must be supported by reference to the text and the image. (6 marks)’ 5 / 6 4 / 6 The question about the image is more open to subjective interpretation, and the difference in marks is minimal. The five spelling mistakes may have influenced the marker.
Question 8: “‘We are,’ he explained to him, ‘a club of officers which currently has three active members who are happy to be benefactors.’ (lines 4–5).
Identify the modifiers of the noun ‘club’ and state the part of speech of each one. (2 marks)”
2 / 2 1 / 2 The examiner penalised the confusion between “clause” and “preposition”, whereas the AI corrected the text itself (once again), making the student’s answer correct.
Question 9b: ‘Identify the grammatical function of this subordinate clause and state at least one method you used to find the answer. (2 marks)’ 1 / 2 / 2 The marker forgot to mark the question and therefore awarded no marks, even though the student had answered it and the mark awarded by the AI was justified.
Question 10b: ‘Explain the meaning of this word [Editor’s note: “Unbearable”] and then find a synonym for it. (1.5 marks)’ 1 / 1.5 0.5 / 1.5 The AI accepted the word “uncontrollable” as a synonym for “unbearable” to some extent, in a show of generosity, whilst acknowledging that “a more precise synonym such as ‘unsustainable’ or ‘intolerable’ would have been better.”

Conclusion on all questions: although the difference in marks is minimal, it nevertheless states the obvious. Indeed, the AI ignores the student’s spelling mistakes, which sometimes results in a higher mark. Furthermore, human markers are not infallible and may unintentionally overlook certain errors in the dictation, or forget to mark or grade certain questions in the first part of the exam (though this is very rare across the entire corpus), or even be irritated by the form, which results in the substance being undervalued. Finally, AI is limited to the instructions provided by the marking scheme and lacks the subjectivity—or rather the sensitivity—of a human marker, who is willing to value a particular turn of phrase or the choice of one word over another. The real added value of AI for this part is the extreme speed of marking (a matter of seconds), but this overlooks all the preparatory time required to enable this digital marking.

b. The dictation

As for the dictation, the first draft was marked more harshly by the AI than it would have been by a human marker.
Indeed, whilst all errors were identified, some were miscategorised and others penalised twice. This resulted in a significant 3-point discrepancy between the AI and the human marker. For the other papers, I had to amend my instructions to ensure an accurate assessment.

Explanation of the marking scheme for the dictation

c. The editorial team

The AI follows the criteria set out in the national marking scheme (originality; organisation and context; expression), and then provides a detailed explanation and justification for its mark. This is an advantage over human markers, who leave very few comments on the paper and even fewer details about the marking scheme used: only the final mark is shown. On the other hand, many markers make their annotations on a separate draft sheet for each essay, or even fill in a table provided by the standardisers, but I do not have these documents.

General comments provided by the examiner
Feedback on the AI-generated text

Here again, there is a noticeable difference in marking between the AI (33/40) and the human marker (38/40), who was likely impressed by the style. Furthermore, from each set of markers’ papers, I selected a paper from the very end of the set: an unconscious comparison of the papers inevitably takes place, and the marker may thus be inclined to mark a good paper more generously if the preceding ones were disappointing.
Out of the entire corpus of seven papers, only one (paper 6) addresses the reflective topic (which is statistically quite consistent with the reality of a standard envelope): I had to specify to the AI that this was a reflective topic, and it applied the marking scheme set out in the national marking scheme. The justification for the marks awarded is once again well-developed and relevant. The difference in marks between the AI and the teacher is minimal (2 marks).

Summary of all answers

Having set out in detail the comparison between human and machine marking on the first exam paper, here is a summary of the results for all seven exam papers:

If there is a significant discrepancy in marks, particularly for the essay, it is possible to ask the AI to re-mark the paper. That is what I did, as the AI had initially awarded 29 out of 40, whereas the examiner had awarded 20 out of 40. Here is the AI’s response:
‘Why my mark was higher
During my first assessment, I took into account the overall coherence and adherence to the instructions (first-person narrative, spatial and temporal context), which may have led me to be more lenient. A human marker, on the other hand, may place greater emphasis on details such as narrative richness, the quality of transitions, and the depth of emotions and reflections.’

Conclusion
The mark of 20/40 likely reflects a strict assessment of the criteria relating to the originality of the ideas, the clarity of the structure and the precision of the writing. It seems that I initially underestimated the impact of the errors and the simplicity of the narrative. From this perspective, I believe the score of 20/40 is justified, given the need for greater depth and better command of the written word.

CONCLUSION

Current trials of AI-based assessment all start with content that is already in digital form, and it is easy to see why: the time wasted on scanning and double-checking the content after it has been read by HTR significantly slows down the processing speed and thus reduces its appeal. However, handwriting remains the norm, particularly for exams.

Once this hurdle has been overcome, the time saved is clear when marking is carried out by AI. However, the generative instructions must be precise, or even refined or modified during the assessment (for example, in the case of a dictation adapted for Year 7), and, above all, it means entrusting a soulless machine with the task of assessing the originality of a pupil’s style. Furthermore, unlike a teacher who can assess specific knowledge and appropriate vocabulary—such as regarding the First World War in the case of this topic—AI does not measure all the skills the pupil has acquired to produce these pieces of work. Moreover, the AI will remain confined to an official marking scheme, which cannot necessarily express all the implicit nuances or even consider every aspect: it merely identifies deviations from the norm. However, across all seven papers, the difference in total marks is only 0.5. Yet this sometimes masks significant discrepancies within a single paper.

Finally, AI often seems more lenient when it comes to the second part of the exam, namely the essay: this is where the difference compared to human markers is most striking. One also senses a certain nervousness on the part of the AI when discrepancies with a human marker are pointed out. Is this a sign of humility, acknowledging human supremacy? In any case, this experiment tends to show that the teacher’s input is truly indispensable and that the AI marking system still has room for improvement. But for how much longer?

Some references:

 


[1] Authorisation granted by DEC 8 on 2 July 2024

[2] OCR: Optical Character Recognition

[3] Three standard OCR programmes were unable to read a single handwritten line

[4] Gemini is an AI developed by Google. You therefore need to have an account and log in to use it

[5] Hallucination: This is a form of generating new text that mimics a human’s style and tone, but without understanding the meaning of the text produced

[6] HTR (Handwritten Text Recognition) has been developing over the past decade using AI

This case comes from the Lyon Academy website (ici).

All use cases

"Explore our full range of use cases—click below to learn more!"