Explainer
Automated marking and AI grading are not the same thing
Every marking tool's pitch is the same shape — papers in, marks out. Underneath that sentence sit two quite different products. One forms a judgement about a piece of writing; the other reads what a student wrote and compares it to an answer you already decided on. They fail in different ways, and they need different questions asked of them.
The first kind: a model forms a judgement
The student has written something open-ended — an essay, an extended response, a paragraph of argument — and there's no single right answer to compare it against. A language model reads the writing against a rubric and proposes a grade, usually with a paragraph of feedback explaining it.
This is judgement work. The model isn't recognising a fact; it's forming an opinion about quality — is the argument sustained, is the evidence used well, does the structure hold. That's the call a human marker makes, and the call two human markers disagree about.
So the value is usually the feedback: a student who gets specific comments on a draft on Tuesday instead of a mark on Friday is better off. The risk is the judgement itself — an opinion from a system that can't show its working in a way you could audit, and one you'd still have to defend if a parent rang about the grade.
The second kind: a definite answer already exists
Now the other situation, which is most of the marking a classroom teacher actually does. Twenty-eight papers, forty questions, and you already know every answer. Q7 is 0.75. Q12 is B. The cell says "endothermic" or it doesn't. Nothing to judge.
The work here isn't judgement, it's reading. Someone has to look at each answer box, work out what the handwriting says, and compare it to the key. Marking a class set by hand is mostly that: hundreds of tiny acts of reading, almost none of them interesting, every one possible to get wrong at eleven at night.
So the failure mode is different. A marker of this kind doesn't form a bad opinion — it misreads: a 7 where a student wrote a 1, "there" for "their", a blank cell missed. Because the answer was definite, a misreading is simply a wrong mark on a real student's record.
Which makes the fix more tractable: check the reading against the page, and keep a person on the decision. Verify every reading against the actual scan, hand anything that doesn't verify to the teacher with the handwriting beside it, and the error class that matters mostly disappears. The teacher isn't re-marking — they're confirming a reading, which is exactly what a human eye is good at.
Why "the AI marks it" hides the important question
That phrase covers both of the above, which is the problem. It answers "what does it do" and skips the two things you need to know.
Who decides? There's a real difference between a tool that proposes a mark and one that records it. If marks land in a record without a person agreeing to them, the teacher has moved from marker to spectator — and is still the one who has to explain the result.
Can you see why? A mark you can't trace is a mark you can't defend. For reading work, "why" has a concrete form: show me the piece of the scan this came from. For judgement work it's harder, and the honest tools say so. Be sceptical of any tool that hands you a tidy certainty number about its own output — a number like that is easy to print and hard to earn.
Where AutoMark sits, plainly
AutoMark is the second kind, and only the second kind.
You build the test and the answer key in the app, print the papers from it, and scan the completed class set back in. AutoMark reads the handwriting, marks each answer against your key, and checks every reading back against the scan. That check either passes or it doesn't — no dial, no score to interpret. Readings that pass mark straight through; anything that doesn't is flagged with the cropped handwriting beside it, and you confirm or override in a keystroke. Then you commit. You confirm every mark before it's recorded, and no mark is ever produced by a solely automated decision.
It marks structured answers — multiple choice, true or false, numbers, short written answers, fill-the-blanks, tables cell by cell, matching, ordering, selections from a passage, checklists, diagram labels. More on the reading itself is in marking handwritten tests.
It does not mark essays, and we'd rather say so than sell you something it isn't. A tool built to verify readings against a page has nothing useful to say about whether a thesis is well supported. Set the long-answer question aside and mark it by hand — there's a setting for that, so the rest of the paper still marks. (If you're searching: an unrelated essay-marking product shares our name, and we've written up the difference.)
Where the essay tools sit
If what you want is feedback on drafts, look at the first kind. We won't name products, but judge them on this: the good ones are honest that they produce a proposal for a human marker, let you edit the rubric, and don't pretend a judgement is a measurement. The two kinds aren't competitors — one could take the extended response, the other the remaining forty questions.
Who the marker is
Worth holding onto: in Australian schooling, assessment policy generally treats a person — the teacher, or the examiner — as the authoritative marker, with the judgement and the accountability sitting with them. Specifics vary by state, sector and kind of assessment, so check your own school or authority's policy rather than anything we say here. But the direction is consistent enough to plan around: a tool's output is an input to your decision, not a substitute for it.
That's a large part of why AutoMark is built this way: a correction after commit is recorded against the result, never a silent overwrite.
Questions to ask any marking tool
- Which kind is this? Judgement about open-ended writing, or reading a definite answer? If the marketing won't say, that's informative.
- Who records the mark? Can a result reach a student's record without a person agreeing to it? If it can, you're accountable for a decision you didn't make.
- Can I see the evidence? For reading work that means the actual scan of what the student wrote, beside what the tool read — not a summary of it.
- What does it do when it isn't sure? Flag it, or guess and move on? A tool that never asks you anything isn't more accurate — it's less honest.
- What can I change afterwards? Can you correct a mark, or move a paper to the right student — and is the change recorded, or does it quietly overwrite?
- Where does the student work go? Which country, who else processes it, how long is it kept, does anyone train a model on it? Ours is on student data and privacy and in the privacy policy.
- What does it refuse to do? Every honest tool has a list. Ask for it.
If you keep one thing: automated marking should reduce how much reading you do, not how much deciding you do. The moment a tool makes the decision instead of the reading, you've bought something other than what you thought.
More on the review step is in the FAQ; how to scan a class set for marking covers the part that most affects how much needs your eyes.
Mark your next class set in minutes.
The beta is full for now — leave your address and we'll tell you when a place opens. Free for two months when it does.
Join the waiting list