Effort lands in the wrong place
Teachers spend their scarcest resource (subject-matter judgement) on formatting, layout and answer-key bookkeeping rather than on pedagogical decisions.
PaperCraft generates A2 Key and B1 Preliminary practice material, and it is built to measure its own use. Every teacher decision (approve, edit, reject) is retained alongside the untouched model draft, so the system can report where and how much human correction each item type actually requires. Once learners have answered, per-item evidence goes back to the teacher who set the practice.
15
Cambridge item types
KET Parts 1–7 · PET Reading & Writing
3
telemetry instruments
intervention · item analysis · skill aggregation
7
agent pipeline stages
specification → teacher adjudication
A2 · B1
CEFR wordlists in use
2,317 and 4,443 entries, checked separately
01 · Problem
A A2 Reading Part 2 item requires every testable fact to be uniquely locatable in one text, and questions must paraphrase rather than echo distinctive vocabulary. Part 4 requires all options to share a word class so grammar alone cannot eliminate a candidate. A language model can satisfy the surface form while breaking these rules, and the defect is invisible until students have already answered.
Teachers spend their scarcest resource (subject-matter judgement) on formatting, layout and answer-key bookkeeping rather than on pedagogical decisions.
A model will produce a distractor nobody could choose, or a question whose key is also true of a second person. Both read correctly on the page.
Tools report what the AI produced. They rarely record what the teacher had to change, so the case for keeping a human in the loop remains anecdotal.
Strand
What it establishes
What it leaves open
Refs
Template-based automatic item generation
Items can be produced at scale with controlled psychometric properties from expert-authored item models.
Item models must be written per construct, which does not fit CEFR-banded language tasks authored on demand.
[4, 9]
Neural and LLM item / exercise generation
Language models produce fluent, plausible exercises, including whole language-learning activities.
Evaluation is largely model-centric (fluency and expert rating of output) rather than workflow-centric.
[11, 14, 15, 18, 29]
Teacher-in-the-loop content and feedback
Keeping educators in the loop improves acceptability and output quality, and teacher preferences can even be optimised against.
The human contribution is usually reported qualitatively; editing behaviour is rarely instrumented as a dependent variable.
[1, 16, 19, 20, 30]
Distractor quality
Distractor plausibility is the hardest part of multiple-choice generation and the main threat to item usefulness.
Little linkage between generation-time constraints and empirical item behaviour once learners answer.
[3, 5]
Prompting by non-experts
Non-experts systematically under-specify prompts and over-generalise from a single successful output.
Suggests moving the specification burden out of free-text prompting into a structured, domain-encoded pipeline.
[21, 26, 27]
CEFR control of generated text
Requested CEFR levels are not reliably achieved by current models.
Motivates an explicit lexical audit as a visible guardrail rather than trusting the instruction.
[17]
03 · Research questions
The questions were fixed before data collection, together with the instrument and the analysis for each. RQ1 and RQ2 are answered by product telemetry; RQ3 by the teacher study, which is qualitative and fully unmoderated; RQ4 is out of scope for this round, because isolating the rules' effect needs a randomised arm an unmoderated study cannot run.
How does the amount of teacher correction vary across Cambridge item types, and which types consume the most human editing?
Measured by
Word-level edit distance and typed change flags between the frozen AI draft and the approved version, aggregated per exam part.
What per-item evidence does the system return to teachers once learners have answered, and what will that evidence support?
Measured by
Classical item analysis on collected responses: difficulty (p), point-biserial discrimination, option choice frequencies, and requested against observed difficulty. Students answer the paper the teacher approved, so these statistics describe that paper and are reported as a teacher-facing feature, not as evidence about the generator.
How do practising teachers work with the drafts, and where do they judge that the tool helps or gets in the way?
Measured by
The written answers each teacher returns in the self-guided form are the primary source, read against the form's own timings and against what the system logged. Task timing against each teacher's own manual baseline, SUS, NASA-TLX and a content-quality rubric are collected in the same form as supporting measures.
Exploratory: does retrieving the rules a teacher has accepted into the prompt reduce subsequent editing for that teacher?
Measured by
Out of scope for this round. Isolating the rules' effect means withholding a teacher's accepted rules from some of their generations at random, and that needs a facilitator: consent must be verified before the withholding starts, and a teacher meeting deliberately weaker drafts alone would reasonably conclude the tool is unreliable and rate it accordingly, so the arm is not used. Comparing a teacher's own drafts before and after they accept a rule needs nobody and settles nothing, because familiarity grows in the same direction across the same session. Answering this needs a moderated session or a design running across weeks.
04 · Measurement design
Approval, editing and student answering are captured at item level. The system can therefore report how much correction each Cambridge part costs, how the items on an approved paper behaved once a cohort answered them, and where that cohort is currently least accurate. Each instrument is stated with its limits.
Instrument built for this project
The frozen AI draft and the approved version are canonicalised into comparable strings and compared with a word-level edit distance, alongside typed change flags for the passage, question stems, option sets and answer keys. Results aggregate per exam part.
Reads out: Which item types are drafted acceptably, and which consistently require human correction.
Limit: The metric has not yet been validated against human coding of edit severity; it measures amount of change, not its pedagogical importance.
Classical Test Theory
For auto-graded items the system computes difficulty as the proportion correct, point-biserial discrimination, the frequency with which each option was chosen, distractors that were never selected, and the AI-requested difficulty band against observed difficulty.
Reads out: How the items on one paper behaved with one cohort: which were too easy, which failed to separate stronger from weaker learners, and which distractors nobody chose.
Limit: Discrimination is only computed when at least three responses exist and scores vary. Students answer the version the teacher approved, so the figures describe generator and teacher together and are not read as evidence about the generator (see limitations).
Descriptive aggregation
Item types map onto skill areas (scanning and matching, detailed comprehension, vocabulary and collocation, grammatical function words); grammar focus tags and open-cloze word classes map onto finer skill points. The lowest-accuracy area produces a deep link that pre-fills the authoring form.
Reads out: Where a cohort is currently least accurate, and a one-click route to generating targeted practice.
Limit: This is a descriptive threshold over observed accuracy, not a mastery model: no latent-trait estimation or prediction is performed.


05 · Agent system design
The generation path is not a single call to a large model. It is a staged pipeline in which each stage can be validated, repaired and logged independently, and in which the teacher's decision is an input to the system rather than only an output of it.
Each stage has an explicit contract: a validated request object, a per-part instruction set, a schema describing the exact shape of a legal item, and a compliance report. A stage can fail and be repaired without discarding the run, so behaviour is inspectable rather than anecdotal.
When generated JSON violates the schema (wrong number of questions, an option set mixing word classes, a missing answer key) the validator's error text is appended to the next prompt and the model is asked to repair its own output, up to three attempts.
A shape contract cannot express a syllabus. A A2 Reading Part 3 carrying three questions instead of five parses cleanly and is still not that part, so schema-only validation let out-of-specification material reach teachers. The task tables from the two handbooks for teachers are transcribed into a separate specification covering question and mark counts, options per question, the letters used, one word per gap in an open cloze, the word minimum for each writing task, and no two people matched to the same text. An exercise that contradicts it is regenerated with the discrepancy named in the retry prompt. Anything still unresolved is shown to the teacher in the words of the handbook rather than hidden. Applied to the existing bank, this gate found six exercises that every earlier check had passed.
Every other check judges one question on its own, so none can see whether the key varies: that is a property of the set. Across 33 generations of A2 Reading Part 4 produced on one day, 26 had one letter taking at least two thirds of the answers and 13 put every answer on option A; over the 198 questions the split was 147 on A, 41 on B and 10 on C. In those 13 a candidate who never read the passage and always answered A would have scored full marks. Instructing the model to spread the key cut the rate without removing the fault: 12 later generations still gave 5 skewed sets and 2 entirely on A. So the spread is now counted in code at generation time and reported to the teacher, which is a check no wording of a prompt can replace.
Distillation stays off the generation request path, because a prompt assembled from raw events grows without bound and cannot be reproduced. For a while that was conflated with running on a schedule, and the schedule left the gap one day wide: a teacher who had just corrected six exercises of one part was told to come back at 03:00. A teacher can now ask for the run. It is the same function the schedule calls, so the grouping that keeps one teacher's evidence out of another's rules and the carry-forward that protects their earlier verdicts cannot differ between the two routes, and asking makes the system look rather than adopt. What it reports when it finds nothing matters as much: evidence too thin for a pattern and no pattern at all are the same blank list and opposite instructions, so the reply names which happened and how many more exercises of one part a pattern needs.
The first two attempts use the primary text model; the final attempt routes to a different model with different structured-output behaviour, so a systematic formatting failure is not simply repeated. The model used is recorded with every generation.
Prompt assembly retrieves the rules this teacher has accepted for that part, a measurement from their own edits of how often a draft of that part needs fixing and which part of it [34], the corrections they made to earlier drafts, a preference summary mined from their approved items, recent rejection reasons as soft avoidance hints, and up to two approved exercises as exemplars. The accepted rules are the only material that has passed a repetition threshold, an audit and the teacher's approval; everything else is merely recent. The material that does not change between one draft and the next is placed before the task description, so a provider can reuse its cached copy of it rather than recomputing it on every draft [33]. This is in-context retrieval, not fine-tuning; no weight updates are involved.
The system freezes the untouched draft beside the approved version in order to measure intervention. For a long time that comparison fed the research metric only: generation saw the teacher's finished exercises but never the corrections that produced them, so it was learning what an acceptable item looks like and not what it had got wrong. The difference is now decomposed into field-level before-and-after pairs, such as a rewritten stem, a replaced distractor or a corrected answer key, and a handful reach the next prompt for that part, answer-key corrections first, because a changed key means the draft was wrong rather than clumsy. Measurement, learning and application are checked as three separate links against live data.
The three panels are produced in order, and every panel after the first is generated with the previous panel supplied back to the model as a reference image, so the same characters, clothing and drawing style carry through the story. Separate text-only calls cannot hold character identity: a prompt asking for the same two children returned different children in each frame, so the panels were three unrelated scenes and no single model answer could agree with all of them. That defect was found in supervisor review. The reference copy is downscaled before it is sent back, because it only has to establish who the characters are, not print resolution.
A vision model is asked for four observable facts about each finished panel: whether any letters are legible, how many separately framed sub-scenes it contains, how many people are in it, and whether it is a drawing or a photograph. A panel that fails is redrawn, preferring a different provider, because the same endpoint tends to reproduce the same fault. Three faults were caught this way that no structural check could see: legible handwriting inside the picture, a whole storyboard grid drawn inside one frame, and one provider ignoring the line-art instruction and returning a photograph of real children. Greyscale is applied in code rather than asked for in the prompt, because two of the three providers add colour regardless, and the handbook's picture prompts are line art.
Each panel gets more than one attempt inside a budget for the series as a whole, so one hung call cannot consume the request. Only one provider accepts a reference image, which made every panel after the first a single point of failure: with that endpoint returning "model load is too high", the exercise collapsed. A panel that cannot be drawn from the previous one is now drawn from a written description of the cast instead, and the teacher is told which panels were produced that way and that the faces may not match. Enforcing the caller's timeout across all three providers was what made this safe: two of them applied their own, so a fallback measured at 287 seconds against a 300-second limit would have returned a platform timeout rather than an exercise. It now completes in about 140.
Once the panels exist, a vision model writes the writing prompt, the hints and the model answer from what it can see, and is instructed to describe only what is actually visible rather than inventing people, places or events. The model answer's word count is recomputed on the server rather than taken from the model, and a model answer that comes back too short is asked for again.
Exam and worksheet rendering is a typed layout pass: row heights measured from real text metrics, table headers repeated across page breaks, vector answer boxes, and image blocks that never split from their captions. Presentation fidelity does not depend on model sampling.
The audit tolerates up to 10% out-of-list content words. The allowance is deliberate rather than arbitrary: proper nouns, topic-specific vocabulary and productive morphology legitimately fall outside a base wordlist, and the figure is reported to the teacher rather than used to silently block an item.


Rendering layer
Exported papers are produced by a deterministic layout pass rather than by asking a model for markup. Row heights are measured from real text metrics, matching tables repeat their headers after a page break, multiple-choice cloze items use horizontal option rows in the Cambridge convention, and the one picture-based part lays its three panels out three-across. Printed examples are in the next section.
06 · Item types in print
Picture-story writing is the hardest part to automate, because the prompt, the three pictures and the sample answer must all describe the same thing. It is also the only illustrated task in either handbook: every other KET part and all of PET Reading and Writing are text only, and an earlier version of this system got that wrong by drawing portraits for B1 Reading Part 2. These pages are real exports. The pictures were synthesised by the pipeline and the task was written from those pictures by a vision model.
4 printed examples. Drag, scroll sideways, or use the arrows.
The pipeline covers KET Parts 1–7 and PET Reading and Writing. Exactly one of those parts is illustrated, KET Writing Part 7, and it is the reason the multimodal branch exists. Generating the text separately from the pictures produces hints that do not match what is shown, and generating the pictures separately from each other produces three unrelated scenes instead of one story. B1 Reading Part 2 was illustrated here for a while, which was a misreading of the handbook: it matches five written descriptions to eight texts and shows no pictures at all.
07 · Classroom delivery
Print and PPTX assume the artefact leaves the system and is finished elsewhere. That covers preparation and handouts but not the twenty minutes when the item is on the wall and the teacher is working through it with thirty students. The HTML deck is a third path: one self-contained file, projectable, with answer reveal under teacher control, shareable as a link when the class needs it afterwards.


A builder maps the fifteen Cambridge parts onto nine slide layouts (cover, part opener, reading text, question set, card grid, picture grid, writing prompt, answer key, model answer). Rendering is a separate pass over that model. A part that the builder does not recognise produces no slide rather than an unstyled one, so coverage gaps are visible instead of silent.
The builder takes a flag. With answers withheld, keys and model answers never enter the slide model, so they are absent from the file rather than concealed in it. For the paper shown here the teacher deck is 32 slides and the student deck is 29, with zero answer-bearing option elements in the markup.
The deck could have been written by a model per paper. It is not, for four reasons: a classroom tool must render the same way every time, a generated file cannot guarantee the answer-exposure property above, generation would add tens of seconds to an action a teacher takes between lessons, and Cambridge part conventions are rules rather than aesthetic choices. Visual variation is a token swap across three themes, so layout and behaviour stay identical.
The exported deck is a single HTML document with inline styles and script, and generated pictures embedded as data URIs. It lays out on a fixed 1280×720 stage that is scaled to the projector by transform, so nothing reflows between the teacher's laptop and the room's screen, and it presents on a school machine with no network.
Tapping an option marks it and exposes the key; a keypress reveals an answer slide when the class has finished. Nothing is recorded. This is deliberate: measured learner evidence comes from the practice link, where one graded attempt per student keeps the item statistics interpretable, and a projected deck that also logged clicks would contaminate that sample.
Slide links and practice links are independent records with independent codes. A practice link is served with every answer key stripped; a slide link may legitimately carry them. Sharing one code for both would let anyone holding a student link read the answers by editing the path.



The builder and the renderer are pure functions, which is what makes this testable rather than eyeballed: 153 automated checks drive real decks in a browser and assert that all fifteen parts produce slides, that every slide stays inside the 16:9 stage, that reveal and navigation work, that teacher-authored text carrying markup is escaped, and that an answer-free deck contains no answers. One production defect was found this way and is worth recording: shared decks were initially served with a sixty-second edge cache, which meant a teacher who republished a link without answers kept serving the answer-bearing copy for several minutes. Shared decks are now uncached; the correctness of who can see the key is worth three database queries per view.
08 · Feedback loops
Most authoring tools stop at export. Here the evidence returns twice: learner performance informs what should be taught and generated next, while a teacher's corrections are distilled into rules that they approve before any later draft is allowed to follow them.

A rule is a claim about someone’s teaching, so the page states where it came from and what would change, and offers refusal as plainly as acceptance. The counters above the list are the same figures the loop produced, and they add up: every proposal is either waiting, in force, turned down by the review, turned down by the teacher, or withdrawn after use.

09 · Claims & rationale
Each row states a claim, the behaviour in the system that follows from it, and what that behaviour does not establish.
Claim
A teacher's edit is evidence, not a failure
Mechanism
Nothing reaches a learner without an explicit teacher decision. Approve, edit and reject are first-class actions, and the rejection reason is stored as data rather than discarded.
Consequence & boundary
Editing is treated as signal, not failure; it is the primary dependent variable of this project.
Claim
Conformance is checked in code, not reviewed afterwards
Mechanism
Each Cambridge part has its own instruction set encoding published item-writing rules: paraphrase instead of repeating distinctive vocabulary, distractors built by reusing and twisting text content, option sets restricted to a single word class.
Consequence & boundary
Content-validity constraints are enforced structurally rather than left to post-hoc review. Construct-validity evidence would require a much larger response pool.
Claim
Per-question evidence has to exist before the exam
Mechanism
Approved items ship as auto-graded practice links; responses are retained at item level, so per-question evidence exists before the exam rather than after it.
Consequence & boundary
Teachers can adjust instruction while it still matters. The system does not yet give individualised feedback to learners.
Claim
Requested and observed difficulty will diverge
Mechanism
Difficulty is requested as a band, audited against the relevant CEFR wordlist, and later compared with observed difficulty from real responses.
Consequence & boundary
Intended and empirical difficulty can diverge, and the divergence is reported. Difficulty remains a teacher decision; there is no automated learner-state adaptation.
Claim
Without the untouched draft, the teacher's work cannot be measured
Mechanism
The pristine model draft is frozen at approval time, so the teacher's final version can always be compared with what the model actually proposed.
Consequence & boundary
The human contribution becomes quantifiable per exam part instead of anecdotal.
10 · Process, ethics & verification
Sole designer and developer of the system, and sole author of the study instruments. Cambridge item specifications were taken from the published handbooks; the pedagogy, architecture, measurement code and evaluation design are my own work, produced as a Final Year Project.
The interface was revised after structured self-audits. Destructive and export actions that were only discoverable on hover were made persistent or keyboard-reachable; matching and cloze pages were rebuilt after export fixtures showed squeezed rows and unsupported glyphs; class-scoped practice links were found to be silently discarded and were fixed only after verifying the write rather than trusting it.
Participation is voluntary and consented in writing, with the right to withdraw at any time. Participants are reported as P1–P6, work on pre-created study accounts that are not linked to their identity, and no sensitive personal data is collected. Collected forms and telemetry are used only for academic analysis of this project.
Behaviour is checked by scripted end-to-end tests against the running application, covering authentication, export contracts, generated-paper layout, and the class-sharing path, so the behaviour described here matches the behaviour that ships.
11 · Evaluation
The design was revised on supervisor advice in September 2026, and the study is run entirely unmoderated: each teacher works alone and remotely, so nothing depends on the researcher and the participant finding the same hour. It is qualitative and teacher-side. The written answers each teacher returns in the form are the primary source, and what teachers write is read against what they did while authoring and against what the system logged. The target is four to six practising KET/PET teachers; that number is chosen for depth of access and is not being optimised for statistical power. Recruitment materials, consent and the self-guided data-collection form are prepared. No results are reported yet.
Findings come from what teachers write in the form they return, read against what they actually did while authoring. Task timings, SUS, NASA-TLX and the per-item quality ratings are collected in the same form and support that account without standing in for it.
Every session is unmoderated, so there is nobody to walk a participant through the interface first. That job moves into the form: the tasks are stated in order with the screens they name, and the list of behaviours that look like faults but are deliberate sits where the participant meets them rather than in a script read aloud. A cold start is the normal case here, not a step to be worked around.
What teachers write (their answers in the form), what they did (the form's own stopwatch timings), and what the system logged: which items they edited and by how much. Each participant is described from all three, so no claim rests on self-report alone.
There is no interview. The researcher sends the package and receives the form back, so every question sits in the form directly after the task it asks about, including what each proposed rule said and why the teacher accepted or refused it. Where a written account contradicts the logged behaviour, the contradiction is reported as a finding and is not settled in favour of one source.
These are the numeric reporting slots that sit alongside the findings from the written answers. Each will be replaced by the observed value, its dispersion and the sample it came from, or by a statement that the data did not support a conclusion.
—
Authoring time change
paired against each teacher's own manual baseline
awaiting data
—
SUS usability score
0–100, interpreted against the standard benchmark
awaiting data
—
NASA-TLX load
six-dimension mean cognitive load
awaiting data
—
Items usable without edits
from the logged comparison with the frozen draft
awaiting data
12 · Limitations & next steps
The evaluation recruits four to six practising KET/PET teachers, each working alone and remotely. Every session is unmoderated, which is what makes recruitment feasible without finding a shared hour, and which costs the randomised arm RQ4 would need. It is designed to surface usability and intervention patterns in depth, not to support population-level claims, and the number of participants is not being chosen for statistical power.
Students answer the version the teacher approved, which they may have edited. Difficulty and discrimination therefore describe the combined artefact of generator plus teacher, and may partly measure the teacher's own item-writing expertise. Per-item analysis is offered as a teacher-facing feature for the papers teachers actually set. A psychometric claim about the generator would need participant numbers well beyond this project.
The intervention metric is computed on the frozen AI draft as it was proposed, before the teacher's changes, so it is not exposed to the confound above. It is the one quantitative measure here that describes the generator rather than the pair.
Difficulty and discrimination are computed with CTT because response counts are small. Item Response Theory calibration would require a substantially larger response pool.
The CEFR audit checks vocabulary membership against the level's wordlist. It cannot verify that syntax, cultural reference or cognitive demand sit at the intended level; that judgement remains with the teacher.
The system measures authoring effort, item behaviour and cohort accuracy patterns. It does not track longitudinal attainment, which would require a controlled classroom study.
A one-word fix to an answer key may matter more than rewriting a paragraph. Weighting edits by pedagogical severity requires human coding that this project has not yet carried out.
Item specifications, teacher expectations and classroom practice are drawn from one exam family and one teaching context, which bounds how far the findings should be generalised.
Planned next
13 · References
Assessment framing follows evidence-centred design and the formative-assessment literature. The psychometrics follow standard classical test theory. The interaction stance draws on work on teacher/AI complementarity and on the difficulty non-experts have in specifying model behaviour. Exam specifications come from the published Cambridge handbooks; the full reference list appears in the project report.
The authoring workspace, the analytics and the export pipeline described above are all live.