Design & research note

Instrumenting a teacher-in-the-loop LLM agent for Cambridge exam-item authoring

PaperCraft generates A2 Key and B1 Preliminary practice material, and it is built to measure its own use. Every teacher decision (approve, edit, reject) is retained alongside the untouched model draft, so the system can report where and how much human correction each item type actually requires. Once learners have answered, per-item evidence goes back to the teacher who set the practice.

15

Cambridge item types

KET Parts 1–7 · PET Reading & Writing

3

telemetry instruments

intervention · item analysis · skill aggregation

7

agent pipeline stages

specification → teacher adjudication

A2 · B1

CEFR wordlists in use

2,317 and 4,443 entries, checked separately

Next.js 16 · React 19 · TypeScriptSupabase Postgres + RLSDual-model text routing + vision modelSchema-constrained generationCEFR wordlist validatorPDF / DOCX / PPTX rendering
Figure 1The system in one figure
Constrained generationTeacher adjudicationEvidenceSCHEMACHECKCEFRAUDITrepairfrozen draftapprove · edit · rejectDIFFICULTYDISCRIMINATIONSKILL AREASteacher history and learner evidenceChart shapes are schematic. No measured values are shown, because the study is still collecting data.
Figure 1. Constrained generation produces a candidate; the teacher adjudicates it and the untouched draft is frozen; learner responses and item statistics form the evidence base. The dashed path is what makes the project a study rather than a tool: both the teacher's decisions and the learners' evidence return to the generator. Chart shapes are schematic; no measured values are shown because data collection is still under way.

01 · Problem

Cambridge item constraints that fluent model output breaks

A A2 Reading Part 2 item requires every testable fact to be uniquely locatable in one text, and questions must paraphrase rather than echo distinctive vocabulary. Part 4 requires all options to share a word class so grammar alone cannot eliminate a candidate. A language model can satisfy the surface form while breaking these rules, and the defect is invisible until students have already answered.

Effort lands in the wrong place

Teachers spend their scarcest resource (subject-matter judgement) on formatting, layout and answer-key bookkeeping rather than on pedagogical decisions.

Plausibility is not quality

A model will produce a distractor nobody could choose, or a question whose key is also true of a second person. Both read correctly on the page.

The human contribution goes unmeasured

Tools report what the AI produced. They rarely record what the teacher had to change, so the case for keeping a human in the loop remains anecdotal.

Template-based automatic item generation

Items can be produced at scale with controlled psychometric properties from expert-authored item models.

Item models must be written per construct, which does not fit CEFR-banded language tasks authored on demand.

[4, 9]

Neural and LLM item / exercise generation

Language models produce fluent, plausible exercises, including whole language-learning activities.

Evaluation is largely model-centric (fluency and expert rating of output) rather than workflow-centric.

[11, 14, 15, 18, 29]

Teacher-in-the-loop content and feedback

Keeping educators in the loop improves acceptability and output quality, and teacher preferences can even be optimised against.

The human contribution is usually reported qualitatively; editing behaviour is rarely instrumented as a dependent variable.

[1, 16, 19, 20, 30]

Distractor quality

Distractor plausibility is the hardest part of multiple-choice generation and the main threat to item usefulness.

Little linkage between generation-time constraints and empirical item behaviour once learners answer.

[3, 5]

Prompting by non-experts

Non-experts systematically under-specify prompts and over-generalise from a single successful output.

Suggests moving the specification burden out of free-text prompting into a structured, domain-encoded pipeline.

[21, 26, 27]

CEFR control of generated text

Requested CEFR levels are not reliably achieved by current models.

Motivates an explicit lexical audit as a visible guardrail rather than trusting the instruction.

[17]

03 · Research questions

Four research questions, each with a stated instrument

The questions were fixed before data collection, together with the instrument and the analysis for each. RQ1 and RQ2 are answered by product telemetry; RQ3 by the teacher study, which is qualitative and fully unmoderated; RQ4 is out of scope for this round, because isolating the rules' effect needs a randomised arm an unmoderated study cannot run.

RQ1

How does the amount of teacher correction vary across Cambridge item types, and which types consume the most human editing?

Measured by

Word-level edit distance and typed change flags between the frozen AI draft and the approved version, aggregated per exam part.

RQ2

What per-item evidence does the system return to teachers once learners have answered, and what will that evidence support?

Measured by

Classical item analysis on collected responses: difficulty (p), point-biserial discrimination, option choice frequencies, and requested against observed difficulty. Students answer the paper the teacher approved, so these statistics describe that paper and are reported as a teacher-facing feature, not as evidence about the generator.

RQ3

How do practising teachers work with the drafts, and where do they judge that the tool helps or gets in the way?

Measured by

The written answers each teacher returns in the self-guided form are the primary source, read against the form's own timings and against what the system logged. Task timing against each teacher's own manual baseline, SUS, NASA-TLX and a content-quality rubric are collected in the same form as supporting measures.

RQ4

Exploratory: does retrieving the rules a teacher has accepted into the prompt reduce subsequent editing for that teacher?

Measured by

Out of scope for this round. Isolating the rules' effect means withholding a teacher's accepted rules from some of their generations at random, and that needs a facilitator: consent must be verified before the withholding starts, and a teacher meeting deliberately weaker drafts alone would reasonably conclude the tool is unreliable and rate it accordingly, so the arm is not used. Comparing a teacher's own drafts before and after they accept a rule needs nobody and settles nothing, because familiarity grows in the same direction across the same session. Answering this needs a moderated session or a design running across weeks.

04 · Measurement design

Three measures, each reading data the system already records

Approval, editing and student answering are captured at item level. The system can therefore report how much correction each Cambridge part costs, how the items on an approved paper behaved once a cohort answered them, and where that cohort is currently least accurate. Each instrument is stated with its limits.

Figure 2From instrumented data to research output
RECORDED DATAMEASUREMENT INSTRUMENTSSUPPORTED OUTPUTAI draft (frozen)Teacher-approved versionLearner responsesItem metadata & tagsTeacher Intervention Metric· word-level edit distance· typed change flags· aggregated per exam partwhere drafts need most editingClassical Test Theory analysis· p-value difficulty· point-biserial discrimination· distractor & dead-option audititem quality evidenceSkill-level accuracy aggregation· item type → skill area· grammar tag → skill point· lowest accuracy → next taskdescriptive, not predictive
Figure 2. The frozen draft, the teacher's approved version, learner responses and item metadata feed the three instruments described below. None requires extra work from the teacher; the measurement is a by-product of normal authoring and practice.

Instrument built for this project

Teacher Intervention Metric

The frozen AI draft and the approved version are canonicalised into comparable strings and compared with a word-level edit distance, alongside typed change flags for the passage, question stems, option sets and answer keys. Results aggregate per exam part.

Reads out: Which item types are drafted acceptably, and which consistently require human correction.

Limit: The metric has not yet been validated against human coding of edit severity; it measures amount of change, not its pedagogical importance.

Classical Test Theory

Item analysis on real responses

For auto-graded items the system computes difficulty as the proportion correct, point-biserial discrimination, the frequency with which each option was chosen, distractors that were never selected, and the AI-requested difficulty band against observed difficulty.

Reads out: How the items on one paper behaved with one cohort: which were too easy, which failed to separate stronger from weaker learners, and which distractors nobody chose.

Limit: Discrimination is only computed when at least three responses exist and scores vary. Students answer the version the teacher approved, so the figures describe generator and teacher together and are not read as evidence about the generator (see limitations).

Descriptive aggregation

Skill-level accuracy & next-task surfacing

Item types map onto skill areas (scanning and matching, detailed comprehension, vocabulary and collocation, grammatical function words); grammar focus tags and open-cloze word classes map onto finer skill points. The lowest-accuracy area produces a deep link that pre-fills the authoring form.

Reads out: Where a cohort is currently least accurate, and a one-click route to generating targeted practice.

Limit: This is a descriptive threshold over observed accuracy, not a mastery model: no latent-trait estimation or prediction is performed.

Analytics view showing cohort skill accuracy bars, a recommended next practice card and authoring insights
Skill-level accuracy, the next-task link it surfaces, and the authoring-insight panel reporting how much of each draft was rewritten. Coverage is stated explicitly, because older items have no frozen draft to compare against.
Per-question analysis for a shared exercise, including accuracy per item and response detail
Per-item evidence for one shared exercise. Difficulty, discrimination and option-level choices are computed only from responses actually collected.

05 · Agent system design

A staged pipeline with typed contracts and a human gate

The generation path is not a single call to a large model. It is a staged pipeline in which each stage can be validated, repaired and logged independently, and in which the teacher's decision is an input to the system rather than only an output of it.

Figure 3Authoring agent pipeline
teacher decisions become contextS1Task specificationExam · part · topic · grammar focus · difficulty bandTyped request, validated beforeany model call is madeS2Personalised prompt assemblyPreference summary + few-shot from approved items+ reject memories + rules the teacher has acceptedReads this teacher's decision history,so the agent adapts per user with no fine-tuningS3Construct-specialised instructionOne system prompt per Cambridge part, encoding handbook validity rulesAnti-word-spotting · reuse-and-twistdistractors · single word class per option setS4Generation with typed repair loopJSON-only decode → schema check → handbook check→ error fed back → retry (≤3, model escalates)Shape and specification are separate gates:a wrong question count parses but is not the partS5Multimodal branchScene decomposition → panels drawn in series→ vision model inspects each → vision model writes the taskEach panel conditioned on the previous one;a rejected panel is redrawn on another providerS6CEFR compliance auditContent-word tokens matched against A2 (2,317) / B1 (4,443) wordlistsCompliance % surfaced to the teacherbefore any decision is takenS7Teacher adjudication gateApprove · edit · reject with reason. Nothing persists without a human decisionThe pristine draft is frozen hereas the measurement baseline
Figure 3. Stages S1–S7. Solid arrows carry item candidates; the dashed path shows personalisation. What travels along it is not the teacher's raw history: their edits are recorded as evidence and distilled offline into rules, and only the rules they have accepted are read during prompt assembly. Figure 6 shows that path in full. S7 is a hard gate: no artefact is persisted for classroom use without an explicit human decision.

The agent is a typed pipeline, not a prompt

Each stage has an explicit contract: a validated request object, a per-part instruction set, a schema describing the exact shape of a legal item, and a compliance report. A stage can fail and be repaired without discarding the run, so behaviour is inspectable rather than anecdotal.

Failure is fed back, not swallowed

When generated JSON violates the schema (wrong number of questions, an option set mixing word classes, a missing answer key) the validator's error text is appended to the next prompt and the model is asked to repair its own output, up to three attempts.

The schema says it parses; a second gate says it is the exam task

A shape contract cannot express a syllabus. A A2 Reading Part 3 carrying three questions instead of five parses cleanly and is still not that part, so schema-only validation let out-of-specification material reach teachers. The task tables from the two handbooks for teachers are transcribed into a separate specification covering question and mark counts, options per question, the letters used, one word per gap in an open cloze, the word minimum for each writing task, and no two people matched to the same text. An exercise that contradicts it is regenerated with the discrepancy named in the retry prompt. Anything still unresolved is shown to the teacher in the words of the handbook rather than hidden. Applied to the existing bank, this gate found six exercises that every earlier check had passed.

A guessable answer key, and why asking the model was not enough

Every other check judges one question on its own, so none can see whether the key varies: that is a property of the set. Across 33 generations of A2 Reading Part 4 produced on one day, 26 had one letter taking at least two thirds of the answers and 13 put every answer on option A; over the 198 questions the split was 147 on A, 41 on B and 10 on C. In those 13 a candidate who never read the passage and always answered A would have scored full marks. Instructing the model to spread the key cut the rate without removing the fault: 12 later generations still gave 5 skewed sets and 2 entirely on A. So the spread is now counted in code at generation time and reported to the teacher, which is a check no wording of a prompt can replace.

Offline does not have to mean tomorrow

Distillation stays off the generation request path, because a prompt assembled from raw events grows without bound and cannot be reproduced. For a while that was conflated with running on a schedule, and the schedule left the gap one day wide: a teacher who had just corrected six exercises of one part was told to come back at 03:00. A teacher can now ask for the run. It is the same function the schedule calls, so the grouping that keeps one teacher's evidence out of another's rules and the carry-forward that protects their earlier verdicts cannot differ between the two routes, and asking makes the system look rather than adopt. What it reports when it finds nothing matters as much: evidence too thin for a pattern and no pattern at all are the same blank list and opposite instructions, so the reply names which happened and how many more exercises of one part a pattern needs.

Model routing, not blind retrying

The first two attempts use the primary text model; the final attempt routes to a different model with different structured-output behaviour, so a systematic formatting failure is not simply repeated. The model used is recorded with every generation.

Prompts are contextualised with the teacher's own history

Prompt assembly retrieves the rules this teacher has accepted for that part, a measurement from their own edits of how often a draft of that part needs fixing and which part of it [34], the corrections they made to earlier drafts, a preference summary mined from their approved items, recent rejection reasons as soft avoidance hints, and up to two approved exercises as exemplars. The accepted rules are the only material that has passed a repetition threshold, an audit and the teacher's approval; everything else is merely recent. The material that does not change between one draft and the next is placed before the task description, so a provider can reuse its cached copy of it rather than recomputing it on every draft [33]. This is in-context retrieval, not fine-tuning; no weight updates are involved.

What the teacher changed is the strongest signal, and it was being discarded

The system freezes the untouched draft beside the approved version in order to measure intervention. For a long time that comparison fed the research metric only: generation saw the teacher's finished exercises but never the corrections that produced them, so it was learning what an acceptable item looks like and not what it had got wrong. The difference is now decomposed into field-level before-and-after pairs, such as a rewritten stem, a replaced distractor or a corrected answer key, and a handful reach the next prompt for that part, answer-key corrections first, because a changed key means the draft was wrong rather than clumsy. Measurement, learning and application are checked as three separate links against live data.

Picture stories are drawn as a series, not one frame at a time

The three panels are produced in order, and every panel after the first is generated with the previous panel supplied back to the model as a reference image, so the same characters, clothing and drawing style carry through the story. Separate text-only calls cannot hold character identity: a prompt asking for the same two children returned different children in each frame, so the panels were three unrelated scenes and no single model answer could agree with all of them. That defect was found in supervisor review. The reference copy is downscaled before it is sent back, because it only has to establish who the characters are, not print resolution.

Every panel is inspected before a teacher sees it

A vision model is asked for four observable facts about each finished panel: whether any letters are legible, how many separately framed sub-scenes it contains, how many people are in it, and whether it is a drawing or a photograph. A panel that fails is redrawn, preferring a different provider, because the same endpoint tends to reproduce the same fault. Three faults were caught this way that no structural check could see: legible handwriting inside the picture, a whole storyboard grid drawn inside one frame, and one provider ignoring the line-art instruction and returning a photograph of real children. Greyscale is applied in code rather than asked for in the prompt, because two of the three providers add colour regardless, and the handbook's picture prompts are line art.

The fallback is a measured degradation, not a promise

Each panel gets more than one attempt inside a budget for the series as a whole, so one hung call cannot consume the request. Only one provider accepts a reference image, which made every panel after the first a single point of failure: with that endpoint returning "model load is too high", the exercise collapsed. A panel that cannot be drawn from the previous one is now drawn from a written description of the cast instead, and the teacher is told which panels were produced that way and that the faces may not match. Enforcing the caller's timeout across all three providers was what made this safe: two of them applied their own, so a fallback measured at 287 seconds against a 300-second limit would have returned a platform timeout rather than an exercise. It now completes in about 140.

The task is written from the pictures and tied to them

Once the panels exist, a vision model writes the writing prompt, the hints and the model answer from what it can see, and is instructed to describe only what is actually visible rather than inventing people, places or events. The model answer's word count is recomputed on the server rather than taken from the model, and a model answer that comes back too short is asked for again.

Layout is deterministic code, not model output

Exam and worksheet rendering is a typed layout pass: row heights measured from real text metrics, table headers repeated across page breaks, vector answer boxes, and image blocks that never split from their captions. Presentation fidelity does not depend on model sampling.

Figure 4What a draft has to pass before a teacher sees it
Raw model outputfree-form completionStructural parsecode-fence stripping · JSON parseSchema contractper-part Zod shape: exact item counts, option sets, answer enumsrejected → repairedTyped repair retryfailure text re-injected · ≤3 attempts · model escalationrejected → repairedCEFR lexical auditcontent words vs A2/B1 wordlist · ≤10% out-of-listTeacher adjudicationapprove · edit · reject with reasonevery surviving item carries its own validation record
Figure 4. Each candidate passes structural parsing, a per-part schema contract, a check against the specification transcribed from the handbooks, a bounded repair loop, and a lexical CEFR audit before adjudication. Rejections inside the loop are recoverable: the validator's message becomes prompt context for the retry.

The audit tolerates up to 10% out-of-list content words. The allowance is deliberate rather than arbitrary: proper nouns, topic-specific vocabulary and productive morphology legitimately fall outside a base wordlist, and the figure is reported to the teacher rather than used to silently block an item.

PaperCraft authoring workspace showing exam, exercise type, topic, difficulty and grammar-focus controls
The specification surface (S1). Exam, Cambridge part, topic, difficulty band and grammar focus are explicit parameters, so a request is reproducible and loggable.
Exercise review drawer with the generated item, options, answer key and sharing controls
The adjudication surface (S7). The teacher reads the draft, edits any part of it, and only then approves, at which point the untouched draft is frozen for later comparison.

Rendering layer

Exported papers are produced by a deterministic layout pass rather than by asking a model for markup. Row heights are measured from real text metrics, matching tables repeat their headers after a page break, multiple-choice cloze items use horizontal option rows in the Cambridge convention, and the one picture-based part lays its three panels out three-across. Printed examples are in the next section.

06 · Item types in print

Item types in print, including the picture-based task

Picture-story writing is the hardest part to automate, because the prompt, the three pictures and the sample answer must all describe the same thing. It is also the only illustrated task in either handbook: every other KET part and all of PET Reading and Writing are text only, and an earlier version of this system got that wrong by drawing portraits for B1 Reading Part 2. These pages are real exports. The pictures were synthesised by the pipeline and the task was written from those pictures by a vision model.

Figure 5Item-type coverage against skill areas
Signs & shortmessagesScanning &matchingDetailedcomprehensionTextcohesionVocabulary &collocationGrammar —function wordsExtendedwritingA2 KeyP1 NoticesP2 MatchingP3 Reading MCP4 Cloze MCP5 Open clozeP6 EmailP7 Picture storyB1 PreliminaryR1 Short textsR2 MatchingR3 Reading MCR4 Gapped textR5 Cloze MCR6 Open clozeW1 EmailW2 Article / Storyprimary skill area · auto-graded, item statistics availableprimary skill area · not auto-graded, teacher markedimage-based itemMapping transcribed from the skill-diagnosis and grading code, so the figure cannot drift from the implementation.
Figure 5. Every Cambridge part the system authors, and the skill area each one is mapped to. Solid marks are auto-graded, so item statistics exist for them; dashed marks are the four writing tasks a teacher marks, which is why no discrimination figure can be reported for those parts. The one picture-based part is flagged. The mapping is transcribed from the skill-diagnosis and grading code rather than redrawn by hand, so the figure cannot drift from the implementation.

4 printed examples. Drag, scroll sideways, or use the arrows.

Generated A2 Writing Part 7 paper page with three sequential line-drawing pictures and writing hints

A2 Writing Part 7 · picture story

The agent decomposes the topic into an ordered three-scene sequence, draws the panels in order with each one conditioned on the last so the children stay the same, then a vision model writes the task from the images it produced. The page renders the sequence three-across in the Cambridge convention.

Generated B1 Reading Part 2 paper page with five described people, eight option texts and an answer box for each person

B1 Reading Part 2 · text matching

An earlier version of the generator drew a portrait for each of the five people here. The Cambridge handbook describes them in words and shows no faces, so the pictures made the page unlike the paper it rehearses; they were removed and the layout now pairs each description with its answer box alone. Only A2 Writing Part 7 is illustrated.

Generated A2 Reading Part 2 paper page with three source texts and an A/B/C answer grid

A2 Reading Part 2 · text matching

Three source texts followed by a measured answer grid whose header repeats after a page break.

Generated B1 Reading Part 5 paper page with horizontal A/B/C/D option rows

B1 Reading Part 5 · multiple-choice cloze

One row per gap with four horizontal options and vector answer boxes, matching how the printed paper is laid out.

The pipeline covers KET Parts 1–7 and PET Reading and Writing. Exactly one of those parts is illustrated, KET Writing Part 7, and it is the reason the multimodal branch exists. Generating the text separately from the pictures produces hints that do not match what is shown, and generating the pictures separately from each other produces three unrelated scenes instead of one story. B1 Reading Part 2 was illustrated here for a while, which was a misreading of the handbook: it matches five written descriptions to eight texts and shows no pictures at all.

07 · Classroom delivery

How an approved exercise reaches a class, in print and on screen

Print and PPTX assume the artefact leaves the system and is finished elsewhere. That covers preparation and handouts but not the twenty minutes when the item is on the wall and the teacher is working through it with thirty students. The HTML deck is a third path: one self-contained file, projectable, with answer reveal under teacher control, shareable as a link when the class needs it afterwards.

Slides panel in the paper workspace with theme selector, an include-answers checkbox, present and download actions, and a share link
The delivery surface. One decision (are the answers in this deck?) and three destinations: present immediately without publishing anything, download a single self-contained file, or publish a link. Presenting produces no share record, so a draft paper can be rehearsed without becoming reachable.
A projected multiple-choice cloze slide showing three numbered gaps, each with three options, the correct option marked in green and the key printed beneath it
A question set on the stage, keys revealed. Gaps carry their number from the reading text, options are large enough to read from the back of a room, and the key appears on a keypress rather than being on screen from the start. This is the sequence a teacher needs when working through answers with a class.

The deck is a data structure, not a document

A builder maps the fifteen Cambridge parts onto nine slide layouts (cover, part opener, reading text, question set, card grid, picture grid, writing prompt, answer key, model answer). Rendering is a separate pass over that model. A part that the builder does not recognise produces no slide rather than an unstyled one, so coverage gaps are visible instead of silent.

Answer visibility is decided before rendering, not hidden afterwards

The builder takes a flag. With answers withheld, keys and model answers never enter the slide model, so they are absent from the file rather than concealed in it. For the paper shown here the teacher deck is 32 slides and the student deck is 29, with zero answer-bearing option elements in the markup.

Deterministic layout, deliberately not model-generated

The deck could have been written by a model per paper. It is not, for four reasons: a classroom tool must render the same way every time, a generated file cannot guarantee the answer-exposure property above, generation would add tens of seconds to an action a teacher takes between lessons, and Cambridge part conventions are rules rather than aesthetic choices. Visual variation is a token swap across three themes, so layout and behaviour stay identical.

One file, no runtime dependencies

The exported deck is a single HTML document with inline styles and script, and generated pictures embedded as data URIs. It lays out on a fixed 1280×720 stage that is scaled to the projector by transform, so nothing reflows between the teacher's laptop and the room's screen, and it presents on a school machine with no network.

Interaction serves adjudication, not scoring

Tapping an option marks it and exposes the key; a keypress reveals an answer slide when the class has finished. Nothing is recorded. This is deliberate: measured learner evidence comes from the practice link, where one graded attempt per student keeps the item statistics interpretable, and a projected deck that also logged clicks would contaminate that sample.

Publishing a deck is a separate capability from setting practice

Slide links and practice links are independent records with independent codes. A practice link is served with every answer key stripped; a slide link may legitimately carry them. Sharing one code for both would let anyone holding a student link read the answers by editing the path.

Deck navigation panel listing all 32 slides of the teacher copy, including answer key and model answer entries
Teacher copy, 32 slides. The jump panel doubles as evidence of structure: each Cambridge part contributes an opener, its stimulus, its questions, and where the part is objectively scored, an answer key.
The same deck exported without answers, listing 29 slides with no answer key or model answer entries
Student copy of the same paper, 29 slides. The three answer-key and model-answer slides are not present, and no option element carries a correctness attribute. The difference is a property of the exported file, which is what makes the distinction safe to hand out.
A projected slide showing the three generated line-drawing pictures of a KET picture-story task
The multimodal branch on the stage. The pictures the pipeline synthesised for a A2 Writing Part 7 task are embedded in the deck as data URIs, so the sequence projects identically on a machine with no network: the same images the printed page uses, in the format the room needs.

The builder and the renderer are pure functions, which is what makes this testable rather than eyeballed: 153 automated checks drive real decks in a browser and assert that all fifteen parts produce slides, that every slide stays inside the 16:9 stage, that reveal and navigation work, that teacher-authored text carrying markup is escaped, and that an answer-free deck contains no answers. One production defect was found this way and is worth recording: shared decks were initially served with a sixty-second edge cache, which meant a teacher who republished a link without answers kept serving the answer-bearing copy for several minutes. Shared decks are now uncached; the correctness of who can see the key is worth three database queries per view.

08 · Feedback loops

Two feedback loops: one instructional, one generative

Most authoring tools stop at export. Here the evidence returns twice: learner performance informs what should be taught and generated next, while a teacher's corrections are distilled into rules that they approve before any later draft is allowed to follow them.

Figure 6The teacher-in-the-loop cycle
Two bands. The upper band, headed While the teacher waits: accepted rules (all options for a gap share a word class) and the exam and part (B1 Reading Part 5) feed Generate, which drafts four options: quickly, silence, brightly, softly. Plain code with no model then checks the draft against the schema and the handbook, with a dashed arrow back to Generate labelled repair, names the fault. The teacher reads the draft and changes three options to stillness, brightness and calm. An arrow labelled what changed leads down to the lower band, headed Off the generation path, which runs right to left: the edit log records quickly to stillness, adverb to noun; Distil proposes that all options for a gap share a word class; an audit labelled other vendor passes it; and the teacher decides to use, refuse or retire it. An arrow labelled accept or retire returns to the accepted rules.
Figure 6. The same diagram as the final report, from one source. The top row is the online path, which authors and records but never rewrites the system. The bottom row is the offline path, which only proposes: a pattern must recur across several exercises, survive an audit run on a different vendor's model, and be accepted by the teacher before any later generation reads it. It runs on a daily schedule and whenever a teacher asks for it, through one shared function. The audit can only refuse or pass, never write or apply, and it fails closed, so a silent or truncated audit refuses. A proposal it refuses is never offered to the teacher for approval. Raw edits are never placed in a prompt: they are evidence, and a teacher's typed note is treated as untrusted input rather than as instructions.

Where the teacher sees it, and can say no

A rule is a claim about someone’s teaching, so the page states where it came from and what would change, and offers refusal as plainly as acceptance. The counters above the list are the same figures the loop produced, and they add up: every proposal is either waiting, in force, turned down by the review, turned down by the teacher, or withdrawn after use.

The What it learned page. A row of counters reads 37 corrections recorded, 10 became proposals, 6 refused by the audit, 2 waiting for you, 1 in force now, 1 you stopped using. Below it two proposed rules sit side by side, each with the rule, the reason it was proposed, the exam part and how many exercises it came from, and buttons reading Start using this and No thanks. A further section lists the rule currently shaping the teacher's exercises.
What it learned, on the demonstration account. Nothing here was written by hand: the corrections came from real authoring sessions, the proposals from the scheduled run and from the button beside the heading, and the six refusals from the audit. Two proposals are waiting because the audit did not refute them and the teacher has not answered yet, and neither is shaping a draft while it waits.

09 · Claims & rationale

Five claims, how each is implemented, and what each does not cover

Each row states a claim, the behaviour in the system that follows from it, and what that behaviour does not establish.

Claim

A teacher's edit is evidence, not a failure

Mechanism

Nothing reaches a learner without an explicit teacher decision. Approve, edit and reject are first-class actions, and the rejection reason is stored as data rather than discarded.

Consequence & boundary

Editing is treated as signal, not failure; it is the primary dependent variable of this project.

Claim

Conformance is checked in code, not reviewed afterwards

Mechanism

Each Cambridge part has its own instruction set encoding published item-writing rules: paraphrase instead of repeating distinctive vocabulary, distractors built by reusing and twisting text content, option sets restricted to a single word class.

Consequence & boundary

Content-validity constraints are enforced structurally rather than left to post-hoc review. Construct-validity evidence would require a much larger response pool.

Claim

Per-question evidence has to exist before the exam

Mechanism

Approved items ship as auto-graded practice links; responses are retained at item level, so per-question evidence exists before the exam rather than after it.

Consequence & boundary

Teachers can adjust instruction while it still matters. The system does not yet give individualised feedback to learners.

Claim

Requested and observed difficulty will diverge

Mechanism

Difficulty is requested as a band, audited against the relevant CEFR wordlist, and later compared with observed difficulty from real responses.

Consequence & boundary

Intended and empirical difficulty can diverge, and the divergence is reported. Difficulty remains a teacher decision; there is no automated learner-state adaptation.

Claim

Without the untouched draft, the teacher's work cannot be measured

Mechanism

The pristine model draft is frozen at approval time, so the teacher's final version can always be compared with what the model actually proposed.

Consequence & boundary

The human contribution becomes quantifiable per exam part instead of anecdotal.

10 · Process, ethics & verification

Process, ethics, and verification

Figure 7Evidence provenance and the boundary of each claim
RECORDED DATADERIVED MEASURESUPPORTED CLAIMFrozen AI draftTeacher-approved versionLearner responsesItem metadata & tagsEdit distance & change flagsWhere human correction concentratesDifficulty & discriminationHow the teacher's paper behavedSkill-area accuracyWhere a cohort is least accurateOut of scope: learning gainswould need a controlled classroom studyEach claim is limited to the data on its own row; no row licenses a claim about attainment.
Figure 7. Read in the direction of the arrows: a claim is only as strong as the row of data behind it. Edit-based claims come from the draft/approved pair, item-behaviour claims from learner responses, and cohort-accuracy claims from responses plus item tags. Attainment is deliberately outside the diagram; nothing recorded here licenses a claim about learning gains.

Author's role

Sole designer and developer of the system, and sole author of the study instruments. Cambridge item specifications were taken from the published handbooks; the pedagogy, architecture, measurement code and evaluation design are my own work, produced as a Final Year Project.

Design iteration, with evidence

The interface was revised after structured self-audits. Destructive and export actions that were only discoverable on hover were made persistent or keyboard-reachable; matching and cloze pages were rebuilt after export fixtures showed squeezed rows and unsupported glyphs; class-scoped practice links were found to be silently discarded and were fixed only after verifying the write rather than trusting it.

Ethics & data handling

Participation is voluntary and consented in writing, with the right to withdraw at any time. Participants are reported as P1–P6, work on pre-created study accounts that are not linked to their identity, and no sensitive personal data is collected. Collected forms and telemetry are used only for academic analysis of this project.

Verification

Behaviour is checked by scripted end-to-end tests against the running application, covering authentication, export contracts, generated-paper layout, and the class-sharing path, so the behaviour described here matches the behaviour that ships.

11 · Evaluation

Evaluation: a qualitative study with practising teachers

The design was revised on supervisor advice in September 2026, and the study is run entirely unmoderated: each teacher works alone and remotely, so nothing depends on the researcher and the participant finding the same hour. It is qualitative and teacher-side. The written answers each teacher returns in the form are the primary source, and what teachers write is read against what they did while authoring and against what the system logged. The target is four to six practising KET/PET teachers; that number is chosen for depth of access and is not being optimised for statistical power. Recruitment materials, consent and the self-guided data-collection form are prepared. No results are reported yet.

Figure 8Evaluation design and analysis plan
SESSION, PER PARTICIPANTRECORDEDANALYSIS1 · Orientation in the formthe steps are stated in order,with the screens they name2 · Authoring taskdone alone and remotely, timed bythe form's own stopwatch3 · Written answersin the same form, returned byemail: the primary sourceWhat teachers writeanswers in the returned formWhat teachers didthe form's own task timingsWhat the system loggeditems edited · amount changedRead against each otherEach participant is described from allthree sources, so no finding rests onwhat the teacher reports alone.The written account beside the logWhat a teacher wrote about a draft or arule is set beside what they did to it.Account contradicts the log?The contradiction is the finding.Target sample: 4–6 practising KET/PET teachers, chosen for depth of access rather than for statistical power.SUS, NASA-TLX and the per-item quality ratings are collected alongside as supporting measures.Study in progress. Results will be reported here.
Figure 8. There is no facilitator in this round, so the orientation a facilitated session would have given is written into the form: the tasks are stated in order with the screens they name, and the briefing on deliberate behaviours sits where the participant meets it. Three sources are recorded per participant and read against each other, and where a teacher's written account and the log disagree, the disagreement is itself a finding.

The written answers carry the argument

Findings come from what teachers write in the form they return, read against what they actually did while authoring. Task timings, SUS, NASA-TLX and the per-item quality ratings are collected in the same form and support that account without standing in for it.

The form gives the orientation a facilitator would have given

Every session is unmoderated, so there is nobody to walk a participant through the interface first. That job moves into the form: the tasks are stated in order with the screens they name, and the list of behaviours that look like faults but are deliberate sits where the participant meets them rather than in a script read aloud. A cold start is the normal case here, not a step to be worked around.

Three sources per participant

What teachers write (their answers in the form), what they did (the form's own stopwatch timings), and what the system logged: which items they edited and by how much. Each participant is described from all three, so no claim rests on self-report alone.

The questions are asked straight after the work

There is no interview. The researcher sends the package and receives the form back, so every question sits in the form directly after the task it asks about, including what each proposed rule said and why the teacher accepted or refused it. Where a written account contradicts the logged behaviour, the contradiction is reported as a finding and is not settled in favour of one source.

Study in progress

These are the numeric reporting slots that sit alongside the findings from the written answers. Each will be replaced by the observed value, its dispersion and the sample it came from, or by a statement that the data did not support a conclusion.

—

Authoring time change

paired against each teacher's own manual baseline

awaiting data

—

SUS usability score

0–100, interpreted against the standard benchmark

awaiting data

—

NASA-TLX load

six-dimension mean cognitive load

awaiting data

—

Items usable without edits

from the logged comparison with the frozen draft

awaiting data

12 · Limitations & next steps

Limitations and next steps

Small, purposive sample

The evaluation recruits four to six practising KET/PET teachers, each working alone and remotely. Every session is unmoderated, which is what makes recruitment feasible without finding a shared hour, and which costs the randomised arm RQ4 would need. It is designed to surface usability and intervention patterns in depth, not to support population-level claims, and the number of participants is not being chosen for statistical power.

Item statistics describe the teacher's paper, not the generator

Students answer the version the teacher approved, which they may have edited. Difficulty and discrimination therefore describe the combined artefact of generator plus teacher, and may partly measure the teacher's own item-writing expertise. Per-item analysis is offered as a teacher-facing feature for the papers teachers actually set. A psychometric claim about the generator would need participant numbers well beyond this project.

Why intervention is the quantitative signal instead

The intervention metric is computed on the frozen AI draft as it was proposed, before the teacher's changes, so it is not exposed to the confound above. It is the one quantitative measure here that describes the generator rather than the pair.

Classical Test Theory, not IRT

Difficulty and discrimination are computed with CTT because response counts are small. Item Response Theory calibration would require a substantially larger response pool.

Lexical, not semantic, compliance

The CEFR audit checks vocabulary membership against the level's wordlist. It cannot verify that syntax, cultural reference or cognitive demand sit at the intended level; that judgement remains with the teacher.

No learning-outcome claim

The system measures authoring effort, item behaviour and cohort accuracy patterns. It does not track longitudinal attainment, which would require a controlled classroom study.

Intervention amount is not intervention importance

A one-word fix to an answer key may matter more than rewriting a paragraph. Weighting edits by pedagogical severity requires human coding that this project has not yet carried out.

Single institutional context

Item specifications, teacher expectations and classroom practice are drawn from one exam family and one teaching context, which bounds how far the findings should be generalised.

Planned next

  • Validate the intervention metric against human coding of edit severity, so amount of change can be weighted by pedagogical importance.
  • Calibrate item difficulty with IRT once the response pool is large enough to support it.
  • Tag every item with a cognitive level and language point to visualise construct coverage across a whole paper.
  • Add an agent self-critique stage that flags suspect distractors before the teacher sees them, and test whether it reduces intervention.

13 · References

References

Assessment framing follows evidence-centred design and the formative-assessment literature. The psychometrics follow standard classical test theory. The interaction stance draws on work on teacher/AI complementarity and on the difficulty non-experts have in specifying model behaviour. Exam specifications come from the published Cambridge handbooks; the full reference list appears in the project report.

  1. [1]Balaji, V., Ogange, B. O., & Mays, T. (2025). Teacher in the Loop AI (TiL-AI): A Strategy for Empowering Educators in Developing Countries through OER Adaptation. Journal of Learning for Development, 12(2), 439–445.
  2. [2]Black, P., & Wiliam, D. (1998). Assessment and Classroom Learning. Assessment in Education, 5(1), 7–74. doi:10.1080/0969595980050102
  3. [3]Awalurahman, H. F., & Budi, I. (2024). Automatic distractor generation in multiple-choice questions: a systematic literature review. PeerJ Computer Science, 10, e2441. doi:10.7717/peerj-cs.2441
  4. [4]Gierl, M. J., Lai, H., & Tanygin, V. (2021). Advanced Methods in Automatic Item Generation. Routledge. doi:10.4324/9781003025634
  5. [5]Haladyna, T. M., & Rodriguez, M. C. (2013). Developing and Validating Test Items. Routledge. doi:10.4324/9780203850381
  6. [6]Hattie, J., & Timperley, H. (2007). The Power of Feedback. Review of Educational Research, 77(1), 81–112. doi:10.3102/003465430298487
  7. [7]Holstein, K., McLaren, B. M., & Aleven, V. (2019). Designing for Complementarity: Teacher and Student Needs for Orchestration Support in AI-Enhanced Classrooms. AIED 2019. doi:10.1007/978-3-030-23204-7_14
  8. [8]Holstein, K., & Aleven, V. (2022). Designing for human–AI complementarity in K-12 education. AI Magazine, 43(2). doi:10.1002/aaai.12058
  9. [9]von Davier, M. (2018). Automated Item Generation with Recurrent Neural Networks. Psychometrika, 83(4), 847–857. doi:10.1007/s11336-018-9608-y
  10. [10]Mislevy, R. J., Steinberg, L. S., & Almond, R. G. (2003). Focus Article: On the Structure of Educational Assessments. Measurement, 1(1), 3–62. doi:10.1207/s15366359mea0101_02
  11. [11]Kurdi, G., Leo, J., Parsia, B., Sattler, U., & Al-Emari, S. (2020). A Systematic Review of Automatic Question Generation for Educational Purposes. IJAIED, 30, 121–204. doi:10.1007/s40593-019-00186-y
  12. [12]Hart, S. G., & Staveland, L. E. (1988). Development of NASA-TLX (Task Load Index). Advances in Psychology, 52, 139–183. doi:10.1016/S0166-4115(08)62386-9
  13. [13]Shute, V. J. (2008). Focus on Formative Feedback. Review of Educational Research, 78(1), 153–189. doi:10.3102/0034654307313795
  14. [14]Huang, X., Jiang, F., & Xiao, J. (2025). Generating Reading Comprehension Exercises with Large Language Models for Educational Applications. arXiv:2511.18860.
  15. [15]Gupta, S., Singh, S., Joshi, A., & Kim, M. (2025). LangLingual: A Personalised, Exercise-oriented English Language Learning Tool Leveraging Large Language Models. arXiv:2510.23011.
  16. [16]Zhao, R., Bobrov, A., Li, J., Aloisi, C., & He, Y. (2025). LearnLens: LLM-Enabled Personalised, Curriculum-Grounded Feedback with Educators in the Loop. arXiv:2507.04295.
  17. [17]Uchida, S. (2025). Generative AI and CEFR levels: evaluating the accuracy of generated text. Vocabulary Learning and Instruction, 14(1), 2078. doi:10.29140/vli.v14n1.2078
  18. [18]Rosá, A., Góngora, S., Filevich, J. P., Sastre, I., Musto, L., Carpenter, B., et al. (2025). A Platform for Generating Educational Activities to Teach English as a Second Language. arXiv:2504.20251.
  19. [19]Burke, C. M. (2025). AI-Assisted Exam Variant Generation: A Human-in-the-Loop Framework for Automatic Item Creation. Education Sciences, 15(8), 1029. doi:10.3390/educsci15081029
  20. [20]Woodrow, J., Koyejo, S., & Piech, C. (2025). Improving Generative AI Student Feedback: Direct Preference Optimization with Teachers in the Loop. EDM 2025. doi:10.5281/zenodo.15870266
  21. [21]Zamfirescu-Pereira, J. D., Wong, R. Y., Hartmann, B., & Yang, Q. (2023). Why Johnny Can't Prompt: How Non-AI Experts Try (and Fail) to Design LLM Prompts. CHI 2023, 1–21. doi:10.1145/3544548.3581388
  22. [22]Bangor, A., Kortum, P. T., & Miller, J. T. (2008). An Empirical Evaluation of the System Usability Scale. IJHCI, 24(6), 574–594. doi:10.1080/10447310802205776
  23. [23]Lewis, J. R. (2018). The System Usability Scale: Past, Present, and Future. IJHCI, 34(7), 577–590. doi:10.1080/10447318.2018.1455307
  24. [24]Cambridge Assessment English (2023). A2 Key and A2 Key for Schools: Handbook for Teachers.
  25. [25]Cambridge Assessment English. B1 Preliminary for Schools: Handbook for Teachers.
  26. [26]Wu, T., Terry, M., & Cai, C. J. (2022). AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model Prompts. CHI 2022. doi:10.1145/3491102.3517582
  27. [27]Amershi, S., Weld, D., Vorvoreanu, M., Fourney, A., et al. (2019). Guidelines for Human-AI Interaction. CHI 2019. doi:10.1145/3290605.3300233
  28. [28]Bonner, E., Lege, R., & Frazier, E. (2023). Large Language Model-Based Artificial Intelligence in the Language Classroom. Teaching English with Technology, 23(1), 23–41. doi:10.56297/bkam1691/wieo1749
  29. [29]Xu, H., Gan, W., Qi, Z., Wu, J., & Yu, P. S. (2024). Large Language Models for Education: A Survey. arXiv:2405.13001.
  30. [30]Natarajan, S., Mathur, S., Sidheekh, S., Stammer, W., & Kersting, K. (2025). Human-in-the-loop or AI-in-the-loop? Automate or Collaborate? AAAI-25, 28594.
  31. [31]UNESCO (2021). AI and Education: Guidance for Policy-Makers.
  32. [32]Pavlova, A., Gerazov, B., & Barreiro, A. (2024). Large Language Models and OpenLogos: An Educational Case Scenario. Open Research Europe, 4, 110. doi:10.12688/openreseurope.17605.1
  33. [33]Li, B. (2026). AI Agents in Depth: Design Principles and Engineering Practice. Open-source edition, chapters 2 and 9. https://github.com/bojieli/ai-agent-book
  34. [34]Ding, W., Tomlin, N., & Durrett, G. (2026). Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents. arXiv:2602.16699.
  35. [35]Shinn, N., Cassano, F., Gopinath, A., Narasimhan, K., & Yao, S. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. Advances in Neural Information Processing Systems, 36. arXiv:2303.11366.
  36. [36]Zhao, A., Huang, D., Xu, Q., Lin, M., Liu, Y.-J., & Huang, G. (2024). ExpeL: LLM Agents Are Experiential Learners. Proceedings of the AAAI Conference on Artificial Intelligence, 38(17). arXiv:2308.10144.
  37. [37]Wang, Z. Z., Mao, J., Fried, D., & Neubig, G. (2024). Agent Workflow Memory. arXiv:2409.07429.

The running system

The authoring workspace, the analytics and the export pipeline described above are all live.