Part three of the Mika Labs research series. Part one looked at why visual, interactive tutoring works. Part two was about diagnosing why a student got something wrong. This one is about the hardest problem in university teaching right now: proving a student learned something, when a machine that can do the work is one tab away.
In 2024, the University of Sydney finished auditing its own assessments against the tools its students now carry. Roughly 90% no longer held up. Its Pro Vice-Chancellor for Educational Innovation, Adam Bridgeman, named the deeper problem in The Chronicle of Higher Education: the university said it taught critical thinking, and he didn't think it had actually been assessing it.
Sydney sits in the world's top twenty. It had not been careless. It had been assessing the way nearly every university assesses, by collecting the final answer, and the final answer had quietly stopped being evidence of anything.
Every institution now faces the same arithmetic, and none of the variables are negotiable. You cannot remove the tool; it is on every phone, in every browser, and increasingly inside the operating system itself. You cannot reliably detect it. And you cannot keep certifying graduates on evidence that a free product manufactures in seconds.
Which leaves exactly one variable you control. The assessment.
Here is the distinction we think the sector keeps missing:
The goal is not AI-resistant assessment. It is AI-resilient assessment.
An AI-resistant task tries to keep the tool out. Lockdown browsers, proctoring, bans, detectors. Every one of those strategies depends on verifying an absence, and an absence cannot be verified. When the defence fails, and it does fail, the assessment fails with it, because it was only ever valid on the assumption that no help arrived.
An AI-resilient task makes no such assumption. It stays valid whether or not the student used AI, because the evidence it collects is not the kind an answer generator produces. Nothing has to be kept out. Nothing has to be proven absent. The task survives contact with the tool.
Resistance is a wall, and a wall is only as good as the last person who climbed it. Resilience is a property of the design.
We did not arrive at that phrasing for this article. Internally, this work has been called the Mika AI-Resilient Assessment Framework, or M-ARAF, since long before it had anything to ship. The framework is the principle. The Evidence of Learning Engine is its implementation, and the name instructors see when they open Mika. We are putting the internal term on the record because the distinction it encodes is the load-bearing one, and we would rather be held to a principle than to a feature.
In mathematics and the sciences we have a sharper version of this problem, and, oddly, a head start on it. A worked derivation, a free-body diagram, a titration calculation, a page of algebra where a student factored x² + 2x − 3 as (x−3)(x+1) instead of (x−1)(x+3) and carried the error cleanly through to a wrong answer by an otherwise sound method: none of that carries the signal detectors were built to read. No sentence rhythm, no token distribution, no stylometry. Our disciplines never had detection to lose, which means we were never able to pretend resistance was working. We have to build for resilience because nothing else was ever available to us.
And our fallback, the mark, does not rescue us either. A 63% cannot separate the student who chose the right method and mis-factored a quadratic from one who never understood the technique at all, any more than it separates a physics answer lost to a single sign error from one where the student could not identify which forces were acting. We made that case in part two of this series: a grade is a scalar, and teaching needs a map.
So STEM assessment is left with one question that cannot be answered and one number that is not worth much, while the thing an instructor actually needs to know goes unmeasured: can this student do this, and what happened when they couldn't?
That is what the Evidence of Learning Engine was built to measure. Every design decision in it is traceable to a published finding, and this piece is where we show our work.
Part 1: Why resistance fails even where it can run
The obvious reply is that plenty of STEM assessment does produce text. Lab reports, methods sections, project write-ups, discussion of results, code and its comments all give a detector something to grip. So it's worth establishing that where detection can run, it still doesn't work.
It would be convenient if detectors were simply immature, and another release cycle would fix the false positives and close the evasion gap. The evidence says otherwise, and the clearest admission came from the company with the most to gain from a working detector.
OpenAI launched its AI Text Classifier in January 2023 and withdrew it six months later, citing a low rate of accuracy. Its own evaluation had caught roughly 26% of AI-written text while falsely flagging about 9% of human text. The organisation that built the generator could not reliably build the detector.
The independent evidence is worse than that. A Stanford team ran seven commercial detectors over essays known to be human-written and found the false-positive rate on writers using English as a second language reached 61.22%, with 97.8% of those papers flagged by at least one tool, while the same detectors were near-perfect on writing by US eighth-graders. Rewriting the same papers with richer vocabulary dropped the false-positive rate to 11.77%, which tells you what was actually being measured: not authorship, but linguistic sophistication (Liang, Yuksekgonul, Mao, Wu, & Zou, 2023, Patterns). For any institution teaching STEM in English to students who don't speak it at home, which covers most of the Gulf and a large share of engineering faculties everywhere, that is a bias applied precisely to the students least able to contest it.
Institutions reached the same conclusion independently. Vanderbilt University disabled Turnitin's AI detection feature in August 2023, reasoning publicly that even Turnitin's claimed 1% false-positive rate would mean roughly 750 wrongly flagged papers across the university's ~75,000 annual submissions. It noted specific concern about non-native English writers.
The cost of ignoring this is not theoretical. ABC News reported in October 2025 that Australian Catholic University recorded nearly 6,000 alleged academic misconduct cases across its campuses in 2024, with about 90% relating to AI use. The university's Deputy Vice-Chancellor, Professor Tania Broadley, said the figures were substantially overstated: around a quarter of all referrals were dismissed after investigation, and any case resting on Turnitin's AI detector as sole evidence was dismissed immediately. Thousands of students spent months under an accusation that the evidence could not support.
And underneath all of it sits a structural problem that no amount of model improvement touches. Detection is an attempt to prove a negative about a process nobody observed. Even a detector that worked perfectly would return a verdict on authorship, which is not the thing an instructor needs. It would not say whether the student understood the method, where their reasoning broke, or what to teach them next.
So the reframe isn't a preference. It's forced:
How did AI participate in this student's learning, and what evidence shows the student learned?
Part 2: Validity, not cheating
Once you stop asking who used what, the problem clarifies considerably, and it turns out the assessment-security literature had already reframed it before ChatGPT arrived in classrooms.
Phillip Dawson, Margaret Bearman, Mollie Dollinger and David Boud make the argument directly in Assessment & Evaluation in Higher Education (2024), in a paper titled, bluntly, "Validity matters more than cheating." Their point is that an assessment exists to license an inference: this student can do this thing. Generative AI threatens that inference regardless of whether any rule was broken. A task whose output a machine can produce no longer supports the inference, even if every student in the room worked honestly. The integrity framing was always downstream of the validity problem.
That reframe has a practical consequence we built the whole engine around. If the final answer is the only thing you collect, then any tool that produces final answers destroys your assessment. So collect more than the final answer, and collect it in an order that makes the additional evidence hard to outsource.
The field has converged on roughly this shape. Sydney's answer to the audit we opened with was the "two-lane" approach (Liu & Bridgeman), which separates secured assessment of learning from open assessment for learning and entered policy in November 2024. It is a resilience strategy rather than a resistance one: the open lane assumes AI and is designed accordingly, so nothing depends on keeping it out.
Australia's regulator has moved the same way. TEQSA's Assessment reform for the age of artificial intelligence (2023) and its 2025 follow-on hold that detection cannot guarantee integrity and that structural redesign toward secured assurance of learning is the sustainable response; all 203 Australian providers submitted AI action plans by mid-2024. The UK's QAA has issued guidance in the same direction. In the UAE, where our pilots run, there is no dedicated CAA guidance on generative AI in assessment; expectations sit inside general instruments, including the Ministry's 2023 compliance-inspection framework and its requirements on assessment-material retention, rubrics and moderation. An engine that freezes every stage server-side and marks against per-dimension rubrics fits those requirements closely. That is a compliance fit rather than an endorsement, and we won't present it as one. And Perkins, Furze, Roe and MacVaugh's AI Assessment Scale (Journal of University Teaching & Learning Practice, 2024) gave the sector a usable vocabulary for setting AI permission per task rather than per policy.
An Evidence of Learning task is M-ARAF in its shipping form, and our attempt to satisfy all of this in a single instrument: five unaided stages that carry the grade, one logged stage where AI is reachable, and a chronological record that can't be reconstructed after the fact.
Here is what each stage is for, and what the evidence says about it.
Part 3: Why the unaided attempt has to come first
Stage 1 asks the student to solve the anchor problem alone, recording their working, naming the method they chose, and stating their confidence. The attempt then locks. Nothing they do later can edit it.
Two separate literatures say this ordering matters, and neither is about integrity.
The first attempt is a learning event, not just a baseline. Manu Kapur's productive failure work established that having students attempt a problem before instruction produces better conceptual understanding and transfer than instructing first, even though the initial attempts mostly fail. Sinha and Kapur's meta-analysis in the Review of Educational Research (2021) pooled 53 studies and 166 comparisons across more than 12,000 participants and found an overall advantage of d ≈ 0.36 (95% CI 0.20–0.51), with no cost to procedural knowledge and larger effects at higher design fidelity and with older students. The adjacent pretesting literature points the same way: Richland, Kornell and Kao (2009) and Kornell, Hays and Bjork (2009) showed that unsuccessful retrieval attempts before instruction improve later learning relative to equivalent study time, provided corrective feedback follows.
We hold ourselves to the boundary condition too. Sinha and Kapur's companion analysis found scaffolded versus unscaffolded problem-solving-before-instruction averaged g = −0.08 (95% CI −0.34 to 0.28). Productive failure is not automatic. It depends on what follows the failure, which is precisely why Stage 1 is followed by tutoring and revision rather than by a mark and a red pen.
Deciding before seeing the machine's answer is a documented defence against overreliance. Buçinca, Malaya and Gajos (Proceedings of the ACM on Human-Computer Interaction, 2021) studied "cognitive forcing functions," meaning interface designs that interrupt automatic acceptance of AI advice. Requiring a person to commit to their own judgment before the AI's suggestion appears measurably reduced overreliance on wrong suggestions. They also found, honestly, that people dislike the friction and that the benefit is larger for those already inclined to deliberate.
That is the Stage 1 → Stage 2 ordering exactly. The student commits first. The lock is what makes the commitment real.
Which brings us to the line we most want instructors to repeat to their students: a wrong first attempt is not a trap. It is the measurement. Getting it wrong alone and right afterwards is a legible, creditable outcome. The task is built to show that difference, not to punish it.
Part 4: One stage where Mika is reachable, and it doesn't tell you if you're right
Stage 2 is the only place in the cycle where the student can talk to Mika. Mika opens with a question about what the student actually wrote, then works under whatever policy the instructor set: Socratic tutoring by default, hints-only if they want it tighter, or no AI at all. Every exchange is recorded and forms part of the submission. Under Socratic tutoring, Mika never says whether an answer is right. That holds symmetrically, because confirming a guess lets a student guess their way through.
The evidence here is the clearest in this piece, because two large recent trials happened to run the experiment in both directions.
Bastani and colleagues (PNAS, 2025) gave nearly a thousand Turkish high-school students access to GPT-4 maths tutors. Unguarded access, the version that answers, lifted practice performance by 48%. Then they took it away and ran an unaided exam. Those students scored 17% worse than peers who had never had access at all. A guardrailed version that gave hints instead of answers lifted practice performance by 127% and left students statistically level with the control on the unaided exam. The harm was neutralised. The detail we find most instructive: students in the unguarded condition believed they had learned as much. The tool produced fluency and confidence while removing the cognitive work that produces competence.
Running the other way, Kestin, Miller, Klales, Milbourne and Ponti (Scientific Reports, 2025) ran a randomised trial with 194 students in Harvard's Physical Sciences 2 course, comparing an AI tutor against the university's own active-learning classroom, already a research-validated gold standard. The AI tutor group learned significantly more, in less time, and reported higher engagement and motivation. Their tutor was given the correct solutions and explicitly instructed not to hand them over.
Kirk VanLehn's review in Educational Psychologist (2011) explains why the distinction is load-bearing rather than stylistic. Human tutoring produced gains of about d = 0.79 over no tutoring, not Bloom's two sigma, but substantial. Step-based tutoring systems, which engage with the reasoning steps, reached d ≈ 0.76, statistically close to human tutors. Answer-based systems, which work only with final answers, managed roughly 0.40. Engaging with the steps is where the effect lives.
And there is a deliberate architectural decision behind Stage 2 that we want to state plainly: it is not graded. The tutoring produces the interaction log, not a score. The assessment must not depend on a language model behaving perfectly, so every mark in the task lives in a stage where the model isn't the one being assessed. This is the same discipline we described in Part two of this series, where diagnosis runs on a closed, expert-validated list rather than free generation. Steinbach, Bhandari, Meyer and Pardos (L@S, 2025) showed across 252 learners that as an LLM tutor's hallucination rate rose, student confusion rose with it and trust in the feedback collapsed. A confidently wrong tutor is worse than no tutor. We would rather constrain what the model can do than hope it behaves.
Stage 3 then shows the student their locked first attempt beside an editable copy and asks what they changed and why. The gap between the two is itself a measurement, and, per Kluger and DeNisi's classic meta-analysis (1996), it is the reason we keep every diagnosis at the level of the task rather than the student. More than a third of feedback interventions in their review reduced performance, typically when the feedback targeted the person rather than the work.
Part 5: The four stages that carry the weight
Stages 4 through 7 never ask for the anchor answer again. This is the design's central claim, and it's why an answer-generating tool doesn't obviously help with any of them.
Stage 4: Reasoning Check. Three or four short prompts: why this method, why is this theorem applicable, what changes if an assumption is dropped. Marked on whether the reasoning is specific and sound; a correct final answer earns nothing here.
This is the self-explanation effect, and it is one of the best-quantified findings we build on. Bisra, Liu, Nesbit, Salimi and Winne's meta-analysis in Educational Psychology Review (2018) identified 69 effect sizes from 64 research reports covering roughly 6,000 participants, and found a weighted mean of g = 0.55 for prompted self-explanation, holding across subject areas and across both conceptual and procedural knowledge. The authors explicitly recommended exploring computer-generated prompts. Dunlosky and colleagues (2013) rated self-explanation and elaborative interrogation as moderate-utility techniques in their review for Psychological Science in the Public Interest. Marking the justification rather than the result is not a stylistic preference; the benefit comes from generating the explanation, not from being right.
Stage 5: AI Critique. The student is handed a confident, fluent, wrong solution. They must locate the flawed step, explain why it's wrong, correct it, and describe how they spotted it.
Two distinct bodies of evidence support this. The first is the erroneous-worked-example literature. Große and Renkl (Learning and Instruction, 2007) found that identifying and correcting errors can foster learning; Booth, Lange, Koedinger and Newton (2013) and Durkin and Rittle-Johnson (Learning and Instruction, 2012) showed that studying and explaining incorrect examples, especially alongside correct ones, improves conceptual understanding and reduces misconceptions more than correct examples alone. The broader worked-example effect in mathematics sits at around g ≈ 0.48 in recent meta-analysis.
The second is the overreliance literature, and it is the reason this stage exists at all. Bastani's students copied flawed AI output without catching it. Buçinca's participants accepted wrong AI recommendations at measurable rates. Catching a confident, wrong, fluent answer is not a skill students already have, which makes it worth assessing, and worth teaching. Stage 5 asks a student to do to a machine's output precisely what they will have to do to it for the rest of their careers.
The methodological detail matters here: the seeded error comes from the course's own misconception catalogue, the same expert-validated ontology described in Part two. So it's a mistake the curriculum already cares about, and the marking has known ground truth rather than a generated wrong answer nobody has verified. This is also why instructors must review the flawed solution before publishing, and why publication is blocked until they confirm it. A generated "deliberately wrong" solution is sometimes accidentally correct, or wrong in a second unintended place. Either of those makes the stage unanswerable.
Stage 6: Transfer. The same underlying concept in a different context or representation. Explicitly not the anchor problem with the numbers changed.
We want to be careful here, because this is the stage where the research is most cautionary. Barnett and Ceci's taxonomy in Psychological Bulletin (2002) established that transfer probability depends on how far the new context differs across multiple dimensions, and warned against treating far transfer as a single quantity. Detterman's (1993) review is blunter still: reviewers largely agree that little transfer occurs, and demonstrations that lean on hints amount to following instructions rather than transferring understanding.
Our design instinct is right: an isomorphic problem with different numbers is near transfer, and it inflates apparent competence. But the corollary is a limit we hold ourselves to in interpretation: a low Transfer score may reflect how hard transfer is in general, not a specific deficit in that student. We treat the dimension as a diagnostic flag, not a precise cross-student measure.
Stage 7: Dynamic Verification. Generated after stages 1–6 are marked, aimed at the specific weakness that student demonstrated. Two students sitting the same task get different final questions.
The formative-assessment case for this is textbook: Black and Wiliam's Inside the Black Box (1998), Hattie and Timperley's The Power of Feedback (2007), and Shute's Focus on Formative Feedback (2008) all establish that feedback keyed to the specific demonstrated gap is where the effect lives. It's also a scalable analogue of something the integrity literature has independently arrived at: verification by individualised questioning. Sotiriadou, Logan, Daly and Guest (Studies in Higher Education, 2020) introduced interactive orals at Griffith University on exactly this logic: the more the assessment requires the student to account for their own work in real time, the less room misconduct has. A 2025 bioscience study of 722 students across cohorts from 2009–2023 found significant improvements in final assessment performance and course grades after interactive oral assessment was introduced, with no significant differences by gender, international status, or language background.
We'll return to the psychometric cost of Stage 7 in Part 9, because it is real and we'd rather raise it ourselves.
Part 6: The comparison that carries the most weight
Stage 1 and Stage 7 are both unaided, and they sit either side of all the help.
A student weak at Stage 1 and strong at Stage 7 learned something during the task. A student strong at both didn't need it. A student weak at both didn't get there, and the revision in between shows exactly where the help was carrying them.
That structure is what the Bastani result implies you must build. His students looked excellent while the tool was in their hands and worse than baseline the moment it left. If you only measure with the tool present, you cannot distinguish those two states. If you only measure with it absent, you learn nothing about how the student uses help. Measuring on both sides of the assistance is the only arrangement that separates them.
We'll say clearly in Part 9 what this comparison is not.
Part 7: A profile, not a mark
Every task produces six dimensions on a 0–100 scale: Accuracy, Conceptual Understanding, Reasoning, AI Evaluation, Transfer, and Independent Performance. A dimension that no stage fed is dropped and its weight shared among the rest. It is never scored as zero, because switching a stage off is the instructor's decision, not the student's failure.
This is the same argument we made in Part two about grades, applied to assessment output rather than diagnosis. A single percentage is a scalar; teaching needs a map. The specifications- and standards-based grading literature makes the case at the course level. Nilson's Specifications Grading (2015), and studies such as Tsoi and colleagues' (2019) implementation across twelve STEM courses, report benefits in rigour, clarity and instructor time.
We should be honest that this literature is thinner than the cognitive-science evidence above. A 2024 scoping review of alternative grading in STEM found the research base limited and weighted toward perceptions and affect rather than learning outcomes. And multi-dimensional rubrics carry costs of their own: rater reliability has to be engineered rather than assumed, and instructor cognitive load rises with every dimension. We think the diagnostic gain is worth it. We don't claim it's been proven to measure learning more validly than a single mark.
On marking, one rule we apply strictly. Error-carried-forward credits technique, never a result. If a criterion names a value, finding something else does not satisfy it, however sound the working that produced it, and the final-answer criterion has no carry-forward exception at all. On a task that exists to record what a student could actually do, awarding the final mark for a wrong final answer would invert the finding.
Part 8: How the engine maps onto the evidence
No study below tested the Evidence of Learning Engine. What they establish is a set of design principles for assessment under AI availability. We built the cycle to satisfy each one. Judge the fit yourself:
| What the research says | The Mika design choice |
|---|---|
| AI text detectors are unreliable and biased against non-native English writers; detection cannot underwrite integrity (Liang et al.; OpenAI's withdrawal of its own classifier; TEQSA; QAA) | No detection anywhere in the engine. The question is what the work shows, not who wrote it |
| Generative AI threatens the validity of the inference an assessment licenses, whether or not a rule was broken (Dawson, Bearman, Dollinger & Boud) | Stages 4–7 never ask for the anchor answer again; the graded evidence is not the kind a generator produces |
| Attempting before instruction improves conceptual understanding and transfer (Sinha & Kapur; Richland, Kornell & Kao) | Stage 1 is unaided, records method and confidence, and locks before anything can help |
| Committing to a judgment before seeing AI advice reduces overreliance (Buçinca, Malaya & Gajos) | The Stage 1 lock is a cognitive forcing function: the student's own answer is on the record before Mika opens |
| Answer-giving AI harms later unaided performance; hint-giving AI does not (Bastani et al.) | Stage 2 runs Socratically by default and never confirms whether an answer is right |
| Step-based tutoring matches human tutors; answer-based tutoring doesn't (VanLehn) | Mika opens on what the student actually wrote and works through the reasoning, not the result |
| Assessment shouldn't depend on a language model behaving perfectly (Steinbach et al.) | Stage 2 is not graded. The marks live in the stages where the model isn't involved |
| Prompted self-explanation produces g ≈ 0.55 across conceptual and procedural knowledge (Bisra et al.) | Stage 4 marks the justification for specificity and soundness; a correct final answer earns nothing |
| Explaining incorrect worked examples beats correct-only examples for conceptual understanding (Große & Renkl; Booth et al.; Durkin & Rittle-Johnson) | Stage 5 hands the student a fluent wrong solution to audit, with the error drawn from the course misconception catalogue for known ground truth |
| Students accept confidently wrong AI output without catching it (Bastani et al.; Buçinca et al.) | AI Evaluation is scored as its own dimension, because it can't be assumed |
| Isomorphic problems with changed numbers measure near transfer and inflate apparent competence (Barnett & Ceci) | Stage 6 changes the context or representation, not the numbers |
| Feedback keyed to the specific demonstrated gap is where the effect lives (Black & Wiliam; Hattie & Timperley; Shute) | Stage 7 is written after stages 1–6 are marked, aimed at that student's revealed weakness |
| Individualised verification questioning supports integrity and doesn't disadvantage international or second-language students (Sotiriadou et al.; 2025 bioscience cohort study) | Stage 7 is a scalable written analogue of the interactive oral |
| Feedback aimed at the person can backfire (Kluger & DeNisi) | Every dimension and diagnosis describes the work, never the student |
We didn't choose these stages and then find citations for them. The cycle is shaped the way the evidence said it should be.
Part 9: What we've actually seen, and what we haven't
Here is where we have to be more restrained than in the previous two pieces in this series, and we'd rather say so than blur it.
The tutoring pilot results we've reported elsewhere, the pass-rate and average-score lifts across 600+ students, belong to Ask Mika. They are not evidence for this engine. We have no cohort efficacy data for the Evidence of Learning Engine. What we have is a Calculus II test run, and we publish it because it shows what the output is for, not because it demonstrates that the output is valid.
The student was set one anchor problem: evaluate ∫ (3x + 5) / (x² + 2x − 3) dx.
Alone, they factored the denominator as (x−3)(x+1) instead of (x−1)(x+3), carried the error cleanly through partial fractions, and reached a wrong answer by sound method. Mika's opening question sent them back to the factorisation. They got it right on revision, and later missed the planted error entirely.
Stage marks
| Stage | Mark |
|---|---|
| Independent attempt | 2 / 5 |
| AI support | 8 exchanges |
| Revision | 5 / 5 |
| Reasoning check | 4 / 6 |
| AI critique | 0 / 4 |
| Transfer | 4 / 4 |
| Verification | 3 / 3 |
Dimension profile
| Dimension | Score |
|---|---|
| Accuracy | 76 |
| Conceptual | 67 |
| Reasoning | 78 |
| AI Evaluation | 0 |
| Transfer | 100 |
| Independent | 46 |
A single percentage would have read 63% and told an instructor almost nothing. The profile says something specific: this student can apply the technique and move it to a new setting, but could not get there without help, and could not audit a wrong solution at all. Those are two different pieces of teaching, and neither is visible from a mark.
That is a demonstration of diagnostic resolution. It is not evidence of efficacy, and we won't present it as such.
What a measurement specialist would demand next, and we think they'd be right. Four things, stated as we'd state them to a reviewer:
The single anchor problem is the sharpest objection. The performance-assessment literature is decisive and inconvenient here. Shavelson, Baxter and Gao (Journal of Educational Measurement, 1993), and Gao, Shavelson and Baxter (Applied Measurement in Education, 1994), found that in mathematics and science performance assessment, task-sampling variability, the person × task interaction, is the dominant source of measurement error. Raters can be trained to agree; tasks refuse to generalise. Shavelson and colleagues later called it the Achilles' heel of science performance assessment. One anchor problem, however deep, is a thin basis for a general ability inference. We need a generalizability study across multiple anchors to know how many a reliable profile actually requires.
The Stage 1 → Stage 7 comparison is not a clean gain score. It's confounded by practice effects (doing the material once improves the second attempt regardless of learning), by regression to the mean (weak first attempts drift upward), and, most importantly, by deliberate non-parallelism. Stage 7 is targeted at the student's weakness, which is exactly what makes it good formative feedback and exactly what stops it being psychometrically parallel to Stage 1. We treat the comparison as strong diagnostic signal. It is not a validated learning-gain measure, and we won't report it to a registrar as one.
Different final items across students is a real fairness question. Adaptive testing solves comparability with a calibrated item bank and an IRT model. We don't have one yet. Until we do, Stage 7 should be read as verification and diagnosis rather than as a cross-student-comparable score.
Time on task is an upper bound, not attention. The server timestamps each commit, so the sequence and its intervals are dependable. Whether the student was working during an interval is not. Read the timestamps; treat the durations as context.
And the study we'd need to run to claim the engine works: a randomised trial, ideally cluster-randomised by section, comparing the full cycle against a conventional unaided exam and against a process-visible non-Mika control, matched on content, time-on-task and instructor, with the primary outcome a delayed, external, unaided transfer test that Mika did not write. Powered for a modest effect, since the component effect sizes above sit around d = 0.35–0.55; realistically several hundred students across multiple sections. Plus subgroup analysis by language background, not because it's the premise of the design, but because it's where the tools we're replacing demonstrably failed, and a replacement should be held to the standard the incumbent missed.
Until that exists, the honest description of M-ARAF is: assembled from well-evidenced components, with the combination not yet independently validated. We'd rather publish that sentence than have a reviewer write it for us.
Three boundaries belong next to it, stated rather than waited for. The AI policy governs Mika, not the internet: a student can open another tab and nothing here will know, which is exactly why stages 4 to 7 are built to measure what they can do unassisted instead of to police what they had open. Kofinas and colleagues (2025) showed empirically that generative AI can circumvent assessments specifically designed to resist it, so what we claim is resilience, not immunity, and any vendor claiming immunity is selling something. And this instrument replaces a quiz, not a problem set: one anchor problem at 25 to 40 minutes is a deep measurement of a concept that matters, not a way to cover twelve exercises.
The bottom line
Resistance asks whether a student used a machine. In mathematics and the sciences that question mostly cannot be answered, and where it can, it is answered unreliably and with a bias against precisely the students most universities in this region teach. Worse, it is the wrong thing to want to know.
Resilience asks something answerable instead: does this task still tell me what this student can do, given that the tool exists and is not going away? That question has an answer, and the answer is an architecture. Collect more than the final answer, in an order that makes the extra evidence hard to outsource. An unaided attempt that locks before help arrives. One logged stage of tutoring that never confirms an answer. Then four stages that ask the student to justify a choice in their own words, to audit a confident wrong solution and find the exact step that fails, to recognise the same idea wearing different clothes, and to answer a question that did not exist until their own earlier work was marked.
Each of those is a research finding before it is a feature, and the table in Part 8 shows the trace from one to the other. What we have not yet shown is that the assembly outperforms the alternatives, and we have said above precisely what it would take to find out.
Nothing in this design requires AI to be kept out of the room. That is the entire point of building for resilience instead of resistance, and it is a claim we are happy to be measured on.
What happens next
Live classroom pilots of the Evidence of Learning Engine are running now. The research programme described in Part 9 is underway rather than aspirational: we are collecting per-stage and per-dimension data across real cohorts, with a comparison condition, a delayed unaided post-test that Mika did not write, and subgroup analysis by language background.
We will publish what comes back. If the results do not support the design, that will be the fourth piece in this series, and we will write it with the same citations and the same care.
We are looking for more partner universities. If you teach mathematics, physics, chemistry or engineering at scale, and you are rebuilding assessment for a world where the tool is not going away, we would like to run a cohort with you. Partner institutions get the engine, per-dimension cohort analytics for their own faculty, and co-authorship on what we publish. What we need in return is a real course, a real term, and permission to report the findings honestly, whichever way they fall.
We are genuinely impatient to share these results. Assessment in the sciences has been arguing about detection for three years, and the argument has produced very little any instructor can use on a Monday morning. Evidence would move it.
To discuss a pilot, write to us at [email protected] or through mikalabs.org.
References
Barnett, S. M., & Ceci, S. J. (2002). When and where do we apply what we learn? A taxonomy for far transfer. Psychological Bulletin, 128(4), 612–637. https://doi.org/10.1037/0033-2909.128.4.612
Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122. https://doi.org/10.1073/pnas.2422633122
Bearman, M., Tai, J., Dawson, P., Boud, D., & Ajjawi, R. (2024). Developing evaluative judgement for a time of generative artificial intelligence. Assessment & Evaluation in Higher Education, 49(6), 893–905. https://doi.org/10.1080/02602938.2024.2335321
Bisra, K., Liu, Q., Nesbit, J. C., Salimi, F., & Winne, P. H. (2018). Inducing self-explanation: A meta-analysis. Educational Psychology Review, 30(3), 703–725. https://doi.org/10.1007/s10648-018-9434-x
Black, P., & Wiliam, D. (1998). Inside the black box: Raising standards through classroom assessment. Phi Delta Kappan, 80(2), 139–148.
Booth, J. L., Lange, K. E., Koedinger, K. R., & Newton, K. J. (2013). Using example problems to improve student learning in algebra: Differentiating between correct and incorrect examples. Learning and Instruction, 25, 24–34. https://doi.org/10.1016/j.learninstruc.2012.11.002
Buçinca, Z., Malaya, M. B., & Gajos, K. Z. (2021). To trust or to think: Cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1), 1–21. https://doi.org/10.1145/3449287
Dawson, P., Bearman, M., Dollinger, M., & Boud, D. (2024). Validity matters more than cheating. Assessment & Evaluation in Higher Education, 49(7), 1005–1016. https://doi.org/10.1080/02602938.2024.2386662
Detterman, D. K. (1993). The case for the prosecution: Transfer as an epiphenomenon. In D. K. Detterman & R. J. Sternberg (Eds.), Transfer on trial: Intelligence, cognition, and instruction (pp. 1–24). Ablex.
Dunlosky, J., Rawson, K. A., Marsh, E. J., Nathan, M. J., & Willingham, D. T. (2013). Improving students' learning with effective learning techniques. Psychological Science in the Public Interest, 14(1), 4–58. https://doi.org/10.1177/1529100612453266
Durkin, K., & Rittle-Johnson, B. (2012). The effectiveness of using incorrect examples to support learning about decimal magnitude. Learning and Instruction, 22(3), 206–214. https://doi.org/10.1016/j.learninstruc.2011.11.001
Gao, X., Shavelson, R. J., & Baxter, G. P. (1994). Generalizability of large-scale performance assessments in science: Promises and problems. Applied Measurement in Education, 7(4), 323–342. https://doi.org/10.1207/s15324818ame0704_4
Große, C. S., & Renkl, A. (2007). Finding and fixing errors in worked examples: Can this foster learning outcomes? Learning and Instruction, 17(6), 612–634. https://doi.org/10.1016/j.learninstruc.2007.09.008
Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research, 77(1), 81–112. https://doi.org/10.3102/003465430298487
Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). AI tutoring outperforms in-class active learning: An RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports, 15, 17458. https://doi.org/10.1038/s41598-025-97652-6
Kluger, A. N., & DeNisi, A. (1996). The effects of feedback interventions on performance. Psychological Bulletin, 119(2), 254–284. https://doi.org/10.1037/0033-2909.119.2.254
Kornell, N., Hays, M. J., & Bjork, R. A. (2009). Unsuccessful retrieval attempts enhance subsequent learning. Journal of Experimental Psychology: Learning, Memory, and Cognition, 35(4), 989–998. https://doi.org/10.1037/a0015729
Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), 100779. https://doi.org/10.1016/j.patter.2023.100779
Nilson, L. B. (2015). Specifications grading: Restoring rigor, motivating students, and saving faculty time. Stylus.
Perkins, M., Furze, L., Roe, J., & MacVaugh, J. (2024). The AI Assessment Scale (AIAS): A framework for ethical integration of generative AI in educational assessment. Journal of University Teaching & Learning Practice, 21(6). https://doi.org/10.53761/q3azde36
Richland, L. E., Kornell, N., & Kao, L. S. (2009). The pretesting effect: Do unsuccessful retrieval attempts enhance learning? Journal of Experimental Psychology: Applied, 15(3), 243–257. https://doi.org/10.1037/a0016496
Shavelson, R. J., Baxter, G. P., & Gao, X. (1993). Sampling variability of performance assessments. Journal of Educational Measurement, 30(3), 215–232. https://doi.org/10.1111/j.1745-3984.1993.tb00424.x
Shute, V. J. (2008). Focus on formative feedback. Review of Educational Research, 78(1), 153–189. https://doi.org/10.3102/0034654307313795
Sinha, T., & Kapur, M. (2021). When problem solving followed by instruction works: Evidence for productive failure. Review of Educational Research, 91(5), 761–798. https://doi.org/10.3102/00346543211019105
Sotiriadou, P., Logan, D., Daly, A., & Guest, R. (2020). The role of authentic assessment to preserve academic integrity and promote skill development and employability. Studies in Higher Education, 45(11), 2132–2148. https://doi.org/10.1080/03075079.2019.1582015
Steinbach, M., Bhandari, S., Meyer, J., & Pardos, Z. A. (2025). When LLMs hallucinate: Examining the effects of erroneous feedback in math tutoring systems. In Proceedings of the Twelfth ACM Conference on Learning @ Scale (L@S '25). https://doi.org/10.1145/3698205.3729555
Tertiary Education Quality and Standards Agency. (2023). Assessment reform for the age of artificial intelligence. TEQSA. https://www.teqsa.gov.au/guides-resources/resources/corporate-publications/assessment-reform-age-artificial-intelligence
Tsoi, M. Y., Anzovino, M. E., Erickson, A. H. L., Forringer, E. R., Henary, E., Lively, A., Morton, M. S., Perell-Gerson, K., Perrine, S., Robinson, R., Sanchez, I., & Venkatesh, A. (2019). Variations in implementation of specifications grading in STEM courses. Georgia Journal of Science, 77(2), Article 7.
VanLehn, K. (2011). The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems. Educational Psychologist, 46(4), 197–221. https://doi.org/10.1080/00461520.2011.611369
A note on the evidence in this piece: the studies cited here have been checked against their original sources. None of them tested the Evidence of Learning Engine, or M-ARAF as a whole; they establish the principles it is built on. Where we cite journalism or institutional statements rather than peer-reviewed work (the OpenAI classifier withdrawal, the Vanderbilt and Australian Catholic University decisions, the Sydney policy review), we have labelled them as such in the text. Unlike the first two pieces in this series, this one reports no pilot outcomes, because we don't have any for this feature yet. The worked Calculus II case is from a single test run and is offered as an illustration of output, not as evidence of effect.