MIT's committee on AI in teaching and learning reported in August 2026. Mika was already built. This piece reads one against the other — where they converge, where we are ahead, and where the report is a challenge we have not yet met.
On 13 August 2026, MIT's Ad Hoc Committee on AI Use in Teaching, Learning, and Research Training published its final report. It runs to three sections and roughly thirty recommendations, and it is the most consequential institutional statement on AI in higher education so far — not because MIT's conclusions are novel, but because MIT wrote them down and attached a budget.
Mika shipped in September 2025, with the Evidence of Learning Engine and the design principles behind it in place. The committee was charged in January 2026 and reported in August.
We are stating those dates once and then not returning to them, because precedence is the least interesting thing in this piece. What is interesting is that a committee of MIT faculty, students and staff, working from listening sessions and surveys across five schools, arrived at an architecture that matches one built independently in Abu Dhabi from the assessment literature. Neither party consulted the other. That convergence is worth more than either document alone, because it is the ordinary sign that a design conclusion is real rather than fashionable.
This piece maps the overlap, names the places we think we are structurally ahead, and then spends the rest of its length on where the report is a challenge to us.
Part 1: Four places the report and the engine reach the same conclusion
Redesign, not AI-proofing. The report's central assessment move is to urge instructors past "AI-proofing" toward assessments that build in the productive struggle that produces durable understanding — oral exams, semester portfolios, out-of-class work paired with in-class conversation (§3.1.2). It reaches this by way of a prior observation: AI can now produce credible responses to almost any written assignment in the undergraduate curriculum, including proofs and problem sets.
That is the argument we made in AI-Resilient, Not AI-Resistant, phrased differently. We called it resilience over resistance. The report calls it AI-aware course design. The structural consequence is identical: stop collecting only the final answer. Stages 4 through 7 of an Evidence of Learning task never ask for the anchor answer again.
No detectors. §3.1.9 recommends against relying on AI detection, on three grounds: an arms race with humanizers that serves nobody, false positives against non-native English writers and neurodivergent students, and the adversarial classroom atmosphere that policing produces. It notes that MIT's own Committee on Discipline does not treat detector output alone as sufficient evidence.
We reached the same position from the same literature — Liang et al.'s 61.22% false-positive rate on second-language writers, OpenAI's withdrawal of its own classifier, Vanderbilt's arithmetic. There is no detection anywhere in the engine. For institutions teaching STEM in English to students who do not speak it at home, which describes most of the Gulf, this is not a philosophical preference.
Policy per course, with a stated reason. §2.6 rejects a uniform institution-wide rule as inevitably too permissive for some contexts and too restrictive for others, and asks instead for a shared menu of options that departments and instructors select from. §3.1.8 adds that every subject should carry a clear, prominently posted policy — and that the policy must include a rationale tied to the course's learning goals, because a rule students understand is one they are more likely to keep.
Mika ships a three-policy menu — no AI, hints only, Socratic — set per course and refined per stage, with course-level instructions on top. The policy prompt is assembled server-side, so a student cannot weaken it by changing their own settings. The report asks for a menu; we built the menu and the enforcement.
A profile instead of a number. §3.1.6 asks MIT to reconsider what role grades play at all, and encourages exploring competency-based and mastery-based paradigms, portfolios, and feedback channels that tell a student more than a mark can. It observes that real-time feedback during an oral exam probably conveys more about mastery than any grade.
Every Evidence of Learning task outputs six dimensions rather than one mark: Accuracy, Conceptual Understanding, Reasoning, AI Evaluation, Transfer, Independent Performance. Our argument for it is the one we made in The Diagnostic Loop — a grade is a scalar and teaching needs a map. We have also said, and repeat here, that the alternative-grading evidence base is thinner than the cognitive-science evidence, and that remains true whoever is recommending it.
Part 2: Two places the mechanism already exists
§3.2.4 asks that students learn to verify an AI output and to recognise when a model is likely to hallucinate. §3.1.9 suggests something more concrete: that instructors require students to work on platforms that capture a version history, and to submit that history alongside the work, as process evidence.
Both of those are requests for a mechanism the report does not have. Disclosure, in almost every institutional AI policy written so far, is self-reported — a line in a cover sheet declaring which tools were used and how. It is unverifiable by construction, and it asks the least honest student for the most honesty.
Stage 2 of an Evidence of Learning task records every exchange with Mika as part of the submission. Stage 1 locks before any help arrives and cannot be edited afterward. The stages commit server-side with timestamps. The student sees, on the task cover screen, which stages Mika is reachable in and that their instructor sees every step including anything they ask. Disclosure is not a declaration; it is the artifact.
The logging is not aimed only at students. Every actor on the platform is recorded — students, instructors, coordinators, institutional IT, and Mika Labs. There is no privileged role that acts unobserved. This matters more than it first appears, because the usual objection to student-side process capture is that it is surveillance pointed downward: the people with the least power are the most instrumented. On Mika the audit trail is symmetric, and an institution can inspect ours as readily as its own.
One gap remains, and we would rather name it than let it be found. §3.2.3 asks for something adjacent but distinct: that instructors be transparent to their students about their own AI use, because students notice — and read as a double standard — when faculty use AI for slides, feedback or grading while restricting student use. A symmetric audit log satisfies governance. It does not satisfy that complaint, because the student never sees it. A per-course instructor disclosure line, surfaced to students on the same cover screen that already tells them their own policy, would close it. We have not built that yet.
And stage 5 is §3.2.4 as a graded act. The student is handed a confident, fluent, wrong solution, with the error drawn from the course's own misconception catalogue, and must find the failing step and say how they spotted it. Verifying model output is not a skill students arrive with — Bastani's cohort copied flawed output without catching it — which is exactly why it is worth assessing rather than assuming.
These are the only claims in this piece we would describe as being ahead, and we make them narrowly: not that we understood something MIT did not, but that the things they ask for are already built.
Part 3: Where the report is a challenge to us
This is the part of the piece that makes it worth publishing, and we would rather write it than have a reviewer write it.
The in-person recommendation is the one no platform can satisfy alone. §3.1.4 asks that every subject include a structured in-person social component, and the report's first listed harm is behavioural: office-hour attendance down, online discussion down, study groups thinning in dorms and libraries. This is the deepest charge against AI tutoring, and it is aimed at products like ours.
Our design answer is that Mika routes toward the instructor rather than substituting for them. Under Socratic policy it never confirms whether an answer is right; on defined conditions it hands the student off; the Consultant surfaces the flagged student to the instructor, who decides whether to reach out and schedule time. The claim we want to make is stronger than "we don't displace contact" — it is that diagnostic reach can manufacture contact that would not otherwise have occurred, because a student who would never have booked an office hour gets identified and invited.
We cannot yet evidence it. The report's harm is a measured change in behaviour; a feature list is not a rebuttal to it. What would be is countable, and we are collecting it: referrals raised, share acted on by the instructor, share converted to a scheduled session, and the comparison between referred students and matched peers on the concept they were stuck on. Until that exists, the honest description is design intent, and a reader should treat it as such.
Environment, models, and the parts of §3.2.5 and §3.3.10 we can only partly answer. The report asks institutions to publish information about the real environmental and financial cost of AI use, and records that some community members object to how training data was collected.
What we can state: our contracts prohibit training on institutional or student data, as contract language rather than as an inference from our certifications. Our inference runs on Google Cloud infrastructure, which is currently the only major AI serving stack to publish production-metered rather than modelled environmental figures — for its own assistant, Google reports 0.24 Wh of energy, 0.03 gCO₂e and 0.26 mL of water for a median text prompt, on a measurement boundary that includes host CPU and memory, provisioned idle capacity and data-centre overhead. And our Vector open-weight line exists partly so that institutions who want inference under their own roof can have it.
What we will not claim: that any of this makes us the low-impact option. Google's published figure describes its own product, not a Mika session, which carries vision, long-context curriculum grounding and extended reasoning; it is a reference point for what an honest measurement looks like, not a number we are entitled to quote as ours. And per-prompt efficiency coexists with absolute growth: the hyperscalers' total emissions have risen sharply since 2019 with AI a named driver. We have not published a per-student footprint for a Mika course. We should, and that is on the list below.
Part 4: What the report asks for that nobody has yet
The report is disciplined about its own uncertainty — its first guiding principle is humility, and it expects major course corrections. Read against the limits we have already published on our own work, four gaps are shared rather than ours alone.
Nobody has cohort efficacy data for AI-resilient assessment. We have said in print that we have none for the Evidence of Learning Engine, and that the study required is a cluster-randomised trial with a delayed, external, unaided transfer test that Mika did not write, plus subgroup analysis by language background. The report recommends alternative assessment without pointing to that evidence either, because it does not exist yet. Whoever produces it first will have done the sector a service.
Nobody has behavioural data on whether AI platforms can restore in-person contact rather than erode it. §3.3.6 asks MIT to start tracking metrics on AI use, campus engagement and student satisfaction. We are instrumenting the referral loop. These are the same measurement, run from two ends.
Nobody has resolved the logging question. §3.3.9 sets out the genuine conflict: students bring personal and sensitive material to these systems, instructors have a legitimate interest in seeing whether learning goals are being met, and no policy yet reconciles them. Our boundary is that assessment interaction is part of the submission and disclosed on the cover screen. General study conversation is a different category and we think it should stay that way. We do not claim to have solved the problem, only to have drawn a line where we could defend it.
And nobody in this sector publishes a per-student environmental footprint. We have not. We intend to.
We are looking for partner universities that want to be at the leading edge of AI in education to pilot Mika with us.
A note on sources: the MIT report is cited by section number throughout, from the published text at aiandeducation.mit.edu. Section numbers refer to that version. No part of this piece should be read as MIT endorsing Mika or Mika Labs; the report makes no reference to any commercial product, and its §3.2.5 records substantial community wariness about vendors in this market. The research underlying our own design decisions is cited in full in AI-Resilient, Not AI-Resistant and is not repeated here.