The homework fight is about proof

Pumulo SikanetaAugust 26, 2026educationartificial intelligenceassessmentpolicy
AI impact on student learning
AI impact on student learning

I dropped my daughter back at college two weeks ago for her second year. Somewhere on the drive home I realized I could tell you what she is studying and not what she is learning, and that those are two different questions.

The second one is the one I could not answer. She will hand in work this year that a machine could have produced in ninety seconds. So will every student in her building, and every ninth-grader in the school down my road. What does a grade mean now.

A high school in Turkey went looking for that answer with nearly a thousand of its own students.

Grades nine to eleven. About fifty classrooms. Four math lessons, ninety minutes each, covering about 15 percent of that semester's syllabus. Some classes practiced the ordinary way. Some got ChatGPT. And some got a version the researchers had wrapped in guardrails, built around problems and worked solutions the teachers had supplied, so that it gave hints and refused to hand over a finished answer.

Then the researchers switched the machines off and gave everybody an exam.

The students who had practiced with plain ChatGPT had done 48 percent better than the control group while they still had it. On the exam without it, they scored 17 percent worse.

Hamsa Bastani and her colleagues at the University of Pennsylvania published that in June 2025, in the Proceedings of the National Academy of Sciences. It is the cleanest fact in this whole argument and it points two ways at once. The tool made the work better. It made the student worse.

The third group is the part people skip. The students with the guarded tutor, the one that withheld answers, practiced 127 percent better than the control group. Then they walked into the exam and came out roughly level with everyone else. No damage. No gain either. The guardrails stopped the harm. They did not manufacture learning.

Keep those three numbers somewhere near the front of your mind, because most of what you will read about AI in schools this year is written by somebody who has only met one of them.

Students are being told two opposite things

At work, people are being told to learn these tools or fall behind. Employers want staff who can research faster, draft quicker, find patterns and clear routine work off the desk before lunch.

At school, a student doing the same thing may be called a cheat.

The worry underneath that is fair. A student who asks a machine to write the book report, invent the sources and polish the final draft can collect a grade without learning anything. This article acknowledges this concern and the focus is on enhancing learning rather than curtailing it.

School is where a young person builds the muscles for hard questions. Attention. Patience. Memory. Suspicion of a confident answer. The willingness to change your mind when the evidence tells you to. Those muscles do not grow when somebody else does every lift.

But a flat ban has its own cost, and the cost is that students are being prepared for a world that no longer exists.

Look at what is actually happening in houses right now. Common Sense Media surveyed 1,204 children aged nine to seventeen in March 2026. Eighty-six percent use or interact with AI. Of the ones who use it, 85 percent use it for schoolwork, and one in five does so every single day.

Now look at the adults. Gallup surveyed 2,069 American public school teachers in February and March 2026. Eighteen percent had received formal written guidance on using AI at work. Thirty-four percent had received none of any kind.

So the rule that students are expected to follow has, in most cases, never been written down for the person enforcing it.

The calculator story is not the story you were told

Somebody always brings up calculators. The version they bring up is usually wrong in three specific ways, and the real history is more useful than the myth.

The myth says the math establishment panicked and was proved wrong. It did not panic. The National Council of Teachers of Mathematics put out its first position statement in 1974, while pocket calculators were still expensive, and it said teachers "should recognize its potential contribution as a valuable instructional aid." By 1980 the Council was recommending that math programs "take full advantage of the power of calculators and computers at all grade levels." The resistance came from parents, school boards and individual teachers. The profession was in front of it, not behind it.

The myth also says the research came back clean. It came back conditional. Ray Hembree and Donald Dessart pooled 79 studies in 1986 and found that calculators improved paper-and-pencil skills, with one exception. Grade four. At grade four, sustained calculator use appeared to hinder basic skills in average students. Aimee Ellington pooled 54 studies in 2003 and found the gains held when calculators were part of the teaching and part of the test. Nobody ever found that calculators improved unaided arithmetic, because nobody ever claimed they would.

And the myth says everything turned out fine. Something did happen to computation, and the honest answer is that we do not know what caused it. Tom Loveless went through the national math assessment items for Brookings in 2004 and found that American arithmetic scores rose through the 1980s and then fell through the 1990s, most clearly for nine-year-olds across all four operations. He named classroom calculator use as one plausible contributor and then wrote, plainly, that "no one can be certain what caused the decline."

Loveless also found the single most useful number in the calculator literature. Nine-year-olds in 1999, answering the same multiplication questions. Without a calculator, 42.5 percent got them right. With one, 87.9 percent did. He called that the difference between signaling mastery and signaling incompetence.

That is the real lesson, and it is not about calculators at all. Give a child a tool during the test and you have changed what the test measures. You have not changed what the child knows.

The historian Bronwen Everill made the sharpest version of this point in 2025. The calculator argument was about whether to teach the basics or use technology for higher-order thinking. The AI argument is about whether to use technology to skip the thinking altogether. Those are not the same question, and pretending they are is how a school ends up with a policy that does not work.

Four students, and what the evidence says about each

Before arguing about whether students should use AI, it helps to remember who is in the classroom.

So here are four of them. They are broad caricatures, drawn that way on purpose, and they are meant to show different styles and stages of learning. No child is only one of them for long. The point is that thirty students do not arrive at the same worksheet with the same obstacles.

Most of us have been all four at different times in the same week.

The Archivist wants the answer that matches

The Archivist does the reading, highlights the parts that look important, follows the instructions and can give back the chapter with real accuracy. Ask for the definition from page 42 and page 42 arrives intact.

School accidentally rewards this. A student can produce a perfect definition of osmosis, collect full marks, and have no idea why a slug dies in salt.

Think of somebody who has memorized the driver's manual. They know the stopping distances. They know every road sign. They can pass every practice test in the book. Then they hit a wet road at night and see brake lights coming up fast, and the manual is not the thing that helps them.

Yizhou Fan and colleagues, including Dragan Gašević, published a study of 117 university students in 2025 and split them four ways. One group had ChatGPT. One had a human expert. One had writing analytics. One had nothing extra. The ChatGPT group produced the best essays by a clear margin. On knowledge gained and knowledge transferred to a new problem, they were no better than anyone else. The researchers named the pattern metacognitive laziness, which is a long phrase for a short thing. The tool did the noticing, so the student stopped noticing.

There is a more startling version of the same finding, and I want to hand it to you with the caveats attached rather than after. Nataliya Kosmyna and a team at the Massachusetts Institute of Technology (MIT) put a cap of electrodes on people writing essays. The group using a chatbot showed the weakest brain connectivity, and 83 percent of them could not quote a single line from the essay they had just submitted, against 11 percent in the other groups. That study is still not peer reviewed. The headline part rests on eighteen people. The authors have published a page asking journalists to stop describing it with words like brain damage and brain rot, which tells you how it was covered.

So use it as a suggestion, not a finding. It matches Fan's result, which is better built, and that is why it is worth a mention at all.

The Archivist should stop asking the tool to compress the chapter. Paste in your own notes and ask something like this instead.

Give me three real situations where this rule would not work as expected. Do not explain them yet. Ask me what I think first.

That turns the machine into a sparring partner. A history student can ask it to argue as somebody who disagreed with the textbook. A science student can ask for the same concept explained through a scientific analogy. A literature student can ask it to attack their reading of a character.

The point is to leave the harbor of recall and go out into the rougher water where judgment lives. In working life, a person who can recite the company policy is useful. A person who can identify when the policy will fail is worth a great deal more.

The Spark needs a doorway, and access is not a doorway

The Spark gets called lazy. Sometimes that is accurate, because sometimes people avoid work, adults very much included. But the label buries a better question, which is what would make this student want to start.

Imagine they are learning percentages. The examples involve imaginary bags of apples and a shop nobody has ever visited. Nothing they would ever really think about.

The same student follows basketball and wants to know whether a player's shooting is genuinely better than last season. Or plays a game with in-app purchases and wants to know what a 20 percent discount really saves over a month. Or cares about a local tax proposal and wants to know what it does to a household budget.

The math did not change. The doorway did.

Here is where a lot of people leap to the wrong conclusion, so watch what happened when somebody built the doorway and left it open.

Philip Oreopoulos and Nina Low studied Khanmigo, the Khan Academy AI tutor, in 18 Tennessee middle schools over two school years. They published what they found in August 2026.

Almost every student tried it. Ninety-six percent opened it at least once.

Then they stopped talking to it. On two days out of every three that a student sat down to practice, they sent the tutor nothing at all. When a student got a question wrong, they asked the tutor for help about one time in six. And most of the messages they did send were not about math.

The learning gains were real and very small. The researchers concluded that a class with the AI tutor ended up looking much like a class using Khan Academy without it.

Access was never the bottleneck. Wanting to was.

A tutor sitting in the corner of the screen that nobody talks to is not a tutor. It is a widget. The teacher who finds the basketball question is doing the part that the software cannot do, and that part has not been automated and is not close to being automated.

Personalizing a lesson does not mean lowering the bar. It means finding a road that arrives at the same place. The old model treated a student's interests as a distraction. The better model treats them as kindling.

The Climber is the student these tools can hurt most

She started differential equations last week. She called home after the first lesson and said it was confusing. The material was fine. What she could not follow was her teacher's way of explaining it.

So she sat down that evening and worked through it from other explanations, and it took her about an hour. By the end of the week she had learned to translate her teacher, and the class was fine.

That translation was the cost, and she paid it once, in a week, because she is confident enough to go looking for another explanation. A student who is not confident would have sat in that class until December, believing the problem was differential equations.

The Climber is the one we fail quietly. They nod during the lesson because they do not want to slow thirty people down. They copy every worked example, because from across the room copying looks like learning. At home they stare at the page until the words stop meaning anything.

One explanation, at one speed, in one style, does not reach everybody. This is where a patient machine could do the most good, and it is also where the evidence gets uncomfortable.

Lehmann and colleagues ran three studies of university students learning Python, one in the field with 113 people and two in a lab with 107 and 69. The students with AI covered more topics and understood less. The gap widened for the ones who had started with the least prior knowledge. In a separate study of 273 German ninth-graders working on a physics problem, Becker and colleagues found the chatbot significantly reduced how hard the work felt. That was reported as a benefit. In a subject you are trying to learn, it is not obviously a benefit at all.

Meanwhile the children who need scaffolding are the ones embracing it most enthusiastically. In the Common Sense survey, 56 percent of kids who report trouble focusing use AI weekly for homework, against 45 percent of those who do not.

But there is one result that jumps out to me, and it points somewhere else entirely.

Rose Wang and colleagues at Stanford built Tutor CoPilot, which does not tutor the student at all. It coaches the human tutor, live, while the lesson is happening. Nearly eight hundred tutors and about a thousand students in low-income schools. Students were four percentage points more likely to master the topic. For the least experienced tutors, seven points. For the lowest-rated tutors, nine points.

The tool worked best when it was pointed at the adult.

That is the finding I would put on a wall in every district office. AI aimed at replacing the helper produced small gains and some harm. AI aimed at making an ordinary helper better produced the biggest gains for the students furthest behind.

For the Climber personally, the useful move is to make the machine slow down.

Explain why three quarters is bigger than two thirds using pizza first. Then give me a different example with no pizza in it. Then ask me one question at a time and wait for my answer before you carry on.

That is not cheating. That is tutoring, and a good tutor never grabs the pencil.

The Runner needs a harder question, not a longer worksheet

The Runner has finished. They understood it the first time, they did the work, and now they are waiting.

The standard response is more problems that look exactly like the first problems. That is not enrichment. That is the academic version of clearing your plate quickly and being handed three more bowls of peas.

Here is the best evidence anyone has that AI can help this student. Gregory Kestin and Kelly Miller at Harvard built a tutor on GPT-4 for an introductory physics course and ran a proper trial, with 194 of the 233 enrolled students taking part, each doing one lesson with the tutor and one in class. The AI group's median score was 4.5 against 3.5 for the class. They finished in a median 49 minutes against the hour the lesson took.

Now the caveats, because they matter more than the headline. It was introductory material. It was tested immediately, with no measurement of what anyone remembered a month later. And the authors wrote themselves that the approach "may not always outperform in-class active learning in all contexts, for example, those requiring complex synthesis of multiple concepts and higher-order critical thinking."

Which is a careful way of saying it worked on the easy half.

For the Runner, the useful prompt is never for more homework.

I understand the basic idea. What is a harder question that experts still argue about? Give me the sources and the terms I should look up, then ask me to take a position and push back on it.

The Runner's real lesson is that speed is not depth. The world does not reward the fastest person in the room. It rewards the ones who keep going after the obvious answer has already turned up.

The lab result and the hallway result are not evaluating the same thing

Put the last two sections side by side and you have the most important pattern in this field.

In a Harvard lab, with a tutor built by physicists and a lesson designed around it, the effect was large. In eighteen Tennessee middle schools, with a product off the shelf and real twelve-year-olds, the effect was small enough that you would struggle to see it in a report card.

That gap is the story. Almost every confident claim you will hear about AI and learning is a lab number being quoted as if it were a hallway number.

It gets worse. In 2026, a Stanford group called SCALE, short for Systems Change Advancing Learning and Equity, went through more than 800 papers on AI in schools. Twenty had strong causal evidence. Eight percent of the whole pile were randomized trials. And their finding, in their own words, was that "we did not identify any high-quality causal studies in K-12 settings in the U.S. for students."

Not one. In a debate this loud.

Then there is the paper everybody quoted. In 2025, a meta-analysis by Wang and Fan reported that ChatGPT improved learning performance with an effect size of 0.867, which would be enormous. It went into ed-tech sales decks immediately and was cited hundreds of times. Two researchers at the Arctic University of Norway, Magnus Ingebrigtsen and Marko Lukic, filed a detailed objection one month after publication. It was retracted on 22 April 2026. Among the problems: studies weighted wrongly, an already-retracted study included, and a true figure that should have been nearer 0.5.

Ten months between the objection and the retraction. In education technology, ten months is a full sales cycle.

And then there is the number that will not die. Benjamin Bloom's two sigma, from 1984, the one every AI tutoring company puts on slide three. The claim is that one-to-one tutoring moves a student two standard deviations. A standard deviation is a way of saying how far a group has moved compared with how spread out it already was. Two of them would be extraordinary.

The evidence was two doctoral dissertations by Bloom's own students, with cells of roughly twenty to thirty children per condition. Both ran three weeks. Both tested with instruments the researchers wrote themselves, covering only those three weeks of material. The tutoring condition also bundled in repeated quizzing with corrective feedback, which Paul von Hippel calculates was worth about 1.1 of the two sigmas on its own.

It has never replicated. Of 96 tutoring studies reviewed in 2020, none produced a two sigma effect. Von Hippel's summary, in 2024, is that "about one-third of a standard deviation seems to be the typical effect of an intense, well-designed program evaluated against broad tests."

I have had Bloom's number quoted at me in three separate product demonstrations. Nobody in any of them had read the dissertations. Neither had I, until last month.

The skill is judgment, not prompting

There is a fashionable claim that what students need to learn is prompting, which means writing good instructions for an AI system.

Prompting is useful. A vague question gets a vague answer. But it is a small skill wearing a large coat, because a student can write a beautiful instruction and then accept a terrible answer without blinking.

The durable skill is judgment, and it comes down to a short list of questions.

  • Is this claim actually true?
  • Where did the information come from?
  • Is that source real?
  • What is missing from this answer?
  • Whose point of view is not here?
  • Does the conclusion fit the evidence, or just follow it?
  • What happens if I act on this and it is wrong?
  • Can I say the answer in my own words with the screen shut?

Those are not school skills. They are the skills of every adult who has ever had to sign something.

A young employee will one day use AI to read a contract for them, write code, compare investment options or draft a message to a customer. The tool will save time. It will also, sooner or later, make a confident mistake that costs money or trust or somebody's safety.

The person who does well will not be the one who pasted the fastest answer into an email. It will be the one who knew what to check before pressing send.

Think of AI as the world's fastest summer intern. It reads a mountain overnight. It gives you a first draft before you have finished your coffee. It turns messy notes into a clean list and offers ten ideas before lunch. It also misunderstands the brief, invents a detail, misses the obvious thing and writes a fluent paragraph about a source that does not exist.

No sensible manager hands that intern the keys and leaves for the weekend. No sensible school should teach students to.

The new label on the lunchbox

On 14 August 2026, Anthropic, the company behind Claude, described a new way of marking what its models produce.

Claude models launched on or after 2 August 2026 weave a hidden, machine-readable watermark into generated text. Supported files, things like images and drawings, get signed provenance metadata attached. Provenance means a record of where something came from and what has been done to it since. The metadata follows an open standard called C2PA, short for the Coalition for Content Provenance and Authenticity, which is a long name for a shared way of stapling a history to a file. Its steering committee includes Adobe, Amazon, the BBC, Google, Meta, Microsoft, OpenAI, Sony and TikTok, so it is not a fringe effort.

It is also, worth saying, not an education initiative. Anthropic points to the European Union's AI Act and its transparency code, which took effect on 2 August 2026. Schools are downstream of a Brussels compliance deadline.

Now the part that matters, and I am going to quote Anthropic against the headlines about Anthropic.

From the company's own support page: "A detected mark provides a signal that content was processed by Claude, but is not fully conclusive." And: "Claude may not be the original author." And: "Lack of a detected mark doesn't mean the content wasn't AI-generated or processed."

From the technical write-up: the watermark "cannot distinguish 'Claude wrote this' from 'Claude heavily edited this.'" It also works badly on short passages, because a short passage does not contain enough word choices to carry a signal, and it thins out on factual writing, where there are fewer choices available in the first place.

Two more things a teacher should know before anyone builds a policy on this.

There is no public detection tool. Anthropic says an interface for checking is coming and has not given a date. As of today, a teacher cannot check anything.

And the file half of the scheme can be removed with the print screen key. C2PA says so itself, in its own security document, in a sentence I would print out and pin up: "C2PA does not offer any protection against the complete removal of C2PA manifests from assets." A screenshot strips it. So does a format conversion. So does re-saving.

Schools have seen this film before. When copy machines arrived, the promise was less paper. Anyone who worked in an office in the 1980s knows how that went. One unnecessary memo became five hundred unnecessary memos before lunch. When the internet arrived, some people expected every student to become a researcher overnight. Instead a generation learned the hard way that the first result in a search was not a wise elder. Sometimes it was a stranger with a website and vitamins to sell.

Every new information tool creates a fresh chance to confuse speed with wisdom. A watermark can show part of the road a document traveled. It cannot tell you whether the student learned anything, whether they checked a single fact, or whether they can defend one sentence of it.

What the detector years actually cost

Before anyone treats a watermark as evidence, look at what happened the last time schools trusted a machine to spot a machine.

Weixin Liang, James Zou and colleagues at Stanford took 91 essays written by non-native English speakers for a university entrance English test, and published what they found in July 2023. Every one was written before ChatGPT existed. Every one was definitely human. They ran them through seven commercial AI detectors.

The average false alarm rate was 61 percent. All seven detectors flagged eighteen of the 91 essays at once. At least one detector flagged eighty-nine of them.

American eighth-graders' essays, run through the same detectors, came back with a false alarm rate of about 5 percent.

Then the researchers did something that should have ended the industry. They fed those human essays to ChatGPT with an instruction to improve the word choices so they sounded more like a native speaker. The false alarm rate fell from 61 percent to 12 percent.

Running a human essay through a chatbot made it look more human to the detectors. The machines were not detecting machines. They were detecting a limited vocabulary and punishing it.

OpenAI, meanwhile, launched its own detector in January 2023. It correctly identified 26 percent of AI text and wrongly flagged human text 9 percent of the time. The company retired it on 20 July 2023, under six months in, with a notice that still reads "no longer available due to its low rate of accuracy." The organization that built the model could not reliably recognize the model's own output.

None of this stayed theoretical. Turnitin announced in April 2024 that it had reviewed over 200 million papers, and that more than 22 million of them scored at 20 percent AI writing or above. Its own published false positive rate for a whole document is under one percent, and it only claims that rate above the 20 percent line. So do the arithmetic on the page, and do it fairly. Under one percent of 22 million is still up to 220,000 papers wrongly marked, in the range where the company says it is at its most accurate.

Common Sense Media surveyed teenagers and parents in 2024. Twenty percent of Black teens said they had been falsely accused of using AI on an assignment, against 7 percent of white teens. The researchers were careful about why, and so am I. That gap could come from the detection tools, or from the adults reading the output, or from both. It is a real gap either way.

RAND, which is a nonprofit research organization that runs regular surveys of American schools, reported in September 2025 that half of students worry about being falsely accused of using AI.

And at Australian Catholic University, the ABC, which is Australia's public broadcaster, reported in October 2025 that the university had logged nearly 6,000 alleged academic misconduct cases in 2024, around 90 percent of them AI-related. About a quarter were dismissed after investigation. One final-year nursing student had "results withheld" printed on her transcript for six months while she was applying for graduate positions, and believes it cost her a job. A paramedic student had 84 percent of an essay flagged, submitted dozens of pages of evidence, and was still asked for handwritten notes and search histories. The university eventually stopped accepting the detector as sole evidence and dropped it entirely in March 2025.

That is the bill for building a school around catching people. Students learn to hide. Teachers learn to suspect. Everyone spends their energy on "can we prove this" and none on "did anyone learn anything."

From gotcha to proof

There is a better direction, and it is not complicated. Ask students to show their work, their choices and their understanding.

For any major assignment, a school can require a short declaration.

Tools I used:

What I used them for:

What I accepted, rejected, or changed:

Sources I checked myself:

What I can explain without the tool:

That is not busywork. It teaches the actual lesson, which is that using a tool is not shameful and hiding the process is.

It also hands the teacher a much better conversation. "You used AI to generate three counterarguments. Which one did you throw out, and why?" A student who understands their own paper answers that in about four seconds. A student who outsourced the whole thing does not.

Pair permitted AI with small proof-of-learning moments. Two minutes talking about the argument. A short in-class reflection written cold. A source-checking exercise. A handwritten outline before a long paper. A live explanation of one step in a calculation. A version history. A discussion where the student has to defend a choice they made.

This is not a return to quill pens and candlelight. It is assessment catching up with the tools students already have in their pockets.

The students, as it happens, have already asked for exactly this.

In July 2026, 98 high school juniors and seniors from all 50 states spent three days on the UMass Boston and MIT campuses and produced a proposed national framework. They called it the Students First Act and passed it 82 to 16. It bans AI on graded tests and on written and artistic assignments. It permits conditional use from ninth grade, with teacher permission, for brainstorming and editing. It requires students to cite AI use and to prove mastery.

And then it does something the adults have mostly failed to do. It writes in due process. A suspected case must go to two school officials and to a human reviewer, rather than being settled by software. The student can appeal. A teacher must personally investigate before flagging anyone. Teachers have to post their written AI policy every semester.

Jeff Riley, who used to run education for the state of Massachusetts and helped convene it, put the awkward part out loud. "The adults themselves haven't been able to come to consensus."

Read those two provisions together. The students asked for stricter rules than most districts have, and for the right not to be convicted by a piece of software. Both at once. That is a more coherent position than almost anything I have read from a school board.

The science fair test

A ninth-grader tests the water quality in a local stream.

The old model: write the report, make the poster, hand it in.

The better model: the student may use AI to build a question list, organize research notes or suggest clearer chart labels. They must say so. They must also explain how they collected the data, name the source of every important fact, and answer questions about why their results might be wrong.

Then comes presentation day, and a judge asks why they tested after rain instead of only on dry days.

The student who did the project talks about runoff and soil and how rain moves things into water that were not in the water yesterday. The student who downloaded a finished report looks at their own poster as though it were written in Finnish.

No watermark replaces that moment. A real question from a real person still reveals an enormous amount, and it takes ninety seconds.

Not every assignment needs judges and a poster board. Every student needs a few moments a term where they have to bring their own mind into the building.

The rule should fit the task

There should be places where AI is not allowed. If a teacher is checking whether a student can read, calculate, write a paragraph, remember a fact or work a problem alone, the student does it alone. That is not harsh. That is the only way anyone finds out what a child can actually do.

There should be places where it is allowed and must be named. Researching a hard topic, comparing viewpoints, revising a draft, preparing for a debate, exploring a design. Disclose the use, verify the claims, own the result.

And there should be places where it helps teachers. Different practice questions. A reading passage simplified for one student. A lesson outline. A translation for a family. Same rules apply. Check the output, protect private information, use human judgment.

I went looking for evidence that drawing the line works, and I found one study that supports it and one that does not, so here are both. Kreijkes and colleagues studied 344 English Year 10 students in 2026 and found that note-taking, on its own or alongside a chatbot, beat using the chatbot alone. Then Fischer, Rau and Rilke studied 334 German university students and found that unrestricted AI access beat restricted access by a clear margin, which is the opposite of what I expected to find.

The pen still does something. So, sometimes, does getting out of the way. Anyone who tells you the evidence points cleanly in one direction has not read enough of it.

So the rule is not "AI is good" or "AI is bad." The rule is that the tool should match the learning goal. A closed-book vocabulary quiz stays tool-free. A research project can permit AI for brainstorming and still demand cited sources and a student explanation. A writing class can use AI output as something to attack and improve rather than something to submit.

A good policy stops pretending that every task has the same purpose.

What a parent can ask at the kitchen table

You do not need a computer science degree for this. Six questions do most of the work.

  • What did you use it for?
  • Can you explain this answer without looking at the screen?
  • Which facts did you check yourself?
  • What did it get wrong or leave out?
  • Did your teacher say this kind of use was allowed?
  • If you had to teach this to someone younger, what would you say?

Those beat trying to work out whether a sentence "sounds like AI." Plenty of teenagers write formally. Plenty of adults write badly. Plenty of machine text reads perfectly naturally. Guessing from style is how you get the Stanford result, and the Stanford result was 61 percent wrong.

There is a gap worth closing at home. In the Common Sense survey, 73 percent of children said their school had told them what they may and may not use AI for. Only 51 percent had been taught anything about checking whether the answer was true. And only 56 percent said a parent had ever talked to them about using AI safely. Among the children using it every day, a third had never had that conversation with a parent.

What a school should publish

Every school and district should make its AI rules easy to find and written in a language a fourteen-year-old can read.

No student should have to guess whether making flashcards is allowed, whether grammar help counts as editing, or whether a disclosure note is expected.

Four questions. Answer them in public.

  1. When is AI allowed?
  2. When is it not allowed?
  3. When must a student say they used it?
  4. How will a student show they understood the work?

Add privacy. Students should know not to paste personal details, school records, medical information or family business into a public AI tool, and they should be told why rather than just told.

And publish it before a student makes a mistake, not after.

Some places are further along than the arguing suggests. The Associated Press reported on 21 August 2026 that 37 states have now published official AI guidance. Utah appointed Matt Winters as a full-time state AI education specialist in 2024, the first in the country, and has since trained more than 7,000 teachers, close to a third of the state's public school teaching workforce. Every Utah district must have a policy by July 2027. Maine, West Virginia and Georgia have copied the post.

Amanda Bickerstaff, who trains schools on this, described her favorite classroom demonstration to the AP. "If you ever want kids not to over trust these tools, try the map demo." You ask a chatbot for a map. It gives you one, confidently, and it is wrong, and thirty teenagers watch it be wrong. That does more for skepticism than a term of warnings.

Where this argument thins

Here is where I am on soft ground. The evidence is thin in both directions, and anyone telling you otherwise is selling something. The Stanford reviewers went through 800 papers and found no high-quality causal work on student-facing AI in American schools. The strongest harm result, the Turkish study, is one school in one country. The strongest help result is one Harvard physics course tested the same afternoon.

The much-quoted brain study rests, in its headline part, on eighteen people, is still not peer reviewed, and has a formal published critique against it. The much-quoted Nigeria result lost 36 percent of its treatment group and 50 percent of its control group before the final assessment, which is enough to make me put it down.

I am also arguing for disclosure forms and two-minute conversations without a randomized trial showing that those work either. What I have instead is that they are cheap, they are human, and they fail safely. A detector that gets it wrong takes six months off a nursing student's life. A two-minute conversation that gets it wrong takes two minutes.

Which brings me to the part that matters more than any of the above.

Waiting for better evidence is itself a decision, and it is the worst one available. The tools got materially more capable between the start of this school year and the end of it, and they will do so again. A school that holds still until the research settles will be holding still for years, while every student in the building uses these tools anyway, with no guidance, no disclosure and no practice at checking anything.

So publish the imperfect policy. Try the disclosure form for one term. Run the two-minute conversation on one assignment and see what it tells you. Some of it will be wrong, and finding out quickly is the point. A rough approach that a school can correct in November beats a perfect one that arrives in three years, and it beats no approach by a distance that is not close.

The receipt and the dinner

The point of school has not changed. It is still where a young person learns to make sense of things, work through difficulty, hold a fact, express an idea, listen to somebody else and take responsibility for a choice.

The Archivist still has to learn that knowing the words is not knowing the idea. The Spark still needs a doorway, and then has to learn to keep going after the first burst of interest wears off. The Climber still needs another explanation and another go, not a quiet message that they are behind. The Runner still needs a deeper question, not another bowl of peas.

AI can help with every one of those. It can hand the Archivist a harder question, help the Spark find a reason to care, give the Climber a patient explanation at eleven at night when no adult is awake, and show the Runner a horizon further out than the worksheet.

It can also make it easier for all four of them to skip the part that would have made them stronger.

That is the actual choice in front of schools. Not learning or technology. Better learning with the tools, or easier hiding behind them.

Anthropic's watermark is a receipt. A receipt is useful. It tells you what was bought and where. It does not tell you whether dinner was any good.

Proof of learning lives somewhere harder to scan. In the question the student asks next. In the source they decided to check. In the mistake they caught before anyone else did. In the explanation they can still give with the laptop closed.

So try one thing this term. Pick a single assignment. Let them use whatever they like, and require them to say what they used. Then ask them to talk about it for two minutes with the screen shut.

You will know inside thirty seconds. That is the whole assessment technology, and it has been available since Socrates, who managed without a detection tool.


Sources and notes

0 comments
Checking sign-in status…

No comments yet. Be the first.