← Back to Blog

Reimagining Homework in the Age of AI by Prioritizing Engagement With the Material

• By John Britton

A recent study of students using AI in China found that they completed homework faster and scored higher on it, while their later exam performance fell. My hypothesis is that the 30% drop in time spent engaging with the material is central, and that if the time is approximately equal and AI is used, both the quality of what people produce and their ability with or without AI could improve. Schools and workplaces can adapt by setting expectations for engagement, using a tailored process rubric, and then bringing whatever paper, PowerPoint, Word document, image, website, or other output comes from that work into a second stage of live discussion.

The recent study of 26,811 students in China found that after students adopted generative AI, their homework scores increased 18%, their completion time fell 30%, and their monthly closed-book exam scores fell 20%. The losses were concentrated among the roughly 80% of AI users whose very short completion times and high homework scores looked like outsourcing. Students who maintained completion times similar to students who did not use AI had only small losses.

My actual hypothesis is that if the time spent with the material is equal and AI is used, then everything goes up. The quality of what the person produces goes up because AI can raise that level of quality, and their ability with or without AI also goes up. When AI compresses two hours of reading, searching, deciding, writing, and revising into ten minutes, that engagement disappears.

A newer randomized study of undergraduates is closer to what I am actually talking about. Total learning time stayed approximately equal. Students with AI scored higher on unaided knowledge tests immediately afterward and a week later, while shifting time away from drafting and toward reading and searching.

I don’t think it totally doesn’t matter how someone uses AI. But once someone is in the ballpark of engagement, I think most of that is by degree, so it doesn’t matter as much. The first thing to ask is whether AI ended the engagement or helped the person spend that time differently.

Engagement as the Assignment

A teacher could estimate approximately how much time students would normally spend on a project and then create an AI equivalent of that engagement. The student could read, ask questions, request summaries, find details they are interested in, have AI quiz them, apply the material somewhere else, and check whether they understand the pieces they think they understand.

There are a million easy ways to do this without building a perfectly designed AI tutor. The student could prompt AI to be a debater that always takes the other point of view. AI could insert incorrect facts on purpose and have the student keep looking until they find all of them or the engagement time ends. It could ask for examples, present a competing explanation, or require the student to make decisions as the work develops.

The nice thing is that this could all happen in a chatbot window. The teacher, supervisor, or individual grading themselves could create a rubric for whatever three or five process-related things should occur. The criteria might include the questions asked, decisions made, attempts to test an idea, use of evidence, corrections, or application to a new situation. The rubric can be tailored to the class or assignment, and the amount of time invested in creating it can depend on how substantial the assignment is.

AI could then report total time, turn-taking, questions, changes in direction, the main points addressed, and how much it did without interruption, prompting, or human decision-making. It could rate the person’s engagement in the turns and the person’s contribution to the final paper, PowerPoint, Word document, image, or website. Those are two separate things.

What This Looked Like for Case Conceptualization

I did this recently with case conceptualization. I started with de-identified general symptoms and the diagnoses I was thinking about, then looked at all the differential diagnoses and talked through Occam’s razor. What single questions would rule each diagnosis in or out?

I spent about half an hour looking through the DSM, connecting symptoms with the history, putting those factors in, and discussing the case conceptualization. At the end, I had a comprehensive document listing all of that, but the document was reflective of the engagement I had already put in.

I could take that document to a supervisor, talk through the options, and answer questions that were not predetermined. The AI rated the engagement as 75% me and 25% AI, and the final document as 55% me and 45% AI. The AI wrote the actual document, but I changed the order and categorization, worked through some of the questions, chose the diagnosis, and supplied much of the wording through my earlier prompts.

Impromptu Questions and Discussion

The output could be a paper, PowerPoint, Word document, image, website, or something else. I don't care what it is. In most domains, AI can already create artifacts of higher quality in less time. With continual AI progress, I don't think it will be many more months before almost all of them can be created more easily by AI. These outputs used to be the main thing students were judged on for a grade. At least when the goal is learning, engagement should now get priority, even if our institutions and past grading habits don't agree.

Someone might spend thirty seconds creating an image that represents the ideas discussed or thirty minutes organizing a paper. If the total engagement time was four hours, all of that goes toward it.

The next stage happens in class, supervision, or a work discussion. People can bring in what they created, but presenting a PowerPoint or following another prewritten script does not demonstrate much by itself. Taking impromptu questions, answering them, explaining decisions, comparing what they found, applying the material, and engaging in discussion demonstrates whether they understand it.

AI capabilities will continue to develop, so let’s develop strategies that both take advantage of and are immune to continued progress.

AI contribution note: For this article, Codex estimated that the ideas and argument were 90% mine and 10% AI, while the final written output was 85% mine and 15% AI. AI helped find and check the studies, organize my spoken prompts, and lightly edit wording and structure.

Subscribe for future posts

If you want new writing at the intersection of AI and psychology, ethics, and implementation of AI in clinical practice, subscribe on Substack.

Subscribe on Substack

The views expressed here are my own and do not necessarily reflect the views of any current or future employer, training site, academic institution, or affiliated organization.