The Wharton School and Penn Engineering: AI and math learning

High school students who practiced math with an unguided chatbot did worse on the exam once it was taken away.

Authors
Bastani et al., The Wharton School and Penn Engineering
Published
PNAS, 2025
Peer review
Peer-reviewed
Participants
Nearly 1,000 students in grades 9 to 11, in Turkey
Method
Field experiment in real math classes
Groups
Standard ChatGPT ("GPT Base"), a guardrailed tutor ("GPT Tutor"), or no AI

What they did

Researchers from the Wharton School and Penn Engineering ran a field experiment with nearly 1,000 high school students (grades 9 to 11) in Turkey, during real math classes. For practice sessions, some students got a standard ChatGPT-style assistant ("GPT Base"), some got a version with guardrails built to guide them rather than hand over answers ("GPT Tutor"), and some got no AI. Then everyone took an exam without AI.

What they found

On the exam taken without AI, students who had practiced with the unguided assistant scored 17% lower than students who never had AI. The guardrailed tutor version avoided that drop.

“Students relying on the technology may underperform when access to AI is subsequently removed, indicating reduced skill acquisition.”

Bastani et al., PNAS 2025, significance statement

“GPT Base diminished the average control student's performance on the unassisted exam by 17%.”

Bastani et al., PNAS 2025, results

What it doesn't show

  • The effect is about unrestricted chatbot help. A tutor version with guardrails did not show the drop, so the design of the tool matters a great deal.
  • It covers high school math in one country. Other subjects and ages may differ.
  • Of the four studies here, this is the strongest evidence of cause and effect, because students were randomly assigned in real classes.

Read the paper

Read the paper in PNAS“Generative AI without guardrails can harm learning: Evidence from high school mathematics.”

How MindMargin relates

This study is the closest to what MindMargin is built around: the same AI hurt learning when it did the work and did not when it made students do the thinking. MindMargin never gives you the answer to what you are studying. It asks, and you answer from memory.

All four studies, and what they add up to