If you give the exact same test to two comparable classes of students, once digitally and once on paper, it's likely you'll see measurably different results. This has nothing to do with what the students know. It's a property of the medium itself.
This phenomenon is called the Mode effect. The mode effect is the systematic difference in scores caused purely by how a test is delivered - paper, screen, or spoken - rather than by what the student knows.
A research-backed difference
This difference has been measured in a large body of research over the last four decades.
The easiest way to quantify the exact difference, is by looking at international standardised testing, like the Programme for International Student Assessment (PISA), which looks at various academic skills of 15-year-old students across around 90 countries.
In 2015, the PISA test moved from paper-based assessment to online assessment. Before making that switch, the OECD ran in which pupils in the same schools were randomly assigned to sit the test either on paper or on screen. In all three countries studied, every subject came out lower on screen:
A similar thing can be seen in the for the years 2010-2020, which shows a leaning towards students scoring better on paper tests. Across 42 studies, 21 found paper easier, 17 found no meaningful difference, and only 4 found students doing better online:
Why it happens
When we learn, we learn in a specific context: we remember something we've handwritten more easily when handwriting it a second time, versus using a different medium - speaking the text out loud, or typing it on a keyboard.
In addition, there are four specific mechanisms we see when comparing paper tests and digital tests:
- A different spatial awareness. Our brains are better in remembering where something was located in a physical book with pages, rather than scrolling and clicking through screens.
- Typing fluency is generally lower in students compared to handwriting.
- Differential speededness: screens make some students slower, so a timed test penalises them twice.
- Simple familiarity with the interface: most students only get to see a particular digital testing interface once they start taking the test.
Interestingly, there doesn't seem to be a difference here between "digital natives" and older generations of students; younger cohorts show no advantage when taking digital assessments, and the paper advantage measured in reading studies has actually grown between 2000 and 2017, not shrunk.
The penalty isn't constant, either. A found it concentrated almost entirely in informational text (for narrative text it disappears altogether) and roughly three times larger when students are reading against a clock:
What it means for you
Unlike the examples above from organisations like PISA, you're probably not comparing tests taken across different modes with each other. Everyone in your class sits the same test in the same mode, so the mode effect moves the whole group together, and a shift that applies to everyone equally is one you can absorb when you set your grade boundaries.
The catch is that it doesn't apply to everyone equally. The mechanisms above land harder on some students than on others: the slow typist, the student who finds reading on screen tiring, the one who has never seen the interface before. Grading relative to the class average corrects the group shift, but you might also want to minimize the spread of scores.
There's a few things you can do to compensate for it:
- Let students practise in the mode they'll be tested in: if you're assessing digitally, see if you can also let students practice digitally, like with Examplary's Practice Spaces.
- Vary modes across the term: combine digital and written assessments, so that any imbalance between students performing better in one mode than the other is cancelled out. If you're using Examplary, you'll review and grade answers the same way, no matter if a student took the assessment online or on paper.
- Be careful with time limits on the screen: this is exactly where the penalty is largest, so apply time limits in digital assessments carefully, and make sure students who work more slowly on screen aren't disadvantaged twice.
- Watch the reading load: long informational passages are where screens cost the most. If a question hangs off a page of text, consider shortening the stem.
Conclusion
Does that mean we should avoid using digital assessment? No. There are real valid reasons to test online, including security, traceability, consistency and efficiency.
Moreover, all of the other modes also have an effect, there is no 'neutral' delivery method. On paper, digital, oral: all of these options are valid, as long as we are aware of the fact that the way we deliver tests has an effect on outcomes.
