Why pre-/post-tests are failing your training evaluations ... and what to do instead
- Zahra Miller
- Jul 21
- 4 min read
Updated: Jul 28
I want to tell you about a project that humbled me — not because the training was poor, but because the evidence we were required to produce could not reflect what the training actually did.

We were delivering training as a core intervention in a capacity-building project for a network of organisations. The results framework tied two outcome-level indicators directly to pre- and post-test scores: the percentage of participants demonstrating knowledge gain and the average score improvement across cohorts. The logic seemed clean: measure what people know, teach them, measure again, report the difference.
We Had Richer Data. It Just Did Not Count.
The sessions themselves were well designed. A mid-module wheel-spin activity had small teams classifying real results as outputs, outcomes, or impacts in real time — applied thinking under mild time pressure, with no facilitator scaffolding — and an end-of-session quiz game added another comprehension check. A comprehensive workshop evaluation captured the relevance of each subtopic to participants' roles, self-rated confidence against every learning objective, the practicality of shared tools, and willingness to recommend the session: granular, role-specific data with high response rates.
That data told a credible and consistent story. The pre- and post-test scores did not. Some sessions showed modest gains; in others, participants scored worse after well-facilitated workshops with application exercises. We even shared correct answers immediately after each pre-test, expecting transparency to sharpen performance. It didn't move the needle in any consistent way.
Why the Scores Made No Sense...
The instrument was fighting how learning actually works. Knowledge does not consolidate in a single session: research on the forgetting curve (Ebbinghaus, 1885; Murre & Dros, 2015) shows that without reinforcement, people lose up to 67% of newly learned information within 24 hours. There is a deeper problem than retention, though. Much of what training delivers is not new information filling an empty space — it contradicts something participants already believe or habitually practice. And changing a held belief is not a learning event; it is a slow, effortful process of cognitive restructuring. Learners must first feel dissatisfied with their existing understanding before a new idea can take hold (Posner et al., 1982), and of the seven documented responses to contradictory information, only one is a genuine change in belief — the other six protect the prior understanding (Chinn & Brewer, 1993). A post-test score cannot tell you which one occurred.
The numbers were also structurally fragile. With small cohorts and a 12-question instrument — capped to avoid overburdening participants — each question carried roughly 8 percentage points, so one person's bad morning could move the group average by 15 to 20 points: alarming on a dashboard, meaningless as a signal of training quality. Matched pairs also eroded when people arrived late or left early, sometimes resulting in a handful of unusable data points. Add end-of-day fatigue, and the instrument was measuring the environment as much as the learning.
Finally, a hard truth we must accept as facilitators: not everyone in the room chose to be there. Motivation is a primary driver of whether understanding genuinely shifts (Limón, 2001), yet pre- and post-tests treat all respondents as equally motivated learners — an assumption that quietly corrupts the data before a single question is answered.
This Is Not a Tool Problem. It Is a Results Framework Problem.
The in-session evidence was doing what the post-test could not. Application activities surfaced misunderstandings in the moment, while I could still respond with clarification. Confidence ratings have no correct answer to game; they reveal how participants feel about their ability to perform — a more honest, more actionable signal. Together, this data showed which learning objectives had likely shifted and which modules needed refinement. Yet the pre- and post-test scores were the outcome metrics.
Kirkpatrick's Four-Level Model makes the gap explicit. External facilitators can produce excellent Level 1 evidence — reaction, relevance, confidence, engagement. But Levels 2 through 4 require delayed knowledge checks, observed on-the-job behaviour, and sustained organisational access that facilitation contracts almost never include. In practice, external facilitators operate almost entirely at Level 1, yet are asked to produce Level 2 outcome evidence using an instrument that cannot deliver it.
What to Do Instead
The tools do not need to be invented — they need to be elevated from supplementary feedback to recognised indicators.
Treat observed practice during application exercises as direct evidence of conceptual understanding, not a proxy for it.
Bring the workspace into the room: a programme officer applying a new framework to their actual project produces observable evidence no multiple-choice question can. Include checks on whether participants felt that examples given in the session directly map to capacity for routine tasks in their workflow.
Use self-rated confidence mapped to learning objectives — with the caveat that novices often overrate themselves (Kruger & Dunning, 1999), so confidence complements evidence of application rather than replacing it.
Close with structured reflection tied to a named task or real deadline, and consider a small post-workshop challenge with committed individual feedback for post-training support: uptake may be low, but any uptake may yield more meaningful evidence of transfer than a full-cohort post-test at the end of a tired day.
The Conversation That Has to Happen Earlier...
The most important conclusion is the most direct: the measurement of learning cannot reasonably be placed on an external trainer, nor can it be expected to occur within a workshop setting. Where training serves a coalition of organisations, the access problem is structural — no coalition can mandate post-workshop follow-through in organisations it does not manage. What results frameworks should honestly claim is robust in-session evidence, done well. Claiming more without the architecture to back it up does not strengthen a programme's credibility; it undermines it.
Hinging a project's success on a tool that cannot distinguish genuine understanding from end-of-day noise is not a measurement decision. It is a risk — and an avoidable one, if credible learning evidence is valued from the start.
Ready to Move Beyond the Inadequate Pre-/Post-Test?
At EvaluCore, we will work with you to design training evaluation approaches that recognise how learning actually works, or other appropriate measurement frameworks across your programme that produce evidence capable of standing up to scrutiny.




This was a really insightful read. As I was reading, I found myself reflecting on all the pre- and post-tests I’ve either conducted or participated in as a trainee in the past. Motivation is 100% a factor that is often overlooked, and it makes complete sense to observe the learning that happens during a workshop or training itself. That learning is meaningful and should count. In many ways, it also creates less pressure to "perform" when the expectation of demonstrating increased knowledge immediately after the training is removed. The final takeaway resonated with me most: the measurement of learning cannot reasonably be placed solely on an external trainer, nor can it be expected to occur within the confines of a…