Education, Testing & Data
A didactic test result is produced by the combination of a student’s knowledge, the selection of items, scoring, and the pass threshold. A year-on-year difference in the failure rate therefore does not by itself prove that one generation was better at mathematics than another.
1. What Can Be Substantiated
CERMAT publishes test papers, correct-answer keys, and assessment criteria. The 2025 mathematics didactic test had a maximum score of 50 points and a pass threshold of 33 percent, meaning at least 17 points.[1][2]
CERMAT’s data portal publishes the number of test takers, failure rate, mean score, and item-level results for individual tests. It also defines the group used for year-on-year comparisons.[3]
Even an identical percentage threshold does not by itself ensure that two tests are equally difficult. Such a conclusion would require a comparison of test design, individual items, and the population being studied.[1][3]
2. How to Read the Claims in Context
A change in the failure rate can have several causes: a different mix of students, different item difficulty, a different distribution of points, or a genuine change in preparedness. Without separating these effects, a story about a “better” or “worse” generation is stronger than the data support.
A controversy over a particular item must be assessed against the question and the published answer key. It is not valid to generalize a disputed interpretation of one item to the entire test without evidence of its impact on the results.
Transparency increases the scope for independent scrutiny. Publishing a table is not enough, however; an analyst must describe precisely who is being compared and what uncertainty remains.
- Test papers, answer keys, and aggregate statistics are available for each year.
- The pass threshold forms part of the criteria published in advance.
- The year-on-year failure rate is not in itself a causal explanation.
- How much of the difference between years is due to test difficulty and how much to the composition of the student population.
- How the same student would perform in every year being compared.
- Whether a particular disputed item changed the result for a significant share of the population.
3. Five Questions to Ask
- Are we comparing first-time candidates taking a compulsory examination?
- Were the threshold and point scale the same in both years?
- Which items produced the largest difference?
- Are both the answer key and the rule for accepting alternative solutions published?
- Which claims are descriptive, and which already require causal analysis?
4. Conclusion
An examination figure is the result of a measurement process, not a direct snapshot of a generation’s intelligence. The stronger the conclusion we wish to draw, the more precisely we must describe the test, the population, and comparability across years.
— Jiný Kontext
