Scientific foundation

The foundation behind EvaluatieScore

EvaluatieScore assesses whether a project evaluation report delivers a learning moment or remains a box-ticking exercise, across 7 touchstones derived from four authoritative, independently developed frameworks: OECD DAC, UK Gateway Gate 5, PRINCE2 and the CIPP evaluation model. In the Netherlands, unlike before a project starts, there is no mandatory instrument for evaluation afterwards — even the Netherlands Court of Audit found that the government's learning capacity on this point is still in its infancy. The rubric therefore rests on a solid normative foundation, but EvaluatieScore does not claim that a good evaluation is scientifically proven to lead to better future projects — that direct evidence is thin and, even after deeper research, remains the weakest link in this foundation. As the youngest instrument in the Dutchmind suite (€15), that modesty fits: EvaluatieScore checks against the frameworks, not against a guarantee of results.

How this assessment is grounded

The underlying body of knowledge was built through deep research across six search lines plus a targeted follow-up round, with adversarial verification through three independent checks per claim (37 claims tested, 20 confirmed). Only confirmed claims are included in this rationale; rejected or unverified material is explicitly excluded.

What this does not measure

EvaluatieScore checks the evaluation report against four normative frameworks (OECD DAC, UK Gateway Gate 5, PRINCE2, CIPP), but no validated scientific instrument exists that directly links the quality of an ex-post evaluation report to future project success — the only quantitative evidence for that link is over 25 years old and rests on a single source, so EvaluatieScore deliberately does not make that claim. The tool assesses whether the report structurally meets the frameworks (independence, benefit comparison, concrete lessons), not whether the organisation actually applies those lessons afterwards — research shows organised learning is in fact weak without a dedicated, assigned implementation process. EvaluatieScore also does not assess whether the report's author looks back with a self-serving slant: the strongest available peer-reviewed source on bias in project management places that mechanism exclusively before a project starts, not in the evaluation phase, so this is not detected in the text. Like PlanScore, EvaluatieScore has no separate, published measurement of score consistency or rank-order calibration such as RapportageScore has; every analysis does run twice by default, with a third run and the median result if the two diverge too much. As the newest instrument in the suite, EvaluatieScore is not a substitute for a mandatory, independent meta-evaluation such as the Dutch RPE requires for subsidies above €10 million or the external Gate 5 review in the UK — this product's calibration set and broader scientific foundation, as noted elsewhere on the site, are still being built out.

The touchstones and their foundation

1 Scope, purpose & evaluation questions

OECD DAC (EvalNet, revised 2019) stresses that evaluation criteria must not be applied mechanically but should be tailored to the purpose and context of the evaluation. The Dutch Regeling Periodiek Evaluatieonderzoek (RPE 2022) requires evaluation questions to match the project's phase and explicitly address both effectiveness and efficiency.

2 Benefits realisation vs. business case/plan

UK Gateway Gate 5 (Cabinet Office/IPA, now NISTA), as an independent external review, explicitly assesses whether the benefits from the business case are actually being realised and whether governance for future benefits measurement exists. In addition, the OECD sustainability criterion points to the institutional and financial capacity to sustain benefits over time, and PRINCE2 practice (End Project Report) describes a comparison of realised versus expected benefits.

3 Effectiveness & efficiency

OECD DAC lists effectiveness and efficiency among its six core evaluation criteria. The RPE 2022 requires departments to periodically assess policy and projects on both effectiveness and efficiency, each substantiated separately rather than folded into one vague conclusion.

4 Methodological quality & independence

Stufflebeam's CIPP model requires that an evaluation itself be tested against meta-evaluation standards, preferably through an independent meta-evaluation. UK Gateway Gate 5 is itself an independent, external assurance layer on top of the internal Post Implementation Review, and the RPE 2022 mandates independent experts for subsidy evaluations above €10 million — three independently developed frameworks that each name independence as a quality safeguard.

5 Performance on time, cost, quality, scope, risk

PRINCE2 practice (End Project Report and Lessons Report) describes a factual comparison of estimate versus actual on time, cost and effort, and an assessment of the effectiveness of management controls. This is a secondary source (prince2.wiki, not an official AXELOS manual) and is therefore framed as common PRINCE2 practice, not as an official requirement.

6 Lessons learned & follow-up actions

Stufflebeam's CIPP model requires final evaluation reports to explicitly state lessons, including mistakes to avoid. Williams (2007, University of Southampton ePrints) shows that lessons from individual projects are rarely incorporated into an organisation's broader policy in practice without dedicated effort and a designated process — documenting is not the same as learning.

7 Coherence, wider impact & sustainability

In its 2019 revision, OECD DAC added the coherence criterion: the compatibility of an intervention with other policy and other interventions in the same sector or organisation. The impact and sustainability criteria ask about wider, long-term effects and the institutional capacity to sustain benefits.

Key claims from the research

Every claim below has been adversarially verified: three independent checks per claim, and only what held up was included.

Four authoritative frameworks, developed independently of each other — OECD DAC, UK Gateway Gate 5, PRINCE2 and the CIPP model — converge on almost the same building blocks for a good evaluation report: testing benefits against the original promise, independent validation, pre-set criteria, and action-oriented lessons learned.

Source: OECD DAC (2019); HM Government/Cabinet Office-IPA Gate 5 (2021); PRINCE2-praktijk; Stufflebeam CIPP (2015)

The OECD DAC Evaluation Criteria (revised 2019) distinguish six criteria — relevance, coherence, effectiveness, efficiency, impact, sustainability — with an explicit warning against mechanical application.

Source: OECD DAC Network on Development Evaluation (EvalNet), Evaluation Criteria, herzien 2019 (DCD/DAC(2019)58 FINAL)

UK Gateway Gate 5 (Cabinet Office/IPA, now NISTA) is an independent, external assurance review that explicitly assesses whether the benefits from the business case are actually being realised and whether governance for future benefits measurement exists.

Source: HM Government/Cabinet Office-IPA (nu NISTA), Gate 5: Operations Review and Benefits Realisation (Assurance Portfolio Standard, V1.0, juli 2021)

Stufflebeam's CIPP model requires that an evaluation itself be tested against meta-evaluation standards (utility, feasibility, propriety, accuracy, evaluator accountability), preferably through an independent meta-evaluation.

Source: Stufflebeam, D.L. (2015). CIPP Model checklist (2e ed.)

The Elias parliamentary committee concluded that before the start of ICT projects there is insufficient consideration of the how and why; the resulting ten basic rules are exclusively ex-ante oriented — there is no comparable mandatory Dutch government instrument for ex-post evaluation after an ICT project ends.

Source: Commissie-Elias, Eindrapport 'Grip op ICT' + kabinetsreactie (2014/2015)

In 2013, the Netherlands Court of Audit stated literally that the mutual learning process based on general lessons from executed Gateway Reviews is still in its early stages, and concluded that the realisation of benefits from business cases is still insufficiently monitored.

Source: Algemene Rekenkamer (2013), Aanpak van ICT door het Rijk 2012

Williams (2007) shows that lessons from individual projects are rarely incorporated into an organisation's broader policy in practice; without a dedicated effort and a designated process, lessons are lost and mistakes are repeated.

Source: Williams, T. (2007). Post-Project Reviews to Gain Effective Lessons Learned (via University of Southampton ePrints)

The Regeling Periodiek Evaluatieonderzoek (RPE 2022) requires Dutch government departments to periodically evaluate policy and projects on effectiveness and efficiency, in cycles of 4 to 7 years, with mandatory independent expert validation for subsidy schemes above €10 million.

Source: Regeling periodiek evaluatieonderzoek 2022 (RPE), wetten.overheid.nl

Key sources

  • OECD DAC Network on Development Evaluation (EvalNet). Evaluation Criteria (herzien 2019, DCD/DAC(2019)58 FINAL). one.oecd.org
  • OECD. Evaluation criteria (institutionele samenvattingspagina). oecd.org
  • HM Government / Cabinet Office–IPA (nu NISTA). Gate 5: Operations Review and Benefits Realisation (Assurance Portfolio Standard, V1.0, juli 2021). assets.publishing.service.gov.uk
  • Stufflebeam, D. L. CIPP Evaluation Model Checklist (Second Edition). The Evaluation Center, Western Michigan University — wmich.edu
  • Commissie-Elias. Eindrapport "Grip op ICT" + kabinetsreactie (2014/2015). pianoo.nl
  • Regeling periodiek evaluatieonderzoek 2022 (RPE). wetten.overheid.nl
  • Kotnour, T. (2000). Organizational learning practices in the project management environment. International Journal of Quality & Reliability Management. researchgate.net
  • Algemene Rekenkamer (2013). Aanpak van ICT door het Rijk 2012. rekenkamer.nl + zoek.officielebekendmakingen.nl/kst-33584-2.html
  • Williams, T. (2007). Post-Project Reviews to Gain Effective Lessons Learned. eprints.soton.ac.uk
  • Flyvbjerg, B. (2021/2022). Top Ten Behavioral Biases in Project Management: An Overview. Project Management Journal 52(6), SAGE; arxiv.org/pdf/2202.00125

Based on our Body of Knowledge, v1.1 (11 juli 2026)