ERP finalists can all look capable in a demonstration. The selection risk appears when a polished workflow, a low subscription quote or one high composite score hides a control gap, a difficult conversion, an unowned interface or an implementation plan that depends on people the supplier has not committed.
This guide starts after the organisation has chosen ERP as the relevant system class and formed a credible shortlist. It does not compare ERP with accounting software or rank products. It gives a finance transformation lead one evidence standard for testing finalists and taking a recommendation to executive approval.
Quick answer
A blended score can hide control gaps, migration exposure, weak delivery ownership or unaffordable life-cycle cost unless non-compensable gates are tested first.
Decision: Recommend, condition or reject an ERP finalist for executive approval using non-compensable gates, evidence-backed weighted scoring and a reconciled life-cycle cost case.
Key takeaways
- Apply non-compensable approval gates before comparing weighted scores; a serious control, migration, security, delivery or exit failure should not disappear inside an average.
- Convert requirements into buyer-scripted scenarios with expected results and retained evidence, then cap scores when support is limited to claims, roadmap items or canned demonstrations.
- Use a 100-point structure that separates finance fit, controls, architecture, migration, implementation, vendor risk, life-cycle cost, adoption and evidence quality.
- Moderate independent scores, test weight and cost sensitivity, and record every condition in the contract or decision log before approval.
What ERP evaluation should decide
ERP evaluation is the controlled process for deciding whether a shortlisted platform, implementation team and commercial proposal can support the organisation’s required finance outcomes at an acceptable level of control, delivery risk and life-cycle cost. It is narrower than the full sourcing process and broader than a feature comparison.
SAP’s guide moves from requirements through a request for proposal, comparison and selection, and asks buyers to describe volumes, process complexity and distinctive needs. That broad ERP evaluation sequence is useful; the framework below adds finance gates, evidence caps and decision ownership.
- Recommend: all mandatory gates pass, and the adjusted score, delivery plan and life-cycle cost case are supportable. Retain the approved evaluation, residual risks and named owners.
- Recommend with conditions: no gate fails, but approval depends on specified evidence, remediation or contract language. Record the condition, owner, due date, closure evidence, remedy and closure authority.
- Do not recommend: a gate fails, material evidence is missing, or delivery or cost is not supportable. Record the failed test, unresolved risk and reconsideration trigger.
Use approval gates before weighted scoring
A weighted score is compensatory: strength in one area offsets weakness in another. That is useful for trade-offs, but unsafe for requirements that cannot be traded away. Start with a gate register. A finalist enters weighted comparison only after each gate is marked pass, conditional, fail or unverified.
| Gate | Pass condition | Owner and minimum evidence |
|---|---|---|
| G1 finance control | Critical posting, approval, access, audit, close and reconciliation scenarios pass, or an approved compensating design exists. | Controller: scripts, role output, audit records, exceptions and sign-off. |
| G2 data and migration | Required populations load, reconcile, remain accessible and fit the cutover window. | Migration owner: inventory, mappings, mock loads, reconciliations and archive plan. |
| G3 integration and architecture | Each object has one authoritative owner by lifecycle state; transfers recover, reconcile and support exit. | Systems architect: architecture, interface register, failure tests and export proof. |
| G4 delivery accountability | Supplier, partner and buyer roles, dependencies, acceptance rights and defect duties are named. | Programme sponsor: responsibility matrix, committed team, plan and acceptance schedule. |
| G5 security and supplier risk | Required access, assurance, resilience, incident, subcontractor and continuity evidence is accepted or treated. | Security and vendor risk: due diligence, assurance, incident terms and continuity plan. |
| G6 commercial and exit | Scope, pricing, data rights, renewal exposure, termination support and remedies are acceptable. | Procurement, legal and Finance: cost model, contract schedules, extraction test and exit plan. |
The gate design should reflect the organisation’s own policy, materiality and regulatory context. The GAO Green Book treats internal control as a management process supporting operations, reliable reporting and compliance. Although it is a US federal standard, its control-design concepts are a useful reference when adapted rather than copied. Legal, tax, accounting and sector-specific requirements still need the relevant specialist owner.
Turn requirements into testable finance scenarios
Requirements become decision-grade only when the evaluation team can trace each one to a business scenario, expected result and evidence. GAO’s review of a US financial-management programme cites two-way traceability from higher-level requirements to lower-level requirements and back. That principle is directly useful in selection: every scored criterion should point to its source requirement, and every critical requirement should appear in at least one test. GAO’s requirements-traceability discussion provides the source context.
- ID and owner: stable ID, accountable role, and the source policy, process issue or target-state decision.
- Scope and volume: entities, countries, currencies, ledgers, users, transactions, history and peak loads.
- Decision treatment: gate, weighted criterion or context, with the reason for any non-compensable status.
- Scenario: starting state, roles, representative data, steps, exceptions and expected result.
- Evidence and delivery: required artifact plus standard, configured, extension, third-party, custom, manual or roadmap treatment.
Use end-to-end scenarios rather than isolated questions such as “Does the system support approvals?” A close scenario might create an adjusting journal, enforce preparer and approver separation, reject a closed-period posting, preserve the before-and-after values, route an exception and show the resulting ledger and audit evidence. A source-to-pay scenario should carry the approved need, supplier, contract, purchase order, receipt and invoice references through the handoffs. The controlled source-to-pay stage model can supply the objects and exit evidence for that test.
Use the same script for every finalist
Send every finalist the same script and sample data. State what must be performed live and which evidence files must be returned. Record deviations, configuration and workarounds during the session.
Record the scenario ID, expected result, required evidence, observed result, raw score, evidence cap, gate effect and proposed condition. A failed step cannot be replaced by a slide, unrelated customer example or roadmap statement.
Build the 100-point ERP evaluation scorecard
The structure below is a Finance Circuit starting model, not an industry benchmark. Change the weights before supplier responses are opened, document the reason and freeze them for the scored round. Criteria within each domain must add to the domain weight, and all domain weights must total 100.
| Domain | Weight | Starting criterion split |
|---|---|---|
| Finance and operating requirements | 18 | Record-to-report 5; payables and procurement 3; receivables and billing 3; treasury 2; multi-entity and localisation 3; reporting 2. |
| Finance controls and auditability | 15 | Roles 4; approvals and change 3; audit records 3; close and reconciliation 3; privileged access 2. |
| Integration and target architecture | 12 | Boundaries 3; interfaces 3; transfer controls 3; extensibility and export 3. |
| Data migration and reconciliation | 12 | Source quality 3; mapping 3; repeatable loads 2; reconciliation 3; archive 1. |
| Implementation and operating readiness | 13 | Plan 3; accountable team 3; configuration and testing 3; change and support 2; cutover 2. |
| Vendor and product viability | 8 | Corporate capacity 2; product policy 2; assurance 2; delivery capacity 1; continuity 1. |
| Total cost and commercial terms | 12 | Software 3; services 3; internal effort 2; data and interfaces 2; operations 1; exit 1. |
| Adoption and change burden | 5 | Role fit 2; process change and training 2; accessibility and localisation 1. |
| Reference and evidence quality | 5 | Comparable references 2; evidence completeness 2; unresolved assumptions 1. |
| Total | 100 | Freeze the approved split before final demonstrations and supplier scoring. |
The compact table is the central comparison asset. A tab-separated appendix at the end adds working columns for the nine weighted domains, gate links, owners, evidence and conditions, ready to paste into a spreadsheet without repeating the explanatory sections.
Score capability and evidence separately
Use a 0-to-5 raw capability score, then apply an evidence cap. The adjusted criterion score is the lower of the two.
adjusted raw score = min(capability score, evidence cap)
weighted points = criterion weight × adjusted raw score ÷ 5
- 0: not met, contradicted or not tested.
- Evidence ceiling: no usable support or contradictory evidence.
- 1: uncommitted roadmap item or material manual workaround.
- Evidence ceiling: questionnaire, presentation or assertion.
- 2: material custom work, untested third party or unresolved design.
- Evidence ceiling: generic documentation or canned demonstration.
- 3: baseline met through standard capability or bounded configuration.
- Evidence ceiling: configured demonstration with returned artifacts.
- 4: requirement met with a relevant operating or control advantage.
- Evidence ceiling: buyer-scripted test with representative data.
- 5: material advantage survives exception and scale tests.
- Evidence ceiling: buyer proof plus a contractual commitment, prototype or closely matched reference.
Store the source of every score. The minimum scorecard columns are: criterion ID, weight, gate link, owner, raw score, evidence cap, adjusted score, weighted points, evidence reference, exception, cost effect and proposed condition. Keep each evaluator’s original score before moderation so the group can distinguish genuine evidence from consensus pressure.
Test finance controls with scripted evidence
Do not treat a list of security and audit features as proof that the configured finance process will be controlled. NIST SP 800-53 separates control families including access control, audit and accountability, configuration management, system acquisition and supply-chain risk. Its controls are customisable and distinguish functionality from assurance. NIST SP 800-53 Revision 5 is not an ERP scorecard, but it is a useful source for the control questions and evidence types the relevant specialists should test.
- Segregation of duties: attempt to create and approve the same supplier, journal, payment or configuration change. Retain the role design, blocked action, exception route and review output.
- Approval authority: submit below, at and above limits, then change an approved amount. Retain workflow version, identity, time, reapproval and rejection evidence.
- Audit records: change master data, workflow, posting fields and configuration, then retrieve the history. Retain before-and-after values, actor, time, object, reason and export.
- Period control: post into open, soft-closed and closed periods using permitted and prohibited roles. Retain status, blocked attempt, override and accounting result.
- Reconciliation: create a partial, duplicate and out-of-balance transfer. Retain control totals, exceptions, ageing, owner and correction evidence.
- Configuration change: change an account rule, approval route or interface mapping through the promoted path. Retain the request, approval, test, deployment and rollback record.
For reconciliation ownership, evidence and exception design, use the site’s account reconciliation control standard as a separate operating reference. The evaluation should test whether the ERP can support that control design; it should not assume the product defines the organisation’s accounting policy.
Evaluate integrations and system-of-record boundaries
An ERP can pass a functional demonstration and still fail the target architecture. Define which application may create or change each business object and status, then test the boundary rather than the connector label.
Start with the Finance Technology Stack reference architecture to assign authoritative ownership by business-object lifecycle before a finalist is scored. That prevents a product demonstration from quietly changing which system owns a customer, invoice, journal, payment or bank-confirmed state.
Then use the Finance Systems Integration Map to specify transfer IDs, timing, retries, acknowledgements, control totals, suspense ownership and reconciliation evidence for every material interface. Those records become the acceptance baseline for the ERP evaluation.
- Who owns the object? Record the system of record by lifecycle state, write rights and stable identifier.
- How does it move? Record the supported method, data contract, versioning, authentication, volume and timing.
- What proves completion? Require a transfer ID, counts, values, acknowledgement, posting status and reconciliation.
- What happens on failure? Test duplicate, timeout, partial, stale, out-of-order, replay and suspense states.
- How can the organisation exit? Prove export format, history, extraction time, cost, assistance and deletion evidence.
The UK government’s Open Standards Principles connect documented, publicly available standards with interoperability and supplier flexibility. Use them as a design reference, not a corporate mandate. Score interfaces and extractability against the buyer’s actual boundary and exit needs.
Score data migration and reconciliation risk
Migration evaluation should establish whether the buyer and delivery team can produce a controlled conversion, not whether the product has an import utility. Build the data population before scoring: objects, source systems, owners, retention rules, history, volumes, quality issues, transformations and the financial totals that must bridge source and target.
An Ofgem programme data-migration plan required test-cycle reconciliation for completeness, correctness, integrity and reliability, and required reconciliation reporting to support load validity, defect management and audit evidence. Those are useful acceptance dimensions even though the document belongs to a specific UK energy programme. The Ofgem migration plan shows the underlying reconciliation approach.
- Population: reconcile in-scope master, open, historical, document and configuration data to extraction. Evidence should name owners, counts, values, ranges and exclusions.
- Mapping: map entities, accounts, dimensions, parties, tax, currency, status and identifiers. Retain versioned rules, rejects, defaults and approval.
- Mock conversion: repeat representative loads with defects, volume and timing. Retain run ID, input, duration, rejects and regression results.
- Financial reconciliation: bridge subledger, open-item, ledger and balance populations. Retain counts, debits, credits, balances, exceptions and tolerances.
- Cutover and archive: prove freeze, delta, final reconciliation, rollback and retained history. Retain the runbook, authority, timing, rollback result and archive search.
A failed opening-balance reconciliation, an unowned data exception or an untested cutover window should affect the migration gate, not just remove a few weighted points. Tolerances must be approved for the specific object and reporting consequence; there is no universal acceptable percentage.
Assess implementation capacity and delivery ownership
Evaluate the people and delivery model presented for the engagement, not only the supplier’s general methodology. Record who will design, configure, convert, test, train, approve, cut over and repair defects. Separate the software vendor, implementation partner and buyer because shared delivery can otherwise leave acceptance and warranty boundaries unclear.
NIST SP 800-53A Revision 5 provides assessment procedures for security and privacy controls across system-development phases and guidance for analysing results. Its scope is security and privacy, but the evidence principle applies here: accept a configured control from test results, not from a design statement alone.
- Named team: lead roles, allocation, location, relevant delivery record, substitution rights and interview access.
- Solution boundary: standard, configured, extended and custom components, with owners, upgrade effects and acceptance evidence.
- Buyer capacity: period-by-period finance, data, testing, control, change, support and backfill effort.
- Test plan: entry and exit criteria for integration, conversion, security, user, performance, regression and cutover tests.
- Defect model: severity, response, correction, warranty, paid-change boundary and commercial remedy.
Evaluate vendor viability, roadmap and support
Vendor viability is not a single credit score. It combines the supplier’s capacity to support the product, the product’s direction, the delivery ecosystem, security and assurance, concentration dependencies and the buyer’s ability to continue or exit if conditions change.
NIST SP 1305 recommends communicating supplier cybersecurity requirements. NIST SP 1326 adds research on ownership, provenance, resilience, cyber practices and supply-chain tiers. These are cyber sources, not complete corporate-viability tests, but they prevent assurance from becoming a checkbox.
- Corporate capacity: ownership, available financial evidence, insurance, material disputes and continuity arrangements.
- Product direction: supported versions, release and end-of-support policy, deprecations and roadmap. Score current capability separately from future commitments.
- Security and assurance: applicable reports, scope, exceptions, subprocessors, recovery evidence and customer duties.
- Delivery ecosystem: the committed partner, named skills, escalation, subcontractors and quality oversight.
- Continuity and exit: data export, documentation, transition support and service continuity terms.
Build a reconciled life-cycle cost model
Compare vendors on the same scope, volumes, environments, implementation boundary, contract term, currency, tax basis, inflation treatment and internal resource assumptions. A subscription line is not total cost.
The GAO Cost Estimating and Assessment Guide calls for a defined purpose and scope, technical baseline, work breakdown structure, assumptions, data, estimating methods, sensitivity and risk analysis, documented results and updates with actual costs. The guide is written for programme estimates, but those disciplines fit an ERP cost comparison well.
- Software: users, modules, entities, transactions, environments, storage and integration services. Test growth tiers, premium APIs, retention and price increases.
- Implementation: design, configuration, extensions, data, interfaces, reports, testing, training and cutover. Test change orders, subcontractors, defects and excluded work.
- Internal capacity: Finance, technology, data, controls, change, testing, backfill and overtime. Test business-owner time and temporary close support.
- Transition: legacy licences, parallel running, archive, decommissioning and termination. Test delayed shutdown, access costs and stranded commitments.
- Ongoing operation: administration, releases, regression, support, assurance and enhancement. Test mandatory upgrades, support tiers and custom maintenance.
- Exit: extraction, documentation, transition help, termination and deletion. Test format limits, volume charges and missing history.
Model a base case and at least one downside case. Test the assumptions most likely to reverse the result: implementation effort, internal backfill, migration cycles, interface count, custom work, user growth, price escalation, parallel running and exit timing. Show both undiscounted cash by year and the organisation’s approved present-value view where that is part of its investment policy.
The UK government’s Digital, Data and Technology Playbook says contracts should require an exit plan joining the outgoing supplier’s exit with the mobilisation of the replacement or in-house provision, and that exit preparation should be reviewed before contract end. The playbook’s exit guidance is public-sector guidance, but the commercial question is useful for any long-lived ERP commitment.
Run reference checks as evidence, not testimonials
Reference calls should test claims that remain decision-relevant after demonstrations and due diligence. Ask for organisations with comparable modules, legal-entity complexity, geography, transaction profile, implementation partner and time since go-live. Add independently identified users where policy and access permit.
- What scope, entities, modules and interfaces went live? Test comparability.
- What changed between proposal, design and delivery? Test fit and change-order exposure.
- What failed during mock conversion or reconciliation? Test migration effort and repeatability.
- Which controls or reports needed extensions or manual work? Test whether standard claims survived detailed design.
- Was the named team retained through cutover? Test staffing reliability.
- Which costs and internal roles exceeded the case? Test downside assumptions.
- What would you contract or govern differently? Test conditions for the current decision.
For each answer, record corroborated, contradicted, not comparable or unresolved. A positive reference should not lift a score unless it addresses the same claim, scope and delivery model. A contradictory answer should reopen the related criterion, cost assumption or approval gate.
Govern the final recommendation and contract conditions
Complete scoring independently before moderation. The moderation chair should show the evidence for material differences, record every changed score and retain the original submissions. Evaluation participants should disclose conflicts, gifts, prior vendor relationships and advisory roles under the organisation’s policy.
- Gate register: status, evidence, owner, exception, condition and closure authority for G1 to G6.
- Scorecard: frozen weights, evaluator scores, evidence caps, moderation changes, weighted total and unresolved items.
- Cost case: common scope, baseline, assumptions, base and downside cash flows, sensitivity and approval limit.
- Delivery case: named teams, responsibilities, schedule, dependencies, acceptance plan, cutover authority and defect model.
- Due diligence and references: supplier findings, assurance limits, matched references, contradictions and required treatments.
- Contract conditions: commitment, owner, due date, evidence, remedy, price treatment and the right to stop or terminate.
Stress-test whether the ranking is stable
Recalculate the result under disclosed alternatives before recommending a winner. Increase and decrease the largest material weights within a range approved by the steering group, remove disputed criteria, apply the downside cost case and test whether a conditional gate closes as assumed. If small, plausible changes reverse the ranking, report the finalists as economically or operationally close rather than presenting false precision.
A policy can require all gates to pass, a minimum adjusted score and an affordable downside case for an unconditional recommendation. Conditional approval needs a named owner, date, evidence and remedy. Set thresholds before scoring, and retain controller, security, data, technology, procurement and legal authority over their own risks.
Do not let contract signature become the first time Finance sees the final scope. Reconcile the selected configuration, implementation statement of work, cost model, control conditions, data rights and exit obligations to the version that was scored. Any material variance returns to the accountable criterion and gate owner before approval proceeds.
Copyable ERP evaluation scorecard appendix
Copy the tab-separated block into a spreadsheet. Add criterion rows when the approved split needs more detail, but keep each domain total unchanged. A failed mandatory gate remains separate from the weighted result.
Use adjusted score = MIN(raw score, evidence cap) and weighted points = weight × adjusted score ÷ 5. Point each evidence reference to a returned artifact, test result, contract clause or reference-call note.
Copy the scorecard starter
ID Domain Weight Gate link Gate status Owner Raw score (0-5) Evidence cap (0-5) Adjusted score Weighted points Evidence reference Exception or condition
D1 Finance and operating requirements 18 G1 Finance transformation
D2 Finance controls and auditability 15 G1/G5 Controller
D3 Integration and target architecture 12 G3/G6 Systems architect
D4 Data migration and reconciliation 12 G2 Migration owner
D5 Implementation and operating readiness 13 G4 Programme sponsor
D6 Vendor and product viability 8 G4/G5 Vendor risk
D7 Total cost and commercial terms 12 G6 Finance and procurement
D8 Adoption and change burden 5 Change lead
D9 Reference and evidence quality 5 Evaluation chair The nine domain weights total 100. Adapt the split before supplier responses are opened, then freeze the approved version for every finalist.