By
Logiks Lab
Published on
August 9, 2026
Updated on
August 14, 2026

Buying an AI solution in 2026: due diligence on data, models, costs and exit strategy

Structure AI supplier due diligence around evidence, responsibilities, total cost and a realistic exit strategy.

A decision-maker assessing several technology options in a restrained setting.
Type
Practical guide
Level
Intermediate
Reading time
17
Progress0 %

“Buying an AI solution” should lead to evidence, not merely deployment: the expected effect must be measurable and reversible.
Define “purpose and regulatory role”, control the “contractual terms for input data and retention”, then decide against an explicit baseline.

Key figures

Figure What it establishes Source, date and scope What it means for you
4 functions The NIST AI RMF organises AI risk management around Govern, Map, Measure and Manage. NIST — AI Risk Management Framework, updated in 2026, AI systems and services An assessment must cover deployment conditions, monitoring and documentation
6 months The AI Act requires deployers of certain high-risk systems to retain logs under their control for at least six months. European Commission — AI Act Article 26, official text, accessed 11 July 2026, deployers of high-risk AI systems in the European Union Traceability must be designed into operations, not reconstructed when an incident occurs
5 human capabilities Article 14 provides that human oversight should enable people to understand, monitor and interpret the system, disregard or override its output, and interrupt it. European Commission — AI Act Article 14, official text, accessed 11 July 2026, high-risk AI systems A human in the loop is useful only if they have information, competence, authority and a genuine means of stopping the system
2 phases before commitment GOV.UK calls for discovery followed by alpha before committing to an off-the-shelf product. GOV.UK — Commercial off-the-shelf products, updated 4 July 2025, procurement of digital products and services Tool selection should follow an understanding of the problem and trials of the available options
3 exit assets The DDaT Playbook stresses technology-neutral requirements, clarified intellectual property and maintained documentation to reduce supplier lock-in. GOV.UK — Digital, Data and Technology Playbook, accessed 11 July 2026, digital procurement and contracts Reversibility must be negotiated before the contract and tested during the relationship

These reference points set boundaries for decisions about AI supplier due diligence; they do not make the decision for you. A published figure describes a specific scope, date and sometimes a population different from your own. Treat it as a constraint to test, not as a promise of automatic impact. The timetable supports the evidence.

After production begins, the comparison retains the previous state. For this topic, the first source leads to the following operational interpretation: “An assessment must cover deployment conditions, monitoring and documentation.” The second reference in the table must likewise be tested against your scope and a local measurement. Distinguishing an external reference from local evidence protects the analysis from easy extrapolation.

How to read sources without overinterpreting them

At the next milestone, stopping must remain an option: a source is useful when readers can understand at the same time what it states, the scope it covers and the limits of extrapolation. The five reference points below must therefore be read as decision boundaries, never as causal promises.

Within the scope of “AI supplier due diligence”, external data should inform a decision only when its scope, date, unit and limitations are explicit. The review must separate what the source establishes, what the team infers from it and what a local test still needs to demonstrate.

In practice, the evidence record retains the organisation, title, URL, access date, population, unit, method and interpretive caveat. It then records the decision informed by the reference point and the local observation capable of contradicting it. In this case, link that record to “purpose and regulatory role” and assign its review to “Business functions”. Data without a documented owner ages silently; data with a revision condition remains controllable and can be cited without losing its context.

Reference point 1

The “4 functions” framework published by NIST — AI Risk Management Framework applies to “AI systems and services”. It helps formulate a testable hypothesis without turning an external figure into an automatic target. Exit planning starts early.

Reference point 2

European Commission — AI Act Article 26 documents “6 months”. Its exact scope appears in the preceding table; retain it when comparing the figure with your own operations, populations and time periods. This evidence is local.

Reference point 3

European Commission — AI Act Article 14 gives “5 human capabilities”. This information sheds light on a choice; it does not, by itself, prove that the same effect will appear in your context. Reversibility is decisive.

Reference point 4

GOV.UK — Commercial off-the-shelf products sets out “2 phases before commitment”. Before using that figure to make a decision, check the date, the population covered and whether the measurement can be reproduced locally. The test must withstand scrutiny.

Reference point 5

GOV.UK — Digital, Data and Technology Playbook places the “3 exit assets” boundary within “digital procurement and contracts”. It provides an external reference for the diagnosis; it replaces neither a local baseline nor an analysis of exceptions. This reference point does not decide.

Reusable citation card

Between reviews, changes are versioned: a robust citation should be reusable without losing its author, date, scope or limitation. The card below separates those elements and connects them to a precise decision, preventing a correct figure from becoming misleading once removed from context.

Field Content to retain
Verifiable claim The NIST AI RMF organises AI risk management around Govern, Map, Measure and Manage.
Attribution NIST — AI Risk Management Framework, updated in 2026
Declared scope AI systems and services
Value or boundary 4 functions
Operational interpretation An assessment must cover deployment conditions, monitoring and documentation.
Decision concerned Link “purpose and regulatory role” to a local observation before arbitration
Review owner Business functions — Avoid digitising friction that has not been questioned
Revision condition Re-examine the citation if the source, scope or “portability of prompts, data and logs” changes

Introduction: define the primary risk

The topic appears technical until the first disputed decision. Yet “purpose and regulatory role”, “contractual terms for input data and retention”, “independent edge-case evaluation” and “portability of prompts, data and logs” all belong to the same decision path.

The practical risk is a purchase based on a demonstration that does not use representative data. Neither enabling another option nor adding another dashboard will fix it; the issue requires a defined scope, an accountable owner and evidence that can withstand challenge.

Our position is therefore clear: the system has value only if its claimed effect can be observed. Compare the situation before the change with the same segments after the test. Context requires evidence.

Stakeholders and responsibilities

Stakeholder Responsibility in the decision Point to watch
Business functions Describe the real work, exceptions and value Avoid digitising friction that has not been questioned
AI and data team Designs data, evaluations, models and observability Measure the complete task and failure cases
IT and security Manages identities, tools, risks and continuity Limit scopes, secrets and irreversible actions
Model providers Supply capabilities, limitations and updates Monitor costs, versions, retention and dependency

This allocation avoids confusing execution with accountability. The “Business functions” carry the first operational responsibility; the “AI and data team” provides an independent control. The decision is defensible only if every stakeholder knows what they measure, what they authorise and what they restore when the accepted limit is crossed. The answer depends on the cycle.

Definition: AI supplier due diligence

In this guide, “AI supplier due diligence” combines “purpose and regulatory role”, “contractual terms for input data and retention”, “independent edge-case evaluation” and “portability of prompts, data and logs”. Its aim is to reach a sustainable choice with known limitations and responsibilities; the decision rests on performance and cost measured against a buyer-controlled test set.

The definition is therefore operational: it names the components, desired effect, metric and limitation. Readers can quote it without reconstructing its meaning from the rest of the page. Exceptions reveal maturity.

Why this issue is becoming central

The sources converge on three boundaries—4 functions, 6 months and 5 human capabilities. They do not describe a universal average; they specify thresholds, obligations or operating conditions. Here, the third source leads to the following operational interpretation: “A human in the loop is useful only if they have information, competence, authority and a genuine means of stopping the system.”

This reading turns figures into decision questions: what scope do they cover, what uncertainty remains and who can act when the measurement moves outside the accepted threshold? In AI supplier due diligence, that accountability determines whether the intended effect can be achieved. The risk is tangible.

Comparing four levels of commitment

Level What it optimises Decision criterion Limitation to make visible
Observation without a baseline Apparent speed Purpose and regulatory role The outcome cannot be attributed
Bounded pilot Learning on one workflow Difference from the baseline The tested case may remain too simple
Governed deployment Demonstrated impact across the useful scope “Independent edge-case evaluation” and “portability of prompts, data and logs” controls Recurring cost must remain explicit
Reduction or shutdown Control of the primary risk Documented exit threshold Preserve data, evidence and reversibility

For AI supplier due diligence, this comparison does not identify a universal winner. It makes visible the cost of missing evidence, an overly simple pilot or premature expansion. The right level depends on how critical the workflow is, the quality of the “contractual terms for input data and retention” and the practical ability to recover “portability of prompts, data and logs”. The threshold remains explicit.

Applied to AI supplier due diligence, the following method draws on public and operational good practice. It is not presented as a proprietary Logiks method: its value lies in the order of the controls and in a third party's ability to verify every deliverable.

1. Frame the decision

First describe the expected outcome and connect it to “purpose and regulatory role”. Use neither an ideal demonstration nor a global average; examine the actual decision that remains open and the value that justifies it. The useful deliverable is a framing note that names the decision, limitation and accountable owner.

2. Measure the starting point

If the measurement diverges, the assumptions must remain reviewable: at this stage, observe the decision metric before making any change. Involve the person who handles exceptions, then compare the outcome against the starting position and its variation across segments. You should be able to give a decision-maker who was not involved in the project an initial, dated measurement broken down by relevant segment.

3. Map the critical path

The task here is to connect “contractual terms for input data and retention” to the relevant data, teams and dependencies. Run the control on a normal case and a degraded case, using the exceptions encountered by the teams operating the system as criteria. The concrete output is a map of exceptions, dependencies and owners.

4. Set safeguards

This stage turns intent into control: surround “independent edge-case evaluation” with limits, access rights and a recovery procedure. Measure what actually changes in the limits, rights to act and ability to roll back, including human recovery. Document everything in a control matrix that makes cost and reversibility visible.

5. Test the difficult case

To move forward without concealing deferred cost, test “portability of prompts, data and logs” in a representative scenario and then in a degraded one. Compare before and after across nominal behaviour, induced failure and recovery quality, then ask a stakeholder who did not design the test to review the record of the nominal scenario, failure and human recovery.

6. Build the evidence base

As long as uncertainty remains, external dependency is documented. Required action: compare the outcome, errors, interventions and full cost against the starting point. Start within a scope from which the team can still step back. The evidence must address the gap between the initial promise and the recorded facts; keep it in a file containing logs, discrepancies and decisions that a third party can audit.

7. Decide and review

Because the context evolves, break the signal down by segment: first assign the review and track the metric on an explicit schedule. Use neither an ideal demonstration nor a global average; examine the threshold that triggers correction, expansion or shutdown. The useful deliverable is a review rule with correction and exit thresholds.

Logiks recommendations: evidence, control and reversibility

Our priority is the following risk: a purchase based on a demonstration that does not use representative data. Start where this weakness already creates delay, loss or a disputed decision; the prestigious scope can wait.

Once the baseline has been established, record the measurement date. Keep the baseline at a level where a team can act. A quarterly average does not replace an observation by journey, cohort or type of exception; the reference point must remain actionable.

Treat “purpose and regulatory role” as a documented decision. An accountable owner, hypothesis, limitation and review date are more valuable than a configuration setting whose origin nobody knows.

Test “independent edge-case evaluation” together with “portability of prompts, data and logs”, then test degraded recovery. The exercise must reveal how the system operates and what it costs to run—not merely confirm that the demonstration works.

Expand only when the observed facts support the intended effect and “contractual terms for input data and retention” remain controllable by someone outside the project.

These recommendations express a judgement about sequence: make the risk observable, test the hypothesis concerning “independent edge-case evaluation”, then commit resources. Sophistication follows proof of the claimed effect; it does not replace it. An average can mislead.

Decision framework

Status Observed signal Expected evidence Prudent decision
To be framed “Purpose and regulatory role” exists without a named outcome Dated baseline Do not commit the full scope
In pilot “Contractual terms for input data and retention” are tested on a real workflow Difference from the starting point Include a representative exception
Governed “Independent edge-case evaluation” has an owner and a review Stability, cost and incidents Document degraded mode
Ready to expand or stop “Portability of prompts, data and logs” supports a decision Net value and residual risk Apply the exit rule

The framework does not automatically decide the outcome for AI supplier due diligence. It does force teams to reveal their assumptions about “purpose and regulatory role”, their thresholds and their responsibilities; disagreement then becomes explicit and can be resolved. Scope is decisive.

Common mistakes

1. Confusing activation with outcome

Defining the “purpose and regulatory role” does not prove that the expected effect has been achieved. This mistake shifts the discussion towards the tool when the decision concerns an observable change.

2. Optimising the first available metric

With a named owner, the fallback procedure is accessible: a convenient proxy may improve while the decisive metric deteriorates. Connect every signal to a decision and a safeguard.

3. Ignoring exceptions

When data is incomplete, the budget limit is recorded: the nominal journey often conceals the weakness described above. Test an edge case, a failure and the way the team regains control.

4. Leaving a dependency without an owner

When “contractual terms for input data and retention” belong to everyone, nobody resolves the incident or cost. Assign decision ownership before deployment.

5. Treating risk as a formality

Documenting “independent edge-case evaluation” without correcting the system produces compliance theatre. The evidence file must show a control that was actually performed and its result.

6. Expanding without an exit rule

If “portability of prompts, data and logs” does not support a decision, the pilot continues through inertia. Define continuation, correction and shutdown thresholds in advance.

30 / 60 / 90-day action plan

Days 1 to 30: establish the starting point

  • describe the decision, scope and person accountable for it;
  • record the initial metric value before making any change;
  • inventory the dependencies and their exceptions;
  • state the primary risk and how it will be detected.

When a discrepancy appears, the unit of measurement does not change: the first phase is designed to make disagreement visible. By day thirty, leadership should know the baseline, missing data and precise case against which progress will be judged.

Days 31 to 60: test the critical path

  • implement the primary control on a representative workflow;
  • test recovery under normal and then degraded conditions;
  • record errors, human interventions, delays and costs;
  • compare observations with the starting scenario.

Depending on the selected hypothesis, that hypothesis may be disproved: this pilot is not merely intended to prove that the technology works. It must establish whether the system improves the selected metric without shifting a disproportionate burden onto operations, users or a supplier.

Days 61 to 90: decide and organise what follows

  • consolidate the evidence and have its limitations reviewed;
  • assign every recurring control to a named function;
  • confirm the next review date and exit procedure;
  • expand only if the facts support the effect stated at the outset.

In routine operation, the local check can be reproduced: by day ninety, the initial hypothesis must be demonstrated or disproved. Three decisions remain legitimate—expand, correct or stop the scope. Continuing without a threshold is not a fourth option.

FAQ

How should AI supplier due diligence be defined?

It is a decision framework applied to AI supplier due diligence. It connects “purpose and regulatory role” to the controls for “independent edge-case evaluation” and “portability of prompts, data and logs”, with a baseline, accountable owners and an exit rule.

Where should you begin?

During the audit, operations must be able to recover: begin with a real decision, a baseline and an already observed manifestation of the primary risk. The tool comes after that framing exercise.

What budget should be used?

On the business side, rights to act are documented: add preparation, integration, operations, control, training, incidents and exit costs. Compare that full cost with the expected value, not just the licence or campaign price.

How long should the test run?

The test must cover a complete measurement cycle and at least one exception associated with “independent edge-case evaluation”. Its duration follows from that observation, not from an arbitrary standard.

When should you scale?

Scale when progress remains stable, “portability of prompts, data and logs” is controlled, and responsibilities, costs and exit conditions are documented.

Conclusion

Along the critical path, the full cost becomes visible: a decision is robust when a shared metric connects technical, operational and financial choices. The number of activated options matters less than the ability to explain discrepancies, handle exceptions and reverse a choice that has become costly.

The shift is simple: “AI supplier due diligence” should no longer be treated as a project to deliver, but as a capability to govern so that it produces the intended effect. The trade-off is clear.

Primary sources