You open a proposal to modernize a business platform. The document contains twelve AI functions: predictive analysis of the order pipeline, search agent across supplier archives, automatic classification of support tickets, report generation at month-end, client conversation summaries, data extraction from invoices. Each technical block is illustrated, priced, documented. The question is not whether there is AI, the question is what changes if you remove it.
The standard reflex is to ask for the ROI of each module. That is necessary, but insufficient: an ROI is calculated after the scope is fixed, and the scope is precisely what you are validating. Before calculating return on investment, you need an instrument that separates, function by function, what is load-bearing from what is decorative. This test exists, it holds in four questions, and it applies to any AI block in a proposal, regardless of sector or technology.
This article delivers the four-question test, the rarely articulated symmetric counter-test (ask the vendor where AI adds nothing in your case), and a scoring grid ready to attach to your RFP. It is written strictly from the buyer's side: the goal is not to evaluate the vendor, but to evaluate the proposal itself, and to conclude if necessary that a good vendor has proposed three AI functions too many. At ClaroDigi, we support this evaluation through our AI transformation service and our digital consulting offering which includes critical audit of vendor proposals before signature.
The question is not 'is there AI' but 'what do we lose by removing it'
A proposal contains twelve AI functions. You mentally remove function number 3. What happens? If the answer is "someone has to take this task back manually, but we keep moving forward," the function is accessory. If the answer is "a business process stops, a decision can no longer be made, or a regulatory deadline is missed," the function is load-bearing.
The distinction does not rest on theoretical usefulness, it rests on the fact that a task or a decision disappears without this function. An email summarization module is useful if you receive 200 emails per day; it is load-bearing if without it your support team cannot process requests within contractual SLAs. A search agent across archives is handy if you often look through old files; it is load-bearing if without it you cannot retrieve the information needed to respond to an audit or a compliance notice.
The removal test immediately reveals two categories of functions that inflate proposals without structural value:
-
Technological duplicates: two different modules that cover the same decision or the same critical task. Example: a product recommendation engine and a chatbot that also recommends products from the same catalog. You remove one of the two, the business process continues. One suffices.
-
Comfort modules: a function that improves user experience without blocking any decision. Example: an automatic summary at the top of each report. If this summary disappears, the reader spends three more minutes reading the full report, but no decision is delayed and no SLA is threatened. It is comfort, not a load-bearing function.
The removal criterion does not eliminate these functions permanently, it defers them to phase 2. A phase 1 must concentrate budgets and risks on the functions without which the system cannot operate. Comfort modules can come later, when you have verified that the load-bearing functions hold up under load.
The four-question test, applied function by function
The removal test is the first filter. If a function passes it (something stops without it), you then apply four questions that separate a load-bearing AI function from one that will end up disabled after three months. These four questions apply to each AI function in the proposal, module by module.
1. What decision or task disappears if we remove this function?
If the answer is vague or starts with "it facilitates," "it accelerates," "it improves," that is a warning signal. A load-bearing function has a clear, factual answer. Examples of good answers:
- "Without this module, the procurement team cannot identify suppliers behind on delivery within 24 hours."
- "Without this module, the controller cannot reconcile budget commitments and received invoices before monthly close."
- "Without this module, the customer service team cannot retrieve the commercial terms applicable to each case within the contractual timeframe."
If the answer does not identify a precise task or decision that disappears, the function is not load-bearing. It may be useful, but it should not enter the phase 1 scope.
2. Who triggers this function, and at what real frequency?
An AI function that is never called in production costs exactly as much as a function running continuously, but it creates zero value. You must therefore require from the vendor a nominal use case scenario with an identified actor and a realistic frequency.
Examples of realistic frequencies to demand:
- A search agent across supplier archives: triggered by whom (buyer, contract manager, internal auditor)? how many times per week?
- An automatic support ticket classification module: how many incoming tickets per day? does this volume justify automation, or can a support manager still sort manually in 20 minutes?
- An automatic summary of monthly reports: how many reports per month? how many readers per report? do these readers actually read the summary, or do they go straight to the data tables?
If the vendor answers "it depends," "according to needs," or "we will see in use," the function has no nominal use case. You are paying for an option that may never be used. Defer it to phase 2 or remove it.
3. What happens when this function makes a mistake?
Every AI function makes mistakes. The question is not whether it makes mistakes, but what happens when it does. A load-bearing function has a documented answer to this question, with a detection and correction plan. A decorative function has no plan, because no one thought about the risk.
Examples of error scenarios to demand:
- A contract generation agent: if the generated contract contains an incorrect clause or a miscalculated value, who detects it before signature? within what timeframe? with what impact if the error is not detected before signature?
- An automatic ticket classification module: if an urgent ticket is wrongly classified as low priority, how soon is it detected? is an SLA threatened?
- An automatic summary of a monthly report: if the summary omits a critical point or contradicts a figure in the report, who notices? what is the impact if the erroneous summary circulates internally or externally?
If the vendor answers "the model is very reliable," "we will monitor the metrics," or "we can always correct afterward," you have no error management plan. An AI function that has no error detection and correction plan is not ready for production. Defer it to phase 2, until this plan is built.
4. How much does this function cost in use, not in construction?
Most quotes describe the construction cost of an AI function (development days, configuration, testing, training). Very few describe the usage cost: how much each API call to the model costs, how much storing input and output data costs, how much human supervision costs to detect and correct errors.
The usage cost can be ten times the construction cost over three years, and it often appears as a surprise in the run phase. Before validating a function, you must require two separate lines in the quote:
- Construction cost: how much to develop, test and deploy the function (fixed price or time-and-materials).
- Estimated usage cost: how much per month in production, based on the nominal usage volume described in question 2.
If the vendor cannot estimate the usage cost, they have not modeled the call volume nor the unit cost of APIs. You are buying a function whose total cost you do not know. Ask for a low and high range, and plan a revision mechanism if the real volume diverges by more than 30% from the estimated volume.
A function that passes all four questions (identified task or decision, realistic actor and frequency, documented error plan, estimated usage cost) is a load-bearing function. A function that fails on any of the four is a function that must exit the phase 1 scope or be clarified before signature.
Real usage frequency: the criterion no one verifies before paying
You have an AI module that costs $30,000 to develop and $500 per month to run. The vendor tells you it will be "very useful for complex cases." You sign. Three months after go-live, you look at the usage logs: the module was called four times. Cost per real use: $8,000. You just paid for a module that could have been replaced by an email to the vendor each time the complex case arose.
This scenario is common, and it is avoidable if you ask the real frequency question before signing. The question is not "is this module useful," the question is "how many times per week will this module actually be called in production." If the vendor cannot answer, they have not modeled the usage volume.
Here is how to obtain a realistic frequency estimate:
-
Identify the triggering actor: who calls this function? An end user? An automated process? An administrator? How many actors of this type do you have in the organization?
-
Reconstitute the current volume: how many times today is this task performed manually or via another tool? If you automate a task that no one does today, the function will not be used after deployment.
-
Apply a friction coefficient: even if the function automates a frequent task, not all users will adopt it immediately. A coefficient of 50% in the first year is realistic (half of theoretical uses become real uses). A coefficient of 100% assumes perfect training and instant adoption, which never happens.
-
Verify alignment with usage cost: if the function will be called 10 times per month and each call costs $10 in API, the monthly cost is $100. If it will be called 500 times per month, the monthly cost is $5,000. The two scenarios do not justify the same construction budget.
If the vendor resists this exercise, that is a warning signal. A serious vendor has modeled usage volumes, if only to properly size their infrastructure. If they cannot give you a range, you are buying a function whose usage no one knows.
Usage cost versus construction cost: the two lines to separate in the quote
A standard quote for an AI project presents the development cost: so many days of design, so many days of development, so many days of testing, so many days of training. What the quote does not show is how much the function costs once it runs in production. This usage cost is composed of three line items:
-
API cost: each call to a language model (GPT-4, Claude, Gemini) costs between $0.01 and $2 depending on model size and length of inputs and outputs. If a function is called 1,000 times per month with an average cost of $0.60 per call, the monthly API cost is $600, or $7,200 per year.
-
Storage cost: if the AI function generates or ingests documents (product sheets, reports, invoices), these documents must be stored. A volume of 10 TB costs about $250 per month in standard cloud storage, or $3,000 per year. If the volume grows by 30% per year, the storage cost in year 3 is $5,070 per year.
-
Human supervision cost: any AI function that produces content for a client, a partner, or a regulatory authority must be supervised. A half-time supervisor costs about $30,000 per year. If the function generates 50 documents per month and each document requires 10 minutes of review, the supervision cost is 8.3 hours per month, or 0.2 FTE, or about $6,000 per year.
Total usage cost for this function over three years: ($7,200 + $3,000 + $6,000) × 3 + storage adjustment = about $52,000. If the construction cost was $40,000, the total cost over three years is $92,000, of which 57% is usage costs. If you had only budgeted the $40,000 construction, you are 130% over budget by year 3.
Before signing, require from the vendor a quote separated into two sections:
- Section 1: construction cost (fixed price or time-and-materials, with phase breakdown).
- Section 2: estimated usage cost (monthly and annual, with API + storage + supervision breakdown, based on the usage volume described in question 2 of the test).
If the vendor cannot fill section 2, you have no visibility on the total project cost. Refuse to sign until this section is documented. A good vendor has already calculated these figures to size their infrastructure; they have no reason to hide them.
What must exit phase 1, and why scope refusal is a good signal
A phase 1 must concentrate on the functions without which the system cannot operate. Everything else, however useful, must exit the scope or move to phase 2. This sorting is the buyer's mission, not the vendor's: a vendor has an interest in maximizing scope, you have an interest in minimizing it.
Here are three categories of functions that must systematically exit a phase 1:
Comfort modules
A comfort module improves user experience without unblocking any critical decision. Examples: an automatic summary at the top of each report, an internal FAQ chatbot, automatic title generation for shared documents. These modules are useful, but they do not structurally change the organization's capacity to function.
Sorting criterion: if you remove this module, does a business process stop or is an SLA threatened? If the answer is no, the module goes to phase 2. You will do it after verifying that the load-bearing functions hold up under load.
Technological duplicates
Two different modules that cover the same decision or the same critical task. Example: a product recommendation engine based on purchase history, and a chatbot that also recommends products from stated preferences. Both modules are load-bearing if you take them in isolation, but together they are redundant. One suffices.
Sorting criterion: identify all decisions and all critical tasks covered by the proposal. If two modules cover the same decision or the same task, keep the one with the best ratio (business value / cost + complexity), and defer the other to phase 2 or remove it.
Functions without an identified actor
An AI function that has no clear triggering actor (who calls it, when, why) will never be used in production. Example: an advanced report generation module "for exceptional cases" without specifying who requests them or at what frequency. If the vendor cannot name the actor and estimate the frequency, the function has no nominal use case.
Sorting criterion: apply question 2 of the test (who triggers this function, and at what real frequency?). If the vendor cannot answer with a named actor and a quantified frequency, the function exits the scope.
Scope refusal is often interpreted as a sign of rigidity. In reality, it is a signal of buyer maturity: you have understood that phase 1 is not a catalog of all possible functions, but a minimal foundation that must work in production before adding anything else. A vendor who accepts reducing their phase 1 scope without resistance shows you they have understood this. A vendor who resists any scope reduction is maximizing their short-term revenue, not maximizing your chances of success.
The counter-test: ask the vendor where AI adds nothing in your case
The four-question test identifies load-bearing AI functions. The rarely articulated symmetric counter-test consists of asking the vendor: "In my case, where does AI add nothing?" A vendor who answers "everywhere we proposed it, it adds something" is selling, not diagnosing.
A good vendor has a precise answer to this question. Examples of good answers:
-
"Your procurement team already does a good job of supplier tracking with an Excel file updated weekly. Automating this tracking with AI will only save you 2 hours per week, which does not justify the $22,000 development cost. We can keep the Excel and invest this budget elsewhere."
-
"You have 15 support tickets per week. An automatic classification agent is useless at this volume, your support manager can sort manually in 20 minutes. We remove this module."
-
"Your product catalog contains 80 references. An AI recommendation engine does not have enough data to be better than a simple business rule (for example: recommend the three best-selling products in the same category). We can do this with a SQL query, no need for a model."
This counter-test has two virtues:
-
It reveals the vendor's posture. A vendor who refuses to answer or dismisses the question ("everything is useful," "we will see in use") is not advising you, they are maximizing their scope. A vendor who spontaneously identifies two or three useless functions in your context is making an honest diagnosis. You can trust them on the functions they recommend.
-
It forces the vendor to argue their choices. An AI function that survives the counter-test ("no, in your case this one is really necessary, and here is why") has been validated twice: once in the positive direction (it adds something), once in the negative direction (removing the AI here would be a mistake). This double filter eliminates decorative functions.
Systematically ask this question in your RFPs: "In our case, where does AI add nothing, and why?" If the vendor cannot answer, put them in competition with a vendor who can.
The scoring grid to attach to your RFP
You launch an RFP to modernize a business platform with AI functions. You receive five proposals. Each proposes between eight and fifteen AI functions. How do you compare these proposals objectively, without being impressed by the volume of proposed functions?
The grid below scores each AI function from 0 to 4 points according to the answers to the four test questions. A function that totals 4 points is load-bearing. A function that totals 0 or 1 point is decorative or insufficiently specified. A function that totals 2 or 3 points is in the gray zone: it can become load-bearing if the vendor clarifies certain points.
| Criterion | 0 point | 1 point |
|---|---|---|
| 1. Identified task or decision | No task or decision identified, or vague answer ("it facilitates," "it improves") | Task or decision clearly identified ("without this module, team X cannot do Y within deadline Z") |
| 2. Actor and frequency | No named actor, or unquantified frequency ("according to needs," "it depends") | Named actor (role or function) and quantified frequency (per day, per week, per month) with applied friction coefficient |
| 3. Error plan | No error detection or correction plan, or generic answer ("we monitor the metrics") | Documented detection and correction plan: who detects the error, within what timeframe, with what impact if it is not detected |
| 4. Estimated usage cost | No estimated usage cost, or cost not separated from construction cost | Estimated usage cost (monthly or annual) with API + storage + supervision breakdown, based on nominal usage volume |
Scoring a proposal: add up the points of all proposed functions. A proposal with 12 functions at 2 points on average (24 points total) is worse than a proposal with 6 functions at 4 points (also 24 points total, but with half the complexity and risks).
Qualification threshold: any function that totals less than 3 points must be clarified before signature, or removed from the phase 1 scope. If a vendor refuses to clarify, eliminate the proposal.
This grid can be attached to your RFP as a qualification criterion: "Each proposal must document, for each proposed AI function, the following four elements: covered task or decision, triggering actor and usage frequency, error management plan, estimated usage cost. Proposals that do not document these elements will not be evaluated."
FAQ
Can we apply this test to a SaaS platform that already integrates AI?
Yes. A SaaS platform with integrated AI (for example Salesforce Einstein, HubSpot AI, or Monday.com AI) must pass the same test function by function. Identify each AI function proposed by the platform, ask the four questions (what task disappears, who uses it and at what frequency, what happens in case of error, how much does it cost in use), and decide whether you activate this function or not. The fact that a function is included in the subscription does not mean it is useful in your context. Many SaaS platforms charge supplements for advanced AI modules; apply the test before paying.
How to evaluate an AI function that improves the quality of a decision without making it possible?
An AI function that improves the quality of a decision already made manually (for example: a recommendation engine that suggests three suppliers instead of one) must be evaluated according to the measurable quality gain. Ask the question: how many incorrect or suboptimal decisions are made today without this function, and what is the cost of these errors? If this cost is greater than the total cost of the function over three years, the function is justified. If this cost is lower or unmeasurable, the function goes to phase 2.
A vendor tells us that certain AI functions cannot be priced before testing in production. Is this normal?
No. A vendor can have a margin of uncertainty on the usage cost (for example: between $400 and $800 per month depending on real volume), but they cannot refuse to give a range. If the vendor refuses to price, they have not modeled the call volume nor the unit cost of APIs. You are buying a function whose total cost you do not know. Require a low and high range with volume assumptions, and plan a revision clause if the real volume diverges by more than 30% from the assumptions.
What to do if a function scores 2 or 3 points out of 4 on the test?
A function at 2 or 3 points is in the gray zone. It has value elements, but it is not completely specified. Ask the vendor to clarify the missing points before validating it. For example: if the function has 1 point on the "actor and frequency" criterion (no quantified frequency), ask for a frequency estimate with calculation assumptions. If the vendor refuses or cannot, remove the function from the phase 1 scope. A gray-zone function can become load-bearing if clarified, or decorative if it remains vague.
How to handle a vendor who resists scope reduction?
A vendor who systematically resists any scope reduction is maximizing their revenue, not maximizing your chances of success. Two options: either you impose the reduced scope and verify that the vendor accepts without friction (a good vendor understands the interest of a concentrated phase 1), or you put this vendor in competition with another who accepts the phase 1 / phase 2 split. Resistance to splitting is a warning signal: a mature vendor knows that an overloaded phase 1 increases failure risks, and they have an interest in your success to win phase 2. Our AI Transformation Sprint offering adopts precisely this incremental scoping approach.
