Evaluation workbook. Artificial intelligence for your distribution business.
Three steps to decide whether artificial intelligence belongs in your sales operation, with the arithmetic that actually applies to a distributor.

Treat this as a field manual. Fill it in as you go and bring your team into the process. By the end you will have a number you can defend, a test you can run and a rollout plan with dates on it.
Three steps: the math that matters, the evaluation and the rollout.
Step 1 · The math that matters
1.1 Agreement first, tooling second
The finding that is hardest to ignore: RAND interviewed 65 practitioners about why artificial intelligence projects fail, and 84% named a leadership-driven cause as the primary reason, ahead of data and ahead of the maturity of the technology. The most common one of all: leadership did not understand how to set the project on a path to succeed.
The first problem is never the model. It is that three people at the same company believe they bought something different.
Fill this in with your team before you sit through a single demo.
| Question | Examples | Your answer |
|---|---|---|
| What problem are you solving? | Requests that arrive after hours and are still unanswered in the morning · Quotes that depend on two people · Customers who leave because someone else answered first · Long lists nobody gets around to translating | — |
| Which company goal does it serve? | Growing without hiring · Recovering the sale you lose to response time · Ending your dependence on the person about to retire · Opening a new line without a new specialist | — |
| Which number has to move, and who owns it? | Quotes issued per month · Requests answered same day · Dollars quoted · Requests answered with "we don't carry it" | — |
The third question is the disqualifying one. If nobody at your company owns the number you agree to move, ninety days from now there will be no way to tell whether it worked, and the conversation turns into opinions. It is what kills the most projects, and it has nothing to do with technology.
1.2 The arithmetic that actually applies to a distributor
This is where the borrowed customer-support model falls apart. Dividing the cost of a department by the number of conversations says nothing in distribution, because conversations are not the product: quotes are, and most of the money is not in what it costs to answer, but in what never got answered at all.
There are three lines, and they have to be counted separately because they are not the same thing.

Line A · What comes in and never gets quoted
Requests per month × % not quoted today × Average quote value × Historical win rate
The percentage that never gets quoted is the number almost nobody has and almost everybody can pull in an afternoon: count the requests that came in by email and messaging last month, then subtract the quotes that went out. The difference is your number, and it usually comes as a surprise.
What inflates it, and is worth breaking out: the ones that arrived after hours, the ones waiting on a detail nobody asked for, and the ones waiting on the person who knows.
Line B · What gets answered "we don't carry it"
Requests where the brand asked for is out × Average quote value × Rate at which your customer accepts a valid alternative
For a multi-line distributor this line is money already sitting in your warehouse. Answering "we don't carry it" gives away the whole sale, and it gives away the next one too, because the customer has already found someone else to ask.
And the criterion matters: an alternative does not qualify because it looks similar. NAHAD's hose assembly guidelines, aligned with ISO/TS 17165-2, exist precisely for this: they fix the data you need before deciding which product meets the specification.
Line C · The time of the person who knows
Hours per month your specialist spends on the repeatable × Their fully loaded hourly cost
Fully loaded, not salary: benefits, taxes, equipment, the cost of whoever supervises them, and the cost of replacing them when they leave.
Do not add the three lines as if they were the same thing. A and B are recovered revenue. C is freed capacity.
C only turns into money if that person spends the time on something that produces. A calculation that adds them together without saying so is a calculation someone will take apart in the meeting.
| Variable | Where to get it | Your number |
|---|---|---|
| Requests per month | Email and messaging inboxes, counted by hand for one month | — |
| Quotes issued per month | Your ERP | — |
| % never quoted | The subtraction of the two above | — |
| Average quote value | Your ERP; use the median, not the mean, if you have a few large orders | — |
| Historical win rate | Your ERP, quotes against orders | — |
| Requests where the brand was out | Your team knows; if it is not recorded anywhere, that is the first signal | — |
| Specialist hours on the repeatable | Ask them. Do not estimate it | — |
| Fully loaded hourly cost | Payroll plus benefits, taxes, equipment and supervision | — |
1.3 What does not fit in the math, and decides anyway
Two public data points that explain why the technical step is worth more than the spreadsheet shows.
operating margin on two segments of the same company, the same year and with the same customers
DXP · fiscal year 2025increase in EBITDA margin from a single point of price
McKinsey · 130 publicly traded distributorswhere price ranks among the criteria for choosing a distributor
McKinsey · survey of 200+ customersThe first: inside the same distribution business, operating margin changes with what it sells. DXP, which is publicly traded and reports by segment, earned 18.0% in Innovative Pumping Solutions and 8.7% in Supply Chain Services in fiscal year 2025. Nine points of difference inside a single company. Engineering content pays. Fulfillment does not.
The second: McKinsey measured that one point of price produces a 22% increase in EBITDA margin across 130 publicly traded distributors. And in the same article, in a survey of more than two hundred distribution customers, price came in sixth among the criteria for choosing a supplier.
Translated: your customers do not pick you because you are cheap, and you earn more where there is judgment than where there is fulfillment. Any system that moves requests from the second group into the first is working on the line that leaves the most margin, and none of that shows up in a savings calculation.
Read alsoFour operations measured, each with its own period and its own denominator.
Step 2 · The evaluation
2.1 Two filters, in this order
Do not mix the two questions. First, whether the system can work at your company, and only then how well it does it. A system that cannot read your catalog does not deserve a quality test.
Entry filter: can it operate here?
| Area | What to check | Your list |
|---|---|---|
| Capability | Takes the request as it arrives, without asking your customer to use your format · Asks when a missing detail would change the answer · Builds multi-line solutions, not just isolated items · Proposes the valid alternative · Escalates to a person with everything it has already gathered | — |
| Technical fit | Reads the catalog, that customer's pricing and inventory wherever they live · Writes the quote and the order into the system you already use · Connects to what runs today, with no migration project · Your IT group understands which permissions it asks for and why | — |
| Data and security | Where your data lives and who can see it · What happens to your customers' information · What the vendor keeps if you leave | — |
Evaluation filter: how well does it do it?
| Area | What to measure | Your list |
|---|---|---|
| Commercial result | Quotes issued · Dollars quoted · Requests answered same day · Requests that used to be answered "we don't carry it" and now go out with an alternative | — |
| Answer quality | Did it pick the product that meets the spec, or the one that looks similar? · Did it ask for the missing detail? · Did it stop when it was not sure? · Would your specialist sign that quote without correcting it? | — |
Sit your specialist in front of twenty of the system's answers and ask how many they would sign as they are.
That is the best question there is, and it is free. That number tells you more than any percentage anyone puts in front of you.
Read alsoThe questions distributors ask before deciding, answered without selling anything.
2.2 How to run the test
Three steps, and it works with a single vendor. You do not need a four-way bake-off to decide well: one thorough test with the best fit gives more signal than four shallow ones.

One · Build the exam before you grade it. Gather real requests from the last three months. Not the easy ones: the ones that actually cost you.
| Type of request | Why it belongs | Yours |
|---|---|---|
| The photo with no part number | It is the normal case, not the exception | — |
| The after-hours voice note | Measures whether the operation moves when nobody is there | — |
| The twenty-line list in the customer's own words | Measures whether it translates, or whether you have to translate for it | — |
| The brand that is out of stock | Measures whether it recovers the sale or gives it away | — |
| The one your best person took a long time to solve | Measures the ceiling, not the floor | — |
| The one that ended badly | Measures whether it knows how to stop | — |
One fact worth keeping in mind as you build the exam: the STAMPED method in NAHAD's hose assembly guidelines defines seven fields for capturing a requirement, and not one of the seven is a part number. Size, temperature, application, material, pressure, ends and delivery. An entire industry is built on the assumption that the customer does not bring the code, but the problem.
Two · Give it the same material you would give a new hire. If the system is going to work with your portfolio, it needs your portfolio. Spec sheets, manufacturer catalogs, the rules your team applies without writing them down.
| Area | Question | Your answer |
|---|---|---|
| Coverage | Is what you are about to hand over enough to answer what customers actually ask? | — |
| Currency | Is it up to date, or are there three versions of the same catalog in three folders? | — |
| Judgment | The thing that decides which product goes on each line: is it written down anywhere, or does it live in two heads? | — |
The third question is usually uncomfortable and it is the most useful one in this workbook. If the answer is that it lives in two heads, you have found your biggest risk, and it has nothing to do with software.
Three · Grade it, and ask what makes people uncomfortable. Run your requests and grade them with the two tables above. Then ask these four questions in the session — they are answered with a demonstration, not a slide.
- Take this photo and this twenty-line list, written in my customer's own words. Show me what comes out. Not what could come out after a data project: what comes out today.
- When the brand they asked for is out, what do you answer? And on what basis do you decide the alternative works.
- Where does the result get written? If it does not land in the system your operation actually checks, somebody will re-key it, and that somebody will be your best person.
- What happens when the system is not sure? The only acceptable answer is that it stops and escalates to a person with everything it has already gathered.
And a warning about percentages. If a vendor offers you an accuracy rate, ask for the period, the denominator and the customer it was measured on. A number without those three things is not data, it is a phrase. We do not publish accuracy percentages for that reason, and the claim we do stand behind is about behavior: when there is no certainty, the system stops and escalates to a person with what it has gathered.
Read alsoEvery integration, with what each one reads and what each one writes.
2.3 Evaluate the vendor too
What you are buying is not a file you install. It is someone who will have to understand your portfolio, your rules and how you make money.
Two profiles fail on their own, and they are worth recognizing. The consultant or agency that "adds AI for you": they know process and they know the industry, but they do not build systems; they deliver recommendations, and a recommendation does not execute. The AI expert who has never produced that result: they build well, but they have never moved the number you are asking them to move, at a company in your industry, not even without AI.
| Question for the vendor | What you are looking for |
|---|---|
| Who is going to understand my portfolio, and how long will it take them? | If the answer is "you explain it to us," the work is still yours |
| What happens when my catalog or a rule changes? | If every change is billed hourly, the price you were quoted is not the price |
| What do you show me that did not work? | A vendor with nothing but perfect cases has not operated long enough |
| What figures do you publish, and with what period and denominator? | It is the same yardstick you are holding your own team to |
| What happens if I want out? | What you take with you, what the vendor keeps |
Step 3 · The rollout
3.1 The mode, before the scope
This is the decision that reassures a team the most and the one almost nobody makes on purpose. It is not whether to use the system, it is how much rope you give it.

| Mode | What it does | When |
|---|---|---|
| Watches | Reads and writes nothing. You see what it would have answered, at no risk | To convince whoever is not convinced, and to calibrate without exposing a customer |
| Drafts, your specialist signs | The system does the heavy lifting, the person reviews, corrects and sends | The normal starting point. Keeps human judgment on the last line |
| Resolves within your limits | Answers only within what you authorized, and escalates the rest | Once the numbers from the first two have earned your confidence |
You advance when your team wants to, not when the calendar says so. And you advance by type of request, not all at once: the ones already coming out right in the previous mode.
3.2 The plan
| Area | What to define | Your plan |
|---|---|---|
| Scope | Which requests go in now, which later and which not yet | — |
| Phases | Phase 1: the repeatable and low-risk · Phase 2: more lines and more channels · Phase 3: the complex and the edge cases | — |
| Owners | Who owns the content that feeds the system · Who owns configuration and connections · Who owns the number | — |
| Checkpoints | If the target is ninety days, what has to be visible at week 2, week 4 and week 8 | — |
| What shuts it down | What result, by what date, would make you stop. Write it down before you start | — |
That last row is what separates a project from a bet. Writing down in advance what would stop it forces you to define what working means, and it takes the emotional weight out of the decision at day sixty.
3.3 Before you open it to your customers
| Area | Check | Ready? |
|---|---|---|
| Content | Does it have what it needs to answer what customers actually ask? | — |
| Data | Does it read that customer's pricing, inventory and purchase history? | — |
| Behavior | Does it sound like your company? Does it know what it must not promise? | — |
| Escalation | Who does each case reach, with what context, and what happens if that person is out? | — |
| Your team | Do they know what it does and what it does not? Has anyone told them it is not here to replace them? | — |
| Measurement | Is the before counted? With no baseline there is no way to prove the after | — |
The last row is the one most often forgotten and the one that costs the most later. Count the before in the week leading up to the rollout, even by hand. A baseline reconstructed three months later never convinces anyone.
Where this is not the right answer
Two cases, and they are worth stating in full because they save you an evaluation.
When nobody has to decide anything. If a code arrives, it ships and it gets invoiced, and no request requires understanding what it will be used for or combining product with service, there is no judgment to institutionalize. A system that reasons will not give you anything a good order-entry module would not.
When nobody owns the number. It is already above and it is repeated on purpose. It is what kills the most projects.
And a warning about the test almost everyone runs first: putting a general-purpose model in front of the catalog and asking it questions. In BEAVER, a public evaluation run against real enterprise data warehouses, the same method that scores 62.9% on the public benchmark drops to 11.4% when it faces a company's own data, with its names, its abbreviations and its mess. Reading is not solving.
The workbook, on one page
If you take away one thing, let it be this list.
- Agree on the problem and on who owns the number before you see a demo.
- Count what comes in and never gets quoted. That is your business case.
- Count what gets answered "we don't carry it." That is money already sitting in your warehouse.
- Build the exam from your hard requests, not the easy ones.
- Ask to see what comes out today, not what would come out after a data project.
- Ask what happens when the system is not sure.
- Sit your specialist in front of twenty answers and count how many they would sign.
- Choose the mode before the scope.
- Write down what result would shut the project down.
- Count the before in the week leading up to the rollout.
Your ERP knows everything that became an order. This workbook is about everything that never did.
Every figure, with its origin.
The figures in this workbook come from public documents. These are the places where each one is published, so you can verify at the source and not in this article.
- RANDThe Root Causes of Failure for Artificial Intelligence ProjectsAcross 65 practitioners interviewed, 84% named a leadership-driven cause as the primary reason projects fail; the most common one, not understanding how to set the project on a path to succeed.
- NAHADHose Safety Institute · assembly guidelinesThe STAMPED method defines the seven fields that specify an assembly, and not one of them is a part number. Aligned with ISO/TS 17165-2:2018.
- DXP EnterprisesFourth quarter and fiscal year 2025 results18.0% operating margin in Innovative Pumping Solutions against 8.7% in Supply Chain Services, the same fiscal year.
- McKinsey & CompanyPricing: Distributors' most powerful value-creation leverOne point of price lifts EBITDA margin by 22% across 130 publicly traded distributors; in a survey of more than 200 customers, price came in sixth. September 2019.
- BEAVERBEAVER: An Enterprise Benchmark for Text-to-SQLThe same method that scores 62.9% on the public benchmark drops to 11.4% against real enterprise data warehouses.
Peaking's own figures come from a production audit with a cutoff of September 23, 2026, not from a survey.
What people ask before evaluating.
How do you calculate the return on an AI system for a distribution business?
Not with the cost-per-conversation model used in customer support. In distribution the unit is the request that has to be solved before anyone can quote it, and the return has three lines that do not add up to each other: what comes in and never gets quoted (requests per month × percentage not quoted × average quote value × win rate), what gets answered "we don't carry it" and could have been recovered with a valid alternative, and the specialist hours freed. The first two are recovered revenue; the third is capacity, which only turns into money if that person spends it on something that produces.
How long does it take to evaluate a system like this?
The evaluation itself is two to three weeks if you arrive with the test requests already gathered, which is the work that actually takes time. What usually stretches the process is not the technical test but the internal agreement on which number you want to move and who answers for it. That is why the agreement comes first in this workbook and not at the end.
Do I need to compare several vendors?
Not necessarily. One thorough test with the best fit gives more signal than four shallow ones, because it lets you load your portfolio for real and watch the behavior on hard cases. If you do evaluate several, give them all the same material and the same requests, or the comparison says nothing.
What if my catalog is a mess?
That is normal and it is not a barrier to evaluating, but it is the first job of the project. The question that matters is not whether the catalog is clean, but whether the judgment that decides which product goes on each line is written down anywhere or lives in two heads. If it lives in two heads, that is your biggest risk, and it exists with or without artificial intelligence.
Do I have to change ERP?
No. What is being evaluated here works on top of the system your company already has: it reads the catalog, that customer's pricing and inventory wherever they live, and it writes the quote and the order right there. If a vendor makes the project conditional on changing ERP, they are evaluating two things at once and should separate them.
How do I measure whether it worked at ninety days?
With the number you agreed on in step 1 and with your baseline counted before you started. The measures that tend to work in distribution are quotes issued, dollars quoted, requests answered same day and requests that used to be answered "we don't carry it." Reconstructing the baseline afterward never convinces anyone, so count it the week before, even by hand.
What if my team thinks it is here to replace them?
It is the most common objection and it has to be answered head-on, not sidestepped. The rollout mode helps: starting in watch mode or in draft-and-sign makes clear that the judgment still belongs to the person. And it is worth being explicit about which work goes away: the repeatable, not the work that requires judgment. The specialist is still needed; what stops being needed is their availability for the forty requests that come in every day.
You've got this.
Book a demo and see, with your own catalog, what this workbook describes.