AI REVENUE AND GOVERNANCE
AI in Sales: Where It Helps and Hurts
A Harvard and BCG field experiment with 758 consultants found AI users were 19% less likely to be correct on tasks outside its frontier.

A randomized field experiment with 758 consultants found something most AI-in-sales content omits. On tasks outside the tool’s capability, people using AI performed worse than people using nothing at all.
Executive summary
- In a preregistered field experiment run with 758 Boston Consulting Group consultants, published in Organization Science in March 2026, subjects using AI on tasks inside its capability completed 12.2% more tasks, 25.1% faster, at significantly higher quality.
- On a task deliberately chosen to sit outside that capability, subjects using AI were 19% less likely to produce a correct solution than subjects working without it.
- The authors call the boundary a jagged frontier, because it is uneven and invisible. Two tasks that look equally difficult to a manager can sit on opposite sides of it.
- Sales is unusually exposed, because the tasks that decide the number are concentrated on the wrong side of that line.
- One question sorts most sales tasks correctly: is the answer contained in the input? If everything required is in the transcript, the notes, the CRM record or the public web, AI helps. If the answer depends on something nobody wrote down, AI will still answer, and the answer will be confident.
- The fix is not to restrict AI. It is to point it at the input to a decision rather than at the decision.
The finding most AI-in-sales content leaves out
Nearly every article on AI in sales reports the upside. The upside is real and it is well evidenced.
Fabrizio Dell’Acqua, Ethan Mollick, Karim Lakhani and colleagues ran a preregistered randomized experiment with 758 consultants at Boston Consulting Group. Across 18 realistic knowledge tasks inside the capability of the model, consultants using AI completed 12.2% more tasks, completed them 25.1% more quickly, and produced work of significantly higher quality. The peer-reviewed version appeared in Organization Science in March 2026.
Then the researchers did the thing that makes the study useful. They included a task designed to fall outside the model’s capability, and they measured what happened.
Consultants using AI were 19% less likely to reach a correct solution than consultants working without it.
Not equal. Worse. The tool did not simply fail to help. It displaced the reasoning that would have produced the right answer.
The authors named the boundary a jagged technological frontier. Jagged, because it does not follow difficulty. A task that feels hard to a human can sit comfortably inside it, and a task that feels routine can sit outside. Nothing on the surface of the work tells you which is which, and the model gives no signal either.
Why sales is more exposed than most functions
Every function has tasks on both sides of that line. Sales has an unusual distribution, and it is unfavorable.
The tasks that consume the most hours in a sales organization are largely inside the frontier: writing follow-ups, summarizing calls, researching accounts, updating records, reformatting proposals. That is why AI feels transformative in the first month. Hours come back immediately and visibly.
The tasks that determine the number are largely outside it: deciding whether a deal is real, choosing where to spend the last hours of the quarter, pricing a non-standard agreement, reading a buying committee, deciding when to walk away.
Those decisions turn on information that was never written down. The pause before an answer. The person who stopped replying. The budget conversation that happened without you. A CRM record contains what was said. Sales judgment is mostly about what was not.
So the function gets a fast, real, visible win on the hours, and takes on a quiet, unmeasured risk on the decisions. Those two effects show up in different places on the P&L and at different speeds, which is part of why they are rarely connected.
The dividing line: is the answer in the input?
One question sorts most sales tasks correctly.
Is everything needed to answer this contained in what I can give the model?
If yes, the task is inside the frontier. Transcripts, notes, public company information, CRM history, prior proposals, an example of a message that already works. The model is reorganizing, extracting, summarizing or recombining information that is present. It is good at this and it is getting better.
If no, the task is outside. And the critical property of outside-frontier tasks is not that the model refuses. It answers anyway, fluently, in the format you requested, with no indication that the basis for the answer does not exist.
The map
| Sales task | Side | Why |
|---|---|---|
| Summarizing a call from a transcript | Inside | Everything needed is in the transcript |
| Drafting a first-pass follow-up from notes | Inside | The content is the notes, reorganized |
| Researching an account before a meeting | Inside | The facts are public and retrievable |
| Extracting CRM fields from unstructured notes | Inside | Transformation, not judgment |
| Producing variants of a message that already works | Inside | The pattern is in the example you provide |
| Translating or adjusting tone for a segment | Inside | Both the content and the target are specified |
| Deciding whether a deal is real | Outside | The decisive signals are usually what the buyer avoided saying |
| Choosing which two deals get Friday afternoon | Outside | Requires knowing people, not records |
| Pricing a non-standard deal | Outside | Depends on precedent, relationship, and what you are willing to lose |
| Reading the buying committee | Outside | The org chart is not the power map |
| Setting a close date for a specific deal | Outside | Depends on the buyer’s internal calendar, which is not in your system |
| Deciding when to walk away | Outside | Requires accepting a sunk cost, which is a judgment, not an inference |
The pattern is worth stating plainly, because it is the useful part. Inside the frontier, the model handles the material. Outside it, the model is being asked to supply the missing information, and it does so by producing something plausible.
The failure mode nobody watches for
The experiment’s result is counterintuitive until you consider the mechanism.
A confident, well-structured answer is more persuasive than an uncertain human judgment. When a rep asks whether a deal will close and receives a clean paragraph explaining that it will, with three supporting reasons, that output competes directly with the rep’s own unease. The output usually wins, because it looks like analysis and the unease looks like a feeling.
The unease was the signal. It was built from a hundred prior deals and it was carrying information the CRM never captured.
This is why performance falls rather than staying flat. The tool does not add a bad input to a good process. It replaces the process.
In sales this concentrates exactly where it costs the most. Qualification, prioritization, pricing and walk-away decisions are the four levers with the largest effect on the number, and all four sit outside the frontier.
Point AI at the input, not at the decision
The response is not to restrict access or to run a policy exercise. It is a change in how the question is asked.
For outside-frontier tasks, AI should widen the input to a human decision and stop there.
| Instead of asking | Ask |
|---|---|
| Should I keep this deal in the forecast? | List every commitment this buyer made, and every question they avoided answering |
| What should I price this at? | Show me every deal in the last 18 months with this shape, and what we charged |
| Who is the decision maker? | List everyone who appeared in any thread, and what each one asked about |
| Is this account worth pursuing? | Compare this account against our last 20 closed-won and closed-lost on the criteria we defined |
| When will this close? | List every date the buyer has mentioned, and what had to happen before each one |
Every question on the right is answerable from material that exists. Each one makes the human decision better without making it for them, and each one surfaces the specific thing a rep would otherwise have to remember to check.
That is the whole adjustment. The model prepares the evidence. The person makes the call. It costs nothing to implement and it moves a task from the wrong side of the frontier to the right one.
What this changes about how you measure AI in sales
Most sales organizations track AI adoption: seats activated, weekly active users, percentage of reps using the tool.
That metric assumes usage is good. The experiment shows usage is good on one side of a line and harmful on the other, which makes an undifferentiated adoption number not merely incomplete but directionally misleading. A team with 90% adoption concentrated on qualification decisions is in worse shape than a team with 40% adoption concentrated on call summaries, and the dashboard will show the opposite.
Measure by task class instead.
There is also a diagnostic worth running, and it is cheap. Compare win rates on deals where AI-generated qualification was accepted without human revision against deals where a human revised it. If the unrevised group wins less, you have located your frontier with your own data, in your own market, which is the only version that matters.
That comparison requires only one thing you may not have: a record of which outputs were revised. Start capturing it now and the answer arrives in a quarter.
Five ways this goes wrong
- Treating adoption as the goal. Adoption on outside-frontier tasks has negative value. The metric needs a denominator of task type.
- Rolling out to qualification first. It is the most attractive use case in a demo and the worst one to start with, because errors compound quietly through the pipeline for a quarter before anyone sees them.
- Assuming better models move the line. Models improve, the frontier moves, and it stays jagged. Tasks that depend on unwritten information do not become answerable because the model got larger.
- Letting confidence stand in for evidence. Output fluency carries no information about whether the basis for the answer exists. Reps are not trained to make that distinction, and nothing in the interface helps them.
- Skipping the baseline. Without recorded win rates and cycle times from before the rollout, you cannot tell the difference between AI working and a good quarter. This is the same discipline that makes any AI investment reach earnings rather than stopping at productivity.
Where this leaves a sales leader
The hours saved on inside-frontier work are real, immediate, and worth capturing. Nothing here argues for slowing that down.
The discipline is in the second list. Name the decisions in your sales process that depend on information nobody wrote down, and make it explicit that those stay with people, supported by better inputs rather than replaced by generated answers.
That exercise takes an afternoon with your sales leadership. It is the difference between a function that gets faster and a function that gets faster at reaching the wrong conclusions.
Talk it through
If you are deploying AI across a sales organization and have not yet drawn the line between the tasks it should own and the decisions it should only inform, that is worth an hour before the rollout rather than after.
Frequently asked questions
Does AI actually make salespeople worse at anything?
On tasks outside the model’s capability, yes. In a preregistered field experiment with 758 BCG consultants published in Organization Science in 2026, subjects using AI were 19% less likely to produce a correct solution than subjects working without it. On tasks inside its capability the same subjects completed 12.2% more work, 25.1% faster, at higher quality. The effect depends entirely on which task the tool is pointed at.
What is the jagged frontier?
It is the boundary between tasks AI performs well and tasks it degrades, named by the authors of that study. It is called jagged because it does not follow difficulty. Two tasks that appear equally hard can sit on opposite sides, and nothing about the task’s surface tells you which.
Which sales tasks should AI handle?
Tasks where the answer is contained in the input: call summaries, first-draft follow-ups, account research, CRM data extraction, message variants, tone adjustment. Tasks where the answer depends on unwritten information should stay with people: qualification, prioritization, pricing non-standard deals, reading a buying committee, and deciding when to walk away.
How do I find the frontier in my own sales organization?
Compare win rates on deals where AI-generated qualification was accepted unrevised against deals where a human revised it. If the unrevised group wins less, you have located the boundary using your own data. This requires recording which outputs were revised, so start capturing that now.
Should we stop using AI for forecasting then?
Not stop, redirect. Asking a model to predict whether a specific deal closes is an outside-frontier question. Asking it to list every date the buyer mentioned and what had to happen before each one is inside, and it makes the human forecast better. The distinction is between generating the judgment and assembling the evidence.
Sources
- Fabrizio Dell’Acqua, Edward McFowland III, Ethan Mollick, Hila Lifshitz, Katherine C. Kellogg, Saran Rajendran, Lisa Krayer, Francois Candelon and Karim R. Lakhani, “Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality,” Organization Science, Vol. 37, No. 2, published online March 11, 2026. DOI 10.1287/orsc.2025.21838
- Working paper version of the same study, Harvard Business School and Boston Consulting Group, 2023. SSRN
- McKinsey and Company, State of AI survey, figures reported August 25, 2026, survey of 1,719 professionals. mckinsey.com
APPLY THE THINKING
Turn this analysis into an accountable operating decision.
Start with the revenue, GTM, RevOps, governance, or AI constraint that matters most.
Request an AI Revenue Diagnostic