Show Me the Evals, AIforce
Salesforce says Claude works better with its new layer than without it. I’d like to see that.
By Chris Pearson
Claude working through Salesforce’s newly announced reasoning layer, AIforce, will do better CRM work than Claude left to sniff around your org, and the rest of your enterprise data estate, by itself was the big bold claim announced last week.
I’m prepared (and cautiously excited) to believe it. I’ve watched Claude sniff around an org by itself. It guesses, confidently, and it burns tokens doing it. I just want to see it work – like I expect it to.
I’ve experimented with asking Claude various “end user” types of requests. Simple queries using standard fields and functionality usually work fine. The more customized your org is (and by that measure, the more tech debt you have) – the less of a chance Claude has of doing what you expect.
This leaves me with more questions than answers. Notably: why hasn’t anyone shown us by comparison how poorly Claude performs without AIforce? That comparison is the whole sale, and it is the easiest kind of claim to test. The materials I was given offer an adoption count instead.
I’d like to see the evals.

Figure 1. The comparison at the center of the AIforce pitch, and the two numbers it would take to settle it.
AIforce was a headline announcement at Dreamforce 2026. The idea: people don’t have to come to Salesforce anymore. Salesforce comes to them, in Claude, in Slack and in Lightning, carrying the data, business logic, permissions and governance already built into the org. Loosely translated: ask an LLM about your pipeline and the answer is supposed to be grounded in your org instead of a plausible-sounding guess. It shipped as Claudeforce, Slackforce and Agentforce Coworker. Amazon, Google and Microsoft are on the partner list, not in the box.
That’s a problem worth solving – but it is also not a new one.
Table of contents
Everyone has been building layers
Anthropic published the Model Context Protocol on November 25, 2024. Seven weeks later, on January 16, 2025, Tapas Mukherjee released a community-built Salesforce MCP server, and a lot of us, me included, were running Claude against sandbox orgs before the month was out. Salesforce made it official with Headless 360 at TDX in April 2026 (a year and some change, later) and expanded it on August 19.

So the roll-your-own path is 20 months old, and the commercial path is crowded. Gearset’s Cam and Copado’s Agentia put org-aware “builder” agents inside your existing release pipeline. Laminar, Myko AI and Ressl AI come at it from the user side, with a plain-language front door and org-specific context making the output correct.
Different products, same job: take a general-purpose model, make it right for your org, and keep it inside the governance you already have. Everyone is building and defending their layer and who their target audience is, Salesforce just named theirs.
The doom loop
The impact is real. So is the risk. A Salesforce org is a decade-plus of metadata, thousands of (questionably named) fields, rules and automations that interact in ways nobody fully documented. Hand a model an ungrounded, poorly written prompt and you get a token-hungry doom loop: every response is an educated guess assembled from partial information, delivered in a tone of complete authority. The model doesn’t know you have that shadow object called Contact__c or you have 3 different “Strategic?” fields on your Opportunity, but only 1 of them is really used. It just answers.
That is the “without” case. AIforce is supposed to be the “with”: requests run on existing permissions, every action routes back through Salesforce, and business logic travels with the data instead of being reconstructed on every call. If that works as advertised, the doom loop shrinks.
If.
Salesforce knows how to test this
AI Hallucination isn’t a dirty little secret. It’s a measured phenomenon with a benchmark industry around it: TruthfulQA, HaluEval, SimpleQA, Vectara’s Hallucination Leaderboard. The AI companies conditioned us to read and pay attention to these. We even added a word to the vocabulary: “evals.” Ethan Mollick has argued for a year that the ones that matter are the ones you build yourself: real tasks from your actual job, scored by experts.
Salesforce’s own researchers agree. Their CRMArena-Pro benchmark (May 2025) found leading agents completing about 58% of single-turn CRM tasks and roughly 35% of multi-turn ones, with near-zero confidentiality awareness. Salesforce also runs a public CRM benchmark, and on the day it announced AIforce it pointed that benchmark at Koa, its new reasoning model, claiming “three times fewer errors.” A ratio, not a pass rate. But at least it’s a claim.

AIforce didn’t even get that (and I read the announcement twice). It has plenty of numbers: 37 prebuilt sales skills, 100,000 Coworker activations in 35 days, one bank with roughly 3,000 users. Great stats for adoption, however, none measures whether an answer was correct or whether a permission boundary held.

Salesforce didn’t forget how to run evals. It left them out of the AIforce launch. And the absence of that data leads directly to a lack of trust (layer).
The test I’d run
A layer is hard to demo. It doesn’t have a screen. It just sits there as ambient intelligence, with nothing to build, activate or control, which may explain the (lack of) excitement at Dreamforce among people who have spent decades building things they can point to. Here is the test I’d propose:
- Convene a panel of community experts to define a core set of agentic CRM tasks, simplest first – ending with complex, multi-turn request that today’s AI rarely gets right.
- Include tasks the agent should refuse, such as a request for records the user can’t see.
- Agree in writing on what an acceptable response looks like, and turn that into a rubric (what’s a good answer, great answer, poor answer, etc).
- Run every task three ways against the same reference org: a raw community-built MCP connection, one of the existing hosted 360 MCP Servers and then Salesforce’s new Headless 360 MCP Server powered by AIforce.
- Run each task five times per configuration and grade blind.
- Pin the model version, log the token spend, and publish the tasks, the rubric and the raw scores.

I’d expect to see similar results between configurations on simple/straight-forward reads and writes, and a murkier picture as requests get more complex. Will AIforce show a clear performance boost? That’s what the test is for.
Who should ask
If you’re a customer, ask how the same model scored without AIforce. If you sell a layer of your own, expect the same question. And, if you’re Salesforce, you own the benchmark, the researchers and the orgs, so publish the with-and-without scores.
One ask for anyone with an AIforce pilot on the calendar: before it starts, get one number in writing. The pass rate with the layer and without it, same model, same org.
Show me the evals.





