Your AI strategy should outlast your favorite model

Pauline PhamPauline Pham
-October 7, 2026
Dust x Index
Introducing the Dust Index: a study grounded in real-world business use of AI, examining model choice, shared playbooks, and what it means for a company to own its AI.
An AI model can top a leaderboard and still be the wrong choice for your company’s next task. What matters is whether it can do the work reliably, at a cost you understand, in a way your colleagues can use.
The stakes change when that work becomes part of the daily routine. A sales team relies on an agent to prepare account research. Finance uses another agent to investigate a revenue figure. They are building something valuable around the model.
That investment should survive the next model release. Companies need to be able to adopt a better model, respond to a price change, or move to another provider while preserving the knowledge and playbooks they have built. Agent traces (records of the steps an agent took, the tools it used, and the results it produced) help teams check what changes when they switch models. As AI takes on more responsibility, understanding and controlling it becomes an operating requirement.
Yet a model benchmark cannot tell us on its own how well companies are making that transition. Which agents become useful beyond their creators? Does access to several providers translate into meaningful choice? Can teams connect AI spending to the work it produces? These questions determine whether better models turn into better businesses.
That is why we’re introducing the Dust Index, our first study of AI models and workflows. It examines that transition through three dimensions: whether useful agents become essential for collaborative work, whether companies can change the models they use, and whether they can understand the cost of their workflows.
The Dust Index combines evaluations on real work inside Dust, aggregated and anonymized customer usage data, and interviews with the people building and maintaining agents, whom we call AI Operators. Together, these give us a way to examine both what models can do and how companies create compounding and transformative value with AI.
In other words, how companies put AI agents to work.

On Dust, 65% of all messages sent to agents go to custom agents used by more than one person.

This is one sign that the knowledge and instructions built into an agent become useful beyond its creator, giving colleagues a shared resource for their own work.

1. Model choice starts with the work

The question for an AI Operator is specific: can this model do this job accurately, quickly enough, and at an acceptable cost? Dust gives teams access to models from OpenAI, Anthropic, Google, Mistral, and all the best open-weight models, so they can compare them in real life on real work data and choose the one that works best for each task, and switch as their needs change.
For the first Dust Index, we compared Anthropic’s Opus 5.5, OpenAI’s GPT-6 Astra, GLM 5.3, and xAI’s Grok 4.7 on three kinds of work: synthesis across multiple sources, financial data analysis, and generating presentations and live dashboards. We also compared different reasoning levels.
Cross-source synthesis tests whether a model can turn scattered company information into a reliable summary of what is happening. For example, it might need to reconcile project updates, distinguish completed work from promises, identify unresolved dependencies, and recommend a next step. A fluent summary that misses the main blocker fails the job.
The Finance tasks demand a different kind of precision. Can the model reconcile a revenue bridge? Calculate retention using the right starting cohort? Keep new business out of a retention calculation? A plausible number with the wrong denominator is still wrong.
The Frames tasks ask models to turn supplied information into dashboards, interfaces and presentations, or make constrained changes to existing ones. They also reveal how models handle missing or contradictory inputs and preserve the content they have been asked to keep.
Cross-source synthesis produced the clearest differences. Claude Opus 5.5 had the highest preference estimate with marginal difference between medium and max efforts. GLM 5.3 (max effort) and GPT-6 Astra (max effort) followed. These two models are nearly tied in our benchmark on real-world business use cases.
We were not surprised to see Anthropic’s and OpenAI’s models stand out from the crowd. However, it’s worth noting that GLM 5.3 could be considered as a credible synthesis challenger worth testing on a company’s own work.
In finance, Claude Opus 5.5 (max effort) was the strongest candidate once again. GPT-6 Astra (medium effort) and Grok 4.7 (xhigh effort) followed suit. As you can see in the chart above, the complete finance ranking was much less predictable. And the reason for that is quite simple.
In our testing, we’ve noticed that models have come a long way over the past few months, especially when it comes to calculation, math and data processing in general. That’s why it’s harder to differentiate one model from another for financial analysis as they are all adequate choices for these tasks.
Our Frames benchmark did not produce a definitive ranking. All models were able to generate interactive dashboards and presentations. Once again, this was not the case a few months ago. Choosing one model over another came down to taste and design. Teams should test lower-cost options before assuming a more expensive configuration is worth the extra cost.
But charts only tell part of the story. At the end of the day, users themselves can feel when a model is not nearly smart enough for the task at hand.
At 1Password, Liam Ehrlich, a staff GTM engineer, builds Patch Prospector, a workflow that helps salespeople identify promising accounts and find reasons to reach out. It brings together research from several agents, then lets salespeople ask follow-up questions about the recommendations.
“I had to step that one up recently from [Claude] Sonnet to Opus just because it was working with a lot of context,” Ehrlich says.
The task had grown beyond producing a report. The agent also needed to explain its recommendations and help salespeople decide what to do next. Ehrlich found that the more capable model handled those follow-up questions better.
A model that is sufficient for one task may fall short when the scope expands.

2. The important agent is the one your colleagues use

An agent starts to become a company asset when someone other than its creator can rely on it. The builder’s knowledge becomes reusable.
Among our enterprise customers, only 1.8% of custom agents had 50 different users or more. However, these agents represented 43.3% of the total volume of messages and interactions.
Once colleagues rely on an agent, the person who built it has to keep it working well. That means checking that its information is up to date, that model updates haven’t made its answers worse, and that it still does the job people need it to do.
At 1Password, Tré Bembry, a senior IT engineer, describes a distinction between experimentation and publishing something for wider use. Employees can build their own agents while team champions review use cases before publishing shared agents.
“Anyone can build an agent, but we limit who can actually publish agents.” — Tré Bembry, 1Password.
1Password’s Liam Ehrlich explains how colleagues find the right agents to use: an enablement colleague has built starter packs that point salespeople to agents and skills for their role and sales stage. The aim is to give people a maintained route through a task, while leaving room to experiment.
Argon & Co, a consultancy with more than 1,000 employees worldwide that helps companies improve their operations, draws a similar boundary around business processes. Alexandre Starck, an associate partner at the firm, argues that the team responsible for a CRM should own the agents that interact with it, because that team carries the business rules. For personal productivity, he is comfortable with people building agents that fit their own habits.
Shared use therefore needs judgment. Some agents are useful to just one person, and that’s fine. When several people do the same task, a shared agent can help them follow the same instructions and use the same information, with someone responsible for keeping it up to date.

3. Model freedom means being able to change course

A company should be able to choose models for the work it wants to make possible, then revisit those choices as its ambitions and the models’ capabilities change.
In our analysis, AI operators who created some of the most used AI agents in their own companies consistently tried several models and providers. 68% of them tried at least two different providers. On average, they used more than 5 different AI models for their AI agents.
That suggests that companies are cautious about the dependence that can be created when relying too much on one provider.
Styleheads, a Berlin-based marketing agency, is a useful example of the distinction. Kimberley Haverland, its marketing data analyst and automations manager, says model selection is largely left to individuals. People draw on previous experience, compare answers, and share recommendations. She describes Gemini becoming popular for research through colleagues’ recommendations, rather than a mandated allocation of providers to tasks.
“I only change it if I’m really unhappy with the results that I’m getting from a certain model.” — Kimberley Haverland, Styleheads.
That is a different picture from a company constantly optimizing a model portfolio. Familiarity and satisfactory results can be enough to keep a choice in place. 1Password’s Liam Ehrlich describes a similar pull toward the Anthropic models he knows well, while acknowledging that he needs to experiment more with alternatives.

4. An AI bill needs a unit of work

An AI bill can rise because teams are getting more done or because agents are wasting money. The total alone cannot tell you which. Until a company can connect its spending to the work delivered, it risks cutting useful tools and paying for ones that don’t earn their keep.
Tré describes the next step at 1Password as matching models more carefully to tasks. The problem runs in both directions: an expensive model can be excessive for a simple job, while a weaker model can consume more resources trying to do work it cannot handle well.
“Right sizing models is definitely something that we’re looking towards for the future.” — Tré Bembry, 1Password.
At Styleheads, Kimberley likewise says being mindful about credit usage could prompt more guidance on model use. However, she hasn’t implemented a system for attributing cost to outcomes.
The next question is whether a company can follow the cost through an entire workflow: the agent, the models it calls and the tools it uses. That attribution, paired with an assessment of the result, can help an operator distinguish a costly but valuable process from one that needs redesigning. And this is something we offer on Dust.
For many companies, even that basic visibility is still missing.
"What agents are deployed? What are they doing? How are they touching customer data? Who has access to the outputs? The answer to a lot of those questions is 'I don't know' at most organizations. The first step is understanding what's going on,"  Vanta co-founder and CEO Christina Cacioppo said.
Dust provides this visibility through its governance tools.
What starts with one agent becomes a company asset when colleagues can use and improve it, the business can change the models behind it without losing what it has built, and leaders can see what the work costs and delivers. Together, those capabilities let a company build on its AI investment over time.
Your company’s AI should become more useful as people improve it. That means retaining the knowledge, instructions, and workflows they build, while keeping the ability to change the models underneath. The Dust Index tracks how far companies are getting and where the evidence challenges our expectations.

About the study

The Dust Index combines model evaluations inside Dust, analysis of aggregated and anonymized customer usage, and customer interviews. These sources serve different purposes: controlled tasks assess model performance, usage data describes observed behavior, interviews explain individual decisions.