The Ash AI Tax

For months I harbored a nagging suspicion I couldn’t prove. Every time I pointed a coding agent at our Ash codebase, it consumed tokens noticeably faster than when working on plain Phoenix. I initially dismissed this as mere confirmation bias and filed it away.

I finally tested this by building the same feature against the passing test suite in two stacks—idiomatic Ash and plain Phoenix + Ecto—and having an agent build it five times in each. The result was bigger than my hunch:

Building an identical feature, Claude (Opus 4.7) used +162% more tokens in Ash than in plain Ecto. Codex (GPT-5.5) used +90% more. This effect showed up on both agents.

Bar chart: Claude used 1.76M tokens on Ash vs 672k on Ecto (+162%); Codex used 1.02M on Ash vs 533k on Ecto (+90%). 20 runs, all green.

I’d have bet on 25%. It was closer to 2.6x.

Why this is worth measuring

We have crossed a line where most code is now written with an agent in the loop. Tokens are not free; they represent money, latency, and pressure on a finite context window. This means framework choice must account for a dimension nobody includes in comparison tables: how expensive is it for an AI to work in?

This is an uncomfortable question because the answer may contradict everything that makes a framework pleasant for humans. Ash is genuinely delightful for certain problems, but the core issue isn’t whether it’s good. It’s whether the feature that helps me costs the robot extra—and by how much.

The experiment

I wanted a number I could defend, not a vibe. So I pinned everything I could.

The goal was building a small Support-ticket domain: a Ticket with validations, a Comment relationship, an assignee, custom close/reopen actions with state-transition guards, and authorization (assignee or admin can close; admin only can reopen). This provided enough surface area to exercise the parts of Ash people actually use, moving beyond simple CRUD.

The key was defining the contract in a plain Elixir module boundary (Support.create_ticket/2, Support.close_ticket/2, etc.) and writing one framework-agnostic ExUnit suite against it. This boundary is idiomatic for both sides—an Ash code_interface and an Ecto context function expose the exact same six functions. Success meant the suite went green, the project compiled, and the linter was happy. No partial credit allowed.

The controls were strict:

I repeated this process using two different agents—Claude Code and Codex—because “it replicates across models” is what separates a finding from an anecdote.

The entire harness—including the specification, shared suite, both reference implementations, the runner, and raw per-run token logs—is open source for reproduction or inspection: github.com/tkolsto/ash-ai-tax-bench.

How much more does Ash cost in tokens?

About 2× more — verified across 20 runs, every one reached green:

AgentEcto (avg tokens)Ash (avg tokens)Ash taxTurns: Ecto → Ash
Claude (Opus 4.7)671,8861,758,599+162%22.4 → 31.0
Codex (GPT-5.5)533,0011,015,218+90%23.2 → 32.8

A quick disclaimer: Per-run variance is high—Codex’s Ash runs, for example, ranged from 526k to 1.44M tokens—so treat these as rough estimates (“roughly 2x”) rather than precise percentages. However, the overall direction is rock solid: on neither agent did Ash come out cheaper. Furthermore, the deltas barely budged between my initial 3-run pilot and the full 5-run matrix (Claude went from 179% to 162%; Codex went from 85% to 90%), which is exactly what you want to see before trusting a number.

Both agents needed about 40% more turns to get Ash green. This difference matters greatly, as it unravels the whole thing.

Where do the extra tokens actually go?

The most interesting part is that it wasn’t what I expected. The issue isn’t that Ash code is bigger; the agents’ diffs were comparable in size—sometimes even smaller. The cost isn’t in what the agent writes, but in what it has to understand before it can write anything correct.

I saw this concentration of effort in three places:

Context loading is a major hurdle. Frameworks like Ash require agents to pull in usage rules, resource definitions, and extensions for correct resources. Claude handles this poorly because it re-reads its cached context on every single turn. This means that a large starting context multiplied by many turns results in an exponential cost increase. That explains why Claude’s tax (+162%) significantly exceeds Codex’s (+90%): it is the same underlying phenomenon, just different token accounting.

Reasoning, not typing. The agent spent significant effort thinking about Ash, determining which declarative knob produces the desired imperative behavior for a test. Ecto simply requires you to “write a query.” Ash, however, involves considering the action, change, validation, policy, and whether it will survive atomic execution.

The “that’s not how Ash works” tax. This is the retry loop, and I have receipts. Watching the reference build, the friction was rarely domain logic; it was bending the framework to specific contracts:

None of this is hard once you know it. That’s the point: it’s a large surface of structured knowledge, and every gap in the agent’s understanding costs tokens.

So is it worth it?

I need to be fair here. The tension is this: for a human, Ash often means writing less code; for an AI, it means understanding more before writing anything. These pull in opposite directions, and which one dominates depends on who—or what—is at the keyboard.

The tax buys real value: declarative actions, policies living next to their data, generated APIs, and a mountain of correctness you don’t have to hand-roll or maintain. On a long-lived system with a team, that leverage is worth far more than a pile of tokens. Furthermore, this tax is front-loaded comprehension, not ongoing bloat—it’s the cost of the agent learning the framework’s shape, not a per-line penalty forever.

However, if you are spinning up something small, leaning hard on agents for speed, or watching your token bill, “2x the tokens, every feature” is a number that belongs in the decision, right next to all the reasons Ash is lovely.

How to shrink the Ash AI tax

The most expensive element is comprehension, but it’s also the easiest to fix: provide the agent with understanding rather than making it pay to rediscover it.

What I’d still want to know

Remember this is just a snapshot: one task, two models. Model representation will only improve, so today’s results aren’t permanent. It remains unclear if larger, multi-feature tasks will amortize or compound comprehension cost. Also, note that I measured processed tokens; because Claude bills cache reads cheaply, the dollar gap is smaller than the token gap—“2.6x the tokens” doesn’t mean “2.6x the bill.”

If you reproduce this and get a different number, please let me know. That’s what the repo is for.

The thing I actually took away

As agents become the default way code gets written, my focus shifted from Ash’s specific token cost to something larger: how comprehensible is this framework to an AI? This dimension—the ease with which a system can learn its shape in a fresh context window—is rapidly becoming a critical axis of design that we have barely started measuring.

The frameworks that will win the agent era might not be the most elegant for humans; they will be those whose structure is cheapest to learn, over and over, by an AI. This is precisely the problem usage_rules is quietly inventing solutions for.

Ash is paying a tax today. But with its commitment to usage-rules, it stands as one of the few frameworks actively building the mechanism that pays that comprehension debt back down.

The benchmark, warts and all, is on GitHub. Tell me where I’m wrong.