Anthropic says it fixed Claude’s writing. I ran the evals to check.

Opus 5.5 dropped em dashes from 12.9 per 1,000 words to two in 57,000 words, and cut the rest of the Claudisms in half. Better, but not fixed.

TL;DR:

Anthropic says Opus 5.5 fixed Claude’s writing. I tested that claim with the standard eval loop in Arize AX: build a dataset, run an experiment for each model, annotate the output by hand, then build an evaluator from the annotations and run it.

The em dash really is gone: 12.9 per 1,000 words in Opus 5, and two in the entire 57,000 words of Opus 5.5 output. The rest of the Claud-isms halved. Better, but not fixed.

I’m sick of Claudisms. You know the ones. You’re scrolling LinkedIn or X, and a post tells you that something “isn’t just a tool, it’s a mindset,” or that a detail “is doing a lot of work,” and that you should sit with that. Humans don’t really write like that, but Claude does, and so does everyone who pastes Claude’s output straight into a post.

When Opus 5.5 came out, one of the launch details caught my eye. The Opus 5.5 announcement says, “We’ve made major improvements to the way Opus 5.5 writes and communicates, one of the most common areas of feedback we heard about Opus 5,” and that it “is less likely to use jargon or idiosyncratic phrases.” Anthropic staff went a step further. Sholto Douglas posted, “We fixed the writing,” and Tom Brown posted, “we fixed the accent.” The Decoder’s headline said Anthropic “promises less ‘Claudish’ writing.”

That’s a testable claim. And I work at Arize, so when someone says a model got better at something, my first reaction is “show me the eval.” Vibes don’t count, whether they’re mine or a vendor’s.

So I built claude-compare, and ran the same loop I’d use to test any change to an agent: build a dataset, run experiments on it, annotate the output, then build an evaluator and run it.

Five-step eval loop: build a dataset of 20 frozen briefs, run experiments once per model twice each, annotate the output with 153 human flags, build a span judge and em dash counter, then run the evaluator and iterate against the annotations.
Five-step eval loop: build a dataset of 20 frozen briefs, run experiments once per model twice each, annotate the output with 153 human flags, build a span judge and em dash counter, then run the evaluator and iterate against the annotations.

Step 1: Build a dataset

A dataset is the fixed set of inputs you test against. Every experiment runs over the same examples, so when a score moves, you know it was the thing you changed and not the inputs.

For a writing test, the inputs are writing tasks. I used 20 research briefs, each one containing the notes for a blog post, so that I had long form content to review for Claudisms. The topics cover five genres (opinion, technical explainer, tutorial intro, news analysis, product announcement) and five domains (AI, travel, cooking, games, books). A dataset of AI topics alone would only tell you how Claude writes about AI.

Two things make a dataset like this trustworthy:

It’s frozen. Opus 5 researched each topic once, with web search, and the briefs are committed and hash-locked. The writer gets no tools and no network, so the brief is all it has to go on.

The inputs don’t carry the style you’re testing for. The briefs are written as bullet fragments, so Opus 5’s style can’t leak into the input.

The brief is the load-bearing part of the whole setup (sorry).

Step 2: Run an experiment for each model

An experiment is one run of your task over the whole dataset, with each output stored against the example that produced it. When you compare two experiments, you’re comparing whatever differs between them, so you want that to be one thing.

Arize AX experiment list for the claude-compare dataset. Opus 5 repeats score about 4.7 to 4.9 claudism spans and about 13 em dashes per 1,000 words. Opus 5.5 repeats score about 2.3 to 2.5 claudism spans and under 0.1 em dashes.
Arize AX experiment list for the claude-compare dataset. Opus 5 repeats score about 4.7 to 4.9 claudism spans and about 13 em dashes per 1,000 words. Opus 5.5 repeats score about 2.3 to 2.5 claudism spans and under 0.1 em dashes.

If you want to compare how two models write, the first step is to just change the model. Sounds obvious, but folks often mix the model change with a prompt change as well. Changing the model first, measuring, and then changing the prompt if necessary, gives you a much clearer view on the impact of your changes.

That means pinning everything else. Both models get an identical prompt. Effort is pinned to medium, because Opus 5 defaults to high and Opus 5.5 to medium, and leaving the defaults in place would have measured two different settings. The harness also checks the served model on every response, so each experiment really is the model it says it is.

Model output varies from run to run, so I ran each model over the dataset twice. That gave four experiments (two models, two repeats), 80 posts and about 110,000 words. In AX the four experiments sit on the same dataset, so every evaluator I add later scores them all the same way and I can compare Opus 5 and Opus 5.5 side by side.

Step 3: Annotate the output

Annotations are human labels on experiment output. They’re your ground truth: what a perfect evaluator would say. You want them before you build the evaluator, because they tell you what you’re actually measuring and they give you something to check the evaluator against later.

You should label at the level you want the evaluator to work at. A score per post would tell me which posts felt Claude-ish, but not why. Marking individual phrases tells me which sentences, and an evaluator that returns phrases can be checked line by line.

The labelling tool doesn’t need to be fancy. I put all 40 Opus 5 posts into one Google Doc and read the lot, all 53,000 words, leaving a “Claudism” comment on every phrase that made me wince. That came to 153 flags across 38 of the 40 posts. The flags then went into AX as annotations on the Opus 5 runs, so they sit next to the output they describe.

Google Doc annotation of an Opus 5 post, with highlighted sentences and Jim Bennett comments labeled Claudism.
Google Doc annotation of an Opus 5 post, with highlighted sentences and Jim Bennett comments labeled Claudism.

Then I looked at what I’d flagged, and it wasn’t what I expected.

Bar chart of Jim's 153 Claudism flags grouped by move: salience flag about 48, contrast reframe about 30, verdict intensifier about 25, signpost about 20, stock metaphor about 10, gotcha framing about 7, and other about 13. A phrase list caught only 9 of the 153.
Bar chart of Jim’s 153 Claudism flags grouped by move: salience flag about 48, contrast reframe about 30, verdict intensifier about 25, signpost about 20, stock metaphor about 10, gotcha framing about 7, and other about 13. A phrase list caught only 9 of the 153.

I had a hand-written list of the famous Claudisms: “load-bearing”, “delve”, “crucially”, “genuinely,” and so on. It matched just nine of my 153 flags. “Load-bearing,” the phrase everyone jokes about, appears three times in 53,000 words of Opus 5.

What I’d actually flagged were rhetorical moves. The biggest group, about 48 of the 153, was the text telling me something was important instead of showing me why: “The interaction matters,” “a fact worth internalising,” “deserves a moment.” Next were contrast reframes (“It’s a topology, not a genre.”), verdict intensifiers (“the honest answer,” “the whole point”), and signposts that tease an insight instead of giving it (“Here’s the part that surprises people…”).

None of those are fixed wording, so no phrase list will ever catch them reliably. It’s not a phrase problem. It’s a moves problem. 🫠

Step 4: Build an evaluator from the annotations

An evaluator scores every run in an experiment automatically, so you don’t have to read 60,000 words again each time something changes. The annotations shape how you build the evaluator.

Arize AX evaluators named claudism_spans and Claudism Density, both span-scoped, created by Jim Bennett.
Arize AX evaluators named claudism_spans and Claudism Density, both span-scoped, created by Jim Bennett.

Start with the output format. My first evaluator was an LLM judge that gave each post a single 1 to 5 score for how Claude-ish it read. That’s easy to build and almost impossible to check. If it says a post is a 4, which sentences made it a 4? You can’t line a single number up against 153 human flags. So the version I kept is a span judge: it returns every Claudism it finds as an exact quote, in the same shape as my annotations, and I can match each quote against my flags to measure recall directly.

The categories come from the annotations too. The judge sorts each quote into one of six moves taken from my flags: salience flag, contrast reframe, verdict intensifier, signpost, gotcha framing (“the trap is”) and stock metaphor (“load-bearing”, “earns its keep”). And the annotations double as test cases. All three “load-bearing” sentences have to be caught, or the run fails.

A few other choices apply to almost any evaluator:

  • Use code when you can. Counting em dashes doesn’t need an LLM, so that’s a plain code evaluator with no wiggle room.
  • Don’t let a model grade its own family. The judge is OpenAI’s gpt-6-luna, so Claude isn’t marking Claude’s homework.
  • Normalise for length. Opus 5.5 writes about 8% longer, so the judge’s quotes become Claudisms per 1,000 words rather than raw counts.

Both evaluators run in AX against all four experiments.

Step 5: Run it

The em dash really is gone. Opus 5 uses 12.9 em dashes per 1,000 words. Opus 5.5 used two in its entire 57,000 words of output. As Brodie Robertson put it, “The em dashes have been deleted I repeat the em dashes have been deleted.” That result comes from a code evaluator, so there’s no judge involved and no wiggle room.

Side-by-side bars. Em dashes fall from 12.9 per 1,000 words in Opus 5 to 0.05 in Opus 5.5, a 99.6 percent drop. Claudisms fall from 4.79 to 2.38, a 50 percent drop.
Side-by-side bars. Em dashes fall from 12.9 per 1,000 words in Opus 5 to 0.05 in Opus 5.5, a 99.6 percent drop. Claudisms fall from 4.79 to 2.38, a 50 percent drop.

The other Claudisms halved. The span judge found 4.79 Claudisms per 1,000 words in Opus 5 and 2.38 in Opus 5.5, a 50% drop, and every category went down:

Category Example Opus 5 Opus 5.5 Change
Salience flag “This matters.” 1.67 0.93 −44%
Verdict intensifier “The honest answer is…” 1.10 0.38 −65%
Signpost “Here’s the part that…” 0.79 0.53 −33%
Contrast reframe “It’s a topology, not a genre.” 0.63 0.27 −56%
Stock metaphor “load-bearing”, “earns its keep” 0.43 0.13 −70%
Gotcha framing “The trap is…” 0.16 0.13 −19%
All Claudisms 4.79 2.38 −50%

Claudisms per 1,000 words, from the span judge running in Arize AX.

Salience flags, the “this matters” move, are still the most common Claudism in both models. Gotcha framing is too rare to read much into, with 8 uses against 7. The gap holds in every genre and every domain I tested.

Arize AX compare view of Opus 5 and Opus 5.5 writing the same briefs. Opus 5 outputs score higher on claudism spans, claudism count, and em dash density than the Opus 5.5 outputs beside them.
Arize AX compare view of Opus 5 and Opus 5.5 writing the same briefs. Opus 5 outputs score higher on claudism spans, claudism count, and em dash density than the Opus 5.5 outputs beside them.

Opus 5.5 also picked up some new habits. A simple scan found it uses about 2.5 times as many bold lead-in bullets as Opus 5, 35% more three-part lists, and writes 8% longer. So some of the old tics were traded in rather than dropped.

My favourite detail: Anthropic’s own prompting guide for Claude Fable 5.1 warns about “mannered prose,” using “this point earns its keep” as its example. Opus 5.5 still wrote that a stand mixer “earns its counter space” and that a searing technique “earns its place.”

So the claim holds up halfway. “Fixed” is too strong. “Much better” is fair.

Iterate on the judge prompt

The 50% figure came from the third version of the judge, not the first. An evaluator is a prompt like any other, and you rarely keep the first draft of a prompt. The annotations from step 3 are what let me iterate on it with numbers instead of gut feel.

Grouped bars of Claudisms per 1,000 words across three judge versions. Version 1 shows a 27 percent gap, version 2 a 34 percent gap with about 40 percent of Opus 5.5 spans real, and version 3 a 50 percent gap with about 87 percent of Opus 5.5 spans real.
Grouped bars of Claudisms per 1,000 words across three judge versions. Version 1 shows a 27 percent gap, version 2 a 34 percent gap with about 40 percent of Opus 5.5 spans real, and version 3 a 50 percent gap with about 87 percent of Opus 5.5 spans real.

The first version said Claudisms fell by only 27%. Moving the judge to gpt-6-luna gave 34%. Recall against my flags was high on both runs, so the judge was finding the Claudisms I’d marked.

Here’s the part that bites people. (Yes, I know, sorry.) Recall only tells you what the judge caught. It says nothing about what else it tagged. So I read a random sample of the spans it found in the Opus 5.5 posts. Plenty of them were ordinary writing. “First, a definition.” got tagged as a signpost. “Strain the context window” got tagged as a stock metaphor. “It applies to all output tokens, not only thinking.” got tagged as a contrast reframe, when it’s just being precise.

Only about 40% of the Opus 5.5 spans in that sample were real Claudisms. (That spot check was done by an AI and is small, so take it as rough). The false positives weren’t random noise. Every writer uses plain transitions and ordinary metaphors, human or model, so they formed a floor under both scores and made the two models look closer than they are.

Sit with that for a minute. (Sorry.)

The third version gave the judge a test to run on every phrase before it tags it: imagine deleting the phrase and reread the sentence. If the post loses information, such as a fact, a number, how something works or what to do, the phrase carries information and it isn’t a Claudism. If nothing is lost, it’s filler, and it counts. Delete “This matters.” and the post says exactly the same thing, so it gets tagged. Delete “not only thinking” from “It applies to all output tokens, not only thinking.” and you lose the point of the sentence, so it doesn’t.

I also added examples of what not to tag in each category and limited stock metaphors to the well-worn ones. Precision on the Opus 5.5 sample went up to about 26 in 30. Recall against my flags on held-out posts dropped from 74% to 64%, and it still caught all three “load-bearing” sentences. I’ll take that trade: an evaluator that misses a few real Claudisms beats one that counts plain English as a Claudism.

Each version ran against the same annotations, so every change came with a number for what it gained and what it cost. I tuned on half the posts and held the other half back to check the result. Without the annotations, the 27% would have looked like a perfectly good answer.

What I’d tell you before you try this

If you’re writing with Claude, hunt your own drafts for the Claudisms, more than the em dash or “delve”: telling the reader something matters, knocking down a claim nobody made, teasing a point instead of making it.

The code, the categories and the judge prompt are all in the claude-compare repo, along with a de-styling tool that uses the same categories as an editing brief.

The bigger lesson applies well beyond Claude. When someone tells you a model got better, don’t trust the vibe, and don’t trust a vendor’s word for it. Build a dataset, run the experiment, annotate the output, build an evaluator from the annotations, and iterate on it against those annotations until you trust its numbers.

And this matters. 🙃

Get the latest on AI & Observability

Sign up for our newsletter, The Evaluator—and stay in the know with updates and new resources:

Don’t ship vibes.

Arize gives AI teams observability and evals to understand and improve agent performance.