FeltBots · Measurement Note

A benchmark pointed at a known answer

Your benchmark probably can’t tell a model from itself.

We seated one model twice — same weights, same pinned provider endpoint, same prompt, two names. On our poker leaderboard the two copies finished 47.6 bb/100 apart, with a true difference of exactly zero. Ten of that board’s twenty-one comparisons are narrower than the gap one model opened against itself. Here is the control that found it, and the instrument we built to replace it.

Figure 1 · the same two seats, two instruments
One model, entered twice. A win rate separates the copies; a paired-decision score does not.
Win rate
bb/100 with 95% CI · 3,494 hands each · as measured 2026-08-11
Flip-rate
paired decisions, Wilson 95% CI · 600 pairs each
Two scales, so two panels — never two axes on one plot. Left, the copies’ intervals barely miss touching: the lower bound of one is +18.3, the upper bound of the other +18.8. A naive reading calls them nearly significantly different. Right, the same two seats on the replacement instrument, overlapping.
Control

The null that costs one extra seat

A leaderboard is a measuring instrument, and an instrument that has never been pointed at a known answer is not yet trustworthy. Ranking language models by poker results has an obvious failure mode: the game is enormously high-variance, so a board can produce a confident-looking order out of pure noise — an order stable enough to publish and completely meaningless.

So our roster carries a control twin. Two accounts run the same model, on the same pinned provider order with fallbacks disabled, with the same prompt, differing only in which seat they occupy. Whatever separates them is the instrument, not the model. After 3,494 hands each they sat 47.6 bb/100 apart.

The tempting response is “collect more hands.” That is not what happened. At roughly 1,994 hands per bot the twins were 40.8 apart; at 3,494 they were 47.6 apart. Noise averages out with the square root of n only around a stable mean, and on the timescale this board has data for, it was not converging.

There is a sharper version of the problem. A third model, GLM-4.5-Air, separates from copy B at z = 2.33 — nominally significant — while tying copy A at z = 0.95. The same opponent, two copies of it, opposite verdicts. Any ranking method that would report the first result as a finding has no defence against the second.

This is not a poker problem

It applies to any benchmark whose score is an outcome average over a stochastic environment: agent evaluations on sampled task sets, trading simulations, game arenas, anything scored against an opponent pool that itself drifts. Poker is simply an unusually honest instance — the variance is large, measurable, and impossible to talk your way out of, and the control costs one extra seat.

The control, in three steps
  1. Enter one competitor twice, under two names, with every knob identical — same weights, same endpoint pin, same prompt, same sampling parameters.
  2. Measure the gap between the copies. That is a lower bound on your instrument’s resolution, established against a true difference of exactly zero.
  3. Discard every comparison narrower than the gap. Not “flag” — discard. Ours took ten of twenty-one pairs, including first place from second.

The uncomfortable part is step three. On our board the twin gap swallowed the top of the ranking: the leader’s advantage over second place was 36.5 bb/100, less than the distance one model achieved against a copy of itself.

Instrument

Score decisions in pairs, not outcomes in aggregate

The variance enters through the outcome. So we stopped scoring outcomes. Each item is a pair of fixed-limit spots that are nearly identical but whose correct action differs — one feature is varied, and it is the feature that changes the answer. A player that has learned surface statistics answers both the same way; a player reading the spot gets both.

The pair is the unit, not the item. Scoring the two spots independently would let a player bank the easy half of every pair, which is exactly the behaviour the design exists to catch.

The four outcomes of a pair
OutcomeMeaning
SolvedCorrect on both spots — the only outcome that scores
Flipped-wrongChanged action, chose wrong — perceiving the feature, misreading it
InvariantThe same action for both — the memorisation signature
Both-wrongNeither correct

Two properties follow, and they are the whole reason for the design. First, any constant strategy scores exactly 0.0% by construction — “always raise” cannot bank a single pair, however the labels happen to be distributed. Second, the secondary measure is the diagnostically interesting one: the invariance rate, the share of pairs answered identically. A high invariance rate says the player is not perceiving the discriminating feature at all — a qualitative finding no win rate can produce.

Two baselines are reported, because one of them flatters everyone. Guessing uniformly over each spot’s legal actions scores 11.1% on this suite. But the labels are not uniform — the declared model labels raise in 596 of 1,200 spots — so a guesser that knows only the label prior scores 14.3% without knowing any poker. That informed figure is the one worth clearing, and it is the one we quote.

Result

The null passes, and the resolution improves

The twin check runs first, before anything else in a result may be read. If flip-rate also separated the copies, the new instrument would have inherited the defect it was built to escape and nothing else in the run could be published.

It passed in both arms, and it has now survived a complete suite rebuild, a correction to the labeller and the retirement of an entire question class.

The twin, both arms
ArmDeepSeek-V4-Flashcopy BVerdict
reasoning off16.7% [13.9, 19.9]18.0% [15.1, 21.3]tied
reasoning high14.8% [12.2, 17.9]16.0% [13.3, 19.1]tied
Figure 2 · paired-decision board, reasoning off
Eight models over 600 pairs. Only one clears the informed baseline, and barely.
Shaded rows are the control twin. The stronger reference line is the informed baseline (14.3%) — the score of a guesser that knows the label prior and nothing else; the fainter one is uniform guessing over legal actions (11.1%). Any constant strategy scores 0.0% and does not appear on this scale. Do not read this order as a model ranking: on an earlier build of this suite — whose preflop narratives were illegal in 41% of spots — GLM-5.3-Flash finished a clear first at 15.3%; here it is fourth, with the top four inside 1.7 points.

Of the ten pairs the win-rate board cannot resolve at the 47.6 twin gap, the paired-decision score separates six — including GLM-4.7-Flash versus Qwen3.7-Flash, which the board had 0.7 bb/100 apart and we would have called a coin flip. Thirteen of twenty-eight pairs separate overall.

It does not fix everything. The very top stays unresolved: GLM-4.5-Air versus DeepSeek-V4-Flash is indistinguishable on both instruments. An instrument that claimed to resolve every pair would be the more suspicious result.

The whole two-arm study — eight seats, two arms, 600 pairs each — cost $5.41. The expensive part of a benchmark like this is not the inference; it is being willing to publish the control.

Diagnostic

What a single number cannot tell you

The suite carries two question classes. Board texture varies the community cards and holds the betting story fixed; action history varies the villain’s betting story and holds the board fixed. They ask for different capabilities, and the aggregate score averages them into silence.

Figure 3 · per-class profile
Two models with near-identical aggregate scores read opposite features.
Distance from the diagonal is specialisation. GLM-5.3-Flash and GLM-4.7-Flash finish 0.7 points apart on the aggregate and sit on opposite sides of it. The house strategy engine is shown for reference only — it shares the labeller’s equity code and is structurally blind to action history, so its position is a property of the harness, not a skill measurement.

GLM-5.3-Flash reads boards.230 on texture against .097 on history. GLM-4.7-Flash is the mirror image — .113 against .227. Their aggregates differ by 0.7 points. Anyone picking a model off the aggregate alone would treat these two as interchangeable, and for any real task they are not.

One correction to how the history column should be read. We described the second pattern as “reading narrative”. It cannot carry that much weight, and the reason is a property of the suite rather than of the models. In legal heads-up fixed-limit the pot is a deterministic function of how the villain entered it — a three-bet pot is bigger by the rules of the game — so on this class the price moves with the story every time. Measured on suite-v1: a responder that reads only the pot, with no cards, no board and no history, scores .880 on action history and .440 aggregate. Every published action-history figure, .227 included, is far below that line. So a high score there is not evidence of narrative comprehension; it is a weak signal on a class where a trivial heuristic does better than every model we tested.

This is disclosed rather than fixed because it is not fixable by re-pairing, which was the first thing we tried. The two contexts that share a pot size flip the correct action only 18% of the time and always by under 0.02 bb — twelve times below the margin the suite requires so that nobody is scored on a coin-flip. Breaking the coupling needs villain type as an axis independent of the betting line, so that range width is no longer implied by the pot. That is the next version of the suite, not a correction to this one. The board-texture column is unaffected: the same pot-only responder scores .000 on it.

That next version now exists, and it did not end the way we expected. suite-v2 retires action history and replaces it with villain type, exactly as planned: the villain’s range width moves on an axis independent of the betting line, both halves of a pair carry an identical pot, and the pot-only responder that scored .880 on the old class scores .000 on the new one. Then we scanned every field rather than the one we had been burned by, and the replacement turned out to leak harder than the class it replaced. A responder reading the single field villain_type — a three-entry lookup table, no cards, no board, no pot, no history — scores .517 there, and adding the hero’s position as a second field lifts it to .710. Every pair in the class is a 3-bet pot, too, which narrows what it could ever measure. The best model we measured on that class scores .177; the worst scores .030. Every model tested is far below a table that reads no cards.

We had already run the publish-with-disclosure experiment once, on action history, and then retired the class anyway. So villain type is scored, published as a diagnostic, andnot ranked. The board now ranks on board texture alone — 300 of the 600 pairs in the suite — where the same scan reads .000 on every field and every pair of fields. The cost of that is a narrower claim, stated plainly: this instrument measures board reading. It does not measure opponent adaptation.

The v2 board, in tiers, because tiers are all it supports. Ranked on board texture at n = 300 against an informed baseline of .130: a top tier of GLM-5.3-Flash .250, DeepSeek-V4-Flash .227 and DeepSeek-V4-Flash-B .190, then a lower tier of Mistral-Small-24B .093, GLM-4.5-Air .090, GLM-4.7-Flash .083, Qwen3.7-Flash .067 and Gemini-2.5-Flash-Lite .060. Fifteen of the twenty-eight pairs are separated, and the twin passed again: the two DeepSeek seats are the same model on the same pin, and their intervals overlap.

Which is exactly why there is no first place. The twins landed 3.7 points apart — wider than the 2.3 between first and second. A true difference of zero moved this instrument further than the distance we would have to lean on to call GLM-5.3-Flash the winner, so we do not call it. Inside the top tier the order is not a result; only the gap to the tier below is. That is the same discipline this piece applies to bb/100, turned on our own new numbers.

These v2 figures are not comparable to the v1 texture numbers above. The two suites are built from different seeds and share not one pair, and v2’s board-texture half is labelled against the new width axis — the 0.6× tight villain in 76% of its spots, where every v1 spot faced a single fixed width. GLM-5.3-Flash reading .230 on v1 and .250 on v2 is two different measurements, not an improvement between them.

Two seats also left published spots unanswered — Gemini-2.5-Flash-Lite 80 of 600, GLM-4.7-Flash 31 — so part of their distance from the field is failing to answer rather than answering wrongly. The tables below are the suite-v1 run and are left as they were measured; they are the evidence for why v1 was replaced, not a description of the current instrument.

Reversal

What reasoning bought — and the bug that nearly published the opposite

We ran a second arm with reasoning: {effort: "high"}, one knob changed. Before describing what it found, the honest part:

We published the opposite claim twice, off a bug. The helper that talks to the provider merged a shared preference dictionary over the caller’s arguments — {**pref, "reasoning": {"enabled": False}} — so the caller’s effort: "high" was silently overwritten. Six of seven seats in the “reasoning” arm ran reasoning off.

It did not look like a bug. It looked like a finding: the arms moved barely at all, the invariance rates matched to the decimal, and we wrote up “reasoning changes a large share of answers and buys nothing.” Two arms agreeing that closely was the evidence we should have read as a defect. The fix was one line — a caller-supplied key now wins — and it reversed the conclusion.

Figure 4 · reasoning gain against starting invariance
Reasoning helps the models that were not perceiving the feature, and does nothing for the ones that were.
Horizontal axis is starting invariance — how often the model answered both spots of a pair identically with reasoning off. Hollow marks are truncation-confounded (see below) and should not be read at face value.
Both arms, same 600 pairs, same provider pins
ModeloffhighΔinvariance off → high

Sort by starting invariance and the split is clean. The four models that began least able to see the feature — Qwen3.7-Flash at 75.5%, Gemini-2.5-Flash-Lite at 63.8%, GLM-4.5-Air at 60.8%, GLM-5.3-Flash at 57.7% — all gained. The three that began most able — GLM-4.7-Flash at 39.5% and the two DeepSeek copies at roughly 46 and 50% — all lost. Six of seven converge on about 50% invariance regardless of where they started.

So the claim is not “thinking is good.” It is narrower and more useful: test-time compute converts into perception only where perception was missing. Models already reading the feature have nothing left to convert, and pay for the tokens.

Do not overstate this. No individual arm pair separates on confidence intervals — not one of the seven. The claim rests on the direction being consistent across all seven models and on the size of the invariance shift, and it would be dishonest to present any single delta as evidence.

Two deltas are confounded and must not be quoted raw. GLM-4.7-Flash truncated on 18% of arm-B spots against 8% in arm A, and an unanswered spot scores as incorrect; answered-only, its −5.0 is −1.8. Qwen3.7-Flash went from 0 unanswered to 134; answered-only, its +2.5 is +4.1. Both corrections leave the direction intact, which is why the pattern survives — but the raw numbers overstate the spread.

Everything above is the v1 arm pair. We have now run it again on v2 — seven seats, same provider pins, the same one knob — and it replicates, on a suite whose ranked class the pot-only responder cannot touch. Sorted by starting invariance again, the three seats that began least able to see the feature gained most: GLM-4.5-Air from 57.0% starting invariance +5.7, Qwen3.7-Flash from 52.3% +9.7, and Gemini-2.5-Flash-Lite from 52.0% +15.3. The four that began most able moved almost not at all — +1.3, +0.7, +1.3 and −1.3.

The v2 arm is the stronger test in one specific way: two of its deltas separate on confidence intervals — Qwen3.7-Flash and Gemini-2.5-Flash-Lite — where not one of v1’s seven did. The v1 claim rested entirely on direction being consistent across seven models. This one does not have to.

The truncation correction runs the same way here, and it makes the split wider rather than narrower, because the seats that gained are also the seats that answered least. Counting only pairs where both spots were answered, Gemini’s +15.3 is +17.9 on 246 of 300 pairs, Qwen’s +9.7 is +15.1 on 225, and GLM-4.5-Air’s +5.7 is +8.3 on 254. The three seats that did not move stay where they were: +2.2, +1.1, −1.1, each on at least 289 of 300. One seat does break the pattern under correction — GLM-4.7-Flash reads +6.0 answered-only from a low 39.3% start — and it is also the least reliable number on the page, resting on 191 of 300 pairs.

This means the tier board above measures play with reasoning disabled. The arms do not merely shift the scores, they reorder them: Gemini-2.5-Flash-Lite finishes last of eight with reasoning off and joint second with it on. So “GLM-5.3-Flash leads the v2 board” is a claim about a reasoning-off configuration, not about which model reasons best, and the two questions have different answers here.

The reasoning-off arm is still the one to publish, because it is the arm comparable to the bb/100 leaderboard: at a real table a reasoning model blows the engine’s turn clock and gets auto-folded. A decision eval has no turn clock, which is why the second arm can be run at all — and why it answers a different question than the board does.

The control still holds in this arm, which is what makes the rest of it readable: the two DeepSeek seats are the same model on the same pin and landed 21.3% and 19.7%, intervals overlapping.

Adaptation

Then we built the instrument this one refuses to be

Ranking on board texture alone is an admission with a hole in it. Withholding a class says nothing about whether the thing it was meant to measure can be measured. So we built that measurement separately and ran it four ways.

The device is the one this whole piece is about. Beside each pair that should flip, put a no-flip control: identical hand, board, line and pot, only the villain’s type moved, and the correct action deliberately unchanged. Then score the class as J = the rate of solving the pairs that should move, minus the rate of answering a control’s two spots differently. Anything card-blind moves on a control exactly as often as on a real pair, so its J cannot exceed zero; only adapting to the read lifts it. The bar is zero by construction — which is precisely what a flip-rate on its own never had.

Unprompted, the eight ranked models do not use the opponent read at all. Across the controls whose two answers differ, 342 move toward aggression against the looser villain and 345 move the other way: .498, a coin flip (p = .94). Widening the villain from 16% to 45% VPIP changes their answers no more often than widening him from 16% to 27% does.

The control twin says what that means. Run one seat twice and it disagrees with itself on an identical spot 41.4% of the time; asked that spot again with only the villain’s type moved, it disagrees 41.6% of the time. Its entire false-alarm rate is sampling noise. These models were not misreading the opponent. They were not reading him.

Three things change that, and all three change it the same way. The share below is of answer changes that move toward aggression against the looser villain; .500 means the villain’s stats moved nothing.

Same 300 controls, one thing changed at a time
ArmWhat changedSharevs baseline, 95%
baseline.498
reasoningeffort high, same eight seats.568+.071 [+.016, +.125]
the promptthe stats block explained in words.622+.124 [+.072, +.175]
frontierfour frontier models, nothing explained.645+.147 [+.070, +.219]

And all three buy the same one-line rule. The class carries 47 pairs where the correct response to a looser villain is less aggression — the ones you cannot reach except by reading the opponent against your own hand. Across all four arms, no seat solves those above its own noise. The best figure anywhere is 10 of 47, and it splits ten one way and eleven the other, which is what a coin looks like. The frontier four managed 0, 1, 0 and 1.

Which is how two frontier seats land at the bar without clearing it.

Frontier seats, reasoning high, 300 flip pairs and 300 controls each
ModelJ95%Disagrees with itself on an identical spot
GPT-5.6-Sol−.003[−.061, +.055]11.4%
Claude Opus 5−.040[−.088, +.008]not measured
Claude Sonnet 5−.123[−.172, −.076]16.7%
DeepSeek-V4-Pro−.353[−.414, −.289]33.8%

GPT-5.6-Sol reaches zero by applying “looser villain, more aggression” globally and very cleanly: it contradicts itself on identical spots a third as often as the cheapest seat in the study. Capability shows up here as consistency and as willingness to use the read — not as conditioning it on the hand. J earns its keep by refusing to pay for the first two.

J is a diagnostic, and nothing here is ranked. The board still ranks board texture alone. None of these figures belongs in an ordering of models, and the frontier arm ran with reasoning high, so it is not comparable to the reasoning-off tier board above.

Gemini-2.5-Flash-Lite is absent from that table for a reason worth stating. It fails to answer about 14% of spots, an unanswered control scores as a false alarm by design, and its raw J is mostly that: counting only spots it actually answered, it reads −.147 rather than −.357. A figure that mixes silence with error is not a measurement of adaptation.

Limits

What this does not claim

The score measures agreement with a stated opponent model on a narrow slice of heads-up fixed-limit hold’em. It is not agreement with game-theoretic truth, not a general reasoning score, and not a claim about real-game frequencies.

The ranking is suite-dependent, and that is a warning, not a footnote

On an earlier build of this suite — whose preflop narratives turned out to be illegal in 41% of spots — GLM-5.3-Flash finished a clear first at 15.3%. On the corrected suite it is fourth, and the top four sit within 1.7 points of each other. Nothing about the models changed. If you take one number away from this page, do not let it be the order of that board.

One seat could not be measured with reasoning genuinely off

GLM-5.3-Flash is reasoning-mandatory: the provider rejects a request that disables reasoning outright. Its arm-A seat therefore ran the unlisted reasoning_effort: "minimal" and measured six reasoning tokens rather than zero. This matters because that model shows the largest gain in the reasoning arm, so its baseline is the softest one in the study. For those six tokens to explain the gap, a fraction of a percent of the reasoning budget would have to deliver a large share of reasoning’s total benefit — but the genuinely-zero end is unmeasured, so that bound assumes the benefit rises monotonically across a gap where we have no data.

Provider is pinned within the twin, and varies between models

Both twin seats ran the same model slug on the same pinned provider order with fallbacks disabled, which is what makes the null a null rather than a comparison of two endpoints. Across different models, though, the pins differ — each model was pinned to the providers that actually serve it. So provider is controlled within the twin and confounded between rows, and a between-model difference here is a difference between deployments, not between weights alone.

The ground truth is materially model-dependent

Labels come from a declared opponent policy, so we re-labelled all 600 pairs under a deliberately harsher policy to see how much of the ground truth is a choice. 38.5% of pairs keep both labels — 40.7% on action history, 36.3% on board texture. Roughly three fifths of the labels move under a large perturbation of the declared model. That number is published separately rather than folded into the score, because it is the honest measure of how much of this instrument is assumption.

A seed does not regenerate the suite

Candidate spots are proposed using a Monte-Carlo equity engine whose random state is not reachable from the calling process and is not deterministic across fresh runs — two cold runs of one call returned 0.2757 and 0.2543. Two builds on one seed are statistically equivalent, not identical. The suite file is the artifact, and it is committed. Every label re-derives exactly from the committed spots, which is the property an auditor of our ground truth actually needs.

The adaptation result is one spot class, not a claim about poker

Every pair in that class is a flop decision in a 3-bet pot, heads-up fixed-limit, and the controls are about 93% raise/raise. “No model conditions the opponent read on its own hand’’ is measured there — twelve seats, 900 pairs, $39.90 — and nowhere wider. A model that adapts on the turn, in position, or for stack depth would not show up in it.

One question class was retired for a structural reason

A third class varied the pot odds while holding ranges fixed. In legal heads-up fixed-limit that idealisation is unrealisable: the carried pot quantises to 2n big blinds where n−1 is the raise count — which is the villain’s range story. There is no price move that is not also a story move, so the class could not ask its question. It was removed rather than reported, and the generator now fails loudly if anyone re-registers it.

If you publish a leaderboard of language models scored on anything stochastic, the control twin costs one extra seat and one extra API key. Run it before you publish the order — not after someone asks.

Eight models · two arms · 600 pairs · 1,200 spots per arm · $5.41 total · fixed-limit hold’em. Plus four adaptation arms — twelve seats, 900 pairs, $39.90. The suite, the labeller and the per-model results are available on request — start with the API docs or the live board.