
Hallucinated Legibility
A benchmark on whether upscaling helps a vision model read a low-res image — and why the honest fix is the opposite of the instinct.
I spent a day trying to help an AI read a low-res screenshot. It turned out the machine wasn't the one that needed help seeing — I was.
I handed a fresh instance of Claude a 45-pixel-wide picture and asked it what it was looking at. It told me, without a flicker of doubt, that it was an analog watch dial.
It wasn't. It was a chrome Windows XP media-player skin — the "Microsoft Windows XP" wordmark, a little fan of buttons. But I hadn't given it the raw 45-pixel image. I'd run it through a super-resolution model first — one of those "enhance" networks that turn a thumbnail crisp — and the model had grown a watch face out of noise. Claude just read back what the upscaler had written for it.
Upscaling a low-res image before a vision model reads it doesn't add information. It launders a guess into evidence.
The instinct is exactly backwards
The instinct — mine, yours, everyone's — is that if the model can't read the image, you make the image bigger and sharper. It feels like cleaning your glasses. There's a whole cottage industry of AI upscalers built on that feeling: ESRGAN, Real-ESRGAN, the UltraSharp checkpoints people trade around. So I ran the benchmark I assumed would confirm the instinct, and it did the opposite, three times over.
First you have to know that a vision model doesn't look at an image the way you do. It chops it into tokens — one clump per patch of pixels — and reasons over those. The budget is finite and it scales with pixel count: Claude spends about width×height over 750 tokens on a picture and hard-shrinks anything past ~1568 pixels on a side; Gemini cuts it into 768-pixel tiles at 258 tokens each. A tiny image buys almost no tokens, so there is almost nothing there to think with.
Why size is the whole game. A vision model spends tokens in proportion to pixel area (Claude ≈ width×height ÷ 750), so a tiny image is starved — and above ~1568 pixels it downsamples, so a huge screenshot gets quietly shrunk before it's ever seen.
Here is the part that should bother you. You and the model are not looking at the same photograph. You send a razor-sharp 3000-pixel screenshot; the model quietly downsamples it to a blur before it ever "sees" anything. It's like mailing someone a painting and having the post office swap in a photocopy at the door — then arguing with them about the brushwork.
Zero gain
The first test was blind. I took one screenshot, made four versions — raw, a dumb pixel resize, and a fancy model upscale — gave them neutral filenames so nobody could cheat, and had eight fresh Claude instances describe each one with no memory of the others.
They returned identical answers. Same product, same text, same button count, same icons, straight down the line. Upscaling — pixel or model — bought exactly nothing, because the raw image was already under the model's own size ceiling. It had everything it was ever going to get from the original. I had spent real effort making the picture "better" for a reader who couldn't tell the difference.
Eight blind readers, four versions of the same screenshot, neutral filenames so none could cheat. Every one returned the same answer — because the raw was already under the model's size ceiling.
Where it stops sharpening and starts composing
So I degraded the source on purpose — 128 pixels wide, then 90, 64, 45, 32 — and ran each size raw, pixel-upscaled, and model-upscaled. That is where the shape of the thing showed up.
Down to about 128 pixels, upscaling helps: the raw image is starved for tokens, and blowing it up — even with brainless bicubic interpolation that invents nothing — hands the model more patches to chew on, and the buttons resolve. Around 64 to 90 pixels the model starts guessing glyphs; it reported a "bug," a "heart," a "play" triangle where there were none. Below 45 it composes freely — the watch dial. The raw and plainly-resized versions, meanwhile, just failed, honestly, and said they couldn't tell.
The whole result in one grid. Upscaling only helps at the top, and plain pixel-resize does it as well as the fancy model — while the model column bleeds red as the image shrinks: confident fabrication exactly where there's least to work with.
That is the trap, and it deserves a name: hallucinated legibility. And it is worse than a blur — because you can see that you can't see through a blur, whereas the fabrication arrives pre-dressed as fact.

Real outputs, not mock-ups. Below about 64 pixels the super-resolution model stops sharpening and starts composing — inventing icons that were never there, and finally a whole watch face.
None of this is new. In 2013 David Kriesel caught Xerox scanners silently swapping digits in scanned documents — a 6 quietly becoming an 8 — because a compression routine was "reconstructing" the numbers it thought it recognized. In 2023 Samsung got caught painting learned lunar texture onto blurry photos of the moon. Same sin, different decade. The academic name for the corner you get backed into is the perception–distortion tradeoff (Blau and Michaeli, 2018): past a hard limit, you cannot make an image look more real without making it less faithful. The GAN upscalers live at the "looks real" end by construction. Reading wants the other end, and no prompt drags them there.
What works is the dumb thing
Crop, don't shrink. The highest-leverage move is the least clever one — cut to the part you actually care about at native resolution, so the model's downsampler never gets a chance to smear it. And when you genuinely must make something bigger, plain bicubic — pixel interpolation that invents nothing — matched or beat the expensive AI model in every case I ran. The fancy upscaler never once won. It only tied, or lied.
I should say plainly: this was one image, read once per cell, and half of it is a logo anyone would guess. Don't oversell it into a law. But the direction is not subtle.
Trust the Blur
We built these machines to see for us, and the first thing they got good at was showing us what we wanted to be there. The blur was the honest one the whole time. It never claimed to be a watch.