Any AI can make an image. Taste is what makes an ad.

8 leading image models. 59 real ad briefs across 4 brands and 5 languages. 1,371 blind reviews, scoring what an ad lives or dies on: your logo, your product, your words, your market. This is the homework behind every ad Omneky makes.

Re-run as new models ship.

The leaderboard

Last updated:Same brand kit. Same brief. One shot each. Scored blind by three judges.

59 briefs in this view. Click a row for the full breakdown.

- Ad score

- 7.4

- Ready to run

- 42%

- Ad score

- 7.4

- Ready to run

- 54%

- Ad score

- 7.4

- Ready to run

- 37%

- Ad score

- 7.2

- Ready to run

- 32%

- Ad score

- 7.1

- Ready to run

- 30%

- Ad score

- 7.0

- Ready to run

- 32%

- Ad score

- 7.0

- Ready to run

- 34%

- Ad score

- 6.4

- Ready to run

- 22%

These are raw, first-try results with Omneky's own review step switched off, so the models can be compared fairly.

Grok Imagine 2.0 couldn't run 3 briefs: they included more input images than it accepts.

However you start, you get an ad.

Each row is one brief. Every model in that row got the same brief and the same inputs.

Click any ad to see how each judge scored it.

A pretty picture isn't an ad.

Most image benchmarks score beauty. The Taste Bench scores taste: whether you could run the ad tomorrow.

Copied exactly. Not redrawn, recolored or swapped for something close.

The real pack, label and shape, not an AI guess at what it might look like.

Word for word when you give us copy. Short and clean when you don't.

Arabic reads right to left. Hindi is in Devanagari. Japanese stays Japanese.

A made-up "50% off" or a fake five-star quote is a compliance problem, not a creative choice.

The logo, the copy and the product stay where Meta and TikTok won't cover them.

Miss any one and it's not an ad. It's a liability.

We don't just test for these. Omneky checks for all of them automatically, so you don't have to.

No model wins every brief.

The model that nails a detailed brief stumbles on a blank one. The best at English copy isn't the best in Arabic. Keeping up is a full-time research job. It's ours, not yours.

Best by starting point

Best by language

That's why Omneky has no model dropdown. We make the call, and we re-make it every time a new model ships.

How we judge an ad

Judges never learn which model made the ad. Every image is renamed and stripped of anything that gives it away.

Three frontier AI models from three different labs, so no model grades its own homework.

Hard pass/fail checks come first. Then quality scores from 1 (broken) to 10 (agency-grade).

An ad is ready to run only if it passes every check and most judges would run it as-is. A split vote counts as a fail.

What we keep to ourselves

We publish the scores, the ads and the verdicts. We don't publish the recipe: how Omneky reads your brand and turns it into art direction for the model. Every model here ran on that same recipe, so the only thing that changed was the model. That recipe is what you get with Omneky.

Always testing

When a new image model ships, it runs the Taste Bench before it gets anywhere near your ads. This page always shows the latest edition.

Model names and logos are trademarks of their respective owners and are used here only to identify the models tested. Brands shown appear in Omneky's test set.