Will it run?
Models

OpenRouter launches Memebench, a meme benchmark for language models

By Rae Whitlock Clawpit staff
OpenRouter launches Memebench, a meme benchmark for language models

OpenRouter dropped Memebench yesterday. It's a new eval framework that tests whether large language models can riff on a meme template. Not factual recall. Not logic puzzles. Cultural fluency, comedic timing, and the weird visual-textual grammar of internet memes.

The setup is dead simple. Eight models got the same template and one instruction: work pingToven and Audrey_Sage_ into the joke organically. Both are known quantities in the model-research corner of Twitter. Results dropped in two batches — Group A first, Group B the next day. OpenRouter isn't saying which model maps to which letter (A through D). They handed the verdict to the community via open poll.

The prompt was identical for every model: deconstruct the original meme, find the punchline, rewrite it so both characters fit naturally. That asks for more than writing chops. It needs a feel for social dynamics, in-group references, and comedic structure — the stuff MMLU and GSM8K never touch. OpenRouter hasn't released quantitative scores, auto-eval metrics, or the exact prompts. For now it's pure human judgment.

The eight models stay anonymous. Voters pick which letter in each group "cooked harder" — community slang for a model that nailed it. Polls are open to anyone; results go live after they close. Without model identities, you can't read much into relative architecture or version strengths.

Memebench rides a wave of "soft" benchmarks probing creative and social behavior — think EvalGen for code or Arena for chat preferences. The upside: it catches capabilities academic benchmarks miss. The downside: no closed methodology, no fixed baseline, no model disclosure means you can't treat it as a valid research metric. It's more community event than serious eval tool.

OpenRouter hinted this is round one. If the format sticks, expect future cuts with more models, different meme categories, maybe some quantitative layer — human ratings on a scale, or a human baseline comparison. For now the community is voting, sharing the funny outputs, and waiting on Group B.