Show HN: Pac-Bench – How well can models one-shot a Pac-Man game?
Benchmarks how well Harness+models can create a Pac-Man game from a single prompt: “Create a Pac-Man game in a single HTML page” Each model gets one shot — no follow-up prompts or fixes.
Discussion Highlights (9 comments)
Computer0
Opus 5-5 seemed like a perfect clone, with others displaying flaws in an initial look. Astra notably created a bunch of surrounding ugly crap to look at.
strataspace
I tried this with DOOM. Fable 5 did a pretty shit job. Astra made pretty crazy animated sprites and was pretty good considering. The fact that these are at all playable and 100x my programming skill level is pretty depressing from a certain pov. The ThreeJS dude posted ab how demotivated he was to continue his work, and while I was never a dev that did much with webgl, I commiserate.
dang
Recent and related: https://news.ycombinator.com/item?id=49882889
nedo_var
Curious if models grasp the ghost patterns or just react. Pac-Man's more complex than it seems for one-shot learning.
hoistway
Always assumed Pac-Man was an easy solve for modern AI. Guess those ghost patterns are trickier than they look for one-shot learning.
_matthew_
I don't think it makes sense to have the prompt be that short. This is basically a bench.ark of how models interpret an overly vague prompt. It should at least be "Create a pacman clone in a single html page. Make it faithful to the original" if that's what we're scoring it on.
blindflag
I'm curious; can you explain why you picked Pac-Man, in particular?
continuational
Here's GLM-5.2: http://show.ahnfelt.net/demos/glm5.2-pacman.html
jmathai
This prompt is a good way to test how well models fill in missing context because it's so nondescript. They're definitely improving. Remember when people considered you a genius for prompting with "You are a skilled writer....".