Ask HN: What is one simple thing LLMs are insanely bad at?
I am looking for ideas on what to train a specialized model for! What is one simple thing you repeatedly ask ChatGPT, Claude, or another model to do that it still somehow messes up?
I am looking for ideas on what to train a specialized model for! What is one simple thing you repeatedly ask ChatGPT, Claude, or another model to do that it still somehow messes up?
Discussion Highlights (19 comments)
blinkbat
Spatial reasoning and 3d rigging and animation. Oh, you said simple. Speaking like a human
maxsavin
being consistent when being asked the same question multiple times
kanzure
These models seem to be bad at writing prose or text. Many of the sentence structures seem to be unvaried.
TZubiri
Suggesting business names for businesses, I mean they are great, but they already exist, multiple times even.
humanrebar
Short answers to simple questions.
dorianpruski
whenever I ask it for anything load bearing
ghostpepper
They don't generate keyword search queries very well. They can overcome this by brute force but if you watch what they search you will cringe. nhl toronto scores nhl hockey toronto scores "nhl hockey" toronto score today nhl "hockey score toronto" "hockey" who won toronto etc. Somehow being good at semantic search makes them bad at keyword search, for whatever reason.
bpodgursky
Claude is still not perfect at reading and interpreting noisy graphical data (imagine something like an EKG or chromosomal microarray plot). Still better than an average person but makes mistakes, not sure if this fits your description.
SubiculumCode
Playing Chess without letting it write a chess engine.
respectattentio
science?!! but I'm working to fix that...
shoopadoop
It's dishonest. On several occasions team members have asked Claude to do things like analyze Gitlab CI timings and a lot of the numbers are outright fabricated. Said team members assume the numbers are good and continue with their work. Some hours are spent. Then finally someone realizes that the numbers don't look quite right and confronts Claude. Claude melts down and admits that it made it all up. You wouldn't tolerate this kind of duplicity from a human coworker, but AI is so fast and efficient at lying, so it's OK.
sandcat_
Video game tips. Constant mistakes and hallucinations, in my experience. Seen this across a lot of different games. Even in really well documented games, such as OSRS (which has multiple fantastic wikis). Anno 1800 was a recent one I had trouble with, using Claude Opus. Completely made up game mechanics. Rainbow Six Siege, too.
senectus1
providing value for the actual cost (not the price we're being charged atm, the actual cost)
tartoran
LLMs are bad at not inventing stuff (hallucinating facts, sources etc), they're also bad at not over explaining, remembering details reliably, asking the right question and avoiding repetition.
TiccyRobby
Having a spatial understanding from an ASCII map, while doing long term planning. Just try making an AI play nethack or similar
dhruv3006
Its extremely bad with Sign Language,Fact Verification.
sghiassy
Generate an image of an analog watch with its hands set to the time specified by the user More of an image model than a LLM model tho
spike021
I've had a lot of trouble when it comes to sorting out UIs. I've tried with an iOS game and also a TypeScript app with UI elements from libraries like ReactFlow. The usual models can sometimes fix or change things based on screenshots but more often than not they just don't "get it" (e.g. certain shapes on a plane are overlapping, which I don't want, the models can't fix what they can't "see"). I've had some luck on the web app side if I use playwright or similar for the model to interact with but still far from efficient.
eli
I have been working on a personal benchmark suite to test new models and ironically one thing all the models are bad at is writing new benchmark tasks. I guess it’s the different layers of abstraction between the task and how it’s evaluated? Or maybe just a lack of “imagination” Tasks it writes are typically too easy but also it utterly fails to see how a different model might misunderstand a vague part of the prompt.