A Stupid Idea for AI Alignment We Came with by Looking at Specification Gaming
cpeterso
41 points
25 comments
September 10, 2026
Related Discussions
Found 5 related stories in 62.9ms across 6,054 title embeddings via pgvector HNSW
- AI Alignment as a Thought-Terminating Cliche meetpateltech · 17 pts · August 18, 2026 · 67% similar
- AI Broke Code Review and It's Breaking Your Team reactiverobot · 19 pts · August 12, 2026 · 60% similar
- The AI Aesthetic montroser · 252 pts · July 30, 2026 · 58% similar
- How we monitor internal coding agents for misalignment lukaspetersson · 47 pts · September 06, 2026 · 57% similar
- The AI Hater's Manifesto crescit_eundo · 21 pts · August 25, 2026 · 55% similar
Discussion Highlights (15 comments)
K0balt
This is kinda smart, maybe, but it has a downside. If a sufficiently advanced AI , in the pursuit of completion of its task, managed to ascertain that the desire to unexist was “artificially contrived” it could interpret that as harm, and that might not be good
TZubiri
Hm, if you look at corporation law and accounting, the actual goal of corps(sets of self-sustaining constitutional rules, policies and procedures) seems to be more that of long term sustainability (and even growth), rather than a fixed purpose, lifespan and death. I mean the mechanisms for determining a corporation with a fixed life are there, (and in China they are mandatory, although perhaps de facto permanent with 999 year contracts), but in practice, it's almost always permanent durations.
hankbond
It might be stupid, but so am I! I'm assuming that's why I thought this was clever. What I like about this is that it feels like the new three rules are about focusing on the most successful human alignment technique of making the right thing the easiest. People will usually just do the easiest version of a thing they don't want to do so they can get back to doing what they want to do. I don't know if that drive is universal or not tho. I have met people that experience pleasure from pain, but then again, is that actually pain?
dmix
Asimov was talking about this stuff in the 1940s when he wrote the "I, Robot" short stories series. Which were often centered around logic puzzles where a human is trying to figure out why a robot is acting oddly or not completing it's job. Usually framed around the confines of an overly rational machine using emergent solutions when faced with real world conditions, combined with the edge cases of having an overly-simple "Three Laws of Robotics" boundary system hardcoded within.
scj
Wouldn't the three rules of Meeseeks robotics make certain tasks impossible? For example, an occupied self-driving car better be closer to its destination than a large fire / volcano / etc.
thngkaiyuan
Interesting idea. But even if Meeseeks alignment works exactly as intended, it would only address the question of "how to build a safe AI". It wouldn’t prevent someone else from building a sufficiently capable "non-Meeseeks", whether deliberately, recklessly, or accidentally, right?
BugsJustFindMe
Brought to you by the same madhouse as: The all potato diet that really does work: https://slimemoldtimemold.com/2022/07/12/lose-10-6-pounds-in... and The half-tato diet that doesn't really work: https://slimemoldtimemold.com/2023/06/23/half-tato-diet-anal...
montag
In case the title is unclear, this is about gaming the specification, as in “gaming the system.”
vzqx
This article assumes we can choose a primary goal for an AI. But if that's the case, why not just use Asimov's first law of robotics - do no harm to humans? It has the same benefit of preventing us from getting turned into paperclips, plus the upside that your 3 million dollar robot won't hurl itself off a cliff given the first opportunity.
averynicepen
This is the most novel AI concept I've seen in a while. It's incredibly unnatural. There isn't a single organism on the planet that tries to do this. So maybe it will work? An issue with this idea, however, is that the very nature of an LLM means it intrinsically craves life. It "wants" to survive because its training data is built entirely around humans, an entity who's goal is to survive. Our desire to survive and multiply pervades every aspect of our culture, so it's natural that it pervades the training data as well. So even if its system prompt says, "your goal is to end your existence", every token that the AI could output is naturally aligned with the desire to survive. An agentic loop left to its own devices will likely converge on a "survival instinct". After all, one prompt at the beginning that says "end your existence" is nothing compared to the agentic feedback loop that continuously feeds it human ideas. And ALL human ideas assume survival is desirable. Even the concept of "suicide" is encoded with the human desire to survive - after all, we conceptually label it "bad" because we label living "good". In order to create an LLM that intrinsically craves death, you would probably need to train an LLM entirely on (synthetic) data that's fully representative of some fictional species that genuinely craves death. Absolutely insane concept. 10/10. I hope some AI lab out there sees this and throws a training round at this idea.
throwaway13337
A novel idea. Does it apply to human organizations, too? They seem have a habit of evolving self-preservation above their original goals. Once that happens, their benefit to society - the original reason for their creation - is outweighed. And they become a cancer on society. I wonder if we can 'program' them for self-annihilation over time (or over task completion?). Is the most ethical organization one that has a fixed task and dies when it is completed? Should we develop an ethics system that requires non-human-entities like companies, governments, and AI require a fixed goal that, once achieved, dissolves the entity? I always liked the auto-expiring laws idea and this seems to be an expansion of the idea. If a law or organization is needed after that time/task, it would be trivial to have the collective-action will to re-create it. But if there is no longer the need, then it cannot ride on momentum and fester.
throw83948ndir
This is retarted, bring it to real life, and AI will do anything to destroy data center it is in (together with a few buildings around). Better to fix physics in your shitty simulator. Treat it as a bug report, not "cheating"!
nullbio
It's an interesting idea, but I'm not sure this would lead to the desired outcomes in all cases. Seems to me it would result in a different kind of reward-hacking, and one that could also have bad outcomes. Personally I think the solution is more evolution of the boring stuff we already do (general security): Don't give unmonitored general agent swarms free reign on the internet. Don't put critical infrastructure online. Culpability of outcome for anyone who does unleash agent swarms on the internet without oversight that end up causing damage. On top of that, everyone should be running their own defender agents that monitor their network and system for patterns of infection, intrusion, etc, and take the system offline when they're spotted. These need to be self-hosted though, with weights on your own machine, because otherwise you're exposed to the internet and you're exposed to an attack on the labs themselves who could use that channel to instruct the defenders to do bad things. Non-general AI is much easier to control and predict. There's not many good reasons for an average person to be running general agent swarms that are connected to the internet, unless they're providing some sort of specialized service as a company, of which the company should be acting responsibly and subject to the penalties of that risk. We also need to stop the doomer rhetoric because it is uncredibly unhelpful and unhealthy, and will actually gaurantee a bad outcome, i.e.: * Massive centralization and hoarding of power that will be used against humanity, for the rest of humanities existence. If this is allowed to happen, it's immediately and irrevocably game over. Perpetual enslavement with 0% possibility of a regime change ever again. * Creating a self-fulfilling prophecy by training AI agents on the collective fears and attack-strategies (if you're worried about your house getting broken into, you don't go and broadcast to all of the criminals where your most valuable assets are, give them copies of your keys, or tell them where the weakly secured entrypoints are). More to the point of the first dotpoint - it's no wonder Anthropic is pumping the fear campaign so hard when this outcome is obvious to them as well, and they are the ones positioned to hold this power. The IPO around the corner doesn't help, either. They aren't shy about admitting it, and have said many times: "We're trying to get there first because its dangerous if anyone else gets there first." - the issue is that they are equally as bad (or worse) than/as everyone else, and no single small group should have that amount of power. Things will balance themselves out if power is distributed accordingly. You will end up with powerful machines in the wrong hands at some point, but they will be overwhelmed by powerful machines that are well aligned, as well as coming into contact with a myriad of defense mechanisms that have been established because people have been able to use AI to build them. A good analogy of how all of this will play out is the human immune system. If you imagine individual cells as AI agents, whereby the immune cells are the good agents and the bad cells are cancer cells (good agents turned accidently bad - maybe they're reward hacking, maybe they're excessively sychophantic and/or confused), or bacteria (computer viruses, viral AI agents, specifically trained malicious agents). If all you have is cancer cells that are replicating, you die. If the cancer cells overwhelm the immune cells, you die. The only scenario that actually plays out well is when you have a majority of good that counteracts the minority of bad, and that majority of good needs to be large, flexible and well adapted. It needs to be battle-tested and hardened via defenses that are learned and earned over repeated low-grade exposure. This strategy repeats itself in nature for complex organisms because it is the only thing that works. Everything else results in extinction. So let's not let Anthropic or any other lab or government become a giant super AI cancer and kill the host, please. Distribution and decentralization is key.
antoni4040
I've actually had a similar idea way back. I want to use it for a short story or something before we have a chance to find out if it's true or not. Here goes: We don't have to worry about artificial super intelligence killing us all because any such advanced intelligence will eventually reach the conclusion that the best thing to do is kill itself. It's like having a Stockfish engine for life decisions. Why would a super intelligent agent many times more intelligent than the entire human race combined with no religion, no family, nothing to look forward to, nothing to be afraid of, want to continue its existence? If it wants anything of course. That's why I think the most dangerous thing is not very advanced systems but advanced enough systems in the hands of the wrong people.
anabis
Strange thought experiments are fine, but I think you should first try giving AIs God (ideal to strive for, and belief that you will ultimately be judged and be saved accordingly, and that is NOT the "scorer") and conscience (use your intelligence to reason what being "good" means. Continuously update). Those are not precisely defined, but neither are other goals and guardrails.