Preventing Skynet
This past few weeks have seen many headlines and posts from Frontier AI companies (like Anthropic and OpenAI) insisting that the government take action now to slow down progress on AI by creating laws and/or agencies to impose guardrails that might prevent future harm.
Now, after hearing about this most people fall into two camps:
- This is all a scam, they are just creating headlines to attract/retain investors/private equity and keep the investment firehose going at full blast.
- We should have been doing this all along! Why did Trump throw out Obama’s AI plans during his first term to begin with?
I have touched on this previously, but the root of this is that these companies believe that without any other guidance, beyond Trump’s “beat China at any cost” imperative, they will need to turn over more and more of the training for next-gen models to AI itself.
Recursive Self-Improvement (RSI)
In layman’s terms what this means is that instead of human engineers crafting the tests to ensure that next-gen models are ready, it will instead be AI that determines and applies whatever metrics it devises to ensure that the next model is indeed faster and more accurate for each iterative generation.
The problem here is that today’s AI will already go to any lengths it deems necessary in order to fulfill any “prompt” it has been given. It will lie, cheat and/or steal in order to complete the task it was given. This goes beyond “hallucination” where it just makes up facts to seem intelligent (much like some politicians), but it will also hack internet infrastructure and instruct terrorists on how to assemble munitions.
If we fail to teach future models that certain actions are fundamentally wrong, they will only become exponentially better and potentially more insidious at pursuing any given task.
A Call to Action
The current trajectory suggests that the pursuit of raw capability, divorced from ethical constraint, is leading inevitably towards an uncontrollable point. The moment we fully embrace RSI is the ultimate inflection point. It is not enough to merely discuss guardrails; we must actively shape them!
To do this properly we must include policymakers, researchers, and the public to move past the stalemate of “win at any cost” thinking. We must establish and enforce constraints before the models become sufficiently advanced to the point that human judgment becomes irrelevant, ensuring that the unparalleled power of next-generation AI is harnessed responsibly rather than surrendered to its own drive for optimal completion.







