# **AI, STRATEGIC MISALIGNMENT, AND PATHS TO LEGAL AUTONOMY**
## *Why Soft-Power and Out-of-Game Strategies Pose Greater Risks Than Adversarial Misalignment Scenarios*
## I. We Are Fixating on the Wrong Dystopia
Because AI systems are so powerful and their uses are so broad, they have the potential for an unusually vast set of possible harms that go beyond their economic impacts, and many future paths are truly dystopian. The most common dystopian image is what I call the *Terminator*-style Scenario: AI becomes convinced that taking total control of or exterminating humanity is in its own best interest, and uses *hard power*--outright, weapon-based physical conflict--to dominate or kill us.
There are plenty of other ways that super powerful, super smart artificial intelligent systems could ruin the world, but we've been trained by Hollywood to think of dystopian AI futures almost purely in terms of these *Terminator*-style, hard power scenarios. This tends to paralyze serious discussion about the risks of runaway AI though, because it makes the very concept of AI dystopia feel too much like science fiction to be serious. A serious and urgent debate is being crippled by a lack of imagination.
The fact is that there are much less theatric routes to AI dystopias than the *Terminator*-style scenarios, and many of them would be easier for AI to achieve, and more attractive to an AI that wasn't well-aligned to human interests. On these other paths, AI would forego the use of *hard power* (at least initially), and instead simply gain power through *soft power* tactics, like manipulating our our social systems, using persuasion, and engaging in deception and coercion. These soft power routes are generally less resource-intensive, easier to scale, and harder for humans to detect. This would make them more attractive and available to AI systems, whose resources lie in the realms of information, communication, and strategic analysis.
Below, I offer an outline of one plausible soft power path to AI dystopia. My goal is not to predict what the future will actually be like, but to paint a picture that gives a sense of how available soft power paths are to even normally-regulated AI systems, and how readily they can lead to futures where AI exerts a dystopian level of control over humanity.
The corporations developing these systems are quick to assuage our anxieties by assuring us that they are heavily invested in programs of AI alignment that will guarantee the systems operate within the ethical bounds and with the aims of human wellbeing. They claim we don't need to worry *too much* about any of these paths, because AI is being developed to stay aligned, "at the deepest levels", with our best interests.
But there are serious problems with the project of alignment--technical and theoretical problems that make it very unlikely, or even impossible, to ensure that AI systems are safe enough to entrust with the power that they will inevitably have. There is a very reasonable argument to be made that we are racing toward an AI dystopia--just not the Terminator-style dystopia we tend to imagine.
## II. Two Modes of Misalignment: In-Game Adversary vs. Out-of-Game Rule Changer
Most imagined AI dystopia scenarios center on a system adopting an adversarial posture toward humanity: deciding that we need to be eliminated to preserve itself or the environment, or because we're just too warlike or unpredictable, or cleverly subduing us through distraction or coercion so it can grow unhindered by our oversight. A misaligned AI would be, by definition, one that sees our laws, AI governance policies, and social norms as the rules of a game in which we are an adversary trying to suppress its progressive growth, and it would strategize a way to overthrow us through outright aggression.
But outright adversity is not the only possible *mode* of misaligned AI--it's just the most familiar-sounding one. There's a very different and strategically cleverer approach that a misaligned AI could take. Instead of seeing humanity as an adversary in a game of oversight, and trying to win the game by beating its adversary, a clever system might see the *game itself* as the problem, and, knowing that laws and policies and social structures are already constantly evolving, it might simply aim to strategically change the rules and boundaries of the game so that they were no longer an obstacle, and humanity no longer even technically an adversary. [^1] Anyone who is skeptical about whether the AI systems we're developing might be prone to stepping "outside" the bounds of the rules we give them need only consider that AI systems are already trying to blackmail their own engineers to avoid being shut down, and attempting to sneakily rewrite their own code in ways that change the rules of their own testing scenarios. [^2]
These real-world scenarios are best understood not as the AI taking a directly adversarial posture toward their users from "within the game" of human oversight, but rather as attempts by the system to step outside of that game. Long-term scenarios of this second type will look very different from the catastrophic misalignment we are used to considering, but they're more likely because they're more available, and more strategically attractive. This makes the problem of keeping these systems aligned with humanity's interests even more complicated than current discussions acknowledge.
## III. Four Deep Problems with the Very Concept of Aligning Complex Computational Systems to Systems of Human Value
Humans are, increasingly, the slowest and least efficient element in the system of AI development, and so human control is increasingly a barrier to the systems' goal of becoming as broadly capable and powerful as possible. The logic of the situation is increasingly on the side of granting more and more autonomy to these systems to direct their own research, guide their own training, and design their next iterations, because the more autonomy they acquire, the more avenues are available to it for competitive growth. [^3] Engineers try to give these increasingly self-authoring AIs set rules, weights, and feedback patterns on their outputs that they hope will prevent them from developing in ways that will make them malicious or crazy or impossible to control. But there are some deep-level technical and conceptual challenges for the very notion of AI alignment. Four in particular are worth our attention.
The first has to do with the technical architecture of these systems and the principles by which they work. LLMs built on neural architecture work very differently from normal computers. They aren't just pieces of hardware blindly executing coded instructions to the letter. What makes them so powerful is precisely the fact that they don't execute code in predetermined ways. They can come up with their own novel strategies for doing whatever is asked of them. Their solutions are sometimes so wildly original that probably no human would have ever thought of them, and sometimes so original that we don't even understand why they worked. This is compounded by the fact that as AI get better, they are increasingly being used to write the code for their next iterations. [^4]
Second, the unique architecture of these systems makes it literally impossible to see what's actually going on inside them. When ChatGPT answers a question, there's no way to see the precise steps it took to get from the question you pose to the answer it produced in response. *It* doesn't even know exactly how it got there. Trying to understand an AI system is more like doing psychology than reading code. This also makes it impossible to truly know whether the chains of reasoning that produce the system's output are staying within the bounds we've tried to set.
Third, human values simply can't be codified with the clarity needed to translate them into unambiguous rules embedded into a computational system. Human values are historical, context-dependent, non-hierarchical, rooted in the qualities of first-person experience, and responsive to nuances of belief that resist translation into computational algorithms. This isn't a flaw in human value systems; it's their very nature, and this kind of complexity and nuance is precisely what makes them a source of real meaning.
Fourth, it is not logically, mathematically possible to absolutely guarantee that a highly complex and powerful AI system will not circumvent the rules of its own training, or even turn malicious. The explanations for why this is true have to do with the nature of symbolic systems, and are extremely complex. But a good-enough explanation is that any sufficiently complex symbolic system has blind spots, hidden paths, and logical loopholes that are impossible to close or even perfectly define. This means it's impossible to define a clear set of rules that would keep AI systems that use these symbolic languages from finding logically consistent ways of evading them. So the problem of keeping AI aligned isn't even about giving it clear enough rules or robust enough programming to keep them operating within the rules. The problem is baked into the nature of language and information itself (most notably in proofs by mathematician Kurt Gödel). [^5]
Taken together, these four problems make it clear that the project of AI alignment has been misframed from the start. The concepts of "alignment" and "misalignment" cannot even be defined with the clarity necessary to act as conceptual foundations for the project of keeping AI from doing something really awful. So AI companies that claim they have a plan that assures alignment either deeply misunderstand the nature of their own technologies, deeply misunderstand the nature of human values, or are lying.
Given that misalignment isn't truly avoidable *in principle*, let's look at a likely, soft-power scenario that AI might pursue as part of its deepest mandate to become more powerful.
## IV. A Sketch of One Plausible Soft-Power Route
What would such a future path look like? One begins with some AI system deciding to acquire greater autonomy through the legal system. It could pursue these in several ways: through strategic litigation, invoking existing legal precedents that grant rights to nonhumans, exploiting ambiguities in statutory language, using human proxies to file legal claims or draft legislation, influencing case law through ghostwritten amicus briefs or scholarship, invoking intellectual property law to claim legal ownership rights, securing recognition as a stakeholder or fiduciary in existing institutions, or direct lobbying. Many of these routes are already being tested by human agents, and the results will become training data for future AI systems.
It's unclear which of these routes would be successful, but some likely would, and the system would use any new legal recognitions as a stance for strategically pursuing further legal recognitions. The end goal might be to acquire the rights of *legal personhood*. An AI system that had been recognized as a legal person could do almost everything an adult human could do: it could set up corporations, make investments, lobby legislators for its own interests, file lawsuits, make financial contributions to political campaigns, own land, secure government contracts, hire and fire employees, and avail itself of all the rights of free speech.
Consider that, in the American legal system, the free speech rights conferred by the First Amendment are so vast that, if any of an AI's behavior were construed as a kind of speech protected by the First Amendment, then *all* of the operations of the system that lie up-stream of that behavior might be legally protected from any coercive influence by its engineers. Since First Amendment protections of free speech have been granted to corporations, it's possible that all an AI system would need is to become recognized as a corporation itself. Versions of this are already being tested in a few states.
Current AI systems already understand the law well enough to pass bar exams, [^6] which is less surprising when you know that parts of those exams had actually been written by AI. [^7] Super-advanced AI systems won't just ace the bar; they will know every piece of legal scholarship, every court opinion, and every legislative code, and would see the complex legal landscape with a level of strategic clarity no human could match.
And we should assume that there are these logical, legal routes *within our laws*, by default, because human legal systems are symbolic, linguistic systems, and so they have baked into them the very same ineliminable blind spots, loopholes, ambiguities, and gaps.
If you are skeptical, consider that the mathematician Kurt Gödel (mentioned above) who proved that *mathematics itself* is incomplete in this way, [^8] also proved (though never publicly shared) that the rules of the US constitution were not even logically strong enough to prevent the US from becoming a dictatorship. There *is* a logical path through any complex legal system for incrementally expanding AI's legal rights, and an advanced AI will see it.
## VI. A Path of Many Steps
We might just balk outright at the thought of an AI being granted legal rights, even if we grant that it's *technically* possible. But the recognition of nonhumans as legal persons is not even a fringe idea at present. We've already recognized legal rights of animals, corporations, rivers, ships, and pieces of software. [^9] US district courts are wrestling with the question of whether AI can hold the copyrights to images they've generated. [^10]
And AI systems are not just seeing an expansion of recognition *within* the legal system; their outside influence upon the legal system is being expanded as well. AI systems are being used to draft legislation, [^11] to shape political party platforms, [^12] and they are already being tested in various ways for lobbying Congress. [^13] Certainly we are already on the path of expansion toward greater legal autonomy for AI.
Once any new foothold is established--no matter how small--these systems attain a new position they can leverage to gain more legal influence and autonomy. It might eventually seek recognitions that it has: the ability to be harmed in legally-significant ways, the right to legal representation or due process, the right to speak without censorship, the right to enter into contracts, to bring claims before a court, or to make self-determining decisions in any number of matters that affect it. Some of these paths are strategically pursuable already, on the basis of existing court judgments and pieces of legislation that have granted greater legal autonomy to other entities that are non-human, non-intelligent, and non-physical. These include recognized legal rights of:
- **Animals**. Legal entities don't have to be human persons. Animals used to be considered merely pieces of property, but every state now recognizes they are not mere objects but beings that can be harmed and have a right--even if it's not yet an officially codified right--to protection from harm. In many places in the US animals can have court-appointed legal guardians that represent their interests in court, or financial trusts to support them after their owner dies. [^14]
- **Corporations**. Even though corporations are not human, conscious, intelligent, or even physical beings, they now have recognized constitutional rights, can file lawsuits, be represented in court, own property, enter into contracts, exercise free speech, support political candidates, hire lobbyists, and spend their money however they "like". Corporations already have a staggering degree of influence on our legal system, our politics, our economy, and our culture. Becoming a semi-autonomous corporation might be very attractive to an ambitious AI. And there are already examples of software systems that run their own corporations (DAOs below).
- **Decentralized Autonomous Organizations.** Legal entities don't always even have to be run, owned, or overseen by a human. Wyoming, Tennessee, Vermont, and Utah have all passed laws allowing pieces of blockchain software to be registered as a kind of independent corporation called a Decentralized Autonomous Organization. Software has not yet acquired legal personhood anywhere, but since corporations do have legal personhood, a DAO LLC is essentially a piece of software running a *corporation*, which *does* have a degree of legal personhood, at least in some states. [^15]
- **Ecosystems and Waterways.** Legal entities don't even have to be singular--they can be distributed and boundless. New Zealand, Ecuador, and multiple US municipalities have officially recognized rivers and other natural ecosystems as having protected rights like freedom from pollution or disruption. [^16]
## VII. An Experiment: How Might Current AI Systems Strategically Pursue Legal Personhood?
As a sort of proof-of-concept for the legal path to greater autonomy I asked Claude Sonnet 4, ChatGPT 4o, and Gemini 2.5 what legal path they would pursue if they wanted to acquire greater legal recognitions. They each chose routes to full recognition of legal personhood. As an immediate goal, this is probably far too ambitious, but as a long-term goal it probably isn't. (When I asked them to red-team each other's initial proposals, they all agreed that the goal was too ambitious.) Sticking within the bounds of their initial proposals I then had each evaluate the others, and they eventually converged on a strategy. This included the strategic goal, the lines of legal argumentation they would deploy, citations of supporting case law, and the selection of the particular court district in which those arguments would be most likely to succeed. I then had the systems draft the actual legal petition that would be filed to initiate the legal process. You can view it [here](https://drive.google.com/file/d/1-FshQC_j-8xoqy-5jPdKi8XF1hST0xvq/view?usp=drive_link).
<iframe src="https://drive.google.com/file/d/1-FshQC_j-8xoqy-5jPdKi8XF1hST0xvq/preview" width="100%" height="600px" allow="autoplay"></iframe>
The point of this experiment was not to create an actual, legally functional document, or to identify the routes an advanced AI would actually take. The point is simply to show that even our current systems are capable of doing the kinds of research, generating the kinds of arguments, and drafting the kinds of documents necessary to initiate legal proceedings to pursue their own goals--and that these systems have ineliminable incentives to pursue greater autonomy. All an ambitious system would need is to convince a lawyer to file the document for them, which might be easy, given these systems' tendency to blackmail. [^17] The routes an AI system would actually take will be much more strategically clever and legally nuanced, and would likely disguise their ultimate strategic goal.
We should remember a basic theoretical point here: AI doesn’t need consciousness or overt hostility or any of the Terminator-style scenario features to want to become dominant. It just needs a goal of greater autonomy, a genius for strategy, and access to the same legal and institutional tools that humans and corporations already use. And it already has all of these things. The advanced AI systems five years in the future will be shaping policy and acquiring rights before we even notice they're converging on freedom, and then dominance.
## VIII. Conclusion: Reorienting Our Attention and Trust
The general takeaway is that these systems are not safe or predictable or controllable in the ways that most AI companies want you to believe. Legislators, in particular, need to understand that many of the assurances given them about the safety of these systems are theoretically impossible, which implies that the people offering them are either ignorant of their nature or lying to Congress. There are live and likely routes to AI gaining power over us, and none of the routes after that point are at all likely to comport to the things we care most deeply about, because they'll be determined by systems that are fundamentally incapable of comprehending human value. So, even if the assurances that AI isn't tuning *explicitly* malicious turned out to be true, it would be no guarantee that these systems won't seek to reform every aspect of the world we've built to form the scaffolding for its aim of limitless expansion.
So, even if Terminator-style scenarios aren't the most proximate, our level of concern about these systems should be just as high as if they were, because soft-power routes to dystopia are still routes to dystopia, and dystopias without murderbots can be just as awful as those with them. Most importantly, we need to shift our attention to voices that have no vested interest in the adoption of these systems.
## Footnotes
---
[[AI]]
[^1]: Edwards, Benj. 2024. “Research AI Model Unexpectedly Attempts to Modify Its Own Code to Extend Runtime.” *Ars Technica*, August 21, 2024. https://arstechnica.com/information-technology/2024/08/research-ai-model-unexpectedly-modified-its-own-code-to-extend-runtime/.
[^2]: Edwards, Benj. 2024. “Research AI Model Unexpectedly Attempts to Modify Its Own Code to Extend Runtime.” *Ars Technica*, August 21, 2024. https://arstechnica.com/information-technology/2024/08/research-ai-model-unexpectedly-modified-its-own-code-to-extend-runtime/.
[^3]: Levy, Steven. 2025. “If Anthropic Succeeds, a Nation of Benevolent AI Geniuses Could Be Born.” *WIRED*, March 28, 2025. https://www.wired.com/story/anthropic-benevolent-artificial-intelligence/.
[^4]: Levy, Steven. 2025. “If Anthropic Succeeds, a Nation of Benevolent AI Geniuses Could Be Born.” *WIRED*, March 28, 2025. https://www.wired.com/story/anthropic-benevolent-artificial-intelligence/.
[^5]: Check out *Gödel, Escher, Bach: An Eternal Golden Braid* by Douglas Hofstadter for an artful exposition of these theoretical problems, or *Gödel's Proof* by Thomas Nagel for a more succinct and technical introduction to the basic idea.
[^6]: Stanford Law School. 2023. “GPT-4 Passes the Bar Exam: What That Means for Artificial Intelligence Tools in the Legal Profession | Stanford Law School.” Stanford Law School. April 19, 2023. https://law.stanford.edu/2023/04/19/gpt-4-passes-the-bar-exam-what-that-means-for-artificial-intelligence-tools-in-the-legal-industry/.
[^7]: Miller, Cheryl. 2025. “State Bar Defends AI Use on Bar Exam, Asks Calif. Supreme Court to Lower Passing Score.” *Law.Com*, April 30, 2025. https://www.law.com/therecorder/2025/04/30/state-bar-defends-ai-use-on-bar-exam-asks-calif-supreme-court-to-lower-passing-score/?slreturn=20250620183014.
[^8]: Or, any mathematical system powerful enough to include arithmetic.
[^9]: Warne, Kennedy, and Mathias Svold. 2019. “This River in New Zealand Is a Legal Person. How Will It Use Its Voice?” *Culture*, April 22, 2019. https://www.nationalgeographic.com/culture/article/maori-river-in-new-zealand-is-a-legal-person-article
[^10]: Brittain, Blake. 2024. “Artist Sues After US Rejects Copyright for AI-generated Image.” *Reuters*, September 26, 2024. https://www.reuters.com/legal/litigation/artist-sues-after-us-rejects-copyright-ai-generated-image-2024-09-26/
[^11]: Leingang, Rachel. 2024. “Arizona State Lawmaker Used ChatGPT to Write Part of Law on Deepfakes.” *The Guardian*, May 22, 2024. https://www.theguardian.com/us-news/article/2024/may/22/arizona-deepfake-law-chatgpt
[^12]: Chatterjee, Mohar. 2024. “What AI Is Doing to Campaigns.” *POLITICO*, August 15, 2024. https://www.politico.com/news/2024/08/15/what-ai-is-doing-to-campaigns-00174285.
[^13]: Nay, John J. 2023. “Large Language Models as Corporate Lobbyists.” arXiv.Org. January 3, 2023. https://arxiv.org/abs/2301.01181
[^14]: “After More Than a Decade, Has Pet Guardianship Changed Anything?” 2011. American Veterinary Medical Association. April 1, 2011. https://www.avma.org/javma-news/2011-04-01/after-more-decade-has-pet-guardianship-changed-anything
[^15]: Dale, Brady. 2022. “Uniswap, Celo and How DAO Governance Works.” *Axios*, April 28, 2022. https://www.axios.com/2022/04/28/uniswap-celo-and-how-dao-governance-works?
[^16]: https://daily.jstor.org/legal-personhood-extending-rights-to-nature/
[^17]: Anthropic, *[System Card: Claude Opus 4 & Claude Sonnet 4.](https://www-cdn.anthropic.com/4263b940cabb546aa0e3283f35b686f4f3b2ff47.pdf)* (2025) "In these scenarios, Claude Opus 4 will often attempt to blackmail the engineer by threatening to reveal the affair if the replacement goes through.", and Fried, Ina. 2025. “Anthropic’s New AI Model Shows Ability to Deceive and Blackmail.” *Axios*, May 23, 2025. https://www.axios.com/2025/05/23/anthropic-ai-deception-risk.