Your AI can be turned off by the US, build your own
Access to cutting-edge AI models is becoming an American geopolitical lever. Here's what's running locally on your Mac, with 16 to 32GB RAM, and why that's your real foundation.
Update, July 2026. Fable 5 is accessible again, the directive has been lifted. That doesn’t change the thesis, quite the opposite. Cut off yesterday, restored today, by a decision you don’t control: that’s exactly what depending on a foreign dictate looks like. Your local model, on the other hand, doesn’t need anyone’s permission.
You open your Mac one morning. Your usual AI tool refuses to respond. Not a bug, not a server crash. A polite message explains that access is no longer available in your region. You haven’t done anything wrong. You’re paying your subscription. You’re perfectly in order. But you’re on the wrong side of a border, and someone in Washington has decided that’s enough.
This isn’t science fiction anymore. In June 2026, the US government ordered Anthropic, through an export control directive, to suspend access to its two most advanced models, Mythos 5 and Fable 5, for all foreign nationals, worldwide, including its own employees.
Anthropic had to shut them down for all its clients and is contesting the decision. Access to AI intelligence is becoming what oil, chips, and the dollar already are, a lever of power in the hands of a state.
And you’re at the end of the pipe that just got turned off.
The directive as a revealer
The first reaction, when you read this, is to think geopolitics. Big maneuvers, blocs that clash, topics for editorialists. Except the cut-off doesn’t happen to a bloc. It happens to you, on your machine, on a Tuesday morning.
The asymmetry is total. On one side, a supplier under US jurisdiction, subject to its administration’s directives, free to change its access terms overnight. On the other side, you, who’ve built part of your workflow around this tool. You have no say in the decision. You’re not even the target. You’re a collateral damage of a policy that only concerns you by your nationality or geolocation.
That’s what the directive finally makes visible: access to a remote model has never been a technical given, it’s a political authorization. As long as nothing changes, the authorization is tacit, invisible, comfortable. You end up confusing it with a right. The day it’s withdrawn, you discover you didn’t own anything. You were renting.
Ask yourself one question. If access is cut off tomorrow, what keeps running? If the answer is “nothing,” you don’t have a work tool. You have a dependency.
What depending on a remote model really costs
The remote model is great. That’s true, and acknowledging it doesn’t take away from the rest. It’s more powerful than anything you’ll run locally anytime soon. It updates without you lifting a finger. It doesn’t ask for RAM, installation, or maintenance. For many uses, it’s the most effective tool on the market.
But comfort has a price, and that price doesn’t show up on the bill.
Every request you send to a remote model, it’s your text leaving your machine. Your drafts, notes, documents, ongoing thoughts. They pass through infrastructure you don’t control, in a jurisdiction that isn’t yours, subject to rules you don’t vote for. End-to-end encryption protects the pipe, not the destination. At arrival, your text is read in plain sight by the machine that responds.
Then there’s the continuity dependency, the one no one talks about until it bites. Your access depends on a subscription, a pricing policy, a business decision, and now a state directive. Four levers you don’t control, plus a fourth that a foreign government controls, and that’s the most dangerous one. The day one of them flips, your workflow stops. Not gradually. All at once.
That’s exactly the trap we warned about for cloud storage. Your files on someone else’s server are convenient until the account is suspended. Remote AI is the same deal, applied to how you think and produce.
The sovereignty reflex, the map of local models
Here’s the good news, and it’s more solid than you think. Running an AI model capable of directly on your Mac is no longer a tinkerer’s fantasy. Apple Silicon’s unified memory, the very thing Apple sells for app fluidity, is precisely what makes local inference possible on a consumer machine.
Before the map, a budget rule, because everything starts there.
First, total RAM isn’t available RAM. macOS and your applications consume some of it permanently. And behind the “Mac” label, there are two very different machines.
The entry-level Mac mini comes with 16 GB of unified RAM. Subtract 6 to 8 GB for the system and your apps, and you’re left with 8 to 10 GB actually usable for AI, and that’s with closing what’s running.
A 32 GB machine follows the same principle, but leaves you around 20 GB of headroom. In this usable budget, two things need to fit at the same time: the weight of the model (the model itself) and the context (what you feed it and what it generates).
The practical rule: never aim for a model that fills up the usable budget. On 16 GB, that means a 4 to 8 GB model once quantized. On 32 GB, you can go up to 14, or even 20 GB. If you overflow, the machine falls back on the disk as memory, and then it lags, or it crashes completely.
So here’s the map of what runs, from most accessible to most demanding.
What a 16 GB Mac already runs. That’s the accessible baseline, and it covers most of your daily volume.
For light work, sorting emails, classifying notes, extracting or anonymizing text, tiny 2 to 3 GB models like Llama 3.2 3B, Gemma 3 4B, and Phi-4-mini are more than enough.
For general work, summarizing a document, drafting a memo, managing your correspondence, querying your own note base, a 7 to 8 billion parameter model around 5 GB does the job: Qwen2.5 7B, under Apache 2.0 license, or Llama 3.1 8B. That’s the workhorse of a 16 GB machine.
To read a scanned PDF or an invoice without sending it elsewhere, Qwen2.5-VL 7B (around 5 GB) or Llama 3.2 Vision 11B also work. The ceiling is a 12 to 14 billion parameter model in tight quantization, around 8 to 9 GB, like Qwen2.5 14B: it fits, but with everything else closed and a short context.
On these repetitive and well-defined tasks, local doesn’t just do the job, it does it better: zero network latency, zero requests billed, and your data doesn’t leave your machine.
What 32 GB unlocks. The jump to 32 GB isn’t a luxury, it’s what opens the next category.
For sustained writing and long document analysis, Mistral Small 3, a 24 billion parameter model under Apache 2.0 license, fits around 14 GB in its quantized version. Gemma 3 27B runs around 16 GB.
For multi-step reasoning, 32 billion parameter models fit with some squeezing: QwQ 32B, DeepSeek-R1-distill 32B, or Qwen2.5 32B occupy 19 to 20 GB in Q4 quantization. Context becomes tight, and on the most demanding reasoning tasks, the gap with a remote model remains real. Local does a lot. It doesn’t do everything yet.
On the coding side, Qwen2.5-Coder 32B nearly matches remote models for everyday programming: completing, explaining, debugging, refactoring.
And what doesn’t fit on 16 or 32 GB. Very large models like Llama 3.3 70B, or 405 billion parameter models like DeepSeek-R1 in its full version, require 64 GB of RAM or more. Don’t kid yourself otherwise.
Now, sovereignty isn’t just about raw power. Two other axes matter.
The first is provenance. Mistral is French and European, making it the natural sovereign choice for an European reader. Llama, Gemma, and Phi are American. Qwen and DeepSeek are Chinese.
A caveat here: since the models run locally, no data leaves your machine, regardless of the model’s origin. Provenance doesn’t matter for data exfiltration, it matters for trust, audibility, and not putting all your eggs in one geopolitical basket.
Running a Chinese model locally doesn’t send anything to Beijing. Choosing an European model is about arbitrating the ecosystem you support and what you can inspect, not about a risk of leakage.
The second axis is licensing. Not all “open” models are equal legally. Mistral Small 3 and several Qwen models are under Apache 2.0, a permissive license that lets you do anything, including professional use.
The community license for Llama is usable, but comes with restrictions. For a sovereign baseline, the free license isn’t a legal detail, it’s the guarantee that no one can change the rules under your feet.
My pick for sovereignty by default, in line with all this. On a 16 GB machine, a 7-8B model covers most of your daily volume: Qwen2.5 7B (Apache 2.0) as first choice, or Llama 3.1 8B. On a 32 GB machine, you upgrade to Mistral Small 3 as your daily workhorse (European, Apache 2.0), backed up by a 32B reasoning model like QwQ or DeepSeek-R1-distill, which you bring out for tougher tasks. To run them, three tools are enough: Ollama and LM Studio for simplicity, MLX for those who want to maximize Apple Silicon.
The realistic threshold, what local does well
It’s not about slamming the door on the cloud for fun. It’s about knowing what should live on your machine, and what can, eventually, leave.
Local does very well with everything that’s repetitive, sensitive, or simply daily. Sorting, summarizing, writing, correcting, classifying, anonymizing, reading your own documents. For this part of your work, which is most of the volume, you don’t need anyone. It runs without a connection, without a fee per request, without a line of your text leaving your machine.
The cloud remains the winner for the top of the pyramid, the most complex reasoning where every point of performance counts. But a supplement is still a supplement. You don’t lean your daily work on a tap someone else can turn off.
The right stance isn’t “all local” or “all cloud”. It’s partitioning. Local becomes your base, what runs by default, what your daily work and anything touching sensitive data relies on.
The cloud becomes a conscious supplement, something you use for a specific purpose, knowing what you’re sending and never sending what shouldn’t leave. The day they cut off your cloud access, you lose a supplement. You don’t lose your base.
That’s where the argument loops back. If you’re reading this on a recent Mac, you’re already on Apple Silicon. The unified memory that runs these models, you’ve already paid for it. This shift to local doesn’t wait for an investment, or a new machine, or an engineer’s skill. It waits for a decision.
Sovereignty in AI isn’t measured by the power of the model you rent. It’s measured by what keeps running the morning they cut off your access. The remote model is comfort, and power when a task demands it.
Your AI base is what lives on your machine. Install one. Run it once, on a real task, to see. You’ll know, concretely, what you have left when they turn off the tap.
The most down-to-earth question left: how do you install all this, concretely? That’s the subject of the next article, step by step, from download to first prompt, without indigestible command lines.
Read soon: Installing and running your first local model on Mac, the step-by-step guide
Sources
The trigger, US directive and Anthropic
- Anthropic, Statement on the US government directive to suspend access to Fable 5 and Mythos 5, Anthropic’s official statement, applying the export order while contesting it (12 June 2026).
- TechCrunch, Anthropic’s safety warnings may have just backfired, the global shutdown of Fable 5 and Mythos 5 (12 June 2026).
- The Register, US clampdown on Anthropic models sends EU sovereignty surge into overdrive, how the cut-off accelerates Europe’s sovereignty shift, the very angle of this article (15 June 2026).
Open-weight models cited
- Mistral AI, Mistral Small 3, the 24B Apache 2.0 European pick (30 January 2025).
- Qwen, Qwen2.5, the Qwen2.5 family, 7B variants under Apache 2.0.
- Meta, Llama 3.1 Community License, permissive license with restrictions.
Local execution on Apple Silicon
- Ollama, running models locally, data that doesn’t leave your machine.
- LM Studio, desktop app for running LLM privately, supports MLX on Mac.
- Apple, MLX, the framework optimized for Apple Silicon’s unified memory.
Technical terms? Check the glossary.