You can have the best AI model on the market and lose it overnight, without having done anything wrong!

Share
You can have the best AI model on the market and lose it overnight, without having done anything wrong!

On the night of June 12th, Italian time, the US government suspended Fable 5, the most powerful publicly available artificial intelligence model, for all non-US citizens. Not due to a malfunction. Due to a decree.

Imagine a European company that had built a critical process on top of it. It would have found itself without a model, without notice, because of a decision made in a room where it had no seat.

When I shared an initial reflection on this topic, a friend wrote me something that struck me: this, at its core, is the story of Memori in general, not just of this latest episode. Anticipating the future. He was right, and it's worth explaining why.

A history built on anticipation

Memori was founded in 2017 in Altedo, near Bologna. Even back then we were talking about agents that simulated people and completed tasks, at a time when almost no one used the word "agent" in relation to AI.

From there on, we kept building by looking at where the technology was heading, not where it had already arrived. We gave agents persistent memory, back when models still forgot everything with each conversation. We created the Board of Experts, a structure where multiple specialized agents collaborate within a single interface, before multi-agent systems became a trending topic. And we developed our own connectors to bring agents into corporate systems, connectors that have since evolved into MCP, the standard the market only adopted later.

Every time, the pattern was the same: spot a need first, and build on it before it became obvious to everyone. Fable's shutdown, rather than scaring me, confirmed our direction. Because we've been dealing with the issue of single vendor dependency in practice for years.

No backup exists? That's not the right question

In a comment, someone raised the most honest objection possible: an equivalent backup to Fable doesn't exist today, it's first by a wide margin. True. But we're conflating two different things.

Having the single most powerful model in the world is a race we will always lose, because the top performer will belong to someone else again within a few months. Not being able to be cut off at the plug is a different matter, and that's a battle we can win. In fact, we already hold the pieces.

So the question isn't "does an identical twin of Fable exist?" The question is: what do you do in the hours when that model isn't there? And above all, as a country: can we afford for the answer to depend every time on a signature in Washington?

The software for independence already exists, and it's open

In the last two months, two models have been released that we didn't have a year ago. DeepSeek V4, 1.6 trillion parameters, MIT license, the second most capable open model in the world. And NVIDIA Nemotron 3 Ultra, 550 billion parameters, released on June 4th with fully open weights, training data, and recipes, built specifically for agents that reason over long horizons.

Neither beats Fable on extreme tasks. But both can be downloaded and run on your own premises. No one can shut them down remotely.

The missing piece is the hardware. And here we stop talking in slogans, because the numbers are well known.

Three tiers of machines, with real prices

Tier 1, the departmental workstation. A model like Nemotron Super or DeepSeek V4-Flash, in quantized form, runs on a machine with a couple of high-end GPUs. Order of magnitude: 50,000 to 120,000 euros. It's the box that lets an SME secure its internal use cases without sending a single piece of data outside.

Tier 2, the serious single node. To serve DeepSeek V4-Flash at full capacity, with a one-million-token context and multiple concurrent users, you need around 170 GB of GPU memory — that's two to four H200 accelerators. A turnkey 8-H200 node costs between 350,000 and 500,000 euros today. It's the machine for a mid-sized company that means business.

Tier 3, the cluster. To serve the largest models, at 1.6 trillion parameters, you need over a terabyte of GPU memory, so eight to sixteen H200s in a cluster with fast interconnect. Starting price: 500,000 euros and up. This isn't something for a single company. It's exactly the level where a supply-chain investment (public or consortium-based) makes sense.

The tax lever that changes the math in Italy

The 180% hyper-depreciation, in effect until September 2028, applied to a 150,000-euro investment in Industry 4.0-compliant hardware and software, generates an increased depreciable base of an additional 270,000 euros. Translated into cash, with corporate tax (IRES) at 24%, that's about 64,800 euros in reduced taxes, spread over the depreciation period. The net cost of the infrastructure thus drops below 90,000 euros.

It's not an immediate discount, it's an increased depreciation allowance spread over time. But it substantially changes the economics of building in-house.

The plan, in four moves

First, train people, because without internal skills any infrastructure is just an unused monument. Second, validate use cases in the cloud, where you measure the return before buying a single GPU. Third, choose the right tier of machine based on actual volume, without over-provisioning for the sake of trend-chasing. Fourth, bring what's critical in-house, governed by a platform that decides which model to use for which task, and that performs automatic fallback when a source goes down.

But a model, alone, is not enough

There's a point that's often overlooked in these discussions. An AI agent is only as useful as the systems it can access. The model is the engine, but without a connection to company data and processes, it remains something that talks, not something that does.

That's why, at AIsuru, we have built a collection of over twenty MCP connectors, orchestrated by a proprietary gateway that manages security, identity, and permissions. Without writing a single line of code, an AIsuru agent can read and send emails in Outlook, including attachments, browse SharePoint documents while respecting each user's permissions, query CRMs like Salesforce and Dynamics 365 in natural language, read production databases — including legacy ones — create real Word and Excel files with formulas and letterhead, trigger workflows in n8n or Zapier, and autonomously execute recurring tasks through the internal Scheduler.

The piece that makes the difference between a demo and a production-ready project, however, is the one you don't see: the gateway. Every connector passes through a proprietary layer that manages per-user login, encrypted credentials and tokens, identity verification on every call, data isolation, logging, and automatic expiration. It's the difference between AI that talks and AI that acts, within the security boundaries a company demands.

In conclusion

Technological sovereignty doesn't mean having the most powerful model in the world, because that will always belong to someone else for a few months. It means not being able to be cut off at the plug. It means open models on our own hardware, governed by a platform that orchestrates, connects, tracks, and enforces corporate policies.

And here's the most important point: for the first time, we already have, today, on the market, all the pieces needed to get started. Solid open models, purchasable hardware, tax incentives in force. All that's missing is the decision to begin.

My view, as a European and as an Italian, is that it should be made now.

If you want to understand what an agent connected to your systems could do, and on what infrastructure, write to us. The most interesting use cases always start with a concrete question.


#AI #DigitalSovereignty #AIGovernance #MultiLLM #OnPremise #EnterpriseAI #AIAct #Industry40 #MCP #DeepSeek #Nemotron #AIsuru #Memori

Read more