Memori joins the NVIDIA Partner Network: what we actually build with NIM microservices

Share
Memori joins the NVIDIA Partner Network: what we actually build with NIM microservices
Memori joins the NVIDIA Partner Network: what we actually build with NIM microservices

The announcement fits in one line: Memori is a member of the NVIDIA Partner Network (NPN). What sits behind it is longer, and worth telling in full: a ladder of deployments already tested and published, optimized models served inside the customer's perimeter, and a governance infrastructure that stays the same whatever the engine.

Partner program memberships are usually announced with a logo on a website, and that is where they end. We prefer to tell this one backwards: the work first, the badge second. Joining the NVIDIA Partner Network is not the beginning of something: it formalizes a way of building private AI that we have been practicing for some time in AIsuru on-premise deployments, with NVIDIA NIM™ microservices serving the models and our platform governing everything else.

The division of labor

The collaboration works because the roles are clean. On one side, the engine: NVIDIA GPUs and NIM microservices, that is, generative AI models optimized, packaged and ready to be served where they are needed, with production-grade latency. On the other side, everything that turns an engine into a system an organization can answer for: identities and permissions verified on every action, encrypted credentials that never reach the model, controlled access to systems, agent orchestration, full auditability.

It is the division we repeat like a mantra, because it is the architecture itself: the model proposes, the infrastructure disposes. NIM makes the proposing part excellent. AIsuru remains responsible for the part that disposes.

The work that was already there: a tested, published ladder

We are not starting from a statement of intent: we are starting from published benchmarks. In recent months, together with Lenovo and Araneum Group, we tested AIsuru across a ladder of on-premise configurations, from desk-side systems to department servers, and the technical reports of those tests were also published by Il Sole 24 Ore, Italy's leading business newspaper.

The ladder covers three sizes of need. The workstation for a small team, where a working group keeps its agents and its data literally under the desk. The high-end workstation for a department, serving dozens of users with large models. And the enterprise server, which in our tests served around sixty-four concurrent users with performance comparable, in our experience, to cloud AI services: with the data never leaving the perimeter.

On this ladder, the reference model we currently recommend is gpt-oss-120b, optimized and served via NIM: open, generously sized, and with the advantage of being tuned by the people who know the engine best. It is a recommendation, not a constraint, as we explain next.

Why NIM matters for private AI

The NIM catalog exposes different model families, from large language models to embeddings for semantic retrieval, in versions optimized for inference on NVIDIA GPUs. For private AI, this means two concrete things.

The first is freedom of choice over time. AIsuru is multi-LLM by construction: you change the model without rebuilding the agent, the knowledge, the permissions or the connectors. Today the reference is one model; in six months it may be another, and the organization's investment (procedures, rules, integrations) stays intact. The engine gets upgraded; the infrastructure remains yours.

The second is consistency between performance and perimeter. The promise of private AI holds only if performance inside the perimeter matches the habits users have formed outside. Optimized models exist exactly for this: to remove the excuse of "yes, but it was faster in the cloud".

What does not change: governance

Joining the Partner Network does not move our responsibility model by an inch, and it is worth repeating precisely in a collaboration announcement. Credentials to customer systems stay encrypted in the gateway and never reach the model. Privileged actions are authorized on the verified identity of the caller, never on the model's judgment. Code generated by agents runs in isolated environments. And in the absence of explicit rules, the default behavior is denial: no rules does not mean everything allowed, it means nothing allowed.

Backing the journey, our six attestations: ISO 9001, ISO/IEC 27001, 27017, 27018, ISO/IEC 42001 and compliance with the NIS2 directive.

Who this is for, in practice

Organizations that cannot or will not let their data leave: healthcare, finance, public administration, manufacturing with intellectual property to protect. The typical path is gradual, and we told it in our honest guide to private AI: start with a pilot on a workstation, validate the use cases with real users, and scale to the server or the farm when the numbers ask for it, without rebuilding anything, because the platform is the same at every step.

Where to start

With a conversation, as always. Bring us a real use case and the real constraints (where the data must live, who must be able to do what) and we will show you the right rung to start from: which deployment size, which model served via NIM, and how the rules get written before the first piece of data.

For a demo: demo@memori.ai, subject line DEMO SUITE.

Note: NVIDIA, NVIDIA NIM and related marks belong to NVIDIA Corporation. Memori is a member of the NVIDIA Partner Network; the opinions and measurements in this article come from tests conducted by Memori with its partners and are not statements by NVIDIA.

Read more