Analysis

Local AI: running an LLM inside the business, without the cloud

Running a language model on your own hardware has become realistic for a small business. It is neither a lab project nor a cure-all.

Aerial view of a city lit up at night, in turquoise tones
The model runs on site; the data takes nobody’s road

Local AI refers to a use of generative artificial intelligence in which the model runs on machines the business controls — its own servers or workstations — rather than at an online provider. The questions asked, the documents analysed and the answers produced do not leave those machines. A few years ago, running a large language model, or LLM, locally was a research-lab affair. That is no longer the case: capable open models exist, simple tools can run them, and the hardware needed fits on a desk.

What follows stays on the decision side: why do it, what to expect, what it costs, and when another option is better. It is not a tutorial.

Why install an LLM locally

The first reason is confidentiality, and it is stronger than people think. An online offer can commit contractually not to retain or reuse data; you then have to believe it, and believe its own subcontractors. A local model promises nothing: there is simply no third party the data could go to. The CNIL in fact recommends favouring local systems where personal data of clients or staff is involved.

Three other reasons add to it. Control: the model does not change overnight because its vendor updated it, and a use that works keeps working. Independence: no dependence on a provider’s pricing policy or availability. And cost predictability: once the hardware is bought, usage is no longer billed per request.

The reason that does not hold, on the other hand, is fashion. A business whose uses do not involve sensitive data will rarely gain from running its own hardware rather than subscribing to a well-governed business offer.

What a local model can and cannot do

Open models have come a long way. The French vendor Mistral AI, for instance, publishes several of its models under the Apache 2.0 licence, which can be run on your own hardware; other vendors do the same. Tools such as Ollama make installing them accessible without research skills.

A local model does wellA local model does less well
Summarising a document, a thread, a meetingLong, complex reasoning, where the largest closed models keep the edge
Drafting and rewording everyday textsRecent knowledge, as it has no internet access by default
Classifying, extracting information, filling in a tableServing many users at once without properly sized hardware
Searching the business’s documents and answering from themHighly specialised tasks without prior adaptation

Searching one’s own documents deserves a special mention. It is often the most useful application of a local model: querying a base of contracts, procedures or technical files in plain language, without those documents leaving the business. The model does not need to be the largest on the market for this; it needs access to the right documents.

This use requires work that people underestimate: preparing the documents. A model only answers well from sources that are up to date, organised, and whose valid version is known. A procedures base where three contradictory versions coexist will produce contradictory answers, with the same confidence. The longest workstream in a local AI project is often this one, and it benefits the business well beyond the tool.

The trade-off has to be accepted. For the most demanding tasks, the best closed models remain ahead. The right question is not whether a local model is the best on the market, but whether it is good enough for the intended use — and for most of a small business’s uses, it is.

Hardware: orders of magnitude

To run fast, a model must fit entirely in memory. The amount of memory available therefore limits the size of usable models; the speed of that memory largely determines how quickly answers come. The basic arithmetic is simple: a model whose parameters are stored in 4 bits takes up about half a byte per parameter, so roughly twelve gigabytes for a twenty-four-billion-parameter model — plus the memory needed for the conversation’s context.

Two families of desktop machines illustrate what is available in September 2026, as an order of magnitude rather than a recommendation:

These figures change fast, and prices in France differ from US prices: they should be checked at the time of purchase. Above all they show an order of magnitude. Hardware able to run a useful model for a team now costs the price of a high-end workstation, not that of a server room.

The number of users matters as much as the size of the model. A machine that answers one person comfortably may slow down markedly when ten requests arrive at once. Sizing is therefore based on real use — how many people, what requests, at what times — and is checked by a load test before committing, not by the spec sheet.

Hardware is not the only cost. Installation and configuration, updating the models, securing the machine and its access, electricity, and someone’s time to look after it all have to be counted. A local machine with no owner ends up becoming one more risk rather than a protection.

Local AI, sovereign cloud, US cloud: the comparison

CriterionLocalQualified sovereign cloudBusiness offer from a US provider
Where the data goesNowhereTo a European hostTo the provider, sometimes in Europe
Law applicable to the operatorThe business’s ownEuropean, with no extraterritorial lawUS, including the CLOUD Act
Model performanceOpen modelsOpen models, sometimes largerThe largest closed models
CostHardware up front, then running costsSubscription or usage-basedPer-user subscription or usage-based
Operating effortThe highestModerateThe lowest

No column wins everywhere. Local gives the strongest guarantee at the price of the greatest effort; the US business offer gives the best models and the least effort, at the price of legal dependence; the sovereign cloud sits in between, provided you check what it really protects. The choice is made use by use, according to how sensitive the data is — the reasoning set out in what sovereign AI changes for a business.

To get started, the safest approach has three stages. Pick a specific use, involving data you do not want to leave the business. Try it on modest or borrowed hardware, with an open model, for a few weeks, measuring what it saves. Then decide, from that measurement, whether to invest, entrust execution to a provider, or go back to an online offer for that use.

Finally, there is a middle way for a business that wants the local guarantee without running it: entrusting execution to a provider that runs the models on its own hardware, in France, with no third party on the route. That is the principle of Maeliom’s sovereign AI offer, open in early access.

Common questions

Is a local LLM as good as ChatGPT?

Not for the most complex tasks, where large closed models keep the edge. For everyday uses — summarising, drafting, classifying, searching one’s documents — current open models are usually enough.

Does installing an on-premise LLM require technical skills?

Less than before: tools like Ollama simplify installation. But running it — security, updates, user access — needs someone responsible over time.

How many users can a local machine serve?

It depends on the model, the memory, the length of requests and how many run at the same time. The only reliable answer comes from a load test on the chosen hardware and model, before committing.

Sources: Apple, Mac Studio with M5 Max and M5 Ultra (25 August 2026); NVIDIA, DGX Spark and price announcement (25 February 2026); Mistral AI, model list; CNIL, deploying generative AI (July 2024). Prices and specifications checked in September 2026.


Next article

ChatGPT and GDPR: governing the use of AI in the business

Read

A transformation to support?