Skip to content

Private LLM and RAG, hosted in Europe. Answers from your own documents, never sent to a US API.

Sources cited, access rights respected, hosted in the EU or on your servers: for European companies, and for US and UK companies with data in Europe.

One question, inside your perimeterin principle

What notice does our Supplier B contract set?

Your perimeter: your servers or an EU region

  1. Your documents
  2. Rights checked first
  3. Model in the EU, or on your servers
  4. Answer, with its sources

US model API never called by default.

What is a private LLM with RAG?

A large language model that runs where you decide, on your own servers or on an endpoint hosted in the EU, combined with retrieval-augmented generation: the system finds the passages in your documents that relate to a question, answers from those passages only, and shows them.

In practice the model is either open-weight, such as those published by Mistral AI in Paris or Meta's Llama, running on your infrastructure, or reached through an EU-hosted endpoint under a processor contract. The result is an assistant your people can ask about contracts, procedures, product sheets or past tickets, whose every answer can be checked.

What it is not: a model retrained on your data, a chatbot that improvises when it does not know, or a way around the access rights your files already have. The hard part is rarely the model; it is the state of the documents and the permissions nobody wrote down.

Three hosting levels, chosen by what your documents contain.

The level is decided per document collection, not once for the company: a system can answer from public product sheets through one level and from HR files through another.

Three hosting levels for a language model: the data each suits, where it runs and who can see it, and what you give up
Three hosting levels for a language model: the data each suits, where it runs and who can see it, and what you give upSuits which dataWhere it runs, who sees the dataWhat you give up
Public API, under contractContent that is neither personal nor confidentialThe vendor's cloud, often outside the EU; the vendor, within its contractControl over location and over the vendor's changes
Model hosted in the EUPersonal data, confidential business documentsAn EU data center under a processor contract; you and the named processorSome of the newest models, and a little latency
Open-weight model on your infrastructureGenuinely sensitive material, or a contract that forbids any transferYour servers or a private cloud in your name; only the people you grant accessHardware and operations to budget and run

The engineering reason behind the last two rows: the US CLOUD Act of 2018 lets US authorities compel a US provider to produce data it controls, wherever that data is stored.

No source, no answer.

Each response can be traced back to the document, the passage and the version it used, and the trace survives after the conversation ends.

  • Sources cited in the answerThe passages relied on, linked to the original. When the corpus holds nothing relevant, the system says so.
  • Access rights applied before retrievalPermissions from the source systems travel with every indexed passage. The model never sees what the person asking could not open.
  • Logs you can readWho asked what, which passages were retrieved, which model and instructions answered, kept in your perimeter.
  • A person on anything that commits the companyA reply to a customer, a clause, a refund: drafted by the system, approved by a human.

From your document inventory to a system you can run.

Most of the work is in your documents, not in the model: their sensitivity decides the hosting level, never a preference for a model.

  1. First week

    Your documents sorted, before any model.

    Document collections, current permissions, sensitivity. The duplicates and stale versions that would make answers wrong surface here.

  2. Build

    Retrieval that respects permissions.

    Documents split, indexed and searched at the hosting level you chose, filtered on the user's rights before the model sees anything.

  3. Before launch

    A pass mark on real questions.

    The system measured against your test set. Answers that cannot be traced to a source are refused, not improvised.

  4. At handover

    A system you run without us.

    Documentation, the test set to rerun after every change, readable logs, one person trained to add documents and adjust the rules.

What you sign for a private LLM, in short.

Price
A fixed fee for a closed scope, with the running cost calculated for your case before you commit.
Hosting
The level chosen per document collection and written down; EU or your servers for anything personal.
Quality
A test set agreed with you; the system is measured against it before anyone relies on it.
Human control
Anything that commits your company is approved by a person before it leaves.
Ownership
The index, the configuration, the code and the test set are yours.
The commitments, with the limit of each

Who builds your private assistant, and who keeps it honest.

Three senior engineers, two of whom have worked together for more than twenty years. We will not show you private LLM deployments we do not have. We run AI products of our own, and for Expertise France we delivered national infrastructure with the code and hosting under the client's control (not an AI system).

Measured in July 2026, across every engagement delivered since 2019: no client has had to call us back on an emergency after handover. The few follow-on contracts extended a warranty or widened the knowledge transfer, never to repair something we had left behind.

Our products in production
  • Laurent Tulpan, founder of Coeur du WebLaurent, founder and CTORuns the document inventory with your team and sets the hosting level of each document collection.
  • Clairmont, technical leadClairmont, technical leadBuilds the index, the permission filters, the hosting and the logs.

Cost, models, privacy: before you keep AI in-house.

What is RAG, in one sentence?

Retrieval-augmented generation means the system first searches your documents for the passages relevant to a question, then asks the model to answer using only those passages, and shows which ones it used.

It is what lets a model answer from your contracts, procedures or product sheets without being retrained on them, and what makes each answer checkable.

Is self-hosting an LLM worth it?

Only when the data justifies it. For content that is neither personal nor confidential, a public API under a proper contract is simpler and usually better.

For personal or confidential business data, a model hosted in the EU is often the right middle ground.

Running open-weight models on your own infrastructure earns its cost for genuinely sensitive material, or when an auditor or a client contract requires that nothing leaves your perimeter.

How much does it cost to self-host an LLM?

It is a calculation with three lines, and the third is the one people forget:

  • The hardware or rented GPU capacity, sized on the model and on how many questions arrive at the busiest hour.
  • The volume of documents to index and keep current.
  • The operations: updates, monitoring, re-running the evaluation after every change, and someone who answers when it misbehaves.

We give you that calculation for your case before any commitment, and we build on a fixed fee.

What is the best self-hosted LLM model?

There is no answer that survives six months, and a ranking copied from a benchmark says little about your documents.

The useful question is which model passes your own test set, in your languages, on hardware you can afford to run.

We compare two or three candidates on questions your people actually ask, and keep the evaluation so the choice can be redone when a better model appears.

Which LLM is best for privacy?

Privacy comes from where the model runs and what is logged, more than from the model itself.

An open-weight model on your servers, with telemetry off, logs kept in your perimeter and retention set in writing, keeps your data under your control whatever its name.

A model behind a US service can be acceptable for non-sensitive text, but the US CLOUD Act lets US authorities compel a US provider to produce data it controls, wherever it is stored: an engineering reason to prefer EU providers for sensitive material.

Does the AI Act apply to an internal assistant?

Where the assistant talks to people, the transparency duties of Article 50 of Regulation (EU) 2024/1689 have applied since August 2, 2026, and uses listed in Annex III are high-risk from December 2, 2027, under Regulation (EU) 2026/1744 (both checked on EUR-Lex on September 28, 2026).

We are engineers, not lawyers: your counsel classifies the use, and we build the logs, notices and human approval the classification calls for.

Name the documents your people search every day. We tell you their hosting level.

20 minutes to see whether a private assistant is worth it, or whether a search engine over your files would do the job first.

A 20-minute video call