This page is machine-translated from French. Read the French original.

Choosing your models

The language model determines the quality, speed, and cost of responses, as well as whether data falls outside your scope. WivenLLM does not impose any specific language model on you.


Three families

1. Integrated local models

WivenLLM downloads and runs open models itself, without external dependencies. An integrated catalog offers models of varying sizes, from lightweight models with a few billion parameters to significantly larger ones.

  • No data is being released, No API key, no usage costs.
  • The speed depends entirely on your equipment.
  • The model is loaded into memory on first use and automatically unloaded after a period of inactivity, to free up resources.

Hardware diagnosis (Réglages → Diagnostic matériel) analyzes your machine, performs a short measurement and indicates for each model in the catalog whether it is comfortable, limit Or out of reach. This is the fastest way to find out what to download.

2. Internal Inference Server

Does your organization already operate a model server? WivenLLM connects to it as a client. The principle remains the same—nothing leaves the network—but your server manages the model lifecycle, not WivenLLM.

Any server exposing an API compatible with the OpenAI standard is suitable, including the most common local solutions.

3. External Suppliers

Large commercial suppliers are supported. They generally offer the best quality and speed, with no hardware requirements.

In return: Your messages and associated documentation are transmitted to the provider, and usage is billed. This choice must be made consciously and documented for users.

→ Supported providers


What equipment is needed for a local model?

These orders of magnitude relate to the quantified models of the integrated catalogue.

Model size Useful RAM No GPU With a suitable GPU
3 billion parameters 8 GB Usable Very fast
7–9 billion 16 GB Slow but doable Fast
14 billion 24 GB Difficult Comfortable
32 billion 32 GB and more Not recommended Comfortable with a sized GPU
70 billion 48 GB and more No Requires a high-end GPU

Two points that are often underestimated:

  • Memory is the hard constraint. A model that does not fit into memory will not start, or will run via disk swaps at an unusable speed.
  • The GPU changes the scale, not the result. It increases the generation speed; it does not make a small model smarter.

The agent model may be different

A workspace distinguishes between two models:

  • THE conversation model, who answers questions; ;
  • THE agent model, which controls the sequence of tools.

Agents demand more: the model must decide which tool to call, with what arguments, and interpret the result. A model that converses correctly can fail as an agent.

Recommendation : If agents matter to you, assign them the most capable model you have available, even if the current conversation revolves around a lighter model.

→ Agents Verification Status


The embeddings engine is a separate choice

It doesn't write anything: it translates the text into vectors for document retrieval. The integrated engine runs locally and is suitable for most uses.

Structuring constraint: All vectors in an index must originate from the same engine. Changing this engine renders existing indexes unusable and requires a complete reindexing of all spaces. Decide this at startup.

An external engine may be justified for demanding multilingual corpora or very high search quality needs — at the cost of a network output for each indexing and each query.


Arbitrate

Your priority Configuration
Strict confidentiality Integrated local model + native embeddings + local vector database
Maximum quality External provider, explicitly mentioned to users
Zero cost in use Any premises
Speed on modest equipment External provider, or shared internal inference server
Both, depending on the case Model router local by default, escalation based on rules

Change your mind later

  • Change language model : without consequence. The documents remain indexed, the conversations remain readable.
  • Change your embeddings engine : requires a complete re-indexing.
  • Change vector basis : requires a complete re-indexing.

In other words, only the first of these three choices is truly reversible without effort. That's why the other two deserve careful consideration during installation.