This page is machine-translated from French. Read the French original.

Images, videos and voices

Beyond text, WivenLLM can produce images and short videos, read answers aloud, and transcribe speech. These functions are optional : they only exist if an administrator has configured them.


Image generation

Use

In a conversation:

/image un plan de coupe schématique d'une pompe à chaleur, style technique

The image appears in the conversation and can be downloaded. An agent can also generate an image during a task, if the corresponding skill is enabled.

Configuration — administrator

Réglages → Génération d'images. Two types of suppliers:

Local — The models run on your machine; no data is sent out. This requires suitable hardware, in practice a graphics card. Supported local installations include both an integrated engine managed by WivenLLM and image generation servers that you already use.

External — The main commercial image generation services are supported. Your descriptions are then sent to the chosen provider.

Good to know

  • The generated images are attached to the conversation.
  • Describe the style as much as the subject: "technical diagram in black and white, white background" gives a very different result from "illustration".
  • An image generation model cannot reliably write readable text; do not expect visuals with accurate labels.

Video generation

Same principle, with the command /video.

/video une vue aérienne lente d'un chantier de construction, 5 secondes

Realistic expectations: These are short sequences of a few seconds. Generation is long — from several minutes to much longer depending on the supplier and the equipment — and much more demanding than the image.

The configuration is done in Réglages → Génération de vidéos, with the same choice between local execution and external service.


Listen to the responses (text-to-speech)

A play button appears under each answer.

Available options:

  • Navigator's voice (by default) — uses the system's built-in text-to-speech function. No configuration required, no data sent anywhere, quality varies depending on the operating system.
  • High-quality local voice — a downloaded engine that runs in the browser, more natural than system voice, still without network output.
  • External services — the main speech synthesis providers are supported for superior quality; the text of the responses is then transmitted to the provider.

Configuration in Réglages → Préférences audio.


Dictating messages (voice recognition)

A microphone button in the input area transcribes your speech into text. By default, the browser's speech recognition is used.

Important confidentiality point to be aware of: Browser-integrated speech recognition typically transmits audio to the browser publisher's servers. On a system requiring a purely local environment, this is a data output that shouldn't be overlooked—check the behavior of your browser, or disable the feature.


Transcription of audio and video files

Distinct from dictation: when an audio or video file is imported as a document, it is transcribed then indexed like any other text.

The built-in transcription engine is working locally, without a key or network output. An external service can be substituted for it in Réglages → Préférences de transcription.

Common uses: Recordings of meetings, interviews, training sessions, voice messages. The transcript becomes searchable like an ordinary document — "what was decided about the budget at the meeting on the 12th?".

To record a meeting directly from the application and obtain a structured summary, see Meetings.

Quality : It depends on the clarity of the recording, the number of speakers, and any overlap in speech. A meeting recording captured by a laptop microphone will produce inconsistent results.


What comes out of the network

Function In local configuration In external configuration
Image generation Nothing The description is sent to the supplier
Video generation Nothing The description is sent to the supplier
Text-to-speech Nothing The text of the response is sent to the supplier
Voice recognition Depends on the browser — to be verified The audio goes to the supplier
File transcription Nothing The file is sent to the supplier