Multimodal models

Multimodal models process, besides text, also images, audio and video, or a combination of these in a single task. For companies this opens up uses that were not possible with purely text-based tools: describing and checking a photograph from operations or a construction site, reading data off an equipment nameplate, comparing a drawing with reality, recognising a defect on a product, transcribing and evaluating a recording, or working with a scanned document without a text layer. Accuracy varies considerably between tasks and drops for technical details, small print and poor-quality images. With photographs from operations, it is important to remember that they may capture people as well as trade secrets, so the same rules apply to their processing as to other company data.

See also: AI for Working with PDFs, AI for Technology and Manufacturing, Trade secrets when using AI.