Large Multimodal Model (LMM)
能同时处理和生成文本、图像、音频等多种模态信息的大规模基础模型,常作为界面操控型 AI 智能体的底座。
English Definition
"A large-scale foundation model capable of processing and generating multiple modalities (e.g., text, images, audio) within a unified representation, often serving as the backbone for embodied or interface-manipulating AI agents."
用法说明 · 针对中文母语者
LMM 与 LLM(Large Language Model)相对,但注意文献中也常写作 MLLM(Multimodal Large Language Model),本篇用 LMM。中文母语者常与「多模态大模型」混用。
真实用例 · 3 条
"The recent advent of Computer-Using Agents (CUA) based on Large Multimodal Models, which can directly manipulate graphical user interfaces, offers new opportunities for such accessibility."
"Recent advances in Large Multimodal Model (LMM) have led to the emergence of Computer-Using Agent (CUA)"
"Our paper contributes by presenting one of the first empirical studies exploring how LMM-based agents can facilitate accessibility in visually intensive online shopping contexts."