DesignLex
/人机交互/Large Multimodal Model (LMM)
论文专业

Large Multimodal Model (LMM)

大型多模态模型(Large Multimodal Model, LMM),可同时处理和生成多种模态数据(文本、图像、音频、视频等)的大规模基础模型,例如 GPT-4V、Gemini。

English Definition

"A large-scale foundation model capable of processing and generating multiple data modalities simultaneously, such as text, images, audio, and video (e.g., GPT-4V, Gemini)."

用法说明 · 针对中文母语者

与 LLM(Large Language Model,纯文本)相区分;LMM 是 CUA 得以直接'看懂'屏幕并操作 GUI 的关键能力来源。

真实用例 · 2

"The recent advent of Computer-Using Agents (CUA) based on Large Multimodal Models, which can directly manipulate graphical user interfaces, offers new opportunities for such accessibility."

论文Pipeline extracted — review needed·指出 CUA 的技术基础是 LMM

"Our paper contributes by presenting one of the first empirical studies exploring how LMM-based agents can facilitate accessibility in visually intensive online shopping contexts."

论文Pipeline extracted — review needed·强调本研究是基于 LMM 的代理在视觉密集场景的早期实证
由 pipeline 自动采集,待人工 review