论文专业
Large Multimodal Model (LMM)
大型多模态模型(Large Multimodal Model, LMM),可同时处理和生成多种模态数据(文本、图像、音频、视频等)的大规模基础模型,例如 GPT-4V、Gemini。
English Definition
"A large-scale foundation model capable of processing and generating multiple data modalities simultaneously, such as text, images, audio, and video (e.g., GPT-4V, Gemini)."
用法说明 · 针对中文母语者
与 LLM(Large Language Model,纯文本)相区分;LMM 是 CUA 得以直接'看懂'屏幕并操作 GUI 的关键能力来源。
真实用例 · 2 条
"The recent advent of Computer-Using Agents (CUA) based on Large Multimodal Models, which can directly manipulate graphical user interfaces, offers new opportunities for such accessibility."
论文Pipeline extracted — review needed·指出 CUA 的技术基础是 LMM
"Our paper contributes by presenting one of the first empirical studies exploring how LMM-based agents can facilitate accessibility in visually intensive online shopping contexts."
论文Pipeline extracted — review needed·强调本研究是基于 LMM 的代理在视觉密集场景的早期实证
由 pipeline 自动采集,待人工 review