跳到正文
The Decoder· Matthias Bastian·· 4 小时前AI 评分57

Reka AI 全模态模型 Rho-1 在单一模型内处理文本、图像、视频与机器人控制

Reka AI's omni-model Rho-1 handles text, images, video, and robot control in a single model

AI 导读

Reka AI 发布了 19B 参数全模态模型 Rho-1 的研究预览,该模型在单个神经网络中处理和生成文本、图像、视频与机器人控制动作。所有模态作为 token 在一个共享上下文窗口中运行,无需工具调用或外部模型,并能实时生成连续视频、即时响应新指令。为弥补机器人数据稀缺,Reka AI 用逆向动力学模型从普通互联网视频提取控制信号,Rho-1 在 320 块 H100 GPU 上训练约三个月。

来源:The Decoder · the-decoder.com