NLNEXT LAB RADAR
← 返回首页
话题

多模态

463 条相关资讯 · 来自历史归档

行业信号
8/14 15:13
阿里开源 Qwen3.8-27B 模型,编程、办公场景表现超越 Qwen3.7-Plus

IT之家 8 月 14 日消息,今天(14 日)晚间,阿里“千问大模型”公众号宣布,Qwen3.8-27B 模型正式开源,所有开发者、科研机构和企业均可自由下载、部署和使用。 IT之家附体验地址: Hugging Face 魔搭社区 根据介绍,27B(270 亿参数)是 全球 AI 社区呼声最高 的模型尺寸,Qwen3.8-27B 是原生多模态稠密(Dens…

AI 点评 · 开源27B尺寸兼顾性能与部署成本,编程办公双场景越级表现,或成中小开发者首选。

IT之家
设计 / 产品NEW
8/14 03:35
oil-oil/dsh-vision

Near-native image understanding for DeepSeek Harness

可信度 74交叉信源 1
GitHub
设计 / 产品
8/13 18:57
ysr666/dsh-vision-router

Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG t…

可信度 74交叉信源 1
GitHub
设计 / 产品
8/10 14:32
cobusgreyling/Muse-Glimmer

Introducing Muse Glimmer: open-weight 30B agentic multimodal model that runs on your device (Meta). Interactive local agent lab + guide. Apache 2.0 · on-device…

GitHub
Skill / 资源
8/10 05:44
not much happened today

**Meta** re-enters the open-weight frontier with the release of **Muse Glimmer**, a **30B dense**, multimodal, agent-focused model under **Apache 2.0**, optimized for always-on loc…

AI News
设计 / 产品
8/9 11:01
unknowlei/minimax-h3-opencode-skills

OpenCode skill suite for MiniMax H3 directing, routing, multishot planning, prompt generation, and review.

可信度 74交叉信源 1
GitHub
论文 / 方法
8/6 20:00
An AI4AI Framework for Visual Token Pruning

Visual-token pruning can substantially reduce the inference cost of multimodal large language models (MLLMs), yet existing methods largely rely on fixed, handcrafted heuristics and costly expert trial…

可信度 74交叉信源 1
HuggingFace Papers
设计 / 产品
8/5 06:51
KuaaMU/mcp-vision-bridge

MCP server that gives text-only LLM coding agents vision — analyze images via any multimodal model (mimo, Claude, Gemini, OpenAI-compatible). Works with Claude…

GitHub
模型 / Agent一手源
8/4 12:00
Introducing Shieldstral.

Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size.

可信度 92交叉信源 2
Mistral AI
模型 / Agent
8/4 06:32
Kimi K3与DeepSeek V4之间,隔着原生多模态的时间差

文 | 李炤锋 编辑 | 张雨忻 “长链任务如果只通过代��层面的反馈,误差可能会不断累积,最终效果会非常差。”谈及原生多模态的意义,一位多模态研究员表示,“视觉是一种更准确的反馈,也更贴近用户意图。” 过去一年,Coding与Agent能力不断改写大模型的排名,也成为AI最快兑现商业价值的场景之一。与此同时,随着Agent开始接管更多长链任务,越来越多的通…

36氪
Skill / 资源
8/4 05:44
not much happened today

**Alibaba** launched **Qwen3.8-Max**, enhancing multimodal capabilities and agent ecosystem integration. **NVIDIA** introduced **Alpamayo 2 Super** for autonomous vehicle reasoning…

AI News
模型 / Agent
8/3 05:44
Qwen 3.8 Max

**Alibaba** launched **Qwen3.8-Max**, a **2.4T-parameter** open-weight model emphasizing autonomous coding, long-horizon execution, and multimodal feedback, with aggressive pricing…

可信度 78交叉信源 2
AI News
行业信号
7/31 15:23
国内唯一做多模态长记忆的公司,融资数千万,押注主动智能|涌现新项目

文|王欣逸 编辑|张雨忻 一句话介绍 国内唯一做多模态长记忆的公司——丘脑智能,推出原生多模态记忆基座,押注AI从通用走向个性化,最终走向主动智能。 主动智能,指的是AI能在足够了解用户的基础上,在合适的时间、以恰当的方式主动跟用户交互。要实现主动智能,Memory是必须要跨过的门槛。 融资情况 近日,丘脑智能已完成数千万元种子轮融资,投资方包括深圳一线基金…

AI 点评 · 多模态长记忆是主动智能关键门槛,资本押注稀缺赛道,看点十足。

36氪
论文 / 方法一手源
7/29 17:32
Anatomy Contextualized Adaption of CT Foundation Models

CT vision-language foundation models have demonstrated promising performance across downstream tasks, but are typically trained with whole-volume representations that dilute fine-grained anatomical si…

arXiv
论文 / 方法
7/28 20:00
Metis: Memory Foundation Model

Recent advances in AI agents have increasingly internalized native capabilities into their underlying foundation models, giving rise to multimodal foundation models and large reasoning models. However…

HuggingFace Papers
论文 / 方法
7/27 20:00
Shieldstral

We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7times its size on text safety benchmarks and sets a new state of the ar…

HuggingFace Papers
论文 / 方法一手源
7/27 17:46
ERUnderstand: Evaluating Vision-Language Models on Structured ER Diagrams

Entity-Relationship Diagrams (ERDs) are central to conceptual database design, yet they are typically available only as rendered images rather than machine-readable schemas, limiting AI-assisted datab…

AI 点评 · 评测视觉语言模型理解结构化ER图的能力,填补了AI辅助数据库设计的评估空白。

arXiv
论文 / 方法
7/26 20:00
Data Pyramid for Embodied Manipulation

Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states…

HuggingFace Papers
设计 / 产品
7/24 10:16
亚马逊升级 Alexa+ AI 助手,可完成购物、餐厅预订、叫车等任务

IT之家 7 月 24 日消息,亚马逊官方今日宣布升级 Alexa+ AI 助手, Alexa+ 支持自然对话、信息记忆和任务执行 ,可帮助用户管理日程、总结邮件、生成播客、控制智能家居,并完成购物、餐厅预订、叫车等日常任务。 Alexa+ 已从单纯的语音助手升级为一个具备多模态能力(视觉识别、生成式 AI 创作)、超强记忆力,并打通了众多第三方现实服务接口…

AI 点评 · AI助手从对话升级到行动,打通现实服务,智能助手终于能真正“办事”了。

IT之家
行业信号
7/24 00:00
8点1氪丨段永平称10年内大概率不会卖泡泡玛特;中国数学家王虹、邓煜获得菲尔兹奖;宜家回应甩卖8处物业:不代表退出中国市场

今日热点导览 混元多模态理解负责人胡瀚离职创业,原团队或将聚焦世界模型 极氪回应“海外锁车”事件 客服回应滔搏暴力打折甩卖耐克库存:没有收到降价通知 哈兰德和亚马尔2.2亿欧元身价破纪录 张雪峰女儿再接手三家公司股份 TOP 3 大新闻 段永平:10年内大概率不会卖泡泡玛特 7月23���,段永平在社交媒体平台雪球上发表了他对近期投资操作的最新想法。雪球上有…

AI 点评 · 段永平罕见长线看好,泡泡玛特投资逻辑获大佬背书,市场风向标意义显著。

36氪
论文 / 方法
7/23 17:59
3D-Aware VLMs with Implicit and Explicit Geometries

Despite rapid progress, most existing vision-language models (VLMs) built from 2D visual inputs often struggle when handling various 3D tasks that require fine-grained spatial understanding and reason…

AI 点评 · 将隐式与显式几何融合进视觉语言模型,突破2D局限,实现精细3D空间推理。

arXiv
论文 / 方法
7/23 17:35
MIRROR: Learning from the Other View for Multi-Modal Reasoning

Unlike large language models (LLMs) that exhibit strong reasoning capabilities, vision-language models (VLMs) struggle with visual reasoning, even on geometry problems that admit equivalent text, diag…

AI 点评 · 多模态推理新突破,利用跨视角学习提升视觉语言模型几何问题解决能力。

arXiv
行业信号
7/23 08:07
独家|混元多模态理解负责人胡瀚离职创业,原团队或将聚焦世界模型

文 | 周鑫雨 编辑 | 张雨忻 《智能涌现》独家获悉,近期,腾讯混元多模态理解负责人胡瀚提出了离职。 此前,他曾担任微软亚洲研究院视觉计算组首席研究员。2025 年初加入腾讯后,负责视觉大模型的研究。在后续的调整中,他加入大语言模型部旗下的“Frontier”前沿技术研究组,负责多模态理解的相关研究,汇报给姚顺雨。 据了解,胡瀚还曾承担世界模型的研发工作。…

AI 点评 · 顶级多模态人才离职创业,折射出世界模型赛道竞争白热化。

36氪
论文 / 方法
7/21 20:00
Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning

Reinforcement learning with verifiable rewards (RLVR) has substantially improved language-model reasoning, yet its extension to vision-language models remains constrained by the lack of training data…

AI 点评 · 用分类学引导多域视觉推理,突破强化学习的训练数据瓶颈,值得关注。

HuggingFace Papers
论文 / 方法
7/21 20:00
ReferTrack: Referring Then Tracking for Embodied Visual Tracking

Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural language using only onboard vision. While recent vision-language-action (VLA) polic…

AI 点评 · 将自然语言描述与移动追踪结合,突破传统视觉追踪限制,提升具身智能的实用性与交互性。

HuggingFace Papers
Skill / 资源
7/21 14:22
Nativ: Run AI models locally on your Mac

Nativ: Run AI models locally on your Mac Prince Canuma is the developer behind the excellent MLX-VLM Python library for running vision-LLMs using MLX on a Mac. I'm really excited a…

AI 点评 · 本地运行AI模型,保护隐私且无需联网,苹果用户的新利器。

Simon Willison
设计 / 产品
7/21 07:32
Today-Hbw/rag-platform

Enterprise multi-source RAG platform with multimodal embedding, hybrid vector/BM25 retrieval, MCP & RBAC support

GitHub
行业信号
7/17 10:45
2026最受投资人关注人工智能/具身智能企业50揭晓

人工智能正在进入一个新的产业周期。 过去一年,大模型能力持续演进,生成式AI、多模态交互、智能体等技术方向快速推进;而具身智能也从早期的技术探索阶段,逐渐步入产业验证的深水区,机器人开始成为人工智能与现实世界的重要载体。 市场率先给出了回应。据36氪研究院测算,中国具身智能市场规模已从2018年的2133亿元增长至2025年的9150亿元,2026年有望突破…

36氪
行业信号
7/17 07:22
腾讯发布具身 VLM 基座模型 Hy-Embodied-VLM-1.0,A3B 规模整体性能接近上一代 A32B 模型

IT之家 7 月 17 日消息,腾讯 Robotics X 实验室、福田实验室联合腾讯混元打造的第二代具身 VLM 基座模型 Hy-Embodied-VLM-1.0 昨日正式发布。 官方表示,在覆盖 37 个评测任务的具身能力评测体系中,Hy-Embodied-VLM-1.0 在物理状态理解、动作 — 变化推理、时序与自适应推理三大维度分别取得 68.6、6…

AI 点评 · 参数规模锐减十倍,性能逼近上一代大模型,具身智能走向高效实用化。

IT之家
行业信号
7/17 02:39
36氪首发 | 港科大博士创业做机器人全身触觉系统,红杉、瓴智、智元共同押注

作者 | 乔钰杰 编辑 | 袁斯来 硬氪获悉,全身多模态融合触觉解决方案公司模感科技(MoSense)近日完成数千万元天使轮融资,投资方包括红杉中国、高瓴创投及智元机器人。本轮融资资金将主要用于加速研发、团队扩充、算力投入及量产测试体系建设。 模感科技成立于2026年5月,总部注册于上海,在深圳前海设有研发中心,聚焦机器人全身多模态触觉感知系统研发。公司正式…

AI 点评 · 红杉、高瓴、智元联手押注,机器人触觉赛道技术壁垒高、应用前景广。

36氪
论文 / 方法
7/16 17:38
Beyond the Leaderboard: Design Lessons for Trustworthy Multimodal VQA

Healthcare multimodal AI must combine visual and textual evidence while remaining reliable and interpretable. Using MediaEval Medico 2025 as a retrospective GI endoscopy case study, we analyze design…

AI 点评 · 多模态医疗AI的可靠性设计突破,用胃肠镜案例揭示可解释性关键。

arXiv
行业信号
7/13 08:34
字节探索自动驾驶,Seed世界模型团队负责|36氪独家

36氪从多位产业人士处获悉,字节跳动正探索进入自动驾驶领域。这一项目目前由Seed旗下周畅的世界模型团队负责。据了解,Seed旗下不仅有周畅的多模态模型、世界模型等团队,还有大语言模型方向。 而自动驾驶与世界模型的技术路线有交叠之处。 另有消息人士告诉36氪,业务方向上,字节有意布局的自动驾驶场景有无人物流,这一业务隶属于字节旗下的火山引擎汽车行业线。 部分…

36氪
行业信号
7/13 02:39
对话Om AI赵天成:多年坚守,押注物理AI原生的「流式」未来

一个从未见过监控画面的多模态模型,却比在监控数据上练了多年的小模型“老将”更懂监控。这不是科幻电影,这是2023年Om AI联汇的一场“无心插柳”,也是CEO兼首席科学家赵天成博士更加坚信“多模态训练方式能为物理开放世界带来泛化性”的关键节点。彼时,AI行业正在追求以大语言模型为核心的生成式AI。 三年后,这个多模态模型演变成了VLX——全球首个面向物理AI…

AI 点评 · 押注物理AI原生流式架构,多模态泛化性突破传统小模型局限。

36氪
设计 / 产品
7/12 15:12
Meta 发布多模态推理模型 Muse Spark 1.1,强化 AI 智能体任务能力

IT之家 7 月 12 日消息,Meta 于 7 月 9 日正式发布适用于 AI 智能体的多模态推理模型 Muse Spark 1.1 版本,重点提升了模型在智能体任务中的规划、协同与执行能力,并增强了工具调用、代码开发、应用操作能力。 Meta 表示,Muse Spark 1.1 强化了多智能体协作机制,由主智能体负责收集信息、制定计划,再将任务拆分并分配…

AI 点评 · 多智能体协作机制是AI落地的关键突破,Meta这次强化了任务拆解与分工能力。

IT之家
论文 / 方法
7/10 16:42
PAC-ACT: Post-training Actor-Critic for Action Chunking Transformers

Precision industrial contact manipulation requires reliable robot policies under pose perturbations and contact-force constraints. Vision-language-action models offer broad generalization but often in…

AI 点评 · 用强化学习微调动作块,提升机器人接触操作的鲁棒性和泛化能力。

arXiv
设计 / 产品
7/7 08:38
caseclose/cma-harness

Cognitive-structured Multimodal Agent (CMA-Harness): a memory-centric agent for long-horizon multimodal understanding, generation, and editing — externalizing v…

GitHub
行业信号
7/7 01:05
用AI“复刻”人类细胞、预判药效,「华源智因」获千万级人民币种子轮融资|36氪首发

文|胡香赟 编辑|海若镜 36氪获悉,AI虚拟细胞(AIVC)企业华源智因近期已完成千万级人民币种子轮融资。本轮融资由水木创投领投,募集资金将主要用于多模态测序底层技术迭代,进一步拓展与头部三甲医院的合作,以及团队扩充等。此外,华源智因团队已计划启动新一轮融资。 华源智因创始团队由资深医药产业从业者、计算生物学研发人员组成,并邀请到深圳国家基因库等单位专家组…

36氪
论文 / 方法
7/6 20:00
Vision as Unified Multimodal Generation

We formulate computer vision as unified multimodal generation, where heterogeneous visual tasks are expressed in the native text and image generation spaces of a unified multimodal model, without task…

HuggingFace Papers
设计 / 产品
7/6 02:52
EPFL-VILAB/Modus

MODUS: Decoder-only Any-to-Any Modeling of Diverse Modalities [ICML 2026]

GitHub
设计 / 产品
7/5 05:38
zhiweio/EagleRAG

Search knowledge by what documents mean and how they look — not one or the other.

GitHub
设计 / 产品
7/4 08:20
JT-Sun/UAVReason

🚁 Can Vision-Language Models Think from the Sky? UAVReason for Aerial Reasoning and Generation

GitHub
论文 / 方法
7/1 20:00
Gemma 4 Technical Report

We introduce Gemma 4, a new generation of open-weight, natively multimodal language models in the Gemma model family. Designed to advance compute efficiency and reasoning, the Gemma 4 model suite feat…

HuggingFace Papers
模型 / Agent
6/30 23:29
谷歌发布全新AI创作工具,加速多模态内容生成

谷歌在近期举行的I/O开发者大会上宣布了一系列面向开发者的AI创作工具升级,旨在通过最新的Gemini模型家族,降低多媒体内容的生成门槛并提升效率。在视频和多模态创作领域,谷歌发布了全新的Gemini Omni模型。该模型能够理解并处理文本、图像、音频和视频输入,并生成连贯的视频内容。其最突出的特点是支持对话式编辑,用户只需用自然语言描述修改需求,如更换角色…

AI 点评 · 多模态对话式编辑降低视频创作门槛,自然语言交互革新内容生产流程。

36氪
行业信号
6/30 15:27
华为官宣全球首个商用多模态文旅大模型规模化应用

IT之家 6 月 30 日消息,华为中国宣布,2026 年 6 月 29 日,全球首个商用多模态文旅大模型 ——“博观文旅大模型”在西安规模应用。截至今年 3 月, “博观”支撑开发的 AI 伴游智能体已覆盖超 400 万用户 。其打造的非遗数字 IP,衍生产品销售超 200 万。 IT之家查询获悉,陕文投与华为等于 2025 年 9 月联合开发的“博观文旅…

IT之家
论文 / 方法
6/29 20:00
Xiaomi-GUI-0 Technical Report

Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interface actions such as tapping, swiping, text entry, and navigat…

HuggingFace Papers
论文 / 方法
6/28 20:00
Orca: The World is in Your Mind

We introduce Orca, an initial instantiation of a general world foundation model. Orca learns a unified world latent space from multimodal world signals and exposes it through multimodal readout interf…

HuggingFace Papers
设计 / 产品
6/27 06:12
odebo/CC-Vision

Give non-multimodal Claude Code main models the ability to see pasted screenshots — a ~200-line UserPromptSubmit hook.

GitHub
行业信号
6/26 23:30
追赶FSD V14,理想在补哪些课?|最前线

过去几年,智能驾驶行业的竞争重心经历了几次明显变化。 最早比的是硬件:激光雷达要不要上、摄像头装几个、算力做到多少 TOPS;随后进入大模型时代,竞争开始转向端到端、VLA(Vision-Language-Action)、World Model(世界模型)等路线。 到了今天,越来越多公司发现,仅仅拥有更大的模型已经不足以形成代际优势,真正决定上限的,开始变成…

AI 点评 · 解析理想追赶特斯拉FSD V14的技术短板,揭示智驾竞争从模型规模转向系统整合的新趋势。

36氪
设计 / 产品
6/26 12:23
FudanCVL/Unison

[ICML 2026] Unison: Benchmarking Unified Multimodal Models via Synergistic Understanding and Generation

GitHub
设计 / 产品
6/25 07:42
RoboScience机器科学发布Visics通用具身大模型,实现跨本体、跨物体、跨任务|最前线

作者|黄楠 编辑|袁斯来 6月24日,通用具身智能企业RoboScience机器科学通用具身大模型发布,首次完整披露自研Visics大模型的技术架构VLOA(Vision-Language-Object-Action),并展示了模型在家具拼装、灵巧抓取、动态流水线等多项真实场景的应用。 大语言模型有标准的文本Token,自动驾驶有统一的视觉或点云表征,这些基…

36氪
设计 / 产品
6/23 07:11
vancyland/DataClaw0

DataClaw: Agentic Tailoring Multimodal Data from Raw Streams — coming soon (code, weights, dataset & DataClaw-val upon acceptance).

GitHub
设计 / 产品
6/23 07:11
vancyland/DataClaw0

DataClaw: Agentic Tailoring Multimodal Data from Raw Streams — coming soon (code, weights, dataset & DataClaw-val upon acceptance).

GitHub
论文 / 方法
6/22 17:58
AIR: Adaptive Interleaved Reasoning with Code in MLLMs

Following the paradigm shift initiated by OpenAI o3, interleaved reasoning with code to enhance multimodal large language models (MLLMs) has become a pivotal research frontier. The existing literature…

arXiv
Skill / 资源
6/22 16:32
Embed the world: Multimodal AI for searchable aerial imagery at scale

In this post, we walk through the problem space, our architecture on Amazon Bedrock and Amazon OpenSearch Serverless, the evaluation methodology we built on OpenStreetMap ground tr…

AI 点评 · 多模态AI将航拍图像转化为可搜索数据,实现地理空间信息的规模化智能检索。

AWS ML
设计 / 产品
6/17 05:11
kaistmm/SeeandSniff

[ECCV 2026] Official Pytorch implementation for See & Sniff: Learning Visuo-Olfactory Representations

GitHub
模型 / Agent
6/16 23:13
谷歌推送 Android 17 正式版,深度集成 AI 功能

IT之家 6 月 17 日消息,谷歌于当地时间周二正式推送了 Android 17 正式版,同时发布了智能手表操作系统 Wear OS 7。本次新版系统将率先搭载于谷歌自家 Pixel 系列设备,同步上线 Pixel 专属功能更新包,新增多项 AI 相关功能,包括对最新人工智能模型的支持,如音乐生成模型 Lyria 3、多模态大模型 Gemini Omni,…

AI 点评 · AI深度融入系统底层,Android 17标志移动平台正式进入AI原生时代,看点在于其生态影响力。

IT之家
行业信号
6/15 22:59
招商银行推出“运通工程师信用卡”,新用户办卡提供“专属 AI 权益”单月可享 18 亿 Token M3 用量

IT之家 6 月 16 日消息,招商银行宣布推出一款“运通工程师信用卡”,强调相应信用卡拥有“专属 AI 权益”。 IT之家参考官方介绍获悉,新用户办卡首次参与活动达标后,至高可享每月 18 亿 Token M3 用量,可直接用于文档、图像、音视频等多模态模型调用,以及 MaxClaw 龙虾部署等 AI 高频场景。具体来看,相应信用卡“AI 权益体系”提供三…

AI 点评 · 银行信用卡与AI算力打包,精准切中工程师群体的高频需求,开创了金融+AI的跨界权益新玩法。

IT之家
论文 / 方法
6/15 17:59
Context-Aware RL for Agentic and Multimodal LLMs

Large language models (LLMs) often fail when answering requires identifying a small but decisive piece of evidence within a long or complex context, such as a single line in a tool trace or a subtle d…

arXiv
设计 / 产品
6/15 09:15
volcengine/ark-cli

The fastest way to put Volcengine Ark in your terminal and your AI agent — go from prompt to generated media, multimodal answer, or deployed endpoint in a sin…

GitHub
论文 / 方法
6/14 20:00
Thinking with Visual Grounding

Visual thinking should not only sound right; it should show its evidence. While recent vision-language models (VLMs) can produce natural-language reasoning traces, these traces often leave the support…

HuggingFace Papers
设计 / 产品
6/14 13:26
Egoist-Machines/LodeDB

World's fastest and most compact embedded vector database: exact by default, multimodal, local-first, and GPU-accelerated

GitHub
设计 / 产品
6/13 06:50
ratschlab/DeepSpotM

Multimodal foundation model predicting transcriptome-wide virtual spatial transcriptomics from histology.

GitHub
论文 / 方法
6/12 17:59
Gaze Heads: How VLMs Look at What They Describe

How a vision-language model internally solves the task of describing an image is far from obvious. We find that the model develops a specific mechanism for this: a small set of attention heads in its…

arXiv
模型 / Agent
6/10 20:10
小米 MiMo Code V0.1.0 探索性 AI 编程助手发布并开源:基于 OpenCode 二次开发,采用 MIT 协议

IT之家 6 月 11 日消息,小米 MiMo 官方今日凌晨正式发布并开源 MiMo Code V0.1.0 —— 一款运行在终端里的探索性 AI 编程助手。 据介绍, MiMo Code 基于开源项目 OpenCode 二次开发,发布并开源,采用 MIT 协议 。它还内置限时免费多模态模型 MiMo-V2.5,同时支持接入 DeepSeek、Kimi 和…

AI 点评 · 小米开源AI编程助手,降低开发者门槛,推动生态共建,MIT协议利于广泛商用。

IT之家
论文 / 方法
6/10 20:00
Self-Evolving Visual Questioner

Vision-language models (VLMs) are typically trained as passive answerers, while their ability to actively ask diverse, non-trivial, visual-centric and grounded questions remains underexplored. Existin…

HuggingFace Papers
设计 / 产品
6/9 11:20
wanshuiyin/ARIS-Movie-Director

Agentic, long-horizon visual generation: a fuzzy story → a cross-model-audited image-based movie. Brings ARIS's research-wiki + multi-agent debate to multimodal…

GitHub
设计 / 产品
6/9 07:08
kenchan0226/multimodal-docs-public

[EMNLP 2025] M-LongDoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework

GitHub
论文 / 方法
6/8 20:00
Kwai Keye-VL-2.0 Technical Report

We introduce Kwai Keye-VL-2.0-30B-A3B, an open-source Mixture-of-Experts (MoE) multimodal foundation model designed to advance long-video understanding and agentic intelligence. To address the challen…

HuggingFace Papers
设计 / 产品
6/6 22:06
GaoxiangLuo/MM-FM

[CVPR 2026] Flow Matching for Multimodal Distributions

GitHub
论文 / 方法
6/5 17:59
MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism

Current Vision-Language Models struggle with hours-long videos because processing full-length visual sequences induces prohibitive token explosion and attention dilution. To overcome this, we introduc…

AI 点评 · 分层记忆架构破解长视频理解瓶颈,用图记忆与智能检索分离感知推理,显著降低计算成本。

arXiv
论文 / 方法
6/3 20:00
Robots Need More than VLA and World Models

Generalist robot intelligence is often framed as a policy-scaling problem: collect more robot demonstrations, train larger Vision-Language-Action (VLA) models, and expect broader generalisation. In th…

HuggingFace Papers
设计 / 产品
6/3 14:37
Able-rip/cc-VisionRouter

Transparent proxy for Claude Code that auto-routes image-bearing requests to a multimodal model — so a non-multimodal primary model never crashes your long-runn…

GitHub
行业信号
6/2 10:22
字节Seed架构调整:周畅管理范围扩大,具身业务纳入核心

字节跳动多模态负责人周畅管理范围再次扩大,原由李航负责的SeedRobotics团队已向周畅汇报月余,李航现以顾问身份负责学术合作方向。字节也正在招聘具身智能技术负责人,负责机器人业务整体规划,职级定位为L8,对标阿里P10-P11,将向周畅汇报。该岗位候选人主要来自头部具身智能创业公司技术负责人。(晚点 LatePost)

AI 点评 · 架构调整显示字节加速整合资源,具身智能成战略核心,技术负责人招聘透露行业人才争夺升级。

36氪
设计 / 产品
6/1 22:38
阿里发布 Qwen3.7-Plus 模型,升级多模态交互混合 AI 智能体

IT之家 6 月 2 日消息,阿里千问大模型今天(6 月 2 日)发布博文,宣布推出 Qwen3.7-Plus 模型, 定位为多模态交互混合智能体。 Qwen3.7-Plus 是 Qwen3.7 的多模态升级版,核心定位是视觉与语言统一的智能体基座。 它保留文本、编码、工具使用和生产力工作流能力,同时强化视觉理解、视觉推理和跨模态任务处理。 模型已通过阿里云…

AI 点评 · 多模态与智能体融合,或加速AI从“对话”迈向“行动”的关键一步。

IT之家
论文 / 方法
6/1 17:56
AdaCodec: A Predictive Visual Code for Video MLLMs

Video is temporally redundant: adjacent frames usually share most objects, background, and layout. Yet existing video multimodal large language models (video MLLMs) usually encode each sampled frame a…

AI 点评 · 提出视频时序冗余新视角,用预测编码压缩帧,有望大幅降低视频多模态模型算力成本。

arXiv
模型 / Agent
6/1 03:36
MiniMax M3 正式发布:前沿 Coding 能力、1M 上下文、原生多模态

MiniMax M3 今日正式发布。 MiniMax M3 在编程和智能体等专业任务上达到了前沿的能力。它使用了全新注意力架构 MSA (MiniMax Sparse Attention),最高支持 1M 超长上下文。它也是一个原生多模态模型,支持图片和视频的输入,并能操作电脑桌面。 在衡量 Coding 能力的 SWE-Bench Pro 上,MiniMa…

开源中国
论文 / 方法
5/29 17:48
Giving Sensors a Voice: Multimodal JEPA for Semantic Time-Series Embeddings

Transformer-based architectures have advanced sequence modeling in language and vision, yet general-purpose representation learning for heterogeneous multivariate time series remains underexplored. We…

AI 点评 · 多模态联合嵌入让传感器数据“开口说话”,突破时间序列通用表征瓶颈。

arXiv
论文 / 方法
5/29 17:20
Vision-Language Models Suppress Female Representations Under Ambiguous Input

Alignment teaches vision-language models (VLMs) to avoid expressing demographic biases, and when gender is clearly visible they largely succeed. Far less is known about ambiguous inputs (a worker in f…

AI 点评 · 揭示视觉语言模型在模糊情境下仍会抑制女性表征,暴露了AI公平性研究的深层盲区。

arXiv
设计 / 产品
5/29 07:25
StarTrail-org/PixelRAG

The end of web parsing. The beginning of scalable pixel-native search.

AI 点评 · 将网页解析转向像素级原生搜索,为多模态检索开辟全新路径。

GitHub
论文 / 方法
5/28 20:00
Task-Focused Memorization for Multimodal Agents

Long-term memory is essential for multimodal agents to build coherent experience, accumulate world knowledge, and achieve continual learning. However, constructing effective memory goes beyond memory…

AI 点评 · 聚焦多模态智能体的长期记忆构建,突破传统记忆局限,实现持续学习与知识积累。

HuggingFace Papers
论文 / 方法
5/28 20:00
Linear Scaling Video VLMs for Long Video Understanding

Video vision-language models (VLMs) are increasingly used in long-horizon and streaming settings, yet most video encoders still rely on spatiotemporal self-attention, causing compute and latency to gr…

AI 点评 · 突破视频模型计算瓶颈,实现线性缩放,为长视频实时理解铺平道路。

HuggingFace Papers
论文 / 方法
5/28 20:00
Representation Forcing for Bottleneck-Free Unified Multimodal Models

Unified multimodal models (UMMs) aim to handle perception and generation in a single model. Yet existing UMMs still rely on a frozen, separately pretrained VAE for image generation, imposing a structu…

AI 点评 · 打破多模态模型依赖预训练VAE的瓶颈,实现真正统一感知与生成,是迈向高效AI的关键一步。

HuggingFace Papers
论文 / 方法
5/28 20:00
How can embedding models bind concepts?

Humans easily determine which color belongs to which shape in multi-object scenes, an ability known as concept binding. Vision-language embedding models such as CLIP struggle with binding: they recogn…

HuggingFace Papers
论文 / 方法
5/28 20:00
MineExplorer: Evaluating Open-World Exploration of MLLM Agents in Minecraft

Multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and action generation. However, their ability to sustain exploration in dynamic open worlds remains un…

AI 点评 · 首个用《我的世界》评估多模态大模型开放世界探索能力的基准,填补了该领域测试空白。

HuggingFace Papers
设计 / 产品
5/28 01:46
modelstudioai/cli

Official Model Studio CLI(阿里云百炼 CLI)built for AI Agent frameworks, exposing models, search, multimodal, and workflow capabilities as structured tool calls.

GitHub
论文 / 方法
5/27 20:00
VLM3: Vision Language Models Are Native 3D Learners

Vision Language Models (VLMs) enable a unified model to solve various vision tasks through prompting. They have shown promising performance in semantic understanding. However, 3D understanding still l…

AI 点评 · 打破视觉语言模型对3D理解的局限,开启原生3D学习新范式。

HuggingFace Papers
论文 / 方法
5/27 20:00
Why Far Looks Up: Probing Spatial Representation in Vision-Language Models

Vision-language models (VLMs) achieve strong performance on spatial reasoning benchmarks, yet it remains unclear whether this reflects structured 3D understanding or reliance on statistical shortcuts…

AI 点评 · 研究揭示视觉语言模型空间推理的盲点,质疑其是否真正具备三维理解能力,对AI可靠性提出关键挑战。

HuggingFace Papers
论文 / 方法
5/26 20:00
Which Pretraining Paradigm Better Serves Spatial Intelligence? An Empirical Comparison of Vision-Language and Video Generation Models

Spatial intelligence requires visual representations that capture both semantic objects and geometric structure in the physical world. To support this, two major pre-training schemes are now widely us…

AI 点评 · 对比视觉语言与视频生成模型,揭示哪种预训练范式更利于空间智能发展。

HuggingFace Papers
论文 / 方法
5/22 20:00
Silent Failures in Physical AI: A Literature Review of Runtime Action Authorization for Autonomous Systems

Physical AI systems increasingly map multimodal observations, language instructions, and learned world representations into physically consequential actions. Robotics foundation models, vision-languag…

AI 点评 · 聚焦物理AI安全盲区,系统梳理运行时动作授权机制,为自主系统风险防控提供关键学术支撑。

HuggingFace Papers
设计 / 产品
5/21 11:14
wangchuxiaoji-oss/doubao2api

Reverse-engineered Doubao (豆包) API → OpenAI-compatible REST service. Free multimodal chat, image/video/music generation, and file hosting for AI agents.

AI 点评 · 逆向工程将豆包API转为OpenAI兼容接口,免费提供多模态功能,大幅降低AI开发门槛。

GitHub
设计 / 产品
5/19 03:58
fudan-generative-vision/PromptReinjection

[ICML 2026] Alleviating Prompt Forgetting in Multimodal Diffusion Transformers

AI 点评 · 用强化学习缓解多模态扩散模型的提示遗忘,为提升AI生成质量开辟新路径。

GitHub
设计 / 产品
5/13 07:17
GoodQ02/goodq4all

Local-first multimodal epistemic memory for scene-level video, audio, and text intelligence.

GitHub
设计 / 产品
5/7 13:28
InternLM/ETCHR

A question-conditioned, reasoning-aware image editor designed to serve as a decoupled visual reasoning assistant for Multimodal Large Language Models (MLLMs).

AI 点评 · 将推理能力融入图像编辑,为多模态大模型提供解耦式视觉助手,拓展了AI交互边界。

GitHub
设计 / 产品
5/3 13:37
VeniVeci/VLM-wiki

A full-modal personal knowledge base built on the Karpathy LLM Wiki concept.

GitHub
设计 / 产品
5/3 11:04
shawn0728/OpenSearch-VL

🔍 OpenSearch-VL provides a fully open recipe for training strong multimodal deep search agents through high-quality data curation, diverse visual/search tools,…

GitHub
Skill / 资源
6/4 14:00
AGI Is Not Multimodal

"In projecting language back as the model for thought, we lose sight of the tacit embodied understanding that undergirds our intelligence." –Terry Winograd The recent successes of…

AI 点评 · 挑战语言中心主义,揭示具身认知对通用智能的核心价值,重塑AI发展路径。

可信度 74交叉信源 1
The Gradient
模型 / Agent一手源
1/26 11:08
Qwen2.5 VL! Qwen2.5 VL! Qwen2.5 VL!

QWEN CHAT GITHUB HUGGING FACE MODELSCOPE DISCORD We release Qwen2.5-VL, the new flagship vision-language model of Qwen and also a significant leap from the previous Qwen2-VL. To tr…

可信度 88交叉信源 1
通义千问 Qwen