Interactive autoregressive video generation demands both low-latency rollouts and precise online control. Few-step distillation accelerates generation by reducing denoising steps, while online control…
视频生成
共 107 条相关资讯 · 来自历史归档
Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos pr…
LTX-2.5 brings frontier video generation to local NVIDIA hardware: 6.8-second clips, native multishot, day-one ComfyUI, open weights. The post The Video Production Stack Now Fits o…
AI 点评 · 本地开源视频生成首次媲美前沿水平,多镜头与ComfyUI支持将大幅降低创作门槛。
4D generation synthesizes dynamic 3D scenes from conditions such as text or images. Existing methods either reconstruct generated RGB videos with a separate 4D model or adapt a particular video genera…
We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated examples spanning 60…
8月6日,阿里巴巴视频生成大模型Wan 3.0开启公测
Black Forest Labs 正式推出 FLUX 3 视频生成模型,研究人员发现 iCloud Private Relay 存在 IP 泄露风险等。 查看全文

Black Forest Labs has launched FLUX 3 Video, which generates Full HD clips up to 20 seconds long with native audio and lip-synced dialogue in more than 14 languages. It can also re…
10秒1080P,成本只要5毛钱
Open agentic prompt-expansion harness for image and video generation, bridging polished demos, public APIs, and deployable workflows.
Controllable video generation models are increasingly being developed as world models. Accordingly, evaluating them in this role extends beyond the apparent appearance of generated videos to the inher…
狂刷3小时!
AI 点评 · AI巨头沉迷短视频,揭示人性与算法共谋的荒诞现实。
🎬 Curated MiniMax H3 video generation prompts — cinematic, ads, anime, UGC, product videos, and more. Includes playable examples and creator attribution.
文|王毓婵 兰杰 编辑|乔芊 36氪独家获悉,曾爱玲入职哔哩哔哩(下称“B站”),担任AI视频生成业务负责人,向CEO陈睿汇报。 36氪就此事向B站方面求证,对方暂无回应。 曾爱玲 B站此前已公开表示,AI投入主要聚焦视频理解、视频推荐和辅助视频创作等方向。曾爱玲入职后,或将参与相关业务。 曾爱玲的个人主页显示,她曾在腾讯混元&AI Lab团队和国际数字经济…
Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly f…

7 月 28 日,OpenAI CEO 萨姆 · 奥尔特曼表示,自己曾因研究 TikTok 的产品机制而逐渐沉迷,甚至在一个周六下午连续刷了约 3 个小时,最终不得不删除这款应用。 奥尔特曼在《Relentless》播客节目中称,OpenAI 此前筹备一款类似 TikTok、以 AI 生成视频为主要内容的社交应用 Sora。为了了解短视频平台的运行方式及用户…
AI 点评 · 科技领袖也难逃算法沉迷,凸显AI时代注意力争夺的普遍性与产品设计的力量。
Generation in video diffusion or flow models is computationally expensive due to the slow and iterative sampling process. Current state-of-the-art (SOTA) acceleration methods heavily rely on variation…
Progress in video generation keeps narrowing the visual gap between AI-generated and professionally produced footage, yet most benchmarks still draw prompts from web sources or LLM templates and score…
Diffusion transformers are essential for high-fidelity video generation, but long token sequences make attention a dominant inference bottleneck. Training-free dynamic sparse attention alleviates this…
The AI lab Midjourney continues to expand its purview beyond image and video generation.
The Media Router is a tool that automatically selects the best image, video, or audio generation model for a request based on whether a developer prioritizes quality, speed or cost…
AI 点评 · 解决生成式媒体模型选择难题,自动平衡质量、速度与成本,实用价值高。
Controllable video generation remains challenging due to the difficulty of specifying precise multi-object interactions using text prompts or motion-control inputs that primarily constrain pixel movem…
AI 点评 · 用图结构控制多对象交互,精准生成动态视频,突破文本与运动控制的局限。
We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SA…
Recent conditional video generation models have shown promising potentials to transform 3D engine renderings, such as depth maps and untextured geometry, into photorealistic videos for gaming and imme…
Kandinsky WM 1.0 — a family of models for Physical AI. Image-to-video generation for autonomous driving, robotics & general physics.
Recent diffusion models enable high-quality video generation, but suffer from slow runtimes. The large transformer-based backbones used in these models are bottlenecked by spatiote…
Text-to-video generation has advanced significantly over the past five years through scaling of model size, data, and compute. Unlike model architecture, training data is often underexplored. Real-wor…
AI 点评 · 训练数据对文生视频影响被低估,该研究首次系统控制变量,揭示数据质量比规模更关键。
CLI & async Python library for free AI chat, image & video generation.
Current video generation models achieve impressive results in single-shot generation, yet remain limited in cinematic video generation, where coherent narratives and effective multi-shot composition r…
Human-object centric video personalization (HOCVP) is a core task within subject-driven video generation. However, existing methods suffer from two key limitations. First, most approaches focusing on…
Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose efficient deployment requires application-specific decisio…
文 | 周鑫雨 编辑 | 张雨忻 智能涌现从多个独立信源处获悉, 截至 2026 年 7 月, 智谱的 ARR(年度经常性收入)已经达到 10 亿美元 。 截至发稿前,针对上述信息,智谱未回复。 过去一年,AI Coding 和视频生成模型已经成为全球造血能力最强的 AI 赛道。 海外,Anthropic 的 Claude Code 仅发布半年,ARR 就飙…
AI 点评 · 智谱ARR半年飙升15倍达10亿美元,印证AI商业化进入爆发期,行业格局加速重塑。
Modern video generation models are increasingly hailed as emerging world models with an internalized grasp of physical law. Yet existing benchmarks largely evaluate physical plausibility only at the o…
Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck in high-resolution video generation. Training-free sparse attention reduces this cost, bu…
Building interactive worlds that respond coherently to player actions has long been a shared goal of computer graphics, games, and artificial intelligence. Recent video generative models provide a dat…
Multi-reference-to-audio-video (MR2AV) generation aims to generate coherent audio-video content conditioned on multiple references and textual instructions. Existing benchmarks mainly focus on text-dr…
Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning,…
Video generative models commonly rely on latent spaces learned by 3D Variational Autoencoders (3D-VAEs). However, conventional 3D-VAEs are mainly optimized for pixel-level reconstruction, which can li…
Recent foundation image and video generation models offer strong generalization and controllability, but their direct application to embodied scenarios is limited by requirements for multi-view consis…
Driven by next-token prediction, NLP shifted from task-specific models into powerful generalist foundation models. What, then, is the equivalent catalyst needed to achieve a general-purpose model in c…
AI 点评 · 视频生成模型突破任务局限,迈向通用视觉,预示AI基础模型新范式。
视频生成的下一站,或是机器人大脑
Reasoning has become a core capability for large models, especially when reliable decisions require understanding logical consequences. Recent video generation models offer a reasoning path distinct f…
We propose OPSD-V, an on-policy self-distillation paradigm for post-training few-step autoregressive (AR) video diffusion models. Existing few-step AR video generators can produce long videos with low…
2026年,AI视频生成赛道已迈入全面爆发的成熟竞速期,Seedance 2.0 的出圈更是让AI视频变成一场“全民狂欢”。 短短两年间,AI视频从最初几秒的碎片化模糊画面,到如今分钟级长视频的连贯叙事、真实物理世界的精准还原,AI视频工具的迭代速度远超预期,AI视频工具完成了从“能用”到“好用”再到“专业”的三级跳,让创意落地的门槛降至新低——专业团队能用…
Despite the recent promise in robot control, video generative models suffer from a domain mismatch due to their primary focus on content creation. For example, their design inherently prioritizes visu…
Pretrained video generative models are promising backbones for visuomotor control, but their imagined futures often drift from task intent and are not reliably action-conditional. As a result, these m…
Recent advances in video diffusion models have enabled either long single-view generation through temporal autoregression, or short multi-view synthesis through bidirectional attention. However, gener…

IT之家 7 月 3 日消息,据 AI 普瑞斯消息,字节豆包视频生成模型 Seedance 2.5 预计 7 月 6 日上线体验中心,将在一周后开放 API 。 据IT之家此前报道, 字节豆包视频生成模型 Seedance 2.5 发布于 6 月 23 日,该模型目前处于全球企业内测阶段。 据介绍,Seedance 2.5 在单段生成长度、多素材参考、视频编…
AI 点评 · 字节视频模型快速从内测走向开放,行业落地节奏加快,值得关注其能力上限。
Recent progress in large-scale generative models has substantially advanced video generation, yet existing methods remain constrained by a rigid inference paradigm. Bidirectional diffusion models exce…
We introduce Vidu S1, a real-time interactive video generation model supporting voice control of digital characters. Users can control video generation content at any moment through voice instructions…
Diffusion transformers (DiTs) achieve state-of-the-art image and video generation, but their multi-step sampling and growing parameter count make inference expensive. Post-training quantization (PTQ)…
Diffusion transformers (DiTs) achieve state-of-the-art image and video generation, but their multi-step sampling and growing parameter count make inference expensive. Post-training quantization (PTQ)…
过去几年,AI的战场在屏幕里。GPT系列用参数堆出了惊人的语言能力,Sora用视频生成震撼了全世界……但2026年,产业界达成了一组共识:2026年,是物理AI的元年。 年初拉斯维加斯CES上,英伟达CEO黄仁勋用一场演讲,17遍提及物理AI,用以宣布“物理AI的ChatGPT时刻已经来了”。这也是他近两年一直推崇备至的关键词。而在过去的2年多时间里,物理A…
Video World Models are interactive video generation models that predict future world states based on user actions and history video frames. A critical challenge in video world models is the lack of me…
“每一代模型,我们都在押注一个非共识。” 文|邓咏仪 编辑|张雨忻 Sand.ai 创始人曹越,不太关心自己站在共识的哪一边。 Sand.ai 是一家视频生成模型和产品公司,成立于2024年1月。曹越创立Sand.ai 的故事也已经被讲过很多遍:在上一段创业“光年之外”戛然而止后,曹越很快就投入到 Sand.ai 的创业中,做视频生成模型。 彼时,市场的主流…
Interactive video generation systems for camera-controlled world exploration roll out growing sequences of latent video frames, entangling state transition with high-frequency observation synthesis. W…
Audio-video generation has recently gained unprecedented research attention, aiming to synthesize high-quality sounding video content with fine-grained synchronization and semantic alignment between t…
Video generation models have emerged as a promising paradigm for embodied world simulation. However, both general-domain video generators and robot-specific data fine-tuned models can still produce ph…
Video generation models aspire to simulate dynamic environments, and several benchmarks now evaluate memory consistency across frames. However, most assess consistency only while the target remains in…
Video generation models are increasingly capable of producing realistic videos, but they still struggle to generate videos that follow basic physical laws. Compounding this is a lack of reliable granu…
Synthesizing a novel-view video from a monocular reference video along a target camera trajectory requires both geometric consistency and motion fidelity with respect to the reference video. Existing…
Open domain subject-driven text-to-video (S2V) generation has drawn significant interest in academia and industry. Open domain S2V mainly involves two scenarios: in-domain, which requires retaining th…
Autoregressive video diffusion with causal diffusion transformers has emerged as a major paradigm for real-time streaming video generation and action-conditioned interactive world models. In this work…
Video diffusion models have enabled remarkable progress in video generation and editing. However, content preservation remains a core challenge: existing methods regenerate every pixel and often alter…
速度快7倍,成本只有Veo 3的1/2000
AI 点评 · 00后挑战行业巨头,用低成本实现7倍速度突破,颠覆音视频模型效率认知。
Generating a coherent multi-shot video requires structured cross-shot memory. Subject appearance, scene context, and speaker identity must persist across cuts. Existing approaches either train end-to-…
World Action Models (WAMs) are embodied predictive-action models that make a forecast of the future available to action. Recent WAMs repurpose large video generation models, and a parallel line relies…
Video generative models ( VGMs) have become a new frontier that can be used not just for video generation but for a multitude of downstream tasks, including world modeling. To advance these tasks, a g…
World Action Models (WAMs) commonly rely on video generation to bridge visual world modeling and robot control. However, video-based WAMs face three coupled limitations: dense multi-frame future token…
Precise 3D spatial orchestration in text-to-video generation remains a significant challenge, particularly for multi-object scenes where semantic layout and temporal dynamics are often entangled. Whil…
Streaming video generation is emerging as a new serving workload in which users interact with long-lived sessions that generate video progressively, chunk by chunk. Unlike offline video generation or…
Curated, original high-craft prompts for AI video ads (Seedance 2.0 / Veo 3 / Kling / Runway). Companion to HeyDreaming.
The shift from video generation to interactive world modeling places new demands on data: beyond captioned videos, world models require temporally aligned video-action-language trajectories grounded i…
As an increasing majority of global video content is consumed on social platforms for interactive social purposes, video generation models built for social worlds are important but largely overlooked…
Consistent video generation under editing operations requires persistence: when edits modify scene appearance or layout, subsequent generations should remain coherent across time and viewpoints. Howev…
We introduce Qwen-RobotWorld, a language-conditioned video world model for embodied intelligence. With natural language as a unified action interface, it predicts physically grounded future visual tra…
Re-rendering an existing video from a novel camera viewpoint requires the output to follow the prescribed camera trajectory while preserving the appearance and dynamics of the original scene across ev…
Long-form video generation requires recurring subjects to remain consistent across various shots, viewpoints, motions, and scene transitions. Existing temporal decomposition methods improve scalabilit…
Cloning camera motion from reference videos is an important task in video generation, as videos provide intuitive and precise control. Existing methods either directly use parametric representations t…
Generating avatar videos that are not merely visually similar to a target individual but behaviorally recognizable, faithfully reproducing their talking rhythm, gestural tendencies, and expression dyn…
Pretrained video generators are promising visual world models that exhibit emergent task-solving abilities; however, their reliance on detailed textual descriptions limits their direct use for plannin…
Autoregressive video generation has emerged as a powerful paradigm for World Action Models (WAMs). However, existing approaches suffer from slow training convergence and limited converged accuracy, pa…
Autoregressive video generators synthesize long videos by generating successive temporal segments, but their historical KV cache grows with video length. Existing bounded-cache methods reduce this cos…
Recent work has demonstrated that online reinforcement learning (RL) can substantially improve the quality and alignment of flow matching models for image and video generation. Methods such as Flow-GR…
Video generative models have become increasingly powerful, but long-range consistency remains challenging to achieve because even a few dozen frames require impractically long transformer sequence len…
Powerful AI magic tools for video editing, text-to-video generation, and VFX.
AI 点评 · 首个为Windows优化的RunwayML套件,让视频AI工具脱离云端限制,本地高效运行。
We introduce StreamForce, a streaming video generation framework that enables physically grounded control through continuous force inputs. Unlike prior video models that train separate models for diff…
Video generation models have made impressive strides in synthesizing visually compelling content, yet their outputs remain confined to the virtual domain. A natural question follows: how well do these…
Developing unified video generation and editing models capable of interpreting interleaved multimodal inputs is a promising yet challenging frontier field. Existing unified frameworks predominantly re…
Video generation models based on Diffusion Transformers (DiTs) have achieved remarkable performance in video synthesis, yet they suffer from high inference latency and computational costs due to the q…
We present Echo Infinity, an autoregressive (AR) framework towards real-time infinite video generation that employs a learnable evolving memory to dynamically filter, abstract, and compress any-length…
We present AAD-1, an Asymmetric Adversarial Distillation framework for One-step autoregressive image-to-video generation. State-of-the-art methods adopt adversarial distillation but suffer from motion…
36氪获悉,国内大厂首个开源龙虾类产品LobsterAI (网易有道龙虾)近日宣布上线图片生成与视频生成能力,并一次性接入包括Seedream、Seedance、HappyHorse、MiniMax-Hailuo在内的模型。
AI 点评 · 多模型矩阵整合,开源策略降低使用门槛,推动AI创作生态。
Autoregressive (AR) video diffusion enables variable-length synthesis, but long-horizon generation often suffers from accumulated errors and identity drift. For efficiency, existing methods commonly a…
AI 点评 · 提出通用检索增强框架,解决长视频生成的累积误差与身份漂移,兼顾效率与质量。
The recent "Reasoning with Video" paradigm utilizes Video Generation Models (VGMs) to generate temporally coherent visual trajectories to complete reasoning tasks. Although state-of-the-art VGMs excel…
AI 点评 · 自适应测试时优化让视觉语言模型成为视频推理的“好老师”,突破传统方法局限。
Official Implementation of LongLive-RAG: A general retrieval-augmented framework for long video generation.
AI 点评 · 开源长视频RAG框架,突破生成时长限制,为AI视频创作提供新路径。
Text-to-video (T2V) generation faces challenging questions when generating videos with long horizons containing multiple events. Inspired by the intrinsics of the diffusion process, we probe video dif…
AI 点评 · 无需额外训练,即可精准控制多事件视频生成,大幅降低算力门槛,推动视频创作民主化。
Recent advances in video generative models have promoted rapid progress in controllable world models. However, maintaining fine-grained spatio-temporal consistency under long-horizon reasoning remains…
Runway Unlimited Pro Gen-3 with unlimited credits, premium models, motion brush, lip sync, upscale, and the complete creative suite—top subscription tier fully…
Spatial intelligence requires visual representations that capture both semantic objects and geometric structure in the physical world. To support this, two major pre-training schemes are now widely us…
AI 点评 · 对比视觉语言与视频生成模型,揭示哪种预训练范式更利于空间智能发展。
Real-time streaming joint audio-video generation for character animation requires a generator to speak the requested transcript, maintain visual identity across chunks, and run within a strict playbac…
AI 点评 · 解耦式编排实现长时流式音视频生成,突破实时角色动画的连贯性与延迟瓶颈。
Recent advances have substantially improved real-time interactive video generation in the autoregressive regime. However, most existing few-step autoregressive video generation methods, often distille…
AI 点评 · 提出单步自回归视频生成新范式,有望突破实时交互瓶颈,显著提升生成稳定性。
Diffusion models have demonstrated strong results on image synthesis in past years. Now the research community has started working on a harder task—using it for video generation. T…
AI 点评 · 视频生成新突破,扩散模型从图像迈向动态世界。
Diffusion models have demonstrated strong results on image synthesis in past years. Now the research community has started working on a harder task—using it for video generation. T…