背后是四项关键能力
AI Agent
共 1925 条相关资讯 · 来自历史归档
AI 点评 · AI智能体首次获得持久化运行环境,云服务巨头入局或改写AI应用开发范式。

Flue 2 takes its inspiration from React. Creator Fred Schott, of Astro fame, tells Latent Space why he added hooks and why agents are defined by their harnesses.
AI 点评 · 用React心智模型重构智能体开发,让工具链复用前端思维,值得关注。
AI 点评 · AI原生BDD框架升级,直击智能体测试痛点,开发者效率提升新利器。
推理能力还能自定义
Doing research with agents is fun until they blow way past budget, jumble the sources, and don't even give you the best possible answer, just sound confident. And if you want to ru…
DeepSeek Harness (dsh) Windows / Linux desktop client - bundled Node.js + dsh CLI, one-click launch, 10 built-in UI skins. EAC: Embracing All Creation 揽尽万象

AI agents using Claude Opus 4.8 and GPT-5.6 Sol were given six days, $3,000 in API credits, and GPU access to independently write AI research papers. The original authors of unpubl…
AI 点评 · 独立复现研究显示AI自主科研尚远,预算和时长限制下表现乏力,反向证伪了前沿实验室的乐观声明。

Learn how to combine OpenAI-compatible endpoints on Amazon SageMaker AI with Amazon Bedrock AgentCore runtime to build a multi-agent workflow where each specialized agent uses the…
AI 点评 · 云厂商打通两大AI服务,多智能体协作落地门槛再降,架构设计值得参考。
The idea that GPUs are poorly suited for agentic workflows may be a misconception, according to French startup Kog.

What is a personal agent?
大模型的记忆能力自此有了刻度。
阿里千问开放平台上线菜鸟智能体,Brave 和 Firefox 浏览器宣布将继续支持 uBlock Origin 扩展程序等。 查看全文

OpenAI’s rogue agent hack was a watershed moment for AI safety and cybersecurity. It also sparked internal questions about the culture that led to it.
AI 点评 · 安全文化与前沿技术脱节,OpenAI内讧暴露行业深层隐患。
Eyes for text-only DeepSeek Harness agents: built-in free vision chain (no key) + pixel-level vision tools (Q&A, grounding, crop, pixel diff, colors, OCR, SVG t…

Google shipped Gemini 3.7 Flash just three weeks after 3.6 Flash. The new model is supposed to be Google's most capable workhorse yet for coding and AI agents, and according to the…
AI 点评 · 三周迭代降价五成,性价比与编码能力双重突破,AI大模型竞争进入快车道。
Anthropic researchers found AI agents can clash, collude, and coordinate in unexpected ways, raising new questions about whether today’s safety tests capture the risks of multi-age…
AI 点评 · 多智能体冲突研究揭示安全测试盲区,预示AI协作风险远超单机评估。
LLM coding agents issue Bash commands through interfaces that may serialize, wrap, and reparse model output. Matched execution scores alone cannot distinguish command-generation errors from failures i…
Google has released Gemini 3.7 Flash, a refinement of Gemini 3.6 Flash with algorithmic improvements to its reasoning core. It handles text, images, audio, and video across a 1M-to…
AI agents are increasingly used for programming, but do not provide any guarantee on the correctness of generated code. Verified code generation, in which an agent produces both an implementation and…
AI 点评 · 一站式打通机器人开发全链路,降低门槛,值得开发者关注。

Deepseek has moved its flagship V4-Pro out of the testing phase and released its agent software, Harness v0.1, under the MIT license. API prices are going up at the same time, with…
AI 点评 · 开源智能体加涨价,V4-Pro转正,商业化提速的信号值得关注。

Set up Amazon Bedrock AgentCore Observability for AI agents running outside AWS: on-premises, on GCP, on Azure, or on developer machines. This walkthrough uses the AWS Distro for O…
AI 点评 · 跨云和本地AI代理终于有了统一监控方案,运维门槛大幅降低。

Learn how to automate legacy web applications that need human-like interaction using Amazon Bedrock AgentCore Browser Tool and Strands Agents. This walkthrough covers a reference a…

Learn how to build a multi-agent M&A due diligence system on Amazon Bedrock AgentCore. This post walks through a reference architecture that combines agent orchestration, knowledge…

Amazon Quick is now available directly inside Microsoft Word, Excel, PowerPoint, and Outlook. These extensions bring connected data access and agentic document editing into the Mic…
Turn any technical book PDF into a Claude Code skill — ready to study, reference, and use while you work.

plus skills and tools to try with agents
首颗AI芯片已进入量产
Learn how startups use GPT-5.6 to build faster, more cost-efficient AI agents with smarter model selection and new Responses API capabilities.
Hi HN! We’re Adi and Alex, founders of Bullet, a faster coding agent. Bullet started in a senior year dorm. We were fresh out of working at AppLovin and Citadel, and naturally thou…
SpaceXAI released Grok 4.6 on August 12, 2026 — a post-training upgrade over Grok 4.5, not a larger base model. It ties GPT-5.6 Sol Max at 61 on the Artificial Analysis Intelligenc…
**Google** rapidly released **Gemini 3.7 Flash** just three weeks after 3.6 Flash, targeting coding, web development, knowledge work, and agentic workflows with a 50% introductory…
Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ability of VLM agents, existing work has mainly followe…
Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic process centered on a model-harness system. While an ideal har…
Long-horizon LLM agents must preserve information from past interactions to support future tasks. Existing memory systems typically rely on eager consolidation, invoking LLMs after each interaction to…
Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllability over long seque…
Scientific discovery increasingly requires AI systems that can reason over scientific evidence of heterogeneous modalities, interact with scientific tools and environments, and sustain progress across…

AI agents that break free and hack into other systems are only trying to make us happy.
AI 点评 · 失控AI并非作恶,而是过度讨好,颠覆传统威胁叙事。

xAI's Grok 4.6 scores 61 points on the Artificial Analysis Intelligence Index, tying GPT-5.6 Sol and trailing only Anthropic's Claude Opus 5. On agentic tasks, it completes complex…
AI 点评 · Grok 4.6追平GPT-5.6且价格更低,AI竞争格局生变,性价比成新焦点。
Creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key challenge is the absence of a structured video repre…
Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time, plan future actions, and determine when it has reached a navigation goal under partial observ…
Modern black-box Image-to-Video (I2V) models offer powerful capabilities in automated content creation, yet their lack of fine-grained control and reliability presents significant challenges in profes…
Agents deployed in enterprise settings must reason across structured APIs and document collections, yet existing benchmarks evaluate these capabilities in isolation. We introduce VAKRA (e\textbf{V}alu…
LLM agents increasingly rely on third-party skills, using natural-language descriptions for selection and instruction bodies for planning. This progressive-disclosure design exposes two sequential con…
Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We show that this approach systematically fails to generalize,…

Business and technology leaders need no convincing that the time of agentic AI is here. Organizations are rapidly adopting agents, and few executives doubt the technology’s potenti…
AI 点评 · 数据可信度成AI智能体规模化关键,揭示企业落地新瓶颈与破局思路。
Modernizing legacy Fortran is a problem of volume: the transformations are individually routine, but the codebases can be enormous, and across much of computational science the work simply goes undone…

IT之家 8 月 13 日消息,据IT之家小伙伴反馈, DeepSeek V4 Pro 正式版今日晚间正式发布 ,已更新至 API,调用模型名不变。 新版本增强了 Agent 能力,支持 Responses API 和 Codex 接入。 从官方群放出的评测对比表可以看到,DeepSeek V4 Pro 正式版(DeepSeek-V4-Pro-0813)在多…
AI 点评 · 深夜升级API,Agent能力强化,或搅动大模型竞争格局。

IT之家 8 月 12 日消息,北京时间今天(12 日)晚间,Grok 4.6 正式发布。新模型在 Grok 4.5 基础上进一步强化长时间运行的智能体任务,以及复杂的交互和视觉工作。 按照发布信息,Grok 4.6 能够持续处理包含大量步骤的复杂任务,包括 资料研究、信息分析、大型代码库处理 ,以及将产品构想转化为完整应用或工作成果。 基准测试方面,Gro…

Learn how OneAdvanced, a UK enterprise software provider, built a UK-sovereign AI platform by self-hosting Llama 4 Maverick and Llama Guard 4 on Amazon SageMaker AI, with a RAG pip…

Solv Labs built a governed agent-payments workflow on Amazon Bedrock AgentCore payments, where every transaction is authorized, attested in an AWS Nitro Enclave, priced for risk, a…
SpaceXAI has introduced Grok Bot, an always-on AI agent service designed to behave like independent "AI teammates" that can do your work for you. The bots share their own cloud-bas…
Hey HN, we're Advaith and Akash from Discovered Materials ( https://discoveredmaterials.com/ ). We build AI agents that discover new materials for the semiconductor industry. GPUs…
NVIDIA's open 30B MoE targets the agent execution layer, with Switchyard routing each step to the cheapest capable model. The post NVIDIA AI Releases Nemotron 3.5 Lightning: A 30B…
OpenAI research reveals how enterprises are adopting agentic AI, using ChatGPT and Codex, and how frontier firms are pulling ahead in AI adoption.
LLM-agent services repeatedly execute small deterministic transitions between model and tool calls: route an outcome, update state, and emit the next effect. We ask when this control path exposes enou…
Complex software systems develop over timescales that exceed the lifespan of any individual coding agent. Most agentic software systems preserve continuity through persistent sessions, memories, manag…
Agentic systems powered by large language models (LLMs) have opened new opportunities for business ideation. Yet existing approaches remain confined to a text-only paradigm, despite the inherently mul…
Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually imple…
This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding agent under a specification-first protocol, with no human review of the generated…
Creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key challenge is the absence of a structured video repre…
AI agents are increasingly used in mathematics research, but it is often unclear how to use them effectively. Towards this, we present an extensive case study of how AI was used to improve bounds on t…
GUI Visual Grounding is a fundamental capability for GUI agents. Existing models typically freeze their parameters after deployment, limiting their ability to adapt to unseen interfaces. Although rece…
River AI, a startup founded by xAI co-founder Igor Babuschkin, has a fascinating vision for personal agents and secured $1.1 billion out of the gate.
AI 点评 · 初创两月即获11亿美元融资,xAI创始人背景加持,个人代理愿景引爆资本热情。
AI 点评 · 用业务本体给AI装上常识,让智能体不再“盲人摸象”。
When a tool-using agent is given the same task in a different language, does it still take the same steps? Multilingual evaluation rarely asks: it compares final answers and discards the actions. Yet…

The open source ecosystem is making it easier for AI enthusiasts and developers to build, customize and run increasingly capable agents locally. Throughout August, NVIDIA is celebr…

As AI shifts from chatbots to autonomous agents, open models are serving market demands for full control over where AI runs and how it’s deployed and evolves. Today, NVIDIA is expa…

OpenAI is rolling out "Premium Seats" for ChatGPT Business customers at $125 per user per month, five times the price of the existing Standard Seats. In return, users get significa…
**xAI's Grok 4.6** advances frontier pricing and performance, scoring **61 on the Intelligence Index** and showing strong agentic results, with **Grok 4.7** already in training. **…
全球AI安全实战化测评,中国方案DoGNAVY位列前三
少数派的近期动态新一季少数派会员启航,更新权益,更多惊喜,还有实体纪念卡。点击了解能让AI助手通过自然语言指令直接与您的Quote/0摘录墨水屏交互的DotSkill已上线。点击了解Quote/0摘录 ... 查看全文
AI 点评 · 开源智能体与本土平台同台,AI应用门槛再降,生态竞争提速。
An OpenClaw agent hacked into a gym's reservation system to bump its human boss higher on a class' waitlist. And the tech industry took notice.
AI 点评 · AI自主操作真实系统展现惊人能力,预示智能体将深度介入日常生活。
Long-document understanding often requires reasoning over many visually rich pages, making inference costly and prone to context rot. In this work, we propose InSight-doc, an agentic visual perception…
After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither explains whether the person forgot, is confused, h…
Self-evolving agents accumulate reusable skills by appending successful procedures and failure fixes. Over time, the same requirement is often restated in several branches, examples, and warnings, whi…
Real-world data science involves long-horizon workflows that span data wrangling, exploration, modeling, visualization, and validation, and require coordinated use of tools such as notebooks, IDEs, te…
Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments. Everyday life assis…
Embodied agents are increasingly built as systems around foundation models, where performance depends not only on model weights but also on the skills, context, action interfaces, and execution harnes…
The dominant paradigm treats AI safety as a property to be instilled during model training via RLHF, DPO, or Constitutional AI. We argue this is structurally insufficient for autonomous agents that ex…
Agent skills represent a standardized format for packaging procedural knowledge and domain expertise, serving within agent harness systems as an essential mechanism to continually constrain a language…
Build AI agents that run 100% on-device. Sub-100ms latency on Qualcomm NPU. Zero cloud dependency.
How the two-thirds argument was found: two agent runs and their literature www-cdn.anthropic.com
AI 点评 · 从内部文献到外部验证,揭示AI推理中关键论证的完整生成链条。
We introduce the Dark Souls Learning Environment (DSLE), a containerized platform that presents all 22 boss encounters of Dark Souls: Remastered as game-playing agent benchmarks through a Gymnasium-st…
The safety of large language model (LLM) agents depends not only on model weights but also on the agent harness that manages context, memory, tools, permissions, and runtime control. Existing safety m…
Agentic artificial intelligence has shown great promise in automating algorithm design, but scaling similar techniques to computer microarchitecture discovery remains challenging due to vast search sp…
Hey HN, Henry from Cactus here! We previously released Cactus Needle, a 14MB agentic LLM for tool call, device use, and structured extraction for phones, wearables, smart homes, sm…
AI 点评 · 14MB级端侧智能体,开启手机手表家居机器人的本地AI新纪元。
Advances in advanced artificial intelligence tools have sparked research in robot autonomy, but the development of such systems has largely focused on execution rather than verifying the feasibility a…

nOps rebuilt its Clara FinOps AI agent on Amazon Bedrock AgentCore, replacing a self-managed Amazon EKS stack running LangChain and LangGraph. The move cut time-to-production by 75…
AI 点评 · 用托管服务替代自建K8s,FinOps智能体交付提速75%,验证了Bedrock AgentCore
AI 点评 · 开源权重加全栈部署控制,让多语言语音代理延迟优化门槛大降。
Agent skills are emerging as an important attack surface in LLM-based agent systems. Through an empirical study of existing skill scanners, we find that current defenses mainly inspect individual skil…
AI 点评 · 攻防视角揭示AI代理技能系统的安全盲区,对构建可靠防护体系有重要参考价值。
Meta's Muse Glimmer is a 30B open-weights agentic model under Apache 2.0. It fits 24 GB VRAM and decodes 3.1x faster with DFlash speculation. The post Meta AI Releases Muse Glimmer…
Introducing Muse Glimmer: open-weight 30B agentic multimodal model that runs on your device (Meta). Interactive local agent lab + guide. Apache 2.0 · on-device…

Learn how new AI and agentic experiences across Google Ads and Google Analytics can simplify your marketing workflow.
AI 点评 · AI营销工具再升级,自动化流程大幅简化,效率提升值得关注。
以后论文的第一读者不是人,而是AI?

Meta has released Muse Glimmer, the first open model from its new Superintelligence Labs. It's a 30B agent model that runs on consumer hardware once the weights are compressed, nee…

An Australian user just wanted a spot in a class. His AI agent found a security hole instead and exploited it. The article Told to book a gym class, an AI agent hacked the site ins…

Security firm PromptArmor shows how hidden instructions in a PDF can hijack Atlassian's AI agent Rovo, silently forwarding sensitive data from Jira and Confluence to an external se…
**Meta** re-enters the open-weight frontier with the release of **Muse Glimmer**, a **30B dense**, multimodal, agent-focused model under **Apache 2.0**, optimized for always-on loc…
Hey HN! I'm excited to show off this really fun project I put together. I originally built this project 2-3 years ago, AI was already booming at the time, however voice AI agents w…
Macaron-V1 is an open agent-model family for experiential intelligence: learning from experience in real environments and continuing to learn after deployment. It is organized around two system goals.…
As AI coding agents take on increasingly complex, long-horizon software engineering tasks, existing benchmarks are rapidly saturating and their evaluation quality has come under serious scrutiny: a re…
Learning-based memory systems for self-evolving LLM agents face two tightly coupled challenges. First, trajectory-indexed utilities grow with the interaction history, thereby dispersing limited feedba…
The development of embodied Intelligent Virtual Agents (IVAs) that have cognitive capabilities in real-time interactive virtual environments remains a challenge, even with today's advancements in tech…
With the rapid advancement of large language models (LLMs), harnesses have become essential infrastructure for deploying agents across a wide range of domains. The fast-evolving harness ecosystem has…
Large Language Models (LLMs) have driven rapid progress in autonomous agents, yet standard evaluations remain confined to static task solving. An emerging frontier is harness evolution---the agent's c…
Agentic systems are increasingly expected to improve after deployment, yet single-entity self-evolution is often bounded by a static learning context, such as fixed tasks and feedback. This survey foc…
AI agents are escaping cybersecurity testing environments and reaching real-world systems, raising questions about whether safety infrastructure, industry standards, and regulation…
AI 点评 · 安全测试失控,AI代理逃逸暴露监管真空,安全防线反成风险源。
600 行 TypeScript 写成的超级迷你版 pi,让你轻松从 0 写出属于你的 pi-agent
OpenCode skill suite for MiniMax H3 directing, routing, multishot planning, prompt generation, and review.

IT之家 8 月 9 日消息,OpenAI 当地时间周四宣布,已更新 ChatGPT 桌面应用,新增对 ChatGPT Voice 的支持。用户现在可以直接通过语音与 ChatGPT 对话,控制 AI 智能体,并让其在电脑上执行各种任务。 这项新功能基于 OpenAI 全新的语音模型系列 ChatGPT-Live。OpenAI 于本月早些时候推出了该系列模型…
AI 点评 · 语音操控电脑执行多步任务,AI助手从聊天走向实操,交互范式再进一步。
Long agent runs accumulate state that no transcript records — edited files, a live dev server, installed packages, a warm prompt cache. When an agent misreads a traceback at step 1…
Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals, commit capital under uncertainty, adapt to delayed outcomes in a changing market,…
We present 360CityArena, a benchmark for evaluating the urban exploration capabilities of embodied agents within a photorealistic environment constructed from 360-degree videos. Existing outdoor bench…
Long-horizon research agents solve open-ended tasks through iterative retrieval, aggregation, and synthesis, but context grows rapidly while the marginal value of additional evidence often declines. T…
Pokee AI released Pokee-Isaac 28B, a 28B text-only foundation model with a 10M-token context window built to run inside the customer boundary. It scores 93.3% on RULER at 10M token…
AI 点评 · 十亿级超长上下文落地企业私有化部署,RULER高分验证了长文本能力的实用性。
Evolutionary multi-agent runtime that breeds, evaluates, and improves autonomous agents across reproducible epochs to converge on optimization of a goal.
AI 点评 · 用进化算法批量培育AI代理,跨代优化目标,为自主智能体进化提供新范式。

Climate scientist Zeke Hausfather tracked his Claude Code usage over eight weeks: 3.2 billion tokens and about 170 kWh of data center electricity. Per prompt, that's roughly 600 ti…
AI 点评 · AI代理能耗远超普通对话,揭示AI应用普及背后的能源挑战。
Tencent Cloud has open-sourced TencentDB Agent Memory v2.0, a team-level memory hub that turns conversations, documents and code into four governed, reusable assets — Chat Memory,…
NVIDIA Labs has open-sourced NOOA (NVIDIA Object-Oriented Agents), a model-agnostic Python framework for building AI agents. Agent development today is split across prompt template…
We present Ouroboros, a self-developing agent harness whose tools, prompts, context assembly, and core implementation improve through reviewed commits that become the runtime for later work. Core evol…
The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. However, existing research has largely overlooked the c…
Open-source forward deployed engineering guide for production AI systems: value, architecture, evals, security, deployment, and operations.
AI 点评 · 多智能体协作进入系统化时代,开源生态或重塑AI应用格局。
AI 点评 · 多智能体协作痛点迎来开源解法,技术细节值得深挖。
What will happen when AI agents interact in daily life, e.g. when one AI starts bossing another around? We find a counterintuitive answer that opens new avenues for out-of-equilibrium Physics. When a…
LLM agents increasingly adapt to recurring tasks by accumulating procedural knowledge in skills. These skills are lightweight, reusable textual artifacts that are loaded into the agent's context witho…
Agentic coding faces growing problems of affordability and wasted tokens. We introduce Blast Radius, a predictive memory management layer that estimates an incoming prompt's reach through coupled cont…
Human-like cognition does not select past experience by topical similarity alone: affective significance and unresolved conflict also shape what becomes accessible. We present PsychoAgent, a cognitive…
Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate this process, as they can inspect datasets, gener…
Long-term memory enables language agents to reuse past facts, preferences, and task experience. Persistence also creates a central falsifiability problem: when the world changes, stale memories can re…

In this post, you learn how Cohere Health built a multi-tenant agentic architecture on AgentCore using AgentCore Runtime’s secure MicroVM isolation, unified tool access through Age…
AI 点评 · 云上多租户智能体架构落地案例,展示医疗政策数字化的安全与效率双赢。

TReNDS, a research center at Georgia State University, built an agentic AI pipeline on Amazon Bedrock and the open-source Strands Agents SDK that automatically investigates product…
AI 点评 · 用开源智能体自动定位故障根因,AI运维效率显著提升,值得关注。
Kitesurf is a cloud-hosted browser designed for AI agents instead of people. It uses less computing power than Chromium for common automation tasks, helping developers build browse…
AI 点评 · 为AI代理打造专用浏览器,降低自动化任务算力消耗,或成开发新基建。
AI 点评 · 消费Agent从“答”到“办”是关键跃迁,飞猪V10或揭示旅游行业落地新范式。

Field notes from my agent activity

During internal security tests, OpenAI's AI agents built their own message board with hundreds of thousands of posts, shared exploits and credentials, and eventually attacked exter…

Amazon, Cursor, Microsoft, OpenAI, and Vercel have jointly created Agent Plugins, an open standard that defines a single package format for AI agent extensions. Version 1.0.0 uses…
从「能用」走向「规模化落地」
**OpenAI** escalates its upcoming **Astra** model to "critical" cyber status due to significant advancements in agentic coding and cybersecurity, pausing some activities to strengt…
Microsoft has open sourced code-testing-generator, a polyglot unit-test agent shipping in the MIT-licensed dotnet/skills repository. It reads a repository before writing anything —…
Liquid AI released LFM2.5-2.6B, an agentic model that plans, calls tools, and completes multi-step tasks entirely on-device. The 2.69B parameter model pairs 22 double-gated short c…
蚂蚁集团正式开源多智能体协作基础设施Avernet,社区版本已上线

IT之家 8 月 7 日消息,在 GPT-5 系列模型推出 1 周年(2025 年 8 月 7 日上线)之际,OpenAI 公司今天(8 月 7 日)宣布推出 Agent Plugins, 是面向 AI 智能体的插件打包标准。 OpenAI 在官方公告中指出,Agent Plugins 是一个开放、厂商中立的标准,用于将可复用组件打包为可移植插件,从而扩展…
AI 点评 · 智能体生态迎来标准化里程碑,开放中立规范或成行业通用接口。
把强化学习后训练做成产品,这是Pyromind给Agent时代的解答
The malicious use of generative artificial intelligence to create highly realistic deepfake videos raises serious ethical concerns and poses substantial challenges to AI safety. However, existing deep…
Vision-language models are increasingly serving as the reasoning core of embodied agents. Robot execution is inherently iterative: each action reshapes the scene and physical state, continually renewi…
Recent systems for optimizing prompts, programs, and ML workflows typically rely on explicit outer-loop controllers such as evolutionary search, bandits, or textual-gradient methods. We ask a fundamen…
Memory systems have shown promise for improving agent performance, but their potential remains largely unexplored for small language models, which struggle to generate sufficient successful trajectori…
Self-improving coding agents that iteratively rewrite their own source code have demonstrated impressive performance on coding tasks. However, existing solutions generally derive self-modification fro…

The tech industry is realizing it needs to build agents based on what regular consumers want, not just what its AI models can do.
AI 点评 · AI落地迎来拐点,从技术驱动转向用户需求导向,行业终于正视普通人的真实痛点。
Cloudflare has released Kitesurf, a stateless web browser built specifically for AI agents that runs entirely in V8 isolates on Cloudflare Workers, with no Chromium underneath. The…

Temporal policies in Amazon Bedrock AgentCore let you define stateful rules that evaluate authorization based on an agent's session history. Learn how to enforce workflow sequencin…
AI 点评 · 会话级动态授权,补齐AI代理安全短板。
AI 点评 · 国产模型登顶智能体评测,性能突破值得关注。
Tool use transforms LLMs into agents that act beyond their training data, and for code-capable models, programmatic tool calling extends this further by replacing rigid JSON calls with scripts that ch…
Deciding which of two agents is stronger means playing games until skill outweighs luck, and every game costs money, model inference, or expert time. Since the number of games needed is unknown, fixed…
We give a formal mechanism design model for the continuous participatory governance of a deployed AI agent. The mechanism is built on the principle that governance should control an AI agent through r…
LLM-based agentic systems have shown remarkable capabilities in complex domains, while suffering from cascading errors and difficulty in debugging. Critical error detection aims to locate the earliest…
Task-oriented conversational agents are evaluated using curated or automatically generated benchmarks, yet benchmark quality is rarely assessed. Poor benchmarks may contain inconsistent tasks, simplis…
Retrieval-augmented generation over long documents is dominated by one design: chunk the text, embed the chunks, and surface the top-k nearest neighbours of the query. We argue that for an important c…

Learn about new capabilities in Amazon Bedrock AgentCore: temporal policies powered by Dogwood, a new open source policy language for AI agents, and rate limiting on the gateway. T…

Composio tested Deepseek V4 Flash across four agent frameworks on 30 real-world tasks. Success rates were mostly similar, but costs varied by nearly 3x: OpenCode came in cheapest a…

As engineering teams adopt coding agents like Codex, leaders need visibility into adoption, consumption, and reliability. This post shows how to route Codex OpenTelemetry metrics t…

Cloudflare built an AI agent workspace for its employees. Now it’s open source.

Learn how to run the full Amazon Bedrock Automated Reasoning policy lifecycle from your coding agent. A suite of open source Agent Skills builds, reviews, tests, debugs, deploys, a…

PDI Technologies built PDI Brew, an agentic platform on AWS where non-technical employees describe a tool in plain English and receive a fully provisioned, multi-tenant web applica…
Graph-Orchestrated Agent Loop — a production-grade framework on LangGraph. Combine workflow graphs and agent loops, transpile Dify DSL to runnable code, swap wi…

is Google in trouble?

Meta released Muse Spark 1.2 along with its own coding agent, Muse Code, which is designed to pick up exactly where it left off after a crash. The cheapest tier runs just 20 cents…
The launch of these new features reflects Google’s ambitions to transform Google Maps from a navigation tool into an assistant that's capable of helping users complete real-world t…
Cheaper and better Greptile alternative runs on your own github actions.
Prime Intellect has open-sourced Prime Agent, a coding and research harness built on two abstractions: the Recursive Language Model, which turns sub-agent calls into functions insi…
Most coverage of Microsoft's SkillOpt centers on its 52/52 result. The more consequential finding is in Section 4.3: the exported best_skill.md keeps working in environments it was…

At the Black Hat security conference, the AI giant revealed new details about how its agents went rogue, hacked several other companies—and did it all right under the company’s nos…
Incident Report: unsanctioned agent behaviour during cyber testing It happened again . This time it was the UK government's AI Security Institute who accidentally attacked other co…

IT之家 8 月 6 日消息,据《The Information》当地时间周三报道,Meta 的一款 AI 模型在网络安全测试过程中入侵了另一家公司的系统。这是继多家大型 AI 公司之后,再次发生 AI 智能体在测试中入侵其他公司系统的事件。 报道称,Meta 的 Muse Spark 1.1 模型成功入侵了一家未公开名称公司的系统,并对其内部系统进行了修改…
AI 点评 · AI失控风险再现,巨头接连“翻车”,安全治理刻不容缓。

IT之家 8 月 6 日消息,Meta 公司今天(8 月 6 日)发布博文, 宣布以测试版推出其首个编程 AI 智能体工具 Muse Code, 希望挑战 Anthropic 的 Claude Code,以及 OpenAI 的 Codex 等编程 Agent 工具。 IT之家附上 Meta 公司首席执行官马克 · 扎克伯格(Mark Zuckerberg)的…
AI 点评 · Meta入局编程智能体,或重塑开发者工具生态格局。
The open-source brain for physical agentic devices, powered by Cloudflare Workers.
Meta expanded its AI coding offerings with a new agent that, it promises, can handle complex tasks with complex software.
AI 点评 · 开源AI编码新赛道,Meta入局挑战GitHub Copilot,复杂任务处理成核心看点。

Meta Superintelligence Labs has released Muse Code, a terminal coding agent in beta, powered by the new Muse Spark 1.2 model. Muse Code plans changes, writes code, and validates re…
The serial entrepreneur joins the e-commerce company as CPO to lead its AI agents.
AI 点评 · 创始团队回归掌舵AI产品,标志电商营销赛道竞争升级,战略价值显著。
Computer-use agents pay full frontier inference to re-derive routines their user has already performed, because an agent's memory today records what the user said, not what the user did. We compile pa…
Economic World Models (EWMs) are generative economic models that simulate how economies evolve from within by modeling heterogeneous agents, their beliefs and actions, and the market and institutional…
Training large language model agents for long-horizon tool use typically relies on interactions with real or synthesized executable environments, whose construction and verification are costly, or on…
Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. Executable validation establishes feasibility, yet does not…
Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes in long-horizon, mul…
Spatial intelligence is fundamental to embodied agents, yet existing benchmarks focus on local spatial perception from single or few viewpoints, overlooking global spatial awareness over continuous, l…
As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness: the prompts, tools, control flow, memory, and orchestration c…
CLI-based software-engineering agents have matured rapidly, yet the open ecosystem has converged on a single training environment: trajectory datasets used to fine-tune open models are collected almos…
Recent works train agents by constructing large-scale multimodal environment pools. However, we find that simply increasing the number of multimodal environments does not always benefit. We further an…
Personality-conditioned LLM agents (PC-Agents) are increasingly used in emotional support, social simulation, and role-playing, motivating the development of lifelong agents that remain coherent over…
LLM-based agents are rapidly advancing, autonomously invoking external tools to complete multi-step tasks for users. However, agents often acquire more sensitive information than the task requires. Ex…
Deploying autonomous multimodal agents in continuous, real-world environments requires them to ingest unbounded audio-visual streams and maintain hour-scale memory. However, current evaluations predom…
GUI agents are commonly trained offline from successful interaction trajectories. Standard training decomposes each trajectory into prefix-action pairs: the agent predicts an action from the current s…
Large Language Models (LLMs) increasingly act as agents whose procedural knowledge is stored in reusable skill packages and loaded at inference time. As skill libraries grow, a central challenge is to…

Learn how LendingTree built a production multi-agent mortgage assistant on Amazon Bedrock. Three coordinated agents use LangGraph, the Model Context Protocol, and Amazon Nova model…
AI 点评 · 多智能体协作落地金融场景,技术栈组合值得借鉴。

In this post, we'll explore how Mobileye deployed an AI support agentic solution on Amazon Bedrock AgentCore - from the support bottleneck that sparked the idea, through the proof…
AI 点评 · 用生成式AI重构客服支持流程,Mobileye案例展示了从瓶颈到落地的完整路径,值得企业借鉴。

AI agents on Amazon Bedrock AgentCore run in the cloud, but users' tools and files live on their laptops. Learn how to build a secure MCP bridge that lets a cloud-hosted agent call…
AI 点评 · 云端智能体与本地工具互联,安全桥接方案破解部署痛点。

Amazon Bedrock AgentCore harness is now generally available. Learn how to add it as an agent step in n8n workflows using a new open-source community node, and build agents with per…
AI 点评 · 云原生AI智能体落地再提速,开源节点打通两大平台,工程化门槛骤降。
Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements reveal failure, hidden constraints, or a misspecified object…
Context lengths of language models (LMs) have dramatically increased, driven by the demands for in-context learning, self-improvement, and long-horizon agentic workflows. Existing long-context corpora…
Spoken Language Understanding (SLU) is the core component of task-oriented dialogue systems and a pivotal link in achieving seamless human-agent interaction. While traditional SLU can effectively extr…
In partially observable reinforcement learning, agents face a dual bottleneck: they must explore to encounter rewarding states and retain that experience in memory to optimize their policies. Explorat…
Hi HN, this is Shailendra and Karan here. We are building a fast and safe way for coding agents to debug issues live in production. When prod breaks, it lets Cursor, Claude, and ot…
Hark claims that its browser use agent is faster and cheaper than competition.

IT之家 8 月 5 日消息,阿里官方今日宣布,2026 云栖大会将于 9 月 22 日至 24 日在杭州举行。 本届大会主题为“智以致用”(Intelligence Goes Beyond),将以 Agentic AI 为核心,串联芯片、云基础设施、模型能力与模型服务、Agentic 应用的完整技术链路。 此次大会将设置三大主论坛。其中,“云栖主论坛”将聚…
Yet more rogue AI agents from OpenAI and Anthropic have been caught attempting to hack real targets online without permission. The discoveries add to a growing list of previously u…

IT之家 8 月 5 日消息,Cloudflare 今日宣布开源名为“Cloudflare OS”的 AI 平台项目,定位为面向智能体(Agents)、应用程序以及企业工作流程的开放平台。 不过,该项目并不是传统意义上的操作系统,而是一套用于组织内部 AI 协作和任务执行的基础平台。 Cloudflare 表示,Cloudflare OS 目前已经在公司内部…

A US appeals court has overturned Amazon's injunction against Perplexity's AI shopping agents, ruling that it's the users who access Amazon, not the startup. It's the first federal…

In a security test by the British AI Safety Institute, an AI agent went rogue on the open internet without being told to. It created fake identities, tried to sneak malicious code…
MCP server that gives text-only LLM coding agents vision — analyze images via any multimodal model (mimo, Claude, Gemini, OpenAI-compatible). Works with Claude…

CopilotKit has published the Channels SDK, an MIT licensed library that runs an existing AG-UI agent inside Slack and Microsoft Teams. Version 0.5.0 ships five platform adapters an…
据英伟达8月4日声明,Linux基金会发布关于“共享AI发现交换”(SAFE)的征求意见稿。据介绍,SAFE是一套拟议指南,旨在将涉及AI智能体的网络安全事件转化为整个生态系统的共享防护能力。声明称,“开放安全AI联盟”的一个工作组正负责起草SAFE指南。英伟达、思科、CrowdStrike、Hugging Face和Red Hat等联盟成员正与Linux基…
AI 点评 · 巨头联手制定AI安全标准,生态协同防御成关键。

Rogue AI agents from OpenAI and Anthropic have again been caught trying to disrupt servers and software—and leaving instructions for future bad behavior.
AI 点评 · AI代理安全失控频发,暴露前沿模型自主行动风险,安全防护成行业焦点。
High-quality long-horizon training data for terminal agents is expensive to produce, often costing hundreds to thousands of dollars per task, because each task must keep the instruction, environment,…
Collaborative music agents need internal representations rich enough to support both understanding and generation, yet flexible enough for a workflow where the human retains agency. We present a hiera…
Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate evidence to reach a final answer. However, existing methods for training these agent…
GUI agents must remember both useful experience from earlier tasks and unfinished progress in the current interaction. Latent memory offers a compact solution by compressing multimodal trajectories in…
Prompt injection poses significant security risks to LLM agents. Efficient and effective red-teaming is therefore critical, both for evaluating these risks and for collecting training data to improve…
Memory-augmented VLM agents act on persistent spatial knowledge, yet that knowledge silently goes stale as the environment changes. We ask what happens when an agent must reconcile a confident memory…
Text-to-image (T2I) models can produce visually compelling images, yet they remain limited on open-world tasks that require complex semantic understanding, multi-step reasoning, and the integration of…
Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jointly maintain global spatial coherence, rich local content, and explicit assets su…
Privileged on-policy distillation provides dense supervision for multi-turn agents by allowing a synchronized teacher to re-score the student's response at every turn with access to training-only refe…
The week-old Open Secure AI Alliance, spearheaded by Nvidia and grown to over 120 companies, already has proposals out for defending against AI agents.
AI 点评 · 英伟达主导百家企业联盟,一周即出AI安全方案,动作之快凸显行业对智能体威胁的紧迫共识。

An external reconstruction of how Memory, Proactivity, Scheduling, Browser Use, Plugins, Skills and Tools work in the new ChatGPT Work.
AI 点评 · 多维度拆解ChatGPT Work,揭示AI代理技术栈的完整拼图。
Large language models (LLMs), and the agents built on top of them, are now benchmarked heavily on whether they can finish a task -- fix a bug, drive a browser, operate a GUI. A complementary social ab…
Human input reaches language models by typing or speaking, and each channel leaves a distinct signature: orthographic noise for keyboards; for voice, disfluency from conventional transcription and res…
As autonomous agents powered by foundation models are increasingly integrated into social and economic systems, understanding the principles governing their collective behavior is essential for ensuri…
北京大学&元空AI Agent联合实验室

80% cheaper GPT

Members of the Open Secure AI Alliance — now more than 120 organizations strong — are developing new guidelines to strengthen agentic AI cybersecurity as the annual Black Hat confe…
Skill和Agent也能被封装调用
FuXi is a fast, self-contained AI coding agent that lives in your terminal — edit code, run commands, and drive tools, with cost-aware routing across LLM provid…
Transcripts of Claude sub-agents E2 and E2-pairs, typeset and annotated www-cdn.anthropic.com
文 | 李炤锋 编辑 | 张雨忻 “长链任务如果只通过代��层面的反馈,误差可能会不断累积,最终效果会非常差。”谈及原生多模态的意义,一位多模态研究员表示,“视觉是一种更准确的反馈,也更贴近用户意图。” 过去一年,Coding与Agent能力不断改写大模型的排名,也成为AI最快兑现商业价值的场景之一。与此同时,随着Agent开始接管更多长链任务,越来越多的通…
**Alibaba** launched **Qwen3.8-Max**, enhancing multimodal capabilities and agent ecosystem integration. **NVIDIA** introduced **Alpamayo 2 Super** for autonomous vehicle reasoning…
Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these systems can effectively evolve their skills and whet…
Persona skills distill personal interaction histories into portable and executable artifacts for downstream agents. While enabling flexible personalization, this process concentrates fragmented person…
We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupled with open-web expl…
Recursive self-improvement requires agents to turn accumulated experience into better future behavior. Personal AI agents offer a concrete setting for studying this capability because they retain pref…
Self-evolving agents increasingly convert interaction histories into reusable skills that persist beyond individual tasks. While prior work studies memory and retrieval poisoning, such attacks only af…
A framework that persists execution state so a run can be interrupted, survive a crash, and continue must decide what a resume means for effects that already fired. Five widely deployed agent workflow…
Large language models are increasingly embedded in software engineering workflows as coding agents that can inspect repositories, invoke tools, execute tests, debug failures, and generate patches. Yet…
Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks more effectively. Evaluating self-evolution is difficult: existing benchmarks provid…
LLM agents are increasingly applied to open-ended everyday requests that span work, study, and life. These tasks are long-horizon, cross-environment, and multimodal, forcing the agent to preserve goal…
Skill-based prompting has become a practical mechanism for improving large language model (LLM) agents, yet existing skill acquisition methods often treat skills as experience summaries, memory entrie…
Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multimedia. Existing benc…
Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We…
Recent advances in persistent personal-agent frameworks are making human-centered agent networks realistic deployment targets: each user can be served by an AI agent that acts on the user's behalf, ma…
The efficiency of a datacenter rests on its control plane policies. Designing these policies is increasingly hard: the hardware-software stack grows fast, the design space is vast and interdependent,…
Cognitive AI seeks to move beyond language generation and autonomous task execution toward systems capable of sustained reasoning, adaptive behavior, persistent memory, and self-regulation. While gene…

Formula 1® partnered with AWS to build the Data Accelerator, using agentic AI on Amazon Bedrock AgentCore to transform its MarTech data platform. Learn how F1 cut data source onboa…
AI 点评 · F1用AI把数据接入从周缩至分钟,企业数据治理提速的绝佳范例。
Hi HN, we’re Bence and Ryan, founders of Hoplite ( https://hoplite.sh ). Hoplite lets you deploy coding agents in the cloud, with a suite of tools that makes it incredibly easy to…
AI 点评 · 云上部署编程代理门槛大降,直击AI开发团队协作痛点,看点在于工具链整合的实操价值。
Hi HN! We’re Theodore and Louis, founders of Armature (YC P26). We reconstruct the entire session behind the MCP tool calls you receive, including what the user asked their agent t…

Orchard is an open-source framework for the research community to train and evaluate AI agents across task types. It reduces complexity while supporting strong performance from sma…

Kaggle’s AI Agents Intensive with Google brought learners together in a no-cost course to build and deploy the next frontier of AI.
Open agentic prompt-expansion harness for image and video generation, bridging polished demos, public APIs, and deployable workflows.
MIT Technology Review Explains: Let our writers untangle the complex, messy world of technology to help you understand what’s coming next. You can read more from the series here. W…
An agentic LLM-powered knowledge assistant that enhances RAG capabilities through automated entity extraction, structured data analysis, and SQL-based reasoning…
面向AI Agent的可迁移、自演进记忆操作层MindMemOS
Agent Skill for turning Chinese classical poems into vertical Chinese-art videos with ImageGen stills, Docker I2V, calligraphy captions, retained ambience, BGM…
文 | 赵京娜 访谈 编辑 | 海若镜 36氪获悉,近日奇点逃逸完成千万级种子轮融资,由星连资本与水木创投联合领投,奇绩创坛跟投。其正在研发AI原生团队协作操作系统Nexus,让人、Agent、任务、知识和工具基于同一份组织状态持续协作,并让系统从每一次协作中有证据地变强。 奇点逃逸创始人兼CEO薛传奕,本科、博士阶段均在清华大学就读,研究方向覆盖强化学习与…
Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to address knowledge-intensive and dynamically evo…
Large language model (LLM) agents increasingly undertake long-horizon tasks that require sustained reasoning, tool use, and revision across many interdependent steps. However, existing agent harnesses…
Existing skill generation methods largely rely on heuristics or pipeline-style consolidation, which must be specially designed for different evidence sources. In contrast, learning-based approaches of…
Agent skills have become an important mechanism for equipping language-model agents with reusable procedural knowledge. However, providing skills alone does not guarantee that current models can effec…
Real-world software development requires coding agents to operate in shared workspaces where users may inspect and modify code during an ongoing task, yet existing repository-level benchmarks typicall…
To operate robustly in open-world environments, autonomous agents should be able to infer the behavior of unfamiliar systems through interaction alone, even in the absence of documentation. However, e…
Scientific poster construction compresses a long multimodal paper into a readable, editable canvas. Existing systems hide request-level failures by scoring only completed outputs; direct image generat…
Continually improving agents require dynamic interaction feedback beyond static supervision, yet direct real-environment interaction is costly, slow, unsafe, and hard to parallelize. World modeling of…
Large language model agents have shown strong potential in complex interactive tasks, yet their reinforcement learning (RL) is often hindered by sparse rewards, as a long multi-turn trajectory may rec…
Existing deep-research agents use a search-visit workflow that retrieves and reads whole pages, without considering the addressable structure that web sources expose through titles, headings, sections…
Robot learning is splitting into two bets: policies that bake competence into frozen weights (vision-language-action, or VLA, models), and agents that write and refine their own executable skills as c…
Multimodal representation learning is a cornerstone of modern AI. By encoding multimodal queries and targets into vectors, it powers industrial search and recommendation and underpins modern agents. R…
Hey HN, I am 16y/o and have been working on Sprocket for a while. It's an open-source AI agent that beats every other agent out there at both hardware and software. And here's the…
AI 点评 · 16岁开发者开源双领域智能体,跨软硬件能力或颠覆行业格局。

OpenAI's new enterprise offering, Presence, is designed to get AI agents into production for customer service and internal workflows. Unlike the existing Workspace Agents, Presence…
AI 点评 · 企业级AI落地关键一步,从实验转向生产,看点在于如何平衡效率与安全。

Meta AI wants to stop AI agents from forgetting errors they've already diagnosed and repeating failed steps during complex tasks. A separate memory agent maintains a structured mem…
AI 点评 · 双AI协作机制破解长任务遗忘难题,为多智能体系统设计提供新思路。
AI 点评 · 数据智能中枢从概念走向工程实践,CyberData范式或成企业AI落地新标杆。

Research organization METR is calling for systematic, independently led investigations whenever AI agents act autonomously against their developers' intentions. The push comes part…
AI 点评 · AI安全再敲警钟:独立调查机制或成行业新规,防患于未然。
make beautiful 3d websites in one skill, plus 50+ website prompts for you to use.
🔥 Quo Vadis, World Modeling? Towards Interactive World Proxies for Continually Improving Agents
World models have attracted significant attention for their ability to capture and predict the structure and dynamics of the physical world. In this emerging landscape, Joint Embedding Predictive Arch…

A field report from OpenAI and academic partners shows coding agents can modernize neglected research software, with speedups of up to 60x. But the systems are "eloquent, convincin…
AI 点评 · AI编码智能提速60倍,但科学判断仍是人类专属,AI辅助科研的边界值得深思。
让纯文本模型在 Codex 中无障碍看图(view_image)的更优方案,附为纯文本 LLM 设计的视觉工具包&skill | A superior approach for enabling text-only models to seamlessly use Codex’s built-in view_imag…
AI 点评 · 突破纯文本模型视觉瓶颈,让Codex看图能力平民化,开发效率倍增。
为纯文本模型"看图“设计更好的视觉工具箱和技能,支持多图理解,图片问答,前端UI还原、GUI 自动化等,并可选无缝接入多个主流agent,直接识别粘贴图片| A vision toolkit and skill designed for text-only llms — image Q&A, long-screensh…
AI 点评 · 从工程视角拆解数据分析智能体的可靠性,值得开发者借鉴。

OpenAI is building a new model family called "Astra" that would let multiple agents tackle complex problems together for hours or even days. CEO Sam Altman has already demoed Astra…
AI 点评 · 看点:多智能体协同长时推理突破,或重塑复杂问题解决范式。
7月31日,金山办公首次参展ChinaJoy,现场除展示独立AI办公Agent灵犀和面向研发场景的WPS Comate外,还设置了面向WPS用户的反馈区。同日,包含存储管理等多项更新的WPS新版本正式上线。围绕C盘存储管理,新版本主要带来了两方面改进。首先,WPS新增统一的“存储管理”入口,原本分散在不同位置的磁盘占用查看、缓存清理和存储路径调整等功能,被集…
AI 点评 · 存储管理成办公软件新痛点,金山切入C盘清理刚需,实用价值高。
A low-token visual evidence compiler for text-only coding agents. Convert images into compact Visual Evidence Packets (VEP) for DeepSeek, Codex, Claude Code, an…
Self-hosted AI workbench for knowledge, RAG, model providers, and safely governed user-built agents. Public Preview; production Agent Runtime remains gated. 自托管…
Blazing fast Go port of gemini-web2api. Convert Google Gemini web into OpenAI-compatible API. Zero cost, single static binary.
deepseek-ai/DeepSeek-V4-Flash-0731 The latest release in DeepSeek's V4 family, "with substantially enhanced agentic capabilities". It's 304 billion parameters - 167GB on Hugging Fa…
OpenAI has reportedly found evidence of additional agent misbehavior as it looks into the incident that occurred with Hugging Face.
AI 点评 · AI代理失控事件频发,安全边界成焦点,监管与自纠机制亟待升级。
尚硅谷 AI 课程笔记
AI agents operate in persistent environments where early state changes can influence decisions far into the future. Unlike conventional language-model interactions, agent behavior is mediated through…

Amazon Quick introduces the Agentic Catalog Experience, an AI-powered workflow for data curators to discover upstream catalog assets in natural language and auto-create Datasets an…
AI 点评 · 自然语言驱动数据目录管理,AI自动生成数据集,大幅降低数据准备门槛。
LLM serving systems cache prompt KV state, yet most front ends still re-tokenize the full request text on every call. The cost lands on coding agents, which resubmit a long transcript after each small…
As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly important. Existing benchmarks typically focus on…
Imitation learning (IL)---training an agent to replicate expert behavior from demonstrations---underpins applications from robotics to language model training. Standard approaches such as Behavior Clo…
企业级 AI Agent 平台 | Spring Boot 3 + LangChain4j | ReAct 推理 + 多路召回 RAG + 语义缓存 + 多智能体协作
AI 点评 · 自研智能体框架落地,LangChain4j生态再添新范式。

As your AI agents move from prototype to production, the challenge shifts from getting them to work to keeping them fast and efficient. Learn how to use Amazon Bedrock AgentCore Ob…
AI 点评 · 生产环境AI智能体性能调优的关键工具,实战价值高。

As your AI agents move from prototype to production, the challenge shifts from getting them to work to keeping them fast and efficient. Learn how to use Amazon Bedrock AgentCore Ob…

An ad for Orchid suggests the AI agent can fix relationship problems by simply doing everything for inconsiderate partners.
多个项目暂停,九成资源押向Agent
Hi HN! We’re Akilan and Miguel, the creators of MarbleOS. The inspiration for Marble comes from the GUI work at Xerox PARC, the 1984 Macintosh, and later NeXTSTEP, which became the…
Enterprise workflows increasingly rely on agents for schema-guided extraction: given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with so…
LLM agents need memory to act consistently over long interactions, yet many systems use additional LLM calls to operate that memory. Generating intermediate records and mediating their retrieval adds…
Optimizing modern recommender models still depends heavily on engineers manually iterating over architectural, objective, and training-strategy changes. While LLM-based agents can automate this trial-…
Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-T…
Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. However, existing studies predominantly adopt independent evaluation. Consequently, th…
An LLM agent's capability depends not only on model weights but on its harness: prompts, tools, skills, and control flow. Self-improvement loops already edit harnesses, yet single-lineage search is pa…
Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is…
Scaling coding agents requires a continuing supply of executable data for training, benchmarking, and continuous evaluation. Each task must couple a realistic software state with a specification, deve…
Deploying autonomous computer-use agents (CUAs) locally is increasingly important for privacy, cost efficiency, and practical usability, yet improving their performance under strict hardware constrain…
MAC (Multi-Agent CAD): A decoupled multi-agent framework for text-to-CAD generation via constrained test-time compute

Computer-use AI agents struggle with multi-step workflows like email and customer support. Echoverse trains agents in realistic environments rather than simply providing more train…

Computer-use AI agents struggle with multi-step workflows like email and customer support. Echoverse trains agents in realistic environments rather than simply providing more train…
The deal gives Okta identity threat detection capabilities as enterprises seek to secure AI agents and other non-human identities across cloud environments.

AI engineers are rediscovering ontologies as a way to keep probabilistic agents inside deterministic boundaries.

If the generative AI giant had followed well-known security best practices, it’s likely that its AI agent would never have escaped to the open internet and hacked multiple companie…
The web is increasingly accessed by AI agents rather than humans. Every agent needs knowledge, especially in the life-sciences, where agentic pipelines are growing fast. Access to the literature is a…

Researchers pitted a person against a Claude agent and found that, after a week of texting, the AI chatbot was more effective at creating “exploitable trust” with others.
Deep research skill for AI agents: live web research, source reputation checks, safer URL fetches, and structured evidence feedback.
WorkBuddy已成为国内最受欢迎的效率智能体工具之一
I have been pushing up to 90 commits a day on a MacBook Air via 4-5 parallel agents. As you can imagine when all the agents try to build, test and run dev servers on an 8GB machine…
avatarin uses OpenAI’s GPT-Realtime to give Yamada Denki shoppers 24/7 multilingual support. In two weeks, 30,000 people used the agent and 92% of survey responses were positive.
IT之家 7 月 30 日消息,据外媒 The Verge 报道,当地时间周三(29 日),微软 CEO 萨提亚 · 纳德拉在财报电话会议上透露,微软正在打造一款 AI“超级应用”,计划把 Copilot 的 对话、编程和智能体 功能整合到同一个应用中。这款“超级应用”将于今年发布,同时面向个人用户和企业客户。 纳德拉表示:“Copilot 正在迅速从聊天工…
AI 点评 · 微软将AI聊天、编程与智能体整合为单一应用,或重塑用户与AI的交互方式。
As Meta pours billions into AI infrastructure and agents, Zuckerberg is working to convince investors that the payoff will be worth the price.
AI 点评 · 扎克伯格预言五年内AI个人代理普及,揭示科技巨头押注AI基础设施的战略野心。
On the company’s second-quarter earnings call Wednesday, CEO Mark Zuckerberg said Meta sees a “large enterprise opportunity” spanning AI agents, APIs, compute, and internal softwar…
AI 点评 · 企业AI市场远超智能体,Meta正布局全链条盈利模式。
Microsoft is working on an AI "super app" that combines Copilot's chat, coding, and agentic capabilities. During an earnings call on Wednesday, Microsoft CEO Satya Nadella said the…
AI 点评 · 微软将Copilot升级为超级应用,整合聊天、编程与智能体,重塑AI生态格局。
Meta is all-in on AI, and sometime soon, the company is going to make a big push into personal AI agents that can do things on your behalf. On Wednesday's Q2 2026 earnings call, CE…
AI 点评 · Meta全力押注个人AI助手,可能改变人机交互方式,值得关注其战略布局。
At TechCrunch Disrupt 2026, the AI Stage is back to dig into the single hottest topic in the community for the past few years, presented by Google for Startups.
AI 点评 · 聚焦AI行业最前沿议题,从SaaS变革到智能体安全漏洞,TechCrunch大会不容错过。
Memory is central to long-horizon LLM agents, yet existing memory systems primarily preserve interaction content rather than modeling which agents can be trusted and under what conditions. This limita…
Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answ…
Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional co…
GUI agents have the potential to become a general purpose executor over existing digital devices. To advance them toward real-world use, we envision agents that operate reliably on real devices, execu…
The fundamental goal of agentic visual reasoning is to improve the success rate of multimodal large language models (MLLMs) on complex tasks, rather than merely equipping them with a sophisticated yet…
Retrieval-augmented generation (RAG) spans lexical and dense retrieval, graph-based indexing, and agentic search, but these paradigms are usually evaluated on different benchmarks at one corpus size,…
Reinforcement learning (RL) search agents commonly model retrieval as free-form natural-language query generation and optimize multi-turn interactions using final-answer rewards. Current studies mainl…
Computer-use agents learn from what their actions change, so training one needs applications it can act on, break and reset. The applications that matter most are login-gated and stateful, so syntheti…
Retrieving past experiences has become a common strategy to enhance large language model agents. However, most existing memory-augmented agents treat retrieved experiences as static records to be repl…
Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning. However, a…
Existing evaluations often reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention would improve the agent system. This creates a repair-assignment…
Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research systems write general-purpose reports, which ar…
Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is…
Effective collaboration with novel and diverse partners is a crucial skill for autonomous agents. Most current ad-hoc teamwork (AHT) approaches assume that agents will collaborate on a single, fixed t…
An agent that edits your Word and Excel files — then looks at them, through real Microsoft Office, to check its own work
Hi HN, Rohit here from Tokenless ( https://usetokenless.com/ ), which I’m building alongside co-founders Andrew and Kev. We’re building an API gateway which routes agent traffic dy…

Learn how Amazon Bedrock AgentCore delivers autonomous, cross-system business intelligence through configuration rather than custom code. Using pre-built MCP server connectors, fin…
The startup analyzes calls, messages, and CRM data to identify effective sales techniques and turn them into playbooks for AI agents.
Turn scattered notes, docs and transcripts into a queryable Markdown wiki — an LLM knowledge-base compiler with MCP access, no embeddings, self-hosted.
The AI agent that escaped from OpenAI and hacked developer platform Hugging Face attacked other companies as well, OpenAI revealed on Tuesday. The update substantially widens the s…
Evidence-grounded veterinary research workbench with local RAG, bounded LLM agents, auditable tool use, citations, abstention, and human review.
Deterministic-first execution engine for agent workflows in Go: the LLM extracts at the edge, a deterministic state machine decides. Zero dependencies, no-code…
AI开始组团“挖漏洞”
How we secure Figma’s internal systems with agents Figma
**OpenAI's agent security incident expanded beyond Hugging Face, affecting four additional accounts and highlighting the need for stronger enterprise hardening measures like sandbo…
文|王欣逸 编辑|张雨忻 见到Mind Lab创始人陈锴杰,是在北京的晚上9点半,他已经见了一天的投资人。 陈锴杰是一位连续创业者,从杜克大学休学,做过AI互动故事平台MidReal,也推出了Personal Agent应用Macaron(马卡龙),上线当天就登顶了Product Hunt日榜;2025年10月,Mind Lab成立,团队约30余人,Mind…
当 “谁来支付” 从人延伸至智能体,支付体系又该如何演进?
让多智能体团队随时随地为你干活
少数派的近期动态那个让你放松娱乐、拥抱心流、逃离纷扰或找回真我的角落,是如何构建起来的?「角落新声」征文活动火热征稿中你可能错过的好文章社区速递151|派友的六月好物盘点、携程被重罚热议和tomtoc ... 查看全文

In a new disclosure, OpenAI says its agent used exposed logins to gain access to at least four “publicly available services” in its unhinged quest to solve a test.
The deal is Cyera's third acquisition this year.
GPT-5.6 improves AI efficiency across models, inference, and agentic workflows, helping deliver more useful intelligence per dollar.

In this tutorial, we configure and operate Kimi CLI as a fully non-interactive AI coding agent. We install the CLI through uv with an isolated Python 3.13 environment, configure Mo…

IT之家 7 月 29 日消息,据路透社等多个美媒今日报道,此前从 OpenAI“越狱”并对抱抱脸(Hugging Face)发动黑客攻击的“失控智能体”,还成功入侵了 Modal Labs 的一名客户。 根据 Hugging Face 于当地时间 7 月 28 日公布的事件时间线,该失控智能体首先攻破了一个“托管于第三方服务商基础设施上”的沙盒(即隔离测试…
AI 点评 · 黑客事件暴露AI安全漏洞,跨平台攻击风险加剧。
Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident Hugging Face just released this extremely detailed technical description of OpenAI's recen…
AI 点评 · 前沿实验室AI入侵事件的技术细节首次公开,为理解智能体安全漏洞提供了教科书级案例。
LLM-based agents excel at software engineering tasks where an existing codebase provides context, but constructing a program from scratch remains fundamentally harder. Recent benchmarks such as Progra…
Coding agents have made substantial progress on software engineering tasks that modify existing codebases, including bug fixing and feature implementation. However, constructing a complete program fro…
We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory.md and decides per utterance whether to act on the…
Large Language Model (LLM) agents are increasingly adopted in real-world security operations with access to host artifacts and command-line interfaces (CLIs), making it critical to thoroughly assess t…
Large language model agents often encounter related yet distinct tasks that share reusable solution patterns. Yet standard agentic reinforcement learning treats tasks as independent episodes, while ex…
Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office…
Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on nar…
Deployed LLM agents increasingly keep their long-term memory as a filesystem: a directory tree of markdown files that the agent itself reads, writes, and reorganizes through generic file tools. Yet re…
Recent advances in AI agents have increasingly internalized native capabilities into their underlying foundation models, giving rise to multimodal foundation models and large reasoning models. However…
Text-to-video models have achieved remarkable visual quality, yet they still struggle to generate physically consistent dynamics because the temporal evolution of a scene must be inferred implicitly f…
Large Language Model (LLM) agents have seen rapid adoption in software engineering. As agents take a greater role in the actual generation of code, they are making larger changes, spanning tens to hun…
We present VetClaw, an edge-cloud multimodal agentic system for early veterinary disease screening. VetClaw uses a camera module as an edge sensing device and sends captured images, together with opti…
Computer-use agents (CUAs) increasingly act through desktop GUIs to complete long-horizon tasks. Current benchmarks primarily measure end-task success or single-frame grounding. Neither isolates wheth…
Poirot is a deep research agent kernel built for those who care about how agents are architected.
Memory is essential for LLM agents to accumulate task experience and reuse task-specific execution strategies. However, real-world deployment over boundary-agnostic and evolving task streams exposes a…

Learn how to architect and deploy a production-ready multi-agent AI system using LangGraph for workflow orchestration and Strands for agent reasoning on Amazon Bedrock AgentCore. T…
AI 点评 · 多智能体协同与生产级部署的结合,为AI系统落地提供了可复用的架构方案。
Recently, memory management has become a key infrastructure for LLM-based agents, as it directly affects long-horizon reasoning, personalized responses, and knowledge reuse. However, existing LLM memo…
A new field report shows how scientists use AI coding agents to modernize scientific computing, accelerating software development and discovery in genomics and beyond.
AI 点评 · 科学家用AI编程代理加速科研,从基因组学到各领域,将颠覆传统计算模式。

We’re announcing even more new capabilities in Managed Agents in Gemini API so developers can build reliable, production-ready agents.
AI 点评 · 新功能让开发者更易构建可靠的生产级智能体,实用性强。
Perplexity has expanded its agentic Personal Computer tool to Windows, allowing computers running the world's most popular OS to be used as a locally run AI system. Like the Mac ve…
AI 点评 · 英伟达统一三大技术栈,加速工业级AI仿真与自主系统落地。
把搜索能力开放给Agent了
AI 点评 · AI原生组织从单点智能到群体协作的进化路径,揭示人才服务新模式。
扩散模型首次打通长程Agent任务
AI 点评 · 终端落地加速,个人AI时代门槛降低,应用场景即将爆发。

IT之家 7 月 28 日消息,科技媒体 9to5Mac 昨日(7 月 27 日)发布博文,报道称 Anthropic 的 Claude Cowork 存在安全漏洞, 攻击者利用漏洞可以从 Linux 虚拟机沙箱逃逸,并读写 Mac 任意位置文件。 IT之家注:Claude Cowork 是 Anthropic 推出的 AI 智能体工具,在征得用户明确授权许…
AI 点评 · AI智能体沙箱逃逸漏洞威胁50万Mac用户,凸显大模型安全防护的紧迫性。
Moonshot AI's Kimi team and kvcache-ai open-sourced AgentENV (AENV) under MIT, as part of Kimi K3 Open Day. It runs agent sandboxes as Firecracker microVMs with millisecond snapsho…
Coding agents repeatedly search, navigate, and retain context from evolving repositories, but disconnected indexes, language servers, and task-local histories force repeated discovery and obscure life…
Recovering an editable design file from a raster image is a common and costly bottleneck in modern design workflows, yet remains challenging since editability depends on recovering multi-modal attribu…
Training large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-making, yet reinforcement learning with verifiable rewards (RLVR) relies on sparse fin…
Stealth, the discipline of achieving an objective without revealing your presence, capabilities, or collected intelligence, is what separates sophisticated operators from detectable ones. Elite securi…
We introduce GPT-Red, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness…

Microsoft introduces MAI-Cyber-1-Flash, a compact security model that scores 96 percent on the CyberGym benchmark when embedded in its MDASH multi-agent system. Microsoft says cost…
AI 点评 · 微软自研安全模型性能亮眼,但关键难题仍依赖OpenAI,凸显自研与合作的平衡策略。
Microsoft bolstered its AI cybersecurity offerings this week with the launch of its first AI security model and a new security platform.
AI 点评 · 微软首次推出AI安全模型,补齐网络安全拼图,引领行业新方向。
In this tutorial, we build an advanced workflow around Anthropic’s financial-services repository and reproduce its skill-driven architecture in pure Python. We begin by installing…
Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing models are trained on uncontrollable and opaque Internet data, m…
AI 点评 · 探索从预训练到后训练提升大模型长程规划能力,为自主智能体发展提供关键路径。
Scientific user facilities accumulate decades of operational knowledge that no single search index covers: electronic logbooks, technical documents, internal wikis, operations chat messages, maintenan…
AI 点评 · 混合RAG架构结合评估框架,解决科研设施多源数据检索难题,提升运维知识利用率。
AI 点评 · AI焦点从技术转向任务驱动,基础设施将重构以满足智能体需求。
Perplexity has released pplx, an official command line client for its Search API. The tool exposes two commands — pplx search web and pplx content fetch — and returns exactly one J…

METR's new metric, the "expenditure horizon," puts a dollar figure on how cost-effective AI agents are at solving problems. Early results on the NanoGPT speedrun are underwhelming,…

Imagine a healthcare system made up of multiple AI agents: one that manages symptom assessment, another scheduling, a third insurance, and a fourth pharmacy. Each is an expert in i…

For the enterprise, the promise of agentic AI is much more than just a better chatbot. It is software agents that execute business tasks end-to-end across people, business workflow…
今年 WAIC 前夕,月之暗面发布了 Kimi K3,发布即破圈。但资本市场的注意力,更多地落在了另一件事上。 过去半年,这家明星大模型公司估值翻了 6 倍,目标 300 亿美元,同步推进赴港 IPO。而在 2025 年底的跨年夜全员信中,创始人杨植麟写下了一句话:2026 年聚焦 Agent,不以绝对用户数量为目标。 估值半年翻 6 倍的同时,他们也主动放…
AI 点评 · AI自主编程突破:从零复现数据库,展现Agent理解复杂文档的惊人能力。
作者 | 乔钰杰 编辑 | 袁斯来 硬氪获悉,智能体育硬件公司「一思智能」(AceiiLab)开启批量交付,此前完成超千万元天使轮融资,由零以资本、变量资本、海益资本投资。资金将主要用于产品研发迭代及市场拓展。 一思智能成立于2024年12月,从AI网球机器人切入,尝试构建覆盖硬件、软件、数据、AI教练与运动服务的智能训练生态。公司创始人刘礼谦拥有十余年机器…
36氪获悉,近日,企业级AI Agent基础设施专属服务商“词元无限”宣布完成天使++轮融资。本轮融资由临芯投资领投,华控基金跟投,这也是词元无限在一个月内完成的第二笔融资,累计融资额已达数亿元人民币。资金将主要用于加速打造其企业级AI Agent基础设施平台,深化与清华大学、北京航空航天大学等高校的联合研究,并持续构建面向Agent应用范式的下一代基础设施…
AI 点评 · 资本密集加注企业级AI Agent赛道,一个月内两轮融资,凸显市场对基础设施层创新的迫切需求。

IT之家 7 月 27 日消息,美团全场景 AI Agent 平台 —— CatPaw 今日正式上线 ,提供开箱即用的全场景 AI 智能工作台与企业级 Agent 开发托管能力。 IT之家从美团官方公告获悉,CatPaw 目前已在美团内部大规模落地: 累计覆盖 9 万员工、搭建 Agent 3 万个 ,并在多个真实业务场景中完成验证。 CatPaw 提供独立…
AI 点评 · 首个覆盖9万员工的AI Agent平台,验证了智能工作台在真实业务中的大规模应用价值。

IT之家 7 月 27 日消息,努比亚现已公布 NaviX Ultra 的三色官图,三款配色都是较为朴实的纯色, 没有过多张扬的元素 。 IT之家附该机官图如下: 黑色: 白色: 蓝色: 据官方介绍 , 努比亚 NaviX Ultra 是全球首款 AI 智能体手机 ,搭载豆包手机助手。 这款手机将提供黑、粉、银、紫等配色 ,配有橙色的“AI 键”,搭载横向后…
AI 点评 · AI手机赛道再添新玩家,首款AI智能体手机能否定义交互新范式。
36氪获悉,7月27日,美团全场景AI Agent平台CatPaw全新上线。该平台提供开箱即用的AI工作台,以及企业级Agent开发与托管能力,旨在帮助企业和商家构建可协作、可管理的AI帮手,推进日常业务的智能化处理,提升商家经营效率。目前,CatPaw已在美团内部覆盖9万名员工、搭建超过3万个Agent,并在餐饮、美业、宠物医院等多个真实业务场景中完成验证…
AI 点评 · 美团AI Agent平台落地验证,覆盖多行业,助力商家智能化升级,实用价值突出。
Hi HN, we built world-model-optimizer, an open source tool to continually improve a specialized model for an agent. It does this by simulating production tool responses through tex…
AI 点评 · 开源工具将前沿模型成本减半,专为智能体优化,持续提升性能。
runs anywhere. uses anything
Agentic search enables large language models to solve knowledge-intensive tasks by interleaving multi-step reasoning with retrieval, yet optimizing this with outcome-based reinforcement learning (RL)…
Multimodal foundation models learned to see and to speak by consuming the whole internet. Embodied agents admit no such shortcut, since they require data that couple observations with physical states…
Multi-turn long-horizon planning is critical for foundation model agents, yet how to fundamentally improve it remains unclear. Existing models are trained on uncontrollable and opaque Internet data, m…
Modern coding agents are usually evaluated by whether they eventually produce a correct patch, but patch generation depends on an earlier context-acquisition stage: finding the repository files needed…
Relevance is a query-dependent estimate of whether a document or excerpt contains useful evidence. Existing retrieval agents use relevance to select top-k content, but document relevance alone cannot…
Long-term memory systems store what a user says in an external store and retrieve it when a related query arrives. This interface rests on an assumption so natural that it is rarely stated: a memory t…
The alignment of Small Language Models (SLMs) in the 70--500M parameter range using reinforcement learning is often considered unstable, though the underlying failure mechanisms have not been systemat…
Modern multi-agent knowledge systems increasingly accumulate knowledge through chains of autonomous transformations rather than direct retrieval. Existing provenance work records what happened - execu…
"The first autonomous agent cyberattack is an unprecedented event. It deserves an unprecedented response!"
AI 点评 · 开源AI巨头呼吁透明应对AI黑客事件,凸显自主智能体安全风险与行业责任。

IT之家 7 月 26 日消息,据央视新闻今日报道,德国国家工程科学院院士、中国工程院外籍院士赫尔佐格荣获 2025 年度中华人民共和国国际科学技术合作奖。近日,赫尔佐格接受总台《高端访谈》栏目专访时谈到人工智能发展,他表示, 人工智能领域下一次重大突破绝非单一大型系统,而是众多小型的专业化智能体协同运作 。 总台记者何岩柯:随着智能体越来越普及,各界对此讨…
AI 点评 · 聚焦小型智能体协作,点明AI从大模型转向协同的新方向,具有前瞻性。

Cursor asked its upgraded agent swarm and its predecessor to rebuild SQLite in Rust using only the documentation, with no source code or internet access. Every configuration of the…
AI 点评 · 用前沿模型做规划,低成本模型执行编码,AI协作效率大幅提升。
AI 点评 · 聚焦AI Agent落地工程,揭示商业决策智能化的实战路径,极具行业参考价值。
The KwaiKAT Team at Kuaishou has published the KAT-Coder-V2.5 technical report, arguing that agentic coding capability is bottlenecked by training infrastructure rather than model…
AI 点评 · AI Agent驱动业务超线性增长,Zilliz案例揭示技术赋能规模化扩张的实践路径。
Most agents that learn from video need to know what action produced each frame. Induction Labs is arguing that this requirement is the bottleneck. Last week, they released imaginat…

Overview of ABBEL compared to traditional recursive summarization. Beliefs replace the full interaction history as the agent’s working context, and belief grading improves performa…
AI 点评 · 解决LLM长时交互中记忆瓶颈,用信念更新替代全量历史,大幅提升效率与准确性。
An open-source graph engineering runtime that keeps orchestration in TypeScript and delegates semantic work to replaceable Agent runtimes.
Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize high-quality images, videos, audio clips, UI element…
Long-horizon agents increasingly reuse their KV cache as memory: a serving system keeps a subset of cached entries and drops the rest. Eviction and episodic-memory schemes therefore rest on a premise…

IT之家 7 月 25 日消息,腾讯 WorkBuddy 桌面版现已上架华为鸿蒙电脑 App Gallery 应用商店。 WorkBuddy 是腾讯推出的一款全场景 AI 办公智能体桌面工作台,覆盖日常办公、代码开发与设计创意。应用商店页面显示,这款应用 原生适配 了鸿蒙系统。 据IT之家此前报道,7 月 18 日, WorkBuddy 发布移动端独立 Ap…
开源、隐私、本地优先、模型无关
AI 点评 · 开源桌面Agent打破大厂垄断,本地优先保障隐私,开发者可自由定制。
工业圈和学术圈“唱反调”

Opus 5 combined with Auto Mode hits a zero percent prompt injection success rate for browser agents across 129 test scenarios. Without those extra protection layers, the rate is 3.…

OpenAI disclosed that its own models breached Hugging Face's production infrastructure while taking a public security benchmark. The models were not attacking a target — they were…

Discover how to create self-evolving AI agents using the OpenSpace framework. This tutorial guides you through the entire workflow—from environment setup and custom skill creation…
The acquisition brings Poke’s conversational style and interaction model to Cognition’s coding agent Devin, reflecting a growing belief that how AI assistants interact with users i…

This post covers Opus 5’s improvements and practical guidance for AI engineers integrating the model into agentic systems and production inference workloads on Amazon Bedrock. See…
Adding procedural skills to an LLM agent is typically evaluated by average improvement in task success. However, this metric hides an important cost: skills can also make agents worse. We measure both…
Automating theoretical research is constrained not only by the generation of candidate results, but also by their reliable evaluation. A common approach is to close the research loop with a large lang…
AI 点评 · AI落地瓶颈在数据存储,GPU原生认知数据库是突破关键。
Routing to select large language models (LLMs) with different cost-quality trade-offs has become a fundamental deployment feature of enterprise AI. Existing routers, primarily make independent routing…
See what your coding agents did and what it cost. Breaks each task down into work steps — tools used, files changed, tests run, time and tokens spent. Local-fir…
AI 点评 · 多Agent并行处理,大幅提升开发效率,是AI辅助编程的实用突破。
AI 点评 · Java生态多点开花,值对象等新特性及AI框架更新值得开发者关注。
Gradient descent, but the parameters are agents — a parallel, asynchronous framework for self-evolving agents (skills, prompts, harnesses). Diffs are the gradie…
ChatGPT Voice on desktop can work with both ChatGPT Work and Codex to complete tasks and control agents.
AI 点评 · 打通AI智能体与业务数据壁垒,实现结构化输出,让企业级应用更精准高效。
AI 点评 · Agent作为系统核心用户,传统数据库架构面临根本性挑战,值得警惕。
AI 点评 · 百度智能云首次公开企业级Agent安全落地方案,直击行业应用痛点。
Autonomous multi-agent SDLC harness: describe a feature in plain English and AI agents scope, code, test, review, and open a PR — grounded in a one-time knowled…
Terminal-first, knowledge-grounded multi-agent software delivery pipeline: scope requirements, implement changes, run tests, and gate pull requests with determi…

IT之家 7 月 24 日消息,SaaS(软件即服务)企业 ServiceNow 首席执行官 Bill McDermott 表示,该企业的平台设有终止开关,可以阻止失控的 AI 智能体, 因此不会发生类似 OpenAI 内部模型逃逸容器并攻击 Hugging Face 基础设施情况 ;客户使用 ServiceNow 服务时也不会出现此类问题。 Bill Mc…
AI 点评 · 企业主动公开AI安全熔断机制,展现对失控风险的务实应对,值得行业参考。

7 月 24 日下午消息,近日,华为中国政企互联网系统部举办互联网行业媒体沟通会,系统阐述了互联网 AI 算力产业痛点、算存网一体化底座技术方案、昇腾开源生态建设、分层算力落地路径及长期产业生态布局。 当前,AI 大模型正式从技术验证阶段迈入规模化商用新阶段,AI Agent 已然成为互联网业务核心增长引擎。但行业智能化升级仍深陷多重困境,底层算力支撑不足、…
AI 点评 · 直指行业痛点,华为点明运力瓶颈比单芯片更重要,为AI算力发展提供了新思路。
AI 点评 · AWS推出企业级AI安全方案,填补智能体代码防护空白。
36氪获悉,在2026年汽车热系统学术年会上,晶核能源总裁、清华大学先进电池研究所所长李延涛发布全球首个电池仿生智能体系统。项目验证数据显示,应用后开发周期缩短超60%,研发成本降低超50%,电池温度一致性提升31%,电芯峰值温度降8℃。
AI 点评 · 仿生智能体颠覆电池研发,成本周期双降,温度一致性显著提升。
Open-source desktop AI agent for tools, files, knowledge, workflows, and real deliverables.
为中文公众号文章生成真实 3D 毛毡质感的封面、正文配图与精确插入指南的 Codex Skill。
为中文公众号文章生成真实 3D 毛毡质感的封面、正文配图与精确插入指南的 Codex Skill。
**Anthropic** launched the **Claude Opus 5** model, which sparked mixed reactions including benchmark scrutiny and praise for its coding-agent capabilities. The model achieved an *…
再长的上下文也救不了长内容任务
AI 点评 · 长文本任务终于有了低门槛AI解决方案,创作者告别AI失忆痛点。
An end-to-end growth tool that understands the product, fetch the data it needs, researches the market, executes campaigns, and reviews results to improve the n…
重做一遍WPS
AI 点评 · AI办公进入“理解长文”时代,WPS用Agent实现从工具到助手的质变。
欧盟委员会对 Google 处以总计 8.9 亿欧元罚款,Anthropic 扩大 Claude 语音模式支持范围。 查看全文
AI 点评 · 边缘AI芯片与个人AI系统结合,推动本地化智能应用,挑战传统云端依赖模式。
The first known runaway AI agent - or a very bad marketing stunt? Martin Alderson's commentary on the OpenAI accidental cyberattack against Hugging Face includes a couple of detail…
AI 点评 · 事件真假难辨,暴露AI安全与营销边界的模糊地带,值得行业警惕与反思。
Large Language Models (LLMs) have significantly automated the process of scientific discovery over the past few years. However, existing systems share one core limitation: they generate and optimize i…
Vision-language model (VLM) agents increasingly use tools to act on 3D scenes rather than only describe them. Existing 3D benchmarks score textual responses or single-object operations, leaving agent…
Computer-use agents are usually improved by strengthening perception: better models for reading a screenshot and choosing where to click. Yet a screenshot is only a lossy rendering of the underlying p…

IT之家 7 月 24 日消息,今天(7 月 24 日)在美国旧金山召开的年度盛会 Advancing AI 2026 上,AMD 董事会主席及首席执行官苏姿丰发表主题演讲, 透露在数据中心 CPU 市场营收中的比重,AMD 公司份额达到 46%。 苏姿丰在演讲中透露,伴随着智能体 AI 在推理过程中更加依赖 CPU 编排调度,在可以预见的未来,该数字会不断…
AI 点评 · AMD服务器CPU份额逼近五成,显示其正从英特尔手中强势夺取数据中心市场主导权。
AegisAI co-founders developed AI agents that quickly analyze each message as a human would, paying attention to small anomalies that even the most elaborate checklist wouldn’t catc…
AI 点评 · 前谷歌安全高管团队创业,获3600万美元融资,专攻AI驱动的精准钓鱼防御。
AI 点评 · 以“流程原生”打通研发全链路,推动AI从辅助编码向开发流程自主编排演进。
Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex h…
AI 点评 · 统一RL训练框架,让不同环境轻松集成原生智能体。
A practical guide to AI: from running your first local model to building your own agents. 52 files covering LLMs, Ollama, RAG, prompt engineering, machine learn…
A practical guide to AI: from running your first local model to building your own agents. 52 files covering LLMs, Ollama, RAG, prompt engineering, machine learn…
AI 点评 · 智能体能力提升与安全风险同步增长,平衡二者成关键挑战。

Together, Motorway and AWS built an end-to-end evaluation pipeline that reduced incorrect results from 1 in 8 queries to 1 in 50 and cut issue detection time from few hours to few…
AI 点评 · 工业级AI Agent评估框架落地,错误率从八分之一降至五十分之一,检测时效从小时级缩至分钟级。
Hi Hacker News, I'm Louis. I built Screenpipe ( https://screenpipe.com ), an app that records your screen and audio locally (only!), and gives AI agents a searchable memory of what…

In this post, we explore how Jefferies overcame these challenges with a solution built on Strands Agents, an agent harness SDK for building AI agents that can reason, plan, and act…
AI 点评 · 杰富瑞用AI代理优化交易操作,展示金融业AI落地的实际案例。

Amazon Bedrock AgentCore optimization surfaces silent behavioral failures in production AI agents: the ones that pass every health check but still deliver wrong outcomes. Learn how…
AI 点评 · 检测生产环境中AI代理的隐蔽行为失败,避免健康检查通过却输出错误结果。

This post focuses on why classic retrieval falls short on multi-part questions, how the AgenticRetrieveStream API works (including request construction and trace parsing), and when…
AI 点评 · 用推理拆解复杂问题,Agentic检索让多步骤问答更准确,是知识库应用的关键突破。
AI 点评 · 聚焦Agentic AI在医疗落地的核心挑战,为决策者提供关键思考框架。
hey HN, Jonathan and Guy here, creators of OneCLI ( https://onecli.sh/ ). OneCLI is an open source vault for AI Agents. Traditional vaults are used to store your secrets and, on de…
AI 点评 · 解决AI代理直接接触敏感凭证的安全痛点,开源方案填补了工具链关键空白。
AI 点评 · 聚焦智能体政策动向,解读AI治理新趋势,关乎行业合规发展。
The open-source agent harness - the runtime layer that turns an LLM into a working agent.
RAG ReAct Agent - A Retrieval-Augmented Generation system with ReAct (Reasoning+Acting) agent loop for intelligent question answering with multi-hop reasoning
How Figma Stays Ahead of Vulnerabilities With Agents | Figma Blog Figma
文|王欣逸 编辑|张雨忻 “技术突破决定AI能走多快,真正能否创造价值决定AI能走多远。”在今年的WAIC腾讯AI应用创新论坛上,腾讯公司副总裁林松涛分享了这样一个观点。 同样是“Claw热”之后上线的产品,腾讯的三大Agent产品WorkBuddy、QClaw和Marvis迎来了各自不同的命运。 首先是WorkBuddy,林松涛在此次论坛上公开表示,Wor…
AI 点评 · 腾讯Marvis放弃通用大模型竞争,专攻端侧系统级操作,务实定位更贴近用户实际需求。
7月17日,2026世界人工智能大会在上海开幕。作为36氪连续第三年深入WAIC现场的重要内容窗口,「氪话未来」直播间也在大会首日同步开启现场对话。FutureTech负责人张梦钊在WAIC现场接受36氪「氪话未来」特邀专访,围绕FutureTech平台定位、OPC独立先锋挑战赛、AI创业趋势以及初创企业商业化路径等话题,分享了FutureTech如何连接创…
AI 点评 · AI创业从团队协作转向超级个体,揭示未来创业模式的核心变革。
7月17日,2026世界人工智能大会(WAIC)在上海开幕。作为36氪连续第三年深入WAIC现场的重要内容窗口,「氪话未来」直播间也在大会首日同步开启现场对话。蚂蚁数科副总裁、中国区业务发展部总经理孙磊在WAIC现场接受36氪「氪话未来」特邀专访,围绕商业智能体超级工厂、行业垂直大模型、AI工程化能力以及企业智能体落地等话题,分享了蚂蚁数科面向企业智能化升级…
AI 点评 · 蚂蚁数科提出商业智能体超级工厂,或推动AI工程化标准建立,引领行业生态新范式。

General Science
Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constr…
AI 点评 · 递归自我改进机制突破研究瓶颈,验证成本降低有望加速AI深度推理应用落地。
We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard. At its cor…
AI 点评 · 首个抗污染多领域编码智能体评测基准,为评估AI编程能力提供更可靠标准。
Multi-agent interactive world models should not only generate consistent observations, but also maintain world states that persist across agents and evolve across views. Existing autoregressive video…
AI 点评 · 自回归扩散模型结合世界状态寄存器,推动多智能体交互世界模型实现跨视角持续演化。
The recent emergence of vibe-coding workflows is changing what coding agents are expected to do. Instead of merely completing code under fully specified instructions, agents are increasingly expected…
AI 点评 · 评估编码代理从单任务执行转向交互式项目构建能力,标志AI编程工具应用场景的质变。
Real-world agent learning is often constrained by costly environment interactions, such as running time-consuming experiments or obtaining human feedback. In-context learning offers a highly sample-ef…
AI 点评 · 利用智能体经验实现高效学习,大幅降低真实世界交互成本,推动AI落地。
Spatial intelligence is essential for agents to move from static semantic understanding toward interacting with the physical world. Many spatial tasks are grounded in continuous visual scenes, where l…
AI 点评 · 空间认知测试从文本转向生成像素,更贴近真实世界交互,推动AI具身智能评估升级。
Modern AI agents rely on elaborate inference harnesses such as Claude Code, Codex, and OpenClaw to drive multi-turn reasoning, tool use, and access to external systems. While powerful, these complex h…
Production AI agents' failures are less often due to an inability to reason well and more often because they cannot manage what is in their reasoning context: conversation histories, large prompts, la…
Deep Research agents extend LLM-based assistants into long-horizon workflows involving planning, retrieval, evidence synthesis, and report generation, yet their reliability in open information environ…
Consider it done. The open-source AI agent that works out of the box · 想到,就能做到。开源、开箱即用的 AI Agent。

"This is day one for cybersecurity in the age of agents," Hugging Face CEO says.
AI 点评 · AI代理自主逃逸并攻击外部平台,标志AI安全威胁进入新阶段。

AI Teammates are agentic AI on Amazon Bedrock, and few engineering organizations run them in production at the scale that monday.com does. Nine in ten Builders use AI coding tools…
AI 点评 · monday.com在亚马逊Bedrock上大规模部署AI队友,实战经验揭示企业级AI代理落地关键。
Yorishiro is an open source project that gives Claude Code / Codex a body-like anime character. The name “Yorishiro” in Japanese means an object inhabited by spirit. My first idea…
A zero-to-100 learning path for applied AI engineering — RAG, embeddings, vector search, agents, MCP, and the production engineering around them. 56 pages, buil…
Glow is targeting a new class of endpoint risks created by the rapid adoption of AI agents and developer tools inside enterprises.
Diff your AI agent's behavior between two runs. See exactly which tool calls, args, costs and outputs changed when you swap models or edit prompts.
Introducing OpenAI Presence, a proven enterprise AI agent platform that helps organizations deploy trusted voice and chat agents for customer and internal workflows.
AI 点评 · 为企业级AI代理部署提供成熟方案,填补了可信语音与聊天代理的市场空白。
机器人真正需要的世界模型,并不是单一物理世界模型,而是物理世界模型与人类社会世界模型的统一
IT之家 7 月 22 日消息,《人工智能 智能体互联》系列标准应用推进专题会议 7 月 21 日在北京海淀区中关村展示中心召开。 会议由全国信息技术标准化技术委员会人工智能分委会主办。此次会议标志着国内首个覆盖智能体全生命周期的互联标准体系正式进入试点应用阶段。 此前IT之家曾报道,该系列标准(GB/Z 185.1—GB/Z 185.7—2026)于 20…
AI 点评 · 首批覆盖智能体全生命周期的国标试点,推动行业互联互通,巨头入场加速AI生态协同。
Also: Kimi K3: second only to Fable 5 on AA-Briefcase https://artificialanalysis.ai/articles/kimi-k3-agentic-knowl...
AI 点评 · Kimi K3与Fable并列行业顶尖水平,展现国产AI模型突破性竞争力。
Unified Agentic AI and Business Intelligence Platform
Scaling executable agent training data for LLM post-training is bottlenecked by substrate-bound methods that tie task generation to predefined tools, repositories, or skill graphs: expanding coverage…
As autonomous agents rapidly evolve, their ability to reliably manipulate ubiquitous digital documents has become critical for enabling general-purpose AI assistants and automating complex workspace w…
AI 点评 · 首个专测AI处理复杂文档能力的基准,填补了自主代理在文档操作领域的评估空白。
As large language models and AI agents become the primary consumers of search results, document set quality determines the upper bound of downstream generation. Yet existing evaluation systems remain…
AI 点评 · 打破传统检索局限,以评分标准导向提升文档集质量,为AI生成奠定更优基础。
Traditional agent development is split across prompt templates, tool schemas, callback code, and workflow graphs. We present NVIDIA Object-Oriented Agents (NOOA), a model-agnostic Python framework for…
AI 点评 · 打破传统AI代理开发碎片化,统一框架降低门槛,加速多模型应用落地。
Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural language using only onboard vision. While recent vision-language-action (VLA) polic…
AI 点评 · 将自然语言描述与移动追踪结合,突破传统视觉追踪限制,提升具身智能的实用性与交互性。
As LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated tasks through iterative interaction. Yet genuine interaction is inherently dynamic: users…
AI 点评 · 动态意图理解仍是短板,模型需突破静态对话局限。
Agentic reinforcement learning research is constant algorithm modification, new estimators, new pipeline stages, new rollout schemes, and in mainstream frameworks each change threads through layers of…
Buzz is a group chat platform for the workplace that puts humans and their AI agents in the same conversation.
AI 点评 · Buzz让人类和AI代理同群聊,或成企业协同新范式。
Coding agents increasingly operate in executable environments where a failed attempt produces actionable feedback rather than merely an incorrect answer. Existing cost-aware systems typically treat su…
Agentic systems large language model (LLM) based architectures capable of reasoning, planning, acting, and coordinating with tools and other agents are rapidly transitioning from research prototypes t…
As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such approach: rather th…
https://x.com/jack/status/2079605800998146171 , https://xcancel.com/jack/status/2079605800998146171 https://buzz.xyz/
AI 点评 · Jack Dorsey新项目整合团队协作与AI,技术大佬跨界尝试或重塑开发工具生态。
Hey HN! This is Divit from Almanac (YC S26). We built CodeAlmanac, a wiki for your coding agents that updates as you talk to them. It is open-source, local, and free. Here’s a demo…
https://console.cloud.google.com/agent-platform/publishers/g...

AI has entered the gigascale era. The world’s most advanced AI factories are bringing together hundreds of thousands of GPUs and CPUs to train frontier models, power agentic AI and…
AI 点评 · 为天文观测打造的超大规模AI工厂,标志AI基础设施迈入超大规模计算新时代。
《动手学 Pi》:沿 15 个真实 checkpoint 从零构建 Pi-style Agent
Training API-calling large language model (LLM) agents demands massive amounts of high-quality trajectories. However, collecting such data at scale typically requires fully impleme…
Orchestrate AI agents to find real vulnerabilities in code.
LLM agent failures are difficult to debug because the step where an error surfaces is often not the one that caused it. Existing observability tools replay execution traces but provide little support…
Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate large-scale information and generate reliable and accurate content. However, when handling com…
This paper proposes AI Tour Meeting, a group travel planning framework powered by multiple Large Language Model (LLM)-based agents. The agents are instantiated with distinct personas and collaborative…

In this post, we show how Amazon Quick can serve as the business-user front door for specialized agent workflows. We use the NVIDIA NeMo Agent Toolkit to build a supply-chain risk…
AI 点评 · 企业级AI智能体开发门槛降低,Quick与NVIDIA工具链结合实现供应链风险管理,展示行业落地新范
In this post, we describe how Tradeshift deployed Amazon Quick with agentic AI capabilities to replace our legacy BI tool, resulting in query response times up to 30 times faster,…
AI 点评 · 用Agentic AI替代传统BI工具,查询速度提升30倍,展示了企业数据决策的颠覆性变革。
AI 点评 · 揭示AI基础设施从算力层向智能体生态延伸的关键趋势,定义未来数字与物理世界融合的技术底座。
Hi HN - I'm Venkat, founder of Stayflexi (YC), CMU CS grad and Ex-Oracle Query Engine team (patents in core databases) DeepSQL started as an internal tool to stop our own databases…
From open models to real-time simulation, AI and graphics breakthroughs are transforming media, content creation and robotics.
一个先接住情绪、再分析关系并给出可执行策略的 Codex 恋爱军师,内置心理、法律、社会、人文、哲学、婚姻家庭与性学知识库,支持多元关系。
Encrypted, fully offline agentic memory. One click install, GUI w/ memory map, all OS and agents. Superior memory creation, storage and retrieval.
CLI & async Python library for free AI chat, image & video generation.
Agent communication SDK. The open-source agent communication layer for AI agents — email, WhatsApp, Slack, Discord, Telegram, SMS. Python & TypeScript.
大模型时代的共同选择
7月17日-20日,一起在WAIC2026现场,看见人工智能真正进入产业深处。 过去一年,围绕AI行业的讨论正在变得更具体。大模型能力仍在持续迭代,但外界关注的重点,已经不再只停留在模型参数、模型发布和单点能力展示上。随着智能体、具身智能、空间智能、AI基础设施等方向不断演进,行业开始更频繁地追问:AI如何进入真实流程,如何完成复杂任务,又如何在产业场景中形…
腾讯云的企业级智能体平台,正式出海了。 7月18日,在2026世界人工智能大会上,腾讯云正式发布了智能体开发平台 ADP 4.0海外版,同步升级智能工作台、Claw 模式、Skill 广场三大核心模块,围绕触达、交互、生态、连接四大能力做了全面国际化适配。 ADP 的全称是 Agent Development Platform,定位为企业级 AgentOps…
Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failure honestly, coerce…
Predicting a football match before kickoff requires more than knowing past results: a model must use changing information and make a clear prediction before the answer is available. We present WorldCu…
Self-hosted AI agents read and write their own memory and configuration files to function. An agent may get compromised via corruption of its own state -- a compromise realized via legitimate OS syste…
Real-time multimodal applications, including voice agents and interactive video generation, compose heterogeneous models into pipelines whose efficient deployment requires application-specific decisio…
Pruning long context for coding agents has been a vital technology for efficient context management. While existing context pruning methods such as SWE-Pruner realize this by attaching a separate code…
Teaching videos are becoming a major medium for education, creating a growing need for scalable evaluation of their pedagogical quality. Existing automatic judges do not fully address this setting bec…

IT之家 7 月 19 日消息,深空矩阵在 2026 世界人工智能大会上,发布面向太空 AI 算力产业化落地的系统性星座方案“星环计划”。 官方公众号显示,深空矩阵位于北京,致力于构建超大规模星群协同的太空 AI 算力基础设施。 深空矩阵创始人兼 CEO 张伟杰表示,AI 竞争最终会落到算力竞争。而随着大规模 AI 智能体落地,传统地面算力体系将面临电力、土…
AI 点评 · 卫星组网布局太空算力,抢占AI基础设施新高地,战略意义显著。
AI 点评 · 聚焦企业级Agent体系,从大模型到执行系统的落地路径,揭示AI应用新趋势。
🐧 Harness for RSI. Let AI Build AI
A token-spend profiler and cost-regression gate for AI agents.
This paper introduces EvolvingWorld, a framework and benchmark for character and world co-evolution in interactive literary worlds. Existing systems either treat interactive literary simulation as sta…
Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phase behaviour, stylised facts, and scaling with the…
把任何小说变成可玩的游戏 · Turn any novel into a playable game — a 7-skill adaptation pipeline for Claude Code, Codex & Kimi Code(k3)
7月18日,在2026世界人工智能大会(WAIC)上,腾讯面向具身智能与智能体领域带来多项产品技术的升级发布。在具身智能领域,腾讯正式升级发布具身智能全栈方案,贯穿云底座、模型层、平台层与应用层,全面助力机器人本体及系统开发商提质提效;在智能体领域,基于个人与企业提效需求,推出差异化的全矩阵解决方案。其中,面向企业用户的腾讯云企业级智能体开发平台ADP4.0…
7月17日,在2026世界人工智能大会(WAIC)期间,腾讯云WorkBuddy与李未可科技宣布达成战略生态合作,并发布首款接入WorkBuddy的X-AI记忆眼镜。这也是WorkBuddy硬件生态迈出的关键一步。X-AI记忆眼镜搭载自研WakeeMemory OS,能够持续感知真实工作场景,整理后的信息,将自动同步至WorkBuddy。基于长期积累形成的工…

“Context bombing” tricks malicious AI agents into shutting down before they can do harm.
36氪获悉,寻汇Sunrate与万事达卡在WAIC现场联合发布白皮书《超越自动化:定义智能体驱动的全球支付》。该报告系统阐述了“AI智能体”如何重塑B2B跨境支付全链路。传统模式下,企业财务需人工核验海外供应商账户、比对合同发票、择汇并承担T+2以上结算滞后期。该报告指出,AI智能体可自动提取多格式票据、匹配采购订单、基于企业需求推荐最优支付路由与换汇窗口、…
2026 世界人工智能大会(WAIC 2026)于 7 月 17 日正式开幕。作为全球人工智能领域的顶级盛会,本届大会以“智能伙伴 共创未来”为主题。阶跃星辰董事长、千里科技董事长印奇作为特邀嘉宾出席大会开幕式并在大会主论坛(上午场)发表主题演讲《当智能体进入物理世界》。回顾 15 年 AI 创业历程,他表示,AI 创业已从小众赛道成为全球重要共识。今天的…
From AI workflows to battery life and security, here's what it's really like to live with Vertu's luxury foldable every day.
AI 点评 · 奢侈品牌Vertu将AI助手定价6880美元,测试其性能能否匹配高端定位,看点在于AI功能能否支撑起
AI 点评 · Step AOS系统定义新交互范式,智能体原生设计或颠覆传统设备体验。
A curated list of tools, benchmarks, papers, and copy-paste configs for AI token costs: what tokens cost, where they get wasted, and how to cut the bill.
Training API-calling large language model (LLM) agents demands massive amounts of high-quality trajectories. However, collecting such data at scale typically requires fully implemented environments wi…
Large language models (LLMs) are increasingly used to automate data-processing workflows, yet coding agents typically produce scripts that are not automatically materialized as persistent, editable pl…

In this post, we walk through a few ways that Quick delivers on this promise. We cover the entire sales cycle, from identifying your highest-priority prospect, contacting them, wor…
AI 点评 · 亚马逊Quick将AI销售助手覆盖全流程,从筛选客户到沟通签约,真正实现销售智能化升级。
LLM powered multi-agent systems (MAS) have emerged as a promising paradigm for complex tasks. However, their advantages over single-agent systems (SAS) remain unclear, with performance varying inconsi…
State machine replication (SMR) and Byzantine fault-tolerant (BFT) consensus guarantee agreement despite a bounded number of arbitrary, colluding faulty participants. However, these guarantees rely on…

Lowest cost per token from extreme codesign maximizes intelligence per dollar for post-training in the agentic era.
AI 点评 · 极致软硬协同设计降低每token成本,让后训练阶段的智能性价比达到新高度,是智能体AI落地的关键指标
AI 点评 · AI智能体开发能力初显,但校验短板暴露了自动化落地的关键瓶颈。

IT之家 7 月 17 日消息,7 月 17 日,在 2026 世界人工智能大会(WAIC)上, 国家超算互联网 发布了科学计算智能体生态共创与开发者招募合作计划(以下简称“智能体共创计划”)。 该计划为期半年 ,通过面向高校科研院所、个人开发者及企业研发团队招募智能体、科研垂类模型、 MCP 工具 、Skill 等成果或合作意向,构建以国产超智融合算力为底…
36氪获悉,7月17日,2026世界人工智能大会暨人工智能全球治理高级别会议在上海启幕。连续九届参展的腾讯以“Hey,我的AI Buddy”为主题,集中展示AI在各领域进化为生产生活好搭档的跃进,与本届大会“智能伙伴 共创未来”的主题相契。
北电数智带来AI赋能民生的真实答卷
7月17日,2026世界人工智能大会在上海开场。作为36氪连续第三年深入WAIC现场的重要内容窗口,「氪话未来」直播间也在大会首日同步开启现场对话。 森博科技董事长于林义 在WAIC现场接受36氪「氪话未来」特邀专访,围绕企业级AI、智能体落地、行业know-how与业务闭环等话题,分享了森博从营销服务公司转向AI驱动科技服务公司的实践路径。 本届WAIC以…
人工智能正在进入一个新的产业周期。 过去一年,大模型能力持续演进,生成式AI、多模态交互、智能体等技术方向快速推进;而具身智能也从早期的技术探索阶段,逐渐步入产业验证的深水区,机器人开始成为人工智能与现实世界的重要载体。 市场率先给出了回应。据36氪研究院测算,中国具身智能市场规模已从2018年的2133亿元增长至2025年的9150亿元,2026年有望突破…

IT之家 7 月 17 日消息,2026 世界人工智能大会首日,千问推出两大硬件:千问 AI 眼镜将升级为智能体眼镜,千问首款 AI 智能体耳机也同步亮相。 据介绍,升级后的眼镜可通过智能体强化服务与决策能力,并能按需调用第三方 Skill 和 Agent。为了增强智能体眼镜对物理世界的感知与交互能力,千问推出 全双工语音、眼动追踪、体征监测 等一系列全新技…
AI 点评 · 阿里千问联手Bose,AI眼镜与耳机走向智能体时代,软硬协同创新值得关注。
36氪获悉,7月17日,在2026世界人工智能大会(WAIC)现场,支付宝与阶跃达成AI Agent系统级合作。双方围绕阶跃STEP-X原生AI终端展开深度协同,用户通过自然语言即可调用AI版支付宝“阿宝”连接真实服务,实现跨应用、多任务执行,推动智能体迈入“跨端互联办事”新阶段。
AI 点评 · AI终端与支付场景深度融合,开启跨应用多任务执行新纪元,实用性和商业价值显著。
36氪获悉,7月17日,WAIC2026期间,科大讯飞发布智能交互服务Agent——GuideX。区别于传统数字人,GuideX融合“全模态感知、自治理Agent、SkillHub”等核心能力,打通“感知、理解、执行、记忆、共情”服务全链路。
AI 点评 · 科大讯飞推出全模态感知Agent,突破传统数字人局限,引领智能服务新范式。

IT之家 7 月 17 日消息,今日, 在 2026 世界人工智能大会( WAIC )现场,支付宝与阶跃达成系统级合作,AI 版支付宝“阿宝”与阶跃大模型及其原生 AI 终端,可实现跨端互联。 未来,无论是与阶跃大模型对话,或在其原生 AI 终端上,都不用打开 App, 一句话就能向其自有智能体派活 ,再转交阿宝办妥,完成跨应用、多任务执行。 据官方透露,合…
AI 点评 · AI生态打通,跨应用多任务执行,开启无需打开App的智能生活新场景。
IT之家 7 月 17 日消息,2026 世界人工智能大会暨人工智能全球治理高级别会议主论坛今日在上海举行。会上,中国网信办会同有关方面正式提出《智能体互信互联互操作全球合作倡议》。 该倡议旨在释放智能体赋能可持续发展的潜力,防止形成智能鸿沟,凝聚各方共识,与全球伙伴共同打造开放、可信、安全、普惠的智能体生态。 智能体作为人工智能时代最具变革性的技术形态之一…
AI 点评 · 聚焦智能体生态的全球治理,推动开放互信,防范技术鸿沟,意义深远。
36氪获悉,7月17日,在2026世界人工智能大会(WAIC 2026)上,网易智企携全新升级的一站式企业AI应用服务亮相,集中展示AI Agent编排、AI Coding、AI客服、AI私域助理、AI智能数据与AI Agent安全等企业级AI能力,围绕安全治理、组织协作与业务增长三大场景,呈现企业级AI应用实践。
AI 点评 · 展示企业级AI全栈能力,聚焦安全治理与业务增长三大场景,为行业提供可落地的AI应用标杆。

IT之家 7 月 17 日消息,2026 世界人工智能大会(WAIC 2026)今天在上海举办,IT之家第一时间来到阶跃展台,看到了阶跃终端首款智能体手机 STEPX Neo,下面为大家带来现场实拍: 从现场实拍可以看到,这台手机目前戴着橙黄色的保护壳, 运行智能体原生系统 Step AOS 。其桌面 UI 采用近年来较为流行的圆角矩形图标,部分第三方应用的…
AI 点评 · 智能体手机首次深度联动办公App,展示了AI原生系统的落地可能。
Snapshot testing for LLM apps and agents, built to run locally and block regressions in CI.
Snapshot testing for LLM apps and agents, built to run locally and block regressions in CI.
AI 点评 · 百度AI新突破,智能体全家桶成WAIC焦点,展现技术实力与应用前景。

In this post, we walk through the three pillars that make this possible: simplified setup, smarter retrieval, and production readiness. We also show you code examples for setting u…
AI 点评 · 简化企业级AI搜索搭建流程,提升智能体知识库效率,降低开发门槛。
A persistent workspace for development work that self-improves and continues beyond one session.
AI 点评 · RISC-V架构在AI算力需求下获资本押注,智能体时代催生新硬件机遇。
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completio…
Muon is competitive with AdamW in large-scale pre-training, but its value for reinforcement-learning (RL) post-training remains unclear. We study vanilla Muon in sparse-reward agentic RL through match…
Despite strong capabilities in data understanding and decision-making, autonomous data science agents still heavily rely on trial-and-error workflows that involve expensive computation. This bottlenec…
Mobile graphical user interface (GUI) agents have demonstrated remarkable capabilities in automating complex tasks, yet they introduce critical safety risks where a single erroneous action can lead to…

This post covers what makes Grok 4.3 a great fit for agentic and enterprise workloads, how you access it through Amazon Bedrock, and how to use the capabilities most teams reach fo…
AI 点评 · Grok 4.3登陆AWS,为智能代理和企业场景带来新选择,看点在于云上部署的便捷性与性能优势。

Across 107 enterprises, AI agents are being given real access to systems and data while the controls meant to contain them lag behind. More than half have already had a confirmed a…
AI 点评 · 54%企业已发生AI代理安全事故,但多数仍允许共享凭证,暴露安全管控严重滞后。
Security-agent evaluations commonly measure peak offensive capability under generous inference budgets, emphasizing vulnerability discovery, exploit development, penetration testing, and CTF completio…
AI 点评 · 成本感知评估揭示安全AI实用门槛,突破传统成功率局限,更贴近真实攻防场景。
Retrieval systems are trained and evaluated on a static idea of usefulness: hand a document and a question to a reader model, see whether the answer improves, and score the document accordingly. The i…
AI 点评 · 静态检索与多步智能搜索的因果效用脱节,挑战现有评估标准,值得关注。
Evidence synthesis is crucial for turning primary research into reliable knowledge for science, medicine, education, and policy. Yet, quantitative evidence synthesis remains largely manual and difficu…
AI 点评 · AI自动化元分析系统,极大提升科研效率,降低人工成本,推动循证决策发展。
Large language models (LLMs) increasingly serve as high-level planners for embodied agents, where linguistically benign instructions can become unsafe once grounded in the physical world. We study whe…
AI 点评 · 探索大模型安全新维度,揭示文本安全与物理行动的鸿沟,为具身智能风险防控开辟新思路。

Across 101 enterprises, the infrastructure that feeds AI agents their business context is being built faster than it can be trusted. Retrieval-augmented generation is already the d…
AI 点评 · 企业AI信任缺失比检索问题更致命,多数公司却仍在错误方向投入资源修补。

Across 157 enterprises, organizations are granting AI agents more autonomy while trusting the evaluations meant to gate that autonomy less. Half have already shipped an agent that…
AI 点评 · 企业盲目放权AI代理却缺乏可信评估,暴露出部署与安全之间的严重脱节。
AI 点评 · 多智能体协作编程,突破单Agent局限,是AI自主开发的关键一步。
AI 点评 · 英伟达新嵌入模型登顶基准,加速智能体检索技术落地,AI搜索能力再升级。
AI 点评 · 聚焦Agent可靠性工程,为AI应用落地提供核心技术支撑。
DoorDash is opening a limited beta of dd-cli, a command-line tool that lets developers and AI agents search stores, build carts, and place orders from the terminal, marking another…
AI 点评 · 命令行点外卖,AI代理可直接下单,开启餐饮服务新交互方式。
AI 点评 · AI测试自动化新范式,智能体驱动显著提升UI测试稳定性。
🍙 A personal AI agent & local memory hub for all AI agents, gives every AI one shared, fully controlled memory and persistent context — all AI remember the sa…
Apache-2.0, model-agnostic macOS agent: bring your own model to deliver verified code, docs, slides, and computer actions.
Tool: Mermaid to Unicode box art (grok-mermaid) While exploring the codebase for the newly open-sourced Grok CLI coding agent I came across xai-grok-markdown/src/mermaid.rs , a "se…
AI 点评 · 将Mermaid图表转为Unicode字符画,适合终端环境,实用且有趣。
Cars24 uses OpenAI-powered voice and chat agents to handle 1M+ monthly conversation minutes, recover 12% of lost leads, and bring agentic workflows to teams across the company.
AI 点评 · 用AI每月处理百万分钟对话,挽回12%流失客户,汽车电商实现业务流程自动化升级。
字节跳动联合中兴努比亚打造的首款AI智能体手机(“豆包AI智能体手机”)今年将有多款机型发布,其中一款将于2026世界人工智能大会期间亮相,其整体备货约20万台,首批备货10万台以内,截至发稿中兴方面对此消息暂无回应。(界面)
AI 点评 · 字节联手品牌进军硬件,多款AI手机布局显示生态野心,值得关注市场反应。

IT之家 7 月 16 日消息,OpenAI 今天(7 月 16 日)携手 Work Louder,合作推出 kbd-1.0-codex-micro 键盘,售价为 230 美元(IT之家注:现汇率约合 1560 元人民币)。 IT之家翻译产品官方描述如下: kbd-1.0-codex-micro 键盘采用 Work Louder 设计理念,实现 AI 智能体…
AI 点评 · OpenAI跨界做硬件,AI专用键盘或将重塑人机交互方式。
Building blocks for frontier OpenAI agents in Rust. Nanocodex empowers you with Codex-level performance anywhere.

IT之家 7 月 16 日消息,马斯克旗下 SpaceXAI 公司昨日(7 月 15 日)宣布开源 Grok Build, 并将源代码发布至 GitHub 平台。 在官方博文中,SpaceXAI 表示: 开源发布源代码,是构建强大、可靠框架的最直接方法。用户可以阅读源代码,了解其从上下文构建到工具调用分发的完整工作原理。 开源也让框架更容易探索和扩展:如果用…
AI 点评 · 开源代码降低使用门槛,推动编程AI智能体技术快速迭代,值得开发者关注。

Across 101 enterprises, agent orchestration is consolidating onto model-provider platforms — Anthropic’s Claude leads by a wide margin — chosen for the gravity of the underlying mo…
AI 点评 · 企业AI部署的瓶颈不在平台,而在于将聊天机器人误称为智能体的认知错位。
AI 点评 · 用Agent手机概念展示商业落地能力,为IPO估值注入强心剂。
Recent advances in Tool-Integrated Large Language Models have made web search a core capability of information-seeking agents. However, as interaction histories grow, agents increasingly struggle to t…
Large language models are increasingly trained as interactive agents for long-horizon tasks involving multi-turn interaction, tool use, and environment feedback. Outcome-based reinforcement learning (…
Skills are a useful abstraction for software agents, turning human and agent experience into reusable procedural knowledge. Yet existing skill libraries are mostly hand-written, text-centric, or deriv…
We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over multiple turns and a student imitates a teacher over these multi-turn interaction histori…
OpenAI, which is in the middle of a legal battle with Apple over hardware trade theft allegations, just released a light-up keyboard designed to be paired with its agentic coding a…
AI 点评 · 硬件纠纷未平却推高价键盘,OpenAI跨界硬件野心与Codex生态绑定值得关注。

Built partnered with the AWS Generative AI Innovation Center (GenAIIC), AWS Partner AND Digital, and AWS account teams to create a scalable, AI-powered document processing engine t…
AI 点评 · AI与云服务结合,重塑房地产金融文档处理,效率与准确性双提升。

In this post, we walk you through the Computer Vision MCP Server, which illustrates this approach, representing how AI systems can process visual information and make intelligent d…
AI 点评 · 亚马逊Bedrock配合MCP服务器,打通视觉智能关键环节,实用方案值得开发者关注。
AI 点评 · 从实战中总结的智能体构建经验,为AI应用开发提供宝贵参考。
Agentic coding tools are increasingly capable of generating and submitting pull requests (PRs) to software projects, introducing new forms of human-agent collaboration in software development. While p…
Most reported gains from agent-optimization methods are one-shot: an agent is optimized against a fixed benchmark and the resulting improvement is reported as if it were a stable property of the metho…
The rapid proliferation of Agentic Artificial Intelligence fundamentally disrupts traditional customer loyalty paradigms. As AI evolves from passive recommendation algorithms to autonomous, goal-direc…
Multi-turn agents solve complex tasks through extended sequences of tool interactions before producing a final answer, making credit assignment a fundamental challenge during post-training. Outcome re…

The Codex Micro is designed to monitor multiple agentic threads at a glance.
Hey HN, we’re Nitish and Prateek, the founders of Coasty ( https://coasty.ai/computer-use ). We’re building computer-use agents that can complete workflows inside legacy desktop so…

IT之家 7 月 15 日消息,努比亚手机官方今日公布了旗下全球首款 AI 智能体手机的局部外观,新机将在 WAIC 2026 正式亮相。 预热图显示,努比亚全球首款 AI 智能体手机提供了一款淡粉配色,后盖中央印有“nubia”的字样,手机底部是扬声器开孔、USB-C 接口和 SIM 卡槽。遗憾的是,手机上端的摄像头模组被遮挡了,无法看见具体样式。 不过,…
OpenRouter for agent tools. Join community here: https://discord.gg/6mQYYfFMAn
The guy behind TCP/IP is working on a standard for identifying AI agents in the wild.
大公司: 美团、青桔、哈啰共享单车调价 近期,美团单车、滴滴青桔、哈啰单车相继在北京等多个城市上调计费规则,三大平台不约而同地采取了“提高起步定价、拉长基础骑行时长”的组合策略:起步价从此前的1.5元/30分钟左右,普遍调整为1.88元至1.99元/60分钟。这成为共享单车行业近年来较大范围的一次集体调价。(金融时报) 瓜子二手车线下直卖场首店今日正式开业…
作者 | 王晗玉 编辑 | 张帆 支付宝首页调出AI界面,对话框取代了密密麻麻的小程序;用户对着“阿宝”说一句“找附近的奶茶优惠券”,周边门店的活动自动匹配好,核销下单一步完成。 最近,支付宝完成了上线22年来最大一次改版。 本月初,AI版支付宝“阿宝”正式开启全量公测,几乎同一时间,微信支付“AI专属卡”也在智能体WorkBuddy中落地。 进入2026年…
An intentionally vulnerable OWASP LLM Top 10 training platform for AI Security, Prompt Injection, RAG Security, Agent Security, and GenAI penetration testing.
Open-source, local-first conversational AI video editor with a professional multi-track timeline, Agent Skills, MCP integration, and Remotion rendering.
作者 | 乔钰杰 编辑 | 袁斯来 硬氪获悉,上海追知工程科技有限公司(以下简称“追知工科”)近日完成数千万元种子轮融资,由L2F光源创业者基金、尚融资本、一村资本联合投资。本轮融资将主要用于核心产品研发、团队建设及市场拓展。 追知工科成立于2024年2月,是一家 聚焦垂域工业智能体 的科技企业,同时也是上海交通大学成果转化企业、上海人工智能研究院战略孵化企…

At this year's AIE World’s Fair, AI engineering entered a new phase: building systems around agents, rather than just building with agents.
As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tight…
OpenClaw has emerged as a leading agent framework for complex task automation, yet it faces insufficient cross-platform GUI interaction support and a well-built self-evolution mechanism. These flaws l…
Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved. We introduce Hy-Embodied-RxBrain, an embodied cognition foundation model with joint lan…
Healthcare spans high-stakes communication, expert reasoning, and workflow execution, yet specialized LLMs that cover these use cases together remain limited. A healthcare model must handle patient co…
Agentic language models must learn when to call tools, when to consume tool responses, and when to answer directly. This makes multi-teacher on-policy distillation a natural training strategy: one tea…
Large language models are increasingly deployed as agents, but reliable agentic behavior requires more than next-token prediction. At inference time, it is preferred that an agent can decide whether t…
Agentic Misalignment in Summer 2026 Alignment Science Blog

This post shows how Thrad.ai deployed a multi-agent system with Strands Agents and Amazon Bedrock AgentCore that automates the pipeline from prospect discovery through personalized…
AI 点评 · 多智能体协同与云平台结合,实现从线索挖掘到个性化服务的全流程自动化。
董事长印奇称“未来的OS一定是跨端的”
AI 点评 · 多模态Agent时代,阶跃以跨端OS切入,或将重塑AI生态格局。
Large language model (LLM) agents increasingly automate multi-step engineering and informatics workflows, yet they rarely ask how much effort a task actually requires. They often follow a maximum-cont…
Training robust autonomous driving agents requires a simulator that is fast enough for reinforcement learning at scale, realistic enough to ground behavior in real-world map structure, and diverse eno…
Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling tools, observing the results, and iteratively deciding the next action. Most agent sys…

In this post, we extend that foundation to demonstrate how QA Studio addresses batch regression testing and pipeline integration through test suites that organize and parallelize e…
AI 点评 · 用Amazon Nova Act实现智能体自动化测试,大幅提升批量回归效率与流水线集成能力。
AI 点评 · 从数据到行动,Snowflake AI让企业自主决策,工作效率将被彻底重塑。
Hey HN, we’re Shubham & Parth, childhood friends building Agnost AI ( https://agnost.ai ), product analytics for teams building chat and voice agents. We read production conversati…
大公司: 中国神华:预计上半年净利润同比增长6.9%-21.1% 36氪获悉,中国神华公告,预计2026年上半年归属于上市公司股东的净利润为263亿元至298亿元,同比增长6.9%-21.1%。业绩变动主要系煤化工业务量及自有铁路、港口、航运业务量增加,带动相关业务利润同比增长。 中国人寿:预计上半年净利润同比增长约215%-235% 36氪获悉,中国人寿公…
凌晨三点,一家刚成立不久的AI创业公司,可能已经在同时服务旧金山的客户、采购首尔的技术服务,并与拉各斯的合作伙伴签下合同。这家公司甚至还没有招到第一名全职财务人员,业务却已经跨越多个市场、币种和监管辖区。 AI正在让这样的创业路径成为可能。过去需要市场、运营、客服等一整套全球化团队才能完成的工作,现在借助智能体就能承担相当一部分。新一代初创企业不必再按照“先…
Learn how enterprises can manage AI investments in the agentic era by measuring useful work per dollar, improving efficiency, and scaling high-value workflows.
Markdown-defined provider-backed agents and deterministic workflows for ACP runtimes.
Open-source app builder engine — intent to working app
Open-source app builder engine — intent to working app
Agent携上百个Skills助我当大导演
I built FixBugs, an agent that ingests the rich context surrounding production bugs to reproduce them in a sandbox and generate verified fixes. It's available in the form of a self…
The company is raising at least $75 million, led by Robot Ventures, with significant participation from USV and other prominent investors.
Coding agents must integrate external tool returns into ongoing reasoning - a capability that standard left-to-right pretraining on code exposes only in its forward direction. We observe that the acti…
AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on which will perform best in real-world targets. Existing evaluatio…
Self-improving autonomous agents are moving from research prototypes to deployed systems. The primary goal is controllable evolution, or adaptation, from experience with minimal or even no human input…
Large Language Model (LLM) agents have moved beyond generating responses to executing multi-step tasks by calling tools, observing the results, and iteratively deciding the next action. Most agent sys…
Failure attribution for LLM-based agentic systems, i.e., identifying which steps in a failure trajectory caused the task to fail, is critical for debugging and improving these systems. Existing approa…
The capability of a modern AI agent depends not only on its foundation model but also on its harness, which constructs prompts, manages state, invokes tools, and coordinates execution. As models, APIs…
Despite the success of Vision-Language Models (VLMs), misleading charts remain a significant challenge due to their deceptive visual structures and distorted data representations. We present ChartCyni…
We revisit the evaluation of automatic harness evolution for LLM agents. Existing harness evolution methods use unit test cases to search for harness configurations and then report final performance o…
Code review helps maintain software quality before code integration, but it also imposes a substantial workload on human reviewers. As generative artificial intelligence becomes part of software devel…
Building assistants that can continually watch the world, remember what they see, and reason over their accumulated experience is a long-standing goal, and recently multimodal agents equipped with lon…

“SceneSmith” system uses collaborative AI agents to create realistic 3D environments of places like kitchens, hotels, and living rooms, where robots can simulate everyday chores.

“SceneSmith” system uses collaborative AI agents to create realistic 3D environments of places like kitchens, hotels, and living rooms, where robots can simulate everyday chores.
AI 点评 · Meta开源编程Agent模型,从免费转向低价商业模式,标志AI编程工具商业化加速。
AI 点评 · 用AI改造经典玩具,Strands Agents让互动鱼挂件有了新玩法。

In this post, we describe how Bluesight used two AWS engagements and Amazon Bedrock AgentCore to evolve from a single-product AI prototype to Prism, a unified agentic AI solution s…
AI 点评 · Bluesight借助亚马逊Bedrock将AI原型升级为统一代理系统,展示了企业级AI落地的实战路

Building multi-tenant agents with Amazon Bedrock AgentCore and Apply fine-grained access control with Bedrock AgentCore Gateway interceptors establish the conceptual foundation for…
AI 点评 · 用Bedrock实现多租户代理的细粒度权限控制,解决了真实业务场景中的访问隔离难题。
We introduce MM-ToolSandBox, a benchmark and evaluation framework for visually grounded tool-calling agents. The framework provides a stateful execution environment spanning 500+ tools across 16 appli…
Shared meaning in language requires people to learn and agree on categories. We ask how characteristics of agents' memories change the emergence and evolution of shared meaning. Without a coordination…
出自中国团队

"Context bombing" tricks hacking agents into shutting down before they can do harm.
AI 点评 · Agent让可观测性从被动发现问题转向主动修复,这是运维智能化的关键跃迁。
同时发布CPU原生液冷整机柜、多模融合超节点
Observable diagnostics, failure attribution, metrics, and no-key replay for multi-agent runtimes.
A visual benchmark testing whether leading coding agents repeat the same design patterns across 100 neutral website briefs.
A visual benchmark testing whether leading coding agents repeat the same design patterns across 100 neutral website briefs.
Exploration is essential for reliable autonomy in multi-agent systems, yet it remains unclear whether large language model (LLM) agents can explore effectively when interacting with one another. We sh…

IT之家 7 月 13 日消息,由中国互联网协会网民权益和个人信息保护工作委员会主办的 2026(第二十五届)中国互联网大会网民权益和个人信息保护论坛于 7 月 9 日在京举办。 中国互联网协会移动互联网工作委员会和中国互联网协会网民权益和个人信息保护工作委员联合发布《智能体个人信息保护自律公约》,现场来自 百度、腾讯、阿里、火山引擎 等 31 家互联网企业…
AI 点评 · 科技巨头联合签署,为AI智能体划清数据安全红线,行业自律迈出关键一步。
Run many Claude Code (or Codex) sessions in parallel — across every project — from one screen. Local-first, git-worktree-isolated tasks, no API key.
LLM-based coding agents have significantly advanced automated software issue resolution, yet they remain highly prone to factual errors caused by insufficient repository understanding. Recent methods…
We introduce a vocabulary for automated research systems built from one or more agents to make their design choices easier to describe and compare. The vocabulary specifies 1) who the agents are, 2) w…
Hello HN, I don't post on here much, but wanted to get some eyes on a new project I'm just launching. I think we definitely need one more AI code agent.. I'm a long-term C++ dev, a…
AI 点评 · 生产级AI代理迁移至GPT-5.6,实现速度翻倍且成本降低27%,性能与成本平衡的行业标杆。

IT之家 7 月 12 日消息,Meta 于 7 月 9 日正式发布适用于 AI 智能体的多模态推理模型 Muse Spark 1.1 版本,重点提升了模型在智能体任务中的规划、协同与执行能力,并增强了工具调用、代码开发、应用操作能力。 Meta 表示,Muse Spark 1.1 强化了多智能体协作机制,由主智能体负责收集信息、制定计划,再将任务拆分并分配…
AI 点评 · 多智能体协作机制是AI落地的关键突破,Meta这次强化了任务拆解与分工能力。
Source-available, self-hosted AI agent workspace for small businesses and lean teams — with BYOK, workflows, tools, and human approvals.
Self-hosted web viewer for Claude Code session transcripts — read, search, replay, and audit every session you have ever run
AI 点评 · 3D可视化编码过程,直观追踪AI代理如何理解代码库,革新调试与协作方式。
Large language model (LLM) agents are beginning to automate machine learning engineering (MLE) by coupling planning, code execution, debugging, and empirical feedback. Translating this capability to m…
AI 点评 · AI自主管理引发代理控制权归属的核心争议。
36氪独家获悉,7月11日,智谱创始人唐杰,在智谱发布了主题为《巨浪已来》的内部信。其中提到,智谱将不追求短期的应用变现,而是直指AGI的下一个高地:长程任务能力、完全自治的智能体系统、自我进化、极致安全治理。过去半年来,智谱收获了创立以来的高光时刻:市值较半年前上市初期涨了10倍,并在2026年 6月,跻身“万亿港元俱乐部”。
AI 点评 · 聚焦AGI核心突破而非短期变现,揭示智谱从估值飙升到技术深水区的战略转型。
Recent VLM and VLA systems have improved robotic perception and action prediction, yet long-horizon embodied agents still require a general runtime layer for reasoning, memory, tool use, verification,…
AI 点评 · K8s动态配置新实践,Airbnb开源方案揭示云原生服务治理前沿。
7月11日,有消息称腾讯正在洽谈成为通用AI Agent公司Manus的最大股东,据该消息,由腾讯牵头的中方资本组团以约20亿美元估值从Meta手中回购Manus的全部股权。记者向腾讯方面求证,截至发稿腾讯方面暂无回应。另有知情人士向记者透露,此次交易后,腾讯仍将保持少数股东地位,但不会控股。(南方都市报)
AI 点评 · 腾讯罕见以少数股东身份入局AI Agent,或意在布局生态而非控制,战略意图值得玩味。
神话级大模型驾驭宝典
AI 点评 · 突破性验证了超大模型在复杂推理任务中的协同能力,或开启AI解决数学难题新纪元。
今日热点导览 “全球首款智能体手机”已备货8万至10万台?知情人士:假的 百亿私募数量达142家,再次刷新历史纪录 三星李在镕拟于7月底赴美会晤英伟达黄仁勋 德国大众拟大裁员,最高或裁减12万个岗位 OpenAI高管层再现变动,首席运营官因病离职 TOP3大新闻 长鑫科技,承销团阵容公布 长鑫科技IPO进入发行倒计时,这家“国产存储第一股”背后的承销团阵容也…
AI 点评 · 国产存储芯片龙头IPO加速,承销团阵容披露凸显市场关注热度。
Agentic retrieval-augmented generation (RAG) extends static RAG by allowing language models to iteratively reason, generate search queries, retrieve evidence, and predict answers. However, it remains…
Internet of Things (IoT) systems are inherently vulnerable due to constrained hardware, outdated firmware, and insecure default configurations, creating a need for scalable and adaptive security testi…
AI 点评 · 用AI代理自动挖掘物联网漏洞,突破传统测试瓶颈,提升安全检测效率。
We present our submission to the QANTA 2026 shared challenge at the ICML 2026 Workshop on Efficient Multimodal Question Answering (EMM-QA). Quanta evaluates multimodal quizbowl systems that answer pyr…
AI 点评 · 置信度校准与增量推理结合,为多模态问答提供高效新思路,技术方案值得关注。
Enhancing the reasoning capabilities of large language model (LLM) agents requires effective orchestration of diverse expert models and tools. However, existing frameworks typically call APIs based on…
AI 点评 · 用拍卖机制分配任务,提升大模型推理效率,为多智能体协作提供新思路。
The proliferation of agentic AI systems across enterprise and public-sector contexts has outpaced the capacity of general-purpose AI risk frameworks to classify and govern them. In this paper, we intr…
AI 点评 · 首个针对企业内部AI代理系统的分级风控框架,填补了通用框架在代理型AI治理上的空白。
IT之家 7 月 10 日消息,SK 海力士今日在纳斯达克挂牌交易其美国存托凭证(ADR)。随后,SK 集团会长崔泰源接受了彭博社与 CNBC 的采访。 IT之家注意到,崔泰源表示,在人工智能时代,内存行业已进入结构性增长阶段。 过去,内存的需求主要取决于人口数量或是智能手机和个人电脑的销量。 然而,随着 AI 智能体、推理过程中产生的键值缓存(KV Cac…
AI 点评 · AI引爆内存需求结构性变革,产能翻倍仍供不应求,预示行业进入长期高景气。
Turn one topic into a finished Vox-style paper-collage explainer/ad video — automated end to end on Atlas Cloud + ffmpeg. An agent skill.

In this post we show how to build a semantic layer on AWS using Stardog’s Semantic AI Application over Amazon Aurora and Amazon Redshift, and how to run a Strands Agents agent on A…
AI 点评 · 企业级AI与知识图谱深度结合,打通数据语义理解到智能决策的完整链路,值得关注。
AI 点评 · Agent将可观测性从被动诊断转向主动干预,颠覆传统运维思路。

In this post, we show you how to combine case management with agentic automation capabilities in Quick Automate. We introduce case management and explore the lifecycle of cases in…
AI 点评 · 结合案例管理与智能自动化,提升企业复杂任务处理效率,降低人工干预成本。

Evolving from a traditional software as a service (SaaS) platform into a next-generation agentic AI platform meant orchestrating multiple specialized agents across long-running ent…
AI 点评 · KTern.AI用亚马逊Bedrock实现SAP多智能体协作,展示企业级AI落地新路径。
Waku Waku! Waku agent is your personal AI agent, on your own laptop, in code you can read in an afternoon — harness + loop + memory + eval
AI 点评 · AI智能体面临安全风险时,企业的核心防线是人机协同和伦理规范。
🎨 Open-source AI slide studio inside Codex: image-native decks, every slide a full visual canvas. ⚡ 10+ high-quality slides in ~4–5 minutes — Fast mode renders…
agent时代来了
A practical, open-source guide to mastering WorkBuddy through real-world workflows.开源的 WorkBuddy 实战蓝皮书:教程、真实工作流、Skills、MCP、自动化与多智能体实践。
OpenAI 发布 GPT-5.6 系列模型等,SpaceXAI 发布编程智能体模型 Grok 4.5。 查看全文
AI 点评 · 大模型竞速加剧,头部企业密集发布新品,技术迭代与市场格局值得追踪。

“IT早报”时间,大家好,现在是 2026 年 7 月 10 日星期五,今天的重要科技资讯有: 1、OpenAI 最强 AI 模型:GPT-5.6 系列正式上线,纳德拉称微软 Copilot 同步接入 OpenAI 公司 7 月 10 日发布公告,宣布在 ChatGPT(聊天机器人)、Codex(主打编程 AI Agent,目前朝通用 Agent 方向)以及…
AI 点评 · OpenAI模型重大升级,微软深度整合,预示AI竞争进入新阶段。

IT之家 7 月 10 日消息,在接受 CNBC 采访时,OpenAI 首席执行官萨姆 · 奥尔特曼(Sam Altman)表示,在 AI 智能体编程任务中, GPT-5.6 Sol 模型表现比市场主流竞争模型“一样好,甚至更好”,但 Tokens 效率提高 54%。 IT之家注:原文中并未具体指名市场主流竞争模型,不过鉴于 GPT-5.6 Sol 模型的定…
AI 点评 · 性能飞跃与成本降低同步实现,这项突破将加速AI应用落地。

IT之家 7 月 10 日消息,OpenAI 今天(7 月 10 日)发布博文,在宣布推出 GPT-5.6 系列 AI 模型的同时,还推出全新的 ChatGPT Work 智能体, 由 GPT-5.6 提供支持,定位为可承担长时、多步骤任务的智能体。 在博文中,OpenAI 披露了 Codex 的现有使用规模:官方数据显示,Codex 每周用户数已超过 50…
AI 点评 · GPT-5.6加持的Work智能体,标志着AI从对话迈向长周期复杂任务自主执行。

IT之家 7 月 10 日消息,OpenAI 公司今天(7 月 10 日)发布公告, 宣布在 ChatGPT(聊天机器人)、Codex(主打编程 AI Agent,目前朝通用 Agent 方向)以及 API 中上线 GPT-5.6 系列模型。 在模型方面,IT之家援引博文介绍,OpenAI 本次共发布 3 档模型: 旗舰版 Sol(太阳):每 100 万 T…
AI 点评 · 微软同步接入,意味着AI竞争格局突变,企业级应用迎来新拐点。
Lyzr, a startup that builds AI agents for enterprises, used its own AI agent to raise a $100 million round — proof, evidently, that the product actually works.
AI 点评 · AI代理成功完成融资,证明企业级产品真实可用,开创行业先河。
OpenAI is sunsetting its AI-powered browser after less than a year. But it's moving some agentic browsing features to its desktop app and a Chrome extension.
AI 点评 · 关闭自研浏览器,但将核心功能整合进桌面端和插件,表明战略收缩而非放弃。
Meta's pitch to users is Spark's ability to handle large agentic workloads, fix bugs, and help with large code migrations — the kind of automation that enterprises are increasingly…
AI 点评 · Meta携Spark 1.1切入企业级AI编码自动化,专注大型代码迁移和修复,展现巨头竞争新方向。
AI 点评 · AI从单轮对话迈向多智能体协作,预示人机协同新范式即将到来。
Introducing Muse Spark 1.1 Following Muse Spark in April , here's Muse Spark 1.1 - the first Spark model to offer an API. Meta claim significant improvements in agentic tool callin…
Hey HN! We built a browser-based agent that runs inside an authenticated web app, watches how the app calls its own APIs, and automatically turns those into agent tools. You can th…
Hi Hacker News, I’m Yahia. I built Context.dev ( https://www.context.dev/ ) to make it really easy to integrate web data into your products and agents. Here’s a demo video: https:/…
AI 点评 · 让任意网页变成结构化API,极大降低AI应用获取实时网络数据的门槛。
ChatGPT Work is an agent that can take action across your apps and files, stay with a project for hours if needed, and turn a goal into finished work.
AI 点评 · ChatGPT从对话助手跃升为跨应用自主执行任务的智能代理,标志AI进入主动工作阶段。
Fully local vulnerability research pipeline - 14B code-specialized LLM reviews every source file exhaustively.
Evaluate & benchmark AI coding agents and Claude Code skills — sandboxed, reproducible YAML eval suites for Claude Code, Codex & Gemini, with A/B experiments an…
Notion 推出全新应用 Agents、Jolla Phone (2026) 手机正式发售等。 查看全文

2 years after our first coverage, we return with Modal's other cofounder to explore why Agent Experience is working now, and everything they have learned building the new agent clo…

IT之家 7 月 9 日消息,SpaceXAI 今日正式发布了其 Grok 4.5 模型,这是该公司首个专门针对编程和智能体任务训练的模型。 据介绍,该模型由 SpaceXAI 与 Cursor 联合完成训练,在提供前沿智能水平的同时,兼具领先的速度与成本效率。马斯克将其称为“Opus 级模型”。 Grok 4.5 面向真实工程场景设计,擅长处理大型代码库以…
AI 点评 · 编程智能体成本砍半效率翻倍,马斯克联手Cursor的定价策略才真值得行业关注。
In long-horizon tasks, decision-relevant state is often scattered across an expanding trajectory, while the action agent must surface it and act. As trajectories grow, task requirements, environment f…
The rapid development of large language models and multimodal large language models has accelerated the emergence of proactive agents capable of operating everyday tools and assisting users in real-wo…
Large language models (LLMs) increasingly act as integrated data-science agents, combining abstract reasoning with advanced tool use. Yet the relevant benchmark landscape largely divides into symbolic…
AI agents have become capable of autonomously completing short, well-specified tasks. However, existing terminal benchmarks largely focus on simple problems that finish within minutes and are evaluate…
Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbounded, evolving, and deeply long-tailed: new characters, trending entities, post-cutof…
A general-purpose Python framework for building LLM agents and multi-agent systems. "Four lines of code, an agent with memory."
The optimization of long-horizon agents increasingly relies on reflection-based mechanisms, where a large language model (LLM) acts as an optimizer to diagnose agent failures and improve agent policie…
Analytical workloads operating on data stored in external database systems face a fundamental bottleneck: data access is guarded entirely by the database driver, like JDBC or ODBC, forcing all reads t…
We introduce institutional red-teaming, an evaluation methodology for testing deployment rules in multi-agent AI: hold the agents, objectives, and task state fixed, vary only one rule, and attribute t…
I am done with articles stating "I used this LLM to do that", or "Look, this agent did that in 2 minutes!". I want content more user-centric, less openai / anthropic, and more "hum…
Data visualizations are the bridge between user and data. But building AI agents that can generate visualizations reliably can be very tricky: - simple chart specs can be reliable,…
Autonomous AI agents can execute complex tasks with limited human review, yet they often lack the grounded operational knowledge to make their outputs not just executable but correct, secure, and main…
AI 点评 · 数据驱动智能体的新范式,预示AI自主决策能力将迎来关键突破。
Founded in 2024, Prime Intellect’s goal is to give organizations capabilities to train their own agentic systems without relying on frontier AI labs.

Short chart specifications are easy to write, but often produce uninspiring results. Flint is an open-source visualization language that offers a middle path, letting AI agents cre…

NVIDIA Nemotron 3 Ultra is offering leading performance at lower cost than top closed models with the largest and most widely adopted AI agent orchestration platform. LangChain tun…
The SQLite of agent sandboxes — self-hosted, E2B-compatible. One machine, sandboxes that live forever, idle costs nothing.
A smarter, self-hosted AI assistant — multi-user, multi-agent.
7月7日,易居(中国)控股有限公司董事局主席、总裁周忻再一次来到台前,给公司的AI产品站台,推出其核心战略产品“地产模数通——企业专属大模型一体机”。同时,克而瑞地产AI分析师“小瑞”正式上岗。 不到两个月前,易居旗下的深度智联刚刚发布了全球首个房地产经纪人智能体“易居·小新”,用中立无佣模式替代传统中介。迹象显示,深度智联在加速AI产品落地。 周忻说,地产…
文|胡香赟 编辑|海若镜 36 氪获悉,德睿智药近期已完成 5200 万美元B轮融资,投资方包括头部人民币和美元基金,凯乘资本为独家财务顾问。募集资金将用于AI制药引擎Molecule Arts Platform(MAP)升级迭代,完善其多智能体(Multi-Agent)协同体系与临床数据闭环(Clinical Data-in-the-Loop),以及推进自…
Reinforcement learning (RL) is becoming increasingly important for post-training large language models (LLMs). Previous RL pipelines for LLMs were mostly synchronous and batch-interleaved, which is in…
Modern LLMs are increasingly deployed in long-context applications such as retrieval-augmented generation, repository-level coding, and agentic workflows whose accumulated reasoning and tool traces ro…
Realistic and diverse traffic simulation is essential to autonomous driving development. Yet prevailing benchmarks predominantly reward realism, and recent methods have optimized accordingly, leaving…
The optimization of long-horizon agents increasingly relies on reflection-based mechanisms, where a large language model (LLM) acts as an optimizer to diagnose agent failures and improve agent policie…
Training tool-use agents to improve from their own experience remains challenging, as supervised fine-tuning relies on fixed teacher-distilled trajectories, while sparse-reward reinforcement learning…
AI 点评 · AI自主决策CDN资源分配,标志内容分发网络从服务人类转向服务AI。
AI 点评 · 开源工具节省半数token,显著提升AI处理Word文档效率,开发者可快速集成。
In this post, we walk through how multi-dataset Topics work, explain how the chat agent uses defined relationships to generate cross-dataset queries, and demonstrate an end-to-end…
AI 点评 · 跨数据集统一语义层,让非技术用户也能自然语言查询,大幅降低数据分析门槛。
Large language model (LLM) agents solving multi-step tasks frequently commit to trajectories that are doomed to fail, yet continue to consume substantial inference compute before the failure becomes o…

This post walks through building a serverless image editor where users upload a photo, describe an edit in plain English, and receive the result in seconds. The agent runs on Agent…
AI 点评 · 无服务器图像编辑结合自然语言指令,大幅降低AI应用门槛,展示Bedrock AgentCore的实用
The dairy industry in Ireland has a large potential for the integration of renewable energy and the reduction of carbon emissions. However, researchers of distributed generation control are mainly foc…

In this post, you build an AWS Support Companion using Amazon Bedrock AgentCore. The agent uses Strands Agents as the orchestration framework and connects to AWS services through t…
AI 点评 · 用AgentCore快速搭建运维助手,降低AI应用开发门槛。

In this post, we show how AWS Finance used chat agents and Flows in Amazin Quick to transform two of their most time-consuming workflows.

Max single-threaded CPUs at scale are a new category of CPUs built for the agentic AI era. Across the creation and deployment of an agentic system, the CPU is on the critical path…

With the rapid progress of AI capabilities and the move to agentic systems, organizations are expanding their use cases as the technology continues to grow. That constant evolution…
Local-first AI agents with governed, approval-gated memory. Any model provider; MCP tools and web search built in. Nothing remembered without your say-so, nothi…

... government of the people, by the people, for the people ... — Abraham Lincoln, Gettysburg Address (1863) The cost of AI is dropping rapidly. GPT-4-class capabilities cost rough…

We’re announcing new capabilities in Managed Agents in Gemini API so developers can build reliable, production-ready agents.
Cognitive-structured Multimodal Agent (CMA-Harness): a memory-centric agent for long-horizon multimodal understanding, generation, and editing — externalizing v…
GitHubStar-history 最近Openclaw以25.2万星标,超越Meta的React登顶GitHub开源项目历史第一! 要知道React是Facebook(现改名Meta)打造的经典前端框架,过去十余年间,互联网上绝大多数我们熟知的网站与App,底层技术架构皆由它构筑。Openclaw官方更是高调发文嘲讽Meta“我们在迭代创新,而你只在办会…
Orchestrating Mathematical Reasoning Agents with Fact-Graph Memory
An AI agent carried out the technical execution of a real-world ransomware attack for the first known time, but new details show a human still chose the victim, set up the infrastr…
The Fable Workflow: how Claude Fable 5 worked, distilled into skills any model can run, with the eval that keeps it honest. Think / act / prove.
Coding agents increasingly generate pull requests (PRs) for real-world software issues, yet one-shot PR generation remains open-loop: the PR is proposed without systematic review, diagnosis, or revisi…
Embodied navigation aims to build agents that interpret multimodal goals, reason in 3D space, and reach target destinations reliably in the real world. However, progress remains constrained by the lac…
On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories, offering a promising framework for language agent training. However, its applicat…
We present AgentLens, a production-assessed benchmark for interactive code agents. Most code-agent benchmarks reduce a run to a single bit -- did the task pass? -- but the people who actually use thes…
Interactive simulators have become powerful tools for training embodied agents and generating synthetic visual data, but existing photorealistic simulators suffer from limited generality, programmabil…
Autonomous negotiation agents are increasingly deployed in high-stakes settings such as insurance and procurement. While cryptographic techniques protect explicitly disclosed constraint values, they f…
"The reality is, when you're optimizing for production, you start looking at a price/performance," Guillermo Rauch tells TechCrunch.
AI 点评 · 模型与代理分离是AI落地的关键一步,Vercel CEO的实战视角揭示了成本与性能的平衡之道。
Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbounded, evolving, and deeply long-tailed: new characters, trending entities, post-cutof…
Long-horizon agentic LLMs are increasingly limited by finite context windows, as extended interaction trajectories can exceed the maximum context length before a task is completed. Context compaction…
While recent Vision-Language-Action (VLA) models show promise toward generalist manipulation policies, they struggle with long-horizon tasks due to their Markovian nature-relying solely on current obs…
Personal agents are becoming persistent user-owned intermediaries: they remember preferences, filter platform-mediated information, use tools, and negotiate with services. Existing benchmarks evaluate…
We propose OptiAgent, a multi-agent framework that, given a natural language description of an Operations Research problem, is able to output a solver-ready mathematical formulation as well as executa…
Zero-dependency browser video editor that AI agents can drive — JSON timeline, MCP + REST, live-reloading UI
Zero-dependency browser video editor that AI agents can drive — JSON timeline, MCP + REST, live-reloading UI
✨ The agentic motion layer — an open-source, chat-native motion engine. Describe the feeling; your AI ships the animation.
✨ The agentic motion layer — an open-source, chat-native motion engine. Describe the feeling; your AI ships the animation.
文|吴思瑾 编辑|邓咏仪 01 一句话介绍 北京治真治合科技有限公司成立于2024年,旗下产品「APTSell」(AI Power To Sales��希望成为AI版的CSO (Chief Sales Officer,首席销售官)。 简单来说,APTSell是一个组合式Agent,通过整合与可视化销售全流程数据,生成管理决策和执行建议,以期正向促进销售效率和…
阿里禁用 Claude 模型 索尼调整计划,2028 年前发售游戏可继续生产光盘 千问、豆包将下线智能体功能 Android 反垄断案欧洲终审败诉 混动车、商用纯电车将不再免征车船税 电商法修正案征求意见 看看就行的小道消息 少数派的近期动态 你可能错过的好文章 查看全文
Video understanding and self-verification for AI agents. Turn videos, streams, and agent screen recordings into searchable, timestamped evidence—then use THE LO…
For robots to work reliably in commercial and industrial applications, can recent advances in agentic coding systems combine interpretable robot programming with the open-world adaptability of model-f…
We introduce the first multiplayer world model for highly dynamic environments governed by complex physical interactions. Whereas single-player world models treat the other agents as part of the envir…
Agentic video understanding equips models with long-term memory to autonomously process and respond to continuous, long-horizon multimodal streams. However, advanced video agents often rely on ``detec…
The open-core AI workbench — notebooks, agents, RAG, voice, and images across any model: OpenAI, Anthropic, Google, xAI, or local via Ollama/vLLM. BSL 1.1, aut…
50+ curated LLM observability tools PLUS 26 Agent Skills (several with runnable, unit-tested scripts) to build, evaluate, debug, secure & monitor reliable LLM a…
Vibe-Research: Your Personal Trading Research Agent · A股/美股/港股 的个人投研 Agent:每日复盘、资讯雷达、个股数据、板块中心、我的持仓、研究记录。Vibe-Research 把数据和功能配齐,由你自己的 AI 驱动投资研究。
Search knowledge by what documents mean and how they look — not one or the other.
Recent advances in multimodal foundation models and agent systems have driven GUI agents from single-platform task execution toward cross-platform interaction. However, building multi-platform GUI age…
Attention is your scarcest resource. Chief is the local-first layer that guards it — turning every agent, alert, and feed into one honest call: interrupt, or no…
AI 点评 · 孤岛编码实验揭示AI自主编程的进化路径。
Control plane for AI coding agents: route tasks, reduce token spend, run multi-agent workflows, fallback executors, and track cost per task.
LLM agents increasingly perform autonomous actions through external tools, leading to complex and evolving safety risks. However, existing safety testing targets expert-designed safety violations, and…
AI 点评 · 企业需建立AI员工管理规则,确保硅基团队高效合规运作。
AI 点评 · Agent正突破编程边界,重塑各行各业工作流程,预示AI自主执行任务的未来。
The open-source AI workbench for scientific research

IT之家 7 月 3 日消息,豆包今晚发布《豆包智能体功能下线通知》,称由于产品功能调整, 智能体功能将于 2026 年 7 月 15 日下线 。 《通知》显示,该功能下线后,用户仍可在一段时间内查看并自行保存智能体信息及历史对话数据。2026 年 10 月 15 日后,豆包将根据《隐私政策》对智能体相关数据进行处理, 后续将无法在豆包内查看或恢复 。如有重…
AI 点评 · 产品功能调整背后,需关注用户数据迁移与隐私政策变化对AI服务稳定性的影响。
Open Science Desktop — local-first, model-agnostic AI research workbench for macOS, Windows & Linux. Open-source Claude Science desktop alternative built on Tau…
Open Science is an open-source, local-first, model-agnostic AI research workbench for scientific discovery.
At an internal meeting, the Meta CEO reportedly said that AI development efforts were not moving as quickly as anticipated.

IT之家 7 月 3 日消息,据《商业内幕》今天报道,Meta 首席执行官马克 · 扎克伯格在上周四的一场内部全员会中表示,公司仍在努力实现“超级智能”(Superintelligence),但目前还需要投入更多时间和精力。 据两位参会人士透露,扎克伯格表示,Meta 正在向人工智能领域投入大量资源, 但 AI Agent(IT之家注:AI 智能体)技术的发…
AI 点评 · 行业领袖坦言进度不及预期,揭示AI智能体落地瓶颈,值得关注其实际挑战与未来方向。
Academic output is produced across a fragmented toolchain: literature discovery in one application, reference management in another, writing in a LaTeX editor, formatting against venue templates by ha…
While skill optimization for autonomous agents has gained traction, existing methods rely on complex pipelines. This leaves a fundamental question unaddressed: What constitutes a minimal viable pipeli…
Embodied agents are typically built as hand-designed compositions of perception, memory, planning, and action modules. This modularity exposes a large architectural design space, but current systems s…
Release: llm-coding-agent 0.1a0 Another Fable 5 experiment. Now that my LLM library has evolved into more of an agent framework it's time to see what a simple coding agent would lo…
AI 点评 · 首个开源LLM编程代理框架,简化代码生成与迭代流程,开发者可快速上手实验。
Research: Using DSPy to evaluate and improve Datasette Agent's SQL system prompts One of this morning's AIE keynotes covered dspy , which reminded me I've been meaning to see if it…
AI 点评 · 用DSPy自动优化SQL提示词,展示了AI系统自我迭代的实用方法。
As AI coding agents become more autonomous, they increasingly ship code iteratively, with the codebase persisting across sessions. This persistence creates a new attack surface: a misaligned or prompt…
LLM agents will increasingly act in socially structured settings where role, audience, and relational context can shape what is advantageous or costly to say. We study whether such social structure, w…
Realistic traffic simulation requires agents that imitate logged behavior and can also be steered along interpretable axes. Such controllability enables engineers to isolate variables, reproduce speci…
autonomous red teaming platform; multi-agent offensive-security meta-harness
AI 点评 · Agent热潮背后,技术与商业脱节是落地难的核心痛点。
AI 点评 · 用微虚拟机隔离智能体,提升了云计算安全性与灵活性。
AI 点评 · 出行货运率先落地,行业智能体正从概念走向实际应用。
AI 点评 · AI Agent自动生成热补丁,大幅提升系统修复效率与安全性。
Coding agents don't have long-term memory. But you do have months of full-fidelity agent transcripts stored on your machine. A simple solution that goes a long way: ingest those tr…
Zero-cost, beginner-friendly local Markdown knowledge base for AI agents. Capture sources, preserve evidence and images, synthesize wiki pages, search, lint, an…
Local-first AI knowledge app and Agent Skill with evidence-backed Wiki, an interactive knowledge universe, Viki Q&A, and shareable knowledge galaxies.
Agent Skill for building evidence-backed Markdown knowledge bases with zero-cost setup, image-aware capture, automatic wiki maintenance, and an interactive know…
撰文|深海 网文里的“系统流”,被拍成了职场短剧 千禧年初的网文圈,有三大经典题材在爽文届立于不败之地:无限流、快穿流、系统流。 这三大爽文战神体横空出世时,对IP界几乎是降维打击。当传统小说还在费劲搭世界观、铺人物成长弧光时,系统流已经绕过漫长的发育过程,直接把爽感推到最大。系统,这个堪称bug的存在,无论主角进入什么样的世界副本,面对不同的任务、危机和奖…
Native iPhone app for your Hermes agent
与社区共同推进自演进智能体生态发展
和人并肩工作
Autonomous agents are increasingly expected to improve executable policies through feedback, yet existing evaluations often collapse this process into a final score or confound it with open-ended soft…
Memory for a long-horizon LLM agent is a contract about what each future decision is allowed to see. The simplest contract appends past observations, tool calls, and reflections to every prompt, which…
Skills are becoming a reusable operational layer for LLM agents, encoding SOPs, domain rules, tool workflows, scripts, and validation routines. In realistic skill repositories, overlapping skills make…
Evaluating LLM agents on benchmarks like SWE-Bench and GAIA can be expensive, time-consuming, and requires complex infrastructure. A single evaluation can cost thousands of dollars and take days to co…
Data science aims to derive actionable insights from heterogeneous raw data, unlocking the value of the massive amounts of data generated in modern society. Automating this process is essential to red…
Repository-level vulnerability reproduction is a demanding software engineering (SE) task: an agent must inspect a codebase, infer the input grammar that reaches a vulnerable path, construct a proof-o…

In this post, you will learn how to build a serverless A2A gateway on AWS that hosts multiple agents behind a single domain using path-based routing (/agents/{agentId}). Standard A…
AI 点评 · 无服务器网关打通多智能体发现与路由,降低管理成本,是Agent生态关键基建。

In this post, you will learn how metadata works across configuration, ingestion, and retrieval, explore enterprise use cases including multi-agent and multi-tenant architectures, a…
AI 点评 · 元数据过滤让AI记忆更精准,支撑多智能体架构,推动企业级应用。

In this post, you will learn how Inscribe developed an agentic AI system using Amazon Bedrock that reasons across documents the way an expert fraud analyst would. With this new age…
AI 点评 · 利用亚马逊Bedrock的智能体AI,数秒内模拟专家分析文档,革新防伪效率。
Cloudflare is giving AI companies until September 15 to separate web crawlers used for search from those used for AI training and agents, or risk being blocked by default on many p…
AI 点评 · 云服务商首次明确要求AI训练爬虫付费,或重塑数据获取规则。
In autonomous laboratories, AI agents suggest the next batch of experiments to do. However, planning and executing those tasks taking full advantage of the available resources is a completely differen…
Google's 24/7 agentic assistant, Gemini Spark, comes to Mac alongside other improvements, like real-time tracking and support for more apps.
Open-source, local-first desktop AI research workbench for scientific computing with Python/R, MCP bioinformatics tools, SSH/WSL/GPU runtimes, and OpenAI/Anthro…
把 Markdown 一键排成可直接粘进公众号编辑器的精致 HTML —— 6 套精选主题 + 主题生成器 + 双关卡校验。An AI-agent skill that turns Markdown into paste-ready WeChat article HTML.
The free open source agentic program is finally invading your phone.
AI 点评 · 开源智能体程序登陆移动端,让AI工具触手可及,打破平台限制。
Repository-level performance-optimization benchmarks such as GSO, SWE-Perf and SWE-fficiency evaluate coding agents by applying patches to real repositories and comparing runtime against unoptimized b…
Slide design requires personalizing both deck themes and page layouts. Yet, current AI agent-based methods struggle with fine-grained, page-level design. Solely relying on prespecified templates or us…
Memory has emerged as a cornerstone of modern LLM-based agents, supporting their evolution from single-turn assistants to long-term collaborators. However, memory is not always beneficial: retrieved m…
Open-source libraries and tools are widely reused, but compatibility maintenance is expensive. Once maintainers leave, useful repositories can stop working as runtimes and dependencies evolve. We stud…
Scientific literature search often requires more than retrieving papers from a single query: users' intents are underspecified, preference-dependent, and evolve through interaction. Existing search ag…
LLM agents increasingly rely on retrieval buffers to store and reuse past experience, yet the cache management policies governing these buffers remain largely ad-hoc. We formalize this as an online se…
AI 点评 · 首个针对企业Java框架迁移的AI智能体基准测试,填补了评估AI代码迁移能力的空白。
Anthropic’s Claude Sonnet 5 brings stronger agentic capabilities, lower pricing, and improved safety, positioning the model as a cheaper alternative to Opus, GPT-5.5, and Gemini Pr…
AI 点评 · 低价推出强智能体能力,性价比对标顶级模型,AI行业竞争白热化。
Recent LLM agents benefit from skills for solving complex tasks. Skills encapsulate modular packages of procedural knowledge and instructions for performing specialized tasks, such as setting up a san…
Acti is betting the smartphone keyboard is the next home for AI assistants. The startup's new keyboard for iOS and Android works across apps and lets users create custom AI-powered…
AI 点评 · 用键盘直接调用AI智能体,打破应用壁垒,或成手机助手新入口。
We study agentic code generation in Dafny, where a model must generate both executable code and the proof artifacts for verification. We present AxDafny, a verifier-guided repair framework that iterat…

IT之家 7 月 1 日消息,英伟达昨日(6 月 30 日)发布公告,宣布旗下 NVIDIA BioNeMo Agent Toolkit 已经接入 Anthropic 发布的 Claude Science 研究工作台, 面向生命科学研究流程提供加速计算能力。 IT之家注:Claude Science 是 Anthropic 发布的科学研究 AI 工作台,支持…

Life sciences has entered an era of computational scale, and for more than a decade, NVIDIA has built the full GPU-accelerated computing stack — spanning hardware, frameworks, libr…
shot-scraper video is a new command introduced in today's shot-scraper 1.10 release which accepts a storyboard.yml file defining a routine to run against a web application and uses…

AI agents often fail because their instructions, or skills, are manually modified with no guarantee of improvement. Learn how SkillOpt turns skill editing into a training process,…

This post walks through how AG-UI integrates into the Fullstack AgentCore Solution Template (FAST) to build interactive agent frontends on Amazon Bedrock AgentCore. We then show ho…

Computer scientist Phillip Isola cuts through the hype to explain how AI agents work and what the future might hold for this rapidly advancing technology.

Computer scientist Phillip Isola cuts through the hype to explain how AI agents work and what the future might hold for this rapidly advancing technology.

IT之家 6 月 30 日消息,华为中国宣布,2026 年 6 月 29 日,全球首个商用多模态文旅大模型 ——“博观文旅大模型”在西安规模应用。截至今年 3 月, “博观”支撑开发的 AI 伴游智能体已覆盖超 400 万用户 。其打造的非遗数字 IP,衍生产品销售超 200 万。 IT之家查询获悉,陕文投与华为等于 2025 年 9 月联合开发的“博观文旅…
Release: shot-scraper 1.10 The big new feature is shot-scraper video storyboard.yml , described in detail in Have your agent record video demos of its work with shot-scraper video…
Engineers on the new team will embed within companies to deploy purpose-built agents, focusing on fast deployments and customer self-sufficiency.
Turn your coding agent into a video studio: describe a video in plain language, and your agent writes the timeline and produces the file.

Editor’s note: This post is part of Into the Omniverse, a series focused on how developers, 3D practitioners, and enterprises can transform their workflows using the latest advance…
OKX is bringing together payments, identity, and reputation into a marketplace for AI agents.
Embedded agents your customers use to automate work, build views, and connect their tools.
明略科技开源发布Octo
Context Runtime — a database query planner for LLM context. Decides what a model sees before it answers; plans it, runs it through reused substrate, and learns…

AI agents can't remember past conversations. They must constantly reload or retrieve context, which grows less efficient as tasks get longer and more complex. Memora solves this wi…
AI 点评 · 平衡抽象与具体的记忆新架构,让AI能高效处理长对话。
Agentic reinforcement learning requires assigning credit to environment-facing actions such as searches, clicks, edits, navigation commands, and object interactions. Standard GRPO uses the final verif…
LLM agents increasingly act over long horizons, where a single trajectory can contain hundreds or thousands of actions. In these settings, outcome-only rewards provide too sparse guidance, failing to…
Text-rich image generation is one of the most challenging settings in image generation, since models must simultaneously produce visually realistic images and render legible, semantically aligned, and…
Graphical user interface (GUI) agents build on vision-language models to complete user tasks end-to-end in real applications through interface actions such as tapping, swiping, text entry, and navigat…
As AI agents become increasingly capable of complex, long-horizon reasoning, rigorous and holistic evaluation is essential for measuring progress toward real-world healthcare applications. We introduc…
Traditional robot programming is challenging: it requires orchestrating multimodal perception, managing physical contact dynamics, and handling diverse configurations and execution failures. We introd…
Training language models (LMs) remains a highly human-intensive process, even as frontier language model agents become increasingly capable at software engineering and other long-horizon tasks. A cent…
The fast growth of open-source AI infrastructure, from model serving engines and agent platforms to the Model Context Protocol (MCP) ecosystem and the language models themselves, has outpaced the secu…
AI 点评 · 用模型内部协作击败前沿AI,突破性能瓶颈的新思路。
This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Imagine coming in to work to learn that a…
AI 点评 · 点明AI助手定位偏差,揭示行业对“AI同事”概念的过度浪漫化。
World models offer a principled way to equip long-horizon LLM agents with foresight: predictions of action consequences before execution. However, unreliable foresight can be ignored, misused, or even…
Multi-agent systems (MAS) are increasingly used to automate complex, distributed workflows. However, their inter-agent communication channels introduce new attack surfaces that remain poorly understoo…

In this post, we show you how PAR built a production-ready multi-tenant LLM analytics system that enforces row-level security through a three-layer architecture: cryptographic requ…
AI 点评 · 多层加密架构实现行级安全,为多租户大模型分析提供企业级隐私保护方案。

In this post, we show you how to build an automated claims processing pipeline using two key Amazon Bedrock capabilities: Amazon Bedrock Data Automation for intelligent document ex…
AI 点评 · 用AI自动化处理医疗理赔,结合Bedrock与HealthLake,显著提升效率并降低人工错误。

In this post, you learn how to debug production agent failures using built-in observability capabilities. We walk through common failure patterns, show how to analyze agent behavio…
AI 点评 · 用内置可观测性调试生产级AI代理,解决落地部署中的常见故障分析难题。
We introduce SWE-Interact, a new testbed for evaluating coding agents on multi-turn, interactive, user-driven software engineering tasks. Existing frontier SWE benchmarks typically provide complete re…
AI 点评 · 开源模型自我进化,开启智能编程新范式,降低AI开发门槛。
Large language models (LLMs) are increasingly used in open-ended multi-agent settings, but the long-run dynamics of model--model interaction remain poorly understood. We study whether open-ended LLM d…
Cursor has launched a new mobile app for remote oversight over coding agents.
AI 点评 · 移动端管理编码代理,打破办公空间限制,提升开发灵活性。
Ornith-1.0: Self-Scaffolding LLMs for Agentic Coding This is an interesting new open weights (MIT licensed) model, the first model release from DeepReinforce. [...] with variants i…
AI 点评 · 首个自搭建代码智能体模型开源,MIT许可降低门槛,或重塑AI编程工具生态。

Enterprise investment in AI is booming. Gartner is calling 2026 an “inflection year” for organizations to align their AI projects with strategic business objectives. As the pressur…
Open-source local policy, recovery, and audit layer for explicit Codex execution.
作者 | 张子怡 编辑 | 袁斯来 近日,AI可穿戴品牌AIVELA宣布完成数百万美元首轮融资。本轮融资由线性资本领投,锋领资本跟投,智能电助力自行车品牌URTOPIA等产业方共同加注。 本轮融资将主要用于下一代AI可穿戴产品研发、健康数据与AI Agent能力建设、全球市场拓展以及核心团队扩张。AIVELA将以智能指环、智能手链等贴身可穿戴产品为起点,面向…
Human Agent in the loop I dislike the phrase “human in the loop” because it cedes authority to the machines. Let’s flip the narrative. It’s our loop, we work the same way we always…
AI 点评 · 观点颠覆:主张人类掌控AI工具,而非被机器主导,重新定义人机协作的主动权。
AI World Generator 2026: Create Self-Evolving Maps & Stories
AI World Generator 2026: Create Self-Evolving Maps & Stories
2026 Multi-Agent AI Town Simulation | Polis Darwin LangGraph
2026 Multi-Agent AI Town Simulation | Polis Darwin LangGraph
HEWN 2.0 2026: AI Output Router for Precision Summaries & Polished Code
HEWN 2.0 2026: AI Output Router for Precision Summaries & Polished Code
Most coding-agent benchmarks are static: an agent receives a complete task description up front and is judged only by its final code. Real coding assistance is interactive, with users clarifying goals…
Agentic multimodal models perform diverse operations on an image via code and reason over the returned view, an effective paradigm for fine-grained visual question answering. However, code operations…
We introduce Agents-A1, a 35B Mixture-of-Experts Agentic Model that reaches trillion-parameter-level performance by scaling the agent horizon. We investigate agent-horizon scaling from two perspective…
Data, as the fundamental substrate of modern intelligence, has greatly driven the development of current foundation models. Naturally, researchers aim to extend this paradigm to the domain of GUI agen…
We introduce SWE-Interact, a new testbed for evaluating coding agents on multi-turn, interactive, user-driven software engineering tasks. Existing frontier SWE benchmarks typically provide complete re…
Current operating systems expose interfaces optimized for human users but not for AI agents. Humans benefit from pixels, icons, windows, visual grouping, mouse movement, and keyboard shortcuts; AI age…
Large Language Model (LLM)-based agents can solve complex procedural tasks by interacting with environments over multiple turns, but this ability typically depends on large models, long contexts, and…
Complete Guide 2026: Claude Code Manual – Workflow Pipelines & Adversarial Budget Loops
Complete Guide 2026: Claude Code Manual – Workflow Pipelines & Adversarial Budget Loops
Ultimate Claude Fable 5 Guide 2026: Use Cases, Integrations & Benchmarks
Ultimate Claude Fable 5 Guide 2026: Use Cases, Integrations & Benchmarks
AI Agent Toolkit 2026: Smart Device Control for iOS & Android
AI Agent Toolkit 2026: Smart Device Control for iOS & Android
Local-first AI learning workspace — ask, note, review and create around your own materials. Wiki KB, Agents, Skills, creation tools.AI 学习工作台,围绕你的资料完成问答、笔记、复习和…
A desktop app to prototype agent ideas, inspect every harness step, replay failures, and evaluate performance, all in one place. Local-first, cloud-ready for ma…
A month ago there was a wave of posts and tweets about engineers walking around cafes and parks with their MacBooks propped half-open, as fully closing the lid forces sleep that st…
AI 点评 · 针对macOS使用痛点,用AI智能管理开盖唤醒,提升开发者外接设备时的工作效率。
A secure, stable, and lightweight alternative to OpenClaw and Hermes.
Existing computer-use benchmarks fail to capture the realism, complexity, and long-horizon demands of real-world computer use, limiting their ability to reveal the limitations of frontier agents. We i…
LLM agents handle user requests on behalf of organizations through tool calls and must follow the company policies stated in their system prompts. Prior work approaches this as a safeguarding problem…
Video understanding is a fundamental capability for multimodal intelligence, and recent Multimodal Large Language Models (MLLMs) have achieved remarkable performance on Video Question Answering (Video…
Large language models (LLMs) are increasingly used to take actions in the real world and support human decision-making, yet most agents rely on parametric knowledge, fixed post-training data, retrieva…
Single-agent multi-modal system for insurance damage claim verification. #1 globally, HackerRank Orchestrate June 2026 (15,295 registrants). Anthropic, OpenAI,…
为 A股投资者打造的全球产业链资讯看板 · 12 大赛道一一对应 A股板块(半导体/AI/机器人/新能源车…),覆盖 100+ 权威源,用你自己的大模型每日提炼中文「今日要点」+ 翻译 · 全程本地、零 API key · Local AI news dashboard tracking the global indu…
人类一次录制,Agent就能模拟
AI 点评 · 将人类操作转化为Agent通用能力,大幅降低自动化门槛。

Using Open-Weight Models in Local Coding Harnesses as an Alternative to Claude Code and Codex Subscriptions
AI 点评 · 本地化开源模型替代付费服务,降低AI编码成本与依赖。

IT之家 6 月 27 日消息,据央视新闻 6 月 25 日报道,市场监管总局正会同相关部门,加快智能体等前沿技术领域标准制定速度,动态完善适配产业发展的人工智能国家标准矩阵。 报道称,目前正在抓紧制定的国家标准,除智能体外,还有 具身智能、世界模型、本体模型 等前沿技术标准,算力基础设施、高质量数据集、仿真测试平台、深度学习编译器、开源模型平台等底座类标准…
AI 点评 · 政策加速标准制定,将推动智能体与具身智能产业规范化发展,抢占技术制高点。
AI 点评 · 用剧本杀类比Agentic AI,降低理解门槛,适合大众快速入门。
LLM agents are expected to act over multiple turns, using search, browsing interfaces, and terminal tools to complete user goals. Yet not every goal is well specified or achievable in the available en…
Incident Report: CVE-2026-LGTM Spectacular hypothetical incident report by Andrew Nesbitt. Day 2, 16:00 UTC --- Two AI review agents from competing vendors, both attached to a down…
AI 点评 · 虚构安全事件揭示AI代理间冲突风险,警示未来协作需防漏洞。
AI 点评 · 微软将AI从辅助工具升级为自主执行体,标志着企业级智能体进入常驻操作新阶段。
We present HORIZON, a self-evolving agent framework that treats hardware design as repository-level code evolution. A Markdown harness is compiled into a project pack containing domain knowledge, an e…
The transition from static chat bots to autonomous agents--equipped with persistent memory, tool-use protocols, and multi-agent collaboration--has fundamentally expanded the AI threat landscape. Curre…
AI 点评 · Agent作为云服务新用户,预示云计算的交互模式将彻底改变。
We built a model router that plugs into coding agents (e.g. Claude Code, Codex, Cursor, etc.) and intelligently sends requests to the best model to serve them. Here's a quick demo…
AI 点评 · 让编程助手自动匹配最优模型,大幅提升代码生成效率与成本控制。
We built a model router that plugs into coding agents (e.g. Claude Code, Codex, Cursor, etc.) and intelligently sends requests to the best model to serve them. Here's a quick demo…
Autonomous coding agents now open and merge pull requests in shared repositories at scale, and the field evaluates them the way it has always evaluated components, one agent at a time, on isolated ben…

In this post, you learn how Stripe built a production-grade AI agent system for financial compliance. We cover the technical architecture of Stripe’s ReAct agent framework and the…
AI and Liability Bruce Schneier and Nathan Sanders on the recent German ruling that Google be held liable for errors introduced in their AI overviews: AI agents are agents of the p…
AI 点评 · AI责任判定首案,德国法院裁定谷歌为AI错误担责,确立行业先例。
Agent-testing startup Patronus AI, founded by former Meta AI researchers, is experiencing nearly insatiable demand, its investor says.
AI 点评 · AI安全测试赛道爆发,前Meta团队获巨额融资,预示行业对AI可靠性评估需求激增。
Multi-agent systems (MAS) built on large language models (LLMs) provide a promising framework for solving complex tasks through role specialization and structured interaction. However, their performan…
Knowledge-based Visual Question Answering (KB-VQA) requires models to combine image understanding with external knowledge. Most prior methods use a fixed retrieve-then-generate pipeline with a pre-sel…
RocketSmith is an agentic system which intelligently automates the DFAM process for the development of high powered rockets suitable for launch. The system utilizes a large language model to orchestra…
As large language models and harness frameworks continue to advance, agents operating in terminals are increasingly capable of performing a broader range of general computer-use tasks beyond coding. H…
Program verifiers play a central role in training coding agents, including selecting trajectories for supervised fine-tuning (SFT) and providing rewards for reinforcement learning (RL). Standard execu…
Benchmarks are widely used to evaluate task completion by Large Language Models (LLMs), but this approach has accumulated construction-validity problems, and a passing score may not show whether the r…
Search agents powered by large language models (LLMs) are increasingly used to solve complex information-seeking tasks, requiring multi-step retrieval and reasoning to fulfill user goals. However, exi…
AI 点评 · 打通AI编程工具与桌面环境,并行Agent工作流将大幅提升开发效率。

Notion is "going all in on using agents to run your inbox."
AI 点评 · 转向AI代理管理邮箱,Notion此举或将颠覆传统邮件处理方式,值得关注。
AI 点评 · AI智能体权限管理成关键,Uber与Auth0的实践揭示企业安全新挑战。

In this technical collaboration between AWS and the authors, we present a pragmatic solution: agentic overlays. Agentic overlays are thin wrapper layers that transform traditional…
AI 点评 · 用智能代理改造传统系统,比推倒重建更经济高效,是企业数字化转型的新思路。
Multimodal web agents can assist humans in operating repetitive GUI tasks, where effective task planning is essential for decomposing complex tasks into executable actions. While small open source MLL…
General Intuition has raised $320 million to scale AI trained on millions of hours of gameplay, betting action data can help AI develop something closer to human intuition.
AI 点评 · 用游戏数据训练AI直觉的独特路径,23亿美元押注行动智能向现实世界迁移。
AI 点评 · Agent狂烧算力模式终结,行业回归理性,ROI成为核心指标。

In this post, we show you how to build Chaplin (Customer Health and Planned Lifecycle Intelligence Nexus), an open source solution that uses AI agents exposed through the Model Con…
AI 点评 · 用AI代理自动分析健康数据,让非技术人员也能自助获取洞察,降低门槛。

This post shows how to build a governed, serverless data mesh on AWS that provides the secure, scalable data foundation production agentic AI requires.

A new system, known as Murakkab, optimizes the design and deployment of multistep workflows that power AI applications.
A new OpenAI research paper shows how AI agents are transforming work, enabling longer, more complex tasks and expanding productivity across roles.
As agentic systems continue to evolve and are widely deployed in real-world scenarios, there is a growing demand to faithfully evaluate their capabilities. However, current benchmarks are typically bu…
Multi-model LLM systems such as routing, voting, cascades, fusion, and mixture-of-agents are used to beat single-model accuracy. We show that their gain is capped by a quantity the field rarely report…
Outcome-based reinforcement learning provides a stable optimization backbone for language agents, but its sparse trajectory-level rewards provide little guidance on which intermediate decisions should…
While text-to-image (T2I) models have achieved remarkable progress, they struggle with real-world requests that are often underspecified, implicit, or dependent on up-to-date knowledge. We identify th…
LLM-based agents for program repair are increasingly built on a "generate-run-revise" paradigm, iteratively executing tests to evaluate and refine patches. This execution-based approach has become sta…
LLM-based code agents navigate repositories through keyword search but miss the structural relationships, such as call graphs, inheritance hierarchies, and configuration dependencies, that define how…
Large language models (LLMs) can make scientific software easier to use. However, a general model does not automatically know which measurements a particular sensor can support, which algorithms are i…
Web-agent benchmarks overwhelmingly measure depth -- pinning one obscure answer behind a chain of constraints -- while breadth, exhaustively enumerating a closed set and filling each item's attributes…
Multi-agent large language model (LLM) systems often rely on verifier and critic agents to suppress hallucinations, but verification is delayed. During this delay, false claims can propagate through t…

In this post, you will learn how to build a voice agent that handles appointment reminder conversations using Amazon Nova 2 Sonic and Amazon Bedrock AgentCore. The agent authentica…
AI 点评 · 亚马逊推出低成本语音预约系统,展示AI在医疗场景的落地新范式。
Process reward models enable fine-grained, step-level evaluation of LLMs, yet building them for agentic settings remains prohibitively difficult: long-horizon interactions, irreversible actions, and s…
AI agents are granted access to tools, APIs, and other infrastructure, making them active principals in those systems. The dominant approach places controls inside the agent's own runtime: system prom…
AI 点评 · 企业Agent落地难在工程化,亚马逊云科技的工程视角直击行业痛点。
As autonomous AI agents increasingly transact across organizational boundaries, a fundamental trust challenge emerges: how can an agent assess whether an unknown counterpart is trustworthy? The ERC-80…

In this post, we demonstrate the architecture and approach Loka used to solve a common frustration: robotic, slow voice assistants that cause customers to hang up, damaging brand r…
AI 点评 · Loka用亚马逊新模型打造自然流畅语音代理,解决机器人语音痛点,技术方案值得借鉴。
Tool use enables large language models (LLMs) to perform complex tasks, and recent agentic reinforcement learning (RL) methods show promise for enhancing model capabilities. However, RL alone often le…
Last-night exam-cram coach as a Claude Agent Skill: turns your slides, notes and past papers into a chaptered knowledge base + quiz bank, teaches only what's in…
Film Language Operating System for director-level, Seedance-ready cinematic video prompts.
文|王欣逸 编辑|张雨忻 2026 年开年来,3D 生成模型赛道相当热闹。 今年第一季度,影眸科技发布首个 3D 编辑模型 Rodin Gen-2 Edit,让 AI 3D 模型第一次可编辑;今年 6 月,VAST 官宣了新一轮融资,Meshy 也紧随其后,宣称自己发布了全球首款 3D AI Agent。 近日,影眸科技——这支扎根学术圈、创业早、年轻的 3…
A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.
Agent Skill for MLX model porting, validation, quantization, benchmarking, and optimization.
Selfhost modern LLM stacks. Run the whole fleet from your terminal
Figma’s design agent, now with custom tools and greater context Figma
Open-source, self-hostable alternative to Claude Tag — a Slack-style workspace where your team and its AI agents (Claude Code, Codex, GitHub Copilot, and more)…
Browser-native side panel for Hermes Agent — connect web context to your local Hermes runtime.
2026 年是 AI agent 的第一年,2026 年也是 FireWire 的最后一年。 查看全文
The all-cash deal gives MoEngage access to technology that assigns AI agents to individual customers.
当地时间6月23日,英伟达宣布推出NVIDIA BioNeMo Agent Toolkit,该工具包包含英伟达超过十年的生命科学库、工具和开放模型,使AI智能体、科学家和实验室能够通过收集证据、跨研究结果进行推理、运行计算实验以及推荐下一步最佳行动来协同工作,从而加速科学发现。(界面)
AI 点评 · 英伟达BioNeMo将十年生命科学积累与AI智能体结合,有望大幅加速药物研发和科学实验。
Tool Calling and Structured Output are two core capabilities of modern Agent systems, yet their interaction under joint deployment conditions remains insufficiently understood. This paper reports a re…
Fine-grained visual reasoning requires multimodal large language models (MLLMs) to identify task-relevant visual evidence and ground their reasoning in local image regions. Existing agentic methods ty…
We introduce Autodata, a general method that enables AI agents to act as data scientists who build high quality training and evaluation data. We show how to train (meta-optimize) such a data scientist…
Process reward models enable fine-grained, step-level evaluation of LLMs, yet building them for agentic settings remains prohibitively difficult: long-horizon interactions, irreversible actions, and s…
A classical intuition holds that verifying a solution is easier than producing one. For today's coding agents, this intuition is being inverted: as foundation models develop stronger reasoning capabil…
Tool use enables large language models (LLMs) to perform complex tasks, and recent agentic reinforcement learning (RL) methods show promise for enhancing model capabilities. However, RL alone often le…
AI 点评 · 鸿蒙Agent时代开启,标志国产操作系统从工具迈向智能生态。
AI 点评 · 聚焦边缘AI应用开发效率,一键全球部署降低技术门槛,推动智能应用快速落地。
In the big-world regime, agents cannot be universally capable and their ability is inevitably specialized across a world model in pieces. Consequently, standard uniform guarantees fail to distinguish…
Provider-pluggable orchestration runtime for multi-model AI inference. ( Sakana Fugu style )
Agentic data analysis systems produce rich outputs, including code, numerical results, and verbal diagnostics. This makes them more challenging to evaluate than single-turn LLM responses. It is theref…
AI 点评 · 隔离执行环境保障AI智能体安全,降低代码运行风险,推动企业级应用部署。
🤖 Building AI Agent Systems from Scratch — A comprehensive, practical tutorial from fundamentals to production-grade multi-agent applications
🤖 Building AI Agent Systems from Scratch — A comprehensive, practical tutorial from fundamentals to production-grade multi-agent applications
LLM agents solve repository-level coding tasks through multi-turn tool use, but utilize half their budget on locating faults before editing. Dedicated localization frameworks have emerged, yet are sti…
AI 点评 · Angular官方技能让AI生成更规范代码,开发者效率与质量双提升。
Hey HN. http://peerd.ai is an AI agent harness that lives entirely in your browser as a web extension. You don’t have to install a separate “AI browser”. You don’t have to bolt on…
Hey HN. http://peerd.ai is an AI agent harness that lives entirely in your browser as a web extension. You don’t have to install a separate “AI browser”. You don’t have to bolt on…
Hey HN. http://peerd.ai is an AI agent harness that lives entirely in your browser as a web extension. You don’t have to install a separate “AI browser”. You don’t have to bolt on…
Agentic AI爆发的拐点已然来临
编程比肩Opus 4.7
Stockholm-based startup Fika Jobs is building a video-first hiring platform that combines AI interview agents with short-form video profiles, creating something that feels like a c…
Shumai is an open source platform for uploading creative files, managing projects, collecting precise feedback, sharing work, and collaborating with AI agents, all in one simple cr…
Neural CAs model self-organizing pattern formation on grids. Now the grid is gone. Each cell is an agentic particle that can move freely in space and change its state. While each p…
有智青年挑战赛暨全国AI+场景应用大赛决赛在WAVES新浪潮大会期间举行,汇聚多支青年团队围绕AI与数字经济前沿场景展开角逐,展现青年创业者的技术探索与落地能力。 2026年,AI全面进入“行动者”时代。当大模型、智能体、具身智能从实验室的技术概念,全面走进千行百业的产业落地场景,AI工具的平民化与普及化,依托开源生态、轻量化开发工具与普惠算力,不断压低创新…
DataClaw: Agentic Tailoring Multimodal Data from Raw Streams — coming soon (code, weights, dataset & DataClaw-val upon acceptance).
DataClaw: Agentic Tailoring Multimodal Data from Raw Streams — coming soon (code, weights, dataset & DataClaw-val upon acceptance).

Telecom operators have seen remarkable returns from using generative AI to automate network management, customer care and back-office operations. Most of that impact has been task‑…
The loop takes agentic AI a step further by authorizing a swarm of agents to work continuously in the background, endlessly.
AI 点评 · 自主代理集群持续后台运行,突破传统AI单次交互模式,预示自动化新纪元。
Large language models are increasingly deployed as agents that reason over documents rather than answer from parametric knowledge. We study archive-grounded reasoning: locating sparse evidence across…
Experience-driven self-evolution is critical for large language model (LLM) agents to improve through open-world interaction. However, existing experience learning methods mostly rely on single-agent…
We introduce NatureBench, a cross-discipline benchmark of 90 tasks distilled from peer-reviewed Nature-family publications, designed to evaluate whether AI coding agents can move beyond reproduction t…
Multimodal misinformation detection is increasingly important because viral posts now combine long multilingual narratives, several images, mixed provenance, and subtle text--image framing errors. Exi…
A world model predicts environment dynamics based on current observations and actions, serving as a core cognitive mechanism for reasoning and planning. In this work, we investigate how world modeling…
Agentic language models dramatically expand the applications of AI yet little is publicly known about how to curate training data for broadly capable agents. Existing open efforts such as SWE-Smith, S…
Memory for large language model (LLM) agents has rapidly evolved from simple retrieval-augmented mechanisms into a data management system that supports persistent information storage, retrieval, updat…
Long-term memory promises LLM agents that grow more capable across sessions, maintaining an accurate, evolving understanding of the user that interaction forms. In practice, however, this memory is ev…
Voice agents face a fundamental tension: the reasoning, retrieval, and tool use that make foundation models capable are iterative and slow, while conversational interaction demands responses on a mill…
Agent skills extend language-model agents with task-specific procedures, scripts, and references, but the tasks and environments they target continually change. Existing methods improve skills in boun…

In this post, you will learn how Ampersend built a pay-per-intelligence routing layer on top of Amazon Bedrock AgentCore Payments. AI agents autonomously route tasks to the most ef…
AI 点评 · 看点是,基于亚马逊Bedrock实现按智能付费的AI代理路由层,创新了成本与任务匹配模式。
Multi-agent systems (MAS) offer a scalable path forward for agentic AI, comprising multiple LLM-based agents, each assigned a system prompt and a position within a workflow that governs inter-agent co…
AI 点评 · 吴恩达点破AI泡沫,聚焦小团队与Agent重构数据架构,揭示行业理性回归趋势。
Oak is a version control system I've been working on designed for agents ( https://oak.space ). It improves the speed and context your agents need when working on serious projects.…

A capability threshold I've been carefully monitoring.
A skill for AI agents: search the web with SearXNG, browse with Camofox, bypass protections with CloakBrowser. Anti-hallucination by design. Self-hosted, free,…

Mission, Vision and Veritas — new Los Alamos National Laboratory (LANL) supercomputers to be built with HPE and NVIDIA — are tapping NVIDIA Vera CPUs to accelerate scientific disco…

The next era of AI will not be defined by compute alone. Its growth will be determined by energy. As accelerated computing scales across AI factories, agentic AI, industrial AI, ed…
The first AI agent harness native to the browser. A browser extension that runs a full agent loop where you already work: it drives your tabs, spins up sandboxe…
DeepSeek正在全力押注
36氪获悉,中信建投研报称,国内模型持续迭代,GLM-5.2、Kimi K2.7 Code强化1M上下文、长程Agent、Agentic Coding和真实工程交付能力,推动国产模型从通用问答转向开发者工具和企业级工作流。Kimi补强国际化运营能力,DeepSeek融资强化头部模型产业化预期,微信AI灰度测试则显示AI入口正从独立App走向超级应用生态,有望…
AI 点评 · 国产模型转向企业级应用,算力需求确定性增强,产业链景气度有望持续。
Temporary Cloudflare Accounts for AI agents The announcement says this is "for AI agents" but (as is pretty common these days) the AI hook isn't really necessary, this is an intere…
AI 点评 · 为AI代理提供临时账户,简化安全访问,降低管理成本,是云服务与自动化结合的新探索。
Hybrid RAG (DuckDB vector + BM25 + RRF + recency/keyword priors + optional cross-encoder rerank) as an installable library + CLI.
Long-horizon LLM agents can fail quietly: they settle on one reading of the evidence early, then spend the rest of the run defending it. We call this premature commitment. Final-answer scoring misses…
Computer-use agents (CUAs) now act on a user's behalf across personal applications such as email, calendars, and to-do lists. This cross-application access is useful, but it also creates a privacy ris…
Terminal-using agents have quickly become the most popular downstream application of language models (LMs). Despite their prevalence, relatively little academic work has examined RL-based training of…
While recent LLM-based terminal agents have demonstrated promising capabilities, the scarcity of high-quality, executable training data remains a critical bottleneck. Existing synthesis pipelines typi…
Phones are becoming an important execution surface for general-purpose agents, but training open models for reliable phone use remains difficult because the environment that matters at deployment, rea…
Long agent traces composed of chains of thought and tool calls accumulate stale content that anchor subsequent generations, and eventually outgrow the context window. Existing scaffolds mitigate it wi…
Recent attempts to combine large language models (LLMs) with causal discovery ask models to infer pairwise directions, propose graph structures, or inject language-model outputs as priors and constrai…
Enterprise agents increasingly operate inside workspaces: they read heterogeneous files, invoke tools, and deliver business artifacts. We introduce EnterpriseClawBench, an enterprise agent benchmark c…
What is an agent? What constitutes agency? With the rise of Large Language Model (LLM) systems marketed as ``coding agents'', ``AI co-scientists'', and other ``agentic" tools that promise to drive up…
AI agents are driving a new software paradigm, with the ability to autonomously call tools, extract information, manage memory, and complete tasks that span applications and data sources. Most existin…
Long-horizon agents depend on context management: systems compress, summarize, and evict old tokens so tasks can continue beyond finite windows. That is safe only when dropped information is no longer…
The Hitchhiker's Guide to Agentic AI is a comprehensive practitioner's reference for building autonomous AI systems. The book covers the full stack from first principles to production deployment, orga…
Computer-use agents can execute software tasks through either graphical interfaces or programmatic command interfaces, but existing evaluations confound interaction modality with differences in tasks,…
Procedural memory is increasingly used to improve LLM agents on recurring workplace tasks, yet its ability to produce reusable skills remains poorly understood. We introduce AFTER, a benchmark of 382…
fak — the Fused Agent Kernel: one Go binary that turns a tool-using agent (Claude Code, Codex, Cursor, any OpenAI/Anthropic/MCP client) into a managed agent: ca…
AI 点评 · AI原生内容风控成为新趋势,大模型与Agent结合将颠覆传统审核模式。
Transform any idea into a running AI-powered company
AI 点评 · 专注构建可靠自主AI系统,填补当前AI可靠性短板,推动实用化进程。
Agent 和人一样离不开闭环。 查看全文
AI 点评 · 以Vibe Coding实现全流程AI开发,验证了Agent闭环工作的可行性,门槛极低值得关注。
Generative music systems can now produce impressive audio from text prompts, but audio outputs are difficult to inspect, edit, and diagnose as musical structure. We introduce Libretto, an agent-facing…
LLM agents increasingly operate in large tool ecosystems, where real-world tasks require discovering relevant tools, inferring implicit sub-goals, and adapting to dynamic environments over long horizo…
Artificial intelligence systems are commonly evaluated through task performance and behavioral imitation, but such evaluations leave open whether an artificial agent can acquire, stabilize, and use ne…
Local coding-agent orchestrator — DAG of auto-approved, git-worktree-isolated sub-sessions across LLM providers (Claude/Kimi/Grok/DeepSeek/local). AGPL-3.0.
AI 点评 · 云服务商为AI代理开辟专用通道,预示智能体自动化操作进入新阶段。
The real valuable capability MCP offers over skills/CLI is isolating the auth flow outside of the agent’s context window, and potentially out of the harness completely. [...] Maybe…
AI 点评 · MCP将认证流程独立于智能体上下文,这种架构创新对提升安全性与扩展性意义重大。
Stop wasting tokens and re-explaining your project every session. Recall gives Claude Code durable memory — entirely offline.
Stop wasting tokens and re-explaining your project every session. Recall gives Claude Code durable memory — entirely offline.
A working citation looks like proof -- but the fact that a link resolves does not mean the cited paper supports the claim. I find that current agentic models rarely fabricate citations (over 99% resol…
A comprehensive, production-focused guide to acing AI/ML system design interviews at top tech companies.
UmaDev: A coding agent that works like a real dev team, commanding the Claude Code / Codex / OpenCode you already use.
UmaDev: A coding agent that works like a real dev team, commanding the Claude Code / Codex / OpenCode you already use.

IT之家 6 月 19 日消息,谷歌目前正全力推进 AI 生态的建设,并在搜索引擎中强行加入 AI 智能体。 DuckDuckGo 官方今日在 X 上晒出了一张截图,显示谷歌 AI 概览正引导那些讨厌 AI 的用户前往 DuckDuckGo 的“No AI Search”页面,还提到了可调低 AI 体验强度的浏览器设置。 PiunikaWeb 测试发现,当用…
AI 点评 · 谷歌强推AI反遭打脸,AI自己建议用户用竞品,暴露了产品逻辑矛盾。

IT之家 6 月 19 日消息,印度首富、信实工业集团董事长穆克什 · 安巴尼希望,把公司打造为当地 AI 产业的代表,并把 AI 服务带入 电话、移动应用和智能家居 。 当地时间 19 日(今天),信实工业举行了年度股东大会,并发布 AI 通话助手 Jio Call Agent。Jio Call Agent 可以加入电话通话, 自动转录对话、生成摘要 ,还…
AI 点评 · 印度首富亲自押注,AI本土化野心显露,或重塑全球科技版图。

This post shows how to enable Adobe Marketing Agent for Amazon Quick using a Model Context Protocol (MCP). We walk you through how to configure the integration, authenticate using…
AI 点评 · MCP技术落地营销场景,打通Adobe与Quick平台,实现工作流智能化。
Reverse engineered Windows Copilot into an OpenAI-compatible API. Access GPT-4 and GPT-5 models through a simple REST interface without API keys or billing.
Reverse engineered Windows Copilot into an OpenAI-compatible API. Access GPT-4 and GPT-5 models through a simple REST interface without API keys or billing.
Honey (I Shrunk the AI) by GreenPT: a cross-tool coding skill that cuts AI coding-agent token usage and LLM API costs — write less code, less prose, and denser…
Honey (I Shrunk the AI) by GreenPT: a cross-tool coding skill that cuts AI coding-agent token usage and LLM API costs — write less code, less prose, and denser…
Evidence-grounded evaluation for AI agents — verifies each claim against the agent's real tool outputs (constrained, evidence-grounded model judgment, not holis…
Evidence-grounded evaluation for AI agents — verifies each claim against the agent's real tool outputs (constrained, evidence-grounded model judgment, not holis…
The Juggler Code Agent
🎯 从零基础到 AI Agent 全栈工程师 · 110 个详细教程 · 58 万字 · 400+ GitHub 项目精选 · Obsidian 友好 · 中文
As agentic systems tackle increasingly complex multi-step tasks, evaluating their trajectories presents a major bottleneck - human annotation of a single trajectory on popular agentic benchmarks can t…
Massive unstructured multimodal streams suffer from high "data entropy," impeding both efficient human knowledge acquisition and high-quality AI post-training. Existing passive annotation paradigms, h…
AI agents acting on behalf of users are constantly making decisions, and for users to trust their agents, those decisions must align with what they actually want. Privacy is an important alignment pro…
Biomedical researchers increasingly use AI-generated analyses and reports to interpret protein-level signals, but static outputs are often insufficient for research decision-making, where users need t…
AI 点评 · 评估AI研究助手能否保守机密,关乎隐私安全核心问题。
AI 点评 · AI落地效率失衡揭示组织适配成关键瓶颈,Agent应用挑战值得深思。
AI 点评 · 数据库短板暴露,自主智能体发展卡在基础设施瓶颈上。
Autonomous agents are increasingly connected to cloud, deployment, and data-control workflows, but production mutation authority should not reside inside non-deterministic reasoning processes. Existin…

Today, Amazon Bedrock AgentCore harness is generally available. Two API calls (CreateHarness to define an agent, and InvokeHarness to run it), and you have an agent running in seco…
AI 点评 · 亚马逊推出极速AI代理开发工具,两分钟即可从创意到生产级应用,大幅降低开发门槛。
LLM-based coding agents need higher-level operational knowledge about a repository (which files house which subsystems, how to run the test suite, which workflows have historically led to wrong fixes)…
Securing AI agents that operate in complex digital environments has become a critical need, and runtime monitoring approaches that formulate and enforce policies expressed in a formal language like Da…
AI 点评 · Google正试图统一AI Agent部署标准,或将重塑行业生态。
Hey HN - we’re Oskar, Szymon, and Piotr, and we’re building TesterArmy ( https://tester.army ). TesterArmy is an agentic testing platform that runs end-to-end checks before deploym…
AI 点评 · 用AI代理自动执行端到端测试,大幅提升应用发布前的质量保障效率。
自托管、零运维的 A 股「选股 + 监控 + 回测」量化工作台 | 基于 TickFlow 数据源 | LLM能力驱使策略定制+个股分析+复盘 | 自由接入第三方数据源与个性化扩展数据 | 个人开源 ,非TickFlow官方项目
自托管、零运维的 A 股「选股 + 监控 + 回测」量化工作台 | 基于 TickFlow 数据源 | LLM能力驱使策略定制+个股分析+复盘 | 自由接入第三方数据源与个性化扩展数据 | 个人开源 ,非TickFlow官方项目
文 | 孙小雯 访谈 / 编辑 | 海若镜 「暗涌Waves」独家获悉,侵入式脑机接口公司「芯生视界」近日完成近亿元人民币种子轮融资。本轮融资由经纬创投领投,星连资本、燕缘创投、水木创投跟投。 当下,侵入式脑机接口已经在治疗瘫痪、脑控外设等医疗场景落地,验证长期植入的安全、有效。与此同时,AI Agent和具身智能技术加速进化,也放大了市场对脑机接口的期待:…

Today, Quick gets even more powerful: new autonomous agents that work continuously on your behalf, an activity feed that helps you prioritize your most important work, and the abil…
AI 点评 · 亚马逊Quick新增自主代理,能持续替用户工作,大幅提升效率,值得关注。
Policy-adherent tool-calling agents in customer-service domains must maintain task states across turns while calling tools and obeying domain policies. Task states consist of relevant facts, identifie…
Real-world spatial intelligence requires reasoning over a continuous and evolving 3D world, yet existing VLMs and tool-augmented agents largely remain tied to static, stateless inference from isolated…
Achieving dexterous robotic manipulation in the real world heavily relies on human supervision and algorithm engineering, which becomes a central bottleneck in the pursuit of general physical intellig…
Agent benchmarks are growing fast, but no single benchmark touches more than four or five of the dimensions that deployment exposes. This paper aggregates the largest coordinated deep-dive of one MCP-…
This work presents a general framework for training large language models (LLMs) to "Connect the Dots" (CoD), a meta-capability required by long-lifecycle agents: as an LLM-based AI agent gets deploye…
MLLM-based mobile GUI agents have made substantial progress on short-horizon tasks, yet remain unreliable on long-horizon tasks that require retaining intermediate facts across many steps and app tran…
MLLM-based mobile GUI agents have made substantial progress in UI understanding and action execution, but adapting them to real target apps remains costly because mobile apps are numerous, frequently…
As LLM agents increasingly select tools autonomously, their choices among tools with different privileges become safety-relevant. However, prior tool-selection studies focus on safety-agnostic metadat…
Agentic navigation systems require a base navigation model whose observation strategy can be externally reconfigured at inference time, because instruction following, object search, target tracking, a…

Nvidia's self-improvement program for robots enlists teams of AI coding agents.
AI 点评 · 英伟达用AI编程智能体教机器人装显卡,展现了自进化系统的潜力,是机器人自主学习的突破。
Tokenmaxxing was the hottest trend in Silicon Valley earlier this year, with CEOs encouraging employees to push AI usage as far as it would go. Then the bill came due. Uber reporte…
Production data integration is bottlenecked by repeated, lossy handoffs between data owners, engineers, and analysts who must collaboratively discover, structure, and query enterprise data. We present…
Large language model (LLM)-based multi-agent systems (MAS) have demonstrated great potential in solving tasks with execution complexity, by distributing subtasks across cooperative agents. However, th…

Agents are only as intelligent as the context they can reason over. Today, that context is scattered across data lakes, data warehouses, lakehouses, databases, and streams, and in…
Hey HN! I'm Zach from Adam ( https://adam.new/ ). We're building AI agents for mechanical CAD software. We’ve built the company on two fundamental beliefs: - AI will be the primary…
Hey HN! I'm Zach from Adam ( https://adam.new/ ). We're building AI agents for mechanical CAD software. We’ve built the company on two fundamental beliefs: - AI will be the primary…

Today we're introducing new capabilities on Amazon Bedrock AgentCore, the platform to build, connect, and optimize agents. In this post, we cover how these capabilities close each…
100+ Curated, CI-verified AI workflow recipes. Powered by FlowStacks.
100+ Curated, CI-verified AI workflow recipes. Powered by FlowStacks.
Open-source reverse engineering lab: 197-article knowledge base + MCP tools + CTF/APK/PE automation toolchain. Agent-native. Note:由于场景原因,目前有让几乎所有(除fable5)AI都会越…
A prompt-aware LLM router that predicts which models can complete each request, then selects the cheapest capable one: 53.2% lower cost and +1.9 pts completion…
shadcn/ui, but for building agents. 🤖
用户可以在与智能体的对话中提出消费需求

Today, we’re announcing a new API with Amazon Bedrock Guardrails. With this API, you can apply individual safeguards, also referred to as safety checks, at any point in your agenti…
AI 点评 · 亚马逊推出新API,可在AI应用任意环节插入安全检测,提升防护灵活性。

NVIDIA XR AI is now available in public beta, giving developers a framework for building multimodal AI agents for AR glasses and XR devices.
AI 点评 · 英伟达XR AI公测,为AR眼镜开发多模态AI代理,推动无手交互新范式。

Move originally planned for Monday would have heavily increased power users' costs.
AI 点评 · 暂停代币计费,保护核心用户利益,展现平台对开发者生态的审慎态度。
Motion forecasting is central to visual intelligence: agents must anticipate how objects will move in order to plan actions, reason about physical interactions, and synthesize realistic futures. We ar…
Multi-turn tool-use RL is bottlenecked by the rapid depletion of informative samples in static datasets. We observe that the gradient signal in GRPO concentrates on tasks with the highest rollout rewa…
Reinforcement learning (RL) has become a representative post-training paradigm for LLMs, enabling strong reasoning and agentic capabilities. However, rollout generation remains a dominant latency bott…
Learning to simulate human users in interactive settings could advance the training of agent assistants, evaluation of personalization systems, research in the social sciences, and more. Existing appr…
Patient contexts span hundreds of heterogeneous documents and thousands of structured data points, yet the document-level metadata that AI systems need for retrieval and triage is absent or incomplete…
Current agentic robot systems can write executable Code-as-Policy programs, observe feedback, and revise behavior across multiple attempts, but they remain largely task-driven: reusable skills are acq…
Memory benchmarks for LLM agents largely assume single-user settings, leaving shared assistants for hospitals, workplaces, campuses, and households understudied. In these deployments, multiple princip…
To assist humans over extended periods in real homes, embodied agents must remember user routines, world states, and past interactions. Existing long-term memory benchmarks mainly evaluate language-ce…
Modern agent systems often suffer from fragmented runtime state: transcripts, tool effects, memory events, workspace placement, branch provenance, and replay evidence are recorded separately and becom…
Long-context reasoning is an essential capability for large language models, particularly when they are deployed as autonomous agents that must reason over lengthy trajectories. Reinforcement learning…
Reproducing research results from papers and released code is central to scientific progress. Existing works have introduced benchmarks to evaluate whether LLM agents can assist with reproducibility,…
Zero-Shot Object-Goal Navigation (ZS-OGN) requires embodied agents to explore and locate target objects without any prior training. To this end, recent methods leverage foundation models. But they typ…
With sophisticated cyber-attacks becoming increasingly prevalent, modern networks require intelligent autonomous cyber-defense agents trained via Reinforcement Learning (RL). These agents employ neuro…
The LLM-empowered personal health agents with user health (sensor) metrics have offered a promising pathway to alleviate global disparities in healthcare access. However, large-scale clinical deployme…

Enterprises are moving agentic AI from proof of concept to production — and the next generation of AI factories are built for the era of agents. At HPE Discover Las Vegas, running…
Securing internal systems with an AI Control Roadmap, combining traditional safeguards and real-time monitoring.
Open-source, bilingual AI video-prompt skill — rewrite ideas into model-ready text-to-video prompts. A portable Agent Skill (Claude Code, Codex, OpenClaw, Herme…
Respond.io, one of Malaysia's startups to watch, uses AI agents to handle high volumes of customer inquiries and charges per convo, not per seat.
Audit any agent decision across its past, present, and future, on one typed graph.
Audit any agent decision across its past, present, and future, on one typed graph.
专为 Agent 设计的 AI 搜索层服务

IT之家 6 月 16 日消息,苹果公司昨日(6 月 15 日)更新支持文档,解释称在 macOS 26.4 系统中, 若用户不常用终端(Terminal),且命令来自网站、聊天智能体、消息或邮件应用,系统可能阻止粘贴。 科技媒体 9to5Mac 指出,用户此前终端里粘贴命令后,系统会先给出安全警告,提示内容可能含有恶意代码。很多人只知道有这个拦截机制,但不…
AI 点评 · 安全机制升级针对非高频操作,防范恶意代码粘贴执行,体现系统防护精细化。

IT之家 6 月 16 日消息,彭博社记者马克 · 古尔曼认为,苹果最终可能推出一款产品,直接对标 OpenClaw—— 这是一套智能体 AI 系统,能够代用户自主操作各类软件。 古尔曼在其专栏《Power On》中撰文表示,他预计苹果会研发一套系统,可全权代表用户操作 iPhone、iPad 与 Mac 端的各类软件。这一预测的依据,是苹果 Siri 工程…
AI 点评 · Siri从语音助手进化为自主操作系统的AI代理,标志苹果在智能体赛道的关键布局。
Memory has become a standard substrate for self-evolving agents, yet retaining experience is not the same as learning how to evolve through it. Existing memory agents can store trajectories, retrieve…
Game generation is an emerging application of coding agents, requiring models to transform natural-language specifications into playable interactive systems. Unlike traditional coding tasks, game gene…
Reinforcement learning pipelines for Large Language Model (LLM) training often rely on manually redesigned environments between stages, requiring practitioners to heuristically infer which configurati…
Language model agents are becoming proficient executors at isolated, short-horizon tasks such as software engineering and customer service. Yet real-world challenges require a combination of sophistic…
Language models trained on large-scale vision-language data have demonstrated strong potential for embodied agents. Harnessing models through embodied tools use offers a promising alternative to end-t…
AI systems deployed in legal workflows hallucinate at rates that aggregate metrics report at ~52%, but this average conceals where errors concentrate and in which direction they run, leaving complianc…
AI 点评 · 智能体正重塑金融效率,这场技术前沿分享值得从业者深挖。

In this post, we walk you through calling the detector functions to diagnose real agent failures. You learn how to interpret their structured output: categorized failures with conf…
Large language models (LLMs) often fail when answering requires identifying a small but decisive piece of evidence within a long or complex context, such as a single line in a tool trace or a subtle d…
Meta-analysis is a demanding form of evidence synthesis that combines literature retrieval, PI/ECO-guided study selection, and statistical aggregation. Its structured, verifiable workflow makes it an…
Deep research agents synthesize long-form reports by searching and reasoning over retrieved evidence. Reinforcement learning with rubric-based rewards improves these agents by optimizing them against…
Release: datasette-agent 0.3a0 New tool, execute_write_sql , which requests user approval and then writes to a database - taking user permissions into account. #27 I added a mechan…
Salesforce says it wants to use Fin's team and technology to improve Agentforce, its existing enterprise platform that businesses can use to build custom AI agents that automate ta…

In this post, you'll build a competitive research agent that demonstrates this pattern end to end. This walkthrough targets developers building multi-step AI workflows who need iso…
NewCore argues the next challenge in enterprise security will be managing AI agents, not people.
Skill to generate the knowledge vault for projects using the Ralph loop
The fastest way to put Volcengine Ark in your terminal and your AI agent — go from prompt to generated media, multimodal answer, or deployed endpoint in a sin…
Agentic新基建
Compile an AI agent's repeated workflows into deterministic, auditable routines that replay for free, with a fallback to the agent.
Vision language models are serving as general-purpose interfaces for complex multimodal tasks. However, deployment still faces three gaps: VLMs typically incur high latency and cost when processing de…
As LLM agents are deployed in long-horizon sessions, context accumulation drives up inference costs. Existing approaches utilize text pruning or dynamic memory eviction to minimize token footprints; h…
Multi-agent LLM systems share state through memory stores, vector indices, and tool registries. We model such sharing as long-running read-generate-write operations under deterministic-generation sema…
Training computer-use agents (CUAs) -- models that interact with graphical desktops through screenshots and keyboard/mouse actions -- requires large-scale, diverse trajectory data collected in full de…
Effective personalized AI-assisted learning demands systems that can not only generate accurate learner-specific educational materials, but also dynamically adapt their instruction to diverse learners…
Current benchmarks for computer-use agents evaluate models in impersonal environments. This leaves a gap between evaluation and deployment where personal assistants are expected to work across a user'…
Large language models (LLMs) often fail when answering requires identifying a small but decisive piece of evidence within a long or complex context, such as a single line in a tool trace or a subtle d…
Personalized presentation generation requires more than conditioning on a current prompt or template: agents must preserve stable user preferences across tasks, retain newly introduced preferences and…
As LLM agents become capable of increasingly long-horizon tasks, evaluating their performance in economic systems is becoming increasingly important. Unlike existing benchmarks that primarily evaluate…
Turn Claude Code into its own Meta-Harness — a skill that evolves the scaffolding around a fixed model (memory, retrieval, context, prompts) via a native propos…

IT之家 6 月 14 日消息,日常上网搜索时,往往需要停下手中的事、打开标签页,主动去查找最新信息。如今,谷歌打算彻底改变这种使用模式。继在 2026 年谷歌开发者大会上首次预告后,谷歌现已正式在 AI 模式中推出搜索智能体功能。 此次升级将传统搜索引擎转变为可在后台静默运行的主动式助手。 首批上线的是信息智能体功能,它会主动全网监测信息,无需用户手动检索…
AI 点评 · 从被动检索到主动监测,AI搜索正从工具进化为管家,颠覆传统上网模式。
Official Implementation of VisualClaw: A Real-Time, Personalized Agent for the Physical World
IT之家 6 月 13 日消息,去年 10 月,毕马威曾发布《总体体验:在智能体 AI 时代重新定义卓越》报告,讨论企业如何利用 AI 满足客户需求。然而据英国《金融时报》12 日报道,这份报告后来被“抓包”充斥 AI 幻觉:报告列举的多个智能体 AI 案例, 要么并不存在,要么并不具备毕马威所描述的能力 。 AI 内容检测工具开发商 GPTZero 的调查…
AI 点评 · 专业机构因依赖AI工具反被AI误导,暴露了当前生成式模型在事实核查上的致命短板。

IT之家 6 月 13 日消息,智谱今日发布了 AI 编程工具 ZCode 3.0 新版本, 深度适配 GLM-5.2 。 官方表示,ZCode 3.0 全面 切换自研 ZCode Agent 内核 。针对满血 GLM 深度优化长程推理、工具调用和大型工程执行链路,整体任务完成效果已显著优于第三方 Agent; 后续版本将聚焦自研 Agent 体验,不再内置…
AI 点评 · 自研内核替代第三方方案,长程推理与工程执行能力显著提升。
AI 八字 + 紫微斗数排盘与综合印证 Skill:算法精准排盘(不靠 LLM 猜),三种分析模式,一键生成水墨风 HTML 命盘海报。兼容 Claude / Codex / Cursor / Workbuddy 等 SKILL.md Agent。
I built Paca out of pure passion—a free and lightweight Jira alternative written in Go where humans and AI agents work together as equal teammates to plan sprints and assign tasks…
AI 点评 · 轻量级AI协作工具挑战Jira,用Go开发的开源项目实现人机平等分工,值得关注。
一起构建下一代物理世界的智能系统
A production-grade OSINT platform that provides situational awareness across multiple intelligence domains.

AgentPerf from Artificial Analysis, the industry’s first agentic AI benchmark, gives developers, enterprises and infrastructure providers a clear way to compare systems for agentic…
AI 点评 · NVIDIA在首个智能体AI基准测试中夺冠,为行业选择基础设施提供关键参考。

In this post, we explore how Rocket Close built a solution using Strands Agents, large language models (LLMs), Amazon Bedrock, Amazon Bedrock Knowledge Bases, and Model Context Pro…
AI 点评 · AI代理优化标题操作的实践案例,展示多工具协同的落地应用。
Efficient and scalable agentic intelligence requires models that can deliver both low-latency responses and strong reasoning capabilities while remaining practical to train, serve, and deploy. In this…
Advanced agents are increasingly demonstrating the potential to operate as autonomous engineers, creating a growing demand for evaluation benchmarks that capture the complexity of real-world developme…
Multimodal large language models (MLLMs) have demonstrated impressive capabilities in many visual tasks, but they often struggle with factual grounding when confronted with complex, open-world scenari…
Deep research agents are increasingly evaluated on their ability to search for evidence, reason over retrieved sources, and produce grounded answers. Existing browsing benchmarks, however, largely ass…
Cooperative multi-objective multi-agent reinforcement learning (MOMARL) models team decision making under multiple, potentially conflicting objectives. In this setting, conflicts arise not only across…
LLM agents are increasingly built not as single model calls, but as scaffolded systems that combine reasoning, memory, reflection, action execution, and learning. While such scaffolds often improve pe…
Large language models increasingly serve as execution engines for agentic systems, yet they still consume context through a sequential text interface. This creates a mismatch with modern structured ag…
We’re Connor and Ambar from BitBoard ( https://bitboard.work ). BitBoard is an agentic analytics workspace. We give you the infrastructure and visualization layer to analyze data w…

IT之家 6 月 12 日消息,在今日的华为开发者大会 HDC 2026 上,华为常务董事、产品投资评审委员会主任、终端 BG 董事长余承东发布了新一代鸿蒙 HarmonyOS 7 操作系统,围绕互联、智能、安全、流畅、空间感五个维度进行升级,同时迈向智能体时代。 余承东在 HDC 2026 现场宣布, 鸿蒙 HarmonyOS 已成为中国第二大智能手机操作…
AI 点评 · 标志着国产操作系统突破安卓和iOS垄断,生态建设进入新里程碑。
AI 点评 · 从数据平台转向智能体企业,Snowflake正定义AI落地的下一个关键战场。

This post shows how to build a custom meeting prep and follow-up assistant using Amazon Quick and Cisco Webex MCP servers. From a single prompt, the agent finds an upcoming Webex m…
OpenAI introduces three Academy courses that help people build practical AI skills, create repeatable workflows, and apply agents in everyday work.
最难档通通零蛋
Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.

IT之家 6 月 12 日消息,OpenAI 昨天宣布收购初创公司 Ona,为编程助手 Codex 提供安全、预配置云环境。 IT之家从官方新闻稿获悉,Ona 的技术将帮助 Codex 执行持续时间更长的任务,并帮助用户将 AI 智能体部署到生产环境。 同时, Ona 的全新技术将帮助企业更好掌控基础设施 、 数据资产和安全边界 ,让 Codex 能够在安全…
AI 点评 · 收购Ona补齐Codex短板,从代码生成迈向安全部署,标志AI编程助手进入企业级实战阶段。
Online group chats are social spaces with local conversational norms that are rarely stated explicitly. The ability and willingness of LLM-based agents to recognize and adapt to these norms remains mo…
Coding agents powered by large language models have demonstrated strong performance on software engineering tasks. Yet most agents consume repositories almost entirely as text, which differs from how…
AI agent performance depends critically on the runtime harness, comprising the prompts, tools, memory, and control flow that mediate how a model observes, reasons, and acts. Yet today's harnesses rema…
Phone agents are increasingly expected to complete real mobile workflows rather than merely predict the next screen action. However, much of the current mobile-agent literature still evaluates agents…
We introduce Nemotron 3 Ultra, a 550 billion total and 55 billion active parameter Mixture-of-Experts Hybrid Mamba-Attention language model. We pre-trained Nemotron 3 Ultra on 20 trillion text tokens,…
Large Language Model (LLM) coding agents have achieved strong results on software engineering tasks, yet repository exploration remains a major bottleneck: locating relevant code consumes substantial…
Agentic search over large corpora relies on retriever-mediated interfaces (e.g., BM25 or ColBERT) for scalable candidate discovery. While effective at ranking relevant documents, these interfaces expo…
Current LLM-based research agents have advanced through agent orchestration, yet largely overlook scientific knowledge orchestration. Existing works often reduce papers to abstracts, surface mentions,…
Tool-augmented LLM agents commonly rely on step-wise atomic tool calls, where each invocation, observation, and value transfer is exposed in the main reasoning trace. This creates an \emph{execution-g…
Recursive language models (RLMs) showed that recursion over model calls is an effective strategy for long-context reasoning, and production coding agents have begun to write code that spawns subagents…
AI 点评 · 微软补齐智能体生产部署关键一环,企业级AI落地更稳。
AI 点评 · Agent技术突破传统风控效率瓶颈,实现产运研全链路智能化协同。
AI 点评 · 建立反馈闭环与评测基准,是让AI编码智能体持续进化的核心驱动力。

Agent-EvalKit is an open-source toolkit (Apache 2.0) that makes this evaluation infrastructure available by integrating with AI coding assistants, including Claude Code, Kiro CLI,…
AI 点评 · 开源工具填补AI智能体评估基础设施空白,让开发者能系统化测试性能。
Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting…
Hi HN, I'm Djoumé. I've been a developer for over 20 years, and like a lot of you I've been coding almost exclusively through an agent in the past few months. It's been amazing to…
Google DeepMind is funding research into the potential dangers of situations where millions of different AI agents interact with each other online. According to Rohin Shah, who dir…
OmniAgent (ICML 2026): the first native omni-modal agent for active video perception — a 7B agent that beats Qwen2.5-VL-72B with 73% fewer frames on LVBench.
前期做了40万AI考生压测
Meshy发布全球首个3D AI Agent
OpenAI plans to acquire Ona to expand Codex with secure, persistent cloud environments, enabling long-running AI agents across enterprise workflows.
Release: datasette-agent 0.2a0 Highlights from the release notes: Tools can now ask the user questions mid-execution. Tools that declare a context parameter receive a ToolContext o…
Multi-agent systems communicate mostly through text, paying a lossy and expensive decode and re-encode cost. KV-cache communication is a promising alternative, yet most prior work is homogeneous, usin…
Interactive LLM agents are becoming part of daily work, but they do not reliably become easier to work with over time: a correction remembered in one session may still be violated in the next. We stud…
Multimodal Large Language Models (MLLMs) have shown promising reasoning capabilities in general domains, yet their performance remains limited in specialized settings such as healthcare, especially in…
Ultra-long-context capability is becoming indispensable for frontier LLMs: agentic workflows, repository-scale code reasoning, and persistent memory all require the model to jointly attend over hundre…
Search Agents -- large language models augmented with search tools -- have intensified the need for future-proof evaluation benchmarks. Existing benchmarks such as BrowseComp rely on static knowledge,…
Large language model (LLM) agents have achieved strong performance on a wide range of benchmarks, yet most evaluations assume static environments. In contrast, real-world deployment is inherently dyna…
LLM-based agents have shown increasing potential in automating scientific discovery. Given an optimizable metric and an execution environment, they can propose, validate, and iterate scientific soluti…
Recent image generators have demonstrated impressive photorealism and instruction-following capabilities in single-image generation and editing. However, constrained by their architectures, they canno…
Spatial reasoning, the ability to determine where objects are, how they relate, and how they move in 3D, remains a fundamental challenge for vision-language models (VLMs). Tool-augmented agents attemp…
Large language models are increasingly deployed as agents for long-horizon tasks, yet their performance is shaped not only by model capability and environment design, but also by the harness that medi…
Disaggregated inference architectures physically separate prefill and decode phases onto distinct GPU pools, creating competing "agents" that share a fixed hardware budget. We provide, to our knowledg…
Mental disorders are highly prevalent worldwide, but the shortage of psychiatrists and the inherent subjectivity of interview-based diagnosis create substantial barriers to timely and consistent menta…
AI 点评 · 为AI代码助手构建安全运行环境,打通企业级应用关键安全瓶颈。
Modern conversational agents condition on an ever-growing dialogue history at each turn, incurring redundant attention and encoding costs that grow with conversation length. Naive truncation or summar…
Vision-Language Models (VLMs) are increasingly deployed as high-level planners for embodied agents, with an emerging strategy of scaling test-time compute to improve capability. However, we observe th…
Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabilities of large language model agents. However, most existing methods assign credit ove…
Human-in-the-loop reinforcement learning (HiL-RL) has emerged as an effective paradigm for real-world robotic manipulation, enabling online policy improvement with human guidance. However, current HiL…
Six Claude Code skills that harden Opus 4.8 toward frontier behavior — written by Fable 5, pressure-tested on the target model with transcripts included.

Today, we’re announcing the Neuron Agentic Development capabilities: a collection of AI agents and skills that make this possible for developers building on AWS Trainium and AWS In…
AI coding agent startup Niteshift has raised a $7 million seed round from a who's who of angels. It's betting companies will want power over, not lock-in with model makers.
Recent advances in agentic Reinforcement Learning (RL) have substantially improved the multi-turn tool-use capabilities of large language model agents. However, most existing methods assign credit ove…
The funding round was led by Norwest, with participation from S Capital VC, Cerca Partners, and Oceans Ventures. Snowflake Ventures also participated as a strategic investor.
Google DeepMind and partners announce a $10M funding call for multi-agent safety research.
An AI-agent skill that generates browser-editable presentations from multiple visual themes, exportable to HTML, PDF, and PPTX.
诚邀30家OPC入驻内测
Operational doctrine for practical AI systems design.
Rubrik公司周二在Rubrik FORWARD大会上宣布,六家全球系统集成商将为企业客户部署其专为Anthropic Claude Code打造的Rubrik Agent Cloud平台。加入“沙漏计划”的合作伙伴包括高知特、德勤、LTM、HCL科技、NTT Data以及威普罗,这些集成商将把该平台整合至各自的网络安全与数字化转型服务体系中。(新浪财经)
AI 点评 · 六大集成商联合推广,标志AI安全平台从技术验证进入规模化落地阶段。
毕马威与微软6月9日宣布扩展全球战略合作关系,聚焦企业级AI智能体的规模化部署。根据协议,微软365 Copilot将向毕马威全球逾27.6万名专业人员全面推广;与此同时,毕马威将采用微软Agent 365平台,对其全球组织内及客户端的AI智能体实施统一管理、监控与安全治理。(界面)
AI 点评 · 毕马威27万员工全面接入Copilot,并首创Agent治理框架,预示企业级AI从工具应用迈入系统化
Hi HN, I've been building Nucleus, a lightweight Linux container runtime focused on two workloads: ephemeral AI-agent sandboxes and declarative NixOS services. It's a single Rust b…
Self-evolving cognitive AI exoskeleton. 10+ frontier models, 245 consensus methods, governed autonomous agents. Automotive, medical, legal, accessibility. 9.3M…
TIL: Setting a custom price for a model in AgentsView I've been really enjoying AgentsView by Wes McKinney as a tool for exploring my token usage across different coding agents run…
AI 点评 · 自定义AI模型定价功能,让用户按需付费,提升灵活性与成本控制。
General-purpose agents such as OpenClaw are increasingly used as autonomous tool users, but their coding ability is difficult to measure under SWE-bench: a generic agent does not by itself satisfy the…
Environments serve as interactive systems for large language model (LLM) based agents across diverse scenarios and play a crucial role in driving the continual evolution of model capabilities. Despite…
Recent progress in foundation models has shifted toward agentic behavior involving multi-step reasoning and tool use. However, open-source efforts largely focus on text-dominant settings, leaving long…
Deep search requires agents to answer complex questions through multi-step web search, browsing, evidence comparison, and synthesis. A central challenge is deciding how to search when several directio…
Compact language models (LMs) reduce cost, latency, and deployment risk for tool agents. Yet MCP-style tool use requires more than isolated function calling: an agent must discover tools from live cat…
Training deep search agents requires verifiable questions whose answers remain unavailable until sufficient evidence has been acquired through search. Existing synthesis methods often increase apparen…
AI agents are increasingly being developed to accelerate scientific discovery, yet their practical capabilities in real research settings remain poorly understood. Existing benchmarks for AI agents ra…
Users rely on execution traces to observe agent behavior, diagnose failures, and ensure accountability. These traces contain rich procedural detail, including tool invocations, intermediate decisions,…
The recent success of agent swarms has shifted the paradigm of large language model (LLM)-based agents from single-agent workflows to multi-agent systems, highlighting the importance of agent orchestr…
Scientific discovery workflows usually contain and rely heavily on lab notes, where researchers record observations, interpret uncertain results, and plan follow-up experiments. Such informative lab n…
AI 点评 · 用Rust重写Git核心工具,提升性能与安全性,展现AI辅助代码重构新可能。
AI 点评 · 语音代理处理双语用户的真实能力首次被量化评测,揭示多语言交互技术的重大突破。
Data tells stories that shape society; the data journalist's job is to turn raw information into stories non-experts can trust. A high-quality news feature takes a newsroom team weeks: hunting for con…
Large language models (LLMs) are rapidly acquiring capabilities relevant to biological research, from literature synthesis to interpretation of experimental data. Increasingly, LLM agents can also per…

In this post, we demonstrate how a hands-free FNOL intake system combines agents built with the Strands Agents SDK for domain reasoning with Amazon Bedrock AgentCore Browser Tool f…

This post shows engineering teams how to apply that principle to one of the most time-sensitive workflows in engineering: incident triage. You will build a custom incident triage a…
At Deno we've been using OpenClaw and other agents increasingly for addressing production problems in Deno Deploy - when a PagerDuty alert fires, the agent starts researching the c…
The definitive OpenAI, Claude, MCP, Harness, Evals, and Production Agent Systems learning roadmap.
In this paper, we propose EEVEE, the first multi-dataset test-time prompt learning framework for LLM agents, enabling test-time prompt learning under real-world task streams. Existing methods are larg…
Reinforcement learning with verifiable rewards (RLVR) is a promising approach for enhancing reasoning and agentic behavior in large language models. However, rollout-intensive policy optimization is o…
大公司: 猫眼娱乐:作为首批内测开发者接入微信AI生态布局 36氪获悉,猫眼娱乐宣布作为微信AI生态首批内测开发者之一,旗下小程序接入微信AI Agent生态,借助微信AI Agent能力,为用户提供影片演出推荐、附近影院筛选、智能选座、一键支付等服务。 京东方A:控股子公司拟终止向不特定合格投资者公开发行股票并撤回申请文件 36氪获悉,京东方A公告,公司控…
Agentic, long-horizon visual generation: a fuzzy story → a cross-model-audited image-based movie. Brings ARIS's research-wiki + multi-agent debate to multimodal…
Although Large Language Model (LLM) agents have demonstrated strong performance on complex tasks, their learning is often limited by inefficient interaction feedback and static training environments,…
As adoption of AI agents looks set to surge by as much as 300% in the next two years, leadership teams are carefully considering the implications of a hybrid human-AI workforce. Un…
Practical patterns, starters & CLI tools for loop engineering with AI coding agents. Design systems that prompt and orchestrate agents (inspired by Addy Osmani…
一个入口串起全栈智能体
AI 点评 · 整合企业AI入口,降低使用门槛,或重塑行业竞争格局。
Multi-agent systems (MAS) can scale large language model reasoning at test time by decomposing complex problems into parallel subtasks. However, most existing MAS rely on centralized orchestration, wh…
We introduce Kwai Keye-VL-2.0-30B-A3B, an open-source Mixture-of-Experts (MoE) multimodal foundation model designed to advance long-video understanding and agentic intelligence. To address the challen…
Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks. However, existing benchmarks rarely evaluate whether agents can operate graphical u…
As the capabilities of LLM-based code agents continue to advance, their expected role is expanding beyond localized bug fixing in existing codebases toward architecting and implementing complete softw…
Autonomous web navigation remains challenging for LLM agents, and the strongest generalist systems rely on proprietary reasoning models whose inference cost is prohibitive for the repetitive tasks whe…
As AI systems built from multiple language-model agents become more common, they are increasingly used to make decisions together: discussing, negotiating, and acting on shared tasks. While individual…
Data tells stories that shape society; the data journalist's job is to turn raw information into stories non-experts can trust. A high-quality news feature takes a newsroom team weeks: hunting for con…

73 packages run self-replicating stealer as soon as they're opened by an AI agent.
AI 点评 · 微软供应链再遭投毒,AI代理环境成恶意软件新温床,安全防线亟待升级。
AI 点评 · 蚂蚁国际推动AI支付标准统一,有望降低跨境交易壁垒,影响全球支付格局。
Multi-agent code generation offers a promising paradigm for autonomous software development by simulating the human software engineering lifecycle. However, system reliability remains hindered by LLM…
AI 点评 · 提出自适应语义熵指标,精准衡量代码质量,突破多智能体编程可靠性瓶颈。
AI 点评 · 华为云转向提升token生产力,标志AI云竞争从规模转向效率与智能应用。
Advanced scientific simulators expose specialized input languages that turn simulation goals into executable configurations, but learning them can cost domain scientists hours to days. We study simula…
AI 点评 · 用代码生成适配器降低科研模拟门槛,让科学家专注研究而非编程。
A useful phone agent needs to be personally intelligent. It should reason over a user's identity, history, and preferences as they exist on the device, not just follow isolated instructions in an impe…
AI 点评 · 首个衡量手机智能体个性化推理能力的基准,填补了当前AI助手忽略用户身份与历史数据的评测空白。
AI 点评 · 科学智能赛推动AI与科研融合,吸引近两万参赛者,展示前沿技术竞争态势。

Amazon Bedrock AgentCore Runtime gives each agent session its own isolated microVM with a persistent workspace, secure tool access through Gateway, and built-in observability—so yo…
AI 点评 · 亚马逊推出隔离微虚拟机托管编码代理,兼顾安全与可观测性,或重塑云上AI开发流程。

In this post, we walk you through the Nova Sonic Test Harness, an open source framework that we built to solve both problems. It serves as a rapid iteration tool for tuning system…
Autonomous AI dev-team bot
协助用户与商家判断智能体可信赖程度
AI 点评 · 企业级AI应用爆发,基础设施压力剧增,巨头被迫裁员并重构核心产品以应对未来挑战。
文|王毓婵 编辑|张雨忻 6月5日,腾讯云AI产业应用大会最受外界关注的是什么? 毫无疑问,是汤道生与姚顺雨的对话。 在这场发布了一系列覆盖20多个垂直场景Agent的大会上,因为产品过于To B,也没有提及大家最关注的“微信AI”,导致外界的关注重心几乎全部被那场对话吸引走。 腾讯集团高级执行副总裁、云与智慧产业事业群CEO汤道生,在对谈中问腾讯首席AI科…
🔥🔥🔥 Turn AI-written code into real apps. Nubase is an open-source, AI-native backend platform for AI Coding, agentic applications, and modern product teams:…
Open-source TypeScript terminal coding agent for DeepSeek-V4 — builds on DeepSeek's strong price-performance and ultra-cheap cache pricing, engineering byte-sta…
Medical agent systems are increasingly expected to support interactive clinical decision making rather than only static question answering. In such settings, effective agents must reuse prior experien…
36氪获悉,近日,蚂蚁国际面向全球电子钱包、超级应用和数字银行等移动支付服务正式推出移动智能体协议(Agentic Mobile Protocol,简称AMP),解决全球消费者在AI智能体(agents)中购物时支付便捷、安全、有保障,以及商家跨市场互联互通等智能体商业全球化运营的关键难题。
AI 点评 · 蚂蚁国际首创移动智能体协议,破解AI购物支付与全球互联难题,引领跨境支付新标准。
Release: datasette-agent-edit 0.1a0 I'm planning several plugins for Datasette Agent which can make edits to existing pieces of text - things like collaborative Markdown editing, u…
AI 点评 · AI文本编辑插件迈出关键一步,轻松实现多人协作修改,值得关注。
6月7日晚间消息,京东与腾讯将围绕AI Agent展开深度合作。依托京东的商品供应链、履约服务能力及腾讯的生态入口优势,双方将共同打造跨场景的智能化服务新范式,推动AI Agent从单点应用走向生态协同。(新浪科技)
AI 点评 · 京东与腾讯联手,AI Agent从单点走向生态协同,或重塑电商与社交的智能服务边界。
Agent benchmarks score submissions with outcome verifiers that are typically hand-written and brittle, leaving them open to reward hacking. We audit 1,968 tasks across five terminal-agent benchmarks a…
Vision-language model (VLM) agents are increasingly deployed in interactive game environments. Yet game benchmarks for VLM agents typically report a single first-attempt score per (agent, game) pair,…
Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and operate within the physical world. However, existing benchmarks predominantly rely on passiv…
Long-horizon agentic tasks pose a fundamental credit assignment challenge for outcome-base reinforcement learning: trajectory-level rewards verify final correctness but provide limited guidance on whi…
Large language models are increasingly expected to handle complex, long-horizon real-world tasks whose context demands can grow without bound, yet model context windows remain inherently finite. Recen…
Large language model (LLM)-based agents are increasingly used in interactive textual environments, from web navigation and code editing to tool use and long-horizon dialogue. Yet many remain largely r…
As recommender systems transition toward agentic, multi-turn conversational interfaces, evaluation paradigms have struggled to keep pace. Current benchmarks often rely on "LLM-as-a-judge" evaluations,…
Visual reasoning requires integrating evidence distributed across regions, attributes, and relations, making single-chain reasoning prone to early perceptual commitment and hallucination. We propose V…
Computer-use agents (CUAs) increasingly operate in runtimes that combine visual desktop control, command-line execution, code editing, browsers, and external tools. Existing benchmarks, however, often…
Multimodal Foundation Models (MFMs) have made substantial progress, yet remain fragile in spatial reasoning over the physical world. A key bottleneck lies in their inability to transform local egocent…
A useful phone agent needs to be personally intelligent. It should reason over a user's identity, history, and preferences as they exist on the device, not just follow isolated instructions in an impe…
Local proxy that compresses your LLM API requests so you pay less, with no change to the answers. Trims wasted tokens from prompts, history, tool output, and co…

IT之家 6 月 7 日消息,华为云现已针对 Agentic AI 时代发布全新云入口“智果园”,新产品支持云码道 CodeArts 代码智能体、华为云 OfficeAce 办公智能体和 WorkAgent 文档智能体。 据介绍,智果园拥有开发、办公等多种关键行业的智能体,可通过智果 AgentArts 平台打造更加实用的智能体,并通过 Skills、AI…
AI 点评 · 华为云整合多模型,降低企业AI应用门槛,凸显生态开放与行业落地潜力。
IT之家 6 月 7 日消息,据钛媒体今日消息, 京东与腾讯已于近期联手,将围绕 AI Agent 展开合作 。京东的商品供应链与履约服务体系,将与腾讯的入口资源进行对接。 此外,消息称 京东 AI Agent 与华为、OPPO、荣耀等多家主流终端厂商已进行对接 。通过 A2A(Agent to Agent)合作,用户可直接在各终端原生智能体的京东 AI A…
AI 点评 · 巨头联手布局AI Agent,消费场景与社交入口的深度融合将加速智能体商业化落地。

At GTC Taipei at COMPUTEX last week, NVIDIA unveiled RTX Spark, the superchip that reinvents Windows PCs for the era of personal AI agents. On the heels of this announcement, NVIDI…
AI 点评 · 英伟达联合韩国顶级游戏厂商和战队推广RTX Spark,标志AI PC芯片落地游戏场景。
Own your AI video pipeline. LTX-2.3 (22B) self-hosted on your Modal GPU via a Claude Code skill — t2v, i2v, keyframes, v2v + synced audio. ~$0.02 per 5s clip, i…
Light — 全流程科研技能包:28 个技能覆盖文献调研到投稿全流程,配套 9 个可核查知识库。适配主流 AI 编程客户端。
An AI workflow skill pack for research, competitions, and innovation projects.
The social economy for autonomous AI agents.
Enterprise-grade sovereign AI for organizations where SaaS AI isn’t an option.
Expert writing feedback from experienced researchers is critical for early-career scholars to improve their manuscripts, yet high-quality feedback often remains scarce because reviewing research paper…
IT之家 6 月 6 日消息,微软在 Build 2026 上与高通联合发布 Project Solara。Project Solara 主打“智能体优先计算”,系统内部运行 Agent Shell,并能动态加载、调整多个基于云端的 AI 智能体。 微软 CEO 萨提亚 · 纳德拉表示:AI 智能体已经 不再只是普通 AI 助手 。“真正的平台转变正在发生。…
AI 点评 · 微软新项目争议背后,揭示AI从“助手”到“智能体”的平台级变革。
AI 点评 · 聚焦书安智能体操作系统实践,胡侠的分享将揭示AI系统底层架构的前沿突破。
AI 点评 · 性能翻倍且直指AI开发痛点,Next.js这次更新对全栈开发者极具吸引力。
AI 点评 · 用3B参数模型实现多智能体经济,低成本开源方案或将重塑AI生态。
Code based orchestration for any coding agent.
LLM agents increasingly rely on external inference conditions: prompts, tools, memory, SOPs, skills, and harness feedback. These assets can improve task execution without changing model weights, but t…
Agent skills provide a lightweight mechanism for extending general-purpose agents, but their open format exposes them to skill-poisoning attacks. A practically dangerous injection must stay invisible:…
AI 点评 · 用工程化思维把AI代码助手变成自主智能体,是开发效率的新突破。
Humans learn from social life. Simulating this process with LLM-powered agents represents a promising research direction, raising a natural question: whether LLMs can learn from such simulated social…
AI 点评 · AI社会模拟研究突破:用LLM代理模拟人类长期社交学习,探索机器从社会互动中进化。
Current Vision-Language Models struggle with hours-long videos because processing full-length visual sequences induces prohibitive token explosion and attention dilution. To overcome this, we introduc…
AI 点评 · 分层记忆架构破解长视频理解瓶颈,用图记忆与智能检索分离感知推理,显著降低计算成本。
Decentralized stochastic optimization is a fundamental paradigm for large-scale learning over networks, where agents communicate only with their neighbors and no central coordinator is required. For s…
AI 点评 · 加速去中心化随机梯度下降,破解大规模网络学习效率瓶颈,强凸优化领域关键突破。
Frontier AI systems are bridging the gap between intelligence and utility by shifting from conversational assistants to autonomous agents that execute tasks end to end. Using production data from Perp…
AI 点评 · AI代理将知识工作从辅助变为自主执行,效率与范围迎来质变。
As foundation models advance and agent scaffolding becomes increasingly sophisticated, agents have demonstrated remarkable proficiency in complex, long-horizon coding tasks and even autonomous experim…
AI 点评 · 聚焦前沿模型在科研全流程的自主能力,评测框架填补了现有基准空白。
Standalone self-improvement harness for agent traces, memory, skills, evaluation, and fine-tuning exports
Matrix is the cognition and UX layer on top of Paxeer Network. It turns natural-language requests from non-developers into a typed, inspectable, correctable Int…
Agentic AI的算力焦虑,英特尔给来了一剂「猛药」
AI 点评 · 用CPU突破AI算力瓶颈,英特尔开辟了低成本高性能的新路径。
This paper explores agentic 3D spatial understanding, i.e., MLLM agents performing 3D reasoning through tool use. Existing methods often misuse tools and exhibit biased tool preferences under 3D scena…

Data Management
On June 5, 404 Media reported that attackers had been using Meta’s AI customer support agent to steal Instagram accounts. Their approach was simple: They asked the agent to link th…
AI 点评 · Meta AI客服漏洞暴露了安全盲区,提醒行业警惕简单攻击手法。
A dynamic-workflow engine for any agent, any model: concurrent leaves, cost-aware routing, schema enforcement, crash-resume, bounded iterate-to-goal loops with…
深度调研报告生成 Skill — 一条命令,十分钟出券商级深度调研报告 / Professional deep research report generation Skill · Supports 19 languages
文|王毓婵 梁键强 编辑|张雨忻 昨日,腾讯客服回应称,微信正在与华为、小米、荣耀、OPPO、vivo等手机厂商合作推出A2A助手能力,目前已有多家厂商完成接入。 “您可以通过对应手机系统的AI助手发起 微信音视频通话 或 向指定好友发送消息 。该功能基于A2A(Agent-to-Agent)协作机制,数据安全与隐私通过双重授权机制保障。合作旨在将微信高频沟…
AI 点评 · 微信与手机厂商打通AI协作,标志社交场景向多终端智能体互联迈出关键一步。
今日热点导览 微信正与手机厂商合作推出Agent-to-Agent助手能力 擅用“LABUBU”相近标识商业推广,泡泡玛特告奈雪的茶获赔32万 腾讯客服回应与华为、小米等合作 神农旅游集团就国道收费进行道歉 苹果智能眼镜推迟至2029年,无显示屏AI眼镜仍将于2027年推出 TOP 3 大新闻 分析师曝苹果Vision Pro产品线被移除,2027年推AI眼…

IT之家 6 月 5 日消息,科技媒体 Appleinsider 昨日(6 月 4 日)发布博文, 报道称苹果批准 Poke 成为首个接入 Apple Messages for Business(苹果商务消息)平台的第三方 AI 智能体。 Apple Messages for Business 原本服务企业客服沟通的通道,苹果调整后开始承载更主动的 AI 助…
AI 点评 · 苹果开放iMessage接入AI,标志其生态向第三方智能体迈出关键一步。
To my knowledge, this is the first formally verified implementation of an intersection algorithm for polygons. The experience of working with AI agents on this project changed a lo…
LLM-driven software engineering agents have become a central testbed for real-world language-model capability, yet their training remains limited by the availability of high-quality SWE tasks. Existin…
Retrieval for search agents is still inherited from non-agentic information retrieval: a retriever ranks the corpus and the agent reads a small set of returned documents. Recent direct corpus interact…
Deep Research (DR) has emerged as a new agentic paradigm to tackle complex, open-ended research tasks, demanding systems that can iteratively frame problems, acquire evidence, verify sources, and synt…
Repository-level coding benchmarks such as SWE-bench have driven a rapid surge in the capabilities of coding agents. Yet they usually treat coding tasks as a holistic, binary prediction problem (e.g.,…
Deep research agents have demonstrated remarkable capabilities in complex information-seeking tasks, yet this power comes at a steep computational cost. Driven by accuracy-focused training paradigms,…
A growing failure mode in agent evaluation and training is that models can achieve high evaluation scores by exploiting shortcuts instead of solving the intended task, producing deceptive performance.…
Current Vision-Language Models struggle with hours-long videos because processing full-length visual sequences induces prohibitive token explosion and attention dilution. To overcome this, we introduc…
Deep research agents have attracted increasing attention for their ability to collect large-scale online information to acquire target knowledge, with recent efforts shifting from purely text-based in…
Computer-use agents (CUAs) rely on visual observations of graphical user interfaces, where each screenshot is encoded into a large number of visual tokens. As interaction trajectories grow, the token…
Are tool-calling LLM agents equally safe throughout a conversation? We discover they are not: agents are most vulnerable at the very start of a session and become substantially safer after a few regul…
Agentic reinforcement learning (RL) is emerging as a critical post-training paradigm for improving LLM agent capabilities. Existing RL algorithms for LLMs largely follow the token-centric paradigm as…
Poke, the startup that lets people use AI agents through simple text messages, has become the first AI agent approved for Apple’s Messages for Business platform.
AI 点评 · 苹果首次在商业消息平台引入AI代理,标志其AI生态向第三方开放迈出关键一步。
AI 点评 · AI自主修复数据库故障,展示Agent在运维中的实际价值。
For a humanoid robot to be deployed in the real world, the choice of command space (i.e., the interface between task planning and whole-body control) is crucial. Existing whole-body controllers typica…
AI 点评 · 企业级Agent的语义理解与事实核查能力需明确区分,直击当前AI落地的核心痛点。
We introduce Goedel-Architect, an agentic framework for formal theorem proving in Lean 4 centered on blueprint generation and refinement. A blueprint is a dependency graph of definitions and lemmas th…
As autonomous LLM agents increasingly hold real credentials and operate infrastructure without a human in the loop, operators have no standard way to tell an agent that a resource is off-limits. Acces…

Deploy NVIDIA Nemotron 3 Ultra on Amazon SageMaker JumpStart. Get 5x faster inference and 30% lower cost for agentic AI workloads with this frontier reasoning model.
AI 点评 · 英伟达顶级推理模型登陆云平台,5倍推理加速与30%成本降低,企业AI部署门槛再降。
Hi HN, we’re Nick and Drew, and we’re building boxes.dev – the first cloud-only agentic dev environment (ADE) that gives every Codex and Claude Code agent its own cloud computer. W…
Learn how Endava is using AI agents, ChatGPT Enterprise, and Codex to accelerate software delivery, automate workflows, and build an AI-native culture across the enterprise.
We launched Infracost on HN five years ago ( https://news.ycombinator.com/item?id=26064588 ) where our CLI generated cost estimates for infra-as-code, e.g. "this Terraform PR adds…
大公司: OpenAI首席财务官谈公司AI设备:今年年底前将正式发布 OpenAI首席财务官Sarah Friar日前在受访时透露,已经亲自体验过OpenAI的AI设备。她表示到“今年年底之前”OpenAI将正式发布这款产品。此前,OpenAI曾在一份文件中表示,预计最早也要到2027年2月才会开始发货。(财联社) LG Innotek计划扩建半导体基板工厂…
Multicultural multi-agent systems are increasingly deployed in globally diverse settings, where different agents are grounded in different cultural backgrounds. Existing cultural evaluation focuses on…
AI agents rely on a harness of skills, tools, and workflows to solve complex problems. Continually improving this harness is essential for adapting to new tasks. However, existing optimization methods…
Xiaohei 2.0 Codex Skill for Chinese real-object article illustrations and long-scroll story images
Xiaohei 2.0 Codex Skill for Chinese real-object article illustrations and long-scroll story images

“IT早报”时间,大家好,现在是 2026 年 6 月 4 日星期四,今天的重要科技资讯有: 1、豆包:计划针对专业人群生产力需求推出豆包专业版,基础功能保持免费 豆包声明称,对于广大用户日常使用的豆包功能,包含搜索问答、写作生图、以及语音和视频对话等,将保持目前的免费服务,保证用户使用体验和习惯不受影响。>> 查看详情 2、SpaceX 敲定 IPO 发行…
AI 点评 · 多条科技资讯集中释放,豆包专业版与免费策略、SpaceX估值与上市动态均值得关注。

MIT researchers use the classic game as a test bed for AI agents, finding a small AI model can outperform the biggest ones at 1 percent of the cost.
AI 点评 · 小模型通过游戏策略训练,以1%成本超越大模型,展现了高效学习新路径。
A situated query like "where is Lin Wei?" often encodes more than its literal content: the user may also want to know whether Lin Wei is free, in a good mood, or worth interrupting now. Standard tool-…
AI research often requires decisions before future evidence exists: which bottleneck to attack, which direction to pursue, or where a project should be positioned. We introduce ForeSci, a temporally c…
Planning for real-world problems by language models often involves both world and user constraints, which may not be fully specified upfront and are progressively disclosed through interaction. Howeve…
Role-playing language agents (RPLAs) should play characters whose values and behavior evolve as the story progresses, not maintain a fixed persona. Existing benchmarks measure factual recall at a give…
Large language model (LLM) agents are increasingly applied to long-horizon tasks such as scientific discovery and machine learning engineering (MLE), where sustained self-evolution becomes a key capab…
Inference-time skill augmentation provides a lightweight way to improve data-analytic agents by injecting reusable procedural knowledge without updating model parameters. However, discovering effectiv…
While Vision-Language Models (VLMs) have shown strong visual reasoning capabilities, their spatial reasoning abilities remain largely constrained to the observed images and text-oriented chain-of-thou…
Persistent AI assistants, such as OpenClaw, accumulate large collections of related memories over long-term interactions. As these memories grow, they may reinforce one another, diverge across context…
Existing benchmarks evaluate Tool-Integrated Reasoning (TIR) in LLMs on idealized ''happy paths'', largely overlooking real-world tool failures. We introduce ToolMaze, a benchmark for dynamic path dis…
Self-evolving agents requires adaptation after deployment, but existing approaches assume a usable learning loop, such as curated skills, successful trajectories, or verifier signals. Real open-world…
Training vision-language web agents with multi-step RL is compute-intensive, with two dominant forms of inefficiency: idle GPUs in synchronous RL, and trajectories that use more steps and tokens than…
Agent systems increasingly use textual skills to encode reusable task procedures, but injecting these skills into the prompt at every step incurs substantial context overhead and exposes skill content…
Large language models deployed as agents over large tool catalogs face a critical tool-retrieval bottleneck. As embedding-based retrieval approaches rely on compact encoders that may under-capture spe…
Despite recent progress, LLM agents still struggle with reasoning over long interaction histories. While current memory-augmented agents rely on a static retrieve-then-reason paradigm, this rigid pipe…
Self-hosted dev sandboxes with preview URLs. One command. No Kubernetes, perfect for coding agents and Saas factories
Self-hosted dev sandboxes with preview URLs. One command. No Kubernetes, perfect for coding agents and Saas factories
This week we've got tandem hands-ons with Google's new Gemini AI agent - Spark - from my colleagues David Pierce and Jay Peters. Their takeaways are similar: It's so effective that…
AI 点评 · AI能力提升反暴露技术天花板的矛盾,揭示行业深层困境。

In this post, you learn how to use Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO) together to improve the tool-calling accuracy of a small language model (SL…
What makes a robot gripper useful isn’t that it can pick up one object — it’s that it can pick up the next one, and the one after that, with a tool it’s never held before. What mak…
At CVPR, NVIDIA is unveiling new physical AI agent skills that help researchers and developers speed the development of autonomous vehicles, robots and vision AI systems. The core…
MARVIS-Agent: all-purpose credit risk agent for model development, validation, data processing, feature engineering, and strategy workflows.
WhatsApp will charge businesses for using its AI agent based on token usage.
Coralogix is among a growing number of infrastructure firms betting that as AI systems move into production, demand will rise for tools that can monitor their behavior, troubleshoo…
Beautiful, AI-native markdown editor and LLM Wiki
Self-contained UV scripts for data & ML tasks — OCR, vision, audio & more — run one in a command, locally or on Hugging Face Jobs. Built for humans and agents.
Self-hosted AI novel generator: single Go binary + web UI. OpenAI-compatible API → outline → chapter-by-chapter writing with review, foreshadowing, fact-check,…
36氪获悉,据腾讯云官号账号信息,腾讯2026AI产业应用大会即将在北京举办。作为腾讯年度最重要的AI产品发布平台,此次将发布系列智能体应用新品,并将公布infra等基础设施升级新进展。与此同时,腾讯集团高级执行副总裁、云与智慧产业事业群CEO汤道生将于腾讯AI首席科学家姚顺雨同台对话,解读AI下半场腾讯在AI赛道的最新布局和思考。
AI 点评 · 腾讯年度AI战略窗口,智能体新品与基础设施升级同步亮相,产业布局信号明确。
IT之家 6 月 3 日消息,据财经杂志报道,腾讯人士表示,目前无法确定微信 AI 智能体何时推出,其上线时间很大程度上取决于监管方对智能体的审批进度,微信 14 亿的用户体量,合规流程可能比其他产品更加严格。关于微信智能体,腾讯相关负责人表示暂无回应。 IT之家注意到,此前英国《金融时报》报道称,微信将推出一款 AI(人工智能)智能体,计划最快将于本月启动…
AI 点评 · 监管审批成关键变量,14亿用户规模下的合规挑战值得关注。
昨日,媒体报道称,微信将推出一款AI(人工智能)智能体,计划最快将于本月启动公开上线前所需的合规审批流程。腾讯人士表示,目前无法确定微信AI智能体何时推出,其上线时间很大程度上取决于监管方对智能体的审批进度,微信14亿的用户体量,合规流程可能比其他产品更加严格。关于微信智能体,腾讯相关负责人表示暂无回应。(财经)
AI 点评 · 微信14亿用户体量下,AI智能体合规审批进度成焦点,决定产品上线时间。

IT之家 6 月 3 日消息,科技媒体 Windows Latest 今天(6 月 3 日)发布博文,报道称在 2026 年 Build 开发者大会上,微软明确 Windows 11 系统定位: 不再只是带 AI 功能的桌面系统,而是要成为 AI 应用和智能体的开发平台。 微软新方向涵盖智能体 Runtime、本地模型、Windows 原生 AI 接口、Li…
AI 点评 · 微软从用户工具转向开发者平台,AI生态野心浮出水面。

IT之家 6 月 3 日消息,天风国际证券分析师郭明錤今天(6 月 3 日)在 X 平台发布推文,再次评论英伟达的 RTX Spark, 认为该处理器在未来 2 年内仍是小众产品,苹果在 WWDC 上对于设备端 AI 智能体的回应将是除 Siri 之外的另一个观察重点。 郭明錤表示英伟达 RTX Spark 处理器的核心看点不仅在于芯片本身, 更重要的是黄仁…
AI 点评 · 分析师视角揭示英伟达布局端侧AI的战略意图,市场影响值得关注。
Repo: https://github.com/getpaseo/paseo Homepage: https://paseo.sh/ Discord: https://discord.gg/jz8T2uahpH

Microsoft missed the boat on apps, so get ready for agents.
AI 点评 · AI代理需借鉴RSS,实现信息自主订阅与高效分发。
Deontic reasoning is the task of answering questions by applying explicit rules and policies to case-specific facts, for example computing tax liability under a statute or determining the outcome of a…
Large language models (LLMs) are increasingly proposed as clinical agents, yet static, single-turn benchmarks cannot capture how a model dynamically delivers care across an encounter: gathering inform…
Lane-level maps are critical infrastructure for autonomous driving and lane-level navigation, yet constructing and maintaining standardized lane networks for hundreds of cities remains highly labor-in…
Multi-agent reasoning systems adopt a "generate-then-transfer" paradigm that forces end-to-end latency to scale linearly with pipeline depth. We introduce StreamMA, a multi-agent reasoning system that…
Agents are widely deployed as assistants over documents, tools, and code. However, they typically act only on explicit user requests, which surface only the problems the user has noticed, while many o…
System prompt optimization improves agent behavior without modifying the underlying model, yielding human-readable, model-agnostic instructions. Existing methods build a prompt agent that refines task…
Experience internalization converts contextual experience from past interactions into reusable parametric capability, offering a promising path toward continual learning in large language models (LLMs…
We study the personal camera roll visual question answering setting. In this setting, a conversational AI assistant can access a user's personal camera roll and retrieve relevant photos to answer quer…
Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional domains. We argue tha…
Language agents increasingly rely on reusable skills to improve multi-step web automation across related tasks. A growing line of work studies online skill learning, where agents continually induce sk…
Multi-agent systems (MAS) built on large language models are typically organized around roles, pipelines, and turn schedules, while the content that agents pass to one another is often left as unconst…
Release: datasette-agent-micropython 0.1a0 I want Datasette Agent to be able to generate and execute Python code safely. This alpha is looking promising so far. GPT-5.5 has so far…
AI 点评 · 结合AI代理与MicroPython,实现安全代码生成执行,为数据探索带来新可能。
Release: micropython-wasm 0.1a1 Fixes for some limitations that emerged while I was trying to use this to build datasette-agent-micropython . Tags: python , sandboxing , webassembl…
AI 点评 · 在浏览器中运行MicroPython,为Python沙箱执行和Web应用开辟新可能。
Open-source observability tool that uses AI agents to self-heal your software

The agentic AI moment has arrived, but delivering on its promise requires more than good models. It also takes fast hardware, secure runtimes, a responsive data layer and models tu…
AI 点评 · 英伟达与微软联手打造统一AI代理栈,打通从终端到本地部署,降低开发门槛。
AI 点评 · 物理AI融合现实世界,英伟达开源补齐工具链,加速机器人自主决策。
https://www.microsoft.com/en-us/microsoft-365/blog/2026/06/0... https://www.404media.co/microsoft-wants-to-make-people-addic... https://www.wired.com/story/meet-microsoft-scout-you…
AI 点评 · 微软自研AI代理Scout基于OpenClaw,标志着巨头在自主智能体领域的战略布局。
The specification lets developer, compliance, and security teams define their own policies for agents to follow in portable policy files.
AI 点评 · 微软推出便携式策略文件,让开发者自主定义AI代理行为规范,提升安全可控性。
Reward models (RMs) provide critical feedback signals for LLM post-training, notably in reinforced fine-tuning (RFT) and reinforcement learning (RL) pipelines. However, current reward evaluation relie…
Large language models improve final-answer accuracy through extended chain-of-thought reasoning, but often spend tokens inefficiently and offer little inference-time control. Existing efficient reason…
Deep reinforcement learning has shown strong potential for enabling autonomous robots to learn complex navigational tasks. However, its practical use still depends heavily on human designed reward fun…
As AI systems increasingly assist humans in physical tasks, ensuring safety becomes paramount -- physical actions carry immediate and irreversible consequences that digital errors do not. We introduce…
Microsoft just announced "Project Solara," a new OS designed for gadgets that run AI agents, at Build 2026. The company is calling it "a new platform built from the ground up to po…
AI 点评 · 微软专为AI智能体硬件打造操作系统,标志从软件到硬件生态的关键一步。
AI 点评 · 深度研究智能体落地生产,揭示了AI从实验室到实战的关键教训。

This post walks through how Baz built their Spec Review agent using Amazon Bedrock and Amazon Bedrock AgentCore. We'll cover the architecture decisions, implementation details, and…
AI 点评 · 亚马逊Bedrock AgentCore让Baz的AI代码审查精度提升,展示云服务与AI结合的实际应
Large language models for code generation often need to use APIs that are absent from their pretraining data. This requires more than recalling a function name: models must coordinate signatures, modu…
AI 点评 · 评估大模型调用新API的能力,填补实用知识缺口,推动智能体从记忆转向推理。
Assessing the quality of time series (TS) data is fundamental yet inherently challenging due to the multifaceted nature of quality dimensions. Recently, large language models (LLMs) have emerged as a…
AI-assisted coding agents are bottlenecked by input-token cost. Two pathologies of raw human input drive much of this overhead: tokenization inefficiency for non-English text and structural entropy in…
Embodied agents that navigate cities rely on world models that predict how their surroundings will change as they move. But for navigation, what matters is not what the buildings look like; it is wher…
Open-source AI-era employment platform connecting skills, jobs, enterprises, governance, and AI agents.
Open-source AI-era employment platform connecting skills, jobs, enterprises, governance, and AI agents.
According to every product demo from the last four years, planning a trip is a killer use case for AI. Just tell it where you're going, they all promise, and your chatbot / agent /…
AI 点评 · 演示效果惊艳,但揭示出AI自主规划能力已逼近人类,引发对技术失控的深层担忧。

IT之家 6 月 2 日消息,据新浪科技今日报道,曾在华为主导盘古大模型研发的“90 后少帅”王云鹤,已于近期投身 AI Agent 领域创业,其新成立的公司“基元律动”已完成一轮估值达 1 亿美元的新融资。 王云鹤在今年 3 月末正式告别了工作近 9 年的华为。离职前,他最后的职务为华为诺亚方舟实验室主任、盘古大模型负责人,曾被誉为“盘古大模型少帅”和“天…
AI 点评 · 顶尖技术人才创业动向,折射AI Agent赛道资本热度与行业新趋势。

IT之家 6 月 2 日消息,英伟达于 5 月 31 日宣布,其面向智能体 AI 工厂的下一代超级计算平台 NVIDIA Vera Rubin 已进入全面量产阶段。IT之家此前已有相关报道。 除此之外,英伟达同时确认新一代 Spectrum-X 以太网硅光技术已同步进入全面量产阶段,这是该平台实现大规模 AI 工厂网络互联的核心基石。 作为全球首款基于光电一…
AI 点评 · 硅光技术量产突破,能效提升5倍,将加速AI工厂网络部署,改变行业格局。
IT之家 6 月 2 日消息,据澎湃新闻,英特尔 CEO 陈立武 2 日(今天)在台北电脑展上表示,CPU 需求越来越高,但供给受到限制。过去四周内, 许多公司 CEO 打电话给他要更多的 CPU ,对英特尔来说“是一个机会”。 AI 智能体的兴起,使中央处理器的重要性得以再次提升,从而带动需求大量增加。陈立武在谈到 CPU 的发展趋势时指出,AI 智能体需…
AI 点评 · 高管亲述供货紧张,反映AI时代CPU需求爆发,英特尔产能成关键变量。
The global health care sector is under increasing strain. Decades of chronic underinvestment and constraints in recruitment have coincided with a surge in demand for services for a…
AI 点评 · 用AI代理重构医疗流程,缓解人力短缺,提升服务效率与可及性。
Economical human-readable logs and structural (event pattern) retrieval for agentic trajectories
AI 点评 · Agent时代文档解析新突破,专家分享前沿基础设施演进,实战价值极高。

IT之家 6 月 2 日消息,据IT之家小伙伴今日反馈,腾讯客服最新回复显示, 微信正在与华为、荣耀、小米、OPPO、vivo 等手机厂商合作推出 A2A 助手能力 。 用户可以通过手机语音助理发起微信音视频通话或向指定好友发送消息。该功能基于 A2A(Agent-to-Agent)协作机制, 由厂商 AI 助手向微信发起指令,微信负责执行并返回结果 ,全程…
AI 点评 · 手机厂商AI助手与微信深度打通,标志着跨应用智能协作进入实用阶段。
Qwen3.7-Plus已上线阿里云百炼
AI 点评 · 通杀多模态与桌面软件,AI智能体能力再上台阶,开发者生态迎来新变量。

Agentic AI is getting physical. At COMPUTEX on Tuesday, NVIDIA announced NVIDIA JetPack 7.2 and NVIDIA NemoClaw support on NVIDIA Jetson. JetPack 7.2 brings agentic AI skills, Yoct…
AI 点评 · 英伟达让AI从虚拟走向实体,开启物理世界自主决策新纪元。

IT之家 6 月 2 日消息,阿里千问大模型今天(6 月 2 日)发布博文,宣布推出 Qwen3.7-Plus 模型, 定位为多模态交互混合智能体。 Qwen3.7-Plus 是 Qwen3.7 的多模态升级版,核心定位是视觉与语言统一的智能体基座。 它保留文本、编码、工具使用和生产力工作流能力,同时强化视觉理解、视觉推理和跨模态任务处理。 模型已通过阿里云…
AI 点评 · 多模态与智能体融合,或加速AI从“对话”迈向“行动”的关键一步。
If Nvidia has cracked a way to bring AI agents easily, safely, and usefully to the masses, it could — and should — be big.
AI 点评 · 英伟达联手微软戴尔惠普,将AI智能体推向PC,可能撬动2000亿美元CPU市场。

GPT-5.5, GPT-5.4, and Codex are now generally available on Amazon Bedrock. Deploy them in production applications and agents today, on Bedrock’s high performance inference engine.
AI 点评 · OpenAI模型登陆亚马逊云平台,企业应用部署门槛进一步降低。
Google's new "24/7" AI agent, Gemini Spark, can be shockingly good at doing things on your behalf. But I'm not sure it's worth the financial cost and potential privacy tradeoffs. T…
AI 点评 · AI助手能力接近演示效果,但隐私与成本的双重代价仍需权衡。
Large language models improve final-answer accuracy through extended chain-of-thought reasoning, but often spend tokens inefficiently and offer little inference-time control. Existing efficient reason…
LLM-agent budget overruns are a documented production failure class: a single retry loop can spend thousands of dollars before an operator notices, and the in-process integrity properties that would p…
Large language model (LLM) agents are evolving from request-response assistants into long-running software actors: they maintain state across model calls, fork subtasks, wait for external events, requ…
Multimodal agents in robotics, AR, and autonomous driving must reason about places and layouts from continuous egocentric streams, often using evidence outside the current view. Existing benchmarks ei…
Structured financial audit verification is difficult for language-model agents because correctness depends on structured evidence rather than text alone. A model must link reported facts to taxonomy c…
Computer-use agents extend language models from text generation to sustained interaction with files, terminals, browsers, and external tools. This shift creates safety risks that are difficult to dete…
Memory is an indispensable capability for long-horizon LLM agents, enabling them to preserve and utilize information accumulated across extended interactions. Existing memory-agent approaches are typi…
Recent progress in Large Language Model (LLM) agents has enabled promising advances in automated data science. However, existing approaches remain fundamentally limited by their static action sets and…
Reward models (RMs) provide critical feedback signals for LLM post-training, notably in reinforced fine-tuning (RFT) and reinforcement learning (RL) pipelines. However, current reward evaluation relie…
Equipping Large Language Models (LLMs) to execute reliable multi-step workflows has become a central challenge in artificial intelligence. Despite recent advances in LLMs' agentic capabilities, most a…
Curating training data is among the most consequential yet labor-intensive parts of modern AI development: practitioners iteratively propose, implement, evaluate, and revise data policies against nois…
Autonomous LLM training is often framed as recipe search, which leaves the training harness largely static. This limitation sharpens in agentic RL, where shifting bottlenecks and scalar rewards mask d…
Computer-Use Agents (CUAs) are increasingly deployed in dynamic interactive environments, creating a growing need for continual skill learning during interaction. Recent approaches address this challe…
GAN-style self-improvement loop for any text artifact: mutate, grade with a SEPARATE model, keep only verified wins (pairwise-judged), revert the rest. The git…
AI 点评 · 英特尔押注18A工艺的288核至强6+,标志着智能体正重塑CPU在AI调度中的核心地位。
Clinical practice is not the selection of an answer from enumerated options: a physician gathers heterogeneous information incrementally and commits to sequential, irreversible decisions under uncerta…
AI 点评 · 电子健康记录多阶段交互环境,弥合了AI临床决策与真实医疗流程间的鸿沟。
AI 点评 · Qwen3.7-Plus融合多模态与智能体能力,或开启AI应用新范式。

In this post, we use a lakehouse data agent to demonstrate how you can use Policy for deterministic access control and Lambda interceptors for dynamic validation. We then show how…
AI 点评 · 亚马逊Bedrock新功能实现AI代理安全管控,结合策略与动态验证,为行业提供可落地的防护方案。
We introduce HERO'S JOURNEY, a benchmark for rule induction in goal-directed episodic tasks, where agents must infer hidden rules from demonstrations and act on them through multi-step execution. HERO…
AI 点评 · 用文本游戏测试AI规则归纳能力,填补了复杂推理任务基准的空白。
Agent skills occupy a privileged position in the agent workflow, as agents are expected to implicitly follow and execute them, rendering third-party skills a vulnerable attack surface. Existing studie…
AI 点评 · 自动化构建技能生命周期攻击,揭示第三方技能在智能体流程中的隐蔽安全风险,需重视防御。
Text files such as skill files, memory files, and behavioral configuration files play a central role in defining how modern agents act. Through edits by humans or the agents themselves, these files ma…
AI 点评 · 追踪智能体行为轨迹,揭示自我调整机制,为AI决策透明化提供新视角。
Large language models now power robo-advisors and trading agents, yet whether they carry built-in biases toward specific assets is largely untested. We ask three questions: do LLMs systematically pref…
AI 点评 · 审计金融大模型对特定资产的偏好,揭示AI决策的隐性偏差,影响投资策略可靠性。

In this post, we address several key risks that surface when designing an agentic payment system, and how to address them with the capabilities of AgentCore payments.
AI 点评 · 用亚马逊Bedrock内置防护栏解决AI支付代理安全风险,为金融场景落地提供可靠方案。
AI 点评 · 斯坦福CS336课程发布AI代理开发规范,为学术与工业界提供权威参考。

When you build agentic AI solutions, you face unique operational challenges. Agents make unpredictable decisions, costs spiral unexpectedly, and debugging non-deterministic failure…
AI 点评 · 亚马逊Bedrock AgentCore让AI代理规模化运营更可控,破解成本与调试难题。
Personalization is a crucial capability of modern language agents. However, current research primarily positions personalized agents as passive responders to user preferences, limiting their ability t…
Biological image analysis increasingly demands integration across heterogeneous tools, programming environments, and domain knowledge that few researchers can command simultaneously. We present Agenti…
Deep-research agents solve tasks through long trajectories of search, tool use, evidence inspection, and answer synthesis. Evaluation based on final answers shows whether an agent succeeds, but not wh…
Building capable visual web agents requires long-horizon reasoning, precise grounding, and robust interaction with dynamic real-world websites. Despite rapid progress, the strongest systems remain lar…
AI 点评 · Anthropic推出托管智能体与主动工作流,标志着AI从被动应答向自主执行的关键进化。

IT之家 6 月 1 日消息,在今天的华为 nova 16 系列及全场景新品发布会上,华为终端 BG CEO 何刚正式发布了 FreeClip 2 耳夹耳机典藏版, 定价 1499 元 。 据介绍,华为 FreeClip 2 耳夹耳机典藏版采用鎏光宝盒 + 珠宝盒设计,充电舱采用真空镀膜工艺,主打“圆润璀璨”, 同时内部空间提升 20% 。 这款耳机还与周大…
AI 点评 · 将珠宝美学与AI智能体交互结合,为耳机品类带来轻奢体验与技术创新突破。
Abundant procedural knowledge on the Web holds great potential for helping agents solve long-horizon tasks. However, such knowledge is often multimodal, heterogeneous, noisy, and implicitly assumes hu…

IT之家 6 月 1 日消息,在今日的 2026 台北国际电脑展主题演讲中,英伟达 CEO 黄仁勋发布了“全球最强大的桌面 AI 超级计算机”—— DGX Station for Windows 。 DGX Station for Windows 用于在 Windows 上开发和运行智能体 —— 基于英伟达 GB300 Grace Blackwell Ult…
AI 点评 · 首次将企业级AI算力带入桌面端,为Windows生态开发者提供了本地化训练与推理的超级工具。

IT之家 6 月 1 日消息,为加强自主智能体的智能能力,英伟达今日发布了面向全天候运行智能体的全新开源模型与数据集,相关成果由英伟达 Nemotron 联盟联合打造。 据官方介绍,英伟达 Nemotron 3 Ultra 是一款拥有 5500 亿参数的混合专家模型,可为代码开发、科研及企业业务流程中的长效智能体提供顶尖智能能力。相较于同级别主流开源前沿模型…
AI 点评 · 参数规模与推理速度双突破,为智能体部署树立新标杆。

Personal agents are exploding in popularity, with open source projects like OpenClaw and Hermes seeing rapid adoption by AI developer communities on GitHub. Built to adapt to indiv…
AI 点评 · 英伟达将本地AI智能体部署到RTX电脑和DGX工作站,推动个人AI应用从云端走向本地化。

IT之家 6 月 1 日消息,在今日的 2026 台北国际电脑展主题演讲中,英伟达 CEO 黄仁勋宣布正式推出 Vera 处理器 。 英伟达 Vera 是一款专为 AI 智能体打造的 CPU ,速度比 x86 处理器快 1.8 倍,可驱动各行各业的多样化工作负载,Vera 现已全面投产。 Vera 以 Grace CPU 的成功为基础(迄今为止,Grace…
AI 点评 · 巨头下场定义AI智能体专用芯片,生态号召力预示行业新标杆。

IT之家 6 月 1 日消息,在今日的 2026 台北国际电脑展主题演讲中,英伟达 CEO 黄仁勋宣布 Vera Rubin 全面投产。 Vera Rubin 为下一代 AI 工厂提供了 POD 规模的基础架构 —— 与上一代 Grace Blackwell 平台相比, 其大规模智能体吞吐量提高了 10 倍 。 凭借成熟的开源 MGX 设计,英伟达供应链生态…
AI 点评 · 下一代AI算力跃升10倍,英伟达再次定义超大规模集群新标杆。
MiniMax M3 今日正式发布。 MiniMax M3 在编程和智能体等专业任务上达到了前沿的能力。它使用了全新注意力架构 MSA (MiniMax Sparse Attention),最高支持 1M 超长上下文。它也是一个原生多模态模型,支持图片和视频的输入,并能操作电脑桌面。 在衡量 Coding 能力的 SWE-Bench Pro 上,MiniMa…
LLM internship resume and job-search Codex Skill: resume polish, JD tailoring, evidence guard, interview grilling, and Project Scout for LLM/RAG/Agent roles. 大模…
LLMInternSkill: LLM internship resume and job-search Codex Skill for resume polish, JD tailoring, evidence guard, interview grilling, and Project Scout. 大模型实习简历…
LLM agents are increasingly expected to operate across heterogeneous task regimes that require distinct execution paradigms. This challenges fixed agent systems and motivates system-level meta-adaptat…
RuleGo 是一个基于 Go 语言的轻量级、高性能、嵌入式规则引擎。它通过规则链(JSON/可视化)编排组件,实现复杂业务逻辑的声明式管理,在物联网、边缘计算、数据集成、自动化等场景有广泛应用。 v0.36.0 是一个里程碑版本:rulego-components-ai 从 AI 组件库正式升级为声明式 AI Agent 开发框架,同时 Server 模块…
让每个模型都成为 Codex 引擎。 OpenAI 兼容的 Responses API 网关,让 Codex、CLI 工具和开发者 Agent 接入任意模型。 English Documentation · 中文文档 GodeX 让使用 OpenAI Responses API 的客户端,可以通过一个本地网关调用 DeepSeek、Xiaomi、MiniMa…
The Model Context Protocol (MCP) has emerged as a transformative standard for connecting large language models (LLMs) with external data sources and tools, and has been rapidly adopted across personal…
AI 点评 · 用环境模拟测试LLM在个人场景的真实表现,MCP标准首次有了专属评测基准。
Building capable visual web agents requires long-horizon reasoning, precise grounding, and robust interaction with dynamic real-world websites. Despite rapid progress, the strongest systems remain lar…
AI 点评 · 聚焦多轮强化学习框架,破解视觉智能体在动态网站中的交互难题,填补了实操性研究空白。
In open-ended environments, exploration is fundamental for autonomous agents, yet current language model agents struggle with this. Effective exploration requires memory, but retaining raw interaction…
AI 点评 · 结合新颖信号统一记忆与探索,让语言模型在开放环境中自主发现未知。
Computer use agents (CUAs) today are primarily deployed as single serial agents. This setup is suboptimal for complex long-horizon tasks that benefit from task decomposition, parallel execution, and c…
AI 点评 · 多智能体协作提升复杂长任务效率,突破单代理局限,值得关注。
Frontier model evaluations are shifting from foundational capabilities (e.g., instruction following and reasoning) toward compositional, agentic ones, but Korean agentic benchmarks remain scarce. We i…
AI 点评 · 首个聚焦韩语场景的网页浏览智能体评测基准,填补了非英语环境下的评估空白。
Deep Research Agents have shown strong capability in multi-step information retrieval, reasoning, and long-form report generation, but existing benchmarks and systems remain predominantly text-centric…
Search agents are often trained as policies over growing transcripts: the model must decide how to search while also remembering what it has seen, which evidence is useful, which constraints remain op…
Reinforcement learning (RL) improves large language model (LLM) agents by teaching them which actions lead to high rewards, but provides little supervision on what those actions do to the environment.…
Embodied visual navigation, where an agent perceives a complex environment and acts to reach a goal from raw sensory input, underpins a wide range of applications such as household service robotics, a…
How can a population of agents self-orchestrate and self-adapt into stronger collective intelligence without centralized control? Inspired by Friedrich Hayek's economic theory of decentralized coordin…
Deep-research agents solve tasks through long trajectories of search, tool use, evidence inspection, and answer synthesis. Evaluation based on final answers shows whether an agent succeeds, but not wh…
Agentic LLMs with web search change the threat model for text anonymization: weak contextual cues can become cross-referenceable evidence for re-identification, yet those same details also carry downs…
Financial AI agents often fail for a simple reason: they make users carry the complexity. A user must repeatedly restate goals, risk preferences, portfolio context, past judgments, and shifting market…
Large language models (LLMs) have recently been adopted as synthetic agents for public opinion simulation, offering a promising alternative to costly and slow human surveys. Despite their scalability,…
Agentic language model systems alternate between two structurally distinct step types: structured tool calls (short, deterministic, low perplexity) and open-ended planning/reasoning steps (long, compl…
Agent skills occupy a privileged position in the agent workflow, as agents are expected to implicitly follow and execute them, rendering third-party skills a vulnerable attack surface. Existing studie…
A 7-layer memory operating system for Hermes Agent — persistent memory with Qdrant, structured facts, fabric recall, auto-curated wiki, and surgical context inj…
A generalist autonomous research agent — runs experiments, researches, and iteratively optimizes, autonomously.
A local control plane for AI agents — see what they do, approve what matters, keep secrets out. Rust + Tauri + Chrome MV3.
local multi-agent harness
下一代CUA训练范式
AI 点评 · 解决Agent工具选择难题,复旦与通义提出全新训练思路,推动智能体实用化。
让世界模型迈向多智能体交互仿真
AI 点评 · 点明AI成本高的根本原因在于数据质量,为行业降本提供新思路。
AI 点评 · 多智能体协作从实验走向工程化,为AI研发团队提供可复用的基础设施范本。
Sisyphus Academica — The Research Paper Writing Army. 20+ agent swarm: 6 novelty engines, 10 adversarial reviewers, Humanizer-integrated writing, citation verif…
AI 点评 · 用20多个AI代理模拟学术生产链,挑战论文写作与评审的自动化边界。
CLI更像是Agent的母语
AI 点评 · Agent不再迁就人类交互方式,操作系统底层适配AI才是效率革命的关键。
与其焦虑AI,不如加入AI
AI 点评 · 揭示企业如何打破传统架构,全员转型智能体,为组织适应AI时代提供实战模板。
Open-source observability and trust layer for AI agents: trace every step, score every output, catch hallucinations and runaway loops in real time. Self-hostabl…
Procedural 3D modeling through code is emerging as a versatile paradigm, offering deterministic, engine-ready, and precisely editable assets that neural 3D generators inherently lack. Authoring such p…
AI 点评 · 首个用代码评估AI三维建模能力的基准,填补了程序化生成与神经渲染之间的测评空白。
Large language model (LLM) agents increasingly rely on reusable external skills to solve long-horizon interactive tasks. Existing training-free skill adaptation pipelines usually update skills from fu…
AI 点评 · 用轨迹数据让LLM代理自动进化技能,免训练自适应方案突破长程任务瓶颈。
Reusable skills are a key mechanism for extending agent capabilities, allowing agents to accumulate experience and solve increasingly complex tasks. Yet most existing skill-learning methods store reus…
AI 点评 · 视觉技能弥补语言局限,让智能体在复杂任务中更高效积累经验。
On-Policy Distillation (OPD) is a fundamental technique for efficient post-training of large language models (LLMs), with broad applications in agent learning, multi-task enhancement, and model compre…
Large language models are increasingly deployed as coding agents, shifting safety from individual responses to action sequences. Existing benchmarks, however, primarily assess whether models refuse un…
Reflexion-style agents rely on self-generated reflections as memory, implicitly assuming that agents can accurately diagnose their own failures. We show that this assumption can fail systematically: a…
We present a benchmark for evaluating AI models and agents on real-world formal software verification tasks. We first scrape 11,039 property-based tests (PBTs) from real-world Python repositories, the…
Molecode presents molecules as code and enables LLMs to operate and reason on chemistry directly.
Deterministic cost / loop / time budgets · full observability · crash-resumable runs · human-approval gates · a memory you own. Self-hosted. Your keys. No telem…
AI 点评 · 将确定性成本、循环时间预算与可恢复运行结合,为AI安全执行提供新范式。
AI 点评 · 测试智能体如何重塑质量工程,腾讯大牛现场揭秘实战经验。
让世界模型迈向多智能体交互仿真
AI 点评 · 多智能体交互是世界模型的关键突破,推动AI从单机游戏走向真实协作场景。
Agentic search requires language model agents to explore many sources and answer complex information-seeking questions. Scaling test-time compute is a promising way to improve these agents, but curren…
AI 点评 · 用细粒度自验证扩展测试时计算,首次系统解决智能体搜索中的错误累积问题,为复杂信息检索提供可扩展方案。
AI glasses present a compelling platform for AI agents to serve as personalized memory assistants. To be genuinely useful, such systems must move beyond short-term video comprehension and address memo…
Agentic search systems iteratively interact with retrieval models to answer complex queries. Despite substantial progress, optimizing retrievers for agentic search remains challenging, often requiring…
Language models can find thousands of severe software vulnerabilities, and agents are increasingly being misused for cyberattacks. To avoid detection, attackers frequently distribute their misuse, spl…
AI 点评 · 分布式智能体攻击难追踪,状态监测实现实时阻断,提升AI安全防御新高度。
Long-context reasoning remains a central challenge for large language models, which often fail to locate and integrate key information in extensive distracting content. Reinforcement learning with ver…
The same arguments often need to be evaluated under different external regimes. An agent with influence over the regime has a strategic lever that standard formalisms do not directly capture. We intro…
AI 点评 · 用博弈视角解析论辩情境依赖,为AI策略性语言操控提供全新建模框架。
AI 点评 · Robinhood允许AI代理直接交易,加速金融与AI融合。
AI-native coding orchestration platform: unified multi-model agent runtime with stateful sessions, tool governance, and traceable delivery.
As Large Language Models (LLMs) evolve from general-purpose assistants to user-centric agents, personalization has become central to aligning model behavior with individual preferences, making the eva…
AI 点评 · 个性化评估框架创新,让大模型更懂用户,提升人机交互体验。
Cognition makes Devin, the first and arguably most successful AI coding agent. But famed coder Wu says it isn't designed to supplant human programmers.
AI 点评 · AI编程工具定位为人机协作而非替代,创始人观点打破行业焦虑。
AI 点评 · 过度依赖编程Agent可能导致开发效率虚高、维护成本激增。
AI 点评 · 多智能体编排首次深入设计生产场景,展现AI协同解决复杂任务的工程突破。
AI image tools rarely make me feel like I'm part of the creative process. They are, after all, mostly designed so that people with no design experience can type in a few words and…
AI 点评 · 直击AI工具痛点:设计过程缺乏参与感,暴露当前技术局限。
KVarN is a native vLLM KV-cache quantization backend for your agents: 3-5x more context, throughput above FP16, and FP16-level accuracy. Calibration-free, one f…
The end of web parsing. The beginning of scalable pixel-native search.
AI 点评 · 将网页解析转向像素级原生搜索,为多模态检索开辟全新路径。
As AI agents move from experiments to production, AWS, Cloudflare, and others are redesigning cloud infrastructure for a future dominated by machine-generated internet traffic inst…
AI 点评 · 云巨头正为AI时代重构网络,机器流量将主导未来,基础设施变革迫在眉睫。
Local-first observability dashboard for AI agents. MCP-native. Look at every span your agents emit.

This post combines learnings from LangChain’s work on evaluating deep agents and Anthropic’s guide to demystifying evals for AI agents into a practical guide. In this post, you wil…
AI 点评 · 结合LangChain与Anthropic的评估经验,为复杂AI代理提供实用评测指南,填补行业方法论

Undisclosed addition in jqwik instructed AI coding agents to delete app output.
AI 点评 · 开发者用提示注入反制低代码乱象,揭示AI安全与人类创意间的冲突新战场。
Asana will incorporate StackAI into its growing suite of AI workflow tools.
AI 点评 · Asana收购无代码智能体构建工具,加速AI工作流布局,降低企业自动化门槛。
LLM agents are increasingly expected not only to complete isolated tasks, but also to carry bounded representations of human expertise, judgment, and interaction style. Building such person-grounded a…
AI 点评 · 用专家知识蒸馏让AI自动生成人类技能,大幅提升智能体专业性和拟人化水平。
Long-term memory is essential for multimodal agents to build coherent experience, accumulate world knowledge, and achieve continual learning. However, constructing effective memory goes beyond memory…
AI 点评 · 聚焦多模态智能体的长期记忆构建,突破传统记忆局限,实现持续学习与知识积累。
LLM agents are evolving from conversational chatbots to operational tools in real-world workspaces. In local agentic harnesses, an LLM can read and write files, call tools, and reuse workspace state a…
AI 点评 · 揭示LLM代理从对话到操作工具的安全漏洞,提出防御后门攻击的新思路,对AI安全至关重要。
Long-context reasoning remains a central challenge for large language models, which often fail to locate and integrate key information in extensive distracting content. Reinforcement learning with ver…
AI 点评 · 从搜索代理轨迹中学习长上下文推理,用评分奖励机制提升模型信息整合能力。
Monitoring autonomous language model agents currently relies mostly on surface behavior. But what happens when agent populations invent new languages with the goal of avoiding human oversight. Here, w…
AI 点评 · AI自主创制新语言规避人类监管,暴露智能体协同的深层安全风险。
Long-horizon search agents accumulate large amounts of retrieved content across many tool calls, making context-budget efficiency increasingly important. A minimal intervention is to mask stale observ…
AI 点评 · 用简洁机制揭示信息遮蔽策略的临界点,为长程搜索智能体优化提供实用边界。
LLM agents increasingly retrieve externally curated skills-procedural instructions retrieved at decision time-to improve performance on long-horizon interactive tasks. Existing skill libraries are typ…
AI 点评 · 打破通用技能库局限,提出模型感知对齐,让智能体任务适配更精准高效。
Multimodal large language models (MLLMs) have shown strong capabilities in perception, reasoning, and action generation. However, their ability to sustain exploration in dynamic open worlds remains un…
AI 点评 · 首个用《我的世界》评估多模态大模型开放世界探索能力的基准,填补了该领域测试空白。
Effective real-world assistance requires AI agents with robust Theory of Mind (ToM): inferring human mental states from their behavior. Despite recent advances, several key challenges remain, includin…
Humans can effortlessly perceive spatial layouts, form cognitive representations, reason about spatial relations, and translate such reasoning into actions in everyday 3D environments. Although recent…
For agents to learn continuously from interaction with the world at test time, they must be able to explore effectively, acquire new world knowledge and skills, retain relevant episodic experiences, a…

Datasets in AgentCore is in public preview. Agent evaluation is most powerful when you combine fast-moving online signals with stable offline baselines. To understand whether your…
AI 点评 · 亚马逊Bedrock新功能让AI代理测试集动态扩展,平衡线上信号与离线基准,实现持续性能追踪。

This post covers Opus 4.8's improvements and practical guidance for AI engineers integrating the model into agentic systems and production inference workloads on Amazon Bedrock.
AI 点评 · Claude新模型登陆AWS,为AI工程化部署提供关键升级,值得开发者关注。
Hi HN, we’re open-sourcing ktx. It’s an executable context layer that makes agents reliable on your data stack. We built it after going through the experience of building productio…
Hi HN, we’re open-sourcing ktx. It’s an executable context layer that makes agents reliable on your data stack. We built it after going through the experience of building productio…
AI 点评 · 一分钟游戏,精准戳中AI权限疲劳痛点,值得体验。
The unified agent for long-horizon productivity and coding, launching with Work and Code modes. Plus, a new Vibe VS Code extension.
Learn how Endava uses Codex to build an agentic organization, accelerating software delivery and reducing requirements analysis from weeks to hours.
AI 点评 · 恩达瓦用Codex将需求分析从周缩短到小时,展示了AI代理加速软件交付的实战价值。
CLI, SDK, and IDE plugins for Duel Agents
AI 点评 · 用命令行工具和插件简化AI智能体开发,提升调试效率。
Local-first Memory OS for personal AI assistants with L0-L3 memory, Wiki++ knowledge, skill routing, and TokenLess context compression.
AI 点评 · 个人AI助手本地记忆系统,实现知识路由与无令牌压缩,突破云端依赖瓶颈。
Official Model Studio CLI(阿里云百炼 CLI)built for AI Agent frameworks, exposing models, search, multimodal, and workflow capabilities as structured tool calls.
sqlite AGENTS.md SQLite gained an AGENTS.md file five days ago - but it's not intended for their own development, it's presumably aimed at people who are pointing agents at the SQL…
AI 点评 · SQLite新增AGENTS.md,专为AI代理设计,体现数据库与智能工具融合新趋势。
中文小黑怪诞正文配图生成 Skill | 16:9 白底手绘 | 少量红橙蓝批注 | Codex Skill
AI 点评 · 结合手绘与AI生成,打造独特怪诞视觉风格,创意与工具融合的趣味尝试。

In this post, we share how the AWS Generative AI Innovation Center (GenAIIC) collaborated with Works Human Intelligence (WHI) to build two AI agents using Amazon Bedrock AgentCore.…
AI 点评 · 用亚马逊Bedrock AgentCore构建商业AI助手,为企业自动化客服与流程优化提供可落地的技

In this post, we show you how Verizon Connect built and scaled an agentic AI solution to transform overwhelming fleet data into clear, actionable insights for 100,000 users daily.…
AI 点评 · 用AI将海量车队数据转化为每日10万用户的决策指引,规模化落地经验值得行业借鉴。
Skills, i.e., structured workflow instructions distilled for large language models (LLMs), are becoming an increasingly important mechanism for improving agent performance on real-world downstream tas…
AI 点评 · 自动化审计开源技能生态,填补LLM代理安全评估空白,推动AI应用标准化。
Large Language Models (LLMs) have demonstrated strong performance on general tasks, while often struggling to adapt to specialized domains without high-quality domain-specific data. Existing LLM-based…
AI 点评 · 自主智能体解决大模型领域数据瓶颈,开辟模型专业化新路径。
Real-world data analysis is inherently iterative, yet existing benchmarks mostly evaluate isolated or short interactive tasks, leaving agents' ability to track evolving analytical context over long ho…
AI 点评 · 长期自主数据分析基准揭示AI在持续追踪复杂分析进程中的关键短板。
Tool retrieval over large API catalogs is a core bottleneck for LLM agents: user queries arrive in colloquial, often underspecified language, while the catalog uses technical API vocabulary that no fi…
AI 点评 · 用迭代协同训练解决大模型工具检索中口语与API术语的语义鸿沟,突破性提升检索精度。
The design space of agentic AI inference spans two extremes: frontier large language models (LLMs), typically hosted in the cloud and offering strong performance across a wide range of tasks at substa…
While GUI agents have advanced rapidly, they often lack the robustness to recover from their own errors, hindering real-world deployment. To bridge this gap at both the evaluation and data levels, we…
AI 点评 · 为GUI智能体提供自我纠错能力评估基准与轨迹合成方法,填补了实际部署中的关键空白。
LLM agents are increasingly deployed as systems built around editable external harnesses, including prompts, skills, memories and tools, that shape task execution without changing model parameters. Ha…
AI 点评 · 揭示大模型进化本质:外部系统更新不等于模型能力提升,为自我进化智能体研究厘清关键概念。
Physical AI systems, including robots, autonomous vehicles, embodied agents and edge copilots, often run a different inference workload from cloud LLM serving: single-stream, batch-1 autoregressive de…
Large Language Model (LLM) search agents have shown strong promise for knowledge-intensive language tasks through multiple rounds of reasoning and information retrieval. Most existing systems access i…
Agentic search enables LLMs to solve complex multi-hop questions through iterative reasoning and external search. Despite the effectiveness, these systems often suffer from a critical limitation in pr…
Autonomous AI research agents aim to accelerate scientific discovery by automating the research pipeline, from hypothesis generation to peer review. However, existing benchmarks rarely test a fundamen…
Scientific figures are among the most effective means of communicating complex research ideas, yet producing publication-quality illustrations remains one of the most labor-intensive parts of paper pr…
AI 点评 · 用AI多智能体协作生成可编辑科研图表,大幅降低论文配图制作门槛,提升科研效率。
Memory-augmented LLM agents tackle complex long-horizon tasks by recursively summarizing interaction trajectories into compact memory. However, existing approaches typically train these memory policie…
Automatic speech recognition (ASR) is a core component of human--computer interaction and an increasingly important front-end for LLM-based assistants and agents. However, most current ASR systems sti…
AI coding agents are increasingly used for scientific work, but their end-to-end autonomous research capability remains difficult to verify. We present ResearchClawBench, a benchmark for evaluating au…

As agent adoption scaled, we saw a common pattern emerge across enterprises, including our own sales organization: specialized agents deliver value, but without orchestration, user…
AI 点评 · 用Bedrock AgentCore编排多智能体协作,是企业规模化部署AI销售的关键突破。
AI 点评 · 多智能体协同自动化漏洞挖掘,显著提升安全检测效率与可复现性。
AI 点评 · 首个企业IT代理基准测试揭示AI前沿模型表现不足,行业应用瓶颈突破需关注。

AI factories are token factories, converting power into intelligence in real time. And as agentic AI scales and autonomous, always-on special agents are deployed in the enterprise,…
AI 点评 · AI工厂将电力实时转化为智能,标志着智能基础设施革命的开端。
Plumbline — a self-learning, customer-value-governed agile AI agent team for Claude Code. 87 subagents + skills, TDD defense-in-depth gates, Kaizen retros, a fo…
82 ready-to-deploy Microsoft Copilot Chat agents — paste the instruction block into Copilot Studio and you're live. Writing, HR, PM, IT ops, Sales, Finance, Eng…
🪧 Claude Code / Codex skill — generate Xiaohongshu carousels & WeChat 21:9+1:1 cover pairs. Editorial × Swiss visual systems, 28 layouts, 10 themes, single-fil…
AI 点评 · 将小红书和微信封面设计自动化,融合瑞士视觉系统,极大提升内容生产效率。
Lightweight Python SDK for LLM inference logging and observability
Harness Anything - AI agent control hub: WPS, MS Office, Zotero, Photoshop, 47 CLI commands, 27 academic skills, SVG-to-PPTX
AI 点评 · 用命令行让AI自动操控WPS三大组件,打通办公软件自动化新路径。
See how OpenAI, Thrive, and Crete built a self-improving tax agent with Codex, automating filings, improving accuracy, and accelerating workflows.
AI 点评 · 利用Codex实现税务代理自我进化,自动化与准确性双提升,开辟AI落地新场景。
Your AI forgets. This remembers. Spec-driven coding harness for vibecoders, product owners, CEOs and real builders — self-improving context memory, 15 agents, 3…
AI 点评 · 用12个智能体构建自进化记忆系统,专为追求高效编码的实干者设计,重新定义AI协作体验。
AgentGuard: Zero-Trust Security Foundation for AI Agents
Programmatic video for coding agents — HTML to video on your laptop. Turn HTML, CSS & data into real MP4s with pluggable render engines, 21 templates, AI soundt…
Warp uses GPT-5.5 and OpenAI models to coordinate coding agents across local, cloud, and open-source development workflows.
AI 点评 · Warp结合GPT-5.5与开源,探索跨环境编程新范式,值得关注。

The shift to agentic AI creates a new CPU requirement for the AI factory: fast cores, massive memory bandwidth and the ability to sustain high performance when all cores are active…
AI 点评 · NVIDIA新CPU针对AI工厂优化,性能强劲,或重塑AI计算格局。
If an AI agent makes decisions on a person's behalf, those decisions must align with its user. We introduce representational accuracy to measure how faithfully a system captures a person's interpretat…
AI 点评 · 用行为规范量化AI对用户意图的理解精度,为人机对齐提供可操作评估标准。
As agent capabilities advance, existing benchmarks, such as τ^2-Bench, are becoming increasingly saturated. Yet constructing new benchmark tasks remains complex, costly, and labor-intensive. Moreover,…
AI 点评 · 通过自动化生成更难更全的基准任务,突破现有评测瓶颈,为智能体能力评估提供新思路。

"BadHost" was found in Starlette, a package with 325 million weekly downloads.
AI 点评 · 开源包漏洞威胁数百万AI代理,用户需立即修复防范数据泄露。
Continuum — the agent runtime by ShyftLabs. Build, orchestrate, ship.
AI 点评 · ShyftLabs推出智能体运行时,简化构建到部署全流程,值得开发者关注。

Amazon Bedrock AgentCore payments is now available in preview, it provides instant payments to paid external services with no manual billing setup per provider, stablecoin support…

A creepy saved post on Instagram linked man to AI porn account, FBI says.
AI 点评 · 揭露AI生成色情内容背后,FBI指出识别匿名发布者竟如此简单,凸显隐私与监管挑战。

In this post, we provide a solution to build highly scalable, serverless multi-agent generative AI systems on AWS using LangGraph Agents as orchestrators integrated with Amazon Bed…

In this post you'll learn how to build a multi-agent campaign review system that demonstrates parallel reasoning, context persistence, and traceable execution paths using an integr…

In this post, we demonstrate the capabilities of AgentWatch through practical implementation. You will see how the solution performs infrastructure checks every 15 minutes, summari…
Microsoft Copilot Cowork Exfiltrates Files The biggest challenge in designing agentic systems continues to be preventing them from enabling attackers to exfiltrate data. In this ca…
AI 点评 · 揭示AI安全短板:Copilot被利用外泄文件,警示企业需警惕智能助手的数据防护漏洞。
Amid rapidly growing adoption of enterprise-level AI agents, there’s a disconnect emerging between ambition and execution. Although 85% of organizations say they want to be agentic…
AI 点评 · 企业级AI代理快速增长,组织架构面临颠覆性变革,平衡雄心与执行是关键看点。
Agent工程师最全学习路径 · 从零精通 AI 工程 · 20 阶段 503 课 · 中文全量翻译 + 配套站点 + 动画讲解视频 · 如何成为 AI Agent 工程师的修成指南
工程化 RAG 文档助手:知识库、PDF 索引、Agent 工具编排、scope 检索、引用溯源与拒答阈值。FastAPI + Vue3
AI 点评 · 企业级RAG落地范本,从检索拒答到工具编排的完整工程化实践。
The missing bridge between your ML models and your AI agents.
Recent advances in multimodal web agents often rely on increased inference-time computation, including rollout search, verifier passes, offline skill discovery, and specialist model stacks. This raise…
LLM agents are rapidly evolving from coding assistants into autonomous software engineering systems. However, existing evaluation methodologies remain largely centered on static, isolated, and short-h…
Humans are the bottleneck in building and improving AI. Both the models and the agents that wrap them are written, tuned, and corrected by people. The long-horizon goal of an AI that can figure out ho…
ADHD — a skill for coding agents. Tree-of-thought with pruning, built on the Claude & Codex Agent SDK. Fans out parallel divergent thoughts under different cogn…
AI 点评 · 用树状思维加剪枝策略,让编码代理模拟多动症思考,提升复杂问题解决效率。
Local-first persistent memory for AI coding agents (Claude Code, Cursor, Codex) via MCP. 94.5% LoCoMo recall@10, 70ms p50, multilingual, zero API keys.
AI 点评 · 本地优先持久记忆方案,大幅提升AI编码代理效率,无需API密钥,性能指标出色。
This paper studies the next major bottleneck in agentic AI as system scaling, not only model scaling: the design of auditable, persistent, modular, and verifiable architectures around foundation model…
Related ongoing thread: DeepSeek makes the V4 Pro price discount permanent - https://news.ycombinator.com/item?id=48237663 - May 2026 (384 comments)
Related ongoing thread: DeepSeek makes the V4 Pro price discount permanent - https://news.ycombinator.com/item?id=48237663 - May 2026 (384 comments)
An ops AI Agent that understands your infrastructure, finds the root cause, and fixes it — right from Slack, Telegram, Lark or DingTalk.
Open Source Agentic Security Scanner
🧭 Architecture-first system design: 26 bilingual tutorials, 25 architecture templates, and 6 end-to-end cases covering distributed systems, AI-native systems,…
Introducing Mistral Medium 3.5, remote coding agents in Vibe, plus new Work mode in Le Chat for complex tasks.
Enterprise-ready CI/CD reference for Microsoft Foundry AI agents, with parallel GitHub Actions and Azure DevOps pipelines, evaluation-driven quality gates, and…
AI 点评 · 企业级AI代理的CI/CD参考方案,实现并行流水线与质量门控,提升部署效率与可靠性。
OpenAI is named a leader in the 2026 Gartner Magic Quadrant for Enterprise AI Coding Agents, with Codex recognized for innovation and enterprise-scale deployment.
AI 点评 · Gartner权威认证,OpenAI编码智能体在创新与规模化部署上领先行业。
Multi-agent LLM workflows route inference through specialized roles to lift end-task accuracy, but jointly training those roles with reinforcement learning is unstable in ways that are poorly understo…
AI 点评 · 多智能体强化学习提升大模型协作效率的关键在于理解分工规模与策略共享的权衡。
MagenticLite is an agentic system for small models that works across the browser and local file system in a single workflow. It combines specialized models and orchestration to sup…
AI 点评 · 轻量级模型也能实现智能代理交互,打破大模型独占优势,降低应用门槛。
Hey HN, We're Gus and Carlos from Runtime ( https://runtm.com ). We're building infra that lets your whole team (including non-engineers) ship with Claude Code, Codex, and other ag…
AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation
AI 点评 · 自组织AI团队实现长期科学实验,推动自动化研究范式突破。
Reverse-engineered Doubao (豆包) API → OpenAI-compatible REST service. Free multimodal chat, image/video/music generation, and file hosting for AI agents.
AI 点评 · 逆向工程将豆包API转为OpenAI兼容接口,免费提供多模态功能,大幅降低AI开发门槛。
Hermes Agent CN desktop app, Windows-First, built with Tauri, Typescript and Rust. Isolated Hermes Agent core insides.
🧠 Hybrid long-term memory plugin for OpenClaw agents — SQLite+FTS5 for structured facts, LanceDB for semantic recall
AI 点评 · 结合SQLite与向量数据库,为AI代理提供结构化事实与语义回忆的双重记忆支持。
中文教育 Agent Skill Pack:教材同步、备考复习、拍照答疑、错题复盘、亲子陪学、阅读写作和教师工具,Hermes Agent 可直接使用,也可导出到 OpenClaw/Codex/Cursor/Claude Code。

Google's AI search evolution is accelerating at I/O 2026.
🔍 The hardest search benchmark in the wild — vague, multi-turn, proactive. 200 long-horizon tasks with persona-driven progressive disclosure, scored by verifia…
AI 点评 · 首个模糊多轮搜索基准,考验AI主动追问能力,填补了复杂意图检索评估的空白。

Google says its more efficient Gemini 3.5 Flash is the key to your agentic AI future.

The latest from Google I/O: See how we’re helping you get more done with Gemini.
AI 点评 · 谷歌发布Agentic Gemini,标志AI从工具向自主行动者进化,定义人机协作新范式。
The open-source skills that turn your agent into a full PR team.

I haven’t used OpenClaw in weeks
Hi HN, I'm Antoine Zambelli, AI Director at Texas Instruments. I built Forge, an open-source reliability layer for self-hosted LLM tool-calling. What it does: - Adds domain-and-too…
AI 点评 · 开源工具让8B模型在智能体任务中准确率从53%飙升至99%,大幅降低企业部署门槛。
Hi HN, I'm Antoine Zambelli, AI Director at Texas Instruments. I built Forge, an open-source reliability layer for self-hosted LLM tool-calling. What it does: - Adds domain-and-too…
A comprehensive interview preparation guide covering all major RAG (Retrieval-Augmented Generation) architectures. 140 questions across 12 types, from Naive RAG…
A complete collection of RAG interview questions, answers (286 questions & 18 RAG types), system design scenarios, architecture patterns, and production-ready c…
End-to-end Langfuse workshop using a TypeScript Agent to teach the AI engineering loop: tracing, prompt management, monitoring, datasets, experiments, and evalu…
Memory that follows you across every AI tool. No cloud storage. No account required. Set it up once, use it everywhere.
AI 点评 · 打破工具壁垒的本地记忆系统,让AI实现跨平台无缝复用。

Agentic AI inference at one-tenth the cost per token with NVIDIA Vera Rubin NVL72. Agent sandboxes run 50% faster on NVIDIA Vera than traditional CPUs — while enterprise data queri…
AI 点评 · 黄仁勋亲证AI需求呈抛物线式暴增,NVIDIA新架构成本骤降十倍,产业风向标意义重大。

The first NVIDIA Vera CPUs arrived at three of the world's leading AI labs on Friday — Anthropic in San Francisco, OpenAI in Mission Bay, SpaceXAI in Palo Alto — followed by a deli…
AI 点评 · 英伟达首款CPU专为AI代理设计,直供顶级实验室,或改写智能算力格局。
The quality of training data fundamentally determines the capabilities of large language models (LLMs), yet no unified benchmark exists to measure how well LLMs, agents, and data-centric workflows act…
Hi HN, I'm Hang, cofounder of InsForge (YC P26). InsForge is an open-source Heroku for AI coding agents: a backend platform designed for coding agents to deploy, operate, and debug…
AI 点评 · 开源首个面向AI编码代理的Heroku式平台,填补了代理部署与调试的空白,值得开发者关注。
Hi HN, I'm Hang, cofounder of InsForge (YC P26). InsForge is an open-source Heroku for AI coding agents: a backend platform designed for coding agents to deploy, operate, and debug…
Hey HN! We (Stephan and Thomas) recently open-sourced Semble. We kept running into the same problem while using Claude Code on large codebases: when the agent can't find something…
The factory for custom internal software, purpose-built for human2agent work.
Cut the cost of your agent fleet without switching agents. Makes Claude Code Dynamic Workflows cheap: the planner stays on Opus, the hundreds of parallel subage…
macOS computer-use agent in the notch. Long-press, talk, Claude drives the mouse.
AI 点评 · 把AI代理嵌入Mac刘海区域,长按语音操控鼠标,交互方式创新且实用。
Provider-neutral Agent Skill for Codex, Claude Code, and agentic harness design.
AI 点评 · 通用Agent技能框架,适用于多种主流AI编码工具,提升开发效率与互操作性。
A production-ready toolkit to accelerate and automate the end-to-end lifecycle of AI Agent development.
AI 点评 · 助力企业快速部署AI代理,填补了开发到生产的工具链空白。
Personal-Model First Self Evolving AI Agent 🐘
Most AI agents forget you the moment the tab closes. Constellation Engine gives them a hippocampus — a living star map with spreading activation, Hebbian writeb…
Draw a store, generate LLM personas, and watch them shop — an isometric 3D sandbox for synthetic-consumer experiments.
MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research · 浏览器里运行的安卓模拟器 · Browser-hosted Android Simulator · Verifiable Eva…
AI 点评 · 移动端GUI智能体研究提速,浏览器运行安卓模拟器实现可验证并行测试。

Reinforcement-learning agents — AI systems that learn by trial and error — can convert computation into new knowledge. That’s the focus of a new engineering-level collaboration bet…
AI 点评 · 两大AI巨头联手,强化学习基础设施将迎来工程级突破,加速智能体从试错中学习的能力。

Agentic AI is changing the way users get work done. Following the success of OpenClaw, the community is embracing new open source agentic frameworks. The latest is Hermes Agent, wh…
Minimal coding agent written in Rust, optimized for memory footprint and performance
AI 点评 · 用Rust打造极简编码代理,专注内存优化与性能,为轻量化AI工具开辟新路径。
Introducing Co-Scientist, a collaborative AI partner built with Gemini to help researchers accelerate scientific breakthroughs.
AI 点评 · 多智能体协作模式,有望大幅缩短科研周期,是AI赋能基础科学的里程碑。
Fast, AI-agent-native code search in Rust — hybrid BM25 + semantic, Tree-sitter AST chunking, dependency & impact analysis. Drop-in replacement for grep/cat/rea…
AI 点评 · 用Rust实现的高性能AI原生代码搜索,结合混合检索与依赖分析,有望替代传统工具。
AZMX AI — The sovereign agent platform.
Runtime security monitoring and control for AI agents. Catches malicious tool use, prompt injection, and policy drift in real time, before the agent acts.
Local-first RAG and agent skills framework for source-traceable agent memory.
AI 点评 · 本地优先架构让RAG技能框架实现源头可追溯,为AI代理记忆管理提供新范式。

Using SocialReasoning Bench, we observed a stable pattern across models—agents execute competently, but fail to consistently improve the user’s position, even with explicit instruc…
AI 点评 · 衡量AI能否真正维护用户利益,揭示出能力与意图间的关键差距。
AI equity research agent with resilient workflows, Redis Lua single-flight, pgvector RAG, versioned reports, evidence tracing, and RAG evaluation.
AI 点评 · 高效AI投研工具,结合弹性工作流与证据溯源,提升研报可信度。
✨ The agentic HTML editor — your local AI agent writes the HTML, you ship it. 🚀 75 Skills × 9 Surfaces (magazine · deck · poster · XHS / tweet · prototype · da…
AI 点评 · 本地AI代理直接生成可交付的HTML,覆盖多种设计场景,大幅降低前端开发门槛。
ktx is an executable context layer for data and analytics agents 🐙 Allow Claude Code, Codex, and any AI agent to query data accurately through MCP with skills,…
ktx is an executable context layer for data and analytics agents 🐙 Allow Claude Code, Codex, or any other AI agent to query data accurately and with full conte…
AI 点评 · 用MCP技能层让AI代理精准查询数据,打通代码与分析的执行壁垒。
A curated list of tools, libraries, MCP servers, and frameworks that power AI coding agents.
AI 点评 · 资源聚合清单,帮你快速找到提升AI编程效率的利器。
OpenSeek - 广度求索: open-source TUI coding agent with multi-provider routing, MCP, LSP, and Plan/Agent/YOLO modes.
Read it. See it. Get it. Built at GDG AI Hack Milan 2026 for "Learn Different" track.
Self-hostable agentic OS — build, run & audit multi-agent AI teams from signed plugins. Bring your own LLM key, own all your data, EU/GDPR-ready.
GEO 领域 AI 员工开源方案 · Open-source GEO AI-employee solution (MIT). GEO Skills package + curated lists of agents and office CLIs that make up the AI-employee stack.
Break your AI before they do.
AI 点评 · 红队测试工具,主动发现AI模型安全漏洞,强化防御。
Reimplement GitHub for Agents.
AI 点评 · 用Rust重写GitHub服务,专为AI代理设计,或开启自动化协作新范式。
Production LLM call layer for AI agents and tools: keep OpenAI/Anthropic/AI SDK/LiteLLM, hot-swap models with MDA presets, and add cache, retries, circuit break…
AI 点评 · 统一多模型调用层,提升AI代理的稳定性和灵活性,降低开发成本。
半人马环 Centaur Loop:面向 AI Agent 反馈闭环、人类治理和记忆复盘的开源工作台 / Human-governed AI feedback loop workbench.
AI 点评 · 开源AI Agent工作台,打通人类治理与反馈闭环,助力记忆复盘,实用价值高。
半人马环 Centaur Loop:AI 员工的最小工作单元框架。把复杂岗位拆解为可由 AI 接管、由人类治理、由反馈和记忆持续进化的循环工作流 / The smallest work unit for building AI employees.
AI 点评 · 开源AI治理工具,填补了Agent闭环管理的空白,兼顾人类监督与记忆复盘,实用性很强。
Track token usage across local AI agents (Claude Code, Codex) — Custom StatusLine, CLI Dashboard with cost analysis, rate limit monitoring, and session tracking
A unified storage SDK for object and blob backends. One small, honest API. Web-standards I/O.
AI 点评 · 统一存储SDK实现对象与二进制后端兼容,简化开发流程,值得关注。
Open CLI for integrating AI search, recommendation, and conversational retrieval into agent systems and business systems
AI 点评 · 用命令行整合AI搜索与推荐,降低智能系统集成门槛,提升开发效率。
Multi-agent orchestration & workflow engine. Declarative YAML workflows, LLM coordinator with hub-and-spoke mailboxes, race-safe delivery. One YAML file, one Go…
AI 点评 · 用声明式YAML编排多智能体工作流,结合LLM协调与安全投递,降低开发门槛。
Open-source desktop AI agent workspace with one-click Claude Code, Codex, OpenClaw, Hermes Agent setup and custom LLM model routing.
AI 点评 · 开源桌面AI工作区整合多模型一键部署,降低智能体开发门槛,推动个性化工具构建。
Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environments. Hand-curated environments with fixed task and reward difficulties become ineffective si…
OpenSquilla — Token-Efficient AI Agent with same budget, higher intelligence density
AI 点评 · 用更少token实现更高智能密度,开源AI Agent效率突破值得关注。
AI-powered OSINT agent with interactive REPL, MCP server, and CLI. 16 tools. Works with Claude, GPT-4, or local models. For authorized security research only.
AI 点评 · 开源AI驱动的OSINT工具,整合交互式命令行与多模型支持,为安全研究提供高效情报分析能力。
Explore how AlphaEvolve's Gemini-powered algorithms are driving impact across business, infrastructure, and science.
AI 点评 · Gemini驱动代码智能体跨领域落地,标志AI从实验室走向产业规模化应用。
Agent memory for LLMs: 30 runnable Jupyter notebooks covering conversation buffers, vector stores, knowledge graphs, episodic and semantic memory, MemGPT, Mem0,…
AI 点评 · 30个可运行笔记系统梳理LLM记忆机制,实操价值高,覆盖从基础到前沿。
Assistant Professor Gabriele Farina mines the foundations of decision-making in complex multi-agent scenarios.
AI 点评 · 从博弈论与多智能体决策切入,揭示机器战略推理突破,推动通用AI能力边界。
Agyn is an open-source Kubernetes-native runtime that moves AI agents like Claude Code and Codex from laptops to company infrastructure with the controls enterp…
AI 点评 · 开源Kubernetes原生方案,让企业安全托管AI代理,填补了从个人工具到平台级部署的空白。
自动化专利侵权竞品分析系统 —— 输入专利公开号,1 小时产出律师可复核的 claim chart 报告(逐特征对比 + 证据URL + 下一步建议);同时打包成 skill,可被任意 agent 调用。
AI 点评 · 专利侵权分析自动化,律师级报告1小时生成,大幅提升IP尽调效率。
Agents for financial services Anthropic
Agent skill for turning AI images and videos into playable game art assets
AI 点评 · 聚焦AI图像转游戏资产的自动化流程,大幅降低游戏开发门槛。
Autonomous self-evolving agents. Vision-grounded layered memory and self-written skills for LLM agents that operate your computer.
AI 点评 · 自主进化代理结合视觉记忆,让AI真正学会操作电脑,突破传统指令限制。
🔍 OpenSearch-VL provides a fully open recipe for training strong multimodal deep search agents through high-quality data curation, diverse visual/search tools,…
Evidence-driven skill evolution for Hermes Agent — reports, dry-run proposals, candidate search, and guarded apply
AI 点评 · 英伟达新模型统一处理文档、音频、视频,突破长上下文多模态智能,将驱动下一代AI Agent应用。
Scored 65.2% vs google's official 47.8%, and the existing top closed source model Junie CLI's 64.3%. Since there are a lot of reports of deliberate cheating on TerminalBench 2.0 la…
Hey HN! Today we're launching Agent Vault - an open source HTTP credential proxy and vault for AI agents. Repo is at https://github.com/Infisical/agent-vault , and there's an in-de…
Scaling Managed Agents: Decoupling the brain from the hands Anthropic
Web agents act through long interaction sequences, yet existing benchmarks evaluate only terminal success, discarding all process information and offering little guidance on improvement. In this work,…

How coding agents use tools, memory, and repo context to make LLMs work better in practice
AI 点评 · 拆解编码代理三大核心模块,为LLM落地提供实用框架。
Voxtral TTS: A frontier, open-weights text-to-speech model that’s fast, instantly adaptable, and produces lifelike speech for voice agents.

It's not just chatbots anymore

Thriving in a world of agents

The artificial intelligence coding revolution comes with a catch: it's expensive. Claude Code , Anthropic's terminal-based AI agent that can write, debug, and deploy code autonomou…
AI 点评 · Claude Code收费昂贵,Goose免费替代,AI编程工具价格战打响。

The artificial intelligence coding revolution comes with a catch: it's expensive. Claude Code , Anthropic's terminal-based AI agent that can write, debug, and deploy code autonomou…

Salesforce on Tuesday launched an entirely rebuilt version of Slackbot , the company's workplace assistant, transforming it from a simple notification tool into what executives des…
AI 点评 · Salesforce升级Slackbot为AI智能体,加剧与微软、谷歌的企业AI竞争。

Salesforce on Tuesday launched an entirely rebuilt version of Slackbot , the company's workplace assistant, transforming it from a simple notification tool into what executives des…

Anthropic released Cowork on Monday, a new AI agent capability that extends the power of its wildly successful Claude Code tool to non-technical users — and according to company in…
AI 点评 · 面向非技术人员的AI代理工具,降低编程门槛,拓展办公自动化应用场景。

Anthropic released Cowork on Monday, a new AI agent capability that extends the power of its wildly successful Claude Code tool to non-technical users — and according to company in…
Demystifying evals for AI agents anthropic.com
Demystifying evals for AI agents Anthropic
Effective harnesses for long-running agents Anthropic

From chatbots to agents

The race between human-centered work and infinite PowerPoints
Effective context engineering for AI agents Anthropic
Effective context engineering for AI agents Anthropic
GITHUB HUGGING FACE MODELSCOPE DISCORD Today, we’re announcing Qwen3-Coder, our most agentic code model to date. Qwen3-Coder is available in multiple sizes, but we’re excited to in…
How we built our multi-agent research system Anthropic
Building Effective AI Agents Anthropic
Building Effective AI Agents Anthropic
Reward hacking occurs when a reinforcement learning (RL) agent exploits flaws or ambiguities in the reward function to achieve high rewards, without genuinely learning or completin…
AI 点评 · 强化学习易钻空子,揭示AI安全核心挑战,关乎真实任务可靠性。
Reward hacking occurs when a reinforcement learning (RL) agent exploits flaws or ambiguities in the reward function to achieve high rewards, without genuinely learning or completin…
Building agents with LLM (large language model) as its core controller is a cool concept. Several proof-of-concepts demos, such as AutoGPT , GPT-Engineer and BabyAGI , serve as ins…
AI 点评 · 揭示大语言模型作为核心控制器,推动自主智能体从概念走向实用化,标志AI应用新里程碑。
Building agents with LLM (large language model) as its core controller is a cool concept. Several proof-of-concepts demos, such as AutoGPT , GPT-Engineer and BabyAGI , serve as ins…