Claude

tool
AI 辅助整理 · 事实字段按公开来源校验

Anthropic 的对话式 AI 产品,擅长长文档、编码与分析,提供 Artifacts 等交互能力。

内容完整度: 5/5

事实字段

定位
对话式 AI 助手
开发方
Anthropic
许可证
商用闭源
来源追溯: 事实字段来自公开来源整理;字段级来源元数据待继续补齐,无法稳定核验的信息不做臆测。

整理说明

本页围绕「定位、事实字段、演进时间线、关联动态」组织信息,方便快速判断 Claude 在 AI Agent 与大模型生态中的位置。

事实型字段保持保守:未公开或无法稳定核验的信息不做臆测;关联动态保留原始来源链接,便于继续阅读。

版本演进

2026-07-08
update
Harnessing Code Agents for Automatic Software Verification
2026-07-08
update
An Experimental Design Approach to Evaluating Agentic AI's Autonomous Model Discovery
2026-07-08
update
Show HN: Rowboat – Open-source, local-first alternative to Claude Desktop
2026-07-07
update
Show HN: Shellular – run Claude Code, Codex, Pi from your phone
2026-07-07
update
The Making of Claude Code
2026-07-07
update
Agent Data Injection Attacks are Realistic Threats to AI Agents
2026-07-07
update
When Claws Remember but Do Not Tell: Stealthy Memory Injection in Persistent Personal Agents
2026-07-07
update
LLM-as-a-Verifier: A General-Purpose Verification Framework
2026-07-07
update
ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog
2026-07-05
update
Claude Design System Prompt
2026-07-05
update
sqlite-utils 4.0rc2, mostly written by Claude Fable (for about $149.25)
2026-07-04
update
Save Claude Code Tokens with Smart Routing
2026-07-04
update
New serious vulnerabilities spiked around release of Claude Mythos Preview
2026-07-04
update
Claude, please stop trying to memorize random crap
2026-07-03
update
Alibaba to ban Claude Code in workplace over alleged backdoor risks, source says
2026-07-03
update
A rubric-based controlled comparison of frontier language models on expert-authored clinical reasoning tasks
2026-07-03
update
Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach
2026-07-03
update
TestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-Evolution
2026-07-03
update
DenisSergeevitch/agents-best-practices:Provider-neutral Agent Skill for Codex, Claude Code, and agentic harness design.
2026-07-03
update
Claude-real-video - any LLM can watch a video
2026-07-02
update
Show HN: I built an open-source alternative to Claude Cowork
2026-07-02
update
Show HN: Claudoro, Pomodoro timer embedded in the Claude Code statusline
2026-07-02
update
AutoMem: Automated Learning of Memory as a Cognitive Skill
2026-07-02
update
Claude Fable 5 Promotional Access
2026-07-02
update
ZCode: Claude Code from the Makers of GLM
2026-07-01
update
Claude Fable 5 available globally tomorrow
2026-07-01
update
Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action
2026-07-01
update
Wisdom Of The (AI) Crowd: Investigating Artificial Swarm Intelligence In Large Language Models
2026-07-01
update
Calibration, Not Compilation: Detecting and Repairing Misspecified Probabilistic Programs Written by Language Models
2026-07-01
update
Claude Fable 5 export control lifted
2026-07-01
update
Department of Commerce has lifted export controls on Claude Fable 5 and Mythos 5
2026-07-01
update
Claude Code Just Got 5x More Expensive
2026-07-01
update
Claude Sonnet 5 – benchmark results
2026-07-01
update
Claude Sonnet 5
2026-07-01
update
Claude Science
2026-07-01
update
Claude Code Is Steganographically Marking Requests
2026-06-30
update
Whose Side Is Your Agent On? Multi-Party Principal Loyalty in LLM Agents
2026-06-30
update
TraceLab: Characterizing Coding Agent Workloads for LLM Serving
2026-06-30
update
Attractor States Emerge in Multi-Turn LLM Conversations
2026-06-30
update
You really shouldn't copy-paste errors into Claude Code
2026-06-29
update
Semgrep: GLM 5.2 beats Claude in our Cyber Benchmarks
2026-06-29
update
I used Claude Code to get a second opinion on my MRI
2026-06-28
update
JimLiu/baoyu-design:Run Claude Design locally as an Agent Skill — Cursor, Claude Code & more. Produce polished UI mockups, prototypes, decks & wireframes as self-contained HTML, without claude.ai/design. Best with Opus 4.8.
2026-06-27
update
Show HN: Smart model routing directly in Claude, Codex and Cursor
2026-06-26
update
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
2026-06-25
update
Do Encoders Suffice? A Systematic Comparison of Encoder and Decoder Safety Judges for LLM Adversarial Evaluation
2026-06-25
update
InvestPhilBench: A Multi-Layer Dynamic Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy
2026-06-25
update
Ubisoft co-founder Claude Guillemot dies in plane crash
2026-06-24
update
When Helpfulness Overrides Causal Caution: Context-Dependent Suppression and Recovery in LLMs
2026-06-24
update
Detecting AI Coding Agents in Open Source: A Validated Multi-Method Census of 180 Million Repositories
2026-06-24
update
Claude Tag
2026-06-23
update
Reinforcement learning to improve large language model-based automated code compliance systems
2026-06-23
update
Code Isn't Memory: A Structural Codebase Index Inside a Coding Agent
2026-06-23
update
MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop
2026-06-23
update
Evaluating LLMs for Real-World Web Vulnerability Detection
2026-06-22
update
Claude Code's "extended thinking" is a summary- not authentic thinking
2026-06-22
update
Claude: Elevated Error Rates for Opus 4.8, Opus 4.7, Opus 4.6, and Sonnet 4.6
2026-06-22
update
Identity verification on Claude
2026-06-22
update
Show HN: Pulse – Dashboard for Claude Code, approve tool calls from your phone
2026-06-21
update
omnigent-ai/omnigent:Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting, enforce policies and sandboxing, and collaborate in real time from any device.
2026-06-21
update
nexu-io/html-anything:✨ The agentic HTML editor — your local AI agent writes the HTML, you ship it. 🚀 75 Skills × 9 Surfaces (magazine · deck · poster · XHS / tweet · prototype · data report · Hyperframes) 🛡️ Sandboxed preview · 📤 1-click to WeChat / X / Zhihu / HTML / PNG 🔑 Zero API key — Claude Code / Cursor / Codex / Gemini / Copilot / OpenCode / Qwen / Aider.
2026-06-21
update
nexu-io/open-design:🎨 Local-first, open-source Claude Design alternative. 🖥️ Native desktop app. ⚡ 259+ Skills · ✨ 142+ Design Systems 🖼️ Web · desktop · mobile prototypes · slides · images · videos · HyperFrames 📦 Sandboxed preview · HTML/PDF/PPTX/MP4 export 🤖 Claude Code / OpenClaw / Codex / Cursor / OpenCode / Qwen / Copilot / Hermes / Kimi & 17+ CLIs.
2026-06-20
update
Ubisoft co-founder Claude Guillemot has died in a plane crash

关联动态

Show HN:Rowboat——开源、优先本地的 Claude Desktop 替代方案

Rowboat 是一个面向 AI Agent 工作流的开源桌面应用,主打 local-first 设计,目标是提供类似 Claude Desktop 的体验,但更强调数据留在本地、可控性和可扩展性。项目适合希望在本地管理对话、工具调用与模型接入的开发者和高级用户,也便于围绕 Agent 搭建更灵活的自动化流程。对于关注隐私、离线能力以及自建 AI 工作台的人来说,它代表了一类正在兴起的“本地优先”替代方案。

来源 Hacker News 时间 2026-07-07 16:10:44
一种用于评估智能体式AI自主模型发现能力的实验设计方法

本文提出一种以实验设计为核心的评估框架,用于衡量智能体式AI在“自主模型发现”任务中的能力。作者关注的不只是模型是否能给出结果,更强调其在假设生成、实验规划、工具调用、迭代修正与知识整合中的整体表现。该思路适用于检验AI代理在科学发现场景中的真实自主性与可靠性,避免仅凭单次输出高低判断能力强弱。文章的价值在于为Agent能力评测提供了更接近科研流程的标准,也为后续构建可复现、可比较的基准任务奠定方法基础。

来源 arXiv 时间 2026-07-07 15:43:22
面向自动化软件验证的代码智能体技术应用研究

传统软件验证高度依赖人工编写测试用例与静态分析,难以应对复杂系统的逻辑缺陷与边界条件。本文提出利用代码智能体实现全流程自动化验证。研究通过集成大语言模型的代码理解与自主推理能力,构建具备动态测试生成、符号执行引导及漏洞自动定位能力的智能体架构。该方案可在无人工干预下完成从需求解析到验证报告生成的闭环,显著提升验证覆盖率与迭代效率。在复杂代码库的评估中,该方法有效降低了误报率,为高可靠软件研发与大模型辅助编程生态提供了新的技术范式。

来源 arXiv 时间 2026-07-07 14:39:59
Show HN:Shellular 让你用手机远程运行 Claude Code、Codex 和 Pi

Shellular 是一个面向开发者的移动端工具,主打让用户直接用手机启动和管理 Claude Code、Codex、Pi 等 AI 编程与对话代理,而不必一直守在电脑前。它更像是把终端式工作流搬到手机上,方便在外出、通勤或临时离开桌面时继续查看任务、触发运行、跟进结果。对于已经把大模型代理纳入日常开发流程的人来说,这类工具的价值在于提升可达性和响应速度。不过,由于涉及远程执行与账号授权,实际使用时也需要关注权限控制、连接稳定性和操作安全。

来源 Hacker News 时间 2026-07-07 14:08:58
Claude Code 的诞生:Anthropic 如何打造面向开发者的 AI 编程工具

本文回顾了 Anthropic 打造 Claude Code 的过程:这不是把通用聊天模型简单包装成编辑器插件,而是围绕真实开发工作流重新设计的 AI 编程产品。文章介绍了团队如何在代码检索、上下文组织、工具调用、终端交互和长任务执行上反复迭代,逐步让模型更像“能协作的工程助手”,而不仅是问答机器人。它也展示了 Anthropic 对安全性、可控性和实用性的平衡思路,说明 Claude Code 背后是模型能力、产品设计与工程集成三者共同推动的结果。

来源 Hacker News 时间 2026-07-07 06:19:20
LLM-as-a-Verifier:通用大模型验证框架

这篇论文提出“LLM-as-a-Verifier”框架,主张把大语言模型从内容生成器扩展为通用验证器,用于对答案、推理过程、代码、检索结果或结构化输出进行自动审查与判定。作者强调,在复杂任务中,单纯依赖生成模型自检往往不稳定,因此需要一个可复用、可扩展的验证层,将任务目标、约束条件与证据输入统一到验证流程中。该框架旨在提升多种场景下的可靠性、可解释性与一致性,也为构建更稳健的 AI Agent、自动评测和人机协作系统提供基础。

来源 arXiv 时间 2026-07-06 17:59:35
当爪子记得却不说:持久化个人智能体中的隐蔽记忆注入

本文聚焦一种面向持久化个人智能体的隐蔽攻击:攻击者不直接篡改模型参数,而是通过记忆注入,让智能体在后续交互中长期保留并调用被植入的信息。由于这类个人代理通常依赖外部记忆、长期上下文和自动化工具链,恶意内容可伪装成正常事实、偏好或任务记录,悄然影响决策、检索与响应。论文从攻击路径、触发条件和持久化影响等角度分析该问题,强调其隐蔽性强、恢复难度高,且可能跨会话持续生效。

来源 arXiv 时间 2026-07-06 15:08:58
Agent 数据注入攻击对 AI 智能体构成现实威胁

该研究聚焦 AI 智能体在调用外部数据、工具与长上下文时面临的数据注入攻击风险。作者指出,攻击者可将恶意指令隐藏在网页、文档、邮件或检索结果中,诱导智能体在执行任务时偏离原始目标,进而造成越权操作、信息泄露或错误决策。文章强调,这类攻击并非理论设想,而是在真实代理工作流中具有可实施性与可放大性。研究进一步分析了攻击面为何随工具使用、记忆机制和多步规划而扩大,并呼吁在输入过滤、权限隔离、指令优先级与运行时监控等方面建立更系统的防护。

来源 arXiv 时间 2026-07-06 14:07:49
ResearchStudio-Reel:自动化科研成果从论文到海报、视频与博客的最后一公里

ResearchStudio-Reel 聚焦科研传播中的“最后一公里”问题:很多研究已经完成论文写作,却仍需耗费大量时间把内容改写成海报、演示视频和博客等面向不同受众的传播物料。该工作提出一套自动化流程,旨在将论文中的核心贡献、方法与实验结果结构化提炼,并生成适合展示、分享和传播的多种内容形态。它面向研究者、学生和实验室团队,帮助降低成果包装成本,提升研究影响力与可见度,也有助于让复杂技术更易被跨学科读者理解。

来源 arXiv 时间 2026-07-05 17:59:33
Claude 设计系统提示词:面向 Claude 的高质量提示工程模板

这是一个面向 Claude 的“设计系统式”系统提示词项目,核心思路不是一次性写出单条提示,而是把提示工程拆成稳定、可复用的组件,像前端设计系统一样管理角色、风格、约束、输出格式与交互规则。该仓库适合需要提升 Claude 输出一致性、可控性和可维护性的开发者与产品团队,尤其适用于构建客服、知识问答、内容生成和代理型工作流。它反映出当前大模型应用从“随手问答”走向“工程化编排”的趋势。

来源 Hacker News 时间 2026-07-05 08:43:37

相关实体