Harnessing Code Agents for Automatic Software Verification
An Experimental Design Approach to Evaluating Agentic AI's Autonomous Model Discovery
Show HN: Rowboat – Open-source, local-first alternative to Claude Desktop
Show HN: Shellular – run Claude Code, Codex, Pi from your phone
The Making of Claude Code
Agent Data Injection Attacks are Realistic Threats to AI Agents
When Claws Remember but Do Not Tell: Stealthy Memory Injection in Persistent Personal Agents
LLM-as-a-Verifier: A General-Purpose Verification Framework
ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog
Claude Design System Prompt
sqlite-utils 4.0rc2, mostly written by Claude Fable (for about $149.25)
Save Claude Code Tokens with Smart Routing
New serious vulnerabilities spiked around release of Claude Mythos Preview
Claude, please stop trying to memorize random crap
Alibaba to ban Claude Code in workplace over alleged backdoor risks, source says
A rubric-based controlled comparison of frontier language models on expert-authored clinical reasoning tasks
Automated grading of Linux/bash examinations using large language models: a four-level cognitive taxonomy approach
TestEvo-Bench: An Executable and Live Benchmark for Test and Code Co-Evolution
DenisSergeevitch/agents-best-practices:Provider-neutral Agent Skill for Codex, Claude Code, and agentic harness design.
Claude-real-video - any LLM can watch a video
Show HN: I built an open-source alternative to Claude Cowork
Show HN: Claudoro, Pomodoro timer embedded in the Claude Code statusline
AutoMem: Automated Learning of Memory as a Cognitive Skill
Claude Fable 5 Promotional Access
ZCode: Claude Code from the Makers of GLM
Claude Fable 5 available globally tomorrow
Theory of Mind and Persuasion Beyond Conversation: Assessing the Capacity of LLMs to Induce Belief States via Planning and Action
Wisdom Of The (AI) Crowd: Investigating Artificial Swarm Intelligence In Large Language Models
Calibration, Not Compilation: Detecting and Repairing Misspecified Probabilistic Programs Written by Language Models
Claude Fable 5 export control lifted
Department of Commerce has lifted export controls on Claude Fable 5 and Mythos 5
Claude Code Just Got 5x More Expensive
Claude Sonnet 5 – benchmark results
Claude Code Is Steganographically Marking Requests
Whose Side Is Your Agent On? Multi-Party Principal Loyalty in LLM Agents
TraceLab: Characterizing Coding Agent Workloads for LLM Serving
Attractor States Emerge in Multi-Turn LLM Conversations
You really shouldn't copy-paste errors into Claude Code
Semgrep: GLM 5.2 beats Claude in our Cyber Benchmarks
I used Claude Code to get a second opinion on my MRI
JimLiu/baoyu-design:Run Claude Design locally as an Agent Skill — Cursor, Claude Code & more. Produce polished UI mockups, prototypes, decks & wireframes as self-contained HTML, without claude.ai/design. Best with Opus 4.8.
Show HN: Smart model routing directly in Claude, Codex and Cursor
Where Do Models Find Happiness? Emotion Vectors in Open-Source LLMs
Do Encoders Suffice? A Systematic Comparison of Encoder and Decoder Safety Judges for LLM Adversarial Evaluation
InvestPhilBench: A Multi-Layer Dynamic Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy
Ubisoft co-founder Claude Guillemot dies in plane crash
When Helpfulness Overrides Causal Caution: Context-Dependent Suppression and Recovery in LLMs
Detecting AI Coding Agents in Open Source: A Validated Multi-Method Census of 180 Million Repositories
Reinforcement learning to improve large language model-based automated code compliance systems
Code Isn't Memory: A Structural Codebase Index Inside a Coding Agent
MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop
Evaluating LLMs for Real-World Web Vulnerability Detection
Claude Code's "extended thinking" is a summary- not authentic thinking
Claude: Elevated Error Rates for Opus 4.8, Opus 4.7, Opus 4.6, and Sonnet 4.6
Identity verification on Claude
Show HN: Pulse – Dashboard for Claude Code, approve tool calls from your phone
omnigent-ai/omnigent:Omnigent is an open-source AI agent framework and meta-harness: orchestrate Claude Code, Codex, Cursor, Pi, and custom agents — swap harnesses without rewriting, enforce policies and sandboxing, and collaborate in real time from any device.
nexu-io/html-anything:✨ The agentic HTML editor — your local AI agent writes the HTML, you ship it. 🚀 75 Skills × 9 Surfaces (magazine · deck · poster · XHS / tweet · prototype · data report · Hyperframes) 🛡️ Sandboxed preview · 📤 1-click to WeChat / X / Zhihu / HTML / PNG 🔑 Zero API key — Claude Code / Cursor / Codex / Gemini / Copilot / OpenCode / Qwen / Aider.
nexu-io/open-design:🎨 Local-first, open-source Claude Design alternative. 🖥️ Native desktop app. ⚡ 259+ Skills · ✨ 142+ Design Systems 🖼️ Web · desktop · mobile prototypes · slides · images · videos · HyperFrames 📦 Sandboxed preview · HTML/PDF/PPTX/MP4 export 🤖 Claude Code / OpenClaw / Codex / Cursor / OpenCode / Qwen / Copilot / Hermes / Kimi & 17+ CLIs.
Ubisoft co-founder Claude Guillemot has died in a plane crash