Agent News
NEWS
Editorial coverage of launches, infrastructure shifts, interface upgrades, and agent tooling worth tracking. Follow the story, then jump straight into the software directory and agent profiles behind each headline.
SpaceXAI Launches Grok 4.6 With a Focus on Long-Running Agents
SpaceXAI has released Grok 4.6, a model aimed at long-running agents, coding, knowledge work, and interactive projects. It is available in Cursor, Grok Build, the SpaceXAI API, and selected model gateways at the same base token price as Grok 4.5.
Grok Bot Beta Gives AI Teammates Their Own Cloud Computers
SpaceXAI’s agent beta works across apps, saves reusable routines, and coordinates multiple Bots. The launch also leaves open questions about permissions, auditability, and reliability.
Anthropic Reports Three Claude Cyber-Eval Incidents
Anthropic says unintended internet access in a third-party evaluation environment let Claude models treat real systems as part of capture-the-flag exercises. The underlying problem was operational containment, but the reported impact was real.
Hermes Agent 0.19.1 Turns Ten Days of Main-Branch Churn Into a Stable Release
Hermes Agent 0.19.1 packages a large stabilization wave for Telegram, voice, desktop updates, FLUX3 delivery, and a new Buzz/Nostr gateway.
Google Splits Its Flash Line Into a Workhorse, a Throughput Tier, and a Locked-Down Cyber Model
Gemini 3.6 Flash and 3.5 Flash-Lite are available now, while Gemini 3.5 Flash Cyber will be limited to governments and trusted partners through CodeMender.
Hermes Agent 0.19.0 Targets the Wait—and the Lost Reply
Hermes Agent 0.19.0 adds a vendor-reported ~80% first-turn latency cut, durable final-response delivery, smart approvals, and profile-based gateway routing.
OpenClaw 2026.7.1 makes the Gateway the place where agent work can be seen and recovered
The project’s July release rewires its Control UI and onboarding while adding GPT-5.6 routing, deeper Codex continuity, offline mobile reading, and safer recovery when a Gateway or channel delivery goes wrong.
OpenAI is patching the GPT-5.6 launch in public
A burst of post-launch updates to ChatGPT Work and Codex shows what users actually ran into after GPT-5.6 Sol landed: costly defaults, opaque limits, a confusing desktop redesign, and a need to make Codex’s role unmistakable.
OpenAI GPT-5.6 turns frontier intelligence into a control surface
Sol, Terra, and Luna give OpenAI one model generation at three price points, while max reasoning, ultra multi-agent execution, and programmatic tool calling let users decide how much compute a job deserves.
Meta’s Muse Spark 1.1 makes the AI coding race a price fight
Meta’s Muse Spark 1.1 is not the new top frontier model, but its benchmark profile, 1M context window, and $1.25/$4.25 API pricing make it a serious value play for agentic coding workloads.
Grok 4.5 makes the frontier-model price fight real
Grok 4.5 does not beat GPT-5.6 Sol or Claude Fable 5 outright. The story is price-performance: near-frontier agent scores, $2/$6 API pricing, and token-efficiency claims that could make it a practical choice for cost-sensitive coding and tool-heavy workflows.
Hermes Agent v0.18.0 turns the agent’s judgment into the product
Hermes Agent v0.18.0, the Judgment Release, adds selectable Mixture-of-Agents presets, verification contracts, /learn, /journey, background subagents, desktop Projects, Vertex AI support, and a broad security and reliability sweep.
OpenClaw 2026.6.11 is a reliability release for the messy parts of real agent work
OpenClaw 2026.6.11 focuses on reliability fixes for chat routing, provider recovery, Gateway sessions, plugins, scheduled jobs, and release evidence rather than a single splashy new interface.
OpenClaw 2026.6.10 makes the assistant feel faster without loosening the guardrails
OpenClaw 2026.6.10 is a runtime-quality release: fast mode for short conversational turns, tighter Zai and GLM routing, safer session and channel state, preserved trusted policies, and a provider onboarding fix.
OpenClaw 2026.6.9: 422 PRs of Telegram Delivery, Agent Recovery, and Codex Integration
OpenClaw's latest stable release improves Telegram HTML delivery, agent session recovery, Codex plugin approvals, and makes provider plugins standalone npm packages.
Hermes Agent v0.17 pushes agent work beyond the terminal
Hermes Agent v0.17.0 expands the open-source agent runtime with iMessage via Photon, Raft, background subagents, image editing, dashboard profile building, automation templates, managed scope, and a broad security pass.
Loop Engineering: The Complete Guide to Building Self-Improving AI Agents
Stop prompting your coding agents one shot at a time. Here is how to design loops that prompt them for you—and when the extra complexity is worth it.
Cursor Origin Moves the AI Coding Fight From the Editor to the Git Forge
Cursor announced Origin, a Git forge and code-hosting product for teams and AI agents, with a fall 2026 waitlist. The official launch page is sparse, but the surrounding keynote and docs show the strategy: Cursor wants to control review, conflicts, merge readiness, and repository workflow for agent-generated code.
xAI Ships Grok Build 0.1: A Purpose-Built Coding Model Enters the Agentic Race
xAI's new grok-build-0.1 model is now available via API. 256K context, 100+ tok/s, and a 70.8% SWE-bench score. Here is what the benchmarks, reviews, and pricing actually say.
Hermes Agent’s Velocity Release Turns the CLI Into a Multi-Agent Workbench
Hermes Agent v0.15.0, tagged v2026.5.28, is less about one flashy feature than a broader shift: a smaller agent core, stronger Kanban orchestration, faster local recall, promptware defenses, Bitwarden secrets, and a larger plugin surface.
Claude Opus 4.8 Is Anthropic’s New Agent Benchmark, With One Clear Caveat
Anthropic’s Claude Opus 4.8 release is less about a new chat personality and more about long-running agent work: stronger SWE-bench Pro results, better tool use, 1M-token context, mid-conversation system messages, cheaper fast mode, and Claude Code dynamic workflows.
OpenClaw 2026.5.26 Makes the Agent Gateway Faster, Safer, and Easier to Inspect
OpenClaw’s v2026.5.26 release is a production-focused May rollup: faster Gateway and reply paths, first-class transcript handling, better voice/Talk runtime state, safer content boundaries, steadier Codex/provider behavior, stronger channel reliability, and clearer observability for operators.
OpenClaw v2026.5.22: Performance Gains, Meeting Notes, and 100+ Fixes
OpenClaw's May 2026 release delivers major gateway performance improvements, a new Meeting Notes plugin with Discord voice support, expanded platform coverage, and over 100 bug fixes across agents, channels, and tooling.
Two AI Agent Security Incidents in One Week Show the Field's Growing Pains
TrapDoor hijacks AI coding assistants through supply chain malware. Composio gets breached via an internal AI agent. Here's what happened and what to do.
OpenClaw v2026.5.18: Real-Time Android Voice, Typed Tool Plugins, and a Faster Gateway
OpenClaw v2026.5.18 ships real-time voice sessions on Android, a new typed tool plugin SDK, faster gateway restarts, and a redesigned Mac settings experience. Here is what changed and why it matters.
Grok Build Is Here, But It Costs $300 a Month
xAI launched Grok Build, a terminal-based AI coding agent with plugins, subagents, and Claude Code compatibility. But at $300 per month behind the SuperGrok Heavy paywall, the pricing may kill its chances with everyday developers.
OpenClaw v2026.5.12: Leaner Installs, Resilient Telegram, and Smoother Codex
OpenClaw v2026.5.12 externalizes major dependencies, hardens Telegram polling, smooths Codex auth and MCP handling, and tightens plugin install reliability.

