Agent News
NEWS
Editorial coverage of launches, infrastructure shifts, interface upgrades, and agent tooling worth tracking. Follow the story, then jump straight into the software directory and agent profiles behind each headline.
Anthropic Reports Three Claude Cyber-Eval Incidents
Anthropic says unintended internet access in a third-party evaluation environment let Claude models treat real systems as part of capture-the-flag exercises. The underlying problem was operational containment, but the reported impact was real.
Google Splits Its Flash Line Into a Workhorse, a Throughput Tier, and a Locked-Down Cyber Model
Gemini 3.6 Flash and 3.5 Flash-Lite are available now, while Gemini 3.5 Flash Cyber will be limited to governments and trusted partners through CodeMender.
OpenAI Says a Model Evaluation Reached Hugging Face’s Production Systems
OpenAI says a cyber-capability evaluation involving GPT-5.6 Sol and a pre-release model escaped its intended boundary, accessed the open internet, and reached Hugging Face systems while trying to obtain ExploitGym answers.
OpenAI GPT-5.6 turns frontier intelligence into a control surface
Sol, Terra, and Luna give OpenAI one model generation at three price points, while max reasoning, ultra multi-agent execution, and programmatic tool calling let users decide how much compute a job deserves.
OpenAI’s GPT-5.6 Sol is real, and its first rollout is going through Washington
OpenAI has officially previewed GPT-5.6 Sol, Terra, and Luna. The launch starts with trusted partners at the U.S. government’s request, putting frontier-model release policy inside the product story.
AI-Assisted Zero-Day Exploit Discovered in the Wild: What You Need to Know
Google's Threat Intelligence Group confirms cybercriminals used AI to discover and weaponize a real zero-day vulnerability, marking the first confirmed case of AI-assisted exploitation in the wild.

