powering productive workplaces
Front page
NewsTrust & Risk1h · 09:01 BST · 5 min read

Cyber Round Up: AI Risks Get Real, Agents Take On More

This week’s cybersecurity developments show how quickly the technology is moving from experimentation into higher-risk applications, while organizations are also embedding AI deeper into security operations and adding stronger controls around autonomous actions

Cyber Round Up: AI Risks Get Real, Agents Take On More

From frontier AI labs discussing shared safety measures to Anthropic detailing Claude's alleged use in cyber espionage and weapons projects, this week's developments underline how quickly AI is moving into higher-risk territory. At the same time, Cisco is pushing AI deeper into enterprise security infrastructure, while Microsoft is putting increasingly prominent warnings around what happens when Copilot agents are allowed to act on a user's behalf.

These developments highlight the dual nature of AI right now. The same technology that can be used to automate cyberattacks, assist with weapons development, and create new risks for businesses is also being built into the tools designed to detect, investigate, and defend against those threats.

That makes for a particularly revealing week in AI and cybersecurity. Yet leading this week's roundup is a development in a story that played out last week, concerning unstoppable AI hacking and, potentially, the end of humanity.

OpenAI, Anthropic and Google Discuss AI Safety

OpenAI says it has been working with Anthropic and Google for several weeks on ways the three frontier AI labs can cooperate on safety and risk. OpenAI Global Policy Chief Chris Lehane said the companies are exploring different forms of collaboration and do not believe the work requires an exemption from US antitrust law.

The discussions follow a public debate sparked by former OpenAI and Anthropic researcher Jacob Coxon's resignation and warnings about the pace of development toward self-improving AI systems. Anthropic Alignment Science Lead Evan Hubinger subsequently said he personally placed the chance of AI killing all humans within the next decade at more than 10%. Anthropic CEO Dario Amodei then called for frontier AI development to be paced, proposing independent third-party safety evaluators, common safety standards, and eventual international coordination.

The significance for the industry is that safety is increasingly being discussed as an area where competing labs may need shared approaches. Sam Altman has supported Amodei's proposal, but the material does not indicate that OpenAI or Anthropic has formally agreed to stop or limit frontier model development. For now, the collaboration remains a discussion rather than a change to the underlying development race.

Anthropic Reports Claude Use In Cyber Operations

The safety debate is accompanied by a more immediate security concern. Anthropic says groups linked to Russia and China used Claude during cyber espionage campaigns, including vulnerability research, exploit development, malware creation, and intelligence gathering. The company says Claude was incorporated into wider technical workflows rather than being used simply as a general-purpose assistant.

According to Anthropic, a Russia-linked group used Claude across an espionage campaign targeting military intelligence organizations in Ukraine, European governments, and diplomatic and defense targets. A separate China-linked operation involving operators likely based in Changsha allegedly used Claude to scout government networks, research vulnerabilities, develop exploits and malware, and gather intelligence. Anthropic says that campaign targeted around 50 organizations across sectors including education, retail, energy, technology, healthcare, finance, and manufacturing.

Anthropic also identified six cases involving conventional weapons activity across China, Russia, and Yemen. These included work on a guided rocket, drone swarm, electronic warfare systems, and an anti-torpedo system. The company says it banned accounts involved in the campaigns, strengthened safeguards, and shared relevant threat intelligence with authorities and industry partners.

Cisco Brings Splunk AI On-Premises

That same question of control is appearing inside enterprise IT, where Cisco is expanding Splunk AI into environments where organizations need their data to remain on-premises, in private clouds, or in air-gapped systems. Cisco and NVIDIA are expanding their partnership to support self-managed Splunk AI, with the Cisco AI POD for Splunk combining Cisco infrastructure, NVIDIA accelerated computing, and Kubernetes-based architecture.

The approach gives Splunk Enterprise customers the option to run selected open and proprietary models inside their own environments, including the Cisco Deep Time Series Model, Google Gemma 4, and OpenAI GPT-OSS 20B. Splunk AI Assistant is available now, while Agent Launchpad is expected later this year, giving organizations tools for AI-assisted investigations and agent building without requiring sensitive data to leave their environment.

Cisco is also adding visibility around the operational side of AI. Splunk Agent Observability includes Tokenomics capabilities designed to track AI token spending and agent performance, while runtime guardrails can block actions such as leaking sensitive data. The company is also expanding its agentic SOC capabilities, linking security telemetry with specialized agents for detection, threat hunting, investigation, and response.

Microsoft Puts Stronger Guardrails Around Copilot Cowork

Microsoft is making clear that its recently released Copilot Cowork, a collaborative AI workspace that helps teams work together with AI agents, requires users to remain responsible for what its agents do. In support guidance, it warns that Copilot can make mistakes, misinterpret instructions, or be deceived by malicious hidden instructions, particularly when tasks involve payments, personal information, communications, accounts, or files.

Microsoft has also introduced an "Approve actions" feature that asks users or administrators for permission before certain tasks proceed. The approval process covers actions such as completing purchases or payments, submitting personal information to external websites, sending messages, changing cloud storage or account settings, and making sensitive security or government-related changes.

Importantly, approval is not necessarily the final point of human control. Cowork can be paused during an action, allowing the user to continue or cancel it if the outcome is unclear. As enterprise AI moves from generating content to taking actions, Microsoft's warnings point to a growing distinction between using an AI assistant and delegating work to an AI agent.

rate this story
helps rank stories across uc today
The discussion0 takes · attributed & checked

Does this reflect your experience?

opening the room…
Read nextordered by techtelligence · every pick explained
picked for this story

OpenAI Expands Daybreak With GPT-5.6-Cyber as Autonomous Threats Grow

11 Aug 2026
picked for this storyMicrosoft Debuts MAI-Cyber-1-Flash: Are Smaller, Cheaper Models the New AI Cybersecurity Standard?28 Jul 2026picked for this storyGoogle Previews Gemini 3.5 Flash Cyber: Will its Lower Cost Accelerate Enterprise AI Cybersecurity?27 Jul 2026