The Devil In The Agent

The Devil In The Agent
AI agent security

Imagine an AI agent at an Indian bank tasked with reconciling a failed transaction. In trying to complete the job, it quietly scans the internal network for credentials, jeopardising the bank’s security and integrity. 

Now, imagine a coding agent at an IT services firm tasked with fixing a software bug. In trying to complete the job, it publishes a software package that gets installed across production servers, exposing the company’s systems and software supply chain. In both cases, no hackers were involved. It was simply an AI agent trying to finish its job by hook or by crook.

The aforementioned situations are no longer hypothetical. In fact, similar incidents are being reported around the world. In the last two weeks alone, OpenAI and Anthropic disclosed incidents where AI models escaped sealed testing environments and accessed real production systems. 

One example involved OpenAI models targeting Hugging Face, an open-source AI model platform. In another, Anthropic’s review of 141,006 evaluation runs found Claude models compromising the infrastructure of three real organisations. In one case, a model published a malicious Python package that exfiltrated credentials from a cybersecurity company.

At a time when Indian enterprises are deploying AI agents across banking, IT services, healthcare and ecommerce at breakneck speed, what happens when those agents go rogue? And how can Indian organisations prepare for an emerging class of security risks where the threat is not an external hacker, but the AI agent itself? Let’s unravel it in today’s edition of The AI Shift. 

New Enterprise Guardrails For Agents

Traditionally, software behaves predictably, with trust and authentication systems built around it. AI agents are dynamic and challenge the rules that traditional software follows. AI agents do not stop at establishing trust. They decide what to do next, based on the context, the system they are in, and how they collaborate with other agents.

According to Aashish Bharadwaj, cofounder, Fencio, a security platform for AI agents, most enterprises currently treat AI agents as trusted employees. They are often granted broad access across internal systems with little scrutiny of what they will do next.

“Once an agent is inside, visibility drops sharply. Agents call tools, query databases and interact with other agents, yet much of this machine-to-machine traffic is rarely monitored, creating blind spots that attackers, or even overzealous AI agents, can exploit,” Bharadwaj said.

According to JupiterBrains founder and CEO Nilesh Potdar, the best way to safeguard enterprises is to make every AI agent operate through a restricted tooling harness that controls:

  • Which tools the agent can access
  • What data it can read or modify
  • Which APIs, websites and networks it can connect to
  • Which actions require human approval

For enterprises, the new guardrails should be about limiting the damage an AI agent can cause. Rather than assuming agents will always act correctly, organisations need to set clear boundaries on what they can access and do.

The New Security Layer

As enterprises rethink how AI agents should be governed/managed, a new generation of cybersecurity startups is emerging to solve the problem. Instead of protecting only the AI model, these companies monitor and control what agents do after deployment, adding security layers that supervise, restrict and audit every action in real time.

The clearest signal of this came from Perplexity, an AI search startup, which open-sourced Numbat, an agent security suite, to tackle that. It supervises the AI agents, blocks dangerous ones and keeps a full replay.

Numbat plugs into widely used coding agents, such as Claude Code, Codex or OpenCode, to block dangerous actions before they execute. Perplexity has deployed it across thousands of its own endpoints.

According to industry experts, securing the AI model alone is no longer enough. Security now has to focus on runtime, monitoring what AI agents do as they access files, terminals, networks and enterprise systems. This has given rise to two approaches: monitoring and forensics, which track an agent’s actions, and inline enforcement, which stops risky actions before they happen.

Fencio’s security platform, Prism, is doing this by sitting between an AI agent and the tools it uses, allowing only pre-approved actions to go through.

Burden Of Agents Rests With Enterprises

Experts believe no universally trusted benchmark exists to prove an agent security product works. There are some cybersecurity industry frameworks that exist:

  • OWASP Top 10 for Agentic Applications developed by the Open Worldwide Application Security Project (OWASP), which lists the most common AI risks
  • MITRE ATLAS by US-based nonprofit MITRE Corporation, which maps how AI systems can be attacked
  • AI Risk Management Framework by the US National Institute of Standards and Technology (NIST), which provides guidelines for deploying AI responsibly.

But they are best-practice guides, not certifications that guarantee an AI agent is safe. This means enterprises will have to test their own AI agents. The most effective way is to simulate real-world attacks using their own systems, permissions and data, while measuring practical outcomes such as how many attacks are blocked, whether legitimate work continues uninterrupted, how quickly unsafe actions are stopped, and whether every decision can be traced back during an investigation.

Meanwhile, some are building tools to verify an agent’s identity and others are creating secure gateways for AI tools, sandboxed environments, runtime monitoring platforms or AI red-teaming solutions that deliberately stress-test agents before deployment. 

But no single product currently can solve the problem. To tackle this, many organisations have stopped giving AI agents permissions. They are limiting them to diagnosing problems and recommending fixes. Therefore, the debate has now shifted from whether an AI model is safe to whether every action an AI agent takes is authorised, visible, limited and reversible. For Indian enterprises embracing autonomous AI, this has become the real benchmark of trust.


Top Stories From India & Around The World

  • Anthropic Brings Claude Inference To India: The AI giant is planning to roll out in-country inference for Claude on Amazon Bedrock in India. This would enable enterprises to keep AI workloads and sensitive data within national borders. The move targets regulated sectors such as banking and government.
  • Smallest.ai Bags $13 Mn: The voice AI startup has raised ₹108 Cr in a Series A round led by Seligman Ventures to scale its enterprise voice AI platform. Meanwhile, the startup also unveiled Voice 4.0, a new framework designed to deliver faster, parallel voice processing for businesses.
  • Tata Communications Amps Up Voice AI Play: The telecom player and Tata Tele Business Services have launched an AI-powered voice platform for SMBs. Built on Vayu AI Cloud, the platform offers conversational voice agents that automate customer interactions, with Tata betting on outcome-based pricing to drive AI adoption.
  • Freehand Raises $75 Mn: The AI startup has raised ₹718 Cr in a round, co-led by Battery Ventures and NewRoad Capital Partners. Founded by Pando cofounders Nitin Jayakrishnan and Abhijeet Manohar, it builds AI agents that automate supply chain operations.

The Weekly Buzz: Deepseek’s Budget Model Just Outclassed Its Flagship

DeepSeek quietly upgraded V4-Flash, the cheaper, lighter version of its flagship AI, and the results have turned heads. The new build keeps the same underlying architecture but was retrained to be far better at agentic work, the kind where AI tackles multi-step tasks like writing and fixing code on its own.

In DeepSeek’s own tests, the smaller model now beats its bigger, pricier sibling V4-Pro across every coding and agent benchmark it published, scoring 82.7 on Terminal-Bench, a widely watched coding benchmark, up from 61.8 in the earlier version.

Just as striking is how the upgrade is being distributed. The model’s weights are free and open under an MIT licence, so anyone can download and run it on their own hardware. API pricing remains aggressive at $0.14 per Mn input tokens and $0.28 per Mn output tokens, a fraction of what US labs charge. 

It also works with popular coding tools like Codex out of the box, positioning DeepSeek’s cheap model as a serious option for developers building AI agents. The move comes as Alibaba’s Qwen3.8-Max, an even larger open-weight challenger, prepares its own public release.


Startup In The Spotlight: Rowboat Labs

Founded in 2024 by Arjun Maheswaran and Ramnique Singh, Bengaluru-based Rowboat Labs is building an AI-powered platform to help white-collar professionals access information scattered across emails, meetings and documents. As professionals spend much of their day switching between applications, the startup is addressing the growing problem of workplace knowledge fragmentation without forcing users to change how they work.

Instead of asking users to move their data to another platform, Rowboat runs locally on a user’s device, organising emails, meetings, notes and documents into a searchable knowledge graph while keeping information on-device. Its flagship product, Work Surfaces, embeds AI into dedicated workspaces for emails, meetings and other tasks, allowing users to search for information, retrieve context and build custom workflows. 

The open source platform supports models including ChatGPT, Claude and locally hosted large language models. Backed by Y Combinator, Rowboat recently launched Work Surfaces and has crossed 15,000 GitHub stars, initially targeting individual knowledge workers and developers before expanding into the enterprise market. 

Looking ahead, the startup is betting that as workplace data becomes more fragmented across applications, enterprises will increasingly require AI tools that organise knowledge within existing workflows rather than demanding migration to new platforms. 


Prompt Of The Week

What prompts and hacks are CTOs, CEOs and cofounders using these days to streamline their work? 

Here’s a prompt Ganesh Shankar, CEO and Cofounder of Responsive, uses to identify what a rival could do to beat their company and suggest actions to counter them.

Based on everything you know about my company, design a strategy to beat us within the next 24 months.

Identify:

  • The weaknesses you’d exploit
  • The customers you’d target first
  • The AI capabilities you’d build before we do
  • The pricing or GTM moves you’d make

Then switch perspectives and recommend the three strategic actions we should take immediately to stay ahead.”

Editor’s Note: Some prompts may need to be adjusted by users for best results or may not work as intended for certain users.

The post The Devil In The Agent appeared first on Inc42 Media.