From riskmandate.ai, the page as fetched on 2026-09-25 · open the live page ↗Everything on this sheet is the source site's own text; the newsroom's chrome is outside it.
Okay. So now I want to write another article around the idea and the, the hypothesis, which I reckon you can find with a bit of research, quite a lot of data to support it. But the hypothesis is the reason why most or a large number of Gen AI projects and pilots they never make it or never stay in production. Because I think that's the real metric. It's not that they didn't go to production, they don't stay in production. It's because of the risk um, and you know, created by the, the delta between what you want the agent to do and the mandate. And, and so the mandate, what we want it to do, and what you can do, the reach uh, of, the, of the agent. Because I think what happens here is that In a lot of projects, when you're making the proof of value, you're making the proof of concept, you, you're living in a very curated world, right? You, you're living in a very world where you're trying to validate if it's possible. Can, you know, I think a good example was, you know, can the agent help with a loan application? Can he help with an action, you know, to help the users to process something automatically, right? So I think the first phase is always naturally, can we actually do this, right? Can we actually, um, do this in an effective way. The, I think the challenge happens when that works and they start to go into production. And again, it's a typical thing of mixing the difference between an explorer project and a villager and a town planner project and not investing as much in the villager and town planner components as much as sometimes you do on the explorer because it feels like it's the real thing. But the problem in a lot of this is also the controls and especially the way we look at The, the behavior policies, right? The agent behavior policies is what the agent can actually do, right? So, so in this example, for example, of the loan, I, I think the story that I remember was that they, they were all working fine, but I think at two o'clock in the morning, the agent approved a hundred thousand dollars of credit or some amount, and and then in the morning they pull off the agent from live, and it's not that. The agent was wrong in doing that, but I think the understanding is that they it was not something they were expecting the agent to do. And I think the problem with this is that when you then get business asking interesting questions, which is, hold on, what is the liability that we have? How many of these approvals could have happened and how fast they could have happened? It, it gets very big, very, very fast. Because the, the challenge here is that Agents, because of their nature, because of the loop, because of the speed, and because of their ability to operate, they can make lots of actions very, very, very fast, right? And the problem with that is that it suddenly, you know, the question was, hold on, could this agent have issued 500, 1,000 of this, right, in a period of an hour, or, or even five minutes, 10 minutes? And in a lot of systems, the answer would be yes, because Technically, it would be possible to do that. The agent have the speed, the capability, the systems will support it. So suddenly the liabilities grow spectacularly, which is also why the insurance comes into play, right? So, so then the logic of this is that um, I reckon that what has been happening, and again, find some examples, is that when the business has to sign this off, when the business has to say, hey, let's keep now this thing in production, and they understand the damage that could be done, with hallucinations, the damage that can be done with an agent, you know, acting a little bit outside the line. And forget about even being malicious, which is another massive problem, right? Um, but um, even just behaving normally, but, you know, being over enthusiastic or, or not fully understanding the implications or not having a wider mandate or in our case, you know, a policy that controls his behavior more accurately, then it becomes a problem. And I think this is goes to that idea that you know when we talk about the little trifecta of an agent, right? I think there's more variables that are missing because sometimes just having access to a database can be crazy dangerous. And I think some of the variables that are here is the fact that the agent can act at a speed that humans never did. The agent can make requests at the speeds that agents never, you know, humans never did. He can also be creative. He can also find ways. Um, to make a lot of requests and a lot of actions that maybe was not expected. And, and the problem is that all of that is done on the authority of who approved the agent. So now the question becomes, if you cannot control the agent, if you don't have stop mechanisms again, who can stop an agent? Who can pause it? Is there, you know, like in, in the example that we've created in our vault with the insurance, we had this concept of, you know, the agents can go a little bit above the operating limit. But, but then there's a maximum that they can do, right? So then do they lose the, you know, when is the moment they lose the license to operate? And, um, and that fundamentally is the challenge here. So my, my premise is that, you know, risk and uh, the accountability and who is going to be accountable for everything that the agent could do, and, and even, even maybe like damage to customer relationships or commitments or legal commitments. That we had many cases that when something gets sold on the website, if the website says it's sold, then there's a legal commitment to sell it, right? So, um, my premise here is that this is what's preventing a lot of projects from going live. Um, and this is again where we help. And in fact, I would argue that you want to deal with this at design stage, you want to start. moving beyond that, yes, yes, the agent can do this, but then how do you limit, how do you contain it, how do you make sure that the agent is not going to do more than 20 of these an hour, a day, you know, who, who's going to stay within a set of parameters, which are completely business parameters. So you have to have a system that takes that into account. The maximum amount of book orders that can happen in a day, the maximum amount of orders or refunds or actions that can be taken, the maximum amount of, of, of customer records that can be accessed. All of that now becomes very, very business focused. And again, the good news is we can actually use agents to help with this, like what we do with um, the agent behavior policy. But my hypothesis here, and it'd be good to back the data, is that it's risk that is actually preventing, risk and accountability that is preventing a lot of projects that are very successful at pilot stage, are very successful in deployment control environments that when either they get exposed to real world or or they they actually model what can happen the business pulls the plug because they're going hey we're not comfortable with the risk that we're signing off here