“Until you make the unconscious conscious, it will direct your life, and you will call it fate.” (Attributed to Carl Jung.)
Carl Jung had a name for the part of us we do not look at: the shadow. Not the evil part - the unexamined part. The instincts, assumptions and reflexes that shape how we judge and decide without ever surfacing into conscious thought. His point was never that the shadow is dark. It was powerful precisely because it has great influence over our lives but goes unwatched.
Every organisation relying on AI for decision making now has a shadow.
It’s the layer almost nobody looks at directly: the system prompts, the retrieval logic, the agent boundaries and the standing instructions that sit between your people and the models they increasingly rely on.
I call it the AI Mediation Layer, and its defining quality is exactly Jung’s. It shapes how the organisation interprets, judges and decides, silently, at scale, below the level anyone is watching. It doesn’t appear on your architecture diagram. It doesn’t appear in your data-flow map. It doesn’t appear in your threat model. And like any shadow, it is dangerous in direct proportion to how unexamined it is.
Sit with that before going further, because everything else depends on it. The mediation layer is where your organisation’s judgement increasingly lives. Not in the heads of your people but in the instructions that now shape what the AI tells your people what to think. That is a transfer of judgement most organisations have made without realising they made it.
The asset we’ve been mispricing for thirty years
Here is the thing the security industry has had backwards since before any of us started.
Data has no intrinsic value. Data is valuable because it enables judgement. Every market analysis, every customer record, every trade secret, every piece of intelligence you hold exists for one reason: to inform a decision. Strip the decision away and the data is inert. Bytes on a disk.
We built an entire cybersecurity industry around protecting the bytes. Confidentiality, integrity, availability. Encryption at rest and in transit. Data loss prevention. Access controls. We measure breaches in records exposed, in terabytes exfiltrated, in fines incurred. All of it watches the data.
But data is the means. Judgement is the end.
Stealing data is stealing the ore. Stealing judgement is stealing the refined metal, the engineered component, the finished machine. It’s the difference between stealing the blueprints and stealing the architect’s mind.
For thirty years that distinction was academic, because judgement lived in human heads we couldn’t reach at scale. It doesn’t anymore. We have spent the last two years quietly moving judgement out of people and into the mediation layer, and we have protected that layer with almost none of the discipline we lavish on the data sitting one layer below it.
You can secure the data perfectly and still be robbed
This is the shift CISOs in AI-enabled organisations need to grasp, fast.
You can do data security flawlessly. Every byte encrypted, every record access-controlled, every export caught, zero terabytes out the door. And you can still produce systematically poisoned decisions, because in an AI-mediated organisation, sound data and sound judgement are no longer the same thing.
The objection I get every time is that this is just integrity, the ‘I’ in the triad we’ve defended for decades. It isn’t. And the distinction is the whole point of this article.
Classical integrity asks: were the bytes altered? Were records tampered with, files modified, hashes broken?
AI forces a different question: was the interpretation trustworthy?
Because in an AI-mediated system, decisions are not drawn directly from data. They are produced through a mediation process, one that selects context, prioritises instructions, and interprets intent. And that process can be steered two ways:
- By altering data: a poisoned retrieval corpus, a rewritten system prompt.
- Without altering any data at all: runtime prompt injection, adversarial framing, context manipulation.
This is the hardest part to grasp. The data does not need to be corrupted to be weaponised. The malicious document in the retrieval corpus was legitimately created, legitimately stored, properly access-controlled, and its hash is valid. Nothing was tampered with. Every data-integrity check passes green. The bytes are exactly what they claim to be; they were simply chosen, or seeded, to steer the interpretation.
So you can have perfect data integrity at rest, authentic and untampered records, every classical control green, and still produce a corrupted decision, because the interpretation process was manipulated. The conclusion is poisoned. The data it was drawn from remains untouched.
No mature control family in mainstream security architecture currently treats mediation integrity as a first-class security domain.
That is not a data integrity failure. It is a mediation integrity failure. A property that sits one level above everything the triad was built to defend. Data integrity protects the bytes. Mediation integrity protects the interpretation. The two together protect the judgement. And because the existing tooling watches the bytes, none of it fires: no anomalous login, no new process, no volume spike. Nothing is stolen.
The system is simply and consistently wrong.
The danger compounds because AI systems standardise interpretation across thousands of decisions simultaneously, removing the natural variation that once limited organisational error.
Injection is the loud door. The quiet ones are wide open.
This is also why the current discourse badly underestimates the gravity of the problem.
Almost all the attention on AI security has piled onto prompt injection, the crafted input that tricks a model into ignoring its instructions. It’s real, it tops every list, and an entire industry is forming around it. But injection is one door into the shadow. It is not the shadow. It is simply the most visible door, the one everybody’s heard of, the one with a queue of vendors outside it.
The hidden doors are the dangerous ones, precisely because nobody is watching them. A retrieval corpus seeded with authentic-looking poison. A standing instruction rewritten at the source.
This isn’t theoretical. It is exactly what happened when an autonomous agent gained write access to the 95 system prompts governing tens of thousands of consultants at one of the world’s most sophisticated firms.
Had the system prompts changed, nothing would have alerted. Why? Because a write from the application’s own service account is just normal traffic.
There is no queue of vendors at that door. No detection. No discourse. Defend only against injection and you’ve declared the problem solved when you’ve barely seen it.
Let’s make it real
Picture a mid-market Australian commercial builder with a couple of hundred staff, the kind of firm that lives or dies on winning the right tenders at the right price. They roll out an internal AI assistant for their estimators: it draws on the company’s own job history, current supplier pricing, and market feeds, and the team comes to trust it. It’s fast, it’s confident, it sounds like the firm’s best estimator on a good day.
Now suppose a competitor doesn’t try to breach them - they wouldn’t need to. They just need to manipulate the model’s judgement.
They quietly poison the external supply chain feeds or public market reports that the internal assistant ingests and treats as authoritative. They seed it with material that’s entirely real but selectively unflattering: this region’s costs are climbing, that subcontractor is a risk, those margins are thinning.
Every figure checks out. Nothing is tampered with. Because the assistant was tuned to prioritise recent market volatility and external pricing indicators over older internal margin history, the poisoned signals disproportionately influenced the recommendations.
But the assistant’s advice begins to tilt. And because they share the same AI mediation layer, every estimator relying on it receives that exact same bias, presented as independent judgement.
Two quarters later the builder is winning the thin jobs and losing the fat ones. Margins bleed. The board blames a tough market and sharper competition, because that is exactly what it looks like. Nobody examines the assistant; after all, it is “just giving us the numbers,” and the numbers are real. No intrusion or traditional breach was detected, because none occurred. The data was correct. The system was working as designed.
The shadow was simply fed wrong information, and it skewed the outcome.
A breach leaves a body. A poisoned shadow doesn’t.
The shadow doesn’t leave a body.
Here is what makes corrupted judgement more dangerous than any data breach you’ve lived through.
A breach eventually announces itself. Attackers demand a ransom. Records surface. A regulator calls. Corrupted judgement announces nothing. A prompt tilted two degrees produces advice that still reads like your own best people on a good day, flows into the quotes you send and the calls you make with customers, and shapes decisions worth more than the business can afford to get wrong.
And it traces back to nothing.
Because the output is indistinguishable from normal model variance. There is no ground-truth baseline to diff against. No alerting threshold. No forensic artefact.
It just makes you slightly, consistently, invisibly wrong, across everyone who trusts the AI. The CISO doesn’t see it and the CFO just writes the damage off as a bad quarter, a soft market, plain bad luck.
A breach costs you data you can re-secure. A poisoned shadow costs you decisions you will never know you got wrong.
Make the shadow conscious
If I stopped here this would be another entry in the genre of AI scare stories, and I have no interest in adding to that pile. So here is the turn, and it is the most important part of the piece.
Jung never stopped at “the shadow is dangerous.” His point was that the shadow only has power while it remains unconscious. If you examine it, make it visible, and integrate it, it stops running you and becomes a source of strength. The danger was never the shadow. The danger was leaving it unwatched.
- Inventory it: know where your prompts, your retrieval corpora and your agent authorities actually live, and who or what can modify them.
- Govern it: treat those instructions as the high-value assets they are (versioned, signed, attributed to a human, and isolated from your data).
- Baseline it: establish what sound judgement looks like, so drift has something to be measured against and the shadow can no longer move without being seen.
Do that, and the layer that was your largest unmonitored liability becomes a defended, observable, genuinely antifragile asset, one that gets harder to corrupt every time it is tested, because every attempt feeds the baseline and tightens the gates.
The real asset
This is why AI Mediation Security and antifragile architecture are not two ideas but one.
You cannot secure judgement on a brittle foundation. And you cannot build an antifragile organisation while leaving the judgement layer in shadow.
Why? Because an adaptive system with an unguarded mediation layer will faithfully learn and propagate whatever an attacker plants. It weaponises its own antifragility.
The system doesn’t break. It just gets rapidly, efficiently, and confidently wrong.
Each needs the other. Without integration, the system does not just fail. It improves in the wrong direction.
This is the actual stake of the discipline I have been describing. Not data, judgement. Not the ore, the architect’s mind.
The frameworks have not named it because the asset class is younger than the systems already running on it.
The organisations that recognise this first, that deliberately bring their mediation layer into the open rather than waiting for an attacker to expose it, gain a head start that is very difficult to recover.
Because mediation integrity is not a control you can bolt on later.
It is a capability that emerges over time: from observing how your organisation’s decisions are formed, establishing a baseline of what “good” judgement looks like, and detecting when that judgement begins to drift. You cannot build that baseline retroactively. You have to start before you need it.
Key Takeaway
AI does not need to breach your data to corrupt your judgement. The mediation layer is where organisational judgement now lives, and it is undefended precisely because nobody is watching it. Inventory it, govern it, baseline it, before an attacker makes it conscious for you.
In my next piece, I will get specific about how to do this in practice, starting with control family: treating system prompts as what they actually are - production code that runs your business through the judgement of everyone who relies on it.