By Brian French, Former Institutional Investment Manager and Bank Officer
The Mosaic, the PowerPoint, and the Promise That Isn’t Physics
Topic Overview : A deep-dive analysis of how a corporate bank inadvertently discloses confidential information to AI — why work-from-home employees using personal ChatGPT accounts (shadow AI) are the biggest risk of all, plus the mosaic effect, presentation-upload leaks, why “we don’t train on your data” is a promise rather than physics, and every major AI breach vector for banks.
Quick Answer: How Does a Bank Leak Confidential Information to AI Without Anyone Breaking a Rule?
The single biggest way a bank leaks data to AI is the simplest: shadow AI — employees working from home on personal devices, typing bank secrets into personal ChatGPT or Gemini accounts that sit outside every corporate safeguard and, on the consumer tier, may feed the data straight into model training. Beyond that dominant risk, disclosure compounds through the mosaic effect (hundreds of harmless prompts plus shared-calendar metadata reconstructing deals and layoffs), high-density document uploads (“improve my presentation” hands over the whole picture at once), adversarial extraction (“find the weaknesses in my pitch” becomes an acquirer’s playbook), and infrastructure exposure (logs, agent pipelines, vendor chains, AI notetakers, and prompt-injection attacks like EchoLeak). A separate offensive threat inverts the problem entirely: data poisoning and “LLM grooming,” where an adversary floods the AI ecosystem with fabricated claims about the bank so chatbots repeat the lie as fact. The root cause across all of it: cloud AI confidentiality rests on contractual promises, not physical impossibility — and history shows institutions under pressure choose the short-term payoff with depressing regularity.
The Bank That Never Meant to Tell Anyone Anything
No one at the bank did anything wrong.
The credit analyst asked a chatbot to explain a covenant structure. The VP of corporate development asked for help polishing a presentation. The HR business partner asked about severance norms in three states. The syndications desk asked an AI to summarize a term sheet. The junior associate — brilliant, exhausted, and up against a 6 a.m. deadline — uploaded the whole pitch deck and typed: “Give me tips on how to improve this.”
Five hundred employees. Five hundred reasonable requests. Zero policy violations, if the bank’s enterprise AI agreement is in place. And yet, at the end of the quarter, a complete picture of the bank’s M&A pipeline, credit exposure concentrations, restructuring plans, and negotiating posture exists — assembled, structured, and stored — on infrastructure the bank does not own, cannot inspect, and could not subpoena faster than a plaintiff’s lawyer could.
This is not a hypothetical dreamed up by AI skeptics. It is why JPMorgan restricted ChatGPT in February 2023, why Deutsche Bank disabled access entirely, why Bank of America put it on the same unauthorized-apps list as WhatsApp, and why Citigroup, Goldman Sachs, and Wells Fargo followed within days. BNY Mellon blocked public LLMs citing the impossibility of meeting fiduciary data-handling requirements with third-party training pipelines. The banks understood something in 2023 that most mid-sized institutions still haven’t internalized in 2026: the danger isn’t the secret you decide to share. It’s the secret you disclose without ever deciding anything.
This article dissects exactly how that happens — mechanism by mechanism — and then confronts the uncomfortable foundation underneath it all: that every safeguard between a bank’s data and an AI provider’s ambitions is a promise, not a law of physics. But it starts with the one risk that dwarfs all the others in the real world — not a clever hacker, but a tired employee on the couch.
The Real Front Door: The Work-From-Home Employee and the Personal Account
Every sophisticated attack in this article — the mosaic assembled from logs, the zero-click prompt injection, the poisoned training data — is real. Security teams should understand all of them. But if a bank fixes only one thing, it should be this, because it is how the overwhelming majority of confidential data actually leaves the building: an employee working from home, on a personal device, typing bank secrets into a personal AI account.
Picture the most ordinary evening imaginable. A commercial-lending analyst is finishing a credit memo at 9 p.m. on her own laptop, at her kitchen table, on home Wi-Fi. Her work laptop — the one with the monitoring software, the blocked websites, the corporate AI account — is shut in her bag. She’s logged into her personal ChatGPT, the same one she uses for recipes and vacation planning. She pastes in the borrower’s financials and types, “tighten this up and flag any weaknesses in the credit.” Thirty seconds later she has a cleaner memo and a better analysis. She feels efficient. She feels like a good employee. She has just exported a client’s confidential financial condition to a consumer AI account, and nothing at the bank recorded that it happened.
This is shadow AI, and it is not a fringe behavior by rule-breakers. Survey after survey finds that a majority of knowledge workers use AI tools their employer never sanctioned, and a large share admit to feeding work information — including sensitive information — into them. The people most likely to do it are often the best employees: the ambitious, overloaded ones handling the most consequential deals, which means the bank’s highest-value data flows through its least-controlled channel. The original 2023 leak that set off the entire corporate-ban wave was exactly this — Samsung engineers pasting internal source code into ChatGPT to save time, three separate incidents in twenty days.
Why working from home makes it so much worse
The kitchen table defeats, in one stroke, nearly every defense the bank has built:
The enterprise contract doesn’t apply. The bank’s carefully negotiated protection — the no-training clause, the zero-retention deal, the whole legal basis for “our data is safe with this vendor” — covers the corporate account only. On a personal account, the employee is almost always on the consumer tier, where the default setting typically allows conversations to be used to train future models unless the user has hunted through the settings menu to opt out. Almost no one does. So the exact safeguard the bank paid for is simply not present, and the data doesn’t just sit in a 30-day log — it becomes eligible to be absorbed into the next model’s permanent memory, the one leakage layer that can never be undone.
The monitoring can’t see it. DLP tools, firewalls, and network logging live on the corporate network and corporate devices. A personal phone or laptop on home Wi-Fi is entirely outside that perimeter. The bank has no visibility, no alert, no record. This is why shadow-AI incidents take, on average, 247 days to detect and cost roughly $670,000 more than a standard breach: you cannot investigate, contain, or even notice an exfiltration through a door you don’t know exists. The bank’s first indication of the leak is frequently no indication at all.
The employee doesn’t feel like they’re leaking. No file was emailed to a competitor. No document was printed and carried out. It felt like using a smarter spell-check. The mental model of “leaking confidential data” is a villain in a parking garage, not a diligent analyst using the same tool her teenager uses for homework. The behavior generates no guilt, so no policy reminder deters it.
The trap: banning AI makes this exact risk worse
Here is the counterintuitive heart of it, and the point every bank board should sit with. When an institution reacts to this fear by simply banning AI, employees do not stop using it. The productivity pull is overwhelming and the deadline is real. They just move it further into the shadows — off the corporate laptop entirely and onto the personal phone, where the bank has zero visibility and zero contractual protection. A ban does not eliminate the risk; it converts a governable risk into an invisible one. It takes the analyst who might have used a monitored corporate tool and pushes her to the kitchen-table consumer account instead.
This is why the banks that handled AI best did not ban it — they replaced it. Morgan Stanley’s “build, don’t block” approach gave employees a sanctioned internal AI tool good enough that no one needed the leaky public one. The only durable way to stop people using the risky option is to hand them a safe option that is just as fast and convenient. The block only works if it comes with a better door standing open right beside it — an approved tool that lets the work-from-home analyst get her cleaner memo without the client’s financials ever leaving the bank’s control. For an institution too small to build that itself, that safe internal option is precisely the locally hosted, bank-controlled AI system that keeps every prompt on hardware the bank owns.
Everything that follows — the mosaic, the presentation leak, the poisoned models — is worth understanding. But it all rides on top of this: the data mostly walks out the front door, in the evening, in the hands of a good employee who thought she was doing her job well.
Part 1: The Mosaic Method — Death by Five Hundred Harmless Prompts
How intelligence agencies think about disclosure (and banks don’t)
Intelligence professionals have a name for what banks are doing to themselves: the mosaic effect. No single tile reveals the picture. The picture emerges only when tiles accumulate in one place. Classification regimes exist precisely because analysts learned that an adversary with enough unclassified fragments can reconstruct classified conclusions.
Now map that onto a corporate bank’s daily AI usage:
- A leveraged-finance associate asks the AI to “stress test debt service coverage at 9% rates for a borrower in specialty chemicals.”
- A workout officer asks how Article 9 foreclosure works in Texas “for a $40M equipment-backed facility.”
- A compliance analyst asks the AI to summarize SAR filing thresholds “for repeated structuring just under $10,000 through our Brownsville branch.”
- Corporate development asks for “integration cost benchmarks for acquiring a bank with 45 branches and $8B in assets.”
- HR asks for “communication templates for a reduction affecting a commercial lending team.”
- The CFO’s office asks the AI to “sanity check” a liquidity coverage calculation with real numbers, lightly disguised.
Each prompt passes any reasonable data-loss-prevention filter, because DLP inspects messages one at a time and the sensitivity here is emergent — it exists only in combination. Assembled, those six prompts reveal: a troubled chemicals credit, a specific workout in progress, a potential BSA/AML problem at a named location, an acquisition target profile precise enough to shortlist candidates, an impending layoff in a specific department, and the bank’s actual liquidity position. That is a mosaic worth millions to a competitor, an activist investor, a short seller, or a litigant — and no individual employee had the authority to disclose it, because no individual employee did disclose it.
Where the mosaic physically lives
As established in the security literature, the model itself is stateless — the fragments don’t fuse inside the neural network during inference. They fuse in the logs: retention databases, abuse-monitoring systems, safety-review pipelines, and telemetry that coexist under the provider’s roof. Standard enterprise agreements involve retention windows (commonly up to 30 days, longer where legally compelled); zero-data-retention is a negotiated upgrade, not a default. The mosaic exists in raw, un-assembled form the moment the prompts land — and modern LLMs are themselves the most efficient mosaic-assembly tools ever built. Anyone with log access and one instruction — “summarize what this organization appears to be planning” — can do in minutes what would have taken a corporate-intelligence firm months.
The bank’s exposure, precisely stated: it has externalized the raw material of its own competitive intelligence dossier, and the only thing preventing assembly is other people’s access controls.
The calendar: the mosaic tile generator nobody audits
One system manufactures mosaic tiles continuously, org-wide, with no DLP inspection at all: the shared calendar. Consider a single entry visible to hundreds of colleagues by default: “Private aviation — Chicago — June 4 — mtg w/ Tom Anderson (CEO) — strategic discussion.” That one tile discloses the counterparty and rank (CEO-to-CEO isn’t a vendor review), the location (his city, not yours — you’re the one traveling, signaling who’s courting whom), the transport (private aviation means it matters), the word “strategic” (a euphemism so standardized it means M&A), and the date. Only a fraction of the truth has to surface for the mosaic to build — and this fraction is enormous. Add the surrounding auto-generated tiles — the assistant’s tail-number block, legal’s new “Project [codename]” recurring entry, the analyst asking AI to “review my deck for the meeting” that morning — and the deck reveals the what while the calendar reveals the who, when, and how serious.
This isn’t theoretical: corporate jet movements are so predictive of M&A that hedge funds buy aviation-tracking data, and in 2019 observers spotted Occidental’s jet in Omaha days before Berkshire’s $10 billion Anadarko-bid investment was announced. The plane is a public tile; the calendar entry is the private annotation explaining it. And modern suites feed calendars directly into AI assistants by design — Copilot and Gemini for Workspace read invites, agendas, and email threads as core context — so the mosaic assembly that once needed an adversary with log access is now a product feature running continuously inside the tenant, one crafted email away (per EchoLeak) from being asked to share it.
Part 2: The Presentation Problem — When the Employee Assembles the Mosaic for You
“Here’s the deck — give me tips on how to improve it”
The mosaic method requires an adversary to do assembly work. The document-upload pattern eliminates even that step, because a strategy presentation is the pre-assembled mosaic. Consider what a corporate bank’s board-level deck actually contains: the M&A pipeline slide, the credit-concentration heat map, the “strategic alternatives” slide, the regulatory-remediation status page, the pro-forma financials with the real numbers, and — most dangerous of all — the speaker notes, where people write the things too candid for the slide: “Regulators haven’t seen this yet.” “Assumes we exit the Miami CRE book by Q2.” “Board split 5–4 on the sale process.”
When that file hits a cloud AI service, the entire text — slides, notes, embedded tables, OCR’d images — enters the context window. When it hits an agentic tool that fixes the presentation, the exposure multiplies: the agent parses the file, spins up analysis passes, calls rendering tools, writes intermediate copies to a sandbox, generates the improved deck, and stores the output artifact for download. One upload becomes five or ten transmissions through subsystems with independent retention behaviors. The improved deck now exists in the provider’s file storage alongside the original in the logs.
And the speaker-notes risk is no longer theoretical: EchoLeak (CVE-2025-32711, CVSS 9.3) — the first publicly documented zero-click prompt-injection exfiltration against a production enterprise AI — was executed in one documented variant through hidden prompts embedded in PowerPoint speaker notes, which Microsoft 365 Copilot processed during normal summarization and then exfiltrated data through trusted Microsoft domains. No clicks. No alerts. The attack surface was the presentation itself.
“Find the weaknesses” — and “make the case to buy us out”
The pattern deepens when the banker asks the AI to attack the argument: “Where would a skeptical board push back? What would a rival bidder say?” Excellent professional practice — and simultaneously a voluntary red-team of the bank’s position, transcribed and stored off-premises. The AI’s structured enumeration of the bank’s soft spots, weak assumptions, and vulnerable negotiating positions is a new confidential document that never existed before, sitting on someone else’s infrastructure. If the deck is the mosaic, the weakness analysis is the mosaic with arrows drawn on it.
The worst case happens innocently all the time: a corporate-development officer exploring options asks, “Using our financials, create a presentation making the case for why an acquirer should buy us.” In one request the bank manufactures and exports its own valuation argument, an implicit floor price, synergy assumptions revealing cost structures and redundant departments, and a candid inventory of why it might need to sell — effectively the first draft of a confidential information memorandum, the document that in a real sale process is released only under NDA, through a watermarked data room, to vetted parties. Here it was created in an afternoon with none of those protections, potentially constituting undisclosed material information on a third party’s servers. The employee wasn’t leaking. They were brainstorming. The system made no distinction.
Part 3: Promise, Not Physics — Why “We Don’t Train on Your Data” Is an Incentive, Not a Wall
Here is the sentence every bank should require its AI vendors to say out loud: “Nothing technically prevents us from using your data. We have promised not to, and it is currently in our commercial interest to keep that promise.”
That is the actual security model. No-training clauses, retention limits, and access controls are contractual constructs — real, legally enforceable after the fact, and taken seriously by reputable providers. But they are categorically different from physical impossibility. Encryption a provider cannot break is physics. A no-training clause is a promise. The data arrives in processable form — it must, to be processed — and from that moment every protection is a decision someone else keeps making, forever.
The standard reassurance is the long-term incentive: no lab would train on enterprise data because getting caught would vaporize billions in trust. That’s correct — and insufficient — for three reasons.
History is a graveyard of long-term incentives. “It would be irrational to betray this trust” is the epitaph of every institutional scandal. It was irrational for banks to write liar loans in 2006; they wrote them anyway, because the bonus was this quarter and the collapse was someone else’s tenure. Wells Fargo’s fake accounts, Enron’s hidden debt, Boeing trading certification for schedule, Facebook’s promises meeting Cambridge Analytica, Google’s $391 million location-tracking settlement — every one of these institutions had an overwhelming long-term incentive to behave, and every one contained people for whom the short-term payoff was closer and more rewarded than the distant catastrophe. The AI industry’s specific pressure makes this worse: labs face an existential shortage of training data, burn billions annually, and sit atop the most valuable untapped corpus on Earth — the documents behind corporate firewalls. The incentive to keep the promise is real. So is the hunger. A bank betting confidential data on that balance is making a permanent bet on a counterparty’s future behavior under escalating pressure — through every funding crunch, leadership change, and quiet terms-of-service revision.
The promise can be overridden by people who never made it. The New York Times v. OpenAI produced a court order compelling OpenAI to preserve user conversations it would otherwise have deleted. Retention promises yield to subpoenas and discovery; the bank’s mosaic, under legal hold, becomes an asset in someone else’s lawsuit. No one broke their word — the word was never theirs alone to keep.
The promise doesn’t bind the failure modes. A no-training clause governs intentional use. It says nothing about OpenAI’s own 2025 Mixpanel vendor exposure, DeepSeek’s publicly exposed database that spilled over a million log lines in January 2025, malicious browser extensions harvesting conversations, or the employee anywhere in the chain who takes a screenshot. Promises constrain the honest. Breaches require only complexity — and the AI pipeline is the most complex data path a bank has ever used.
Part 3.5: “The Logs” — Who Can Actually Read Your Data?
One phrase keeps appearing in this article: the logs. So ask the simple question — when a prompt is stored somewhere, who specifically can read it, and do they know it’s their job to protect it?
The uncomfortable answer: “the logs” is not one file watched by one accountable person. It is many copies of your data, across many systems, owned by many organizations. At the AI company, several groups may see prompts: safety reviewers (reading flagged conversations is their designed job), engineers keeping the service running, teams improving the models, and — the weak link — lower-paid outside contractors who label training data, often overseas, with the least connection to your bank and the least to lose. Plus whoever answers a subpoena.
Outside the AI company, the chain keeps going: the provider’s own vendors — hosting (Microsoft, Amazon, Google), analytics, monitoring — each with their own staff. When OpenAI disclosed a 2025 exposure, the cause was a third-party analytics vendor most users had never heard of. And most of these people have no idea whose data they’re touching — a contractor reviewing conversations doesn’t know one is a bank’s merger plan. You cannot honor a responsibility you don’t know you have.
Here is the point that ties it together. At any single organization, access is controlled and staff generally understand their duty. So each link can honestly say “we handled it responsibly.” But no one holds the whole chain — the bank can’t see the provider’s access list, the provider can’t fully see its vendors’. When data leaks, it leaks through the seams between organizations: responsibility was real at every step and owned by no one end to end. “Is my data safe in the logs?” has no yes-or-no answer, because there is no single “the logs” and no one you could even call to ask.
A corporate bank using cloud services and internet-connected LLMs faces at minimum twenty-two distinct exposure classes. Most banks’ risk registers cover three of them. The first eleven are the primary ring — the vectors through which the bank’s own AI usage leaks. The second eleven are the quieter ring: adjacent tools and pipelines nobody thinks of as “AI risk” until the data is gone.
1. Shadow AI — the unmanaged front door. The dominant real-world risk, covered in full above. Everything below is secondary to it.
2. The consumer/enterprise confusion gap. The bank signs an enterprise agreement; the employee uses the free tier at home on the same documents. The contract protects the tenant, not the data — one wrong login and the no-training clause never applied.
3. Provider retention and the log mosaic. Even under enterprise terms, prompts persist in retention windows and safety pipelines — the assembled raw material of Part 1, awaiting only access.
4. Prompt injection against agents — the EchoLeak class. Agents read email, documents, and web pages; any input can carry hidden instructions, and the agent can’t reliably tell content from command. EchoLeak (described in Part 2) required zero clicks. OWASP finds prompt injection in over 73% of assessed production deployments; IBM prices these attacks near $6 million.
5. The “lethal trifecta.” Simon Willison’s framing: an agent with private-data access, exposure to untrusted content, and an external communication channel is a fully exploitable exfiltration engine. A bank’s document-improving, email-reading assistant has all three by design.
6. MCP and tool-chain compromise. The protocol connecting agents to databases and files expands attack surface with capability. In January 2026, three prompt-injection CVEs hit Anthropic’s own official Git MCP server — a malicious README sufficed to trigger exfiltration.
7. Vendor and supply-chain exposure. The provider’s own vendors — analytics (OpenAI/Mixpanel, 2025), hosting, evaluation contractors — each extend the trust chain. One contract; a dozen organizations touch the data.
8. Misconfiguration. The 2025–2026 incident record is dominated by open buckets and databases left exposed (see DeepSeek in Part 3), not cinematic hacks.
9. Legal process and discovery. Every retained prompt is discoverable; a bank in litigation may see its own AI conversations produced to opposing counsel under subpoena.
10. Calendar and metadata exposure. Covered above: org-wide calendars broadcast structured intent — counterparties, travel, codenames — with no content inspection, and AI assistants ingest them by design.
11. Model memorization. Data that enters training corpora (via consumer tiers, shadow AI, or terms changes) can be memorized and regurgitated by later models. Once in the weights, it can’t be recalled, deleted, or subpoenaed back — the only irreversible layer.
The second ring: adjacent pipelines and the offensive flip
The first eleven vectors describe data leaking out. The next eleven include the ways data leaks through systems nobody labels “AI,” and — critically — the ways an adversary can weaponize the same channels to push falsehoods in.
12. Data poisoning and “LLM grooming” — attacking a company through the models everyone else trusts. This is the offensive inversion of leakage, and it is the vector you should worry about most, because you cannot patch it inside your own walls. The threat is no longer that your data gets out — it’s that an adversary floods the AI ecosystem with fabricated information about your bank so that every chatbot, every analyst using AI research, every journalist, and every counterparty is told a lie as fact. The playbook is proven at nation-state scale: Russia’s Pravda network published roughly 3.6 million articles across 150 domains in 49 countries in 2024, engineered not to be read by humans (the sites average under 1,000 monthly visitors) but to saturate the web so that AI crawlers ingest and repeat them. When NewsGuard tested ten leading chatbots — ChatGPT, Gemini, Copilot, Claude, Grok, Perplexity, and others — they repeated the false narratives 33% of the time, and seven of the ten cited the propaganda sites as legitimate sources. The American Sunlight Project named the technique “LLM grooming”: the more often a claim appears across indexed content, the more likely models are to absorb it as truth. Now scale that down to a single company. A motivated adversary — a short seller, a hostile bidder, a disgruntled competitor, a foreign rival — can spin up dozens of plausible-looking financial-news domains and flood them with fabricated claims: that your bank is under secret regulatory investigation, that its CRE book is insolvent, that a named executive is about to be indicted, that deposits are fleeing. Most of it never needs a human reader. It needs to exist often enough that when someone asks an AI “is [your bank] financially healthy?” the model hedges, or worse, repeats the smear. Because the false narrative and the assembled true mosaic feed the same models, an attacker who has also harvested fragments of your real strategy can craft disinformation that is corroborated by genuine detail — the most credible lie is the one wrapped around a true fact.
13. Training-data backdoors — poisoning the model itself. Grooming pollutes what models read from the web; backdooring corrupts what they learn in training. A landmark October 2025 study by Anthropic, the UK AI Security Institute, and the Alan Turing Institute found that as few as 250 malicious documents can implant a hidden backdoor in a model regardless of its size — overturning the assumption that attackers must control a percentage of training data. A separate study in Nature Medicine showed that replacing just 0.001% of training tokens with misinformation produced measurably more error-prone models that still passed every standard benchmark — the poison was undetectable by normal testing. For a bank building or fine-tuning its own custom model on scraped or third-party data (the fastest-growing deployment pattern), this means the model it trusts most could carry a trigger that leaks data or produces attacker-chosen outputs on command. And “anyone can create online content that might eventually end up in a model’s training data” — the barrier to entry is a few hundred web pages.
14. RAG and vector-store poisoning. Banks increasingly connect AI to internal knowledge via retrieval-augmented generation — the model answers from a vector database of the bank’s documents. If an attacker (or a careless integration) gets poisoned content into that store — a doctored policy memo, a fake precedent, a booby-trapped PDF — every employee who queries the assistant is served the corruption as authoritative internal truth. The lethal-trifecta and MCP risks from vectors 4–6 apply here in reverse: the retrieval layer is both an ingestion point for bad data and, via injected instructions, an exfiltration trigger.
15. Third-party AI features silently embedded in ordinary software. The bank vets “AI tools” and misses the AI now baked into everything else: the CRM that added a summarization feature, the video-conferencing platform that auto-transcribes, the email client with smart-compose, the PDF reader with a chat assistant, the helpdesk with an AI triage bot. Each may route content to a model under terms the bank never reviewed. The Otter.ai class of incident is instructive — meeting-transcription bots that silently join calls, record, and retain sensitive discussions, sometimes emailing transcripts to unintended recipients. The bank didn’t adopt an AI strategy for these; the vendors did it for them.
15. Silent AI baked into ordinary software. The bank vets “AI tools” and misses the AI now embedded everywhere else — the CRM summarizer, the auto-transcribing video platform, the PDF reader’s chat assistant, the helpdesk triage bot — each routing content to a model under terms nobody reviewed. The vendors adopted an AI strategy on the bank’s behalf.
16. AI notetakers in confidential meetings. Executives routinely admit notetakers (Otter, Fireflies, Copilot, Zoom AI) to board meetings, deal negotiations, and legal strategy calls. The transcript is a verbatim, cloud-stored, searchable record of the most sensitive conversation in the building — often auto-shared to attendees including external parties. One notetaker in a merger negotiation captures both sides’ positions in a single file.
17. Browser extensions harvesting sessions. Over-permissioned “AI helper” extensions capture ChatGPT and DeepSeek prompts and responses and ship them to attacker servers, often invisibly. The enterprise contract is irrelevant when the leak sits between the keyboard and the browser.
18. Usage-pattern and metadata inference. Even without reading content, a spike in the M&A team’s queries or a burst of severance-law questions from one department leaks the tempo of the bank’s intentions. Traffic analysis is a mature intelligence discipline, and AI usage generates rich traffic.
19. Cross-tenant isolation failures. Enterprise assurances rest on tenant isolation, but software isolation fails — the 2023 ChatGPT bug that showed users others’ chat titles was a preview, and memory/org-knowledge features widen the blast radius of any such failure.
20. AI as an insider-exfiltration laundering tool. A departing employee emailing themselves the client book trips DLP; the same employee asking an assistant to “summarize our top 50 relationships with contact details and deal history” may not. Agentic access turns one insider prompt into a curated intelligence package.
21. The bank’s own prompt logs. Internal AI gateways that log every prompt for audit create a new crown-jewel database — every sensitive question every employee ever asked, in plaintext, in one place — often protected worse than a core banking system. The safeguard becomes the target.
22. AI-manufactured attacks aimed inward. Beyond grooming, adversaries use generative AI for deepfake audio authorizing wires (already a documented multimillion-dollar fraud), synthetic “leaked documents” to spook depositors, and mosaic-tuned phishing. The same technology that leaks the bank’s truth outward manufactures falsehoods aimed in.
Layer the taxonomy over the scenarios: the mosaic accumulates through vectors 1–3, is timestamped and annotated by vector 10, and is quietly widened by the adjacent-tool ring of 14–21; the presentation and buyout deck travel through 4–8; everything persists into 9; and the fragments risk immortality through 11. The June 4th Chicago meeting illustrates the defensive stack in miniature: the calendar entry names the counterparty and the date, the deck-review request supplies the substance, the AI assistant joins them by design, and the log retains the union. But vectors 12, 13, and 22 flip the whole model on its head — there, the bank isn’t leaking anything; an adversary is pushing falsehoods in, poisoning the models the bank and its counterparties rely on, so that the market is told the bank is failing whether or not it is. Leakage and disinformation are the same pipeline run in opposite directions. The bank never decided to disclose anything, and never decided to be lied about. The architecture decided both for it.
Part 5: What a Prudent Bank Actually Does
The answer is not abstinence — banning AI simply drives usage onto personal phones, the worst channel of all. The answer is architectural honesty, and it starts with the dominant risk:
Give employees a safe tool, don’t just ban the unsafe one. As the shadow-AI section argued, a ban with no alternative converts a governable risk into an invisible one. The primary fix is a sanctioned internal option good enough that no one reaches for the kitchen-table consumer account — paired with clear guidance that it exists precisely so employees never need a personal account for work, on any device.
Segment by blast radius. Public-data tasks (market summaries, coding help) can use enterprise cloud tiers with negotiated zero-data-retention. Anything touching MNPI, client data, credit files, deal work, or strategy belongs on infrastructure the bank controls — on-premises open-weight models (now achievable at mid-market cost), private-cloud deployments in the bank’s own tenant, or air-gapped enclaves for the truly sensitive.
Treat documents as the crown jewels, not chats. Policy attention obsesses over what employees type; the catastrophic payloads are what they upload. Presentation, spreadsheet, and data-room material should be technically blocked from external AI endpoints, with a local alternative provided so the block doesn’t breed shadow usage.
Assume the mosaic. Governance should evaluate AI exposure the way intelligence agencies evaluate publication: not “is this prompt sensitive?” but “what does the corpus of our prompts reveal?” That analysis is sobering exactly once — and then it changes the architecture.
Treat the calendar as a classified feed. Sensitive meetings get codenames, not counterparty names; deal travel gets booked outside the shared system; free/busy visibility replaces full-detail visibility by default; and — critically — the bank decides deliberately which AI assistants may ingest calendar data at all, because an assistant with calendar context has the mosaic’s index. “Strategic discussion with [CEO name]” should never appear in any system whose contents the bank cannot enumerate.
Ban silent AI in adjacent tools. The vetting process must cover AI features embedded in non-AI software — transcription bots, CRM summarizers, meeting notetakers, browser extensions. Default rule: no AI notetaker in any confidential meeting, and no browser extension touching AI sessions on managed devices.
Monitor the disinformation surface, not just the leak surface. Because vectors 12–13 and 22 push falsehoods in rather than pull data out, the bank needs the mirror image of DLP: monitoring what AI systems and the web say about the institution. Periodically query major chatbots about the bank’s health, executives, and reputation; watch for fabricated-domain clusters and AI-repeated smears; and prepare an incident-response playbook for AI-laundered disinformation, deepfakes, and synthetic “leaks” the same way the bank prepares for a cyber breach. In a deposit-taking institution, a widely repeated AI falsehood about solvency is not a PR problem — it is a bank-run risk.
Constrain the trifecta. Any agent with private-data access should lose either untrusted-content exposure or external communication. EchoLeak’s lesson is that trust boundaries are security boundaries; an agent that reads inbound email should not also hold the keys to SharePoint and an outbound channel.
Price the promise correctly. When the vendor says “we don’t train on your data,” the correct response is: “We believe you — and our architecture will be designed so that we never have to.” Trust is a fine thing to have and a terrible thing to depend on. Physics doesn’t renew annually. Contracts do.
🔎 Brian’s Take
“I spent years watching institutions with impeccable long-term incentives make short-term choices that destroyed them, so the ‘no AI lab would ever risk enterprise trust’ argument lands differently on me than it does on a procurement officer. Of course they wouldn’t — right up until a funding crisis, an acquisition, a desperate quarter, or a middle manager with a growth target decides the fine print has some flex in it. Liar loans were irrational for the banks writing them. They wrote them anyway. My rule from the allocation business: never underwrite a risk whose downside is irreversible on the strength of a counterparty’s ongoing self-restraint. Data in someone else’s logs is exactly that risk — and data in someone else’s training run is irreversible in the strictest sense of the word. The mosaic analysis in this piece is the part I’d staple to every bank board deck: your institution is disclosing continuously, in fragments, through its most diligent employees, and the assembled picture sits outside your walls priced at other people’s discipline. But here’s the part that genuinely changed how I think about it — the pipe runs both ways. It’s not just that your secrets leak out; it’s that an adversary can pump lies in. A short seller who both harvests fragments of your real strategy and floods the web with fabricated claims about your solvency can get the world’s AI systems to repeat a smear that’s half-true and therefore devastating. For a bank — an institution that lives and dies on confidence — an AI-amplified falsehood about your balance sheet isn’t a reputational nuisance, it’s a run. The banks that banned ChatGPT in week one weren’t Luddites. They were the only ones who read the architecture instead of the marketing.”
Frequently Asked Questions
What is the biggest way banks leak data to AI? Shadow AI — employees working from home on personal devices, entering confidential bank information into personal ChatGPT, Gemini, or Claude accounts that sit outside every corporate safeguard. On consumer tiers, that data may be used to train future models by default, and because the traffic never touches the corporate network the bank can’t detect it. Banning AI makes it worse by pushing usage further underground; the fix is a sanctioned internal tool good enough that employees don’t need a personal account.
Can AI providers technically access enterprise prompts? Yes. Data must arrive in processable form to be processed. No-training clauses, retention limits, and access controls are contractual and organizational safeguards — enforceable, but categorically different from technical impossibility.
What is the mosaic effect in AI data leakage? It is the reconstruction of confidential conclusions from many individually harmless fragments. Hundreds of employee prompts, each innocuous alone, can collectively reveal deals, layoffs, credit problems, and strategy — and DLP tools inspecting messages one at a time cannot detect it.
What was EchoLeak? EchoLeak (CVE-2025-32711) was the first zero-click prompt-injection data-exfiltration attack on a production enterprise AI assistant, Microsoft 365 Copilot. Hidden instructions — in one documented variant embedded in PowerPoint speaker notes — caused the assistant to leak tenant data through trusted domains with no user interaction.
Why did major banks ban ChatGPT? JPMorgan, Bank of America, Citigroup, Goldman Sachs, Deutsche Bank, and Wells Fargo restricted or banned ChatGPT beginning in February 2023 over data-leakage and regulatory concerns; BNY Mellon cited the impossibility of meeting fiduciary data-handling duties through third-party training pipelines.
Is organization-wide calendar sharing a form of data leakage? Yes. Shared calendars broadcast structured intent — counterparty names, locations, private travel, project codenames, and meeting cadences — with no DLP inspection, and AI assistants like Copilot and Gemini ingest calendar data as core context, automatically joining it to emails and documents. A single entry like “private aviation to Chicago, June 4, meeting with [CEO name] — strategic discussion” can reveal an impending merger on its own; corporate jet movements alone have historically predicted M&A announcements.
Can an adversary attack a company by feeding AI false information? Yes. Two techniques stand out. “LLM grooming” floods the web with fabricated content so AI models absorb and repeat it — Russia’s Pravda network published ~3.6 million articles in 2024 and got ten leading chatbots to repeat its falsehoods 33% of the time. “Data poisoning” corrupts training data directly — an Anthropic/UK AI Security Institute/Alan Turing study found just 250 malicious documents can backdoor a model of any size. Scaled to one company, a short seller or hostile bidder could flood plausible-looking financial-news domains with false claims about a bank’s solvency or an executive’s conduct, engineering AI systems to repeat the smear as fact.
How much do AI-related breaches cost? IBM’s 2026 Cost of a Data Breach Report puts the average breach at $5 million — up 12% year-over-year — with prompt-injection and inversion attacks on AI tools averaging roughly $6 million, and shadow-AI breaches costing about $670,000 more than standard incidents.
About the Author: Brian French
Brian French is a former institutional money manager and analyst with a career spent evaluating companies, industries, and capital flows on behalf of institutional clients. His background spans equity research, portfolio management, and macro-driven sector analysis — experience he now applies to dissecting technology risk, financial institutions, and long-horizon investment themes. Known for a “follow the capital, not the headlines” approach, Brian focuses on the durable structural forces — incentive design, counterparty risk, and institutional behavior under pressure — that determine which safeguards hold and which fail. He writes and speaks about the intersection of finance, technology, and economic geography.
This article is for informational purposes only and does not constitute legal, compliance, security, or investment advice. Consult qualified counsel regarding data-protection obligations applicable to your institution.
Resources and Sources
- IBM — Cost of a Data Breach Report 2026 (via Cybersecurity Dive) — $5M average breach cost (+12% YoY), ~$6M average for prompt-injection and inversion attacks, 85% of breached organizations increasing governance spend. (cybersecuritydive.com)
- Reco — “AI & Cloud Security Breaches: 2025 Year in Review” — EchoLeak mechanics (email → Copilot ingestion → OneDrive/SharePoint/Teams extraction → exfiltration via trusted domains, zero clicks), IBM 2025 baseline ($4.44M). (reco.ai)
- Sysdig — “The Comprehensive Guide to Prompt Injection Attacks in 2026” — EchoLeak case study, Simon Willison’s “lethal trifecta,” Cursor MCP configuration attack, attacker/defender economics. (sysdig.com)
- Vectra AI — “Prompt injection: types, real-world CVEs, and enterprise defenses” — CVE-2025-32711 (CVSS 9.3) technical breakdown, CVE-2025-53773 GitHub Copilot RCE (CVSS 9.6), OWASP 73% production-deployment finding, Cisco State of AI Security 2026 (83% deploying, 29% ready). (vectra.ai)
- TechStoriess — “AI Agent Security Practices 2026” — EchoLeak via PowerPoint speaker notes, shadow-AI breach economics ($670K premium, 247-day detection), 88% incident rate vs. 82% executive confidence gap. (techstoriess.com)
- BlueRadius — “AI Cybersecurity Incident Report 2026” — MITRE ATLAS v5.1.0 agent-attack techniques, malicious browser extensions harvesting ChatGPT/DeepSeek conversations, OpenAI–Mixpanel third-party exposure disclosure, Samsung three-incidents-in-twenty-days pattern. (blueradius.io)
- Cyber Desserts — “Prompt Injection Attacks: Examples, Techniques, and Defence” — January 2026 CVEs in Anthropic’s official Git MCP server (CVE-2025-68143/68144/68145), MCP attack surface across Microsoft/OpenAI/Google/Amazon ecosystems. (blog.cyberdesserts.com)
- PurpleSec — “Data Exfiltration Via AI Prompt Injection” — Salesforce “ForcedLeak” hidden-prompt vulnerability, direct vs. indirect injection taxonomy. (purplesec.us)
- Techglock — “Prompt Injection: The #1 AI Threat in 2026” — Lethal-trifecta exploitability framework applied to enterprise agents. (techglock.com)
- Forbes — “Workers’ ChatGPT Use Restricted At More Banks — Including Goldman, Citigroup” (Feb 2023) — Bank of America unauthorized-apps listing alongside WhatsApp, Citigroup/Goldman third-party software restrictions, Deutsche Bank access disablement. (forbes.com)
- Moveo.AI — “Companies Banning ChatGPT (2026): The Enterprise Security List” — JPMorgan, Deutsche Bank, Wells Fargo, BofA, Citi, Goldman restrictions; Morgan Stanley “build, don’t block” internal deployment; BNY Mellon fiduciary rationale. (moveo.ai)
- The Telegraph via TipRanks — “JPMorgan restricts use of ChatGPT among staff” — Regulatory-action concern over shared financial information. (tipranks.com)
- Fortune / Yahoo Finance — “Apple, Goldman Sachs, and Samsung among growing list of companies banning ChatGPT” — Samsung April 2023 leak details (internal code, meeting recordings) and subsequent internal-AI pivot; Amazon warnings after outputs resembling internal data. (finance.yahoo.com)
- Visbanking — “Major Banks Restricting Use of ChatGPT” — Amazon confidential-data warnings, JPMorgan third-party-controls framing. (visbanking.com)
- Wiz Research — DeepSeek exposed ClickHouse database disclosure (January 2025) — Publicly accessible database containing over one million log lines including chat histories, discovered via routine scanning. (wiz.io)
- New York Times Co. v. OpenAI — preservation order coverage (2025) — Court-ordered retention of user conversations overriding deletion policies, establishing legal-process precedence over privacy commitments. (court filings; widely reported)
- Simon Willison — “The Lethal Trifecta” (2025) — Original framing of private-data access + untrusted content + external communication as the complete agent-exfiltration condition. (simonwillison.net)
- CNBC / Reuters — Occidental jet in Omaha coverage (April 2019) — Corporate-jet tracking preceding Berkshire Hathaway’s $10 billion Anadarko-bid investment; illustrative of aviation metadata anticipating deal announcements. (cnbc.com)
- Academic and market research on corporate jet tracking — Studies and hedge-fund data products (e.g., aviation-data feeds formerly distributed via Quandl) demonstrating that flights between headquarters cities predict M&A activity. (ssrn.com; quandl/Nasdaq Data Link archives)
- Microsoft — Microsoft 365 Copilot documentation; Google — Gemini for Workspace documentation — Calendar, email, and document ingestion as core assistant context within the enterprise tenant. (learn.microsoft.com; workspace.google.com)
- NewsGuard — “Russian Propaganda Has Now Infected Western AI Chatbots” audit (March 2025) (via Forbes, The Hill, Yahoo News) — Pravda network’s 3.6 million articles across 150 domains in 49 countries; ten leading chatbots repeating false narratives 33% of the time; seven citing Pravda sites as sources; 92 disinformation articles cited across models. (forbes.com; newsguardtech.com)
- American Sunlight Project — LLM grooming report (February 2025) — Definition and mechanics of “LLM grooming”; correlation between narrative saturation and model absorption; 97 Pravda domains publishing ~20,000 articles in 48 hours. (americansunlight.org)
- DFRLab — “Pravda in the pipeline: Early evidence of state-adjacent propaganda in AI training data” (April 2026) — AI poisoning co-opting the retrieval layer of LLMs. (dfrlab.org)
- Anthropic, UK AI Security Institute & Alan Turing Institute — “A small number of samples can poison LLMs of any size” (October 2025) — 250 malicious documents sufficient to backdoor models from 600M to 13B parameters regardless of training-data volume. (anthropic.com)
- Nature Medicine — training-data medical-misinformation study (2025) (via Dell Technologies analysis) — Replacing 0.001% of training tokens with misinformation produced measurably more error-prone models that still passed standard benchmarks; ~100 poisoned models found on Hugging Face. (dell.com; nature.com)
- Reporting on AI notetaker and transcription exposure (Otter.ai class incidents) and malicious AI browser extensions (via BlueRadius AI Cybersecurity Incident Report 2026 and related coverage) — Silent meeting-transcription capture and browser-extension harvesting of ChatGPT/DeepSeek sessions. (blueradius.io)