The Mosaic, the PowerPoint, and the Promise That Isn’t Physics
By Brian French, Former Institutional Investment Manager and Bank Officer
Research Topic: A deep-dive analysis of how a corporate bank can inadvertently disclose confidential information through cloud LLMs and agentic AI tools — the mosaic effect, presentation-improvement leaks, why “we don’t train on your data” is a promise rather than physics, and every major breach vector when banks use cloud AI.
Quick Answer: How Does a Bank Leak Confidential Information to AI Without Anyone Breaking a Rule?
A corporate bank leaks confidential information to cloud AI in four compounding ways: (1) the mosaic effect, where hundreds of individually harmless employee prompts — plus shared-calendar metadata — collectively reconstruct deals, layoffs, and strategy; (2) high-density document uploads, where a single “improve my presentation” request hands over the entire assembled picture at once; (3) adversarial extraction, where the same analytical power employees use for self-critique (“find the weaknesses in my pitch”) produces, in aggregate, an acquirer’s playbook against the bank; and (4) infrastructure exposure, where prompt logs, agent pipelines, vendor supply chains, AI notetakers, browser extensions, and prompt-injection attacks create breach paths no bank firewall can see. A fifth, offensive dimension inverts the whole problem: data poisoning and “LLM grooming,” where an adversary floods the AI ecosystem with fabricated claims about the bank so that chatbots repeat the lie as fact — a tactic already proven at nation-state scale. The core problem: cloud AI confidentiality rests on contractual promises and commercial incentives — not on physical impossibility — and history shows institutions under pressure choose short-term benefit over long-term restraint with depressing regularity. IBM’s 2026 Cost of a Data Breach Report puts the average breach at $5 million, with AI prompt-injection attacks averaging roughly $6 million.
Introduction: The Bank That Never Meant to Tell Anyone Anything
No one at the bank did anything wrong.
The credit analyst asked a chatbot to explain a covenant structure. The VP of corporate development asked for help polishing a presentation. The HR business partner asked about severance norms in three states. The syndications desk asked an AI to summarize a term sheet. The junior associate — brilliant, exhausted, and up against a 6 a.m. deadline — uploaded the whole pitch deck and typed: “Give me tips on how to improve this.”
Five hundred employees. Five hundred reasonable requests. Zero policy violations, if the bank’s enterprise AI agreement is in place. And yet, at the end of the quarter, a complete picture of the bank’s M&A pipeline, credit exposure concentrations, restructuring plans, and negotiating posture exists — assembled, structured, and stored — on infrastructure the bank does not own, cannot inspect, and could not subpoena faster than a plaintiff’s lawyer could.
This is not a hypothetical dreamed up by AI skeptics. It is why JPMorgan restricted ChatGPT in February 2023, why Deutsche Bank disabled access entirely, why Bank of America put it on the same unauthorized-apps list as WhatsApp, and why Citigroup, Goldman Sachs, and Wells Fargo followed within days. BNY Mellon blocked public LLMs citing the impossibility of meeting fiduciary data-handling requirements with third-party training pipelines. The banks understood something in 2023 that most mid-sized institutions still haven’t internalized in 2026: the danger isn’t the secret you decide to share. It’s the secret you disclose without ever deciding anything.
This article dissects exactly how that happens — mechanism by mechanism — and then confronts the uncomfortable foundation underneath it all: that every safeguard between a bank’s data and an AI provider’s ambitions is a promise, not a law of physics.
The Mosaic Method — Death by Five Hundred Harmless Prompts
How intelligence agencies think about disclosure (and banks don’t)
Intelligence professionals have a name for what banks are doing to themselves: the mosaic effect. No single tile reveals the picture. The picture emerges only when tiles accumulate in one place. Classification regimes exist precisely because analysts learned that an adversary with enough unclassified fragments can reconstruct classified conclusions.
Now map that onto a corporate bank’s daily AI usage:
- A leveraged-finance associate asks the AI to “stress test debt service coverage at 9% rates for a borrower in specialty chemicals.”
- A workout officer asks how Article 9 foreclosure works in Texas “for a $40M equipment-backed facility.”
- A compliance analyst asks the AI to summarize SAR filing thresholds “for repeated structuring just under $10,000 through our Brownsville branch.”
- Corporate development asks for “integration cost benchmarks for acquiring a bank with 45 branches and $8B in assets.”
- HR asks for “communication templates for a reduction affecting a commercial lending team.”
- The CFO’s office asks the AI to “sanity check” a liquidity coverage calculation with real numbers, lightly disguised.
Each prompt passes any reasonable data-loss-prevention filter, because DLP inspects messages one at a time and the sensitivity here is emergent — it exists only in combination. Assembled, those six prompts reveal: a troubled chemicals credit, a specific workout in progress, a potential BSA/AML problem at a named location, an acquisition target profile precise enough to shortlist candidates, an impending layoff in a specific department, and the bank’s actual liquidity position. That is a mosaic worth millions to a competitor, an activist investor, a short seller, or a litigant — and no individual employee had the authority to disclose it, because no individual employee did disclose it.
Where the mosaic physically lives
As established in the security literature, the model itself is stateless — the fragments don’t fuse inside the neural network during inference. They fuse in the logs: retention databases, abuse-monitoring systems, safety-review pipelines, and telemetry that coexist under the provider’s roof. Standard enterprise agreements involve retention windows (commonly up to 30 days, longer where legally compelled); zero-data-retention is a negotiated upgrade, not a default. The mosaic exists in raw, un-assembled form the moment the prompts land — and modern LLMs are themselves the most efficient mosaic-assembly tools ever built. Anyone with log access and one instruction — “summarize what this organization appears to be planning” — can do in minutes what would have taken a corporate-intelligence firm months.
The bank’s exposure, precisely stated: it has externalized the raw material of its own competitive intelligence dossier, and the only thing preventing assembly is other people’s access controls.
The calendar: the mosaic tile generator nobody audits
There is one system inside every bank that manufactures mosaic tiles continuously, organization-wide, with no DLP inspection at all: the shared calendar. Org-wide calendar visibility is standard practice — it’s how assistants book meetings and how teams coordinate. It is also a structured, timestamped, machine-readable feed of the institution’s intentions.
Consider a single entry on an executive’s calendar, visible to hundreds of colleagues by default: “Private aviation — Chicago — June 4 — mtg w/ Tom Anderson (CEO) — strategic discussion.” Decompose what that one tile discloses: the counterparty’s identity and rank (CEO-to-CEO meetings are not vendor reviews), the location (his city, not yours — you are the one traveling, which signals who is courting whom), the transport (private aviation means confidentiality and expense are both justified — this meeting matters), the word “strategic” (corporate euphemism so standardized it functions as a synonym for M&A), and the date (a countdown clock for anyone watching the pattern). Only a fraction of the data has to be revealed for the mosaic to build — and this fraction is enormous.
Now add the surrounding tiles the calendar generates automatically: the assistant’s entry blocking the tail number and FBO, the follow-up “diligence workstream kickoff” placeholder two weeks later, legal’s recurring “Project [codename]” block that appeared the same week, and the analyst who — that same morning — asks the AI to “review my pitch deck for the meeting.” The deck upload and the calendar entry cross-confirm each other: the document reveals the what, the calendar reveals the who, when, and how serious. Neither system’s administrators ever see the combination, because the combination doesn’t live in either system. It lives in whatever ingests both.
This isn’t theoretical tradecraft. Corporate jet movements are so predictive of M&A that hedge funds pay for aviation-tracking data, and academic research has confirmed that flights between headquarters cities anticipate deal announcements; in 2019, observers famously spotted Occidental Petroleum’s jet in Omaha days before Berkshire Hathaway’s $10 billion investment in its bid for Anadarko was announced. The plane is a public mosaic tile. The calendar entry is the private annotation that explains it.
And here is where AI changes the calculus: the modern productivity suite now feeds calendars directly into AI assistants by design. Microsoft 365 Copilot and Google’s Gemini for Workspace read calendars, meeting invites, attached agendas, and email threads as core context — that is their value proposition. The assistant that helpfully answers “what should I prepare for the June 4th Chicago meeting?” has already joined the calendar tile to the email tiles to the document tiles. The mosaic assembly that once required an adversary with log access is now a product feature running continuously inside the tenant — and, per the EchoLeak class of attacks in Part 4, an assistant with that assembled context and an inbound channel for untrusted content is one crafted email away from being asked to share it.
The Presentation Problem — When the Employee Assembles the Mosaic for You
“Here’s the deck — give me tips on how to improve it”
The mosaic method requires an adversary to do assembly work. The document-upload pattern eliminates even that step, because a strategy presentation is the pre-assembled mosaic. Consider what a corporate bank’s board-level deck actually contains: the M&A pipeline slide, the credit-concentration heat map, the “strategic alternatives” slide, the regulatory-remediation status page, the pro-forma financials with the real numbers, and — most dangerous of all — the speaker notes, where people write the things too candid for the slide: “Regulators haven’t seen this yet.” “Assumes we exit the Miami CRE book by Q2.” “Board split 5–4 on the sale process.”
When that file hits a cloud AI service, the entire text — slides, notes, embedded tables, OCR’d images — enters the context window. When it hits an agentic tool that fixes the presentation, the exposure multiplies: the agent parses the file, spins up analysis passes, calls rendering tools, writes intermediate copies to a sandbox, generates the improved deck, and stores the output artifact for download. One upload becomes five or ten transmissions through subsystems with independent retention behaviors. The improved deck now exists in the provider’s file storage alongside the original in the logs.
And the speaker-notes risk is no longer theoretical: EchoLeak (CVE-2025-32711, CVSS 9.3) — the first publicly documented zero-click prompt-injection exfiltration against a production enterprise AI — was executed in one documented variant through hidden prompts embedded in PowerPoint speaker notes, which Microsoft 365 Copilot processed during normal summarization and then exfiltrated data through trusted Microsoft domains. No clicks. No alerts. The attack surface was the presentation itself.
“Find the weaknesses in the merits of my presentation”
Now the pattern deepens. The banker doesn’t just ask for formatting help — they ask the AI to attack their own argument: “Find the weaknesses in this pitch. Where would a skeptical board push back? What would a rival bidder say?”
This is excellent professional practice and a genuinely valuable use of AI. It is also, from an information-security standpoint, a voluntary red-team of the bank’s position, transcribed and stored off-premises. The AI’s answer — a structured enumeration of the bank’s soft spots, weak assumptions, overstated synergies, and vulnerable negotiating positions — is itself a new confidential document that never existed before, generated on someone else’s infrastructure, and retained under someone else’s policy. The bank has now disclosed not only its strategy but a professional-grade critique of its strategy. If the original deck is the mosaic, the weakness analysis is the mosaic with arrows drawn on it.
“Create a presentation about why someone should buy us out”
The final step in this escalation is the one that should make any general counsel’s blood run cold, and it happens innocently all the time. A corporate development officer, exploring strategic alternatives, asks the agent: “Using our financials, create a presentation making the case for why an acquirer should buy us.”
Think about what the bank has just manufactured and exported:
- A seller’s own valuation argument, including which metrics management believes are most flattering and which it avoids.
- An implicit floor price and the logic behind it.
- A list of synergy assumptions revealing cost structures, redundant departments, and integration pain points.
- A candid inventory of why the bank might need to sell — the deck’s urgency is itself material non-public information.
- Effectively, the first draft of a CIM (confidential information memorandum) — a document that, in a real sale process, is released only under NDA, to vetted counterparties, through a virtual data room with watermarking and access logs.
In the traditional process, every one of those protections exists because a century of M&A practice taught banks that leaked deal intent moves markets, invites hostile interest, spooks depositors and counterparties, and destroys negotiating leverage. The AI-generated version of that same document was created with none of those protections, in an afternoon, by one employee with a chat window — and under securities law, the bank has potentially created undisclosed material information sitting on a third party’s servers, discoverable in litigation, and exposed to every breach vector in Part 4. The employee wasn’t leaking. They were brainstorming. The system made no distinction.
Promise, Not Physics — Why “We Don’t Train on Your Data” Is an Incentive, Not a Wall
The honest architecture of trust
Here is the sentence every bank should require its AI vendors to say out loud: “Nothing technically prevents us from using your data. We have promised not to, and it is currently in our commercial interest to keep that promise.”
That is the actual security model. Enterprise no-training clauses, retention limits, and access controls are contractual and organizational constructs. They are real, they are legally enforceable after the fact, and reputable providers invest heavily in honoring them. But they are categorically different from physical impossibility. Encryption a provider cannot break is physics. A no-training clause is a promise. The data arrives at the provider’s infrastructure in processable form — it must, to be processed — and from that moment forward, every protection is a decision someone else keeps making, renewal after renewal, forever.
The standard reassurance is the long-term incentive argument: no frontier lab would train on enterprise data because getting caught would vaporize billions in enterprise trust. This argument is correct — and insufficient — for three reasons.
Reason one: history is a graveyard of long-term incentives
The claim “it would be irrational to betray this trust” has been the epitaph of every institutional scandal in modern finance and technology. It was irrational for banks to write liar loans in 2006 — the long-term consequence was the destruction of their own collateral. They wrote them anyway, because the bonus was this quarter and the collapse was someone else’s tenure. It was irrational for Wells Fargo employees to open millions of fake accounts, for Enron to hide debt in SPEs, for Arthur Andersen to shred, for Boeing to trade certification rigor for schedule, for Purdue to trade epidemiology for sales. In technology specifically: Facebook’s data promises met Cambridge Analytica; Google paid $391 million to states over location tracking that continued after users turned it off; Amazon reportedly warned employees off ChatGPT after seeing responses that resembled internal Amazon data. Every one of these institutions had an overwhelming long-term incentive to behave. Every one of them contained quarters, divisions, and individuals for whom the short-term payoff was closer, more vivid, and more personally rewarded than the distant catastrophe.
The AI industry’s specific pressure makes this worse, not better: frontier labs face an existential shortage of high-quality training data, burn billions annually, and sit atop the most valuable untapped corpus on Earth — the professional documents behind corporate firewalls. The incentive to keep the promise is real. So is the hunger. A bank betting its confidential information on that balance is making a permanent bet on the indefinite future behavior of a counterparty under escalating pressure — through every funding crunch, every leadership change, every acquisition, every quiet terms-of-service revision, and every “legitimate business purposes” clause a future lawyer decides to stretch.
Reason two: the promise can be overridden by people who never made it
Even a provider with perfect integrity does not fully control the promise. The New York Times v. OpenAI litigation produced a court order compelling OpenAI to preserve user conversations it would otherwise have deleted — a judge overriding a privacy policy in one ruling. Retention promises yield to subpoenas, national-security process, and discovery. The bank’s mosaic, preserved under legal hold, becomes an asset in someone else’s lawsuit. No one at the provider broke their word; the word was simply never theirs alone to keep.
Reason three: the promise doesn’t bind the failure modes
A no-training clause governs intentional use. It says nothing about the Mixpanel-style vendor exposure OpenAI itself disclosed in 2025 (a third-party analytics provider in the supply chain), nothing about DeepSeek’s publicly exposed ClickHouse database that spilled over a million log lines including chat histories in January 2025, nothing about malicious browser extensions actively harvesting ChatGPT and DeepSeek conversations, and nothing about the employee at any vendor in the chain who takes a screenshot. Promises constrain the honest. Breaches don’t require dishonesty — only complexity, and the modern AI pipeline is the most complex data path a bank has ever put client information through.
“The Logs” — Who Can Actually Read Your Data?
Throughout this article one phrase keeps appearing: the logs. It’s worth stopping to ask the simple question a bank executive would ask — when my employee’s prompt is stored somewhere, who, specifically, can read it? And do they know it’s their job to protect it?
The uncomfortable answer is that “the logs” is not one file in one building watched by one accountable person. It is many copies of your data, spread across many systems, owned by many organizations, touchable by many people — most of whom you will never know exist. Here is the chain in plain terms.
At the AI company itself, several different groups can potentially see prompts:
- Safety reviewers. Their entire job is to read flagged conversations to catch abuse. A human reading your prompt isn’t a malfunction here — it’s the system working as designed.
- Engineers who keep the service running. When something breaks at 2 a.m., the people fixing it can reach the databases where your data lives.
- The teams that improve the models. Data scientists and researchers work with usage data to make the next version better.
- Outside contractors who label data. This is the weak link. Much of the human review and training work is done by lower-paid contract workers, sometimes in other countries, hired through outside firms. They rotate in and out, and they have the least connection to your bank and the least to lose if they mishandle something.
- Whoever answers a subpoena. When a court demands data, someone pulls it.
Outside the AI company, the chain keeps going — and this is what banks miss. The AI provider has its own vendors, and each one has its own employees with their own access: the company that hosts the servers (Microsoft, Amazon, Google), the analytics tools, the monitoring software. When OpenAI disclosed a 2025 data exposure, the cause was a third-party analytics vendor — a company most users had never heard of, holding user data. Every one of these hops adds a new set of people, and here’s the crucial part: most of them have no idea whose data they’re touching. A contractor reviewing conversations doesn’t know one of them is a bank’s merger plan. You cannot honor a responsibility you don’t know you have.
And then there’s your own side. If the bank builds its own internal AI system that records every prompt for monitoring (a sensible-sounding safeguard), the question flips inward: which of your IT staff, admins, and contractors can open that database of every sensitive question every employee ever typed? That internal log is often protected far worse than a core banking system, because it was built quickly by a small team and never got a real security review. The safeguard becomes the target.
Now the point that ties it together. At any single one of these organizations, access is controlled, logged, and governed by policy. The professional staff at a major AI lab generally do understand their responsibility — they sign agreements, take training, and work under monitoring. So each link in the chain can honestly say, “we handled it responsibly.”
But no one holds the whole chain. The bank can’t see the provider’s access list. The provider can’t fully see its vendors’ access lists. Nobody has the end-to-end map. So when data leaks, it leaks through the seams between the organizations — each one followed its own rules, and the data still got out, because responsibility was real at every individual step and owned by no one across the entire path.
That is the plain-English version of the whole problem. “Is my data safe in the logs?” has no yes-or-no answer, because there is no single “the logs” and no single person you could call to ask. There is only a long chain of strangers, each trusting the next, protected by promises that hold best exactly where the data is least sensitive to them — and matters most to you.
A corporate bank using cloud services and internet-connected LLMs faces at minimum twenty-two distinct exposure classes. Most banks’ risk registers cover three of them. The first eleven are the primary ring — the vectors through which the bank’s own AI usage leaks. The second eleven are the quieter ring: adjacent tools and pipelines nobody thinks of as “AI risk” until the data is gone.
1. Shadow AI — the unmanaged front door. Employees blocked from sanctioned tools use personal accounts on personal phones. Consumer tiers train on conversations by default unless opted out. Research cited in 2026 security reporting found shadow-AI breaches cost roughly $670,000 more than standard incidents and take 247 days to detect — because the bank cannot detect exfiltration through a channel it doesn’t know exists. Samsung’s canonical 2023 episode — three leaks in twenty days, including source code and meeting recordings — remains the template.
2. The consumer/enterprise confusion gap. The bank signs an enterprise agreement; the intern uses the free tier at home on the same documents. The contract protects the tenant, not the data class. One employee, one wrong login, and the no-training clause never applied.
3. Provider retention and the log mosaic. Covered in Part 1: even under enterprise terms, prompts persist in retention windows and safety pipelines — assembled raw material awaiting only access.
4. Prompt injection against agents — the EchoLeak class. Agents read email, documents, web pages, and wikis; any of those inputs can carry an attacker’s instructions, and the agent cannot reliably tell content from command. EchoLeak required zero clicks — Copilot processed a crafted email in the background and leaked tenant data through trusted domains with no alert at any layer. GitHub Copilot’s CVE-2025-53773 (CVSS 9.6) achieved code execution from a poisoned code comment. OWASP now finds prompt injection present in over 73% of assessed production AI deployments; Cisco’s 2026 State of AI Security reports 83% of organizations plan agentic deployment while only 29% feel ready to secure it. IBM’s 2026 report prices prompt-injection and inversion attacks at roughly $6 million per incident — 20% above the $5 million all-breach average.
5. The “lethal trifecta.” Security researcher Simon Willison’s framing: an agent with (a) access to private data, (b) exposure to untrusted content, and (c) the ability to communicate externally is a fully exploitable exfiltration engine. A bank’s document-improving, email-reading, web-searching assistant has all three by design. The presentation the agent “fixes” is simultaneously private data (a) and untrusted content (b) — and the rendering pipeline is the outbound channel (c).
6. MCP and tool-chain compromise. The Model Context Protocol that connects agents to databases, files, and terminals — now backed by Microsoft, OpenAI, Google, and Amazon — expands capability and attack surface together. In January 2026, three prompt-injection CVEs were disclosed in Anthropic’s own official Git MCP server; a malicious README was sufficient to trigger code execution or data exfiltration. When the reference implementations of the connective tissue carry vulnerabilities, a bank’s custom integrations should be presumed to carry more.
7. Vendor and supply-chain exposure. The AI provider’s own vendors — analytics (Mixpanel/OpenAI, 2025), logging, hosting, evaluation contractors, RLHF labelers — each extend the trust chain. The bank signed one contract; its data transits a dozen organizations.
8. Misconfiguration. DeepSeek’s exposed database required no attack at all — a security firm found it by scanning. Cloud AI stacks are young, complex, and assembled fast; the 2025–2026 incident record is dominated not by cinematic hacks but by buckets, dashboards, and databases left open.
9. Legal process and discovery. Every retained prompt is discoverable. A bank in litigation may find its employees’ AI conversations — including the “find our weaknesses” critiques and the “why buy us” deck — produced to opposing counsel under subpoena, with the provider legally obligated to comply.
10. Calendar, scheduling, and metadata exposure. Org-wide calendar sharing broadcasts intentions in structured form — counterparties, locations, travel modes, project codenames, meeting cadences — with zero content-inspection anywhere in the stack. AI assistants (Copilot, Gemini for Workspace) ingest calendars as core context, automatically joining scheduling metadata to emails and documents inside the tenant. Externally, the same metadata leaks through flight trackers, visitor logs, and out-of-office replies. Metadata is exempt from every confidentiality control aimed at documents, which is precisely why intelligence services have always preferred it: it’s the truth people forget they’re telling.
11. Model memorization and future-model bleed. The slowest, strangest vector: data that enters training corpora (via consumer tiers, shadow AI, or future terms changes) can be memorized and regurgitated unpredictably by later models. Amazon’s 2023 observation of outputs resembling internal data is the early form. Once in the weights, information cannot be recalled, deleted, or subpoenaed back — the only irreversible layer in the stack.
The second ring: adjacent pipelines and the offensive flip
The first eleven vectors describe data leaking out. The next eleven include the ways data leaks through systems nobody labels “AI,” and — critically — the ways an adversary can weaponize the same channels to push falsehoods in.
12. Data poisoning and “LLM grooming” — attacking a company through the models everyone else trusts. This is the offensive inversion of leakage, and it is the vector you should worry about most, because you cannot patch it inside your own walls. The threat is no longer that your data gets out — it’s that an adversary floods the AI ecosystem with fabricated information about your bank so that every chatbot, every analyst using AI research, every journalist, and every counterparty is told a lie as fact. The playbook is proven at nation-state scale: Russia’s Pravda network published roughly 3.6 million articles across 150 domains in 49 countries in 2024, engineered not to be read by humans (the sites average under 1,000 monthly visitors) but to saturate the web so that AI crawlers ingest and repeat them. When NewsGuard tested ten leading chatbots — ChatGPT, Gemini, Copilot, Claude, Grok, Perplexity, and others — they repeated the false narratives 33% of the time, and seven of the ten cited the propaganda sites as legitimate sources. The American Sunlight Project named the technique “LLM grooming”: the more often a claim appears across indexed content, the more likely models are to absorb it as truth. Now scale that down to a single company. A motivated adversary — a short seller, a hostile bidder, a disgruntled competitor, a foreign rival — can spin up dozens of plausible-looking financial-news domains and flood them with fabricated claims: that your bank is under secret regulatory investigation, that its CRE book is insolvent, that a named executive is about to be indicted, that deposits are fleeing. Most of it never needs a human reader. It needs to exist often enough that when someone asks an AI “is [your bank] financially healthy?” the model hedges, or worse, repeats the smear. Because the false narrative and the assembled true mosaic feed the same models, an attacker who has also harvested fragments of your real strategy can craft disinformation that is corroborated by genuine detail — the most credible lie is the one wrapped around a true fact.
13. Training-data backdoors — poisoning the model itself. Grooming pollutes what models read from the web; backdooring corrupts what they learn in training. A landmark October 2025 study by Anthropic, the UK AI Security Institute, and the Alan Turing Institute found that as few as 250 malicious documents can implant a hidden backdoor in a model regardless of its size — overturning the assumption that attackers must control a percentage of training data. A separate study in Nature Medicine showed that replacing just 0.001% of training tokens with misinformation produced measurably more error-prone models that still passed every standard benchmark — the poison was undetectable by normal testing. For a bank building or fine-tuning its own custom model on scraped or third-party data (the fastest-growing deployment pattern), this means the model it trusts most could carry a trigger that leaks data or produces attacker-chosen outputs on command. And “anyone can create online content that might eventually end up in a model’s training data” — the barrier to entry is a few hundred web pages.
14. RAG and vector-store poisoning. Banks increasingly connect AI to internal knowledge via retrieval-augmented generation — the model answers from a vector database of the bank’s documents. If an attacker (or a careless integration) gets poisoned content into that store — a doctored policy memo, a fake precedent, a booby-trapped PDF — every employee who queries the assistant is served the corruption as authoritative internal truth. The lethal-trifecta and MCP risks from vectors 4–6 apply here in reverse: the retrieval layer is both an ingestion point for bad data and, via injected instructions, an exfiltration trigger.
15. Third-party AI features silently embedded in ordinary software. The bank vets “AI tools” and misses the AI now baked into everything else: the CRM that added a summarization feature, the video-conferencing platform that auto-transcribes, the email client with smart-compose, the PDF reader with a chat assistant, the helpdesk with an AI triage bot. Each may route content to a model under terms the bank never reviewed. The Otter.ai class of incident is instructive — meeting-transcription bots that silently join calls, record, and retain sensitive discussions, sometimes emailing transcripts to unintended recipients. The bank didn’t adopt an AI strategy for these; the vendors did it for them.
16. Meeting transcription and the “AI notetaker” problem. A specific, worsening case of #15. Executives now routinely admit AI notetakers (Otter, Fireflies, Copilot, Gemini, Zoom AI) to confidential calls — board meetings, deal negotiations, legal strategy sessions. The transcript is a verbatim record of the most sensitive conversation in the building, generated automatically, stored in the cloud, often shared to attendees by default (including external participants), and searchable forever. A single notetaker in a merger negotiation captures both sides’ positions in one file.
17. Browser extensions and unofficial “wrapper” apps. Malicious or over-permissioned browser extensions marketed as AI helpers actively harvest ChatGPT and DeepSeek conversations, capturing prompts and responses and transmitting them to attacker servers — often without the user realizing the session is intercepted. Employees install these by the thousands. The bank’s enterprise AI contract is irrelevant when the leak is a Chrome extension between the keyboard and the browser.
18. Inference from usage patterns and query metadata. Even without reading content, a provider (or an attacker with telemetry access) can infer sensitive facts from how the bank uses AI: a sudden spike in the M&A team’s queries, a burst of severance-law questions from one department, unusual after-hours activity on a specific credit. Traffic analysis is a mature intelligence discipline; AI usage generates rich traffic. The metadata leaks the tempo of the bank’s intentions even when the words are protected.
19. Cross-tenant and model-isolation failures. Enterprise assurances rest on tenant isolation — your data stays in your instance. But software isolation fails: misconfigurations, shared-cache bugs, and memory-feature bleed have all produced cross-user exposure in cloud systems. The 2023 ChatGPT bug that showed users other users’ chat titles was a preview. As “memory” and “org knowledge” features grow more powerful, the blast radius of an isolation failure grows with them.
20. Insider misuse of AI as an exfiltration laundering tool. A departing employee who emails themselves the client book triggers DLP. The same employee who asks the AI assistant to “summarize our top 50 relationships with contact details and deal history” and then copies the tidy output may not — the AI has laundered a bulk export into a “normal” query. Agentic tools with broad data access turn a single insider prompt into a curated intelligence package.
21. Prompt and output logging in the bank’s own middleware. Banks building internal AI gateways often log every prompt and response for monitoring and audit — sensibly. But that log is now a new, concentrated crown-jewel database: every sensitive question every employee ever asked, in plaintext, in one place, protected only as well as that internal system is. The safeguard becomes the target.
22. AI-generated content as an inbound attack and reputational vector. Beyond grooming, adversaries use generative AI to manufacture convincing fakes aimed at the bank: deepfake audio of an executive authorizing a wire (already a documented multimillion-dollar fraud pattern), synthetic “leaked documents” designed to move the stock or spook depositors, and AI-written phishing tuned to individual employees using the mosaic harvested elsewhere. The same technology that leaks the bank’s truth outward manufactures falsehoods aimed inward.
Layer the taxonomy over the scenarios: the mosaic accumulates through vectors 1–3, is timestamped and annotated by vector 10, and is quietly widened by the adjacent-tool ring of 14–21; the presentation and buyout deck travel through 4–8; everything persists into 9; and the fragments risk immortality through 11. The June 4th Chicago meeting illustrates the defensive stack in miniature: the calendar entry names the counterparty and the date, the deck-review request supplies the substance, the AI assistant joins them by design, and the log retains the union. But vectors 12, 13, and 22 flip the whole model on its head — there, the bank isn’t leaking anything; an adversary is pushing falsehoods in, poisoning the models the bank and its counterparties rely on, so that the market is told the bank is failing whether or not it is. Leakage and disinformation are the same pipeline run in opposite directions. The bank never decided to disclose anything, and never decided to be lied about. The architecture decided both for it.
What a Prudent Bank Actually Does
The answer is not abstinence — banks that ban AI simply push usage into shadow AI, the worst vector of all. The answer is architectural honesty:
Segment by blast radius. Public-data tasks (market summaries, coding help) can use enterprise cloud tiers with negotiated zero-data-retention. Anything touching MNPI, client data, credit files, deal work, or strategy belongs on infrastructure the bank controls — on-premises open-weight models (the Morgan Stanley “build, don’t block” pattern, now achievable at mid-market cost), private-cloud deployments in the bank’s own tenant, or air-gapped enclaves for the truly sensitive.
Treat documents as the crown jewels, not chats. Policy attention obsesses over what employees type; the catastrophic payloads are what they upload. Presentation, spreadsheet, and data-room material should be technically blocked from external AI endpoints, with a local alternative provided so the block doesn’t breed shadow usage.
Assume the mosaic. Governance should evaluate AI exposure the way intelligence agencies evaluate publication: not “is this prompt sensitive?” but “what does the corpus of our prompts reveal?” That analysis is sobering exactly once — and then it changes the architecture.
Treat the calendar as a classified feed. Sensitive meetings get codenames, not counterparty names; deal travel gets booked outside the shared system; free/busy visibility replaces full-detail visibility by default; and — critically — the bank decides deliberately which AI assistants may ingest calendar data at all, because an assistant with calendar context has the mosaic’s index. “Strategic discussion with [CEO name]” should never appear in any system whose contents the bank cannot enumerate.
Ban silent AI in adjacent tools. The vetting process must cover AI features embedded in non-AI software — transcription bots, CRM summarizers, meeting notetakers, browser extensions. Default rule: no AI notetaker in any confidential meeting, and no browser extension touching AI sessions on managed devices.
Monitor the disinformation surface, not just the leak surface. Because vectors 12–13 and 22 push falsehoods in rather than pull data out, the bank needs the mirror image of DLP: monitoring what AI systems and the web say about the institution. Periodically query major chatbots about the bank’s health, executives, and reputation; watch for fabricated-domain clusters and AI-repeated smears; and prepare an incident-response playbook for AI-laundered disinformation, deepfakes, and synthetic “leaks” the same way the bank prepares for a cyber breach. In a deposit-taking institution, a widely repeated AI falsehood about solvency is not a PR problem — it is a bank-run risk.
Constrain the trifecta. Any agent with private-data access should lose either untrusted-content exposure or external communication. EchoLeak’s lesson is that trust boundaries are security boundaries; an agent that reads inbound email should not also hold the keys to SharePoint and an outbound channel.
Price the promise correctly. When the vendor says “we don’t train on your data,” the correct response is: “We believe you — and our architecture will be designed so that we never have to.” Trust is a fine thing to have and a terrible thing to depend on. Physics doesn’t renew annually. Contracts do.
🔎 Brian’s Take
“I spent years watching institutions with impeccable long-term incentives make short-term choices that destroyed them, so the ‘no AI lab would ever risk enterprise trust’ argument lands differently on me than it does on a procurement officer. Of course they wouldn’t — right up until a funding crisis, an acquisition, a desperate quarter, or a middle manager with a growth target decides the fine print has some flex in it. Liar loans were irrational for the banks writing them. They wrote them anyway. My rule from the allocation business: never underwrite a risk whose downside is irreversible on the strength of a counterparty’s ongoing self-restraint. Data in someone else’s logs is exactly that risk — and data in someone else’s training run is irreversible in the strictest sense of the word. The mosaic analysis in this piece is the part I’d staple to every bank board deck: your institution is disclosing continuously, in fragments, through its most diligent employees, and the assembled picture sits outside your walls priced at other people’s discipline. But here’s the part that genuinely changed how I think about it — the pipe runs both ways. It’s not just that your secrets leak out; it’s that an adversary can pump lies in. A short seller who both harvests fragments of your real strategy and floods the web with fabricated claims about your solvency can get the world’s AI systems to repeat a smear that’s half-true and therefore devastating. For a bank — an institution that lives and dies on confidence — an AI-amplified falsehood about your balance sheet isn’t a reputational nuisance, it’s a run. The banks that banned ChatGPT in week one weren’t Luddites. They were the only ones who read the architecture instead of the marketing.”
Frequently Asked Questions
Can AI providers technically access enterprise prompts? Yes. Data must arrive in processable form to be processed. No-training clauses, retention limits, and access controls are contractual and organizational safeguards — enforceable, but categorically different from technical impossibility.
What is the mosaic effect in AI data leakage? It is the reconstruction of confidential conclusions from many individually harmless fragments. Hundreds of employee prompts, each innocuous alone, can collectively reveal deals, layoffs, credit problems, and strategy — and DLP tools inspecting messages one at a time cannot detect it.
What was EchoLeak? EchoLeak (CVE-2025-32711) was the first zero-click prompt-injection data-exfiltration attack on a production enterprise AI assistant, Microsoft 365 Copilot. Hidden instructions — in one documented variant embedded in PowerPoint speaker notes — caused the assistant to leak tenant data through trusted domains with no user interaction.
Why did major banks ban ChatGPT? JPMorgan, Bank of America, Citigroup, Goldman Sachs, Deutsche Bank, and Wells Fargo restricted or banned ChatGPT beginning in February 2023 over data-leakage and regulatory concerns; BNY Mellon cited the impossibility of meeting fiduciary data-handling duties through third-party training pipelines.
Is organization-wide calendar sharing a form of data leakage? Yes. Shared calendars broadcast structured intent — counterparty names, locations, private travel, project codenames, and meeting cadences — with no DLP inspection, and AI assistants like Copilot and Gemini ingest calendar data as core context, automatically joining it to emails and documents. A single entry like “private aviation to Chicago, June 4, meeting with [CEO name] — strategic discussion” can reveal an impending merger on its own; corporate jet movements alone have historically predicted M&A announcements.
Can an adversary attack a company by feeding AI false information? Yes. Two techniques stand out. “LLM grooming” floods the web with fabricated content so AI models absorb and repeat it — Russia’s Pravda network published ~3.6 million articles in 2024 and got ten leading chatbots to repeat its falsehoods 33% of the time. “Data poisoning” corrupts training data directly — an Anthropic/UK AI Security Institute/Alan Turing study found just 250 malicious documents can backdoor a model of any size. Scaled to one company, a short seller or hostile bidder could flood plausible-looking financial-news domains with false claims about a bank’s solvency or an executive’s conduct, engineering AI systems to repeat the smear as fact.
How much do AI-related breaches cost? IBM’s 2026 Cost of a Data Breach Report puts the average breach at $5 million — up 12% year-over-year — with prompt-injection and inversion attacks on AI tools averaging roughly $6 million, and shadow-AI breaches costing about $670,000 more than standard incidents.
About the Author: Brian French
Brian French is a former institutional money manager and analyst with a career spent evaluating companies, industries, and capital flows on behalf of institutional clients. His background spans equity research, portfolio management, and macro-driven sector analysis — experience he now applies to dissecting technology risk, financial institutions, and long-horizon investment themes. Known for a “follow the capital, not the headlines” approach, Brian focuses on the durable structural forces — incentive design, counterparty risk, and institutional behavior under pressure — that determine which safeguards hold and which fail. He writes and speaks about the intersection of finance, technology, and economic geography.
This article is for informational purposes only and does not constitute legal, compliance, security, or investment advice. Consult qualified counsel regarding data-protection obligations applicable to your institution.
Resources and Sources
- IBM — Cost of a Data Breach Report 2026 (via Cybersecurity Dive) — $5M average breach cost (+12% YoY), ~$6M average for prompt-injection and inversion attacks, 85% of breached organizations increasing governance spend. (cybersecuritydive.com)
- Reco — “AI & Cloud Security Breaches: 2025 Year in Review” — EchoLeak mechanics (email → Copilot ingestion → OneDrive/SharePoint/Teams extraction → exfiltration via trusted domains, zero clicks), IBM 2025 baseline ($4.44M). (reco.ai)
- Sysdig — “The Comprehensive Guide to Prompt Injection Attacks in 2026” — EchoLeak case study, Simon Willison’s “lethal trifecta,” Cursor MCP configuration attack, attacker/defender economics. (sysdig.com)
- Vectra AI — “Prompt injection: types, real-world CVEs, and enterprise defenses” — CVE-2025-32711 (CVSS 9.3) technical breakdown, CVE-2025-53773 GitHub Copilot RCE (CVSS 9.6), OWASP 73% production-deployment finding, Cisco State of AI Security 2026 (83% deploying, 29% ready). (vectra.ai)
- TechStoriess — “AI Agent Security Practices 2026” — EchoLeak via PowerPoint speaker notes, shadow-AI breach economics ($670K premium, 247-day detection), 88% incident rate vs. 82% executive confidence gap. (techstoriess.com)
- BlueRadius — “AI Cybersecurity Incident Report 2026” — MITRE ATLAS v5.1.0 agent-attack techniques, malicious browser extensions harvesting ChatGPT/DeepSeek conversations, OpenAI–Mixpanel third-party exposure disclosure, Samsung three-incidents-in-twenty-days pattern. (blueradius.io)
- Cyber Desserts — “Prompt Injection Attacks: Examples, Techniques, and Defence” — January 2026 CVEs in Anthropic’s official Git MCP server (CVE-2025-68143/68144/68145), MCP attack surface across Microsoft/OpenAI/Google/Amazon ecosystems. (blog.cyberdesserts.com)
- PurpleSec — “Data Exfiltration Via AI Prompt Injection” — Salesforce “ForcedLeak” hidden-prompt vulnerability, direct vs. indirect injection taxonomy. (purplesec.us)
- Techglock — “Prompt Injection: The #1 AI Threat in 2026” — Lethal-trifecta exploitability framework applied to enterprise agents. (techglock.com)
- Forbes — “Workers’ ChatGPT Use Restricted At More Banks — Including Goldman, Citigroup” (Feb 2023) — Bank of America unauthorized-apps listing alongside WhatsApp, Citigroup/Goldman third-party software restrictions, Deutsche Bank access disablement. (forbes.com)
- Moveo.AI — “Companies Banning ChatGPT (2026): The Enterprise Security List” — JPMorgan, Deutsche Bank, Wells Fargo, BofA, Citi, Goldman restrictions; Morgan Stanley “build, don’t block” internal deployment; BNY Mellon fiduciary rationale. (moveo.ai)
- The Telegraph via TipRanks — “JPMorgan restricts use of ChatGPT among staff” — Regulatory-action concern over shared financial information. (tipranks.com)
- Fortune / Yahoo Finance — “Apple, Goldman Sachs, and Samsung among growing list of companies banning ChatGPT” — Samsung April 2023 leak details (internal code, meeting recordings) and subsequent internal-AI pivot; Amazon warnings after outputs resembling internal data. (finance.yahoo.com)
- Visbanking — “Major Banks Restricting Use of ChatGPT” — Amazon confidential-data warnings, JPMorgan third-party-controls framing. (visbanking.com)
- Wiz Research — DeepSeek exposed ClickHouse database disclosure (January 2025) — Publicly accessible database containing over one million log lines including chat histories, discovered via routine scanning. (wiz.io)
- New York Times Co. v. OpenAI — preservation order coverage (2025) — Court-ordered retention of user conversations overriding deletion policies, establishing legal-process precedence over privacy commitments. (court filings; widely reported)
- Simon Willison — “The Lethal Trifecta” (2025) — Original framing of private-data access + untrusted content + external communication as the complete agent-exfiltration condition. (simonwillison.net)
- CNBC / Reuters — Occidental jet in Omaha coverage (April 2019) — Corporate-jet tracking preceding Berkshire Hathaway’s $10 billion Anadarko-bid investment; illustrative of aviation metadata anticipating deal announcements. (cnbc.com)
- Academic and market research on corporate jet tracking — Studies and hedge-fund data products (e.g., aviation-data feeds formerly distributed via Quandl) demonstrating that flights between headquarters cities predict M&A activity. (ssrn.com; quandl/Nasdaq Data Link archives)
- Microsoft — Microsoft 365 Copilot documentation; Google — Gemini for Workspace documentation — Calendar, email, and document ingestion as core assistant context within the enterprise tenant. (learn.microsoft.com; workspace.google.com)
- NewsGuard — “Russian Propaganda Has Now Infected Western AI Chatbots” audit (March 2025) (via Forbes, The Hill, Yahoo News) — Pravda network’s 3.6 million articles across 150 domains in 49 countries; ten leading chatbots repeating false narratives 33% of the time; seven citing Pravda sites as sources; 92 disinformation articles cited across models. (forbes.com; newsguardtech.com)
- American Sunlight Project — LLM grooming report (February 2025) — Definition and mechanics of “LLM grooming”; correlation between narrative saturation and model absorption; 97 Pravda domains publishing ~20,000 articles in 48 hours. (americansunlight.org)
- DFRLab — “Pravda in the pipeline: Early evidence of state-adjacent propaganda in AI training data” (April 2026) — AI poisoning co-opting the retrieval layer of LLMs. (dfrlab.org)
- Anthropic, UK AI Security Institute & Alan Turing Institute — “A small number of samples can poison LLMs of any size” (October 2025) — 250 malicious documents sufficient to backdoor models from 600M to 13B parameters regardless of training-data volume. (anthropic.com)
- Nature Medicine — training-data medical-misinformation study (2025) (via Dell Technologies analysis) — Replacing 0.001% of training tokens with misinformation produced measurably more error-prone models that still passed standard benchmarks; ~100 poisoned models found on Hugging Face. (dell.com; nature.com)
- Reporting on AI notetaker and transcription exposure (Otter.ai class incidents) and malicious AI browser extensions (via BlueRadius AI Cybersecurity Incident Report 2026 and related coverage) — Silent meeting-transcription capture and browser-extension harvesting of ChatGPT/DeepSeek sessions. (blueradius.io)