How to Detect Shadow AI

Shadow AI is detected by combining several partial data sources the organization already holds — expense and procurement records, identity and access logs, software and SaaS inventory, admin data from sanctioned AI platforms, network and browser telemetry where it exists, and structured conversations with teams — into one reconciled list of AI in use. No single source is sufficient, and any method that claims completeness is overselling.

This article describes what enterprises can do with the tooling they have. It is deliberately separate from what Midgentic contributes, which is covered honestly at the end.

Start by defining what you are looking for

"AI tool" is ambiguous enough to derail a discovery effort in its first week. Agree a working definition before you collect anything. A serviceable one: any product, service, API, agent, or feature that uses a machine-learning model to generate, summarize, classify, or decide on work content. Then decide explicitly whether AI features inside already-owned software are in scope. They usually should be, and they are usually forgotten.

Seven detection sources, and what each one actually tells you

1. Expense and corporate card data

The highest-yield first pass in most organizations. Individual AI subscriptions are small, recurring, and charged to a recognizable set of merchant names. Search expense claims and card statements for known AI vendors and for repeating low-value software charges.

Reveals: paid individual and team tools. Misses: free tiers, trials, and anything bought by a contractor.

2. Procurement, contracts, and vendor management

Procurement records show what went through the formal path, which by definition is not shadow. Their value is as the control list: everything discovered elsewhere that does not appear here is a candidate.

Reveals: the sanctioned baseline. Misses: everything below the approval threshold.

3. Identity and access data

If your identity provider brokers sign-in, its application logs show which external services people authenticate to — including AI services signed up for with a work account. Reviewing new applications appearing in identity logs is one of the few genuinely continuous detection methods available.

Reveals: services accessed with corporate identity, including free tiers. Misses: personal-account and unauthenticated use.

4. Software and SaaS inventory / discovery tooling

SaaS management and endpoint inventory tools maintain application lists and can be filtered for AI vendors. Browser extension inventory deserves specific attention — extensions are a common and rarely reviewed route for work content to reach a model provider.

Reveals: installed and connected applications. Misses: anything used purely in a browser on an unmanaged device.

5. Admin data from sanctioned AI platforms

A different but essential question: of the AI you already approved, how much is actually used? Assigned seats with no activity are not Shadow AI, but they distort the same decisions, and the data is straightforward to obtain from platform admin reporting. Covered in how to measure AI adoption.

Reveals: real utilization of sanctioned tools. Misses: everything unsanctioned.

6. API, cloud, and platform consumption

Model API usage is the least visible category because it produces no seats, no application in the identity log, and often a single consolidated cloud bill. Review cloud marketplace charges, model-provider billing accounts, and internally built services that call external models on a schedule.

Reveals: engineering-led AI workloads and agents. Misses: nothing in this category if billing is reviewed thoroughly — but it requires someone who can read the bills.

7. Structured conversations with teams

Technical sources find tools; people explain workflows. A short structured survey or a fifteen-minute conversation with each department head — asking what AI they use, for which task, who pays, and what would break if it stopped — reliably surfaces things no log contains. Response quality depends entirely on whether the exercise is framed as discovery rather than audit.

Reveals: workflow context, free tools, contractor tooling, informal agents. Misses: whatever people forget or prefer not to mention.

Reconciling sources into one list

Detection produces overlapping fragments, not a list. Reconcile them into a single record per AI system with, at minimum: name and vendor, how it was discovered, department, approximate user population, known cost, business purpose, owner, and current governance status. That record is the beginning of an enterprise AI inventory, which is where discovery should always terminate — otherwise the exercise is repeated from scratch next year.

Record confidence alongside each entry. "Confirmed with the owner" and "one card transaction six months ago" should not sit in the same column without distinction.

Turning detection into an operating cadence

  • Re-run expense and identity scans monthly; they are cheap once the queries exist.
  • Refresh departmental conversations twice a year, or after any significant reorganization.
  • Review AI features newly enabled by incumbent vendors at each renewal.
  • Track the trend in newly discovered tools; a rising rate means the sanctioned path is too slow.
  • Treat every discovery as a routing decision: onboard, consolidate, or retire — never simply log it.

What this method cannot do

Be clear with stakeholders about the limits. None of these sources sees personal-device, personal-account use. Free tiers leave a very light trace. And detection tells you a tool exists, not whether sensitive data went into it — that question belongs to security and data protection tooling, not to an inventory exercise.

How Midgentic contributes

Midgentic does not perform network discovery, endpoint scanning, traffic inspection, or content analysis, and it is not a CASB or DLP product. Its role is downstream of detection and is where most programs stall: it maintains the reconciled picture. Automated connectors read seats, active users, and activity from major AI platforms; every other provider you discover can be recorded manually or imported by CSV, with cost and license terms attached. Each figure is labeled with its source, and coverage gaps are shown rather than smoothed over.

The result is that discovery work becomes a living portfolio instead of a spreadsheet that ages out. See Shadow AI, enterprise AI visibility, and the integrations catalog for exactly what is ingested from each platform.

Frequently Asked Questions

How do companies detect Shadow AI?

By combining sources they already hold: expense and card data, procurement records, identity and access logs, SaaS and endpoint inventory, admin data from sanctioned AI platforms, cloud and model-API billing, and structured conversations with departments. The findings are then reconciled into a single AI inventory.

Which detection method finds the most Shadow AI?

In most organizations, expense and corporate card data yields the most in the first pass, because individual and departmental AI subscriptions are paid, recurring, and attributable to recognizable vendors. Identity logs are the best source for ongoing detection.

Can Shadow AI be detected automatically?

Parts of it. Identity logs, SaaS discovery tooling, and expense queries can run continuously. Personal-account use, free tiers, and workflow context still require conversations with teams, so no automated method should be presented as complete.

Does detecting Shadow AI require network monitoring?

No. Network and proxy telemetry helps where it already exists, but a substantive inventory can be built entirely from financial, identity, procurement, and platform admin data plus departmental interviews.

What should happen after Shadow AI is detected?

Each discovered tool should be routed to a decision — onboard and govern, consolidate onto an existing platform, or retire — and recorded in a maintained AI inventory with an owner. Detection that ends in a report changes nothing.

See your AI ROI in one executive view

Try Midgentic free and bring AI adoption, spend, governance, and ROI together across your AI ecosystem.