Crescent Capital Advisors· Technology

The Shadow AI Inventory Checklist

July 16, 2026 · AI Governance · CISO · PE Value Creation

Sujit Maharana · Operating Partner, Crescent Capital Advisors

Ask a portfolio company's CTO for a list of every AI system running inside the business, and the answer is almost never a list. It's a guess, followed by a promise to check. That gap is the actual risk. You cannot govern what you cannot see, and at the next diligence process or board review, "we're looking into it" is not an answer anyone wants to give.

Shadow AI is the ordinary consequence of every SaaS vendor bolting a model onto their product roadmap and every employee with a laptop and a free-tier account. None of it required a purchase order. Most of it never touched procurement, security review, or the CTO's desk. It accumulated the way shadow IT always has, except this time the tool can read sensitive data, generate output that looks authoritative, and act on production systems without anyone signing off.

The fix isn't a policy memo. It's an inventory: the first, unglamorous step of building AI governance that holds up under scrutiny. Below is the checklist for running that sweep.

Where shadow AI hides

SaaS features with embedded LLMs. The support desk, the marketing platform, the analytics tool: most of the SaaS stack a company already pays for has quietly shipped an AI feature in the last eighteen months, often defaulted on. Nobody re-reviewed the data processing terms when it did.

Browser extensions and individual-seat tools. ChatGPT, Copilot, Claude, and a dozen others, used by staff on personal or unmanaged accounts. These are the hardest to see because they leave no line item and no admin console: just a browser extension and a habit.

AI baked into vendor products. The ERP, the CRM, the HR platform: vendors are adding AI-assisted features into contracts that were signed before those features existed. The vendor relationship is sanctioned. The AI feature inside it usually isn't.

Internal scripts, notebooks, and agents wired to production data or APIs. Engineering and data teams build things fast. A notebook that calls a model API against a customer database, or an internal agent that was supposed to be a proof of concept and quietly became load-bearing, rarely appears on any architecture diagram.

Model APIs called directly from the codebase. Grep the repo. An API key for a model provider, called from a service that was never scoped as "an AI system" in anyone's documentation, is one of the most common findings in a first sweep.

Data pipelines feeding any of the above. The system itself is only half the picture. What data flows into it (and whether that includes PII, customer records, or anything contractually restricted) is the other half, and it's usually undocumented in the same places the systems are.

What to capture per system

A list of tool names is not an inventory. For each system found, the catalog needs:

  • Purpose. What business function it serves and who asked for it.
  • Data touched. What it ingests or outputs, flagged explicitly if that includes PII or other sensitive data.
  • Business owner. The person accountable for it, not the person who happened to set it up.
  • Approval status. Sanctioned, informally tolerated, or genuinely unknown until this sweep found it.
  • Vendor or provider. Who operates the model, and whether it's a name-brand provider or a smaller tool with thinner data practices.
  • Whether the vendor trains on your data. This single term, buried in most vendor agreements, determines whether proprietary or customer data is leaking into someone else's model weights.
  • A rough risk tier. Not a full assessment, just enough to know what needs attention first versus what can wait for the next pass.

How to run the sweep this week

This doesn't require a six-month engagement to start.

  1. Pull the SaaS and expense list. Finance and procurement records surface most of the sanctioned and semi-sanctioned tools in an afternoon.
  2. Survey the teams. A short, direct question to department heads ("what AI tools does your team actually use, including the ones nobody approved") gets further than a security bulletin ever will.
  3. Scan the codebase for model-API calls. A search for the major providers' SDKs and API key patterns across every repo takes an engineer a day, not a quarter.
  4. Check identity and SSO logs. Application access logs will show sign-ins to AI tools that never went through procurement, including free-tier and personal-account usage on company devices.

Cross-reference the four, and the list that comes back is almost always longer than expected. That's normal, and it's the point: the sweep is designed to surface what's invisible, not confirm what was already known.

The inventory is step one

Building the list doesn't fix anything by itself. It's the precondition for fixing anything. Once every system is named, owned, and tagged with what data it touches, the estate can be scored against a maturity model, the EU AI Act where it applies, and whatever a board or an acquirer's diligence team is going to ask next.

That's the natural next step: run the AI Governance Quick Scan for a fast read, or the full AI Governance Readiness Assessment to score the estate once it's visible. For companies that need the inventory built and the roadmap that follows it, the AI Governance Program runs that work end to end, starting with Discover.

Working through a version of this?

A 30-minute working conversation - no deck, no pitch. Bring the situation you're sitting with.