Learn / Statistics

Shadow AI statistics, from measurement

Every number on this page comes from one place: our live register of classified AI domains and our ongoing review of vendor terms. No third-party surveys, no invented percentages, and each stat carries its implication.

0AI tool domains classified in the live register
0new domains screened per day for AI signals
0vendor terms documents reviewed for training verdicts
0functional categories in the classification
The headline stats

Six numbers, each with its consequence

A statistic without an implication is decoration. Each tile pairs the measurement with what it means for your network.

85.5%

of AI tools say nothing about training in their public terms

The Policy Silence Index finding, from 13,000+ terms reviewed. Silence, not disclosure, is the market's default state.

Implication: "read the terms before using it" fails as staff guidance. Most terms do not answer the question being asked.
700+

tools train on customer data by default

Confirmed train-by-default verdicts, typically in consumer and free tiers, where shadow usage concentrates.

Implication: any unaudited network of size almost certainly touches several. The question is which ones and who.
3,900+

tools mapped to the model provider actually behind them

Wrapper and white-label products resolved to their backing providers, with OpenAI, Anthropic and Google the most common.

Implication: a niche tool's real data path often ends at a major provider. Risk follows the backend, not the brand.
20,399

classified AI domains, maintained daily

Each with category, risk level, sovereignty flag, abusive-purpose flag and a dated training verdict.

Implication: an approved list of twenty tools governs a market ten thousand times its size. Inventory is not optional.
~300k

domains checked every day

The daily screening stream from a 120 million domain corpus that keeps the register current.

Implication: static blocklists and annual reviews mislabel the present. Freshness is a property, not a feature.
18

functional categories, 100+ subcategories

From text and code assistants to agents, transcription and model infrastructure.

Implication: "AI" is not one policy decision. A transcription bot and an image upscaler deserve different verdicts.
The verdict split

What vendor terms say, when they say anything

Across the reviewed terms, four verdict states cover the market. The bar shows the shape that matters: silence dominates.

Not stated: 85.5%
Terms silent on training Trains by default Trains unless opted out States it does not train

How to read the silent majority

  • Silence is not a safe harbor: it means the vendor has not committed either way, and past prompts have no stated protection.
  • Auditors treat "we used tools with silent terms, knowingly" very differently from "we did not know". Measurement converts the second into the first.

How to read the explicit tail

  • The explicit verdicts are where decisions are easy: contract with the "does not train" tools, negotiate or block the rest.
  • Enterprise tiers frequently flip a consumer verdict; the register tracks tier-level differences where vendors state them.

Segment percentages beyond the 85.5% silence figure shift as reviews continue; the bar shows the current shape rather than a frozen claim. Your own network's split is what the audit measures.

Provider concentration

Many brands, few backends

The wrapper economy means tool counts overstate diversity. The provider mapping shows what your data actually reaches.

What the mapping shows

  • 3,900+ tools resolve to an identifiable model provider behind the product.
  • OpenAI, Anthropic and Google back the largest shares, with a long tail of open-weight and regional providers behind the rest.
  • Rebrands and thin wrappers inherit their backend's data-handling reality regardless of their own marketing.

Why it changes decisions

  • Approving "one chatbot" while blocking five wrappers of the same backend is coherent policy, not inconsistency, when contracts differ.
  • Sovereignty review keys off the backend's jurisdiction as much as the wrapper's.
  • In the audit, provider context feeds the risk level you see per tool.
Methodology

Where these numbers come from

Statistics pages earn trust by explaining themselves. Ours reduce to three processes.

Domain screening

Roughly 300,000 new domains a day from a 120 million domain corpus are checked for AI signals, feeding the register's growth and its freshness.

Terms review

Vendor terms are read for training positions, producing dated verdicts. 13,000+ reviewed to date; the 85.5% silence rate is this process's headline output.

Provider resolution

Products are mapped to the models and providers behind them, resolving the wrapper economy into actual data paths.

Presentation rule for this page: our analysis, our database, stated as such. Nothing here is quoted from third-party surveys, and nothing is a projection.

Using the numbers

Base rates vs your rates

Market statistics set expectations; your audit sets facts. The comparison is where the insight lives.

Market base rateYour audit's counterpartIf yours is worse
85.5% of tools silent on trainingShare of your found tools with "Not stated" verdictsYour stack skews consumer-tier; prioritize enterprise migrations.
700+ train-by-default tools existYour "Trains by default" countAbove a handful, treat as an active leak, not a statistic.
Wrappers concentrate on few providersDistinct tools vs distinct backends in your listConsolidate contracts at the backend level and cut the wrapper tail.
Tools span 18 categoriesCategories present in your networkCategories with no owner in the room need one before next quarter.

Getting your rates costs one log export: uploads are read once and discarded, and the free preview returns your category and risk totals in minutes.

Turn base rates into your rates

The sample report shows how these statistics render as your network's own tiles and verdict columns. Then one export replaces the sample data with yours.

Open the sample report
Citation rules

Quoting these numbers in your own documents

These statistics end up in board decks and policy drafts. Three rules keep the citations defensible.

Date them

The register grows daily, so quote the count with the date you saw it. This page always shows the live figure.

Scope them

"Of AI tools reviewed in a 13,000+ terms analysis" is accurate. "Of all AI tools everywhere" is not what any measurement can claim.

Pair them with your audit

Market rates argue urgency; your report argues specifics. Decks that carry both survive questioning; decks with only base rates do not.

Honesty inventory

Numbers you will not find on this page

The shadow AI content economy runs on unverifiable percentages. We decline four popular ones.

"X% of employees use AI secretly"

Self-reported survey numbers about secretive behavior are structurally unreliable. Your per-user table measures your answer instead.

"Shadow AI costs companies $X billion"

Loss extrapolations stack assumptions we cannot audit. We publish counts we can regenerate on demand.

"X% of prompts contain sensitive data"

We never read prompt content, by design, so we will not quote anyone who claims to have measured it at market scale.

"AI adoption grew X% this year"

Whose adoption, measured how? Our register growth is a supply-side signal we can state; demand claims need your logs.

FAQ

Statistics questions

Where does the 85.5% silence figure come from?

From our review of 13,000+ AI vendor terms documents: 85.5% take no stated position on training with user data.

How current is the register count?

Live. The number on this page is fetched from the register at page load, and screening adds to it daily.

Why do you not quote industry survey percentages?

Because we cannot verify them. Everything here is our own measurement, stated with its scope, which is the standard we would want from anyone else.

Do these numbers apply to my organization?

They are base rates, not your rates. A free audit preview measures your own network's counterparts in minutes.

What share of tools train on user data?

700+ confirmed train-by-default tools, plus opt-out tools where past prompts may already be used. Combined with silent terms, unprotected states dominate the market.

Which model providers are most common behind AI tools?

In our mapping of 3,900+ tools to their backing providers, OpenAI, Anthropic and Google appear most often.

The market's numbers are here. Yours are one upload away.

Run the free preview and put your own network's statistics next to these.

Run the free audit