Surveys lie, app inventories age, agents miss the browser. Your network logs already record every AI tool anyone reached. This page explains how we turn that record into a defensible inventory.
Whatever the tool does, it resolves a hostname. That is the one signal no AI product can avoid emitting.
An employee can skip a survey or shrug at an interview. Their browser still resolved chatgpt.com at 9:14.
Your logs hold weeks or months of the past. Detection starts with a backlog, not from zero on install day.
The DNS filter, proxy or firewall you already run is the sensor. There is nothing to deploy and nothing to maintain.
The method in one sentence: take the hostnames your network already logged, and match them against a live register of 20,399 classified AI tool domains.
Each alternative answers a different question. None of them answers "what did our people actually reach".
A fuller comparison, including when each method genuinely wins, is on the detection methods page.
The match is only as good as the list. The register is a maintained commercial product, not a scraped blocklist.
Our analysis finds 700+ tools that train on customer data by default. Each verdict is dated to when the terms were checked.
85.5% of AI tools say nothing about training in their public terms. Silence is recorded as its own verdict, never guessed over.
3,900+ tools are mapped to the model provider actually behind them. A rebranded wrapper inherits the risk of its backend.
Logs record hostnames at every depth. The matcher starts specific and walks up until it finds the registered tool.
Two real behaviors: walking up to the registered domain, and stopping at an exact subdomain so a whole platform is not condemned for one AI feature.
AI tools scatter their traffic across api, cdn, auth and regional subdomains. Without walk-up, an audit undercounts by whole tools.
The reverse mistake is worse: matching google.com because gemini.google.com is AI would flag all of Google. Depth-aware entries prevent that.
Proxy logs carry full URLs. The parser reduces them to hostnames before matching, so paths and query strings never influence a verdict.
Between your upload and the report, five things happen to the data. All five are visible in the output.
| Stage | What happens | Where you see it |
|---|---|---|
| Format detection | CSV headers, key=value syslog or plain lists are recognized automatically. | The scope line names the detected format. |
| Parsing | Hostname, identity and timestamp fields are extracted. Unparseable lines are skipped and counted. | Lines read vs lines parsed in the report metadata. |
| Deduplication | Repeated queries collapse into hit counts per domain and per identity. | The hits and users columns. |
| Matching | Every distinct hostname runs the walk-up match against the register. | The tool table, plus a count of unmatched domains. |
| Enrichment | Matched tools pull category, risk, sovereignty, abusive and training fields, plus your sanctioned list. | Every column right of the domain. |
The uploaded export is processed in memory and discarded after the run. Reports live 90 days in your account and can be deleted earlier.
A methodology page that lists no limits is a marketing page. These are ours, and how to work around each.
Devices using DoH to an outside resolver bypass your DNS logs. Proxy and firewall logs still catch the connection, which is one reason to audit more than one source.
Home Wi-Fi and cellular data never touch your logs. The audit measures the network you log, and states so in the evidence statement.
A tool launched yesterday may not be classified yet. The daily screening loop and the unmatched-domain count keep this window short and visible.
AI features living entirely on a suite's main domain are hard to separate from normal suite use. Depth-aware entries catch the ones with distinct hostnames.
No identity column means no per-user table. The tool inventory is unaffected.
The audit produces evidence. Blocking, if you choose it, stays in your own stack. Sequencing advice: detect before you block.
The sample evidence pack shows exactly what this methodology produces: tiles, verdicts, the tool table and the per-user breakdown, on sample data.
When someone proposes an alternative in your review meeting, this is the table to have open.
| Property | Survey | CASB catalog | Endpoint agent | Log audit |
|---|---|---|---|---|
| Covers BYOD on your network | No | Partly | No | Yes |
| Sees historical usage | Memory only | Partly | From install day | Yes, full log window |
| Deployment effort | Low | High | High | None |
| Long-tail AI tool coverage | Poor | Lags | Depends on catalog | Register updated daily |
| Training-terms context | No | Rarely | No | Dated verdict per tool |
| Cost to try | Meeting time | Procurement | Procurement | Free preview |
These methods also combine well. Teams with a CASB still run log audits, because the audit catches what the catalog has not classified yet and provides the dated training verdicts.
Every matched tool lands in exactly one category. Categories map findings to the department that owns the conversation.
Subcategories go a level deeper, over a hundred of them, which is what lets policy profiles treat "browser agents" differently from "RPA with AI".
Three hostnames from the same sample export show why per-tool context beats a yes/no AI flag.
All three of these appear in the sample report, so you can see how the verdicts render in the actual document.
Auditors probe methods, and this one holds up. Map their questions straight to report sections.
| The auditor asks | The method's answer | Where it is written |
|---|---|---|
| "What was the data source?" | The organization's own DNS, proxy or firewall export. | Scope line and evidence statement. |
| "What period does this cover?" | The export window you chose, stated with line counts. | Report metadata. |
| "What was it matched against?" | A register of 20,399 classified AI domains, maintained daily. | Scope line, with the register size at run time. |
| "How current are the vendor claims?" | Every training verdict carries the date the terms were checked. | Tool table, per row. |
| "Is it repeatable?" | Same window, same source, re-run any time. Differences between runs are findings. | Evidence statement. |
| "What are the known blind spots?" | Off-network use and third-party DoH, stated rather than hidden. | This page, and your filing notes. |
Not every hostname in your export is an AI tool, and the report says how many were not. That number is useful twice.
For "was this tool reached from this network", it is as accurate as your logs. The register's daily maintenance and depth-aware matching keep false positives down.
When the export carries a user, device or IP column, yes, per identity. Without one you still get the full tool inventory.
API traffic resolves hostnames too, and API subdomains are in the register. Server-to-server AI use shows up like any other source.
It is maintained daily, with roughly 300,000 new domains checked each day from a 120 million domain corpus.
It means the hostname was resolved or the connection was made. Hits and user counts tell you whether it was a one-off lookup or sustained use.
Built-in categories are coarse and update slowly. The audit adds per-tool risk, training verdicts with dates, and the sanctioned split, which category filters cannot express.
Many IT teams start with a hand-kept list of AI domains. The math turns against them within a quarter.
The practical split: let a maintained register carry the classification burden, and spend your team's time on the decisions the report tees up. Costs are on the pricing page.
The free preview applies the full pipeline to your export and shows the totals. Judge the method by its output.
Run the free audit