Three layers, two of which are not mysterious
The surface web is what a search engine has indexed and can return to you.
The deep web is everything else reachable by ordinary means but not indexed. Pages behind a login. Results generated by a database query. Documents in a private share. Your webmail, your banking session, your company intranet, this morning's calendar.
There is nothing sinister in that category. Most of it is yours, and the reason it is not indexed is that it is not supposed to be.
The dark web is a much smaller set of services that cannot be reached with a normal browser at all, because they are hosted on networks requiring specific software to access. In practice this overwhelmingly means Tor onion services.
The distinction is worth holding onto because the two terms are routinely used interchangeably, usually in a sentence implying that the vast unindexed majority of the internet is a criminal underworld. Your email is in that majority.
Where the 96 percent claim came from
You will have seen the iceberg diagram. Surface web above the waterline, deep web as the vast mass below, dark web at the bottom, with a figure somewhere between 90 and 96 percent attached.
That figure has a traceable origin, and it does not survive the trace.
It comes from a white paper by Michael Bergman, published through a company called BrightPlanet, based on data collected in March 2000. It estimated the deep web at 400 to 550 times the size of the surface web, with around 550 billion documents.
Three problems.
First, the data is from 2000. The web then bears no useful resemblance to the web now.
Second, the paper was a commercial white paper for a company selling deep web search, not peer reviewed research.
Third, and most decisively, Bergman has since publicly walked it back. He has written that sizing databases by mean rather than median inflated the result considerably, that he backed off to a range running from a few times to perhaps 100 times the surface web, and, in his own words, that "we didn't know then and we don't know now."
There is also an arithmetic oddity worth noticing. If the deep web were 550 times the surface web, it would be 99.8 percent of the total, not 96. The popular figure is not a rounding of its own source. It is a number that has drifted through retelling.
We are not offering a replacement estimate, because there is not a credible one. The honest position is that nobody has measured this in a way that supports a percentage, and the people who tried first say so themselves.
How Tor works
Worth understanding precisely, because the mechanism explains both the anonymity and its limits.
When you browse the ordinary web through Tor, your traffic is encrypted in layers and routed through three volunteer operated relays. The entry relay knows who you are but not where you are going. The middle relay knows neither. The exit relay knows where you are going but not who you are. No single relay holds both halves, which is the entire design.
Onion services work differently and are frequently confused with the above. An onion service does not have an exit node, because the traffic never leaves the network. Both the visitor and the service are anonymous to each other, and the connection is made through a rendezvous point negotiated via directory infrastructure.
That difference matters for a defender. Browsing a clear web site through Tor hides you from the site. Running or visiting an onion service hides both ends from each other and from observers.
What the network actually looks like
Tor publishes its own metrics, which makes this one of the few areas where solid numbers exist.
On 25 September 2026 the network comprised 9,429 relays and 2,334 bridges, with an estimated 3.18 million daily users. The mean over the preceding four months was approximately 3.28 million.
Two caveats on that user figure, both from Tor's own documentation. It does not count users. It counts directory requests and divides by roughly ten to produce an estimate. And a user in this sense is a client connection, not a person.
Set 3.18 million against the scale of the ordinary internet and the mythology deflates somewhat. This is a small network.
Other anonymising networks exist, principally I2P and Hyphanet, formerly Freenet. Both are substantially smaller than Tor.
What is on it, and how little we know
This is where the confident percentages come from, and they deserve scrutiny.
The most cited content classification of onion services crawled 5,205 sites over five weeks ending February 2016. Of those, over a thousand did not respond at all. The classifier was trained on 634 hand labelled documents. The authors themselves documented misclassifications, including drug forums categorised as social.
The seed list came from public onion directories, which introduces an obvious bias: it finds the sites that want to be found.
None of this makes the study worthless. It was serious work and the authors were candid about its limits. It does mean that percentages derived from it should be quoted with the sample size and date attached, and that anyone presenting them as the current state of the dark web is overstating considerably.
The more interesting finding from our research is what does not exist. No comparable large scale classification study appears to have been published since 2020. Every figure in circulation describes a snapshot taken a decade ago, of a network that has changed substantially since, including a protocol revision that makes the crawling approach used at the time much harder.
That protocol change is worth explaining. Version 3 onion services use blinded descriptors, which means you cannot enumerate services by watching the directory system the way you could before. An observer running around one percent of the directory infrastructure has roughly an eight percent chance of observing any given service. Comprehensive enumeration is not merely difficult now, it is structurally prevented by design.
So when someone tells you what percentage of the dark web is criminal, the correct follow up is to ask how they counted.
Legitimate use
Worth stating plainly, because the framing usually omits it.
Tor is used by journalists communicating with sources, by people in countries that censor the internet, by whistleblowing systems operated by mainstream news organisations, and by ordinary people who would rather not be tracked. Several large organisations operate onion services for exactly these reasons.
The technology is neutral. The population using it is not uniform.
The part that matters commercially
Here is the thing most relevant to anyone buying security services, and it cuts against our own industry's marketing.
Stolen credential trading has largely moved off Tor.
The centre of gravity is Telegram, alongside clear web forums and paste sites. The reasons are practical rather than ideological. Telegram is fast, requires no special software, has a built in audience, supports large file distribution, and reaches buyers who would never install a Tor browser. Vendors' own accounts of their collection reflect this, with one describing Telegram as the primary platform for stealer log distribution and reporting coverage of tens of thousands of channels.
We did not find neutral measurement of the split, and we should be honest that the available numbers are circular: a vendor monitoring forty thousand Telegram channels will naturally report that most of what it finds is on Telegram. What can be said is the direction of travel, and that no credible source puts the majority of this activity on onion services today.
What this means for dark web monitoring
Including ours, so take this as disclosure rather than modesty.
"Dark web monitoring" is a category name, not a description of where the collection happens. In practice it means Telegram channels, criminal forums, paste sites, marketplaces and stealer log distribution points, of which some are onion services and many are on the ordinary internet.
That is not deceptive, but it does mean the word tells you nothing useful when comparing providers. The questions worth asking are which sources are covered, how quickly material is ingested after it appears, whether collection is first hand or licensed from someone else, and what the provider will tell you about their coverage when asked directly.
A provider who answers those questions specifically is telling you something. A provider whose answer is that they monitor the dark web has told you only that they know the phrase.
For what the collected material actually consists of, see stealer log monitoring and combolists.