With the frustrations from using the usual search engines growing I’ve been thinking about this a lot lately.
It seems we’ve poisoned the well by allowing the proliferation of advertising interests to dominate the web.
Like how hard would it be to make your own non-commercial index?
Only human made sites that aren’t related to buying, selling, marketing, etc.
Could that be a federated open-source project?


Thanks for providing that!
Not for the faint of heart resource-wise, but doesn’t sound impossible for a dedicated group.
Oops, realized I didn’t answer your question about actually crawling, dig into common crawl documentation, they provide a bunch of technical data and stats that show you the scale…2-4billion pages per month
And note CC just does a sample of the pages it finds. So the more monthly dumps don’t contain all of the data afaik
And the number above are for one of the monthly dumps
https://commoncrawl.github.io/cc-crawl-statistics/plots/crawlermetrics