Ex-Google-Search engineer here, having also done some projects since leaving that involve data-mining publicly-available web documents. This proposal won't do very much. Indexing is the (relatively) easy part of building a search engine. CommonCrawl already indexes the top 3B+ pages on the web and makes it freely available on AWS. It costs about $50 to grep over it, $800 or so to run a moderately complex Hadoop job.…
>The comments here that PageRank is Google's secret sauce also aren't really true - Google hasn't used PageRank since 2006. That's quite a claim considering they were reporting PageRank in their toolbar until 2016, and toolbar PageRank was visible in Google Directory until 2011. Are you talking about PageRank from the original patent?
https://searchengineland.com/google-has-confirmed-they-are-r...