Perhaps a better approach would be building an open source www index or even a full current cache - as an enabler for people to build their own search engines? Right now it is extremely difficult to build your own web crawler that would compete with Google. And that is not because of the technology, but because multiple sites will prevent your bot from accessing them if you're not Google or Bing - either through robo…
Nah, not a small task but you can break it down into well understood problems that have known solution.
The hard part is ranking everything.