Live data from Hacker News

Guy running a Google rival from his laundry room

fastcompany.com

101–110 of 156 posts

Re: Guy running a Google rival from his laundry room

#101

Well, I created my own domain index. I have not crawled every page inside domains, but it is not my goal. I have 1542766 domains. Might not be much, but it is an honest work. It is available as a github repo, so anybody that wants to start crawling has some initial data to kick off. Links https://github.com/rumca-js/Internet-Places-Database

This is amazing. Thanks for sharing!

Re: Guy running a Google rival from his laundry room

#102

Earlier quoted context omitted.

Not for a CPU but earlier this year I bought a Thinkpad workstation off eBay for $500. It's a machine from 2020 and when it was new cost $5,700. I see this for pretty much all hardware out on eBay, just go back 5 years and watch the price fall 10x.

Has eBay fixed their "and then they ship you a box of rocks" problem? I feel like there was a five year span where everyone I talked to said buying or selling electronics on eBay was a nightmare, so I'm a little curious if I need to re-evaluate my priors.

> Has eBay fixed their "and then they ship you a box of rocks" problem?

I've personally never had that problem after over a decade and hundreds of purchases on eBay. I've had some defective parts, but never outright fraud. IME eBay favors buyers.

Re: Guy running a Google rival from his laundry room

#103

Earlier quoted context omitted.

Have you considered it's a good product that causes its users to become advocates?

Could also be a form of effort justification. [1] [1] https://en.wikipedia.org/wiki/Effort_justification

TIL about effort justification! I think signing up for Kagi is not particularly effort-intensive however.

Re: Guy running a Google rival from his laundry room

#104

I always wondered why someone couldn't do this. Google was invented many years ago by two guys in a dorm room and since then there's been so many white papers and advancements in the public sphere and the actual underlying problem has not changed that much, that it seems like it could be done by a small group or independent person.

I think there are two factors that helped Google. First, the search engine landscape back then was absolutely abysmal. I'm sure someone will chime in saying that it's abysmal today as well, but the reality is that 99%+ of consumer searches get good results today. And that's simply because the nature of search has changed: we have billions of people using the internet, and they overwhelmingly just search for products…

Google maps is probably a big moat that's very hard to replicate. You can't as easily just crawl all of that data. It's not easy to generate directions. The average user doesn't want to use your search engine for one thing and Google for everything else, they just want a one stop shop for search.

Re: Guy running a Google rival from his laundry room

#106

Well, I created my own domain index. I have not crawled every page inside domains, but it is not my goal. I have 1542766 domains. Might not be much, but it is an honest work. It is available as a github repo, so anybody that wants to start crawling has some initial data to kick off. Links https://github.com/rumca-js/Internet-Places-Database

[dead]

Re: Guy running a Google rival from his laundry room

#107

I always wondered why someone couldn't do this. Google was invented many years ago by two guys in a dorm room and since then there's been so many white papers and advancements in the public sphere and the actual underlying problem has not changed that much, that it seems like it could be done by a small group or independent person.

The actual underlying problem has changed altogether. Pagerank is easily gamed by SEO. Search candidates and rankings now require assessment by LLM. Moreover, as a default, users want the results intelligently synthesized into a text response with references rather than as raw results. Crawling too requires innovative approaches to bypass server filters. I doubt any independent person can afford to run a vector datab…

> users want the results intelligently synthesized into a text response with references rather than as raw results

This leads directly to another big change.

People used to submit their sites to search engines and now they might actively block search engines. So a search engine author might have to spend a lot of effort in adversarial games.

Re: Guy running a Google rival from his laundry room

#108

Earlier quoted context omitted.

Not for a CPU but earlier this year I bought a Thinkpad workstation off eBay for $500. It's a machine from 2020 and when it was new cost $5,700. I see this for pretty much all hardware out on eBay, just go back 5 years and watch the price fall 10x.

Has eBay fixed their "and then they ship you a box of rocks" problem? I feel like there was a five year span where everyone I talked to said buying or selling electronics on eBay was a nightmare, so I'm a little curious if I need to re-evaluate my priors.

You don't get that with used old stuff, you get it with unrealistic low prices for new stuff.

A 7532 CPU is now ewaste for all the datacenters out there 1/10 of original price is reasonable, but the latest Nvidia GPU for 200 bucks is obviously a scam.

Re: Guy running a Google rival from his laundry room

#109
post #54

"The beefy CPU running this setup, a 32-core AMD EPYC 7532, underlines just how fast technology moves. At the time of its release in 2020, the processor alone would have cost more than $3,000. It can now be had on eBay for less than $200" why do I never get deals like that when I am shopping for the homelab on eBay?

I searched "AMD EPYC 7532" and there are a ton of listings for $150-$200. Are you just regretful that it wasn't like this when you were shopping parts for your homelab?

I got a 7551p plus motherboard and ram for about 600 bucks from China this January. I may have overpaid but it works great, and gets the job done.

Re: Guy running a Google rival from his laundry room

#110

Earlier quoted context omitted.

You mean all the users of chat services aren't evidence? Chat services increasingly incorporate web links for references in their responses, and this is as the users seek. The tide continues to shift from traditional search to LLM synthesis.

I suspect there are more users of traditional search than there are of llm chat apps.

I suspect that chat apps dominate (80+%?) the under-20 demographic, and have a sizable chunk of the under-30 demographic. Within the next five years it will probably represent 50+% of total search traffic. Maybe it already does. It makes sense that any search site that wants to be in the game tomorrow would keep racing down the AI chat path.
Post reply on HN