Live data from Hacker News

Two upstart search engines are teaming up to take on Google

wired.com

231–240 of 299 posts

Re: Two upstart search engines are teaming up to take on Google

#231

I'm beginning to suspect LLMs' viability for search purposes is dependent on your existing search habits. 40% of my search queries are just copy-pasted error messages. Another 10% are business names, for the sole purpose of finding their hours or phone number. Less than 10% are complete clauses or sentences. I just don't see how an LLM fits my search habits. I tried using ChatGPT for search purposes, and it was dread…

>Have you people been typing complete sentences into Google all these years?

Well somebody has to be responsible for promoting the cancer which converted search boxes from logical set membership definition strings to "Smart™ boxes," designed to piss you off and then direct you to a sponsored product page.

Re: Two upstart search engines are teaming up to take on Google

#233

I'm beginning to suspect LLMs' viability for search purposes is dependent on your existing search habits. 40% of my search queries are just copy-pasted error messages. Another 10% are business names, for the sole purpose of finding their hours or phone number. Less than 10% are complete clauses or sentences. I just don't see how an LLM fits my search habits. I tried using ChatGPT for search purposes, and it was dread…

There are some tasks that I would have used a search engine for in the past, because that was more or less the only option. Say, searching for "ffmpeg transcode syntax" and then spending 15 minutes comparing examples from Stack Overflow with the official documentation to try to make sense of them. Now I can tell Claude exactly what I'm trying to accomplish and it will give me an answer in 30 seconds that's either cor…

Ads delivered via LLMs will cost more to distribute, which means higher cost for the businesses purchasing ads, perhaps high enough to deter a lot of smaller ads customers, so I think we'll see an interesting dynamic appear there. Especially if the ad-laden SEO-boosted sites suffer from further enshittification of Search, which has been spinning in its own vicious cycle lately.

Re: Two upstart search engines are teaming up to take on Google

#234

I'm beginning to suspect LLMs' viability for search purposes is dependent on your existing search habits. 40% of my search queries are just copy-pasted error messages. Another 10% are business names, for the sole purpose of finding their hours or phone number. Less than 10% are complete clauses or sentences. I just don't see how an LLM fits my search habits. I tried using ChatGPT for search purposes, and it was dread…

There are some tasks that I would have used a search engine for in the past, because that was more or less the only option. Say, searching for "ffmpeg transcode syntax" and then spending 15 minutes comparing examples from Stack Overflow with the official documentation to try to make sense of them. Now I can tell Claude exactly what I'm trying to accomplish and it will give me an answer in 30 seconds that's either cor…

>>>Obviously, the free online LLM business model isn't going to last indefinitely.

I think this is one of those points where LLMs have already changed the paradigm

People liked being able to search, but would not pay for it. For many queries , the value wasnt there: users still had to scroll thru pages or tinker for the right query for the required result.

Eventually search turned so much worse by seo spam, that kagi stepped in to fill the void.

LLMs start from a different direction. The value is clearly there. OpenAI etc still have a ton of paying subscribers.

I do think eventually some will incorporate ads, but I think innovation has revealed that theres a market -perhaps a substantial one- for fee-based information search with LLMS

Re: Two upstart search engines are teaming up to take on Google

#235
post #51

Earlier quoted context omitted.

>competing with Google was borderline impossible a decade ago. But in 2024, we have cheap compute, great OSS distributed DBs, powerful new vector search tech. [...] CommonCrawl text-only is ~100TB, Those example 3 bullet points of today's improved 2024 computing power you list isn't even enough to process Google's scale 14 years ago in 2010 when the search index was 100+ petabytes : https://googleblog.blogspot.com/20…

The article you linked doesn't say anything about 100 petabytes

>The article you linked doesn't say anything about 100 petabytes

Excerpt from the article: >[...] Caffeine takes up nearly 100 million gigabytes of storage in one database and adds new information at a rate of hundreds of thousands of gigabytes per day. [...]

Your comment did make me pause and sanity check the math: https://www.google.com/search?q=%22100+million+gigabytes%22

In any case, a lot of people translated "100 million gigabytes" to "100 petabytes" based on that blog : https://www.google.com/search?q=google+search+index+estimate...

What's the current best estimate of its size now in 2024?

Re: Two upstart search engines are teaming up to take on Google

#236
The best way to stop SEO is to build a bot which detects commercial activity on the page (and/or availability of cart/payment controls), and commercial language in the text of the page (does the content look like an ad or product/sales page). Also combine with number of hops to a commercial page (with cart, prices, advertising material, etc), and create a distance metric. Offer the user a commercial activity slider. Sliding to zero commercial activity should result in zero product pages or SEO spam.

Also important are user provided blacklists and whitelists of domains. Experts-Exchange was a site that called for this on search engines like Google, as Experts-Exchange was universally a cash sticky-trap waiting for searchers in pain to pony up. No, and constantly pasting -site:xyz each time is unacceptable.

ADDENDUM COPY&PASTE:

One blind spot I missed: advertising and telemetry activity are to count strongly in the commercial score.

When I slide the Commercial slider to zero, only web pages which there are no obvious detectable ways for the host or author to make money from showing me the page should be returned.

Bad behavior stems from the drive to make money. Make a search where it is impossible to make money by being a result in the search (especially if a Commercial slider is slid to low or zero).

Re: Two upstart search engines are teaming up to take on Google

#237

The best way to stop SEO is to build a bot which detects commercial activity on the page (and/or availability of cart/payment controls), and commercial language in the text of the page (does the content look like an ad or product/sales page). Also combine with number of hops to a commercial page (with cart, prices, advertising material, etc), and create a distance metric. Offer the user a commercial activity slider.…

The presence of affiliate links should be the biggest indicator that the information provided is going to be heavily financially incentivized (ie. bad).

Would be much simpler to calculate and detect. I would be super surprised if Google's algorithm doesn't already know how to detect affiliate links.

Re: Two upstart search engines are teaming up to take on Google

#238

Earlier quoted context omitted.

I'd like to add to your #5. If Google deems you a legitimate threat, then they can just de-crapify their own search for a bit by going back to their old algo. It's extremely easy for them to fight back.

That would be a victory for users.

A temporary reprieve, more like. Google will just wait for the competitor to fail and then go back to their old ways.

Re: Two upstart search engines are teaming up to take on Google

#239

I'm beginning to suspect LLMs' viability for search purposes is dependent on your existing search habits. 40% of my search queries are just copy-pasted error messages. Another 10% are business names, for the sole purpose of finding their hours or phone number. Less than 10% are complete clauses or sentences. I just don't see how an LLM fits my search habits. I tried using ChatGPT for search purposes, and it was dread…

"I want to write EDM like David Guetta. Suggest where to start, books, etc for someone who understands music theory" Good luck wading through whatever Google gives back for that. Also I find they do well on error messages in general.

The results for this particular query are quite okay actually, a relatively recent reddit thread, a quora page, then some forum and youtube results. Funnily, this very HN post is on the first page of results as well.

Re: Two upstart search engines are teaming up to take on Google

#240
post #194

Earlier quoted context omitted.

It's not just that there are few signals to prevent the wrong person being put in charge, but this kind of government bureaucratic actively selects for the wrong person. These kinds of government IT projects are often soul sucking to work on, and so they attract a specific kind of applicant.

It's not only that, they would prefer to give a ton of money to a single entity with "market experience" and all the paperwork that looks good already existing. It would work better if it was lots of small amounts of money to individuals with good "business" (think about workplan, expected value and other stuff) but without paperwork - think a kid that just finished studies or a person that acquired experience freela…

Sadly, our EU kids are more interested in being the next social media influencer or going on a world trip than grounding a business, because the paperwork and legal burden will age them faster than someone on drugs and will likely get them in hefty fine to bankruptcy due to that one arcane reporting requirement they missed about that €20/- to the tax office.

EU needs something akin to Stripe Atlas, but that is not what the politicians want because they want EU to be manufacturing industry only, you can always import tech from other places …

Post reply on HN