Live data from Hacker News

The librarian immediately attempts to sell you a vuvuzela

kaveland.no

31–40 of 346 posts

Re: The librarian immediately attempts to sell you a vuvuzela

#31
I built a portal that makes it easier to query against multiple different search engines (https://allsear.ch/). It's open source, free, all that. I must say, building it really expanded my view of the internet.

I am also a heavy Kagi and Reddit user for search, and usually that's enough. But when it's not, its concerning how much better other search engines can be, especially since non-tech savvy folks will never use them.

Re: The librarian immediately attempts to sell you a vuvuzela

#32
post #2

> These days, I find that I am using multiple search engines and often resort to using an LLM to help me find content. For a few months, I've been wondering: how long until advertisers get their grubby meathooks into the training data? It's trivial to add prompts encouraging product placement, but I would be completely shocked if the big players don't sell out within a year or two, and start biasing the models themse…

There are already companies promising to attack Wikipedia and product LLM-bait YouTube content. Ship's sailed.

Re: The librarian immediately attempts to sell you a vuvuzela

#33
post #8

I am deliberately keeping away from LLMs for search. I'm old enough to remember finally ditching Altavista for the new upstart Google. I did briefly flirt with Ask Jeeves but it was not good enough. I don't think anyone has it sorted yet. LLM search will always be flawed due to being a next token guesser - it cannot be trusted for "facts". A LLM fact is not even a considered opinion, it is simply next token guessing.…

If you ask ChatGPT 4o about a current event it will google things (do some sort of web search) and summarise the result.

Re: The librarian immediately attempts to sell you a vuvuzela

#34
post #2

> These days, I find that I am using multiple search engines and often resort to using an LLM to help me find content. For a few months, I've been wondering: how long until advertisers get their grubby meathooks into the training data? It's trivial to add prompts encouraging product placement, but I would be completely shocked if the big players don't sell out within a year or two, and start biasing the models themse…

> how long until advertisers get their grubby meathooks into the training data You're so right. it's not an if anymore, but when. and when it does, you wouldn't know what's an ad and what isn't. In recent years i started noticing a correlation between alcohol consumption and movies. I couldn't help but notice how many of the movies I've seen in the past few years promote alcohol and try to correlate it with the good…

Ha, now try cigarettes/smoking! At least low level alcohol consumption is only detrimental to the drinker. Cigarettes start poisoning the air from the moment they are lit, and like noise pollution there is no boundary. I hate them or thrir smokers with a vengeance and the foreign satanic cabal that is „hollywood“ sold everyone out for their gold calf tobacco money

Re: The librarian immediately attempts to sell you a vuvuzela

#36

Earlier quoted context omitted.

We already have that and it's Wikipedia.

Can you imagine how much better it would be if people were paid to create the content!

I think it would be worse. Money tends to ruin collaborative communities.

Re: The librarian immediately attempts to sell you a vuvuzela

#37
post #34

Earlier quoted context omitted.

> how long until advertisers get their grubby meathooks into the training data You're so right. it's not an if anymore, but when. and when it does, you wouldn't know what's an ad and what isn't. In recent years i started noticing a correlation between alcohol consumption and movies. I couldn't help but notice how many of the movies I've seen in the past few years promote alcohol and try to correlate it with the good…

Ha, now try cigarettes/smoking! At least low level alcohol consumption is only detrimental to the drinker. Cigarettes start poisoning the air from the moment they are lit, and like noise pollution there is no boundary. I hate them or thrir smokers with a vengeance and the foreign satanic cabal that is „hollywood“ sold everyone out for their gold calf tobacco money

But a drunkard might sit behind the wheel, at which point it becomes detrimental to everyone on the road…

And there are countless books and movies where the hero has drinks, or routinely swigs some whisky-grade stuff from a flask on his belt to calm his nerves, then drives.

Re: The librarian immediately attempts to sell you a vuvuzela

#38
post #36

Earlier quoted context omitted.

Can you imagine how much better it would be if people were paid to create the content!

I think it would be worse. Money tends to ruin collaborative communities.

They should offer more money. Wikipedia invited some professor writing. thye take responsibility for their whole life on job.

Re: The librarian immediately attempts to sell you a vuvuzela

#39
post #2

> These days, I find that I am using multiple search engines and often resort to using an LLM to help me find content. For a few months, I've been wondering: how long until advertisers get their grubby meathooks into the training data? It's trivial to add prompts encouraging product placement, but I would be completely shocked if the big players don't sell out within a year or two, and start biasing the models themse…

The providers can sell inclusion in the system prompt to advertisers. Run some ad-tech on the first message before it goes to the LLM to see whose gets included.

For most advertisers, sure, there's no need to go all the way back to the training data. Advertisers want immediate results. Training takes too long and has uncertain results. Much easier to target the prompt instead.

If you're someone like Marlboro or Coca-Cola, on the other hand, it might be worth your while to pollute the training data and wait for subtle allusions to your product to show up all over the place. Maybe they already did, long before LLMs even existed.

Re: The librarian immediately attempts to sell you a vuvuzela

#40
post #2

> These days, I find that I am using multiple search engines and often resort to using an LLM to help me find content. For a few months, I've been wondering: how long until advertisers get their grubby meathooks into the training data? It's trivial to add prompts encouraging product placement, but I would be completely shocked if the big players don't sell out within a year or two, and start biasing the models themse…

I think for the moment the leading AI companies are strongly incentivized to not succumb to the advertising curse. Their revenue is subscription driven and the competition is ridiculously fierce and immune to collusion. Everyone is trying to one-up everyone else and there is no moat that locks you into a single product. Their incentive is to score as high as possible on benchmarks in order to drive up their user base…

The adds can be outside of the AI reply-pane. Just like adds are outside of Google search results.
Post reply on HN