Earlier quoted context omitted.
We would employ local guides all around the world to craft itinerary plans to visit places, give tips, tricks, recommend experiences and places (we made money by selling some of those through our website) and it was a success. Customers liked the in depth value of that content and it converted to buys (we sold experiences and other stuff, sort of like getyourguide). One day all of our content ended up on Google "what…
I totally get that it killed your traffic. If a thousand people a day typing in "what time is best to visit the Sagrada Familiar" stopped clicking on the link to your page because Google just told them "4 PM on Thursdays" at the top of the page, you lost a bunch of traffic. But why did you want the traffic? Was your revenue from ad impressions, or were you perhaps being paid by the city of Barcelona to provide useful…
Perplexity AI is lying about their user agent
491–500 of 555 posts
Re: Perplexity AI is lying about their user agent
#492Earlier quoted context omitted.
I totally get that it killed your traffic. If a thousand people a day typing in "what time is best to visit the Sagrada Familiar" stopped clicking on the link to your page because Google just told them "4 PM on Thursdays" at the top of the page, you lost a bunch of traffic. But why did you want the traffic? Was your revenue from ad impressions, or were you perhaps being paid by the city of Barcelona to provide useful…
Moreover, if it's the former, then good riddance . An ad-backed site is harming users a little on the margin for the marginal piece of information. Getting the same from a search engine is saving users from that harm. Parent has the right question here: why did you want the traffic? Did you intend for anything good to happen to those people? . I'm going to guess not; there's hardly a scenario where people who complai…
Re: Perplexity AI is lying about their user agent
#493Earlier quoted context omitted.
> (which are notably not the same as crawling) Is that a distinction without a difference? I think the robots.txt RFC was addressed specifically to crawlers ; so technically "ad hoc" requests generated automatically (i.e. by robots) aren't included. But the distinction operators would like to make is between humans and automata. Whether some automaton is a crawler or not isn't relevant.
If explicitly telling it to access a URL is an access by automaton, then isn't every web browser load an access by automaton?
And if we took the analogy to the other end, one could argue that all crawlers have to be kicked off manually at some point...
The problem is here in reality the differentiation is somewhat more understood.
The honor system web is going away, that's for sure.
Re: Perplexity AI is lying about their user agent
#494Earlier quoted context omitted.
It would be a lot better if Google just prioritised concise websites. If Google preferred websites that cut the fluff, then website operators would have an incentive to make useful websites, and Google wouldn't have as much of an incentive to provide the answer in a snippet, and everyone wins. I guess it's hard to rank website quality, so Google just prefers verbose websites.
> Google wouldn't have as much of an incentive to provide the answer in a snippet, and everyone wins. Google has at least two incentives to provide that answer, both of which wouldn't change. The bad one: they want to keep you on their page too, for usual bullshit attention economy reasons. The good one: users prefer the snippets too . The user searching for information usually isn't there to marvel at beauty of rand…
The only reason that users prefer snippets is because websites hide the info you are looking for. The problem is that the top ranked search results are ad-infested SEO crap.
If the top ranked website were actually designed with the user in mind, they would not hide the important info. They would present the most important info at the top, and contain additional details below. They would offer the user exactly what they want immediately, and provide further details that the user can read if they want to.
Think of a well written wikipedia article. The summary is probably all that you need, but it's good that the rest of the article with all the detail is there as well. I'm pretty sure that most people prefer a well designed user-centric article to the stupid Google snippet that may or may not answer the question you asked.
Most people looking for info don't look for just a single answer. Often, the answer leads to the next question, or if the answer is surprising, you might want to check out if the source looks credible, etc. Even ads would be helpful, if they were actually relevant (eg. if I am looking for low profile graphic cards, I'd appreciate an ad for a local retailer that has them in stock).
But the problem is that website operators (and Google) just want to distract you, capture your attention, and get you to click on completely irrelevant bullshit, because that is more profitable than actually helping you.
Re: Perplexity AI is lying about their user agent
#495Earlier quoted context omitted.
I think a concern for people who contribute on Stack Overflow is that an LLM will pollute the water with so many subtly wrong answers that the collective work of answering questions accurately will be overwhelmed by a tsunami of inaccurate LLM-generated answers, more than an army of humans can keep up with checking and debugging (or debunking).
It's nice that people are willing to create content on Stack Overflow so that Prosus NV can make advertising revenue from their free labor. But ultimately only a fool would trust answers from secondary sources like Stack Overflow, Quora, Wikipedia, Hacker News, etc. They can be useful sources to start an investigation but ultimately for anything important you still have to drill down to reliable primary sources. This…
It is not simply "nice", or for internet points, to take time to answer other people's questions.
Being able to pass on knowledge is the glue of society and civilization. Cynicism about the value or reason of doing so is not a replacement for a functioning structure to educate people who want to learn or to point them in the right direction.
Re: Perplexity AI is lying about their user agent
#496There are two different questions at play here, and we need to be careful what we wish for. The first concern is the most legitimate one: can I stop an LLM from training itself on my data? This should be possible and Perplexity should absolutely make it easy to block them from training. The second concern, though, is can Perplexity do a live web query to my website and present data from my website in a format that th…
It's funny I posted the inverse of this. As a web publisher, I am fine with folks using my content to train their models because this training does not directly steal any traffic. It's the "train an AI by reading all the books in the world" analogy. But what Perplexity is doing when they crawl my content in response to a user question is that they are decreasing the probability that this user would come to by content…
Re: Perplexity AI is lying about their user agent
#497Earlier quoted context omitted.
> They’d be fools to buy licenses before it’s been decided. They are willingly ignoring licenses until someone sues them? That's still illegal and completely immoral. There is tons of data to train on. The entirety of Wikipedia, all of StackOverflow (at least previously), all of the BSD and MIT licenses source code on Github, the entire Gutenberg project. So much stuff, freely and legally available, yet their feel th…
The legality of their behavior is not currently well defined, because it's unprecedented. Fair use permits transformative works. It has yet to be decided whether LLMs and their output qualify as transformative, or even if the training is capable of infringing copyright of an individual work in the first place if they're not reproducing it. In fact, there's a good amount of evidence which indicates that fair use _does…
Your take on how all this works is probably more inline with reality than mine, it's just that my brain refuse to comprehend the willingness to take on that type of risk.
You're basically telling investors that your business may be violating all sorts of IP laws, you don't know and have taken no actions to determine that. It's just a gamble that this might work out, while taking billions in funding. There's apparently no risk assessment in VC funding.
Re: Perplexity AI is lying about their user agent
#498Earlier quoted context omitted.
It's funny I posted the inverse of this. As a web publisher, I am fine with folks using my content to train their models because this training does not directly steal any traffic. It's the "train an AI by reading all the books in the world" analogy. But what Perplexity is doing when they crawl my content in response to a user question is that they are decreasing the probability that this user would come to by content…
This is why media publishers went behind paywalls to get away from Google News
Re: Perplexity AI is lying about their user agent
#499Earlier quoted context omitted.
Moreover, if it's the former, then good riddance . An ad-backed site is harming users a little on the margin for the marginal piece of information. Getting the same from a search engine is saving users from that harm. Parent has the right question here: why did you want the traffic? Did you intend for anything good to happen to those people? . I'm going to guess not; there's hardly a scenario where people who complai…
Now think of the 2nd order effects: they paid money to collect that useful information. If it’s no longer feasible to create such high quality content, it won’t magic itself into existence on its own. It’ll all be just crap and slop in a few years.
Except it kind of does. Almost all high-quality free content on the Internet has been made by hobbyists just for the sake of doing it, or as some kind of expense (marketing budget, government spending). The free content is not supposed to make money. An honest way of making money with content is putting up a paywall. Monetizing free content creates a conflict of interest, as optimizing value to publisher pulls it in opposite direction than optimizing for value to consumer. Can't save both masters, and all. That's why it's effectively a bullet-proof heuristic, that the more monetization you see on some free content, the more wrong and more shit it is.
Put another way, monetizing the audience is the hallmark of slop.
Re: Perplexity AI is lying about their user agent
#500Earlier quoted context omitted.
Thanks for the link, that's fantastic to hear! I'm seriously sick of that whole "laundering copyright via AI"-grift - and the destruction of the creative industry is already pretty noticable. All the creatives who brought us all those wonderful masterworks with lots of thought and talent behind, they're all going bankrupt and getting fired right now. It's truly a tragedy - the loss of art is so much more serious than…
> destruction of the creative industry is already pretty noticable. Can you explain what you mean by this? I’d be interested to know what jobs have been lost to AI (or if you are talking about something else)
The most extreme situation is concept artists right now. Essentially, the entire profession has lost their jobs in the last year. Or casual artists making drawings for commission - they can't compete with AI and mostly had to stop selling their art. Similar is happening to professional translators - with AI, the translations are close enough to native that nobody needs them anymore.
The book market is getting flooded with AI-crap, so is of course the web. Authors are losing their jobs.
Currently, it seems to be creeping into the music market - not sure if people are going to notice/accept AI-made music. All the fantastic artists creating dubs are starting to go away as well, after all you can just synthesize their voices now.
It's quite sad, all considered.