Live data from Hacker News

Perplexity is a bullshit machine

wired.com

41–50 of 50 posts

Re: Perplexity is a bullshit machine

#41

Earlier quoted context omitted.

Because technology aside, its a new group of well funded companies who act as parasites on this already struggling industry. I know HN (and probably SV at large) has an anti-journalist bent, but ultimately by draining the individuals and companies that do the reporting, they wont have new material to scrape in the future. (Disclosure - I'm a former journalist building a way for the news industry to fight back. https:…

Well, as of the last 10 years the journalism industry has just been a lazy tweet regurgitation factory. The pipeline looks something like Tweets -> Journalists -> LLMs. So maybe it would be best if we cut out the middlemen?

LLMs are the middlemen. Search engines are the middlemen. Social networks are the middlemen.

Go to the NY Times or LA Times and tell me it's based on Tweets. If you're talking about a place like The Verge which is a tech blog with journalism, they have articles based on tweets but also plenty that are legitimate, researched articles.

Re: Perplexity is a bullshit machine

#42

Earlier quoted context omitted.

Well, as of the last 10 years the journalism industry has just been a lazy tweet regurgitation factory. The pipeline looks something like Tweets -> Journalists -> LLMs. So maybe it would be best if we cut out the middlemen?

LLMs are the middlemen. Search engines are the middlemen. Social networks are the middlemen. Go to the NY Times or LA Times and tell me it's based on Tweets. If you're talking about a place like The Verge which is a tech blog with journalism, they have articles based on tweets but also plenty that are legitimate, researched articles.

If that was true then what do journalists have to be worried about? Why would they need to "fight back" against the LLMs as the above poster said - as you know, ChatGPT can't do original research or investigations on its own. The truth is that the bulk of what is called "journalism" today isn't on-the-ground gumshoes reporting, it's rewording tweets/press releases/PR statements and maybe a couple of phone calls mixed in. A good rule of thumb is that anything that can be replaced by a current-gen AI wasn't worth much to begin with.

Re: Perplexity is a bullshit machine

#43

Earlier quoted context omitted.

LLMs are the middlemen. Search engines are the middlemen. Social networks are the middlemen. Go to the NY Times or LA Times and tell me it's based on Tweets. If you're talking about a place like The Verge which is a tech blog with journalism, they have articles based on tweets but also plenty that are legitimate, researched articles.

If that was true then what do journalists have to be worried about? Why would they need to "fight back" against the LLMs as the above poster said - as you know, ChatGPT can't do original research or investigations on its own. The truth is that the bulk of what is called "journalism" today isn't on-the-ground gumshoes reporting, it's rewording tweets/press releases/PR statements and maybe a couple of phone calls mixed…

> If that was true then what do journalists have to be worried about?

What part are you talking about? Journalists, and everyone else, has to worry that anything they publish then gets served up by the magic bot that billions of dollars are going towards making it an everyday part of life. And Perplexity with $165M in funding is so shady that they can't respect robots.txt and scrape without attribution.

> The truth is that the bulk of what is called "journalism" today isn't on-the-ground gumshoes reporting, it's rewording tweets/press releases/PR statements and maybe a couple of phone calls mixed in

I think that's kind of true but don't neccessarily think that's the majority of major newspapers. It's definitely true for blogs or something like Techcrunch that is 90% press release and 10% articles based on other articles from other outlets. Looking at the LA Times it seems to me like they are reporting the news: https://www.latimes.com

Re: Perplexity is a bullshit machine

#44

Perplexity probably isn't a good company for the reason that they have bad design sense and product. If you go on their main blog such as ( https://www.perplexity.ai/hub/blog/perplexity-raises-series-... ), the monospace choice of font (not that it's monospace, just that from a design sense it doesn't seem right), the excessively large border, and the container being unaligned with the content/image/frontmatter all r…

I don't think that remotely matters to their success. If it does, Exa uses a font in their blog that looks pretty close to Perplexity's choice.

Re: Perplexity is a bullshit machine

#45

Earlier quoted context omitted.

LLMs are the middlemen. Search engines are the middlemen. Social networks are the middlemen. Go to the NY Times or LA Times and tell me it's based on Tweets. If you're talking about a place like The Verge which is a tech blog with journalism, they have articles based on tweets but also plenty that are legitimate, researched articles.

If that was true then what do journalists have to be worried about? Why would they need to "fight back" against the LLMs as the above poster said - as you know, ChatGPT can't do original research or investigations on its own. The truth is that the bulk of what is called "journalism" today isn't on-the-ground gumshoes reporting, it's rewording tweets/press releases/PR statements and maybe a couple of phone calls mixed…

> A good rule of thumb is that anything that can be replaced by a current-gen AI wasn't worth much to begin with.

That's the point, though. It isn't replacing journalists, it's just stealing from them. The reporters are the ones doing the work, the AI is just plagiarizing. It's using their work, while not only removing credit, but more importantly, undermining their livelihoods.

Re: Perplexity is a bullshit machine

#46
It was dumb of Perplexity to specifically say that they follow robots.txt, and then not respect it immediately after. I doubt it was a consciously malicious decision as it wouldn't make any sense to try and get some good will from that commitment and then expect nobody to notice that you weren't actually doing that. I expect that they simply forgot they even made that promise. That's still a problem though. Too many companies sloppily make pledges because they think it will make them look good only to immediately stop caring about the pledge right after they say it.

While it was definitely wrong of them to commit to it and renege, I don't really think robots.txt should be expected to be respected for a live web scraping use case. I think performing a web request immediately after a user requests it is much more similar to a proxy service than it is to an indexing service like search engines. Should archive.is also be expected to respect robots.txt as well, even though it's directly performing a scrape on the user's behalf? How is live scraping materially different than a proxy or a CDN? robots.txt was created in the context of crawlers scraping content on a batch schedule as opposed to a real time on the fly schedule.

Live request services could even leverage the end user's machine to perform the requests directly and then provide the data to them instead of doing the scrape on their behalf, so I feel as though the directly responsible invoker of the data being taken off their sites is more the user than the service that's acting as a proxy.

The current media order is definitely at risk and it's something we need to find solutions to protect in some way to prevent reporting from dying in this paradigm shift. Trying to push against the entirety of AI progress is not going to work and this is really just screaming into the void over technicalities. Even if live scrape powered AI services were banned, the same service would just move to the end user with apps that will perform the requests directly on the user's device. I don't think anyone outside of the industry cares about the technical nuances of how AI services visit websites for the user. There's a bigger problem at play here regarding the future of high quality media, and it needs to be addressed directly.

Re: Perplexity is a bullshit machine

#47
post #33
post #2

https://archive.is/mz5m8

An archive link getting posted so you can read the content for free. Theres a bit of irony to this...

They also scrape and cache sites on behalf of the user too. I understand why media sites are upset about their lack of control over their content, but that's just how the internet works. Sites and services proxy requests for their users all the time, and since this is live scraping directly on behalf of the user from an affirmative user action, it's not really materially different from a proxy.

You could simply shift the site requests to an end user machine instead and then have that data be sent to your server for processing, so all this seems like a technicality of what machine is doing which work. It's very different from an indexer like a search engine that does batched crawling.

Re: Perplexity is a bullshit machine

#48
post #33
post #2

https://archive.is/mz5m8

An archive link getting posted so you can read the content for free. Theres a bit of irony to this...

I don't understand how these archive sites aren't shut down yet. Is it because only a small sliver of internet users know about them?

Re: Perplexity is a bullshit machine

#49
post #44

Perplexity probably isn't a good company for the reason that they have bad design sense and product. If you go on their main blog such as ( https://www.perplexity.ai/hub/blog/perplexity-raises-series-... ), the monospace choice of font (not that it's monospace, just that from a design sense it doesn't seem right), the excessively large border, and the container being unaligned with the content/image/frontmatter all r…

I don't think that remotely matters to their success. If it does, Exa uses a font in their blog that looks pretty close to Perplexity's choice.

I see that they have bad design sense which I think generalizes to bad product sense and bad business sense.
Post reply on HN