Live data from Hacker News

OpenAI Announces SearchGPT

chatgpt.com

141–150 of 203 posts

Re: OpenAI Announces SearchGPT

#141
post #30

OpenAI, up to this point, has shown a willingness to outcompete some of the very companies that rely on their API to function, most famously with the release of GPTs, which had a quite severe impact on many "AI startups" [0]. In that way, they remind me of Apple with Sherlock way back in the early 2000s. Despite this, I was doubtful that they'd go so far as to release a full-on search product due to their relationshi…

[flagged]

Re: OpenAI Announces SearchGPT

#142

Too bad it's only a waitlist. iFixit apparently published fake repair guides [0] that then got crawled by ChatGPT for training [1] [0] https://www.ifixit.com/Guide/Data+Connector++Replacement/147... [1] https://chatgpt.com/share/e52dc4dd-77e6-48a5-a7ca-77e3dfa39e...

So... more vaporware? :( Whatever happened to that new voice interface they "announced" back in May? I wish companies would stick to "releasing" things rather than "announcing" them...

Re: OpenAI Announces SearchGPT

#143
post #136
post #129

Earlier quoted context omitted.

>then got crawled by ChatGPT for training Not really, from the example you provided I think it's pretty clear that a custom system prompt was used and ChatGPT is using its own creativity, it doesn't have anything in common with the iFixit guide.

Actually, you can ask ChatGPT and get the same results without any prompt. No custom prompt is necessary. It takes a few seconds: https://chatgpt.com/share/2f873b1a-e487-469f-a78b-7e57fdd91b...

This is pretty funny, it seems to work with random foods, iFixit doesn't seem to have anything to do with it.

Re: OpenAI Announces SearchGPT

#144
post #125

Earlier quoted context omitted.

The entry here for Perplexity is the one that got a lot of attention but it's also unfair: PerplexityBot is their crawler, which uses that user agent and as far as anyone can tell it respects robots.txt. They also have a feature that will, if a user pastes a URL into their chat, go fetch the data and do something with it in response to the user's query. This is the feature that made a big kerfuffle on HN a while back…

No, thats illogical. The action is indeed prompted by a human, but so is any crawl in some way. At some point they either configured an interval or other trigger to send the script to the Web host to fetch anything it can find. It's inherently different to extensions such as adblockers that just remove elements according to configuration. After all, the users device will never even see the final DOM now. instead it's…

I wrote a reply but you edited out the chunk of text that I quoted, so here's a new reply.

> After all, the users device will never even see the final DOM now. instead it's getting fetched, parsed and processed on a third device, which is objectively a robot.

Sure, but why does it matter if the machine that I ask to fetch, parse, and process the DOM lives on my computer or on someone else's? I, the human being, will never see the DOM either way.

This distinction between my computer and a third-party computer quickly falls apart when you push at it.

If I issue a curl request from a server that I'm renting, is that a robot request? What about if I'm using Firefox on a remote desktop? What about if I self-host a client like Perplexity on a local server?

We live in an era where many developers run their IDE backend in the cloud. The line between "my device" and "cloud device" has been nearly entirely blurred, so making that the line between "robot" and "not robot" is entirely irrational in 2024.

The only definition of "robot" or "crawler" that makes any kind of sense is the one provided by robotstxt.org [0], and it's one that unequivocally would incorporate Perplexity on the "not robot" side:

> A robot is a program that automatically traverses the Web's hypertext structure by retrieving a document, and recursively retrieving all documents that are referenced. ... Normal Web browsers are not robots, because they are operated by a human, and don't automatically retrieve referenced documents (other than inline images).

Or the MDN definition [1]:

> A web crawler is a program, often called a bot or robot, which systematically browses the Web to collect data from webpages. Typically search engines (e.g. Google, Bing, etc.) use crawlers to build indexes.

Perplexity issues one web request per human interaction and does not fetch referenced pages. It cannot be considered a "crawler" by either of these definitions, and the definition you've come up with just doesn't work in the era of cloud software.

[0] https://www.robotstxt.org/faq/what.html

[1] https://developer.mozilla.org/en-US/docs/Glossary/Crawler

Re: OpenAI Announces SearchGPT

#145
post #110
post #93

Earlier quoted context omitted.

Google actually provides means of validating whether a request really came from them, so masquerading as Googlebot would probably backfire on you. I would expect the big CDNs to flag your IP address as malicious if you fail that check. https://developers.google.com/search/docs/crawling-indexing/...

You could maybe still only follow robots.txt rules for Googlebot.

[deleted]

Re: OpenAI Announces SearchGPT

#146

Earlier quoted context omitted.

Sites might exist for reasons other than to "be useful". At a bare minimum, they may be trying to sell eyeballs to advertisers, but they also might be trying to deliver an experience, induce some deeper engagement, make a sale, build a community, whatever. All of that disappears when a bot devours whatever it assesses to be your "content" and then serves it up as a QA response, stripped of any of the surrounding cont…

Because reading nonsense inside an infinite debatable context is fun. I know what you're talking about and frankly I'm not impressed. You know why people like these chat systems? Because it straight up saves time. When a system is made it to indexable, "context dependent", and "creating a certain experience" it just begs to be summarized and made to be something you can use. That interpretable work is.... Pointlessly…

I'm confused how this can be your opinion while you're also spending time on this website responding to people.

Why are you not just asking chatgpt "what's the latest tech news"?

Could it be that there's something else you get from this site other than just it's content being easily searchable in someone elses database?

Re: OpenAI Announces SearchGPT

#147

Earlier quoted context omitted.

Sites might exist for reasons other than to "be useful". At a bare minimum, they may be trying to sell eyeballs to advertisers, but they also might be trying to deliver an experience, induce some deeper engagement, make a sale, build a community, whatever. All of that disappears when a bot devours whatever it assesses to be your "content" and then serves it up as a QA response, stripped of any of the surrounding cont…

Because reading nonsense inside an infinite debatable context is fun. I know what you're talking about and frankly I'm not impressed. You know why people like these chat systems? Because it straight up saves time. When a system is made it to indexable, "context dependent", and "creating a certain experience" it just begs to be summarized and made to be something you can use. That interpretable work is.... Pointlessly…

Sure, and if a chatbot can helpfully summarize factual content being gatekept in a Discord chat, then that's fantastic, but I don't think that's quite what I'm getting at. The internet has room for more than just an infinite queue of fact-seekers interacting with a bank of fact-repositories. Some writing (eg, poetry) is clearly art and the people who have created it are entitled to have a bit of say over how that art is consumed and under what regimes it is summarized/remixed. Or at least us as those consumers should have the discernment required to be able to say "this isn't authentic, let me seek out the original instead."

I'm not normally a purist on these things, but I'm recalling musical artists who bemoaned the destruction of the album format in favour of $0.99/track sales in the early days of the iTunes store. Concept albums in the vein of Sgt Peppers still exist of course, but almost every modern mainstream song is now prepared first and foremost to be listened to in isolation. I didn't care for those arguments at the time they were being made, but years later, I can appreciate how something was lost there and that it might have been appropriate to let artists specify that album X was to be sold only as an album.

Re: OpenAI Announces SearchGPT

#148
post #142

Too bad it's only a waitlist. iFixit apparently published fake repair guides [0] that then got crawled by ChatGPT for training [1] [0] https://www.ifixit.com/Guide/Data+Connector++Replacement/147... [1] https://chatgpt.com/share/e52dc4dd-77e6-48a5-a7ca-77e3dfa39e...

So... more vaporware? :( Whatever happened to that new voice interface they "announced" back in May? I wish companies would stick to "releasing" things rather than "announcing" them...

> Whatever happened to that new voice interface they "announced" back in May?

Yeah, and it was suddenly made just the day before Google announced their models... and do you remember Sora? That was in February.

Re: OpenAI Announces SearchGPT

#149

Earlier quoted context omitted.

Because reading nonsense inside an infinite debatable context is fun. I know what you're talking about and frankly I'm not impressed. You know why people like these chat systems? Because it straight up saves time. When a system is made it to indexable, "context dependent", and "creating a certain experience" it just begs to be summarized and made to be something you can use. That interpretable work is.... Pointlessly…

I'm confused how this can be your opinion while you're also spending time on this website responding to people. Why are you not just asking chatgpt "what's the latest tech news"? Could it be that there's something else you get from this site other than just it's content being easily searchable in someone elses database?

I imagine that something else is conversation.

Note however that HN is not gatekeeping any useful information that may be produced during conversations here; in fact, it's all indexed and searchable.

Re: OpenAI Announces SearchGPT

#150

Maybe SEO will finally die. A man can dream.

I can’t remember the last time I had to wade through blogspam to find an answer to something I wanted. I also haven’t had to endlessly look at stack overflow questions for a couple years now. I’m really grateful to stay away from most of what the modern web has turned into
Post reply on HN