Live data from Hacker News

Introducing deep research

openai.com

371–380 of 445 posts

Re: Introducing deep research

#372

I just gave it a whirl. Pretty neat, but definitely watch out for hallucinations. For instance, I asked it to compile a report on myself (vain, I know.) In this 500-word report (ok, I'm not that important, I guess), it made at least three errors. It stated that I had 47,000 reputation points on Stack Overflow -- quite a surprise to me, given my minimal activity on Stack Overflow over the years. I popped over to the l…

"Pretty neat, but definitely watch out for hallucinations." We'd never hire someone who just makes stuff up (or at least keep them employed for long). Why are we okay with calling "AI" tools like this anything other than curious research projects? Can't we just send LLMs back to the drawing board until they have some semblance of reliability?

> Why are we okay with calling "AI" tools like this anything other than curious research projects?

Because they are a way to launder liability while reducing costs to produce a service.

Look at the AI-based startups y-combinator has been funding. They match that description.

Re: Introducing deep research

#373
post #368

Earlier quoted context omitted.

> Can't we just send LLMs back to the drawing board until they have some semblance of reliability? Well at this point they've certainly proven a net gain for everyone regardless of the occasional nonsense they spew.

"Occasional nonsense" doesn't sound great, but would be tolerable. Problem is - LLMs pull answers from their behind, just like a lazy student on the exam. "Halucinations" is the word people use to describe this. Those are extremely hard to spot - unless you happen to know the right answer already, at which point - why ask? And those are everywhere . One example - recently there was quite a discussion about llm being…

Yeah, what you said represents a 'net gain' over not having any of that at all.

Re: Introducing deep research

#374
post #341

Earlier quoted context omitted.

Interesting You might find it amusing to compare it to: https://hn-wrapped.kadoa.com/timabdulla (Ref: https://news.ycombinator.com/item?id=42857604 )

This is... very uncomfortable. An (expanded) AI summary of my HN and reddit usage would appear to be a pretty complete representation of my "online" identity/character. I remember when people would browse your entire comment history just to find something to discredit you on reddit, and that behavior was _heavily_ discouraged. Now, we can just run an AI model to follow you and sentence you to a hell of being permanen…

throwaway/anonymous.

not just for when discussion of the content not the personality behind it is important.

Re: Introducing deep research

#375
post #341

I just gave it a whirl. Pretty neat, but definitely watch out for hallucinations. For instance, I asked it to compile a report on myself (vain, I know.) In this 500-word report (ok, I'm not that important, I guess), it made at least three errors. It stated that I had 47,000 reputation points on Stack Overflow -- quite a surprise to me, given my minimal activity on Stack Overflow over the years. I popped over to the l…

Interesting You might find it amusing to compare it to: https://hn-wrapped.kadoa.com/timabdulla (Ref: https://news.ycombinator.com/item?id=42857604 )

I put my profile in [0] and it's mostly silly; a few comments extracted and turned into jokes. No deep insights into me, and my "Top 3 Technologies" are hilariously wrong (I've never written a single line of TypeScript!)

[0]: https://hn-wrapped.kadoa.com/dlivingston

Re: Introducing deep research

#376
post #370
post #275

Earlier quoted context omitted.

It's a bit like saying "my kids are going to hit themselves anyway, so it doesn't matter if I give them foam rods or metal rods".

Maybe this would make sense if you saw the whole world as "kids" that you had to protect. As an adult who lives in an adult world, I would like adults to have access to metal tools and not just foam ones.

I guess I can replace "kid" with "toddler" and add "unsupervised" at the end.

Re: Introducing deep research

#377

Earlier quoted context omitted.

That is... debatable. You may be entirely inside the bubble, there.

Not sure if this was posted as humour, but I don't feel that way. In today's world, where I certainly would consider taking the blue pill, I'm having a blast with LLMs! It has helped me learn stuff incredibly faster. Especially I find them useful for filling the gaps of knowledge and exploring new topics in my own way and language, without needing to wait an answer from a human (that could also be wrong). Why does it…

> without needing to wait an answer from a human (that could also be wrong).

The difference is you have some reassurances that the human is not wrong - their expertise and experience.

The problem with LLMs, as demonstrated by the top-level comment here, is that they constantly make stuff up. While you may think you're learning things quickly, how do you know you're learning them "correctly", for lack of a better word?

Until an LLM can say "I don't know", I really don't think people should be relying on them as a first-class method of learning.

Re: Introducing deep research

#378

For “deep research” I’m also reading “getting the answers right”. Most people I talk to are at the point now where getting completely incorrect answers 10% of the time — either obviously wrong from common sense, or because the answers are self contradictory — undermines a lot of trust in any kind of interaction. Other than double checking something you already know, language models aren’t large enough to actually kno…

> They can only sound like they do.

More importantly, I think, is that they are incapable of not doing so. Have we figured out how to make an LLM realize and answer that it doesn't know an answer?

Re: Introducing deep research

#379

I just gave it a whirl. Pretty neat, but definitely watch out for hallucinations. For instance, I asked it to compile a report on myself (vain, I know.) In this 500-word report (ok, I'm not that important, I guess), it made at least three errors. It stated that I had 47,000 reputation points on Stack Overflow -- quite a surprise to me, given my minimal activity on Stack Overflow over the years. I popped over to the l…

"Pretty neat, but definitely watch out for hallucinations." We'd never hire someone who just makes stuff up (or at least keep them employed for long). Why are we okay with calling "AI" tools like this anything other than curious research projects? Can't we just send LLMs back to the drawing board until they have some semblance of reliability?

Yeah, I used to hire people, but then one of them made a mistake, now I'm done with them forever, they are useless. It is not I, who is directing the workers, who cannot create a process that is resistant to errors, it's definitely the fact that all people are worthless until they make no errors as there truly is no other way of doing things other than telling your intern to do a task then having them send it directly to the production line.

Re: Introducing deep research

#380

Earlier quoted context omitted.

Not sure if this was posted as humour, but I don't feel that way. In today's world, where I certainly would consider taking the blue pill, I'm having a blast with LLMs! It has helped me learn stuff incredibly faster. Especially I find them useful for filling the gaps of knowledge and exploring new topics in my own way and language, without needing to wait an answer from a human (that could also be wrong). Why does it…

>It has helped me learn stuff incredibly faster. Especially I find them useful for filling the gaps of knowledge and exploring new topics in my own way and language and then you verify every single fact it tells you via traditional methods by confirming them in human-written documents, right? Otherwise, how do you use the LLM for learning? If you don't know the answer to what you're asking, you can't tell if it's lyi…

I guess it all depends on the topic and levels of trust. How can I be certain that I have a brain? I just have to take something for granted, don't I? Of course I will "verify" the "important stuff", but what is important? How can I tell? Most of the time only thing I need is a pointer in the right direction. Wrong advice? I know when I get there I suppose.

I can remember numerous things I was told while growing up, that aren't actually true. Either by plain lies and rumours or because of the long list of our cognitive biases.

> If you have to look up every fact it outputs after it does, using traditional methods, why not skip to just looking things up the old fashioned way and save time?

What is the old fashioned way? I mean people learn "truths" these days from Tiktok and Youtube. Some of the stuff is actually very good, you just have to distill it based on the stuff I was being taught at school. Nonody has yet declared LLMs as a subtitute for schools, maybe they soon will, but neither "guarantees" us anything. We could as well be taught political agendas.

I could order a book about construction, but I wouldn't build a house without asking a "verified" expert. Some people build anyway and we get some catastrofic results.

Levels of trust, it's all games and play until it gets serious, like what to eat or doing something that involves life threatening physics. I take it as playing with a toy. Surely something great have come up from only a few piece of legos?

> And if you're not verifying literally everything an LLM tells you.. are you sure you're learning anything real?

I guess you shouldn't do it that way. But really, so far the topics I've rigorously explored with ChatGPT for example, have been better than your average journalism. What is real?

Post reply on HN