Live data from Hacker News

Introducing deep research

openai.com

281–290 of 445 posts

Re: Introducing deep research

#282
post #270

Earlier quoted context omitted.

That's a lawyer in an adverserial situation. Business consultants tell their clients what they want to believe, the facts be dammed.

it sounds like ai doesn't really change that situation

But the point is it does if you count making it worse changing the situation.

Re: Introducing deep research

#283
post #45

Earlier quoted context omitted.

Interesting, thanks for highlighting! Did not pick up on that. Re:"leading", tho: Effectiveness in this task environment is well beyond the specific model involved, no? Plus they'd be fools (IMHO) to only use one size of model for each step in a research task -- sure, o3 might be an advantage when synthesizing a final answer or choosing between conflicting sources, but there are many, many steps required to get to th…

I don't believe we have any indication that the big offerings (claude.ai, Gemini, operator, tasks, canvas, chatgpt) use multiple models in one call (other than for different modalities like having Gemini create an image). It seems to actually be very difficult technically and I'm curious as to why. I wonder how much of an impact our being still so early in the productization phase of this all is. Like it takes a ton…

...or its all a farce, for now.

Re: Introducing deep research

#284
post #271

Earlier quoted context omitted.

Also "accuracy" as a measure of model's performance used to mean something objective in the traditional ML world. Now with LLMs it is what human evaluators feel about the LLM output?

Traditional ML is no stranger to measuring accuracy in terms of agreement with human evaluators.

A customer churn model or revenue forecast did have hard objective data (ground truth) to compare against - isn't it?

Re: Introducing deep research

#285

Earlier quoted context omitted.

I mean you want it to grill your steak and eat it for you too? I mean I too can complain that my iPhone doesn’t automatically screen out spammers and send my mom flowers on Mother’s Day.

Why doesn't the iPhone screen spammers yet? Pixel has had this feature for a decade.

Pixel hasn’t even been around for a decade.

Re: Introducing deep research

#286

Earlier quoted context omitted.

They’ve only released o3-mini, which is a powerful model but not the full o3 that is being claimed as too expensive to release. That being said, DeepSeek for sure forced their hand to release o3-mini to the public.

o3 mini was previewed in December. Deepseek maybe made them release it a few weeks early but it was already on its way

I guess the question is, did DeepSeek force them to rethink pricing? It's crazy how much cheaper it (v3 and R1) is, but considering they (Deepseek) can't keep up with demand, the price is kind of moot right now. I really do hope they get the hardware to support the API again. The v3 and R1 models that are hosted by others are still cheap compared to the incumbents, but nothing can compete with DeepSeek on price and performance.

Re: Introducing deep research

#287

Earlier quoted context omitted.

It was expensive as they wanted to charge more for it but deepseek has forced their hand

They’ve only released o3-mini, which is a powerful model but not the full o3 that is being claimed as too expensive to release. That being said, DeepSeek for sure forced their hand to release o3-mini to the public.

no they didn't, this was literally all announced in December with a release date for January

Re: Introducing deep research

#288
post #55

Does anyone actually have access to this? It says available for pro users on the website today - I have pro via my employer but see no "deep research" option in the message composer.

Pro user. No access like everyone else. OpenAI is very much in an existential crisis and their poor execution is not helping their cause. Operator or “deep research” should be able to assume the role of a Pro user, run a quick test, and reliably report on whether this is working before the press release right?

How many times are you going to post this exact same comment here? Are you a Chinese bot or something?

Re: Introducing deep research

#289
Setting aside how well it works, I think this is a pretty nice demonstration of how to do UX for an agentic RAG app. I like that the intermediate steps have been pushed out to a sidebar, with updates that both provide some transparency about the process and make the high latency more palatable.

Re: Introducing deep research

#290
post #30

Feels like only a matter of time before these crawlers are blocked from large swathes of the internet. I understand that they’re already prohibited from Reddit and YouTube. If that spreads, this approach might be in trouble.

While people might attempt that, it's going to be an arms race, just like ads vs adblocks. There's already multiple crawlers that present fake user-agent when their original one is blocked. Temptation of more data is just to irresistible to them
Post reply on HN