Live data from Hacker News

Introducing deep research

openai.com

121–130 of 445 posts

Re: Introducing deep research

#121
post #97
post #67

Earlier quoted context omitted.

This is one of the actual questions: > In Greek mythology, who was Jason's maternal great-grandfather? https://www.google.com/search?q=In+Greek+mythology%2C+who+wa...

This is a hard question for language models since it targets one of their known weaknesses.

Greek mythology? But seriously please elaborate for my less educated self.

Re: Introducing deep research

#122
post #24

If I understood the graphs correctly, it only achieves 20% pass rate on their internal tests. So I have to wait 30min and pay a lot of money just to sift through walls of most likely incorrect text? Unless the possibility of hallucinations is negligible, this is just way too much content to review at once. The process probably needs to be a lot more iterative.

Here's an example of the type of question it is acheiving 20% on;

The set of natural transformations between two functors F,G ⁣:C→DF,G:C→D can be expressed as the end Nat(F,G)≅∫AHomD(F(A),G(A)). Nat(F,G)≅∫A HomD (F(A),G(A)).

Define set of natural cotransformations from FF to GG to be the coend CoNat(F,G)≅∫AHomD(F(A),G(A)). CoNat(F,G)≅∫AHomD (F(A),G(A)).

Let: - F=B∙(Σ4)∗/F=B∙ (Σ4 )∗/ be the under ∞∞-category of the nerve of the delooping of the symmetric group Σ4Σ4 on 4 letters under the unique 00-simplex ∗∗ of B∙Σ4B∙ Σ4 . - G=B∙(Σ7)∗/G=B∙ (Σ7 )∗/ be the under ∞∞-category nerve of the delooping of the symmetric group Σ7Σ7 on 7 letters under the unique 00-simplex ∗∗ of B∙Σ7B∙ Σ7 .

How many natural cotransformations are there between FF and GG?

Re: Introducing deep research

#123
post #97
post #67

Earlier quoted context omitted.

This is one of the actual questions: > In Greek mythology, who was Jason's maternal great-grandfather? https://www.google.com/search?q=In+Greek+mythology%2C+who+wa...

This is a hard question for language models since it targets one of their known weaknesses.

It's categorically more than a weakness.

Re: Introducing deep research

#124

The accuracy of this tool does not matter. This is exclusively designed for box ticking "reports" that nobody reads and a produced for the sake of itself.

The new term for this is "AI Loopidity", highlighting the unintelligent ouroboros nature of one side using AI to generate content and then another side to consume content.

Re: Introducing deep research

#125
post #56

This is terrifying. Even though they acknowledge the issues with hallucinations/errors, that is going to be completely overlooked by everyone using this, and then injecting the outputs into their own powerpoints. Management Consulting was bad enough before the ability to mass produce these graphs and stats on a whim. At least there was some understanding behind the scenes of where the numbers came from, and sources w…

Think of it like a vaccine. The majority of human written consultant reports are already complete rubbish. Low accuracy, low signal-to-noise, generic platitudes in a quantity-over-quality format. LLMs are innoculating people to this kind of low information value content. People who produce LLM quality output, are now being accused of using LLMs, and can no longer pretend to be adding value. The result of this is goin…

This has been downvoted, but I think there’s actually a chance it might become true (until AGI comes along at least).

Re: Introducing deep research

#126

Surprised more comments aren't mentioning deepseek has this feature (for free) already. Assuming this is why OpenAI scrambled to release it. The examples they have on the page work well on chat.deepseek.com with r1 and search options both enabled. Do I blindly trust the accuracy of either though? Absolutely not. I'm pretty concerned about these models falling into gaming SEO and finding inaccurate facts and presentin…

Not really accurate. The "Search" functionality you're describing in DeepSeek is comparable to OpenAI's existing "Search GPT." OpenAI's recent announcement refers to a more advanced capability, similar to Gemini's existing "deep research" feature. DeepSeek's current offerings are significantly more limited in scope.

Re: Introducing deep research

#127
post #86

The accuracy of this tool does not matter. This is exclusively designed for box ticking "reports" that nobody reads and a produced for the sake of itself.

99% of corpo upper management slide deck work. ai only makes more of this useless pencil-neck board of directors slop.

“Pencil-neck” is a strange insult to use here. How are software developers, or hardware design engineers, or finance workers any less “pencil-neck” than “board of directors”?

Re: Introducing deep research

#128
post #44

Eating popcorn while the scaling doubters scramble to move the goalposts for the nth time.

Its number for one of the benchmark has: **with browsing + python tools Maybe we have different definitions of scaling?

I would consider unsupervised tool usage an achievement

Re: Introducing deep research

#129

So much cynicism and hate in these comments, especially as we are likely witnessing AGI come to life. Its still early, but it might be coming. Where is the excitement? This is an interesting time to be alive. HN has a huge cultural problem that makes this website almost irrelevant. All the interesting takes have moved to X/twitter

> "So much cynicism and hate in these comments, especially as we are likely witnessing AGI come to life. Its still early, but it might be coming. Where is the excitement? This is an interesting time to be alive."

Maybe you can define what "AGI" really means and what the end-game and the economic implications are when 'AGI" is some-what achieved? OpenAI somehow believes that they haven't achieved "AGI" yet, which they continue to do this on purpose for obvious reasons.

The first hint I will give you is that it certainly won't be a utopia.

Re: Introducing deep research

#130
post #33

Earlier quoted context omitted.

Do you have any benchmarks to back up your 'feelings'?

I really don't like the snarky tone of the parent comment. Nonetheless, I don't think this is even something that can easily be benchmarked. I'd recommend you take a look at aider [1], and consider how I drew similarities between it and what's presented here. Has ClosedAI presented any benchmarks / evaluation protocols? [1] https://aider.chat/

Yes, they show benchmarks in the article linked here. Did you not read it?
Post reply on HN