Live data from Hacker News

Introducing deep research

openai.com

321–330 of 445 posts

Re: Introducing deep research

#321

For “deep research” I’m also reading “getting the answers right”. Most people I talk to are at the point now where getting completely incorrect answers 10% of the time — either obviously wrong from common sense, or because the answers are self contradictory — undermines a lot of trust in any kind of interaction. Other than double checking something you already know, language models aren’t large enough to actually kno…

My worry is that all these recent capabilities attempt to minimize hallucinations by relying on extensive web search, however web itself is being actively degraded by unfiltered LLM output. After a certain point running your research agent against a ~5-year-old snapshot of the web will be strictly more accurate (for non-current affairs queries) than querying live web.

Re: Introducing deep research

#323
post #303

For “deep research” I’m also reading “getting the answers right”. Most people I talk to are at the point now where getting completely incorrect answers 10% of the time — either obviously wrong from common sense, or because the answers are self contradictory — undermines a lot of trust in any kind of interaction. Other than double checking something you already know, language models aren’t large enough to actually kno…

I mean this all falls down due to the need of verification: "Limitations Deep research unlocks significant new capabilities, but it’s still early and has limitations. It can sometimes hallucinate facts in responses or make incorrect inferences" How do I know which parts are false? It will take as long to verify as to research!

[dead]

Re: Introducing deep research

#324

It is actually interesting for people working in academia. I would like to test it but no way I can afford $200/m right now. Can someone test it with this prompt. "As a research assistant with comprehensive knowledge of particle physics, please provide a detailed analysis of next-generation particle collider projects currently under consideration by the international physics community. The analysis should encompass t…

Ran it for you using o3-high! Here's a link to the results: https://chatgpt.com/share/67a0b227-8ee4-800f-a8ed-882e7bab97...

Hopefully helpful, happy to test others for you :)

-- Raw output --

Understood. I will conduct a detailed technical analysis of next-generation particle collider projects, focusing on the Future Circular Collider (FCC), International Linear Collider (ILC), Compact Linear Collider (CLIC), Muon Collider proposals, and other significant initiatives.

The report will include: - Planned energy ranges and collision types - Estimated construction and operation timelines - Technical advantages and challenges - Approximate costs - Key physics goals - Review of technical design reports and feasibility studies - International collaboration and support - Comparative analysis on feasibility, cost-benefit, scientific potential, infrastructure needs, and environmental impact - Long-term considerations, including upgrade paths and integration with existing facilities

The analysis will be formatted as a structured technical summary, suitable for presentation at a topical meeting of particle physicists, with comparative tables and figures where relevant.

I will update you when the report is ready.

# Next-Generation Particle Collider Projects: A Comparative Technical Analysis

## Introduction High-energy physics is charting plans for new colliders to extend beyond the LHC and its High-Luminosity upgrade (HL-LHC). The discovery of the 125 GeV Higgs boson has motivated proposals for a dedicated “Higgs factory” lepton collider to study its properties with precision, as well as future energy-frontier machines to search for new physics. Several next-generation collider initiatives are under consideration internationally as of 2024. These include CERN’s proposed *Future Circular Collider (FCC)*, Japan’s *International Linear Collider (ILC)*, CERN’s *Compact Linear Collider (CLIC)*, various designs for a *Muon Collider*, China’s *Circular Electron-Positron Collider (CEPC)* and its successor *Super Proton-Proton Collider (SppC)*, among others. Each proposal differs in collision type (electron-positron, proton-proton, muon-muon, etc.), energy scale, technology, timeline, cost, and physics focus. This summary reviews each project’s key parameters – *planned energy ranges, collision types, timeline, technical advantages/challenges, cost, and physics goals* – based on technical design reports and feasibility studies. A comparative analysis then contrasts their *technical feasibility, cost-benefit, scientific potential for discoveries, timeline to first data, infrastructure needs, and environmental impact*, highlighting the relative strengths and weaknesses of each approach. We also discuss long-term implications such as upgrade paths, flexibility for future modifications, and integration with existing infrastructure.

(Citations refer to official reports and peer-reviewed sources using the format 【source†lines】.)

## Future Circular Collider (FCC) – CERN - *Type and Energy:* The FCC is a *proposed 100 km circular collider* at CERN that would be realized in stages. The first stage, *FCC-ee*, is an electron-positron ($e^+e^-$) collider with center-of-mass energy tunable from ~90 GeV up to 350–365 GeV, covering the Z boson pole, WW threshold, Higgs production (240 GeV), and top-quark pair threshold (~350 GeV). A second stage, *FCC-hh*, would use the same tunnel for a proton-proton collider at up to *100 TeV* center-of-mass energy (an order of magnitude above the LHC’s 14 TeV). Heavy-ion collisions (e.g. Pb–Pb) are also envisioned. An *FCC-eh* option (electron-hadron collisions) is considered by adding a high-energy electron injector to collide with the proton beam. This integrated FCC program thus spans both *precision lepton* collisions and *energy-frontier hadron* collisions.

- *Timeline:* The conceptual schedule foresees *FCC-ee construction in the 2030s* and a start of operations by around *2040* (as the LHC/HL-LHC program winds down). According to the FCC Conceptual Design Report, an $e^+e^-$ Higgs factory could begin delivering physics in ~2040, running for 15–20 years. The *hadron collider FCC-hh* would be constructed subsequently (using the same tunnel and upgraded infrastructure), aiming for *first proton-proton collisions in the late 2050s】. This staged approach (lepton collider first, hadron later) mirrors the successful *LEP–LHC sequence*, leveraging the $e^+e^-$ machine to produce great precision data (and to build infrastructure) before pushing to the highest energies with the hadron machine. ...

(Too long for HN to write more)

Re: Introducing deep research

#325
post #317

For “deep research” I’m also reading “getting the answers right”. Most people I talk to are at the point now where getting completely incorrect answers 10% of the time — either obviously wrong from common sense, or because the answers are self contradictory — undermines a lot of trust in any kind of interaction. Other than double checking something you already know, language models aren’t large enough to actually kno…

> Most people I talk to are at the point now where getting completely incorrect answers 10% of the time A year back that number was 30%, and a couple of years back it was 60%. There will be a point where it'll be good enough. There are also better and better ways to verify answers these days. It'll never be a solution for everything, but that's similar to many engineering problems we have: for example, ORMs aren't gr…

It contributes little to discuss a hypothetical future. Maybe we'll have fusion energy, delivery drones, everyone using VR, etc. Maybe we will go into a deep recession due to trade wars, or maybe not.

The meaningful discussion is about how they perform NOW and the edge cases that have persisted since GPT-2 which no one has yet found a good solution for.

Re: Introducing deep research

#326
post #55

Does anyone actually have access to this? It says available for pro users on the website today - I have pro via my employer but see no "deep research" option in the message composer.

I have access as of ~3 hours ago. Using the Win desktop app too, which is behind on some features (Operator, tasks). I open up any of the models and it shows up as a `(Deep research)` tag on the input field next to the web search option. Didn't clear cache or anything.

Re: Introducing deep research

#327
post #317

Earlier quoted context omitted.

> Most people I talk to are at the point now where getting completely incorrect answers 10% of the time A year back that number was 30%, and a couple of years back it was 60%. There will be a point where it'll be good enough. There are also better and better ways to verify answers these days. It'll never be a solution for everything, but that's similar to many engineering problems we have: for example, ORMs aren't gr…

It contributes little to discuss a hypothetical future. Maybe we'll have fusion energy, delivery drones, everyone using VR, etc. Maybe we will go into a deep recession due to trade wars, or maybe not. The meaningful discussion is about how they perform NOW and the edge cases that have persisted since GPT-2 which no one has yet found a good solution for.

We already have delivery drones though.

I disagree though, it is useful as this problem has been whittled down and I think there is expectation that there will be continued effort. Its of course worth discussing but I find that for my workflows, I rarely encounter issues with hallucinations, they certainly exist but its gotten to a point that I don't have major issue with it.

Re: Introducing deep research

#328

For “deep research” I’m also reading “getting the answers right”. Most people I talk to are at the point now where getting completely incorrect answers 10% of the time — either obviously wrong from common sense, or because the answers are self contradictory — undermines a lot of trust in any kind of interaction. Other than double checking something you already know, language models aren’t large enough to actually kno…

> What I’m looking for is therefore not just the correct answer, but the correct answer in an amount of time that’s faster than it would take me to research the answer myself, and also faster than it takes me to verify the answer given by the machine.

This is why I haven't found AI tools very useful. I find my self spending more time verifying and fixing it's answers than I would have just doing or learning the darn thing myself.

Re: Introducing deep research

#329

Not sure if people picked up on it, but this is being powered by the unreleased o3 model. Which might explain why it leaps ahead in benchmarks considerably and aligns with the claims o3 is too expensive to release publicly. Seems to be quite an impressive model and the leading out of Google, DeepSeek and Perplexity.

> Which might explain why it leaps ahead in benchmarks considerably and aligns with the claims o3 is too expensive to release publicly It's the only tool/system (I won't call it an LLM) in their released benchmarks that has access to tools and the web. So, I'd wager the performance gains are strictly due to that. If an LLM (o3) is too expensive to be released to the public, why would you use it in a tool that has to…

>why would you use it in a tool that has to make hundreds of inference calls to it to answer a single question? You'd use a much cheaper model.

The same reason a lot of people switched to GPT-4 when it came out even though it was much more expensive than 3 - doesn't matter how cheap it is if it isn't good enough/much worse.

Re: Introducing deep research

#330

For “deep research” I’m also reading “getting the answers right”. Most people I talk to are at the point now where getting completely incorrect answers 10% of the time — either obviously wrong from common sense, or because the answers are self contradictory — undermines a lot of trust in any kind of interaction. Other than double checking something you already know, language models aren’t large enough to actually kno…

> What I’m looking for is therefore not just the correct answer, but the correct answer in an amount of time that’s faster than it would take me to research the answer myself, and also faster than it takes me to verify the answer given by the machine. This is why I haven't found AI tools very useful. I find my self spending more time verifying and fixing it's answers than I would have just doing or learning the darn…

It is added cognitive load, but there is a lot of value in async tasks if you can trust the output or if the opportunity cost of validating is low.

The challenge with something like this for research, in its current state, is you’ll need to go double check it because you don’t trust it and it will end up effectively being a list of links.

It’s progress though and evidently good enough to find a sweet NSX in Japan, which is all some really need.

Post reply on HN