Live data from Hacker News

Codex for almost everything

openai.com

591–600 of 600 posts

Re: Codex for almost everything

#591
post #508

Earlier quoted context omitted.

In what world would I prefer to give someone access to me via a messaging app rather than a fully-async text SMS message? I don't even love that people can see if you've read their texts now. Fully agree about phone calls though.

I believe that in all of South America people exclusively use WhatsApp to communicate via text because SMS is only used for spam and bad 2FA. Companies are even using WhatsApp for 2FA now instead of SMS and the fact that americans use SMS is viewed as a joke.

Ah yes, those hilarious US Americans, not wanting to trust all their chat communication to Meta...

Edit: literally just saw this article afterwards:

https://www.lbc.co.uk/article/dubai-police-spied-private-wha...

UAE government getting access to private WhatsApp messages (presumably in conjunction with Meta).

Re: Codex for almost everything

#592

Just reading the comments here it's amazing how many people seemingly don't know that Claude Desktop and Cowork basically already does all of this. Codex isn't pioneering these features, it's mostly just catching up.

IMHO no one is really pioneering. A lot more is possible than what is being done. I wrote a blog post about useful agents in a business setting ( https://www.generativestorytelling.ai/blog/posts/useful-corp... ) that highlights AI being proactive. I mean table stakes stuff, why isn't an agent going through all my slack channels and giving me a morning summary of what I should be paying attention to? Why aren't all th…

This makes a lot of sense, but I can't see anyone paying for this because at its simplest layer it's just a Neo4j install + some skills + a local cron job for Claude Desktop. How long will it take for Anthropic to just bake this into Claude Desktop or OpenAI into Codex? Probably not that long.

I keep coming up with good ideas for how to use agents and keep walking away from them because there just is no defensible moat. Everything software related is just going to get totally consumed over the next year.

Re: Codex for almost everything

#593

Earlier quoted context omitted.

IMHO no one is really pioneering. A lot more is possible than what is being done. I wrote a blog post about useful agents in a business setting ( https://www.generativestorytelling.ai/blog/posts/useful-corp... ) that highlights AI being proactive. I mean table stakes stuff, why isn't an agent going through all my slack channels and giving me a morning summary of what I should be paying attention to? Why aren't all th…

THANK YOU. I keep thinking this as well. I'm rolling my own skills to actually make my job easier, which is all about gathering, surfacing, and synthesizing information so I can make quick informed decisions. I feel like nobody is thinking this way and it's bizarre.

The value prop is tenuous and most people still think agents aren't capable of doing this type of work reliably yet (which is... kind of true). You won't get punished too much by users for false positives when summarizing tasks, but you will get absolutely eviscerated for false negatives (e.g. dropping a critical task from the summary). Can you guarantee that your agent won't forget to tell you about something super important?

Re: Codex for almost everything

#594

Earlier quoted context omitted.

IMHO no one is really pioneering. A lot more is possible than what is being done. I wrote a blog post about useful agents in a business setting ( https://www.generativestorytelling.ai/blog/posts/useful-corp... ) that highlights AI being proactive. I mean table stakes stuff, why isn't an agent going through all my slack channels and giving me a morning summary of what I should be paying attention to? Why aren't all th…

You should check out https://pieces.app/ ive been using it for months and I am surprised I have never seen anyone ever talk about it. It does exactly what you are asking for, and it can do it completely locally or with a mixture of frontier models.

Is this just a wrapper on top of Beads?

Re: Codex for almost everything

#595

Earlier quoted context omitted.

Because it's not "PhD-expert level" at all, lol. Even the biggest models (Mythos, GPT-Pro, Gemini DeepThink) are nowhere near the level of effort that would be expected in a PhD dissertation, even in their absolute best domains. Telling it to work out a plan first is exactly how you would supervise an eager but not-too-smart junior coder. That's what AI is like, even at its very best.

I understand that but 1) expert-level performance is how they are being sold; but moreover 2) the level of hand-holding is kind of ridiculous. I'll give another example, Codex decided to write two identical functions linearize_token_output and token_output_linearize. Prompting it not to do things like that feels like plugging holes in a dyke. And through prompting, can you even guarantee it won't write duplicate code…

https://arxiv.org/abs/2604.04721 I like it when researchers confirm my intuition

Re: Codex for almost everything

#596
post #157

Earlier quoted context omitted.

Why would my agent retrieve that lunch menu?

Because it’s hooked up to a microphone in your kitchen & your kid is arguing with you about what lunch they want & they say “Hey [agent], what day is pizza day at [school]?”

I’m not doing that. That would be like giving my child shell access to my system.

Re: Codex for almost everything

#597
post #534
post #422

Earlier quoted context omitted.

Systems have been caught out that review pull requests, that’s a simple and clear one. The more obvious to me for most people is anything you do that interacts with your email without an explicit approve list of emails to read.

Yes, but none of this applies to the local codex agent that runs when I tell it to and has access to my computer. Like: „scan this folder of PDFs and create an excel file with all expenses. Then enter them into my tax software.“ This needs access to very sensitive data and involves a quite complex handling of data. But the only attack vector I see is someone injecting prompts into my invoice files.

Which applies if you were to do this to invoices submitted to you, rather than ones you created, or if you have any way of user info getting into your invoices.

Re: Codex for almost everything

#598
post #500

Earlier quoted context omitted.

We cannot trust identity like we used to here on HN (even pre-LLM-AI I thought we seemed naive.) Unfortunately, we live in a world or anyone or any AI can claim almost anything plausible sounding. Where do we go from here? (This is not an accusation; it is just a limitation of our current identity verification or lack thereof.)

You can confirm that the people who say things are in a position to know.

> You can confirm that the people who say things are in a position to know.

What is the above commenter's sense of how well one can 'confirm' such a thing?

Looking at an HN account and its comment history provides some signal, but this doesn't satisfy me, given the incentives at play here. We're talking about OpenAI, a ~$800B company. Reputation matters a lot. The stakes are higher than e.g. "does so and on really work at e.g. Mozilla and know about the details of a messy Rust governance issue?" (to pick a deliberately lower stakes example).

When we decide what to let into our brains around OpenAI, Anthropic, etc, the bar needs to be higher than i.e. "does a HN account seem to be consistent with someone who works at OpenAI?". (I'm not sure if this is the above commenter's position or close to it?)

We need to be able to have stronger proofs, preferably ones with cryptography and credibility rooted in a legitimate trust model. In 2026, this is certainly possible technically, if a platform made this a priority. The barriers are largely social, cultural, and economic.

HN does not make real-world identity a priority. There might be some workarounds for posting information in one's profile, but practically speaking, I'm not seeing how this would work and what levels of identity it would bolster. Am I missing something?

If I start hand-waving I might dream up something like the following ... Maybe someone could stitch something together with a trusted content time-stamping server and prove they control an OpenAI email address and also provide that cryptographic evidence on their HN profile. It sounds ... practically unappealing at best. I haven't seen this done. Maybe I'm overlooking a good way. I'm all ears. We're going to need better solutions.

Re: Codex for almost everything

#599
post #598

Earlier quoted context omitted.

You can confirm that the people who say things are in a position to know.

> You can confirm that the people who say things are in a position to know. What is the above commenter's sense of how well one can 'confirm' such a thing? Looking at an HN account and its comment history provides some signal, but this doesn't satisfy me, given the incentives at play here. We're talking about OpenAI, a ~$800B company. Reputation matters a lot. The stakes are higher than e.g. "does so and on really wo…

They work at OpenAI, what more do you want? For what it’s worth, I can independently corroborate that the announcement was planned in advance.

Re: Codex for almost everything

#600
post #508

Earlier quoted context omitted.

I believe that in all of South America people exclusively use WhatsApp to communicate via text because SMS is only used for spam and bad 2FA. Companies are even using WhatsApp for 2FA now instead of SMS and the fact that americans use SMS is viewed as a joke.

Ah yes, those hilarious US Americans, not wanting to trust all their chat communication to Meta... Edit: literally just saw this article afterwards: https://www.lbc.co.uk/article/dubai-police-spied-private-wha... UAE government getting access to private WhatsApp messages (presumably in conjunction with Meta).

what's more likely:

- whatsapp/signal encryption was broken without anyone noticing all in order to nab... some guy posting pictures of bomb damage

- someone snitched and the govt lied

Post reply on HN