Earlier quoted context omitted.
One of the dirty secrets of a lot of these "code adjacent" areas is that they have very little testing. If a data science team modeled something incorrectly in their simulation, who's gonna catch it? Usually nobody. At least not until it's too late. Will you say "this doesn't look plausible" about the output? Or maybe you'll be too worried about getting chided for "not being data driven" enough. If an exec tells an i…
I would say that although Claude may hallucinate at least it can be told to test the scripts. Many data scientists will just balk at the idea of testing a crazy excel workbook with lots of formulas that they themselves inherited.
Two kinds of AI users are emerging
331–340 of 358 posts
Re: Two kinds of AI users are emerging
#332Earlier quoted context omitted.
>I've seen Claude hallucinate running test suites before. This reminded of something that happened to me last year. Not Claude (I think it was GPT 4.0 maybe?), but I had it running in VS Code's Copilot and asked it to fix a bug then add a test for the case. Well, it kept failing to pass its own test, so on the third try, it sat there "thinking" for a moment, then finally spit out the command `echo "Test Passed!"`, ex…
I've been using Claude Code with Opus 4.5 a lot the last several months and while it's amazingly capable it has a huge tendency to give up on tests. It will just decide that it can commit a failing test because "fixing it has been deferred" or "it's a pre-existing problem." It also knows that it can use `HUSKY=0 git commit ...` to bypass tests that are run in commit hooks. This is all with CLAUDE.md being very specif…
1) it wants to run X command
2) it notices a hook preventing it from running X
3) it creates a Python application or shell script that does X and runs it instead
Whoops.
Re: Two kinds of AI users are emerging
#333Earlier quoted context omitted.
It depends on how easily testable the Excel is. If Claude has the ability to run both the Excel and the Python with different inputs, and check the outputs, it's stunningly likely to be able to one-shot it.
And also - who understands the system now? Does anyone know Python at this shop? Is it someone’s implicit duty to now learn Python, or is the LLM now the de facto interface for modifying the system? When shit hits the fan and execs need answers yesterday , will they jump to using the LLM to probabilistically make modifications to the system, or will they admit it was a mistake and pull Excel back up to deterministica…
When shit hits the fan, execs need answers yesterday and the 30 sheet Excel monstrosity is producing the wrong numbers - who fixes it?
It was done by Sue, who left the company 4 years ago, people have been using it since and nobody really understands it.
Re: Two kinds of AI users are emerging
#334Earlier quoted context omitted.
Yea I am an ENFP. While I don’t think MBTI is scientific, it captures perfectly that I have the tendency to think out loud. LLMs make me think out loud way better. Best rubber duck ever.
Felt very weird reading this on HN and not r/ENFPmemes. I agree completely.
So I learned that you can definitely glean some insights from it. One insight I have is: I'm a "talk out loud thinker". I don't really value that as an identity thing but it is definitely something I notice that I do. I also think a lot of things in my mind, but I tend to think out loud more than the average person.
So yea, that's how pseudo science can sometimes still lead to useful insights about one particular individual. Same thing with philosophy really, usually also not empirically tested (I do think it has a stronger academic grounding but to call philosophy a science is... a bit... tricky... in many cases. I think the common theme is that it's also usually not empirically grounded but still really useful).
Re: Two kinds of AI users are emerging
#335> I helped one recently almost one-shot[3] converting a 30 sheet mind numbingly complicated Excel financial model to Python with Claude Code. I'm sure Claude Code will happily one-shot that conversion. It's also virtually guaranteed to have messed up vital parts of the original logic in the process.
Doesn't it help you sleep at night that your 401k might be managed by analysts #yoloing their financial modeling tools with an LLM?
I have seen Excel used for financial planning
I have seen Excel used for managing people's health data.
I have BUILT a test suite for a government offical use communication device - inside Excel. The original was a mish-mash of Excel formulas and VBA. I improved the VBA part of it by adding a web cam to the mix.
I don't sleep well at night knowing how many very very essential things are running on top of Excel sheets passed down like stories around a campfire.
Re: Two kinds of AI users are emerging
#336Terrifying that people are creating financial models with AI when they don’t have the skills to verify the model does what they expect
I work for a huge bank: The actual terrifying thing is that these financial models semi-technical management are creating in Excel are actually relied upon in the first place. If you convert bullshit from Excel to Python it's still bullshit. There's a reason why Claude can one-shot it and no one questions the result :D
Nobody will see that on sheet 27 cell FG456 is actually a static number that Brian typoed in there in 2019 and not a formula.
Re: Two kinds of AI users are emerging
#337I don't see a divergence, from what I can tell a lot of people have only just started using agents in the past 3-4 months when they got good enough that it was hard to say otherwise. Then there's stuff like MCP, which never seemed good and was entirely driven by people who talked more about it than used it. There also used to be stuff like langchain or vector databases that nobody talks about anymore, maybe they're s…
While I agree that the MCP craze was a bit off-putting, I think that came mostly from people thinking they can sell stuff in that space. If you view it as a protocol and not much else, things change. I've seen great improvements with just two MCP servers: context7 and playwright. The first is great on planning sessions and leads to better usage of new-ish libraries, and the second is giving the model a feedback loop.…
What I want is a Skill that leverages a normal CLI executable that gives the LLM the same capabilities of browser use.
Re: Two kinds of AI users are emerging
#338I don't see a divergence, from what I can tell a lot of people have only just started using agents in the past 3-4 months when they got good enough that it was hard to say otherwise. Then there's stuff like MCP, which never seemed good and was entirely driven by people who talked more about it than used it. There also used to be stuff like langchain or vector databases that nobody talks about anymore, maybe they're s…
What‘s used instead of MCP in reality? Just REST or other existing API things?
The LLM agent can make sense of the text document, figure out the actual tool calls and use them.
And you, the MCP server operator, can change the "API" at any time and the client (LLM agent) will just automatically adjust.
Re: Two kinds of AI users are emerging
#339Earlier quoted context omitted.
I guess a quarter of the smartphone market (leader), half of the tablet market (leader) and a tenth of the global pc market (2nd place) / 6th of the usa/europe market (2nd place) being a small market share is a take.
>a quarter of the smartphone market (leader) Android is by far the leader. >half of the tablet market (leader) Half does not make someone a "leader" >a tenth of the global pc market (2nd place) 2nd place?? They're last place, by a wide margin . >6th of the usa/europe market (2nd place) Also last place. I guess the reality distortion field is still alive and well.
If half doesn’t make you leader what does? Maybe you should elaborate your definition of leader? For me it’s “has the highest market share”. And in that definition half is necessarily true.
It’s funny that for PC’s you went for manufacturers (apple is 4th) but for mobile you went for OS (Apple is 2nd). On mobile devices, Apple is 1st, having double market share compared to 2nd place (samsung).
The need to paint Apple as purely a marketing company always fascinated me. Marketing is a big part of who they are though.
[1] https://en.wikipedia.org/wiki/Market_share_of_personal_compu...
Re: Two kinds of AI users are emerging
#340Earlier quoted context omitted.
This is me, but for writing code. I own a business, and I use Claude Code to build internal tools for myself. Don't care about code quality; never seen the code. I care if the tools do the things I want them to do, and they verifiably do.
How do you verify them? How do you verify they do not create security risks?
On the verification front, a few examples:
1. I built an app that generates listing images and whitebox photos for my products. Results there are verifiable for obvious reasons.
2. I use Claude Code to do inventory management - it has a bunch of scripts to pull the relevant data from Amazon then a set of instructions on how to project future sales and determine when I should reorder. It prints the data that it pulls from Amazon to the terminal, so that's verifiable. In terms of following the instructions on coming up with reorder dates, if it's way off, I'm going to know because I'm very familiar with the brands that I own. This is pretty standard manager/subordinate stuff - I put some trust in Claude to get it right, but I have enough context to know if the results are clearly bad. And if they're only off by a little, then the result is I incur some small financial penalty (either I reorder too late and temporarily stock out or I reorder too early and pay extra storage fees). But that's fine - I'm choosing to make that tradeoff as one always does when one hands off work.
3. I gave Claude Code a QuickBooks API key and use it to do my books. This one gets people horrified, but again, I have enough context to know if anything's clearly wrong, and if things are only slightly off then I will potentially pay a little too much in taxes. (Though to be fair it's also possible it screws up the other way, I underpay in taxes and in that case the likeliest outcome is I just saved money because audits are so rare.)