Earlier quoted context omitted.
> you'd still have to do 5% of the work No, you still have to do 100% of the work.
You simply do not. You do the math yourself to calculate 2(n) for n in [1, 2, 3, 4] and get [2, 5, 6, 8]. You plug it into your (75% accurate) unreliable calculator and get [3, 4, 6, 8]. You now know that you only need to recheck the first two (50%) of the entries.
Cursor IDE support hallucinates lockout policy, causes user cancellations
521–530 of 635 posts
Re: Cursor IDE support hallucinates lockout policy, causes user cancellations
#522Earlier quoted context omitted.
> Here's a machine that verifies 95% of calculations, but you'd still have to do 5% of the work. The problem is that you don't know which 5% are wrong. The AI is confidently wrong all the time. So the only way to be sure is to double check everything, and at some point its easier to just do it the right way. Sure, some things don't need to be perfect. But how much do you really want to risk? This company thought a li…
>The problem is that you don't know which 5% are wrong This is not a problem in my unreliable calculator use-cases; are you disputing that or dropping the analogy? Because I'd love to drop the analogy. You mention IDEs- I routinely use IntelliJ's tab completion, despite it being wrong >>5% of the time. I have to manually verify every suggestion. Sometimes I use it and then edit the final term of a nested object acces…
If you use an unreliable calculator to sum a list of numbers, you then need to use a reliable method to sum the numbers to validate that the unreliable calculator's sum is correct or incorrect.
Re: Cursor IDE support hallucinates lockout policy, causes user cancellations
#523Earlier quoted context omitted.
Hey both things can be true. It’s a long ways from the AI renaissances of the past. There’s areas LLMs make a lot of sense. I just don’t find them to be great pair programming partners yet.
I think people are kind of kidding themselves here. For Go and Python, two extraordinarily common languages in production software, it would be weird for me at this point not to start with LLM output. Actually building an entire application, soup-to-nuts, vibe-code style? No, I wouldn't do that. But having the LLM writing as much as 80% of the code, under close supervision, with a careful series of prompts (like, "ok…
Re: Cursor IDE support hallucinates lockout policy, causes user cancellations
#524Earlier quoted context omitted.
https://www.anthropic.com/research/tracing-thoughts-language... The section about hallucinations is deeply relevant. Namely, Claude sometimes provides a plausible but incorrect chain-of-thought reasoning when its “true” computational path isn’t available. The model genuinely believes it’s giving a correct reasoning chain, but the interpretability microscope reveals it is constructing symbolic arguments backward from…
> Knowledge emerges from symbolic coherence, linguistic agreement, and social plausibility rather than purely from logical coherence or factual correctness. This just seems like a redefinition of the word "knowledge" different from how it's commonly used. When most people say "knowledge" they mean beliefs that are also factually correct.
Re: Cursor IDE support hallucinates lockout policy, causes user cancellations
#525(Cursor cofounder) Apologies - something very clearly went wrong here. We’ve already begun investigating, and some very early results: * Any AI responses used for email support are now clearly labeled as such. We use AI-assisted responses as the first filter for email support. * We’ve made sure this user is completely refunded - least we can do for the trouble. For context, this user’s complaint was the result of a r…
> * Any AI responses used for email support are now clearly labeled as such. We use AI-assisted responses as the first filter for email support. Don't use AI. Actually care. Like, take a step back, and realise you should give a shit about support for a paid product. Don't get me wrong: AI is a very effective tool, *for doing things you don't care about*. I had to do a random docker compose change the the other day. I…
I agree with this. Also, whenever I care about code, I don’t use AI. So I very rarely use AI assistants for coding.
I guess this is why Cursor is interested in making AI assistants popular everywhere, they don’t want the association that “AI assisted” means careless. Even when it does, at least with today’s level of AI.
Re: Cursor IDE support hallucinates lockout policy, causes user cancellations
#526Earlier quoted context omitted.
> * Any AI responses used for email support are now clearly labeled as such. We use AI-assisted responses as the first filter for email support. Don't use AI. Actually care. Like, take a step back, and realise you should give a shit about support for a paid product. Don't get me wrong: AI is a very effective tool, *for doing things you don't care about*. I had to do a random docker compose change the the other day. I…
The amount paid is still pretty trivial. I wouldn’t expect much human support for most SaaS products costing $20 a month.
Re: Cursor IDE support hallucinates lockout policy, causes user cancellations
#527Earlier quoted context omitted.
> * Any AI responses used for email support are now clearly labeled as such. We use AI-assisted responses as the first filter for email support. Don't use AI. Actually care. Like, take a step back, and realise you should give a shit about support for a paid product. Don't get me wrong: AI is a very effective tool, *for doing things you don't care about*. I had to do a random docker compose change the the other day. I…
They’re like a team of 10 people with thousands, if not hundreds of thousands of users. “Actually care” is not a viable path to success here.
Re: Cursor IDE support hallucinates lockout policy, causes user cancellations
#528For a support agent to actually be useful beyond that, they need some leeway to make decisions unilaterally, sometimes in breach of "protocol", when it makes sense. No company with a significant level of complexity in its interactions with customers can have an actually complete set of protocols that can describe every possible scenario that can arise. That's why you need someone with actual access inside the company, the ability to talk to the right people in the company should the need arise, a general ability(and latitude) to make decisions based on common sense, and an overall understanding of the state of the company and what compromises can be made somewhat regularly without bankrupting it. Good support is effectively defined by flexibility, and diametrically opposed to following a strict set of rules. It's about solving issues that hadn't been thought of until they happened. This is the kind of support that gets you customer loyalty.
No company wants to give an LLM the power given to a real support agent, because they can't really be trusted. If the LLM can make unilateral decisions, what if it hallucinated and gives the customer free service for life? Now they have to either eat the cost of that, or try to withdraw the offer, which is likely to lose them that customer. And at the end of all that, there's no one to hold liable for the fuckup(except I guess the programmers that made the chatbot). And no one wants the LLM support agent to be sending them emails all day the same way a human support agent might. So what you end up with is just a slightly nicer natural language interface to a set of predefined account actions and FAQ items. In other words, exactly what you get from clickfarms in Southern Asia or even a phone tree, except cheaper. And sure, that can be useful, just to filter out the usual noise, and buy your real support staff more time to work on the cases where they're really needed, but that's it.
Some companies, like Netflix and Google(Google probably has better support for business customers, never used it, so I can't speak to it. I've only Bangalored(zing) my head against a wall with google support as a lowly consumer who bought a product), seem to have no support staff beyond the clickfarms, and as a result their support is atrocious. And when they replace those clickfarms with LLMs, support will continue to be atrocious, maybe with somewhat better English. And it'll save them money, and because of that they'll report it as a rousing success. But for customers, nothing will have changed.
This is pretty much what I predicted would happen a few years ago, before every company and its brother got its own LLM based support chatbot. And anecdotally, that's pretty much what has happened. For every support request I've made in the last year, I can remember 0 that were sorted out by the LLM, and a handful that were sorted out by humans after the LLM told me it was impossible to solve.
Re: Cursor IDE support hallucinates lockout policy, causes user cancellations
#529This is extremely funny. AI can't have accountability. Good luck with that. Use AI to augment but don't really replace it as a 100% system if you can't predict and own up the failure rate. My advice would be to use more configurable tools with less interest on selling fake perfection. Aider works.
Sure it can. You just have to bake into the reward function "if you do the wrong thing, people will stop using you, therefore you need to avoid the wrong thing".
Then you wind up at self-preservation and all the wholly shady shit that comes along with it.
I think the AI accountability problem is the crux of the "last-mile" problem in AI, and I don't think you can necessarily solve it without solving it in a way that produces results you don't want.
Re: Cursor IDE support hallucinates lockout policy, causes user cancellations
#530Cursor is trapped in a cat and mouse game against "hacks" where users create new accounts and get unlimited use. The repo was even trending on Github ( https://github.com/yeongpin/cursor-free-vip ). Sadly, Cursor will always be hampered by maintaining it's own VSCode fork. Others in this niche are expanding rapidly and I, myself, have started transitioning to using Roo and Cline.