Earlier quoted context omitted.
The same way you test any system - you find a sampling of test subjects, have them interact with the system and then evaluate those interactions. No system is guaranteed to never fail, it's all about degree of effectiveness and resilience. The thing is (and maybe this is what parent meant by non-determinism, in which case I agree it's a problem), in this brave new technological use-case, the space of possible interac…
> radically different nature of our AI interlocutor It's the training data that matters. Your "AI interlocutor" is nothing more than a lossy compression algorithm.
Mamdani to kill the NYC AI chatbot caught telling businesses to break the law
41–50 of 67 posts
Re: Mamdani to kill the NYC AI chatbot caught telling businesses to break the law
#42Why did NYC release it in the first place? Did they not QA it? Or was it perhaps one of those cases where they found issues, but the only way to really know for sure that the deleterious impact is significant enough by pushing it to prod?
Re: Mamdani to kill the NYC AI chatbot caught telling businesses to break the law
#43Earlier quoted context omitted.
Remember that many people are heavily are happy-path biased. They see a good result once and say "that's it, ship it!" I'm sure they QA'd it, but QA was probably "does this give me good results" (almost certainly 'yes' with an LLM), not "does this consistently not give me bad results".
> almost certainly 'yes' with an LLM LLMs can handle search because search is intentionally garbage now and because they can absorb that into their training set. Asking highly specific questions about NYC governance, which can change daily, is almost certainly 'not' going to give you good results with an LLM. The technology is not well suited to this particular problem. Meanwhile if an LLM actually did give you good…
But as you say that LLM's cant handle search. One of the things that I can't understand and I hope you help in is that this doesn't have to be this way.
Kagi exists (I think I like the product/product idea even though I haven't bought it but I have tried it). Kagi's assistants can actually use Kagi search engine itself which is really customizable and you can almost have a lot of search settings filtered and Kagi overall is considered by many people as giving good search.
Not to be a sponsor of kagi or anything but if this is such a really big problem assuming that NYC literally had to kill a bot because of it & the reason you mention is the garbage in garbage out problem of search happening.
I wonder if Kagi could've maybe helped in it. I think that they are B-corp so they would've really appreciated the support itself if say NYC would've used them as a search layer.
Re: Mamdani to kill the NYC AI chatbot caught telling businesses to break the law
#44> A spokesperson for the mayor, Dora Pekec, confirmed in a text message that the new administration plans to take down the chatbot. She said a member of the Mamdani transition team had seen reporting on the bot from The Markup and THE CITY and presented it to the mayor as a possible place to save funds. Journalism works.
Journalism teed up an easy way for an incoming politician to dunk on his predecessor, if you'll forgive the mixed metaphor. Not that I'm opposed to any part of it, just that this was an easy scenario for "journalism" to "work" in.
https://en.wikipedia.org/wiki/Investigations_into_the_Eric_A...
Re: Mamdani to kill the NYC AI chatbot caught telling businesses to break the law
#45Why did NYC release it in the first place? Did they not QA it? Or was it perhaps one of those cases where they found issues, but the only way to really know for sure that the deleterious impact is significant enough by pushing it to prod?
Remember that many people are heavily are happy-path biased. They see a good result once and say "that's it, ship it!" I'm sure they QA'd it, but QA was probably "does this give me good results" (almost certainly 'yes' with an LLM), not "does this consistently not give me bad results".
https://dl.acm.org/doi/epdf/10.1145/3780063.3780066 (PDF loads slow....)
Re: Mamdani to kill the NYC AI chatbot caught telling businesses to break the law
#46What I do know is that if it was done my way it would be pretty easy for it to do what the Google AI does. Say it's not responsible, give links for humans to fact check it. I've noticed a dramatic drop in hallucinations after it had to provide links to its sources. Still not 0, though.
Re: Mamdani to kill the NYC AI chatbot caught telling businesses to break the law
#47I always ask this question about these bots: is the literature the training data or is the understanding of literature the training data? Meaning, sure you trained the bot on the current rules and regulations. But does that mean the model weights contain them? Such that really is a guess at legal accuracy? Or is it trained to be a lawyer and understand the docs which sit outside the model? Every time I've asked the a…
I thought Gemini just started providing citations in the last few months. Are you saying they should have beaten Google to the punch on this? As part of the $500,000 budget?
Re: Mamdani to kill the NYC AI chatbot caught telling businesses to break the law
#48Earlier quoted context omitted.
The same way you test any system - you find a sampling of test subjects, have them interact with the system and then evaluate those interactions. No system is guaranteed to never fail, it's all about degree of effectiveness and resilience. The thing is (and maybe this is what parent meant by non-determinism, in which case I agree it's a problem), in this brave new technological use-case, the space of possible interac…
> radically different nature of our AI interlocutor It's the training data that matters. Your "AI interlocutor" is nothing more than a lossy compression algorithm.
Re: Mamdani to kill the NYC AI chatbot caught telling businesses to break the law
#49He is turning out to be a benevolent, law-abiding mayor that just happens to be communist.
Re: Mamdani to kill the NYC AI chatbot caught telling businesses to break the law
#50Why did NYC release it in the first place? Did they not QA it? Or was it perhaps one of those cases where they found issues, but the only way to really know for sure that the deleterious impact is significant enough by pushing it to prod?
https://apnews.com/article/eric-adams-crypto-meme-coin-942ba...
I think he is simply not very bright, and got mesmerized by all the shiny promises AI and crypto makes without the slightest understanding of how it actually works. I do not understand how he got into office in the first place.