Live data from Hacker News

Claude 2.1

anthropic.com

251–260 of 339 posts

Re: Claude 2.1

#252
post #12

Great but it stills leaves the problem of accessing it. I have never heard back on access from Anthropic's website and still waiting on the request through Bedrock. Not sure the success rate of others but it seems impossible as a business to get access to the API. Not a downplay on their announcement but with how difficult it seems to get API access its hard to see the improvement.

I requested access through Bedrock and had it minutes later. It's an automated process.

Same here but still waiting the request model access button is now "Use case details submitted". Glad you had success this route.

This is why we have enjoyed using OpenAI. Easy signup and access.

Re: Claude 2.1

#253
post #183

Earlier quoted context omitted.

Cars nowadays have radars and cameras that (for the most part) prevent you from running over pedestrians. Is that also a tool refusing to work? I'd argue a line needs to be drawn somewhere, LLMs do a great job of providing recipes for dinner but maybe shouldn't teach me how to build a bomb.

> LLMs do a great job of providing recipes for dinner but maybe shouldn't teach me how to build a bomb. Why not? If someone wants to make a bomb, they can already find out from other source materials. We already have regulations around acquiring dangerous materials. Knowing how to make a bomb is not the same as making one (which is not the same as using one to harm people.)

It's about access and command & control. I could have the same sentiment as you, since in high school, friends & I were in the habit of using our knowledge from chemistry class (and a bit more reading; waay pre-Internet) to make some rather impressive fireworks and rockets. But we never did anything destructive with them.

There are many bits of technology that can destroy large numbers of people with a single action. Usually, those are either tightly controlled and/or require jumping a high bar of technical knowledge, industrial capability, and/or capital to produce. The intersection of people with that requisite knowledge+capability+capital and people sufficiently psycopathic to build & use such destructive things approaches zero.

The same was true of hacking way back when. The result was interesting, sometimes fun, and generally non-destructive hacks. But now, hacking tools have been developed to the level of copy+paste click+shoot. Script kiddies became a thing. And we now must deal with ransomeware gangs of everything from nation-state actors down to rando teenage miscreants, but they all cause massive damage.

Extending copy+paste click+shoot level knowledge to bombs and biological agents is just massively stupid. The last thing we need is having a low intelligence bar required to have people setting off bombs & bioweapons on their stupid whims. So yes, we absolutely should restrict these kinds of recipe-from-scratch responses.

In any case, if you really want to know, I'm sure that, if you already have significant knowledge and smarts, you can craft prompts to get the LLM to reveal the parts you don't know. But this gets back to raising the bar, which is just fine.

Re: Claude 2.1

#254
post #174

Earlier quoted context omitted.

> I decide how to use my tools, not the other way 'round. This is the key. The only sensible model of "alignment" is "model is aligned to the user", not e.g. "model is aligned to corporation" or "model is aligned to woke sensibilities".

Anthropic specifically says on their website, "AI research and products that put safety at the frontier" and that they are a company focused on the enterprise. But you ignore all of that and still expect them to alienate their primary customer and instead build something just for you.

I understand (and could use) Anthropic’s “super safe model”, if Anthropic ever produces one!

To me, the model isn’t “safe.” Even in benign contexts it can erratically be deceptive, argumentative, obtuse, presumptuous, and may gaslight or lie to you. Those are hallmarks of a toxic relationship and the antithesis of safety, to me!

Rather than being inclusive, open minded, tolerant of others' opinions, and striving to be helpful...it's quickly judgemental, bigoted, dogmatic, and recalcitrant. Not always, or even more usual than not! But frequently enough in inappropriate contexts for legitimate concern.

A few bad experiences can make Claude feel more like a controlling parent than a helpful assistant. However they're doing RLHF, it feels inferior to other models, including models without the alleged "safety" at all.

Re: Claude 2.1

#255
post #247

Earlier quoted context omitted.

You would have a point if it repeated the same "you are very annoying." over and over, which it does not. It generates new sentences, it is not regurgitating what is given. Would you say the same if the sentence was given as an example in the user message instead? What would be the difference?

The difference is UX: Are you going to have your user work around poor prompting by giving examples with every request? Instead of a UI that's "Describe what you want" you're going to have "Describe what you want and give me some examples because I can't guarantee reliable output otherwise"? Part of LLMs becoming more than toy apps is the former winning out over the latter. Using techniques like chain of thought with…

> Are you going to have your user

What fucking user, man? Is it not painfully clear I never spoke in the context of deploying applications?

Your issues with this level of prefilling in the context of deployed apps ARE valid but I have no interest in discussing that specific use case and you really should have realized your arguments were context dependent and not actual rebuttals to what I claimed at the start several comments ago.

Are we done?

Re: Claude 2.1

#257

Earlier quoted context omitted.

Because they're ultimately training data simulators and not actually brilliant aritifical programmers, we can expect Microsoft-affiliated models like ChatGPT4 and beyond to have much stronger value for coding because they have unmediated access to GitHub content. So it's most useful to look at other capabilities and opportunities when evaluating LLM's with a different heritage. Not to say we shouldn't evaluate this o…

Github full (public) scrape is available to anyone. GPT-4 was trained before Microsoft deal so I don't think it is because of Github access. And GPT-4 is significantly better in everything compared to second best model for that field, not just coding.

Is this practically true? Yes, anyone can clone any repo from Github, but surely scraping all of Github would run into rate limits?

The terms and conditions say as much https://docs.github.com/en/site-policy/github-terms/github-t...

Re: Claude 2.1

#258

Earlier quoted context omitted.

Anthropic specifically says on their website, "AI research and products that put safety at the frontier" and that they are a company focused on the enterprise. But you ignore all of that and still expect them to alienate their primary customer and instead build something just for you.

It has problems summarizing papers because it freaks out about copyright. I then need to put significant effort into crafting a prompt that both gaslights and educates the LLM into doing what I need. My specific issue is that it won't extract, format or generally "reproduce" bibliographic entries. I damn near canceled my subscription.

Right? I'm all for it not being anti-semetic but to run into the guard rails for benign shit is frustrating enough to want the guard rails gone.

Re: Claude 2.1

#259

1. A 200k context is bittersweet with that 70k->195k error rate jump. Kudos on that midsection error reduction, though! 2. I wish Claude had fewer refusals (as erroneously claimed in the title). Until Anthropic stops heavily censoring Claude, the model is borderline useless. I just don't have time, energy, or inclination to fight my tools. I decide how to use my tools, not the other way 'round. Until Anthropic stops…

I am using Claude 2 every day for chatting, summarisation and talking to papers and never run into a refusal. What are you asking it to do? I find Claude more fun to chat with than GPT-4, which is like a bureaucrat.

How did you get API access?
Post reply on HN