Live data from Hacker News

LLaMA2 Chat 70B outperformed ChatGPT

tatsu-lab.github.io

111–120 of 135 posts

Re: LLaMA2 Chat 70B outperformed ChatGPT

#111

Earlier quoted context omitted.

It's funny how ChatGPT really does give you the most balanced, middle of the road answers. It feels like a distillation of all human knowledge and sentiments. I use it constantly to get advice on plans, architectures, thoughts, etc.. to get an idea of pretty much what the average person would think. It often points out things I've overlooked which I'll improve my design with and go back and forth with ChatGPT until w…

Would put slight caution around asking it anything more than "what should I read next about this?" For whatever reason*, it is particularly bad at discussing philosophy I find. When I was grading philosophy 101, I would have probably given it a passing grade against the overall curve, but that's about it. Philosophy is a discipline of careful, sometimes jargoney, and always very couched assertions that can be easily…

The reason you feel that way might be that you are familiar enough with philosophy.

After all, LLMs and ChatGPT in particular are are indistinguishable from productised Gell-Mann Amnesia.

Edit: rewrote to be more neutral, sorry

Re: LLaMA2 Chat 70B outperformed ChatGPT

#112
post #29

This was just posted a few hours ago and when I tried it they were neck and neck (for me, LLama 2 won by a 1 question, but it was close): https://llmboxing.com/ It looks like the eval is open sourced so you could easily build a version w/ your own questions for blind testing...

At least when I tried Llama 2's response was always the longer one so it was hard to remain unbiased.

Re: LLaMA2 Chat 70B outperformed ChatGPT

#113

Earlier quoted context omitted.

Would put slight caution around asking it anything more than "what should I read next about this?" For whatever reason*, it is particularly bad at discussing philosophy I find. When I was grading philosophy 101, I would have probably given it a passing grade against the overall curve, but that's about it. Philosophy is a discipline of careful, sometimes jargoney, and always very couched assertions that can be easily…

I asked it where moral relativism fit in with philosophy and it came back with this Philosophy Ethics/Moral Philosophy Meta-Ethics: The study of moral thought, language, and properties Moral Realism: Belief that there are objective moral facts Moral Anti-Realism: Denial of the existence of objective moral facts Moral Relativism: The belief that moral judgments are true or false only relative to some particular standp…

Haha I think it's fine. I think its kinda cheeky answering you so literally, giving it an actual place to fit into :).

I don't doubt it can do, like, Wikipedia type classification ok, but that's not like really getting to the substance of anything! And, either way, its not like there is one decided-upon hierarchy of concepts like this people consciously work within. This is a fine picture to some, but others might contest, perhaps, that Meta-Ethics is the "study of moral thought, language, and properties." What is "moral language" anyway? Why is it meta relative to Moral Philosophy writ-large? Or perhaps one might argue that we need to think of meta-ethics as a sibling rather than child. The whole discipline is a mess of different thoughts and possible rebuttals and grand intellectual overturnings that will not be captured here. Maybe just try pasting that back into the prompt and asking "what's wrong with this picture?".

But like I said, its fine in that its fairly comparable to Wikipedia for utility, (with IMO a worse interface, but I get why people like it more).

Re: LLaMA2 Chat 70B outperformed ChatGPT

#114
post #108

Earlier quoted context omitted.

That's helpful! I've done a lot of work in audio synthesis, which is notoriously difficult measure. The gold-standard is human ratings of audio quality, but it is tough to design good tests (easy to fatigue raters) and the iteration time waiting for results is quite long. Instead, there's now some projects which use neural networks trained on human ratings to predict audio quality, such as ViSQoL: https://github.com/…

just curious, are there any open models doing the opposite of audio synthesis? As in able to generate the stems for a song?

https://github.com/facebookresearch/demucs

Re: LLaMA2 Chat 70B outperformed ChatGPT

#116
post #64

Earlier quoted context omitted.

Hi! Could you please share a few words on what type of data you are cleaning using GPT? It is an intriguing idea and I would love to learn more to see if I could use a similar approach.

Youtube transcripts. It only works with single-person channels at the moment, as I haven't worked on disambiguating multiple speakers. They are very messy if they are just an auto transcription from Google. Practically unusable in most cases. So first step in the pipeline is cleaning up the transcripts for incorrectly transcribed words or sentences. Using the context of the rest of the transcript, it is able to fix t…

How do you represent knowledge in your knowledge graph? Do you use an existing open source ontology?

Re: LLaMA2 Chat 70B outperformed ChatGPT

#119

Earlier quoted context omitted.

Is perfect agreement possible? And what is the definition of agreement? Humans don't agree about much.. are we saying agreement means it matches the truth after intensive investigation by humans?

Ugh, when I was doing my PhD work we were studying creativity in an experiment, and we needed an assessment for how creative different solutions were, and trying to get inter-rater reliability on this quite simple thing was just agonizing. I wound up abandoning the experiment because getting enough reliability would have required screwing down the standards so tightly that it would have ruined the underlying point of…

Kind of hard to consistently evaluate creativity when someone might pull a James T. Kirk (https://en.wikipedia.org/wiki/Kobayashi_Maru), which is only creative the first time and just a cheat thereafter.

Re: LLaMA2 Chat 70B outperformed ChatGPT

#120
post #82

Earlier quoted context omitted.

Maybe "balanced" rather than "average" is a better way of putting it?

What’s the balance between, say, opinions that trans people should be exterminated vs not? What’s the balance between Ukraine sovereignty vs Russian occupation? Etc

It would involve serious consideration of the opinions you don't agree with, rather than just qualifying them in the most hyperbolic and dismissive way possible. Hopefully ai is better capable of this sort of reasoning than people are.
Post reply on HN