Live data from Hacker News

Everything we announced at our first LlamaCon

ai.meta.com

91–100 of 122 posts

Re: Everything we announced at our first LlamaCon

#91

Meta needs to stop open-washing their product. It simply is not open-source. The license for their precompiled binary blob (ie model) should not be considered open-source, and the source code (ie training process / data) isn’t available.

> their precompiled binary blob (ie model)

I agree with you that their license is not open source, but model weights are not binary blobs! Please stop spreading this misconception.

Re: Everything we announced at our first LlamaCon

#92

No new model? Maybe after the Qwen 3 release today they decided to hold back on Llama 4 Thinking until it benchmarks more competitively.

Beyond solid benchmarks, Alibaba's power move was dropping a bunch of models available to use and run locally today. That's disruptive already and the slew of fine tunes to come will be good for all users and builders. https://huggingface.co/collections/Qwen/qwen3-67dd247413f0e2...

> Beyond solid benchmarks, Alibaba's power move was dropping a bunch of models available to use and run locally today.

I agree, the advantage of qwen3's family is a plethora of sizes and architectures to chose from. Another one is ease of fine-tuning for downstream tasks.

On the other hand, I'd say it's "in spite" of their benchmarks, because there's obviously something wrong with either the published results, or the way they measure them, or something. Early impressions do not support those benchmarks at all. At one point they even had a 4b model be better than their prev gen 72b model, which was pretty solid on its own. Take benchmarks with a huge boulder of salt.

Something is messing with recent benchmarks, and I don't know exactly what but I have a feeling that distilling + RL + something in their pipelines is making benchmark data creep into the models, either by reward hacking, or other signals getting leaked (i.e. prev gen models optimised for one benchmark are "distilling" those signals into newer smaller models. No, a 4b model is absoulutely not gonna be better than 4o/sonnet3.7, whatever the benchmarks say).

Re: Everything we announced at our first LlamaCon

#93
post #64
post #59

Earlier quoted context omitted.

I'm working in Quest 3 almost every day. I use Immersed, as it implements virtual displays for my MacBook better than others, but I'm impressed with the Meta ecosystem. Granted, social interaction is still awkward without proper face expressions, but it feels closer each year to the depicted vision. I recently travelled and needed to work (coding and video editing in DaVinci) a lot in hotels and random places. I can'…

How do you power it for extended sessions? Is it constantly connected to power - using wifi/airlink between it and your Mac? Or do you use the link cable?

Yes, it's tethered to my laptop with a USB cable.

Re: Everything we announced at our first LlamaCon

#94
post #65
post #59

Earlier quoted context omitted.

I'm working in Quest 3 almost every day. I use Immersed, as it implements virtual displays for my MacBook better than others, but I'm impressed with the Meta ecosystem. Granted, social interaction is still awkward without proper face expressions, but it feels closer each year to the depicted vision. I recently travelled and needed to work (coding and video editing in DaVinci) a lot in hotels and random places. I can'…

In no way is Quest 3 better than Apple Vision Pro for desktop work. I’ve used both. It’s not close. (There are other critiques of AVP that might not make it the right choice, but desktop work experience isn’t one of them.)

I haven't tried AVP because of the price. But surely would love to.

Re: Everything we announced at our first LlamaCon

#95

Earlier quoted context omitted.

Is it? I am pretty active in those spaces and most folks it seems are using any number of the Chinese models like qwen, qwq, etc...

If by those spaces you mean reddit, then yeah I've also noticed this trend. It has became more egregious with the duality of L4 vs Qwen3 reception. L4 was blamed, mocked and everyone was posting shit about it (well, some of it was relevant since the launch was rushed and many providers had bad implementations for ~2-3 days) in stark contrast to qwen3 which also had inferencing problems (related to 3rd party tools usi…

L4 launched without a thinking model, making it an inferior choice for coding, one of the main LLM use cases. Even in benchmarks it wasn't even competitive at coding with 3-month-old Deepseek R1.

Re: Everything we announced at our first LlamaCon

#96

Earlier quoted context omitted.

If by those spaces you mean reddit, then yeah I've also noticed this trend. It has became more egregious with the duality of L4 vs Qwen3 reception. L4 was blamed, mocked and everyone was posting shit about it (well, some of it was relevant since the launch was rushed and many providers had bad implementations for ~2-3 days) in stark contrast to qwen3 which also had inferencing problems (related to 3rd party tools usi…

L4 launched without a thinking model, making it an inferior choice for coding, one of the main LLM use cases. Even in benchmarks it wasn't even competitive at coding with 3-month-old Deepseek R1.

Agreed, coding is not a strong point of L4. However the "hive mind" in some places thinks L4 is "a failure" and "useless". In reality it is an "ok" model, and most 3rd party benchmarks done after the inference lib updates were in line with whatever meta announced.

Re: Everything we announced at our first LlamaCon

#97
post #59
post #48

Its impressive that Llama and the Ai teams in general survived the meta-verse push at Facebook. Congrats to the team for keeping their heads down and saving the company from itself. Its all Ai all the time now though, not seen any mention of our reimagined future of floating heads hanging out together in quite some time.

I'm working in Quest 3 almost every day. I use Immersed, as it implements virtual displays for my MacBook better than others, but I'm impressed with the Meta ecosystem. Granted, social interaction is still awkward without proper face expressions, but it feels closer each year to the depicted vision. I recently travelled and needed to work (coding and video editing in DaVinci) a lot in hotels and random places. I can'…

[deleted]

Re: Everything we announced at our first LlamaCon

#98
post #66

Earlier quoted context omitted.

Would you happen to know of any resources for how to distill a ModernBERT model out of a larger one? I'm interested in doing exactly what you did, but I don't know how to start.

I was trying to identify "evergreen" and "time-sensitive" kinds of writing -- basically, I wanted to figure out if web pages captured in 2016 would still have content that's interesting to read today or if the passage of time would have rendered them irrelevant. Here's the training code that I used to fine-tune ModernBERT from the ~5000 pages I had labeled with Llama 3.3. It should be a good starting point if you hav…

Thank you!

Re: Everything we announced at our first LlamaCon

#99
post #59
post #48

Its impressive that Llama and the Ai teams in general survived the meta-verse push at Facebook. Congrats to the team for keeping their heads down and saving the company from itself. Its all Ai all the time now though, not seen any mention of our reimagined future of floating heads hanging out together in quite some time.

I'm working in Quest 3 almost every day. I use Immersed, as it implements virtual displays for my MacBook better than others, but I'm impressed with the Meta ecosystem. Granted, social interaction is still awkward without proper face expressions, but it feels closer each year to the depicted vision. I recently travelled and needed to work (coding and video editing in DaVinci) a lot in hotels and random places. I can'…

How do you work laying down?

Do you put a keyboard on your stomach or something?

Re: Everything we announced at our first LlamaCon

#100
post #59

Earlier quoted context omitted.

I'm working in Quest 3 almost every day. I use Immersed, as it implements virtual displays for my MacBook better than others, but I'm impressed with the Meta ecosystem. Granted, social interaction is still awkward without proper face expressions, but it feels closer each year to the depicted vision. I recently travelled and needed to work (coding and video editing in DaVinci) a lot in hotels and random places. I can'…

How do you work laying down? Do you put a keyboard on your stomach or something?

Yes, on laps/stomach. In VR mode, almost all these work-oriented apps have a concept of "desk view" or "portal" - i.e. window into "reality" to show keyboard. Meta's Horizon OS detects and shows physical keyboards automatically I believe. But a lot of time, I just use Mixed Reality mode - so only my screens floating in the corners of the room are rendered on top of "reality". If you forget that you're in VR goggles, it's the same as just lying down and working on huge screens on your ceilings or walls.
Post reply on HN