Live data from Hacker News

Liquid AI reveals 8B-A1B MoE trained on 38T

liquid.ai

81–90 of 101 posts

Re: Liquid AI reveals 8B-A1B MoE trained on 38T

#81
post #69

At some point we have to be running into some inherent mathematical limits of knowledge compression, right? No way the knowledge benchmarks on these 8B models will keep getting better without overfitting on these benchmarks

If you give the model access to specialized tools (e.g. web search for question answering) the knowledge doesn't have to be stored in the model weights, which leaves some room for improvement. You'd still be overfitting to benchmarks (since different tasks might require different tools) but not necessarily to specific benchmark questions, so within-domain generalization could be quite good. As an example for a simila…

good point I have the feeling larger models (20b+) rely too much about their stored knowledge and sometimes fail to use tools because they think they know the answer. smaller specialized tool calling models could be the smart route for the future

Re: Liquid AI reveals 8B-A1B MoE trained on 38T

#82
post #52

Earlier quoted context omitted.

I've never thought of it as a fun term before. We use "park" as "I will park the car" not park as in "amusement park"

Why isn’t the “pantry” called the “food store”?

Does your pantry have a cashier and let you buy stuff there?

Because a food store sounds like it does.

Re: Liquid AI reveals 8B-A1B MoE trained on 38T

#83
post #49

I just tested this on a bug fixing benchmark I'm working on. It did not perform as well as I expected. Qwen2.5-Coder-3B (2 years old) outperformed it by a wide range -> fixing ~50% of bugs whereas this model only fixed ~12%. Granted, it's not a coder specific model, but given its benchmark performance to Gemma models, and that it's two years newer, and that it's an MoE with 8B total params, I expected it to be more c…

I will test it when it's accessible via OpenRouter, but the previous LFM2 model (lfm-2-24b-a2b) didn't do well on my tests, it got only 1/20 questions/tasks right, way below Gemma 31B or Qwen 35b-a3b (those get like 10/20 right)

I tested it against Gemma 4 31B and it's expectedly not favorable for world knowledge.

But even against E4B it's shaky, which is surprising given how many tokens they trained on. I guess it was on a lot of synthetic data.

Re: Liquid AI reveals 8B-A1B MoE trained on 38T

#84

Earlier quoted context omitted.

yes, I have, I use both. 27B slower in tok/s due to density, obviously, 35B-A3B for speed on simpler tasks.

You should enable MTP now that its available. LLamaCPP has had some massive updates in the last week or so.

Yes, Qwen 3.6 MoE is hitting like 80-90tk/s on Strix halo. On R9700 I had like 170t/s. It was not possible to keep up. But MoE is circling very often. I switch then to dense model and have 20-30t/s but it is able to solve quite a lot of tasks.

Re: Liquid AI reveals 8B-A1B MoE trained on 38T

#85

Earlier quoted context omitted.

You should enable MTP now that its available. LLamaCPP has had some massive updates in the last week or so.

Yes, Qwen 3.6 MoE is hitting like 80-90tk/s on Strix halo. On R9700 I had like 170t/s. It was not possible to keep up. But MoE is circling very often. I switch then to dense model and have 20-30t/s but it is able to solve quite a lot of tasks.

I get 50-60t/s tg on my r9700 with the dense, unsloth MTP quant UD-Q5_K_XL, K@8/V@4 256k context.

Using Vulkan backend.

``` llama-server -fa on -t 7 -ngl 999 --mlock --fit off --kv-offload --no-webui --metrics --chat-template-kwargs {"preserve_thinking": true} -b 2048 -ub 1024 -m /mnt/models/unsloth/Qwen3.6-27B-MTP-GGUF/Qwen3.6-27B-UD-Q5_K_XL.gguf --mmproj /mnt/models/unsloth/Qwen3.6-27B-MTP-GGUF/mmproj-F16.gguf -c 262144 --kv-unified -ctk q8_0 -ctv q4_0 --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-ngl 99 --alias unsloth/Qwen3.6-27B-MTP-GGUF --temp 0.60 --top-k 20 --top-p 0.95 --min-p 0.00 --presence-penalty 0.00 --repeat-penalty 1.00 ```

Re: Liquid AI reveals 8B-A1B MoE trained on 38T

#87

I just tested this on a bug fixing benchmark I'm working on. It did not perform as well as I expected. Qwen2.5-Coder-3B (2 years old) outperformed it by a wide range -> fixing ~50% of bugs whereas this model only fixed ~12%. Granted, it's not a coder specific model, but given its benchmark performance to Gemma models, and that it's two years newer, and that it's an MoE with 8B total params, I expected it to be more c…

That's not all that surprising, IMO. From what I understand, LiquidAI is focusing pretty narrowly on building models that operate as the "agentic core" of a larger system.

If I were going to use this model, I'd be looking to use it more as is the primary chat interface of a larger system, and having it orchestrate & delegate tasks to other places via tool calls. It's not quite as exciting on the surface as a local "do it all" model, but it does enable some pretty neat use-cases, IMO.

I'm imagining a local agent that is super low latency, works entirely offline, and capable of queuing up complex tasks for larger/smarter cloud agents which execute them asynchronously.

Re: Liquid AI reveals 8B-A1B MoE trained on 38T

#88
Beware the license. They misleadingly state on the blog post "Open-weight — Download, fine-tune, and deploy without restrictions". But if you read their license https://huggingface.co/LiquidAI/LFM2.5-8B-A1B/blob/main/LICE...> it has significant restrictions for any org with other $10M in revenue.

Re: Liquid AI reveals 8B-A1B MoE trained on 38T

#89
post #57

Earlier quoted context omitted.

Oh, I'm interested - do you have any docs with human responses to that?

“Car Wash” test with 53 models https://news.ycombinator.com/item?id=47128138 This article has a graph of the human response rates. About 70% correct on average. Accuracy depends on the country (maybe a language barrier?). See also original thread on the car wash thing. I want to wash my car. The car wash is 50 meters away. Should I walk or drive? https://news.ycombinator.com/item?id=47031580

"Correct" is pushing it, the question is too vague if approached as a genuine question and not a gotcha. I've actually had literal experiences where I wanted to wash my car and walked to a car wash in the past. That was me collecting the car, and there is an argument that would be a valid walk answer.

If we require logical rigour there isn't enough context in the question. If we allow for informal language then there are absolutely situations where cars get washed and people walk 50 meters to the car wash. It is a reasonable guess that the car is already at the wash and you have a 2nd car, given the question is being asked. It's a slight leap, but it is an inference that makes the question meaningful and so it is one that could be made.

I'd assume the LLMs are just failing at spatial reasoning, because AFAIK they're terrible at it. But both answers are justifiable because we don't know where the car is and have to make assumptions.

Re: Liquid AI reveals 8B-A1B MoE trained on 38T

#90

I just tested this on a bug fixing benchmark I'm working on. It did not perform as well as I expected. Qwen2.5-Coder-3B (2 years old) outperformed it by a wide range -> fixing ~50% of bugs whereas this model only fixed ~12%. Granted, it's not a coder specific model, but given its benchmark performance to Gemma models, and that it's two years newer, and that it's an MoE with 8B total params, I expected it to be more c…

It's not intended to be a coding model, however.
Post reply on HN