Live data from Hacker News

Mistral AI Releases Forge

mistral.ai

31–40 of 210 posts

Re: Mistral AI Releases Forge

#31

I am rooting for Mistral with their different approach: not really competing on the largest and advanced models, instead doing custom engineering for customers and generally serving the needs of EU customers.

their ocr model is goated

Better than Qwen? I guess the best overall is Gemini, right?

Re: Mistral AI Releases Forge

#33
post #12

Earlier quoted context omitted.

[flagged]

[flagged]

I don't think RAG is dead, and I don't think NFTs have any use and think that they are completely dead.

But the OP's blog is more about ZK than about NFTs, and crypto is the only place funding work on ZK. It's kind of a devil's bargain, but I've taken crypto money to work on privacy preserving tech before and would again.

Re: Mistral AI Releases Forge

#34

How many proprietary use cases truly need pre-training or even fine-tuning as opposed to RAG approach? And at what point does it make sense to pre-train/fine tune? Curious.

rag basically gives the llm a bunch of documents to search thru for the answer. What it doesn't do is make the algorithm any better. pre-training and fine-tunning improve the llm abaility to reason about your task.

Re: Mistral AI Releases Forge

#35

I am rooting for Mistral with their different approach: not really competing on the largest and advanced models, instead doing custom engineering for customers and generally serving the needs of EU customers.

first, there was .ai

next, it sounds like it's going to be .eu

but what about ai.eu

Re: Mistral AI Releases Forge

#38
post #28

> Pre-training allows organizations to build domain-aware models by learning from large internal datasets. > Post-training methods allow teams to refine model behavior for specific tasks and environments. How do you suppose this works? They say "pretraining" but I'm certain that the amount of clean data available in proper dataset format is not nearly enough to make a "foundation model". Do you suppose what they are…

Probably marketing speak for full fine-tuning vs PEFT/LoRA.

Re: Mistral AI Releases Forge

#39

How many proprietary use cases truly need pre-training or even fine-tuning as opposed to RAG approach? And at what point does it make sense to pre-train/fine tune? Curious.

You can fine tune small, very fast and cheap to run specialized models ie. to react to logs, tool use and domain knowledge, possibly removing network llm comms altogether etc.

Re: Mistral AI Releases Forge

#40
Don't sleep on Mistral. Highly underrated as a general service LLM. Cheaper, too. Their emphasis on bespoke modelling over generalized megaliths will pay off. There are all kinds of specialized datasets and restricted access stores that can benefit from their approach. Especially in highly regulated EU.

Not everyone is obsessed with code generation. There is a whole world out there.

Post reply on HN