Live data from Hacker News

Mistral AI Releases Forge

mistral.ai

121–130 of 210 posts

Re: Mistral AI Releases Forge

#121
post #106

Earlier quoted context omitted.

No they haven't. Proof: Most big EU companies use Claude or Gemini or OpenAI, not Mistral. That choice was made recently. Things have changed in the loud echo chambers of the internet, maybe (but not really, since people were saying that EU data sovereignty was happening any time now since 2016).

My _feeling_ is that a lot of EU/European politicians has talked a lot more about the need to be independent from the US after Trump threaten Greenland. At least in the nordic countries. Not only concerning data & privacy, but defence, communications, space etc. All areas. The wheel has started to turn. You will not see it if you look around. But in 10 years time, maybe more, Europe will have stopped depending on the…

The politicians can talk, but they needed to set up an environment that would've let a European company have a decent shot at competing with the best AI models. But they didn't. Should've thought of that before being proud of setting up those strict tech regulations.

Re: Mistral AI Releases Forge

#122
post #40

Don't sleep on Mistral. Highly underrated as a general service LLM. Cheaper, too. Their emphasis on bespoke modelling over generalized megaliths will pay off. There are all kinds of specialized datasets and restricted access stores that can benefit from their approach. Especially in highly regulated EU. Not everyone is obsessed with code generation. There is a whole world out there.

I also think that this is the best approach for businesses wanting to adopt AI to automate, streamline, etc their business. The problem they have is that this is not a moat - their approach is easily reproducible. If they can pull ahead in having the most number of pre-trained models (one for this ERP, one for that CRM, etc) and then being able to close sales to companies using these products and sell them on post-tr…

Except the evidence today rather points to SOTA model + harness than fine tuned models.

Re: Mistral AI Releases Forge

#123

Earlier quoted context omitted.

I also think that this is the best approach for businesses wanting to adopt AI to automate, streamline, etc their business. The problem they have is that this is not a moat - their approach is easily reproducible. If they can pull ahead in having the most number of pre-trained models (one for this ERP, one for that CRM, etc) and then being able to close sales to companies using these products and sell them on post-tr…

Except the evidence today rather points to SOTA model + harness than fine tuned models.

> Except the evidence today rather points to SOTA model + harness than fine tuned models.

I have not seen that, actually. I still see most companies who want to jump into AI for the business sort of try RAG, but more often they just buy Chat accounts for their users.

The only place that harnesses appear to be used is in software development, but most companies aren't doing that either.

Re: Mistral AI Releases Forge

#125

Earlier quoted context omitted.

They use those because the decision to use them was made years ago. Things have changed since then

No they haven't. Proof: Most big EU companies use Claude or Gemini or OpenAI, not Mistral. That choice was made recently. Things have changed in the loud echo chambers of the internet, maybe (but not really, since people were saying that EU data sovereignty was happening any time now since 2016).

I consult for various companies and have definitely seen a trend. It's not quite the rupture that some expect but clearly not nothing either. Until very recently, the risk assessment of using US providers was considered very hypothetical. Today it still doesn't feel imminent, but it does feel very real.

Of course, it will be slow and painful and Europeans will need to use their own services for them to grow and mature.

Re: Mistral AI Releases Forge

#126
post #28

> Pre-training allows organizations to build domain-aware models by learning from large internal datasets. > Post-training methods allow teams to refine model behavior for specific tasks and environments. How do you suppose this works? They say "pretraining" but I'm certain that the amount of clean data available in proper dataset format is not nearly enough to make a "foundation model". Do you suppose what they are…

Pre-training mean exposing an already-trained model to more raw text like PDF extracts etc (aka continued pre-training). You wouldn't be starting from scratch, but it's still pre-training because the objective is just next token prediction of the text you expose it to.

Post-training means everything else: SFT, DPO, RL, etc. Anything that involves things like prompt/response pairs, reward models, or benefits from human feedback of any kind.

Re: Mistral AI Releases Forge

#127

Earlier quoted context omitted.

Is this the best Grok alternative?

Any model is.

This sounds like an ideology based reply. Grok is underrated and I think has a better chance of long term success than most. The current growth strategy means (for me) their chat harness is not up to par for serious work.

Their API is consistently among the most used on OpenRouter. While I can’t vouch for it myself, I think this is a decent proxy for capability. You can definitely see glimmers of greatness in their chat interface, it just feels like the system prompts are focused on something that doesn’t interest me.

Re: Mistral AI Releases Forge

#128
post #126
post #28

> Pre-training allows organizations to build domain-aware models by learning from large internal datasets. > Post-training methods allow teams to refine model behavior for specific tasks and environments. How do you suppose this works? They say "pretraining" but I'm certain that the amount of clean data available in proper dataset format is not nearly enough to make a "foundation model". Do you suppose what they are…

Pre-training mean exposing an already-trained model to more raw text like PDF extracts etc (aka continued pre-training). You wouldn't be starting from scratch, but it's still pre-training because the objective is just next token prediction of the text you expose it to. Post-training means everything else: SFT, DPO, RL, etc. Anything that involves things like prompt/response pairs, reward models, or benefits from huma…

Er, then what is the "already trained" model? I thought pre-training was the gradient descent through the internet part of building foundational models.

Re: Mistral AI Releases Forge

#129
> Forge enables enterprises to build models that internalize their domain knowledge. Organizations can train models on large volumes of internal documentation, codebases, structured data, and operational records. During training, the model learns the vocabulary, reasoning patterns, and constraints that define that environment.

I'm probably really out of date at this point, but my impression was that fine tuning never really worked that well for knowledge acquisition, and that don't variety of RAG is the way to go here. Fine tuning can affect the "voice", but not really the knowledge.

Re: Mistral AI Releases Forge

#130

I like Mistral, it hits the exact sweet spot between cost and my data staying in the EU, withouth a significant drop in quality, but man are their model naming conventions confusing af. They mention they have a model called Devstral 2, which is neither Codestral nor Devestral. I want to use it, but the api only lists devstral-2512, devstral-latest, devstral-medium-latest, devstral-medium-2507, devstral-small, devstra…

>data staying in the EU

This is really why Mistral has any support.

The models are bottom barrel, but its the best Europe has...

Although you could use Chinese models on European servers.

Post reply on HN