Pretty cool someone is still doing this. Training in house LLMs was extremely popular in 2023-2024, back when domain-specific LLMs could easily top GPT in their field. In my field alone (tax/HR tech) I remember that Intuit, Workday, Indeed, LinkedIn were all training internal models. It eventually stopped making sense because of inference costs. Running something internal with 30% GPU utilization is just too cost ine…
Thomson Reuters Launches Its Own Frontier Model
61–65 of 65 posts
Re: Thomson Reuters Launches Its Own Frontier Model
#62I don't trust that they'll be able to make back that $40M. This feels very much like a news agency getting into crypto or launching its own NFT line. Or IBM selling Watson. Or Mozilla chasing every which thing. They're not stakeholders in the future of work. They're just wanting to stay relevant and pattern matching against what they see. Reuters is too important for this. If they were trying to use this as a narrati…
I doubt that 40 million are a lot of money for TR. I also doubt that R&D spending having to "make back" directly is a winning strategy.
Re: Thomson Reuters Launches Its Own Frontier Model
#63Earlier quoted context omitted.
So it's a qwen fine-tune? I mean that's a reasonable thing to do, but then the press release shouldn't be written the way it is written. They're not as detached from the rest as the industry as the writing suggests. __ > It is obtained by repurposing the open-weight Qwen3.6-35B-A3B model and substantially improving it on a wide range of performance domains. nice wording on the HF page tho. "Repurposing". Lmao
the press release says this explicitly
Re: Thomson Reuters Launches Its Own Frontier Model
#64This is going to increasingly happen over the years to come. Big organizations will become more sophisticated with operationalizing their data, training and running LLMs will continue to be demystified and accessible, and over time we'll get more and more specialized / industry-specific models. It's going to become another way to monetize your informational assets if you're a big older enterprise with troves of data.…
40M$ to get a marginally better model is surprising, why not just use the free weight models
China is fucking smart. They have technical expertise to know that by feeding the model certain data, they can shape the outcome of questions posed.
Yeah, in the end, maybe both US and Chinese models can solve Fizz Buzz, or some Erdős problem.
And they can answer our inquisitive minds when we wonder, "Why is the sky blue?".
But deep questions about what is normal in the world they can change the outcome of.
And "Why is the sky blue?" is a deeper question than we think. Because it isn't blue everywhere in the world right now. And whether it is the fault of the US automotive industry, or aggressive investment by China in their own manufacturing base matters.
Of course, both have been responsible for pollution at different times. Growing up in Michigan near Dow Chemical I know this.
But depending on how they select training data answers can be nudged one way or another.
The US is smart too, and they have people working on the same things. It's ironically easier for me to talk about how China's security services are likely to shape their models' perception of the world. Which is a damn shame, because we deserve unbiased answers so we can all help our country improve.
Or countries improve if you want to take a global shared world view.