Pretty cool someone is still doing this. Training in house LLMs was extremely popular in 2023-2024, back when domain-specific LLMs could easily top GPT in their field. In my field alone (tax/HR tech) I remember that Intuit, Workday, Indeed, LinkedIn were all training internal models. It eventually stopped making sense because of inference costs. Running something internal with 30% GPU utilization is just too cost ine…
At the Thomson Reuters family of companies (technically then: Refinitiv Ltd. sold to LSEG), the first foundational model (in the sense of "trained entirely from scratch") was trained already in 2018 (i.e., pre-ChatGPT); it would even have been earlier, but the electricity wires and fuses in the rented 5 Canada Sq, Canary Wharf office had to be replaced first at the time to deal with the current needed to serve the GPUs.