At the end of his post, Simon mentions translation between human languages. While maybe not directly related to token limits, I just did a test in which both R1 and o3-mini got worse at translation in the latter half of a long text. I ran the test on Perplexity Pro, which hosts DeepSeek R1 in the U.S. and which has just added o3-mini as well. The text was a speech I translated a month ago from Japanese to English, pr…
This is a great anecdote and I hope others can learn from it. R1, o1, and o3-mini work best on problems that have a “correct” answer (as in code that passes unit tests, or math problems). If multiple professional translators are given the same document to translate, is there a single correct translation?
Notes on OpenAI o3-mini
41–50 of 81 posts
Re: Notes on OpenAI o3-mini
#42At the end of his post, Simon mentions translation between human languages. While maybe not directly related to token limits, I just did a test in which both R1 and o3-mini got worse at translation in the latter half of a long text. I ran the test on Perplexity Pro, which hosts DeepSeek R1 in the U.S. and which has just added o3-mini as well. The text was a speech I translated a month ago from Japanese to English, pr…
Re: Notes on OpenAI o3-mini
#43So finally ChatGPT catches up with Claude which has a 200,000 token input limit ever since.
Claude with its projects feature is my go to tool for working on projects that I have to work on for weeks and months. Now I see a possible alternative.
Re: Notes on OpenAI o3-mini
#44Earlier quoted context omitted.
Programmers have always made a living by automating ourselves out of business. Somehow, we're still doing pretty well.
[flagged]
Digging through peoples past posts to look for gotcha’s is a pretty crass practice. It’s the kind of thing people used to do to try and dunk on each other on Reddit. It should not become the norm here.
Re: Notes on OpenAI o3-mini
#45At the end of his post, Simon mentions translation between human languages. While maybe not directly related to token limits, I just did a test in which both R1 and o3-mini got worse at translation in the latter half of a long text. I ran the test on Perplexity Pro, which hosts DeepSeek R1 in the U.S. and which has just added o3-mini as well. The text was a speech I translated a month ago from Japanese to English, pr…
Re: Notes on OpenAI o3-mini
#46At the end of his post, Simon mentions translation between human languages. While maybe not directly related to token limits, I just did a test in which both R1 and o3-mini got worse at translation in the latter half of a long text. I ran the test on Perplexity Pro, which hosts DeepSeek R1 in the U.S. and which has just added o3-mini as well. The text was a speech I translated a month ago from Japanese to English, pr…
Re: Notes on OpenAI o3-mini
#47At the end of his post, Simon mentions translation between human languages. While maybe not directly related to token limits, I just did a test in which both R1 and o3-mini got worse at translation in the latter half of a long text. I ran the test on Perplexity Pro, which hosts DeepSeek R1 in the U.S. and which has just added o3-mini as well. The text was a speech I translated a month ago from Japanese to English, pr…
How far off was o3 from the level of a professional translator (before it started to go off track)?
For more than a year, regular LLMs, when properly prompted, have been able to produce translations that would be indistinguishable from those of some professional translators for some types of translation.
General-purpose LLMs are best for translating straight expository prose without much technical or organization-specific vocabulary. Results are mixed for texts containing slang, dialogue, poetry, archaic language, etc.—partly because people’s tastes differ for how such texts should be translated.
Because most translators are freelancers, it’s hard to get a handle on what impact LLMs have been having on their workloads overall. I have heard reports from experienced translators who have seen work drop off precipitously and have had to change careers, while others report an increase in their workloads over the past two years.
Many translation jobs involve confidential material, and some translators may be hanging onto their jobs because their clients or employers do not allow the use of cloud-based LLMs. That safety net won’t be in place forever, though.
I suspect that those who work directly with translation clients and who are personally known and trusted by their clients will be able to keep working, using LLMs as appropriate to speed up and improve the quality of their work. That’s the position I am fortunate to be in now.
But translators who do piecework through translation agencies or online referrers like Fiverr will have a hard time competing with the much faster and cheaper—and often equally good—LLMs.
I made a few videos about LLMs and translation a couple of years ago. Parts of them are out of date, but my basic thinking hasn’t changed too much since then. If you’re interested:
“Translating with ChatGPT”
“Can GPT-4 translate literature?”
“What do translators think about GPT?”
https://www.youtube.com/watch?v=8JUepj7wIl0
I’m planning to make a few more videos on the topic soon, this time focusing on how I use LLMs in my own translation work.
Re: Notes on OpenAI o3-mini
#48At the end of his post, Simon mentions translation between human languages. While maybe not directly related to token limits, I just did a test in which both R1 and o3-mini got worse at translation in the latter half of a long text. I ran the test on Perplexity Pro, which hosts DeepSeek R1 in the U.S. and which has just added o3-mini as well. The text was a speech I translated a month ago from Japanese to English, pr…
This is a great anecdote and I hope others can learn from it. R1, o1, and o3-mini work best on problems that have a “correct” answer (as in code that passes unit tests, or math problems). If multiple professional translators are given the same document to translate, is there a single correct translation?
Re: Notes on OpenAI o3-mini
#49Earlier quoted context omitted.
[flagged]
I find this comment and most of your others very distasteful. Digging through peoples past posts to look for gotcha’s is a pretty crass practice. It’s the kind of thing people used to do to try and dunk on each other on Reddit. It should not become the norm here.
And I didn't go back that far, spent about 10 seconds looking.
As for my other comments, at least 80-90% of them have a positive vote ratio, even though I deliberately voice unpopular opinions. Most of the negatives have to do with politics, or Elon Musk, which people here tend to feel strongly about.
Considering that heated debates are not unwelcome here, I don't think that makes me a terrible person. That being said I'll try to be a little more careful before posting in the future, and I broadly agree with the point you were aiming for with your comment. Apologies to the GP if I caused any unwelcome feelings.
Re: Notes on OpenAI o3-mini
#50Earlier quoted context omitted.
Are you implying it isn't? (evidence please, everyone)
Simple example: o3-mini-high gets this [1] right, whereas Gemini 2.0 Flash 01-21 gets it wrong. [1] https://chatgpt.com/share/679d9579-5bb8-8008-ac4a-38cef65b45...