Live data from Hacker News

AI Does NYT Connections

mikehearn.notion.site

11–20 of 22 posts

Re: AI Does NYT Connections

#12
Why is Gemini given 4 Xs for #547?

It has "Group 3" correct. It should be marked as having 1/4 groups correct.

Same thing happened on #535, Gemini actually got "Group 1" correct but was marked 0/4 correct.

Re: AI Does NYT Connections

#14

Beyond being able to solve Connections, can a LLM generate (good/challenging/solvable) connections? Would be pretty cool to be able to generate a test set.

Took a bit of prompting and a few awful ones, but ultimately it's not bad.

https://chatgpt.com/share/67570ab1-b2c0-8006-b5d2-d3fa7132de...

Going to try to feed this into some other LLMs like qwq and see if they can solve them.

Re: AI Does NYT Connections

#16

Beyond being able to solve Connections, can a LLM generate (good/challenging/solvable) connections? Would be pretty cool to be able to generate a test set.

I’d be surprised if they’re not using AI or some sort of rule-based generator at this point.

But creating one must be even funnier than trying to solve it, right?

Re: AI Does NYT Connections

#18
> Correct group with the wrong connection

This seems highly subjective. We should not care about this. The game is to connect the words, not find the connection. For human players, it doesn't matter if you get the connection or not.

Re: AI Does NYT Connections

#19

Beyond being able to solve Connections, can a LLM generate (good/challenging/solvable) connections? Would be pretty cool to be able to generate a test set.

Took a bit of prompting and a few awful ones, but ultimately it's not bad. https://chatgpt.com/share/67570ab1-b2c0-8006-b5d2-d3fa7132de... Going to try to feed this into some other LLMs like qwq and see if they can solve them.

QwQ gets it wrong, Gemini gets it wrong. o1 gets it right, R1 gets a pretty good not originally intended set of 4... tempted to give it partial credit. 4o gets it wrong. Will update with claude once my usage limits are up lol.

Re: AI Does NYT Connections

#20
This is very cool. It seems like the prompt is asking the LLM to one shot an answer. Have you tried asking it to make a group, confirm whether it's correct, and repeat with the remaining words? (like a human would)
Post reply on HN