Earlier quoted context omitted.
If we want to test these beasts in logic, we should probably start using actual formalized logic, rather than English. In just one test, Gemini flopped hard, while GPT-4-Turbo nailed it. Here is my prompt: Below is a well-typed CoC function: foo : ∀(P: Nat -> *) ∀(s: ∀{n} -> ∀(x: (P n)) -> (P (n + 1))) ∀(z: (P 0)) (P 3) = λP λs λz (s (s (s z))) Below is an incomplete CoC function: foo : ∀(P: Nat -> *) ∀(f: ∀{n} -> ∀(…
> I think we have a winner... It makes me sad that the complete and total lack of an objective way to measure these products means that the coming decades will be filled with this kind of hyper-specific gotcha test made in inappropriately confident internet posts. Literally this could have been down to one extra book in someone's training corpus, or a tokenizer that failed to understand λ as a non-letter. But no matt…
An AI system that produces right answers 90% of the time but 10% of the time drives your car into a lane divider, or says "there are 4 US states that start with 'K'" or "Napoleon was defeated at the Battle of Gettysburg" is worse than useless: It's dangerous.
As long as we call it a bullshit parlor trick, no problem. But unfortunately people are making important decisions based on these things.