IQ tests results for AI
trackingai.org
IQ tests results for AI
1–10 of 359 posts
Re: IQ tests results for AI
#2Re: IQ tests results for AI
#3Re: IQ tests results for AI
#4Babe wake up. New benchmark to overfit models just dropped.
Can’t wait for CEOs to start saying “why would we hire a 120 IQ person who works 9-5 with a lunch break when we can hire a 170 IQ worker who works 24x7 for half the cost??”
Re: IQ tests results for AI
#5Re: IQ tests results for AI
#6They then hypothesized a general factor, “g,” to explain this pattern. Early tests (e.g., Binet–Simon; later Stanford–Binet and Wechsler) sampled a wide range of tasks, and researchers used correlations and factor analysis to extract the common component, then norm it around 100 with a SD of 15 and call it IQ.
IQ tend to meaningfully predicts performance across some domains especially education and work, and shows high test–retest stability from late adolescence through adulthood. It is also tend to be consistent between high quality tests, despite a wide variety of testing methods.
It looks like this site just uses human rated public IQ tests. But it would have been more interesting if an IQ test was developed specifically for AI. I.e. a test that would aim to Factor out the strength of a model general cognitive ability across a wide variety of tasks. It is probably doable by doing principal component analysis on a large set of benchmarks available today.
Re: IQ tests results for AI
#7Really need to use a CDN before you get #1 on HN
Re: IQ tests results for AI
#8Re: IQ tests results for AI
#9Babe wake up. New benchmark to overfit models just dropped.
They’re definitely going to overfit on this, but this will be much better from a marketing perspective. Normies don’t know wtf an MMLU is, but they do know what IQ is and that 140 is a big number. Can’t wait for CEOs to start saying “why would we hire a 120 IQ person who works 9-5 with a lunch break when we can hire a 170 IQ worker who works 24x7 for half the cost??”
Re: IQ tests results for AI
#10Babe wake up. New benchmark to overfit models just dropped.